Research2w ago
Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
A study evaluating multimodal LLMs as peer reviewers for ICLR 2026 found they gave inflated scores, failed to catch most inserted errors, and struggled with…
#LLM#Peer Review#Multimodal#Research Evaluation