Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

automated peer reviewconference guidelinesllmreviewer guidelinesrubric scoringsubjective evaluation

Abstract


Peer review remains a cornerstone of scientific research, but the increasing volume of submissions has driven interest in its automation. This study investigates how the design of reviewer guidelines—specifically official conference guidelines versus reviewer-imitating guidelines generated from high-quality human reviews using large language models (LLMs)—affects the quality of automated peer review. Our experiments reveal that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines are generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degrades performance, underscoring the importance of allowing subjective and holistic scoring in automated review systems.


Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)


Comments: 18 pages, 2 figures, ACL 2026 Findings


Cite as: arXiv:2607.22553 [cs.CL] (or arXiv:2607.22553v1 [cs.CL] for this version)


Submission history: Submitted on 16 May 2026, by Masafumi Oyamada. [v1] Sat, 16 May 2026 02:12:55 UTC (405 KB)

via ArXiv CL

Related