Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
评估审稿人指南设计对基于大语言模型的自动化同行评审的影响
Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review.
摘要: 同行评审是科学研究中不可或缺的环节,但日益增长的工作量使得自动化评审变得愈发必要。在本研究中,我们分析了不同类型的审稿人指南(例如官方会议指南,以及利用大语言模型从高质量人类评审中生成的模仿型指南)如何影响自动化同行评审的效果。
Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well.
我们的实验表明,官方会议指南所产生的评审结果与人类判断最为一致,这表明通过会议实践所提炼出的评估标准,同样可以作为自动化评审的有效指导。
In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.
相比之下,模仿型指南的效果普遍不如官方会议指南。此外,强制执行严格的评分细则(rubric-style scoring)会持续降低模型表现,这凸显了允许主观和整体性评分的重要性。
Paper Details:
- Authors: Haowen Li, Yoichi Ishibashi, Masafumi Oyamada
- arXiv ID: 2607.22553
- Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Submission Date: 16 May 2026
论文详情:
- 作者: Haowen Li, Yoichi Ishibashi, Masafumi Oyamada
- arXiv ID: 2607.22553
- 学科分类: 计算与语言 (cs.CL);人工智能 (cs.AI)
- 提交日期: 2026年5月16日