Do Methods Support the Claims? Intra-Paper Verification for Peer Review
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
方法是否支持声明?同行评审中的论文内部验证
Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper’s claimed contributions against prior literature, implicitly assuming that these contributions are accurately realized in the work itself. Human reviewers, however, frequently challenge novelty claims not because similar ideas already exist, but because the methodological evidence presented in the paper does not adequately support them. This internal mismatch between claimed contributions and methodological realization is rarely examined by current LLM-based review systems.
摘要: 科学投稿量的日益增长激发了人们利用大语言模型(LLM)辅助同行评审的兴趣。现有的自动化创新性评估方法通常将论文声称的贡献与现有文献进行比较,并隐含地假设这些贡献在论文本身中得到了准确实现。然而,人类审稿人经常质疑创新性声明,原因并非因为存在类似的想法,而是因为论文中提出的方法论证据无法充分支持这些声明。这种声称的贡献与方法论实现之间的内部不匹配,在当前的基于 LLM 的评审系统中很少被审查。
To address this gap, we introduce intra-paper claim verification, a framework that evaluates whether novelty claims articulated in a paper are substantiated by the methods used to realize them. The framework employs an LLM to extract novelty claims from the introduction, retrieve claim-relevant methodological evidence, and assess whether the methods substantiate the stated contributions. Assessment is guided by reviewer-inspired evaluation criteria derived inductively from human peer reviews collected from 182 ICLR 2025 papers. These criteria capture recurring reviewer concerns related to novelty, methodology, clarity, and other issues and are used to generate structured reviewer-style assessments of claim substantiation.
为了解决这一空白,我们引入了“论文内部声明验证”(intra-paper claim verification),这是一个评估论文中阐述的创新性声明是否由实现它们的方法所支撑的框架。该框架利用 LLM 从引言中提取创新性声明,检索与声明相关的方法论证据,并评估这些方法是否支撑了所陈述的贡献。评估过程由受审稿人启发的评估标准指导,这些标准是从 182 篇 ICLR 2025 论文的人类同行评审中归纳得出的。这些标准捕捉了审稿人关于创新性、方法论、清晰度及其他问题的常见顾虑,并被用于生成结构化的、审稿人风格的声明支撑评估。
We evaluate the framework by comparing LLM-generated review comments against human reviewer concerns on a balanced subset of accepted and rejected papers. Human evaluation demonstrates significant alignment between framework-generated assessments and human reviewer concerns, particularly for novelty-related issues. BERTScore further distinguishes corresponding human-LLM review pairs from mismatched controls, indicating that the framework captures concerns consistent with human reviewer observations.
我们通过在已录用和被拒论文的平衡子集上,对比 LLM 生成的评审意见与人类审稿人的顾虑,对该框架进行了评估。人类评估表明,框架生成的评估与人类审稿人的顾虑之间存在显著的一致性,特别是在与创新性相关的问题上。BERTScore 进一步将对应的人类-LLM 评审对与不匹配的对照组区分开来,这表明该框架捕捉到的顾虑与人类审稿人的观察结果是一致的。