What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
我们对大语言模型有何期待?大语言模型基准测试设计图谱
Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance.
摘要: 基准测试是评估和交流大语言模型(LLM)进展的核心。然而,单纯的模型排名并不能揭示评估需求本身是如何变化的。基准测试种类的不断扩大提供了另一个视角:研究人员期望 LLM 完成什么任务,以及他们将什么样的表现视为成功。
We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.
我们系统地梳理了 2022 年 1 月至 2026 年 8 月间 arXiv 提交的 14,767 篇引入或更新评估资源的论文。通过分阶段筛选和自动化全文编码,我们考察了目标系统与领域、评估材料与条件以及评分机制方面的变化。
The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts.
研究结果显示,业界对行动(action)、交互(interaction)和专业应用(professional applications)的重视程度日益提高,同时既有的设计元素与较新的设计元素经常并存。模型的参与程度发展也不均衡:基于 LLM 的评分在智能体(agent)和非智能体组中均有所增长,而模型生成的评估材料在近期研究中并未表现出类似的持续增长趋势。
These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?
这些发现阐明了公共研究如何将能力预期转化为具体的测试和成功标准。随着人工智能参与到构建测试、执行任务和评判回复的过程中,它们也提出了一个问题:不断扩大的评估体系究竟是提供了更独立的证据,还是冒着复制其参与模型偏见与盲点的风险?