AI Evaluation Should Work With Humans

AI Evaluation Should Work With Humans

AI 评估应致力于人机协作

Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.

摘要: 本立场论文指出,当前主流的 AI 评估范式(侧重于超越人类的自主表现,从而隐含地以取代人类为目标)正在将 AI 的发展引向错误的方向。

Instead, the AI community should pivot to evaluating the performance of human—AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.

相反,AI 社区应转向评估“人机协作团队”的表现。我们认为,这种协作模式的转变将促进 AI 系统真正成为人类能力的补充,从而比当前的评估流程带来更好的社会效益。


Paper Details:

  • Title: AI Evaluation Should Work With Humans
  • Authors: Jan Kulveit, Gavin Leech, Tomáš Gavenčiak, Raymond Douglas
  • arXiv ID: 2608.13577
  • Date: 6 Jul 2026

论文详情:

  • 标题: AI Evaluation Should Work With Humans (AI 评估应致力于人机协作)
  • 作者: Jan Kulveit, Gavin Leech, Tomáš Gavenčiak, Raymond Douglas
  • arXiv ID: 2608.13577
  • 日期: 2026 年 7 月 6 日