Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

找到棋步不等于赢得比赛:用于大模型智能体闭环评估的 XiangqiBench

Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender.

摘要: 静态评估通常会因为语言模型给出了正确的棋步而给予其肯定,但智能体必须在对手做出回应的同时,将计划执行到底并获得验证结果。我们引入了 XiangqiBench,这是一个用于衡量中国象棋中这种差异的可执行基准测试:从 119 个由引擎或仅限将军搜索支持的强制将死残局开始,大模型智能体必须在对抗引擎防守者的情况下完成将死。

An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success.

一个交互式的 REPL 界面将真实棋步、状态查询和前向模拟分离开来,我们记录了 12 个前沿大模型在两种观察协议下的 8,568 条多轮轨迹。三个看似代表能力的信号实际上都高估了闭环成功率。

(i) The Conversion Gap: models play the stored reference first move in 26.1% of Sighted trials, yet only 13.9% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3% of accepted simulation calls stop on an illegal move, and in 49.3% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent.

(i) 转化差距:模型在 26.1% 的“有视野”(Sighted)试验中走出了存储的参考首步,但其中只有 13.9% 的试验最终获胜。(ii) 一致性差距:领先模型达到了 38.7% 的 pass@3,但 pass^3 仅为 5.9%;在其获胜的 46 个局面中,只有 7 个局面在三次试验中全部获胜。(iii) 模拟差距:32.3% 的被接受模拟调用在非法棋步上停止,且在 49.3% 的可比案例中,真实防守者的回应与智能体模拟的路径不同;自主生成的推演虽然能检查合法性,但无法预判对手的行动。

Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.

找到棋步并不等于赢得比赛:智能体评估应当对闭环结果进行评分,并在报告覆盖率的同时报告可靠性。