TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression.
摘要: 角色扮演评估不应仅仅给出一个单一的分数,而应揭示测试了哪些角色要求、哪些要求失败了,以及哪些对话证据支持了这一判断。我们提出了 TRACE Bench,这是一个任务驱动的智能体清单评估框架。它在离线状态下将每个角色档案分解为固定的检查清单,然后利用一个用户智能体(User Agent)与目标角色扮演模型进行自然对话,同时根据模型的回复私下更新检查清单的状态。因此,评分可以追溯到具体的检查项和支持性的对话轮次,而不是基于黑盒的整体印象。
Content: For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement.
内容: 为了进行覆盖率交叉验证,我们根据相同的角色衍生清单,对 MiniMax 角色扮演基准测试中发布的 M2 自由对话记录进行了审计。已发布的自由聊天记录仅覆盖了关键角色档案要点的 73.74%,而 TRACE Bench 在更少的对话轮次中达到了 99.91% 的覆盖率。鲁棒性实验表明,在重复运行和更换用户智能体的情况下,排名保持稳定。
Content: Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.
内容: 在 26 个模型中,TRACE Bench 报告了整体排名以及能力细分和检查清单追踪记录。它还支持闭环基准进化(Closed-Loop Benchmark Evolution),提炼在失败追踪中被证明有效的验证方法,从而使后续的评估能够更可靠地诱导和检查已观察到的失败模式。