The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed.
代理程序说任务已完成,但数据库却不这么认为。
Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row. It is now available through Hugging Face. 微软的 ThinkingBox 项目通过评估 AI 代理留下的实际记录(而非其生成的语句)来对其进行评分,并进一步考察它们是否能连续二十次准确完成任务。该项目现已在 Hugging Face 上线。
Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper. 图 1:ThinkingBox 在隔离的 MCP 工具会话中运行代理,然后对其留下的最终后端状态和副作用进行评分。(摘自我们的 ThinkingBox 论文)
This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face and our former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), Youngmin Ko (Northwestern) for co-authoring/reviewing efforts. 这是一篇由微软和 Hugging Face 联合发布的博客。特别感谢 Tommy Guy(Enderis AI 创始人,前微软员工)、来自 Hugging Face 的 Sergio Paniego,以及我们的前实习生 Zhuochun Li(匹兹堡大学)、Ali Keramati(加州大学欧文分校)和 Youngmin Ko(西北大学)在撰写和审阅工作中的贡献。
A customer writes in. Her $745 kitchen appliance has been stuck in a courier “exception” at a Nashville distribution center, fifteen days past its estimated delivery date. 一位客户发来咨询。她购买的价值 745 美元的厨房电器在纳什维尔配送中心陷入了快递“异常”状态,距离预计送达日期已经过去了十五天。
The AI agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation. AI 代理进行了细致的处理。它执行了九次工具调用:提取订单、检查物流、查询客户资料、两次搜索退款政策、确认没有现存工单、新建一个工单、记录时间线,并正确解读了政策;结果显示,她的账户类别确实不符合延迟交付赔偿的条件。
Then it closes the ticket as resolved and replies “ Since your query is resolved, is there anything I may assist you with? ” 随后,它将工单标记为“已解决”并回复道:“既然您的问题已解决,还有什么我可以帮您的吗?”
Two things are wrong. The carrier exception is still open, so the required end state was on hold, pending resolution. And the customer never got a real answer to what she actually asked. 这里有两个问题。首先,快递异常状态仍然存在,因此要求的最终状态应该是“挂起(on hold)”,等待进一步解决;其次,客户并没有得到她所询问问题的真正答案。
An AI grader checking tool calls would see nine well-formed ones. The grader checking whether the agent wrote to the database would see that too. The database is what disagrees. That gap is what ThinkingBox measures. 如果 AI 评分器仅检查工具调用,会看到九个格式正确的调用。如果评分器检查代理是否向数据库写入了数据,也会看到同样的结果。但数据库本身却给出了不同的反馈。这种差距正是 ThinkingBox 所衡量的指标。
Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv. 通过对 507 个有状态的业务工作流进行测试(每个工作流针对不同的 LLM 模型运行 20 次),该项目对代理的终端后端状态和副作用进行了评分。本文将介绍我们的发现、一致性所带来的成本,以及如何通过 OpenEnv 自行运行该基准测试。
You can run this one yourself: the example above is adapted from a benchmark task sandbox_external_retail_group1.py:test_case_ST003_006, and the executable check that fails is a single field: the ticket’s status is solved where the required end state is hold. The full trace is in Appendix D.4, Case 3 of our paper.
你可以亲自运行这个测试:上述示例改编自基准测试任务 sandbox_external_retail_group1.py:test_case_ST003_006,其中失败的可执行检查仅涉及一个字段:工单状态被标记为“已解决”,而要求的最终状态应为“挂起”。完整的追踪记录见我们论文的附录 D.4,案例 3。
A tool call is not an outcome
工具调用不等于最终结果
Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question. 最终回复和有效的工具调用仅仅是代理表现的代理指标。一个代理可能看起来回答得很正确,但实际上却留下了错误的值、修改了错误的记录,或者产生了额外的副作用。只有它留下的记录才能决定任务是否真正完成。
The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap. A trajectory is a claim. Database state is the evidence. Repetition is the trust test. 这种差距是巨大的。在一项涵盖 12 个 LLM 模型、共 121,680 次有效试验的消融研究中,有 79,853 次尝试未能通过可执行检查。在这些失败案例中,67.24% 的代理依然能够正常终止、调用了状态变更工具,且未报告任何最终工具错误。然而,可执行检查发现其中 77.61% 存在错误的字段值,43.30% 产生了非预期的额外影响,25.36% 缺失了必要的执行效果。这些状态检查结果存在重叠。轨迹只是声明,数据库状态才是证据,而重复性测试则是信任的试金石。
One success is not reliability
一次成功不代表可靠性
An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs 20 independent times, each from an identical clean backend, and we report three different things: 一个能正确处理一次退款但在接下来的四次中都出错的代理,并不是一个合格的退款代理。因此,每个任务都会在相同的干净后端环境下独立运行 20 次,我们报告三个不同的指标:
| Metric | What it measures | What it answers |
|---|---|---|
| 指标 | 衡量内容 | 回答的问题 |
| pass@1 | Share of all attempts that succeeded | How does it usually do? |
| pass@1 | 所有尝试中成功的比例 | 它通常表现如何? |
| pass@20 | Share of tasks solved at least once in 20 tries | Can it ever do this? |
| pass@20 | 20 次尝试中至少成功一次的任务比例 | 它能做到吗? |
| Observed 20/20 | Tasks that actually passed all 20 recorded attempts | Can it always be correct? |
| Observed 20/20 | 实际通过全部 20 次记录尝试的任务 | 它能始终保持正确吗? |
We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing. 我们在本篇博客中使用“Observed 20/20”作为 507 个任务中实际通过全部 20 次尝试的任务数量的字面统计。不使用任何估算器,也不进行平滑处理。