TomasuLLM: Out-of-Order Speculative Execution for LLM Agents

TomasuLLM: Out-of-Order Speculative Execution for LLM Agents

TomasuLLM:面向 LLM Agent 的乱序推测执行技术

Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors — a sequential interface hides work that can be predicted and started early, but a speculative result may become visible only after it and every earlier step have been validated.

长时间运行的工具往往会主导编程 Agent 的延迟:编译器、测试套件和仓库命令可能需要数秒到数分钟的运行时间,而此时 Agent 只能处于空闲状态。这种观察到的停顿与推动乱序处理器发展的核心矛盾如出一辙——顺序接口掩盖了那些可以被预测并提前启动的工作,但推测性的结果只有在它本身及所有前序步骤都经过验证后,才能被正式采纳。

We present TomasuLLM, a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness. It drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state.

我们提出了 TomasuLLM,这是一个能够在保持任务执行正确性的前提下,实现 Agent 工具调用乱序执行的运行时系统。它会预判未来的动作,在隔离的“写时复制”(copy-on-write)沙箱中运行这些动作,追踪其依赖关系和影响,并仅在通过已提交状态的验证后,才按轨迹顺序提交结果。

Across three benchmarks spanning sub-second to minutes-long tool calls, TomasuLLM improves the reported benchmark means and scales with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. Across 4,010 audited commit-validation records, it produces zero false accepts.

在涵盖亚秒级到分钟级工具调用的三个基准测试中,TomasuLLM 提升了基准测试的平均表现,并随工具延迟的增加而展现出良好的扩展性:在 100 个 SWE-bench Verified 任务上提升了 1.31 倍,在 28 个 Terminal-Bench 2.0 任务上提升了 1.35 倍,在 18 个 SWE-Marathon 会话中实现了 1.27 倍的进度匹配。在 4,010 条经过审计的提交验证记录中,该系统实现了零误判(false accepts)。