Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
快速模型,缓慢证据:针对大模型智能体框架中“系统 1”决策模型的配对与自审计评估
Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls.
摘要: 大模型智能体(Agent)框架在执行每个任务时都会做出许多细小的、特定类型的决策:调用哪个模型、使用哪个工具、检索到的文本是否相关、输入内容是否包含注入攻击等。“系统 1”(System-1)决策模型通过单次前向传播即可输出类别概率来回答这些问题,相较于调用大语言模型(LLM),有望大幅降低成本并减少延迟。
We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks.
我们对一个开源权重模型(Laya)和一个托管模型(Jev)进行了配对评估,这些模型涵盖了基于 18 个公共来源构建的 11 个智能体决策点。评估数据集包含 7,283 个基础案例和 6,640 个鲁棒性变体,并采用了字节级相同的输入、配对测试以及跨硬件和跨日的复现性检查。
Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool).
在 11 个决策点中的 9 个上,Jev 的准确率显著更高(提升了 10.8 到 46.0 个百分点)。两个模型在零样本模型路由任务上的表现均未超过随机猜测,且在 RAG(检索增强生成)相关性门控任务上表现持平。当选项顺序颠倒时,Laya 会改变 30% 的答案,且在面对大量或相似候选对象时性能急剧下降(在 50 个最近邻工具的情况下准确率为 31%,而 Jev 在具有唯一正确工具的项目上准确率为 98%)。
We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a “channel effect” on injection false positives that vanishes with channel-native the content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at this URL.
我们还对自身的评估流程进行了审计。三项分析错误和一项设计混淆因素扭曲了最初的部署结论:遗漏了预筛选成本(报告节省 23.9%,实际为 4.3%)、将门控准确率误报为端到端质量(58% 对比 98%)、样本内阈值问题(目标 5%,但在留出集上漏报率高达 17%),以及在注入攻击误报中存在的“通道效应”(该效应在通道原生内容中消失)。另外两个被怀疑的混淆因素并未改变最终结论。所有案例、原始输出和分析代码均可在该链接获取。