From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
从提议到验证效果:Praxa,一种用于受控 AI 智能体执行的证据绑定框架
Abstract: Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion.
摘要: 大语言模型智能体能够提议并执行操作,但“提议”、“授权”、“调度”、“验证外部效果”以及“服务推广”是截然不同的概念。我们提出了 Praxa,这是一个智能体框架,通过确定性准入、代理执行、外部回读、对账和审查推广,明确地呈现了这些状态。
We report four evidence lanes. First, an author-run repository-local audit at a pinned revision passed 1,027/1,027 unit tests and 89/89 Workerd tests, instrumented all 363 expected source files, and met four coverage floors; raw per-test transcripts and independent reproduction are unavailable.
我们报告了四个证据维度。首先,作者在固定版本上进行的本地仓库审计通过了 1,027/1,027 项单元测试和 89/89 项 Workerd 测试,覆盖了所有 363 个预期的源文件,并达到了四个覆盖率基准;但原始的单项测试记录和独立复现结果目前尚不可用。
Second, in a provider-backed Terminal-Bench Core 0.1.1 pilot across 12 curated tasks, baseline and reliability-layer arms each passed 17/36 strict trials. The reliability layer used 37.49% more input and 50.73% more output tokens, so the pilot does not support superiority.
其次,在基于供应商的 Terminal-Bench Core 0.1.1 试点中,针对 12 项精选任务,基准组和可靠性层组各通过了 17/36 次严格试验。可靠性层多消耗了 37.49% 的输入 Token 和 50.73% 的输出 Token,因此该试点结果并不支持其具有优越性。
Third, in a post-debug, two-order coordination-proxy development comparison, baseline and a source-authored candidate each completed 180/180 trials with equal measured accuracy, full hermetic crash recovery, and zero protected violations. The candidate used 37.11% fewer tokens, 33.84% lower estimated endpoint cost, and 11.63% fewer steps; this does not establish improved quality, latency, or production behavior.
第三,在调试后的双阶协调代理开发对比中,基准组和作者开发的候选组均完成了 180/180 次试验,测量准确率相同,具备完全的封闭式崩溃恢复能力,且零受保护违规。候选组少消耗了 37.11% 的 Token,预估端点成本降低了 33.84%,步骤减少了 11.63%;但这并不能证明其在质量、延迟或生产行为方面有所提升。
Fourth, deployed source/configuration evidence shows bounded reflection, recall accounting, memory compilation, and tool-health paths, but no production outcome lift. Praxa’s supported contribution is an evidence-bound architecture that makes authority-to-effect transitions explicit and testable. Current evidence does not establish adversarial security, production safety, general specialist superiority, autonomous recursive optimization, or user benefit.
第四,部署的源代码/配置证据显示了有界反射、召回核算、内存编译和工具健康路径,但并未带来生产成果的提升。Praxa 的主要贡献在于提供了一种证据绑定架构,使“授权到效果”的转换变得明确且可测试。目前的证据尚不能证明其在对抗性安全、生产安全性、通用专家优越性、自主递归优化或用户收益方面的有效性。