Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

有限主权与控制税:当部署者不拥有模型时,如何为 AI 监管定价

Abstract: AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs.

摘要: AI 控制研究探讨了如何在模型可能出现对齐偏差的情况下安全地部署它们,但许多控制协议都假设部署者能够对模型及其周边流水线进行监测。对于通过 API 或托管端点使用前沿模型的受监管组织而言,这一假设往往难以成立,因为部署者可能控制业务流程,却无法控制模型权重、服务基础设施、内部追踪记录、更新过程或完整的交互日志。

This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through contracts, architecture, audit, vendor assurance, residual risk, or reduced system scope.

本文引入了“有限主权”(bounded sovereignty)的概念:即在 AI 技术栈的数据、模型、基础设施和交互层面上,仅拥有部分技术和合同访问权限。文章指出,这些访问条件决定了哪些控制协议能够在实践中执行。本文提出了一个四层访问类型学、一个协议与层级的需求矩阵,以及“主权折扣成本”(sovereignty discount cost)的概念:这是控制税的一部分,用于通过合同、架构、审计、供应商保证、剩余风险或缩小系统范围来弥补缺失的访问权限。

It also reports a synthetic access-ablation experiment over 1.35 million synthetic case simulations and interprets the findings through an anonymised national-payments-infrastructure scenario. The experiment is not real-world payment-system evidence; it is a construct-validity exercise. The results show that complete logs improve diagnosis, a pre-execution gateway enables intervention, trace access and model-version control strengthen post-incident explanation, and scope restriction can improve safety while reducing usefulness. Control protocols proposed as general safety solutions should therefore state their access assumptions explicitly.

本文还报告了一项针对 135 万次合成案例模拟的“访问消融实验”,并通过一个匿名的国家支付基础设施场景对研究结果进行了阐释。该实验并非现实世界支付系统的证据,而是一项构念效度验证练习。结果表明,完整的日志有助于诊断,执行前网关能够实现干预,追踪访问和模型版本控制增强了事后解释能力,而范围限制可以在提高安全性的同时降低实用性。因此,作为通用安全解决方案提出的控制协议,应明确说明其对访问权限的假设。