Position: AI Is Not Ready for Strategic Conflicts
Position: AI Is Not Ready for Strategic Conflicts
Title: Position: AI Is Not Ready for Strategic Conflicts 标题:立场:人工智能尚未准备好应对战略冲突
Abstract: Open-ended strategic wargames are high-stakes LM-based social simulations: they model adversaries, institutions, escalation, plan brittleness, doctrine, and crisis response. Language models (LMs) are attractive because they can play agents, generate scenario branches, adjudicate ambiguous actions, and summarize lessons, but the same affordances make open-ended roles dangerous: model language determines both what an actor attempts and what becomes simulated reality. 摘要: 开放式战略兵棋推演是基于语言模型(LM)的高风险社会模拟:它们对对手、机构、升级、计划脆弱性、准则和危机应对进行建模。语言模型之所以具有吸引力,是因为它们可以扮演代理人、生成场景分支、裁决模糊行为并总结经验教训,但同样的特性也使得开放式角色变得危险:模型语言既决定了参与者尝试做什么,也决定了什么会成为模拟现实。
This position paper argues that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, and that the proper use of open-ended wargames today is to stress-test decision-influencing LM agents. We identify five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. 本立场文件认为,在没有可审计的安全案例的情况下,任何基于语言模型的兵棋推演都不应为规划、准则、政策或危机应对提供参考;目前开放式兵棋推演的正确用途是对影响决策的语言模型代理进行压力测试。我们确定了五种失效模式:决策洗白(decision laundering)、裁决不透明、角色崩溃、通过裁决导致的升级,以及战略想象力的缺失。
Ordinary benchmarks cannot establish safety for these settings. Wargames can expose failures as stress tests; they are not themselves safety cases for consequential use. 普通的基准测试无法为这些环境建立安全性。兵棋推演可以作为压力测试来暴露故障;它们本身并不是用于重要决策的安全案例。