Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit

立场:AI 对齐社区正在无意中构建一套审查工具包

Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. 摘要: 本立场论文指出,现代人工智能对齐方法(最初旨在防止有害输出)属于两用技术,极易被恶意行为者滥用于审查和操纵。

By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a “perfectly aligned” model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. 通过将当前的对齐技术与滥用的可能性及实际案例进行映射,我们展示了对“完美对齐”模型的追求,在无意中也为恶意行为者提供了一种不断改进的信息主导工具。

We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. 我们现在必须讨论这种两用潜力,因为随着用户迅速将人工智能作为信息提供者,加之经济权力不对称以及政治格局日益向威权主义倾斜,其风险正在加剧。

We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential. 我们最后呼吁社区关注对人工智能对齐机制的蓄意滥用,并提出缓解策略,以防范这种两用潜力带来的风险。


Paper Details:

  • Authors: Sarah Ball, Phil Hackemann
  • Published: Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), Seoul, South Korea.
  • arXiv ID: 2608.12346

论文详情:

  • 作者: Sarah Ball, Phil Hackemann
  • 发表: 第 43 届国际机器学习会议 (ICML 2026),韩国首尔。
  • arXiv ID: 2608.12346