Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
保持冷静(CALM):分析文本生成图像中全局不安全性(Global Unsafety)的局限性
Abstract: Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts.
摘要: 文本生成图像的免训练安全防护机制通常依赖于可复用的安全信号(例如不安全方向或全局有害子空间),并将其广泛应用于各类提示词(prompts)。我们对这种“全局不安全性”假设进行了受控的几何分析,并揭示了一个始终存在的“覆盖率-选择性”权衡问题:紧凑的不安全子空间无法覆盖异构的不安全语义,而更广泛的聚合则会日益扭曲与安全相关的良性提示词。
Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components.
受此发现启发,我们提出了 CALM(反事实自适应局部调制),这是一种免训练的安全防护机制,它用提示词局部的反事实修正取代了统一的全局移除。通过使用匹配的不安全-良性锚点,CALM 将每个提示词引导至活跃的不安全类别,仅对违规的标记(token)表示进行最小程度的修正以使其趋向安全侧,并抑制正向对齐的不安全残差分量。
Across broad evaluation, CALM significantly improves unsafe content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe signal removal.
通过广泛的评估,CALM 在显著提升不安全内容抑制效果的同时,保持了良性内容的实用性,证明了局部反事实修正为全局不安全信号移除提供了一种更具选择性的替代方案。