LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

大语言模型知晓约束却无法应用:语用约束推理中的激活瓶颈

Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail — but aggregate accuracy conflates genuine constraint inference with conservative defaulting.

摘要: 当显著的表面线索与隐含的可行性约束发生冲突时,大语言模型(LLMs)往往会失败——但总体的准确率指标混淆了真正的约束推理与保守的默认行为。

We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair).

我们将这种区别形式化为“条件约束激活”:约束在存在约束和不存在约束的提示词中被对称地内部编码(知识),但有时仅被路由到决策中(路由),并可通过供体激活进行修复(修复)。

A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above 88%, yet activation patching repairs one (+6.4 nats) and not the other (-0.07).

通过对 14 个模型进行的四重诊断揭示了两种失效模式;对两个开源权重模型的探测显示,其对约束的解码准确率超过 88%,然而激活修补(activation patching)仅修复了其中一个(+6.4 nats),而对另一个无效(-0.07)。

On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway — prerequisite mention.

在缓解策略方面,没有任何提示词干预手段能达到“修复”的效果:所有方法都通过单一的中介路径——即前提提及(prerequisite mention)——加剧了保守偏见。

Hidden-constraint failure is a routing problem, not a knowledge problem.

隐含约束的失效是一个路由问题,而非知识问题。


Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) 学科分类: 计算与语言 (cs.CL);人工智能 (cs.AI)

Cite as: arXiv:2608.12321 [cs.CL] 引用格式: arXiv:2608.12321 [cs.CL]