Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

大语言模型中的稳定错误校准:高置信度错误的实践视角

Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations.

摘要: 大语言模型中的高置信度错误通常被视为内部推理脆弱的证据。我们研究了另一种可能性:稳定错误校准(stable miscalibration),即一个置信度很高的错误答案在受到微小扰动时仍保持局部稳定。

We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement.

我们结合了两种诊断方法:一种是标签感知的输出级审计评分,它根据强制回答基准下的置信度变化和过度自信的错误对领域进行排名;另一种是测量隐藏状态移动的内部敏感性探测。

On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly.

在一个多领域二元事实审计集上,该审计评分能够追踪到“感知弃权(abstention-aware)”的自我批判在何处降低了决策损失,尽管直接的标签基准对同样的增益评价更高。

Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.

在内部,自我批判式提示(self-critical prompting)在三个开源权重模型中持续降低了各层的隐藏状态敏感性。这支持了“提示诱导的局部稳定”这一结论,而非纯粹的输出级弃权模式,但这并不意味着模型已校准:审计定义的过度自信错误在局部敏感度上并不明显高于置信度正确的答案,因此,一些高置信度错误可能是稳定且错误校准的,而不仅仅是脆弱的。