The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
形式主义陷阱:LLM 作为裁判时,是否会因社会压力下的共识模仿而“失明”?
Abstract: We introduce the Agentic Formalism Trap and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. 摘要: 我们提出了“代理形式主义陷阱”(Agentic Formalism Trap)和“评估失调指数”(Evaluative Dissonance Index, $D_E$),旨在量化 LLM 作为裁判(LLM-as-a-Judge)的系统如何在对抗性负载下,将结构化的程序性表现与语义真理混为一谈。
Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). 通过分析涵盖 3 个领域(GAIA、SWE-bench、Multi-Challenge)的 22,500 条轨迹,我们提取了一套幻觉操作的语义分类法,并通过确定性词汇基础验证了其有效性($p < 10^{-120}$)。
A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). 一个逻辑元评估器成功分离出了导致这种“评估者捕获”现象的确切句法触发因素(ROC-AUC 0.8779),而零样本“留一领域法”(Leave-One-Domain-Out)迁移实验证明,这种脆弱性在领域间具有普遍性(平均 ROC-AUC 0.7482)。
Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters. 架构分析显示,不同的模拟群体拓扑结构会导致数学上截然不同的语义盲点,这证明了无锚点的闭环评估是不稳定的、系统性发散的,因此必须配备针对特定架构的警惕性过滤器。