Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
Knowing the Form, Not the Function: Automatically Auditing Answer—Authority Decoupling in Legal Benchmarks
知其形式,不知其功能:法律基准测试中答案与权威依据脱钩的自动审计
Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items.
法律基准测试通常只对最终答案进行评分,即使模型同时也陈述了法律依据。我们测试了答案的正确性是否可以作为权威依据(authority grounding)的代理指标。在未要求引用法条的常规推理提示下,四个大语言模型(LLM)在 238 道台湾司法考试题目中自发地生成了权威标记。
Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0—42.4% of valid responses were answer-correct but missed the gold authority, while 15.2—21.7% were answer-incorrect but cited it.
由于每道题目都有经过验证的适用条款,我们对答案正确性和权威依据进行了联合自动审计。这两个维度在两个方向上均出现了脱钩现象。在刑法领域,24.0% 至 42.4% 的有效回答虽然答案正确,但未能引用标准法律依据;而 15.2% 至 21.7% 的回答虽然答案错误,却引用了正确的法律依据。
A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes.
一项独立的法条检索探测实验和一项允许引用弃权的干预实验进一步表明,答案生成与引用行为在输出层面是可以独立变化的。由于这种不匹配并非源于对抗性或诱导不一致的提示,仅基于答案的评分方式会将自然发生的“遗漏标准法律依据”的情况视为完全成功的基准测试结果。
Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer—authority evaluation for statute-grounded legal benchmarks.
由于法律权威依据在结构上是可提取且可外部验证的,这种失败是可以自动测量的。一项针对中国民法的初步扩展研究也观察到了在未要求引用情况下的权威标记行为,这促使我们进行全面的跨司法管辖区联合审计。因此,我们建议针对基于法条的法律基准测试,采用答案与权威依据的联合评估方法。