When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
多智能体代码评估器何时才算真正有据可依?两种无需标签的度量方法及一个懂得“拒绝猜测”的评估器
Abstract: When one language model judges whether another’s code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. 摘要: 当一个语言模型判断另一个模型的代码是否正确时,它并不会报告证据的缺失。相反,它会给出一个带有推理过程的自信结论,这与它确实有据可依时给出的结论别无二致。
Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. 多智能体验证(Multi-agent verification)将判断过程分解为可检查的声明,并针对证据逐一验证,这是一种很有前景的方案,且在证据为一组检索文档时效果良好。我们认为,此类方法对证据有两点要求:证据必须独立于被审查的答案,且在比较两个候选对象时,证据必须有所差异。
The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions equally good on 78 to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly reaches 43.7%. Neither easier problems nor a larger judge changes this. 第二个条件在处理检索文档时会自动满足,但在代码评估中却不再成立。我们在两个代码评估基准测试上运行了未经修改的已发布框架 MARCH,进行了 80 次条件单元测量。结果发现,在 78% 到 95% 的比较中,该框架认为两个解决方案同样优秀,准确率仅为 4.4%,而直接询问同一模型时准确率可达 43.7%。无论是更简单的问题还是更大的评估模型,都无法改变这一现状。
Two measurements taken from the pipeline’s own logs explain it without needing labels. Gating on one of them, the pipeline declines the comparisons it cannot make and raises its accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is not a more accurate judge, but a label-free way to tell when a judge has no basis for its answer. 通过从流水线自身的日志中提取两个度量指标,我们无需标签即可解释这一现象。通过对其中一个指标进行门控(Gating),流水线能够拒绝无法判断的比较,从而将准确率从 20.7% 提升至 36.9%,同时仍能完成一半的比较任务。本文的贡献不在于构建了一个更精确的评估器,而在于提供了一种无需标签的方法,用以判断评估器何时对其答案缺乏依据。