More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
更多程序还是更多尝试?在大语言模型(LLM)工具链中区分覆盖率与专业化
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. 自动化生成大语言模型(LLM)工具链有望通过任务专业化来提升推理能力。然而,额外的答案覆盖率可能源于对同一程序的重复执行,这使得专业化程度难以界定。我们引入了一种受控评估方法,将答案覆盖率、可重复的任务优势以及预执行选择带来的增益区分开来。
On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. 在 386 个 MATH-500 任务上,我们将八个生成的工具链和一个基准模型,与九个字节完全相同的基准副本进行了对比,每个成员执行三次。完全相同的程序产生了 2.16 个百分点的重复平均预言机(oracle)提升空间。生成的程序表现出明显更具可重复性的评分模式,但这主要揭示了持续存在的弱点:在 100 个任务中,相对于基准的损失在所有三次重复中持续存在,而持续的获胜仅出现在一个任务中,且对答案提取过程非常敏感。
The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. 固定选择器(frozen selector)的增益为 0.00 个百分点,且两个群体在 27 次工具链执行后均达到了 98.70% 的预言机覆盖率。在三次重复执行后,稳定的互补性仍未得到解决。支持性的 BIRD 追踪定位了机制实现、激活和输出有效性方面的失败。
Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets. 总之,这些发现确立了为什么仅凭覆盖率和可重复性不足以证明所谓的“有效专业化”。这些发现推动了针对工具链多样性的评估标准的建立:任务优势应在多次执行中保持一致,能够指导可用的决策,并且在匹配的推理预算下,优于额外的固定程序执行。