Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
托管大语言模型中的“无持久性复制”:行动时刻信念评估中的测量敏感性
Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence).
摘要: 托管大语言模型的行为评估结果可能会产生差异,原因在于评估的服务、测量工具或两者在不同运行期间发生了变化。我们区分了三个验证问题:先前的研究结果在历史配置下使用新数据时是否会重现(复制);当在相同标识符下重建评估与推理配置时,评估终点是否会发生变化(测量敏感性);以及该研究结果在同一工具下对后续测试的标识符是否具有持续性(持久性)。
We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]).
我们在 Regent Chess 中研究了这些问题。这是一个序列化环境,其中隐藏且可变的状态被精确记录,从而允许在行动时刻根据事实真相(ground truth)对陈述的信念进行评分;正的终点值意味着性能比匹配的均匀比较器(matched-uniform comparator)更差。此前报告的 Gemini 3.1 Flash-Lite 的性能缺陷在历史配置下的新对局中再次出现(+0.0530,95% 置信区间 [+0.0329, +0.0714])。
In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern.
在同一公共标识符下进行的背靠背、同日 H/R(历史/重建)对比中,重建配置下的“模型减去均匀”终点值降低了 0.0429(H 减 R 对比的 95% 置信区间为 [+0.0182, +0.0667]);由于所有六个配置组件共同变化,因此无法隔离出单一组件的影响。在重建的 R 配置下,预先冻结、交错的同窗口 4K 对比在 Gemini 3.1 和 Gemini 3.7 之间出现了符号反转,这两个标识符在发布版本和产品层级上有所不同;额外的描述性和探索性单元格显示了相同的方向模式。
Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.
任何额外的服务周期贡献仍未得到解决(-0.0166,[-0.0483, +0.0157])。因此,复制、测量敏感性和持久性在同一评估中可能会得出不同的结论,这促使我们有必要通过测试标识符、服务周期、测量工具和推理配置,对托管模型的行为声明进行明确的索引。