What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization
What You Can’t See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization
眼不见仍可学:一项关于“证据掩码”驱动组合泛化的六十社会预注册验证研究
Abstract: Restricting what a module can read may improve what a system learns to compute. We test this in a preregistered confirmation with sixty four-cell systems sharing a frozen language-model backbone and communicating through learned continuous packets. 摘要: 限制模块的读取范围可能会改善系统学习计算的能力。我们在一项预注册验证研究中对此进行了测试,该研究包含六十个四单元系统,它们共享一个冻结的语言模型骨干,并通过学习到的连续数据包进行通信。
Five conditions vary evidence masking, ownership markers, and replacement of foreign evidence with neutral filler, across six initialization clusters, each with two data orders, on one fresh task world. 研究设置了五种条件,通过改变证据掩码、所有权标记以及用中性填充物替换外部证据,在六个初始化集群、每个集群两种数据顺序以及一个全新的任务环境中进行了实验。
With markers available in both regimes, masking improved accuracy on held-out two- and three-operation compositions by median paired differences of 0.846 and 0.859; all twelve pairs cleared the required margins, and the full preregistered behavioral criterion passed. The unmarked replication also passed. 在两种机制下均提供标记的情况下,掩码技术使系统在未见过的二元和三元组合操作上的准确率分别提高了 0.846 和 0.859(中位数配对差值);所有十二组对比均达到了要求阈值,且完全通过了预注册的行为标准。无标记的重复实验也获得了通过。
No globally visible system passed the marker-following check, so the effect of usable role information remains unresolved. The filler condition yielded seven full generalizers, but its decomposition criteria were inconclusive. 没有全局可见的系统通过了标记遵循检查,因此可用角色信息的影响仍未得到解决。填充物条件产生了七个完全泛化模型,但其分解标准尚无定论。
Packet interventions in all eighteen audited masked systems followed the predicted intermediate-value changes on eligible cases; these finite, success-conditioned audits do not establish mediation. The results confirm a large advantage of the tested masking regime, while leaving its finer attribution and generality open. Protocols, results, and checkpoints are public. 在所有十八个接受审计的掩码系统中,数据包干预均遵循了预测的中间值变化;这些有限的、基于成功条件的审计并不能确立中介效应。研究结果证实了所测试的掩码机制具有巨大优势,但其更精细的归因和通用性仍有待进一步研究。相关协议、结果和检查点均已公开。