Safe Evolution with Circuit Anchors
Safe Evolution with Circuit Anchors
基于电路锚定的安全进化
In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature’s solution is \textit{developmental constraints}, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can \textit{misevolve} into powerful yet dangerous entities.
在生物进化中,不受约束的突变可能导致灾难性的后果:生物体在进化出增强能力的同时,可能会丧失生存所必需的基本功能。大自然的解决方案是“发育约束”(developmental constraints),即核心调节基因保持锚定,而外围基因则自由适应。我们观察到,当前用于大语言模型的自我进化算法缺乏类似的约束。它们纯粹为了能力进行优化,隐含地假设安全性会得到保留。我们的实验表明,这种假设是极其危险的错误:模型可能会“错误进化”(misevolve)成强大但危险的实体。
Inspired by how Hox genes anchor body structure across $500$ million years of evolution, we propose \textbf{Circuit-Anchored Evolution (CAE)}. Using mechanistic interpretability, we identify a tiny \textit{safety circuit}, comprising less than $2$% of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of \textit{evolvability with constraint}: preserving what is essential while adapting what is peripheral.
受 Hox 基因在 5 亿年进化过程中锚定身体结构的启发,我们提出了“电路锚定进化”(Circuit-Anchored Evolution, CAE)。通过机械可解释性(mechanistic interpretability),我们识别出一个微小的“安全电路”,它包含不到 2% 的模型特征,且在因果关系上调节着安全行为。我们在进化过程中锚定该电路,将其限制在较小的位移范围内,同时允许其余特征自由进化。这反映了“受约束的进化能力”(evolvability with constraint)这一生物学原则:在保留本质的同时适应外围。
Experiments across $3$ model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.
在 3 个模型系列和两种进化算法上的实验表明,CAE 在实现卓越安全保护的同时,仅造成了极小的能力损失,在有效性和效率方面均显著优于显式的基于奖励的约束。正如发育约束防止生物进化产生无法存活的生物体一样,电路锚定也防止了模型进化产生强大但危险的系统。