An Anthropic researcher just gave us a peek at self-improving AI

An Anthropic researcher just gave us a peek at self-improving AI

Anthropic 研究员带我们一窥 AI 的自我进化

Training AI models with other AI models has become a very popular goal for neolabs — and now, a researcher in Anthropic’s fellows program has given us an early look at what it might look like in practice. 利用其他 AI 模型来训练 AI 模型已成为新兴实验室的热门目标。现在,Anthropic 研究员项目的一位研究人员让我们初步了解了这一目标在实践中可能的样子。

On Friday, Anthropic published a new paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” detailing how AI systems could reliably improve a model’s performance on a set of alignment benchmarks. When given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance. 周五,Anthropic 发表了一篇题为《自动化研究人员可以可靠地缓解对齐失败》(Automated Researchers Can Reliably Mitigate Alignment Failures)的新论文,详细介绍了 AI 系统如何可靠地提升模型在一系列对齐基准测试中的表现。在针对 10 种特定对齐偏差行为的基准测试中,这些自动化系统能够在不降低整体性能的情况下,改善每一项测试的表现。

Led by Anthropic fellow Chen Yueh-Han, the system replicates much of the traditional approach to research. Each automated system searches the available literature, proposes a method, and trains the model using that method for 30 minutes, gradually increasing the benchmark over several iterations. Effective methods are preserved while ineffective ones are discarded, allowing the system to operate quickly and at a great scale. 该系统由 Anthropic 研究员 Chen Yueh-Han 领导,复制了大部分传统的研究方法。每个自动化系统都会搜索现有文献,提出一种方法,并使用该方法对模型进行 30 分钟的训练,通过多次迭代逐步提高基准测试水平。有效的方法被保留下来,无效的方法则被剔除,这使得系统能够快速且大规模地运行。

“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term,” the paper reads. The paper is a step toward recursive self-improvement, which many see as the next significant step in AI progress. If models can improve their own alignment training, it’s plausible they could improve training practices more broadly — at which point, human AI researchers might soon become obsolete. 论文写道:“总的来说,这些结果提供了初步证据,表明自动化对齐后训练在短期内可能变得切实可行。”这篇论文是迈向递归自我改进的一步,许多人将其视为 AI 进步的下一个重要里程碑。如果模型能够改进自身的对齐训练,那么它们很有可能更广泛地改进训练实践——届时,人类 AI 研究人员可能很快就会变得过时。

The paper isn’t shy about addressing this idea, explicitly comparing the Automated Alignment Researcher (AAR) to its human equivalent. “The best AAR method beats what experienced humans propose, on average within six hours,” the paper reads. “Human guided research directions do not lead to stronger performance.” 论文毫不避讳地探讨了这一观点,明确将“自动化对齐研究员”(AAR)与人类研究员进行了对比。论文指出:“最好的 AAR 方法平均在六小时内就能胜过经验丰富的人类所提出的方案。人类引导的研究方向并不能带来更强的性能。”

There’s even a cost comparison, in case anyone wasn’t convinced. “An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.” 为了更有说服力,论文甚至进行了成本对比:“AAR 的 API 推理成本约为每小时 4 美元,而我们支付给人类研究人员的费用为每小时 150 美元。”

In fairness, the paper also points out a few limitations to this approach. The automated system only works insofar as the benchmarks reflect the actual alignment goals, and even then there’s significant work to be done in establishing and maintaining those benchmarks — not to mention maintaining and expanding on the literature the automated researchers are drawn from. 公平地说,论文也指出了这种方法的一些局限性。自动化系统只有在基准测试能够反映实际对齐目标时才有效,即便如此,在建立和维护这些基准测试方面仍有大量工作要做——更不用说维护和扩展自动化研究员所引用的文献库了。