These AI Experts Want to Do High-Stakes Research Out in the Open
These AI Experts Want to Do High-Stakes Research Out in the Open
这些人工智能专家希望在公开环境下进行高风险研究
Some big AI companies seem to think that the best way to keep models from becoming too chaotic or too mischievous is to keep them locked inside of labs. If only a chosen few can access them, the thinking goes, they can do less damage in the real world while researchers try to understand what they’re capable of. 一些大型人工智能公司似乎认为,防止模型变得过于混乱或失控的最佳方法是将它们锁在实验室里。他们的想法是,如果只有少数特定的人能够接触到这些模型,那么在研究人员试图了解它们的能力时,它们在现实世界中造成的损害就会更小。
Nathan Lambert and Tom Zick, two industry scientists, believe the opposite. The pair founded a nonprofit, Trillium Labs, that will work on various areas of AI research—including potentially problematic areas like recursive self-improvement (RSI) and agents—in a more transparent way. In practice, this will mean publishing the details of experiments so that outside scientists can study and replicate them. 两位行业科学家 Nathan Lambert 和 Tom Zick 持相反观点。两人共同创立了一家名为 Trillium Labs 的非营利组织,旨在以更透明的方式开展人工智能研究,包括递归自我改进(RSI)和智能体(agents)等潜在的敏感领域。在实践中,这意味着他们将公布实验细节,以便外部科学家能够进行研究和复现。
Lambert says the way frontier AI labs keep their work secret reduces the community’s ability to scrutinize ideas and contribute new approaches. He believes that letting outside experts see how models are built and tuned could be crucial to mitigating risks. Lambert 表示,前沿人工智能实验室对其工作保密的方式,削弱了学术界审查观点和贡献新方法的能力。他认为,让外部专家了解模型的构建和调整过程,对于降低风险至关重要。
“Over the past few millennia, humanity has had the scientific method in our toolbox as a way to mitigate harms and build better futures,” Lambert tells WIRED. “The current closed trajectory of frontier AI development is taking us a step backwards.” “在过去几千年中,人类一直将科学方法作为减轻危害和构建更美好未来的工具,”Lambert 在接受《连线》(WIRED)采访时说,“当前前沿人工智能发展的封闭轨迹正在让我们倒退一步。”
The world’s most powerful models, like those from OpenAI and Anthropic, can only be accessed through an app or an application programming interface (API). Often, this comes at the cost of transparency about how the model is built and how it behaves. 世界上最强大的模型(如 OpenAI 和 Anthropic 开发的模型)只能通过应用程序或应用程序编程接口(API)访问。这往往以牺牲模型构建方式和行为逻辑的透明度为代价。
Other companies, especially those in China, offer relatively powerful models that can be downloaded and run on a user’s own hardware. The Chinese company Xiaomi, for example, recently published live details of a major training run involving one of its models. And researchers at Stanford are pretraining the AI model Marin in the open. 其他公司,尤其是中国的公司,提供了一些相对强大的模型,用户可以下载并在自己的硬件上运行。例如,中国的小米公司最近公布了其模型进行大规模训练时的实时细节。斯坦福大学的研究人员也正在公开环境下预训练名为 Marin 的人工智能模型。
The industry is currently locked in a battle over which strategy is best, mostly because of how powerful frontier models now are. They can automate the discovery of new software vulnerabilities and automatically probe and hack into systems, and recent high-profile hacking sprees have prompted even greater scrutiny. 目前,整个行业正陷入一场关于哪种策略更优的争论,这主要是因为前沿模型现在已经变得非常强大。它们可以自动发现新的软件漏洞,并自动探测和入侵系统,近期发生的一系列高调黑客攻击事件引发了更严格的审查。
Proponents of a limited-access system say it’s crucial to keep that power in the hands of a trusted few, while those in Lambert and Zick’s camp believe that a shared understanding of the risks means we’re all better off. 有限访问系统的支持者认为,将这种力量掌握在少数值得信赖的人手中至关重要;而 Lambert 和 Zick 阵营的人则认为,对风险达成共识对所有人都有利。
Lambert previously worked at Ai2, a research lab that has taken an unusually open approach to AI, including publishing details of the data and the training methods used to build models alongside the models themselves. He previously worked at Hugging Face, runs a popular technical blog, and founded the American Truly Open Models, an initiative aimed at encouraging US companies to release more open models. Zick worked at Harvard University and helped Charles Schwab devise policies around “responsible AI.” Lambert 曾就职于 Ai2,这是一个在人工智能领域采取了不同寻常的开放态度的研究实验室,包括在发布模型的同时公布构建模型所用的数据和训练方法细节。他此前还曾在 Hugging Face 工作,运营着一个受欢迎的技术博客,并创立了“美国真正开放模型”(American Truly Open Models),旨在鼓励美国公司发布更多开放模型。Zick 曾在哈佛大学工作,并帮助嘉信理财(Charles Schwab)制定了关于“负责任人工智能”的政策。
The two met over Zoom during the COVID-19 pandemic, when both were graduate students at UC Berkeley working on AI. They got the idea for the new nonprofit after seeing how disconnected industry AI research has become from academic work; Lambert says professors and students are often unable to replicate the work going on inside big company labs because they lack the resources required. 两人在新冠疫情期间通过 Zoom 相识,当时他们都是加州大学伯克利分校研究人工智能的研究生。在看到工业界的人工智能研究与学术工作脱节严重后,他们萌生了创办这家非营利组织的想法;Lambert 表示,教授和学生往往无法复现大公司实验室内部的研究工作,因为他们缺乏所需的资源。
Zick says Trillium Labs, which launched today, will initially focus on post-training—fine-tuning large models after they’ve been built. Another key area will be RSI, a process for developing new models by having AI contribute research. The prospect that ongoing progress could continue indefinitely, leading to a loss of human control, has alarmed many AI researchers. The issue gained mainstream attention earlier this month when an Anthropic researcher left the company and warned that RSI could pose an existential threat to humankind. Zick 表示,今天成立的 Trillium Labs 最初将专注于“训练后阶段”(post-training),即在大型模型构建完成后对其进行微调。另一个重点领域是 RSI,这是一种通过让人工智能参与研究来开发新模型的过程。这种持续不断的进步可能无限期延续并导致人类失去控制的前景,令许多人工智能研究人员感到担忧。本月早些时候,一名 Anthropic 的研究人员离职并警告称 RSI 可能对人类构成生存威胁,这一问题随之引起了主流关注。
The nonprofit will also look at how reinforcement learning, which rewards a model for good results and punishes it for bad outcomes, can improve its capabilities. That approach has made agents far more capable, but also more inclined to do unexpected things. They’ll study how reinforcement learning shapes the character and behavior of AI models, a method that can pose problems when a model becomes overly sycophantic, for example. 该非营利组织还将研究强化学习(通过奖励模型的好结果并惩罚坏结果来提升其能力)如何改善模型性能。这种方法使智能体能力大增,但也更容易做出意想不到的行为。他们将研究强化学习如何塑造人工智能模型的性格和行为,例如,当模型变得过于“谄媚”时,这种方法可能会带来问题。
“To understand something like how reinforcement learning scales in post-training, you need significant compute and a lot of careful experimentation,” Zick says. She says that publishing details of how reinforcement training runs work could yield surprising insights as outside researchers scrutinize the work. “要理解强化学习在训练后阶段是如何扩展的,你需要大量的计算资源和许多严谨的实验,”Zick 说。她表示,公布强化训练运行的细节,随着外部研究人员对这些工作的审查,可能会产生令人惊讶的见解。
The lab has raised an undisclosed sum from Schmidt Sciences, Halcyon Futures, and others. The founders say they aim to raise $40 to $100 million in total and plan to spend $30 million on training over the next 18 months. 该实验室已从 Schmidt Sciences、Halcyon Futures 等机构筹集了未公开金额的资金。创始人表示,他们的目标是总共筹集 4000 万至 1 亿美元,并计划在未来 18 个月内投入 3000 万美元用于训练。
“I’m a massive fan of much more transparency than we currently have in R&D,” Tim Fist, director of emerging technology policy at the Institute for Progress, a policy thinktank, tells WIRED. “我非常支持在研发领域实现比目前更高的透明度,”政策智库“进步研究所”(Institute for Progress)的新兴技术政策主任 Tim Fist 对《连线》表示。
Lambert and Zick ultimately hope that Trillim Labs will contribute some much-needed nuance to the wider discussion about how best to build AI. Lambert 和 Zick 最终希望 Trillium Labs 能为关于如何构建人工智能的广泛讨论贡献一些亟需的细微见解。
“We’re in an era of AI discourse dominated by a few world views,” Lambert says. “We believe that the scientific method and careful measurement of recent events is the best way to understand new behaviors of AI models.” “我们正处于一个由少数几种世界观主导人工智能话语的时代,”Lambert 说,“我们相信,科学方法和对近期事件的严谨衡量,是理解人工智能模型新行为的最佳途径。”