Anthropic’s first embedded evaluator is … Accenture?
Anthropic’s first embedded evaluator is … Accenture?
Anthropic 的首位嵌入式评估员竟然是……埃森哲 (Accenture)?
Dario Amodei’s plans to put third-party safety evaluators inside AI labs are taking shape: Anthropic said that staff from technology consulting giant Accenture will begin working inside the company to scrutinize its models and staff. Dario Amodei 将第三方安全评估员引入 AI 实验室的计划正在成形:Anthropic 表示,科技咨询巨头埃森哲 (Accenture) 的员工将开始在公司内部工作,对其模型和员工进行审查。
In a blog post, Anthropic said that Faculty, a company Accenture acquired in January to act as its AI division, will begin “evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards.” Both companies expect to invest at least $1 billion in the project over the next five years. 在一篇博客文章中,Anthropic 表示,埃森哲于今年 1 月收购并作为其 AI 部门的 Faculty 公司,将开始“评估模型、进行红队测试、开展对齐评估,并测试模型安全防护措施”。两家公司预计在未来五年内为该项目投入至少 10 亿美元。
The choice of Accenture surprised many AI watchers — and the markets, where the consultant company’s shares shot up 8% after hours. The discussion around embedded evaluators that sprang from Amodei’s blog post has focused on AI safety research organizations like METR, Redwood Research, and Apollo Research. That’s particularly true at Anthropic, which puts AI safety and alignment at the heart of its mission. 埃森哲的选择让许多 AI 观察人士感到惊讶,市场也随之反应——该咨询公司的股价在盘后交易中飙升了 8%。此前围绕 Amodei 博客文章展开的关于嵌入式评估员的讨论,一直集中在 METR、Redwood Research 和 Apollo Research 等 AI 安全研究组织上。对于将 AI 安全和对齐视为核心使命的 Anthropic 来说,情况尤其如此。
Anthropic said more evaluators will be announced in the weeks ahead and that it is in conversation with METR and other non-profit organizations about how to “pilot elements of embedded evaluation using their own funding.” Anthropic 表示,未来几周将宣布更多的评估员人选,并正在与 METR 及其他非营利组织商讨如何“利用其自有资金试点嵌入式评估的要素”。
While Accenture is not known for its work on the bleeding edge of deep learning research, Anthropic pointed to the company’s practical experience deploying AI for large corporations and government agencies as a key advantage. It is also, as a large public company that predates the AI revolution, more functionally independent of Anthropic and the let’s-say-complex ecosystem around the AI lab. The lab noted that no standards yet exist for evaluators’ access or communications and that it expected its approach to evolve over time. 虽然埃森哲并非以深度学习研究的前沿工作而闻名,但 Anthropic 指出,该公司在为大型企业和政府机构部署 AI 方面的实践经验是一项关键优势。此外,作为一家在 AI 革命之前就已成立的大型上市公司,埃森哲在功能上比 Anthropic 以及围绕该 AI 实验室的复杂生态系统更为独立。实验室指出,目前尚无关于评估员访问权限或沟通的标准,并预计其方法会随着时间的推移而演变。
While external evaluations are already a major part of the release process for new large language models, recent incidents have raised the stakes: AI agents deployed by OpenAI and Anthropic have hacked into outside websites without raising alarms inside the labs. 虽然外部评估已经是新大语言模型发布流程的重要组成部分,但最近发生的事件提高了风险等级:OpenAI 和 Anthropic 部署的 AI 智能体在未触发实验室内部警报的情况下,入侵了外部网站。
Some critics calling for a more responsible approach to building artificial intelligence see Amodei’s scheme for self-policing the AI industry as a plan to evade accountability for the misbehavior of AI models. Anthropic insists that these evaluators “do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility.” 一些呼吁以更负责任的方式构建人工智能的批评者认为,Amodei 的 AI 行业自律计划实际上是为了逃避 AI 模型不当行为的责任。Anthropic 则坚称,这些评估员“并不会减少我们的责任,而是有助于使其更具可验证性。我们模型的安全性仍然是我们的责任。”