Partnering with Accenture on embedded evaluation
Partnering with Accenture on embedded evaluation
与埃森哲(Accenture)合作开展嵌入式评估
We’re partnering with Accenture on independent evaluation of frontier AI. This is an important step toward the commitment, made in our CEO’s essay “We Must Pace the Frontier,” to embed evaluators within Anthropic. 我们正与埃森哲(Accenture)合作,对前沿人工智能进行独立评估。这是我们朝着首席执行官在文章《我们必须跟上前沿步伐》(We Must Pace the Frontier)中所作承诺迈出的重要一步,即在 Anthropic 内部嵌入评估人员。
The partnership will be led by Faculty, Accenture’s specialist AI business, and will include evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards. Accenture helps businesses and governments deploy AI across many industries. Their understanding of how enterprises use AI in practice informs their safety approach, and they will bring that perspective to evaluating our models. 此次合作将由埃森哲旗下的专业人工智能业务部门 Faculty 领导,工作内容包括评估模型、进行红队测试、开展对齐评估以及测试模型安全防护措施。埃森哲致力于帮助各行各业的企业和政府部署人工智能。他们对企业如何实际应用人工智能的理解,为其安全方法提供了支撑,他们也将把这一视角带入到对我们模型的评估中。
Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years. Anthropic 和埃森哲预计在未来五年内,双方将各自投入至少 10 亿美元,用于构建该领域的评估能力。
Embedded evaluation is new, and many of the details about how it will operate are still being worked out. Unlike today’s external evaluators, embedded evaluators will work inside AI companies, with access comparable to an employee’s. That access allows them to watch models take shape in training, follow the decisions that govern how those models are built and deployed, and speak directly to employees. From this vantage point, embedded evaluators can assess how a company operates, verify that it is keeping its safety commitments, and identify blind spots. They can also report incidents and give the public a more informed account of benefits and risks. “嵌入式评估”是一个新概念,其运作细节仍在制定中。与当今的外部评估机构不同,嵌入式评估人员将在人工智能公司内部工作,并拥有与员工相当的访问权限。这种权限使他们能够观察模型在训练过程中的成型过程,追踪模型构建和部署的决策过程,并直接与员工沟通。从这一视角出发,嵌入式评估人员可以评估公司的运营方式,核实其是否履行了安全承诺,并识别潜在的盲点。他们还可以报告事故,并为公众提供关于人工智能收益与风险的更知情报告。
To be clear, independent embedded evaluators do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility. 需要明确的是,独立的嵌入式评估人员并不会减轻我们的责任,而是有助于使我们的责任更具可验证性。我们模型的安全性依然由我们负责。
There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation. Long-term, we think funding should come from pooled or government sources, as we called for in our Advanced AI Framework in June. As neither exists today, we plan to work with different evaluators under different funding arrangements. 目前,关于嵌入式评估人员应具备哪些信息访问权限,以及应如何报告其发现,尚无统一标准。此外,独立评估的资金来源也尚未形成固定体系。从长远来看,我们认为资金应来自联合基金或政府来源,正如我们在六月份发布的《先进人工智能框架》(Advanced AI Framework)中所呼吁的那样。由于目前这两者均不存在,我们计划在不同的资金安排下与不同的评估机构合作。
Given the importance and urgency of this work, Anthropic will fund Accenture’s work directly. We are also in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding. Ultimately, we believe frontier AI needs an ecosystem of evaluators operating with shared standards. 鉴于这项工作的重要性和紧迫性,Anthropic 将直接资助埃森哲的工作。我们同时也正在与 METR 及其他非营利性评估机构进行对话,以试点使用其自有资金开展嵌入式评估的相关要素。最终,我们相信前沿人工智能需要一个在共享标准下运作的评估生态系统。
We expect frontier labs to work with several organizations at once. Our partnership is non-exclusive; Anthropic will work with other evaluators to be announced in the coming weeks, and Accenture will work with other AI developers in similar capacities. 我们预计前沿实验室将同时与多家机构合作。我们的合作是非排他性的;Anthropic 将在未来几周内宣布与其他评估机构的合作,而埃森哲也将以类似的能力与其他人工智能开发商合作。
We’ll continue to train and release frontier models, and we want independent evaluators working alongside us as we do. We’re sharing these early efforts now so people and other AI developers can see our process. We expect our approach to evolve as the field matures, and we’ll share more as our work begins and as we bring on additional evaluators. 我们将继续训练并发布前沿模型,并希望在这一过程中有独立评估人员与我们并肩工作。我们现在分享这些初步努力,是为了让公众和其他人工智能开发商了解我们的流程。我们预计随着该领域的成熟,我们的方法也会不断演进,随着工作的开展和更多评估人员的加入,我们将分享更多信息。