Funding better evaluations of AI’s impact on wellbeing

Funding better evaluations of AI’s impact on wellbeing

资助对人工智能福祉影响的更佳评估

We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how our models affect those who use them. Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.

我们正在启动一项 500 万美元的资助计划,旨在资助关于人工智能如何影响用户福祉的独立研究。该计划将为受资助者提供直接资金、模型访问权限以及技术支持,帮助他们构建开源评估工具,从而协助人工智能行业衡量我们的模型对用户的影响。受资助者将完全独立开展工作,并将其研究成果作为开源项目发布,供任何开发者使用。

AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But as an industry, we are still working towards developing clear standards for how models should behave in these conversations, for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.

人工智能系统已成为许多人工作、学习和解决问题的核心。它们也成为了对话伙伴,并在困难时期提供情感支持。但作为一个行业,我们仍在努力制定明确的标准,规范模型在这些对话中应如何表现,例如当用户开始寻求模型的陪伴,或使用人工智能来应对心理健康危机时。

Furthermore, wellbeing is a particularly difficult area to evaluate. For most model behaviors, we can look at a single answer and determine whether it is accurate and appropriate. But assessing wellbeing requires much more context. For example, a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation. And a response that might be reasonable in one context might be harmful in another. For example, Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate, and potentially actively harmful.

此外,福祉是一个特别难以评估的领域。对于大多数模型行为,我们可以通过查看单一答案来判断其是否准确和恰当。但评估福祉需要更多的背景信息。例如,处于困境中的用户可能不会立即表达自残的想法;只有在长时间的对话过程中,才可能显现出需要更谨慎回应的必要性。在一种语境下合理的回答,在另一种语境下可能是有害的。例如,当用户询问如何减肥时,Claude 可能会提供均衡饮食和锻炼计划的建议,但如果该用户有饮食失调史,这种回应可能是不恰当的,甚至可能造成实质性伤害。

We work to develop safeguards to identify such conversations and help ensure Claude responds appropriately, and we publish research into the types of conversations people have with Claude to better inform how we develop our safeguards, how we evaluate them, and other measures we can take to protect users’ wellbeing. But these are nuanced considerations, and the stakes are significant. The right approach will need to evolve alongside our models and their uses.

我们致力于开发安全防护措施来识别此类对话,并确保 Claude 做出适当的回应。我们还发布了关于人们与 Claude 对话类型的研究,以更好地指导我们如何开发防护措施、如何进行评估,以及采取其他措施来保护用户的福祉。但这些都是细微的考量,且影响重大。正确的方法需要随着我们的模型及其应用场景的发展而不断演进。

By funding the creation of independent evaluations and benchmarks of user wellbeing, we hope to invite more people to lend their expertise to this emerging and critical field, including clinicians, psychologists, methodologists, and others.

通过资助创建独立的用户福祉评估和基准测试,我们希望邀请更多人贡献他们的专业知识,共同参与到这个新兴且关键的领域中来,包括临床医生、心理学家、方法论专家等。

Towards more effective wellbeing evaluations and benchmarks

迈向更有效的福祉评估与基准测试

As part of this program, we’re sharing guidance from our Safeguards team on what we believe makes a wellbeing evaluation rigorous enough to build on, along with the common challenges that can limit an evaluation’s usefulness.

作为该计划的一部分,我们分享了来自安全团队的指导意见,阐述了我们认为什么样的福祉评估才足够严谨,并指出了可能限制评估有效性的常见挑战。

In brief, we’re seeking evaluations that: 简而言之,我们寻求的评估应具备以下特点:

  • State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters); 明确说明评估内容(即什么算作通过或失败,以及为什么这很重要);
  • Involve clinical and subject-matter experts in the design and validation; 在设计和验证过程中引入临床医生和学科专家;
  • Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal); 同时测试预防措施和潜在危害(即评估过度顺从和过度拒绝的风险);
  • Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations, where risk escalates and context shifts over the course of a long conversation); 反映用户实际使用人工智能的方式(这通常意味着构建代表多轮对话的场景,在这些场景中,风险会随着长时间对话的进行而升级,语境也会发生变化);
  • Validate their graders against real subject-matter experts. 通过真实的学科专家来验证评估者的准确性。

To learn more about the grant program and apply, see our application form. For more on building strong wellbeing evaluations and benchmarks, read our guidance. Applications are due by September 21; applicants who are selected to submit full proposals will be notified by October 5.

欲了解更多关于资助计划的信息并进行申请,请查看我们的申请表。如需了解更多关于构建稳健的福祉评估和基准测试的信息,请阅读我们的指导文档。申请截止日期为 9 月 21 日;被选中提交完整提案的申请者将在 10 月 5 日前收到通知。