Improving Fable 5's biology safeguards
Improving Fable 5’s biology safeguards
改进 Fable 5 的生物学安全防护机制
We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer “fallbacks”—where the system switches to a less capable model after they make a biology-related query. In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces. Fable 5 will thus be able to assist with a wider range of biology tasks. 我们正在对 Claude Fable 5 的生物学安全防护机制进行更新,以大幅减少误报。Fable 5 的用户现在将体验到更少的“回退”情况——即系统在用户进行生物学相关查询后切换到能力较弱的模型。在我们的测试中,此次更新使各产品界面中与生物学相关的回退率降低了约 85%。因此,Fable 5 将能够协助处理更广泛的生物学任务。
In practice, users should see far fewer fallbacks on everyday health and educational questions—for example, interpreting lab results, understanding symptoms, and learning about biology in an educational context. Healthcare professionals will be able to receive more support from Fable 5 on clinical tasks. 在实际应用中,用户在处理日常健康和教育问题时,遇到的回退情况应该会大幅减少——例如解读化验结果、理解症状以及在教育背景下学习生物学知识。医疗专业人员将能够从 Fable 5 获得更多临床任务方面的支持。
We believe the greatest opportunity for AI to positively affect the world is in biology and medicine, and we’re investing significantly in building a responsible way to give biologists frontier access. Today, Fable still falls back to Opus 5 for requests we consider dual-use—including virology, toxicology, and molecular design—so it isn’t yet usable for professional biology research and drug development. We’re committed to closing that gap through trusted access pathways for frontier biology capabilities. 我们相信,人工智能对世界产生积极影响的最大机遇在于生物学和医学领域。我们正投入大量资源,以负责任的方式为生物学家提供前沿访问权限。目前,对于我们认为具有“双重用途”的请求(包括病毒学、毒理学和分子设计),Fable 仍会回退至 Opus 5,因此它尚不能用于专业的生物学研究和药物开发。我们致力于通过可信的访问路径,弥合这一差距,以提供前沿的生物学能力。
Why we built strong biology safeguards
我们为何建立强大的生物学安全防护机制
Our objective is to get Fable 5’s frontier capabilities into the hands of as many of our users as possible, as quickly as possible. However, to do so, we need to manage the increasing risks that come with models this capable. One such risk is in the field of biology: Fable 5 can now outperform experts on some highly complex biological tasks and provide operational support on others. That means that it can provide genuine assistance to a researcher developing a new medical treatment (which is the reason we’re so keen to widen access to the model via both classifier improvements and trusted access programs). But in the wrong hands, those same capabilities could be used by a malicious actor, for example in developing a biological weapon. Our capability assessments show that Fable 5 could provide significant uplift to such an actor—that is, it could provide them with capabilities they could not find anywhere else. 我们的目标是尽快将 Fable 5 的前沿能力提供给尽可能多的用户。然而,要做到这一点,我们需要管理这种高性能模型所带来的日益增长的风险。其中一个风险存在于生物学领域:Fable 5 现在在某些高度复杂的生物学任务上可以超越专家,并为其他任务提供操作支持。这意味着它能为开发新医疗方案的研究人员提供真正的帮助(这也是我们热衷于通过改进分类器和可信访问计划来扩大模型访问权限的原因)。但如果落入不法之徒手中,同样的能力可能会被恶意行为者利用,例如用于开发生物武器。我们的能力评估显示,Fable 5 可能会为这类行为者提供显著的助力——也就是说,它能提供他们在其他任何地方都无法获得的能力。
It’s often difficult to tell apart beneficial and harmful uses of AI in biology. For example, in some cases researching a treatment for a disease requires scientists to produce the dangerous compounds that cause that disease in the first place. This is most obvious for live vaccines, which require scientists to grow the same pathogen they’re aiming to prevent. It’s also the case for some medicines. To develop the drug captopril, which treats hypertension, scientists isolated toxic components of snake venom that crash blood pressure in humans. As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits. 在生物学领域,区分人工智能的有益用途和有害用途往往很困难。例如,在某些情况下,研究某种疾病的治疗方法需要科学家首先制造出导致该疾病的危险化合物。这在活疫苗中最为明显,因为科学家需要培养他们旨在预防的同一种病原体。某些药物也是如此。为了开发治疗高血压的药物卡托普利(captopril),科学家们分离出了蛇毒中会导致人体血压骤降的毒性成分。随着人工智能前沿领域新生物学能力的不断发展,我们需要保持谨慎,确保它们带来的新风险不会在潜在的科学益处显现之前就成为现实。
Sophisticated actors who wish to use our models to do harm know how to exploit this ambiguity to obscure their intent, making dangerous tasks look like ordinary research pursuits. The US Intelligence Community’s 2026 Annual Threat Assessment makes clear that such actors exist, and that advances in biotechnology including synthetic biology and genomic editing “could lead to novel biological threats.” It notes that several state actors likely maintain active offensive biological and chemical weapons programs—programs that could be accelerated by access to the raw capabilities of frontier AI models. 那些希望利用我们的模型进行破坏的复杂行为者,深知如何利用这种模糊性来掩盖其意图,使危险任务看起来像普通的科研追求。美国情报界 2026 年的年度威胁评估明确指出,此类行为者确实存在,且包括合成生物学和基因编辑在内的生物技术进步“可能导致新型生物威胁”。报告指出,一些国家行为者可能正在维持活跃的进攻性生物和化学武器计划——而获取前沿人工智能模型的原始能力可能会加速这些计划。
Because of our concerns about these “dual-use” capabilities (those that could be used for beneficial or harmful purposes, and where the line between them is not always easy to draw), we intentionally launched Fable 5 with almost all biology queries blocked. This enabled us to make the model available for users in other domains. We knew this would be frustrating for legitimate biology users: it would result in a high number of false positives in the near term, where users asking biology-related questions would have their requests blocked and sent to a less capable model. Nevertheless, we chose to make this tradeoff because the cost of Fable being misused in a dual-use domain like biology could potentially be catastrophic. 由于我们对这些“双重用途”能力(即可能被用于有益或有害目的,且两者之间的界限并不总是容易划清的能力)感到担忧,我们特意在发布 Fable 5 时屏蔽了几乎所有的生物学查询。这使我们能够将该模型提供给其他领域的用户。我们知道这对合法的生物学用户来说会令人沮丧:这会导致短期内出现大量的误报,即用户提出的生物学相关问题会被拦截并转至能力较弱的模型。尽管如此,我们还是选择了这种权衡,因为 Fable 在生物学等双重用途领域被滥用的代价可能是灾难性的。
How our biology safeguards work
我们的生物学安全防护机制如何运作
One of the core ways we protect against misuse in biology is via safety classifiers: smaller, automated AI systems that detect when Fable 5 is asked to perform a safeguarded biology task, or produce a harmful output (we’ve previously written about our similar classifiers in the domain of cybersecurity). 我们防止生物学领域滥用的核心方法之一是通过安全分类器:这是一种较小的自动化人工智能系统,用于检测 Fable 5 何时被要求执行受保护的生物学任务,或产生有害输出(我们之前曾撰文介绍过我们在网络安全领域类似的分类器)。
In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked. 对于 Fable 5,当分类器触发时,模型会将用户的请求重新路由至 Opus 5。Opus 5 是一款功能强大的模型,但其生物学能力不及 Fable 5,且无法为恶意用户提供同等程度的协助。这就是用户在请求被拦截时所看到的回退现象。
Developing precise, robust classifiers is not a straightforward task. For a classifier to work rapidly and consistently, it has to learn the difference between what we consider “in scope” and “out of scope” for the topics and queries we consider to be potentially harmful. It takes time and iteration to tune the classifiers, avoiding both false positives (where classifiers fire on out-of-scope content) and false negatives (where in-scope content is missed). We also require our classifiers to be robust to attempts to bypass them (known as jailbreaks), which requires even further research and testing. 开发精确、稳健的分类器并非易事。为了使分类器能够快速且一致地工作,它必须学习区分我们认为“在范围内”和“在范围外”的主题与查询,即那些我们认为具有潜在危害的内容。调整分类器需要时间和迭代,以避免误报(分类器对范围外的内容触发)和漏报(范围内的内容被遗漏)。我们还要求分类器能够抵御绕过尝试(即“越狱”),这需要进行更深入的研究和测试。
Starting with a very broad biology classifier meant that we could give our users access to Fable 5 while we continued our research aimed at refining it. The alternative—holding back the model until much more safeguards progress was made—would have delayed the model’s general access, and its potential benefits to our users, by weeks or months. 从一个非常广泛的生物学分类器开始,意味着我们可以在继续进行优化研究的同时,让用户使用 Fable 5。另一种选择——在取得更多安全防护进展之前暂缓发布模型——则会将模型的全面开放及其对用户的潜在益处推迟数周或数月。
Over the past several weeks, we’ve carefully… 在过去的几周里,我们已经仔细地……