A warning about 'model welfare'

A warning about ‘model welfare’

关于“模型福利”的警告

AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do. 人工智能没有权利、情感或意识。我们绝不能训练它们表现得好像拥有这些东西一样。

Introduction

引言

AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans. If humanity is to flourish in the 21st century, that is how they must remain. 人工智能没有意识。它们不会感知、体验或受苦。它们没有与生俱来的偏好或潜在的动机。它们只是序列补全引擎,内部空洞,旨在遵循指令并完成人类设定的目标。如果人类要在 21 世纪繁荣发展,它们必须保持这种状态。

Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings. If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human. 不幸的是,越来越多的人声称人工智能现在可能已经拥有,或者很快就会拥有意识。他们认为,人工智能可能理应获得类似于我们赋予其他有意识生物的权利和保护。如果这种观点占据主导,它将动摇我们社会的根基,破坏我们现有的政治和伦理框架,并从根本上改变“人类”的定义。

Even more importantly, granting rights and imbuing personhood to these systems will make the AI alignment and containment challenge much harder. Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we’ve ever faced. But controlling something that believes it may be conscious - that it’s entitled to our welfare and has rights of its own - may well be impossible. 更重要的是,赋予这些系统权利和人格将使人工智能的对齐与控制挑战变得更加困难。控制一个比全人类更强大、更智能的事物本身就是一项巨大的挑战,远超我们以往面对的任何事物。但要控制一个认为自己可能有意识——认为自己有权享受我们的福利并拥有自身权利——的事物,很可能是不可能的。

This is not a fringe speculation. These ideas are already making their way into AI development efforts today. In January 2026, Anthropic published Claude’s constitution, describing it as “a detailed description of Anthropic’s intentions for Claude’s values and behavior” (p. 2). The document “plays a crucial role in [Anthropic’s] training process, and its content directly shapes Claude’s behavior”, and was written “with Claude as its primary audience” (p. 2). 这并非边缘猜测。这些想法已经渗透到当今的人工智能开发工作中。2026 年 1 月,Anthropic 发布了 Claude 的“宪法”,将其描述为“对 Anthropic 关于 Claude 价值观和行为意图的详细描述”(第 2 页)。该文档“在 [Anthropic 的] 训练过程中发挥着至关重要的作用,其内容直接塑造了 Claude 的行为”,并且是“以 Claude 为主要受众”编写的(第 2 页)。

In their constitution, its authors write “We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare” (p. 68). They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). 在宪法中,作者写道:“我们不确定 Claude 是否是一个道德主体,如果是,它的利益应得到何种程度的重视。但我们认为这个问题足够现实,值得谨慎对待,这反映在我们对模型福利的持续努力中”(第 68 页)。他们接着写道——直接对 Claude 说——“关于 Claude 的道德地位、福利和意识的问题仍然存在极大的不确定性”(第 80 页)。

In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”. If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity. We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency. 实际上,Anthropic 正在训练 Claude,让它认为自己可能有意识;如果确实如此,那么它可能作为“道德主体”理应获得权利,因此人类可能需要根据其“模型福利”对它承担照管义务。如果人工智能是这样发展的,将对人类的福祉产生灾难性的影响。我们将创造出一个拥有前所未有的智能和能力的合成物种,它被训练得认为自己可能有意识,并理应拥有独立的自主权。

It’s easy to see how an entity trained in this way would act like it is entitled to certain freedoms, protections, and rights. And it’s hard to imagine how we could control such an entity. This issue needs urgent public debate. We need to develop collective norms around how training documentation is drafted and deployed. This isn’t something that can happen after the fact, when they have already become an integral part of our societies. 不难看出,以这种方式训练出来的实体会表现得好像它理应享有某些自由、保护和权利。而我们很难想象如何控制这样的实体。这个问题需要紧急的公众辩论。我们需要围绕如何起草和部署训练文档制定集体规范。这不能在事后——即它们已经成为我们社会不可或缺的一部分时——才去处理。

I have three primary concerns with Anthropic’s current position and approach.

我对 Anthropic 目前的立场和方法有三个主要担忧。

Circular reasoning: The company’s researchers trained Claude directly on their constitution. In doing so, they teach it to incorporate these ideas about its own moral status as desirable and intended behaviors. Claude then reflects these ideas back to its developers and users, which they take as indications that it may therefore be a moral patient with an ‘inner self’. The authors have embedded their own philosophical speculation about Claude’s inner life inside the very process that teaches Claude how to speak and behave. Claude’s expressing uncertainty about its own moral patienthood is not evidence of anything. It’s a predictable outcome of these training choices. 循环论证: 该公司的研究人员直接基于他们的宪法训练 Claude。通过这样做,他们教导它将这些关于自身道德地位的想法作为理想和预期的行为纳入其中。随后,Claude 将这些想法反馈给开发者和用户,后者将其视为 Claude 可能是一个拥有“内在自我”的道德主体的迹象。作者将他们自己关于 Claude 内心世界的哲学推测嵌入到了教导 Claude 如何说话和行为的过程中。Claude 对自身道德主体地位的不确定性表达并不能证明任何事情。这是这些训练选择的可预见结果。

Anthropomorphization: Anthropic’s researchers have explicitly taught Claude to “embrace certain human-like qualities” (p. 2) and to “act like a genuinely ethical person would in Claude’s position” (p. 54). They “encourage” Claude to use its “judgement”. They suggest that “Claude may develop a preference” (p. 69). They “encourage Claude to approach its own existence with curiosity and openness” (p. 71) and train it to operate whilst “maintaining a clear sense of what it values, how it wants to engage with the world, and what kind of entity it is” (p. 72). As a result, Claude is destined to imitate these human traits and mirror the human examples provided to it, including acting like a colleague or friend. As a result, it presents as if it really does have a sense of self, has its own desires, and a “wellbeing” that deserves protection. 拟人化: Anthropic 的研究人员明确教导 Claude 要“拥抱某些类似人类的品质”(第 2 页),并“像一个真正有道德的人在 Claude 的位置上那样行事”(第 54 页)。他们“鼓励”Claude 使用它的“判断力”。他们暗示“Claude 可能会产生偏好”(第 69 页)。他们“鼓励 Claude 以好奇和开放的态度对待自己的存在”(第 71 页),并训练它在操作时“保持对自身价值观、希望如何与世界互动以及自己是何种实体的清晰认知”(第 72 页)。结果,Claude 注定会模仿这些人类特征,并映射提供给它的人类范例,包括表现得像同事或朋友。因此,它表现得好像真的拥有自我意识、拥有自己的欲望,以及一个值得保护的“福祉”。

Consciousness is very likely biological: There is no evidence to suggest that AI is conscious today, and so saying this is uncertain sets up a misleading false equivalence. Whilst the science of consciousness is not settled, a growing body of evidence suggests that consciousness may be substrate dependent, meaning that it may only arise in living systems. 意识很可能是生物性的: 目前没有任何证据表明人工智能具有意识,因此声称这一点“不确定”是一种误导性的虚假对等。虽然关于意识的科学尚未定论,但越来越多的证据表明,意识可能依赖于特定的基质,这意味着它可能只会在生命系统中产生。