How Claude’s text watermark works

How Claude’s text watermark works

Claude 的文本水印是如何运作的

Future Claude models will generate text that contains a watermark. This is a way of determining the likelihood that Claude was involved in writing the text, and we, along with several other major AI providers, are implementing this change to comply with the EU AI Act. 未来的 Claude 模型生成的文本将包含水印。这是一种确定 Claude 是否参与编写该文本的可能性(概率)的方法。我们与其他几家主要的 AI 提供商正在实施这一变更,以遵守《欧盟人工智能法案》(EU AI Act)。

In this article, we share answers to some of the questions we’ve received about how our chosen watermarking method works, whether it affects Claude’s outputs, and why we’re making this change. To summarize: 在本文中,我们将回答一些关于我们所选水印方法如何运作、它是否会影响 Claude 的输出,以及我们为何进行此项变更的问题。总结如下:

  • We use a method of watermarking that does not have any practical impact on the quality or content of Claude’s outputs;
  • 我们使用的水印方法对 Claude 输出的质量或内容没有任何实际影响;
  • The difference between watermarked and un-watermarked text will not be distinguishable to readers;
  • 读者无法区分带水印和不带水印的文本;
  • Nothing is added to the text and there are no hidden characters;
  • 文本中没有添加任何内容,也没有隐藏字符;
  • Watermarking doesn’t require extra tokens, and will not be more expensive;
  • 水印不需要额外的 Token,也不会增加成本;
  • Watermarking carries no identifying information and can’t be traced to a specific person, organization, or chat;
  • 水印不包含任何识别信息,无法追踪到特定的人、组织或聊天记录;
  • Watermarking won’t be specific to Claude. As of August 2, the EU requires AI providers serving its market to mark AI-generated content. Other major model developers have signed the same Code of Practice and will be implementing their own watermarks.
  • 水印并非 Claude 所独有。自 8 月 2 日起,欧盟要求为其市场提供服务的 AI 提供商标记 AI 生成的内容。其他主要模型开发商也签署了相同的《实践准则》,并将实施各自的水印。

What is watermarking?

什么是水印?

Large language models like Claude work by generating one word at a time. Each time the model decides on the next word, it chooses among a list of potential candidates, ultimately selecting the most sensible or likely based on the preceding text. Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number. 像 Claude 这样的大型语言模型通过一次生成一个词来工作。每当模型决定下一个词时,它会在一系列潜在候选词中进行选择,最终根据前面的文本选出最合理或最可能的词。以句子“The weather today was cold and…”(今天天气很冷,而且……)为例。下一个词不太可能是“sugary”(甜的),但很有可能是“overcast”(阴天的)或“grey”(灰蒙蒙的)。在大多数情况下,模型最终选择这两个词中的哪一个对读者来说并不重要——无论选哪一个,句子的含义基本相同。在这种情况下,选择由一个随机数决定。

Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude. 水印利用这些低风险的选择(在一段生成的文本中会多次出现)在 Claude 的回复中留下一种模式。这种模式对读者来说是不可见的,但对于拥有编码密钥的人来说是可以检测到的。当使用水印时,选择仍然是随机的,但随机性的来源不同。水印不是使用任意的随机数生成器来选择下一个词,而是使用密钥和前面的几个词来确定模型应该选择什么词。也就是说,Claude 选择的词仍然是随机的,但现在,人们可以检查词序列,看看它是否与 Claude 在使用密钥时会做出的选择一致。如果一致,就可以赋予该文本由 Claude 生成的概率。

Importantly, it isn’t that the model will now always be biased toward overcast or grey. Just as with non-watermarked text, overcast might be selected in one sentence, grey in the next, depending on the words that came before. And it’s not the case that the watermarking method pushes Claude to choose a word it wouldn’t have considered anyway (for instance, it wouldn’t make Claude pick a word like “nubilous”—an obscure synonym for overcast or grey that Claude almost certainly wouldn’t use under normal circumstances). 重要的是,这并不意味着模型现在总是会偏向于“overcast”或“grey”。就像没有水印的文本一样,根据前面的词,可能在一个句子中选择“overcast”,在下一个句子中选择“grey”。水印方法也不会强迫 Claude 选择它本来不会考虑的词(例如,它不会让 Claude 选择像“nubilous”这样晦涩的同义词,这是“overcast”或“grey”的生僻词,Claude 在正常情况下几乎肯定不会使用)。

How does watermarking affect Claude’s outputs?

水印如何影响 Claude 的输出?

Watermarking does not impact the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one (in this way, AI watermarks differ substantially from their namesakes on banknotes, other physical objects, and some digital documents, which are visible to the naked eye). 水印不会影响 Claude 的输出质量。对于读者来说,带水印的回复与不带水印的回复无法区分(从这一点来看,AI 水印与钞票、其他实物和某些数字文档上的同名水印有本质区别,后者是肉眼可见的)。

In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text. In the SynthID-Text paper, which introduced the technique we use, Google DeepMind tested this impact by serving a model that used watermarking to a portion of their Gemini traffic and comparing thumbs-up and thumbs-down ratings. They found no statistically significant differences from the unwatermarked model. And in a controlled study, human raters comparing watermarked and unwatermarked answers side-by-side saw no difference in quality. 在内部测试中,我们没有发现水印对 Claude 文本的内容、创造力水平或可读性有任何影响。在介绍我们所用技术的《SynthID-Text》论文中,Google DeepMind 通过将使用水印的模型应用于其部分 Gemini 流量,并比较点赞和点踩评分来测试这种影响。他们发现与不带水印的模型相比,没有统计学上的显著差异。在一项对照研究中,人类评估员并排比较带水印和不带水印的答案,也没有发现质量上的差异。

A useful analogy is to imagine you’re playing a game like Monopoly. On each turn, each player moves a random number of spaces around the board according to the roll of a die. Suppose that, instead of rolling the die to get this randomness, we decided to use a book of the digits of pi. We start from a randomly-chosen digit (say, the 1,012,845th after the decimal place, which happens to be a 6), and from that point on each player simply uses the next digit in the sequence as their next “roll”. 一个有用的类比是想象你在玩大富翁(Monopoly)之类的游戏。在每一轮中,每个玩家根据掷骰子的结果在棋盘上移动随机数量的格子。假设我们不掷骰子来获得这种随机性,而是决定使用一本记录圆周率(pi)数字的书。我们从一个随机选择的数字开始(比如小数点后第 1,012,845 位,恰好是 6),从那时起,每个玩家只需使用序列中的下一个数字作为他们的下一次“掷骰”结果。

For all intents and purposes, the moves are still random: it makes no difference to the players—or to the outcome of the game—whether the randomness comes from pi or from dice rolls each time. But if we could see the sequence of all the moves after the game (and we knew the value of pi), we could work out whether this was a game that likely used pi to determine its moves. The game that used pi is, in a sense, “watermarked”. 在所有实际目的中,移动仍然是随机的:随机性是来自圆周率还是每次掷骰子,对玩家或游戏结果都没有影响。但如果我们能在游戏结束后看到所有移动的序列(并且我们知道圆周率的值),我们就能算出这是否是一个可能使用圆周率来决定移动的游戏。从某种意义上说,使用圆周率的游戏就是“带水印的”。

It’s the same for Claude-generated text. Watermarking doesn’t change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so. Claude 生成的文本也是如此。水印不会改变阅读者的含义或体验,但如果你想事后检查该文本是否可能是由 Claude 生成的,水印允许你这样做。

Which specific method of watermarking do you use?

你们使用哪种具体的水印方法?

Claude’s text watermark is a version of the SynthID-Text approach published by Google DeepMind in a Nature paper in 2024. It belongs to a family of approaches that go back to a proposal by Scott Aaronson in 2022, all of which share the same design principle that we described above—the watermark only changes the source of the randomness used to pick among words. There are limitations to the effectiveness of watermarking. Using our key, one can only answer the question “What is the likelihood this was partly written by Claude?” It doesn’t confirm whether the text was human-written, and it can’t tell whether the text… Claude 的文本水印是 Google DeepMind 在 2024 年发表于《自然》(Nature)杂志的论文中所述的 SynthID-Text 方法的一个版本。它属于一系列可以追溯到 Scott Aaronson 在 2022 年提出的建议的方法,所有这些方法都共享我们上面描述的相同设计原则——水印只改变用于在词语之间进行选择的随机性来源。水印的有效性存在局限性。使用我们的密钥,人们只能回答“这段文本由 Claude 部分编写的可能性有多大?”这个问题。它不能确认文本是否由人类编写,也无法判断文本是否……