Anthropic shares more details about how Claude’s new watermarks will work
Anthropic shares more details about how Claude’s new watermarks will work
Anthropic 分享了关于 Claude 新水印工作原理的更多细节
Anthropic published a blog post Friday seeking to answer some basic questions about how it will watermark the text generated by its chatbot Claude. Such as: How will the watermarking actually work? Can it be hidden with editing? And how does this affect code?
Anthropic 周五发布了一篇博文,旨在回答关于其如何为聊天机器人 Claude 生成的文本添加水印的一些基本问题。例如:水印究竟是如何工作的?通过编辑可以隐藏吗?这对代码有什么影响?
Claude users have been debating the move since the company revealed earlier this week that it would be doing this watermarking to comply with the EU AI Act’s Transparency Code, which requires AI companies to use systems that make it possible to identify AI-generated content. On Reddit, for example, one poster characterized this as a conspiracy against innocent Claude users, while another claimed, “The only reason you wouldn’t want this is to lie to people.” And Business Insider reports that “dozens” of users on X have claimed to cancel their Claude subscriptions as a result.
自本周早些时候该公司透露将实施此项水印措施以遵守《欧盟人工智能法案》(EU AI Act)的透明度准则(该准则要求人工智能公司使用能够识别 AI 生成内容的系统)以来,Claude 的用户一直在争论这一举措。例如,在 Reddit 上,一位发帖者将其描述为针对无辜 Claude 用户的阴谋,而另一位则声称:“你不想这样做的唯一原因就是为了欺骗他人。”据《商业内幕》(Business Insider)报道,X 平台上有“数十名”用户声称因此取消了他们的 Claude 订阅。
Anthropic’s new post starts with a general overview of the watermarking concept, explaining that when making “low-stakes choices” — like choosing between the words “overcast” and “grey” to describe the weather — Claude can create a pattern in its responses that is “undetectable to the reader, but is detectable to anyone who has a key that encodes it.”
Anthropic 的新文章首先概述了水印概念,解释说当做出“低风险选择”时——比如在描述天气时选择“阴天(overcast)”还是“灰色(grey)”——Claude 可以在其回复中创建一种模式,这种模式“对读者来说是不可察觉的,但对于任何拥有编码密钥的人来说都是可检测的。”
“Watermarking does not impact the quality of Claude’s output,” the company said. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”
“水印不会影响 Claude 的输出质量,”该公司表示。“对于读者来说,带有水印的回复与没有水印的回复是无法区分的。”
More specifically, Anthropic said it will be using the SynthID-Text approach that the Google DeepMind team outlined in 2024, and that it plans to release a watermark detection API. It also noted that watermarking is distinct from the AI detection approaches offered by companies like Pangram that look for “tells” in the writing (like the construction “his isn’t [X], it’s [Y]”) to reveal AI usage: “Picking up on these patterns is fundamentally different from checking for a watermark.”
更具体地说,Anthropic 表示将使用 Google DeepMind 团队在 2024 年概述的 SynthID-Text 方法,并计划发布一个水印检测 API。它还指出,水印与 Pangram 等公司提供的 AI 检测方法不同,后者通过寻找写作中的“破绽”(如“这不是 [X],而是 [Y]”这种结构)来揭示 AI 的使用:“捕捉这些模式与检查水印有着本质的区别。”
Could someone just rewrite the text to hide the watermark? Anthropic said it’s possible, but “light editing probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.” “In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said.
有人可以通过重写文本来隐藏水印吗?Anthropic 表示这是可能的,但“轻微的编辑可能无法完全去除水印”,而“替换掉每一个词的彻底重写则可以”。该公司表示:“当然,在后一种情况下,该文本是否还能被称为 AI 生成的内容是有争议的。”
As for whether the watermark will be detectable in text that was only proofread or edited by Claude, Anthropic said that will depend on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author and “there’s very little (if anything) for the watermark to attach to.”
至于仅由 Claude 校对或编辑的文本是否能检测到水印,Anthropic 表示这取决于“文本的长度以及 Claude 编辑的程度”。如果只是轻微编辑,“几乎所有的词”都是由人类作者撰写的,“水印几乎没有(甚至完全没有)附着的地方。”
Code, meanwhile, should have less of a watermark than other text, because the model will need to create working code and won’t have the freedom to choose between a variety of equally valid options. “Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code,” Anthropic said. “But by definition, it will have a negligible effect on the actual code produced.”
与此同时,代码所含的水印应该比其他文本少,因为模型需要创建可运行的代码,而没有自由在多种同样有效的选项之间进行选择。Anthropic 表示:“话虽如此,在代码中某些词汇或术语可以任意选择的地方,水印是可以使用的,例如代码中的注释。但根据定义,它对实际生成的代码影响微乎其微。”
Anthropic also said that Claude won’t be the only AI chatbot to generate watermarked text, as “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”
Anthropic 还表示,Claude 不会是唯一生成带水印文本的 AI 聊天机器人,因为“其他主要模型开发商也签署了相同的行为准则,并将实施各自的水印。”