Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks
Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks
程序员称已找到绕过 Claude 隐形水印的方法
Within four hours of Anthropic confirming that Claude models would globally embed invisible, machine-readable watermarks into any AI-generated content, developer Guillaume Meyer had published his override. 在 Anthropic 确认 Claude 模型将在全球范围内为所有 AI 生成内容嵌入机器可读的隐形水印后,仅过了四个小时,开发者 Guillaume Meyer 就发布了他的破解方案。
His code to remove watermarks from Claude-generated text has since gone viral on GitHub, has been bookmarked more than 20,000 times on X, and has drawn more than 100 contributors, with many more incorporating the technology into their own projects. “Anthropic is embedding watermarks in its Claude texts … the issue is practically history just one day later,” wrote one AI specialist, accompanied by an image of Meyer breaking out of chains and standing on crumpled EU and Anthropic flags. 他编写的用于移除 Claude 生成文本水印的代码随后在 GitHub 上走红,在 X(原推特)上被收藏超过 2 万次,并吸引了 100 多名贡献者,还有更多人将该技术整合到自己的项目中。一位 AI 专家写道:“Anthropic 正在其 Claude 文本中嵌入水印……仅仅一天后,这个问题实际上就已经成为历史了。”并配上了一张 Meyer 挣脱锁链、站在揉皱的欧盟和 Anthropic 旗帜上的图片。
Meyer and others started investigating how watermarking works after Anthropic announced last week that Claude would adopt it in order to comply with the European Union’s AI Act. 在 Anthropic 上周宣布 Claude 将采用水印技术以遵守欧盟《人工智能法案》(AI Act)后,Meyer 和其他人开始研究水印的工作原理。
Some are trying to evade the watermarking because they disagree with the idea that all AI-generated content should be labeled as such, Meyer told WIRED, while others, including himself, say they simply relish the technical challenge. Freelance content writers and social media creators have also contacted Meyer asking for assistance using the code, he says. Meyer 告诉《连线》(WIRED)杂志,有些人试图规避水印是因为他们不同意“所有 AI 生成的内容都应被标记”这一观点;而包括他自己在内的另一些人则表示,他们只是单纯享受这种技术挑战。他说,自由职业内容创作者和社交媒体博主也联系了 Meyer,寻求使用该代码的帮助。
The new rules, which came in earlier this month, stipulate that model providers like Anthropic and OpenAI must label synthetic audio, image, video, or text so that this material can be detected by a machine as AI-generated—or face fines of up to 3 percent of annual turnover. While the rules say providers cannot market circumvention tools, there is no legal restriction on independent tools. 本月初生效的新规规定,Anthropic 和 OpenAI 等模型提供商必须对合成音频、图像、视频或文本进行标记,以便机器能够识别这些材料是由 AI 生成的,否则将面临最高相当于年营业额 3% 的罚款。虽然规定指出提供商不得销售规避工具,但对独立工具并无法律限制。
“I’m not against transparency, and I’m all for content attribution,” says Meyer. “I just think watermarking in itself is a really bad solution, because it has major drawbacks and risks.” He is concerned about the risk of false positives and that the watermarking might not distinguish between light or heavy AI use, especially since, as a native French speaker, he often uses Claude and other AI tools like Grammarly to edit his writing. Using the watermark as evidence–when even Anthropic admits it can only generate a probability that the text has been touched by Claude–could lead to employers unfairly rejecting candidates or overblown accusations of researchers using artificial intelligence just because the detector flags it, he says. “我不反对透明度,也完全支持内容归属,”Meyer 说,“我只是认为水印本身是一个非常糟糕的解决方案,因为它有重大的缺陷和风险。”他担心误报风险,以及水印可能无法区分 AI 的轻度或重度使用。特别是作为母语为法语的人,他经常使用 Claude 和 Grammarly 等 AI 工具来编辑文章。他说,当连 Anthropic 都承认水印只能生成文本被 Claude 处理过的“概率”时,若将水印作为证据,可能会导致雇主不公平地拒绝候选人,或者仅仅因为检测器标记了某项研究,就对研究人员进行过度的指责。
Anthropic watermarks text invisibly by leaving a pattern in Claude’s choice of words and phrases that is indiscernible to a human reader but would be detectable by a machine that knows how to look for it. Because this influences Claude’s output, some users are concerned this will degrade the quality of Claude’s responses, though Anthropic insists this won’t be the case. The technique, called SynthID, was developed by Google, which has been using it to watermark its AI-generated content since 2023. Computer scientist Scott Aaronson proposed a similar method when working at OpenAI but says the firm never deployed it because the company was worried that watermarks would put customers off its product. Anthropic 通过在 Claude 的词汇和短语选择中留下一种人类读者无法察觉、但懂得如何查找的机器可以检测到的模式,来对文本进行隐形水印处理。由于这会影响 Claude 的输出,一些用户担心这会降低 Claude 的回复质量,尽管 Anthropic 坚称不会出现这种情况。这项名为 SynthID 的技术由谷歌开发,自 2023 年以来一直用于为其 AI 生成的内容添加水印。计算机科学家 Scott Aaronson 在 OpenAI 工作时曾提出过类似的方法,但他表示该公司从未部署过,因为担心水印会使客户对产品望而却步。
Meyer’s removal method uses a non-watermarking large language model to generate multiple rewrites, swapping in synonyms and slightly reorganizing content. Of course, this relies on using other large language models which do not insert watermarks—possibly not a safe bet since 190 organizations—providers OpenAI, Microsoft, and Meta among them—have signed the EU’s transparency code of practice. It remains to be seen how many of these laboratories are going to implement their watermarks, which must be included in all new models released from August and must be integrated into existing models by December. Meyer 的移除方法是使用一个不加水印的大型语言模型来生成多个重写版本,替换同义词并稍微重组内容。当然,这依赖于使用其他不插入水印的大型语言模型——这可能并非万全之策,因为包括 OpenAI、微软和 Meta 在内的 190 家机构已经签署了欧盟的透明度行为准则。目前尚不清楚这些实验室中有多少会实施水印,根据规定,8 月起发布的所有新模型都必须包含水印,且必须在 12 月前集成到现有模型中。
While there’s no certainty this tool works until Anthropic releases the software it uses to detect a watermark, understanding the basic SynthID-text approach underpinning Claude’s watermarking makes them fairly sure the method works, says Wayne Pan, chief technology and cofounder at Silicon Valley–based sovereign AI startup Haimaker. He incorporated Meyer’s open-source tool into his platform because he similarly disliked the idea of Claude watermarking content even when it’s only been lightly edited and disagreed with the watermark being invisible to the user. 硅谷主权 AI 初创公司 Haimaker 的首席技术官兼联合创始人 Wayne Pan 表示,虽然在 Anthropic 发布用于检测水印的软件之前,无法确定该工具是否有效,但了解 Claude 水印背后的 SynthID-text 基本方法,让他们相当确信该方法是有效的。他将 Meyer 的开源工具整合到了自己的平台中,因为他同样不喜欢 Claude 在内容仅经过轻微编辑时就添加水印的做法,也不赞同水印对用户不可见。
Other coders have developed their own removal tools: Software engineer Erik Hughes took 15 minutes to knock up a tool with Claude that removes invisible and look-alike characters, reorders sentences within paragraphs, and swaps several words for synonyms. Leon Chlon, a Visiting Fellow at the University of Oxford, says the watermarks can be removed by condensing Claude’s response, translating it into a dialect like Arabic, which has very different semantics compared to English, and then translating it back. Anthropic itself acknowledged that heavily edited, paraphrased, or translated content might not carry a watermark. 其他程序员也开发了自己的移除工具:软件工程师 Erik Hughes 用了 15 分钟,利用 Claude 开发了一个工具,可以删除隐形字符和相似字符,重新排列段落中的句子,并将几个词替换为同义词。牛津大学访问学者 Leon Chlon 表示,可以通过压缩 Claude 的回复、将其翻译成与英语语义差异巨大的方言(如阿拉伯语),然后再翻译回来来移除水印。Anthropic 本身也承认,经过大量编辑、改写或翻译的内容可能不会带有水印。
In a statement to WIRED, a spokesperson for Anthropic said: “We’re adding marking to Claude’s output to comply with the EU AI Act, and other labs are taking similar steps. It’s hard to identify AI-generated text, and this gives people better tools for identification. Text from supported Claude models, including output from Claude Code, will carry an invisible watermark, and it doesn’t change the meaning, quality, or readability of Claude’s responses. We also plan to ship a text-detection API so users can do more of this themselves.” Anthropic 的发言人在给《连线》的一份声明中表示:“我们正在 Claude 的输出中添加标记以遵守欧盟《人工智能法案》,其他实验室也在采取类似措施。识别 AI 生成的文本很困难,这为人们提供了更好的识别工具。来自受支持的 Claude 模型的文本(包括 Claude Code 的输出)将带有隐形水印,这不会改变 Claude 回复的含义、质量或可读性。我们还计划发布一个文本检测 API,以便用户可以更多地自行进行此类操作。”
Anthropic says it’s working out how to implement watermark detection for text and plans to release a tool to do so soon—at which point developers will finally be able to see whether their methods are foolproof. It’s also continuing to work on improving the watermarking system. “I think they wanted to show that they’re in good faith doing it,” says Pan, “but I don’t think you can ever have a watermark that will withstand everything.” Anthropic 表示,正在研究如何实现文本水印检测,并计划很快发布相关工具——届时开发者将最终能够验证他们的方法是否万无一失。该公司也在继续致力于改进水印系统。“我认为他们想表明自己是在真诚地履行义务,”Pan 说,“但我认为,不存在一种能够抵御一切的水印。”