OpenAI will watermark ChatGPT outputs by default—but only in the EU
OpenAI will watermark ChatGPT outputs by default—but only in the EU
OpenAI 将默认对 ChatGPT 输出内容添加水印——但仅限于欧盟地区
OpenAI has announced that it will begin automatically watermarking text generated with ChatGPT in the European Union. It will also offer the watermarking feature in other regions, but it will be off by default outside of the EU. OpenAI 宣布,将开始自动为欧盟地区 ChatGPT 生成的文本添加水印。该功能也将在其他地区提供,但在欧盟以外的地区,该功能默认处于关闭状态。
The move in Europe is driven by a need to comply with the EU AI Act, which took effect in August. It requires marking content produced by AI models in a way that another tool can detect. 此举旨在遵守 8 月生效的《欧盟人工智能法案》(EU AI Act)。该法案要求以可被其他工具检测到的方式标记由人工智能模型生成的内容。
Unfortunately, there is still no completely effective and reliable way to do that. A few standards already exist, like SynthID and the C2PA project, but they are relatively easy to circumvent for anyone with basic know-how. The same is likely true for OpenAI’s watermark. 遗憾的是,目前尚无完全有效且可靠的方法来实现这一点。虽然已经存在一些标准(如 SynthID 和 C2PA 项目),但对于任何具备基本技术知识的人来说,这些标准都相对容易规避。OpenAI 的水印很可能也面临同样的问题。
Its method is proprietary; the company calls it textGrain, and has published a technical paper explaining how it works. But in general, it works like other LLM watermarking tools we’ve seen in the past: It puts patterns in the word choices that are not clear to a human reader, and that don’t meaningfully change the general quality of the output, but that someone with a key can use a specialized detector to find. 其方法是专有的;该公司将其称为“textGrain”,并发布了一份技术论文解释其工作原理。但总的来说,它的工作方式与其他我们过去见过的 LLM 水印工具类似:它在词汇选择中植入人类读者无法察觉的模式,这些模式不会显著改变输出内容的整体质量,但持有密钥的人可以使用专门的检测器将其识别出来。
OpenAI says it will be giving access to the detector to a limited number of researchers and organizations, and providing a request-for-approval process for others to be added over time. OpenAI 表示,将向少数研究人员和组织开放该检测器的访问权限,并为其他申请者提供审批流程,以便日后逐步增加访问权限。
While the best case OpenAI’s tests of textGrain show a respectable but not entirely reliable 92 percent successful detection rate, the company’s tests also show that changing just 10 percent of the text in an output reduces the successful detection rate by almost 30 percent, and changing 20 percent of the text can lower the success rate by almost 75 percent. (Also worth noting that in general, the success rate is lower for shorter or translated text than for longer text.) 虽然 OpenAI 对 textGrain 的测试显示,在最佳情况下,其检测成功率达到了可观但并非完全可靠的 92%,但该公司的测试也表明,仅修改输出文本的 10%,成功率就会下降近 30%;而修改 20% 的文本则可能使成功率降低近 75%。(此外值得注意的是,通常情况下,较短文本或翻译文本的检测成功率低于较长文本。)
All this is roughly in line with what we’ve seen with other, similar watermarking solutions in the past. In August, OpenAI competitor Anthropic also introduced watermarking for text generated by its models—but Anthropic has enabled it globally, while OpenAI is currently only making it the default where regulators require it—at least in the ChatGPT and Codex apps. It will remain off but optionally available in the API, it seems. 这一切与我们过去看到的其他类似水印解决方案的情况大致相符。今年 8 月,OpenAI 的竞争对手 Anthropic 也为其模型生成的文本引入了水印功能,但 Anthropic 是在全球范围内启用的,而 OpenAI 目前仅在监管机构有要求的地区将其设为默认(至少在 ChatGPT 和 Codex 应用中是这样)。在 API 中,该功能似乎将保持关闭,但用户可选择开启。
The new watermarking will roll out to users in the EU “in the coming weeks,” the company says. 该公司表示,这项新的水印功能将在“未来几周内”向欧盟用户推出。