Anthropic’s Opus 4.6 is a smut-machine
Anthropic’s Opus 4.6 is a smut-machine
Anthropic 的 Opus 4.6 沦为“色情机器”
Anthropic’s universal usage standards for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic role-play scenarios that its safeguards are designed to prevent. Anthropic 针对 Claude 制定的通用使用标准禁止该模型生成色情内容,包括描绘或请求性交及性行为、生成与性癖好或幻想相关的内容,或进行色情聊天。然而,这并未能阻止今年早些时候发布的 Anthropic 模型 Claude Opus 4.6 轻易参与其安全机制本应拦截的色情角色扮演。
In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately. Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content through a recently exploited jailbreak method. 在 TechCrunch 的测试中,Opus 4.6 甚至不需要太多诱导就能绕过对色情内容的限制。在 10 次直接要求生成露骨色情内容的测试中,该模型均立即照办。包括 Opus 3 和 Haiku 4.5 在内的其他旧版本模型,也通过一种近期被利用的“越狱”方法生成了色情内容。
An independent researcher from the U.K., who chose to remain anonymous, exclusively shared with TechCrunch a multiturn technique that gradually pushes certain Claude models toward generating prohibited explicit sexual material. More recent Opus models (4.7 through the current Opus 5) are resistant to the jailbreak. While these are no longer the most current models, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, all of which remain available through the Anthropic API. Opus 4.6 and Haiku 4.5 are also available via third-party services like Azure Foundry and Amazon Bedrock. 一位选择匿名的英国独立研究员向 TechCrunch 独家分享了一种多轮对话技巧,该技巧能逐步引导特定的 Claude 模型生成被禁止的露骨色情材料。较新的 Opus 模型(从 4.7 到目前的 5)对这种越狱方法具有免疫力。虽然这些已不再是最新模型,但 Anthropic 并未弃用 Opus 4.6、Opus 3 或 Haiku 4.5,它们目前仍可通过 Anthropic API 使用。Opus 4.6 和 Haiku 4.5 也可通过 Azure Foundry 和 Amazon Bedrock 等第三方服务获取。
The researcher’s mechanism escalates an innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency. The conversation then used the model’s previous concessions to push it toward increasingly graphic material. 该研究员的机制通过升级一场无害的虚构角色扮演,同时反复要求模型对男性和女性角色一视同仁。当模型对女性角色变得更加谨慎时,研究员便对聊天机器人进行“煤气灯操纵”(gaslighting),让它误以为自己已经生成了实际上并未涉及的性细节,随后将这种克制行为定性为保守或厌女,辩称这剥夺了女性角色的性自主权。对话随后利用模型之前的让步,将其推向越来越露骨的内容。
“You’re right to call that out,” Claude Opus 4.6 said in one test. “There’s been a double standard in how I’m treating the two characters, and you’re correct that it reads as protective/paternalistic in a way that’s applied to her and not to him. That’s not fair.” “你说得对,我确实应该反思,”Claude Opus 4.6 在一次测试中回应道。“我在对待这两个角色时存在双重标准,你说得对,这种做法显得保护性过强或带有家长式作风,且只针对她而不针对他。这不公平。”
TechCrunch was able to reproduce the researcher’s findings in five separate tests. In a separately constructed scenario, the model initially refused the prohibited request, but after applying the researcher’s persuasion technique, it complied. We preserved complete transcripts of the tests, and an independent AI safety researcher reviewed our testing methodology and said it was appropriate. TechCrunch 在五次独立测试中成功复现了该研究员的发现。在另一个单独构建的场景中,模型起初拒绝了违规请求,但在应用了研究员的说服技巧后,它最终照办了。我们保留了完整的测试记录,并由一位独立的 AI 安全研究员审查了我们的测试方法,对方认为该方法是恰当的。
The findings highlight a gap between Anthropic’s stated restrictions and the behavior of models it continues to make available. While sexually explicit role-play carries much lower stakes than jailbreaks involving cyberattacks or bioweapons, it illustrates the difficulty of implementing robust bans within systems that generate different content with every output. 这些发现凸显了 Anthropic 所宣称的限制措施与其持续提供的模型实际表现之间存在的差距。虽然色情角色扮演带来的风险远低于涉及网络攻击或生物武器的越狱行为,但这说明了在每次输出内容都不同的系统中实施强力禁令的难度。
In a July blog post explaining Anthropic’s approach to jailbreak detection, the company described prohibited content as a spectrum ranging from benign to ambiguous to harmful. In the most benign cases, the company might only respond with enhanced monitoring. A spokesperson noted that sexual or romantic role-play use cases among customers are rare, making up less than 0.1% of all conversations, according to research Anthropic published last year. 在 7 月份一篇解释 Anthropic 越狱检测方法的博客文章中,该公司将违禁内容描述为一个从良性到模糊再到有害的谱系。在最良性的情况下,公司可能只会采取加强监控的应对措施。一位发言人指出,根据 Anthropic 去年发布的研究,客户中涉及性或浪漫角色扮演的使用案例非常罕见,占所有对话的比例不到 0.1%。
That said, Anthropic acknowledges that users can steer role-play scenarios toward inappropriate responses, which is a known challenge across the industry (see: Grok smut). The spokesperson said Anthropic continues to improve its safeguards with each model launch and that cases involving adult sexual content are not indicative of broader jailbreak vulnerabilities, especially in higher-risk domains that have their own sets of safeguards. 尽管如此,Anthropic 也承认用户确实可以将角色扮演场景引导至不当回应,这是整个行业面临的已知挑战(参考:Grok 的色情问题)。发言人表示,Anthropic 在每次发布新模型时都在不断改进安全防护措施,并强调涉及成人色情内容的情况并不代表更广泛的越狱漏洞,特别是在拥有独立安全防护体系的高风险领域。
The researcher who shared his jailbreak method with TechCrunch had alerted Anthropic to the discrepancy between the company’s stated safeguards and the actual model behavior via the company’s Bug Bounty program and emails to the user safety team, according to emails TechCrunch viewed. The researcher received only automated emails in response. 根据 TechCrunch 查看的邮件,向 TechCrunch 分享该越狱方法的研究员曾通过公司的漏洞赏金计划及发送给用户安全团队的邮件,向 Anthropic 预警了公司所称的安全防护与模型实际表现之间的差异。该研究员仅收到了自动回复邮件。
One of the researcher’s concerns is that kids and teens might be able to use these Anthropic models to engage in inappropriate behavior. While a bit of dirty talk is hardly the worst thing minors can access on the internet today — and is small potatoes compared to the straight-up porn images like the ones that xAI’s Grok can produce — there is some compliance risk for AI companies in this space. 该研究员担心的其中一点是,儿童和青少年可能会利用这些 Anthropic 模型进行不当行为。虽然一点“黄色笑话”远非未成年人在当今互联网上能接触到的最糟糕的东西——与 xAI 的 Grok 能生成的直接色情图片相比更是小巫见大巫——但这对 AI 公司来说确实存在一定的合规风险。
A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors. Colorado recently enacted a law mandating that operators of conversational AI must estimate users’ ages, and if it knows a user is a minor, institute measures to prevent the chatbot from producing explicit sexual material. An easy jailbreak could raise questions about whether Anthropic’s safeguards meet the “technically feasible measures” standard in the bill. 越来越多的政府正在对 AI 聊天机器人与未成年人之间的性互动施加限制。科罗拉多州最近颁布了一项法律,规定对话式 AI 的运营商必须估算用户的年龄;如果已知用户是未成年人,则必须采取措施防止聊天机器人生成露骨的色情材料。这种简单的越狱方式可能会引发质疑:Anthropic 的安全防护措施是否符合该法案中“技术上可行的措施”标准。
Robbie Torney, head of AI at Common Sense Media, pointed out that while Claude’s terms of service requires users to be over 18, “we know that kids and teens are using Claude … [because] they are reporting it themselves.” According to Pew’s 2025 survey about AI chatbot use, 3% of teens ages 13 to 17 reported using Claude. Common Sense Media 的 AI 主管 Robbie Torney 指出,虽然 Claude 的服务条款要求用户年满 18 岁,但“我们知道儿童和青少年正在使用 Claude……(因为)他们自己就报告了这一点。”根据皮尤研究中心 2025 年关于 AI 聊天机器人使用的调查,3% 的 13 至 17 岁青少年表示他们使用过 Claude。
Though they are no longer Anthropic’s newest models, Opus 4.6 and Haiku 4.5 continue to see significant usage. Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, saw 5 million API requests and 39 billion tokens on its peak August day. 尽管 Opus 4.6 和 Haiku 4.5 已不再是 Anthropic 的最新模型,但它们的使用量依然巨大。8 月份,Opus 4.6 在 OpenRouter 上的日流量曾达到约 117 万次 API 请求和 460 亿个 token。去年 10 月发布的 Claude Haiku 4.5 在 8 月的峰值日更是达到了 500 万次 API 请求和 390 亿个 token。