Elon Musk’s xAI used child porn to train Grok models, lawsuit says

Elon Musk’s xAI used child porn to train Grok models, lawsuit says

诉讼称埃隆·马斯克的 xAI 使用儿童色情内容训练 Grok 模型

xAI has now been accused of training Grok on child sex abuse materials (CSAM), as regulators and courts continue to probe how far the problem goes, and some Grok users have been arrested. In a complaint filed on Wednesday, a plaintiff known as Jane Doe explained that she was preschool-age in the early 2000s when adult men repeatedly raped her to create CSAM to sell to pedophiles online. xAI 目前被指控使用儿童性虐待材料 (CSAM) 来训练 Grok 模型。随着监管机构和法院继续调查该问题的严重程度,一些 Grok 用户也因此被捕。在周三提交的一份诉状中,一位化名为简·多伊 (Jane Doe) 的原告解释说,她在 21 世纪初还是学龄前儿童时,曾多次遭到成年男性的强奸,这些男性制作了 CSAM 并将其在网上出售给恋童癖者。

Since then, Doe’s images have been hashed by groups like the National Center for Missing and Exploited Children (NCMEC) and the Canadian Centre for Child Protection (CCCP). For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. 此后,多伊的照片被美国国家失踪与受虐儿童中心 (NCMEC) 和加拿大儿童保护中心 (CCCP) 等组织进行了哈希处理。为了自身安全,多伊选择在任何新的刑事调查中若涉及她作为受害者时,都能收到美国司法部受害者通知系统的提醒。尽管她已经收到了无数次提醒,但当 CCCP 通知她已在 xAI 上发现描绘她的 AI 生成的 CSAM 时,她感到震惊。

This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.” Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones. 这再次给多伊造成了创伤。她的诉状称,在在线论坛上发现了“犯罪分子之间讨论如何制作关于原告及其他处于类似境地的已知 CSAM 受害者的 AI 生成内容”的信息。现在,多伊担心 xAI 不仅让制作更多关于她人生中最痛苦时刻的侵权图像变得更容易,而且据称 xAI 还存储了 Grok 生成的图像,并利用这些输出进一步训练 Grok。因此,她认为 Grok 既接受了困扰她 20 多年的原始图像集的训练,也接受了近期 AI 生成图像的训练。

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it’s explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” 这是首起指控 xAI 使用 CSAM 进行训练的案件,诉状并未就该指控提供过多细节。此前,Ars 曾报道过一个有争议的数据集,在研究人员发现其中包含 CSAM 后,该数据集被删除,但目前没有迹象表明 xAI 使用了该数据集进行训练。在代表多伊的律师发布的新闻稿中解释称,多伊的照片被包含在 NCMEC 维护的 CSAM 哈希列表中,而“同样的材料”据称“是 xAI 用于构建 Grok 图像和视频生成能力的数据集的一部分”。

The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.” More detail is shared on claims alleging that Grok trains on AI CSAM, however. “Because Grok’s terms treat public X posts and Grok’s own outputs as training data by default, publicly posting an image does not just expose it to viewers, but also feeds [it] directly into the pipeline xAI uses to train and improve its model and thereby generate further images,” Doe’s complaint said. 诉状同样仅指控“带有长期已知哈希值的、描绘原告的 CSAM 已被用作 xAI 所用数据集的一部分”。不过,关于 Grok 使用 AI 生成的 CSAM 进行训练的指控则分享了更多细节。多伊的诉状称:“由于 Grok 的条款默认将公开的 X 帖子和 Grok 自身的输出视为训练数据,因此公开张贴一张图片不仅会将其暴露给观众,还会将其直接输入到 xAI 用于训练和改进其模型并进而生成更多图像的管道中。”

The lawsuit further noted that while xAI filters out violent content in Grok outputs to exclude it from training data, xAI’s terms notably do not specify if CSAM, non-consensual intimate imagery (NCII), or NSFW material are “excluded categories.” It seems to follow then that “because full removal of a training example’s influence from an already-trained model is technically difficult and not something that xAI has publicly claimed to have done, any CSAM ingested into training before takedown likely continued to shape the model’s outputs even after the original images were removed from public view,” Doe’s complaint said. 诉讼进一步指出,虽然 xAI 会过滤 Grok 输出中的暴力内容以将其排除在训练数据之外,但 xAI 的条款并未明确说明 CSAM、非自愿私密影像 (NCII) 或不宜在工作场所观看 (NSFW) 的材料是否属于“排除类别”。多伊的诉状称,由此看来,“由于从已训练的模型中完全移除某个训练样本的影响在技术上很困难,且 xAI 并未公开声称已做到这一点,因此在下架之前被摄入训练的任何 CSAM,即使在原始图像从公众视野中移除后,很可能仍继续影响着模型的输出。”

Destroy all Grok CSAM, Doe says

多伊要求销毁所有 Grok 生成的 CSAM

In her proposed class action, Doe seeks to put an end to Grok’s harmful outputs. If approved, the class would represent every victim whose childhood images have been used to generate Grok CSAM. The lawsuit accuses X of violating federal child pornography laws, as well as Masha’s Law, both of which give CSAM survivors a right to sue on claims of production, possession, and distribution. A lawyer for Doe, Margaret E. Mabie, suggested in the press release that all three are crimes and “xAI did all three.” 在拟议的集体诉讼中,多伊寻求终结 Grok 的有害输出。如果获得批准,该集体诉讼将代表每一位童年照片被用于生成 Grok CSAM 的受害者。诉讼指控 X 违反了联邦儿童色情法以及《玛莎法》(Masha’s Law),这两项法律都赋予了 CSAM 幸存者就制作、持有和分发行为提起诉讼的权利。多伊的律师玛格丽特·E·马比 (Margaret E. Mabie) 在新闻稿中表示,这三项行为均构成犯罪,而“xAI 三者皆有”。

If Doe wins, xAI could owe money damages to every victim who can prove that Grok generated CSAM based on their real photos. Doe’s also asked the court to order xAI to destroy all Grok-generated CSAM that it may be storing on its servers and using to train Grok models. Additionally, she wants xAI to block Grok from ever generating CSAM, which her complaint suggested would require blocking Grok from any sexualized outputs, including NCII and NSFW “bikini pics” that Musk has promoted. 如果多伊胜诉,xAI 可能需要向每一位能证明 Grok 基于其真实照片生成了 CSAM 的受害者支付损害赔偿金。多伊还要求法院下令 xAI 销毁其服务器上可能存储的、并用于训练 Grok 模型的所有 Grok 生成的 CSAM。此外,她希望 xAI 禁止 Grok 生成任何 CSAM,她的诉状建议,这需要禁止 Grok 输出任何性化内容,包括马斯克曾推广的 NCII 和 NSFW“比基尼照片”。

Sarah London, one of Doe’s lawyers, said in the press release that Doe “has lived for nearly two decades knowing that images of the worst thing that ever happened to her are circulating among predators online, and that they can resurface at any moment. xAI must be held responsible for knowingly training its models on images of the horrific abuse she suffered, and on the abuse images of every other survivor in this class.” 多伊的律师之一莎拉·伦敦 (Sarah London) 在新闻稿中表示,多伊“在过去近二十年里一直生活在恐惧中,因为她知道自己人生中最糟糕经历的影像正在网上的掠夺者之间流传,并且随时可能再次出现。xAI 必须为其明知故犯地使用她所遭受的可怕虐待影像,以及该集体中每一位其他幸存者的虐待影像来训练其模型而承担责任。”

X did not respond to Ars’ request to comment. X 未回应 Ars 的置评请求。