He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

他曾为 AI 抓取了他们所有的艺术作品,如今却在协助开发保护艺术家的工具

Since early 2023, photographer Jingna Zhang and a small crew of volunteers have worked tirelessly to maintain an image-sharing social media and portfolio app called Cara. So far, it has attracted about 1.5 million artists. What drew them to the platform? A shared opposition to the unauthorized use of their work to train AI models and a desire to publicize their art while avoiding exploitation by Big Tech.

自 2023 年初以来,摄影师张晶娜(Jingna Zhang)和一小群志愿者一直在不懈地维护一个名为 Cara 的图片分享社交媒体及作品集应用。到目前为止,该平台已吸引了约 150 万名艺术家。是什么吸引了他们?是对未经授权使用其作品训练 AI 模型的一致反对,以及在推广艺术作品的同时避免被大型科技公司剥削的愿望。

But while Cara filters out AI images and offers some minimal protective features, preventing scrapes themselves is nearly impossible. And just this month, beginning on August 13, Cara was subjected to three major scrapes, which spiked its server fees and alarmed creators who had migrated there from platforms like Instagram, where all content is explicitly available to Meta as training data.

尽管 Cara 过滤了 AI 生成的图像并提供了一些基本的保护功能,但要完全防止数据抓取几乎是不可能的。就在本月,从 8 月 13 日开始,Cara 遭遇了三次大规模的数据抓取,这导致其服务器费用激增,并引起了从 Instagram 等平台迁移过来的创作者们的恐慌——在那些平台上,所有内容都被明确规定可供 Meta 用作训练数据。

The first of these incidents came to light when the individual responsible took to the subreddit r/DefendingAIArt to announce he had obtained a 12-terabyte archive of 12 million works from Cara—more or less its entire library of publicly available images. “It was a fun project,” he wrote under the handle MandarinDawnPoppy994 in his since-deleted post, saying the process cost him less than $10.

第一起事件曝光时,责任人前往 Reddit 的 r/DefendingAIArt 版块宣布,他已从 Cara 获取了一个 12TB 的档案,包含 1200 万件作品——这几乎涵盖了该平台所有公开的图像库。“这是一个有趣的项目,”他在后来被删除的帖子中以 MandarinDawnPoppy994 的网名写道,并称整个过程花费不到 10 美元。

“We actually found out about it through our users tagging us,” Zhang tells WIRED, since the scraper was “gloating and looking for other people to join him to do something with the dataset on Reddit,” sparking a fierce debate across AI-related forums about the ethics of what he had done. “I just feel it’s targeted and very hurtful,” she adds, noting that “laws are not caught up on” protections against such data harvests, meaning that scrapers can often justify it as technically legal. (Zhang is separately part of two ongoing class actions brought by visual artists, one against Stability AI, Midjourney, and others and the second against Google, alleging that the companies’ image generator tools were trained on their copyrighted work.)

“我们实际上是通过用户标记才发现这件事的,”张晶娜告诉《连线》(WIRED),因为抓取者当时“正在沾沾自喜,并试图在 Reddit 上寻找其他人加入,一起利用这个数据集做点什么”,这在各大 AI 相关论坛引发了关于其行为伦理的激烈辩论。“我只觉得这是针对性的,而且非常伤人,”她补充道,并指出“法律尚未跟上”针对此类数据采集的保护措施,这意味着抓取者往往能将其辩解为技术上的合法行为。(张晶娜目前还参与了两起由视觉艺术家发起的集体诉讼,一起针对 Stability AI、Midjourney 等公司,另一起针对谷歌,指控这些公司的图像生成工具使用了受版权保护的作品进行训练。)

In a surprising turn of events, however, the person who grabbed all the art off Cara would turn out to regret his stunt and agree to collaborate with Zhang on a new open-source tool to protect artists.

然而,令人惊讶的是,这位从 Cara 抓取了所有艺术作品的人最终对自己的行为感到后悔,并同意与张晶娜合作开发一种新的开源工具来保护艺术家。

In the meantime, unfortunately, other scrapers continued to take advantage of Cara’s vulnerabilities and minimal resources. While a number of AI proponents objected to going after Cara, a few were apparently emboldened by MandarinDawnPoppy994 to carry out what Zhang sees as “copycat” attacks.

与此同时,不幸的是,其他抓取者继续利用 Cara 的漏洞和有限的资源。虽然许多 AI 支持者反对攻击 Cara,但显然有一些人受到 MandarinDawnPoppy994 的鼓舞,实施了张晶娜所认为的“模仿”攻击。

A second scraper pulled about 8.5 million links from Cara, as well as metadata like usernames, titles, and tags, and uploaded these to Hugging Face, the AI developer platform. After Hugging Face was bombarded with takedown requests, it responded in a statement that while it would issue a notice to the user, “CaptiveDreamer,” to remove the personal metadata, it could not do the same for the URLs, since “no copies of the artworks are hosted here,” and the links “point to the copies the artists published on Cara.” The company concluded that “further copyright reports on the same basis will not change this outcome.”

第二名抓取者从 Cara 提取了约 850 万个链接,以及用户名、标题和标签等元数据,并将这些内容上传到了 AI 开发平台 Hugging Face。在 Hugging Face 遭到大量下架请求轰炸后,该公司发表声明称,虽然会向用户“CaptiveDreamer”发出通知要求删除个人元数据,但无法对 URL 执行相同操作,因为“这里没有托管任何艺术作品的副本”,且这些链接“指向的是艺术家在 Cara 上发布的副本”。该公司总结称,“基于相同理由的进一步版权投诉将不会改变这一结果。”

Finally, on August 22, a third scraper obtained 123,000 images from Cara, along with text posts and user bios that included personal information, sharing it all on a site called Academic Torrents. Zhang then launched a GoFundMe for legal fees, setting a goal of $120,000, explaining that the money would go toward exploring any and all strategies of defending Cara through cyber and copyright laws. As of Thursday, she has raised more than $100,000, and she says Cara is actively looking for any additional legal assistance.

最后,在 8 月 22 日,第三名抓取者从 Cara 获取了 12.3 万张图片,以及包含个人信息的文字帖子和用户简介,并将所有内容分享到了一个名为 Academic Torrents 的网站上。随后,张晶娜发起了 GoFundMe 筹款以支付法律费用,目标金额为 12 万美元,并解释说这笔钱将用于探索通过网络法和版权法保护 Cara 的所有策略。截至周四,她已筹集了超过 10 万美元,并表示 Cara 正在积极寻求任何额外的法律援助。

Zhang is frustrated not only by the scraping but also by the confusion around what she and the Cara team can realistically do to shield artists from malicious actors, saying that some users have already deleted their portfolios and abandoned the site. “We have done the right things within limits without making it horrible to use,” Zhang says of the app’s current safeguards, including some new temporary measures like login gates—which, she adds, aren’t really a solution to an ongoing, internet-wide problem.

张晶娜不仅对数据抓取感到沮丧,还对她和 Cara 团队在保护艺术家免受恶意行为者侵害方面能做些什么感到困惑,她说一些用户已经删除了他们的作品集并放弃了该网站。“我们在不影响使用体验的前提下,在能力范围内做了正确的事,”张晶娜在谈到该应用目前的保障措施时说,包括一些新的临时措施,如登录门槛——但她补充说,这并不能真正解决一个持续存在的、全互联网范围的问题。

And Zhang worries that people blaming her for the string of attacks may not understand that Cara has almost certainly been scraped before, like any other site; it cannot guarantee complete security. “If it makes them feel better, deleting your work and leaving Cara, I support that,” Zhang says of those leaving. “But I don’t want to give people the misconception that if they go somewhere else, they are safer, because they’re not. Bigger platforms get scraped more, so that makes me feel worse.”

张晶娜还担心,那些指责她导致这一系列攻击的人可能不明白,Cara 和其他任何网站一样,几乎肯定以前就被抓取过;它无法保证绝对的安全。“如果删除作品并离开 Cara 能让他们感觉好受些,我支持,”张晶娜谈到那些离开的人时说。“但我不想让人们产生一种误解,认为去别的地方就更安全了,因为事实并非如此。更大的平台被抓取得更多,这反而让我感觉更糟。”

Yet Zhang, who is quick to remind WIRED that she is not a tech founder except by accident, has a newfound ally in this battle: the person who kicked off this month’s scraping frenzy. After a Cara user confronted him and eventually persuaded to delete his dataset, she got in touch to understand his motivations and confirm the deletion. “He felt very bad to see how hurt people were,” Zhang says. “So he decided to help us.”

然而,张晶娜——她很快提醒《连线》说她只是偶然成为科技创业者——在这场战斗中获得了一位新盟友:那个引发本月抓取狂潮的人。在一名 Cara 用户与他交涉并最终说服他删除数据集后,她与他取得了联系,以了解他的动机并确认删除情况。“看到人们受到如此大的伤害,他感到非常难过,”张晶娜说。“所以他决定帮助我们。”

“Heft” is a student in North America with a background in software and an interest in digital preservation and archival projects. (He requested that we refer to him only by one of his screen names due to doxing and death threats he says he received over the Cara situation.) In a conversation over Discord, Heft tells WIRED that scraping Cara was originally nothing more than a technical project and that he had no intention of making the data public.

“Heft”是一名在北美的学生,拥有软件背景,对数字保存和档案项目感兴趣。(由于他称自己在 Cara 事件后收到了人肉搜索和死亡威胁,他要求我们仅以他的一个网名来称呼他。)在 Discord 的一次对话中,Heft 告诉《连线》,抓取 Cara 最初仅仅是一个技术项目,他并没有打算将数据公开。

Heft says he nevertheless “made a foolish decision to attempt to ragebait with the dataset on Reddit” and “was carried away by trolling in the comments.” He knew it would provoke artists on Cara but did not anticipate the sheer anguish in that community. He saw people “sharing how they were having panic attacks over the scrape, how they deleted their entire portfolios from the internet.” In direct conversations with artists, he gained a greater appreciation of how personal their work was to them and how much…

Heft 表示,尽管如此,他还是“做出了一个愚蠢的决定,试图在 Reddit 上利用该数据集来激怒他人”,并且“被评论区的挑衅冲昏了头脑”。他知道这会激怒 Cara 上的艺术家,但没想到该社区会产生如此巨大的痛苦。他看到人们“分享他们因为这次抓取而产生恐慌发作,以及他们如何从互联网上删除了整个作品集”。在与艺术家的直接对话中,他更深刻地体会到他们的作品对他们个人而言有多重要,以及……