Amazon Can Use Your Twitch Content to Train Its AI—Unless You Opt Out

Amazon Can Use Your Twitch Content to Train Its AI—Unless You Opt Out

亚马逊可以使用你的 Twitch 内容来训练其 AI——除非你选择退出

Twitch has updated its account settings to let streamers opt out of having their content used to help train the artificial intelligence models of Twitch’s parent company, Amazon. Twitch 更新了其账户设置,允许主播选择退出,不再将其内容用于训练其母公司亚马逊(Amazon)的人工智能模型。

Although the move has reassured some creators, it remains unclear exactly when the streaming platform began using the posts, streams, and videos of its users to train Amazon’s systems. The revelation is raising new concerns about how big tech companies handle their users’ data. 尽管此举安抚了一些创作者,但目前尚不清楚该直播平台究竟是从何时开始使用用户的帖子、直播和视频来训练亚马逊系统的。这一披露引发了人们对大型科技公司如何处理用户数据的新担忧。

The opt-out process is simple. From the Twitch website or mobile app, click on your account avatar, select Settings, and then go to the Security and Privacy section in the menu. There, you’ll find the Generative AI Training option, where you can disable the use of your content for that purpose. 退出流程很简单。在 Twitch 网站或移动应用上,点击你的账户头像,选择“设置”(Settings),然后进入菜单中的“安全与隐私”(Security and Privacy)部分。在那里,你会找到“生成式 AI 训练”(Generative AI Training)选项,你可以从中禁用你的内容被用于该目的。

In the language next to the toggle, Twitch notes that “disabling this option does not prevent Twitch and Amazon from using your channel’s content for other purposes described in Twitch’s Privacy Notice.” These include AI-powered platform features designed to facilitate streamers’ growth and monetization, such as real-time assistance for sponsorship campaigns, viewer discovery through recommendations, and community safety via tools like AutoMod. 在开关旁边的说明中,Twitch 指出:“禁用此选项并不能阻止 Twitch 和亚马逊将你的频道内容用于 Twitch 隐私声明中所述的其他目的。”这些目的包括旨在促进主播增长和变现的 AI 驱动平台功能,例如赞助活动的实时协助、通过推荐进行观众发现,以及通过 AutoMod 等工具维护社区安全。

The new setting aims to recognize creators’ right to decide how the content they produce is used. However, its launch has also raised questions about how Twitch and Amazon have used that content up to now. 这项新设置旨在尊重创作者决定其所生产内容如何被使用的权利。然而,它的推出也引发了关于 Twitch 和亚马逊迄今为止是如何使用这些内容的质疑。

In a forum dedicated to the topic, more than 16,000 creators expressed opposition to having their content being used by default to train Amazon’s AI systems—a practice that only came to light following the update to the settings. 在一个专门讨论该话题的论坛上,超过 16,000 名创作者表达了反对意见,抗议其内容在默认情况下被用于训练亚马逊的 AI 系统——这种做法直到设置更新后才为人所知。

The backlash arose after a livestream in which Mary Kish, Twitch’s head of community, explained the introduction of the changes. The executive acknowledged that the change would provoke a negative reaction. Mike Minton, Twitch’s head of product, also noted—in what he described as “a candid response”—that keeping the option to use content for AI training enabled by default was necessary, since otherwise “no one would participate” in the process. 此次反弹源于 Twitch 社区负责人 Mary Kish 在一次直播中解释了这些变更的引入。这位高管承认,这一改变会引发负面反应。Twitch 产品负责人 Mike Minton 也指出——他称之为“坦诚的回应”——将“用于 AI 训练的内容使用”选项设为默认开启是必要的,因为否则“没有人会参与”这个过程。

Minton pointed out that these mechanisms are not unique to Twitch and considered it likely that other companies developing AI systems are also extracting content from Twitch and other services to use in training their models. “I don’t know for sure,” he said, “but I think it’s quite reasonable to assume that almost any publicly available content is used to train models in one way or another, with or without permission. So I think we also need to acknowledge that there’s a lot here that’s beyond even our direct control.” Minton 指出,这些机制并非 Twitch 所独有,并认为其他开发 AI 系统的公司很可能也在从 Twitch 和其他服务中提取内容来训练其模型。“我不确定,”他说,“但我认为可以合理地推测,几乎任何公开可用的内容都在以某种方式被用于训练模型,无论是否获得许可。因此,我认为我们也需要承认,这里有很多事情甚至超出了我们的直接控制范围。”

These statements raised new questions: Since when has Twitch content been used to train AI models? Is Amazon the only company using this data, or are its business partners also involved? To what extent and in what ways is the authorship of content published by creators respected? 这些声明引发了新的问题:Twitch 的内容从何时开始被用于训练 AI 模型?亚马逊是唯一使用这些数据的公司,还是其商业合作伙伴也参与其中?创作者发布内容的署名权在多大程度上以及以何种方式得到了尊重?

Twitch’s Terms of Service, in effect since March 2024, have stipulated that users grant Twitch and its sublicensees the right to use, reproduce, modify, adapt, distribute, and create derivative works from their content. However, until now, they have not explicitly stated that such materials could be used to train generative AI models. 自 2024 年 3 月起生效的 Twitch 服务条款规定,用户授予 Twitch 及其分许可方使用、复制、修改、改编、分发其内容并据此创作衍生作品的权利。然而,直到现在,条款中并未明确说明这些材料可被用于训练生成式 AI 模型。

The Training Data Problem

训练数据难题

This case highlights one of the major challenges facing AI system developers: the growing shortage of high-quality data for training models. 此案例凸显了 AI 系统开发者面临的主要挑战之一:用于训练模型的高质量数据日益短缺。

Although companies like OpenAI have reached agreements with various publishers (including WIRED’s corporate parent, Condé Nast) to use some of their content for this purpose, available data sources are dwindling as technology advances and the demand for data increases. As a result, various alternatives have emerged that are reigniting the debate over the ethics and transparency of training processes. 尽管 OpenAI 等公司已与多家出版商(包括《连线》杂志的母公司康泰纳仕)达成协议,将部分内容用于此目的,但随着技术进步和数据需求增加,可用的数据源正在减少。因此,各种替代方案应运而生,重新点燃了关于训练过程伦理和透明度的争论。

For some time now, Meta has been using posts and images that users share on Facebook and Instagram to train its AI models. Recently, it also came to light that the company was using its employees’ activity and screenshots captured by users of its smart glasses for the same purpose. Similar cases have been documented involving Google and YouTube. 一段时间以来,Meta 一直在使用用户在 Facebook 和 Instagram 上分享的帖子和图片来训练其 AI 模型。最近还有消息披露,该公司正在使用其员工的活动记录以及智能眼镜用户拍摄的截图来实现同样的目的。谷歌和 YouTube 也被记录有类似的案例。

Thus, high-quality training data has become another strategic resource—along with chips, memory, and electricity—that the leading developers of artificial intelligence models have found running short as they struggle to sustain the pace of innovation demanded by the market. All signs point to a portion of that demand being met with information generated by users themselves, in exchange for free access to everyday digital services. 因此,高质量的训练数据已成为继芯片、内存和电力之后的又一战略资源。领先的人工智能模型开发者发现,在努力维持市场要求的创新步伐时,这些资源正变得短缺。种种迹象表明,这种需求的一部分正通过用户自身生成的信息来满足,作为交换,用户可以免费使用日常数字服务。

This story originally appeared in WIRED en Español and has been translated from Spanish. 本文最初发表于《连线》西班牙语版,由西班牙语翻译而来。