We still don’t know how people are really using AI

We still don’t know how people are really using AI

我们仍然不知道人们究竟是如何使用人工智能的

EXECUTIVE SUMMARY 执行摘要

AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. 人工智能研究人员表示,像 Anthropic 和 OpenAI 这样的 AI 公司会定期发布关于人们如何使用 Claude 和 ChatGPT 等产品的报告,但他们只发布他们希望我们看到的数据。斯坦福大学可信人工智能研究实验室(STAIR Lab)的计算机科学博士候选人 Anka Reuel 表示:“目前没有独立的来源来证实这些数据。”

Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The intent is to provide independent sources of information that can help researchers and policymakers assess how people are using generative AI. Reuel 是一个名为“AI 观察站”(AI Observatory)的新研究项目的共同负责人,该项目旨在填补这一空白。这是一个公共平台,它汇总并分析了通过七个现有数据集、在用户同意下收集的与 Claude 和 Gemini 等流行模型的真实 AI 对话。其目的是提供独立的信息来源,帮助研究人员和政策制定者评估人们如何使用生成式 AI。

Highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited data, says Reuel. The AI Observatory found that AI use differs significantly across models and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use. Reuel 指出,目前关于 AI 益处和风险的重大决策都是基于非常有限的数据做出的。AI 观察站发现,不同模型之间的 AI 使用情况存在显著差异,且随时间推移而发生了变化。其研究显示,人们表现出的敏感行为远多于大型 AI 公司报告中所记录的内容,因为这些公司往往更关注工作用途而非个人使用。

The Anthropic Economic Index is one of the best-known and most widely cited sources of AI usage data, but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses. When the AI Observatory researchers applied Anthropic’s methods to their dataset, they found that nearly half the conversations—48%—would have been filtered out. Anthropic 经济指数是 AI 使用数据中最著名、引用最广泛的来源之一,但它存在盲点。顾名思义,它专注于 Claude AI 在工作和生产力方面的用途,并过滤掉了与这些用途无关的对话。当 AI 观察站的研究人员将 Anthropic 的方法应用于他们的数据集时,发现近一半(48%)的对话会被过滤掉。

Those non-work-related conversations were more likely to involve health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT, similarly, found that only 30% of consumer use was related to work.) 这些非工作相关的对话更有可能涉及健康和人际关系(占 44.2%,而 Anthropic 的分析中仅为 31.2%)、成人或非法话题(7.9% 对 2.1%)、骚扰和仇恨言论(27.5% 对 5.66%)以及色情内容(16.7% 对 2.4%)。(OpenAI 关于 ChatGPT 的 2025 年报告同样发现,只有 30% 的消费者使用与工作相关。)

Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, but “having [the AI Observatory’s] bird’s-eye-view analysis” rather than leaving that information “sectioned off into a separate report” helps researchers understand the different uses more consistently, says David Widder, an assistant professor at the University of Texas at Austin, who researches how people interact with AI systems and is not involved with the AI Observatory. Anthropic 曾发布过单独的博客文章,讨论人们如何使用 Claude 获取支持或陪伴,甚至生成儿童性虐待材料(CSAM),但德克萨斯大学奥斯汀分校助理教授 David Widder 表示,拥有“AI 观察站”这种鸟瞰式的分析,而不是将这些信息“分割在单独的报告中”,有助于研究人员更连贯地理解不同的使用方式。Widder 研究人们如何与 AI 系统互动,并未参与 AI 观察站项目。

The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded. Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, as indicated by growing numbers of prompt tokens, response tokens, and conversation turns. AI 观察站研究的数据集包含了 2023 年至 2025 年间的对话,研究发现人们使用 AI 的方式以及各 AI 平台的响应方式都存在差异。作为 AI 观察站研究中规模最大、最详细的数据集之一,WildChat 中的对话随着时间的推移变得更长、更复杂,这体现在提示词 Token、响应 Token 和对话轮数的增加上。

There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile, the AI assistants’ self-disclosure (i.e., admitting to being a chatbot) decreased. Additionally, exchanges that the researchers labeled as sensitive—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—became less frequent. That might suggest that platforms were generally deploying more effective safeguards. 随着时间的推移,闲聊的内容也显著增加。这表明 AI 陪伴的需求正在增长;与此同时,AI 助手进行自我披露(即承认自己是聊天机器人)的频率有所下降。此外,研究人员标记为“敏感”的交流(即包含潜在有害或受限内容,包括性骚扰和仇恨言论)变得不那么频繁了。这可能表明各平台普遍部署了更有效的安全防护措施。

The AI Observatory also found that topics, interaction styles, conversation structures, and the likelihood and type of sensitive use cases differed from one model to another. For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.) AI 观察站还发现,不同模型在话题、交互风格、对话结构以及敏感用例的可能性和类型上存在差异。例如,研究人员发现人们更频繁地使用 Grok 和 Gemini 进行信息检索。Grok 尤其在新闻和政治信息方面很受欢迎,但它也是错误信息容易集中的地方。(这与其他研究一致,即错误信息在 Grok 上极易传播。xAI 未回应置评请求。)

Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. There were even differences between different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction. 与此同时,人们更倾向于使用 Anthropic 进行编程,使用 Gemini 进行社交和角色扮演,使用 ChatGPT 辅助完成家庭作业。即使是同一模型的不同版本之间也存在差异。研究人员发现,当 ChatGPT 由 GPT-3.5 驱动时,人们的对话较短;而使用 GPT-4o 时,对话更长且更具迭代性——考虑到该版本因容易导致情感依赖而闻名,这也就不足为奇了。

Companies’ reports, however, didn’t tend to capture these nuances between or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel. 然而,各公司的报告往往无法捕捉到这些模型之间甚至模型内部的细微差别。麻省理工学院媒体实验室的应届博士毕业生、与 Reuel 共同领导该研究的 Shayne Longpre 表示:“没有任何一家公司的报告能讲述完整的故事。”

To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions aggregated 85,633 conversational turns (that is, the user prompt and corresponding AI response) across 24,521 conversations from seven real-world datasets collected in previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025. 为了创建 AI 观察站,Reuel 与来自麻省理工学院、斯坦福大学、数据溯源倡议(Data Provenance Initiative)及其他机构的研究人员汇总了 85,633 个对话轮次(即用户提示和相应的 AI 回复),这些数据来自先前研究中收集的 24,521 次对话和七个真实世界数据集。这些对话来自 5,000 名用户,他们在 2023 年至 2025 年间与包括 ChatGPT、Gemini、Claude 和 Grok 在内的 52 种不同模型进行了互动。

But these conversations are a drop in the proverbial bucket compared with the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations. An Anthropic representative said the company’s published research reflects its research teams’ specific questions and interests and that it’s important to support external independent research. OpenAI did not respond to requests for comment. The fact that the AI Observatory’s dataset draws from voluntarily provided sources means it’s probably underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all. 但与大型实验室自身拥有的数据相比,这些对话只是沧海一粟。例如,最新的 Anthropic 经济 AI 指数基于对 100 万次 Claude 对话的分析;OpenAI 关于人们如何使用 ChatGPT 的报告分析了 150 万次对话。Anthropic 的一位代表表示,该公司发布的报告反映了其研究团队的具体问题和兴趣,支持外部独立研究非常重要。OpenAI 未回应置评请求。由于 AI 观察站的数据集来源于自愿提供的信息,这意味着它可能低估了敏感用途,因为人们不太愿意分享这些内容。因此,研究人员提醒说,其研究结果并不能代表全部情况。