AI detectors are creating a new era of distrust
AI detectors are creating a new era of distrust
AI 检测工具正在开启一个充满不信任的新时代
Tools that promise to detect AI writing can be dubious, but educators and publishers are using them anyway. 那些承诺能够检测 AI 写作的工具可能并不可靠,但教育工作者和出版商们依然在使用它们。
How it started
起源
Long before ChatGPT became a thing, educators and editors frequently used anti-plagiarism tools to see if writers were being honest about their work. These tools work by comparing a written work against a database filled with content from across the web, scholarly articles, and more to check for matching sentences and phrases. Some, like Turnitin, offer a percentage that claims to illustrate how much of the student’s writing overlaps with other works. Between potential false positives and uncertainty about whether work was duplicated intentionally, some educators have backed away from using that particular tool. 在 ChatGPT 问世之前,教育工作者和编辑们就经常使用反剽窃工具来核实作者是否诚实地完成了作品。这些工具通过将书面作品与包含全网内容、学术文章等在内的数据库进行比对,来检查是否存在重复的句子和短语。一些工具(如 Turnitin)会给出一个百分比,声称能展示学生作品与他人作品的重合程度。由于存在潜在的误报以及难以确定作品是否为故意抄袭,一些教育工作者已经不再使用该工具。
But now, the hunt for copied content is evolving into a war on AI-generated work. As quickly as students have picked up ChatGPT, Google Gemini, and Microsoft Copilot, teachers have adopted so-called AI detectors just as fast. A survey from the Center for Democracy and Technology found that 43 percent of sixth to 12th grade teachers in the US regularly used AI detectors between 2024 and 2025. Some universities already using Turnitin in their learning management systems found that the service automatically enabled AI detection when the tool launched in 2023. 但现在,对抄袭内容的搜寻正演变成一场针对 AI 生成内容的战争。学生们使用 ChatGPT、Google Gemini 和 Microsoft Copilot 的速度有多快,老师们采用所谓“AI 检测器”的速度就有多快。民主与技术中心(Center for Democracy and Technology)的一项调查发现,2024 年至 2025 年间,美国 43% 的 6 至 12 年级教师经常使用 AI 检测工具。一些在学习管理系统中已经使用 Turnitin 的大学发现,当该工具在 2023 年推出 AI 检测功能时,系统自动启用了这一功能。
Instead of comparing pieces of written content, AI detectors like GPTZero, Pangram, and the one created by Turnitin rely on their own AI models to guess whether something might not be human-written — a process that’s arguably even murkier than matching text on the web. As noted by GPTZero, AI detectors use an algorithm to analyze a text’s wording, rhythm, and structure, as well as to pick up on patterns in length and tone that may be more common in AI-written text. This relatively subjective evaluation isn’t as solid as something you could back up by matching text online, and can get tripped up by writers who speak English as a second language. Despite this, Turnitin has said that its AI detector falsely flags less than 1 percent of human-written content as AI, while Pangram claims its false positive rate is just 1 in 10,000. GPTZero claims it has a similarly low rate of mistaking human content for AI. 与比对书面内容不同,GPTZero、Pangram 以及 Turnitin 开发的 AI 检测器依赖于它们自己的 AI 模型来推测某段文字是否非人类所写——这个过程可以说比在网上匹配文本更加模糊。正如 GPTZero 所指出的,AI 检测器使用算法来分析文本的措辞、节奏和结构,并捕捉 AI 文本中可能更常见的长度和语气模式。这种相对主观的评估不如在线文本匹配那样可靠,而且很容易误伤那些将英语作为第二语言的作者。尽管如此,Turnitin 称其 AI 检测器将人类创作内容误判为 AI 的比例不到 1%,而 Pangram 声称其误报率仅为万分之一。GPTZero 也声称其将人类内容误认为 AI 的比例同样很低。
How it’s going
现状
People online are already accusing each other of “sounding like AI,” but the ready availability of AI detection tools is only adding fuel to the LLM witch hunt. In some high-profile cases, AI writing accusations have directly impacted people’s livelihoods and reputations. Last month, the publisher Minotaur dropped a $2 million book deal over concerns that its author, Jerry Falade, used AI — something he vehemently denies. 网络上的人们已经开始互相指责对方“听起来像 AI”,而 AI 检测工具的唾手可得只会为这场针对大语言模型(LLM)的“猎巫行动”火上浇油。在一些备受瞩目的案例中,关于 AI 写作的指控已经直接影响了人们的生计和声誉。上个月,出版商 Minotaur 取消了一份价值 200 万美元的图书合同,原因就是担心作者 Jerry Falade 使用了 AI——而他对此坚决否认。
There’s Thierry Rignol, a French national who sued Yale last year after a professor accused him of writing portions of his final exam with AI, resulting in a failing grade and a one-year suspension. The professor used GPTZero to scan Rignol’s writing for signs of AI, but the lawsuit argues that “AI surveillance and detection tools are known to unfairly target non-native English speakers” like Rignol. In February, a student at Adelphi University won a lawsuit against the school after his professor similarly claimed he used AI to write an essay. Though the lawsuit doesn’t say which AI tool the professor used to examine the student’s essay, Adelphi University has a licensing agreement with Turnitin. 法国公民 Thierry Rignol 去年起诉了耶鲁大学,此前一位教授指控他在期末考试中使用 AI 写作,导致他不及格并被停学一年。该教授使用 GPTZero 扫描 Rignol 的试卷以寻找 AI 痕迹,但诉讼指出,“众所周知,AI 监控和检测工具会不公平地针对像 Rignol 这样非英语母语的人”。今年 2 月,阿德尔菲大学(Adelphi University)的一名学生在起诉学校后胜诉,此前他的教授也声称他使用 AI 撰写论文。虽然诉讼中没有说明教授使用了哪种 AI 工具来检查该学生的论文,但阿德尔菲大学与 Turnitin 签有许可协议。
A 2023 Stanford study found that AI detectors falsely flagged essays written by non-native English speakers as AI more often than native speakers. (Many services still argue that their tools are accurate when dealing with text written by non-native speakers.) These tools may also be biased against neurodivergent writers. 斯坦福大学 2023 年的一项研究发现,AI 检测器将非英语母语者撰写的论文误判为 AI 的频率高于母语者。(许多服务商仍然坚称,在处理非母语者撰写的文本时,他们的工具是准确的。)这些工具还可能对神经多样性(如自闭症、多动症等)作者存在偏见。
As pointed out by the University of California, Los Angeles, AI detection tools are trained to pick up on patterns that could indicate AI use, such as repetitive terms and phrases, text that sounds too formal or informal, and nonsensical phrasing. Some, like QuillBot, also measure the “unpredictability” of text, as “AI tends to make the most ‘obvious’ or most common language choices as compared with human-produced writing,” according to UCLA. They may also look for sentence structure that remains the same throughout as another sign of AI. But these measurements aren’t indicative of AI on their own, as some people may just have a writing style with these qualities. 正如加州大学洛杉矶分校(UCLA)所指出的,AI 检测工具经过训练,旨在捕捉可能表明 AI 使用的模式,例如重复的术语和短语、听起来过于正式或非正式的文本,以及毫无意义的措辞。据 UCLA 称,一些工具(如 QuillBot)还会测量文本的“不可预测性”,因为“与人类创作的写作相比,AI 倾向于做出最‘明显’或最常见的语言选择”。它们还可能寻找贯穿全文且保持不变的句子结构,将其作为 AI 的另一个标志。但这些指标本身并不能证明就是 AI 所为,因为有些人的写作风格可能恰好具备这些特质。
Even though Turnitin touts low false positive rates, it maintains that its tool “may not always be accurate” and shouldn’t be used to take actions against a student. Grammarly warns that users “should never rely on the results of an AI detector alone,” while GPTZero says “no AI detector can ever truly be 100% perfect.” OpenAI even shut down its own AI writing detector in 2023 due to low accuracy. 尽管 Turnitin 标榜其误报率低,但它也承认其工具“可能并不总是准确的”,不应仅凭此对学生采取惩戒措施。Grammarly 警告用户“绝不应仅依赖 AI 检测器的结果”,而 GPTZero 则表示“没有任何 AI 检测器能做到 100% 完美”。OpenAI 甚至在 2023 年因准确率过低而关闭了自家的 AI 写作检测器。
But AI writing accusations are still being flung across the web. Last week, in a video broadcast to the more than 3.5 million followers across his social channels, Ozzy Osbourne’s son, Jack, accused journalist and Verge contributor Kat Tenbarge of using AI to write an article. 然而,关于 AI 写作的指控仍在网络上四处蔓延。上周,奥兹·奥斯本(Ozzy Osbourne)的儿子杰克(Jack)在向其社交渠道超过 350 万粉丝发布的视频中,指责记者兼《The Verge》撰稿人 Kat Tenbarge 使用 AI 撰写文章。