Opus 5.5 loves to tell you ‘this matters’ (and other AI writing tells)

Opus 5.5 loves to tell you ‘this matters’ (and other AI writing tells)

Opus 5.5 喜欢告诉你“这很重要”(以及其他 AI 写作的蛛丝马迹)

Now that LLM-generated prose is everywhere, human beings are eager for ways to sniff it out. While early tells like em-dashes and “delve” are long gone, researchers say there are still plenty of telltale habits that AI models fall back on when writing prose. 如今,大语言模型生成的文章随处可见,人类迫切想要找到识别它们的方法。虽然像破折号和“delve”(钻研)这类早期的识别特征早已消失,但研究人员表示,AI 模型在撰写文章时仍然会依赖许多显而易见的习惯。

A new study from the marketing firm Graphite looked at the writing habits of frontier models, sussing out each model’s favorite words and phrases. While old tells like em-dash use have been stamped out, models still fall back on contrast-heavy constructions, with each model version showing its own unique quirks. 营销公司 Graphite 的一项新研究调查了前沿模型的写作习惯,梳理出了每个模型最喜欢的词汇和短语。虽然像使用破折号这样的旧特征已被消除,但模型仍然倾向于使用对比强烈的句式,且每个模型版本都表现出其独特的怪癖。

The biggest surprise is how broad the scope of tells turns out to be. Graphite found 13,000 phrases that were at least twice as common in the AI content as human content — their definition of a “tell.” 最令人惊讶的是,这些识别特征的范围竟如此广泛。Graphite 发现了 13,000 个短语,它们在 AI 内容中出现的频率至少是人类内容的两倍——这就是他们对“识别特征”(tell)的定义。

“It turns out that Claude models are actually getting closer to the human word distribution over time,” Graphite’s chief AI officer Greg Druck told TechCrunch. “And for the GPT models, it’s getting further away.” “事实证明,Claude 模型随着时间的推移,其词汇分布实际上正越来越接近人类,”Graphite 的首席 AI 官 Greg Druck 告诉 TechCrunch,“而对于 GPT 模型,这种差距反而越来越大。”

Studying AI-generated writing at scale required a careful study design. Graphite started with a corpus of 10,000 articles published before the release of ChatGPT, serving as the human-generated control group. Then researchers had different AI models rewrite the articles from summaries, hoping to eliminate as much source bias as possible. With matching samples from both humans and each model, they could compare how often certain words and phrases appeared in AI writing, as well as broader patterns in sentence construction. 大规模研究 AI 生成的写作需要精心的实验设计。Graphite 首先选取了 ChatGPT 发布前发表的 10,000 篇文章作为语料库,作为人类生成的对照组。随后,研究人员让不同的 AI 模型根据摘要重写这些文章,以尽可能消除源偏差。通过对比人类和每个模型生成的匹配样本,他们能够比较特定词汇和短语在 AI 写作中出现的频率,以及更广泛的句式结构模式。

According to Graphite’s results, Claude Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples. While Opus 5.5 now avoids the “it’s not X, it’s Y” sentence construction, it still tends to say something “is more than an X, it’s a Y.” Above all, Opus loves to tell you why things matter, using the phrase “this matters” 116 times more often than human writing, while “why X matters” occurs 92 times more often. 根据 Graphite 的研究结果,Claude Opus 5.5 最明显的特征是“dependable”(可靠的)一词,其出现频率比人类样本高出 23 倍。虽然 Opus 5.5 现在避免使用“it’s not X, it’s Y”(不是 X,而是 Y)的句式,但它仍然倾向于说某事物“is more than an X, it’s a Y”(不仅仅是 X,更是 Y)。最重要的是,Opus 喜欢告诉你为什么事情很重要,使用“this matters”(这很重要)这一短语的频率比人类写作高出 116 倍,而“why X matters”(为什么 X 很重要)的出现频率则高出 92 倍。

OpenAI’s Astra has a different set of tip-offs. This model loves to describe “another dimension” of whatever it’s talking about, and tends to hedge claims by saying an action “may provide” or “can provide” a particular benefit. Its biggest tell is what Graphite calls the “corrective framing,” where a topic is defined as “not simply X” or offered as an alternative, “rather than relying on X.” According to graphite’s research, those constructions were more than 100 times more common in Astra-generated prose than in human writing. OpenAI 的 Astra 则有另一套识别特征。该模型喜欢描述它所讨论事物的“另一个维度”(another dimension),并倾向于通过说某项行动“may provide”(可能提供)或“can provide”(能够提供)某种特定好处来规避风险。它最明显的特征是 Graphite 所称的“纠正性框架”(corrective framing),即把一个主题定义为“not simply X”(不仅仅是 X),或者作为一种替代方案提出,即“rather than relying on X”(而不是依赖 X)。根据 Graphite 的研究,这些句式在 Astra 生成的文章中出现的频率比人类写作高出 100 多倍。

Notably, all the frontier labs seem to have responded to the idea that models overuse em-dashes. In Graphite’s samples, Opus 5.5 used the punctuation mark 99% less often than Opus 5. Astra now uses it 88% less than human samples, whereas Gemini 3.1 Pro has almost completely eliminated the em-dash from its writing. 值得注意的是,所有前沿实验室似乎都对“模型过度使用破折号”这一观点做出了回应。在 Graphite 的样本中,Opus 5.5 使用该标点符号的频率比 Opus 5 降低了 99%。Astra 现在的使用频率比人类样本低 88%,而 Gemini 3.1 Pro 则几乎完全从其写作中消除了破折号。

But while individual tells change, Graphite says the overall number is mostly holding steady. “It’s not like the tells are decreasing,” Druck told TechCrunch. “They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own.” 然而,尽管具体的识别特征在改变,Graphite 表示总数基本保持稳定。“识别特征并没有减少,”Druck 告诉 TechCrunch,“他们设法去除了最广为人知的特征,但其他特征又冒了出来。而且每个模型版本都有其独特的特征。”

It’s surprising that tells are so persistent, given the labs’ focus on human-like writing styles. In the Opus 5.5 release, Anthropic boasted that the model “communicates more naturally than prior models,” saying early users “found its writing clearer and easier to follow.” OpenAI made similar claims when releasing the GPT-6 versions of Sol and Luna, saying users could “expect to see more clarity, less jargon, [and] fewer odd turns of phrase.” 考虑到各实验室对类人写作风格的追求,这些识别特征如此顽固令人惊讶。在发布 Opus 5.5 时,Anthropic 吹嘘该模型“比之前的模型交流更自然”,并称早期用户“发现其写作更清晰、更容易理解”。OpenAI 在发布 GPT-6 版本的 Sol 和 Luna 时也提出了类似的说法,称用户可以“期待看到更清晰的表达、更少的术语,以及更少奇怪的措辞”。

But Druck is skeptical about how much the labs can do to completely eliminate telltale construction or phrases. “A general hypothesis I have is that the labs are less able to control some of these things than you might expect,” Druck says. “These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through.” 但 Druck 对实验室能在多大程度上彻底消除这些识别性的句式或短语持怀疑态度。“我的一个普遍假设是,实验室对这些事物的控制能力可能不如你预期的那样强,”Druck 说,“这些都是拥有数十亿参数的巨型模型。他们能进行的测试数量有限,总会有一些东西漏网。”