The revolt of the reader
The revolt of the reader
读者的反叛
Reading is important to me. While I’m not a quick reader nor an especially voracious one, I have found that long-form reading has had a profound influence on me over my life. And because reading is important to me, writing is too: writing not only allows us to convey our ideas, but the very act forces us to test and distill them — at once making our ideas more robust while providing the vehicle by which to share them.
阅读对我而言很重要。虽然我阅读速度不快,也不是那种特别贪婪的读者,但我发现长篇阅读对我的一生有着深远的影响。正因为阅读对我重要,写作也同样重要:写作不仅让我们能够传达思想,而且这一行为本身迫使我们去检验和提炼这些思想——在使我们的观点更加稳固的同时,也提供了分享它们的载体。
But I am writing this now because — speaking as a reader — we are exasperated: too many people — people we otherwise like and respect! — are writing (or otherwise putting their name to) pieces that are clearly LLM-authored. We readers are left with pointed questions for those promoting LLM-authored pieces: do you think readers can’t tell? Or do you think readers don’t care?
但我现在写下这些是因为——作为一名读者——我们感到非常愤怒:太多人——那些我们本来喜欢和尊重的人!——正在撰写(或以其他方式署名)那些明显由大语言模型(LLM)生成的文章。我们读者不禁要向那些推广 LLM 生成内容的人提出尖锐的问题:你认为读者看不出来吗?还是你认为读者根本不在乎?
To answer the first question, readers can absolutely tell. To those who read broadly, the hand of the LLM is so clear it’s as if the writer’s intellectual fly is open. In fact, it’s so jarring that I have to believe that those writing with LLMs are either not reading enough to see the LLM’s obvious structural tells — or (and?) they aren’t even reading their own content. (A confession: with particularly egregious pieces, I have fantasized about sentencing the author to read them aloud, certain that they themselves will be unable to endure the slop that they are foisting upon the rest of us.)
回答第一个问题:读者绝对看得出来。对于博览群书的人来说,LLM 的痕迹是如此明显,简直就像作者的“智力拉链没拉上”一样尴尬。事实上,这种感觉非常刺耳,以至于我不得不相信,那些使用 LLM 写作的人要么是阅读量不足,看不出 LLM 明显的结构性破绽,要么(或者两者皆有?)他们甚至根本没有阅读自己发布的内容。(坦白说:遇到特别糟糕的文章时,我曾幻想判处作者大声朗读它们,我确信他们自己也无法忍受那些强加给我们的垃圾内容。)
As to the second question, readers emphatically care. In fact, the tells of LLM writing are so grating (“and here’s why that framing matters!”) that our brains pull an LLM-triggered ejection handle, bailing us out mid-sentence in an act of self-preservation. And the first person plural there is deliberate: as Cynthia Dunlop writes, of the 668 developers that replied to her survey, 78% “stop reading immediately” when they detect an LLM. And there are consequences that outlast the piece: 71% of the respondents in Cynthia’s survey also “avoid the author in the future” (!!).
至于第二个问题,读者非常在乎。事实上,LLM 写作的特征是如此令人厌烦(比如“这就是为什么这种框架很重要!”),以至于我们的大脑会拉动“LLM 触发的弹射手柄”,在句子读到一半时为了自我保护而选择退出。这里使用第一人称复数是刻意的:正如 Cynthia Dunlop 所写,在她调查的 668 名开发者中,78% 的人在察觉到 LLM 内容时会“立即停止阅读”。而且后果远不止于此:Cynthia 调查中 71% 的受访者还会“在未来避开该作者”(!!)。
Revealingly, the respondents are not after linguistic perfection, but rather authenticity: 98% reported preferring an author’s own (imperfectly) written piece over an LLM-polished one. Finally, be wary of dismissing Cynthia’s respondents as a self-selecting group: active readers on social media are exactly the folks most likely to repost or otherwise promote writing they like — the early adopters and the tastemakers of online prose.
耐人寻味的是,受访者追求的并非语言的完美,而是真实性:98% 的人表示更喜欢作者自己(不完美)写的文章,而不是经过 LLM 润色的作品。最后,不要轻易将 Cynthia 的受访者视为一个自我选择的群体:社交媒体上的活跃读者正是最有可能转发或推广他们所喜爱文章的人——他们是早期采用者,也是网络散文的品味引领者。
Why do people have this reaction? Beyond having to endure aggravating stylistic tics, when reading a piece that has had substantial LLM assistance, we — the readers — don’t know what is real and what isn’t. As I wrote in RFD 576, to use an LLM to write is to void the social contract between writer and reader: we readers shouldn’t be expected to labor to understand a sentence that the writer themselves didn’t work to create.
为什么人们会有这种反应?除了要忍受令人恼火的文体怪癖外,当阅读一篇有大量 LLM 辅助的文章时,我们——读者——不知道什么是真实的,什么是虚构的。正如我在 RFD 576 中所写,使用 LLM 写作等同于废除了作者与读者之间的社会契约:我们读者不应该被要求去费力理解一个连作者自己都没有用心创作的句子。
Does the revolt of the reader matter? If it needs to be said, when you are using an LLM to author a public piece, you are no longer writing for yourself or to otherwise pressure-test your own ideas — the only purpose is to serve the reader. If readers choose to ignore you (perhaps forever!), you will have undermined yourself: instead of attracting readers you will be actively repelling them. So other judgement about its use aside, using an LLM will increasingly become simply… ineffective.
读者的反叛重要吗?如果非要说的话,当你使用 LLM 撰写公开文章时,你不再是为了自己写作,也不是为了压力测试自己的想法——唯一的目的就是服务读者。如果读者选择无视你(甚至永远无视!),你就是在破坏自己:你不仅没能吸引读者,反而是在积极地排斥他们。因此,抛开对使用 LLM 的其他评价不谈,使用 LLM 将越来越变得……无效。
In this regard, I am reminded of the arc of e-mail spam. There was a time in the early 2000s when people were (reasonably!) afraid that the explosion of spam would mean the end of e-mail. This was an era largely before social networking, and e-mail was the canonical killer app of the Internet; that we were losing e-mail to spam felt deeply dispiriting. But sometime in the late 2000s, we turned a corner: spam filtering improved so much that the economics of spam were undermined.
在这方面,我想到了电子邮件垃圾邮件的发展历程。在 21 世纪初,人们曾(合理地!)担心垃圾邮件的泛滥意味着电子邮件的终结。那是一个社交网络尚未普及的时代,电子邮件是互联网的典型杀手级应用;眼看着电子邮件被垃圾邮件淹没,人们感到非常沮丧。但在 21 世纪 2000 年代末的某个时候,我们迎来转机:垃圾邮件过滤技术得到了极大的改进,以至于垃圾邮件的经济效益被瓦解了。
Moreover, as spam became effectively contained, the consequences of being labeled as spam became increasingly dire. Today, legitimate businesses are very careful about how they use bulk e-mail; to be labeled as spam is to effectively destroy your domain name and tarnish your brand. The war on spam started to turn when we could identify it at scale; could something similar happen to LLM-authored writing?
此外,随着垃圾邮件得到有效遏制,被标记为垃圾邮件的后果变得越来越严重。如今,正规企业在进行群发邮件时非常谨慎;一旦被标记为垃圾邮件,实际上就等于摧毁了你的域名并玷污了你的品牌。反垃圾邮件战争是在我们能够大规模识别垃圾邮件时开始转折的;LLM 生成的文章是否也会发生类似的情况?
Like spam, an LLM’s influence is readily identifiable to humans reading it; surely this is a solvable problem? Up until somewhat recently, the results on this problem had been decidedly mixed. I had tried to use LLMs themselves for LLM identification, but I found that their false negative rate was far too high: they were chipper in accepting stuff that I was certain was LLM-authored. Other services seemed to look for basic LLM tells, but as an avid (and unapologetic!) user of the em-dash, these superficial techniques make me shift nervously in my seat (and I found them to be so broadly unreliable that they didn’t earn regular use).
像垃圾邮件一样,LLM 的影响对于人类读者来说是很容易识别的;这肯定是一个可以解决的问题吧?直到最近,针对这一问题的结果还相当参差不齐。我曾尝试使用 LLM 本身来识别 LLM 内容,但我发现它们的漏报率(false negative rate)太高了:它们总是兴高采烈地接受那些我确信是由 LLM 生成的内容。其他服务似乎在寻找基础的 LLM 特征,但作为一个热衷于(且毫不掩饰!)使用破折号的人,这些肤浅的技术让我坐立难安(而且我发现它们非常不可靠,根本无法投入常规使用)。
Late last year, Pangram Labs launched their Pangram 3 model. I found it to be a huge leap over other detectors, and became an avid user. Importantly, over months of use (and especially on samples that I otherwise knew the origins of), I found its false positive rate to be very low: when Pangram identified a text as being largely AI written, I could say with some certainty that an LLM was heavily involved. (I found its false negative rate to be higher than I would like, but it was a small price to pay for a low false positive rate.)
去年年底,Pangram Labs 推出了他们的 Pangram 3 模型。我发现它比其他检测器有了巨大的飞跃,并成为了它的忠实用户。重要的是,经过几个月的使用(特别是在我已知来源的样本上),我发现它的误报率(false positive rate)非常低:当 Pangram 将一段文本识别为主要由 AI 编写时,我可以相当肯定地说,其中有大量的 LLM 参与。(我发现它的漏报率比我预期的要高,但为了获得极低的误报率,这是一个很小的代价。)
A little over a month ago, they introduced Pangram 4, which I found to be a step-function improvement over the already-impressive Pangram 3: in my experience it has an astonishingly low false positive rate and low false negative rate (which is to say: very high accuracy!). I am finding it to be so effective (and the loss of trust in voices that use LLMs to be so precipitous) that I recently extended RFD 576 to be explicit about public writing, specifically mandating that public Oxide writing be reported by Pangram as human-authored. As I explained in the RFD, the standard for our public wri
一个多月前,他们推出了 Pangram 4,我发现它比已经令人印象深刻的 Pangram 3 有了阶跃式的提升:根据我的经验,它有着惊人的低误报率和低漏报率(也就是说:准确率非常高!)。我发现它非常有效(而且人们对使用 LLM 的声音的信任度下降得如此之快),以至于我最近扩展了 RFD 576,明确了关于公开写作的规定,特别要求 Oxide 的公开文章必须经由 Pangram 检测并显示为人类创作。正如我在 RFD 中所解释的那样,我们公开写作的标准……