AI-Generated Papers and Journal Integrity
AI-Generated Papers and Journal Integrity
AI 生成论文与期刊诚信
Two quite different things are discussed under one heading, and almost all the confusion comes from that. One is a researcher using a model to draft, edit or translate work they did. The other is fabricated content submitted to inflate a publication record. The first is a disclosure question. The second is fraud, and it is not new. 在同一个标题下,人们讨论着两件截然不同的事情,而几乎所有的困惑都源于此。一种情况是研究人员使用模型来起草、编辑或翻译他们自己的工作;另一种情况则是为了增加发表记录而提交的虚假内容。前者是一个披露问题,而后者是欺诈,且这并非新鲜事。
Two problems wearing one name. A non-native English speaker using a model to make their methods section readable has done nothing wrong and has improved the literature. A paper mill generating plausible manuscripts at volume has committed fraud, and would have done so with or without a language model — mills existed, using image manipulation, template text and fabricated data, long before this technology arrived. 两个问题共用一个名称。一位非英语母语者使用模型来使他们的方法部分更具可读性,这并没有做错什么,反而改进了文献质量。而一个大规模生成看似合理手稿的“论文工厂”则犯了欺诈罪,无论有没有语言模型,它们都会这样做——早在这项技术出现之前,论文工厂就已经存在,它们利用图像处理、模板文本和伪造数据进行造假。
Keeping them apart matters because they call for opposite responses. The first needs a disclosure norm and nothing else. The second needs content verification, and content verification does not care what tool produced the content. Any policy built around detecting machine text will punish the first group and miss most of the second, because fabricated research that has been lightly rewritten is indistinguishable from careful assisted writing. 将两者区分开来至关重要,因为它们需要截然不同的应对措施。前者只需要一个披露规范,别无他求。后者则需要内容验证,而内容验证并不关心内容是由什么工具生成的。任何围绕检测机器文本制定的政策,最终都会惩罚第一类群体,却漏掉第二类中的大多数,因为经过轻微改写的虚假研究与精心辅助写作的内容是无法区分的。
It is also worth being clear about where the demand comes from, because it explains why no technical measure will resolve this. Paper mills exist because publication counts are used as a proxy for research contribution in hiring, promotion and institutional ranking, in systems large enough that buying an authorship is a rational purchase for some buyers. Generative tools lowered the cost of supplying that demand; they did not create it. 同样值得明确的是,这种需求从何而来,因为它解释了为什么没有任何技术手段能解决这个问题。论文工厂之所以存在,是因为在招聘、晋升和机构排名中,发表数量被用作衡量研究贡献的指标,在这些庞大的系统中,购买作者身份对某些买家来说是一种“理性”的消费。生成式工具降低了满足这种需求的成本,但它们并没有创造这种需求。
A detector, even a perfect one, sits downstream of an incentive that would simply route around it — which is why the interventions with the best track record are the ones that attack verifiability, such as requiring data and code, rather than the ones that attack production. 即使是一个完美的检测器,也处于激励机制的下游,而这种激励机制只会绕过它——这就是为什么记录中效果最好的干预措施是那些针对“可验证性”的手段(例如要求提供数据和代码),而不是那些针对“生产过程”的手段。
What the artefacts look like
伪造痕迹的表现形式
-
Leftover interface text. Phrases that belong to a chat interface rather than to a paper — a preamble about being a language model, or a stray instruction to continue — have appeared verbatim in published articles. What this reveals is not that a model was used. It is that the author, the reviewers, the editor and the production process all failed to read the text, which is a far worse finding. 残留的界面文本。 属于聊天界面而非论文的短语——例如关于“作为语言模型”的开场白,或是一条残留的“继续”指令——曾逐字出现在已发表的文章中。这揭示的不是模型被使用了,而是作者、审稿人、编辑和出版流程都没有认真阅读文本,这是一个糟糕得多的发现。
-
Nonsense figures. Generated images with garbled labels and anatomically impossible content have been published and subsequently retracted. Same lesson: the figure was never examined. 荒谬的图表。 带有乱码标签和解剖学上不可能内容的生成图像曾被发表,随后又被撤回。同样的教训:这些图表从未被检查过。
-
Tortured phrases. Standard terms replaced by odd synonyms — the sort of substitution a paraphrasing tool makes to evade plagiarism software. This artefact predates language models entirely and has been catalogued by integrity screening projects for years, which is a useful reminder that the underlying behaviour is older than the current tooling. 扭曲的短语。 标准术语被奇怪的同义词替换——这是改写工具为了规避查重软件而进行的替换。这种伪造痕迹在语言模型出现之前就已存在,并已被诚信审查项目记录多年,这提醒我们,这种潜在的行为比当前的工具出现得更早。
-
Fabricated references. The most consequential of the four, because it survives copy-editing, looks completely normal, and corrupts the citation graph. 伪造的参考文献。 这是四者中后果最严重的,因为它能通过编辑校对,看起来完全正常,并破坏了引文图谱。
-
Confident, specific, unsupported claims. The hardest to spot and the reason the others matter: the failure mode of a language model is fluency without grounding, which is precisely what peer review is least well equipped to catch quickly. 自信、具体但缺乏支持的断言。 这是最难发现的,也是其他问题之所以重要的原因:语言模型的失效模式是“流利但缺乏依据”,而这恰恰是同行评审最不擅长快速捕捉到的。
Why detection is the wrong instrument
为什么检测是错误的手段
Statistical detectors for machine-generated text have two problems that are not going to be engineered away, and a third that is arithmetic. They are biased against non-native English writers. Published work evaluating detectors on essays by non-native English speakers found them flagged at high rates, for a straightforward reason: the features detectors key on — limited vocabulary variety, regular sentence structure, low unpredictability — describe careful second-language writing as well as they describe generated text. 针对机器生成文本的统计检测器存在两个无法通过技术手段消除的问题,以及第三个算术层面的问题。首先,它们对非英语母语作者存在偏见。评估检测器在非英语母语者论文上表现的研究发现,这些论文被高频率标记,原因很简单:检测器所依赖的特征——词汇多样性有限、句子结构规律、不可预测性低——既描述了谨慎的第二语言写作,也描述了生成的文本。
They are trivially defeated. Paraphrasing, editing or asking for a different style moves text out of the detected region, so the detector selects against exactly the people who did not try to hide anything. And the base rate defeats what remains. Run any detector with a small false-positive rate across a large submission stream and the number of wrongly flagged honest authors will be large in absolute terms — potentially larger than the number of genuine cases, depending on how rare the genuine cases are. A flag is therefore not evidence about an individual manuscript, and using it as though it were is how careers get damaged over a probability score with no chain of custody. 其次,它们很容易被击败。改写、编辑或要求改变风格就能使文本脱离检测区域,因此检测器恰恰针对的是那些没有试图隐藏任何东西的人。最后,基准率问题让剩下的检测手段也失效了。在大量的投稿流中运行任何具有低误报率的检测器,被错误标记的诚实作者在绝对数量上会非常大——根据真实造假案例的稀有程度,错误标记的数量甚至可能超过真实案例。因此,一个标记并不能作为单篇手稿的证据,将其当作证据使用,会导致人们仅仅因为一个缺乏监管链的概率分数而毁掉职业生涯。
Where policy has settled
政策现状
There is now broad agreement among major publishers, editors and publication ethics bodies on a small set of points. 目前,主要出版商、编辑和出版伦理机构在一小部分问题上达成了广泛共识。
| Position | Description |
|---|---|
| Not an author | A model cannot be listed as an author. Authorship entails accountability for the work and the ability to approve the final version, and a tool can do neither. |
| Disclose use | Substantive use in producing the manuscript is disclosed, usually in methods or acknowledgements. Venues differ on whether routine language editing needs declaring, so the specific policy has to be read. |
| Full accountability | Authors are responsible for everything in the paper, including anything a model produced. “The tool generated it” is not a mitigation for a fabricated citation or an invented result. |
| Review is separate | Using these tools during peer review is governed by different and generally stricter rules, for confidentiality reasons. |
| 立场 | 描述 |
|---|---|
| 非作者 | 模型不能被列为作者。作者身份意味着对工作的问责能力以及批准最终版本的能力,而工具两者都无法做到。 |
| 披露使用 | 在撰写手稿过程中的实质性使用必须披露,通常在方法或致谢部分。不同期刊对于是否需要声明常规语言编辑存在差异,因此必须阅读具体政策。 |
| 完全问责 | 作者对论文中的一切负责,包括模型生成的任何内容。“工具生成的”不能作为伪造引用或捏造结果的减刑理由。 |
| 评审独立 | 出于保密原因,在同行评审期间使用这些工具受到不同且通常更严格的规则约束。 |
The checks that actually work
真正有效的检查手段
Every one of these targets whether the content is true rather than how it was produced, which is the right target and is also the only one that is defensible when challenged. Resolve every reference. Mechanical, cheap, and… 所有这些检查手段的目标都是内容是否真实,而不是它是如何产生的。这是正确的方向,也是在受到质疑时唯一站得住脚的方法。核实每一条参考文献。机械、廉价,且……