Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
审计合成回忆录:针对大语言模型生成的自传,根据其描述的生活记录衡量场景级虚构
Abstract: When a large language model (LLM) is asked to write a person’s life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. 摘要: 当要求大语言模型(LLM)撰写一个人的生平时,它所写的内容中有多少是真实发生的?我们提出了一项场景级的案例研究审计——据我们所知,这是基于非系统性文献检索,首次针对大语言模型生成的自传与特定主题的真实语料库进行的量化审计。
The subject and the author of this paper are the same person: a 366-day “page-a-day” book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day’s quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. 本文的研究对象与作者为同一人:我们利用一个对话式大语言模型起草了一本为期 366 天的“每日一页”第一人称轶事日记。该模型所使用的输入仅限于一个模板、两个示例日以及每日引言,而非作者本人的语料库。随后,我们使用分析前设定的四级评估标准,针对独立的验证语料库,对每一天的轶事场景进行了审计。
We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. 我们将验证失败率定义为未被评为“已验证”(即场景得到正面证实)的天数占比:366 天中有 354 天验证失败,失败率为 96.7%(Wilson 95% 置信区间为 94.4-98.1%)。仅有 12 天包含得到证实的场景;19 天(5.2%)所陈述的内容与记录直接矛盾;主要的失败模式是“接地漂移”(grounded drift)——即在虚构的场景中混入真实的人物、雇主和背景,尽管不同评估者测得的比例有所差异。
Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject’s corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). 独立的重新评估重复了上述主要结论(没有证据表明原始失败率被夸大),同时也显示该四级分类法的可靠性仅处于尚可至中等水平。使用当前主流模型在相同输入下重新生成这些日记,验证失败率高达 100%;而将生成过程建立在研究对象的语料库基础上,虽然显著提高了验证率,但仍存在大量的残留失败(83.3%)。
We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect. 我们贡献了这一测量方法、一种可重复使用的审计工具(我们证明了该工具在“弱/未验证”边界上的不可靠性),以及一种具有量化效果的接地(grounding)补救措施。