TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking
TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking
TimeCapsule:作为历史意义构建方法的生成式幻觉
Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliable narrators of the past.
摘要: 大型语言模型(LLMs)存在时间上的过度暴露问题:由于在海量的当代语料库上进行训练,它们编码了现代概念,这使得它们在叙述过去时变得不可靠。
We present TimeCapsule, a 1.2B-parameter LLaMA-style causal model trained exclusively on Victorian texts (1800-1875) as an epistemologically isolated generative archive.
我们提出了 TimeCapsule,这是一个拥有 12 亿参数的 LLaMA 风格因果模型。该模型仅使用维多利亚时代(1800-1875 年)的文本进行训练,旨在作为一个认识论上相互隔离的生成式档案库。
Quantitative evaluation shows a 45.4% perplexity reduction over a GPT-2 baseline on held-out Victorian prose, while larger contemporary causal models achieve lower raw perplexity through broader pretraining but lack temporal isolation.
定量评估显示,在维多利亚时代的留存散文测试集上,该模型相比 GPT-2 基准线的困惑度(perplexity)降低了 45.4%。虽然更大的当代因果模型通过更广泛的预训练实现了更低的原始困惑度,但它们缺乏时间上的隔离性。
TimeCapsule exhibits computational sensemaking, generating historically plausible analogical explanations for unfamiliar modern concepts (e.g., describing a computer as a “hypertrophied lung”).
TimeCapsule 展示了计算意义构建的能力,能够为陌生的现代概念生成历史上看似合理的类比解释(例如,将计算机描述为“肥大的肺”)。
A qualitative hermeneutic probe with two humanities scholars revealed a crisis of authenticity, as both misclassified approximately 40% of genuine Victorian excerpts as machine-produced.
通过与两位人文学者进行的定性诠释学探究揭示了一场“真实性危机”,因为这两位学者都将约 40% 的真实维多利亚时代文本片段误认为是机器生成的。
We argue that structural ignorance of the future transforms hallucinations into interpretive probes of nineteenth-century ontologies.
我们认为,对未来的结构性无知将(模型的)幻觉转化为对十九世纪本体论的解释性探索。