What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces
What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces
什么是“好”?从大语言模型推理轨迹中提取并测试文学质量的隐性理论
Abstract: What makes writing “good” remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality.
摘要: 什么样的写作才算“好”?这在文学研究和计算语言学中始终是一个悬而未决的问题。我们通过两项研究,探讨了具备推理能力的大语言模型(LLM)如何评估文学质量。
In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model’s implicit theory of quality from its reasoning traces. Across five DeepSeek replications, the model achieves 79.3% mean tier-classification accuracy. The traces reveal a consistent stated theory: the model values intentionality over correctness, prioritizing craft, depth, and distinctive voice. A familiarity experiment with style-matched but unrecognizable passages suggests that source recognition may inflate scores, although this is confounded by genuine quality differences between canonical originals and researcher-written pastiches.
在研究一中,我们构建了一个包含 30 篇真实文本的基准测试集,涵盖了从经典文学到匿名论坛帖子的六个质量等级,并从模型的推理轨迹中提取了其对质量的隐性理论。在五次 DeepSeek 复现实验中,模型达到了 79.3% 的平均等级分类准确率。推理轨迹揭示了一个一致的理论:模型更看重“意图性”而非“正确性”,优先考虑写作技巧、深度和独特的文风。一项针对风格匹配但无法识别出处的段落的熟悉度实验表明,对来源的识别可能会提高评分,尽管这受到经典原作与研究人员仿写之间真实质量差异的干扰。
In Study 2, we probe this theory through systematic degradation of five canonical prose passages. We apply six manipulations - vocabulary simplification, rhythm flattening, imagery removal, voice genericization, structure simplification, and combined degradation - and reevaluate each version. Vocabulary simplification causes the smallest quality loss (0.41 +/- 0.46 points), far below structure (2.78) or voice (2.34) loss. Combined degradation is devastating (-5.64) but subadditive. An exploratory comparison with Qwen QwQ shows the same broad qualitative pattern.
在研究二中,我们通过对五篇经典散文段落进行系统性的“降级”处理来验证这一理论。我们应用了六种操作——词汇简化、节奏平淡化、意象移除、文风通用化、结构简化以及组合降级——并对每个版本进行了重新评估。词汇简化导致的质量损失最小(0.41 +/- 0.46 分),远低于结构(2.78 分)或文风(2.34 分)的损失。组合降级的影响是毁灭性的(-5.64 分),但具有次可加性。与 Qwen QwQ 的探索性对比显示,两者呈现出相同的宏观定性模式。
Together, these studies suggest that LLM judgments of writing quality are holistic, author-specific, and more sensitive to structural than lexical features, with implications for automated writing feedback and computational aesthetics.
综上所述,这些研究表明,大语言模型对写作质量的判断是整体性的、针对特定作者的,并且对结构特征的敏感度高于词汇特征。这一发现对自动化写作反馈和计算美学领域具有重要意义。