Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry
Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry
Neo-Classic:评估古典诗歌语言审美推理能力的基准测试
Abstract: While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that combines a constructionist Out-of-Sample (OOS) dataset with a suite of reverse understanding probes. Unlike traditional benchmarks that rely on verification or generation over historical corpora, Neo-Classic comprises strictly metrical poetry authored by contemporary experts, reducing the possibility of direct retrieval.
摘要: 尽管大型语言模型(LLMs)在现有的古典诗歌基准测试中取得了很高的准确率,但要区分模型是真正具备可迁移的“语言审美推理能力”,还是仅仅依赖于熟悉的预训练模式,仍然极具挑战。为了解决这一问题,我们推出了 Neo-Classic,这是一个结合了建构主义样本外(OOS)数据集与一系列反向理解探测任务的评估基准。与依赖历史语料库进行验证或生成的传统基准不同,Neo-Classic 包含由当代专家创作的严格格律诗,从而降低了模型直接检索的可能性。
We evaluate state-of-the-art models, including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2, across five behavioral probes designed to test hierarchical constraint satisfaction. Our results reveal two primary limitations. First, a performance gap of 20 to 50 percent emerges when models transition from historical to contemporary texts. Second, models exhibit substantial difficulties in discourse-level ordering tasks, with standard accuracy remaining low (0 to 13 percent).
我们评估了包括 Qwen3-Max、Gemini-3-Pro 和 DeepSeek-V3.2 在内的顶尖模型,通过五项旨在测试层级约束满足能力的探测任务进行了评估。研究结果揭示了两个主要局限性:首先,当模型从处理历史文本转向当代文本时,性能会出现 20% 到 50% 的下滑;其次,模型在篇章级的排序任务中表现出显著困难,标准准确率依然处于低位(0% 到 13%)。
Although expert-level guidance improves the performance of reasoning-enhanced models to 36 percent, a notable gap with human experts persists. These findings suggest that while current LLMs capture local formal patterns, they struggle with global hierarchical planning required for robust Linguistic-Aesthetic Reasoning.
尽管专家级的引导将推理增强型模型的性能提升到了 36%,但与人类专家相比仍存在显著差距。这些发现表明,虽然当前的 LLMs 能够捕捉局部的形式规律,但在实现稳健的语言审美推理所需的全局层级规划方面,它们仍面临挑战。