You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
You Really Didn’t Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
你真的没听懂吗?针对中文网络评论中含蓄与戏谑语用推理的基准测试
Abstract: Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic categories, leaving open whether models can distinguish plausible readings of what a naturally occurring comment is doing in a particular exchange.
摘要: 中文网络评论往往通过含蓄和戏谑的语言传达社交含义,若脱离语境则难以解读。现有的评估方法大多围绕预定义的现象或受控的语用类别来组织项目,这使得模型是否能够辨别自然发生的评论在特定交流中所表达的合理含义,仍是一个悬而未决的问题。
We introduce a benchmark for evaluating whether LLMs can recover such situated pragmatic meanings. From more than 200,000 public Chinese social media interaction records, we construct 4,735 human-validated diagnostic items, each pairing a target comment with reconstructed preceding context and plausible misreadings.
我们引入了一个基准测试,旨在评估大语言模型(LLM)是否能够还原此类情境化的语用含义。我们从超过 20 万条公开的中文社交媒体互动记录中,构建了 4,735 个经人工验证的诊断项目,每个项目都将一条目标评论与重构的前置语境及合理的误读选项进行配对。
We evaluate eight LLMs as both question writers and solvers in a cross-writer setting. The task is challenging: the strongest model achieves 81.42% leave-writer-out accuracy. Across all eight models, the mean leave-writer-out accuracy is 68.70% while human accuracy was 90.8%. Case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.
我们在跨作者设置下评估了八个大语言模型,让它们分别担任出题者和解题者。这项任务极具挑战性:表现最强的模型在“留一作者”测试中达到了 81.42% 的准确率。在所有八个模型中,平均“留一作者”准确率为 68.70%,而人类的准确率则为 90.8%。案例分析表明,模型通常能够识别出大范围的讽刺或戏谑,但在识别具体的语用机制或互动意图时往往会出现偏差。