Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
具有意义和表达替代方案灵活生成的语用推理计算模型
Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification.
摘要: 语用语言的使用需要对替代方案进行推理:即说话者可能选择的其他表达方式,或听者可能考虑的其他解释。因此,语用的形式化和计算模型必须明确对话者所推理的替代方案集合,而这通常是通过人工指定来完成的。
Here we propose a framework, ScAffolded Generative models for Explanation (SAGE), that combines the explanatory transparency of cognitive models with the generative flexibility of language models (LMs). SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives (e.g., their semantics, complexity, or typicality); and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis.
在此,我们提出了一个名为“解释性支架生成模型”(SAGE)的框架,它结合了认知模型在解释上的透明度与语言模型(LM)在生成上的灵活性。SAGE 将语用过程分解为三种模块:提议者(proposers),利用语言模型生成开放式的候选替代方案空间;评估者(evaluators),对这些替代方案进行评估(例如评估其语义、复杂性或典型性);以及选择者(selectors),执行基于规则的计算步骤,以实现认知驱动的任务分析。
We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner (M-)implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data.
我们在三个案例研究中评估了 SAGE,涵盖了语用生成与解释——指称表达生成、方式(M-)含义以及格莱斯(Gricean)会话含义。我们使用计算认知建模中既定的方法对 SAGE 模型进行了严格评估,包括消融实验、基线对比以及与人类数据的定量拟合。
Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.
在各项研究中,SAGE 模型均达到了高准确率,且表现往往优于基线模型。但组件层面的分析揭示了一种不对称性:语言模型提议者能够可靠地生成非常适合语用建模的替代方案,而语言模型评估者更擅长提供直觉判断,而非基于理论或形式化指标的判断。我们讨论了神经符号模型作为人类语用语言使用解释性方案的潜力与局限性。