Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

通过隐含意义识别与撤销评估大语言模型中的交际信念更新

Abstract: Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. 摘要: 人类语言由未言明的信念和信念更新所驱动,这使得它们对于大语言模型(LLM)与用户之间的成功交流至关重要。

In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance’s implied meaning is weakened or negated. 在本文中,我们评估了大语言模型识别通过隐含意义(implicatures)表达的未言明信念的能力,并考察了它们通过“隐含意义撤销”(implicature cancellation)理解信念更新的能力——这是一种语用现象,即话语的隐含意义被削弱或否定。

We create the first expert-annotated implicature cancellation dataset, [DatasetName], crowdsourced for human judgements of implicatures and their corresponding cancellations. 我们创建了首个经专家标注的隐含意义撤销数据集 [DatasetName],该数据集通过众包方式收集了人类对隐含意义及其相应撤销情况的判断。

We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. 我们发现,大语言模型在理解信念更新方面落后于人类,尤其是在更自然的场景中表现更为明显。

Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. 额外的对照实验表明,大语言模型在信念更新上的成功可能部分源于对先验信念的依赖,而信念更新的失败则可能取决于其类型和形式。

Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at this https URL. 总的来说,我们的研究表明,当前的大语言模型尚未达到人类对未言明信念和信念更新的理解水平。代码和数据可在该链接获取。