Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

大语言模型辩论行为基准测试:针对人身攻击防御策略的研究

Abstract: Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors.

摘要: 大语言模型(LLMs)正越来越多地被部署为说服性对话中的辩论主体,因此有必要对其相对于人类对话者的辩论能力进行严格评估。

In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content.

在本研究中,我们重点关注人身攻击(ad hominem arguments)。这类论证传统上被视为谬误,但在政治说服性对话中却起着关键作用,因为在这些语境下,人格(ethos)的影响力往往与命题内容不相上下。

Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game.

具体而言,我们调查了现代大语言模型是否能够复制人类在战略性使用和应对此类攻击时的能力。我们分析了一个自然语言政治对话语料库,以识别人类对话者在以人格为中心的辩论中自然采用的防御策略,并将其构建为一种对话博弈模型。

Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters’ repertoire of defensive strategies with those of artificial agents.

在实证方面,我们将大语言模型生成的对话与美国总统辩论的“ElecDeb60to16-fallacy”语料库进行了基准对比,分析了人类辩手与人工智能体在防御策略库上的差异。

Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse.

研究结果显示出显著差异:大多数大语言模型僵化地优先考虑逻辑防御,未能将人格反击作为政治话语中的有效手段加以利用。

We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.

我们认为,当前的安全微调限制了这些大语言模型的战略行动空间,使其无法在那些将人格质疑视为规范性预期而非单纯谬误的领域中,进行充分的自然化互动。