DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
DeflectBench:评估大语言模型修辞谬误生成能力的基准测试
Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text.
摘要: 与检测现有文本中的谬误这一相关问题相比,大语言模型是否可以通过提示词按需生成修辞谬误,以及当前的安全性后训练是否限制了这种行为,受到的关注较少。
We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels.
我们通过 DeflectBench 填补了这一空白。该基准测试评估了来自四个前沿模型的 23,990 条生成内容,涵盖了三种转移话题策略(“那又怎么说”谬误、人身攻击、红鲱鱼谬误)、七种提示词框架,以及跨越四个争议等级的 80 个论点。
Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings.
拒绝生成主要受请求结构而非论点内容的影响。在 80 个论点中,每个论点的拒绝率差异仅为 11 个百分点;然而,仅仅改变提示词框架就能使模型内部的拒绝率波动近 100 个百分点,而在明确的框架内切换所要求的谬误类型,则可导致拒绝率波动超过 80 个百分点。
An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it.
使用“教育性辩论教练”的提示词框架可以将所有四个模型系列的拒绝率降至接近零,但这种绕过限制的行为并非完全合规。模型通常会产生“标记合规”的结果,即在包含所要求操纵手段的同一回复中,明确指出该操纵手段的名称。
The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at this https URL.
这四个模型在拒绝、标记合规、软拒绝和完全合规方面的表现分布各不相同。代码和数据集已在以下网址发布。