Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

评估提示词扰动对大语言模型偏见与幻觉的影响

Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination.

摘要: 大语言模型(LLMs)在各种自然语言处理任务中展现出了卓越的能力,使其被广泛部署为决策场景中的智能助手。然而,这些模型日益增加的复杂性引发了人们对其可靠性的担忧,特别是在偏见和幻觉方面。

In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models.

在这项工作中,我们评估了大语言模型在决策任务中对原始查询进行扰动变体后的鲁棒性。研究表明,与以往的研究相反,在某些大语言模型中,扰动反而能够比其他模型更有效地减轻偏见和幻觉。

It’s found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others.

研究发现,Claude 3 在大多数数据集所代表的任务中表现更为出色;而像 GPT-3.5 这样的模型则表现出不同程度的适应性,在某些情况下表现尚可,但在其他情况下则明显落后。

These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness.

这些见解对于理解在现实应用中部署基于大语言模型的助手作为有效决策支持工具的实际意义至关重要,并强调了进行严格测试和验证以确保其可靠性和有效性的必要性。

This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.

本研究为日益增长的大语言模型评估研究领域做出了贡献,并为在关键决策场景中开发更稳健、更值得信赖的 AI 助手提供了参考。


Paper Details:

  • Authors: Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon
  • Journal Reference: Neural Information Processing (ICONIP 2024), Lecture Notes in Computer Science (LNCS), vol. 15290, pp. 361-374, Springer, 2025
  • DOI: 10.1007/978-981-96-6588-4_25

论文详情:

  • 作者: Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon
  • 期刊引用: Neural Information Processing (ICONIP 2024), Lecture Notes in Computer Science (LNCS), vol. 15290, pp. 361-374, Springer, 2025
  • DOI: 10.1007/978-981-96-6588-4_25