A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models

A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models

虚假评论检测综述:从预训练语言模型到大语言模型

Abstract: Online reviews shape consumer decisions, platform governance, and corporate reputation. Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms. The rise of large language models, or LLMs, has changed the problem in two ways. LLMs can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for detection.

摘要: 在线评论影响着消费者的决策、平台治理以及企业声誉。虚假评论通过向评分系统、推荐流程和公众信任机制中注入欺骗性证据,破坏了这一信息渠道。大语言模型(LLM)的兴起从两个方面改变了这一问题:一方面,LLM 可以生成流畅且具有上下文感知能力的欺骗性评论;另一方面,预训练语言模型(PLM)和 LLM 也为检测提供了更强大的语义表示。

This survey reviews fake review detection from an information fusion perspective, covering 211 studies published from 2018 to early 2026. We organize existing work by evidence source and fusion level, covering review text, sentiment, rating behavior, temporal metadata, user-product graphs, multimodal content, external knowledge, and LLM-generated features.

本综述从信息融合的角度回顾了虚假评论检测,涵盖了 2018 年至 2026 年初发表的 211 项研究。我们根据证据来源和融合层级对现有工作进行了梳理,涵盖了评论文本、情感、评分行为、时间元数据、用户-产品图、多模态内容、外部知识以及 LLM 生成的特征。

We trace the development from traditional machine learning and deep learning to PLM-based and LLM-based methods, and examine how different approaches combine textual, behavioral, structural, and multimodal data. We also analyze reported performance trends on widely used Amazon, Yelp, and OpSpam benchmark families, while noting the limitations caused by different label construction procedures, data splits, and evaluation metrics.

我们追踪了从传统机器学习和深度学习到基于 PLM 和 LLM 方法的发展历程,并考察了不同方法如何结合文本、行为、结构和多模态数据。我们还分析了在广泛使用的 Amazon、Yelp 和 OpSpam 基准系列上报告的性能趋势,同时指出了由不同的标签构建程序、数据划分和评估指标所带来的局限性。

Finally, we identify open problems in adversarial generation, cross-domain transfer, uncertainty-aware fusion, missing-source robustness, interpretability, and trustworthy evaluation for AI-generated deceptive content.

最后,我们指出了在对抗性生成、跨域迁移、不确定性感知融合、缺失源鲁棒性、可解释性以及针对 AI 生成的欺骗性内容进行可信评估等方面存在的开放性问题。


Paper Details:

  • Authors: Fanji Yang, Huiyao Chen, Xi Yu, Meishan Zhang, Xiaohong Xiao, Mingsen Deng
  • Affiliations: Guizhou University of Finance and Economics, Harbin Institute of Technology (Shenzhen), Guizhou University of Commerce
  • arXiv ID: 2609.30292
  • DOI: 10.1016/j.inffus.2026.104715

论文详情:

  • 作者: Fanji Yang, Huiyao Chen, Xi Yu, Meishan Zhang, Xiaohong Xiao, Mingsen Deng
  • 所属机构: 贵州财经大学、哈尔滨工业大学(深圳)、贵州商学院
  • arXiv ID: 2609.30292
  • DOI: 10.1016/j.inffus.2026.104715