Reviewing Model Collapse and Countermeasures

Reviewing Model Collapse and Countermeasures

模型崩溃及其对策综述

Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors. The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models. 摘要: 在海量网络规模数据的驱动下,生成式人工智能(GenAI)取得了显著进展,并在多个领域实现了广泛应用。GenAI 的进步促使从业者开始使用 AI 合成数据来训练下一代 AI 模型。

Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply. Unfortunately, it also introduces a new critical issue: in a self-consuming cycle between model and data, the model ultimately collapse, raising more trustworthiness concerns to GenAI. 不可否认,使用合成数据缓解了日益严峻的数据供应需求。然而不幸的是,这也引入了一个新的关键问题:在模型与数据之间的自我消耗循环中,模型最终会发生“崩溃”,从而引发了对 GenAI 可信度的更多担忧。

In recent years, increasingly more studies have investigated the phenomenon of model collapse (MC) and explored potential solutions to mitigate it. However, the review of the phenomenon of MC still remains blank. To fill this gap, this paper provides an up-to-date overview of these studies for consolidating and reviewing the progress of MC in different application scenarios and countermeasures for mitigating MC. We also highlight challenges and future research opportunities. 近年来,越来越多的研究开始调查模型崩溃(Model Collapse, MC)现象,并探索缓解该问题的潜在解决方案。然而,针对 MC 现象的综述研究目前仍处于空白状态。为了填补这一空白,本文对相关研究进行了最新的概述,旨在整合并回顾 MC 在不同应用场景下的进展,以及缓解 MC 的对策。同时,我们也强调了当前面临的挑战及未来的研究机遇。


Paper Information:

  • Authors: Xihao Xie, Beichen Hu
  • Submission Date: 17 Jun 2026
  • Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
  • DOI: 10.48550/arXiv.2608.21366

论文信息:

  • 作者: Xihao Xie, Beichen Hu
  • 提交日期: 2026 年 6 月 17 日
  • 学科分类: 人工智能 (cs.AI);机器学习 (cs.LG)
  • DOI: 10.48550/arXiv.2608.21366