Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

基于 GMM 和 LLM 的目标数据增强技术实现不平衡数据聚类

Abstract: In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics.

摘要: 在自然语言处理(NLP)领域,处理代表性不足的主题是一项挑战,尤其是在无监督任务中,聚类算法往往无法充分捕捉到这些少数派主题。

To tackle this challenge, our paper presents a novel unsupervised data augmentation method that integrates Gaussian Mixture Models (GMMs) and Large Language Models (LLMs).

为了应对这一挑战,本文提出了一种新颖的无监督数据增强方法,该方法结合了高斯混合模型(GMM)和大语言模型(LLM)。

Due to their flexibility and robustness, GMMs can detect clusters corresponding to underrepresented areas in the data, while LLMs create synthetic documents to enrich these clusters and improve their representation.

得益于其灵活性和鲁棒性,GMM 能够检测出数据中代表性不足区域所对应的聚类,而 LLM 则负责生成合成文档,以丰富这些聚类并提升其表征能力。

Experiments on various imbalanced text datasets demonstrate that our approach preserves clustering performance in all cases and often enhances cluster interpretability, offering a robust and scalable solution for improving data representation in unsupervised NLP tasks.

在多种不平衡文本数据集上的实验表明,我们的方法在所有情况下都能保持聚类性能,并经常能增强聚类的可解释性,为改善无监督 NLP 任务中的数据表征提供了一种稳健且可扩展的解决方案。


Paper Details:

  • Authors: Noor Khalal, Abdallah Alaa-Eddine Djamai, Imed Keraghel, Mohamed Nadif
  • Submission Date: 19 May 2026
  • Journal Reference: Advances in Intelligent Data Analysis: 23rd International Symposium on Intelligent Data Analysis; IDA 2025; Proceedings; pp 246-260
  • DOI: 10.48550/arXiv.2607.28635

论文详情:

  • 作者: Noor Khalal, Abdallah Alaa-Eddine Djamai, Imed Keraghel, Mohamed Nadif
  • 提交日期: 2026 年 5 月 19 日
  • 期刊参考: 《智能数据分析进展:第 23 届智能数据分析国际研讨会;IDA 2025;会议论文集;第 246-260 页》
  • DOI: 10.48550/arXiv.2607.28635