Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
超越模式坍缩:通过生成流网络(GFlowNets)生成多样化的合成专家对话
Abstract: High quality synthetic data is central to post-training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end-use scenarios yields low diversity data that collapses onto dominant modes.
摘要: 高质量的合成数据对于大语言模型(LLM)的训练后阶段至关重要,它能支持自适应人工智能应用,从而体现对话中多样化的专家策略与决策。直接提示(Prompting)大语言模型或基于最终使用场景对其进行条件化处理,往往会产生多样性较低的数据,并导致模型坍缩至主导模式。
We propose a method to generate diverse high-quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate latent conversation structure using a Gaussian mixture density over key interaction features (e.g., confusion episode dynamics, scaffolding directive balance) enables sampling expert strategies in proportion to their prevalence in the training data.
我们提出了一种利用生成流网络(GFlowNets)生成多样化高质量合成数据的方法。研究表明,通过在关键交互特征(如困惑片段动态、支架指令平衡)上使用高斯混合密度来训练 GFlowNets 生成潜在对话结构,能够按专家策略在训练数据中的占比进行采样。
Across two structurally distinct domains, tutoring and emotional support dialogues, our GFlow-based synthetic data generation approach offers a better balance of fidelity, mode coverage, and authenticity than reinforcement-learning and end-to-end LLM baselines, without copying training data.
在辅导和情感支持对话这两个结构迥异的领域中,与强化学习和端到端大语言模型基准相比,我们基于 GFlow 的合成数据生成方法在保真度、模式覆盖率和真实性之间取得了更好的平衡,且无需复制训练数据。
Evaluated on three downstream outcome prediction tasks, classifiers trained on synthetic GFlowNet-generated conversations provide a stronger training signal than competitive synthesis baselines.
在三项下游结果预测任务的评估中,使用 GFlowNet 生成的合成对话进行训练的分类器,比竞争性的合成基准提供了更强的训练信号。