ExploreNet: Learning Where to Explore in Diffusion GRPO
ExploreNet: Learning Where to Explore in Diffusion GRPO
Abstract: Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally.
摘要: 诸如 Flow-GRPO 之类的组相关强化学习(Group-relative RL)方法,通过在每个去噪步骤中添加各向同性高斯噪声来进行探索,从而对图像生成器进行后训练。这种噪声决定了模型从哪些展开(rollouts)中学习,但它对潜在空间(latent)的每个通道和空间位置的扰动是均等的。
In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution.
在本文中,我们指出潜在空间的元素对生成图像的影响程度各不相同,因此探索策略应当适应这些差异。我们引入了 EXPLORENET 来学习一种自适应的探索分布。
EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged.
EXPLORENET 是一种策略模型,它在观察到任何奖励之前,根据当前的潜在向量、去噪步骤和提示词(prompt),为每个潜在元素预测噪声尺度;该模型通过每个展开组的奖励分布进行训练,并在训练完成后被丢弃,因此不会改变推理过程。
On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate.
在 Stable Diffusion 3.5 Medium 模型上,EXPLORENET 在留出的 GenEval2 测试集上比 Flow-GRPO 提升了 14%,并成功迁移到两个独立的组合性基准测试以及五个偏好和图像质量模型中,达到了 67.2% 的人类偏好胜率。
Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.
总的来说,通过我们的组相关扩散强化学习实验,我们发现探索过程是可以学习的,探索分布的形状比其幅度更为重要,且展开的质量比展开的数量更为有效。