Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
用于蛋白质扩散模型测试时对齐的谱反馈方法
Abstract: Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections.
摘要: 针对离散扩散模型的奖励最大化对齐方法,主要集中于引导反向生成过程,通常通过影响标记(token)的逻辑值(logits)或在中间步骤选择有利的序列来实现。这些方法大多将推理视为单向过程,缺乏对不理想标记选择进行重新审视的机制。
We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process.
我们引入了“谱反馈”(Spectral Feedback),这是一种在反馈循环中选择编辑位置的算法,允许模型迭代地修正其自身的生成结果。该方法利用离散扩散模型的掩码结构,通过对标记进行重新掩码和重新采样,类似于图像编辑中重新引入噪声潜变量并重新运行反向过程的方法。
While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously.
虽然先前的对齐方法侧重于分配哪些标记标签以最大化目标奖励,但我们将“重新审视哪些标记”视为核心对齐问题。选择编辑位置具有挑战性,因为编辑效果是相互依赖的:修改一个标记的影响取决于同时被编辑的其他标记。
We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection.
我们将“编辑集”(edit-set)定义为一组需要重新掩码和重新采样的标记位置。受先前关于生物系统中稀疏相互作用研究的启发,我们通过实验发现,蛋白质反向折叠的编辑集价值函数具有稀疏的傅里叶表示。这种结构使得谱反馈能够高效地学习并优化用于选择编辑位置的价值函数。
Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process.
谱反馈与模型无关,可应用于预训练模型、测试时对齐模型以及微调后的扩散模型。对于所有这些模型,该算法在不修改底层生成过程的前提下,提升了对齐性能。
Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.
在结合蛋白质稳定性奖励预言机的反向折叠任务中,该方法使预训练模型的稳定蛋白质比例提高了 32.3%,Best-of-10 方法提高了 24.8%,而对于最先进的强化学习微调扩散模型,则提升了 5.8%。