Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

Persistence Forcing:在像素空间扩散模型中利用特征专业化

Abstract: Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations.

摘要: 像素空间扩散 Transformer (DiTs) 直接对高维视觉数据进行操作,但其隐藏表示通常在整个深度上经历均匀的细化过程。然而,自然图像本质上是以不同的粒度级别组织的。全局结构通常可以用紧凑的方式表示,而局部纹理和精细细节则需要更丰富的表示。

Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively.

受此启发,我们在像素空间 DiTs 中引入了异构细化机制,为不同特征组分配跨深度的不同细化预算。因此,出现了一种有序的特征专业化:稀疏细化的特征主要编码全局视觉结构,而更频繁细化的特征则逐渐专门化于局部的高频细节。我们将这两组特征分别称为“持久特征”(persistent features)和“活跃特征”(active features)。

Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent—active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details.

基于这种涌现的专业化,我们引入了 Persistence Forcing (PerF),它明确利用这种“持久-活跃”特征组织来进行像素空间图像生成。这使得持久特征能够持续调节活跃细化的特征,从而允许稳定的全局信息引导更精细视觉细节的持续细化。

During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.

在生成采样过程中,这种交互进一步诱导出一个有意义的引导方向,促进了全局结构的连贯性,并自然地补充了无分类器引导(classifier-free guidance)。在 ImageNet $256\times256$ 数据集上,PerF-L 达到了 $1.91$ 的 FID,仅用一半的参数就接近了 JiT-H 的 $1.86$;而 PerF-H 在 ImageNet $256\times256$ 和 $512\times512$ 上分别进一步达到了 $1.63$ 和 $1.76$ 的 FID。