Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models

Computer Science > Computer Vision and Pattern Recognition arXiv:2610.06945 (cs) [Submitted on 3 Oct 2026] Title: Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models Authors: Muhammad Atif Butt, Paweł Skierś, Joost Van De Weijer, Kamil Deja.

计算机科学 > 计算机视觉与模式识别 arXiv:2610.06945 (cs) [提交于 2026 年 10 月 3 日] 标题:超越线性表示假设:文本到图像模型中的非线性激活引导 作者:Muhammad Atif Butt, Paweł Skierś, Joost Van De Weijer, Kamil Deja。

Abstract: Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear visual transition: between sunny and stormy lies an intermediate weather state such as a sky with a few white clouds, not simply a weaker storm; between a caterpillar and a butterfly, the progression is not a caterpillar with continuously growing wings.

摘要:机械可解释性通常依赖于线性表示假设(LRH),该假设认为高层概念在激活空间中被编码为线性方向。然而,自然的视觉概念并不一定需要线性的视觉过渡:在晴天和暴风雨之间存在着中间天气状态,例如带有少量白云的天空,而不仅仅是较弱的暴风雨;在毛毛虫和蝴蝶之间,演变过程也不是毛毛虫长出不断生长的翅膀。

This raises the question of whether such true intermediate states are also represented nonlinearly by the model. Indeed, when we prompt text-to-image models directly for intermediate attributes, their activations rarely fall along the straight direction connecting the endpoints. Therefore, we propose KANSteer, which models concept traversal as a curve passing through its intermediate states.

这引发了一个问题:这些真实的中间状态是否也由模型以非线性方式表示。事实上,当我们直接提示文本到图像模型获取中间属性时,它们的激活值很少落在连接端点的直线上。因此,我们提出了 KANSteer,它将概念遍历建模为一条穿过其中间状态的曲线。

Seeking a representation that is both simple and interpretable, we propose to use Kolmogorov-Arnold Networks (KANs), which provide a one-dimensional coordinate whose learned functions define the trajectory. This allows the steering direction to vary along the concept while preserving an interpretable representation.

为了寻求一种既简单又可解释的表示方法,我们建议使用柯尔莫哥洛夫-阿诺德网络(KANs),它提供了一个一维坐标,其学习到的函数定义了轨迹。这使得引导方向可以在概念沿线上变化,同时保持可解释的表示。

Across several concepts and text-to-image diffusion transformers, we find that their activation trajectories substantially deviate from straight lines, and that KANSteer provide a closer fit and smoother traversal of intermediate attributes than linear steering.

通过对多个概念和文本到图像扩散变换器的研究,我们发现它们的激活轨迹与直线有显著偏差,并且与线性引导相比,KANSteer 提供了更紧密的拟合和更平滑的中间属性遍历。