FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law

FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law

FourierQK:滤波器形状、容许性与泄露-覆盖定律

Abstract: Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties — DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage — using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT).

摘要: 频率坍缩注意力机制(Frequency-collapse attention)[Zeris, 2026e] 通过将 Q/K 点积替换为在学习频率下的带通滤波内积,在标准点积注意力机制的基础上取得了显著提升。一个自然的后续问题是:哪种滤波器形状效果最好,原因何在?我们通过在字符级语言建模(TinyShakespeare,6 层 GPT)上进行受控消融实验,测试了关于滤波器属性的五个假设——直流(DC)抑制、奈奎斯特(Nyquist)抑制、带宽、中心频率以及多尺度覆盖。

Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage — narrowband filters (gap > +4) are clean, wideband filters (gap < +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f].

我们的主要发现包括:(1) 直流和奈奎斯特分量具有明显的负面影响(val ~= 2.0,相当于相位随机化),这证实了振荡带通结构是必不可少的,而不仅仅是某种低维频谱摘要;(2) 最优的单尺度带宽为 sigma ~= 2 个 bin,中心位于段落尺度(约 70 个 token),相比 BASE-DOT 获得了 +1.15 nats 的纯增益;(3) 容许滤波器(零均值,墨西哥帽 DOG m = 2)在相同尺度下优于非容许高斯滤波器,并能针对双边 FFT 泄露提供部分保护;(4) 双边 FFT 泄露随频谱覆盖范围单调增加——窄带滤波器(间隙 > +4)表现良好,而宽带滤波器(间隙 < +2)则存在泄露;(5) 字符尺度的因果时域 Morlet 无法超越 BASE-DOT(K=128 个抽头覆盖了 T=256 上下文的 50%),这促使我们在配套的 MorletQK 论文 [Zeris, 2026f] 中进行词级实验。

Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT).

综上所述,发现 (1)-(5) 表明 FourierQK 在双向注意力设置(编码器风格,如 BERT)中非常有效,因为在这种设置下,训练和推理时均可获得全序列上下文;而自回归生成则需要像 MorletQK [Zeris, 2026f] 这样的因果频谱变体(解码器风格,如 GPT)。