Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

用于掩码语言模型的流形投影与迭代自编码器细化

Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity.

摘要: 在基于 Transformer 的掩码语言模型中,注意力机制是上下文混合的主要手段,但同时也存在其他跨 Token 的数据混合方式。近期出现的无注意力混合器(Attention-free mixers)通过固定或由超网络生成的 MLP 来替代注意力机制,通过交替使用动态的、依赖于内容的加权方式,以实现计算上的简化。

We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect.

我们构建了一种替代方案,通过低秩瓶颈自编码器实现了相同的特性。我们用一系列基于自编码器的混合模块取代了注意力机制,这些模块分别作用于局部邻域、完整序列以及注意力头之间。每个模块都通过瓶颈层对输入进行压缩和重构,其宽度是一个超参数,而非训练过程的产物。

In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder.

在掩码位置,我们引入了一种包含两个明确步骤的迭代细化过程。第一步是“拉取”(Pulling),将嵌入表示向其邻居的加权平均值拉近;第二步是“校正”(Correcting),通过自编码器将结果投影回已学习的流形空间。

Our architecture achieves a significant portion of attention’s performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.

当在 C4 数据集上进行预训练,并使用参数匹配的 BERT 基线进行评估时,我们的架构在 FLOPs 减少约 1.9 倍的情况下,实现了与注意力机制相当的性能。通过使用一种频率感知训练计划(在掩码任务中对稀有 Token 的采样频率高于均匀分布),我们的模型在最稀有 Token 频率区间上的表现与参数匹配的 BERT 和 TinyBERT 基线持平。