Skip to content

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

This work replaces attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect.

Abstract

In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.

View source

Similar papers

Preprint Sep 2026

Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning

Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most succ...

A. Fuller, Scott C. Lowe, Daniel G. Kyrollos et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
#natural language process... Preprint Sep 2026

MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation

Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns...

Mehdi Jafari, Hao Xue, F. Salim · 0 citations
#natural language process... Preprint Sep 2026

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing via gated addition. Inserted between the lower and upper halves of a 6-layer transformer...

E. Hering · 0 citations
#natural language process... Preprint Oct 2026

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fix...

Zhuo-Kun Chen, Xi Lin, Xi-Yu Wu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.