Skip to content

TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

The TANGO model (Token-Aggregated Nonlinear Gating Operators), which replaces these two sublayers with one cross-token gated residual update, obtains the lowest mean validation negative log-likelihood on FineWeb-Edu, Lean, and DeepMind Mathematics, although it has the largest analytical forward-pass operation count.

Abstract

A standard Transformer block separates cross-token interaction in self-attention from a nonlinear feed-forward network applied independently at each position. We introduce the TANGO model (Token-Aggregated Nonlinear Gating Operators), which replaces these two sublayers with one cross-token gated residual update. Each source token produces a SwiGLU gate vector. Query-key similarities determine a weighted average of source gates for each destination, and the resulting gate rescales projected destination features. TANGO assigns a separate weight to every causally visible source and is quadratic in sequence length. The WANGO model (Windowed Aggregation of Nonlinear Gating Operators) retains the same unnormalized scores within a recent window and uses positive feature-map prefix statistics for older sources, giving linear sequence-length complexity for fixed window and feature dimensions. We compare TANGO and WANGO with Recurrent and Untied Transformer++, full-attention GAU, and FLASH. All models have approximately 44.3M nonembedding parameters and are trained in three matched runs. TANGO, WANGO, and Recurrent Transformer++ apply one shared block four times; the other architectures use four independent blocks. TANGO obtains the lowest mean validation negative log-likelihood on FineWeb-Edu, Lean, and DeepMind Mathematics, although it has the largest analytical forward-pass operation count. WANGO obtains the lowest mean FineWeb-Edu NLL among the architectures with computation linear in sequence length and outperforms Recurrent Transformer++ at nearly the same analytical forward-pass multiply-accumulate count.

View source

Similar papers

Jul 2026

Convolution for Large Language Models

These results support depthwise convolution as a lightweight complement to self-attention for modeling short-range token interactions and suggest that the convolution makes repeated token IDs more sensitive to their immediate context.

Yu-Chuan Tian, Yingte Shu, Wei He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation

Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling, provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and...

Ke-Wei Li, Rong Zhang, Xuelin Wang et al. · 0 citations
Preprint Aug 2026

DeltaFlow: Noise-Adaptive Bidirectional Gated Delta Networks for Embedded Language Flows

DeltaFlow is a promising alternative to dense attention for efficient continuous language denoising and noise-adaptive memory control and scheduled Temporal State Consistency to stabilize hidden representations across nearby noise levels are introduced.

Guang-Fu Guo, Xiao-Qian Lu, Linsey Pang et al. · 1 citation
#natural language process... Preprint Sep 2026

Line-Coupled Language Model

Autoregressive language models generate one token per decoding step, limiting the useful output of each forward pass. Although diffusion models, insertion-based decoding, and multi-token prediction enable parallel generation, they either incur additional training-time token traffic or struggle to predict strongly depen...

Shi-Yuan Li, Shao-Rong Zhang, Zhaorui Yang et al. · 0 citations
#machine learning Preprint Sep 2026

Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss

Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level text spans. To address this, we present an information-weighted cross-entropy loss that rescales token-level contributi...

Zhi-Jian Li, Stefan Larson, Kevin Leach · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.