Skip to content
Preprint

From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does, and a two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does.

Abstract

Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carried token state that is non-participating at the attention boundary. We assign each token hidden carrier \(h_i\) an active-presence coefficient \(p_i=\lVert h_i\rVert^2/(\tau+\lVert h_i\rVert^2)\). The same coefficient has two roles: it gates information emitted by token \(i\), and it determines the mass with which token \(i\) enters computations shared with other tokens. OAttention is the support-coupled attention realization of this rule. It gates the receiver output by \(p_i\) and weights source \(j\) by \(p_j\) in both the attention numerator and partition, while retaining the standard score, visibility relation, exponential competition, and value aggregation. This makes the zero-vector token a zero element and yields exact null-receiver, null-source insertion, self-attention insertion, and empty-support properties. The same token-level presence gives local O-components (OFFN, ONorm, and OInject), presence-weighted OStandardize, the O-Closure law \(M(H\oplus0)=M(H)\oplus0\), and an OTransformer by residual and compositional closure. The canonical operator is checked by contract tests and a GPU evaluation. In a zero-fine-tuning retrofit of a cloned pretrained TabPFN v3 regressor, calibrated hidden-carrier OAttention and Full-O variants change mean RMSE by $+0.088\%$ and $+0.177\%$, respectively, over 18 matched dataset--seed cases. A two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does. These are scoped tests of exactness, active-path compatibility, and compositional necessity; they do not establish universal no-loss, arbitrary-host closure, learned attraction to the origin, or a general semantics for missing values.

View source

Similar papers

#machine learning Preprint Sep 2026

Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces

In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the...

N. Mysore · 0 citations
Preprint Aug 2026

L\'evy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

This work shows the attention layer itself can close that gap between deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps: with the right stochastic formulation, the pass that makes each prediction also reports how far it should be trusted.

S. Chatzis, Loukas Papadoulas · 0 citations
Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations
#natural language process... Preprint Sep 2026

Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context

Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which laye...

Mu-Yu He, Yu-Chen Liu, Ran Tao et al. · 0 citations
Preprint Aug 2026

TANGO: Treating Tokens as Operators

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO compu...

Joshua Nunley · 0 citations
Preprint Aug 2026

Toward a First-Principles Update Geometry for the Language-Model Head

Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional change. Softmax removes shared logit shifts, whereas the spectral norm can assign arbitraril...

Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.