Skip to content
Preprint

Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

Role-Decoupled Attention Residuals (RD-AttnRes), a minimal extension that shares one depth route between queries and keys while learning an independent value route over the same residual sources, is introduced and suggests that, within the evaluated training regime, attention matching and content retrieval benefit from distinct reads over the residual hierarchy.

Abstract

Depth-routing residual architectures allow Transformer layers to retrieve earlier representations instead of inheriting only the immediately preceding state. Existing Block Attention Residuals, however, use a single content-dependent depth mixture to construct the inputs to queries, keys, and values. This design couples two functionally different decisions: queries and keys determine where attention matches, whereas values determine what content is retrieved. We therefore ask whether matching and content retrieval should be forced to read from the same depth. We introduce Role-Decoupled Attention Residuals (RD-AttnRes), a minimal extension that shares one depth route between queries and keys while learning an independent value route over the same residual sources. Tying the two routing queries exactly recovers the parent architecture, while decoupling them adds only one model-width vector per layer and introduces no additional token-to-token attention operation. We evaluate RD-AttnRes using a frozen, paired pretraining protocol on FineWeb-Edu with five matched seeds for both 120M- and 343M-parameter models and a 2.0B-token training budget. RD-AttnRes improves validation negative log-likelihood in all 10 matched comparisons. The mean reductions are 0.0301 and 0.0247, corresponding to perplexity reductions of 2.97 percent and 2.43 percent at 120M and 343M parameters, respectively. Early-budget controls indicate that neither the additional parameter count, duplicated routing execution, nor a fixed value route reproduces the improvement. Routing diagnostics further reveal persistent divergence between the query-key and value depth distributions. These results suggest that, within the evaluated training regime, attention matching and content retrieval benefit from distinct reads over the residual hierarchy.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

This work introduces SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction and consistently reduces attention-reconstruction error across four heterogeneous video generation and world models.

P. Taghavi, Reza Langari, Gaurav Pandey · 1 citation
#artificial intelligence Preprint Sep 2026

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

Semantic-Geometric Decoupled Routing is proposed, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances and yields an explicit closed-form block routing score without token-level search...

Xin-Wei Long, Wei-Gao Sun, Wei-Bo Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
#machine learning Preprint Aug 2026

LoGo: Token-Level Dynamic Local-Global Attention

LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.

Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al. · 1 citation
#artificial intelligence Preprint Oct 2026

LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing princip...

Zhao-Hui Wang, Zhi-Xin Pan, Fan-Xu Meng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention o...

Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.