Skip to content

Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four Approaches.

Sep 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP · 0 citations
Medicine

Abstract

Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a systematic comparison of four PE approaches-Learned, Sinusoidal, Rotary Position Embedding (RoPE), and a 1D-ALiBi-style linear-bias variant-together with a targeted 2D-ALiBi-style diagnostic intervention motivated by a raster-distance mismatch diagnosis. ViT-Base models are trained on three low-to-intermediate data-regime datasets (CIFAR-100, TinyImageNet, and ImageNet-100), using a primary cross-method matrix and a canonical paired protocol for the targeted 2D-vs-1D ALiBi-style comparison. We apply a two-track diagnostic suite: embedding-space analyses (per-dimension variance, entropy, PCA, and probes) for additive PE, and attention-space intrinsic analyses (bias-tensor rank and entropy, slope schedule, and RoPE short-wavelength band counting) for attention-space PE, combined with attention-position mutual information (MI), noise ablation, and PE removal. The results reveal four qualitatively distinct encoding regimes: Learned PE is "quiet and ubiquitous," Sinusoidal PE "loud and structured," RoPE attains the highest accuracy with a graceful mid-depth MI decay, and the ALiBi-style variant exhibits a persistent attention-space bias. Direct attention-space analysis shows that the 1D-ALiBi-style bias tensor has full intrinsic rank. PE removal reveals a ${\sim }10\times$ dependency spectrum-from Sinusoidal collapse to near-chance to Learned PE retaining most of its accuracy-indicating extensive implicit positional learning. Linear probes show that Sinusoidal PE gives $0\%$ held-out column accuracy-a protocol-level generalisation failure for the modulo-column label rather than absence of column information-while Learned PE partially recovers 2D structure. Identifying the ALiBi-style raster-scan distance $|i{-}j|$ as mismatched to the 2D geometry of image patch grids, we introduce 2D-ALiBi-style, which replaces it with the 2D Euclidean patch-grid distance. On the canonical CIFAR-100 paired cohort ($n{=}12$ seeds), fixed-slope 2D-ALiBi-style yields a statistically reliable $+0.45$ pp accuracy improvement ($p{=}0.013$), is directionally positive on TinyImageNet, and shows no statistically reliable difference on ImageNet-100. It also lowers CLS-excluded patch-only attention-position MI. The mean-magnitude-matched 2D arm reverses this MI effect and reaches $70.28\pm 0.35\%$, the highest accuracy among the three canonical ALiBi-style arms; a secondary within-2D paired contrast ($+1.94$ pp, $12/12$ positive) establishes bias magnitude as a substantial lever without supporting a magnitude-rather-than-geometry interpretation. We therefore frame the contribution as a controlled characterisation plus a modest, statistically supported intervention, not a state-of-the-art PE method. Overall, PE choice affects accuracy, robustness, and attention-space information flow beyond what accuracy alone captures.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.