Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four Approaches.
Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a syste...