A Survey on Vision Transformers: Design Pressures, Architectural Taxonomy, and Future Outlook
Abstract
Vision Transformers (ViTs) have developed from a patch-based alternative to convolutional backbones into a diverse architectural family spanning tokenization, attention reformulation, hierarchy construction, hybrid design, token economy, and training-oriented regularization. As this literature has matured, the main question is no longer whether Transformers work for vision, but how the ViT design space should be interpreted. This survey presents a pressure-driven and trade-off-aware synthesis of Vision Transformer architectures. The central thesis is that ViT evolution is best understood through six recurring design pressures: representation, interaction, structure, inductive bias, efficiency, and optimization/generalization. These pressures are mapped to six corresponding taxonomy axes: 1) tokenization and representation formation; 2) attention operator reformulation; 3) hierarchical and multi-scale feature construction; 4) convolution-Transformer hybridization; 5) efficiency, compression, and token economy; and 6) training strategies and structural regularization. Beyond categorization, the survey derives cross-axis design principles that explain how token quality, interaction structure, hierarchy, inductive bias, computational economy, and training support constrain one another. It also distinguishes nominal efficiency, such as lower asymptotic complexity or reduced arithmetic workload, from realized efficiency in terms of memory behavior, latency, throughput, hardware regularity, and conditional-execution overhead. The analysis connects these design strategies to task requirements and clarifies why architecture and training methodology must be interpreted jointly. Beyond classification-oriented narratives, the survey situates ViTs within foundation-scale visual representation learning, vision-language and multimodal systems, diffusion models, embodied artificial intelligence, and neuroscience-inspired visual computation. This perspective explains why major ViT directions emerged, when their trade-offs are favorable, and which design tensions remain central to scalable, efficient, and transferable visual representation learning.