A Survey on Vision Transformers: Design Pressures, Architectural Taxonomy, and Future Outlook
Vision Transformers (ViTs) have developed from a patch-based alternative to convolutional backbones into a diverse architectural family spanning tokenization, attention reformulation, hierarchy construction, hybrid design, token economy, and training-oriented regularization. As this literature has matured, the main que...