This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design, and critically examines how different mechanisms influence scalability, computational efficiency, and task‐specific accuracy.
Abstract
The rapid evolution of deep learning has positioned Vision Transformers (ViTs) as a powerful alternative to convolutional neural networks (CNNs) in computer vision. By leveraging self‐attention to model global dependencies, ViTs achieve state‐of‐the‐art performance across tasks such as image classification and segmentation. This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design. We critically examine how different mechanisms, such as sparse and linear attention, influence scalability, computational efficiency, and task‐specific accuracy. A dedicated analysis of hybrid CNN‐Transformer architectures reveals how they effectively reintroduced crucial inductive biases, such as locality and translation equivariance, to mitigate the data inefficiency and slow convergence issues inherent in pure ViTs. Furthermore, we extend our review to specialized domains, including 3D analysis, video inpainting, and low‐level vision, correlating architectural choices with performance gains. Here, we show that while ViTs excel at capturing global context, their practical deployment is often constrained by quadratic computational complexity and substantial data requirements. By synthesizing recent breakthroughs and critically evaluating performance trade‐offs, this survey not only provides a structured reference for researchers but also identifies open challenges and delineates promising directions for future research, particularly in developing data‐efficient and resource‐aware Transformer models for real‐world visual computing systems.
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
Vision Transformers (ViTs) have developed from a patch-based alternative to convolutional backbones into a diverse architectural family spanning tokenization, attention reformulation, hierarchy construction, hybrid design, token economy, and training-oriented regularization. As this literature has matured, the main que...
C. Lee, E. Ooi, Nicole Kai Ning Loh et al.· IEEE Access· 0 citations
A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al.· Discover Artificial Intellig...· 0 citations
A new iterative pruning algorithm is proposed, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways in deep neural networks.
A. Samarin, Artem A. Nazarenko, E. Kotenko et al.· Machine Learning and Knowled...· 0 citations
This survey comprehensively evaluates the underlying mechanisms, inherent strengths, and specific weaknesses of each FSIC methods into three categories: Metric Learning, Optimization/Meta-Learning, and Transfer Learning with Large Model Fine-Tuning.
A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.
Jean-Marc Boutay, Damian Boquete, Deniz Köprülü et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.