Skip to content
Review

The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains

Sep 2026 · Expert systems · Vol 43 · 0 citations · 191 references

TL;DR

This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design, and critically examines how different mechanisms influence scalability, computational efficiency, and task‐specific accuracy.

Abstract

The rapid evolution of deep learning has positioned Vision Transformers (ViTs) as a powerful alternative to convolutional neural networks (CNNs) in computer vision. By leveraging self‐attention to model global dependencies, ViTs achieve state‐of‐the‐art performance across tasks such as image classification and segmentation. This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design. We critically examine how different mechanisms, such as sparse and linear attention, influence scalability, computational efficiency, and task‐specific accuracy. A dedicated analysis of hybrid CNN‐Transformer architectures reveals how they effectively reintroduced crucial inductive biases, such as locality and translation equivariance, to mitigate the data inefficiency and slow convergence issues inherent in pure ViTs. Furthermore, we extend our review to specialized domains, including 3D analysis, video inpainting, and low‐level vision, correlating architectural choices with performance gains. Here, we show that while ViTs excel at capturing global context, their practical deployment is often constrained by quadratic computational complexity and substantial data requirements. By synthesizing recent breakthroughs and critically evaluating performance trade‐offs, this survey not only provides a structured reference for researchers but also identifies open challenges and delineates promising directions for future research, particularly in developing data‐efficient and resource‐aware Transformer models for real‐world visual computing systems.

View source

Similar papers

Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations
Review Open access 2026

A Survey on Vision Transformers: Design Pressures, Architectural Taxonomy, and Future Outlook

Vision Transformers (ViTs) have developed from a patch-based alternative to convolutional backbones into a diverse architectural family spanning tokenization, attention reformulation, hierarchy construction, hybrid design, token economy, and training-oriented regularization. As this literature has matured, the main que...

C. Lee, E. Ooi, Nicole Kai Ning Loh et al. · 0 citations
Open access Sep 2026

A data efficient pyramid vision transformer for image classification

A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.

Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al. · 0 citations
Open access Aug 2026

Flow-Guided Neural Pruning: Signal-Flow Framework for Multi-Architecture Model Compression

A new iterative pruning algorithm is proposed, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to identify and eliminate non-essential parameters while preserving critical information pathways in deep neural networks.

A. Samarin, Artem A. Nazarenko, E. Kotenko et al. · 0 citations
Conference Open access Sep 2026

Evolution of Few-Shot Image Classification: A Survey from Metric Learning to Large Model Fine-Tuning

This survey comprehensively evaluates the underlying mechanisms, inherent strengths, and specific weaknesses of each FSIC methods into three categories: Metric Learning, Optimization/Meta-Learning, and Transfer Learning with Large Model Fine-Tuning.

Yu-Hang Ji · 0 citations

A Patch-Masking Approach to Explain CNNs and Vision Transformers through Class-Specific Impacts

A new method is presented for explaining CNNs and ViTs image classification through a set of interpretable rules composed of one or more antecedents, which combine pixel-level properties and patch-level features, analyzing the impact of each region on the model’s classification.

Jean-Marc Boutay, Damian Boquete, Deniz Köprülü et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.