Skip to content
Open access

Classification performance of Vision Transformer models on medical cancer image datasets: An investigation on a multi-image set

Sep 2026 · Gümüşhane Üniversitesi Fen Bilimleri Enstitüsü Dergisi · 0 citations · 24 references

Abstract

The purpose of this study is to systematically examine the performance boundaries and characteristic behaviors of six prominent Vision Transformer models (ViT, DEiT, Swin, BEiT, PVT, and CvT) in the deep learning landscape, with the goal of optimizing the classification accuracy and computational efficiency of medical cancer images. In the research, a total of six medical image datasets—comprising four original and two augmented histopathological and MRI datasets obtained from the Kaggle platform—were utilized. To establish a fair evaluation infrastructure, all models were trained on the same CUDA-supported hardware using rigorous regulation methods such as 5-fold stratified cross-validation, AdamW optimization, controlled learning rates, and early stopping, and were subsequently subjected to feature extraction and classification tests. While ViT delivered the highest accuracy on large datasets, the PVT model achieved a substantial time advantage by operating 43% faster. On the smallest dataset (BPH Raw), ViT and DEiT exhibited full stability with 100% accuracy and 0.00 standard deviation, whereas the CvT and PVT models displayed higher fluctuation with lower performance. Data augmentation radically improved the models' performance; when the BPH dataset was expanded fivefold, all models reached 100% accuracy. Overall, the BEiT model lagged behind its competitors in terms of accuracy and variance values. On the other hand, although the CvT model outperformed pure transformer models on certain raw data, it remained the slowest model with the highest training times across all scenarios. Moving beyond standard accuracy-focused comparisons in the literature, this study offers a holistic guide that reveals the direct relationship between dataset scale and model architecture alongside a time/cost matrix. The most unique contribution of this work is that it defines clear architectural selection boundaries based on hardware constraints for the real-time integration of artificial intelligence into clinical decision support systems (ViT/Swin for maximum accuracy; PVT for speed-performance balance; DEiT for limited data).

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.