Skip to content
Open access

Integrating Local and Global Feature Learning for Retinal Disease Detection: An EfficientNetB3-ViT Hybrid Model

Sep 2026 · Fırat Üniversitesi Mühendislik Bilimleri Dergisi · 0 citations · 30 references

Abstract

Undiagnosed ocular conditions in initial phases can progressively impair vision, potentially leading to critical visual dysfunctions. Fundus imaging has enabled the identification of several retinal disorders, notably diabetic retinopathy, glaucoma, and age-related macular degeneration. However, manual evaluation of these images is both time-consuming and dependent on expert judgment. In this context, artificial intelligence-based automatic diagnosis systems have become an important need in the field of medical image analysis. In this study, deep learning-based models were compared for multi-class eye disease detection from fundus images, and a unique hybrid model was proposed. In this study, current deep learning architectures, including EfficientNetB3, DenseNet-121, ResNet-50, MobileNetV2, and ViT, were employed. All models were independently trained and tested, and their effectiveness was evaluated using commonly employed classification metrics. Additionally, a hybrid model based on EfficientNetB3 and ViT was designed, combining the strengths of Transformer and CNN architectures. All models were subjected to hyperparameter optimization using the GridSearch method, allowing for a fair comparison. The results obtained showed that the hybrid model successfully learned both local and global features in fundus images and exhibited higher performance compared to other models. In this context, the developed system has the potential to be integrated into future clinical decision-support systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.