MergeNeXt: a hybrid CNN–transformer model for retinal OCT image classification
Abstract
Retinal disease recognition is a critical component of computer-aided diagnosis (CAD). Despite advancements, existing models often struggle to capture both fine-grained local details and complex global relationships in optical coherence tomography (OCT) images. This study aims to develop a high-precision deep learning architecture to improve automated classification accuracy for retinal lesions. This study proposes MergeNeXt, an innovative hybrid model that integrates the strengths of convolutional neural networks (CNNs) and Transformers. The architecture leverages two key components: ConvMerge Block (CMB), which uses a parallel branch structure to independently model spatial and channel features, and Adaptive Double-Channel Attention (ADCA), which employs a dual-pooling strategy to adaptively adjust features. The model was evaluated on retinal OCT datasets and compared against state-of-the-art architectures, including ConvNeXt and RepViT. MergeNeXt achieved an accuracy of 99.58%, a precision of 99.58%, a recall of 99.56%, and an F1 score of 99.57% on the institutional dataset. Five-fold cross-validation yielded an average accuracy of 99.25 ± 0.34%, while evaluation on the public OCT-C8 dataset achieved an accuracy of 98.29%. MergeNeXt contained 27.57 M parameters and required 6.6 ms per image under the reported hardware configuration. It also achieved the highest point estimates across the reported classification metrics among the evaluated models. These findings support the effectiveness, stability, architecture-level cross-dataset generalizability, and computational efficiency of MergeNeXt for retinal OCT image classification. However, prospective and task-matched multicenter and multi-device validation remains necessary before its clinical utility can be established.