Optimizing multi-dimensional retinal segmentation: a performance-practicality analysis
Abstract
This study introduces a multi-dimensional framework to systematically evaluate the generalization and computational efficiency of deep learning models for retinal vessel segmentation. Using the DRIVE and STARE datasets, this research evaluates the performance of SegNet, U-Net, and DeepLab across in-domain, cross-domain, and mixed-domain scenarios. Our experiments identify an asymmetric generalization phenomenon, where knowledge transfers effectively from STARE to DRIVE, but encounters a performance bottleneck in the opposite direction. Architectural analysis reveals that SegNet is notably more robust to domain shifts than U-Net; it exhibited a more stable transition between datasets, with a performance degradation (Δshift) of only 0.1603, compared to U-Net’s 0.2344. However, under mixed-dataset training, U-Net leverages its superior hierarchical feature extraction to achieve the highest Dice score. By integrating analysis of parameters, FLOPs, and latency, this research defines key evaluation criteria for clinical use, demonstrating that lightweight models often offer the optimal balance between performance and operational efficiency.