Synthetic Data Augmentation via Class-Conditional Latent Diffusion Models for Fundus Image Quality Assessment
Abstract
Reliable automated fundus image quality assessment (FIQA) is a critical prerequisite for large-scale diabetic retinopathy (DR) screening, yet clinical datasets typically suffer from severe class imbalance that undermines classifier performance on minority quality grades. We present a comprehensive comparative study of deep learning architectures for three-class FIQA good, usable, and poor under three training regimes: no augmentation (NOAUG), traditional data augmentation (TRADAUG), and synthetic augmentation (SYNTHAUG) generated by a novel class-conditional Latent Diffusion Model (LDM). The proposed LDM is conditioned on CLIP-encoded quality-grade text prompts, enabling the generation of semantically coherent synthetic fundus images from pure Gaussian noise without using any real image as input. A selective oversampling strategy substantially reduces class imbalance while fully preserving the fidelity of original training samples. Eight state-of-the-art backbone architectures EfficientNet-B0/B3, SWIN-T/S/B, and ResNet-50/101/152 are systematically benchmarked under identical experimental conditions. SYNTHAUG consistently outperforms both baselines: SWIN-B achieves 90.85% accuracy and κ = 0.8399, representing a gain of 2.16 percentage points over NOAUG. A systematic validation–test accuracy inversion observed under SYNTHAUG is consistent with improved generalisation rather than overfitting to the training distribution. Generative model quality is assessed using FID, IS, LPIPS, SSIM, and PSNR. Our results establish a strong evidence base for LDM-based synthetic augmentation as a practical and effective approach to mitigating class imbalance in medical image analysis.