Multimodal Deep Learning For Eye Disease Recognition From Fundus And OCT Images With Synthetic OCT Generation
Abstract
Multimodal learning that integrates Color Fundus Photography (CFP) and Optical Coherence Tomography (OCT) can enhance retinal disease recognition by combining surface appearance with depth-resolved structural cues. However, the limited availability of synchronized 2D-3D CFP-OCT datasets hinders the deployment of multimodal AI in real-world screening settings where OCT devices are often unavailable. This study investigates whether synthetic 3D OCT volumes can serve as a practical auxiliary modality to support multimodal classification in fundus-only scenarios. We synthesize 3D OCT from fundus images using three generative paradigms-a 3D GAN (DCGAN/WGAN-GP style), a 3D VAEGAN, and a fundus-conditioned latent diffusion model (3D-DDPM with classifier-free guidance)-and integrate the generated volumes into an EyeMoST- based uncertainty-aware fusion framework. We evaluate the pipeline on RFMiD and ODIR-5K across multiple backbone configurations. Empirically, diffusion-generated OCT yields the most stable fusion performance. In contrast, synthetic OCT does not consistently outperform strong fundus-only models due to domain gap and imperfect biomarker preservation.