A unified multi-modal latent diffusion framework with modality-dropout training
Abstract
Introduction Multi-modal conditioning in latent diffusion models—combining text, structural, and spatial guidance signals—substantially improves controllable image synthesis, yet two limitations persist across most existing frameworks. First, conditioning modalities are fused using fixed architectural weights that remain constant regardless of the input, preventing any input-adaptive rebalancing of guidance. Second, most existing multi-modal approaches are designed under the assumption that all conditioning signals will be present at inference time, leaving the model without any way to degrade gracefully when some modalities are unavailable. This paper introduces a unified latent diffusion framework that addresses both limitations simultaneously. Methods Our architecture integrates three complementary conditioning modalities—semantic text embeddings via CLIP, structural boundary maps derived from Canny edge detection, and regional semantic layout from segmentation masks-as co—equal inputs into a single U-Net denoising backbone. The framework has two core components: Softmax-Normalized Adaptive Guidance Fusion (SNAGF), which replaces fixed fusion weights with three learnable, softmax-normalized modality importance scores optimised end-to-end alongside the diffusion objective; and Modality-Dropout Training (MDT), a structured regularisation strategy that randomly zeroes individual guidance representations with probability pdrop = 0.3 during training, training the model to produce coherent outputs under any available subset of conditioning inputs. Results Evaluated on CIFAR-10, the full SNAGF+MDT framework achieves SSIM 0.93, FID 17.2, PSNR 34.1 dB, CLIP Score 0.88, LPIPS 0.09, and IS 10.4. MDT alone yields an average 15.5% FID improvement across partial-conditioning inference configurations, and learned SNAGF weights converge to λtext = 0.370, λedge = 0.363, λseg = 0.267, demonstrating stable convergence from epoch 30 onward. Results are averaged over three random seeds. Discussion Taken together, they indicate that adaptive modality weighting outperforms fixed conditioning strategies and that modality-dropout training provides an effective partial-conditioning robustness mechanism. All findings are established at 32 × 32 resolution on a single benchmark with algorithmically constructed conditioning signals; validation on higher-resolution datasets with naturally paired text and structural annotation remains necessary before these conclusions can be generalised.