SSDiffNet: Semantic-Guided and Structure-Consistent Diffusion for Aerial Visible-to-Infrared Image Translation
Abstract
Visible-to-infrared image translation in aerial scenarios is fundamentally challenged by severe cross-modal appearance discrepancies, complex land-cover semantics, and large variations in brightness distribution. Existing methods, largely developed for natural scenes, often fail to jointly preserve semantic correspondence and fine-grained structural consistency when applied to aerial imagery, leading to semantic misalignment and local structural distortion despite visually plausible global appearance. To address these challenges, we propose SSDiffNet, a diffusion-based framework that explicitly enforces semantic alignment and structural consistency for aerial visible-to-infrared image translation. Instead of relying on global appearance matching, SSDiffNet guides the diffusion process with cross-modal semantic priors while constraining structural stability throughout generation. Specifically, an infrared image adaptive module (IIAM) adapts the diffusion decoder to infrared imaging characteristics under joint spatial- and frequency-domain constraints, enabling more faithful infrared-style reconstruction. A cross-modal semantic feature guidance module (CSFGM) injects learned visible-to-infrared semantic priors into multilevel denoising features, effectively enhancing regional semantic correspondence. In addition, a dual structural consistency constraint module (DSCCM) enforces invariant structural representations under equivalent transformations, substantially reducing structural deviation and local distortion during generation. Extensive experiments conducted on three public aerial visible-to-infrared datasets demonstrate that SSDiffNet consistently outperforms existing state-of-the-art methods in both visual quality and quantitative metrics.