Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network
Abstract
Hyperspectral–multispectral image fusion reconstructs high-spatial-resolution hyperspectral images (HR-HSIs) by combining low-resolution hyperspectral images (LR-HSIs) with high-resolution multispectral images (HR-MSIs). Many methods only support integer resolution ratios; when the LR-HSI/HR-MSI ratio is fractional, inputs are often resampled to a nearby integer ratio, altering observations and introducing interpolation error. We propose SCDF-Net, an unsupervised Scale-Conditioned Dual-domain Fusion Network that treats the spatial ratio as an explicit conditioning variable and fuses directly on native grids without HR-HSI labels. A degradation network first estimates the point spread function and spectral response function in a self-supervised manner; the learned operators are then frozen as physical priors. Conditioned on a continuous scale embedding, SCDF-Net integrates HSI spectral features and MSI spatial features through coupled frequency- and spatial-domain branches, trained with dual observation-domain consistency and spectral/frequency regularizations. On four benchmarks, baseline comparisons at scale factors ×1.5, ×2.0, ×2.4, and ×4.0 show that SCDF-Net obtains the lowest SAM on all datasets and improves PSNR/RMSE in most reported settings, with clear advantages under fractional scale factors. In addition, the multi-scale self-evaluation is extended to ×5.0 and ×6.0 to assess robustness under more severe spatial degradation.