Latent Diffusion Prior-Enhanced Frequency-Aware Deep Unfolding Network for Hyperspectral Image Denoising
Abstract
Deep unfolding networks (DUNs) offer an iterative paradigm that unrolls optimization procedures into a cascaded network structure for hyperspectral image (HSI) denoising. However, existing DUNs for HSI denoising suffer from two notable limitations: 1) the ill-posed inverse problem of handling severely degraded observations, calling for a more powerful degradation-free prior that encodes the spatial–spectral structure of HSIs and 2) existing discriminative denoisers are deterministic and tend to produce oversmoothed results. On the contrary, stochastic latent diffusion models (LDMs) have demonstrated the potential to synthesize high-frequency details with high perceptual quality for HSIs, but often at the cost of sacrificing data fidelity. To this end, we propose a latent diffusion prior-enhanced frequency-aware DUN (Diff-DUN) for HSI denoising. It builds on a half-quadratic splitting formulation that reconciles the observation consistency of deterministic DUNs with the high perceptual quality of a stochastic LDM. In particular, we employ a two-phase training strategy to first learn degradation-free priors from clean HSIs, followed by training the LDM to generate these priors conditioned on noisy inputs. Furthermore, we propose a frequency-aware Tetra Transformer (TT) block as the core component within the denoiser, introducing explicit frequency flow and cross-flow interactions to jointly model spatial, spectral, and frequency characteristics for more faithful recovery. Comprehensive evaluations across multiple datasets demonstrate that Diff-DUN achieves a favorable tradeoff between denoising accuracy and computational overhead, performing competitively with state-of-the-art (SOTA) methods.