Skip to content
Open access

Stage-Aware Robust Multimodal Prior Guidance for Diffusion-Based Image Super-Resolution

Jul 2026 · Electronics · 0 citations · 64 references

Abstract

Diffusion-based image super-resolution (SR) has recently achieved impressive perceptual quality by progressively generating plausible high-resolution details. However, its restoration performance still depends strongly on the reliability of the conditioning signal derived from degraded low-resolution inputs. Under severe or complex degradations, LR-derived conditions may become incomplete or ambiguous, leading the denoising trajectory toward visually plausible but input-inconsistent reconstructions. This work focuses on a central question: how to construct reliable multimodal prior guidance for a diffusion backbone that commonly adopts a hierarchical U-shaped architecture. To this end, we propose STMP-DiT, a stage-aware text-aligned multimodal prior-guided Diffusion Transformer for image super-resolution. From a multimodal data mining perspective, STMP-DiT aims to discover, align, and organize complementary semantic and structural priors from heterogeneous foundation-model representations. To improve the reliability of semantic guidance, STMP-DiT first aligns LLaVA-derived LR prompts with the frozen CLIP HR-image embedding space, producing visually grounded textual priors for restoration. These aligned textual priors are complemented by hierarchical DINO features, where deep features guide coarse semantic layout, intermediate features support structural recovery, and shallow features refine local edges and textures in the U-shaped DiT backbone. Rather than treating textual and visual priors as a single homogeneous condition, STMP-DiT assigns hierarchical DINO priors to different restoration stages according to their representational granularity. The fused condition is then injected through bounded feature modulation, enabling controlled stage-aware guidance while reducing redundant conditioning and improving the parameter efficiency of the conditioning modules. Experimental results on widely used SR benchmarks indicate promising improvements in perceptual quality and distributional realism with a compact trainable parameter scale, suggesting the value of reliable, stage-aware, and parameter-efficient multimodal prior integration for diffusion-based image super-resolution.

Read PDF