Skip to content

RealFusion: Rethinking Infrared and Visible Image Fusion from A Modality-Invariant Perspective.

Sep 2026 · IEEE Transactions on Image Processing · Vol PP · 0 citations
Medicine

Abstract

Existing fusion methods integrate complementary information from source images to generate a mixed-modality fused image, causing visually ambiguous representations that degrade downstream task performance (e.g., segmentation and detection). Therefore, we rethink infrared and visible image fusion from a modality-invariant perspective that reconstructs two real-modality fusion images (RealFusion), where the two challenges of cross-modality feature gap and modality inconsistency are addressed. Firstly, we introduce a scene prompt generation that leverages a semantic segmentation task to extract holistic scene information from source images, thereby bridging the gap between visible and infrared features. Secondly, we design a real-modality fusion image reconstruction that constructs the modality-guided vectors between the embedding space and each modality-specific space as the projection directions, thus projecting the fused features to their respective modality spaces. Thus, the fusion network is forced to learn a consistent representation from real-modality images, which enables the fusion results to keep the modality invariant. Additionally, we introduce a differential interaction fusion module that explicitly formulates the infrared and visible features as common and differential components, and then dynamically enhances the differential information that obtains a complementary fusion representation. Extensive experiments on three datasets demonstrate that RealFusion outperforms the state-of-the-art methods on both the fusion task and the downstream segmentation task. The code will be released at: https://github.com/HengshuaiCui/RealFusion.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.