Skip to content

Visual-Prompt Bidirectional Interactive Learning for Remote Sensing Infrared and Visible Image Fusion

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5639416-5639416 · 0 citations · 59 references

Abstract

Remote sensing infrared and visible image fusion (IVIF) aims to leverage the modality advantages of infrared and visible images to enhance the comprehensiveness and precision of visual representations. However, existing text-guided methods are only constrained by the unidirectional text-to-vision semantic guidance paradigm, where textual description biases directly propagate to fusion results. This may lead to deviations in the fusion outputs, thereby suffering from the loss of original visual information. Therefore, this article proposes a visual-prompt bidirectional interactive learning (VP-BIL) method for remote sensing IVIF. The method constructs a bidirectional semantic alignment mechanism based on visual features and prompt representations generated by the text encoder, enabling efficient cross-modal fusion without requiring accurate text annotations. In particular, a text-visual semantic correlation is first established that ensures both semantic alignment accuracy and the preservation of original visual information. Subsequently, a visual-driven text correction module is designed to dynamically calibrate textual description biases, enhancing the cross-modal semantic alignment between the text and vision. Extensive experiments demonstrate that the proposed method effectively overcomes the unidirectional dependence limitation of the text-guided paradigm, outperforming the state-of-the-art (SOTA) methods.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.