Skip to content

The PID-Aware Geometric Semantic Representation Learning for Infrared-Visible Image Fusion

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 4414714-4414714 · 0 citations · 51 references

Abstract

To relieve the problems of the lack of semantic alignment in the geometric representation space and insufficient modality-specific instruction calibration in infrared-visible image fusion (IVIF), we first propose a dynamic instruction-aware geometric representation (DIGR) for IVIF, which achieves semantic alignment between fusion features and semantic features within the geometric space and progressively injects textual instructions into the deep-layer features. Second, to alleviate the semantic gap between fusion features and semantic features, we propose a cross-graph neighborhood interaction (CGNI) module. In this module, each node in one graph is updated by aggregating information from the neighbor nodes of its position-aligned counterpart in the other graph, thereby enabling cross-level semantic alignment. Third, to alleviate the absence of modality-related text correction in the fusion network, we introduce an instruction-guided proportional–integral–derivative (IPID) calibration module, which uses modality-aware texts as external priors. By modeling the discrepancy between instructions and deep-level modality features using the proportional–integral–derivative (PID) control principle, IPID achieves progressive and controllable instruction calibration. Extensive experiments on three public datasets, namely, MFNet, Potsdam, and WHU, demonstrate that the proposed DIGR achieves superior fusion results and segmentation performance. The code of DIGR is publicly available at https://github.com/YongzheWang666/DIGR

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.