Skip to content
Open access

Structure-Based Feature Representation for Robust Multi-Modal Image Matching

Aug 2026 · Remote Sensing · 0 citations · 32 references

TL;DR

This paper proposes a robust feature-based matching framework that reduces reliance on intensity information while enhancing structural representation and demonstrates that MIHOG can provide dense and reliable correspondences under complex cross-modal radiometric and geometric variations.

Abstract

Multi-modal image matching (MIM) remains a challenging problem due to nonlinear radiometric variations and geometric distortions across heterogeneous sensors. This paper proposes a robust feature-based matching framework that reduces reliance on intensity information while enhancing structural representation. The filter with local normalization is applied to transform the input images into a common intermediate domain. A block-based strategy is then employed to enforce a uniform spatial distribution of keypoints using the ORB (Oriented FAST and Rotated BRIEF) detector. To further suppress intensity variations and improve discriminability, a novel Max-Index-based HOG (MIHOG) is developed. This descriptor integrates multi-scale feature representations and encodes dominant structural information through discrete max-index mapping. Finally, correspondences are established using a brute-force matching strategy. Extensive experiments are conducted on two multi-modal datasets covering eight diverse scenarios. The proposed method achieves an average NCM of 224.52, RMSE of 3.4788, and SR of 92%. MIHOG obtains the highest NCM on 4/8 test scenarios and improves the average NCM by 18.3% compared with the second-best method. Meanwhile, it maintains competitive computational efficiency, with an average running time of 10.20s. These results demonstrate that MIHOG can provide dense and reliable correspondences under complex cross-modal radiometric and geometric variations.

Read PDF

Similar papers

Preprint Sep 2026

Radiation, Rotation and Scale Invariant Feature Descriptor for Multimodal Image Matching

A radiation, rotation, and scale invariant (RRSI) feature descriptor that enables feature encoding, interaction, and fusion across intra-modal, dual-head sampled, and inter-modal regions, and introduces a bidirectional cross-modal generative reconstruction constraint during training.

Yuan-Xin Ye, Teng-Feng Tang, Tao Peng et al. · 0 citations
Conference Open access 2026

Dense Image Matching Method Based on Transformer and Multi-Scale Feature Fusion

A dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet) is proposed, which outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness.

Shi-Xiong Liu · 0 citations
Open access 2026

Robust Optical-to-SAR Image Registration via Dense Tukey-Weighted Gradient Histogram and Structural Saliency Weight

Optical-to-SAR image registration is a fundamental prerequisite for multisource remote sensing applications, yet it remains challenging due to severe nonlinear radiometric differences, complex speckle noise, and structural blurring. Existing area-based matching methods often impose uniform spatial weighting. This unifo...

Wenhao Tong, An-Xi Yu, Huatao Yu et al. · 0 citations
Open access 2026

Fusing Joint Multiplexed Curvature Gradient and Color Features for Image Retrieval

Content-based image retrieval has become a popular research area, and image retrieval algorithms using the multi-channel mechanism or multi-feature extraction strategy usually can achieve competitive results. However, the robustness against rotation and scale variation is unsatisfactory in those algorithms. This paper...

Gang Zhang, Haoxiang Zhang · 0 citations
Open access 2026

Infrared and Visible Image Fusion Based on Gaussian Weighted Standard Deviation Filter

Infrared and visible image fusion (IVIF) aims to generate a comprehensive and informative fused image by combining complementary thermal radiation and texture details from dual-modal source images. However, current fusion techniques still suffer from inadequate preservation of thermal targets and fine-grained details,...

Lian Liu, Xiao-L. Cheng, Jin-Liang Huang et al. · 0 citations
Conference Aug 2026

Noise-Robust Multimodal Remote Sensing Image Matching Network with Edge-Aware Attention Fusion

Multimodal remote sensing image matching is challenged by nonlinear radiometric differences, geometric deformation, sensor noise, and weak cross-modal feature repeatability. We present an edge-aware coarse-to-fine Transformer network. A Feature Enhancement Transformer (FET) performs linear self- and cross-attention, wh...

Ping Liu, Bo Xu, Da-Peng Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.