Skip to content
Conference Open access

Dense Image Matching Method Based on Transformer and Multi-Scale Feature Fusion

2026 · ITM Web of Conferences · 0 citations · 1 references

TL;DR

A dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet) is proposed, which outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness.

Abstract

Dense image matching is crucial in applications such as 3D reconstruction, autonomous driving, and remote sensing mapping; however, weak textures, occlusions, and large-disparity scenes remain challenging. To address these issues, this paper proposes a dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet). First, Swin Transformer is used to model global context in feature maps, enhancing the feature discriminability in weak texture regions. Then, a multi-scale cost volume is constructed, and adaptive fusion is achieved through deformable convolution to accommodate disparity variations of different ranges. Finally, an attention- guided iterative optimization module is introduced to improve the matching accuracy in occluded regions. Experimental results on the Scene Flow, KITTI-2015, and Middlebury datasets show that TMSMNet outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness. Ablation studies also confirm the effectiveness of each module. In summary, the method in this paper provides a feasible approach for dense matching. Future work will explore model lightweighting to support real-time applications and attempt to combine generative models to handle completely textureless regions, further enhancing its performance in complex scenes.

Read PDF

Similar papers

2026

A Subpixel Accurate Image-Matching Method Integrating SuperGlue and Local Optimization

Subpixel accurate image matching is critical for high-precision applications across diverse scenarios, including multi-temporal satellite image registration, underwater inspection, and rigorous photogrammetric processing. Deep learning–based matching methods demonstrate strong robustness. However, applying them to out–...

Zhenling Ma, Zheng-Jie Wang, Xu Zhong et al. · 0 citations
Open access Aug 2026

Structure-Based Feature Representation for Robust Multi-Modal Image Matching

This paper proposes a robust feature-based matching framework that reduces reliance on intensity information while enhancing structural representation and demonstrates that MIHOG can provide dense and reliable correspondences under complex cross-modal radiometric and geometric variations.

Yameng Hong, Chengcai Leng, Zhao Pei · 0 citations
Open access Sep 2026

A Lightweight Deep Learning Framework for Parallax-Tolerant Image Stitching

A transformer-based channel attention block improves the discriminative capability of fused features in low-texture regions and enhances global consistency in a lightweight deep stitching framework that integrates multi-scale feature fusion with attention-enhanced matching.

Yi-Liang Wu, Hua-Wang Huang, Zong-Kai Huang et al. · 0 citations
Preprint Aug 2026

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging...

Zhi-Hua Xu, Run-Yu Zhu, Rong Qin · 0 citations
Conference Aug 2026

A multimodal BEV 3D object detection method with depth uncertainty and geometric saliency

High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this...

Jie Hu, Xinghao Cheng, Shuaidi He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.