Skip to content

A Coarse-to-Fine Progressive Fusion Network for Multimodal Semantic Segmentation of Remote Sensing Images

2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 7506705-7506705 · 3 citations · 14 references

Abstract

Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and inconsistent spatial responses. To address the aforementioned challenges, this letter proposes a coarse-to-fine progressive fusion network (CFPFNet). Specifically, a visual state space model extracts a Mamba-derived global structural prior to guide the coarse-grained context enhancement (CGCE) module for preliminary cross-modal alignment. Then, the fine-grained adaptive frequency-spatial fusion (FGAF) module performs amplitude-phase collaboration and adaptive spatial cross-gating for multiscale semantic refinement. Experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness of CFPFNet. On Vaihingen, CFPFNet improves mIoU and mF1 by 2.22% and 1.39% over the simple dual-stream baseline, respectively.

View source

Similar papers

Open access Sep 2026

HDSMNet: Height-Guided Sparse Cross-Modal Fusion for High-Resolution Remote Sensing Semantic Segmentation

HDSMNet is proposed, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data that enhances discriminative dense feature representations in high-resolution remote sensing images through interaction with a compact set of geometry-guided anchors.

Han-Xun Gu, Jiang-Jie Hu, Li Wang et al. · 0 citations
Open access 2026

Scale-Aware Fusion and Spatial-Frequency Collaborative Network for Remote Sensing Imagery Semantic Segmentation

Semantic segmentation of remote sensing images is crucial for serial earth observation tasks. However, significant scale variations in remote sensing scenes and insufficient exploitation of frequency-domain information often cause small-scale objects to be overwhelmed by large backgrounds under imbalanced multiscale fe...

Lu Wang, Chenxuan Lou, Jing Wang et al. · 0 citations
Open access 2026

Bridging Global Context and Local Detail in State Space Models for Fine-Grained Landslide Segmentation in Remote Sensing Imagery

This work proposes landslide state-space Mamba, a landslide segmentation network that augments a VSSD backbone with two lightweight local enhancement modules: a structural feature calibration (SFC) module that adaptively calibrates spatial structural discrepancies after global interaction, and a local detail enhancemen...

Kang-Ning Wang, Hao-Peng Zhang, Zhi-Guo Jiang · 0 citations
2026

LiEAF-Net: A Lightweight Multiscale Elevation-Aware Fusion Network for Multimodal Semantic Segmentation of High-Resolution Remote Sensing Imagery

Semantic segmentation of high-resolution remote sensing imagery is computationally demanding, particularly for multimodal fusion of RGB and normalized digital surface model (nDSM) data. Existing multimodal networks improve segmentation accuracy but often introduce substantial computational overhead. This letter present...

Furkat Sultonov, Mu-Gyeong Gong, Sang-Jae Park et al. · 0 citations
Open access Aug 2026

GCF-Net: Stage-Aligned Optical–Elevation Fusion for Aerial Remote Sensing Semantic Segmentation

Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation repres...

Yi-Fan Yu, Song Deng, Yang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.