2026· Photogrammetric Engineering & Remote Sensing· 0 citations
TL;DR
The gated multi-scale interaction network (GMSINet), a U-shaped encoder–decoder framework with a hierarchical shifted-window self-attention Transformer backbone with a hybrid loss function is introduced to balance pixel-level supervision stability and region-level structural consistency, is proposed.
Abstract
Building extraction from very-high-resolution (VHR) remote sensing imagery remains challenging because buildings exhibit large variations in scale and are often embedded in complex textured backgrounds. These characteristics increase visual similarity and spatial confusion between buildings and their surrounding environments, making it difficult for building extraction models to achieve both multi-scale representation and detail preservation. To address these issues, we propose the gated multi-scale interaction network (GMSINet), a U-shaped encoder–decoder framework with a hierarchical shifted-window self-attention Transformer backbone. To reduce feature redundancy, background-noise propagation, and unstable multi-scale fusion in skip connections, a gated multi-scale interaction module is designed. This module adaptively filters enhanced information through SimAM-based recalibration, multi-receptive-field interaction, and channel-wise gating, thereby improving feature fusion quality in the decoding stage without changing feature resolution. In addition, a hybrid loss function is introduced to balance pixel-level supervision stability and region-level structural consistency. Experiments on the Massachusetts, INRIA, and WHU building data sets show that GMSINet achieves intersection over union scores of 74.50%, 77.35%, and 90.52%, respectively, outperforming the second-best models by 0.95%, 0.17%, and 0.25%, while maintaining stable overall accuracy and F1 score. These results demonstrate the effectiveness of GMSINet for building extraction from VHR remote sensing imagery.
As fundamental elements of urban landscapes, accurate building detection from high-resolution remote sensing imagery is an essential application of remote sensing information interpretation. However, this task remains challenging due to significant variations in building orientation, scale, and geometric shape. Existin...
Tong Gao, Yi-Peng Rao, Bin Bai et al.· IEEE Transactions on Geoscie...· 0 citations
Semantic segmentation of high-resolution remote sensing images faces three major challenges in frequency-spatial feature fusion: background clutter mixed into high-frequency components, semantic discontinuities within large homogeneous regions, and loss of fine rigid boundaries caused by convolutional downsampling. Tra...
Qi-Yuan Zhang, Jian-Shun Liu· Italian National Conference...· 0 citations
Road extraction from high-resolution remote sensing images is crucial for urban planning and geographic information systems (GIS). However, complex background interference, severe occlusions, and the inherent morphological complexity of roads often lead to discontinuities and insufficient accuracy in extraction results...
Jia-Jia Liu, Xuan Zhao, Wen-Xiang Dong et al.· Frontiers in Computing and I...· 0 citations
In remote-sensing scene classification (RSSC), persistent challenges such as high interclass similarity and substantial intraclass diversity remain key obstacles to accurate recognition. Although vision Transformer (ViT) has demonstrated outstanding performance, it tends to smooth out high-frequency discriminative deta...
Hui-Hui Dong, Tong Wang, Zong-Fang Ma et al.· IEEE Transactions on Geoscie...· 0 citations
Building extraction from optical remote sensing (RS) imagery is fundamental to urban mapping, yet existing methods are often dataset-specific and generalize poorly to unseen domains. Their practical use is also limited by insufficient detail recovery and weak geometric regularization, leading to blurred boundaries, irr...
Wei Huang, Chen-Ying Liu, Yi-Lei Shi et al.· 0 citations
This work presents a Dehazing Enhanced Multi-branch Attention Network (DEMANet) for effective remote sensing image dehazing that outperforms existing algorithms in haze removal, while simultaneously preserving intricate image details and color fidelity.
Pei-Xue Liu, Shu Liu, Peng-Fei He et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.