Skip to content
Preprint

BASeg: Boundary-Aware Remote Sensing Segmentation with Structural Penalties

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.

Abstract

Semantic segmentation is a core computer vision task in the remote sensing field, accelerating advancements in ur- ban development, agriculture, ecology, water resources, and environmental monitoring. However, recent methods usually struggle to capture fine-grained object features and bound- ary details. Besides, current widely used datasets often lack city morphology diversity and segmentation on generative im- ages remains largely unexplored. To address these issues, we propose a Mahalanobis-Angle Boundary Loss (MABL) that explicitly enhances boundary and shape consistency. MABL jointly models structural importance and boundary orientation through Mahalanobis distance-based weighting and angle- aware penalty. It can be readily integrated into diverse seg- mentation architectures and consistently improves their accu- racy. Built upon MABL, we introduce BASeg, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties. BASeg integrates a Global Visual State Space module (GSM) with a Cross-Feature Fusion module (CFM) to capture both long-range contextual dependencies and fine- grained local details. Additionally, we establish a global 10- city benchmark dataset (GCD-25k) to facilitate accurate build- ing and road segmentation. Extensive experiments on four remote-sensing benchmarks demonstrate that BASeg consis- tently outperforms existing methods, achieving up to a 2.8% improvement in mIoU while producing more accurate object boundary segmentation across diverse scenes. Moreover, integrating MABL into multiple existing segmentation archi- tectures consistently improves performance across datasets, demonstrating its robustness and broad applicability.

View source

Similar papers

Open access Aug 2026

Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation

Remote sensing image semantic segmentation (RSISS) has attracted significant attention due to the growing demand for fine-grained land cover information. The Segment Anything Model (SAM), proposed as a foundation vision model, offers strong segmentation performance and generalization capabilities for RSISS tasks. However, existing SAM-based approaches face two limitations: (1) Insufficient adaptation of SAM's features to the diverse characteristics of land cover types. (2) Semantic ambiguity at object boundaries, which hinders accurate delineation. To address these limitations, we propose Frequency and Edge-guided SAM (FE-SAM), a scalable and efficient framework for RSISS. Specifically, we introduce a Frequency-Modulated Adapter (FMA) that adaptively decomposes and modulates frequency-domain features based on the input data. It selectively enhances informative high- and low-frequency components corresponding to different land cover types. Furthermore, to improve SAM's ability to capture fine-grained details, we design EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image. Extensive experiments on three benchmark datasets demonstrate that FE-SAM outperforms state-of-the-art methods. The source codes are available at: https://github.com/oucailab/FE-SAM.

Feng Gao, Zizhe Pan, Haoting Wang et al. · 0 citations
2026

MBANet: Multiscale Boundary Aware Network for Landslide Identification on Remote Sensing Imagery

Semantic segmentation of remote sensing imagery has been widely applied in landslide identification, effectively addressing the time-consuming and labor-intensive nature of manual visual interpretation. However, existing models still face challenges in extracting multiscale features and accurately delineating boundaries under complex background conditions. To overcome these limitations, this study proposes a multiscale boundary aware network (MBANet) for landslide identification in remote sensing imagery. Specifically, we design a multiscale cross-interaction convolution (MCC) module that captures local details and broader contextual cues through heterogeneous receptive-field branches, and recalibrates the concatenated multiscale features via an adaptive cross-branch interaction strategy. In addition, a boundary sensitive refinement attention (BSRA) module is introduced to enhance boundary localization by combining a Sobel-based gradient prior, learnable boundary estimation, and region-context enhancement for fine-grained boundary refinement. These modules are integrated into an encoder–decoder architecture to jointly achieve semantic consistency and boundary precision. Experimental results on two public datasets show that MBANet outperforms other comparison models in overall segmentation performance and maintains competitive performance in boundary delineation. On the Bijie dataset, it achieves a recall of 83.84% and an $F1$ -score of 85.39%; on the Palu dataset, it reaches a recall of 75.83% and an $F1$ -score of 78.17%, highlighting its superior performance.

Zixun Xie, Chuang Song, Xingmin Cai et al. · 0 citations
Open access Aug 2026

AB-SAM: A SAM-Based Asymmetric Boundary-Aware Model for the Semantic Segmentation of Small and Medium-Sized Landslides

Small- and medium-sized landslides frequently occur in clusters and exhibit fragmented morphologies, irregular boundaries, and spectral characteristics similar to surrounding roads, bare soil, and sparsely vegetated surfaces, making their automated extraction from remote sensing imagery challenging. Although the Segment Anything Model (SAM) provides strong general-purpose segmentation capabilities, its direct application to landslide mapping is limited by the geoscience domain gap and its dependence on external prompts. This study proposes the Asymmetric Boundary-aware Segment Anything Model (AB-SAM), a parameter-efficient adaptation of SAM for automated landslide semantic segmentation. AB-SAM integrates three task-specific components. First, the offline Multi-Feature Variation-Guided Prompting (MF-VGP) module generates cached auxiliary bounding boxes from registered pre- and post-event images without accessing ground-truth masks. Second, the Asymmetric Feature Augmentation (AFA) strategy combines geometric perturbation, CutMix, and asymmetric dual-branch supervision, in which a Hint-free branch serves as the primary optimization pathway and a lower-weight box-guided branch provides auxiliary spatial supervision. Third, the Boundary-Aware Morphological Prompting (BAMP) module injects trainable boundary-aware morphological information into the largely frozen SAM image encoder. During validation, testing, and application, only the Hint-free branch is retained, enabling inference using post-event imagery without external point, box, or mask prompts. On the fixed, spatially disjoint Zixing test set, AB-SAM achieved an overall accuracy of 96.171%, a precision of 68.149%, a recall of 60.011%, an F1-score of 63.822%, a landslide-class Intersection over Union of 46.867%, and a mean Intersection over Union of 71.452%. Repeated experiments with three random seeds showed low run-to-run variation. Direct evaluation without retraining on the Hokkaido Iburi-Tobu dataset yielded a mean Intersection over Union of 66.136%, providing evidence of cross-region and cross-event transferability. These results demonstrate that AB-SAM provides a practical parameter-efficient framework for automated, hint-free landslide segmentation, although further evaluation across additional regions, sensors, and landslide-size distributions remains necessary.

Jiting Tang, Zhiwei Liang, SuLi Guo et al. · 0 citations
Open access Aug 2026

BCNet: Boundary-Constrained Remote Sensing Change Detection Network Based on Vision Foundation Models

Limited by the diversity and complexity of real-world scenes, existing remote sensing change detection methods often suffer from insufficient fine-grained semantic understanding and blurred boundaries of change targets. To address these issues, this paper proposes a boundary-constrained remote sensing change detection network based on vision foundation models (BCNet). BCNet employs a differential modeling approach and multi-branch guidance mechanism to design a differential detail enhancement module, amplifying fine-grained semantic information. Through cross-layer feature alignment, stepwise fusion, and edge-sensitive modeling, it constructs a multi-scale edge enhancement module that enhances perception of minute variations and edge details, fully leveraging the universal semantic representation capabilities of the vision foundation model. In addition, an edge feature constraint mechanism is introduced that applies dual guidance and supervision during the feature fusion and output stages. This mechanism achieves refined delineation of change region boundaries and significantly mitigates the issue of boundary blurring. Experimental results on four mainstream datasets, namely LEVIR-CD, WHU-CD, NJDS and MSRS-CD, demonstrate that BCNet outperforms 13 state-of-the-art methods in terms of key metrics including F1 and IoU. Against the best VFM-based baseline, BCNet obtains F1 score gains of 0.21%, 0.71%, 6.33% and 0.63% on the above four datasets. Specifically, the proposed method exhibits superior detection accuracy and edge detail preservation capabilities in complex regions.

Shenbo Liu, Dongxue Zhao, Huang He et al. · 0 citations
Open access 2026

SAPLNet: State-Aware Prototype Learning for Remote Sensing Segmentation

Semantic segmentation of high-resolution remote sensing images remains challenging due to complex spatial structures, multiscale object variations, fine-grained category differences, and high interclass similarities. Conventional segmentation methods usually rely on fixed convolutional heads or single feature representations, which makes it difficult to effectively model both intraclass appearance variations and interclass texture similarities, often leading to category confusion, missed objects, and incomplete segmentation in complex scenes. To address these challenges, we propose a state-aware prototype learning network, termed SAPLNet. Specifically, a cross-stage state refiner is introduced to progressively refine multilevel features by integrating the input features with the outputs of different stages through state-aware gated normalization. Then, a weighted feature pyramid decoder performs top-down fusion of the refined hierarchical features, combining high-level semantic information with low-level spatial details. Furthermore, a state-aware multiprototype classifier is designed to construct multiple semantic prototypes for each class via ground-truth-guided local class-center extraction and momentum-based prototype memory updating. A global state vector derived from the refined cross-stage features is used to adaptively modulate decoder features, improving the matching reliability between pixel features and class prototypes. In addition, prototype compactness loss, prototype diversity loss, and lightweight boundary loss are employed to enhance intraclass consistency, prototype discriminability, and boundary awareness. Experimental results demonstrate the effectiveness and superiority of SAPLNet.

Zeyu Zhao, Zhaolong Gao, Jun Feng · 0 citations
2026

RSUS: A Novel Upsampling Layer for Semantic Segmentation Network of Remote Sensing Images

Semantic segmentation of remote sensing images (RSIs) often struggles with boundary blurring, structural discontinuity, and category confusion due to limitations in conventional interpolation and dynamic upsampling methods. This article proposes RSUS, a new upsampling layer designed to preserve semantic consistency and spatial structure. RSUS consists of three components: 1) global context vector aggregation (GCVA) for content-aware prediction kernels that introduce global priors into local feature reconstruction; 2) cross-scale anisotropic implicit positional encoding (CAIPE) for direction-sensitive spatial deformation and structural alignment; and 3) adaptive high-frequency structure gating (AHSG) to enhance boundary-related frequency responses. In addition, persistent homology (PH) is used as a topological analysis tool to assess connectivity and structure preservation beyond pixel-level metrics. Extensive experiments on five datasets (International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen, ISPRS Potsdam, LoveDA, UAVid, and Xining) demonstrate that RSUS improves performance across mainstream segmentation networks, outperforming advanced upsampling methods. Analyses show RSUS alleviates structural misalignment, improves category consistency, and recovers fine-grained boundaries in remote sensing segmentation. The source code is available at: https://github.com/Ronin-711/RSUS

Yaning Liu, Ronghao Yang, Shaoda Li et al. · 0 citations