Accurate agricultural parcel delineation from remote sensing imagery is essential for farmland registration, land-use monitoring, and field-level agricultural management. However, reliable parcel delineation remains challenging due to irregular field geometries, weak or ambiguous boundary cues, and substantial geographic heterogeneity across agricultural landscapes. To address these challenges, we propose SBCL-Net, a discrepancy-guided semantic–boundary interaction framework for agricultural parcel delineation. The framework consists of a semantic stream for region-level parcel representation and an independent boundary stream for structure-sensitive contour perception. To enhance semantic–boundary coupling, a Discrepancy-Guided Selective Modulation Module explicitly models shared and discrepant contexts between the two streams and selectively refines target features at discrepancy-sensitive locations. In addition, a consistency-constrained multi-task supervision strategy jointly optimizes mask prediction, boundary prediction, and their structural agreement. Experiments on FHAPD, FTW-France, and AI4Boundaries demonstrate strong and competitive performance across region-, boundary-, and object-level evaluations, with IoUs of 96.08% and 96.41% on FHAPD-JS and FHAPD-XJ, respectively, 71.49% on FTW-France, and 74.06% on AI4Boundaries. Cross-region evaluations further demonstrate the robustness of SBCL-Net under both within-country and cross-continental domain shifts. The code is available at https://github.com/Sabo-D/SBCL-Net
Yi-Fan Yu, Song Deng, Yang Yang et al.· IEEE Access· 0 citations
Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit cross-modal structural inconsistency, frequency–spatial response imbalance, and decoder-stage structural attenuation. To address these challenges, we propose GCF-Net, a stage-aligned optical–elevation fusion network that matches different cross-modal processing objectives to the evolving representation states of the encoder–decoder pipeline. A Structure-Guided Cross-Modal Correction Module first performs structure-conditioned correction of modality-specific features before fusion. A Frequency–Spatial Cross-Modal Fusion Module then constructs joint representations through bounded cross-modal magnitude conditioning, frequency-to-spatial reconstruction, and spatial recalibration. During decoding, a Geometry-Aware Cross-Scale Refinement Module reintroduces elevation-derived structural guidance into multiscale fused features. Experiments on ISPRS Vaihingen, ISPRS Potsdam, and MMHunan yield mIoU scores of 72.37%, 75.38%, and 52.23%, respectively, achieving the highest mIoU among the evaluated unimodal, multimodal, and SAM-based methods under the unified protocol. Ablation and replacement experiments verify the complementary roles of the three stage-specific components, while sensitivity, elevation perturbation, and complexity analyses indicate architectural flexibility, tolerance to moderate elevation degradation, and a balanced accuracy–efficiency trade-off.
Yi-Fan Yu, Song Deng, Yang Yang et al.· Remote Sensing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.