Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

SBCL-Net: Discrepancy-Guided Semantic–Boundary Interaction for Robust Agricultural Parcel Delineation From Remote Sensing Imagery

Accurate agricultural parcel delineation from remote sensing imagery is essential for farmland registration, land-use monitoring, and field-level agricultural management. However, reliable parcel delineation remains challenging due to irregular field geometries, weak or ambiguous boundary cues, and substantial geographic heterogeneity across agricultural landscapes. To address these challenges, we propose SBCL-Net, a discrepancy-guided semantic–boundary interaction framework for agricultural parcel delineation. The framework consists of a semantic stream for region-level parcel representation and an independent boundary stream for structure-sensitive contour perception. To enhance semantic–boundary coupling, a Discrepancy-Guided Selective Modulation Module explicitly models shared and discrepant contexts between the two streams and selectively refines target features at discrepancy-sensitive locations. In addition, a consistency-constrained multi-task supervision strategy jointly optimizes mask prediction, boundary prediction, and their structural agreement. Experiments on FHAPD, FTW-France, and AI4Boundaries demonstrate strong and competitive performance across region-, boundary-, and object-level evaluations, with IoUs of 96.08% and 96.41% on FHAPD-JS and FHAPD-XJ, respectively, 71.49% on FTW-France, and 74.06% on AI4Boundaries. Cross-region evaluations further demonstrate the robustness of SBCL-Net under both within-country and cross-continental domain shifts. The code is available at https://github.com/Sabo-D/SBCL-Net

Yi-Fan Yu, Song Deng, Yang Yang et al. · 0 citations
Open access Aug 2026

Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation

Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage framework that follows a correction-before-fusion strategy. At each feature scale, SPGNet estimates a learned structure prior from cross-modal agreement and discrepancy. The Cross-Modal Correction Module (CCM) uses this prior to regulate bidirectional information transfer, suppressing unreliable responses while retaining complementary cues. The Dual-branch Enhancement Fusion Module (DEF) then enhances the corrected RGB and depth features and integrates them through shared-representation-guided interaction, after which a lightweight multi-scale decoder produces the segmentation output. Under a unified training and evaluation protocol, SPGNet achieved three-run mean Intersection over Union (mIoU) scores of 50.845% on NYU Depth V2 and 48.457% on SUN RGB-D. Compared with the best reproduced baseline on each dataset, SPGNet improved mean mIoU by 2.111 and 0.899 percentage points, respectively. These results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance.

Yi-Fan Yu, Zhiwei Zhong, Fan Min et al. · 0 citations
Open access Aug 2026

GCF-Net: Stage-Aligned Optical–Elevation Fusion for Aerial Remote Sensing Semantic Segmentation

Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit cross-modal structural inconsistency, frequency–spatial response imbalance, and decoder-stage structural attenuation. To address these challenges, we propose GCF-Net, a stage-aligned optical–elevation fusion network that matches different cross-modal processing objectives to the evolving representation states of the encoder–decoder pipeline. A Structure-Guided Cross-Modal Correction Module first performs structure-conditioned correction of modality-specific features before fusion. A Frequency–Spatial Cross-Modal Fusion Module then constructs joint representations through bounded cross-modal magnitude conditioning, frequency-to-spatial reconstruction, and spatial recalibration. During decoding, a Geometry-Aware Cross-Scale Refinement Module reintroduces elevation-derived structural guidance into multiscale fused features. Experiments on ISPRS Vaihingen, ISPRS Potsdam, and MMHunan yield mIoU scores of 72.37%, 75.38%, and 52.23%, respectively, achieving the highest mIoU among the evaluated unimodal, multimodal, and SAM-based methods under the unified protocol. Ablation and replacement experiments verify the complementary roles of the three stage-specific components, while sensitivity, elevation perturbation, and complexity analyses indicate architectural flexibility, tolerance to moderate elevation degradation, and a balanced accuracy–efficiency trade-off.

Yi-Fan Yu, Song Deng, Yang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.