Skip to content
Preprint

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

Aug 2026 · 0 citations · 62 references
Computer Science

TL;DR

HAFR-Net is presented, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation and shows consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.

Abstract

Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.

View source

Similar papers

Open access Sep 2026

HDSMNet: Height-Guided Sparse Cross-Modal Fusion for High-Resolution Remote Sensing Semantic Segmentation

HDSMNet is proposed, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data that enhances discriminative dense feature representations in high-resolution remote sensing images through interaction with a compact set of geometry-guided anchors.

Han-Xun Gu, Jiang-Jie Hu, Li Wang et al. · 0 citations
2026

A Coarse-to-Fine Progressive Fusion Network for Multimodal Semantic Segmentation of Remote Sensing Images

Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and in...

Di Zhang, Yuhang Yan, Q. Niu et al. · 3 citations
2026

GMSINet for Building Extraction from Very-High-Resolution Remote Sensing Imagery

The gated multi-scale interaction network (GMSINet), a U-shaped encoder–decoder framework with a hierarchical shifted-window self-attention Transformer backbone with a hybrid loss function is introduced to balance pixel-level supervision stability and region-level structural consistency, is proposed.

Guobiao Yao, Ze-Yu Zhang, Qing-Dong Wang et al. · 0 citations
Open access Aug 2026

Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement

A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing s...

Jing Li, Yu-Lin Cao, Xian-Tao Jiang et al. · 0 citations
2026

Collaborative Context-Affine Perception Network for Remote Sensing Small Object Detection

Small object detection in remote sensing images (RSIs) is challenging because imaging degradation weakens object textures, reduces contrast, and blurs boundaries. These effects are further aggravated by hierarchical feature extraction, where repeated downsampling weakens shallow spatial cues before they reach deeper se...

Wei He, Yun-Tao Xu, Qi Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.