Skip to content

Global and Neighbor-Aware Token Learning for Weakly Supervised Remote Sensing Image Semantic Segmentation

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 4706315-4706315 · 0 citations · 52 references

Abstract

Image-level weakly supervised remote sensing semantic segmentation aims to learn pixel-level land-cover prediction using only image-level labels, greatly reducing the annotation cost of fully supervised methods. Class activation map (CAM)-based methods are widely used for this task, but they usually focus on the most discriminative regions, leading to incomplete activation and inaccurate boundaries. Recently, vision Transformer (ViT)-based methods have been introduced to alleviate the limitation of CAMs by exploiting token relations and attention mechanisms. However, remote sensing images often contain dense land-cover regions with subtle interclass differences, and patch tokens in deep ViT layers may become oversmoothed without explicit patch-level supervision, weakening local semantic discrimination. Moreover, large intraclass variations and frequent category co-occurrence make image-specific class tokens prone to semantic drift across different remote sensing images. To address these problems, we propose a global and neighbor-aware token learning (GNATL) framework. GNATL contains two complementary modules: neighbor-aware patch token learning (NPTL) and global class token learning (GCTL). NPTL exploits overlapping regions between neighboring crops to construct implicit patch-level constraints, thereby alleviating patch token oversmoothing. Global class token learning (GCTL) dynamically maintains global class tokens as category-level prototypes to guide image-specific class tokens toward stable category semantics. Experiments on the International Society for Photogrammetry and Remote Sensing (ISPRS) Potsdam, ISPRS Vaihingen, and DeepGlobe Land Cover datasets show that GNATL achieves mean Intersection over Union (mIoU) scores of 56.38%, 47.74%, and 62.75%, outperforming the best compared methods by 2.63%, 4.57%, and 2.36%, respectively.

View source

Similar papers

Open access Aug 2026

Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement

A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing s...

Jing Li, Yu-Lin Cao, Xian-Tao Jiang et al. · 0 citations
Open access Sep 2026

Reliability-Aware Dual-Stream Self-Training for Semi-Supervised Semantic Segmentation of High-Resolution Remote Sensing Imagery

RAPST is proposed, a reliability-aware dual-stream self-training framework that combines reliable offline supervision with full-data online learning and achieves the best performance in most dataset–annotation settings.

Wan-Zhaohe-Wei-Shunwu-Shi-Fang-Wang-Rui-Fang Liu, Yan Lin, Wei Liu et al. · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations
Aug 2026

MSPD-net: structural–appearance prototype decoupling for weakly supervised semantic segmentation

A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.

Wei Cao, Yong Jiang, Ruiying Wang · 0 citations
2026

Heterogeneous Multihead Decoding With Spatial Guidance for Remote Sensing Weakly Supervised Semantic Segmentation

Weakly supervised remote sensing semantic segmentation aims to achieve pixel-level prediction with limited annotation costs. Recently, CLIP-based methods have shown promising potential by leveraging vision-language alignment for semantic localization; however, they often suffer from ambiguous semantic representations a...

Bei Cheng, Quan-Li Deng, Tao Shen et al. · 0 citations
Open access Aug 2026

Implementation of Partial Cross Entropy Loss for Point-supervised Remote Sensing Image Segmentation

This study presents an implementation framework for point-supervised remote sensing image segmentation using Partial Cross Entropy Loss and provides a simple, modular, and reproducible implementation and establishes a foundation for future investigations involving weakly supervised and semi-supervised remote sensing im...

Loraine Mutune, Tecla Mutave Kyalo, J. Mutinda et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.