Skip to content

Heterogeneous Multihead Decoding With Spatial Guidance for Remote Sensing Weakly Supervised Semantic Segmentation

2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 2506605-2506605 · 0 citations · 18 references

Abstract

Weakly supervised remote sensing semantic segmentation aims to achieve pixel-level prediction with limited annotation costs. Recently, CLIP-based methods have shown promising potential by leveraging vision-language alignment for semantic localization; however, they often suffer from ambiguous semantic representations and incomplete pseudolabels in complex remote sensing scenes. To address these issues, we propose a spatial cue-guided framework for weakly supervised remote sensing semantic segmentation. In particular, a confidence-aware prototype alignment module is designed to extract reliable semantic cues from high-confidence regions, enhance feature discrimination through contrastive learning, and improve object completeness by exploiting structural information. An adaptive pseudolabel completion (APLC) strategy is developed to progressively improve pseudolabel coverage while reducing noise propagation. In addition, a structure-aware heterogeneous multihead attention decoder is introduced to effectively fuse global semantic context, local spatial details, and cue information for refined prediction. Experimental results on two benchmark remote sensing datasets demonstrate the superiority of the proposed method.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.