Aug 2026· Asian Journal of Research in Computer Science· 0 citations
TL;DR
This study presents an implementation framework for point-supervised remote sensing image segmentation using Partial Cross Entropy Loss and provides a simple, modular, and reproducible implementation and establishes a foundation for future investigations involving weakly supervised and semi-supervised remote sensing image analysis.
Abstract
Weakly supervised semantic segmentation has emerged as a promising approach for reducing the annotation burden associated with dense pixel-level labelling. Among weak supervision strategies, point supervision offers an attractive compromise by requiring only a small subset of labelled pixels while preserving meaningful semantic information. However, effective optimisation under sparse supervision requires loss functions capable of excluding unlabelled regions from the training process. This study presents an implementation framework for point-supervised remote sensing image segmentation using Partial Cross Entropy Loss. Dense segmentation masks obtained from the LoveDA dataset were converted into sparse point annotations through random pixel sampling with a point ratio of 1%. A subset containing 200 image-mask pairs was constructed and partitioned into training and validation sets. A U-Net architecture with a ResNet-34 encoder was employed as the segmentation backbone, while Partial Cross Entropy Loss was implemented using the ignore-index mechanism available in PyTorch to restrict optimization to labelled pixels only. Functional validation was performed through forward propagation, loss computation, and gradient backpropagation. Successful parameter updates and finite loss values confirmed the correct integration of sparse point supervision with the encoder-decoder segmentation network. The proposed framework provides a simple, modular, and reproducible implementation for point-supervised semantic segmentation and establishes a foundation for future investigations involving weakly supervised and semi-supervised remote sensing image analysis.
Image-level weakly supervised remote sensing semantic segmentation aims to learn pixel-level land-cover prediction using only image-level labels, greatly reducing the annotation cost of fully supervised methods. Class activation map (CAM)-based methods are widely used for this task, but they usually focus on the most d...
Mansu Gu, Jing Bai, Rui-Zhe Guan et al.· IEEE Transactions on Geoscie...· 0 citations
RAPST is proposed, a reliability-aware dual-stream self-training framework that combines reliable offline supervision with full-data online learning and achieves the best performance in most dataset–annotation settings.
Wan-Zhaohe-Wei-Shunwu-Shi-Fang-Wang-Rui-Fang Liu, Yan Lin, Wei Liu et al.· Remote Sensing· 0 citations
Multimodal fusion methods have shown great potential in remote sensing image analysis, but existing approaches rely heavily on massive amounts of annotated data. This is not only costly and time-consuming but also prone to subjective bias. To address this issue, we propose a category-prior-based self-supervised framewo...
Jia-Hang Liu, Jian Cui, Mao-yin Guo et al.· IEEE Transactions on Geoscie...· 0 citations
A three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training, Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM) is proposed, demonstrating the effectiveness and generalizability of the proposed framework across diverse remote sensing s...
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
Accurate building segmentation from high-resolution unmanned aerial vehicle (UAV) imagery is essential for urban mapping, three-dimensional modeling, and related remote sensing applications. However, constructing precise pixel-level annotations for such imagery is labor-intensive, whereas unlabeled imagery from new sur...