Skip to content
Open access

A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison

Aug 2026 · Applied Sciences · 0 citations · 12 references

Abstract

As a key pixel-level analysis technology, semantic segmentation is widely deployed in autonomous driving and medical imaging. Fully supervised segmentation relies on labour-intensive pixel-wise annotations, so weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted wide attention. Existing Vision Transformer (ViT)-based WSSS methods suffer from two critical limitations: ViT’s global self-attention mechanism leads to insensitivity to local small target features and incomplete foreground activation; its class-agnostic attention maps frequently misactivate background regions as foreground objects, introducing heavy noise. To tackle these two issues, this paper proposes a single-stage weakly supervised segmentation algorithm based on local–global class labelling comparison. First, we design a local–global class labelling comparison (LTG) module. By feeding both original images and randomly cropped local patches into ViT, we adopt InfoNCE contrastive loss to align local class tokens with global class tokens, enhancing the feature integrity of local target regions and suppressing background false activation. Second, a class-aware stimulus module (CSM) is embedded into ViT’s multi-head attention branch. It injects category semantic constraints into self-attention to generate class-aware attention maps, guiding the model to focus on real foreground targets and reduce background interference. Finally, we construct a feature fusion class-aware activation map (FFCAM) by fusing ViT global output features and CSM class-aware attention features to generate high-quality pseudo-labels for segmentation training. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 datasets. Our method achieves 78.5% mIoU on the validation set of PASCAL VOC 2012 and 50.9% mIoU on MS COCO 2014, showing competitive numerical performance among the compared single-stage ViT-based WSSS approaches. Ablation experiments verify the independent and joint effectiveness of LTG, CSM and FFCAM. The proposed method effectively improves the completeness of target activation regions and suppresses background noise, and it maintains strong generalization for slender, small and texture-sparse objects. In future work, we will further lightweight the ViT backbone to reduce computational overhead for embedded deployment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.