Skip to content

Robust webly supervised fine-grained recognition via decoupled global-local fusion and geometric-semantic consensus.

Aug 2026 · Neural Networks · Vol 205 Pt C, pp. 109571 · 0 citations · 45 references
Medicine

TL;DR

A novel framework named Decoupled Global-local Consensus Learning (DGCL) is proposed that identifies and leverages noisy-labeled samples through multi-perspective analysis, enhancing both robustness and representation power.

Abstract

Webly Supervised Fine-Grained Recognition demands robust learning mechanisms to mitigate the impact of noisy labels and misleading annotations. To address this, current methods typically rely on prediction probabilities to filter noisy labels. However, such paradigms often struggle with mislabeled yet high-confidence samples, leading to confirmation bias. In this paper, we propose a novel framework named Decoupled Global-local Consensus Learning (DGCL) that identifies and leverages noisy-labeled samples through multi-perspective analysis, enhancing both robustness and representation power. Motivated by the high-quality dense representations of foundation models, we first introduce the Decoupled Global-Local Fusion (DGLF) framework, which leverages LoRA to calibrate the encoder's attention, enabling precise identification of discriminative regions. Subsequently, a dual-stream bilinear mechanism is designed to facilitate global-local interactions for enhanced detail capture. To mitigate confirmation bias, we introduce the Geometric-Semantic Consensus (GSC) strategy, which integrates geometric consistency and semantic confidence to partition samples into three mutually exclusive subsets. Finally, a Noise-Aware Supervised Contrastive (NASC) loss is introduced to convert filtered noise into repulsion signals, effectively compressing intra-class variance. Extensive experiments demonstrate that DGCL achieves a new state-of-the-art average accuracy of 91.73% on three web-supervised benchmark datasets, surpassing the previous best method by 1.63 percentage points. Our source code will be made publicly available at: https://github.com/YT3DVision/DGLF.

View source

Similar papers

Open access 2026

HCCL: Hierarchical Consistency-Driven Contrastive Learning for Semi-Supervised Learning

This work introduces an optimal transport-based online clustering mechanism to automatically discover latent coarse-grained and fine-grained semantic priors without external supervision and introduces a cross-granularity alignment module that enforces consistency between these discovered latent structures and the targe...

Tian-Zhu Cong, Wei Liu, He-Feng Yin · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations
2026

Reliability-Calibrated Posterior Supervision for Weakly Supervised Semantic Segmentation

Weakly supervised semantic segmentation (WSSS) with image-level labels is largely limited by the reliability of dense seed supervision. Existing CAM- and CLIP-based methods provide complementary localization cues, but their predictions are biased in different ways: classification-oriented cues are usually precise but i...

Xiao-Ya Sun, Xin Xu · 0 citations
Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al. · 0 citations
Open access 2026

Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval

Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...

Zi-Xuan Zhang, Lu-Ming Xiao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.