Skip to content

CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification

Jul 2026 · arXiv.org · Vol abs/2607.13437 · 0 citations · 25 references
Computer Science

TL;DR

A unified CLIP-guided label-free region scoring framework for fine-grained classification is proposed, and random-crop-based pseudo-label scoring consistently outperforms SAM-based scoring across all datasets, suggesting that random crops preserve surrounding information and provide more stable semantic context when pseudo-labels are noisy.

Abstract

Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is to compare the representations of segmented image regions with the text prompt embeddings of the corresponding labels. However, it remains unclear how different local regions and CLIP-based scoring strategies affect the selection of discriminative evidence, especially when ground-truth labels are unavailable. In this paper, we propose a unified CLIP-guided label-free region scoring framework for fine-grained classification. The framework evaluates cosine similarity-based, margin-based, and entropy-based scoring strategies using both SAM-generated masks and random crops, and introduces two label-free pseudo-label variants based on global image embeddings and local region embeddings. We conduct experiments on five fine-grained classification datasets to systematically compare different region generation methods and scoring strategies. The results show that Soft Negative Margin scoring achieves the strongest performance, and pseudo-label scoring closely approximates true-label performance. Although SAM produces semantically meaningful masks, random-crop-based pseudo-label scoring consistently outperforms SAM-based scoring across all datasets, suggesting that random crops preserve surrounding information and provide more stable semantic context when pseudo-labels are noisy. In addition, SAM masks benefit from aggregating embeddings from all regions, whereas random crops tend to perform better with a smaller top-k subset. These findings provide new insights for fine-grained classification.

View source

Similar papers

Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations

Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning

Class-wise Covariance Regularization is proposed, which aligns the predicted covariance structure of class confidences with the semantic correlations encoded in pretrained text embed-dings with the geometric consistency of the class space throughout fine-tuning, resulting in more stable and interpretable confidence dis...

Ao Zhou, Zhi-Wei Jiang, Zi-Feng Cheng et al. · 0 citations
Open access Aug 2026

Few-shot image classification algorithm based on deep learning and feature fusion

Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...

Huie Zhang, Mary Jane C. Samontet · 0 citations
Open access Jul 2026

CANNE: CLIP-Based ANNE Selection for Noisy-Label Learning

Experimental results on CIFAR-10, CIFAR-100, Animal-10N, and Mini-WebVision, together with additional evaluation under open-set noise, show that the proposed CANNE method achieves competitive performance across diverse noisy-label settings.

Ge Jin, Qian Zhang, Li Huang et al. · 0 citations
Aug 2026

Text-guided zero-shot localization of unseen object categories

A text-guided Zero-Shot Localization framework for unseen object categories (ZSOL) for addressing the aforementioned challenges, which can be guided by prompt words to identify and localize unseen object categories in images by transferring localization knowledge learned from supervised base categories.

Jingjing Wang, Xing-Lin Piao, Zongzhi Gao et al. · 0 citations
Aug 2026

MSPD-net: structural–appearance prototype decoupling for weakly supervised semantic segmentation

A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.

Wei Cao, Yong Jiang, Ruiying Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.