A unified CLIP-guided label-free region scoring framework for fine-grained classification is proposed, and random-crop-based pseudo-label scoring consistently outperforms SAM-based scoring across all datasets, suggesting that random crops preserve surrounding information and provide more stable semantic context when pseudo-labels are noisy.
Abstract
Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is to compare the representations of segmented image regions with the text prompt embeddings of the corresponding labels. However, it remains unclear how different local regions and CLIP-based scoring strategies affect the selection of discriminative evidence, especially when ground-truth labels are unavailable. In this paper, we propose a unified CLIP-guided label-free region scoring framework for fine-grained classification. The framework evaluates cosine similarity-based, margin-based, and entropy-based scoring strategies using both SAM-generated masks and random crops, and introduces two label-free pseudo-label variants based on global image embeddings and local region embeddings. We conduct experiments on five fine-grained classification datasets to systematically compare different region generation methods and scoring strategies. The results show that Soft Negative Margin scoring achieves the strongest performance, and pseudo-label scoring closely approximates true-label performance. Although SAM produces semantically meaningful masks, random-crop-based pseudo-label scoring consistently outperforms SAM-based scoring across all datasets, suggesting that random crops preserve surrounding information and provide more stable semantic context when pseudo-labels are noisy. In addition, SAM masks benefit from aggregating embeddings from all regions, whereas random crops tend to perform better with a smaller top-k subset. These findings provide new insights for fine-grained classification.
G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.
Zehua Hao, Fang Liu, Qinliang Wang et al.· 0 citations
Class-wise Covariance Regularization is proposed, which aligns the predicted covariance structure of class confidences with the semantic correlations encoded in pretrained text embed-dings with the geometric consistency of the class space throughout fine-tuning, resulting in more stable and interpretable confidence dis...
Ao Zhou, Zhi-Wei Jiang, Zi-Feng Cheng et al.· 0 citations
Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...
Huie Zhang, Mary Jane C. Samontet· International journal of com...· 0 citations
Experimental results on CIFAR-10, CIFAR-100, Animal-10N, and Mini-WebVision, together with additional evaluation under open-set noise, show that the proposed CANNE method achieves competitive performance across diverse noisy-label settings.
Ge Jin, Qian Zhang, Li Huang et al.· Entropy· 0 citations
A text-guided Zero-Shot Localization framework for unseen object categories (ZSOL) for addressing the aforementioned challenges, which can be guided by prompt words to identify and localize unseen object categories in images by transferring localization knowledge learned from supervised base categories.
A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.