Skip to content
Preprint

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

Aug 2026 · 0 citations · 63 references
Computer Science

TL;DR

This work proves the mechanism and bound its share of the damage at 26% by a counterfactual, and finds the ceiling in the model class, not the hyperparameter: the ratio can be set near-optimally for free, and it is still not where the performance is.

Abstract

Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and the mean of the K labelled image features, with a single blending ratio routinely tuned on held-out labels, often on the test set itself. We ask what the family's own bias-variance justification invites: what is the right ratio, can it be estimated without validation data, and is finding it where the performance is? First, the ratio minimising prototype mean-squared error has a closed form whose support-set plug-in is exactly a positive-part James-Stein coefficient shrinking towards the text prototype. Across 4,800 cells (ten datasets, five backbones including SigLIP, five shot counts, five seeds, four prompt tiers) this theoretically optimal ratio is a reliable estimate of the wrong quantity: on the 950 primary-tier cells where it is defined it trails a test-set-oracle ratio by 8.5 points. It saturates near 1, discarding the text prior for a nearest-class-mean classifier, because 78% of the text-image prototype distance it treats as bias is a class-independent offset that the arg max largely cancels. We prove the mechanism and bound its share of the damage at 26% by a counterfactual. Second, leave-one-out on the support set alone sets a ratio landing within 0.9 points of the oracle blend, so it is estimable without validation data. Third, validation-free linear probes beat even the oracle-tuned blend: CLAP by +1.9 points and LP++ by +1.5 on average, and at K>= 4 all four validation-free baselines sit above the oracle, the linear probes by margins excluding zero. These results locate the ceiling in the model class, not the hyperparameter: the ratio can be set near-optimally for free, and it is still not where the performance is. Code, cached features, per-cell records: https://huggingface.co/datasets/Liangzhi-Li/clipbench-blending

View source

Similar papers

Preprint Sep 2026

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting...

Shangzhe Di, Zhaokai Wang, Wei-Di Xie · 1 citation
Conference Aug 2026

Robust Zero-Shot Learning with Distribution-Preserving Feature Generation and Bias-Calibrated Classification

The ability to recognize objects out of our environment with the use of additional information such as attribute vectors or text embeddings-this is what people call Zero-Shot Learning. Believe me, there are many more ZSL techniques that get wrong in real-life context - they're not always semantically correct with visua...

I. K, A. Meenakshi · 0 citations
#artificial intelligence Preprint Sep 2026

Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting...

Gautam Rajendrakumar Gare, Si-Ying Li, He-Wei Wang et al. · 0 citations
Review Jul 2026

Explaining AI-Image Detection: What the Heatmap Actually Shows

This work builds a detector and attaches an attribution map as its evidence, then measures what that pair delivers on 186,527 images under controls designed to change the authors' conclusions when something is wrong, finding none.

Leonid Kuturin, Ilya Sotnikov, Mark Khusnutdinov et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sample-Conditioned Representation Selection for Audio Few-Shot Learning

Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propos...

Feng-Rui Liu, Ning-Xin Shen, Yi Li et al. · 0 citations
Preprint Aug 2026

Scalable Black-Box Model Attribution for Images

A lightweight CNN that attributes more models at higher accuracy than prior work, reaching 98.9% on 25-class DRAGON and 95.0% on 27-class OpenFake; runs in a few milliseconds at a cost nearly independent of candidate-set size; and remains robust to transformations encountered in the wild.

Asaf Livne, Amir Jevnisek, S. Avidan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.