Skip to content
Preprint

CAST: Closed-form Analytic Semantic Transfer for Zero-Shot Classifier Extension

Aug 2026 · 0 citations · 54 references
Computer Science

TL;DR

This work introduces CAST (Closed-form Analytic Semantic Transfer), a training-free, image-free framework for extending a pre-trained classifier to previously unseen classes through weight injection and derives a finite-sample error decomposition that identifies the semantic extrapolation residual.

Abstract

Large pre-trained models have become foundational components of modern machine learning systems. Yet adapting these models to novel categories typically requires examples from the target distribution. In many domains, however, such data are unavailable. Zero-shot learning (ZSL) permits recognition under these limitations through relying on auxiliary semantic information such as textual descriptions. We introduce CAST (Closed-form Analytic Semantic Transfer), a training-free, image-free framework for extending a pre-trained classifier to previously unseen classes through weight injection. We provide a theoretical foundation for CAST and derive a finite-sample error decomposition that identifies the \emph{semantic extrapolation residual} $\rho_u$. The residual is a computable, model-agnostic measure and provides a principled criterion for dataset curation and benchmark design. Experiments on standard zero-shot learning benchmarks demonstrate that CAST matches or exceeds existing image-free approaches and approaches the performance of few-shot adaptation methods, while requiring neither iterative optimization nor examples from the target distribution.

View source

Similar papers

2026

Toward Zero-Forgetting: A Training-Free Multimodal Framework for Remote Sensing Class-Incremental Learning

Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...

Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster...

Youngeun Seol, Jiyon Shin, Hee Suk Yoon et al. · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations
Preprint Open access Sep 2026

Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs

Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard ad...

Hung-Jen Chen, Yuek F. Ho, Ting-Yao Huang et al. · 0 citations
Preprint Sep 2026

Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification

Few-shot recognition with frozen visual features is especially fragile under domain shift and one-shot supervision, where a single labelled image is an unreliable estimate of its class. We ask how far this fragility can be reduced purely at test time, without retraining the encoder or augmenting the source domain. We p...

F. Rahman, S. Rohan, Mahmound Sayed et al. · 0 citations
Sep 2026

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Xilang Huang, Seon-Han Choi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.