BiReg: Bilateral Regularized Kernel Adaptation for Training-Free Few-Shot Adaptation of Vision–Language Models
Abstract
Large vision-language models (VLMs) demonstrate impressive zero-shot capabilities but exhibit limited adaptability in few-shot learning and domain transfer scenarios. Current adaptation methods present a critical tradeoff: gradient-based fine-tuning achieves strong performance but requires extensive computation and risks overfitting with scarce training samples, while training-free approaches sacrifice modeling of inter-class semantic relationships for efficiency. We propose Bilateral Regularized Kernel Adaptation (BiReg), a closed-form, training-free framework that formulates few-shot VLM adaptation as multi-output kernel ridge regression. BiReg introduces bilateral regularization that simultaneously constrains the visual feature space and the semantic label embedding space. This dual regularization enables the model to exploit inter-class correlations while preserving alignment with the pretrained zero-shot prior. Experiments across ten diverse benchmarks demonstrate that BiReg achieves superior accuracy and robustness while retaining the efficiency of a closed-form, training-free solution.