This paper uses a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model, and proposes a framework that leverages foundation models (LFM) for SF-UniDA.
Abstract
Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.
COSMO replaces expert-to-expert guidance with co-adaptation through an anchored shared consensus and achieves state-of-the-art performance under matched VLM backbones, indicating that it better balances the retention of valid source-derived evidence with the absorption of complementary VLM evidence.
Bo Li, Junjie Peng, Xiaohua Xie et al.· 0 citations
A Peer-level Heterogeneous Perception Framework is proposed that departs from such paradigms by enabling balanced collaboration between heterogeneous models by introducing an auxiliary domain that is significantly different from the target domain and employ an auxiliary model with the same architecture as the source mo...
Zhi-Ze Wu, Yu-Tao Fu, Huan-Xin Zou et al.· Multimedia Systems· 0 citations
This work proposes a novel criterion, termed Maximum Refinement Decision score (MRD-score), which replaces softmax with median centering to normalize source models’ predictions along both positive and negative axes, thereby harnessing both affirmative and complementary guidance.
Bing-Tao Zhou, Mian Xiang, Qian Ning· Journal of King Saud Univers...· 0 citations
This paper proposes a novel SFDA with high-confidence sample selection and feature disentanglement for machinery fault diagnosis, which effectively alleviates the adverse influence of noisy pseudo-labels during the stage of adaptation.
Yi-Ming Yuan, Kang Wu, Xing-Xing Jiang et al.· Measurement science and tech...· 0 citations
This work introduces CAST (Closed-form Analytic Semantic Transfer), a training-free, image-free framework for extending a pre-trained classifier to previously unseen classes through weight injection and derives a finite-sample error decomposition that identifies the semantic extrapolation residual.
W. Heyden, Habib Ullah, M. Siddiqui et al.· 0 citations
This paper proposes ADA-CS, a plug-and-play module compatible with any ADA or ASFDA framework, and introduces a CSS metric to quantify the Concept Shift Severity across domains, revealing that non-negligible concept shift exists in many transfer tasks.