MDRAN: a multi-source dual-refining adaptation network for source-free domain adaptation guided by vision-language-model
Abstract
Source-Free Domain Adaptation (SFDA) adapts a pre-trained source model to an unlabeled target domain without accessing source data, alleviating the need for direct access to source data during adaptation, which reduces the risk of data transmission. Existing methods leverage pre-trained Vision-Language (ViL) models to provide supervisory signals, but their predictions are often noisy, limiting adaptation performance. Moreover, current approaches refine ViL predictions only within individual source domains, overlooking the complementary discriminative evidence available across multiple sources. To better select the most decisive refinement signal from each source model, we propose a novel criterion, termed Maximum Refinement Decision score (MRD-score), which replaces softmax with median centering to normalize source models’ predictions along both positive and negative axes, thereby harnessing both affirmative and complementary guidance. Building upon MRD-score, we present Multi-source Dual-Refining Adaptation Network (MDRAN), which, to the best of our knowledge, is among the first attempts to exploit ViL-based guidance for multi-source-free domain adaptation (MSFDA). MDRAN first aggregates the MRD-score-based predictions from multiple source models to construct Selected Source-Logit Set (SLS), which is then used to customize a collective ViL model via prompt learning. To mitigate noise in the collective ViL models’ outputs, we design a Dual-View Refinement (DVR) module that refines ViL model’s predictions using source model predictions. Extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art performance on the more challenging Office-Home, VisDA-C, and DomainNet-126, while remaining highly competitive on the comparatively easy Office-31. The code can be found in https://anonymous.4open.science/r/MDRAN-098E/.