Unlike test-time adaptation (TTA), open-set TTA (OSTTA) aims to robustly adapt to in-distribution (ID) domain shifts while suppressing the detrimental impact of out-of-distribution (OOD) samples encountered at test time. To this end, we propose a novel OSTTA approach named Noise-Immune Self-Purification (NISP). By exploiting the zero-shot priors of pre-trained vision-language models (VLMs), NISP advances from coarse pseudo-labeling to fine-grained, noise-resilient adaptation. Technically, we first introduce a dual-cluster Bayesian Gaussian mixture model to fit VLMs-derived scores, achieving coarse pseudo-ID/OOD separation via posterior-risk thresholding. Subsequently, NISP constructs a noise-immune fine-grained adaptation where the adapter enforces consensus-discrepancy constraints to refine coarse pseudo-labels and suppress noise propagation. We then devise the Jaccard Consistency Score for OOD discrimination. Overall the coarse-to-fine pipeline enables rigorous self-purification and robust online adaptation. Theoretically we show that consensus-discrepancy losses mitigate the deleterious effects of noise. Empirically, NISP achieves state-of-the-art results across multiple OSTTA benchmarks, validating its efficacy. The code is available at https://github.com/njustkmg/IJCAI26-NISP.
Ming-Xu Feng, Feng-Qiang Wan, Yang Yang· Proceedings of the Thirty-Fi...· 0 citations
Despite significant advancements in multimodal learning (MML), it has been unexpectedly shown to underperform compared to unimodal approaches in practice, largely due to the modality imbalance problem, ultimately affecting the overall performance of the model. Naturally, most existing methods aim to rebalance optimization speeds across different modalities to avoid performance degeneration caused by modality imbalance. However, in addition to task-oriented modality fusion, we experimentally find that multimodal learning requires explicit modality alignment to stimulate weak modal capabilities so that they can be fully exploited, which is ignored by existing works. Therefore, in this paper, we explore the impact of modality fusion and alignment on multimodal learning from a unified perspective, and develops a dynamic strategy that jointly optimizes both, with particular emphasis on addressing modality imbalance. Concretely, we initially design a soft alignment strategy to impose the positive intervention from the prediction level by integrating modality fusion and alignment into a unified framework. We further extend this strategy to the representation level and hybrid level, enabling compatibility with a wider range of architectures. Subsequently, we design a heuristic strategy to dynamically integrate fusion and alignment. Furthermore, we develop a learning-based strategy using a bi-level optimization framework and theoretically prove the convergence of the learning algorithm to ensure its reliability. These two dynamic integration strategies are incorporated into a unified framework applicable to both supervised and semi-supervised scenarios, further enhancing performance. We conduct a series of experiments to demonstrate the effectiveness of our method on diverse datasets. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art multimodal learning approaches, achieving accuracy improvements of 1.30%, 2.69%, and 0.60% on representative bimodal benchmarks, namely KSounds, CREMA-D, andSarcasm, respectively, as well as gains of 1.35% and 0.65% on trimodal datasets, namely NVGesture and IEMOCAP.
Yang Yang, Fengqiang Wan, Qingjun Jiang et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.