Preprint
Jul 2026
Robustifying Vision-Language Models via Test-Time Prompt Adaptation
This work proposes RITA, a Robust test-tIme prompt-TAdaptation framework that shifts from sample-level estimates to distribution-level alignment, and employs optimal transport to align the distribution of augmented visual features with textual prototypes, mitigating adversarial outliers and rectifying cross-modal semantic misalignment.
Xingyu Zhu, Huanshen Wu, Shuo Wang et al.
· 1 citation