Understanding Cross-Modal Contributions in Continual Vision-Language Models: A Theoretical Perspective
This paper presents a new theoretical perspective to understand the cross-modal (vision-language) contributions to consecutive environments, and provides deeper insights into continual VLMs, highlighting their contribution robustness to varying task orders and inter-task similarities, and their improved generalization...