Alignment-Consistent Multimodal Learning under Uncertain Correspondence.
Multimodal learning aims to integrate heterogeneous observations such as images, text, depth, and radar to improve perception and reasoning. However, most existing multimodal models implicitly assume that cross-modal observations are well aligned, an assumption that rarely holds in real-world scenarios due to viewpoint...