2025· Neural Information Processing Systems· pp. 65631-65657· 3 citations· 47 references
Computer Science
TL;DR
This work proposes a plug-and-play feature causality decomposition method for multimodal representation learning from causality perspective, which can be integrated into existing models with no affects on the original model structures.
Abstract
Multimodal representation learning is critical for a wide range of applications, such as multimodal sentiment analysis. Current multimodal representation learning methods mainly focus on the multimodal alignment or fusion strategies, such that the complementary and consistent information among heterogeneous modalities can be fully explored. However, they mistakenly treat the uncertainty noise within each modality as the complementary information, failing to simultaneously leverage both consistent and complementary information while eliminating the aleatoric uncertainty within each modality. To address this issue, we propose a plug-and-play feature causality decomposition method for multimodal representation learning from causality perspective, which can be integrated into existing models with no affects on the original model structures. Specifically, to deal with the heterogeneity and consistency, according to whether it can be aligned with other modalities, the unimodal feature is first disentangled into two parts: modality-invariant (the synergistic information shared by all heterogeneous modalities) and modality-specific part. To deal with complementarity and uncertainty, the modality-specific part is further decomposed into unique and redundant features, where the redundant feature is removed and the unique feature is reserved based on the backdoor-adjustment. The effectiveness of noise removal is supported by causality theory. Finally, the task-related information, including both synergistic and unique components, is further fed to the original fusion module to obtain the final multimodal representations. Extensive experiments show the effectiveness of our proposed strategies.
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these ta...
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes int...
Multimodal affective analysis benefits from combining textual, acoustic, and visual cues, yet many methods implicitly mix modality-invariant information with modality-specific factors, which can reduce robustness when modalities vary in reliability. We propose DiMoE, a representation-first framework that integrates fea...
M. S. Sran, S. Ramanna, K. Kotecha· Algorithms· 0 citations
Recent advances in multimodal foundation models have intensified the need to understand how different modalities share, preserve, and complement information. Mutual Information (MI), the Information Bottleneck (IB), and Partial Information Decomposition (PID) provide complementary perspectives, yet existing studies oft...
Liang-Jian Wen, Lin Li, Jiang Duan et al.· 1 citation
The increasing availability of heterogeneous data sources, including text, images, and structured records, has intensified the need for robust multimodal artificial intelligence systems. However, existing multimodal learning approaches often rely on simplistic fusion strategies and struggle to capture deep semantic rel...
I. Ibrahim, M. Alshar'e, I. Sanjaya et al.· Journal of Data Science· 0 citations