Skip to content

Plug-and-play Feature Causality Decomposition for Multimodal Representation Learning

2025 · Neural Information Processing Systems · pp. 65631-65657 · 3 citations · 47 references
Computer Science

TL;DR

This work proposes a plug-and-play feature causality decomposition method for multimodal representation learning from causality perspective, which can be integrated into existing models with no affects on the original model structures.

Abstract

Multimodal representation learning is critical for a wide range of applications, such as multimodal sentiment analysis. Current multimodal representation learning methods mainly focus on the multimodal alignment or fusion strategies, such that the complementary and consistent information among heterogeneous modalities can be fully explored. However, they mistakenly treat the uncertainty noise within each modality as the complementary information, failing to simultaneously leverage both consistent and complementary information while eliminating the aleatoric uncertainty within each modality. To address this issue, we propose a plug-and-play feature causality decomposition method for multimodal representation learning from causality perspective, which can be integrated into existing models with no affects on the original model structures. Specifically, to deal with the heterogeneity and consistency, according to whether it can be aligned with other modalities, the unimodal feature is first disentangled into two parts: modality-invariant (the synergistic information shared by all heterogeneous modalities) and modality-specific part. To deal with complementarity and uncertainty, the modality-specific part is further decomposed into unique and redundant features, where the redundant feature is removed and the unique feature is reserved based on the backdoor-adjustment. The effectiveness of noise removal is supported by causality theory. Finally, the task-related information, including both synergistic and unique components, is further fed to the original fusion module to obtain the final multimodal representations. Extensive experiments show the effectiveness of our proposed strategies.

View source

Similar papers

#machine learning Preprint Sep 2026

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these ta...

Wan-Ting Huang, Sanvesh Srivastava, Wei-Ran Wang · 0 citations
#machine learning Preprint Sep 2026

SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes int...

Yavuz Yarici, Ghassan AlRegib · 0 citations
Open access Sep 2026

DiMoE: Disentangled Representation Learning with Mixture of Experts Fusion for Sentiment Intensity Prediction and Emotion Classification

Multimodal affective analysis benefits from combining textual, acoustic, and visual cues, yet many methods implicitly mix modality-invariant information with modality-specific factors, which can reduce robustness when modalities vary in reliability. We propose DiMoE, a representation-first framework that integrates fea...

M. S. Sran, S. Ramanna, K. Kotecha · 0 citations
Review Sep 2026

Dependency, Compression, and Synergy: A Unified Information-Theoretic View of Multimodal Learning

Recent advances in multimodal foundation models have intensified the need to understand how different modalities share, preserve, and complement information. Mutual Information (MI), the Information Bottleneck (IB), and Partial Information Decomposition (PID) provide complementary perspectives, yet existing studies oft...

Liang-Jian Wen, Lin Li, Jiang Duan et al. · 1 citation
Open access Sep 2026

Cross-Modal Representation Learning for Integrating Heterogeneous Data in AI Systems

The increasing availability of heterogeneous data sources, including text, images, and structured records, has intensified the need for robust multimodal artificial intelligence systems. However, existing multimodal learning approaches often rely on simplistic fusion strategies and struggle to capture deep semantic rel...

I. Ibrahim, M. Alshar'e, I. Sanjaya et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.