This work introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities and designs a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations.
Abstract
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.
A trainable reliability-aware evidential fusion framework that estimates not only sentiment predictions but also modality-specific evidence, predictive uncertainty, observable input quality, cross-modal disagreement, and normalized sample-dependent fusion weights is presented.
A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.
Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al.· International Journal of Dat...· 0 citations
A Variational Autoencoder-Based Multi-Granularity Missing-Aware Dual-Path Fusion Network (VMMD) with contrastive learning-driven Cross-Path Semantic Consistency Constraint (CSCC), which mitigates semantic representation inconsistency caused by granularity heterogeneity between supplementary information and proxy featur...
Jia-Xu Li, Wei Liu· International Journal of Mac...· 0 citations
SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment, is proposed.
Wen-Hao Li, Zhi-Bin Wu, Chong-Yao Xiao et al.· 0 citations
Multimodal sentiment analysis faces several fundamental challenges in real-world applications, including modality heterogeneity, noisy and unreliable signals, and suboptimal fusion strategies. To address these issues, a text-guided and confidence-aware dynamic fusion framework (TGCF) is proposed. The method introduces...
Ye Zhang, Hai-Tao Yang, Yong-Lin Leng· 2026 6th International Confe...· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026