Skip to content

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

Sep 2026 · 0 citations · 70 references
Computer Science

TL;DR

This work introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities and designs a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations.

Abstract

Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.

View source

Similar papers

Open access Aug 2026

Reliability-Aware Multi-Modal Sentiment Analysis Under Missing and Corrupted Modalities

A trainable reliability-aware evidential fusion framework that estimates not only sentiment predictions but also modality-specific evidence, predictive uncertainty, observable input quality, cross-modal disagreement, and normalized sample-dependent fusion weights is presented.

Yu-Bin Wu, Xian-Xun Zhu, Huilin Liu · 0 citations
Sep 2026

Reliability-aware disentangled adaptive network for multimodal sentiment analysis

A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.

Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al. · 0 citations
Open access Sep 2026

Register-augmented attention and self-calibrated fusion for robust multimodal sentiment analysis

Results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis, and RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification...

Bin Xu, Wei-Yang Wang, Ao Ding · 0 citations
Sep 2026

Dual-path fusion network for multimodal sentiment analysis under uncertain missing modalities

A Variational Autoencoder-Based Multi-Granularity Missing-Aware Dual-Path Fusion Network (VMMD) with contrastive learning-driven Cross-Path Semantic Consistency Constraint (CSCC), which mitigates semantic representation inconsistency caused by granularity heterogeneity between supplementary information and proxy featur...

Jia-Xu Li, Wei Liu · 0 citations
Conference Aug 2026

TGCF: Text-Guided and Confidence-Aware Dynamic Fusion for Multimodal Sentiment Analysis

Multimodal sentiment analysis faces several fundamental challenges in real-world applications, including modality heterogeneity, noisy and unreliable signals, and suboptimal fusion strategies. To address these issues, a text-guided and confidence-aware dynamic fusion framework (TGCF) is proposed. The method introduces...

Ye Zhang, Hai-Tao Yang, Yong-Lin Leng · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.