This paper introduces the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations, and introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations.
Abstract
Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase, which leads to insufficient generalization capabilities and unstable performance. In this paper, we introduce the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations. To address this challenge, we propose the model named $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) for IMSAUMC. It introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations. Additionally, we design modality-combination prompts with a soft router to facilitate better learning of various modality combinations. Furthermore, we introduce three prompt contrastive learning strategies, which enable effective learning of prompts corresponding to unseen modality combinations, thereby significantly strengthening the model's generalization capabilities in diverse testing scenarios. Extensive experiments on three widely used datasets demonstrate that $\textsf{CMPL}$ achieves more than a 5% improvement in accuracy compared to state-of-the-art approaches.
The framework first introduces learnable sentiment prototypes as semantic anchors to provide explicit sentiment-discriminative guidance for feature completion, and a gradient decoupling strategy is designed to separate the optimization paths of unimodal and multimodal objectives, preventing fusion gradients from interfering with unimodal encoders, thereby synergistically enhancing both discriminative representation learning and multimodal fusion.
Shan Tao, Haipeng Chen, Yu Liu et al.· Multimedia Systems· 0 citations
Multimodal sentiment analysis aims to extract affective signals from multiple modalities and perform feature analysis, modality fusion and sentiment prediction. However, in practical application scenarios, input modality information is highly prone to be missing due to various objective factors, which further causes the model to produce biased or even completely erroneous sentiment judgment results. To address this critical issue, a model named Modal Experts and Missing Prompt Generation (MEMPG) is proposed. The model employs a brand-new dual-residual modality expert architecture to integrate the knowledge of hybrid modality experts. This architecture can not only retain the original information but also preserve the contextual information learned by the attention mechanism, thus exhibiting better robustness. The model adopts a twostage training strategy: In the first stage, each modality expert branch is independently pre-trained to extract low-level basic features of a single modality and strengthen the feature extraction capability of modality experts. In the second stage, all input modality features are fed into all modality expert branches for processing, and adaptive weights are assigned to modality features via an adaptive gating mechanism to obtain fused modality features with stronger representation ability. Subsequently, the enhanced features are input into the missing modality prompt generation module, which incorporates prompt learning to guide the reconstruction of missing modalities. Extensive comparative experiments are conducted on two standard multimodal sentiment datasets, namely MOSI and MOSEI, under various common modality missing scenarios. Experimental results demonstrate that, compared with current mainstream baseline models, the proposed MEMPG model achieves superior sentiment prediction performance.
Shuai Liu, Xuan-Yu Wu· 2026 IEEE International Conf...· 0 citations
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.
Bin Xu, Wei-Yang Wang, Ao Ding· Multimedia Systems· 0 citations
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al.· International Conference on...· 0 citations
Consistency-Aware Gated Fusion (CAGF), a lightweight and fusion module tailored to Mamba-based architectures that achieves state-of-the-art performance, outperforming strong multimodal baselines such as CLIP, MISA, DLF, AoM, and SFTTR, while remaining more efficient and interpretable.
Jian Hu· Poster Volume 0008 The 2026...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.