2026· Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada· 0 citations
TL;DR
A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.
Abstract
Multimodal Sentiment Analysis (MSA) aims to integrate text, audio, and visual modalities to achieve accurate sentiment modeling. Existing methods often rely on shallow interaction structures, making it difficult to jointly capture fine-grained local dynamics and high-level semantic dependencies. In addition, dominant modalities may suppress the learning of weaker modalities, leading to insufficient cross-modal semantic alignment. To address these issues, we pro-pose a Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR). Specifically, a Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness. Based on this, a Hierarchical Local Inter-action (HLI) module models multimodal sequences through a multi-layer pro-gressive structure to capture local dynamic features at different semantic levels. Within the HLI module, a Cross-Modal Synergistic Learning (CMSL) mecha-nism explicitly models cross-modal semantic consistency and gradually aligns information during interaction. Furthermore, a Global Representation Refine-ment (GRR) module introduces learnable global representations and iteratively updates them in a multi-layer structure to aggregate long-range semantic de-pendencies and form stable high-level semantic representations. Experimental results on CMU-MOSI and CMU-MOSEI demonstrate the effectiveness of the proposed framework across multiple evaluation metrics.
DualScope is proposed, a novel model that combines a global-local fusion strategy with bidirectional image-text generation for semantically consistent data augmentation and introduces both label contrastive learning and data contrastive learning to align heterogeneous modalities and enhance model robustness.
Bing Zhang, Junteng Wang, Bin Sun et al.· Memetic Computing· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address these challenges, this paper proposes a framework for learning adaptive cross-modal interactions for multimodal sentiment analysis. The proposed framework consists of three stages: modality-aware preprocessing, heterogeneous representation learning, and adaptive multimodal fusion. First, a unified preprocessing strategy is designed to improve cross-modal consistency through textual normalization, speaker-aware visual alignment, and utterance-level acoustic representation enhancement. Second, modality-specific encoders are constructed to capture complementary semantic, spatial, and utterance-level acoustic characteristics from textual, visual, and acoustic modalities, respectively. Third, an adaptive fusion framework is introduced to explicitly model cross-modal interactions, dynamically estimate the importance of different modality combinations, and further calibrate discriminative feature channels through channel attention. By jointly performing modality-level interaction learning and channel-wise feature refinement, the proposed framework effectively enhances multimodal representation capability for sentiment classification. Extensive experiments conducted on the CMU-MOSI and MELD benchmark datasets demonstrate that our framework consistently outperforms previous methods. In particular, the proposed model achieves 90.27% accuracy and 90.26% F1-score on CMU-MOSI, together with 66.57% accuracy and 66.21% F1-score on MELD. Additional ablation studies and qualitative analyses further validate the effectiveness of the proposed preprocessing strategy, modality-specific representation learning, and adaptive fusion mechanism.
Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al.· International Conference on...· 0 citations
MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.