TGCF: Text-Guided and Confidence-Aware Dynamic Fusion for Multimodal Sentiment Analysis
Abstract
Multimodal sentiment analysis faces several fundamental challenges in real-world applications, including modality heterogeneity, noisy and unreliable signals, and suboptimal fusion strategies. To address these issues, a text-guided and confidence-aware dynamic fusion framework (TGCF) is proposed. The method introduces a text-guided gating mechanism to refine acoustic and visual representations by suppressing irrelevant and noisy information. Meanwhile, a confidence-aware adaptive weighting strategy is designed to estimate modality reliability at the sample level and dynamically adjust modality contributions during fusion, resulting in more robust multimodal representation learning. Extensive experiments conducted on the CMU-MOSI and CMU-MOSEI datasets demonstrate that TGCF achieves improvements in both Acc-2 and F1-score, validating the effectiveness of text-guided noise suppression and confidence-based adaptive fusion.