Skip to content
Book Open access

ADMC: Attention-based Diffusion model for Missing modalities Completion

Oct 2026 · 0 citations

Abstract

Multimodal emotion and intent recognition is essential for automated human-computer interaction, and it aims to analyze users’ speech, text, and visual information to predict their emotions or intent. One of the significant challenges is modality missingness caused by sensor malfunctions or incomplete data. Traditional methods that attempt to reconstruct missing information often suffer from over-coupling and imprecise generation processes, leading to suboptimal outcomes. To address these issues, we introduce an Attention-based Diffusion model for Missing Modalities feature Completion (ADMC). Our framework independently trains feature extraction networks for each modality, preserving their unique characteristics and avoiding over-coupling. The Attention-based Diffusion Network (ADN) generates missing modality features that closely align with authentic multimodal distribution, enhancing performance across all missing-modality scenarios. Moreover, ADN’s cross-modal generation offers improved recognition even in full-modality contexts. Our approach achieves state-of-the-art results on the IEMOCAP and MIntRec datasets, with up to 9.4% improvement in Average Weighted Accuracy, demonstrating its effectiveness.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.