XSentiFusionNet: An Explainable Cross-Modal Attention Framework For Audio-Visual Sentiment Analysis Using Hybrid Deep Learning
Abstract
The exponential growth of audio-visual content created through social media, communication platforms as well as human-computer interaction has led to a need for effective multimodal sentiment analysis. Most of the multimodal frameworks have limitations in cross-modal interaction modeling, fusion strategy adaptability, interpretability, and robustness to noise. This paper introduces XSentiFusionNet, an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI. Our approach uses CNNs, BiLSTMs, and ViTs to capture emotional features while adaptively combining acoustic and visual modalities based on their reliability levels. Model transparency is achieved by integrating SHAP, LIME, Grad-CAM, and attention mechanisms. Comprehensive experiments were performed on CMU-MOSEI, MELD, and RAVDESS benchmark datasets using comparative analysis, ablation studies, cross-dataset generalization experiments, confusion matrix analysis, explainability, and robustness to noise experiments. Our framework achieved an accuracy of 94.82%, F1-score of 94.11% and a ROC-AUC score of 96.04% on the benchmark datasets performing better than existing approaches such as CNN-LSTM, Transformer Fusion, and Multimodal BERT. Additionally, the framework showed higher robustness to noise and generalization ability on different multimodal datasets while the XAI module demonstrated interpretability by highlighting key speech and facial features used for predictions. These results show that XSentiFusionNet is a reliable and efficient framework for audio-visual sentiment analysis and can be used in real-world multimodal processing and affective computing scenarios.