Skip to content
Open access

Bridging Visual and Textual Cues: Cross-Attention Fusion for Multi-Modal Sentiment Analysis

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 27 references

TL;DR

A hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN) is presented, a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals that outperforms single-modal sentiment analysis.

Abstract

Uni-modal studies on sentiment analysis (SA) has made significant strides, emotions in the real world, which are typically multi-modal, encompassing audio, images, video and other contents more than text. The several modes contribute to mutual development. The accuracy of sentiment evaluation will be enhanced even more if it is possible to mine the connections between different modalities. This work presents a hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN), a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals. Text features are first extracted using the pre-training model; text context features are then extracted CNN; image characteristics are extracted using the network model; and specific emotion-related areas in images are extracted using DAE+CNN. The retrieved characteristics of the text and image are then fused using multi-modal cross-attention, and the output is classified to ascertain the emotional polarity. The model presented in this research outperforms the baseline model in the comparison study of the Fashion dataset and Deep Fashion datasets with accuracy of 86.5 % and 85.5% and F1-scores of 75.3% and 76.7%, respectively,. Furthermore, authors carried out ablation studies, which verified that multi-modal fusion sentiment evaluation outperforms single-modal sentiment analysis.

Read PDF

Similar papers

Conference Jul 2026

Multimodal Sentiment Analysis Through Deep Learning: Leveraging Early Fusion Using Translation Alignment

As the amount of text and image data on social media continues to increase, multimodal sentiment analysis has emerged as a critical area of study. However, a "modality gap" that lowers classification accuracy is frequently caused by the domain disparities between textual and visual modalities. In this paper, a Translat...

Revano Fabiansyah Priadi, Arie Ardiyanti Suryani · 0 citations
Open access Jul 2026

XSentiFusionNet: An Explainable Cross-Modal Attention Framework For Audio-Visual Sentiment Analysis Using Hybrid Deep Learning

XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.

M. Kidwai, C. Author, Dr. Faiyaz Ahmad · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations
Open access Aug 2026

Attention-Enhanced Multimodal Sentiment Analysis Using Resnet50-Cbam, BERT, And Graph Neural Networks

The present study proposes such a multimodal sentiment analysis framework with attention-enhanced properties, a combination of ResNet50 and Convolutional Block Attention Module (CBAM), a textual encoder with BERT, and refinement of relational features via Graph Neural Networks (GNN). The model is designed to address th...

K. Mounika, B. V. RamNaresh Yadav · 0 citations
Aug 2026

TSRB: Transformer-based Semantic Refinement Block for Sentiment Analysis using Scene Text Images

This work aims to use scene text images for sentiment analysis to assist in understanding the intentions of captured scenes and presents TSRB (Transformer-based Semantic Refinement Block), which comprises a multimodal approach and semantic gating.

Soutik Mukherjee, Shivakumara Palaiahnakote, Umapada Pal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.