Aug 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 1181-1190· 0 citations
Abstract
The present study proposes such a multimodal sentiment analysis framework with attention-enhanced properties, a combination of ResNet50 and Convolutional Block Attention Module (CBAM), a textual encoder with BERT, and refinement of relational features via Graph Neural Networks (GNN). The model is designed to address the vulnerability of uni-modal sentiment analysis and integrate related visual and textual evidence. CBAM enhances the visual feature representation with the assistance of channel and spatial attention, but BERT proposes text embeddings in their context. Another model similar to multimodal representations and similarity-based sampling relationships is a Graph Neural Network. It is experimentally demonstrated that the proposed framework is characterized by a total classification accuracy of 0.7676 compared to baseline and conventional attention-based models. The additional outcomes of Precision show that enhanced retrieval performance was obtained, which highlights the fact that multimodal fusion that is strengthened by mental attention can be effective in the sentiment classification task.
Consistency-Aware Gated Fusion (CAGF), a lightweight and fusion module tailored to Mamba-based architectures that achieves state-of-the-art performance, outperforming strong multimodal baselines such as CLIP, MISA, DLF, AoM, and SFTTR, while remaining more efficient and interpretable.
Jian Hu· Poster Volume 0008 The 2026...· 0 citations
This study focuses on the problem of insufficient accuracy in sentiment classification in complex texts and multimodal scenarios. A sentiment classification optimization method is proposed, which integrates a bidirectional encoder representation Transformer model for sentiment analysis (Senti-BERT)and a bidirectional L...
Jie Zhang, Jian-Bang Liu, Zhao-Sheng Xu· PLoS ONE· 0 citations
This paper presents a context-aware hybrid deep learning approach by integrating the Robustly Optimized BERT Pretraining Approach (RoBERTa) with Bidirectional Long Short-Term Memory (BiLSTM) networks to generate rich contextual word embeddings.
V. Gayatri, Rajani Rajalingam· International Journal for Re...· 0 citations
A hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN) is presented, a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals that outperforms single-modal sentiment analysis.
M. Yuvaraja, Dr. C. Kumuthini· Journal of Intelligent Decis...· 0 citations
XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.
M. Kidwai, C. Author, Dr. Faiyaz Ahmad· Journal of Intelligent Decis...· 0 citations
: Nowadays, with the integration of multiple information carriers into daily communication, emotional transmission shows a trend of cross-modal fusion. Multimodal sentiment analysis has become a cutting-edge direction in Artificial Intelligence (AI). According to the study, there are three basic categories into which t...
Jiajia Li· Proceedings of the 3rd Inter...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.