Aug 2026· International journal of pattern recognition and artificial intelligence· 0 citations
TL;DR
This work aims to use scene text images for sentiment analysis to assist in understanding the intentions of captured scenes and presents TSRB (Transformer-based Semantic Refinement Block), which comprises a multimodal approach and semantic gating.
Abstract
Sentiment analysis is essential for several real-world applications, such as opinion mining and predicting a person's intent and personality. Most existing work aims to address challenges of sentiment analysis using normal text and images uploaded on social media. This work aims to use scene text images for sentiment analysis to assist in understanding the intentions of captured scenes. We present TSRB (Transformer-based Semantic Refinement Block), which comprises a multimodal approach and semantic gating. The proposed method constructs hierarchically fused image and text representations and then routes them through a TSRB and a learned three-way Semantic Gating module. The image branch encodes both the full meme image and text image extracted from the input image through a convolutional network with spatial attention; the text branch encodes OCR text, raw tweet text, and image captions via three independent Distil-BERT+CNN encoders and hierarchically fuses them. The resulting visual and textual embeddings are jointly refined by three stacked Transformer encoder layers within the proposed TSRB and then selectively blended by a softmax-weighted Semantic Gate that dynamically arbitrates among the post-attention, visual, and textual streams. Experiments are conducted on two standard datasets (MVSA-Single and Memotion) and compared with state-of-the-art models to demonstrate the effectiveness of the proposed method.
As the amount of text and image data on social media continues to increase, multimodal sentiment analysis has emerged as a critical area of study. However, a "modality gap" that lowers classification accuracy is frequently caused by the domain disparities between textual and visual modalities. In this paper, a Translat...
Experimental results on the MVSA-Single and MVSA-Multiple datasets show that the proposed method improves performance in image–text multimodal sentiment classification, thereby validating the effectiveness of combining semantic enhancement with difference-aware modeling.
Hengyuan Zhang, Aizihaierjiang Yusufu, Jiang Liu et al.· Electronics· 0 citations
A hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN) is presented, a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals that outperforms single-modal sentiment analysis.
M. Yuvaraja, Dr. C. Kumuthini· Journal of Intelligent Decis...· 0 citations
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhi Liang et al.· Signal, Image and Video Proc...· 0 citations
Robustly Optimized Bidirectional Encoder Representations from Transformers Approach (RoBERTa) model, fine-tuned for domain adaptation with library domain data, and introducing the Aspect-Based Sentiment Analysis (ABSA) framework to extract and classify fine-grained sentiment related to collection services in reader fee...
C. Song· Advanced Electromagnetics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.