Jul 2026
Multimodal sentiment analysis based on image captions and aspect-guided soft prompts
This work proposes a novel framework that employs the pre-trained vision-language model BLIP (Bootstrapping Language-Image Pre-training) to generate descriptive image captions and introduces an aspect-guided soft prompt mechanism that enables dynamic interaction between aspect terms and multimodal features, thereby mitigating the effects of structural irregularities.
Yuchen Li, Changhong Yu, Binxiao Yu
· Multimedia Systems · 0 citations