Skip to content
Open access

Multimodal Perception Based Recognition of Emotional Visual Narratives

Sep 2026 · Applied and Computational Engineering · 0 citations

Abstract

The recognition of emotional visual stories is an issue that involves simultaneous comprehension of semantic transformation between successive frames, the distribution of visual attention by the viewers, and observable emotional response. Information carried by visuals is not enough to completely characterize the perception process. For that reason, a trimodal computation model is designed with visual semantics, visual attentional behavior, and facial response. The visual stimuli consist of 120 four frame emotional stories. Sixty subjects watch the stories and identify the main emotion. Tobii Pro Fusion tracks eye movements with 120 Hz sampling frequency, whereas 30 fps video recording tool captures the facial videos simultaneously. CLIP ViT B 32 and Transformer compute the semantics of consecutive frames. BiLSTM computes the temporal representation of eye movements and facial action units individually. Next, cross-modal attention mechanism fuses the three types of representation together. With subject independent five fold cross validation, the trimodal model achieves an accuracy score of 0.824 ± 0.013 and macro F1 score of 0.817 ± 0.014, compared with scores 0.724 ± 0.017 and 0.706 ± 0.019 of visual only model. Fear and Surprise categories still have cross classification rate of 9% and 8%, respectively.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.