Skip to content
Review Open access

Towards Unified Affective AI: A Cross-Modal Survey of Emotion Recognition, Generation, and Control

Sep 2026 · Applied Sciences · Vol 16, pp. 9333 · 0 citations · 146 references

TL;DR

This survey provides a comprehensive review of state-of-the-art methodologies for emotion recognition and generation across facial, speech, and textual modalities, covering preprocessing techniques, datasets, deep learning architectures, evaluation metrics, and emotion control mechanisms.

Abstract

Emotion recognition and generation have become increasingly important research areas in artificial intelligence, underpinning applications in healthcare, education, customer service, and human–computer interaction. Human emotion is inherently multimodal, being conveyed through facial expressions, speech, and language, while modern affective AI systems increasingly integrate these complementary modalities to improve robustness and expressiveness. Furthermore, emotion recognition and generation are becoming increasingly interconnected, with recognition models frequently supporting the training, control, and evaluation of generative systems, yet existing surveys typically examine these tasks independently or focus on individual modalities. This survey provides a comprehensive review of state-of-the-art methodologies for emotion recognition and generation across facial, speech, and textual modalities, covering preprocessing techniques, datasets, deep learning architectures, evaluation metrics, and emotion control mechanisms. Recent advances are categorised according to their underlying methodological approaches, enabling a systematic comparison of current techniques, their strengths, and their limitations. The quantitative comparisons show that Transformer-based Poster++ achieved 92% accuracy on RAF-DB, while XLM-EMO obtained an F1 score and accuracy of 0.85 on Affect in Tweets. For speech emotion recognition, the highest reported accuracy on RAVDESS was 92.88%. Among facial expression generation methods, EAT achieved an emotion-recognition accuracy of 75.43% on MEAD. For speech emotion generation, Seq2seq-VC obtained the lowest word error rate and equal error rate on VCTK, at 2.9% and 1.0%, respectively, while the diffusion-based DDDM-VC achieved the lowest character error rate of 1.0% on VCTK. These results indicate strong performance by Transformer-based facial and textual emotion-recognition methods; however, no single architectural family demonstrates consistent superiority across the reported generation metrics, datasets, and evaluation settings. Finally, the survey discusses current challenges and future research directions, providing a unified perspective on the development of multimodal affective AI systems.

Read PDF

Similar papers

Review Open access Aug 2026

A Systematic Review of Emotion Recognition: From Unimodal Signals to Multimodal Integration

These findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets.

B. Bashir, Zayyanu Yunusa · 0 citations
Conference Aug 2026

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combin...

Gnanaseelan Dharshika, A. Ramanan · 0 citations
Book Open access Oct 2026

Recognition of Affective States Using Multimodal and Adaptive AI for Individuals with Motor Disabilities

People with motor disabilities often display atypical emotional expressions that elude conventional recognition models, leading to communication difficulties, isolation, and restricted access to emotional support. Multimodal AI affective recognition systems offer promising avenues to overcome these barriers, but existi...

Maram Djebbi, Kirmene Marzouki, Fernando Marmolejo-Ramos et al. · 0 citations
Sep 2026

Mamba-CrossMod: a multimodal affective analysis framework based on selective state space model

The proposed Mamba-CrossMod is a novel multimodal feature fusion framework that introduces Mamba-ATT, an enhanced attention mechanism based on a selective state-space model for capturing long-range dependencies with theoretically linear complexity.

Yi-Wen Tong, Jing Mu, Wen-Xin Chang et al. · 0 citations
Open access Sep 2026

Multimodal fusion for emotion classification based on GSR, BVP, and SKT

Emotion recognition in next-generation healthcare is revolutionizing patient care by fostering empathetic and personalized interactions. Advanced systems that analyse facial expressions, speech, and physiological signals enhance our understanding of emotional states, ultimately improving patient experiences. However, i...

Adhitya Velip, H. Virani, Amita Dessai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.