Towards Unified Affective AI: A Cross-Modal Survey of Emotion Recognition, Generation, and Control
TL;DR
This survey provides a comprehensive review of state-of-the-art methodologies for emotion recognition and generation across facial, speech, and textual modalities, covering preprocessing techniques, datasets, deep learning architectures, evaluation metrics, and emotion control mechanisms.
Abstract
Emotion recognition and generation have become increasingly important research areas in artificial intelligence, underpinning applications in healthcare, education, customer service, and human–computer interaction. Human emotion is inherently multimodal, being conveyed through facial expressions, speech, and language, while modern affective AI systems increasingly integrate these complementary modalities to improve robustness and expressiveness. Furthermore, emotion recognition and generation are becoming increasingly interconnected, with recognition models frequently supporting the training, control, and evaluation of generative systems, yet existing surveys typically examine these tasks independently or focus on individual modalities. This survey provides a comprehensive review of state-of-the-art methodologies for emotion recognition and generation across facial, speech, and textual modalities, covering preprocessing techniques, datasets, deep learning architectures, evaluation metrics, and emotion control mechanisms. Recent advances are categorised according to their underlying methodological approaches, enabling a systematic comparison of current techniques, their strengths, and their limitations. The quantitative comparisons show that Transformer-based Poster++ achieved 92% accuracy on RAF-DB, while XLM-EMO obtained an F1 score and accuracy of 0.85 on Affect in Tweets. For speech emotion recognition, the highest reported accuracy on RAVDESS was 92.88%. Among facial expression generation methods, EAT achieved an emotion-recognition accuracy of 75.43% on MEAD. For speech emotion generation, Seq2seq-VC obtained the lowest word error rate and equal error rate on VCTK, at 2.9% and 1.0%, respectively, while the diffusion-based DDDM-VC achieved the lowest character error rate of 1.0% on VCTK. These results indicate strong performance by Transformer-based facial and textual emotion-recognition methods; however, no single architectural family demonstrates consistent superiority across the reported generation metrics, datasets, and evaluation settings. Finally, the survey discusses current challenges and future research directions, providing a unified perspective on the development of multimodal affective AI systems.