Light-ED: Lightweight Multimodal Emotion Detection using Enhanced EfficientNet
Abstract
Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency. In this paper, we propose Lightweight Multimodal Emotion Detection (Light-ED), a unified and resource-efficient framework that integrates multimodal fusion and emotion classification within a single architecture. We introduce a lean cross-modal cross-attention (CMCA) mechanism to enable effective interaction between audio and visual modalities with low computational cost. Building on this, we design a unified model that augments an EfficientNet backbone with CMCA, streamlining feature extraction, fusion, and classification. We evaluate the proposed approach on two benchmark datasets, RAVDESS and CREMA-D. Our model achieves state-of-the-art performance with accuracies of 97.92% and 84.42%, respectively, while maintaining strong efficiency in terms of model size, latency, and GPU usage. The results demonstrate that jointly addressing fusion and classification leads to a better balance between accuracy and computational cost, making the approach suitable for real-world multimodal emotion recognition systems.