Skip to content
Open access

Evaluating the robustness of specialized and general-purpose facial expression recognition systems across varied scenarios

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 55 references
Medicine

Abstract

Introduction This work presents a comprehensive evaluation of facial expression recognition (FER) systems across four benchmarked datasets of varying complexity, ranging from controlled static images (ADFES, WSEFEP) to more realistic dynamic recordings (RAVDESS, CREMA-D). Methods Three categories of models were evaluated: traditional FER neural networks models, general-purpose vision language models (VLMs), and the commercial software FaceReader© 10. Results The results show that performance on controlled datasets substantially overestimates real-world FER capability, with average weighted and unweighted average recall values decreasing from approximately 72% in static datasets to below 30% in naturalistic settings. All tested models exhibited a marked bias toward happiness, with negative emotions frequently misclassified, a trend particularly pronounced in VLMs, where categories such as fear or anger often received F1-scores near zero. Among the neural networks, the DAN model trained on the AfectNet dataset achieved the strongest generalization, outperforming all VLMs and confirming that AfectNet provides a more realistic training distribution than the RAF-DB database. FaceReader© delivered excellent performance under ideal conditions but experienced substantial degradation in dynamic scenarios, falling below a random classifier in CREMA-D. Discussion These findings highlight the limitations of general-purpose VLMs and commercial tools for real-world FER and underscore the need for models explicitly designed to handle naturalistic variability. Furthermore, the reported performance of FaceReader© 10 in their manual on ADFES and WSEFEP was corroborated in this study.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.