Nov 2026· Journal of Engineering, Project, and Production Management· 0 citations
TL;DR
Key contributions include proposing a multi-task learning framework for jointly optimizing visual quality and emotion, establishing the inaugural VAWE-Art dataset comprising 5,000 AI-generated images with 20-dimensional emotional annotations, and providing computational foundations for emotion-controllable generative art systems.
Abstract
The rapid advancement of generative artificial intelligence has enabled significant breakthroughs in the visual realism and artistic expressiveness of AI-generated artworks. However, challenges persist in objectively evaluating their visual quality and emotional impact. Existing evaluation methods either focus on underlying image quality or rely on subjective user surveys, lacking a unified computational framework. This paper proposes a bimodal, multi-dimensional, and interpretable computational evaluation method. The Visual Quality and Emotion (VQ-Emo) framework is designed to jointly optimize visual quality scores and multidimensional emotional predictions. The framework comprises three core modules: a visual quality assessment module (a multi-branch convolutional neural network incorporating style perception, quantifying composition, color, texture, and lighting); an emotional impact computation module (a 20-dimensional multi-label emotion classifier based on the VAWE emotion model), and a visual-emotion association module (using attention mechanisms to identify emotion-driven visual regions). This model was trained and validated on the AGIQA-1K, ArtEmis, and self-constructed VAWE-Art datasets. Experimental results demonstrate that VQ-Emo outperforms existing methods in both visual quality assessment and emotion recognition, achieving an SRCC of 0.912 and a mAP of 0.678. Key contributions include proposing a multi-task learning framework for jointly optimizing visual quality and emotion, establishing the inaugural VAWE-Art dataset comprising 5,000 AI-generated images with 20-dimensional emotional annotations. This reveals quantifiable correlations between visual features and emotional responses across diverse artistic styles and provides computational foundations for emotion-controllable generative art systems.
The rapid development of AI image generation technology has created an urgent need for systematic aesthetic evaluation of generated visual content. Existing computer vision assessment methods are often limited to technical image parameters and insufficiently consider multidimensional artistic judgment, texture details,...
Lingbo Yang, N. Yang· Advanced Electromagnetics· 0 citations
A real-time visual art emotion recognition framework that combines a MHAI feature extraction module with transfer learning based on a pre-trained Inception-V3 network is proposed to enhance the extraction of emotionally salient features while improving robustness to stylistic diversity in artworks.
ProFocus is a novel framework that models affective experience in artistic images via progressive visual focusing inspired by a hierarchical cognitive theory of human aesthetic appreciation and consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation.
Zhiyan Zhang, Zi-Qing Yan, Jianqi Chen et al.· 0 citations
Modeling aesthetic perception from dynamic visual information requires simultaneous understanding of human motion and the physical behavior of textile materials, making multimodal feature fusion a critical challenge in intelligent perception systems. Similar information integration strategies have also attracted increa...
By enabling machines to comprehend, interpret, and react to human affective states, visual mapping of multimodal emotion traits is crucial to intelligent interface design. Nevertheless, current multimodal emotion detection algorithms primarily focus on predictive performance, providing little insight into interpretabil...
Wan-Bao Ge, Zhen-Hua Yang, Zhen-Hu Liu et al.· Journal of Visualized Experi...· 0 citations
This survey provides a comprehensive review of state-of-the-art methodologies for emotion recognition and generation across facial, speech, and textual modalities, covering preprocessing techniques, datasets, deep learning architectures, evaluation metrics, and emotion control mechanisms.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.