Skip to content
Open access

An Interpretable Vision-Language Framework for Evaluating the Uncanny Valley Effect of XR Humanoid Characters

Jul 2026 · Electronics · Vol 15, pp. 2959 · 0 citations

TL;DR

UVE-PCoT improves affinity prediction and cue-level explanation over general-purpose multimodal large language models and ablations and operationalizes this perceptual-conflict-oriented perspective into an interpretable framework, advancing UVE evaluation from black-box scoring to explanatory analysis and providing cue-level insights for XR character assessment and revision.

Abstract

As AI-generated humanoid characters are increasingly used in virtual, augmented, and mixed reality applications, evaluating the Uncanny Valley Effect (UVE) is crucial for immersive user experience. Existing evaluation methods map visual features to affective scores, offering limited interpretability regarding which visual cues are associated with affinity judgments. Among the theoretical perspectives proposed to explain the UVE, perceptual conflict provides a visual-cue-oriented perspective for analyzing whether local-feature realism supports a coherent overall human-likeness impression and how this is reflected in affinity judgments, yet this perspective is rarely incorporated into interpretable UVE assessment. Thus, we propose UVE-Perception Chain-of-Thought (UVE-PCoT), a vision-language framework for interpretable UVE evaluation from a perceptual-conflict-oriented perspective. UVE-PCoT organizes assessment through a structured perceptual decomposition, including assessments of overall human-likeness, local-feature realism, perceptual conflict, and affinity. To provide supervision, we construct UVE-R, a structured rationale dataset with image-grounded, rating-consistent rationales linking visual cue observations, cue-level inconsistency analysis, and affinity judgments. Results show that UVE-PCoT improves affinity prediction and cue-level explanation over general-purpose multimodal large language models and ablations. Our approach operationalizes this perceptual-conflict-oriented perspective into an interpretable framework, advancing UVE evaluation from black-box scoring to explanatory analysis and providing cue-level insights for XR character assessment and revision.

Read PDF

Similar papers

Preprint Aug 2026

Human-AI Perceptual Alignment by Playing Hues and Cues

Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained semantic and cultural nuances. In this work, we propose a novel evaluation framework that leverages the gamified, discrete color space of the bo...

Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver et al. · 0 citations
Preprint Aug 2026

ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing

ProFocus is a novel framework that models affective experience in artistic images via progressive visual focusing inspired by a hierarchical cognitive theory of human aesthetic appreciation and consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation.

Zhiyan Zhang, Zi-Qing Yan, Jianqi Chen et al. · 0 citations
Preprint Aug 2026

A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dime...

Ana Naveriani, Jakob Suchan, S. Zoia et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

Though absolute performance remains low, finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance.

Akhila Yerukola, Fabrice Y. Harel-Canada, Simran Khanuja et al. · 0 citations
Review Sep 2026

Organization of Valence and Arousal in Vision-Language Representations of Built Environments: Insights from the EMOIS Dataset

Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of...

Madoka Yonekura, Katsunori Kohda, Nobuhiko Muramoto et al. · 0 citations
Preprint Aug 2026

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

MMArt is introduced, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations.

Shuai Wang, Wang-Yuan Ding, Yixian Shen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.