Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in Persian, a low-resource language, across multiple evaluation strategies and prompt formulations. We find that LLM-human...
Mohammad Reza Modarres, Armin Tourajmehr, Yadollah Yaghoobzadeh et al.· 0 citations
In everyday reasoning and scientific inquiry, explanations help people make sense of their experiences and observations. We investigate how people generate new possible explanations. Specifically, we study how the phrasing of "why" questions shapes exploration of the space of possible explanations. We hypothesized that...
Kara Kedrick, Russell Golman· Cognition· 0 citations
An understanding gap is illuminated: user attributions are partly guided by epistemic orientation and experiential ascriptions that make sophisticated simulation appear as understanding, especially in affective interaction, raising urgent questions about epistemic trust, relational vulnerability, and the ethics of AI c...
Erez Firt, Rinat B. Rosenberg-Kima· AI & SOCIETY· 0 citations
Judicial judgments are increasingly available, yet dense language and distributed relationships among facts, evidence, reasoning, and rulings remain difficult for non-experts to interpret. Through a mixed-methods formative study with Chinese non-expert readers (survey N=34; interviews N=6), we identified structural, in...
Xin-Yi Chen, Rui-Ji Li, Yue-Lu Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.