Skip to content
Preprint

Psychological Competence as a Missing Dimension in AI Evaluation

Jul 2026 · 0 citations · 43 references
Computer Science

TL;DR

It is argued that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

Abstract

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

View source

Similar papers

Review Open access Jul 2026

Evaluating AI “Understanding” with Cognitive-Psychological Criteria: Evidence, Gaps, and Structural Limits

With the rapid development of large-scale language models, the performance of artificial intelligence in language understanding and reasoning tasks has become increasingly strong. As a result, many people have gradually regarded the success of tasks as the same as true understanding. However, from the perspective of cognitive psychology, understanding is not merely about giving correct or fluent responses; it is actually an internal psychological process. It includes aspects such as meaning construction, context model update, reasoning generation, and metacognitive regulation. Under such circumstances, this paper systematically studies the cognitive understanding ability of artificial intelligence from the perspective of cognitive psychology. It employs the methods of literature review and qualitative case analysis. First, it clarifies the concept definitions and evaluation criteria of understanding in cognitive psychology by combining classic theories and experimental evidence, and then uses these criteria as the analytical framework. Study the performance of “similar understanding” in contemporary artificial intelligence systems. It was found that although artificial intelligence can approach human understanding at the behavioral level, it has not reached the core cognitive psychological standards. AI lacks experience-based situational models, causal and goal-oriented reasoning mechanisms, and inherent metacognitive monitoring. These limitations are structural, not quantitative issues. It reflects the fundamental difference between artificial intelligence systems and human cognition. This article can help clarify the theoretical distinction between surface manifestations and true understanding, and also provide a cognitive psychology foundation for a more cautious explanation of artificial intelligence capabilities.

Ruo Qin · 0 citations
Review Open access Aug 2026

Artificial Intelligence in eWOM: Roles, Mechanisms, and an Integrative Framework

Artificial intelligence (AI) is changing electronic word-of-mouth by participating directly in the production, revision, circulation, evaluation, and governance of consumer information. This systematic literature review synthesizes 71 peer-reviewed studies on how AI functions within electronic word-of-mouth (eWOM), how consumers respond to those roles, and when those effects vary. Using SPAR-4-SLR and PRISMA 2020 procedures, the review draws on database, supplemental, and citation-based searches. It identifies five recurring AI roles -- Creator, Co-Creator, Curator, Intermediary, and Regulator -- and shows that their effects depend largely on perceptions of authenticity, credibility, usefulness, humanness, transparency, and manipulation risk. These effects are context dependent, varying across products, tasks, message formats, consumer orientations toward AI, levels of AI involvement, and platform environments. The review also argues that AI reshapes not only how eWOM is presented, but also what consumers are likely to contribute afterward. Building on these findings, the review proposes an AI-enabled lifecycle of Creation, Transformation, Transmission, Evaluation, and Governance and develops an AI-eWOM fit perspective. A TCCM-organized research agenda identifies priorities for future research.

A. Joyal · 0 citations
Review Open access Aug 2026

Cognitive Readiness for Human-AI Collaboration.

ObjectiveThis narrative review examines the cognitive, metacognitive, and team competency requirements that may contribute to productive and reliable collaboration between human and AI to address two questions: What capabilities make AI a competent collaborator? What makes humans ready for AI collaboration?BackgroundAs AI systems are increasingly integrated into workplaces and framed as teammates rather than tools, humans face challenges that include maintaining situation awareness, calibrating trust, and working with systems that may surpass them cognitively. We analyzed Human-Agent Teaming (HAT) readiness around two complementary levels: operational team competencies (communication, coordination, and adaptability) and regulatory capacities (trust calibration and metacognitive awareness).MethodWe conducted a structured narrative review of literature from 2010 through January 2026, searching Google Scholar, Scopus, PsycINFO, IEEE Xplore, ACM Digital Library, and Semantic Scholar, complemented by forward citation tracking. After screening 572 records, 192 articles were included for synthesis.ResultsCommunication inflexibility, limited shared understanding, and trust miscalibration emerge as recurring barriers to HAT, while regulatory capacities (trust calibration and metacognitive awareness) represent particularly critical dimensions of HAT readiness that remain to be fully operationalized.ConclusionHAT requires mutual readiness, with both humans and AI developing metacognitive and adaptive capabilities. Despite methodological heterogeneity limiting clear conclusions, cross-training and co-learning methods offer a promising avenue for building shared understanding and calibrated collaboration.ApplicationThis review provides practical principles for designing AI systems that support calibrated collaboration and for preparing humans to work adaptively with AI, thereby enhancing team effectiveness, reliability, and resilience in collaborative work environments.

Sébastien Tremblay, Delphine De Hemptinne, Gabrielle Teyssier-Roberge et al. · 0 citations
Review Jul 2026

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment"essential"when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.

Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins et al. · 0 citations
Book Open access Jul 2026

Intelligence Is Not All You Need: A Case for Artificial Wisdom in Conversational Agents

This paper argues that intelligence is the wrong primary design goal for conversational systems. As agents become more capable, proactive, and socially embedded, their most consequential failures stem less from weak task performance than from poor judgment: misframing problems, mishandling uncertainty, overlooking stakeholder interests, and steering users in ways that undermine autonomy and well-being. I propose artificial wisdom as a needed corrective for conversational systems. By artificial wisdom, I mean context-sensitive, morally grounded, meta-cognitively regulated judgment oriented toward human flourishing under uncertainty. This perspective shifts attention from what systems can do to how they should act when advising, persuading, and shaping decisions. I argue that intelligence alone cannot determine appropriate goals, guide action under uncertainty, or ensure beneficial human outcomes. On this basis, I sketch a wisdom-oriented agenda for conversational AI centered on role awareness, meta-cognition, deliberation, self-explanation, and calibrated proactivity. Based in this, I examine its promise, difficulty, and risks.

Matthias Kraus · 0 citations
Preprint Jul 2026

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment

Sudhir Alladi Venkatesh · 0 citations