Jun 2025· Annual Review of Psychology· 11 citations· ⚡ 1 influential· 91 references
MedicineComputer Science
TL;DR
This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant, developing a dual-validity framework in which evidentiary demands scale with scientific ambition.
Abstract
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms-statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use p...
The replication crisis in psychology is now widely recognized as a crisis of measurement itself, not merely of questionable research practices. Recent critical work has revealed that many psychological scales may be measuring semantic associations rather than the psychological phenomena they purport to assess. This Hyp...
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpret...
Alona Strugatski, Licol Zeinfeld, Jason Cooper et al.· 0 citations
The results reveal that, although LLMs occasionally succeed in decoding communicative intentions, their performance is not attributable to human-like ToM reasoning, and offers insight into their interpretive biases, contributing to a deeper understanding of their linguistic capabilities.
A. Lombardi, Alessandro Lenci· Transactions of the Associat...· 0 citations
Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response be...
Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building...
Oliver A. Guidetti, Reza Ryan· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.