2026· International Conference on Language Resources and Evaluation· pp. 8365-8386· 1 citation· 71 references
Computer Science
TL;DR
This paper presents R.U.Psycho, a framework for designing and running robust and reproducible psychometric experiments on generative language models that reduces the required coding expertise and demonstrates the capability of the framework on a variety of psychometric questionnaires.
This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant, developing a dual-validity framework in which evidentiary demands scale with scientific ambition.
Zhicheng Lin· Annual Review of Psychology· 12 citations· ⚡1
The central contributions of this paper articulate the conditions under which distributional predictability threatens the internal validity of an experiment and provide concrete recommendations for how to control for this potential confound.
Sean Trott, James A. Michaelov, Cameron R. Jones et al.· Open Mind· 0 citations
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...
Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al.· 0 citations
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use p...
Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated adminis...
Yu Sha, Jun-Qi Tao, Di-Xin Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.