Skip to content

R.U.Psycho? A Framework for Robust Unified Psychometric Testing of Language Models

2026 · International Conference on Language Resources and Evaluation · pp. 8365-8386 · 1 citation · 71 references
Computer Science

TL;DR

This paper presents R.U.Psycho, a framework for designing and running robust and reproducible psychometric experiments on generative language models that reduces the required coding expertise and demonstrates the capability of the framework on a variety of psychometric questionnaires.

View source

Similar papers

Review Open access Jun 2025

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology.

This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant, developing a dual-validity framework in which evidentiary demands scale with scientific ambition.

Zhicheng Lin · 12 citations · ⚡1
Review Open access Jul 2026

Large Language Models as Distributional Baselines for Language Tasks

The central contributions of this paper articulate the conditions under which distributional predictability threatens the internal validity of an experiment and provide concrete recommendations for how to control for this potential confound.

Sean Trott, James A. Michaelov, Cameron R. Jones et al. · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Large-scale factor analysis shows machine intelligence is only partially interpretable

A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...

Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi · 0 citations
Preprint Aug 2026

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use p...

Alona Strugatski, Licol Zeinfeld, Giora Alexandron · 1 citation
#artificial intelligence Preprint Sep 2026

Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated adminis...

Yu Sha, Jun-Qi Tao, Di-Xin Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.