A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.
This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant, developing a dual-validity framework in which evidentiary demands scale with scientific ambition.
Zhicheng Lin· Annual Review of Psychology· 11 citations· ⚡1
In education, cognitive models are often representations of knowledge and processes required to solve a task. These models can be used as foundations of pedagogy, curricula, and educational software. Diagnostic classification modeling (DCM) is a family of statistical approaches designed to combine responses to an asses...
Adam Coates· Review of Educational Resear...· 0 citations
Conceptual learning in childhood is challenging, particularly because children’s intuitive theories often conflict with scientifically accurate concepts, and these beliefs rarely revise spontaneously. Prior research has largely focused on structured instructional settings, leaving uncertainty about how learning unfolds...
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use p...
Different capacities for mentalization across LLMs are demonstrated, and cognitive computational modeling is highlighted as a formal method for assessing comparative intelligence across humans and machines.
Aamir Sohail, Xintong Zhong, Arkady Konovalov et al.· 0 citations
Introduction Magic, understood as a set of effects designed to produce experiences of impossibility, offers a distinctive methodological potential for cognitive science because it makes it possible to manipulate, with ecological validity, processes such as attention, perception, reasoning, and memory. Despite increasin...
Miguel Modino de Lucas, Jon Andoni Duñabeitia· Frontiers in Psychology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.