Skip to content

Predicting Program Comprehension with Foundation Models of Human Cognition

Jul 2026 · arXiv.org · Vol abs/2607.11372 · 0 citations · 66 references
Computer Science

TL;DR

Centaur, a foundation model trained on 160 general psychological experiments, is evaluated and it is found that Centaur more closely aligns with human response patterns than its base model, is significantly less reliant on information from prior trials and responses, and benefits more from task-related information.

Abstract

Software engineering depends on the ability of developers to understand code, yet predicting how they do so remains an open challenge despite decades of research. Existing approaches rely either on simplified proxy measures that limit accuracy or on non-trivial measurements requiring elaborate experimental setups that are difficult to scale and apply in practice. In contrast, recent work in psychology suggests an alternative perspective: Instead of modeling task-specific phenomena directly, human behavior can be captured through cognitive regularities learned from large-scale behavioral data. This idea treats complex human behavior as the observable outcome of underlying cognitive processes that manifest consistently across tasks and domains. In this paper, we explore this perspective in the context of program comprehension. We evaluate Centaur, a foundation model trained on 160 general psychological experiments, on 9 previously published program-comprehension studies. We assess how well its predicted response distributions align with human response data and compare Centaur's performance to its base model, Llama 3.1. To better understand the source of its performance, we conduct ablation studies to isolate the contribution of different sources of information, such as the code artifacts, task-related context, and prior trials and participant responses. In a nutshell, we find that Centaur more closely aligns with human response patterns than its base model, is significantly less reliant on information from prior trials and responses, and benefits more from task-related information. These findings suggest that behavioral patterns learned from general psychological data can transfer to complex software engineering tasks such as program comprehension. More broadly, they point toward foundation models of human cognition as a basis for modeling developer behavior in software engineering.

View source

Similar papers

Book Open access Aug 2026

How (and How Not) Do Code Complexity Measures Predict Cognitive Load?

Background and Context. Code complexity measures have been used to guide the design of various activities within computing education, such as instructional sequencing and assessment. However, empirical evidence for the link of these measures to actual cognitive difficulties remains mixed, with studies suffering from sm...

Sverrir Thorgeirsson, Jan Vahrenhold · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
Open access 2026

From Code Generation to Code Auditing: Constrained AI Scaffolding, Transfer, and Cognitive Efficiency in Debugging

Generative artificial intelligence is reshaping programming education, yet its effects on skill development depend partly on how learners interact with artificial intelligence-supported systems. This study introduces the Artificial Intelligence-Scaffolding Interaction Framework, which conceptualizes constrained, questi...

M. Yenidogan, Zeynep Cömert, Duygu Çakır · 0 citations
Preprint Aug 2026

On the use of foundation models in cognitive science

A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. Ho...

Raj Sanjay Shah, Alex Warstadt, Michael C. Frank et al. · 0 citations
Review Open access Jun 2025

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology.

This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant, developing a dual-validity framework in which evidentiary demands scale with scientific ambition.

Zhicheng Lin · 12 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.