Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 47 references
TL;DR
This work provides the first comprehensive empirical investigation of humans' internal models play a mediating role in feedback behaviour through a randomized controlled trial and shows that this relationship is invariant across visual contexts and is robust to three common feedback types.
Abstract
Reward learning via human feedback is a crucial capability for beneficial AI. Current methods are built on decision-making theories that assume a matched dynamics model between the learning agent and the feedback provider. However, humans often form imperfect internal dynamics models, and their feedback reflects these misconceptions. While this relationship has long been hypothesised, its manifestation in sequential decision-making remains largely an assumption. Our work provides the first comprehensive empirical investigation of this relationship through a randomized controlled trial (N=211). We followed a two-stage design where we first initialized the participants' understanding of the dynamics in a grid-world navigation domain and then manipulated it using text-based instructions. Causal mediation analysis revealed that humans' internal models play a mediating role in feedback behaviour. We show that this relationship is invariant across visual contexts and is robust to three common feedback types: pairwise preferences, trajectory corrections, and off-switch interventions. These findings confirm a critical limitation of current reward learning methods and establish the missing psychological foundation for approaches that incorporate dynamics understanding.
IMPLIED is an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time, and predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label...
Qiping Zhang, Kate Candon, Debasmita Ghose et al.· 0 citations
Humans are adept at accurately estimating the value of available choices from accumulated experience. However, cognitive processing also incorporates irrelevant information during deliberation, undermining decision accuracy. Here, we show that credit assignment operates automatically, allowing irrelevant action feature...
Ido Ben-Artzi, Maayan Pereg, R. Luria et al.· Nature Communications· 1 citation
Observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes, and these findings show that calibration analyses of repeated human–AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional...
Shreyan Biswas, Alexander Erlei, U. Gadiraju· Proceedings of the 2026 ACM...· 0 citations
Avoidance behaviour is fundamental for survival but can become maladaptive in clinical conditions. A large body of literature has accumulated on the dynamics of human avoidance learning. However, current theories and overviews do not provide an exhaustive account of this evidence. In this systematic review, we identify...
Federico Mancinelli, Dominik R. Bach· Neuroscience and Biobehavior...· 0 citations
The proposed SeekJudge framework, in which four role-specialized agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek--Analyze loop over the trajectory, is the first practical model-based reward to match or surpass native rule-based supervision in online RL.
Yang Wan, Zhenhao Zhang, Jie-Rui Wang et al.· arXiv.org· 0 citations
Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability, which allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload.
Yi-Hsin Ku, Xin Ou, Ji-Qiang Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.