Mice trained to play a zero-sum game against an opponent that exploited statistical regularities in their choices and rewards, and tracked their behaviour and dorsal cortical dynamics across learning found that mice transitioned from structured, predictable strategies towards a near-optimal stochastic strategy as they learned.
A conserved latent reward-learning strategy is identified in mice and humans characterized by the greatest feedback sensitivity, reward efficiency, and adaptation after reversal, defining a translational framework for studying how adaptive decision-making is shaped by task experience, stress, affective processes, and n...
Zahra Rostami, Hadi Choubdar Parvin, E. Iyer et al.· bioRxiv· 0 citations
Effective decision making in stochastic environments requires balancing flexible, value-based learning with a stabilising influence of habitual action selection. While dopamine-mediated reward prediction errors (RPEs) are a well-established component of value learning, the mechanisms underlying habit- like behaviour re...
Yu-Hao Wang, L. Burgeno, J. Cerpa et al.· bioRxiv· 0 citations
This work comprehensively investigates the concept of constant-memory strategies in stochastic games, and uncovers the connection between decision models in single-agent and multi-agent contexts.
Feng-Ming Zhu, Fang-Zhen Lin· Proceedings of the Thirty-Fi...· 0 citations
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\tim...
Behaviour of human participants in a prey-pursuit task reflects dynamic blending of goal-specific control policies, with hippocampus estimating latent states, anterior cingulate cortex orchestrating policy switches and orbitofrontal cortex supplying value-based contextualization rather than driving continuous policy up...
Assia Chericoni, Justin M. Fine, Taha S. Ismail et al.· Nature· 2 citations
Experimental findings indicate that reinforcement learning agents achieve superior consistency and long-term optimization in structured settings, while human players demonstrate greater flexibility and adaptability under uncertain or novel conditions.
Ziad Karim Tabbouche· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.