A conserved latent reward-learning strategy is identified in mice and humans characterized by the greatest feedback sensitivity, reward efficiency, and adaptation after reversal, defining a translational framework for studying how adaptive decision-making is shaped by task experience, stress, affective processes, and neural circuit function.
Abstract
Adaptive behavior requires that organisms learn which actions are rewarded and to update action selection when the environment changes. While human and non-human animals exhibit adaptive behavior, whether apparently similar behavior reflects common decision strategies remains unclear. Probabilistic reversal learning provides a cross-species assay of reward-guided choice, yet standard metrics such as accuracy or reward rate can obscure underlying strategies that generate choices. Here, we applied parallel probabilistic reversal learning tasks in mice and humans and used a generalized linear model–hidden Markov model to infer latent decision strategies from trial-by-trial behavior. Across species, choices were organized into stable behavioral states with differing reliance on choice history, reward history, and response bias. Among these latent states, we identify a conserved reward-learning strategy in mice and humans characterized by the greatest feedback sensitivity, reward efficiency, and adaptation after reversal. Simulating choice behavior using state-specific decision policies reproduced the empirical hierarchy of performance, confirming that the latent states capture meaningful behavioral strategies. Although mice and humans differ in the temporal dynamics of reward learning, both species ultimately converge on the same optimized strategy. These findings identify a conserved latent reward-learning strategy in mice and humans, defining a translational framework for studying how adaptive decision-making is shaped by task experience, stress, affective processes, and neural circuit function. Teaser A shared latent behavioral strategy underlies flexible decision-making in mice and humans.
Mice trained to play a zero-sum game against an opponent that exploited statistical regularities in their choices and rewards, and tracked their behaviour and dorsal cortical dynamics across learning found that mice transitioned from structured, predictable strategies towards a near-optimal stochastic strategy as they...
Joanna Aloor, Timothy P. H. Sit, Oliver M. Gauld et al.· bioRxiv· 0 citations
Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically...
Effective decision making in stochastic environments requires balancing flexible, value-based learning with a stabilising influence of habitual action selection. While dopamine-mediated reward prediction errors (RPEs) are a well-established component of value learning, the mechanisms underlying habit- like behaviour re...
Yu-Hao Wang, L. Burgeno, J. Cerpa et al.· bioRxiv· 0 citations
Using predictive models to explain cognition requires more than accurate behavioral predictions. Input ablations offer an appealing route: remove information from a model and interpret the resulting performance change as evidence of its importance for behavior. Yet this inference assumes that the model's dependence on...
Behavioral flexibility—the ability to adapt behavior in response to changing conditions—is commonly measured using reversal learning tasks, where individuals must inhibit a previously rewarded response and adopt a new one after contingencies change. Despite widespread use, task comparability across species remains un...
N. Alessandroni, Rachael Miller, D. Altschul et al.· Open Mind· 1 citation· ⚡1
Real-world decision-making rarely occurs with perfect information. Instead, individuals must constantly weigh potential rewards against the probability of adverse outcomes.1 Failures of this process can lead to maladaptive decisions associated with reduced lifetime success, and numerous psychiatric disorders such as ga...
T. Price, A. Liu, Rhiannon L. Cowan et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.