A biologically plausible, multi-region neural circuit model of songbird vocal learning is presented and it is shown that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating.
Abstract
Reinforcement learning is a key means by which animals learn appropriate actions in a given context. A large body of work suggests that such learning depends on interactions between cortico-basal ganglia circuits and the midbrain dopaminergic system, yet the underlying circuit mechanisms and plasticity rules are not fully understood. Here we present a biologically plausible, multi-region neural circuit model of songbird vocal learning and map it onto the actor-critic framework of reinforcement learning. In this model, stochastic spiking activity in the cortico-basal ganglia pathway implements action selection and drives behavioral exploration, while the pathways driving midbrain dopaminergic signaling evaluate behavioral outcomes and support a reward prediction error based learning rule that approximates stochastic gradient ascent. The model achieves millisecond-scale precise learning that matches observed behavior. We further use the model to examine two fundamental constraints on biological reinforcement learning. First, dopaminergic reinforcement signals are temporally imprecise, which can cause interference between neurons controlling actions that occur close in time. Second, dopaminergic signals are spatially imprecise, which can cause interference between neurons controlling different aspects of behavior but receiving a common reinforcement signal. By jointly modeling the actor and critic components of the circuit, we show that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating. These results suggest a circuit-level mechanism by which biological systems achieve reinforcement learning despite the temporal and spatial limitations of global neuromodulatory signals.
Learning depends on both reward- and error-based feedback, yet how the brain integrates these distinct signals to guide behaviour remains fundamentally unclear. Here, using a normative computational framework, we derive credit assignment rules for both reward-based learning (RBL) and error-based learning (EBL). In cont...
Michele Garibbo, C. Filipe, L. Aitchison et al.· bioRxiv· 0 citations
These results demonstrate that spiking navigation circuits can learn goal-directed behavior using local plasticity, but robust multi-goal learning benefits from context-specific evidence-based consolidation.
Samuel A Neymotin, Hananel Hazan, Gozde Unal et al.· Research Square· 0 citations
In vivo whole-cell membrane potential recordings, simultaneous monitoring and bidirectional manipulation of dopamine signaling in awake, behaving mice to examine how dopamine shapes corticostriatal circuits and identify learning related plasticity as its principal mechanism for shaping striatal circuits in vivo.
Mélanie Druart, Yun C. Yang, N. Tritsch et al.· bioRxiv· 0 citations
Connectomics is used to map the cell types and synaptic connections underlying a form of multi-layer continual learning that cancels predictable sensory responses in a cerebellum-like structure in electric fish, highlighting the potential of connectomics, in combination with cell-type-specific physiological recordings...
Krista E. Perks, Mariela D. Petkova, Salomon Z. Muller et al.· Nature· 0 citations
It is suggested that medial prefrontal cortical (mPFC) implements a default strategy based on prior knowledge, which actively hinders the expression of more efficient strategies, and a decentralized multiexpert competition model best predicts behavior and causal perturbations.
Kai Lu, K. Wong, Cheng Yang et al.· Science Advances· 1 citation
Learning requires adaptive changes in neuronal circuits, but how neurons encode learning content in their activity patterns to construct memories remains poorly understood. Using longitudinal multi-site recordings in freely moving male mice performing a sensory discrimination task, we discover the emergence of burst-co...
Filippo Heimburg, Nadin Mari Saluti, Lars-Lennart Oettl et al.· Nature Communications· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.