Skip to content
Open access

Constraints for spatially and temporally precise learning in a neural circuit model of reinforcement learning

Jul 2026 · bioRxiv · 1 citation · 94 references
Biology

TL;DR

A biologically plausible, multi-region neural circuit model of songbird vocal learning is presented and it is shown that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating.

Abstract

Reinforcement learning is a key means by which animals learn appropriate actions in a given context. A large body of work suggests that such learning depends on interactions between cortico-basal ganglia circuits and the midbrain dopaminergic system, yet the underlying circuit mechanisms and plasticity rules are not fully understood. Here we present a biologically plausible, multi-region neural circuit model of songbird vocal learning and map it onto the actor-critic framework of reinforcement learning. In this model, stochastic spiking activity in the cortico-basal ganglia pathway implements action selection and drives behavioral exploration, while the pathways driving midbrain dopaminergic signaling evaluate behavioral outcomes and support a reward prediction error based learning rule that approximates stochastic gradient ascent. The model achieves millisecond-scale precise learning that matches observed behavior. We further use the model to examine two fundamental constraints on biological reinforcement learning. First, dopaminergic reinforcement signals are temporally imprecise, which can cause interference between neurons controlling actions that occur close in time. Second, dopaminergic signals are spatially imprecise, which can cause interference between neurons controlling different aspects of behavior but receiving a common reinforcement signal. By jointly modeling the actor and critic components of the circuit, we show that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating. These results suggest a circuit-level mechanism by which biological systems achieve reinforcement learning despite the temporal and spatial limitations of global neuromodulatory signals.

Read PDF

Similar papers

Open access Aug 2026

Unifying error and reward action learning: a cerebello-basal ganglia theory

Learning depends on both reward- and error-based feedback, yet how the brain integrates these distinct signals to guide behaviour remains fundamentally unclear. Here, using a normative computational framework, we derive credit assignment rules for both reward-based learning (RBL) and error-based learning (EBL). In cont...

Michele Garibbo, C. Filipe, L. Aitchison et al. · 0 citations
Open access Aug 2026

Context-Aware Evidence-Gated Plasticity for Multi-Goal Learning in Spiking Neural Networks

These results demonstrate that spiking navigation circuits can learn goal-directed behavior using local plasticity, but robust multi-goal learning benefits from context-specific evidence-based consolidation.

Samuel A Neymotin, Hananel Hazan, Gozde Unal et al. · 0 citations
Open access Aug 2026

The Cellular and Synaptic Actions of Dopamine During Behavior

In vivo whole-cell membrane potential recordings, simultaneous monitoring and bidirectional manipulation of dopamine signaling in awake, behaving mice to examine how dopamine shapes corticostriatal circuits and identify learning related plasticity as its principal mechanism for shaping striatal circuits in vivo.

Mélanie Druart, Yun C. Yang, N. Tritsch et al. · 0 citations
Open access Sep 2026

Connectome analysis of a cerebellum-like circuit for sensory prediction.

Connectomics is used to map the cell types and synaptic connections underlying a form of multi-layer continual learning that cancels predictable sensory responses in a cerebellum-like structure in electric fish, highlighting the potential of connectomics, in combination with cell-type-specific physiological recordings...

Krista E. Perks, Mariela D. Petkova, Salomon Z. Muller et al. · 0 citations
Open access Aug 2026

Neural competition between prefrontal and auditory cortex constrains novel sound strategy learning

It is suggested that medial prefrontal cortical (mPFC) implements a default strategy based on prior knowledge, which actively hinders the expression of more efficient strategies, and a decentralized multiexpert competition model best predicts behavior and causal perturbations.

Kai Lu, K. Wong, Cheng Yang et al. · 1 citation
Open access Aug 2026

Thalamocortical bursts encode reward contingencies and drive associative learning

Learning requires adaptive changes in neuronal circuits, but how neurons encode learning content in their activity patterns to construct memories remains poorly understood. Using longitudinal multi-site recordings in freely moving male mice performing a sensory discrimination task, we discover the emergence of burst-co...

Filippo Heimburg, Nadin Mari Saluti, Lars-Lennart Oettl et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.