Skip to content
Open access

Biologically Plausible Dopamine-Modulated STDP Model of Pavlovian Learning in Spiking Neural Networks

Jul 2026 · bioRxiv · 0 citations · 20 references
Biology

TL;DR

A DA-modulated STDP rule is proposed in which increasing DA progressively biases plasticity toward potentiation while receptor saturation limits further DA effects beyond a critical concentration, providing a biologically grounded model of DA-dependent plasticity and offering new insight into how abnormal dopamine signaling can impair learning in neurological disorders.

Abstract

Spike-timing-dependent plasticity (STDP) and dopamine (DA) are fundamental to reward-based learning and memory formation. A widely used DA-modulated STDP model explains how neural networks associate stimuli with delayed dopaminergic rewards through an eligibility trace. However, we show that this model supports learning even at unrealistically high DA concentrations because DA simply scales the magnitude of STDP without changing its temporal profile. In contrast, experiments demonstrate that DA nonlinearly reshapes the STDP window, converting long-term depression (LTD) into long-term potentiation (LTP) at high DA levels. We therefore propose a DA-modulated STDP rule in which increasing DA progressively biases plasticity toward potentiation while receptor saturation limits further DA effects beyond a critical concentration. Simulations of recurrent networks of Izhikevich neurons show that the proposed rule supports robust conditioning only within a biologically realistic DA range (0.04–0.70 µM). Successful learning produces a hybrid network architecture consisting of a strong feedforward backbone embedded within recurrent circuitry and generates enhanced burst responses selectively to reward-associated stimuli. At the upper limit of the biologically plausible DA range, the network passes through a narrow bistable regime, converging to one of two distinct stable configurations. At higher DA concentrations, conditioning fails altogether. These results provide a biologically grounded model of DA-dependent plasticity and offer new insight into how abnormal dopamine signaling can impair learning in neurological disorders.

Read PDF

Similar papers

Open access Aug 2026

Context-Aware Evidence-Gated Plasticity for Multi-Goal Learning in Spiking Neural Networks

Background / Introduction: Biologically inspired spiking neural networks can model adaptive behavior, but learning multiple goals is difficult because synaptic updates for different targets can interfere. We tested whether multi-timescale plasticity and context-specific credit assignment could improve continual multi-goal learning in a spiking navigation system inspired by entorhinal-hippocampal circuitry. Methods: We developed a closed-loop spiking model containing grid-like, place-like, target-related, association, and motor-output populations. An agent navigated in a two-dimensional environment with randomized starting locations and learned through reward-modulated spike-timing dependent plasticity (STDP/RL) and a novel evidence-gated plasticity (EGP) framework. EGP accumulates candidate synaptic modifications, evaluates them using reward evidence, and consolidates only changes that improve performance. A target-context variant maintained separate proposal stores and reward evaluation for each target. Results: STDP/RL learned and retained a single-target navigation policy, but multi-target training produced substantial interference, including attraction to incorrect targets after learning. Across 10 connectivity seeds, target-context EGP achieved higher late-stage reward than global EGP, improved weakest-target performance, and increased the fraction of targets achieving positive reward. In a longer continual-learning simulation, reward increased for all targets, TEST-phase performance increasingly exceeded TRAIN-phase performance, and proposal magnitudes grew over learning. Dwell-time confusion analyses showed that target-context EGP reduced wrong-target attraction and improved target selectivity relative to multi-target STDP/RL. Conclusions: These results demonstrate that spiking navigation circuits can learn goal-directed behavior using local plasticity, but robust multi-goal learning benefits from context-specific evidence-based consolidation. Target-context EGP provides a biologically motivated mechanism for reducing interference during continual reinforcement learning in spiking neural networks.

Samuel A Neymotin, Hananel Hazan, Gozde Unal et al. · 0 citations
Open access Aug 2026

Learning in spiking neural networks with a calcium-based Hebbian rule for spike timing-dependent plasticity

This work presents a Hebbian local learning rule that models synaptic modification as a function of calcium traces tracking neuronal activity and demonstrates how spike timing and rate can be complementary in their role of shaping the connectivity of spiking neural networks.

Willian Soares Girāo, Nicoletta Risi, Caroline Geisler et al. · 0 citations
Open access Jul 2026

Constraints for spatially and temporally precise learning in a neural circuit model of reinforcement learning

Reinforcement learning is a key means by which animals learn appropriate actions in a given context. A large body of work suggests that such learning depends on interactions between cortico-basal ganglia circuits and the midbrain dopaminergic system, yet the underlying circuit mechanisms and plasticity rules are not fully understood. Here we present a biologically plausible, multi-region neural circuit model of songbird vocal learning and map it onto the actor-critic framework of reinforcement learning. In this model, stochastic spiking activity in the cortico-basal ganglia pathway implements action selection and drives behavioral exploration, while the pathways driving midbrain dopaminergic signaling evaluate behavioral outcomes and support a reward prediction error based learning rule that approximates stochastic gradient ascent. The model achieves millisecond-scale precise learning that matches observed behavior. We further use the model to examine two fundamental constraints on biological reinforcement learning. First, dopaminergic reinforcement signals are temporally imprecise, which can cause interference between neurons controlling actions that occur close in time. Second, dopaminergic signals are spatially imprecise, which can cause interference between neurons controlling different aspects of behavior but receiving a common reinforcement signal. By jointly modeling the actor and critic components of the circuit, we show that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating. These results suggest a circuit-level mechanism by which biological systems achieve reinforcement learning despite the temporal and spatial limitations of global neuromodulatory signals.

Yiheng Wang, J. Kornfeld, M. Fee et al. · 1 citation
Open access Aug 2026

Irregular and bursty neuronal firing extend spike-timing dependent plasticity window to behavioral timescales

Summary Synaptic plasticity, which is thought to underlie learning and memory, is commonly induced in experimental settings with regular activity patterns. However, such regularity strongly differs from natural in vivo firing statistics. Therefore, it remains unclear how in vivo-like patterns, such as irregular and bursty activity, shape synaptic plasticity. We combined mathematical modeling and ex vivo patch-clamp experiments inducing naturalistic spike-timing-dependent plasticity (STDP) at cortico-striatal synapses. We found that irregular spike-pair stimulation diminished LTD occurrence and widened the LTP temporal window compared to regular patterns at low firing rate. Furthermore, increasing the firing rate abolished LTD to the profit of LTP. Our modeling and experimental data indicate that bursts of action potentials are key contributors to this effect. These results highlight the importance of naturalistic firing statistics and show that irregular and bursty firings extend the typical compressed STDP expression temporal window toward timescales relevant for behavior.

Yulia Dembitskaya, Silvana Valtcheva, Yihui Cui et al. · 0 citations
Open access Aug 2026

The Cellular and Synaptic Actions of Dopamine During Behavior

Dopamine signaling in the striatum is essential for a wide range of functions, from reward learning to motor vigor and behavioral flexibility. Although dopamine signals fluctuate on sub-second timescales, how these signals are translated into lasting changes in striatal circuit function remains unknown. Resolving this requires cell-type-specific measurements of synaptic and intrinsic properties during behavior, a longstanding technical challenge. Here, we combined in vivo whole-cell membrane potential recordings, simultaneous monitoring and bidirectional manipulation of dopamine signaling in awake, behaving mice to examine how dopamine shapes corticostriatal circuits. Acute manipulations of dopamine over seconds to minutes produced only modest effects on corticostriatal synaptic transmission and no detectable changes in membrane potential dynamics or intrinsic excitability. By contrast, associative learning robustly strengthened identified corticostriatal synapses onto both D1- and D2-expressing spiny projection neurons, yet only D1-SPN plasticity required dopamine signaling. These findings challenge models in which dopamine acts rapidly to tune striatal excitability and identify learning related plasticity as its principal mechanism for shaping striatal circuits in vivo.

Mélanie Druart, Yun C. Yang, N. Tritsch et al. · 0 citations
Book Open access Jul 2026

EvoPlaSNN: Evolving Reward-modulated ANN-based Plasticity Rule for Spiking Neural Networks

Neuroevolution can indirectly train a neural network by optimising the underlying plasticity rules that adapt the network weights. In Spiking Neural Networks (SNNs), evolvable meta-learning studies have been applied to unsupervised learning settings. In our experiment, we let evolution search for rules that can train an SNN to solve a Reinforcement Learning maze. These rules are encoded as an Artificial Neural Network (ANN), taking as inputs synaptic variables such as weight, reward and eligibility traces. We found that using synaptic weight as inputs has no effect on rule performance, while utilising either pre-before-post or post-before-pre eligibility trace is better than a trace that combines both. Although there is still too much variation in fitness both within individuals and across generations to allow for conclusive results, there is potential for evolving plasticity rules to solve RL tasks in the future.

Napat Sahapat, S. Chevtchenko, Y. Bethi et al. · 0 citations