bioRxiv Science⌕ Search

Biology subjects

Louka, M.

Publications and source records attributed to Louka, M..

2 recordsLinked to original sources

VTA dopamine neuron activity produces spatially organized value representations

What is the neural architecture by which dopamine (DA) determines choice? Reinforcement learning (RL) has suggested an algorithmic chain: prediction errors based on reward train predicted values for stimuli and actions, and thereby determine choice. Elements of this chain have been identified, but have yet to be convincingly assembled. Here we used optogenetic stimulation of putative prediction error neurons at the top of this chain - i.e., DA neurons in the ventral tegmental area (VTA) - and recorded the activity of thousands of neurons in their target structure - the striatum - in mice performing a task that dissociates VTA DA signalling, stimulus value, action value, and choice. Conventional RL models captured some features of the data, including DA-driven choice reinforcement and value correlates, but failed on two key observations: i) action and state value correlates were relatively weak in the VTAs main target, the nucleus accumbens (NAc), and ii) VTA DA could not reinforce actions in the absence of a reward-predictive cue (CS+). Instead, we find that VTA DA produces stimulus value representations in the NAc, which in turn generates action value representations downstream (e.g. the CP). We formalize these observations in a Conditioned Reinforcer model, where VTA DA specifically acts on stimuli rather than actions, and these stimulus value representations are used to drive choice reinforcement.

neuroscience↗

Pre-existing visual responses in a projection-defined dopamine population explain individual learning trajectories

Learning a new task is challenging because the world is high dimensional, with only a subset of features being reward-relevant. What neural mechanisms contribute to initial task acquisition, and why do some individuals learn a new task much more quickly than others? To address these questions, we recorded longitudinally from dopamine (DA) axon terminals in mice learning a visual task. Across striatum, DA responses tracked idiosyncratic and side-specific learning trajectories. However, even before any rewards were delivered, contralateral-side-specific visual responses were present in DA terminals only in the dorsomedial striatum (DMS). These pre-existing responses predicted the extent of learning for contralateral stimuli. Moreover, activation of these terminals improved contralateral performance. Thus, the initial conditions of a projection-specific and feature-specific DA signal help explain individual learning trajectories. More broadly, this work implies that functional heterogeneity across DA projections serves to bias target regions towards learning about different subsets of task features, providing a mechanism to address the dimensionality of the initial task learning problem.

neuroscience↗