Error-driven representation learning in the mesolimbic system

  1. Program in Neuroscience, Harvard Medical School, Boston, United States
  2. Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Cambridge, United States
  3. Department of Psychiatry and Psychotherapy, University Medical Center, Johannes Gutenberg University Mainz, Mainz, Germany
  4. Department of Psychiatry and Psychotherapy, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany
  5. Department of Psychology and Center for Brain Science, Harvard University, Cambridge, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Rui Ponte Costa
    University of Oxford, Oxford, United Kingdom
  • Senior Editor
    Kate Wassum
    University of California, Los Angeles, Los Angeles, United States of America

Reviewer #1 (Public review):

Summary:

This manuscript reports on simultaneous neural recordings in the olfactory tubercle (OTu) and ventral tegmental area in head-fixed mice performing a simple go-nogo odor-guided reversal task. In the task, there were 3 odors that predicted lick-spout water at 0%, 50%, and 100% probability, with 0% and 100% odors reversing at some point each session. The authors fitted this neural data to a value function approximator in which reward prediction errors fed back onto state representations, allowing the optimal set of representations to be learned. They found that model predictions of adjustments to state representations were correlated with trial-by-trial changes in OTu neural activity, from which the authors conclude that such a system, with dopaminergic errors feeding back onto OTu representations of states, which then generate reward predictions, is biologically plausible.

Strengths:

This is a novel and creative modeling approach that has important implications. It seems to be showing the biological plausibility of a model that can learn state representations rather than relying on a fixed set of states that are programmed into the model. This makes tremendous sense, because the real world is much less well-defined than the kinds of tasks conventionally used by neuroscientists to probe reinforcement learning. As such, it is an important demonstration.

Weaknesses:

The task used in this study is quite simple in its state space, and in particular in how it maps sensory stimuli (odors) onto states, such that the task would seem not to require a system that can learn state representations, or at least that it would not be ideal for testing such a model. This mismatch raises some questions about why this model would perform as well as it appears to be doing here.

The model has two updating functions, both using dopaminergic RPE's. One of these maps raw stimuli to state representations using a parameter termed theta; the second maps state representations to a value prediction, using a parameter termed w. The interaction of these two updating functions seems to be giving the model its interesting characteristics. But it is critical to test what the first of these updates is doing in this model, given that odor stimuli appear to map straightforwardly onto states. For example, one might test the effect of ablating this part of the model, leaving only updates of what the authors term w. Relatedly, one might test the extent to which OTu neurons show simple odor selectivity before and after reversals, asking whether these neurons reflect state representations in this model merely by being selective for a particular odor, or if they develop a more complex kind of responsiveness.
The authors compare the full model, which uses a gradient descent update, to a series of alternatives. The fact that the full model performs significantly better than any of the alternatives leads to the conclusion that the brain is using something like this model in this task. But these alternative models are all reduced or simplified versions of the primary model. This suggests that the full model is the best version within the basic framework posed by the authors. But to draw the conclusion that this model is capturing what is occurring in the brain, one would want to test how this model would perform compared to a different class of model, in particular one that assumes a fixed set of states.
A second weakness of the paper is that the authors do not show the behavioral or raw neural data, which would be important to summarize for the sake of transparency and to help readers get an intuitive sense of what is going on in the task and why the model performs as well as it does. One essential issue is: how much training do mice receive before neural data used in the analysis are collected? Do mice get pre-exposure to contingency reversals before analyzed neural data is collected? Relatedly, how quickly (i.e., in how many trials) do mice show behavioral evidence of having learned initial contingencies and then reversed contingencies? What is the behavioral criterion? With regard to the number of trials mice take to learn the reversals, this can change enormously over training, and such changes could have a big effect on how the model performs. Regarding the neural data, one would want to show some measure of odor selectivity of SPN's and DAN's and how each population responds to delivery and omission of reward in different conditions.

Reviewer #2 (Public review):

Summary:

In this paper, the authors use electrophysiological recordings from the olfactory tubercle (OTu) and ventral tegmental area (VTA) of mice learning an olfactory Pavlovian conditioning task to demonstrate the consistency of changes in OTu odour responses with gradient descent updates minimising reward prediction error (RPE). The paper is clearly written and the work well motivated. The authors address a gap in the literature on animal reinforcement learning by providing neural evidence for gradient-based representation learning, something that had been proposed but not yet tested. The results are convincing and the limitations comprehensively addressed. Of particular interest is the proposal that OTu SPNs could solve the weight transport problem through knowledge of the sign of their downstream connections from their expression of either D1/D2 receptors. This makes a concrete experimental prediction that future research could test.

Strengths:

The paper provides one of the first demonstrations backed by neural recordings that representation learning in the brain is consistent with gradient descent. It shows how, although weight transport may be biologically implausible, the brain appears to find other ways to compute a gradient in multilayered networks for efficient learning. The study builds nicely on recent work in systems neuroscience and provides evidence for a concrete implementation in the OTu-VTA circuitry of mice.

The paper demonstrates that changes in OTu striatal projection neuron (SPN) activity over trials are proportional not only to the RPE relayed by VTA dopamine neurons, but also account for the influence of each particular SPN on the RPE. If an SPN decreases the RPE when active, its activity will increase on the next trial after a positive RPE. The activity will instead decrease for an SPN that increases the RPE. This relation is encapsulated in the update rule of Equation 5.

Weaknesses:

The main weakness of the paper is that it provides only indirect evidence for the update mechanism by inferring synaptic weights based on the (justified) assumption of VTA dopamine neurons encoding RPE. This is still a substantial contribution to understanding representation learning in the brain, though I do think that the authors could provide some additional evidence to further convince readers. One idea could be to analyse the distribution of inferred weights and validate whether it agrees with known statistics of connectivity between OTu SPNs and their downstream projections (e.g. fraction of D1/D2 SPNs).

Further, as dopamine neurons are known to have asymmetrical responses for positive and negative RPEs (with the dips in activity related to negative RPEs being generally smaller) I'd expect an improvement in the correlation of the learning updates particularly after the reversal if the authors account for this in the model.

Reviewer #3 (Public review):

Most models of reinforcement learning in the brain treat the question of how the external world is represented as an afterthought. In "Error driven representation learning in the mesolimbic system", the authors begin by calling attention to this limitation of previous work, then proceed to show that neural activity in part of the ventral striatum evolves in a manner consistent with dopaminergic RPE-driven representation learning. Overall, the question is interesting, the modeling is well done, and the claims are bold. While I am not convinced that olfactory tubercle outputs _mainly_ reflect state features acquired through error-driven learning (see main points below), I am now more willing to believe that they might. I am confident that this work will spark discussion.

Strengths:

The latent weight trajectory inference approach is a nice application of Kalman filtering, the model validation and comparison steps are well executed and thorough, the discussion is nicely written and includes an appropriate caveat about the weight transport problem, and the general idea of striatal output being value-like but also having state-like aspects that evolve over time is thought-provoking.

Weaknesses:

In my view, the main limitation of this work is that the authors do not clearly rule out (1) non-representation learning and (2) non-error-driven learning. This fits with the narrow research question stated at the end of the introduction (L71-72), which is confirmatory in nature and does not claim to exclude alternatives, but other aspects of the framing are less consistent.

Main points of criticism

(1) Representation vs. value learning

The first two pages of the manuscript gave me the impression that the authors wish to draw a clear distinction between representation and value learning, and that they would squarely position this paper as a study of representation learning. The substance of the work does not seem consistent with this positioning, a mismatch that could be addressed either by incorporating new analysis and discussion or by changing the framing.

Examples of emphasis on representation:

- Abstract L14-17 defines value and representation learning and states why they are different.

- First three paragraphs of introduction explain the power of learned representations.

- Results L103-108 attribute state representations to OTu and value to downstream regions.

For this level of emphasis, it would be good to see a convincing argument that OTu MSN activity is better understood as a state representation than as the output of a value function.

Options for establishing OTu as state and not value:

- Explicitly claim in the introduction that previous work has established this, and explain why evidence of stimulus valence being encoded in this area (citations on L68-70) does not favour the value interpretation. Discussing how OTu differs from other parts of ventral striatum that are canonically seen as value coding would also help.

- Directly compare the extent of state vs. value encoding in the present data. The $w=1$ control is a good step in this direction, but its connection to value coding is mentioned only in passing.

(2) Error driven learning vs. other types of learning

Similar to my previous point, the authors seem to claim that the representational changes they study are specifically error driven. While the authors include a good number of controls and ablations, it was not obvious to me that any of them correspond to a form of non-error driven learning that could plausibly generate useful representations. Adding Hebbian learning or a sparsifying learning rule would strengthen this aspect of the work.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation