Neural circuit model for learning values of odor-cued states.

(A) Stimulus information is transmitted to the olfactory tubercle (OTu) of the ventral striatum via the olfactory bulb and/or piriform cortex. Striatal projection neurons (SPNs) in OTu represent states s(x) = θ⊤x, and a learned weighted sum of SPN activity estimates the value vϕ(x) = w ⋅ s(x). The reward prediction error (RPE) δ = r − vϕ(x) is projected back from the ventral tegmental area (VTA) to update both feature representation parameters θ and the linear approximation parameters w.(B) Linear-Gaussian dynamical system model of weight updates. Shaded variables are observed while unshaded variables are latent.

Validation of weight inference on synthetic data.

(A) Correlation between the true and inferred weight trajectory. (B) Correlation between true and reconstructed representation updates from sampling weight trajectories and reconstructing RPEs and representation trajectories. C Correlation between true and inferred representation updates using weight trajectory mean and synthetic RPEs. Bars show mean and standard error across neurons.

Experimental setup.

(A) Simultaneous recording in the olfactory tubercle (OTu) and ventral tegmental area (VTA). Licking and neural data were recorded in a head-fixed configuration during stimulus and reward presentation. (B) Behavioral paradigm. Mice were trained on a go/no-go task with three odor stimuli that predicted reward with 0%, 50%, and 100% probability. ITI: intertrial interval.

Correlations between predicted and observed changes in olfactory tubercle activity during (A) Pavlovian conditioning sessions (411 neurons) and (B) reversal sessions (284 neurons).

Bars show mean and standard error across neurons. The gradient model (Eq. 5) was compared to several alternative representation updates: fixing δ = 1, fixing w = 1, replacing w by its element-wise sign (sign(w)), replacing w by a fixed Gaussian vector (feedback alignment), and randomly sampling the sign of the update. The t-values and p-values shown above each alternative model are derived from paired t-tests.

Change in SPN response between successive representations of an odor as a function of the decile of predicted change according to (A) gradient descent, (B) δ = 1, (C) w = 1, (D) sign(w), (E) feedback alignment, (F) random sign update rules.

Points denote mean and error bars denote standard error. A total of n = 99902 updates were considered across all SPNs and trials and sessions of both types.

Comparison between alternative weight dynamics models on the ability to explain the SPN and DAN response data using the Bayesian information criterion.

Each point is subtracted from the minimum BIC value for that session and offset by 1 for visualization purposes on a logarithmic scale. Box denotes Q1, Q2, Q3, and whiskers denote Q1 − 1.5IQR, Q3 + 1.5IQR where Qi denotes ith quartile.

Hyperparameter bounds and initial values for marginal likelihood optimization procedure.

Full histogram of correlations between successive model-inferred and recorded changes in SPN response.

Top row: Pavlovian conditioning sessions (411 neurons). Bottom row: Reversal sessions (284 neurons). Our alternative representation updates are: (A, F) Gradient descent rule, (B, G) fixing δ = 1, (C, H) fixing w = 1, (D, I) replacing w by its element-wise sign, (E, J) replacing w by a fixed Gaussian vector. Results of paired t-tests against a random-sign shuffle null of each model’s updates are shown above each plot.

Normalization window robustness of update correlation results across both Pavlovian conditioning and reversal sessions.

(A) 2 s pre-CS window, 0.5 s offset from stimulus and response window for SPNs and DANs respectively. (B) 1 s pre-CS window, 0.5 s offset from stimulus window for both SPNs and DANs. Bars show mean and standard error across neurons.

Robustness to minimal SPN and DAN count thresholds.

Results are robust to requiring at least (A) 16 SPNs and 8 DANs and (B) 8 SPNs and 4 DANs per session. Bars show mean and standard error across neurons.

Correlations between predicted and observed changes in olfactory tubercle activity pooled across (A) Pavlovian conditioning sessions and (B) reversal sessions when predictions are evaluated on data held out with respect to weight inference hyperparameter fitting.

Bars show mean and standard error across neurons. The gradient model (Eq. 5) was compared to several alternative representation updates: fixing δ = 1, fixing w = 1, replacing w by its element-wise sign (sign(w)), replacing w by a fixed Gaussian vector (feedback alignment), and randomly sampling the sign of the update (null). The t-values and p-values shown above each alternative model are derived from paired t-tests. Bars show mean and standard error across neurons.

Predicted SPN changes computed on trials that were held out from hyperparameter fitting in the weight inference procedure.

Change in SPN response between successive representations of an odor as a function of the decile of predicted change according to (A) gradient descent, (B) δ = 1, (C) w = 1, (D) sign(w), (E) feedback alignment, (F) random sign update rules. Points denote mean and error bars denote standard error. A total of n = 20443 updates were considered across all SPNs and trials.

Correlation between each SPN update and the number of trials between successive measurements of SPN response to a particular odor stimulus.

Pavlovian conditioning (A) and reversal (B) sessions are shown. We compute a 1-sample t-test against the alternative hypothesis that the mean is greater than 0.

Weight inference synthetic recovery (A-C) and update predictions (D) with inferred constant weight learning rate β.

(A) Correlation between the true and inferred weight trajectory. (B) Correlation between true and reconstructed representation updates from sampling weight trajectories and reconstructing RPEs and representation trajectories. C Correlation between true and inferred representation updates using weight trajectory mean and synthetic RPEs. (D) Predicted update correlation on the data. Results are aggregated across both Pavlovian conditioning and reversal sessions. (C-D) Bars show mean and standard error across neurons.