The Kalman Filter requires seeding with values for reward stochasticity/volatility.

It is seeded with the initial values, for which it is statistically optimal. When the values change due to unmodelled, unrepresentable largeworld effects, it performs poorly. A Bayes-optimal learner can be fragile to large-world effects that break its foundational assumptions. Hetlearn is suboptimal for any particular, fixed reward stochasticity/volatility, but robust to changes.

A: Convergence of ALS algorithm’s stochasticity and volatility estimates over time, compared to their (dotted) true values. One run plotted for illustration. B: Faster convergence of the weightings of Hetlearn sublearners over time. One run plotted for illustration C: 100 samples of the final ALS estimates of stochasticity and volatility, after 1000 timesteps. Estimation errors are anticorrelated. D: Mean MSE of Kalman learner vs Hetlearn over 100 repeats. Bands represent standard error. E: Mean trajectory (100 repeats) of the difference in cumulative MSE for the Kalman learner and Hetlearn. Green (positive) values indicate that the Kalman learner had higher cumulative MSE than Hetlearn, while purple (negative) values indicate the opposite. F: Learning rate of the two algorithms as compared to the optimal formula of Eq. (3) over a single trial.

A: Volatility and stochasticity for four different environments not encompassed by a single generative model. B: Hetlearn (top) and Piray et al (bottom) tracking the the optimal learning rate across all four environments. C: Cumulative error over 10 trials for Hetlearn and Piray et al. D: Cumulative mean squared error (MSE) over 10 trials for Hetlearn and Piray et al. For B, C, and D, and the average of 10 runs were plotted for Hetlearn and Piray et al.

A: Schematic of an individual MBON compartment. The effect of sensory representations, transmitted through KC axons, on MBON activity, is adapted through plasticity induced by co-occurring DAN activation. Adapted from Figure 1A of [34]. B: Topology of the 16 compartments comprising the MB. KC axons arising from the three KC cell types (α/β, α′/β, γ) are depicted as blue lines, segregated according to anatomical layer. Example MBON and DAN included. Adapted from Figure 1C of [3] C: Schematic of proposed qualitative mechanism. Separate MBON compartments receive highly overlapping sensory representations from KCs. However, their valences are different, and are updated by anatomically and functionally distinct DAN teaching signals. Overall valence is a weighted sum of the individual compartments’ opinions.

A: Hetlearn model for behavioural conditioning in Drosophila. Different engrams strengths compete to update from a prediction error, reducing interference with previously learned information. Predictions themselves do not invoke this competition, allowing the system to make rich, context-dependent predictions that make use of the representional capacity of multiple engrams. B: Proposed model of mechanism underlying second order conditioning. Compare to C: Mechanisms underlying second-order conditioning as proposed in [58]

Model predictions vs behavioural experiments.

A: Extinction protocol adapted from [18]: S1 is associated with punishment (reward of −1) 30 times, followed by 30 punishment omissions. Stimulus preference of the model to S1 over time is graphed, along with the time at which the first (rewarding) and second (omitting) contexts are automatically formed in the model. B: Reversal protocol taken from [56]. Each block depicts ten pairings of S1 or S2 with punishment (reward of −1) or omission. Bar plot depicts difference between valences for S1 and S2 at the end of each block. C: Experimental protocol used in [58], and our model, for second-order conditioning. We modelled DAN activation as a reward on the timestep after the stimulus was presented. In context 2, S2 and S1 were presented on successive timesteps. D: Drosophila preference to S1 and S2 over time. Experimental data (Figure 1H of [58]) on leftmost pane, depicting preference index. Our model graphs valence on other panes, depicting (left) standard conditioning, (centre) conditioning under SMP108 knockout, and (right) conditioning when S1 is presented before S2 in context 2.