Small circuit principles for learning in large worlds

  1. Department of Informatics, University of Sussex, Brighton, United Kingdom
  2. Department of Engineering, University of Cambridge, Cambridge, United Kingdom

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Gordon Berman
    Emory University, Atlanta, United States of America
  • Senior Editor
    Albert Cardona
    University of Cambridge, Cambridge, United Kingdom

Reviewer #1 (Public review):

The authors describe a learning algorithm that is meant to deal with the problem of inference in large, complex environments. They argue that simpler but ensemble learning systems with heterogeneous learning rates often perform better than more complex approaches. They also relate the ideas of their model to a circuit implementation in the mushroom body, an associative learning area in insects.

I found the argument about the value of sublearners remaining simple to be thoughtful and intriguing. I also appreciated the connection to biology and, although there remain details to be worked out, it seems that the proposed approach has what I'd call a "biologically plausible" instantiation.

Comments on the major claims:

The paper claims that algorithms that learn stochasticity and volatility "necessarily result in worse short-term performance than Hetlearn." The results with Autocovariance Least Squares and the model of Piray and Daw are consistent with this statement, but the authors seem to be making a more generic claim. Is there a rigorous justification of this? In particular, I found the section "Competing sublearners should stay simple" to straddle a line between opinion and fact, and I wasn't sure where it intended to be regarding that line.

Relatedly, regarding the statement just before that section: "The exponential convergence of Hetlearn weightings stands in contrast to the slow convergence of stochasticity/volatility estimates (Figures 2A and C), which result in poor initial performance of the Kalman learner." Is the claim here that the Kalman learner doesn't converge exponentially, or that it converges more slowly?

The paper would benefit from a comparison to other models of multiple learning systems with different timescales. The classic one would be the complementary learning systems ideas about hippocampus and neocortex (McClelland, McNaughten, O'Reilly 1995), and the many subsequent follow-ups. Two other relevant ones to consider are Roxin & Fusi 2013 and Lindsey & Litwin-Kumar 2024, which also consider multiple memory systems with different plasticity rates.

Comments on clarity:

(1) Page 2: There could be some more unpacking of the relationship between x and z in equations 1a, 1b. If z_t are time-varying, shouldn't theta, if it represents what is learned about "large-world factors," also depend on time?

The description of second-order conditioning on page 6 is confusing. Second-order conditioning tests whether S2 acquires a positive predicted value after pairing with S1, which had previously been paired with reward. The text makes it seem like the goal of second-order conditioning is to learn S2->S1->no reward.

(2) Page 6: "Representations decay on a slower timescale than stimulus presentation. Hence, the representation x2 after experiencing S2 → S1 holds traces of the S2 representation." I couldn't understand this. Maybe an equation would help?

(3) Page 7: The phrase "unconditioned stimulus" is introduced without being described. Is this a reward or punishment? What does it have to do with valence lower-case v, or value upper-case V? The authors seem to mix notations from different subfields, leading to a rather confusing description of the model.

Comments on relationship to biology:

(1) Page 6: It is assumed that Kenyon cells interact recurrently through plastic weights. This is an unconventional assumption for mushroom body models, and it is made without much discussion. The authors cite papers that argue for changes in the calyx during learning (although these effects are generally much weaker than the well-documented plasticity onto MBONs). However, KC-to-KC interactions are largely in the lobes, not the calyx. The authors should provide more discussion of this assumption. What happens if the recurrence is weak or absent? A relevant reference to consider is Manoim et al. Current Biology 2022, which shows that KC-to-KC interactions are inhibitory.

The model seems to imply that each MBON compartment encodes only one memory, or did I misunderstand? Is this necessary? Typically, it is thought that because of the sparse, decorrelated representations in the KCs, an individual MBON could store multiple associations without too much interference.

Reviewer #2 (Public review):

Summary:

The authors propose Hetlearn, a novel learning algorithm for unstable environments: instead of using statistically optimal methods to estimate the variability of the environment, simply have a range of "sub-learners" that learn at different rates, and rely more heavily on the sub-learners with a recent track record of success. This strategy is particularly helpful when the environment has unknown and fluctuating stochasticity (randomness around a constant mean reward) and volatility (the mean itself changes over time). For high stochasticity and low volatility, it's better to learn slowly to smooth out randomness; for low stochasticity and high volatility, it's better to learn quickly to catch onto the new trend. The authors show that existing methods work well when stochasticity and volatility are constant, or when methods can learn them over a long time, but that Hetlearn does better when stochasticity and volatility change over time or an agent has limited time or data to learn from.

They then propose that the fly mushroom body implements a form of Hetlearn, since it has different compartments that form parallel memory traces at different learning and forgetting rates. Here, the conclusions are less clear because their model doesn't appear to implement the key feature of Hetlearn (the differential weighting of sublearners based on recent success), and because their model has features that don't have an obvious biological equivalent in the real mushroom body.

Strengths:

The idea behind Hetlearn is clever and interesting, and it makes sense that a heuristic method like Hetlearn would outperform a theoretically optimal algorithm under conditions of limited data. The model is clearly defined, and they test it across a range of tasks and compare it to a range of different Kalman/particle filter learners. The application to the mushroom body is also an inspiring connection and takes advantage of the discovery of empirically described sublearners. The work will stimulate new ways of understanding how biological systems might implement heuristic tricks that may not be optimal under ideal conditions, but that succeed in "real world" conditions, and it will stimulate new thinking about the computational role of parallel memory traces in the mushroom body.

Weaknesses:

In the tests of Hetlearn vs. Kalman filters, it would be useful to more fully explore the parameter space. Specifically, in Figures 1-2, it would help define the boundaries of when Hetlearn is better or worse than Kalman filters, by systematically changing the parameters of the task (e.g. increasing/decreasing the amount of data available to the learner or the frequency of changing stochasticity/volatility), and changing the parameters of Hetlearn (the learning rates of the sublearners and the recency bias, i.e. how far back to track the recent success of the sublearners when weighting their outputs). For Figure 3, while they do check the performance when varying the recency bias, there are no conditions presented in which the particle filter does better. To define when Hetlearn is better, it is important to show when it is worse.

In the mushroom body model, there are several elements that have no apparent biological equivalent, and it's not clear whether they are incidental (in which case they could be removed to enhance biological realism) or essential to the model's performance (in which case the discrepancy should be discussed; perhaps the model can be reframed as a prediction for as-yet-undiscovered features). In addition, the model seems to lack the defining feature of Hetlearn (up/down-weighting sublearners according to their recent performance). Finally, the authors argue that Hetlearn explains experimental data better than existing models, but existing models could be easily and parsimoniously tweaked to accommodate these data, so the explanatory advantage of Hetlearn is unclear.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation