Introduction

Everyday decisions often involve uncertainty. From selecting a route home, to trying a new restaurant, to accepting a job offer, individuals must act without explicit knowledge of each choice’s outcomes or their probabilities. Major depressive disorder has been associated with alterations in how uncertainty is evaluated, how rewarding and aversive associations are learned, and how these experiences guide future actions1–4. Such alterations may contribute to the emergence and maintenance of symptoms, particularly motivational symptoms such as anhedonia5. However, it remains unclear whether these differences reflect vulnerability to depression, arise only during an active episode, or persist following remission–a question with direct implications for when and how learning and decision might be effectively targeted.

When outcomes and their probabilities are not explicit, individuals must learn through experience. Reinforcement learning (RL) models provide a formal framework for characterising this process3,6, allowing the identification of multiple latent processes: sensitivity to rewards and punishments (how outcomes are subjectively valued), learning rates (how rapidly expectations are updated by outcomes) and choice stochasticity (randomness in choice).

Meta-analytic evidence indicates that depression (and anxiety) disrupts learning from rewards and punishments, while leaving sensitivity to those outcomes largely intact3. That is, the hedonic experience of receiving a reward or punishment remain broadly preserved–the chocolate bar tastes just as good7–9, failing the exam feels similarly bad-but depressed individuals may differ in how much they update future behaviour based on these outcomes1,3,10. Specifically, current depressive episodes appear to reduce approach learning following rewards while amplifying updating following negative outcomes. In volatile environments, where choice-outcome contingencies change over time, such patterns may promote excessive avoidance leading individuals to prematurely abandon options that are only temporarily worse11–13. This reframes anhedonia and apathy (core symptoms of depression) not as blunted hedonic experiences per se, but as disruptions in learning and motivation9,10,14,15.

Beyond reinforcement learning, decisions under explicit risk, where probabilities are stated and learning is not required, offer a complementary window into reward and punishment processing. Prospect theory provides a theoretical model of such decisions and explains that individuals typically prefer certainty over probabilistic options, due to risk aversion (diminishing sensitivity to outcome value as value increases) and loss aversion (the tendency to weigh potential losses more than potential gains)16. Generalised anxiety disorder (GAD) has been associated with increased risk but not loss aversion in such contexts17,18; work in depression remains limited, and we aimed to extend this here.

For these findings to inform treatment, prevention or relapse interventions, it is necessary to establish whether RL and decision-making alterations reflect causes, correlates or consequences of depression5,19,20. A cross-sectional design comparing currently depressed, remitted, and at-risk individuals to healthy controls allows competing accounts to be tested: state-dependent alterations should be present only during active illness; trait-like vulnerability markers should be shared across depressed, remitted and at-risk groups; and compensatory mechanisms should be unique to remitted individuals.

To date, however, studies including “at-risk” groups (e.g. first-degree relatives), have produced mixed findings: some report elevated risk aversion (following negative outcomes) or blunted reward/punishment learning consistent with trait vulnerability, but others do not21–26. Evidence for compensatory mechanisms is similarly inconsistent: some evidence suggests partial normalisation of reward learning and avoidance (lose-shift behaviour) following cognitive behavioural therapy27, but it is not clear whether this underlies successful maintenance of remission. Methodological heterogeneity likely contribute: participants are often medicated, confounding illness and treatment; remitted groups often report residual symptoms; and variability in task design and modelling limits comparability across studies.

Here, we compared four unmedicated groups: individuals with current depression, individuals remitted from depression, never-depressed individuals at familial risk, and healthy controls. Remitted and at-risk participants had minimal symptoms at the time of testing. We assessed reinforcement learning using a volatile four-armed bandit task and decision making under explicit risk using a task that dissociates risk from loss aversion. This allowed robust estimation of separate reward and punishment learning rates, sensitivities, and choice noise via hierarchical modelling and Bayesian Model Averaging, to enhance robustness. This design enables differentiation between vulnerability markers, state/trait effects, and features associated with maintained recovery. We additionally examined associations between task parameters and symptom dimensions–including apathy and anhedonia–derived from confirmatory factor analysis, to investigate how learning and decision-making alterations map onto specific features of depression.

Methods and Materials

Participants

Participants from all groups were recruited via multiple channels: university mailing lists, flyers posted around university campus and student therapeutic services, and online recruitment platforms (e.g. MQ Mental Health Research).

We conducted two studies: Study 1 was a pilot study with healthy volunteers, used to refine the study protocol and select and validate our computational models. In Study 1 (pilot), 102 healthy participants were recruited; data from n=99 were analysed after exclusions. Exclusion criteria included current or past psychiatric diagnoses (self-reported), marijuana use within four weeks, other recreational drug use within one week, or alcohol consumption within 24 hours.

For Study 2, 195 participants were recruited: 60 healthy controls, 36 individuals with a first-degree relative with depression, 50 individuals in remission from major depressive disorder, and 49 individuals with current major depressive disorder (MDD). With N=195 (harmonic mean n=47.14 per group), alpha=0.05, and 80% power, the minimum detectable effect size is f=0.24.

Diagnoses were established using the Mini International Neuropsychiatric Interview (MINI-v5). Current MDD required meeting MINI criteria and a Hamilton Depression Rating Scale (HAM-D) score ≥17. Remitted participants had a past MDD episode and current HAM-D ≤7. Controls and first-degree relatives had no personal history of depression and HAM-D ≤7; relatives additionally had at least one first-degree relative with depression (established using the Family Interview for Genetic Studies, FIGS).

Participants completed the Digit Span (forwards and backwards) and Wechsler Test of Adult Reading to estimate IQ.

Participants were instructed that they would be compensated £30 plus performance-based bonuses (up to £20, based on their performance on a random selection of trials from the task battery). Ethical approval was granted by the UCL Research Ethics Committee (fMRI/2013/005) and the London Queen Square NHS Research Ethics Committee (10/H0716/2).

Procedure

Participants attended two sessions. At session one, clinical interviews, questionnaires and cognitive tests (digit span, WTAR) were completed. Within two weeks, participants attended session two, where they completed cognitive tasks (including the 4AB and Gambling tasks) in a randomized order.

Psychiatric symptom measures

Symptoms were assessed using validated questionnaire measures of depression, anxiety, apathy, anhedonia, optimism/pessimism and dysfunctional attitudes (see Supplement).

Confirmatory factor analysis (non-depressed groups only; to avoid biasing factor structure with extreme scores) was conducted on questionnaire subscale scores (listed in Supplement) and identified four latent factors: low mood, apathy, anhedonia and dysfunctional attitudes (in line with previous work10), with acceptable fit (CFI=0.901, RMSEA=0.08, SRMR=0.064). Correlations with task parameters are reported for Study 2; two outliers (>3SD above the mean) were excluded.

4-Armed Bandit Task (4AB)

On each trial, participants selected one of four stimuli and received one of four outcomes: reward, punishment, both, or neither. Reward and punishment probabilities for each stimuli fluctuated independently via a slow random walk. Participants were instructed that reward and punishment contingencies were dynamic and independent across stimuli and were instructed to maximize rewards and minimize punishments. The task comprised 200 trials (~15 minutes).

Gambling Task

Participants chose between a probabilistic gamble (50/50 probability) and a certain outcome. Participants were not shown the outcome of their choices. Trials were either mixed (50/50 gain vs loss OR certain 0-points) or gain-only (50/50 gain vs 0-points OR certain gain). This allows distinction of risk-versus loss-aversion.

The first task phase (50 mixed, 40 gain-only trials) determined participant’s individual indifference point (i.e., difference in expected value between the gamble and sure option such that the participant selected either option at chance); the second phase presented individually calibrated offers (64 mixed, 56 gain-only trials) based on this indifference point (see17 for further details). The task lasted ~15 minutes.

Model-Agnostic Statistical Analysis

For the 4AB task, the probability of repeating a choice following win-only, loss-only, no outcome (neutral), and mixed (simultaneous win and loss) was calculated. Effects of outcome (within-subjects: win, loss, neutral, mixed) and group (between-subjects: controls, relatives, remitted, depressed) were assessed using mixed-effects ANOVA. For the gambling task, the probability of choosing the gamble in mixed and gain-only trials was calculated.

Computational Modelling

For the 4AB task, seven reinforcement learning models were fitted using hierarchical Bayesian estimation (hBayesDM in Stan28), using hierarchical priors for each participant group, in line with prior studies10,29 (see Supplementary Materials).

To select models, we searched previous literature modelling the same or similar learning and decision-making tasks. This uncovered two sets of learning models: classical Rescorla-Wagner reinforcement learning models; and models integrating RL with prospect utility functions from Prospect Theory. Most models were selected based on prior work demonstrating good fit to 4-AB data29,30 and relevance to mood disorders6. A model with an additional decay term that accounts for forgetting of unchosen options was also included, following 31The second set of models were selected and compared to more classical RL models, based on evidence from Ahn et al. 32,33 of their efficacy for capturing decisions in similar tasks (Iowa Gambling Task, Soochow Gambling Task; although they did not perform well for our data).

Models were compared using leave-one-out information criterion (LOOIC). Although the six-parameter lapse-decay model showed the lowest LOOIC in Study 1, model checking indicated that the simpler five-parameter lapse model performed better (Table S1; Figure S1, S2; Supplementary Materials).

For the gambling task, three Prospect Theory models were fitted to the calibrated second phase. Our modelling approach was derived from previous work17, from which the gambling task was directly adapted. As in previous work, we compared a full three-parameter model including risk aversion, loss aversion and inverse temperature parameters to simpler models without risk aversion or without loss aversion. The best-fitting model included all three terms.

Results were robust to the use of either a single prior for all participants or separate priors for each group (see Supplementary Materials).

Bayesian Model Averaging

To avoid bias from single-model selection, we applied Bayesian Model Averaging (BMA). Model weights were derived from information criteria and used to generate participant-level parameter estimates by sampling from weighted posterior distributions. Only parameters present in at least two models were retained. Full model specifications and predictive accuracy analyses are provided in the Supplement (see Figure S1, S2).

Results

Demographic and clinical characteristics are presented in Table 1. There were no significant differences across groups in demographics (age, gender, years in education) or digit span performance. However, the groups significantly differed in IQ (WTAR; see Table 1); thus, we repeated all analyses with IQ included as a covariate. IQ was not significantly associated with any task parameters, and did not moderate associations between group and parameters, therefore, we report models without including IQ (see Table S3 for results including IQ as covariate).

Demographic and clinical characteristics of the included subjects.

4-Armed Bandit Task

Model-Agnostic Variables

As expected, participants were more likely to repeat a choice following a win than after other outcomes in both studies (Study 1: F(3,294)=114.00, p<0.001, η2=0.538; Study 2: F(3, 573)=310.73, p<0.001, η2=0.619). In Study 2, no significant effect of participant group was detected (F(3,191)=0.18, p=0.91, η2=0.003), nor was there a significant group×outcome interaction (F(3,582)=1.31, p=0.27, η2=0.027).

Model-agnostic analyses of the 4-Armed Bandit Task.

Note. Violin plots showing probability to stay after (A) win, (B) loss, (C) both, (D) neither. CTR: healthy control; REL: participants with a first-degree relative with depression; REM: remitted depressed participants; MDD: currently depressed participants.

Group differences in parameter estimates

One-way ANOVA analyses on the BMA-derived parameters in Study 2 revealed significant group differences in punishment learning rate (F(3,191)=5.44, p=.001) and lapse (F(3,191)=4.37, p=.005).

Post-hoc tests indicated that the punishment learning rate was lower in the remitted group, compared with all other groups: depressed group (t(96)=-3.86, p<.001, d=-.77), relatives (t(70)=-2.58, p=.012, d=-.57), and healthy controls (t(97)=-2.22, p=.029, d=-.43); additionally, the depressed group had a higher punishment learning rate than healthy controls (t(101)=1.99, p=.050, d=.38). The lapse rate was also lower in the remitted group, compared with all other groups (depressed: t(79)=-3.44, p<.001, d=-.69; relatives: t(52)=-2.76, p=.008, d=-0.66; controls: t(106)=-2.09, p=.038, d=-.39).

Analyses conducted using parameters derived from the two best-performing models (the “lapse” model and the “lapse_decay” model) yielded similar results, supporting the robustness of these findings. No other parameters demonstrated significant group differences (see Table 2).

Group differences in computational parameters for the 4-Armed Bandit task.

Note. (A) Reward learning rate; (B) Punishment learning rate; (C) Reward sensitivity; (D) Punishment sensitivity; (E) Lapse. Note: CTR: healthy controls; REL: participants with at a first-degree relative with Depression; REM: remitted depressed participants; MDD: currently depressed participants. *p<.05 **p<.01***p<.001.

Parameter estimates and group comparisons on the 4 parameters in the “lapse” model for the 4-Armed Bandit task and the “ra_prospect” model in Gambling task.

Associations between parameters and symptoms

Associations between symptom factor scores and computational parameters were examined for Study 2 in the non-depressed groups only. Punishment learning rate was significantly negatively associated with apathy (r=-0.18, p=0.033) and anhedonia scores (r=-0.18, p=0.035).

Association between factor scores and punishment learning rate.

Note. Points represent individual participants. Solid lines show linear regression fits with 95% confidence intervals (shaded). Left panel: anhedonia factor scores (r = -.18, p=.035); right panel: apathy factor scores (r = -.18, p=.033).

Gambling task

Model-agnostic analysis

As expected, participants were more likely to gamble on mixed compared to gain-only gamble trials in both studies (Study 1: t(97)=1.96, p=0.026; Study 2: F(1,194)=7.39, p=0.007, η2=0.038). No significant groupxgamble type interaction was observed in Study 2 (F(3, 194)=0.61, p=0.61, η2=0.009).

Model-agnostic analyses of the Gambling task.

Note. Violin plots of the gambling task showing probability to take gamble choice under gain-only gamble (A) and mixed gamble (B) conditions in Study 2. CTR: healthy controls; REL: participants with at a first-degree relative with depression; REM: remitted depressed participants; MDD: currently depressed participants.

Group differences in parameter estimates

One-way ANOVA on parameters derived from the winning model in Study 2 revealed a significant group effect in inverse temperature (Table 2, Figure 5C). Post-hoc comparisons indicated that, compared with controls, depressed participants (t(70)=2.61, p=.011, d=0.53) and first-degree relatives (t(43)=2.32, p=.025, d=0.58) had higher inverse temperature values. No other significant parameter-factor-score associations were observed.

Group differences in computational parameters for the Gambling task.

Note. (A) Risk aversion; (B) Loss aversion; (C) Inverse temperature. CTR: healthy controls; REL: participants with at a first-degree relative with depression; REM: remitted depressed participants; MDD: currently depressed participants.*p<.05

Discussion

Using complementary paradigms assessing reinforcement learning and decision-making under uncertainty, we compared individuals with current depression, remitted depression, at familial risk for depression, and healthy controls. Computational modelling allowed us to extract latent cognitive processes and provided greater sensitivity when detecting between-group differences, compared with model-agnostic analyses. Remitted individuals uniquely showed reduced punishment learning rates, which may reflect a compensatory mechanism that maintains remission in unmedicated individuals. Among non-depressed participants, higher punishment learning rates were associated with apathy- and anhedonia-related factor scores. Remitted participants also showed less noisy (or non-modelled), more deterministic choice behaviour in the reinforcement learning (4-armed bandit) task, and depressed and those at familial risk showed noisier choice behaviour in the value-based decision-making (gambling) task.

Reduced punishment learning in remitted individuals relative to all other groups suggests that recovery may involve qualitatively distinct learning patterns, rather than simple normalisation to healthy control levels. Although this pattern requires confirmation in longitudinal studies, lower punishment learning implies that negative outcomes exert weaker influence on future expectations, reducing rapid avoidance following losses. Because the remitted group was largely asymptomatic and unmedicated, these patterns likely reflect recovery-related processes rather than residual symptom burden. Although causality cannot be inferred, given limited information about participants’ treatment histories, this pattern aligns with principles emphasized in cognitive behavioural therapy and behavioural activation20,27,34–37, which encourage sustained engagement despite occasional negative outcomes and discourage overgeneralization from unlikely adverse events. The fact that punishment learning rate is lower in remitted individuals than in all other groups suggests that it may be a compensatory mechanism. While computational psychiatry has begun to inform treatment design5,19,38, its relevance for (relapse) prevention remains underexplored. If replicated longitudinally, punishment learning dynamics may represent a candidate target for early or recoverymaintaining intervention.

Our results were aligned with previous meta-analytic work and a now largely replicated effect, showing that participants with depression, compared with healthy controls, demonstrate elevated punishment learning rates. Further, as in previous work linking exaggerated updating following negative outcomes to motivational symptoms14,15, we found an association between punishment learning rate and anhedonia and apathy, among nondepressed groups. This supports further evidence that alterations in learning from negative outcomes may underlie core symptoms of depression.

Group differences in the lapse and inverse temperature parameters further suggest alterations in how learned values guide action. Remitted participants showed more deterministic choice and less random responding in the learning task, whereas depressed and at-risk individuals exhibited greater variability in the risky decision-making task. While these may reflect (dis)engagement with the task, (in)effective exploitation of accumulated knowledge or consistency or rigidity in decision making, it is also possible that these group differences partly reflect variation in model fit across groups. Although fitting a common model to all participants is necessary for direct comparison, different groups may rely on partially distinct decision strategies that are not fully captured within the current modelling framework.

We observed no group difference in reward or punishment sensitivity. This supports prior evidence that depression-related alterations are more consistently expressed in learning dynamics and choice variability, rather than the subjective evaluation of outcomes3,15. Some earlier studies have reported altered reward sensitivity or response biases in individuals at familial risk for depression, which we did not find here26,39; however, many of these tasks employed stable reward contingencies. The 4AB task’s volatility may attenuate sensitivity differences observed in more stable paradigms, as optimal behaviour involves placing greater weight on recent outcomes and flexibly updating expectations in response to environmental change11,12.

Several limitations of our study warrant consideration. Groups were heterogeneous; for example, older individuals in the familial risk group may represent a more resilient subset, while remitted participants potentially achieved recovery through diverse treatments (which were also not fully characterised). We also are unable to determine whether the remitted group had elevated punishment learning rates during a previous symptomatic episode. Our study design rests on the assumption that individuals in the currently depressed group are representative of those that will go on to remit, and that never-depressed relatives are representative of those who will go on to develop depression. While our groups are well-matched on key demographic variables, we cannot rule out the possibility that unmeasured factors differ systematically between groups in ways that violate this assumption. This is an inherent limitation of cross-sectional designs. Future longitudinal studies following individuals before, during, and after depressive episodes will be essential to determine whether these learning dynamics predict episode onset, relapse or sustained recovery.

Additionally, the present study employed two complementary paradigms assessing distinct computational processes — reinforcement learning under uncertainty and value-based decision-making under explicit risk — analysed separately given their fundamentally different theoretical frameworks and model structures. Future work might consider joint estimation approaches that share information across tasks and participant samples, which could improve parameter precision and allow direct examination of cross-task concordance; however, the benefits of such an approach would need to be demonstrated empirically, particularly given that RL and prospect theory parameters have been shown to be largely uncorrelated across tasks 40.

In summary, our results support a compensatory model: remitted depression was associated with reduced punishment learning and more consistent choice behaviour, compared with all other groups; we also partly support a state model, in that, depressed participants showed elevated punishment learning rates compared to healthy controls, but the differences between depressed participants and the at-risk group failed to reach significance. Speculatively, these findings suggest that sustained recovery from depression may involve adaptive changes in how negative experiences are integrated into future expectations. Longitudinal replication is needed to establish whether these computational phenotypes confer resilience and represent mechanistic targets for prevention.

Data availability

Data is available upon request to the authors.

Supplementary materials

Symptom questionnaires

Beck Depression Inventory (Beck, Steer & Brown, 1996)

Apathy Evaluation Scale (four subscales: cognitive apathy, behavioural apathy, emotional apathy and other; Marin et al., 1991)

Snaith-Hamilton Pleasure Scale (Snaith et al., 1995), State-Trait Anxiety Inventory (state and trait subscales; Spielberger, 1983)

Life Orientation Test-Revised (to measure dispositional optimism and pessimism; Scheier, Carver & Bridges, 1994)

Temporal Experience of Pleasure Scale (anticipatory and consummatory hedonia subscales; TEPS; Gard, 2006)

Hamilton Depression Rating Scale (HAMD; Hamilton, 1960)

Dysfunctional Attitudes Scale (social approval and perfectionism subscales; DAS; Weissman & Beck, 1978).

Model-agnostic analyses on 4-arm bandit task

Outcome measures for behavior based on summary statistics:

p(stay) after win (Pstay/Win): number of repeated choices after wins only/total number of win trials;

p(stay) after loss (Pstay/Loss): number of repeated choices after losses only/total number of loss trials;

p(stay) after neither (Pstay/Neither): number of repeated choices after neither only/total number of neither trials;

p(stay) after neither (Pstay/Both): number of repeated choices after both win(loss)/total number of both trials.

Model and prior fits

Model analysis on 4-arm bandit task

The bandit4arm models were calculated by inputting reward (rew) and punishment (pun) values separately into the following equations, where i refers to a given bandit and t refers to trial.

Choice probability was determined by passing the reward and punishment values through a Softmax function in the _4par model, where n represents the total number of bandits:

Regarding the _lapse model, the addition of an irreducible noise parameter (lapse) was added to account for the possibility of decisions made at random, irrespective of the inferred values of the bandits.

For the _2par_lapse model, there are no sensitivity parameters in equations (3) and (4). For the _singleA_lapse model, there is a single learning rate across equations (1) and (2). For the lapse-decay model, a decay rate gradually reduced the values of the unchosen bandits to zero:

Beyond the above models, two IGT_pvl models were also used, which are prospect valence learning models that integrate aspects of reinforcement learning and prospect theory models (for details, please see Ahn et al., 2008; Ahn et al., 2014).

Model-agnostic analyses on Gambling task

P(gamble) on mixed trials: number of gamble choices/total number of mixed trials;

P(gamble) on gain-only trials: number of gamble choices/total number of gain-only trials.

Model analysis on Gambling task

The winning gambling task model, ra_prospect, was estimated with the following equations:

On each trial t the subjective expected value (EV) of the gamble and sure option was calculated. These subjective expected values were passed through a Softmax function to calculate the estimated probability of choosing the gamble option:

Details of Model Fitting procedures

Model fitting was conducted using four Markov Chain Monte Carlo (MCMC) chains, with 1,000 burn-in samples and 4,000 post-burn-in samples per chain.

Bayesian model averaging

Bayesian Model Averaging (BMA) was used to derive parameters, which integrates information from multiple models by weighting their contributions according to their fit to the data for each participant. Each model was initially fitted using separate hierarchical priors for each group: healthy controls (Study 1) and depressed, remitted, first-degree relatives, and healthy controls (Study 2); and the winning model was defined as the one with the lowest leave-one-out information criterion (LOOIC) across groups. However, this single-model selection method follows a “winner-takes-all” approach, which may introduce bias; for example, if different models are favoured by different groups. Furthermore, model checking analyses indicated that while the best-fitting model for the 4AB task (“lapse decay”) provided the lowest LOOIC, its parameter recovery was substantially inferior to that of the “lapse” model.

Model weights were computed using a pseudo-BMA weighting procedure. Bayesian Information Criterion (BIC) values were normalized by subtracting each BIC from the maximum BIC and ensuring that the sum of all weights equaled one. Final parameter estimates for each participant were derived by sampling from the posterior distributions of all models, weighted by their respective BMA weights. To ensure parameter stability, only parameters present in at least two models were estimated. In cases where a parameter was not defined in a given model (i.e., “NA” values), it was imputed as 0.001 to maintain consistency in the parameter estimation process.

Four-armed bandit and gambling task procedures.

Note. (A) Example trial of the four-armed bandit task. On each trial, participants chose one out of four bandits and received one out of four possible outcomes: reward (green token), punishment (red token), neither reward nor punishment (empty box) or both reward and punishment (red and green token). (B) On each gambling task trial, participants chose between a 50–50 gamble and a sure (guaranteed amount of points) option. Trials were either mixed gambles (50–50 chance of winning or losing points or sure option of 0 points) or gain-only trials (50–50 chance of winning or receiving nothing or sure gain).

Group difference in parameters for the lapse model, using a single prior for the 4-arm bandit task

Note. (A) Reward learning rate; (B) Punishment learning rate; (C) Reward sensitivity; (D) Punishment sensitivity; (E) Lapse. (F) Decay rate. Note: CTR: Control Participants; REL: Participants with at a first-degree relative with Depression; rMDD: Remitted depressed participants; MDD; Currently depressed participants.

Group difference in parameters of the ra_prospect model using single prior in the Gambling task.

Note. (A) Risk aversion; (B) Loss aversion; (C) Inverse temperature. Note: CTR: Control Participants; REL: Participants with at a first-degree relative with Depression; rMDD: Remitted depressed participants; MDD; Currently depressed participants.

Acknowledgements

This research was funded by a Wellcome Trust grant (101798/Z/13/Z) to JPR. PD was funded by the Max Planck Society and the Humboldt Foundation. OJR was funded by an MRC Senior Non-Clinical Fellowship. The authors report no biomedical financial interests or potential conflicts of interest.

Additional information

Funding

Wellcome Trust (WT)

https://doi.org/10.35802/101798

  • Jonathan Roiser

Humboldt Area Foundation (HAF)

  • Peter Dayan

Medical Research Foundation (MRF)

  • Oliver J Robinson