Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorJian LiPeking University, Beijing, China
- Senior EditorMichael FrankBrown University, Providence, United States of America
Reviewer #1 (Public review):
Summary:
This manuscript investigates state-, trait-, and recovery-related computational phenotypes in Major Depressive Disorder (MDD) using an unmedicated, four-group cross-sectional design comprising currently depressed participants, remitted individuals, first-degree relatives at familial risk, and healthy controls. Utilizing a volatile four-armed bandit task and an explicit risk gambling task alongside hierarchical Bayesian modeling, the authors report that remitted participants uniquely display a lower punishment learning rate relative to all other groups. The authors conclude that recovery from MDD does not represent a simple normalization to healthy baseline levels, but rather involves a protective computational recalibration that dampens reactivity to negative outcomes to sustain remission.
Strengths:
(1) Highly Valuable Clinical Sample: Evaluates a rare, unmedicated sample across four distinct clinical stages (MDD, REM, REL, CTR), providing an exceptionally controlled framework to disentangle state, trait, and recovery markers.
(2) Combines dynamic reinforcement learning under environmental volatility with prospect-theoretic decision-making under explicit risk to capture multiple dimensions of value-based choice.
(3) Challenges the conventional assumption of clinical "normalization," offering a compelling hypothesis that psychiatric recovery may depend on active, compensatory recalibrations of cognitive parameters.
(4) Utilizes hierarchical Bayesian parameter estimation and leverages Bayesian Model Averaging (BMA) to mitigate single-model selection bias.
Weaknesses:
(1) The Discussion characterizes MDD and familial-risk groups as exhibiting "noisier" choice behavior in the gambling task, which directly contradicts the reported higher inverse temperature values that mathematically denote more deterministic choices.
(2) Framing a reduced punishment learning rate as a "recovery mechanism" overinterprets single-timepoint data, which cannot differentiate an acquired post-episode adaptation from a pre-existing resilience trait.
(3) The theoretical claim that lower punishment learning is protective in remission directly conflicts with the authors' dimensional findings, where lower punishment learning correlates with worse subclinical apathy and anhedonia in non-depressed participants.
(4) Model-agnostic choice repetition yielded no significant group effects, contrasting sharply with the robust group differences in model-derived parameters and necessitating posterior predictive checks.
(5) Fails to provide parameter recovery analyses to demonstrate that punishment learning rates can be reliably disentangled from lapse rates and outcome sensitivities across a 200-trial task structure.
(6) Selects a lower-ranked model under LOOIC without sufficient quantitative justification, and lacks sensitivity analyses to confirm that group-specific hierarchical priors did not skew estimates given unequal group sizes.
(7) Relies on several marginal p-values bordering across multiple parameters and symptom correlations without establishing a clear family-wise error or FDR correction strategy.
Reviewer #2 (Public review):
This manuscript reports on reinforcement learning in participants with current depression, remitted depression (without current depression), in people without depression but with first-degree relatives with depression, as well as healthy controls. Participants completed two common tasks measuring reinforcement learning and risk aversion, and their behavior was fit to computational models assessing processes on these tasks.
Participants with remitted depression showed a lower punishment learning rate and more value-concordant decisions on the reinforcement learning task. Relatives of depressed participants, as well as people who are currently depressed, had higher inverse temperature, indicating more value-driven choices. Within the non-depressed participants, the punishment learning rate was negatively associated with symptoms of anhedonia and apathy.
I have reviewed this paper at a previous journal. This revised version is responsive to most of my concerns, particularly in terms of placing the manuscript more in the context of other related literature and providing more details on methods. There are some remaining concerns about sample size and the appropriateness of some of the methods (e.g., interpreting participant-level parameter estimates from hierarchical models estimated using BMA), but the latter has been adequately addressed with sensitivity analyses.
Reviewer #3 (Public review):
Summary:
This manuscript describes an interesting study (preceded by a pilot study) that combined computational modeling with two tasks to try to tease apart state, trait, and vulnerability-related effects of depression. Specifically, the authors administered a bandit task and a gambling task to healthy controls, currently depressed adults, formerly depressed adults, and adults at familial risk of depression, and then used a suite of computational models to identify group differences in key model parameters. Key results included the detection of lower punishment learning rates (during the bandit task) in the remitted depressed group, which the authors hypothesize may be a compensatory mechanism to counteract the over-reaction to negative feedback that often characterizes depression (and indeed, punishment learning rates were elevated in currently depressed adults). Among non-depressed adults, lower punishment learning rates were associated with more apathy and anhedonia. The remitted depressed group also showed lower lapse rates in the bandit task. In the gambling task, the inverse temperature parameter was elevated in currently depressed adults and in adults at high familial risk for depression; exactly how to interpret this last result seems unclear. This is a revised manuscript, and from what I can see it appears that the authors were responsive to earlier comments.
Strengths (and summary of weaknesses):
I think the study has several noteworthy strengths. Testing four groups is a strength, as there is great interest in teasing apart risk factors of depression vs. "scars" of the illness, and the use of unmedicated individuals removes a common confound. The work is hypothesis-driven, the modeling is sophisticated, and the paper is well-written. Moreover, as the authors note, modeling allowed the authors to identify group differences that were not evident in raw behavioral analyses.
However, I think the manuscript could be further improved because some aspects of the methodology are a bit confusing and/or do not seem optimal. I list these issues in the comments below, but to briefly summarize: (a) I did not understand why the authors estimated separate reward and punishment values for each bandit; (b) the rationale for Bayesian Model Averaging could be strengthened; (c) examining relationships between model parameters and symptom scores only in the control group is suboptimal given the goal to better understand depression; and (d) the group differences in inverse temperature seem like they may depend on potential outliers. I think addressing these concerns would improve the paper and ensure that the study has a strong impact on the field.
Details regarding weaknesses/concerns:
(1) I found a basic aspect of the 4-arm bandit modeling confusing - namely, the use of separate reward and punishment value estimates (see equations 1 and 2 in the supplement). The participants are choosing among the bandits (presumably) based on value estimates for each bandit. Clearly, delivery of rewards and punishments affect those estimates, but it is not clear to me why or how there are separate value terms for rewards and punishments; typically, rewards and punishments influence one overall value estimate per bandit. Can the authors clarify? (I see that similar models were used in references 29 and 30, but additional clarification for readers of the current manuscript would be helpful)
(2) Bayesian Model Averaging (BMA) is new to me and may be new to many readers, and it would be helpful to provide a stronger rationale for the approach. The manuscript argues that model selection amounts to a "winner-takes-all" approach that may introduce bias; maybe, but typically the goal is to figure out which mechanism(s) best explain behavior, and so winner-takes-all is often appropriate. My limited understanding is that BMA is often used when prediction-rather than mechanistic understanding-is the goal. Why is BMA the right choice here?
(3) The authors performed a confirmatory factor analysis on questionnaire data from healthy controls out of concern that extreme scores in the other groups might bias the factor structure; consequently, they can only relate model parameters to their latent factors in the healthy controls. This does not seem optimal. While the negative relationships between punishment learning rate and both anhedonia and apathy in controls are interesting, the controls are not struggling with anhedonia or apathy. It would be valuable to know if similar relationships obtain in the other groups, particularly because this would speak to the study's goal of distinguishing between risk for depression and state/trait aspects of depression.
(4) Figure 5 gives the impression that group differences in inverse temperature may depend on potential outliers in the relative and MDD groups. Is that correct?
(5) The paper notes a group difference in IQ as estimated from the WTAR, and looking at Table 1 it appears that the group difference is driven by a lower WTAR score in the healthy volunteers (HV) from the pilot study. First things first: the score in the HV group is too low to be a standardized WTAR score. Like IQ scores, standardized WTAR scores typically have a mean around 100; scores for the four Study 2 groups look alright, but the Study 1 mean score of 40 is much too low. Can the authors clarify?