Author response:
The following is the authors’ response to the original reviews.
Public Reviews:
Reviewer #1 (Public reviews):
Weaknesses:
Reliance on self-reports
A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.
We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”
Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.
Statistical analysis and interpretation of the dynamics matrix
Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.
We thank the reviewer for raising this important methodological point. We address it in three parts.
(i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.
(ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T2, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.
(iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”
Terminology (controllability, stability, sensitivity, Gramian)
Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.
This is an excellent and important point. We have made the following changes throughout the manuscript.
(i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.
(ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”
(iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.
(iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.
(v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.
Reviewer #2 (Public reviews):
Online recruitment, selection effects, and remuneration
Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).
We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”
Intervention ongoing during the second block
Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.
This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”
Observation noise
As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.
We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.
Reliance on formal model comparison
Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.
We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.
We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.
Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.
Statistical limitations; Bayesian inference
The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.
We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.
Reviewer #3 (Public reviews):
Dual meanings of controllability
An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”
We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.
Strategy reminder and possible reappraisal
As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.
This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.
Mechanism of distancing effects (eye movement and oculomotor avoidance)
Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.
We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”
Recommendations for the Authors:
Reviewer #1 (Recommendations for the Authors):
(1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?
Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.
(2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?
We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.
(3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.
References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.
(4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.
The typo has been corrected.
(5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.
We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.
(6) Could the authors develop the rationale behind the bias in Equation 1?
We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.
(7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.
We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.
(8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).
We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.
(9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.
The sentence has been rephrased: ”Cumulative model weights (wj) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”
(10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.
We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.
(11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.
We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.
(12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.
We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.
<(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.
Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.
We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.
(14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?
The regression equation has been corrected to DV = β0 + β1IV + β2G + β3(IV × G) + ϵ, making the interaction term explicit.
(15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?
The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.
Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.
Beyond the visual trajectory overlays in Figure 4C, we have now added R2 and peak cross-correlation as a quantitative measure of individual fit quality. Mean R2 across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R2 values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.
(16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?
Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.
(17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.
Interaction effects are now included in the relevant figure caption (Figure 5).
(18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?
Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?
This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.
(19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?
This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.
(20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?
Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.
Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.
(21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?
The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.
(22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).
Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.
(23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.
We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.
(24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.
The supplementary tables have been restructured to a hierarchical format.
(25) Table I9 is not referenced in the manuscript.
A reference to this table has been added in the appropriate Results section.
(26) It’s a shame that neither data nor analysis code has been made available.
Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).
Reviewer #2 (Recommendations for the Authors):
(1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.
The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.
(2) p 5: ’on [not in] the recruiting platform’.
Corrected.
(3) p 8: Clarify notation of xtt = 1T (and same for u). What does this mean?
The notation has been corrected.
(4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.
The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.
(5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.
References have been added at each key definition.
(6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.
This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.
(7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?
A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.
(8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?
Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.
(9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?
We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.
(10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.
Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.
(11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].
This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.
Reviewer #3 (Recommendations for the Authors):
(1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.
More information about the power analysis have been added to the Participants section of the main text.
(2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).
We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.
(3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”
Thank you for this comment. Please see our response to the public review comments above, which we hope address this.
(4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?
See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.
(5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.
We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.