Evidence for flexible regulation of movement and decision vigor in a reward-oriented task

  1. Université Paris-Saclay, Inria, CIAMS, Gif-sur-Yvette, France
  2. Université Paris Cité, CNRS UMR 8002, INCC - Integrative Neuroscience and Cognition Center, Paris, France
  3. Centre national d’études spatiales (CNES), Paris, France

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Alaa Ahmed
    University of Colorado Boulder, Boulder, United States of America
  • Senior Editor
    Tamar Makin
    University of Cambridge, Cambridge, United Kingdom

Reviewer #1 (Public review):

Summary:

The authors introduce a foraging task designed to test whether movement and harvesting durations are governed by a common objective combining reward, effort, and time or by separate objectives. Participants generate wrist torque to move a cursor to a rewarding patch and then hold it within the patch while reward accumulates at a diminishing rate. The task allows reach and harvest effort, as well as delays before or between the two phases, to be manipulated separately.

Reach effort strongly affected reach duration but had little effect on harvest duration, whereas a delay imposed after reaching altered harvest duration without affecting the completed reach. Manipulation of harvest effort did not affect either duration. The authors also found consistency across individuals within each type of duration but no reliable association between reach and harvest durations. The separate model fit the behavioral results better than the common-utility model, particularly for the post-reach delay condition. The findings therefore support partly dissociable, but coupled optimization of harvest and movement in this task, while not excluding common optimization in other contexts.

Strengths:

The task provides a controlled way to manipulate the effort and delays associated with reaching and harvesting while measuring both durations. The control experiments alter the timing of delays, allowing the experimenter to test whether their effects depend on when they occur. These manipulations produce clear behavioral dissociations, particularly the selective effect of a post-reach delay on harvest duration. The delay and effort manipulations were verified, and most participants reported noticing the changes in effort and time.

The separate model extends a previous cost-of-time framework to a foraging task involving movement and reward harvesting. Its assumptions have a clear relationship to the experimental design: applied delays and actions enter the time offset for each separate cost function. This provides an intuitive account of why a delay occurring after the reach can affect harvesting without affecting the completed movement. The separate model also provides a better fit than the common utility model in both modeled experiments.

The individual-level behavioral analyses are also useful. Reach durations are correlated across reach conditions, and harvest durations are correlated across harvest conditions, indicating stable individual differences within each measure. The lack of a corresponding relationship between reach and harvest durations raises an important question about whether the two forms of vigor necessarily reflect a common individual tendency in their task.

Weaknesses:

The block design and familiarization made the changes in effort and delay learnable, and most participants reported awareness of these manipulations. The observed dissociation may reflect adaptation to this task structure. This does not weaken the behavioral result, but it limits how directly the findings can be generalized to tasks in which changes in reward, effort, or time are less explicit.

The main modeling limitation is that both models are fitted to ten group-level median observations in each experiment, rather than to individuals' behavior. The separate model has seven fitted parameters compared with three in the common model. Although the reported AIC penalizes additional parameters, the analysis includes no participant-level fitting, resampling, or cross-validation.

Reviewer #2 (Public review):

Summary:

The manuscript by Conessa and colleagues tests whether movement time and decision time are co-regulated in humans using a novel harvesting paradigm and computational modelling of the behavioral data. The data contradict the common utility model and are better supported by an alternative model where the costs of the harvesting and movement phases are optimized separately - rather than jointly.

Strengths:

The major strength of the study is the development of the novel paradigm where the forces applied during the distinct phases of the task (reaching and harvesting) can be perfectly controlled and measured.

This elegant manipulation allows precise estimation of subjective effort costs, and later can be used to better understand how these costs shape behavior.

The novel separate-cost model is also an interesting development.

Weaknesses:

In my opinion, the manuscript suffers from too many weaknesses to solidly support the general conclusions:

(1) The manuscript oversells the significance of the results;

(2) The novel model appears to be built ad hoc based on the results, and its ingredients are poorly motivated

(3) There are important methodological shortcomings. I develop all of these points below. The paper also lacks clarity, particularly when presenting the key concepts of the novel model.

Importantly, I also recommend that the authors follow open science policy and publish their data and code on a public repository.

(1) Significance

As mentioned above, the question of the co-regulation of the vigor of different behaviors is a very interesting topic, but I believe the significance of the results is overstated, for different reasons:

a) This is a harvesting task, which means that decisions are specifically about the duration of the reward period. In other words, more time directly provides more (immediate) reward, whereas in other types of tasks more time allows gathering and processing more information, and so reaching a more informed decision, which is a more indirect way of securing better outcomes - see e.g. Drugowitsch, Moreno-Bote et al. JNeurosci 2012 for optimizing decision time in this most common setting. This major difference is not discussed in the manuscript, and it is not clear whether the results found here would hold for the most common type of decisions. This is also different from the time to initiate a response, as explored by Morvan et al. in their cited preprint.

b) The title of the study mentions "flexible" co-regulation, but it's hard to see direct evidence for flexibility in the results. This flexibility is extrapolated in the Discussion, but in order to make such a statement, it is necessary to show that the co-regulation differs between different conditions, as predicted from the model.

c) The lack of impact of the harvesting effort on harvest time is very surprising, and the null value of the sensitivity to effort (d=0) in the model is intellectually uncomfortable. As some value of effort, it SHOULD matter. I am not very satisfied with the explanation provided in the Discussion ("possible hypothesis would be that the perceived effort of harvest was outweighed by the reward value"): this trade-off should be precisely dealt with by the model.

(2) Separate-Cost Model

The separate-cost model is poorly motivated and not clearly presented, in particular the indirect coupling between the two phases.

a) Some key components are motivated by the constraint to obtain a well-behaved optimization problem ('we need to find a cost that increases with time' ) rather than based on cognitive principles. Some other key components seem to have been placed there just to fit the data better, but their plausibility is not even discussed. Notably:
i- The fact that the offset o_h includes the reaching delay is key to the coupling between t_h and t_r, but is not properly motivated. Why should this offset include all preceding intervals within the trial (leading to double counting of the cost of the reaching phase, in both G_r and G_d)? In this cyclic paradigm where participants aim at maximizing cumulative reward, it is not even clear that they parse behavior into different trials starting as they need to reach for a new target.
ii- Why is there a free parameter d for the cost of harvest effort? If time and motor costs are converted into the same currency, and that conversion holds for all types of phases, then I believe we should have d=1. This free parameter seems to have been introduced a posteriori to accommodate the lack of effect of harvest effort on harvest time, but is not motivated.

b) The two models differ in two fundamental aspects: whether optimization is run separately on the different phases of the trial, or jointly; and how the different costs and rewards are assembled in the to-be-optimized metric. Which aspect is more important to capture behavior is not assessed.

c) The presentation of the "Separate-cost model" is very obscure. For example, the notion of a `cost of reward` is rather odd, and consequently the formula for R(t_h) appears arbitrary. I understand that this is the difference between the maximum reward that can be collected within a trial and the actual reward that is collected, but it is not even presented as such. The effort cost of harvesting seems to be discarded in the model presentation but does appear in equation 14.

(3) Methodological Weaknesses

I believe that there are numerous methodological weaknesses that really limit the inferences that can be drawn from the experimental and modelling results.

a) Sample sizes are small, especially for control experiments, and it is not clear how they have been selected. This is especially problematic because the only manipulation that shows no significant effect on behavior (harvest duration) is only tested with the smallest sample size (n=10), severely limiting the significance of this negative result.

b) The rationale for the different manipulations in different experiments is not clear. Notably, why were modulations of the harvest effort not included in the main experiment? Because not a single participant sees a modulation of both reach and harvest efforts, we cannot evaluate whether the model can capture both effects at the same time (i.e., with the same parameter values). The number of data points used for fitting directly depends on the number of conditions performed by the participant, and is very low compared to the number of parameters to be fitted. For example, the model used very different shapes for the costs of time (Figure 8) for different sets of conditions, but this cost of time should be fixed, so it's not clear whether a single cost of time function can accommodate all manipulations.

c) Best practice in behavioral modelling is to model single-subject data, as fitting the model jointly to all subjects completely obliterates subject heterogeneity. Fitting on single-subject data would also allow investigating whether the models reproduce the consistency of harvest durations and reach durations across conditions, and most importantly the absence of correlation across subjects between reach and harvest duration. At the moment, we're left wondering what the implications of these analyses for the original question are.

d) What is the rationale for comparing reach duration in one condition to harvest duration in another condition?

e) The study does not take into account the likely impact of delays on muscular fatigue and perceived effort: the effort associated with maintaining the cursor on target during harvest might be perceived as less effortful after a longer break, partially explaining the impact of delay on harvest duration.

f) I am confused by the reference to indirect effects in the separate-cost model (L692-7, Table 2). To fit this model, was the experimental value or theoretical value of the reach duration t_d used to determine the optimal harvest value (from Equation 14)? In both cases, the reach duration DOES increase with reach effort, so we should expect an impact on the harvest duration. This is indeed seen in Figure 7A. So why does Table 2 suggest that reach effort should not impact harvest duration in this model?

g) Could the effect of pre-reach delay on reach time be alternatively explained simply by longer reaction times as reported in Figure S2, e.g., a longer period to reactivate the muscles after a longer break? In other words, as far as I understood, here reaction times are subsumed in the motor time, against the tradition of splitting them into two independent (non-overlapping) behavioral metrics.

h) The EMG-time integral in the harvest phase should be largely proportional to the harvest duration, given that the applied torque is constrained to remain on target. Is this the case? If so, then how to interpret the significant effect of reach effort on this value, given that the main analyses did not identify an effect of reach effort on harvest duration?

Author response:

Reviewer #1 (Public review):

Summary:

The authors introduce a foraging task designed to test whether movement and harvesting durations are governed by a common objective combining reward, effort, and time or by separate objectives. Participants generate wrist torque to move a cursor to a rewarding patch and then hold it within the patch while reward accumulates at a diminishing rate. The task allows reach and harvest effort, as well as delays before or between the two phases, to be manipulated separately.

Reach effort strongly affected reach duration but had little effect on harvest duration, whereas a delay imposed after reaching altered harvest duration without affecting the completed reach. Manipulation of harvest effort did not affect either duration. The authors also found consistency across individuals within each type of duration but no reliable association between reach and harvest durations. The separate model fit the behavioral results better than the common-utility model, particularly for the post-reach delay condition. The findings therefore support partly dissociable, but coupled optimization of harvest and movement in this task, while not excluding common optimization in other contexts.

Strengths:

The task provides a controlled way to manipulate the effort and delays associated with reaching and harvesting while measuring both durations. The control experiments alter the timing of delays, allowing the experimenter to test whether their effects depend on when they occur. These manipulations produce clear behavioral dissociations, particularly the selective effect of a post-reach delay on harvest duration. The delay and effort manipulations were verified, and most participants reported noticing the changes in effort and time.

The separate model extends a previous cost-of-time framework to a foraging task involving movement and reward harvesting. Its assumptions have a clear relationship to the experimental design: applied delays and actions enter the time offset for each separate cost function. This provides an intuitive account of why a delay occurring after the reach can affect harvesting without affecting the completed movement. The separate model also provides a better fit than the common utility model in both modeled experiments.

The individual-level behavioral analyses are also useful. Reach durations are correlated across reach conditions, and harvest durations are correlated across harvest conditions, indicating stable individual differences within each measure. The lack of a corresponding relationship between reach and harvest durations raises an important question about whether the two forms of vigor necessarily reflect a common individual tendency in their task.

We thank the reviewer for acknowledging the relevance of our study.

Weaknesses:

The block design and familiarization made the changes in effort and delay learnable, and most participants reported awareness of these manipulations. The observed dissociation may reflect adaptation to this task structure. This does not weaken the behavioral result, but it limits how directly the findings can be generalized to tasks in which changes in reward, effort, or time are less explicit.

The results show that, in a controlled environment, the common-utility model did not account for certain behaviors observed in our participants. In variable environments, because the previous history of reward and effort affect the current trial (Yoon et al., 2018; Sukumar et al., 2024), the behavioral regulation could be less stable and apparent. Here, we provided a favorable environment to assess the common-utility model, with stable manipulations of time and effort in a block-wise design allowing clear estimation of the capture rate by the participants. We agree with the reviewer that it is not necessarily generalizable to other types of environments involving uncertainty for instance. However, demonstrating that movement and decision vigor can be dissociated in a specific task is good evidence to suggest that it is not always co-regulated, and should therefore be flexible. This potential adaptation resulting from the development of explicit knowledge is a problem present in most studies of this type, where conditions are introduced in a blocked manner (Saleri Lunazzi et al., 2021; Sukumar et al., 2024). We will discuss this point in more detail in the revised version of the manuscript.

The main modeling limitation is that both models are fitted to ten group-level median observations in each experiment, rather than to individuals' behavior. The separate model has seven fitted parameters compared with three in the common model. Although the reported AIC penalizes additional parameters, the analysis includes no participant-level fitting, resampling, or cross-validation.

We acknowledge that fitting the models to participants’ individual behavior could provide better insight into the ability of the evaluated models to accurately capture the effects of time and effort on behavior. We will perform these analyses during the review.

Reviewer #2 (Public review):

Summary:

The manuscript by Conessa and colleagues tests whether movement time and decision time are co-regulated in humans using a novel harvesting paradigm and computational modelling of the behavioral data. The data contradict the common utility model and are better supported by an alternative model where the costs of the harvesting and movement phases are optimized separately - rather than jointly.

Strengths:

The major strength of the study is the development of the novel paradigm where the forces applied during the distinct phases of the task (reaching and harvesting) can be perfectly controlled and measured.

This elegant manipulation allows precise estimation of subjective effort costs, and later can be used to better understand how these costs shape behavior.

The novel separate-cost model is also an interesting development.

We thank the reviewer for acknowledging the strengths of our study and for the constructive suggestions. Below, we provide preliminary responses with some planned adjustments.

Weaknesses:

In my opinion, the manuscript suffers from too many weaknesses to solidly support the general conclusions:

(1) The manuscript oversells the significance of the results;

(2) The novel model appears to be built ad hoc based on the results, and its ingredients are poorly motivated

(3) There are important methodological shortcomings. I develop all of these points below. The paper also lacks clarity, particularly when presenting the key concepts of the novel model.

The model relies on a large part of the literature that assumes a cost of time affecting movement and decision-making. For instance, the urgency-gating model associate a time-growing signal with the pace of decisions (Cisek et al., 2009; Thura, 2020; Saleri Lunazzi et al., 2021); and a cost of time underlying human reaching vigor has been identified in previous studies (Shadmehr et al., 2010a; Choi et al., 2014; Berret et al., 2018; Berret and Baud-Bovy, 2022; Verdel et al., 2023). Our model is simply grounded on those studies. We will detail further the rationale behind the separate-cost model and address as much as possible the points highlighted by the reviewer.

Importantly, I also recommend that the authors follow open science policy and publish their data and code on a public repository.

The authors agree with the reviewer. Data are available in Supplementary Data and code used will be published on a public repository.

(1) Significance

As mentioned above, the question of the co-regulation of the vigor of different behaviors is a very interesting topic, but I believe the significance of the results is overstated, for different reasons:

a) This is a harvesting task, which means that decisions are specifically about the duration of the reward period. In other words, more time directly provides more (immediate) reward, whereas in other types of tasks more time allows gathering and processing more information, and so reaching a more informed decision, which is a more indirect way of securing better outcomes - see e.g. Drugowitsch, Moreno-Bote et al. JNeurosci 2012 for optimizing decision time in this most common setting. This major difference is not discussed in the manuscript, and it is not clear whether the results found here would hold for the most common type of decisions. This is also different from the time to initiate a response, as explored by Morvan et al. in their cited preprint.

One feature of the present study, consistent with previous work (Yoon et al., 2018; Sukumar et al., 2024), is the use of a harvesting task to study the decision-making process (Hayden, 2018). The context of optimal foraging provides an interesting approach to study decision-making, as it is directly linked to ecological contexts. Remaining in a food patch does indeed result in an immediate increase in reward, but staying there too long risks reducing the capture rate, which, according to the utility model, represents the actual value that matters. The subject has to gather information about the current state of the patch and the global capture rate of the environment to decide how long he should harvest. Here, the aim of the article was not to generalize the results to all types of decisions, but to test the utility model in a very specific and controlled environment to better dissociate the effects of time and effort. In our task, movement and decision vigor is not co-regulated, which can be seen as a counterexample of a systematic co-regulation. However, there are a lot of possible tasks combining movement and decision-making in different forms.  As such, we argue that behavior may benefit from a certain degree of flexibility that allows for the separation of movement and decision-making vigor when relevant to a given task. Depending on the context, optimal behavior may require moving quickly while deciding quickly, moving quickly while deciding slowly, or other combinations (see Thura et al., 2025 for a review). In this respect, the separate-cost model allows flexible regulation (coupled or not), hence capturing which regulation is optimal, as long as the objective costs (defined by the task) and the participant’s subjective costs (specific to the participant) are appropriately combined. The common-utility model is comparatively more constrained because movement and decision costs are intrinsically integrated.  We agree that this distinction warrants further discussion and will clarify it, together with the resulting limits on generalization, in the revised manuscript.

b) The title of the study mentions "flexible" co-regulation, but it's hard to see direct evidence for flexibility in the results. This flexibility is extrapolated in the Discussion, but in order to make such a statement, it is necessary to show that the co-regulation differs between different conditions, as predicted from the model.

The term flexible is used to highlight that the rigid co-regulation predicted by the common-utility model does not apply to all manipulations of effort and time in this controlled design, nor to all types of decision-making (harvest vs reaction time). Decisions in the form of harvest durations do not seem to co-vary with reach durations, contrary to decisions performed during reaction times. Accordingly, and based on recent evidences (see Thura et al. 2025 for a review), this co-regulation appears to be flexible, depending on the actual optimal behavior required by the task and the type of decisions involved. In our paper, we found a co-regulation between reaction times and movement times but an independence of harvest times and movement times. Together with existing results from the literature, we conclude about the flexibility of co-regulation. However, we acknowledge that the (co-)regulations have not been directly compared in our paper. We plan to further assess this point.

c) The lack of impact of the harvesting effort on harvest time is very surprising, and the null value of the sensitivity to effort (d=0) in the model is intellectually uncomfortable. As some value of effort, it SHOULD matter. I am not very satisfied with the explanation provided in the Discussion ("possible hypothesis would be that the perceived effort of harvest was outweighed by the reward value"): this trade-off should be precisely dealt with by the model.

We agree with the reviewer and the lack of an effect of harvest effort on its duration was surprising to us too. However, the null value of sensitivity to harvest effort should not be interpreted as the effort does not influence the harvest or the decision in general. This effort is indeed likely not significant enough to necessitate a change in behavior. The free weighting parameters capture the relative sensitivity of the costs involved. Although effort was noticeable for the participants and of the same magnitude as variations in movement effort, it is likely they did not find it disruptive enough to adjust their behavioral strategy, demonstrating that it was not a critical factor in these conditions. We note that both the cost of time and the harvesting cost of effort increase with time, which suggests that in this timescale and this level of effort, the cost of time may have had a greater influence on behavior. Yet, our model could have captured effort-related variations of harvest duration if there were observed experimentally. To observe such effects, we should have increased harvesting effort but it might have led to behavioral changes due to muscular fatigue or maximal isometric torque limitations (which are not considered in the current modeling) rather than to the optimization of subjective utility. We will clarify what this sensitivity of harvest effort represents, and examine alternative constraints on the free parameters.

(2) Separate-Cost Model

The separate-cost model is poorly motivated and not clearly presented, in particular the indirect coupling between the two phases.

a) Some key components are motivated by the constraint to obtain a well-behaved optimization problem ('we need to find a cost that increases with time' ) rather than based on cognitive principles. Some other key components seem to have been placed there just to fit the data better, but their plausibility is not even discussed. Notably:

i) The fact that the offset o_h includes the reaching delay is key to the coupling between t_h and t_r, but is not properly motivated. Why should this offset include all preceding intervals within the trial (leading to double counting of the cost of the reaching phase, in both G_r and G_d)? In this cyclic paradigm where participants aim at maximizing cumulative reward, it is not even clear that they parse behavior into different trials starting as they need to reach for a new target.

Precisely, as the reviewer pointed out, this offset is key to the coupling of th and tr, and it actually constitutes an original extension to previous `cost of time` models. Including tr in the offset oh allows the separate-cost model to co-regulation to some extent, without leading to double-counting the reach cost. It must be noted that Gr and Gh are distinct time cost functions, and the value of Gh at tr can be different from that of Gh at </>tr. The offset allows reach duration to influence harvest duration because it shifts its cost of time. Recent evidence show that subjects rely only to a limited extent on their past experience when assessing the quality of their environment (Sukumar et al., 2024). Hence, our hypothesis is that only recent delays could influence the cost of time, in the form of an offset that creates a bias affecting the vigor. For this reason, we chose to define it as the cumulative time elapsed from the start of each trial. Due to the cyclic nature, the models are then applied to the average behavior within a condition, allowing us to reduce variability. We plan to make the rationale clearer in the revised manuscript.

ii) Why is there a free parameter d for the cost of harvest effort? If time and motor costs are converted into the same currency, and that conversion holds for all types of phases, then I believe we should have d=1. This free parameter seems to have been introduced a posteriori to accommodate the lack of effect of harvest effort on harvest time, but is not motivated.

The number of free weighting parameters is defined by the minimization of a linear combination of cost functions. The reaching phase comprises only two costs in this model, leading to the possibility to have only one free weighting parameter that capture the relative sensitivity to time and effort. However, regarding the harvest phase, as there are three different costs, it was necessary to have at least two free weighting parameters to account for relative sensitivity and convert all costs in the same currency. Alternatively, we could indeed have applied a free parameter to the reward and set d = 1. However, the purpose of the study was to assess the contribution of time and effort changes, and ultimately participants’ sensitivity to changes in these factors. As such, it appeared appropriate to evaluate the sensitivity of the harvest effort captured by the models, rather than the reward (which is not manipulated). It was also to keep some consistency with the common utility model, which includes a free weighting parameter for the costs of effort and time (and none for the reward function) as well. We will add more details regarding the motivation behind the separate-cost model and the use/choice of these free parameters. We agree that the rationale for this parameterization was not sufficiently described in the manuscript and will clarify it in the revised version. We will also examine alternative constraints on the free parameters.

b) The two models differ in two fundamental aspects: whether optimization is run separately on the different phases of the trial, or jointly; and how the different costs and rewards are assembled in the to-be-optimized metric. Which aspect is more important to capture behavior is not assessed.

We thank the reviewer for raising this interesting point. It is difficult to determine which aspect is most important: the separate versus joint optimization, or how the costs are assembled (additive vs multiplicative). However, the two are closely linked. It is the switch from a multiplicative association of the cost (common-utility model) to an additive one (separate-cost) that makes the separate optimization possible. For example, it is not possible to achieve an optimal movement duration using the costs interaction from the common-utility model, since there would simply be a discounting of the cost of effort by the cost of time . It would be necessary to add another element justifying why the movement should not be infinitely slow, which would therefore require changing the costs or adapting them to capture a logical behavior. This is what led to the design of the separate-cost model, by switching to an additive model allowing an optimal duration to emerge for each phase. We will further examine this issue in the revision, in particular by testing whether the separate-cost model can also account for observed behaviors under a joint optimization of reaching and harvesting durations.

c) The presentation of the "Separate-cost model" is very obscure. For example, the notion of a `cost of reward` is rather odd, and consequently the formula for R(t_h) appears arbitrary. I understand that this is the difference between the maximum reward that can be collected within a trial and the actual reward that is collected, but it is not even presented as such. The effort cost of harvesting seems to be discarded in the model presentation but does appear in equation 14.

The terminology `cost of reward` is admittedly odd but is used because it is the counterpart of the harvest function used in the common-utility fh(th). Instead of using it in a form subject to maximization, we use it in a form subject to minimization. The purpose of this function is to maintain a certain consistency between the models and the various types of costs implied; as such it was crucial to incorporate a function capturing the harvested reward in the separate-cost model. We will detail further this point in the revised version, as we agree that our previous explanation was insufficient. The effort cost of harvesting was not presented again in the separate-cost model section since it was already identified (as was the effort cost of reaching) in the common-utility section and referenced to the corresponding equation.

(3) Methodological Weaknesses

I believe that there are numerous methodological weaknesses that really limit the inferences that can be drawn from the experimental and modelling results.

a) Sample sizes are small, especially for control experiments, and it is not clear how they have been selected. This is especially problematic because the only manipulation that shows no significant effect on behavior (harvest duration) is only tested with the smallest sample size (n=10), severely limiting the significance of this negative result.

We acknowledge that some of the sample sizes are small. Importantly, the main goal of the study was to determine the effect of time and effort manipulations on the behavioral vigor. The first main experiment had already provided us with lot of insights to reach this objective, but to go further and explore our subsequent hypotheses, we decided to examine the role of the delay position and another type of effort cost. The change in harvesting effort is not the only manipulation that showed no effect on behavior. In the main experiment, we observed that changes in effort required for reaching does not affect the harvest duration, even though this is predicted by the common-utility model. In the same experiment, we also observed that changes in delays do not affect the reach duration, even though this is also predicted by common-utility model. Nevertheless, we will discuss the limit of the control experiments further in the revised version.

b) The rationale for the different manipulations in different experiments is not clear. Notably, why were modulations of the harvest effort not included in the main experiment? Because not a single participant sees a modulation of both reach and harvest efforts, we cannot evaluate whether the model can capture both effects at the same time (i.e., with the same parameter values). The number of data points used for fitting directly depends on the number of conditions performed by the participant, and is very low compared to the number of parameters to be fitted. For example, the model used very different shapes for the costs of time (Figure 8) for different sets of conditions, but this cost of time should be fixed, so it's not clear whether a single cost of time function can accommodate all manipulations.

We agree that the rationale for the different manipulations across experiments was not sufficiently explained. Traditionally, optimal foraging has not been studied in conjunction with movement vigor (Yoon et al., 2018), but a recent framework accounts for both decision-making and motor control (Shadmehr et al., 2016). In the main experiment, we chose to manipulate reaching effort because our primary aim was to test whether variations in the cost involved in motor control influence (i.e., co-regulate) the decision-making process. The manipulations of effort and time in the main experiment were a good basis for evaluating the common-utility model in a controlled environment. Showing that reaching effort would affect harvest duration would have already been great evidence to illustrate a form of co-regulation. Moreover, we had no certainty that the participants would exhibit stable behavior, and adding twice as many conditions risked introducing further biases related to fatigue and attentional focus. We decided to run control experiments upon seeing the results of this initial experiment, when other questions arose regarding the effect of the delay position and of the type of effort. We acknowledge that performing all manipulation in one large experiment would indeed have significantly increased the significance of the models’ predictions. Nonetheless, we made sure that the number of parameters fitted were taken into account by using the AIC to evaluate models’ performance. Regarding the assumption that the shape of the cost of time should be fixed: adding a waiting period changes (slightly) the task, and it is conceivable that a different task should be represented by a different cost of time. Reaching and harvesting are too distinct tasks, and the subjective cost of elapsed time may therefore differ depending on whether participants are moving or remaining stationary. Different cost of times have been proposed for both arm-reaching movements and saccades (Shadmehr et al., 2010b; Berret et al., 2018), in particular because these movements occur over very different timescales. This suggests that how the passage of time is valued may depend on the task at hand. Altogether, this led us to consider letting the cost of time have different shapes for each phase and each experiment (although fixed within the sigmoidal family of function). We will clarify the rationale for the different manipulations and the use of different cost of time for each phase in the revised manuscript.

c) Best practice in behavioral modelling is to model single-subject data, as fitting the model jointly to all subjects completely obliterates subject heterogeneity. Fitting on single-subject data would also allow investigating whether the models reproduce the consistency of harvest durations and reach durations across conditions, and most importantly the absence of correlation across subjects between reach and harvest duration. At the moment, we're left wondering what the implications of these analyses for the original question are.

We agree that fitting these models to the individual behavior of the participants would be interesting to assess the models’ ability to capture participants’ variability. We will explore this during the revision.

d) What is the rationale for comparing reach duration in one condition to harvest duration in another condition?

If behavioral vigor (either of movement or decision) is a trait-like characteristic supported by shared processes, some degree of consistency between movement and decision vigor is expected. For instance, participants who tend to move faster might also tend to decide faster, if their vigor relates to traits like impulsivity for instance (Choi et al., 2014; Berret et al., 2018; Labaune et al., 2020). As such, we decided to perform all correlations to provide a comprehensive overview of inter-individual consistency of vigor in a foraging-like task.

e) The study does not take into account the likely impact of delays on muscular fatigue and perceived effort: the effort associated with maintaining the cursor on target during harvest might be perceived as less effortful after a longer break, partially explaining the impact of delay on harvest duration.

The idea that delays may influence muscular fatigue and perceived effort is intriguing. However, we believe that if perceived effort was the reason leading to changes in harvest duration, changes would also have been elicited by manipulation of harvest effort. This point will be discussed further during the revision.

f) I am confused by the reference to indirect effects in the separate-cost model (L692-7, Table 2). To fit this model, was the experimental value or theoretical value of the reach duration t_d used to determine the optimal harvest value (from Equation 14)? In both cases, the reach duration DOES increase with reach effort, so we should expect an impact on the harvest duration. This is indeed seen in Figure 7A. So why does Table 2 suggest that reach effort should not impact harvest duration in this model?

We agree that this point was not sufficiently clear. The purpose of Table 2 was to present the specific effect of each modulation. Thus, the prediction that reach effort does not affect harvest duration in the separate-cost model refers specifically to a change in reach effort at a fixed reach duration. An example with our experimental design would be to display a virtual movement of fixed duration during which participants have to produce a torque (isometric contraction) superior to a certain threshold that could vary between conditions. The inter-dependence of reach and harvest are based on duration changes; manipulation of the magnitude of effort exerted during movement at fixed duration would therefore not affect harvest duration. Table 2 was intended to distinguish these direct and indirect effects rather than to predict that reach effort should have no effect on harvest duration in our experiment. To fit the model to the data, the experimental value of the reach duration was used. However, since Table 2 shows only the directions of the variations, an arbitrary fixed value for reach duration (of the same magnitude) was used to determine the direct effect represented. We will clarify this distinction and the purpose of Table 2 in the revised manuscript.

g) Could the effect of pre-reach delay on reach time be alternatively explained simply by longer reaction times as reported in Figure S2, e.g., a longer period to reactivate the muscles after a longer break? In other words, as far as I understood, here reaction times are subsumed in the motor time, against the tradition of splitting them into two independent (non-overlapping) behavioral metrics.

This is an interesting point raised by the reviewer and will be discussed further in the revision. Here, we assume that a shorter reaction times are associated with a less steep cost-of-time function. According to drift-diffusion models, it brings the decision-related neural activity more slowly toward the commitment threshold, thereby delaying movement initiation. Since the cost of time is less steep, the movement duration will also be increased, thus implying this co-regulation between movement and reaction time.

Here, reaction time was not separated, since the protocol is not appropriately designed to its study, particularly in the main experiment. The detection of reaction time is highly conservative (see text of Time-Related supplementary analyses) and it is not entirely clear whether we can properly refer to a reaction time in the main experiment, since the two phases follow one another without any delay. Hence to avoid any uncertainty, we chose not to separate these presumed reaction times from motor times.

h) The EMG-time integral in the harvest phase should be largely proportional to the harvest duration, given that the applied torque is constrained to remain on target. Is this the case? If so, then how to interpret the significant effect of reach effort on this value, given that the main analyses did not identify an effect of reach effort on harvest duration?

We thank the reviewer for pointing this interesting effect. It is possible that a greater muscle activation during the reaching phase was carried over to the harvest phase (due to increased co-contraction for instance), leading to an increase in the EMG-time integral in the harvest despite manipulation of reach effort. We will explore this during the revision. In particular, we will examine how the EMG signals recorded during the reaching relate to those recorded during the harvesting, as well as the relationship between the EMG signals and the torque exerted by the participants.

References

Berret B, Baud-Bovy G (2022) Evidence for a cost of time in the invigoration of isometric reaching movements. J Neurophysiol 127:689–701.

Berret B, Castanier C, Bastide S, Deroche T (2018) Vigour of self-paced reaching movement: cost of time and individual traits. Sci Rep 8:10655.

Choi JES, Vaswani PA, Shadmehr R (2014) Vigor of Movements and the Cost of Time in Decision Making. J Neurosci 34:1212–1223.

Cisek P, Puskas GA, El-Murr S (2009) Decisions in Changing Conditions: The Urgency-Gating Model. J Neurosci 29:11560–11571.

Hayden BY (2018) Economic choice: the foraging perspective. Curr Opin Behav Sci 24:1–6.

Labaune O, Deroche T, Teulier C, Berret B (2020) Vigor of reaching, walking, and gazing movements: on the consistency of interindividual differences. J Neurophysiol 123:234–242.

Saleri Lunazzi C, Reynaud AJ, Thura D (2021) Dissociating the Impact of Movement Time and Energy Costs on Decision-Making and Action Initiation in Humans. Front Hum Neurosci 15:715212.

Shadmehr R, Huang HJ, Ahmed AA (2016) A Representation of Effort in Decision-Making and Motor Control. Curr Biol 26:1929–1934.

Shadmehr R, Orban de Xivry JJ, Xu-Wilson M, Shih T-Y (2010a) Temporal discounting of reward and the cost of time in motor control. J Neurosci Off J Soc Neurosci 30:10507–10516.

Shadmehr R, Xivry JJO de, Xu-Wilson M, Shih T-Y (2010b) Temporal Discounting of Reward and the Cost of Time in Motor Control. J Neurosci 30:10507–10516.

Sukumar S, Shadmehr R, Ahmed AA (2024) Effects of reward and effort history on decision making and movement vigor during foraging. J Neurophysiol 131:638–651.

Thura D (2020) Decision urgency invigorates movement in humans. Behav Brain Res 382:112477.

Thura D, Haith AM, Derosiere G, Duque J (2025) The integrated control of decision and movement vigor. Trends Cogn Sci:S1364-6613(25)00185-8.

Verdel D, Bruneau O, Sahm G, Vignais N, Berret B (2023) The value of time in the invigoration of human movements when interacting with a robotic exoskeleton. Sci Adv 9:eadh9533.

Yoon T, Geary RB, Ahmed AA, Shadmehr R (2018) Control of movement vigor and decision making during foraging. Proc Natl Acad Sci 115:E10476–E10485.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation