Figures and data

Contrast comparison task and behavioural results.
(A) Stimulus and evidence time courses: The stimulus consisted of two interleaved gratings flickering on and off in alternation, whose relative contrast had to be judged. 0.6 sec after achieving fixation and triggering stimulus onset, the fixation point turned yellow to cue participants to prepare for the onset of contrast-difference evidence another 0.6 sec later. In all trials, the stimulus remained on for a full 1.6 sec following evidence onset, after which it disappeared and the fixation turned green to prompt the subject to make a left/right-hand button press to indicate that the left- or right-tilted grating had higher-contrast, within a 0.5-sec deadline. In 80% of trials, a weak (±10%) contrast increment/decrement was oppositely applied to the two gratings for 0.2, 0.4, 0.8, or 1.6 sec, returning to equal baseline contrast thereafter until stimulus offset. In the remaining 20% of trials, an easy contrast increment/decrement of ±40% was applied for the entire 1.6-sec period. (B) Accuracy significantly increased with the duration of the stimulus for low-contrast trials (blue dots) and was at the ceiling for the high-contrast condition (red dot). Individual participants’ accuracies are shown in grey lines. (C) The mean reaction time (RT; relative to evidence onset) slightly decreased in trials with longer durations and higher contrast levels. Correct responses (solid, with individuals also shown in grey) were also slightly faster than error responses (dashed) in this delayed response task. For b and c, data are mean ± s.e.m. after between-participant variance was factored out (subtracting each participant’s mean across conditions and adding the grand mean), in keeping with the repeated-measures experimental design.

Alternative models and their fits to behaviour only.
(A) Model schematics. In the unbounded integration model (‘UIntg;’ top left), the sign of the accumulated evidence at stimulus offset determines the choice, while in bounded integration (‘BIntg;’ top right), the decision is either terminated early by reaching the correct or error bound or otherwise determined in the same way as in the unbounded model. In the unbounded Extremum-tracking model (‘UExtE;’ middle left), the decision is based on the sign of the evidence sample that attains the maximum absolute value, which by making use of the full stimulus-assessment period can produce accuracy improvements with increasing evidence duration. In the bounded Extrema detection model (middle right), a choice is determined as soon as a single sample exceeds a bound or, if neither bound is reached, the choice is determined by: i) a random guess (‘BExtG’); ii) the sign of the last sample (‘BExtL’); or iii) the sign of the sample that attains the most extreme value (‘BExtE’). Finally, the Snapshot model, where the choice is based on a single sample chosen at random at any time (highlighted with the larger dot), is shown in the bottom row. In the unbounded version (‘USpst’), the decision is made based on the sign of the sample. In the bounded version (‘BSpst’), the decision is based on whether the sample surpasses the bound; otherwise, the decision is guessed. (B) Observed (black dots) and model-predicted (bars) accuracy, showing that the five best-fitting models can capture the observed accuracy improvement with duration well. (C) AIC values across the seven most competitive behaviour-only models show that models without bounds are favored due to parsimony, and selection among integration and most of the extrema detection mechanisms is largely inconclusive. Given that we computed G2 values through Monte-Carlo simulation, we conducted the fits with 10 different instantiations of noise, each shown here as a separate point (see Methods).

Goodness of fit metrics and parameter estimates of models fit to behaviour only.
Models beginning with U were unbounded, and beginning with B were bounded, with the bound parameter denoted by B. Drift rates for low and high contrast difference were denoted by D1 and D2, respectively; single-drift rate variants (denoted ‘_1d’) of the unbounded models provided a worst-case benchmark for each mechanism and s refers to common scaling parameter. G2 quantifies how closely the model quantitatively fits the data, while Akaike’s Information Criteria (AIC) adds a penalty for complexity. For both metrics, lower values indicate better fits. Colour shading highlights the competitive models depicted in Figure 2C.

Neural Signals.
(A) SSVEP showed a robust representation of sensory evidence encoding. Note that the signal ramps rather than suddenly steps simply due to the width of the window for spectral estimation (230 ms), and lags the physical contrast changes simply due to transmission delays to the visual cortex. Topographies show the average SSVEP amplitude from 300 ms to 1500 ms (shaded period) for the high contrast conditions (top) and the longest low contrast condition (bottom), with a primary focus at standard site Oz. (B) Event-related potential (ERP) at centroparietal sites averaged over trials (regardless of outcome) within each condition, showing that the centro-parietal positivity (CPP) undergoes a gradual buildup and sustained elevation for lower contrast conditions, and a peak at around 530 ms for the high contrast trials, at more than twice the amplitude. Topographies show ERP amplitude for the high contrast (top) and the average over low contrast conditions (bottom), for the shaded time ranges 210-530 ms (left) and 1280-1600 ms (right). (C) Motor preparation was measured by subtracting ipsilateral from contralateral mu/beta-band activity (MB; 8-30 Hz) with respect to the correct side. The topographies show the difference between trials favouring a left minus right response, again shown for both high contrast (top) and averaged low contrast (bottom) conditions in the shaded windows 570-890 ms (left) and 1280-1600 ms (right).

Neurally constrained models and their predictions.
Among the models with various additional parameters, starting-point variability and collapsing bounds stood out as helping to achieve good CPP waveform fits without substantially sacrificing behavioural fits (Figure 4 - Figure Supplement 1), and given that both of these features were exhibited empirically in motor preparation signal effects (Figure 5), we show these models here (but see Figure 4 - Figure Supplements 4-6 for basic versions of the models without the extra parameters). (A) Single-trial sensory evidence traces were simulated for each of the five conditions using boxcar functions (stepping up during positive evidence) with added Gaussian noise as in behaviour-only models (one example for each condition shown). (B-D) single-trial examples of the decision variable in each of the three competing models, again simulating one trial per condition. (B) In the Integration model, the evidence was integrated over time until it reached a bound, after which it fell back to zero. Where the decision variable did not hit the bound the decision was made based on the sign of accumulated evidence at 1.6 sec. (C) In the Extremum-tracking model, the DV keeps track of the most extreme value so far, until it reaches the bound, after which it falls to zero. The sign of the extremum determined the choice, and the most extreme sub-bound value was used in the event that it did not cross the bound. (D) In the Extremum-flagging model, an all-or-nothing Half-Sine-profile signal marking choice termination onsetting at the bound crossing time (see main text). As in the other models, one example DV trace per condition is shown; for three of these (red and two dark blue traces) the bound was reached, while in the others it was not, so no signal was generated (flat light blue lines), and the choice was made as a random guess. In all 3 models in this figure, starting point variability and collapsing bound parameters are included, and they have an equal number of free parameters. (E) Akaike’s Information Criterion (AIC) derived from only the behavioural component of the neurally-constrained fit, plotted as a function of the weighting of CPP-waveform constraint. With increasing emphasis on neural data in the fit, the fit to behaviour was compromised to varying extents, and the Extremum-flagging model was best at retaining good behavioural fits at higher neural weights. (F) Overall neurally-constrained objective function (behavioural G2 value plus weighted neural penalty quantifying divergence in observed vs simulated CPP) as a function of neural constraint weighting. The inset compares the objective function values with starting point variability and collapsing bound parameters (labelled ‘Extra’ in G) to the ‘basic’ models without either feature, at w = 10. (G) Neural fit (R2) plotted against behavioural fit (G2), illustrating a trade-off between how well the models can capture behaviour and the CPP up to w = 100. The closer the joint values come to the origin the better they are at simultaneously fitting both. (H-J) Predicted behavioural accuracy for a range of neural constraint weightings (‘w’). (K-M) Simulated average CPP (where the DV is rectified on each single trial in keeping with CPP’s always-positive nature; see Methods) for the best fits of each model at a neural weighting of w=10 (see Figure 4 - Figure Supplement 2 for simulations from all weightings, and Figure 4 - Figure Supplement 3 for predicted bound crossing frequencies). The simulated, normalised average CPP (solid lines) is shown for all five conditions. The normalised empirical data is also shown (dashed lines; see Methods for details) with the waveforms for the four low-contrast conditions averaged together, as they were for the neurally-constrained fitting. (N-P) Simulated differential motor preparation. This is based on the same simulated decision variable as in (K-M) but now retaining its differential form by not taking the absolute value and having the signal sustain at the bound level after it is crossed. This is not normalised since it is not directly fit to empirical data. Smoothing is applied equivalent to the Fourier windowing in the Mu/Beta measurements.

Fit metrics and parameter estimates for neurally-constrained models with a neural-constraint weighting of w = 10.
B refers to the bound; D1 and D2 refer to the drift rates in the low- and high-contrast conditions, respectively. FlagWidth refers to the width of the half-sine, sz refers to starting-point variability, K refers to the temporal slope of the collapsing bound and s refers to within-trial noise, the common scaling parameter (see Tables S3 for fits using a nonlinear collapse function, or Extremum-flagging with last-sample default, with similar findings). Models including starting point variability and collapsing bounds are distinguished from basic models without those features, by the model name ending ‘_szC.’

Observed and simulated motor preparation and bound dynamics.
(A) Mu/beta (MB) decreases reflecting building motor preparation, contralateral and ipsilateral to the correct response, in correct (top) and error trials (bottom), shown aligned to evidence onset (left) and response (right). Contralateral traces, on correct trials, and ipsilateral traces, on error trials, reach a fixed threshold prior to response (marked by horizontal grey dashed line). Interestingly, on easy, high-contrast trials, this threshold (marked by the grey dashed horizontal line to aid comparison) is reached at approximately 800 ms and sustains at just below that level until the response. Note that on high-contrast trials, contralateral MB appears to reach a less decreased level prior to response compared to low-contrast trials. The topography inset shows the difference in MB between high-contrast and low-contrast conditions, revealing that this difference is unlikely motor-related, and more likely reflects greater attentional engagement in the harder low-contrast trials, known to be reflected in posterior alpha activity (Foxe & Snyder, 2011; Kelly & O’Connell, 2013). (B) Collapsing bound estimated from the empirical data and the model fits. Linear (solid) and non-linear (‘NL,’ dashed) collapsing bounds were calculated for Integration and Extremum-flagging models. All traces were normalised to the initial bound height. (C) Observed Mu/beta lateralisation for correct and error trials, demonstrating baseline choice-predictive effects that arise from starting-point variability. This can be compared to simulated average relative motor preparation at a neural constraint weighting of w = 10 for the Integration (D) and Extremum-flagging (E) models. Note again that models do not include non-decision delays, and so evidence-dependent dynamics begin at time zero in the simulated traces, whereas they are clearly delayed as usual in the empirically observed data. All simulations are from models containing both starting point variability and collapsing bounds.

Histograms of Bound-crossing times per condition in behaviour-only models that included a bound.
The integration model, BIntg, allows 55-60% of trials to exceed the bound in the low-contrast conditions and nearly all trials in the high-contrast conditions, with the relative proportion of correct vs error bound crossing scaling with duration. A similar pattern is observed in the BExtE and BExtL extrema detection models. The BExtG model produces a good fit by setting a relatively high bound and assuming a drift rate for low-contrast conditions that is almost as high as the high-contrast condition. This produces no error bound crossings and an almost uniform distribution of correct bound crossing times during the evidence periods for hard trials, which produces a linearly increasing ratio of correct bound crossings to guesses across durations, in turn yielding a linear increase in accuracies. The Snapshot model, BSpt, by construction, produces a bound crossing distribution that is uniform across all time for which there is evidence.

Simulated decision variable waveforms of the unbounded integration (A; UIntg) and non-integration (B; UExtE) models that produced good fits to behaviour alone, illustrating that neither of these models can reproduce the peak and falling-down of the empirical CPP signal (dashed) in the high-contrast condition.
For consistency with the neurally-constrained modelling analysis, empirical traces show the mean CPP segments that are used as constraints, with low-contrast conditions averaged (‘cppL’; See Neurally-constrained model-fitting in Results), whereas all five conditions are simulated for the models.

(A) Centroparietal positivity (CPP) waveform without the 5-Hz cutoff smoothing filter applied as in Fig 3 and Fig 4, demonstrating a bimodal morphology that is not present in the low-contrast conditions (shown here averaged across durations). For comparison, the waveforms from bilateral occipital sites are superimposed in dashed lines, showing that the initial peak in centroparietal buildup is contemporaneous with the peak of the N2 component. We verified that this pattern was visible in several individual subjects. (B) Series of topographies from the onset of activity through the dip in centroparietal buildup, for the high-contrast (top) and averaged low-contrast conditions (bottom), again indicating that the initial peak and dip of the high-contrast centroparietal waveform likely contains a contribution from the positive end of dipolar activity generating the occipital N2. Each labelled timepoint represents the centre of a 60-ms window across which the amplitude is averaged.

Impact of including additional, plausible free parameters to the basic neurally-informed, integration and Extremum-flagging non-integration models, taking the neural constraint weighting of 100 as a representative case.
(A) Full objective function value (G2+neural penalty term) across models at weighting of 100. For the integration models, both starting point variability (‘sz’) and the collapsing bound improved the fit, while the competitive Extremum-flagging model (BExtG-flag) underwent the greatest fit improvement with the addition of a collapsing bound. B,C) neural-behavioural fit plots illustrating the trade-off of fitting to both behaviour and the CPP for the additional-parameter models within each model class. For all models, the lowest weighting starts in the upper left part of the plot, and the trace moves downwards as better neural fits are achieved by weighting that aspect more, at some point curving rightwards as behavioural fit is compromised in favour of neural fit. (B) The integration model including starting point variability achieved a better trade-off in fitting jointly to behaviour and the CPP waveforms than any other single additional parameter (fit points reaching closer to the origin). Additionally adding a collapsing bound further improves neural fits (lower R2). (C) The Extremum-flagging model achieved a comparable trade-off, with good neural and behavioural fits, in this case benefitting more from a collapsing bound than starting point variability. “sampT” denotes the basic model with extra parameter as “sampling onset time”; similarly “eta” represents between-trial drift rate variability; “sz” indicates starting point variability; “Dur” corresponds to accumulation duration; “C” refers to models with a collapsing decision bound; “szC” represents models incorporating both starting point variability and a collapsing bound; and “Leak” denotes models with leaky accumulation.

Simulated average CPP using the estimated parameters of the bounded Integration (top row), Extrema-tracking (middle), and Extremum-flagging model (bottom) for the full range of neural constraint weightings, all with starting-point variability and collapsing bounds included (same models as shown in Figure 4).

Bound crossing frequencies predicted by the Integration and Extremum-flagging models with starting-point variability and collapsing bounds.
(A) Bound crossing frequencies for all easy and hard trials of the Integration model, with a zoomed-in version of just the low-contrast conditions in (B) for better visibility. Error bound crossings are shown in dashed lines against the negative y-axis. In general, the overall proportion of hard trial with bound crossings decreased markedly with increasing neural constraint weighting for the Integration model, from approximately 0.95 at w=0.1 to 0.42 at w=1000. (C & D) Bound crossings of the collapsing Extremum-flagging model. These decreased much less with increasing neural constraint weighting, from 0.92 at w=0.1 to 0.85 at w=1000.

Neurally constrained for the basic models (with no additional parameters beyond two drift rates and a constant bound).
(A) Akaike’s Information Criterion (AIC) derived from only the behavioural component of the neurally-constrained fit, plotted as a function of the weighting of CPP-waveform constraint. With increasing emphasis on neural data in the fit, the fit to behaviour was compromised to varying extents, and the Extremum-flagging model was best at retaining good behavioural fits at lower neural weights. (B) Overall neurally-constrained objective function (behavioural G2 value plus weighted neural penalty quantifying divergence in observed vs simulated CPP) as a function of neural constraint weighting. For visualisation purposes, the inset compares the objective function values to the ‘basic’ models at w = 10. The integration model outperforms at the neural weights w>1 (C) Neural fit (R2) plotted against behavioural fit (G2), illustrating a trade-off between how well the models can capture behaviour and the CPP up to w = 100. The closer the joint values come to the origin the better they are at simultaneously fitting both. (D-F) Predicted behavioural accuracy for a range of neural constraint weightings (‘w’) for (D) Integration (E) the Extremum-tracking and (F) the Extremum-flagging model.

Simulated CPP signals for all basic models (with no additional parameters beyond two drift rates and a constant bound), for each of the neural constraint weightings tested.
The simulated waveforms over the neural weights showed that both the integration model, BIntg, and Extremum-flagging model, BExtG-flag, are qualitatively successful in reproducing the CPP waveforms. The BExtE-track model sacrifices the fit to the easy condition to have a better prediction on the weak contrast conditions (The blue line fits better to the real CPP over the higher weights although the falling-down peak in the easy condition in red is not captured).

Bound crossing frequencies in correct and error trials of “basic” neurally constrained models with no additional parameters beyond two drift rates and a constant bound.
(A) Bound crossings of all easy and hard conditions for the Integration model. The portion of the hard trials crossing the bound decrease across the neural weights. In contrast, easier trials are able to surpass the bound at higher neural constraint weightings. The drastic change occurred at w > 1. Further, the number of error trials reduces by focusing more on the neural weights. (B) The bound crossings of only the low-contrast conditions in the bounded integration, BIntg, are shown again on their own for better visibility. The bound crossings are delayed in the higher neural weights toward the later time in the trials. (C) Bound crossings for all conditions of the Extremum-flagging model. Similar to the Integration model, a transition happened at w>1 where reproducing the early CPP peak improved. (D) A replotting of only the low-contrast conditions of the BExtG-flag model for visibility.

Neurally constrained models and their predictions for models with 4-fold scaled drift rate to match physical contrasts.
(A-B) Sensitivity analysis showed that the Extremum-flagging models outperformed the integration and Extremum-tracking models in capturing behaviour at higher neural weights. In line with the fact that the Integration model estimates a much higher ratio of drift rates when allowed to do so, it provides a poor fit for a constrained 4-fold scaling of drift rate. The inset in panel B compares the G2 plus penalty of these one-drift-rate-scaled (“1Ds”) models compared to the corresponding models with two free drift rates (“2D”) at w = 10. (C) The trade-off shows that the Integration and Extremum-flagging models with two free drift rates outperform their corresponding 4-fold scaled models, but the 4-fold scaled Integration model has particular difficulty capturing the CPP dynamics. (D-F) All models compensate for accuracy at higher weights to provide better fit to the neural data. The 4-fold scaled drift rate is not enough for the Integration model to capture accuracy in the high contrast trials. (G-I) Simulated normalised average CPP for the models at w = 10. The simulated normalised average CPP (solid lines) is shown for all five conditions. The empirical data is also shown (dashed lines) with the waveforms for the four low-contrast conditions averaged together. (J-L) Simulated motor preparation. Same as (G-I) except showing the trial-averaged differential DV (without taking the absolute value) and having the signal sustain beyond bound crossing.

Comparison of Extremum-flagging model with a default of taking the last sample (‘BExtL’)) vs taking a random guess (‘BExtG’), when a bound is not reached during the stimulus.
Apart from a tendency to sacrifice the fit to accuracy more at higher neural weightings, the model works very similarly for the two alternative default strategies. Parameter estimates were also similar at w = 10.

Pairwise comparisons of accuracy, showing a significant difference between each pair of consecutive durations.

Pairwise comparisons of mean RT.
There was a significant difference in reaction time between each pair of consecutive durations except for the shortest durations (0.2s vs 0.4s).

the mean and standard deviation parameter estimates of models across the 10 fits to behaviour only.
‘UIntg’ refers to the unbounded Integration model, ‘BIntg’ refers to the bounded Integration model, ‘UExtE’ refers to the unbounded Extremum-tracking, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a default guess. ‘B’ refers to the bound and D1 and D2 are the drift rates in the low- and high-contrast conditions, respectively. The parameter estimates tend to vary by no more than 2-5% except for the bound estimates in the Snapshot and Extremum-tracking models, in line with their inability to distinguish whether there is a bound or not.


Neurally constrained Integration, Extremum-tracking and Extremum-flagging models with various extra model features at neural constraint weighting w = 10.
‘BIntg’ refers to the bounded Integration model, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a default guess. ‘FlagWidth’ refers to the width of half-sine while t05 and K are semi-saturation constant and collapsing rate parameters of the collapsing bound function.

Allowed parameter ranges for model fits, i.e.
the minimum and maximum values that each parameter can take in the SIMPLEX (fminsearch) algorithm. ‘Default value’ indicates parameter settings used when the parameter is fixed rather than free. An asterisk * marks the parameters with 5 grid points in the models with both starting point variability and a collapsing bound.


Neurally constrained Integration, Extremum-tracking and Extremum-flagging models with one extra feature at neural constraint weighting w = 0.1
‘BIntg’ refers to the bounded Integration model, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a guess default. ‘FlagWidth’ refers to the width of half-sine. t05 and K are semi-saturation constant and collapsing rate parameters of the nonlinear collapsing bound function, whereas the linear collapsing bound function only has a slope parameter K.


Neurally constrained Integration, Extremum-tracking and Extremum-flagging models with one extra feature at neural constraint weighting w = 1.
‘BIntg’ refers to the bounded Integration model, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a default guess. ‘FlagWidth’ refers to the width of half-sine while t05 and K are semi-saturation constant and collapsing rate parameters of the collapsing bound function.


Neurally constrained Integration, Extremum-tracking and Extremum-flagging models with one extra feature at neural constraint weighting w = 100.
‘BIntg’ refers to the bounded Integration model, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a guess default. ‘FlagWidth’ refers to the width of half-sine while t05 and K are semi-saturation constant and collapsing rate parameters of the collapsing bound function.



Neurally constrained Integration, Extremum-tracking and Extremum-flagging models with one extra feature at neural constraint weighting w = 1000.
‘BIntg’ refers to the bounded Integration model, ‘BExtE’ refers to Extremum-tracking, and ‘BExtG’ refers to Extremum-flagging which has a guess default. ‘FlagWidth’ refers to the width of half-sine while t05 and K are semi-saturation constant and collapsing rate parameters of the collapsing bound function.