Author response:
The following is the authors’ response to the original reviews.
Public Reviews:
Reviewer #1 (Public review):
Summary:
This paper examines whether humans use protracted temporal integration in a noise-free, deferred-response contrast discrimination task, using a covert evidence-duration manipulation combined with EEG (SSVEP, CPP, Mu/Beta). The key finding is that evidence for protracted sampling is behaviorally and neurally supported, but even joint CPP + behaviour fitting cannot fully discriminate a standard integration (DDM) model from a novel "extremum-flagging" non-integration model. The paper is transparent about this outcome.
Strengths:
This is a well-conducted and well-written study that makes a genuine contribution to the perceptual decision-making literature by introducing a clean experimental design for probing temporal integration without participants adapting their strategy and demonstrating for the first time that a non-integration model (extremum-flagging) can replicate CPP waveform dynamics that have long been considered hallmarks of evidence accumulation. The transparent treatment of equivocal modelling outcomes is commendable.
Weaknesses:
My main concerns relate to statistical power, the under-specification of the and the extremum-flagging mechanism. Addressing these would greatly strengthen the paper.
(1) The sample of 16 participants (15, after the exclusion of one participant) is described as "close to similar EEG studies" with no formal power analysis. Given that the paper's core claim rests on subtle quantitative differences between two model classes - differences that are, by the authors' own admission, not sufficient to declare a winner - even a modest increase in sample size might yield a more decisive outcome. At a minimum, the authors should report a sensitivity analysis or post-hoc power calculation to indicate what effect sizes the current N could reliably detect, particularly for the rmANOVA comparisons and the neural constraint fitting.
We appreciate the reviewer’s concern regarding sample size and statistical sensitivity. To address statistical robustness throughout the paper, we have now reported effect sizes for our statistical tests (e.g. η2 for rmANOVA; including Tables S1 and S2), and we provide error-shading around the ERP waveforms to indicate the reliability of the key patterns our models are aimed at capturing (i.e. the dramatically higher and earlier CPP peak for high-contrast, and very little systematic differences across the four low-contrast durations - see revised Figure 3). We also conducted an indicative post-hoc power analysis using the G*Power software based on the behavioural data. Using the observed partial η2 = 0.44 for the test of duration effect on accuracy among only the low-contrast conditions, and the final sample of 15 participants, this amounts to a statistical power of 0.998.
On the model comparison, while we agree that larger sample sizes are generally beneficial for population-level inferences, we respectfully maintain that our current sample size is sufficient to support the core claim that qualitative dynamics of neural signatures of decision formation, usually assumed to reflect temporal integration, can be successfully reproduced using non-integration models in the delayed-response task conditions we examine here. Any marginal changes in quantitative fit resulting from having a higher N contribute to grand averages are unlikely to substantively alter this conclusion of the model comparison. The statistical reliability of the data to which our models are fitted is also bolstered by the number of trials (about 256 per condition per participant). It is common in behavioural modelling studies for data to be collected from a much smaller sample (e.g., fewer than 10 subjects) but with a high trial yield - a relevant precedent for us being Stine et al., (2020), who provided a compelling demonstration of similar model fits for Extrema and Integration models using only 6 subjects. In sum, the key qualitative data patterns of accuracy improvements with duration and broader, lower and duration-invariant low-contrast CPPs are statistically robust and provide a strong basis to reveal the fundamental principle that the Extremum-flagging and Integration models are both able to produce these key qualitative dynamics.
(2) The Extremum-flagging model is the paper's most novel contribution, yet its physiological basis is underspecified. The model posits that each decision-terminating bound-crossing triggers a stereotyped, half-sine-shaped centroparietal signal, but no neural circuit or computational mechanism is proposed for how the brain could detect the first bound-crossing event in a non-accumulating evidence stream or generate a temporally precise, fixed-amplitude signal in response. Possible connections to P3b theories of context updating and response facilitation are acknowledged, but these are vague functional descriptions rather than mechanistic accounts. I think the discussion should engage more directly with potential neural substrates that could generate this flagging signal, and whether these are consistent with the known generators of the CPP/P3b. Without this, the extremum-flagging model risks being viewed as a mathematical convenience rather than a biologically plausible alternative.
We thank the reviewer for this constructive comment. While the focus of this paper was indeed on simple mathematical descriptions in the spirit of classical cognitive modelling, we agree that expanding on the potential neural substrates of the Extremum-flagging model strengthens its utility as an alternative framework. We have revised the Discussion to engage with potential biological mechanisms, particularly those previously proposed to underlie the P300/P3b, such as Nieuwenhuis’ (2005) proposal that it reflects a phasic arousal response mediated by the LC/NE system that serves to activate task-relevant areas following completion of a decision. We agree that aside from this, accounts of ERP component functions over the years have often been vague and non-mechanistic, but the idea that they reflect discrete neural activations marking an internal cognitive event in a stereotyped way persists, and remains a basic assumption of several new and influential ERP signal analysis toolboxes (e.g. Ehinger 2019; Weindel 2024). If the flagging signal’s fixed amplitude seems physiologically implausible, all-or-nothing neural activation events are not generally unheard of in neurophysiology, and, again, we are taking an approach favouring parsimony in the spirit of cognitive modelling, and we found that we did not need to assume any variation in the amplitude of the flagging signal in order to capture the key decision signal dynamics alongside behavioural accuracies in this particular case.
We also discuss the study of Latimer et al. (2015), who demonstrated that discrete, step-function state transitions that on single trials may mark extrema detection events, can produce ramp-like signals when trial-averaged. While the biological plausibility of such step-function dynamics remains a subject of debate, it serves as another example of how continuous evidence integration is not the only way to reproduce the ramping neural signals traditionally observed in grand-average neural signals.
(3) The Integration model at the preferred neural weighting estimates a high-to-low contrast drift rate ratio of 8.7, whereas the empirical Mu/Beta lateralization slopes suggest a ratio of approximately 3.5. The authors attribute this discrepancy to the nonlinear contrast response function of early visual cortex and the salience of the high-contrast evidence onset, but these explanations are speculative. These outcomes are arguably the most quantitatively damaging result for the integration model, so they deserve more than a brief discussion. I would recommend that the authors (a) estimate what range of contrast response nonlinearities would be required to close this gap, (b) test whether an alternative drift rate parameterization (e.g., scaling drift rates directly by SSVEP amplitude rather than contrast) reduces the discrepancy, or (c) be more explicit about treating this as a point against the Integration account.
We agree that the quantitative discrepancy we demonstrated between the empirically observed buildup rate ratio in motor preparation signals (3.5) and the greater drift-rate ratio (8.7) required by the Integration model to fit the CPP waveforms is an important one that should be emphasised and discussed with greater depth and clarity. As we said, nonlinear contrast response functions and a boosting effect of the salient high-contrast step-change are two plausible ways that a drift rate might scale disproportionately more steeply with contrast, but in principle, assuming straightforward transmission of evidence accumulation to the motor level, Mu/Beta lateralization slopes should then reflect this steeper drift rate scaling, or at least approach it even when allowing for some temporal blurring. We have thus put more emphasis on the discrepancy by confirming that if we constrain the drift rates to be directly proportional to contrast, the Integration model is indeed significantly hampered in its ability to produce the much steeper CPP buildup for higher-contrast trials, much more so than the Extremum-flagging model (Figure 4 - Supplement 7). We have also applied a temporal blurring equivalent to the short-time Fourier Transform to the simulated motor preparation waveforms (convolving with a boxcar of the same duration as the Fourier window) in Figure 4N-P so that the real and simulated traces are on an equal footing in this respect. We have also revised the Discussion to elaborate on how this quantitative discrepancy represents a point against the Integration account, and possible ways it might be reconciled with an Integration account. One reason, for example, why the relative steepness of the Centroparietal ERP in the high-contrast condition so far exceeds that of Mu/Beta might be that additional processes are evoked by the very salient step-change, which may make a positive-polarity contribution to the centroparietal ERP waveform and hence cause overestimation of how early and steeply the underlying, high-contrast CPP decision signal rises. We looked into this by carefully examining time courses and topographies through the initial period of buildup, with no additional smoothing low-pass filter applied, now presented in Figure 3 - Supplementary Figure 1. While the smoothed waveforms that we show in the main paper and to which we fit models could be seen to have a brief inflection during the main buildup for the high-contrast condition, removing the smoothing shows that this arises not from random noise but from a distinct bimodal morphology, with a distinct early peak and lull during the buildup, which temporally coincides with a very strong bilateral occipital N2 (associated with a low-level evidence-onset detection or ‘target selection’ process - Loughnane et al 2016), in a way that suggests that the positive tail-end of the dipolar neural generators of the N2 may contribute to the initial part of the positive centro-parietal buildup. It is difficult to estimate the extent to which the neurally-constrained model estimate of high-contrast drift rate is inflated by this initial overlapping potential, because we can’t precisely know the ground truth of the N2 tail’s contribution, but this analysis provides a potential explanation that can be explored in future (e.g. through softening the strong-evidence onset with a ramp or use of auditory evidence). We thank the reviewer for raising this as we feel that this extra discussion positively adds to the theme of the paper to highlight methodological challenges with neurally-constrained modelling. In the process, we have updated the methods section to present in full detail the centroparietal electrode selection and waveform smoothing that was applied to provide the models with a relatively uninterrupted buildup signal to capture, which is important for readers to appraise the potential impact of this overlapping potential.
(4) The sensitivity analysis over neural constraint weightings (w = 0.1 to 1000) is thoughtful, but the paper ultimately acknowledges that the preferred weighting is w=10, chosen because it achieves "a good fit to CPP dynamics without substantively sacrificing behavioral fit" - a qualitative criterion. No principled statistical framework is used to select the optimal weighting or to compare models at a given weighting. A Bayesian model comparison could provide a more formal framework for combining behavioral and neural fit components, and would allow a clearer statement about the relative posterior probability of each model.
We agree with the reviewer that theoretically, the Bayesian framework provides a principled way to combine behavioural and neural evidence by weighting each source according to its statistical reliability. However, a Bayesian formulation typically quantifies reliability through across-trial variance, which applies quite differently for accuracy and EEG data. While the precision of EEG measurements can be estimated empirically (e.g., from noise characteristics), we currently lack a formal measure of uncertainty for the linking function itself, that is, the theoretical mapping between neural signatures and latent decision processes. This represents an unresolved methodological issue rather than a straightforward parameter estimation problem. Second, although hierarchical Bayesian approaches are well established for standard diffusion models, the mechanisms examined here for extremum flagging do not currently have tractable closed-form formulations suitable for Bayesian integration. Developing a dedicated hierarchical Bayesian framework for these non-standard mechanisms would require substantial methodological work and is beyond the scope of this research.
Thus, rather than imposing a single assumed reliability relationship between neural and behavioural data, we chose to perform a systematic sweep across weighting values. We view this approach as a transparent sensitivity analysis that accommodates different scientific priors regarding the relative contribution of neural versus behavioural constraints. By presenting the full range of w (including in supplemental tables and figures), readers can directly evaluate how model behaviour changes when emphasis is shifted between behavioural data and neural data, transparently revealing how the behavioural and neural signal fits can trade against one another.
Reviewer #2 (Public review):
Summary:
The manuscript by Hajimohammadi, Mohr, O'Connell and Kelly is intended to demonstrate that participants integrate evidence over time to make a decision, even in a noise-free, static decision context. This is validated by the observation that (1) participant accuracy improves with increased exposure to the stimulus; and (2) there is a correlation between participant accuracy and a neural index of evidence accumulation, as measured by centro-parietal positivity (CPP).
Strengths:
(1) Joint modelling of accuracy and CPP dynamics is a significant achievement, as behaviour alone often cannot distinguish between competing theories of decision-making. In the case of protracted sampling in particular, the absence of reaction times (RT) due to the delayed nature of the response makes this method highly appealing.
(2) The experimental manipulations and the method used to extract the different neural indices are well chosen, enabling the mapping of putative cognitive processes such as evidence accumulation and motor preparation onto the recorded EEG with clarity.
(3) The in-depth discussion of the results clearly articulates those reported by the authors and in previous works.
Weaknesses:
(1) One main issue to support the interpretation of the authors toward the need for protracted sampling is the timing of the evidence. By design, participants believe that the signal is present for 1.6 seconds (reinforced by the fact that easy trials were displayed for 1.6 seconds). However, the difference in stimuli is turned off either 1.4, 1.2, 0.8 or 0 seconds before the cue to respond. While this makes sense in the context of the authors' question, it also raises the possibility that participants will focus on the last samples before answering. Even if participants apply equal weighting, this still favours them delaying evidence accumulation until they are sufficiently certain that the evidence should be present (e.g. participants might start accumulating after the stimulus has disappeared in the 0.2 condition). I do not see an easy way to test these alternative explanations outside of running a study in which the evidence is always offset before the go cue.
This is a reasonable question about the design - if participants were under the impression that they had a whole 1.6 sec of stimulation, couldn’t they afford to wait until later into the stimulus to start sampling? However, the task was designed to be so difficult that participants would be deterred from ignoring any initial evidence, and the fixed and explicitly instructed lead-in period as well as the interleaved easy trials, would have continually reinforced their ability to time their sampling onset quite precisely. Indeed, key aspects of the data confirm they did not appreciably delay sampling. First, accuracy in even the shortest (0.2 s) condition was reliably above chance (t(15) = 2.60, p = 0.0201) and improved steadily across evidence durations (Figure 1B). This places an upper bound on the accumulation onset: participants cannot have delayed accumulation until after the evidence disappeared and still achieve above-chance performance; if they only used the ‘last samples,’ at the end of the stimulus, they would have performed at chance level for all durations except 1.6 sec. Second, we fit a model that allowed for such a delayed sampling onset, captured in the parameter ‘sampT,’ which, across the range of neural weightings (Tables S3, S5-8), consistently landed within a few tens of msec of evidence onset (often slightly before rather than delayed), and improved the overall model fit very little relative to the addition of starting point variabilities or collapsing bound. The Methods section now addresses these aspects of task design.
(2) Regarding the behavioural models, are these identifiable based on accuracy data alone? This should be addressed using a parameter recovery study, in which a set of parameters is used to generate data, and the same fitting routine used for the real data is used to estimate the parameters. This would enable us to determine what can be inferred from the model comparison presented. This is not a serious problem for the manuscript, as it specifically aims to go beyond behaviour. It is, however, worth noting that such a parameter recovery addition could be used to demonstrate the need for a joint modelling framework to answer the question of protracted sampling on delayed response times (RT).
As the reviewer notes, we did have the specific aim of going beyond behaviour, and the need to do so is demonstrated in the inability to adjudicate between the alternative models based on behaviour alone. We took this as sufficient justification without a formal parameter recovery test to assess the degree to which behaviour-only models could accurately estimate parameter values. Still, we agree that it is valuable to address parameter identifiability in some way. Since a full parameter recovery covering the full possible parameter space for each of the many models would be too great in volume to add to this paper, we can instead address identifiability somewhat indirectly through parameter estimate consistency across the 10 fits we conducted with different instantiations of noise; we now provide the standard deviations alongside the mean parameter values for the D1, D2 and B parameters of each of the behaviour-only models in Table 1 - Table Supplement 1, which indicates that the parameter estimates were reliable across 10 different instantiations.
Minor comments:
(1) I would advise authors to fix the D1 parameter and use it as a scaling parameter across all models. Currently, as I understand it, the models are scale-free, meaning the same fit is achieved by multiplying all parameters by two, for example. This makes the fit more complex (bounds on parameter values are required) and means that the models are less comparable in terms of their estimates. Perhaps I'm missing something, but I would have thought that fixing D1 (the common parameter across all models) would solve these issues.
The models are not scale-free because they are constrained relative to a fixed sampling noise parameter value of s = 0.1; All tables in the main text have now been updated to make this more immediately clear. Aside from this being standard in diffusion modelling (Ratcliff & Smith, 2004), this enabled us to replicate the observation by Stine et al., (2020) that since the non-integration models depend on the magnitude of individual evidence samples rather than an integration of many, the drift rate values must be set much higher to achieve the same choice accuracy as the integration models (Table 1).
(2) Why is the snapshot model so bad despite being a good model in Stine et al 2020? Can the authors speculate in the discussion?
We thank the reviewer for querying this. We had originally thought that the poor performance of the snapshot model made sense because the continued presentation of zero contrast difference for short-evidence trials renders it a bad strategy. Because our main purpose was to briefly substantiate the principle that accuracies alone are an insufficient basis for model comparison and move on to the main goal of jointly modelling accuracies and CPP dynamics, we did not take the same level of care to ensure we attained the very best fit of the behaviour-only models, as we did for the neurally-constrained models. In the neurally-constrained modeling, we took care to check for every parameter whether the range of allowed values (Table S4) was narrow enough to avoid the optimisation algorithm getting lost in untenable parts of parameter space, yet wide enough to include the optimum point, and wherever we saw parameter values landing at or near the edge of the allowed range we expanded that range and re-ran the model fit. Applying these same checks to the behaviour-only fitting, we found that the SnapShot model needed a wider range on drift rate and when we applied this, the fit was much more competitive, in line with Stine et al., (2020), though it remained the worst-fitting model among all two-drift-rate behaviour-only models (see updated Table 1). We similarly conducted these checks across all behaviour-only models and re-ran them. The extrema detection model with last-sample default when no bound is hit also improved its fit, though again it did not fit better than the version with guess default. Thus, the point we were making with this section, that behaviour alone can be captured competitively by a range of integration and non-integration models, is bolstered by the updated model fits. Since the last-sample default was competitive in the behaviour-only fits, we also ran a version of the Extremum-flagging model jointly fit to accuracies and CPP dynamics with a last-sample rather than random guess default when a bound was not reached, and show in new Figure 4 - Figure Supplement 8 that the conclusions are the same. Again, thank you for prompting us to look back at those fits.
(3) The meaning of the flag width is unclear. Figure 4 provides the reader with an intuitive understanding of the model that the authors have in mind. However, the tables in the appendices report values between 0.2 and 0.9. I understand that these values represent the width of the half-sine in seconds. This suggests that the actual estimated values for these flag events are much broader than those displayed in Figure 4. While this is probably fine for most models, it can be problematic for the extremum-flagging model, as it means that the rise to the peak takes between 0.1 and 0.45 seconds. While strictly speaking, this is still a 'flag' model, such a slow rise to the peak, given the usual expectation of evidence accumulation, would place this model closer to a smooth integration model than to a boundary-crossing flagging mechanism.
We thank the reviewer for raising this about the flag width parameter. In so doing, they enabled us to catch that our schematic depiction of the model in Figure 4 was misleading, and have now revised it to make clear that the flag signal is a post-decision one triggered by the bound crossing, and we have updated explanations accordingly (in ‘Neurally-constrained models’ and Discussion). The reviewer is correct that the reported values in the supplemental materials (approximately 0.2–0.9 s) correspond to the width of the half-sine kernel used to model the post-decision flag event. However, the flag signal is stereotyped, evidence-independent, and is triggered once the decision threshold has already been crossed, so it does not share the key characteristics of evidence integration, regardless of how wide the model estimates it. In the extremum-flagging model, the boundary crossing remains a discrete event. The width parameter instead captures the temporal extent of the neural process that follows this commitment event. Such a post-decision neural process unfolding over several hundred milliseconds is in line with some classic theories of the centroparietal P300/P3b component, and we now expand our discussion point on this to address proposed neural substrates (e.g. Nieuwenhuis et al’s (2005) implication of a phasic noradrenaline system response).
(4) In the modelling section, it is not clear overall (i.e. for G<sup>2</sup> and R<sup>2</sup>) how the participant dimension is taken into account. Are these individually fitted models, and if so, how are the secondary statistics generated from the individual estimates? Or were these fitted over all participants?
All models were fitted to the grand-average neural and behavioural data across participants, rather than to individual participant data. We chose this approach as the CPP signal at the individual level is highly noisy, which can introduce substantial instability and noise into the model fitting procedure. We have revised the Modelling section to explicitly state that the reported G<sup>2</sup> and R<sup>2</sup> values are derived from models fitted to the grand-average data, and in the revised discussion acknowledged this as a limitation of the current modelling framework.
(5) On page 7, in the last sentence of the first paragraph of the section titled 'Decision-Related Neural Signals', the authors state that 'this stable contrast-difference encoding suggests that a constant (i.e. non-adapting) drift rate is a reasonable simplifying model assumption'. However, I am not sure how this is true given that SSVEP quantifies encoding, yet the drift rate can vary due to non-sensory aspects (e.g. attention).
The reviewer makes a good point - even if sensory encoding is stable, non-sensory factors like attention could cause dynamic changes in the effective drift rate independently of the sensory representation itself. However, our point in that section, which we have revised to put more clearly, was to test for one particular well-known time-varying effect that could impact drift rate, namely sensory adaptation, a well-established phenomenon behaviorally and at the level of sensory neuronal responses, where prolonged stimulation produces reductions over time in sensory neural activity. If strong adaptation were present in the sensory evidence representation indexed by the SSVEP, we would expect corresponding temporal changes in the signal. The absence of such changes lends support to the simplifying assumption (as in most accumulation models) that the drift rate is approximately stationary over time, even if we cannot be sure there isn’t a time-varying effect downstream.
(6) The mu/beta lateralisation does indeed favor the integration model more, but in terms of boundary estimation and starting-point analyses, both models are pretty far apart. Providing an interpretation of this observation, e.g. regarding alternative linking functions for mu/beta, would add to the manuscript.
In response to this comment, we revised the manuscript in the Discussion to say that in the current analyses, we implicitly assume an approximately linear mapping between Mu/Beta amplitude and decision units. However, the true relationship may instead reflect another monotonic transformation (e.g., involving power rather than amplitude, logarithmic scaling such as dB units, or a nonlinear saturating function). This uncertainty could affect the apparent correspondence between the neural signal and the model-derived estimates of boundary position or urgency dynamics. While our analyses support a close relationship between Mu/Beta lateralisation and the evolving decision process, the precise quantitative mapping remains uncertain. One possibility is that urgency itself evolves nonlinearly (e.g., decelerating over time), even if the measured neural trajectory appears approximately linear under the current transformation assumptions.
Reviewer #3 (Public review):
Summary:
The authors aim to compare proposal models of perceptual decision making using a joint modeling approach, where they fit models to both behavioral outcomes as well as CPP. Most notably, they compare a standard evidence accumulation model with models that track the evidence without integrating it over time (extrema detection). The authors report that the joint CPP-behavioral data do not discriminate between two of their proposals.
Strengths:
This is an interesting finding that reinforces the idea that what we believe to see based on aggregation over trials may not be what happens on every single trial. The models are creative, and the simulations are convincing, relating the models to multiple neural markers of decision formation. These include the CPP but also mu/beta power spectra.
Weaknesses:
The paper makes some strong points, and the work seems generally well-executed. The weaknesses that I identified are twofold:
(1) Embedding in the literature/exposition of the main argument.
The focus in the introduction is on the noise-free nature of the stimulus and the prolonged presentation time. However, after reading the paper, I felt these were mostly experimental design choices that enable comparison of the different models using the CPP. Perhaps my misreading of the goals of the paper stems from two other observations:
(a) The fact that the stimulus is noise-free does not entail that perception is noise-free. Thus, the argument that using a noise-free stimulus precludes the necessity of temporal integration seems not completely valid. Of course, one could argue that noise is limited in this case, but that makes a noise-free stimulus more of a design choice.
(b) The focus on prolonged stimulus presentation, but at the same time the contrast with expanded judgement, did not make sense to me. Perhaps, as a non-native speaker, I am misreading the subtle difference between "protracted sampling" and "longer sampling", but again, the longer duration seems mostly a design choice.
We thank the reviewer for this impression, which has helped us revise the introduction to more clearly motivate the paradigm as an interesting case for close examination. The primary driver of our choice of stimulus and task parameters was not to enable model comparison using the CPP; it was to examine a decision scenario that exists in everyday life but that has not been examined in terms of underlying decision mechanisms because it offers only sparse behavioural data - the scenario in which plainly visible objects (without noise or stochasticity, as in daylight conditions) need to be examined for a subtle feature difference to guide a later action. The reviewer echoes our point in the Intro, that despite the absence of physical noise in the stimulus, perceptual processing itself is not noise-free. Therefore, temporal integration is certainly not precluded, but its benefit is minimised and less obvious to the decision maker. Given the examples we raise where integration was found to not be employed to its optimal extent (e.g. bound setting foregoing accuracy improvements with duration), and the various theoretical accounts citing energy costs associated with integration and the fleeting nature of many natural environments where prolonged deliberation about a static stimulus is not the norm (e.g. Uchida et al 2006), it is quite hard to guess a priori whether humans will engage in protracted sampling and integration in this case, in practice, even if it is optimal under basic assumptions. As we make clear in our revised Intro, this theoretical interest in the uncertain case of long, noise-free stimuli where perfect, unbounded integration may be optimal but seems doubtful given extant empirical findings, is coupled with a methodological interest in the extent to which neural signatures of decision formation can ‘come to the rescue’ and provide grounds for reliable adjudication between competing mathematical models, when behavioural data fall short.
More could be said about the optimality of the extrema detection methods. In particular, decades of work (centuries?) have shown that evidence integration is an optimal decision-making procedure: For example, the Sequential Probability Ratio Test is Bayes-optimal wrt mean RT (Wald, 1946); evidence accumulation together with collapsing threshold serves to maximize rewards in repeated choices (e.g., Bogacz et al., PsychRev, 2006; Boehm et al. APP, 2020). Given all this work, why would the brain have evolved to adopt a different mechanism? I realize that the paper is not about optimal decision making, but some discussion of this point seems warranted.
We had a similar impression initially when reading Stine et al., (2020) where extrema-detection was pitted against integration - is extrema detection so suboptimal that it is too implausible to even consider? Ditterich (2006) argued that signal-to-noise ratio would have to be implausibly high for extrema-detection to produce the behaviour observed on typical decision tasks. However, the fact is, we do not know the effective signal-to-noise ratio, nor can we precisely quantify the costs associated with prolonged evidence accumulation, such as attentional or energetic costs (Drugowitsch et. al., 2012). Even if the extrema detection strategy appears implausibly suboptimal, it is an important principle to demonstrate how not only behavioural but also neural decision signal dynamics can be so nicely consistent with integration yet technically can be quantitatively captured with non-integration mechanisms.
(2) Modeling choices.
The authors introduce a parameter, sampT, that represents uncertainty in the sampling onset time. It was not clear to me whether this parameter represented an offset of all trials, or a distribution (probably the latter). I wonder how exactly this parameter was integrated into the models, and in particular, if and how it interacts with the starting-point parameters. My intuition is that on a single-trial, IF early sampling occurs, you can model that with either a negative sampT and z at 0, or with sampT at 0 but a shift in z. This would suggest trade-offs between these parameters, making them hard to estimate independently. Since the paper does not depend on the identification of parameter estimates, this may not be a huge problem, but nevertheless it is good to explore the consequences.
We thank the reviewer for raising an important question regarding the relationship between sampT and starting-point variability (sz). Mechanistically, early accumulation onset can indeed generate effects that resemble starting-point variability: if accumulation begins during a period containing only zero-mean noise, then by the time informative evidence appears, the decision variable will already have diffused away randomly from zero. In this sense, negative sampT can induce variability in the state of the accumulator at evidence onset. However, the two mechanisms are not mathematically equivalent. The sz parameter assumes a uniform distribution over starting points, whereas the variability induced by early accumulation onset would instead reflect the distribution resulting from integrating zero-mean Gaussian noise over variable durations. Aside from this distinction between distribution shapes, the reviewer is correct that these parameters could partially trade off with one another when sampT takes negative values. In our model, however, sampT was allowed to take either positive or negative values. Positive values delay the onset of evidence integration relative to the evidence, thereby ignoring the first samples, very different from the effect of starting point variability. Nevertheless, to the extent that they can partially trade off each other to some degree, the consequent problem this might cause to accurately estimating both parameters is part of the reason we do not fit a model that includes both simultaneously.
The way the Bounded Integration model (BIntg) is formulated seems very close to the EZ-diffusion model (Wagenmakers et al., PBR, 2007). This model states that the proportion of correct responses Pc = 1/(1+exp(-B*D/s^2), with B and D the bound and drift rate parameters, respectively. However, filling in the numbers for the high contrast condition from Table 2, and assuming that s=2 (because the model description states that dt=2, with s undefined), I get a Pc of 80% for the 1.6H condition. This seems substantially less than what Figure 2 suggests.
As we had stated in the Methods section, the model used “a standard deviation of 0.1 arbitrary units” for the evidence. We now make it more explicitly clear that this corresponds to setting within-trial noise s = 0.1 as the scaling parameter (the first paragraph of ‘Model Fits to behaviour only’ and the first paragraph of ‘Integration models’ in Methods). Replacing (s = 0.1) in the suggested calculation yields a predicted accuracy close to 1 for the high-contrast condition, consistent with both the behavioural data and the model predictions shown in Figure 2. We have now clarified this parameter explicitly in the revised Methods and updated Table 1 and Table 2 to avoid confusion.
On some occasions, it is unclear to me what modeling choices are being made:
(a) It seems as if the models are fit on accuracy data alone (before introducing the neural data). This seems suboptimal given that the authors do report differences in RT.
Because of the delayed-report feature of the task, RTs do not directly reflect decision termination time, which is the basis of the use of RT in cognitive modelling normally. Here, whether the decision process has concluded during the stimulus or not, indeterminate response-cue detection and motor execution processes intervene between the stimulus and RT, which would necessitate complicating the models with additional mechanisms, of which there are several possibilities as reflected in the response to the reviewer’s final comment about the RT effects below. This could potentially obscure the core mechanisms of decision formation during the stimulus itself, which was the focus of the study.
(b) Are the models fit on all data combined, or on the data of individual participants? Fitting individual participant data is preferred, as combined or aggregated data may be distorted by individual differences.
Because of the noise and variability of EEG data at the single-participant level, we model data averaged across participants, which we have ensured is clear in the revised paper. We provide individual accuracy trends in Figure 1, to verify that the accuracy improvements with increasing evidence duration seen on average are representative of the vast majority of individual subjects. We also added a comment on the limitation this incurs regarding individual difference analysis in the revised discussion.
(c) The authors seem to suggest that the diffusion coefficient s is estimated (in the section "Integration models"). Most likely, however, this is set to a fixed value. Obviously, it matters for the model comparison using AIC whether this parameter was freely estimated or not.
As noted in our response to an earlier Comment, the diffusion coefficient was fixed at s=0.1, and to make this explicit, we have entered it for all models in the revised Table 1 and Table 2.
Not really a weakness, but I wondered about the effect of stimulus duration on RT. In particular, what hypothesis (or post hoc explanation) do the authors have for these RT effects? I could think of at least three hypotheses that are consistent with the behavioral data:
(a) H1: The shorter the evidence duration, the more likely participants are to require a double-check before response execution, reflecting their uncertainty about their decision.
(b) H2: There is a collapsing threshold that initiates at stimulus offset, leading to quicker responses on trials where there is more evidence.
(c) H3: motor preparation is correlated with the evidence signal, which leads to faster responses on trials with more evidence.
We thank the reviewer for these hypotheses. We agree that the RT effects admit multiple possible interpretations, and while we are cautious not to overinterpret them mechanistically in the paper given our focus on the decision process during the stimulus preceding these response-cue-triggered responses, we do take them to signify that the decision process has not always fully completed and been transformed to a finalised action plan by the time of response cue (start of Results section). To consider these interesting possibilities further:
We agree that the longer RTs for shorter duration, more uncertain trials could reflect a “double-checking” process (H1), but it could alternatively reflect the fact that if a bound has already been reached during the stimulus, this commitment can be translated to fully-selected action plan that only needs triggering, whereas if a bound has not been reached by stimulus offset, which would occur more often for shorter evidence-duration trials, more of the motor action-selection process would yet need to be completed to initiate the action, causing the slight delay in RT. In other words, on longer-duration trials, the accumulated evidence is more likely to have already reached the bound before the response cue, allowing participants to both commit to a choice and prepare the associated motor response in advance. Therefore, RTs would be shorter as the remaining processes after the cue primarily involve cue detection and motor execution.
This interpretation is broadly compatible with the reviewer’s H3 account, in the sense that motor preparation may track the evolving decision variable/evidence state. It is also possible that collapsing bounds are set on a post-stimulus, cue-evoked process (H2), which is not mutually exclusive with the above possibilities. It would be hard to determine whether such a process is primarily a response cue-detection decision process that is modulated by uncertainty state at stimulus offset, or a cue-triggered “double-check” process that perhaps operates on the iconic memory of the evidence, and our paradigm does not allow these alternatives to be cleanly dissociated.
Recommendations for the authors:
Reviewing Editor Comments:
As you can see, the reviewers are positive about the work and highlight several strengths, while at the same time offering recommendations for improvement. Once these points are satisfactorily addressed, this may also lead to a revision of the eLife assessment below.
Reviewer #1 (Recommendations for the authors):
As outlined in my public review, my most important recommendations relate to sample size and possible neural bases of the extremum-flagging model.
Minor Comments:
(1) p. 4 rmANOVA statistic is listed as 12.35.31.
This has been corrected, with thanks for spotting it.
(2) The stable d-SSVEP amplitude during the evidence period is used to justify a constant (non-adapting) drift rate assumption. This is a reasonable inference, but the SSVEP reflects early sensory encoding rather than the decision variable per se. Neural adaptation or gain changes at later processing stages could still produce a non-constant effective drift rate even with a stable sensory representation. This inference should be qualified.
This is true. We have clarified that these checks for one important potential source of a time-varying drift rate, namely adaptation at the level of early sensory representation, but admit that other effects may happen downstream.
(3) The Methods describe a leaky accumulation extension; the Results note leak ≈ 0.0002 at w=10, which is effectively zero. This is a positive result (evidence against leaky integration) that should be stated more explicitly in the Results or Discussion rather than appearing only in supplementary tables.
We thank the reviewer for highlighting this. We have pointed to this result now in the first paragraph of the Discussion.
Reviewer #2 (Recommendations for the authors):
(1) The panels in Figure 4 H, I, J are not discussed in the Neurally-constrained models Section, while I believe they are probably more informative than Figure 4E alone.
The manuscript has been revised to explain Figure 4H,I, J fully under ‘Neurally-constrained models’.
(2) In the method section, we don't know how many trials were rejected based on the chosen threshold.
The Method section is updated with the rejected trials after preprocessing. It now reads, “After preprocessing, 12017 trials remained across all conditions, with an average rejection rate of 15% (± 12.9%) across participants.”
(3) Figure 1: For b and c, data are mean {plus minus} s.e.m. after between-participant variance was factored out. -> reference or detailed method.
We have now explained in the caption that this is done by subtracting the overall mean of each individual from their data and adding back the grand mean, retaining the between-condition differences - that is, we remove the component of variance that repeated-measures tests ignore.
(4) Figure 2: Shouldn't the snapshot model only feature one sample, as in Stine et al. 2020? This figure and others would also benefit from a better resolution.
The reviewer is correct regarding the schematic of the snapshot model presented in Stine et al., (2020). We used small, light orange dots to represent the evidence samples presumably being encoded and a larger orange dot to indicate the single randomly chosen sample used as evidence for the decision. In our revised figure, we have increased the visual distinction between the evidence dots and the selected dot, and pointed this out in the caption. We have also improved the figure resolutions.
(5) Typo:
- Semi-saturation Table S4.
- pi missing in the text of the G^2 equation.
Thank you for catching these.
Reviewer #3 (Recommendations for the authors):
Small, random points:
(1) How was it ensured that participants indeed did not detect the change in contrast throughout the 1.6s interval? In previous work (Winkel et al., PBR, 2014) we did something similar in a random-dot motion task, but observed that participants always observed the change, if we did not slowly change the coherence of the stimulus (unfortunately, Winkel et al., 2014 is not explicit about the exact parameters of the change, but Figure 2 suggests that the change in coherence lasted 50ms, independent of stimulus strength).
It is true that abrupt changes in stimulus strength are more salient and detectable than a ramped change. The experimenters tested the stimulus during the task design subjectively, to satisfy themselves that they could not tell when the contrast stepped back to baseline, but this was not verified systematically with psychometrics, nor can we be sure that a more sensitive observer couldn’t sometimes detect the change. However, based on the task design, stimulus properties, and both behavioural and neural data, we are confident that participants are very unlikely to have been sufficiently confident in detecting the step-back in contrast to ceasing their contrast-comparison decision process at that point:
(1) Participants were naive to the underlying manipulation. They were informed that trials would naturally vary in difficulty. It was normal for them to perceive some trials as harder than others without suspecting a mid-trial structural change.
(2) The contrast difference in the hard condition was very subtle, and though it is possible that the step-down in contrast could be detected with above-chance accuracy if instructed to do so, given there was no instruction on whether and when the step-down would happen, it is very unlikely they could be detected with sufficient confidence to be certain there is no remaining evidence in the stimulus. Furthermore, the rapid, flickering nature of the stimulus would have helped to mask the transition point, in comparison with a sudden change in a continuously-playing random-dot motion stimulus (as in Winkel et al., 2014).
(3) Our behavioural post-cue RT data imply the participants did not cease decision formation at evidence offset. If they had, then they would have been afforded the most time to prepare their chosen action in advance of the response cue in the case of the earlier evidence offset (i.e. shorter durations), yet these were the conditions with the longest, not the shortest post-cue RTs.
(4) The low-contrast CPP traces remained elevated for the full 1600 ms interval, unperturbed by the evidence offsets. If participants were explicitly detecting a sudden change in contrast, we would expect to see a transient evoked response locked to that change, marking that detection. Instead, the sustained elevation of the CPP is characteristic of a continuation of the decision process, uninterrupted, through the subtle offsets.
(2) Could you include the regression coefficients of the statistical modeling of the behavioral data?
The regression coefficient is now added to the second paragraph of the Results section.
(3) I felt Figure 5B was a bit confusing: Are the dashed lines here the ipsilateral sides or the non-linear bounds? This was confusing because "data" only has a solid line in the legend.
We agree that the subtle nonlinearity, which appears visually close to the linear model, may have caused this confusion. To improve clarity, we have added arrows to Figure 5B to explicitly indicate which legend refers to which panel, and distinguish the dashed nonlinear bounds from the other traces.
(4) "amplitude variations [...] to be used as an independent evaluation of model fit": Could you refer to where these model predictions are presented? I think this is in Figure 4 - Sup 5?
We thank the reviewer for raising this. The model-predicted waveforms showing amplitude variations across durations for all neural weightings are presented in Figure 4 - Supplementary Figures 2 and 5. We have clarified this in the revised manuscript.
References
Ditterich, J. (2006). Evidence for time‐variant decision making. European Journal of Neuroscience, 24(12), 3628–3641. https://doi.org/10.1111/j.1460-9568.2006.05221.x
Drugowitsch, J., Moreno-Bote, R., Churchland, A. K., Shadlen, M. N., & Pouget, A. (2012). The cost of accumulating evidence in perceptual decision making. Journal of Neuroscience, 32(11), 3612–3628.
Ehinger, B. V., & Dimigen, O. (2019). Unfold: An integrated toolbox for overlap correction, non-linear modeling, and regression-based EEG analysis. PeerJ, 7, e7838.
Latimer, K. W., Yates, J. L., Meister, M. L. R., Huk, A. C., & Pillow, J. W. (2015). Single-trial spike trains in parietal cortex reveal discrete steps during decision-making. Science, 349(6244), 184–187. https://doi.org/10.1126/science.aaa4056
Loughnane, G. M., Newman, D. P., Bellgrove, M. A., Lalor, E. C., Kelly, S. P., & O’Connell, R. G. (2016). Target selection signals influence perceptual decisions by modulating the onset and rate of evidence accumulation. Current Biology, 26(4), 496–502.
Nieuwenhuis, S., Aston-Jones, G., & Cohen, J. D. (2005). Decision making, the P3, and the locus coeruleus–norepinephrine system. Psychological Bulletin, 131(4), 510.
Stine, G. M., Zylberberg, A., Ditterich, J., & Shadlen, M. N. (2020). Differentiating between integration and non-integration strategies in perceptual decision making. Elife, 9, e55365.
Uchida, N., Kepecs, A., & Mainen, Z. F. (2006). Seeing at a glance, smelling in a whiff: Rapid forms of perceptual decision making. Nature Reviews Neuroscience, 7(6), 485–491.
Weindel, G., van Maanen, L., & Borst, J. P. (2024). Trial-by-trial detection of cognitive events in neural time-series. Imaging Neuroscience, 2, imag–2.
Winkel, J., Keuken, M. C., Van Maanen, L., Wagenmakers, E.-J., & Forstmann, B. U. (2014). Early evidence affects later decisions: Why evidence accumulation is required to explain response time data. Psychonomic Bulletin & Review, 21(3), 777–784. https://doi.org/10.3758/s13423-013-0551-8