Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorMichael FrankBrown University, Providence, United States of America
- Senior EditorMichael FrankBrown University, Providence, United States of America
Reviewer #1 (Public review):
This manuscript examines EEG activity during face-to-face Diapix conversations and asks whether neural activity preceding a participant's next speaking turn predicts the latency and duration of that turn. The authors report sustained ERP and alpha/beta effects, particularly associated with upcoming response duration, together with temporal-generalization decoding beginning approximately 1-1.3 s before speech onset. They interpret these results as evidence for two stages of conversational speech planning: an early process that activates and maintains the forthcoming response, and a later process related to commitment and motor execution. The naturalistic dyadic EEG setting, relatively large number of conversational turns, trial-level analyses, and convergence across ERP, oscillatory, and multivariate approaches are important strengths.
I nevertheless have substantial concerns about the interpretation. Most importantly, the EEG analyzed during this period is recorded while participants are processing their partner's speech. Therefore, neural activity necessarily includes auditory, linguistic, semantic, pragmatic, and turn-boundary processing, in addition to any potential response-planning activity. Because properties of the incoming utterance are themselves strongly related to what the listener subsequently says, predicting subsequent response duration or latency does not by itself isolate production planning. The authors acknowledge some of these limitations, but the current title, theoretical framing, Figure 8, and conclusions remain considerably stronger than the evidence allows. Indeed, the Discussion itself acknowledges that the signals cannot be determined to reflect linguistic content, general planning demands, or motor preparation.
I therefore think the study could become valuable, but substantial additional analyses and conceptual reframing are required.
Major concerns:
(1) The central problem is dissociating speech planning from listening/comprehension
This is my main concern with the manuscript and also with the claims in the Discussion. The authors repeatedly infer that because EEG during the partner's turn predicts the duration or latency of the participant's subsequent response, this activity reflects speech planning. For example, the manuscript states that examining the period before self-speech onset "isolat[es] preparatory neural activity before articulation." I do not think this inference is justified.
During precisely this interval, the participant is also listening to and comprehending the partner. The incoming utterance is not independent of the forthcoming response; it is what determines the response. A complex statement, several questions, a request for clarification, or an information-rich description will produce different auditory/linguistic/comprehension activity and may also systematically change the length and timing of the subsequent response. Consequently, both partner speech → listening/comprehension EEG and partner speech → subsequent response properties can generate an EEG-response relationship without requiring that the measured EEG activity itself reflect production planning.
The authors control partner-turn duration, which is useful, and they show that self-duration, self-latency, and partner duration are behaviorally related. However, partner duration is only a very coarse property of the input and cannot control for what participants are actually processing. The analysis should consider, at minimum, partner word count/speech rate, acoustic properties, information density or lexical/syntactic complexity, dialogue act, question versus statement, and preferably semantic/discourse content. For example, the number of questions or propositions in the partner's turn may directly determine both comprehension demands and how much the listener subsequently needs to say.
Given that transcripts and word-level timing are already available, the authors should be able to substantially improve this analysis. One particularly useful approach would be to construct an encoding model of the EEG driven by acoustic and linguistic features of the partner's speech and then test whether the residual neural activity continues to predict subsequent response properties. At minimum, richer input-related covariates are necessary.
Without such analyses, I think the main conclusion should be reframed from "neural signatures of speech planning" to something closer to "pre-response neural activity during conversational listening predicts properties of the subsequent turn." The latter is supported by the data; the former requires stronger dissociation.
(2) It is not clear that the analyzed period is actually a "listening interval".
Relatedly, the definition of the analyzed interval needs careful reconsideration. Epochs are aligned to the participant's own speech onset and extend back 2.2 s, whereas a turn transition is accepted whenever self-onset falls between −1 and +1 s relative to partner offset.
Therefore, the interval from −2 to 0 s relative to self-onset is not necessarily an interval during which the partner is continuously speaking. For positive response latencies, part of this interval contains silence following partner offset. If the preceding partner IPU is relatively short, earlier portions of the epoch could potentially contain a different conversational state altogether. Yet the manuscript describes the whole pre-onset period as being "while they are listening to their partner."
This matters particularly for the comparison between the early and late periods. Different response-latency conditions will necessarily place the partner offset at systematically different positions within an epoch aligned to self-onset. The authors already demonstrate this problem for the anterior latency ERP, where its inflection tracks the distribution of partner offsets. This is an important observation, but I do not think the problem is restricted to that single component.
The authors should quantify, separately for TW1 and TW2, how much time is occupied by: active partner speech, silence after partner offset, overlapping speech, and potentially other conversational states.
I would strongly recommend repeating the critical analyses for periods/trials during which the partner is actually speaking, and/or explicitly modeling partner-speech presence and the temporal distance to partner offset. Analyses aligned with both partner offset and self-onset would also help separate responses to the conversational boundary from activity specifically preceding self-production.
(3) The Early-versus-Late Planning hypotheses are not sufficiently differentiated
I was confused by the theoretical contrast (e.g., Lines 53-62). The Early Planning account is described as proposing that response preparation begins once sufficient information is available, whereas the Late Planning account explicitly allows "some aspects of conceptual preparation" during listening but proposes that full articulatory planning is delayed until near the turn boundary.
Under these definitions, both hypotheses allow for early planning. They differ mainly in the level of production planning that occurs early, that is, conceptual/content preparation versus articulatory/motor preparation. This is quite different from a simple early-versus-late timing hypothesis.
This is problematic because the current EEG measures do not determine whether the early activity reflects conceptual planning, lexical formulation, articulatory planning, comprehension, or general cognitive demands. Indeed, the authors explicitly acknowledge this later. Consequently, observing EEG-behavior relationships 1 s before speech cannot adjudicate between the two hypotheses as currently formulated.
The theoretical framework should therefore distinguish explicitly among at least: conceptual/message planning → linguistic formulation → articulatory/motor preparation → initiation.
If the "Late Planning" hypothesis specifically concerns late articulatory planning, then the authors need evidence that distinguishes articulatory planning from earlier conceptual processes. Otherwise, I would describe the findings more conservatively as showing early behavior-related neural activity followed by stronger onset-proximal sensorimotor activity, rather than claiming that Early and Late Planning accounts have been reconciled.
(4) Response duration is not equivalent to the amount or content of a response that has already been planned
The strongest empirical finding concerns upcoming speaking duration. I agree that this is interesting. However, throughout the Discussion, duration is gradually interpreted as the "extent," "scale," or even the amount of content already prepared. For example, the authors argue that longer responses involve maintaining a larger set of propositions, lexical items, or planned sequences. Figure 8 goes even further and explicitly contrasts "High Content" and "Low Content," although no measure of speech content has been analyzed.
I do not think response duration alone supports this inference.
A 2-s utterance is not necessarily twice as much planned content as a 1-s utterance. Duration can vary because of speech rate, hesitations, lexical retrieval, syntactic formulation, repetitions, interactive feedback, or decisions made after speech has already begun. An eventual IPU duration is therefore not necessarily specified before articulation.
This issue is especially important because the manuscript explicitly acknowledges that the analyses cannot establish whether the signals encode lexical or semantic content. It is then inconsistent to conclude later that early activity supports "what and how much to say."
Since transcripts are available, the authors could directly test several alternatives: word count, content-word count, number of propositions, speech rate, syntactic complexity, semantic similarity between partner and self turns, and possibly dialogue-act categories. If the early signal predicts eventual linguistic content/amount rather than merely duration, the claim of maintained response specification would become considerably stronger. If not, the interpretation should remain at the level of response duration.
(5) The ROI/time-window linear mixed-model (LMM) analysis appears circular
I have a methodological concern about the relationship between the cluster analyses and subsequent LMMs. The authors first identify electrodes and time ranges showing significant differences between long versus short responses or fast versus slow responses. They then use those same data-driven clusters to define anterior/posterior ROIs and TW1/TW2 and extract EEG values from the same trials, after which the extracted EEG measures are tested as predictors of the same behavioral variables.
This creates a potentially serious double-dipping/selection bias problem. The continuous LMM outcome is not statistically independent of the median-split outcome used to select the neural features. FDR correction at the LMM stage does not correct for this feature-selection procedure.
I think the confirmatory analyses need independent feature definition, for example, predefined electrodes/time windows from previous literature, split-half analysis, leave-one-participant-out feature selection, or another cross-validated procedure in which the data used to identify the ROI/window are independent from those used to estimate and test the EEG-behavior relationship.
This is particularly important because the LMM results are used as evidence that the effect survives behavioral covariates. That conclusion should come from statistically independent tests.
(6) Several statistical and validation issues need additional attention
There are several related concerns here. First, the trial-level LMMs include only (1 | participant), despite the fact that the observations come from dyadic interactions, multiple 4-min conversations, and temporally adjacent conversational turns. Participants within the same dyad do not generate independent observations because the speech of one participant directly determines the context for the other participant. I would expect the hierarchy to include at least a dyad/conversation block, and random slopes for critical within-participant predictors should be considered. Robustness to conversational autocorrelation should also be examined.
Second, the authors describe response duration as approximately log-normal, but the model description indicates only centering/scaling rather than transformation. The authors should report residual diagnostics and consider modeling log-duration or using an appropriate non-Gaussian mixed model.
Third, the decoding uses stratified 10-fold cross-validation, but it is unclear how conversational dependence is handled. Randomly assigning turns from the same 4-min conversation to training and testing sets may permit classifiers to exploit slow neural states, discourse context, or block-specific characteristics shared by temporally adjacent turns. A more convincing analysis would use blocked cross-validation, such as leave-one-conversation/run-out. The authors should also clarify whether decoding is performed independently within participants and whether median thresholds are participant-specific or global.
These analyses are particularly important because temporal generalisation establishes stable predictive information, but not the functional identity of that information. A stable partner-speech/context signal could also produce a broad temporal-generalisation pattern.
(7) ERP baseline correction and potential speech/muscle contamination require clarification
The ERP preprocessing raises another concern. ERPs were baseline corrected using the entire −2200 to +500 ms epoch. This interval includes the first 500 ms after self-speech onset, the period most vulnerable to speech-related muscular activity. Because the main grouping variable is upcoming response duration, post-onset activity may itself systematically differ between long and short responses. Subtracting a baseline that includes this activity could introduce condition-dependent shifts in the pre-speech ERP.
I therefore think the critical ERP results should be demonstrated without using post-speech activity in baseline estimation, for example, by using an appropriate pre-speech/regression-based baseline strategy, or by analyses demonstrating that the result is insensitive to baseline choice.
There is also an intriguing inconsistency concerning EMG. The participant section states that facial hair was an exclusion criterion to ensure electrode contact for EMG recordings, whereas the Limitations section suggests that combining EEG with EMG would be useful in future work. If EMG was in fact recorded, it should be described and, ideally, analyzed to establish when peripheral speech preparation begins and whether the onset-proximal alpha/beta results remain significant after controlling for muscle activity. This would be particularly valuable for the proposed late "motor commitment" interpretation.
Reviewer #2 (Public review):
This is an excellent paper, well researched, correctly referenced and very well presented. The main addition to the existing literature, apart from supporting earlier results (see also Roberts et al. 2015), is the ability to detect the length of the upcoming conversational contribution from the neural signal. This is interesting because it shows very early and sustained planning of response during listening to the incoming turn from the interlocutor. Why is that interesting? Because this simultaneous listening and response planning is utilising the very same mental machinery - a kind of dual-tasking that humans generally find very hard, given extreme capacity constraints on short-term and working memory. This then talks to the issue of whether speech production and comprehension are more independent than current theory holds. Clearly, human language is one of the elite skills of the species, but this dual-task ability in just this domain may be a specialization evolved over deep evolutionary time. I would like to see this background sketched more clearly in the introduction - it would make the paper of more interest to the general reader.
The other main contribution is methodological, namely the ability to extract meaningful EEG out of unconstrained verbal interaction, using, e.g. machine learning to filter out movement artefacts. Earlier work was held back by scepticism that this was possible, but this paper points clearly to one way to do this. A further contribution is the suggestion that Beta modulation is associated with the duration of the turn under planning, while Alpha modulation was in addition associated with preparation for execution. This needs to be tested in further studies, but could be a valuable addition to the analytical armoury.
In the final discussion, a conceptual model is outlined in Figure 8. This looks very similar to that given in Levinson & Torreira 2015 and Levinson 2016 - at least if there is a difference, it would be good to see it clarified. Note that in that model we hypothesized that preparation of the upcoming turn was continued right up to completion and then held if necessary for ending cues from the interlocutor. Note that quite late delivery preparations have been noted both in Torreira et al. 2015 (breathing signal) and in Bögels & Levinson 2023; in the latter, there is evidence of tongue preparation sometimes occurring not tied time-wise to delivery (suggesting inhibition till incoming turn is ending). So our model already had an early and late component - the news here is that the early component can be robustly discerned in naturalistic verbal interaction.
Also, in the discussion, the Pickering & Garrod IAM model is favourably discussed, even though that model imagines the production machinery is used to predict the unfolding comprehension of the incoming turn: in that case, the production machinery would have to be doing two things at once - doing comprehension and planning one's own turn. In a way, the main lesson of the findings of the submission is precisely that this is implausible.
In general, I have no hesitation recommending publication of an important contribution to the neuroscience of human performance in ecologically more valid settings.
References:
Bögels S, Levinson SC. Ultrasound measurements of interactive turn-taking in question-answer sequences: Articulatory preparation is delayed but not tied to the response. PLoS One. 2023 Jul 5;18(7):e0276470. doi: 10.1371/journal.pone.0276470. PMID: 37405982; PMCID: PMC10321606.
Griffin, Z.M. and Bock, K. (2000) What the eyes say about speaking. Psychol. Sci. 4, 274-279.
Levinson SC. Turn-taking in Human Communication--Origins and Implications for Language Processing. Trends Cogn Sci. 2016 Jan;20(1):6-14. doi: 10.1016/j.tics.2015.10.010. Epub 2015 Dec 1. PMID: 26651245.
Levinson SC, Torreira F. Timing in turn-taking and its implications for processing models of language. Front Psychol. 2015 Jun 12;6:731. doi: 10.3389/fpsyg.2015.00731. PMID: 26124727; PMCID: PMC4464110.
Roberts SG, Torreira F, Levinson SC. The effects of processing and sequence organization on the timing of turn taking: a corpus study. Front Psychol. 2015 May 13;6:509. doi: 10.3389/fpsyg.2015.00509. PMID: 26029125; PMCID: PMC4429583.