Listening to the room: disrupting activity of dorsolateral prefrontal cortex impairs learning of room acoustics in human listeners

  1. Department of Linguistics, The Australian Hearing Hub, Macquarie University, Sydney, Australia
  2. National Acoustic Laboratories, Macquarie Park, Australia
  3. Psychological and Brain Sciences Department, The University of Iowa, Iowa City, United States
  4. School of Science, Auckland University of Technology, Auckland, New Zealand

Peer review process

Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Björn Herrmann
    Baycrest Hospital, Toronto, Canada
  • Senior Editor
    Andrew King
    University of Oxford, Oxford, United Kingdom

Reviewer #3 (Public review):

Summary:

This manuscript presents a well-designed study examining human adaptation to room acoustics, building on prior work. The psychophysical results are convincing and add meaningful knowledge to our understanding of reverberation learning. The transcranial magnetic stimulation (TMS) component shows a role of prefrontal cortex in this listening task, targeting dorsolateral prefrontal cortex (dlPFC). Cautious interpretation of the TMS results is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, and the limited ability of TMS to precisely target dlPFC in individual subjects. A surprising and interesting finding in the study is that listeners performed the speech recognition task more poorly in anechoic conditions than in those with naturalistic levels of reverberation. This is likely due to contributions of spatial release from masking provided by reverberation acoustics, which may counteract the detrimental effects of reverberation on speech perception itself. Overall, the experiments are well performed and clearly presented, improving our understanding of how the brain copes with reverberant environments.

Strengths:

(1) Well designed acoustical stimuli and psychophysical task.

(2) Comparisons across room combinations is well conducted.

(3) Virtual acoustic environment is impressive and applied well here.

(4) Timely study with interesting behavioural results.

(5) Causal evidence of a role for dlPFC in reverberation learning.

Weaknesses:

(1) Poorer performance in anechoic environments than some reverberant environments suggests and interplay of spatial release from masking and speech intelligibility that are not fully unpicked here. This could be controlled in future experiments, for example by comparing monaural and binaural listening conditions.

(2) Lack of evidence for targeting TMS to dlPFC in individual participants. This is simply a limitation of the technique which the reader should keep in mind.

(3) Most interesting effect of TMS is a null result compared to a weak statistical effect for "meta-adaptation"

Reviewer #4 (Public review):

The authors use a d' defined for 2-alternative forced choice experiments, but their data are 4-alternative (for color) and 8-alternative (for number) forced-choice. So, the d' is not computed correctly. For mAFC experiments, the authors should use the Hacker-Ratcliff (1979) method, also defined in chapter 10 of the Macmillan & Creelman textbook.

Normalization of the stimuli was arbitrary, and consequently the unexpected improvement in performance in reverberation compared to anechoic condition is still not explained. A natural normalization across different environments is to take the direct portion of the BRIR (and HRTF for the anechoic condition) and make sure that that is scaled identically across the different simulated rooms (with the reverberant tails scaled naturally). That corresponds to the situation when the sources are emitting the sound at the same level in each environment. The current study scaled the overall levels. As a minimum, it should be reported how this scaling boosted/attenuated the targets and maskers in each environment.

The potential that the listeners are tuning to individual voices, as opposed to rooms, has not been eliminated. The authors suggest that a lack of interaction with different voices is evidence that that is not the case. This is not correct: lack of this interaction just means that there are no differences in tuning between the voices. But that does not mean that the same amount of tuning is happening for each voice, as observed in previous studies. Unless the authors provide a follow-up data with randomly varying voices within each trial, these claims should be tuned down.

Author response:

The following is the authors’ response to the original reviews.

Public Reviews:

Reviewer #1 (Public review):

Summary:

This manuscript describes the results of an experiment that demonstrates a disruption in statistical learning of room acoustics when transcranial magnetic stimulation (TMS) is applied to the dorsolateral prefrontal cortex in human listeners. The work uses a testing paradigm designed by the Zahorik group that has shown improvement in speech understanding as a function of listening exposure time in a room, presumably through a mechanism of statistical learning. The manuscript is comprehensive and clear, with detailed figures that show key results. Overall, this work provides an explanation for the mechanisms that support such statistical learning of room acoustics and, therefore, represents a major advancement for the field.

Strengths:

The primary strength of the work is its simple and clear result, that the dorsolateral prefrontal cortex is involved in human room acoustic learning.

Weaknesses:

A potential weakness of this work is that the manuscript is quite lengthy and complex.

Reviewer #2 (Public review):

Summary:

This study investigated how listeners adapt to and utilize statistical properties of different acoustic spaces to improve speech perception. The researchers used repetitive TMS to perturb neural activity in DLPFC, inhibiting statistical learning compared to sham conditions. The authors also identified the most effective room types for the effective use of reverberations in speech in noise perception, with regular human-built environments bringing greater benefits than modified rooms with lower or higher reverberation times.

Strengths:

The introduction and discussion sections of the paper are very interesting and highlight the importance of the current study, particularly with regard to the use of ecologically valid stimuli in investigating statistical learning. However, they could be condensed into parts. TMS parameters and task conditions were well-considered and clearly explained.

Weaknesses

(1) The Results section is difficult to follow and includes a lot of detail, which could be removed. As such, it presents as confusing and speculative at times.

(2) The hypotheses for the study are not clearly stated.

(3) Multiple statistical models are implemented without correcting the alpha value. This leaves the analyses vulnerable to Type I errors.

(4) It is confusing to understand how many discrete experiments are included in the study as a whole, and how many participants are involved in each experiment.

(5) The TMS study is significantly underpowered and not robust. Sample size calculations need further explanation (effect sizes appear to be based on behavioural studies?). I would caution an exploratory presentation of these data, and calculate a posteriori the full sample size based on effect sizes observed in the TMS data.

Reviewer #3 (Public review):

Summary:

This manuscript presents a well-designed and insightful behavioural study examining human adaptation to room acoustics, building on prior work by Brandewie & Zahorik. The psychophysical results are convincing and add incremental but meaningful knowledge to our understanding of reverberation learning. However, I find the transcranial magnetic stimulation (TMS) component to be over-interpreted. The TMS protocol, while interesting, lacks sufficient anatomical specificity and mechanistic explanation to support the strong claims made regarding a unique role of the dorsolateral prefrontal cortex (dlPFC) in this learning process. More cautious interpretation is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, the imprecise targeting of dlPFC (which is not validated), and the lack of knowledge about the timescale of TMS effects in relation to the behavioural task. I recommend revising the manuscript to shift emphasis toward the stronger behavioural findings and to present a more measured and transparent discussion of the TMS results and their limitations.

Strengths:

(1) Well-designed acoustical stimuli and psychophysical task.

(2) Comparisons across room combinations are well conducted.

(3) The virtual acoustic environment is impressive and applied well here.

(4) A timely study with interesting behavioural results.

Weaknesses:

(1) Lack of hypotheses, particularly for TMS.

(2) Lack of evidence for targeting TMS in [brain] space and time.

(3) The most interesting effect of TMS is a null result compared to a weak statistical effect for "meta adaptation"

Reviewer #4 (Public review):

Summary:

Several behavioral experiments and one TMS experiment were performed to examine adaptation to room reverberation for speech intelligibility in noise. This is an important topic that has been extensively studied by several groups over the years. And the study is unique in that it examines one candidate brain area, dlPFC, potentially involved in this learning, and finds that disrupting this area by TMS results in a reduction in the learning. The behavioral conditions are in many ways similar to previous studies. However, they find results that do not match previous results (e.g., performance in anechoic condition is worse than in reverberation), making it difficult to assess the validity of the methods used. One unique aspect of the behavioral experiments is that Ambisonics was used to simulate the spaces, while headphone simulation was mostly used previously. The main behavioral experiment was performed by interleaving 3 different rooms and measuring speech intelligibility as a function of the number of words preceding the target in a given room on a given trial. The findings are that performance improves on the time scale of seconds (as the number of words preceding the target increases), but also on a much larger time scale of tens to hundreds of seconds (corresponding to multiple trials), while for some listeners it is degraded for the first couple of trials. The study also finds that the performance is best in the room that matches the T60 most commonly observed in everyday environments. These are potentially interesting results. However, there are issues with the design of the study and analysis methods that make it difficult to verify the conclusions based on the data.

Strengths:

(1) Analysis of the adaptation to reverberation on multiple time scales, for multiple reverberant and anechoic environments, and also considering contextual effects of one environment interleaved with the other two environments.

(2) TMS experiment showing reduction of some of the learning effects by temporarily disabling the dlPFC.

Weaknesses:

While the study examines the adaptation for different carrier lengths, it keeps multiple characteristics (mainly talker voice and location) fixed in addition to reverberation. Therefore, it is possible that the subjects adapt to other aspects of the stimuli, not just to reverberation. A condition in which only reverberation would switch for the target would allow the authors to separate these confounding alternatives. Now, the authors try to address the concerns by indirect evidence/analyses. However, the evidence provided does not appear sufficient.

The authors use terms that are either not defined or that seem to be defined incorrectly. The main issue then is the results, which are based on analysis of what the authors call d', Hit Rate, and Final Hit rate. First of all, they randomly switch between these measures. Second, it's not clear how they define them, given that their responses are either 4-alternative or 8-alternative forced choice. d', Hit Rate, and False Alarm Rate are defined in Signal detection theory for the detection of the presence of a target. It can be easily extended to a 2-alternative forced choice. But how does one define a Hit, and, in particular, a False Alarm, in a 4/8-alternative? The authors do not state how they did it, and without that, the computation of d' based on HR and FAR is dubious. Also, what the authors call Hit Rate, is presumably the percent correct performance (PCC), but even that is not clear. Then they use FHR and act as if this was the asymptotic value of their HR, even though in many conditions their learning has not ended, and randomly define a variable of +-10 from FHR, which must produce different results depending on whether the asymptote was reached or not. Other examples of usage of strange usage of terms: they talk about "global likelihood learning" (L426) without a definition or a reference, or about "cumulative hit rate" (L1738), where it is not clear to me what "cumulative" means there.

There are not enough acoustic details about the stimuli. The authors find that reverberant performance is overall better than anechoic in 2 rooms. This goes contrary to previous results. And the authors do not provide enough acoustic details to establish that this is not an artefact of how the stimuli were normalized (e.g., what were the total signal and noise levels at the two ears in the anechoic and reverberant conditions?).

There are some concerns about the use of statistics. For example, the authors perform two-way ANOVA (L724-728) in which one factor is room, but that factor does not have the same 3 levels across the two levels of the other factor. Also, in some comparisons, they randomly select 11 out of 22 subjects even though appropriate test correct for such imbalances without adding additional randomness of whether the 11 selected subjects happened to be the good or the bad ones.

Details of the experiments are not sufficiently described in the methods (L194-205) to be able to follow what was done. It should be stated that 1 main experiment was performed using 3 rooms, and that 3 follow-ups were done on a new set of subjects, each with the room swapped.

We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and insightful comments. We greatly appreciate the time and expertise invested in reviewing our work. The feedback has been invaluable in improving the clarity, rigor, and overall presentation of the manuscript.

In response to the reviewers’ comments, we have carefully revised the manuscript throughout. The revisions include clarification of the study hypotheses, re-analysis of the TMS data using a mixed ANOVA framework, additional methodological details regarding the TMS procedures and behavioural analyses, expanded justification of the statistical modelling approach, clarification of the room-acoustics paradigm, revision of figures and figure legends, additional discussion of study limitations, and a more balanced interpretation of the TMS findings. We have also substantially revised the Discussion section and improved the overall structure and readability of the manuscript.

Recommendations for the authors:

Reviewer #1 (Recommendations for the authors):

(1) It is understood that this topical area is necessarily detail-heavy, but if there are ways to streamline the manuscript to more quickly arrive at key results (Figure 3?), the work might have an even greater overall impact.

We appreciate the reviewer’s feedback and have carefully revised the manuscript to address all comments from all reviewers. However, we have retained the existing order of figures and results to maintain consistency and avoid extensive structural changes that could compromise the clarity and flow of the manuscript.

(2) Minor point: I believe Equation 1 should be d' = z(H) – z(F).

Thanks for noticing this, we have fixed the equation.

Reviewer #2 (Recommendations for the authors):

(1) Line 115: Towards the end of the introduction, the hypotheses for the current study remain unclear. Please explicitly outline each hypothesis before the Methods section.

We have now specified the hypotheses tested at the end of the Introduction.

(2) Line 229: Please state the minimum MEP amplitude criterion used during TMS thresholding (usually 50 µV). Please also reference the EMG hardware and software used to measure MEPs.

We are grateful to the reviewer for noticing that a few details regarding our TMS procedure were missing in the “Continuous theta-burst stimulation” section of the Methods. We have extensively revised this section to include the required details. We did not, however, use MEP amplitude as a criterion for estimating motor thresholds. For single-pulse TMS-induced motor threshold determination, we used visual observation of the first dorsal interosseous (FDI) muscle twitch (Varnava et al., 2011).

(3) Line 248: The Coordinate Response Measure corpus contains combinations of callsigns, colours, and numbers. What was the rationale in asking participants to only identify the colour and the number spoken, and not the callsign?

The rationale for reporting only Color and Number was to ensure an equal number of keyword identifications across phrase lengths—for example, CP0 and CP1 do not contain a callsign. This has now been clarified in the Procedure section of the Methods.

(4) Lines 327–334: The analyses outlined here are unclear – it appears that there are multiple different statistical tests being conducted on the same outcome variable in this study. Given that the design includes a between-subjects factor (TMS condition) and two within-subjects factors (Rooms, Carrier Length Phrase), could the analyses be simplified by employing a mixed ANOVA as opposed to separate repeated-measures and univariate ANOVAs? If this is, in fact, the analysis that was conducted, please improve the wording. Clarification/correction is also needed surrounding the use of the term “univariate” ANOVA, as this can refer to different statistical tests.

We appreciate the reviewer pointing this out. We have re-analysed the TMS data using a mixed ANOVA design and reported it as such in the Results section. Although the numerical values have changed, the significant findings and conclusions remain the same.

(5) Line 340: Whilst the authors identify that an alpha value of 0.05 and Bonferroni corrections were used for statistical inference in two-tailed t-tests, there is no indication of the alpha value used in the interpretation of the ANOVA results. Please include this before the Results section. Furthermore, given that multiple tests are being conducted in this project, the alpha value used to infer statistical significance should be corrected in accordance with the number of hypotheses, to reduce Type I error rates (e.g., 4 hypotheses would result in an alpha inference criterion of α = 0.0125).

We appreciate the reviewer’s observation. The corrected alpha value was not explicitly reported because all statistical analyses were conducted using IBM SPSS Statistics for Windows, Version 29.0.2.0 (IBM Corp., Armonk, NY; RRID: SCR_002865). In SPSS, Bonferroni corrections are applied by adjusting the p-values rather than the alpha threshold itself. For transparency, we have now clarified in the manuscript the factors included in each statistical comparison to make the tested hypotheses fully explicit.

(6) Line 346: The sample size calculation used in the current study could be improved. It is unclear why a sample size estimation of ≥18 is used when the actual sample recruited is significantly greater than this (62). Is this because 62 participants were divided across multiple experiments in this paper? Further clarification is needed. The alpha value used in this calculation should also be corrected to account for multiple statistical models.

We thank the reviewer for noting this. We have clarified that the initial sample size estimation (n = 18) referred to individual ANOVA analyses. In the revised manuscript, we specify in each experimental section (Identity of Sound Environments and Continuous Theta Stimulation) the exact number of participants recruited per experiment, which together sum to a total of 74 participants across all experiments (also clarified in the Participants section of the Materials and Methods).

(8) Line 438: Many of the statistical tests presented in the Results section have not been outlined in the Methods section or had the rationale explained in the Introduction – this makes the analyses feel confusing and speculative. Explicitly identifying the core hypotheses earlier in the manuscript and clearly stating which hypotheses are confirmatory or exploratory would be an important improvement.

We appreciate the reviewer noticing this. We have clarified the hypotheses tested in the Introduction (4th and 5th paragraphs) and provided a more detailed account of the statistical analyses in the Materials and Methods (Statistical Analysis section).

(9) Line 565: It is unclear how the 62 participants recruited in the study were divided across each of the experimental conditions. I would recommend expanding the Participants section in Methods to outline the number of participants involved in each stage of the study.

Thanks for noticing this discrepancy—this calculation was indeed confusing and incorrect. We ultimately tested a total of 74 participants. We have specified the sample size per experiment in both the “Identity of selected sound environments” and “Continuous theta-burst stimulation” sections of the Methods, and reiterated this in the Results section to prevent confusion.

(9) Line 792: The Results section as a whole is incredibly complex and lacks structure. This could be condensed significantly – details regarding previous research should be removed from Results, as this should already be outlined in the Introduction as rationale for the current project. It would be clearer to outline each confirmatory and exploratory hypothesis in the Introduction, then identify at each stage in the Results section where it is being tested.

We appreciate the reviewer’s point. We have rewritten the 4th and 5th paragraphs of the Introduction to outline the confirmatory and exploratory hypotheses. While we have streamlined parts of the Results section to enhance clarity, we have retained contextual information for each analysis to help readers follow the logic of the findings, given the complexity and scope of the study.

Reviewer #3 (Recommendations for the authors):

(1) Introduction – Overinterpretation and Hypothesis Clarity. The final paragraph of the Introduction (page 4) discusses the experimental findings and their interpretation, which belongs in the Discussion. This section should instead clearly state hypotheses for both the behavioural and TMS experiments. In particular, the TMS experiment lacks a clear rationale: what mechanism is being tested, and what behavioural outcome is predicted? Please revise this section to focus on the theoretical motivation, clearly defined hypotheses, and expected results, and less on summarizing the results.

We appreciate this point and have accordingly deleted the final paragraph of the Introduction, as it is already incorporated into the Discussion section. We have stated the hypotheses/rationale more clearly in the Introduction.

(2) TMS – Mechanism, Timeframe, and Clarity. The manuscript does not adequately explain how TMS produces long-lasting effects relevant to the task, which occur minutes (or possibly longer) after stimulation.

(a) What is the specific timeframe between stimulation and behavioural testing?

The timeframe between stimulation and behavioural testing has been clarified in the Methods section, e.g.: “Participants exposed to ‘real’ or ‘sham’ TMS completed the familiarization and behavioural task right after cTBS procedures.”

(b) What is the evidence that TMS to the prefrontal cortex affects function on this timescale?

We have added the following paragraph to the Discussion to address the timescale of continuous theta-burst stimulation (cTBS): cTBS, as used in our study, typically induces aftereffects lasting 20–50 minutes (Huang et al., 2005; Wischnewski & Schutter, 2015). While these effects are well established in the motor cortex—with motor-evoked potential changes persisting for up to one hour—recent evidence suggests that similar durations of cortical modulation can also occur in the prefrontal cortex (Taylor et al., 2025). Specifically, studies applying inhibitory rTMS to the dorsolateral prefrontal cortex (dlPFC) during cognitive tasks have demonstrated functional effects lasting up to one hour in healthy participants (Wagner et al., 2006). Furthermore, Tupak et al. (2013) showed that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels, reflecting decreased cortical activity, for at least 45 minutes—the same duration as the experimental task in our study. Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), the fNIRS-measured alterations in local cerebral blood oxygenation provide an indirect but reliable indicator of TMS-induced neural modulation within this timescale.

(c) Can post-stimulation effects be objectively measured or confirmed?

Although no objective post-stimulation measures were collected in the present study, we acknowledge this as a limitation. However, previous research has shown that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels—reflecting decreased cortical activity—for at least 45 minutes (Tupak et al., 2013). Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), these findings indicate that fNIRS can serve as an indirect but reliable method for confirming TMS-induced neural modulation. We plan to incorporate such objective measures in future studies.

We thank the reviewer for raising these important points and have addressed them by updating the Methods and Results sections and adding a paragraph to the Discussion. We would like to clarify, however, that the TMS effects observed in our study are not weak: the statistically significant differences between sham and TMS conditions were accompanied by large effect sizes, indicating that bilateral inhibitory stimulation of the dlPFC produced a robust and consistent effect across participants—specifically, a reduction in performance, reflecting decreased improvement in speech understanding with increasing exposure to the reverberant environment.

(3) Figure 3A – Anatomical Specificity and Interpretation. Figure 3A implies precise stimulation of dlPFC and its projections to auditory cortex (A1), but the authors cannot actually target dlPFC or its connections specifically with this approach. Rather, the TMS protocol disrupts an undetermined region of PFC, with diffuse downstream effects. This should be clearly acknowledged in the figure legend and main text.

We appreciate this comment and have acknowledged this point in the figure legend and Discussion, e.g., Figure 3 legend: “Although the TMS protocol was intended to target the dlPFC, it likely affected adjacent prefrontal regions, leading to diffuse downstream effects that may have included modulation of A1.”

(4) Additionally, the Discussion overstates the evidence for a specific dlPFC → AC role in reverberation learning. The weak and poorly localized TMS effect does not support strong claims about this pathway. Please scale back this interpretation and focus more on the robust psychophysical results, which are the manuscript's stronger contribution.

We thank the reviewer for this comment. We have revised our interpretation to clarify that the proposed dlPFC–auditory cortex link is speculative, and have added caveats regarding the limited spatial precision of TMS targeting and individual variability in its effects, in the Discussion section “A role for dlPFC in statistical learning of room acoustics.”

(4a) Line 152: Extra comma after “of”? Also, why are there square brackets around “callsigns” etc.?

Fixed.

(4b) Line 156: “RRID:SCR_001622” is unexplained and likely unnecessary—consider removing.

Removed.

(4c) Line 159: Why was no ramping applied at the end of the noise? Please clarify.

Similar to Brandewie & Zahorik (2013), no ramping was applied. This has been clarified in the Methods (“Acoustic Stimuli”).

(4d) Line 207: Methods do not describe the sham TMS protocol—please add this information.

Thanks for noticing this. Information on the sham TMS protocol has been added.

(4e) Line 352: Fix bracket formatting.

Fixed.

(4f) Line 382: Sentence is grammatically incorrect—please revise.

Fixed.

(4g) Line 394: Unclear use of square brackets—clarify or standardize.

Fixed.

(4h) Line 401: It is unclear how interleaving the talker and length ensures a different room each trial. Aren't these variables independent?

The reviewer is correct: the only variable that was pseudorandomized was room order, to prevent carry-over effects, similar to Brandewie & Zahorik (2013). This has been clarified in the Methods section.

(4i) Figure 1: Clarify that AI-generated images refer only to the room images, not other components.

Fixed.

(4j) Figure 2A: Confirm that “overall” includes all speakers and durations—clarify in legend.

Fixed.

(4k) The interesting duration effects in Figure 1D are not discussed in the text and appear before overall room effects in Figure 2A—please reorder and comment on these results.

Fixed.

(4l) Supplementary Figure 2: Caption contains a typo (“Lecture Room/Open-Plan Office”).

Typo has been fixed.

Also, I recommend adding this result to the main figure set—e.g., include overall d′ for all six talkers in Figure 2 alongside rooms (2A) and lengths (2D).

We thank the reviewer for the suggestion. Including overall performance for all six talkers in Figure 2 would require substantial restructuring and risk making the figure crowded. We have therefore retained these results in the Supplementary Materials (Supplementary 2 and 3), as originally presented.

(4m) Line 555: Phrase “to better understand” could be clearer—consider rewording.

This section has been reworded.

(4n) Lines 587–592: The lack of main effect of TMS is helpful, but more important is whether interactions between TMS and room/length variables occur. Please report these interactions, as they are central to interpreting the TMS effects.

We appreciate the reviewer highlighting this. We have reviewed this analysis and reported the interaction Condition x CP length as follows: “A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], and no significant interaction Condition x CP length was observed: [F (3,93) =0.48, p=0.69, ŋp2 = 0.01]; confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population”. 

(4o) Line 601: Reiterate the timeframe of the TMS-behaviour gap. Is there supporting evidence that TMS can affect behaviour over this duration? Could null effects reflect fading TMS efficacy?

We appreciate the reviewer pointing this out. We have clarified the timeframe of TMS stimulation in both the Methods and Results, e.g.: “The procedure began with TMS manipulation, and although the behavioural task lasted 45 minutes, the inhibitory effects of TMS extended for at least 60 minutes post-stimulation (Huang et al., 2005; Gamboa et al., 2010; Hoogendam, Ramakers, & Di Lazzaro, 2010; Romero et al., 2022).” However, we cannot dismiss individual differences in the duration of TMS effects, nor differences in efficacy duration between anatomical areas (motor cortex vs. dlPFC). We have noted this in the Discussion section “A role for dlPFC in statistical learning of room acoustics” (Pallant, 2011).

(4p) Figure 3I: The “meta-adaptation” effect is marginal in both Exp 1 (p = 0.03) and Exp 2 sham (p = 0.04). These should be interpreted cautiously, given their statistical fragility.

We appreciate this comment. We have now calculated effect sizes for all Wilcoxon signed-rank tests (Pearson’s r) and report them. For the two comparisons noted by the reviewer, the effect sizes are medium (Exp 1) and large (Exp 2). We are therefore confident that, even where the p-values are not extremely low, the statistical differences are reliable.

(4q) Line 696: Reverberation is described as “common,” but it is nearly universal. Consider rephrasing to reflect this.

We appreciate this suggestion and have rephrased this line.

(4r) Line 816: The authors state that TMS reduced overall performance, but the earlier ANOVA (lines 587–592) shows no such effect. Please correct this discrepancy.

We appreciate the reviewer noticing this. This section has been clarified: the lack of statistical significance at lines 587–592 relates to the comparison between a subset of ‘no-TMS-exposed’ listeners and ‘sham’-TMS-exposed listeners, made only to demonstrate the absence of placebo effects in the sham sample. Following the reviewers’ suggestions, we also re-analysed the data using a one-way ANOVA; this slightly changed the numerical values of the reported main effect but did not change the statistical outcome.

Reviewer #4 (Recommendations for the authors):

(1) Lines 201–202: It's not clear what is meant by combination and by carrier length here.

This section has been rewritten for clarity.

(2) Line 330: What is meant by “Univariate” here? I think this was a mixed ANOVA, with a betweensubject factor of TMS exposure and the remaining factors within-subject.

We appreciate this suggestion; we have re-analysed this section to use a mixed-ANOVA design. The numerical results differ, but the statistical outcome remains the same.

(3) Lines 335–343: This is impossible to follow if one does not understand that there were 3 followup experiments.

Thank you for highlighting that this section was confusing. We have rewritten it to clarify the following: “Three follow-up experiments were performed (univariate ANOVA) to assess whether speech understanding was affected by room context (i.e., the third room in which Open-Plan Office and Underground Car Park were learnt), with one between-subjects factor: room context (levels: Anechoic Room, Living Room, Lecture Room, and Highly Reflectant Room).” We have also added the following earlier in the Methods: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).”

(4) Lines 414–416: The review of Tsironis et al. (2024) (doi:10.1177/23312165241273399) does not provide strong evidence that there are multiple scales for adaptation to room reverberation (most adaptation effects stabilize within 1 sec).

We apologize for this mistake, which arose from an issue with our reference manager. It has been corrected to: Robinson, Harper, & McAlpine (2016), Nature Communications, and Simpson, Harper, Reiss, & McAlpine (2014), Journal of Neuroscience.

(5) Line 427: The supplementary figure shows that many subjects did not achieve asymptotic performance.

We appreciate the reviewer pointing this out. The fittings have been extensively reviewed; please see the Methods and Results for the new fitting analysis. Indeed, some participants, although very slowly, keep improving over time without reaching clearly asymptotic behaviour. This section, however, referred specifically to the point at which performance stabilised within ±10% of final performance.

(6) Line 428: The ±10% statistic is random (as discussed below). And why switch to HR now? And what is its meaning when the HR did not converge by the end of the run?

We appreciate the reviewer raising the inconsistent use of d′ versus HR. d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis; time-course analysis does not allow us to calculate d′ at each trial or time point, owing to the lack of HR and FA values for single trials. We have clarified this in the Methods: “Given how d′ was calculated for |Color| and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data for each |Color| and |Number| affecting the temporal resolution of any generated curve. We therefore analysed the development of individual and average performance in each acoustic environment using cumulative hit rates, applying a 5-point moving average (~7 seconds) to each trace and plotting performance as a function of mean cumulative exposure time.”

We have also extensively reviewed our fits, following these steps: (i) we compared single- vs. double-exponential fits across 22 participants (Supplementary Figure 1), which showed that double exponentials provided a better R2 for the majority of participants; (ii) we re-ran all analyses forcing the fits to the HR endpoint; (iii) characterizing taus was not informative for our sample, given flat-like performance for some participants (Supplementary Figure 1)—in these cases, taus do not aid understanding of how performance stabilizes over time, particularly given the use of two taus; and (iv) cutting initial points differs by participant.

(7) Line 434: Or that they learned/adapted to other characteristics that were fixed.

Thank you—we have added “adapted to” in the sentence.

(8) Supplementary tables often show differences, but the actual values are not shown. Also, the tables and figures randomly switch between d' and HR.

Supplementary tables are intended only to show additional detail not reported in the main text or figures, to avoid redundancy; means (referred to in the Supplementary tables) are always shown in the main figures. We appreciate the reviewer raising the inconsistent use of d′ versus HR, addressed above, and have clarified this throughout the Methods and Results.

(9) Lines 436–442: There seem to be a lot of issues with the fitting shown in Supplemental Figure 1 and Figure 2B:

(a) It does not seem to converge, especially for the green line. So, presumably, the asymptotic value obtained for tau_slow was the upper bound set to 2000 s for many subjects' conditions. But those values are never shown—they should be in Supplemental Figure 1.

(b) Then the FHR value, derived from that, is completely dependent on what the bound was set to, and is therefore arbitrary. And its value of 10% is also arbitrary. Why do this when tau itself of an exponential model represents the time it takes to reach 67% of the asymptotic value, from which one can derive whatever time it should take to reach the final 10%?

(c) Even the use of the model specified by Equation 2 seems arbitrary. Average data in Figure 2B do not provide strong evidence for two time scales. If the authors are worried about the instability of the data at the beginning, a simple exponential with a weighted fit that prioritizes the later portions seems sufficient.

(d) Lines 430–434: This conclusion seems wrong, based only on the arbitrary measure chosen for “global likelihood learning.” Looking at Figure 2B, there is no evidence that the green graph reached any asymptote, while for the yellow and blue it appears to have. The authors should try fitting a simple exponential function to it to show that tau is larger.

We appreciate the reviewer’s comments and have significantly revised these sections of the Methods and Results. In summary, we fitted the data with single- and double-exponential functions. Double exponentials were fitted to the full time course. Single exponentials were fitted to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) to mitigate initial variability (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments. Compared against the truncated single-exponential fit, the double-exponential model retained a lower mean AIC (−420.7 vs. −399.1) and remained the preferred model for approximately half the subjects. The double-exponential model was preferred not only for its automated nature (requiring no manual truncation) but also for the magnitude of improvement: when the single model was superior, the advantage was marginal (ΔAIC = 9.7 ± 1.4), whereas the double model’s advantage was substantial (ΔAIC = 52.8 ± 8.9). We therefore used the double-exponential fit for further analysis.

(10) Lines 436–446: How can this analysis be performed if asymptotic performance was not achieved in any of the conditions (nothing has plateaued in Figure 2B)? Also, the 10% FHR measure is dependent on the FHR estimate; correlating two measures based on the same measure is, by definition, expected to be correlated. This result seems to reflect that if one's learning is faster within a fixed number of trials (150), one has more opportunity to reach a higher final PCC even if asymptotic performance is identical.

To clarify, the variables being correlated are not the FHR and ±10% of the FHR values themselves, but rather the time points at which each participant reached ±10% of their individual FHR during the task. This analysis therefore does not involve two measures derived directly from the same estimate. The timing of reaching ±10% of the FHR reflects the learning-settling trajectory rather than the FHR magnitude, so there is no a priori reason for the two measures to be intrinsically correlated. While asymptotic performance was not reached within 150 trials for some participants, the estimated FHR still provides a consistent individual marker of learning rate, allowing comparison of relative learning dynamics across participants and conditions.

(11) Lines 448–461: Brandewie & Zahorik (2013) show that a large portion of that improvement is due to tuning to the voice and location of the speaker. Also, in the current study, there are some issues with the anechoic condition (see below).

We thank the reviewer for this comment. This section refers specifically to results related to carrier phrase length, not to speaker identity (addressed separately below) or location, both of which were fixed in our study and therefore unlikely to account for the observed effects. We address the reviewer’s concerns about the anechoic condition in our responses below.

(12) Line 491: What were the average trial numbers for the steady and initial trials? Also for the anechoic condition?

We appreciate the reviewer raising this. We analysed the average trial number at which initial and steady trials occurred across a total of 360 trials: for all 22 participants, initial trials mean = 6 ± 4 and steady trials mean = 37 ± 9; sham TMS: initial trials mean = 5 ± 4, steady trials mean = 34 ± 6; real TMS: initial trials mean = 10 ± 9, steady trials mean = 38 ± 12; anechoic condition: initial trials mean = 6 ± 3, steady trials mean = 36 ± 8; and the 11 randomly selected subjects: initial trials mean = 6 ± 5, steady trials mean = 38 ± 7. This information has been added to the relevant Results sections.

(13) Line 498: It's still not clear when the anechoic condition was performed. Lines 200–205 talk about combinations in which the Lecture Room was swapped, but it's impossible to follow when and how often that occurred. Given that the anechoic room was not included in the same way as the main three rooms, the conclusion at lines 500–504 is questionable.

Thank you for noticing this. We have rewritten the relevant section of the Methods (“Identity of sound environments”) as follows: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics—i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).” The anechoic room was therefore explored in the same manner as the main three rooms.

(14) Lines 508–521: Tuning to the talker's voice/location would not predict that the effect would be different for a different voice.

We thank the reviewer for this point. If the improvement in speech understanding were due to tuning to a specific talker’s voice or location, we would expect the effect to differ across talkers. However, our analysis across six talkers (three female, three male) showed no significant interaction between talker, carrier phrase length, and room (F(30,630) = 0.73, p = 0.85, ηp2 = 0.034). Although overall performance differed across talkers (main effect of talker: F(5,105) = 27.19, p < 0.001, ηp2 = 0.56), these differences did not modulate the carrier phrase effect. We therefore conclude that the improvement in speech understanding with increasing carrier phrase length is consistent across talkers.

(15) Lines 523–536: Neither of these tests addresses the question directly. That would require switching the talker randomly between the carrier and target (or throughout the sentence).

We thank the reviewer for this comment. We respectfully disagree that our analyses fail to address the question. While our experiment was not specifically designed to test the effect of switching talkers between the carrier and target segments, we examined whether adaptation to a talker could explain the improvement in performance with increasing carrier phrase length through three complementary analyses: (1) a repeated-measures ANOVA testing for interactions between talker and carrier phrase length (see response above); (2) an analysis of potential carry-over effects across consecutive same-talker trials; and (3) an assessment of talker-learning effects in the absence of reverberation (anechoic condition). As detailed in the Results section “Improvements in performance are explained by exposure to the environment, not talker idiosyncrasies,” none of these analyses revealed evidence that talker identity influenced the observed improvement in speech understanding. We therefore conclude that the performance improvements with increasing carrier phrase length are better explained by adaptation to the acoustic environment than to specific talkers.

(16) Line 544: Why is FHR used in this measure when d' is used for the standard analysis in Figure 2D?

As noted above, d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis. Given how d′ was calculated for | Colour | and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data affecting temporal resolution. We therefore used cumulative hit rates, with a 5-point moving average (~7 seconds), plotted against mean cumulative exposure time. This has been clarified in the Methods.

(17) Also, why is FHR, as opposed to HR (which I assume is really PCC), computed across the whole experiment?

FHR refers to the Final Cumulative Performance. This naming was used to distinguish it from trial-by-trial Hit Rate used in the time-course analysis. FHR is the final data point of the cumulative hit rate—i.e., after all responses have been accumulated in that listening environment. This has been clarified throughout the manuscript.

(18) Still worse, it's also not clear when these anechoic trials were measured.

We appreciate the reviewer noting a lack of clarity here. We performed three follow-up experiments in different, naïve groups of listeners, assessing performance across combinations of three rooms, including Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners), assessed in the same way as Brandewie & Zahorik (2013). The anechoic room was therefore explored in the same manner as the main three rooms; this has been clarified in the Methods and reiterated in the Results.

(19) And specifically, from Supplemental Table 4, it looks like the improvement was considerable (up to 15%), supporting that the effect is occurring. Also, note that there seems to be something numerically wrong in Supplemental Table 4: the improvement CP0–CP1 is −6, CP1–CP2 is −9.667, and CP2–CP3 is −2.5. Based on this, CP0–CP2 is expected to be −15.667 (which it is), but CP0–CP3 is expected to be −18.167, yet it's stated as −13.167.

We appreciate the reviewer pointing this out. Our statistical analysis does not match the calculations the reviewer derived from Supplemental Table 4. For transparency, we report below the means for each carrier phrase in the anechoic room, exported directly from SPSS, which are the values reported in the manuscript. We have reviewed this section to improve clarity and have included a link to the raw supplemental data.

CP0: Mean 40.500, SE 4.548, 95% CI [30.212, 50.788]

CP1: Mean 46.500, SE 5.296, 95% CI [34.519, 58.481]

CP2: Mean 56.167, SE 3.777, 95% CI [47.624, 64.710]

CP3: Mean 53.667, SE 3.966, 95% CI [44.695, 62.638]

(20) Lines 555–557: This sentence seems grammatically incorrect.

It has been corrected.

(21) Lines 559–560: The sentence “a brain region implicated in listening performance in noise (Houtgast & Steeneken, 1973; Knudsen, 1929; Lochner & Burger, 1961)” seems to imply that the cited studies support dlPFC being the brain region implicated in hearing in noise. None of these studies does that.

Thank you for noticing this—this was an error introduced by our reference manager and has been corrected.

(22) Lines 586–592: What was the “overall performance” measure—d′, HR, or PCC? Also, what is “univariate” analysis here? A mixed ANOVA with a between-group factor of condition and withingroup factors of room and CP length would be appropriate, and the whole group of 22 subjects should be used for the “no-exposure” group, rather than a random selection of an 11-subject subgroup.

We appreciate this comment and we have revised this analysis to include a mix ANOVA as suggested by the reviewer. It reads as follows in the Manuscript: “Given the potential placebo effects of a ‘sham’ TMS stimulation, we first tested whether our sample of 11 ‘sham’ TMS participants exhibited similar behavioural performance to the larger sample of 22 participants who had not been exposed to any TMS manipulation. A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population.”

(23) Lines 600–601: By “univariate ANOVA” is meant one-way ANOVA? And why wasn't it a two-way ANOVA with factors of room and sham/real TMS? More importantly, the 10% of FHR measure is arbitrary and should be replaced by standard fitting, as discussed earlier. Looking at Figure 3B, the black line appears near an asymptote while the red one is still growing toward the end, and that should be reflected in tau.

The ANOVAs in this and other sections have been revised based on the reviewers’ suggestions; mixed ANOVAs have instead been performed and reported, yielding similar results. The fittings and 10% FHR calculations have also been extensively revised. We now show that double exponentials are better suited to our dataset, and that two-tau parameters are not informative about when performance reaches a stable point during the task.

(24) Also, why is the exposure time on the x-axis different in Figure 3B from Figure 2B (150 vs 500)? And it would be good to see where the across-room average no-TMS data would lie here (or show the equivalent average in Figure 2B).

We appreciate the reviewer noticing this mismatch. Figure 2B shows the time course for each environment (150 s of exposure to each), whereas Figure 3B shows all environments collapsed (150 s × 3). This is because, for the 22 listeners without TMS exposure, a Rooms main effect was observed, justifying separate time courses per room; however, for listeners exposed to sham and real TMS, no Rooms × TMS interaction was observed, so separating time courses per room was not statistically justified. As the only significant effect was TMS condition, we grouped the time spent across all environments by TMS condition.

(25) Lines 606–617: This analysis and Figure 3C have the same issues as described for Figure 2C— asymptotic performance was not achieved for many conditions, so the 10% measure is arbitrary, as is the resulting correlation.

We appreciate the reviewer raising these fitting issues. This part of the manuscript has been extensively revised, including new analyses and figures, although our results have not changed. Additional detail has been added to the Methods (“Speech Performance Analysis and Timecourse Fittings of Mean Cumulative Hit Rates”), and the following summary has been added to the Results (“Statistical learning of reverberant environments occurs over long and short time courses”): we fitted data with single- and double-exponential functions; double exponentials were fitted to the full time course, and single exponentials to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments, and remained preferred when compared against the truncated single-exponential fit (−420.7 vs. −399.1, preferred for roughly half the subjects). The double-exponential model was preferred for both its automated nature and the magnitude of improvement (marginal ΔAIC = 9.7 ± 1.4 when the single model won, versus substantial ΔAIC = 52.8 ± 8.9 when the double model won). We therefore used the double-exponential fit for further analysis.

(26) Lines 619–620: Was d′ really calculated using FHR (the final value) and a non-final False Alarm Rate? This would be arbitrary. It is still unclear how HR and FAR are defined here. There is no apparent benefit to switching between HR (Figure 3B/C, presumably overall percent correct, PCC), d′ (D, E, F), and back to HR (H, I).

We appreciate the reviewer raising this. We have clarified in the Methods (“Speech performance analysis and time-course fittings of mean cumulative hit rates”) how Hit Rate, False Alarm Rate, and d′ were calculated.

(27) Line 655: Figure 3E should be Figure 3F.

Corrected.

(28) Lines 661–670: Why was Number only analyzed for initial trials, while Color was analyzed for both initial and steady trials? Also, the choice of trials 9–10 for “steady” is arbitrary and should be shown somewhere in Figure 3B.

We analysed performance for |Number| on initial trials (1–2) only, for CP0, because performance for this speech token could only improve if positively influenced by short-term, within-trial accumulation of information (acknowledging that | Colour | precedes |Number|). To determine how much knowledge accumulated over repeated exposures — i.e., metaadaptation—we instead needed to compare performance on a speech token whose improvement could only stem from knowledge gained across previous trials, not within a single carrier phrase. We compared | Colour | performance on CP0 between initial trials (1–2) and later, steady trials (9–10). If this hypothesis is supported, it suggests that | Colour | performance for CP0 benefits from meta-adaptive information conveyed across trials as knowledge of the environment’s global structure accumulates—our proxy for meta-adaptation (Figure 3G). The choice of trials 9–10 as “steady” follows work on animal models of meta-adaptation (Robinson, Harper, & McAlpine, 2016), which described a faster adaptation rate after the eighth presentation of an environment. This has been clarified in the manuscript.

(29) Lines 724–728: This description is confusing. The main effect of “Lecture Room” vs. “Highly Reflectant” context is that performance is very good in the Lecture Room (green line) and poor in the Highly Reflectant Room (purple). Averaging that with OPO and CP and reporting “mean difference = 20.09” (in what units?) as “overall performance” distracts from the main point. Moreover, how can that be entered into an ANOVA when the room contexts differ (LR+OPO+CP vs. HR+OPO+CP)? That ANOVA seems incorrect; it should only be performed on OPO+CP across the two contexts.

This section has been revised and re-analysed as suggested. Redundant and unnecessary statistical comparisons were removed, retaining only those that show the effect of context on OPO and CP when comparing the different contexts in which these common environments were learned.

(30) Lines 740–750: Again, it is not surprising that when Living Room replaces Lecture Room—and performance in Living Room is worse than in Lecture Room—the average of LiR+OPO+CP is lower than LER+OPO+CP, if OPO+CP performance is unchanged. The interesting question is whether anything changed in OPO+CP performance, as suggested for the previous point.

This section has been revised as suggested by the reviewer.

(31) Lines 752–771: Again, the same issue—the main effect is that performance in Anechoic trials is worse than in Lecture Room or Living Room trials, which alone explains the group difference. I am also sceptical of the finding that Anechoic performance is worse than reverberant performance, contrary to typical spatial-release-from-masking results, where reverberation degrades performance by adding noise energy at the better ear and reducing binaural benefit through decorrelation. This may be an artefact of how target and noise levels were normalized after convolution with HRTFs/BRIRs (or the use of Ambisonics); no acoustic analysis of the stimuli is provided. At minimum, the total received level at the two ears for target and masker in every environment should be reported. Brandewie & Zahorik (2013), using equivalent anechoic and reverberant conditions, never observed reverberant performance to exceed anechoic, contrary to what is stated here (lines 754–755).

We appreciate the reviewer raising these points. The statistical analysis in this section has been revised: only the common rooms across the three-room conditions (Open-Plan Office and Car Park) were directly compared. Performance in Living Room and Lecture Room was not statistically different (Results, paragraph 4, “Statistical learning of room acoustics is tuned to universally experienced reverberation times”). However, Anechoic and Lecture Room performance remained significantly different (mean difference = 11.08, t(9) = 2.66, p = 0.013, Cohen’s d = 0.84), as did performance in the common rooms when learned in the context of Lecture Room versus other contexts (mean difference = 9.7, F(1,41) = 13.24, p < 0.001, ηp2 = 0.26).

While this setup resembles many masking studies, Brandewie & Zahorik (2013) tested four rooms simultaneously, whereas we tested three-room conditions explicitly designed to test environment-mix adaptation. We observed a synergistic relationship between performance in ‘good’ reverberant rooms (Lecture Room, Living Room) and the common but less favourable rooms (Open-Plan Office, Car Park, with longer RT60): in the absence of a ‘good-reverb anchor,’ performance in the common rooms improved less over time, possibly because participants had less to leverage in anechoic environments.

We agree that verifying at-ear acoustic levels is critical to ruling out a normalization artefact. Our stimuli were normalized in the 41-channel sound field, not at the listener’s ears: source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses, the 41-channel noise energy was scaled to a target of 70 dB, and the 41-channel speech field was scaled to the target SNR; the 41-channel signals were then rendered to two channels using a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related acoustic effects such as head shadow.

To verify that this did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (see supplied table, Summary Reverb Data). In the anechoic condition, the noise (positioned to the left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, raising noise level at the right ear to 55–56 dB depending on room, substantially lowering the ear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming this is not a normalization artefact but rather a genuine perceptual spatial release from masking, likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added the at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2, and updated the Discussion to clarify this mechanism.

(32) Lines 755–759: Describing anechoic spaces as “rare” and as rooms whose “walls are treated” states the facts backwards. Open spaces (e.g., a grass lawn) are largely anechoic fields, and people spend considerable time in such environments. An anechoic room may be artificial, but an anechoic (or near-anechoic) space is very common, and the room is simply an attempt to simulate that within an enclosure.

We appreciate this point and have rewritten this section to reflect it.

(33) Lines 808–812: This sentence appears incorrect. It refers to “the ability to correctly report keywords spoken in environments with the more extreme—lower or higher—RIRs,” presumably meaning OPO and CP, but these are not the environments with extreme RIRs; or, if referring to An and HR, those were not “encountered in experimental blocks also containing the moderately reverberant Lecture Room or Living Room.”

Thank you for noting this. We have rephrased this section to refer only to the extreme high-RIR environments encountered.

(34) Line 843: Should Fig 3Ai be Fig 3A? Also, in that figure there are arrows between dlPFC and A1, and between A1 and (the cerebellum?)—it's unclear what these represent.

The arrows were intended to represent feedforward and feedback information flow to lower auditory brain centres. We acknowledge they were confusing and have removed them from the figure.

(35) Lines 929–949: The authors did not account for listeners tuning to voice and location (as now cited via Best et al.), and their own and Brandewie & Zahorik's data show improvement due to carrier phrase even in the anechoic case (with the inconsistency in Supplemental Table 4 noted earlier). A direct test—switching the environment between carrier phrase and target phrase, as in Brandewie & Zahorik and Vlahou et al.—would be needed to fully attribute the effect to reverberation rather than other factors.

We thank the reviewer for this detailed comment and agree that directly manipulating the environment between carrier and target phrase would provide the most direct test of environment-specific adaptation. While our study did not implement this manipulation, our data provide converging evidence: (1) listeners showed improvement with longer carrier phrases even in the anechoic condition, consistent with previous reports, but this improvement did not interact with talker identity, carrier phrase length, or room, indicating it is not driven by tuning to specific voices or locations; and (2) regarding Supplemental Table 4, the calculations suggested by the reviewer do not match our statistical analysis—we have reported the SPSS-exported means directly (shown above) and reviewed this section for clarity, including a link to the raw supplemental data. Taken together, while we cannot fully quantify the proportion of adaptation attributable to reverberation versus other factors without the direct environment-switch manipulation, our results indicate that the observed improvements primarily relate to exposure to the environment rather than talker-specific effects.

(36) Lines 951–953: It is unclear what about “understanding speech in background noise” distinguishes this study from previous studies of adaptation to reverberation, many of which also examined speech in noise (as reviewed in Tsironis et al., 2024). Rather than reviewing pertinent studies on adaptation to reverberation for speech tokens, the authors cite abstract noise-texture studies that are only partially relevant, given the prevalence of speech in everyday listening (lines 955–958).

We appreciate the reviewer raising this point. Our intention was to refer specifically to statistical learning of implicit environmental acoustic features such as reverberation, rather than to speech-in-noise perception per se. We have revised the text accordingly: “A key feature of our study, which distinguishes it from previous investigations of statistical learning of acoustic features in human listeners, is the use of an ethologically valid listening task—understanding speech in background noise while listeners implicitly learn repeated acoustic features.”

(37) Lines 978–980: In what way? For environments with large T60, a simpler explanation than “ecological validity” is that there is more late reverberant energy in the target acting as a masker, predictable from DRR.

This section has been rewritten to clarify that it is the decline in performance at longer RT60 that is reminiscent of the decline observed under rTMS.

(38) Line 981: What does “the better to understand speech in reverberant background noise” mean?

This sentence has been revised.

(39) Lines 987–988: When did “performance decline over the course of an experimental session”? Figures 2B, 3B, and 4A all show performance improving over the session.

We have rephrased this sentence to refer to a decline in overall performance.

(40) Lines 990–992: Many previous studies report better adaptation to reverberation for some rooms than others (e.g., Brandewie & Zahorik, 2010; Vlahou et al., 2021), but none have reported decreased performance for an anechoic space relative to a reverberant one. This anomaly should be explained and reconciled with the existing literature before invoking ecological explanations such as “ethologically relevant environments.”

We thank the reviewer for raising this important point. We agree that verifying the at-ear acoustic levels is critical to ruling out a normalization artefact, particularly given our finding that reverberant performance exceeded anechoic performance. To address this directly: our stimuli were normalized in the 41-channel sound field, not at the listener’s ears. The source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses; the 41-channel noise field was scaled to a target of 70 dB and the 41-channel speech field scaled to the target SNR; the 41-channel signals were then rendered to two channels via a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related effects such as head shadow because normalization preceded binaural rendering.

To verify that this sound-field normalization did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (Summary Reverb Data table). In the anechoic condition, the noise (positioned left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, increasing right-ear noise level to 55–56 dB depending on room, substantially worsening the atear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming the finding is not a normalization artefact but instead reflects a genuine perceptual spatial release from masking—likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added these at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2 to clarify this mechanism.

(41) Discussion: Given the questions about the results, the discussion might need to be rewritten to only discuss claims that are actually supported.

The Discussion section has indeed been extensively revised.

(42) The hippocampus and other areas have been proposed for statistical learning, and studies also show that disruption of DLPFC can boost statistical learning (https://doi.org/10.1016/j.jml.2020.104144).

We appreciate the reviewer raising this. We have cited Ambrus et al. (doi:10.1016/j.jml.2020.104144) in the Discussion (line 888) as evidence of opposing effects of dlPFC stimulation on statistical learning, and have further revised lines 891–903 of the Discussion to more clearly describe the known projections and functional interactions between dlPFC, hippocampus, striatum, and basal ganglia that support implicit and statistical learning.

(43) Line 1739: What is “cumulative” here?

“Cumulative” has been deleted.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation