Author response:
The following is the authors’ response to the previous reviews
Summary of revision for all reviewers:
We are encouraged that the reviewers recognized the importance of the central question addressed by this study, the value of comparing acoustic- and cochlear-implant-evoked cortical responses within a common framework, and the potential relevance of these analyses for auditory neuroscience and neuroprosthetic design.
At the same time, the reviewers made clear that the manuscript would be strengthened by:
(1) Clearer calibration of claims regarding spatial organization
(2) Clarified explanation of the Figure 8 cross-modal analyses
(3) More explicit discussion of the methodological and interpretive limitations of the present dataset, particularly the acute, anesthetized, and partially indirectly validated nature of the experiments
In response, we substantially revised the manuscript to improve methodological clarity, narrow claims where appropriate, and align the title, abstract, results, discussion, figures, and legends with the precision of the data.
Summary of major changes to revised manuscript:
We revised the manuscript throughout to distinguish non-random spatial organization, coarse topographic/cochleotopic structure, and locally graded tonotopy/cochleotopy, and we softened claims accordingly, especially for the cochlear implant and TCA-derived analyses.
We substantially clarified the Figure 8 cross-modal decoding framework, including:
- The within-modality normal-hearing control
- The normal-hearing-trained / cochlear-implant-tested analysis
- The shuffled baseline used to estimate chance-level information transfer
We expanded the methods and discussion to make the scope and limitations more explicit, including:
- The acute and anesthetized nature of recordings
- The use of monopolar stimulation
- The lack of direct deafening validation in the main iEEG cohort
- That ECAP forward-masking measurements were obtained in a separate acute cohort
- The possibility that the observed acoustic–electrical mismatch is transient rather than fixed
We also improved the consistency of our terminology, panel labeling, figure legends, and cohort descriptions.
Public Reviews:
Reviewer #1 (Public review):
We thank the reviewer for recognizing the importance of the question addressed by this study and for pushing us to sharpen the distinction between non-random spatial structure, coarse topography, and graded cochleotopy. These comments also helped us better calibrate our claims about TCA-derived maps and cross-modal generalization, and prompted clearer discussion of the limits of the present dataset.
(1) The main weakness is that the evidence for spatial organization remains difficult to interpret. In Figure 2, the authors argue that both tone-evoked and cochlear implant-evoked responses are spatially organized, but the slope analyses are not significant for the cochlear implant condition. The revised vector-strength analysis supports the presence of non-random spatial structure, but this is not the same as demonstrating a clear graded cochleotopic organization. The manuscript would be strongest if it consistently distinguished between non-random spatial structure, coarse topography, and true graded tonotopy or cochleotopy.
We agree. We revised the manuscript to distinguish these levels of interpretation more carefully and consistently. Specifically, we now reserve “non-random spatial organization” for results supported by the vector strength/shuffle analyses, and describe cochlear-implant-evoked maps as showing coarse and variable spatial structure rather than uniformly robust graded cochleotopy.
We now also emphasize more explicitly that implanted animals varied considerably: some showed clear nonrandom organization, whereas others exhibited effectively random global maps. At the same time, the aggregate implant-evoked data remained non-random and showed decreasing spatial correlation with increasing electrode separation, which we interpret as evidence for coarse cochleotopic structure at the population level, rather than strong local graded cochleotopy in every animal.
Accordingly, we revised the manuscript to distinguish:
“Non-random spatial organization”
“Coarse topographic/cochleotopic structure”
“Locally graded tonotopy/cochleotopy”
(2) A related issue is that some figure titles and interpretive statements still appear stronger than the data justify. For example, the TCA results in Figure 7 are described as revealing topographically organized latent spatial factors, but the statistical support appears strongest for normal-hearing high-gamma responses, with weaker or non-significant results in other conditions. These data remain interesting, but they would be better framed as evidence for weak or coarse spatial structure rather than robust topographic organization across all modalities.
Now in our updated manuscript we revised the Figure 7 title, legend, and results to avoid implying robust topographic organization across all conditions. The manuscript now describes the TCA-derived spatial factors as showing coarse, non-random spatial structure, with the strongest support for local tonotopy in the normal hearing high-gamma condition and weaker or non-significant support for robust local topography in the other conditions.
(3) The decoder analyses are improved, especially with the added tone-to-tone control. This control supports the conclusion that poor acoustic-to-CI transfer is not simply a failure of the TCA/LDA pipeline. However, the analysis remains model-dependent, and the absolute information transfer values are low. It would be helpful either to include an analogous analysis using raw ERP/high-gamma features or to explain more explicitly why the TCA-based approach is the appropriate primary test. The data support poor generalization between acoustic and implant-evoked cortical responses, but claims about perceptual qualities should remain speculative because perception is not directly measured in these experiments.
Thanks, we now revised the Figure 8 results section to first explain that TCA is especially appropriate here because it constrains the decomposition into separable spatial, temporal, and trial factors. This makes it possible to learn spatial and temporal factors in the normal-hearing condition, hold those fixed, and re-optimize only trial factors in withheld normal-hearing data or cochlear-implant-evoked data, providing an interpretable test of within-modality recovery and cross-modal generalization.
Thus, our emphasis on TCA is primarily conceptual, not merely computational. Raw-feature decoding could also be informative, but for the specific cross-modal question addressed here, TCA provides the more interpretable framework.
We also now state more explicitly that the absolute information-transfer values are low, and we correspondingly limit our conclusion: normal-hearing-trained representations generalize poorly to acute cochlear-implant-evoked responses in this framework. We further and thoroughly revised the abstract, results, and discussion so that any perceptual implications remain explicitly speculative, since perception was not measured directly in these experiments.
(4) Finally, although methodological reporting is much improved, some verification remains indirect. The authors provide useful implantation criteria and cite prior validation of their deafening approach, but the manuscript would be clearer if it explicitly distinguished between validation performed in the present animals and validation based on previous cohorts. This distinction is important because surgical variability, implantation efficacy, and deafening completeness can influence the interpretation of cochlear implant experiments.
We revised the manuscript to make this distinction explicit. In the methods, we now state clearly that we did not obtain ABR measurements, hair-cell counts, or behavioral confirmation of deafening in the animals used for the present iEEG dataset, and that support for deafening efficacy in this cohort therefore relies on prior validation of the same procedure in separate cohorts.
We also now clarify that the ECAP forward-masking measurements were obtained in a separate acute cohort (N=3) and were not recorded from the animals used in the main iEEG dataset.
Our goal in these revisions was to make the provenance of each validation step fully explicit and to avoid implying that all validation measures were obtained in the same animals used for cortical recording.
Reviewer #2 (Public review):
We thank the reviewer for highlighting the value of the decoder-based analyses while also challenging us to frame the study’s novelty more precisely and to make Figure 8 substantially clearer. In response, we narrowed our novelty claims, improved the explanation of the cross-modal decoding framework, and clarified the logic of the shuffled baseline.
(1) The observation that responses to cochlear implant stimulation (stimulation) is spatially organized is not new (e.g. Adenis et al. 2024)
Thanks, good point. We revised the Introduction to make clear that prior studies have already demonstrated spatial organization of cochlear-implant-evoked cortical responses, including prior animal work cited in the manuscript. Our intended contribution is therefore not the demonstration of spatial organization per se, but rather the use of a shared analytical framework to test whether acoustic and electrical stimulation produce overlapping or transferable spatiotemporal cortical population representations, including single-trial decoding and cross-modal generalization analyses.
(2) The claim that spatial and temporal dimensions contribute information about the sound is also not new there is a large literature on this topic.
We revised the manuscript to avoid implying novelty for this general principle. Our intended point is narrower: in the present work, spatial and temporal dimensions were recorded simultaneously across a 60-channel cortical surface array and integrated into trial-by-trial population analyses that could be directly compared across normal-hearing and cochlear-implant conditions within a shared framework.
(3) The analyses supporting the claim that there is a mismatch between cochlear implant and sound representation are still unclear, particularly in Fig. 8.
We revised both the Figure 8 legend and the results section to explain the analysis step-by-step. In particular, we now distinguish clearly among:
The normal-hearing to normal-hearing re-optimization control, which tests whether the TCA/LDA framework can recover stimulus-related structure when modality is unchanged;
The normal-hearing-trained / cochlear-implant-tested analysis, which tests cross-modal generalization; the shuffled baseline, which estimates chance-level information transfer by destroying any structured relationship between predicted labels and actual implant channels while preserving matrix dimensions.
We now also explain explicitly why the shuffled baseline can equal or slightly exceed the measured normal hearing-to-cochlear-implant transfer. Because observed cross-modal transfer was extremely small, shuffling does not restore meaningful structure; rather, it shows that the measured transfer lies at or below chance level. We believe that this substantially improves the clarity and interpretability of Figure 8.
Reviewer #3 (Public review):
We thank the reviewer for recognizing the strengths of the study design, the single-trial decoding analyses, and the potential clinical relevance of this approach, while also pressing us to better address the limitations of monopolar stimulation, possible place-frequency mismatch, and the acute nature of the recordings. These comments prompted important revisions to both interpretation and discussion.
(1a) The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results as the authors ignored fundamental limitations of CI related stimulation. First, the authors stimulated in a Monopolar mode which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.
We agree that monopolar stimulation is an important limitation, especially in the rodent cochlea, where current spread may be broader than in human subjects. We now acknowledge this more explicitly in the discussion.
At the same time, our ECAP forward-masking measurements in a separate acute cohort provided supportive evidence for spatially and temporally tuned peripheral activation under the stimulus intensities used here. We therefore interpret the data as indicating that some peripheral selectivity was present, while also acknowledging that broader current spread under monopolar stimulation may have contributed to the coarse cortical organization observed in the implant condition.
We also now state explicitly that the observed coarse cortical organization likely reflects a combination of peripheral current spread, downstream cortical pooling, and the spatial resolution limits of mesoscale surface iEEG, rather than any single factor alone.
(1b) Comparing the averaged BF maps for iEEG (Fig-2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies might reveal a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Fig2F (and to some extend 2A) most of CI electrodes elicited responses around the 4kHz regions and averaged maps show a predominance of CI-3-4 across cortex (Fig-2C, H and Sup Fig. 3) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.
We appreciate this point and agree that a precise one-to-one alignment between cortical best-frequency maps and intracochlear electrode position cannot be established from the present data. We now clarify this limitation in the manuscript, noting that the iEEG-based maps are relatively coarse and appear to overrepresent mid-frequency regions, limiting direct inference from implant electrode number to a precise acoustic-frequency equivalent.
We also emphasize that our central conclusion does not depend on assigning each implant electrode a specific best frequency, but rather on comparing the spatial organization and cross-modal generalization of cortical population responses.
(1c) Moreover, Supplemental figure 3 shows that only a couple of CI electrodes are predominately represented at the level of the cortex. Thus, it seems possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the placecoding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses.
We agree that the predominance of only a subset of implant electrodes in the cortical maps is consistent with relatively coarse peripheral and/or cortical place coding. We revised the discussion to acknowledge more explicitly that possible current spread, broad peripheral excitation, and the limited spatial resolution of surface iEEG could all contribute to the coarse spatial organization observed in the implant condition.
We also emphasize in the revised text that this possibility does not negate the evidence for non-random organization, but it does limit the strength of any claim about finely graded cochleotopy.
(2) Second, although the authors acknowledge that post-lingual CI users always have an adaptation period, their conclusion is based on measurements that are relatively "early" in the CI-use timeline so to speak since iEEG were collected a) acutely right after mono-aural implantation and stimulation, b) under anesthesia, c) using unmodulated pulse train fixed at 900pps regardless of the electrode used and thus lacking any temporal information shifts in relationship to electrode cochleotopic placement. Basically, all CI electrodes had the same rate whereas you would expect basal CI electrodes to be amplitude modulated at higher frequencies than apical electrodes.
We agree. We now emphasize more clearly that the present experiments probe only an early and simplified stage of cochlear implant use. The recordings were acute, performed under anesthesia, and used short constant-rate pulse trains chosen to isolate fundamental organizing features of primary cortical responses rather than to reproduce the full complexity of clinical stimulation.
We therefore frame our conclusions specifically in terms of acute cortical encoding under these conditions, and now state explicitly that these data do not address how chronic use, wakefulness, adaptation, or more complex modulation-rich stimuli may alter more global cortical representations.
(3) As much as the reviewer likes the overall approach with the use of PCA-LDA and TCA, and agrees that information transfer seems inexistant at time of measurement, authors should be more careful in their strong conclusion that two distinct encoding exist. The non-overlapping between sound and electric stimulation representations might exist only transiently and this should be acknowledged a bit more in the discussion. Without repetition of iEEG measurement at later period with chronic use of the CI, it is not possible to definitively claim that two distinct, non-overlapping coding co-exist at all times.
We agree and appreciate this point. We revised the Discussion to make this much more explicit. Our data support poor overlap or poor cross-modal generalization at the time of acute implant activation, but they do not establish that this relationship is fixed over chronic implant use.
We now state directly that the observed mismatch may be transient, and that future longitudinal studies will be required to determine whether cortical representations become more aligned with experience, whether downstream readout adapts to a novel code, or whether both processes contribute.
Recommendations for the authors:
Reviewer #2 (Recommendations for the authors):
(1) The authors have provided new analyses to support the claim that CI and NH representations do not match. However, the analyses performed are presented in an unclear manner. Most particularly, it is very hard to understand what has been done in Fig. 8D and Fig. 8G. The legend in D is incomplete. The legend for G does not explain why there are lines with different color and what are the different lines. The reasoning behind the shuffled CI dataset is not explained. Why does shuffling restore information transfer, this is very counter intuitive and raises again the question of the value of this analysis.
Agreed, we have revised the Figure 8 legend and rewritten the Figure 8 results section to initially explain the analysis workflow more clearly.
We now explain that the shuffled condition is used to estimate a chance-level baseline for information transfer in the normal-hearing-trained / cochlear-implant-tested confusion matrix. Specifically, we compute mutual information after randomizing the predicted labels in the confusion matrix, thereby destroying any structured relationship between predicted tone labels and actual implant channels while preserving matrix dimensions and marginal structure (results, Fig. 8 section).
We also now state explicitly that while this shuffled condition can yield slightly higher mutual information than the non-shuffled NH→CI condition, the point is not that shuffling restores meaningful decoding. Rather, the point is that the observed NH→CI transfer is so low that it does not exceed this chance-level baseline (results, Fig. 8 section).
(2) Note that the legends of fig. 8 are mislabel (goes up to H although the last panel is G, a legend for E is missing).
Thank you for catching this error. We have corrected the Figure 8 legend so that all panels are accurately labeled and described.
Reviewer #3 (Recommendations for the authors):
(1) Fig. 2C and 2H are from 2 to 16kHz (as announced in the previous responses) but then Sup. Fig. 3 goes from 2 to 32kHz. Still on Sup. Fig. 3, color code for CI electrodes is inverted compared to the rest of the MS.
Thanks, good catches. We have updated Figures 2C and 2H to represent the full tested frequency range, and we have corrected the inverted color code for cochlear-implant electrodes in Supplemental Figure 3.
(2) Fig. 2A. says n=1, Fig. 2F should say the same. Fig. 2C and 2H should also have n=1 on bottom map. Fig. 2D and 2I should mention n=1. Same thing with Fig. 3A/3C, Fig. 3B/3D (n=7), Fig. 7A/7B/7C.
We have revised the relevant figure legends to make clear that data are from a single animal unless otherwise noted, while minimizing visual crowding in the figure panels.
(3) The reviewer once again thinks it would be better to not truncate the axis of Figs. 4C, 6C, 7D as it creates confusion with Figs. 5D and 8G, where suddenly, there are 8 electrodes. The legend justification isn't enough.
We appreciate this concern and have clarified the issue in the revised manuscript. In some animals (N=3), some electrodes in the 8-channel array were non-functional prior to implantation, so those animals contributed six rather than eight usable implant channels. We now make this explicit in the manuscript and supplemental figures, including the number of channels used in the respective animals in the methods under “Cochlear implant programming” and in Supplemental Figure 3.
(4) Although the reviewer appreciated the justification of 15 PCA for their analysis, this justification should be provided as is in the methods, especially as the Sup. Fig. 4 does not even explain the significance of the linear regression of components 16 to 30. Some other reviewers would say that 10 components were more than enough without this justification.
We agree and have provided further justification of the 15 PCA components to provide justification as is in the Methods “Principal component analysis” and in the results describing Figure 4.
(5) Sup. Fig. 1 and the eCAP measurements are a nice and needed addition to the MS but the authors should make it clearer that it came from 3 animals that are not part of the cohort presented in the rest of the paper / in Sup. Fig. 2. The method section certainly not make that distinction, giving the impression that all animals have been tested for channel interactions and temporal recovery.
We agree and have revised the methods section to state clearly that the ECAP forward-masking measurements were obtained in a separate group of acutely implanted rats (N=3), rather than the main iEEG cohort in Supplemental Figure 1.
(6) As said in the public review, the authors should acknowledge more the potential transiency of the nonoverlapping representation in the discussion, especially since most of the recordings are acute.
We agree, and we have revised the discussion accordingly. As described in our public response above, we now state directly that the poor overlap between acoustic and electrical representations was observed under acute recording conditions and may not persist unchanged with chronic implant use or behavioral adaptation.