Author response:
The following is the authors’ response to the original reviews.
Public Reviews:
Reviewer #1 (Public review):
Summary:
Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that, contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to detect TE insertions from RNAseq accurately. In this short response, Choucri and Treiber clearly demonstrate that differences in the tools used between their study and that of Azad et al. likely account for the contrasting results, along with RT-PCR failure in designing primers that would match the chimeric transcript, and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.
Strengths:
The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. Finally, the discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.
Weaknesses:
I think it is necessary to add more detail to this article, for instance, the differences between TEchim and Tidal could be laid out more precisely.
We thank the reviewer for this helpful suggestion and agree that a more explicit comparison improves the clarity of the manuscript. Briefly, TIDAL and TEChim are designed to answer different questions. TIDAL is an insertion/deletion caller that was primarily built to determine whether a TE insertion exists at a genomic locus and how it is distributed across strains or populations. It was not designed to resolve splice junctions between exons and TEs. TEChim, by contrast, is purpose-built to detect breakpoint-spanning reads that span exon junctions and putative splice sites within a TE. This difference in design has several concrete consequences in the algorithms used:
(1) TIDAL clusters reads that support a candidate breakpoint within a window of twice the sequencing read length (e.g. 300nt for 150nt reads). This works well for calling genomic insertions, but it can be too restrictive for a splice junction, which may fall at a variable position within the gene and TE. TEChim does not rely on fixed-window clustering, but instead maps all reads, groups them to individual nt positions within the genome, and then filters for events that were detected in more than one biological replicate.
(2) TEChim reconstructs long in-silico reads from overlapping paired-end reads, using FLASH, prior to alignment. In-silico paired-end reads are analysed, and full-read merged fragments provide single-nucleotide resolution for a breakpoint. This increases accuracy and confidence in split reads.
(3) TEChim explicitly intersects candidate breakpoints with annotated exon/intron/UTR. Structures and canonical splice donor- and acceptor sites.
We have now added a concise summary of the underlying principles of TEChim in the Methods section, and added details to the comparison of TIDAL with TEChim in the discussion. We hope that this addition makes it clearer why the two pipelines produce different results from the same input data.
Regarding the roo example, one of the caveats of this family, along with others, is the presence of simple repeats. It would be important to show that the simple repeats are not interfering with the read mapping.
We thank the reviewer for raising this important point. We agree that simple repeats can complicate read mapping and that this requires careful consideration. We have now added to the results the exact locations and lengths of the three known and annotated repeat regions within roo (Domínguez, 2021) and show that the breakpoints we report map more than 4kb away from these regions. In addition, the splice junctions we identify (at positions 5190 and 5462) recur at the same position relative to the roo consensus sequence across multiple genomic insertions, and biological replicates. If these calls were artifacts of copyspecific simple repeats that interfered with read mapping, then we would expect breakpoints to vary between with each insertions local sequence, rather than converge on the same breakpoint. We therefore consider it unlikely that simple repeats account for the observed splice junctions. We have expanded our discussion to make this reasoning clearer.
Regarding the experiments, if we are looking for a standardized protocol, then we should have a detailed material and methods section, with every experiment, replicate, and PCR temperature clearly defined.
We thank the reviewer for this suggestion and agree that a more detailed description of the experimental procedures will improve the reproducibility of the study. We have substantially expanded the Materials and Methods section, including the number of biological replicates, primer information, PCR conditions and other methodological details relevant for reproducing the experiments. In addition, we have written a detailed manual for the updated version of TEChim that we used here.
Finally, and in my opinion, more importantly, the use of RT negative controls on the RT PCRs, along with DNA PCRs to show insertion presence, is mandatory for testing the presence of chimeric genes. Of course, water negative PCR controls are also needed, and unfortunately, absent from Figure 3.
We thank the reviewer for this helpful suggestion. We have repeated the RNA extraction on 3 new samples, and this time also included minus-RT aliquots for each of the three biological replicates. We have run our PCR for the chimeric transcript between Beadex and opus on all these samples, and in addition on a water control. All these new results are now shown in Figure 3, and confirm our previous conclusions.
We now also provide results from our DNA testing for the opus insertion in Beadex. We conduct these at regular intervals in our lab to ensure the insertion remains stable in our stock, and mentioned the results in the original version of this manuscript, but we agree that it is important to also show the raw data of this in this study. We use primers at the up- and downstream end of the opus insertion. The downstream pair gave a single band at the predicted size, which we confirmed by Sanger sequencing. This data is now presented in Figure 3 – Figure supplement 1. The PCR around the upstream end resulted in the expected band of 888bp, and two additional bands. Sanger sequencing of the 888bp band produced signals for both Beadex and opus, but we did not get a reliable signal across the precise breakpoint (see Author response image 1). We think this might partly be due to a tandem repeat at the beginning of the opus LTR, which may interfere with Sanger sequencing. Taken together, our data provides strong evidence that the opus insertion is present in our flies.
Author response image 1.
Sanger sequencing results of the 888bp band: Two segments of the raw Sanger sequencing trace are shown; the intervening, unmapped section is omitted. The left segment (grey) shows clean, high-confidence signal matching the Bx locus. The right segment (pink) shows the signal falling to near baseline within the opus LTR, so base calls in this region are log-confidence, and the trace does not resolve the breakpoint. Numbers above the sequence indicate position within the raw sequencing read.

Reviewer #2 (Public review):
Summary:
This study by Choucri and Treiber aims to directly address a recent critique regarding the role of transposable elements (TEs) in diversifying the neural transcriptome of Drosophila. The authors seek to demonstrate that TEs are not merely genomic "noise" but are frequently and reliably "exonized" into brain-specific mRNA. By introducing an upgraded computational pipeline, TEChim, and conducting precise experimental validations, the authors set out to show that TE-mediated splicing represents a genuine biological phenomenon that expands the molecular repertoire of the nervous system.
Strengths:
The study's primary strength lies in its rigorous technical "forensic" analysis of previous failed replication attempts. The authors convincingly demonstrate that the lack of signal in the opposing study stemmed from a fundamental methodological mismatch: the software used by the critics (TIDAL) is logically incapable of detecting splice sites located within TE sequences. Importantly, the authors complement this computational clarification with definitive experimental evidence through an effective "experimental rescue." By employing correctly designed primers and matching the genetic backgrounds of the fly strains, thereby accounting for genomic polymorphisms, they successfully validated all seven loci that were previously reported as undetectable. This dual-pronged strategy, addressing both algorithmic bias and experimental design, establishes a more robust technical benchmark for the detection and validation of TE-derived exons in neural tissues.
Weaknesses:
While the technical rebuttal is highly convincing, the scope of the study remains primarily defensive. As a response to a prior critique, the work focuses on establishing the existence and detectability of chimeric TE-derived transcripts rather than exploring their broader functional consequences. As a result, there is limited new insight into how these TEmodified isoforms influence neural circuit function or organismal behavior.
We agree with the reviewer that the primary focus of this study is to establish a robust protocol for the detection and validation of TE-derived chimeric transcripts, and to resolve discrepancies raised by Azad et al. We believe this provides an important technical and conceptual framework for future studies investigating the functional impact of TE-driven genetic variation.
In addition, the detection and validation of these events remain technically demanding, requiring deep sequencing and specialized bioinformatic expertise, which may limit broader adoption by laboratories without dedicated computational resources.
We agree that the robust detection and validation of TE-derived chimeric transcripts is technically demanding, requiring both high-quality sequencing data and specialised computational analyses. However, we believe that these methodological challenges are justified by the biological insights that can be gained. By providing an updated computational approach together with experimental validation, we hope to facilitate further research into this phenomenon.
Reviewer #3 (Public review):
Summary:
This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.
Strengths:
The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.
Weaknesses:
The authors should mention that combining PCR-amplified cDNA generation with shortread sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct longread ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv). Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.
We thank the reviewer for this excellent suggestion and agree that long-read RNA sequencing represents a powerful approach to characterise chimeric transcripts, because it can capture full-length transcripts and reduce ambiguity of mapping short reads onto repetitive sequences. We have now expanded the Discussion to highlight the advantages of these technologies and to cite the suggested studies. At present, long-read RNA sequencing remains challenging for Drosophila brain samples, because the total amount of input RNA is limited. As a consequence, amplification is usually required, which itself can introduce artefacts, which we showed previously (Treiber and Waddell, 2017). While we agree that long-read sequencing will be an important approach for future studies, we believe that the combination of computational analysis and targeted experimental validation presented here provides robust evidence for the existence of chimeric TE-gene transcripts.
Recommendations for the authors:
Reviewer #2 (Recommendations for the authors):
To maximize the impact of the work, the authors should consider adding a technical guide that outlines the implementation of the updated TEChim pipeline and provides a standardized protocol for designing chimeric RT-PCR primers. This section should include a clear software workflow and a precise primer design strategy, emphasizing junction spanning probes and the necessity of genomic confirmation, to establish these methods as the definitive technical standard for studying transposable elements in the nervous system.
We thank the reviewer for this constructive suggestion and want to be upfront about what can and cannot be standardized here. Because TE sequences are repetitive, the specific primer pair that works best for a given locus cannot necessarily be predicted in advance, and some empirical testing of candidate primer pairs is unavoidable. What we can standardise, and now describe explicitly in the Methods section, is, firstly the use of TEChim to identify candidate splicing events, and secondly, the logic behind how we choose primers. Several candidate primer pairs are tested for a given genomic locus, and all resulting bands are confirmed using Sanger sequencing. Together with an expanded TEChim manual on GitHub, we believe this study gives other research groups a clear and reproducible starting point, while remaining honest that, as with most repeat-adjacent primer design, some locus-specific optimization remains necessary.