Peer review process

Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Xiaohua Shen
    Tsinghua University, Beijing, China
  • Senior Editor
    David Ron
    University of Cambridge, Cambridge, United Kingdom

Reviewer #1 (Public review):

[Editors' note: This revised version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the substantive comments raised in the previous round of review. The qualifications and limitations are now more clearly acknowledged in the revised manuscript.]

Summary:

This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

Strengths:

In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

Comments on previous version:

The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

Reviewer #2 (Public review):

Summary:

Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

Strengths:

(1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

(2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

Weaknesses:

(1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

Author response:

The following is the authors’ response to the previous reviews.

Public Reviews:

Reviewer #1 (Public review):

This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

Strengths:

In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

Weaknesses:

Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

(1) Bias and representations of data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

(2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TErelated mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

(3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated. Comparisons with other noncoding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

(4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in drosophila, worms, mouse, and human. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied, for their functions and corss-species conservations. The authors should explicitly show what is new here in their analyses.

(5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

(6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

(7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

Comments on revisions:

The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

We thank Reviewer #1 for the constructive comments and recognition of the value of our ncORF atlas. We have addressed the key concerns by strengthening the analyses and clarifying the framing and limitations of our conclusions. We appreciate the reviewer’s support for publication and believe these revisions have further improved the manuscript.

Reviewer #2 (Public review):

Summary:

Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

Strengths:

(1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

(2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

Weaknesses:

(1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

(2) Some analytical methods and standards were not clearly presented in the manuscript.

We thank Reviewer #2 for the positive assessment of our comprehensive dataset and standardized analytical framework. We have clarified the analytical methods and criteria throughout the manuscript and better defined the scope and limitations of our bioinformatics-based analyses. We appreciate the reviewer’s constructive suggestions, which have helped improve the clarity and rigor of the manuscript.

Recommendations for the authors:

Reviewing Editor:

We have evaluated the revision together with the reviewers' second-round assessments and your responses. The reviewers agree that the manuscript has improved in clarity and that the standardized integration of large-scale Ribo-seq datasets provides a valuable resource for the field. However, several important concerns remain insufficiently resolved. In multiple cases, the revision relies primarily on acknowledgment or reframing of limitations rather than additional analyses or clearer methodological justification, leaving the evidential support for several conclusions incomplete.

Because the study is entirely computational, all analytical procedures, criteria, and thresholds should be explicitly defined and adequately justified to meet the expected standard of technical rigor. In particular, key components of the analytical framework require clearer description, including the definitions and criteria used for co-translation and ncORF classification. The limitations of the dataset should also be discussed more explicitly, especially regarding the heterogeneity of public Ribo-seq datasets, technical factors influencing detection sensitivity, and the interpretation of variable detection frequencies across samples.

In addition, several conclusions remain largely descriptive, and the distinction between novel findings and confirmation of previous observations should be clarified more carefully. Conclusions should be framed strictly within the limits of the presented data and positioned appropriately relative to prior work, with suitable citation to avoid overstating novelty.

We therefore ask the authors to refine the technical descriptions, ensure that all methods and analytical criteria are presented unambiguously, and expand the Discussion to clearly articulate the limitations of the dataset and analysis. The conclusions should also be revised to reflect an appropriately cautious interpretation of the findings.

The primary strength of this study lies in the scale and standardization of the dataset as a community resource. Given this substantial resource value, we believe the manuscript could become suitable for publication provided that the issues outlined above are addressed clearly and transparently. With these revisions, the work will provide a useful foundation for future studies in this area.

We therefore invite you to submit a final revised version addressing the points described above.

We thank the Editor for the careful assessment and constructive guidance. In the final revision, we have clarified all key methodological definitions and analytical criteria, expanded the Discussion of dataset and detection limitations, and revised the conclusions to avoid overstatement. We believe these changes improve the technical rigor, transparency, and overall value of the manuscript as a community resource.

Reviewer #1 (Recommendations for the authors):

The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

We appreciate the reviewer’s support for publication.

Reviewer #2 (Recommendations for the authors):

While the authors have made commendable efforts to revise the manuscript and address previous concerns, the revised version still falls short of fully resolving several critical issues regarding data interpretation and methodological transparency. I recommend the following modifications before the manuscript can be considered for publication:

(1) Although the authors have annotated the detection frequencies of individual sORFs in the revised

Supplementary Table 3 and Fig. S1B, the biological and technical implications of these data require further clarification:

(a) As the data demonstrates, even the most abundant sORFs were detected in no more than half of the samples. If the authors attribute this low detection rate to technical limitations (e.g., batch effects, sequencing depth, or threshold stringency), this must be explicitly discussed and annotated in the main text to prevent readers from misinterpreting this as low biological penetrance.

We thank the reviewer for this constructive comment. We have added a paragraph to the Discussion explicitly addressing the potential technical factors underlying the variable detection frequencies to avoid overinterpreting these frequencies as biological penetrance.

(b) The authors did not fully address my previous query regarding tissue specificity. Given the diverse and complex origins of the analyzed ribo-seq datasets, it is crucial to know whether any of these sORFs are tissue-specifically translated. The authors should analyze and state whether certain sORFs are exclusively detected in specific sample categories (tissues/organs), and whether this translation pattern aligns with the tissue-specific expression of their corresponding host transcripts.

We thank the reviewer for raising this important point. Strict tissue-exclusive translation is difficult to establish from heterogeneous public Ribo-seq datasets, as gene expression is inherently stochastic and most genes have some probability of being expressed across tissues, although expression levels may vary substantially between tissues. Moreover, failure to detect an ncORF in a given tissue may reflect low expression or insufficient sequencing depth rather than true biological absence. We therefore avoid making definitive claims about tissue-specific translation and instead quantify variation in ncORF expression across tissues using a tissue specificity index.

(2) Echoing the concerns raised by Reviewer #1, I remain concerned that the observed uORF-CDS cotranslation might be a computational artifact or false positive. The current Methods section lacks sufficient detail on how "co-translation" was strictly defined and quantified. I strongly recommend that the authors include a schematic diagram (e.g., in Figure 6 or supplementary figures) that explicitly details their analytical strategy, statistical thresholds, and evaluation criteria for defining co-translation. As was pointed out, it is well-established that uORFs typically exert an inhibitory effect on the translation of the main CDS. To validate the accuracy and robustness of their analytical pipeline, the authors should use their collected dataset to demonstrate the prevalence and nature of this canonical inhibitory phenomenon. Successfully capturing this expected repression would serve as a crucial positive control for their methodology. While the authors provided a theoretically acceptable mechanistic model in the text to reconcile co-translation with uORF-mediated repression, this hypothesis currently lacks literature support. The authors must cite relevant prior studies that support this specific regulatory dynamic to strengthen their argument.

We thank the reviewer for this important comment. We have further clarified the definition, detection criteria, and statistical framework for co-translation in the Methods, and have made the complete analysis code publicly available to facilitate reproducibility and independent evaluation. We have also added relevant literature supporting the proposed regulatory interpretation clarifying the relationship between our observations and the established inhibitory effects of uORFs.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation