A retrospective analysis of 400 publications reveals patterns of irreplicability across an entire life sciences research field

  1. Department of Epidemiology, Gillings School of Global Public Health, University of North Carolina at Chapel Hill, Chapel Hill, United States
  2. Global Health Institute, School of life science, EPFL, Lausanne, Switzerland
  3. Manipal Institute of Regenerative Medicine, Bengaluru, Manipal Academy of Higher Education, Manipal, India
  4. College of Plant Science and Technology, Huazhong Agricultural University, Wuhan, China
  5. Centre for Ecology and Conservation, University of Exeter, Penryn, United Kingdom
  6. Department of Entomology and Plant Pathology, Oklahoma State University, Stillwater, United States
  7. Dalhousie University, Department of Microbiology and Immunology, Halifax, Canada
  8. Department of Human Biology, Faculty of Natural Sciences, University of Haifa, Haifa, Israel
  9. BioInformatics Competence Center of EPFL-UNIL, School of Life Sciences, EPFL, Lausanne, Switzerland

Peer review process

Revised: This Reviewed Preprint has been revised by the authors in response to the previous round of peer review; the eLife assessment and the public reviews have been updated where necessary by the editors and peer reviewers.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Peter Rodgers
    eLife, Cambridge, United Kingdom
  • Senior Editor
    Peter Rodgers
    eLife, Cambridge, United Kingdom

Reviewer #1 (Public review):

Summary:

The authors set out on the ambitious task of establishing the reproducibility of claims from the Drosophila immunity literature. Starting out from a corpus of 400 articles from 1959 and 2011, the authors sought to determine whether their claims were confirmed or contradicted by previous or subsequent publications. Additionally, they actively sought to replicate a subset of the claims for which no previous replications were available (although this set was not necessarily representative of the whole sample, as the authors focused on suspicious and/or easily testable claims). The focus of the article is on inferential reproducibility; thus, methods don't necessarily map exactly to the original ones.

The article is a large-scale analysis of the individual replication findings, which are presented in a companion article (Westlake et al. doi.org/10.1101/2025.07.07.663442). In their retrospective analysis, the authors find that 61% of the original claims were verified by the literature, 7.5% were partially verified, and only 6.8% was challenged, with 23.8% having no replication available. This is in stark contrast with the result of their prospective replications, in which only 16% of claims were successfully reproduced.

The authors proceed to investigate correlates of replicability, with the most consistent finding being that findings stemming from higher-ranked universities were more likely to be challenged - a finding that is observed in a multivariable model and multiple sensitivity analyses. Other trends observed in the initial descriptive analysis, such as less replicability in findings from "trophy" journals and duration of engagement with the field, seem less consistent and fail to reach significance in the multivariable model, and should thus be considered tentative.

Although this is a major contribution to the field, it is important to note that the high replication rates are mostly based on a retrospective survey of the literature, while prospective replication rates in the subset of articles directly replicated by the authors were much lower. Thus, the possibility of biases relating both to the field's decisions about what to replicate (and what to publish) and to the authors' criteria in evaluating claims should be considered when comparing these results with prospective replicability estimates obtained in other fields of science.

Strengths:

The work presents a large-scale, in-depth analysis of a particular field of science that includes authors with deep domain expertise of the field. This is a rare endeavour to establish the reproducibility of a particular subfield of science, and I'd argue that we need many more of these in different areas.

The project was built on a collaborative basis (https://ReproSci.epfl.ch/), using an online database (https://ReproSci.epfl.ch/), which was used to organize the annotations and comments of the community about the claims. The website remains online and can be a valuable resource to the Drosophila immunity community. A companion article details the individual replication findings (Westlake et al. doi.org/10.1101/2025.07.07.663442).

The revised version of the preprint includes a number of sensitivity analyses as supplementary material, which help to evaluate how the main findings change according to analytical decisions.

Main concerns:

(1) Although a number sensitivity analyses have now been conducted in order to address reviewer comments, these have been relegated to the supplementary material. Their existence is acknowledged in the multivariable model section, but the actual results of the main analysis are not mentioned at all in the main text and not taken into account in the discussion. Thus, the casual reader is not made aware of how particular results are dependent on analysis choices.
Although the main finding of the correlation between replicability and university ranking holds over nearly all the analyses conducted, others are not as consistent. The association of irreplicability with trophy journals, for instance, is not statistically significant in the multivariate model (suggesting that it may have been partially due to confounding of journal tier with institutional ranking). It becomes even weaker when unchallenged claims are removed from the analysis (see point 2 below). Similarly, the association of irreplicability with "exploratory" career style is already weak in the univariate analysis and becomes even more marginal in the multivariable model. The effect of the lead author having been a first author in the field, on the other hand, is stronger than that of journal impact factor or that of the continuity/exploratory style (even though it fails to reach formal significance), but seems to receive less attention in the discussion.
These differences in strength of evidence could be made more clear both in the abstract and in the discussion, which seem to prioritize some trends over others when discussing results. Once more, presenting a brief description of how results are affected by the sensitivity analyses in the results would be important to make these nuances clearer to the reader.

(2) I still do not understand the authors' choice of including unchallenged claims along with verified ones in the multivariate model used to investigate predictors of replicability. In particular, the justification given for this choice in the supplementary material (i.e. the fact that a finding remaining unchallenged is not a random event and might be influenced by variables such as journal tier and institutional ranking) actually seems like an additional reason for not including unchallenged claims along with verified ones: if findings from low-tier journals or less prestigious institutions are more likely to remain unchallenged, this could lead to a spurious claim of "less irreplicability" in these journals and institutions, which could be purely due to receiving less attention. Predictably, both associations are reduced when unchallenged claims are removed (although the reduction is much more marked in the journal case) in the sensitivity analysis (whose results, again, are only mentioned in the supplementary material).
Moreover, multiple other statements by the authors themselves concerning unchallenged claims seem to contradict their own decision. In multiple points in the text, they reiterate that remaining unchallenged is associated with less replicability (as suggested by the fact that replicability rates were lower in prospective replications than in the literature, even for non-suspicious claims). The decision to include unchallenged claims along with verified ones thus seems unjustified - if anything, I'd have expected the opposite approach after reading the discussion. That said, I'd argue that the best decision would be to exclude them from the analysis, as done in Figure S4, and presenting this model as the main one.

Reviewer #3 (Public review):

Summary:

The authors of this paper were trying to identify how reproducible, or not, their subfield (Drosophila immunity) was since its inception over 50 years ago. This required identifying not only the papers, but the specific claims made in the paper, assessing if these claims were followed up in the literature, and if so whether they supported or refuted the original claim. In addition to this large manually curated effort, the authors further investigated some claims that were left unchallenged in the literature by conducting replications themselves. This provided a rich corpus of the subfield that could be investigated into what characteristics influence reproducibility.

A major strength of this study is the focus on a subfield, the detailing of identifying the main, major, and minor claims - which is a very challenging manual task - and then cataloging not only their assessment of if these claims were followed up in the literature, but also what characteristics might be contributing to reproducibility, which also included more manual effort to supplement the data that they were able to extract from the published papers. While this provides a rich dataset for analysis, there is a major weakness with this approach, which is not unique to this study.

The main weakness is relying heavily on the published literature as the source of whether a claim was determined to be verified or not. Nonetheless it is understandable why the authors took this approach - it is the only way to get at a breadth of the literature. However, there are many documented issues with this stemming from every field of research - such as publication bias, selective reporting, all the way to fraud. Importantly, this limitation is highlighted throughout the paper. At the same time, it is not reasonable to expect this study to have conducted independent experimental replications for all 400 papers identified as this would have been a laborious and costly effort. To overcome this weakness, the authors leveraged their own expertise, and crowdsourcing within the Drosophila immunity community, to assess the validity of the claims. While this helps mitigate some of this weakness, it does introduce new challenges in interpretation, which is acknowledged in the limitations section of the discussion. Overall, this situates this study more as an expert assessment of the replicability of the published literature opposed to an assessment of independent experimental replications.

The authors should be applauded for the monumental effort they put into this project, which does a wonderful job of having experts within a subfield engage their community to understand the connectiveness of the literature and attempt to understand how reliable specific results are and what factors might contribute to them. This project provides a nice blueprint for others to build from as well as leverage the data generated from this subfield and thus should have an impact in the broader discussion on reproducibility and reliability of research evidence.

Author response:

The following is the authors’ response to the original reviews.

eLife Assessment:

This important study presents an impressive large-scale effort to assess the reproducibility of published findings in the field of Drosophila immunity. The authors analyse 400 papers published between 1959 and 2011, and assess how many of the claims in these papers have been tested in subsequent publications. In a companion article they report the results of experiments to test a subset of the claims that, according to the literature, have not been tested. The present article also explores if various factors related to authors, institutions and journals influence reproducibility in this field. The evidence supporting the claims is solid, but there is considerable scope for strengthening and extending the analysis. The limitations inherent to evaluating reproducibility based on the published literature should also be acknowledged.

We have included in the abstract the limitation of evaluating replicability based on the published literature

Public Reviews:

Reviewer #1 (Public review):

Summary:

The authors set out on the ambitious task of establishing the reproducibility of claims from the Drosophila immunity literature. Starting out from a corpus of 400 articles from 1959 and 2011, the authors sought to determine whether their claims were confirmed or contradicted by previous or subsequent publications. Additionally, they actively sought to replicate a subset of the claims for which no previous replications were available (although this set was not representative of the whole sample, as the authors focused on suspicious and/or easily testable claims). The focus of the article is on inferential reproducibility; thus, methods don't necessarily map exactly to the original ones.

The authors present a large-scale analysis of the individual replication findings, which are presented in a companion article (Westlake et al., 2025. DOI 10.1101/2025.07.07.663442). In their retrospective analysis of reproducibility, the authors find that 61% of the original claims were verified by the literature, 7.5% were partialy verified, and only 6.8% were challenged, with 23.8% having no replication available. This is in stark contrast with the result of their prospective replications, in which only 16% of claims were successfully reproduced. The authors proceed to investigate correlates of replicability, with the most consistent finding being that findings stemming from higher-ranked universities (and possibly from very high impact journals) were more likely to be challenged.

Strengths:

(1) The work presents a large-scale, in-depth analysis of a particular field of science that includes authors with deep domain expertise of the field. This is a rare endeavour to establish the reproducibility of a particular subfield of science, and I'd argue that we need many more of these in different areas.

(2) The project was built on a collaborative basis (https://ReproSci.epfl.ch/), using an online database (https://ReproSci.epfl.ch/), which was used to organize the annotations and comments of the community about the claims. The website remains online and can be a valuable resource to the Drosophila immunity community.

(3) Data and code are shared in the authors' GitHub repository, with a Jupyter notebook available to reproduce the results.

We thank the reviewer for their positive comments.

Main concerns:

(1) Although the authors claim that "Drosophila immunity claims are mostly replicable", this conclusion is strictly based on the retrospective analysis - in which around 84% of the claims for which a published verification attempt was found. This is in very stark contrast with the findings that the authors replicate prospectively, of which only 16% are verified. Although this large discrepancy may be explained by the fact that the authors focused on unchallenged and suspicious claims (which seems to be their preferred explanation), an alternative hypothesis is that there is a large amount of confirmation bias in the Drosophila immunity literature, either because attempts to replicate previous findings tend to reach similar results due to researcher bias, or because results that validate previous findings are more likely to be published.

Both explanations are plausible (and, not being an expert in the field, I'd have a hard time estimating their relative probability), and in the absence of prospective replication of a systematic sample of claims - which could determine whether the replication rate for a random sample of claims is as high as that observed in the literature -, both should be considered in the manuscript.

I still believe that most of ‘verified’ claims are indeed solid, consistent with the feeling of increased knowledge in the Drosophila immunity. Of note, when I started in the field in the early 90, nearly all the articles in Drosophila developmental biology, the dominant field at that time, were solid. I agree that this may represent a past period. Still, perceived irreproducibility is high because affecting claims published in high-impact articles. To address the reviewer’s comments, we have included the possibility of confirmation bias in the discussion of the article.

(2) The fact that the analysis of factors correlating with reproducibility includes both prospective and retrospective replications also leads to the possibility of confusion bias in this analysis. If most of the challenged claims come from the authors' prospective replications, while most of the verified ones come from those that were replicated by the literature, it becomes unclear whether the identified factors are correlated with actual reproducibility of the claims or with the likelihood that a given claim will be tested by other authors and that this replication will be published.

We agree that the ReproSci is a human biased endeavor and may be subjected to bias, notably the one mentioned. Some verified claims might be considered as challenged and some challenged claims as verified. But this is likely not affecting most of our results. Our papers have been posted on a public archive and that the community could scrutinize our data (most authors have checked their claims!). Based on community feedback, we have had to change the status of three claims only: one verified claim became challenged, one challenged claim became unchallenged and one unchallenged became verified. We therefore did not expect major changes in our conclusions. We have provided an estimation of irreproducibility ranging from 10-20% and we believe this is a rather safe interval. This is still significant. The fact that the ReproSci project is a human endeavor, and therefore subject to inherent limitations, is already emphasized in the last section of the discussion.

(3) The methods are very brief for a project of this size, and many of the aspects in determining whether claims were conceptually replicated and how replications were set up are missing.

Some of these - such as the PubMed search string for the publications and a better description of the annotation process - are described in the companion article, but this could be more explicitly stated. Others, however, remain obscure. Statements such as "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category" summarize a very complex process for which more detail should be given - in particular because what constitutes inferential reproducibility is not a self-evident concept. And although I appreciate that what constitutes a replication is ultimately a case-by-case decision, a general description of the guidelines used by the authors to determine this should be provided. As these processes were done by one author and reviewed by another, it would also be useful to know the agreement rates between them to have a general sense of how reproducible the annotation process might be.

The same gap in methods descriptions holds for the prospective replications. How were labs selected, how were experimental protocols developed, and how was the validity of the experiments as a conceptual replication assessed? I understand that providing the methods for each individual replication is beyond the scope of the article, but a general description of how they were developed would be important.

We have extended the first section of the Result to mention that an extended Material and Methods can be found in the complementary article Westlake et al., eLife 2026 and on the ReproSci website indicated the resources described in the companion articles. Disagreement between the two authors were limited and mostly dealt with textual interpretation (the definition of the major claims, which can affect if they are verified or not). As mentioned above, our articles are now public since more 6 months and we got little request to change the status of a claim (only 3 over more than 1000). Of note many claims from some ancient articles are part of the common knowledge, are verified routinely by many labs in the field, and therefore were easy to categorize.

We also added a sentence to also underline that the scientific value of an article does not fully correlate with its status. A number of problematic articles filled with exaggeration or confused claims were still categorized as verified. Thus, we screen for replicability but they are many other publication issues we could not capture in our study. This point is now mentioned in the discussion.

(4) As far as I could tell, the large-scale analysis of the replication results was not preregistered, and many decisions seem somewhat ad hoc. In particular, the categorization of journals (e.g. low impact, high impact, "trophy") and universities (e.g. top 50, 51-100, 101+) relies on arbitrary thresholds, and it is unclear how much the results are dependent on these decisions, as no sensitivity analyses are provided.

Particularly, for analyses that correlate reproducibility with continuous variable (such as year of publication, impact factor or university ranking, I'd strongly favor using these variables as continuous variables in the analysis (e.g. using logistic regression) rather than performing pairwise comparisons between categories determined by arbitrary cutoffs. This would not only reduce the impact of arbitrary thresholds in the analysis, but would also increase statistical power in the univariate analyses (as the whole sample can be used in at once) and reduce the number of parameters in the multivariate model (as they will be included as a single variable rather than multiple dummy variables when there are more than two categories).

We did not pre-registered our study, which is quite different from other reproducibility projects. We did not know at the start how the project would unfold. Of note, I believe that ‘cleaning’ an extensive field by experts was by itself an achievement and represent another way to assess replicability, with its own limitation.

We agree that the cut-off can be seen as arbitrary. However, publication year has been modeled as a continuous variable through a spline (to allow for non-linearity) in the multivariate analysis. For journal and institutional ranking, we would like to defend and keep the categorical specifications in the main text for two reasons:

- Institution ranking: For institutional ranking, the largest institutional category is “Not Ranked” which we cannot assign a rank and ranks institutions individually only through the top 100. Below that it publishes bands (101-150, 151-200, 201-300, 301400, 401-500) of variable width, which are not straightforward to use continuously without making strong assumptions (e.g midpoint) unsupported by the data.

- Prestige: As mentioned by another reviewer there are career consequences attached to Nature, Science and Cell that are different from other journal, and the effect of moving from a journal from an IF 3 to 8 is different than from 30 to 35. We provide a sensitivity analysis in the supplementary information where we refit the multivariate model with log-transformed impact factor continuous predictor. We obtain very similar results than these in the main text, with the exception that the impact factor of the journal becomes significant. We describe and discuss this results in the supplementary information (section S5), while justifying why we keep the categorical model in the main text.

We are grateful to the reviewer for prompting these additions, which we agree strengthen the paper.

The text added in the SI states:

“As a sensitivity analysis, we refitted the multivariable hierarchical logistic model by replacing journal-impact-factor categories with log₂-transformed continuous impact factor, while retaining all other covariates and random intercepts from the primary analysis. University ranking was retained as a categorical variable because the 2010 Shanghai Academic Ranking reports individual ranks only for institutions ranked 1– 100 and otherwise reports rank bands (101–150, 151–200, 201–300, 301–400, and 401–500); moreover, 441 of 1,006 claims (43.8%) were from unranked institutions, which constitute a distinct group.

Each doubling of journal impact factor was associated with 1.51-fold higher odds that a claim was challenged (OR 1.51, 94% HDI 1.13–2.08). At the mean impact factors of the low-impact and trophy-journal tiers (4.4 and 61, respectively; approximately 3.8 doublings), the continuous model implies a trophy-versus-low-impact odds ratio of approximately 4.8. By contrast, the primary model directly estimated the trophy-versus-low-impact odds ratio as 1.75 (94% HDI 0.59–4.91). Although this wide interval includes 4.8, the point estimates differ between the two models. We interpret the larger implied trophy-journal effect in the continuous model cautiously because it assumes that the log-linear slope estimated predominantly from the many journals with lower impact factors (around 2–12) applies unchanged to the sparsely represented trophy journals (62 claims across Science, Nature, and Cell). We therefore retained the categorical journal analysis as the primary analysis because it directly estimates the trophy-journal comparison using claims published in those journals. Estimates for all other covariates were broadly similar to those in the primary model. In particular, the association with Top 50 university affiliation remained evident (OR 4.04, 94% HDI 1.57–10.08, compared with OR 3.68, 94% HDI 1.50–9.28 in the primary model)”.

[This response is referred later as “Response to categorical variable encoding”]

(5) The multivariate model used to investigate predictors of replicability includes unchallenged claims along with verified ones in the outcome, which seems like an odd decision. If the intention is to analyze which factors are correlated with reproducibility, it would make more sense to remove the unchallenged findings, as these are likely uninformative in this sense. In fact, based on the authors' own replications of unchallenged findings, they may be more likely to belong the "challenged" category than to the "unchallenged" one if they were to be verified.

We thank the reviewer for raising this comment, we agree that the classification of the unchallenged claims as not challenged is difficult given their nature. We refitted the multivariable analysis excluding unchallenged claims, with a classification of Challenged and Unchallenged (Verified, Partially Verified, Mixed). The results are shown in the supplementary section S4 and are extremely similar to the main model.

The text added in the SI states:

“We fit the multivariate model from the main text while excluding unchallenged claims from the analysis. The methods are similar and the convergence and diagnosis checks are good and not reported here for brevity. The model simultaneously adjusts for author demographics, laboratory attributes and journal or institutional prestige and predict a binary outcome for each claim: challenged or not challenged (including Verified, Partially Verified, Mixed). The re-analysis included 659 claims, of which 60 (9.1%) were challenged. The non-challenged reference group comprised of Verified (n=525), Partially Verified (n=64), Mixed (n=10). The Adjusted odds ratios and 94% highest-density intervals are reported in Figure S4. The results are very similar to the analysis with the unchallenged claims in the main text, with only university ranking (Top 50) being a significant predictor (OR 3.34 (1.28–9.01) vs the main text OR 3.68 (94% HDI 1.50– 9.28). The other variable with the biggest changes shows only modest differences:

- High-impact journal OR 1.12 (vs 1.31)

- Trophy journal OR 1.47 (vs 1.75)

- Post-doc OR 1.44 (vs 1.16) and their intervals still include 1. Note that the excluding unchallenged decision condition on follow-up, which is not random and might be influenced by our key variable (journal tier, institutional ranking) and bias results for this model.”

However, we did not replace the main model of the article by this particular model because when you chose only claim that are either verified or challenged (so you exclude unchallenged), the condition that made these claims included in this analysis could be influenced by a cofactor, so our odds ratios might be wrong.

[This response is referred later as “Response to unchallenged exclusion”]

Reviewer #2 (Public review):

Summary:

Lemaitre et al. conducted an analysis of 400 publications in the Drosophila immunity field (1959-2011), performing both univariable and multivariable analyses to identify factors that correlate with or influence the irreproducibility of scientific claims. Some of the findings are unexpected, for instance, neither the career stage of the PI nor that of the first author appears to matter that much, while others, such as the influence of institutional prestige or publication in "trophy journals," are more predictable. The results provide valuable insight into patterns of irreproducibility in academia and may help inform policies to improve research reproducibility in the field.

Strengths:

This study is based on a large, manually curated dataset, complemented by a companion paper (Westlake et al., 2025. DOI 10.1101/2025.07.07.663442) that provides additional details on experimentally documented cases. The statistical methods are appropriate, and the findings are both important and informative. The results are clearly presented and supported by accessible documentation through the ReproSci project.

We thank the reviewer for their assessment.

Weaknesses:

The analysis is limited to a specific field (immunity) and model system (Drosophila). Since biological context may influence reproducibility -- for example, depending on whether mechanisms are more hardwired or variable -- and the model system itself may contribute to these effects (as the authors note), it remains unclear to what extent these findings generalize to other fields or organisms. The authors could expand the discussion to address the potential scope and limitations of the study's generalizability.

We have added a limitation section at the end of the manuscript to discuss these points. We wrote:

“Finally, the findings we report may be specific to a particular field, model organism, and time period, and may be strongly influenced by factors idiosyncratic to their development. Because this analysis is limited to a single field (immunity) and model system (Drosophila), it remains unclear to what extent these findings generalize to other fields, organisms, or more recent studies.”

However, we believe that our study adds on the discussion on reproducibility by providing another approach. Our study also aligns with the feeling that knowledge of Drosophila innate immunity has increased so much in the last decades, and that science is still quite robust in many fields.

Reviewer #3 (Public review):

Summary:

The authors of this paper were trying to identify how reproducible, or not, their subfield (Drosophila immunity) was since its inception over 50 years ago. This required identifying not only the papers, but the specific claims made in the paper, assessing if these claims were followed up in the literature, and if so whether the subsequent papers supported or refuted the original claim. In addition to this large manually curated effort, the authors further investigated some claims that were left unchallenged in the literature by conducting replications themselves. This provided a rich corpus of the subfield that could be investigated into what characteristics influence reproducibility.

Strengths:

A major strength of this study is the focus on a subfield, the detailing of identifying the main, major, and minor claims - which is a very challenging manual task - and then cataloging not only their assessment of if these claims were followed up in the literature, but also what characteristics might be contributing to reproducibility, which also included more manual effort to supplement the data that they were able to extract from the published papers. While this provides a rich dataset for analysis, there is a major weakness with this approach, which is not unique to this study.

Weaknesses:

The main weakness is relying heavily on the published literature as the source for if a claim was determined to be verified or not. There are many documented issues with this stemming from every field of research - such as publication bias, selective reporting, all the way to fraud. It's understandable why the authors took this approach - it is the only way to get at a breadth of the literature - however the flaw with this approach is it takes the literature as a solid ground truth, which it is not. At the same time, it is not reasonable to expect the authors to have conducted independent replications for all of the 400 papers they identified.

However, there is a big difference trying to assess the reproducibility of the literature by using the literature as the 'ground truth' vs doing this independently like other large-scale replication projects have attempted to do. This means the interpretation of the data is a bit challenging.

We acknowledge that our analysis necessarily relies on the published literature; however, two important points should be considered. First, we assessed replicability more than a decade after the original studies, allowing us to evaluate whether their claims have withstood subsequent scrutiny. Second, the evaluation draws on the authors’ strong expertise in the field. Accordingly, the validity of published claims was not accepted uncritically but subjected to informed expert assessment. This is exemplified by the case of the proposed microbicidal role of Duox, which we classified as challenged despite the original article accumulating more than 600 citations.

Overall, our replicability analysis, like similar efforts, relies not only on literature but also on expert judgment. Of course it has inherent limitations, which are now more clearly articulated in the discussion.

Below are suggestions for the authors and readers to consider:

(1) I understand why the authors prefer to mention claims as their primary means of reporting what they found, but it is nested within paper, and that makes it very hard to understand how to interpret these results at times. I also cannot understand at the high-level the relationship between claims and papers. The methods suggest there are 3-4 major claims per paper, but at 400 papers and 1,006 claims, this averages to ~2.5 claims per paper. Can the authors consider describing this relationship better (e.g., distribution of claims and papers) and/or considering presenting the data two ways (primary figures as claims and complimentary supplementary figures with papers as the unit). This will help the reader interpret the data both ways without confusion. I am also curious how the results look when presented both ways (e.g., does shifting to the paper as the unit of analysis shift the figures and interpretation?). This is especially true since the first and last author analysis shows there is varying distribution of papers and claims by authors (and thus the relationship between these is important for the reader).

Articles can contain claims that we found replicable and other that we did not. Thus, we believe than an evaluation by claim is more precise. The extraction of the different claims relies on the reading of the article, but is a human enterprise. There are clearly some unavoidable biases at this step. Of note, the articles analyzed in this study are very different one from another (size, novelty, amount of data…).

However, we agree that the analysis at the article level is interesting, and we followed the reviewer suggestion and provided in Supplementary Section S6 a sensitivity analysis of using article-level effects for the multivariate analysis. We note that because only 43 articles (12.5%) contained at least one challenged claim, this analysis is less robust than the claim analysis. The text added in the SI states:

“Two claims in the same paper might not be independent, so we might want to incorporate an article random effect to the primary Bayesian formulation. However, due to the low number of claim per article (mostly one to four) and only 43 articles (12.5%) contained at least one challenged claim (and most articles contains only one to four claims), we could not manage to fit this model properly (we observe MCMC divergences). As a work-around sensitivity analysis, we repeated the analysis with the article as the unit. Each of 345 articles with complete covariate data (representing 869 major claims, 60 of the 69 challenged claims in the full dataset; the remaining 9 fell in articles excluded for missing covariates) contributed as outcome the proportion of its major claims classified as challenged. The model was a (frequentist) binomial-link fractional-response generalized linear model with HC3 standard errors (the intervals reported here are 94% confidence intervals rather than posterior credible intervals).

Each article contributed equally regardless of its number of claims. The Top-50 institutional effect remained significant under both analysis (OR 5.15, 94% CI 2.18– 12.15; versus a claim-level OR 3.68, 94% HDI 1.50–9.28). In this new model Trophy journals also showed had higher odds of challenged claims than low-impact journals (OR 2.90, 94% CI 1.24–6.78 vs a claim-level OR 1.75, 94% HDI 0.59–4.91). This impact was not significant in the primary model but became significant at using article-level analysis. Other estimates are consistent in direction and magnitude with the primary claim-level model. Model diagnostics were satisfactory.”

Moreover, we added Figure S6 which display visually the relationship between the claims and the article.

[This response is referred later as “Response to claim vs article level analysis”]

(2) As mentioned above, I think the biggest weakness is that the authors are taking the literature at face value when assigning if a claim was validated or challenged vs gathering new independent evidence. This means the paper leans more on papers, making it more like a citation analysis vs an independent effort like other large-scale replication projects. I highly recommend the authors state this in their limitations section.

We have now explicitly acknowledged this limitation in a dedicated section at the end of the revised manuscript. However, as noted above, we do not take the literature at face value. Notably, some of the claims we identify as challenged originate from highly cited articles, which could lead a naïve reader to assume their validity.

Ultimately, this replicabilty project remains a human, expert-driven effort aimed at determining which claims withstand scrutiny within a field. The broader community was invited to engage with and respond to these assessments and we did get some feedbacks. Importantly, claims were not evaluated in isolation but in the context of the entire body of knowledge, and all evaluations were made transparent and explicitly justified in the ReproSci database.

On top of that, I have questions that I could not figure out (though I acknowledge I did not dig super deep into the data to try). The main comment I have is How was verified (and challenged) determined? It seems from the methods it was determined by "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category". If this is true, and all claims were done this way - are verified claims double counted then? (e.g., an original claim is found by a future claim to be verified - and thus that future claim is also considered to be verified because of the original claim).

When two articles reported the same main claim and this claim was judged as “verified,” each instance was counted in the analysis. However, when validation came from studies published after 2011, the claim was counted only once. Thus, our study specifically evaluates the replicability of claims originating from a defined period (pre-2011).

We note that, in principle, repeated publication of the same finding across many articles could artificially inflate estimates of replicability. However, such cases appeared to be rare in practice, and studies reporting similar claims typically provided complementary evidence—for example, elucidating gene function through either genetic or biochemical approaches.

Related, did the authors look at the strength of validation or challenged claims? That is, if there is a relationship mapping the authors did for original claims and follow-up claims, I would imagine some claims have deeper (i.e., more) claims that followed up on them vs others. This might be interested to look at as well.

We did not perform a systematic mapping between original claims and subsequent follow-up studies. While this is an important and interesting aspect, it was considered beyond the scope of the present work. The current study already represents a substantial investment of time and effort, and the resulting database provides a valuable resource for future analyses. In particular, it could support a dedicated bibliometric study at a later stage.

(3) I recommend the authors add sample sizes when not present (e.g., Fig 4C).

The cited panel appears to correspond to Figure 5C in the revised numbering, as the current Figure 4 contains only panels A and B. We systematically checked the denominators of all figures and revised the captions to report the numbers of claims, authors or laboratories represented, together with the reasons for exclusions. We also added the numbers of authors and claims directly within the author-level scatter plots, with the observational unit explicitly identified.

I also find that the sample sizes are a bit confusing, and I recommend the authors check them and add more explanation when not complete, like they did for Fig 4A. For example, Fig 7B equals to 178 labs (how did more than 156 labs get determined here?), and yet the total number of claims is 996 (opposed to 1,006).

There are 156 labs total, 23 labs contributed publications during both junior and senior stages, and 11 claims lacked a junior/senior classification. This has been explicited in the caption.

Another example, is why does Fig 8B not have all 156 labs accounted for?

The panel does not include all 156 laboratories because we could obtain information about first author training for only 146 PI and laboratories while prior first-author status was unavailable or indeterminate for ten PI. This has been added to the caption.

(Related to Fig 8B, I caution on reporting a p value and drawing strong conclusions from this very small sample size - 22 authors).

The 22 authors are not the total sample size; they constitute the subgroup of PIs who had previously published a first-author paper on Drosophila immunity in another laboratory. In the submitted panel, they were compared with 117 PIs without this training history, giving 139 classified laboratories and 774 claims.

As a last example, Fig 8C has al 156 labs and 1,006 claims - is that expected? I guess it means authors who published before 1995 (as shown in Figure 8A continued to publish after 1995?) in that case, it's all authors? But the text says when they 'set up their lab' after 1995, but how can that be?

Thank you for identifying this inconsistency. Figure 8C intentionally includes the full cohort of 156 PIs and all 1,006 claims. Panel C compares research styles and does not impose a restriction based on year of entry into the field or year of claim publication. Thus, PIs who entered the field before 1995 are included together with all claims attributed to them; the panel is not restricted to claims published after 1995.

The statement that only PIs who established their laboratories after 1995 were included was an erroneous carryover from an earlier version of the analysis. We removed this restriction from the Results and revised the caption to state that exploratory PIs (n = 91) contributed 366 claims, whereas continuity PIs (n = 65) contributed 640 claims.

We also replaced references to the date when PIs “established their laboratory” with “year of entry into the field,” defined as the year of their first Drosophila-immunity publication as either first or last author. During the same check, we revised Panel B to include all 146 PIs eligible for the pre-PI training comparison, contributing 793 claims.

(4) Finally, I think it would help if the authors expanded on the limitations generally and potential alternative explanations and/or driving factors. For example, the line "though likely underestimated' is indicated in the discussion about the low rate of challenged claims, it might be useful to call out how publication bias is likely the driver here and thus it needs to be carefully considered in the interpretation of this. Related, I caution the authors on overinterpreting their suggestive evidence. The abstract for example, states claims of what was found in their analysis, when these are suggestive at best, which the authors acknowledge in the paper. But since most people start with the abstract, I worry this is indicating stronger evidence than what the authors actually have.

We have added a dedicated section on limitations at the end of the Discussion to explicitly acknowledge key constraints, and we also address these points throughout the Discussion where relevant. In addition, these limitations are now in the abstract.

The authors should be applauded for the monumental effort they put into this project, which does a wonderful job of having experts within a subfield engage their community to understand the connectiveness of the literature and attempt to understand how reliable specific results are and what factors might contribute to them. This project provides a nice blueprint for others to build from as well as leverage the data generated from this subfield, and thus should have an impact in the broader discussion on reproducibility and reliability of research evidence.

Thank you for this positive assessment!

Recommendations for the authors:

Reviewer #1 (Recommendations for the authors):

Comments and recommendations are provided in the order that they appear in the manuscript.

Introduction:

- In the second paragraph of the introduction, I don't think all the references refer to "molecular life sciences" (e.g. one of them is from ecology and evolution, for example). I'd also argue that Macleod et al. 2014 refers to risk of bias rather than to reproducibility per se.

We have corrected this in the revised version.

- The author use both "reproducibility" and "replicability", apparently as synonyms, but as definitions for these words vary across sources (e.g. https://arxiv.org/abs/1802.03311, https://www.nationalacademies.org/our-work/reproducibility-and-replicability-in-science, https://osf.io/br9sp/) it may be worth stating their definitions upfront.

Thank you for the references. We agree that the literature contains multiple, sometimes conflicting definitions of these terms, which can lead to confusion. In response to the reviewer’s comment, we have clarified our terminology. Specifically, we now use the term “conceptual replicability” in the abstract and consistently refer to “replicability” and “irreplicability” throughout the manuscript. We hope this revised terminology improves clarity.

- Why are analysis described as "exploratory" and "multivariate"? If it has not been preregistered, shouldn't the multivariate analysis be considered exploratory as well?

Thanks, we agree. The word “exploratory” is confusing as we use it for its structural meaning and not its epistemological meaning. We have replaced it throughout by “descriptive”, as to not imply that the multivariate analysis is preregistered by contrast.

Results:

Methodology and claim assessment

- The methods for the selection, annotation and classification of studies are described very cursorily here, and the Methods section doesn't add that much, but I'll keep these observations to the methods. Still, it's worth mentioning that it's hard to know how reproducible the annotation and classification process is from the information provided.

We have provided more information on the method section on the way to annotate and class claims. Of note, there is more information on the methods on the associated publication Westlake et al., 2025 and on the ReproSci website.

- "Then, experimental work was performed in several laboratories...". Once again, a much better description of how replications were set up and performed is needed.

We have provided information on the number of laboratories (nine) involved in the reproduction of 45 major claims. The laboratories involved in replication are listed in the author list. All the files associated with experimental replication can be found in the supplementary materials of Westlake et al 2025 or the ReproSci website.

- "These claims subjected to experimental validation were selected according to criteria such as appearing suspicious due to the absence of direct follow up or being straightforward to test experimentally". This is rather vague, but if no explicit criteria were set upfront this may be inevitable. Still, it would be useful to know who made this decision.

A first list of claims to be replicated was made by Hannah Westlake and Bruno Lemaitre. The claims that were selected to be replicated were tested by the labs and experimental workers that had the expertise required to test them. The choice of the claims to be tested was not random. Claims that were suspicious are claims that were already tested in the host laboratory without success, claims whose validity are debated at meeting, claims from highranked journal with no follow-up. This is now indicated in the method of the revised version.

Drosophila immunity claims are mostly reproducible

- The percentages given in the initial description of the results pool the results of the retrospective analysis of the literature and the prospective replications performed. As mentioned in the Public Review, I think these are rather distinct ways to assess reproducibility, and would recommend that these two analyses are described separately.

We agree with the reviewer but we believe that we have already well separated the results before and after experimental work in the revised version. This is clearly shown in Table 1. Our text says ‘Importantly, 6.8% of claims (69 out of 1006) were challenged. Among these, 44.9% (31 out of 69) were contested by published articles including 7.2% by the same authors, and 55.1% (38 out of 69) were challenged experimentally as part of the ReproSci project.’

- "These results may reflect the robust scientific standards and methodological rigor of Drosophila research...". Once again, as mentioned in the Public Review, I think publication bias is a plausible explanation, so this should be mentioned as an alternative.

We have mentioned publication bias in the revised version in the section on limitation. But an important point to underline is that this analysis did not take publication results as face fact. This is a critical analysis by experts in the field. As already mentioned, some claims like ‘Duox produces microbicidal ROS” ‘NO is a signaling molecule…’that have been mentioned in multiple articles were considered as challenged.

A second point is that it is likely that publications in some fields are more replicable than in other. Actually, we did not discuss ‘reproducibility issues’ when I started in the field 30 year ago because this was not perceived as an issue. It is normal that reproducibility rate varies according to fields. Nevertheless, we do find important claims with high visibility that are not replicable in our sample. Thus, our findings nuance the ‘reproducibility crisis narrative’ but we still believe that there is concern with reproducibility in life science. Last point, the criterium we have used in our study did not map all the problematic papers, notably some claims might be fundamentally verified but exaggerated in the original paper.

- It would be useful to break down the "partially verified" category somewhere. Illustrative examples of what is meant by "insightful data were accompanied by incomplete interpretations" or "incomplete data were paired with insightful interpretations" would be useful as well.

All the claims are listed in the public database with an explanation as to why they got affected to a category. The categories mixed and partially verified refer to claims that are more complex to categorize.

Here are are two examples from the ReproSci website:

Claim: Relish Rel mutants are very susceptible to bacterial and fungal infection. Partially verified

Assessment: Relish mutants are primarily susceptible to Gram-negative bacterial infections, and do not typically have increased susceptibility to Gram-positive bacteria and fungi which are primarily rebuffed by the Toll pathway, although Imd signaling may provide a minor contribution to defense against these. The high susceptibility of Relish mutant to fungi observed in this article is likely due to the presence of the ebony marker (Lemaitre et al., 1996) [here only one part of the statement has been verified]

Claim: PGRP-SC1a is required for Lys-type peptidoglycan recognition and Toll activation. Mixed

Assessment: A possible role for PGRP-SCs in Toll activity has been strongly debated. Challenged by (Bischoff et al., 2006) who found no effect of PGRP-SC RNAi on activation of Toll (Drs expression) in adult flies (although note that their data show a minor reduction of Drs expression in flies in response to E. faecalis). Supported by (Costechareyre et al., 2016) who used single -SC1 and -SC2 mutants to show that Toll activation (Drs expression) was reduced in -SC1 but particularly -SC2 mutants (~50%) in response to E. faecalis (although survival was not affected). The supplementary data of (Paredes et al., 2011) similarly show that Drs expression was reduced (50%) in response to M. luteus septic infection in PGRP-SCdelta flies (PGRP-SC1A/B -SC2 triple mutant), but concluded that this was due to a secondary effect of Imd overactivation. Note that the phenotype found by (Garver et al., 2006) was much stronger (complete ablation of Drs expression), which is not consistent with any subsequent results and argues for a secondary mutation in this line, although this is not consistent with the successful rescue of Drs expression by transgenic replacement of PGRP-SC1a. Downregulation of Toll in these mutants could be explained if the Toll pathway requires amidase activity of PGRP-SC for immunogenicity (e.g. the sugar backbone without stem peptides is a stronger elicitor than whole peptidoglycan), whereas similar cleavage reduces recognition by the Imd pathway (consistent with the demonstrated requirement of stem peptides for full stimulation of the Imd pathway (Chang et al., 2005; Stenbak et al., 2004)). See annotation for (Mellroth et al., 2003). [here the claim is mostly challenged but there are observations that go in the same directions]

As shown by these examples, category assessment was complex process and relied on human expertise. However, there was very little contestation from the community after the submission of our articles and the opening of the ReproSci website.

A significant fraction of unchallenged claims is non-reproducible

- Excluding the 45 unchallenged major claims that were experimentally tested, we categorized the remaining 240 unchallenged claims into three groups". No information is provided on that classification process (either here or in the methods). Who made these assessments (which seem quite subjective and dependent on field expertise), and how do you know if such judgments are reproducible across different evaluators?

This categorization was done by Hannah Westlake and reviewed by Bruno Lemaitre with few disagreements. This is now indicated in the revised version. The category ‘unchallenged logically consistent” and ‘‘unchallenged logically consistent’ refer to claims that although not directly verified are corroborated or in with current literature. See ReproSci for justifications.

Example from ReproSci website:

Claim: The minimal structure needed to activate the Toll pathway is a muropeptide dimer.

Assesment: Unchallenged logically consistent This is consistent with (Park et al., 2007) who show that on a linearized strand of peptidoglycan, the minimal motif is at least 3 dimers. But it is expected that a non-linearized dimer linked by the peptide bridge could serve as the minimal motif by clustering peptidoglycan.

Claim: eater null flies are susceptible to oral infection with Serratia marcescens Unchallenged logically consistent Inconsistent with the observation that Eater mutants successfully phagocytose Serratia marcescens (Bretscher et al., 2015). This may indicate that another factor affected by eater mutation is required for defense against S. marcescens, such as formation of lamellopodia and filopodia or adhesion of hemocytes to the body wall, or that Eater contributes to binding of Gram-negative bacteria but does not trigger phagocytosis in response to them as it does for Gram-positive bacteria.

Higher representation of challenged claims in trophy journals and from top universities

- Both impact factor and the Shanghai university ranking are continuous variables, but the authors opt to use them as categorical variables in the analysis (i.e. "low-impact", "highimpact", "trophy"; "Top 50", "51-100", "101+"). While there may be legitimate reasons to do this if they feel that these categories are a better descriptor of the underlying reality (e.g. perhaps Cell, Science and Nature are indeed in a category by themselves), it is an unusual decision that leads the comparison to ignore the distinctions between journals/universities within a category. Moreover, it also opens up the opportunity to analysis bias as categories can be set up in many different ways. Thus, if the analysis was not preregistered, I'd recommend that analyses using impact factor and ranking as quantitative variables are added as sensitivity analyses, as these seem to me to be the most natural/less ad hoc way to look at the issue.

See “Response to categorical variable encoding” above.

- As mentioned in the Public Review, the analysis here pools the results from the retrospective analysis based on the literature and the prospective one based on the performed replications. Although this may be justified from a sample size perspective (as the number of challenged claims is not that high), it would be useful to perform this analysis separately on the two sets of results as well, as different trends may be noticed. There are many biases that might come into play here (e.g. results from top institutions being more or less likely to have challenged/unchallenged claims published) and looking at the results separately could help in tearing apart these hypotheses.

The ReproSci project analyzes all the papers of a community during a period of time. As such, the total number of claims is limited and we believe that pooling the data was justified. We provide as sensitivity analysis in Supplement S7, a multivariate model with claims in their pre-experimental classification state:

“As a sensitivity analysis, we refitted the multivariable hierarchical logistic model using only the claim classifications available before experimental validation by the ReproSci project, as the choice of which claims to further validate could have been influenced by covariate (impact factor, university status). The 45 claims tested prospectively were therefore restored to their original “Unchallenged” classification, while all covariates and author-level random intercepts were retained from the primary analysis. The complete-case analysis included 869 claims, of which 28 (3.2%) were challenged. No predictor had a 94% highest-density interval excluding 1, including publication in trophy journals (OR 1.11, 94% HDI 0.33–3.77) and affiliation with a Top50 institution (OR 1.44, 94% HDI 0.48–3.94). Thus, the associations observed in the pooled analysis were not evident when using classifications based exclusively on the retrospective literature review, although estimates were imprecise because of the small number of challenged claims”

The irreproducibility rate has increased over time as the field has grown in popularity

- Again, why not use year as a continuous variable rather than using 5-year windows (as this would effectively include more information)? I think the categorization here is less ad hoc than in the journal/university case, but it is still an unusual decision. More important than this, however, is the fact that pairwise comparisons between periods is probably not an appropriate strategy to look for a time trend, as it excludes all information not related to the pair in each analysis. I'd strongly suggest substituting this for a straightforward regression with publication year as a continuous variable.

We thank the reviewer for this important suggestion. We agree that categorizing publication year into five-year periods may obscure temporal trends.

We therefore modelled publication year continuously using a spline in the multivariable analysis, thereby retaining the full temporal information while allowing for non-linearity. A straightforward linear regression would impose a linear trend, which may not adequately capture changes over time. We retained the five-year groupings solely for visualization.

- "The subsequent increase in unchallenged claims may reflect the rapid conceptual expansion of the field, which likely outpaced the growth in the number of researchers". Do we have data on the growth in the number of researchers in the field? If so, it might be interesting to cite this.

We do not have the number of researchers to assess the growth of the field, however we can see in Figure 8A an increase in the number of laboratories (as identified by last author name) after 1995. This together with the increase in the number of articles (Figure 3A) clearly show the expansion on of the field.

- Isn't a "rise in unchallenged claims" over time expected by chance, as older findings will have more time to have been verified by someone else? I understand that the 14-year window probably mitigates this effect, but it still probably exists in some degree and should be mentioned.

The reviewer is correct that the % unchallenged claims should increase over time, but this does not explain the very low number of unchallenged claims between 1992-2001.

First-author patterns of irreproducibility

- With 69 challenged claims and 289 first authors, it would be impossible to achieve "perfect equality" (e.g. a Gini index of 0.88). To verify how far the observed coefficient deviates from chance, it would be useful to obtain (perhaps via simulations) the Gini coefficient expected by chance given the number of challenged claims/authors, and perhaps derive a p value as well (e.g. the proportion of random permutations in which the index exceeds the observed value).

Very good observation, thank you, we provide this analysis for both first and last author, in the main text and in a supplementary table.

Added first author text:

“The distribution of challenged claims among first authors was highly unequal with a Gini coefficient of 0.881 (inequality index ranging from 0 -perfect equality- to 1 extreme inequality-, Figure 4B): the top 10% of first authors accounted for 73.9% of challenged claims, and the top 20% accounted for all challenged claims. This inequality exceeded that expected by chance (reassigning randomly the 69 challenged labels across individual claims, while preserving each author’s number of claims, p = 0.000001). It also remained greater than expected when claims from the same paper were kept together (p = 0.015; Supplementary Table Sx). The concentration of challenged claims among first authors cannot be explained solely by differences in the number of claims or by clustering within papers.”

Added last author text:

“Challenged claims were also unevenly distributed among leading authors (Gini coefficient = 0.856): the top 10% accounted for 71.0% of challenged claims and the top 20% accounted for 94.2%. The observed inequality exceeded that expected when challenged labels were reassigned across individual claims (p = 0.00081), but not when claims from the same paper were kept together (p = 0.127; Supplementary Table S9). The apparent concentration among leading authors could be explained by multiple challenged claims arising from the same papers.”

Added method text:

“We compared the observed Gini coefficients with two permutation analyses, each based on 1,000,000 permutations. First, we randomly reassigned the challenged labels across individual claims while keeping every claim attached to its original author. This preserved the number of claims contributed by each author. Second, we kept the challenged-claim pattern of each paper together and reassigned these patterns among papers containing the same number of claims. This additionally accounted for the possibility that claims from the same paper were challenged because of a shared problem. For each permutation, we recalculated the Gini coefficient across authors and calculated the p-value as the proportion of simulated coefficients that equalled or exceeded the observed coefficient. Full results are reported in Supplementary.”

- Claims in a single paper may not be fully independent from each other (as both could be irreproducible due to the same error), so it could be worth adding the article as a random variable here.

See answer above: [“Response to claim vs article level analysis”]

- In Fig. 4B, what defines the order of the dots between 0 and 250 (i.e. authors that have no challenged claims)? Is this merely arbitrary? This should be stated more clearly.

Thank you, the order was arbitrary, it has now been fixed by using a secondary sorting (verified claim proportion), and Authors tied on challenged-claim proportion are ordered by verified-claim proportion has been added the legend.

- In Fig. 5C, there are clearly less dots than there are authors. This is likely due to superposition (i.e. there are probably multiple circles with 1 article and 0% verified claims). Nevertheless, it means that the graph conveys an erroneous message. As the x axis is a discrete variable, it may be worth turning it into five categories (e.g. 1 to 5) and adding some jitter to each of them in order to let the reader know how many dots are in each category/%.

Thank you, this was indeed due to superposition. After trying to add a jitter, the graph was still confusing, so we changed the representation to show circle whose size depends on the number of authors. The graph now conveys the size of each block properly.

Lead-author patterns of irreproducibility

- The same comment concerning the Gini index made for the first authors also holds here.

We added a text, see answer above.

- In Fig. 6B, the same comment made for figure 4B also holds.

Thank you, the order was arbitrary, it has now been fixed by using a secondary sorting (verified claim proportion), and Authors tied on challenged-claim proportion are ordered by verified-claim proportion has been added the legend.

- The division between "senior" and "junior" PIs is rather ad hoc here. Once more, why not use "years from first last-author publication" as a continuous variable as a more neutral way to analyze this (as done in Fig. 8)? Also note that it is not clear whether "published a lastauthor article at least five years prior to the considered publication" means any article or one about Drosophila immunity (as in Fig. 8A), and which database was used to examine this (PubMed? Other?).

Senior and junior researchers were classified based on whether they had published a last-author article five years or more using Pubmed (for seniors). We did not specify articles in Drosophila immunity when considering seniority but only having a last author article five year before the publication.

- In Fig. 7C, the same comment made for Fig. 5C also holds, although the solution here is less obvious as there are more possible numbers for "number of articles".

Thank you. As in Fig. 5C, superposition obscured multiple authors occupying the same coordinates. We therefore revised Fig. 7C so that authors with the same number of articles and proportion of challenged claims are represented by a single circle. Circle size and the number shown inside indicate how many authors are represented.

- As far as I could tell, the data on Fig. 8A refers to when authors published their first first/last author papers on Drosophila immunity (which makes it different from Fig. 7, which refers to any article, but I could be mistaken). If this is the case, I'm not sure it's correct to talk about "the period when principal investigators established their laboratories" as mentioned in the text, as (a) they could have published a first author paper in somebody else's lab or (b) they could have started a lab and only later published a paper on Drosophila immunity. "Year of entry in the field" as in the figure legend seems more appropriate.

We have changed for ‘year of entry in the field’ as suggested by the reviewer

- In the same figure, the 1995 cutoff seems completely arbitrary. Why not analyze this as a regression with year of entry as a continuous variable?

The 1995 boundary is historical as it marks the expansion of Drosophila immunity from a marginal subject into a popular field. The expansion of the field at this specific time was probably driven by a combination of factors: major advances in Drosophila genetics and genomic resources and growing knowledge of antimicrobial peptides in the early 1990, a strong interest on innate immunity that drew attention on the power of Drosophila to answer key questions in this new area. These developments attracted new researchers to insect immunity, and the later discoveries of the IMD and Toll pathways in 1995 and 1996 gave the field an additional (and probably even stronger) boost. Our corpus holds 29 articles for 1959– 1994 (0.8 per year) against 371 for 1995–2011 (21.8 per year), and only 13 of the 156 PIs entered before 1995. The comparison is therefore a cohort contrast more than a contrast or search for a cutting point. Moreover, a regression would impose a linear relationship between the different years, which we believe is too constrained on regard of the observed data.

- "We hypothesized that these authors, having gained prior hands-on experience on Drosophila immunity, would be less prone to publishing irreproducible claims.". This is a possibility, but given that most of the sample is retrospective, it's also possible that the field is more prone to publicly challenging findings from newcomers.

Since we do not observe major difference in replicability between senior and junior PI, we still believe that training as first author in a traditional immunity laboratory compared to no training is significant.

Irreproducibility according to research styles

- The objective definition of "continuity" and "exploratory" PIs should be stated here for this to be interpretable (note that this is not clear in the Methods either).

We have better defined how we separate "continuity" and "exploratory" PIs in the methods and in the result section.

Multivariable analysis of predictors of claim irreproducibility

- As mentioned in the Public Review, does it make sense to include "unchallenged" along with "verified" in the outcome? If the objective is to reduce the analysis to a "verified/not verified" claim, wouldn't it make more sense to remove the unchallenged findings (as these are likely uninformative, and based on the authors' own replications may be more related to the "challenged" category than to the "unchallenged" one?

See above, “Response to unchallenged exclusion”

- As stated previously, why not include journal impact factor, university ranking, year of first paper and year of first Drosophila immunity paper, as continuous variables rather than categorical ones (which would effectively lead to a model with less parameters when there are more than 2 categories). Particularly, there is more information to be gained from adding a single variable than from performing individual comparisons between categories when these have a natural order.

See above, “Response to categorical variable encoding”

Discussion:

- As "conceptual reproducibility" is somewhat of a vague concept, it seems important to discuss (both in the Methods and Discussion) how this was operationalized. Although it's obvious some decisions of what constitutes a direct replication will vary on a case-by-case basis, general guidelines on how these criteria were set are needed.

We have defined in the methods used to categorize article claims as well as the link with the other companion article. In the ReproSci database, all the assessment are justified allowing to see how we classified claims.

- "Contrary to the more dramatic narratives...". Here, it is important to emphasize the key differences between this and the cited references: (a) the fact that the majority of the replicability of the sample was found on the basis of a retrospective sample and (b) the fact that the study is dealing with conceptual rather than direct replications. Also, the possibility of confirmation bias should be mentioned as an alternative hypothesis in the last sentence of this paragraph.

This is a good point and we have highlighted that methodological differences between our study and other studies could explain differences in replicability rates.

- "findings that, despite being published, have never been independently tested". A more accurate description may be "have never been independently confirmed or refuted in the published literature", as many (and perhaps most - see Baker 2016) replication attempts may go unpublished, as the following sentences themselves indicate.

We have changed the text according to the reviewer’s suggestion.

- The description of how findings came to be regarded as "suspicious" here is interesting - and an important part of how the sample was determined. In this sense, I'd consider this as part of the methods. Even though defining what makes something "suspicious" may not be completely systematic, a general description of the method that led the researchers to arrive at this list (e.g. the process that seems to be hinted at in this paragraph) deserves a thorough description.

The description of how findings came to be regarded as "suspicious" is now detailed in the methods.

- "Our results suggest the status of a claim being "unchallenged" is not a reliable proxy for its validity". I agree, but this is statement is in direct contradiction with the authors' decision to include unchallenged claims along with validated ones in the binary outcome of the multivariate model.

See above, “Response to unchallenged exclusion”

- "Trophy journals are more likely to publish articles with challenged claims than high or lowimpact, although the difference was not significant." Again, a single analysis using a continuous measure of impact should provide more statistical evidence than the pairwise comparisons.

See above, “Response to categorical variable encoding”

- "However, this higher rate of follow-up work cannot fully explain by itself the higher proportion of challenged claims in trophy journals." Why not, exactly? I don't remember seeing an objective analysis of it.

If we remove the unchallenged claims, we still observe higher rate of irreplicability in trophy journal: 15.85% in trophy journal versus 9% in high-impact and 8.1% in low-impact. So we believe that the higher rate of irreplicability in trophy journals cannot be explained by a lower level of unchallenged.

- The very large paragraph on pages 21-22 could be split into two, one about journals and the other about universities.

This has been done.

- "Our data align with the broader narrative of increasing rates of non-reproducible science.". Does it? The time trend did not seem very clear to me (and I would argue that it was not analyzed properly).

As indicated in the text, we observed an increase in the rate of irreplicable claims over time; however, this trend did not reach statistical significance. We are confident in the robustness of our analysis and therefore did not pursue this question further. While our findings align with concerns about non-replicable science, they also offer a more nuanced perspective, showing that in some fields, replicability remains significantly high.

- "A similar but more acute pattern was observed during the SARS-CoV-2 pandemic." References should be provided here.

We have added this reference to suggest that articles published on SARS-CoV-2 pandemic are overall less reliable than others: An alarming retraction rate for scientific publications on Coronavirus Disease 2019 (COVID-19) Nicole Shu Ling Yeo-The https://doi.org/10.1080/08989621.2020.1782203

- Some of the trends discussed here (e.g. time, continuity vs. exploratory style are not supported (even as a trend) by the multivariate model, and this caveat should be mentioned. An odds ratio of 0.89 with a very wide confidence interval may be too weak in terms of evidence strength to merit a whole paragraph in the introduction discussing it.

We did not mention the question of time horizon or the distinction between continuity and exploratory research styles in the Introduction; however, we believe these issues merit consideration in the Discussion. In particular, the contrast between continuity-driven and exploratory approaches is noteworthy. Although we found no significant difference in the rate of challenged claims between these approaches, there was a significant difference in the rate of unchallenged claims, which are subsequently more likely to be challenged. This finding is important in the current funding landscape, where many agencies prioritize short-term, exploratory projects. Such incentives may inadvertently contribute to issues of irreproducibility.

- There are many other limitations beyond those mentioned, in particular the fact that much of the analysis is retrospective and based on a potentially biased literature. The fact that the analysis does not seem to be preregistered and is dependent on a lot of ad hoc decisions also merits discussion. This should definitely be explored in more detail in the Limitations section.

These limitations are now discussed in the limitation section at the end of the discussion. We agree on the fact that our analysis was not preregistered and that some of our conclusions are raised after analyzing the data. This article is part of a more global project including the databases with all the information. However, full objectivity when analyzing literature is not unachievable, and these critics are inherent to all reproducibility projects.

Methods:

Selected articles, annotation and experimental validation:

- "In brief, a list of 400 publications published before 2011 was generated using a curated search string on the publicly available PubMed database.". Please state that the search string is available in the companion article (e.g. https://doi.org/10.1101/2025.07.07.663442) or include it here.

This is now stated.

- "Selected primary articles were annotated by a single researcher".

By "annotation", do the authors mean the extraction of claims? This is not self-evident. Also, what does the "review" process entails? Do the authors have any data on agreement?

We do not have data on agreement but overall, there were few discrepancies. We agree that the way to section article in separate claims could affect the conclusion in a number of cases. All the data are public and we did not have any request from the community. The fact that this project depends from arbitrage from the two annotators is mentioned in the limitations.

- "Claims were cross-checked with evidence from previous, contemporary and subsequent publications and assigned a verification category."

How was this process performed? This seems to be quite complex and there's hardly any information about it. And once again, do the authors have any data on how reproducible this process would be when performed by different people?

All the data, notably claim assessment and justifications, are available on the website. Findings cross-checking the claim could be identified by knowledge of the authors, checking articles that quotes the articles. This was an enormous amount of work with human arbitrage. We expected more feedback from the community. The reaction of the community will be described in a companion article later.

- The authors mentioned that annotations and verifications were made available for comments on the community, but how were these comments incorporated if different opinions were voiced? Can the authors provide data on how frequent these comments were, and how often they were incorporated?

All data is available on the website, where members of the community can publicly comment on our assessments. The authors can see these comments on the ReproSci website

Table 1:

- Again, the three "partially verified" categories in Table 1 seem to refer to very different situations and it would be interesting to break these down somewhere.

We agree that categories ‘mixed’ and ‘partially verified’ are complex. They represent situations where a decision was not easy. We prefer not to break down those categories in multiple sub-categories.

First and last author classification:

- What does "status of first author" mean?

Status means their position: technician, PhD student, Post-doc, PI….

- "PIs were manually classified as..." - there seems to be something missing in this sentence (e.g. "as senior or junior on the basis of whether they have...")

We have added PIs: PIs were manually classified as senior or junior PIs on the basis on having published an article in Pubmed as last author more than five year ago.

- The distinction between continuity and exploratory is not explained in objective terms. Is there any objective definition of what it means to "continue to work in the field". Publishing a paper in the last X years? And what counts as a "transient" incursion?

We have better explained in the methods this distinction.

Statistical analysis:

- Can the authors define exactly what they mean by "weakly informative priors"?

Weakly informative prior are Bayesian prior distribution that are meant to be vague as to let the data dominate the results (vs the prior belief of the scientist). They can be seen as very close to flat priors, which are not used here because they cause numerical convergence issues.

- The authors mention a lot of variables included in the multivariable analysis, but little information is provided on the categorization. What are "low, high and trophy journals"? What university rankings are used?

We used Shanghai Ranking’s 2010 Academic Ranking of World Universities as our ranking of university. Our binning of ranking in categories make the analysis less susceptible to changes in ranking system. We selected the Shanghai ranking as it primarily evaluates research output and awards (where Times also evaluate teaching, and QS also employer reputation), which we believe represent better the variable than may affect irreplaceability. For the categorization, please see answer above: “Response to categorical variable encoding”

- "Exact formulas (...) are available in the public repository". What repository do authors mean? The project website? The GitHub repository? Please specify and provide a direct link if possible.

We added a link to the main text (it was only in the supplementary document). The Supplementary document contains information on how to reproduce the results.

Reviewer #2 (Recommendations for the authors):

Specific comments:

Some methodological details related to the main conclusions of the paper are missing.

We have extended the methodological section and also better link this article to the companion article.

- Impact factor: It is unclear which year(s) of journal impact factors were used. Trophy journals are defined as those with an impact factor >50 (Science, Nature, and Cell), but according to Figure 2B, there are only three such journals. Are "trophy journals" limited to these three, or are others included that meet the >50 impact factor threshold?

These are the 2022 impact factor, we made that clear in the main text. Moreover, only 3 journals in our datasets had an impact factor >50: Nature, Cell, and Science (so the Trophy Journal category is limited to these 3, but not by definition). This is explicit in Table S1: List of journals with impact factor and claim assessment.

- University ranking: Similarly, please clarify which year(s) of university ranking data were used. Since rankings vary depending on the system (e.g., QS, Times, Shanghai), were the conclusions consistent across multiple ranking sources?

We used Shanghai Ranking’s 2010 Academic Ranking of World Universities as our ranking of university. Our binning of ranking in categories make the analysis less susceptible to changes in ranking system. We selected the Shanghai ranking as it primarily evaluate research output and awards (where Times also evaluate teaching, and QS also employer reputation), which we believe represent better the variable than may affect irreplaceability.

Reviewer #3 (Recommendations for the authors):

Table S4 - for leading author there is a variable of 'historical lab after 1998 continuity' - should that be 1995?

Correct, we changed this label to Trained in Historical laboratory (Comparison restricted to claims published after 1995) to make it more explicit.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation