Author response:
eLife Assessment
This study presents a useful database resource containing protein conformations generated through molecular dynamics simulations, with extensive quality evaluation and benchmarking. While the database is well-constructed and professionally organized, the evidence supporting its claimed representation of protein conformational landscapes is incomplete, as the short simulation times and starting structure bias prevent true Boltzmann sampling of the conformational space.
We thank the editors for recognizing the usefulness of ProteinConformers and the value of its quality evaluation and benchmarking. We will revise the manuscript to clarify that ProteinConformers provides large-scale, energetically profiled descriptions of protein conformational landscapes, with broad coverage of locally stereochemically valid and energetic compatible structures from non-native to near-native regions, rather than a complete equilibrium sampling of all conformational states. These revisions will better define the scope of the resource while preserving its intended use for benchmarking and data-driven studies of protein conformational variability.
Public Reviews:
Reviewer #1 (Public review):
Summary:
The authors describe a new database that rigorously explores protein conformations.
Strengths:
It is extremely well done, using state-of-the-art tools by a group at the top of the field of structural modeling. The evaluation of qualities and the benchmarking of the structures are outstanding, and it is expected that the new database will have a significant impact on the field.
We thank Reviewer #1 for the positive evaluation of our work and for recognizing the potential impact of the ProteinConformers resource.
Weaknesses:
The authors are using MD simulation to generate some of the structure, and therefore should have access to standard MD energies. I am surprised that no evaluation is provided based on these energies that can be extended to free energies.
We thank the reviewer for this helpful suggestion. We reprocessed the original MD energy files and extracted three MD-derived energy terms, including total energy, system potential energy, and protein-only potential energy. These energy terms have been added to the ProteinConformers resource and the web portal. We will update the manuscript to describe these additional energetic annotations.
Reviewer #2 (Public review):
Summary:
The authors developed a dataset of protein conformations by running molecular dynamics simulations starting from both native and decoy conformations for a large number of proteins. These conformations were put together as a dataset for querying and downloading, along with their energies under different force fields. The authors suggest that such conformations represent the proteins' conformational landscape, so that they will be useful for evaluating methods generating multiple conformations of proteins.
Strengths:
The dataset is online and working. It has good documentation for others to use.
We appreciate Reviewer #2’s positive assessment of the online resource and documentation.
Weaknesses:
The biggest weakness is that the collected conformations very likely do not represent the true conformational landscape. To represent the conformational landscape, the structures need to be sampled based on the Boltzmann distribution. However, in this study, conformations are generated by running very short (125ps to 375ps) MD simulations starting from near-native conformations and decoys. Such short simulations will produce small fluctuations around the starting conformations, so the distribution of conformations is largely dominated by the distribution of the initial conformations, which by one means are Boltzmann distributed. A conformation might be physically plausible, but it might have very small weight in the Boltzmann distribution. On the other hand, conformations with large weights might not be in the dataset.
We thank the reviewer for this important and constructive comment. We agree that the conformations in ProteinConformers should not be interpreted as an equilibrium ensemble sampled according to the Boltzmann distribution. Because the MD simulations used here are short, the resulting snapshots around each seed mainly reflect local relaxation and limited thermal fluctuation from that seed, rather than exhaustive equilibrium sampling. Therefore, the relative population of conformations in our dataset should not be interpreted as a Boltzmann weight, and some thermodynamically important states may be underrepresented or absent.
Our goal in this work is different from conventional long-timescale MD studies that aim to estimate equilibrium populations from one or a few initial structures. ProteinConformers was designed as a large-scale, multi-seed, MD-refined conformer resource. The broad structural coverage comes primarily from initiating simulations from many diverse seed decoys for each protein, while the short all-atom MD protocol is used to relax structures under a molecular mechanics force field, remove structures that fail to converge, reduce steric clashes and unrealistic local geometries, and generate energetically annotated conformers. We will revise the manuscript to make this distinction clearer and to avoid implying that ProteinConformers provides a rigorous Boltzmann-sampled representation of the underlying thermodynamic landscape.
We also agree that longer simulations and enhanced sampling methods, such as replica-exchange MD, metadynamics, or umbrella sampling, would be necessary to estimate equilibrium populations and improve sampling of rare but thermodynamically relevant states. We will add this point as a limitation and future direction in the revised manuscript. Thus, ProteinConformers should be viewed as a broad, energetically annotated, MDrefined conformer library for benchmarking, data-driven modeling, and descriptions of protein conformational landscapes, rather than as a complete equilibrium ensemble itself.
Reviewer #3 (Public review):
Summary:
This manuscript describes a web-based tool that allows researchers to compare large numbers of representative ("plausible") conformations of proteins. It also includes energetic analysis from multiple widely used structure-prediction methods.
Strengths:
This tool will likely be useful for students who want to learn more about the ensemble properties of proteins. The resource is well organized and it represents a large amount of computing resources.
We thank Reviewer #3 for the positive assessment of the ProteinConformers resource and for recognizing its potential value for community education.
Weaknesses:
It is not entirely clear how the database may be utilized by other groups to advance research. It could be helpful if the authors add a short section that provides example use cases that illustrate how this database can support new strategies for studying protein dynamics.
We thank the reviewer for this constructive suggestion. We agree that the manuscript should more explicitly explain how other groups can use ProteinConformers to advance research. In the revised Discussion, we will add a short section describing concrete use cases of the database. In particular, we will emphasize that the benchmark analysis already presented in this manuscript provides a worked example of how ProteinConformers can be used by other groups. ProteinConformers-lite, together with the released evaluation metrics and codes, can serve as a standardized reference set for testing new multi-conformation or protein ensemble generation methods. Other groups can generate conformational ensembles for the same targets, compare their coverage of low-energy regions using the diversity metrics reported in Table S1, and evaluate the agreement of residue-pair geometric statistics using the plausibility metrics reported in Table S2.
We will also describe additional use cases enabled by the full ProteinConformers resource. The dataset can be used as a training, validation, or pretraining resource for conformation generators, energy-aware ranking models, and model quality assessment methods, because each conformer is paired with structural similarity annotations and multiple energetic scores. The broad coverage from non-native to near-native conformations also enables systematic analysis of how local stereochemical validity, global structural similarity, and energetic evaluations co-vary across diverse conformational perturbations, and may provide useful structural proxies or starting points for modeling flexible or disordered-like protein states, where experimentally resolved structural data are often limited. In addition, the interactive portal allows users to filter protein-specific conformers by structural similarity, energetic annotations, and secondary-structure features, making it possible to construct customized subsets for downstream biomolecular modeling, hypothesis generation, and educational exploration.