The structural context of mutations in proteins predicts their effect on antibiotic resistance
Figures
Proteins with the highest frequency of mutation events are associated with antibiotic resistance.
(A) Workflow used to create our combined dataset of missense substitutions and in-frame indel mutations mapped to protein 3D structures, for 92% of the M. tuberculosis H37Rv proteome. Using a dataset of homoplastic mutations from 31,428 Mycobacterium tuberculosis complex (MTBC) isolates (Green et al., 2023; Vargas et al., 2023), we mapped mutations to protein sequences and 3D structures based on a combination of experimentally determined (RCSB PDB; Berman et al., 2000) and computationally predicted (AlphaFold; Varadi et al., 2022) structures. (B) The total number of mutation events in our dataset per protein, versus the percent of the amino acids in the protein’s structure that have been mutated at least once. Marginal histograms are displayed along both axes. Proteins with the highest frequency of mutations are those associated with resistance to antibiotics, according to the WHO catalog of resistance-associated mutations (Walker et al., 2022).
G-statistic reveals clustering of mutations in antibiotic-resistance-conferring proteins.
(A) Workflow to compute residue-wise Getis–Ord statistic for proteins in M. tuberculosis. (B) Results for proteins: KatG is shown in complex with heme (orange), PDB ID = 4C51 chain A (Zhao et al., 2013). RpoB is shown in complex with rifampin (orange), PDB ID = 5UH6 chain C (Lin et al., 2017) aligned as described in Methods. PncA is shown in complex with Fe2+, PDB ID = 3PL1 chain A (Petrella et al., 2011). RsmG (encoded by the gidB gene) structure from Thermus thermophilus is shown in complex with ligand adenosine monophosphate (please note the streptomycin-binding site is unknown in M. tuberculosis and thus is not shown here), PDB ID = 3G8A chain A (Gregory et al., 2009).
Benchmarking the ability of G-scores to find significant clustering in protein structures.
The procedure for generating downsampled true positive and true negative examples from real proteins. We then test five scores for their ability to distinguish true positives and negatives.
Hits of proteome-wide screen for clustering of mutations.
(A) Pipeline for detecting hits in 3687 proteins. (B) GO terms (abbreviated for space) with top fold enrichment (all significant at FDR <0.05). (C) Examples of proteins with significant clustering. All structures shown are from AlphaFold with low-confidence residues filtered out (note that relative domain orientation for PknB and PknH is low confidence).
Homoplastic inframe insertion at position 3131469 in the Cas10 gene (Rv2823c).
Screenshot from Mycobrowser (https://mycobrowser.epfl.ch/) of relevant genomic region beginning at 3131469. Table showing the number of mutation events per lineage and total in the dataset, as well as total number of isolates with the alternate allele. Note that the inserted sequence is similar to but not exactly the same as the H37Rv reference sequence at that location, and that similar motifs recur throughout the sequence region.
Mutations and GeO clustering for RS12 (RS12_MYCTU).
Two residues, shown in orange (K43 and K88) are highly mutated in the protein RpsL, leading to significant G-score clustering in that region of the protein.
Distance between top pairs of high G-score residues in proteins with significant clustering.
Of the 499 proteins with significant hits, we analyze what number are still significant after filtering to ensure that the minimum inter-atom distance between the top 2 residues with high G-score is less than a defined distance threshold. For 90.6% (452 of 499) of the significant hits, the top G-score pair of residues are within 15 Ångstroms in 3-D space. Decreasing the distance threshold to 8 and 5 Ångstroms results in 74.7% (373) and 61.1% (305) hits whose top pairs are close in 3-D, respectively.
Relationship between protein length and distance between top pair of residues.
Of the 499 proteins with significant hits, we analyze whether there is a statistically significant relationship between the length of the protein (filtered for residues that pass our structure quality thresholds) and the distance between the two residues with highest G-score. Using Ordinary Least Squares regression implemented with default parameters in statsmodels v0.14.4, we find a weak negative relationship: R2 = 0.008, beta = -3.1806, p-value = 0.048, indicating that distance between top pairs does not increase with protein length, hence our method is likely capturing real signal for 3-D clustering.
Tables
Performance of classification models on predicting whether mutations are R-conferring from the mutation catalog.
The 1D-proximity model was trained using just the distance in primary sequence to the nearest known R mutation, the 3D-proximity model was trained using the distance in 3D to the nearest known R mutation, and G-score was trained using the G-score calculated in this manuscript. Reported values are calculated on the held-out test set.
| Feature set | F1 | Precision/PPV | Sensitivity |
|---|---|---|---|
| 3D-proximity | 96.5 | 98.2 | 94.9 |
| 1D-proximity | 89.1 | 98.3 | 81.3 |
| G-score | 80.8 | 99.2 | 68.2 |
| Reagent type (species) or resource | Designation | Source or reference | Identifiers | Additional information |
|---|---|---|---|---|
| Other | M. tuberculosis H37Rv | UniProt | UP000001584 | Reference proteome |
| Other | M. tuberculosis H37Rv | Cole et al., 1998. | H37Rv | Reference genome |
Top 10 GO categories significantly enriched in the clustered protein set.
Categories with identical members and FDR (e.g., GO:0071103 DNA conformation change and GO:0006265 DNA topological change) have only one representative category shown. See Supplementary file 6 for complete table.
| GO ID | GO term label | UniProt identifiers | Fold enrich. | FDR | Category |
|---|---|---|---|---|---|
| GO:0035635 | Entry of bacterium into host cell | Q6MX51_MYCTU, GLMU_MYCTU, SAHH_MYCTU | 11.04 | 0.01 | Host cell entry |
| GO:0006265 | DNA topological change | GYRA_MYCTU, GYRB_MYCTU, TOP1_MYCTU | 11.04 | 0.01 | DNA topology |
| GO:0016539 | Intein-mediated protein splicing | DNAB_MYCTU, RECA_MYCTU, Y1461_MYCTU | 11.04 | 0.01 | Protein maturation |
| GO:0046349 | Amino sugar biosynthetic process | GLMM_MYCTU, MURA_MYCTU, GLMU_MYCTU | 11.04 | 0.01 | Amino sugar |
| GO:0006047 | UDP-N-acetylglucosamine metabolic process | GLMM_MYCTU, GLMS_MYCTU, GLMU_MYCTU | 11.04 | 0.01 | Amino sugar |
| GO:0062014 | Negative regulation of small molecule metabolic process | PKNB_MYCTU, GARA_MYCTU, PKNA_MYCTU, PKNE_MYCTU, PKND_MYCTU | 9.2 | <0.01 | Fatty acid |
| GO:0042304 | Regulation of fatty acid biosynthetic process | PKNB_MYCTU, PKNA_MYCTU, PKNE_MYCTU, PKND_MYCTU | 8.83 | 0.01 | Fatty acid |
| GO:0006040 | Amino sugar metabolic process | GLMM_MYCTU, GLMS_MYCTU, MURA_MYCTU, GLMU_MYCTU | 8.83 | 0.01 | Amino sugar |
| GO:0015990 | Electron transport coupled proton transport | NUOM_MYCTU, COX1_MYCTU, NUOL_MYCTU | 8.28 | 0.04 | ETC |
| GO:0046890 | Regulation of lipid biosynthetic process | PKNB_MYCTU, PKNA_MYCTU, PKNE_MYCTU, P71814_MYCTU, PKND_MYCTU | 7.88 | <0.01 | Fatty acid |
Additional files
-
Supplementary file 1
Table with IDs of all M. tuberculosis complex isolates used in this study, including their internal identifier, their external database identifier, and their dataset of origin.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp1-v1.csv
-
Supplementary file 2
Table of all genomic mutations analyzed and relevant statistics.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp2-v1.csv
-
Supplementary file 3
Table of all protein structures used and their database of origin.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp3-v1.csv
-
Supplementary file 4
Table containing results of mutation downsampling and shuffling experiment used to calibrate protein-level score.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp4-v1.csv
-
Supplementary file 5
Table containing protein-level mutation clustering score for all proteins in M. tuberculosis.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp5-v1.csv
-
Supplementary file 6
Table with results of GO-enrichment experiment.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp6-v1.csv
-
Supplementary file 7
Table with data input to mutation classification study.
- https://cdn.elifesciences.org/articles/109450/elife-109450-supp7-v1.csv
-
MDAR checklist
- https://cdn.elifesciences.org/articles/109450/elife-109450-mdarchecklist1-v1.docx