The structural context of mutations in proteins predicts their effect on antibiotic resistance

  1. Anna G Green  Is a corresponding author
  2. Mahbuba Tasmin
  3. Roger Vargas Jr
  4. Maha Reda Farhat  Is a corresponding author
  1. Department of Biomedical Informatics, Harvard Medical School, United States
  2. Manning College of Information and Computer Sciences, University of Massachusetts, United States
  3. Division of Pulmonary & Critical Care, Massachusetts General Hospital, United States
8 figures, 3 tables and 8 additional files

Figures

Proteins with the highest frequency of mutation events are associated with antibiotic resistance.

(A) Workflow used to create our combined dataset of missense substitutions and in-frame indel mutations mapped to protein 3D structures, for 92% of the M. tuberculosis H37Rv proteome. Using a dataset of homoplastic mutations from 31,428 Mycobacterium tuberculosis complex (MTBC) isolates (Green et al., 2023; Vargas et al., 2023), we mapped mutations to protein sequences and 3D structures based on a combination of experimentally determined (RCSB PDB; Berman et al., 2000) and computationally predicted (AlphaFold; Varadi et al., 2022) structures. (B) The total number of mutation events in our dataset per protein, versus the percent of the amino acids in the protein’s structure that have been mutated at least once. Marginal histograms are displayed along both axes. Proteins with the highest frequency of mutations are those associated with resistance to antibiotics, according to the WHO catalog of resistance-associated mutations (Walker et al., 2022).

G-statistic reveals clustering of mutations in antibiotic-resistance-conferring proteins.

(A) Workflow to compute residue-wise Getis–Ord statistic for proteins in M. tuberculosis. (B) Results for proteins: KatG is shown in complex with heme (orange), PDB ID = 4C51 chain A (Zhao et al., 2013). RpoB is shown in complex with rifampin (orange), PDB ID = 5UH6 chain C (Lin et al., 2017) aligned as described in Methods. PncA is shown in complex with Fe2+, PDB ID = 3PL1 chain A (Petrella et al., 2011). RsmG (encoded by the gidB gene) structure from Thermus thermophilus is shown in complex with ligand adenosine monophosphate (please note the streptomycin-binding site is unknown in M. tuberculosis and thus is not shown here), PDB ID = 3G8A chain A (Gregory et al., 2009).

Benchmarking the ability of G-scores to find significant clustering in protein structures.

The procedure for generating downsampled true positive and true negative examples from real proteins. We then test five scores for their ability to distinguish true positives and negatives.

Hits of proteome-wide screen for clustering of mutations.

(A) Pipeline for detecting hits in 3687 proteins. (B) GO terms (abbreviated for space) with top fold enrichment (all significant at FDR <0.05). (C) Examples of proteins with significant clustering. All structures shown are from AlphaFold with low-confidence residues filtered out (note that relative domain orientation for PknB and PknH is low confidence).

Appendix 1—figure 1
Homoplastic inframe insertion at position 3131469 in the Cas10 gene (Rv2823c).

Screenshot from Mycobrowser (https://mycobrowser.epfl.ch/) of relevant genomic region beginning at 3131469. Table showing the number of mutation events per lineage and total in the dataset, as well as total number of isolates with the alternate allele. Note that the inserted sequence is similar to but not exactly the same as the H37Rv reference sequence at that location, and that similar motifs recur throughout the sequence region.

Appendix 1—figure 2
Mutations and GeO clustering for RS12 (RS12_MYCTU).

Two residues, shown in orange (K43 and K88) are highly mutated in the protein RpsL, leading to significant G-score clustering in that region of the protein.

Appendix 1—figure 3
Distance between top pairs of high G-score residues in proteins with significant clustering.

Of the 499 proteins with significant hits, we analyze what number are still significant after filtering to ensure that the minimum inter-atom distance between the top 2 residues with high G-score is less than a defined distance threshold. For 90.6% (452 of 499) of the significant hits, the top G-score pair of residues are within 15 Ångstroms in 3-D space. Decreasing the distance threshold to 8 and 5 Ångstroms results in 74.7% (373) and 61.1% (305) hits whose top pairs are close in 3-D, respectively.

Appendix 1—figure 4
Relationship between protein length and distance between top pair of residues.

Of the 499 proteins with significant hits, we analyze whether there is a statistically significant relationship between the length of the protein (filtered for residues that pass our structure quality thresholds) and the distance between the two residues with highest G-score. Using Ordinary Least Squares regression implemented with default parameters in statsmodels v0.14.4, we find a weak negative relationship: R2 = 0.008, beta = -3.1806, p-value = 0.048, indicating that distance between top pairs does not increase with protein length, hence our method is likely capturing real signal for 3-D clustering.

Tables

Table 1
Performance of classification models on predicting whether mutations are R-conferring from the mutation catalog.

The 1D-proximity model was trained using just the distance in primary sequence to the nearest known R mutation, the 3D-proximity model was trained using the distance in 3D to the nearest known R mutation, and G-score was trained using the G-score calculated in this manuscript. Reported values are calculated on the held-out test set.

Feature setF1Precision/PPVSensitivity
3D-proximity96.598.294.9
1D-proximity89.198.381.3
G-score80.899.268.2
Key resources table
Reagent type (species) or resourceDesignationSource or referenceIdentifiersAdditional information
OtherM. tuberculosis H37RvUniProtUP000001584Reference proteome
OtherM. tuberculosis H37RvCole et al., 1998.H37RvReference genome
Appendix 1—table 1
Top 10 GO categories significantly enriched in the clustered protein set.

Categories with identical members and FDR (e.g., GO:0071103 DNA conformation change and GO:0006265 DNA topological change) have only one representative category shown. See Supplementary file 6 for complete table.

GO IDGO term labelUniProt identifiersFold enrich.FDRCategory
GO:0035635Entry of bacterium into host cellQ6MX51_MYCTU, GLMU_MYCTU, SAHH_MYCTU11.040.01Host cell entry
GO:0006265DNA topological changeGYRA_MYCTU, GYRB_MYCTU, TOP1_MYCTU11.040.01DNA topology
GO:0016539Intein-mediated protein splicingDNAB_MYCTU, RECA_MYCTU, Y1461_MYCTU11.040.01Protein maturation
GO:0046349Amino sugar biosynthetic processGLMM_MYCTU, MURA_MYCTU, GLMU_MYCTU11.040.01Amino sugar
GO:0006047UDP-N-acetylglucosamine metabolic processGLMM_MYCTU, GLMS_MYCTU, GLMU_MYCTU11.040.01Amino sugar
GO:0062014Negative regulation of small molecule metabolic processPKNB_MYCTU, GARA_MYCTU, PKNA_MYCTU, PKNE_MYCTU, PKND_MYCTU9.2<0.01Fatty acid
GO:0042304Regulation of fatty acid biosynthetic processPKNB_MYCTU, PKNA_MYCTU, PKNE_MYCTU, PKND_MYCTU8.830.01Fatty acid
GO:0006040Amino sugar metabolic processGLMM_MYCTU, GLMS_MYCTU, MURA_MYCTU, GLMU_MYCTU8.830.01Amino sugar
GO:0015990Electron transport coupled proton transportNUOM_MYCTU, COX1_MYCTU, NUOL_MYCTU8.280.04ETC
GO:0046890Regulation of lipid biosynthetic processPKNB_MYCTU, PKNA_MYCTU, PKNE_MYCTU, P71814_MYCTU, PKND_MYCTU7.88<0.01Fatty acid

Additional files

Supplementary file 1

Table with IDs of all M. tuberculosis complex isolates used in this study, including their internal identifier, their external database identifier, and their dataset of origin.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp1-v1.csv
Supplementary file 2

Table of all genomic mutations analyzed and relevant statistics.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp2-v1.csv
Supplementary file 3

Table of all protein structures used and their database of origin.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp3-v1.csv
Supplementary file 4

Table containing results of mutation downsampling and shuffling experiment used to calibrate protein-level score.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp4-v1.csv
Supplementary file 5

Table containing protein-level mutation clustering score for all proteins in M. tuberculosis.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp5-v1.csv
Supplementary file 6

Table with results of GO-enrichment experiment.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp6-v1.csv
Supplementary file 7

Table with data input to mutation classification study.

https://cdn.elifesciences.org/articles/109450/elife-109450-supp7-v1.csv
MDAR checklist
https://cdn.elifesciences.org/articles/109450/elife-109450-mdarchecklist1-v1.docx

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Anna G Green
  2. Mahbuba Tasmin
  3. Roger Vargas Jr
  4. Maha Reda Farhat
(2026)
The structural context of mutations in proteins predicts their effect on antibiotic resistance
eLife 14:RP109450.
https://doi.org/10.7554/eLife.109450.3