Mechanistic transmission models inform frequency dynamics.

(A) Genomic surveillance reveals changes in the frequency of genetic variations over time. (B) These frequency changes can arise due to differences in phenotypes related to transmission (e.g, immune escape, transmissibility, binding) and changes in population immunity due to recent exposure. (C) Despite being instrumental for real-time analysis, both variant phenotype and population immunity are rarely observed in real time. (D) We use mechanistic transmission models to infer relative fitness from frequency data alone, taking advantage of known structure in transmission dynamics. This enables us to quantify trade-offs between variant phenotypes and develop new methods for estimating fitness in populations undergoing antigenic evolution.

Trade-off between degree of immune escape and increased transmissibility.

A. Relative fitness for a transmissibility increasing variant T with ρ = 0.2 and an immune escaping variant E with η = 0.3 for R0,W = 2.8 and 1/γ = 3.0 days. The intersection point shows that after 40% of the population has wildtype immunity, the escape variant has higher fitness. B. The critical exposure proportion is shown for various escape fraction and transmissibility increase. Above the critical exposure proportion, we expect dominance of escape variants. C. The minimum escape fraction needed for second waves to be comprised of escape variant assuming competition with transmissibility increase variants and first wave with a given R0.

Relative fitness is correlated with vaccination levels in the absence of immune escape.

We simulate the growth of a pure transmissibility increased variant at varying levels of vaccination. Darker colors represent lower vaccine uptake. We identify an early growth period where relative fitness is at its highest; the cutoff for this period is denoted with a vertical dashed line. A. Prevalence of variant, each line is its own simulation. B. Frequency of variant. C. Relative fitness for variant over time. D. Estimated log growth advantage using linear regression of log relative frequency of variant over wildtype using only data before the early cutoff. E. Same as D. but using data from the entire period shown.

Predicting epidemic growth rate using estimated selective pressure.

A. Variant frequency estimated using the Gaussian process relative fitness model between January 2021 and November 2022 for sequence count data from Washington state. B. Case counts from Washington state. C. Selective pressure computed using estimated variant frequencies and relative fitnesses from Washington state. D-F. Predictions for empirical growth rate from selective pressure for selected US states. The light gray period is the training period and the darker gray is the testing period. G-I. Predictions for empirical growth rate from selective pressure for countries South Africa, South Korea and the UK. J. Prevalence estimates for England from ONS Infection Survey. K. Estimated selective pressure in England. L. Empirical growth rates (gray) computed from prevalence estimates and predictions from our model (green) computed from selective pressure.

Latent factor models of immunity describe variant dynamics.

We fit our latent immunity factor model for D = 8 pseudo-immune groups using only SARS-CoV-2 sequence count data. A. Variant frequency. Lines are colored to show 4 variants of interest (of 53 total variants) with the style of the line denoting 3 countries of interest (of 18 total countries). B. Estimated relative fitness for selected variants and countries. These variant-specific relative fitnesses are similar across countries, but not identical. C. Estimated pseudo-immunity cohorts (PIC) over time for multiple countries ordered by decreasing share in the first geography using sequence data alone. D, E. Dimensionality-reduced pseudo-escape rates using multidimensional scaling (MDS). F. Estimated pseudo-escape rates for each variant relative to pivot variant “other”.

Latent factor models of immunity predict titers across varying exposure histories.

Using pseudo-escape values from our latent factor model (Fig. 5) and human titer data, we show that pseudo-escape values predict antigenic distance and titers. (A-B) Principal component analysis of subject-level titers, colored by learned immune clusters (A) and subject infection history (B). (C) Co-occurrence matrix between learned immune clusters and infection histories. (D) Comparing pairwise distance between variants in the pseudo-immune space to observed distances in human titer data. (E-F) Correlation between log mean titer against a variant and that variant’s pseudo-escape by group. (G) Estimated escape burden by immune cluster, showing variant escape potential against individual immune clusters.

Simulated variant dynamics in a mechanistic model.

Mechanistic transmission models constrain variant frequency dynamics by specifying a functional form for relative fitnesses. Simulations of a three-variant model including wildtype W , an intrinsic transmission variant T , and an immune escape variant E show the relationship between population-level transmission and selection. We begin the simulation with initial wildtype prevalence IW (0) = 1, effective reproduction number R0,W = 1.4, and duration of infection 1/γ = 3.0 days. We introduce transmissibility variant T at t = 20 with frequency fT (20) = 10−5 and a 50% increase in transmissibility ρ = ρT = 0.5. We introduce escape variant E at t = 70 with frequency fE(70) = 10−6 that infects 5% of hosts possessing wildtype immunity η = ηE = 0.05. A. Prevalence I by variant. B. Exponential growth rate r by variant. C. Variant frequency f . D. Fitness relative to wildtype λ. E. Underlying immune pools. F. Effective reproduction number Rt by variant.

Estimating relative fitness with Gaussian processes.

Gaussian processes allow us a non-parametric estimate of the relative fitness for variants through time. This figure uses Gaussian processes to model the 3 variant example shown in Fig. S1. A. Synthetic sequence counts generated using a multinomial distribution with frequencies from Fig. S1C. B. Frequencies and posterior frequencies according to Gaussian process model. Intervals show the 80% credible interval. C. Posterior relative fitnesses. Intervals show the 80% credible interval. Dashed line shows true relative fitnesses from underlying mechanistic model.

Differences in fitness mechanisms impact frequency and prevalence in the short-term.

Comparing simulations from two independent two-variant systems with either an escape variant E (orange) or a transmissibility variant T (purple). We fix the initial relative fitness for the two variants using Equation 4 and simulate dynamics for 365 days. A. The prevalence for the variants. B. The relative fitness from the variants. C. The cumulative wildtype incidence as a function of the initial wildtype frequency. D. The difference between the cumulative incidence between the escape variant and the transmissibility variant as a function of wildtype incidence.

Estimated variant frequencies, relative fitnesses, and selective pressure.

Alabama through Georgia.

Estimated variant frequencies, relative fitnesses, and selective pressure.

Hawaii to Maryland.

Estimated variant frequencies, relative fitnesses, and selective pressure.

Massachusetts to New Jersey.

Estimated variant frequencies, relative fitnesses, and selective pressure.

New Mexico to South Carolina.

Estimated variant frequencies, relative fitnesses, and selective pressure.

South Dakota to Wyoming.

Predictions for empirical growth rate using selective pressure for all locations.

Cross-validation error by model

We compare the errors between models fit on 10 time series cross-validation splits.

Comparing latent factor model by number of latent immune dimensions.

A. Maximum a posteriori loss by number of latent immune dimensions. B. Spearman correlation between pseudo-immune and titer distance by number of latent immune dimensions. C. Number of parameters by number of latent immune dimensions. D. Bayesian Information Criterion (BIC) by number of latent immune dimensions.

Latent factor model with D = 2 pseudo immune dimensions.

Latent factor model with D = 4 pseudo immune dimensions.

Latent factor model with D = 6 pseudo immune dimensions.

Latent factor model with D = 10 pseudo immune dimensions.

Choosing the number of immune clusters (k = 5) using inertia and silhouette score.

This analysis uses only titer data from Jian et al. [7].

Bootstrapping pseudo-immune distance and human titer distance analysis (Nreplicate = 1, 000).

Comparing pseudo-escape distance and titer distance between exposure groups.

Titers predicted with pseudo-escape values and observed titers within infection cohorts.

Titers predicted with pseudo-escape values and observed titers within immune clusters.