ClustIRR analysis of five TCR repertoires from Dataset 1.

(A) Pairwise cosine similarity (CS) matrix between TCR repertoire’s CJ occupancy vectors (cell counts per CJ). Tile color intensity and labels indicate similarity values. (B) Mean β coefficients (y-axis) for 7,505 CJs (dots) across repertoires (x-axis). CJs categorized as EBV specific (+, blue) or non-EBV specific (–, orange). (C) Mean β coefficients (y-axis) for 7,505 CJs (dots) across repertoires (x-axis). CJs categorized as MART1 specific (+, purple) or non-MART1 specific (–, orange). Dot sizes indicate numbers of antigen specific T cells. Horizontal jitter width scales with point density. (D) Cumulative probability (y-axis) for all T cells (black) and T cells specific for EBV (blue) or MART1 (purple) across CJs ranked by β coefficients (x-axis) in each TCR repertoire (panels).

Integration of gene expression and TCR repertoires.

(A) Annotated bubbletree (tree structure) of dataset 1. The bubbletree has five bubbles (white circles) as leaves. Bubble radii scale linearly with the number of cells in the bubbles. Bubble indices, absolute and relative cell frequencies are shown as labels between bubbletree and heatmap. Gray branch labels indicate perfect support of branches in 1,000 bootstrap iterations. (B) Heatmap columns are relative frequencies of cells from different samples, integrating to 100%. (C) Heatmap rows are within-bubble relative frequencies of different cell lines, integrating to 100%. (D) Distribution of normalized expression sum of all marker genes associated with T cell activation (x-axis) in cells from each bubble (y-axis). (E) Heatmap rows are within-bubble relative frequencies of cell cycle phases, integrating to 100%. (F) Mean T cell activation score (y-axis) versus β coefficient (x-axis) for Communities on the Joint graph (CJs) (dots) across five repertoires (panels). Orange contours show 2D density of CJs. (G) Mean T cell activation score (y-axis) for T cells in CJs grouped by β coefficients (x-axis; sliding window: midpoints β = 1 to 6 in 0.5 increments, window size 2). Error bars: 95% HDIs from 1,000 bootstraps. Colors: TCR repertoires.

Analysis of longitudinal TCR data in lung cancer patients treated with CTLA-4 checkpoint inhibitors.

(A) Changes of Differentially occupant CJs (DCJs) between day 0 and day 22. Percentage of expanding (Ex, blue dots) and contracting (Co, orange dots) CJs in responders (CR/PR) and non-responders (SD/PD), including the overall mean percentages (large black dot) and 95% HDIs (black error bars) in each group. (B/C) Mean CJ intensity (β, y-axis) and 95% HDI intervals (vertical error bars) for expanding (B) and contracting (C) DCJs in patient Pt4 across five timepoints (x-axis). Dashed and solid lines show CJs with (+) and without (-) tumor TCRs, respectively. (D) Network visualization of 41 expanding and 14 contracting DCJs in Pt4. Nodes represent clonotypes (colored by timepoint; sized by relative clonal expansion on logarithmic scale). Squares: tumor-infiltrating TCRs (T+); circles: non-tumor-infiltrating TCRs (T-). Red labels: CDR3β sequences of T+ clonotypes. The largest clonotype in each CJ is labeled with its CDR3β sequence. Edges connect clonotypes with similar CDR3βs.

Analysis of TCR repertoires from humans and mice.

(A) Heatmap with clonotype-centric cosine similarity (CS) between pairs of TCR repertoires. (B) Heatmap with community-centric CS between pairs of TCR repertoires. Color-coded tiles represent CS values. (C) Log-log plot of normalized rank abundance distributions (NRADs) of TCR repertoires from humans (H, purple) and mouse (M, orange) repertoires. Each line represents an individual repertoire.

Differentially occupant CJs (DCJs) between humans and mice.

(A) Relative occupancy of CJs (dots) in humans and mice, represented by the mean of µ (y-axis). Horizontal dashed lines split the CJs into six groups used in panel B. Two most expanded in humans and mice are labeled (a–d). (B) Mean percentage of clonotypes recognizing specific antigens (color-coded) within CJ groups (x-axis) and the corresponding 95% HDIs (vertical error bars). The groups are ordered by increasing µ. Groups with no antigen-specific clonotypes were omitted. (C) Network of four DCJs, shared across all repertoires of one species (a–b in humans, c–d in mice) but not the other. Nodes represent TCR clonotypes from humans (reds) and mice (blues). Gray edges connect clonotypes sharing similar CDR3 sequences. Antigen-specific clonotypes are shown as nodes with black borders and labeled with the names of the antigen species. (D) Sequence logos of more conserved CDR3α motifs and more divergent CDR3βs within each CJ (rows: CJs a–d) (E) Relative usage (y-axis) of TRBV/TRBJ/TRAV/TRAJ gene segments (panels) in each CJ (x-axis). The most frequently used gene segments are labeled.

ClustIRR workflow.

(1) TCR sequencing (TCRseq) data are collected from multiple repertoires across longitudinal time points or biological conditions. 2) The main ClustIRR input is a table of TCR clonotypes for each repertoire, defined by CDR3α and CDR3β amino acid sequences and clonal sizes (cell counts). (3) ClustIRR constructs a joint similarity graph where nodes represent clonotypes and weighted edges connect clonotypes based on pairwise CDR3α and CDR3β sequence similarity. This graph integrates clonotypes from all repertoires into a single structure. (4) Graph-based clustering algorithms (e.g., Leiden or Louvain) identify densely connected groups of clonotypes, defined as Communities on the Joint graph (CJs). The number of cells in each CJ is quantified to form a CJ occupancy matrix k×r, where k indexes CJs and r indexes repertoires. This matrix is used to fit a hierarchical Bayesian Dirichlet-Multinomial model. The model estimates posterior distributions for relative CJ occupancies (β coefficients), providing mean estimates along with 95% credible intervals (CIs).