Skip to content

Repertoire Analysis

After preprocessing and quality filtering, TCR repertoire analysis answers questions about which T cells are present, how diverse the repertoire is, which clones have expanded and what antigens they might recognise.

The first analysis step is V(D)J assignment. Each TCR sequence is mapped back to the germline V, D and J gene segments used in its recombination.

IgBLAST handles this for TCR sequences just as it does for BCR. It aligns each sequence against the IMGT reference database of germline TCR gene segments. The output includes:

  • Which V gene was used (for example TRBV6-1)
  • Which D gene was used (for example TRBD1)
  • Which J gene was used (for example TRBJ2-1)
  • The CDR3 nucleotide and amino acid sequence
  • Framework and CDR region boundaries

MiXCR is another popular tool for TCR V(D)J assignment. It performs alignment, error correction and clonotype assembly in a single workflow. MiXCR is particularly common in studies using the Adaptive immunoSEQ platform.

V gene usage analysis counts how often each V gene appears in the repertoire.

In a healthy human, the TRBV gene usage distribution is not uniform. Some V genes are used far more often than others. Common dominant families include TRBV20-1, TRBV5-1 and TRBV7-2.

V gene usage shifts in disease:

  • TRBV20-1 is frequently enriched in CMV-specific CD8+ T cell responses.
  • TRBV19 is associated with responses to the influenza M1 peptide presented by HLA-A2.
  • TRBV6-5 and TRBV27 are found in many public TCR responses to SARS-CoV-2 epitopes.

V gene usage is typically visualised as bar charts comparing V gene frequencies across samples, conditions or timepoints. The chart below comes from the committed toy subject used across this section: one AIRR table before a vaccination and one twenty-eight days after, shaped like airrflow output. Every TRBV gene detected in either sample gets two bars, and the post-vaccination column drops for several genes because the expanded clones now dominate the counts:

library(immunarch)
library(ggplot2)
# The airrflow changeo tables for one subject, t0 and t28.
imm <- repLoad(c("sample01_airr.tsv", "sample02_airr.tsv"))
# Sequences per germline V gene in each sample.
gu <- geneUsage(imm$data, .gene = "hs.trbv")
ggsave(
"outputs/vgene-usage.png",
vis(gu),
width = 7.5, height = 4.5, dpi = 130, bg = "white"
)

V gene usage of the toy subject, paired bars per TRBV gene before and after vaccination. Most genes hold between fifty and ninety sequences at both timepoints, a handful drop in the second sample as the expanded clones take over the count, and TRBV2 sits near zero on three records.

The CDR3 region is the most variable part of the TCR. It spans the V-D-J junction and makes the most contacts with the peptide in the peptide-MHC complex. CDR3 sequence largely determines antigen specificity.

Key CDR3 properties:

  • Length distribution: TCR beta CDR3 is typically 12 to 16 amino acids. The distribution is narrower than BCR heavy chain CDR3. CDR3 length affects which peptide-MHC complexes the TCR can bind.
  • Amino acid composition: charge, hydrophobicity and specific residue preferences vary by V gene and by the antigen being recognised.
  • Public CDR3 sequences: identical CDR3 sequences found in unrelated individuals. Public TCRs are far more common than public BCRs. They arise from biased V(D)J recombination and convergent selection by common pathogens.

Public TCR databases like VDJdb and McPAS-TCR catalog known TCR-antigen associations. If your expanded clonotypes match entries in these databases, you have candidate antigen assignments without any experimental validation.

The spectratype plot shows the length distribution directly. TCR beta CDR3s concentrate between twelve and seventeen amino acids, and both timepoints of the toy subject land in exactly that window:

# CDR3 amino acid length distributions per sample.
sp1 <- as.data.frame(spectratype(imm$data[[1]], .col = "aa"))
sp2 <- as.data.frame(spectratype(imm$data[[2]], .col = "aa"))
sp1$Sample <- "day 0"
sp2$Sample <- "day 28"
spectra <- rbind(sp1, sp2)
p <- ggplot(spectra, aes(factor(Length), Val, fill = Sample)) +
geom_col(position = "dodge") +
labs(x = "CDR3 length, amino acids", y = "Sequences")
ggsave(
"outputs/cdr3-lengths.png",
p,
width = 7, height = 4, dpi = 130, bg = "white"
)

CDR3 amino acid length distributions for the two timepoints, dodged bars per length. Both samples span twelve to seventeen residues with the peak at fourteen, and the post-vaccination bars sit slightly lower because its sequences concentrate inside a few expanded clones.

A clonotype is defined by its CDR3 sequence and V/J gene usage. T cells sharing the same clonotype descended from the same V(D)J recombination event.

Clonal expansion occurs when a T cell recognises antigen and proliferates. In sequencing data, expanded clones appear as groups of many sequences with identical CDR3 regions.

Key metrics:

  • Clone size distribution: plot as a rank-abundance curve. Most clonotypes appear once. A few dominate.
  • Top clone fraction: the percentage of the repertoire occupied by the largest clone. Values above 5% indicate strong expansion.
  • Clonality index: 1 minus the normalised Shannon entropy. Ranges from 0 (perfectly even) to 1 (single clone dominates). Also called the productive clonality in the Adaptive immunoSEQ framework.
  • Gini index: measures inequality in clone sizes. Near 0 means all clones are equal size. Near 1 means a few clones dominate.

Clonal expansion is expected during active infection, after vaccination and in the tumour microenvironment. In T cell lymphomas, a single malignant clone can dominate the entire repertoire.

Two immunarch calls put numbers behind those metrics on the toy subject. The top-clone table shows what share of the repertoire each rank class holds, and the clonal-space plot bins every sequence by clone size:

# Share of the repertoire held by ranked clone groups.
rc <- repClonality(imm$data, .method = "top")
ggsave(
"outputs/top-clones.png",
vis(rc),
width = 5.5, height = 4.5, dpi = 130, bg = "white"
)
# Repertoire space held by each expanded-clone category.
cs <- repClonality(imm$data, .method = "homeo")
ggsave(
"outputs/clonal-space.png",
vis(cs),
width = 6.5, height = 4.5, dpi = 130, bg = "white"
)

Top-clone proportions for the two timepoints as stacked bars. The ten largest clonotypes hold about two percent of the repertoire at day zero and about ten percent at day twenty-eight, visible as the growth of the red rank class.

Clonal space homeostasis for the two timepoints as stacked bars. The hyperexpanded class appears only in the post-vaccination sample, where it holds about seven percent of the repertoire space.

A major advantage of TCR sequencing is longitudinal clonal tracking. Because TCR sequences are fixed and do not mutate, you can track the exact same clone across timepoints.

Common tracking analyses:

  • Pre/post vaccination: which clones expand after vaccination? Do they persist at later timepoints?
  • Tumour vs blood: which TIL clones are also found in peripheral blood? These could serve as liquid biopsy biomarkers.
  • Treatment response: which clones expand or contract during immunotherapy?

The toy subject makes the tracking plot small enough to read at a glance. Four clonotypes, three of them real VDJdb records, all rise by roughly an order of magnitude between the two timepoints:

# Follow four clonotypes across both timepoints.
tracked <- trackClonotypes(
imm$data,
c(
"CASSLAPGATNEKLFF", "CAASTGIYGYTF",
"CAGGTGETSPGELFF", "CAAGTRTDTQYF"
),
.col = "aa"
)
tracked_df <- as.data.frame(tracked)
names(tracked_df)[1] <- "cdr3"
long_df <- data.frame(
cdr3 = rep(tracked_df$cdr3, 2),
timepoint = rep(c("day 0", "day 28"), each = nrow(tracked_df)),
proportion = c(tracked_df[[2]], tracked_df[[3]])
)
p <- ggplot(long_df, aes(timepoint, proportion, group = cdr3)) +
geom_line() +
geom_point(aes(color = cdr3), size = 2.5) +
labs(x = NULL, y = "Fraction of repertoire", color = "CDR3")
ggsave(
"outputs/tracked-clones.png",
p,
width = 7, height = 4.5, dpi = 130, bg = "white"
)

Slope chart following four expanded clonotypes across the two timepoints. Every line rises from below half a percent of the repertoire to between three tenths and three point seven percent, with each clonotype in its own colour.

Repertoire diversity measures how varied the T cell pool is. The same metrics used for BCR analysis apply to TCR data.

  • Richness: total number of unique clonotypes. Sensitive to sequencing depth.
  • Shannon index: accounts for both the number of clonotypes and their relative abundance.
  • Simpson index: the probability that two randomly chosen sequences belong to different clones. Less sensitive to rare clonotypes.
  • Hill diversity profiles: a family of indices parameterised by q. Plotting across q values gives a complete picture from richness-weighted to evenness-weighted diversity.

On the toy subject the Hill profile is the clearest picture of what vaccination did. Day zero holds a broad repertoire whose effective size barely moves with q; day twenty-eight starts lower and collapses towards the handful of expanded clones as q grows:

# Diversity as a function of the order parameter q.
dv <- repDiversity(imm$data, .method = "hill")
ggsave(
"outputs/hill-diversity.png",
vis(dv),
width = 7, height = 4.5, dpi = 130, bg = "white"
)

Hill diversity profiles for both timepoints. The day-zero curve starts near seven hundred and eighty effective clones at q=1 and eases to about six hundred by q=6, while the day-twenty-eight curve collapses from about six hundred and forty to around sixty by q=3 and stays there.

Sequencing depth affects diversity estimates. Always compare diversity between samples at the same rarefied depth.

Rarefaction curves subsample to equal depth and recalculate diversity. If the curve has plateaued, you have captured most of the diversity. If it is still rising, deeper sequencing would find more clonotypes.

This is where TCR analysis goes beyond what BCR analysis can currently offer. Computational tools can predict what antigen a TCR recognises based on its CDR3 sequence. This emerging field connects repertoire sequencing to functional immunology.

TCRs with similar CDR3 sequences often recognise the same peptide-MHC complex. This is because the CDR3 loop makes direct contact with the peptide. Structurally similar CDR3 loops bind similar peptides. This principle of TCR convergence enables sequence-based prediction of specificity.

BCR specificity prediction is harder because antibodies use all six CDR loops to bind antigen and somatic hypermutation continuously changes the binding surface. TCR specificity is more concentrated in CDR3, and the sequence is fixed after thymic selection.

Prediction tools rely on databases of experimentally validated TCR-antigen pairs.

Database Contents Access
VDJdb Curated TCR-peptide-MHC associations from literature vdjdb.cdr3.net
McPAS-TCR Pathology-associated TCR sequences friedmanlab.weizmann.ac.il/McPAS-TCR
IEDB Immune epitope database with some TCR annotations iedb.org
10x Genomics datasets Public single-cell V(D)J datasets 10xgenomics.com/datasets
PIRD Pan immune repertoire database db.cngb.org/pird

VDJdb is the most widely used. It contains tens of thousands of TCR-peptide-MHC triplets. Most entries are for viral epitopes from CMV, EBV, influenza and SARS-CoV-2.

These tools group TCRs by CDR3 similarity and look for clusters enriched for a specific antigen.

GLIPH2 (Grouping of Lymphocyte Interactions by Paratope Hotspots) identifies TCR specificity groups by finding shared CDR3 motifs. It looks for short amino acid patterns of 2 to 4 residues that are statistically enriched compared to a background repertoire. TCRs sharing a motif are grouped into a specificity cluster. GLIPH2 can process large datasets and does not require pre-existing antigen labels.

Terminal window
# Example GLIPH2 input: tab-separated CDR3b and TRBV gene
# CDR3b TRBV patient_id
# CASSLAPGATNEKLFF TRBV5-1 patient01
# CASSYSRGSGANVLTF TRBV6-1 patient01

TCRdist3 calculates a distance metric between TCRs based on CDR1, CDR2 and CDR3 amino acid similarity. It uses a substitution matrix trained on TCR structural data. TCRs below a distance threshold are grouped into neighbourhoods. TCRdist3 can identify convergent TCR responses across individuals and map them to known epitopes.

ClusTCR uses the Faiss library to perform fast, approximate clustering of millions of CDR3 sequences. It is designed for large-scale repertoire datasets where pairwise distance calculations are computationally prohibitive.

Machine learning models trained on TCR-antigen binding data can predict specificity for individual TCR sequences.

DeepTCR is a deep learning framework for TCR repertoire analysis. It uses convolutional neural networks to learn sequence features from CDR3 and V gene usage. DeepTCR can classify TCR repertoires by disease state and predict antigen specificity at the single-sequence level.

ERGO-II predicts whether a given TCR-peptide pair will bind. It takes a CDR3 beta sequence and a peptide sequence as input and returns a binding probability. ERGO-II uses autoencoder-based representations of both the TCR and the peptide. It was trained on VDJdb data and performs best on epitopes well represented in the training set.

NetTCR-2.1 predicts TCR-peptide binding for specific HLA alleles. It uses convolutional neural networks trained on peptide-MHC tetramer sorting data. NetTCR is epitope-specific: you train or select a model for each target peptide-MHC combination.

pMTnet predicts TCR binding to peptide-MHC class I complexes. It can handle novel peptides not seen during training by using a transfer learning approach from protein structure. pMTnet is particularly useful for predicting T cell recognition of tumour neoantigens.

TITAN (TCR Interaction with Transformer Attention Networks) uses attention-based neural networks to predict TCR-epitope binding. The attention mechanism highlights which CDR3 positions are most important for binding each specific peptide.

Key limitations to keep in mind:

  • Training data bias: most validated TCR-antigen pairs come from a handful of immunodominant viral epitopes. Predictions for understudied epitopes are less accurate.
  • MHC restriction: specificity depends on both the peptide and the MHC allele. Most tools require MHC information for accurate predictions. Some tools like ERGO-II work without it but have lower accuracy.
  • Paired chain information: the alpha chain contributes to specificity. Tools using only the beta chain CDR3 miss this contribution. Paired alpha-beta data improves predictions.
  • Validation: experimentally confirm computational predictions using peptide-MHC tetramers, ELISPOT or functional killing assays.
Tool Purpose Input
IgBLAST V(D)J assignment FASTA sequences
MiXCR V(D)J assignment and clonotype assembly FASTQ files
immunarch (R) Repertoire analysis, diversity, gene usage AIRR tables, immunoSEQ output
Scirpy (Python) Single-cell TCR analysis with Scanpy AnnData objects
GLIPH2 Specificity group clustering CDR3 + V gene tables
TCRdist3 Distance-based clustering CDR3 + V/J genes
ClusTCR Fast large-scale CDR3 clustering CDR3 sequences
DeepTCR Deep learning repertoire classification CDR3 + V gene, optional HLA
ERGO-II TCR-peptide binding prediction CDR3 + peptide sequences
NetTCR-2.1 Epitope-specific binding prediction CDR3 + peptide + HLA
pMTnet Pan-peptide binding prediction CDR3 + peptide-MHC
TITAN Attention-based binding prediction CDR3 + peptide sequences
  • V(D)J assignment with IgBLAST or MiXCR identifies germline gene usage and CDR3 sequences.
  • V gene usage and CDR3 properties describe the composition of the T cell repertoire.
  • Clonal expansion analysis identifies antigen-driven T cell responses. Clones can be tracked longitudinally because TCR sequences do not mutate.
  • Diversity metrics quantify how broad or focused the repertoire is. Always account for sequencing depth with rarefaction.
  • Antigen specificity prediction uses CDR3 sequence similarity and deep learning to predict what a TCR recognises. Tools like GLIPH2, TCRdist3, DeepTCR, ERGO-II and pMTnet connect sequence data to function.
  • Reference databases like VDJdb and McPAS-TCR provide experimentally validated TCR-antigen associations for comparison.