Skip to content

TCR Affinity Prediction Tools

Given a CDR3 sequence, can a computer tell you what that T cell recognises? A growing stack of tools says it can, at least sometimes. This page walks through how they work, what they are trained on, and where the whole enterprise runs into walls that better models will not move. It expands on the survey in the repertoire analysis guide, and it is deliberately skeptical: the field moves fast, the benchmarks flatter the tools, and the failure modes are structural.

Two related tasks hide under one name. Specificity grouping asks which TCRs recognise the same thing, without naming the thing. Binding prediction takes a TCR sequence and a candidate peptide and returns a probability. The first is a clustering problem and mostly works. The second is a supervised learning problem, and everything hard about it follows from its training data.

Almost every predictor learns from VDJdb, a curated database of TCR sequences with experimentally verified antigen specificity, supplemented by IEDB and McPAS-TCR. The June 2026 release used on this site holds over a hundred thousand human TRB records alone. Three properties of this data shape every model built on it.

First, coverage is wildly uneven. A handful of immunodominant viral epitopes dominate: CMV pp65 NLVPMVATV, influenza M1 GILGFVFTL, EBV BMLF1 GLCTLVAML, HTLV-1 Tax SLLMWITQV, most of them restricted by HLA-A*02:01. Hundreds of other epitopes have a few records each, or none.

Second, the negative labels do not exist. A TCR-epitope pair is in the database because someone tested it and got a hit. Pairs that were never tested, or tested and failed, are simply absent. Every predictor therefore invents negatives, usually by random pairing, and a random peptide is so obviously not the binder that the task is easier than reality.

Third, the records carry redundancy that inflates benchmark scores. Public clonotypes recur across donors, and near-identical CDR3s appear many times. A model can score well by memorising the abundant sequences rather than learning the recognition rule.

The oldest and still most defensible approach sidesteps learning entirely: group receptors with similar CDR3s and borrow annotations within groups.

GLIPH2 clusters TCRs by shared local motifs, short amino acid patterns at fixed positions in the CDR3, together with V-gene similarity. It needs no antigen labels to run, scales to millions of sequences, and reports specificity groups you can then annotate from a reference database or from convergent responses across unrelated donors. Its 2017 predecessor produced one of the field’s founding observations: T cells from different people carrying near-identical CDR3s against the same epitope.

TCRdist3 computes a distance between receptors from CDR1, CDR2 and CDR3 similarity plus V-gene identity, with radii tuned against structural data. Neighborhoods of similar receptors inherit each other’s known specificities. It is transparent: when two TCRs end up together, you can see exactly which positions drove it.

ClusTCR trades some accuracy for speed, using approximate nearest-neighbour search to cluster very large repertoires in minutes rather than days.

All three share the same ceiling: they only transfer labels between receptors that look alike, so they say nothing about a receptor whose neighbours are unannotated.

A worked example: annotation by nearest neighbour

Section titled “A worked example: annotation by nearest neighbour”

The airrflow guide covers exact matching against VDJdb. This example uses the next annotation tier, similarity transfer.

It counts mismatches between CDR3s of the same length against the complete excerpt. The example uses no new packages.

import pandas as pd
def count_mismatches(cdr3: str, reference_cdr3: str) -> int:
"""Return the count of nonmatching positions in two CDR3 sequences.
Args:
cdr3: Query CDR3 amino acid sequence.
reference_cdr3: Reference CDR3 amino acid sequence.
Returns:
Count of amino acid mismatches.
"""
return sum(a != b for a, b in zip(cdr3, reference_cdr3))
def best_match(
cdr3: str,
reference: pd.DataFrame,
) -> tuple[int | None, str | None]:
"""Return distance and epitope of the closest same-length record.
Args:
cdr3: Query CDR3 amino acid sequence.
reference: VDJdb excerpt with CDR3 and epitope columns.
Returns:
Closest mismatch distance and its epitope, or two None values.
"""
hits = reference[reference["cdr3"].str.len() == len(cdr3)]
if hits.empty:
return None, None
distances = pd.Series(
[count_mismatches(cdr3, ref) for ref in hits["cdr3"]],
index=hits.index,
)
best = distances.idxmin()
return distances.loc[best], hits.loc[best, "antigen.epitope"]
airr_df = pd.read_csv("sample01_airr.tsv", sep="\t")
vdjdb = pd.read_csv("vdjdb_excerpt.txt", sep="\t")
exact = airr_df.merge(vdjdb, left_on="junction_aa", right_on="cdr3")
print(f"exact matches: {exact['sequence_id'].nunique()}")
rows = []
for cdr3 in airr_df["junction_aa"].unique():
distance, epitope = best_match(cdr3, vdjdb)
if distance is not None and distance <= 1:
rows.append({"cdr3": cdr3, "distance": distance, "epitope": epitope})
annotated = pd.DataFrame(rows).sort_values(["distance", "cdr3"])
print(annotated.to_string(index=False))
exact matches: 3
cdr3 distance epitope
CAAGTRTDTQYF 0 NLVPMVATV
CAASTGIYGYTF 0 GILGFVFTL
CAGGTGETSPGELFF 0 GLCTLVAML
CAAGTRTDSQYF 1 NLVPMVATV
CAATTGIYGYTF 1 GILGFVFTL
CAGGTGETTPGELFF 1 GLCTLVAML

The run finds 3 exact matches. It lists 6 CDR3 sequences at distance 0 or 1, with 3 exact records and 3 one-residue variants. Each variant receives its parent epitope, either NLVPMVATV, GILGFVFTL, or GLCTLVAML.

NetTCR is the cleanest baseline of the learned models: a small convolutional network over the CDR3 beta sequence and the peptide, trained epitope by epitope for specific HLA-A*02:01-restricted targets. Version 2 added paired alpha-beta input and showed that pairing helps, modestly.

ERGO-II scores TCR-peptide pairs with embeddings learned from both sequence data and, indirectly, protein language models, and accepts either chain alone or both.

pMTnet adds transfer learning from protein structure: an LM embedding of the peptide feeds a trained module that outputs binding odds, which lets it reach peptides with little or no training data, including neoantigens.

TITAN applies attention across the paired CDR3s and the peptide, and its attention maps give a rough view of which loop positions matter for which target, which is useful for triage even when the probability itself is not.

Reported performance on well-covered epitopes is genuinely good, with AUROCs around 0.9 in-distribution. Two caveats follow immediately. Performance degrades steeply as epitope coverage thins, and the headline numbers partly reflect the redundancy described above.

A 2024 reanalysis by Meynard-Piganeau and colleagues made the uncomfortable point precisely: once you remove near-duplicate CDR3s between train and test sets, a simple nearest-neighbour method matches or beats the deep learning models on VDJdb-derived benchmarks. The deep models were largely rediscovering sequence similarity in the training data. This does not make them useless; it means their apparent sophistication outruns their demonstrated mechanism, and any paper reporting a new predictor should be read with deduplicated benchmarks as the price of entry.

The second evaluation trap is epitope-level generalization. Models score well when the test epitope appeared during training with other TCRs, and much worse when it did not. Predicting whether a TCR binds a truly novel peptide remains close to an open problem, whatever the marketing slides say.

Some constraints are not model defects but facts of biology and data:

  • MHC restriction. Specificity is a property of a TCR-peptide-MHC triple, not a pair. Tools that ignore the MHC allele are averaging over different molecular contexts.
  • Missing chains. Bulk sequencing gives beta only. Alpha contributes to recognition, so predictions on beta-alone inputs carry an irreducible error floor.
  • Unknown negatives. Absence of evidence is not evidence of non-binding, so all invented negatives encode assumptions somebody should state out loud.
  • Conformation and affinity. Most databases record binding or activation assays at arbitrary thresholds, not affinities. Calling these tools affinity predictors overstates what was measured.
  • Bias in what gets studied. Models reflect decades of A*02-heavy, virus-heavy research. Autoimmune, tumour and rare-HLA space is thin.

None of these improve with more layers.

A workflow that stays honest looks like this. Start with exact matching against high-confidence VDJdb records, score two or higher, which is cheap and interpretable and is what the airrflow guide demonstrates on real excerpt data. Then use GLIPH2 or TCRdist3 to lift annotations onto similar unannotated receptors, checking that the transferred label makes mechanistic sense. Only then reach for NetTCR, ERGO-II or pMTnet, as a ranker of hypotheses among many candidates rather than an oracle. Whatever comes out goes to the bench: tetramer staining, ELISPOT or activation assays are the only ground truth, and a prediction that survives them has earned its keep.

  • Specificity grouping works because TCRs recognising the same epitope really do resemble each other; GLIPH2, TCRdist3 and ClusTCR exploit that directly.
  • Binding predictors learn from VDJdb, whose uneven coverage, absent negatives and internal redundancy inflate apparent performance.
  • After removing train-test redundancy, simple nearest-neighbour methods rival deep models; treat new benchmarks skeptically.
  • MHC restriction, missing alpha chains and unknown negatives set floors no architecture removes.
  • Use exact matching and similarity transfer first, learned predictors as hypothesis rankers, and experimental validation always.