Skip to content

Library Prep & Sequencing

The library preparation strategy determines what you can learn from TCR sequencing data. Different methods target different chains, introduce different biases and require different analysis approaches.

TCR sequencing starts with T cells. The most common sources are:

  • PBMCs (peripheral blood mononuclear cells): easy to collect from a blood draw. Contains circulating T cells. The most common starting material.
  • Sorted T cells: flow-sorted populations like CD4+ helper T cells, CD8+ cytotoxic T cells or regulatory T cells. Gives a focused view of a specific compartment.
  • Tumour biopsies: fresh or FFPE tumour tissue. Captures tumour-infiltrating lymphocytes (TILs) that may not circulate in blood.
  • Other tissues: cerebrospinal fluid, synovial fluid, bronchoalveolar lavage. Captures T cells at disease sites.

You can extract either RNA or DNA from these cells.

RNA-based methods capture expressed TCR genes. They are more sensitive because T cells actively transcribe their TCR genes at high levels. Most TCR sequencing uses RNA.

DNA-based methods capture all rearranged TCR genes, including those from cells that are not actively transcribing. DNA is more stable and works better with degraded samples like FFPE tissue. The Adaptive Biotechnologies immunoSEQ platform uses a DNA-based approach.

Most studies sequence the TCR beta chain only. The beta chain has higher diversity than the alpha chain because it uses V, D and J segments instead of just V and J. The D segment adds substantial junctional diversity to CDR3.

There is also a practical reason. Each T cell has one productive beta chain rearrangement but can have two productive alpha chain rearrangements. This means a single cell can express two alpha chains. Linking a specific alpha chain to its correct beta chain is only possible with single-cell sequencing.

Target Diversity Chains per cell Best for
Beta chain only Higher 1 Repertoire profiling, clonal tracking
Alpha chain only Lower 1 or 2 Specific research questions
Paired alpha-beta Highest Linked Antigen specificity, TCR cloning
Gamma-delta Lower 1 Specialised gamma-delta T cell studies

For antigen specificity prediction, paired alpha-beta sequences are ideal. Both chains contribute to peptide-MHC recognition. However, many prediction tools work with beta chain CDR3 alone because paired data is less available.

The most common approach uses multiplex PCR to amplify the TCR variable region.

Primers target two regions:

  1. Forward primers: bind to the V gene framework regions. You need many forward primers because there are dozens of V gene families. A typical human beta chain primer set contains 20 to 30 forward primers.
  2. Reverse primers: bind to the constant region. The beta chain has two constant region genes (TRBC1 and TRBC2), so you need at least two reverse primers.

The PCR product spans from the V gene through CDR3 to the beginning of the constant region. This is typically 300 to 400 base pairs for the beta chain.

5’RACE (Rapid Amplification of cDNA Ends) avoids the need for V gene primers.

The protocol:

  1. Reverse transcribe mRNA using a constant region primer.
  2. The reverse transcriptase adds a template-switching oligo at the 5’ end of the cDNA.
  3. PCR amplifies using the template-switching sequence as the forward primer and a constant region primer as the reverse primer.

Because no V gene primers are used, there is no primer bias in V gene capture. This makes 5’RACE better for quantitative V gene usage analysis.

The tradeoff is lower sensitivity compared to multiplex PCR. 5’RACE captures fewer unique sequences from the same amount of starting material.

UMIs are short random nucleotide sequences added to each cDNA molecule during reverse transcription. A typical UMI is 12 to 15 nucleotides long.

Every original mRNA molecule gets a unique UMI tag before PCR amplification. After sequencing, reads sharing the same UMI came from the same original molecule.

UMIs solve two problems:

  1. PCR duplicate removal: collapse PCR duplicates into a single consensus sequence per original molecule.
  2. Error correction: reads from the same UMI can be used to build a consensus. Sequencing errors are random, so they get voted out.

Bulk TCR sequencing extracts RNA from a pool of T cells and sequences them together. You get individual beta chain sequences and individual alpha chain sequences, but you cannot determine which alpha chain was paired with which beta chain in a given cell.

Single-cell TCR sequencing captures paired alpha and beta chains from the same cell. The 10x Genomics Chromium 5’ V(D)J kit is the most widely used platform. Each cell is captured in a droplet with a barcoded gel bead. The cell barcode links the alpha and beta chain sequences.

Feature Bulk Single-cell
Chain pairing No Yes
Cells per experiment 100,000+ 1,000 to 10,000
Cost per cell Very low Higher
Depth per cell High Lower
Best for Repertoire diversity, clonal tracking, MRD Antigen specificity, paired chain analysis

Single-cell TCR sequencing can be combined with single-cell gene expression profiling. The same droplet captures both the TCR sequence and the transcriptome. This links TCR identity to T cell phenotype: you can see which clonotype is a cytotoxic effector, which is exhausted and which is a memory cell.

Several commercial platforms provide standardised TCR sequencing workflows:

Platform Method Input Chain Notes
Adaptive immunoSEQ Multiplex PCR from DNA DNA Beta Synthetic spike-in calibration, quantitative
10x Genomics Chromium Single-cell 5’ V(D)J Cells Paired alpha-beta Combined with gene expression
Takara SMARTer 5’RACE from RNA RNA Alpha, beta or both UMI-based, low primer bias
iRepertoire Multiplex PCR RNA or DNA Alpha, beta, gamma, delta Semi-quantitative

The Adaptive immunoSEQ platform is the most widely used for bulk TCR-beta sequencing in clinical and translational research. It reports calibrated clone frequencies and has a large reference database for comparison.

Illumina MiSeq with 2x300 bp paired-end reads is the standard platform for amplicon-based TCR sequencing. The paired reads overlap in the middle of the amplicon, allowing assembly of the full variable region. A single MiSeq run produces 15 to 25 million read pairs.

Illumina NovaSeq or NextSeq provide higher throughput for large sample numbers. Read lengths of 2x150 bp are common and may be sufficient for shorter TCR amplicons, but check whether your full amplicon length is covered.

Oxford Nanopore long reads can cover the entire variable region in a single read. Higher error rates can be compensated by UMI consensus correction.

The Adaptive Immune Receptor Repertoire (AIRR) community defines data standards for both TCR and BCR sequencing. The same AIRR data format and MiAIRR metadata standards apply to TCR data.

Key points for TCR data:

  • The locus field specifies TRB (beta), TRA (alpha), TRG (gamma) or TRD (delta).
  • V/D/J gene names follow IMGT nomenclature: TRBV, TRBD, TRBJ for beta chain genes.
  • The junction and junction_aa fields contain the CDR3 sequence.
  • RNA-based amplicon PCR is the most common library prep for TCR sequencing. It captures expressed receptors with good sensitivity but introduces primer bias.
  • 5’RACE methods eliminate primer bias at the cost of lower sensitivity.
  • Most studies sequence the beta chain only. Single-cell methods provide paired alpha-beta chain information.
  • UMIs are essential for quantitative analysis and error correction.
  • Single-cell sequencing links TCR identity to gene expression phenotype.
  • Commercial platforms like immunoSEQ and 10x Genomics provide standardised, validated workflows.
  • Follow AIRR standards for data formatting and metadata reporting.