Library Prep & Sequencing
The library preparation strategy determines what you can learn from BCR sequencing data. Different methods capture different parts of the antibody sequence, introduce different biases and require different analysis pipelines.
Starting material
Section titled “Starting material”BCR sequencing starts with B cells. The most common sources are:
- PBMCs (peripheral blood mononuclear cells): easy to collect from a blood draw. Contains circulating B cells.
- Sorted B cells: flow-sorted populations like naive B cells, memory B cells or plasmablasts. Gives a focused view of a specific compartment.
- Tissue biopsies: lymph node, gut or tumour tissue. Captures tissue-resident B cells not found in blood.
You can extract either RNA or DNA from these cells.
RNA-based methods capture expressed antibodies. They reflect what the B cell is actively producing. RNA is more abundant than DNA for highly expressed genes, so you get better sensitivity. Most BCR sequencing uses RNA.
DNA-based methods capture all rearranged immunoglobulin genes, including non-expressed rearrangements. DNA is more stable than RNA. Some clinical assays for minimal residual disease use DNA.
Amplicon-based methods
Section titled “Amplicon-based methods”The most common approach uses multiplex PCR to amplify the antibody variable region.
Primers target two regions:
- Forward primers: bind to the leader sequence or framework region 1 of V genes. You need many forward primers because there are dozens of V gene families. A typical human heavy chain primer set contains 6 to 12 forward primers.
- Reverse primers: bind to the constant region. One primer per isotype. This captures isotype information along with the variable region sequence.
The PCR product spans from the V gene through CDR3 to the beginning of the constant region. This is typically 400 to 500 base pairs for the heavy chain.
5’RACE-based methods
Section titled “5’RACE-based methods”5’RACE (Rapid Amplification of cDNA Ends) avoids the need for V gene primers entirely.
The protocol works like this:
- Reverse transcribe mRNA using a constant region primer.
- The reverse transcriptase adds a template-switching oligo at the 5’ end of the cDNA.
- PCR amplifies using the template-switching sequence as the forward primer and a constant region primer as the reverse primer.
Because no V gene primers are used, there is no primer bias in V gene capture. Every V gene is amplified with equal efficiency. This makes 5’RACE methods better for quantitative V gene usage analysis.
The tradeoff is lower sensitivity compared to multiplex PCR. 5’RACE captures fewer unique sequences from the same amount of starting material.
Protocols from Galson and colleagues and from the Briney lab use this approach.
UMIs: Unique Molecular Identifiers
Section titled “UMIs: Unique Molecular Identifiers”UMIs are short random nucleotide sequences added to each cDNA molecule during reverse transcription. A typical UMI is 12 to 15 nucleotides long.
Every original mRNA molecule gets a unique UMI tag before PCR amplification. After sequencing, you group reads by their UMI. All reads sharing the same UMI came from the same original molecule.
UMIs solve two problems:
- PCR duplicate removal: without UMIs, you cannot distinguish true biological duplicates from PCR artifacts. UMIs let you collapse PCR duplicates into a single consensus sequence per original molecule.
- Error correction: reads from the same UMI can be used to build a consensus sequence. Sequencing errors are random, so they get voted out. The consensus is more accurate than any single read.
Bulk vs single-cell sequencing
Section titled “Bulk vs single-cell sequencing”Bulk BCR sequencing extracts RNA from a pool of B cells and sequences them together. You get individual heavy chain sequences and individual light chain sequences, but you cannot determine which heavy chain was paired with which light chain in any given cell.
Single-cell BCR sequencing captures heavy and light chains from the same cell. The 10x Genomics Chromium 5’ V(D)J kit is the most widely used platform. Each cell is captured in a droplet with a barcoded gel bead. The cell barcode links the heavy and light chain sequences.
| Feature | Bulk | Single-cell |
|---|---|---|
| Chain pairing | No | Yes |
| Cells per experiment | 100,000+ | 1,000 to 10,000 |
| Cost per cell | Very low | Higher |
| Depth per cell | High | Lower |
| Best for | Repertoire diversity, clonal tracking | Antibody discovery, paired chain analysis |
Bulk sequencing captures far more cells and provides deeper coverage of the repertoire. Single-cell sequencing gives paired chain information that bulk cannot provide.
For many studies, the right choice is both. Use bulk sequencing for broad repertoire profiling and single-cell for targeted antibody discovery.
Sequencing platforms
Section titled “Sequencing platforms”Illumina MiSeq with 2x300 bp paired-end reads is the standard platform for amplicon-based BCR sequencing. The 300 bp reads from each end overlap in the middle of the amplicon, allowing assembly of the full variable region. A single MiSeq run produces 15 to 25 million read pairs.
Illumina NovaSeq or NextSeq can be used for higher throughput when sequencing many samples. Read lengths of 2x150 bp are common on these platforms but may not be long enough to cover the full variable region. Check whether your amplicon length fits within the read length.
Oxford Nanopore produces long reads that cover the entire variable region in a single read. Error rates are higher than Illumina, but UMI consensus correction can compensate. Long reads are especially useful for full-length heavy chain sequencing including the constant region.
PacBio also produces long, high-accuracy reads. The HiFi mode achieves accuracy above 99.9%. Cost per read is higher than Illumina.
AIRR standards
Section titled “AIRR standards”The Adaptive Immune Receptor Repertoire (AIRR) community defines data standards for BCR and TCR sequencing. These standards ensure that data and results are interoperable between different tools and labs.
Key AIRR standards include:
- AIRR Data Format: a tab-separated file format for annotated receptor sequences. Each row is one sequence. Columns include V gene call, D gene call, J gene call, CDR3 sequence, mutation count and more.
- MiAIRR: Minimum Information about an AIRR-seq Experiment. Defines the metadata you should record and report. Includes sample source, cell type, library method, primer sets and sequencing platform.
- AIRR Data Commons: repositories for sharing AIRR-seq data. iReceptor and the VDJServer provide searchable databases of published repertoires.
Summary
Section titled “Summary”- RNA-based amplicon PCR is the most common library prep method. It captures expressed antibodies with good sensitivity but introduces primer bias.
- 5’RACE methods eliminate primer bias at the cost of lower sensitivity. Use them when quantitative V gene usage matters.
- UMIs are essential for accurate quantification and error correction. Always use a UMI-based protocol if your analysis requires clonal frequency comparisons.
- Bulk sequencing captures more cells at lower cost. Single-cell sequencing provides paired heavy and light chain information.
- Illumina MiSeq 2x300 bp is the standard sequencing platform for BCR amplicon data.
- Follow AIRR standards for data formatting and metadata reporting.