Skip to content

Running nf-core/airrflow

nf-core/airrflow is a Nextflow pipeline that processes raw BCR sequencing data from FASTQ files to annotated clonotypes. It handles UMI processing, V(D)J annotation, quality filtering and clonal analysis in a single reproducible workflow.

The pipeline takes raw FASTQ files from BCR sequencing experiments and produces AIRR-formatted output with full V(D)J annotations, mutation analysis and clonal assignments. It wraps well-established tools from the Immcantation framework into a standardised Nextflow pipeline.

airrflow supports multiple library preparation protocols. It handles both UMI-based and non-UMI protocols, multiplex PCR and 5’RACE approaches, and paired-end or single-end reads. You specify your protocol through pipeline parameters, and airrflow selects the correct processing steps.

The pipeline requires a CSV samplesheet listing your samples. Each row is one sample with paths to the FASTQ files and metadata.

sample_id,filename_R1,filename_R2,subject_id,species,tissue,pcr_target_locus,treatment,collection_time_point_relative
sample01,sample01_R1.fastq.gz,sample01_R2.fastq.gz,subject01,human,blood,ig,none,0
sample02,sample02_R1.fastq.gz,sample02_R2.fastq.gz,subject01,human,blood,ig,none,7
sample03,sample03_R1.fastq.gz,sample03_R2.fastq.gz,subject02,human,blood,ig,none,0

The pcr_target_locus column specifies what was sequenced. Use ig for immunoglobulin (BCR) or tr for T cell receptor. The subject_id column is important because clonal analysis groups sequences within the same subject.

The most important parameter is --library_generation_method. This tells the pipeline which protocol was used for library preparation. Common values include:

Value Protocol
specific_pcr_umi Multiplex PCR with UMIs
race_uid 5’RACE with UMIs
specific_pcr Multiplex PCR without UMIs
dt_5p_race 5’RACE with template switching

Other key parameters:

  • --umi_length: length of the UMI sequence. Depends on your protocol. Common values are 12 or 15.
  • --vprimers: FASTA file with V gene primer sequences.
  • --cprimers: FASTA file with constant region primer sequences.
  • --race_linker: linker sequence for RACE protocols.
  • --index_file: set to true if UMIs are in a separate index read.

airrflow runs four major processing stages.

pRESTO handles the raw read processing. For UMI-based protocols, it extracts UMI sequences from the reads, groups reads by UMI and builds consensus sequences. This collapses PCR duplicates so each consensus represents one original mRNA molecule.

pRESTO also masks primer sequences, filters reads by quality score and assembles paired-end reads into full-length sequences. The output is a set of high-quality, deduplicated sequences ready for annotation.

Each preprocessed sequence is aligned against the IMGT germline reference database using IgBLAST. This identifies which V, D and J gene segments were used in the original V(D)J recombination event.

IgBLAST also identifies the CDR and framework regions, determines the reading frame and checks for productive rearrangements. A non-productive sequence has a stop codon or frameshift and cannot encode a functional antibody.

Change-O parses the IgBLAST output into a standardised AIRR format. It filters out non-functional sequences, removes sequences that fail quality thresholds and calculates mutation frequencies.

An important step here is determining the clonal distance threshold. Change-O uses a nearest-neighbour distance approach. It calculates the distance between every pair of sequences based on their junction regions and fits a Gaussian mixture model to distinguish clonal relationships from background noise. The valley between the two distributions becomes the threshold.

Sequences from the same subject are grouped into clones. Two sequences belong to the same clone if they share the same V gene, same J gene and junction sequences within the distance threshold determined in the previous step.

Within each clone, the pipeline constructs germline sequences and calculates per-sequence mutation counts. This information feeds into downstream analyses of somatic hypermutation and affinity maturation.

Here is an example command for running airrflow on a multiplex PCR dataset with UMIs:

Terminal window
nextflow run nf-core/airrflow \
-r 5.1.0 \
-profile podman \
--input samplesheet.csv \
--library_generation_method specific_pcr_umi \
--umi_length 12 \
--vprimers vprimers.fasta \
--cprimers cprimers.fasta \
--outdir results

For running on AWS Batch with Spot instances:

Terminal window
nextflow run nf-core/airrflow \
-r 5.1.0 \
-profile awsbatch \
--input s3://mybucket/samplesheet.csv \
--library_generation_method specific_pcr_umi \
--umi_length 12 \
--vprimers s3://mybucket/vprimers.fasta \
--cprimers s3://mybucket/cprimers.fasta \
--outdir s3://mybucket/results \
-work-dir s3://mybucket/work

The results directory contains several subdirectories:

results/
├── presto/ # pRESTO preprocessing logs and stats
├── igblast/ # Raw IgBLAST output
├── changeo/ # Parsed AIRR tables, threshold plots
├── clonal_analysis/ # Clone assignments
├── repertoire_analysis/ # Gene usage, diversity plots
├── multiqc/ # Summary report
└── pipeline_info/ # Execution logs and resource usage

The most important output files are the AIRR-formatted TSV files in changeo/. Each row is one sequence with columns for V/D/J gene assignment, CDR3 sequence, mutation count, clone ID and isotype. These files are ready for downstream analysis with Alakazam, immunarch, pandas or any other tool that reads AIRR tables.

To show what those tables hold, this page works with one committed to the companion repository: 656 heavy-chain sequences from one sample, carrying IMGT gene calls, CDR3 amino acids, isotypes, mutation frequencies and clone identifiers.

Loading it needs nothing more than pandas:

import pandas as pd
# One AIRR table as airrflow writes it into changeo/.
airr_df = pd.read_csv("sample01_airr.tsv", sep="\t")
# The annotation columns every row carries.
print(airr_df.columns.tolist())
# Ten most frequent V genes.
v_usage = airr_df["v_call"].value_counts().head(10)
print(v_usage)
['sequence_id', 'sequence', 'rev_comp', 'sequence_alignment', 'germline_alignment', 'locus', 'v_call', 'd_call', 'j_call', 'c_call', 'junction', 'junction_aa', 'duplicate_count', 'productive', 'mu_freq', 'clone_id', 'cigar', 'v_cigar', 'd_cigar', 'j_cigar']
v_call
IGHV3-23 84
IGHV3-30 69
IGHV3-33 69
IGHV1-69 61
IGHV3-13 59
IGHV1-2 59
IGHV3-11 55
IGHV4-39 46
IGHV3-48 46
IGHV4-34 43
Name: count, dtype: int64

V gene usage counts how many consensus sequences each germline gene produced. In this toy repertoire the IGHV3 family sits at the top, which matches what a healthy adult heavy-chain repertoire usually looks like.

The mutation frequency column separates naive from experienced antibodies. Grouping it by isotype shows the effect of class switching directly:

# Class-switched sequences carry more somatic hypermutation.
mutation_by_isotype = (
airr_df.groupby("c_call")["mu_freq"].agg(["count", "mean"]).round(4)
)
print(mutation_by_isotype)
count mean
c_call
IGHA1 72 0.0671
IGHD 73 0.0108
IGHE 89 0.0658
IGHG1 132 0.0641
IGHG3 75 0.0678
IGHM 217 0.0098

The IgM rows carry almost no mutations. The switched isotypes carry several times more, because those B cells passed through germinal centres where AID introduced mutations across the variable region.

Export the AIRR-formatted TSV files into R or Python for detailed repertoire analysis. The Immcantation suite provides specialised tools:

  • Alakazam: diversity analysis, gene usage plots, CDR3 physicochemical properties.
  • SHazaM: somatic hypermutation analysis, selection pressure estimation.
  • SCOPer: refined clonal clustering using spectral methods.
  • Dowser: B cell phylogenetics and lineage tree construction.

In Python, you load the AIRR tables directly with pandas, as the examples above do, and reach for packages like scirpy when single-cell context matters. The TCR section runs a full scirpy example in its airrflow guide.

Wrong UMI length. If you specify the wrong --umi_length, pRESTO will extract the wrong bases as UMIs. This leads to poor consensus building and low sequence recovery. Check your protocol documentation carefully.

Incorrect primer sequences. The primer FASTA files must match exactly what was used in the library preparation. Mismatched primers cause pRESTO to fail at the primer masking step. You will see a large drop in read counts at this stage.

Clonal threshold selection. The Gaussian mixture model sometimes fails to find a clear bimodal distribution, especially with small datasets or low diversity samples. Check the threshold plot in changeo/. If the threshold looks wrong, you can set it manually with --set_cluster_threshold.

Memory for large datasets. The clonal assignment step loads all sequences from one subject into memory. For subjects with millions of sequences, this can exceed available RAM. Consider downsampling or using a machine with more memory.

Database version mismatches. The IMGT reference database version should match what your collaborators or previous analyses used. Different database versions can assign different V gene names to the same sequence.

  • nf-core/airrflow processes BCR sequencing data from raw FASTQ to annotated clonotypes.
  • The --library_generation_method parameter must match your library preparation protocol.
  • The pipeline runs pRESTO for preprocessing, IgBLAST for V(D)J annotation and Change-O for clonal assignment.
  • Output is AIRR-formatted TSV files ready for downstream analysis with Alakazam, immunarch or pandas.
  • Always check the clonal distance threshold plot and primer masking statistics to catch common errors early.