Skip to content

Bioinformatics File Formats

Every pipeline is a chain of tools, and each tool reads one format and writes another. A sequencer writes FASTQ, an aligner turns it into BAM, a variant caller turns that into VCF, and a quantification tool turns it into a count matrix. The same pattern holds off the sequencer. A cytometer writes FCS, a qPCR machine writes a vendor export or RDML, a mass spectrometer writes a vendor raw file that becomes mzML, and an Olink or Luminex run arrives as a plate reader export.

If you do not know the format, you cannot tell whether a file is correct or corrupted, and you cannot debug a pipeline that fails silently.

The table below is the map of this section. File sizes vary with sequencing depth, panel size and organism.

Format What it stores Typical size Written by Read by
FASTQ Raw sequencing reads 1-50 GB Sequencer (Illumina, ONT) Aligners (STAR, BWA)
SAM/BAM Aligned reads 2-100 GB Aligners Variant callers, featureCounts
CRAM Aligned reads, reference compressed 0.5-30 GB samtools Same as BAM
VCF/BCF Variants (SNPs, indels) 10 MB-10 GB GATK, bcftools Annotation and filtering tools
BED Genomic intervals KB-MB Peak callers, scripts bedtools, IGV
GFF/GTF Gene annotations 50-500 MB Ensembl, GENCODE STAR, featureCounts
H5AD/MTX Single-cell count matrices 50 MB-10 GB Cell Ranger, Scanpy Scanpy, Seurat
FCS Per-event cytometry measurements 1-500 MB Cytometers flowCore, FlowJo
RDML, vendor exports qPCR amplification and Cq 0.1-50 MB qPCR instruments RDML, vendor software
CEL, IDAT Microarray intensities 5-50 MB Array scanners affy, oligo, limma
mzML, mzXML Mass spectra 0.1-10 GB Vendor converters mzR, Spectra
NPX export Olink protein concentrations 1-100 MB Olink software OlinkAnalyze
xPONENT export Luminex bead fluorescence 1-50 MB xPONENT, Bio-Plex drLumi, base R
POD5, BAM Long-read signal and reads 10-500 GB ONT sequencers pod5, dorado, samtools

You do not need to memorise this table. Each step in a workflow has specific input and output formats, and the pages below cover each one.

The first seven pages cover the sequencing formats. The rest cover the assay formats that arrive from bench instruments.

Page What it covers
FASTQ: raw reads the four-line record, Phred scores, paired-end files, gzip
SAM, BAM and CRAM the 11 alignment fields, FLAG, MAPQ, CIGAR, sorting and indexing
VCF: variants the variant columns, INFO and FORMAT fields, genotypes, gVCF
BED, GFF and GTF intervals and gene models, and the 0-based versus 1-based trap
Single-cell formats the 10x matrix directory, H5AD, Loom and Seurat RDS
Inspecting files the command line checks to run before any analysis
FCS: flow cytometry the binary layout, the keyword segment, reading with flowCore
qPCR files RDML and the Bio-Rad, QuantStudio and LightCycler exports
Microarray files CEL and IDAT binaries, and the GEO series matrix text format
Mass spectrometry mzML and mzXML, the vendor raw formats, reading with mzR
Olink NPX the NPX export columns, LOD and QC flags, reading with OlinkAnalyze
Luminex the xPONENT export, bead counts and median fluorescence
Long-read sequencing POD5 signal, modified-base BAM tags, PacBio HiFi

The sequencing pages use command line tools such as samtools and bcftools, and their shell examples are illustrative. The assay pages carry R examples, and every R example on them runs in a pinned Podman container in the companion code repository under guides/file-formats/. Every printed output on those pages comes from that run. The readers for these formats live in R and Bioconductor, so the section is R only. Python is named in the prose where it is the standard tool.