Bioinformatics File Formats
Every pipeline is a chain of tools, and each tool reads one format and writes another. A sequencer writes FASTQ, an aligner turns it into BAM, a variant caller turns that into VCF, and a quantification tool turns it into a count matrix. The same pattern holds off the sequencer. A cytometer writes FCS, a qPCR machine writes a vendor export or RDML, a mass spectrometer writes a vendor raw file that becomes mzML, and an Olink or Luminex run arrives as a plate reader export.
If you do not know the format, you cannot tell whether a file is correct or corrupted, and you cannot debug a pipeline that fails silently.
The formats at a glance
Section titled “The formats at a glance”The table below is the map of this section. File sizes vary with sequencing depth, panel size and organism.
| Format | What it stores | Typical size | Written by | Read by |
|---|---|---|---|---|
| FASTQ | Raw sequencing reads | 1-50 GB | Sequencer (Illumina, ONT) | Aligners (STAR, BWA) |
| SAM/BAM | Aligned reads | 2-100 GB | Aligners | Variant callers, featureCounts |
| CRAM | Aligned reads, reference compressed | 0.5-30 GB | samtools | Same as BAM |
| VCF/BCF | Variants (SNPs, indels) | 10 MB-10 GB | GATK, bcftools | Annotation and filtering tools |
| BED | Genomic intervals | KB-MB | Peak callers, scripts | bedtools, IGV |
| GFF/GTF | Gene annotations | 50-500 MB | Ensembl, GENCODE | STAR, featureCounts |
| H5AD/MTX | Single-cell count matrices | 50 MB-10 GB | Cell Ranger, Scanpy | Scanpy, Seurat |
| FCS | Per-event cytometry measurements | 1-500 MB | Cytometers | flowCore, FlowJo |
| RDML, vendor exports | qPCR amplification and Cq | 0.1-50 MB | qPCR instruments | RDML, vendor software |
| CEL, IDAT | Microarray intensities | 5-50 MB | Array scanners | affy, oligo, limma |
| mzML, mzXML | Mass spectra | 0.1-10 GB | Vendor converters | mzR, Spectra |
| NPX export | Olink protein concentrations | 1-100 MB | Olink software | OlinkAnalyze |
| xPONENT export | Luminex bead fluorescence | 1-50 MB | xPONENT, Bio-Plex | drLumi, base R |
| POD5, BAM | Long-read signal and reads | 10-500 GB | ONT sequencers | pod5, dorado, samtools |
You do not need to memorise this table. Each step in a workflow has specific input and output formats, and the pages below cover each one.
The pages
Section titled “The pages”The first seven pages cover the sequencing formats. The rest cover the assay formats that arrive from bench instruments.
| Page | What it covers |
|---|---|
| FASTQ: raw reads | the four-line record, Phred scores, paired-end files, gzip |
| SAM, BAM and CRAM | the 11 alignment fields, FLAG, MAPQ, CIGAR, sorting and indexing |
| VCF: variants | the variant columns, INFO and FORMAT fields, genotypes, gVCF |
| BED, GFF and GTF | intervals and gene models, and the 0-based versus 1-based trap |
| Single-cell formats | the 10x matrix directory, H5AD, Loom and Seurat RDS |
| Inspecting files | the command line checks to run before any analysis |
| FCS: flow cytometry | the binary layout, the keyword segment, reading with flowCore |
| qPCR files | RDML and the Bio-Rad, QuantStudio and LightCycler exports |
| Microarray files | CEL and IDAT binaries, and the GEO series matrix text format |
| Mass spectrometry | mzML and mzXML, the vendor raw formats, reading with mzR |
| Olink NPX | the NPX export columns, LOD and QC flags, reading with OlinkAnalyze |
| Luminex | the xPONENT export, bead counts and median fluorescence |
| Long-read sequencing | POD5 signal, modified-base BAM tags, PacBio HiFi |
How the code runs here
Section titled “How the code runs here”The sequencing pages use command line tools such as samtools and bcftools, and
their shell examples are illustrative. The assay pages carry R examples, and every R
example on them runs in a pinned Podman container in the companion code repository
under guides/file-formats/. Every printed output on those pages comes from that
run. The readers for these formats live in R and Bioconductor, so the section is R
only. Python is named in the prose where it is the standard tool.