Skip to content

VCF: Variants

VCF is the standard format for storing genetic variants. It records SNPs, insertions, deletions and structural variants. Every variant calling pipeline produces VCF output, and every downstream annotation or filtering tool reads it.

A variant is a position where a sample differs from the reference sequence. The three common types are SNPs, where one base differs; insertions, where the sample carries extra bases; and deletions, where the sample lacks bases the reference has. Callers such as GATK HaplotypeCaller, DeepVariant and FreeBayes compare aligned reads with the reference and write the differences into a VCF file.

A VCF file has three sections: meta-information lines, one header line, and data lines. The meta-information lines start with ## and define the file version, the reference, and the meaning of the fields used in the data lines.

##fileformat=VCFv4.2
##FILTER=<ID=PASS,Description="All filters passed">
##INFO=<ID=DP,Number=1,Type=Integer,Description="Total Depth">
##INFO=<ID=AF,Number=A,Type=Float,Description="Allele Frequency">
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
##FORMAT=<ID=AD,Number=R,Type=Integer,Description="Allelic depths">
##FORMAT=<ID=GQ,Number=1,Type=Integer,Description="Genotype Quality">

The header line starts with #CHROM and names the columns. Each data line is one variant.

#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1
chr1 12345 rs12345 A G 99 PASS DP=50;AF=0.5 GT:AD:GQ 0/1:25,25:99
Column Name Description
1 CHROM Chromosome name
2 POS 1-based position on the chromosome
3 ID Variant identifier from dbSNP or another database
4 REF Reference allele
5 ALT Alternate allele observed in the sample
6 QUAL Phred-scaled quality of the variant call
7 FILTER PASS when the variant passed all filters
8 INFO Semicolon-separated key-value annotations for the variant
9 FORMAT Colon-separated list of the per-sample fields
10+ SAMPLE Per-sample genotype and quality values

VCF uses 1-based coordinates, so position 12345 is the 12,345th base on the chromosome. BED uses 0-based coordinates, and the difference matters.

INFO packs variant-level annotations into semicolon-separated key-value pairs. The common fields:

Key Meaning Example
DP Total read depth at the position DP=50
AF Allele frequency of the alternate allele AF=0.5
MQ Root mean square mapping quality of the reads MQ=60
QD Quality by depth (QUAL divided by DP) QD=15.2
FS Fisher strand bias p-value FS=0.0
SOR Strand odds ratio SOR=0.7

Different callers add different fields. The ##INFO lines at the top of the file define each one.

The FORMAT column declares which per-sample fields follow, and the sample columns carry the values.

GT is the called genotype, written as allele indices separated by / or |. The index 0 is the REF allele, 1 is the first ALT allele, and so on.

Value Meaning
0/0 Homozygous reference
0/1 Heterozygous
1/1 Homozygous alternate
./. No call (insufficient data)

AD is the allelic depth, the read count supporting each allele. AD=25,25 means 25 reads support the reference and 25 support the alternate. GQ is the Phred-scaled confidence in the genotype call, and GQ 30 means a 1 in 1000 chance the genotype is wrong. DP is the depth for that one sample, which can differ from the INFO DP summed across samples.

FILTER records whether a variant passed quality control. PASS is the value you want. Any other value names the filter that flagged the variant, such as LowQual or StrandBias, and a dot means no filters were applied. GATK’s recommended thresholds for SNPs include QD below 2.0, FS above 60.0 and MQ below 40.0, applied with gatk VariantFiltration and recorded here.

One VCF can carry genotypes for many samples, one column each after FORMAT.

#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1 SAMPLE2 SAMPLE3
chr1 12345 . A G 99 PASS DP=150;AF=0.33 GT:AD:GQ 0/1:25,25:99 0/0:50,0:99 1/1:0,50:99

Here SAMPLE1 is heterozygous, SAMPLE2 is homozygous reference and SAMPLE3 is homozygous alternate. Multi-sample VCFs are the standard output of cohort studies.

A standard VCF records only the positions where a variant was found, so a missing position is ambiguous: the sample may be homozygous reference, or there may be no data at all. A gVCF resolves the ambiguity by also recording the non-variant positions, compressed into blocks with an END tag.

chr1 10001 . A <NON_REF> . . END=10100 GT:DP:GQ 0/0:30:90

This line says positions 10001 to 10100 are all homozygous reference at about 30x depth. gVCFs feed GATK’s joint calling workflow: individual gVCFs combine with GenomicsDBImport or CombineGVCFs, and GenotypeGVCFs genotypes the cohort.

Compress VCF files with bgzip and index them with tabix.

Terminal window
bgzip variants.vcf
tabix -p vcf variants.vcf.gz

This creates variants.vcf.gz and variants.vcf.gz.tbi. The index lets tools jump to a region without reading the whole file, and most tools expect both.

bcftools is the standard toolkit for VCF files.

Terminal window
# View the header
bcftools view -h variants.vcf.gz
# Summary statistics: SNP and indel counts, Ts/Tv ratio, per-sample stats
bcftools stats variants.vcf.gz
# Extract specific fields
bcftools query -f '%CHROM\t%POS\t%REF\t%ALT\t%QUAL\n' variants.vcf.gz | head
# Keep only PASS variants with QUAL of 30 or more
bcftools view -f PASS -i 'QUAL>=30' variants.vcf.gz -o filtered.vcf.gz -Oz
# Extract one region
bcftools view -r chr1:10000-20000 variants.vcf.gz
# Extract two samples
bcftools view -s SAMPLE1,SAMPLE2 variants.vcf.gz
  • VCF stores SNPs, indels and structural variants, one data line per variant.
  • INFO carries variant-level annotations. FORMAT and the sample columns carry per-sample genotypes.
  • GT encodes the genotype: 0/1 is heterozygous, 1/1 is homozygous alternate.
  • FILTER set to PASS is the only value you usually keep.
  • gVCF also records non-variant blocks and feeds joint calling workflows.
  • Compress with bgzip and index with tabix, and drive everything with bcftools.