VCF: Variants
VCF is the standard format for storing genetic variants. It records SNPs, insertions, deletions and structural variants. Every variant calling pipeline produces VCF output, and every downstream annotation or filtering tool reads it.
A variant is a position where a sample differs from the reference sequence. The three common types are SNPs, where one base differs; insertions, where the sample carries extra bases; and deletions, where the sample lacks bases the reference has. Callers such as GATK HaplotypeCaller, DeepVariant and FreeBayes compare aligned reads with the reference and write the differences into a VCF file.
VCF structure
Section titled “VCF structure”A VCF file has three sections: meta-information lines, one header line, and data
lines. The meta-information lines start with ## and define the file version, the
reference, and the meaning of the fields used in the data lines.
##fileformat=VCFv4.2##FILTER=<ID=PASS,Description="All filters passed">##INFO=<ID=DP,Number=1,Type=Integer,Description="Total Depth">##INFO=<ID=AF,Number=A,Type=Float,Description="Allele Frequency">##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">##FORMAT=<ID=AD,Number=R,Type=Integer,Description="Allelic depths">##FORMAT=<ID=GQ,Number=1,Type=Integer,Description="Genotype Quality">The header line starts with #CHROM and names the columns. Each data line is one
variant.
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1chr1 12345 rs12345 A G 99 PASS DP=50;AF=0.5 GT:AD:GQ 0/1:25,25:99The fixed columns
Section titled “The fixed columns”| Column | Name | Description |
|---|---|---|
| 1 | CHROM | Chromosome name |
| 2 | POS | 1-based position on the chromosome |
| 3 | ID | Variant identifier from dbSNP or another database |
| 4 | REF | Reference allele |
| 5 | ALT | Alternate allele observed in the sample |
| 6 | QUAL | Phred-scaled quality of the variant call |
| 7 | FILTER | PASS when the variant passed all filters |
| 8 | INFO | Semicolon-separated key-value annotations for the variant |
| 9 | FORMAT | Colon-separated list of the per-sample fields |
| 10+ | SAMPLE | Per-sample genotype and quality values |
VCF uses 1-based coordinates, so position 12345 is the 12,345th base on the chromosome. BED uses 0-based coordinates, and the difference matters.
The INFO field
Section titled “The INFO field”INFO packs variant-level annotations into semicolon-separated key-value pairs. The common fields:
| Key | Meaning | Example |
|---|---|---|
| DP | Total read depth at the position | DP=50 |
| AF | Allele frequency of the alternate allele | AF=0.5 |
| MQ | Root mean square mapping quality of the reads | MQ=60 |
| QD | Quality by depth (QUAL divided by DP) | QD=15.2 |
| FS | Fisher strand bias p-value | FS=0.0 |
| SOR | Strand odds ratio | SOR=0.7 |
Different callers add different fields. The ##INFO lines at the top of the file
define each one.
The genotype fields
Section titled “The genotype fields”The FORMAT column declares which per-sample fields follow, and the sample columns carry the values.
GT is the called genotype, written as allele indices separated by / or |. The
index 0 is the REF allele, 1 is the first ALT allele, and so on.
| Value | Meaning |
|---|---|
| 0/0 | Homozygous reference |
| 0/1 | Heterozygous |
| 1/1 | Homozygous alternate |
| ./. | No call (insufficient data) |
AD is the allelic depth, the read count supporting each allele. AD=25,25 means
25 reads support the reference and 25 support the alternate. GQ is the
Phred-scaled confidence in the genotype call, and GQ 30 means a 1 in 1000 chance
the genotype is wrong. DP is the depth for that one sample, which can differ from
the INFO DP summed across samples.
The FILTER column
Section titled “The FILTER column”FILTER records whether a variant passed quality control. PASS is the value you
want. Any other value names the filter that flagged the variant, such as LowQual
or StrandBias, and a dot means no filters were applied. GATK’s recommended
thresholds for SNPs include QD below 2.0, FS above 60.0 and MQ below 40.0, applied
with gatk VariantFiltration and recorded here.
Multi-sample VCF
Section titled “Multi-sample VCF”One VCF can carry genotypes for many samples, one column each after FORMAT.
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1 SAMPLE2 SAMPLE3chr1 12345 . A G 99 PASS DP=150;AF=0.33 GT:AD:GQ 0/1:25,25:99 0/0:50,0:99 1/1:0,50:99Here SAMPLE1 is heterozygous, SAMPLE2 is homozygous reference and SAMPLE3 is homozygous alternate. Multi-sample VCFs are the standard output of cohort studies.
A standard VCF records only the positions where a variant was found, so a missing position is ambiguous: the sample may be homozygous reference, or there may be no data at all. A gVCF resolves the ambiguity by also recording the non-variant positions, compressed into blocks with an END tag.
chr1 10001 . A <NON_REF> . . END=10100 GT:DP:GQ 0/0:30:90This line says positions 10001 to 10100 are all homozygous reference at about 30x
depth. gVCFs feed GATK’s joint calling workflow: individual gVCFs combine with
GenomicsDBImport or CombineGVCFs, and GenotypeGVCFs genotypes the cohort.
Compression and indexing
Section titled “Compression and indexing”Compress VCF files with bgzip and index them with tabix.
bgzip variants.vcftabix -p vcf variants.vcf.gzThis creates variants.vcf.gz and variants.vcf.gz.tbi. The index lets tools
jump to a region without reading the whole file, and most tools expect both.
bcftools basics
Section titled “bcftools basics”bcftools is the standard toolkit for VCF files.
# View the headerbcftools view -h variants.vcf.gz
# Summary statistics: SNP and indel counts, Ts/Tv ratio, per-sample statsbcftools stats variants.vcf.gz
# Extract specific fieldsbcftools query -f '%CHROM\t%POS\t%REF\t%ALT\t%QUAL\n' variants.vcf.gz | head
# Keep only PASS variants with QUAL of 30 or morebcftools view -f PASS -i 'QUAL>=30' variants.vcf.gz -o filtered.vcf.gz -Oz
# Extract one regionbcftools view -r chr1:10000-20000 variants.vcf.gz
# Extract two samplesbcftools view -s SAMPLE1,SAMPLE2 variants.vcf.gzSummary
Section titled “Summary”- VCF stores SNPs, indels and structural variants, one data line per variant.
- INFO carries variant-level annotations. FORMAT and the sample columns carry per-sample genotypes.
- GT encodes the genotype: 0/1 is heterozygous, 1/1 is homozygous alternate.
- FILTER set to PASS is the only value you usually keep.
- gVCF also records non-variant blocks and feeds joint calling workflows.
- Compress with bgzip and index with tabix, and drive everything with bcftools.