BED, GFF & GTF
BED, GFF and GTF describe regions and annotations on a genome. BED defines intervals such as peaks or capture regions. GFF and GTF define gene models with exons, introns and coding sequences. You meet all three regularly.
BED format
Section titled “BED format”BED is a tab-separated format for genomic intervals. The minimum file has three columns: chromosome, start and end.
chr1 1000 2000chr1 3000 3500chr2 5000 6000BED uses zero-based, half-open coordinates. The start is 0-based and the end is
excluded, so chr1 1000 2000 covers bases 1001 to 2000 in 1-based counting. This
is the single most common source of off-by-one errors in bioinformatics.
BED3, BED6 and BED12
Section titled “BED3, BED6 and BED12”BED supports extra columns beyond the required three.
| Variant | Columns | Extra fields |
|---|---|---|
| BED3 | 3 | chrom, start, end only |
| BED6 | 6 | adds name, score, strand |
| BED12 | 12 | adds thickStart, thickEnd, itemRgb, blockCount, blockSizes, blockStarts |
A BED6 file looks like this:
chr1 1000 2000 peak_1 500 +chr1 3000 3500 peak_2 300 -chr2 5000 6000 peak_3 800 +The name labels the region, the score carries a numeric value such as a peak
signal, and the strand is + or -. BED12 adds a block structure that represents
exons within a transcript through the blockCount, blockSizes and blockStarts
fields.
You meet BED files everywhere: MACS2 writes ChIP-seq and ATAC-seq peaks as BED, exome capture kits ship a target BED that variant callers use to restrict the search space, the ENCODE blacklist is a BED of artefact-prone regions to exclude, and any custom set of regions of interest travels as BED.
GFF3 format
Section titled “GFF3 format”GFF3 describes gene structures and other genomic features in nine tab-separated columns.
##gff-version 3chr1 ensembl gene 1000 5000 . + . ID=gene1;Name=BRCA1chr1 ensembl mRNA 1000 5000 . + . ID=tx1;Parent=gene1chr1 ensembl exon 1000 1200 . + . Parent=tx1chr1 ensembl CDS 1050 1200 . + 0 Parent=tx1chr1 ensembl exon 3000 3300 . + . Parent=tx1chr1 ensembl CDS 3000 3300 . + 2 Parent=tx1| Column | Name | Description |
|---|---|---|
| 1 | seqid | Chromosome or scaffold name |
| 2 | source | Program or database that produced the feature |
| 3 | type | Feature type: gene, mRNA, exon, CDS |
| 4 | start | 1-based start position |
| 5 | end | 1-based end position, inclusive |
| 6 | score | Numeric score or . |
| 7 | strand | + or - |
| 8 | phase | For CDS: 0, 1 or 2, the reading frame offset |
| 9 | attributes | Semicolon-separated key=value pairs |
GFF3 is hierarchical. A gene contains mRNA transcripts, and each transcript
contains exons and CDS features. The Parent attribute links a child feature to
its parent. GFF3 uses 1-based, fully closed coordinates: both the start and the
end base belong to the feature.
GTF format
Section titled “GTF format”GTF is a close relative of GFF with a different attribute syntax, and it is the standard input for RNA-seq tools.
chr1 ensembl gene 1000 5000 . + . gene_id "ENSG00000012048"; gene_name "BRCA1";chr1 ensembl transcript 1000 5000 . + . gene_id "ENSG00000012048"; transcript_id "ENST00000357654";chr1 ensembl exon 1000 1200 . + . gene_id "ENSG00000012048"; transcript_id "ENST00000357654"; exon_number "1";The difference sits in column 9. GTF uses key "value"; pairs, GFF3 uses
key=value. GTF requires the gene_id and transcript_id attributes. GTF also
uses 1-based, fully closed coordinates.
GFF or GTF
Section titled “GFF or GTF”Most RNA-seq alignment and quantification tools expect GTF. STAR, HISAT2, featureCounts and Cell Ranger all take GTF annotations, and GENCODE ships GTF by default. GFF3 is more common at NCBI and in non-model organism databases, and Ensembl provides both. When a tool asks for a gene annotation file, check its documentation, but in most cases it wants GTF.
Where to download annotations
Section titled “Where to download annotations”| Source | Site | Formats | Notes |
|---|---|---|---|
| GENCODE | gencodegenes.org | GTF, GFF3 | Human and mouse. The standard for RNA-seq. |
| Ensembl | ensembl.org | GTF, GFF3 | Many species, Ensembl gene IDs. |
| UCSC | genome.ucsc.edu | Various | Table Browser for custom downloads, chr prefix. |
| NCBI | ncbi.nlm.nih.gov | GFF3 | RefSeq annotations, broad species coverage. |
bedtools basics
Section titled “bedtools basics”bedtools is the standard toolkit for interval operations, and it reads BED, GFF, GTF, VCF and BAM.
Intersect, merge, closest
Section titled “Intersect, merge, closest”# Find peaks that overlap promoter regionsbedtools intersect -a peaks.bed -b promoters.bed > overlapping_peaks.bed
# Report the peak and the overlapping promoter togetherbedtools intersect -a peaks.bed -b promoters.bed -wa -wb > peaks_with_promoters.bed
# Keep peaks that do NOT overlap blacklist regionsbedtools intersect -a peaks.bed -b blacklist.bed -v > filtered_peaks.bed
# Merge overlapping peaks (merge needs sorted input)sort -k1,1 -k2,2n peaks.bed > peaks_sorted.bedbedtools merge -i peaks_sorted.bed > merged_peaks.bed
# Merge peaks within 500 bp of each otherbedtools merge -i peaks_sorted.bed -d 500 > merged_peaks_500bp.bed
# The nearest gene for each peakbedtools closest -a peaks.bed -b genes.bed > peaks_nearest_gene.bedThe -a flag takes the query file and -b the reference file. The -v flag
inverts the match and reports entries with no overlap.
Other frequent operations
Section titled “Other frequent operations”# Remove blacklisted areas from peaksbedtools subtract -a peaks.bed -b blacklist.bed > clean_peaks.bed
# 1000 bp upstream of each genebedtools flank -i genes.bed -g genome.sizes -l 1000 -r 0 > upstream_1kb.bedSummary
Section titled “Summary”- BED defines intervals with 0-based, half-open coordinates, in 3 to 12 columns.
- GFF3 describes hierarchical gene models with 1-based coordinates and is common at NCBI and Ensembl.
- GTF is the annotation format RNA-seq tools expect. Download it from GENCODE for human and mouse work.
- Coordinate systems differ between formats. Always check which one you hold.
- bedtools covers interval work:
intersect,mergeandclosesthandle most of it. - Chromosome names must match between annotation, reference genome and BAM files.