Skip to content

FASTQ: Raw Reads

FASTQ is the starting point for almost all sequencing analysis. It stores raw reads straight from the sequencer, and every read occupies exactly four lines.

@SRR123456.1 1 length=150
ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG
+
IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIII

Line 1 is the header. It starts with @ and carries the read identifier. The exact format varies between sequencers, but the read name always comes first.

Line 2 is the nucleotide sequence of the read. The characters are A, T, C, G and N for ambiguous bases.

Line 3 is the plus line. It always starts with +. It sometimes repeats the header, but usually carries the + character alone.

Line 4 holds the quality scores, one ASCII character per base. Each character encodes the confidence that the base call is correct, as a number called a Phred quality score.

The Phred score Q relates to the probability of error P by a logarithm:

Q = -10 * log10(P)
Phred score Error probability Accuracy ASCII character (Phred+33)
10 1 in 10 90% +
20 1 in 100 99% 5
30 1 in 1,000 99.9% ?
40 1 in 10,000 99.99% I

Q30 is a common quality threshold. A base with Q30 has a 99.9% chance of being correct, and modern Illumina sequencers produce the majority of bases at Q30 or above.

The ASCII encoding adds 33 to the Phred score. A score of 40 becomes ASCII character 73, which is I. A score of 0 becomes ASCII 33, which is !. This convention is called Phred+33 and every modern sequencer uses it.

Most Illumina sequencing is paired-end. Each DNA fragment is sequenced from both ends, which gives two reads per fragment: read 1 (R1) and read 2 (R2). The two reads arrive in two separate files.

sample_R1.fastq.gz # forward reads
sample_R2.fastq.gz # reverse reads

The read order is identical in both files. The first read in R1 pairs with the first read in R2, and so on, which means both files must hold exactly the same number of reads in the same order. Two reads from one fragment improve alignment accuracy, span repetitive regions and give an estimate of the fragment size.

Raw FASTQ files are enormous. A single lane of a NovaSeq S4 produces around 800 million read pairs, and uncompressed that can exceed 500 GB of sequence and quality characters. Always store FASTQ files compressed with gzip, under the .fastq.gz extension. Compression reduces the size by a factor of three to five.

Terminal window
# Compressed files are the default from most sequencing facilities
ls sample_R1.fastq.gz sample_R2.fastq.gz
# View a compressed FASTQ without decompressing it
zcat sample_R1.fastq.gz | head -8
# Count the reads in a compressed FASTQ
zcat sample_R1.fastq.gz | wc -l
# Divide the line count by 4

Most tools accept .fastq.gz directly, so you almost never decompress on disk. If a tool refuses compressed input, pipe through zcat:

Terminal window
zcat sample_R1.fastq.gz | some_tool -

A transfer that fails midway leaves a truncated FASTQ. The file ends in the middle of a read, which breaks the four-line structure and causes cryptic errors in downstream tools. Check integrity after every transfer.

Terminal window
# Check the gzip stream is complete
gzip -t sample_R1.fastq.gz
# Verify the read counts match between R1 and R2
echo $(( $(zcat sample_R1.fastq.gz | wc -l) / 4 ))
echo $(( $(zcat sample_R2.fastq.gz | wc -l) / 4 ))

If the counts differ, one file is corrupted. Download it again from the sequencing facility.

Some quality filtering tools remove reads from one file but not the other, which breaks the pairing. STAR and BWA expect R1 and R2 to hold matching reads in the same order. Use tools that handle pairs together: fastp and trimmomatic in paired-end mode filter both files at once and keep them synchronised.

Older Illumina data used Phred+64 encoding instead of Phred+33. If the quality lines show characters like h, i and j but nothing below @, the file may be Phred+64. Almost all modern data is Phred+33, and FastQC detects the encoding and flags the issue.

  • FASTQ stores raw reads in four lines: header, sequence, plus line and quality scores.
  • Phred scores encode base call confidence, and Q30 means 99.9% accuracy.
  • Paired-end reads arrive in two files with reads in matching order.
  • Always store FASTQ compressed, and always verify integrity after a transfer.
  • The next page covers what an aligner writes: SAM, BAM and CRAM.