FASTQ: Raw Reads
FASTQ is the starting point for almost all sequencing analysis. It stores raw reads straight from the sequencer, and every read occupies exactly four lines.
@SRR123456.1 1 length=150ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG+IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIILine 1 is the header. It starts with @ and carries the read identifier. The exact
format varies between sequencers, but the read name always comes first.
Line 2 is the nucleotide sequence of the read. The characters are A, T, C, G and N for ambiguous bases.
Line 3 is the plus line. It always starts with +. It sometimes repeats the header,
but usually carries the + character alone.
Line 4 holds the quality scores, one ASCII character per base. Each character encodes the confidence that the base call is correct, as a number called a Phred quality score.
Phred quality scores
Section titled “Phred quality scores”The Phred score Q relates to the probability of error P by a logarithm:
Q = -10 * log10(P)| Phred score | Error probability | Accuracy | ASCII character (Phred+33) |
|---|---|---|---|
| 10 | 1 in 10 | 90% | + |
| 20 | 1 in 100 | 99% | 5 |
| 30 | 1 in 1,000 | 99.9% | ? |
| 40 | 1 in 10,000 | 99.99% | I |
Q30 is a common quality threshold. A base with Q30 has a 99.9% chance of being correct, and modern Illumina sequencers produce the majority of bases at Q30 or above.
The ASCII encoding adds 33 to the Phred score. A score of 40 becomes ASCII
character 73, which is I. A score of 0 becomes ASCII 33, which is !. This
convention is called Phred+33 and every modern sequencer uses it.
Paired-end FASTQ files
Section titled “Paired-end FASTQ files”Most Illumina sequencing is paired-end. Each DNA fragment is sequenced from both ends, which gives two reads per fragment: read 1 (R1) and read 2 (R2). The two reads arrive in two separate files.
sample_R1.fastq.gz # forward readssample_R2.fastq.gz # reverse readsThe read order is identical in both files. The first read in R1 pairs with the first read in R2, and so on, which means both files must hold exactly the same number of reads in the same order. Two reads from one fragment improve alignment accuracy, span repetitive regions and give an estimate of the fragment size.
Compressed FASTQ
Section titled “Compressed FASTQ”Raw FASTQ files are enormous. A single lane of a NovaSeq S4 produces around 800
million read pairs, and uncompressed that can exceed 500 GB of sequence and quality
characters. Always store FASTQ files compressed with gzip, under the .fastq.gz
extension. Compression reduces the size by a factor of three to five.
# Compressed files are the default from most sequencing facilitiesls sample_R1.fastq.gz sample_R2.fastq.gz
# View a compressed FASTQ without decompressing itzcat sample_R1.fastq.gz | head -8
# Count the reads in a compressed FASTQzcat sample_R1.fastq.gz | wc -l# Divide the line count by 4Most tools accept .fastq.gz directly, so you almost never decompress on disk. If
a tool refuses compressed input, pipe through zcat:
zcat sample_R1.fastq.gz | some_tool -Common FASTQ issues
Section titled “Common FASTQ issues”Truncated files
Section titled “Truncated files”A transfer that fails midway leaves a truncated FASTQ. The file ends in the middle of a read, which breaks the four-line structure and causes cryptic errors in downstream tools. Check integrity after every transfer.
# Check the gzip stream is completegzip -t sample_R1.fastq.gz
# Verify the read counts match between R1 and R2echo $(( $(zcat sample_R1.fastq.gz | wc -l) / 4 ))echo $(( $(zcat sample_R2.fastq.gz | wc -l) / 4 ))If the counts differ, one file is corrupted. Download it again from the sequencing facility.
Mismatched R1 and R2
Section titled “Mismatched R1 and R2”Some quality filtering tools remove reads from one file but not the other, which
breaks the pairing. STAR and BWA expect R1 and R2 to hold matching reads in the
same order. Use tools that handle pairs together: fastp and trimmomatic in
paired-end mode filter both files at once and keep them synchronised.
Wrong quality encoding
Section titled “Wrong quality encoding”Older Illumina data used Phred+64 encoding instead of Phred+33. If the quality
lines show characters like h, i and j but nothing below @, the file may be
Phred+64. Almost all modern data is Phred+33, and FastQC detects the encoding and
flags the issue.
Summary
Section titled “Summary”- FASTQ stores raw reads in four lines: header, sequence, plus line and quality scores.
- Phred scores encode base call confidence, and Q30 means 99.9% accuracy.
- Paired-end reads arrive in two files with reads in matching order.
- Always store FASTQ compressed, and always verify integrity after a transfer.
- The next page covers what an aligner writes: SAM, BAM and CRAM.