Microarray Files
Before RNA-seq, expression was measured on microarrays, and the archives still hold hundreds of thousands of them. A scanner reads the fluorescence of each probe on the chip and writes one file per array. Two vendor families dominate: Affymetrix wrote CEL files, and Illumina wrote IDAT files. On top of those sits the GEO series matrix, the text format in which processed array data are deposited.
Affymetrix CEL
Section titled “Affymetrix CEL”A CEL file holds the probe intensities of one array: the grid position of each cell, its intensity, and its standard deviation. The format exists in two encodings. Version 3 is ASCII text with bracketed sections, and version 4 is a binary container. The affydata package ships one of each, from the same HG-U133A chip type.
library(affy)
cel_dir <- system.file("celfiles", package = "affydata")list.files(cel_dir)
# The text variant opens with a bracketed section header.text_lines <- readLines(file.path(cel_dir, "text.cel"), n = 8)cat(text_lines, sep = "\n")
# The binary variant opens with a magic number and a version.con <- file(file.path(cel_dir, "binary.cel"), "rb")magic <- readBin(con, "integer", n = 2)close(con)cat("magic:", magic[1], " version:", magic[2], "\n")[1] "binary.cel" "text.cel"[CEL]Version=3
[HEADER]Cols=712Rows=712TotalX=712TotalY=712magic: 64 version: 4Both encodings read through the same call, and both describe the same probe grid. The two example files are different scans, so their intensities differ, but their dimensions and chip definition match.
binary_cel <- ReadAffy(filenames = file.path(cel_dir, "binary.cel"))text_cel <- ReadAffy(filenames = file.path(cel_dir, "text.cel"))
# The probe grid: 712 rows by 712 columns on this chip.dim(binary_cel)cdfName(binary_cel)
# Perfect match and mismatch probe intensities.length(pm(binary_cel))head(pm(binary_cel), 3)head(mm(binary_cel), 3)Rows Cols 712 712[1] "HG-U133A"[1] 247965 binary.cel129340 1582213420 1502396671 3111 binary.cel130052 519214132 1902397383 4229An Affymetrix probeset measures a transcript with 11 to 20 probe pairs. Each pair is a perfect match probe, whose sequence matches the target exactly, and a mismatch probe, whose central base is flipped to estimate cross-hybridisation. The CEL file stores the raw intensities only. The mapping from probe to gene lives in a separate CDF, the chip definition file, and affy resolves it from the chip name.
From probes to an expression matrix
Section titled “From probes to an expression matrix”CEL files are the input. The expression matrix comes from preprocessing:
background correction, normalisation across arrays, and summarisation of each
probeset to one value. rma does all three. The affydata package bundles the
classic Dilution dataset, four arrays from a dilution experiment, as an
AffyBatch already read from CEL files.
data(Dilution, package = "affydata")pData(Dilution)
eset <- rma(Dilution)
# One row per probeset, one column per array, log2 scale.dim(exprs(eset))head(exprs(eset)[, 1:2], 4) liver sn19 scanner20A 20 0 120B 20 0 210A 10 0 110B 10 0 2Background correctingNormalizingCalculating Expression[1] 12625 4 20A 20B100_g_at 8.039969 7.8915221000_at 7.790839 7.6612091001_at 5.175568 5.0670231002_f_at 5.916532 5.754500RMA values are on the log2 scale, so a one-unit difference is a two-fold change. This matrix of features by samples is the same shape of object an RNA-seq count matrix becomes after normalisation, and everything downstream, PCA, clustering, differential expression with limma, works on it.
Illumina IDAT
Section titled “Illumina IDAT”Illumina BeadArrays write IDAT files, a binary format with one file per sample
per channel. A BeadChip holds several samples, so a run produces a folder of
IDATs rather than one file per chip. The R readers are illuminaio for the raw
format, and minfi or beadarray for methylation and expression analysis on
top of it. IDAT is binary and proprietary in structure, so treat it as read-only
and keep the originals.
The GEO series matrix
Section titled “The GEO series matrix”Processed array data are deposited at GEO, the Gene Expression Omnibus, and the
format you download is the series matrix: a text file of metadata lines starting
with !, then a tab-delimited table of expression values between two marker
lines. GEOquery parses it into an ExpressionSet, the Bioconductor container that
pairs the matrix with its sample metadata.
library(GEOquery)
gse <- getGEO(filename = "/opt/data/gse_series_matrix.txt", getGPL = FALSE)class(gse)dim(gse)head(exprs(gse), 3)pData(gse)[, c("title", "treatment:ch1")][1] "ExpressionSet"attr(,"package")[1] "Biobase"Features Samples 8 6 GSM0000001 GSM0000002 GSM0000003 GSM0000004 GSM0000005 GSM0000006PROBE_01 5.97 6.27 5.42 5.91 5.57 5.65PROBE_02 7.01 7.22 7.19 9.45 9.17 8.80PROBE_03 5.83 5.61 6.21 6.13 5.43 5.78 title treatment:ch1GSM0000001 control_1 controlGSM0000002 control_2 controlGSM0000003 control_3 controlGSM0000004 treated_1 treatedGSM0000005 treated_2 treatedGSM0000006 treated_3 treatedThe !Sample_characteristics lines carry the experimental design, here a
treatment column parsed into the phenotype table. On real deposits the
characteristics are free text written by the submitter, so expect inconsistent
spelling and missing values. The example file is a synthetic fixture; a real
series matrix for a twenty-sample study runs to tens of megabytes and downloads
with getGEO("GSExxxxx").
Summary
Section titled “Summary”- Affymetrix CEL files hold probe intensities for one array, in a text (v3) or
binary (v4) encoding, and affy reads both with
ReadAffy. - The probe-to-gene mapping lives in the CDF, a separate chip definition file.
rmaturns a set of CEL files into a log2 expression matrix of probesets by arrays.- Illumina writes one binary IDAT per sample per channel; read it with illuminaio, minfi or beadarray.
- GEO series matrix files are the deposit format:
!metadata lines plus a tab-delimited expression table, parsed by GEOquery.