Skip to content

Microarray Files

Before RNA-seq, expression was measured on microarrays, and the archives still hold hundreds of thousands of them. A scanner reads the fluorescence of each probe on the chip and writes one file per array. Two vendor families dominate: Affymetrix wrote CEL files, and Illumina wrote IDAT files. On top of those sits the GEO series matrix, the text format in which processed array data are deposited.

A CEL file holds the probe intensities of one array: the grid position of each cell, its intensity, and its standard deviation. The format exists in two encodings. Version 3 is ASCII text with bracketed sections, and version 4 is a binary container. The affydata package ships one of each, from the same HG-U133A chip type.

library(affy)
cel_dir <- system.file("celfiles", package = "affydata")
list.files(cel_dir)
# The text variant opens with a bracketed section header.
text_lines <- readLines(file.path(cel_dir, "text.cel"), n = 8)
cat(text_lines, sep = "\n")
# The binary variant opens with a magic number and a version.
con <- file(file.path(cel_dir, "binary.cel"), "rb")
magic <- readBin(con, "integer", n = 2)
close(con)
cat("magic:", magic[1], " version:", magic[2], "\n")
[1] "binary.cel" "text.cel"
[CEL]
Version=3
[HEADER]
Cols=712
Rows=712
TotalX=712
TotalY=712
magic: 64 version: 4

Both encodings read through the same call, and both describe the same probe grid. The two example files are different scans, so their intensities differ, but their dimensions and chip definition match.

binary_cel <- ReadAffy(filenames = file.path(cel_dir, "binary.cel"))
text_cel <- ReadAffy(filenames = file.path(cel_dir, "text.cel"))
# The probe grid: 712 rows by 712 columns on this chip.
dim(binary_cel)
cdfName(binary_cel)
# Perfect match and mismatch probe intensities.
length(pm(binary_cel))
head(pm(binary_cel), 3)
head(mm(binary_cel), 3)
Rows Cols
712 712
[1] "HG-U133A"
[1] 247965
binary.cel
129340 1582
213420 1502
396671 3111
binary.cel
130052 519
214132 1902
397383 4229

An Affymetrix probeset measures a transcript with 11 to 20 probe pairs. Each pair is a perfect match probe, whose sequence matches the target exactly, and a mismatch probe, whose central base is flipped to estimate cross-hybridisation. The CEL file stores the raw intensities only. The mapping from probe to gene lives in a separate CDF, the chip definition file, and affy resolves it from the chip name.

CEL files are the input. The expression matrix comes from preprocessing: background correction, normalisation across arrays, and summarisation of each probeset to one value. rma does all three. The affydata package bundles the classic Dilution dataset, four arrays from a dilution experiment, as an AffyBatch already read from CEL files.

data(Dilution, package = "affydata")
pData(Dilution)
eset <- rma(Dilution)
# One row per probeset, one column per array, log2 scale.
dim(exprs(eset))
head(exprs(eset)[, 1:2], 4)
liver sn19 scanner
20A 20 0 1
20B 20 0 2
10A 10 0 1
10B 10 0 2
Background correcting
Normalizing
Calculating Expression
[1] 12625 4
20A 20B
100_g_at 8.039969 7.891522
1000_at 7.790839 7.661209
1001_at 5.175568 5.067023
1002_f_at 5.916532 5.754500

RMA values are on the log2 scale, so a one-unit difference is a two-fold change. This matrix of features by samples is the same shape of object an RNA-seq count matrix becomes after normalisation, and everything downstream, PCA, clustering, differential expression with limma, works on it.

Illumina BeadArrays write IDAT files, a binary format with one file per sample per channel. A BeadChip holds several samples, so a run produces a folder of IDATs rather than one file per chip. The R readers are illuminaio for the raw format, and minfi or beadarray for methylation and expression analysis on top of it. IDAT is binary and proprietary in structure, so treat it as read-only and keep the originals.

Processed array data are deposited at GEO, the Gene Expression Omnibus, and the format you download is the series matrix: a text file of metadata lines starting with !, then a tab-delimited table of expression values between two marker lines. GEOquery parses it into an ExpressionSet, the Bioconductor container that pairs the matrix with its sample metadata.

library(GEOquery)
gse <- getGEO(filename = "/opt/data/gse_series_matrix.txt", getGPL = FALSE)
class(gse)
dim(gse)
head(exprs(gse), 3)
pData(gse)[, c("title", "treatment:ch1")]
[1] "ExpressionSet"
attr(,"package")
[1] "Biobase"
Features Samples
8 6
GSM0000001 GSM0000002 GSM0000003 GSM0000004 GSM0000005 GSM0000006
PROBE_01 5.97 6.27 5.42 5.91 5.57 5.65
PROBE_02 7.01 7.22 7.19 9.45 9.17 8.80
PROBE_03 5.83 5.61 6.21 6.13 5.43 5.78
title treatment:ch1
GSM0000001 control_1 control
GSM0000002 control_2 control
GSM0000003 control_3 control
GSM0000004 treated_1 treated
GSM0000005 treated_2 treated
GSM0000006 treated_3 treated

The !Sample_characteristics lines carry the experimental design, here a treatment column parsed into the phenotype table. On real deposits the characteristics are free text written by the submitter, so expect inconsistent spelling and missing values. The example file is a synthetic fixture; a real series matrix for a twenty-sample study runs to tens of megabytes and downloads with getGEO("GSExxxxx").

  • Affymetrix CEL files hold probe intensities for one array, in a text (v3) or binary (v4) encoding, and affy reads both with ReadAffy.
  • The probe-to-gene mapping lives in the CDF, a separate chip definition file.
  • rma turns a set of CEL files into a log2 expression matrix of probesets by arrays.
  • Illumina writes one binary IDAT per sample per channel; read it with illuminaio, minfi or beadarray.
  • GEO series matrix files are the deposit format: ! metadata lines plus a tab-delimited expression table, parsed by GEOquery.