Skip to content

Mass Spectrometry Files

A mass spectrometer records spectra. Each scan is a list of peaks, and a peak is a pair of numbers: a mass-to-charge ratio and an intensity. A run of an hour holds tens of thousands of scans, and the file formats split into two families. The vendor formats are what the instrument writes, and mzML is the open standard the community converts them into.

Every vendor writes its own container, and none of them are documented for outside parsing.

Vendor Format Shape
Thermo Fisher .raw Single binary file
Bruker .d Directory holding .baf or .tdf files
Agilent .d Directory of binaries
Waters .raw Directory of binaries
SCIEX .wiff Binary file plus a .scan companion

The conversion tool is msconvert from ProteoWizard. It reads every vendor format above and writes mzML, optionally with peak picking and filters applied. The conversion is where the analysis starts, and keeping the raw files matters: mzML with peak picking applied cannot be un-picked.

mzML itself is XML defined by the HUPO Proteomics Standards Initiative, with the peak lists stored as base64-encoded binary arrays inside the XML. The head of a real file shows the structure: controlled vocabulary declarations, then a description of the file and its source.

mzml_file <- system.file("microtofq", "MM14.mzML", package = "msdata")
file.size(mzml_file)
# mzML is XML. The first lines declare the vocabularies and the source file.
mzml_lines <- readLines(mzml_file, n = 15)
cat(substr(mzml_lines, 1, 100), sep = "\n")
[1] 3651990
<?xml version="1.0" encoding="ISO-8859-1"?>
<mzML xmlns="http://psi.hupo.org/ms/mzml" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:
<cvList count="2">
<cv id="MS" fullName="Proteomics Standards Initiative Mass Spectrometry Ontology" URI="http://psid
<cv id="UO" fullName="Unit Ontology" URI="http://obo.cvs.sourceforge.net/obo/obo/ontology/phenotyp
</cvList>
<fileDescription>
<fileContent>
<cvParam cvRef="MS" accession="MS:1000294" name="mass spectrum" />
</fileContent>
<sourceFileList count="1">
<sourceFile id="sf_ru_0" name="analysis.baf" location="MM14_20uM_2-A%2c4_01_1745.d/">
<cvParam cvRef="MS" accession="MS:1000569" name="SHA-1" value="" />
<cvParam cvRef="MS" accession="MS:1000564" name="PSI mzData file" />
<cvParam cvRef="MS" accession="MS:1000824" name="no nativeID format" />

This file records its own provenance. The source was analysis.baf inside a directory ending in .d, which is a Bruker raw file, converted to mzML before analysis. The cvParam elements are controlled vocabulary terms: rather than free text, every property is a numbered term from a shared ontology, so two programs that both say MS:1000294 both mean a mass spectrum.

The example is a MALDI-TOF run on a Bruker Microtof Q, shipped by the msdata package and used across the Bioconductor mass spectrometry documentation.

mzR is the Bioconductor reader for mzML and its predecessors. It gives the run header, one row per scan, and the peak matrix of any scan.

library(mzR)
ms <- openMSfile(mzml_file)
hd <- header(ms)
nrow(hd)
table(hd$msLevel)
range(hd$peaksCount)
head(hd[, c("seqNum", "msLevel", "retentionTime", "peaksCount",
"basePeakMZ")], 5)
[1] 112
1
112
[1] 1301 1695
seqNum msLevel retentionTime peaksCount basePeakMZ
1 1 1 270.334 1378 0
2 2 1 270.671 1356 0
3 3 1 271.007 1404 0
4 4 1 271.343 1496 0
5 5 1 271.680 1525 0

The run holds 112 scans, all at MS level 1, each with around fifteen hundred peaks. In a tandem run the header also carries MS2 rows, and the precursor columns say which MS1 ion each MS2 scan fragmented.

# The scan with the most peaks, as an mz by intensity matrix.
scan_i <- which.max(hd$peaksCount)
scan_i
peaks_scan <- peaks(ms, scan_i)
dim(peaks_scan)
head(peaks_scan, 4)
close(ms)
[1] 25
[1] 1695 2
mz intensity
[1,] 100.0274 34
[2,] 105.1012 30
[3,] 110.0472 39
[4,] 116.0829 21

Scan 25 is the densest of the run. The peak matrix is what every downstream tool consumes, and drawn as stems it is the picture every mass spectrometrist recognises.

library(ggplot2)
spectrum_plot <- ggplot(data.frame(peaks_scan),
aes(x = mz, xend = mz, y = 0, yend = intensity)) +
geom_segment(linewidth = 0.2) +
labs(x = "m/z", y = "Intensity",
title = paste("Scan", scan_i, "of MM14.mzML"))
dir.create("outputs", showWarnings = FALSE, recursive = TRUE)
ggsave("outputs/mass_spec_scan.png", plot = spectrum_plot,
width = 7, height = 4, dpi = 110, bg = "white")

Scan 25 of the MALDI-TOF run, the scan with the most peaks, drawn as stems. Peaks rise from a baseline of noise across the mass range, and the tallest signals cluster below m/z 250.

mzML covers the raw spectra. The steps after identification have their own standards, all from the same initiative.

Format What it stores Read by
mzXML Spectra, the older ISB schema mzR
mzIdentML Peptide and protein identifications mzR, mzID
mzTab Identification and quantitation results base R, it is tab-delimited
mgf Peak lists as plain text, for search engines MSnbase, search engines

The mgf is the one you can read with no tools at all, which is why database search engines took it as input.

BEGIN IONS
TITLE=scan 1234
PEPMASS=842.51
CHARGE=2+
RTINSECONDS=271.3
110.14 1023
111.09 441
...
END IONS
  • Vendor raw formats are proprietary. Convert to mzML with msconvert, and keep the originals.
  • mzML is XML with base64 peak arrays, defined by numbered controlled vocabulary terms, and it records the source file it was converted from.
  • mzR reads mzML: header for the per-scan table, peaks for the peak matrix of one scan.
  • Identifications travel as mzIdentML, quantitation as mzTab, and search engine peak lists as mgf.