Skip to content

Single-Cell Formats: H5AD, MTX & Loom

Single-cell genomics produces millions of measurements across thousands of cells, and specialised formats store that data efficiently. This page covers the formats you meet in single-cell RNA-seq work. For runnable analysis code on these objects, see the single-cell section, which works through AnnData step by step.

The central data structure is the count matrix. Rows are genes or features, columns are cells identified by their barcodes, and each entry is a UMI count: how many unique transcripts of that gene were detected in that cell.

Most entries are zero. A typical cell expresses 2,000 to 5,000 genes out of 20,000 or more, and that sparsity is why the formats use sparse storage rather than dense tables.

The Cell Ranger pipeline produces two main outputs.

The filtered feature barcode matrix directory

Section titled “The filtered feature barcode matrix directory”

The filtered_feature_bc_matrix/ directory holds three files.

File Contents
matrix.mtx.gz Sparse count matrix in Matrix Market format
barcodes.tsv.gz One cell barcode per line
features.tsv.gz Gene IDs, gene names and feature types

Matrix Market stores only the non-zero entries. After a two-line header, each line carries a row index, a column index and a value, which is efficient for sparse data. You can read the dimensions without loading anything.

Terminal window
# The third line gives genes, cells and non-zero entries
zcat matrix.mtx.gz | head -3
# Barcode count is the cell count, feature count is the gene count
zcat barcodes.tsv.gz | wc -l
zcat features.tsv.gz | wc -l

10x also ships filtered_feature_bc_matrix.h5, an HDF5 file with the same data as the three-file directory. One file is easier to transfer and archive than three, and both Scanpy and Seurat read it directly.

H5AD is the file format of AnnData, the data structure at the centre of Scanpy and the Python single-cell ecosystem. One H5AD file holds the whole analysis state.

Slot Contents Example
adata.X The main data matrix Raw or normalised counts
adata.obs Cell metadata, one row per cell Cell type labels, sample IDs, QC metrics
adata.var Gene metadata, one row per gene Gene names, highly variable flags
adata.obsm Cell embeddings, matrices keyed by name PCA and UMAP coordinates
adata.varm Gene embeddings PCA loadings
adata.obsp Cell-cell pairwise data Neighbour graphs, distance matrices
adata.uns Unstructured dictionaries Colour palettes, analysis parameters
adata.layers Alternative matrices, same shape as X Raw counts beside log-normalised values

H5AD bundles everything into one file, so you can save after each step and reload later without losing results. For very large datasets, backed mode reads from disk on demand instead of loading the matrix into memory.

Terminal window
# A one-line summary without opening an analysis session
python -c "import anndata; adata = anndata.read_h5ad('file.h5ad'); print(adata)"
AnnData object with n_obs × n_vars = 8000 × 20000
obs: 'cell_type', 'n_genes', 'percent_mito'
var: 'gene_ids', 'highly_variable'
obsm: 'X_pca', 'X_umap'
layers: 'raw_counts'

The printout gives the cell and gene counts and lists which metadata fields the file carries.

Loom is an older HDF5-based format built for large single-cell datasets. A Loom file holds a main expression matrix, row attributes for genes, column attributes for cells, and graph objects for nearest-neighbour graphs. Most of the Python ecosystem has standardised on H5AD, and you meet Loom mainly when you run RNA velocity with Velocyto or when you work with older datasets.

In R, Seurat is the dominant analysis package, and Seurat objects are saved with R’s native serialisation as .rds files. RDS is not interoperable with Python tools, so moving an analysis between R and Python means converting formats.

Conversion Tool Language
Seurat to H5AD SeuratDisk R
H5AD to SingleCellExperiment zellkonverter R
Seurat to AnnData sceasy R
AnnData to Seurat SeuratDisk R
Dataset Cells Genes H5AD size
Small experiment 5,000 20,000 about 30 MB
Standard experiment 10,000 20,000 about 50 MB
Large experiment 100,000 20,000 about 500 MB
Cell atlas 1,000,000 20,000 5-15 GB

These sizes assume sparse storage. The same data as a dense CSV would be 10 to 100 times larger, which is why the specialised formats exist.

  • Working in Python with Scanpy: use H5AD. One file holds the matrix, the metadata, the embeddings and the results.
  • Working in R with Seurat: use RDS. It preserves every Seurat slot.
  • Receiving 10x Genomics data: start with either the MTX directory or the H5 file. Both carry the same counts.
  • Running RNA velocity: Velocyto writes Loom, and you convert to H5AD for the rest.
  • Sharing between R and Python: convert to H5AD as the common exchange format.
  • The count matrix is the core structure, and it is sparse because most genes are not detected in most cells.
  • 10x Genomics outputs a three-file MTX directory or a single H5 file with the same data.
  • H5AD is the standard in the Python ecosystem and stores the full analysis state in one file.
  • Loom appears mainly in RNA velocity work.
  • Seurat RDS files are R-specific, and SeuratDisk or zellkonverter convert between the two ecosystems. Verify every conversion.