Single-Cell Formats: H5AD, MTX & Loom
Single-cell genomics produces millions of measurements across thousands of cells, and specialised formats store that data efficiently. This page covers the formats you meet in single-cell RNA-seq work. For runnable analysis code on these objects, see the single-cell section, which works through AnnData step by step.
The count matrix
Section titled “The count matrix”The central data structure is the count matrix. Rows are genes or features, columns are cells identified by their barcodes, and each entry is a UMI count: how many unique transcripts of that gene were detected in that cell.
Most entries are zero. A typical cell expresses 2,000 to 5,000 genes out of 20,000 or more, and that sparsity is why the formats use sparse storage rather than dense tables.
10x Genomics output
Section titled “10x Genomics output”The Cell Ranger pipeline produces two main outputs.
The filtered feature barcode matrix directory
Section titled “The filtered feature barcode matrix directory”The filtered_feature_bc_matrix/ directory holds three files.
| File | Contents |
|---|---|
matrix.mtx.gz |
Sparse count matrix in Matrix Market format |
barcodes.tsv.gz |
One cell barcode per line |
features.tsv.gz |
Gene IDs, gene names and feature types |
Matrix Market stores only the non-zero entries. After a two-line header, each line carries a row index, a column index and a value, which is efficient for sparse data. You can read the dimensions without loading anything.
# The third line gives genes, cells and non-zero entrieszcat matrix.mtx.gz | head -3
# Barcode count is the cell count, feature count is the gene countzcat barcodes.tsv.gz | wc -lzcat features.tsv.gz | wc -lThe H5 file
Section titled “The H5 file”10x also ships filtered_feature_bc_matrix.h5, an HDF5 file with the same data as
the three-file directory. One file is easier to transfer and archive than three,
and both Scanpy and Seurat read it directly.
H5AD format
Section titled “H5AD format”H5AD is the file format of AnnData, the data structure at the centre of Scanpy and the Python single-cell ecosystem. One H5AD file holds the whole analysis state.
| Slot | Contents | Example |
|---|---|---|
adata.X |
The main data matrix | Raw or normalised counts |
adata.obs |
Cell metadata, one row per cell | Cell type labels, sample IDs, QC metrics |
adata.var |
Gene metadata, one row per gene | Gene names, highly variable flags |
adata.obsm |
Cell embeddings, matrices keyed by name | PCA and UMAP coordinates |
adata.varm |
Gene embeddings | PCA loadings |
adata.obsp |
Cell-cell pairwise data | Neighbour graphs, distance matrices |
adata.uns |
Unstructured dictionaries | Colour palettes, analysis parameters |
adata.layers |
Alternative matrices, same shape as X | Raw counts beside log-normalised values |
H5AD bundles everything into one file, so you can save after each step and reload later without losing results. For very large datasets, backed mode reads from disk on demand instead of loading the matrix into memory.
# A one-line summary without opening an analysis sessionpython -c "import anndata; adata = anndata.read_h5ad('file.h5ad'); print(adata)"AnnData object with n_obs × n_vars = 8000 × 20000 obs: 'cell_type', 'n_genes', 'percent_mito' var: 'gene_ids', 'highly_variable' obsm: 'X_pca', 'X_umap' layers: 'raw_counts'The printout gives the cell and gene counts and lists which metadata fields the file carries.
Loom format
Section titled “Loom format”Loom is an older HDF5-based format built for large single-cell datasets. A Loom file holds a main expression matrix, row attributes for genes, column attributes for cells, and graph objects for nearest-neighbour graphs. Most of the Python ecosystem has standardised on H5AD, and you meet Loom mainly when you run RNA velocity with Velocyto or when you work with older datasets.
Seurat RDS format
Section titled “Seurat RDS format”In R, Seurat is the dominant analysis package, and Seurat objects are saved with
R’s native serialisation as .rds files. RDS is not interoperable with Python
tools, so moving an analysis between R and Python means converting formats.
Converting between formats
Section titled “Converting between formats”| Conversion | Tool | Language |
|---|---|---|
| Seurat to H5AD | SeuratDisk | R |
| H5AD to SingleCellExperiment | zellkonverter | R |
| Seurat to AnnData | sceasy | R |
| AnnData to Seurat | SeuratDisk | R |
File sizes
Section titled “File sizes”| Dataset | Cells | Genes | H5AD size |
|---|---|---|---|
| Small experiment | 5,000 | 20,000 | about 30 MB |
| Standard experiment | 10,000 | 20,000 | about 50 MB |
| Large experiment | 100,000 | 20,000 | about 500 MB |
| Cell atlas | 1,000,000 | 20,000 | 5-15 GB |
These sizes assume sparse storage. The same data as a dense CSV would be 10 to 100 times larger, which is why the specialised formats exist.
Choosing a format
Section titled “Choosing a format”- Working in Python with Scanpy: use H5AD. One file holds the matrix, the metadata, the embeddings and the results.
- Working in R with Seurat: use RDS. It preserves every Seurat slot.
- Receiving 10x Genomics data: start with either the MTX directory or the H5 file. Both carry the same counts.
- Running RNA velocity: Velocyto writes Loom, and you convert to H5AD for the rest.
- Sharing between R and Python: convert to H5AD as the common exchange format.
Summary
Section titled “Summary”- The count matrix is the core structure, and it is sparse because most genes are not detected in most cells.
- 10x Genomics outputs a three-file MTX directory or a single H5 file with the same data.
- H5AD is the standard in the Python ecosystem and stores the full analysis state in one file.
- Loom appears mainly in RNA velocity work.
- Seurat RDS files are R-specific, and SeuratDisk or zellkonverter convert between the two ecosystems. Verify every conversion.