Guides
My aim is to support scientists who want to make sense of their own data.
This section is free training material to get you started with the tools that matter most for everyday bioinformatics work.
Available now
Section titled “Available now”- Linux command line. Installing Ubuntu, navigating the filesystem, files, pipes, text processing, shell scripts, SSH.
- Git and version control. Why version control matters, basics, branching and merging, GitHub for scientists.
- Containers with Podman. What containers are, building images, running containers, volumes and networking, Podman compose.
- Programming with R. Variables, data structures, control flow, functions, debugging, testing, packages, tidyverse, S4 classes.
- Programming with Python. Variables, data structures, control flow, functions, debugging, testing, packages, pandas, polars, NumPy, classes.
- LLM assisted coding. Why coding agents matter, choosing one, configuring Claude Code, Codex, OpenCode, and Gemini, common pitfalls.
- Reproducibility. Why it matters, Nextflow and nf-core, AWS Batch, real RNA seq runs, Snakemake.
- RNA-seq analysis. The workflow from count matrix to enriched pathways, grouped into background, processing, differential expression, and pathway analysis.
- Batch correction. What a batch effect is, how to see one with PCA, and the design formula versus ComBat on the pasilla dataset.
- Single-cell RNA-seq. The Scanpy workflow on PBMC datasets, from QC and normalization through clustering and annotation to pseudobulk differential expression and integration, following sc-best-practices.
Questions or feedback?
Section titled “Questions or feedback?”See the About page.