Skip to content

Publication-Ready Figures

Most analysis ends in a figure, and a small set of figures does most of the work in a biology methods section. A box plot with a p-value. A survival curve. A clustered heatmap. This section builds each one twice, in R and in Python, from the same tidy data, with the right statistical test wired into the plot rather than bolted on after.

The approach is the one a bench scientist recognizes from GraphPad Prism: pick the test the data supports, draw the comparison, and put the result on the figure where a reviewer will see it. The difference is that here the figure, its code, and its numbers are all reproducible. Every figure on these pages is produced by a real run in a pinned container, the same one for both languages.

Figure What it compares Test on the plot
Two-group comparison one outcome, two groups t-test or Mann-Whitney, chosen from the data
Multi-group comparison one outcome, three or more groups one-way ANOVA + Tukey HSD
Factorial ANOVA one outcome, two crossed factors Type II ANOVA with interaction
Kaplan-Meier survival time to event, two arms log-rank
Cox forest plot adjusted hazard per covariate Cox proportional hazards
Clustered heatmap expression, features by samples hierarchical clustering
UpSet plot set membership across five assays intersection counts
Raincloud plot one outcome, distribution shape kernel density, mode count
FactoMineR PCA figures many variables, samples in a plane variance explained, contributions
ggstatsplot two groups, and paired data test, effect size and Bayes factor on the plot

Most pages show the same figure in both languages, in synced tabs, so you can read whichever you know and see its twin.

  • R uses ggpubr and rstatix for the comparison plots, survminer for survival, ComplexHeatmap for the heatmap, ComplexUpset and ggrain for the set and distribution figures, and FactoMineR with factoextra for the multivariate ones. All build on ggplot2.
  • Python uses seaborn with statannotations and pingouin for the comparison plots, lifelines for survival, seaborn.clustermap for the heatmap, and upsetplot and ptitprince for the set and distribution figures.

The first six pages run in the pinned pubplot container. The UpSet, raincloud and FactoMineR pages need packages no pubplot image carries, so they run in a second pinned image, figextra, built from a Containerfile in the code repo. The ggstatsplot page runs in a third, ggstats, for the ggstatsplot dependency tree. Which container a page used is stated on the page.

Two pages are R only. FactoMineR and factoextra have no mainstream Python counterpart, and neither does ggstatsplot; a hand-rolled imitation would not be one.

Where both languages compute the same statistic, they agree: the two-group t-test matches to nine significant figures, the ANOVA and the log-rank agree to the digit. The figures themselves are drawn each library’s idiomatic way, so they look alike without being pixel-identical. The runnable code and fixtures for every page live in the companion code repo under guides/figures/.