Skip to content

ADaM Datasets with admiral

SDTM organises what a trial collected. ADaM organises what the trial concludes, and it does it in two dataset shapes. ADSL, the subject-level analysis dataset, holds one row per subject with everything an analysis needs to know about them: treatment, population flags, dates, stratification. Every other analysis dataset is BDS-shaped, basic data structure: long, one row per subject per parameter per timepoint, with the analysis value in AVAL, the baseline in BASE, and the change in CHG. A lab dataset, a vitals dataset and a questionnaire dataset all share that skeleton, which is why one set of derivation functions covers most of an ADaM submission.

admiral is the pharmaverse toolbox for those derivations. It was started by Roche and GSK and is now a wider collaboration, and it is deliberately a programming library rather than a one-button generator: you state each derivation as a function call, and the call records the decision. This page builds a small ADSL from the pilot study’s SDTM domains, then derives two lab parameters into BDS form, and at each stage compares the result with pharmaverseadam, the reference ADaM datasets the admiral team regenerates from its own template programs on every release.

library(admiral)
library(dplyr)
library(pharmaversesdtm)
# SDTM codes missing values as empty strings; turn them into real NAs.
dm <- convert_blanks_to_na(pharmaversesdtm::dm)
ex <- convert_blanks_to_na(pharmaversesdtm::ex)
dim(dm)
[1] 306 28

ADSL: treatment dates and population flags

Section titled “ADSL: treatment dates and population flags”

The planned and actual treatment columns come straight from DM. The work is in the dates: exposure starts and ends live in EX as character datetimes, possibly partial, so they are parsed to real datetimes first, then the earliest start and latest end per subject become TRTSDTM and TRTEDTM. The filter in the merge is the study’s definition of treated: a dose above zero, or a zero dose whose name says placebo, which is how infusion-placebo records look in this study.

adsl <- dm %>%
select(-DOMAIN) %>%
mutate(TRT01P = ARM, TRT01A = ACTARM)
# Parse the exposure datetimes, then take the first start and the last end.
ex_ext <- ex %>%
derive_vars_dtm(dtc = EXSTDTC, new_vars_prefix = "EXST") %>%
derive_vars_dtm(dtc = EXENDTC, new_vars_prefix = "EXEN",
time_imputation = "last")
adsl <- adsl %>%
derive_vars_merged(
dataset_add = ex_ext,
filter_add = (EXDOSE > 0 |
(EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO"))) & !is.na(EXSTDTM),
new_vars = exprs(TRTSDTM = EXSTDTM, TRTSTMF = EXSTTMF),
order = exprs(EXSTDTM, EXSEQ),
mode = "first",
by_vars = exprs(STUDYID, USUBJID)
) %>%
derive_vars_merged(
dataset_add = ex_ext,
filter_add = (EXDOSE > 0 |
(EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO"))) & !is.na(EXENDTM),
new_vars = exprs(TRTEDTM = EXENDTM, TRTETMF = EXENTMF),
order = exprs(EXENDTM, EXSEQ),
mode = "last",
by_vars = exprs(STUDYID, USUBJID)
) %>%
derive_vars_dtm_to_dt(source_vars = exprs(TRTSDTM, TRTEDTM)) %>%
derive_var_trtdurd()
adsl %>%
filter(!is.na(TRTSDT)) %>%
select(USUBJID, TRT01A, TRTSDT, TRTEDT, TRTDURD) %>%
head(4)
# A tibble: 4 × 5
USUBJID TRT01A TRTSDT TRTEDT TRTDURD
<chr> <chr> <date> <date> <dbl>
1 01-701-1015 Placebo 2014-01-02 2014-07-02 182
2 01-701-1023 Placebo 2012-08-05 2012-09-01 28
3 01-701-1028 Xanomeline High Dose 2013-07-19 2014-01-14 180
4 01-701-1033 Xanomeline Low Dose 2014-03-18 2014-03-31 14

The population flag is the same decision expressed as an existence check: a subject is in the safety population when at least one treated exposure record exists. Screen failures never reach EX, so they get SAFFL = "N" by construction.

adsl <- adsl %>%
derive_var_merged_exist_flag(
dataset_add = ex,
by_vars = exprs(STUDYID, USUBJID),
new_var = SAFFL,
false_value = "N",
missing_value = "N",
condition = (EXDOSE > 0 |
(EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO")))
)
table(adsl$SAFFL)
N Y
52 254

Because pharmaverseadam was generated by running admiral’s own ADSL template on the same SDTM inputs, the derivation above should land on the same numbers. Comparing the two is the traceability story in miniature: not “trust the package”, but “here is the derivation, and here is the reference it reproduces”.

ref_adsl <- pharmaverseadam::adsl
cmp <- adsl %>%
select(USUBJID, TRTSDT, SAFFL) %>%
inner_join(
select(ref_adsl, USUBJID, TRTSDT_REF = TRTSDT, SAFFL_REF = SAFFL),
by = "USUBJID"
)
cat("subjects compared:", nrow(cmp), "\n")
cat("SAFFL differing:", sum(cmp$SAFFL != cmp$SAFFL_REF), "\n")
cat("TRTSDT differing:", sum(cmp$TRTSDT != cmp$TRTSDT_REF, na.rm = TRUE) +
sum(xor(is.na(cmp$TRTSDT), is.na(cmp$TRTSDT_REF))), "\n")
cat("TRTSDT missing in both:",
sum(is.na(cmp$TRTSDT) & is.na(cmp$TRTSDT_REF)), "\n")
subjects compared: 306
SAFFL differing: 0
TRTSDT differing: 0
TRTSDT missing in both: 52

A BDS dataset starts from a findings domain and adds four things: the ADSL variables every analysis needs, a real analysis date, a parameter layer that renames the SDTM test codes into analysis parameters, and the analysis value. The reference adlb maps the cholesterol test to PARAMCD = "CHOLES", so the lookup here does the same to stay comparable.

lb <- convert_blanks_to_na(pharmaversesdtm::lb)
adsl_vars <- exprs(TRTSDT, TRTEDT, TRT01A, TRT01P)
param_lookup <- tibble::tribble(
~LBTESTCD, ~PARAMCD, ~PARAM,
"CHOL", "CHOLES", "Cholesterol (mmol/L)",
"ALT", "ALT", "Alanine Aminotransferase (U/L)"
)
adlb <- lb %>%
filter(LBTESTCD %in% param_lookup$LBTESTCD) %>%
derive_vars_merged(
dataset_add = adsl,
new_vars = adsl_vars,
by_vars = exprs(STUDYID, USUBJID)
) %>%
derive_vars_dt(new_vars_prefix = "A", dtc = LBDTC) %>%
derive_vars_dy(reference_date = TRTSDT, source_vars = exprs(ADT)) %>%
derive_vars_merged(
dataset_add = param_lookup,
new_vars = exprs(PARAMCD, PARAM),
by_vars = exprs(LBTESTCD)
) %>%
mutate(AVAL = LBSTRESN) %>%
# Analysis visits: screening collapses to Baseline, weeks keep their number.
mutate(
AVISIT = case_when(
stringr::str_detect(VISIT, "SCREEN") ~ "Baseline",
!is.na(VISIT) ~ stringr::str_to_title(VISIT),
TRUE ~ NA_character_
),
AVISITN = case_when(AVISIT == "Baseline" ~ 0, TRUE ~ VISITNUM)
)
adlb %>%
filter(USUBJID == "01-701-1015", PARAMCD == "CHOLES") %>%
select(USUBJID, PARAMCD, VISIT, AVISIT, ADT, ADY, AVAL) %>%
head(5)
# A tibble: 5 × 7
USUBJID PARAMCD VISIT AVISIT ADT ADY AVAL
<chr> <chr> <chr> <chr> <date> <dbl> <dbl>
1 01-701-1015 CHOLES SCREENING 1 Baseline 2013-12-26 -7 5.95
2 01-701-1015 CHOLES WEEK 2 Week 2 2014-01-16 15 5.48
3 01-701-1015 CHOLES WEEK 4 Week 4 2014-01-30 29 4.99
4 01-701-1015 CHOLES WEEK 6 Week 6 2014-02-12 42 5.64
5 01-701-1015 CHOLES WEEK 8 Week 8 2014-03-05 63 5.53

SDTM’s LBDTC is a character string and its LBDY counts from a study-specific reference; ADaM’s ADT is a real date and ADY counts days from the first dose, with day one being the day of first dose and no day zero. The visit mapping beside it is the same idea applied to timepoints: the collected visit name becomes a controlled analysis visit with a number that sorts.

The baseline rule for this study is the usual one: the last non-missing measurement on or before the first dose. restrict_derivation runs the extreme-record flag only on the rows that pass that filter, then derive_var_base copies the flagged value into every row of the parameter. Change from baseline is computed on post-baseline analysis visits only, which leaves CHG missing at baseline: the template programs the pharmaverseadam reference is generated from made that producer choice, and this page makes the same one so the comparison below is exact.

adlb <- adlb %>%
restrict_derivation(
derivation = derive_var_extreme_flag,
args = params(
by_vars = exprs(STUDYID, USUBJID, PARAMCD),
order = exprs(ADT, VISITNUM, LBSEQ),
new_var = ABLFL,
mode = "last"
),
filter = !is.na(AVAL) & ADT <= TRTSDT
) %>%
derive_var_base(
by_vars = exprs(STUDYID, USUBJID, PARAMCD),
source_var = AVAL,
new_var = BASE
) %>%
restrict_derivation(
derivation = derive_var_chg,
filter = AVISITN > 0
) %>%
restrict_derivation(
derivation = derive_var_pchg,
filter = AVISITN > 0
)
adlb %>%
filter(USUBJID == "01-701-1015", PARAMCD == "CHOLES") %>%
arrange(ADY) %>%
select(USUBJID, PARAMCD, AVISIT, ADY, AVAL, ABLFL, BASE, CHG) %>%
head(6)
# A tibble: 6 × 8
USUBJID PARAMCD AVISIT ADY AVAL ABLFL BASE CHG
<chr> <chr> <chr> <dbl> <dbl> <chr> <dbl> <dbl>
1 01-701-1015 CHOLES Baseline -7 5.95 Y 5.95 NA
2 01-701-1015 CHOLES Week 2 15 5.48 <NA> 5.95 -0.465
3 01-701-1015 CHOLES Week 4 29 4.99 <NA> 5.95 -0.957
4 01-701-1015 CHOLES Week 6 42 5.64 <NA> 5.95 -0.310
5 01-701-1015 CHOLES Week 8 63 5.53 <NA> 5.95 -0.414
6 01-701-1015 CHOLES Week 12 84 5.64 <NA> 5.95 -0.310

The same reference check applies to the derived columns. The shipped adlb carries extra derived records beyond the collected rows, extreme values and last on-treatment observations flagged in DTYPE, which this minimal derivation does not create, so the comparison keeps only the collected rows on both sides.

ref_adlb <- pharmaverseadam::adlb %>%
filter(PARAMCD %in% c("CHOLES", "ALT"), is.na(DTYPE))
cmp_lb <- adlb %>%
select(USUBJID, PARAMCD, ADT, AVAL, ABLFL, BASE, CHG) %>%
inner_join(
ref_adlb %>%
select(USUBJID, PARAMCD, ADT, ABLFL_REF = ABLFL,
BASE_REF = BASE, CHG_REF = CHG),
by = c("USUBJID", "PARAMCD", "ADT")
)
# NA-aware disagreement: NA on one side only, or unequal non-missing values.
cat("records compared:", nrow(cmp_lb), "\n")
cat("baseline flags differing:",
sum(xor(is.na(cmp_lb$ABLFL), is.na(cmp_lb$ABLFL_REF))) +
sum(cmp_lb$ABLFL != cmp_lb$ABLFL_REF, na.rm = TRUE), "\n")
cat("BASE differing:",
sum(xor(is.na(cmp_lb$BASE), is.na(cmp_lb$BASE_REF))) +
sum(cmp_lb$BASE != cmp_lb$BASE_REF, na.rm = TRUE), "\n")
cat("CHG differing:",
sum(xor(is.na(cmp_lb$CHG), is.na(cmp_lb$CHG_REF))) +
sum(cmp_lb$CHG != cmp_lb$CHG_REF, na.rm = TRUE), "\n")
records compared: 3642
baseline flags differing: 0
BASE differing: 0
CHG differing: 0

A submission ADSL goes further: disposition status from DS, death variables from AE and DS, age and region groupings, and the last-known-alive date assembled from every domain at once, all with the same derivation verbs. The therapeutic-area extensions (admiralonco, admiralvaccine, admiralophtha and siblings) package exactly that kind of extra derivation for their endpoints, and each admiral package ships its template programs via use_ad_template, so the programs that generated pharmaverseadam are the same ones you would start from. The tables page turns these datasets into the submission’s tables.