ADaM Datasets with admiral
SDTM organises what a trial collected. ADaM organises what the trial concludes, and
it does it in two dataset shapes. ADSL, the subject-level analysis dataset, holds one
row per subject with everything an analysis needs to know about them: treatment,
population flags, dates, stratification. Every other analysis dataset is BDS-shaped,
basic data structure: long, one row per subject per parameter per timepoint, with the
analysis value in AVAL, the baseline in BASE, and the change in CHG. A lab
dataset, a vitals dataset and a questionnaire dataset all share that skeleton, which
is why one set of derivation functions covers most of an ADaM submission.
admiral is the pharmaverse toolbox for those derivations. It was started by Roche
and GSK and is now a wider collaboration, and it is deliberately a programming
library rather than a one-button generator: you state each derivation as a function
call, and the call records the decision. This page builds a small ADSL from the pilot
study’s SDTM domains, then derives two lab parameters into BDS form, and at each
stage compares the result with pharmaverseadam, the reference ADaM datasets the
admiral team regenerates from its own template programs on every release.
library(admiral)library(dplyr)library(pharmaversesdtm)
# SDTM codes missing values as empty strings; turn them into real NAs.dm <- convert_blanks_to_na(pharmaversesdtm::dm)ex <- convert_blanks_to_na(pharmaversesdtm::ex)
dim(dm)[1] 306 28ADSL: treatment dates and population flags
Section titled “ADSL: treatment dates and population flags”The planned and actual treatment columns come straight from DM. The work is in the
dates: exposure starts and ends live in EX as character datetimes, possibly partial,
so they are parsed to real datetimes first, then the earliest start and latest end
per subject become TRTSDTM and TRTEDTM. The filter in the merge is the study’s
definition of treated: a dose above zero, or a zero dose whose name says placebo,
which is how infusion-placebo records look in this study.
adsl <- dm %>% select(-DOMAIN) %>% mutate(TRT01P = ARM, TRT01A = ACTARM)
# Parse the exposure datetimes, then take the first start and the last end.ex_ext <- ex %>% derive_vars_dtm(dtc = EXSTDTC, new_vars_prefix = "EXST") %>% derive_vars_dtm(dtc = EXENDTC, new_vars_prefix = "EXEN", time_imputation = "last")
adsl <- adsl %>% derive_vars_merged( dataset_add = ex_ext, filter_add = (EXDOSE > 0 | (EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO"))) & !is.na(EXSTDTM), new_vars = exprs(TRTSDTM = EXSTDTM, TRTSTMF = EXSTTMF), order = exprs(EXSTDTM, EXSEQ), mode = "first", by_vars = exprs(STUDYID, USUBJID) ) %>% derive_vars_merged( dataset_add = ex_ext, filter_add = (EXDOSE > 0 | (EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO"))) & !is.na(EXENDTM), new_vars = exprs(TRTEDTM = EXENDTM, TRTETMF = EXENTMF), order = exprs(EXENDTM, EXSEQ), mode = "last", by_vars = exprs(STUDYID, USUBJID) ) %>% derive_vars_dtm_to_dt(source_vars = exprs(TRTSDTM, TRTEDTM)) %>% derive_var_trtdurd()
adsl %>% filter(!is.na(TRTSDT)) %>% select(USUBJID, TRT01A, TRTSDT, TRTEDT, TRTDURD) %>% head(4)# A tibble: 4 × 5 USUBJID TRT01A TRTSDT TRTEDT TRTDURD <chr> <chr> <date> <date> <dbl>1 01-701-1015 Placebo 2014-01-02 2014-07-02 1822 01-701-1023 Placebo 2012-08-05 2012-09-01 283 01-701-1028 Xanomeline High Dose 2013-07-19 2014-01-14 1804 01-701-1033 Xanomeline Low Dose 2014-03-18 2014-03-31 14The population flag is the
same decision expressed as an existence check: a subject is in the safety population
when at least one treated exposure record exists. Screen failures never reach EX, so
they get SAFFL = "N" by construction.
adsl <- adsl %>% derive_var_merged_exist_flag( dataset_add = ex, by_vars = exprs(STUDYID, USUBJID), new_var = SAFFL, false_value = "N", missing_value = "N", condition = (EXDOSE > 0 | (EXDOSE == 0 & stringr::str_detect(EXTRT, "PLACEBO"))) )
table(adsl$SAFFL) N Y 52 254Checking against the shipped reference
Section titled “Checking against the shipped reference”Because pharmaverseadam was generated by running admiral’s own ADSL template on the
same SDTM inputs, the derivation above should land on the same numbers. Comparing
the two is the traceability story in miniature: not “trust the package”, but “here
is the derivation, and here is the reference it reproduces”.
ref_adsl <- pharmaverseadam::adsl
cmp <- adsl %>% select(USUBJID, TRTSDT, SAFFL) %>% inner_join( select(ref_adsl, USUBJID, TRTSDT_REF = TRTSDT, SAFFL_REF = SAFFL), by = "USUBJID" )
cat("subjects compared:", nrow(cmp), "\n")cat("SAFFL differing:", sum(cmp$SAFFL != cmp$SAFFL_REF), "\n")cat("TRTSDT differing:", sum(cmp$TRTSDT != cmp$TRTSDT_REF, na.rm = TRUE) + sum(xor(is.na(cmp$TRTSDT), is.na(cmp$TRTSDT_REF))), "\n")cat("TRTSDT missing in both:", sum(is.na(cmp$TRTSDT) & is.na(cmp$TRTSDT_REF)), "\n")subjects compared: 306SAFFL differing: 0TRTSDT differing: 0TRTSDT missing in both: 52ADLB: the BDS skeleton
Section titled “ADLB: the BDS skeleton”A BDS dataset starts from a findings domain and adds four things: the ADSL variables
every analysis needs, a real analysis date, a parameter layer that renames the SDTM
test codes into analysis parameters, and the analysis value. The reference adlb
maps the cholesterol test to PARAMCD = "CHOLES", so the lookup here does the same
to stay comparable.
lb <- convert_blanks_to_na(pharmaversesdtm::lb)
adsl_vars <- exprs(TRTSDT, TRTEDT, TRT01A, TRT01P)
param_lookup <- tibble::tribble( ~LBTESTCD, ~PARAMCD, ~PARAM, "CHOL", "CHOLES", "Cholesterol (mmol/L)", "ALT", "ALT", "Alanine Aminotransferase (U/L)")
adlb <- lb %>% filter(LBTESTCD %in% param_lookup$LBTESTCD) %>% derive_vars_merged( dataset_add = adsl, new_vars = adsl_vars, by_vars = exprs(STUDYID, USUBJID) ) %>% derive_vars_dt(new_vars_prefix = "A", dtc = LBDTC) %>% derive_vars_dy(reference_date = TRTSDT, source_vars = exprs(ADT)) %>% derive_vars_merged( dataset_add = param_lookup, new_vars = exprs(PARAMCD, PARAM), by_vars = exprs(LBTESTCD) ) %>% mutate(AVAL = LBSTRESN) %>% # Analysis visits: screening collapses to Baseline, weeks keep their number. mutate( AVISIT = case_when( stringr::str_detect(VISIT, "SCREEN") ~ "Baseline", !is.na(VISIT) ~ stringr::str_to_title(VISIT), TRUE ~ NA_character_ ), AVISITN = case_when(AVISIT == "Baseline" ~ 0, TRUE ~ VISITNUM) )
adlb %>% filter(USUBJID == "01-701-1015", PARAMCD == "CHOLES") %>% select(USUBJID, PARAMCD, VISIT, AVISIT, ADT, ADY, AVAL) %>% head(5)# A tibble: 5 × 7 USUBJID PARAMCD VISIT AVISIT ADT ADY AVAL <chr> <chr> <chr> <chr> <date> <dbl> <dbl>1 01-701-1015 CHOLES SCREENING 1 Baseline 2013-12-26 -7 5.952 01-701-1015 CHOLES WEEK 2 Week 2 2014-01-16 15 5.483 01-701-1015 CHOLES WEEK 4 Week 4 2014-01-30 29 4.994 01-701-1015 CHOLES WEEK 6 Week 6 2014-02-12 42 5.645 01-701-1015 CHOLES WEEK 8 Week 8 2014-03-05 63 5.53SDTM’s LBDTC is a character string and its LBDY counts from a study-specific
reference; ADaM’s ADT is a real date and ADY counts days from the first dose,
with day one being the day of first dose and no day zero. The visit mapping beside it is the same idea applied to timepoints: the
collected visit name becomes a controlled analysis visit with a number that sorts.
Baseline and change from baseline
Section titled “Baseline and change from baseline”The baseline rule for this study is the usual one: the last non-missing measurement
on or before the first dose. restrict_derivation runs the extreme-record flag only
on the rows that pass that filter, then derive_var_base copies the flagged value
into every row of the parameter. Change from baseline is computed on post-baseline
analysis visits only, which leaves CHG missing at baseline: the template programs
the pharmaverseadam reference is generated from made that producer choice, and this
page makes the same one so the comparison below is exact.
adlb <- adlb %>% restrict_derivation( derivation = derive_var_extreme_flag, args = params( by_vars = exprs(STUDYID, USUBJID, PARAMCD), order = exprs(ADT, VISITNUM, LBSEQ), new_var = ABLFL, mode = "last" ), filter = !is.na(AVAL) & ADT <= TRTSDT ) %>% derive_var_base( by_vars = exprs(STUDYID, USUBJID, PARAMCD), source_var = AVAL, new_var = BASE ) %>% restrict_derivation( derivation = derive_var_chg, filter = AVISITN > 0 ) %>% restrict_derivation( derivation = derive_var_pchg, filter = AVISITN > 0 )
adlb %>% filter(USUBJID == "01-701-1015", PARAMCD == "CHOLES") %>% arrange(ADY) %>% select(USUBJID, PARAMCD, AVISIT, ADY, AVAL, ABLFL, BASE, CHG) %>% head(6)# A tibble: 6 × 8 USUBJID PARAMCD AVISIT ADY AVAL ABLFL BASE CHG <chr> <chr> <chr> <dbl> <dbl> <chr> <dbl> <dbl>1 01-701-1015 CHOLES Baseline -7 5.95 Y 5.95 NA2 01-701-1015 CHOLES Week 2 15 5.48 <NA> 5.95 -0.4653 01-701-1015 CHOLES Week 4 29 4.99 <NA> 5.95 -0.9574 01-701-1015 CHOLES Week 6 42 5.64 <NA> 5.95 -0.3105 01-701-1015 CHOLES Week 8 63 5.53 <NA> 5.95 -0.4146 01-701-1015 CHOLES Week 12 84 5.64 <NA> 5.95 -0.310The same reference check applies to the derived columns. The shipped adlb
carries extra derived records beyond the collected rows, extreme values and last
on-treatment observations flagged in DTYPE, which this minimal derivation does not
create, so the comparison keeps only the collected rows on both sides.
ref_adlb <- pharmaverseadam::adlb %>% filter(PARAMCD %in% c("CHOLES", "ALT"), is.na(DTYPE))
cmp_lb <- adlb %>% select(USUBJID, PARAMCD, ADT, AVAL, ABLFL, BASE, CHG) %>% inner_join( ref_adlb %>% select(USUBJID, PARAMCD, ADT, ABLFL_REF = ABLFL, BASE_REF = BASE, CHG_REF = CHG), by = c("USUBJID", "PARAMCD", "ADT") )
# NA-aware disagreement: NA on one side only, or unequal non-missing values.cat("records compared:", nrow(cmp_lb), "\n")cat("baseline flags differing:", sum(xor(is.na(cmp_lb$ABLFL), is.na(cmp_lb$ABLFL_REF))) + sum(cmp_lb$ABLFL != cmp_lb$ABLFL_REF, na.rm = TRUE), "\n")cat("BASE differing:", sum(xor(is.na(cmp_lb$BASE), is.na(cmp_lb$BASE_REF))) + sum(cmp_lb$BASE != cmp_lb$BASE_REF, na.rm = TRUE), "\n")cat("CHG differing:", sum(xor(is.na(cmp_lb$CHG), is.na(cmp_lb$CHG_REF))) + sum(cmp_lb$CHG != cmp_lb$CHG_REF, na.rm = TRUE), "\n")records compared: 3642baseline flags differing: 0BASE differing: 0CHG differing: 0What a production ADSL adds
Section titled “What a production ADSL adds”A submission ADSL goes further: disposition status from DS, death variables from AE
and DS, age and region groupings, and the last-known-alive date assembled from every
domain at once, all with the same derivation verbs. The therapeutic-area extensions
(admiralonco, admiralvaccine, admiralophtha and siblings) package exactly that
kind of extra derivation for their endpoints, and each admiral package ships its
template programs via use_ad_template, so the programs that generated
pharmaverseadam are the same ones you would start from. The tables
page turns these datasets into the submission’s
tables.