CDISC and SDTM Data in R
Clinical trial data in pharma moves between two CDISC standards. SDTM, the Study Data Tabulation Model, organises what was collected on the trial: every observation lands in a named domain with fixed variable names, fixed types and controlled terminology. ADaM, covered on the next page, organises what gets analysed. Both are required for submission to the FDA and to the PMDA in Japan, so a statistician at one company reads another company’s submission datasets without a translator.
An SDTM study is a set of two-letter domains. Three classes cover most of the data.
Interventions hold what was given to the subject: EX for study treatment, CM for
concomitant medications. Events hold what happened: AE for adverse events, MH for
medical history, DS for disposition. Findings hold what was measured: LB for
laboratory results, VS for vital signs, EG for ECG readings. Beside them sit the
special-purpose domains, of which DM, one row per subject with the demographics and
the arm assignment, is the anchor every other domain joins on through USUBJID.
Reading a real SDTM study needs nothing more than a data package. pharmaversesdtm
ships example domains for the whole pharmaverse, several of them sourced from the
CDISC pilot
project, a real,
anonymised study of xanomeline in Alzheimer’s disease that CDISC published so the
industry could test its tooling. It holds 64 datasets in version 1.5.0: the
therapy-agnostic core named after the domain itself, plus therapy-area variants such
as rs_onco for oncology response and oe_ophtha for ophthalmology.
library(pharmaversesdtm)library(dplyr)library(purrr)
# The therapy-agnostic core domains, with their sizes.domains <- c("dm", "vs", "ae", "lb", "ex", "mh", "ds", "cm", "eg", "ts")map_dfr(domains, function(d) { data(list = d, package = "pharmaversesdtm", envir = environment()) x <- get(d, envir = environment()) data.frame(domain = d, rows = nrow(x), columns = ncol(x))}) domain rows columns1 dm 306 282 vs 29643 243 ae 1191 354 lb 59580 235 ex 591 176 mh 1818 287 ds 850 138 cm 7510 229 eg 26717 2310 ts 33 6DM: one row per subject
Section titled “DM: one row per subject”DM is where every analysis starts, because it carries the arm. The pilot study,
CDISCPILOT01, randomised subjects to placebo or to one of two xanomeline doses;
the screen failures stay in the data because SDTM records everyone who entered the
study, not everyone who was treated. Note the two arm pairs: ARM is the planned
arm, ACTARM the arm the subject actually received, and the two can differ.
data(dm)
# Planned arm, and the identifiers every other domain joins on.table(dm$ARM)sub_dm <- dm[1:4, c("USUBJID", "SUBJID", "SITEID", "AGE", "SEX", "ARM", "ACTARM")]print(as.data.frame(sub_dm), row.names = FALSE) Placebo Screen Failure Xanomeline High Dose 86 52 84 Xanomeline Low Dose 84 USUBJID SUBJID SITEID AGE SEX ARM ACTARM 01-701-1015 1015 701 63 F Placebo Placebo 01-701-1023 1023 701 64 M Placebo Placebo 01-701-1028 1028 701 71 M Xanomeline High Dose Xanomeline High Dose 01-701-1033 1033 701 74 M Xanomeline Low Dose Xanomeline Low DoseThe TS domain is smaller and stranger: one row per trial-level fact rather than per subject or per observation. It is how the study says, in machine-readable form, that it was double-blind, placebo-controlled, and restricted to subjects over fifty.
data(ts)
# A trial summary is one row per trial-level fact.ts[ts$TSPARMCD %in% c("AGEMIN", "AGEMAX", "TBLIND", "TCNTRL", "ADDON"), c("TSPARMCD", "TSVAL")]# A tibble: 5 × 2 TSPARMCD TSVAL <chr> <chr>1 ADDON Y2 AGEMAX No maximum3 AGEMIN 50 years4 TBLIND DOUBLE BLIND5 TCNTRL PLACEBOFindings are long, and split original from standard
Section titled “Findings are long, and split original from standard”The findings domains are the long tables of SDTM: one row per measurement, keyed by
the test code and the visit. LB shows the design pattern clearly. Each record carries
the result twice: LBORRES, the result as collected, a character string in the
original unit, and LBSTRESN, the same result as a number in the standardised unit.
Analysis needs the number; traceability keeps the string.
data(lb)
# One subject's albumin records, original result beside the standardised one.alb <- lb[lb$USUBJID == "01-701-1015" & lb$LBTESTCD == "ALB", ]alb[, c("USUBJID", "LBTESTCD", "LBORRES", "LBORRESU", "LBSTRESN", "LBSTRESU", "VISIT", "LBDY")]# A tibble: 10 × 8 USUBJID LBTESTCD LBORRES LBORRESU LBSTRESN LBSTRESU VISIT LBDY <chr> <chr> <chr> <chr> <dbl> <chr> <chr> <dbl> 1 01-701-1015 ALB 3.8 g/dL 38 g/L SCREENING 1 -7 2 01-701-1015 ALB 3.9 g/dL 39 g/L WEEK 2 15 3 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 4 29 4 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 6 42 5 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 8 63 6 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 12 84 7 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 16 126 8 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 20 140 9 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 24 16810 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 26 182Events use the same skeleton with controlled terminology on top. The AE domain
codes every term with MedDRA, from the verbatim report (AETERM) to the coded
preferred term (AEDECOD) and the system organ class (AESOC), and fixes severity
to a three-level set. The pilot study records 1,191 events in 225 subjects.
data(ae)
# Controlled terminology: severity is a fixed three-level set.table(ae$AESEV)
# One event, from the verbatim term to the coded hierarchy.one_ae <- ae[1, c("AETERM", "AEDECOD", "AEBODSYS", "AESEV", "AESTDTC", "AESTDY")]print(as.data.frame(one_ae), row.names = FALSE) MILD MODERATE SEVERE 770 378 43 AETERM AEDECOD APPLICATION SITE ERYTHEMA APPLICATION SITE ERYTHEMA AEBODSYS AESEV AESTDTC AESTDY GENERAL DISORDERS AND ADMINISTRATION SITE CONDITIONS MILD 2014-01-03 2When a variable does not fit the model
Section titled “When a variable does not fit the model”SDTM is strict about which variables a domain may hold, so data that does not fit
goes into a supplemental domain: SUPPDM holds extra facts about DM records, keyed by
RDOMAIN, IDVAR and IDVARVAL, with the fact itself in QNAM and QVAL. The
pattern is called SUPP–, and every domain can have one. In the pilot study the
supplemental facts are the population and completer flags, ITT, SAFETY and the
COMPLT8/16/24 week markers, which the analysis datasets then promote to real
columns.
data(suppdm)
# Supplemental qualifiers: name-value pairs hung off DM records.count(suppdm, QNAM)sub_supp <- suppdm[1:3, c("STUDYID", "RDOMAIN", "USUBJID", "QNAM", "QVAL", "QLABEL")]print(as.data.frame(sub_supp), row.names = FALSE)# A tibble: 6 × 2 QNAM n <chr> <int>1 COMPLT16 1472 COMPLT24 1183 COMPLT8 1904 EFFICACY 2345 ITT 2546 SAFETY 254 STUDYID RDOMAIN USUBJID QNAM QVAL CDISCPILOT01 DM 01-701-1015 COMPLT16 Y CDISCPILOT01 DM 01-701-1015 COMPLT24 Y CDISCPILOT01 DM 01-701-1015 COMPLT8 Y QLABEL Completers of Week 16 Population Flag Completers of Week 24 Population Flag Completers of Week 8 Population FlagMaking a domain: sdtm.oak
Section titled “Making a domain: sdtm.oak”Reading SDTM is the common task. Writing it is the expensive one: raw EDC extracts
arrive with sponsor-specific names and formats, and someone must map every collected
field to its SDTM variable with the right controlled terminology. Roche open-sourced
its engine for that mapping as sdtm.oak, which treats the conversion as a sequence
of small, reusable algorithms: assign a collected value to a variable, assign it
through a codelist, hardcode a value under a condition, or assemble an ISO 8601
datetime from date and time pieces.
The package ships a raw concomitant-medication extract and a codelist table, so the
four algorithm types run end to end. The key detail is generate_oak_id_vars, which
tags every raw record with an identifier the target rows keep, which is what makes a
derived variable traceable back to the raw record it came from.
library(sdtm.oak)
cm_raw <- read.csv( system.file("raw_data/cm_raw_data.csv", package = "sdtm.oak"))study_ct <- read.csv(system.file("raw_data/sdtm_ct.csv", package = "sdtm.oak"))
# The raw extract as the EDC system exports it.cm_raw[1:3, c("PATNUM", "MDRAW", "DOSU", "MDBDR", "MDBTM", "MDPRIOR")] PATNUM MDRAW DOSU MDBDR MDBTM MDPRIOR1 375 BABY ASPIRIN mg 12 375 CORTISPORIN g 15-Sep-20 03 376 ASPIRIN 17-Feb-21 8:00 0cm_raw <- generate_oak_id_vars(cm_raw, pat_var = "PATNUM", raw_src = "cm_raw")
cm <- assign_no_ct( raw_dat = cm_raw, raw_var = "MDRAW", tgt_var = "CMTRT") %>% assign_ct( raw_dat = cm_raw, raw_var = "DOSU", tgt_var = "CMDOSU", ct_spec = study_ct, ct_clst = "C71620", id_vars = oak_id_vars() ) %>% assign_datetime( raw_dat = cm_raw, raw_var = c("MDBDR", "MDBTM"), tgt_var = "CMSTDTC", raw_fmt = c(list(c("d-m-y", "dd mmm yyyy")), "H:M"), raw_unk = c("UN", "UNK"), id_vars = oak_id_vars() ) %>% hardcode_ct( raw_dat = condition_add(cm_raw, MDPRIOR == "1"), raw_var = "MDPRIOR", tgt_var = "CMSTRTPT", tgt_val = "BEFORE", ct_spec = study_ct, ct_clst = "C66728", id_vars = oak_id_vars() )
cm[1:4, c("oak_id", "patient_number", "CMTRT", "CMDOSU", "CMSTDTC", "CMSTRTPT")]# A tibble: 4 × 6 oak_id patient_number CMTRT CMDOSU CMSTDTC CMSTRTPT <int> <int> <chr> <chr> <iso8601> <chr>1 1 375 BABY ASPIRIN "mg" NA BEFORE2 2 375 CORTISPORIN "g" 2020-09-15 <NA>3 3 376 ASPIRIN "" 2021-02-17T08:00 <NA>4 4 377 DIPHENHYDRAMINE HCL "mg" 2020-10-04T09:00 <NA>The four calls map the four problems a mapping specification actually states.
assign_no_ct copies a collected value into the topic variable CMTRT.
assign_ct maps the collected dose unit through the CDISC codelist C71620, so
whatever the site typed becomes the submission term. assign_datetime parses a
day-month-year date and an hour-minute time, tolerating the UN and UNK unknown
codes, into one ISO 8601 string. hardcode_ct sets CMSTRTPT to BEFORE, but only
on the records condition_add keeps, where the prior flag is set. Each target row
carries its oak_id, so the mapping stays auditable record by record.
The rest of a production mapping is more of the same shape: study identifiers and
sequence numbers get derived, study days are computed against DM with
derive_study_day, and the result is written out per domain. A mapping
specification reads as a table of these small decisions, and sdtm.oak makes each
decision one testable function call.
The checks that catch a bad mapping at scale, a library of cross-domain consistency
rules, live in the companion package sdtmchecks. The next step in the pipeline is
the analysis model: ADaM datasets with
admiral derives ADSL and ADLB from the domains
read here, and checks the result against the reference outputs the admiral team
ships.