Skip to content

CDISC and SDTM Data in R

Clinical trial data in pharma moves between two CDISC standards. SDTM, the Study Data Tabulation Model, organises what was collected on the trial: every observation lands in a named domain with fixed variable names, fixed types and controlled terminology. ADaM, covered on the next page, organises what gets analysed. Both are required for submission to the FDA and to the PMDA in Japan, so a statistician at one company reads another company’s submission datasets without a translator.

An SDTM study is a set of two-letter domains. Three classes cover most of the data. Interventions hold what was given to the subject: EX for study treatment, CM for concomitant medications. Events hold what happened: AE for adverse events, MH for medical history, DS for disposition. Findings hold what was measured: LB for laboratory results, VS for vital signs, EG for ECG readings. Beside them sit the special-purpose domains, of which DM, one row per subject with the demographics and the arm assignment, is the anchor every other domain joins on through USUBJID.

Reading a real SDTM study needs nothing more than a data package. pharmaversesdtm ships example domains for the whole pharmaverse, several of them sourced from the CDISC pilot project, a real, anonymised study of xanomeline in Alzheimer’s disease that CDISC published so the industry could test its tooling. It holds 64 datasets in version 1.5.0: the therapy-agnostic core named after the domain itself, plus therapy-area variants such as rs_onco for oncology response and oe_ophtha for ophthalmology.

library(pharmaversesdtm)
library(dplyr)
library(purrr)
# The therapy-agnostic core domains, with their sizes.
domains <- c("dm", "vs", "ae", "lb", "ex", "mh", "ds", "cm", "eg", "ts")
map_dfr(domains, function(d) {
data(list = d, package = "pharmaversesdtm", envir = environment())
x <- get(d, envir = environment())
data.frame(domain = d, rows = nrow(x), columns = ncol(x))
})
domain rows columns
1 dm 306 28
2 vs 29643 24
3 ae 1191 35
4 lb 59580 23
5 ex 591 17
6 mh 1818 28
7 ds 850 13
8 cm 7510 22
9 eg 26717 23
10 ts 33 6

DM is where every analysis starts, because it carries the arm. The pilot study, CDISCPILOT01, randomised subjects to placebo or to one of two xanomeline doses; the screen failures stay in the data because SDTM records everyone who entered the study, not everyone who was treated. Note the two arm pairs: ARM is the planned arm, ACTARM the arm the subject actually received, and the two can differ.

data(dm)
# Planned arm, and the identifiers every other domain joins on.
table(dm$ARM)
sub_dm <- dm[1:4, c("USUBJID", "SUBJID", "SITEID", "AGE", "SEX",
"ARM", "ACTARM")]
print(as.data.frame(sub_dm), row.names = FALSE)
Placebo Screen Failure Xanomeline High Dose
86 52 84
Xanomeline Low Dose
84
USUBJID SUBJID SITEID AGE SEX ARM ACTARM
01-701-1015 1015 701 63 F Placebo Placebo
01-701-1023 1023 701 64 M Placebo Placebo
01-701-1028 1028 701 71 M Xanomeline High Dose Xanomeline High Dose
01-701-1033 1033 701 74 M Xanomeline Low Dose Xanomeline Low Dose

The TS domain is smaller and stranger: one row per trial-level fact rather than per subject or per observation. It is how the study says, in machine-readable form, that it was double-blind, placebo-controlled, and restricted to subjects over fifty.

data(ts)
# A trial summary is one row per trial-level fact.
ts[ts$TSPARMCD %in% c("AGEMIN", "AGEMAX", "TBLIND", "TCNTRL", "ADDON"),
c("TSPARMCD", "TSVAL")]
# A tibble: 5 × 2
TSPARMCD TSVAL
<chr> <chr>
1 ADDON Y
2 AGEMAX No maximum
3 AGEMIN 50 years
4 TBLIND DOUBLE BLIND
5 TCNTRL PLACEBO

Findings are long, and split original from standard

Section titled “Findings are long, and split original from standard”

The findings domains are the long tables of SDTM: one row per measurement, keyed by the test code and the visit. LB shows the design pattern clearly. Each record carries the result twice: LBORRES, the result as collected, a character string in the original unit, and LBSTRESN, the same result as a number in the standardised unit. Analysis needs the number; traceability keeps the string.

data(lb)
# One subject's albumin records, original result beside the standardised one.
alb <- lb[lb$USUBJID == "01-701-1015" & lb$LBTESTCD == "ALB", ]
alb[, c("USUBJID", "LBTESTCD", "LBORRES", "LBORRESU",
"LBSTRESN", "LBSTRESU", "VISIT", "LBDY")]
# A tibble: 10 × 8
USUBJID LBTESTCD LBORRES LBORRESU LBSTRESN LBSTRESU VISIT LBDY
<chr> <chr> <chr> <chr> <dbl> <chr> <chr> <dbl>
1 01-701-1015 ALB 3.8 g/dL 38 g/L SCREENING 1 -7
2 01-701-1015 ALB 3.9 g/dL 39 g/L WEEK 2 15
3 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 4 29
4 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 6 42
5 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 8 63
6 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 12 84
7 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 16 126
8 01-701-1015 ALB 3.7 g/dL 37 g/L WEEK 20 140
9 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 24 168
10 01-701-1015 ALB 3.8 g/dL 38 g/L WEEK 26 182

Events use the same skeleton with controlled terminology on top. The AE domain codes every term with MedDRA, from the verbatim report (AETERM) to the coded preferred term (AEDECOD) and the system organ class (AESOC), and fixes severity to a three-level set. The pilot study records 1,191 events in 225 subjects.

data(ae)
# Controlled terminology: severity is a fixed three-level set.
table(ae$AESEV)
# One event, from the verbatim term to the coded hierarchy.
one_ae <- ae[1, c("AETERM", "AEDECOD", "AEBODSYS", "AESEV",
"AESTDTC", "AESTDY")]
print(as.data.frame(one_ae), row.names = FALSE)
MILD MODERATE SEVERE
770 378 43
AETERM AEDECOD
APPLICATION SITE ERYTHEMA APPLICATION SITE ERYTHEMA
AEBODSYS AESEV AESTDTC AESTDY
GENERAL DISORDERS AND ADMINISTRATION SITE CONDITIONS MILD 2014-01-03 2

SDTM is strict about which variables a domain may hold, so data that does not fit goes into a supplemental domain: SUPPDM holds extra facts about DM records, keyed by RDOMAIN, IDVAR and IDVARVAL, with the fact itself in QNAM and QVAL. The pattern is called SUPP–, and every domain can have one. In the pilot study the supplemental facts are the population and completer flags, ITT, SAFETY and the COMPLT8/16/24 week markers, which the analysis datasets then promote to real columns.

data(suppdm)
# Supplemental qualifiers: name-value pairs hung off DM records.
count(suppdm, QNAM)
sub_supp <- suppdm[1:3, c("STUDYID", "RDOMAIN", "USUBJID", "QNAM",
"QVAL", "QLABEL")]
print(as.data.frame(sub_supp), row.names = FALSE)
# A tibble: 6 × 2
QNAM n
<chr> <int>
1 COMPLT16 147
2 COMPLT24 118
3 COMPLT8 190
4 EFFICACY 234
5 ITT 254
6 SAFETY 254
STUDYID RDOMAIN USUBJID QNAM QVAL
CDISCPILOT01 DM 01-701-1015 COMPLT16 Y
CDISCPILOT01 DM 01-701-1015 COMPLT24 Y
CDISCPILOT01 DM 01-701-1015 COMPLT8 Y
QLABEL
Completers of Week 16 Population Flag
Completers of Week 24 Population Flag
Completers of Week 8 Population Flag

Reading SDTM is the common task. Writing it is the expensive one: raw EDC extracts arrive with sponsor-specific names and formats, and someone must map every collected field to its SDTM variable with the right controlled terminology. Roche open-sourced its engine for that mapping as sdtm.oak, which treats the conversion as a sequence of small, reusable algorithms: assign a collected value to a variable, assign it through a codelist, hardcode a value under a condition, or assemble an ISO 8601 datetime from date and time pieces.

The package ships a raw concomitant-medication extract and a codelist table, so the four algorithm types run end to end. The key detail is generate_oak_id_vars, which tags every raw record with an identifier the target rows keep, which is what makes a derived variable traceable back to the raw record it came from.

library(sdtm.oak)
cm_raw <- read.csv(
system.file("raw_data/cm_raw_data.csv", package = "sdtm.oak")
)
study_ct <- read.csv(system.file("raw_data/sdtm_ct.csv", package = "sdtm.oak"))
# The raw extract as the EDC system exports it.
cm_raw[1:3, c("PATNUM", "MDRAW", "DOSU", "MDBDR", "MDBTM", "MDPRIOR")]
PATNUM MDRAW DOSU MDBDR MDBTM MDPRIOR
1 375 BABY ASPIRIN mg 1
2 375 CORTISPORIN g 15-Sep-20 0
3 376 ASPIRIN 17-Feb-21 8:00 0
cm_raw <- generate_oak_id_vars(cm_raw, pat_var = "PATNUM", raw_src = "cm_raw")
cm <- assign_no_ct(
raw_dat = cm_raw,
raw_var = "MDRAW",
tgt_var = "CMTRT"
) %>%
assign_ct(
raw_dat = cm_raw,
raw_var = "DOSU",
tgt_var = "CMDOSU",
ct_spec = study_ct,
ct_clst = "C71620",
id_vars = oak_id_vars()
) %>%
assign_datetime(
raw_dat = cm_raw,
raw_var = c("MDBDR", "MDBTM"),
tgt_var = "CMSTDTC",
raw_fmt = c(list(c("d-m-y", "dd mmm yyyy")), "H:M"),
raw_unk = c("UN", "UNK"),
id_vars = oak_id_vars()
) %>%
hardcode_ct(
raw_dat = condition_add(cm_raw, MDPRIOR == "1"),
raw_var = "MDPRIOR",
tgt_var = "CMSTRTPT",
tgt_val = "BEFORE",
ct_spec = study_ct,
ct_clst = "C66728",
id_vars = oak_id_vars()
)
cm[1:4, c("oak_id", "patient_number", "CMTRT", "CMDOSU", "CMSTDTC", "CMSTRTPT")]
# A tibble: 4 × 6
oak_id patient_number CMTRT CMDOSU CMSTDTC CMSTRTPT
<int> <int> <chr> <chr> <iso8601> <chr>
1 1 375 BABY ASPIRIN "mg" NA BEFORE
2 2 375 CORTISPORIN "g" 2020-09-15 <NA>
3 3 376 ASPIRIN "" 2021-02-17T08:00 <NA>
4 4 377 DIPHENHYDRAMINE HCL "mg" 2020-10-04T09:00 <NA>

The four calls map the four problems a mapping specification actually states. assign_no_ct copies a collected value into the topic variable CMTRT. assign_ct maps the collected dose unit through the CDISC codelist C71620, so whatever the site typed becomes the submission term. assign_datetime parses a day-month-year date and an hour-minute time, tolerating the UN and UNK unknown codes, into one ISO 8601 string. hardcode_ct sets CMSTRTPT to BEFORE, but only on the records condition_add keeps, where the prior flag is set. Each target row carries its oak_id, so the mapping stays auditable record by record.

The rest of a production mapping is more of the same shape: study identifiers and sequence numbers get derived, study days are computed against DM with derive_study_day, and the result is written out per domain. A mapping specification reads as a table of these small decisions, and sdtm.oak makes each decision one testable function call.

The checks that catch a bad mapping at scale, a library of cross-domain consistency rules, live in the companion package sdtmchecks. The next step in the pipeline is the analysis model: ADaM datasets with admiral derives ADSL and ADLB from the domains read here, and checks the result against the reference outputs the admiral team ships.