This R Markdown document illustrates example usage of the RESIDE
package with multiple related tables, using the Demographics (DM) and
Adverse Events (AE) SDTM domains from the pharmaversesdtm
package. The tables are linked by a subject identifier
(USUBJID), DM has one row per subject and AE has one row
per adverse event.
Load the RESIDE package and set a seed for reproducibility and store the folder directory for export / import.
Store the tables in a named list, the names are used to identify the tables in the marginal distributions.
# Store the tables in a named list
dfs <- list(
dm = pharmaversesdtm::dm,
ae = pharmaversesdtm::ae
)
# Number of rows in each table
sapply(dfs, nrow)
#> dm ae
#> 306 1191
# Number of subjects in each table
sapply(dfs, function(df) length(unique(df$USUBJID)))
#> dm ae
#> 306 225
# Treatment arms of the subjects
table(dfs$dm$ARM)
#>
#> Placebo Screen Failure Xanomeline High Dose
#> 86 52 84
#> Xanomeline Low Dose
#> 84
# Severity of the adverse events
table(dfs$ae$AESEV)
#>
#> MILD MODERATE SEVERE
#> 770 378 43Join the adverse events to the demographics of the subject and fit a logistic regression model for the odds of an adverse event being moderate or severe, by age, sex and treatment arm. The same function is used for the original and synthesised data.
# Join the adverse events to the demographics of each subject
prepare_ae_data <- function(dm, ae) {
ae |>
select(USUBJID, AESEV) |>
inner_join(select(dm, USUBJID, AGE, SEX, ARM), by = "USUBJID") |>
# Remove screen failures and adverse events without a severity
filter(ARM != "Screen Failure", AESEV != "") |>
mutate(
MOD_SEV = AESEV %in% c("MODERATE", "SEVERE"),
ARM = relevel(factor(ARM), "Placebo")
)
}
ae_original <- prepare_ae_data(dfs$dm, dfs$ae)
# Proportion of moderate or severe adverse events by treatment arm
prop.table(table(ae_original$ARM, ae_original$MOD_SEV), 1)
#>
#> FALSE TRUE
#> Placebo 0.7275748 0.2724252
#> Xanomeline High Dose 0.6725275 0.3274725
#> Xanomeline Low Dose 0.5632184 0.4367816
# Fit a logistic regression model
glm.original <- glm(
MOD_SEV ~ AGE + SEX + ARM,
data = ae_original,
family = binomial
)
# Output the coefficients of the model
summary(glm.original)$coefficients
#> Estimate Std. Error z value Pr(>|z|)
#> (Intercept) -0.305106216 0.611013964 -0.4993441 6.175370e-01
#> AGE -0.006640127 0.007838388 -0.8471291 3.969232e-01
#> SEXM -0.435382516 0.126327520 -3.4464582 5.679865e-04
#> ARMXanomeline High Dose 0.331225289 0.167062099 1.9826477 4.740679e-02
#> ARMXanomeline Low Dose 0.740406571 0.162802809 4.5478734 5.419071e-06NB Adverse events from the same subject are not independent, a mixed effects model would be more appropriate, however a logistic regression model is used here for simplicity.
Use the get_marginal_distributions() function to get the
marginal distributions of the tables, specifying the column that
identifies each subject using the subject_identifier
parameter.
# Get the Marginal Distributions of both tables
marginals <- get_marginal_distributions(
dfs,
subject_identifier = "USUBJID"
)
#> Registered S3 method overwritten by 'butcher':
#> method from
#> as.character.dev_topic generics
# Summarise the Marginal Distributions
summary(marginals)
#> Summary of Marginal Distributions
#> Number of Data Frames: 2
#> Subject Identifier: USUBJID
#> Number of Subjects: 306
#> Common Columns: STUDYID
#>
#> Data Frame Rows Subjects Variables Categorical Binary Continuous Dates Missing
#> dm 306 306 27 17 0 10 8 10
#> ae 1191 225 34 28 0 6 3 12Columns that are present in every table (other than the subject
identifier), are identified as common columns, in this case the study
identifier STUDYID.
Export the marginal distributions using the
export_marginal_distributions() function, using the
force parameter to override any existing files.
# Export the Marginal Distributions
export_marginal_distributions(marginals,
folder_path = folder_path,
force = TRUE)
#> Exporting Categorical variables to: C:\Users\Ryan\AppData\Local\Temp\RtmpMRdU12/categorical_variables.csv
#> Exporting Continuous variables to: C:\Users\Ryan\AppData\Local\Temp\RtmpMRdU12/continuous_variables.csv
#> Exporting Summary to: C:\Users\Ryan\AppData\Local\Temp\RtmpMRdU12/summary.csvImport the exported marginal distributions using the
import_marginal_distributions() function
Synthesise data from the imported marginals using the
synthesise_data function, for multiple tables a named list
of data frames is returned.
# Synthesise the tables from the imported Marginal Distributions (without correlations)
sim_dfs <- synthesise_data(imported_marginals)
# Number of rows in each table
sapply(sim_dfs, nrow)
#> dm ae
#> 306 1191
# Number of subjects in each table
sapply(sim_dfs, function(df) length(unique(df$USUBJID)))
#> dm ae
#> 306 225The number of rows and subjects of each table are maintained, synthesised subjects are identified by a number, which links the subjects between the tables.
Fit the same logistic regression model as earlier except this time on the synthesised data.
ae_sim <- prepare_ae_data(sim_dfs$dm, sim_dfs$ae)
# Proportion of moderate or severe adverse events by treatment arm
prop.table(table(ae_sim$ARM, ae_sim$MOD_SEV), 1)
#>
#> FALSE TRUE
#> Placebo 0.6265664 0.3734336
#> Xanomeline High Dose 0.6234310 0.3765690
#> Xanomeline Low Dose 0.5913043 0.4086957
# Fit a logistic regression model on the synthesised data
glm.sim <- glm(
MOD_SEV ~ AGE + SEX + ARM,
data = ae_sim,
family = binomial
)
# Output the coefficients of the model
summary(glm.sim)$coefficients
#> Estimate Std. Error z value Pr(>|z|)
#> (Intercept) -0.630833056 0.601786974 -1.0482664 0.2945159
#> AGE 0.002193439 0.007856429 0.2791903 0.7800988
#> SEXM -0.108603407 0.131271564 -0.8273186 0.4080565
#> ARMXanomeline High Dose 0.017967552 0.169169777 0.1062102 0.9154156
#> ARMXanomeline Low Dose 0.147867959 0.151204267 0.9779351 0.3281064Without correlations the variables are synthesised independently, so there is no relationship between the treatment arm and the severity of the adverse events.
Synthesise data from the imported marginals with assumed
correlations, using the correlations parameter to specify a
list of correlations created with the correlation()
function. Categorical variables are correlated using a single category,
specified with factor_name.x or factor_name.y,
and the tables of the variables can be specified with
df_name.x and df_name.y.
# Synthesise the tables specifying assumed correlations
sim_dfs_cor <- synthesise_data(
imported_marginals,
correlations = list(
# Subjects on placebo are more likely to have mild adverse events
correlation(
"ARM",
"AESEV",
0.2,
df_name.x = "dm",
df_name.y = "ae",
factor_name.x = "Placebo",
factor_name.y = "MILD"
),
# Male subjects are younger
correlation("AGE", "SEX", -0.2, factor_name.y = "M")
)
)Correlated variables are synthesised together, one row per subject, and joined to each table by subject. Therefore correlations between tables are between subjects, and a correlated variable takes a single value for each subject within a table, in this case every adverse event of a subject has the same severity.
Using the synthesised data (with correlations) fit the same logistic regression model as earlier.
ae_sim_cor <- prepare_ae_data(sim_dfs_cor$dm, sim_dfs_cor$ae)
# Proportion of moderate or severe adverse events by treatment arm
prop.table(table(ae_sim_cor$ARM, ae_sim_cor$MOD_SEV), 1)
#>
#> FALSE TRUE
#> Placebo 0.7005988 0.2994012
#> Xanomeline High Dose 0.6027778 0.3972222
#> Xanomeline Low Dose 0.6548387 0.3451613
# Fit a logistic regression model on the synthesised data (with correlations)
glm.sim.cor <- glm(
MOD_SEV ~ AGE + SEX + ARM,
data = ae_sim_cor,
family = binomial
)
# Output the coefficients of the model
summary(glm.sim.cor)$coefficients
#> Estimate Std. Error z value Pr(>|z|)
#> (Intercept) -1.434536276 0.653056342 -2.1966501 0.028045445
#> AGE 0.006320768 0.008338343 0.7580365 0.448429162
#> SEXM 0.240313969 0.136579469 1.7595175 0.078489647
#> ARMXanomeline High Dose 0.417586363 0.161300094 2.5888786 0.009628903
#> ARMXanomeline Low Dose 0.244232900 0.170323799 1.4339329 0.151591410With the assumed correlation, adverse events of subjects on placebo are less likely to be moderate or severe, allowing the analysis to be tested before access to the original data is granted.
NB The correlation specified is that of the underlying multivariate normal distribution (Copula), the correlation observed in the synthesised data will differ, particularly for categorical variables. It is not possible to entirely maintain all the marginal distributions when specifying correlations.