Data and Code for “Cluster trials inference with CARE”
Replication package for Cluster Trials Inference with CARE This repository contains the data, code, and documentation supporting the manuscript Cluster Trials Inference with CARE by Sergey Alexeev and Rachael L. Morton. The package provides materials used to gene
Replication package for Cluster Trials Inference with CARE
This repository contains the data, code, and documentation supporting the manuscript Cluster Trials Inference with CARE by Sergey Alexeev and Rachael L. Morton.
The package provides materials used to generate the empirical examples and simulation evidence reported in the paper. It is designed to allow readers and reviewers to inspect the analysis workflow and reproduce the main results where the underlying data are publicly available.
Contents
The archive contains two main components:
1. Empirical examples
These reproduce the case studies discussed in the manuscript:
Case Study 1 – Ten Hoor et al. (2018)
Case Study 2 – Tannenbaum et al. (2013)
Case Study 3 – Mudge et al. (2022)
Case Study 4 – Kaaya et al. (2022)
For the Stata-based case studies, the package provides:
cleaned public-facing run scripts
corresponding log files documenting the outputs
structured folders separating
data,code,output, and supporting documentation
The empirical examples illustrate the CARE workflow (Clarify, Apply, Refine, Evaluate) for cluster-randomised trials.
Important note:
Case Study 3 cannot be independently rerun from this archive because the original dataset is not publicly available. The analysis outputs included in the package were generated in collaboration with the original study authors.
2. Simulation materials
The repository also includes the simulation code and outputs underlying the manuscript’s simulation study. These simulations examine the performance of inference methods for cluster-randomised trials under varying numbers of clusters and cluster-size imbalance.
Simulation components correspond to:
Figure 2
Figures 3–6
Figure 7
Where possible, simulation outputs are included so that readers can inspect results without rerunning computationally intensive simulations.
Package structure
The repository is organized into two main directories:
empirical examples/– case study analyses and replication scriptssimulation/– simulation scripts and outputs
Top-level manifest files (manifest_empirical_examples.csv and manifest_simulations.csv) map repository files to the corresponding tables and figures in the manuscript.
Software requirements
The empirical analyses were implemented primarily in Stata, with additional simulations implemented in R. The i
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.