Oncology is experiencing a growing interest in the use of external control (EC) data for the design and analysis of clinical trials [2,6]. Data from EC patients may be used in a variety of tasks, for example for the analysis of single-arm studies [5] or for early futility stopping of randomized trials [10]. The rationale for incorporating EC data into clinical trials is typically to increase the efficiency of the drug development process. For example, consider newly diagnosed glioblastoma (GBM) patients. Most Phase 2 trials in this population are standard single-arm trials. Typically in these studies the primary analyses directly compare the objective response rate or other outcomes to historical estimates that were previously published [5]. However, the populations may be different across studies, and these discrepancies can compromise the validity of the trial results. The use of single-arm trials early in the drug development process and its shortcomings have been associated in a series of articles with the high rate of confirmatory Phase 3 trial failures in GBM in the past two decades [1,9]. The slow progress of therapeutics in GBM and other cancers is a major motivation for methodological efforts to develop novel trial designs, from early stage to registration studies, that leverage individual patient level EC data, including pre-treatment clinical profiles, treatments, and outcomes [7].
It is well understood that integrating EC data into a trial comes with some risks of bias, potentially leading to false positive results and more general limitations in the scientific rigor. Possible distortion mechanisms include confounding, missing data, selection bias, and different standards to measure biomarkers and outcomes across institutions [3].
To understand how ECs have been used in clinical studies we surveyed ClinicalTrials.gov, the public trial reporting database of the US National Institutes of Health. Our survey included all interventional trials that began in the years 1996–2023, including both completed and ongoing trials. To identify studies we searched for the following key words: “external control”, “historic control”, “external comparator”, “synthetic control”, “historical comparator”, and “historical control.” For each of these search terms we examined all records, except for “historical control” where we examined a random sample of 126 of the 460 records. We excluded trials that: 1) did not leverage EC data, 2) did not evaluate a drug or biologic therapy, or 3) ended after enrolling fewer than 10 patients.
We notice an increasing number of clinical trials that leverage ECs (Fig. 1A). Our survey identified 130 trials using EC data in ClinicalTrials. gov. Cancer trials represented 31.5% of these studies. The top five cancer types were leukemia (22.0%), pancreatic cancer (9.8%), lung cancer (7.3%), liver cancer (7.3%), and neuro-oncology (7.3%). Among non-oncology trials, the most common disease areas were orthopedics (13.5%), cardiology (12.4%), infectious disease (11.2%), neurology (7.9%), and pediatrics (6.7%).
Fig. 1.

A) Number of trials using ECs in our survey of ClinicalTrials.gov that began in each year (total of N = 130). B) Percent of trials in the sample (out of total N = 130) using ECs in our survey by type of data source.
For most of these studies there was little to no publicly reported information on the EC data or statistical methods used to analyze the trial. Of the 130 trials in our survey, only 59.2% provided any description of the EC data source (e.g., if the EC data originate from past trials or if they are real world data such as electronic health records for administrative purposes; see the green and blue bars of Fig. 1B), only 23.8% mentioned a statistical strategy to account for potential confounding or other distortions (e.g. propensity score methods, or matching procedures), and only 10% provided study protocol or statistical analysis plan documents. This information is also unavailable or difficult to access from other sources beyond ClinicalTrials.gov, such as sponsors’ websites, press releases, and other public repositories.
We are convinced that more transparent, accessible, and detailed information on the EC data and statistical procedures would be beneficial in the future of clinical research for a wide range of reasons. From the sponsors’ point of view, better public information would help identify valuable datasets as well as datasets that are becoming obsolete (e.g. due to inadequate biomarker profiles). It would also reduce errors [8,4], such as planning single-arm studies when EC data are insufficient to infer treatment effects, or post-hoc “cherry picking” of the EC dataset among candidates, which can create biased and overly optimistic trial results. We note that cherry picking is not necessarily malicious; it can be the result of multiple analyses led at different times by multiple analysts. From a methodological point of view, better public information would also allow biostatisticians to reach consensus on good standards for EC data analysis. In addition, transparent and accessible information on the EC data and trial design makes patients, investors, and other trial stakeholders more confident about the rigor and scientific validity. Moreover, transparency is essential to identify disease-specific barriers to the integration of EC data, such as new biomarker-specific standard of care therapeutics.
To inform stakeholders and the scientific community, regulators could incentivize study sponsors to make public limited and potentially standardized information about the EC data and the analysis methods. We propose that when reporting to public databases, such as ClinicalTrials.gov in the United States and the Clinical Trials Information System in the European Union, sponsors include more information about the use of ECs, including:
rationale for using ECs (e.g., to shorten study duration, to alleviate recruitment needs for a rare disease, etc.),
description of data source (e.g., past clinical trials, specific patient registries, electronic health records),
any eligibility differences between the EC group and the trial population,
EC sample size,
the list of relevant pre-treatment covariates (available for both EC and trial patients), and measurement standards,
statistical method to incorporate EC data (e.g. propensity score matching, power prior, linear regression, none, etc.), and
planned sensitivity analyses to evaluate the robustness of the primary findings.
Acknowledgements
We gratefully acknowledge support from the National Institutes of Health under grants R01LM013352 (LT and DES) and T32CA009337 (DES).
Footnotes
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributor Information
Daniel Evan Schwartz, Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, United States; Department of Data Science, Dana-Farber Cancer Institute, Boston, MA, United States.
Hanna Essaouabi, Northeastern University, Boston, MA, United States.
Lorenzo Trippa, Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, United States; Department of Data Science, Dana-Farber Cancer Institute, Boston, MA, United States.
References
- [1].Bagley SJ, Kothari S, Rahman R, Lee EQ, Dunn GP, Galanis E, et al. Glioblastoma clinical trials: current landscape and opportunities for improvement. Clin Cancer Res 2022;28:594–602. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Carrigan G, Whipple S, Capra WB, Taylor MD, Brown JS, Lu M, et al. Using electronic health records to derive control arms for early phase single-arm lung cancer trials: proof-of-concept in randomized controlled trials. Clin Pharmacol Ther 2020; 107:369–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Friends of Cancer Research Whitepaper. Characterizing the use of external controls for augmenting randomized control arms and confirming benefit; 2019.
- [4].Ghadessi M, Tang R, Zhou J, Liu R, Wang C, Toyoizumi K, et al. A roadmap to using historical controls in clinical trials - by Drug Information Association Adaptive Design Scientific Working Group (DIA-ADSWG). Orphanet J. Rare Dis 2020;15:69. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Grossman SA, Schreck KC, Ballman K, Alexander B. Point/counterpoint: randomized versus single-arm phase II clinical trials for patients with newly diagnosed glioblastoma. Neuro-Oncology 2017;19:469–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Mishra-Kalyani PS, Amiri Kordestani L, Rivera DR, Singh H, Ibrahim A, DeClaro RA, et al. External control arms in oncology: current use and future directions. Ann Oncol 2022;33:376–83. [DOI] [PubMed] [Google Scholar]
- [7].Rahman R, Ventz S, McDunn J, Louv B, Reyes-Rivera I, Polley M-YC, et al. Leveraging external data in the design and analysis of clinical trials in neuro-oncology. Lancet Oncol 2021;22:e456–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].U.S. Food and Drug Administration. Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft Guidance; 2023. [Google Scholar]
- [9].Vanderbeek AM, Rahman R, Fell G, Ventz S, Chen T, Redd R, et al. The clinical trials landscape for glioblastoma: is it adequate to develop new treatments? Neuro-Oncology 2018;20:1034–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Ventz S, Lai A, Cloughesy TF, Wen PY, Trippa L, Alexander BM. Design and evaluation of an external control arm using prior clinical trials and real-world data. Clin Cancer Res 2019;25:4993–5001. [DOI] [PMC free article] [PubMed] [Google Scholar]
