Observational research promises to add evidence where randomized trials are underpowered, unethical, unrealistic, or—most often—simply too resource intensive to meet the huge demand. Clinical guidelines that are intended to be evidence-based are forced to rely on expert opinion,1 and of the tens of thousands of side effects that could result from each of the thousands of drugs on the market, only a fraction of these potential side effects has been formally studied. The emergence of enormous databases, increasing computing power, and new analytic methods should propel observational research forward, but researchers and the public have become increasingly aware of the unreliability of published real-world evidence.2,3
For observational research to reach its potential, we must recognize where the challenges are and take concrete steps to address them. We focus here on observational research that tests clinical hypotheses with an intent to publish in top clinical journals and affect the care of millions of patients. We recognize that other observational research is hypothesis generating and may be treated differently.
Observational research can be roughly divided into several components: picking an appropriate and well-formulated hypothesis, developing a proper study design, accessing data that are relevant and accurate, executing the study rigorously, and interpreting the results correctly. Several initiatives have identified concrete steps we can take and criteria we can use to improve reliability. For example, the Observational Health Data Sciences and Informatics (OHDSI) Large-scale Evidence Generation and Evaluation across a Network of Databases (LEGEND) framework4 identifies 10 criteria for improving the reliability of observational research and was used in published cardiovascular safety and effectiveness studies.5 The criteria can be grouped as openness and verification, shown in Figure 1. Openness is well-known but poorly followed. The study protocol must be prespecified and published, all study software and clinical definitions must be made available, all diagnostics must be shared before unblinding the results, and all results must be published in some format. Verification is through formal diagnostics that assess every component of the process, including data quality, accuracy of the definitions of outcomes and key covariates, analytic diagnostics such as achieved balance in confounding adjustment, a test for consistency among multiple databases, and large-scale negative control subjects to substantiate the claim of minimal bias. The Sentinel Initiative provides an overlapping set of system requirements, focusing on openness, data properties, and standardized tools.6 It has been applied extensively in the work of the U.S. Food and Drug Administration. The RCT-DUPLICATE (Randomized Controlled Trials Duplicated Using Prospective Longitudinal Insurance Claims: Applying Techniques of Epidemiology) initiative7 overlaps the other 2, focusing on study design. Study design is critical to reliability, including specifying the population, the intervention, the comparator, and the causal contrast, and employing an analytic approach appropriate to causal inference. The RCT-DUPLICATE team adopted target trial emulation,8 in which a hypothetical randomized trial is designed first and then the observational trial is built to emulate the randomized trial, avoiding biases that commonly occur in observational designs.
FIGURE 1. OHDSI LEGEND Principles for Reliable Observational Research.

The Observational Health Data Sciences and Informatics (OHDSI) Large-scale Evidence Generation and Evaluation across a Network of Databases (LEGEND) principles for reliable research can be summarized in 2 main themes, openness and verification, which are applied through the course of a study. In this piece, we focus on openness as being sufficiently mature for adoption in hypothesis-testing observational research, while progress is being made on the concrete criteria for verification.
Of the various criteria cited, some are straightforward and, we argue, should be adopted immediately. Openness has relatively little cost and is well understood. Academic journals are in an ideal position to enforce its adoption. Researchers today pull from a multiplicity of paths to answer a question, so that virtually any answer is possible and preconceived notions of truth are easily generated,2 especially if the researcher iterates on the design. This is compounded by publication bias, in which 81% of published observational studies are positive.3 The result can be chance signals and biased results. Academic incentives compound the problem as researchers strive to advance their careers through obtaining results that are most likely to be published and as journals strive for success with appealing topics that are new rather than confirmatory or expected. Openness adopted today would go far in addressing multiplicity, publication bias, and counterproductive incentives. Specifically, we suggest a prespecified, publicly available study protocol; a discussion of the appropriateness of the study design, potentially using target trial emulation as a model for the description; open-source software or commercially available statistical software with release of all locally developed scripts; detailed assessment of the data and populations used; computable description of the interventions, outcomes, other covariates, and study procedures; publicly available diagnostics with results initially blinded; and all results passing diagnostics made publicly available in some form.
In parallel, we can continue to work on empirical verification. Verification is obvious as a concept, but the exact choice of analytic methods for a given hypothesis and the choice of how to verify them remain subject to ongoing discussion and research. We believe that the use of multiple databases—looking for differences among databases as evidence of systematic error that should be addressed—is already warranted for high-impact studies that may affect the care of millions of patients. Although it may be premature to require the following at the editorial level, LEGEND also employs the following components9: data quality assessment; accuracy of the definitions of interventions, outcomes, and other covariates; success of the statistical procedures in addressing sources of bias like confounding and censoring and in achieving sufficient power; and detection of residual bias using methods like negative control subjects. Research on the best methods to control bias caused by sources like confounding and censoring should continue. We favor the use of large-scale methods for creating balanced samples to address confounding10 and the use of negative-control calibration to account for potential residual bias in study results.3
Observational research promises to add evidence where randomized trials are impractical, unethical, or underpowered, and it offers a glimpse of real-world practice outside of strict research protocols. Given a successful cultural shift to support standards for design, diagnostics, and transparency, we believe that observational research can augment what we can practically do with randomized trials and provide evidence where opinion reigns today.
FUNDING SUPPORT AND AUTHOR DISCLOSURES
This study was partially supported by National Institutes of Health grant R01 LM006910. Drs Hripcsak and Suchard have received grant funding from Johnson and Johnson through their universities to support methods research. Drs Schuemie and Ryan are employees of Johnson and Johnson. Johnson and Johnson did not have input in the design, execution, interpretation of results, or decision to publish.
Footnotes
The authors attest they are in compliance with human studies committees and animal welfare regulations of the authors’ institutions and Food and Drug Administration guidelines, including patient consent where appropriate. For more information, visit the Author Center.
REFERENCES
- 1.Fanaroff AC, Califf RM, Windecker S, Smith SC, Lopez RD. Levels of evidence supporting American College of Cardiology/American Heart Association and European Society of Cardiology guidelines, 2008–2018. JAMA. 2019;321:1069. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Hoffmann S, Schönbrodt F, Elsas R, Wilson R, Strasser U, Boulesteix A-L. The multiplicity of analysis strategies jeopardizes replicability: lessons learned across disciplines. R Soc Open Sci. 2021;8:201925. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Schuemie MJ, Ryan PB, Hripcsak G, Madigan D, Suchard MA. Improving reproducibility by using high-throughput observational studies with empirical calibration. Philosophical Transactions A. 2018;376:20170356. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Schuemie MJ, Ryan PB, Pratt N, et al. Principles of large-scale evidence generation and evaluation across a network of databases (LEGEND). J Am Med Inform Assoc. 2020;27:1331–1337. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Khera R, Aminorroaya A, Dhingra LS, et al. Comparative effectiveness of second-line anti-hyperglycemic agents for cardiovascular outcomes: a multinational, federated analysis of LEGEND-T2DM. J Am Coll Cardiol. 2024;84(10): 904–917. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Brown JS, Maro JC, Nguyen M, Ball R. Using and improving distributed data networks to generate actionable evidence: the case of real-world outcomes in the Food and Drug Administration’s Sentinel system. J Am Med Inform Assoc. 2020;27(5):793–797. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Heyard R, Held L, Schneeweiss S, Wang SV. Design differences and variation in results between randomised trials and non-randomised emulations: meta-analysis of RCT-DUPLICATE data. BMJ Med. 2024;3(1):e000709. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183:758–764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Conover MM, Ryan PB, Chen Y, Suchard MA, Hripcsak G, Schuemie MJ. Objective study validity diagnostics: a framework requiring pre-specified, empirical verification to increase trust in the reliability of real-world evidence. J Am Med Inform Assoc. 2025;32(3):518–525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zhang L, Wang Y, Schuemie M, Blei D, Hripcsak G. Adjusting for indirectly measured confounding using large-scale propensity score. J Biomed Inform. 2022;134:104204. [DOI] [PMC free article] [PubMed] [Google Scholar]
