Skip to main content
Lippincott Open Access logoLink to Lippincott Open Access
. 2023 May 31;34(5):619–623. doi: 10.1097/EDE.0000000000001637

Start with the Target Trial Protocol, Then Follow the Roadmap for Causal Inference

Lauren E Dang a,, Laura B Balzer a
PMCID: PMC10392882  PMID: 37255259

Before taking flight, pilots complete a preflight checklist. Before starting surgery, operating room staff complete a presurgical checklist. These checklists synthesize the knowledge of a field into a series of steps to be considered every time a job is performed. What about when researchers want to infer causality?

A century’s worth of literature regarding the design and analysis of observational studies is available, yet common guidance is not always followed. Bykov et al.1 reviewed a random sample of 75 observational studies on cardiovascular disease, diabetes, or osteoporosis and found that 95% “had at least one avoidable methodological issue known to incur bias.” It is, thus, critical to distill this literature into concrete instructions so that best practices are consistently implemented.

Many useful frameworks operationalize aspects of the causal and statistical inference literature into steps for researchers to follow.212 Target trial emulation is a popular approach that has led to high-profile studies in recent years.1315 The basic structure shared by target trial emulation studies follows Hernán and Robins’7 statement that “at the very least, we need to specify the following key components of the [target trial] protocol: eligibility criteria, start and end of follow-up, treatment strategies, outcomes of interest, causal contrast, and data analysis plan.”16 The assignment procedure of the target trial is generally included, as well.7

Pearce and Vandenbroucke17 raised concerns about certain studies that use target trial emulation protocol components. First, they suggest that target trial emulation seems to focus on mimicking a conditionally randomized controlled trial, as opposed to considering a variety of study designs and identification assumptions. Second, target trial emulation is sometimes paired with a matching approach for statistical estimation; matching can change the target population and is often less efficient than other methods.18,19 Third, after emulating a target trial, one may be tempted to over-interpret the results; we still have an observational study, subject to multiple potential biases, including intractable confounding. Ultimately, Pearce and Vandenbroucke17 conclude that target trial emulation is not the optimal starting point for causal inference.

In contrast, we believe the target trial emulation protocol is a useful starting place for analyses seeking to infer causality. In other words, we agree with Hernán and Robins’7 statement: “specifying the protocol of the target trial is a useful device to clarify the causal question of interest.”16 However, we disagree with their assertion: “once the causal question is stated with little ambiguity, study design and data analysis flow naturally.”16 Instead, a given research question can lead us down many potential paths, including alternative study designs, identification assumptions, estimation approaches, and interpretations. What guidance should investigators follow after they have defined their causal question?

The Roadmap for Causal and Statistical Inference8,9,2024 (originally described in Petersen and van der Laan9 and hereafter, the Roadmap) provides step-by-step directions to aid in the design, analysis, and interpretation of studies—randomized or observational. The Roadmap

  1. can accommodate any study design and all identification assumptions,

  2. emphasizes the importance of selecting the best statistical estimator for a given problem based on theoretical and finite sample properties, and

  3. includes steps to explicitly protect against over-interpretation, while providing a path forward when we cannot estimate a causal effect.

Below, we describe how the Roadmap (summarized in the Table) accomplishes these goals.

TABLE.

The Steps of the Roadmap8,9,20-24* with Example Questions to Ask

The Roadmap for Causal and Statistical Inference
1. Specify the research question, including the target population, exposure(s), outcome(s), time period, and context of interest.
▪ What do we want to learn from data that have been collected or will be collected? To whom do we want to apply the results?
2. Specify a causal model, such as a directed acyclic graph,10 to describe relationships between variables.
▪ Are there unmeasured confounders or time-dependent confounders? Are the outcomes missing or censored?
3. Define the causal parameter of interest with counterfactual outcomes.
▪ What hypothetical change to the causal model, even if impossible, would we make to generate counterfactuals and answer our research question? How do we want to summarize the distributions of counterfactual outcomes?
4. Describe the observed data and the statistical model.
▪ What data did we or will we actually observe? Are there functional form assumptions or can the relationship between the outcome(s), exposure(s), and adjustment variables take any form?
5. Assess identifiability.
▪ What modifications can we make to reduce the causal gap?
6. Define the statistical parameter.
▪ What function of the observed data are we aiming to estimate?
7. Choose and implement a statistical estimator based on statistical properties; obtain 95% confidence intervals.
▪ What are the theoretical properties of potential estimators (e.g., robustness)? How do they perform in finite sample simulations according to objective criteria?
8. Conduct sensitivity analyses.
▪ What can existing evidence tell us about plausible magnitudes of the causal gap?
9. Interpret the results, accounting for the prior steps.
▪ Have we estimated an association or a causal effect? What are the real-world implications?

*The order in which the steps are presented differs in different versions of the Roadmap8,9,2024 because these steps inform one another and are generally specified through an iterative process.

FIRST, CHOOSE YOUR IDEAL DESTINATION

Both the Roadmap (Step 1) and target trial emulation prompt us to specify the question that we would ideally like to answer. A strength of target trial emulation is its use of language familiar to researchers from a wide variety of backgrounds. This step tells us our ideal destination. Given the myriad options for study design, analysis, and interpretation, how should we plot a course from our question to an answer with real-world relevance? How will we evaluate whether we landed on the island of causality, and what do we do if we find ourselves in the sea of statistical associations? The remainder of the Roadmap’s steps help us to answer these questions.

“MIND THE [CAUSAL] GAP!”

Step 2 of the Roadmap is to specify a causal model, such as a directed acyclic graph,10,25 to describe existing knowledge and uncertainty about the relationships between the covariates (baseline or time-varying), exposures of interest, censoring, and outcomes. By intervening on this causal model, we then generate counterfactual outcomes corresponding to our ideal, but often impossible, experiment (Step 3). After describing which variables are actually measured (Step 4), the Roadmap’s identification step requires us to transparently state and critically evaluate the assumptions needed to infer causality (Step 5).

Most commonly, these assumptions include no unmeasured confounding (i.e., conditional exchangeability), missing/censoring at random, and sufficient data support (i.e., positivity). The Roadmap also accommodates alternative approaches such as difference-in-differences, whose alternate set of identification assumptions must be explicitly discussed; see, for example, Weber et al.26 Although Hernán and Robins’7 cover causal models and various identification assumptions at length in their book Causal Inference: What If,16 the “assignment procedures” protocol component of the target trial emulation focuses on conditional exchangeability. In contrast, the Roadmap explicitly requires researchers to delineate all necessary identification assumptions every time they aim to estimate a causal effect. Doing so can minimize the potential for severely misleading inferences arising from violations of assumptions related to missingness or data support, among others.2730 Additionally, by transparently stating the identification assumptions (Step 5) and including a causal graph (Step 2), the Roadmap empowers subject-matter experts to evaluate whether these assumptions are reasonable.

TRY TO BRIDGE THE GAP

Without identification, the true value of the statistical parameter that we will estimate from the observed data will not be equivalent to our wished-for causal parameter; in other words, there will be a so-called causal gap.21 When working with observational data, the Roadmap’s identification step generally leads us to conclude that we are only able to estimate an association, rather than the desired causal effect. What should we do when our causal gap is nonzero? Rudolph and Keyes31 review approaches to improve the plausibility of the identification assumptions, including changing the research question, definition of the exposure or outcome, or target population.7,9,22,27,28 Roadmap Steps 1–5 should be viewed as an iterative process.

RED LIGHT [CHOOSE AN ESTIMATOR]; GREEN LIGHT [ESTIMATE!]

After taking these steps, we are still usually stuck with a causal gap and the need to provide timely and actionable answers to our research questions. The Roadmap’s Step 6 and the “analysis plan” protocol component of target trial emulation guide us to choose a statistical parameter that is motivated by our underlying question.79,16 This parameter is often a complex function of the observed data distribution—even in trials due to censoring or noncompliance—and is rarely equal to a coefficient in a regression model.22,32,33

An additional benefit of the Roadmap is that it guides us to choose an estimation approach based on statistical properties (Step 7). First, the estimator should have a low bias and variance relative to available choices.8 Not all target trial emulation studies use matching, but the Roadmap would generally lead away from matching due to inefficiency and the risk of changing the target population. Second, the estimator should not rely on unrealistic functional form assumptions; machine learning may help minimize bias due to reliance on misspecified parametric regressions, but must be incorporated in a manner that yields valid 95% confidence intervals.8,34 For these reasons, the Roadmap often leads us to doubly robust estimators, such as the targeted maximum likelihood estimator.8,35

Of course, there are many possible implementations of a given approach (e.g., targeted maximum likelihood estimator with different Super Learner libraries or sample-splitting schemes). Candidate algorithms should be compared using finite sample simulations reflecting the challenges in the real data. See, for example, Montoya et al.36, who based their simulation on the empirical confounder distribution but with blinded outcomes. The process of using simulations to select an estimation approach and its precise implementation, according to prespecified criteria, is integral when developing the statistical analysis plan. Afterward, we have the green light to estimate!

LOOK BEFORE YOU LEAP

Once we have a point estimate and confidence intervals, how should we interpret our results? As discussed above, if the causal gap is nonzero, then we must be transparent that we have estimated a statistical association, not a causal effect. Yet just as there is a spectrum from a perfectly conducted randomized trial to a messy observational study, there is a spectrum of possible interpretations, each with value to science.9

In their inaugural Roadmap paper, Petersen and van der Laan9 provide a hierarchy of interpretations, ranging from purely statistical to replicating the results of a hypothetical trial. Importantly, they highlight that increasing the strength of our interpretations requires increasing the strength of our assumptions. Specifically, concluding that we have emulated a randomized trial requires (1) the statistical estimator has negligible bias and its variance is well-estimated, (2) the causal gap is zero, (3) the intervention is feasible and applicable to a real-world population, and (4) the intervention could have been randomized to that population.9 Roadmap Steps 1–7 help us to select and implement an estimator that supports valid statistical inference (#1) and to evaluate whether there is a causal gap (#2). The goal of replicating the results of a randomized trial (#3 and 4) may or may not be desirable, depending on the context.

Like a real map, the Roadmap has evolved over time and now includes sensitivity analyses (Step 8) to guide interpretation when the causal gap is not zero.21 We emphasize that the difference in adjusted and unadjusted estimates does not reveal whether the adjusted approach has accounted for all confounders.10 To assess potential violations of untestable causal identification assumptions (e.g., residual confounding), we can consider available evidence regarding the plausible magnitude and direction of the causal gap.37,38 Such sensitivity analyses are also applicable in studies using matching.

Completing Roadmap Steps 1–8 helps us to interpret our results appropriately (Step 9). First, the plausible causal gap could be small enough to allow for a close approximation of a causal effect. See, for example, Skeem et al.’s39 study of alternate probation protocols for people with mental illness. Second, the confidence interval bounds could be far enough from the null relative to the plausible magnitude of the causal gap that researchers may conclude with reasonable certainty that there is a causal effect. Examples include Cornfield et al.’s40 assessment that smoking causes lung cancer and Diaz et al.’s37 assessment that Nifurtimox is effective in treating Chagas disease.

In many scenarios, however, we cannot conclude that we have estimated a causal effect. Nonetheless, the estimated association may still provide useful real-world information.31,41,42 For example, even if we cannot prove that state-level mask mandates caused lower mortality rates from COVID-19, knowledge that the mortality rate, adjusted for measured state-level confounders, was lower with early versus delayed implementation of mask mandates is still valuable.42 Altogether, the Roadmap steps of specifying a causal model, discussing the plausibility of identification assumptions, conducting sensitivity analyses, and ensuring that the interpretation is consistent with these prior steps help to generate estimates with real-world implications.

CONCLUSION

No single framework can substitute for the full value of the literature on causal and statistical inference. Yet just as checklists have improved the quality of services in aviation and medicine,43,44 we believe in the power of the Roadmap to improve the design and analysis of observational studies seeking to infer causality. Of course, any tool could be misinterpreted or instill false confidence in some users. Yet, the Roadmap supports structured choices that researchers are already making, including picking a study design, selecting adjustment sets, choosing and implementing an estimator, and deciding how to interpret results. While the Roadmap cannot prevent all missteps, it makes the reasoning behind these necessary choices more transparent, empowering readers of Roadmap papers to make informed interpretations of the study results in both randomized trials and observational studies.9,26,29,30

ABOUT THE AUTHORS

Lauren Eyler Dang, MD, PhD recently completed her PhD in Biostatistics at the University of California, Berkeley. She will soon join the National Institute of Allergy and Infectious Diseases Biostatistics Research Branch as a Mathematical Statistician. Laura B. Balzer, PhD, is an Associate Professor of Biostatistics at the University of California, Berkeley. Both Lauren and Laura love developing, evaluating, applying, and teaching methods for causal inference.

Footnotes

Supported, in part, by a philanthropic gift from the Novo Nordisk corporation to the University of California, Berkeley to support the Joint Initiative for Causal Inference. The funders had no role in the conceptualization or writing of the manuscript.

L.E.D. reports tuition and stipend support from a philanthropic gift from the Novo Nordisk corporation to the University of California, Berkeley to support the Joint Initiative for Causal Inference. The other author has no conflicts to report.

REFERENCES

  • 1.Bykov K, Patorno E, D’Andrea E, et al. Prevalence of avoidable and bias-inflicting methodological pitfalls in real-world studies of medication safety and effectiveness. Clin Pharmacol Ther. 2022;111:209–217. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Neyman J. Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes. English translation by D.M. Dabrowska and T.P. Speed (1990). Stat Sci. 1923;5:465–480. [Google Scholar]
  • 3.Rubin DB. Estimating causal effects of treatments in randomized and nonrandomized studies. J Educ Psychol. 1974;66:688–701. [Google Scholar]
  • 4.Patient-Centered Outcomes Research Institute. PCORI Methodology Standards. Available at: https://www.pcori.org/research/about-our-research/research-methodology/pcori-methodology-standards. Accessed March 22, 2023. [Google Scholar]
  • 5.Gatto NM, Reynolds RF, Campbell UB. A structured preapproval and postapproval comparative study design framework to generate valid and transparent real-world evidence for regulatory decisions. Clin Pharmacol Ther. 2019;106:103–115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Wang SV, Pinheiro S, Hua W, et al. STaRT-RWE: structured template for planning and reporting on the implementation of real world evidence studies. BMJ. 2021;372:m4856. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183:758–764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.van der Laan MJ, Rose S. Targeted Learning: Causal Inference for Observational and Experimental Data. Springer; 2011. [Google Scholar]
  • 9.Petersen ML, van der Laan MJ. Causal models and learning from data: integrating causal modeling and statistical estimation. Epidemiology. 2014;25:418–426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Pearl J. Causality: Models, Reasoning, and Inference. 2nd ed. Cambridge University Press; 2009. [Google Scholar]
  • 11.Richardson T, Robins J. Single World Intervention Graphs (SWIGs): A Unification of the Counterfactual and Graphical Approaches to Causality. University of Washington Center for Statistics and the Social Sciences; 2013. Available at: https://csss.uw.edu/Papers/wp128.pdf [Google Scholar]
  • 12.Rosenbaum PR. Observational Studies. 2nd ed. Springer; 2002. [Google Scholar]
  • 13.Ioannou GN, Locke ER, O’Hare AM, et al. COVID-19 vaccination effectiveness against infection or death in a national U.S. health care system: a target trial emulation study. Ann Intern Med. 2022;175:352–361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Barda N, Dagan N, Cohen C, et al. Effectiveness of a third dose of the BNT162b2 mRNA COVID-19 vaccine for preventing severe outcomes in Israel: an observational study. Lancet. 2021;398:2093–2100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Monge S, Rojas-Benedicto A, Olmedo C, et al. ; IBERCovid. Effectiveness of mRNA vaccine boosters against infection with the SARS-CoV-2 omicron (B.1.1.529) variant in Spain: a nationwide cohort study. Lancet Infect Dis. 2022;22:1313–1320. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Hernan MA, Robins JM. Causal Inference: What If. Boca Raton: Chapman & Hall/CRC; 2020. [Google Scholar]
  • 17.Pearce N, Vandenbroucke J. Are target trial emulations the gold standard for observational studies? Epidemiology. 2023;34. [DOI] [PubMed] [Google Scholar]
  • 18.Kupper LL, Karon JM, Kleinbaum DG, Morgenstern H, Lewis DK. Matching in epidemiologic studies: validity and efficiency considerations. Biometrics. 1981;37:271–291. [PubMed] [Google Scholar]
  • 19.Shiba K, Kawahara T. Using propensity scores for causal inference: pitfalls and tips. J Epidemiol. 2021;31:457–463. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Balzer L, Petersen M, van der Laan M. Tutorial for Causal Inference. In: Buhlmann P, Drineas P, Kane M, van der Laan M, eds. Handbook of Big Data. Chapman & Hall/CRC Press; 2016:361–386. [Google Scholar]
  • 21.Gruber S, Phillips RV, Lee H, Ho M, Concato J, van der Laan MJ. Targeted learning: Towards a future informed by real-world evidence. Stat Biopharm Res. 2023:1–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Petersen ML. Commentary: applying a causal road map in settings with time-dependent confounding. Epidemiology. 2014;25:898–901. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Tran L, Yiannoutsos CT, Musick BS, et al. Evaluating the impact of a HIV low-risk express care task-shifting program: a case study of the targeted learning roadmap. Epidemiol Methods. 2016;5:69–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Saddiki H, Balzer LB. A primer on causality in data science. J Soc Fr Statistique. 2020;161:67–90. [Google Scholar]
  • 25.Pearl J. Causal diagrams for empirical research. Biometrika. 1995;82:702669–702710. [Google Scholar]
  • 26.Weber AM, van der Laan MJ, Petersen ML. Assumption Trade-Offs when choosing identification strategies for pre-post treatment effect estimation: an illustration of a community-based intervention in Madagascar. J Causal Inference. 2015;3:109–130. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Petersen ML, Porter KE, Gruber S, Wang Y, van der Laan MJ. Diagnosing and responding to violations in the positivity assumption. Stat Methods Med Res. 2012;21:31–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Rudolph KE, Gimbrone C, Matthay EC, et al. When effects cannot be estimated: redefining estimands to understand the effects of naloxone access laws. Epidemiology. 2022;33:689–698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Balzer LB, Ayieko J, Kwarisiima D, et al. Far from MCAR: obtaining population-level estimates of HIV viral suppression. Epidemiology. 2020;31:620–627. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Nugent JR, Marquez C, Charlebois ED, Abbott R, Balzer LB. Blurring cluster randomized trials and observational studies using Two-Stage TMLE to address sub-sampling, missingness, and minimal independent units. 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Rudolph KE, Keyes KM. Voluntary firearm divestment and suicide risk: real-world importance in the absence of causal identification. Epidemiology. 2023;34:107–110. [DOI] [PubMed] [Google Scholar]
  • 32.Robins J. A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Math Model. 1986;7:1393–1512. [Google Scholar]
  • 33.Robins JM, Hernán MA. Estimation of the causal effects of time-varying exposures. In: Fitzmaurice G, Davidian M, Verbeke G, Molenberghs G, eds. Longitudinal Data Analysis. Chapman & Hall/CRC Press; 2009:553–597. [Google Scholar]
  • 34.Balzer LB, Petersen ML. Invited commentary: machine learning in causal inference—how do I love thee? let me count the ways. Am J Epidemiol. 2021;190:1483–1487. [DOI] [PubMed] [Google Scholar]
  • 35.van der Laan MJ, Rubin D. Targeted maximum likelihood learning. [published online ahead of print December 28, 2006]. Int J Biostat. 2006;2. doi: 10.2202/1557-4679.1043. [Google Scholar]
  • 36.Montoya LM, Kosorok MR, Geng EH, Schwab J, Odeny TA, Petersen ML. Efficient and robust approaches for analysis of sequential multiple assignment randomized trials: Illustration using the ADAPT-R trial. [published online ahead of print December 9, 2022]. Biometrics. doi: 10.1111/biom.13808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Díaz I, van der Laan MJ. Sensitivity analysis for causal inference under unmeasured confounding and measurement error problems. Int J Biostat. 2013;9:149–160. [DOI] [PubMed] [Google Scholar]
  • 38.Lash TL, Fox MP, MacLehose RF, Maldonado G, McCandless LC, Greenland S. Good practices for quantitative bias analysis. Int J Epidemiol. 2014;43:1969–1985. [DOI] [PubMed] [Google Scholar]
  • 39.Skeem JL, Manchak S, Montoya L. Comparing public safety outcomes for traditional probation vs specialty mental health probation. JAMA Psychiatry. 2017;74:942942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Cornfield J, Haenszel W, Hammond EC, Lilienfeld A, Shimkin M, Wynder E. Smoking and lung cancer: recent evidence and a discussion of some questions. J Natl Cancer Inst. 1959;22:173–203. [PubMed] [Google Scholar]
  • 41.Swanson SA, Studdert DM, Zhang Y, Prince L, Miller M. Handgun divestment and risk of suicide. Epidemiology. 2023;34:99–106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wong AK, Balzer LB. State-Level masking mandates and COVID-19 outcomes in the United States: a demonstration of the causal roadmap. Epidemiology. 2022;33:228–236. [DOI] [PubMed] [Google Scholar]
  • 43.Ross P. Human factors issues of the aircraft checklist. JAAER. 2004;13:4. [Google Scholar]
  • 44.Anwer M, Manzoor S, Muneer N, Qureshi S. WHO surgical safety checklist, compliance and its effectiveness: a JPMC audit. Pak J Med Sci. 2016;32:831–835. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Epidemiology (Cambridge, Mass.) are provided here courtesy of Wolters Kluwer Health

RESOURCES