Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2022 Dec 23;42(4):487–516. doi: 10.1002/sim.9628

Dealing with confounding in observational studies: A scoping review of methods evaluated in simulation studies with single‐point exposure

Anita Natalia Varga 1,✉, Alejandra Elizabeth Guevara Morel 1, Joran Lokkerbol 2, Johanna Maria van Dongen 1, Maurits Willem van Tulder 1,3, Judith Ekkina Bosmans 1
PMCID: PMC10107671  PMID: 36562408

Abstract

The aim of this article was to perform a scoping review of methods available for dealing with confounding when analyzing the effect of health care treatments with single‐point exposure in observational data. We aim to provide an overview of methods and their performance assessed by simulation studies indexed in PubMed. We searched PubMed for simulation studies published until January 2021. Our search was restricted to studies evaluating binary treatments and binary and/or continuous outcomes. Information was extracted on the methods' assumptions, performance, and technical properties. Of 28,548 identified references, 127 studies were eligible for inclusion. Of them, 84 assessed 14 different methods (ie, groups of estimators that share assumptions and implementation) for dealing with measured confounding, and 43 assessed 10 different methods for dealing with unmeasured confounding. Results suggest that there are large differences in performance between methods and that the performance of a specific method is highly dependent on the estimator. Furthermore, the methods' assumptions regarding the specific data features also substantially influence the methods' performance. Finally, the methods result in different estimands (ie, target of inference), which can even vary within methods. In conclusion, when choosing a method to adjust for measured or unmeasured confounding it is important to choose the most appropriate estimand, while considering the population of interest, data structure, and whether the plausibility of the methods' required assumptions hold.

Keywords: bias, causal inference, confounding, observational study, simulation, treatment effect

1. INTRODUCTION

The increasing availability of routinely collected data (eg, through patient registries and insurance claims) has led to a growing interest in using observational data for treatment effect estimation. 1 However, an important disadvantage of routinely collected data is that treatment is not randomly allocated, resulting in an increased risk of the results being biased by confounding. Confounding occurs when a variable, also known as a confounding variable or confounder, is correlated with both the treatment and the outcome. Confounding is an important issue when evaluating treatment effects as it may bias estimates so that they appear greater or smaller than their true magnitude. Consequently, confounding may lead to wrong conclusions regarding a treatment's impact on the outcome of interest. 1

Depending on whether the confounding variable is observed (eg, disease severity is registered in a patient record) or unobserved (eg, motivation is not registered in a patient record), confounding can be classified as measured confounding or unmeasured confounding, respectively. If possible, confounding should be prevented. However, in real‐world data this is generally not possible meaning that confounding needs to be dealt within the analysis. Which approach is most appropriate for handling confounding is highly dependent on the type of confounding the researcher is facing.

Many methods exist for dealing with measured and unmeasured confounding, but technical properties and the circumstances under which they perform well differ. The aim of this paper was to perform a scoping review of simulation studies evaluating methods available for dealing with confounding when analyzing the effect of health care treatments with a single‐point exposure in observational data. We aim to provide an overview of the methods and their performance based on simulation studies indexed in PubMed. Simulation studies are key to understanding the technical properties of statistical methods, but few studies have systematically summarized such information about methods available for dealing with confounding.

2. METHODS

2.1. Search strategy

PubMed was searched for studies published from inception to January 2021 using search terms related to observational data (eg, nonrandomized, real‐world data, database) and confounding (eg, confounding by indication, omitted variable bias). Our search was restricted to PubMed, since we were primarily interested in methods for biomedical research. Additionally, the search was limited to simulation studies with binary treatment and binary and/or continuous outcomes. The search strategy was developed in close cooperation with an experienced information specialist of the Vrije Universiteit Amsterdam. The electronic search was supplemented by searching reference lists of retrieved full texts and searching the authors' databases. A logbook was kept with the date, keywords, and search results. An online electronic library (Rayyan) was used for storing titles and abstracts. The studies were managed using Endnote X8.

2.2. Study selection

As we were primarily interested in the relative performance of the methods under specific circumstances (eg, rare outcomes, skewed data distribution, failing assumptions), our review was restricted to simulation studies. Simulation studies allow for the careful observation of a method's behavior in relation to the underlying data generation mechanism, whereas in empirical case studies using real data, the “true” value is unknown. Simulation studies that did not test a specific method's performance in dealing with confounding were excluded.

In biomedical research, three types of outcomes are mainly used: continuous, binary, and time‐to‐event outcomes. 2 Time‐to‐event outcome analyses were excluded because they present additional methodological challenges, such as the time‐dependent covariate distribution, which is beyond the scope of our study. We excluded studies that evaluated non‐binary treatments, because they are less frequently conducted and require different methods. Studies were also excluded if the number of individuals exceeded the number of covariates (eg studies with genetic data, neuroimaging, and high dimensional confounding). Studies dealing with time‐dependent confounding were also excluded because Daniel et al have already provided extensive guidance on this topic. 3 Reviews, tutorials, and guidelines were excluded if they did not include a quantitative assessment of the evaluated methods' performance using simulated data. For the sake of generalizability and conciseness, in‐depth investigations of specific scenarios (eg, comparing different estimands resulting from the same method, multilevel data) were also excluded.

2.3. Data extraction

Information was extracted from the included studies by one author (ANV) using a structured data extraction form. Data were extracted about the methods and performance measures (ie, bias, precision, empirical coverage probability). Furthermore, we extracted in‐depth information on the methods' strengths and limitations, and the authors' recommendations and remarks related to each method. To optimally support researchers in choosing the most suitable method for their specific situation, we also extracted information regarding the different methods' assumptions, the methods' performance, and the estimands (ie, the outcome of interest) generated by the methods. The different kinds of estimands are described in below Box*. Technical remarks were collected to summarize the methods' performance under specific circumstances (eg, rare treatment or outcome, or varying strength of confounding). A second author (AEGM) randomly checked five percent of the extracted data. Possible inconsistencies were discussed, all of which were settled during face‐to‐face meetings. As there were few inconsistencies, we did not expand the data check.

sim9628-gra-1000-b

3. RESULTS

We identified 28 548 potential studies via the electronic search and 55 through other sources. After removing 104 duplicates, 28 499 studies were screened based on title and abstract resulting in 28 091 exclusions. The most common reason for exclusion was that a study did not evaluate a method's performance (see Figure 1 for other exclusion reasons). Eventually, 127 studies were included, of which the majority were published during the last decade (see Figure 2).

FIGURE 1.

SIM-9628-FIG-0001-b

PRISMA flow chart of identification, screening, eligibility and inclusion of studies

FIGURE 2.

SIM-9628-FIG-0002-b

Included simulation studies per year of publication

3.1. Identified methods for dealing with measured confounding

Of the 127 included studies, 84 (66%) dealt with measured confounding. The main characteristics of the studies are displayed in Table 1. A more detailed description of the assessed methods can be found in Appendix Table A1. Most studies assessed continuous outcomes (54/84 studies, 64%). Performance measures were selected and extracted based on the study of Morris et al 11 The most frequently reported performance measures were bias (83/84 studies, 99%) and mean squared error (MSE) or root mean squared error (RMSE) (58/84 studies, 69%). Most studies reported three or more performance measures (66/84 studies, 79%) and most studies (48/84 studies, 57%) contained both a simulation study and an empirical application of the method(s) of interest.

TABLE 1.

Summary of the properties of the included simulation studies per type of confounding

Measured confounding (N = 84) Unmeasured confounding (N = 43) Total (N = 127)
Outcome type N % N % N %
Continuous 45 54% 27 63% 72 57%
Binary 30 36% 9 21% 39 31%
Both 9 11% 7 16% 16 13%
Reported simulation performance measures a N % N % N %
Bias 83 99% 40 93% 123 97%
MSE/RMSE 58 69% 20 47% 78 61%
Empirical SE 41 49% 17 40% 58 46%
Model SE 37 44% 4 9% 41 32%
Coverage 31 37% 12 28% 43 34%
Confidence interval length 20 24% 12 28% 32 25%
Type I error 11 13% 3 7% 14 11%
Power 9 11% 8 19% 17 13%
Convergence 7 8% 6 14% 13 10%
Total of the above measures N % N % N %
One 9 11% 7 16% 16 13%
Two 9 11% 12 28% 21 17%
Three or more 66 79% 24 56% 90 71%
Type of data used in the study N % N % N %
Simulation and application of empirical data 48 57% 27 63% 76 59%
Simulation only 36 43% 14 33% 50 39%

Abbreviations: MSE, mean squared error; RMSE, root mean squared error; SE, standard error.

a

Please note that one study can report more than one performance measure.

3.1.1. Traditional regression methods

Regression adjustment

Traditional regression methods (eg, ordinary least squares, maximum likelihood estimation) allow for an unbiased estimation of treatment effects if the model is correctly specified and all variables associated with the outcome are included in the regression model. No simulation study assessed the performance of traditional regression methods beyond using it as a comparator method. It is likely, however, that regression adjustment has been studied in ways other than simulation studies.

3.1.2. Propensity score (PS) methods

The general building blocks of the performance of PS methods will be discussed below, followed by an in‐depth discussion of the identified methods.

General building blocks of PS methods' performance

PS estimation was introduced to balance measured covariates between groups and formed the basis of several methods for dealing with measured confounding. 1 The general idea behind PS estimation is that the true PS is unknown, but can be estimated from the data. 12 The PS is estimated for all individuals and is the scalar function of covariates reflecting the probability of being assigned to one of the treatments conditional on observed baseline covariates. 1 , 13 The causal interpretation of PS estimation relies on the following five assumptions: “positivity”, “SITA”, “SUTVA”, “consistency”, and a “correct specification of the PS model” (see Appendix Box A). 14 , 15

When estimating the PS, three points are critical. First, it is important to achieve a good balance between the treatment and control groups in observed covariate distributions, as consistency is only warranted if such balance exists. 12 , 16 Currently, there is no single best criterion to define a good balance of covariate distribution between treatment groups, but the standardized mean difference is the most commonly used statistic to examine it. 17

Second, it is important to consider the overlap in propensity distributions. 18 The level of overlap can easily be examined visually by comparing propensity distribution plots. 19 Inadequate overlap can be addressed by restricting the study's conclusions to the sufficiently represented group of individuals in the overlap region. While removing individuals from the region of non‐overlap can reduce the bias of estimators, discarding particularly rare outcomes may lead to increased standard errors. 20

Third, selecting appropriate variables for estimating the PS is critical, because all PS‐based estimators are prone to bias when the PS estimator is misspecified. 21 Four of the included studies addressed this issue. 2 , 22 , 23 , 24 Austin found that including only true confounding variables in the PS model resulted in the least biased estimation of the treatment effect and the lowest MSE † . 22 However, the inclusion of a variable that is strongly related to treatment, but independent of the outcome may have the undesirable consequence of reducing precision without decreasing bias. 23 , 24 It is therefore suggested to include all covariates that are believed to be associated with the outcome variable, regardless of whether they are related to the treatment. 2 , 22 , 23 This is because inclusion of variables that are unrelated to the exposure but related to the outcome will increase the precision of the estimated exposure effect without increasing bias. 23

Identified PS methods estimated with logistic regression

PS matching

We identified 47 simulation studies on PS matching. Its most popular implementation is pair‐matching (eg, 1:1 matching), where treated and untreated individuals are matched based on their similarity in PSs, allowing for the estimation of ATT or ATE (Table 2). 13 , 27 Alternative matching techniques included many‐to‐one matching, variable‐ratio matching, and full matching ‡ . 13 Full matching was evaluated in 5 simulation studies 2 , 13 , 28 , 29 , 30 and has the advantage that it does not suffer from bias due to incomplete matching and enables the estimation of both ATT and ATE. The performance of full matching is comparable to that of IPTW. 13 Full matching was found to perform well for weak to moderate levels of confounding, and for the estimation of odds ratios, relative risks, and risk differences when the ATT was the target estimand. 13 It did, however, result in a bias for strong levels of confounding. This likely occurred due to extreme weights and was found to decrease when caliper restriction was introduced. 13 , 28

TABLE 2.

Reported estimands per method reported for measured and unmeasured confounding

Estimand
Measured confounding methods a , c ATE ATT ATO LATE Additional
PS matching b 26 17 3 NA SATT, SATE
IPTW b 25 7 3 NA NA
Overlap weights NA NA 3 NA SWATE
Matching weights 2 NA NA NA SATE
Covariate adjusted PS b 9 2 NA NA NA
PS stratification b 15 3 NA NA NA
Genetic matching 4 2 NA NA NA
Covariate balancing PS 3 2 NA NA NA
G‐computation 3 1 NA NA NA
Doubly Robust estimator b 12 1 NA NA NA
Bayesian methods b 4 NA NA NA CATE
Unmeasured confounding methods a ATE ATT ATO LATE Additional
Two‐stage least squares 8 NA NA 6 SLATE
Two‐stage residual inclusion 5 NA NA 3 NA
Two‐stage predictor substitution 3 NA NA 1 NA
Difference‐in‐Differences NA 8 NA NA NA
Difference in Differences with matching NA 5 NA NA NA
Synthetic control 2 5 NA NA NA
Lagged dependent variable NA 1 NA NA NA
Regression discontinuity NA NA NA 3 NA

Abbreviations: ATE, average treatment effect; ATO, average treatment effect in the overlap population; ATT, average treatment effect on the treated; CATE, conditional average treatment effect; IPTW, inverse probability treatment weighting; LATE, local average treatment effect; PS, PS; SATE, sample average treatment effect; SATT, sample average treatment effect on the treated; SLATE, super local average treatment effect; SWATE, sample weighted average treatment effect.

a

Please note that while some studies performed several estimands, others did not report the estimands clearly. The papers with non reported estimands (n=27) are not part of this table.

b

Please note that the methods are represented in broad categories. For instance, for PS matching no distinction is made between caliper and matching specifications, nor the PS estimating function, as it has no influence on the estimand.

c

Prognostic scores are not displayed, because they lead to the exact same estimands as their PS alternative.

When matching treated and untreated individuals, several considerations need to be taken into account, including the optimal matching ratio (ie, the ratio of treated versus untreated individuals, eg, 1:1,1:2,…,1:n). Several studies showed that, on average, increasing the number of untreated individuals matched to each treated subject increases the bias of the estimated treatment effect, while increasing precision. 31 , 32 To minimize bias, only one untreated subject for each treated subject seems to be the optimal approach. 32 , 33 Matching more than one untreated subject to each treated subject can be considered if the number of untreated subjects largely exceeds the number of treated subjects. One study found that when using nearest‐neighbor matching or caliper matching, the MSE was smallest in most scenarios (84%) when, at most, two untreated subjects were matched to each treated subject. 32

Four studies 20 , 34 , 35 , 36 evaluated solutions for handling poor overlap of PS distributions (ie, substantially different distributions of the observed covariates between treated and control samples). These studies evaluated caliper restriction, which pre‐specifies the allowed within‐pair difference in PSs and therefore reduces the number of dissimilar matches. 34 Caliper restriction may help to reduce bias, but the exclusion of some treated individuals may also reduce precision, and affect the generalizability of results, as the treatment effect estimate only applies to the matched individuals. 20 , 35 To prevent the exclusion of too many individuals, Austin suggested that matching on the logit of the PS using calipers with a width equal to 0.2 of the standard deviation of the logit of the PS would be optimal when estimating differences in means or risk differences. 34 When imposing calipers, it is essential to indicate to which sub‐population the resulting inferences apply. 35 Another solution for reducing poor overlap of PS distributions would be to discard observations that fall outside the overlapping region (ie, trimming) of the PS distributions. 36

The matching algorithm can be “greedy” or “optimal.” A greedy algorithm performs a stepwise local optimization procedure by finding the closest matches (eg, nearest neighbor matching). With an optimal algorithm (eg, optimal subset matching, cardinality matching), the average within‐pair difference in PSs is globally minimized at the cost of computational power. 33 Optimal matching might be helpful when insufficient control matches are available for the treated individuals, as it does not discard treated individuals. Resa and Zubizarreta showed that cardinality matching § performed better in terms of covariate balance, matched sample size, and treatment effect estimates than comparator methods, such as optimal subset matching and nearest‐neighbor matching. 36

Matching can be performed with replacement (ie, matching one untreated to several treated individuals, meaning that controls similar to several treated individuals can be used multiple times) or without replacement. 33 Austin found that matching with replacement did not result in less bias than the best‐performing methods based on caliper matching without replacement. 33 Moreover, matching with replacement tends to result in estimates with greater variability and a larger MSE than caliper matching without replacement. 33 , 34 The use of matching with replacement is therefore discouraged.

Inverse probability treatment weighting (IPTW)

IPTW was evaluated in 30 simulation studies. With the help of re‐weighting by the inverse probability of receiving the treatment, a synthetic sample is created that is representative of the population in which treatment assignment is independent of the observed baseline covariates. 15 Basu and Polsky found that when a PS estimand is correctly specified, IPTW is most likely to be unbiased and outperforms other PS methods for estimating ATE under alternative data generating processes (eg, cost data), but is inefficient compared to an unbiased regression estimator. 21 Two studies found PS matching and IPTW to perform similarly in terms of reducing baseline imbalances. 4 , 37 Compared to PS matching and genetic matching (see paragraph 3.1.3. below), IPTW provided the least biased, most efficient estimates for small sample sizes (n=100), but all methods performed poorly. In a simulation study by Austin, IPTW yielded smaller residual differences in baseline characteristics between treated and untreated subjects than stratification, or PS adjustment. 37 Zagar et al, however, found no superior performance of IPTW for estimating ATE compared to stratification or PS matching. 38 Austin also compared IPTW and full matching and found similar performance in terms of bias and MSE when ATT was the target estimand. 13 Both methods showed convergence issues at a sample size of 300. 2 While showing good performance under weak or moderate confounding, IPTW resulted in biased estimation of the true ATT estimand when the treatment‐selection process was strong, even under correct PS model specification. 13

Two studies found IPTW to be unstable when PSs of some individuals approached 0 or 1, resulting in adverse finite sample consequences. 8 , 39 Such extreme weights may disproportionately influence (ie, distort) the weighting and increase the variance of the treatment effect estimate, resulting in a decreased precision and poor balance. 8 , 39 , 40 One possible remedy for this issue is weight trimming at the tails of the PS distribution, which refers to a symmetric or asymmetric removal of weights outside a pre‐specified value. 7 Trimming, however, may distort the original population, meaning that the population becomes more selected, which will likely affect the generalizability of the results. 7 Another alternative is truncation, which defines a maximum weight and replaces extreme probabilities with the prespecified maximum. 20 , 41 While truncation removes variance in IPTW, this is at the expense of increased bias. 20 , 41

Alternative weighting methods beyond IPTW: overlap weights and matching weights

Some alternative weighting methods beyond IPTW are worth mentioning. To overcome the limitations of truncation and trimming, which are often sensitive to the choice of cutoff points or discard a large proportion of the sample, overlap weights were developed by Li et al This method was assessed in four simulation studies 7 , 8 , 42 , 43 and utilizes individual‐level weights representing the probability of a specific subject being assigned to the opposite group. 7 , 8 Overlap weights are bounded and lead to smaller bias, asymptotic variance, and good coverage compared to IPTW with trimming. 7 Li and Greene proposed matching weights, which can be considered an analogue to one‐to‐one pair matching without replacement and with a caliper restriction. 44 The method overcomes the difficulty of the variance estimation of pair matching while providing better efficiency. 44 For rare binary outcomes, the method shows good performance in terms of bias and MSE. Overlap weights and matching weights lead to an estimand that deviates from ATE (eg, average treatment effect in the overlap population [ATO] in the case of overlap weights). Mao and Li found both methods to improve power in poor overlap, while overlap weights showed better performance in terms of power. 42 Yang et al extended the use of overlap weights in the context of subgroup analyses. 43

Covariate adjustment using the PS

Thirteen studies assessed covariate adjustment using the PS, a simple method of incorporating the PS as a covariate into the regression model for estimating the treatment effect. 37 However, this combines data analysis with study design and, therefore, precludes an analysis blinded to the outcome. 1 , 13 Moreover, covariate adjustment using the PS does not result in balanced covariates across treated and untreated groups but tries to deal with confounding bias by constructing a prediction model. 45 Finally, the estimate of interest becomes highly model dependent since this method requires both the PS model and the outcomes regression model to be correctly specified, which increases the risk of misspecification. 46 Austin found that IPTW and PS matching can remove more systematic differences (treatment‐selection bias) than covariate adjustment. 37 Moreover, PS adjustment led to low bias and MSE for rare binary outcomes. 20

PS stratification (or subclassification)

PS stratification (or subclassification) was assessed in 26 studies and most commonly used five PS quintile strata to group individuals based on their PSs. 47 This method compares the within‐stratum distribution of the baseline covariates, which should be similar between the treated and control groups within the same quintile of the PS. 37 The overall treatment effect is then generally obtained by averaging across the stratum‐specific standardized differences by weighting by the proportion of observations in each subclass (stratum). 1 , 37 Combining stratum‐specific estimates by weighting by the inverse variance is an alternative approach to estimate the ATE. Rudolph et al found weighting by the inverse variance to perform similarly to weighting by the proportion in the absence of positivity violations and with a constant treatment effect, with weighting by the inverse variance performing slightly better. 48 The inverse variance weighting, however, is not recommended for heterogeneous treatment effects. 48

The performance of PS stratification is also dependent on the choice of strata. 24 , 47 , 49 , 50 , 51 Jill et al found that quintiles outperformed the other stratification schemes in the majority of their simulations. 24 When choosing the number of strata, it is recommended to consider sample size (eg, opt for more strata with a larger sample size), 52 the number of covariates, 24 and the data generating mechanism. 48 It is also recommended to examine the robustness of the results to the number of strata. 47 , 49 , 50 , 51 The choice of cutoff points for the strata was found to influence the variance and bias of the estimate. As a rule of thumb, wide strata produce low variance but high potential bias, while the opposite is true for narrow strata. 53 One should also be aware that PS stratification is likely to suffer from residual confounding; that is, it seems unable to fully remove bias as it is impossible to classify individuals into several strata with a common PS. 40

3.1.3. Alternative methods for PS estimation

PSs are typically estimated using logistic regression. 54 While they are easy to interpret and tend to converge on parameter estimates relatively easily compared to semi‐parametric or non‐parametric models, they might not be flexible enough to capture complex relationships. This may result in model misspecification. Logistic regression models also bear a high risk of misspecification when error terms are not independent and are not identically distributed from their known parametric distributions. 55 , 56 Below, the identified alternative methods for PS estimation will be described.

Genetic matching

Genetic matching (a non‐parametric matching method) was assessed in seven studies and algorithmically optimizes covariate balance. 12 By allowing the covariates to enter the model with different weights, genetic matching has additional flexibility compared to other PS methods. 12 Additionally, by taking the covariance structure of the variables into account in the distance calculation between covariates, some modeling difficulties caused by collinearity between covariates can be eliminated. 12 , 57 Five studies compared genetic matching to PS matching, and found that genetic matching achieved better covariate balance between groups and produced more stable and less biased treatment effect estimates with a lower MSE compared with PS matching. 12 , 57 , 58 , 59 , 60 Genetic matching also performed better compared to PS in terms of bias and precision matching and compared with IPTW for estimating subgroup effects when subgroup‐specific treatment assignment was ignored. 60 , 61 Additionally, the method was found to be more robust to PS misspecification than IPTW and PS matching. 60 However, genetic matching was found to perform poorly when noise covariates dominated the data because genetic matching aims to balance all covariates, including the noise. 38 This property of genetic matching also leads to an increase in the dimensionality of the matching problem, which may, in turn, result in a loss of precision.

Generalized additive models (GAM)

One simulation study evaluated GAMs, which replace the linear component of a logistic regression with a flexible additive function. Woo et al found that estimating PSs with GAMs can improve covariate balance and reduce bias compared to estimating PSs with logistic regression in situations where there is adequate overlap in the covariate distributions between treated and untreated individuals. 62

Generalized boosted regression (GBM) trees

GBM trees were evaluated in 3 studies and aimed to estimate the function of covariates in a more flexible manner than logistic regression by averaging the PSs of small regression trees. 63 GBM tends to produce better results than logistic regression “when the true PS model is considerably non‐linear or contains important interactions in the log‐odds scale, or the outcome model is not linear in the observed confounders”. 40 Moreover, GBMs were found to work well in situations where there are complex relationships between variables, and because they implicitly select and estimate interaction effects, they are more efficient when it is unclear which interactions to include in a model. 62

Covariate‐balancing PSs

Five studies evaluated covariate‐balancing PSs, which combine PS estimation with covariate balancing to simultaneously optimize treatment assignment prediction and covariate balance. 64 , 65 , 66 The technique was found to be relatively robust to PS model misspecification in terms of balancing covariates and reducing bias, compared with logistic PS matching, PS weighting, boosted classification, and regression trees. 64 , 65 , 67 However, its success is conditional on identifying a complete set of confounders. 67 The method performs well if the parametric assumptions of a logistic regression model are satisfied and the outcome model is linear in the covariates. 63 However, when substantial non‐linearity is present in the true PS model or the outcome model, GBM tends to provide better results.

3.1.4. Doubly robust (DR) estimators

Thirteen simulation studies evaluated the performance of DR estimators, which were developed to protect against the adverse consequences of model misspecification. 68

Augmented inverse probability of treatment weighted (AIPTW)

The AIPTW estimator was assessed in eight studies, and achieves the doubly‐robust property by combining outcome regression with weighting by the PS. 69 , 70 AIPTW provides consistent estimations of the ATE if either the outcome regression model or the PS model are correctly specified. 40 However, if both models are misspecified, the DR estimator is prone to severe bias. 39 , 40 , 69 , 71 , 72 Basu and Polsky found that this property of double robustness comes at the expense of efficiency, which is in between that of regression methods and IPTW. 21 Because AIPTW is not a substitution estimator # , it may yield negative risk estimates. 74 Also, similarly to the IPTW, the AIPTW estimator is expected to perform poorly in the presence of extreme probabilities. 39

Stratified DR estimator

The stratified DR estimator was assessed in one simulation study, and is a hybrid of outcome regression with PS weighting and stratification. 40 This method relies on two separate PS models and an outcome regression. It appears that the incorporation of stratification protects against bias due to the misspecification of the first PS compared to other DR estimators. 40 It was also found that existing DR estimators may suffer from considerable bias when the link function is misspecified, whereas the stratified DR is robust against misspecification of the link function in the second PS, which is an inherited property of PS stratification. 40

Targeted maximum likelihood estimation (TMLE)

TMLE was assessed in two studies and is a semiparametric DR method that allows for flexible estimation using (nonparametric) machine‐learning methods. 68 The estimator guarantees consistent estimation if either the PS or the conditional mean outcome is correctly estimated. 74 Unlike AIPTW, TMLE is a substitution estimator, meaning it will always result in point estimates within the parameter range. 74

Collaborative TMLE

Since all previous methods are based on PS estimation, they also rely on its assumptions and are associated with the same problems. Amongst others, near‐violations of the positivity assumption might be problematic. Lendle et al suggested that data‐adaptive estimation methods (ie, collaborative TMLE) may provide a solution to near‐violations of the positivity assumption in DR estimation, but further research into this issue is warranted. 75

3.1.5. Bayesian approaches

Bayesian PS approaches aim to overcome the limitation of frequentist methods that they ignore the uncertainty around the estimated PS and instead treat it as a known quantity in the outcome stage. 76 Thus, standard errors for the treatment effect estimate are usually estimated without acknowledging uncertainty in the estimated PSs, resulting in falsely precise confidence intervals around the treatment effect in frequentist PS models. 77 , 78 Findings regarding the direction of this imprecision are inconsistent. While McCandless et al and Kaplan et al found that frequentist methods underestimate the variability of the treatment effect estimate, An stated the opposite. 77 , 78 , 79 This discrepancy might be due to a difference in the utilized variance estimator. 77 , 79

One‐step Bayesian PS or joint Bayesian PS

McCandless et al and An described a one‐step Bayesian PS or joint Bayesian PS approach, which was assessed in three simulation studies. The method models the joint distribution of the data and the parameters ‖ . 78 The model parameters are sampled from the joint posterior distribution using a Markov Chain Monte Carlo algorithm. 78 The marginal posterior for the treatment effect ensures that uncertainty is incorporated in the PS estimation. 78 Simulation results showed that accounting for uncertainty in the estimated PSs does not significantly improve coverage probabilities. 78 McCandless et al found that weak associations between the covariates and the treatment resulted in wider confidence interval estimates for the joint Bayesian PS analysis method than for the frequentist PS analysis. 78 Moreover, in a setting where the covariate‐treatment and covariate‐outcome relationships are different for every covariate, joint Bayesian estimation may not adequately estimate the treatment effect. Another potential weakness of the method is that since joint Bayesian PS with feedback uses information from the outcome model to estimate the PS, it may fail to adjust for confounding leading to biased treatment effect estimates. 76 Corwin et al showed that PS adjustment with additional covariate adjustment in the outcome model using Bayesian methods can accurately estimate the treatment effect and correct for this distortion. 76 However, the implementation of joint Bayesian PS implies a significant computational burden compared to frequentist methods. 76

Two‐step Bayesian approach

Kaplan and Chen argued that joint estimation of the data and the parameters is conceptually problematic, as the predictive distribution of PSs will possibly be affected by the outcome. 79 Therefore, they addressed the problem of joint modeling by separating the PS model from the outcome model, proposing the two‐step Bayesian approach, which was assessed in two studies. 79 The method seems to slightly outperform frequentist methods for small samples. 77 , 79

Bayesian model averaging approach

Kaplan and Chen provided a fully Bayesian model averaging approach within the PS context and compared it with an existing approximate approach to Bayesian model averaging. 80 Since the differences between both methods were small, the use of the new approach is not likely affect inferences dramatically.

Intermediate Bayesian approach

The intermediate Bayesian approach was assessed in two studies and addresses the aforementioned joint modeling problem by specifying a Bayesian PS model first, followed by a conventional ordinary least squares outcome model. 77 , 79 This half‐Bayesian method requires less computing power than the joint Bayesian method. 77 An found that the standard errors produced by the intermediate approach were considerably larger than those of the fully Bayesian or frequentist PS. 77 Kaplan et al, however, showed that, on average, the method performed well in terms of the closeness between the estimated and true variability. 79

3.1.6. Alternative methods

G‐computation

Four simulation studies evaluated the performance of G‐computation, a maximum likelihood substitution estimator of the G‐formula. 81 Chatton et al argued that the current predominance of PS methods over G‐computation methods is not justified due to their comparable performance. This predominance might be explained by the lack of an explicit variance formula and didactic explanations of G‐computation. 2 , 81 Porter et al found G‐computation to be more robust to non‐positivity compared to IPTW and DR estimators. 82 Moreover, while PS methods were found to be particularly sensitive to positivity assumption violations, and even near‐violation can cause a decrease in performance, G‐computation appears to be more robust to non‐positivity. 2 , 75 G‐computation, however, is generally not an efficient estimator, as it optimizes the bias‐variance trade‐off for the expectation of the outcome given the treatment and covariates instead of the parameter of interest. 83 Even though we identified four studies that compared G‐computation to other methods, the information regarding the method's performance was limited, which we believe has three reasons: (1) G‐computation was typically used as a comparator method and was not the primary interest of the study, (2) the authors focused on specific aspects of the method rather than its statistical performance in general, and (3) we did not include time‐dependent confounding studies, where the method might be most promising. 3 , 6

Prognostic scores (disease risk scores)

Prognostic scores were assessed in seven studies. Prognostic scores are considered the prognostic analog of PS methods, 29 as they model the association of covariates with potential outcomes. 84 The difference between the two methods is that the prognostic score includes covariates based on their ability to predict response, whereas the PS includes covariates that predict treatment assignment. Compared to the PS, the “positivity” assumption is not required for prognostic score models. 30 For the prognostic scores it is sufficient that there are no levels of disease risk at which treatment or control is received with certainty. 84 Hansen suggested using prognostic scores complementary to PSs or in situations where PS estimation is infeasible. 84 , 85 Leacy and Stuart showed that for binary treatments and continuous outcomes, methods combining the estimated PS (eg, matching on a Mahalanobis distance combining the estimated propensity and prognostic scores) and prognostic scores are preferred over PS estimation alone for the estimation of ATT. 85 Arbogast and Ray recommended the use of prognostic scores for cohort studies for which covariates are not strongly correlated with treatment. 86 Wyss et al found that prognostic scores lead to larger overlap in disease risk across treatment groups compared to PS matching, and therefore a potentially larger number of matches. 87

3.2. Identified methods for dealing with unmeasured confounding

Of the 127 included studies, 43 (34%) dealt with unmeasured confounding. A description of the assessed methods can be found in Appendix Table A2. The main characteristics of the studies are displayed in Table 1. Similar to measured confounding, continuous outcomes were most commonly reported (34/43 studies, 79%) amongst unmeasured confounding simulation studies. The most frequently reported simulation performance measures were bias (40/43 studies, 93%) and MSE or RMSE (20/43 studies, 47%). Most studies reported three or more performance measures (24/43 studies, 56%). The majority of the studies (27/43 studies, 63%) included both a simulation study and an empirical application of the method(s) of interest.

3.2.1. Instrumental variable approach (IVA) and analogue methods

The IVA was assessed in seventeen studies. All IV methods rely on the existence of a strong instrument ** (ie, variable that is strongly correlated with the endogenous treatment variable but uncorrelated with the outcome) that should not be contaminated (ie, correlated with the observed residuals) to guarantee the consistency of the estimation. 88 , 91 , 92 Crown et al found that a high degree of contamination of the instrument resulted in a biased IVA estimator and that the maximally acceptable level of contamination is relative to both the instrument's strength and the strength of confounding. 91 With IVA, LATE is estimated, which is the effect of the treatment on the treated for those who comply with the encouragement, and their treatment status changes due to the instrument (see Section 2.3). 92

Two‐stage least squares (2SLS)

IVA can be estimated with 2SLS, which were evaluated in 11 studies and consists of two consecutive least‐squares regression models. 89 The first model uses the instrument and additional covariates to predict treatment uptake. 89 The second regression uses the predicted value of treatment and additional covariates to estimate the outcome. 89 Conventional 2SLS for continuous outcomes with a linear first stage leads to consistent estimates when the assumptions required for identifying LATE using IVs are met (see Appendix Box B), and the instrument is strong. 93 Three studies found that 2SLS resulted in less bias than the conventional OLS, at the cost of higher variances, as long as the instrument is sufficiently strong. 88 , 94 , 95 When outcomes were rare, the 2SLS estimates were somewhat unstable. 94 Boef found that for small samples, OLS estimates were, on average, closer to the true effect than 2SLS estimates due to their smaller variance. 95 The minimum sample size above which the performance of 2SLS analysis exceeds that of OLS regression depends upon the strength of the instrument and the level of unmeasured confounding. 95 The authors present an equation for approximating the threshold sample size beyond which 2SLS estimates are closer on average to the true effect than OLS but also emphasize that the presence of a strong instrument is not very likely in medical research. 95 Uddin et al reported unstable and biased treatment effect estimates for both binary and continuous instruments when the instrument was weakly associated with the treatment. 88 The variability of the estimates between simulation runs in the presence of weak instruments was more pronounced for rare outcomes (<1%). 88 For binary outcomes, 2SLS is implemented using a linear probability model †† . This implies that bias might be present because the linear probability model does not respect the boundary conditions, and therefore imposes no constraints on the parameter values. 96 Uddin et al found that 2SLS performed well for both binary and continuous outcomes, but only in scearios with a strong instrument. 88 Basu et al showed that 2SLS produces consistent estimates of LATE in the context of binary outcomes even for rare treatments or outcomes. 97 Bhattacharya et al showed, however, that 2SLS estimators are generally not consistent. 98 That is, linear probability models perform well only if “the data generating process is normal or approximately normal, and the average probability that the dependent variable is one is not near 0 or 1, and there is only one treatment” (in addition to the reference treatment). 98

Two‐stage predictor substitution (2SPS)

The 2SPS was evaluated in five studies. The method replaces the second stage equation with logistic regression, aiming to overcome the restrictions of 2SLS for binary outcomes. In general, 2SPS tends to lack consistency, and bias does not seem to attenuate with increasing sample size. 99 , 100 Cai et al showed that the 2SPS estimator is asymptotically biased even when no unmeasured confounding is present, and that bias increases with an increasing magnitude of unmeasured confounding. 101 , 102 Wan et al showed that 2SPS lacks consistency in non‐collapsible models even given a perfect IV but is consistent when estimating collapsible effect measures (such as risk rates or risk differences) using Poisson or additive hazards models. 103

Two‐stage residual inclusion (2SRI)

The 2SRI was evaluated in 8 studies and adjusts for endogeneity by using the predicted or residual values from the first stage regression as covariates in the second stage outcome model ‡‡ . 99 , 101 , 103 Conclusions regarding the performance of 2SRI are not consistent across simulation studies, which might be explained by differences in simulation settings. Terza et al found 2SRI to be consistent for estimating ATE in models with binary outcomes without relying on the assumption of treatment effect heterogeneity §§ . 99 Chapman and Brooks warn, however, that this property may be the consequence of a specific set of simulated individuals that comply with the assumptions required for the estimation of the ATE. 105 Their study showed that a non‐linear 2SRI could only guarantee unbiasedness when the distribution of treatment effectiveness for the marginal patients differs from that of the general patient population only to a small extent. 105 Wan et al demonstrated that the consistency of 2SRI only holds under specific conditions ¶¶ . 103 Basu et al argued that in specific cases, for instance, cases with very rare outcomes, 2SRI (with Anscombe residuals) may be close to the unbiased estimator of ATE, and therefore may provide a better estimate than the 2SLS. 97 Whether this finding constitutes a general functional characteristic of the 2SRI or is a result of the unique simulation setting of Basu et al remains unclear. Cai et al estimated the causal odds ratio with 2SRI without bias when no unmeasured confounding was present, while bias increased with an increasing level of unmeasured confounding. 101 This finding is in line with Koladjo et al for weak instruments. 90 However, when a strong instrument is available Koladjo et al suggested that 2SRI can lead to good results in terms of bias and confidence interval estimate. 90

IV‐based generalized structural mean models (IV‐GSMM)

IV‐GSMMs were evaluated by one study and aim to overcome the difficulty of estimating treatment effects for binary outcomes and show more promising results than 2SRI or 2SPS. 102 A known weakness of GSMM, however, is its reliance on the assumption that the treatment‐outcome association is correctly modelled for each level of the IV. Moreover, in simulations where the logistic estimator was biased upward due to the large amount of unmeasured confounding, the estimation of the treatment effect sometimes failed to converge. 102

Instrumental PS (IPS)

The IPS was evaluated in two studies and aims to reduce the dimensionality of the measured confounders while also dealing with unmeasured confounders by combining PS matching with IVA. 106 Lu and Marcus showed that this method performs well as long as the magnitude of unmeasured confounding remains low or a strong instrument is present. 107 Jing and Winston found that incorporating IPS into a semiparametric method improved the estimation of treatment effects for skewed or binary outcomes in terms of bias, compared to 2SLS. 106

3.2.2. Difference‐in‐differences and analogous methods

Difference‐in‐differences (DiD)

DiD analysis was assessed in seven studies. The method requires data on study outcomes for groups exposed and not exposed to the treatment both before and after implementation of the treatment, and obtains the treatment effect estimate by comparing the average change over time in the outcomes. 108 A confounder variable in a DiD study is any variable related to both treatment assignment and the change in the outcome over time (ie, the trend). 109 The association between an unobserved confounder and the outcome should be equal in the treatment and control groups and constant over time. 108 , 109

DiD seems highly sensitive to the identifying assumption of parallel trends (see Appendix Box C). When the parallel trend assumption fails, the DiD estimator becomes biased, because changes unrelated to the treatment are confounded by the effect of the treatment ## . 111 Two studies showed that a common pitfall of DiD is not taking correlations, such as serial correlations or within‐cluster correlations, into account. Bertrand et al showed that serial correlation introduces severely biased standard errors around the treatment effect estimate and hence might result in false rejections of the null hypothesis of no effect. 112 To adjust for this, in large sample sizes, the authors proposed variance‐covariance matrix corrections or aggregating the time‐series information. 112 Rokicki et al found that grouped/clustered observations may also lead to correlated errors across individuals within groups. Ignoring this may lead to misleadingly small standard errors and false rejection of the null hypothesis. In comparison, aggregation, permutation tests, bias‐adjusted and wild cluster bootstrap produced coverage rates close to the nominal rate and were sufficiently powered when within‐cluster correlation was present. 113

Matching enhanced DiD

DiD models that include matching were assessed in six studies and may result in less biased point estimates than DiD when the parallel trend assumption fails. Matching, however, cannot eliminate bias induced by the failing counterfactual assumption (see Appendix Box C), but can decrease its magnitude. 109 , 111 Matching on pre‐treatment measurements can be done for three kinds of variables: (1) outcome trends, (2) outcome levels ‖‖ , 114 and (3) covariates.

Lindner and McConnell found that DiD combined with matching on pre‐treatment outcome levels tends to suffer from at least two sources of bias, either coming from time‐invariant differences among observations (ie, differences in the distribution of the level effect) or increasing standard deviation of the error term (ie, unobserved, random error term; bias due to short‐term fluctuations). 111 Daw and Hatfield, found that attempting to “match‐away” pre‐treatment outcome level differences in time‐varying variables, may introduce regression to the mean. When treatment assignment is correlated with outcome level in the pre‐treatment period only, matching on pre‐treatment outcome levels can lead to additional bias. 109 The magnitude of the resulting bias depends on the magnitude of the regression to the mean, which is influenced by (a) how extreme the pre‐treatment variation is in the matching variables in the matched sample relative to those of the unmatched sample, and (b) the correlation between the matching variable and the post‐treatment outcome (ie, the serial correlation between measures of the outcome). 109 The bias appears to be most severe in cases where few pre‐treatment periods are used to match, and there is large outcome variation and low serial correlation. 115 If treatment assignment is correlated with pre‐treatment trend only, matching on pre‐treatment level does not introduce additional bias.

Matching on pre‐treatment outcome trends performed well when high serial correlation of the outcome was present and when the error variance was small. 111 Daw and Hatfield showed, however, that the true bias reduction is marginal. For instance, highly unstable trends over time will result in a trivial or no reduction in bias, while the researchers may think that the effect estimate is accurate because the assumptions seem to hold. 116 This is because matching on pre‐treatment trends produce a matched sample with similar trends in the treatment and control groups by construction, which is in turn considered as an indication of parallel trends. 116 However, if the trends in the outcome over time are unstable, reversion to the population mean trend in the post‐treatment period occurs, thus the matching does not result in a meaningful reduction of bias. 116

Ryan et al argued that only a narrow set of scenarios would cause matching to produce biased effects due to regression to the mean, namely when matching variables are unstable (ie, modest serial correlation in the matching variable) and the overlap between groups at baseline is large. 116 , 117 Daw and Hatfield showed, however, that these scenarios are not unlikely. The difference between the studies is that Daw and Hatfield simulated overlapping but separate distributions for the matching variables in the treatment and comparison groups, whereas the samples of Ryan et al were drawn from the same distribution and the probability of being part of the treated group increased for high pre‐treatment outcomes. This explains why the matching was a success in the study of Ryan et al, while it was found to induce a larger bias in the study of Daw and Hatfield. The authors agree that matching can be beneficial in combination with DiD if: (1) the matching variables are correlated with future outcome trends, or (2) strong theoretical or subject matter evidence supports that treatment and control groups come from the same population. 115 , 116 Illenberger et al added to the research of Daw and Hatfield by developing a novel correction for regression to the mean bias that can be used in a sensitivity analysis to check how robust inference may be to the effect of regression to the mean.

Synthetic control method (SCM)

Another possible approach when the parallel trend assumption fails is the SCM, which was assessed in seven studies. SCM uses a time‐invariant weighted average of the available controls, which prior to the treatment, have similar pre‐treatment characteristics and outcome trajectories to the treated. 118 , 119 This is because the combination of units often provides a better comparison for the unit exposed to the treatment than any single unit alone. 120 The synthetic control is constructed as a weighted average of the available control units. 114 The choice of weights ensures that prior to the treatment, levels of covariates and outcomes are similar over time to those of the treated unit, which was first defined as being restricted to be non‐negative and summing to one. 114 , 120 The failure of the assumption that a weighted set of comparison units can represent the treatment group's counterfactual outcomes leads to biased treatment effect *** . 121 Powell relaxed the strict assumptions by allowing outcomes to be functions of transitory shocks, referring to this approach as imperfect SCM, which resulted in less bias than the traditional SCM. 121

Ferman et al provided evidence from simulations that specification searching when choosing the predictor variables and covariates for weight estimation may be a problem in real SC applications. 122 SCM is robust to specification searching if a large number of pre‐treatment periods is available and a restriction to a specific subset of specifications is imposed. 122 If the number of controls is large, SCM may result in low bias with high variance (“overfitting”). 123 The overfitting can be controlled with an imposed weight restriction, which reduces variance; however, it does lead to an increase in bias. 123 O'Neill et al found that the SCM is relatively inefficient compared to the LDV estimator, which may be biased when few pre‐treatment periods are available. 114 Kinn argued that neither DiD nor traditional SCM is well suited for estimating the counterfactual in a high dimensional setting, where more flexibility (eg, machine learning techniques) could capture the data generating process better. 123 Illenberger et al found that similarly to DiD with matching, synthetic control may also suffer from regression to the mean, resulting in inflated type I error rates and bias toward the null. 124 The authors, therefore, proposed a method to determine the sensitivity of matched estimators of the ATT to the effect of regression to the mean bias. 124

Xu and O'Neill et al proposed the generalized SCM, which combines SC with the efficiency gains of interactive fixed effects. 125 , 126 The method requires the existence of repeated observations of the same units over time (ie, panel data) and multiple pre‐treatment periods (one more than the specified number of interactive fixed effects to include). 114 Xu found the method to be less biased than DiD and more efficient than the original SCM estimator under the premise of correct model specifications. 125 O'Neill et al found the method to perform better than its comparators when nonparallel trends were present, treatment effects were heterogeneous, and only a few pre‐treatment periods were available. 126 Arkhangelsky et al proposed and evaluated the synthetic DiD method, which implements two‐way (both unit and time) fixed effects and weights and can be interpreted as a unit and time‐weighted version of the standard DiD estimator. 127 The method has a promising, double robust property that allows for consistency if either the unit and time weights are well chosen or the outcome model is correctly specified. 127

Lagged dependent variable approach (LDV)

The LDV approach adjusts for pre‐treatment outcomes and covariates with a parametric regression model. 114 Up until now, the LDV has only been assessed in one simulation study, in which it was found to perform better than the synthetic control or matching enhanced DiD methods in terms of bias and efficiency when the parallel trend assumption does not hold. When the parallel trend assumption fails, the inclusion of a LDV increases bias, instead of diminishing it. 114 O'Neill et al showed that the synthetic control and matching approaches report even greater bias than LDV when the parallel trends assumption holds. 114

3.2.3. Alternative methods

Trend‐in‐trend approach

The trend‐in‐trend approach was assessed in one study. The method performed well when there was a strong time trend in treatment (eg, a newly approved drug or medical treatment). 128 The method has only been compared to ignored unmeasured confounding, but not comparator methods.

Prior event rate ratio (PERR)

Three simulation studies assessed PERR. The method is considered a quasi‐experimental method that derives a flexible pairwise Cox likelihood function, which is used to estimate treatment effects. 129 Similar to DiD, this method also strongly relies on the constant confounding effect assumption. Yu et al showed that the PERR method performs well when the effect of treatment (on the outcome) is relatively large in comparison with the confounder‐treatment interaction (on the outcome). 130 Uddin et al found that PERR does not perform well in situations when prior events strongly influence the probability of subsequent treatment. 131

PS calibration

PS calibration combines PS and regression calibration when confounders are observed in a validation study. 132 To our knowledge, the method has only been assessed in a single study, that found the method to perform well when the surrogacy assumption holds, meaning that the PS estimate of the main study contains no additional information on the outcome. 132

3.2.4. Regression discontinuity (RD)

The RD was assessed in four studies. It requires a specific study setting where the treatment is initiated based on a certain threshold level, which allows for a comparison of individuals around the threshold ††† . Assumptions are relatively weak compared to other unmeasured confounding methods (see Appendix Box D). 133 RD can be categorized as sharp RD if treatment assignment is deterministic and discontinuous at the threshold and as fuzzy RD otherwise ‡‡‡ . In the following, only the performance of sharp RD will be discussed, as the fuzzy RD requires specific adjustments that are outside the scope of this review.

Sharp RD

Estimation of the treatment effect with sharp RD can be based on parametric or non‐parametric (eg, local linear regression) estimation techniques §§§ . While the parametric approach may offer greater precision due to its utilization of all available data, the potential model misspecification may lead to an increase in bias. 134 The performance of RD strongly relies on the appropriate choice of bandwidth (eg, the area around the cutoff). Imbens proposed an optimal, data dependent, bandwidth choice rule, which showed good performance in simulations for estimation by local linear regression. 9 Barreca et al found that non‐random heaping in the data will lead to biased RD estimates. 133 The authors proposed two approaches to establish the validity of RD designs if the distribution of the treatment has heaps and suggested separately investigating non‐heaped and heaped data when possible.

4. DISCUSSION

4.1. Main findings

The aim of this paper was to perform a scoping review of methods available in PubMed for single‐point exposure to deal with measured and unmeasured confounding when evaluating treatment effects in observational studies. The most frequently tested methods to adjust for measured confounding were PS matching, IPTW, and PS stratification. Based on the studies reviewed, we found that there is sufficient literature supporting the use of PS adjustment, matching, and stratification methods to adjust for measured confounding. One of the most critical assumptions required for the valid use of PS methods is a correct specification of the PS model because misspecification leads to biased treatment effect estimates. The most commonly tested method for estimating PSs in simulation studies was logistic regression, but more flexible alternatives for PS estimation have also been introduced, which may improve the model specification. Another way of overcoming potential misspecification of either the outcome regression model or the PS model is using DR methods, which allow for consistent estimation when either model is correctly specified. Bayesian techniques go beyond frequentist methods and address the uncertainty in the PS estimation. However, the Bayesian PS methods require more research as an in‐depth investigation of such methods is currently lacking. Prognostic scores provide an alternative to PSs, but the simulation literature on prognostic score methods is sparse, and this method is only recommended as a complementary method to PSs rather than a standalone method.

The most frequently tested methods to adjust for unmeasured confounding in observational studies were IVA, DiD, and DiD analogues. 135 These methods aim to mimic the random treatment allocation done in experiments with the help of specific data characteristics. All these methods rely on strong assumptions or availability of a particular type of data, such as an instrumental variable, parallel trends, pre‐treatment outcomes, or a threshold for treatment initiation. Therefore, the choice of the method should be primarily based on the specific setting in which the study is conducted. Among IV‐type methods, there is sufficient evidence that 2SLS performs well for continuous outcomes, but 2SPS and 2SRI methods, while researched extensively, did not prove to be valid. Further research is required in the area of matching enhanced IV approaches, for example, near‐far matching for binary outcomes. Until now, however, simulation studies assessing the performance of this method are lacking. We mention this method explicitly here, because it may allow the researcher to strengthen the instrument by matching similar individuals with sufficiently different levels of IV, and thereby overcomes the problem of a weak instrument. 136 However, while acquiring greater internal validity, this advantage comes at the expense of generalizability. 136 The use of DiD, DiD with matching, and SCM is supported by substantial evidence. While having promising properties, both DiD with matching and SCM have their pitfalls, such as potential regression to the mean bias.

Conclusions regarding the identified methods' performance were mostly consistent across studies, but contradictory findings were reported for 2SRI and DiD combined matching. These inconsistencies most likely occurred due to differences in simulation settings.

4.2. Comparison with previous studies

Streeter et al provided an extensive overview of methods for dealing with unmeasured confounding in longitudinal observational studies. 135 They included both simulation and applied studies and provided a solid overview of method‐specific obstacles to implementation and a summary of application practices. Our paper includes the latest literature and complements their research by adding methods for dealing with measured confounding as well as information about the identified methods' assumptions, estimands, and performance. Webster‐Clark et al performed an extensive review of PS methods and described the alignment with the potential outcomes' framework, implications for study design, estimation procedures, implementation options, and reporting. They also demonstrated the evolution of balance assessment techniques over time and discussed evolving applications of PSs. Although Webster‐Clark et al included manuscripts from PubMed, Embase, Web of Science, and Scopus, they only selected a random sample of manuscripts from 2014 to 2018 for evaluation. Our review only included PubMed, but we systematically extracted all simulation papers identified. Our review complements this previous review in two ways. First, we focused on simulation studies to compare the performance of the identified methods, whereas they focused on application studies. Second, our paper does not only focus on PS methods, but also on other viable methods, thereby providing a more complete overview of available methods.

4.3. Strengths and limitations

This review was the first to provide a comprehensive overview of methods for dealing with both measured and unmeasured confounding. Our review includes information on both the methods' assumptions and the main technical remarks related to the methods' performance. A limitation of the review is that we included PubMed and simulation studies only. This may have led to the exclusion of some potentially relevant papers. The focus on simulation studies also means that there is a selection effect on the methods that are picked up and emphasized, which is reflected in an over‐representation of PS methods and an under‐representation of other methods (eg, outcome adjusted, g‐computation). Moreover, in contrast to what we expected, our search did not identify many simulation studies assessing the application of RD. This might be explained by the fact that the method was introduced before simulation studies became a popular evaluation tool of performance. However, we are confident that we did not miss any appropriate methods completely that would satisfy our inclusion criteria. Another potential limitations is that because all simulation findings are focused on specific case scenarios (eg, simulation settings), the generalizability of our results might be limited. This means that the observed results and conclusions are specific to the simulation scenarios considered, and one should avoid generalizing results to settings that have not yet been evaluated. Furthermore, even though continuous outcomes dominated the simulation studies, the presence of binary and continuous outcomes was well balanced within the methods, except for those that were developed for a specific type of outcome (ie, 2SRI, 2SPS). Lastly, there is currently no available guideline for evaluating the quality of simulation studies, like the CONsolidated Standards of Reporting Trials (CONSORT) statement 137 for clinical trials and the Consolidated Health Economic Evaluation Reporting Standards (CHEERS) statement 138 for economic evaluations. While Morris et al developed a tool for designing and reporting a simulation study, this is not a consensus‐based checklist. 11 When reviewing the studies, we did observe some subjective differences in quality between studies, such as the number of scenarios considered, the complexity of the data generating mechanism, the number of reported performance measures and os on. However, objectively all of these can be considered study‐specific assumptions and choices.

4.4. Recommendations for research practice and further research

We did not identify a single best‐performing method to deal with confounding that performs well under all circumstances. There is a broad selection of appropriate methods, and we recommend choosing a specific analytical method based on the research question, the properties of the data, and most importantly, the population's characteristics. It is also important to emphasize that methods for measured and unmeasured confounding produce different estimands. Policy‐making, health technology assessment, and clinical outcome evaluation may prioritize certain estimands over others. For instance, a common criticism against the IV method is the difficulty in interpreting the LATE. Because the population of compliers ¶¶¶ is unidentifiable when using real‐world data, converting the effect estimate into an outcome of practical relevance for decision‐makers is somewhat problematic. 139 On top of that, the represented population by LATE might be highly selective and may depend on the chosen instrument. 140 Methods estimating ATEs ### may also be unsuitable for policy decisions as they provide the treatment effect among those for whom the treatment was never intended. 4 ATT, on the other hand, assesses the treatment effect on individuals who were actually treated. 4 The estimand and the targeted population must be fully understood before choosing an estimation method. 4 It is important to realize that while this is a crucial decision in observational studies, RCTs by design produce ATEs that are equal to ATTs, since the treated and untreated populations do not differ systematically.

Although there is little difference in the assumptions between methods for dealing with measured confounding, they do play a key a role in the choice of method when dealing with unmeasured confounding. For example, DiD analysis cannot be performed without fulfilling the parallel trend assumption, and the IV method not without fulfilling the exchangeability assumption. Finally, the available data is decisive when choosing a method, especially in case of unmeasured confounding. While IV requires a strong instrument, DiD needs data prior to and after treatment for both treated and untreated individuals. Regression discontinuity can be used if treatment uptake is defined based on a certain threshold. Moreover, while there is a wide range of literature evaluating the performance of methods for measured confounding (eg, PS matching, PS stratification, and IPTW), more research is needed in the field of unmeasured confounding. For unmeasured confounding, most research focuses on IV and DiD methods. Since both methods require very strict assumptions, more research on the weakening of assumptions is awaited. While several papers focus on rare outcomes and nonlinear relationships, more research is also needed on methods for skewed distributions.

Finally, the methods identified in this paper rely on correct variable selection and a good understanding of the causal path. This is essential to avoid introducing new biases such as collider bias by conditioning on the common effect.

5. CONCLUSION

Many methods are available for dealing with measured and unmeasured confounding. First, researchers need to identify whether the presence of unmeasured confounding is likely in their data. Second, based on the research question, outcome type, and the aim of the study, one needs to decide which estimand is of interest. We showed that for providing consistent estimates, the methods strongly rely on specific assumptions, some of which cannot be empirically tested. Thus, when choosing between methods, it is important to consider the plausibility of assumptions that the method requires. Violations of these assumptions undermine the reliability of the causal effect estimates. Instead of claiming that one method is better than the rest (which is not supported by data), our review based on simulation studies presents an overview of when one method is expected to perform better. With this article, we hope that we managed to provide a toolbox and context‐dependent advice.

ACKNOWLEDGEMENTS

We would like to thank Reviewers for taking the time and effort necessary to review the manuscript. We sincerely appreciate all valuable comments and suggestions, which helped us to improve the quality of the manuscript.

1.

sim9628-gra-1001-b

sim9628-gra-1002-b

sim9628-gra-1003-b

sim9628-gra-1004-b

TABLE A1.

Measured confounding methods

Method Description of the method Study References
1 PS matching (N = 47) a Treated and untreated individuals are matched based on their propensity score‐similarity. After creating comparable groups of treated and untreated individuals the effect of the treatment can be estimated. 2, 4, 12, 13, 16, 20, 21, 22, 27, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 42, 44, 49, 55, 57, 58, 59, 60, 61, 62, 71, 72, 74, 75, 79, 85, 87, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155
2 IPTW (N = 30) a With the help of re‐weighting by the inverse probability of receiving the treatment, a synthetic sample is created which is representative of the population and in which treatment assignment is independent of the observed baseline covariates. Over‐represented groups are downweighted and underrepresented groups are upweighted. 2, 4, 7, 8, 13, 20, 21, 29, 37, 38, 39, 42, 44, 49, 51, 60, 61, 64, 69, 70, 71, 72, 83“74, 79, 147, 150, 153, 156, 157
3 Overlap weights (N = 4) a Overlap weights were developed to overcome the limitations of truncation and trimming for IPTW, when some individual PSs approach 0 or 1. 7, 8, 42, 43
4 Matching weights (N = 2) a Matching weights is an analogue weighting method for IPTW, when some individual PSs approach 0 or 1. 42, 44
5 Covariate adjustment using PS (N = 13) a The estimated PS is included as covariate in a regression model of the treatment. 20, 21, 22, 27, 37, 49, 53, 146, 147, 155, 158, 159
6 PS stratification (N = 26) a First the subjects are grouped into strata based upon their PS. Then, the treatment effect is estimated within each PS stratum, and the ATE is computed as a weighted mean of the stratum specific estimates. 16, 20, 21, 22, 24, 27, 37, 38, 39, 47, 48, 49, 50, 51, 53, 69, 76, 78, 79, 84, 85, 152, 153, 155, 156, 160
7 GAM (N = 1) GAMs provide an alternative for traditional PS estimation by replacing the linear component of a logistic regression with a flexible additive function. 62
8 GBM (N = 3) GBM trees provide an alternative for traditional PS estimation by estimating the function of covariates in a more flexible manner than logistic regression by averaging the PSs of small regression trees. 40, 62, 63
9 Genetic matching (N = 7) This matching method algorithmically optimizes covariate balance and avoids the process of iteratively modifying the PS model. 12, 38, 57, 58, 59, 60, 61
10 Covariate‐balancing PS (N = 5) Models treatment assignment while optimizing the covariate balance. The method exploits the dual characteristics of the PS as a covariate balancing score and the conditional probability of treatment assignment. 63, 64, 65, 66, 80
11 DR estimation (N = 13) Combines outcome regression with with a model for the treatment (eg, weighting by the PS) such that the effect estimator is robust to misspecification of one (but not both) of these models. 21, 39, 40, 49, 68, 69, 70, 72, 74, 75, 156, 157, 161
11.1 AIPTW (N = 8) This estimator achieves the doubly‐robust property by combining outcome regression with weighting by the PS. 21, 39, 40, 69, 70, 72, 74, 75
11.2 Stratified DR estimator (N = 1) Hybrid DR method of outcome regression with PS weighting and stratification. 40
11.3 TMLE (N = 2) Semi‐parametric double‐robust method that allows for flexible estimation using (nonparametric) machine‐learning methods. 2, 68, 74, 75, 83
11.4 Collaborative TMLE (N = 1) Data‐adaptive estimation method for TMLE. 75
12.1 One step joint Bayesian PS (N = 3) Jointly estimates quantities in the PS and outcome stages. 76, 78, 79
12.2 Two‐step Bayesian approach (N = 2) b A two‐step modeling method is using the Bayesian PS model in the first step, followed by a Bayesian outcome model in the second step. 79, 80
12.3 Bayesian model averaging (N = 1) b Fully Bayesian model averaging approach. 80
12.4 An's intermediate approach (N = 2) b Not fully Bayesian insofar as the outcome equation in An's approach is frequentist. 77, 79
13 G‐computation (N = 4) The method interprets counterfactual outcomes as missing data and uses a prediction model to obtain potential outcomes under different treatment scenarios. The entire set of predicted outcomes is then regressed on the treatment to obtain the coefficient of the effect estimate. 2, 75, 81, 83
14 Prognostic scores (N = 7) b Prognostic scores are considered to be the prognostic analog of the PS methods. the prognostic score includes covariates based on their predictive power of the response, the PS includes covariates that predict treatment assignment. 29, 30, 64, 84, 85, 86, 87
a

These remarks are based on PS estimated with logistic regression.

b

Can be implemented for PS stratification, weighting, and optimal full matching.

TABLE A2.

Unmeasured confounding methods

Method Description of the method Study References
1 IV approach (N = 17) Post‐randomization can be achieved using a sufficiently strong instrument. IV is correlated with the treatment and only affects the outcome through the treatment. 88, 89, 90, 91, 92, 94, 95, 97, 98, 100, 103, 105, 154, 159, 162, 163, 164
1.1 2SLS (N = 11) Linear estimator of the IV method. Uses linear probability for binary outcome and linear regression for continuous outcome. 88, 89, 91, 94, 95, 97, 98, 100, 105, 162, 163
1.2 2SPS (N = 5) Non‐parametric estimator of the IV method. Logistic regression is used for both the first and second stages of 2SPS procedure. The predicted or residual values from the first stage logistic regression of treatment on the IV are used as covariates in the second stage logistic regression: the predicted value of treatment replaces the observed treatment for 2SPS. 99, 100, 101, 102, 103
1.3 2SRI (N = 8) Semi‐parametric estimator of the IV method. Logistic regression is used for both the first and second stages of the 2SRI procedure. The predicted or residual values from the first stage logistic regression of treatment on the IV are used as covariates in the second stage logistic regression. 90, 97, 99, 100, 101, 102, 103, 105
1.4 IV based on generalized structural mean model (GSMM) (N = 1) Semi‐parametric models that use instrumental variables to identify causal parameters. IV approach 102
2 Instrumental PS (Matching enhanced IV) (N = 2) Reduces the dimensionality of the measured confounders, but it also deals with unmeasured confounders by the use of an IV. 106, 107
3 DiD (N = 7) DiD method uses the assumption that without the treatment the average outcomes for the treated and control groups would have followed parallel trends over time. The design measures the effect of a treatment as the relative change in the outcomes between individuals in the treatment and control groups over time. 109, 110, 111, 112, 113, 114, 115
4 Matching combined with DiD (N = 6) Alternative approach to DiD. (2) Uses matching to balance the treatment and control groups according to pre‐treatment outcomes and covariates 109, 110, 111, 114, 115, 124
5 SCM (N = 7) This method constructs a comparator, the synthetic control, as a weighted average of the available control individuals. The weights are chosen to ensure that, prior to the treatment, levels of covariates and outcomes are similar over time to those of the treated unit. 114, 121, 122, 123, 124, 126, 127
5.1 Imperfect SCM (N = 1) Extension of SCM method with relaxed assumptions that allow outcomes to be functions of transitory shocks. 121
5.2 Generalized SCM (N = 2) Combines SC with fixed effects. 114, 126
5.3 Synthetic DiD (N = 1) Both unit and time fixed effects, which can be interpreted as the time‐weighted version of DiD. 127
6 LDV regression approach (N = 1) Adjusts for pre‐treatment outcomes and covariates with a parametric regression model. Alternative approach to DiD. 114
7 Trend‐in‐trend (N = 1) The trend‐in‐trend design examines time trends in outcome as a function of time trends in treatment across strata with different time trends in treatment. 128
8 PERR (N = 3) PERR adjustment is a type of self‐controlled design in which the treatment effect is estimated by the ratio of two rate ratios (RRs): RR after initiation of treatment and the RR prior to initiation of treatment. 129, 130, 131
9 PS calibration (N = 1) Combines PS and regression calibration to address confounding by variables unobserved in the main study by using variables observed in a validation study. 132
10 RD (N = 4) Method used for policy analysis. People slightly below and above the threshold for being exposed to a treatment are compared. 9, 133, 134, 165

Varga AN, Guevara Morel AE, Lokkerbol J, van Dongen JM, van Tulder MW, Bosmans JE. Dealing with confounding in observational studies: A scoping review of methods evaluated in simulation studies with single‐point exposure. Statistics in Medicine. 2023;42(4):487–516. doi: 10.1002/sim.9628

Funding information ZonMw, Grant/Award Number: 91717368

ENDNOTES

*

Please note that the box contains only the main estimands that were identified in at least three of the reviewed articles. Some estimands are excluded from this due to their rare occurrence (eg, Super‐local average treatment effect (SLATE), Sample average treatment effect (SATE), or Average treatment effect in the reference population (ATR)).

†

Please note that this has also been extensively studied in simulation studies in RCT context. Prognostic covariates not correlated with treatment were studied by Williamson et al and Zeng et al. 25 , 26

‡

“Full matching constructs strata consisting of either one treated subject and at least one control subject or one control subject and at least one treated subject.” 13

§

“Optimal matching method that maximizes the cardinality or size of the matched sample subject to constraints on covariate balance.” 36

Which can be achieved by optimizing the number of trees.

#

Estimator that respects the global constraints of the statistical model . 73

‖

This means that the models for the outcome variable and treatment assignment are estimated simultaneously rather than sequentially.

**

Instrument can only be related to the outcome through the treatment. 88 Assuming perfect compliance, the randomization mechanism of an RCT can be interpreted as being the “perfect” instrument. 89 Commonly used instruments are, for instance, geographical variation, calendar time, and physician's prescription preference. 90

††

The probability of dichotomous events is often modeled as a linear function of covariates. 90

‡‡

The first stage is the same as for 2SPS. 99

§§

This assumption is necessary for estimating ATE. 99 Heterogenous treatment effect implies that the treatment effect is homogenous across heterogeneous patient characteristics. 104

¶¶

The conditions are the following: “(i) when the null hypothesis is true; (ii) when the outcome model is collapsible; or (iii) when estimating the nonnull causal effect from Cox or logistic regression models, the strong and unrealistic assumption that the effect of the unmeasured covariates on the treatment is proportional to their effect on the outcome needs to hold.”

##

The parallel trends assumption is often substituted with an empirical test of parallel trends in the pre‐treatment period, which ensures that pre‐treatment, all unobserved factors that affect outcomes, are equal across groups. This does not warrant that the parallel trend effect does not change over time, making the common shocks assumption necessary (see Appendix Box C). Ryan et al developed a checklist for the application of DiD that contains (a) the critical points and assumptions of the analysis that need to be assessed, (b) the recommendation on how to test, and (c) the consequences of the violation of each assumption. 110

‖‖

Correction for pre‐treatment outcome level differences requires the assumption that conditional on pre‐treatment outcome levels and covariates, potential outcomes are independent of treatment status For this assumption to hold, matching has to balance out both level (eg, time‐invariant fixed effect) and trend effects (eg, time‐varying fixed effect). 111

***

The method requires the assumption that the pre‐treatment outcomes of the treated can be reconstructed as a weighted average of pre‐treatment outcomes of the controls (ie, that the treated unit belongs to the convex hull of the control units). 121

†††

As long as characteristics related to outcomes are smoothly distributed around the treatment threshold, differences in outcomes across the threshold can be attributed to the treatment.

‡‡‡

This can be distinguished with the help of graphical analysis. 134

§§§

Robin et al suggested basing the decision on the examination of graphical illustration of the outcome variable against the rating variable.

¶¶¶

Individuals who comply with the treatment if and only if assigned to it.

###

Represents the expected difference of treatment effect on the outcome if treatment uptake randomly assigned.

Contributor Information

Anita Natalia Varga, Email: a.n.varga@vu.nl.

Alejandra Elizabeth Guevara Morel, Email: a.e.guevaramorel@vu.nl.

DATA AVAILABILITY STATEMENT

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.

REFERENCES

  • 1. Austin P. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivar Behav Res. 2011;46:399‐424. doi: 10.1080/00273171.2011.568786 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Chatton A, Le Borgne F, Leyrat C, et al. G‐computation, propensity score‐based methods, and targeted maximum likelihood estimator for causal inference with different covariates sets: a comparative simulation study. Sci Rep. 2020;10(1):9219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Daniel R, Cousens S, De Stavola B, Kenward M, Sterne J. Methods for Dealing with Time‐Dependent Confounding. Stat Med. 2013;32(9):1584‐1618. doi: 10.1002/sim.5686 [DOI] [PubMed] [Google Scholar]
  • 4. Pirracchio R, Carone M, Rigon MR, Caruana E, Mebazaa A, Chevret S. Propensity score estimators for the average treatment effect and the average treatment effect on the treated may yield very different estimates. Stat Methods Med Res. 2016;25(5):1938‐1954. [DOI] [PubMed] [Google Scholar]
  • 5. Imbens GW, Angrist JD. Identification and estimation of local average treatment effects. Econometrica. 1994;62(2):467‐475. [Google Scholar]
  • 6. Wang A, Nianogo RA, Arah OA. G‐computation of average treatment effects on the treated and the untreated. BMC Med Res Methodol. 2017;17(1):3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Li F, Thomas LE, Li F. Addressing extreme propensity scores via the overlap weights. Am J Epidemiol. 2018;188(1):250‐257. doi: 10.1093/aje/kwy201 [DOI] [PubMed] [Google Scholar]
  • 8. Li F, Morgan KL, Zaslavsky AM. Balancing covariates via propensity score weighting. J Am Stat Assoc. 2018;113(521):390‐400. doi: 10.1080/01621459.2016.1260466 [DOI] [Google Scholar]
  • 9. Imbens G, Kalyanaraman K. Optimal bandwidth choice for the regression discontinuity estimator. Rev Econom Stud. 2011;79(3):933‐959. doi: 10.1093/restud/rdr043 [DOI] [Google Scholar]
  • 10. Wendling T, Jung K, Callahan A, Schuler A, Shah NH, Gallego B. Comparing methods for estimation of heterogeneous treatment effects using observational data from health care databases. Stat Med. 2018;37(23):3309‐3324. doi: 10.1002/sim.7820 [DOI] [PubMed] [Google Scholar]
  • 11. Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38(11):2074‐2102. doi: 10.1002/sim.8086 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Diamond A, Sekhon JS. Genetic matching for estimating causal effects: a general multivariate matching method for achieving balance in observational studies. Rev Econ Stat. 2013;95(3):932‐945. doi: 10.1162/REST_a_00318 [DOI] [Google Scholar]
  • 13. Austin P, Stuart E. Estimating the effect of treatment on binary outcomes using full matching on the propensity score. Stat Methods Med Res. 2017;26(6):2505‐2525. doi: 10.1177/0962280215601134 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Cole S, Hernán M. Constructing inverse probability weights for marginal structural models. Am J Epidemiol. 2008;168:656‐664. doi: 10.1093/aje/kwn164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Austin PC, Stuart EA. Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Stat Med. 2015;34(28):3661‐3679. doi: 10.1002/sim.6607 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Austin P, Grootendorst P, Anderson G. A comparison of the ability of different propensity score models to balance measured variables between treated and untreated subjects: a Monte Carlo study. Stat Med. 2007;26:734‐753. doi: 10.1002/sim.2580 [DOI] [PubMed] [Google Scholar]
  • 17. Zhang Z, Kim HJ, Lonjon G, Zhu Y. Balance diagnostics after propensity score matching. Ann Transl Med. 2019;7(1):16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Stuart E. Matching methods for causal inference: a review and a look forward. Stat Sci. 2010;25:1‐21. doi: 10.1214/09-STS313 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Harder V, Stuart E, Anthony J. Propensity score techniques and the assessment of measured covariate balance to test causal associations in psychological research. Psychol Methods. 2010;15:234‐249. doi: 10.1037/a0019623 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Franklin J, Eddings W, Austin P, Stuart E, Schneeweiss S. Comparing the performance of propensity score methods in healthcare database studies with rare outcomes. Stat Med. 2017;36(12):1946‐1963. doi: 10.1002/sim.7250 [DOI] [PubMed] [Google Scholar]
  • 21. Basu A, Polsky D, Manning WG. Estimating treatment effects on healthcare costs under exogeneity: is there a ‘Magic Bullet’? Health Serv Outcomes Res Methodol. 2011;11(1‐2):1‐26. doi: 10.1007/s10742-011-0072-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Austin PC, Grootendorst P, Normand SL, Anderson GM. Conditioning on the propensity score can result in biased estimation of common measures of treatment effect: a Monte Carlo study. Stat Med. 2007;26(4):754‐768. [DOI] [PubMed] [Google Scholar]
  • 23. Brookhart MA, Schneeweiss S, Rothman KJ, Glynn RJ, Avorn J, Stürmer T. Variable selection for propensity score models. Am J Epidemiol. 2006;163(12):1149‐1156. doi: 10.1093/aje/kwj149 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Adelson JL, McCoach DB, Rogers HJ, Adelson J, Sauer TM. Developing and applying the propensity score to make causal inferences: variable selection and stratification. Front Psychol. 2017;8:1413. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Williamson EJ, Forbes A, White IR. Variance reduction in randomised trials by inverse probability weighting using the propensity score. Stat Med. 2014;33(5):721‐737. doi: 10.1002/sim.5991 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Zeng S, Li F, Wang R, Li F. Propensity score weighting for covariate adjustment in randomized clinical trials. Stat Med. 2021;40(4):842‐858. [DOI] [PubMed] [Google Scholar]
  • 27. Austin P. The performance of different propensity score methods for estimating marginal odds ratios. Stat Med. 2007;26:3078‐3094. doi: 10.1002/sim.2781 [DOI] [PubMed] [Google Scholar]
  • 28. Austin PC, Stuart EA. The performance of inverse probability of treatment weighting and full matching on the propensity score in the presence of model misspecification when estimating the effect of treatment on survival outcomes. Stat Methods Med Res. 2017;26(4):1654‐1670. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Stuart EA, Lee BK, Leacy FP. Prognostic score‐based balance measures can be a useful diagnostic for propensity score methods in comparative effectiveness research. J Clin Epidemiol. 2013;66(8 Suppl):S84‐S90. doi: 10.1016/j.jclinepi.2013.01.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Nguyen TL, Debray TP. The use of prognostic scores for causal inference with general treatment regimes. Stat Med. 2019;38(11):2013‐2029. doi: 10.1002/sim.8084 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Rassen J, Shelat A, Franklin J, Glynn R, Rothman K, Schneeweiss S. One‐to‐Many Propensity Score Matching in Cohort Studies. Pharmacoepidemiol Drug Saf. 2012;21(Suppl 2):69‐80. doi: 10.1002/pds.3263 [DOI] [PubMed] [Google Scholar]
  • 32. Austin P. Statistical criteria for selecting the optimal number of untreated subjects matched to each treated subject when using many‐to‐one matching on the propensity score. Am J Epidemiol. 2010;172:1092‐1097. doi: 10.1093/aje/kwq224 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Austin P. A comparison of 12 algorithms for matching on the propensity score. Stat Med. 2014;33(6):1057‐1069. doi: 10.1002/sim.6004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Austin P. Optimal caliper widths for propensity‐score matching when estimating differences in means and differences in proportions in observational studies. Pharm Stat. 2011;10:150‐161. doi: 10.1002/pst.433 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Lunt M. Selecting an appropriate caliper can be essential for achieving good balance with propensity score matching. Am J Epidemiol. 2014;179(2):226‐235. doi: 10.1093/aje/kwt212 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Resa M, Zubizarreta J. Evaluation of subset matching methods and forms of covariate balance. Stat Med. 2016;35(27):4961‐4979. doi: 10.1002/sim.7036 [DOI] [PubMed] [Google Scholar]
  • 37. Austin PC. The relative ability of different propensity score methods to balance measured covariates between treated and untreated subjects in observational studies. Med Decis Making. 2009;29(6):661‐677PMID: 19684288. doi: 10.1177/0272989X09341755 [DOI] [PubMed] [Google Scholar]
  • 38. Zagar A, Kadziola Z, Lipkovich I, Faries D. Evaluating different strategies for estimating treatment effects in observational studies. J Biopharm Stat. 2017;27(3):535‐553. doi: 10.1080/10543406.2017.1289953 [DOI] [PubMed] [Google Scholar]
  • 39. Kang JDY, Schafer JL. Demystifying double robustness: a comparison of alternative strategies for estimating a population mean from incomplete data. Stat Sci. 2007;22(4):523‐539. doi: 10.1214/07-STS227 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Hattori S, Henmi M. Stratified doubly robust estimators for the average causal effect. Biometrics. 2014;70(2):270‐277. [DOI] [PubMed] [Google Scholar]
  • 41. Xiao Y, Moodie E, Abrahamowicz M. Comparison of approaches to weight truncation for marginal structural Cox models. Epidemiol Methods. 2013;2:1‐20. doi: 10.1515/em-2012-0006 [DOI] [Google Scholar]
  • 42. Mao H, Li L, Greene T. Propensity score weighting analysis and treatment effect discovery. Stat Methods Med Res. 2018;28(8):2439‐2454. doi: 10.1177/0962280218781171 [DOI] [PubMed] [Google Scholar]
  • 43. Yang S, Lorenzi E, Papadogeorgou G, Wojdyla DM, Li F, Thomas LE. Propensity score weighting for causal subgroup analysis. Stat Med. 2021;40(19):4294‐4309. doi: 10.1002/sim.9029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Li L, Greene T. A weighting analogue to pair matching in propensity score analysis. Int J Biostat. 2013;9(2):215‐234. [DOI] [PubMed] [Google Scholar]
  • 45. Elze MC, Gregson J, Baber U, et al. Comparison of propensity score methods and covariate adjustment: evaluation in 4 cardiovascular studies. J Am Coll Cardiol. 2017;69(3):345‐357. [DOI] [PubMed] [Google Scholar]
  • 46. Rubin DB. On principles for modeling propensity scores in medical research. Pharmacoepidemiol Drug Saf. 2004;13(12):855‐857. doi: 10.1002/pds.968 [DOI] [PubMed] [Google Scholar]
  • 47. Neuhäuser M, Thielmann M, Ruxton GD. The number of strata in propensity score stratification for a binary outcome. Arch Med Sci AMS. 2018;14(3):695‐700. doi: 10.5114/aoms.2016.61813 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Rudolph KE, Colson KE, Stuart EA, Ahern J. Optimally combining propensity score subclasses. Stat Med. 2016;35(27):4937‐4947. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Austin PC. The performance of different propensity‐score methods for estimating differences in proportions (risk differences or absolute risk reductions) in observational studies. Stat Med. 2010;29(20):2137‐2148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Hullsiek KH, Louis TA. Propensity score modeling strategies for the causal analysis of observational data. Biostatistics. 2002;3(2):179‐193. doi: 10.1093/biostatistics/3.2.179 [DOI] [PubMed] [Google Scholar]
  • 51. Linden A. A comparison of approaches for stratifying on the propensity score to reduce bias. J Eval Clin Pract. 2017;23(4):690‐696. doi: 10.1111/jep.12701 [DOI] [PubMed] [Google Scholar]
  • 52. Drake C. Effects of misspecification of the propensity score on estimators of treatment effect. Biometrics. 1993;49(4):1231‐1236. [Google Scholar]
  • 53. Myers JA, Louis TA. Comparing treatments via the propensity score: stratification or modeling? Health Serv Outcomes Res Methodol. 2012;12(1):29‐43. doi: 10.1007/s10742-012-0080-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Westreich DJ, Lessler J, Funk MLJ. Propensity score estimation: machine learning and classification methods as alternatives to logistic regression. 2010. [DOI] [PMC free article] [PubMed]
  • 55. Lehrer S, Kordas G. Matching using semiparametric propensity scores. Empirical Econom. 2013;44:13‐45. doi: 10.1007/s00181-012-0591-3 [DOI] [Google Scholar]
  • 56. Horowitz JL. 2 Semiparametric and nonparametric estimation of quantal response models. Handbook Stat. 1993;11:45‐72. [Google Scholar]
  • 57. Tsai KT, Peace K. Genetic matching: an efficient algorithm to adjust covariate imbalance for data analysis and modeling. http://worldcomp‐proceedings.com/proc/p2011/BIC3060.pdf; 2012.
  • 58. Sekhon J, Grieve R. A new non‐parametric matching method for covariate adjustment with application to economic evaluation. Experiments in Political Science 2008 Conference Paper. SSRN. 2009. https://ssrn.com/abstract=1301767. doi: 10.2139/ssrn.1301767 [DOI]
  • 59. Sekhon JS, Grieve RD. A matching method for improving covariate balance in cost‐effectiveness analyses. Health Econ. 2012;21(6):695‐714. [DOI] [PubMed] [Google Scholar]
  • 60. Kreif N, Grieve R, Radice R, Sadique Z, Ramsahai R, Sekhon JS. Methods for estimating subgroup effects in cost‐effectiveness analyses that use observational data. Med Decis Making. 2012;32(6):750‐763. [DOI] [PubMed] [Google Scholar]
  • 61. Radice R, Ramsahai R, Grieve R, Kreif N, Sadique Z, Sekhon JS. Evaluating treatment effectiveness in patient subgroups: a comparison of propensity score methods with an automated matching approach. Int J Biostat. 2012;8(1):25. [DOI] [PubMed] [Google Scholar]
  • 62. Woo MJ, Reiter JP, Karr AF. Estimation of propensity scores using generalized additive models. Stat Med. 2008;27(19):3805‐3816. doi: 10.1002/sim.3278 [DOI] [PubMed] [Google Scholar]
  • 63. Setodji C, Mccaffrey D, Burgette L, Almirall D, Griffin BA. The right tool for the job: choosing between covariate‐balancing and generalized boosted model propensity scores. Epidemiology. 2017;28:1. doi: 10.1097/EDE.0000000000000734 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Wyss R, Ellis A, Brookhart M, et al. The role of prediction modeling in propensity score estimation: an evaluation of logistic regression, bCART, and the covariate‐balancing propensity score. Am J Epidemiol. 2014;180(6):645‐655. doi: 10.1093/aje/kwu181 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Imai K, Ratkovic M. Covariate balancing propensity score. J R Stat Soc Series B Stat Methodology. 2014;76(1):243‐263. doi: 10.1111/rssb.12027 [DOI] [Google Scholar]
  • 66. Xie Y, Zhu Y, Cotton CA, Wu P. A model averaging approach for estimating propensity scores by optimizing balance. Stat Methods Med Res. 2019;28(1):84‐101. doi: 10.1177/0962280217715487 [DOI] [PubMed] [Google Scholar]
  • 67. Imai K, Dyk vDA. Causal inference with general treatment regimes: generalizing the propensity score. J Am Stat Assoc. 2004;99(467):854‐866. doi: 10.1198/016214504000001187 [DOI] [Google Scholar]
  • 68. Luque‐Fernandez MA, Schomaker M, Rachet B, Schnitzer ME. Targeted maximum likelihood estimation for a binary treatment: a tutorial. Stat Med. 2018;37(16):2530‐2546. doi: 10.1002/sim.7628 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Li J, Handorf E, Bekelman J, Mitra N. Propensity score and doubly robust methods for estimating the effect of treatment on censored cost. Stat Med. 2016;35(12):1985‐1999. doi: 10.1002/sim.6842 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Funk MJ, Westreich D, Wiesen C, Stürmer T, Brookhart MA, Davidian M. Doubly robust estimation of causal effects. Am J Epidemiol. 2011;173(7):761‐767. doi: 10.1093/aje/kwq439 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71. Schafer J, Kang J. Average Causal Effects From Nonrandomized Studies: A Practical Guide and Simulated Example. Psychol Methods. 2009;13:279‐313. doi: 10.1037/a0014268 [DOI] [PubMed] [Google Scholar]
  • 72. Waernbaum I. Model misspecification and robustness in causal inference: comparing matching with doubly robust estimation. Stat Med. 2012;31(15):1572‐1581. doi: 10.1002/sim.4496 [DOI] [PubMed] [Google Scholar]
  • 73. van der Laan MJ, Rose S. Causal Inference for Observational and Experimental Data: Targeted Learning. Springer Series in Statistics. New York: Springer; 2011. ISBN 9781441997821. [Google Scholar]
  • 74. Balzer L, Ahern J, Galea S, Laan vM. Estimating effects with rare outcomes and high dimensional covariates: knowledge is power. Epidemiol Methods. 2016;5(1):1‐18. doi: 10.1515/em-2014-0020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Lendle SD, Fireman B, Laan vMJ. Targeted maximum likelihood estimation in safety analysis. J Clin Epidemiol. 2013;66(8 Suppl):S91‐S98. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76. Zigler CM, Watts K, Yeh RW, Wang Y, Coull BA, Dominici F. Model Feedback in Bayesian Propensity Score Estimation. Biometrics. 2013;69(1):263‐273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. An W. Bayesian propensity score estimators: incorporating uncertainties in propensity scores into causal inference. Sociolog Methodol. 2010;40(1):151‐189. doi: 10.1111/j.1467-9531.2010.01226.x [DOI] [Google Scholar]
  • 78. McCandless LC, Gustafson P, Austin PC. Bayesian propensity score analysis for observational data. Stat Med. 2009;28(1):94‐112. doi: 10.1002/sim.3460 [DOI] [PubMed] [Google Scholar]
  • 79. Kaplan D, Chen J. A two‐step Bayesian approach for propensity score analysis: simulations and case study. Psychometrika. 2012;77(3):581‐609. doi: 10.1007/s11336-012-9262-8 [DOI] [PubMed] [Google Scholar]
  • 80. Kaplan D, Chen J. Bayesian model averaging for propensity score analysis. Multivar Behav Res. 2014;49(6):505‐517. doi: 10.1080/00273171.2014.928492 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81. Snowden JM, Rose S, Mortimer KM. Implementation of G‐computation on a simulated data set: demonstration of a causal inference technique. Am J Epidemiol. 2011;173(7):731‐738. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82. Porter KE, Gruber S, Laan vMJ, Sekhon JS. The relative performance of targeted maximum likelihood estimators. Int J Biostat. 2011;7(1). doi: 10.2202/1557-4679.1308 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83. Schuler MS, Rose S. Targeted maximum likelihood estimation for causal inference in observational studies. Am J Epidemiol. 2017;185(1):65‐73. doi: 10.1093/aje/kww165 [DOI] [PubMed] [Google Scholar]
  • 84. Hansen BB. The prognostic analogue of the propensity score. Biometrika. 2008;95(2):481‐488. doi: 10.1093/biomet/asn004 [DOI] [Google Scholar]
  • 85. Leacy F, Stuart E. On the joint use of propensity and prognostic scores in estimation of the average treatment effect on the treated: a simulation study. Stat Med. 2014;33(20):3488‐3508. doi: 10.1002/sim.6030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Arbogast PG, Ray WA. Performance of disease risk scores, propensity scores, and traditional multivariable outcome regression in the presence of multiple confounders. Am J Epidemiol. 2011;174(5):613‐620. [DOI] [PubMed] [Google Scholar]
  • 87. Wyss R, Ellis AR, Brookhart MA, et al. Matching on the disease risk score in comparative effectiveness research of new treatments. Pharmacoepidemiol Drug Saf. 2015;24(9):951‐961. doi: 10.1002/pds.3810 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88. Uddin MJ, Groenwold R, Boer A, et al. Performance of instrumental variable methods in cohort and nested case‐control studies: a simulation study. Pharmacoepidemiol Drug Saf. 2014;23(2):165‐177. doi: 10.1002/pds.3555 [DOI] [PubMed] [Google Scholar]
  • 89. Johnston KM, Gustafson P, Levy AR, Grootendorst P. Use of instrumental variables in the analysis of generalized linear models in the presence of unmeasured confounding with applications to epidemiological research. Stat Med. 2008;27(9):1539‐1556. [DOI] [PubMed] [Google Scholar]
  • 90. Koladjo F, Escolano S, Tubert‐Bitter P. Instrumental variable analysis in the context of dichotomous outcome and exposure with a numerical experiment in pharmacoepidemiology. BMC Med Res Methodol. 2018;18:61. doi: 10.1186/s12874-018-0513-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91. Crown WH, Henk HJ, Vanness DJ. Some cautions on the use of instrumental variables estimators in outcomes research: how bias in instrumental variables estimators is affected by instrument strength, instrument contamination, and sample size. Value Health. 2011;14(8):1078‐1084. doi: 10.1016/j.jval.2011.06.009 [DOI] [PubMed] [Google Scholar]
  • 92. Li Y, Lee Y, Wolfe R, et al. On a preference‐based instrumental variable approach in reducing unmeasured confounding‐by‐indication. Stat Med. 2015;34(7):1150‐1168. doi: 10.1002/sim.6404 [DOI] [PubMed] [Google Scholar]
  • 93. Angrist JD. Estimation of limited dependent variable models with dummy endogenous regressors: simple strategies for empirical practice. J Bus Econom Stat. 2001;19(1):2‐16. [Google Scholar]
  • 94. Ionescu‐Ittu R, Delaney J, Abrahamowicz M. Bias‐variance trade‐off in pharmacoepidemiological studies using physician‐preference‐based instrumental variables: a simulation study. Pharmacoepidemiol Drug Saf. 2009;18:562‐571. doi: 10.1002/pds.1757 [DOI] [PubMed] [Google Scholar]
  • 95. Boef AG, Dekkers OM, Vandenbroucke JP, Cessie lS. Sample size importantly limits the usefulness of instrumental variable methods, depending on instrument strength and level of confounding. J Clin Epidemiol. 2014;67(11):1258‐1264. [DOI] [PubMed] [Google Scholar]
  • 96. Constantine Gatsonis SCM. Methods in Comparative Effectiveness Research. 1st ed. Boca Raton, FL: CRC Press; 2017. [Google Scholar]
  • 97. Basu A, Coe N, Chapman C. 2SLS versus 2SRI: appropriate methods for rare outcomes and/or rare exposures. Health Econ. 2018;27(6):937‐955. doi: 10.1002/hec.3647 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98. Bhattacharya J, Goldman D, McCaffrey D. Estimating probit models with self‐selected treatments. Stat Med. 2006;25(3):389‐413. doi: 10.1002/sim.2226 [DOI] [PubMed] [Google Scholar]
  • 99. Terza JV, Basu A, Rathouz PJ. Two‐stage residual inclusion estimation: addressing endogeneity in health econometric modeling. J Health Econ. 2008;27(3):531‐543. doi: 10.1016/j.jhealeco.2007.09.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100. Zhang Z, Uddin MJ, Cheng J, Huang T. Instrumental variable analysis in the presence of unmeasured confounding. Ann Trans Med. 2018;6:182‐182. doi: 10.21037/atm.2018.03.37 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101. Cai B, Small DS, Have TRT. Two‐stage instrumental variable methods for estimating the causal odds ratio: analysis of bias. Stat Med. 2011;30(15):1809‐1824. doi: 10.1002/sim.4241 [DOI] [PubMed] [Google Scholar]
  • 102. Cai B, Hennessy S, Flory JH, Sha D, Have TRT, Small DS. Simulation study of instrumental variable approaches with an application to a study of the antidiabetic effect of bezafibrate. Pharmacoepidemiol Drug Saf. 2012;21(Suppl 2):114‐120. [DOI] [PubMed] [Google Scholar]
  • 103. Wan F, Small D, Mitra N. A general approach to evaluating the bias of 2‐stage instrumental variable estimators. Stat Med. 2018;37(12):1997‐2015. doi: 10.1002/sim.7636 [DOI] [PubMed] [Google Scholar]
  • 104. Velentgas P, Dreyer NA, Nourjah P, Smith SR, Torchia MM. Developing a Protocol for Observational Comparative Effectiveness Research: A User's Guide. Rockville, MD: Agency for Healthcare Research and Quality; 2013. [PubMed] [Google Scholar]
  • 105. Chapman C, Brooks J. Treatment effect estimation using nonlinear two‐stage instrumental variable estimators: another cautionary note. Health Serv Res. 2016;51(6):2375‐2394. doi: 10.1111/1475-6773.12463 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106. Cheng J, Lin W. Understanding causal distributional and subgroup effects with the instrumental propensity score. Am J Epidemiol. 2017;187(3):614‐622. doi: 10.1093/aje/kwx282 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107. Lu B, Marcus S. Evaluating long‐term effects of a psychiatric treatment using instrumental variable and matching approaches. Health Serv Outcomes Res Methodol. 2012;12:288‐301. doi: 10.1007/s10742-012-0101-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108. Sofer T, Richardson D, Colicino E, Schwartz J, Tchetgen E. On negative outcome control of unobserved confounding as a generalization of difference‐in‐differences. Stat Sci. 2016;31:348‐361. doi: 10.1214/16-STS558 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109. Daw J, Hatfield L. Matching and regression to the mean in difference‐in‐differences analysis. Health Serv Res. 2018;53(6):4138‐4156. doi: 10.1111/1475-6773.12993 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110. Ryan A, Burgess J, Dimick J. Why we should not be indifferent to specification choices for difference‐in‐differences. Health Serv Res. 2015;50(4):1211‐1235. doi: 10.1111/1475-6773.12270 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111. Lindner S, McConnell KJ. Difference‐in‐differences and matching on outcomes: a tale of two unobservables. Health Serv Outcomes Res Methodol. 2019;19(2):127‐144. doi: 10.1007/s10742-018-0189-0 [DOI] [Google Scholar]
  • 112. Bertrand M, Duflo E, Mullainathan S. How Much Should We Trust Differences‐In‐Differences Estimates? The Quarterly Journal of Economics. 2003;119:249‐275. doi: 10.1162/003355304772839588 [DOI] [Google Scholar]
  • 113. Rokicki S, Cohen J, Fink G, Salomon JA, Landrum MB. Inference with difference‐in‐differences with a small number of groups: a review, simulation study, and empirical application using share data. Med Care. 2018;56(1):97‐105. [DOI] [PubMed] [Google Scholar]
  • 114. O'Neill S, Kreif N, Grieve R, Sutton M, Sekhon JS. Estimating causal effects: considering three alternatives to difference‐in‐differences estimation. Health Serv Outcomes Res Methodol. 2016;16:1‐21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115. Ryan AM, Kontopantelis E, Linden A, Burgess JF. Now trending: coping with non‐parallel trends in difference‐in‐differences analysis. Stat Methods Med Res. 2019;28(12):3697‐3711. [DOI] [PubMed] [Google Scholar]
  • 116. Daw JR, Hatfield LA. Matching in difference‐in‐differences: between a rock and a hard place. Health Serv Res. 2018;53(6):4111‐4117. doi: 10.1111/1475-6773.13017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117. Ryan AM. Well‐Balanced or too Matchy–Matchy? The Controversy over Matching in Difference‐in‐Differences. Health Serv Res. 2018;53(6):4106‐4110. doi: 10.1111/1475-6773.13015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118. Athey S, Imbens G. The state of applied econometrics ‐ causality and policy evaluation. J Econom Perspect. 2017;31(2):3‐32. doi: 10.1257/jep.31.2.3 [DOI] [Google Scholar]
  • 119. Kreif N, Grieve R, Hangartner D, Turner A, Nikolova S, Sutton M. Examination of the synthetic control method for evaluating health policies with multiple treated units. Health Econ. 2016;25(12):1514‐1528. doi: 10.1002/hec.3258 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120. Abadie A, Diamond A, Hainmueller J. Synthetic control methods for comparative case studies: estimating the effect of California's Tobacco Control Program. J Am Stat Assoc. 2010;105:493‐505. doi: 10.1198/jasa.2009.ap08746 [DOI] [Google Scholar]
  • 121. Powell D. Imperfect Synthetic Controls: Did the Massachusetts Mealth Care Reform Save Lives? Health Economics eJournal. Santa Monica, CA: RAND; 2018. https://www.rand.org/pubs/working_papers/WR1246.html [Google Scholar]
  • 122. Ferman B, Pinto C, Possebom V. Cherry picking with synthetic controls. J Policy Anal Manage. 2020;39(2):510‐532. doi: 10.1002/pam.22206 [DOI] [Google Scholar]
  • 123. Kinn D. Synthetic control methods and big data. https://arxiv.org/pdf/1803.00096.pdf; 2018.
  • 124. Illenberger NA, Small DS, Shaw PA. Impact of regression to the mean on the synthetic control method: bias and sensitivity analysis. Epidemiology. 2020;31(6):815‐822. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125. Xu Y. Generalized synthetic control method: causal inference with interactive fixed effects models. Political Anal. 2017;25(1):57‐76. doi: 10.1017/pan.2016.2 [DOI] [Google Scholar]
  • 126. O'Neill S, Kreif N, Sutton M, Grieve R. A comparison of methods for health policy evaluation with controlled pre‐post designs. Health Serv Res. 2020;55(2):328‐338. doi: 10.1111/1475-6773.13274 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127. Arkhangelsky D, Athey S, Hirshberg DA, Imbens G, Wager S. Synthetic difference in differences. Am Econ Rev. 2021;111(12):4088‐4118. doi: 10.3386/w25532 [DOI] [Google Scholar]
  • 128. Ji X, Small DS, Leonard CE, Hennessy S. The trend‐in‐trend research design for causal inference. Epidemiology. 2017;28(4):529‐536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129. Lin NX, Henley WE. Prior event rate ratio adjustment for hidden confounding in observational studies of treatment effectiveness: a pairwise Cox likelihood approach. Stat Med. 2016;35(28):5149‐5169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130. Yu M, Xie D, Wang X, Weiner MG, Tannen RL. Prior event rate ratio adjustment: numerical studies of a statistical method to address unrecognized confounding in observational studies. Pharmacoepidemiol Drug Saf. 2012;21(Suppl 2):60‐68. [DOI] [PubMed] [Google Scholar]
  • 131. Uddin MJ, Groenwold RH, van Staa TP, et al. Performance of prior event rate ratio adjustment method in pharmacoepidemiology: a simulation study. Pharmacoepidemiol Drug Saf. 2015;24(5):468‐477. [DOI] [PubMed] [Google Scholar]
  • 132. Stürmer T, Schneeweiss S, Rothman K, Avorn J, Glynn R. Performance of propensity score calibration–a simulation study. Am J Epidemiol. 2007;165:1110‐1118. doi: 10.1093/aje/kwm074 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133. Barreca AI, Lindo JM, Waddell GR. Heaping‐induced bias in regression‐discontinuity designs. Econ Inq. 2016;54(1):268‐293. doi: 10.1111/ecin.12225 [DOI] [Google Scholar]
  • 134. Robin J, Pei Z, Marie‐Andrée S, Howard B. A Practical Guide to Regression Discontinuity. New York City: MDRC; 2012. http://bibpurl.oclc.org/web/47981. http://www.mdrc.org/sites/default/files/regression_discontinuity_full.pdf [Google Scholar]
  • 135. Streeter A, Lin N, Crathorne L, et al. Adjusting for unmeasured confounding in nonrandomized longitudinal studies: a methodological review. J Clin Epidemiol. 2017;87:23‐34. doi: 10.1016/j.jclinepi.2017.04.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136. Baiocchi M, Small D, Yang L, Polsky D, Groeneveld P. Near/far matching: a study design approach to instrumental variables. Health Serv Outcomes Res Methodol. 2012;12(4):237‐253. doi: 10.1007/s10742-012-0091-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137. Moher D, Hopewell S, Schulz KF, et al. CONSORT 2010 explanation and Elaboration: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c869. doi: 10.1136/bmj.c869 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138. Husereau D, Drummond M, Augustovski F, et al. Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) statement: updated reporting guidance for health economic evaluations. Int J Technol Assess Health Care. 2022;38(1):e13. doi: 10.1017/S0266462321001732 [DOI] [PubMed] [Google Scholar]
  • 139. Lousdal ML. An introduction to instrumental variable assumptions, validation and estimation. Emerg Themes Epidemiol. 2018;15:1. doi: 10.1186/s12982-018-0069-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140. Wang L, TE. Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables. J R Stat Soc Series B Stat Methodology. 2018;80(3):531‐550. doi: 10.1111/rssb.12262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141. Desai RJ, Franklin JM. Alternative approaches for confounding adjustment in observational studies using weighting based on the propensity score: a primer for practitioners. BMJ. 2019;367:l5657. [DOI] [PubMed] [Google Scholar]
  • 142. Brady H, Collier D, Sekhon J. The Neyman‐Rubin model of causal inference and estimation via matching methods. Oxford Handbook Political Methodol. Oxford: Oxford Academic; 2008; online ed. doi: 10.1093/oxfordhb/9780199286546.003.0011 [DOI] [Google Scholar]
  • 143. VanderWeele T, Hernán M. Causal inference under multiple versions of treatment. J Causal Inference. 2013;1:1‐20. doi: 10.1515/jci-2012-0002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144. Stock JH, Wright JH, Yogo M. A survey of weak instruments and weak identification in generalized method of moments. J Bus Econ Stat. 2002;20(4):518‐529. [Google Scholar]
  • 145. Oldenburg CE, Moscoe E, Bärnighausen TW. Regression discontinuity for causal effect estimation in epidemiology. Current Epidemiol Rep. 2016;3:233‐241. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146. Brakenhoff T, Moons K, Kluin J, Groenwold R. Investigating risk adjustment methods for health care provider profiling when observations are scarce or events rare. Health Services Insights. 2018;11:117863291878513. doi: 10.1177/1178632918785133 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147. Ertefaie A, Stephens D. Comparing approaches to causal inference for longitudinal data: inverse probability weighting versus propensity scores. Int J Biostat. 2010;6:14. doi: 10.2202/1557-4679.1198 [DOI] [PubMed] [Google Scholar]
  • 148. Thoemmes FJ, West SG. The use of propensity scores for nonrandomized designs with clustered data. Multivariate Behav Res. 2011;46(3):514‐543. [DOI] [PubMed] [Google Scholar]
  • 149. Austin PC, Small DS. The use of bootstrapping when using propensity‐score matching without replacement: a simulation study. Stat Med. 2014;33(24):4306‐4319. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 150. Pirracchio R, Resche‐Rigon M, Chevret S. Evaluation of the propensity score methods for estimating marginal odds ratios in case of small sample size. BMC Med Res Methodol. 2012;12:70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151. Setoguchi S, Schneeweiss S, Brookhart MA, Glynn RJ, Cook EF. Evaluating uses of data mining techniques in propensity score estimation: a simulation study. Pharmacoepidemiol Drug Saf. 2008;17(6):546‐555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152. Desai R, Rothman K, Bateman B, Hernandez‐Diaz S, Huybrechts K. A propensity‐score‐based fine stratification approach for confounding adjustment when exposure is infrequent. Epidemiology. 2016;28:1. doi: 10.1097/EDE.0000000000000595 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153. Brooks JM, Ohsfeldt RL. Squeezing the balloon: propensity scores and unmeasured covariate balance. Health Serv Res. 2013;48(4):1487‐1507. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154. Cnossen MC, Cnossen MC, van Essen TA, Ceyisakar IE, et al. Adjusting for confounding by indication in observational studies: a case study in traumatic brain injury. Clin Epidemiol. 2018;10:841‐852. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155. Austin PC. The performance of different propensity‐score methods for estimating relative risks. J Clin Epidemiol. 2008;61(6):537‐545. [DOI] [PubMed] [Google Scholar]
  • 156. Lunceford JK. Stratification and weighting via the propensity score in estimation of causal treatment effects: a comparative study. Stat Med. 2017;36(14):2320. [DOI] [PubMed] [Google Scholar]
  • 157. Linden A. Improving causal inference with a doubly robust estimator that combines propensity score stratification and weighting. J Eval Clin Pract. 2017;23(4):697‐702. [DOI] [PubMed] [Google Scholar]
  • 158. Mitra N, Indurkhya A. A propensity score approach to estimating the cost–effectiveness of medical therapies from observational data. Health Econ. 2005;14(8):805‐815. doi: 10.1002/hec.987 [DOI] [PubMed] [Google Scholar]
  • 159. Landsman V, Pfeiffer RM. On estimating average effects for multiple treatment groups. Stat Med. 2013;32(11):1829‐1841. [DOI] [PubMed] [Google Scholar]
  • 160. Cepeda M, Boston R, Farrar J, Strom B. Comparison of logistic regression versus propensity score when the number of events is low and there are multiple confounders. Am J Epidemiol. 2003;158:280‐287. doi: 10.1093/aje/kwg115 [DOI] [PubMed] [Google Scholar]
  • 161. Liu J, Ma Y, Wang L. An alternative robust estimator of average treatment effect in causal inference. Biometrics. 2018;74(3):910‐923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 162. Ertefaie A, Small D, Flory J, Hennessy S. Selection bias when using instrumental variable methods to compare two treatments but more than two treatments are available. Int J Biostat. 2016;12(1):219‐232. doi: 10.1515/ijb-2015-0006 [DOI] [PubMed] [Google Scholar]
  • 163. Swanson S, Robins J, Miller M, Hernán M. Selecting on treatment: a pervasive form of bias in instrumental variable analyses. Am J Epidemiol. 2015;181(3):191‐197. doi: 10.1093/aje/kwu284 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 164. Brooks JM, Fang G. Interpreting treatment‐effect estimates with heterogeneity and choice: simulation model results. Clin Ther. 2009;31(4):902‐919. [DOI] [PubMed] [Google Scholar]
  • 165. Leeuwen vN, Lingsma H, Craen T, et al. Regression discontinuity design simulation and application in two cardiovascular trials with continuous outcomes. Epidemiology. 2016;27(4):503‐511. doi: 10.1097/EDE.0000000000000486 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.


Articles from Statistics in Medicine are provided here courtesy of Wiley

RESOURCES