Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2021 Apr 1.
Published in final edited form as: Clin Pharmacol Ther. 2020 Jan 24;107(4):834–842. doi: 10.1002/cpt.1754

Analytic and data sharing options in real-world multi-database studies of comparative effectiveness and safety of medical products

Sengwee Toh 1
PMCID: PMC7093255  NIHMSID: NIHMS1065591  PMID: 31869442

Abstract

A wide range of analytic and data sharing options are available in non-experimental multi-database studies designed to assess the real-world benefits and risks of medical products. Researchers often consider six scientific domains when choosing among these options – study design, exposure type, outcome type, covariate summarization technique, covariate adjustment method, and data sharing approach. This article reviews available analytic and data sharing options and discuss key scientific and practical considerations when choosing among these options in multi-database studies of comparative effectiveness and safety of medical products. The scientific considerations must be balanced against what the data-contributing sites are able or willing to share. While pooling of person-level datasets remains the most familiar and analytically flexible approach, newer analytic and data sharing approaches that share less granular summary-level information may be equally valid and preferred in some multi-database studies, especially when sharing of person-level data is challenging or infeasible.

Keywords: disease risk scores, distributed research networks, multi-center studies, multi-database studies, pharmacoepidemiology, privacy protection, propensity scores, real-world data, real-world evidence

Introduction

The use of real-world data to generate real-world evidence on the benefits and risks of medical products is not a completely new concept, but it attracted much greater interest and scrutiny following the passage of the 21st Century Cures Act, which required the U.S. Food and Drug Administration to incorporate real-world evidence in its regulatory decision-making process (13). Many of the real-world data studies may require large sample sizes and diverse study populations from multiple sources to generate robust and generalizable evidence. In these multi-database or multi-center studies (collectively called multi-database studies hereafter), it is critical to protect sensitive person-level and institution-level data in the pursuit of advancing scientific knowledge. In practice, researchers analyzing data from multiple databases often have to balance the scientific needs with what can realistically be shared by the data-contributing sites. In this article, we review available analytic and data sharing options and discuss key scientific and practical considerations when choosing among these options in real-world, non-experimental multi-database studies of comparative safety and effectiveness of medical products.

General overview of real-world multi-database studies

In a typical non-experimental multi-database study designed to assess the real-world benefits and risks of medical products, the analysis center, which can also be a data-contributing site, receives data from all participating sites and performs the desired statistical analysis using the pooled data (Figure 1). In principle, the research question informs the study design and statistical analysis, which in turn determine the type of dataset to be shared and combined. In practice, researchers often have to balance the analytic needs with what the data-contributing sites are able or willing to share due to a multitude of reasons, including concerns about patient privacy, unauthorized uses of transferred data, and disclosures of sensitive or proprietary institution-level information (47). Contractual agreements between some organizations (e.g., delivery systems, health plans) and their patients or members may restrict sharing of person-level data with other entities for secondary purposes, such as research.

Figure 1.

Figure 1.

The organizational structure of a typical multi-database study of comparative effectiveness and safety of medical products. Panel A shows a multi-database study in which the data-contributing sites each shares a person-level dataset with the analysis center. In a typical person-level dataset, each row represents an observation of a patient and each column represents a variable. Panel B shows a multi-database study in which the data-contributing sites each shares a summary-level dataset with the analysis center. The information included in the summary-level dataset varies by the analytic and data sharing option used. See Tables S1S7 for some examples of person-level and summary-level datasets.

Key analytic and data sharing considerations in real-world multi-database studies

As shown in Figure 2, there are six main scientific domains of analytic and data sharing considerations in multi-database studies designed to assess the real-world comparative effectiveness and safety of medical products. We briefly discuss them below.

Figure 2.

Figure 2.

Six scientific domains of analytic and data sharing considerations in real-world multi-database studies of comparative effectiveness and safety of medical products.

Study design

There are two families of study designs in non-experimental studies, one enables between-person comparisons and the other within-person (8, 9). The between-person designs compare distinct groups of individuals defined by their exposure status. The most well-known design in this family is the cohort design, including the commonly used new-user cohort design (10, 11). All other between-person comparison designs, including the case-cohort design and the case-control design, can be viewed as variants of the cohort design with more efficient sampling of the study cohort. The within-person designs compare the same group of individuals at different time periods during which the exposure status might vary (12). They include the case-crossover design (13), the case-time-control design (14), the self-controlled case series design (15), the self-controlled risk interval design (16), and their variants (e.g., the case-case-time-control design (17)).

Exposure type

In most studies of medical products, the exposure of interest is a binary (e.g., saxagliptin versus sitagliptin) or non-ordinal categorical (e.g., adalimumab versus etanercept versus infliximab) variable. However, certain exposures of interest can be classified as an ordinal variable (e.g., morphine-equivalent dosage classified into none [0.00 mg], low [0.01–49.99 mg], moderate [50.00–99.99 mg], and high [≥100 mg]), a count variable (e.g., 0, 1, 2, 3, and ≥4 dispensings of oral corticosteroids), or a continuous variable (e.g., morphine-equivalent dosage with potentially infinite and non-integer values). In practice, researchers may re-classify a count exposure to a categorical or continuous exposure in the analysis. Researchers may be interested in studying time-invariant or time-varying exposures.

Outcome type

Most studies of comparative safety and effectiveness of medical products focus on one of the following outcome types: binary outcome (e.g., occurrence of acute myocardial infarction), non-ordinal categorical outcome (e.g., occurrence of emergency department visits versus occurrence hospitalizations versus occurrence of intensive care unit admissions), ordinal outcome (e.g., absolute change in hemoglobin A1c level classified into none [0.00%], low [0.01–0.99%], moderate [1.00–2.99%], and high [≥3.00%]), count outcome (e.g., number of all-cause hospitalizations), continuous outcome (e.g., change in body mass index from baseline), and time-to-event outcome (e.g., time to first occurrence of acute myocardial infarction). In principle, these outcome types can be measured once or multiple times within an individual during the study period.

Covariate summarization technique

Exposure status, outcome status, confounders, and effect modifiers are often represented as individual variables in an analysis. Researchers can generally remove or mask direct patient identifiers and other potentially identifiable data elements, such as the protected health information per the U.S. Health Insurance Portability and Accountability Act, without compromising the validity of the analysis (18). It is possible to further de-identify the data elements using methods like homomorphic encryption, a form of encryption that allows researchers to perform analysis on the encrypted data and generate results identical to those obtained from the unencrypted version of the same data (19, 20). For the purposes of confounding adjustment, it is also possible to summarize multiple covariates using confounder summary scores, such as the propensity scores and the disease risk scores. Briefly, the propensity scores are the probabilities of having the study exposure given patients’ measured baseline covariates (21, 22). The disease risk scores are the probabilities or hazards of having the study outcome based on patients’ measured baseline characteristics (23, 24). In general, the propensity scores are preferred when the exposure is common and the outcome is rare; the disease risk scores are more advantageous when the exposure is rare (e.g., newly approved medical products) and there are more data to estimate the scores (e.g., from appropriate historical data) (24). As a data dimension reduction technique, these summary scores condense the information from a large number of covariates into a single, less identifiable variable.

Covariate adjustment method

Researchers can adjust for measured confounders, either as individual covariates or as confounder summary scores, through matching, stratification, restriction, weighting, or modelling. Although non-experimental multi-database studies are susceptible to several biases (e.g., confounding, selection bias, information bias), researchers are often interested in drawing causal inference from these studies. It is therefore important to note that these covariate adjustment methods produce treatment effects with different causal interpretations (25, 26).

Data sharing approach

Depending on the analysis, each data-contributing site produces and shares with the analysis center a dataset that contains person-level data, summary-table data, risk-set data, effect-estimate data, or intermediate statistics. Tables S1S7 provide some example datasets for each of these data sharing approaches.

Analytic and data sharing options in real-world multi-database cohort studies

With four exposure types, five outcome types, three covariate summarization techniques, five covariate adjustment methods, and five data sharing approaches, there would be a total of 1,500 (4 × 5 × 3 × 5 × 5) possible analytic and data sharing combinations in a multi-database study that employs a specific study design. To narrow down the scope of the discussion, we focus on real-world multi-database studies with the following characteristics in this section: (1) the studies employ a cohort design, (2) the studies examine the association between a time-invariant binary or categorical exposure and a one-time, non-correlated, and non-clustered outcome, and (3) each data-contributing site provides data on a distinct group of individuals. We focus on studies with these characteristics because they are among the most common types of multi-database studies of comparative effectiveness and safety of medical products.

Not all analytic and data sharing combinations are valid or currently feasible for the multi-database cohort studies of interest. For example, adjusting for disease risk scores through weighting is not an established confounding adjustment method, either in single-database or multi-database studies. In addition, these options offer different levels of confounding adjustment. For example, matching or stratification on individual covariates often suffers from the “curse of dimensionality”, i.e., researchers generally are only able to match or stratify on a handful of covariates, which may not provide adequate confounding adjustment compared to methods that use confounder summary scores. We describe some of the most commonly used or promising analytic and data sharing options below. We also provide examples of multi-database studies that have employed these options to examine the real-world comparative effectiveness and safety of medical products.

Person-level data with individual covariates

Among all the available analytic and data sharing options in multi-database studies, requesting person-level data with individual covariates from the data-contributing sites is the most analytically flexible approach (Figure 3). In person-level data sharing, each data-contributing site sends the analysis center an analytic dataset that includes one row per person and one column per variable (e.g., exposure status, outcome status, individual confounders, effect modifiers). The pooled person-level dataset allows researchers at the analysis center to perform both pre-specified and post hoc analyses, often without the need to go back to the data-contributing sites to request additional data elements.

Figure 3.

Figure 3.

Trade-off between analytic flexibility and privacy protection, by analytic and data sharing option. Analytic flexibility is broadly defined by the ability to perform pre-specified and post hoc analyses and the complexity of these analyses.

Using data from four health plans in the United States, Cooper et al investigated the risk of serious cardiovascular events in children and young adults aged 2–24 years treated with attention deficit-hyperactivity disorder medications (27). Each data-contributing site created a person-level dataset with individual covariates using an analytic program developed by the lead site and then transferred the dataset to the lead site for final statistical analysis. The study identified 81 outcome events among 1,200,438 eligible children and young adults. It did not observe an association between current use of an attention deficit-hyperactivity disorder medication and an increased risk of serious cardiovascular events.

Person-level data with confounder summary scores

Should the use of confounder summary scores be deemed appropriate for the study, researchers can reduce the dimension of the person-level data by summarizing the measured confounders using the propensity scores or the disease risk scores. In its simplest form, the dataset will only include the exposure variable, the outcome variable, and the confounder summary score. When estimated correctly, methods that use these summary scores produce results comparable to those from individual covariate adjustment (2831). Researchers can request other individual covariates needed for any pre-specified analysis, e.g., sex if they wish to perform sex-stratified analysis. With this option, researchers at the analysis center can perform essentially all the pre-specified analyses, but it may not be able to accommodate all post hoc analyses, without going back to the data-contributing sites to request more data elements. For example, if sex is included in the estimation of the propensity scores but not requested as an individual covariate, researchers will not be able to perform a post hoc sex-stratified analysis with the pooled data at hand.

Grijalva et al used data from four electronic healthcare databases in the United States to examine the association between initiation of anti-TNFα agents and the risk of hospitalizations for serious infections among patients with autoimmune diseases (32). To minimize sharing of a large number of covariates, they first estimated propensity scores within each data source and then requested a person-level dataset with variables indicating the exposure, outcome, select key covariates (e.g., age in decades, sex), and the estimated propensity scores for centralized pooling (33). The analytic and data sharing approach allowed the researchers to identify 10,484 propensity score-matched treatment episodes of anti-TNFα agents and non-biologic comparator medications among patients with rheumatoid arthritis, 2,323 matched episodes among patients with inflammatory bowel disease, and 3,215 matched episodes among patients with psoriasis and spondyloarthropathies. Based on 1,172 outcome events within the matched population, the authors found that initiation of anti-TNFα agents was not associated with an elevated risk of hospitalizations for serious infections among patients with autoimmune diseases.

Summary-level analytic and data sharing options that use confounder summary scores

For certain multi-database analyses, researchers can avoid sharing of person-level data and only use summary-level information in the form of a summary-table or risk-set data structure by leveraging the properties of confounder summary scores. For example, in a 1:1 propensity score-matched analysis of a binary exposure and a binary outcome, researchers can first estimate the propensity scores and perform propensity score matching within each data-contributing site. Each site then sends a summary-table dataset that includes the number of patients and the number of outcome events within each exposure group in the matched cohort to the analysis center. Using the summary-table datasets from all sites, researchers at the analysis center will be able to estimate the adjusted effect estimate (e.g., odds ratio) and its 95% confidence interval. It is also possible to perform diagnostics on the propensity score model without sharing person-level data. For example, database-specific standardized differences that compare the balance of baseline patient characteristics (34) and the propensity score distributions before and after matching can be obtained using only summary-level information (35).

If researchers choose to perform propensity score-stratified analysis, each stratum will contain the same summary-level information as described in the matched analysis. If the study is investigating the association between a binary exposure and a time-to-event outcome using a propensity score-matched or propensity score-stratified analysis, researchers can request a summary-level risk-set dataset to perform the desired analysis (36, 37). Results from this risk set-based approach have been shown to be identical to the results from the corresponding pooled person-level Cox regression model (3840).

Toh et al examined the risk of hospitalized heart failure in association with the use of saxagliptin and sitagliptin using data from 18 electronic healthcare databases in the United States (41). They identified 78,553 saxagliptin users and 298,124 sitagliptin users. Using propensity score matching and a risk set-based data sharing approach, they did not observe an elevated risk of hospitalized heart failure in either saxagliptin users or sitagliptin users when compared with users of pioglitazone, sulfonylureas, or long-acting insulin.

Distributed regression using summary-level intermediate statistics

Distributed regression is a statistical technique that performs the same numerical computation as the standard person-level regression analysis (42, 43). Unlike the standard regression analysis, distributed regression only requires summary-level intermediate statistics produced as part of the modelling process (e.g., sums of squares and cross products matrix) from the data-contributing sites to perform the same analysis. Using the summary statistics, researchers at the analysis center either calculate the effect estimates or if an iterative process is needed, update the parameter estimates and send them back to each site to further update their summary statistics. This iterative process continues until it meets a pre-specified convergence criterion or reaches the maximum number of iterations. In general, researchers can use logistic regression for binary or categorical outcomes, Poisson regression for count outcomes, linear regression for continuous outcomes, and Cox proportional hazards regression for time-to-event outcomes.

Real-world applications of the distributed regression approach to investigate the safety and effectiveness of medical products are limited. Although not conducted using real-world data, El Emam et al described how distributed logistic regression can be employed to detect rare adverse drug events (44). Wu et al used two clinical datasets to develop and test the distributed logistic regression approach (45). Pastorino et al used distributed regression to examine the association between maternal physical activity and the birth size of the infant using data from eight databases in the Europe and the United States (46). Using the association between antibiotic treatment early in life and the body mass index z-score later in life among children in 34 electronic healthcare databases in the United States, Toh et al conducted a proof-of-concept study on how distributed linear regression can be used to examine the benefits and risks of medical products (47).

Meta-analysis of database-specific effect estimates

Perhaps the most intuitive and easily understood summary-level analytic and data sharing option is the meta-analysis of database-specific effect estimates. With this option, each data-contributing site performs its own statistical analysis to produce the effect estimate and its 95% confidence interval (or other measures that allow researchers to calculate the database-specific weights) (48, 49). Researchers at the analysis center then pool the database-specific effect estimates to produce an overall estimate using a fixed-effect or random-effects model. This approach has been shown to produce results similar to those obtained from the corresponding pooled person-level data analysis (5052). There is no consensus about whether a fixed-effect or random-effects model is more appropriate in multi-database studies. Although random-effects meta-analysis accommodates treatment effect heterogeneity by database (more below), it has been argued that the data-contributing sites are generally known in advance and not a random sample from some potential larger universe of sites, so the assumption of a random-effects model may not hold in most multi-database studies (53).

Requena et al investigated the association between benzodiazepine use and the risk of hip or femur fracture using three databases from Spain, the United Kingdom, and the Netherlands (54). Using a common protocol approach, they identified 7,079 outcome events among 1,272,957 eligible patients across the three databases. In a random-effects meta-analysis of the three site-specific effect estimates, they observed an increased risk of hip or femur fracture with current use of benzodiazepines compared with past use.

Treatment effect heterogeneity by database in real-world multi-database studies

When there is treatment effect heterogeneity by database, the issues around the appropriateness of combining data across sites apply to all the analytic and data sharing options discussed in this article. To the extent that researchers have a way to distinguish the data-contributing sites, which is often the case, all the analytic and data sharing options described above allow them to examine potential treatment effect heterogeneity by database. In the presence of substantial heterogeneity by database, it is generally not advisable to pool the data from various sites and only produce an overall effect estimate (55). With the exception of random-effects meta-analysis, methods that share summary-level information described in this article primarily consider data source as a fixed effect, a setting in which pooling of multiple databases is less controversial. Although random-effects meta-analysis explicitly allows treatment effect heterogeneity by database, it is important to note that the approach is not designed to remove the intrinsic heterogeneity but rather to properly account for the uncertainty in the effect estimate of interest, which is generally the mean or center of the random effects distribution. Clinical interpretation of the pooled effect estimate from a random-effects meta-analysis can sometimes be challenging (56, 57), and the number of sites required to support the analysis may need to be sufficiently large to provide valid results (58).

It is worth noting that unlike the conventional pooling of independently conducted randomized controlled trials or observational studies, well-designed multi-database studies often have more opportunities to remove or minimize several “preventable” sources of heterogeneity, such as differences in study design, analytic method, cohort eligibility criteria, exposure definition, outcome definition, confounder definitions, and follow-up length. Specifically, researchers can employ the same study design and statistical analysis and define the study variables as consistently as possible across all databases. If there is treatment effect heterogeneity by database after removal of those sources, researchers can then focus on more substantive contributors of heterogeneity, e.g., differences in distributions of effect modifiers, and differences in formulary, clinical practice, or health system.

Other considerations

Ease of implementation

In practice, ease of implementation is a major consideration when choosing among available analytic and data sharing options. For example, even though distributed regression has appealing theoretical properties because it produces results identical to those obtained from the corresponding pooled person-level data analysis, the method is not routinely used in multi-database studies because some regression models involve multiple rounds of information exchange between the analysis center and the data-contributing sites. When done manually, the process can be labor-intensive and error-prone. To make distributed regression a more practical analytic option in practice, some researchers have developed statistical packages and stand-alone software to partially or fully automate the file transfer process (44, 45, 5961).

The amount of data processing and statistical analysis done at the data-contributing site versus the analysis center also varies by analytic and data sharing option (Figure 4). In principle, standardizing the format and definition of the data elements across databases improves the ease of implementation of the analytic and data sharing options discussed in this article. There has been an increase in the number of distributed data networks or multi-center collaboratives that employ a common data model, such as the Sentinel System (62), the Canadian Network for Observational Drug Effect Studies (63), and networks that leverage the infrastructure from the Observational Health Data Sciences and Informatics (64). In these networks, data-contributing sites transform their source data into standardized formats in advance to make the conduct of subsequent multi-database studies more efficient (65). Specifically, the analysis center can first develop and test an analytic program against the common data model, and then distribute the program to the data-contributing sites to create the necessary dataset for the chosen analytic and data sharing option. In principle, the data-contributing sites can simply execute the analytic program without substantial ad hoc programming, which helps reduce programming burden and improve consistency.

Figure 4.

Figure 4.

Figure 4.

The amount of data processing and statistical analysis done at the data-contributing site versus the analysis center, by analytic and data sharing option

Database-specific versus study-wide confounder summary scores

For researchers considering using confounder summary scores, it is generally recommended that these scores be estimated within each database if sample size allows. The prevalence of the study exposure may vary across databases due to differences in clinical practice and other factors. Patients from two different databases may not be comparable even if they have the same estimated propensity score. Estimating database-specific confounder summary scores also allows a given covariate (e.g., age) to have different effects on the probability of being exposed or having the outcome at different sites. It is also generally recommended to handle the database-specific confounder summary scores and data source simultaneously in the analysis, e.g., matching on propensity scores within site (66) and stratifying jointly on site-specific propensity scores and site.

However, some or even all of the data-contributing sites may be too small to allow robust estimation of database-specific confounder summary scores. An alternative is to fit a study-wide confounder summary score estimation model that includes the confounders, site indicators, and additional interaction terms between confounders and site indicators. This allows all sites to contribute to the estimation of the confounder summary scores. It is straightforward to fit such a model if the data-contributing sites share person-level data with individual covariates with the analysis center. It is also possible to use distributed regression to fit a study-wide model using only summary-level information.

Database-specific confounder lists versus common confounder list across databases

In many multi-database studies, it is common for some sites to have richer covariate information than other sites. For example, some databases may have laboratory or genetic testing results. For simplicity and consistency, researchers can choose to adjust for a common set of confounders that are available at all data-contributing sites. Alternatively, researchers can choose to include all measured confounders at a given site even though they are not available at other sites. More data-adaptive covariate selection methods, such as the high-dimensional propensity score approach (67), can also be used to identify additional covariates that are available at certain sites but not pre-specified by the researchers. Alternatively, researchers can conceptualize this as a missing data problem, with some covariates partially or completely missing at some data-contributing sites, and use methods like multiple imputation by chained equations that appropriately account for between-site heterogeneity to handle the missing data (68, 69). Whenever possible, researchers should specify their primary analysis and perform sensitivity analyses to examine the robustness of their results.

Transparency and reproducibility

In multi-database studies, researchers often have to specify the study design and analysis plan in advance in order to request the necessary data elements from the data-contributing sites. This is especially true for analytic and data sharing options that require only summary-level information because most of the data processing steps occur at the data-contributing sites. Failure to request the correct person-level data or summary-level information may result in another data request, creating delays and additional burdens. By being explicit about the analysis, multi-database studies can minimize the risk of data dredging and selective reporting (70). Regardless of the analytic and data sharing option used, researchers conducting multi-database studies should make the data models, study protocols, analytic code, and results publicly available whenever possible to improve transparency and encourage reproducibility (71). Several large distributed data networks currently adopt this practice (6264).

Institutional privacy

In principle, researchers should present the overall and database-specific effect estimates (e.g., using forest plots) in multi-database studies (55, 72). However, certain data-contributing sites may be concerned about this practice even if the site identities are masked. For example, one of the largest (or smallest) organizations may have a higher mortality risk compared to other sites simply because it serves an older or sicker population. The organization may be concerned about presenting database-specific results without proper context because the width of the 95% confidence interval around the effect estimate may make it identifiable. In practice, researchers have to balance the trade-off between institutional privacy and transparency of the multi-database study.

Methodological gaps

Methods that analyze summary-level information are relatively well-developed for time-invariant binary or categorical exposures and one-time, non-correlated, and non-clustered outcomes. More complicated statistical analysis that involves time-varying exposures or repeated outcomes may still require sharing of person-level data, but there has been some methodological work initiated to address this issue (73). There is limited research in developing summary-level analytic and data sharing options for multi-database studies with continuous exposures, although this type of exposure is less commonly used in comparative effectiveness and safety research of medical products. Except for random-effects meta-analysis, most methods that share summary-level information consider database as a fixed effect, which restricts their use when there is treatment effect heterogeneity by database. All of these methodological gaps require additional research. Missing data within and between sites is common in multi-database studies. Methods developed in individual participant data meta-analysis to handle multilevel or clustered missing data can be applied to multi-database studies of comparative effectiveness and safety of medical products (68, 69), but additional enhancements that allow missing data analysis to be performed without pooling of person-level data is necessary.

Conclusions

Researchers can choose from a suite of analytic and data sharing methods in multi-database studies of comparative effectiveness and safety of medical products based on the study design, exposure type, outcome type, covariate summarization technique, covariate adjustment method, and data sharing approach. These scientific considerations must be balanced against what the data-contributing sites are able or willing to share. Pooling of person-level datasets remains the most familiar and analytically flexible approach. Newer analytic and data sharing approaches that share less granular summary-level information may be equally valid compared to their corresponding pooled person-level data analysis. These summary-level analytic and data sharing approaches may be preferred in some multi-database studies and make otherwise infeasible studies possible due to their more privacy-protecting properties.

Supplementary Material

Supp TableS1-7

Acknowledgments

The author thanks Rui Wang and Xiaojuan Li in the Department of Population Medicine, Harvard Medical School and Harvard Pilgrim Health Care Institute for their comments on an earlier version of the manuscript. The author also thanks Xiaojuan Li and Di Shu in the Department of Population Medicine, Harvard Medical School and Harvard Pilgrim Health Care Institute for their help with the appendices.

Funding: The author is supported by the National Institutes of Health (U01EB023683), the Agency for Healthcare Research and Quality (R01HS026214), and a Harvard Pilgrim Health Care Institute Robert H. Ebert Career Development Award.

Footnotes

Conflict of interest: The authors declared no competing interests for this work.

References

  • (1).Sherman RE et al. Real-World Evidence - What Is It and What Can It Tell Us? N Engl J Med 375, 2293–7 (2016). [DOI] [PubMed] [Google Scholar]
  • (2).Jarow JP, LaVange L & Woodcock J Multidimensional Evidence Generation and FDA Regulatory Decision Making: Defining and Using “Real-World” Data. JAMA 318, 703–4 (2017). [DOI] [PubMed] [Google Scholar]
  • (3).Corrigan-Curay J, Sacks L & Woodcock J Real-World Evidence and Real-World Data for Evaluating Drug Safety and Effectiveness. JAMA 320, 867–8 (2018). [DOI] [PubMed] [Google Scholar]
  • (4).Simon GE et al. Data Sharing and Embedded Research. Ann Intern Med 167, 668–70 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (5).Brown JS, Holmes JH, Shah K, Hall K, Lazarus R & Platt R Distributed health data networks: a practical and preferred approach to multi-institutional evaluations of comparative effectiveness, safety, and quality of care. Med Care 48, S45–51 (2010). [DOI] [PubMed] [Google Scholar]
  • (6).Toh S, Platt R, Steiner JF & Brown JS Comparative-effectiveness research in distributed health data networks. Clin Pharmacol Ther 90, 883–7 (2011). [DOI] [PubMed] [Google Scholar]
  • (7).Mazor KM et al. Stakeholders’ views on data sharing in multicenter studies. J Comp Eff Res, (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (8).Gagne JJ et al. Design considerations in an active medical product safety monitoring system. Pharmacoepidemiol Drug Saf 21 Suppl 1, 32–40 (2012). [DOI] [PubMed] [Google Scholar]
  • (9).Rothman KJ, Greenland S & Lash TL (eds.). Modern Epidemiology (Lippincott Williams & Wilkins, Philadelphia, PA, 2008). [Google Scholar]
  • (10).Ray WA Evaluating medication effects outside of clinical trials: new-user designs. Am J Epidemiol 158, 915–20 (2003). [DOI] [PubMed] [Google Scholar]
  • (11).Lund JL, Richardson DB & Sturmer T The active comparator, new user study design in pharmacoepidemiology: historical foundations and contemporary application. Curr Epidemiol Rep 2, 221–8 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (12).Hallas J & Pottegard A Use of self-controlled designs in pharmacoepidemiology. J Intern Med 275, 581–9 (2014). [DOI] [PubMed] [Google Scholar]
  • (13).Maclure M The case-crossover design: a method for studying transient effects on the risk of acute events. Am J Epidemiol 133, 144–53 (1991). [DOI] [PubMed] [Google Scholar]
  • (14).Suissa S The case-time-control design. Epidemiology 6, 248–53 (1995). [DOI] [PubMed] [Google Scholar]
  • (15).Whitaker HJ, Farrington CP, Spiessens B & Musonda P Tutorial in biostatistics: the self-controlled case series method. Stat Med 25, 1768–97 (2006). [DOI] [PubMed] [Google Scholar]
  • (16).Glanz JM et al. Four different study designs to evaluate vaccine safety were equally validated with contrasting limitations. J Clin Epidemiol 59, 808–18 (2006). [DOI] [PubMed] [Google Scholar]
  • (17).Wang S et al. Future cases as present controls to adjust for exposure trend bias in case-only studies. Epidemiology 22, 568–74 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (18).Sarpatwari A, Kesselheim AS, Malin BA, Gagne JJ & Schneeweiss S Ensuring patient privacy in data sharing for postapproval research. N Engl J Med 371, 1644–9 (2014). [DOI] [PubMed] [Google Scholar]
  • (19).Lu WJ, Yamada Y & Sakuma J Privacy-preserving genome-wide association studies on cloud environment using fully homomorphic encryption. BMC Med Inform Decis Mak 15 Suppl 5, S1 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (20).McLaren PJ et al. Privacy-preserving genomic testing in the clinic: a model using HIV treatment. Genet Med 18, 814–22 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (21).Rosenbaum PR & Rubin DB The central role of the propensity score in observational studies for causal effects. Biometrika 70, 41–55 (1983). [Google Scholar]
  • (22).Rosenbaum PR & Rubin DB Reducing bias in observational studies using subclassification on the propensity score. J Am Stat Assoc 79, 516–24 (1984). [Google Scholar]
  • (23).Miettinen OS Stratification by a multivariate confounder score. Am J Epidemiol 104, 609–20 (1976). [DOI] [PubMed] [Google Scholar]
  • (24).Arbogast PG & Ray WA Use of disease risk scores in pharmacoepidemiologic studies. Stat Methods Med Res 18, 67–80 (2009). [DOI] [PubMed] [Google Scholar]
  • (25).Kurth T et al. Results of multivariable logistic regression, propensity matching, propensity adjustment, and propensity-based weighting under conditions of nonuniform effect. Am J Epidemiol 163, 262–70 (2006). [DOI] [PubMed] [Google Scholar]
  • (26).Austin PC An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies. Multivariate Behav Res 46, 399–424 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (27).Cooper WO et al. ADHD drugs and serious cardiovascular events in children and young adults. N Engl J Med 365, 1896–904 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (28).Cook EF & Goldman L Performance of tests of significance based on stratification by a multivariate confounder score or by a propensity score. J Clin Epidemiol 42, 317–24 (1989). [DOI] [PubMed] [Google Scholar]
  • (29).Arbogast PG & Ray WA Performance of disease risk scores, propensity scores, and traditional multivariable outcome regression in the presence of multiple confounders. Am J Epidemiol 174, 613–20 (2011). [DOI] [PubMed] [Google Scholar]
  • (30).Sturmer T, Joshi M, Glynn RJ, Avorn J, Rothman KJ & Schneeweiss S A review of the application of propensity score methods yielded increasing use, advantages in specific settings, but not substantially different estimates compared with conventional multivariable methods. J Clin Epidemiol 59, 437–47 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (31).Cadarette SM, Gagne JJ, Solomon DH, Katz JN & Sturmer T Confounder summary scores when comparing the effects of multiple drug exposures. Pharmacoepidemiol Drug Saf 19, 2–9 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (32).Grijalva CG et al. Initiation of tumor necrosis factor-alpha antagonists and the risk of hospitalization for infection in patients with autoimmune diseases. JAMA 306, 2331–9 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (33).Rassen JA, Solomon DH, Curtis JR, Herrinton L & Schneeweiss S Privacy-maintaining propensity score-based pooling of multiple databases applied to a study of biologics. Med Care 48, S83–9 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (34).Mamdani M et al. Reader’s guide to critical appraisal of cohort studies: 2. Assessing potential for confounding. BMJ 330, 960–2 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (35).Connolly JG et al. Development and application of two semi-automated tools for targeted medical product surveillance in a distributed data network. Curr Epidemiol Rep 4, 298–306 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (36).Fireman B, Lee J, Lewis N, Bembom O, van der Laan M & Baxter R Influenza vaccination and mortality: differentiating vaccine effects from bias. Am J Epidemiol 170, 650–6 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (37).Toh S, Gagne JJ, Rassen JA, Fireman BH, Kulldorff M & Brown JS Confounding adjustment in comparative effectiveness research conducted within distributed research networks. Med Care 51, S4–10 (2013). [DOI] [PubMed] [Google Scholar]
  • (38).Toh S, Shetterly S, Powers JD & Arterburn D Privacy-preserving analytic methods for multisite comparative effectiveness and patient-centered outcomes research. Med Care 52, 664–8 (2014). [DOI] [PubMed] [Google Scholar]
  • (39).Yoshida K, Gruber S, Fireman BH & Toh S Comparison of privacy-protecting analytic and data-sharing methods: A simulation study. Pharmacoepidemiol Drug Saf 27, 1034–41 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (40).Li X et al. Validity of Privacy-Protecting Analytical Methods That Use Only Aggregate-Level Information to Conduct Multivariable-Adjusted Analysis in Distributed Data Networks. Am J Epidemiol 188, 709–23 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (41).Toh S et al. Risk for hospitalized heart failure among new users of saxagliptin, sitagliptin, and other antihyperglycemic drugs: A retrospective cohort study. Ann Intern Med 164, 705–14 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (42).Karr AF, Lin X, Sanil AP & Reiter JP Secure regression on distributed databases. J Comput Graph Stat 14, 263–79 (2005). [Google Scholar]
  • (43).Fienberg SE, Fulp WJ, Slavković AB & Wrobel TA “Secure” log-linear and logistic regression analysis of distributed databases. Lect Notes Comput Sci 2006, 277–90 (2006). [Google Scholar]
  • (44).El Emam K, Samet S, Arbuckle L, Tamblyn R, Earle C & Kantarcioglu M A secure distributed logistic regression protocol for the detection of rare adverse drug events. J Am Med Inform Assoc 20, 453–61 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (45).Wu Y, Jiang X, Kim J & Ohno-Machado L Grid Binary LOgistic REgression (GLORE): building shared models without sharing data. J Am Med Inform Assoc 19, 758–64 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (46).Pastorino S et al. Associations between maternal physical activity in early and late pregnancy and offspring birth size: remote federated individual level meta-analysis from eight cohort studies. BJOG 126, 459–70 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (47).Toh S et al. Privacy-protecting multivariable-adjusted distributed regression analysis for multi-center pediatric study. Pediatr Res, (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (48).DerSimonian R & Laird N Meta-analysis in clinical trials. Control Clin Trials 7, 177–88 (1986). [DOI] [PubMed] [Google Scholar]
  • (49).Borenstein M, Hedges LV, Higgins JP & Rothstein HR A basic introduction to fixed-effect and random-effects models for meta-analysis. Res Synth Methods 1, 97–111 (2010). [DOI] [PubMed] [Google Scholar]
  • (50).Rassen JA, Avorn J & Schneeweiss S Multivariate-adjusted pharmacoepidemiologic analyses of confidential information pooled from multiple health care utilization databases. Pharmacoepidemiol Drug Saf 19, 848–57 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (51).Toh S et al. Multivariable confounding adjustment in distributed data networks without sharing of patient-level data. Pharmacoepidemiol Drug Saf 22, 1171–7 (2013). [DOI] [PubMed] [Google Scholar]
  • (52).Tudur Smith C et al. Individual participant data meta-analyses compared with meta-analyses based on aggregate data. Cochrane Database Syst Rev 9, MR000007 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (53).Platt RW, Dormuth CR, Chateau D & Filion K Observational Studies of Drug Safety in Multi-Database Studies: Methodological Challenges and Opportunities. EGEMS (Wash DC) 4, 1221 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (54).Requena G et al. Hip/femur fractures associated with the use of benzodiazepines (anxiolytics, hypnotics and related drugs): a methodological approach to assess consistencies across databases from the PROTECT-EU project. Pharmacoepidemiol Drug Saf 25 Suppl 1, 66–78 (2016). [DOI] [PubMed] [Google Scholar]
  • (55).Madigan D et al. Evaluating the impact of database heterogeneity on observational study results. Am J Epidemiol 178, 645–51 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (56).Greenland S Can meta-analysis be salvaged? Am J Epidemiol 140, 783–7 (1994). [DOI] [PubMed] [Google Scholar]
  • (57).Riley RD, Higgins JP & Deeks JJ Interpretation of random effects meta-analyses. BMJ 342, d549 (2011). [DOI] [PubMed] [Google Scholar]
  • (58).Guolo A & Varin C Random-effects meta-analysis: the number of studies matters. Stat Methods Med Res 26, 1500–18 (2017). [DOI] [PubMed] [Google Scholar]
  • (59).Lu CL et al. WebDISCO: a web service for distributed cox model learning without patient-level data sharing. J Am Med Inform Assoc 22, 1212–9 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (60).Gaye A et al. DataSHIELD: taking the analysis to the data, not the data to the analysis. Int J Epidemiol 43, 1929–44 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (61).Her QL et al. A query workflow design to perform automatable distributed regression analysis in large distributed data networks. EGEMS (Wash DC) 6, 11 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (62).Ball R, Robb M, Anderson SA & Dal Pan G The FDA’s sentinel initiative--A comprehensive approach to medical product surveillance. Clin Pharmacol Ther 99, 265–8 (2016). [DOI] [PubMed] [Google Scholar]
  • (63).Suissa S et al. CNODES: the Canadian Network for Observational Drug Effect Studies. Open Med 6, e134–40 (2012). [PMC free article] [PubMed] [Google Scholar]
  • (64).Hripcsak G et al. Observational Health Data Sciences and Informatics (OHDSI): Opportunities for Observational Researchers. Stud Health Technol Inform 216, 574–8 (2015). [PMC free article] [PubMed] [Google Scholar]
  • (65).Toh S, Pratt N, Klungel O, Gagne JJ & Platt RW Chapter 25: Distributed networks of databases analyzed using common protocols and/or common data models In: Pharmacoepidemiology (eds. Strom BL, Kimmel SE and Hennessy S) 617–38 (Wiley-Blackwell, Chichester, UK, 2019). [Google Scholar]
  • (66).Gayat E, Thabut G, Christie JD, Mebazaa A, Mary JY & Porcher R Within-center matching performed better when using propensity score matching to analyze multicenter survival data: empirical and Monte Carlo studies. J Clin Epidemiol 66, 1029–37 (2013). [DOI] [PubMed] [Google Scholar]
  • (67).Schneeweiss S, Rassen JA, Glynn RJ, Avorn J, Mogun H & Brookhart MA High-dimensional propensity score adjustment in studies of treatment effects using health care claims data. Epidemiology 20, 512–22 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (68).Resche-Rigon M, White IR, Bartlett JW, Peters SA, Thompson SG & Group P-IS Multiple imputation for handling systematically missing confounders in meta-analysis of individual participant data. Stat Med 32, 4890–905 (2013). [DOI] [PubMed] [Google Scholar]
  • (69).Resche-Rigon M & White IR Multiple imputation by chained equations for systematically and sporadically missing multilevel data. Stat Methods Med Res 27, 1634–49 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (70).Platt RW, Platt R, Brown JS, Henry DA, Klungel OH & Suissa S How pharmacoepidemiology networks can manage distributed analyses to improve replicability and transparency and minimize bias. Pharmacoepidemiol Drug Saf, (2019). [DOI] [PubMed] [Google Scholar]
  • (71).Wang SV et al. Reporting to Improve Reproducibility and Facilitate Validity Assessment for Healthcare Database Studies V1.0. Pharmacoepidemiol Drug Saf 26, 1018–32 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (72).Bate A, Chuang-Stein C, Roddam A & Jones B Lessons from meta-analyses of randomized clinical trials for analysis of distributed networks of observational databases. Pharm Stat 18, 65–77 (2019). [DOI] [PubMed] [Google Scholar]
  • (73).Jones EM, Sheehan NA, Gaye A, LaFlamme P & Burton P Combined analysis of correlated data when data cannot be pooled. Stat 2, 72–85 (2013). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp TableS1-7

RESOURCES