Abstract
Background
Distributed Research Networks (DRNs) offer significant opportunities for collaborative multi-site research and have significantly advanced healthcare research based on clinical observational data. However, generating high-quality real-world evidence using fit-for-use data from multi-site studies faces important challenges, including biases associated with various types of heterogeneity within and across sites and data sharing difficulties. Over the last 10 years, Privacy-Preserving Distributed Algorithms (PDA) have been developed and utilized in numerous national and international real-world studies spanning diverse domains, from comparative effectiveness research, target trial emulation, to healthcare delivery, policy evaluation, and system performance assessment. Despite these advances, there remains a lack of comprehensive and clear guiding principles for generating high-quality real-world evidence through collaborative studies leveraging the methods under PDA.
Objective
The paper aims to establish 10 principles of best practice for conducting high-quality multi-site studies using PDA. These principles cover all phases of research, including study preparation, protocol development, analysis, and final reporting.
Discussion
The 10 principles for conducting a PDA study outline a principled, efficient, and transparent framework for employing distributed learning algorithms within DRNs to generate reliable and reproducible real-world evidence.
Keywords: clinical observational data, distributed research network, principled Privacy-Preserving Distributed Algorithms (PDA), reliable clinical evidence generation
Introduction
Over the past decades, there has been a substantial increase in the number of distributed research networks (DRNs) across healthcare systems. DRNs refer to collaborative infrastructures in which multiple institutions pursue shared scientific objectives while retaining patient-level data locally under their own governance. As a result, analyses require coordinated cross-site procedures rather than direct pooling of individual-level data. These networks across healthcare systems include the Observational Health Data Sciences and Informatics (OHDSI),1 the National Patient-Centered Clinical Research Network (PCORnet),2 the NIH-funded Health Care Systems Research Collaboratory,3,4 the Biologics Effectiveness and Safety (BEST) Initiative,5 and the Sentinel Initiative of the Food and Drug Administration (FDA).6,7 These research networks have greatly expanded the prospects for multi-site research and public health surveillance activities across multiple data partners. In addition, the COVID-19 pandemic has further facilitated the development of various national and international multi-institutional research consortia, such as the 4CE consortium,8 the RECOVER initiative,9 N3C,10 among others. Although these networks differ in scope, governance structures, and analytic pipelines, they share core structural characteristics, including decentralized data storage, site-level control of access and disclosure, reliance on harmonized data models or shared phenotyping standards.
The concept of large-scale collaborative research is not new. Historic examples such as the Veterans Affairs (VA) Cooperative Studies Program11,12 demonstrate that multi-center clinical collaborations have long generated high-impact evidence through coordinated study design and standardized protocols. Contemporary DRNs extend this tradition into the era of electronic health records (EHR) and decentralized data governance, introducing new methodological challenges related to privacy, heterogeneity, and distributed computation.
However, leveraging multi-site clinical observational data, such as EHR, administrative claims, and disease registries for clinical evidence generation presents significant challenges.13,16 Key barriers include restrictions on patient-level data sharing, the need for potentially iterative and resource-intensive communications across sites within the networks, the presence of various types of heterogeneity, and the reliance on labor-intensive human-in-the-loop processes. Ensuring analytic transparency and reproducibility further compounds these challenges, making multi-site research both methodologically intricate and operationally demanding.
In addition to operational and methodological complexities, regulatory and governance constraints further limit direct data integration across institutions. Requirements related to patient confidentiality, institutional review board oversight, and data use agreements often prohibit the transfer of individual-level data beyond local sites. As a result, collaborative analyses must be designed to operate within decentralized data environments while preserving institutional autonomy and compliance with privacy regulations.
In response to these challenges, substantial efforts have been devoted to developing federated learning algorithms for healthcare. Early examples include Grid Binary Logistic Regression (GLORE),14 a distributed algorithm for conducting logistic regression, and WebDISCO (a web service for distributed Cox model learning), which enables fitting the Cox proportional hazards models across sites.15 In particular, Privacy-Preserving Distributed Algorithms (PDA, https://pdamethods.org/)16,17 have been developed to further facilitate the use of clinical observational data from DRNs. PDA refer to statistical learning methods designed for DRN settings, in which model estimation and inference are achieved through the exchange of carefully constructed summary statistics rather than raw patient-level data, with performance comparable to pooled analyses under institutional privacy constraints. The PDA framework encompasses a broad family of such distributed and federated algorithms tailored to the methodological and operational demands of multi-site research. Throughout this paper, we use the terms “distributed” and “federated” interchangeably to refer to settings in which patient-level data remain stored locally and are not shared across sites. The “PDA framework” refer collectively to all algorithms developed under the PDA paradigm, including those currently implemented in the PDA package18 and future ones.
Although various distributed learning algorithms have been developed and implemented under PDA, there are no comprehensive or clear guiding principles for conducting collaborative studies using these algorithms. Herein, we define a multi-site study employing any PDA methods as a “PDA study.” This paper outlines 10 foundational principles we propose to guide the design, execution, and reporting of PDA studies, spanning all stages from study preparation to dissemination of results. We also provide an overview of a typical PDA study to illustrate these principles in practice. Looking forward, PDA studies have the potential to expand the scope of clinical evidence generation to additional medical domains and downstream tasks, ultimately strengthening the foundation for trustworthy and generalizable real-world evidence.
Materials and methods
Privacy-Preserving Distributed Algorithms (PDA)
The algorithms within the PDA framework are specifically designed to protect patients’ confidentiality by only requiring the sharing of summary statistics, not individual patient-level data, and are implemented under data usage agreements and institutional review board (IRB) protocols. These algorithms are theoretically derived to produce estimates identical to those obtained from pooled analyses under the same model specification, and their equivalence has been demonstrated through mathematical proofs, simulation studies, and empirical benchmarking against pooled analyses when feasible.14,19 Furthermore, the PDA framework supports a broad spectrum of analytical tasks, including comparative effectiveness research, causal inference, target trial emulation, health disparities assessment, and health policy evaluation. From a modeling standpoint, the PDA framework supports diverse data structures and outcome types, including binary, count, continuous, and time-to-event outcomes, as well as extensions such as competing risks, multivariate associations, and high-dimensional feature modeling. In practice, the framework provides methodological components that allow investigators to explicitly model or evaluate between-site differences in data provenance, distributions, data capture processes, and study design, rather than assuming homogeneity across sites.
What distinguishes PDA from previous efforts is its emphasis on streamlining communication rounds, that is the exchange of summary statistics across participating data partners. In multi-site studies operating using decentralized data networks, the most time-intensive steps involve repeated communication rounds that require coordination, synchronization, and secure data transfer across institutions. By substantially reducing the number of these exchanges, especially in modeling fitting, PDA improves the computational efficiency, lowers latency, and reduces the risk of miscommunication during summary statistics transfer. This design makes distributed algorithms more practical, scalable, and efficient for real-world healthcare and clinical research applications. Importantly, while emphasizing the minimization of communication rounds for sharing summary statistics, PDA ensures that data-sharing constraints do not compromise the accuracy and validity of the estimates.
Beyond communication efficiency, the PDA framework is specifically designed to be robust to, or explicitly account for, the potential existence of heterogeneity in treatment effects or other parameters of scientific interest. Ignoring heterogeneity within and across sites often leads to biased estimates and misleading quantification of evidence. Common types of heterogeneity encountered in DRNs include differences in data provenance, distributions, modalities, quality, and design. The current PDA framework features a comprehensive set of algorithms tailored to address or accommodate these variations.
Each algorithm within the PDA framework has been evaluated in a diverse set of studies with clinical observational data and has been widely adopted across diverse clinical fields including opioid use disorder (OUD),20 dementia and other aging-related conditions,21 myocardial infarction,22 pediatric conditions,23 fetal loss,19,24 COVID-19,25–27 long COVID,23,28 pediatric Crohn’s disease (PCD),29 health disparities and fairness,30 and health policy.30–32 To date, the PDA R package has been downloaded more than 13 300 times since its release in 2020. The framework has been utilized in more than 25 national and 5 international studies, has contributed to more than 35 methodological papers, and is currently being applied in 8 ongoing studies. Furthermore, the PDA framework has been adopted by institutions and organizations worldwide, including the International Agency for Research on Cancer (IARC) in France, Peking University Health Science Center, Ajou University Graduate School of Medicine, The Information System for Research in Primary Care (SIDIAP), Erasmus University Medical Center, and Hospital del Mar Research Institute.
Ten guiding principles
The PDA framework is strategically designed to ensure the scalability, applicability, and transparency of distributed learning algorithms for analyzing large-scale data from DRNs in the context of generating clinical evidence. The implementation of a PDA study adheres to 10 foundational principles that guide investigators through essential steps and key considerations in multi-site studies–from the study preparation and protocol development to analysis and reporting. The framework provides a flexible and comprehensive set of distributed learning algorithms, each tailored to meet the practical challenges of data integration, such as reducing communication overheads, protecting patient confidentiality, and understanding and accounting for between-site heterogeneity. Users can select the algorithms that best align with their specific statistical objectives and underlying assumptions of their PDA study. In the following Table 1, we present the 10 guiding principles for conducting a PDA study, covering all phases of research, including study preparation, protocol development, analysis, to final reporting.
Table 1.
Ten Guiding principles of implementing a PDA study.
|
|
|
|
|
|
|
|
|
|
Phase I: Preparation Phase; Phase II: Analysis Phase; Phase III: Reporting Phase.
Overview of a principled PDA study
Figure 1 illustrates the overall workflow of a typical PDA study, which is organized into 3 phases: Preparation, Analysis, and Reporting.
Figure 1.

Overview of a typical PDA study and its 10 guiding principles. The PDA workflow consists of 3 sequential phases: (I) Preparation, which includes defining the research protocol and assessing data quality and potential biases (Principles 1–2); (II) Analysis, where participating sites retain patient-level data locally and collaborate through secure exchange of summary statistics using distributed learning algorithms that leverage heterogeneity and ensure accuracy (Principles 3–8); and (III) Reporting, which emphasizes transparency, open research dissemination, and sustained collaboration (Principles 9–10).
In the Preparation phase, the research protocol is collaboratively defined by the participating sites, including specification the research question and identification of the aggregation and distribution units (Principle 1). To ensure the data are fit-for-use, potential risks of bias are assessed through quality assessment of cohort, such as cohort diagnostics33 (Principle 2).
During the Analysis phase, patient-level data remain securely stored at each site to protect patient confidentiality (Principle 3). The selected PDA algorithm leverages heterogeneity across sites (Principle 4) to strengthen the robustness of evidence, meanwhile, maintaining accuracy through systematic and reproducible statistical learning methods (Principle 5). Sites collaborate at scale by executing analyses locally and sharing only summary statistics (Principle 6). The PDA framework further supports flexible downstream analytic tasks (Principle 7) and support implementation readiness for real-world applications (Principle 8).
Finally, in the Reporting phase, PDA promotes open science is promoted by publicly disseminating pre-specified protocols, analysis code, and R packages through publicly accessible repositories such as GitHub and CRAN (Principle 9). This approach ensures transparency and encourages widespread collaboration to critically interpret findings and assess the implications within the research community (Principle 10).
Preparation phase: study protocol design and cohort diagnostics (Principles 1 and 2)
Before implementing statistical model fitting within the PDA framework, the protocol development stage is a crucial step in multi-site collaborative studies. During this stage, investigators develop the study analysis plan and protocol by specifying the study design, study cohort with inclusion/exclusion criteria, and selecting appropriate distributed learning algorithms. Equally important is the explicit definition of the aggregation and distribution units. The Aggregation Unit represents to the unit of study interest in the analytical models, whose index is included in the model to distinguish the subjects across different units. This can be an individual hospital, clinical site, or a collective network, each distinctly marked by a site/study identification number for analytical inclusion. The Distribution Unit defines the operational unit of communication, including the data partners who exchange the summary statistics for collaborative modeling tasks. The distribution unit may coincide with, or differ from, the aggregation unit depending on study configuration. The PDA framework provides flexible solutions for multi-site data integration and analysis, accommodating various configurations of aggregation and distribution units and ensuring operability and interpretability of the results.
To enhance cross-site comparability before any federated modeling is conducted, participating sites typically align their data to a common data model, such as the OMOP Common Data Model (CDM),1,34 and implement a shared cohort construction and variable extraction specification. This harmonization process ensures that key variables, coding systems, and temporal definitions are consistently defined across sites prior to distributed estimation.
Following the specifications of a PDA study, Principle 2 emphasizes the importance of a thorough assessment of study cohorts. This step involves confirming their suitability and ensuring that clinical observational data at each unit is fit-for-use before initiating data analysis, which is performed locally. Specifically, prior to deploying the analytical methods, the PDA framework undertakes detailed evaluation such as the cohort diagnostics33 across multiple dimensions, such as the cohort definition level and aggregation unit level, to pinpoint and address any potential misclassification errors within the phenotype data. These diagnostics yield critical insights into the availability of cohorts across participating databases, outcome incidence rates, distributions of participant characteristics, and overlaps among cohorts, among other key quality metrics.
Analysis phase: protect patient confidentiality, unfold heterogeneity, and ensure reliability (Principles 3, 4, 5)
In collaborative multi-site studies, safeguarding confidentiality for both patients and participating institutions is a central priority throughout the analytic lifecycle. Aligned with common privacy and data governance requirements, such as HIPAA35,36 and GDPR.37 the PDA framework safeguards 2 aspects of privacy: hospital-level privacy and patient-level data privacy (Principle 3). To project the identities of hospitals from disclosure, PDA meticulously safeguards the identities of our data partners, particularly in projects where confidentiality is required. For example, in a hospital profiling project aimed at evaluating and ranking hospital performance, where institutions may be hesitant to publicize their rankings, access to sensitive information is restricted to the coordinating center. Individual sites can view only their own rankings and are not provided with the identities or performance results of other participating institutions. Importantly, while hospital identities may be masked in sensitive settings, aggregated study-level characteristics, model specifications, and overall effect estimates remain available to participating sites, allowing investigators to assess applicability to their own populations without compromising institutional confidentiality. Regarding the protection of patient-level data, depending on the protection level required at each DRN, PDA can integrate various protection techniques such as homomorphic encryption38–40 and differential privacy41,42 to ensure security of the summary statistics of individual data transferred across sites, as dictated by the protocol or required by our data partners’ practical needs. The summary statistics will only be used for agreed-upon purposes based on the pre-specified study protocol.
Within the scope of DRNs, a core methodological goal is to understand, systematically account for, and, where appropriate, leverage between-site heterogeneity (Principle 4). This heterogeneity may arise from distribution shift, variations in coding practices, intrinsic population differences, modality disparities, data quality variations, sampling mechanisms, and temporal factors as shown in Figure 2. The heterogeneity is typically assumed at the aggregation unit level. The goal is to produce reliable evidence that accurately represents the intricate nature of real-world healthcare scenarios, thus ensuring that the research findings are not only statistically robust but also widely applicable to a range of clinical settings.
Figure 2.

Examples of between-site heterogeneity. Between-site heterogeneity may arise from differences in population distributions, outcome domains, data modalities, data quality, sampling mechanisms, and study time periods across participating sites. Accounting for these variations is essential for generating robust and generalizable real-world evidence within the PDA framework.
Importantly, the implementation of federated learning methods requires methodological safeguards beyond those in a conventional pooled analysis, which is often infeasible under real-world data-sharing constraints. Because distributed computation and limited information exchange can introduce additional sources of statistical error, Principle 5 emphasizes the use of reliable federated methods to ensure statistical accuracy. Accordingly, PDA studies should prioritize distributed learning approaches with clear theoretical guarantees and strong empirical validation. A key expectation is that, when targeting the same estimand, PDA results should be comparable to those from a benchmark pooled analysis that assumes patient-level data can be centralized. In practice, this expectation is addressed through comprehensive empirical evaluations that explicitly quantify any accuracy loss attributable to distribution, aggregation, or privacy constraints. Across published applications, these benchmarking exercises indicate that PDA methods can achieve near-pooled performance with minimal degradation even under strict privacy and data-sharing constraints, supporting their robustness and suitability for multi-site evidence generation.43–46 In addition, PDA studies routinely incorporate cross-site consistency checks, sensitivity analyses, and replication across independent data partners to evaluate the stability of findings under heterogeneous real-world conditions. The PDA framework also incorporates advanced methods to accommodate between-site heterogeneity, such as ODACH,47 ODACoRH,29 and COLA-GLM-H,43 allowing site-level variation to be explicitly modeled rather than assuming homogeneous data-generating processes across participating units.
Analysis phase: Scalable and efficient communication rounds and implementation readiness with workflow management (Principles 6 and 8)
As multi-site collaborations scale, communication must remain efficient and well-structured to minimize coordination overhead while ensuring reproducibility. Within a PDA study, participating sites share only prespecified summary-level statistics, transmitted through an efficient communication protocol. This design reduces implementation and coordination burden as the number of sites grows (Principle 6). In contrast, a benchmark pooled analysis typically requires transferring patient-level data to a central location. The summaries used in PDA are generally computed at the agreed aggregation-unit level, thereby reducing disclosure risk while preserving the information needed for valid inference. As a result, PDA aligns well with common privacy and governance requirements, including HIPAA35,36 and GDPR.37 Additionally, the PDA framework is designed for scalability and the minimization of iterative burden on participating sites. It is capable of handling analyses that range from just a few units to thousands and accommodating unit sizes from a single patient to one hundred million. This flexibility ensures robust performance across a diverse array of study configurations. The frequency of communication is tailored to the specific demands of each task, depending on the selected algorithm and its unique requirements.
Within the PDA framework, a number of algorithms are purposefully “light-touch,” typically completing in 2 communication rounds: an initialization round led by a designated lead site to coordinate the distributed analysis, followed by a synthesis round in which sites enables the synthesis of results with each site serving as the lead. Such design improves numerical stability and yields more robust estimates. When cross-site synthesis of initial values is beneficial, a brief aggregation round before the distributed analysis can further stabilize initialization.22 Even more efficient are “one-shot” methods, where only a single communication round is needed, such as DLMM,25 COLA-GLM,44,45 and COLA-GLMM.46 By collapsing human-in-the-loop steps and sharply reducing network traffic, these approaches lower latency and cost, preserve privacy, and enable scalable, routine analyses across large, heterogeneous clinical networks.
A key feature of PDA is that it facilitates the implementation readiness and workflow management (Principle 8). The PDA-OTA (Privacy-preserving Distributed Algorithms Over the Air) web-based interface platform has been developed.48 This platform serves as the operational backbone, facilitating smooth and close interactions among collaborators. As illustrated in Figure 3, the PDA-OTA interface provides an operational dashboard for a project involving 8 participating sites in Round 1. The platform allows sites to upload and track summary statistics, monitor participation status, and manage data submission progress in real time. By centralizing communication and managing summary-statistic transfers, the platform streamlines multi-site coordination and supports efficient deployment of distributed learning workflows across diverse study settings. We also acknowledge parallel efforts aimed at principled evidence generation from biomedical data, including approaches developed for multi-site settings, such as the Predictability-Computability-Stability (PCS) framework, Rhino Health, and pSCANNER (patient-centered Scalable National Network for Effectiveness Research).49–51
Figure 3.

Screenshot of PDA-OTA website. An example of transferring summary statistics and monitoring project status in the context of conducting a PDA study.
Analysis phase: Enable capacity of conducting downstream tasks (Principles 7)
Beyond offering distributed learning analytical models and synthesizing information from multiple databases within DRNs, PDA offers users flexible downstream capabilities to extend analyses into diverse and high-impact areas of healthcare research. All downstream analyses are implemented in accordance with the guidelines outlined within the study protocol, ensuring methodological consistency and transparency. In other words, PDA enables adaptive deployment and flexible execution of distributed learning algorithms, tailored to the specific objectives and requirements of each project.
For example, in the context of hospital profiling, which evaluates and compares healthcare delivery across hospitals, PDA supports a “Distributed Hospital Comparer” framework. This framework consists of 2 primary modules: a distributed learning module for fitting Generalized Linear Mixed Effects Model (GLMM), which is certified model for hospital profiling by National Quality Forum,31 and a counterfactual modeling module designed to address the patient-mix variation among different hospitals. Another example under PDA equipped with downstream analysis module is dGEM-disparity,30 a decentralized GLMM designed to quantify health disparities attributable to site-of-care differences.
In addition, PDA supports causal inference and target trial emulation by replicating clinical trial analyses using observational data, providing a scalable and effective alternative to traditional randomized controlled trials. Together, these applications demonstrate PDA’s capacity to deliver rigorous, policy-relevant evidence that informs healthcare quality assessment, guides interventions, and advances equity-focused decision-making.
Reporting phase: transparency (Principle 9)
The PDA framework requires the adherence to transparency principles to facilitate replication, evaluation, and community evaluation. The algorithms developed under the PDA framework are available as open-source resources through an R package and the GitHub repository.18,52 Users are encouraged to make their analysis code and study protocols accessible to data partners, collaborators, and the broader scientific community to enable thorough review and precise reproducibility. In addition, summary statistics and analysis results are required to be made transparent among all participating units within a PDA study, reinforcing openness and scientific integrity throughout the research process.
Reporting phase: Promotion of a collaborative community on clinical implication and results dissemination (Principle 10)
PDA facilitates the dissemination of analysis results, promoting a collaborative community focused on exploring clinical implications and broadening the reach of research findings. The PDA-OTA platform integrates a structured dissemination pipeline to support transparent result sharing and collaborative interpretation across distributed research teams. Within the platform,48 each project instance includes a dedicated “Final Result” module, which serves as a centralized endpoint for posting and synchronizing analysis outputs. The coordinating center or lead site can upload the final results in multiple data-interchange formats (eg, .json, .pdf, .png, as shown in Figure 4), enabling both machine-readable integration and human-readable review. Once uploaded, results are automatically propagated to all participating sites, ensuring version consistency, traceability, and open access within the project network.
Figure 4.

Example of “Final Result” module within PDA-OTA: result for a PDA study on hospital profiling.
From a workflow perspective, the project lead initiates the publication process by compiling the aggregated results into a draft manuscript or technical report. This document undergoes collaborative review and sign-off from all participating sites. The framework also highly encourages the inclusion domain experts to contribute clinical interpretations and insights during post-analysis review. This integrated dissemination architecture not only supports transparency and reproducibility but also accelerates the translation of distributed analytical outputs into actionable clinical and policy knowledge.
Use cases
To demonstrate the feasibility of truly decentralized analyses under the planned distributed analysis (PDA) framework, 3 real-world use cases have been conducted using newly developed federated learning methods: COLA-GLM,45 COLA-GLM-H,44 and COLA-GLMM.46 These studies examined COVID-19 mortality using EHR and medical claims data contributed by multiple partners, including the IBM MarketScan Commercial Database, the Japan Medical Data Center, and the Optum de-identified EHR dataset, collectively representing millions of patients.
The work began with a structured preparation phase in close collaboration with site leads and data stewards. Together, the partners refined the research question, defined the target population and outcomes, and finalized an analysis protocol specifying inclusion and exclusion criteria, covariates, model forms, and planned sensitivity analyses. Governance and privacy requirements were established up front, including a clear delineation of what information would be shared (site-level aggregates only), how transfer would occur, and how analytic decisions would be documented to support transparency and reproducibility.
To ensure harmonization across sites, each partner implemented a common data model workflow (OMOP CDM) and executed a shared extraction specification to create comparable cohorts and variables. Using the PDA R package, each site then ran the prespecified local computations to produce the required summary-level statistics (rather than sharing patient-level data). These site-level aggregates were transmitted to the coordinating center (University of Pennsylvania) and combined using the PDA-OTA aggregation procedure to obtain the overall estimates.
Throughout the process, partners reviewed data quality diagnostics, confirmed that the shared aggregates were consistent with the protocol, and validated model outputs through prespecified checks (for example, plausibility of effect directions, stability across sensitivity analyses, and concordance across data sources). The resulting estimates, such as fixed-effect associations between key risk factors and COVID-19 mortality among hospitalized patients, were reviewed jointly by all data partners and were consistent with existing evidence in the literature. For example, in the COLA-GLM use case, older age, male sex, and histories of diabetes and hypertension were identified as significant risk factors for COVID-19 mortality, consistent with findings reported in the existing literature.53–56 This end-to-end workflow illustrates how PDA enables rigorous multi-site inference while preserving local control of patient-level data, improving transparency of analytic decisions, and supporting reproducible, privacy-preserving evidence generation at scale.
Other use cases that follow the PDA principles have been conducted in additional clinical areas and scientific domains. For example, DLMM25 has been applied in a multi-site international study examining associations between demographic and clinical characteristics and length of hospital stay among patients with COVID-19. In addition, the dGEM-COVID31,57 study used an international, multi-partner setting to perform hospital profiling and evaluate variation in hospital performance. Collectively, these use cases demonstrate the feasibility of implementing PDA in real-world multi-site collaborations and highlight its ability to support rigorous inference and profiling without sharing patient-level data.
Discussion
In response to the challenges of integrating clinical observational data from multiple databases, this paper provides 10 guiding principles for conducting PDA studies aimed at enhancing clinical evidence generation. Adherence to these principles promotes a principled, efficient, collaborative, and effective approach to employing distributed learning algorithms within DRNs for clinical evidence generation. Evidence from prior PDA studies, which have demonstrated strong consistency with existing clinical evidence and pooled analyses, suggests that systematic implementation of these key steps strengthens the reliability and interpretability of multi-site findings. Collectively, these principles also support informed decision-making for policymakers and contribute to the advancement of healthcare research.
Guided by these principles, the PDA framework integrates advanced distributed algorithms designed to tackle practical challenges frequently encountered in real-world applications. For instance, it incorporates methodologies capable of addressing various forms of heterogeneity, such as discrepancies in data capture timing and variations in data quality. A key focus of PDA moving forward is the inclusion of more algorithms that exhibit both lossless and one-shot properties–ideal for collaborative federated learning studies. The lossless feature ensures that estimation and prediction accuracy remain identical to analyses using pooled data, with no loss of accuracy due to data-sharing constraints. Meanwhile, one-shot algorithms require only a single round of communication between participating data partners, further enhancing efficiency. There are several one-shot, lossless algorithms with real-world applications that have been developed, and their implementations are publicly available online.25,45,46 These innovations empower real-time insights in clinical investigations and public health surveillance, improve the efficiency of healthcare services, and accelerate decision-making processes through large-scale collaborative studies.
The use of heterogeneous multi-site data inevitably introduces methodological complexity. The reliability of PDA studies rests on a layered set of safeguards, including prior data harmonization and cohort diagnostics, explicit modeling and evaluation of between-site heterogeneity, and empirical benchmarking against pooled or gold-standard analyses when feasible. These mechanisms are intended to ensure that distributed estimation does not compromise validity. Nevertheless, in settings characterized by extreme distributional shifts, incompatible data structures, or highly specialized research questions, additional modeling strategies or tailored methodological extensions may be required.
At the same time, it is important to distinguish structural barriers to multi-site data integration from intrinsic limitations of observational research. While PDA methods are designed to address privacy constraints, communication burden, and cross-site heterogeneity, they do not eliminate challenges such as residual confounding, selection bias, or measurement error that are inherent to real-world data. These limitations are particularly pronounced in EHR-derived data, which are known to contain substantial data quality issues, including incomplete or inaccurate medication lists, errors and inconsistencies within clinical narrative notes, and clinically relevant information embedded in unstructured formats such as scanned PDFs or free-text documents.58–60 Because these weaknesses originate at the point of data capture rather than at the analytic or data-sharing stage, they propagate into any downstream analysis regardless of whether patient-level data are pooled or analyzed through distributed methods. Addressing them requires upstream safeguards such as phenotype validation, data quality assessment, and triangulation across complementary data sources, in combination with rigorous study design and sensitivity analyses at the analysis stage. To summarize the major structural and methodological barriers to multi-site clinical evidence generation and the corresponding mitigation strategies within the PDA framework, we provide Table S1. Ongoing methodological development will continue to focus on strengthening robustness and interpretability in increasingly complex observational settings.
Trust in clinical evidence generated through distributed methods does not rest solely on adherence to principles, but on empirical validation, transparency, and demonstrated performance in real-world applications. Large-scale collaborative efforts such as the OHDSI LEGEND-T2DM program have shown that rigorously designed real-world evidence studies can reproduce findings consistent with randomized controlled trials across diverse treatment comparisons.61,62 Similarly, PDA-based analyses have been applied in multi-site clinical studies addressing outcomes such as COVID-19,63 opioid use disorder,20 and kidney graft failure,30 with findings evaluated by domain experts and aligned with existing clinical knowledge.
Importantly, the objective of this manuscript is not to claim that distributed studies are inherently infallible, but to articulate a transparent and disciplined framework grounded in methodological theory, large-scale empirical benchmarking, and practical deployment experience across multiple national and international collaborations. By consolidating lessons learned from prior collaborations, statistical method development, and multi-site implementation efforts, these principles aim to reduce the risk of spurious conclusions and to promote reproducible, clinically interpretable evidence generation.
Looking ahead, we in the PDA framework community are committed to generating clinical evidence that clinicians can rely on for patient care decisions, and to expanding across a wide range of fields to address diverse clinical and regulatory questions as part of its future agenda. To connect the evidence generated through PDA, based on its guiding principles, with its clinical and regulatory applications, the framework acknowledges the critical role of collaborating closely with domain experts and regulatory agencies and engaging in in-depth discussions about the insights derived from PDA studies. The knowledge gained from these multi-site studies should enhance patient-centered outcomes analyses, inform clinical and regulatory decision-making processes, and benefit all stakeholders in the healthcare system.
Conclusion
In this paper, we present 10 guiding principles for users to conduct multi-site PDA studies using the algorithms within the PDA framework. Through the strategic synthesis of large-scale clinical observational data from distributed research networks, PDA is dedicated to generating reliable evidence capable of addressing practical research questions or regulatory questions simultaneously. This is achieved with a commitment to transparency, reproducibility, and a systematic approach to data analysis. A number of publications using the PDA framework highlight the successful application of this framework for clinical evidence generation, demonstrating the production of high-quality evidence. The evidence produced using the PDA framework is promising to fill critical gaps in current medical knowledge, offering valuable insights to inform and enhance medical and regulatory decision-making processes.
Supplementary Material
Contributor Information
Yong Chen, The Center for Health AI and Synthesis of Evidence (CHASE), Perelman School of Medicine, The University of Pennsylvania, Philadelphia, PA, 19104, United States; Department of Biostatistics, Epidemiology, and Informatics, Perelman School of Medicine, The University of Pennsylvania, Philadelphia, PA, 19104, United States; The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania, Philadelphia, PA, 19104, United States; Penn Institute for Biomedical Informatics (IBI), Philadelphia, PA, 19104, United States; Leonard Davis Institute of Health Economics, Philadelphia, PA, 19104, United States; Penn Medicine Center for Evidence-based Practice (CEP), Philadelphia, PA, 19104, United States.
Jiayi Tong, The Center for Health AI and Synthesis of Evidence (CHASE), Perelman School of Medicine, The University of Pennsylvania, Philadelphia, PA, 19104, United States; Department of Biostatistics, Epidemiology, and Informatics, Perelman School of Medicine, The University of Pennsylvania, Philadelphia, PA, 19104, United States; Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, 21205, United States.
Yiwen Lu, The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania, Philadelphia, PA, 19104, United States; Penn Institute for Biomedical Informatics (IBI), Philadelphia, PA, 19104, United States.
Rui Duan, Department of Biostatistics, School of Public Health, Harvard University, Boston, MA, 02115, United States.
Chongliang Luo, Division of Public Health Sciences, Department of Surgery, Washington University in St. Louis, St. Louis, MO, 63110, United States.
Marc A Suchard, Department of Veterans Affairs Informatics and Computing Infrastructure, Tennessee Valley Healthcare System VA, Nashville, TN, 37212, United States; Department of Internal Medicine, University of Utah School of Medicine, Salt Lake City, UT, 84112, United States; Department of Biostatistics, University of California, Los Angeles, CA, 90095, United States.
Patrick B Ryan, Epidemiology, Janssen Research & Development, Titusville, NJ, 08560, United States.
Andrew E Williams, Center for Advanced Healthcare Research Informatics, Tufts University School of Medicine, Boston, MA, 02111, United States.
John H Holmes, Department of Biostatistics, Epidemiology, and Informatics, Perelman School of Medicine, The University of Pennsylvania, Philadelphia, PA, 19104, United States.
Jason H Moore, Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, 90048, United States.
Hua Xu, Department of Biomedical Informatics and Data Science, Yale University, New Haven, CT, 06510, United States.
Yun Lu, Center for Biologics Evaluation and Research, Food and Drug Administration, Silver Spring, MD, 20993, United States.
Raymond J Carroll, Department of Statistics, Texas A&M University, College Station, TX, 77840, United States.
Scott L Zeger, Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, 21205, United States.
George Hripcsak, Department of Biomedical Informatics, Columbia University, New York City, NY, 10032, United States.
Martijn J Schuemie, Epidemiology, Janssen Research & Development, Titusville, NJ, 08560, United States.
Author contributions
Yong Chen (Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Writing—original draft, Writing—review & editing), Jiayi Tong (Conceptualization, Software, Supervision, Visualization, Writing—original draft, Writing—review & editing), Yiwen Lu (Visualization, Writing—review & editing), Rui Duan (Conceptualization, Writing—review & editing), Chongliang Luo (Conceptualization, Writing—review & editing), Marc A. Suchard (Writing—review & editing), Patrick Ryan (Writing—review & editing), Andrew E. Williams (Writing—review & editing), John H. Holmes (Writing—review & editing), Jason H. Moore (Writing—review & editing), Hua Xu (Writing—review & editing), Raymond J. Carroll (Writing—review & editing), Yun Lu (Writing—review & editing), Scott L. Zeger (Writing—review & editing), and George Hripcsak (Writing—review & editing), Martijn Schuemie(Writing—review & editing)
Supplementary material
Supplementary material is available at Journal of the American Medical Informatics Association online.
Funding
This work was supported in part by National Institutes of Health (U01TR003709, U24MH136069, U24AG098157).
Conflicts of interest
None declared.
Data availability
This article is a perspective and does not involve primary data. Example resources and tools referenced in this work are publicly available as described in the manuscript.
References
- 1. Hripcsak G, Duke JD, Shah NH, et al. Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform. 2015;216:574-578. [PMC free article] [PubMed] [Google Scholar]
- 2. Collins FS, Hudson KL, Briggs JP, et al. PCORnet: turning a dream into reality. J Am Med Inform Assoc. 2014;21:576-577. 10.1136/amiajnl-2014-002864 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Smelser D, Tromp G, Elmore J, et al. PS2-9: The NIH health care systems research collaboratory. Clin Med Res. 2013;11:152-152. 10.3121/CMR.2013.1176.PS2-9 [DOI] [Google Scholar]
- 4. Health Care Systems (HCS) Research Collaboratory (NIH Collaboratory) | NIH Common Fund. Accessed April 11, 2024. https://commonfund.nih.gov/hcscollaboratory
- 5. BEST Initiative. Accessed January 21, 2025. https://bestinitiative.org/
- 6. FDA’s Sentinel Initiative | FDA. Accessed January 21, 2025. https://www.fda.gov/safety/fdas-sentinel-initiative
- 7. Platt R, Brown JS, Robb M, et al. The FDA Sentinel Initiative—an evolving national resource. N Engl J Med. 2018;379:2091-2093. [DOI] [PubMed] [Google Scholar]
- 8. Brat GA, Weber GM, Gehlenborg N, et al. International electronic health record-derived COVID-19 clinical course profiles: the 4CE consortium. NPJ Digit Med. 2020;3:109-109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. About the Initiative | RECOVER COVID. Accessed July 27, 2022. https://recovercovid.org/
- 10. Haendel MA, Chute CG, Bennett TD, N3C Consortium, et al. The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deployment. J Am Med Inf Assoc. 2021;28:427-443. 10.1093/JAMIA/OCAA196 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Boden WE, O’Rourke RA, Teo KK, et al. ; COURAGE Trial Research Group. Optimal medical therapy with or without PCI for stable coronary disease. N Engl J Med. 2007;356:1503-1516. 10.1056/nejmoa070829 [DOI] [PubMed] [Google Scholar]
- 12. Duckworth W, Abraira C, Moritz T, VADT Investigators, et al. Glucose control and vascular complications in veterans with type 2 diabetes. N Engl J Med. 2009;360:129-139. 10.1056/nejmoa0808431 [DOI] [PubMed] [Google Scholar]
- 13. Blumenthal D, Glaser JP. Information technology comes to medicine. N Engl J Med. 2007;356:2527-2534. 10.1056/NEJMhpr066212 [DOI] [PubMed] [Google Scholar]
- 14. Wu Y, Jiang X, Kim J, et al. Grid Binary LOgistic REgression (GLORE): building shared models without sharing data. J Am Med Inform Assoc. 2012;19:758-764. 10.1136/amiajnl-2012-000862 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Lu C-L, Wang S, Ji Z, et al. WebDISCO: a web service for distributed cox model learning without patient-level data sharing. J Am Med Inform Assoc. 2015;22:1212-1219. 10.1093/jamia/ocv083 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Lu Y, Zhang B, Tong J, et al. Meta-analysis and federated learning over decentralized distributed research networks. Annu Rev Biomed Data Sci. 2025;8:405-421. 10.1146/ANNUREV-BIODATASCI-103123-094441 [DOI] [PubMed] [Google Scholar]
- 17. Li R, Romano JD, Chen Y, et al. Centralized and federated models for the analysis of clinical data. Annu Rev Biomed Data Sci. 2024;7:179-199. 10.1146/ANNUREV-BIODATASCI-122220-115746/CITE/REFWORKS [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. CRAN - Package pda. Accessed June 20, 2022. https://cran.r-project.org/web/packages/pda/index.html
- 19. Duan R, Boland MR, Liu Z, et al. Learning from electronic health records across multiple sites: a communication-efficient and privacy-preserving distributed algorithm. J Am Med Inform Assoc. 2020;27:376-385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Tong J, Chen Z, Duan R, et al. Identifying clinical risk factors for opioid use disorder using a distributed algorithm to combine real-world data from a large clinical data research network. AMIA Annu Symp Proc. 2021;2020:1220. [PMC free article] [PubMed] [Google Scholar]
- 21. Duan R, Chen Z, Tong J, et al. Leverage real-world longitudinal data in large clinical research networks for Alzheimer’s Disease and Related Dementia (ADRD). AMIA Annu Symp Proc. 2020;2020:393-401. [PMC free article] [PubMed] [Google Scholar]
- 22. Duan R, Luo C, Schuemie MH, et al. Learning from local to global-an efficient distributed algorithm for modeling time-to-event data. J Am Med Inform Assoc. 2020;27:1028-1036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Zhang D, Tong J, Jing N, et al. Learning competing risks across multiple hospitals: one-shot distributed algorithms. J Am Med Inform Assoc. 2024;31:1102-1112. 10.1093/JAMIA/OCAE027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Duan R, Boland MR, Moore JH, et al. ODAL: A one-shot distributed algorithm to perform logistic regressions on electronic health records data from multiple clinical sites. Pac Symp Biocomput. 2019;24:30-41. [PMC free article] [PubMed] [Google Scholar]
- 25. Luo C, Islam M, Sheils NE, et al. DLMM as a lossless one-shot algorithm for collaborative multi-site distributed linear mixed models. Nat Commun. 2022;13:1678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Luo C, Islam MN, Sheils NE, et al. dPQL: a lossless distributed algorithm for generalized linear mixed model with application to privacy-preserving hospital profiling. J Am Med Inform Assoc. 2022;29:1366-1371. 10.1093/JAMIA/OCAC067 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Edmondson MJ, Luo C, Nazmul Islam M, et al. Distributed Quasi-Poisson regression algorithm for modeling multi-site count outcomes in distributed data networks. J Biomed Inform. 2022;131:104097. 10.1016/J.JBI.2022.104097 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Hu J, , TongJ, , NingY, et al. Federated feature selection with false discovery rate control. Journal of the Royal Statistical Society Series B: Statistical Methodology. 2025; 10.1093/jrsssb/qkaf074 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Zhang D, Tong J, Stein R, et al. One-shot distributed algorithms for addressing heterogeneity in competing risks data across clinical sites. J Biomed Inform. 2024;150:104595. 10.1016/j.jbi.2024.104595 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Tong J, Shen Y, Xu A, et al. Evaluating site-of-care-related racial disparities in kidney graft failure using a novel federated learning framework. J Am Med Inform Assoc. 2024;31:1303-1312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. dGEMcovid/. Accessed October 4, 2023. https://github.com/ohdsi-studies/dGEMcovid/blob/master/extras/dGEM_Decentralized_Algorithm_for_Generalized_Linear_Mixed_Model_v0.3.docx
- 32. Wang Y, , ZhangD, , TongJ, et al. A communication-efficient federated learning algorithm to assess racial disparities in post-transplantation survival time. J Am Med Inform Assoc. 2025;3212:1916-1926. 10.1093/jamia/ocaf138 40990064 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Rao GA, Shoaibi A, Makadia R, et al. CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization. PLoS One. 2025;20:e0310634. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Marc Overhage J, Ryan PB, Reich CG, et al. Validation of a common data model for active safety surveillance research. J Am Med Inform Assoc. 2012;19:54-60. 10.1136/AMIAJNL-2011-000376 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. The HIPAA Privacy Rule and Adolescents: Legal Questions and Clinical Challenges on JSTOR. Accessed April 23, 2025. https://www.jstor.org/stable/3181198 [DOI] [PubMed]
- 36. Benitez K, Malin B. Evaluating re-identification risks with respect to the HIPAA privacy rule. J Am Med Inform Assoc. 2010;17:169-177. 10.1136/jamia.2009.000026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Gruschka N, Mavroeidis V, Vishi K, et al. Privacy issues and data protection in big data: a case study analysis under GDPR. Proceedings—2018 IEEE International Conference on Big Data, Big Data 2018. 2018:5027-5033. 10.1109/BIGDATA.2018.8622621 [DOI]
- 38. Cheon JH, Kim A, Kim M, et al. Homomorphic encryption for arithmetic of approximate numbers. Advances in cryptology–ASIACRYPT 2017: 23rd international conference on the theory and applications of cryptology and information security, Hong kong, China, December 3-7, 2017, proceedings, Part I 23. Springer 2017:409–437.
- 39. Zhang C, Li S, Xia J, et al. {BatchCrypt}: Efficient homomorphic encryption for {Cross-Silo} federated learning. In: 2020 USENIX Annual Technical Conference (USENIX ATC 20). 2020:493–506.
- 40. Aono Y, Hayashi T, Wang L, et al. Privacy-preserving deep learning via additively homomorphic encryption. IEEE Trans Inf Forensics Secur. 2017;13:1333-1345. [Google Scholar]
- 41. Dwork C. Differential privacy. In: International Colloquium on Automata, Languages, and Programming. Springer, 2006:1-12. [Google Scholar]
- 42. Dwork C. Differential privacy: a survey of results. In: International Conference on Theory and Applications of Models of Computation. Springer, 2008:1-19. [Google Scholar]
- 43. Zhang B, Wu Q, Reps JM, et al. A lossless one-shot distributed algorithm for addressing heterogeneity in multi-site generalized linear models. J Am Med Inform Assoc. 2026;33:700-709. 10.1093/jamia/ocaf198 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Wu Q, Reps JM, Li L, et al. COLA-GLM: collaborative one-shot and lossless algorithms of generalized linear models for decentralized observational healthcare data. NPJ Digit Med. 2025;8:442. 10.1038/s41746-025-01781-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Tong J, Reps JM, Luo C, et al. Unlocking efficiency in real-world collaborative studies: a multi-site international study with one-shot lossless GLMM algorithm. NPJ Digit Med. 2025;8:457. 10.1038/s41746-025-01846-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Chen Y, Dong G, Han J, et al. Regression cubes with lossless compression and aggregation. IEEE Trans Knowl Data Eng. 2006;18:1585-1598. 10.1109/TKDE.2006.196 [DOI] [Google Scholar]
- 47. Luo C, Duan R, Naj AC, et al. ODACH: a one-shot distributed algorithm for Cox model with heterogeneous multi-center data. Sci Rep. 2022;12:6627. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. PDA-OTA. Accessed June 20, 2022. https://pda-ota.pdamethods.org/login
- 49. Rewolinski ZT, Yu B. PCS Workflow for Veridical Data Science in the Age of AI. Published Online First: 3 December 2025. 10.1098/rsta.2024.0605 [DOI]
- 50. Rhino Health Raises $5 Million to Improve AI Workflows in Healthcare Using Federated Learning – ProQuest. Accessed February 10, 2026. https://www.proquest.com/docview/2488020210?pq-origsite=gscholar&fromopenview=true&sourcetype=Scholarly%20Journals
- 51. Ohno-Machado L, Agha Z, Bell DS, et al. ; pSCANNER Team. pSCANNER: patient-centered scalable national network for effectiveness research. J Am Med Inform Assoc. 2014;21:621-626. 10.1136/amiajnl-2014-002751 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. PDA website. https://pdamethods.org/
- 53. Bonanad C, García-Blas S, Tarazona-Santabalbina F, et al. The effect of age on mortality in patients with COVID-19: a meta-analysis with 611,583 subjects. J Am Med Dir Assoc. 2020;21:915-918. 10.1016/j.jamda.2020.05.045 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Nguyen NT, Chinn J, de Ferrante M, et al. Male gender is a predictor of higher mortality in hospitalized adults with COVID-19. PLoS One. 2021;16:e0254066. 10.1371/journal.pone.0254066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Corona G, Pizzocaro A, Vena W, et al. Diabetes is most important cause for mortality in COVID-19 hospitalized patients: Systematic review and meta-analysis. Rev Endocr Metab Disord. 2021;22:275-296. 10.1007/s11154-021-09630-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Du Y, Zhou N, Zha W, et al. Hypertension is a clinically important risk factor for critical illness and mortality in COVID-19: a meta-analysis. Nutr Metab Cardiovasc Dis. 2021;31:745-755. 10.1016/j.numecd.2020.12.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Call for participation in OHDSI network study: dGEM covid—Researchers—OHDSI Forums. Accessed October 4, 2023. https://forums.ohdsi.org/t/call-for-participation-in-ohdsi-network-study-dgem-covid/16485
- 58. Hersh WR, Weiner MG, Embi PJ, et al. Caveats for the use of operational electronic health record data in comparative effectiveness research. Med Care. 2013;51:S30-S37. 10.1097/MLR.0B013E31829B1DBD [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Kahn MG, Callahan TJ, Barnard J, et al. A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. EGEMS (Wash DC). 2016;4:1244. 10.13063/2327-9214.1244 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Weiskopf NG, Weng C. Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research. J Am Med Inform Assoc. 2013;20:144-151. 10.1136/AMIAJNL-2011-000681 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Khera R, Aminorroaya A, Dhingra LS, et al. Comparative effectiveness of second-line antihyperglycemic agents for cardiovascular outcomes: a multinational, federated analysis of LEGEND-T2DM. J Am Coll Cardiol. 2024;84:904-917. 10.1016/J.JACC.2024.05.069 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Khera R, Dhingra LS, Aminorroaya A, et al. Multinational patterns of second line antihyperglycaemic drug initiation across cardiovascular risk groups: federated pharmacoepidemiological evaluation in LEGEND-T2DM. BMJ Med. 2023;2:e000651. 10.1136/bmjmed-2023-000651 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Tong J, Luo C, Islam MN, et al. Distributed learning for heterogeneous clinical data with application to integrating COVID-19 data across 230 sites. NPJ Digit Med. 2022;5:76. 10.1038/s41746-022-00615-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
This article is a perspective and does not involve primary data. Example resources and tools referenced in this work are publicly available as described in the manuscript.
