Abstract
Background
The use of health data supports knowledge-based decision-making in healthcare. Common Data Models (CDMs) and data standards facilitate the integration of diverse data sources and enable federated analysis by harmonizing data formats and terminologies.
Methods
To determine the best approaches to harmonizing patient data, we undertook a comprehensive literature search, which allowed us to identify the most popular and established CDMs (i2b2, Sentinel CDM, PCORnet CDM, OMOP CDM) and data standards (CDA, HL7 version 2, FHIR, openEHR). We established a set of criteria across the categories of Suitability, Popularity, Adaptability, Interoperability, and Support.
Results
The CDMs and data standards are evaluated based on the defined criteria. Overall criteria the OMOP CDM and FHIR scored best. We highlight the strongest CDM and data standard for each criteria category.
Conclusion
Given the unique characteristics, strengths, and weaknesses of each CDM and data standard, no single global representation can be selected. To promote broad adoption of CDMs and data standards, it is essential to enable transformation between different representations and utilize various formats within a single tool to facilitate their interoperability. Only then seamless data exchange and research across borders can be achieved.
Clinical trial number
Not applicable.
Supplementary information
The online version contains supplementary material available at 10.1186/s12911-025-03267-2.
Keywords: Common data model, FAIR-Principles, Data standard, Interoperability, Data harmonization
Background
The use of health data has the potential to significantly improve patient-centered care [1]. This is particularly important given the many challenges faced by healthcare providers and researchers in exchanging health data and conducting large-scale studies across multiple institutions [2]. As such, it is becoming increasingly clear that research on medical data beyond institutional borders is crucial for identifying risk groups, establishing a decision-making framework for pandemic situations, and detecting adverse drug reactions, among other aspects [1, 3]. Moreover, the necessity for a Common Data Model (CDM) for federated learning is emphasized [4]. To achieve these goals, it is necessary to facilitate the reuse and sharing of health data across borders, research institutions and different data providers, including hospitals and medical practices. However, the realization of this goal is not without its difficulties. Various formats, terminologies, and information scopes of collected data make the process of data sharing and reuse highly complex [5]. Moreover, the need to safeguard patient data privacy and security complicates the process. These challenges are being addressed through the development and implementation of FAIR (findability, accessibility, interoperability, and reusability) principles [6]. To this end, there has been a growing interest in the topic of FAIR principles, as evidenced by the increasing number of publications on the subject; see Fig. 1. By highlighting the strengths and weaknesses of various CDM and data standards, we take a step towards the FAIR principles.
Fig. 1.
A keyword search of “fair principles” on scopus [7]. The number of publications focusing on fair principles increased strongly within the last years
A data model represents data elements, its properties and relations, and it serves as a blueprint for organizing data in a structured way. It includes a unified set of metadata that harmonizes data from various sources in a standardized manner. They enable consistent information content across different applications and organizations, promoting integrated and coordinated data, as well as facilitating the federation of research analyses and the aggregation of results. If a data model is a conceptual framework that goes beyond a single use case, it is referred to as a CDM. [8–10].
Data standards can be categorized into two main types: syntactic standards and semantic standards. Semantic standards focus on the meanings of terms, such as specific terminologies. For example, the Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) is a widely recognized and commonly used semantic standard. CDMs often incorporate these semantic standards as standard concepts, significantly enhancing the interoperability of diverse data sources.
Data standards facilitate information flow within national health information infrastructures, enabling clinical and patient safety systems to effectively collect, share, retrieving, and integrate data. A data standard must address the following: defining data elements, standard formats for electronically encoding these data elements, medical terminologies for classification and coding, and methods for knowledge representation for decision support [11]. In this paper the focus lies on the syntactic standard, which defines structure and format, though possible semantic inclusions are discussed for both, syntactic standards and CDMs. By using CDMs and data standards, interoperability can be achieved and data can be shared among different systems [8, 12–20].
Instead of trying to separate the treatment of CDMs and Data Standards, we acknowledge their tight interdependencies and present them with their complementary roles. We advocate for the necessity of bridging these two domains, as their joint application is crucial for effective data management. While the criteria proposed by [15, 21–23] are invaluable, we elaborated popularity, data exchange, and extract, transform and load (ETL) processes more detailed. Transitioning from one data standard or CDM to another can be costly in terms of both time and resources, and may also lead to potential information loss. Therefore, selecting a widely adopted data standard or CDM that allows for easy transformation is of paramount importance. Additionally, we have taken a holistic approach by not relying on a single data set or source type to evaluate the CDMs and data standards. Instead, we conducted a thorough literature search and systematically reviewed the findings. This methodology allows us to encompass a diverse array of use cases, data sources, health systems, and expert fields. Furthermore, we acknowledge that invaluable findings represented in [15, 21–23] may have been outdated, as the field has evolved with new models, tools, and networks.
On these grounds, we discuss CDMs and data standards identified by a literature search. In Section “Methods” the identified CDMs and data standards, are briefly introduced. We separate them based on their purpose into two groups: CDMs for storing data and data analysis (Section “CDMs for Storing Data and Data Analysis”), and data standards for data acquisition and data exchange (Section “Data Standards for Data Acquisition and Data Exchange”). After an initial overview of their characteristics, we present in Section “Criteria Catalog to assess the CDMs and Data Standards” a set of criteria organized into five categories: Suitability, Popularity, Adaptability, Interoperability, and Support. We evaluate each CDM and data standard against these criteria and present the results in Section “Results”. In Section “Discussion” we discuss and illuminate the derived Results.
Methods
We conducted a comprehensive literature search to identify the most popular CDMs and data standards based on their prevalence in Scopus [7]. Following, we briefly introduce the CDMs in Section “CDMs for Storing Data and Data Analysis” and data standards in Section “Data Standards for Data Acquisition and Data Exchange” [15].
Assessment approach
We undertook a comprehensive review of relevant literature, with detailed information provided in the Appendix A. The identified papers were filtered and exclusion criteria were applied; see Fig 2. For each CDM and data standard, a list of references was collected. If a paper covers multiple data standards or CDMs, it was incorporated into the relevant CDMs and data standards.
Fig. 2.
Approach of selecting collection of literature
The CDMs and data standards have been randomly assigned among the authors, allowing us to review our designated list of literature. We have documented the information and references that met the established criteria. The final assessment of the literature was conducted by the authors, based on the collected data.
CDMs for storing data and data analysis
We first introduce the CDMs for storing and analysis, namely the Sentinel CDM (SCDM), PCORnet CDM, i2b2 CDM and the OMOP CDM. Each of these CDMs is a development targeting a specific research question.
Sentinel
In 2007, the US Food and Drug Administration (FDA) requested that drugs be tested based on real-patient data after the drug has been released to the market [24–27]. To this end, the FDA launched the Sentinel Initiative and the Mini-Sentinel CDM was developed in 2008 [28, 29]. Over time, it evolved into the full SCDM. The SCDM focuses on rapid adverse drug event detection, drug safety, and monitoring in the pharmaceutical industry and had been used for hundreds of privacy-preserving analyses [30]. The SCDM structures data in tables, organized in domains. The SCDM and its tools are SAS based [31]. The latest extension of the SCDM was released in 2022 and can be accessed from the Sentinel repository [32].
Pcornet CDM
The Patient Centered Outcomes Research Network (PCORnet) CDM was developed based on the Mini-Sentinel [33, 34] and is funded by the Patient Centered Outcomes Research Institute [35]. With the PCORnet CDM, first released in 2014, PCORnet aims to conduct studies across multiple networks. Its latest version can be accessed on [36]. The table based model can be queried with SAS or SQL.
i2b2
The Informatics for Integrating Biology & the Bedside (i2b2) model is an openly accessible CDM used for data integration and standardization. It was developed in 2004 and has since been used by over 200 institutions [37, 38]. The i2b2 model is constructed as a star schema, i.e., a schema with one central table surrounded by linked tables. Model description, code links and further documentation is available at [39].
OMOP CDM
The Observational Medical Outcomes Partnership (OMOP) CDM was established in response to the same request from the US FDA as Sentinel [10, 40]. Yet, it has expanded to include Electronic Health Records (EHR), medical notes, and other types of health-related data. As a result of the OMOP, a public-private partnership, the Observational Health Data Sciences and Informatics (OHDSI) was established in 2014 [40–42]. In 2011, Overhage et al. identified the first requirements of the OMOP CDM [43]. The SQL-based model is divided into domain-oriented concepts.
Data standards for data acquisition and data exchange
Data standards are used in clinical daily life to record, exchange, and request information on patient level. The four introduced data standards are syntactic standards, i.e., defining the structure and format of the data, and are freely available.
HL7 version 2
Health Level Seven (HL7) is a non-profit, American national standards institute [44]. HL7 has developed multiple data standards. The data standard HL7 version 2 supports hospital workflows from patient registration to hospital logistics. HL7 version 2 is used in more than 35 countries and by 95% of the US healthcare organizations [44]. It is the most widely used healthcare interoperability standard and is heavily utilized by hospitals and healthcare IT suppliers [45].
CDA
The Clinical Document Architecture (CDA), released in 2000, is a version of the HL7 version 3, Reference Information Model, that was redefined and enhanced [46]. The CDA is an XML-based document markup standard that specifies the structure and semantics of clinical documents [44, 45, 47–49]. The main objective of CDA is standardizing the clinical documents that already exist as free text, and its main purpose is to exchange information among those involved in the care of a patient [50]. A CDA document consists of a header and a body. The header provides structured metadata about the document itself, like the context in which the document was created [50]. The header serves to facilitate clinical document exchange and management within and across institutions, and over the course of a patient lifetime. The body contains the informational (factual) statements that make up the actual content of the document.
FHIR
The Fast Healthcare Interoperability Resources (FHIR) were developed by HL7 in 2011. The human-readable standard describes its format and components, also known as resources [45, 51, 52]. FHIR serves in the daily clinical routine by easing the exchange of patient data between parties [19, 53, 54]. With a Representational State Transfer (RESTful) Application Programming Interface (API) the system can be fed with health information. Information can be functional operations (services), a fixed set of information (message), or a fixed package of information (document). As a web API, FHIR has a strong foundation in web standards, such as XML, JSON, HTTP, and OAuth.
openEHR
The openEHR public standard was established in 2003 by the non-profit organization openEHR International and standardized to the EN/ISO 13606 standard series by CEN and ISO. OpenEHR has more than 3200 users worldwide and approximately 30 partners [55]. It is used for reporting EHR as well as research analytics [17, 56]. It applies a 3-level approach, which contains of the reference model, reusable content element definitions, and context-specific data set definitions [57–61]. The content element definitions are building blocks, known as archetypes. Combining archetypes forms a template, which defines a use case specific data structure [62–64]. Examples like the pathology report template can be found on the openEHR website [55]. Its implementation technology specifications include XSDs, JSON-schema, and REST APIs [55].
Criteria catalog to assess the CDMs and data standards
To evaluate the introduced CDMs and data standards, we first define criteria to compare the CDMs and data standards. We have carefully selected three review papers [15, 22, 23] along with the guidance principles from the European Medicines Agency (EMA) [21] that define criteria for CDMs and data standards. This selection allows us to establish a comprehensive set of criteria to evaluate both CDMs and data standards.
The criteria are detailed in Table 1. The phrasing and grouping of criteria has been considered along the following dimensions. Notably, use case-based criteria have been excluded from our analysis, as our work does not rely on a single use case. Additionally, criterion “costs” has not been included, given that all CDMs and data standards are freely accessible. Yet, the SCDM is based on the fee-based Statistical Analysis System (SAS). While the criterion of scalability was mentioned by González-Ferrer et al. [23], it was not evaluated in their work. The criterion for backend-frontend communication was only addressed by González-Ferrer et al. [23]. All other relevant criteria are thoroughly encompassed in our analysis.
Table 1.
Overview of selected criteria within a collection of review paper and guidance principles [15, 21–23]
| Criteria | Source | ||||
|---|---|---|---|---|---|
| González-Ferrer | Garza | Schneeweiss | EMA | Our | |
| et al. [23] | et al. [15] | et al. [22] | [21] | paper | |
| Suitability | ![]() |
![]() |
![]() |
||
| Spread | ![]() |
![]() |
![]() |
||
| Available networks | ![]() |
![]() |
![]() |
||
| Data exchange (ETL) | ![]() |
![]() |
|||
| Freedom to extend | ![]() |
![]() |
![]() |
![]() |
![]() |
| Evolution and maintenance | ![]() |
![]() |
|||
| Terminologies and concepts | ![]() |
![]() |
![]() |
![]() |
|
| Governance | ![]() |
![]() |
![]() |
||
| Data validation | ![]() |
![]() |
![]() |
||
| Ease of Use | ![]() |
![]() |
![]() |
![]() |
![]() |
| Tools | ![]() |
![]() |
![]() |
![]() |
|
| Version Control | ![]() |
![]() |
![]() |
||
| Integrity (use case based) | ![]() |
not applicable | |||
| Querability (use case based) | ![]() |
not applicable | |||
| Cost | ![]() |
not applicable | |||
| Scalability |
(not measured) |
||||
| Backend-frontend | ![]() |
||||
| communication | |||||
After a thorough review of the proposed criteria, we have identified the following criteria organized into five distinct categories:
Suitability to various data sources and use cases: The chosen CDM or data standard should address a wide range of data sources (e.g., EHR, claims data, survey data) to enable cross-cutting use cases and facilitate data reuse.
-
Popularity
-
3.Spread: Even the most suitable CDMs can only be established if it is widely known and used.
-
4.Data networks and secure data exchange: Data networks enable cross institutional research, while satisfying data protection and security declarations [65]. Data networks with many data partners influence the choice of the CDM for storing data and data analysis. On the other hand, data standards should enable an easy, fast, and secure exchange of data between parties. Therefore, standards which intend to exchange sensitive data must provide security standards.
-
5.Data exchange with other CDMs and data standards: The ability to easily transform data between CDMs increases the number of data sources that can be utilized in a study.
-
3.
-
6.
Adaptability
-
7.Freedom to extend: The ability to extend or adapt the CDM minimizes information loss.
-
8.Evolution and maintenance: Continuous evolution and regular maintenance are required to accommodate changing healthcare demands and prevent errors.
-
7.
-
9.
Interoperability
-
10.Terminologies and concepts: Mapping data to standard concepts enables easy unification of data from various sources but may require significant resources and effort, potentially leading to information loss.
-
11.Governance: Strong model governance eases running the same analysis on various data sources or joining data from several sources. It is in conflict with Criterion 3a.
-
12.Data validation: Comprehensive validation tools help to keep data quality.
-
10.
-
13.
Support
-
14.Ease of use: The CDM should be understandable and usable by individuals with different backgrounds, such as data collectors, maintainers, and researchers. Transparent and clear definitions of concepts and rules are necessary to handle data transformation consistently. A strong community support with experience in edge cases and common practices is crucial.
-
15.Tools: Provided tools to support developer and user, for instance, to generate the CDM or message, to develop an ETL process, or to create analysis queries.
-
16.Version control: The ability to track changes in patient history. For instance, conducting a study with the cohort based on their residence might be inaccurate without versioning.
-
14.
Results
Based on the defined criteria, we evaluate each of the CDMs and data standards listed in Section “CDMs for Storing Data and Data Analysis” and Section “Data Standards for Data Acquisition and Data Exchange”. The results of the CDMs are summarized in Table 2, while the evaluation of the data standards is shown in Table 3. In the subsequent Section “Discussion” we discuss and justify all CDM and data standard ratings we have aggregated in Tables 2 and Table 3. For easier accessibility, we have cross-linked the respective detail discussions from the table entries.
Table 2.
Overview of CDM meeting defined criteria
| Criterion | SCDM | PCORnet | i2b2 | OMOP |
|---|---|---|---|---|
| 1 Suitability | + | + | + | ++ |
| 2a Popularity: Spreading | ++ | ++ | ++ | ++ |
| 2b Popularity: Data Networks | ++ | ++ | ++ | ++ |
| 2c Popularity: ETL | 0 | ++ | ++ | ++ |
| 3a Adaption: Freedom to extend | + | + | ++ | + |
| 3b Adaption: Evolvement | ++ | ++ | ++ | ++ |
| 4a Interoperability: Terminologies | + | ++ | + | ++ |
| 4b Interoperability: Governance | + | ++ | + | ++ |
| 4c Interoperability: Interop.: Data validation | ++ | ++ | + | ++ |
| 5a Support: Ease of use | ++ | ++ | + | + |
| 5b Support: Tools | ++ | + | ++ | ++ |
| 5c Support: Version control | 0 | + | 0 | + |
Table 3.
Overview of data standards meeting defined criteria
| Criterion | HL7 v 2 | CDA | FHIR | openEHR |
|---|---|---|---|---|
| 1 Suitability | ++ | ++ | ++ | ++ |
| 2a Popularity: Spreading | ++ | ++ | ++ | ++ |
| 2b Popularity: Data Networks | + | + | ++ | ++ |
| 2c Popularity: ETL | + | + | ++ | ++ |
| 3a Adaption: Freedom to extend | ++ | ++ | ++ | ++ |
| 3b Adaption: Evolvement | + | ++ | ++ | ++ |
| 4a Interoperability: Terminologies | ++ | ++ | ++ | ++ |
| 4b Interoperability: Governance | + | + | ++ | ++ |
| 4c Interoperability: Data validation | ++ | ++ | ++ | ++ |
| 5a Support: Ease of use | + | + | ++ | + |
| 5b Support: Tools | + | ++ | ++ | ++ |
| 5c Support: Version control | 0 | + | ++ | ++ |
As shown in Table 2, the OMOP CDM scored the highest in the overall assessment of the CDMs with a score of 21+, followed by PCORnet with 20+, i2b2 with 17+, and SCDM with 16 + . Criterion 3a (Freedom to extend) is led by i2b2, while Criterion 5a is led by both PCORnet and SCDM. Multiple criteria have received identical scores.
Table 3 shows that FHIR achieved the maximum score (24+) in each category for the data standards. OpenEHR is presented as a strong alternative to FHIR with a score of 23 + . As a predecessor of FHIR, CDA scored 19+, while HL7 v2 scored 16+, indicating that they are less competitive in these categories. Several criteria have obtained identical scores.
Discussion
Criterion 1 (Suitability to various data sources and use cases)
In Table 4 an overview of possible data sources and reported use cases is given. Additionally, Fig. 3 shows the number of tables and fields of the four considered CDMs. The number of tables and the number of fields are first indicators of the domain scope and details. Yet, it is not a direct measurement. Garza et al. [15] obtained a comparable result. They found that the OMOP CDM achieved the highest domain coverage, followed by the Study Data Tabulation Model, which is not part of our study. PCORnet rated third, while SCDM came in fourth. However, unstructured data cannot be included in the PCORnet [93].
Table 4.
Overview of CDM suitability. suitability of data types that can be stored and reported use cases
| CDM | Data | Common Use Cases |
|---|---|---|
| SCDM | EHR, pharmacy service, laboratory results, stroke registries, Medicare Fee-for-Service, text, link to external sources possible, e.g., registries [27, 66–71] | Observational cohort studies, embedded randomized trials, drug safety and monitoring, drug exposure analysis of neonates and infants, analyses and prediction of relevance to nephrology, hospital outcomes (e.g. hospitalized stroke) [30, 68, 69] |
| PCORnet | EHR, administrative data, inpatient, outpatient, emergency department, ancillary service, patient-reported outcomes, area-level social and behavioral determinants [72, 73] | Computable phenotypes for resistant hypertension, stable controlled hypertension, longitudinal EHR studies, drug safety and monitoring, monitoring rates of COVID-19 illness, complications, and correlations [72, 74] |
| i2b2 | EHR, biospecimen data, case report, survey data, cancer registry, anatomic pathology, clinical findings, oncological medication, unstructured text and reports, images, and genomics [39, 75–77] | Clinical trials, monitoring drug safety, epidemiology research, phenotyping, genotyping, registry studies, anatomic pathology studies, medication studies [39, 75–78] |
| OMOP | Clinical data, health system data, health economics data, derived elements, molecular markers, hematology data, notes, images [15, 79–81] | Analyzing longitudinal data, identify high-risk patients, treatment pathway planning for pediatric epilepsy, studies on depression, hypertension, diabetes type II, dementia, drug safety studies, automatic trial matching, predicting risk of smoking score, epidemiology studies, preventing cardio-cerebrovascular disease [10, 15, 41, 59, 80, 82–92] |
Fig. 3.
The number of tables (left) and number of fields (right) in the CDMs. The omop CDM counts the most tables and PCORnet CDM the most fields
In Table 5 data types for the standards are listed. Use cases are not considered since they are developed for clinical daily routine and data exchange.
Table 5.
Overview of data standard suitability. suitability of data types that can be stored and reported use cases
| Data Standard | Data | Additional Information |
|---|---|---|
| HL7 v2 | Clinical data, EHR, laboratory results, genomic data [94] | A message is comprised of segments, which contain mandatory and optional fields [45] |
| CDA | EHR, laboratory reports, and administrative data, tables, lists, and any relevant multimedia data [46, 48–50, 95, 96] | CDA consists of a header, body, and sections. The structure and semantics is defined in templates [46, 48–50] |
| FHIR | Human and veterinary care, clinical trials, public health, genomic, lab results, vital sign, medication request, allergy intolerance [45, 97, 98, 98–100] | Approx. 145 resources are categorized into five groups: Foundation, base, clinical, financial, and specialized [45] |
| OpenEHR | Clinical daily routine, public health data, pharmacy, genomics, clinical registries, EHR, administration [55, 61] | Has a clinical decision-making system [55] |
Taking into account the findings presented by Garza et al. [15], we evaluate the SCDM using a ranking of
. The minimalist i2b2 CDM also receives a
rating due to its limited number of domains. In contrast, while the OMOP CDM is awarded a
, the PCORnet CDM cannot accommodate unstructured data, resulting in a
evaluation for PCORnet. All data standards can hold a wide field of data sources to comply their purpose, evaluated by
.
Criterion 2a (Spread)
Given that the selected CDMs and data standards have been identified based on their popularity, Criterion 2a is trivially satisfied. Further details can be found in Appendix A.
Criterion 2b (Data networks and secure data exchange)
CDMs are categorized into those for analysis and storing, and data standards. According to the European medicines agency’s report [21] the access to the Sentinel network is restricted to the FDA. Additionally, the Sentinel operations center coordinates the network of Sentinel data partners. The Sentinel Distributed Database operates as a distributed network, ensuring that data remains within institutional firewalls. If patient-level data is shared, it is fully anonymized and restricted to essential information. The data is shared using the PopMedNet platform. If needed, text records can be provided, and data can be linked to external registries [71]. For each query, the Sentinel operation center creates an analytic package, which is executed by the data partners. The aggregated data is then returned to the operation center for final analysis [69, 70]. The Sentinel Distributed Database spans 500.1 million unique patient identifiers from 2000 to 2024 and contains 1.3 billion person-years of data with 22.3 billion pharmacy dispensings and 24 billion unique medical encounters [32]. Within the Precise4Q a harmonization framework was developed and used to integrate various heterogeneous stroke-related datasets from institutions throughout Europe [66]. Huang and colleagues adapted the SCDM to the National Health Insurance Research Database in Taiwan [101]. The Oxford Royal College of General Practitioners Clinical Informatics Digital Hub, ORCHID uses the SCDM as well as the OMOP CDM to satisfy the FAIR principles [102].
Like Sentinel, PCORnet is based on the PopMedNet, too. It consists of eleven research networks and one coordination center. The clinical research networks connect among others 337 hospitals, 3,564 primary care practices, and 1,024 community clinics [36, 72, 103–106]. For data security reasons, data is maintained locally by the institutions [107]. Only requests are sent through the distributed research network query portal and only results are returned [108]. Carnahan et. al used billing data, and EHR of three PCORnet networks (Greater Plains Collaborative, OneFlorida, STAR) for Assessing Use of Molecular-Guided Cancer Treatment. Although some limitations were found, several studies suggest that EHR at PCORnet networks were highly complete for their use case [109]. PCORnet offers encrypted, keyed secure hash tokens to match patient records [110].
i2b2 offers a Shared Health Research Information Network, called SHRINE [111]. On that account, i2b2 offers authentication and authorization based on HIPAA guidelines [39, 75, 112]. Some institutions established SHRINE during projects, e.g., Harvard Catalyst, which was decommissioned and augmented by the ACT Network. Additionally, tranSMART was established for sharing, integration, standardization and analysis of diverse data [39, 113, 114]. i2b2 CDM is used by over 200 organizations worldwide [78, 115]. González et.al [76] report from the project InSite, a global clinical research network, with over 130 healthcare providers, Gardner and colleagues [77] introduce a de-identified i2b2 clinical data warehouse with multiple hospitals and more than half a million patients and Castro et. al [75] introduce the Mass General Brigham Biobank Portal with 125,645 patients. Based on tranSMART and i2b2 Johns et. al introduce a data warehousing portal to provide access to a range of data warehouses [116]. The German project Data Integration for Future Medicine makes use of the i2b2 CDM, as well as of tranSMART. Cross-site selection between university hospitals is based on SHRINE [117]. The Accrual to Clinical Trials network of 21 National Clinical and Translational Science Award sites. It deploys a set of i2b2 repositories and complements the PCORnet [111].
To conduct studies involving multiple institutions, with the OMOP CDM the OHDSI network can be utilized. The OHDSI distributed data network comprises over 331 data sources, containing more than 2.1 billion patient records across 34 countries [118]. For network research studies, researcher check the available databases and get in touch with potential collaborators [42]. Aggregated results are shared across the network while patient-level data remains within each institution. The OHDSI GitHub platform provides access to the analysis and aggregated study results of open OHDSI network studies. The Arachne platform automates the network study process [119]. Since 2020, OHDSI collaborates with the European Health Data and Evidence Network (EHDEN) [10, 120, 121].
Data exchange within all HL7 standards are possible. HL7 version 2 allows the exchange of clinical data among systems using messages. HL7 version 2 does not provide security applications or authentication tools. The standard leaves the implementation to the end-user [51]. Whereas FHIR and CDA offer security features, such as security applications and authentication tools [45, 51, 54, 122]. The FHIR API does not impose specific security rules on operations. However, it includes essential building blocks for encryption, TLS (Transport Layer Security), access control, user management, OAuth (Open Authorization) and provenance tracking to support various security approaches [52, 123]. FHIR enables easy exchange of small pieces of information, i.e., resources [124], while HL7 version 2 and CDA usually exchange whole reports and large amounts of data as message or document. The United States, Canada, Argentina, Brazil, Chile, and Colombia are using FHIR for health data exchange [125]. The World Health Organization recommends FHIR as a standard for structuring SMART guidelines [125].
OpenEHR supports individual and population-level queries and enables federated querying across datastores [126]. It has minimal security policies, with the latest API version covering authentication and authorization. For instance, openEHR supports digital signature and access control [55].
Since CDA and HL7 version 2 usually exchange rather whole exports and due to missing security applications and authentication tools for HL7 version 2, we have evaluated them with
, while for all other CDMs and data standards the criterion is completely fulfilled.
Criterion 2c (Data exchange with other CDMs and data standards)
The adoption of CDMs enables efficient data transformation and collaboration. Three primary approaches can be identified in the ETL landscape:
-
17.
Direct ETL Processes: These involve the direct transformation and mapping of data from one CDM to another, such as from i2b2 to FHIR or OMOP. Tools developed for these processes aim to simplify the conversion and ensure data completeness.
-
18.
Query Translation: This approach allows users to query data across different CDMs without the need for data transformation. It facilitates unified querying and analysis, maintaining data in its original format while enabling cross-model compatibility.
-
19.
Layered CDM Integration: Some tools build upon existing CDMs to create new standards or knowledge graphs, integrating various models into a cohesive framework. This method emphasizes interoperability and the ability to leverage multiple data standards simultaneously.
For a comprehensive list of available ETL tools and processes, refer to Table 6.
Table 6.
Tools supporting interoperability between CDMs
| Direct ETL Processes | |||
| Source | Target | Example Project(s) | Publication(s) |
| PCORnet | OMOP CDM | [8, 110] | |
| i2b2 | FHIR | CAMP FHIR, i2b2 | [19, 39, 127–129] |
| RDF | i2b2 | i2b2 | [130, 131] |
| CDA | OMOP | [132] | |
| FHIR | PCORnet, OMOP | [133] | |
| FHIR | OMOP | MIRACUM | [97, 134–137] |
| openEHR | i2b2 | [64] | |
| openEHR | OMOP CDM | [138] | |
| EHR messages | HL7 v2 | [139] | |
| Query Translation | |||
| Source | Target | Example Project(s) | Publication(s) |
| i2b2, openEHR | openEHR, i2b2 | [140] | |
| OMOP | i2b2 | [141] | |
| openEHR, FHIR | FHIR, openEHR | [142] | |
| PCORnet, OMOP | i2b2 | SHRINE | [115, 143] |
| Layered CDM Integration | |||
| Base CDM | Layered CDM | Example Project(s) | Publication(s) |
| OMOP | FHIR | FHIR-Ontop-OMOP, OMOPonFHIR | [144, 145] |
| openEHR | FHIR | GECCO | [63, 146, 147] |
In addition to the ETL processes and query translation tools between two CDMs, the Common Data Model Harmonization (CDMH) project aims to standardize various CDMs like PCORNet, i2b2, OMOP, and Sentinel by mapping them to the Biomedical Research Integrated Domain Group (BRIDG) model, facilitating easier sharing and interpretation of research data. The CDMH FHIR Implementation Guide (IG) focuses on translating observational data into FHIR format, enabling efficient data publication through RESTful APIs and leveraging FHIR tools for enhanced data extraction from clinical systems [44, 148, 149]. Table 6 exposes a notable scarcity of solutions for CDA, HL7 version 2, and Sentinel. Furthermore, no solutions are offered for the more generalized solutions: query translation and layered CDM Integration. Leading to an evaluation of
for CDA and HL7 version 2, and a neutral evaluation
for the SCDM.
Criterion 3a (Freedom to extend)
Criteria 1 address the general defined scope of the CDM, whereas Criterion 3a addresses the ability to extend, which might be a local solution but hamper interoperability (Criteria 4).
In the SCDM project specific data elements can be extracted [21]. However, the project specific data elements are outside the Active Risk Identification and Analysis (ARIA) system [71].
PCORnet offers, beside its core tables, supplementary tables for specific use cases, which do not need to be populated in general [34]. Additionally, user might extend their CDM outside of supplementary tables , e.g., Hornik et. al developed four custom data domains because of their importance to the trial [150].
For OMOP CDM, we also found multiple publications that reported the possibility of extending the model for specific applications, e.g., time-based comorbidity [151], clinical next-generation sequencing data [152], chemotherapy regimens [153], and microbiology lab results [154]. Yet, these extensions can be only used locally. PCORnet as well as OMOP are suggesting to extend the CDM only if it is crucial [34, 42].
The i2b2 CDM is more flexible than the other three CDMs [19, 76, 151]. Making adaptations is easy and intended [75, 78, 115].
CDA, FHIR and openEHR are intended to be adapted [45]. HL7 version 2 offers the z-Segment, which can be designed by the user to satisfy user-specific needs [45]. FHIR is even more flexible than its predecessors [53, 125].
Garza et al. [15] classify the extensibility of new domains within SCDM, PCORnet, and OMOP as straightforward. However, because these extensions might restrict the use of other tools and networks, we evaluate them with
. i2b2 and the four data standards are rated with
.
Criterion 3b (Evolvement and maintenance)
In Fig. 4 and Fig. 5 the number of releases in the last five years are visualized. We notice that all CDMs and data standards were updated regularly, rated with
, except the HL7 version 2, which was only updated once within this five year, suggesting a
evaluation [32, 39, 44, 55, 79].
Fig. 4.
Releases of CDMs in the last 5 years. Regular releases are available [32, 34, 39, 79]
Fig. 5.
Releases of data standards in the last 5 years. With the exception of HL7 version 2, regular releases are available [44, 55]
Criterion 4a
The use of standard concepts, although time-consuming and occasionally leading to generalization, is a significant step towards interoperability.
The SCDM primarily preserves original data values. The guiding principles emphasize minimizing data transformation, mapping, manipulation, and combination to ensure that information loss is kept to a minimum [71, 155]. Sentinel utilizes standard terminologies, like SNOMED CT and the ontology framework BioTopLite2 available at all data partners to effectively use the available data [66, 71]. To the best of our knowledge, there is currently no mapping tool available that supports the creation of mappings specifically tailored to SCDM.
PCORnet CDM users are asked to utilize standard vocabulary and terminologies. PCORnet offers several standard terminologies which can be mapped [156], such as SNOMED CT [157], Logical Observation Identifiers Names and Codes (LOINC) [158], rxNorm [159], and International Statistical Classification of Diseases and Related Health Problems (ICD) [160]. A vocabulary mapping software can support the mapping. If a code can be directly mapped to one of the PCORnet CDM standard concept, there is no need to keep the source code [21, 34].
With i2b2, source terminologies are not needed to be mapped to standard terminologies. However, standard concepts, like ARCH (Accessible Research Commons for Health), can be utilized. Yet, i2b2 does not distinguish between standard and local terminologies [39, 143, 161]. SHRINE enables mappings to standard concepts. For instance, Rasmussen et. al [113] performed a real-time mapping of ontology terms to the local i2b2 instance with SHRINE. Within the EuCanImage projects, the OMOP CDM and i2b2 were assessed with a non-image cancer use case [162]. For i2b2 93.6% could be mapped to a standard vocabulary. For the OMOP CDM 92.6% of the codes could be mapped to standard vocabulary.
To ensure institutional interoperability, OMOP CDM users must use the defined OMOP CDM standard vocabulary, for instance SNOMED CT for diagnoses and rxNorm for drugs, but original codes are kept in a provided source field [10]. Many terminologies are integrated in the OMOP CDM as standard concepts, with implemented hierarchies that often facilitate linkages. [10]. The web application ATHENA provides freely available mappings of commonly used vocabularies [42, 163–165]. However, due to Kent and colleagues [10] the greatest challenges is the absence of mapping for local vocabularies. OHDSI encourages the development of custom mappings to the standard concept if necessary [21, 119, 143]. The literature supports the coverage of terminologies used in OMOP CDM for different use cases from various countries, breast cancer in a multi-centric European study [162], cancer registry data in Germany [166]. Within the Criterion 2c we mentioned an ETL process of Yu and colleagues. In [110] they present mapping results on terminologies and concepts from PCORnet CDM to OMOP CDM. Overall, the vocabulary based mapping results were good. The largest vocabulary based loss, with 0.61% was recorded within the drug concept. Garza et al. obtained similar results, limited on their use case, with a terminology coverage of 100 % [15]. To ensure institutional interoperability, OMOP CDM users must use the defined OMOP CDM standard vocabulary, for instance SNOMED CT for diagnoses and rxNorm for drugs, but original codes are kept in a provided source field [10]. Many terminologies are integrated in the OMOP CDM as standard concepts, with implemented hierarchies that often facilitate linkages. [10]. The web application ATHENA provides freely available mappings of commonly used vocabularies [42, 163–165]. However, due to Kent and colleagues [10] the greatest challenges is the absence of mapping for local vocabularies. OHDSI encourages the development of custom mappings to the standard concept if necessary [21, 119, 143]. The literature supports the coverage of terminologies used in OMOP CDM for different use cases from various countries, breast cancer in a multi-centric European study [162], cancer registry data in Germany [166]. Within the Criterion 2c we mentioned an ETL process of Yu and colleagues. In [110] they present mapping results on terminologies and concepts from PCORnet CDM to OMOP CDM. Overall, the vocabulary based mapping results were good. The largest vocabulary based loss, with 0.61% was recorded within the drug concept. Garza et al. obtained similar results, limited on their use case, with a terminology coverage of 100 % [15].
HL7 provides common vocabulary and allow inclusion of external standards and own coding systems for all three standards [45, 167, 168]. Standard coding systems such as SNOMED CT and LOINC enabling compatibility of the CDA with various health and medical fields [132].
OpenEHR allows binding of well-known terminologies, e.g., SNOMED CT, LOINC, provides a set of own mappings, and allows inclusion of external terminologies [138, 142, 169].
Since the SCDM encourages to keep mapping to a minimum and no mapping tool is available, and i2b2 does not distinguish between local and standard concept, we evaluated them with
. All the other CDMs and data standards could be rated with
, given the terminology policy shown in Table 7.
Table 7.
CDM terminology availability and recommendations
| CDM | Keep Source Voc. | Map to Standard | Mapping Recommendation | Tools |
|---|---|---|---|---|
| SCDM | ![]() |
![]() |
Minimize mapping | ? |
| PCORnet | ![]() |
Use standard terms | ![]() |
|
| i2b2 | either source or standard | Not required | ![]() |
|
| OMOP | ![]() |
![]() |
Mandatory mapping | ![]() |
| HL7 v2 | ![]() |
![]() |
![]() |
|
| CDA | ![]() |
![]() |
![]() |
|
| FHIR | ![]() |
![]() |
![]() |
|
| OpenEHR | ![]() |
![]() |
![]() |
|
Criterion 4b (Governance of the CDM)
Schneeweiss et al. categorize CDMs into two types: organizing CDMs, which aim to preserve the data in its original form as much as possible, with examples including the SCDM, i2b2, and PCORnet. In contrast, mapping CDMs, such as the OMOP CDM, are designed to transform a given dataset into a standardized set of constructs and analyzable variables. [22]. With strong governance, the OMOP CDM perfectly satisfy Criterion 4b when guidelines are followed [113]. PCORnet CDM is also characterized by strong governance, which emphasizes the importance of not adapting or extending the CDM unless it is absolutely necessary [34]. As organizing CDMs the SCDM and i2b2 CDM are more flexible and less governed.
The HL7 version 2 standard, often referred to as the “non-standard standard,” is widely used but still requires customized adaptations. D’Amore et al. discuss the upsides and barriers of data interoperability with CDA, which offers an improvement over HL7 version 3 with its 3-level definition [45, 170]. Due to the multi-level architecture, data with the based on different templates can still have the same archetype path [171]. Within a use case Lingtong et al. [61] present the semantic interoperability from health care data from different countries. The human and machine readable CDA was employed for standardization and interoperability of the message structure [46, 99]. Existing and reusable CDA templates provide mandatory and optional fields [46, 172].
FHIR addresses challenges faced by CDA and enhances interoperability [53, 173]. FHIR4FAIR, a FHIR implementation guide, supports a FAIR implementation and assessment [100]. Mandatory resources ensure interoperability, but specific requirements can be met by custom resources [45, 174].
The separation of data representation and concept expression in openEHR ensures synaptic interoperability [175]. Archetypes are recognized as the gateway to semantic interoperability [142, 176, 177]. Templates re-use archetypes as building blocks. Additionally, archetype nodes can be limited [138].
With the request to not extend the OMOP CDM or PCORnet CDM, we categorize them with a stronger governance as the i2b2 CDM and SCDM. The HL7 version 2 and CDA lack underlying structure, which is given by openEHR and FHIR, evaluated by
.
Criterion 4c (Data validation)
Criterion 4c is satisfied by Sentinel, which implements a consistent process with over 1,200 data checks and offers a data quality toolkit, including Data Quality Metrics [21, 32, 72, 178]. Every time data is updated, it must undergo extensive quality assurance checks [30, 155].
The PCORnet CDM provides a data curation package with validation and quality checks, employing a two-stage process for data ingestion. Data accessible through PCORnet must conform completeness, including diagnosis codes, plausibility, and persistence. Quality checks include conformance to the PCORnet CDM and its standard terminologies [21, 36].
Wagholikar et al. developed an open-source application for importing electronic health data into the i2b2 platform, which includes data validations among other functionalities [38]. However, a standalone validation tool is to our best knowledge not available.
OHDSI offers various software packages for validation and quality checks, such as Achilles and the Data Quality Dashboard [42]. PEDSnet is a national learning health system and network of children’s hospitals, that offers advanced data quality assessment tools. The OMOP CDM can utilize these tools. Additionally, PEDSnet participate in the data characterization process managed by PCORnet [171, 179, 180].
HL7 version 2 offers a conformance class generator, while validation tools are provided for CDA and FHIR [44, 51, 181].
OpenEHR offers conformance testing to evaluate the quality of solutions and data validation conformance [55]. Additionally, data quality assessment tools like openCQA are introduced by the community [171].
Besides i2b2, all other CDMs and data standards providing a validation or conformance tool are rated with
. Wagholikar et al. enable a validation check for i2b2 only within their well maintained ETL pipeline [38] and is therefore evaluated by
.
Criterion 5a
Comprehension can be reached by clear structures, user interfaces, structured and detailed documentation and training. All CDMs for storing and data analysis are table-based and generally easy to understand [15]. Instead of offering a Graphical User Interface (GUI), PopMedNet is integrated into the SCDM. PopMedNet is an open-source informatics platform, designed to support the implementation and operation of distributed health data networks [103]. To explore the SCDM, a synthetic public use file in the SCDM format can be downloaded from the Sentinel website. To be able to run the script, SAS license is required [32, 71]. Sentinel offers documentation within its repository, providing an overview of each table with definitions and data type specifications, along with general guidance [32]. Training materials are freely available on its webpage, and more training can be requested. Sentinel strives to provide the best possible clarity in study planning, execution and reporting [70]. Programs of conducted studies can be accessed through Sentinel [182].
PCORnet offers a centralized access point called “Front Door” that facilitates the submission of research queries to diverse data networks within the PCORnet initiative. Webinars can be accessed without registration [34]. PCORnet CDM guidance is available in the PCORnet forum, which is a GitHub repository of PDF documentations [183]. Additionally, PCORnet states to handle some edge cases, e.g., if the day specific date is not given in the source data [34].
As a patient-centric star schema, the i2b2 is intuitive [114]. All observations, such as diagnoses and medications, are stored in one single table and a public use file is provided [39, 78]. Without comprehensive coding knowledge, i2b2 user can share and query data across institutions with SHRINE and explore data, test hypothesis, and discovery cohorts with the analysis tool tranSMART [143, 184]. i2b2 provides a community wiki that is under development. The i2b2 installation guideline offers detailed information and guidance for some chapters, particularly for developers, while other chapters may be brief or empty. i2b2 encourages participation in working groups and community meetings [39].
OHDSI offers a freely available book, “The Book of OHDSI,” free training through the “EHDEN Academy,” and documentation of the CDM versions [79, 119]. The documentation of some tables, such as the cost table of version 5.4, may be incomplete, and documentation of certain tools may appear outdated, but the majority is detailed and maintained. There are active forums and regular meetings are taking place [42]. Compared to the other CDMs, OMOP CDM contains many tables with linkages between tables. Additionally, OMOP user are encouraged to use the OMOP specific concepts. Consequently, familiarization and transformation is complex [22]. Schneeweiss et. al [22] indicate that analytical tools have faced challenges in generating consistent and valid results.
Data standards have a more intricate structure. HL7 provides specifications and freely available implementation guidance, but the amount of information can be overwhelming. Fee-based training and books are offered by HL7, such as “Principles of Health Interoperability FHIR, HL7 and SNOMED CT” [45]. HL7 version 2 provides predefined XML messages and schemes online [44]. For the CDA, Templates, such as the Continuity of Care Document (CCD) or Consolidated CDA, are provided. Additionally, a freely available implementation guides like ART-DECOR are provided for the CDA [44, 46, 49]. However, it is rather complex and difficult to understand [185]. FHIR is claimed to be more intuitive than CDA and HL7 version 2, with its RESTful API facilitating data input [100, 124, 186]. Public test servers, like HAPI FHIR, a Java API, are provided to support FHIR implementations [60, 181]. A community chat is available for FHIR [44].
openEHR provides clear responsibility alignment through a 3-level approach and offers pre-built data models and pre-integrated platforms for a straightforward and simple architecture [187]. The openEHR Foundation provides a comprehensive information architecture for representing the entirety of EHR content, including pre-built templates [61]. OpenEHR offers the freely available “Guideline Definition Language” and provides a white paper, two short tutorials, and active discussion rooms for support, along with information for development and integration. The provided RESTful API support openEHR users [171, 175, 177]. Low code development of openEHR applications is possible, e.g., a User Interface can be used to form data sets [55, 142]. Even though openEHR offers detailed support and predefined elements, it is known for its complexity. To build a clinical information model, contextual data elements or sensible constraints for archetypes domain and technical experts are needed [171]. Archetypes must be written in the Archetype Definition Language and can be queried by the Archetype Query Language [171, 175].
In summary, SCDM and PCORNet CDM are easy to understand and provide good documentations, whereas i2b2 documentation is incomplete and the OMOP CDM documentation for newer versions sometimes is incomplete and tool documentations are sometimes slightly outdated. HL7 version 2 provides an overwhelming amount of information, CDA and openEHR provide comprehensive documentations and guides but are known for their complexity, therefore we have evaluated them by
. As FHIR is more intuitive and supports the implementations, therefore rated by
.
Criterion 5b (Tools)
Sentinel provides routine querying tools, reporting tools, toolkits, software packages, a web-based data visualization, and the Sentinel Query Builder based on SAS [32]. As base of ARIA, Sentinels routine analytic frameworks, known as Sentinel tools, contain reusable SAS-modules [26, 32, 68, 101, 155, 188]. It was shown that the Sentinel tools are capable of reproducing product-outcome association [69, 70]. Some examples tools are: Medical Product Use Overlap and Data Quality Review and Characterization [26, 101]. Lu et. al [26] developed a publicly available reusable interrupted time series (ITS) tool to assess the impact of regulatory actions on drug use of longitudinal data formatted to the SCDM. Cocoros and colleagues developed a publicly available tool to examine patterns and trends in medication such that the adherence to safe drug use can be assessed [30]. Another tool to analyze manufacturer-level drug utilization was developed by Gagne et. al [188]. Additionally, a data source translation code is available publicly [69].
The PCORnet CDM can make use of tools developed for the Mini SCDM v4.0. In addition, PCORnet provides a curation package, which needs to be passed to become a network partner [36]. Self-developed queries can be distributed through the network by the PopMedNet Query Tool [107]. Waitman et al. [106] introduce a federally compliant, cloud-based data environment for PCORnet and i2b2. i2b2 offers a powerful queering tool [115]. The i2b2 API can be adapted to other CDMs too. For instance, the i2b2 API can be used with PCORnet and OMOP. This offers the great opportunity to use data of these three CDMs together [115, 184]. SHRINE enables population-based research and vocabulary mapping [111, 115].
The i2b2 enable research without comprehensive programming skills and without data leaving the data warehouse [76]. The toolkit Time-based Elixhauser Comorbidity Index (TECI) can be used with i2b2 and OMOP, among others, it can be used for creating mappings between them [151].
OHDSI provides tools to support the ETL process (e.g., USAGI, “White Rabbit”, and “Rabbit in a Hat”) and to apply standardized methods and pipelines for data analysis (e.g., ATLAS and HADES) [10, 42, 189–192].
HL7 offers CDA tools such as C-CDA Online, C-CDA Scorecard, CDA Validator and HL7 C-CDA Online Search Tool [44, 49].
For HL7 version 2 an HL7 application programming interface (HAPI) is provided by HL7 [44].
For FHIR, HL7 offers a range of tools, e.g., for editing, validation, testing, servers, and development, including the Firely editor for constructing FHIR profiles and a tool for mapping of natural language annotations to FHIR [51, 174, 193].
OpenEHR offers open-source and commercial tools for template design and executable medical workflows, e.g., “Better”, a web-based visual design tool [55]. Applications and Software solutions are available for openEHR [171]. For instance HiGHmed introduced a smart infection control system, which connects to an openEHR server and uses standardized, open queries [177].
All CDMs and data standards provide tools to support users and developers. However, to the best of our knowledge, PCORnet supports self-development but lacks complete analysis packages. Additionally, HL7 version 2 is supported solely by HAPI. As a result, we rated these two accordingly by
.
Criterion 5c (Version control)
PCORnet CDM and OMOP CDM version 6 provide limited versioning through history tables [34, 79]. No history tables are available for the SCDM and i2b2 CDM. Sentinel states in their documentation that postal code should be updated by overwriting [32].
HL7 version 2 does not actively support versioning, while CDA provides elements and identifiers for tracking changes [45, 194]. FHIR supports version tracking to prevent conflicts when multiple users edit a resource simultaneously. For large datasets, the “Bulk Data API” is provided as an alternative [51, 195]. When a resource is deleted, a new history entry is created with a “deleted” status, but the resource can be reconstituted as a valid resource. OpenEHR offers change control packages and version control as an integrated part of its architecture [55].
Both FHIR and openEHR present comprehensive solutions, earning them a strong evaluation of
. In contrast, PCORnet CDM, OMOP CDM, and CDA provide more limited capabilities, evaluated by
. Finally, SCDM, i2b2, and HL7v2 do not support versioning features.
Limitations
Our comprehensive literature search was done on Scopus, which holds data from a wide range of journals. It is one of the largest and comprehensive academic databases with more than 97.3 million records, 28,300 active serial titles, and 368,000 books [7]. Details regarding the literature search can be found in the Appendix A. Additionally, we restricted our search to the field of medicine and a selected list of search words. Another choice of citation database, categories or keywords might influence the results. Nevertheless, restrictions were kept to a minimum and the choice of threshold is set low to ensure a comprehensive and unbiased literature search.
We assert that we have gathered all pertinent information to the best of our knowledge, utilizing references identified through a comprehensive literature search. However, we acknowledge that the completeness of the literature and information cannot be guaranteed. Nevertheless, we are confident that the most relevant, popular, and frequently utilized references, tools, networks, and related materials have been considered.
Furthermore, we note that a quantitative assessment was often impractical. For example, the number of tools counted does not necessarily reflect their quality and scope. Specifically, generic ETL solutions tend to have a broader impact compared to ETL solutions designed for smaller, less frequently used databases. Therefore, we have decided against quantitative measurements and assessed based on the knowledge gathered from the comprehensive collection of literature.
Infrastructure or network receptively could not be established for all data standards and CDMs. Therefore, the set up was not included in the list of criteria.
Conclusion
In this article, we discussed various well-established CDMs and data standards, each with its own history, purpose, strengths, and weaknesses. The need for a common data representation has become increasingly apparent in recent years, particularly as diverse health systems across countries require seamless integration.
The COVID-19 pandemic highlighted the critical importance of swift and seamless data exchange and research across borders. For example, FHIR was utilized for the exchange of health information among countries and the Pan American Health Organization, enhancing the surveillance system for adverse events following immunization across the Americas [125]. Chai et al. [196] transformed nine databases from six countries into the OMOP CDM to investigate the reduction of incident cases and the incidence rates of mental health diagnoses during the pandemic.
However, agreeing on a single global common data representation or a national health information exchange network does not seem feasible [185]. Besides the varying scope of application, every CDM and data standard has its strengths and weaknesses, making it better suited to specific use cases or data as highlighted by Schneeweiss et al. [22]. For example, while OMOP CDM accommodates a wide range of data and use cases, its strict adherence to concepts can lead to limited accuracy and information loss during mapping the data with complex relationships [15, 197, 198]. Depending on the specific application and source data, the importance of criteria may vary. If the data has already been standardized into a CDM or a data standard, it is advisable to consider bridging solutions between them rather than opting for costly transformations.
In Table 2 and Table 3 , we indicate whether a CDM or data standard meets various criteria. Some CDMs and data standards particularly excel in certain categories. In its original form, the OMOP CDM accommodates a wide range of data and use cases, while openEHR stands out for its capability to store complex data and support clinical decision-making (Suitability). The popularity of CDMs is largely driven by their accessibility, networking capabilities, and the availability of existing ETL processes. While Sentinel boasts a substantial database, it is primarily limited to the USA. In contrast, OHDSI has collaborators across six continents and maintains a database with over one billion patient records.
Additionally, several ETLs and software solutions for integrating the OMOP CDM are available, contributing to its exceptional Popularity. FHIR places a strong emphasis on data exchange, rigorously tested during the COVID-19 pandemic, providing essential building blocks for security measures. In the category of Adaptability, the i2b2 CDM excels with its flexible representation, enabling easy modifications to meet specific local needs. All data standards are adaptable, with openEHR and FHIR being particularly extendable through their RESTful APIs.
Interoperability presents a critical challenge that CDMs must tackle. The use of standardized terminologies, such as SNOMED CT and LOINC, ensures compatibility. All models we examined support standard terminologies, and ETL processes enhance interoperability, facilitating seamless data transformation between models. In the category of Interoperability, only PCORnet and OMOP meet all three subcriteria within the CDMs. Especially, Garza et al. [15] highlight the strengths of the OMOP CDM within this category. PCORnet provides standard terminologies but does not retain the source code. Conversely, the OMOP CDM enforces strict adherence to its concepts requiring developers to map to these concepts. If local vocabularies cannot be mapped to standard terminologies, due to missing expertise or capacity, this may lead to information loss and failed quality assessments using the data within a distributed network. The data standards FHIR and openEHR also perform equally well in this category. Both data standards enable the inclusion of common standard terminologies and concepts as well as the inclusion of external coding systems. Mandatory FHIR resources enhance interoperability compared to its predecessor, and openEHR eases interoperability by separation of data representation and concept expression. Both offer conformance tools.
Schneeweiss et al. [22] discuss the concept of rule-based transformation and analysis of mapping CDM and its associated complexity. This discussion highlights the need for a comprehensive supporting tool. Therefore, in the final category, Support, the OMOP CDM stands out due to its extensive selection of tools, comprehensive documentation, and limited version control. Among all data standards, FHIR is notable for its user-friendliness, which is crucial for broad adoption among healthcare providers.
Influence of current and future projects
Ongoing and upcoming projects are expected to influence the development and adoption of the CDMs and data standards discussed. In the context of the German and European Union (EU) landscape, we note that the SCDM and PCORNet CDM are not widely used. The tranSMART tool, developed by i2b2, is utilized by the Medical Informatics Initiative [199]. Within the Data Analysis and Real World Interrogation Network (DARWIN) project, the EMA and the European Medicines Regulatory Network have established a coordination center aimed at providing timely and reliable evidence on the use, safety, and effectiveness of medicines, leveraging real-world healthcare databases across the EU [10, 200]. All data within this initiative is standardized to the OMOP CDM. Additionally, EHDEN, a federated network across the EU, adopts the OMOP CDM [120].
Notably, FHIR is prominently featured in several projects and initiatives across Germany and Europe. For instance, the German electronic patient file, the elektronische Patientenakte, is implemented via a FHIR-based Clinical Data Repository. In collaboration with HiGHmed, the Health-X legitimized, open, and federated data platform (dataLOFT) also utilizes FHIR alongside openEHR. HiGHmed highlights the complementary nature of FHIR and openEHR [201, 202]. In contrast, HL7 version 2 and CDA are less common in projects, primarily due to the significant success of FHIR [203].
Outlook
CDMs and data standards are employed to achieve FAIRness and enhance health research beyond borders. However, health systems vary from country to country CDMs and data standards are influenced by the specific data they are designed for, which might challenge the transformation. For instance, transforming billing information of German claims data into the OMOP CDM might result in information loss [204]. Additionally, frequent changes in national terminologies pose significant challenges. To ensure high CDM adoption, qualitative and comprehensive mappings are essential.
Publicly available ETL processes facilitate simple transformations and encourage collaboration among institutions [22]. However, ETLs also have drawbacks, such as increased effort, additional storage requirements, and the risk of information loss. We have discussed methods to maintain data within its representation while utilizing them together. For example, Shrine enables querying across the i2b2 CDM, OMOP CDM, and PCORnet CDM [161]. OMOPonFHIR allows users to treat an OMOP database like a FHIR server [145]. However, we have encountered some challenges using OMOPonFHIR, including performance issues, incorrect vocabulary mappings, and failures when posing new resources. Addressing these drawbacks can make OMOPOnFHIR a powerful tool for collaborating among institutions that use OMOP CDM for data storage and FHIR for transferring the data.
In the future, tools must be enhanced, and performance improved to advance towards FAIR healthcare. Additionally, the integration of emerging technologies, such as artificial intelligence and machine learning, into CDMs and data standards presents exciting opportunities for more robust data analysis and decision-making. Overall, continued collaboration and innovation will be essential for realizing the full potential of CDMs in healthcare.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
We would like to thank the entire project team for their fruitful discussions and invaluable support, particularly Dörte Corr and Johannes Bunk.
Abbreviations
- ARIA
Active Risk Identification and Analysis
- ARCH
Accessible Research Commons for Health
- BRIDG
Biomedical Research Integrated Domain Group
- CDMH
Common Data Model Harmonization
- CDM
Common Data Model
- CCD
Continuity of Care Document
- CDA
Clinical Document Architecture
- dataLOFT
Data legitimized, open and federated
- EHR
Electronic Health Records
- EMA
European Medicines Agency
- ETL
Extracted, Transformed, and Loaded
- EU
European Union
- FAIR
Findability, Accessibility, Interoperability, and Reusability
- FHIR
Fast Healthcare Interoperability Resources
- FDA
US Food and Drug Administration
- HAPI
HL7 Application Programming Interface
- HL7
Health Level Seven
- ICD
International Statistical Classification of Diseases and Related Health Problems
- IG
Implementation Guide
- i2b2
Integrating Biology & the Bedside
- ISO
International Organization for Standardization
- ITS
Interrupted Time Series
- LOINC
Logical Observation Identifiers Names and Codes
- OMOP
Observational Medical Outcomes Partnership
- OHDSI
Observational Health Data Sciences and Informatics
- OAuth
Open Authorization
- SCDM
Sentinel CDM
- SNOMED
Systematized Nomenclature of Medicine
- TECI
Time-based Elixhauser Comorbidity Index
- TLS
Transport Layer Security
- PCORnet
Patient Centered Outcomes Research Network
Author contributions
Authors M.F. and E.T. have collected the data, investigated and reviewed the references. Authors M.F. , and E.T. have done conceptual work. Authors M.F. and E.T. wrote the main manuscript text, and tables. Author M.F. prepared the figures. Authors E.T. and M.W. have supervised. Authors M.F. , E.T., M.W. have reviewed/edited. All three authors have contributed substantially to the manuscript.
Funding
Open Access funding enabled and organized by Projekt DEAL. This work was funded based on a resolution of the German parliament by the German Federal Ministry of Health (BMG, Project KI-FDZ, Funding Code: 2521DAT01B). The funding bodies played no role in the design of the study and collection, analysis, and interpretation of data and in writing the manuscript.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Melissa Finster and Markus Wenzel contributed equally to this work.
References
- 1.Safran C, Bloomrosen M, Hammond WE, Labkoff S, Markel-Fox S, Tang PC, Detmer DE. Toward a national framework for the secondary use of health data: an American medical informatics association white paper. J. Am. Med. Inf. Assoc. 2007;14(1):1–9. 10.1197/jamia.M2273. [DOI] [PMC free article] [PubMed]
- 2.Martin-Sanchez FJ, Aguiar-Pulido V, Lopez-Campos GH, Peek N, Sacchi L. Secondary use and analysis of big data collected for patient care. Yearb Med Inf. 2017;26:28–37. 10.15265/IY-2017-008. [DOI] [PMC free article] [PubMed]
- 3.Trinh NT, Houghtaling J, Bernal FL, Hayati S, Maglanoc LA, Lupattelli A, Halvorsen L, Nordeng HM. Harmonizing Norwegian registries onto omop common data model: mapping challenges and opportunities for pregnancy and covid-19 research. Int J Multiling Med Inf. 2024;191. 10.1016/j.ijmedinf.2024.105602. [DOI] [PubMed]
- 4.Cremonesi F, Planat V, Kalokyri V, Kondylakis H, Sanavia T, Miguel Mateos Resinas V, Singh B, Uribe S. The need for multimodal health data modeling: a practical approach for a federated-learning healthcare platform. J Educ Chang Biomed Inf. 2023;141:104338. 10.1016/j.jbi.2023.104338. [DOI] [PubMed]
- 5.Dugas M, Neuhaus P, Meidt A, Doods J, Storck M, Bruland P, Varghese J. Portal of medical data models: information infrastructure for medical research and healthcare. Database. 2016;2016:1–9. 10.1093/database/bav121. [DOI] [PMC free article] [PubMed]
- 6.Wilkinson MD, Dumontier M, Aalbersberg IJJ, Appleton G, Axton M, Baak A, Blomberg N, Boiten J-W, Silva Santos LB, Bourne PE, Bouwman J, Brookes AJ, Clark T, Crosas M, Dillo I, Dumon O, Edmunds S, Evelo CT, Finkers R, Gonzalez-Beltran A, Gray AJG, Groth P, Goble C, Grethe JS, Heringa J, Hoen PAC, Hooft R, Kuhn T, Kok R, Kok J, Lusher SJ, Martone ME, Mons A, Packer AL, Persson B, Rocca-Serra P, Roos M, Schaik R, Sansone S-A, Schultes E, et al. The fair guiding principles for scientific data management and stewardship. Scientific data 3. 2016. 10.1038/sdata.2016.18.
- 7.Scopus. 2023. https://www.scopus.com/search/form.uri?display=basic#basic.
- 8.Danese MD, Halperin M, Duryea J, Duryea R. The generalized data model for clinical research. Bmc Med Inf Decis. 2019;19(1):1–13. 10.1186/s12911-019-0837-5. [DOI] [PMC free article] [PubMed]
- 9.Ahmadi N, Zoch M, Kelbert P, Noll R, Schaaf J, Wolfien M, Sedlmayr M. Methods used in the development of common data models for health data: scoping review. JMIR Med Inf. 2023;11:45116. 10.2196/45116. [DOI] [PMC free article] [PubMed]
- 10.Kent S, Burn E, Dawoud D, Jonsson P, Torup Østby J, Hughes N, Rijnbeek P, Bouvy J. Common problems, common data model solutions: evidence generation for health technology assessment. PharmacoEconomics. 2020;39. 10.1007/s40273-020-00981-9. [DOI] [PMC free article] [PubMed]
- 11.Aspden P, Corrigan JM, Wolcott J, A, eds.: Patient safety: achieving a new standard for care. Washington, DC: National Academies Press (US; 2004). Chapp. 4. [PubMed]
- 12.Bönisch C, Kesztyüs D, Kesztyüs T. Harvesting metadata in clinical care: a crosswalk between FHIR, omop, CDISC and openEHR metadata. Sci Data. 2022;9. 10.1038/s41597-022-01792-7. [DOI] [PMC free article] [PubMed]
- 13.Chapman M, Curcin V, Sklar. E.I.: a semi-autonomous approach to connecting proprietary ehr standards to FHIR. CoRR abs/1911.12254 (2019.
- 14.Gamal A, Barakat S, Rezk A. Standardized electronic health record data modeling and persistence: a comparative review. J Educ Chang Biomed Inf. 2021;114:103670. 10.1016/j.jbi.2020.103670. [DOI] [PubMed]
- 15.Garza M, Del Fiol G, Tenenbaum J, Walden A, Zozus MN. Evaluating common data models for use with a longitudinal community registry. J Educ Chang Biomed Inf. 2016;64:333–41. 10.1016/J.JBI.2016.10.016. [DOI] [PMC free article] [PubMed]
- 16.Kahn M, Batson D, Schilling L. Data model considerations for clinical effectiveness researchers. Med Care 50 Suppl. 2012;60–67. 10.1097/MLR.0b013e318259bff4. [DOI] [PMC free article] [PubMed]
- 17.Kalra D. Electronic health record standards. Methods Inf Med. 2006;45(1):136–80. [PubMed]
- 18.Liyanage H, Liaw ST, Jonnagaddala J, Hinton W, De Lusignan S. Common data models (CDMs) to enhance international big data analytics: a diabetes use case to compare three CDMs. Stud Health Technol Inf. 2018;255:60–64. 10.3233/978-1-61499-921-8-60. [PubMed]
- 19.Pfaff E, Champion J, Bradford R, Clark M, Xu H, Fecho K, Krishnamurthy A, Cox S, Chute C, Taylor C, Ahalt S. Fast healthcare interoperability resources (fhir) as a meta model to integrate common data models: development of a tool and quantitative validation study. JMIR Med Inf. 2019;7:15199. 10.2196/15199. [DOI] [PMC free article] [PubMed]
- 20.Xu Y, Zhou X, Suehs B, Hartzema A, Kahn M, Moride Y, Sauer B, Liu Q, Moll K, Pasquale M, Nair V, Bate. A.: a comparative assessment of observational medical outcomes partnership and mini-sentinel common data models and analytics: implications for active drug safety surveillance. Drug Saf. 2015;38. 10.1007/s40264-015-0297-5. [DOI] [PubMed]
- 21.Agency EM. A common data model for europe? Why? which? how? Workshop report. 2017, Dec, European Medicines Agency.
- 22.Schneeweiss S, Brown J, Bate A, Trifirò G, Bartels D. Choosing among common data models for real-world data analyses fit for making decisions about the effectiveness of medical products. Clin Pharmacol Ther. 2019;107. 10.1002/cpt.1577. [DOI] [PubMed]
- 23.González-Ferrer A, Peleg M. Understanding requirements of clinical data standards for developing interoperable knowledge-based dss: a case study. Comput Stand Interface. 2015;42:125–36. 10.1016/j.csi.2015.06.002. [Google Scholar]
- 24.Behrman RE, Benner JS, Brown JDS, Mcclellan M, Woodcock J, Platt R. Developing the sentinel system-a national resource for evidence development developing the sentinel system. NEJM Evid. 2011. 10.1056/NEJMp1014427. [DOI] [PubMed] [Google Scholar]
- 25.Platt R, Madre L, Reynolds R, Tilson H. Active drug safety surveillance: a tool to improve public health. Pharmacoepidemiol Drug Saf. 2008;17(12):1175–82. https://onlinelibrary.wiley.com/doi/pdf/10.1002/pds.1668. [DOI] [PubMed]
- 26.Lu C, Hou L, Kolonoski J, Petrone A, Zhang F, Corey C, Huang T, Bradley M. A new analytic tool for assessing the impact of the us food and drug administration regulatory actions. Pharmacoepidemiol Drug Saf. 2022;32. 10.1002/pds.5552. [DOI] [PubMed]
- 27.Krupka D, Graham J, Wilson N, Li A, Landman A, Bhatt D, Nguyen L, Reich A, Gupta A, Zerhouni Y, Capatch K, Concheri K, Weissman J. Transmitting device identifiers of implants from the point of care to insurers: a demonstration project. J Patient Saf. 2021;17:223–30. 10.1097/PTS.0000000000000828. [DOI] [PMC free article] [PubMed]
- 28.Maguire A, Douglas I, Smeeth L, Thompson M. Determinants of cholesterol and triglycerides recording in patients treated with lipid lowering therapy in Uk primary care. Pharmacoepidemiol Drug Saf. 2007;16:228–228. [Google Scholar]
- 29.Mcgraw D, Rosati K, Evans B. A policy framework for public health uses of electronic health data. Pharmacoepidemiol Drug Saf. 2012;21(Suppl. 1):18–22. 10.1002/PDS.2319. [DOI] [PubMed] [Google Scholar]
- 30.Cocoros N, Wagner A, Haynes K, Petrone A, Fazio-Eynullayeva E, Ding Y, Izem R, Lee J, Major J, Nguyen M, Ju J. A new analytic tool developed to assess safe use recommendations. Pharmacoepidemiol Drug Saf. 2019;28. 10.1002/pds.4724. [DOI] [PubMed]
- 31.Analytics Software and Solutions. SAS Institute Inc. 2023. https://www.sas.com/en us/home.html.
- 32.Sentinel Common Data Model. 2023. https://www.sentinelinitiative.org/methods-data-tools.
- 33.Fleurence R, Curtis L, Califf R, Platt R, Selby J, Brown J. Launching pcornet, a national patient-centered clinical research network. J Am Med Inf Assoc: JAMIA. 2014;21. 10.1136/amiajnl-2014-002747. [DOI] [PMC free article] [PubMed]
- 34.PCORnet Common Data Model. 2023. https://pcornet.org/wp-content/uploads/2022/01/PCORnet-Common-Data-Model-v60-2020_10_221.pdf.
- 35.Jhaveri R, John J, Rosenman M. Electronic health record network research in infectious diseases. Clin Ther. 2021;43. 10.1016/j.clinthera.2021.09.002. [DOI] [PMC free article] [PubMed]
- 36.PCORnet CDM. 2023. https://pcornet.org/data/.
- 37.Murphy S, Weber G, Mendis M, Gainer V, Chueh H, Churchill S, Kohane I. Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2). J Am Med Inf Assoc: JAMIA. 2010;17:124–30. 10.1136/jamia.2009.000893. [DOI] [PMC free article] [PubMed]
- 38.Wagholikar KB, Ainsworth L, Zelle D, Chaney K, Mendis M, Klann J, Blood AJ, Miller A, Chulyadyo R, Oates M, Gordon WJ, Aronson SJ, Scirica BM, Murphy SN. I2b2-etl: python application for importing electronic health data into the informatics for integrating biology and the bedside platform. Bioinformatics. 2022;38(20):4833–36. https://academic.oup.com/bioinformatics/articlepdf/38/20/4833/46535010/btac595.pdf. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.i2b2. 2023. https://community.i2b2.org/wiki/display/BUN/i2b2+Common+Data+Model+Documentation.
- 40.Stang PE, Ryan PB, Racoosin JA, Overhage JM, Hartzema AG, Reich C, Welebob E, Scarnecchia T, Woodcock J. Advancing the science for active surveillance: rationale and design for the observational medical outcomes partnership. Ann Of Intern Med. 2010;153(9):600–06. 10.7326/0003-4819-153-9-201011020-00010. [DOI] [PubMed] [Google Scholar]
- 41.Hripcsak G, Ryan P, Duke J, Shah N, Park RW, Huser V, Suchard M, Schuemie M, Defalco F, Perotte A, Banda J, Reich C, Schilling L, Matheny M, Meeker D, Pratt N, Madigan D. Characterizing treatment pathways at scale using the OHDSI network. Proceedings of the National Academy of Sciences. 2016, 201510502) 10.1073/pnas.1510502113. [DOI] [PMC free article] [PubMed]
- 42.OHDSI Collaborators. 2023. https://www.ohdsi.org/who-we-are/collaborators/.
- 43.Overhage JM, Ryan P, Reich C, Hartzema A, Stang P. Validation of a common data model for active safety surveillance research. J Am Med Inf Assoc: JAMIA. 2011;19:54–60. 10.1136/amiajnl-2011-000376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.HL7 (202). http://www.hl7.org.
- 45.Benson T, Grieve G. Principles of FHIR. Springer Nature; 2021. p. 79–102. 10.1007/978-3-030-56883-2.5.
- 46.Ali DB, Ghorbel I, Gharbi N, Hmida KB, Gargouri F, Chaari L. Consolidated clinical document architecture: analysis and evaluation to support the interoperability of Tunisian health systems. In: Chaari L, editor. Digital health approach for predictive, preventive, personalised and participatory medicine. Cham: Springer; 2019. p. 43–52. [Google Scholar]
- 47.Bossenko I, Linna K, Piho G, Ross P. Migration from HL7 cda to FHIR in infectious disease system of Estonia. 2022;299. 10.3233/SHTI220998. [DOI] [PubMed]
- 48.Gonzalez D, García-Vázquez J, Bravo-Zanoguera M, López-Avitia R, Reyna MA, Zermeño Campos N, Gonzalez-Ramirez ML:. Ecg standards and formats for interoperability between mhealth and healthcare information systems: a scoping review. Int J Environ Res And Public Health. 2022;19:11941. 10.3390/ijerph191911941. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Sabutsch S, Frohner M, Kleinoscheg G, Klostermann A, Svec N, Rainer-Sablatnig S, Tanjga N. Elga outpatient clinic report and elga telehealth note: two HL7-CDA-based modular electronic documents. 2022;293. 10.3233/SHTI220350. [DOI] [PubMed]
- 50.Dolin RH, Alschuler L, Beebe C, Biron PV, Boyer SL, Essin D, Kimber E, Lincoln T, Mattison JE. The HL7 clinical document architecture. J. Am. Med. Inf. Assoc. 2001;8(6):552–69. https://academic.oup.com/jamia/article-pdf/8/6/552/2337897/8-6-552.pdf. [DOI] [PMC free article] [PubMed]
- 51.FHIR. 2023. http://hl7.org/fhir/.
- 52.Sreejith R, Senthil S. Smart contract authentication assisted graphmap-based HL7 FHIR architecture for interoperable e-healthcare system. Heliyon. 2023;9(4):15180. 10.1016/j.heliyon.2023.e15180. [DOI] [PMC free article] [PubMed]
- 53.Torab-Miandoab A, Samad-Soltani T, Jodati A, Rezaei P. Interoperability of heterogeneous health information systems: a systematic literature review. Bmc Med Inf Decis. 2023;23. 10.1186/s12911-023-02115-5. [DOI] [PMC free article] [PubMed]
- 54.Bialke M, Geidel L, Hampf C, Blumentritt A, Penndorf P, Schuldt R, Moser F-M, Lang S, Werner P, Stäubert S, Hund H, Albashiti F, Gührer J, Prokosch H-U, Bahls T, Hoffmann W. A FHIR has been lit on gics: facilitating the standardised exchange of informed consent in a large network of university medicine. Bmc Med Inf Decis. 2022;22. 10.1186/s12911-022-02081-4. [DOI] [PMC free article] [PubMed]
- 55.openEHR. 2023. https://openehr.org.
- 56.Chen R, Klein GO, Sundvall E, Karlsson D, Hlfeldt H. Archetype-based conversion of ehr content models: pilot experience with a regional ehr system. Bmc Med Inf Decis. 2009;9(1):1–13. 10.1186/1472-6947-9-33. [DOI] [PMC free article] [PubMed]
- 57.Kryszyn J, Smolik WT, Wanta D, Midura M, Wróblewski P. Comparison of openEhR and Hl7 FHIR standards. Int J Electron And Telecommun. 2023;69(No 1):47–52. 10.24425/ijet.2023.144330.
- 58.Oliveira D, Miranda R, Hak F, Abreu N, Leuschner P, Abelha A, Machado J. ScienceDirect steps towards an healthcare information model based on openEHR. Procedia Comput Sci. 2021;184:893–98. 10.1016/j.procs.2021.04.015.
- 59.Beale T, Heard S. An ontology-based model of clinical information. Stud Health Technol Inf. 2007;129:760–64. [PubMed]
- 60.Cheng K, Pazmino S, Schreiweis B. Etl processes for integrating healthcare Data. Tools And Archit Patterns. 2022;299. 10.3233/SHTI220974. [DOI] [PubMed]
- 61.Min L, Atalag K, Tian Q, Chen Y, Lu X. Verifying the feasibility of implementing semantic interoperability in different countries based on the OpenEHR approach: comparative study of acute coronary syndrome registries. JMIR Med Inf. 2021;9:31288. 10.2196/31288. [DOI] [PMC free article] [PubMed]
- 62.Bosca D, Moner D, Maldonado JA, Robles M. Combining archetypes with fast health interoperability resources in future-proof health information systems. Stud Health Technol Inf. 2015;210:180–84. 10.3233/978-1-61499-512-8-180. [PubMed]
- 63.Rinaldi E, Thun S. From OpenEHR to FHIR and OMOP data model for microbiology findings. IOS Press; 2021. p. 402–06. 10.3233/SHTI210189. [DOI] [PubMed]
- 64.Haarbrandt B, Tute E, Marschollek M. Automated population of an i2b2 clinical data warehouse from an openEHR-based data repository. J Educ Chang Biomed Inf. 2016;63:277–94. 10.1016/j.jbi.2016.08.007. [DOI] [PubMed]
- 65.Qian F, Zhang A. The value of federated learning during and post-covid-19. J Retailing The Int Soc For Qual In Health Care. 2021;33. 10.1093/intqhc/mzab010. [DOI] [PMC free article] [PubMed]
- 66.Abad-Navarro F, Martínez-Costa C. A knowledge graph-based data harmonization framework for secondary data reuse. Comput Methods And Programs In Biomed. 2024;243:107918. 10.1016/j.cmpb.2023.107918. [DOI] [PubMed]
- 67.Desai R, Matheny M, Johnson K, Marsolo K, Curtis L, Nelson J, Heagerty P, Maro J, Brown J, Toh S, Nguyen M, Ball R, Pan G, Wang S, Gagne J, Schneeweiss S. Broadening the reach of the fda sentinel system: a roadmap for integrating electronic health record data in a causal analysis framework. npj digital medicine 4. 2021. 10.1038/s41746-021-00542-0. [DOI] [PMC free article] [PubMed]
- 68.Wu A, McMahon P, Welch E, McMahill-Walraven C, Jamal-Allial A, Gallagher M, Zhang T, Draper C, Kline A, Koerner L, Brown J, Dyke M. Characteristics of new adult users of mepolizumab with asthma in the Usa. BMJ Open Respir Res. 2021;8. 10.1136/bmjresp-2021-001003. [DOI] [PMC free article] [PubMed]
- 69.Adimadhyam S, Barreto E, Cocoros N, Toh S, Brown J, Maro J, Corrigan-Curay J, Pan G, Ball R, Martin D, Nguyen M, Platt R, Li X. Leveraging the capabilities of the fda’s sentinel system to improve kidney care. J Am Soc Nephrol. 2020;31:2506–16. 10.1681/ASN.2020040526. [DOI] [PMC free article] [PubMed]
- 70.Huang T, Welch E, Shinde M, Platt R, Filion K, Azoulay L, Maro J, Platt R, Toh S. Reproducing protocol-based studies using parameterizable tools-comparison of analytic approaches used by two medical product surveillance networks. Clin Pharmacol Ther. 2019;107. 10.1002/cpt.1698. [DOI] [PubMed]
- 71.Brown JS, Maro JC, Nguyen M, Ball R. Using and improving distributed data networks to generate actionable evidence: the case of real-world outcomes in the food and drug administration’s sentinel system. J. Am. Med. Inf. Assoc. 2020;27(5):793–97. 10.1093/jamia/ocaa028. [DOI] [PMC free article] [PubMed]
- 72.Forrest CB, McTigue KM, Hernandez AF, Cohen LW, Cruz H, Haynes K, Kaushal R, Kho AN, Marsolo KA, Nair VP, Platt R, Puro JE, Rothman RL, Shenkman EA, Waitman LR, Williams NA, Carton. T.W. Pcornet© 2020: current state, accomplishments, and future directions. J Clin Epidemiol. 2021;129:60–67. 10.1016/j.jclinepi.2020.09.036. [DOI] [PMC free article] [PubMed]
- 73.Bian J, Lyu T, Loiacono A, Viramontes TM, Lipori G, Guo Y, Wu Y, Prosperi M, George J, Thomas J, Harle CA, Shenkman EA, Hogan W. Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world data. J. Am. Med. Inf. Assoc. 2020;27(12):1999–2010. https://academic.oup.com/jamia/articlepdf/27/12/1999/34838730/ocaa245.pdf. [DOI] [PMC free article] [PubMed]
- 74.McDonough C, Babcock K, Chucri K, Crawford D, Bian J, Modave F, Cooper-DeHoff R, Hogan W. Optimizing identification of resistant hypertension: computable phenotype development and validation. Pharmacoepidemiol Drug Saf. 2020;29. 10.1002/pds.5095. [DOI] [PMC free article] [PubMed]
- 75.Castro V, Gainer V, Wattanasin N, Benoit B, Cagan A, Ghosh B, Goryachev S, Metta R, Park H, Wang D, Mendis M, Rees M, Herrick C, Murphy S. The mass general brigham biobank portal: an i2b2-based data repository linking disparate and high-dimensional patient data to support multimodal analytics. J. Am. Med. Inf. Assoc. 2021;29. 10.1093/jamia/ocab264. [DOI] [PMC free article] [PubMed]
- 76.González L, Perez-Rey D, Alonso E, Hernández G, Serrano-Balazote P, Pedrera-Jiménez M, Cámara A, Schepper K, Crepain T, Claerhout B. Building an i2b2-based population repository for clinical research. Stud Health Technol Inf. 2020;270:78–82. 10.3233/SHTI200126. [DOI] [PubMed]
- 77.Gardner BJ, Pedersen JG, Campbell ME, McClay JC. Incorporating a location-based socioeconomic index into a de-identified i2b2 clinical data warehouse. J. Am. Med. Inf. Assoc. 2019;26(4):286–93. 10.1093/jamia/ocy172. [DOI] [PMC free article] [PubMed]
- 78.Brat G, Weber G, Gehlenborg N, Avillach P, Palmer N, Chiovato L, Cimino J, Waitman L, Omenn G, Malovini A, Moore J, Beaulieu-Jones B, Tibollo V, Murphy S, L’Yi S, Keller M, Bellazzi R, Hanauer D, Serret-Larmande A, Kohane I. International electronic health record-derived covid-19 clinical course profiles: the 4ce consortium. npj digital medicine. 2020;3:109. 10.1038/s41746-020-00308-0. [DOI] [PMC free article] [PubMed]
- 79.OMOP CDM versions. 2023. https://ohdsi.github.io/CommonDataModel/cdm54.html.
- 80.Kim H, Yoo S, Jeon Y, Yi S, Kim S, Choi SA, Hwang H, Kim KJ. Characterization of anti-seizure medication treatment pathways in pediatric epilepsy using the electronic health record-based common data model. Front Neurol. 2020;11:409. 10.3389/FNEUR.2020.00409/FULL. [DOI] [PMC free article] [PubMed]
- 81.Bardenheuer K, Speybroeck MV, Hague C, Nikai E, Price M. Haematology outcomes network in europe (honeur)—a collaborative, interdisciplinary platform to harness the potential of real-world data in hematology. Eur J Haematol. 2022;109:138–45. 10.1111/EJH.13780. [DOI] [PubMed]
- 82.Wang X, Rao W, Chen X, Zhang X, Wang Z, Ma X, Zhang Q. The sociodemographic characteristics and clinical features of the late-life depression patients: results from the beijing anding hospital mental health big data platform. BMC Psychiatry. 2022;22:1–7. 10.1186/S12888-022-04339-7/FIGURES/1. [DOI] [PMC free article] [PubMed]
- 83.Bae WK, Cho J, Kim S, Kim B, Baek H, Song W, Yoo S. Coronary artery computed tomography angiography for preventing cardio-cerebrovascular disease: observational cohort study using the observational health data sciences and informatics’ common data model. JMIR Med Inf. 2022;10(10):e41503. https://medinform.jmir.org/2022/10/e41503. [DOI] [PMC free article] [PubMed]
- 84.Zhou J, Guo C, Ren L, Zhu D, Zhen W, Zhang S, Zhang Q. Gender differences in outpatients with dementia from a large psychiatric hospital in China. BMC Psychiatry. 2022;22:1–5. 10.1186/S12888-022-03852-Z/TABLES/2. [DOI] [PMC free article] [PubMed]
- 85.Castano VG, Spotnitz M, Waldman GJ, Joiner EF, Choi H, Ostropolets A, Natarajan K, McKhann GM, Ottman R, Neugut AI, Hripcsak G, Youngerman BE. Identification of patients with drug-resistant epilepsy in electronic medical record data using the observational medical outcomes partnership common data model. Epilepsia. 2022;63:2981–93. 10.1111/EPI.17409. [DOI] [PubMed]
- 86.Candore G, Hedenmalm K, Slattery J, Cave A, Kurz X, Arlett P. Can we rely on results from iqvia medical research data Uk converted to the observational medical outcome partnership common data model?: a validation study based on prescribing codeine in children. Clin Pharmacol Ther. 2020;107:915. 10.1002/CPT.1785. [DOI] [PMC free article] [PubMed]
- 87.Yi W, Kim BH, Kim M, Kim J, Im M, Ryang S, Kim EH, Jeon YK, Kim SS, Kim IJ. Heart failure and stroke risks in users of liothyronine with or without levothyroxine compared with levothyroxine alone: a propensity score-matched analysis. 2022;32:764–71. https://home.liebertpub.com/thy. [DOI] [PubMed]
- 88.Zhao Y, Wang Y, Wang H, Yan B, Shen F, Peterson KJ, Rocca WA, Sauver JS, Liu H. Annotating cohort data elements with OHDSI common data model to promote research reproducibility. Proceedings - 2018 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2018. 2019, 1310–17) 10.1109/BIBM.2018.8621269.
- 89.Yu Y, Ruddy KJ, Hong N, Tsuji S, Wen A, Shah ND, Jiang G. Adepedia-on-ohdsi: a next generation pharmacovigilance signal detection platform using the OHDSI common data model. J Educ Chang Biomed Inf. 2019;91:103119. 10.1016/J.JBI.2019.103119. [DOI] [PMC free article] [PubMed]
- 90.Meystre SM, Heider PM, Kim Y, Aruch DB, Britten CD. Automatic trial eligibility surveillance based on unstructured clinical data. Int J Multiling Med Inf. 2019;129:13–19. 10.1016/J.IJMEDINF.2019.05.018. [DOI] [PMC free article] [PubMed]
- 91.Liu H, Carini S, Chen Z, Hey SP, Sim I, Weng C. Ontology-based categorization of clinical studies by their conditions. J Educ Chang Biomed Inf. 2022;135:104235. 10.1016/J.JBI.2022.104235. [DOI] [PubMed]
- 92.Kanbar LJ, Dexheimer JW, Zahner J, Burrows EK, Chatburn R, Messinger A, Baker CD, Schuler CL, Benscoter D, Amin R, Pajor N. Standardizing electronic health record ventilation data in the pediatric longterm Mechanical ventilator-dependent population. Pediatr Pulmonol. 2023;58:433–40. 10.1002/PPUL.26204. [DOI] [PubMed]
- 93.Graham J, Iverson A, Monteiro J, Weiner K, Southall K, Schiller K, Gupta M, Simard EP. Applying computable phenotypes within a common data model to identify heart failure patients for an implantable cardiac device registry. Int J Cardiol Heart Vasc. 2022;39:100974. 10.1016/j.ijcha.2022.100974. [DOI] [PMC free article] [PubMed]
- 94.Dolin RH, Gupta R, Newsom K, Heale BSE, Gothi S, Starostik P, Chamala S. Automated HL7v2 lri informatics framework for streamlining genomics-EHR data integration. J Pathol Inf. 2023;14:100330. 10.1016/j.jpi.2023.100330. [DOI] [PMC free article] [PubMed]
- 95.Shanbehzadeh M, Kazemi-Arpanahi H, Mazhab Jafari K, Haghiri H. Coronavirus disease 2019 (covid-19) surveillance system: development of covid-19 minimum data set and interoperable reporting framework. J Educ Health Promot. 2020;9:203. 10.4103/jehp.jehp_456_ 20. [DOI] [PMC free article] [PubMed]
- 96.Sabutsch S, Weigl G. Using HL7 CDA and LOINC for standardized laboratory results in the Austrian electronic health record. J Lab Med. 2018;42(6):259–66. 10.1515/labmed-2018-0315.
- 97.Peng Y, Henke E, Reinecke I, Zoch M, Sedlmayr M, Bathelt F. An ETLprocess design for data harmonization to participate in international research with German real-world data based on FHIR and OMOP CDM. Int J Multiling Med Inf. 2023;169:104925. 10.1016/J.IJMEDINF.2022.104925. [DOI] [PubMed]
- 98.Jung S, Bae S, Seong D, Oh O, Kim Y, Yi B-K. Shared interoperable clinical decision support service for drug-allergy interaction check: implementation study (preprint). JMIR Med Inf. 2022;10. 10.2196/40338. [DOI] [PMC free article] [PubMed]
- 99.Ayaz M, Pasha MF, Alzahrani M, Budiarto R, Stiawan D. Correction: the fast health interoperability resources (fhir) standard: systematic literature review of implementations, applications, challenges and opportunities. JMIR Med Inf. 2021;9:32869. 10.2196/32869. [DOI] [PMC free article] [PubMed]
- 100.Martínez-García A, Cangioli G, Chronaki C, Löbe M, Beyan O, Juehne A, Parra-Calderón C. Fairness for FHIR: towards making health datasets fair using HL7 FHIR. 2022;290. 10.3233/SHTI220024. [DOI] [PubMed]
- 101.Huang K, Lin F-J, Ou H-T, Hsu C-N, Huang L-Y, Wang C-C, Toh S. Building an active medical product safety surveillance system in Taiwan: adaptation of the u.S. sentinel system common data model structure to the national health insurance research database in Taiwan. Pharmacoepidemiol Drug Saf. 2020;30. 10.1002/pds.5168. [DOI] [PubMed]
- 102.Lusignan S, Jones N, Dorward J, Byford R, Liyanage H, Briggs J, Ferreira FIM, Akinyemi O, Amirthalingam G, Bates C, Bernal J, Dabrera G, Eavis A, Elliot A, Feher M, Krajenbrink E, Hoang U, Howsam G, Leach J, Hobbs FD. Oxford royal college of general practitioners clinical informatics digital hub: rapid innovation to deliver extended COVID-19 surveillance and trial platforms (preprint). JMIR Public Health And Surveillance. 2020;6. 10.2196/19773. [DOI] [PMC free article] [PubMed]
- 103.PopMedNet. 2023. https://www.popmednet.org/.
- 104.Forrest CB, McTigue KM, Hernandez AF, Cohen LW, Cruz H, Haynes K, Kaushal R, Kho AN, Marsolo KA, Nair VP, Platt R, Puro JE, Rothman RL, Shenkman EA, Waitman LR, Williams NA, Carton TW. Pcornet 2020: current state, accomplishments, and future directions. J Clin Epidemiol. 2021;129:60–67. 10.1016/j.jclinepi.2020.09.036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Ma Q, Mack M, Shambhu S, Mctigue K, Haynes K. Characterization of bariatric surgery and outcomes using administrative claims data in the research network of a nationwide commercial health plan. BMC Health Serv Res. 2021;21. 10.1186/s12913-021-06074-3. [DOI] [PMC free article] [PubMed]
- 106.Waitman LR, Song X, Walpitage DL, Connolly DC, Patel LP, Liu M, Schroeder MC, VanWormer JJ, Mosa AS, Anye ET, Davis AM. Enhancing PCORnet clinical research network data completeness by integrating multistate insurance claims with electronic health records in a cloud environment aligned with CMS security and privacy requirements. J. Am. Med. Inf. Assoc. 2021;29(4):660–70. https://academic.oup.com/jamia/article-pdf/29/4/660/42897549/ocab269.pdf. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Block J, Bailey C, Gillman M, Lunsford D, Boone-Heinonen J, Cleveland L, Finkelstein J, Horgan C, Jay M, Reynolds J, Sturtevant J, Forrest C. The PCORnet antibiotics and childhood growth study: process for cohort creation and cohort description. Acad Pediatr. 2018;18. 10.1016/j.acap.2018.02.008. [DOI] [PMC free article] [PubMed]
- 108.Davies M, Erickson K, Wyner ZG, Malenfant JM, Rosen R, Brown JS. Software-enabled distributed network governance: the popmednet experience. eGems 4. 2016. 10.13063/2327-9214.1213. [DOI] [PMC free article] [PubMed]
- 109.Carnahan R, Waitman L, Charlton M, Schroeder M, Bossler A, Campbell W, Campbell J, McDowell B, Smith N, Gryzlak B, Chrischilles E. Exploration of PCORnet data resources for assessing use of molecularguided cancer treatment. JCO Clin Cancer Inf. 2020;4:724–35. 10.1200/CCI.19.00142. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Yu Y, Zong N, Wen A, Liu S, Stone DJ, Knaack D, Chamberlain AM, Pfaff E, Gabriel D, Chute CG, Shah N, Jiang G. Developing an etl tool for converting the PCORnet CDM into the omop CDM to facilitate the COVID-19 data integration. J Educ Chang Biomed Inf. 2022;127. 10.1016/J.JBI.2022.104002. [DOI] [PMC free article] [PubMed]
- 111.Visweswaran S, Becich M, D’Itri V, Sendro E, MacFadden D, Anderson N, Allen K, Ranganathan D, Murphy S, Morrato E, Pincus H, Toto R, Firestein G, Nadler L, Reis S. Accrual to clinical trials (act): a clinical and translational science award consortium network. JAMIA Open. 2018;1. 10.1093/jamiaopen/ooy033. [DOI] [PMC free article] [PubMed]
- 112.Murphy SN, Gainer V, Mendis M, Churchill S, Kohane I. Strategies for maintaining patient privacy in i2b2. J. Am. Med. Inf. Assoc. 2011;18(Suppl 1):103–08. 10.1136/amiajnl-2011-000316. [DOI] [PMC free article] [PubMed]
- 113.Rasmussen L, Brandt P, Jiang G, Kiefer R, Pacheco J, Adekkanattu P, Ancker J, Wang F, Xu Z, Pathak J, Luo Y. Considerations for improving the portability of electronic health record-based phenotype algorithms. AMIA Symp. 2020;2019:755–64. [PMC free article] [PubMed]
- 114.Kothari C, Wack M, Hassen-Khodja C, Finan S, Savova G, O’Boyle M, Bliss G, Cornell A, Horn E, Davis R, Jacobs J, Kohane I, Avillach P. Phelan-mcdermid syndrome data network: integrating patient reported outcomes with clinical notes and curated genetic reports. Am J Med Genet Part B: Neuropsychiatr Genet. 2017;177. 10.1002/ajmg.b.32579. [DOI] [PMC free article] [PubMed]
- 115.Klann J, Phillips L, Herrick C, Joss M, Wagholikar K, Murphy S. Web services for data warehouses: OMOP and PCORnet on i2b2. J Am Med Inf Assoc: JAMIA. 2018;25. 10.1093/jamia/ocy093. [DOI] [PMC free article] [PubMed]
- 116.Johns M, Müller A, Wirth F, Prasser F. A comprehensive portal for clinical and translational data warehouses. Studies in health technology and informatics 281. 2021. 10.3233/SHTI210201. [DOI] [PubMed]
- 117.Prasser F, Kohlbacher O, Mansmann U, Bauer B, Kuhn K. Data integration for future medicine (difuture). Methods of information in medicine. 2018;57:57–65. 10.3414/ME17-02-0022. [DOI] [PMC free article] [PubMed]
- 118.Reich C, Ostropolets A, Ryan P, Rijnbeek P, Schuemie M, Davydov A, Dymshyts D, Hripcsak G. Ohdsi standardized vocabularies—a large-scale centralized reference ontology for international data harmonization. J. Am. Med. Inf. Assoc. 2024;31. 10.1093/jamia/ocad247. [DOI] [PMC free article] [PubMed]
- 119.Sciences OHD. Informatics: the book of OHDSI. 2023.
- 120.EHDEN datapartners. 2023. https://www.ehden.eu/datapartners/.
- 121.Maier C, Lang L, Storf H, Vormstein P, Bieber R, Bernarding J, Herrmann T, Haverkamp C, Horki P, Laufer J, Berger F, Höning G, Fritsch H, Schüttler J, Ganslandt T, Prokosch H, Sedlmayr M. Towards implementation of omop in a German university hospital consortium. Appl Clin Inf. 2018;9:54–61. 10.1055/s-0037-1617452. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Kiourtis A, Mavrogiorgou A, Kyriazis D. Health information exchange through a device-to-device protocol supporting lossless encoding and decoding. J Educ Chang Biomed Inf. 2022;134:104199. 10.1016/j.jbi.2022.104199. [DOI] [PubMed] [Google Scholar]
- 123.Rivera Sánchez YK, Demurjian SA, Baihan MS. A service-based rbac and mac approach incorporated into the FHIR standard. Digit Commun And Networks. 2019;5(4):214–25. 10.1016/j.dcan.2019.10.004. [Google Scholar]
- 124.Monteiro SC, Correia RJC. FHIR based interoperability of medical devices. Studies in health technology and informatics 290. 2022. 10.3233/SHTI220027. [DOI] [PubMed]
- 125.Rizzato Lede D, Molina H, Bertoglia Arredondo M, Otzoy D, Benavides A, Donis J, Osorio V, Aguilar C, Bustos J, Diaz Maffini M, Revirol K, Galindo C, Mansilla J, Campos F, Kaminker D, D’agostino M, Pastor D. Using FHIR to support COVID-19. Vaccine Saf Electron Case Rep In Am. 2022;294. 10.3233/SHTI220558.
- 126.Pohjonen H. Chapter 20 - Norway, Sweden, and Finland as forerunners in open ecosystems and openehr. In: Hovenga E, Grain H, editors. Roadmap to successful digital health ecosystems. Helsinki: Academic Press; 2022. p. 457–71. 10.1016/B978-0-12-823413-6.00011-2. [Google Scholar]
- 127.Boussadi A, Zapletal E. A fast healthcare interoperability resources (fhir) layer implemented over i2b2. Bmc Med Inf Decis. 2017;17:120. 10.1186/s12911-017-0513-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Sinaci AA, Gencturk M, Teoman H, Laleci G, Alvarez-Romero C, Martínez-García A, Poblador-Plou B, Carmona-Pírez J, Löbe M, Parra-Calderon C. A data transformation methodology to create findable, accessible, interoperable, and reusable health data: software design, development, and evaluation study. J Med Internet Res. 2023;25:42822. 10.2196/42822. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129.Wagholikar K, Mandel J, Klann J, Wattanasin N, Mendis M, Chute C, Mandl K, Murphy S. Smart-on-FHIR implemented over i2b2. J. Am. Med. Inf. Assoc. 2016;24:079. 10.1093/jamia/ocw079. [DOI] [PMC free article] [PubMed]
- 130.Stöhr M, Majeed R, Günther A. Metadata import from rdf to i2b2. Studies in health technology and informatics. 2018;253:40–44. [PubMed]
- 131.Fasquelle-Lopez J, Raisaro J. An ontology and data converter from rdf to the i2b2 data model. Studies in health technology and informatics 294. 2022. 10.3233/SHTI220477. [DOI] [PubMed]
- 132.Ji H, Kim S, Yi S, Hwang H, Kim J-W, Yoo S. Converting clinical document architecture documents to the common data model for incorporating health information exchange data in observational health studies: cda to CDM. J Educ Chang Biomed Inf. 2020;107:103459. 10.1016/j.jbi.2020.103459. [DOI] [PubMed]
- 133.Lenert L, Ilatovskiy A, Agnew J, Rudsill P, Jacobs J, Weatherston D, Deans K. Automated production of research data marts from a canonical fast healthcare interoperability resource (fhir) data repository: applications to covid-19 research. J. Am. Med. Inf. Assoc. 2021;28. 10.1093/jamia/ocab108. [DOI] [PMC free article] [PubMed]
- 134.DAF-Research profile list and mappings from FHIR to PCORnet CDM and OMOP CDM. 2023. http://hl7.org/fhir/us/daf-research/2017Jan/ daf-research-profile.html.
- 135.Jiang G, Kiefer R, Sharma D, Prud’hommeaux E, Solbrig. H.: a consensusbased approach for harmonizing the OHDSI common data model with HL7 FHIR. Stud Health Technol Inf. 2017;245:887–91. [PMC free article] [PubMed] [Google Scholar]
- 136.Prokosch H-U, Acker T, Bernarding J, Binder H, Boeker M, Börries M, Daumke P, Ganslandt T, Hesser J, Höning G, Neumaier M, Marquardt K, Renz H, Rothkötter H-J, Schade-Brittinger C, Schmücker P, Schüttler J, Sedlmayr M, Serve H, Storf H. Miracum: medical informatics in research and care in university medicine: a large data sharing network to enhance translational research and medical care. Schattauer. 2018. [DOI] [PMC free article] [PubMed]
- 137.Software FHIR to OMOP. 2023. https://www.toolpool-gesundheitsforschung.de/produkte/fhir-omop-etl-prozess.
- 138.Kohler S, Boscá D, Kärcher F, Haarbrandt B, Prinz M, Marschollek M, Eils R. Eos and omocl: towards a seamless integration of openEHR records into the omop common data model. J Educ Chang Biomed Inf. 2023;144:104437. 10.1016/j.jbi.2023.104437. [DOI] [PubMed] [Google Scholar]
- 139.Ayaz M, Pasha M, Le T, Alahmadi T, Abdullah N, Alhababi Z. A framework for automatic clustering of ehr messages using a spatial clustering approach. Healthcare. 2023;11:390. 10.3390/healthcare11030390. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 140.Fette G, Kaspar M, Liman L, Dietrich G, Ertl M, Krebs J, Puppe F. Query translation between openEHR and i2b2. Studies in health technology and informatics. 2019;258:16–20. [PubMed]
- 141.Majeed R, Fischer P, Günther A. Accessing omop common data model repositories with the i2b2 webclient – algorithm for automatic query translation. Stud Health Technol Inf. 2021. 10.3233/SHTI210077. [DOI] [PubMed]
- 142.Meredith J, Whitehead N, Dacey M. Aligning semantic interoperability frameworks with the foxs stack for fair health data. Methods Inf Med. 2022. 10.1055/a-1993-8036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Klann J, Abend A, Raghavan V, Mandl K, Murphy S. Data interchange using i2b2. J. Am. Med. Inf. Assoc. 2016;23:188. 10.1093/jamia/ocv188. [DOI] [PMC free article] [PubMed]
- 144.Xiao G, Pfaff E, Prud’hommeaux E, Booth D, Sharma DK, Huo N, Yu Y, Zong N, Ruddy KJ, Chute CG, Jiang G. FHIR-Ontop-OMOP: building clinical knowledge graphs in FHIR rdf with the omop common data Model. J Educ Chang Biomed Inf. 2022;134(September):104201. 10.1016/j.jbi.2022.104201. [DOI] [PMC free article] [PubMed]
- 145.OMOPonFHIR. 2023. https://omoponfhir.org/.
- 146.Ladas N, Franz S, Haarbrandt B, Sommer K, Kohler S, Ballout S, Apfel-Starke J, Marschollek M, Gietzelt M. openEHR-to-FHIR: converting openEHR compositions to fast healthcare interoperability resources (FHIR) for the German corona consensus dataset (GECCO. 2022;289. 10.3233/SHTI210963. [DOI] [PubMed]
- 147.Rajput AM, Brakollari I. Mapping of openEHR archetypes to FHIR resources in use case oncology. 2021;285:285–87. 10.3233/SHTI210616. [DOI] [PubMed]
- 148.Laleci Erturkmen GB, Yuksel M, Dogac A. Providing semantic interoperability between clinical care and clinical research domains. IEEE transactions on information technology in biomedicine : a publication of the IEEE Engineering in Medicine and Biology Society 17. 2012. 10.1109/TITB.2012.2219552.
- 149.Common data model harmonization (cdmh) and open standards for evidence generation: final report. Technical report, food and Drug administration (fda). 2020. https://aspe.hhs.gov/sites/default/files/private/pdf/259016/CDMH-Final-Report-14August2020.pdf, National Institutes of Health’s National Library of Medicine (NLM), National Cancer Institute (NCI) and National Center for Advancing Translational Sciences (NCATS), Office of the National Coordinator for Health Information Technology (ONC).
- 150.Hornik C, Atz A, Bendel C, Chan F, Downes K, Grundmeier R, Fogel B, Gipson D, Laughon M, Miller M, Smith M, Livingston C, Kluchar C, Heath A, Jarrett C, McKerlie B, Patel H, Hunter C. Creation of a multicenter pediatric inpatient data repository derived from electronic health records. Appl Clin Inf. 2019;10:307–15. 10.1055/s-0039-1688477. [DOI] [PMC free article] [PubMed]
- 151.Syed S, Baghal A, Prior F, Zozus M, Al-Shukri S, Syeda HB, Garza M, Begum S, Gates K, Syed M, Sexton KW. Toolkit to compute timebased elixhauser comorbidity indices and extension to common data models. Healthc Inf Res. 2020;26:193. 10.4258/HIR.2020.26.3.193. [DOI] [PMC free article] [PubMed]
- 152.Shin SJ, You SC, Roh J, Park YR, Park RW. Genomic common data model for biomedical data in clinical practice. Stud Health Technol Inf. 2019;264:1843–44. 10.3233/SHTI190676. [DOI] [PubMed]
- 153.Warner JL, Dymshyts D, Reich CG, Gurley MJ, Hochheiser H, Moldwin ZH, Belenkaya R, Williams AE, Yang PC. Hemonc: a new standard vocabulary for chemotherapy regimen representation in the omop common data model. J Educ Chang Biomed Inf. 2019;96:103239. 10.1016/J.JBI.2019.103239. [DOI] [PMC free article] [PubMed]
- 154.Theron E, Gorse JF, Gansel X. Usability of omop common data model for detailed lab microbiology results. Stud Health Technol Inf. 2022;294:292–96. 10.3233/SHTI220461. [DOI] [PubMed]
- 155.Platt R, Platt R, Brown J, Henry D, Klungel OH, Suissa S. How pharmacoepidemiology networks can manage distributed analyses to improve replicability and transparency and minimize bias. Pharmacoepidemiol Drug Saf. 2019;29. 10.1002/pds.4722. [DOI] [PubMed]
- 156.Addison J, Razzaghi H, Bailey C, Dickinson K, Corathers S, Hartley D, Utidjian L, Carle A, Rhodes E, Alonso G, Haller M, Gannon A, Indyk J, Arbelaez A, Shenkman E, Forrest C, Eckrich D, Magnusen B, Davies S, Walsh K. Testing an automated approach to identify variation in outcomes among children with type 1 diabetes across multiple sites. Pediatr Qual Saf. 2022;7:602. 10.1097/pq9.0000000000000602. [DOI] [PMC free article] [PubMed]
- 157.SNOMED. 2023. https://www.snomed.org/.
- 158.LOINC. 2023. https://loinc.org/.
- 159.RxNorm. 2023. https://www.nlm.nih.gov/research/umls/rxnorm/index.html.
- 160.International Statistical Classification of Diseases and Related Health Problems. 2023. https://www.who.int/standards/classifications/classification-of-diseases.
- 161.Klann JG, Joss MAH, Embree K, Murphy SN. Data model harmonization for the all of us research program: transforming i2b2 data into the omop common data model. PLoS One. 2019;14. 10.1371/journal.pone.0212463. [DOI] [PMC free article] [PubMed]
- 162.Frid S, Bracons Cucó G, Gil Rojas J, López-Rueda A, Pastor Duran X, Martínez-Sáez O, Lozano-Rubí R. Evaluation of omop CDM, i2b2 and icgc argo for supporting data harmonization in a breast cancer use case of a multicentric European ai project. J Educ Chang Biomed Inf. 2023;147. 10.1016/j.jbi.2023.104505. [DOI] [PubMed]
- 163.Athena. 2023. https://athena.ohdsi.org/search-terms/start.
- 164.Reich C, Ryan PB, Stang PE, Rocca M. Evaluation of alternative standardized terminologies for medical conditions within a network of observational healthcare databases. J Educ Chang Biomed Inf. 2012;45(4):689–96. 10.1016/J.JBI.2012.05.002. [DOI] [PubMed] [Google Scholar]
- 165.Voss E, Makadia R, Matcho A, Ma Q, Knoll C, Schuemie M, Defalco F, Londhe A, Zhu V, Ryan P. Feasibility and utility of applications of the common data model to multiple, disparate observational health databases. J Am Med Inf Assoc: JAMIA. 2015;22. 10.1093/jamia/ocu023. [DOI] [PMC free article] [PubMed]
- 166.Carus J, Nürnberg S, Ückert F, Schlüter C, Bartels S. Mapping cancer registry data to the episode domain of the observational medical outcomes partnership model (omop). Appl Sci. 2022;12(8). 10.3390/app12084010.
- 167.Rafee A, Riepenhausen S, Neuhaus P, Meidt A, Dugas M, Varghese J. Elapro, a loinc-mapped core dataset for top laboratory procedures of eligibility screening for clinical trials. BMC Med Res Methodol. 2022;22. 10.1186/s12874-022-01611-y. [DOI] [PMC free article] [PubMed]
- 168.Wiedekopf J, Ulrich H, Drenkhahn C, Kock-Schoppenhauer A-K. Ingenerf, J.: termicron – bridging the gap between FHIR terminology servers and metadata repositories. Stud Health Technol Inf. 2021;290:71–75. 10.3233/SHTI220034. [DOI] [PubMed]
- 169.Alkarkoukly S, Mateen A. An openEHR virtual patient template for pancreatic cancer. 2021;285. 10.3233/SHTI210618. [DOI] [PubMed]
- 170.D’Amore J, Bouhaddou O, Mitchell S, Li C, Leftwich R, Turner T, Rahn M, Donahue M, Nebeker J. Interoperability progress and remaining data quality barriers of certified health information technologies. AMIA. Annual Symposium proceedings. AMIA Symposium 2018, 358–67 2018. [PMC free article] [PubMed]
- 171.Tute E, Scheffner I, Marschollek M. A method for interoperable knowledgebased data quality assessment. BMC medical informatics and decision making 21. 2021. 10.1186/s12911-021-01458-1. [DOI] [PMC free article] [PubMed]
- 172.Goenaga I, Lahuerta X, Atutxa A, Gojenola K. A section identification tool: towards HL7 CDA/CCR standardization in Spanish discharge summaries. J Educ Chang Biomed Inf. 2021;121:103875. 10.1016/j.jbi.2021.103875. [DOI] [PubMed]
- 173.Traxler B, Helm E, Krauss O, Schuler A, Kueng J. Towards semantic Interoperability in health data management facilitating process mining. 2020;424–36. 10.4018/978-1-7998-1204-3.ch023.
- 174.Scheible R, Caliskan D, Fischer P, Thomczyk F, Zabka S, Schneider H, Boeker M, Schulz S, Prokosch H-U, Gulden C. AHD2FHIR: a tool for mapping of natural language annotations to fast healthcare interoperability resources – A Technical Case Report, vol. 2022;290. 10.3233/SHTI220026. [DOI] [PubMed]
- 175.Bode L, Schamer S, Boehnke J, Group E, Bott O, Marschollek M, Jack T, Wulff A. Tracing the progression of sepsis in critically ill children: clinical decision support for detection of hematologic dysfunction. Appl Clin Inf. 2022;10.1055/a-1950-9637. [DOI] [PMC free article] [PubMed]
- 176.Xudong L, Shan N, Hailing C, Yexuan C, Mengyang L, Li W, Ling-Tong M. Chapter 18 - the road to interoperability: openEHR modelling and implementation. In: Hovenga E, Grain H, editors. Roadmap to successful digital health ecosystems. Zhejiang: Academic Press; 2022. p. 415–35. 10.1016/B978-0-12-823413-6.00027-6.
- 177.Wulff A, Biermann P, Landesberger T, Baumgartl T, Schmidt C, Alhaji A, Schick K, Waldstein P, Zhu Y, Krefting D, Scheithauer S, Marschollek M. Tracing COVID-19 infection chains within healthcare institutions. Another Brick In The Wall Against sars-Cov-2. 2022;290. 10.3233/SHTI220168. [DOI] [PubMed]
- 178.Huser V, Defalco F, Schuemie M, Ryan P, Shang N, Velez M, Park RW, Boyce R, Duke J, Khare R, Utidjian L, Bailey C. Multisite evaluation of a data quality tool for patient-level clinical data sets. eGems (generating evidence and methods to improve patient outcomes. 2016;4. 10.13063/2327-9214.1239. [DOI] [PMC free article] [PubMed]
- 179.PEDSnet Data Quality Program. 2024. https://pedsnet.org/data/data-quality/.
- 180.Khare R, Utidjian L, Razzaghi H, Soucek V, Burrows E, Eckrich D, Hoyt R, Weinstein H, Miller M, Soler D, Tucker J, Bailey C. Design and refinement of a data quality assessment workflow for a large pediatric research network. eGems (generating evidence & methods to improve patient outcomes. 2019;7(36). 10.5334/egems.294. [DOI] [PMC free article] [PubMed]
- 181.HAPI FHIR. 2023. https://hapi.fhir.org/.
- 182.Panozzo C, Welch E, Woodworth T, Huang T-Y, Her L, Gagne J, Sun J, Rogers C, Menzin T, Ehrmann M, Freitas K, Haug N, Toh S. Assessing the impact of the new icd-10-cm coding system on pharmacoepidemiologic studies-an application to the known association between angiotensin-converting enzyme inhibitors and angioedema. Pharmacoepidemiol Drug Saf. 2018;27. 10.1002/pds.4550. [DOI] [PubMed]
- 183.PCORnet Forum. 2023. https://github.com/CDMFORUM.
- 184.Rinner C, Gezgin D, Wendl C, Gall. W. A clinical data warehouse based on omop and i2b2 for Austrian health claims data. Stud Health Technol Inf. 2018;248:94–99. [PubMed]
- 185.Davis B, Swenson A. How carequality, the sequoia project, and ehealth exchange support the interoperable exchange of health data in the Usa. J Digit Imag. 2022;35. 10.1007/s10278-021-00538-y. [DOI] [PMC free article] [PubMed]
- 186.Yang Z, Jiang K, Lou M, Gong Y, Zhang L, Liu J, Bao X, Liu D, Yang P. Defining health data elements under the HL7 development framework for metadata management. J Biomed Semant. 2022;13. 10.1186/s13326-022-00265-5. [DOI] [PMC free article] [PubMed]
- 187.Mukhiya SK, Lamo Y. An HL7 FHIR and graphql approach for interoperability between heterogeneous electronic health record systems. Health Inf J. 2021;27(3):14604582211043920. 10.1177/14604582211043920. PMID: 34524029. [DOI] [PubMed]
- 188.Gagne J, Popovic J, Nguyen M, Sandhu S, Greene P, Izem R, Jiang W, Wang Z, Zhao Y, Petrone A, Wagner A, Dutcher S. Evaluation of switching patterns in fda’s sentinel system: a new tool to assess generic drugs. Drug Saf. 2018;41. 10.1007/s40264-018-0709-4. [DOI] [PubMed]
- 189.Atlas. 2023. https://atlas-demo.ohdsi.org/#/home.
- 190.Hripcsak G, Duke JD, Shah NH, Reich CG, Huser V, Schuemie MJ, Suchard MA, Park RW, Wong ICK, Rijnbeek PR, Lei J, Pratt N, Norén GN, Li Y-C, Stang PE, Madigan D, Ryang PB. Observational health data sciences and informatics (OHDSI): opportunities for observational researchers. IOS Press; 2015. [PMC free article] [PubMed]
- 191.OHDSI White Rabbit. 2023. http://ohdsi.github.io/WhiteRabbit/.
- 192.Stang P, Ryan P, Hartzema AG, Madigan D, Overhage JM, Welebob E, Reich CG, Scarnecchia T. Development and evaluation of infrastructure and analytic methods for systematic drug safety surveillance: lessons and resources from the observational medical outcomes partnership. Mann’s Pharmacovigilance: Third Edition; 2014. p. 453–61. 10.1002/9781118820186.CH29.
- 193.firely. 2023. https://fire.ly/products/forge/.
- 194.Arbor A. HL7 implementation guide for cda release 2: nhsn healthcare associated infection (hai) reports, release 2. Draft Standard For Trial Use, Health Level Seven, Inc. 2009, Feb.
- 195.Liu D, Sahu R, Ignatov V, Gottlieb D, Mandl K. High performance computing on flat FHIR files created with the new smart/HL7 bulk data access standard. AMIA Annu Symp Proc. 2020. [PMC free article] [PubMed]
- 196.Chai Y, Man KKC, Luo H, Torre CO, Wing YK, Hayes JF, Osborn DPJ, Chang WC, Lin X, Yin C, Chan EW, Lam ICH, Fortin S, Kern DM, Lee DY, Park RW, Jang J-W, Li J, Seager S, Lau WCY, Wong ICK. Incidence of mental health diagnoses during the COVID-19 pandemic: a multinational network study. Epidemiol And Psychiatric Sci. 2024;33. 10.1017/S2045796024000088. [DOI] [PMC free article] [PubMed]
- 197.Zhou X, Murugesan S, Bhullar H, Liu Q, Cai B, Wentworth C, Bate A. An evaluation of the thin database in the omop common data model for active drug safety surveillance. Drug Saf. 2013;36:119–34. 10.1007/S40264-012-0009-3/FIGURES/4. [DOI] [PubMed]
- 198.Trigiante L, Beneventano D. Privacy-preserving data integration for health: adhering to OMOP-CDM standard. 2024.
- 199.Toolpool Gesundheitsforschung. Forschungsplattformi (2025;2b2/tranSMART). https://www.toolpool-gesundheitsforschung.de/produkte/i2b2transmart.
- 200.Data Analysis and Real World Interrogation Network. 2025. https://www.ema.europa.eu/en/about-us/how-we-work/data-regulation-big-data-other-sources/real-world-evidence/data-analysis-real-world-interrogation-network-darwin-eu.
- 201.openEHR-Community wächst prominent weiter. 2025. https://www.highmed.org/de/news/openehr-community-waechst-prominent-weiter.
- 202.HEALTH-X dataLOFT; 2025. https://www.health-x.org/en/home.
- 203.HL7 v2.x Nachrichten – IT-Kommunikation im Krankenhaus. 2025. https://hl7.de/themen/hl7-v2-x-nachrichten/.
- 204.Finster M, Moinat M, Taghizadeh E. Etl: from the German health data lab data formats to the omop common data model. PLoS One. 2024. 10.1371/journal.pone.0311511. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
No datasets were generated or analysed during the current study.






































































