Abstract
Objectives:
To evaluate the compatibility of the Society of Critical Care Medicine’s (SCCM) Critical Care Data Dictionary (C2D2) with the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) and initiate a set of steps extending OMOP to accommodate specialized critical care data elements.
Design:
Systematic analysis and mapping study using a three-tiered semantic matching approach to demonstrate technical feasibility and identify fundamental challenges in critical care data standardization.
Setting:
Critical care medicine informatics research environment.
Subjects:
The SCCM’s C2D2 elements.
Interventions:
None.
Measurements and Main Results:
We evaluated the compatibility of C2D2 clinical variables with the OMOP CDM using a three-tier classification system (full match, partial match, and no match). Our analysis of 226 C2D2 elements revealed that 49.6% of concepts had full OMOP equivalents, 46.4% required modification, and 4.0% had no suitable mapping. Key incompatibilities were identified in ventilator parameters, composite scoring systems, and advanced organ support documentation. A large language model-based semantic matching system yielded a precision of 59.5%, recall of 87.0%, and F1 score of 70.7% at an optimized similarity threshold of 0.90. These findings highlight the need to harmonize data standardization approaches within the field of critical care, including how to handle concept stacking within single variables, age-specific criteria, and specialized constructs that were curated through the SCCM Delphi process, but reveal an OMOP mapping incompatibility or missing variables.
Conclusions:
Extending the OMOP CDM for critical care is technically feasible and requires targeted modifications to accommodate composite scores, temporal precision, and specialized critical care concepts as well as the resources needed to support this build. The community acutely faces crucial decisions about whether to pursue OMOP integration, adapt the C2D2 for version 2.0 compatibility, work toward OMOP vocabulary inclusion through Observational Health Data Sciences and Informatics processes, or collaborate with electronic health record vendors for native critical care standards support. These decisions require balancing technical feasibility with long-term sustainability and maintenance considerations.
Keywords: Common Data Model, critical care medicine, data mapping, data standards, Observational Medical Outcomes Partnership
KEY POINTS.
Question: How can the critical care community approach standardization of complex ICU data elements present in the Society of Critical Care Medicine (SCCM) Critical Care Data Dictionary (C2D2) and what are the technical and fundamental challenges in mapping specialized critical care concepts to a Common Data Model (e.g., Observational Medical Outcomes Partnership [OMOP] or Fast Healthcare Interoperability Resources)?
Findings: In this systematic mapping study of the SCCM’s C2D2, automated semantic matching achieved 97% precision and recall when recognizing clinically meaningful partial matches, successfully identifying suitable OMOP relationships for 96% of critical care concepts. Expert review revealed that approximately half of concepts had direct OMOP equivalents, while most remaining concepts required minor vocabulary extensions. The analysis revealed specific challenges including concept stacking requirements, age-specific qualifiers, and specialized constructs, while revealing gaps in C2D2’s current scope, particularly regarding detailed ventilator settings and organ support parameters, highlighting opportunities for C2D2 version 2.0 enhancement.
Meaning: The critical care community must make structured decisions about data standardization approaches sooner rather than later. While OMOP mapping is technically feasible, it requires community consensus about the acceptable level of modification to specialized concepts and whether the benefits of standardization outweigh the complexity of implementation.
Healthcare data standardization enables multicenter collaboration, comparative effectiveness research, and large-scale analytics. The Observational Health Data Sciences and Informatics (OHDSI) and Health level 7 International (HL7; Ann Arbor, MI) have partnered to integrate OHDSI’s Observational Medical Outcomes Partnership (OMOP) Common Data Model with HL7 Fast Healthcare Interoperability Resources (FHIR; HL7) (1). This OMOP-FHIR integration is particularly relevant for critical care as FHIR enables real-time data exchange while OMOP supports retrospective research analytics. OMOP has emerged as the leading research framework for transforming heterogeneous clinical data into a unified structure with standardized vocabularies (2–5). Specialized healthcare domains like critical care medicine present unique challenges for standardization due to the complexity, granularity, and temporal precision of ICU data (6, 7).
Database interoperability in healthcare requires systematic approaches to harmonize disparate data structures and terminologies across institutions through a process often called “crosswalking,” which involves mapping local concepts to standardized vocabularies while preserving clinical meaning. OMOP transforms clinical data using standardized vocabularies (Systematized Nomenclature of Medicine – Clinical Terms, Logical Observation Identifiers Names and Codes, Normalized Drug Naming System), enabling consistent analysis across healthcare systems (4, 8–10) This facilitates multicenter studies, federated research networks, and collaborative analytics without sharing raw patient-level data (9, 11, 12), supporting large-scale real-world evidence generation (13–20).
The Society of Critical Care Medicine (SCCM) developed a Critical Care Data Dictionary (C2D2) through an extensive Delphi consensus process involving critical care experts across multiple domains (21, 22). This systematic consensus approach resulted in specialized constructs and concepts that were curated specifically for critical care research, clinical practice, and quality improvement, reflecting the unique characteristics of ICU data, including age-specific criteria, composite measurements, and specialized interventions that may not align directly with general medical vocabularies (23).
Several parallel efforts address critical care data standardization challenges through different approaches (Table 1). The large-scale ICU Medical Information Mart for Intensive Care (MIMIC) database has been transformed to OMOP format (8), Amsterdam ICU database illustrates OMOP implementation for granular ICU data (24), and the Common Longitudinal ICU data Format (CLIF) offers federated learning approaches (25). This study provides technical evidence for mapping the SCCM C2D2 to OMOP vocabularies using automated semantic matching, quantifies compatibility patterns across clinical domains, and identifies specific structural modifications needed for critical care data standardization.
TABLE 1.
Comparison of Alternative Critical Care Data Standardization Approaches
| Approach | Primary Focus | Implementation Status | Strengths | Limitations | Data Sources |
|---|---|---|---|---|---|
| Medical Information Mart for Intensive Care-OMOP | Transform existing ICU database | Implemented | Large-scale, proven feasibility, historical depth | Limited to single-institution data structure | Beth Israel Deaconess |
| Amsterdam ICU-OMOP | Granular ICU data in OMOP | Implemented | Demonstrates ventilator/device mapping | Single institution focus | Amsterdam University Medical Centers |
| Common Longitudinal ICU data Format framework | Federated minimal elements | Development | Dynamic, adaptable, minimal data requirements | Limited concept coverage, early stage | Multi-institutional |
| Critical Care Data Dictionary-OMOP integration | Expert consensus critical care concepts | Analysis phase | Clinically optimized definitions, multi-domain consensus | Compatibility challenges with general vocabularies | Delphi consensus |
OMOP = Observational Medical Outcomes Partnership.
Analysis of parallel efforts in critical care data standardization, comparing implementation status, strengths, limitations, and data sources across Medical Information Mart for Intensive Care-OMOP, Amsterdam ICU-OMOP, Common Longitudinal ICU data Format framework, and Critical Care Data Dictionary-OMOP integration approaches.
METHODS
This study had two objectives: 1) evaluate automated semantic matching performance for clinical vocabulary mapping and 2) assess fundamental compatibility patterns between C2D2 and OMOP vocabularies. We developed a large language model-based mapping algorithm and used expert clinical informatics review to analyze structural compatibility challenges. Figure 1 illustrates the systematic methodology from C2D2 concept analysis through automated processing to expert validation.
Figure 1.
Study methodology flowchart for Critical Care Data Dictionary (C2D2)-Observational Medical Outcomes Partnership (OMOP) compatibility analysis. This flowchart illustrates the systematic approach used to evaluate compatibility between the Society of Critical Care Medicine’s C2D2 and the OMOP Common Data Model. The methodology employed a three-phase process: 1) Input data preparation using 226 C2D2 concepts at the concept level, 2) Automated processing through large language model-based semantic matching incorporating string matching, linguistic evaluation, and vector embeddings with threshold optimization at 0.90, and 3) Expert validation by clinical informaticists using dual independent review with consensus resolution and structured evaluation protocols. The process yielded classification results showing 112 full matches (49.6%), 105 partial matches (46.4%), and nine unmappable concepts (4.0%). Algorithm performance analysis demonstrated 59.5% precision and 97.2% precision under restrictive and permissive classification approaches, respectively, with an area under the receiver operating characteristic curve (AUROC) of 0.786. LLM = large language model.
Study Design and Scope
We analyzed SCCM’s C2D2 elements for OMOP Common Data Model compatibility and automated mapping feasibility. This work focused on technical feasibility rather than implementation guidance. Key technical terms are defined in Supplemental Table 1 (https://links.lww.com/CCM/H854) for clinical accessibility.
C2D2 Framework and Data Elements
The C2D2 employs a four-tier hierarchical classification system: domain, subdomain, concept, and common data element (with defined collection frameworks including temporality) (21, 22). This enables systematic concept-level evaluation while preserving clinical context.
Automated Semantic Matching Process
While the C2D2 employs a four-tier hierarchical classification for organizing data elements (domain, subdomain, concept, and common data element), our compatibility evaluation used a separate three-tier classification system to assess mapping quality. We employed a novel LLM-based system using three-tiered approaches: exact string matching, linguistic relationship evaluation, and semantic similarity calculation using vector embeddings (26). The semantic matching algorithm used a hierarchical approach: first identifying exact terminology matches, then evaluating linguistic relationships (synonyms, hierarchies), and finally calculating semantic similarity using mathematical vector representations of medical concepts. The 0.90 similarity threshold was selected post hoc as it provided optimal balance between precision (avoiding false matches) and recall (capturing true relationships), as detailed in Supplemental Table 2 (https://links.lww.com/CCM/H854). Two clinical informaticists independently validated all automated suggestions using structured protocols assessing semantic precision, clinical meaning preservation, contextual appropriateness, temporal fidelity, measurement specificity, and domain accuracy. Complete algorithmic specifications and threshold analysis across multiple similarity cutoffs are provided in the Supplemental Methods (https://links.lww.com/CCM/H854).
Classification and Evaluation Framework
C2D2 elements were classified using three tiers: “Full Match” (direct OMOP equivalents), “Partial Match” (requiring modification/extension), and “No Match” (no suitable OMOP equivalent). This illuminated compatibility spectrum between specialized critical care concepts and standardized vocabularies (21, 22).
Validation and Quality Assessment
Quality evaluation used precision, recall, F1 scores, and area under the receiver operating characteristic curve (AUROC). We analyzed 226 distinct C2D2 elements, with expert consensus resolution for disagreements and systematic documentation of failure modes. The original validation methodology was developed using the National Institutes of Health NIH Helping to End Addiction Long-term (HEAL) Initiative clinical trial Common Data Elements and National COVID Cohort Collaborative COVID-19 concepts (27–30).
Quantitative Analysis
Multiple similarity thresholds (0.85–0.95) were evaluated, with 0.90 selected post hoc for optimal precision-recall balance. Complete threshold analysis is in Supplemental Table 2 (https://links.lww.com/CCM/H854). Qualitative analysis examined clinical domains with systematic mapping challenges, assessing clinical significance of mapping deficiencies.
The LLM OMOP tool used in this study was initially developed under Wake Forest University School of Medicine Institutional Review Board (IRB00080548) approval on February 4, 2022. All procedures were followed in accordance with the ethical standards of the responsible committee and with the Helsinki Declaration of 1975. The analysis focused on data dictionary elements and standardized vocabularies without involving individual patient data or identifiable health information.
RESULTS
Automated Mapping Performance
Our analysis reveals that automated semantic matching can effectively bridge critical care vocabularies with standardized models: the algorithm identified appropriate OMOP relationships for 217 of 226 C2D2 elements (96.0%), with expert validation confirming 97.2% precision when recognizing clinically meaningful partial matches. This demonstrates practical feasibility for large-scale vocabulary harmonization in specialized medical domains.
Restrictive classification (only “Full Match” as true positives) yielded 59.5% precision, 87.0% recall, and 70.7% F1 score. However, most “Partial Match” categories represented high-quality semantic alignments requiring minor vocabulary extensions rather than mapping failures. Examples of partial matches requiring minor extensions include: 1) age-qualified severity scores where the C2D2 specifies “Pediatric Risk of Mortality (PRISM) III score (pediatric)” but OMOP contains only generic “PRISM score” without age specification, solvable through metadata annotations; 2) temporal qualifiers like “mean arterial pressure (MAP) at 24 hours” where OMOP captures MAP measurements but requires additional timestamp precision fields; and 3) composite assessments combining multiple OMOP concepts through defined calculation rules rather than stored values. The algorithm identified appropriate OMOP concepts for 217 of 226 C2D2 elements (96.0%), with only nine concepts (4.0%) having no algorithmic suggestion. AUROC was 0.786, indicating strong discriminatory power and comparing favorably with previous medical vocabulary mapping efforts achieving 60–80% precision (31).
Overall Mapping Distribution
Expert review revealed significant domain variations (Supplemental Table 3, https://links.lww.com/CCM/H854). Of 217 OMOP-assignable concepts, 112 (49.6%) achieved full matches, 105 (46.4%) required modification as partial matches, and 9 (4.0%) remained unmappable due to specialized ICU constructs.
Domain-specific patterns reflected fundamental differences between standardized representations and ICU-specific assessments. Procedure domain showed highest compatibility (85.7% full matches), followed by measurement domain (61.9%), condition domain (50.0%), and observation domain (46.8% of 79 elements). Detailed concept-by-concept mapping results with similarity scores and expert rationales are provided in Supplemental Tables 3–7 (https://links.lww.com/CCM/H854), enabling replication and extension of this analysis.
Unmappable Concepts
Nine completely unmappable concepts illuminated fundamental incompatibilities: temporal-specific assessments (“Last Code Status at 24 hr in ICU”), severity-qualified PRISM components (“PRISM WBC low,” “PRISM Platelets–low,” and “PRISM Glasgow Coma Scale–lowest”), and directional physiologic qualifiers (“MAP–highest,” “MAP–low,” “Pao2/Fio2 ratio–Low,” and “oxygen saturation/Fio2 ratio–Low”).
Systematic Challenge Patterns
Analysis revealed systematic incompatibility patterns: concept stacking requirements where multiple OMOP concepts needed combination; age qualifier omissions where pediatric/geriatric criteria could not be preserved; missing medical qualifiers in OMOP vocabularies; temporal precision constraints for specific time windows; severity qualifiers (“mild,” “moderate,” and “severe”) lost in broader OMOP concepts; and specialized ICU assessments lacking appropriate OMOP representations.
Critical Care-Specific Limitations
Ventilator parameters and severity scoring posed substantial challenges, representing 23 of 105 partial matches (21.9%). Temporal precision constraints affected 18 concepts (17.1% of modifications needed), primarily involving time-qualified assessments requiring specific windows (e.g., “first 24 hr”). Age-specific qualifiers affected nine concepts (8.6%), predominantly pediatric severity scoring components. OMOP demonstrates insufficient granularity for critical care contextual metadata. Severity qualifiers embedded within concept definitions are not preserved through standard mapping. Temporal precision requirements for severity scoring exceeded standard OMOP capabilities, reflecting tension between broad interoperability and specialized documentation needs.
These performance metrics compare favorably with previous medical vocabulary mapping efforts achieving 60–80% precision, while the substantial difference between restrictive (59.5%) and permissive (97.2%) precision approaches reveals that most “partial matches” represent high-quality semantic alignments requiring policy decisions about vocabulary extensions rather than algorithmic improvements. This suggests that automated semantic matching provides practical value for vocabulary standardization initiatives by enabling focused expert review on truly challenging concepts.
DISCUSSION
This analysis demonstrates three critical insights for critical care leadership: 1) Automated tools can successfully identify vocabulary relationships for 96% of expert-derived critical care concepts, reducing manual review burden from hundreds to dozens of concepts; 2) Approximately half of specialized critical care concepts can directly integrate with existing healthcare data standards, while most remaining concepts require modification rather than complete reconstruction; and 3) The community faces strategic decisions with distinct trade-offs between clinical precision, research interoperability, and resource requirements. These findings provide technical foundation for informed strategic decisions about critical care data standardization approaches.
OMOP mapping for C2D2 is technically feasible but requires acknowledging two distinct types of compatibility. Nearly half (49.6%) of C2D2 concepts can map directly to existing OMOP vocabularies without modifications, these are immediately usable. An additional 46.4% require vocabulary extensions or modifications that are technically straightforward but require community consensus and OHDSI coordination. Only 4.0% present fundamental incompatibilities requiring new approaches. Thus, the answer to “is OMOP feasible?” depends on the institution’s priorities: organizations seeking immediate implementation can begin with the 50% of direct matches while developing extension strategies; those requiring comprehensive coverage must commit to the multiyear process of vocabulary development and community coordination. For informaticians considering OMOP implementation: direct integration is viable for half of critical care concepts, while the remaining concepts require vocabulary extensions coordinated through OHDSI processes. Institutions can implement immediate solutions for full matches while developing extension strategies for partial matches or await broader community consensus on standardization approaches. Success requires both OMOP technical expertise and understanding of critical care-specific data structures and clinical workflows.
Critical care faces unique data science challenges: data complexity and scale, rapid disease evolution, and profound clinical consequences (7, 21, 22). Our analysis reveals three key findings: automated semantic matching achieves practical vocabulary cross-walking performance; approximately half of expert-derived critical care concepts map directly to existing standards while remainder require modification; and the community faces fundamental strategic choices balancing specialized precision with interoperability benefits.
Context Within Existing Efforts
This analysis complements ongoing initiatives while addressing distinct challenges. MIMIC-IV OMOP demonstrates feasibility for large ICU datasets (8), Amsterdam ICU illustrates granular ventilator extension (24), and CLIF offers federated approaches (25). Our C2D2-OMOP analysis evaluates purpose-built critical care concepts from expert consensus rather than pragmatic electronic health record (EHR) extractions, creating different compatibility challenges with general vocabularies.
Technical Feasibility and Fundamental Limitations
Modern semantic matching effectively bridges specialized ICU vocabularies with standardized models for substantial critical care concept portions. However, findings reveal significant structural limitations in applying general vocabularies to specialized requirements.
Complex composite scores (Acute Physiology and Chronic Health Evaluation II, Sequential Organ Failure Assessment, and PRISM III) present fundamental challenges because OMOP does not natively store precomputed severity assessments requiring dynamic combinations with temporal specifications. This creates competing perspectives: reconstructing scores through standardized OMOP tables using open-source packages vs. direct storage to avoid inconsistency risks. This tension requires community consensus about acceptable trade-offs between standardization benefits and specialized accuracy.
Temporal precision constraints represent substantial limitations. Critical care assessments frequently require specific time windows conflicting with OMOP’s measurement_datetime structure. OMOP lacks explicit preadmission support and does not group measurements by admission phases.
Sustainability requires considering long-term maintenance resources. OMOP integration necessitates ongoing OHDSI vocabulary coordination, requiring sustained funding and personnel. These ongoing costs must be weighed against standardization benefits.
Algorithmic Performance Value
The LLM system’s 97.2% precision with permissive classification demonstrates significant practical value, reducing expert review burden from 226 concepts to nine unmappable concepts. The substantial difference between restrictive (68.3%) and permissive (97.2%) precision highlights that most “partial matches” represent clinically meaningful alignments requiring policy decisions rather than algorithmic improvements.
Strategic Options
The critical care community faces several pathways (detailed in Supplemental Table 4, https://links.lww.com/CCM/H854): pursuing OMOP integration for established analytical access but requiring community effort; developing enhanced C2D2 v2.0 (Society of Critical Care Medicine, Mount Prospect, IL) for compatibility while preserving clinical meaning; maintaining specialized approach preserving Delphi consensus utility; or implementing hybrid strategies with tiered approaches. Selection between these pathways should consider three primary factors: 1) Timeline urgency as OMOP integration requires 3–5 years while specialized approaches can be optimized immediately; 2) Resource availability as OMOP integration demands sustained multi-institutional coordination while independent development requires dedicated informatics teams; and 3) Interoperability priority as institutions prioritizing federated research networks may favor standardization approaches, while those emphasizing clinical precision may prefer specialized solutions. The hybrid approach offers immediate benefits from full matches while preserving development options for specialized concepts.
Beyond standardization, successful implementation requires EHR vendor collaboration. Partnerships with vendors like Epic (Verona, WI) could facilitate native critical care standards support, offering more sustainable solutions than post hoc vocabulary mapping.
Framework for Development
The proposed extension framework (Supplemental Tables 5 and 6, https://links.lww.com/CCM/H854) addresses systematic challenges while preserving expert consensus precision through modular extensions maintaining OMOP compatibility, including structural extensions for temporal precision, vocabulary supplements for specialized concepts, and dedicated modules for high-complexity domains.
Strategic Implications
These reflect broader questions facing medical subspecialties: clinical precision through specialized vocabularies, research opportunities through standardized formats, and resource investments for long-term maintenance. These are strategic choices about community priorities and the balance between specialized utility and broader interoperability.
LIMITATIONS AND FUTURE DIRECTIONS
This focused on technical feasibility rather than implementation guidance. Future work should emphasize community engagement and empirical validation: 1) pilot implementations across 2–3 institutions within 12–18 months, 2) real-world validation using existing OMOP ICU databases, and 3) structured stakeholder input about standardization priorities through surveys and consensus conferences.
CONCLUSIONS
The critical care community possesses a unique opportunity to demonstrate specialized domain standardization while preserving clinical utility essential for advancing patient care. Technical feasibility provides foundation for informed community decisions, but ultimate direction must reflect community priorities regarding specialized functionality vs. broader interoperability benefits.
The critical care community possesses a unique opportunity to demonstrate how specialized medical domains can achieve data standardization while preserving clinical utility essential for advancing patient care. Our technical analysis provides evidence that approximately half of expert-derived critical care concepts integrate directly with existing standards, while most remaining concepts require targeted extensions rather than fundamental reconstruction. Automated semantic matching tools serve as force multipliers for vocabulary standardization, enabling systematic approaches that could accelerate standardization timelines across medical subspecialties. However, technical feasibility provides only the foundation; strategic direction must reflect community priorities about the balance between specialized clinical precision and broader interoperability benefits, with implementation success depending on sustained commitment to whichever pathway the community selects.
Supplementary Material
Footnotes
The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Drs. Adams, Hurley, Khanna, and Bartels were involved in conceptualization. Mr. Perkins, Mr. Hudson, Dr. Topaloglu, and Dr. Adams were involved in methodology and software. Drs. Adams, Hurley, and Khanna were involved in validation and investigation. Drs. Adams, Hurley, and Khanna were involved in data curation. Dr. Adams prepared the writing—original draft. Dr. Adams was involved in supervision and project administration. All authors contributed substantially to review and editing.
Supplemental digital content is available for this article. Direct URL citations appear in the printed text and are provided in the HTML and PDF versions of this article on the journal’s website (http://journals.lww.com/ccmjournal).
This work was supported by the National Institute on Drug Abuse through the National Institutes of Health (NIH) Helping to End Addiction Long-term Initiative of the NIH under award number R24DA055306, U24DA057612, R25DA061740, and R24DA058606 (to Drs. Adams and Hurley). Wake Forest University School of Medicine provided computational resources and institutional support.
Dr. Adams’ institution received funding from the National Institute on Drug Abuse. Drs. Adams, Hurley, Perkins, Topaloglu, and Stocking received support for article research from the National Institutes of Health (NIH). Drs. Bartels’ and Perkins’ institutions received funding from the NIH. Dr. Hurley received funding from Nevro (research funds to the institution, topic-painful diabetic neuropathy and not related to this work) and State Farm (expert/consulting not related to this work). Dr. Bartels’ institution received funding from the Agency for Healthcare Research and Quality. Dr. Topaloglu disclosed government work. Dr. Cobb received funding from Akido Labs and Bauhealth; he is a past Workgroup Co-Chair of the Society of Critical Care Medicine (SCCM) Discovery Data Science Campaign. Dr. Reuter-Rice’s institution received funding from the National Institute of Neurologic Disorders and Stroke and Elsevier; she disclosed that she is co-chair of the SCCM Data Science Campaign. Dr. Stocking’s institution received funding from the National Heart, Lung, and Blood Institute under award number K01HL168222. Dr. Khanna received funding from Medtronic, Edwards Lifesciences, GE Healthcare, Philips Research North America, Sentinel Medical, Bayer Corporation, AOP, Innoviva Therapeutics, and Pharmazz; he disclosed that he is a past chair of the SCCM Discovery Network was involved in the Delphi process for the development of the Critical Care Data Dictionary; and he is currently a member of the SCCM council. Mr. Hudson has disclosed that he does not have any potential conflicts of interest.
The complete Society of Critical Care Medicine Critical Care Data Dictionary used in this analysis is available in the Supplemental Materials (https://links.lww.com/CCM/H854). Individual mapping results and similarity scores are provided in Supplemental Table S7A (https://links.lww.com/CCM/H854) and with discussion of these findings in Supplement 7B (https://links.lww.com/CCM/H854).
The Large Language Model Observational Medical Outcomes Partnership tool used in this study was approved by the Wake Forest University School of Medicine Institutional Review board (IRB00080548) on February 4, 2022.
Contributor Information
Robert W. Hurley, Email: rwhurley2010@gmail.com.
Karsten Bartels, Email: kbartels@med.umich.edu.
Matthew L. Perkins, Email: Matthew.Perkins@Advocatehealth.org.
Cody Hudson, Email: cody.hudson@advocatehealth.org.
Umit Topaloglu, Email: umtopaloglu@gmail.com.
J. Perren Cobb, Email: jpcobb@med.usc.edu.
Karin Reuter-Rice, Email: karin.reuter-rice@duke.edu.
Ashish K. Khanna, Email: ashish.khanna@advocatehealth.org.
REFERENCES
- 1.Observational Health Data Sciences and Informatics: HL7 International and OHDSI Announce Collaboration to Provide Single Common Data Model for Sharing Information in Clinical Care and Observational Research. Available at: https://www.ohdsi.org/ohdsi-hl7-collaboration/. Accessed June 20, 2025
- 2.Carus J, Trübe L, Szczepanski P, et al. : Mapping the oncological basis dataset to the standardized vocabularies of a common data model: A feasibility study. Cancers (Basel) 2023; 15:4059. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Henke E, Zoch M, Peng Y, et al. : Conceptual design of a generic data harmonization process for OMOP common data model. BMC Med Inform Decis Mak 2024; 24:58. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Reich C, Ostropolets A, Ryan P, et al. : OHDSI Standardized Vocabularies—a large-scale centralized reference ontology for international data harmonization. J Am Med Inform Assoc 2024; 31:583–590 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Sathappan SMK, Jeon YS, Dang TK, et al. : Transformation of electronic health records and questionnaire data to OMOP CDM: A feasibility study using SG_T2DM dataset. Appl Clin Inform 2021; 12:757–767 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Kent S, Burn E, Dawoud D, et al. : Common problems, common data model solutions: Evidence generation for health technology assessment. PharmacoEcon 2021; 39:275–285 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Heavner SF, Kumar VK, Anderson W, et al. ; Society of Critical Care Medicine (SCCM) Discovery Panel on Data Sharing and Harmonization: Critical data for critical care: A primer on leveraging electronic health record data for research from Society of Critical Care Medicine’s panel on data sharing and harmonization. Crit Care Explor 2024; 6:e1179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Paris N, Lamer A, Parrot A: Transformation and evaluation of the MIMIC database in the OMOP common data model: Development and usability study. JMIR Med Inform 2021; 9:e30970. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lynch KE, Deppen SA, DuVall SL, et al. : Incrementally transforming electronic medical records into the observational medical outcomes partnership common data model: A multidimensional quality assurance approach. Appl Clin Inform 2019; 10:794–803 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Ward R, Hallinan CM, Ormiston-Smith D, et al. : The OMOP common data model in Australian primary care data: Building a quality research ready harmonised dataset. PLoS One 2024; 19:e0301557. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Papez V, Moinat M, Voss EA, et al. : Transforming and evaluating the UK Biobank to the OMOP common data model for COVID-19 research and beyond. J Am Med Inform Assoc 2022; 30:103–111 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Haberson A, Rinner C, Schöberl A, et al. : Feasibility of mapping Austrian health claims data to the OMOP common data model. J Med Syst 2019; 43:314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Cho S, Mohan S, Husain SA, et al. : Expanding transplant outcomes research opportunities through the use of a common data model. Am J Transplant 2018; 18:1321–1327 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Lee GH, Park J, Kim J, et al. : Feasibility study of federated learning on the distributed research network of OMOP common data model. Healthc Inform Res 2023; 29:168–173 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Overhage JM, Ryan PB, Reich CG, et al. : Validation of a common data model for active safety surveillance research. J Am Med Inform Assoc 2012; 19:54–60 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Voss EA, Makadia R, Matcho A, et al. : Feasibility and utility of applications of the common data model to multiple, disparate observational health databases. J Am Med Inform Assoc 2015; 22:553–564 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Chakrabarti S, Sen A, Huser V, et al. : An interoperable similarity-based cohort identification method using the OMOP common data model version 5.0. J Healthc Inform Res 2017; 1:1–18 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Zhou X, Murugesan S, Bhullar H, et al. : An evaluation of the THIN database in the OMOP common data model for active drug safety surveillance. Drug Saf 2013; 36:119–134 [DOI] [PubMed] [Google Scholar]
- 19.Biedermann P, Ong R, Davydov A, et al. : Standardizing registry data to the OMOP common data model: Experience from three pulmonary hypertension databases. BMC Med Res Methodol 2021; 21:238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Sarrat-González D, Escribà-Montagut X, Houghtaling J, et al. : dsOMOP: Bridging OMOP CDM and DataSHIELD for secure federated analysis of standardized clinical data. Bioinformatics 2025; 41:btaf286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Armaignac DL, Heavner SF, Rausen M, et al. : Guiding principles for data sharing and harmonization: Results of a systematic review and modified Delphi from the Society of Critical Care Medicine data science campaign. Crit Care Med 2025; 53:e619–e631 [DOI] [PubMed] [Google Scholar]
- 22.Murphy DJ, Anderson W, Heavner SH, et al. : Development of a core critical care data dictionary with common data elements to characterize critical illness and injuries using a modified Delphi method. Crit Care Med 2025; 53:e1045–e1054 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Lee YH, Choe YJ, Yoon YS, et al. : Predicting ICU admission risk in children with respiratory syncytial virus. Infect Dis Ther 2025; 14:1277–1286 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Puttmann D, de Groot R, de Keizer N, et al. ; Dutch ICU Data Sharing Against COVID-19 Collaborators: Assessing the FAIRness of databases on the EHDEN portal: A case study on two Dutch ICU databases. Int J Med Inform 2023; 176:105104. [DOI] [PubMed] [Google Scholar]
- 25.Rojas JC, Lyons PG, Chhikara K, et al. : A common longitudinal intensive care unit data format (CLIF) for critical illness research. Intensive Care Med 2025; 51:556–569 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Adams MCB, Perkins ML, Hudson C, et al. : Breaking digital health barriers: Development and validation of an LLM-based tool for automated OMOP mapping. J Med Internet Res 2025; 27:e69004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Adams MCB, Sward KA, Perkins ML, et al. : Standardizing research methods for opioid dose comparison: The NIH HEAL morphine milligram equivalent calculator. Pain 2025; 166:1729–1737 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Adams MCB, Hassett AL, Clauw DJ, et al. : The NIH HEAL pain common data elements (CDE): A great start but a long way to the finish line. Pain Med 2025; 26:146–155 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Adams MCB, Brummett CM, Wandner LD, et al. : Michigan body map: Connecting the NIH HEAL IMPOWR network to the HEAL ecosystem. Pain Med 2023; 24:907–909 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Adams MCB, Hurley RW, Siddons A, et al. : NIH HEAL clinical data elements (CDE) implementation: NIH HEAL Initiative IMPOWR network IDEA-CC. Pain Med 2023; 24:743–749 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Sun JY, Sun Y: A system for automated lexical mapping. J Am Med Inform Assoc 2006; 13:334–343 [DOI] [PMC free article] [PubMed] [Google Scholar]

