Abstract
Objective
To develop and validate a weighting framework to improve representativeness of Oracle Health Real-World Data (OHRWD).
Materials and Methods
We conducted cross-sectional analyses of OHRWD encounters in 2019 and 2022. The primary method (M1) applied design weights based on American Hospital Association (AHA) hospital encounter counts to balance OHRWD encounters across strata of US region, hospital system, bed size, and encounter type (inpatient, emergency department, ambulatory surgery). Comparative methods (M2-M5) used unified structural weights, encounter-type multipliers, iterative proportional fitting, and demographic post-stratification. Weighted OHRWD estimates were validated against the Healthcare Cost and Utilization Project (HCUP) National Inpatient Sample (NIS), Nationwide Emergency Department Sample (NEDS), and Nationwide Ambulatory Surgery Sample (NASS) for demographics, conditions (low back pain, opioid use disorder (OUD), hypertension, diabetes), and procedures (colonoscopy, appendectomy). Equivalence was assessed using two one-sided tests with ±20% margins; standardized differences were summarized using Cohen’s d and h.
Results
M1 produced close alignment between OHRWD and HCUP for age, sex, region, and most race/ethnicity groups, with negligible effect sizes. Low back pain, OUD, and appendectomy prevalences closely matched HCUP benchmarks, whereas hypertension and, in some settings, diabetes and Hispanic ethnicity showed larger deviations. Alternative methods (M2-M5) yielded mixed performance and did not consistently outperform M1.
Discussion
Encounter-type-specific, structurally informed weighting improved alignment with HCUP national benchmarks, highlighting domains, such as hypertension and race/ethnicity estimates, that require additional calibration.
Conclusion
AHA-based, encounter-type-stratified weighting enables OHRWD to better approximate encounter patterns and could support epidemiologic and health services research using large EHR datasets.
Keywords: electronic health records, real-world data, weighting methods, representativeness, selection bias
Introduction
Electronic health record (EHR) databases have become indispensable for clinical and epidemiologic research because they offer large-scale, granular clinical detail. Yet, unlike surveys or administrative claims datasets, EHR-derived data are typically convenience samples drawn from participating health systems rather than probability-based samples of a defined population.1,2 Consequently, EHR databases often lack population representativeness. Patient demographics and clinical characteristics in the EHR can diverge from those of the broader population.2 Because the probability of any given person’s inclusion in an EHR database is unknown, traditional survey-weighting approaches that rely on known selection probabilities cannot be directly applied.2 Without correction, analyses of unweighted EHR data can yield biased estimates that do not generalize to the target population, potentially misrepresenting population-level disease burden, associations, or healthcare outcomes.3
This challenge is especially salient when EHR data are used for population inference across care settings. Probability-based administrative datasets such as the Healthcare Cost and Utilization Project (HCUP) National Inpatient Sample (NIS) are constructed from a defined sampling frame and provide sampling weights to support national inference, whereas EHR networks aggregate records from contributing organizations without a probabilistic design.2 In addition, EHR capture is inherently conditional on healthcare contact within contributing systems; individuals receiving care elsewhere, facing barriers to access, or interacting less frequently with care may be underrepresented.3 Moreover, EHR data are collected as an artifact of care rather than designed for population inference, and “informed presence” can overrepresent individuals with higher encounter frequency or greater comorbidity burden.2–4 Thus, while EHRs are information-rich, their completeness and representativeness are not guaranteed, and methods to mitigate selection and capture bias are needed to support valid generalization.
Recognizing these limitations, prior work has developed weighting strategies to improve EHR representativeness by aligning EHR samples to external benchmarks. One common approach is post-stratification/raking, in which weights are iteratively adjusted so that the weighted EHR sample matches population margins (eg, age, sex, race/ethnicity, geography) from sources such as the US Census or American Community Survey.2,5 This strategy has been applied to large EHR-linked cohorts (eg, All of Us Biobank) to reduce discrepancies with national health statistics,5 and to EHR-based surveillance networks (eg, Multi-State EHR-Based Network for Disease Surveillance [MENDS]) where post-stratification improved alignment of hypertension estimates with survey-based benchmarks.2 Another approach is inverse probability of inclusion weighting, which uses external individual-level data to model selection into an EHR-linked cohort (or biobank) as a function of demographic and health characteristics, then reweights the cohort to better approximate the target population.1 These methods collectively reflect a growing consensus that weighting can improve the transportability of EHR-based findings when selection processes are not ignorable.1–3 However, given the unique nature of large, multi-system EHR networks, spread over distinct encounter types, to our knowledge no prior study has incorporated the specific encounter type explicitly into the weighting design and validated across distinct, encounter-specific, national benchmarks. Such efforts are critical given the rapidly increasing number of health conditions of variable prevalence, presenting to healthcare systems over numerous different settings.
The need for representativeness is particularly acute in research on opioid use disorder (OUD), pain management, and substance use, where disparities vary substantially across demographic and geographic groups. Recent evidence indicates that opioid overdose mortality has increased and shifted in historically marginalized communities, with Black Americans experiencing disproportionately high overdose mortality in many areas.6 At the same time, Black patients remain less likely to receive OUD treatment compared with White patients, despite increasing burden.7 Disparities in pain care have also been documented, including undertreatment of pain among Black and Hispanic patients and the downstream consequences of biased clinical decision-making.7 When EHR samples underrepresent populations with lower access to care or systematically different treatment exposure, unweighted analyses can yield misleading conclusions and may even risk propagating inequities.3 Weighting approaches that improve representativeness can therefore strengthen the validity of EHR-based real-world evidence in these high-stakes domains.
Despite important advances, a scalable framework for weighting a large, multifaceted national EHR database, while considering the explicit encounter type and also validating performance against established external standards that are setting-specific, remains needed. In this study, we address this gap by developing and validating a multivariate, stratified weighting approach for the Oracle Health Real-World Data (OHRWD) EHR dataset. Leveraging external benchmarks (American Hospital Association [AHA] data for weighting and HCUP datasets for validation), we evaluate whether careful weighting can improve alignment between OHRWD encounter distributions and national reference distributions across specific settings. By doing so, we aim to enhance the utility of OHRWD for epidemiologic and outcomes research and improve confidence that findings derived from this large EHR resource can support generalizable inference.2,3
Methods
Data sources
We used OHRWD for weighting and HCUP NIS, Nationwide Emergency Department Sample (NEDS), and Nationwide Ambulatory Surgery Sample (NASS) for validation. OHRWD is a large collection of de-identified EHR from all 50 states. As of February 2025, this dataset included information from 149 US health systems, covering approximately 115 million patients and 2 billion healthcare encounters. OHRWD is extracted from the electronic medical records of hospitals in which Oracle has a data use agreement. Encounters may include pharmacy, clinical and microbiology laboratory, admission, and billing information from affiliated patient care locations. All admissions, medication orders and dispensing, laboratory orders, and specimens are date and time stamped, providing a temporal relationship between treatment patterns and clinical information. Oracle has established Health Insurance Portability and Accountability Act-compliant operating policies to establish de-identification for OHRWD.8
HCUP comprises a family of databases and tools developed for the HCUP, sponsored by the Agency for Healthcare Research and Quality (AHRQ). The HCUP NIS/NEDS are the largest publicly available all-payer inpatient healthcare/emergency department (respectively) databases in the US and the HCUP NASS is the only all-payer ambulatory surgery database in the US. Beginning with 2012 data, the NIS includes discharge-level data from a stratified sample of approximately 20% of all discharges from community hospitals across 47 states. Unweighted, the dataset contains approximately 7 million hospital stays annually; when weighted, it represents an estimated 35 million hospitalizations nationwide.9 Unweighted, NEDS contains data from about 32 million ED visits in 2022. Weighted, it estimates roughly 137 million ED visits.10 Unweighted, the 2022 NASS contains approximately 9.1 million ambulatory surgery encounters and approximately 12.1 million ambulatory surgery procedures. Weighted, it estimates approximately 12.4 million ambulatory surgery encounters and 16.4 million ambulatory surgery procedures.11 A validated weighting method is used to ensure national representativeness across all databases.
Sample and study design
This study employed a serial, cross-sectional, weighted design using data at the encounter level. Analyses were conducted independently for the calendar years 2019 and 2022 to assess for changes over time. The year 2022 was chosen, as this was the most recent year available for HCUP data at the time of analysis, which was required for validation. The year 2019 was chosen as it comprised the largest capture of data prior to COVID-19 to serve as an optimal temporal, pre-pandemic, comparison. Encounters were included if they occurred in each year of study, had an available patient age at encounter, and had available health system characteristic information. Weighting scales encounters so that weighted totals approximate a target universe.12 We applied a multivariable, stratified weighting approach to OHRWD using AHA encounter benchmarks.
Primary method (M1)
M1 followed the HCUP NIS design-weight framework13, modified to estimate separate weights by encounter type. We balanced OHRWD encounters to AHA encounter counts within strata defined by US 1-digit ZIP region, single hospital or larger health-system binary indicator, and bed size, separately for inpatient, emergency, and outpatient encounters. The stratum- and type-specific weight was: Ws,t = [Ns,t(universe)/OHRWDs,t(sample)] × (4/Qi), where Ns,t (universe) is the number of encounters from universe healthcare systems within stratum s and encounter type t; OHRWDs,t (sample) is the number of encounters from sample healthcare systems selected for the OHRWDs; and Qi is the number of quarters of discharge data contributed by healthcare system i to OHRWD (usually Qi = 4).13 When Qi was missing, we set Qi to the number of non-missing quarters. To limit instability from sparse strata, we collapsed strata when needed and, when full-stratum encounter counts were low (eg, below the first quartile), imputed full-stratum weights from reduced-stratum weights. For example, weights produced while collapsing bed size group counts, categorizing one-digit ZIP regions into larger regions, and removing the hospital/health-system indicator from included strata categories, were used to impute weights from low count strata. Final weights were winsorized at the 1st and 99th percentiles. Summary statistics for weights, stratified by encounter type, are listed in Table S1.
Comparative methods (M2-M5)
We evaluated 4 alternative approaches, including M2, a unified structural weighting scheme that ignored encounter type; M3, M2 plus encounter-type multipliers to align encounter-type composition within strata; M4, iterative proportional fitting (IPF; “raking”) to align weighted OHRWD marginal totals with population totals across the structural strata14,15; and M5, M1 followed by a second-stage demographic post-stratification for age, gender, and race/ethnicity using 2020 US Census benchmarks. Figure 1 summarizes M1-M5 and Appendix A, in the supplement, illustrates these methods in further detail.
Figure 1.

Methods overview for weighting approaches (M1-M5) applied to OHRWD using AHA encounter benchmarks, prior to validation against HCUP (NIS/NEDS/NASS).
Validation analysis
Following the implementation of the weighting scheme, we conducted validation analyses by comparing weighted estimates from the OHRWD database, under different encounter type scenarios, to those from the HCUP NIS, NEDS, and NASS, treating HCUP as our gold standard. To approximate the population captured by the HCUP, we excluded OHRWD encounters with lengths of stay exceeding 1 year, as well as those discharged to non-acute settings, including long-term care hospitals, rehabilitation centers, and extended care facilities. For comparability to NIS, we restricted the OHRWD sample to inpatient, acute-care, encounters. For comparability to NEDS, we restricted the sample to emergency department, acute-care, encounters. For comparability to NASS, we restricted the sample to outpatient encounters with ambulatory, invasive and therapeutic, procedures (eg, appendectomy, cholecystectomy, cataract surgery; associated codes identified in Table S2) that were discharged home (ie, removal of encounters discharged to skilled nursing facility, extended care facility, hospital, etc.). This validation comparison was used to assess the performance of the weighting scheme across key demographic characteristics (age, sex, and race/ethnicity, census region), clinical conditions (low back pain, OUD, hypertension, and diabetes mellitus [DM]; associated codes in Table S3), and relevant procedures (physical therapy [PT], colonoscopy, and appendectomy; associated codes in Table S4).
Appropriate International Classification of Diseases, 10th revision, Clinical Modification (ICD-10-CM) and Current Procedural Terminology codes were searched across any position (ie, primary and non-primary diagnoses/procedures). Demographic means/percentages and clinical characteristic prevalences were calculated in HCUP and OHRWD. Formal comparisons of estimates were made between the 2 databases. Continuous estimates were compared with 2 one-sided tests (TOST) for equivalence of means via t-tests and binary estimates were compared with TOST for equivalence of proportions via Fisher’s exact test.16 Equivalence margins were set at ±20% of the referent group. Due to very large samples influencing statistical significance with very little difference, measures of effect size were calculated utilizing Cohen’s d for continuous estimates and Cohen’s h for binary estimates. Absolute and relative differences were also calculated. M1 was the primary weighting method used to compare OHRWD estimates to HCUP, but we additionally compared with M2-M5. All hypothesis tests used a 5% significance level; effect sizes of 0.2, 0.5, and 0.8 were interpreted as small, medium, and large, respectively. All analyses were conducted using R (version 4.0.2; R Foundation for Statistical Computing).
A final, supplemental, validation analysis assessed the performance of M1 weighting in an inferential association between history of OUD and likelihood of PT utilization among those with a new episode of low back pain in an inpatient setting. The analytical year was restricted to 2022. Previous studies have assessed this relationship among patients within a single state with commercial, Medicaid, or Medicare coverage and among a nationally representative Medicare beneficiary cohort.17,18 However, this has yet to be assessed among patients of all beneficiary types, across the entire US, and strictly in an inpatient setting. Propensity score matching accounted for possible confounding and mixed-effects logistic regression calculated the adjusted odds ratio (aOR) with variability of the estimate captured by a 95% CI. The analysis was conducted on the raw, unweighted data and again while incorporating the M1 weights. Appendix B, in the supplementary, gives the full methodological details with accompanying codes and diagnostics listed in Tables S2-S17 and Figure S1.
Results
Demographics
Encounter-level demographics are provided in Tables 1 and 2 with comparisons between HCUP (NIS, NEDS, NASS) and OHRWD M1 for 2019 and 2022. Implementing the M1 weights reduced the mean percentage difference between HCUP and OHRWD percentages by 35.31% in 2019 and 32.25% in 2022 for inpatient encounters (3.88%-2.51% in 2019; 3.38%-2.29% in 2022; Table S18a), 15.52% in 2019 and 10.17% in 2022 for emergency encounters (4.06%-3.43% in 2019; 3.54%-3.18% in 2022; Table S18b), and 30.16% in 2019 and 39.57% in 2022 for outpatient surgery encounters (4.31%-3.01% in 2019; 3.74%-2.26% in 2022; Table S18c). Close alignment was observed in age and sex between OHRWD and all HCUP databases, with significant equivalence and negligible effect sizes observed across all years. By race/ethnicity, OHRWD M1 produced percentages of individuals identifying as White that closely aligned with the NIS, while alignment with NEDS and NASS exhibited slightly more variability.
Table 1.
Encounter-level demographics comparison between HCUPa NISb/NEDSc and weighted (M1) OHRWDd for 2019 and 2022.
| Year | Inpatient |
Emergency |
|||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Weighted estimates |
Database comparison |
Weighted estimates |
Database comparison |
||||||||||
| HCUP NIS n (%)e | OHRWD n (%)e | Zf | hg | Abs diff (%)h | Rel diff (%)i | HCUP NEDS n (%)e | OHRWD n (%)e | Zf | hg | Abs diff (%)h | Rel diff (%)i | ||
| Age j | |||||||||||||
| 2019 | 50.15 | 51.04 | 1385.72k* | 0.03l | 0.90m | 1.79 | 41.77 | 38.91 | 1803.80k* | 0.12l | 2.86m | 6.84 | |
| 2022 | 50.00 | 51.53 | 1183.74k* | 0.06l | 1.53m | 3.06 | 42.30 | 39.94 | 1929.36k* | 0.10l | 2.36m | 5.57 | |
| Female | |||||||||||||
| 2019 | 3 954 285 (55.83) | 2 647 860 (56.22) | 874.56* | 0.01 | 0.39 | 0.70 | 18 232 685 (55.00) | 5 038 450 (55.79) | 1634.46* | 0.02 | 0.79 | 1.43 | |
| 2022 | 3 641 665 (55.37) | 2 362 191 (55.44) | 841.96* | 0.00 | 0.07 | 0.14 | 17 536 023 (54.19) | 5 431 561 (55.25) | 1534.74* | 0.02 | 1.06 | 1.96 | |
| Male | |||||||||||||
| 2019 | 3 128 381 (44.17) | 2 090 611 (43.78) | 685.24* | 0.01 | 0.39 | 0.89 | 14 911 869 (45.00) | 4 097 532 (44.21) | 1314.41* | 0.02 | 0.79 | 1.75 | |
| 2022 | 2 935 293 (44.63) | 1 944 613 (44.56) | 677.53* | 0.00 | 0.07 | 0.17 | 14 757 951 (45.81) | 4 519 474 (44.75) | 1271.79* | 0.02 | 1.06 | 2.32 | |
| White | |||||||||||||
| 2019 | 4 451 889 (64.92) | 3 071 653 (63.02) | 922.66* | 0.04 | 1.91 | 2.94 | 18 063 933 (56.51) | 4 692 419 (53.51) | 1319.59* | 0.06 | 3.00 | 5.31 | |
| 2022 | 3 986 472 (62.81) | 2 699 140 (60.17) | 769.18* | 0.05 | 2.64 | 4.21 | 17 245 585 (55.62) | 5 222 003 (52.32) | 1218.81* | 0.07 | 3.31 | 5.94 | |
| Hispanic | |||||||||||||
| 2019 | 848 807 (12.38) | 733 463 (18.24) | 375.48 | 0.16 | 5.86 | 47.37 | 5 404 397 (15.85) | 2 334 103 (24.23) | 1024.87 | 0.21* | 8.37 | 52.83 | |
| 2022 | 898 194 (14.15) | 769 468 (20.84) | 381.71 | 0.18 | 6.69 | 47.25 | 5 854 273 (17.68) | 2 463 706 (26.09) | 914.74 | 0.20* | 8.41 | 47.58 | |
| Other | |||||||||||||
| 2019 | 1 556 720 (22.70) | 933 355 (18.74) | 57.71* | 0.10 | 3.96 | 17.43 | 8 897 295 (27.64) | 2 109 460 (22.26) | 27.80* | 0.12 | 5.38 | 19.45 | |
| 2022 | 1 461 983 (23.04) | 839 196 (18.99) | 52.53* | 0.10 | 4.04 | 17.55 | 8 371 341 (26.70) | 2 265 326 (21.59) | 42.68* | 0.12 | 5.11 | 19.12 | |
| Northeast | |||||||||||||
| 2019 | 1 281 479 (18.09) | 886 758 (20.12) | 162.95* | 0.05 | 2.03 | 11.20 | 5 896 656 (18.09) | 1 741 488 (19.66) | 415.62* | 0.04 | 1.57 | 8.68 | |
| 2022 | 1 168 095 (17.76) | 958 272 (20.95) | 34.63* | 0.08 | 3.19 | 17.97 | 4 556 368 (17.83) | 1 914 582 (18.34) | 620.22* | 0.01 | 0.51 | 2.86 | |
| Midwest | |||||||||||||
| 2019 | 1 568 100 (22.14) | 1 565 149 (22.42) | 401.41* | 0.01 | 0.28 | 1.27 | 7 278 819 (22.41) | 1 872 985 (16.52) | 286.68 | 0.15 | 5.90 | 26.31 | |
| 2022 | 1 407 981 (21.40) | 1 066 778 (17.95) | 79.71* | 0.09 | 3.45 | 16.13 | 7 057 339 (22.12) | 2 398 081 (17.66) | 6.30 | 0.11 | 4.46 | 20.14 | |
| South | |||||||||||||
| 2019 | 2 815 298 (39.74) | 1 115 017 (41.86) | 478.11* | 0.04 | 2.12 | 5.32 | 13 151 693 (40.69) | 2 788 427 (49.07) | 38.32 | 0.17 | 8.38 | 20.59 | |
| 2022 | 2 675 093 (40.66) | 1 070 126 (45.08) | 285.63* | 0.09 | 4.42 | 10.87 | 13 683 888 (40.25) | 2 930 342 (47.73) | 90.10* | 0.15 | 7.48 | 18.58 | |
| West | |||||||||||||
| 2019 | 1 418 928 (20.03) | 1 171 547 (15.61) | 44.20 | 0.12 | 4.42 | 22.08 | 6 820 083 (18.81) | 2 733 082 (14.76) | 61.98 | 0.11 | 4.05 | 21.54 | |
| 2022 | 1 327 203 (20.18) | 1 211 628 (16.02) | 12.20 | 0.11 | 4.16 | 20.61 | 7 000 525 (19.81) | 2 708 030 (16.27) | 87.71* | 0.09 | 3.53 | 17.83 | |
Healthcare Cost and Utilization Project.
National Inpatient Sample.
Nationwide Emergency Department Sample.
Method 1, Oracle Health Real-World Data.
Unweighted frequency (weighted percentage) unless otherwise noted.
Two one-sided tests (TOST) for equivalence of proportions via Fisher’s exact test, unless otherwise noted; asterisks indicate statistically significant equivalence.
Cohen’s h effect size (0.2 = small, 0.5 = medium, 0.8 = large), unless otherwise noted, bolded values with asterisks indicate differences that have achieved small or greater effect sizes.
Absolute percentage difference; unless otherwise noted.
Relative difference.
Mean (years).
TOST for equivalence of means via t-test.
Cohen’s d effect size (0.2 = small, 0.5 = medium, 0.8 = large).
Absolute mean difference.
Table 2.
Encounter-level demographics comparison between HCUPa NASSb and weighted (M1) OHRWDc for 2019 and 2022.
| Year | Outpatient surgery |
||||||
|---|---|---|---|---|---|---|---|
| Weighted estimates |
Database comparison |
||||||
| HCUP NASS | OHRWD | ||||||
| n (%)d | n (%)d | Ze | hf | Abs diff (%)g | Rel diff (%)h | ||
| Age i | |||||||
| 2019 | 51.35 | 51.66 | 760.32j* | 0.02k | 0.32l | 0.62 | |
| 2022 | 53.16 | 50.89 | 538.16j* | 0.11k | 2.27l | 4.28 | |
| Female | |||||||
| 2019 | 5 027 768 (55.91) | 215 807 (55.95) | 281.10* | 0.00 | 0.04 | 0.07 | |
| 2022 | 5 091 415 (56.01) | 240 081 (55.15) | 250.19* | 0.02 | 0.86 | 1.54 | |
| Male | |||||||
| 2019 | 3 965 432 (44.09) | 135 997 (44.05) | 221.48* | 0.00 | 0.04 | 0.09 | |
| 2022 | 4 004 160 (43.99) | 158 855 (44.85) | 192.02* | 0.02 | 0.86 | 1.96 | |
| White | |||||||
| 2019 | 6 421 160 (73.03) | 272 232 (76.61) | 323.51* | 0.08 | 3.58 | 4.90 | |
| 2022 | 6 413 871 (71.97) | 313 294 (77.34) | 256.75* | 0.12 | 5.36 | 7.45 | |
| Hispanic | |||||||
| 2019 | 898 943 (10.82) | 36 939 (10.72) | 83.29* | 0.00 | 0.10 | 0.94 | |
| 2022 | 981 805 (11.83) | 37 598 (10.28) | 29.41* | 0.05 | 1.62 | 13.68 | |
| Other | |||||||
| 2019 | 1 395 480 (16.15) | 42 633 (12.67) | 9.17 | 0.10 | 3.48 | 21.53 | |
| 2022 | 1 410 349 (16.20) | 48 044 (12.46) | 18.13 | 0.11 | 3.75 | 23.12 | |
Healthcare Cost and Utilization Project.
Nationwide Ambulatory Surgery Sample.
Method 1, Oracle Health Real-World Data.
Unweighted frequency (weighted percentage) unless otherwise noted.
Two one-sided tests (TOST) for equivalence of proportions via Fisher’s exact test, unless otherwise noted; asterisks indicate statistically significant equivalence.
Cohen’s h effect size (0.2 = small, 0.5 = medium, 0.8 = large), unless otherwise noted, bolded values with asterisks indicate differences that have achieved small or greater effect sizes.
Absolute percentage difference; unless otherwise noted.
Relative difference.
Mean (years).
TOST for equivalence of means via t-test.
Cohen’s d effect size (0.2 = small, 0.5 = medium, 0.8 = large).
Absolute mean difference.
When looking at Hispanic individuals, OHRWD M1 produced estimates that diverged from NIS/NEDS in both years with NEDS comparability achieving small differences in effect sizes (h = 0.20-0.21), however were closely aligned with NASS. Other race/ethnicity groups demonstrated negligible effect size differences between OHRWD M1 and all HCUP databases, with significant equivalence achieved when comparing to NIS and NEDS. When comparing census regions between OHRWD M1 and NIS/NEDS, all regions demonstrated negligible effect size differences.
Condition and procedure prevalence
Tables 3 and 4 present the comparisons in encounter-level condition and procedure prevalence between NIS, NEDS, NASS, and OHRWD M1 for 2019 and 2022. OHRWD M1 produced estimates closely aligned with all HCUP databases for low back pain (h = 0.00-0.10) and OUD (h = 0.01-0.06). Although hypertension prevalence diverged between OHRWD M1 and all HCUP databases with effect size differences reaching as high as 0.36, prevalences became, however, more closely aligned in 2022 when comparing with NIS (OHRWD M1 prevalence: 42.82%; NIS prevalence: 49.22%, h = 0.13) and NASS (OHRWD M1 prevalence: 32.43%; NASS prevalence: 36.41%, h = 0.08). Although slightly more variable than low back pain and OUD comparisons, DM exhibited negligible effect size differences between OHRWD M1 and all HCUP databases (h = 0.02-0.18). Although variability existed in procedure prevalence comparisons between OHRWD M1 and all HCUP databases, all effect size differences were negligible. Notably, appendectomy prevalence was closely aligned between OHRWD M1 and NIS in both years (2019—OHRWD M1 prevalence: 0.54%; NIS prevalence: 0.53%, h = 0.00, Z = 49.66*; 2022—OHRWD M1 prevalence: 0.52%; NIS prevalence: 0.49%, h = 0.00, Z = 37.54*) additionally demonstrating significant equivalence.
Table 3.
Encounter-level conditions and procedures comparison between HCUPa NISb/NEDSc and weighted (M1) OHRWDd for 2019 and 2022.
| Inpatient |
Emergency |
||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Weighted estimates |
Database comparison |
Weighted estimates |
Database comparison |
||||||||||
| Year | HCUP NIS n (%)e | OHRWD n (%)e | Zf | hg | Abs diff (%)h | Rel diff (%)i | HCUP NEDS n (%)e | OHRWD n (%)e | Zf | hg | Abs diff (%)h | Rel diff (%)i | |
| Conditions | |||||||||||||
| Low back pain | |||||||||||||
| 2019 | 138 241 (1.95) | 80 154 (1.98) | 104.42* | 0.00 | 0.03 | 1.56 | 918 964 (2.77) | 265 490 (2.99) | 159.13* | 0.01 | 0.22 | 7.90 | |
| 2022 | 115 567 (1.76) | 73 841 (2.17) | 18.22 | 0.03 | 0.42 | 23.81 | 787 863 (2.43) | 262 552 (2.77) | 71.12* | 0.02 | 0.34 | 14.02 | |
| OUDj | |||||||||||||
| 2019 | 189 254 (2.67) | 76 677 (1.78) | 97.06 | 0.06 | 0.88 | 33.19 | 334 936 (1.01) | 62 240 (0.78) | 22.89 | 0.02 | 0.23 | 22.67 | |
| 2022 | 163 853 (2.49) | 76 268 (2.19) | 50.47* | 0.02 | 0.30 | 11.98 | 323 752 (1.01) | 67 473 (0.76) | 36.92 | 0.03 | 0.25 | 24.35 | |
| Hypertension | |||||||||||||
| 2019 | 3 449 833 (48.70) | 1 243 973 (32.38) | 549.14 | 0.33 * | 16.32 | 33.51 | 8 774 621 (26.38) | 1 191 404 (14.13) | 1416.51 | 0.31* | 12.25 | 46.45 | |
| 2022 | 3 238 195 (49.22) | 1 472 566 (42.82) | 263.00* | 0.13 | 6.41 | 13.02 | 8 277 550 (25.32) | 1 329 173 (14.88) | 1069.68 | 0.26* | 10.44 | 41.25 | |
| DMk | |||||||||||||
| 2019 | 1 701 387 (24.02) | 726 601 (18.77) | 44.01 | 0.13 | 5.25 | 21.85 | 4 243 651 (12.75) | 642 756 (7.26) | 794.04 | 0.18 | 5.49 | 43.05 | |
| 2022 | 1 625 772 (24.71) | 763 446 (21.90) | 192.61* | 0.07 | 2.81 | 11.37 | 4 052 025 (12.38) | 665 520 (7.08) | 757.80 | 0.18 | 5.30 | 42.81 | |
| Procedures | |||||||||||||
| PT | |||||||||||||
| 2019 | 32 486 (0.46) | 8890 (0.18) | 134.07 | 0.05 | 0.27 | 60.25 | 2949 (0.01) | 899 (0.01) | 4.13* | 0.00 | 0.00 | 14.53 | |
| 2022 | 31 406 (0.48) | 25 247 (0.24) | 95.21 | 0.04 | 0.24 | 50.60 | 7382 (0.02) | 3665 (0.01) | 23.77 | 0.01 | 0.01 | 38.82 | |
| Colonoscopy | |||||||||||||
| 2019 | 54 355 (0.77) | 24 239 (0.58) | 16.81 | 0.02 | 0.19 | 24.42 | 22 974 (0.07) | 3107 (0.03) | 96.16 | 0.02 | 0.04 | 57.85 | |
| 2022 | 51 726 (0.79) | 25 220 (0.61) | 7.94 | 0.02 | 0.17 | 22.19 | 44 299 (0.14) | 3602 (0.03) | 219.35 | 0.04 | 0.11 | 77.31 | |
| Appendectomy | |||||||||||||
| 2019 | 37 456 (0.53) | 24 154 (0.54) | 49.66* | 0.00 | 0.02 | 2.96 | 17,616 (0.05) | 3371 (0.03) | 58.08 | 0.01 | 0.02 | 47.34 | |
| 2022 | 32 370 (0.49) | 23 171 (0.52) | 37.54* | 0.00 | 0.03 | 5.75 | 31 093 (0.10) | 4272 (0.03) | 144.42 | 0.03 | 0.06 | 67.36 | |
Healthcare Cost and Utilization Project.
National Inpatient Sample.
Nationwide Emergency Department Sample.
Method 1, Oracle Health Real-World Data.
Unweighted frequency (weighted percentage).
Two one-sided tests (TOST) for equivalence of proportions via Fisher’s exact test, asterisks indicate statistically significant equivalence.
Cohen’s h effect size (0.2 = small, 0.5 = medium, 0.8 = large), bolded values with asterisks indicate differences that have achieved small or greater effect sizes.
Absolute difference.
Relative difference.
Opioid use disorder.
Diabetes mellitus.
Table 4.
Encounter-level conditions and procedures comparison between HCUPa NASSb and weighted (M1) OHRWDc for 2019 and 2022.
| Year | Outpatient surgery |
||||||
|---|---|---|---|---|---|---|---|
| Weighted estimates |
Database comparison |
||||||
| HCUP NASS n (%)d | OHRWD n (%)d | Ze | hf | Abs diff (%)g | Rel diff (%)h | ||
| Conditions | |||||||
| Low back pain | |||||||
| 2019 | 128 769 (1.44) | 9851 (2.79) | 96.40 | 0.10 | 1.35 | 93.74 | |
| 2022 | 117 178 (1.29) | 10 267 (2.43) | 80.83 | 0.09 | 1.14 | 88.08 | |
| OUDi | |||||||
| 2019 | 10 293 (0.12) | 566 (0.16) | 8.53 | 0.01 | 0.05 | 40.06 | |
| 2022 | 11 083 (0.13) | 800 (0.16) | 3.20 | 0.01 | 0.03 | 27.36 | |
| Hypertension | |||||||
| 2019 | 3 133 583 (34.93) | 76 397 (19.01) | 313.37 | 0.36 * | 15.92 | 45.59 | |
| 2022 | 3 287 011 (36.41) | 111 890 (32.43) | 95.86* | 0.08 | 3.98 | 10.92 | |
| DMj | |||||||
| 2019 | 1 320 185 (14.72) | 35 150 (12.41) | 27.14* | 0.07 | 2.31 | 15.70 | |
| 2022 | 1 372 462 (15.16) | 46 508 (14.61) | 95.47* | 0.02 | 0.55 | 3.65 | |
| Procedures | |||||||
| Appendectomy | |||||||
| 2019 | 207 595 (2.34) | 8677 (1.68) | 20.90 | 0.05 | 0.66 | 28.27 | |
| 2022 | 196 605 (2.19) | 8150 (2.08) | 31.26* | 0.01 | 0.11 | 5.02 | |
Healthcare Cost and Utilization Project.
Nationwide Ambulatory Surgery Sample.
Method 1, Oracle Health Real-World Data.
Unweighted frequency (weighted percentage).
Two one-sided tests (TOST) for equivalence of proportions via Fisher’s exact test, asterisks indicate statistically significant equivalence.
Cohen’s h effect size (0.2 = small, 0.5 = medium, 0.8 = large), bolded values with asterisks indicate differences that have achieved small or greater effect sizes.
Absolute difference.
Relative difference.
Opioid use disorder.
Diabetes mellitus.
Supplemental comparative methods
Tables S19a-S19c display the additional comparative methods M2-M4 (unified weighting, unified weighting with encounter type multiplier, IPF “raking,” respectively). As determined by effect sizes, M2-M4 demonstrated similar comparability to M1 when looking at low back pain and all procedures among all HCUP databases. Notably, M4 (IPF “raking”) demonstrated even greater alignment to HCUP databases in 2022 for many of these low back pain and procedure outcomes. OUD diverged more slightly across the comparative methods while still maintaining close alignment. DM diverged even more so, with the exception of close alignment in M2 and M3 when comparing against NEDS. Hypertension showed the greatest divergence across methods. M3 and M4 improved alignment only in selected 2019 comparisons, particularly for NASS and marginally for NEDS with M3, but they did not consistently improve on M1. Tables S20a-S20c display the final comparative method M5 (separate encounter type weighting with post-stratification demographic adjustment). M5 generally improved low back pain alignment in several settings and improved OUD alignment for NIS, but OUD results were mixed for NEDS and NASS. These findings support M1 as the primary analytic approach, with M4 serving mainly as a secondary sensitivity analysis rather than a replacement default.
Supplemental inferential validation analysis
Figure S2 displays the association between history of OUD and utilization of PT among patients with a new episode of low back pain in an inpatient setting in 2022. While incorporating M1 weights, those with a history of OUD had a significantly lower likelihood of PT utilization (aOR: 0.63; 95% CI, 0.48-0.83). When unweighted, this association became weaker and lost significance (aOR: 0.85; 95% CI, 0.60-1.21).
Discussion
In this study, by applying a structured weighting framework to a large, convenience-sample EHR database, we substantially improved encounter-level alignment and comparability to national hospital benchmarks. Using the OHRWD, we developed and validated a multivariate, stratified weighting scheme informed by AHA data and benchmarked it against the 3 HCUP national encounter-level databases (NIS, NEDS, and NASS). This encounter-type specific design allowed us to validate the weighting approach across inpatient, emergency, and ambulatory care in a single study, which, to our knowledge, has not been done previously. The primary weighting method (M1) created weights to align OHRWD encounter distributions with national benchmarks. Validation of M1 against HCUP demonstrated close alignment for most demographic and clinical measures across both study years, particularly for age, sex, and conditions such as low back pain, OUD, as well as procedures such as appendectomy. When assessing the inferential association between history of OUD and utilization of PT in an inpatient setting for new episodes of low back pain, the association (aOR: 0.63) closely mirrored the literature (aOR: 0.62-0.65) when incorporating weights and was attenuated and lost significance when ignoring weights.17,18 These findings support the feasibility and value of weighting EHR-based data to improve encounter-level comparability to national hospital benchmarks. The supplemental inferential analysis provides an illustrative example of potential downstream analytic impact, although additional work is needed across outcomes and settings.
Patterns of alignment and divergence were observed in the weighted demographic and clinical characteristic estimates. The strongest alignment in both years was for demographic characteristics that are routinely seen across the majority of hospitals, including age and sex. When looking at race/ethnicity, although White and “Other” racial groups demonstrated greater alignment, evidence of divergence was observed in Hispanic ethnicity patients particularly in emergency settings. This is consistent with prior literature finding greater EHR race/ethnicity classification accuracy in White and Black patient information but higher misclassification in Hispanic/Latinx patient information, generally and particularly in the case of emergency visit data.19,20 Prior research has noted reasons for higher misclassification among Hispanic/Latino patients including a strong identification with the ethnicity field, but not the race category field, and language barriers.21–23 Strong patterns of divergence can lead to residual divergence even after implementing weighting strategies. While negligible effect size differences existed in regional comparisons, slight variation still existed. These demographic differences could partly be attributable to selection biases of health systems included in OHRWD.24 Looking at the distribution of health systems in 2022 alone, approximately 9% of OHRWD systems were considered small, whereas NIS, for example, had ∼23% small. Over half of OHRWD systems were very large, whereas NIS demonstrated more even distribution in size. While the weighting methodology implemented was intended to correct much of this variation seen in all characteristics, some residual variation remained as is common and seen in previous work.24 Prevalence of low back pain and OUD was closely aligned between databases, of which the comparability of OUD post weighting is noteworthy given the known disproportionality of this condition among underserved communities leading to greater misrepresentation when unweighted.25,26
Hypertension prevalence demonstrated noticeable divergence which could be explained by the fact that HCUP may have a higher likelihood of capturing non-acute comorbidities than EHR, which are not likely to be the primary reason for visit yet still be very common and extend across ensuing encounters.27 Other research has noted that EHR data significantly under-document multimorbidity, like conditions such as hypertension, compared with claims.28 In general, claims data are primarily used for billing and thus chronic diseases will be more likely to be continuously recorded in the patient’s claims (in structured fields). In contrast, EHR’s are a compilation of the patient’s health state across numerous structured (diagnostic fields) and unstructured (provider notes) sources, and more formal structured indications are only more likely to be updated when a disease state changes. Thus, the level of capture across visits and ensuing encounter-level prevalence may be different between claims and EHR.29 In general, studies have shown limitations of structured coded data in capturing true comorbidities,30 and while EHR is rich in clinical markers, continued efforts to build phenotypes that leverage all available clinical information to improve accuracy of capturing such comorbidities as hypertension will prove useful without relying on weighting efforts alone.31 Notwithstanding these differences, prevalence comparability improved in 2022 for hypertension and DM, suggesting that expanded network coverage and more complete digitization over time may help reduce discrepancies. OHRWD captured fewer procedures than HCUP generally, which is consistent with prior research possibly due to EHR databases tending to capture fewer elements related to provider orders.24,32 While appendectomy estimates were more closely aligned, other therapies such as PT were more divergent which is consistent with prior research24 and may be driven by nonstructured PT treatment record keeping33 as well as encounter types not largely common with PT. However, this again reiterates the importance of designing standardized methods to translate rich clinical information into structured diagnostic and procedural indicators.
Our study extends previous work by DeShazo and Hoffman24 aligning the EHR database Cerner Health Facts with HCUP NIS. Their work evaluated an earlier generation of Cerner-derived data against a single inpatient benchmark by constructing weights based on bed size and region strata. In contrast, we use a newer Oracle EHR real-world data platform and explicitly stratify weights by encounter type while validating representativeness simultaneously across inpatient, emergency department, and ambulatory surgery encounters using NIS, NEDS, and NASS. Our findings also parallel results from other large-scale initiatives designed to address representativeness in EHR data, such as CDC MENDS and the All of Us Research Program.2,5 Unlike Li et al.’s approach to MENDS, which applied parish-level post-stratification using Census-based distributions of age, sex, and race to correct hypertension prevalence estimates in the single state of Louisiana,2 our method, similar to HCUP, incorporates hospital characteristics into the weighting process and further advances this approach by adding encounter-type stratification. Specifically, we stratified by US 1-digit zip code region, hospital or health-system indicator, bed size, and encounter type (inpatient, emergency, ambulatory), which is a combination of elements not included in prior EHR weighting studies. This structural layer addressed a major limitation of MENDS by capturing heterogeneity in healthcare utilization across facility types. Similar to our study using data from across the whole US, Wang et al.5 addressed the challenge of improving the representativeness of a large non-probability dataset by constructing survey weights for the All of Us Research Program using National Health Interview Survey benchmarks. However, their primary approach was raking based on extensive demographic characteristics, whereas our framework uses stratified weighting based on hospital and encounter structure, with raking included only as a supplemental comparison. Differences between databases do exist in data type in which All of Us couples comprehensive self-reported demographic and socioeconomic survey data with EHR from a modest national sample, whereas OHRWD represents extensive clinical EHR from a much larger national sample albeit with limited socioeconomic information. Nevertheless, both studies demonstrate that weighting, whether demographic raking or multilevel structural adjustment, can improve alignment with external benchmarks and support broader analytic use of non-probability datasets. Beyond Li et al. and Wang et al. recent biobank and EHR-linked studies have applied weighting to address selection bias. In the UK Biobank, Schoeler et al.34 estimated participation probabilities with a lasso model and applied inverse-probability weights, improving representativeness for many socio-behavioral traits. Likewise, Blay et al.35 (GCAT cohort) used raking against multiple Catalonia population benchmarks (SIDIAP, ESCA, IDESCAT, INE) bringing chronic disease prevalences much closer to population values, although conditions linked to lifestyle factors such as alcohol use and smoking remained underestimated. These approaches rely primarily on sociodemographic and behavioral variables. In contrast, our primary weighting strategy (M1) focuses on hospital- and encounter-level structure, while demographic calibration was examined only as a secondary, comparative method (M5).
Comparative methods (M2-M5) demonstrated mixed performance rather than consistent improvement over the primary approach. Unified weighting schemes (M2 and M3) produced results similar to M1 for low back pain and all procedures but diverged more for hypertension and DM when compared to NIS. Raking (M4) aligned similarly to M1 for low back pain but showed mixed performance for hypertension, performing worse in most settings and years and better only in NASS in 2019. Overall, we recommend M1 as the default approach when the goal is encounter-level comparability to national hospital benchmarks because it performed most consistently across settings and outcomes; M4 may be useful as a sensitivity analysis when marginal calibration is the primary objective, but it did not consistently improve alignment in our study. The final comparative method (M5), which added a second stage Census-based demographic correction to M1, produced closer alignment for some outcomes such as OUD and low back pain, but did not improve and in some cases worsened alignment for other conditions and procedures. This indicates that adding demographic post-stratification did not consistently improve alignment. In particular, because OHRWD represents healthcare encounters rather than the general population, national demographic proportions may not necessarily reflect the underlying distribution of patients receiving hospital-based care.36 Future work should investigate how to incorporate demographic information in a way that improves representativeness without diverging estimates further from true national hospital patterns.
Together, these findings provide a strong benchmark for structural and demographic alignment and position our weighting framework as both complementary and additive to existing methods. While demographic raking, as implemented by Wang et al.5 and stratified weighting, as implemented by DeShazo and Hoffman24 effectively reduce bias in population-based or regional EHR datasets, our encounter-type-stratified design uniquely addresses selection and utilization biases specific to large, multi-system EHR networks. At the same time, residual divergence for Hispanic populations and some chronic conditions indicates that structural weighting alone cannot fully overcome characteristic under/over-representation, a limitation shared across EHR weighting efforts.
Limitations
Several limitations should be considered when interpreting these results. First, OHRWD is a convenience sample of participating health systems, and its sampling frame cannot be fully balanced. Second, small cell counts in certain strata occasionally produced extreme weights despite trimming. Third, utilizing more granular hospital regionality and more specific hospital characteristics would have likely improved the weights further compared to what we were allotted in the current study (combination of system and hospital level granularity). Fourth, information on hospital ownership and teaching status was unavailable, preventing exact replication of HCUP’s stratification design. Fifth, residual discrepancies may reflect not only weighting limitations but also differences in how data are captured across the 2 data sources, as prior research has shown claims data (like HCUP) can be very different than patient EHR data captured by providers including differences in documentation practices, billing-focused vs patient health-state capture, and completeness of chronic condition recording.24,30,37–39 OHRWD may include healthcare services that are documented before billing occurs or that are never billed at all, meaning they would not be captured in claims-based datasets such as HCUP and could affect comparisons. Although we attempted to mirror HCUP’s inclusion and exclusion criteria, our interpretation may not have fully matched HCUP’s operational definitions. Moreover, OHRWD uses several different code types for diagnoses or procedures whereas HCUP relies primarily on a single code type in a given year for each. To maintain comparability, we matched HCUP’s coding type which meant that some procedures may be counted differently or not at all in OHRWD. Sixth, differences in the types of services available across hospitals may also contribute to discrepancies between datasets. If one sample includes more hospitals that provide certain specialty services and the other does not, the prevalence of related conditions and procedures will differ regardless of weighting. Finally, HCUP served as the gold standard in this study, but other national benchmarks may yield different alignment patterns.
Conclusion
This study demonstrates that carefully constructed, stratified weighting can improve encounter-level alignment and comparability between OHRWD and national hospital benchmarks. When benchmarked against HCUP databases, weighted OHRWD produced demographic and clinical estimates that were generally closer to national encounter patterns, although residual differences remained for hypertension and Hispanic/Latino ethnicity. The supplemental inferential example further suggests that weighting can materially affect substantive conclusions. These findings may support broader use of OHRWD while indicating that additional calibration may still be needed for selected domains.
Supplementary Material
Acknowledgments
The authors thank Oracle Health and the American Hospital Association for data access and the HCUP team at AHRQ for publicly available comparison datasets.
Contributor Information
Fares Qeadan, Parkinson School of Health Sciences and Public Health, Loyola University Chicago, Maywood, IL 60153, United States.
Benjamin Tingey, Parkinson School of Health Sciences and Public Health, Loyola University Chicago, Maywood, IL 60153, United States.
Mirjana Glisovic Bensa, Parkinson School of Health Sciences and Public Health, Loyola University Chicago, Maywood, IL 60153, United States.
Erin F Madden, Department of Family Medicine and Public Health Sciences, Wayne State University, Detroit, MI 48201, United States.
Pooja Lagisetty, Department of Internal Medicine, University of Michigan, Ann Arbor, MI 48109, United States; Center for Clinical Management and Research, Ann Arbor Veterans Health Administration, Ann Arbor, MI 48105, United States.
Philip J Kroth, Department of Biomedical Informatics, Western Michigan University, Homer Stryker M.D. School of Medicine, Kalamazoo, MI 49007, United States.
Author contributions
Fares Qeadan (Conceptualization, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Writing—original draft, Writing—review & editing), Benjamin Tingey (Formal analysis, Software, Validation, Visualization, Writing—original draft, Writing—review & editing), Mirjana Glisovic Bensa (Investigation, Writing—original draft), Erin F. Madden (Investigation, Writing—review & editing), Pooja Lagisetty (Investigation, Writing—review & editing), and Philip J. Kroth (Investigation, Writing—review & editing)
Supplementary material
Supplementary material is available at JAMIA Open online.
Funding
This work was supported in part by the National Institute on Drug Abuse (NIDA) of the National Institutes of Health (NIH) under grant number R01DA057658. The funder had no role in study design, analysis, or interpretation.
Conflicts of interest
The authors declare no conflicts of interest.
Data availability
The data underlying this article cannot be shared publicly due to data use agreements. Oracle Health Real-World Data (OHRWD) and American Hospital Association datasets are proprietary. HCUP databases are available from the Agency for Healthcare Research and Quality (AHRQ) at https://hcup-us.ahrq.gov. OHRWD could be accessed by signing a data sharing agreement with Oracle Cerner and covering any costs that may be involved (Contact Kendra Stillwell: kendra.stillwell@cernerenviza.com).
Ethics approval
The Loyola University Chicago Institutional Review Board determined this secondary research to be exempt under 45 CFR 46.104(d)(4)(ii-iii). Data were recorded such that subjects could not be identified directly or indirectly, investigators did not contact subjects and did not re-identify subjects, and use of identifiable health information was regulated under HIPAA as described in the exemption determination (LU# 215817).
References
- 1. Salvatore M, Kundu R, Shi X, et al. To weight or not to weight? The effect of selection bias in 3 large electronic health record-linked biobanks and recommendations for practice. J Am Med Inform Assoc. 2024;31:1479-1492. 10.1093/jamia/ocae098 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Li W, Zhang M, Cheng J, et al. Weighted EHR-based prevalence estimates for hypertension at the state and local levels in Louisiana. BMC Public Health. 2025;25:432. 10.1186/s12889-025-21633-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Boyd AD, Gonzalez-Guarda R, Lawrence K, et al. Potential bias and lack of generalizability in electronic health record data: reflections on health equity from the National Institutes of Health Pragmatic Trials Collaboratory. J Am Med Inform Assoc. 2023;30:1561-1566. 10.1093/jamia/ocad115 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Verheij RA, Curcin V, Delaney BC, McGilchrist MM. Possible sources of bias in primary care electronic health record data use and reuse. J Med Internet Res. 2018;20:e185. 10.2196/jmir.9134 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Wang VH-C, Lei J, Shi T, Pagán JA. Weighting the United States All of Us Research Program data to known population estimates using raking. Prev Med Rep. 2024;43:102795. 10.1016/j.pmedr.2024.102795 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Britz JB, O’Loughlin KM, Henry TL, et al. Rising racial disparities in opioid mortality and undertreatment of opioid use disorder and mental health comorbidities in Virginia. AJPM Focus. 2023;2:100102. 10.1016/j.focus.2023.100102 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Lynch S, Katkhuda F, Klepacz L, Towey E, Ferrando SJ. Racial disparities in opioid use disorder and its treatment: a review and commentary on the literature. J Ment Health Clin Psychol. 2023;7:13-18. 10.29245/2578-2959/2023/1.1263 [DOI] [Google Scholar]
- 8. Ehwerhemuepha L, Carlson K, Moog R, et al. Cerner real-world data (CRWD) - a de-identified multicenter electronic health records database. Data Brief. 2022;42:108120. 10.1016/j.dib.2022.108120 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Healthcare Cost and Utilization Project (HCUP). Overview of the National (Nationwide) Inpatient Sample (NIS). Agency for Healthcare Research and Quality. Updated September 10, 2025. Accessed November 21, 2025. https://hcup-us.ahrq.gov/nisoverview.jsp
- 10. Healthcare Cost and Utilization Project (HCUP). Overview of the Nationwide Emergency Department Sample (NEDS). Agency for Healthcare Research and Quality. Updated July 16, 2025. Accessed November 21, 2025. https://hcup-us.ahrq.gov/nedsoverview.jsp
- 11. Healthcare Cost and Utilization Project (HCUP). Overview of the Nationwide Ambulatory Surgery (NASS). Agency for Healthcare Research and Quality. Updated July 16, 2025. Accessed November 21, 2025. https://hcup-us.ahrq.gov/nassoverview.jsp
- 12. Alho J, Spencer B. Statistical Demography and Forecasting. Springer-Verlag; 2005. https://www.springer.com/us/book/9780387235301 [Google Scholar]
- 13. Houchens R, Elixhauser A, Sommers J. Changes in the NIS Sampling and Weighting Strategy for 1998. Agency for Healthcare Research and Quality; 2002. https://www.hcup-us.ahrq.gov/db/nation/nis/reports/Changes_in_NIS_Design_1998.pdf [Google Scholar]
- 14. Deming WE, Stephan FF. On a least squares adjustment of a sampled frequency table when the expected marginal totals are known. Ann Math Stat. 1940;11:427-444. 10.1214/aoms/1177731829 [DOI] [Google Scholar]
- 15. Fienberg SE, Meyer MM. Iterative proportional fitting. Encycl Stat Sci. 2006. https://onlinelibrary.wiley.com/doi/abs/10.1002/0471667196.ess1297.pub2 [Google Scholar]
- 16. Schuirmann DJ. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J Pharmacokinet Biopharm. 1987;15:657-680. 10.1007/BF01068419 [DOI] [PubMed] [Google Scholar]
- 17. Magel JS, Gordon AJ, Fritz JM, Kim J. The influence of an opioid use disorder on initiating physical therapy for low back pain: a retrospective cohort. J Addict Med. 2021;15:226-232. 10.1097/ADM.0000000000000751 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Moyo P, Merlin JS, Gairola R, et al. Association of opioid use disorder diagnosis with management of acute low back pain: a medicare retrospective cohort analysis. J Gen Intern Med. 2024;39:2097-2105. 10.1007/s11606-024-08799-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Johnson JA, Moore B, Hwang EK, Hickner A, Yeo H. The accuracy of race and ethnicity data in US based healthcare databases: a systematic review. Am J Surg. 2023;226:463-470. 10.1016/j.amjsurg.2023.05.011 [DOI] [PubMed] [Google Scholar]
- 20. Pettit NR, Lane KA, Gibbs L, Musey P, Li X, Vest JR. Concordance between electronic health record-recorded race and ethnicity and patient report in emergency department patients. Ann Emerg Med. 2024;84:111-117. 10.1016/j.annemergmed.2024.03.025 [DOI] [PubMed] [Google Scholar]
- 21. Hoang K, Gold J, Powell C, et al. Concordance between electronic health record-recorded race/ethnicity and parental report in hospitalized children. J Hosp Med. 2023;18:610-616. 10.1002/jhm.13140 [DOI] [PubMed] [Google Scholar]
- 22. Thorlby R, Jorgensen S, Siegel B, Ayanian JZ. How health care organizations are using data on patients’ race and ethnicity to improve quality of care. Milbank Q. 2011;89:226-255. 10.1111/j.1468-0009.2011.00627.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Gore A, Truche P, Iskerskiy A, Ortega G, Peck G. Inaccurate ethnicity and race classification of Hispanics following trauma admission. J Surg Res. 2021;268:687-695. 10.1016/j.jss.2021.08.003 [DOI] [PubMed] [Google Scholar]
- 24. DeShazo JP, Hoffman MA. A comparison of a multistate inpatient EHR database to the HCUP nationwide inpatient sample. BMC Health Serv Res. 2015;15:384. 10.1186/s12913-015-1025-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Gondré-Lewis MC, Abijo T, Gondré-Lewis TA. The opioid epidemic: a crisis disproportionately impacting black Americans and urban communities. J Racial Ethn Health Disparities. 2023;10:2039-2053. 10.1007/s40615-022-01384-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Hsu M, Jung OS, Kwan LT, et al. Access challenges to opioid use disorder treatment among individuals experiencing homelessness: voices from the streets. J Subst Use Addict Treat. 2024;157:209216. 10.1016/j.josat.2023.209216 [DOI] [PubMed] [Google Scholar]
- 27. Owens PL, Liang L, Barrett ML, Fingar KR. Comorbidities associated with adult inpatient stays, 2019. Healthcare Cost and Utilization Project (HCUP) Statistical Briefs. Statistical Brief No. 303. Agency for Healthcare Research and Quality (US);2022. Updated December 15, 2022, Accessed December 2, 2025. https://www.ncbi.nlm.nih.gov/books/NBK588380/
- 28. Xiong MC, Pittell H, Kitchen C, Lasser EC, Kharrazi H. Synergy of diagnosis coding between administrative claims and electronic health records of large patient populations across multiple healthcare organizations. JAMIA Open. 2026;9:ooag030. 10.1093/jamiaopen/ooag030 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Rassen JA, Bartels DB, Schneeweiss S, Patrick AR, Murk W. Measuring prevalence and incidence of chronic conditions in claims and electronic health record databases. Clin Epidemiol. 2019;11:1-15. 10.2147/CLEP.S181242 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Pan J, Lee S, Cheligeer C, et al. Assessing the validity of ICD-10 administrative data in coding comorbidities. BMJ HealthCare Inform. 2025;32:e101381. 10.1136/bmjhci-2024-101381 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Hohman KH, Zambarano B, Klompas M, et al. Development of a hypertension electronic phenotype for chronic disease surveillance in electronic health records: key analytic decisions and their effects. Prev Chronic Dis. 2023;20:E80. 10.5888/pcd20.230026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Parsons A, McCullough C, Wang J, Shih S. Validity of electronic health record-derived quality measurement for performance monitoring. J Am Med Inform Assoc. 2012;19:604-609. 10.1136/amiajnl-2011-000557 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Franklin PD, Oatis CA, Zheng H, et al. Web-based system to capture consistent and complete real-world data of physical therapy interventions following total knee replacement: Design and evaluation study. JMIR Rehabil Assist Technol. 2022;9:e37714. 10.2196/37714 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Schoeler T, Speed D, Porcu E, Pirastu N, Pingault JB, Kutalik Z. Participation bias in the UK biobank distorts genetic associations and downstream analyses. Nat Hum Behav. 2023;7:1216-1227. 10.1038/s41562-023-01579-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Blay N, Carrasco-Ribelles LA, Farré X, et al. Weighting health-related estimates in the GCAT cohort and the general population of Catalonia. Sci Rep. 2025;15:16984. 10.1038/s41598-025-01284-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Hull SA, Rivas C, Bobby J, Boomla K, Robson J. Hospital data may be more accurate than census data in estimating the ethnic composition of general practice populations. Inform Prim Care. 2009;17:67-78. 10.14236/jhi.v17i2.718 [DOI] [PubMed] [Google Scholar]
- 37. Peabody JW, Luck J, Jain S, Bertenthal D, Glassman P. Assessing the accuracy of administrative data in health information systems. Med Care. 2004;42:1066-1072. 10.1097/00005650-200411000-00005 [DOI] [PubMed] [Google Scholar]
- 38. Fisher ES, Whaley FS, Krushat WM, et al. The accuracy of medicare’s hospital claims data: progress has been made, but problems remain. Am J Public Health. 1992;82:243-248. 10.2105/AJPH.82.2.243 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Jollis JG, Ancukiewicz M, DeLong ER, Pryor DB, Muhlbaier LH, Mark DB. Discordance of databases designed for claims payment versus clinical information systems: implications for outcomes research. Ann Intern Med. 1993;119:844-850. 10.7326/0003-4819-119-8-199310150-00011 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying this article cannot be shared publicly due to data use agreements. Oracle Health Real-World Data (OHRWD) and American Hospital Association datasets are proprietary. HCUP databases are available from the Agency for Healthcare Research and Quality (AHRQ) at https://hcup-us.ahrq.gov. OHRWD could be accessed by signing a data sharing agreement with Oracle Cerner and covering any costs that may be involved (Contact Kendra Stillwell: kendra.stillwell@cernerenviza.com).
