Skip to main content
JAMIA Open logoLink to JAMIA Open
. 2026 May 29;9(3):ooag084. doi: 10.1093/jamiaopen/ooag084

Assessing data quality of inflammatory bowel disease patients in the All of Us research program

Matthew Spotnitz 1,✉, Adam S Faye 2, John Giannini 3, Tamara R Litwin 4, Yechiam Ostchega 5, Lew Berman 6
PMCID: PMC13220751  PMID: 42220339

Abstract

Purpose

Inflammatory bowel disease (IBD) consists of Crohn’s disease (CD) and ulcerative colitis (UC) and is a spectrum autoimmune disease of the gastrointestinal tract. Large scale real-world evidence studies could provide valuable evidence about IBD for personalized healthcare recommendations. The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) standardizes electronic health record (EHR) data, allowing for research that incorporates multiple data sources. We are interested in whether OMOP CDM data on IBD are fit-for-use.

Methods

We selected IBD diagnosis codes to define the phenotype. We used a data quality checklist to evaluate 5 domains: conformance, completeness, concordance, plausibility, and temporality. We also did sensitivity analyses for CD and UC that consisted of at least 2 diagnosis codes that were at least 30 days apart.

Results

All of the phenotype-defining ICD source codes mapped to SNOMED. Many concept prevalences were low. A total of 78 (30.1%) out of 253 concept correlations were above our strength threshold (⍴ > 0.5). The age distribution of concepts and relative frequency of IBD medications were plausible. The median time between diagnosis and biopsy for the cohort was 4.43 [-0.05, 104.29] weeks. For the subgroup of participants who had sufficient data for the timeline analysis, IBD diagnosis concepts tended to occur first. In our sensitivity analyses, the completeness percentages of many variables in the UC and CD subgroups were similar to IBD, except for disease specific workup and treatment concepts.

Conclusion

We have shown a novel implementation of our data quality framework on IBD cohorts.

Keywords: inflammatory bowel disease, Crohn’s disease, ulcerative colitis, precision medicine, electronic health record, data quality

Introduction

Inflammatory bowel disease (IBD) includes Crohn’s disease (CD) and ulcerative colitis (UC), which are autoimmune diseases that affect the gastrointestinal tract. Both conditions cause long-term morbidity, lack definitive curative medical therapy, and have poorly understood genetic underpinnings.1 The hallmark symptoms include fever, abdominal pain, bloody diarrhea, and weight loss. For moderate to severe IBD cases, corticosteroids are often initiated and used as a bridge to steroid-sparing treatments such as biologics, small molecule inhibitors, and immunomodulators. Despite the growing number of available therapies, current treatments can have limited efficacy in achieving endoscopic remission. Consequently, individuals with IBD can have multimorbidity and adverse outcomes that result from ongoing inflammation. Therefore, precision medicine research may lead to personalized healthcare recommendations and improved treatment strategies for individuals with IBD.

The All of Us Research Program is a multi-institutional cohort study and precision medicine initiative. It aims to enroll one million participants and collect multiple data streams, including electronic health records (EHRs), biospecimens, genomics, and survey data.2 The data repository allows for multiomics analyses, which may accelerate healthcare recommendations that are tailored to an individual.

The All of Us program uses the Observational Medical Outcome Partnership Common Data Model (OMOP CDM) in order to improve the exchangeability and interoperability of source data. The OMOP CDM consists of standardized concepts and relationships, to improve harmonization across different data domains (eg, conditions, procedures, measurements), sources (eg, EHR vendors) and types (eg, EHR, genetic, survey data). OMOP CDM concepts use codes from structured medical terminologies (eg, ICD-09, LOINC) as the source. The standardized concept relationships comprise schema across different data tables.3–6

Combinations of different data types may be uniquely valuable for precision medicine studies about IBD. Previously, we described a data quality framework, which has been used to evaluate cohorts of All of Us participants who had a ductal carcinoma in situ (DCIS) diagnosis, a surgical oncology procedure, or rheumatoid or psoriatic arthritis diagnoses.7–10 In this analysis, we used the same framework to assess the data quality of participants diagnosed with IBD. Specifically, these analyses aim to determine the extent to which data from the All of Us program are fit for use for investigations of IBD, CD, and UC.

Methods

Recruitment and data curation

All of Us participants consent to enroll in the program either independently or at a participating site. Upon enrollment, participants complete a personal and family history health survey. Furthermore, they have the option of submitting retrospective and prospective EHR data, wearable data, and donating biospecimens. The data are curated at participating sites and then submitted to a central repository.

Study design

We designed nested case cohorts of participants with an IBD phenotype. Specifically, our cohorts were defined by at least one occurrence of an International Classification of Diseases, Ninth Revision (ICD-9), or Systematized Nomenclature of Medicine—Clinical Terms (SNOMED CT) diagnosis code for CD or UC. Our study participants were individuals 18 years and older. In this analysis, we restricted ourselves to participants with at least one EHR data point. EHR data that were antecedent to the date of enrollment were included. Our control cohorts consisted of participants aged 18 and older with EHR data who did not have any phenotype defining diagnosis codes for IBD. For each phenotype, the first occurrence of the diagnosis code was used for the analytical sample selection. We did not use any other criteria to match cases to controls. In our sensitivity analysis, we defined the phenotypes by 2 codes that were at least 30 days apart. For both sensitivity analyses, our indeterminate colitis subgroup was defined by participants who had both codes.

Electronic health record (EHR) data quality dimensions (DQD)

For the data quality dimensions (DQD) analysis, we selected OMOP CDM concepts related to risk factors or managing prevalent forms of CD or UC, such as diagnoses, imaging, medications, and laboratory measurements.1,11–19 The set of concepts can be found in the Supplementary Appendix. We used a previously developed DQD framework with 5 distinct dimensions demonstrated in 4 prior works.7–10 The data quality dimensions in our framework included conformance (adherence to data standards), completeness (availability of data), concordance (agreement of data), plausibility (believability of data), and temporality (validity and order of temporal data). All dimensions represented person counts within the cohort, except for conformance, representing the total number of the phenotype defining procedure codes. We used a DQD checklist to evaluate each dimension by 3 criteria: (1) concept selection, (2) internal verification, and (3) external validation.7

Statistical analysis

All analyses presented in this paper used the All of Us Controlled Tier Curated Data Repository (CDR) v8 release, which includes data from electronic health records (EHRs), surveys, biospecimens, wearables, and genomic sequencing for individuals enrolled between May 6, 2017, and October 1st, 2023, with a data cutoff date of November 15, 2023.20 All programming and statistical analyses were performed using Python version 3.7.12 and implemented in Jupyter Notebook version 6.5.4 within the All of Us Researcher Workbench. We utilized descriptive statistics and chi-squared statistics to test for independent association, Spearman coefficients to evaluate bivariate correlations, and data visualizations to explore the application of the DQDs. The significance level was set at P < .05.

The All of Us Institutional Review Board (IRB) has determined that data released to the Researcher Workbench is considered non-human subject research. All study participants gave informed consent. To protect the identity of our participants and follow the All of Us policy, we censored all sociodemographic category counts with fewer than 20 participants.

Results

Sample

In the All of Us v8 release, there were 633 547 participants. Of those, 393 596 (62%) were 18 years and older and had both survey and EHR data. Among the participants with EHR data, there were a total of 239 564 (60.9%) female and 149 867 (38.1%) male participants based on a self-reported question in the “The Basics” questionnaire.

There was a total of 8074 All of Us participants who had a UC or CD diagnosis. Of those, 3800 (47.1%) had UC codes only, 2973 (36.8%) participants had CD diagnoses only and 1301 (16.1%) had both. Within the 3800 participants who had a UC diagnosis, 1909 (50.2%) had multiple diagnosis codes greater than 30 days apart, 648 (17.1%) had multiple diagnosis codes within 30 days apart, and 1243 (32.7%) had exactly one diagnosis code. Within the 2973 participants with a CD diagnosis, 1990 (66.9%) had multiple CD diagnosis codes greater than 30 days apart, 318 (10.7%) had multiple CD diagnosis codes within 30 days apart, and 665 (22.4%) had exactly one CD diagnosis code. Within the 1301 participants with codes for both conditions, 1197 (92.0%) had multiple switches between CD and UC diagnoses. Of the 1197 participants with multiple switches, 361 (30.2%) started and ended with a UC diagnosis, 251 (21.0%) started with a UC diagnosis and ended with a CD diagnosis, 115 (9.6%) started with a CD diagnosis and ended with a UC diagnosis, and 470 (39.3%) started and ended with a CD diagnosis (Figure 1).

Figure 1.

For image description, please refer to the figure legend and surrounding text.

Flowchart of All of Us participants with inflammatory bowel disease diagnosis codes. Ulcerative colitis = UC; Crohn’s disease = CD; diagnosis = dx; days = d; records = rec; Exactly one diagnosis code = Single.

Table 1 provides a breakdown by sex as a biological factor, age at diagnosis, race and ethnicity, education, and income for the IBD phenotype.

Table 1.

Sociodemographic characteristics of All of Us inflammatory bowel disease (IBD) cohorts.

IBD Case No. (%) Non-IBD Control No. (%) P-value
Total 8704 (100) 385 522 (100)
Race/Ethnicitya P < .001
 White 5953 (68.3) 225 013 (58.4)
 Black 1044 (12.0) 74 734 (19.3)
 Hispanic 900 (10.3) 70 895 (18.4)
 Asian 154 (1.8) 14 621 (3.8)
 MENA 112 (1.3) 4149 (1.1)
 NHPI ≤20 1187 (0.3)
 AIAN 302 (3.5) 16 164 (4.2)
 Race ethnicity none of these 93 (1.1) 3949 (1.0)
 Skip/prefer not to answer 149 (1.7) 6847 (1.8)
Sex P = .8
 Female 4920 (61.0) 234 644 (60.9)
 Male 3073 (38.1) 146 794 (38.1)
 None of the above or Skip 75 (0.9) 3867 (1.0)
Education P < .001
 1 through 4 38 (0.5) 3113 (0.8)
 5 through 8 88 (1.1) 8462 (2.2)
 9 through 11 241 (3.0) 22 294 (5.8)
 12 or GED 1258 (15.6) 72 914 (18.9)
 College 1 to 3 2221 (27.5) 100 778 (26.1)
 College graduate 2112 (26.2) 87 228 (22.6)
 Advanced degree 1959 (24.3) 81 269 (21.1)
 Never attended or prefer not to answer 40 (0.5) 2803 (0.7)
 Skip 117 (1.5) 6658 (1.7)
Income (USD) P < .001
 Less 10k 680 (8.4) 49 398 (12.8)
 10k 25k 937 (11.6) 43 594 (11.3)
 25k 35k 568 (7.0) 27 093 (7.0)
 35k 50k 686 (8.5) 30 934 (8.0)
 50k 75k 1001 (12.4) 41 389 (10.7)
 75k 100k 714 (8.8) 32 237 (8.4)
 100k 150k 936 (11.6) 39 900 (10.4)
 150k 200k 455 (5.6) 18 614 (8.8)
 More 200k 642 (8.0) 25 470 (6.6)
Prefer Not to Answer 1012 (12.5) 51 590 (13.4)
 Skip 443 (5.5) 25 300 (6.6)
Age at consent P < .001
 18-39 1903 (23.6) 107 157 (27.8)
 40-59 2706 (33.5) 133 694 (34.7)
 60-79 3165 (39.2) 132 226 (34.3)
 80+ 297 (3.7) 12 301 (3.2)

Note:

a

More than one race or ethnicity category can be selected.

Abbreviations: American Indian/Alaska Native = AI/AN; Middle Eastern or North African = MENA; Native Hawaiian or Pacific Islander = NHPI; United States Dollars = USD.

Data Source: The All of Us Research Program.

In the IBD cohort, the proportion of white patients was greater than in the non-IBD controls (68.3% vs 58.4%, P < .001). There was no major difference in the distribution of sex between the groups (P = .8). There were differences in the distribution of income, education, and age at consent (P < .001 for all).

We calculated the distribution of age at diagnosis for the IBD cohort. There were 135 (1.7%) participants who were diagnosed between the ages of 0 to 17. For the participants who were diagnosed as adults, 2427 (30.1%) were between the ages of 18 to 39, 3042 (37.7%) between the ages of 40 to 59, 2333 (28.9%) between the ages of 60 to 79 and 137 (1.7%) who were at least 80 years old.

Geospatial distribution

Figure 2 shows the geospatial distribution of the cohort. Most of the data came from the following 9 states: Massachusetts, Pennsylvania, Arizona, Wisconsin, California, Illinois, New York, Michigan, and Florida. Counts for those states are shown in Table S1.

Figure 2.

Geospatial distribution of the Inflammatory Bowel Disease (IBD) cohort by state.

Geospatial analysis of the Inflammatory Bowel Disease (IBD) cohort. All of Us = AoU. Data Source: The All of Us Research Program. Due to disclosure risk guidelines, States with fewer than 20 participants were not reported.

Source to standard vocabulary conformance

Our conformance assessment included standards and values for internal verification. No external benchmarks were available to verify these analyses.

We calculated the distribution of all SNOMED CT, ICD10, Columbia International eHealth Laboratory (CIEL) and ICD-9 codes in our cohorts to measure conformance (Figure 3). In the IBD cohort, most source codes were ICD10 (64.9%).

Figure 3.

Sankey diagram that illustrates the number of participants who had different source and standard codes for Inflammatory Bowel Disease.

Sankey Diagram of the International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM), International Classification of Diseases, Tenth Revision, Systematized Nomenclature of Medicine—Clinical Terms (SNOMED), Columbia International eHealth Laboratory (CIEL) codes for Inflammatory Bowel Disease (IBD) cohorts from the source to standard vocabularies. Source Vocabulary = Source; Standard Vocabulary = Standard; Data Source: The All of Us Research Program.

Completeness

Our completeness assessment included concept frequency, combinatorial completeness, and matching completeness. No external benchmarks were available for comparison with these analyses.

Concept prevalence completeness

We compared the percentages of clinical concepts for IBD cases compared to controls (Figure 4). Those groups differed with respect to Colonoscopy with Biopsy: 48.2% vs 13.7%, Corticosteroids: 81.1% vs 50.6%, Immunomodulators: 19.4% vs 2.1%, Advanced Therapeutics: 21.0% vs 1.0% 5-Aminosalicylic Acid (5-ASA) Derivatives: 33.3% vs 0.6% (P < .001). The full set of concept prevalence and counts are in Table S2.

Figure 4.

Bar chart that shows the prevalence of disease specific concept sets for Inflammatory Bowel Disease cases and controls.

Bar charts of clinical concept set for participants with Inflammatory Bowel Disease (IBD) phenotype (blue) and those without (orange). Clostridium Difficile= C Diff; Computerized Tomography= CT; Magnetic Resonance Imaging = MRI; Data Source: The All of Us Research Program.

Combinatorial completeness

We calculated the most frequent combinations of clinical measurements and intervention concepts in each cohort to determine whether the data were sufficient to represent standard clinical practices. For the IBD cohort, no combination comprised more than 4.4% of the cohort. Corticosteroids and a UC diagnosis were the most frequent combination, corticosteroids and a CD diagnosis was the second most frequent combination, and a CD diagnosis only was the third most frequent combination (Table 2).

Table 2.

Combinatorial completeness of inflammatory bowel disease (IBD) cases.

IBD cases (no., %)
Combination 1 Corticosteroids and Ulcerative Colitis diagnosis (354, 4.4)
Combination 2 Corticosteroids and Crohn’s Disease diagnosis (239, 3.0)
Combination 3 Crohn’s Disease diagnosis only (223, 2.8)

Computerized Tomography = CT; Count = No.; Percent = %. Data Source: The All of Us Research Program.

Concordance

Our concordance assessment examined the bivariate correlation between selected concepts to assess similarity or agreement. No external benchmarks were available in comparison with these analyses.

To measure concordance, we calculated 253 bivariate correlations between OMOP CDM concepts for clinical concept sets in the IBD cohort. We then classified the ones with ⍴ > 0.5 because 0.5 is regarded as the threshold between fair and moderate correlations in medicine, and the ones with ⍴ ≤ 0.5 as “weak.”21

Of the 253 bivariate pairs in the IBD analysis, 78 (30.1%) were above the threshold (Figure 5).

Figure 5.

Correlogram that shows the strength of bivariate correlation pairs for medications, measurement and procedure concepts that were used in the Inflammatory Bowel Disease cohort analysis.

Correlogram of medications, measurements, and procedures used in the Inflammatory Bowel Disease (IBD) cohort analysis. Clostridium Difficile = C Diff; Computerized Tomography = CT; Magnetic Resonance Imaging = MRI; Data Source: The All of Us Research Program.

Plausibility

We used a combination of EHR and survey data to assess the plausibility of our cohort using age grouping and self-reported survey questionnaires.

The distribution of concepts was statistically significant across age groups (P < .001). Most concepts occurred before 80 years of age, which was a finding consistent with clinical expectations (Figure S1).

Drug ingredient frequency

We characterized the distributions of immunomodulator ingredient concept frequencies and advanced therapeutic drug classes to evaluate plausibility. The immunomodulator stratification found that there were 793 (9.1%) participants who used azathioprine, 630 (7.2%) who used methotrexate, and 412 (4.7%) who used mercaptopurine. The advanced therapeutics analysis stratification showed that there were 1474 (16.9%) of participants who used Tumor Necrosis Factor (TNF) inhibitors, 416 (4.8%) who used Ant-integrin Agents, 98 (1.1%) who used Janus Kinase Inhibitors, and fewer than 20 participants used Sphingosine-1-Phosphate Receptor (S1PR) Modulators. None of the participants in the IBD cohort used IL-12/23 inhibitors.

Temporality

We calculated the time interval between an IBD diagnosis and colonoscopy with biopsy. There was a total of 3804 (43.7%) participants in our cohort with biopsy data. We used the biopsy that was closest in proximity to the IBD diagnosis to calculate that interval. Participants who had an interval of 15 years or longer were outliers and were excluded. The median time interval was 4.43 [-0.05, 104.29] weeks.

To characterize the sequence of events, we evaluated the subset of the IBD cohort that contained all 4 of the following concept sets: (1) Colonoscopy with Biopsy, (2) Corticosteroids, (3) Disease specific therapy (eg, 5-ASA Derivatives, Immunomodulators, Advanced Therapeutics or Colorectal Surgery), and (4) Diagnosis. We plotted the chronology and spacing of each.

Colonoscopies with biopsies and diagnostic codes were selected because they are standard parts of the diagnostic workup. Corticosteroids, mesalamines, immunomodulators, and advanced therapies (biologics and small molecule inhibitors) were included since they are used to treat both UC and CD.

There were 2030 (23.3%) IBD participants who had all of those clinical concepts. Furthermore, we observed more than one pathway of concepts for each cohort. The time interval between each consecutive concept varied from 14 to 622 days. Furthermore, in one pathway the interval between IBD diagnosis and disease specific therapy was the shortest (14 days) and the intervals between colonoscopy with biopsy and corticosteroids were the longest (622 days) (Figure 6).

Figure 6.

Timeline analysis of the Inflammatory Bowel Disease cohort with representative concepts. For the subset cohorts where participants had all relevant concepts, these timelines showed the most common orderings and median temporal separation between elements. Each arrow color represents a different sequence.

Timeline analysis of the Inflammatory Bowel Disease (IBD) cohort with representative concepts. For the subset cohorts where participants had all relevant concepts, these timelines showed the most common orderings and median temporal separation between elements. Interquartile Range = IQR; Count= no; Days = d. Data Source: The All of Us Research Program.

Sensitivity analysis

We implemented our data quality dimensions analysis on subgroups of participants who had at least 2 UC diagnoses that were at least 30 days apart and at least 2 CD diagnoses that were at least 30 days apart. There were 1909 participants in the UC subgroup and 1990 participants in the CD subgroup. The distribution of sociodemographic variables in the UC and CD subgroups were similar to the overall IBD Cohort (Table S3).

We calculated the distribution of age at diagnosis for the UC and CD subgroups. The distribution of UC subgroup was similar to the IBD cohort. A greater proportion of the CD subgroup (38.2%) was diagnosed between the ages of 18-39 compared to the IBD cohort (30.1%). Otherwise, the age at diagnosis distribution was similar for the CD subgroup and IBD cohort (Table S4). The geospatial distributions of the UC and CD subgroups were similar to the IBD cohort (Figures S2 and S3).

In the conformance analysis, most of the UC and CD diagnosis codes had an ICD10 source (65.2% and 64.9%, respectively). In the completeness analysis, the differences between the subgroup and IBD cohort were less than 10% for most concepts. However, differences in concepts for disease specific workup or treatment were higher. Specifically, in the UC cohort, the clinical measure and intervention concepts with a greater than 5% difference were 5-ASA derivatives (18.8%), CT of the abdomen and pelvis (-7.5%), and MRI of the pelvis (-8.6%). In the CD cohort, the only clinical measure and intervention concepts with a greater than 5% difference were advanced therapeutics (10.7%), immunomodulators (7.8%), and MRI of the pelvis (9.7%). The full concept distribution is in Table S5. In the combinatorial completeness analysis, no combination in the UC or CD subgroups had a prevalence greater than 7%. The most frequent combinations in the subgroups were similar to the most frequent combinations in the IBD cohort (Table S6).

In the concordance analysis for the UC subgroup 83 (35.9%) concepts had bivariate correlations above our threshold and in the CD subgroup 96 (41.6%) concepts had bivariate correlations above that threshold (Figures S4 and S5). In the plausibility analysis, the distribution of concepts across age groups in the CD subgroup was similar to the IBD Cohort (Figures S6 and S7). Furthermore, we characterized the distributions of immunomodulator ingredient concept frequencies and advanced therapeutic drug classes in the UC and CD subgroups to evaluate plausibility. Those findings were consistent with clinical expectations (Table S7).

In the temporality analysis for the UC subgroup, the median time between diagnosis and biopsy was 17.69 [0 to 131.76] weeks and in the CD subgroup, the median time was 40.9 [0.07 to 185.9] weeks. For the 2041 participants who were not in either subgroup, the median time interval was 0 [-0.51 to 48.64] weeks. The relative position of colonoscopy with biopsy on the timeline was later for some pathways in the UC the CD subgroups compared to the overall IBD cohort (Figures S8 and S9).

Discussion

We have shown successful implementation of a DQD framework on an IBD phenotype. This generalizes our method to a new set of conditions than were examined in prior studies.6–9 We chose an IBD diagnosis code to be the main phenotype for this analysis in order to maximize its sensitivity and minimize bias. Furthermore, our results showed that alternative phenotyping approaches could have introduced bias from data missingness. Specifically, approximately 29% of the IBD cohort had exactly one diagnosis code and 13% had diagnosis codes for fewer than 30 days. Consequently, using a phenotype definition that consisted of multiple diagnosis codes that were more than 30 days apart could have reduced sensitivity and skewed the data. Alternative IBD phenotypes have used data from multiple domains (eg, diagnosis codes and visit types).22 However, that approach may have been limited by the amount of data available.

Each dimension provided valuable information about our data quality. The percentages of the cohort with diagnostic workup concepts and disease specific treatment concepts (eg, 5-ASA derivatives) were lower than expected. Furthermore, the 3 most frequent IBD combinations did not include disease specific diagnostic workup or treatment concepts. Specifically, none of those combinations included a colonoscopy or sigmoidoscopy, 5-ASA derivatives, immunomodulators, or advanced therapeutics. The concordance, plausibility, and temporality analyses were limited by the amount of data available. Records from outside hospitals and non-digitized EHR records may have contributed to the data missingness.

The bivariate correlations were performed on participants who had data on both concepts in the pair. There were many bivariate correlations that were above our threshold of weak concordance. Those results suggest that there were subgroups in the IBD cohort that had multiple interrelated concepts and counts that were sufficient for bivariate correlation calculations. Therefore, the correlation results provide indirect evidence of an IBD subgroup that has high data completeness.

In our plausibility analysis, the age distribution of IBD concepts was consistent with expectations for the subgroup of participants who had data on those concepts. We also found that the distribution of advanced therapeutics and immunomodulator drugs was consistent with clinical expectations. Indirectly, results from other parts of our study suggested that the data were plausible. Specifically, the age and sex distribution of our cohort was consistent with prior work.23 Also, our analysis found that switching from UC to CD diagnoses was more frequent than the reverse. Those findings were consistent with another study that reported switching from UC to CD diagnoses and investigated factors associated with diagnostic uncertainty.24 Furthermore, since UC is the more localized illness, switching from that to CD is more plausible than the reverse.

In the temporality analysis, we found that for the subgroup of participants who we analyzed, the IBD diagnosis tended to precede a biopsy. That finding was different from our prior studies which found that the workup or diagnosis tended to precede the diagnosis.3,6 Our concept timeline further showed that the diagnosis was the first concept for many participants. These findings suggest that diagnosis codes can be an inconsistent time anchor across diseases. Therefore, temporal characterizations may provide valuable information for phenotyping different conditions.

The conformance analysis was more independent of completeness than the other data quality domains. It showed that most of the source data for the IBD diagnoses were ICD-10 codes. Those findings imply that our data were recent.

Our sensitivity analysis had results that were similar to the IBD cohort analysis. For example, many of the completeness concepts in the subgroups had similar prevalence, except for disease specific workup and treatment concepts. Those results suggest that increasing the number of diagnosis codes in the phenotype definition had a small effect on the amount of data missingness. Incomplete linkage across data domains is a possible explanation for the small effect. For example, without robust data linkage, increasing the number of diagnosis codes in the phenotype definition may not change the amount of data in the procedure, measurement, or drug domains. Furthermore, the median time between diagnosis and biopsy was longer in the UC and CD subgroups compared to the IBD cohort, which suggests that the diagnosis and biopsy data may have been more uneven in the subgroups than in the main cohort. Alternatively, the fact that diagnoses and biopsies occurred on the same day for participants who were not in either subgroup suggests that the data may have been better linked for indeterminate cases.

Our study had the following limitations. First, there was no data set available for external comparison. Second, in the absence of direct source validation, we were unable to identify true cases, false negatives, or determine to what extent low completeness was due to variance in practice patterns vs data missingness. Consequently, our explanations for data missingness were limited. However, records that were non-digitized or out of network and suboptimal linkage of data from different sources may have contributed to missingness. A multi-institutional chart review that compared source to standard data could provide novel insights into reasons for data missingness. Additionally, there are ongoing federal initiatives to address data quality improvement such as Data Collect Once Use Numerous Times (Data COUNTs) and projects with the Center for Linkage and Acquisition of Data (CLAD). In the interim, a comprehensive and transparent characterization of data missingness can help IBD investigators design rigorous and reproducible studies with All of Us data.

Third, the uneven geospatial distribution of our cohort, and the non-random distribution of sociodemographic variables, may have contributed to selection bias. Fourth, misclassification of IBD cases and switching between CD and UC diagnoses may have contributed to bias in the sensitivity analyses and resulted in inexact prevalence estimates of UC, CD and indeterminate colitis. Fifth, our data had minimal contributions from unstructured data, which contains critical information about disease severity and distribution. Sixth, the data curation process can create a temporal lag. Therefore, the counts of newer therapies (eg, IL-12/23 inhibitors) may have been lower than expected.

This study was a unique data quality dimensions analysis of IBD. We also compared IBD subgroups, CD and UC, with the data quality framework. Our analysis may be informative to researchers who would investigate IBD with All of Us data. Also, our results can be used as a benchmark for IBD data quality at individual EHRs. Our methods may be generalized to evaluate other diseases, and their subgroups.

Conclusion

Our data quality framework was implemented on an IBD cohort to determine its fitness for use. Further research is necessary to improve data quality for those conditions within the OMOP CDM.

Supplementary Material

ooag084_Supplementary_Data

Acknowledgements

We gratefully acknowledge All of Us participants for their contributions, without whom this research would not have been possible. The authors have no acknowledgements, competing interests, or funding to disclose.

Contributor Information

Matthew Spotnitz, All of United States Research Program, National Institutes of Health, Bethesda, MD, United States.

Adam S Faye, Division of Gastroenterology & Hepatology, NYU Langone Health, NYU School of Medicine, New York, NY, United States.

John Giannini, All of United States Research Program, National Institutes of Health, Bethesda, MD, United States.

Tamara R Litwin, All of United States Research Program, National Institutes of Health, Bethesda, MD, United States.

Yechiam Ostchega, All of United States Research Program, National Institutes of Health, Bethesda, MD, United States.

Lew Berman, All of United States Research Program, National Institutes of Health, Bethesda, MD, United States.

Author contributions

Matthew Spotnitz (Conceptualization, Formal analysis, Investigation, Methodology, Writing—original draft, Writing—review & editing), Adam S. Faye (Conceptualization, Methodology, Supervision, Writing—review & editing), John Giannini (Conceptualization, Formal analysis, Investigation, Methodology, Software, Visualization, Writing—review & editing), Tamara R. Litwin (Conceptualization, Formal analysis, Investigation, Methodology, Writing—review & editing), Yechiam Ostchega (Conceptualization, Investigation, Methodology, Supervision, Visualization, Writing—review & editing), and Lew Berman (Conceptualization, Investigation, Supervision, Writing—review & editing)

Supplementary material

Supplementary material is available at [JAMIA Open] online.

Funding

None declared.

Conflicts of interest

None declared.

Ethics statement

The work described here was proposed by Consortium members and confirmed as meeting criteria for non-human subject research by the All of Us IRB. Results reported are in compliance with the All of Us Data and Statistics Dissemination Policy disallowing disclosure of group counts under 20.

Data availability

Data and code used in this study are available as a featured workspace to registered researchers of the All of Us Researcher Workbench.25

References

  • 1. M’Koma AE.  Inflammatory bowel disease: an expanding global health problem. Clin Med Insights Gastroenterol. 2013;6:33-47. 10.4137/CGast.S12731 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Denny JC, Rutter JL, Goldstein DB, et al.  The “All of Us” research program. N Engl J Med. 2019;381:668-676. 10.1056/NEJMsr1809937 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. All of Us Research Program. Data dictionaries. Accessed August 5, 2025. https://support.researchallofus.org/hc/en-us/articles/360033200232-Data-Dictionaries
  • 4. Feurstein JD, Cheifetz AS.  Crohn disease: epidemiology, diagnosis, and management. Mayo Clin Proc. 2017;92:1088-1103. 10.1016/j.mayocp.2017.04.010. [DOI] [PubMed] [Google Scholar]
  • 5. Wilkins T, Jarvis K, Patel J.  Diagnosis and management of Crohn’s disease. Am Fam Physician. 2011;84:1365-1375. [PubMed] [Google Scholar]
  • 6. Veauthier B, Hornecker JM.  Crohn’s disease: diagnosis and management. Am Fam Physician. 2018;98:661-669. [PubMed] [Google Scholar]
  • 7. Berman L, Ostchega Y, Giannini J, et al.  Application of a data quality framework to ductal carcinoma in situ using electronic health record data from the All of Us research program. JCO Clin Cancer Inform.  2024;8:e2400052. 10.1200/CCI.24.00052 [DOI] [PubMed] [Google Scholar]
  • 8. Spotnitz M, Giannini J, Ostchega Y, et al.  Assessing the data quality dimensions of partial and complete mastectomy cohorts in the All of Us research program: cross-sectional study. JMIR Cancer  2025;11:e59298. 10.2196/59298 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Spotnitz M, Giannini J, Clark E  et al.  Assessing the data quality dimensions of surgical oncology cohorts in the All of Us research program. JCO Clin Cancer Inform. 2025;9:e2500078. 10.1200/CCI-25-000 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Spotnitz M, Giannini J, Clark E, et al.  Assessing data quality of rheumatoid and psoriatic arthritis patients in the All of Us research program. JAMIA Open. 2026;9:ooag028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Hripcsak G, Duke JD, Shah NH, et al.  Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform. 2015;216:574-578. [PMC free article] [PubMed] [Google Scholar]
  • 12. Overhage JM, Ryan PB, Reich CG, et al.  Validation of a common data model for active safety surveillance research. J Am Med Inform Assoc. 2012;19:54-60. 10.1136/amiajnl-2011-000376 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Reich C, Ostropolets A, Ryan P, et al.  OHDSI standardized vocabularies-a large-scale centralized reference ontology for international data harmonization. J Am Med Inform Assoc. 2024;31:583-590. 10.1093/jamia/ocad247 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Le Berre C, Honap S, Peyrin-Biroulet L.  Ulcerative colitis. Lancet.  2023;402:571-584. 10.1016/S0140-6736(23)00966-2 [DOI] [PubMed] [Google Scholar]
  • 15. Meier J, Sturm A.  Current treatment of ulcerative colitis. World J Gastroenterol. 2011;17:3204-3212. 10.3748/wjg.v17.i27.3204 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Baumgart DC, Sandborn WJ.  Inflammatory bowel disease: clinical aspects and established and evolving therapies. Lancet. 2007;369:1641-1657. 10.1016/S0140-6736(07)60751-X [DOI] [PubMed] [Google Scholar]
  • 17. Baumgart DC.  The diagnosis and treatment of Crohn’s disease and ulcerative colitis. Dtsch Arztebl Int. 2009;106:123-133. 10.3238/arztebl.2009.0123 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Cushing DR, Higgins PDR.  Management of Crohn disease: a review. JAMA. 2021;325:69-80. 10.1001/jama.2020.18936 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. McGee DL, Liao Y, Cao G, Cooper RS.  Self-reported health status and mortality in a multiethnic US cohort. Am J Epidemiol. 1999;149:41-46. 10.1093/oxfordjournals.aje.a009725 [DOI] [PubMed] [Google Scholar]
  • 20. Langan RC, Gotsch PB, Krafczyk MA, Skillinge DD.  Ulcerative colitis: diagnosis and treatment. Am Fam Physician. 2007;76:1323-1330. [PubMed] [Google Scholar]
  • 21. Chan YH.  Biostatistics 104: correlational analysis. Singap Med J. 2003;44:614-619. [PubMed] [Google Scholar]
  • 22. Eun Y, Culpepper-Morgan J, Akanmode AM, et al.  Comorbidities and systemic steroids drive pneumonia risk in inflammatory bowel disease: propensity score-matched cohort study. World J Gastrointest Pharmacol Ther. 2025;16:105335. 10.4292/wjgpt.v16.i2.105335 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Kappelman MD, Rifas-Shiman SL, Kleinman K, et al.  The prevalence and geographic distribution of Crohn’s disease and ulcerative colitis in the United States. Clin Gastroenterol Hepatol. 2007;5:1424-1429. 10.1016/j.cgh.2007.07.012Epub 2007 Sep 29. [DOI] [PubMed] [Google Scholar]
  • 24. Melmed GY, Elashoff R, Chen GC, et al.  Predicting a change in diagnosis from ulcerative colitis to Crohn’s disease: a nested, case-control study. Clin Gastroenterol Hepatol. 2007;5:602-608; quiz 525. 10.1016/j.cgh.2007.02.015 [DOI] [PubMed] [Google Scholar]
  • 25. Welcome to the All of Us Research Hub. All of Us Research Program. Accessed 09-02-2025. https://www.researchallofus.org

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ooag084_Supplementary_Data

Data Availability Statement

Data and code used in this study are available as a featured workspace to registered researchers of the All of Us Researcher Workbench.25


Articles from JAMIA Open are provided here courtesy of Oxford University Press

RESOURCES