Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 May 16.
Published before final editing as: Clin Gastroenterol Hepatol. 2026 Feb 17:S1542-3565(26)00066-2. doi: 10.1016/j.cgh.2026.01.037

Automated phenotyping with AI predicts future advanced neoplasia risk in colitis-associated low-grade dysplasia

Brian Johnson 1, Hyrum Eddington 1, Misha Kabir 2, Samir Gupta 3,4,5, Shailja C Shah 3,4,5, Kit Curtius 1,3,5,*
PMCID: PMC13178804  NIHMSID: NIHMS2151068  PMID: 41713829

The ability to precisely determine future risk of advanced neoplasia (AN; high-grade dysplasia [HGD] and/or colorectal cancer [CRC]) in patients with ulcerative colitis (UC) and low-grade dysplasia (LGD) is a major unmet need. Given the uncertainty in existing prognostic data, current guidelines advise incorporating expert opinion into management decisions which can be challenging and imprecise1. In addition to supporting clinician decision-making, having quantitative estimates of cancer risk also increases patients’ ability to make informed shared decisions on surgery versus continued intensive surveillance following a LGD diagnosis2.

To address this gap, we used large language models (LLMs) to identify patients with UC-LGD undergoing colonoscopic surveillance in the Veterans Affairs (VA) national database and predicted their future AN risk. To create this cohort, we used open weight LLMs to accurately identify UC and dysplasia/cancer histology from individual-level EHR data, which we previously validated using VA data3. To apply inclusion/exclusion criteria and quantify risk factors, we used LLMs to extract colonoscopy- and lesion-specific variables including morphology, size, location, dysplasia grade, visible vs. “invisible”, completeness of resection of visible lesions, focality (multifocal vs. unifocal), colitis extent and severity, bowel preparation quality, and colonic landmark reached. These variables were derived from the colonoscopy report (e.g., landmark reached), the corresponding pathology report (e.g., dysplasia grade), or both (e.g., completeness of resection). Structured data sources (ICD and CPT codes, VA cancer registry) were integrated with LLM-based abstraction from >5 million free-text notes and pathology reports in our automated pipeline to complete 55,450 longitudinal UC patient clinical histories (Figure 1A).

Figure 1: AI-automated phenotyping pipeline and AN risk forecasting results for 2,939 patients.

Figure 1:

A. Pathology (Path) reports and clinical notes from VA Corporate Data Warehouse were fed to LLMs to create longitudinal patient histories with follow-up to the primary outcome, AN progression, or censoring. The LLM-based approach stratifies each patient based on presence/absence of 4 binary risk factors defined by UC-CaRE at index LGD4. B. Kaplan-Meier (KM) estimation demonstrated statistically significant differences in AN progression-free survival during follow-up from time of index LGD between the 5 risk groups (log-rank p < 0.0001). The risk for AN in the lowest versus highest predicted risk groups ranged from 2.5% (95% CI: 1.6%-3.4%) to 27.8% (95% CI: 0.4%-47.6%) at 5 years. C. Next, the AI pipeline forecasts future risk from index LGD. See Supplementary Methods for time-dependent equation incorporating: any LGD ≥ 1cm (hazard ratio [HR] 2.7, 95% CI 1.2 - 5.9), any incompletely resected or invisible LGD (HR 3.4, 95% CI 1.6 - 7.4), multifocal LGD (HR 2.9, 95% CI 1.3 - 6.2), moderate or severe inflammation (HR 3.1, 95% CI 1.5 - 6.7)4. AI-predicted AN risk (solid bars) aligns with KM data (crosshatch bars, including estimates based on KM when treating colectomy as a competing risk) in each predicted risk group at 1 year post-index LGD. D. The AI-predicted AN risk (black lines) closely matches the KM data (red) in 1,427 patients in the lowest predicted risk group. The AN risk at 10 years is accurately predicted to be 5.16% in these patients.

Veteran patients with UC and an index LGD diagnosis between 1999-2024 were eligible. Index LGD was defined as the patient’s first LGD diagnosis in a VA pathology report located at or distal to the extent of colitis as determined by current and prior colonoscopy and pathology reports. Patients were excluded if they had prior or concurrent colectomy or AN, absence of a linked colonoscopy report, unknown date of CRC, or unknown completeness of LGD resection. Patients were considered at risk of progression from date of index LGD until the earliest date of the primary outcome (AN) or censoring (i.e., colectomy, last available clinical note, or death).

This final study cohort included 2,939 UC-LGD patients with 20,279 total patient-years of follow-up to evaluate risk, which is >6x larger than previous studies predicting AN risk in UC-LGD. Of these patients, 209 (7.1%) progressed to AN (see cohort details, Supplementary Table S1). We followed the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) + AI7 guidelines.

For AN risk prediction, our AI pipeline derived the 4 clinico-pathological variables included in our previous UC-CaRE multivariate model4 for each patient: any LGD ≥ 1cm, any incompletely resected or invisible LGD, multifocal LGD, moderate or severe inflammation. We then applied the corresponding hazard ratios in the UC-CaRE risk forecasting model (webtool available: www.uc-care.uk). We performed additional manual chart review for the patient-specific UC-CaRE variables versus LLM-derived answers and confirmed all had 88-93% accuracy. See Supplementary Methods for details. We compared predicted versus observed risk over time in patients grouped by number of UC-CaRE risk factors present at index LGD (5 risk groups total).

We first validated predicted risk group stratification for future AN development using Kaplan-Meier (KM) survival analysis (p<0.0001, Figure 1B, see Supplementary Figure S1A for data 20+ years post-LGD). This external validation of UC-CaRE independently reproduces expected separation of risk groups from prior UK validation cohorts4, reinforcing the fidelity of the LLM-based variable extraction and the generalizability of the risk model to U.S. patients. Furthermore, unresectable visible LGD conferred a 5-year cumulative risk of AN of 37.4% (95% CI 21.7%-50%), which is consistent with our findings for UK patients4 and nearly twice as high as the perceived risk of these lesions by clinicians in an international survey5.

We next forecasted risk into future years and found accurate AN risk prediction at 1 year post-index LGD for the 5 risk groups when comparing to KM and competing risks estimates (Figure 1C). There was good model calibration when predicting up to 10 years (Supplementary Figure S1B), with accurate risk forecasted >10 years in the lowest predicted risk group (Figure 1D). Importantly, this lowest risk group comprised nearly half of the UC-LGD population, and our AI pipeline accurately predicted that 98.9% of these individuals would not develop AN in 24 months. Our risk model predictions and observed outcomes may warrant revisiting current IBD dysplasia surveillance recommendations6.

To our knowledge, this is the first development and successful application of a fully AI-automated pipeline for GI cancer risk prediction and the largest external validation of time-dependent cancer risk prediction in GI premalignancy. The LLM-derived risk groups are grounded in biological plausibility and previously validated evidence-based risk models rather than ‘black-box’ algorithms, representing a major strength. Study limitations include the retrospective design and validation in a single national database of Veterans. However, rigorous validation in other U.S. healthcare systems may be challenging given that large tertiary centers, which have the most comprehensive datasets, often do not have access to index LGD information and may be missing complete CRC outcome information. The demographics of our VA cohort (96% male, median age 65) also differ from standard UC surveillance populations, potentially impacting generalizability. That said, we still achieved accurate risk prediction using only the clinico-pathological variables originally included in UC-CaRE.

Feasibility for using LLMs to generate guideline-based surveillance recommendations from EHR data input has been demonstrated in the non-IBD population8. In IBD, a major contributor to post-colonoscopy CRC (PCCRC; interval cancers) is delay in timely colonoscopic surveillance due to patient-, clinician-, and system factors6. One potential benefit of this validated AI approach is automatic patient risk stratification and recall for surveillance with minimal human involvement, which may be more efficient and reduce PCCRC. In conclusion, our automated AI pipeline reads in medical notes and provides accurate cancer risk predictions in UC-LGD to aid decision-making. Advances in LLMs and prompt design will enable further improvements in data-driven clinical decisions based on quantitative forecasts of future risk.

Supplementary Material

1

Funding:

This research was supported by a Merit Review Award I01 BX005958 from the United States (U.S.) Department of Veterans Affairs Biomedical Laboratory Research and Development Service. The contents do not represent the views of the U.S. Department of Veterans Affairs or the United States Government. This work was also supported by AGA Research Foundation (AGA Research Scholar Award AGA2022-13-05), NIH grants (R01 CA270235, P30 CA023100), National Library of Medicine Training Grant (NIH grant T15LM011271), and in part by the NIDDK-funded San Diego Digestive Diseases Research Center (P30 DK120515).

Footnotes

Conflict of interest statement: SCS is a paid ad hoc consultant for RedHill Biopharma and Phathom Pharmaceuticals, and unpaid scientific advisory board member for Ilico Genetics, Inc. The other authors do not declare any competing interests.

Publisher's Disclaimer: This is a PDF of an article that has undergone enhancements after acceptance, such as the addition of a cover page and metadata, and formatting for readability. This version will undergo additional copyediting, typesetting and review before it is published in its final form. As such, this version is no longer the Accepted Manuscript, but it is not yet the definitive Version of Record; we are providing this early version to give early visibility of the article. Please note that Elsevier’s sharing policy for the Published Journal Article applies to this version, see: https://www.elsevier.com/about/policies-and-standards/sharing#4-published-journal-article. Please also note that, during the production process, errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES