Abstract
Nonalcoholic fatty liver disease (NAFLD) is the most common global cause of chronic liver disease and remains under‐recognized within healthcare systems. Therapeutic interventions are rapidly advancing for its inflammatory phenotype, nonalcoholic steatohepatitis (NASH) at all stages of disease. Diagnosis codes alone fail to recognize and stratify at‐risk patients accurately. Our work aims to rapidly identify NAFLD patients within large electronic health record (EHR) databases for automated stratification and targeted intervention based on clinically relevant phenotypes. We present a rule‐based phenotyping algorithm for efficient identification of NAFLD patients developed using EHRs from 6.4 million patients at Columbia University Irving Medical Center (CUIMC) and validated at two independent healthcare centers. The algorithm uses the Observational Medical Outcomes Partnership (OMOP) Common Data Model and queries structured and unstructured data elements, including diagnosis codes, laboratory measurements, and radiology and pathology modalities. Our approach identified 16,006 CUIMC NAFLD patients, 10,753 (67%) previously unidentifiable by NAFLD diagnosis codes. Fibrosis scoring on patients without histology identified 943 subjects with scores indicative of advanced fibrosis (FIB‐4, APRI, NAFLD–FS). The algorithm was validated at two independent healthcare systems, University of Pennsylvania Health System (UPHS) and Vanderbilt Medical Center (VUMC), where 20,779 and 19,575 NAFLD patients were identified, respectively. Clinical chart review identified a high positive predictive value (PPV) across all healthcare systems: 91% at CUIMC, 75% at UPHS, and 85% at VUMC, and a sensitivity of 79.6%. Our rule‐based algorithm provides an accurate, automated approach for rapidly identifying, stratifying, and sub‐phenotyping NAFLD patients within a large EHR system.

Abbreviations
- APRI
Aspartate transaminase to Platelet Ratio Index
- CDM
Common Data Model
- CUIMC
Columbia University Irving Medical Center
- DOB
date of birth
- EHR
Electronic Health Record
- FIB‐4
fibrosis‐4
- HbA1c
glycated hemoglobin
- MRN
medical record number
- NAFLD
nonalcoholic fatty liver disease
- NAFLD‐FS
NAFLD Fibrosis Score
- NASH
nonalcoholic steatohepatitis
- NIT
noninvasive test
- OHDSI
Observational Health Data Sciences and Informatics
- OMOP
Observational Medical Outcomes Partnership
- SQL
Structured Query Language
- T2D
type 2 diabetes
- UPHS
University of Pennsylvania Health System
- VUMC
Vanderbilt University Medical Center
Study Highlights.
WHAT IS THE CURRENT KNOWLEDGE ON THIS TOPIC?
NAFLD is the leading form of chronic liver disease with a rising prevalence in the population. Its health impacts include cardiometabolic outcomes independent of liver disease. NAFLD is under‐recognized and under‐treated in affected individuals. The current means of identification and stratification are complex and dependent on provider recognition of clinical risk factors.
WHAT QUESTION DID THIS STUDY ADDRESS?
This study introduces an automated, machine‐based approach for rapid, scalable discovery and stratification of NAFLD patients.
WHAT DOES THIS STUDY ADD TO OUR KNOWLEDGE?
This study demonstrates the successful application of a machine‐based approach to rapidly discover and classify NAFLD patients using EHR data. Our algorithm has high performance (mean PPV = 85%, sensitivity = 79.6%) in NAFLD identification and stratification for fibrosis outcompeting current non‐invasive tests. Further, the majority of algorithmically derived NAFLD patients were previously unidentified within healthcare systems.
HOW MIGHT THIS CHANGE TRANSLATIONAL SCIENCE?
Our algorithmic approach addresses the long‐standing challenge of under‐diagnosis in NAFLD by augmenting human‐based approaches through machine‐based methods for rapid identification of metabolic disease cohorts.
INTRODUCTION
NAFLD is the most common form of chronic liver disease worldwide affecting 25%–30% of the general adult population in industrialized countries. 1 The nomenclature for NAFLD and NASH (Non‐Alcoholic Steatohepatitis) was recently updated to MASLD (Metabolic Dysfunction Associated Steatotic Liver Disease) and MASH (Metabolic Dysfunction Associated Steatohepatitis), respectively. 2 The term NAFLD is used throughout this work to align with the clinical guidelines in use at the time of algorithm development. NAFLD, along with its inflammatory phenotype NASH, is often underdiagnosed 3 due to the cost and invasiveness of liver biopsy, the current gold standard of diagnosis. The recent first drug approval in NASH 4 has underscored the critical importance of identifying and stratifying patients. The impact of the disease ranges from diabetes prevention and treatment to targeted diagnostics, specialist referral, cancer screening, genomic analyses, and intervention for longitudinal assessment and follow‐up. Early recognition is particularly important given disease model projections of a doubling or tripling of end‐stage liver disease patients by 2030 in many parts of the world. 5 Prioritizing interventions at appropriate stages may prevent progression in high‐risk individuals, such as those exhibiting NASH and advanced fibrosis with associated downstream consequences, such as liver transplantation. 6
Emerging and approved therapies for NASH will have limited patient benefit if at‐risk individuals remain difficult to identify in healthcare systems. Given the current limitations inherent in diagnostic coding for this disease, the rapid identification of patients with NAFLD is problematic. Advances in circulating blood biomarkers and imaging biomarkers can assist in risk stratifying patients with identified NAFLD, particularly those with advanced fibrosis; however, these are deployed on a per‐patient basis. Electronic health record (EHR) phenotyping is a rapidly scalable means by which patients can be targeted for diagnosis and risk stratification. EHRs are collected prospectively in a large‐scale, longitudinal manner 7 and can provide the data needed to phenotypically identify and stratify patients. These properties, along with the inclusion of diverse aspects of patients' health‐related information, make EHRs a valuable data source for phenotype discovery. EHRs are limited by the completeness and accuracy of data, which may have confounding effects if not properly addressed in the study design. 8 , 9 , 10 One approach to addressing these inaccuracies is to use a wide range of different data sources available in the EHR (including structured data, such as diagnosis/billing codes and laboratory measures, as well as unstructured elements such as imaging/radiology reports and provider notes) as a means of diagnostic confirmation. Additionally, quality control parameters can be implemented to reduce false‐positive identifications.
Herein, we describe a rule‐based phenotype algorithm developed at Columbia University Irving Medical Center (CUIMC), which expands on earlier work 11 , 12 , 13 , 14 by utilizing a multitude of EHR data sources (structured and unstructured) to identify and stratify NAFLD and NASH patients for clinical intervention. The algorithm queries over 400 diagnosis codes, 100 laboratory and serology measurements, pathology, and various radiology modalities. To demonstrate cross‐institutional utility and performance, we validate the algorithm at two large independent medical centers. We also perform fibrosis scoring tests on all CUIMC identified NAFLD patients without histologically confirmed NASH, identifying patients at highest risk for progressing to end‐stage liver outcomes and demonstrating the feasibility of rapid risk stratification techniques. As this algorithm was developed using the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), it can easily be deployed at healthcare institutions that support the OMOP CDM, which is presently over 90 sites worldwide, including the National Institutes of Health All of Us Research Program. 15
METHODS
Our NAFLD algorithm was developed using EHR data within the CUIMC healthcare center. CUIMC serves the diverse population of New York City, and is composed of approximately 38% Hispanic patients, 37% European American, 21% African American, and 4% other ethnicities. Observational Health Data Sciences and Informatics (OHDSI) is an international initiative with over 3000 collaborators focused on improving the use of healthcare data. 16 , 17 OHDSI maintains the OMOP CDM, 18 a data standard that normalizes observational data to enable efficient, reproducible analyses. CUIMC is a participant of OHDSI and maintains an OMOP database of primary and ancillary EHR systems, pharmacy, and billing systems. The CUIMC clinical data warehouse (CDW) encompasses this OMOP database and other clinical data distributed for research purposes, such as provider notes, unstructured elements, and imaging data. 19 At the time of this study, the CDW contained records for 6.4 million patients dating back to 1985. Figure 1 provides an illustrative depiction of the NAFLD algorithm development and validation process.
FIGURE 1.

Illustration of the NAFLD algorithm development and validation process. The algorithm was developed at Columbia University Irving Medical Center (CUIMC) by clinical and informatics teams. Clinical criteria for the algorithm, provided by medical experts, was used by the bioinformatics team to design queries that produced a set of potential NAFLD patients. The charts of these patients were reviewed by the clinical team to determine true NAFLD status. Clinical criteria, as represented in the EHR system, was adjusted based on chart review results, and the queries were refined. This process was repeated across each step of the algorithm until high accuracy was achieved. Once achieved, the queries were used to code the algorithm. Algorithmic validation was performed at the University of Pennsylvania Health System (UPHS) and Vanderbilt Medical Center (VUMC) where the iterative process described above was repeated by clinical and informatics experts at each site. The final output of the algorithm is a list of NAFLD patients along with clinical characteristics (subset depicted above). A1c, glycated hemoglobin; DOB, date of birth; EHR, electronic health record; N, no; T2D, type 2 diabetes; Y, yes. Stock images for this figure are from BioRender.com.
Code for the NAFLD algorithm predominantly consists of SQL queries of the structured OMOP database coupled with unstructured data parsing of notes from pathology and radiology reports. The algorithm has three main steps which each flow consecutively (Figure 2): (1). inclusion of potential NAFLD patients, (2). removal of non‐NAFLD patients meeting stringent exclusion criteria, and (3). verification of hepatic steatosis. Please see the Data S1 for the “Extended Methods” section.
FIGURE 2.

Three main steps of the NAFLD algorithm. Step 1a lists the NAFLD risk indicator categories that were used for patient identification and lists a few examples of selection criteria. A complete list of codes for selection or exclusion criteria can be found in the following Supplementary Tables: Step 1a (Table S1), Step 1b (Table S2), Step 2 (Table S3), and Step 3 (Tables S4 and S5). dx = diagnosis code (International Classification of Diseases, Ninth and Tenth Revision (ICD9/10)).
Step 1: Identification of NAFLD patients
Step 1 identifies NAFLD patients by the presence of a NAFLD risk indicator and/or a NAFLD diagnosis (see Tables S1 and S2 for codes). NAFLD risk indicators include type 2 diabetes and dysglycemia (Table S1a), obesity (Table S1b), abnormal liver enzymes (Table S1c), hyperlipidemia (Table S1d), or hypertension (Table S1e).
Step 2: Exclusion of patients with confounding diagnoses
Cases meeting specified exclusion criteria (Table S3) were removed in Step 2. Exclusion criteria include excessive alcohol use, confounding liver or extrahepatic conditions that may result in secondary hepatic steatosis, including viral hepatitis, type 1 diabetes, and others. Detailed exclusion, as often undertaken for clinical trial enrollment, was codified. Patients prescribed a hepatotoxic medication associated with steatosis, 20 such as an anti‐retroviral, were also excluded.
Step 3: Verification of hepatic steatosis
Radiology and pathology reports from 1980 to 2019 were used to verify hepatic steatosis. Regular expressions, a powerful pattern search language, and tool 21 were used in conjunction with key terms to identify language and usage context indicative of hepatic steatosis in a string‐matching approach. Language for an indicator of NAFLD and NASH were included (Tables S4 and S5).
Fibrosis scoring
To identify patients at risk for fibrotic NAFLD, we applied three common fibrosis scoring metrics on patients lacking histology: Fibrosis‐4 (FIB‐4), 22 aspartate transaminase to Platelet Ratio Index (APRI), 23 and the NAFLD Fibrosis score 24 calculation (Equation S1–Equation S3). Data for these calculations were extracted from patient clinical records at cross‐sectional timepoints. To ‘rule in’ advanced fibrosis, we required patients to exhibit an elevated score in at least two metrics (FIB‐4 > 3.25, an APRI >1.0, and a NAFLD FS >0.675).
Quality control
To minimize EHR diagnosis errors, we employed quality control (QC) measures requiring patients to have ≥2 NAFLD risk indicators, a risk indicator and a NAFLD diagnosis, or ≥3 unique occurrences of a single given NAFLD risk indicator. The cohort was restricted to patients 18 or older (using the earliest date of hepatic steatosis confirmation).
Algorithm validation
The algorithm was validated at two external, independent healthcare systems, University of Pennsylvania Health System (UPHS) in Philadelphia, Pennsylvania and Vanderbilt University Medical Center (VUMC) in Nashville, Tennessee. UPHS maintains a data warehouse for translational research that combines data (discrete and unstructured text reports) from five hospitals in the greater Philadelphia area and one in Princeton, New Jersey. The CDW contains data for 3 million patients dating back to 2005. This population is 64% European American, 24% African American, and 12% other ethnicities. At the time of this study, UPHS used an Epic EHR platform, and validation was performed using codes and terminologies found in the Supplement. VUMC maintains a de‐identified data warehouse dating to 1990 that contains both structured and unstructured data and describes 3.2 million patients, of which 82% are European American, 13% African American, 4% Hispanic, and 1% other ethnicities. VUMC maintains an OMOP database and validation at this site was performed using data standardized to the OMOP CDM.
Chart review and algorithmic performance
Manual, retrospective chart review was performed to review data elements for cohort construction and to assess algorithmic accuracy. This clinical review served as a critical component of algorithm development, allowing us to evaluate the efficacy of adding or removing selection criteria across the algorithm. It allowed us to adjust NAFLD risk indicator criteria and keyword terminology indicative of hepatic steatosis. Over 150 MRNs were reviewed at CUIMC during development by two clinical research coordinators and verified by a board‐certified transplant hepatologist. Chart review at VUMC and UPHS was similarly performed by board‐certified transplant hepatologists.
Chart review was also used in the calculation of algorithmic positive predictive value (PPV) at each of the assessed sites. PPV is the proportion of patients identified by the phenotyping algorithm as having the condition, determined by expert chart review. We reviewed the charts of 200 patients, independent of those assessed for diagnostic code selection, to calculate PPV at CUIMC. Hepatologists at VUMC and UPHS reviewed 20 charts for PPV calculation. Overall, 390 clinical charts were reviewed during algorithm development and validation.
We performed a sensitivity analysis to determine our algorithm's ability to identify true NAFLD patients using a NAFLD registry maintained by CUIMC hepatology with records from 2006 to 2019. We define sensitivity as the number of patients within this registry who were correctly identified by our algorithm. We assessed sensitivity across three different categories: using all patients within this registry, restricting to patients for whom there was complete data (structured and unstructured) within the CDW, and omitting patients for whom exclusion diagnosis codes could not be physician‐verified. Sensitivity analyses were performed solely at CUIMC as we had complete data access within this institution allowing us to thoroughly query medical records and robustly interact with the EHR.
Ethics statement
This work was performed under IRB approval with a waiver of consent (Protocol AAAL0601).
RESULTS
Our algorithm identified 16,006 NAFLD patients with verified hepatic steatosis at CUIMC with 67% of this cohort previously undiagnosed by NAFLD codes (Figure 3). Over 40% of these patients self‐identified as Black, Hispanic, or Asian (Table 1). The algorithm identified 20,779 NAFLD UPHS patients and 19,575 VUMC patients (Figure 3). Most patients initially meeting the inclusion criteria were dropped from the algorithm during the verification of hepatic steatosis. Of the total potential NAFLD patients, 3.2% at CUIMC, 3.1% at UPHS, and 7.5% of patients at VUMC had algorithmic verified steatosis indicative of NAFLD. This large drop in sample size was primarily due to a lack of available imaging or biopsy data within each system's CDW. All 16,006 NAFLD patients identified at CUIMC with a NAFLD diagnosis code were also diagnosed with a risk indicator (e.g., abnormal liver enzymes, obesity) (Table S6).
FIGURE 3.

Counts of patients at each stage of the algorithm. Data from Columbia University Irving Medical Center (CUIMC) is in blue, that from University of Pennsylvania Healthcare System (UPHS) is in pink/purple, and numbers from Vanderbilt Medical Center (VUMC) are in orange.
TABLE 1.
Demography and summary information for identified NAFLD patients.
| CUIMC | UPHS | VUMC | |
|---|---|---|---|
|
Mean age at diagnosis (+/− Standard Deviation) |
57.4 (+/− 15.7) | 57.0 (+/− 15.4) | 52.3 (+/− 16.8) |
| Biological sex | F = 55.7% M = 44.3% |
F = 44.2% M = 55.8% |
F = 44.4% M = 55.6% U = 0.01% |
| Type 2 diabetes | 53.3% | 31.0% | 24.9% |
| Obesity | 45.2% | 60% | 48.8% |
| Percent race/Ethnicity | |||
| White | 33.9% | 68.4% | 79.3% |
| Black | 9.0% | 16.1% | 12.5% |
| Hispanic | 31.1% |
White = 4.8% Black = 1.2% |
3.7% |
| Asian | 1.7% | 2.7% | 0.9% |
| Other | 0.2% | 4.0% | 1.1% |
| Unknown | 24.1% | 2.9% | 6.0% |
Note: Age of diagnosis is based on the earliest date of verified hepatic steatosis. Mean age is noted with standard deviation in parentheses. Race/ethnicity values may not aggregate to 100% as some sites code race and ethnicity separately.
Abbreviations: F, female; M, male; U, unknown/undeclared.
Previous studies have shown that using a diagnostic combination of obesity, type 2 diabetes (T2D), and abnormal ALT values predicts NAFLD with high accuracy. 25 Using a simpler algorithm consisting only of these three diagnoses, we identified 2107 NAFLD patients at CUIMC, missing 86.9% of the patients discovered by our more complex algorithm. Additionally, 38.8% of these 2107 patients met at least one exclusion criteria included in our algorithm (e.g., alcohol abuse, viral hepatitis infection). Thus, the combination of obesity, DM2, and abnormal ALT is helpful in identifying patients at risk for NAFLD but may not necessarily yield true NAFLD patients. The additive value of a phenotyping algorithm of the complexity presented here is in the recognition of patients for whom elevated ALT, obesity, and DM2 were not observed. In diverse populations, these traditional risk factors may exclude at‐risk groups (i.e., lean NAFLD).
Fibrosis scores
Of the 16,006 NAFLD cohort at CUIMC, we identified 356 patients with a biopsy‐proven NASH by querying pathology reports. We performed fibrosis calculations on the remaining 15,650 patients lacking histology using the FIB‐4, APRI, and NAFLD fibrosis score noninvasive tests (NITs). We identified 943 patients with scores suggestive of advanced fibrosis, as indicated by an elevated score in ≥2 metrics. 204 patients have advanced fibrosis indicators across all three tests and 2245 patients have a high score in at least one metric. Of the patients with advanced fibrosis in ≥2 NITs, 56.5% are diagnosed with DM2 and/or obesity, 46.8% have abnormal liver enzyme levels, 57% are diagnosed with hyperlipidemia, and 78.6% with hypertension (Figure 4). 92.6% of patients with elevated scores in all three calculations have a DM2, obesity, or abnormal liver enzyme diagnosis. The FIB‐4 and APRI calculations were most concordant and together identified 778 (82.5%) patients with elevated NITs in ≥2 tests. Of the discrepancies between the FIB‐4 and APRI scores, 151 patients had FIB‐4 > 3.25 and APRI < 1.0; and 14 had APRI > 1.0 and FIB‐4 < 3.25. We identified 355 patients with a score suggestive of advanced fibrosis using the FIB‐4 and NAFLD‐FS calculators, and 219 patients with high APRI and NAFLD‐FS.
FIGURE 4.

Proportion of patients with select risk indicator diagnoses of those with 1 (n = 2245), 2 (n = 943), or 3 (n = 204) elevated scores (APRI, NAFLD FS, FIB‐4).
PPV and sensitivity
Chart review of 200 random patients performed by clinical experts at CUIMC identified 182 individuals correctly discovered by the algorithm as having NAFLD, a PPV of 91%. At the validation sites, our algorithm correctly discovered 15 NAFLD patients (PPV = 75%) at UPHS and 17 patients (PPV = 85%) at VUMC from a total of 20 patients per site. Algorithmic sensitivity was assessed at CUIMC using clinically verified NAFLD patients within a registry. When considering the 147 patients with complete data, our NAFLD algorithm attains a sensitivity of 79.6%, identifying 117 patients. The algorithm identifies 147 patients following Step 1 (100% sensitivity), 123 after Step 2 (83.7% sensitivity), and 117 following Step 3 (79.6% sensitivity) (Figure 5). Ten of the 24 patients identified as meeting algorithmic exclusion criteria (Step 2) had exclusion codes that could not be clinically verified during chart review suggesting inconsistencies between the CDW and the clinician‐facing system. Our algorithm attains a sensitivity of 85.4% if we drop these 10 patients. When considering all patients within the registry, even those who do not have complete data elements (and are inaccessible to the algorithm), we achieve a sensitivity of 68.3%.
FIGURE 5.

Sensitivity at Columbia University Irving Medical Center (CUIMC) after each stage of the algorithm. Sensitivity was assessed across three categories: (1). using all patients within the NAFLD registry maintained by CUIMC hepatology (blue, n = 167), (2). restricting to patients with complete data within the clinical data warehouse (red, n = 147), and (3). restricting to patients with physician‐validated exclusion codes (orange, n = 137). “Inclusion” refers to the identification of NAFLD patients. “Exclusion” is the removal of patients meeting exclusion criteria. “Verification” refers to the verification of hepatic steatosis, the final step of our algorithm (see Figure 2).
DISCUSSION
Identification and stratification of patients with under‐recognized metabolic disease is critical in addressing the NAFLD public health crisis. Reliance on frontline healthcare recognition may contribute to delays in treatment, monitoring, and referrals for specialty care. Beyond education, awareness campaigns, and resources for disease recognition, additional tools are necessary for effective point‐of‐care intervention. Systemic improvements through EHR‐based methods help streamline the identification of potential disease in patient populations, and with effective linkage to care, may improve access to advanced therapies while unburdening clinical providers. Healthcare systems with interrogatable longitudinal data provide means of discovering patients at risk for NAFLD progression and other disease manifestations. 26 Our work introduces an algorithm that successfully identifies NAFLD and NASH patients across three diverse healthcare systems. 67% of the NAFLD patients discovered at CUIMC were previously undiagnosed by ICD‐9/10 codes alone. Our machine‐based approach first identifies patients at risk of NAFLD using diagnoses and risk indicators, then excludes patients with confounding diagnoses, and finally verifies hepatic steatosis. Institutions supporting the OMOP CDM can implement this approach to identify at‐risk patients for downstream clinical referrals and investigation. Institutions without OMOP can still use the detailed workflow provided with diagnostic codes and search terms for pathology and radiology modalities to assist inpatient identification (as was performed at the UPHS validation site). Our rule‐based algorithm represents the first stage of development in machine learning and AI‐driven methods aimed at predicting, stratifying, and referring for intervention and linkages to care in NAFLD.
Our algorithm aggregates large amounts of clinical data and uses imaging or histologic components to verify hepatic steatosis. The algorithm exhibits a high PPV of 91% at CUIMC, 85% at VUMC, and 75% at UPHS showing generalizability. It also incorporates QC measures to reduce the rate of false positives. We found that patients with only one diagnostic code of NAFLD or a risk indicator were predominantly not true NAFLD patients. QC steps requiring a minimum number of unique diagnoses were employed to remove these patients. The algorithm was designed to prioritize PPV so that patients who truly have NAFLD are selected, reducing the false‐positive rate. This comes with the limitation that not all NAFLD patients will be included, as is reflected by the algorithm's sensitivity of 79.6%. Our sensitivity analysis highlights circumstances that healthcare centers will need to optimize when applying the algorithm to their patient population. For example, we identified 10 patients meeting exclusion criteria for whom exclusion codes could not be verified, highlighting a disconnect between data within the CDW and the clinician‐facing system. Additionally, missing data elements, particularly radiology and histology reports will affect performance but can be overcome with targeted approaches validated for missing data. 12% of NAFLD patients within the CUIMC hepatology registry lacked imaging data within the CDW. Further investigation identified that these patients had imaging performed outside of the healthcare system and were therefore missed by the current iteration of the algorithm. EHRs with centralized data should not have this obstacle. Another limitation of the study is the moderate sample size of patients used for manual chart review at the external validation sites.
Given the explosion of biomarkers in the NASH space, we expect future algorithmic iteration to incorporate additional features. Future directions also include applications to large cohorts such as the All of Us Research Program and integration of genomic sequencing of extreme phenotypes. Iterations of this algorithm may be used to identify rare diseases that resemble NAFLD, such as familial hyperlipidemias, sub‐phenotyping of lean NAFLD for further risk stratification and genomic analysis, and cohort analysis for extrahepatic comorbidities (heart failure, chronic kidney disease). Perhaps the most timely, iterative processing of noninvasive scoring systems (FIB‐4, NAFLD‐FS, APRI) to monitor longitudinal progression in fibrosis will help identify groups for determinants of rapid progression through machine learning methods validated across populations. As circulating and imaging (e.g., elastography) biomarkers for fibrosis become more available, future iterations of this algorithm may identify specific fibrosis stages for clinical trial screening, sub‐phenotype analyses, cancer screening, metabolic weight loss, and bariatric surgery referrals or liver transplant referrals. As clinical guidelines 27 for NAFLD are updated, components may be quickly implemented with machine‐based approaches rather than relying on community adoption and provider education alone. This study lays the groundwork for the development of a human‐centered and machine‐augmented collaborative framework for healthcare delivery in NAFLD. Future research should focus on optimizing this framework and exploring its generalizability to other diseases, paving the way for a more efficient and effective translation of knowledge into improved patient care.
AUTHOR CONTRIBUTIONS
A.O.B., A.A.‐O., M.P.R., M.D.R., and J.W. wrote the manuscript; G.R., N.P.T., and J.W. designed the research; A.O.B., L.A.T., A.V., A.F., B.D., A.K., M.B., and J.W. performed the research; A.O.B., R.M.C., A.S., L.A.T., M.S., A.V., N.P.T., and J.W. analyzed the data.
FUNDING INFORMATION
Funding was provided by Janssen Research and Development in collaboration with Columbia University Irving Medical Center. The sponsor was involved in study concept and design. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health (NIH). N.P.T. and A.O.B. were partially supported by National Center for Advancing Translational Sciences, National Institutes of Health, through grant UL1TR001873. N.P.T. is supported by the US National Institutes of Health grant R35GM131905. R.M.C. is supported by the US National Institutes of Health grant 2R01AA026302. M.D.R. is supported by the National Institute of General Medical Sciences grant R01GM138597; the National Center for Advancing Translational Sciences, National Institutes of Health grant UL1TR001878; National Human Genome Research Institute grants R01HG010067 and R01HG012670.
CONFLICT OF INTEREST STATEMENT
Patent for algorithm to Columbia University Trustees; © 2021 The Trustees of Columbia University in the City of New York. The owner has no objection to reproduction of the work for academic non‐commercial purposes, but otherwise reserves all copyright rights whatsoever. J.W., N.P.T. and A.O.B. are co‐inventors. J.W. has received research support from Janssen, Galectin, Intercept, Genfit, Shire, Conatus, Zydus, AMRA, and is on advisory boards for GlaxoSmithKline and AstraZeneca. R.M.C. has received research support from Intercept Pharmaceuticals and Merck, Inc., and consulting fees from Intercept and AstraZeneca. M.S. is funded by NIDDK R01DK132138, R01DK131547 and has an unrestricted grant from Grifols, SA. A.K. is a speaker and proctor for Intuitive, a reviewer for surgical videos for Crowd Sourced Assessment of Technical Skills (CSATs), and a consultant for Johnson and Johnson and Surgical Specialties Corporation. G.R. retired from Janssen Pharma R&D as Scientific Fellow and Head of Computational Sciences and is currently a Venture Partner in Samsara BioCapital, Palo Alto, CA. All other authors declared no competing interests for this work.
Supporting information
Data S1
ACKNOWLEDGMENTS
Joshua C. Denny MD, MS made scientific contributions to this project as faculty at Vanderbilt University Medical Center before joining NIH.
Basile AO, Verma A, Tang LA, et al. Rapid identification and phenotyping of nonalcoholic fatty liver disease patients using a machine‐based approach in diverse healthcare systems. Clin Transl Sci. 2025;18:e70105. doi: 10.1111/cts.70105
Contributor Information
Nicholas P. Tatonetti, Email: nicholas.tatonetti@csmc.edu.
Julia Wattacheril, Email: jjw2151@cumc.columbia.edu.
DATA AVAILABILITY STATEMENT
Algorithmic code is available for academic, non‐commercial collaborations by request to the corresponding authors.
REFERENCES
- 1. Younossi Z, Anstee QM, Marietti M, et al. Global burden of NAFLD and NASH: trends, predictions, risk factors and prevention. Nat Rev Gastroenterol Hepatol. 2018;15:11‐20. [DOI] [PubMed] [Google Scholar]
- 2. Malhi H, Brown RS, Lim JK, et al. Precipitous changes in nomenclature and definitions‐NAFLD becomes SLD: implications for and expectations of AASLD journals. Hepatology. 2023;78:1680‐1681. [DOI] [PubMed] [Google Scholar]
- 3. Alexander M, Loomis AK, Fairburn‐Beech J, et al. Real‐world data reveal a diagnostic gap in non‐alcoholic fatty liver disease. BMC Med. 2018;16:130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Harrison SA, Pierre B, Guy CD, et al. A phase 3, randomized, controlled trial of Resmetirom in NASH with liver fibrosis. N Engl J Med. 2024;390:497‐509. [DOI] [PubMed] [Google Scholar]
- 5. Estes C, Razavi H, Loomba R, Younossi Z, Sanyal AJ. Modeling the epidemic of nonalcoholic fatty liver disease demonstrates an exponential increase in burden of disease. Hepatology. 2018;67:123‐133. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Sanyal AJ, Van Natta ML, Clark J, et al. Prospective study of outcomes in adults with nonalcoholic fatty liver disease. N Engl J Med. 2021;385:1559‐1569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Carroll RJ, Eyler AE, Denny JC. Intelligent use and clinical benefits of electronic health records in rheumatoid arthritis. Expert Rev Clin Immunol. 2015;11:329‐337. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Basile AO, Ritchie MD. Informatics and machine learning to define the phenotype. Expert Rev Mol Diagn. 2018;18:219‐226. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Corey KE, Kartoun U, Zheng H, Chung RT, Shaw SY. Using an electronic medical records database to identify non‐traditional cardiovascular risk factors in nonalcoholic fatty liver disease. Am J Gastroenterol. 2016;111:671‐676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. DeRouen TA. Promises and pitfalls in the use of “big data” for clinical research. J Dent Res. 2015;94:107S‐109S. [DOI] [PubMed] [Google Scholar]
- 11. Van Vleck TT, Chan L, Coca SG, et al. Augmented intelligence with natural language processing applied to electronic health records for identifying patients with non‐alcoholic fatty liver disease at risk for disease progression. Int J Med Inform. 2019;129:334‐341. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Corey KE, Kartoun U, Zheng H, Shaw SY. Development and validation of an algorithm to identify nonalcoholic fatty liver disease in the electronic medical record. Dig Dis Sci. 2016;61:913‐919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Eremić‐Kojić N, Đerić M, Govorčin M, Balać D, Kresoja M, Kojić‐Damjanov S. Assessment of hepatic steatosis algorithms in non‐alcoholic fatty liver disease. Hippokratia. 2018;22:10‐16. [PMC free article] [PubMed] [Google Scholar]
- 14. CCHMC . Non‐alcoholic fatty liver disease (NAFLD) & alcoholic fatty liver disease (ALD). PheKB; 2016. Accessed September 03, 2021. https://phekb.org/phenotype/588
- 15. Klann JG, Joss MAH, Embree K, Murphy SN. Data model harmonization for the all of us research program: transforming i2b2 data into the OMOP common data model. PLoS One. 2019;14:e0212463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Hripcsak G, Duke JD, Shah NH, et al. Observational health data sciences and informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform. 2015;216:574‐578. [PMC free article] [PubMed] [Google Scholar]
- 17. Hripcsak G, Ryan PB, Duke JD, et al. Characterizing treatment pathways at scale using the OHDSI network. Proc Natl Acad Sci USA. 2016;113:7329‐7336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Reinecke I, Zoch M, Reich C, Sedlmayr M, Bathelt F. The usage of OHDSI OMOP – a scoping review. Ger Med Data Sci. 2021;2021:95‐103. [DOI] [PubMed] [Google Scholar]
- 19. Chelico JD, Wilcox AB, Vawdrey DK, et al. Designing a clinical data warehouse architecture to support quality improvement initiatives. AMIA Annu Symp Proc. 2017;2016:381‐390. [PMC free article] [PubMed] [Google Scholar]
- 20. Chalasani N, Younossi Z, Lavine JE, et al. The diagnosis and management of nonalcoholic fatty liver disease: practice guidance from the American Association for the Study of Liver Diseases. Hepatology. 2018;67:328‐357. [DOI] [PubMed] [Google Scholar]
- 21. Huhdanpaa HT, Tan WK, Rundell SD, et al. Using natural language processing of free‐text radiology reports to identify type 1 Modic endplate changes. J Digit Imaging. 2018;31:84‐90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Sterling RK, Lissen E, Clumeck N, et al. Development of a simple noninvasive index to predict significant fibrosis in patients with HIV/HCV coinfection. Hepatology. 2006;43:1317‐1325. [DOI] [PubMed] [Google Scholar]
- 23. Lin Z‐H, Xin Y‐N, Dong Q‐J, et al. Performance of the aspartate aminotransferase‐to‐platelet ratio index for the staging of hepatitis C‐related fibrosis: an updated meta‐analysis. Hepatology. 2011;53:726‐736. [DOI] [PubMed] [Google Scholar]
- 24. Angulo P, Hui JM, Marchesini G, et al. The NAFLD fibrosis score: a noninvasive system that identifies liver fibrosis in patients with NAFLD. Hepatology. 2007;45:846‐854. [DOI] [PubMed] [Google Scholar]
- 25. Younossi ZM, Golabi P, de Avila L, et al. The global epidemiology of NAFLD and NASH in patients with type 2 diabetes: a systematic review and meta‐analysis. J Hepatol. 2019;71:793‐801. [DOI] [PubMed] [Google Scholar]
- 26. Wattacheril J. Extrahepatic manifestations of nonalcoholic fatty liver disease. Gastroenterol Clin North Am. 2020;49:141‐149. [DOI] [PubMed] [Google Scholar]
- 27. Rinella ME, Neuschwander‐Tetri BA, Siddiqui MS, et al. AASLD practice guidance on the clinical assessment and management of nonalcoholic fatty liver disease. Hepatology. 2023;77(5):1797‐1835. doi: 10.1097/HEP.0000000000000323 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1
Data Availability Statement
Algorithmic code is available for academic, non‐commercial collaborations by request to the corresponding authors.
