Skip to main content
Gastro Hep Advances logoLink to Gastro Hep Advances
. 2026 Aug 11;5(11):101090. doi: 10.1016/j.gastha.2026.101090

Development and External Validation of an Algorithm for Identifying HDV RNA-Positive Patients

Robert J Wong 1, Robert G Gish 2, Ira M Jacobson 3, Joseph K Lim 4, Thomas Debray 5, Scott McDonald 5, Marvin Rock 6, Gary Leung 7, Chong Kim 6,∗
PMCID: PMC13587727  PMID: 42761485

Abstract

Background and Aims

Accurate detection and treatment of hepatitis D virus (HDV)-infected patients can reduce disease-related morbidity and mortality. This study developed and validated real-world evidence-based algorithms to detect ribonucleic acid (RNA)-positive HDV patients from administrative claims data.

Methods

This retrospective observational study identified hepatitis B virus and HDV patients from laboratory testing data linked to administrative claims (HealthVerity; 2015–2022), with external validation performed using electronic health records (TriNetX; 2005–2023). Both diagnosis-based and machine learning algorithms were evaluated. Performance metrics included area under the receiver-operating characteristic curve (AUROC), area under the precision-recall curve, sensitivity, specificity, positive predictive value, and accuracy. Internal validation results showed that diagnosis code-based algorithms identified HDV RNA-positivity with ≥88% accuracy.

Results

The best-performing algorithm-based approach in terms of optimism-adjusted AUROC was a random forest model with the following covariates: age, gender, hepatitis complications, HDV, and hepatitis B virus diagnosis, inpatient/outpatient diagnosis, hepatitis medication, and physician specialty (AUROC 97%). Decision curve analysis showed that the random forest algorithm with covariates excluding physician specialty, both inpatient/outpatient diagnosis, and physician specialty performed best. External validation indicated the random forest algorithm using equivalent covariates but without physician specialty had the best performance (positive predictive value 34%; accuracy 92%).

Conclusion

Algorithms based on HDV diagnosis were more effective for identifying HDV RNA-positive patients than algorithms using HDV-based variables from claims data. Further research into improving such algorithms is needed. Although fair discrimination may occur without claims data, key metrics were not able to replace traditional testing methods. The proposed models could help identify high-risk patients where certain strategies could be prioritized.

Keywords: Machine Learning, Evidence-Based Algorithm, Administrative Claims, Hepatitis Diagnosis

Introduction

Hepatitis D virus (HDV) is a defective ribonucleic acid (RNA) virus which only infects people positive for HBsAg.1 Individuals may be coinfected with HDV, where they are infected with both HDV and hepatitis B virus (HBV) simultaneously, or they can be superinfected, where HDV infection occurs subsequent to HBV infection.2 HDV is considered the most severe form of viral hepatitis.1,3 Compared to infection with HBV alone, infection with both HBV and HDV carries a greater risk of morbidity and mortality and is associated with accelerated progression to liver-related complications such as cirrhosis, decompensation, and HCC.1,3, 4, 5, 6 A meta-analysis estimated that, among patients positive for hepatitis B surface antigen (HBsAg), 18% and 20% of cirrhosis and hepatocellular carcinoma (HCC) cases, respectively, were attributed to concurrent HDV infection.7 Additionally, a claims-based study from the United States reported that HDV superinfection was associated with a 2.6-fold increased risk of cirrhosis and a 93% increased risk of HCC compared with chronic HBV infection.8

Chronic HDV infections are present in approximately 5% of patients with HBV worldwide,2,7 and a recent meta-analysis reported that approximately 75,000 individuals were living with HDV in the United States in 2022.9 However, the prevalence of HDV infection is believed to be underestimated.10,11 Factors contributing to the underestimation of HDV prevalence include insufficient testing among the HBV population and subsequent underdiagnosis,6,10, 11, 12 and heterogeneity in sampling and conducting serological tests.3,7 As a rare disease with limited treatment options, it may be less likely for HDV testing to be conducted, and diagnoses may also be less likely to be documented. Furthermore, current recommendations from the American Association for the Study of Liver Diseases advise initial HDV antibody testing only in individuals with HDV risk factors, such as people with a history of drug infection, men who have sex with men, and those with high-risk sexual behavior.13 Patients who screen positive for anti-HDV antibodies would then be tested for HDV RNA to confirm active HDV infection.13 HDV reporting in the United States is also currently voluntary, with the infection only reportable in 23 of the 50 states.6 Additionally, US estimates of HDV prevalence among the HBV population range from 4.6% to 13%;10,14 the wide variance could be attributed to studies evaluating only select or regional populations and small sample sizes, particularly relating to HDV RNA data.7,10

Artificial intelligence (AI), including machine learning approaches, is being increasingly utilized in healthcare research to analyze the vast and complex patient-level data available from patients in routine clinical care.15 Compared to traditional statistical or mathematical models, AI methods are applied to tasks normally performed by a human expert and are likely to involve considerably less time, expertise, and likelihood of error.15 Machine learning algorithms have been used in studies of viral hepatitis, such as epidemiological studies of hepatitis C virus (HCV) infections, the prevalence of which remains considerably underestimated in the United States.16 Recent studies using machine learning algorithms in detecting undiagnosed HCV patients from electronic medical records (EMR) or electronic health records (EHR) showed that such approaches reduced the number of individuals for screening, compared with universal screening, and hence could enable the prioritization of patients for targeted screening as well as significantly reduce costs associated with HCV screening.17,18 Outside of hepatology, machine learning approaches have also been successfully applied to identify patients from EHR with diverse conditions, including type 2 diabetes, common variable immunodeficiency disease, and mental health crisis.19, 20, 21 The benefits of using such approaches for patient identification include improved allocation of limited healthcare resources, improved rates of diagnosis and treatment, and consequently improved clinical outcomes, identification of hard-to-find patients who may be otherwise overlooked, and identification of high-risk patients.

There is currently a lack of understanding on the prevalence of HDV, defined here as those who test positive for HDV RNA; such information is necessary to inform the natural history of disease, comorbidities, economic burden, and clinical outcomes. The objective of this study was to develop a real-world evidence-based algorithm for the detection of RNA-positive HDV patients from administrative claims data and to validate this algorithm using claims-based variables.

Methods

Study Design and Data Sources

This was a retrospective observational study combining data from laboratory testing and healthcare administrative claims data (HealthVerity), with external validation using harmonized EHR data (TriNetX). HealthVerity Marketplace data provide a US claims dataset with >150 unique payers and >120 million patients and includes fully adjudicated pharmacy hospital and medical claims at the anonymized patient level sourced from commercial payers, Medicare, and Medicaid in the United States. Data from US laboratory testing results were obtained from Quest for the period from January 1, 2015, to July 31, 2022. TriNetX Dataworks Network is a US federated research network hosting EMR, cancer registries, and other forms of data (eg, genomic data from third party genomic testing labs) from >56 healthcare organizations and 91 million patients, and data were sourced from September 23, 2005, to January 20, 2023.

Population

Diagnosis and HealthVerity laboratory data were used to identify patients with HDV infection and monoinfection with HBV. Patients were eligible for the overall HBV cohort if they were adults (aged ≥18 years) with a first HBV diagnosis (ie, HBsAg positivity) during the study period who had ≥1 inpatient (IP) or >2 outpatient (OP) claims 30 days apart with an International Classification of Diseases (ICD)-9-CM or ICD-10-CM diagnosis code for HBV or any laboratory testing for HDV from January 2014 to December 2022. The HDV cohort included adults (aged ≥18 years) who had HDV laboratory tests with linked claims and who tested positive for HDV RNA at any time during the study period. The HBV monoinfection cohort included adults (aged ≥18 years) who were tested for HDV RNA during January 2015 to December 2022, had ≥1 negative HDV RNA result and no positive HDV RNA result.

Analysis

Baseline characteristics were described as categorical variables, expressed as frequencies and proportions, and continuous variables, expressed as means and standard deviations. P values were calculated using chi-squared tests for categorical variables and Mann Whitney U tests for continuous variables (age).

Development of the algorithm to detect RNA-positive HDV patients was conducted in 2 phases, both utilizing the same patient data and modeling methods. Phase 1 analyses primarily served as an exploratory analysis to identify the most promising (“best-performing”) modeling options. Subsequently, phase 2 analyses extended the statistical methodology for model development and validation by accounting for nonlinear effects and overoptimism. Finally, the best-performing models from phase 2 were externally validated using the TriNetX dataset. Model variables and algorithms used in phases 1 and 2 are further described below.

All programming and testing were performed using R.22 All models were generated using the CARET package with 10-fold cross-validation.23

Model variables

Choices of candidate variables for modeling were based on expert opinion, and variables included baseline demographics, clinical characteristics, and treatment utilization (Supplementary Table 1 and Table 1). Demographic variables included age (modeled either as linear or nonlinear covariate) and gender. Diagnosis of HDV and HBV was ascertained using diagnosis-based claims, that is, “≥1 ICD-diagnosis of HBV/HDV” and/or “1 IP or 2 OP claims for HDV” (≥30 days or more apart). Variables relating to hepatitis-related complications included compensated cirrhosis (CC), decompensated cirrhosis (DC), HCC, and liver transplant (LT). Variables on hepatitis-related treatment included use of adefovir, lamivudine, entecavir, tenofovir disoproxil fumarate, tenofovir alafenamide, and pegylated interferon. Variables on physician specialty/setting of medical visits included visits with hepatology, infectious disease, gastroenterologist, and primary care.

Table 1.

Variables Evaluated in All Models (Phase 2)

Covariate set Demographics Hepatitis complications HDV diagnosis HBV diagnosis IP/OP HDV diagnosis Hepatitis medicationsa Physician specialtyb
1 Agec, male gender CC, DC, liver cancer, LT ≥1 ICD-diagnosis of HDV ≥1 ICD-diagnosis of HBV ≥1 IP or ≥2 OP (30 d apart) diagnoses of HDV Yes Yes
2 Agec, male gender CC, DC, liver cancer, LT ≥1 ICD-diagnosis of HDV ≥1 ICD-diagnosis of HBV ≥1 IP or ≥2 OP (30 d apart) diagnoses of HDV Yes -
3 Aged (spline added), male gender CC, DC, liver cancer, LT ≥1 ICD-diagnosis of HDV ≥1 ICD-diagnosis of HBV ≥1 IP or ≥2 OP (30 d apart) diagnoses of HDV Yes -
4 Aged (spline added), male gender CC, DC, liver cancer, LT - - - Yes -
5 Aged (spline added), male gender CC, DC, liver cancer, LT - ≥1 ICD-diagnosis of HBV - Yes -
6 Aged (spline added), male gender CC, DC, liver cancer, LT ≥1 ICD-diagnosis of HDV ≥1 ICD-diagnosis of HBV - Yes -
7 Male gender CC, DC, liver cancer, LT - - - Yes -

GP, general practitioner; PCP, primary care physician; PegIFN, pegylated interferon.

a

Use of PegIFN, tenofovir alafenamide, entecavir, tenofovir disoproxil, lamivudine.

b

Visit with family practitioner, GP, PCP, hepatologist, infectious disease specialist, gastroenterologist, family health practitioner.

c

Age modeled as linear continuous variable.

d

Age modeled as nonlinear covariate.

Model development and internal validation (phase 1)

In phase 1, patient data were divided into training and testing sets (split by 80% and 20%, respectively). Using HDV status ascertained from the HealthVerity laboratory data as outcome variables and claims-based data items as input variables, models were generated using the training dataset via the following algorithms: (1) logistic regression with backward stepwise variable selection (GLM); (2) logistic regression with elastic net regularization (GLMNET), a logistic regression with the lasso penalties; (3) random forest, a bagging ensemble of decision trees; (4) generalized additive model (GAM), a GLM in which the response variable depends linearly on unknown smooth functions of the predictor variables; and (5) GAM with locally estimated scatterplot smoothing to enable a more precise prediction (GAMLSS). Specific sets of candidate predictor variables were included in each model (Supplementary Table 1). Variables on physician specialty of medical visits were only used for models based on the HealthVerity dataset. Two additional diagnosis-based algorithms were included, based on (1) any (≥1) HDV claim and (2) ≥1 IP or ≥2 OP claims (≥30 days apart); these are conventional methods of identifying patients with a certain disease from claims data and were included to enable the comparison between the predictive power of machine learning models against conventional methods.

Models were ordered by rank based on the performance metrics of accuracy, sensitivity, specificity, positive predictive value (PPV), area under the receiver-operating characteristic curve (AUROC), and area under the precision-recall curve (AUPRC), all of which were calculated with respect to a cutoff threshold (0.1–0.9). Each performance metric was ranked individually from 1 to 5 across models (1 indicating best performance). The cumulative score of all of the metrics for each model was generated by the sum of the metric category score; model performance was then ranked with the lowest sum of scores indicating the best-performing model. Models including and excluding physician specialty variables were generated and compared. The 2 best-performing models (with and without physician specialty) were then internally validated using the testing data (20% of the dataset) by generating confusion matrices, which were used to compare the predicted outcomes from a model against the actual outcomes from the data. Performance metrics that were assessed were sensitivity, specificity, accuracy, PPV, AUROC, and AUPRC.

Model development and internal validation (phase 2)

Phase 2 applied the same methodology as phase 1, apart from aspects described here. Instead of using an 80%/20% split sample as in phase 1, all the data were used for development of models and for internal validation in phase 2. A bootstrapping approach was used to obtain optimism-adjusted performance estimates to avoid overfitting of the model24 from using the same data for both purposes. The same modeling methods (GLM, GLMNET, random forest, GAM, and GAMLSS) were evaluated in phase 2, as well as XGBoost (xgbTree and xgbLinear), a scalable machine learning system for tree boosting.25 Additionally, only a restricted set of candidate predictors was used in phase 2, based on input from experts (Table 1). Continuous covariates (eg, age) were not categorized but were instead modeled explicitly as continuous while allowing for nonlinear effects. In total, the 7 modeling methods were applied to all 7 covariate sets, resulting in 49 predictive models being developed and evaluated in phase 2. Models from covariate set 1 were not considered in performance assessment because they cannot be validated or tested with external data using the TriNetX database.

In phase 2, discrimination (optimism-adjusted AUROC) of each model and variant covariate set was the metric of choice to assess model performance, though AUPRC, sensitivity, specificity, PPV, prevalence, and accuracy were also reported. Cutoff thresholds of 0.1 to 0.9 were applied. Decision curves were generated to assess the net benefit over a range of risk thresholds, that is, the net benefit algorithms for treating all patients vs treating none. As the models output a risk that ranges between 0 and 1, these probabilities would need to be converted to discrete, typically binary, decisions, and decision curve analyses allow the evaluation of the predictors with different probability thresholds.

External validation

During both phases of the analyses, the selected models without physician specialty variables were further validated using TriNetX HBV data (2005–2023), using the same procedures to assess laboratory-ascertained HDV status, input variables, and model performance. The same criteria were applied to the TriNetX dataset (test data) to assess covariates. Predictive models were applied to covariates in the test data, and confusion matrices were used to assess performance. Calibration performance in the HealthVerity database was also assessed during phase 2 to assess how well the model generalized to external validation data from TriNetX. Sensitivity, specificity, PPV, prevalence, and accuracy of the best-performing model were compared to the diagnosis code-based algorithms (“any HDV” and “≥1IP/≥2OP claims”).

Results

Patient Characteristics in the HealthVerity Dataset

Of the 127,689,474 individuals included in the HealthVerity database during the study period, 96 patients with HDV and 1138 patients with HBV monoinfection met the eligibility criteria and were included in the analysis (Supplementary Figure 1).

There were more male patients in the HDV cohort (72.92%) compared with the HBV cohort (55.10%; P = .0008) (Table 2). Higher proportions of patients in the HDV cohort than in the HBV cohort had hepatic-related complications of CC (52.08% vs 15.73%; P < .0001), DC (34.38% vs 19.42%; P = .0009), and HCC (13.54% vs 4.22%; P = .0005) and underwent LT (9.38% vs 2.99%; P = .0044). The comorbidity profiles between the HDV and HBV cohorts were comparable, but fewer patients with HDV had hypertension at baseline compared with patients with HBV (33.33% vs 45.69%; P = .0244). More patients in the HDV cohort compared with the HBV cohort were treated with pegylated interferon (11.46% vs 0.09%; P < .0001), tenofovir disoproxil fumarate (36.46% vs 22.93%; P = .0041), and tenofovir alafenamide (23.96% vs 9.93%; P = .0001), while utilization of adefovir, entecavir, and lamivudine was comparable between cohorts.

Table 2.

Baseline Demographics and Disease Characteristics (HealthVerity Dataset)

Baseline characteristics HDV cohort (n = 96) HBV cohort (n = 1138) P value
Age, mean (SD), y 47.51 (13.16) 46.60 (13.33) .7127
Age at start of study period, n (%)
 18–34 y 13 (13.54%) 214 (18.80%) .2197
 35–44 y 31 (32.29%) 268 (23.55%) .0626
 45–54 y 22 (22.92%) 292 (25.66%) .6262
 55–64 y 21 (21.88%) 253 (22.23%) 1.0000
 65–74 y 7 (7.29%) 81 (7.12%) .8386
 ≥75 y 1 (1.04%) 14 (1.23%) 1.0000
Sex, n (%)
 Male 70 (72.92%) 627 (55.10%) .0008
 Female 26 (27.08%) 510 (44.82%) .0008
Hepatic complications/comorbidity profile, n (%)
 Compensated cirrhosis 50 (52.08%) 179 (15.73%) <.0001
 Decompensated cirrhosis 33 (34.38%) 221 (19.42%) .0009
 Hepatocellular carcinoma 13 (13.54%) 48 (4.22%) .0005
 Liver transplant 9 (9.38%) 34 (2.99%) .0044
 Sexually transmitted infections 7 (7.29%) 100 (8.79%) .8496
 Hypertension 32 (33.33%) 520 (45.69%) .0244
 History of smoking 25 (26.04%) 332 (29.17%) .5595
 HCV 27 (28.13%) 242 (21.27%) .1231
 HIV 6 (6.25%) 84 (7.38%) .8389
 Mental health disorder 8 (8.33%) 73 (6.41%) .5165
 Obesity 20 (20.83%) 265 (23.29%) .7051
 Substance abuse 12 (12.50%) 162 (14.24%) .7604
 Alcohol abuse or dependence/alcohol use disorder 14 (14.58%) 126 (11.07%) .3133
 Nonalcoholic steatohepatitis 18 (18.75%) 285 (25.04%) .2162
 Extrahepatic cancers 0 (0.00%) 1 (0.09%) 1.0000
HBV or HDV diagnosis claims, n (%)
 HBV diagnosis 78 (81.25%) 1138 (100.00%) <.0001
 HDV diagnosis 61 (63.54%) 115 (10.11%) <.0001
 1 IP/2 OP HDV diagnosis 51 (53.13%) 47 (4.13%) <.0001
Positive HDV RNA tests, n (%) 96 (100.00%) - -
Treatment utilization, n (%)
 Pegylated interferon 11 (11.46%) 1 (0.09%) <.0001
 Adefovir 0 (0.00%) 4 (0.35%) 1.0000
 Entecavir 13 (13.54%) 147 (12.92%) .8741
 Lamivudine 0 (0.00%) 4 (0.35%) 1.0000
 Tenofovir disoproxil 35 (36.46%) 261 (22.93%) .0041
 Tenofovir alafenamide 23 (23.96%) 113 (9.93%) .0001
Physician specialty on claims
 Family 11 (11.46%) 214 (18.80%) .0747
 Family health 4 (4.17%) 6 (0.53%) .005
 Gastroenterology 47 (48.96%) 506 (44.46%) .3953
 General practice 6 (6.25%) 47 (4.13%) .2955
 Hepatology 10 (10.42%) 64 (5.62%) .0703
 Infectious disease 12 (12.50%) 157 (13.80%) .8771
 Primary care 1 (1.04%) 38 (3.34%) .3583
Insurance
 Commercial 40 (41.67%) 475 (41.74%) 1.0000
 Medicare 7 (7.29%) 108 (9.49%) .5851
 Medicaid 39 (40.63%) 433 (38.05%) .6621
 Other 10 (10.42%) 122 (10.72%) 1.0000

HIV, human immunodeficiency virus; SD, standard deviation.

Internal Validation Results (HealthVerity Dataset)

In phase 1, the GLMNET model without physician specialty was the best-performing model based on the overall rating of 6 and also outperformed both diagnoses-based algorithms (Supplementary Tables 2 and 3).

In phase 2, model outcomes for the diagnosis-based and selected algorithm-based models (using covariate sets 1, 3, and 6) are shown in Table 3, as they represented the machine learning-based models which provided the highest AUROC scores. Outcomes for all machine learning-based algorithms and covariate sets are shown from Supplementary Tables 4–10. Using ICD-10 codes alone identified the HDV diagnosis with 88%–93% accuracy. Initial identification of the best-performing algorithm-based models indicated that the random forest approach with covariate set 1 variables was the best-performing model, based on optimism-adjusted AUROC (97%). The next best performing model was the random forest approach with covariate sets 3 and 6 variables, based on optimism-adjusted AUROC (96%). There were other models which showed specific higher-performing thresholds, including the random forest approach with covariate set 2 (AUROC 94%) variables and the GLMNET approach with covariate set 3 variables (AUROC 93%). Additionally, some models which did not include diagnosis claims records were able to generate meaningful predictions; for example, the random forest approach yielded an AUROC of 80% with covariate set 4 using only limited information relating to age, hepatitis complications, and medications. Notably, certain outcomes relating to sensitivity, specificity, and PPV could not be estimated due to the small sample sizes in the cohort of patients who were positive for HDV RNA.

Table 3.

Optimism-Corrected Model Performance in the HealthVerity Dataset (Phase 2, Selected Results)

Covariate set Algorithm AUROC AUPRC Cutoffa Sensitivity Specificity PPV Accuracy
Dx-based identification of HDVb
 Any HDV N/A N/A N/A 0.64 0.90 0.35 0.88
 ≥1IP/≥2OP HDV (30 d apart) N/A N/A N/A 0.53 0.96 0.52 0.93
Machine learning-based identification of HDVc
 1 Min, max 0.75, 0.97 0.27, 0.91 0.10, 0.90 −0.02, 0.92 0.89, 1.00 0.34, 1.00 0.87, 0.97
GAM 0.85 0.45 0.1 0.65 0.89 0.34 0.87
0.4 0.42 0.98 0.59 0.93
0.5 0.34 0.98 0.64 0.93
0.6 0.27 0.99 0.70 0.93
0.7 0.21 0.99 0.79 0.93
GAMLSS 0.86 0.46 0.1 0.67 0.89 0.35 0.88
0.3 0.49 0.96 0.52 0.93
0.4 0.42 0.98 0.59 0.93
0.5 0.35 0.98 0.65 0.93
0.6 0.28 0.99 0.71 0.93
0.7 0.20 1.00 0.79 0.93
0.8 0.08 1.00 0.76 0.93
GLM 0.93 0.70 0.1 0.77 0.92 0.46 0.91
0.4 0.59 0.98 0.71 0.95
0.5 0.55 0.99 0.76 0.95
0.6 0.50 0.99 0.79 0.95
0.7 0.43 0.99 0.83 0.95
0.8 0.37 1.00 0.90 0.95
GLMNET 0.92 0.73 0.1 0.78 0.94 0.52 0.93
0.3 0.59 0.98 0.70 0.95
0.4 0.58 0.98 0.72 0.95
0.5 0.53 0.98 0.72 0.95
0.6 0.34 1.00 0.94 0.95
RF 0.97 0.91 0.1 0.92 0.90 0.37 0.90
0.4 0.81 0.98 0.81 0.97
0.5 0.77 0.99 0.88 0.97
0.6 0.69 0.99 0.92 0.97
xgbLinear 0.75 0.28 0.1 −0.02 1.00 NaN 0.92
0.2 0.00 1.00 NaN 0.92
0.3 0.00 1.00 NaN 0.92
0.4 0.00 1.00 NaN 0.92
0.5 0.00 1.00 NaN 0.92
0.6 0.00 1.00 NaN 0.92
0.7 0.00 1.00 NaN 0.92
0.8 0.00 1.00 NaN 0.92
0.9 0.00 1.00 NaN 0.92
xgbTree 0.95 0.80 0.1 0.83 0.93 0.48 0.92
0.4 0.68 0.99 0.85 0.97
 3 Min, max 0.83, 0.96 0.45, 0.81 0.10, 0.90 −0.01, 0.85 0.90, 1.00 0.34, 1.00 0.88, 0.96
GAM 0.83 0.45 0.1 0.62 0.90 0.35 0.88
0.6 0.27 0.99 0.72 0.94
GAMLSS 0.84 0.46 0.1 0.65 0.90 0.34 0.88
0.3 0.50 0.96 0.53 0.93
0.4 0.40 0.98 0.59 0.93
0.5 0.33 0.98 0.65 0.93
0.6 0.28 0.99 0.70 0.93
0.7 0.17 0.99 0.73 0.93
0.8 0.08 1.00 0.74 0.93
GLM 0.93 0.72 0.1 0.77 0.92 0.46 0.91
0.3 0.68 0.97 0.64 0.95
0.4 0.57 0.98 0.71 0.95
0.5 0.52 0.98 0.74 0.95
0.6 0.51 0.99 0.79 0.95
0.7 0.42 0.99 0.85 0.95
0.8 0.35 1.00 0.93 0.95
0.9 0.32 1.00 0.97 0.95
GLMNET 0.93 0.72 0.1 0.78 0.93 0.47 0.92
0.3 0.66 0.97 0.68 0.95
0.4 0.60 0.98 0.71 0.95
0.5 0.52 0.98 0.73 0.95
0.6 0.47 0.99 0.83 0.95
0.7 0.39 1.00 0.91 0.95
0.8 0.33 1.00 0.97 0.95
RF 0.96 0.81 0.3 0.66 0.99 0.80 0.96
0.4 0.51 0.99 0.90 0.96
0.8 −0.01 1.00 NaN 0.92
0.9 0.00 1.00 NaN 0.92
xgbLinear 0.87 0.49 0.1 0.00 1.00 NaN 0.92
0.2 0.00 1.00 NaN 0.92
0.3 0.00 1.00 NaN 0.92
0.4 0.00 1.00 NaN 0.92
0.5 0.00 1.00 NaN 0.92
0.6 0.00 1.00 NaN 0.92
0.7 0.00 1.00 NaN 0.92
0.8 0.00 1.00 NaN 0.92
0.9 0.00 1.00 NaN 0.92
xgbTree 0.95 0.79 0.1 0.83 0.93 0.49 0.92
0.4 0.63 0.99 0.80 0.96
0.5 0.60 0.99 0.86 0.96
 6 Min, max 0.83, 0.96 0.42, 0.75 0.1, 0.9 0.00, 0.90 0.87, 1.00 0.29, 0.99 0.85, 0.97
GAM 0.83 0.42 0.1 0.64 0.87 0.30 0.85
0.4 0.36 0.98 0.60 0.93
0.5 0.29 0.99 0.63 0.93
0.6 0.17 0.99 0.76 0.93
0.7 0.11 1.00 0.72 0.93
GAMLSS 0.83 0.43 0.1 0.66 0.87 0.29 0.85
0.4 0.39 0.98 0.63 0.93
0.5 0.28 0.99 0.64 0.93
0.6 0.18 0.99 0.74 0.93
0.7 0.12 1.00 0.73 0.93
GLM 0.93 0.69 0.1 0.78 0.90 0.40 0.89
0.4 0.56 0.98 0.68 0.95
0.5 0.53 0.99 0.76 0.95
0.6 0.46 0.99 0.82 0.95
0.7 0.38 1.00 0.91 0.95
0.8 0.34 1.00 0.97 0.95
GLMNET 0.93 0.70 0.1 0.80 0.90 0.39 0.89
0.5 0.50 0.99 0.80 0.95
0.6 0.41 0.99 0.87 0.95
0.7 0.33 1.00 0.94 0.95
0.8 0.30 1.00 0.98 0.95
RF 0.96 0.70 0.1 0.90 0.90 0.37 0.90
0.4 0.79 0.98 0.82 0.97
0.5 0.77 0.99 0.86 0.97
0.6 0.75 0.99 0.90 0.97
xgbLinear 0.86 0.45 0.1 0.00 1.00 NaN 0.92
0.2 0.00 1.00 NaN 0.92
0.3 0.00 1.00 NaN 0.92
0.4 0.00 1.00 NaN 0.92
0.5 0.00 1.00 NaN 0.92
0.6 0.00 1.00 NaN 0.92
0.7 0.00 1.00 NaN 0.92
0.8 0.00 1.00 NaN 0.92
0.9 0.00 1.00 NaN 0.92
xgbTree 0.94 0.75 0.1 0.82 0.91 0.41 0.90
0.3 0.61 0.98 0.69 0.95
0.4 0.57 0.99 0.80 0.95
0.5 0.51 0.99 0.83 0.95
0.6 0.47 0.99 0.90 0.95
0.7 0.41 1.00 0.92 0.95
0.8 0.32 1.00 0.96 0.95

CI, confidence interval; Dx, diagnosis; LSS, Loess; NaN, not a number (undefined) due to 0 counts within these groups; RF, random forest; xgb, XGBoost.

Results are presented for models which performed the best in terms of AUROC, with results shown for selected cutoffs for each algorithm set based on accuracy.

Bold values indicate the minimum and maximum values across the evaluated options.

a

Cutoff thresholds evaluated from 0.1 to 0.9.

b

Prevalence for both diagnosis-based algorithms was 0.08; 95% CIs for accuracy for “any HDV” was 95% CI 0.86 to 0.90 and for “≥1IP/≥2OP HDV” was 95% CI 0.91 to 0.94.

c

Data presented for 3 covariate sets with the highest AUROC scores (among a total of 7 covariate sets analyzed), the minimum and maximum across all algorithms and cutoff thresholds within these 3 covariate sets, and selected results within each of the 3 covariate sets based on highest and lowest accuracy scores for each algorithm.

Decision Curves

Decision curves for models using covariate sets 2 to 7 based on the HealthVerity dataset are shown in Figure 1. No decision curves were generated for the models using covariate set 1, as the TriNetX data used for external validation did not include data on physician specialty. A random forest approach with covariate set 2 and 6 variables performed the best after adjusting for optimism and remained above the treat-all/treat-none alternatives across all thresholds.

Figure 1.

Figure 1

Decision curves for models using covariate sets 2 to 7 (HealthVerity dataset). RF, random forest; xgb, XGBoost.

External Validation Using the TriNetX Dataset

Patient characteristics

A total of 200 and 3149 patients were included in the HDV and HBV cohorts from the TriNetX dataset, respectively (Table 4). More patients in the HDV cohort than in the HBV cohort had CC (47.00% vs 25.28%; P < .0001), DC (49.50% vs 34.07%; P < .0001), HCC (29.50% vs 18.55%; P = .0003), and underwent LT (7.50% vs 3.24%; P = .0042). There were more patients in the HDV cohort with HCV infection than in the HBV cohort (30.00% vs 19.05%; P = .0003).

Table 4.

Baseline Demographics and Disease Characteristics (TriNetX Dataset)

Baseline characteristics HDV cohort (n = 200) HBV cohort (n = 3149) P value
Age, mean (SD), y 45.73 (13.40) 46.72 (14.01) .3134
Age at start of study period, n (%)
 18–34 y 49 (24.50%) 702 (22.29%) .4843
 35–44 y 52 (26.00%) 789 (25.06%) .8008
 45–54 y 46 (23.00%) 670 (21.28%) .5935
 55–64 y 34 (17.00%) 607 (19.28%) .4596
 65–74 y 15 (7.50%) 309 (9.81%) .3246
 ≥75 y 4 (2.00%) 72 (2.29%) 1.0000
Sex, n (%)
 Male 114 (57.00%) 1836 (58.30%) .7124
 Female 86 (43.00%) 1312 (41.66%) .7122
Hepatic complication/comorbidity profile, n (%)
 Compensated cirrhosis 94 (47.00%) 796 (25.28%) <.0001
 Decompensated cirrhosis 99 (49.50%) 1073 (34.07%) <.0001
 Hepatocellular carcinoma 59 (29.50%) 584 (18.55%) .0003
 Liver transplant 15 (7.50%) 102 (3.24%) .0042
 Sexually transmitted infections 39 (19.50%) 492 (15.62%) .1615
 Hypertension 82 (41.00%) 1206 (38.30%) .4542
 History of smoking 71 (35.50%) 975 (30.96%) .1817
 HCV 60 (30.00%) 600 (19.05%) .0003
 HIV 29 (14.50%) 340 (10.80%) .1040
 Mental health disorder 18 (9.00%) 261 (8.29%) .6924
 Obesity 40 (20.00%) 575 (18.26%) .5113
 Substance abuse 46 (23.00%) 559 (17.75%) .0712
 Alcohol abuse or dependence/alcohol use disorder 32 (16.00%) 413 (13.12%) .2381
 Nonalcoholic steatohepatitis 38 (19.00%) 684 (21.72%) .4247
 Extrahepatic cancers 0 (0.00%) 3 (0.10%) 1.0000
HBV or HDV diagnosis claims, n (%)
 HBV diagnosis 200 (100.00%) 3149 (100.00%)
 HDV diagnosis 110 (55.00%) 244 (7.75%) <.0001
 1 inpatient/2 OP HDV diagnosis 67 (33.50%) 97 (3.08%) <.0001
 Positive HDV RNA tests, n (%) 200 (100.00%) - -
Treatment utilization, n (%)
 Pegylated interferon 1 (0.50%) 0 (0.00%) .0597
 Adefovir 0 (0.00%) 10 (0.32%) 1.0000
 Entecavir 10 (5.00%) 168 (5.34%) 1.0000
 Lamivudine 0 (0.00%) 6 (0.19%) 1.0000
 Tenofovir disoproxil 17 (8.50%) 268 (8.51%) 1.0000
 Tenofovir alafenamide 20 (10.00%) 200 (6.35%) .0541

HIV, human immunodeficiency virus; SD, standard deviation.

Validation results

Model outcomes based on the TriNetX dataset are presented in Table 5, which shows results from confusion matrices by applying the predictive models to external data and comparing the predicted diagnosis against the actual diagnosis. A confusion matrix was not generated using covariate set 1, as the TriNetX database did not include data on physician specialty. The diagnosis-based algorithm of “≥1IP/≥2OP” HDV diagnoses had the best performance overall (PPV of 41% and accuracy of 93% [95% confidence interval: 92–94]). Although both diagnosis-based algorithms performed better than the majority of remaining models, sensitivity and PPV may still be considered fairly low. Among the machine learning-based algorithms, the random forest model using covariate set 2 had the best performance in terms of PPV (34%) and accuracy (92%, 95% confidence interval: 91–93). Sensitivity for the XGBoost algorithms ranged from 92% to 100% across covariate sets 2 to 7, but PPV and accuracy were notably low (6% and range 6%–7%, respectively). Excluding the XGBoost algorithms, sensitivity of most algorithms using covariate set 6 (range, 46%–55%) was comparable to that for “any HDV” (55%).

Table 5.

Model Performance in the TriNetX Dataset (External Validation)

Covariate set Algorithm Sensitivity Specificity PPV Prevalence Accuracy (95% CI)
Dx-based identification of HDV
 Any HDV 0.55 0.92 0.31 0.06 0.90 (0.89, 0.91)
 ≥1IP/≥2OP HDV (30 d apart) 0.34 0.97 0.41 0.06 0.93 (0.92, 0.94)
Machine learning-based identification of HDV
 2 GLM 0.45 0.93 0.28 0.06 0.90 (0.89, 0.91)
RF 0.38 0.95 0.34 0.06 0.92 (0.91, 0.93)
GLMNET 0.10 0.94 0.09 0.06 0.89 (0.88, 0.90)
GAM 0.47 0.90 0.23 0.06 0.88 (0.86, 0.89)
GAMLSS 0.47 0.90 0.23 0.06 0.87 (0.86, 0.88)
xgbLinear 0.93 0.01 0.06 0.06 0.06 (0.06, 0.07)
xgbTree 0.99 0.00 0.06 0.06 0.06 (0.05, 0.07)
 3 GLM 0.46 0.92 0.26 0.06 0.89 (0.88, 0.90)
RF 0.42 0.93 0.27 0.06 0.90 (0.89, 0.91)
GLMNET 0.10 0.94 0.09 0.06 0.89 (0.88, 0.90)
GAM 0.47 0.89 0.21 0.06 0.86 (0.85, 0.87)
GAMLSS 0.45 0.89 0.21 0.06 0.86 (0.85, 0.88)
xgbLinear 0.94 0.01 0.06 0.06 0.06 (0.05, 0.07)
xgbTree 0.99 0.00 0.06 0.06 0.06 (0.05, 0.07)
 4 GLM 0.44 0.75 0.10 0.06 0.73 (0.71, 0.75)
RF 0.36 0.85 0.13 0.06 0.82 (0.81, 0.83)
GLMNET 0.08 0.97 0.13 0.06 0.91 (0.90, 0.92)
GAM 0.40 0.77 0.10 0.06 0.74 (0.73, 0.76)
GAMLSS 0.38 0.77 0.09 0.06 0.75 (0.73, 0.76)
xgbLinear 0.99 0.01 0.06 0.06 0.07 (0.06, 0.08)
xgbTree 1.00 0.00 0.06 0.06 0.06 (0.05, 0.07)
 5 GLM 0.43 0.79 0.11 0.06 0.77 (0.75, 0.78)
RF 0.34 0.87 0.14 0.06 0.84 (0.82, 0.85)
GLMNET 0.71 0.19 0.05 0.06 0.22 (0.20, 0.23)
GAM 0.46 0.74 0.10 0.06 0.72 (0.71, 0.74)
GAMLSS 0.38 0.77 0.09 0.06 0.75 (0.73, 0.76)
xgbLinear 0.99 0.01 0.06 0.06 0.07 (0.06, 0.07)
xgbTree 1.00 0.00 0.06 0.06 0.06 (0.05, 0.07)
 6 GLM 0.54 0.90 0.25 0.06 0.88 (0.86, 0.89)
RF 0.46 0.87 0.19 0.06 0.85 (0.84, 0.86)
GLMNET 0.53 0.25 0.04 0.06 0.27 (0.25, 0.28)
GAM 0.55 0.86 0.20 0.06 0.84 (0.83, 0.86)
GAMLSS 0.53 0.87 0.20 0.06 0.85 (0.84, 0.86)
xgbLinear 0.92 0.01 0.06 0.06 0.07 (0.06, 0.08)
xgbTree 0.98 0.00 0.06 0.06 0.06 (0.05, 0.07)
 7 GLM 0.46 0.74 0.10 0.06 0.73 (0.71, 0.74)
RF 0.09 0.96 0.12 0.06 0.91 (0.90, 0.92)
GLMNET 0.08 0.97 0.13 0.06 0.91 (0.90, 0.92)
GAM 0.50 0.73 0.10 0.06 0.71 (0.70, 0.73)
GAMLSS 0.50 0.73 0.10 0.06 0.71 (0.70, 0.73)
xgbLinear 1.00 0.00 0.06 0.06 0.06 (0.05, 0.07)
xgbTree 1.00 0.00 0.06 0.06 0.06 (0.05, 0.07)

CI, confidence interval; Dx, diagnosis; LSS, loess; RF, random forest; xgb, XGBoost.

Decision curves

Decision curves for the models using covariate sets 2 to 7 based on the TriNetX dataset are shown in Figure 2. For models using covariate set 2 variables, the random forest approach showed a net benefit for a very small range of lower thresholds before this distinction from treating no patients was lost. For models using covariate set 6 variables, the random forest approach showed negative benefits across all thresholds. In contrast to models using covariate sets 2, 3, and 6, those using covariate sets 4, 5, and 7 showed negative benefit on most thresholds.

Figure 2.

Figure 2

Decision curves for models using covariate sets 2 to 7 (TriNetX dataset). RF, random forest; xgb, XGBoost.

Calibration

The calibration curves for models using covariate sets 2 to 7 are shown from Supplementary Figures 2–7. Results of the calibration showed that the calibration slopes for the random forest and GLMNET approaches diverged substantially from 1.0, which suggested overfitting of the random forest and GLMNET models when applied to external validation data. Overfitting leads to predictions with low external validity and may have occurred because the models were too complex for the data.

Discussion

HDV infection is associated with an accelerated progression to liver-related complications compared with chronic HBV alone.1,3, 4, 5, 6 Patients with active HDV infection, that is, those who have a positive HDV RNA test, have an estimated 70% 10-year risk of liver failure, death, LT, or HCC.26 Early detection of RNA-positive HDV patients would enable earlier treatment, which likely leads to improved clinical outcomes and reduced economic burden.

To our knowledge, this study was the first analysis that used a machine learning approach to develop an algorithm to detect HDV patients who are positive for HDV RNA. Such an algorithm, which may likely be based on patient clinical and demographic characteristics, could be used by clinicians to identify individuals that may benefit from HDV RNA testing, as this would inform early diagnosis and subsequent treatment is important. In this study, the diagnosis-based algorithm of “≥1IP/≥2OP” diagnoses had the best performance in identifying HDV RNA-positive patients, with 93% accuracy and a PPV of 41%. However, claims-based diagnoses have limited utility and such data cannot be used for predicting HDV presence at healthcare visits. If such records are not available for prediction, although predictive power may be reduced, performance based on standard information available before initiating any diagnostic procedure, for example, demographic/clinical characteristics, remained reasonably high. Among algorithm-based models, the model using the random forest approach and covariate set 2 had the best predictive performance, with 92% accuracy and a PPV of 34%.

Covariate set 4 did not include HDV/HBV or ≥1IP/≥2OP diagnoses claims while covariate set 5 included only HBV diagnosis claims, which more realistically reflects the information available when making predictions under routine clinical practice settings. Among the models using covariates 4 and 5, the XGBoost approaches had the highest sensitivity (99%–100%), but specificity, PPV, and accuracy were very low (0%–1%, 6%, and 6%–7%, respectively). Although the GLMNET approach using covariate set 4 had the highest specificity (97%) without using claims data indicative of HDV diagnosis, sensitivity and PPV were also low (8% and 13%, respectively).

Overall, the predictive model with the highest AUROC should be prioritized, as this would best discriminate HDV from HBV patients. Since the other metrics (accuracy, sensitivity, specificity, and PPV) vary according to the decision threshold, tuning may be needed to ensure that the model identifies target high-risk subgroups. Moreover, models with good internal validation may perform poorly upon external validation, with performance metrics being highly affected by model calibration. Most prediction models therefore require recalibration, typically by updating the intercept in regression models, to resolve these artifacts.

These results were based on methodology that was optimized using initial analyses showing the net benefit potential over a range of risk thresholds. External validation using the TriNetX dataset further assessed the robustness and limitations of results. However, the presence of overfitting and lower values for PPV and sensitivity when models were applied to the validation data suggest issues with the generalizability of results that should be addressed. There are several areas where such a model could be applied in the field of hepatology and in future real-world studies. In clinical practice, a generalizable model with sufficient sensitivity and net benefit ratio could be used to interrogate data from EHR and may enable clinicians to better target tests and improve diagnosis rates. Real-world data sources could also be leveraged in research to further explore areas of concern, such as demonstrating true disease prevalence, measuring healthcare resource utilization, and researching disease progression, all of which support viral elimination efforts.

AI technologies are increasingly being used to manage viral hepatitis, and applications include detecting liver fibrosis in chronic liver disease, predicting disease recurrence, identifying hepatitis serological markers (eg, HBsAg), detecting precancerous lesions, and identifying high-risk individuals.27 A non-systematic review conducted by Ali and colleagues noted that the implementation of machine learning and deep learning resulted in accuracy of ≥88.3% in the early diagnosis and detection of viral hepatitis, using best-performing algorithms such as random forest, support vector machines, decision trees, and K-nearest neighbors.27 Additionally, a recent study from Chen and colleagues17 used a machine learning approach to identify key clinical attributes and social determinants of health to improve screening of HCV among approximately 300,000 individuals included in an EHR dataset. The predictive model incorporated key features associated with HCV diagnosis, and the addition of a stepwise feature improved precision from 25% using only demographics, to 51% when laboratory data were included, to 73% when diagnostics/treatments were included, and finally to 93% when social determinants of health were included.17 On validation using the target EHR dataset, the model had a PPV (indicating model precision) of 91%, sensitivity of 30%, and AUROC (indicating overall predictive ability) of 86%.17 Furthermore, authors projected that HCV screening would be reduced by 75% if the predictive model was implemented in the healthcare system, which could result in cost-savings.17 Similarly, the retrospective study from Rigg and colleagues18 applied a machine learning algorithm to identify undiagnosed HCV patients from over 28 million individuals in the United States. Predictor selection in the study was based on clinical expert knowledge, which defined events relevant to HCV that spanned diagnoses, prescriptions, procedures, and laboratory tests; predictors were then mapped to clinical codes and extracted from the EMR database.18 Authors used a gradient boosting trees algorithm trained on EMR data which achieved a 101.0-fold, 18.0-fold, and 5.1-fold improvement in precision over universal screening at 5%, 20%, and 50% levels of recall, respectively.18

Further development of these models would involve retaining the current strengths while addressing limitations and issues with generalizability. There may be other avenues to confirm a diagnosis of HDV as an outcome variable, which may allow a larger sample set to be used, such as using claims encounters with HDV as the primary diagnosis paired with specific procedures that may more definitively suggest acute HDV infection. Additionally, simpler model structures may improve calibration with external data sources and support greater generalizability. Further demographic variables could be pursued, as well as laboratory and relevant treatment data and also as those related to social determinants of health; these variables may ensure equity and recognize the disparities and care gaps that exist. Indeed, the study from Chen and colleagues17 noted that a substantial portion of predictive ability was derived from variables of social determinants of health. There could also be potential to incorporate biomarkers into such predictive models and tools. Additional considerations for real-life implementation in both patient care and research include the potential for deployment in EMR and payer-specific claims databases to inform providers and care managers on patient care and increase targeted laboratory testing, ethical considerations, and improved epidemiological identification of cases in research to better estimate true disease burden, progression, and costs.

This study is subject to limitations inherent to claims-based studies. The current analysis was limited by the small sample size, which may be due to suboptimal HDV screening practices, lack of access to laboratory testing, and the possibility that not all sources of laboratory data were included (eg, Quest only and not LabCorp). Notably, in the meta-analysis of 282 studies to evaluate the global prevalence of HDV, Stockdale et al7 suggested that the issue of small sample sizes was a systemic problem for real-world studies of HBV/HDV, particularly those evaluating populations from North America. Data used in the analyses were temporal, and some data could be missing; such data may have value only for retrospective patient identification. Patients who were included in the HBV cohort after having been tested as negative for HDV RNA could have subsequently tested positive and not captured within the HDV cohort, either if the positive test was performed elsewhere or beyond the study period. The lack of censoring and potential gaps in laboratory data (eg, tests occurring outside of Quest would not be captured) could potentially have allowed patients with unobserved HDV infection to be included among the HBV monoinfection cohort. The comprehensiveness of matches between claims and laboratory data may be uncertain, and the included population could be subject to bias. As the nature of claims data and testing is based on encounters and is opportunistic, laboratory testing may have been prioritized for patients with more severe disease. Although nearly 128 million patients are included in the HealthyVerity database, patients are not evenly distributed across categories of baseline characteristics (eg, not all US states may have been represented), and some groups may have been omitted. Other relevant demographic data, social determinants of health, treatment-related data (both approved and unapproved), and behavioral risk information may not have been present in the dataset. Such data are relevant because rates of HBV and HDV diagnosis and treatment can vary greatly with social determinants of health and payer types, and this variation may reflect different reasons for the disparity, for example, low treatment utilization may occur as a result of better health among the population or due to patients having worse access to care. There may also be uncertainty on how representative the patient sample was due to reported health disparities among patients with viral hepatitis in the United States in terms of prevalence and treatment initiation.28,29 The best-performing models were selected based on AUROC. However, different metrics may be appropriate for other applications, for example, clinical vs epidemiological use, and individual metrics may be prioritized for evaluation. Low prevalence of disease could also impact PPV and specificity. There is a lack of guidelines recommending double reflex testing of patients with HBV (ie, an anti-HDV test followed by testing for HDV RNA), which would limit the number of patients that are HDV-positive from the HBV patient pool.

There is increasing international consensus supporting universal screening for HDV among all HBsAg-positive patients, rather than relying solely on risk-based approaches. Clinical Practice Guidelines from the European Association for the Study of the Liver on hepatitis delta virus (2023) explicitly recommend that anti-HDV antibody testing be performed at least once in all HBsAg-positive individuals, with reflex HDV RNA testing in those who are antibody-positive (strong recommendation).30 In contrast, current American Association for the Study of Liver Diseases guidance historically recommends risk-based screening, which has been shown to miss a substantial proportion of HDV cases in real-world practice.13 Several US-based studies and expert reviews have demonstrated that risk-based screening fails to identify a meaningful fraction of HDV-infected patients, leading to delayed diagnosis and more advanced liver disease at presentation, and have therefore called for universal HDV screening in patients with chronic HBV infection.31,32

The findings of the present study further support this shift in screening strategy. Our results indicate that currently available demographic, clinical, and claims-based characteristics alone are insufficient to reliably predict which HBV patients harbor HDV infection, underscoring the limitations of selective screening approaches. In this context, universal HDV screening among HBV patients may be a more effective strategy to address underdiagnosis, align with emerging international guidance, and ensure timely identification and management of HDV infection while additional data are generated to refine risk stratification tools.

Conclusion

This study assessed various models in predicting the identification of HDV RNA-positive patients in claims databases. Identifying patients positive for HDV RNA using diagnosis-based methods (“any HDV” or “≥1IP/≥2OP”) was more accurate than algorithms using HDV-based variables from claims data. External validation using the TriNetX dataset showed similar performance to that obtained from internal validation. However, in real-world practice, diagnosis claims data may not yet be available, and so predictions would more likely need to rely on other patient information that is collected at an earlier stage, for example, demographic and HDV-based clinical characteristics such as hepatitis complications and medications. Using such algorithms consequently reduces performance, although discrimination (AUROC) may still be considered adequately high at approximately 80%. Further investigation is warranted on whether current levels of discriminative performance are sufficient in improving or prioritizing HDV diagnosis and treatment for high-risk patients. However, the limitations of such selective screening approaches indicate that universal HDV screening among HBV patients may be a more effective strategy to address underdiagnosis and ensure timely identification and management of HDV infection.

Acknowledgments

Editing and production assistance were provided by Maple Health Group and Arete Data Labs.

Footnotes

Conflicts of Interest: The authors disclose the following: Robert J. Wong reports research grants to his institution from Gilead Sciences, Inc, Exact Sciences, Durect Corporation, Theratechnologies, and Madrigal Pharmaceuticals and has served as a consultant to Gilead Sciences, Inc, Salix Pharmaceuticals, and Mallinckrodt Pharmaceuticals, all without compensation. Robert G. Gish reports grants from Gilead Sciences, Inc; serves as a consultant and advisor for Abbott, EIT Pharma Inc, Fujifilm Wako Diagnostics, Genlantis, Gerson Lehrman Group, Gilead Sciences, Inc, Helios, HepaTx, HepQuant, Intercept, Quest, Topography Health, and Venatorx; serves on scientific or clinical advisory boards for Genlantis, Gilead Sciences, Inc, Helios, HepaTx, HepQuant, Intercept, and Prodigy; serves as the chair of the clinical advisory board for Prodigy; is a partner in clinical trials alliance with Topography Health; serves on the data safety monitoring boards; is part of the speakers bureau for Gilead Sciences, Inc, and Intercept; is a minor stock shareholder in Riboscience and Cocrystal; and has stock options for Angiocrine, Eiger, Genlantis, HepaTx, and HepQuant. Ira M. Jacobson reports research funding to his institution from AbbVie, Assembly Biosciences, Bristol Myers Squibb, Eli Lilly, Enanta, Gilead Sciences, Inc, Intercept, Janssen, Merck, and Novo Nordisk; and receiving consulting fees from Aligos, Arbutus, Arrowhead, Assembly Biosciences, Galmed, Gilead Sciences, Inc, GSK, Intercept, Janssen, Merck, Roche, Takeda, and VBI Vaccines. Joseph K. Lim has served as a consultant for Gilead Sciences, Inc (without compensation). Thomas Debray and Scott McDonald report receipt of consulting fees from Gilead Sciences, Inc to their institution, Smart Data Analysis and Statistics, for statistical support. Marvin Rock, Gary Leung, and Chong Kim are employees of Gilead Sciences, Inc and may own stock in Gilead Sciences, Inc.

Funding: This study is funded by Gilead Sciences, Inc.

Ethical Statement: The study used de-identified data; institutional review board approval was not required for this analysis.

Data Transparency Statement: Not applicable.

Reporting Guidelines: Reporting Guidelines were not applicable for this article type.

Material associated with this article can be found, in the online version, at https://doi.org/10.1016/j.gastha.2026.101090.

Supplementary Materials

Supplementary Materials
mmc1.pdf (1.5MB, pdf)
Extended PDF
mmc2.pdf (5.9MB, pdf)

References

  • 1.Wranke A., Heidrich B., Deterding K., et al. Clinical long-term outcome of hepatitis D compared to hepatitis B monoinfection. Hepatol Int. 2023;17:1359–1367. doi: 10.1007/s12072-023-10575-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.World Health Organization . Hepatitis D. Switzerland: World Health Organization; Geneva: 2023. [Google Scholar]
  • 3.Miao Z., Zhang S., Ou X., et al. Estimating the global prevalence, disease progression, and clinical outcome of hepatitis delta virus infection. J Infect Dis. 2020;221:1677–1687. doi: 10.1093/infdis/jiz633. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Gish R.G., Wong R.J., Di Tanna G.L., et al. Association between hepatitis Delta virus with liver morbidity and mortality: a systematic literature review and meta-analysis. Hepatology. 2023;10:1097. doi: 10.1097/HEP.0000000000000642. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Lampertico P., Degasperi E., Sandmann L., et al. Hepatitis D virus infection: pathophysiology, epidemiology and treatment. Report from the first international Delta cure meeting 2022. JHEP Rep. 2023;5 doi: 10.1016/j.jhepr.2023.100818. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Pan C., Gish R., Jacobson I.M., et al. Diagnosis and management of hepatitis Delta virus infection. Dig Dis Sci. 2023;68:3237–3248. doi: 10.1007/s10620-023-07960-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Stockdale A.J., Kreuels B., Henrion M.Y.R., et al. The global prevalence of hepatitis D virus infection: systematic review and meta-analysis. J Hepatol. 2020;73:523–532. doi: 10.1016/j.jhep.2020.04.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Elsaid M.I., Mumtaz K., Li N., et al. 2023. Hepatitis delta virus superinfection increases the risk of hepatocellular carcinoma in patients with chronic hepatitis B virus. Presented at AASLD, The Liver Meeting, November 10-14, 2023, Boston. [Google Scholar]
  • 9.Wong R.J., Brosgart C., Wong S.S., et al. Estimating the prevalence of hepatitis delta virus infection among adults in the United States: a meta-analysis. Liver Int. 2024;44:1715–1734. doi: 10.1111/liv.15921. [DOI] [PubMed] [Google Scholar]
  • 10.Gish R.G., Jacobson I.M., Lim J.K., et al. Prevalence and characteristics of hepatitis delta virus infection in patients with hepatitis b in the United States: an analysis of the all-payer claims database. Hepatology. 2023;10:1097. doi: 10.1097/HEP.0000000000000687. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Tharwani A., Hamid S. Elimination of HDV: epidemiologic implications and public health perspectives. Liver Int. 2023;43:101–107. doi: 10.1111/liv.15579. [DOI] [PubMed] [Google Scholar]
  • 12.Papatheodoridi M., Papatheodoridis G.V. Is hepatitis delta underestimated? Liver Int. 2021;41(Suppl 1):38–44. doi: 10.1111/liv.14833. [DOI] [PubMed] [Google Scholar]
  • 13.Terrault N.A., Lok A.S., McMahon B.J., et al. Update on prevention, diagnosis, and treatment of chronic hepatitis B: AASLD 2018 hepatitis B guidance. Hepatology. 2018;67:1560–1599. doi: 10.1002/hep.29800. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Miao Z., Xie Z., Ren L., et al. Hepatitis D: advances and challenges. Chin Med J (Engl) 2022;135:767–773. doi: 10.1097/CM9.0000000000002011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Brnabic A., Hess L.M. Systematic literature review of machine learning methods used in the analysis of real-world data for patient-provider decision making. BMC Med Inform Dec Mak. 2021;21:1–19. doi: 10.1186/s12911-021-01403-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Denniston M.M., Jiles R.B., Drobeniuc J., et al. Chronic hepatitis C virus infection in the United States, national health and nutrition examination survey 2003 to 2010. Ann Intern Med. 2014;160:293–300. doi: 10.7326/M13-1133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Chen L., Goundan P., Bansal A., et al. 2023. Finding undiagnosed hepatitis C cases: using machine learning to identify clinical attributes and social determinants of health to improve screening efficiency and linkage to care. Presented at the EASL Congresss 2023, June 21-24, 2023, Vienna, Austria. [Google Scholar]
  • 18.Rigg J., Doyle O., McDonogh N., et al. Finding undiagnosed patients with hepatitis C virus: an application of machine learning to US ambulatory electronic medical records. BMJ Health Care Inform. 2023;30 doi: 10.1136/bmjhci-2022-100651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Garriga R., Mas J., Abraha S., et al. Machine learning model to predict mental health crises from electronic health records. Nat Med. 2022;28:1240–1248. doi: 10.1038/s41591-022-01811-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Johnson R., Stephens A.V., Mester R., et al. Electronic health record signatures identify undiagnosed patients with common variable immunodeficiency disease. Sci Transl Med. 2024;16 doi: 10.1126/scitranslmed.ade4510. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Zheng T., Xie W., Xu L., et al. A machine learning-based framework to identify type 2 diabetes through electronic health records. Int J Med Inform. 2017;97:120–127. doi: 10.1016/j.ijmedinf.2016.09.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Kuhn M. Building predictive models in R using the caret package. J Stat Softw. 2008;28:1–26. [Google Scholar]
  • 23.R: A Language and Environment for Statistical Computing. Version 4.5.0. R Foundation for Statistical Computing. 2025. https://www.R-project.org/
  • 24.Steyerberg E.W. Clinical prediction models: a practical approach to development, validation, and updating. Springer International Publishing; Cham: 2019. Study design for prediction modeling; pp. 37–58. [Google Scholar]
  • 25.Chen T., Guestrin C. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery; New York, NY: 2016. Xgboost: a scalable tree boosting system; pp. 785–794. [Google Scholar]
  • 26.Gish R.G. Delta hepatitis in the United States: epidemiology, testing, and linkage to care. Gastroenterol Hepatol (N Y) 2023;19:603–605. [PMC free article] [PubMed] [Google Scholar]
  • 27.Ali G., Mijwil M.M., Adamopoulos I., et al. Harnessing the potential of artificial intelligence in managing viral hepatitis. Mesopotamian J Big Data. 2024;2024:128–163. [Google Scholar]
  • 28.Kapadia S.N., Zhang H., Gonzalez C.J., et al. Hepatitis C treatment initiation among US medicaid enrollees. JAMA Netw Open. 2023;6 doi: 10.1001/jamanetworkopen.2023.27326. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Centers for Medicare & Medicaid Services . Office of Minority Health. Hepatitis Disparities in Medicare Fee-for-Service Beneficiaries. Centers for Medicare & Medicaid Services; Baltimore, MD: 2021. [Google Scholar]
  • 30.Brunetto M.R., Ricco G., Negro F., et al. EASL clinical practice guidelines on hepatitis delta virus. J Hepatol. 2023;79:433–460. doi: 10.1016/j.jhep.2023.05.001. [DOI] [PubMed] [Google Scholar]
  • 31.Nathani R., Leibowitz R., Giri D., et al. The Delta Delta: gaps in screening and patient assessment for hepatitis D virus infection. J Viral Hepatitis. 2023;30:195–200. doi: 10.1111/jvh.13779. [DOI] [PubMed] [Google Scholar]
  • 32.Cornberg M., Zoulim F., Gish R., et al. Best practices for screening, testing, diagnosing, and treating patients with hepatitis D (delta) virus based on global expert review and recent guidelines. Antivir Ther. 2025;30 doi: 10.1177/13596535251349380. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Materials
mmc1.pdf (1.5MB, pdf)
Extended PDF
mmc2.pdf (5.9MB, pdf)

Articles from Gastro Hep Advances are provided here courtesy of Elsevier

RESOURCES