Abstract
Clinical practice guidelines recommend hepatocellular cancer (HCC) surveillance in patients with cirrhosis from any etiology and those with chronic hepatitis B virus (HBV) infection and additional risk factors. However, HCC incidence varies across groups. Several risk stratification models using clinical factors and/or biomarkers have been derived to facilitate tailored HCC surveillance. Although risk stratification models are used for patients with hepatitis B, few have been sufficiently validated in patients with cirrhosis. Indeed, many unanswered questions related to the development, validation, and impact evaluation of risk stratification models must be addressed before widespread implementation can be recommended. The National Cancer Institute’s Translational Liver Cancer (TLC) Consortium was established to advance research focused on risk stratification and early detection of liver cancer. The TLC convened a multidisciplinary group, including clinicians, scientists, biostatisticians, and technology experts from the United States, Asia, and Europe, to provide a framework for the development, validation, and implementation of risk stratification models. The framework defines 4 phases of risk stratification model development and validation: phase 1—development and internal validation, phase 2—decision rule development, phase 3—external validation, and phase 4—impact evaluation. The group also defined a set of recommendations to improve the rigor of development and validation of HCC risk stratification strategies. This framework can inform best practices and highlight necessary steps for endorsement by practice guidelines and regulatory agencies, highlighting a path toward implementation in clinical practice.
Keywords: cirrhosis, guidance, hepatitis B, risk-stratified screening
BACKGROUND
Hepatocellular cancer (HCC) is one of the most common and lethal cancers in the world.[1] Chronic liver disease is the main precursor lesion for HCC, with over 90% of HCC occurring in this setting. There is considerable variability in HCC risk among patients with chronic hepatitis B virus (HBV) and those with cirrhosis, with HCC incidence ranging between 1% and 4% per year, and most patients do not develop HCC. Accurate risk stratification can promote shared decision-making between patients and providers by proposing a quantifiable personalized assessment of HCC risk and net benefit of HCC prevention programs, including HCC surveillance.[2]
The current standard of care includes surveillance with liver ultrasonography plus serum alpha-fetoprotein (AFP) every 6 months. Ultrasound is widely available, safe, cheap, and non-invasive; however, it is operator-dependent and has an inadequate sensitivity of 47%–63% for early-stage HCC in recent meta-analyses.[3,4] Alternative imaging modalities, such as multi-phase CT or MRI,[5,6] have higher sensitivity but are not recommended for routine surveillance given concerns related to limited radiology capacity and high costs if broadly applied to all at-risk patients. Accurate risk stratification could enable risk-based surveillance where different HCC surveillance strategies are recommended based on patients’ predicted risk.[7] This approach could also shift other resource-intensive efforts, such as outreach efforts to increase surveillance adherence as well as chemoprevention towards high-risk patients and reduce treatment intensity in low-risk patients.[2,8]
Given these potential advantages of risk-stratified surveillance, numerous HCC risk stratification models, including different combinations of clinical variables and/or biomarkers, have been developed.[9,10] However, the clinical adoption of HCC risk stratification strategies remains limited because of unresolved questions and barriers regarding sufficient validation and impact evaluation.
The National Cancer Institute (NCI)’s Translational Liver Cancer (TLC) Consortium was established to accelerate translational research aimed at improving risk stratification and early detection of liver cancer.[11] The consortium convened a multidisciplinary group including clinicians, scientists, biostatisticians, and technology experts from the United States, Asia, and Europe to create a framework for the development, validation, and implementation of HCC risk stratification models. Beyond defining best practices for HCC risk stratification model development, this framework enables a systematic, comparative evaluation of existing models to identify those most ready for clinical translation. Importantly, it can provide a clear roadmap toward guideline endorsement and regulatory approval, accelerating implementation of HCC risk stratification into routine clinical care.
OVERALL FRAMEWORK FOR RISK STRATIFICATION
The NCI Early Detection Research Network (EDRN) established a widely recognized 5-phase framework to guide the systematic discovery and validation of biomarkers for cancer detection and diagnosis.[12] This framework was subsequently adapted to account for some of the nuances specific to HCC, including appropriate selection of the target population.[13] Herein, TLC’s multidisciplinary expert panel developed a complementary 4-phase framework tailored for the development and validation of risk stratification models. We defined risk stratification models broadly—including clinical, molecular, and imaging-based variables; blood-based assays; a combination of assays with clinical and other variables; and artificial intelligence (AI) derived scores that are associated with the future risk of HCC. For the purposes of this document, we refer to these as risk stratification models henceforth.
The multidisciplinary expert panel was divided into 4 working groups, corresponding to each of the 4 phases. Each working group developed its section independently in a series of virtual meetings, reviewing relevant literature and combining their expert judgment with clinical data, iteratively revising and updating their respective sections throughout the process. The co-chairs (Fasiha Kanwal, Amit G. Singal) then reviewed and collated recommendations for each of the sections, focusing on areas of disagreement, reviewing additional literature, and discussing identified inconsistencies and controversial statements with individual members. The entire panel subsequently met and discussed the combined document in several dedicated virtual meetings, where they reviewed and discussed the range of opinions. After the co-chairs revised the framework based on this combined feedback, the updated final document was subjected to additional rounds of reviews or critiques by each member until consensus was achieved.
The 4-phase framework includes:
Phase 1: Development and internal validation—Initial identification and refinement of promising risk stratification models that warrant external validation.
Phase 2: Decision rule development—Establishment of a decision rule for the risk stratification models that defines predicted risk of future incident HCC, which can inform decision-making regarding clinical management, such as the most appropriate risk-based surveillance or chemopreventive interventions for specific risk strata.
Phase 3: External validation—Evaluation of model performance in independent longitudinal cohorts to demonstrate reproducibility and generalizability across populations.
Phase 4: Impact evaluation—Assessment of the impact of risk stratification on clinical decision-making, patient behavior, health (eg, early detection, receipt of curative treatments, survival benefit), and cost-effectiveness outcomes. This phase is essential for demonstrating real-world utility and justifying clinical implementation of the risk stratification model.
Each phase in the framework is accompanied by methodological and analytical recommendations to ensure rigor, reproducibility, and clinical relevance (Figure 1). Table 1 provides a comprehensive overview of existing HCC risk stratification models and how they map with the 4-phase framework.
FIGURE 1.

Phases of risk stratification markers/models development, validation, and implementation. Abbreviations: HCC, hepatocellular cancer; RCT, randomized controlled trial.
TABLE 1.
Risk-stratification models for hepatocellular cancer in patients with cirrhosis mapped to the 4-phase framework
| Model |
Phase 1: Development and internal validation |
Phase 2: Decision rule development |
Phase 3: External validation (locked thresholds) |
Phase 4: Impact evaluation |
|---|---|---|---|---|
| Prediction model of hepatocarcinogenesis in HCV cirrhosis Ikeda et al., J Hepatol 2006[14] | Completed. Retrospective cohort limited to patients with HCV-related cirrhosis; routinely available clinical variables; internal validation in a temporally distinct cohort; C-index not reported |
Limited. Risk strata defined through simulated 5-y and 10-y absolute HCC risk estimates across combinations of risk factors; no explicit thresholds linked to specific surveillance or prevention decisions |
Not completed. External cohorts were used to validate risk curves, but no pre-specified risk thresholds or decision rules were tested |
Not completed |
| PNPLA3-based HCC risk model in cirrhosis Guyot et al., J Hepatol 2013[15] | Completed. Prospective cohorts of alcoholic and HCV-related cirrhosis; clinical factors and PNPLA3 rs738409 genotype; C-index not reported |
Limited. Risk score used to define low-risk, intermediate-risk, and high-risk groups based on observed 6-year HCC incidence; cutoffs derived but not linked to specific surveillance or prevention decisions |
Not completed. Risk groups evaluated within derivation cohorts; no independent external validation of pre-specified thresholds |
Not completed |
| Cytokine gene variant–augmented model Tarhuni et al., J Hepatol 2014[16] |
Completed. Prospective cohort limited to HCV-related cirrhosis; clinical variables and TNFα-308 genotype; C-index not reported |
Limited. Three risk strata defined using data-driven cutoffs ( < 5%, 5%–15%, > 15% 5-y risk); thresholds not pre-specified for clinical decision-making |
Not completed. No external validation |
Not completed |
| AFP-adjusted laboratory model El-Serag et al., Gastroenterology 2014[17] |
Completed. Retrospective cirrhosis cohort; clinical factors including AFP; C-index 0.81 |
Limited. Only probability cutoffs (4%–30% 6-mo risk); no justified, population-specific actionable thresholds or benefit–harm analysis |
Not completed. No external validation of a locked decision rule; thresholds explored post hoc |
Not completed |
| ANRS CO12 CirVir nomogram Ganne-Carrie et al., Hepatology 2016[18] |
Completed. Prospective multicenter cohort limited to HCV cirrhosis; clinical factors; C-index 0.72 | Limited. Risk score and nomogram define low-risk, intermediate-risk, and high-risk groups based on observed cumulative incidence, from 0% to 30.1% at 5 y. Illustrative thresholds discussed, but not specified as clinical decision rules |
Not completed. No external validation of a locked decision rule; thresholds explored post hoc |
Not completed |
| 4-SNP hepatic fat polygenic risk score Thrift et al., PLOS ONE 2023[19] |
Completed. Prospective multicenter cohorts of U.S. patients with cirrhosis from multiple etiologies enrolled in 2 harmonized studies. A previously described 4-SNP polygenic risk score (PNPLA3, MBOAT7, TM6SF2, GCKR); C-index 0.58 |
Limited. Risk stratification was performed using tertiles of the PRS and not pre-specified as a clinically actionable threshold |
Not completed | Not completed |
| GALAD score for prediction of incident HCC Villa et al., Hepatology Communications 2023[20] |
Completed. Prospective cohort of patients across multiple etiologies and enrolled in a standardized 6-monthly surveillance program; C-index 0.72 | Limited. GALAD was evaluated as a continuous risk score and compared with aMAP and ALBI |
Not completed | Not completed |
| Nine-biomarker serum panel El-Serag et al. Gut 2024[21] |
Completed. Prospective cohort of patients with cirrhosis; 3 and 9 biomarkers with and without clinical factors; C-index = 0.71; with clinical base = −0.73 | Limited. The biomarker panels were evaluated as continuous risk scores and dichotomized at the median for illustrative risk stratification |
Not completed | Not completed |
| ADRESS-HCC Flemming et al., Cancer 2014[22] |
Completed. Large multicenter retrospective cohort; clinical factors; C-index 0.69 |
Limited. Single threshold (~4.67 ≈ 1.5% annual risk) conceptually linked to surveillance cost-effectiveness but without formal decision rule justification or benefit–harm analysis |
Partial. External cohorts evaluated model performance, but risk thresholds were not consistently locked or tested as fixed decision rules |
Not completed |
| HCC risk score in Asian cirrhosis Liang et al., Scientific Reports 2018[23] |
Completed. Multicenter cohorts of patients with cirrhosis, predominantly HBV-related and HCV-related; commonly available clinical variables; using cross-sectional Taiwanese cohort and independent Korean cohort; C-index 0.68 |
Limited. Risk thresholds defined empirically using score cutoffs (R < 0.5, 0.5–0.65, ≥ 0.65) corresponding to increasing short-term HCC incidence; thresholds proposed for risk stratification but not pre-specified as formal clinical decision rules |
Partial. Prospective validation performed using the same score cutoffs in additional cohorts; however, thresholds were derived within the same study and not locked before validation in fully independent cohorts |
Not completed |
| Deep learning HCC risk model in HBV Nam et al., JHEP Reports 2020[24] |
Completed. Retrospective cohort of patients with HBV-related cirrhosis receiving entecavir therapy at a tertiary center; a deep neural network was developed using routine clinical and laboratory variables; internal validation demonstrated good discrimination and calibration; C-index 0.72 |
Limited. Model outputs predicted probabilities of future HCC risk; an illustrative probability cutoff (0.5) was applied to define high-risk and low-risk groups, but no pre-specified, clinically justified risk thresholds or surveillance decision rules were defined |
Partial. Model performance was evaluated in an independent external cohort from a second tertiary hospital using the same trained model; however, risk thresholds were not pre-specified or locked before validation, and validation focused on discrimination rather than testing fixed decision rules |
Not completed |
| MRI-based radiomic model Wei et al., Frontiers in Oncology 2022 |
Completed. Retrospective cohort limited to patients with HBV-cirrhosis from 5 hospitals in China, all undergoing baseline MRI without evidence of HCC. A radiomic signature was derived from whole-liver multi-sequence MRI features with clinical variables (sex and Child–Turcotte–Pugh class); C-index 0.64–0.71 |
Limited. A single optimal Rad-score cutoff was empirically selected to dichotomize patients into high-risk and low-risk groups; the cutoff was statistically optimized within the derivation cohort rather than pre-specified based on clinical risk thresholds |
Partial. The trained radiomic model and Rad-score cutoff were applied to 2 additional hospital-based cohorts labeled as external validation cohorts. However, thresholds were derived within the same study and setting, and validation cohorts were not completely independent |
Not completed |
| ALD- and NAFLD-specific HCC risk models Ioannou et al., Journal of Hepatology 2019[25] |
Completed. Nationwide retrospective cohort limited to patients with alcohol-associated cirrhosis or MASLD-related cirrhosis receiving care in the VA healthcare system; C-index 0.75 |
Completed. Absolute annual and 5-y HCC risk estimates generated; explicit low-risk, medium-risk, and high-risk categories defined ( < 1%, 1%–3%, > 3% annual risk); decision-curve analysis used to evaluate net benefit across pre-specified screening thresholds |
Not completed. No independent external cohort tested pre-specified, locked thresholds or decision rules |
Not completed |
| HCC risk models after sustained virologic response to HCV therapy Ioannou et al., Journal of Hepatology 2018[26] |
Completed. Retrospective cohort of patients with HCV-related cirrhosis who achieved sustained virologic response within the VA healthcare system; post-SVR clinical and laboratory predictors; C-index 0.75–0.80 |
Completed. Absolute annual and multi-year HCC risk estimates reported; candidate low-risk, intermediate-risk, and high-risk categories proposed to reflect clinically meaningful post-SVR risk gradients |
Not completed. Although candidate risk strata were described, thresholds were derived empirically and validated only within the VA system; no independent external cohort tested pre-specified, locked thresholds, or decision rules |
Not completed |
| ASPAM-B score incorporating M2BPGi Chen et al., Cancers 2022[27] |
Completed. Retrospective cohort limited to patients with HBV-cirrhosis receiving entecavir or tenofovir in Taiwan; age, sex, platelet count, AFP, and M2BPGi measured at 12 mo of treatment; C-index 0.71 |
Completed. Empirically derived point thresholds ( ≤ 3.5, 4–7, > 7) were used to define low-risk, intermediate-risk, and high-risk groups based on observed cumulative incidence of HCC; thresholds were data-driven and optimized within the study and not pre-specified based on clinical decision criteria |
Not completed. Validation cohort drawn from the same underlying healthcare systems and study framework using random split-sample methods |
Not completed |
| Alcohol-associated Liver Cancer Estimation score Lee et al., Scientific Reports 2022[28] |
Completed. Retrospective hospital-based cohort limited to patients with alcoholic cirrhosis in Korea; age, AFP, and albumin; C-index not reported; discrimination 0.75 |
Completed. Explicit point-based cutoffs ( ≤ 60, > 60–100, ≥ 100) defined low-risk, intermediate-risk, and high-risk groups; thresholds discussed in relation to surveillance decision-making |
Not completed. Validation cohort drawn from the same underlying cohort |
Not completed |
| AFDA score: ALBI–FIB-4–Diabetes–AFP Wang et al. American Journal of Cancer Research 2023[29] |
Completed. Retrospective cohort of patients with hepatitis B virus–related compensated cirrhosis receiving entecavir or tenofovir therapy at 3 tertiary centers in Taiwan; C-index 0.68 |
Completed. The AFDA score was converted into clinically interpretable risk categories. Patients were stratified into 3 groups based on total score: low risk (0 points), intermediate risk (1–3 points), and high risk (4–6 points). The low-risk group (score 0; 16.1% of the cohort) demonstrated a 5-year cumulative HCC incidence of 3.4% |
Not completed | Not completed |
| AMA score: age, M2BPGi, AFP Chen et al. American Journal of Cancer Research 2024[30] |
Completed. Retrospective cohort of patients with chronic hepatitis B–related cirrhosis receiving long-term entecavir or tenofovir therapy in Taiwan; C-index 0.71 |
Completed. Pre-specified cutoffs were defined to categorize patients into low-risk ( ≤ 2), intermediate-risk (2–5), and high-risk ( ≥ 5) groups | Not completed | Not completed |
| Genetic variants and clinical models Nahon et al. J Hepatol 2023[31] |
Completed. Prospective multicenter cohorts of patients with compensated cirrhosis included in standardized HCC surveillance programs (ANRS CO12 CirVir and CIRRAL cohorts), limited to alcohol-associated cirrhosis and/or cured HCV infection. Clinical variables, with and without 6- and 7-SNP genetic risk scores; C-index 0.78 (aMAP 0.76) |
Completed. Genetic risk scores stratified patients into low-risk, intermediate-risk, and high-risk groups based on empirically defined score categories and illustrative percentile-based cutoffs (eg, 70th, 80th, and 90th percentiles). Decision-curve analyses evaluated net benefit across a range of hypothetical HCC risk thresholds relevant to surveillance |
Not completed | Not completed |
| Toronto HCC Risk Index, THRI Sharma et al., Journal of Hepatology 2018[32] |
Completed. Retrospective clinic-based cohort of patients with cirrhosis across multiple etiologies; clinical factors; C-index 0.77 |
Completed. Pre-specified point thresholds ( < 120, 120–240, > 240) defined low-risk, intermediate-risk, and high-risk groups with separated absolute risk (~0.3%, 1.0%, and 3.2% annual HCC incidence); risk strata discussed in relation to cost-effectiveness thresholds for surveillance and potential risk-based surveillance |
Completed. External validation performed in an independent multinational cohort using the same pre-specified point thresholds and risk strata |
Not completed |
| aMAP score for prediction of HCC. Fan et al., J Hepatol 2020[33] |
Completed. Derivation was performed in international prospective studies (training cohort described as 3688 treated Asian CHB patients), with model performance assessed using discrimination and calibration for HCC development over 3–5 y; C-index 0.82 (0.70–0.76 in external validation studies) | Completed. Two cutoffs to define 3 clinically meaningful risk strata: low risk (0–50), intermediate risk (50–60), and high risk (60–100). These cutoffs are explicitly linked to differing HCC incidence rates to support risk-stratified surveillance strategies |
Completed. External validation was performed in multiple cohorts spanning diverse etiologies, geographic regions, and healthcare settings. The pre-specified aMAP cutoffs (0–50, 50–60, 60–100) applied unchanged across validation cohorts, with preserved discrimination, calibration | Not completed |
| THCC-RI (Texas HCC Risk Index) Kanwal et al., CGH 2023[34] |
Completed. Prospective multicenter cohort with multiple etiologies. Clinical factors including AFP; C-index 0.77 (0.70–0.76 in external validation cohorts) |
Completed. Empirically derived risk thresholds are defined based on the distribution of predicted risk. Patients were stratified into low-risk ( < 20th percentile), intermediate-risk (21st–79th percentile), and high-risk ( > 80th percentile) groups. Absolute cumulative HCC risk was reported at 1, 2, and 3 y by decile and by risk group. In the derivation cohort, the high-risk group demonstrated ~12% cumulative incidence at 2 y and 19% at 3 y, while the low-risk group had near-zero events at 1–2 y. Thresholds proposed to inform potential risk-stratified surveillance strategies |
Completed. External validation in 3 independent cohorts evaluating the performance of pre-specified risk strata, with a locked model and consistent performance |
Not completed |
| PAaM score: integration of PLSec-AFP with aMAP Fujiwara et al. Gastroenterology 2025[35] |
Completed. Prospective cohort of patients with cirrhosis from mixed etiologies enrolled consecutively at the University of Michigan Hospital. PAaM was derived by integrating a molecular score (PLSec-AFP) with aMAP; C-index 0.74 (0.66 and 0.67 in external validation cohorts) | Completed. The low-risk cutoff was set to ensure annual HCC incidence < 1%, the high-risk cutoff was set to maximize annual incidence while maintaining a clinically meaningful prevalence of 20%–30%. In the derivation cohort, the 25th and 75th percentile PAaM cutoffs (4.318 and 5.072) met these criteria and were used for subsequent validations. In the derivation cohort, annual incidence was 0.5%, 3.1%, and 5.9% in low-risk, intermediate-risk, and high-risk groups, respectively |
Completed. External validation was performed in 2 independent prospective multicenter cohorts. In HEDS, PAaM classified 44% as low risk, 41% intermediate, and 15% high risk, with annual incidence 0.8%, 1.8%, and 6.2%, respectively | Not completed |
Notes: The models are rank-ordered based on progression across the 4 phases (from early phase to late phase development). The ranking is based on the checklist in Figure 1. It is not based on the quality of each phase, which is beyond the scope of the manuscript.
Abbreviations: AFP, alpha-fetoprotein; ALBI, Albumin–bilirubin; ALD, alcohol-associated liver disease; FIB-4, fibrosis-4 index; HBV, hepatitis B virus; HCC, hepatocellular cancer; HCV, hepatitis C virus; MASLD, metabolic dysfunction–associated steatotic liver disease; MRI, magnetic resonance imaging; M2BPGi, Mac-2 binding protein glycosylation isomer; NAFLD, nonalcoholic fatty liver disease; SNP, single-nucleotide polymorphism; SVR, sustained virologic response; THRI, Toronto HCC Risk Index; TNFα, tumor necrosis factor alpha; VA, Veterans Affairs.
Phase 1: Development and internal validation
Phase 1 aims to identify promising candidate risk stratification models with preliminary performance estimates that support their potential clinical validity. Model development for HCC should follow broader recommendations for the initial development and reporting of risk prediction models.[36-38] A full description of these recommendations is out of scope, but we outline guidance relevant to HCC risk stratification across the four domains: patient selection, predictor selection, outcome definition, and analysis.
Development and internal validation of the risk model should be performed in the appropriate target patient population, specified by well-defined criteria. For HCC, the appropriate target population is typically patients with cirrhosis or chronic HBV infection. Given data showing that one-fourth of HCC among patients with metabolic dysfunction–associated steatotic liver disease (MASLD) occur in the absence of cirrhosis,[39] this is another potential patient population to consider; however, these patients should be evaluated separately from those with cirrhosis (or those with HBV), as HCC risk and associated risk factors likely differ in these populations.
HCC risk models can be developed using a variety of clinical, molecular, and imaging-based variables (Table 1). Most existing risk stratification models include demographic and clinical risk factors such as age, sex, etiology, and severity of chronic liver disease, metabolic risk factors, and AFP levels. Multi-dimensional risk stratification models, including other domains, such as genetic, molecular, and imaging biomarkers, could potentially improve the accuracy of clinical variable-based risk models.[9,10,35] Regardless of the number and types of predictors included in the model, the predictors should be defined and assessed in the same way for all patients, without knowledge of the outcome. Given that several risk stratification models derived from routine parameters exist, model developers could enrich existing models with newly identified markers instead of building new models.
The initial development and internal validation of risk models is ideally performed using data and/or samples from a prospective cohort of at-risk patients.[40] This evaluation can be performed using data from the full cohort or using a case-cohort design, in which the risk model is assessed at baseline, cases are patients who develop HCC during subsequent follow-up, and controls are patients who remain HCC-free during equivalent or longer follow-up (or a random sample of controls in case-cohort design). Many HCC risk stratification studies use retrospectively developed cohorts (Table 1), but inherent biases in retrospective cohorts (eg, unmeasured or poorly measured con-founders) can result in inaccurate estimates of model performance.
Although case-control studies are often used for the identification of promising early detection biomarkers, this design is not optimal for risk stratification models. Methods used to select controls often result in a selection bias compared with the true target at-risk population. Just as in other cancers, biomarkers and clinical data collected from cases can be biased by the presence of active cancer and may not accurately capture signals that truly predict future risk. Inclusion of non-target populations, such as healthy individuals, especially poses a risk of confounding in studies predicting HCC risk, as performance of risk models could be overestimated because they could be predicting the presence of chronic liver disease instead of predicting the future risk of HCC. Despite these methodological limitations, the case-control design may still be useful in some instances. For example, a lack of association—or only a weak association—in a case-control study strongly suggests that the marker is unlikely to have predictive value in prospective cohort studies, precluding the need for larger, more expensive prospective studies. However, because case-control studies can overestimate associations, positive findings should be interpreted with caution and confirmed in prospective studies of cohorts at risk of HCC.
The outcome (ie, diagnosis) of incident HCC should be defined clearly and consistently to allow replication in external validation studies and future application in clinical practice. The diagnosis of HCC is typically based on histological confirmation or characteristic imaging appearance on contrast-enhanced MRI or multi-phase CT.[41] Notably, the time frame for HCC risk prediction is critical to consider, as patients who develop HCC within 1–2 years of the baseline could have a risk of subclinical, prevalent HCC. Inclusion of these patients in model building can overestimate model performance or bias the predictor selection. Although a 1-year exclusion window may be sufficient, given average tumor doubling times, a sensitivity analysis excluding HCC cases that develop within 2 years of baseline could account for indolent tumors with longer doubling times, as well as early-stage HCC that may be missed due to the limited sensitivity of HCC surveillance tests, that is, ultrasound.[42,43] While these patients could be excluded for model building, they can be included in assessments of model performance to estimate clinical utility in the target patient population.
As with model development in other settings, candidate risk prediction should be evaluated using statistical measures of predictive associations to identify independent predictors, and preliminary assessments of performance metrics such as discrimination, internal calibration, and potential informativeness for clinical decision-making. These metrics should then be re-assessed during internal validation (see below).
Performance metrics for predictive associations can include hazard ratio (HR) in univariable and multivariable Cox regression. When building a regression model for incident HCC, it is important to consider competing events, such as death and liver transplantation, utilizing appropriate methods like Fine–Gray regression to calculate sub-distribution HR with or without adjustment for confounding variables. Indeed, many factors associated with increased risk of HCC are also associated with increased risk of hepatic decompensation and liver-related mortality, so failure to account for these (as in Kaplan–Meier estimator) may inflate the cumulative incidence of HCC, generating biased results.
Discrimination is the extent to which a risk stratification model can differentiate between patients who develop the outcome and those who do not. Discrimination can be assessed using measures like the Harrell concordance index (C-index) and time-specified or dependent AUROC. A model with good discrimination can be used to separate patients at low risk from those at high risk for HCC. A C-index or an area under the ROC curve of < 0.60 reflects poor discrimination; 0.60–0.75, possibly helpful discrimination; and more than 0.75, useful discrimination.[44] Most published HCC risk stratification models demonstrate discrimination in the 0.60–0.75 range, although several achieve values > 0.75, supporting their promise for HCC surveillance programs (Table 1). Notably, prediction models with similar or even lower discriminatory performance—such as the Framingham cardiovascular risk score (C-statistic ~0.63–0.83),[45] the Gail model for breast cancer risk (C-statistic ~0.60),[46] and several lung cancer risk models (C-statistics ~0.65–0.75)[47]—are widely used to inform clinical decision-making across different clinical contexts.
Internal calibration is the model’s ability to estimate the correct absolute risk and assess how well the absolute predicted risks correspond to observed incidence rates. The accuracy of quantitative predictions of risk can be assessed through methods such as calibration plots and the Integrated Brier Score. The lower the Integrated Brier Score, the better the prediction. Rigorous assessment of calibration is essential in the development and validation of HCC risk prediction models,[48] and only well-calibrated models should be used to communicate absolute risk estimates to patients and clinicians. Preliminary informativeness for clinical decision-making is determined by measures such as positive predictive value (PPV) for high-risk predictions and negative predictive value (NPV) for low-risk predictions.
Internal validation is conducted within the same cohort used for initial model development and provides estimates of model performance, including discrimination and calibration. Two common threats can under-mine the reliability of internal validation. These include overfitting, where the model yields overly optimistic results, and unreliable estimates of effect size due to small or imbalanced groups with and without the outcome (incident HCC diagnosis). The latter is a common problem in HCC due to the relatively low incidence of HCC in most prospective studies. To mitigate the risk of overfitting, various analytical methodologies, including k-fold cross-validation, leave-one-out cross-validation, and bootstrapping, may be employed, defined upfront, and conducted once using the best-performing model. Strategies such as penalized regression, Bayesian approaches, or data pooling across cohorts can help mitigate issues related to small samples.
During internal validation, new models can be compared against existing models, and improvement in risk prediction can be gauged using measures of net benefit, including difference in Integrated Brier Score and C-index, misclassification tables, and standardized net benefit. However, it is inappropriate to conclude superiority given the risk of overfitting when the new model is derived in the cohort used for this comparative evaluation.
Based on the results of internal validation, algorithms for the risk score (and analytic parameters, including choice of analytes and detection methods for bio-assays) can be optimized in patients before external validation.
Phase 2: Developing decision rules
Validated models should yield absolute estimates of HCC incidence, which can be used to estimate HCC risk in individual patients given their specific risk factor profile. Although HCC risk is a continuum, categorizing risk into discrete strata using clinically actionable thresholds is necessary to inform clinical decision-making, which is inherently categorical (eg, initiate, intensify, or de-escalate surveillance). A risk-stratified approach to surveillance can allow tailored surveillance strategies where high predicted risk translates into more intensive surveillance (eg, ultrasound + AFP vs. novel imaging such as abbreviated MRI), while lower predicted risk may support less intensive surveillance or, in some patients, discontinuation of surveillance to minimize surveillance-related harms. Beyond the risk thresholds, absolute predicted risk can further inform shared decision-making between patients and clinicians. This hybrid approach—integrating actionable thresholds with continuous risk estimates—is consistent with common clinical practice. For example, although PSA > 4 ng/mL is commonly used as a threshold for prostate biopsy decision, clinicians also look at the absolute values and trajectory over time to make an individualized decision.
Determining these potentially actionable HCC risk thresholds requires inputs from stakeholders—patients, clinicians, public health policymakers, and health economists—based on multidimensional considerations, including not only the model performance but also treatment options, population heterogeneity, patient preferences, and available resources. Below, we provide our recommended approaches that can be employed to identify thresholds for HCC risk stratification. These can then serve as starting points for further input from patients, clinicians, and other stakeholders.
For HCC surveillance, decision analyses can evaluate risk thresholds above which HCC surveillance is cost-effective, considering both the benefits and harms (including unnecessary diagnostic follow-up exams, physical/psychological harms, and direct and indirect costs) of surveillance at different risk thresholds.[49-51] To date, few cost-effectiveness studies have provided data to inform potentially actionable HCC risk thresholds that could guide surveillance strategies, including the cutoffs for intensifying surveillance (the high-risk cutoff) and cutoffs for discontinuing surveillance (the low-risk cutoff). For example, a cost-effectiveness analysis of patients with compensated cirrhosis enrolled in 4 multicenter prospective cohorts (median follow-up 37 mo; annual HCC incidence 2.3%) found that MRI-based surveillance was cost-effective when the baseline annual incidence of HCC was ≥ 3%.[52] Of note, the high cutoff of > 3% annual HCC incidence threshold was selected a priori based on clinical judgment in this study. Another modeling study provides data to support a low-risk cutoff. Using a dedicated threshold analysis, Parikh et al.[53] showed that compared with no screening, screening with ultrasound and AFP became cost-effective at the willingness-to-pay threshold of $100,000/QALY only when the annual incidence of HCC exceeded 0.4%, suggesting that screening may not be cost-effective in patients with an annual HCC risk below 0.4%. Another study found that HCC surveillance in hepatitis C-cured persons with cirrhosis is cost-effective above the annual risk of 0.7%.[54]
These findings highlight the need for additional decision analyses that integrate HCC incidence and natural history, life expectancy, available resources, patient preferences, and costs to establish clinically actionable thresholds and corresponding surveillance strategies. While such data continue to emerge, the following risk thresholds may serve as reasonable starting points for shared decision-making: annual HCC risk > 3% to consider intensified surveillance, 0.4%–3% to continue ultrasound plus AFP, and < 0.4% to consider discontinuing surveillance.
Integrated risk predictiveness curves offer a complementary approach by visually examining model performance across the full risk spectrum and illustrating the potential impact of different decision rules[55] The predictiveness curve plots the distribution of risks in a population from which the cohort was drawn based on the variables and information in the risk model. Figure 2 shows an integrated predictiveness curve using data from the Texas HCC Consortium Cohort (THCCC) to illustrate how this approach can be used to evaluate the impact of selecting different risk thresholds within a given model.[34] Figure 2A displays predicted HCC risk in the cohort from the lowest to the highest percentile based on the THCC-Risk Index. At the 70th percentile on the x-axis, the 1-year HCC risk value is 3.3%—this risk corresponds to the high cutoff from cost-effectiveness models above.[31] Based on the risk model, 30% of the cohort has a calculated risk above 3%. If the annual HCC risk threshold of “ > 3%” is used to intensify surveillance (such as MRI-based surveillance), then ~30% of the population would shift to MRI-based surveillance. Approximately, 10% of the cohort have estimated 1-year HCC risk at or above 5%.[7] While both 3% and 5% thresholds identify higher-risk sub-groups, the 3% cutoff may be preferred because a 5% threshold would reclassify roughly two-thirds of higher-risk patients into the reduced-surveillance group. At the 3% threshold, the true positive fraction exceeds 70% (Figure 2B), meaning that over 70% of HCC cases arise within the 30% of patients whose annual risk is ≥ 3%. By contrast, applying a 5% threshold reduces the true positive fraction to 40%, indicating substantially fewer HCC cases would be captured within the intensified surveillance group. This example illustrates how integrated predictiveness curves can support transparent, data-driven selection of risk thresholds by balancing the proportion of individuals exceeding the risk threshold/s with the proportion of HCC cases captured within the high-risk group, thereby facilitating more objective development of clinically actionable decision rules in phase 2 of the framework.
FIGURE 2.

Integrated predictiveness curve illustrating HCC risk thresholds for clinical decision-making. Panel (A) uses data from the Texas HCC Consortium Cohort (THCCC) and displays the predictiveness of the THCC-Risk Index by displaying predicted HCC risk in the cohort from the lowest to the highest percentile based on the THCC-Risk Index.[52] Panel (B) displays the classification of the THCC-Risk Index. Abbreviations: FPF, false-positive fraction; HCC, hepatocellular cancer; TPF, true-positive fraction.
Few points are noteworthy. Based on published data, providers are generally receptive to risk-stratified indication of surveillance.[56] However, they are more comfortable intensifying surveillance for high-risk patients than reducing or stopping it for those at low risk.[56] Similarly, patients and other stakeholders tend to prioritize the benefits of surveillance over potential harms.[57,58] As a result, the clinical utility of risk models may be greater in distinguishing intermediate-risk from high-risk individuals than in identifying those at low risk, where the impact on clinical decision-making is likely more limited.
HCC risk thresholds may be different for different surveillance tests, different at-risk populations, and different geographic areas; they may also change over time as they depend on the effectiveness, costs, and harms of emerging treatments for HCC.[59] For example, the performance characteristics (true positive, true negative, false positive, false negative), costs, and harms of a surveillance test affect surveillance thresholds in cost–benefit analyses. If an inexpensive, blood-based HCC biomarker became available with excellent performance characteristics (including both sensitivity and specificity), this would potentially lower the HCC risk threshold at which surveillance with this biomarker would become cost-effective.[60] Performance characteristics of a surveillance test also vary by disease [eg, ultrasound performs worse in the setting of obesity/steatosis that characterizes metabolic dysfunction–associated steatohepatitis (MASH)][61] or because the effectiveness of available treatments may be different (eg, higher potential for HCC cure and long-term survival in a patient with non-cirrhotic HBV than in a patient with advanced cirrhosis). The costs of a surveillance test and HCC treatments are different across regions and countries, which affects cost-effectiveness estimates. In the future, it is more likely that precision surveillance will require different thresholds for different populations/settings/surveillance tests, but the approach discussed here still applies. This approach can also extend to decisions around chemoprevention and outreach interventions.
Phase 3: External validation
External validation requires testing of the risk model in an independent validation cohort that is sufficiently distinct from the development or internal validation cohort. It aims to identify and confirm promising risk models that have robust performance in general target populations to warrant further evaluations of clinical utility, impact, and scalability. For external validation, the cohort should be a prospective cohort representing the same target population, the model and preferably the decision rule should be locked-down, and the predictors should be measured without knowledge of the outcome status. The external validation cohort should be sufficiently large (eg, > 1000 to ensure sufficient statistical power and patient demographics representing the target population) and have an adequate duration of follow-up (eg, > 3-year median follow-up) to facilitate enough outcomes of interest (eg, > 100 incident HCC cases). Cohorts of sufficient size can also facilitate subgroup analyses to determine if model performance is consistent across patient factors of particular interest (eg, age, gender, race and ethnicity, liver disease etiology). These cohorts can be selected from prospectively developed large administrative datasets (eg, national Veterans Affairs database) or cohorts of patients included in HCC surveillance programs from individual health systems. Cohorts used for external validation should include the required components of the risk model, ideally measured in a similar manner as the initial development cohort. External validation cohorts should also include assessments for relevant confounders, adherence to surveillance (to assess risk of ascertainment bias), and competing clinical outcomes (eg, liver transplantation and death).
Recruitment from different regions or different types of healthcare settings can be advantageous, as this can evaluate transportability. Transportability assesses if model performance can be replicated in different populations, whereas reproducibility assesses if performance can be replicated in populations that more closely resemble the derivation cohort. Evaluation of risk models in diverse patient populations is particularly important for HCC risk stratification given regional differences in liver disease epidemiology (viral vs. non-viral etiologies), HCC risk factors (eg, metabolic syndrome and alcohol use), and surveillance patterns (contributing to differential missing data mechanisms and risk of ascertainment bias).[62] Notably, reproducibility may be demonstrated in a single external validation cohort, whereas transportability may require evaluation in several external validation studies. For example, risk models initially derived in populations with untreated HBV or HCV infection would require external validation in contemporary populations with hepatitis, based on changes in anti-viral treatment availability and practice patterns resulting in higher proportions of patients with suppressed HBV or cured HCV infections.[63]
External validation is a critical step, as model performance is often poorer than reported in derivation and internal validation cohorts. However, only a small proportion of available risk models progress to external validation, especially using the thresholds identified in phase 2 (Table 1). Specifically, of the 23 included models, only 4 have progressed to phase 3. These include Toronto HCC Risk Index, aMAP score, Texas HCC Risk Index, and PAaM score, suggesting that these models may be candidates for last-phase testing. Data in Table 1 also underscore the importance of external validation as a critical step toward implementation.
The evaluation of risk models during external validation uses similar analytic techniques as internal validation, including assessments of discrimination and calibration. The risk predictiveness generated by the locked-down model score could be plotted in the external cohort data to assess its performance and potential impact for decision-making. Whereas optimization of a risk model can be performed after internal validation, this process is discouraged during external validation, as the model, including component weights, should be locked before external validation. If model optimization is performed at this phase, external validation in a separate cohort would be necessary to determine transportability and reproducibility.
As discussed above, cohorts used for internal validation of a new risk score can concurrently be used for external validation of existing risk models; however, it is not appropriate to conclude that the new model has superior performance given the risk of overfitting. Conversely, comparisons between 2 existing models in external validation cohorts can provide important insights regarding which model is best suited to move forward for further consideration in clinical practice.
There is increasing interest in AI-based risk stratification models.[64,65] Although data sharing of large amounts of data, with or without protected health information, historically made validation of AI models difficult, recent techniques, including federated learning, have mitigated some of these challenges. The general principles of external validation apply equally to AI-based models; however, AI models leveraging large amounts of input data have some unique complexities. Most notably, AI methods, especially those based on imaging, can be influenced by spurious correlations or exhibit limited robustness to data variations.[66,67] AI-based biomarkers can include established biomarkers (class A), indirect measures of known biomarkers (class B), or novel AI-derived biomarkers (class C).[68] The potential for spurious associations and requirements for validation increase from class A to class C AI-based models. In addition to the validation steps described above, it is also essential to evaluate the performance of AI-based models across subgroups introducing potential data variation, such as equipment manufacturer and acquisition protocols.
Phase 4: Impact evaluation—Modeling and prospective studies
Impact evaluation, also referred to as clinical utility, goes beyond clinical validity and predictive accuracy. Clinical utility is a multidimensional construct and its evaluation requires understanding the benefits, harms, accessibility, and acceptability of risk driven clinical decisions and/or interventions.[69,70] The goal of impact evaluation is to address questions such as whether the use of risk stratification using a specific model improves patient outcomes, provides a favorable balance of harms versus benefits, offers a cost-effective solution, and significantly alters clinical decision-making.
Randomized controlled trials (RCTs) remain the most rigorous approach for assessing the clinical utility of risk assessment and risk-based surveillance. There are currently no RCTs for risk-stratified surveillance in HCC (Table 1). However, 2 RCTs are currently evaluating the effectiveness of risk-stratified surveillance for breast cancer and can provide the blueprint for the impact evaluation of risk-stratified surveillance for HCC. WISDOM is a pragmatic, adaptive, RCT in women aged 40–74 years[71] who are being stratified into 4 risk groups: highest risk, elevated risk, average risk, and lowest risk based on a clinical risk prediction tool enhanced with a polygenic risk score. Based on risk, the surveillance strategies vary in starting age, frequency, and modality of surveillance, ranging from annual mammography with MRI in the highest risk group to deferred surveillance until the age of 50 years in the lowest risk group. MyPeBS[72] is a pragmatic RCT that is underway in 5 countries (UK, France, Belgium, Israel, Italy) using a similar pragmatic design and endpoints as WISDOM.
For HCC, two ongoing RCTs are examining the effectiveness of different HCC surveillance modalities and will shed some light on the role of HCC risk stratification. FASTRAK trial (FAST-MRI for HCC suRveillance in pAtients with high risK of liver cancer) is a multicentre, French RCT comparing surveillance using semi-annual ultrasound and abbreviated MRI versus semi-annual ultrasound alone in patients at high risk for HCC (annual HCC incidence > 3%).[73] This RCT applies an HCC risk stratification model to identify patients with an annual HCC incidence over 3% for more precise inclusion and a lower required sample size.[52] The PREMIUM trial (PREventing liver cancer Mortality through Imaging with Ultrasound vs. MRI) is an RCT of ultrasound plus AFP versus abbreviated MRI plus AFP among patients with cirrhosis who have a high risk of HCC. PREMIUM uses a number of eligibility criteria to identify participants (eg, FIB-4 score > 3.25, estimated annual HCC risk > 2.5% based on a VA-validated model, or presence of clinically significant portal hypertension using non-invasive vibration-controlled transient elastography and platelet criteria) to ensure that the study is adequately powered.[74,75] Data from these trials will provide information on the benefits of HCC surveillance in high-risk patients, but will not directly examine the impact of risk- based surveillance. Other trials, such as the National Liver Cancer Screening Trial (TRACER), plan to conduct secondary analyses to evaluate differential benefits of surveillance strategies across risk strata to inform risk-stratified surveillance approaches.[76]
RCTs provide the strongest evidence of efficacy for risk-stratified strategies. However, RCTs have limitations. They require large sample sizes, long duration, and intensive follow-up, rendering them expensive to conduct. For instance, the two breast cancer trials referenced above will recruit between 85,000 and 100,000 women across multiple centers and countries and will extend over 10 years. They also required a multi-year stakeholder engagement process that brought together consumers, advocates, primary care physicians, specialists, policymakers, technology companies, and payers, underscoring the large investments needed to conduct such studies.
Risk-stratified surveillance RCTs for HCC will need a multidisciplinary approach with engagement of multiple stakeholders to ensure a systems approach to its implementation in real-world settings. Hybrid effectiveness–implementation research studies, especially those that use adaptive designs, can reduce the time lag between the generation of evidence on the effectiveness of a program and its implementation. Other alternative trial designs may address the limitations of traditional large-scale RCT including emulated clinical trials within a large, ongoing prospective cohort, adaptive trials, use of synthetic control arms, cluster-randomized design, Zelen design or random invitation single-arm trial.[77,78] Future approaches for RCTs can also benefit from AI to streamline candidate identification, thereby enhancing trial efficiency and reducing costs.[79,80] AI can potentially help refine patient selection, optimize surveillance intervals, and improve prediction of clinical endpoints. AI-driven analytics can also offer potential enhancements to biomarker evaluation and clinical integration.
Mathematical modeling studies can also play a crucial role in studying long-term outcomes (eg, mortality), harms, and cost-effectiveness of various risk-stratified surveillance strategies. Modeling enables projection of surrogate endpoints, like early detection rates, into meaningful clinical outcomes such as improved survival or quality-adjusted life years (QALYs). Economic evaluations integrated within these models can highlight cost-effectiveness, optimize resource allocation, and guide health policy decisions. Modeling studies can precede and inform RCTs of surveillance interventions. Modeling studies can also guide population-surveillance policies by extrapolating evidence beyond the time horizon of RCTs. Models that account for regulatory differences between regions (such as the European Union and the United States) can be especially beneficial to guide subsequent studies or policy. Modeling has been advocated to augment the EDRN 5-phase guidelines for cancer early detection.[81] While modeling studies can inform key data gaps, they have limitations given their reliance on many assumptions, which may overlook or underestimate critical confounding factors. A remedy is to use an array of models from different modeling groups used different assumptions and check the robustness of the conclusions.
Impact evaluations must prioritize outcomes that are clinically meaningful but also feasible in the setting of a research program. Primary outcomes can include benefits (reduction in late-stage HCC detection rate, HCC-specific deaths prevented, health-related quality of life gained), harms (reduction in unnecessary diagnostic tests or liver biopsies due to false-positive results, overdiagnosis, and harms of associated treatments), and cost-effectiveness. A reduction in the proportion of late-stage HCC represents a good surrogate for overall survival, given the lower risk of overdiagnosis. Overdiagnosis relates to detection of indolent HCC or tumors that are unlikely to impact survival given high competing risks of mortality.[82] From a societal standpoint, the main objective of these studies should likely be to evaluate the cost per QALY. Secondary outcomes of importance could be patient adherence to recommended surveillance protocols and reduction in healthcare resource utilization.
SCALABILITY AND IMPLEMENTATION
While predictive accuracy and internal validity are essential for evaluating the potential clinical utility of a risk model, scalability—the ability to implement the model broadly and sustainably across diverse settings—is equally important for forecasting its clinical and population-level impact. Scalability depends on a range of interrelated factors, including:
Generalizability across populations and settings. Models that have been externally validated across diverse populations, geographic regions, healthcare settings, and disease subtypes demonstrate greater robustness and adaptability. As noted earlier, generalizability is particularly important for HCC risk stratification given differences in liver disease epidemiology and care patterns.
Model complexity. Simpler HCC models that rely on routinely collected clinical variables (eg, age, platelet counts, liver enzymes, ultrasound findings) are easier to apply within existing cirrhosis care work-flows and are more likely to scale. In contrast, models requiring advanced omics (proteomics, genomics), specialized biomarkers, may be less amenable to widespread implementation, especially in low-resource or publicly funded healthcare systems. Model developers should consider examining the incremental predictive value of advanced biomarkers compared with readily available clinical measures when considering clinical translation.
Cost of deployment and integration. For HCC, models that require minimal training, provide clear risk stratification and guidance, and integrate into routine liver clinic and electronic health record workflow are more scalable than those requiring specialized expertise, certification, or frequent technical support.
Regulatory, ethical, and legal considerations. Scalable models must comply with data protection regulations (eg, HIPAA, GDPR), ethical standards for transparency and equity, and avoid algorithmic biases, particularly given known disparities in HCC burden and outcomes across groups. Securing regulatory approvals, including coding for reimbursement, and ensuring adherence to local and international liver society guidelines, is essential for broad implementation.
By proactively addressing these considerations during model development and validation, researchers and developers can enhance the likelihood that HCC risk prediction models will be successfully implemented and sustained across multiple clinical and public health settings.
CONCLUSION
Our proposed 4-phase framework for the development and validation of HCC risk stratification models can inform best practices and help to improve the rigor of development, validation, and implementation of HCC risk stratification strategies. Our review of the current landscape showed that only a small number of HCC risk stratification studies have progressed beyond early-phase evaluation—few have undergone phase 2, fewer reached phase 3, and none have completed phase 4 impact evaluation. Hence, this framework enables a structured, comparative assessment of existing models by systematically examining key performance metrics and explicitly mapping them to their stage of development, thereby identifying those most ready for clinical translation. Importantly, this framework can serve as a roadmap for practice guidelines and regulatory agencies to support deliberate endorsement of HCC risk stratification models, accelerating their integration into routine clinical practice.
FUNDING INFORMATION
This work was supported by NIH U01 CA283935, U01 CA230997, and U24CA230144. Fasiha Kanwal’s research is additionally supported by NIH R01 CA256977, CPRIT RP200633, and P01 CA263025. Vincent Wong’s research is additionally supported by GRF 14106923 and GRF 14106824. Yujin Hoshida’s research is additionally supported by NIH R01 CA233794, R01 CA292930, U01 CA288375, P50 CA295495, European Commission ERC-AdG-2020-101021417, and CPRIT RR180016. Amit Singal’s research is additionally supported by U01 CA271887, CPRIT RP200554, and P50 CA295495. Ziding Feng’s research is additionally supported by R01CA277133 and U24CA086368. Pierre Nahon’s research is funded in part by the European Union (GENIAL, Grant agreement ID: 101096312), French Agence Nationale de la Recherche (France 2030 DELIVER ANR-21-RHUS-0001), and by France 2030 RHU LIVER-TRACK (ANR-23-RHUS-0014). Guillermo Marquez, Sidney W. Fu, and Fasiha Kanwal are federal employees. The opinions expressed in this article are the authors’ own and do not reflect the views of the NIH, the Department of Health and Human Services, or the United States government.
Abbreviations:
- AFP
alpha-fetoprotein
- AI
artificial intelligence
- AUROC
area under the receiver operating characteristic curve
- EDRN
Early Detection Research Network
- FASTRAK trial
FAST-MRI for HCC suRveillance in pAtients with high risK of liver cancer
- HBV
hepatitis B virus
- HCC
hepatocellular cancer
- HR
hazard ratio
- MASH
metabolic dysfunction–associated steatohepatitis
- MASLD
metabolic dysfunction–associated steatotic liver disease
- NCI
National Cancer Institute
- NPV
negative predictive value
- PPV
positive predictive value
- PREMIUM trial
PREventing liver cancer Mortality through Imaging with Ultrasound vs. MRI
- QALY
quality-adjusted life year
- RCT
randomized controlled trial
- ROC
receiver operating characteristic curve
- THCCC
Texas HCC Consortium Cohort
- TLC
Translational Liver Cancer
- VA
Veterans Affairs
Footnotes
CONFLICTS OF INTEREST
Ziding Feng received grants from Fujifilm Medical Sciences and Exact Sciences. Yujin Hoshida advises and owns stock in Espervita and Alentis. He advises Helio Genomics and Roche. Jagpreet Chhatwal owns stock in Value Analytics Labs. Vincent Wai-Sun Wong consults for, advises, is on the speakers’ bureau for, and received grants from Gilead. He consults for, advises, and is on the speakers’ bureau for AbbVie, Echosens, and Novo Nordisk. He consults for and advises AstraZeneca, Boehringer Ingelheim, Lilly, Merck, Pfizer, and TARGET PharmaSolutions. He is on the speakers’ bureau for Abbott. He owns stock in and is a co-founder of Illuminatio Medical Technology. Pierre Nahon consults for, received grants from, and/or received honoraria from AstraZeneca, BMS, and Eisai. He consults for and/or received honoraria from Roche. William Lotter consults for and owns stock in SpringTide Ventures. He has other interests in Dekang Healthcare Services. Bachir Taouli consults for and received grants from Bayer, Guerbet, and Siemens. He consults for Bracco, RedDress, and Ascelia. He received grants from Echosens, Regen-eron, Takeda, Helio Genomics, and Perspectum. Amit G. Singal consults for or advises Genentech, AstraZeneca, Eisai, Exelixis, Bayer, Merck, Elevar, Boston Scientific, Sirtex, Fujifilm Medical Sciences, Curve Biosciences, Exact Sciences, Glycotest, Helio Genomics, Roche, Mursla, DELFI, Universal Dx, ImCare, and Abbott. The remaining authors have no conflicts to report.
REFERENCES
- 1.Bray F, Ferlay J, Soerjomataram I, Siegel RL, Torre LA, Jemal A. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2018;68:394–424. [DOI] [PubMed] [Google Scholar]
- 2.Suzuki H, Fujiwara N, Singal AG, Baumert TF, Chung RT, Kawaguchi T, et al. Prevention of liver cancer in the era of next-generation antivirals and obesity epidemic. Hepatology. 2025. doi: 10.1097/hep.0000000000001227. Epub ahead of print. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Colli A, Nadarevic T, Miletic D, Giljaca V, Fraquelli M, Štimac D, et al. Abdominal ultrasound and alpha-foetoprotein for the diagnosis of hepatocellular carcinoma in adults with chronic liver disease. Cochrane Database Syst Rev. 2021;2021:CD013346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Tzartzeva K, Obi J, Rich NE, Parikh ND, Marrero JA, Yopp A, et al. Surveillance imaging and alpha fetoprotein for early detection of hepatocellular carcinoma in patients with cirrhosis: A meta-analysis. Gastroenterology. 2018;154:1706–718.e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Gupta P, Soundararajan R, Patel A, Kumar MP, Sharma V, Kalra N. Abbreviated MRI for hepatocellular carcinoma screening: A systematic review and meta-analysis. J Hepatol. 2021;75:108–19. [DOI] [PubMed] [Google Scholar]
- 6.Rhee H, Kim MJ, Kim DY, An C, Kang W, Han K, et al. Noncontrast magnetic resonance imaging vs ultrasonography for hepatocellular carcinoma surveillance: A randomized, single-center trial. Gastroenterology. 2025;168:1170–77.e12. [DOI] [PubMed] [Google Scholar]
- 7.Kao SZ, Sangha K, Fujiwara N, Hoshida Y, Parikh ND, Singal AG. Cost-effectiveness of a precision hepatocellular carcinoma surveillance strategy in patients with cirrhosis. EClinicalMedicine. 2024;75:102755. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Singal AG, Chen Y, Sridhar S, Mittal V, Fullington H, Shaik M, et al. Novel application of predictive modeling: A tailored approach to promoting HCC surveillance in patients with cirrhosis. Clin Gastroenterol Hepatol. 2022;20:1795–1802.e2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Singal AG, Sanduzzi-Zamparelli M, Nahon P, Ronot M, Hoshida Y, Rich N, et al. International Liver Cancer Association (ILCA) white paper on hepatocellular carcinoma risk stratification and surveillance. J Hepatol. 2023;79:226–39. [DOI] [PubMed] [Google Scholar]
- 10.Lee YT, Fujiwara N, Yang JD, Hoshida Y. Risk stratification and early detection biomarkers for precision HCC screening. Hepatology. 2023;78:319–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.National Cancer Institute-Division of Cancer Prevention. Translational Liver Cancer (TLC) Consortium. Accessed September 8, 2025. https://prevention.cancer.gov/research-areas/net-works-consortia-programs/tlc
- 12.Pepe MS, Etzioni R, Feng Z, Potter JD, Thompson ML, Thornquist M, et al. Phases of biomarker development for early detection of cancer. J Natl Cancer Inst. 2001;93:1054–61. [DOI] [PubMed] [Google Scholar]
- 13.Singal AG, Hoshida Y, Pinato DJ, Marrero J, Nault JC, Paradis V, et al. International Liver Cancer Association (ILCA) White Paper on biomarker development for hepatocellular carcinoma. Gastroenterology. 2021;160:2572–84. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ikeda K, Arase Y, Saitoh S, Kobayashi M, Someya T, Hosaka T, et al. Prediction model of hepatocarcinogenesis for patients with hepatitis C virus-related cirrhosis. Validation with internal and external cohorts. J Hepatol. 2006;44:1089–97. [DOI] [PubMed] [Google Scholar]
- 15.Guyot E, Sutton A, Rufat P, Laguillier C, Mansouri A, Moreau R, et al. PNPLA3 rs738409, hepatocellular carcinoma occurrence and risk model prediction in patients with cirrhosis. J Hepatol. 2013;58:312–8. [DOI] [PubMed] [Google Scholar]
- 16.Tarhuni A, Guyot E, Rufat P, Sutton A, Bourcier V, Grando V, et al. Impact of cytokine gene variants on the prediction and prognosis of hepatocellular carcinoma in patients with cirrhosis. J Hepatol. 2014;61:342–50. [DOI] [PubMed] [Google Scholar]
- 17.El-Serag HB, Kanwal F, Davila JA, Kramer J, Richardson P. A new laboratory-based algorithm to predict development of hepatocellular carcinoma in patients with hepatitis C and cirrhosis. Gastroenterology. 2014;146:1249–55.e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ganne-Carrié N, Layese R, Bourcier V, Cagnot C, Marcellin P, Guyader D, et al. Nomogram for individualized prediction of hepatocellular carcinoma occurrence in hepatitis C virus cirrhosis (ANRS CO12 CirVir). Hepatology. 2016;64:1136–47. [DOI] [PubMed] [Google Scholar]
- 19.Thrift AP, Kanwal F, Liu Y, Khaderi S, Singal AG, Marrero JA, et al. Risk stratification for hepatocellular cancer among patients with cirrhosis using a hepatic fat polygenic risk score. PLoS One. 2023;18:e0282309. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Villa E, Donghia R, Baldaccini V, Tedesco CC, Shahini E, Cozzolongo R, et al. GALAD outperforms aMAP and ALBI for predicting HCC in patients with compensated advanced chronic liver disease: A 12-year prospective study. Hepatol Commun. 2023;7:e0262. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.El-Serag H, Kanwal F, Ning J, Powell H, Khaderi S, Singal AG, et al. Serum biomarker signature is predictive of the risk of hepatocellular cancer in patients with cirrhosis. Gut. 2024;73. doi: 10.1136/gutjnl-2024-332034. Epub ahead of print. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Flemming JA, Yang JD, Vittinghoff E, Kim WR, Terrault NA. Risk prediction of hepatocellular carcinoma in patients with cirrhosis: The ADRESS-HCC risk model. Cancer. 2014;120:3485–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Liang KH, Ahn SH, Lee HW, Huang YH, Chien RN, Hu TH, et al. A novel risk score for hepatocellular carcinoma in Asian cirrhotic patients: A multicentre prospective cohort study. Sci Rep. 2018;8:8608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Nam JY, Sinn DH, Bae J, Jang ES, Kim JW, Jeong SH. Deep learning model for prediction of hepatocellular carcinoma in patients with HBV-related cirrhosis on antiviral therapy. JHEP Rep. 2020;2:100175. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Ioannou GN, Green P, Kerr KF, Berry K. Models estimating risk of hepatocellular carcinoma in patients with alcohol or NAFLD-related cirrhosis for risk stratification. J Hepatol. 2019;71:523–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Ioannou GN, Green PK, Beste LA, Mun EJ, Kerr KF, Berry K. Development of models estimating the risk of hepatocellular carcinoma after antiviral treatment for hepatitis C. J Hepatol. 2018;69:1088–98. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Chen CH, Hu TH, Wang JH, Lai HC, Hung CH, Lu SN, et al. A Mac-2 binding protein glycosylation isomer-based risk model predicts hepatocellular carcinoma in HBV-related cirrhotic patients on antiviral therapy. Cancers. 2022;14:5063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Lee K, Choi GH, Jang ES, Jeong SH, Kim JW. A scoring system for predicting hepatocellular carcinoma risk in alcoholic cirrhosis. Sci Rep. 2022;12:1717. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Wang HW, Chen CY, Lai HC, Hu TH, Su WP, Lu SN, et al. Prediction model of hepatocellular carcinoma in patients with hepatitis B virus-related compensated cirrhosis receiving antiviral therapy. Am J Cancer Res. 2023;13:526–37. [PMC free article] [PubMed] [Google Scholar]
- 30.Chen CH, Wang JH, Lai HC, Hu TH, Hung CH, Lu SN, et al. Mac-2 binding protein glycosylation isomer at 5 years of antiviral therapy predict hepatocellular carcinoma and mortality beyond year 5 in chronic hepatitis B patients with cirrhosis. Am J Cancer Res. 2024;14:2465–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Nahon P, Bamba-Funck J, Layese R, Trépo E, Zucman-Rossi J, Cagnot C, et al. Integrating genetic variants into clinical models for hepatocellular carcinoma risk stratification in cirrhosis. J Hepatol. 2023;78:584–95. [DOI] [PubMed] [Google Scholar]
- 32.Sharma SA, Kowgier M, Hansen BE, Brouwer WP, Maan R, Wong D, et al. Toronto HCC risk index: A validated scoring system to predict 10-year risk of HCC in patients with cirrhosis. J Hepatol. 2017;68:S0168-8278(17)32248-1. doi: 10.1016/j.jhep.2017.07.033. Epub ahead of print. [DOI] [PubMed] [Google Scholar]
- 33.Fan R, Papatheodoridis G, Sun J, Innes H, Toyoda H, Xie Q, et al. aMAP risk score predicts hepatocellular carcinoma development in patients with chronic hepatitis. J Hepatol. 2020;73:1368–78. [DOI] [PubMed] [Google Scholar]
- 34.Kanwal F, Khaderi S, Singal AG, Marrero JA, Asrani SK, Amos CI, et al. Risk stratification model for hepatocellular cancer in patients with cirrhosis. Clin Gastroenterol Hepatol. 2023;21:3296–3304.e3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Fujiwara N, Lopez C, Marsh TL, Raman I, Marquez CA, Paul S, et al. Phase 3 validation of PAaM for hepatocellular carcinoma risk stratification in cirrhosis. Gastroenterology. 2025;168:556–567.e7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): The TRIPOD statement. Ann Intern Med. 2015;162:55–63. [DOI] [PubMed] [Google Scholar]
- 38.Efthimiou O, Seo M, Chalkou K, Debray T, Egger M, Salanti G. Developing clinical prediction models: A step-by-step guide. BMJ. 2024;386:e078276. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Mittal S, El-Serag HB, Sada YH, Kanwal F, Duan Z, Temple S, et al. Hepatocellular carcinoma in the absence of cirrhosis in United States veterans is associated with nonalcoholic fatty liver disease. Clin Gastroenterol Hepatol. 2016;14:124–31.e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Pepe MS, Feng Z, Janes H, Bossuyt PM, Potter JD. Pivotal evaluation of the accuracy of a biomarker used for classification or prediction: Standards for study design. J Natl Cancer Inst. 2008;100:1432–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.van der Pol CB, Lim CS, Sirlin CB, McGrath TA, Salameh JP, Bashir MR, et al. Accuracy of the liver imaging reporting and data system in computed tomography and magnetic resonance image analysis of hepatocellular carcinoma or overall malignancy—A systematic review. Gastroenterology. 2019;156:976–86. [DOI] [PubMed] [Google Scholar]
- 42.Nathani P, Gopal P, Rich N, Yopp A, Yokoo T, John B, et al. Hepatocellular carcinoma tumour volume doubling time: A systematic review and meta-analysis. Gut. 2021;70:401–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Rich NE, John BV, Parikh ND, Rowe I, Mehta N, Khatri G, et al. Hepatocellular carcinoma demonstrates heterogeneous growth patterns in a multicenter cohort of patients with cirrhosis. Hepatology. 2020;72:1654–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Alba AC, Agoritsas T, Walsh M, Hanna S, Iorio A, Devereaux PJ, et al. Discrimination and calibration of clinical prediction models: Users’ guides to the medical literature. JAMA. 2017;318:1377–84. [DOI] [PubMed] [Google Scholar]
- 45.D’Agostino RB Sr, Grundy S, Sullivan LM, Wilson P. Validation of the Framingham coronary heart disease prediction scores: Results of a multiple ethnic groups investigation. Jama. 2001;286:180–7. [DOI] [PubMed] [Google Scholar]
- 46.Gail MH, Brinton LA, Byar DP, Corle DK, Green SB, Schairer C, et al. Projecting individualized probabilities of developing breast cancer for white females who are being examined annually. J Natl Cancer Inst. 1989;81:1879–86. [DOI] [PubMed] [Google Scholar]
- 47.Bhardwaj M, Schöttker B, Holleczek B, Brenner H. Comparison of discrimination performance of 11 lung cancer risk models for predicting lung cancer in a prospective cohort of screening-age adults from Germany followed over 17 years. Lung Cancer. 2022;174:83–90. [DOI] [PubMed] [Google Scholar]
- 48.Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: The Achilles heel of predictive analytics. BMC Med. 2019;17:230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Narasimman M, Hernaez R, Cerda V, Lee M, Sood A, Yekkaluri S, et al. Hepatocellular carcinoma surveillance may be associated with potential psychological harms in patients with cirrhosis. Hepatology. 2024;79:107–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Narasimman M, Hernaez R, Cerda V, Lee M, Yekkaluri S, Khan A, et al. Financial burden of hepatocellular carcinoma screening in patients with cirrhosis. Clin Gastroenterol Hepatol. 2024;22:760–67.e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Singal AG, Patibandla S, Obi J, Fullington H, Parikh ND, Yopp AC, et al. Benefits and harms of hepatocellular carcinoma surveillance in a prospective cohort of patients with cirrhosis. Clin Gastroenterol Hepatol. 2021;19:1925–932.e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Nahon P, Najean M, Layese R, Zarca K, Segar LB, Cagnot C, et al. Early hepatocellular carcinoma detection using magnetic resonance imaging is cost-effective in high-risk patients with cirrhosis. JHEP Rep. 2022;4:100390. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Parikh ND, Singal AG, Hutton DW, Tapper EB. Cost-effectiveness of hepatocellular carcinoma surveillance: An assessment of benefits and harms. Am J Gastroenterol. 2020;115:1642–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Chhatwal J, Hajjar A, Mueller PP, Nemutlu G, Kulkarni N, Peters MLB, et al. Hepatocellular carcinoma incidence threshold for surveillance in virologically cured hepatitis C individuals. Clin Gastroenterol Hepatol. 2024;22:91–101.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Pepe MS, Feng Z, Huang Y, Longton G, Prentice R, Thompson IM, et al. Integrating the predictiveness of a marker with its performance as a classifier. Am J Epidemiol. 2008;167:362–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Kim NJ, Rozenberg-Ben-Dror K, Jacob DA, Rich NE, Singal AG, Aby ES, et al. Provider attitudes toward risk-based hepatocellular carcinoma surveillance in patients with cirrhosis in the United States. Clin Gastroenterol Hepatol. 2022;20:183–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Singal AG, Yang JD, Jalal PK, Salgia R, Mehta N, Hoteit MA, et al. Patient-perceived risk of hepatocellular carcinoma and net benefit of surveillance: A multicenter survey study. Am J Gastroenterol. 2025;120:1800–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Singal AG, Quirk L, Boike J, Chernyak V, Feng Z, Giamarqo G, et al. Value of HCC surveillance in a landscape of emerging surveillance options: Perspectives of a multi-stakeholder modified Delphi panel. Hepatology. 2024;82:794–809. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Pepe MS, Janes H, Li CI, Bossuyt PM, Feng Z, Hilden J. Early-phase studies of biomarkers: what target sensitivity and specificity values might confer clinical utility? Clin Chem. 2016;62:737–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Singal AG, Chhatwal J, Parikh N, Tapper E. Cost-effectiveness of a biomarker-based screening strategy for hepatocellular carcinoma in patients with cirrhosis. Liver Cancer. 2024;13:643–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Schoenberger H, Chong N, Fetzer DT, Rich NE, Yokoo T, Khatri G, et al. Dynamic changes in ultrasound quality for hepatocellular carcinoma screening in patients with cirrhosis. Clin Gastroenterol Hepatol. 2022;20:1561–69.e4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Toyoda H, Hiraoka A, Olivares J, Al-Jarrah T, Devlin P, Kaneoka Y, et al. Outcome of hepatocellular carcinoma detected during surveillance: Comparing USA and Japan. Clin Gastroenterol Hepatol. 2021;19:2379–388.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Singal AG, Kanwal F, Llovet JM. Global trends in hepatocellular carcinoma epidemiology: Implications for screening, prevention and therapy. Nat Rev Clin Oncol. 2023;20:864–84. [DOI] [PubMed] [Google Scholar]
- 64.Audureau E, Carrat F, Layese R, Cagnot C, Asselah T, Guyader D, et al. Personalized surveillance for hepatocellular carcinoma in cirrhosis—Using machine learning adapted to HCV status. J Hepatol. 2020;73:1434–45. [DOI] [PubMed] [Google Scholar]
- 65.Dana J, Meyer A, Paisant A, Rode A, Sartoris R, Séror O, et al. Improving risk stratification and detection of early HCC using ultrasound-based deep learning models. JHEP Rep. 2025;7:101510. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.DeGrave AJ, Janizek JD, Lee SI. AI for radiographic COVID-19 detection selects shortcuts over signal. medRxiv. 2020. doi: 10.1101/2020.09.13.20193565 [DOI] [Google Scholar]
- 67.Lotter W, Hippe DS, Oshiro T, Lowry KP, Milch HS, Miglioretti DL, et al. Influence of mammography acquisition parameters on AI and radiologist interpretive performance. Radiol Artif Intell. 2025;7:e240861. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Aldea M, Salto-Tellez M, Marra A, Umeton R, Stenzinger A, Koopman M, et al. ESMO basic requirements for AI-based biomarkers in oncology (EBAI). Ann Oncol. 2025;37:414–30. [DOI] [PubMed] [Google Scholar]
- 69.Sanderson S, Zimmern R, Kroese M, Higgins J, Patch C, Emery J. How can the evaluation of genetic tests be enhanced? Lessons learned from the ACCE framework and evaluating genetic tests in the United Kingdom. Genet Med. 2005;7:495–500. [DOI] [PubMed] [Google Scholar]
- 70.Smart A. A multi-dimensional model of clinical utility. Int J Qual Health Care. 2006;18:377–82. [DOI] [PubMed] [Google Scholar]
- 71.Esserman LJ. The WISDOM Study: Breaking the deadlock in the breast cancer screening debate. NPJ Breast Cancer. 2017;3:34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.My Personalized Breast Screening (MyPeBS). Updated September 19, 2024. Accessed July 2025. [Google Scholar]
- 73.Nahon P, Ronot M, Sutter O, Natella PA, Baloul S, Durand-Zaleski I, et al. Study protocol for FASTRAK: A randomised controlled trial evaluating the cost impact and effectiveness of FAST-MRI for HCC suRveillance in pAtients with high risK of liver cancer. BMJ Open. 2024;14:e083701. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.HCC risk calculator. Hepatocellular Carcinoma (HCC) Risk calculator. 2025. https://hccrisk.com/
- 75.Kaplan DE, Ripoll C, Thiele M, Fortune BE, Simonetto DA, Garcia-Tsao G, et al. AASLD Practice Guidance on risk stratification and management of portal hypertension and varices in cirrhosis. Hepatology. 2024;79:1180–211. [DOI] [PubMed] [Google Scholar]
- 76.Singal AG, Parikh ND, Kanwal F, Marrero JA, Deodhar S, Page-Lester S, et al. National Liver Cancer Screening Trial (TRACER) study protocol. Hepatol Commun. 2024;8:e0565. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Janiaud P, Ioannidis JPA, Kasenda B, Fretheim A, Goodman SN, Hemkens LG. Single-arm trials can provide randomized real-world evidence: The random invitation single-arm trial design. Ann Intern Med. 2025;178:1150–6. [DOI] [PubMed] [Google Scholar]
- 78.Zelen M. A new design for randomized clinical trials. N Engl J Med. 1979;300:1242–5. [DOI] [PubMed] [Google Scholar]
- 79.Jin Q, Wang Z, Floudas CS, Chen F, Gong C, Bracken-Clarke D, et al. Matching patients to clinical trials with large language models. ArXiv. 2024;15:9074. doi: 10.1038/s41467-024-53081-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Unlu O, Varugheese M, Shin J, Subramaniam SM, Stein DWJ, St Laurent JJ, et al. Manual vs AI-assisted prescreening for trial eligibility using large language models—A randomized clinical trial. JAMA. 2025;333:1084–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Etzioni R, Gulati R, Patriotis C, Rutter C, Zheng Y, Srivastava S, et al. Revisiting the standard blueprint for biomarker development to address emerging cancer early detection technologies. J Natl Cancer Inst. 2024;116:189–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Rich NE, Parikh ND, Singal AG. Overdiagnosis: An under-studied issue in hepatocellular carcinoma surveillance. Semin Liver Dis. 2017;37:296–304. [DOI] [PMC free article] [PubMed] [Google Scholar]
