Skip to main content
Lippincott Open Access logoLink to Lippincott Open Access
. 2023 Dec 13;119(5):814–822. doi: 10.14309/ajg.0000000000002629

Performance of Prediction Models for Esophageal Squamous Cell Carcinoma in General Population: A Systematic Review and External Validation Study

Hao Jiang 1, Ru Chen 2,3, Yanyan Li 4, Changqing Hao 5, Guohui Song 6, Zhaolai Hua 7, Jun Li 8, Yuping Wang 1, Wenqiang Wei 1,2,3,
PMCID: PMC11062607  PMID: 38088388

Abstract

INTRODUCTION:

Prediction models for esophageal squamous cell carcinoma (ESCC) need to be proven effective in the target population before they can be applied to population-based endoscopic screening to improve cost-effectiveness. We have systematically reviewed ESCC prediction models applicable to the general population and performed external validation and head-to-head comparisons in a large multicenter prospective cohort including 5 high-risk areas of China (Fei Cheng, Lin Zhou, Ci Xian, Yang Zhong, and Yan Ting).

METHODS:

Models were identified through a systematic review and validated in a large population-based multicenter prospective cohort that included 89,753 participants aged 40–69 years who underwent their first endoscopic examination between April 2017 and March 2021 and were followed up until December 31, 2022. Model performance in external validation was estimated based on discrimination and calibration. Discrimination was assessed by C-statistic (concordance statistic), and calibration was assessed by calibration plot and Hosmer-Lemeshow test.

RESULTS:

The systematic review identified 15 prediction models that predicted severe dysplasia and above lesion (SDA) or ESCC in the general population, of which 11 models (4 SDA and 7 ESCC) were externally validated. The C-statistics ranged from 0.67 (95% confidence interval 0.66–0.69) to 0.70 (0.68–0.71) of the SDA models, and the highest was achieved by Liu et al (2020) and Liu et al (2022). The C-statistics ranged from 0.51 (0.48–0.54) to 0.74 (0.71–0.77), and Han et al (2023) had the best discrimination of the ESCC models. Most models were well calibrated after recalibration because the calibration plots coincided with the x = y line.

DISCUSSION:

Several prediction models showed moderate performance in external validation, and the prediction models may be useful in screening for ESCC. Further research is needed on model optimization, generalization, implementation, and health economic evaluation.

KEYWORDS: esophageal squamous cell carcinoma, screening, prediction model, systematic review, external validation

INTRODUCTION

Esophageal cancer ranks seventh and sixth in global cancer incidence and mortality with 604,000 new cases and 544,000 deaths estimated in 2020, respectively (1). China has the highest burden of esophageal cancer, accounting for more than half of cases (53.7%) and deaths (55.3%) worldwide. Esophageal squamous cell carcinoma (ESCC) is the predominant histological subtype, accounting for approximately 90.4% of the incidence of esophageal cancer in China (2). Population-based endoscopic screening that enable early detection of precancerous lesions and early diagnosis of ESCC have significantly reduced the burden in high-risk areas of China over the past decades (35). However, as the prevalence of precancerous lesions decreases in the high-risk areas, less than 3% of the participants were detected with severe dysplasia and above lesions (SDA, requiring immediate treatment), leading to a large number of unnecessary endoscopies (6,7). In addition, the risk of mechanical injury, bleeding, and other complications during unnecessary endoscopy and biopsy may impair physical and mental health.

Prediction models that estimate the risk of developing ESCC in the general population can greatly enhance the implementation efficiency of population-based screening by risk stratification. This means that endoscopy is used only for high-risk individuals, while lifestyle modifications are recommended for low-risk individuals. Furthermore, these models can also be used in hospitals and physical examination centers to provide recommendations for individualized endoscopy screening.

There are several prediction models for ESCC or precancerous lesions in the general population (8,9). However, many models have not undergone external validation, and the ability of these models to identify potential cases in the target population is unknown. This information is important for assessing the model screening effectiveness and determining which models could be considered for further optimization.

Above all, prediction models need to be proven effective in the target population before they can be applied to population-based screening practices to improve cost-effectiveness. Thus, we have systematically reviewed ESCC prediction models suitable for the general population. The searched models underwent quality appraisal, followed by validation and head-to-head comparisons in a large multicenter prospective cohort.

METHODS

Literature search

We comprehensively searched the PubMed and EMBASE up to April 1, 2023, and updated our previously published systematic review on esophageal cancer prediction models (8) to identify prediction models for ESCC. The search strategy is summarized Supplementary Table S2 (see Supplementary Digital Content 1, http://links.lww.com/AJG/D141). Studies were qualified if they met the following criteria: (i) published as an original research article; (ii) including prediction model of ESCC and/or its precancerous lesions; (iii) based on the general population; and (iv) considering more than 1 predictor. The appropriate models were then extracted from the studies. Two independent reviewers (H.J. and R.C.) searched and screened the search results to find qualified studies and models. Disagreements between the 2 reviewers were fully discussed and agreed upon, with the senior researcher making the final decisions.

Data extraction and quality assessment

Data were systematically extracted from the qualified studies according to the criteria of the checklist for critical appraisal and data extraction for systematic reviews of prediction modeling studies (CHARMS) (10). The information mainly covered the following 9 domains: general information (first author, year of publication, and country), data sources (study design), participants (inclusion and exclusion criteria), outcomes (definition), predictors (number, definition, and type of predictors), sample size (number of outcomes and events), model performance (calibration and discrimination), model evaluation (methods of validation), and results (predictive weight or regression coefficients and intercepts). The quality of the studies was assessed using Prediction model Risk Of Bias Assessment Tool (PROBAST) (11). This tool contains 20 signaling questions to judge the risk of bias (ROB) in 4 domains: participants, predictors, outcome, and analysis.

External validation

The eligible models were externally validated in a large multicenter prospective cohort in high-risk areas of China. Models with genetic predictors were not validated due to a lack of such information. If the model parameters were not reported in the original literature, we contacted the corresponding author to obtain this information. If unsuccessful, we did not validate them or provided validation after recalibration. We applied the inclusion criteria (see Supplementary Table S5, Supplementary Digital Content 1, http://links.lww.com/AJG/D141) of the development population to filter the validation cohort to maintain homogeneity between them as much as possible. The external validation was reported in accordance with the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis guideline (12) (see Supplementary Table S1, Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Validation cohort

The validation cohort was from the National Cohort of Esophageal Cancer–Prospective Cohort Study of Esophageal Cancer and Precancerous Lesions based on High-Risk Population in China (13). It included 5 centers in high-risk areas (Fei Cheng, Lin Zhou, Ci Xian, Yang Zhong, and Yan Ting) and contained 89,753 participants registered between April 2017 and March 2021. The prevalence of SDA and ESCC at baseline was 1,190 per 100,000 and 442 per 100,000, respectively. A total of 87,667 participants who were free of any tumors and severe dysplasia at baseline were followed up until December 31, 2022, with a median follow-up time of 1,389 days (3.8 years) and a total of 293 new cases with ESCC.

The recruitment was conducted village by village in each center. Village doctors and local staff informed residents of the benefits of screening and invited all permanent residents aged 40–69 years in the selected villages to the designated hospitals for endoscopic examination. Those who were willing to be screened were registered and scheduled for screening. If the hospitals were far from the villages, we arranged vehicles to transport the participants. If the response rate in a village was below 30%, we conducted a second mobilization. Each participant underwent face-to-face informed consent signing, an endoscopic contraindication survey, and a questionnaire by trained staff before endoscopy. Participants were excluded if they met the following criteria: (i) history of cancer or mental disorder; (ii) contraindications for endoscopic examinations, and (iii) unable to provide informed consent.

Screening procedures followed the recommendations of the Chinese Expert Consensus on Early Esophageal Cancer Screening and Endoscopic Diagnosis and Treatment. All endoscopic examinations and therapies were performed by trained physicians at local hospitals. The entire esophagus was examined, and all lesions were biopsied. Two pathologists checked the biopsy sections independently. Inconsistencies in diagnosis were resolved by discussion.

A uniform questionnaire was developed based on the Chinese Kadoorie Biobank questionnaire (14) and was used to collect information, covering demographic factors, socioeconomic status, smoking, alcohol and tea consumption, diet, indoor air pollution, physical activity, reproductive history (female), sleep status, medical and family history, risk factors associated with esophageal cancer (history of digestive disorders, family history of cancer), drinking water, dietary habits (hot food, food tenderness, and speed of eating), and oral hygiene. Physical information (blood pressure, heart rate, height, and weight) was also measured.

Both active and passive follow-ups were conducted for all participants (see Supplementary Figure S2, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). Active follow-ups were required according to the diagnosis. For patients with SDA, we took treatment immediately, and those who refused treatment were followed up at least once a year. For patients diagnosed with esophageal mild dysplasia, a reexamination was required in 3 years, and for those with esophagus moderate dysplasia, an annual reexamination was required. Annual interviews with participants who were diagnosed with SDA during the screening were conducted by telephone or home visit to collect information on outcomes. Passive follow-up was conducted by matching cases through cancer registries, local hospitals, and death surveillance systems once a year. The last follow-up was up to December 31, 2022. All participants would be followed up for at least 10 years.

Outcomes and predictors

The outcomes for external validation were SDA and ESCC, which would be adjusted for different models based on their original definitions. The outcome of logistic models (1522) was defined as the endoscopic diagnosis at baseline and 2 of the models, Liu et al (2022) (15) and Chen et al (2021) (19), also included cases with ESCC diagnosed within 1 and 3 years of follow-up, respectively. The outcome of Cox proportional hazards models (2325) was defined as new cases with ESCC during follow-up in participants without any tumors and severe dysplasia at baseline. Esophageal cancer was coded as C15 according to the International Classification of Diseases, 10th Revision. SDA included severe dysplasia, squamous carcinoma in situ, and squamous cell carcinoma, and ESCC excluded severe dysplasia compared with SDA.

Regarding the definitions of the predictors, we initially tried to match the original definitions of the predictors with the variables in the validation cohort. Second, we kept as close as possible to the original definitions depending on the relevant variables in the validation cohort when there were no variables that could be directly matched. Finally, we applied a typical value to all participants if no suitable alternatives were available. The definitions of predictors in the development population and the validation cohort were detailed in Supplementary Table S6 (see Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Statistical analysis

We conducted complete case validation because the proportion of missing values for the variables used for external validation was very small in the validation cohort (see Supplementary Table S7, Supplementary Digital Content 1, http://links.lww.com/AJG/D141) and the missing values were mostly concentrated (i.e., a participant had multiple variables missing at the same time). No more than 4% of the participants were excluded because of missing values in external validation. Characteristics of the validation cohort were summarized as the median and interquartile range (IQR) for continuous variables and the frequency and proportion for categorical variables.

The model performance in external validation was estimated based on discrimination and calibration. Discrimination was assessed by C-statistic (concordance statistic) and its 95% confidence interval (CI), and calibration was assessed in 2 ways: plotting calibration plots to compare predicted and observed probabilities and performing Hosmer-Lemeshow (HL) test. The model was considered well calibrated if the calibration plot coincided with the x = y line and the HL test was not significant (P > 0.05) (26,27). The stratified models were also assessed as a whole by combining their respective predicted probabilities. We also recalibrated all models based on the prevalence of outcomes in the validation cohort to eliminate the impact of different prevalence between development and validation populations on calibration. Prognostic index was calculated as i=1nβixi, where xi was the ith predictor and βi was the corresponding coefficient. The recalibrated model was then constructed by refitting with prognostic index as the only covariate. This recalibration method only adjusted the calibration and has no effect on the discrimination (28,29). All analyses were performed in R (version 4.2.1), using the lrm function in rms (6.3-0) and coxph function in survival (3.2-13) for model validation and recalibration, the hoslem.test in ResourceSelection (0.3-5) for the HL test, and the ggplot2 (3.3.5) for plotting.

RESULTS

Systematic review results

In all, 2,648 of 2,662 studies screened on title and abstract were excluded, with 14 remaining for full-text review. Three studies were identified after screening the full text of the studies. Finally, 15 studies were qualified of which 3 were newly identified and 12 were selected from previous systematic review, and 15 models were extracted from these studies. The flowchart of the systematic review and model selection is presented in Supplementary Figure S1 (see Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Characteristics of models in the systematic review

Of the 15 eligible models, two-thirds were constructed based on Chinese populations (Table 1). SDA was used as the outcome in 4 models and ESCC in 11 models. They were based on 3 study design: cross-sectional (5 models), case-control (6 models), and cohort (4 models) and 2 model types: logistic model (11 models) and Cox proportional hazards model (4 models). The sample size ranged from 868 to 115,686, and the cross-sectional and cohort designs had much larger sample size than the case-control design. The number of predictors varied from 3 to 11, and the most commonly used were age, sex, smoking, alcohol, body mass index, and family history (see Supplementary Table S3, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). All models reported discrimination through C-statistic, and the values ranged from 0.68 (95% CI 0.62–0.74) to 0.87 (0.84–0.95) in SDA models and 0.71 (0.66–0.78) to 0.88 (0.85–0.90) in ESCC models. The calibration plot and HL test were the principal calibration methods, while many models did not report calibration results. Fourteen models underwent internal validation (study types of 1b, 2a, 2b, and 3), but only half were externally validated in separate populations (study type of 3).

Table 1.

Characteristics of the qualified models in the systematic review

graphic file with name acg-119-814-g001.jpg

The result of the quality assessment is summarized in Supplementary Table S4 (see Supplementary Digital Content 1, http://links.lww.com/AJG/D141). While most models had low ROB on participants, predictors, and outcome domain, all models were considered to have high ROB according to PROBAST because they all showed high ROB in the analysis domain. The high ROB in the analysis domain was mainly due to the small number of participants with outcomes and lack of methods to handle missing data and to assess calibration.

Characteristics of the validation cohort

Overall, 89,753 participants were included from 5 centers (Fei Cheng, Lin Zhou, Ci Xian, Yang Zhong, and Yan Ting). A total of 671 participants were detected with severe dysplasia and 397 with ESCC (Table 2) at baseline; thus, the prevalence of SDA and ESCC were 1,190 per 100,000 and 442 per 100,000, respectively. The median age was 55 years (IQR 50–62), and the proportion of male participants was 41.3% (n = 37,032). The smoking and alcohol consumption rates reached 21.6% and 12.2%, respectively. Approximately 4.3% of participants had a family history of esophageal cancer among first-degree relatives. Height (n = 1,974) and weight (n = 2,015) had the highest proportions of missing values, but both were less than 3%, and most of them were missed simultaneously; hence, the proportion of missing body mass index (n = 2,046) was also low. The complete description of the variables used in the external validation is summarized in Supplementary Table S7 (see Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Table 2.

Characteristics of the validation cohorta

graphic file with name acg-119-814-g002.jpg

External validation

In summary, 11 of the 15 models were externally validated, and each model was assessed for discrimination (Table 3). All SDA models were assessed for calibration before and after recalibration. Seven ESCC models were assessed only for calibration after recalibration due to a lack of parameter information. Four ESCC models were excluded for external validation owing to the inclusion of genetic predictors or lack of information on parameters (see Supplementary Table S8, Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Table 3.

The sample size, discrimination, and calibration in the external validation

graphic file with name acg-119-814-g003.jpg

The number of participants in the validation sets ranged from 81,271 (1,030 cases) to 81,291 (1,185 cases) in SDA models and 85,454 (285 cases) to 89,666 (396 cases) in ESCC models (Table 3). The discrimination was reported through C-statistic. Values ranged from 0.67 (0.66–0.69) to 0.70 (0.68–0.71) of the SDA models, and the highest was achieved by Liu et al (2020) (18) and Liu et al (2022) (17). Values of the ESCC models varied considerably, ranging from 0.51 (0.48–0.54) to 0.74 (0.71–0.77), and Han et al (2023) (24) had the best discrimination. There was a significant decrease in C-statistic in external validation compared with that in development validation (Figure 1). The calibration was reported through calibration plot and the HL test. All SDA models had significant HL test and underestimation before recalibration (Figure 2), while calibration improved after recalibration as the calibration plots coincided with the x = y line (see Supplementary Figure S3, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). Of the 7 ESCC models for which calibration was assessed after recalibration, most models performed well based on calibration plots and HL tests (see Supplementary Figures S4 and S5, Supplementary Digital Content 1, http://links.lww.com/AJG/D141), but the calibration of Shen et al (2021) (21) and Wang et al (2019) (22) were poor because HL tests were significant and calibration plots deviated from the x = y line.

Figure 1.

Figure 1.

Discrimination of the prediction models in development and external validation.

Figure 2.

Figure 2.

Calibration plots of the logistic models for severe dysplasia and above lesion (SDA) before recalibration.

DISCUSSION

We have conducted a systematic review and external validation of prediction models for SDA and ESCC. Overall, most models showed moderate discrimination in external validation, but the C-statistics were significantly lower compared with the original models (Figure 1). Although Liu et al (2022) (17) added 2 new predictors (gender and family history of ESCC), compared with Liu et al (2020) (18), its discrimination was not improved in external validation (Table 3). Liu et al (2022) (15) was an updated version of Liu et al (2017) (16) that incorporated cases with ESCC diagnosed within 1 year after screening, added the quadratic term for age as a predictor, and fitted an all-age model. While this update slightly improved the discrimination, the calibration was much worse than the previous model (Figure 2). The calibration of all SDA models in the external validation improved after recalibration (see Supplementary Figure S3, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). Of the 7 ESCC models, the discrimination of models with cohort and cross-sectional design was generally stronger than case-control design in external validation (Table 3). Yang et al (2021) (20) and Han et al (2023) (24) had the best calibration after recalibration (see Supplementary Figures S4 and S5, Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

Screening for ESCC is a challenging task in China due to the large population in high-risk areas. SDA might be the more appropriate outcome for screening because the primary purpose of both population-based screening and opportunistic screening in hospitals is to detect precancerous lesions and early-stage cancers. In general, the logistic models of cross-sectional design (1519) were diagnostic and more suitable for prescreening before endoscopic screening, and the Cox proportional hazards models of cohort design (2325) were prognostic and more suitable for predicting long-term risk. The event rate of the prediction models varied depending on the outcome and study design. SDA models had higher event rate than ESCC models because the prevalence of SDA was higher than ESCC in the general population. Unlike models with cross-sectional design, the event rate of the models with case-control design was determined by researcher, so we performed recalibration based on the event rate in the validation cohort to make the models comparable.

The prediction models from China and externally validated in this study were based on population-based screening programs in high-risk areas, and a few models (17,18,21) also included hospital outpatients. This external validation reflected the performance of the models in different high-risk areas and provided aid in optimizing population-based screening programs. Moreover, the models could be extended to opportunistic screening in hospitals to predict precancerous lesions and early cancers compared with models based on clinical patients who were mostly diagnosed with ESCC at advanced stages. In addition, 2 models based on Nordic populations (22,25) were also validated, which assessed the generalization ability of the models in different populations and might provide some indication for other countries using prediction models for screening.

The PROBAST gave a preliminary assessment of the prediction models in systematic review (see Supplementary Table S4, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). The 4 models (3033) excluded for external validation were at high ROB in almost all domains. The risks of the remaining models were mainly in the analysis domain. First, there were small number of participants with outcome due to the low prevalence of outcomes in the population. Second, there was a lack of methods to handle missing data, which was avoided by the fact that there were few missing values in the validation cohort. Third, more than half of the models did not report calibration (Table 1), and we assessed calibration after recalibration for those models.

Predictors are critical to the model applicability. We aim to achieve higher model performance with fewer and simpler predictors. Uncommon variables, such as the use of wood and coal as the main fuel (15,16), oral hygiene (20), family wealth score (20), pesticide exposure (16), and place of residence during childhood (22), require more evidence to avoid including them based on statistical significance alone. Incorporating genetic predictors (30,33) and endoscopic findings (number of lesions and size of distinct lesion) (24) may reduce the applicability of models for population-based screening because they need additional measurements and costs.

Our study showed that the calibration of SDA models varied significantly before and after recalibration. According to the logistic regression principle, the predicted probability is closely related to the prevalence of the development population. If the prevalence of the validation population is significantly higher than that of the development population, the predicted probability is underestimated and vice versa (Figure 2 and see Supplementary Figure S3, Supplementary Digital Content 1, http://links.lww.com/AJG/D141). In the cancer screening context, overestimation may lead to overscreening by getting more people to screen, and underestimation may result in missed diagnosis. In the absence of perfect prediction, the selection of model and cutoff should be judged on practical considerations.

The common approach to implementing these prediction models for ESCC screening is to select cutoff to identify high-risk subgroups while retaining a high sensitivity of SDA and ESCC. A study noted that the risk-stratified endoscopic screening improved the SDA and ESCC detection rates by approximately 88.9%, saved approximately half of the number needed to screen to detect 1 case, and reduced the average cost per SDA by 40.9% compared with the universal screening, suggesting that this screening strategy in high-risk areas of China would be more efficient and cost-effective (34). Despite the substantial benefits of this strategy, 7.5% of patients with SDA would be missed compared with universal screening. Moreover, screening performance varied across study centers although the same model was used. The challenges in model implementation included the selection of cutoff, generalization of prediction model, and consideration of screening efficiency (detection rate), effectiveness (number needed to screen), and cost (average cost per case detection). In future, more prospective studies are needed on model performance optimization and health economic evaluation in practical screening.

This study has some strengths and limitations. To our best knowledge, this is the first study to systematically identify prediction models for ESCC and evaluate their performance. Second, the validation cohort is the latest, most representative, and large-scale multicenter prospective cohort in the world. Third, we comprehensively gathered data on participants' risk factors and have been following them for a long period.

The first limitation is that some models could not be externally validated or only be assessed after recalibration due to a lack of information. The second limitation is that there were some biases in the validation process. For the models of case-control design, we assessed after recalibration rather than calculated the 5-year risk as the original study, which requires age-specific and sex-specific incidence of development population and population attributable risk (PAR) (2022). In addition, the definitions of predictors were not fully consistent with the original model. For example, we used family history of esophageal cancer substitute for ESCC (see Supplementary Table S6, Supplementary Digital Content 1, http://links.lww.com/AJG/D141).

We identified 15 SDA and ESCC prediction models through a systematic review and externally validated 11 models in a large multicenter prospective cohort. Several models showed moderate discrimination, such as Liu et al (2022) (17) for the SDA model and Han et al (2023) (24) for the ESCC model. The calibration of SDA models improved after recalibration. In summary, our study hinted that the prediction models may be potentially useful in screening for ESCC, and further research is needed on model optimization, generalization, implementation, and health economic evaluation.

CONFLICTS OF INTEREST

Guarantor of the article: First author Hao Jiang, MS, corresponding author Ru Chen, MD, PhD, and Wenqiang Wei, MD, PhD, take full responsibility for the conduct of the study.

Specific author contributions: W.Q.W., R.C., and H.J.: designed this study. Y.L., C.Q.H., G.H.S., L.H.Z., and J.L.: were responsible for endoscopy, diagnostic pathology, epidemiology investigation, and surveillance at each center. H.J., R.C., and W.Q.W.: contributed to data analysis. H.J.: completed the initial drafting of the manuscript. H.J., R.C., W.Q.W., and Y.P.W.: critically revised the manuscript for important intellectual content.

Financial support: This study was supported by CAMS Innovation Fund for Medical Sciences (2021-I2M-1-010, 2021-I2M-1-013), National Natural Science Foundation of China (81974493, 81903403), National Key R&D Program of China (2016YFC0901400), and National Science & Technology Fundamental Resources Investigation Program of China (2019FY101101). The study funders had no role in the design of the study; the collection, analysis, or interpretation of the data; the writing of the manuscript; or the decision to submit the manuscript for publication.

Potential competing interests: None to report.

Study Highlights.

WHAT IS KNOWN

  • ✓ Early detection of esophageal squamous cell carcinoma (ESCC) and its precancerous lesions in the general population would provide substantial public health benefits.

  • ✓ There are many models available to predict the risk of developing ESCC in the general population, but they have not all been validated and compared in external populations.

  • ✓ External validation is essential for model evaluation, optimization, and implementation in screening practice.

WHAT IS NEW HERE

  • ✓ This study identified 15 prediction models for ESCC and validated 11 of them in an external cohort, providing a basis for model selection and implementation in future screening.

  • ✓ Several prediction models showed moderate discrimination in the external validation, and most of the models were well calibrated after recalibration.

Supplementary Material

acg-119-814-s001.pdf (643.5KB, pdf)

Footnotes

SUPPLEMENTARY MATERIAL accompanies this paper at http://links.lww.com/AJG/D141

Contributor Information

Hao Jiang, Email: chiang_hao@qq.com.

Ru Chen, Email: chenru1900@163.com.

Yanyan Li, Email: 71354440@qq.com.

Changqing Hao, Email: haochq@126.com.

Guohui Song, Email: sghui2009@163.com.

Zhaolai Hua, Email: cnyzhzl@qq.com.

Jun Li, Email: liji0326@163.com.

Yuping Wang, Email: wyppumc@163.com.

Wenqiang Wei, Email: weiwq2006@126.com.

REFERENCES

  • 1.Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2021;71(3):209–49. [DOI] [PubMed] [Google Scholar]
  • 2.Morgan E, Soerjomataram I, Rumgay H, et al. The global landscape of esophageal squamous cell carcinoma and esophageal adenocarcinoma incidence and mortality in 2020 and projections to 2040: New estimates from GLOBOCAN 2020. Gastroenterology 2022;163(3):649–58.e2. [DOI] [PubMed] [Google Scholar]
  • 3.Wei WQ, Chen ZF, He YT, et al. Long-term follow-up of a community assignment, one-time endoscopic screening study of esophageal cancer in China. J Clin Oncol 2015;33(17):1951–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Zhang N, Li Y, Chang X, et al. Long-term effectiveness of one-time endoscopic screening for esophageal cancer: A community-based study in rural China. Cancer 2020;126(20):4511–20. [DOI] [PubMed] [Google Scholar]
  • 5.Chen R, Liu Y, Song G, et al. Effectiveness of one-time endoscopic screening programme in prevention of upper gastrointestinal cancer in China: A multicentre population-based cohort study. Gut 2021;70(2):251–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Li F, Li X, Guo C, et al. Estimation of cost for endoscopic screening for esophageal cancer in a high-risk population in rural China: Results from a population-level randomized controlled trial. Pharmacoeconomics 2019;37(6):819–27. [DOI] [PubMed] [Google Scholar]
  • 7.He Z, Ke Y. Precision screening for esophageal squamous cell carcinoma in China. Chin J Cancer Res 2020;32(6):673–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Chen R, Zheng R, Zhou J, et al. Risk prediction model for esophageal cancer among general population: A systematic review. Front Public Health 2021;9:680967. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Li H, Sun D, Cao M, et al. Risk prediction models for esophageal cancer: A systematic review and critical appraisal. Cancer Med 2021;10(20):7265–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Moons KG, de Groot JA, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: The CHARMS checklist. PLoS Med 2014;11(10):e1001744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Wolff RF, Moons KGM, Riley RD, et al. PROBAST: A tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med 2019;170(1):51–8. [DOI] [PubMed] [Google Scholar]
  • 12.Moons KG, Altman DG, Reitsma JB, et al. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): Explanation and elaboration. Ann Intern Med 2015;162(1):W1–W73. [DOI] [PubMed] [Google Scholar]
  • 13.Chen R, Ma S, Guan C, et al. The National Cohort of Esophageal Cancer-Prospective Cohort Study of Esophageal Cancer and Precancerous Lesions based on High-Risk Population in China (NCEC-HRP): Study protocol. BMJ Open 2019;9(4):e027360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Chen Z, Lee L, Chen J, et al. Cohort profile: The Kadoorie Study of Chronic Disease in China (KSCDC). Int J Epidemiol 2005;34(6):1243–9. [DOI] [PubMed] [Google Scholar]
  • 15.Liu M, Zhou R, Liu Z, et al. Update and validation of a diagnostic model to identify prevalent malignant lesions in esophagus in general population. EClinicalMedicine 2022;47:101394. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Liu M, Liu Z, Cai H, et al. A model to identify individuals at high risk for esophageal squamous cell carcinoma and precancerous lesions in regions of high prevalence in China. Clin Gastroenterol Hepatol 2017;15(10):1538–46.e7. [DOI] [PubMed] [Google Scholar]
  • 17.Liu Z, Zheng H, Liu M, et al. Development and external validation of an improved version of the diagnostic model for opportunistic screening of malignant esophageal lesions. Cancers (Basel) 2022;14(23):5945. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Liu Z, Guo C, He Y, et al. A clinical model predicting the risk of esophageal high-grade lesions in opportunistic screening: A multicenter real-world study in China. Gastrointest Endosc 2020;91(6):1253–60. [DOI] [PubMed] [Google Scholar]
  • 19.Chen W, Li H, Ren J, et al. Selection of high-risk individuals for esophageal cancer screening: A prediction model of esophageal squamous cell carcinoma based on a multicenter screening cohort in rural China. Int J Cancer 2021;148(2):329–39. [DOI] [PubMed] [Google Scholar]
  • 20.Yang X, Suo C, Zhang T, et al. A nomogram for screening esophageal squamous cell carcinoma based on environmental risk factors in a high-incidence area of China: A population-based case-control study. BMC Cancer 2021;21(1):343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Shen Y, Xie S, Zhao L, et al. Estimating individualized absolute risk for esophageal squamous cell carcinoma: A population-based study in high-risk areas of China. Front Oncol 2021;10:598603. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wang QL, Lagergren J, Xie SH. Prediction of individuals at high absolute risk of esophageal squamous cell carcinoma. Gastrointest Endosc 2019;89(4):726–32.e2. [DOI] [PubMed] [Google Scholar]
  • 23.Han J, Wang L, Zhang H, et al. Development and validation of an esophageal squamous cell carcinoma risk prediction model for rural Chinese: Multicenter cohort study. Front Oncol 2021;11:729471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Han J, Guo X, Zhao L, et al. Development and validation of esophageal squamous cell carcinoma risk prediction models based on an endoscopic screening program. JAMA Netw Open 2023;6(1):e2253148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wang QL, Ness-Jensen E, Santoni G, et al. Development and validation of a risk prediction model for esophageal squamous cell carcinoma using cohort studies. Am J Gastroenterol 2021;116(4):683–91. [DOI] [PubMed] [Google Scholar]
  • 26.Steyerberg EW, Vergouwe Y. Towards better clinical prediction models: Seven steps for development and an ABCD for validation. Eur Heart J 2014;35(29):1925–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ramspek CL, Jager KJ, Dekker FW, et al. External validation of prognostic models: What, why, how, when and where? Clin Kidney J 2020;14(1):49–58. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Vergouwe Y, Nieboer D, Oostenbrink R, et al. A closed testing procedure to select an appropriate method for updating prediction models. Stat Med 2017;36(28):4529–39. [DOI] [PubMed] [Google Scholar]
  • 29.Janssen KJ, Moons KG, Kalkman CJ, et al. Updating methods improved the performance of a clinical prediction model in new patients. J Clin Epidemiol 2008;61(1):76–86. [DOI] [PubMed] [Google Scholar]
  • 30.Chang J, Huang Y, Wei L, et al. Risk prediction of esophageal squamous-cell carcinoma with common genetic variants and lifestyle factors in Chinese population. Carcinogenesis 2013;34(8):1782–6. [DOI] [PubMed] [Google Scholar]
  • 31.Etemadi A, Abnet CC, Golozar A, et al. Modeling the risk of esophageal squamous cell carcinoma and squamous dysplasia in a high risk area in Iran. Arch Iran Med 2012;15(1):18–21. [PMC free article] [PubMed] [Google Scholar]
  • 32.Kunzmann AT, Thrift AP, Cardwell CR, et al. Model for identifying individuals at risk for esophageal adenocarcinoma. Clin Gastroenterol Hepatol 2018;16(8):1229–36.e4. [DOI] [PubMed] [Google Scholar]
  • 33.Yokoyama T, Yokoyama A, Kumagai Y, et al. Health risk appraisal models for mass screening of esophageal cancer in Japanese men. Cancer Epidemiol Biomarkers Prev 2008;17(10):2846–54. [DOI] [PubMed] [Google Scholar]
  • 34.Li H, Ding C, Zeng H, et al. Improved esophageal squamous cell carcinoma screening effectiveness by risk-stratified endoscopic screening: Evidence from high-risk areas in China. Cancer Commun (Lond) 2021;41(8):715–25. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from The American Journal of Gastroenterology are provided here courtesy of Wolters Kluwer Health

RESOURCES