Skip to main content
Frontiers in Artificial Intelligence logoLink to Frontiers in Artificial Intelligence
. 2026 Sep 9;9:1837887. doi: 10.3389/frai.2026.1837887

Artificial intelligence-assisted endoscopic diagnosis of esophageal squamous cell carcinoma

Jie Mao 1,2,†, Kexun Li 2,3,*,†, Zilong Qian 2,†, Jianzhe Zhang 2,†, Shengguai Gao 2,†, GuoMin Tian 4, Daiheng Yang 2, Xin Tang 2, Xin Yang 2, Can Li 4, Yapeng Xing 4, Jian Xu 4, Chengwei Bi 4,*
PMCID: PMC13598749  PMID: 42780771

Abstract

Esophageal squamous cell carcinoma (ESCC) remains a major cause of cancer-related mortality, and prognosis depends strongly on detection at a curable stage. Endoscopy is central to screening and diagnosis, but subtle flat lesions, operator dependence, cognitive fatigue, lesion-location blind spots, and variability in interpretation contribute to missed or delayed diagnosis. Artificial intelligence (AI), particularly deep learning applied to white-light imaging, narrow-band imaging, blue-light imaging, magnifying endoscopy, Lugol chromoendoscopy, video endoscopy, and endocytoscopy, has shown clinically meaningful potential for lesion detection, margin delineation, invasion-depth estimation, and microvascular-pattern classification. Following peer-review feedback, this article has been reframed as a structured mini-review and evidence appraisal rather than a formal systematic review or meta-analysis. We summarize representative studies published mainly from 2019 to 2026, describe the literature-identification scope and eligibility criteria, and critically appraise the evidence using domains adapted from diagnostic-accuracy, prediction-model, and medical-imaging AI reporting frameworks. In this review, multimodal AI refers specifically to complementary endoscopic inputs rather than routine integration of histopathological or genomic data into deployed ESCC endoscopic systems. Current evidence suggests that many AI systems achieve high sensitivity in enriched image datasets, and several video-based, prospective, or randomized studies support translational feasibility. However, major limitations remain: most studies are retrospective; many use single-center or high-quality still-image datasets; confidence intervals, calibration, uncertainty quantification, subgroup analysis, failure-mode reporting, and latency benchmarks are inconsistent; and external validation across devices, operators, patient spectra, and live-video workflows remains limited.

Keywords: artificial intelligence, convolutional neural networks, early detection, endoscopic diagnosis, esophageal squamous cell carcinoma

1. Introduction

Esophageal cancer (EC) is a major global health burden. In East Asia, including China and Japan, ESCC is the predominant histological subtype and accounts for most EC cases (Siegel et al., 2023; Han et al., 2024). Although overall incidence rates vary by region and birth cohort, the absolute number of older patients remains substantial because of population aging (Lu et al., 2025). The survival gap between early and advanced disease is striking: early-stage ESCC can often be treated endoscopically with favorable outcomes, whereas advanced-stage disease carries a poor prognosis (Zhao et al., 2024; Pennathur et al., 2013). Therefore, improving early detection remains a high-value clinical objective.

Endoscopic screening and diagnosis rely on white-light imaging (WLI), chromoendoscopy, narrow-band imaging (NBI), blue-light imaging (BLI), magnifying endoscopy (ME), and endocytoscopy. Endoscopic screening has demonstrated population-level clinical benefit in randomized evidence (Liu et al., 2024). These modalities are complementary, but all are constrained by lesion subtlety, image quality, operator experience, withdrawal technique, and cognitive fatigue. Reported missed-lesion rates and post-endoscopy cancer events indicate that even high-quality endoscopy is not immune to error (Zhang, 2017; Chadwick et al., 2014; Zhou et al., 2025; Hassan et al., 2025; Zhang et al., 2020; Shimizu et al., 2008; Muto et al., 2010; Rodriguez de Santiago et al., 2019). AI-based image analysis has therefore been investigated as a second observer or decision-support tool to improve consistency in detection, characterization, and therapeutic planning. Prior AI diagnostic reviews and umbrella reviews have reported promising accuracy across endoscopic tasks, while also emphasizing heterogeneity, reporting limitations, and the need for robust validation (Zhang et al., 2021; Zha et al., 2024). This review focuses on AI-assisted endoscopic diagnosis of ESCC, including lesion detection, margin delineation, depth prediction, microvascular classification, multimodal integration, explainability, failure modes, and regulatory translation.

This review is a structured mini-review and evidence appraisal designed for clinically oriented synthesis rather than quantitative evidence aggregation (Figure 1).

Figure 1.

Infographic illustrating artificial intelligence-assisted endoscopic diagnosis of esophageal squamous cell carcinoma, detailing convolutional neural networks for real-time analysis, multimodal data integration, dynamic diagnostic algorithms, and an AI-assisted diagnostic workflow with a timeline and a sensitivity and specificity comparison chart.

Virtual abstract of the structured mini-review and evidence appraisal of AI-assisted endoscopic diagnosis of esophageal squamous cell carcinoma.

2. Review methodology and evidence-appraisal approach

2.1. Scope and literature-identification strategy

The objective of this structured mini-review was to summarize and critically appraise AI-assisted endoscopic diagnosis of ESCC and squamous precancerous lesions. The review was informed by PRISMA 2020 transparency principles for search reporting (Page et al., 2021), while acknowledging that the article is not a registered systematic review and does not include pooled diagnostic-accuracy estimates. To reduce selection bias from a PubMed-only approach, the literature-identification strategy was broadened beyond PubMed/MEDLINE to include biomedical and engineering-oriented sources relevant to medical AI, including Embase, Web of Science Core Collection, Scopus, IEEE Xplore, Google Scholar, reference-list checking, and citation tracking of key ESCC AI studies. Search concepts combined disease, modality, and AI terms: “esophageal squamous cell carcinoma”, “oesophageal squamous cell carcinoma”, “precancerous lesion,” “high-grade intraepithelial neoplasia”, “endoscopy,” “white-light imaging,” “narrow-band imaging,” “blue-light imaging,” “Lugol,” “magnifying endoscopy,” “endocytoscopy,” “artificial intelligence,” “machine learning,” “deep learning,” “convolutional neural network”, “YOLO”, “single-shot multibox detector”, “transformer”, “vision transformer”, “computer-aided diagnosis”, “real-time”, “video”, “multimodal”, “explainable AI”, “self-supervised learning”, and “foundation model.” The primary focus was on studies published from January 2019 to May 2026, with older landmark studies retained when necessary for clinical or technical context (Figure 2).

Figure 2.

Flowchart labeled PRISMA diagram outlining a systematic review process: 694 records identified, 218 removed before screening, 476 screened, 389 excluded, 87 sought for retrieval, 2 not retrieved, 85 assessed for eligibility, 47 excluded for specific reasons, and 38 studies included in qualitative synthesis.

Flowchart of literature inclusion and screening.

2.2. Eligibility criteria

Studies were considered eligible if they evaluated AI, machine learning, or deep learning systems for endoscopic detection, classification, lesion localization, margin delineation, invasion-depth estimation, IPCL/microvascular classification, real-time video assistance, or multimodal endoscopic diagnosis of ESCC or squamous precancerous lesions. Eligible data sources included WLI, NBI, BLI, ME-NBI/BLI, Lugol chromoendoscopy, endocytoscopy, still images, video clips, and prospectively acquired endoscopic examinations. We included retrospective development/validation studies, external-validation studies, prospective observational studies, randomized trials, and clinically relevant technical studies with explicit performance metrics or translational relevance. Exclusion criteria were: non-endoscopic imaging only; esophageal adenocarcinoma without separable ESCC data; non-human or purely phantom studies; papers without identifiable AI methods or diagnostic outcomes; abstracts without sufficient methodological detail; duplicated reports from the same dataset unless a later study provided additional validation; and publications not relevant to clinical endoscopic diagnosis.

2.3. Data extraction, synthesis, and quality appraisal

For each representative study, we extracted the diagnostic task, imaging modality, model family, dataset composition, study design, validation design, unit of analysis, performance metrics, comparison with endoscopists, and translational limitations. Because modalities, case mix, annotation standards, outcome definitions, and thresholds differed markedly across studies, results were summarized narratively rather than pooled. When 95% confidence intervals (CIs) were unavailable in the source reports, the tables identify the limitation as “CI NR” or describe the absence of precision estimates rather than implying statistical certainty. Evidence quality and risk of bias were appraised narratively using domains adapted from QUADAS-2, CLAIM 2024, STARD-AI, CONSORT-AI, SPIRIT-AI, and TRIPOD+AI: patient/case selection, reference standard, index-test definition, separation of training and testing data, patient-level rather than image-level splitting, external validation, class balance, annotation process, robustness to artifacts and low-quality frames, calibration/uncertainty reporting, interpretability, real-time workflow evaluation, and post-deployment governance (Tejani et al., 2024; Sounderajah et al., 2025; Whiting et al., 2011; Liu et al., 2020; Rivera et al., 2020; Collins et al., 2024). To address the need for stronger methodological contextualization, we emphasized transparent source selection, extractable validation data, prespecified thresholds, external validation, model/code availability where feasible, and explicit reporting of bias, uncertainty, calibration, and clinical workflow context.

To make the evidence tables reproducible, studies were featured in Table 1 only when the source report provided, at minimum, an ESCC-relevant diagnostic task, endoscopic modality, model family, validation material or sample size, and at least one extractable performance benchmark such as sensitivity, specificity, accuracy, AUC, F1 score, PPV/NPV, or boundary-delineation accuracy. Table 2 was restricted to studies with video-based, real-time, prospective, randomized, or workflow-simulation evidence relevant to clinical translation. No absolute minimum sample size was imposed because several early ESCC AI studies are proof-of-concept investigations; instead, sample size, validation design, and missing precision estimates are displayed as methodological limitations. Studies without extractable validation material or diagnostic outcomes were excluded from the summary tables and discussed only narratively when they provided relevant technical context.

Table 1.

Standardized summary of representative retrospective and validation studies in AI-assisted endoscopic diagnosis of ESCC.

Study Task/modality Architecture Dataset/validation Performance reported Main limitation or risk-of-bias note
Horie et al. (2019) WLI; EC/superficial vs. advanced cancer CNN 8,428 images from 384 patients; retrospective Accuracy 98%; all lesions <10 mm detected Curated still images; limited benign controls; CI NR
Cai et al. (2019) WLI; early ESCC screening CNN-based CAD 2,428 training images; 187 validation images Sensitivity 97.8%; specificity 85.4%; accuracy 91.4%; AUC > 0.96 Small validation set; controls mainly normal mucosa; CI NR
Tang et al. (2021) WLI; early ESCC vs. reflux/normal DCNN 4,002 training/cross-validation images; 1,033 validation images Sensitivity 97.9%; specificity 88.6%; AUC 0.954 Multicenter validation but retrospective; device/operator shift not fully explored; CI NR
Feng et al. (2023) WLI; superficial ESCC CNN 5,892 training images; 4,529 validation images AI assistance improved accuracy 75.12 to 84.95% Retrospective; patient-spectrum and site-generalization concerns
Li et al. (2021) NM-NBI versus WLI; early ESCC CAD-NBI and CAD-WLI 2,167 abnormal and 2,568 normal NM-NBI images for training CAD-NBI: sensitivity 91.0%; specificity 96.7%; accuracy 94.3%; CAD-WLI: sensitivity 98.5%; specificity 83.1% Useful modality comparison; retrospective; CI NR
Wang et al. (2021) WLI/NBI; histological grade classification SSD 936 images; 264 test images Tumor detection sensitivity 96.2%; specificity 70.4%; accuracy 90.9%; grading accuracy 92% Pilot size; moderate specificity; benign mimics remain challenging
Everson et al. (2019) ME-NBI; IPCL classification CNN 7,046 images from 17 patients Accuracy 93.7%; sensitivity 89.3%; specificity 98.0% Proof-of-concept; very small patient count; CI ranges reported in limited form
Everson et al. (2021) ME-NBI; clinically interpretable IPCL prediction Interpretable CNN 67,742 high-quality ME-NBI images; expert comparison CNN F1 94.0%; experts 97.0–98.0% High image volume but still-image emphasis; dependence on expert-defined reference
Kumagai et al. (2019) Endocytoscopy; ESCC diagnosis GoogLeNet CNN 4,715 training images; 1,520 validation images AUC 0.85 overall; sensitivity 92.6%; specificity 89.3%; accuracy 90.9% ECS-specific; image quality and device availability limit generalization
Kumagai et al. (2022) Endocytoscopy; modified type classification DeiT transformer 7,983 training images; 114 test images from 38 cases AUC 0.92; patient-level accuracy 94.7% Small test set; limited direct comparison with CNNs
Ohmori et al. (2020) WLI, NBI/BLI, ME; superficial ESCC detection Deep neural network 9,591 non-ME and 7,844 ME training images WLI: sensitivity 90.0%; specificity 76.0%; non-ME NBI/BLI: sensitivity 100.0%; specificity 63.0%; ME-NBI/BLI: sensitivity 98.0%; specificity 56.0% High sensitivity but modest specificity; thresholds and modality-specific errors need analysis
Meng et al. (2022) WLI/NBI; superficial ESCC/HGIN vs. benign YOLOv5 CAD Training: 622 ESCC/HGIN cases and 215 non-cancer cases; independent testing AUC 0.982; accuracy 92.9%; sensitivity 91.9%; specificity 94.7% Stronger benign-control inclusion; retrospective and dataset-specific
Liu et al. (2022) WLI; lesion-margin delineation Segmentation/localization AI 13,083 images; internal, external, and endoscopist comparison sets Internal/external boundary accuracy 93.4%/95.7% Boundary reference standard and interobserver variability should be explicit
Yuan X. L. et al. (2023) NBI; detection and extent delineation AI detection/delineation system 10,047 still images and 140 videos from four hospitals; prospective evaluation Detection accuracy 92.4%/89.9% internal/external; delineation 88.9%/87.0%; prospective 91.4%/85.9% Video and prospective data strengthen evidence; workflow and false-alert burden need further reporting
Nakagawa et al. (2019) Non-ME/ME; invasion-depth classification SSD Training: 804 patients; validation: 155 patients Sensitivity 90.1%; specificity 95.8%; PPV 99.2%; NPV 63.9%; accuracy 91.0% Low NPV limits rule-out use; retrospective design
Shimamoto et al. (2020) Video; invasion-depth assessment Video AI system 23,977 images; 102 video clips; expert comparison Sensitivity 71.0% vs. experts 50.0%; specificity 99.0% vs. experts 95.0% in one comparison Video-based evidence; clip selection and latency/generalization require more detail
Urabe et al. (2025) Histological surface images; invasive vs. non-invasive regions Deep learning patch model Surface histology images from early ESCC specimens AUC 0.869; accuracy 78.8% Mechanistic insight rather than direct endoscopic deployment

Table 2.

Real-time, video-based, prospective clinical, and translational framework evidence relevant to ESCC AI translation.

Study Clinical/technical setting Validation material Key finding Translational limitation
Guo et al. (2020) Real-time detection of precancerous lesions and early ESCC; NBI images and video clips 6,473 training images; 6,671 testing images; 80 video clips Probability heatmaps closely matched endoscopist-marked areas Requires prospective workflow validation and systematic false-positive analysis
Fukuda et al. (2020) Real-time video detection and characterization of superficial ESCC; non-ME and ME-NBI/BLI 28,333 images from 862 videos; 144 short test clips AI exceeded experts in detection sensitivity (91.0% vs. 79.0%) and characterization sensitivity (86.7% vs. 74.4%) Specificity varied by modality; deployment latency and alert stability require reporting
Shimamoto et al. (2020) Real-time video assessment of invasion depth 23,977 images; 102 video clips; comparison with 14 experts Higher sensitivity and specificity than experts in selected comparisons Clinical decision thresholds and external prospective impact remain unclear
Tajiri et al. (2022) Simulated clinical classification of ESCC and non-cancerous esophageal lesions; NBI videos 25,048 ESCC images; 4,746 non-cancer images; 147 NBI clips Accuracy 80.9%; sensitivity 85.5%; specificity 75.0%; exceeded endoscopist averages Moderate specificity indicates false positives in benign conditions
Yang et al. (2021) Real-time AI for early ESCC diagnosis Clinical real-time endoscopy setting Demonstrated feasibility of real-time assistance Single-study workflow; standard latency/artifact benchmarks needed
Shiroma et al. (2021) T1 ESCC detection from videos and effect of real-time assistance Endoscopic videos Reported video-detection feasibility and assistance effect Video selection and patient-level clinical outcomes need standardization
Waki et al. (2021) Detection of overlooked ESCC in videos simulating miss scenarios Video-simulation evaluation Supported use of AI as a second observer Simulated setting may not represent true screening prevalence
Tani et al. (2023) Prospective real-time usefulness study Single-center prospective setting Prospective workflow assessment Single-center design limits generalizability
Yuan et al. (2024) Multicenter tandem double-blind randomized trial of AI-assisted superficial ESCC/precancerous-lesion diagnosis Multicenter real clinical setting; WLI and NM-NBI Primary outcome: lesion and patient miss rate; trial completed Among strongest translational evidence, but cost–benefit, implementation burden, frame-rate/latency reporting, and long-term outcomes remain to be assessed
Kelly et al. (2019) Broader clinical-AI translation framework; not ESCC-specific Narrative synthesis of clinical AI deployment challenges Emphasizes prospective evaluation, representative test sets, clinically meaningful metrics, dataset shift, generalization, bias assessment, and post-market surveillance Framework reference only; does not provide ESCC-specific diagnostic performance
Nakao et al. (2025) Randomized controlled trial of an AI diagnostic system in clinical practice Clinical-practice randomized design Represents emerging RCT-level ESCC AI evidence Details of performance by lesion subtype, device, and operator level should guide deployment

2.4. Methodological limitations of this review

This review has several methodological limitations that should be made explicit. First, because it is a structured mini-review rather than a registered systematic review, it does not provide duplicate independent screening statistics or meta-analytic pooling. Second, heterogeneity in primary-study design and reporting limits direct comparison: many studies report image-level metrics from enriched datasets, whereas clinical adoption requires patient-level and lesion-level outcomes in consecutive real-world procedures. Third, full reproducibility is limited because training code, trained weights, annotation manuals, negative-case composition, endoscope-device metadata, threshold-setting rules, and CIs are often unavailable; this caveat is consistent with AI reporting guidance that emphasizes complete model, data, validation, calibration, and uncertainty documentation (Tejani et al., 2024; Sounderajah et al., 2025; Collins et al., 2024).

3. Conventional endoscopic diagnosis and unresolved clinical needs

The standard diagnostic sequence for suspected early ESCC commonly begins with WLI or non-magnifying NBI/BLI to identify suspicious mucosal changes, followed by chromoendoscopy or ME-NBI/BLI to characterize lesion extent, microvascular morphology, and invasion risk. WLI is widely available and efficient but can miss subtle flat lesions or lesions with minimal color contrast. Lugol iodine staining increases sensitivity but has practical limitations, including longer procedure time, mucosal irritation, patient discomfort, and lower specificity in inflammatory or regenerative mucosa. ME-NBI improves visualization of IPCL patterns, yet interpretation depends heavily on training and experience (Hassan et al., 2025; Zhang et al., 2020; Shimizu et al., 2008; Muto et al., 2010; Rodriguez de Santiago et al., 2019; Inoue et al., 2015).

These clinical limitations create several AI-relevant use cases. First, a detection model could act as a real-time second observer for subtle lesions. Second, a classification model could reduce variability in distinguishing non-neoplastic inflammation from dysplasia or superficial cancer. Third, segmentation and heatmap models could support margin delineation before endoscopic resection. Fourth, invasion-depth prediction could inform the choice between endoscopic resection, surgery, and chemoradiotherapy. Finally, quality-control algorithms could identify blind spots, poor focus, excessive motion, or inadequate inspection time.

4. AI-assisted endoscopic diagnosis of ESCC

4.1. Single-modality image diagnosis

Most early ESCC AI studies used CNN-based architectures on still endoscopic images. WLI-based models demonstrated high sensitivity in retrospective datasets. Horie et al. trained a CNN on 8,428 EC images from 384 patients and reported 98% accuracy in distinguishing superficial from advanced cancer, with detection of small lesions under 10 mm (Horie et al., 2019). Cai et al. (2019) developed a WLI CAD system that achieved sensitivity of 97.8%, specificity of 85.4%, accuracy of 91.4%, and AUC greater than 0.96 in a validation set, improving performance especially among junior endoscopists. Tang et al. (2021) subsequently included reflux esophagitis and normal mucosa in multicenter validation and reported sensitivity of 0.979, specificity of 0.886, and AUC of 0.954. Feng et al. (2023) further reported a generalized WLI-based CNN system for superficial ESCC detection, showing that AI assistance improved diagnostic accuracy and specificity in endoscopist comparison experiments. These studies support diagnostic feasibility, but many were retrospective, used curated high-quality frames, and did not consistently report CIs, calibration, or patient-level generalization (Table 1).

NBI and BLI provide enhanced mucosal and vascular contrast. Li et al. (2021) compared CAD systems for NM-NBI and WLI and found that CAD-NBI achieved higher specificity and accuracy, whereas CAD-WLI showed higher sensitivity and negative predictive value. Wang et al. (2021) used SSD to classify histological grades from WLI and NBI images, reporting high sensitivity but only moderate specificity, highlighting the challenge of differentiating cancer from benign or inflammatory mimics. ME-NBI studies further focused on IPCL pattern recognition. Everson et al. demonstrated proof-of-concept CNN classification of normal and abnormal IPCL patterns and later expanded to a larger dataset, with performance approaching expert endoscopists (Everson et al., 2019; Everson et al., 2021). Other ME-NBI or microvascular studies have evaluated CAD, microvessel classification, and AI-assisted IPCL identification for early-stage ESCC (Zhao et al., 2019; Uema et al., 2021; Yuan et al., 2022b; Wang et al., 2023). Endocytoscopic studies using GoogLeNet and later DeiT further showed the potential of AI-assisted optical histology, but external validation and clinical workflow studies remain limited (Kumagai et al., 2019; Kumagai et al., 2022).

4.2. Multimodal endoscopic AI and fusion strategies

Multimodal AI is particularly relevant to ESCC because WLI, NBI, BLI, ME, iodine staining, endocytoscopy, still images, and continuous video emphasize different but complementary lesion features. In this review, multimodal should be understood as multimodal endoscopic imaging rather than combined endoscopy-histopathology-genomics modeling. Histopathology remains the reference standard for diagnosis and depth assessment, and genomic or molecular data may become valuable for future risk stratification, but the currently reviewed ESCC diagnostic systems are designed primarily for endoscopic image or video decision support. Ohmori et al. (2020) integrated WLI, NBI, BLI, and ME data and achieved high sensitivity across modalities, although specificity varied by modality. Meng et al. (2022) developed a YOLOv5-based CAD system for WLI and NBI that achieved AUC 0.982, accuracy 92.9%, sensitivity 91.9%, and specificity 94.7%, with improved performance among non-expert endoscopists when AI assistance was provided. Additional studies have expanded this direction through multiple endoscopic imaging modalities, YOLO-based clinical validation, hybrid modeling, transfer learning, and risk stratification for multiple Lugol-voiding lesions (Ikenoyama et al., 2021; Yuan et al., 2022a; Wang et al., 2022; Chou et al., 2023).

Technically, multimodal fusion can be organized at several levels. Early fusion concatenates aligned image channels or modality-specific feature maps before deep feature extraction, but it requires reliable registration and is vulnerable to missing or asynchronous modalities. Intermediate fusion uses parallel modality-specific encoders followed by shared attention, cross-attention, gated feature exchange, feature-pyramid integration, or modality-dropout training; this approach can preserve modality-specific information while learning complementary representations and robustness to absent modalities. Late fusion combines independent model outputs by voting, stacking, calibration-weighted averaging, or uncertainty-aware decision rules; it is simpler to deploy but may not fully exploit cross-modal interactions. For ESCC, temporal fusion across consecutive video frames is also important because a suspicious area may only be visible briefly during peristalsis, insufflation, or angle changes. General multimodal machine-learning frameworks distinguish early, intermediate, and late fusion and emphasize that fusion strategies should be chosen according to modality alignment, missing-modality robustness, and output-level calibration; ESCC models should therefore justify the level of fusion and report whether fusion improves robustness rather than only point accuracy (Baltrusaitis et al., 2019).

4.3. Lesion localization, margin delineation, invasion-depth estimation, and microvascular classification

Beyond binary detection, AI systems are increasingly designed to localize lesions and support therapeutic decision-making. Guo et al. (2020) generated probability heatmaps to delineate suspected lesion boundaries on NBI images and video clips. Yuan X. et al. (2023). reported real-time multimodal delineation of small flat-type ESCC. Liu et al. (2022) reported boundary-delineation accuracies above 90% in internal and external validation. Yuan X. L. et al. (2023) subsequently evaluated detection and delineation under NBI using still images and videos from multiple hospitals, with promising internal, external, and prospective performance. These studies are clinically important because incomplete lesion-margin assessment can affect endoscopic resection planning. However, boundary “truth” is not always straightforward; histopathology, expert consensus, iodine staining, and endoscopic resection specimens may disagree, and inter-annotator variability should be reported (Table 2).

Invasion-depth estimation is another high-impact task because submucosal invasion changes treatment strategy. Conventional IPCL-based depth assessment is clinically useful but has observer-dependence and reproducibility constraints (Sato et al., 2015). Nakagawa et al. (2019) used SSD-based modeling and reported sensitivity of 90.1%, specificity of 95.8%, and accuracy of 91.0% for distinguishing mucosal microinvasive cancer from submucosal invasive cancer. Related CNN-based and human-like ME-NBI systems have also been developed for ESCC invasion-depth prediction (Tokai et al., 2020; Zhang et al., 2023). Shimamoto et al. (2020) evaluated video-based invasion-depth assessment and reported higher sensitivity and specificity than expert endoscopists in some comparisons. Urabe et al. (2025) used AI to detect histological differences between invasive and non-invasive regions and showed that vascular features were associated with invasive patches. These findings suggest that AI may identify subtle morphologic and vascular correlates of depth, but negative predictive value, calibration, lesion selection, and decision thresholds require further validation before AI can guide definitive treatment selection.

4.4. Real-time video diagnosis and clinical deployment constraints

Real-time AI is the key translational step from retrospective image performance to clinical utility. Still-image datasets often exclude the most difficult frames: motion blur, defocus, bubbles, mucus, bleeding, overexposure, underexposure, peristalsis, folds, rapid scope movement, variable insufflation, and partial lesion visualization. Video-based models must therefore operate robustly under temporal noise and must distinguish true lesions from transient artifacts. Fukuda et al. trained an SSD-based model using images extracted from hundreds of videos and evaluated short video clips, reporting higher sensitivity than expert endoscopists for detection and characterization (Fukuda et al., 2020). Tajiri et al. (2022) evaluated an AI system on NBI video clips simulating clinical use and found improved accuracy and sensitivity compared with endoscopists, but specificity remained moderate. Additional real-time and video studies have examined early ESCC diagnosis, T1 ESCC video detection, and overlooked-lesion simulation, reinforcing both the feasibility and the need for standardized latency, artifact, and false-alert reporting (Yang et al., 2021; Shiroma et al., 2021; Waki et al., 2021).

Deployment also requires engineering, human-factors, and workflow validation. The model must process frames fast enough for the AI overlay to remain synchronized with endoscopic motion; bounding boxes or heatmaps should be temporally stable rather than flickering; false positives should not create alarm fatigue; and the interface must allow endoscopists to accept, reject, or override suggestions without disrupting inspection technique. Latency should be reported as an end-to-end quantity that includes image capture, preprocessing, inference, post-processing, display rendering, and network or hardware delays. Hardware integration differs across processors, endoscope vendors, image-enhancement modes, image resolutions, hospital networks, and computing accelerators. Continuous monitoring is needed because model drift can occur when devices, software versions, patient populations, procedural techniques, or training programs change. Prospective and randomized ESCC studies, including clinical-practice and multicenter tandem designs, are therefore more informative than image-level retrospective testing alone (Yuan et al., 2024; Tani et al., 2023; Nakao et al., 2025); broader clinical-AI translation guidance further emphasizes representative test sets, clinically meaningful endpoints, generalization, bias assessment, and post-market surveillance (Kelly et al., 2019). Broader randomized trials of AI-assisted gastric neoplasm and colon polyp detection provide additional translational precedent for workflow-integrated evaluation beyond offline image accuracy (Wu et al., 2021; Wang et al., 2019).

4.5. CNNs, transformers, and hybrid architectures

CNNs remain the dominant architecture in ESCC endoscopic AI because they efficiently learn local texture, color, vascular, and edge features and can be adapted to detection frameworks such as SSD and YOLO. Their locality bias and parameter efficiency are advantageous when datasets are relatively small, as is common in medical imaging. However, CNNs can be limited in modeling long-range spatial dependencies, global context, and relationships between non-contiguous mucosal regions. Transformer-based models, including ViT, DeiT, and Swin Transformer variants, use self-attention to model global dependencies and may better capture context across a broad field of view, but they are data-hungry, computationally demanding, and more prone to overfitting when training cohorts are small or device-specific (Kumagai et al., 2022; Shamshad et al., 2023; Li et al., 2023). General multimodal and transformer literature provides useful methodological vocabulary, but domain-specific gains cannot be extrapolated to ESCC without direct endoscopic validation (Baltrusaitis et al., 2019; Shamshad et al., 2023; Li et al., 2023).

Only limited ESCC-specific evidence directly compares CNNs and transformers under identical data partitions. Kumagai et al. reported a DeiT-based endocytoscopic system with AUC 0.92 and patient-level accuracy 94.7%, suggesting feasibility but not proving architectural superiority (Kumagai et al., 2022). In broader medical imaging, transformer and hybrid CNN-transformer architectures are increasingly used for classification, segmentation, and multimodal diagnosis (Shamshad et al., 2023; Li et al., 2023). Reporting guidance for medical-imaging AI and AI prediction models shows why architectural comparisons should include external validation, locked testing, calibration and uncertainty, reproducibility, and deployment information, not only accuracy (Tejani et al., 2024; Collins et al., 2024). For ESCC, the most plausible near-term direction is hybrid design: CNN stems for efficient local texture extraction, attention modules for global context, temporal modules for video stability, and uncertainty-aware heads for calibrated clinical output; hybrid transfer-learning studies in esophageal endoscopic detection are therefore particularly relevant (Chou et al., 2023). Claims that transformers outperform CNNs should be reserved for studies using the same patient-level external test sets, the same threshold-locking rules, and the same real-time hardware constraints.

Architecture selection should therefore be matched to the diagnostic use case rather than listed as interchangeable model names. Standard CNN patch-classifiers are well suited to offline still-image classification, optical-histology classification, or confirmation of cropped regions because they can learn local color, texture, and vascular features with relatively modest data requirements. However, patch classifiers usually require a separate candidate-generation or sliding-window step and may process only selected frames; this can increase end-to-end latency and create unstable outputs when the endoscope is moving. By contrast, one-stage detectors such as YOLO and SSD predict class probabilities and bounding boxes in a single forward pass, which makes them more appropriate for live video streams where the system must process sequential frames, maintain a stable overlay, and avoid perceptible delay. Their speed advantage is clinically meaningful only when reported as end-to-end FPS and latency on the actual deployment hardware, including capture, preprocessing, inference, post-processing, display rendering, and any network transmission. The trade-off is that aggressive model compression, lower input resolution, or high detection thresholds may miss very small, flat, or partially visualized lesions, whereas segmentation networks may better support margin delineation but often require heavier computation and denser annotation.

The same speed-accuracy trade-off should be handled as a reporting requirement rather than as an architecture claim. For live ESCC workflows, lightweight backbones, feature-pyramid or multiscale modules, temporal smoothing, and one-stage detector heads may reduce latency, but their clinical value depends on locked validation in artifact-rich videos and transparent reporting of end-to-end performance. Future ESCC studies should therefore report parameter count, model size, FLOPs, input resolution, FPS, hardware accelerator, and latency alongside diagnostic accuracy, so that clinicians can judge whether a model is suitable for dynamic endoscopy rather than only for retrospective image benchmarking (Tejani et al., 2024; Collins et al., 2024; Kelly et al., 2019).

4.6. Dataset quality, annotation variability, and class imbalance

Dataset construction is a central determinant of AI validity. Many ESCC datasets overrepresent clear cancer images and underrepresent real-world negative or confounding conditions such as reflux esophagitis, radiation injury, post-endoscopic resection scars, varices, glycogenic acanthosis, candidiasis, vascular ectasia, submucosal lesions, white plaques, and benign Lugol-voiding areas. Models trained on such enriched data may show inflated sensitivity and specificity but fail under screening conditions where disease prevalence is low and benign abnormalities are common.

Annotation is also complex. Detection labels may be lesion-level, frame-level, or pixel-level; segmentation labels may be drawn from expert consensus, iodine staining, ME-NBI findings, or post-resection pathology; and invasion-depth labels depend on histopathological sampling and classification thresholds. Interobserver variability should be quantified using agreement statistics, and discordant labels should be adjudicated transparently. Patient-level splitting is essential to avoid leakage from near-duplicate frames or multiple frames from the same lesion appearing in both training and testing sets. Class imbalance should be addressed through transparent sampling, loss-function design, threshold tuning, and evaluation under clinically plausible disease prevalence. Future studies should report not only sensitivity, specificity, accuracy, and AUC, but also PPV, NPV, lesion miss rate, false alerts per procedure, calibration metrics, decision-curve analysis, and 95% CIs at patient or lesion level. Comparable practices are needed for ESCC endoscopic datasets, including transparent data provenance, locked test sets, external validation, and availability of code, models, or reproducible implementation details when feasible (Tejani et al., 2024; Collins et al., 2024).

4.7. Interpretability and explainable AI

Clinical trust requires more than high point estimates. Heatmaps, Grad-CAM, saliency maps, attention visualization, prototype-based explanations, uncertainty estimates, and counterfactual analysis can help clinicians understand which image regions influenced model output. For ESCC, interpretable outputs are most useful when they correspond to known endoscopic features such as demarcation lines, IPCL irregularity, brownish areas, microvascular changes, flat depressed morphology, or Lugol-voiding margins. However, explainability methods can be visually persuasive without being faithful to model reasoning. Broader critiques of explainable AI in health care caution that visually persuasive heatmaps or post-hoc explanations should not be treated as evidence of safety or faithful model reasoning (Ghassemi et al., 2021). XAI outputs should therefore be validated against expert lesion and margin annotations, tested for stability under image perturbation, and reported together with uncertainty or reject-option mechanisms. Interpretability should support—but never substitute for—prospective clinical validation.

4.8. Failure modes and safeguards

Failure-mode analysis is underreported in the ESCC AI literature. Potential false negatives include very small flat lesions, lesions at the upper esophageal sphincter or near folds, partially visualized lesions, low-contrast lesions under WLI, lesions obscured by mucus or bubbles, and frames affected by motion blur, poor focus, bleeding, overexposure, underexposure, or rapid scope withdrawal. Potential false positives include reflux esophagitis, radiation injury, inflammation, scars, glycogenic acanthosis, candidiasis, white plaques, vascular congestion, varices, benign submucosal lesions, and iodine-unstained non-neoplastic mucosa. False localization can occur when heatmaps highlight specular reflection, shadows, borders of folds, bubbles, or non-lesion texture rather than true pathologic tissue. Broader clinical-AI translation and robustness literature reinforces the need to describe rare but clinically consequential failure modes, including artifacts, domain shift, lesion mimics, unintended bias, and adversarial or spurious visual cues, rather than reporting average-case performance alone (Kelly et al., 2019; Finlayson et al., 2019).

Several mitigation strategies are being actively explored for these failure modes. Technical approaches include hard-negative mining with reflux, scar, radiation-injury, candidiasis, and specular-reflection examples; artifact-aware augmentation that simulates motion blur, defocus, mucus, bubbles, bleeding, variable illumination, and partial visualization; frame-quality classifiers that suppress or flag unusable frames; temporal smoothing, optical-flow tracking, or recurrent/transformer memory to require persistence across consecutive frames before an alarm is displayed; uncertainty-aware calibration and reject options for low-confidence predictions; and domain adaptation or federated learning to reduce device- and center-specific drift. For margin and depth tasks, combining detection with segmentation heatmaps, IPCL-aware features, and modality-specific confirmation under NBI/BLI or Lugol staining may reduce false localization and improve clinically interpretable outputs.

Clinical safeguards are equally important. Endoscopists can mitigate false negatives by cleaning mucus and bubbles, optimizing insufflation and focus, slowing withdrawal at the upper esophageal sphincter and near folds, switching between WLI and image-enhanced modes for suspicious low-contrast areas, and using AI as a second observer rather than as a replacement for systematic inspection. Recommended deployment safeguards include external validation across institutions and devices, stress testing with artifact-rich videos, stratified performance reporting by lesion size, morphology, location, and imaging modality, uncertainty thresholds or reject options, human override, audit logs, incident reporting, post-market monitoring, and periodic model updating under change-control governance. Such safeguards are consistent with broader clinical-AI translation, robustness, and regulatory lifecycle recommendations (Kelly et al., 2019; Finlayson et al., 2019; International Medical Device Regulators Forum, 2017; U.S. Food and Drug Administration, 2025; European Parliament and Council of the European Union, 2024; World Health Organization, 2021; U.S. Food and Drug Administration, 2021). Clinicians must retain responsibility for final diagnosis, biopsy/resection decisions, management of discordant AI-human judgments, and patient communication (Table 3).

Table 3.

Narrative quality and risk-of-bias appraisal using domains adapted from QUADAS-2, CLAIM, STARD-AI, CONSORT-AI, SPIRIT-AI, and TRIPOD+AI.

Domain Observed status in ESCC AI literature Risk/applicability concern Implication for this review and future studies
Search and selection transparency Earlier ESCC reviews often relied heavily on PubMed and narrative selection. Selection bias; incomplete coverage of engineering literature and conference-derived AI methods. Search scope, databases, terms, inclusion/exclusion criteria, and non-meta-analytic design are now stated; reproducibility-related reporting expectations are explicitly acknowledged (Collins et al., 2024).
Patient/case selection Many studies use enriched image datasets with high disease prevalence and high-quality frames. Spectrum bias; performance may be overestimated relative to screening. Interpret results as feasibility evidence; prioritize patient-level external validation.
Reference standard Histology is often used, but margin and invasion labels may vary by biopsy, resection specimen, iodine staining, and expert consensus. Mislabeling and verification bias; uncertain ground truth for boundaries. Report reference-standard hierarchy and interobserver agreement.
Training/testing separation Some papers report image-level splits; multiple frames from one lesion can be near-duplicates. Data leakage if patient-level separation is not enforced. Require patient-level and center-level separation for testing.
External validation External and multicenter validation is increasing but remains inconsistent. Domain shift across endoscope vendors, processors, centers, operators, and patient populations. Use multicenter external test sets and prospective deployment monitoring.
Metrics and precision Sensitivity, specificity, accuracy, and AUC are commonly reported, but CIs, calibration, and false alerts per procedure are inconsistently available. Uncertain precision and clinical utility; threshold-dependent results. Tables define CI NR when unavailable; future studies should report 95% CIs, calibration, decision thresholds, and false-alert burden.
Class imbalance and benign mimics Negative controls may be normal mucosa only; inflammatory and scar-like mimics may be underrepresented. False positives and false reassurance under real-world prevalence. Include clinically plausible negative-case spectra and benign mimics; report class balance, sampling strategy, prevalence assumptions, PPV/NPV under expected screening prevalence, and subgroup performance rather than relying on enriched image-level accuracy alone.
Video and real-time robustness Still-image performance may not survive motion, blur, mucus, bubbles, lighting changes, and rapid scope movement. Reduced real-time reliability and alert fatigue. Stress-test with real procedure videos and report end-to-end latency, frame rate, false alerts per procedure, overlay stability, artifact-specific performance, failure-mode taxonomies, and monitoring for domain shift or spurious visual cues (Kelly et al., 2019; Finlayson et al., 2019).
Interpretability Heatmaps and attention maps are sometimes shown but rarely validated quantitatively. Misleading explanations can increase unwarranted trust. Validate XAI outputs against expert lesion/margin annotations and combine heatmaps, attention maps, post-hoc explanation audits, and uncertainty estimation; visually persuasive explanations should be audited rather than accepted at face value (Ghassemi et al., 2021).
Clinical impact Prospective trials are emerging but remain fewer than retrospective studies. Unclear effect on miss rate, procedure time, biopsy/resection decisions, cost, and training. Use prospective, randomized, and workflow-integrated studies with predefined clinical outcomes, including lesion miss rate, procedure time, biopsy/resection decisions, false-alert burden, training effect, cost-effectiveness, and unintended consequences (Kelly et al., 2019; Vasey et al., 2022).
Regulatory and ethics Data governance, privacy, cybersecurity, post-market monitoring, and liability are variably discussed. Barriers to deployment and patient trust. Treat diagnostic AI as regulated SaMD; define human oversight, audit trails, model-update governance, and safety monitoring.

4.9. Regulatory, ethical, and implementation considerations

AI endoscopic systems intended for diagnosis or clinical decision support will generally be regulated as software as a medical device or as part of an integrated medical device, depending on jurisdiction and intended use. Regulatory expectations include analytical validity, clinical validity, usability, cybersecurity, risk management, human oversight, post-deployment monitoring, change-control plans for adaptive models, and transparent labeling of intended use. International and regional frameworks emphasize clinical evaluation of SaAI/ML device oversight, predetermined change-control planning, responsible governance for AI in health, and risk-based obligations for high-impact clinical AI systems (International Medical Device Regulators Forum, 2017; U.S. Food and Drug Administration, 2025; European Parliament and Council of the European Union, 2024; World Health Organization, 2021; U.S. Food and Drug Administration, 2021).

Ethically, ESCC AI systems must avoid widening disparities. Training and validation datasets should represent different patient populations, endoscope platforms, operator skill levels, clinical settings, and screening-risk profiles. Privacy-preserving data sharing, federated learning, secure annotation platforms, and de-identification pipelines may help expand multicenter datasets without exposing patient data. Liability should be defined for scenarios in which AI misses a lesion, generates excessive false positives, or influences an inappropriate clinical decision. Implementation should also address consent, transparency to patients where required, cybersecurity, auditability, fairness monitoring across subgroups, and clinician deskilling risk. These considerations are essential for responsible translation from retrospective image performance to safe clinical use.

5. Future directions

The next phase of ESCC AI development should move from proof-of-concept image classification toward validated clinical systems. First, multicenter prospective datasets should capture consecutive procedures rather than selected images. Metadata should include endoscope model, processor, imaging mode, resolution, frame rate, operator experience, lesion size, location, morphology, histology, treatment, and follow-up. Second, benchmarking should standardize patient-level, lesion-level, and frame-level metrics; report 95% CIs; separate tuning, validation, and locked testing; and include calibration, false alerts per procedure, time-to-detection, and decision thresholds. Third, real-time trials should measure workflow effects, including procedure duration, endoscopist attention, biopsy decisions, lesion miss rate, patient outcomes, cost-effectiveness, and unintended consequences such as over-biopsy or alert fatigue. These priorities are aligned with recent calls for transparency, reproducibility, workflow-integrated evaluation, explainable outputs, and explicit failure-mode reporting in healthcare AI (Tejani et al., 2024; Collins et al., 2024; Kelly et al., 2019; Vasey et al., 2022; Ghassemi et al., 2021; Finlayson et al., 2019).

Methodologically, self-supervised learning and foundation-model approaches may reduce dependence on expensive pixel-level annotations by learning representations from large unlabeled endoscopic video archives (Azizi et al., 2021). Such models could be fine-tuned for ESCC detection, margin delineation, depth prediction, microvascular classification, and quality control. Hybrid CNN-transformer networks, temporal transformers, multimodal cross-attention, federated learning, domain adaptation, synthetic minority oversampling, and uncertainty-aware calibration are also promising. At the same time, clinical deployment requires explicit attention to model efficiency and lifecycle governance: lightweight backbones, feature-pyramid or multiscale modules, temporal smoothing, and detector heads may help balance parameter count, multiscale feature capture, and real-time FPS in ESCC video workflows, but they must be validated under locked testing, documented change-control procedures, and post-deployment monitoring (Kelly et al., 2019; U.S. Food and Drug Administration, 2021). Nevertheless, technical novelty should be subordinated to clinical validity: a simpler model with transparent external validation and stable real-time performance may be preferable to a complex model that lacks generalization (Table 4).

Table 4.

Recommended reporting, benchmarking, and deployment checklist for future ESCC AI studies.

Domain Minimum information to report Why it matters for ESCC AI Reviewer concern addressed
Study design Retrospective/prospective design; single- or multicenter setting; consecutive versus enriched sampling; patient-, lesion-, and image-level units; prespecified threshold-locking rule; early clinical-evaluation stage and human-factor context where applicable (Vasey et al., 2022). Clarifies whether reported performance reflects real screening workflow or curated image feasibility. Methodological transparency and reproducibility (Collins et al., 2024)
Dataset composition Endoscope vendor, processor, imaging mode, resolution, frame rate, lesion size/location/morphology, benign mimics, negative-case spectrum, class balance, and cross-dataset validation strategy; broader medical-image AI examples should be interpreted with disease-domain differences in mind (Tejani et al., 2024; Collins et al., 2024). Prevents overestimation from high-quality cancer-enriched datasets and supports domain-shift analysis, while clarifying whether external datasets truly reflect ESCC screening conditions. Dataset quality, class imbalance, and generalizability
Reference standard and annotation Histology source; biopsy versus resection specimen; expert consensus method; number and experience of annotators; interobserver agreement; adjudication of discordant labels. Margin, IPCL, and invasion-depth labels are clinically complex and can introduce verification or labeling bias. Inclusion criteria, annotation variability, and risk of bias
Model and training details Architecture family; use-case mapping (classification, detection, segmentation, depth prediction, or quality control); preprocessing; augmentation; hyperparameters; training/validation/test split; patient-level leakage prevention; calibration method; parameter count, model size, FLOPs, input resolution, FPS, hardware environment, code/model availability or reasons for restricted sharing, and change-control assumptions for deployed models (Tejani et al., 2024; Collins et al., 2024; U.S. Food and Drug Administration, 2021). Allows meaningful comparison among CNN, transformer, YOLO/SSD, segmentation, lightweight, and hybrid systems and supports reproducibility; multimodal designs should report fusion strategy, missing-modality handling, and cross-modal contribution (Baltrusaitis et al., 2019). Architecture comparison, model-size trade-offs, multiscale feature capture, real-time feasibility, and reproducibility
Performance metrics Sensitivity, specificity, accuracy, AUC, PPV, NPV, lesion miss rate, false alerts per procedure, calibration, decision-curve analysis, and 95% CIs at patient or lesion level. Point estimates alone are insufficient for clinical adoption, especially under low-prevalence screening conditions. Statistical precision and standardized benchmarking
Real-time deployment End-to-end latency, FPS, overlay stability, hardware accelerator, memory footprint, artifact stress testing, motion/blur/mucus/bubble performance, user-interface design, temporal smoothing rules, lesion-detection failure-mode taxonomy, and post-deployment monitoring plan (Kelly et al., 2019; Finlayson et al., 2019; U.S. Food and Drug Administration, 2021). Retrospective still-image accuracy may not translate to live endoscopy; robustness failures can arise from artifacts, benign mimics, device shift, transient partial lesion visualization, spurious visual cues, or insufficient frame-rate handling (Kelly et al., 2019; Finlayson et al., 2019). Real-time AI diagnosis, robustness, and deployment constraints
Interpretability and safeguards Heatmap/XAI validation against expert annotations; post-hoc explanation audit where applicable; uncertainty estimates; reject options; human override; audit logs; and safeguards for discordant AI-human judgments (Ghassemi et al., 2021; U.S. Food and Drug Administration, 2021). Clinicians need actionable explanations and safety mechanisms for uncertain or discordant predictions; XAI outputs should be treated as testable explanations rather than proof of model reasoning (Ghassemi et al., 2021). Explainability and clinical safeguards
Regulatory and ethics Intended use, SaMD status, privacy and cybersecurity controls, subgroup fairness, post-market monitoring, predetermined change-control plan, model-update governance, and liability pathway (International Medical Device Regulators Forum, 2017; U.S. Food and Drug Administration, 2025; European Parliament and Council of the European Union, 2024; World Health Organization, 2021; U.S. Food and Drug Administration, 2021). Responsible deployment requires safety, accountability, and continuous monitoring beyond initial accuracy testing. Regulatory and ethical considerations

Implementation should follow a human-centered pathway. AI output should be designed for endoscopist comprehension, with adjustable alert thresholds, stable visual cues, clear confidence information, and easy override. Training programs should teach clinicians both how to use AI and how to recognize its limitations. Hospitals should define responsibility for local validation, software updates, adverse-event review, and model-performance surveillance. Post-market monitoring should track performance drift, failure modes, false-alert burden, missed lesions, cybersecurity incidents, and equity across patient subgroups.

6. Conclusion

AI-assisted endoscopic diagnosis of ESCC has progressed rapidly from retrospective still-image classification to multimodal endoscopic, video-based, and prospective clinical evaluation. Current evidence indicates that AI can improve sensitivity in selected datasets and may support less-experienced endoscopists, especially for subtle superficial lesions and precancerous changes. However, current evidence is limited by retrospective design, dataset enrichment, single-center development, inconsistent CI reporting, insufficient calibration and failure-mode analysis, and incomplete evaluation of real-time deployment constraints such as model size, frame-rate stability, and end-to-end latency. AI should therefore be implemented as an assistive second observer under clinician oversight rather than as an autonomous diagnostic authority.

Glossary

Glossary

EC

esophageal cancer

ESCC

esophageal squamous cell carcinoma

AI

artificial intelligence

CNN

convolutional neural networks

WLI

white light imaging

ME

magnifying endoscopy

NBI

narrow-band imaging

YOLO

you only look once

SSD

single-shot multibox detector

BLI

blue light imaging

CAD

computer-aided detection

AUC

area under the receiver operating characteristic curve

DCNN

deep convolutional neural network

NM-NBI

non-magnified narrow-band imaging

ME-NBI

magnifying narrow-band imaging

IPCL

the intrapapillary capillary loop

ESC

the endocytoscopic system

DeiT

the vision transformer, data-efficient image transformer

HGIN

high-grade intraepithelial neoplasia

HrEL

high-risk esophageal lesion

FPN

feature pyramid network

FPS

frames per second

FLOPs

floating-point operations

SaMD

software as a medical device

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by the Yunnan Provincial University Science and Technology Project for Serving Key Industries (FWCY-BSPY2024084); Yunnan Provincial Department of Science and Technology-Kunming Medical University Joint Fund for Applied Basic Research (202401AY070001-349); the Doctoral Student Educational Innovation Fund Project of Kunming Medical University (2025B014); Graduate Innovation Fund of Kunming Medical University (2024B026); Kunming Medical University “Hengrui Pharmaceutical Innovation and Development Clinical Transformation Special Project” (YQHR2025-M24).

Footnotes

Edited by: Andrea Cusano, University of Sannio, Italy

Reviewed by: Olaolu Olabintan, King's College Hospital NHS Foundation Trust, United Kingdom

Vaibhav C. Gandhi, Charutar Vidya Mandal University, India

Asif Raza, University of Mianwali, Pakistan

Author contributions

JM: Conceptualization, Data curation, Visualization, Writing – original draft, Writing – review & editing. KL: Conceptualization, Validation, Visualization, Writing – original draft, Writing – review & editing. ZQ: Visualization, Writing – review & editing. JZ: Conceptualization, Visualization, Writing – review & editing. SG: Conceptualization, Writing – review & editing. GT: Conceptualization, Writing – review & editing. DY: Conceptualization, Writing – review & editing. XT: Writing – review & editing. XY: Conceptualization, Writing – review & editing. CL: Conceptualization, Writing – review & editing. YX: Conceptualization, Writing – review & editing. JX: Conceptualization, Writing – review & editing. CB: Conceptualization, Validation, Visualization, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was used in the creation of this manuscript. During the preparation of this work, the author(s) used ChatGPT 5.2 to polish sentences. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  1. Azizi S., Mustafa B., Ryan F., Beaver Z., Freyberg J., Deaton J., et al. (2021). Big self-supervised models advance medical image classification Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) IEEE/CVF. 3478–3488 [Google Scholar]
  2. Baltrusaitis T., Ahuja C., Morency L. P. (2019). Multimodal machine learning: a survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41, 423–443. doi: 10.1109/TPAMI.2018.2798607, [DOI] [PubMed] [Google Scholar]
  3. Cai S. L., Li B., Tan W. M., Niu X. J., Yu H. H., Yao L. Q., et al. (2019). Using a deep learning system in endoscopy for screening of early esophageal squamous cell carcinoma (with video). Gastrointest. Endosc. 90, 745–753.e2. doi: 10.1016/j.gie.2019.06.044, [DOI] [PubMed] [Google Scholar]
  4. Chadwick G., Groene O., Hoare J., Hardwick R. H., Riley S., Crosby T. D. (2014). A population-based, retrospective, cohort study of esophageal cancer missed at endoscopy. Endoscopy 46, 553–560. doi: 10.1055/s-0034-1365646, [DOI] [PubMed] [Google Scholar]
  5. Chou C. K., Nguyen H. T., Wang Y. K., Chen T. H., Wu I. C., Huang C. W., et al. (2023). Preparing well for esophageal endoscopic detection using a hybrid model and transfer learning. Cancers (Basel) 15:3783. doi: 10.3390/cancers15153783, [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Collins G. S., Moons K. G. M., Dhiman P., Riley R. D., Beam A. L., Van Calster B., et al. (2024). TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385:e078378. doi: 10.1136/bmj-2023-078378, [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. European Parliament and Council of the European Union Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying down Harmonised rules on artificial Intelligence (Artificial Intelligence Act). Official Journal of the European Union (2024). Available online at: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (Accessed June 4, 2026).
  8. Everson M., Garcia-Peraza-Herrera L. C., Li W., Luengo I. M., Ahmad O., Banks M., et al. (2019). Artificial intelligence for the real-time classification of intrapapillary capillary loop patterns in the endoscopic diagnosis of early oesophageal squamous cell carcinoma: a proof-of-concept study. United European Gastroenterol J 7, 297–306. doi: 10.1177/2050640618821800 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Everson M. A., Garcia-Peraza-Herrera L. C., Wang H. P., Lee C. T., Chung C. S., Hsieh P. H., et al. (2021). A clinically interpretable convolutional neural network for the real-time prediction of early squamous cell cancer of the esophagus: comparing diagnostic performance with a panel of expert European and Asian endoscopists. Gastrointest. Endosc. 94, 273–281. doi: 10.1016/j.gie.2021.01.043, [DOI] [PubMed] [Google Scholar]
  10. Feng Y., Liang Y., Li P., Long Q., Song J., Li M., et al. (2023). Artificial intelligence assisted detection of superficial esophageal squamous cell carcinoma in white-light endoscopic images by using a generalized system. Discov. Oncol. 14:73. doi: 10.1007/s12672-023-00694-3, [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Finlayson S. G., Bowers J. D., Ito J., Zittrain J. L., Beam A. L., Kohane I. S. (2019). Adversarial attacks on medical machine learning. Science 363, 1287–1289. doi: 10.1126/science.aaw4399, [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Fukuda H., Ishihara R., Kato Y., Matsunaga T., Nishida T., Yamada T., et al. (2020). Comparison of performances of artificial intelligence versus expert endoscopists for real-time assisted diagnosis of esophageal squamous cell carcinoma (with video). Gastrointest. Endosc. 92, 848–855. doi: 10.1016/j.gie.2020.05.043, [DOI] [PubMed] [Google Scholar]
  13. Ghassemi M., Oakden-Rayner L., Beam A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit. Health 3, e745–e750. doi: 10.1016/S2589-7500(21)00208-9, [DOI] [PubMed] [Google Scholar]
  14. Guo L., Xiao X., Wu C., Zeng X., Zhang Y., Du J., et al. (2020). Real-time automated diagnosis of precancerous lesions and early esophageal squamous cell carcinoma using a deep learning model (with videos). Gastrointest. Endosc. 91, 41–51. doi: 10.1016/j.gie.2019.08.018, [DOI] [PubMed] [Google Scholar]
  15. Han B., Zheng R., Zeng H., Wang S., Sun K., Chen R., et al. (2024). Cancer incidence and mortality in China, 2022. J. Natl. Cancer Cent. 4, 47–53. doi: 10.1016/j.jncc.2024.01.006, [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Hassan C., Antonelli G., Chiu P. W., Emura F., Goda K., Iyer P. G., et al. (2025). Position statement of the world endoscopy organization: role of endoscopy in screening, diagnosis, and treatment of esophageal superficial squamous neoplasia. Dig. Endosc. 37, 470–489. doi: 10.1111/den.14967 [DOI] [PubMed] [Google Scholar]
  17. Horie Y., Yoshio T., Aoyama K., Yoshimizu S., Horiuchi Y., Ishiyama A., et al. (2019). Diagnostic outcomes of esophageal cancer by artificial intelligence using convolutional neural networks. Gastrointest. Endosc. 89, 25–32. doi: 10.1016/j.gie.2018.07.037, [DOI] [PubMed] [Google Scholar]
  18. Ikenoyama Y., Yoshio T., Tokura J., Naito S., Namikawa K., Tokai Y., et al. (2021). Artificial intelligence diagnostic system predicts multiple Lugol-voiding lesions in the esophagus and patients at high risk for esophageal squamous cell carcinoma. Endoscopy 53, 1105–1113. doi: 10.1055/a-1334-4053, [DOI] [PubMed] [Google Scholar]
  19. Inoue H., Kaga M., Ikeda H., Sato C., Sato H., Minami H., et al. (2015). Magnification endoscopy in esophageal squamous cell carcinoma: a review of the intrapapillary capillary loop classification. Ann. Gastroenterol. 28, 41–48. [PMC free article] [PubMed] [Google Scholar]
  20. International Medical Device Regulators Forum Software as a Medical Device (SaMD): Clinical Evaluation IMDRF/SaMD WG/N41FINAL: 2017 (2017) Available online at: https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation (Accessed June 4, 2026).
  21. Kelly C. J., Karthikesalingam A., Suleyman M., Corrado G., King D. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17:195. doi: 10.1186/s12916-019-1426-2, [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Kumagai Y., Takubo K., Kawada K., Aoyama K., Endo Y., Ozawa T., et al. (2019). Diagnosis using deep-learning artificial intelligence based on the endocytoscopic observation of the esophagus. Esophagus 16, 180–187. doi: 10.1007/s10388-018-0651-7, [DOI] [PubMed] [Google Scholar]
  23. Kumagai Y., Takubo K., Sato T., Ishikawa H., Yamamoto E., Ishiguro T., et al. (2022). AI analysis and modified type classification for endocytoscopic observation of esophageal lesions. Dis. Esophagus 35:doac010. doi: 10.1093/dote/doac010, [DOI] [PubMed] [Google Scholar]
  24. Li B., Cai S. L., Tan W. M., Li J. C., Yalikong A., Feng X. S., et al. (2021). Comparative study on artificial intelligence systems for detecting early esophageal squamous cell carcinoma between narrow-band and white-light imaging. World J. Gastroenterol. 27, 281–293. doi: 10.3748/wjg.v27.i3.281, [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Li J., Chen J., Tang Y., Wang C., Landman B. A., Zhou S. K. (2023). Transforming medical imaging with transformers? A comparative review of key properties, current progresses, and future perspectives. Med. Image Anal. 85:102762. doi: 10.1016/j.media.2023.102762, [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Liu X., Rivera S. C., Moher D., Calvert M. J., Denniston A. K., SPIRIT-AI and CONSORT-AI Working Group (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. BMJ 370:m3164. doi: 10.1136/bmj.m3164, [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Liu M., Yang W., Guo C., Liu Z., Li F., Liu A., et al. (2024). Effectiveness of endoscopic screening on esophageal cancer incidence and mortality: a 9-year report of the endoscopic screening for Esophageal Cancer in China (ESECC) randomized trial. J. Clin. Oncol. 42, 1655–1664. doi: 10.1200/JCO.23.01284, [DOI] [PubMed] [Google Scholar]
  28. Liu W., Yuan X., Guo L., Pan F., Wu C., Sun Z., et al. (2022). Artificial intelligence for detecting and delineating margins of early ESCC under WLI endoscopy. Clin. Transl. Gastroenterol. 13:e00433. doi: 10.14309/ctg.0000000000000433, [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Lu S., Li K., Wang K., Liu G., Han Y., Peng L., et al. (2025). Global trends of esophageal cancer among individuals over 60 years: an epidemiological analysis from 1990 to 2050 based on the global burden of disease study 1990-2021. Oncol. Rev. 19:1616080. doi: 10.3389/or.2025.1616080, [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Meng Q. Q., Gao Y., Lin H., Wang T. J., Zhang Y. R., Feng J., et al. (2022). Application of an artificial intelligence system for endoscopic diagnosis of superficial esophageal squamous cell carcinoma. World J. Gastroenterol. 28, 5483–5493. doi: 10.3748/wjg.v28.i37.5483, [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Muto M., Minashi K., Yano T., Saito Y., Oda I., Nonaka S., et al. (2010). Early detection of superficial squamous cell carcinoma in the head and neck region and esophagus by narrow band imaging: a multicenter randomized controlled trial. J. Clin. Oncol. 28, 1566–1572. doi: 10.1200/JCO.2009.25.4680, [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Nakagawa K., Ishihara R., Aoyama K., Ohmori M., Nakahira H., Matsuura N., et al. (2019). Classification for invasion depth of esophageal squamous cell carcinoma using a deep neural network compared with experienced endoscopists. Gastrointest. Endosc. 90, 407–414. doi: 10.1016/j.gie.2019.04.245, [DOI] [PubMed] [Google Scholar]
  33. Nakao E., Yoshio T., Kato Y., Namikawa K., Tokai Y., Yoshimizu S., et al. (2025). Randomized controlled trial of an artificial intelligence diagnostic system for the detection of esophageal squamous cell carcinoma in clinical practice. Endoscopy 57, 210–217. doi: 10.1055/a-2421-3194, [DOI] [PubMed] [Google Scholar]
  34. Ohmori M., Ishihara R., Aoyama K., Nakagawa K., Iwagami H., Matsuura N., et al. (2020). Endoscopic detection and differentiation of esophageal lesions using a deep neural network. Gastrointest. Endosc. 91, 301–309.e1. doi: 10.1016/j.gie.2019.09.034, [DOI] [PubMed] [Google Scholar]
  35. Page M. J., McKenzie J. E., Bossuyt P. M., Boutron I., Hoffmann T. C., Mulrow C. D., et al. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 372:n71. doi: 10.1136/bmj.n71, [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Pennathur A., Gibson M. K., Jobe B. A., Luketich J. D. (2013). Oesophageal carcinoma. Lancet 381, 400–412. doi: 10.1016/S0140-6736(12)60643-6, [DOI] [PubMed] [Google Scholar]
  37. Rivera S. C., Liu X., Chan A. W., Denniston A. K., Calvert M. J., SPIRIT-AI and CONSORT-AI Working Group (2020). Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. BMJ 370:m3210. doi: 10.1136/bmj.m3210, [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Rodriguez de Santiago E., Hernanz N., Marcos-Prieto H. M., de-Jorge-Turrion M., Barreiro-Alonso E., Rodriguez-Escaja C., et al. (2019). Rate of missed oesophageal cancer at routine endoscopy and survival outcomes: a multicentric cohort study. United Eur. Gastroenterol. J. 7, 189–198. doi: 10.1177/2050640618811477, [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Sato H., Inoue H., Ikeda H., Sato C., Onimaru M., Hayee B., et al. (2015). Utility of intrapapillary capillary loops seen on magnifying narrow-band imaging in estimating invasive depth of esophageal squamous cell carcinoma. Endoscopy 47, 122–128. doi: 10.1055/s-0034-1390858, [DOI] [PubMed] [Google Scholar]
  40. Shamshad F., Khan S., Zamir S. W., Khan M. H., Hayat M., Khan F. S., et al. (2023). Transformers in medical imaging: a survey. Med. Image Anal. 88:102802. doi: 10.1016/j.media.2023.102802, [DOI] [PubMed] [Google Scholar]
  41. Shimamoto Y., Ishihara R., Kato Y., Shoji A., Inoue T., Matsueda K., et al. (2020). Real-time assessment of video images for esophageal squamous cell carcinoma invasion depth using artificial intelligence. J. Gastroenterol. 55, 1037–1045. doi: 10.1007/s00535-020-01716-5, [DOI] [PubMed] [Google Scholar]
  42. Shimizu Y., Omori T., Yokoyama A., Yoshida T., Hirota J., Ono Y., et al. (2008). Endoscopic diagnosis of early squamous neoplasia of the esophagus with iodine staining: high-grade intra-epithelial neoplasia turns pink within a few minutes. J. Gastroenterol. Hepatol. 23, 546–550. doi: 10.1111/j.1440-1746.2007.04990.x, [DOI] [PubMed] [Google Scholar]
  43. Shiroma S., Yoshio T., Kato Y., Horie Y., Namikawa K., Tokai Y., et al. (2021). Ability of artificial intelligence to detect T1 esophageal squamous cell carcinoma from endoscopic videos and the effects of real-time assistance. Sci. Rep. 11:7759. doi: 10.1038/s41598-021-87405-6, [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Siegel R. L., Miller K. D., Wagle N. S., Jemal A. (2023). Cancer statistics, 2023. CA Cancer J. Clin. 73, 17–48. doi: 10.3322/caac.21763, [DOI] [PubMed] [Google Scholar]
  45. Sounderajah V., Guni A., Liu X., Collins G. S., Karthikesalingam A., Markar S. R., et al. (2025). The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat. Med. 31, 3283–3289. doi: 10.1038/s41591-025-03953-8, [DOI] [PubMed] [Google Scholar]
  46. Tajiri A., Ishihara R., Kato Y., Inoue T., Matsueda K., Miyake M., et al. (2022). Utility of an artificial intelligence system for classification of esophageal lesions when simulating its clinical use. Sci. Rep. 12:6677. doi: 10.1038/s41598-022-10739-2, [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Tang D., Wang L., Jiang J., Liu Y., Ni M., Fu Y., et al. (2021). A novel deep learning system for diagnosing early esophageal squamous cell carcinoma: a multicenter diagnostic study. Clin. Transl. Gastroenterol. 12:e00393. doi: 10.14309/ctg.0000000000000393, [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Tani Y., Ishihara R., Inoue T., Okubo Y., Kawakami Y., Matsueda K., et al. (2023). A single-center prospective study evaluating the usefulness of artificial intelligence for the diagnosis of esophageal squamous cell carcinoma in a real-time setting. BMC Gastroenterol. 23:184. doi: 10.1186/s12876-023-02788-2, [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Tejani A. S., Klontzas M. E., Gatti A. A., Mongan J. T., Moy L., Park S. H. (2024). Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol. Artif. Intell. 6:e240300. doi: 10.1148/ryai.240300, [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Tokai Y., Yoshio T., Aoyama K., Horie Y., Yoshimizu S., Horiuchi Y., et al. (2020). Application of artificial intelligence using convolutional neural networks in determining the invasion depth of esophageal squamous cell carcinoma. Esophagus 17, 250–256. doi: 10.1007/s10388-020-00716-x, [DOI] [PubMed] [Google Scholar]
  51. U.S. Food and Drug Administration (2021). Health Canada Medicines and Healthcare Products Regulatory Agency Good Machine Learning Practice for Medical Device Development: Guiding Principles. Silver Spring: U.S. Food and Drug Administration. [Google Scholar]
  52. U.S. Food and Drug Administration Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions: Guidance for Industry and Food and Drug Administration Staff (2025). Available online at: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence (Accessed June 28, 2026).
  53. Uema R., Hayashi Y., Tashiro T., Saiki H., Kato M., Amano T., et al. (2021). Use of a convolutional neural network for classifying microvessels of superficial esophageal squamous cell carcinomas. J. Gastroenterol. Hepatol. 36, 2239–2246. doi: 10.1111/jgh.15479, [DOI] [PubMed] [Google Scholar]
  54. Urabe A., Adachi M., Sakamoto N., Kojima M., Ishikawa S., Ishii G., et al. (2025). Deep learning detected histological differences between invasive and non-invasive areas of early esophageal cancer. Cancer Sci. 116, 824–834. doi: 10.1111/cas.16426, [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Vasey B., Nagendran M., Campbell B., Clifton D. A., Collins G. S., Denaxas S., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ 377:e070904. doi: 10.1136/bmj-2022-070904, [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Waki K., Ishihara R., Kato Y., Shoji A., Inoue T., Matsueda K., et al. (2021). Usefulness of an artificial intelligence system for the detection of esophageal squamous cell carcinoma evaluated with videos simulating overlooking situation. Dig. Endosc. 33, 1101–1109. doi: 10.1111/den.13934, [DOI] [PubMed] [Google Scholar]
  57. Wang P., Berzin T. M., Glissen B. J., Bharadwaj S., Becq A., Xiao X., et al. (2019). Real-time automatic detection system increases colonoscopic polyp and adenoma detection rates: a prospective randomised controlled study. Gut 68, 1813–1819. doi: 10.1136/gutjnl-2018-317500 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Wang S. X., Ke Y., Liu Y. M., Liu S. Y., Song S. B., He S., et al. (2022). Establishment and clinical validation of an artificial intelligence YOLOv5l model for the detection of precancerous lesions and superficial esophageal cancer in endoscopic procedure. Zhonghua Zhong Liu Za Zhi 44, 395–401. doi: 10.3760/cma.j.cn112152-20211126-00877 [DOI] [PubMed] [Google Scholar]
  59. Wang J., Long Q., Liang Y., Song J., Feng Y., Li P., et al. (2023). AI-assisted identification of intrapapillary capillary loops in magnification endoscopy for diagnosing early-stage esophageal squamous cell carcinoma: a preliminary study. Med. Biol. Eng. Comput. 61, 1631–1648. doi: 10.1007/s11517-023-02777-3, [DOI] [PubMed] [Google Scholar]
  60. Wang Y. K., Syu H. Y., Chen Y. H., Chung C. S., Tseng Y. S., Ho S. Y., et al. (2021). Endoscopic images by a single-shot multibox detector for the identification of early cancerous lesions in the esophagus: a pilot study. Cancers (Basel) 13:321. doi: 10.3390/cancers13020321, [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Whiting P. F., Rutjes A. W. S., Westwood M. E., Mallett S., Deeks J. J., Reitsma J. B., et al. (2011). QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 155, 529–536. doi: 10.7326/0003-4819-155-8-201110180-00009, [DOI] [PubMed] [Google Scholar]
  62. World Health Organization (2021). Ethics and Governance of artificial Intelligence for Health: WHO Guidance. Geneva: World Health Organization. [Google Scholar]
  63. Wu L., Shang R., Sharma P., Zhou W., Liu J., Yao L., et al. (2021). Effect of a deep learning-based system on the miss rate of gastric neoplasms during upper gastrointestinal endoscopy: a single-Centre, tandem, randomised controlled trial. Lancet Gastroenterol. Hepatol. 6, 700–708. doi: 10.1016/S2468-1253(21)00216-8, [DOI] [PubMed] [Google Scholar]
  64. Yang X. X., Li Z., Shao X. J., Ji R., Qu J. Y., Zheng M. Q., et al. (2021). Real-time artificial intelligence for endoscopic diagnosis of early esophageal squamous cell cancer (with video). Dig. Endosc. 33, 1075–1084. doi: 10.1111/den.13908, [DOI] [PubMed] [Google Scholar]
  65. Yuan X. L., Guo L. J., Liu W., Zeng X. H., Mou Y., Bai S., et al. (2022a). Artificial intelligence for detecting superficial esophageal squamous cell carcinoma under multiple endoscopic imaging modalities: a multicenter study. J. Gastroenterol. Hepatol. 37, 169–178. doi: 10.1111/jgh.15689, [DOI] [PubMed] [Google Scholar]
  66. Yuan X. L., Liu W., Lin Y. X., Deng Q. Y., Gao Y. P., Wan L., et al. (2024). Effect of an artificial intelligence-assisted system on endoscopic diagnosis of superficial oesophageal squamous cell carcinoma and precancerous lesions: a multicentre, tandem, double-blind, randomised controlled trial. Lancet Gastroenterol. Hepatol. 9, 34–44. doi: 10.1016/S2468-1253(23)00276-5, [DOI] [PubMed] [Google Scholar]
  67. Yuan X. L., Liu W., Liu Y., Zeng X. H., Mou Y., Wu C. C., et al. (2022b). Artificial intelligence for diagnosing microvessels of precancerous lesions and superficial esophageal squamous cell carcinomas: a multicenter study. Surg. Endosc. 36, 8651–8662. doi: 10.1007/s00464-022-09353-0, [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Yuan X., Zeng X., He L., Ye L., Liu W., Hu Y., et al. (2023). Artificial intelligence for detecting and delineating a small flat-type early esophageal squamous cell carcinoma under multimodal imaging. Endoscopy 55, E141–E142. doi: 10.1055/a-1956-0569, [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Yuan X. L., Zeng X. H., Liu W., Mou Y., Zhang W. H., Zhou Z. D., et al. (2023). Artificial intelligence for detecting and delineating the extent of superficial esophageal squamous cell carcinoma and precancerous lesions under narrow-band imaging (with video). Gastrointest. Endosc. 97, 664–672.e4. doi: 10.1016/j.gie.2022.12.003, [DOI] [PubMed] [Google Scholar]
  70. Zha B., Cai A., Wang G. (2024). Diagnostic accuracy of artificial intelligence in endoscopy: umbrella review. JMIR Med. Inform. 12:e56361. doi: 10.2196/56361, [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Zhang H. (2017). A systematic review and meta-analysis of missed squamous cell esophageal carcinoma after esophagogastroduodenoscopy [abstract]. Gastrointest. Endosc. 85:AB568. doi: 10.1016/j.gie.2017.03.1309 [DOI] [Google Scholar]
  72. Zhang R., Lau L., Wu P., Yip H. C., Wong S. H. (2020). Endoscopic diagnosis and treatment of esophageal squamous cell carcinoma. Methods Mol. Biol. 2129, 47–62. doi: 10.1007/978-1-0716-0377-2_5, [DOI] [PubMed] [Google Scholar]
  73. Zhang L., Luo R., Tang D., Zhang J., Su Y., Mao X., et al. (2023). Human-like artificial intelligent system for predicting invasion depth of esophageal squamous cell carcinoma using magnifying narrow-band imaging endoscopy: a retrospective multicenter study. Clin. Transl. Gastroenterol. 14:e00606. doi: 10.14309/ctg.0000000000000606, [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Zhang S. M., Wang Y. J., Zhang S. T. (2021). Accuracy of artificial intelligence-assisted detection of esophageal cancer and neoplasms on endoscopic images: a systematic review and meta-analysis. J. Dig. Dis. 22, 318–328. doi: 10.1111/1751-2980.12992, [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Zhao Y. Y., Xue D. X., Wang Y. L., Zhang R., Sun B., Cai Y. P., et al. (2019). Computer-assisted diagnosis of early esophageal squamous cell carcinoma using narrow-band imaging magnifying endoscopy. Endoscopy 51, 333–341. doi: 10.1055/a-0756-8754, [DOI] [PubMed] [Google Scholar]
  76. Zhao Y. X., Zhao H. P., Zhao M. Y., Yu Y., Qi X., Wang J. H., et al. (2024). Latest insights into the global epidemiological features, screening, early diagnosis and prognosis prediction of esophageal squamous cell carcinoma. World J. Gastroenterol. 30, 2638–2656. doi: 10.3748/wjg.v30.i20.2638, [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Zhou N., Yuan X., Liu W., Luo Q., Liu R., Hu B. (2025). Artificial intelligence in endoscopic diagnosis of esophageal squamous cell carcinoma and precancerous lesions. Chin. Med. J. 138, 1387–1398. doi: 10.1097/CM9.0000000000003490, [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Frontiers in Artificial Intelligence are provided here courtesy of Frontiers Media SA

RESOURCES