Abstract
Objectives
To evaluate the performance of machine learning models in predicting pathological complete response (pCR) to neoadjuvant chemoradiotherapy (nCRT) in patients with rectal cancer using magnetic resonance imaging.
Methods
We searched PubMed, Embase, Cochrane Library, and Web of Science for studies published before March 2024. The Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) was used to assess the methodological quality of the included studies, random-effects models were used to calculate sensitivity and specificity, I2 values were used for heterogeneity measurements, and subgroup analyses were carried out to detect potential sources of heterogeneity.
Results
A total of 1699 patients from 24 studies were included. For machine learning models in predicting pCR to nCRT, the meta-analysis calculated a pooled area under the curve (AUC) of 0.91 (95% CI, 0.88-0.93), pooled sensitivity of 0.83 (95% CI, 0.74-0.89), and pooled specificity of 0.86 (95% CI, 0.80-0.91). We investigated 6 studies that mainly contributed to heterogeneity. After performing meta-analysis again excluding these 6 studies, the heterogeneity was significantly reduced. In subgroup analysis, the pooled AUC of the deep-learning model was 0.93 and 0.89 for the traditional statistical model; the pooled AUC of studies that used diffusion-weighted imaging (DWI) was 0.90 and 0.92 in studies that did not use DWI; the pooled AUC of studies conducted in China was 0.93, and was 0.83 in studies conducted in other countries.
Conclusions
This systematic study showed that machine learning has promising potential in predicting pCR to nCRT in patients with locally advanced rectal cancer. Compared to traditional machine learning models, although deep-learning-based studies are less predominant and more heterogeneous, they are able to obtain higher AUC.
Advances in knowledge
Compared to traditional machine learning models, deep-learning-based studies are able to obtain higher AUC, although they are less predominant and more heterogeneous. Together with clinical information, machine learning-based models may bring us closer towards precision medicine.
Keywords: meta-analysis, rectal cancer, artificial intelligence, radiomics, deep learning, machine learning, neoadjuvant chemoradiotherapy
Introduction
Colorectal cancer (CRC) is the third leading cause of cancer death worldwide and rectal cancer accounts for approximately 30% of total CRC cases.1,2 For locally advanced rectal cancer, neoadjuvant chemoradiotherapy (nCRT) followed by mesorectal excision and adjuvant chemotherapy is the standard treatment guideline.3 nCRT is an effective regimen that reduces tumour size prior to surgery and increases tumour resection rate with reduced side effects.4,5 However, responses to nCRT are highly variable among patients. While approximately 10-15% of patients achieve pathological complete response (pCR) with no residual tumour cells on histological examination, a significant proportion of patients merely respond to treatment.6,7 Since inappropriate nCRT dosage may cause associated toxicities, it is essential to select the appropriate treatment dosage.8 To achieve such personalized treatment and thus optimize patient outcomes, effective prediction of treatment response to nCRT has become a fundamental step.
A substantial number of studies have been conducted to identify factors that predict response to nCRT based on clinical and histopathological information. However, such approaches lack a macroscopic view of tumour information such as lesion form and metastasis.9 With advances in diagnostic imaging and surgical planning, magnetic resonance imaging (MRI) has also become an efficient tool in assessing tumour response to nCRT, as it can provide detailed information, such as tumour morphology and texture. As machine learning has shown promising results in healthcare, clinical images are viewed more than just pictures, but as data that can be analysed through mathematical algorithms.10,11 Utilizing machine learning enables higher-level imaging analysis on MRI images and can further assist in predicting treatment response to nCRT to a more precise stage.
In this study, we aim to evaluate the performance of machine learning models in predicting pCR to nCRT using MRI in patients with rectal cancer. We assessed effectiveness of the models, evaluated methodological quality and bias risk in the machine learning workflows, and further performed subgroup analysis to identify potential important factors in predicting treatment response to nCRT.
Methods
Search strategy and selection criteria
We searched publications prior to March 2024, in PubMed, Embase, IEEE Xplore, and Cochrane Library databases using the following key words: rectal cancer, neoadjuvant, artificial intelligence, MRI, and machine learning. In addition, we also searched for publications using “radiomics,” “deep learning,” “neoadjuvant,” and “rectal neoplasms/rectal cancer” to avoid missing relevant articles. Furthermore, we also went through the reference lists of every article that meets our criteria to search for more relevant articles. Our search was restricted to publications in the English language.
Articles that met the following criteria were included: (1) all the patients with rectal cancer underwent nCRT; (2) developed machine learning models (deep learning, radiomics, etc.) to predict pCR to nCRT using MRI; and (3) provided sufficient information to construct a 2 × 2 contingency table. The exclusion criteria included: (1) the study included less than 10 patients; (2) abstracts, case reports, review articles, editorials, letters, and comments; (3) only performed training, without validation; (4) did not provide sensitivity or specificity; (5) did not provide any of the following quantifications: accuracy, positive predictive value (PPV), or negative predictive value (NPV).
To avoid missing articles and misinformation, the literature screening process was done by 2 of the authors independently. Selected articles were compared and finalized upon discussion. The 2 authors then extracted data independently and finalized the extracted data upon discussion.
The extraction of data and assessment of quality
After carefully selecting 24 studies that meet the inclusion criteria, we extracted the study and machine leaning algorithm characteristics from these articles. The study characteristics include: year of publication, region of the study population (country), study type (prospective or retrospective), number of patients included in the study, MRI field intensity, MRI sequences, machine learning architecture, and validation type. The machine learning algorithm characteristics include: segmentation method, feature extraction software, image feature used for analysis, algorithm architecture, sensitivity, specificity, accuracy, PPV, and NPV.
The methodological quality (“bias and applicability risks”) of the included studies was assessed using the QUADAS-2 tool12 by 2 authors independently. Disagreements were resolved upon consensus. Publication bias across the included studies was evaluated by the Deek test, diagnostic odds ratio (DOR), and by checking the symmetry of the funnel plot.
Statistical analysis
Statistical analyses were performed using StataSE 16. All the analysis was performed on the test or validation cohorts of the studies. We used random-effects models to calculate pooled sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), with 95% confidence intervals (CIs). To measure heterogeneity, we performed inconsistency index (I2) calculations. To identify key studies that caused heterogeneity, we further conducted the meta-analysis 24 times using the leave-one-out strategy and performed meta-analysis again after excluding identified studies. Furthermore, subgroup analyses that regard to MRI sequences, algorithm, country, validation cohort, and sample size were carried out to detect potential sources of heterogeneity. All the statistical analyses, including the subgroup analyses, were planned in advance.
Results
Study selection
A total of 1123 records were screened and 55 articles that use machine learning to predict tumour response to nCRT use were selected for possible inclusion. Due to the lack of validation, sensitivity or specificity, or accuracy-related information, 31 articles were excluded. Finally, 24 studies were selected following the inclusion and exclusion criteria. Detailed information regarding the article selection process is presented in Figure 1.
Figure 1.
Flow diagram of the study selection process.
Study characteristics
The included 24 studies were published between 2019 and March 2024. A total of 1699 patients were included in these 24 studies, with sample size ranging from 14 to 306. Thirteen studies were based on populations from China, 4 from Italy, 2 from Brazil and South Korea, and 1 from Japan and the United States. Three studies were prospective studies, while the remaining studies were conducted retrospectively. Sixteen studies acquired MRI before nCRT, 1 study acquired MRI after nCRT, and 5 of the studies acquired MRI before and after nCRT. Ten studies used both 1.5- and 3.0-T MRI scanners, 7 studies used only a 3.0-T MRI scanner, and 3 studies only used a 1.5-T MRI scanner, with image slice thicknesses ranging from 2.0 to 5.0 mm. Twenty-three studies included T2-weighted MRI and eleven studies utilized diffusion-weighted imaging (DWI). Over half of the studies (14/24) used 2 or more sequences to build predictive models.
The most commonly used segmentation software was ITK-SNAP (9/24) and 3D-slicer (5/24).13,14 Most studies performed manual segmentation (18/24) and most of the commonly used image feature extraction software is PyRadiomics (7/24).15 Fourteen studies used texture features to construct predictive models. Seven studies used convolutional neural networks to construct the models, while the remaining 17 studies used statistical models to predict responses to nCRT, and the most common statistical model was logistic regression. Eleven studies used external validation and thirteen studies used cross-validation. Table 1 summarizes details of the study characteristics.
Table 1.
Characteristics and details of the studies included in the meta-analysis.
| Author | Country | Study type | No. of patients | Field intensity | Sequences | Algorithm architecture | Validation |
|---|---|---|---|---|---|---|---|
| Antunes 2021 | USA | Retrospective | 104 | 1.5-3.0 T | T2 | RF | External (3 centres) |
| Arianna 2022 | Italy | Retrospective | 95 | 1.5 T | T2, DWI | Logistic regression | External (3 centres) |
| Boldrini 2022 | China | Retrospective | 220 | 1.5-3.0 T | T2 | Logistic regression | External (2 centres) |
| Bulens 2021 | Belgium | Retrospective | 125 | 3 T | T2, DWI | LASSO regression | External (2 centres) |
| Cheng 2021 | China | Retrospective | 193 | 3 T | T1, T2, T2FS | Logistic regression | Split sample (1 centre) |
| Cui 2019 | China | Retrospective | 186 | 3 T | CET1, T2, ADC | LASSO regression | Split sample (1 centre) |
| Feng 2022 | China | Prospective | 100 | Not reported | T1, T2, DWI | SVM | External (4 centres) |
| Ferrari 2019 | Italy | Retrospective | 55 | Not reported | T2 | RF | Split sample (1 centre) |
| Horvat 2018 | Brazil | Retrospective | 114 | 1.5-3.0 T | T2, DWI | RF | Split sample (1 centre) |
| Horvat 2022 | Brazil | Retrospective | 164 | 1.5-3.0 T | T2, DWI | RF | External (2 centres) |
| Huang 2020 | China | Retrospective | 270 | Not reported | Not reported | ANN | Split sample (1 centre) |
| Jang 2021 | Korea | Retrospective | 466 | 1.5-3.0 T | T1, T2 | CNN + LSTM | Split sample (1 centre) |
| Jin 2021 | China | Prospective | 141 | 1.5-3.0 T | T1, T2, DWI | 3D RP-Net | External (2 centres) |
| Nardone 2022 | Italy | Retrospective | 100 | 1.5 T | T2, ADC, DWI | Logistic regression | External (3 centres) |
| Ouyang 2022 | China | Retrospective | 118 | 3.0 T | T2 | Logistic regression | Split sample (1 centre) |
| Qiurong 2022 | China | Retrospective | 151 | 3.0 T | T1, T2, DWI | RF | External (2 centres) |
| Rengo 2022 | Italy | Retrospective | 95 | 1.5-3 T | T2 | KNN | External (2 centres) |
| Shin 2022 | South Korea | Retrospective | 898 | 1.5-3.0 T | T2, DWI | Radiomics | Split sample (1 centre) |
| Wei 2023 | China | Retrospective | 151 | 3 T | T2, DWI, DCE-T1, T1 | Radiomics | Internal and external (2 centres) |
| Wen 2023 | China | Retrospective | 126 | 1.5 T, 3 T | T2 | Radiomics | Split sample (1 centre) |
| Yardimci 2023 | Japan | Retrospective | 76 | 1.5 T | T2 | ANN | Split sample (1 centre) |
| Yi 2019 | China | Retrospective | 134 | 1.5-3.0 T | T2 | SVM | Split sample (1 centre) |
| Zhang 2019 | China | Prospective | 104 | 3.0 T | T1, T2, DKI | CNN | External (2 centres) |
| Zhu 2022 | China | Retrospective | 472 | 3 T | DWI | CNN | Split sample (1 centre) |
Quality assessment and publication bias
The methodological quality of the included studies is summarized in Figure 2. According to the analysis based on QUADAS-2, in patient selection, all the studies had low risk of bias, 2 studies showed unclear applicability concerns, and 22 studies showed no applicability concerns. In the index test, 9 (38%) studies showed unclear risk of bias and 15 (62%) studies had low bias risk; 1 study showed unclear applicability concerns and 23 showed no applicability concerns. In reference standard, all studies showed low risk of bias; 1 showed high risk of applicability concerns, 1 showed unclear risk of applicability concerns, while the remaining showed low applicability concerns. With regard to flow and timing, 1 study showed high bias risk, while 23 (96%) studies showed low bias risk. This study assessed publication bias across the included studies by performing Deeks funnel plot asymmetry test (P = .94). The funnel plot in Figure 3 shows no significant statistical asymmetry, which indicates a low risk of publication bias.
Figure 2.
Summary of QUADAS-2 assessments of bias and applicability risks of included studies: proportion of studies with low, high, or unclear risk of bias.
Figure 3.

Funnel plot of publication bias.
Meta-analysis
A total of 24 studies16–38 were included in the meta-analysis and we only evaluated the test or validation cohorts of these studies. For each study, we summarized sensitivity, specificity, and AUROC, with 95% CIs on a per-patient basis. The result is summarized in Figure 4: the pooled sensitivity was 0.83 (95% CI, 0.74-0.89), the pooled specificity was 0.86(95%CI, 0.80-0.91), the pooled positive likelihood ratio was 6.0 (95% CI, 4.0-8.9), the pooled negative likelihood ratio was 0.20 (95% CI, 0.13-0.30), and DOR was 30 (95% CI, 16-54).
Figure 4.
Coupled forest plot, with all 24 studies.
The summary receiver operating characteristics (SROC) analysis showed an area under the curve (AUC) of 0.91 (95% CI, 0.88-0.93). According to the SROC analysis summarized in Figure 5, the studies showed noticeable discrepancies between the 95% CI and the 95% prediction areas from the curve, which indicated significant variation probabilities among these studies. Through Q-test, our study found statistically significant heterogeneity in the pooled sensitivity (I2 = 79.70) and pooled specificity (I2 = 83.94) when calculating the above metrics.
Figure 5.

SROC curve and heterogeneity analysis. SROC, summary receiver operating characteristics.
To identify key studies that caused heterogeneity, we conducted the meta-analysis 24 times using the leave-one-out strategy where 1 study was left out each time when conducting the meta-analysis. The results of each analysis are summarized in Table 2 and Figure 6. Using the leave-one-out strategy, we found that heterogeneity was significantly reduced after removing 6 studies from the meta-analysis. Furthermore, we performed the meta-analysis again after removing these 6 studies. As shown by the results in Figure 5, the pooled sensitivity reduced to 0.82, the pooled specificity reduced to 0.84, and the Q-test result shows no statistically significant heterogeneity in pooled sensitivity (I2 = 35.06) or in pooled specificity (I2 = 65.25).
Table 2.
Summary of the results of the leave-one-out meta-analysis to detect source of heterogeneity.
| Author | No. of patients | No. of patients analysed | Sensitivity | Q | I 2 | Specificity | Q | I 2 | AUC |
|---|---|---|---|---|---|---|---|---|---|
| Total | 1699 | 1690 | 0.83 (0.74-0.89) | 113.3 | 79.70 (72.10-87.30) | 0.86 (0.80-0.91) | 143.26 | 83.94 (78.31-89.58) | 0.91 (0.88-0.93) |
| Zhu2022 | 472 | 1227 | 0.82 (0.73-0.89) | 109.18 | 79.85 (72.15-87.55) | 0.86 (0.79-0.91) | 137.63 | 84.02 (78.29-89.74) | 0.91 (0.88-0.93) |
| Zhang2019 | 104 | 1595 | 0.81 (0.73-0.87) | 101.42 | 78.31 (69.85-86.76) | 0.85 (0.78-0.90) | 118.42 | 81.42 (74.48-88.36) | 0.90 (0.87-0.92) |
| Yi2019 | 134 | 1565 | 0.83 (0.74-0.89) | 133.86 | 80.68 (73.38-87.97) | 0.86 (0.79-0.91) | 141.38 | 84.44 (78.91-89.97) | 0.91 (0.88-0.93) |
| Yardimci2023 | 76 | 1623 | 0.83 (0.74-0.89) | 113.47 | 80.61 (73.28-87.94) | 0.86 (0.79-0.91) | 143.5 | 84.67 (79.24-90.10) | 0.91 (0.89-0.93) |
| Wen2023 | 126 | 1573 | 0.82 (0.73-0.89) | 113.07 | 80.54 (73.18-87.90) | 0.87 (0.80-0.91) | 143.12 | 84.63 (79.18-90.07) | 0.91 (0.89-0.94) |
| Wei2023 | 151 | 1548 | 0.83 (0.74-0.89) | 113.74 | 80.66 (73.35-87.96) | 0.86 (0.79-0.91) | 144.01 | 84.72 (79.32-90.13) | 0.91 (0.89-0.94) |
| Shin2022 | 898 | 801 | 0.83 (0.74-0.89) | 107.75 | 79.58 (71.76-87.41) | 0.87 (0.80-0.92) | 115.73 | 80.99 (73.84-88.14) | 0.92 (0.89-0.94) |
| Rengo2022 | 95 | 1604 | 0.82 (0.74-0.89) | 110.28 | 80.05 (72.45-87.65) | 0.85 (0.79-0.90) | 132.6 | 83.41 (77.40-89.41) | 0.91 (0.88-0.93) |
| Qiurong2022 | 151 | 1548 | 0.83 (0.74-0.89) | 113.5 | 80.62 (73.29-87.94) | 0.86 (0.79-0.91) | 141.64 | 84.47 (78.95-89.99) | 0.91 (0.89-0.93) |
| Ouyang2022 | 118 | 1581 | 0.83 (0.75-0.89) | 114.29 | 80.75 (73.49-88.01) | 0.86 (0.79-0.91) | 144.06 | 84.73 (79.33-90.13) | 0.91 (0.89-0.94) |
| Nardone2022 | 100 | 1599 | 0.82 (0.74-0.89) | 113.41 | 80.60 (73.27-87.93) | 0.87 (0.80-0.91) | 145.41 | 84.87 (79.53-90.21) | 0.91 (0.88-0.93) |
| Jin2021 | 141 | 1558 | 0.82 (0.73-0.88) | 107.75 | 79.58 (71.76-87.41) | 0.86 (0.79-0.91) | 133.99 | 83.58 (77.65-89.51) | 0.91 (0.88-0.93) |
| Jang2021 | 466 | 1233 | 0.84 (0.77-0.89) | 77.24 | 71.52 (59.58-83.46) | 0.85 (0.78-0.90) | 122.36 | 82.02 (75.37-88.68) | 0.91 (0.89-0.94) |
| Huang2020 | 270 | 1429 | 0.82 (0.73-0.88) | 100.02 | 78.01 (69.40-86.61) | 0.87 (0.82-0.91) | 120.48 | 81.74 (74.95-88.53) | 0.92 (0.89-0.94) |
| Horvat2022 | 164 | 1535 | 0.84 (0.76-0.90) | 106.17 | 79.28 (71.30-87.25) | 0.86 (0.79-0.90) | 134.56 | 83.65 (77.76-89.54) | 0.91 (0.89-0.94) |
| Horvat2018 | 114 | 1585 | 0.82 (0.74-0.89) | 111.06 | 80.19 (72.66-87.72) | 0.86 (0.79-0.91) | 140.7 | 84.36 (78.80-89.93) | 0.91 (0.88-0.93) |
| Ferrari2019 | 55 | 1644 | 0.82 (0.74-0.89) | 112.11 | 80.38 (72.93-87.82) | 0.87 (0.80-0.91) | 144.31 | 84.75 (79.37-90.14) | 0.91 (0.89-0.94) |
| Feng2022 | 100 | 1599 | 0.82 (0.73-0.89) | 111.16 | 80.21 (72.69-87.73) | 0.87 (0.80-0.91) | 143.93 | 84.71 (79.31-90.12) | 0.91 (0.89-0.94) |
| Cui2019 | 186 | 1513 | 0.81 (0.73-0.88) | 101.16 | 78.25 (69.77-86.74) | 0.86 (0.79-0.91) | 137.52 | 84 (78.27-89.73) | 0.91 (0.88-0.93) |
| Cheng2021 | 193 | 1506 | 0.82 (0.73-0.88) | 108.85 | 79.79 (72.06-87.52) | 0.86 (0.79-0.91) | 145.9 | 84.92 (79.61-90.23) | 0.91 (0.88-0.93) |
| Bulens2021 | 125 | 1574 | 0.84 (0.76-0.90) | 96.45 | 77.19 (68.16-86.20) | 0.85 (0.78-0.90) | 132.22 | 83.36 (77.33-89.39) | 0.91 (0.89-0.94) |
| Boldrini2022 | 220 | 1479 | 0.83 (0.74-0.89) | 114.93 | 80.86 (73.65-88.07) | 0.87 (0.81-0.92) | 132.53 | 83.40 (77.39-89.41) | 0.92 (0.89-0.94) |
| Arianna2022 | 95 | 1604 | 0.83 (0.74-0.89) | 114.1 | 80.72 (73.44-87.99) | 0.87 (0.80-0.91) | 145.03 | 84.83 (79.48-90.18) | 0.92 (0.89-0.94) |
| Antunes2020 | 104 | 1595 | 0.83 (0.75-0.89) | 113.05 | 80.54 (73.18-87.90) | 0.86 (0.80-0.91) | 146.46 | 84.98 (79.69-90.27) | 0.92 (0.89-0.94) |
Figure 6.
Coupled forest plot after removing the 6 studies that caused heterogeneity.
Subgroup analysis
The meta-analysis results indicated heterogeneity in the combined data. To further investigate sources of heterogeneity, we performed a subgroup analysis of the 24 included studies based on 5 factors: MRI sequence, machine learning algorithm, country, validation cohort, and sample size. We performed 10 subgroup analyses, with each subgroup showing a different yet critical diagnostic value (Table 3). The subgroup analysis is summarized as follows:
Table 3.
Results of the subgroup analysis.
| Num. of study | Sensitivity | I 2 | Specificity | I 2 | PLR | NLR | DOR | AUROC | |
|---|---|---|---|---|---|---|---|---|---|
| Overall | 24 | 0.83 (0.74-0.89) | 79.70 (72.10-87.30) | 0.86 (0.80-0.91) | 83.94 (78.31-89.58) | 6.0 (4.0-8.9) | 0.20 (0.13-0.30) | 30 (16-55) | 0.91 (0.88-0.93) |
| Country | |||||||||
| China | 11 | 0.66 (0.52-0.77) | 77.15 (63.93-90.38) | 0.87 (0.77-0.93) | 83.19 (74.21-92.16) | 4.9 (3.1-7.8) | 0.40 (0.29-0.54) | 12 (8-20) | 0.83 (0.80-0.86) |
| Others | 13 | 0.89 (0.83-0.93) | 47.87 (14.42-81.33) | 0.84 (0.74-0.91) | 83.88 (76.10-91.67) | 5.5 (3.3-9.3) | 0.13 (0.09-0.20) | 41 (19-91) | 0.93 (0.90-0.95) |
| Sample size | |||||||||
| >100 | 6 | 0.85 (0.62-0.95) | 92.46 (87.98-96.95) | 0.90 (0.79-0.95) | 94.58 (91.65-97.51) | 8.2 (3.8-17.8) | 0.17 (0.06-0.49) | 49 (10-237) | 0.94 (0.91-0.96) |
| ≤100 | 18 | 0.81 (0.72-0.88) | 69.61 (54.61-84.29) | 0.84 (0.75-0.90) | 73.31 (60.85-85.78) | 5.1 (3.3-7.8) | 0.22 (0.15-0.33) | 23 (13-40) | 0.89 (0.86-0.92) |
| Sequences | |||||||||
| Not DWI | 13 | 0.86 (0.74-0.93) | 83.686 (76.06-91.66) | 0.86 (0.73-0.93) | 86.49 (80.28-92.7) | 6.0 (3.0-11.7) | 0.17 (0.09-0.32) | 36 (13-102) | 0.92 (0.90-0.94) |
| DWI | 11 | 0.79 (0.66-0.88) | 74.6 (59.48-89.71) | 0.87 (0.79-0.92) | 82.16 (72.48-91.84) | 6.0 (3.9-9.4) | 0.24 (0.15-0.40) | 25 (12-49) | 0.90 (0.87-0.93) |
| Segmentation | |||||||||
| Not manual | 7 | 0.80 (0.63-0.90) | 89.31 (82.87-95.75) | 0.85 (0.64-0.95) | 93.02 (89.30-96.74) | 5.4 (2.1-13.5) | 0.24 (0.13-0.44) | 23 (8-67) | 0.89 (0.86-0.91) |
| Manual | 17 | 0.84 (0.74-0.91) | 72.31 (58.78-85.76) | 0.86 (0.80-0.91) | 73.19 (60.28-86.10) | 6.1 (4.0-9.3) | 0.19 (0.11-0.31) | 33 (16-70) | 0.92 (0.89-0.94) |
| Algorithm | |||||||||
| Statistic model | 18 | 0.79 (0.69-0.87) | 67.62 (50.75-84.50) | 0.84 (0.76-0.89) | 72.82 (59.27-86.37) | 4.9 (3.3-7.3) | 0.25 (0.16-0.38) | 20 (10-38) | 0.89 (0.86-0.91) |
| Neural network | 6 | 0.86 (0.71-0.94) | 89.73 (84.05-95.41) | 0.88 (0.74-0.95) | 90.46 (85.29-95.62) | 7.3 (3.2-16.6) | 0.16 (0.07-0.34) | 46 (15-145) | 0.93 (0.91-0.95) |
| Validation | |||||||||
| Internal | 12 | 0.85 (0.73-0.92) | 84.28 (76.41-92.15) | 0.83 (0.72-0.90) | 84.77 (77.22-92.32) | 4.9 (3.0-7.8) | 0.18 (0.11-0.32) | 26 (13-54) | 0.90 (0.88-0.93) |
| External | 12 | 0.80 (0.66-0.89) | 73.30 (57.93-88.68) | 0.89 (0.80-0.94) | 82.67 (73.76-91.58) | 7.1 (3.8-13.4) | 0.23 (0.13-0.40) | 31 (12-84) | 0.92 (0.89-0.94) |
Abbreviations: AUROC = area under the receiver operating characteristics; DOR = diagnostic odds ratio; NLR = negative likelihood ratio; PLR = positive likelihood ratio.
In the population nationality analysis, the I2 of the Chinese population analysis was 77.15 for sensitivity, which indicates considerable heterogeneity in this population;
Studies that used DWI (not limited to DWI) to predict complete response to nCRT had lower heterogeneity in terms of sensitivity and specificity, although including DWI did not improve the AUC.
No significant heterogeneity in sensitivity and specificity was observed across studies that used traditional machine learning models;
No significant heterogeneity among studies with a small training cohort was observed, while studies with relatively large training cohorts showed marked heterogeneity.
Heterogeneity was lower across studies that used external validation than in those that used internal validation. Studies that used external validation did not show significant heterogeneity.
Discussion
NCRT is a standard treatment for locally advanced rectal cancer. However, responses to nCRT are highly variable among patients. Thus, effective prediction of treatment response to nCRT is essential for high-quality prognosis and personalized treatment to improve patient outcomes. In this study, we evaluated the effectiveness of machine learning models in predicting complete patient response to nCRT. Our study showed a high AUC of 0.91 (95% CI, 0.88-0.93), a pooled sensitivity of 0.83 (95% CI, 0.74-0.89), and a pooled specificity of 0.86 (95% CI, 0.80-0.91), when using machine-learning-based models to predict pCR to nCRT. These results were consistent with the findings in the study conducted by Jia et al11 which showed an AUC of 0.91 (95% CI, 0.88-0.93), a pooled sensitivity of 0.82 (95% CI, 0.71-0.90), and a pooled specificity of 0.86 (95% CI, 0.80-0.91). Compared to Jia et al’s11 study that included 21 studies using AI models to predict patient complete response to nCRT with MRI, our study demonstrates superiority from 3 aspects: (1) our included more (24) studies; (2) our study performed additional subgroup analysis. We performed an additional analysis on studies that used DWI and found that studies using DWI (not limited to DWI) to predict complete response to nCRT have lower heterogeneity in terms of sensitivity and specificity; (3) to identify key studies that caused heterogeneity, we conducted the meta-analysis 24 times using the leave-one-out strategy and identified 6 studies17,22,28,29,31,35 that mainly contributed to the heterogeneity.
Machine learning has been effectively applied to colorectal-related diseases such as colonic polyps, adenomas, CRC, ulcerative colitis, and intestinal motility disorders.39–42 In recent years, the rapid development of deep learning has further advanced colorectal diagnosis and treatment. Up to March 2024, 6 of our selected studies26–28,35,37,38 employed deep learning and these studies obtained a higher AUC (0.93) than studies using traditional machine learning models (0.89). However, heterogeneity test indicated higher heterogeneity among studies that used deep-learning-based models in terms of sensitivity and specificity. After finding significant heterogeneity in the pooled sensitivity and pooled specificity, we conducted the meta-analysis 24 times using the leave-one-out strategy and identified 6 studies that mainly contributed to the heterogeneity. Two studies28–35 used deep-learning-based models, while 4 studies used radiomics. Such results confirmed that although deep-learning-based models are able to obtain higher AUC, they are subject to heterogeneity. This may be due to variations in different convolutional neural network models, varying computation resources, lack of generalizations, as well as varying training strategies.
In addition, DWI has been widely applied to assess tumour response to neoadjuvant treatment. Thus, our study further investigated the role DWI plays in rectal cancer treatment prediction. We found that studies that used DWI (not limited to DWI) had lower heterogeneity in terms of sensitivity and specificity; however, the AUC was relatively lower compared to T2-MRI. Although promising results have also been reported when using DWI for tumour response prediction and tumour prognostication, DWI analysis protocols require standardization and lack of standardization may lead to different outcomes.43
As for the limitations of the study, first, while our analysis showed that deep-learning-based models are able to produce outcomes with higher AUC, the sample size of deep-learning-based models is relatively small. As more deep-learning-based studies are getting published, future analysis with more deep-learning-based algorithms is desirable. Second, we only analysed studies that used MRI, even though other modalities such as CT, endoscopy, as well as clinical features are also commonly used in identifying cPR to nCRT. Third, over 50% of the studies are conducted in China and 87.5% of the studies are retrospective study, which may lead to regional bias and selection bias. Since negative results are more challenging to get published, it may further add to publication bias. Lastly, noticeable heterogeneities are present among the selected studies, which means the results need to be interpreted more cautiously. Heterogeneity limits the ability translating machine learning results into clinical practice and it might be due to the lack of standardization in the imaging diagnosis protocols since scanners from different manufactures and different parameter settings when acquiring the images. Machine learning models also vary from image preprocessing to training strategies, which makes generalizing the whole diagnosis process challenging. Furthermore, different hospitals have different treatment regimes, which may lead to various patients outcomes. Thus, standardizing imaging, diagnosis, and treatment protocol in the future is essential to improve robustness and clinical applicability in machine learning models. With the fast development of large pretrained models such as the Segment Anything model,44 we believe that standardized protocol will enable unlimited possibilities in precision medicine.
Conclusion
This systematic study showed that machine learning has promising potential in predicting tumour response to nCRT in patients with locally advanced rectal cancer. Compared to traditional machine learning models, deep-learning-based studies are able to obtain higher AUC, although they are less predominant and more heterogeneous. Together with clinical information, machine learning-based models may bring us closer towards precision medicine.
Supplementary Material
Appendix
| Section and topic | Item # | Checklist item | Location where item is reported |
|---|---|---|---|
| Title | |||
| Title | 1 | Identify the report as a systematic review. | Line1-3 |
| Abstract | |||
| Abstract | 2 | See the PRISMA 2020 for Abstracts checklist. | Line5-40 |
| Introduction | |||
| Rationale | 3 | Describe the rationale for the review in the context of existing knowledge. | Line41-66 |
| Objectives | 4 | Provide an explicit statement of the objective(s) or question(s) the review addresses. | Line67-72 |
| Methods | |||
| Eligibility criteria | 5 | Specify the inclusion and exclusion criteria for the review and how studies were grouped for the syntheses. | Line74-95 |
| Information sources | 6 | Specify all databases, registers, websites, organisations, reference lists and other sources searched or consulted to identify studies. Specify the date when each source was last searched or consulted. | Line75-76 |
| Search strategy | 7 | Present the full search strategies for all databases, registers and websites, including any filters and limits used. | Line75-82 |
| Selection process | 8 | Specify the methods used to decide whether a study met the inclusion criteria of the review, including how many reviewers screened each record and each report retrieved, whether they worked independently, and if applicable, details of automation tools used in the process. | Line92-95 |
| Data collection process | 9 | Specify the methods used to collect data from reports, including how many reviewers collected data from each report, whether they worked independently, any processes for obtaining or confirming data from study investigators, and if applicable, details of automation tools used in the process. | Line92-95 |
| Data items | 10a | List and define all outcomes for which data were sought. Specify whether all results that were compatible with each outcome domain in each study were sought (eg, for all measures, time points, analyses), and if not, the methods used to decide which results to collect. | Line112-123 |
| 10b | List and define all other variables for which data were sought (eg, participant and intervention characteristics, funding sources). Describe any assumptions made about any missing or unclear information. | Line112-123 | |
| Study risk of bias assessment | 11 | Specify the methods used to assess risk of bias in the included studies, including details of the tool(s) used, how many reviewers assessed each study and whether they worked independently, and if applicable, details of automation tools used in the process. | Line115-167 |
| Effect measures | 12 | Specify for each outcome the effect measure(s) (eg, risk ratio, mean difference) used in the synthesis or presentation of results. | Line112-123 |
| Synthesis methods | 13a | Describe the processes used to decide which studies were eligible for each synthesis (eg, tabulating the study intervention characteristics and comparing against the planned groups for each synthesis (item #5)). | Line74-95 |
| 13b | Describe any methods required to prepare the data for presentation or synthesis, such as handling of missing summary statistics, or data conversions. | Line112-123 | |
| 13c | Describe any methods used to tabulate or visually display results of individual studies and syntheses. | Line106-110 | |
| 13d | Describe any methods used to synthesize results and provide a rationale for the choice(s). If meta-analysis was performed, describe the model(s), method(s) to identify the presence and extent of statistical heterogeneity, and software package(s) used. | Line112-123 | |
| 13e | Describe any methods used to explore possible causes of heterogeneity among study results (eg, subgroup analysis, meta-regression). | Line116-117 | |
| 13f | Describe any sensitivity analyses conducted to assess robustness of the synthesized results. | Line104 | |
| Reporting bias assessment | 14 | Describe any methods used to assess risk of bias due to missing results in a synthesis (arising from reporting biases). | Line112-123 |
| Certainty assessment | 15 | Describe any methods used to assess certainty (or confidence) in the body of evidence for an outcome. | Line112-123 |
| Results | |||
| Study selection | 16a | Describe the results of the search and selection process, from the number of records identified in the search to the number of studies included in the review, ideally using a flow diagram. | Line126-131 |
| 16b | Cite studies that might appear to meet the inclusion criteria, but which were excluded, and explain why they were excluded. | Line127-129 | |
| Study characteristics | 17 | Cite each included study and present its characteristics. | Line133-153 |
| Risk of bias in studies | 18 | Present assessments of risk of bias for each included study. | Line155-167 |
| Results of individual studies | 19 | For all outcomes, present, for each study: (a) summary statistics for each group (where appropriate) and (b) an effect estimate and its precision (eg, confidence/credible interval), ideally using structured tables or plots. | Line170-177 |
| Results of syntheses | 20a | For each synthesis, briefly summarise the characteristics and risk of bias among contributing studies. | Line155-167 |
| 20b | Present results of all statistical syntheses conducted. If meta-analysis was done, present for each the summary estimate and its precision (eg, confidence/credible interval) and measures of statistical heterogeneity. If comparing groups, describe the direction of the effect. | Line170-194 | |
| 20c | Present results of all investigations of possible causes of heterogeneity among study results. | Line178-184 | |
| 20d | Present results of all sensitivity analyses conducted to assess the robustness of the synthesized results. | Line182-184 | |
| Reporting biases | 21 | Present assessments of risk of bias due to missing results (arising from reporting biases) for each synthesis assessed. | Line187-190 |
| Certainty of evidence | 22 | Present assessments of certainty (or confidence) in the body of evidence for each outcome assessed. | Line196-215 |
| Discussion | |||
| Discussion | 23a | Provide a general interpretation of the results in the context of other evidence. | Line217-237 |
| 23b | Discuss any limitations of the evidence included in the review. | Line263-285 | |
| 23c | Discuss any limitations of the review processes used. | Line263-285 | |
| 23d | Discuss implications of the results for practice, policy, and future research. | Line238-262 | |
| Other information | |||
| Registration and protocol | 24a | Provide registration information for the review, including register name and registration number, or state that the review was not registered. | The review was not registered |
| 24b | Indicate where the review protocol can be accessed, or state that a protocol was not prepared. | A protocol was not prepared | |
| 24c | Describe and explain any amendments to information provided at registration or in the protocol. | None | |
| Support | 25 | Describe sources of financial or non-financial support for the review, and the role of the funders or sponsors in the review. | Title page |
| Competing interests | 26 | Declare any competing interests of review authors. | None |
| Availability of data, code, and other materials | 27 | Report which of the following are publicly available and where they can be found: template data collection forms; data extracted from included studies; data used for all analyses; analytic code; any other materials used in the review. | None |
From: Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
Contributor Information
Jia He, Department of Radiology, The First Affiliated Hospital of Hunan Normal University, Hunan Provincial People’s Hospital, Changsha 410002, China.
Shang-xian Wang, Bayer Radiology, Chengdu 610000, China.
Peng Liu, Department of Radiology, The First Affiliated Hospital of Hunan Normal University, Hunan Provincial People’s Hospital, Changsha 410002, China.
Author contributions
Jia He and Shang-xian Wang: substantial contributions to the conception and design of the study, acquisition of data, and analysis and interpretation of data. Drafting the article critically for important intellectual content. Peng Liu: Final approval of the version to be published. Revising it critically for important intellectual content. Agreement to be accountable for all aspects of the work in ensuring the integrity of any part of the work is appropriately investigated and resolved.
Supplementary material
Supplementary material is available at BJR online.
Funding
This work was supported by the Changsha Municipal Natural Science Foundation [grant number: kq2014201].
Conflicts of interest
The authors report no competing interests.
Data availability
All data generated during this study are included in this article and supplementary material.
Statement
In this manuscript, the guidelines of the PRISMA 2020 Statement have been adopted.
References
- 1. Rawla P, Sunkara T, Barsouk A. Epidemiology of colorectal cancer: incidence, mortality, survival, and risk factors. Prz Gastroenterol. 2019;14(2):89-103. 10.5114/pg.2018.81072 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Siegel RL, Miller KD, Fuchs HE, Jemal A. Cancer statistics, 2022. CA Cancer J Clin. 2022;72:7-33. [DOI] [PubMed] [Google Scholar]
- 3. Benson AB, Venook AP, Al-Hawary MM, et al. Rectal cancer, version 2.2018, NCCN clinical practice guidelines in oncology. J Natl Compr Canc Netw. 2018;16(7):874-901. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Feeney G, Sehgal R, Sheehan M, et al. Neoadjuvant radiotherapy for rectal cancer management. World J Gastroenterol. 2019;25(33):4850-4869. 10.3748/wjg.v25.i33.4850 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Li Y, Wang J, Ma X, et al. A review of neoadjuvant chemoradiotherapy for locally advanced rectal cancer. Int J Biol Sci. 2016;12(8):1022-1031. 10.7150/ijbs.15438 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Dossa F, Chesney TR, Acuna SA, Baxter NN. A watch-and-wait approach for locally advanced rectal cancer after a clinical complete response following neoadjuvant chemoradiation: a systematic review and meta-analysis. Lancet Gastroenterol Hepatol. 2017;2(7):501-513. [DOI] [PubMed] [Google Scholar]
- 7. Maas M, Nelemans PJ, Valentini V, et al. Long-term outcome in patients with a pathological complete response after chemoradiation for rectal cancer: a pooled analysis of individual patient data. Lancet Oncol. 2010;11(9):835-844. 10.1016/S1470-2045(10)70172-8 [DOI] [PubMed] [Google Scholar]
- 8. Li M, Xiao Q, Venkatachalam N, et al. Predicting response to neoadjuvant chemoradiotherapy in rectal cancer: from biomarkers to tumor models. Ther Adv Med Oncol. 2022;14:17588359221077972. 10.1177/17588359221077972 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. de Wilt JH, Vermaas M, Ferenschild FT, Verhoef C. Management of locally advanced primary and recurrent rectal cancer. Clin Colon Rectal Surg. 2007;20(3):255-263. 10.1055/s-2007-984870 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Gillies RJ, Kinahan PE, Hricak H. Radiomics: images are more than pictures, they are data. Radiology. 2016;278(2):563-577. 10.1148/radiol.2015151169 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Jia L-L, Zheng Q-Y, Tian J-H, et al. Artificial intelligence with magnetic resonance imaging for prediction of pathological complete response to neoadjuvant chemoradiotherapy in rectal cancer: A systematic review and meta-analysis. Front Oncol. 2022;12:1026216. 10.3389/fonc.2022.1026216 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Whiting PF, Rutjes AW, Westwood ME, et al. ; QUADAS-2 Group. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;55(8):529-536. 10.7326/0003-4819-155-8-201110180-00009 [DOI] [PubMed] [Google Scholar]
- 13. Yushkevich PA, Piven J, Hazlett HC, et al. User-guided 3D active contour segmentation of anatomical structures: significantly improved efficiency and reliability. Neuroimage. 2006;31(3):1116-1128. [DOI] [PubMed] [Google Scholar]
- 14. Kikinis R, Pieper SD, Vosburgh K. 3D Slicer: a platform for subject-specific image analysis, visualization, and clinical support. In: Jolesz FA, ed. Intraoperative Imaging Image-Guided Therapy, Vol 3; 2014:277-289. ISBN: 978-1-4614-7656-6 [Google Scholar]
- 15. Van Griethuysen JJM, Fedorov A, Parmar C, et al. Computational radiomics system to decode the radiographic phenotype. Cancer Res. 2017;77(21):e104-e107. 10.1158/0008-5472.CAN-17-0339 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Antunes JT, Ofshteyn A, Bera K, et al. Radiomic features of primary rectal cancers on baseline T(2)-weighted MRI are associated with pathologic complete response to neoadjuvant chemoradiation: a multisite study. J Magn Reson Imaging. 2020;52(5):1531-1541. 10.1002/jmri.27140 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Defeudis A, Mazzetti S, Panic J, et al. MRI-based radiomics to predict response in locally advanced rectal cancer: comparison of manual and automatic segmentation on external validation in a multicentre study. Eur Radiol Exp. 2022;6(1):19. 10.1186/s41747-022-00272-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Boldrini L, Lenkowicz J, Orlandini LC, et al. Applicability of a pathological complete response magnetic resonance-based radiomics model for locally advanced rectal cancer in intercontinental cohort. Radiat Oncol (London England). 2022;17(1):78. 10.1186/s13014-022-02048-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Bulens P, Couwenberg A, Intven M, et al. Predicting the tumor response to chemoradiotherapy for rectal cancer: model development and external validation using MRI radiomics. Radiother Oncol. 2020;142:246-252. 10.1016/j.radonc.2019.07.033 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Cheng Y, Luo Y, Hu Y, et al. Multiparametric MRI-based radiomics approaches on predicting response to neoadjuvant chemoradiotherapy (Ncrt) in patients with rectal cancer. Abdominal Radiol (New York). 2021;46(11):5072-5085. 10.1007/s00261-021-03219-0 [DOI] [PubMed] [Google Scholar]
- 21. Cui Y, Liu H, Ren J, et al. Development and validation of a MRI-based radiomics signature for prediction of kras mutation in rectal cancer. Eur Radiol. 2020;30(4):1948-1958. 10.1007/s00330-019-06572-3 [DOI] [PubMed] [Google Scholar]
- 22. Feng L, Liu Z, Li C, et al. Development and validation of a radiopathomics model to predict pathological complete response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer: a multicentre observational study. Lancet Digital Health. 2022;4(1):e8-e17. 10.1016/s2589-7500(21)00215-6 [DOI] [PubMed] [Google Scholar]
- 23. Ferrari R, Mancini-Terracciano C, Voena C, et al. MR-based artificial intelligence model to assess response to therapy in locally advanced rectal cancer. Eur J Radiol. 2019;118:1-9. 10.1016/j.ejrad.2019.06.013 [DOI] [PubMed] [Google Scholar]
- 24. Horvat N, Veeraraghavan H, Nahas CSR, et al. Combined artificial intelligence and radiologist model for predicting rectal cancer treatment response from magnetic resonance imaging: an external validation study. Abdominal Radiol (New York). 2022;47(8):2770-2782. 10.1007/s00261-022-03572-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Horvat N, Veeraraghavan H, Khan M, et al. MR imaging of rectal cancer: radiomics analysis to assess treatment response after neoadjuvant therapy. Radiology. 2018;287(3):833-843. 10.1148/radiol.2018172300 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Huang CM, Huang MY, Huang CW, et al. Machine learning for predicting pathological complete response in patients with locally advanced rectal cancer after neoadjuvant chemoradiotherapy. Sci Rep. 2020;10(1):12555. 10.1038/s41598-020-69345-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Jang B-S, Lim YJ, Song C, et al. Image-based deep learning model for predicting pathological response in rectal cancer using post-chemoradiotherapy magnetic resonance imaging. Radiother Oncol. 2021;161:183-190. 10.1016/j.radonc.2021.06.019 [DOI] [PubMed] [Google Scholar]
- 28. Jin C, Yu H, Ke J, et al. Predicting treatment response from longitudinal images using multi-task deep learning. Nat Commun. 2021;12(1):1851. 10.1038/s41467-021-22188-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Nardone V, Reginelli A, Grassi R, et al. Ability of delta radiomics to predict a complete pathological response in patients with loco-regional rectal cancer addressed to neoadjuvant chemo-radiation and surgery. Cancers (Basel). 2022;14(12):3004. 10.3390/cancers14123004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Ouyang G, Yang X, Deng X, et al. Predicting response to total neoadjuvant treatment (TNT) in locally advanced rectal cancer based on multiparametric magnetic resonance imaging: a retrospective study. Cancer Manag Res. 2021;13:5657-5669. 10.2147/CMAR.S311501 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Wei Q, Chen Z, Tang Y, et al. External validation and comparison of MR-based radiomics models for predicting pathological complete response in locally advanced rectal cancer: a two-centre, multi-vendor study. Eur Radiol. 2023;33(3):1906-1917. 10.1007/s00330-022-09204-5 [DOI] [PubMed] [Google Scholar]
- 32. Rengo M, Landolfi F, Picchia S, et al. Rectal cancer response to neoadjuvant chemoradiotherapy evaluated with MRI: development and validation of a classification algorithm. Eur J Radiol. 2022;147:110146. 10.1016/j.ejrad.2021.110146 [DOI] [PubMed] [Google Scholar]
- 33. Shin J, Seo N, Baek S-E, et al. Mri radiomics model predicts pathologic complete response of rectal cancer following chemoradiotherapy. Radiology. 2022;303(2):351-358. 10.1148/radiol.211986 [DOI] [PubMed] [Google Scholar]
- 34. Wen L, Liu J, Hu P, et al. MRI-based radiomic models outperform radiologists in predicting pathological complete response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer. Acad Radiol. 2023;30(Suppl 1):S176-S184. 10.1016/j.acra.2022.12.037 [DOI] [PubMed] [Google Scholar]
- 35. Yardimci AH, Kocak B, Sel I, et al. Radiomics of locally advanced rectal cancer: machine learning-based prediction of response to neoadjuvant chemoradiotherapy using pre-treatment sagittal T2-weighted MRI. Jpn J Radiol. 2023;Jan41(1):71-82. 10.1007/s11604-022-01325-7 [DOI] [PubMed] [Google Scholar]
- 36. Yi X, Pei Q, Zhang Y, et al. Mri-based radiomics predicts tumor response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer. Front Oncol. 2019;9:552. 10.3389/fonc.2019.00552 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Zhang X-Y, Wang L, Zhu H-T, et al. Predicting rectal cancer response to neoadjuvant chemoradiotherapy using deep learning of diffusion kurtosis mri. Radiology. 2020;296(1):56-64. 10.1148/radiol.2020190936 [DOI] [PubMed] [Google Scholar]
- 38. Zhu HT, Zhang XY, Shi YJ, Li XT, Sun YS. The conversion of MRI data with multiple b-values into signature-like pictures to predict treatment response for rectal cancer. J Magn Reson Imaging. 2022;56(2):562-569. 10.1002/jmri.28033 [DOI] [PubMed] [Google Scholar]
- 39. Luo Y, Zhang Y, Liu M, et al. Artificial intelligence-assisted colonoscopy for detection of colon polyps: a prospective, randomized cohort study. J Gastrointest Surg. 2020;25(8):2011-2018. 10.1007/s11605-020-04802-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Kudo S-E, Ichimasa K, Villard B, et al. Artificial intelligence system to determine risk of T1 colorectal cancer metastasis to lymph node. Gastroenterology. 2021;160(4):1075-1084 e1072. 10.1053/j.gastro.2020.09.027 [DOI] [PubMed] [Google Scholar]
- 41. Gubatan J, Levitte S, Patel A, Balabanis T, Wei MT, Sinha SR. Artificial intelligence applications in inflammatory bowel disease: emerging technologies and future directions. World J Gastroenterol. 2021;27(17):1920-1935. 10.3748/wjg.v27.i17.1920 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Bedrikovetski S, Dudi-Venkata NN, Kroon HM, et al. Artificial intelligence for pre-operative lymph node staging in colorectal cancer: a systematic review and meta-analysis. BMC Cancer. 2021;21(1):1058. 10.1186/s12885-021-08773-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Schurink NW, Lambregts DMJ, Beets-Tan RGH. Diffusion-weighted imaging in rectal cancer: current applications and future perspectives. Br J Radiol. 2019;92(1096):20180655. 10.1259/bjr.20180655 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Kirillov A, Mintun E, Ravi N, et al. 2023. Segment anything, arXiv, abs/2304.02643, preprint: not peer reviewed.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data generated during this study are included in this article and supplementary material.




