ABSTRACT
Background
Pathological complete response (pCR) following neoadjuvant chemoradiotherapy (nCRT) in locally advanced rectal cancer (LARC) is a key prognostic marker with implications for response‐adapted management. Although magnetic resonance imaging (MRI) is central to response assessment, differentiating residual tumour from treatment‐related changes remains challenging. Artificial intelligence (AI) and machine learning (ML) models applied to MRI show promise in predicting pCR; however, variability in methodology and performance limits clinical translation.
Methods
A search of Embase, Medline, Cochrane and Web of Science was conducted in April 2025 in accordance with Preferred Reporting Items for Reviews and Meta‐Analysis (PRISMA) guidelines. Eligible studies used MRI‐only AI or ML models to predict pCR following chemoradiotherapy in adults with rectal cancer. Screening, full‐text review and data extraction were performed independently by two reviewers. Risk of bias was assessed using Quality Assessment of Diagnostic Accuracy Studies (QUADAS‐2).
Results
Twenty‐two studies comprising 94 predictive models were included. Most studies were retrospective, used T2‐weighted MRI and demonstrated variability in MRI protocols, modelling methods and validation strategies. Only five studies conducted external validation. Median AUC was 0.801, with performance ranging from poor to excellent (AUC 0.49–0.997).
Conclusion
MRI‐based AI models demonstrate moderate discriminative performance for predicting pCR following neoadjuvant therapy in LARC. However, methodological heterogeneity, inconsistent reporting and limited external validation currently hinder generalisability. Greater methodological standardisation and multicentre external validation are required before clinical implementation.
Keywords: artificial intelligence, colorectal neoplasms, machine learning, MRI, neoadjuvant therapy
1. Introduction
The standard treatment for locally advanced rectal cancer (LARC) (Stages II–III) is neoadjuvant therapy, either conventional neoadjuvant chemoradiotherapy (nCRT) or total neoadjuvant therapy (TNT), followed by total mesorectal excision (TME) [1]. Following chemoradiotherapy, approximately 15%–30% of patients achieve pCR, which is associated with better oncological outcomes, including improved disease‐free survival and lower recurrence rates [2]. However, confirmation of pCR requires surgery, and TME carries a significant risk of morbidity, including permanent stoma formation and low anterior resection syndrome (LARS).
Contrastingly, patients with a clinical complete response (cCR), defined by the absence of residual tumour on digital rectal exam, endoscopy and magnetic resonance imaging (MRI), may undergo a non‐operative “watch‐and‐wait” strategy and avoid surgery [3]. However, cCR remains an imperfect surrogate for pCR. Some patients with apparent cCR have residual microscopic disease, while others without cCR may still achieve pCR.
This diagnostic uncertainty has driven interest in imaging‐based methods to predict pathological response prior to surgery. Although MRI is central to evaluating post‐treatment response, differentiating residual tumour from post‐treatment fibrosis or inflammation remains challenging. Artificial intelligence (AI) and machine learning (ML) techniques have emerged as promising tools to enhance MRI interpretation and quantitatively predict pCR, supporting treatment stratification and organ‐preserving strategies.
ML is a subset of AI that uses algorithms to learn patterns from complex datasets [4]. Radiomics is a quantitative imaging technique involving the extraction of high‐dimensional features from medical images. These features include tumour shape, intensity, texture and spatial heterogeneity and can be used to train ML models [4, 5] to predict pCR. By capturing quantitative imaging characteristics undiscernible to the human eye, radiomics may reflect underlying tumour biology and improve the prediction of treatment response [5]. However, the clinical readiness of AI‐based prediction models remains uncertain.
This systematic review aims to evaluate MRI‐only AI and ML models for predicting pCR in response to neoadjuvant therapy in LARC, focusing on discriminative performance and methodological quality, while summarising patient characteristics, modelling approaches and key limitations to guide further development and translation to clinical use.
2. Methods
2.1. Search Strategy
The literature search was conducted in April 2025 using Embase, Medline, Cochrane Library and Web of Science. Searches were limited to English‐language studies, with no date restrictions. Search strategies used Boolean operators to combine variations of the key terms ‘rectal cancer’, ‘neoadjuvant therapy’, ‘machine learning’ and ‘pathological complete response’. Full search strategies are provided in Tables S1–S4.
2.2. Inclusion and Exclusion Criteria
This review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta‐Analysis (PRISMA) guidelines. Eligible studies included adult patients (≥ 18 years) with LARC undergoing nCRT or TNT, with subsequent surgical resection and histopathological assessment of pCR. Studies involving short‐course radiotherapy were excluded. Included studies developed AI/ML models using MRI‐derived imaging features to predict pCR. Models incorporating clinical, biochemical or multimodal imaging variables were excluded to isolate the predictive contribution of MRI alone. The primary outcome was discriminative performance, reported as the area under the receiver operating characteristic curve (AUC). Secondary outcomes included methodological characteristics of studies, including modelling approach, feature engineering, neoadjuvant therapy regimen, validation strategy and risk of bias. Detailed inclusion and exclusion criteria are outlined in Table 1.
TABLE 1.
Inclusion and exclusion criteria.
| Inclusion criteria | Exclusion criteria |
|---|---|
|
|
2.3. Selection Process
Studies were imported into Covidence, and duplicates were removed. Title/abstract and full‐text review were performed independently by W.L. and H.N., with disagreements resolved by consensus or deferral to the senior author, J.Y.
2.4. Data Extraction
Data extraction was performed independently using a customised Covidence template by W.L. and E.E. Disagreements were resolved by consensus discussion. Data extracted included the categories of study characteristics, nCRT regimens, MRI protocols, feature engineering, modelling, validation methods and diagnostic performance.
3. Results
3.1. Study Selection and Characteristics of Included Studies
Our research identified 1046 studies, with 84 studies included for full‐text review and 22 included in the study cohort. The literature search and review process is outlined in Figure 1.
FIGURE 1.

PRISMA flow diagram.
Table 2 provides a summary of the characteristics of eligible studies published between 2019 and 2025, predominantly originating from China (11 studies). Most studies were retrospective (17 studies). The median number of patients in each paper was 156 (range: 40–912). The median pCR rate was 20.8% (range: 16%–73.3%).
TABLE 2.
Study characteristics.
| First author and year | Country | Number of centres | Study design | pCR definition | Total number of patients | Training set | pCR (%) training set | Validation set/test set | pCR (%) validation set | pCR (%) total | Validation method |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Antunes (2020) [6] | United States of America | 3 | Retrospective | AJCC | 104 | NR | NR |
T + IV 60 EV 44 |
T + V 13 (21.67%) EV 10 (22.73%) |
23 (22.1%) | IV—50 runs of threefold CV, EV |
| Bellini (2022) [7] | Italy | 1 | Prospective | Dworak | 40 | 40 | 13 (32.5%) | NR | NR | 13 (32.5%) | No validation |
| Bulens (2020) [8] | Belgium | 2 | Prospective | AJCC | 125 | 70 | 12 (17.1%) | EV 55 | NR | 20 (16%) | IV fivefold CV, EV |
| Crimi (2024) [9] | Italy | 1 | Prospective | NR | 102 | 72 | 20 (27.8%) | IV 30 | 4 (13.3%) | 24 (23.5%) | IV—nested 10‐fold CV with split sample |
| (Ferrari) 2019 [10] | Italy | 1 | Retrospective | Dworak | 55 | 28 | NR | IV 27 | NR | 16 (29.1%) | IV—split sample |
| Horvat (2018) [11] | Brazil | 1 | Retrospective | AJCC | 114 | NR | NR | NR | NR | 21 (18.4%) | IV—fivefold CV |
| Jang (2021) [12] | South Korea | 1 | Retrospective | Dworak | 466 | NR | NR |
IV 113 Test 107 |
20 (17%) | 60 (17%) | IV—split sample |
| Jayaprakasam (2022) [13] | USA | 1 | Retrospective | AJCC | 236 | NR | NR | NR | NR | 41 (17.4%) | IV—fivefold CV |
| Lee (2021) [14] | South Korea | 1 | Retrospective | NR | 912 | 592 | 114 (19.3%) | Test 320 | 78 (24.4%) | 192 (21.1%) | IV—split sample |
| Li (2019) [15] | China | 1 | Retrospective | AJCC | 131 | 87 | 18 (20.69%) | IV 44 | 9 (20.45%) | 27 (20.6%) | IV—split sample |
| Li (2021) [16] | China | 1 | Retrospective | AJCC | 80 | NR | NR | NR | NR | 15 (18.8%) | IV—leave one out CV |
| Ma (2024) [17] | China | 1 | Retrospective | NCCN | 120 | 84 | NR | IV 36 | NR | 38 (31.7%) | IV—split sample 70:30 |
| Miranda (2023) [18] | Brazil | 1 | Retrospective | AJCC | 180 | 126 | NR | IV 54 | NR | 33 (18%) | IV—10‐fold CV with 70:30 split |
| Nie (2016) [19] | China | 1 | Prospective | AJCC | 48 | 36 | NR | IV 12 | NR | 11 (23%) | IV—fourfold CV |
| Pang (2021) [20] | China | 2 | Retrospective | NCNN | 187 | 107 | 36 (33.64%) |
IV 46 EV 34 |
IV 8 (17.4%) EV 6 (17.6%) |
50 (73.3%) | IV—fivefold CV, EV |
| Shi (2019) [21] | China | 1 | Retrospective | AJCC | 51 | NR | NR | NR | NR | 10 (19.6%) | IV—fourfold CV for ANN, 10‐fold CV for CNN |
| Shi (2023) [22] | China | 3 |
Retrospective (training) Prospective (testing) |
AJCC | 224 | 147 | 41 (27.9%) | EV 77 | 16 (20.8%) | 57 (25.5%) | IV—fivefold CV, EV |
| Shin (2022) [23] | Republic of Korea | 1 | Retrospective | Mandard | 898 | 592 | NR | IV 306 | NR | 189 (21%) | IV—10‐fold CV with chronological split |
| Wu (2025) [24] | China | 3 | Retrospective | AJCC/CAP | 768 | 401 | 127 (31.6%) | EV 367 | EV 104 (28.3%) | 231 (30%) | IV—10‐fold CV, EV |
| Xia (2025) [25] | China | 1 | Retrospective | AJCC | 373 | 264 | 50 (18.9%) | IV 144 | IV 22 (15.3%) | 72 (19.3%) | IV—split sample |
| Zhang (2020) [26] | China | 1 | Prospective | NCNN | 383 | 290 | 56 (19.3%) | IV 93 | 18 (19.4%) | 74 (19.3%) | IV |
| Zhu (2022) [27] | China | 1 | Retrospective | NR | 472 | 200 | 42 (21%) |
IV 72 Test 200 |
IV 11 (15.3%) Test 44 (22%) |
97 (20.6%) | IV—chronological split |
Abbreviations: AJCC, American Joint Committee on Cancer; CAP, College of American Pathologists; CV, cross‐validation; EV, external validation; IV, internal validation; NCCN, National Comprehensive Cancer Network; NR, not reported.
3.2. MRI Protocols and Modelling Algorithms
Model construction parameters are summarised in Table S5.
All studies except one utilised T2‐weighted (T2W) imaging. Studies also employed diffusion‐weighted imaging (DWI), T1‐weighted imaging (T1W), dynamic contrast‐enhanced imaging (DCE) and diffusion kurtosis imaging (DKI), in decreasing frequency. Most studies used a 3 T MRI field strength. MRI acquisition timing varied across studies, with some models using pre‐treatment imaging alone, post‐treatment imaging alone or pre‐ and post‐treatment MRI. For MRI segmentation software, most studies used ITK‐SNAP, then 3D‐Slicer and MIM Maestro.
Fifteen studies used ML algorithms alone, including support vector machines (SVM), random forest (RF), least absolute shrinkage and selection operator (LASSO) regression, k‐nearest neighbour (kNN), decision trees (DT), gradient boosting methods and logistic regression, in order of most to least common. Three studies used Deep Learning (DL) algorithms alone, using convolutional neural networks (CNN), whilst four studies used both ML and DL models.
3.3. Validation Methods
Only five studies conducted external validation. The remaining studies conducted internal validation, with k‐fold cross‐validation being the most common, including one study using nested cross‐validation and one using leave‐one‐out cross‐validation. Many studies utilised a split sample for validation, for example, 70:30 training and validation splits or temporal splits.
3.4. Diagnostic Performance
Across 22 studies, 94 models were identified. The mean and median AUC values were 0.778 and 0.801, respectively, with a range of 0.49–0.997 and an IQR of 0.71–0.85. Using commonly accepted AUC interpretation thresholds [28], 13 models were considered excellent (> 0.9), 34 models good (0.8–0.9), 27 models fair (0.7–0.8) and 20 models poor or unacceptable (< 0.7).
3.5. Quality Assessment
Figure 2 shows the QUADAS‐2 assessments. Index test risk of bias had the greatest concern, given that only one paper had prespecified thresholds [9]. The remaining papers either specified thresholds being post hoc or were unclear regarding this. Concerns regarding patient selection bias were present, as four studies either did not state or had unclear inclusion/exclusion criteria. In addition, two studies could have avoided inappropriate exclusions [11, 22].
FIGURE 2.

QUADAS‐2 risk of bias assessment traffic light plot and bar plot.
Concerns regarding the reference standard were generally low, given that all papers used surgical histopathology, the gold standard for determining pCR. Concerns in flow and timing were common, as seven papers were unclear on the interval between nCRT and TME. Applicability concerns were low as the inclusion criteria were already highly specific.
4. Discussion
This review found a median AUC of 0.801, which, although it is considered ‘good’ by conventional thresholds [28], needs to be assessed in context.
4.1. Strengths and Limitations of Selected Studies
External validation is widely regarded as a prerequisite for demonstrating clinical usefulness and generalisability. Only five studies conducted external validation, while the remainder used internal validation. Regarding internal validation types, train‐validation split samples allow validation to be done on unseen data; however, this reduces the effective sample size and increases the risk of variance and instability [29]. Conversely, most studies used resampling techniques such as cross‐validation to maximise the use of limited datasets; however, these models still yield optimistically biased performance estimates [29]. Given the high‐stakes implications of pCR prediction, robust external validation is required for reliability and patient safety.
Our review identified substantial methodological heterogeneity, including differences in MRI protocols, tumour segmentation approaches, feature engineering methods, ML and DL models and validation strategies. This variability limits comparison and generalisability, making it unclear whether differences in reported performance reflect better model design or pipeline variation.
Inconsistent reporting of post‐treatment imaging timing is another limitation. As tumour regression and radiological appearance continue to evolve over time following nCRT [30], variation in imaging timing may influence radiomic feature extraction and model calibration. Consequently, models may learn timing‐related imaging characteristics rather than true biological predictors of response.
To maintain the consistency of the reference standard, only studies using surgical histopathology to show pCR were included, whilst studies including sustained cCR as a surrogate for pCR were excluded.
Heterogeneity in and inconsistent reporting of neoadjuvant treatment regimens were also observed. Reported treatment regimens included long‐course chemoradiotherapy and TNT, with limited reporting of induction or consolidation sequences. As tumour regression is time and treatment‐dependent [30], different neoadjuvant therapy regimens may result in different biological responses and radiological appearances. Most studies did not evaluate the impact of treatment regimen differences on model performance, suggesting that reported predictive accuracy may be regimen‐specific and not broadly generalisable, limiting comparison of models given heterogeneous treatment cohorts.
4.2. Strengths and Limitations of Methodology
This review focuses on MRI‐based AI models to the exclusion of other imaging modalities and clinical variables. Whilst this improves methodological robustness and comparability, this approach does not reflect real‐world practice, whereby MRI is interpreted alongside physical examination, endoscopy, CT and biochemical markers [31] and in the context of the specific treatment regimen used. While high‐resolution MRI remains the preferred modality for local staging and assessment of tumour relationships with the mesorectal fascia and circumferential resection margin [32], optimal rectal cancer management is reliant on multimodal assessment rather than a single imaging technique.
Furthermore, by excluding clinical variables, this review does not compare the performance of clinical‐only, imaging‐only and combined prediction models. Consequently, the incremental predictive value of MRI‐based AI beyond established clinical predictors of treatment response cannot be determined. Readily available clinical features such as carcinoembryonic antigen (CEA) levels and tumour stage can be easily incorporated into combined models and may substantially influence predictive performance. Direct comparison of clinical, imaging and combined approaches is therefore required to determine whether MRI‐based AI meaningfully improves discrimination and has clinical utility in guiding patient selection for organ preservation strategies.
Significant heterogeneity across studies precluded meta‐analysis. As a result, results were synthesised descriptively, limiting statistical analysis and the strength of evidence compared with meta‐analyses.
ML and DL differ primarily in feature extraction. ML approaches rely on predefined features derived from imaging data, whereas DL approaches use multi‐layered neural networks to automatically learn features directly from imaging data [5]. ML models may perform well in small datasets and offer greater interpretability but are dependent on feature selection and may be sensitive to variation in MRI protocols [33]. In contrast, DL models can capture complex spatial relationships and may be advantageous in distinguishing residual tumour from treatment‐related fibrosis. However, DL approaches require larger datasets and have a higher risk of overfitting in small, single‐centre cohorts [34], which characterises most studies in this review. Consequently, differences in model performance are likely influenced by dataset size, standardisation and validation strategy in addition to algorithm type. The heterogeneity of included studies precluded direct comparison between ML and DL. However, prior reviews have suggested that DL models may demonstrate improved performance [35, 36, 37].
4.3. Comparison With Prior Literature
Prior literature has reported higher AUCs (approximately 0.9–0.91) than our median of 0.801; however, these reviews frequently incorporated clinical features, other imaging modalities, or selectively reported high‐performing models. Jia et al. [35] and Jong et al. [38] demonstrated higher median AUCs whilst incorporating clinical features, whilst Liao et al. [37] and Shen et al. [39] reported higher performance by including clinical features and selection of the validation/test cohorts with ‘superior predictive performance’ or including other imaging modalities, respectively. Subgroup analysis by Jia et al. [35] found that incorporating clinical factors yielded higher diagnostic accuracy than with radiomics models alone. Studies incorporating clinical features or other imaging modalities may achieve higher AUCs and have greater predictive performance. However, inclusion of these factors limits the ability to isolate the predictive contribution of MRI and reduces comparability between MRI‐based modelling pipelines.
Shin et al. [23] demonstrated that AI models surpass radiologists' visual assessment, whilst Zhang et al. [26] demonstrated that AI models can significantly reduce radiologist error rates as an assistive tool. Whilst the predictive performance of imaging alone is modest at 0.801 in this review, it would likely still be a useful assistive tool for radiologists' visual assessment.
4.4. Implications of the Study/Future Directions
This review highlights the need for greater methodological standardisation and transparent reporting to improve reproducibility in MRI‐based radiomics research. Future studies should prioritise standardised radiomics feature extraction and reporting, adhering to established frameworks such as the Image Biomarker Standardisation Initiative (IBSI) and TRIPOD‐AI reporting guidelines [40, 41].
Most studies identified in this review remain at the retrospective model development stage. Future research should prioritise external validation using multicentre cohorts to ensure models are robust and generalisable across different populations and imaging protocols. Subsequent prospective studies are required, including studies to evaluate the clinical utility of these models and their integration into multidisciplinary decision‐making pathways, particularly in guiding patient selection for organ preservation strategies. Future studies should also evaluate whether MRI‐based AI models provide incremental predictive value beyond established clinical predictors.
5. Conclusion
This review demonstrates that MRI‐based AI models achieve moderate discriminative performance for predicting pCR in LARC. However, these findings should be interpreted cautiously, as performance estimates are likely to be optimistically biased due to reliance on small, retrospective cohorts with internal validation strategies. Methodological heterogeneity and inconsistent reporting further limit generalisability. Standardised methodology and robust multicentre external validation are essential for safe clinical translation to position MRI‐based AI models as an assistive component to multimodal assessment of LARC treatment response.
Funding
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Table S1: Ovid Medline search strategy.
Table S2: Embase search strategy.
Table S3: Cochrane search strategy.
Table S4: Web of Science search strategy.
Table S5: Modelling construction parameters.
Table S6: Model performance.
Table S7: Chemoradiotherapy regimens.
Acknowledgements
We thank librarian Evelyn Hutcheon for her support in developing the search strategy. Open access publishing facilitated by The University of Melbourne, as part of the Wiley ‐ The University of Melbourne agreement via the Council of Australasian University Librarians.
Data Availability Statement
The data that supports the findings of this study are available in the Supporting Information of this article.
References
- 1. Benson A. B., Venook A. P., Adam M., et al., “NCCN Guidelines® Insights: Rectal Cancer, Version 3.2024,” Journal of the National Comprehensive Cancer Network 22, no. 6 (2024): 366–375. [DOI] [PubMed] [Google Scholar]
- 2. Maas M., Nelemans P. J., Valentini V., et al., “Long‐Term Outcome in Patients With a Pathological Complete Response After Chemoradiation for Rectal Cancer: A Pooled Analysis of Individual Patient Data,” Lancet Oncology 11, no. 9 (2010): 835–844. [DOI] [PubMed] [Google Scholar]
- 3. Temmink S. J. D., Peeters K., Bahadoer R. R., et al., “Watch and Wait After Neoadjuvant Treatment in Rectal Cancer: Comparison of Outcomes in Patients With and Without a Complete Response at First Reassessment in the International Watch & Wait Database (IWWD),” British Journal of Surgery 110, no. 6 (2023): 676–684. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Wagner M. W., Namdar K., Biswas A., Monah S., Khalvati F., and Ertl‐Wagner B. B., “Radiomics, Machine Learning, and Artificial Intelligence—What the Neuroradiologist Needs to Know,” Neuroradiology 63, no. 12 (2021): 1957–1967. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Lambin P., Rios‐Velazquez E., Leijenaar R., et al., “Radiomics: Extracting More Information From Medical Images Using Advanced Feature Analysis,” European Journal of Cancer 48, no. 4 (2012): 441–446. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Antunes J. T., Ofshteyn A., Bera K., et al., “Radiomic Features of Primary Rectal Cancers on Baseline T2‐Weighted MRI Are Associated With Pathologic Complete Response to Neoadjuvant Chemoradiation: A Multisite Study,” Journal of Magnetic Resonance Imaging 52, no. 5 (2020): 1531–1541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Bellini D., Carbone I., Rengo M., et al., “Performance of Machine Learning and Texture Analysis for Predicting Response to Neoadjuvant Chemoradiotherapy in Locally Advanced Rectal Cancer With 3T MRI,” Tomography 8, no. 4 (2022): 2059–2072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Bulens P., Couwenberg A., Intven M., et al., “Predicting the Tumor Response to Chemoradiotherapy for Rectal Cancer: Model Development and External Validation Using MRI Radiomics,” Radiotherapy and Oncology 142 (2020): 246–252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Crimi F., D'Alessandro C., Zanon C., et al., “A Machine Learning Model Based on MRI Radiomics to Predict Response to Chemoradiation Among Patients With Rectal Cancer,” Life 14, no. 12 (2024): 1530. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Ferrari R., Mancini‐Terracciano C., Voena C., et al., “MR‐Based Artificial Intelligence Model to Assess Response to Therapy in Locally Advanced Rectal Cancer,” European Journal of Radiology 118 (2019): 1–9. [DOI] [PubMed] [Google Scholar]
- 11. Horvat N., Veeraraghavan H., Khan M., et al., “MR Imaging of Rectal Cancer: Radiomics Analysis to Assess Treatment Response After Neoadjuvant Therapy,” Radiology 287, no. 3 (2018): 833EP–843EP. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Jang B. S., Lim Y. J., Song C., et al., “Image‐Based Deep Learning Model for Predicting Pathological Response in Rectal Cancer Using Post‐Chemoradiotherapy Magnetic Resonance Imaging,” Radiotherapy and Oncology 161 (2021): 183–190. [DOI] [PubMed] [Google Scholar]
- 13. Jayaprakasam V. S., Paroder V., Gibbs P., et al., “MRI Radiomics Features of Mesorectal Fat Can Predict Response to Neoadjuvant Chemoradiation Therapy and Tumor Recurrence in Patients With Locally Advanced Rectal Cancer,” European Radiology 32, no. 2 (2022): 971–980. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Lee S., Lim J., Shin J., Kim S., and Hwang H., “Pathologic Complete Response Prediction After Neoadjuvant Chemoradiation Therapy for Rectal Cancer Using Radiomics and Deep Embedding Network of MRI,” Applied Sciences 11, no. 20 (2021): 9494. [Google Scholar]
- 15. Li Y., Liu W., Pei Q., et al., “Predicting Pathological Complete Response by Comparing MRI‐Based Radiomics Pre‐ and Postneoadjuvant Radiotherapy for Locally Advanced Rectal Cancer,” Cancer Medicine 8, no. 17 (2019): 7244–7252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Li Z., Ma X., Shen F., Lu H., Xia Y., and Lu J., “Evaluating Treatment Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer Using Various MRI‐Based Radiomics Models,” BMC Medical Imaging 21, no. 1 (2021): 30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Ma Q., Liu Z., Zhang J., et al., “Multi‐Task Reconstruction Network for Synthetic Diffusion Kurtosis Imaging: Predicting Neoadjuvant Chemoradiotherapy Response in Locally Advanced Rectal Cancer,” European Journal of Radiology 174 (2024): 111402. [DOI] [PubMed] [Google Scholar]
- 18. Miranda J., Horvat N., Assuncao A. N., et al., “MRI‐Based Radiomic Score Increased mrTRG Accuracy in Predicting Rectal Cancer Response to Neoadjuvant Therapy,” Abdominal Radiology 48, no. 6 (2023): 1911–1920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Nie K., Shi L., Chen Q., et al., “Rectal Cancer: Assessment of Neoadjuvant Chemoradiation Outcome Based on Radiomics of Multiparametric MRI,” Clinical Cancer Research 22, no. 21 (2016): 5256–5264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Pang X., Wang F., Zhang Q., et al., “A Pipeline for Predicting the Treatment Response of Neoadjuvant Chemoradiotherapy for Locally Advanced Rectal Cancer Using Single MRI Modality: Combining Deep Segmentation Network and Radiomics Analysis Based on “Suspicious Region”,” Frontiers in Oncology 11 (2021): 711747. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Shi L., Zhang Y., Nie K., et al., “Machine Learning for Prediction of Chemoradiation Therapy Response in Rectal Cancer Using Pre‐Treatment and Mid‐Radiation Multi‐Parametric MRI,” Magnetic Resonance Imaging 61 (2019): 33EP–40EP. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Shi L., Zhang Y., Hu J., et al., “Radiomics for the Prediction of Pathological Complete Response to Neoadjuvant Chemoradiation in Locally Advanced Rectal Cancer: A Prospective Observational Trial,” Bioengineering 10, no. 6 (2023): 634. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Shin J., Seo N., Baek S. E., et al., “MRI Radiomics Model Predicts Pathologic Complete Response of Rectal Cancer Following Chemoradiotherapy,” Radiology 303, no. 2 (2022): 351–358. [DOI] [PubMed] [Google Scholar]
- 24. Wu X., Wang J., Chen C., et al., “Integration of Deep Learning and Sub‐Regional Radiomics Improves the Prediction of Pathological Complete Response to Neoadjuvant Chemoradiotherapy in Locally Advanced Rectal Cancer Patients,” Academic Radiology 6 (2025): 3384–3396. [DOI] [PubMed] [Google Scholar]
- 25. Xia S. J., Wang Z. N., Wu J. Q., et al., “Rectal‐Radiosam: Large Model‐Assisted Multi‐Parametric MRI Pipeline for Predicting Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer via no Human Intervention,” Physics and Imaging in Radiation Oncology 35 (2025): 100797. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Zhang X. Y., Wang L., Zhu H. T., et al., “Predicting Rectal Cancer Response to Neoadjuvant Chemoradiotherapy Using Deep Learning of Diffusion Kurtosis MRI,” Radiology 296, no. 1 (2020): 56–64. [DOI] [PubMed] [Google Scholar]
- 27. Zhu H. T., Zhang X. Y., Shi Y. J., Li X. T., and Sun Y. S., “The Conversion of MRI Data With Multiple b‐Values Into Signature‐Like Pictures to Predict Treatment Response for Rectal Cancer,” Journal of Magnetic Resonance Imaging 56, no. 2 (2022): 562–569. [DOI] [PubMed] [Google Scholar]
- 28. Nahm F. S., “Receiver Operating Characteristic Curve: Overview and Practical Use for Clinicians,” Korean Journal of Anesthesiology 75, no. 1 (2022): 25–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Steyerberg E. W., Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating (Springer, 2009). [Google Scholar]
- 30. Seo N., Kim H., Cho M. S., and Lim J. S., “Response Assessment With MRI After Chemoradiotherapy in Rectal Cancer: Current Evidences,” Korean Journal of Radiology 20, no. 7 (2019): 1003–1018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Stefanou A. J., Dessureault S., Sanchez J., and Felder S., “Clinical Tools for Rectal Cancer Response Assessment Following Neoadjuvant Treatment in the Era of Organ Preservation,” Cancers 15, no. 23 (2023): 5535. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. MERCURY Study Group , “Diagnostic Accuracy of Preoperative Magnetic Resonance Imaging in Predicting Curative Resection of Rectal Cancer: Prospective Observational Study,” BMJ 333, no. 7572 (2006): 779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Gillies R. J., Kinahan P. E., and Hricak H., “Radiomics: Images Are More Than Pictures, They Are Data,” Radiology 278, no. 2 (2016): 563–577. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Litjens G., Kooi T., Bejnordi B. E., et al., “A Survey on Deep Learning in Medical Image Analysis,” Medical Image Analysis 42 (2017): 60–88. [DOI] [PubMed] [Google Scholar]
- 35. Jia L.‐L., Zheng Q.‐Y., Tian J.‐H., et al., “Artificial Intelligence With Magnetic Resonance Imaging for Prediction of Pathological Complete Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer: A Systematic Review and Meta‐Analysis,” Frontiers in Oncology 12 (2022): 1026216. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. He J., Wang S. X., and Liu P., “Machine Learning in Predicting Pathological Complete Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer Using MRI: A Systematic Review and Meta‐Analysis,” British Journal of Radiology 97, no. 1159 (2024): 1243–1254. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Liao Z., Luo D., Tang X., Huang F., and Zhang X., “MRI‐Based Radiomics for Predicting Pathological Complete Response After Neoadjuvant Chemoradiotherapy in Locally Advanced Rectal Cancer: A Systematic Review and Meta‐Analysis,” Frontiers in Oncology 15 (2025): 1550838. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Jong B.‐K., Yu Z.‐H., Hsu Y.‐J., Chiang S.‐F., You J.‐F., and Chern Y.‐J., “Deep Learning Algorithms for Predicting Pathological Complete Response in MRI of Rectal Cancer Patients Undergoing Neoadjuvant Chemoradiotherapy: A Systematic Review,” International Journal of Colorectal Disease 40, no. 1 (2025): 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Shen H., Jin Z., Chen Q., et al., “Image‐Based Artificial Intelligence for the Prediction of Pathological Complete Response to Neoadjuvant Chemoradiotherapy in Patients With Rectal Cancer: A Systematic Review and Meta‐Analysis,” La Radiologia Medica 129, no. 4 (2024): 598–614. [DOI] [PubMed] [Google Scholar]
- 40. Zwanenburg A., Vallières M., Abdalah M. A., et al., “The Image Biomarker Standardization Initiative: Standardized Quantitative Radiomics for High‐Throughput Image‐Based Phenotyping,” Radiology 295, no. 2 (2020): 328–338. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Collins G. S., Moons K. G. M., Dhiman P., et al., “TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods,” BMJ 385 (2024): e078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Table S1: Ovid Medline search strategy.
Table S2: Embase search strategy.
Table S3: Cochrane search strategy.
Table S4: Web of Science search strategy.
Table S5: Modelling construction parameters.
Table S6: Model performance.
Table S7: Chemoradiotherapy regimens.
Data Availability Statement
The data that supports the findings of this study are available in the Supporting Information of this article.
