Skip to main content
The Journal of International Medical Research logoLink to The Journal of International Medical Research
. 2026 Jul 3;54(7):03000605261460311. doi: 10.1177/03000605261460311

Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Jia Yang 1,#, Junyu Jiang 2,#, Jin Peng 1, Jie Li 1, Peng Mi 3, Guangwen Chen 1,
PMCID: PMC13332255  PMID: 42396618

Abstract

Objective

This systematic review and meta-analysis evaluates the diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in Chinese patients with lung adenocarcinoma.

Methods

Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines and prospectively registered in the International Prospective Register of Systematic Reviews (CRD420251273027), a systematic search of PubMed, Embase, Web of Science, the Cochrane Library, Scopus, China National Knowledge Infrastructure, Wanfang, VIP, and Chinese Biomedical Literature Database was conducted from inception to 31 October 2025. Two reviewers independently screened studies, extracted data, and assessed bias using the Quality Assessment of Diagnostic Accuracy Studies-2 tool. A bivariate random-effects model was used to synthesize the data. Subgroup analyses were conducted for three factors: (a) imaging modality (computed tomography vs. positron emission tomography-computed tomography); (b) algorithm type (deep learning vs. conventional machine learning); and (3) validation strategy (external vs. internal).

Results

Thirteen studies encompassing 6628 patients were included. The pooled sensitivity was 71% (95% confidence interval: 68–74), the pooled specificity was 81% (95% confidence interval: 78–84), and the summary area under the curve was 0.85 (95% confidence interval: 0.82–0.88). Deep learning models significantly outperformed conventional machine learning models (area under the curve: 0.871 vs. 0.798; P = 0.012). Computed tomography-based models yielded higher accuracy than positron emission tomography-computed tomography-based models (area under the curve: 0.879 vs. 0.828; P = 0.038). Models validated on independent external cohorts demonstrated superior performance compared with those relying solely on internal validation (area under the curve: 0.922 vs. 0.841; P = 0.006). Imaging modality was a significant source of heterogeneity (P < 0.05). No threshold effect or publication bias was detected.

Conclusion

Machine learning-based radiomics models exhibit promising diagnostic accuracy for the noninvasive prediction of epidermal growth factor receptor mutations in Chinese patients with lung adenocarcinoma. Computed tomography-based deep learning models subjected to independent external validation represent the current optimal approach. However, the retrospective nature and substantial heterogeneity of the included studies necessitate large-scale, prospective, multicenter trials with standardized workflows before clinical translation.

Keywords: Lung adenocarcinoma, epidermal growth factor receptor mutation, radiomics, machine learning, deep learning, diagnostic accuracy, meta-analysis

Introduction

Lung cancer remains the leading cause of cancer-related mortality worldwide. Non-small cell lung cancer (NSCLC) constitutes 80%–85% of all cases, with lung adenocarcinoma (LUAD) being the predominant histological subtype.1,2 In Asian populations, epidermal growth factor receptor (EGFR) mutations occur in nearly 50% of patients with LUAD, a prevalence markedly higher than the 15% observed in Western cohorts. 3 Tumors harboring these sensitizing mutations respond favorably to tyrosine kinase inhibitors (TKIs), which significantly prolong progression-free and overall survival.4,5

Currently, tissue biopsy coupled with molecular genetic testing remains the gold standard for determining EGFR status. 6 However, this invasive approach carries inherent risks of sampling error, procedural complications, and delays in obtaining results, fundamentally limiting its utility for real-time monitoring.

Radiomics offers a noninvasive alternative by extracting high-dimensional quantitative features from routine medical images. These features are used to develop predictive models based on machine learning (ML) algorithms, ranging from conventional methods (e.g. support vector machines and random forests) to deep learning (DL) architectures such as convolutional neural networks. 7 While numerous studies have explored computed tomography (CT)- or positron emission tomography-computed tomography (PET/CT)-based radiomics for predicting EGFR mutations in LUAD,8,9 most have been small-scale, single-center, retrospective studies. Substantial methodological variations in algorithm selection, feature processing, and validation strategies have resulted in considerable heterogeneity in performance, underscoring the need for a systematic synthesis of the existing evidence.

Synthesizing these data specifically for Chinese populations is not merely beneficial but imperative. The approximately 50% prevalence of sensitizing EGFR mutations in this population—more than three times that observed in Western cohorts—dramatically alters the pre-test probability. This baseline difference fundamentally changes the clinical utility and performance characteristics of any predictive model. Furthermore, emerging evidence suggests that EGFR mutation subtypes and co-mutation patterns vary across ethnic groups, which may manifest as differences in imaging phenotypes. 10 A pan-ethnic meta-analysis risks obscuring these population-specific signals, potentially yielding summary estimates that are disconnected from the realities of Chinese clinical practice, where the demand for accurate, noninvasive prediction is greatest. Consequently, this review focuses specifically on sensitizing EGFR mutations (e.g. exon 19 deletions and exon 21 L858R mutations), excluding rare variants or primary resistance mutations.

Materials and methods

Literature search strategy

This systematic review strictly adhered to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement and was prospectively registered with the International Prospective Register of Systematic Reviews (PROSPERO) (CRD420251273027).

A systematic search of PubMed, Embase, Web of Science, the Cochrane Library, Scopus, China National Knowledge Infrastructure (CNKI), Wanfang, VIP Database, and Chinese Biomedical Literature (CBM) Database was conducted to retrieve all relevant records from inception to 31 October 2025. The search strategy used a combination of Medical Subject Headings (MeSH) terms and free-text keywords across five domains: disease (“lung adenocarcinoma,” “LUAD,” “non-small cell lung cancer”), methodology (“radiomics,” “radiogenomics”), technology (“machine learning,” “deep learning,” “artificial intelligence”); imaging (“computed tomography,” “CT,” “PET/CT”), and biomarker (“EGFR,” “epidermal growth factor receptor mutation”). Manual screening of reference lists was also performed to supplement the electronic search.

Inclusion and exclusion criteria

Retrieved records were managed using EndNote 2025 (Clarivate Analytics). Studies were selected based on predefined eligibility criteria.

The inclusion criteria are as follows:

  • Population. Pathologically confirmed LUAD patients without prior anti-tumor therapy;

  • Index test. Radiomics models using ML algorithms derived from preoperative CT or PET/CT images;

  • Reference standard. EGFR mutation status confirmed by tissue-based molecular assays (e.g. next-generation sequencing, Sanger sequencing, or amplification refractory mutation system);

  • Study design. Diagnostic accuracy studies providing sufficient data to construct a 2 × 2 contingency table (true positive (TP), false positive (FP), true negative (TN), false negative (FN));

  • Language. English or Chinese.

The exclusion criteria are as follows:

  • Conference abstracts, reviews, meta-analyses, case reports, animal studies, and editorials;

  • Studies lacking extractable diagnostic accuracy data;

  • Duplicate publications (only the most complete or most recent version retained);

  • Studies without full-text availability or outside the scope of the review.

Study selection and data extraction

Two investigators independently screened titles, abstracts, and full texts. Any discrepancies were resolved through discussion or by consultation with a third reviewer (G.C.). A standardized data extraction form was used to collect the following information: (a) baseline study characteristics (first author, publication year, and country); (b) patient demographics (sample size, age, sex, and tumor stage); (c) technical parameters (imaging modality, region of interest (ROI) segmentation, ML algorithm, and validation strategy); and (d) diagnostic performance metrics (sensitivity, specificity, accuracy, area under the curve (AUC), TP/FP/TN/FN values). To reduce overfitting bias, data were preferentially extracted from independent test sets or external validation cohorts rather than training sets.

Risk of bias assessment

Two reviewers independently evaluated methodological quality using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS)-2 tool. Risk of bias and applicability concerns were evaluated across four domains: patient selection, index test, reference standard, and flow and timing. Each domain was classified as having a low, high, or unclear risk of bias.

Statistical analysis

Statistical analyses were performed using Meta-Disc 1.4 and Stata 18.0. The presence of a threshold effect was evaluated using Spearman's correlation between the logit of sensitivity and the logit of (1−specificity), with P <0.05 indicating statistical significance. Heterogeneity was assessed using Cochran's Q test (P < 0.10) and the I2 statistic, with I2 >50% indicating substantial heterogeneity.

A bivariate random-effects model was used to pool diagnostic accuracy estimates in the presence of significant heterogeneity; otherwise, a fixed-effects model was applied. Pooled sensitivity, specificity, positive likelihood ratio (PLR), negative likelihood ratio (NLR), diagnostic odds ratio (DOR), and corresponding 95% confidence intervals (CIs) were calculated. A summary receiver operating characteristic (SROC) curve was constructed, with the AUC used as the overall performance metric.

Predefined subgroup analyses were conducted stratified by three factors: (a) imaging modality (CT vs. PET/CT); (b) algorithm type (DL vs. conventional ML); and (c) validation strategy (external vs. internal). In this study, ML refers to the broad category, including conventional ML (e.g. support vector machines (SVMs) and random forests (RFs)) and DL(e.g. convolutional neural networks (CNN) and residual networks (ResNet)). Between-group differences were statistically compared. Univariate meta-regression was performed to explore imaging modality and algorithm type as potential sources of heterogeneity. Galbraith plot analysis was used to assess the robustness of pooled estimates. Publication bias was assessed using Deeks’ funnel plot asymmetry test, P <0.05 indicating significant bias.

Results

Study selection and characteristics

The initial search yielded 1684 records. After removing 412 duplicates, 1272 records were screened based on titles and abstracts. Full-text review of 98 articles resulted in the inclusion of 13 studies in the meta-analysis. Figure 1 illustrates the PRISMA flow diagram.

Figure 1.

Figure 1.

PRISMA 2020 flow.

PRISMA: Preferred Reporting Items for Systematic Reviews and Meta-Analyses.

All included studies were conducted in China and encompassed 6628 patients (4763 in training cohorts and 1865 in testing cohorts). Seven studies used CT, while six used PET/CT. Three studies employed conventional ML, whereas 10 adopted DL approaches, predominantly CNN. Table 1 summarizes the baseline characteristics.1123

Table 1.

Characteristics and diagnostic performance of included studies.

Study (year) Modality Algorithm Validation TP FP FN TN Sensitivity (%) Specificity (%) AUC (95% CI)
Shao et al.11 PET/CT DL (CNN) Internal 48 14 23 27 67.6 65.9 0.73 (0.629–0.830)
Zhang et al.12 CT DL (SE-CNN) External 51 13 33 108 60.7 89.3 0.841
Ge et al.13 PET/CT ML (SVM) Internal 23 24 23 70 50.0 74.5 0.672 (0.586–0.741)
Huang et al.14 CT DL (ResNet) Internal 41 2 6 19 87.9 91.4 0.916
Zhang et al.15 CT DL (3D-CNN) External 49 7 16 60 75.4 89.6 0.874 (0.820–0.924)
Gao et al.16 PET/CT ML (RF) Internal 56 18 14 23 80.0 56.1 0.704
Huang et al.17 PET/CT DL (Inc-v3) Internal 52 13 6 57 89.4 81.4 0.921 (0.904–0.937)
Wang et al.18 CT DL (3D-CNN) Internal 38 11 10 41 78.6 78.6 0.781 (0.638–0.923)
Deng et al.19 PET/CT ML (SVM) Internal 14 1 3 16 82.0 94.0 0.94
Weng et al.20 CT DL (ViT) External 9 4 1 16 90.0 80.0 0.885
Yin et al.21 PET/CT DL (SE-ResNet) Internal 41 10 10 42 80.4 80.8 0.84 (0.75–0.90)
Song et al.22 CT DL (EML) Internal 100 23 51 93 66.2 80.2 0.813 (0.763–0.864)
Xiong et al.23 CT DL (CNN) Internal 53 6 40 62 57.1 91.0 0.776 (0.702–0.849)

TP: true positive; FP: false positive; FN: false negative; TN: true negative; AUC: area under the curve; CT: computed tomography; PET/CT: positron emission tomography-computed tomography; DL: deep learning; ML: machine learning; CNN: convolutional neural network; SE-CNN: squeeze-and-excitation convolutional neural network; ResNet: residual network; Inc-v3: Inception-v3; 3D-CNN: three-dimensional convolutional neural network; ViT: vision transformer; SVM: support vector machine; RF: random forest; EML: ensemble machine learning.

Risk of bias assessment

Figure 2 presents the QUADAS-2 assessment. Most studies showed a high risk of bias in patient selection, primarily due to retrospective study designs and non-consecutive patient enrollment. Applicability concerns related to the index test were generally low.

Figure 2.

Figure 2.

QUADAS-2.

QUADAS-2: Quality Assessment of Diagnostic Accuracy Studies-2.

Meta-analysis of diagnostic performance

No threshold effect was detected (Spearman ρ = −0.129, P = 0.670). However, significant heterogeneity remained (Cochran's Q = 24.29, P < 0.001; I2 = 92%), necessitating the use of a bivariate random-effects model.

The pooled sensitivity was 71% (95% CI: 68%–74%), and the pooled specificity was 81% (95% CI: 78%–84%). The PLR was 4.1 (95% CI: 3.0–5.5), the NLR was 0.31 (95% CI: 0.24–0.41), and the DOR was 12.5 (95% CI: 7.7–20.4). The SROC curve indicated favorable overall diagnostic performance, with a summary AUC of 0.85 (95% CI: 0.82–0.88) (Figure 3).

Figure 3.

Figure 3.

Forest plots of diagnostic accuracy by subgroup. (a) Subgroup by algorithm (deep learning vs. conventional machine learning); (b) subgroup by imaging modality (PET/CT vs. CT). Squares represent point estimates for each study with size proportional to study weight. Horizontal lines represent 95% confidence intervals. Diamonds represent pooled estimates. Heterogeneity: I2 = 73.2%, P < 0.001 (random-effects model).

ET-CT: positron emission tomography-computed tomography; CT: computed tomography.

Subgroup analysis and meta-regression

Figures 3 and 4 display the forest plots and hierarchical SROC (HSROC) curves. Algorithm type. DL models achieved a significantly higher summary AUC (0.871, 95% CI: 0.832–0.918) than conventional ML models (0.798, 95% CI: 0.517–1.000; P = 0.012). However, the wide CI for conventional ML should be interpreted with caution due to the limited number of studies (n = 3).

Figure 4.

Figure 4.

Pooled diagnostic performance: forest plots and SROC curve. (a) Forest plots of pooled sensitivity and specificity; (b) SROC curve. Squares denote point estimates with 95% CIs (diamonds: pooled estimates). In (b), circle size reflects sample size; the dashed line indicates random chance.

CI: confidence interval; SROC: summary receiver operating characteristic.

Imaging modality. CT-based models demonstrated higher diagnostic accuracy than PET/CT-based models (AUC: 0.879, 95% CI: 0.820–0.939 vs. 0.828, 95% CI: 0.721–0.935; P = 0.038).

Validation strategy. Studies incorporating independent external validation exhibited superior performance (AUC: 0.922, 95% CI: 0.850–0.993) compared with those restricted to internal validation (AUC: 0.841, 95% CI: 0.776–0.907; P = 0.006).

Univariable meta-regression was performed to explore potential sources of heterogeneity. Imaging modality emerged as a significant moderator of the DOR, with a coefficient of 0.45 (95% CI: 0.12–0.78; P < 0.05), explaining 37% of the between-study variance (R2 = 0.37). Consistent with this finding, subgroup analysis showed that CT-based studies yielded a higher pooled OR than PET/CT-based studies (1.86 vs. 1.03). In contrast, ML algorithm type (DL vs. conventional ML) was not significantly associated with the DOR (coefficient = −0.08, 95% CI: −0.35–0.19; P = 0.56), as reflected by overlapping 95% CIs between subgroups. Similarly, subgroup analysis stratified by validation type (internal vs. external) showed no significant difference in pooled OR (1.54 vs. 1.09), consistent with the meta-regression results indicating no significant moderating effect of validation approach.

Sensitivity analysis and publication bias

Galbraith plot revealed no significant outliers. Deeks’ funnel plot asymmetry test indicated no significant publication bias (P = 0.46) (Figure 5).

Figure 5.

Figure 5.

Deeks’ funnel plot asymmetry test for publication bias. The x-axis represents the natural logarithm of the diagnostic odds ratio, and the y-axis represents the inverse square root of the effective sample size (1/√ESS).

ESS: effective sample size.

Discussion

To our knowledge, this is the first meta-analysis specifically evaluating ML-based radiomics models for predicting EGFR mutations in Chinese patients with LUAD. The findings demonstrate the potential of this noninvasive approach, with a summary AUC of 0.85 (95% CI: 0.82–0.88), alongside balanced sensitivity (71%) and specificity (81%).

The exclusive focus on the Chinese population, which carries the highest global burden of EGFR-mutated lung cancer, substantially enhances the clinical relevance of these findings.24,25 The high pre-test probability in this cohort suggests that radiomics models validated in this setting provide a more robust and clinically actionable benchmark than those derived from heterogeneous populations.

Furthermore, although EGFR mutations commonly occur across exons 18–21 in this population, the lack of disaggregated data prevented subgroup analyses based on specific mutation subtypes. This represents an important knowledge gap that should be addressed in future primary studies.

Imaging modality: rethinking the value proposition

A key finding of this study is the superior accuracy of CT-based models compared with PET/CT-based models, a difference that contributes significantly to observed heterogeneity. This challenges the assumption that multimodal metabolic imaging is inherently necessary for EGFR mutation prediction. For this task, the high-dimensional morphologic and textural features derived from routine CT appear not only sufficient but potentially more generalizable in clinical practice. 26 This observation is consistent with the principle that model performance is constrained by the intrinsic “information density” of the input data relative to the clinical prediction task. 27 Although PET/CT provides additional metabolic information, its incremental predictive value may be limited by higher cost and lower inter-center standardization. However, definitive conclusions require prospective intra-patient comparative studies.

Algorithm comparison: a nuanced perspective

As anticipated, DL models outperformed conventional ML models. This finding is consistent with the theoretical advantage of DL in automated feature representation learning. However, the observed difference (AUC: 0.871 vs. 0.798) should be interpreted with caution, as the small number of conventional ML studies resulted in wide CIs, limiting the strength of any definitive conclusion. More importantly, this comparison may underestimate the potential of both approaches. The future of predictive oncology is likely to involve multimodal data integration, combining imaging, clinical, and genomic information. Within such high-dimensional frameworks, the ability of DL methods to learn abstract representations and integrate heterogeneous data types is likely to provide a comparative advantage over conventional ML approaches.28,29 Nevertheless, conventional ML remains valuable due to its interpretability, whereas DL is expected to play a central role in the development of more complex, integrated predictive models.

The imperative of rigorous validation

The notable finding of this study is the substantial performance difference between models using independent external validation and those relying solely on internal validation (AUC: 0.922 vs. 0.841). This difference highlights a critical principle in medical artificial intelligence: robust validation strategies are more indicative of real-world clinical utility than high performance metrics derived from internal datasets alone. A model achieving an AUC of 0.95 in internal validation but failing in external validation has limited clinical applicability, a limitation that has been frequently observed in lung cancer radiomics research.30,31 Therefore, independent external validation should not be considered an optional enhancement but rather a fundamental requirement for assessing the clinical readiness of predictive models.32,33

Limitations and future directions

This study has several limitations that should be considered when interpreting the findings. First, although a bivariate random-effects model was used, the high heterogeneity (I2 = 92%) and wide CIs in some subgroups indicate that the pooled estimates reflect a highly variable evidence base rather than a precise or universal effect estimate. Differences in radiomics workflows—including manual versus semi-automated segmentation and the use of different feature extraction platforms (e.g. PyRadiomics versus in-house software)—limited direct comparability across studies and hindered identification of an optimal technical pipeline. Second, the exclusive reliance on retrospective studies conducted in a single country (China) introduces potential selection and spectrum bias, which may have inflated diagnostic performance estimates. 34 Although focusing on Chinese patients enhances the clinical relevance for this high-prevalence population, it limits the generalizability of the findings to other ethnic groups with different EGFR mutation profiles. Third, the relatively small number of included studies, particularly within the conventional ML subgroup, reduced the statistical power of meta-regression analyses, making these findings exploratory in nature. Furthermore, insufficient primary data prevented subgroup analyses based on specific EGFR mutation subtypes (e.g. exon 19 deletions versus L858R mutations) or co-mutation patterns. The predictive value of radiomics may vary significantly across these molecularly distinct entities, representing a substantial knowledge gap. Finally, from a clinical translation perspective, the “black-box” nature of many advanced ML and DL models, together with the absence of cost-effectiveness analyses and integration studies within real-world clinical decision-making pathways, highlights substantial practical barriers that cannot be addressed by diagnostic accuracy metrics alone.35,36

To bridge these gaps, we propose a targeted agenda for future research: (a) large-scale, prospective, multicenter trials using pre-registered, standardized protocols; (b) head-to-head, intra-patient comparisons of CT and PET/CT radiomics to definitively quantify their comparative value; (c) mandatory independent external validation as a core reporting requirement for all predictive model studies; and (d) expansion beyond binary EGFR prediction toward models capable of distinguishing specific sensitizing mutations, resistance mechanisms, and alterations in other driver genes (e.g. anaplastic lymphoma kinase (ALK) and ROS proto-oncogene 1 (ROS1)). Only through such rigorous methodological advancement can radiomics be fully integrated into precision oncology.

Conclusion

ML-based radiomics models demonstrate robust diagnostic accuracy for predicting EGFR mutation status in Chinese patients with LUAD, achieving a summary AUC of 0.85. CT-based DL models subjected to independent external validation represent the most promising technical configuration currently available. However, the retrospective design and substantial methodological heterogeneity across the included studies preclude immediate clinical deployment. High-quality prospective multicenter trials are essential to validate these findings and standardize radiomics workflows before routine clinical implementation.

Supplemental Material

sj-docx-1-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-1-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research

sj-docx-2-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-2-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research

sj-docx-3-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-3-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research

Acknowledgements

The authors thank the journal editorial office for linguistic assistance. We acknowledge the use of Deepseek AI-assisted tools for grammar and language polishing during the preparation of this manuscript.

Footnotes

Author contributions: Conceptualization: Jia Yang

Data curation: Jia Yang and Junyu Jiang

Formal analysis: Jia Yang

Funding acquisition: None

Investigation: Junyu Jiang

Methodology: Jia Yang and Junyu Jiang

Project administration: Junyu Jiang

Resources: Jia Yang and Junyu Jiang

Software: Jin Peng

Supervision: Junyu Jiang and Guangwen Chen

Validation: Jia Yang, Junyu Jiang, Jin Peng, Jie Li, and Peng Mi

Visualization: Jia Yang and Junyu Jiang

Writing—original draft: Jia Yang

Writing—review and editing: Jia Yang, Junyu Jiang, and Guangwen Chen.

Funding: The authors received no financial support for the research, authorship, and/or publication of this article.

Declaration of conflicting interests: The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Data availability statement: The data supporting the findings of this study are available from the corresponding author upon reasonable request.

Supplemental material: Supplementary material for this article is available online.

References

  • 1.Gridelli C. Histology-based treatment: a new scenario in the management of advanced nonsmall cell lung cancer. Curr Opin Oncol 2009; 21: 97–98. [DOI] [PubMed] [Google Scholar]
  • 2.. Blandin Knight SB, Crosbie PA, Balata H, et al. Progress and prospects of early detection in lung cancer. Open Biol 2017; 7: 170070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lee SH, Kim WS, Choi YD, et al. Analysis of mutations in epidermal growth factor receptor gene in Korean patients with non-small cell lung cancer: summary of a nationwide survey. J Pathol Transl Med 2015; 49: 481–488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Marrocco I, Yarden Y. Resistance of lung cancer to EGFR-specific kinase inhibitors: activation of bypass pathways and endogenous mutators. Cancers 2023; 15: 5009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Linardou H, Dahabreh IJ, Bafaloukos D, et al. Somatic EGFR mutations and efficacy of tyrosine kinase inhibitors in NSCLC. Nat Rev Clin Oncol 2009; 6: 352–366. [DOI] [PubMed] [Google Scholar]
  • 6.Weber B, Meldgaard P, Hager H, et al. Detection of EGFR mutations in plasma and biopsies from non-small cell lung cancer patients by allele-specific PCR assays. BMC Cancer 2014; 14: 294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Scott IA, Cook D, Coiera EW, et al. Machine learning in clinical practice: prospects and pitfalls. Med J Aust 2019; 211: 203–205.e1. [DOI] [PubMed] [Google Scholar]
  • 8.Chang C, Zhou S, Yu H, et al. A clinically practical radiomics-clinical combined model based on PET/CT data and nomogram predicts EGFR mutation in lung adenocarcinoma. Eur Radiol 2021; 31: 6259–6268. [DOI] [PubMed] [Google Scholar]
  • 9.Zhang J, Zhao X, Zhao Y, et al. Value of pre-therapy 18F-FDG PET/CT radiomics in predicting EGFR mutation status in patients with non-small cell lung cancer. Eur J Nucl Med Mol Imaging 2020; 47: 1137–1146. [DOI] [PubMed] [Google Scholar]
  • 10.Qin BM, Chen X, Zhu JD, et al. Identification of EGFR kinase domain mutations among lung cancer patients in China: implication for targeted cancer therapy. Cell Res 2005; 15: 212–217. [DOI] [PubMed] [Google Scholar]
  • 11.Shao X, Ge X, Gao J, et al. Transfer learning-based PET/CT three-dimensional convolutional neural network fusion of image and clinical information for prediction of EGFR mutation in lung adenocarcinoma. BMC Med Imaging 2024; 24: 54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Zhang B, Qi S, Pan X, et al. Deep CNN model using CT radiomics feature mapping recognizes EGFR gene mutation status of lung adenocarcinoma. Front Oncol 2021; 10: 598721. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ge XY. Based on PET/CT three-classification machine learning model for predicting EGFR mutation subtypes in lung adenocarcinoma. Master’s thesis, Soochow University, Suzhou, 2025.
  • 14.Huang LY, Xu L, Wen LC, et al. CT radiomics model and deep learning technique for predicting EGFR mutation in lung adenocarcinoma. Radiol Pract 2022; 37: 971–976. [Google Scholar]
  • 15.Zhang G, Shang L, Cao Y, et al. Prediction of epidermal growth factor receptor (EGFR) mutation status in lung adenocarcinoma patients on computed tomography (CT) images using 3-dimensional (3D) convolutional neural network. Quant Imaging Med Surg 2024; 14: 6048–6059. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Gao J, Niu R, Shi Y, et al. The predictive value of [18F]FDG PET/CT radiomics combined with clinical features for EGFR mutation status in different clinical staging of lung adenocarcinoma. EJNMMI Res 2023; 13: 26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Huang L, Kong W, Luo Y, et al. Predicting epidermal growth factor receptor mutation status of lung adenocarcinoma based on PET/CT images using deep learning. Front Oncol 2024; 14: 1458374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Wang XY, Wu SH, Ren J, et al. Predicting gene comutation of EGFR and TP53 by radiomics and deep learning in patients with lung adenocarcinomas. J Thorac Imaging 2025; 40: e0817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Deng Z, Jin D, Huang P, et al. Predictive models of epidermal growth factor receptor mutation in lung adenocarcinoma using PET/CT-based radiomics features. Med Phys. 2025; 52: 3697–3710. [DOI] [PubMed] [Google Scholar]
  • 20.Weng L, Xu Y, Chen Y, et al. Using vision transformer for high robustness and generalization in predicting EGFR mutation status in lung adenocarcinoma. Clin Transl Oncol 2024; 26: 1438–1445. [DOI] [PubMed] [Google Scholar]
  • 21.Yin G, Wang Z, Song Y, et al. Prediction of EGFR mutation status based on 18F-FDG PET/CT imaging using deep learning-based model in lung adenocarcinoma. Front Oncol 2021; 11: 709137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Song Q, Li X, Song B, et al. An online explainable ensemble machine learning model for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma. Transl Lung Cancer Res 2025; 14: 2670–2687. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Xiong JF, Jia TY, Li XY, et al. Identifying epidermal growth factor receptor mutation status in patients with lung adenocarcinoma by three-dimensional convolutional neural networks. Br J Radiol 2018; 91: 20180334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhang Q, Cui Y, Zhang J. et al. Comparison of the characteristics of uncommon epidermal growth factor receptor (EGFR) mutations and EGFR-tyrosine kinase inhibitor treatment in patients with non-small cell lung cancer from different ethnic groups. Exp Ther Med 2020; 20: 2707–2714. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Carbonnaux M, Souquet PJ, Meert AP, et al. Inequalities in lung cancer: a world of EGFR. Eur Respir J 2016; 47: 1502–1509. [DOI] [PubMed] [Google Scholar]
  • 26.Mei D, Luo Y, Wang Y, et al. CT Texture analysis of lung adenocarcinoma: can Radiomic features be surrogate biomarkers for EGFR mutation statuses. Cancer Imaging 2018; 18: 52. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yu M, Sun A. Dataset versus reality: understanding model performance from the perspective of information need. J Assoc Inf Sci Technol 2023; 74: 1293–1306. [Google Scholar]
  • 28.Deng Y, Ren Z, Kong Y, et al. A hierarchical fused fuzzy deep neural network for data classification. IEEE Trans Fuzzy Syst 2016; 25: 1006–1012. [Google Scholar]
  • 29.Sanakoyeu A, Ma P, Tschernezki V, et al. Improving deep metric learning by divide and conquer. IEEE Trans Pattern Anal Mach Intell 2021; 44: 8306–8320. [DOI] [PubMed] [Google Scholar]
  • 30.Van Calster B, Steyerberg EW, Wynants L, et al. There is no such thing as a validated prediction model. BMC Med 2023; 21: 70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Collins GS, de Groot JA, Dutton S, et al. External validation of multivariable prediction models: a systematic review of methodological conduct and reporting. BMC Med Res Methodol 2014; 14: 40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.La Roi-Teeuw HM, van Royen FS, De Hond A, et al. Don’t be misled: 3 misconceptions about external validation of clinical prediction models. J Clin Epidemiol 2024; 172: 111387. [DOI] [PubMed] [Google Scholar]
  • 33.Cabitza F, Campagner A, Soares F, et al. The importance of being external. Methodological insights for the external validation of machine learning models in medicine. Comput Methods Programs Biomed 2021; 208: 106288. [DOI] [PubMed] [Google Scholar]
  • 34.Talari K, Goyal M. Retrospective studies–utility and caveats. J R Coll Physicians Edinb 2020; 50: 398–402. [DOI] [PubMed] [Google Scholar]
  • 35.London AJ. Artificial intelligence and black-box medical decisions: accuracy versus explainability. Hastings Cent Rep 2019; 49: 15–21. [DOI] [PubMed] [Google Scholar]
  • 36.Nwanosike EM, Conway BR, Merchant HA, et al. Potential applications and performance of machine learning techniques and algorithms in clinical practice: a systematic review. Int J Med Inform 2022; 159: 104679. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

sj-docx-1-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-1-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research

sj-docx-2-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-2-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research

sj-docx-3-imr-10.1177_03000605261460311 - Supplemental material for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis

Supplemental material, sj-docx-3-imr-10.1177_03000605261460311 for Diagnostic performance of machine learning-based radiomics models for predicting epidermal growth factor receptor mutation status in lung adenocarcinoma in Chinese patients: A systematic review and meta-analysis by Jia Yang, Junyu Jiang, Jin Peng, Jie Li, Peng Mi and Guangwen Chen in Journal of International Medical Research


Articles from The Journal of International Medical Research are provided here courtesy of SAGE Publications

RESOURCES