Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2025 Dec 18;26:25. doi: 10.1186/s12911-025-03321-z

Machine learning in early screening for high-grade cervical intraepithelial neoplasia using blood testing

Congbo Yue 1,2,#, Shichao Liu 3,#, Wenhua Wang 1,4, Yu Zhao 5, Xiaofeng Zhang 1,2, Guanghui Zhao 1,2,✉
PMCID: PMC12829010  PMID: 41413554

Abstract

Background

High-grade cervical intraepithelial neoplasia (CIN2/3) is a critical precursor to cervical cancer, yet current screening methods (e.g., HPV testing, colposcopy) face challenges in accessibility and invasiveness, especially in resource-limited settings. We aimed to develop a non-invasive, machine learning (ML)-based model using routine blood biomarkers. This model is intended to assess the risk of high-grade CIN and potentially serve as a triage tool before colposcopy.

Methods

Data were collected from two groups: 128 high-grade CIN (CIN2/3) and 120 low-grade CIN (CIN1) patients. A total of 29 clinical characteristics and blood test measurements were considered for use in model development. Four feature selection algorithms (F-test, LASSO regression, decision tree, and random forest) were used to identify key predictors, and 11 machine learning algorithms were employed for model training. The dataset was split into training (70%) and testing (30%) cohorts. Model performance was evaluated using learning curves, receiver operating characteristic curves (ROC), area under the curve (AUC), Brier score, calibration curves, Precision-Recall (PR) curves, and Decision Curve Analysis (DCA). A web-based calculator was developed for clinical deployment. We assessed feature importance using the SHapley Additive exPlanation (SHAP) approach.

Results

Key features selected for the model included creatinine (CREA), red blood cell count (RBC), neutrophil ratio (NEU%), direct bilirubin (DBIL), and monocyte count (MON). The Support Vector Machine (SVM) algorithm achieved the best predictive performance, with an AUC of 0.75 (95% CI: 0.69–0.80) and a Brier score of 0.21 (95% CI: 0.17–0.28). By employing the SHAP method, we identified the variables that contributed to the model. The web tool (https://dvhl6xsf29zmdewixjx7kz.streamlit.app) provides real-time risk stratification.

Conclusions

The model demonstrated strong performance across various validation metrics, with the SVM algorithm achieving an AUC of 0.75, indicating potential clinical utility. We also developed a web-based calculator to estimate high-grade CIN.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12911-025-03321-z.

Keywords: Cervical intraepithelial neoplasia, Machine learning, Blood biomarkers, Prediction model, Decision support tool

Background

Cervical intraepithelial neoplasia (CIN) is a precursor to cervical cancer. The progression from low-grade lesions (CIN1) to high-grade lesions (CIN3) represents a critical step towards invasive cervical cancer [1]. High-grade CIN, particularly CIN2 and CIN3, is a precursor of cervical cancer, which is among the leading causes of morbidity and mortality in women worldwide [2]. In 2020, there were approximately 604,000 new cases of cervical cancer and 342,000 related deaths worldwide, with disproportionately high rates in certain regions [3]. Although cervical cancer rates have declined in high-income countries due to widespread screening programs, the disease remains a major public health challenge in low- and middle-income countries [4]. Early detection and accurate histopathological classification of cervical high-grade squamous intraepithelial lesions (HSIL) are critical to preventing progression to invasive cervical cancer.

Several techniques have been employed to detect and diagnose cervical HSIL, including human papilloma virus (HPV) DNA testing, cytology, colposcopy, and biopsy [5, 6]. HPV testing is a non-invasive and highly sensitive method for identification of women at risk of developing cervical cancer [7, 8]. Colposcopy, an optical examination of the cervix, can provide more definitive information but involves subjective interpretation and often requires specialized training [9]. Biopsy, the gold standard for diagnosing cervical HSIL, is invasive and unsuitable for widespread use due to its cost and need for skilled professionals [10]. Moreover, although these methods are in widespread use, they are often not seamlessly integrated in an automated and cost-effective process, especially in resource-limited regions.

Artificial intelligence and machine learning have demonstrated a powerful capacity to advance early detection and risk stratification in healthcare. Predictive models utilizing routine clinical data, such as blood biomarkers and lifestyle factors, have been successfully applied to forecast outcomes like COVID-19 hospitalization [11] and type 2 diabetes mellitus (T2DM) [12, 13]. Such work confirms the utility of machine learning for addressing complex health risks, frequently by extracting compact, predictive feature sets from a wide array of initial variables. In the specific context of high-grade CIN, machine learning advances show promise for enhancing the accuracy of detection and prediction. Machine learning techniques have shown effectiveness in analyzing large-scale datasets, enabling identification of complex patterns and features overlooked by traditional diagnostic methods [14]. Numerous studies have demonstrated the potential of machine learning models in medical diagnostics, particularly in the fields of imaging and pathology, in which they contribute to increased diagnostic precision and reproducibility [15–18]. However, there is a research gap regarding specific applications of machine learning for prediction of high-grade CIN, with few comprehensive studies having developed and validated models for this purpose [19]. Furthermore, the implementation of such AI solutions in low-resource settings presents unique challenges and opportunities that warrant further exploration [20].

The present study aimed to develop a blood-based ML model to identify high-risk patients most likely to benefit from colposcopy and biopsy, thus addressing a critical gap in current screening techniques. By employing a machine learning approach, we sought to provide an effective, accessible, and standardized method for high-grade CIN prediction that could reduce reliance on subjective interpretation of cytology and histology results and promote early detection. The predictive model developed in this study was trained and validated on separate datasets. Its performance was evaluated using key indicators such as sensitivity, specificity, and AUC to optimize diagnostic accuracy. We thus present an evidence-based, machine-learning-driven diagnostic tool that could facilitate earlier intervention. Ultimately, this may reduce the incidence and mortality of cervical cancer, particularly in resource-limited settings [21, 22].

Methods

Study population and design

For this study, we retrospectively recruited two diagnostic cohorts from the gynecology department of Peking University People’s Hospital, Qingdao, China. The recruitment period was from January 1, 2024, to May 31, 2025, and the cohorts included a low-grade CIN group (CIN1) and a high-grade CIN group (CIN2/3). The inclusion criteria were as follows: (1) patients diagnosed with high-grade CIN or low-grade CIN based on histopathological examination of biopsy specimens; (2) patients who had not undergone any prior treatment for cervical lesions (e.g., loop electrosurgical excision procedure, conization, or cryotherapy) [23]; and (3) patients with complete and available relevant blood test data. The exclusion criteria were: (1) previous treatment for cervical lesions; (2) insufficient or missing data; or (3) other gynecological conditions that could interfere with the assessment of cervical lesions (e.g., severe pelvic inflammatory disease or endometrial cancer) [24]. A total of 260 cases were initially considered, with 12 excluded due to incomplete data (> 25% of variables missing). No imputation was performed; only complete cases were included in the final analysis. This study was approved by the Ethics Committee of Peking University People’s Hospital, Qingdao.

Data collection and feature selection

For this study, we collected baseline characteristic data, including patient age, CIN grade, and various predictor variables comprising routine hematological and biochemical parameters. The routine hematological parameters included red blood cell count (RBC), white blood cell count (WBC), hemoglobin (HGB), platelet count (PLT), neutrophil ratio (NEU%), lymphocyte ratio (LYM%), monocyte ratio (MON%), neutrophil count (NEU), lymphocyte count (LYM), monocyte count (MON), hematocrit (HCT), mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), mean platelet volume (MPV), and platelet-large cell ratio (P_LCR). The routine biochemical parameters included indirect bilirubin (IBIL), direct bilirubin (DBIL), total bilirubin (TBIL), total protein (TP), globulin (GLB), albumin (ALB), aspartate aminotransferase (AST), alanine aminotransferase (ALT), gamma-glutamyl transferase (GGT), alkaline phosphatase (ALP), creatinine (CREA), and urea (UREA). All hematological parameters were measured using a Mindray BC-6800Plus series auto hematology analyzer (Shenzhen Mindray Bio-Medical Electronics Co., Ltd.). Biochemical parameters were measured using a Mindray BS-2800 M automatic biochemical analyzer (Shenzhen Mindray Bio-Medical Electronics Co., Ltd.). All laboratory data, as well as age, were standardized and used for analysis.

After data collection, initial feature selection was performed using four algorithms (F-test, LASSO regression, decision tree, and random forest) to identify candidate features. Subsequently, Pearson correlation analysis and LASSO regression were applied to select the most suitable and non-redundant features for constructing a simplified predictive model.​​.

Machine learning algorithms

A total of 11 machine learning algorithms were used for model development and evaluation: Naive Bayes (NB), K-Nearest Neighbors (KNN), Logistic Regression (LR), Random Forest (RF), Decision Tree (DT), Artificial Neural Network (ANN), Support Vector Machine (SVM), Gradient Boosting Decision Trees (GBDT), Light Gradient Boosting Machine (LightGBM), Adaptive Boosting (AdaBoost), and Extreme Gradient Boosting (XGBoost).

Model establishment and evaluation

The entire model was implemented using Python (3.10). All samples were randomly divided into a training cohort and a testing cohort in a 7:3 ratio. Scikit-learn (version 1.2.2) was used for data splitting, with stratified sampling to ensure the class distribution remained consistent between the training and testing cohorts. Eleven machine learning algorithms were built using Python libraries (scikit-learn 1.2.2, XGBoost 1.7.4, LightGBM 4.0.0). Grid search and ten-fold cross-validation were used to select the optimal hyperparameters for each algorithm. We compared 11 distinct machine learning algorithms using the selected predictor features and identified the best-performing model.

The effectiveness of different algorithms was evaluated using ROC curves plotted with matplotlib (3.7.1) and machine learning curves plotted with scikit-plot (0.3.7). The optimal diagnostic model was selected by combining the AUC and the performance of the algorithm in machine learning. Predictive accuracy was evaluated using the ROC AUC and calibration curve. Decision curve analysis (DCA) was employed to assess the clinical usefulness and net benefit of the models. Precision–recall (PR) curves were also used, as they provide a more informative indication of performance than accuracy or ROC evaluation when positive class prediction is of primary interest. Further, Shapley Additive Expla nations (SHAP 0.44.0) values were applied to ascertain the variable importance of each predictor and to visualize their association with risk stratification. Details of the study design are displayed in Fig. 1.

Fig. 1.

Fig. 1

Flow chart of the survey design

Web application development

To bridge the gap between research and clinical practice, we developed a user-friendly web application based on the final prediction model using the Python Streamlit framework. This tool allows healthcare professionals to input patient data and receive real-time predictions of the probability of high-grade CIN. It also generates a force plot for each participant, visually showing how different features contribute to the prediction.

Statistical processing

Initial statistical analysis (e.g., descriptive statistics, normality tests) was conducted using SPSS 23.0, while all machine learning model development and evaluation were performed in Python (3.10). The normality of data distribution was assessed using the Shapiro-Wilk test. Normally distributed data are expressed as mean standard deviation (x̄±SD), and inter-group comparisons were made using independent-samples t-tests. The assumption of homogeneity of variance was verified by Levene’s test; if violated, Welch’s t-test was applied. Non-normally distributed data are presented as median and interquartile range (IQR) and were compared using the Mann-Whitney U test. Statistical significance for all tests was determined at a threshold of P < 0.05.

Results

Clinical characteristics of the training cohort and testing cohort

A total of 248 patients, comprising two distinct diagnostic groups (128 high-grade CIN and 120 low-grade CIN cases), were included in the study. These patients were randomly divided into a training cohort and testing cohort. The baseline characteristics of patients in the training and testing cohorts, which were crucial for assessing group comparability, are shown in Table 1. The training cohort consisted of 173 cases (approximately 69.8% of the total sample), whereas the testing cohort comprised 75 cases (approximately 30.2%). The distribution of high-grade and low-grade CIN cases was well balanced between the two cohorts, ensuring consistent outcome proportions across both groups. It is noteworthy that UREA levels showed a statistically significant difference (p = 0.033) between the training and testing cohorts (Table 1). This is likely a chance finding from the random split. Crucially, UREA was not selected as a predictive feature in our final model. Furthermore, we confirmed that UREA levels did not differ significantly between high-grade and low-grade CIN groups within the entire cohort or within each sub-cohort (all p > 0.05), indicating this imbalance does not compromise our model’s conclusions.

Table 1.

Perioperative statistical data of participants

Variables Cohort Statistic P
Train (n = 173) Test (n = 75)

WBC

(Mean ± SD)

5.54 ± 1.50 5.59 ± 1.31 t=-0.241 0.810

RBC

(Mean ± SD)

4.29 ± 0.38 4.38 ± 0.35 t=-1.630 0.104

HGB

(Mean ± SD)

129.35 ± 12.40 132.33 ± 10.72 t=-1.809 0.072

PLT

(Mean ± SD)

245.38 ± 55.27 249.43 ± 62.35 t=-0.510 0.611

NEU%

(Mean ± SD)

60.04 ± 8.67 59.63 ± 8.86 t = 0.344 0.731

LYM%

(Mean ± SD)

32.29 ± 7.81 32.70 ± 8.32 t=-0.372 0.710

MON%

(Mean ± SD)

5.42 ± 1.47 5.35 ± 1.35 t = 0.338 0.735

LYM

(Mean ± SD)

1.74 ± 0.47 1.77 ± 0.43 t=-1.773 0.584

HCT

(Mean ± SD)

38.40 ± 3.17 39.15 ± 2.77 t=-1.773 0.077

MCV

(Mean ± SD)

89.70 ± 6.01 89.66 ± 5.02 t = 0.048 0.962

MCH

(Mean ± SD)

30.21 ± 2.49 30.31 ± 2.11 t=-0.309 0.758

MCHC

(Mean ± SD)

336.51 ± 8.82 337.75 ± 7.83 t=-1.049 0.295

MPV

(Mean ± SD)

9.84 ± 0.95 9.90 ± 1.15 t=-0.455 0.649

TP

(Mean ± SD)

73.34 ± 4.29 73.64 ± 4.19 t=-0.515 0.607

ALB

(Mean ± SD)

45.32 ± 2.43 45.76 ± 2.29 t=-1.339 0.182

GLB

(Mean ± SD)

28.02 ± 3.22 27.88 ± 3.02 t = 0.316 0.752

UREA

(Mean ± SD)

4.84 ± 1.13 5.18 ± 1.19 t=-2.139 0.033

CREA

(Mean ± SD)

56.63 ± 8.47 55.66 ± 10.64 t = 0.764 0.446

Age

M (Q1, Q3)

38.00(32.00,49.50) 37.00(31.00,51.00) Z=-0.781 0.435

NEU

M (Q1, Q3)

3.28(2.54,3.98) 3.26(2.54,4.01) Z=-0.220 0.826

MON

M (Q1, Q3)

0.28(0.23,0.33) 0.28(0.24,0.33) Z=-0.441 0.659

TBIL

M (Q1, Q3)

11.40(9.00,14.45) 10.90(9.50,13.30) Z=-0.765 0.444

DBIL

M (Q1, Q3)

3.20(2.60,4.25) 3.10(2.60,4.30) Z=-0.786 0.432

P_LCR

M (Q1, Q3)

24.20(20.05,28.65) 24.80(19.30,31.00) Z=-0.384 0.701

IBIL

M (Q1, Q3)

8.00(6.27,10.20) 7.71(6.20,9.60) Z=-0.589 0.556

ALT

M (Q1, Q3)

15.40(11.25,20.35) 16.70(12.90,22.10) Z=-1.869 0.062

AST

M (Q1, Q3)

17.30(15.00,20.00) 18.55(15.10,22.50) Z=-1.706 0.088

GGT

M (Q1, Q3)

17.40(13.75,22.55) 17.00(14.20,22.80) Z=-2.78 0.781

ALP

M (Q1, Q3)

60.00(48.65,76.55) 59.60(49.00,73.00) Z=-0.09 0.993

t: t-test, Z: Mann-Whitney test

SD: standard deviation, M: Median, Q1: 1st Quartile, Q3: 3rd Quartile

Characteristic indicators and screening results from different algorithms

We generated a Venn diagram based on the results of the four algorithms (F-test, LASSO regression, decision tree, and random forest) and found that the following seven features were selected by three or more algorithms: CREA, RBC, HCT, NEU%, DBIL, TBIL, and MON (Fig. 2A). Pearson correlation coefficient analysis was performed to evaluate pairwise correlations among these seven features. As shown in Fig. 2B, strong positive correlations were observed between HCT and RBC, as well as DBIL and TBIL, with Pearson correlation coefficients exceeding 0.6. Figure 2C presents the results of the LASSO regression analysis, which was used to reduce overfitting by selecting the most relevant features (CREA, RBC, NEU%, DBIL, and MON) for subsequent model construction.

Fig. 2.

Fig. 2

Comprehensive screening of the included features. (A) F test, LASSO regression, decision tree and random forest intersection of the four algorithms selected features. (B) Pearson correlation coefficient of 7 features. (C) The fitting degree of LASSO regression analysis parameters

Predictive accuracy of the developed model for triaging high-grade CIN

The five selected indicators were modeled and evaluated using 11 different algorithms. Participants were randomly divided into a training set of 173 cases and a test set of 75 cases at a ratio of 7:3. As shown in Fig. 3A and B, the model showed satisfactory overall performance, even after a reduction in the number of indicators. SVM was identified as the best learning algorithm; as shown in Fig. 4A–E, the learning curve, calibration curve, decision curve, PR curve, and ROC curve, all demonstrated the strong performance of the model constructed using the SVM algorithm. Specifically, the AUC was 0.75 (95% CI: 0.69–0.80), and the Brier score was 0.21 (95% CI: 0.17–0.28). The performance of 11 machine learning algorithms on five indicators: AUC, sensitivity, specificity, accuracy, and F1 score have been included in Supplementary Table 1.

Fig. 3.

Fig. 3

Comparison set of the area under the receiver operating curve before (A) and after (B) index screening

Fig. 4.

Fig. 4

The evaluation of machine learning algorithms by machine learning curve (A), calibration curve (B), DCA curve (C), PR curve (D), and ROC curve (E). (F) Interpretation of the model constructed by the SVM algorithm with 5 variables

To better elucidate the contributions of different factors to the prediction ability of our SVM-based model, we performed further evaluation using SHAP. The results indicated that RBC made the most significant contribution to the performance of the model, followed by CREA, NEU%, MON, and DBIL (Fig. 4F).

Clinical applications

To facilitate clinical applications of the SVM model, we developed a user-friendly web application (https://dvhl6xsf29zmdewixjx7kz.streamlit.app). This application allows clinicians to input values for the five key features and automatically predicts the risk of high-grade CIN for individual patients. A screenshot of the generalized model is shown in Fig. 5.

Fig. 5.

Fig. 5

An example output of the web application

Discussion

In this study, we developed and validated a machine learning model using five routine blood biomarkers (RBC, CREA, NEU%, DBIL, and MON) for the non-invasive risk stratification of high-grade CIN. The SVM-based model achieved an AUC of 0.75 (95% CI: 0.69–0.80) and a well-calibrated Brier score of 0.21 (95% CI: 0.17–0.28), demonstrating robust diagnostic potential.

We selected the SVM algorithm for our final model due to its demonstrated superiority in performance and particular suitability for our clinical dataset. SVM is renowned for its significant advantages in handling high-dimensional data with limited sample sizes. The algorithm effectively manages non-linear classification problems by employing kernel functions to map input features into a higher-dimensional space, where an optimal separating hyperplane can be constructed [25]. Compared to traditional models like logistic regression, SVM exhibits greater robustness against outliers and overfitting, making it particularly suitable for clinical datasets, which often contain inherent noise and potential missing values [26]. Furthermore, we used SHAP to achieve transparent interpretation of the model’s predictions. Grounded in the concept of Shapley values from cooperative game theory, SHAP provides a unified framework for quantifying the contribution of each feature to individual prediction outcomes, thereby offering both global and local interpretability [27]. This approach effectively addresses critical limitations of current screening paradigms, which are often hindered by limited accessibility, subjectivity, and invasiveness, particularly in resource-constrained regions [28].

The biomarkers identified here are all relevant to the pathophysiology of cervical cancer. A low RBC count is indicative of chronic anemia, a well-documented paraneoplastic syndrome in cervical cancer often resulting from chronic blood loss and cancer-associated inflammation that suppresses erythropoiesis [29]. MON and NEU% reflect chronic systemic inflammation, which actively promotes cancer development [30]. In HPV-related cervical lesions, local IL-8 overexpression is directly associated with an increase in NEU%, while CCL2-driven monocyte infiltration into the tumor stroma leads to a decrease in the exhaustion of peripheral blood MON count [31, 32]. Meanwhile, subtle alterations in CREA and DBIL may signal broader metabolic dysregulation associated with advanced disease, such as cancer cachexia, or reflect systemic inflammatory effects on metabolic and liver function [33, 34]. The integration of these routine biomarkers suggests a host profile marked by chronic inflammation, anemia, and metabolic shifts, providing a coherent clinical rationale for the model’s predictive logic.

The SVM-based model developed in this study achieved an AUC of 0.75 (95% CI: 0.69–0.80) for high-grade CIN risk stratification. This performance is moderate when compared to other machine learning-based prediction models in the same domain. For instance, Yuan et al. developed a classification model for CIN based on colposcopic images, achieving an AUC of 0.93 [35]. Similarly, Li et al. leveraged clinical parameters—such as HPV infection status, cytology results, and transformation zone type—to predict the risk of misdiagnosing low-grade squamous intraepithelial lesions (LSIL) as high-grade squamous intraepithelial lesions (HSIL), with an AUC of 0.936 [36]. However, such high-performance models are inherently dependent on specialized equipment, expert interpretation, and substantial costs, limiting their scalability in resource-constrained settings [4]. In contrast, our non-invasive approach leverages routine blood tests, generating an objective risk score that minimizes inter-observer variability and standardizes triage decisions [37]. Rather than replacing existing methods, our model serves as a complementary tool for secondary triage—particularly for HPV-positive patients—by identifying individuals who may benefit from intensified surveillance rather than immediate colposcopy, thereby reducing unnecessary invasive procedures [38, 39]. The model’s reliance on low-cost, widely available blood tests ensures high scalability, making it a viable strategy for resource-limited settings [4, 40]. Furthermore, its strong calibration (Brier score: 0.21) underscores its clinical utility by providing reliable probability estimates for individual risk assessment.

We developed an intuitive web application using the Streamlit framework to facilitate the diagnosis of high-grade CIN for healthcare providers. By utilizing simple and commonly available clinical indicators, the tool provides fast, real-time predictions without requiring specialized hardware or high-end computing equipment. Its user-friendly interface also enables smooth integration into standard clinical workflows, thereby promoting early detection and timely intervention for high-grade CIN across diverse healthcare environments, including those in low- and middle-income countries where the need for cost-effective triage tools is most acute [20]. Furthermore, the real-time capability of our web-based tool aligns with the growing trend of developing efficient AI models for clinical applications, as demonstrated in other gynecological contexts [41].

Although the model shows promise, it is important to note some limitations. The primary limitations relate to the sample size and the lack of external validation. The total of 248 participants, though sufficient for a preliminary study, remains modest for machine learning applications. As demonstrated by Rahimi et al. [42], large and diverse datasets are fundamental to ensuring the generalizability and robustness of predictive models. Due to the limited and homogeneous sample, our model’s performance (AUC 0.75) may be prone to overfitting, limiting its generalizability to broader, independent populations. To address this critical issue, future research must prioritize recruiting larger and more diverse cohorts.

Secondly, and of equal important, our model was developed and validated internally without an external cohort. The absence of external validation is a major constraint that affects the translational potential of our findings. To bridge this gap, we have initiated plans for a prospective, multi-center study aimed at externally validating the model’s performance across several institutions. This step will be crucial for confirming the model’s robustness and advancing it towards clinical application.

In addition, the model relies solely on blood-based features, whereas other critical factors, including HPV status and histopathological characteristics, are known to influence CIN progression [38]. Future research should therefore focus on expanding the cohort through multi-center collaborations to increase sample size and diversity and on integrating a wider range of clinical and molecular markers to enhance the accuracy, robustness, and external validity of the model.

Conclusions

In this study, we successfully developed and validated a non-invasive, machine-learning-based model that utilizes readily available, routine blood test parameters (specifically CREA, RBC, NEU%, DBIL, and MON) to assess the risk of high-grade CIN (CIN2/3) as a triage tool. The SVM model demonstrated moderate predictive accuracy (AUC of 0.75) and calibration (Brier score of 0.21) and thus represents a cost-effective and accessible means of identifying women at higher risk. A key future direction is the integration of our blood-based model with established biomarkers, notably HPV status and cytology. This multi-modal strategy promises to yield a more robust and accurate screening tool, potentially outperforming single-method assessments and significantly refining triage strategies. This approach, which has been translated into a web-based calculator, shows significant potential to enhance cervical cancer prevention by enabling early detection and prioritizing women needing more intensive follow-up, particularly in resource-limited settings.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (10.7KB, xlsx)

Acknowledgements

This study was supported by the ​​National High Performance Medical Device Innovation Center Project​​ (No. NMED2025KF-01-005) and the ​​Application of Intelligent Mutual Recognition of Inspection Results Project​​ (No. JYHRXZ2025B06). We sincerely thank the medical staff at ​​Peking University People’s Hospital Qingdao Hospital​​ for their assistance in data collection, as well as all patients who participated in this research.

Abbreviations

ML

Machine Learning

CIN

Cervical Intraepithelial Neoplasia

RBC

Red Blood Cell

NEU%

Neutrophil Ratio

MON

Monocyte

DBIL

Direct Bilirubin

CREA

Creatinine

SVM

Support Vector Machine

ROC

Receiver Operating Characteristic Curves

AUC

Area Under the Curve

DCA

Decision Curve Analysis

PR

Precision-Recall

SHAP

Shapley Additive Explanation

Author contributions

C.Y. and S.L. designed the study, performed data analysis, and wrote the original manuscript. X.Z. collected and curated clinical data. W.W. developed the methodology and validated results. Y.Z. reviewed pathological classifications. G.Z. supervised the project, acquired funding, and revised the manuscript. All authors reviewed the manuscript. Note: C.Y. and S.L. contributed equally as co-first authors. Note C.Y. and S.L. contributed equally as co-first authors.

Funding

This study was sponsored by the National High Performance Medical Device Innovation Center Project (No. NMED2025KF-01-005) and the Application of Intelligent Mutual Recognition of Inspection Results Project (No. JYHRXZ2025B06).

Data availability

The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

This study was conducted in accordance with the Declaration of Helsinki and was approved by the Ethics Committee of Peking University People’s Hospital Qingdao Hospital (Approval No.: 2025PHQDB022-01). The need for informed consent was waived by the Ethics Committee of Peking University People’s Hospital Qingdao Hospital because this was a retrospective analysis of anonymized data.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Congbo Yue and Shichao Liu contributed equally to this work.

References

  • 1.Soergel P, Dahl GF, Onsrud M, Hillemanns P. Photodynamic therapy of cervical intraepithelial neoplasia 1–3 and human papilloma virus (HMV) infection with Methylaminolevulinate and hexaminolevulinate–a double-blind, dose-finding study. Lasers Surg Med. 2012;44(6):468–74. [DOI] [PubMed] [Google Scholar]
  • 2.Saslow D, Solomon D, Lawson HW, Killackey M, Kulasingam SL, Cain J, Garcia FAR, Moriarty AT, Waxman AG, Wilbur DC, et al. American cancer society, American society for colposcopy and cervical pathology, and American society for clinical pathology screening guidelines for the prevention and early detection of cervical cancer. Cancer J Clin. 2012;62(3):147–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J Clin. 2021;71(3):209–49. [DOI] [PubMed] [Google Scholar]
  • 4.Petersen Z, Jaca A, Ginindza TG, Maseko G, Takatshana S, Ndlovu P, Zondi N, Zungu N, Varghese C, Hunting G, et al. Barriers to uptake of cervical cancer screening services in low-and-middle-income countries: a systematic review. BMC Womens Health. 2022;22(1):486. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Mao C, Balasubramanian A, Koutsky LA. Should liquid-based cytology be repeated at the time of colposcopy? J Low Genit Tract Dis. 2005;9(2):82–8. [DOI] [PubMed] [Google Scholar]
  • 6.Zhang L, Tian P, Li B, Xu L, Qiu L, Bi Z, Chen L, Sui L. Risk-stratified management of cervical high-grade squamous intraepithelial lesion based on machine learning. J Med Virol. 2024;96(10):e70016. [DOI] [PubMed] [Google Scholar]
  • 7.Arbyn M, Weiderpass E, Bruni L, de Sanjosé S, Saraiya M, Ferlay J, Bray F. Estimates of incidence and mortality of cervical cancer in 2018: a worldwide analysis. Lancet Global Health. 2020;8(2):e191–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Castle PE, Cremer M. Human papillomavirus testing in cervical cancer screening. Obstet Gynecol Clin North Am. 2013;40(2):377–90. [DOI] [PubMed] [Google Scholar]
  • 9.Steinauer JE, Turk JK, Pomerantz T, Simonson K, Learman LA, Landy U. Abortion training in US obstetrics and gynecology residency programs. Am J Obstet Gynecol. 2018;219(1):e8681–6. [DOI] [PubMed]
  • 10.Stoler MH, Schiffman M. Group ftASCoUSL-gSILTS: interobserver reproducibility of cervical cytologic and histologic interpretationsrealistic estimates from the ASCUS-LSIL triage study. JAMA. 2001;285(11):1500–5. [DOI] [PubMed] [Google Scholar]
  • 11.Salehnasab Z, Mousavizadeh A, Ghalamfarsa G, Garavand A, Salehnasab C. Predictive modeling of COVID-19 hospitalization using twenty Machine Learning classification algorithms on cohort data. Front Health Inf. 2023;12.
  • 12.Ghaderzadeh M, Salehnasab C. Filter-Based feature selection for type II diabetes prediction. J Clin Care Skills. 2025;6(3):121–8. [Google Scholar]
  • 13.Rafie Z, Talab MS, Koor BEZ, Garavand A, Salehnasab C, Ghaderzadeh M. Leveraging XGBoost and explainable AI for accurate prediction of type 2 diabetes. BMC Public Health. 2025;25(1):3688. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Esteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, van der Laak JAWM, van Ginneken B. Sánchez CI: a survey on deep learning in medical image analysis. Medical Image Analysis. 2017;42:60–88. [DOI] [PubMed]
  • 16.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44–56. [DOI] [PubMed] [Google Scholar]
  • 17.Rajpurkar P, Irvin J, Ball RL, Zhu K, Yang B, Mehta H, Duan T, Ding D, Bagul A, Langlotz CP, et al. Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists. PLoS Med. 2018;15(11):e1002686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Ehteshami Bejnordi B, Veta M, van Johannes P, van Ginneken B, Karssemeijer N, Litjens G, van der Laak JAWM. Consortium atc: diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. JAMA. 2017;318(22):2199–210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Mantula F, Toefy Y, Sewram V. Barriers to cervical cancer screening in africa: a systematic review. BMC Public Health. 2024;24(1):525. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Taherikia S, Firouzjaeian Galougah A, Ahmadi Lashkenari M, Lotfi S. Perspectives of gynecologic oncologists on AI for women’s cancers in Low-Resource settings: A survey in Iran. Infosci Trends. 2025;2(8):1–10. [Google Scholar]
  • 21.Rahimi M, Akbari A, Asadi F, Emami H. Cervical cancer survival prediction by machine learning algorithms: a systematic review. BMC Cancer. 2023;23(1):341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Liu Y, Chen PC, Krause J, Peng L. How to read articles that use machine learning: users’ guides to the medical literature. JAMA. 2019;322(18):1806–16. [DOI] [PubMed] [Google Scholar]
  • 23.Castle PE, Sideri M, Jeronimo J, Solomon D, Schiffman M. Risk assessment to guide the prevention of cervical cancer. Am J Obstet Gynecol. 2007;197(4):e356351–6. [DOI] [PMC free article] [PubMed]
  • 24.Massad LS, Einstein MH, Huh WK, Katki HA, Kinney WK, Schiffman M, Solomon D, Wentzensen N, Lawson HW, for the ACGC.: 2012 updated consensus guidelines for the management of abnormal cervical cancer screening tests and cancer precursors. J Lower Genit Tract Dis. 2013;17. [DOI] [PubMed]
  • 25.Noble WS. What is a support vector machine? Nat Biotechnol. 2006;24(12):1565–7. [DOI] [PubMed] [Google Scholar]
  • 26.Huang S, Cai N, Pacheco PP, Narrandes S, Wang Y, Xu W. Applications of support vector machine (SVM) learning in cancer genomics. Cancer Genomics Proteom. 2018;15(1):41–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Lundberg SM, Lee S-I. A Unified approach to interpreting model predictions. In: Neural Information Processing Systems: 2017.
  • 28.Arbyn M, Sankaranarayanan R, Muwonge R, Keita N, Dolo A, Mbalawa CG, Nouhou H, Sakande B, Wesley R, Somanathan T, et al. Pooled analysis of the accuracy of five cervical cancer screening tests assessed in eleven studies in Africa and India. Int J Cancer. 2008;123(1):153–60. [DOI] [PubMed] [Google Scholar]
  • 29.Knight K, Wade S, Balducci L. Prevalence and outcomes of anemia in cancer: a systematic review of the literature. Am J Med. 2004;116(Suppl 7):s11–26. [DOI] [PubMed] [Google Scholar]
  • 30.Mantovani A, Allavena P, Sica A, Balkwill F. Cancer-related inflammation. Nature. 2008;454(7203):436–44. [DOI] [PubMed] [Google Scholar]
  • 31.Chen Z, Zhao B. The role of tumor-associated macrophages in HPV induced cervical cancer. Front Immunol. 2025;16:1586806. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Yang Y, Yang Y, Yang J, Zhao X, Wei X. Tumor microenvironment in ovarian cancer: function and therapeutic strategy. Front Cell Dev Biol. 2020;8:758. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hanahan D, Weinberg RA. Hallmarks of cancer: the next generation. Cell. 2011;144(5):646–74. [DOI] [PubMed] [Google Scholar]
  • 34.Wagner KH, Wallner M, Mölzer C, Gazzin S, Bulmer AC, Tiribelli C, Vitek L. Looking to the horizon: the role of bilirubin in the development and prevention of age-related chronic diseases. Clin Sci (Lond). 2015;129(1):1–25. [DOI] [PubMed] [Google Scholar]
  • 35.Yuan C, Yao Y, Cheng B, Cheng Y, Li Y, Li Y, Liu X, Cheng X, Xie X, Wu J, et al. The application of deep learning based diagnostic system to cervical squamous intraepithelial lesions recognition in colposcopy images. Sci Rep. 2020;10(1):11639. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Li D, Wang Z, Liu Y, Zhou M, Xia B, Zhang L, Chen K, Zeng Y. Assessing the risk of high-grade squamous intraepithelial lesions (HSIL+) in women with LSIL biopsies: a machine learning-based study. Infect Agents Cancer. 2024;19(1):61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Massad LS, Einstein MH, Huh WK, Katki HA, Kinney WK, Schiffman M, Solomon D, Wentzensen N, Lawson HW. 2012 updated consensus guidelines for the management of abnormal cervical cancer screening tests and cancer precursors. J Low Genit Tract Dis. 2013;17(5 Suppl 1):S1–27. [DOI] [PubMed] [Google Scholar]
  • 38.Castle PE, Sideri M, Jeronimo J, Solomon D, Schiffman M. Risk assessment to guide the prevention of cervical cancer. Am J Obstet Gynecol. 2007;197(4):e356351–356. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Arbyn M, Weiderpass E, Bruni L, de Sanjosé S, Saraiya M, Ferlay J, Bray F. Estimates of incidence and mortality of cervical cancer in 2018: a worldwide analysis. Lancet Glob Health. 2020;8(2):e191–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Saslow D, Solomon D, Lawson HW, Killackey M, Kulasingam SL, Cain J, Garcia FA, Moriarty AT, Waxman AG, Wilbur DC, et al. American cancer society, American society for colposcopy and cervical pathology, and American society for clinical pathology screening guidelines for the prevention and early detection of cervical cancer. Am J Clin Pathol. 2012;137(4):516–42. [DOI] [PubMed] [Google Scholar]
  • 41.Rostami G, Hosseini Berneti SH, Habibzadeh N, Bazir M. From colon to uterus: potential of YOLOv7 for Real-Time polyp detection in hysteroscopy. Infosci Trends. 2025;2(4):48–57. [Google Scholar]
  • 42.Mohammad-Rahimi H, Sohrabniya F, Ourang SA, Dianat O, Aminoshariae A, Nagendrababu V, Dummer PMH, Duncan HF, Nosrat A. Artificial intelligence in endodontics: data preparation, clinical applications, ethical considerations, limitations, and future directions. Int Endod J. 2024;57(11):1566–95. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (10.7KB, xlsx)

Data Availability Statement

The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES