Skip to main content
BMC Cancer logoLink to BMC Cancer
. 2026 Jun 30;26:1113. doi: 10.1186/s12885-026-16427-y

Artificial intelligence for cervical cancer screening and diagnosis using Pap smear images: a systematic review

Aynaz Esmailzadeh 1, Asma Rashki Kemmak 2, Alireza Rasoulian 1, Fatemeh Sadat Alizadeh Tabatabaei 3, Mohammad Reza Mazaheri Habibi 1,✉
PMCID: PMC13587484  PMID: 42380812

Abstract

Background

Cervical cancer continues to be the second most prevalent type of cancer in women globally, especially in the less developed areas. Among several screening methods Pap test, more popularly known as the Papanicolaou test or Pap smear, is one of the most efficient ones for the early detection of this type of cancer. AI has recently made possible the large-scale screening of the whole process. This whole thing has been done in order to increase the early detection rates, which is the long-term aim of reducing the number of cases and deaths that are due to cervical cancer.

Objective

This systematic review aimed to investigate the role of artificial intelligence in the diagnosis of cervical cancer based on Pap smear results.

Methods

A comprehensive search was conducted in PubMed, Web of Science, Scopus, Cochrane Library, and Google Scholar from database inception to January 2025, with the final search update performed on January 30, 2025. The search strategy was designed using relevant keywords and their synonyms related to “artificial intelligence,” “diagnosis,” and “cervical cancer.”This review included only the English-language studies that had investigated the application of AI in the diagnosis of cervical cancer using Pap test data. The titles and abstracts were initially reviewed by two independent reviewers, and subsequently, full-text assessment was carried out. Data extraction followed the use of standardized forms that collected information on study title, country, number of participants, study purposes, AI technique, error rate, accuracy, and performance outcomes.

Results

The initial search identified 844 studies, of which 22 met the inclusion criteria and were included in the final analysis. Most studies reported that AI-based algorithms improved the accuracy and efficiency of cervical cancer detection using Pap smear images. Deep learning and machine learning approaches demonstrated high diagnostic performance, with several studies reporting accuracy rates above 90%.

Conclusion

AI-based approaches show considerable potential for improving the accuracy and timeliness of cervical cancer diagnosis using Pap smear analysis. However, further high-quality studies are required to validate these tools and support their integration into clinical practice.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1186/s12885-026-16427-y.

Keywords: Artificial intelligence, Diagnosis, Screening, Cervical cancer, Pap Smear Images

Introduction

Cervical cancer remains one of the most common cancers among women all over the world. In the year 2020 alone, there were already around 604,000 new cases and 342,000 death reports [1]. GLOBOCAN indicates that cervical cancer was responsible for nearly 3.2% of the total and around 8% of the death cases among women in that year [2, 3]. What is shocking is that almost 94% of the cervical cancer deaths occur in the less developed countries, which also reiterates the inequitable distribution of health care resources as well as preventive measures between the rich and the poor [3]. According to the World Health Organization (WHO) and GLOBOCAN 2018 estimates, cervical cancer remains one of the most common cancers among women worldwide, particularly in low- and middle-income countries [4]. In Iran, studies denote an incidence rate of 5.4 per 100,000 women along with a death rate of approximately 9 per 100,000 among those affected [5]. Cervical cancer is considered one of the most preventable and treatable cancers when identified at an early stage through regular screening and appropriate management [4]. Therefore, it not only the HPV vaccination for girls that serves as the core of primary prevention but also Pap smear screening is still advised for women who are sexually active [6]. Artificial Intelligence (AI), a branch of computer science, is the application of specific algorithms to perform tasks that are similar to human actions. Machine Learning (ML), which is one of the AI techniques, is the technology behind making algorithms learn from the data and improving their performance through repeated trials without human coding [7].

Several previous studies and systematic reviews have investigated the application of artificial intelligence in cervical cancer screening using Pap smear images [8–11]. These studies indicate that deep learning approaches, particularly convolutional neural networks (CNNs), are the most widely used methods for automated cytology analysis and can achieve high diagnostic performance in cervical cell classification [9, 10]. For instance, Jiang et al. reviewed more than 120 studies and highlighted the rapid expansion of deep learning in computational cytology, including classification, detection, and segmentation tasks [11]. Similarly, systematic reviews have reported that AI-based systems can achieve performance comparable to expert cytopathologists while reducing inter- and intra-observer variability [8, 9]. However, despite these promising findings, existing literature consistently reports limitations such as small and imbalanced datasets, lack of external validation, and limited generalizability across populations [8–11].

Artificial intelligence has increasingly become integrated into medical diagnostics, offering improved accuracy and efficiency across various clinical workflows, particularly in oncology and cervical cancer screening [12–14]. In recent years, deep learning has emerged as an end-to-end solution for many biomedical image analysis tasks [15–19]. Cervical cancer screening currently relies on multiple clinical techniques, including cytology (Pap smears), HPV testing, and colposcopy, each with its own advantages and limitations [20, 21]. Despite the promising role of AI in improving screening performance, its clinical adoption remains in the early stages [22]. Therefore, further development and validation are required to support its integration into large-scale screening programs aimed at reducing cervical cancer incidence and mortality [23].

Despite the growing body of literature on artificial intelligence in cervical cancer screening, the evidence remains scattered across different methodologies, datasets, and evaluation metrics. Previous reviews in this field have generally adopted a broad scope, often covering digital pathology or multiple imaging modalities rather than focusing exclusively on Pap smear–based AI applications. In addition, recent advances in deep learning techniques have not been fully and systematically integrated into a unified synthesis.

Therefore, this study systematically reviews and critically evaluates the existing literature on artificial intelligence applications in cervical cancer diagnosis using Pap smear images. This review specifically focuses on AI-based approaches applied to Pap smear cytology and provides an updated synthesis of recent studies regarding methodologies, datasets, and diagnostic performance.

This study aims to assess the effectiveness of artificial intelligence techniques in improving the accuracy and efficiency of cervical cancer screening using Pap smear images. In addition, it seeks to identify current limitations and research gaps in the existing literature and to provide insights that may support the future development and clinical integration of AI-based diagnostic systems.

Methods

Study design and protocol registration

The review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines [24, 25]. The study protocol was prospectively registered in the International PROSPERO database (Registration ID: CRD420251178458). All stages of the review were performed in accordance with established international standards to ensure transparency, reproducibility, and methodological rigor.

Research question and objectives

The principal goal of this study was to systematically review the existing evidence regarding the diagnostic performance of artificial intelligence–based algorithms, including machine learning (ML) and deep learning (DL), applied to Pap smear images for cervical cancer detection.

The secondary objectives were to analyze research trends, critically evaluate study methodologies by identifying strengths and limitations, and provide recommendations for the future integration of AI-based cervical cancer screening tools.

Inclusion and exclusion conditions

Inclusion conditions

The studies were accepted only if they satisfied the following criteria:

  1. Diagnosis confirmation by the use of a valid reference standard, e.g., pathology or cytology results.

  2. Quantitative metrics of diagnostic performance such as sensitivity, specificity, accuracy, area under the curve (AUC), recall, or F1-score were reported.

  3. Only studies published in the English language were considered for inclusion.

Exclusion conditions

The studies were rejected if any of the following situations were identified:

  1. Non-original publications, including conference abstracts, book chapters, review articles, editorials, and letters to the editor, were excluded.

  2. Studies not relevant to the aim of the review or not focusing on AI applications in cervical cancer detection using Pap smear images were excluded.

  3. Articles without full-text availability were excluded.

Search methodology

A comprehensive literature search without time restrictions was conducted in PubMed, Web of Science, Scopus, Cochrane Library, and Google Scholar from database inception until January 2025. The final search update was performed on January 30, 2025. The search queries comprised the words and MeSH terms “artificial intelligence,” “diagnosis,” “cervical cancer,” and “Pap smear.” (Table 1). For Google Scholar, only the first 15–20 pages of results were screened, as relevant studies are typically indexed within the initial pages. The entire search process, including query strings, dates of execution, and retrieved results, was recorded to ensure reproducibility.

Table 1.

Search strategy for each database

#1 PubMed database search approach Results
((((((“Artificial Intelligence“[Mesh]) OR “Machine Learning“[Mesh]) OR “Deep Learning“[Mesh]) OR (((“Artificial Intelligence“[Title/Abstract]) OR (“Machine Learning“[Title/Abstract])) OR (“Deep Learning“[Title/Abstract]))) AND ((“Papanicolaou Test“[Mesh]) OR (“Papanicolaou Test“[Title/Abstract]))) AND (((“Diagnosis“[Mesh]) OR “Early Diagnosis“[Mesh]) OR ((“Diagnosis“[Title/Abstract]) OR (“Early Diagnosis“[Title/Abstract])))) AND ((“Uterine Cervical Neoplasms“[Mesh]) OR (“Uterine Cervical Neoplasms“[Title/Abstract])) 83
#2 Scopus database search approach Results
(TITLE-ABS-KEY (“artificial intelligence” OR “machine learning” OR “deep learning” OR “neural network*” OR “computer-aided diagnosis” OR “CAD”)) AND (TITLE-ABS-KEY (“cervical cancer” OR “cervical neoplasia” OR “cervical carcinoma” OR “uterine cervical neoplasia”)) AND (TITLE-ABS-KEY (“Pap smear” OR “Pap test” OR “cervical cytology” OR cytology)) AND (TITLE-ABS-KEY (diagnosis OR detection OR screening)) 81
#3 Google Scholar database search approach Results
(“Cervical cancer” OR “Cervical carcinoma”) AND (“Artificial Intelligence” OR “Machine Learning” OR “Deep Learning”) AND (“Pap smear” OR “Papanicolaou test”) AND (“diagnosis”) 6670
#4 Web of Science database search approach Results
TS=(“artificial intelligence” OR “machine learning” OR “deep learning” OR “neural network*” OR “computer-aided diagnosis” OR CAD) AND TS=(“cervical cancer” OR “cervical neoplasia” OR “cervical carcinoma” OR “uterine cervical neoplasia”) AND TS=(“Pap smear” OR “Pap test” OR “cervical cytology” OR cytology) AND TS=(diagnosis OR detection OR screening) 37
#5 Cochrane Library database search approach Results
(“Cervical cancer” OR “Cervical carcinoma”) AND (“Artificial Intelligence” OR “Machine Learning” OR “Deep Learning”) AND (“Pap smear” OR “Papanicolaou test”) AND (“diagnosis”) 0

Study selection

Process The process of selecting studies started with the removal of duplicate records through the use of EndNote X21. The titles and abstracts of the resulting studies were then screened independently by two reviewers following the eligibility criteria. Full-text articles of potentially eligible studies were retrieved and assessed for final inclusion. A standardized data extraction form was developed that contained study objectives, publication year, first author, country, dataset size, AI algorithms or techniques employed, and model characteristics as its main fields. The data extraction form was created using Microsoft Excel to ensure systematic and consistent data collection. Data extraction was performed independently by two reviewers using a standardized form. Any discrepancies were resolved through discussion and consensus. A third reviewer was involved when agreement could not be reached.

Quality assessment and risk of bias

The QUADAS-2 tool [26] was utilized to assess the methodological quality and potential bias of the selected studies. Four significant areas of concern were examined: selection of patients, the index test, reference standard, and flow/timing. The risk of bias (low, high, or uncertain) and concerns about applicability were noted for each domain. The evaluations were performed by two reviewers independently, and any discrepancies were settled through consensus. The findings were presented in tabular form and illustrated with the ROBVIS tool (Fig. 1 and Fig. 2).

Fig. 1.

Fig. 1

ROBVIS tool to assess the methodological quality and potential bias of the selected studies

Fig. 2.

Fig. 2

Diagram of article selection based on flowchart PRISMA

Data synthesis

Due to the heterogeneity of AI algorithms, imaging modalities, dataset sizes, and reporting metrics, data synthesis was performed qualitatively. The results were organized into three main categories: types of algorithms, data characteristics, and performance measures. Comparative tables and narrative synthesis were used to identify key patterns, research gaps, and the most promising methodological approaches in the field. No statistical meta-analysis was conducted due to variability in study designs and outcome reporting. Where studies reported incomplete or heterogeneous data, only the available reported metrics were extracted. No imputation of missing data was performed. When necessary, performance measures were converted into percentage form to ensure consistency across studies.

The primary outcomes of this review were the diagnostic performance measures of AI algorithms for cervical cancer detection using Pap smear images, including accuracy, sensitivity, specificity, AUC, recall, and F1-score.

Ethical considerations

Since this review was based exclusively on publicly available data, it was not necessary to seek ethical approval or informed consent.

Results

Study selection

A total of 6,841 records were initially identified. After title and abstract screening, 198 studies were selected for full-text review. At the full-text stage, 176 studies were excluded based on predefined eligibility criteria due to lack of relevance to Pap smear-based AI applications, absence of reported diagnostic performance metrics, non-original study designs (e.g., reviews, conference abstracts, and editorials), unavailability of full text, or use of non–AI-based methods. Studies that initially appeared eligible after abstract screening were subsequently excluded during full-text assessment following detailed evaluation against the inclusion criteria. Finally, 22 studies met the eligibility criteria and were included in the qualitative synthesis. The study selection process is illustrated in the PRISMA flow diagram (Fig. 3).

Fig. 3.

Fig. 3

Domain-based risk of bias assessment (QUADAS-2) of studies included in cervical cancer diagnostic accuracy analysis

Although 26 studies underwent risk-of-bias assessment, 4 studies were excluded at the data synthesis stage due to insufficient or incomplete outcome data required for synthesis, resulting in 22 studies included in the final analysis.

Study characteristics (Appendix 1)

This systematic review included studies published between 2014 and 2025, covering research conducted across multiple geographical regions, including India [27–36], the United States [37–39], Thailand [40, 41], Jordan [42, 43], Iran [44], Malaysia [45], Peru [46], Saudi Arabia [47], and Uganda [48]. The distribution of studies reflects a growing global interest in applying artificial intelligence to cervical cytology analysis, particularly in both developed and developing healthcare settings.

Across the 22 included studies, deep learning (DL) approaches accounted for a substantial proportion of the literature. Approximately 15 studies specifically focused on DL-based methods [27–29, 31, 32, 34, 37–40, 42, 43, 46, 48]. Among these, ResNet-based architectures were the most frequently used, appearing in several studies [28, 31, 32, 34, 36, 38, 45], highlighting their strong adaptability in medical image classification tasks.

Machine learning (ML) techniques were applied in 8 studies [30, 33, 35, 41, 43, 44, 47, 48], with Random Forest emerging as the most commonly used algorithm [30, 33, 35, 43]. Only one study implemented a hybrid DL–ML framework [43], indicating that integrative approaches remain relatively underexplored in this research area. Overall, most studies aimed to develop automated AI-based systems for the detection and classification of cervical cancer using Pap smear images.

Data characteristics and Pap smear types

The datasets used in the included studies were primarily based on Pap smear cytology images. Conventional Pap smear samples were the most frequently used data source, while liquid-based cytology (LBC) datasets were also reported in some of the included studies.

The dataset sizes varied substantially, ranging from approximately 170 to 6,000 images. This variation introduces significant heterogeneity in model training and evaluation, which may affect performance comparability across studies. In particular, widely used benchmark datasets such as Herlev, SIPaKMeD, and LBC-based datasets differ in imaging conditions, staining protocols, and class distribution, which should be considered when interpreting reported results.

Performance of deep learning models (Appendix 2)

Deep learning models demonstrated consistently strong performance in cervical cytology classification tasks, with reported accuracy values generally ranging from approximately 90% to 99%, depending on architecture, dataset, and classification complexity.

In study [31], six deep learning architectures (AlexNet, VGG16, VGG19, ResNet50, ResNet101, and GoogLeNet) were evaluated for multiclass cervical lesion classification using conventional, LBC, and Herlev datasets. GoogLeNet and ResNet101 achieved the highest performance, exceeding 90% accuracy.

Chauhan et al. [28] proposed a hybrid HDFCN model combining VGG16, ResNet152, and DenseNet169, achieving 99.29% binary accuracy and 97.45% multiclass accuracy using SIPaKMeD and LBC datasets.

Zammataro et al. [39] developed the CINNAMON-GUI tool based on CNN architectures, where Model B outperformed Model A with a validation accuracy of 0.95.

Kaur et al. [36] evaluated 16 pretrained models and found that ResNet50 achieved 95% accuracy on the Herlev dataset, while VGG16 reached 99.95% on SIPaKMeD. DenseNet121 also achieved strong performance with 97.65% accuracy in three-class classification.

Sompawong et al. [40] applied Mask R-CNN, achieving 91.7% precision, sensitivity, and specificity for Pap smear screening, and 89.8% precision for single-cell classification.

Kalbhor et al. [32] proposed a hybrid deep learning–fuzzy neural network model, where ResNet50 achieved 95.33% accuracy. Parkash et al. [34] introduced the EWR-DNN model, achieving 98.9% accuracy, outperforming conventional VGG, ResNet, and AlexNet architectures.

Overall, DL models consistently demonstrated superior performance compared to traditional approaches; however, their results varied depending on dataset characteristics and evaluation protocols.

Performance of machine learning models (Appendix 2)

Machine learning models generally showed more variability in performance compared to deep learning approaches, although several hybrid and optimized frameworks achieved strong results.

In study [35], decision-tree-based models, ensemble methods, and shallow neural networks were compared with a hybrid ensemble approach, which achieved 96% binary accuracy and 78% multiclass accuracy.

Win et al. [41] reported accuracies of 98.27% (binary) and 94.09% (five-class) using a hybrid Bagging–Random Forest model on SIPaKMeD and Herlev datasets.

Keymasi et al. [44] compared SVM, KNN, and MLP classifiers, where MLP achieved the best performance with 97.83% accuracy.

Waly et al. [47] proposed an IDCNN-CDC framework combining SqueezeNet and Weighted ELM, achieving 97.96% accuracy, along with Precision (97.84%), Recall (98.2%), and F1-score (97.81%).

Another hybrid approach [43] combined features extracted from pretrained CNNs (AlexNet, DarkNet19, NasNet) with PCA-based dimensionality reduction and optimization techniques (PSO, ALO), followed by SVM and Random Forest classification. In this study, PSO–SVM achieved the highest accuracy of 99.5%, outperforming Random Forest (98.9%).

Overall, the findings of this review are consistent with previous studies reporting the superiority of deep learning approaches over traditional machine learning methods in cervical cytology analysis. However, hybrid and ensemble machine learning models also demonstrated competitive performance in several datasets, suggesting that model performance is highly dependent on dataset quality, feature selection, and preprocessing strategies. These results further support the growing body of evidence emphasizing the potential of AI-based systems for improving cervical cancer screening accuracy.

Deep learning models generally achieved higher and more stable performance, particularly ResNet- and VGG-based architectures, with reported accuracies reaching up to approximately 99%. In addition, hybrid deep learning approaches (e.g., CNN combined with feature weighting or ensemble strategies) often outperformed standalone models in several datasets. In contrast, machine learning models showed higher variability in performance, which was largely influenced by feature engineering and preprocessing techniques. Ensemble and hybrid machine learning methods, including PSO-SVM and Random Forest-based models, were reported in several included studies. Performance results varied across studies using different datasets, including SIPaKMeD, Herlev, and LBC datasets.

Discussion

The findings of this systematic review indicate that artificial intelligence (AI), particularly convolutional neural networks (CNNs) and hybrid learning frameworks, has significantly advanced the field of Pap smear-based cervical cancer screening. Across the included studies [27–48], there is a consistent trend toward improved diagnostic performance through deep learning-based architectures, ensemble methods, and optimization-enhanced hybrid systems.

Overall, deep learning approaches demonstrated superior and more stable performance compared to traditional machine learning methods, particularly in multiclass classification tasks. However, performance variability across studies suggests that model effectiveness is strongly influenced by dataset characteristics, preprocessing strategies, and feature representation methods.

Evolution of AI-based cervical cytology systems

The progression of AI applications in cervical cytology shows a clear evolution from classical machine learning pipelines toward deep learning and hybrid architectures. Early studies such as Sarwar et al. [35] demonstrated that ensemble-based classification strategies could improve abnormal cell detection accuracy by combining multiple classifiers.

Similarly, Gupta et al. [30] highlighted the potential of automated Pap smear analysis to achieve diagnostic performance comparable to expert pathologists, supporting the feasibility of AI-assisted screening in clinical workflows, particularly in low-resource settings.

Subsequent developments expanded toward deep learning architectures, where CNN-based models increasingly became the dominant approach. Studies such as Hussain et al. [31] and Chauhan et al. [28] demonstrated that deeper architectures and hybrid feature fusion strategies significantly improved classification accuracy in multiclass cervical lesion detection tasks.

Deep learning architectures and performance trends

Across the reviewed literature, ResNet-based architectures were the most frequently adopted and consistently achieved strong performance [28, 31, 32, 34, 36, 38, 45]. Other architectures such as VGG, DenseNet, and EfficientNet also demonstrated competitive results depending on dataset and task complexity.

For example, Kaur et al. [36] reported that EfficientNet-B4 achieved approximately 99% accuracy, while ResNet50 and VGG16 also performed strongly across different datasets. Similarly, Tan et al. [45] identified ResNet101 as a balanced model in terms of sensitivity and specificity.

Hybrid deep learning approaches further improved performance in several studies. For instance, Chauhan et al. [28] demonstrated that combining multiple CNN architectures improved robustness and classification accuracy, particularly in heterogeneous datasets such as SIPaKMeD and LBC.

In addition, attention-based enhancements [27] improved model robustness in noisy and low-quality imaging conditions, highlighting the importance of feature refinement mechanisms in cytology image analysis.

Role of machine learning and hybrid models

Although deep learning models dominated performance benchmarks, machine learning approaches still demonstrated competitive results when combined with advanced feature engineering techniques.

For example, Win et al. [41] showed that ensemble-based Bagging and Random Forest models achieved strong performance in both binary and multiclass classification tasks. Similarly, Keymasi et al. [44] reported that MLP outperformed traditional classifiers such as SVM and KNN.

Hybrid approaches integrating feature extraction from deep networks with classical classifiers also showed promising results. Alsalatie et al. [42] and Alsalatie [43] demonstrated that optimization techniques such as PSO and ACO, when combined with CNN features and SVM classifiers, improved classification stability and accuracy.

These findings suggest that machine learning performance is highly dependent on feature quality and optimization strategies, and that hybridization can partially bridge the gap between ML and DL approaches.

Clinical translation and regulatory progress

Beyond algorithmic performance, several studies emphasized the transition of AI systems toward clinical application. For instance, Elishaev et al. [46] reported the development of the Hologic Genius Digital Diagnostics system, which received FDA approval and demonstrated a reduction in false negatives in clinical screening workflows.

Similarly, platforms such as CINNAMON-GUI [39] demonstrate efforts toward improving interpretability and usability of AI systems in real clinical environments. These developments indicate a gradual shift from experimental models toward clinically deployable diagnostic tools.

Dataset heterogeneity and methodological considerations

The performance variability observed across studies can be partly attributed to differences in dataset composition, imaging conditions, and annotation protocols. Commonly used datasets such as Herlev, SIPaKMeD, and LBC differ significantly in staining techniques, resolution, and class distribution, which limits direct comparison of reported accuracy values.

In addition, variations in preprocessing methods, feature selection strategies, and evaluation protocols further contribute to inconsistencies across studies. These factors highlight the need for standardized benchmarking frameworks in future research.

Methodological quality (QUADAS-2 assessment)

Methodological quality assessment using the QUADAS-2 tool indicated that most studies demonstrated a generally acceptable level of quality. However, concerns were identified primarily in the reference standard domain, where approximately 40% of studies showed moderate concern and more than one-third exhibited a high risk of bias (Fig. 3).

In contrast, patient selection bias was relatively low in most studies, although around 25% still showed moderate to high concern. The index test and flow/timing domains were predominantly rated as low risk.

Overall, while the methodological quality is acceptable, the observed bias in reference standards suggests that caution is needed when interpreting reported performance metrics.

Strengths

This review provides a comprehensive synthesis of 22 studies on artificial intelligence and deep learning applications in cervical cancer screening, offering an up-to-date overview of the current state of the field. By systematically analyzing diverse AI approaches, the study highlights key methodological trends and performance patterns across different models. The review also discusses clinically relevant developments, including the translational potential of AI-based systems in medical practice. In addition, it examines hybrid and ensemble approaches that have shown improved robustness and diagnostic performance, while also considering key limitations related to data quality and algorithmic variability.

Limitation

The findings of this review may have limited generalizability due to the relatively small sample sizes, limited population diversity, and variability in data quality across the included studies. In addition, the heterogeneity of datasets and methodologies prevented the conduct of a quantitative meta-analysis, resulting in a primarily qualitative synthesis. Furthermore, important implementation-related factors such as model explainability, clinician trust, and interoperability with clinical systems were insufficiently addressed in most of the included studies. Finally, the lack of standardized benchmark datasets specific to diverse populations may introduce potential bias and limit the robustness and comparability of the reported models.

Conclusion

The systematic review of eligible studies finally included 22 revealed that artificial intelligence algorithms, especially convolutional neural networks (CNNs) and hybrid deep learning models, provide better and more reliable performance for detecting cervical cancer from Pap smear images compared to conventional diagnostic methods. Most of the studies reported diagnostic accuracy, sensitivity, and specificity greater than 95% and this emphasizes AI’s role in early detection. This major development can be attributed to deep learning’s ability to distinguish and extract different morphological and textural features within the cytological images. Additionally, the combination of hybrid architectures and optimization algorithms has improved the robustness and generalizability of the models. AI diagnostic systems not only lighten the load of pathologists but they also speed up and increase the accuracy of large-scale cervical cancer screening which in turn might decrease the global mortality rate. On the flip side, the aforementioned challenges still exist, primarily in the areas of data standardization, quality control, and algorithmic transparency which are of utmost importance for bridging the gap between clinical and technological practices. Future researchers will have to come up with multi-center annotated datasets, interpretable (Explainable AI) frameworks, and regulatory validation as their prime focus in the fight for trustworthy and equitable deployment of AI-based cervical cancer screening systems.

Supplementary Information

Supplementary Material 1. (32.4KB, docx)
Supplementary Material 2. (32.4KB, docx)
Supplementary Material 3. (18.5KB, docx)
Supplementary Material 4. (18.4KB, docx)
Supplementary Material 5. (268.4KB, docx)

Acknowledgements

Varastegan Institute for Medical Sciences supported this study.

Abbreviations

WHO

World Health Organization

CC

Cervical cancer

AI

Artificial Intelligence

ML

Machine Learning

DL

Deep Learning

AUC

Area under the curve

F1-score

Harmonic Mean of Precision and Recall

SVM

Support Vector Machine

RF

Random Forest

CIN

Cervical intraepithelial neoplasia

LBC

liquid–based cytology

CNNs

Convolutional neural networks

WSIs

Whole–slide images

PSO

Particle swarm optimization

ACO

Ant colony optimization

Authors’ contributions

Aynaz Esmailzadeh wrote the protocol, assisted with the search strategy, and drafted the article. Asma Rashki Kemmak contributed to the article’s first draft, its revision, and the search strategy. Alireza Rasoulian and Fatemeh Sadat Alizadeh Tabatabaei was involved in the management of the study selection, data extraction, and article quality evaluation. Mohammad Reza Mazaheri Habibi wrote the study protocol, contributed to the study design, and critically revised the first draft of the article for important intellectual content.

Funding

There is no funding source.

Data availability

The data that support the findings of this study are available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA: Cancer J Clin. 2021;71:209–49. 10.3322/caac.21660. [DOI] [PubMed] [Google Scholar]
  • 2.Bokulich NA, Łaniewski P, Adamov A, Chase DM, Caporaso JG, Herbst-Kralovetz MM. Multi-omics data integration reveals metabolome as the top predictor of the cervicovaginal microenvironment. PLoS Comput Biol. 2022;18(2):e1009876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Organization WH. Cervical cancer: key facts Geneva: WHO: Geneva: WHO; 2023. Available from: https://www.who.int/news-room/fact-sheets/detail/cervical-cancer.
  • 4.Arbyn M, Weiderpass E, Bruni L, de Sanjosé S, Saraiya M, Ferlay J, Bray F. Estimates of incidence and mortality of cervical cancer in 2018: a worldwide analysis. Lancet Global Health. 2020;8(2):e191–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Graham SV. The human papillomavirus replication cycle, and its links to cancer progression: a comprehensive review. Clin Sci. 2017;131(17):2201–21. [DOI] [PubMed] [Google Scholar]
  • 6.Phoulady HA. Adaptive Region-Based Approaches for Cellular Segmentation of Bright-Field Microscopy Images. Univ. South Florida; 2017.
  • 7.Kuo RY, Harrison C, Curran TA, Jones B, Freethy A, Cussons D, Stewart M, Collins GS, Furniss D. Artificial intelligence in fracture detection: a systematic review and meta-analysis. Radiology. 2022;304(1):50–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.iang H, Zhou Y, Lin Y, Chan RCK, Liu J, Chen H. Deep learning for computational cytology: A survey. arXiv preprint arXiv:2202.05126; 2022. [DOI] [PubMed]
  • 9.da Silva AL, Nicolli AC, Jr. Araujo MLC, et al. Artificial intelligence use in the routine of cervical-vaginal cytology: A systematic review. Surg Experimental Pathol. 2026;9(1):Article22. [Google Scholar]
  • 10.Valles-Coral MA, Pinedo L, Rodríguez C et al. Application of artificial intelligence in cervical cytology: A systematic review of deep learning models, datasets, and reported metrics. Front Big Data. 2025;8. [DOI] [PMC free article] [PubMed]
  • 11.Jiang H, Zhou Y, Lin Y, Chan RCK, Liu J, Chen H. Deep learning for computational cytology: a survey. arXiv preprint arXiv:2202.05126. 2022. [DOI] [PubMed]
  • 12.Habibi MR, JafariMoghadam A, Norouzkhani N, Nazari E, Imani B, Kheirdoust A, Fatemi Aghda SA. The use of neural networks to determine factors affecting the severity and extent of retinopathy in preterm infants. Int J Retina Vitreous. 2025;11(1):30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Kheirdoust A, Barzanouni F, Rasoulian A, Behrouzi F, Esmailzadeh A, Ghaddaripouri K, Mazaheri Habibi MR. Evaluation of Machine Learning Methods Developed for Prediction and Diagnosis of Pneumonia: A systematic review. Health Sci Rep. 2025. [DOI] [PMC free article] [PubMed]
  • 14.Hou X, Shen G, Zhou L, et al. Artificial intelligence in cervical Cancer screening and Diagnosis. Front Oncol. 2022;12:851367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.William W, Ware A, Basaza-Ejiri AH, Obungoloch J. A review of image analysis and machine learning techniques for automated cervical cancer screening from pap-smear images. Comput Methods Programs Biomed. 2018;164:15–22. [DOI] [PubMed] [Google Scholar]
  • 16.Ma J, Song Y, Tian X, Hua Y, Zhang R, Wu J. Survey on deep learning for pulmonary medical imaging. Front Med. 2020;14(4):450–69. [DOI] [PubMed] [Google Scholar]
  • 17.Wang J, Zhu H, Wang SH, Zhang YD. A review of deep learning on medical image analysis. Mob Networks Appl. 2021;26(1):351–80. [Google Scholar]
  • 18.Manhas J, Gupta RK, Roy PP. A review on automated cancer detection in medical images using machine learning and deep learning based computational techniques: challenges and opportunities. Arch Comput Methods Eng. 2022;29(5):2893–933. [Google Scholar]
  • 19.Suganyadevi S, Seethalakshmi V, Balasamy K. A review on deep learning in medical image analysis. Int J Multimedia Inform Retr. 2022;11(1):19–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yousef R, Gupta G, Yousef N, Khari M. A holistic overview of deep learning approach in medical imaging. Multimedia Syst. 2022;28(3):881–914. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Marth C, Landoni F, Mahner S, et al. Cervical cancer: ESMO clinical practice guidelines for diagnosis, treatment and followup. Ann Oncol. 2017;28(suppl4):iv72–83. 8. Schiffman M,. [DOI] [PubMed] [Google Scholar]
  • 22.Kinney WK, Cheung LC, et al. Relative performance of HPV and cytology components of cotesting in cervical Screening. J Natl Cancer Inst. 2018;110(5):501–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Simms KT, Steinberg J, Caruana M, et al. Impact of scaled up human papillomavirus vaccination and cervical screening and the potential for global elimination of cervical cancer in 181 countries, 2020-99: a modelling study. Lancet Oncol. 2019;20(3):394–407. [DOI] [PubMed] [Google Scholar]
  • 24.Oxfoard u. PRISMA 2015. Available from: http://prisma-statement.org.
  • 25.Abutorabi A, et al. Cost-effectiveness of rivaroxaban versus enoxaparin for prevention of venous thromboembolism after knee replacement surgery in Iran. Med J Islamic Repub Iran. 2023;37:20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Lee J, Mulder F, Leeflang M, Wolff R, Whiting P, Bossuyt PM. QUAPAS: an adaptation of the QUADAS-2 tool to assess prognostic accuracy studies. Ann Intern Med. 2022;175(7):1010–8. [DOI] [PubMed] [Google Scholar]
  • 27.Austin R, Parvathi R. CNN based method for classifying cervical cancer cells in pap smear images. Sci Rep. 2025;15(1):23936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Chauhan NK, Singh K, Kumar A, Kolambakar SB. HDFCN: A robust hybrid deep network based on feature concatenation for cervical cancer diagnosis on WSI pap smear slides. Biomed Res Int. 2023;2023(1):4214817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Gangrade J, Kuthiala R, Gangrade S, Singh YP, Solanki RM. A deep ensemble learning approach for squamous cell classification in cervical cancer. Sci Rep. 2025;15(1):7266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Gupta R, Sarwar A, Sharma V. Screening of cervical cancer by artificial intelligence based analysis of digitized papanicolaou-smear images. Int J Contemp Med Res. 2017;4(5):2454–7379. [Google Scholar]
  • 31.Hussain E, Mahanta LB, Das CR, Talukdar RK. A comprehensive study on the multi-class cervical cancer diagnostic prediction on pap smear images using a fusion-based decision from ensemble deep convolutional neural network. Tissue Cell. 2020;65:101347. [DOI] [PubMed] [Google Scholar]
  • 32.Kalbhor M, Shinde S, Popescu DE, Hemanth DJ. Hybridization of deep learning pre-trained models with machine learning classifiers and fuzzy min–max neural network for cervical cancer diagnosis. Diagnostics. 2023;13(7):1363. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kalbhor M, Shinde SV, Jude H. Cervical cancer diagnosis based on cytology pap smear image classification using fractional coefficient and machine learning classifiers. TELKOMNIKA (Telecommunication Comput Electron Control). 2022;20(5):1091–102. [Google Scholar]
  • 34.Kumar P, Chandrasekaran V, Anitha S. Pap Smear Image Classification with Efficient Weight Regularization for Cervical Cancer Diagnosis. Traitement du Signal. 2025;42(3):1503. [Google Scholar]
  • 35.Sarwar A, Sharma V, Gupta R. Hybrid ensemble learning technique for screening of cervical cancer using Papanicolaou smear image analysis. Personalized Med Universe. 2015;4:54–62. [Google Scholar]
  • 36.Kaur H, Sharma R, Kaur J. Comparison of deep transfer learning models for classification of cervical cancer from pap smear images. Sci Rep. 2025;15(1):3945. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Karasu Benyes Y, Welch EC, Singhal A, Ou J, Tripathi A. A comparative analysis of deep learning models for automated cross-preparation diagnosis of multi-cell liquid pap smear images. Diagnostics. 2022;12(8):1838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Sornapudi S, Brown GT, Xue Z, Long R, Allen L, Antani S, editors. Comparing deep learning models for multi-cell classification in liquid-based cervical cytology image. AMIA annual symposium proceedings; 2020. [PMC free article] [PubMed]
  • 39.Zammataro L. CINNAMON-GUI: Revolutionizing Pap Smear Analysis with CNN-Based Digital Pathology Image Classification. F1000Research. 2024;13:897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sompawong N, Mopan J, Pooprasert P, Himakhun W, Suwannarurk K, Ngamvirojcharoen J, et al. editors. Automated pap smear cervical cancer screening using deep learning. 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC); 2019: IEEE. [DOI] [PubMed]
  • 41.Win KP, Kitjaidure Y, Hamamoto K, Myo Aung T. Computer-assisted screening for cervical cancer using digital image processing of pap smear images. Appl Sci. 2020;10(5):1800. [Google Scholar]
  • 42.Alsalatie M, Alquran H, Mustafa WA, Mohd Yacob Y, Ali Alayed A. Analysis of cytology pap smear images based on ensemble deep learning approach. Diagnostics. 2022;12(11):2756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Alsalatie M, Alquran H, Mustafa WA, Zyout AA, Alqudah AM, Kaifi R, Qudsieh S. A new weighted deep learning feature using particle swarm and ant lion optimization for cervical cancer diagnosis on pap smear images. Diagnostics. 2023;13(17):2762. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Mishra V, Aslan S, Asem MM, editors. Theoretical assessment of cervical cancer using machine learning methods based on pap-smear test. 2018 IEEE 9th Annual Information Technology, Electronics and Mobile Communication Conference (IEMCON); 2018: IEEE.
  • 45.Tan SL, Selvachandran G, Ding W, Paramesran R, Kotecha K. Cervical cancer classification from pap smear images using deep convolutional neural network models. Interdisciplinary Sciences: Comput Life Sci. 2024;16(1):16–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Elishaev E, Harinath L, Ye Y, Matsko J, Colaizzi A, Wharton S, et al. Assessment of the efficacy and accuracy of cervical cytology screening with the Hologic Genius Digital Diagnostics System. Cancer Cytopathol. 2025;133(7):e70022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Waly MI, Sikkandar MY, Aboamer MA, Kadry S, Thinnukool O. Optimal deep convolution neural network for cervical cancer diagnosis model. Computers Mater Continua. 2022;70(2).
  • 48.William W, Ware A, Basaza-Ejiri AH, Obungoloch J. A pap-smear analysis tool (PAT) for detection of cervical cancer from pap-smear images. Biomed Eng Online. 2019;18(1):16. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (32.4KB, docx)
Supplementary Material 2. (32.4KB, docx)
Supplementary Material 3. (18.5KB, docx)
Supplementary Material 4. (18.4KB, docx)
Supplementary Material 5. (268.4KB, docx)

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from BMC Cancer are provided here courtesy of BMC

RESOURCES