ABSTRACT
Background and Aims
Differentiating odontogenic keratocyst (OKC) from other radiolucent jaw lesions like ameloblastoma is clinically important but radiographically difficult. Recent advances in artificial intelligence (AI) show promise for enhancing diagnosis using cone‐beam computed tomography (CBCT). This study aims to systematically evaluate and meta‐analyze the diagnostic accuracy of AI models in detecting OKCs on CBCT imaging.
Methods
A systematic review and meta‐analysis was conducted according to PRISMA‐DTA guidelines. Five electronic databases were searched through July 6, 2025. Studies employing AI models for OKC detection using CBCT were included. Methodological quality was assessed using QUADAS‐2. Pooled estimates were computed using a random‐effects model, with heterogeneity evaluated via I2 and meta‐regression. The Eager test and funnel plot were employed to assess publication bias.
Results
Twelve studies were included. AI models demonstrated high diagnostic accuracy, characterized by a pooled sensitivity of 89% (95% CI: 79%–95%) and specificity of 92% (95% CI: 81%–97%), both exceeding 85%, along with a substantial diagnostic odds ratio (87.06) and a robust discriminative ability (AUC = 0.828). Deep learning (DL) models achieved higher sensitivity (91%) than machine learning (ML) models (86%), while ML models showed slightly higher specificity. Heterogeneity was substantial (I2 = 78%–93%). Publication year explained 57.2% of the variability in sensitivity.
Conclusions
AI‐based models, particularly CNN‐based DL architectures, demonstrate clinically relevant diagnostic performance with high sensitivity, specificity, and diagnostic odds ratios, supporting their potential role as adjunctive tools in CBCT‐based differentiation of OKCs from other odontogenic lesions.
Keywords: artificial intelligence, cone‐beam computed tomography, deep learning, diagnostic accuracy, machine learning, meta‐analysis, odontogenic keratocyst, systematic review
1. Introduction
Artificial intelligence (AI) has rapidly transformed various industries, including healthcare, where it plays an increasingly significant role in disease detection, diagnosis, and prediction [1, 2]. Within AI, machine learning (ML) enables systems to learn from data and improve their performance over time without explicit programming [3, 4]. A more specific subset, deep learning (DL), utilizes multi‐layered neural networks, biologically inspired computational models, that can automatically learn complex representations from large datasets without manual feature extraction [5].
Among these, convolutional neural networks (CNNs) are specifically designed for image analysis and have demonstrated strong performance in medical imaging tasks by capturing spatial hierarchies and local patterns within visual data [2, 4, 5]. Some studies showed that CNN‐based models can achieve diagnostic accuracies comparable to those of experienced clinicians in tasks such as cancer detection (with an accuracy of more than 95%) [6, 7, 8] and normal tissue identification on digitized hematoxylin and eosin‐stained slides (with an accuracy ranging from 72% to 96%) [9, 10]. Recent investigations in thoracic imaging and oncologic CT classification similarly report diagnostic accuracies above 90% [11, 12], highlighting the capability of CNN architectures to detect subtle radiographic patterns and support clinical decision‐making.
Within dental imaging, CNNs have demonstrated strong performance across a broad range of diagnostic tasks [13, 14, 15, 16, 17, 18, 19]. DL models have achieved diagnostic accuracies exceeding 94%–95% for dental caries detection and classify oral cancers from histopathological images [5, 17, 18, 19]. Similarly, CNN‐based systems have been successfully applied to the automated identification of periodontitis and periodontal bone loss, with accuracy rates around 95% [14, 15, 16]. Recent studies have explored the application of DL models for the differentiation of odontogenic lesions, including odontogenic keratocyst (OKC) and ameloblastoma (AME) [20, 21, 22, 23].
OKC is a clinically significant developmental odontogenic cyst characterized by aggressive biological behavior, a tendency for infiltrative growth, and a relatively high recurrence rate compared with other jaw cysts [23, 24, 25]. Epidemiological studies estimate that OKCs account for approximately 5%–15% of odontogenic cysts and occur most frequently in the posterior mandible, particularly the molar–ramus region [26, 27]. Unlike many inflammatory odontogenic cysts, OKCs may expand along cancellous bone with limited cortical expansion, allowing lesions to remain asymptomatic and undetected until they reach a considerable size or are discovered incidentally on radiographic examination [26, 27, 28]. Recurrence rates reported in the literature vary widely, ranging from approximately 15% to over 60%, depending on lesion characteristics and treatment modality [29, 30, 31]. Histopathological features such as satellite cysts, daughter cyst formation, and increased epithelial proliferative activity contribute to this recurrence potential. Furthermore, the occurrence of multiple OKCs may be associated with nevoid basal cell carcinoma syndrome, highlighting the broader diagnostic implications of accurate lesion identification and early detection [31, 32, 33, 34].
In oral and maxillofacial radiology, differentiating OKC from other lesions, particularly AME, remains a persistent diagnostic challenge due to overlapping clinical and radiographic features [26, 27, 28]. In recent years, numerous studies have explored the use of AI for this task across multiple imaging modalities, including panoramic radiographs, computed tomography, and histopathological images [4, 35, 36, 37, 38, 39]. These studies have reported promising diagnostic performance, with several DL models achieving high accuracy in binary and multi‐class classification tasks. However, substantial variability exists among studies in terms of imaging modality, dataset composition, model architecture, and validation strategies, limiting the comparability and generalizability of findings.
A considerable proportion of existing studies rely on two‐dimensional imaging modalities, which are inherently constrained by anatomical superimposition, geometric distortion, and limited representation of lesion depth. These limitations may restrict the ability of DL models to capture subtle spatial characteristics that are critical for accurate lesion characterization. Furthermore, many studies employ internal validation datasets, raising concerns regarding overfitting and the external validity of reported diagnostic performance [35, 36, 37, 38].
Cone‐beam computed tomography (CBCT), in contrast, provides high‐resolution three‐dimensional visualization of craniofacial structures, enabling a more comprehensive assessment of lesion morphology, cortical involvement, and spatial relationships with adjacent anatomical structures [40, 41, 42]. This volumetric information may enhance the differentiation of OKCs from other odontogenic lesions. Nevertheless, despite the growing adoption of CBCT in clinical practice, the evidence regarding the performance of DL models applied to CBCT data remains limited and fragmented, and has not been systematically synthesized.
A recent systematic review and meta‐analysis by Fedato Tobias et al. [39] evaluated the diagnostic performance of AI in detecting and classifying odontogenic cysts and tumors, including OKCs and AMEs. Although the authors reported encouraging diagnostic accuracy, their analysis encompassed heterogeneous lesion types and imaging modalities, and only a limited number of included studies specifically addressed OKC classification. Notably, the CBCT‐based evidence was derived from a small subset of studies, restricting the robustness and generalizability of the conclusions. Consequently, the diagnostic performance of AI‐based models, particularly those utilizing CBCT imaging for OKC detection and differentiation, remains insufficiently characterized.
Therefore, a systematic review and meta‐analysis is warranted to evaluate the diagnostic performance of AI‐based models for the detection of OKCs using CBCT imaging, with the aim of providing more robust and clinically relevant estimates of their diagnostic accuracy.
2. Methods
2.1. Study Design and Protocol Registration
This systematic review protocol was developed in accordance with the PRISMA‐DTA (Preferred Reporting Items for Systematic Reviews and Meta‐Analyses of Diagnostic Test Accuracy Studies) guidelines [43, 44]. However, due to PROSPERO's restriction on registering diagnostic test accuracy systematic reviews, the protocol has not been registered.
2.2. Search Strategy
A comprehensive literature search was performed across five major databases: PubMed, EMBASE, Scopus, ScienceDirect, and Google Scholar from inception to July 6, 2025. The search strategy combined MeSH terms and free‐text keywords related to AI and OKC. Boolean operators (“AND,” “OR”) were used to refine the search. The reference lists of included articles were also screened manually to identify any additional relevant studies. Table 1 presents a detailed overview of the search queries used in each dataset (Table 1).
Table 1.
Search expressions utilized across databases.
| Dataset | Search query | Results |
|---|---|---|
| Pubmed | (“artificial intelligence”[MeSH Terms] OR (“artificial”[All Fields] AND “intelligence”[All Fields]) OR “artificial intelligence”[All Fields] OR (“antagonists and inhibitors”[MeSH Subheading] OR (“antagonists”[All Fields] AND “inhibitors”[All Fields]) OR “antagonists and inhibitors”[All Fields] OR “ai”[All Fields]) OR (“deep learning”[MeSH Terms] OR (“deep”[All Fields] AND “learning”[All Fields] OR “deep learning”[All Fields]) OR (“machine learning”[MeSH Terms] OR (“machine”[All Fields] AND “learning”[All Fields] OR “machine learning”[All Fields]) OR (“convolutional neural networks”[MeSH Terms] OR (“convolutional”[All Fields] AND “neural”[All Fields] AND “networks”[All Fields]) OR “convolutional neural networks”[All Fields] OR (“convolutional”[All Fields] AND “neural”[All Fields] AND “network”[All Fields] OR “convolutional neural network”[All Fields] OR “CNN”[All Fields] AND (“odontogenic cysts”[MeSH Terms] OR (“odontogenic”[All Fields] AND “cysts”[All Fields] OR “odontogenic cysts”[All Fields] OR (“odontogenic”[All Fields] AND “keratocyst”[All Fields] OR “odontogenic keratocyst”[All Fields] OR “OKC”[All Fields] OR (“radicular cyst”[MeSH Terms] OR (“radicular”[All Fields] AND “cyst”[All Fields] OR “radicular cyst”[All Fields] OR (“dental”[All Fields] AND “cysts”[All Fields] OR “dental cysts”[All Fields] AND (“CBCT”[All Fields] OR (“cone beam computed tomography”[MeSH Terms] OR (“cone beam”[All Fields] AND “computed”[All Fields] AND “tomography”[All Fields]) OR “cone beam computed tomography”[All Fields] OR (“cone”[All Fields] AND “beam”[All Fields] AND “computed”[All Fields] AND “tomography”[All Fields] OR “cone beam computed tomography”[All Fields] | 22 |
| Scopus | TITLE‐ABS‐KEY (“artificial intelligence” OR “deep learning” OR “machine learning” OR “CNN” OR “convolutional neural networks”) AND (“odontogenic keratocyst” OR “OKC” OR “Odontogenic cysts”) | 55 |
| ScienceDirect | (artificial intelligence) OR (AI) OR (deep learning) OR (machine learning) OR (convolutional neural network) AND (odontogenic keratocyst) OR (OKC) AND (CBCT) OR (cone beam computed tomography) | 63 |
| Embase | (‘artificial intelligence’/exp OR ‘artificial intelligence’ OR (artificial AND (‘intelligence’/exp OR intelligence) OR ai OR ‘deep learning’/exp OR ‘deep learning’ OR (deep AND (‘learning’/exp OR learning)) OR ‘machine learning’/exp OR ‘machine learning’ OR (‘machine’/exp OR machine) AND (‘learning’/exp OR learning)) OR ‘convolutional neural network’/exp OR ‘convolutional neural network’ OR (convolutional AND neural AND (‘network’/exp OR network)) OR cnn) AND (‘odontogenic keratocyst’/exp OR ‘odontogenic keratocyst’ OR (odontogenic AND (‘keratocyst’/exp OR keratocyst)) OR okc OR ‘dental cysts’ OR (‘dental’/exp OR dental) AND (‘cysts’/exp OR cysts) AND (cbct OR ‘cone beam computed tomography’/exp OR ‘cone beam computed tomography’ OR (cone AND beam AND computed AND (‘tomography’/exp OR tomography) | 33 |
| Google Scholar | allintitle: (“artificial intelligence” OR “deep learning” OR “machine learning” OR “convolutional neural networks”) AND (“odontogenic keratocyst” OR “Odontogenic cysts”) OR (“CBCT” OR “cone beam computed tomography”) | 179 |
2.3. Eligibility Criteria
Studies were included based on the following PICO framework:
Population (P): Human subjects with confirmed OKCs evaluated using CBCT imaging.
Intervention (I): Application of AI models, such as CNNs, ResNet, or other DL architectures, for detection or classification of OKCs.
Comparison (C): Diagnostic performance compared to a reference standard, typically histopathological confirmation or expert radiologist interpretation.
Outcome (O): Reported diagnostic metrics such as sensitivity, specificity, accuracy, precision, recall, F1‐score, or area under the receiver operating characteristic curve (AUC).
Studies were excluded if they did not involve CBCT imaging, were reviews, commentaries, editorials, or case reports, lacked original diagnostic performance data, used only panoramic or 2D imaging modalities, or did not use AI methods for image analysis.
2.4. Study Selection
Two independent reviewers (A.Y. and S.E.) screened all identified articles based on titles and abstracts. Full texts were retrieved for potentially eligible studies. Discrepancies were resolved through discussion or by consulting a third reviewer (S.L.). The study selection process was documented using a PRISMA flow diagram (Figure 1).
Figure 1.

Flow diagram of the study selection process according to PRISMA guidelines, detailing 352 records identified, 64 duplicates removed, 288 records screened, 28 full‐text articles assessed for eligibility, 12 studies included in the systematic review, and 5 studies in the meta‐analysis.
2.5. Data Extraction
A standardized data extraction sheet was used by two reviewers (R.S. and S.E.) to independently collect information from each included study. Extracted data included: general study characteristics (authors, year, and country); sample size and participant demographics; CBCT acquisition protocols and image quality parameters; type and architecture of AI models used; reference standard used for diagnosis; dataset composition (training, validation, and testing); reported diagnostic metrics (e.g., sensitivity, specificity, accuracy, and AUC).
Disagreements during data extraction were resolved by consensus or third‐party adjudication (S.L).
2.6. Risk of Bias Assessment
The methodological quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS‐2) [45]. Four domains were evaluated: patient selection, index test, reference standard, and flow and timing. Two reviewers (R.S. and A.Y.) independently conducted the assessment, with disagreements resolved through discussion.
2.7. Data Synthesis
A meta‐analytical pooling of diagnostic performance was conducted using a random‐effects model to account for the expected variability across studies. Pooled sensitivity, specificity, and diagnostic odds ratios (DORs) were calculated with corresponding 95% confidence intervals (CIs). The DerSimonian and Laird method was applied for estimating summary statistics.
Heterogeneity was quantified using the I 2 statistic and Cochran's Q test, with I 2 > 50% indicating substantial heterogeneity. Where applicable, subgroup analyses were performed based on AI model type, sample size, and publication year. Additionally, meta‐regression was conducted to explore potential sources of heterogeneity.
Publication bias was evaluated using funnel plot symmetry and Egger's test, and adjustments were made using the Trim‐and‐Fill method if bias was detected.
All statistical tests were two‐sided, and values < 0.05 were considered statistically significant. The analyses were performed using R software (version 4.3.1; R Foundation for Statistical Computing, Vienna, Austria). The following R packages were utilized: “meta” for classical meta‐analyses and forest plots; “metafor” for meta‐regression, funnel plots, and DOR analyses; “ggplot2” for visualization of the summary receiver operating characteristic (SROC) curves; and “zoo” for numerical estimation of the area under the curve (AUC).
3. Results
3.1. Study Selection
The PRISMA flow diagram is presented in Figure 1. A comprehensive search across five databases (PubMed, EMBASE, Scopus, ScienceDirect, and Google Scholar) yielded a total of 352 records. After removing 64 duplicates, 288 unique studies remained for screening based on titles and abstracts. Of these, 260 studies were excluded for not meeting the inclusion criteria. The full texts of the remaining 28 articles were assessed for eligibility, resulting in the inclusion of 12 studies in the final systematic review (Figure 1).
3.2. Study Characteristics
The summary of key findings from the included studies is presented in Table 2. The 12 included studies were published between 2017 and 2025. The number of OKC cases per study ranged from 13 to 188. Most studies focused on classification tasks (n = 11) [38, 40, 41, 42, 46, 47, 48, 49, 50, 51, 52], while a subset also addressed segmentation (n = 3) [48, 49, 53] and detection objectives (n = 1) [49].
Table 2.
Summary of key findings from studies assessing AI performance in OKC detection and classification.
| Study reference and year | Research aim | Dataset description | Data split (train/validation/test) | Participant characteristics | Inclusion & exclusion parameters | Annotation & labeling approach | ML task overview | Data preprocessing methods | Model architecture & algorithms | Top performance metrics (OKC‐specific) |
|---|---|---|---|---|---|---|---|---|---|---|
| Sha X, 2025 [40] | To use AI algorithms for the differential diagnosis of OCs, AME, and OKC | 300 patients (100 OC, 100 OKC, 100 AME) | 70% training (n = 210) 30% testing (n = 90) fivefold cross‐validation | 124 Males, 86 Females with 41 yrs mean age (29–55 yrs) in training cohort 49 Males, 41 Females with 43 yrs mean age (30.5–54 yrs) in testing cohort | Inclusion: Histopathologically confirmed diagnoses | Manual segmentation of ROIs | Classification | NA | Soft VotingClassifier (Combination of Random Forest, SVC logistic regression model) | Soft VotingClassifier Training cohort OKC‐vs‐Rest AUC: 92.8% Training cohort OKC‐vs‐Rest Accuracy: 83.3% Training cohort OKC‐vs‐Rest Sensitivity: 74.3% Training cohort OKC‐vs‐Rest Precision: 75.4% Training cohort OKC‐vs‐Rest Specificity: 86.3% Training cohort OKC‐vs‐Rest F1 score: 74.8% Testing cohort OKC‐vs‐Rest AUC: 78.1% Testing cohort OKC‐vs‐Rest Accuracy: 80% Testing cohort OKC‐vs‐Rest Sensitivity: 66.7% Testing cohort OKC‐vs‐Rest Precision: 71.4%Testing cohort OKC‐vs‐Rest Specificity: 84.6%Testing cohort OKC‐vs‐Rest F1 score: 68.9% |
| Muraoka H, 2025 [46] | To use ML models for the classification of odontogenic cysts and tumors | 135 patients (52 DC, 60 OKC, 23 AME) | 70% training 30% testing fivefold stratified cross‐validation within the training set | 99 Males, 36 Females Mean age: 43.57 yrs (9–83 yrs) | Inclusion: Histopathologically confirmed diagnoses Exclusion: Images with sever metal artifacts, small lesions (difficult feature extraction), infection‐associated lesions | Lesion type labeling | Classification | Manual segmentation, ROIs excluded structures like the mandibular canal, tooth roots, and cortical bone, Data augmentation, synthetic samples generated with random noise | Random forest classifier | Overall Precision: 87%Overall recall: 87%Overall F1 score: 87%Cross‐validation accuracy: Training dataset: 90%Testing dataset: 94% |
| Committeri U, 2024 [47] | To use ML models for the differential diagnosis of DCs, AME, and OKC | 103 patients (54 DC, 24 OKC, 25 AME) | 10‐fold cross‐validation | DC: 37 Males, 17 Females, mean age: 46 (range 23–67) OKC: 21 Males, 3 Females, mean age: 47 (range 24–69) AME: 17 Males, 8 Females, mean age: 55 (range 17–86) | Inclusion: Full clinical records, histopathologically confirmed diagnoses, primary lesions (no recurrent), Patients ≥ 18 years Exclusion: Images with artifacts, Syndromic/multiple OKCs, Solid/multicystic ameloblastomas, Inflammatory/autoimmune conditions or medications affecting biomarkers | Delineated ROIs slice‐by‐slice by two expert radiologists histopathology diagnosis‐based labeling | Classification | Manual segmentation | Decision Tree (including Fine Tree), k‐NN, SVM, Medium Neural Network | Medium Neural NetworkAccuracy: 82.5% (train/test)Accuracy: 73.3%% (validation set)AUC: 83% |
| Huang Z, 2024 [53] | To use an auto‐adapting multi‐scaled UNet for the segmentation of odontogenic cysts | 300 CBCT scans (51 DC, 53 RC, 117 OKC, 79 AME) | 80% training 20% testing | NA | Inclusion: Full clinical records, histopathologically confirmed diagnoses Exclusion: Images with excessive artifacts, Poor image quality, Out‐of‐field lesions, Multiple lesions, or coexisting pathology types | Manual labeling, Pixel‐Wise Segmentation Masks | Segmentation | Image resizing to 128 × 128 × 128, removing irrelevant background, and cropping images | A new OCL‐Net | Dice Score: 88.84%IoU: 81.23%AUC: 92.37% |
| Liu W, 2024 [48] | To use nnU‐Net for the classification and segmentation of odontogenic cysts | 368 CBCT (37,168 slices) (75 DC, 73 OKC, 87 AME, 62 PC, 61 OST) | 70% training 20% validation 10% testing | 233 Males, 135 Females 8–91 yrs | Inclusion: Full clinical records, histopathologically confirmed diagnoses, clear lesion regions Exclusion: Previously treated lesions, mixed dentition in lesion areas | Slice‐by‐slice polygon‐based annotation using color codes for each lesion type by two trained oral and maxillofacial surgeons | Classification segmentation | Cropping, normalization, filtering, resampling, augmentation (Rotation, scaling, Gaussian noise/blurring, gamma transform, mirroring) | A modified nnU‐Net | Classification (AI): Sensitivity: 87.1% Specificity: 97.4% Precision: 87.4% Accuracy: 89.1%Classification (OMFR): Sensitivity: 83.3% Specificity: 96.5%Precision: 82.5% Accuracy: 85.5%Classification (OMS): Sensitivity: 78.3%Specificity: 95.6%Precision: 79.5%Accuracy: 81.9%F1‐Score (AI model): 71.4%F1‐Score (OMFRs): 68.4%F1‐Score (OMSs): 0.642F1‐Score improvement (with AI): +6.2% to +12.7% Segmentation:Dice Similarity Coefficient: 89.3% Average Symmetric Surface Distance: 0.752 mm |
| Song Y, 2024 [38] | To use ML algorithms for the differential diagnosis of AME and OKC | 326 patients (152 AMEs, 172 OKCs) | 5‐fold cross‐validation | AME: 94 Males, 58 Females, mean age: 38.3 ± 16.8OKC: 90 Males, 84 Females, mean age: 36.8 ± 17.2 | Inclusion: Full clinical records, histopathologically confirmed diagnoses, and available preoperative CBCT | Manual and semi‐automatic segmentation and the ground truth labels based on postoperative histopathological diagnosis | Classification | Grey‐level discretization, applying image filters, and normalization | XGBoost SVM | XGBoost Precision: 90.0% Recall: 80.7% Accuracy: 84.1% F1‐Score: 84.3% AUC: 87.2% SVM Precision: 67.5% Recall: 62.7% Accuracy: 62.6% F1‐Score: 63.3% AUC: 62.3% |
| Yeshua T, 2023 [49] | To use DL algorithms for benign bone lesion detection and segmentation | 82 CBCT scans:41 benign bone lesions (13 OKC) 41 control scans | 50 training scans 10 validation scans 22 testing scans | Benign bone lesions: 19 Males, 22 Females, mean age: 33.0 ± 18.9 (range 3–68)Control group:15 Males, 26 Females, mean age: 56.0 ± 16.0 (range 19–79) | Inclusion: Full clinical records, histopathologically confirmed diagnoses, and available preoperative CBCT Exclusion: Motion artifacts | Manual Annotation by experienced radiologists (initial) and senior radiologists (+10 yrs experience) (final) | Detection Segmentation Classification | Conversion from DICOM to JPEG, Augmentation (horizontal flipping) | Mask‐RCNN (with ResNet101 backbone) | Classification Accuracy (CBCT‐level): 100% Detection Accuracy (CBCT‐level): 100% Sensitivity:Before improvement: 83.5% After improvement: 95.9% Precision:Before improvement: 88.4% After improvement: 98.8% Dice Coefficient: Average: 83.5% ± 7.8% (range: 66.8%–93.0%) |
| Chai Z, 2022 [41] | To use CNN for the differential diagnosis of AME and OKC | 350 patients (178 AMEs, 172 OKCs) | 272 training cases 78 testing cases | AME: age: 9–81 yrs (mean: 40.3 ± 16.5)OKC: age: 10–70 yrs (mean: 41.5 ± 17.6) | Inclusion: Full clinical records, histopathologically confirmed diagnoses, and available preoperative CBCT Exclusion: Multiple OKCs, nevoid basal cell carcinoma syndrome, images with artifacts | NA | Classification | 150 × 150 cropping | Inception v3 | Inception v3Sensitivity: 87.2% Specificity: 82.1%Accuracy: 84.6% F1 score: 85%7 senior oral and maxillofacial surgeons: Sensitivity: 60% Specificity: 71.4%Accuracy: 65.7% F1 score: 63.6%30 junior oral and maxillofacial surgeons: Sensitivity: 63.9% Specificity: 53.2% Accuracy: 58.5% F1 score: 60.7% |
| Bispo M S, 2021 [42] | To use CNN and Google Inception v3 for the differential diagnosis of AME and OKC | 350 images of 48 patients (22 AMEs, 18 OKCs) | 2500 Augmented images (45% training, 45% validation, 10% testing) | NA | Inclusion: Histopathologically confirmed diagnoses of AME and OKC Exclusion: Lesions > 80 mm, Recent surgical intervention, Evident infection at lesion site, Images with excessive metallic artifacts | Manual ROI segmentation by a calibrated examiner using ImageJ software | Classification | Manual segmentation, File Format Conversion, resizing to 299 × 299 pixels, Data Augmentation (Rotations, Horizontal and vertical flips, Elastic distortion (moderate configuration), Resizing (to regularize different ROI sizes)) | CNN (Google Inception v3) | Accuracy: 90.16%–92.48% |
| Lee JH, 2020 [50] | To use Google Inception v3 for the differential diagnosis of OKC, DC, and PC | 247 patients (986 CBCTs) (188 OKC, 396 DC, 402 PC) | 789 training CBCT cases 592 validation CBCT cases 197 testing CBCT cases | 167 Males, 80 Females Age: < 20 years: 5.7% 20–39 years: 37.6% 40–59 years: 39.7% ≥ 60 years: 17.0% | Inclusion: Histopathologically confirmed diagnoses of OKC, DC, or PC Exclusion: Images with severe distortion, artificial noise, blur, or poor quality | NA | Classification | Cropping and Resizing Brightness and Contrast Normalization Data Augmentation (Horizontal/vertical flipping, rotation, width/height shifting, Shearing, Zooming | CNN (GoogLeNet Inception v3) | AUC: 91.4% Sensitivity: 96.1% Specificity: 77.1% Accuracy: 87.2% |
| Yilmaz E, 2017 [51] | To use SVM for the differential diagnosis of OKC and PC | 50 patients (127 sections per lesion) | 25 training cases 25 testing cases | NA | Inclusion: Clinical, radiographic, and histopathologically confirmed diagnoses CBCT | Manual annotation | Classification | Manual segmentation | SVM, NN, NB, DT, RF, k‐NN | SVM:Accuracy: 100% F1 score: 100%NN:Accuracy: 92% F1 score: 91.67%NB:Accuracy: 98% F1 score: 98%RF:Accuracy: 92% F1 score: 92.86% |
| Abdolali F, 2017 [52] | To use of 2 different automated classifiers for the differential diagnosis of maxillofacial cysts | 96 patients (38 RC, 36 DC, 22 OKC) | NA | NA | Inclusion: Histopathologically confirmed diagnoses CBCT | Manual annotation by two radiologists | Classification | Segmentation | SVM SDA | SVM classifier:Sensitivity: 90.91% Specificity: 95.77% Accuracy: 94.29%SDA classifier:Sensitivity: 95.45% Specificity: 98.59% Accuracy: 96.48% |
Abbreviations: AI, Artificial Intelligence; AME, Ameloblastoma; CBCT, Cone Beam Computed Tomographic; CLAHE, Contrast‐limited Adaptive Histogram Equalization; CNN, Convolutional Neural Network; DC, Dentigerous Cyst; DT, Decision Tree; DL, Deep learning; FCN, Fully Conventional Network; k‐NN, k‐Nearest Neighbors; NA, Not Available; NB, Naïve Bayes; NN, Neural Network; OCs, odontogenic cysts; OKC, Odontogenic keratocyst; OPG, Orthopantomogram; OST, Osteomyelitis; OMFR, oral and maxillofacial radiologists; OMS, oral and maxillofacial surgeons; PC, Periapical Cyst; RC, Radicular Cyst; RF, Random Forest; SBC, Stafne's bone cavity; SDA, Sparse Discriminant Analysis; SVM, Support Vector Machine; ML, Machine Learning; XGBoost, eXtreme Gradient Boosting.
Histopathological confirmation was used as the reference standard in all studies. Common inclusion criteria included the availability of preoperative CBCT scans and complete clinical records. Frequently cited exclusion criteria were image artifacts, syndromic lesions, and the presence of multiple concurrent cyst types.
Annotations were commonly performed manually by trained radiologists or surgeons, often on a slice‐by‐slice basis, using either polygon‐based or region‐of‐interest (ROI) segmentation.
A diverse range of AI algorithms was employed. Classical ML models included support vector machines (SVM), random forests, k‐nearest neighbors (k‐NN), decision trees, XGBoost, and sparse discriminant analysis (SDA). DL models comprised classification CNNs such as Inception v3 and segmentation architectures like modified U‐Nets (e.g., nnU‐Net, OCL‐Net) and Mask R‐CNN (ResNet101 backbone).
Overall, classical ML models were more commonly applied in smaller datasets, while DL approaches were increasingly used in recent years, particularly for larger datasets or tasks requiring complex feature extraction, such as segmentation or detection. This trend reflects the growing availability of computational resources and annotated CBCT datasets.
3.3. Diagnostic AI Performance
3.3.1. Classical ML Models
Six studies employed classical ML methods [38, 40, 46, 47, 51, 52]. Reported diagnostic accuracy ranged from 62.6% to 100%, and F1‐scores from 63.3% to 100%. The lowest performance was reported in the SVM model by Song et al. [38] (Accuracy: 62.6%, AUC: 0.623), while the highest was observed in Yilmaz et al. [51], whose SVM achieved 100% accuracy and F1‐score. These discrepancies likely stem from differences in dataset size, lesion diversity, and validation methodology.
Other high‐performing models included random forest (94% accuracy, 87% F1‐score) [46], XGBoost (84.1% accuracy, AUC 87.2%) [38], and SDA (96.5% accuracy, 98.6% specificity) [52]. Ensemble methods and optimized feature selection strategies tend to yield better diagnostic performance than basic classifiers, such as SVM.
In general, ensemble methods and optimized feature selection consistently outperformed basic classifiers, suggesting that careful feature engineering remains critical in classical ML approaches. Notably, performance variability across studies indicates sensitivity to dataset size, lesion heterogeneity, and validation strategies.
3.3.2. CNN‐Based DL Models for Classification
Three studies utilized CNN‐based classifiers, predominantly based on the Inception v3 architecture [41, 42, 50]. Reported accuracies ranged from 84.6% to 92.5%, with F1‐scores above 85%. The highest classification accuracy was achieved by Bispo et al. [42] using Inception v3 with augmented CBCT datasets (up to 92.5%). Chai et al. [41] demonstrated 84.6% accuracy and 85% F1‐score, outperforming both senior (65.7%) and junior (58.5%) oral surgeons. Lee et al. [50] reported 87.2% accuracy and an AUC of 91.4% using GoogLeNet (Inception v3), with high sensitivity (96.1%) but moderate specificity (77.1%).
Despite differences in preprocessing, dataset sizes, and augmentation methods, Inception‐based CNNs consistently outperformed clinical experts, both junior and senior clinicians. While accuracy differences between studies exist, the overall trend demonstrates the strong potential of CNNs in assisting OKC diagnosis.
3.3.3. Instance‐Level Detection and Segmentation Networks
One study explored instance‐level lesion detection and segmentation using region‐based CNNs (R‐CNNs). Yeshua et al. [49] applied a Mask R‐CNN with a ResNet101 backbone and achieved 100% lesion detection accuracy at the scan level. Sensitivity improved from 83.5% to 95.9%, and precision from 88.4% to 98.8%, following model optimization. The average Dice similarity coefficient was 83.5% (±7.8%), indicating strong spatial agreement with ground truth annotations.
Although only one study explored instance‐level detection, the high sensitivity and precision suggest that region‐based CNNs are capable of accurately localizing lesions at the scan level, highlighting a promising direction for clinical implementation.
3.3.4. Semantic Segmentation Networks (U‐Net Variants)
Two studies utilized U‐Net‐based convolutional architectures for semantic segmentation. Huang et al. [53] developed a multi‐scale model (OCL‐Net), achieving a Dice coefficient of 88.8%, IoU of 81.2%, and AUC of 92.4%. Liu et al. [48] applied a self‐configuring nnU‐Net, reporting slightly higher Dice (89.3%) and superior boundary precision (0.752 mm average surface distance). In addition to segmentation, Liu's pipeline also included a classification component (Accuracy: 89.1%), outperforming the diagnostic performance of clinicians [48].
The consistent high Dice scores and boundary precision across U‐Net variants indicate reliable segmentation performance, which can support subsequent classification and treatment planning. Both studies show that segmentation accuracy can exceed human‐level performance for delineating lesion boundaries.
3.4. Meta‐Analysis
Five studies [40, 41, 48, 50, 52] provided complete sensitivity and specificity data and were included in the meta‐analysis. These studies evaluated six AI models for OKC detection (Figure 2).
Figure 2.

Forest Plots of Pooled Sensitivity and Specificity for AI‐Based OKC Classification; (A) The overall pooled sensitivity was 89% (95% CI: 79%–95%), with deep learning (DL) models achieving 91% (95% CI: 82%–96%) and classical machine learning (ML) models achieving 86% (95% CI: 59%–96%). (B) The overall specificity was 92% (95% CI: 81%–97%), with DL models at 89% (95% CI: 65%–97%) and classical ML models at 95% (95% CI: 81%–99%). No statistically significant subgroup differences were found for either metric.
3.4.1. Sensitivity and Specificity
The pooled sensitivity across all AI models was 89% (95% CI: 79%–95%) with high heterogeneity (I2 = 78.5%; Cochran's Q test p < 0.001), indicating variability beyond chance. DL models showed slightly higher pooled sensitivity (91%, 95% CI: 82%–96%) compared to classical ML models (86%, 95% CI: 59%–96%). However, subgroup differences were not statistically significant (χ2 = 0.42, p = 0.5159) (Figure 2A).
DL models yielded a pooled specificity of 89% (95% CI: 65–97%) with considerable heterogeneity (I2 = 93.6%; Cochran's Q test p < 0.001). Classical ML models demonstrated slightly higher pooled specificity (95%, 95% CI: 81–99%; I2 = 76.5%). Overall pooled specificity was 92% (95% CI: 81–97%), with no significant difference between subgroups (χ2 = 0.57, p = 0.4511) (Figure 2B).
Overall, the pooled sensitivity (89%) and specificity (92%), both exceeding the 85% threshold commonly considered indicative of strong diagnostic performance, together with a high pooled DOR (87.06), indicate that AI models possess substantial discriminative capacity for OKC detection on CBCT imaging.
3.4.2. Meta‐Regression Analysis
Meta‐regression was conducted to assess whether the type of AI model (DL vs. ML) explained variability in diagnostic performance. For sensitivity, residual heterogeneity remained substantial (I2 = 79.9%; Q test p = 0.002), and the AI model type was not a significant moderator (regression coefficient p = 0.42), explaining less than 2% of the between‐study variability. Similar findings were noted for specificity. AI model type was not associated with specificity (regression coefficient p = 0.44), with minimal explained heterogeneity (R2 < 2%) in both models.
Publication year was also evaluated as a covariate. Publication year demonstrated a statistically significant negative association with sensitivity (regression coefficient = –0.21; p = 0.04), explaining 57.2% of between‐study heterogeneity (R2 = 57.18%). This suggests that sensitivity declined slightly in more recent studies, possibly due to increased methodological rigor, stricter validation, or more complex comparator groups. In contrast, no association was observed with specificity (regression coefficient = –0.03; p = 0.90).
3.4.3. SROCs Analysis
The SROC curve (Figure 3) showed an AUC of 0.828, indicating robust overall discriminative ability and confirming that AI‐based models can reliably distinguish OKCs from comparator lesions across varying diagnostic thresholds. Most studies clustered in the upper‐left quadrant, suggesting consistently high sensitivity and specificity. One outlier (Sha X, 2025 [40]) showed lower sensitivity (66.7%) and a higher false positive rate, slightly reducing the curve's performance. The pooled point estimate (black triangle) aligned closely with the study cluster, supporting consistency in diagnostic accuracy.
Figure 3.

Summary Receiver Operating Characteristic (SROC) curve for AI‐based detection of odontogenic keratocysts (OKC). The red curve represents the fitted SROC across six studies. Blue dots indicate individual study estimates, with annotations. The dashed black line denotes random chance. The pooled performance yields an area under the curve (AUC) of 0.828, indicating good overall diagnostic accuracy.
3.4.4. DOR Analysis
To assess the overall diagnostic power of AI models, a DOR meta‐analysis was conducted. Figure 4 shows the DOR forest plot, sub‐grouped by AL models. The pooled DOR under the random‐effects model was 87.06 (95% CI: 28.67–264.31, p < 0.001), indicating strong overall discriminative performance. However, substantial heterogeneity was present across studies (I2 = 83.4%, p < 0.001).
Figure 4.

Forest Plot of Diagnostic Odds Ratios (DOR) for AI Models, Stratified by Model Type. Forest plot displaying the DOR and 95% CI for six AI models from five studies. The overall pooled DOR under the random‐effects model was 87.06 (95% CI: 28.67–264.31), indicating high discriminative capacity. Subgroup analysis revealed a pooled DOR of 129.82 for classical ML models, 47.07 for deep learning CNN models, and 255.11 for a single segmentation‐based CNN. While the test for subgroup differences approached significance (p = 0.0540), considerable heterogeneity was observed (I2 = 83.4%).
Subgroup analysis by AI model group showed notable variation. Classical ML models showed a pooled DOR of 129.82 (95% CI: 7.09–2375.30; k = 3; I2 = 87.1%), DL CNNs reported a DOR of 47.07 (95% CI: 18.18–121.84; k = 2; I2 = 65.6%), and the single segmentation‐based CNN model had the highest DOR of 255.11 (95% CI: 94.78–686.67). The test for subgroup differences approached statistical significance under the random‐effects model (Q = 5.84, df = 2, p = 0.054), suggesting potential variation in diagnostic performance across AI model types.
3.4.5. Publication Bias Assessment
Publication bias was evaluated using a funnel plot of log DORs and Egger's test (Figure 5). The plot was visually symmetrical, and Egger's test did not indicate significant bias (z = 1.64, p = 0.10). However, given the small number of studies (k = 6), statistical power to detect bias remains limited.
Figure 5.

Funnel plot evaluating the risk of publication bias based on log diagnostic odds ratios. Visual inspection suggests approximate symmetry; Egger's regression test indicated no significant asymmetry (p = 0.10).
3.5. Risk of Bias Assessment
Risk of bias assessment using the QUADAS‐2 tool revealed that patient selection was the most problematic domain, with 33.3% (n = 4) of studies rated high risk and an additional 25% (n = 3) rated unclear, primarily due to retrospective designs and insufficient reporting on recruitment methods. The index test domain was generally well‐handled, with 58.3% (n = 7) of studies at low risk, though 33.3% (n = 4) were unclear. The reference standard and flow and timing domains showed uniformly low risk across all studies, supported by consistent use of histopathological confirmation and standardized diagnostic workflows (Figure 6 and Table 3).
Figure 6.

QUADAS‐2 Risk of Bias and Applicability Summary; (A) Risk of bias assessment across four domains: The patient selection domain had the highest risks, with 33.3% (n = 4) rated high and 25% (n = 3) unclear. The index test domain had moderate concerns, with 33.3% (n = 4) unclear. All studies were low risk for the reference standard and flow/timing. (B) Applicability assessment across three domains: All included studies showed low concern in all domains, indicating high relevance to the research question.
Table 3.
QUADAS‐2 Evaluation of Risk of Bias and Study Applicability.
| Risk of Bias Assessment | Applicability Assessment | ||||||
|---|---|---|---|---|---|---|---|
| Study | Patient Selection | Index Test | Reference Standard | Flow and Timing | Patient Selection | Index Test | Reference Standard |
| Sha X, 2025 [40] | High | High | Low | Low | low | low | low |
| Muraoka H, 2025 [46] | Unclear | Low | Low | Low | low | low | low |
| Committeri U, 2024 [47] | Low | Unclear | Low | Low | low | low | low |
| Huang Z, 2024 [53] | High | Low | Low | Low | low | low | low |
| Liu W, 2024 [48] | Unclear | Low | Low | Low | low | low | low |
| Song Y, 2024 [38] | Low | Low | Low | Low | low | low | low |
| Yeshua T, 2023 [49] | Low | Low | Low | Low | low | low | low |
| Chai Z, 2022 [41] | Low | High | Low | Low | low | low | low |
| Bispo M S, 2021 [42] | High | Low | Low | Low | low | low | low |
| Lee JH, 2020 [50] | Low | Low | Low | Low | low | low | low |
| Yilmaz E, 2017 [51] | High | Unclear | Low | Low | low | low | low |
| Abdolali F, 2017 [52] | High | Unclear | Low | Low | low | Low | low |
Applicability concerns were minimal, with all studies rated low concern across all QUADAS‐2 domains (Table 3).
4. Discussion
This systematic review and meta‐analysis evaluated the diagnostic accuracy of AI models in OKCs using CBCT. Our findings indicate that AI models demonstrate high diagnostic performance, with pooled sensitivity and specificity of 89% and 92%, respectively. These results support the growing evidence that AI can assist clinicians in differentiating among odontogenic lesions with overlapping imaging features, such as OKC and AME, and serve as valuable adjuncts in radiologic diagnosis [40, 46, 47, 54].
Our findings are consistent with prior systematic reviews. For example, Shrivastava et al. [54] reported a pooled sensitivity and specificity of 88% for AI applied to CBCT and panoramic images in diagnosing odontogenic cysts and tumors, while Fedato et al. [39] found diagnostic accuracies around 89%, with CBCT outperforming 2D modalities. Shoorgashti et al. [2], focusing on panoramic radiographs, reported high sensitivity (up to 96%) using YOLO‐based models. These consistent results across different imaging modalities underscore the diagnostic utility of AI, particularly when applied to CBCT. However, to our knowledge, this is the first meta‐analysis specifically focused on AI‐driven detection of OKC on CBCT imaging—a clinically relevant distinction, given OKC's aggressive behavior, high recurrence rate, and radiologic similarity to other cystic lesions.
The pooled DOR of 87.06 indicates that patients with OKC had approximately 87‐fold higher odds of being correctly identified by AI models compared with non‐OKC cases, while the AUC of 0.828 further confirms strong diagnostic discrimination. Subgroup analysis revealed that DL models yielded slightly higher sensitivity (91%) compared to classical ML models (86%), whereas ML exhibited marginally higher specificity. This difference likely reflects the inherent advantages of DL, particularly CNNs, which learn hierarchical features directly from raw imaging data [55, 56]. In contrast, traditional ML approaches rely on handcrafted radiomic features, which are limited by human interpretation and susceptible to variability in feature selection [57, 58, 59].
DL architectures are particularly well‐suited for CBCT data, given their ability to capture complex spatial and textural information from volumetric scans. By preserving inter‐slice spatial relationships, DL models can extract subtle morphological features critical for distinguishing OKCs from similar lesions, enhancing diagnostic performance [58, 60, 61].
One notable outlier on the SROC curve was the study by Sha et al. [40], which reported a relatively low sensitivity (66.7%) in the test set. Their study employed a radiomics‐based ensemble of ML classifiers (random forest, SVM, logistic regression) for multi‐class classification (OKC, AME, and other cysts). While clinically relevant, multi‐class tasks introduce greater complexity and often dilute class‐specific performance [62, 63]. Additionally, Sha et al. [40] utilized CBCT data from three different scanners, introducing heterogeneity in image quality and radiomic feature expression, factors known to undermine the reproducibility and robustness of handcrafted radiomic features [55, 56]. This inter‐scanner variability, despite voxel‐size standardization, remains a challenge in radiomics due to differences in acquisition protocols and reconstruction algorithms [64, 65, 66].
Significant heterogeneity was observed across the included studies. Interestingly, this variability was not primarily explained by the type of AI model (DL vs. ML), but rather by clinical and methodological differences. A key contributor was variability in imaging data, particularly differences in scanner types, acquisition protocols, and gray‐level distributions, which can impact how models process input data. Such effects are more pronounced in ML models reliant on handcrafted features than in DL models, which learn features directly from imaging data [64, 65, 66, 67].
The complexity of the diagnostic task also played a role. Studies employing binary classification (e.g., OKC vs. non‐OKC) generally reported higher performance than those using multi‐class frameworks, which are inherently more challenging and sensitive to class imbalance and data sparsity [64, 65, 66].
A particularly insightful finding was the moderating effect of publication year on sensitivity. More recent studies [40, 41, 48] tended to report lower sensitivities, likely reflecting a shift toward more rigorous and transparent research practices. Earlier studies often used internal validation alone, potentially inflating diagnostic metrics. In contrast, newer studies increasingly adopt external validation and use heterogeneous datasets.
Other sources of heterogeneity can include sample size variation and inconsistent evaluation strategies. Additionally, small test sets and isolated outlier studies may have skewed pooled estimates and widened CIs.
Our results align with broader findings in dental imaging AI research [5, 68, 69, 70, 71, 72]. Systematic reviews of AI applications in detecting caries [73], periapical lesions [74], and maxillary sinus pathology [75] have reported pooled AUCs ranging from 0.83 to 0.95, depending on task complexity and modality. However, radiologic differentiation of OKC, particularly from AME, remains more diagnostically challenging due to overlapping features, which may account for the slightly lower AUC in our analysis.
This review offers several strengths. However, several limitations must be acknowledged. All included studies were retrospective, with limited external validation and a lack of standardized CBCT acquisition protocols. Sample sizes were modest, and the reporting of CIs and calibration metrics was inconsistent.
Despite these limitations, our findings have important clinical implications. Accurate imaging‐based differentiation of OKC from other odontogenic lesions can inform biopsy decisions, surgical planning, and patient counseling. Given the consistently high pooled sensitivity and specificity, AI‐assisted CBCT interpretation may serve as a reliable adjunctive diagnostic support tool, particularly for general practitioners or less experienced clinicians, potentially improving diagnostic consistency, reducing misclassification of aggressive lesions such as OKC, and supporting timely referral and management decisions. However, before routine clinical implementation, AI systems must undergo robust external validation, be tested in prospective trials, and incorporate features that promote interpretability and clinician trust.
Future research should prioritize the development of multi‐institutional datasets with standardized imaging and annotation protocols to improve generalizability. Prospective studies comparing AI‐assisted diagnosis against current clinical workflows are essential to establish real‐world utility. Integration of multimodal data, including imaging, clinical, and histopathological information, could further enhance performance. The development of explainable AI models will be crucial for clinician adoption, regulatory approval, and safe integration into diagnostic pathways in oral and maxillofacial radiology.
5. Conclusion
AI models, particularly DL–based CNN architectures, demonstrate high and clinically meaningful diagnostic accuracy for OKC detection on CBCT imaging, with pooled sensitivity and specificity exceeding 85% and strong overall discriminative performance (AUC = 0.828; DOR = 87.06). However, approximately 33% of included studies were judged to have a high risk of bias in the patient selection, primarily related to patient selection and retrospective study design, which may limit the certainty and generalizability of the pooled estimates. Therefore, although the findings support the potential integration of AI systems as adjunctive diagnostic tools, further prospective, externally validated, and methodologically rigorous studies are required before routine clinical implementation.
Author Contributions
Reyhaneh Shoorgashti: conceptualization, methodology, software, data curation, formal analysis, visualization, project administration, writing – original draft, writing – review and editing, validation, investigation, supervision. Simin Lesan: methodology, validation. Sarah Sadat Ehsani: methodology, writing – original draft, investigation. Asma Yazdanfar: methodology, writing – original draft, investigation.
Funding
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Declaration of Generative AI and AI‐Assisted Technologies in the Writing Process
During the preparation of this work, the author(s) used ChatGPT‐5 to enhance language fluency and improve the clarity of the manuscript. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.
Transparency Statement
Dr. Reyhaneh Shoorgashti affirms that this manuscript is an honest, accurate, and transparent account of the study being reported; that no important aspects of the study have been omitted; and that any discrepancies from the study as planned (and, if relevant, registered) have been explained.
Acknowledgments
The authors have nothing to report.
Data Availability Statement
The data sets generated and analyzed during the current research are available from the corresponding author on reasonable request.
All authors have read and approved the final version of the manuscript. Dr. Reyhaneh Shoorgashti had full access to all of the data in this study and takes complete responsibility for the integrity of the data and the accuracy of the data analysis.
References
- 1. Shoorgashti R., Jafari F., and Lesan S., “AI‐Powered Microscopic Diagnostic Techniques for Candida Albicans Detection: A Systematic Review,” Journal of Dentistry (Shiraz) 27 (2026): 1–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Shoorgashti R., Alimohammadi M., Baghizadeh S., Radmard B., Ebrahimi H., and Lesan S., “Artificial Intelligence Models Accuracy for Odontogenic Keratocyst Detection From Panoramic View Radiographs: A Systematic Review and Meta‐Analysis,” Health Science Reports 8 (2025): e70614. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Rokhshad R., Nasiri F., Saberi N., et al., “Deep Learning for Age Estimation From Panoramic Radiographs: A Systematic Review and Meta‐Analysis,” Journal of Dentistry 154 (2025): 105560. [DOI] [PubMed] [Google Scholar]
- 4. Shoorgashti R., Jafari F., Yazdanfar A., et al., “Diagnostic and Prognostic Performance of Artificial Intelligence Models in Detecting Odontogenic Keratocysts From Histopathologic Images: A Systematic Review and Meta‐Analysis,” Middle East Journal of Rehabilitation and Health Studies 12 (2025): e158082. [Google Scholar]
- 5. Rokhshad R., Mohammad‐Rahimi H., Price J. B., et al., “Artificial Intelligence for Classification and Detection of Oral Mucosa Lesions on Photographs: A Systematic Review and Meta‐Analysis,” Clinical Oral Investigations 28 (2024): 88. [DOI] [PubMed] [Google Scholar]
- 6. Alhassan A. M. and Altmami N. I., “Quadra Sense: A Fusion of Deep Learning Classifiers for Mitosis Detection in Breast Cancer Histopathology,” Diagnostics (Basel) 16 (2026): 393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Saha D., Mandal A., Biswas S. K., Das B., Bhattacharya A., and Das A. K., “PCOSFusionNet: Hybrid Deep Feature Fusion Network for PCOS Classification From Ultrasound Images of Ovaries,” Ultrasonic Imaging (New York, NY) 48 (2026): 358–377. [DOI] [PubMed] [Google Scholar]
- 8. Kaya M., Durdag O., Solak M., Coskuncay A., and Burakgazi G., “Lesion Detection Using Artificial Intelligence Models in MR Images of Prostate Cancer and Prostatitis Patients and Comparison of Model Performance,” Frontiers in Urology 5 (2025): 1726795. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Chen S., Parreno‐Centeno M., Booker G., et al., “Normal Breast Tissue (NBT)‐Classifiers: Advancing Compartment Classification in Normal Breast Histology,” NPJ Breast Cancer 12 (2026): 41. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Doğan R. S. and Yılmaz B., “Histopathology Image Classification: Highlighting the Gap Between Manual Analysis and AI Automation,” Frontiers in Oncology 13 (2023): 1325271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Ahmad I. S., Dai J., Xie Y., and Liang X., “Deep Learning Models for CT Image Classification: A Comprehensive Literature Review,” Quantitative Imaging in Medicine and Surgery 15 (2025): 962–1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Hammad M., ElAffendi M., El‐Latif A. A. A., Ateya A. A., Ali G., and Plawiak P., “Explainable AI for Lung Cancer Detection via a Custom CNN on CT Images,” Scientific Reports 15 (2025): 12707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Chen W., Dhawan M., Liu J., et al., “Mapping the Use of Artificial Intelligence‐Based Image Analysis for Clinical Decision‐Making in Dentistry: A Scoping Review,” Clinical and Experimental Dental Research 10 (2024): e70035. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Karobari M. I., Adil A. H., Basheer S. N., et al., “Evaluation of the Diagnostic and Prognostic Accuracy of Artificial Intelligence in Endodontic Dentistry: A Comprehensive Review of Literature,” Computational and Mathematical Methods in Medicine 2023 (2023): 7049360. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Parinitha M. S., Doddawad V. G., Kalgeri S. H., Gowda S. S., and Patil S., “Impact of Artificial Intelligence in Endodontics: Precision, Predictions, and Prospects,” Journal of Medical Signals and Sensors 14 (2024): 25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Issa J., Jaber M., Rifai I., Mozdziak P., Kempisty B., and Dyszkiewicz‐Konwińska M., “Diagnostic Test Accuracy of Artificial Intelligence in Detecting Periapical Periodontitis on Two‐Dimensional Radiographs: A Retrospective Study and Literature Review,” Medicina (Kaunas) 59 (2023): 768. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Ragab M. and Asar T. O., “Deep Transfer Learning With Improved Crayfish Optimization Algorithm for Oral Squamous Cell Carcinoma Cancer Recognition Using Histopathological Images,” Scientific Reports 14 (2024): 25348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Ayhan B., Ayan E., and Atsü S., “Detection of Dental Caries under Fixed Dental Prostheses by Analyzing Digital Panoramic Radiographs With Artificial Intelligence Algorithms Based on Deep Learning Methods,” BMC Oral Health 25 (2025): 216. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Turosz N., Chęcińska K., Chęciński M., Lubecka K., Bliźniak F., and Sikora M., “Artificial Intelligence (AI) Assessment of Pediatric Dental Panoramic Radiographs (DPRs): A Clinical Study,” Pediatric Reports 16 (2024): 794–805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Zhang B., Li Y., Shi J., Liu S., and Liu C., “Deep Learning for Imaging Diagnosis of Jaw Cystic Lesions and Maxillofacial Tumors: A Narrative Review,” Journal of International Medical Research 53 (2025): 3000605251404778. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Esmaeilyfard R., Esmaeeli N., and Paknahad M., “An Artificial Intelligence Mechanism for Detecting Cystic Lesions on CBCT Images Using Deep Learning,” Journal of Stomatology, Oral and Maxillofacial Surgery 126 (2025): 102152. [DOI] [PubMed] [Google Scholar]
- 22. Liang B., Qin H., Nong X., and Zhang X., “Classification of Ameloblastoma, Periapical Cyst, and Chronic Suppurative Osteomyelitis With Semi‐Supervised Learning: The WaveletFusion‐ViT Model Approach,” Bioengineering (Basel) 11 (2024): 571. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Liu X. H., Zhong N. N., Yi J. R., Lin H., Liu B., and Man Q. W., “Trends in Research of Odontogenic Keratocyst and Ameloblastoma,” Journal of Dental Research 104 (2025): 347–368. [DOI] [PubMed] [Google Scholar]
- 24. Zhang X., Yang Y., Zhong C., Li J., and Li G., “Segmentation‐Guided Preprocessing Improves Deep Learning Diagnostic Accuracy and Confidence of Ameloblastoma and Odontogenic Keratocyst in Cone Beam CT Images—A Preliminary Study,” Diagnostics (Basel) 16 (2026): 416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Pacheco A. H. S., Souza M. R. F., Valeriano A. T., et al., “Molecular Insights Into Epithelial Detachment in Odontogenic Keratocyst: The Role of Matrix Metalloproteinases and Effects of Marsupialization,” Journal of Oral Pathology and Medicine 8 (2026): 274–279. [DOI] [PubMed] [Google Scholar]
- 26. Mirzaee S., Shoorgashti R., Sadri D., and Farhadi S., “Comparison of EGFR Expression in Ameloblastoma and Odontogenic Keratocyst,” Journal of Research in Dental and Maxillofacial Sciences 8 (2023): 274–279. [Google Scholar]
- 27. Shoorgashti R., Sadri D., and Farhadi S., “Expression of Epidermal Growth Factor Receptor by Odontogenic Cysts: A Comparative Study of Dentigerous Cyst and Odontogenic Keratocyst,” Journal of Research in Dental Sciences 17 (2020): 127–136. [Google Scholar]
- 28. Shoorgashti R., Moshiri A., and Lesan S., “Evaluation of Oral Mucosal Lesions in Iranian Smokers and Non‐Smokers,” Nigerian Journal of Clinical Practice 27 (2024): 467–474. [DOI] [PubMed] [Google Scholar]
- 29. Mirhosseini N., Shoorgashti R., and Lesan S., “The Evaluation of Clinical Factors Affecting Oral Health Impacts on the Quality of Life of Iranian Elderly Patients Visiting Dental Clinics: A Cross‐Sectional Study,” Special Care in Dentistry 44 (2024): 1219–1227. [DOI] [PubMed] [Google Scholar]
- 30. Alfurhud A. A., “Radiographic Reduction Following Decompression of a Dentigerous Cyst and An Odontogenic Keratocyst: A Comparative Case Report,” Science in Progress 109 (2026): 368504261422277. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Van Cleemput T., Jackers X., Piagkou M., and Politis C., “Recurrence Patterns of Odontogenic Keratocysts in Syndromic and Non‐Syndromic Patients,” Journal of Maxillofacial and Oral Surgery 23 (2024): 152–158. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Dioguardi M., Quarta C., Sovereto D., et al., “Factors and Management Techniques in Odontogenic Keratocysts: A Systematic Review,” European Journal of Medical Research 29 (2024): 287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Gonçalves T. O. F., Rangel R. M. R., Marañón‐Vásquez G. A., et al., “Management and Recurrence of the Odontogenic Keratocyst: An Overview of Systematic Reviews,” Oral and Maxillofacial Surgery 28 (2024): 1457–1478. [DOI] [PubMed] [Google Scholar]
- 34. Bera R. N., Tandon S., Tiwari P., and Mishra M., “Recurrence and Prognosticators of Recurrence in Odontogenic Keratocyst of the Jaws,” Journal of Maxillofacial and Oral Surgery 23 (2024): 1304–1315. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Sim S. Y., Hwang J., Ryu J., Kim H., Kim E. J., and Lee J. Y., “Differential Diagnosis of OKC and SBC on Panoramic Radiographs: Leveraging Deep Learning Algorithms,” Diagnostics (Basel) 14 (2024): 1144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Kise Y., Ariji Y., Kuwada C., Fukuda M., and Ariji E., “Effect of Deep Transfer Learning With a Different Kind of Lesion on Classification Performance of Pre‐Trained Model: Verification With Radiolucent Lesions on Panoramic Radiographs,” Imaging Science in Dentistry 53 (2023): 27–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Lee A., Kim M. S., Han S. S., Park P., Lee C., and Yun J. P., “Deep Learning Neural Networks to Differentiate Stafne's Bone Cavity From Pathological Radiolucent Lesions of the Mandible in Heterogeneous Panoramic Radiography,” PLoS One 16 (2021): e0254997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Song Y., Ma S., Mao B., et al., “Application of Machine Learning in the Preoperative Radiomic Diagnosis of Ameloblastoma and Odontogenic Keratocyst Based on Cone‐Beam CT,” Dento Maxillo Facial Radiology 53 (2024): 316–324. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Fedato Tobias R. S., Teodoro A. B., Evangelista K., et al., “Diagnostic Capability of Artificial Intelligence Tools for Detecting and Classifying Odontogenic Cysts and Tumors: A Systematic Review and Meta‐Analysis,” Oral Surgery, Oral Medicine, Oral Pathology and Oral Radiology 138 (2024): 414–426. [DOI] [PubMed] [Google Scholar]
- 40. Sha X., Wang C., Sun J., et al., “CBCT Radiomics Features Combine Machine Learning to Diagnose Cystic Lesions in the Jaw,” Dento Maxillo Facial Radiology 54 (2025): 381–388. [DOI] [PubMed] [Google Scholar]
- 41. Chai Z. K., Mao L., Chen H., et al., “Improved Diagnostic Accuracy of Ameloblastoma and Odontogenic Keratocyst on Cone‐Beam CT by Artificial Intelligence,” Frontiers in Oncology 11 (2022): 793417. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Bispo, Pierre Júnior M. S., Apolinário M., et al., “Computer Tomographic Differential Diagnosis of Ameloblastoma and Odontogenic Keratocyst: Classification Using a Convolutional Neural Network,” Dento Maxillo Facial Radiology 50 (2021): 20210002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Page M. J., McKenzie J. E., Bossuyt P. M., et al., “The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,” BMJ 372 (2021): n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. McInnes M. D. F., Moher D., Thombs B. D., et al., “Preferred Reporting Items for a Systematic Review and Meta‐Analysis of Diagnostic Test Accuracy Studies: The PRISMA‐DTA Statement,” Journal of the American Medical Association 319 (2018): 388–396. [DOI] [PubMed] [Google Scholar]
- 45. Whiting P. F., Rutjes A. W., Westwood M. E., et al., “QUADAS‐2: A Revised Tool for the Quality Assessment of Diagnostic Accuracy Studies,” Annals of Internal Medicine 155 (2011): 529–536. [DOI] [PubMed] [Google Scholar]
- 46. Muraoka H., Kaneda T., Ito K., Otsuka K., and Tokunaga S., “Machine Learning Approach Using Radiomics Features to Distinguish Odontogenic Cysts and Tumours,” International Journal of Oral and Maxillofacial Surgery 55 (2025): 182–187. [DOI] [PubMed] [Google Scholar]
- 47. Committeri U., Barone S., Arena A., et al., “New Perspectives in the Differential Diagnosis of Jaw Lesions: Machine Learning and Inflammatory Biomarkers,” Journal of Stomatology, Oral and Maxillofacial Surgery 125 (2024): 101912. [DOI] [PubMed] [Google Scholar]
- 48. Liu W., Li X., Liu C., et al., “Automatic Classification and Segmentation of Multiclass Jaw Lesions in Cone‐Beam CT Using Deep Learning,” Dento Maxillo Facial Radiology 53 (2024): 439–446. [DOI] [PubMed] [Google Scholar]
- 49. Yeshua T., Ladyzhensky S., Abu‐Nasser A., et al., “Deep Learning for Detection and 3D Segmentation of Maxillofacial Bone Lesions in Cone Beam CT,” European Radiology 33 (2023): 7507–7518. [DOI] [PubMed] [Google Scholar]
- 50. Lee J. H., Kim D. H., and Jeong S. N., “Diagnosis of Cystic Lesions Using Panoramic and Cone Beam Computed Tomographic Images Based on Deep Learning Neural Network,” Oral Diseases 26 (2020): 152–158. [DOI] [PubMed] [Google Scholar]
- 51. Yilmaz E., Kayikcioglu T., and Kayipmaz S., “Computer‐Aided Diagnosis of Periapical Cyst and Keratocystic Odontogenic Tumor on Cone Beam Computed Tomography,” Computer Methods and Programs in Biomedicine 146 (2017): 91–100. [DOI] [PubMed] [Google Scholar]
- 52. Abdolali F., Zoroofi R. A., Otake Y., and Sato Y., “Automated Classification of Maxillofacial Cysts in Cone Beam CT Images Using Contourlet Transformation and Spherical Harmonics,” Computer Methods and Programs in Biomedicine 139 (2017): 197–207. [DOI] [PubMed] [Google Scholar]
- 53. Huang Z., Li B., Cheng Y., and Kim J., “Odontogenic Cystic Lesion Segmentation on Cone‐Beam CT Using an Auto‐Adapting Multi‐Scaled UNet,” Frontiers in Oncology 14 (2024): 1379624. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Shrivastava P. K., Hasan S., Abid L., Injety R., Shrivastav A. K., and Sybil D., “Accuracy of Machine Learning in the Diagnosis of Odontogenic Cysts and Tumors: A Systematic Review and Meta‐Analysis,” Oral Radiology 40 (2024): 342–356. [DOI] [PubMed] [Google Scholar]
- 55. Archana R. and Jeevaraj P. S. E., “Deep Learning Models for Digital Image Processing: A Review,” Artificial Intelligence Review 57 (2024): 11. [Google Scholar]
- 56. Karakullukcu E., “Leveraging Convolutional Neural Networks for Image‐Based Classification of Feature Matrix Data,” Expert Systems with Applications 281 (2025): 127625. [Google Scholar]
- 57. Chai J., Zeng H., Li A., and Ngai E. W. T., “Deep Learning in Computer Vision: A Critical Review of Emerging Techniques and Application Scenarios,” Machine Learning with Applications 6 (2021): 100134. [Google Scholar]
- 58. Hosny A., Parmar C., Quackenbush J., Schwartz L. H., and Aerts H., “Artificial Intelligence in Radiology,” Nature Reviews Cancer 18 (2018): 500–510. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Bento N., Rebelo J., Barandas M., et al., “Comparing Handcrafted Features and Deep Neural Representations for Domain Generalization in Human Activity Recognition,” Sensors (Basel) 22 (2022): 7324. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Rahman H., Khan A. R., Sadiq T., Farooqi A. H., Khan I. U., and Lim W. H., “A Systematic Literature Review of 3D Deep Learning Techniques in Computed Tomography Reconstruction,” Tomography 9 (2023): 2158–2189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Hu C., Cao N., Li X., He Y., and Zhou H., “CBCT‐to‐CT Synthesis Using a Hybrid U‐Net Diffusion Model Based on Transformers and Information Bottleneck Theory,” Scientific Reports 15 (2025): 10816. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Hughes C. M. L., Zhang Y., Pourhossein A., and Jurasova T., “A Comparative Analysis of Binary and Multi‐Class Classification Machine Learning Algorithms to Detect Current Frailty Status Using the English Longitudinal Study of Ageing (ELSA),” Frontiers in Aging 6 (2025): 1501168. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Gao X., He Y., Zhang M., et al., “A Multiclass Classification Using One‐Versus‐All Approach With the Differential Partition Sampling Ensemble,” Engineering Applications of Artificial Intelligence 97 (2021): 104034. [Google Scholar]
- 64. Jahanshahi A., Soleymani Y., Fazel Ghaziani M., and Khezerloo D., “Radiomics Reproducibility Challenge in Computed Tomography Imaging as a Nuisance to Clinical Generalization: A Mini‐Review,” Egyptian Journal of Radiology and Nuclear Medicine 54 (2023): 83. [Google Scholar]
- 65. Fooladi M., Soleymani Y., Rahmim A., et al., “Impact of Different Reconstruction Algorithms and Setting Parameters on Radiomics Features of PSMA PET Images: A Preliminary Study,” European Journal of Radiology 172 (2024): 111349. [DOI] [PubMed] [Google Scholar]
- 66. Mali S. A., Ibrahim A., Woodruff H. C., et al., “Making Radiomics More Reproducible Across Scanner and Imaging Protocol Variations: A Review of Harmonization Methods,” Journal of Personalized Medicine 11 (2021): 842. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67. Gang G. J., Deshpande R., and Stayman J. W., “Standardization of Histogram‐ and Gray‐Level Co‐Occurrence Matrices‐Based Radiomics in the Presence of Blur and Noise,” Physics in Medicine and Biology 66 (2021): 074004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Turosz N., Chęcińska K., Chęciński M., Brzozowska A., Nowak Z., and Sikora M., “Applications of Artificial Intelligence in the Analysis of Dental Panoramic Radiographs: An Overview of Systematic Reviews,” Dento Maxillo Facial Radiology 52 (2023): 20230284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Khubrani Y. H., Thomas D., Slator P. J., White R. D., and Farnell D. J. J., “Detection of Periodontal Bone Loss and Periodontitis From 2D Dental Radiographs via Machine Learning and Deep Learning: Systematic Review Employing APPRAISE‐AI and Meta‐Analysis,” Dento Maxillo Facial Radiology 54 (2025): 89–108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Negi S., Mathur A., Tripathy S., et al., “Artificial Intelligence in Dental Caries Diagnosis and Detection: An Umbrella Review,” Clinical and Experimental Dental Research 10 (2024): e70004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Zhu J., Chen Z., Zhao J., et al., “Artificial Intelligence in the Diagnosis of Dental Diseases on Panoramic Radiographs: A Preliminary Study,” BMC Oral Health 23 (2023): 358. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Ibraheem W. I., Jain S., Ayoub M. N., et al., “Assessment of the Diagnostic Accuracy of Artificial Intelligence Software in Identifying Common Periodontal and Restorative Dental Conditions (Marginal Bone Loss, Periapical Lesion, Crown, Restoration, Dental Caries) in Intraoral Periapical Radiographs,” Diagnostics (Basel) 15 (2025): 1432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Albano D., Galiano V., Basile M., et al., “Artificial Intelligence for Radiographic Imaging Detection of Caries Lesions: A Systematic Review,” BMC Oral Health 24 (2024): 274. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Pul U. and Schwendicke F., “Artificial Intelligence for Detecting Periapical Radiolucencies: A Systematic Review and Meta‐Analysis,” Journal of Dentistry 147 (2024): 105104. [DOI] [PubMed] [Google Scholar]
- 75. Kim K., Lim C. Y., Shin J., Chung M. J., and Jung Y. G., “Enhanced Artificial Intelligence‐Based Diagnosis Using Cbct With Internal Denoising: Clinical Validation for Discrimination of Fungal Ball, Sinusitis, and Normal Cases in the Maxillary Sinus,” Computer Methods and Programs in Biomedicine 240 (2023): 107708. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data sets generated and analyzed during the current research are available from the corresponding author on reasonable request.
All authors have read and approved the final version of the manuscript. Dr. Reyhaneh Shoorgashti had full access to all of the data in this study and takes complete responsibility for the integrity of the data and the accuracy of the data analysis.
