Abstract
Background
Furcation involvement complicates the management of periodontitis and increases the risk of tooth loss. Conventional methods of detection, such as probing and two‐dimensional radiographs, are limited by operator variability and anatomical complexity. Deep learning has shown a potential to detect furcation involvement on radiographic images. The aim of this review was to systematically evaluate the diagnostic potentials of deep learning models in detecting furcation involvement on radiographic images.
Methods
Systematic search was conducted in PubMed, EMBASE, CENTRAL, ClinicalTrials.gov and ProQuest for studies published from 2010 to September 2025. Two reviewers independently screened studies, extracted data, and assessed quality using QUADAS‐2. Diagnostic metrics (sensitivity, specificity, F1‐score, area under the curve (AUC)) were pooled using random‐effects meta‐analysis. Heterogeneity and publication bias were assessed via I 2 statistics, meta‐regression, and funnel plots.
Results
Eight studies, including 7814 radiographs of 12,373 molars (periapical, panoramic, cone‐beam computed tomography), were analyzed. Deep learning models demonstrated high accuracy: sensitivity 0.93, specificity 0.94, diagnostic odds ratio (DOR) 187, AUC 0.97 with mandibular molars reflecting higher accuracy (sensitivity 0.96, specificity 0.97, DOR 631, AUC 0.99). Fagan plot analysis indicated strong clinical utility. Meta‐regression showed no significant effect of dataset type, augmentation, or number of annotators. No publication bias was detected.
Conclusion
Deep learning models show promising accuracy in detecting furcation involvement, particularly in mandibular molars, comparable to expert clinicians. Further refinement with larger, diverse datasets is needed to reduce false positives and enable safe clinical integration.
Plain language summary
Furcation involvement, a condition where the bone between the roots of a tooth is lost, makes managing gum disease more difficult and increases the risk of tooth loss. Clinical probing remains a reliable and essential method for detecting this condition, while dental X‐rays can provide complementary information, particularly in complex cases. This study reviewed the use of deep learning, a type of artificial intelligence (AI), to detect furcation involvement on dental X‐rays. Data from eight studies, including more than 7,800 dental radiographs of >12,000 molars, were analyzed and the results showed that deep learning models were highly accurate, performing similarly to expert dentists. Accuracy was slightly higher for lower jaw (mandibular) molars. The study suggests that these AI tools could help dentists detect furcation involvement more reliably. However, more research with larger and more diverse datasets is needed before these tools can be safely used in everyday dental practice.
Keywords: artificial intelligence, furcation defect, meta‐analysis, systematic review
1. INTRODUCTION
Periodontitis is a common chronic inflammatory disease characterized by progressive destruction of the supporting structures of the teeth, which can lead to tooth loss, if left untreated. 1 , 2 Accurate diagnosis of bone loss patterns, including horizontal and vertical defects, as well as furcation involvement, is critical for staging, treatment planning, and prognosis of periodontal diseases. 3 , 4 The 2018 classification of periodontal diseases emphasizes severity and complexity, with stage III and stage IV periodontitis associated with furcation involvement, deep attachment loss, and increased risk of tooth loss. 3 , 4
Furcation involvement, defined as the loss of alveolar bone between the roots of multirooted teeth, remains one of the most challenging conditions to manage. 5 The furcation area is anatomically difficult to access and, in many cases, cannot be effectively cleaned using conventional methods. The clinical impact of furcation involvement is substantial. A recent registry‐based cohort study of >2.3 million molars followed for 10 years found that molars with class II furcation involvement exhibited a 22% tooth loss rate, whereas those with class III furcation involvement showed nearly 46% tooth loss. 6 These findings highlight that furcation involvement not only complicates treatment but also markedly undermines long‐term prognosis. 7 , 8
Conventional methods for detecting furcation involvement, such as clinical probing and two‐dimensional radiographs, are limited by operator variability, anatomical complexity, and their inability to fully depict the three‐dimensional bone morphology. 9 , 10 Cone‐beam computed tomography (CBCT), on the other hand, provides enhanced accuracy, but its routine use is constrained by cost, radiation exposure, and accessibility. 11 , 12 These limitations underscore the need for more reliable, efficient, and widely applicable diagnostic aids. Artificial intelligence (AI), particularly convolutional neural networks (CNNs) 13 has demonstrated remarkable performance in medical imaging, including detection of pulmonary nodules, coronary artery calcifications, cerebral aneurysms, and colon polyps. 14 , 15 , 16 In dentistry, AI has been applied in caries detection, 17 assessment of alveolar bone loss, 18 , 19 identification of periapical lesions, 20 and diagnosis of maxillofacial cysts and tumors. 21 , 22 Recently, AI‐based models have been explored for periodontal imaging, including detection of bone loss patterns and furcation involvement, showing promising results. 23 , 24 , 25
Deep learning is a subset of AI that uses multilayered artificial neural networks to automatically learn complex patterns and hierarchical image features from large datasets without the need for manual feature extraction. 26 Among deep learning architectures, CNNs are most commonly applied in medical and dental imaging because of their ability to recognize spatial patterns and structures within radiographs or CBCT scans. These models are trained using large annotated image datasets and can then predict the presence or absence of disease features, such as furcation involvement, on unseen images with minimal human intervention.
Despite these advances, AI‐based studies for the detection of furcation involvement remain heterogeneous in their methodologies, imaging modalities, and outcome measures. Only few have aligned their diagnostic outputs with the 2018 classification, limiting their direct application in clinical practice. Given the high prevalence and prognostic importance of furcation involvement in stage III and stage IV periodontitis, a systematic evaluation of AI in this context is both relevant and timely. This review, therefore, aimed to identify and critically appraise studies applying deep learning models to radiographic images for detecting furcation involvement, and to quantitatively evaluate the diagnostic performance of these models.
2. MATERIALS AND METHODS
This systematic review was conducted according to the Preferred Reporting Items for Systematic Reviews and Meta‐analyses of Diagnostic Test Accuracy Studies (PRISMA‐DTA) guidelines and the recommendations of the Cochrane Collaboration. 27 , 28 The review was registered with the National Institute for Health Research (NHR) under PROSPERO ID CRD420251137286. As this review analyzed previously published data, ethical approval was not required.
2.1. Focused question (PICO framework)
A focused research question was formulated to evaluate the diagnostic performance of deep learning models in detecting furcation involvement on dental radiographs:
Population (P): Adults whose dental radiographs were assessed for furcation involvement.
Intervention (I): Application of deep learning models to detect furcation involvement on periapical, panoramic or CBCT images.
Control: Clinical diagnosis by experienced clinicians interpreting the same radiographs.
Outcomes:
Primary outcomes: Sensitivity (recall) and specificity.
Secondary outcomes: Positive predictive value (precision), negative predictive value, F1 score, and the summary receiver operating characteristic (SROC) curve.
2.2. Search strategy and study selection
The search strategy followed established methodological guidelines. 27 , 29 A comprehensive literature search was conducted across multiple electronic databases, including MEDLINE (via PubMed), EMBASE, The Cochrane Central Register of Controlled Trials (CENTRAL), ClinicalTrials.gov, and ProQuest Dissertation and Theses Global, for both published and unpublished studies up to September 01, 2025 (Table S1 in the online Journal of Periodontology). To ensure the inclusion of studies relevant to deep learning, the search was restricted to publications from 2010 onward, coinciding with the rise of deep learning in image analysis around 2012. 30
Both published and unpublished studies were considered, including retrospective studies using deep learning models to assess periapical, panoramic radiographs, or CBCT for furcation involvement. Only peer‐reviewed studies with sufficient methodological detail for data extraction were included. There were no language restrictions. Studies were excluded if they were not peer‐reviewed or did not report key diagnostic accuracy measures.
Two reviewers (M.A. and N.A.) independently performed the search and screening in duplicate to minimize bias and ensure reproducibility. In addition, the reference lists of all included and potentially eligible studies were manually examined to identify further relevant literature. A hand search was also conducted for the past 5 years in leading dental journals, including Clinical Implant Dentistry and Related Research, Clinical Oral Implants Research, International Journal of Oral and Maxillofacial Implants, International Journal of Periodontics and Restorative Dentistry, Journal of Clinical Periodontology, Journal of Periodontal Research, and Journal of Periodontology.
Two reviewers (M.A. and N.A.) independently screened all retrieved records in duplicate to identify studies that met the predefined inclusion criteria. The screening was conducted in two phases: an initial evaluation of titles, abstracts, and keywords, followed by a full‐text review of potentially relevant articles using a standardized eligibility checklist. Any discrepancies during the screening process were resolved through discussion and, when necessary, a third author (A.H.) was consulted to reach consensus. In instances where multiple publications reported on the same study, the most detailed and informative version was selected for inclusion. All excluded full‐text articles were documented, along with the reasons for exclusion.
2.3. Data collection
Two reviewers (M.A. and N.A.) independently extracted data from all included studies using a standardized data extraction form. The following categories of information were collected: (1) Study characteristics: title, authors’ names and contact details, study location, publication language, year of publication, publication status (published or unpublished), funding source, and study design. (2) Dataset details: inclusion and exclusion criteria, number of radiographs and teeth analyzed, and the number and reasons for any exclusions or dropouts. (3) Diagnostic assessment: number of radiographs and teeth assessed for furcation involvement using the deep learning model. (4) Diagnostic accuracy measures: counts of true positives, false positives, false negatives and true negatives. (5) Technical specifications: annotation tools and procedures used, type of computer vision task (for example: classification, detection, segmentation), and details of deep learning model including the backbone architecture. All extracted data were independently cross‐checked by both reviewers (M.A. and N.A.). Any inconsistencies were resolved through discussion, and if consensus could not be reached, a third reviewer (A.H.) was consulted. For each study, we extracted the raw numbers of true positives, false positives, true negatives, false negatives to calculate pooled sensitivity and specificity, ensuring that the meta‐analysis reflected the actual performance of the deep learning models.
2.4. Quality assessment
The methodological quality of the included studies was assessed using the revised Quality Assessment of Diagnostic Accuracy Studies (QUADAS‐2). 31 This framework evaluates the risk of bias across four main domains: dataset selection, index test, reference standard, and flow and timing. In addition, concerns regarding the applicability of each study were assessed within the first three domains (dataset selection, index test, and reference standard) (Table S2 in the online Journal of Periodontology). To account for the specific challenges of AI research in medical imaging, several domain‐specific questions were adapted to better assess factors such as dataset representativeness, selection bias, and AI‐specific risks like overfitting or data leakage. These modifications also aimed to evaluate the transparency and completeness of reporting related to model development, validation, and testing; for example: studies were rated as having a high risk of bias in dataset selection if they lacked details on data collection procedures, relied on images from a single dental clinic, or failed to describe their validation strategy. In the index test domain, high risk was assigned when models were not reproducible or lacked critical information on their development and evaluation. A high risk in the reference standard domain was noted when the ground truth was poorly defined or determined by only a single assessor. Lastly, the flow and timing domain was judged as high risk if different reference standards were used, not all cases were included in the analysis, or the timing between the AI test and reference diagnosis was inappropriate.
Two reviewers (M.A. and N.A.) independently rated each study for risk of bias and applicability concerns, classifying each domain as low, high, or unclear risk. Discrepancies were resolved through discussion, with input from a third reviewer (A.H.) when needed to reach a consensus.
2.5. Statistical analysis and data synthesis
For each included study, the numbers of true positives, true negatives, false positives, and false negatives were extracted and structured into 2 × 2 contingency tables. These tables served as the basis for calculating key diagnostic performance metrics, including sensitivity (recall), specificity, positive predictive value (precision), negative predictive value, F1 score, positive likelihood ratio (LR+), negative likelihood ratio (LR‐), and diagnostic odds ratio (DOR). To manage instances of zero values in any cell, a continuity correction of 0.5 was applied. 32 LRs were used due to their reduced dependence on disease prevalence and greater clinical interpretability. Values of LR+ above 10 and LR‐ below 0.1 were interpreted as strong indicators of diagnostic accuracy. 33 A random‐effects meta‐analysis was conducted to pool sensitivity, specificity, and LR estimates along with their 95% confidence intervals (CIs). Between‐study heterogeneity was evaluated using forest plots, the Cochran Q test, and the I 2 statistic, with a p value < 0.10 and an I 2 > 50 indicating substantial heterogeneity. 34 To identify potential sources of heterogeneity, meta‐regression analyses were conducted using factors such as dataset characteristics, use of data augmentation, and imaging modality. A sensitivity analysis was conducted by excluding studies rated at high risk to evaluate the robustness of the pooled diagnostic performance estimates. Additionally, a Fagan nomogram was used to estimate post‐test probabilities across varying pre‐test probability levels, providing a visual summary of the clinical relevance of the evaluated deep learning models.
Data derived from periapical, panoramic, and CBCT imaging were pooled in the meta‐analysis, as these modalities are routinely used for radiographic evaluation of furcation involvement and were analyzed using deep learning models addressing the same diagnostic question. Combining results across imaging types provided a broader assessment of overall diagnostic performance and enhanced the robustness of the pooled estimates. While differences in dimensionality and resolution exist among these modalities, their common diagnostic focus justified the integrated analysis. Potential variability associated with imaging modality was further explored through meta‐regression analyses.
An SROC curve was generated to summarize diagnostic performance across studies in terms of sensitivity and specificity. The diagnostic performance of the models was quantified by calculating the area under the curve (AUC), with values interpreted using standard thresholds: 0.5–0.7 indicating low accuracy, 0.7–0.9 moderate accuracy, 0.9–0.99 high accuracy, and 1.0 representing perfect diagnostic performance. 35 The DOR, a single indicator of test performance, was also calculated. It represents the odds of a positive test result in individuals with furcation involvement compared with those without the condition. Higher DOR values indicate better discriminatory ability, ranging from 0 to infinity.
Studies reporting diagnostic metrics at either the image level or tooth level were included. For this meta‐analysis, the tooth was used as the unit of analysis. Publication bias was not assessed because the power to detect publication bias was low (less than 10 papers). All statistical analyses were performed using the MIDAS module in Stata/MP version 14 (StataCorp LLC, College Station, TX, USA), following established methodological guidelines. 36 , 37
3. RESULTS
3.1. Characteristics of study settings
The initial search identified 198 studies (Figure 1). Titles and abstracts were independently screened in duplicate by two reviewers (M.A. and N.A.). Twelve full‐text articles were retrieved for detailed assessment. 25 , 38 , 39 , 40 , 41 , 42 , 43 , 44 , 45 , 46 , 47 , 48 Four studies were excluded after full‐text review. 38 , 40 , 43 , 46 Three of these 38 , 40 , 46 were excluded for not specifically reporting furcation involvement. Additional information was requested from the corresponding author of one study, 43 but it was excluded due to non‐response. Ultimately, eight studies 25 , 39 , 41 , 42 , 44 , 45 , 47 , 48 met the inclusion criteria and were included in the review (Table 1). No additional studies were identified through hand searching.
FIGURE 1.

Flowchart of the search process.
TABLE 1.
Characteristics of included studies.
| Karadsheh and Zabadi 2023 | Kurt‐Bayrakdar et al. 2025 | Kurt‐Bayrakdar et al. 2024 | Mao et al. 2023 | Shetty et al. 2024 | Tajima et al. 2021 | Vilkomir et al. 2024 | Zhang et al. 2025 | |
|---|---|---|---|---|---|---|---|---|
| Aim of study | Develop a deep learning model to detect periodontal bone loss and furcation involvement on periapical radiographs | Develop a deep learning model to detect tooth presence, tooth numbering, periodontal bone defects and furcation involvement on CBCT images | Develop a deep learning model to detect periodontal bone loss and furcation involvement on panoramic radiographs | Detect furcation involvement on periapical radiographs | Assess the accuracy of a deep learning model in detecting furcation involvement on axial CBCT images | Develop a deep learning model to detect furcation involvement on panoramic radiographs | Develop a deep learning algorithm to classify furcation involvement on periapical radiographs | Evaluate the performance of deep learning model in detecting furcation involvement on panoramic radiographs |
| Location | University of Jordan, Amman, Jordan | University of Mississippi Medical Center School of Dentistry, Jackson, United States and Osmangazi University, Faculty of Dentistry, Eskisehir, Turkiye | Osmangazi University, Faculty of Dentistry, Eskisehir, Turkiye | Chang Gung Memorial Hospital, Taoyuan City, Taiwan | University of Sharjah, Sharjah, United Arab Emirates | A total of 19 dental practices including AOI International Hospital, Kawasaki, Kanagawa, Japan | East Carolina University, Greenville, North Carolina, USA | Nanjing Stomatological Hospital, Nanjing University, Nanjing, China |
|
Jaw(s) |
Mandible and maxilla | Mandible and maxilla | Mandible and maxilla | Mandible |
Mandible |
Mandible | Mandible | Mandible and maxilla |
| Age of participants (years) | NR | > 18 | NR | NR | 18‐60 | NR | NR | 20‐70 |
| Index and reference standard | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians | Deep learning model and experienced clinicians |
| Imaging modality | Periapical radiographs | CBCT | Panoramic radiographs | Periapical radiographs | CBCT | Panoramic radiographs | Periapical radiographs | Panoramic radiographs |
| Dataset size (radiographs/teeth) |
1300/2324 Test: 165/310 |
502/502 Training: 400/400 Validation: 52 Test: NR/50 |
1941/2815 Training: 1619/2358 Validation: 161/227 Test: 161/230 |
1500/1500 Training: 960/960 Validation: 240/240 Test: 300/300 |
285/285 Augmented set: 600 used for both training and validation Test set: 85/85 |
702/2552 Augmented set: 10640 images (9044 for training and 1596 for validation) Test set: 170/618 |
1078/1078 Training: 911/911 Validation: 82/82 Test: 85/85 |
506/1568 Augmented set: 7840 images Training: 1096 Validation: 234 Test: 238/238 |
| Annotation method | Segmentation‐based manual labeling | Manual segmentation of CBCT images | Segmentation‐based manual labeling | Image masking to define ROI | Image‐level binary labeling | Bounding box | Image‐level binary labeling | Image‐level binary labeling |
| Annotation tool | Web‐based labeling software (Supervisely OU, Tallinn, Estonia) | Web‐based labeling software (CranioCatch, Eskisehir, Turkey) | NR | NR | NR | NR | NR | NR |
| Annotators | Two experienced dentists | Three experienced oral and maxillofacial radiologists | Three periodontists and one oral and maxillofacial radiologist | Two periodontists | Single examiner with 15 years of clinical experience | Two Experienced dentists | Three experienced dentists and verified by oral and maxillofacial radiologist | Two periodontists and verified by a third periodontist |
| Data augmentation | No | NR | NR | Yes | Yes | Yes | Yes | Yes |
| Model type | CNN‐based classification model | nnU‐Net v2 CNN model | CNN‐based U‐Net | CNN‐based classification model | CNN‐based classification model | Deep CNN classification model (YOLOv3) | CNN‐based binary classification model | ViT‐based model |
| Backbone network | ResNet/DenseNet | U‐Net | U‐Net | GoogLENet (Inception v1) | ResNet101V2 | Darknet‐53 | ResNet‐18 | ViT encoder |
| Sensitivity (recall) (%) | 88.9 | 54.0 | 89.2 | 95.6 | 97.0 | 95.5 | 95.0 |
Maxilla: 89.0 Mandible: 94.0 Overall: 92.0 |
| Specificity (%) | 75.0 | 85.7 | 73.3 | 94.6 | 100 | 97.0 | 98.0 |
Maxilla: 93.6 Mandible: 97.9 Overall: 97.0 |
| PPV (precision) (%) | 84.9 | 62.0 | 93.2 | 91.6 | 100 |
96.2 |
97.0 |
Maxilla: 96.0 Mandible: 99.0 Overall: 98.0 |
| NPV (%) | 81.1 | 81.1 | 62.3 | 97.2 | 97.7 | 96.7 | 96.0 |
Maxilla: 84.6 Mandible: 92.0 Overall: 89.0 |
| F1 score (%) | 86.7 | 57.7 | 91.2 | 93.6 | 98.0 | 95.8 | 96.2 |
Maxilla: 92.0 Mandible: 96.0 Overall: 95.0 |
|
TP: true positive FP: false positive TN: true negative FN: false negative |
169 30 90 21 |
8 5 30 7 |
165 12 33 20 |
109 10 176 5 |
41 0 43 1 |
261 10 335 12 |
38 1 44 2 |
Maxilla: 64 3 44 8 Mandible: 68 1 46 4 Overall: 132 3 92 11 |
Abbreviations: CBCT, cone‐beam computed tomography; CNN, convolutional neural network; ViT, vision transformer; ROI, region of interest; YOLO, you only look once; ResNet, residual network; DenseNet, densely connected convolutional network; PPV, positive predictive value; NPV, negative predictive value.
Of the eight included studies, three 25 , 41 , 48 reported receiving support from university or national funding bodies. Two studies 42 , 44 stated that no funding was received, while three 39 , 45 , 47 did not disclose funding information. One study 45 was published in Japanese, while the remaining studies were published in English. All included studies were conducted in university settings and collectively analyzed a dataset comprising 7,814 radiographic images with 12,373 molar teeth, including both healthy cases and those with furcation involvement.
3.2. Characteristics of dataset
Three types of radiographic modalities were used for assessing molar teeth. A total of 3878 periapical radiographs were analyzed across three studies, 25 , 39 , 47 3149 panoramic radiographs across three studies 42 , 45 , 48 and 787 CBCT images in two studies. 41 , 44 Four studies 39 , 41 , 42 , 48 included both maxillary and mandibular furcation involvement, while the remaining four 25 , 44 , 45 , 47 focused exclusively on mandibular molars.
With respect to periapical radiographic datasets, one study 39 used dental school records but without specifying the time frame; another 25 did not report the source of radiographs; while Vilkomir et al. 47 analyzed radiographs obtained between 2011 and 2023 at a dental school, using similar radiographic units, XCP receptor‐holding devices and standardized exposure parameters to ensure accuracy and reproducibility. For panoramic radiographs, Kurt‐Bayrakdar et al. 42 used images from a single radiography device at a dental school under standardized conditions, Tajima et al. 45 analyzed a multicenter dataset collected from 19 dental practices using various machines, and Zhang et al. 48 analyzed images obtained between 2020 and 2022 at a dental hospital using the same device. Regarding CBCT, Kurt‐Bayrakdar et al. 41 analyzed archived datasets from two dental schools, whereas Shetty et al. 44 used images obtained with a single high‐resolution CBCT unit at a dental school.
The included datasets comprised both patients with periodontal disease and periodontally healthy individuals. 41 Participants ages ranged from >18, 41 , 42 18–60 44 to 20–70 years, 48 with no sex restrictions. 41 , 42 Healthy periodontium was defined as alveolar bone crest up to 2 mm apical to an imaginary line through the cemento‐enamel junction. 42 Rigorous exclusion criteria were applied, including radiographs with dense artifacts 41 , 42 , 48 , poor quality due to patient positioning or movement, 41 , 42 history of orthognathic treatment, 41 , 42 bone metabolism disorders, 42 unusual alveolar bone morphology such as cysts or tumors, 42 cleft lip and palate, 42 crowding‐induced blurring of the alveolar bone region, 42 metal restorations causing diagnostic artifacts, 42 , 44 , 48 root fusion, 48 endodontic treatment, 44 and caries. 44 , 48
Images were anonymized and uploaded to the CranioCatch labeling platform (CranioCatch, Eskisehir, Turkey). 42 Furcation defects were segmented at the tooth level, delineating radiolucencies along root boundaries and lesion margins. 42 CBCT scans (Planmeca Promax 3D Mid, Planmeca, Finland; CS 9600, Carestream Dental, USA) 41 were exported in DICOM format and converted to NlfTI format using a custom Python‐based script, allowing analysis across sagittal, coronal, and axial planes. 41 Tooth segmentation and numbering were performed on axial slices and verified across all planes. Labeling consistency was high, with an Intersection over Union (IoU) > 0.90 and kappa values 0.81–0.99 across periodontal defect categories, including furcation involvement. 41
In Tajima et al., 45 patient identifiers (age, sex, ID, imaging time) were removed to comply with personal information protection guidelines. The images were randomly divided into 532 training and 170 evaluation images. Mandibular molar regions of interest were annotated according to the method of Kwon et al. 49 Preprocessing included noise reduction, contrast enhancement, and conversion of red‐green‐blue (RGB) images to grayscale, followed by Gaussian high‐pass filtering and adaptive thresholding to enhance segmentation quality for subsequent CNN training. 25
Shetty et al. 44 retrospectively screened 3,000 CBCT scans acquired using the Planmeca Viso G7 system, of which 285 scans (143 normal, 142 with furcation involvement) met inclusion criteria. High‐resolution acquisition settings (150 µm voxel size, 100 kVp, 12.5 mA, 5 s exposure), and field of view between Ø3×3–Ø6×6 cm were included. Axial slices were cropped (200 × 400 pixels) below the furcation to standardize inputs. The dataset was split into 85 test scans (43 normal, 42 abnormal), while the remaining 200 were augmented to create 600 images for training and validation. Images were resized to 224 × 224 pixels, and mild data augmentation (rotation ± 5°, zoom 0.1, nearest‐neighbor fill) produced a balanced dataset for ResNet101 v2 model training.
Vilkomir et al. 47 cropped periapical radiographs into individual tooth images (1644 × 643 pixels) and annotated them as “healthy” or “furcation involvement” by three examiners, cross‐verified with electronic patient records. The dataset was partitioned into training, validation, and testing subsets, with challenging cases included across sets to evaluate model robustness. Preprocessing standardized image size, normalized pixel values, and applied PyTorch‐based data augmentation (random flipping, shifting, scaling, shearing, and brightness/contrast adjustments), increasing diversity and reducing overfitting.
Finally, Zhang et al. 48 resized and normalized images for neural network input. Data augmentation, including rotations, zooms, flips, and shear transformations, expanded the dataset and ensured balance. Random partitioning into training, validation and testing sets, combined with cross‐validation, further enhanced model performance for deep learning‐based furcation involvement classification.
3.3. Characteristics of index test and reference standard
All eight included studies developed and validated deep learning models for the detection of furcation involvement, using reference standards established by experienced clinicians. In the studies by Kurt‐Bayrakdar et al. 41 , 42 , image labeling was performed by four observers (three periodontists and one oral and maxillofacial radiologist) after intra‐ and inter‐observer calibration, which involved duplicate labeling of ten radiographs one week apart. The labeled datasets were subsequently reviewed by three experienced oral and maxillofacial radiologists, and any cases without consensus were excluded from model training and validation. 41 In the study by Tajima et al. 45 , two dentists with more than 15 years of clinical experience annotated 702 radiographs for furcation involvement. In Vilkomir and co‐workers’ study, 47 the annotation process involved three experienced clinicians and was subsequently reviewed and verified by an oral and maxillofacial radiologist. Similarly, Zhang et al. 48 defined the reference standard based on classifications provided by two experienced periodontists, with a third expert resolving disagreements.
The datasets were created by pooling all radiographs with relevant labeled parameters and randomly dividing them into training (80%), validation (10%), and testing (10%) sets. For furcation involvement, Kurt‐Bayrakdar et al. 42 used 1619 cropped images (2358 labels) for training, 161 cropped images (227 labels) for validation, and 161 cropped images (230 labels) for testing. In a study by the same research group, Kurt‐Bayrakdar et al. 41 converted 502 CBCT images (251 healthy, 251 with periodontal disease) from DICOM to NlfTl format, with 400 images used for training, 52 for validation, and 50 for testing. To ensure reliability, 10% of the labeled data was reserved for testing, and these scans were completely excluded from training. The model was developed using a U‐Net framework for image segmentation and trained for 800 cycles on a high‐performance computer with a dedicated GPU. 41 , 42
Mao et al. 25 tested several transfer learning CNN models in Matlab, including GoogLeNet, AlexNet, Inception v3, and VGG19, with GoogLeNet performing best due to high accuracy and fewer parameters. Shetty et al. 44 used ResNet101v2 with transfer learning, training on augmented images for 30 epochs with the Adam optimizer. Tajima et al. 45 used YOLOv3 to detect furcation involvement on panoramic radiographs, using an 85/15 training‐validation split and a learning rate of 0.001. Vilkomir et al. 47 applied ResNet‐18 to classify teeth as healthy or with furcation involvement, using AdamW and early stopping. Zhang et al. 48 compared CNNs and Vision Transformer models, trained with Adam and regularization techniques to classify images while preventing overfitting.
3.4. Methodological quality
All included studies used comprehensive image datasets obtained from universities or private clinics, with clearly defined exclusion criteria. The datasets, which included periapical radiographs, panoramic radiographs, and CBCT scans, were generally appropriate for addressing the review question. However, one study 25 did not specify the source of the radiographs, limiting the generalizability of its findings. Additionally, two studies 39 , 45 did not report whether measures were taken to control overfitting, resulting in an unclear risk of bias in this domain (Figure S1 in the online Journal of Periodontology).
Deep learning models across the included studies were generally trained using standardized annotation protocols and without access to reference standard outcomes, thereby minimizing the risk of incorporation bias. However, one study did not apply data augmentation, 39 and two studies 41 , 42 did not report whether augmentation was used, raising concerns about the generalizability and reproducibility of their models. Consequently, these studies were rated as having an unclear risk of bias in this domain. Reference standards were typically established by expert clinicians, radiologists, or periodontists, through consensus and blinded assessments. An exception was one study, 44 in which a single dentist served as the reference standard without mention of blinding or consensus. Accordingly, this study was judged to have a high risk of bias in the reference standard domain.
All studies applied both the index test and reference standard to the full set of images without delays, exclusions, or dropouts, resulting in a low risk of bias in the flow and timing domain. Despite some methodological limitations, the applicability of the included studies was judged to be low concern across all domains, including dataset selection, index test and reference standard, as they were generally aligned with the review question. However, only one study 41 reported an a priori power calculation to justify the sample size. The absence of such calculations in the remaining studies may limit confidence in the robustness and generalizability of their results.
3.5. Results of meta‐analyses
The pooled sensitivity and specificity of deep learning models were high at 0.93 (95% CI 0.88–0.96) and 0.94 (95% CI 0.85–0.97), respectively (Figure 2). Correspondingly, the pooled LR+ was 14.51 (95% CI 5.93–35.48), and the LR‐ was 0.08 (95% CI 0.04–0.14), reflecting strong diagnostic performance. The models were highly effective both at confirming disease when positive and excluding disease when negative. The pooled DOR was 187 (95% CI 47–737), indicating excellent overall diagnostic accuracy, although the wide 95% CI suggests some uncertainty, likely due to heterogeneity among included studies. The AUC was 0.97 (95% CI 0.95–0.98) (Figure S2 in the online Journal of Periodontology), further confirming the very good performance of deep learning models in detecting furcation involvement.
FIGURE 2.

Forest plots of individual/pooled sensitivity and specificity of the included studies (CI, confidence interval; Q, Cochran χ 2 test).
When the analysis was restricted to mandibular molars, performance was slightly higher compared with the overall dataset. Pooled sensitivity and specificity were 0.96 (95% CI 0.93–0.97) and 0.97 (95% CI 0.95–0.98), respectively (Figure 3), with a pooled LR+ of 28.90 (95% CI 19.20–43.60), and a LR‐ of 0.05 (95% CI 0.03–0.07). The DOR was 631 (95% CI 350–1138), indicating excellent overall accuracy, and the AUC was 0.99 (95% CI 0.98–1.00) (Figure S3 in the online Journal of Periodontology). Again, the wide DOR CI reflects heterogeneity across studies, but overall, the findings highlight that deep learning models perform particularly well in detecting furcation involvement in mandibular molars.
FIGURE 3.

Forest plots of individual/pooled sensitivity and specificity for studies on mandibular furcation involvement (CI, confidence interval; Q, Cochran χ 2 test).
Clinical utility was assessed using Fagan plot analysis. Assuming a pre‐test probability of 20%, a positive deep learning model result increased the post‐test probability to 78%, whereas a negative result reduced it to 2% (Figure S4 in the online Journal of Periodontology). This demonstrates that deep learning models provide clinically meaningful information, effectively supporting both ruling in and ruling out furcation involvement. In the sensitivity analysis excluding the high‐risk study, 44 the pooled sensitivity and specificity were similar to the overall estimates, indicating that the main findings are robust and not substantially influenced by this study (Figure S5 in the online Journal of Periodontology).
Meta‐regression was conducted to explore potential sources of heterogeneity. Studies were stratified by dataset characteristics (single‐ vs. multi‐center), use of data augmentation (yes vs. no), and imaging modality (periapical/panoramic radiographs vs. CBCT). Pooled sensitivity and specificity were higher when data augmentation was used, but the differences were not statistically significant. Similarly, whether datasets were sourced from single or multiple centers, or used periapical/panoramic radiographs or CBCT, did not significantly affect deep learning model performance (Table 2, Figure 4). It should be noted that the CBCT subgroup included only two studies, therefore, these results should be interpreted cautiously due to the limited sample and low statistical power.
TABLE 2.
Meta‐regression analyses.
| Covariate | No. of studies | Pooled sensitivity | p value | Pooled specificity | p value |
|---|---|---|---|---|---|
| Dataset characteristics | |||||
| Multi‐center | 2 | 0.87 (0.74–1.00) | 0.07 | 0.93 (0.83–1.00) | 1.00 |
| Single‐center | 6 | 0.94 (0.90–0.98) | 0.93 (0.87–1.00) | ||
| Use of augmentation | |||||
| Yes | 5 | 0.95 (0.93–0.98) | 0.30 | 0.97 (0.95–0.98) | 0.60 |
| No | 3 | 0.85 (0.77–0.93) | 0.78 (0.70–0.85) | ||
| Imaging modality | |||||
| Periapical/panoramic radiographs | 6 | 0.94 (0.90–0.97) | 0.77 | 0.93 (0.85–1.00) | 0.97 |
| CBCT | 2 | 0.84 (0.68–1.00) | 0.98 (0.94–1.00) |
Abbreviation: CBCT, cone‐beam computed tomography.
FIGURE 4.

Meta‐regression and subgroup analyses (DA, data augmentation; IM, imaging modality; PA/PN, periapical/panoramic; CBCT, cone‐beam computed tomography).
4. DISCUSSION
This systematic review, the first meta‐analysis to assess deep learning models for detecting furcation involvement, evaluated eight studies encompassing 7814 radiographs and 12,373 molars from periapical, panoramic, and CBCT imaging. The models consistently demonstrated high diagnostic performance across varying datasets, imaging modalities, and architectures. Using comprehensive literature searches, predefined inclusion criteria, and rigorous data extraction, this review provides a robust overview of model performance in this anatomically complex region, supported by Fagan plots, pooled and subgroup analyses.
Pooled estimates indicated a sensitivity of 0.93 (95% CI 0.88–0.96) and specificity of 0.94 (95% CI 0.85–0.97), with a DOR of 187 (95% CI 47–737), and an AUC of 0.97 (95% CI 0.95–0.98). Subgroup analysis revealed even higher performance for mandibular molars (sensitivity 0.96, specificity 0.97, AUC 0.99), suggesting that these models may approach near‐perfect accuracy in certain clinical contexts. Fagan plot analysis highlighted clinical utility. With a pre‐test probability of 20%, a positive AI result increased the likelihood of furcation involvement to 78%, while a negative result reduced it to 2%. These results indicate that deep learning models can assist in clinical decision‐making, particularly when radiographic findings are ambiguous.
Methodological quality varied among studies. Most used expert consensus as a reference standard, but one relied on a single evaluator without blinding, 44 introducing potential bias. Only one study 41 reported an a priori sample size calculation, limiting confidence in some estimates. Exploratory analyses suggested that dataset origin (single‐ vs. multi‐center), imaging modality, and the use of data augmentation did not significantly affect model performance, although augmentation was associated with slight improvements. Deep learning models generally showed higher diagnostic accuracy for mandibular furcation involvement, likely due to clearer radiographic visualization compared with maxillary molars where superimposition of adjacent structures, such as the zygomatic process and shallow palatal vault, complicates interpretation. 50 CBCT, used in two studies, offered enhanced visualization but did not markedly improve diagnostic accuracy. Routine use of CBCT may not always align with the ALARA (As Low As Reasonably Achievable) principle.
It should be noted that the included studies used radiographic interpretation by clinicians as the reference standard. While radiographs offer a noninvasive assessment, they may under‐detect furcation involvement compared with clinical probing or surgical evaluation, particularly for early or subtle lesions. 51 Meta‐regression showed no significant effect of imaging modality on pooled sensitivity or specificity, suggesting that deep learning models performed consistently across different radiographic inputs despite variations in resolution and image quality. It is important to note that the diagnostic performance reflects detection of furcation presence rather than its severity or grade, which limits the direct translation of these findings.
Variability in reporting further affected interpretability. For example, one study lacked detail on annotation protocols, 39 while another relied on a single clinician as the reference standard without consensus validation, 44 increasing the risk of bias and limiting applicability. These shortcomings highlight the importance of adhering to comprehensive reporting frameworks, such as the updated CLAIM checklist for AI in medical imaging. 52 Beyond annotation protocols, reference standards, and model evaluation, the checklist emphasizes transparent reporting across the entire research pipeline, from data sources and model development to validation, fairness, and clinical implementation, thereby enhancing reproducibility and clinical trust.
Previous systematic reviews on AI in dental imaging have varied considerably in their scope. Khubrani et al., 53 for example, focused on periapical and panoramic radiographs for periodontal bone loss, while Patil et al. 54 and Li et al. 55 examined broader applications of AI in periodontal diagnostics. In contrast, our review uniquely addressed the diagnostic performance of deep learning models for detecting furcation involvement, a condition with distinct anatomical challenges. Translation from proof‐of‐concept to routine clinical use requires further model refinement. Future work should prioritize larger, multi‐center datasets, standardized imaging protocols, consistent reporting of patient demographics and periodontal status, robust augmentation strategies, and advanced or ensemble preprocessing methods to improve reliability and reduce model bias. 56 , 57 Improved detection of maxillary furcation involvement and development of models capable of grading defects, rather than simply detecting them, will also enhance clinical relevance and practical applicability.
Overall, deep learning models demonstrate strong potential for detecting furcation involvement, supporting early diagnosis and treatment planning. However, limitations remain, including the small number of studies, modest sample sizes, and predominance of single‐center datasets. 25 , 39 , 42 , 44 , 47 , 48 Additionally, heterogeneity in imaging modalities, annotation protocols, and outcome measures were noted across the different studies. Patient demographics and periodontal disease severity were often incompletely reported, which could restrict generalizability. Although analyses were conducted for mandibular and combined furcation sites, data for maxillary involvement were limited, which may affect the comparability of results across arches. The apparently lower pooled sensitivity observed for CBCT should be interpreted with caution as it is based on only two studies limiting the reliability of this subgroup estimate. Moreover, the current models face limitations, including dataset diversity, annotation bias, and risk of overfitting, which may increase false positives and compromise clinical decision‐making. Differences in image preprocessing, annotation, and dataset splitting can also hinder reproducibility and limit cross‐study comparability.
5. CONCLUSION
Deep learning models demonstrate high diagnostic accuracy in detecting furcation involvement, particularly in mandibular molars, with performance comparable to that of experienced clinicians. However, challenges remain, particularly in reducing false positives. Future research should prioritize refining these models through larger and more diverse datasets, while addressing current methodological limitations, to support their reliable and safe integration into clinical practice.
AUTHOR CONTRIBUTIONS
Momen A. Atieh: Concept/design; data collection; data analysis/interpretation; drafting article; critical revision of article; approval of article. Maanas Shah: Data analysis/interpretation; critical revision of article; approval of article. Abeer Hakam: Data analysis/interpretation; critical revision of article; approval of article. Omar Al‐Karadsheh: Data analysis/interpretation; critical revision of article; approval of article. Siraj Zabadi: Data analysis/interpretation; critical revision of article; approval of article. Fawaghi AlAli: Data analysis/interpretation; critical revision of article; approval of article. Nisheta Sachdev: Data collection; data analysis/interpretation; critical revision of article; approval of article. Andrew Tawse‐Smith: Critical revision of article; approval of article. Nabeel H. M. Alsabeeha: Data collection; data analysis/interpretation; critical revision of article; approval of article.
CONFLICT OF INTEREST STATEMENT
The authors report no conflicts of interest related to this review.
Supporting information
Supporting Information
ACKNOWLEDGMENTS
The authors have nothing to report.
DATA AVAILABILITY STATEMENT
The data that support the findings of this study are available from the corresponding author upon reasonable request.
REFERENCES
- 1. Petersen PE, Baehni PC. Periodontal health and global public health. Periodontol 2000. 2012;60:7‐14. doi:10.1111/j.1600‐0757.2012.00452.x [DOI] [PubMed] [Google Scholar]
- 2. Bartold PM, Van Dyke TE. Periodontitis: a host‐mediated disruption of microbial homeostasis. Unlearning learned concepts. Periodontol 2000. 2013;62:203‐217. doi:10.1111/j.1600‐0757.2012.00450.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Caton JG, Armitage G, Berglundh T, et al. A new classification scheme for periodontal and peri‐implant diseases and conditions—Introduction and key changes from the 1999 classification. J Periodontol. 2018;89(Suppl 1):S1‐S8. [DOI] [PubMed] [Google Scholar]
- 4. Papapanou PN, Sanz M, Buduneli N, et al. Periodontitis: consensus report of workgroup 2 of the 2017 world workshop on the classification of periodontal and peri‐implant diseases and conditions. J Periodontol. 2018;89(Suppl 1):S173‐S182. [DOI] [PubMed] [Google Scholar]
- 5. Nibali L, Zavattini A, Nagata K, et al. Tooth loss in molars with and without furcation involvement—a systematic review and meta‐analysis. J Clin Periodontol. 2016;43:156‐166. doi:10.1111/jcpe.12497 [DOI] [PubMed] [Google Scholar]
- 6. Trullenque‐Eriksson A, Tomasi C, Petzold M, Berglundh T, Derks J. Furcation involvement and tooth loss: a registry‐based retrospective cohort study. J Clin Periodontol. 2023;50:339‐347. doi:10.1111/jcpe.13754 [DOI] [PubMed] [Google Scholar]
- 7. McFall WT Jr. Tooth loss in 100 treated patients with periodontal disease. A long‐term study. J Periodontol. 1982;53:539‐549. doi:10.1902/jop.1982.53.9.539 [DOI] [PubMed] [Google Scholar]
- 8. Svardstrom G, Wennstrom JL. Prevalence of furcation involvements in patients referred for periodontal treatment. J Clin Periodontol. 1996;23:1093‐1099. doi:10.1111/j.1600‐051X.1996.tb01809.x [DOI] [PubMed] [Google Scholar]
- 9. Abbas F, Hart AA, Oosting J, van der Velden U. Effect of training and probing force on the reproducibility of pocket depth measurements. J Periodontal Res. 1982;17:226‐234. doi:10.1111/j.1600‐0765.1982.tb01149.x [DOI] [PubMed] [Google Scholar]
- 10. Graetz C, Plaumann A, Wiebe JF, Springer C, Salzer S, Dorfer CE. Periodontal probing versus radiographs for the diagnosis of furcation involvement. J Periodontol. 2014;85:1371‐1379. doi:10.1902/jop.2014.130612 [DOI] [PubMed] [Google Scholar]
- 11. Haas LF, Zimmermann GS, De Luca, Canto G, Flores‐Mir C, Correa M. Precision of cone beam CT to assess periodontal bone defects: a systematic review and meta‐analysis. Dentomaxillofac Radiol. 2018;47:20170084. doi:10.1259/dmfr.20170084 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Braun X, Ritter L, Jervoe‐Storm PM, Frentzen M. Diagnostic accuracy of CBCT for periodontal lesions. Clin Oral Investig. 2014;18:1229‐1236. doi:10.1007/s00784‐013‐1106‐0 [DOI] [PubMed] [Google Scholar]
- 13. Schmidhuber J. Deep learning in neural networks: an overview. Neural Netw. 2015;61:85‐117. doi:10.1016/j.neunet.2014.09.003 [DOI] [PubMed] [Google Scholar]
- 14. Hashimoto DA, Rosman G, Rus D, Meireles OR. Artificial intelligence in surgery: promises and perils. Ann Surg. 2018;268:70‐76. doi:10.1097/SLA.0000000000002693 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Thrall JH, Li X, Li Q, et al. Artificial intelligence and machine learning in radiology: opportunities, challenges, pitfalls, and criteria for success. J Am Coll Radiol. 2018;15:504‐508. doi:10.1016/j.jacr.2017.12.026 [DOI] [PubMed] [Google Scholar]
- 16. Seager A, Sharp L, Neilson LJ, et al. Polyp detection with colonoscopy assisted by the GI Genius artificial intelligence endoscopy module compared with standard colonoscopy in routine colonoscopy practice (COLO‐DETECT): a multicentre, open‐label, parallel‐arm, pragmatic randomised controlled trial. Lancet Gastroenterol Hepatol. 2024;9:911‐923. [DOI] [PubMed] [Google Scholar]
- 17. Lee JH, Kim DH, Jeong SN, Choi SH. Detection and diagnosis of dental caries using a deep learning‐based convolutional neural network algorithm. J Dent. 2018;77:106‐111. doi:10.1016/j.jdent.2018.07.015 [DOI] [PubMed] [Google Scholar]
- 18. Lin PL, Huang PW, Huang PY, Hsu HC. Alveolar bone‐loss area localization in periodontitis radiographs based on threshold segmentation with a hybrid feature fused of intensity and the H‐value of fractional Brownian motion model. Comput Methods Programs Biomed. 2015;121:117‐126. doi:10.1016/j.cmpb.2015.05.004 [DOI] [PubMed] [Google Scholar]
- 19. Chang HJ, Lee SJ, Yong TH, et al. Deep learning hybrid method to automatically diagnose periodontal bone loss and stage periodontitis. Sci Rep. 2020;10:7531. doi:10.1038/s41598‐020‐64509‐z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Okada K, Rysavy S, Flores A, Linguraru MG. Noninvasive differential diagnosis of dental periapical lesions in cone‐beam CT scans. Med Phys. 2015;42:1653‐1665. doi:10.1118/1.4914418 [DOI] [PubMed] [Google Scholar]
- 21. Abdolali F, Zoroofi RA, Otake Y, Sato Y. Automated classification of maxillofacial cysts in cone beam CT images using contourlet transformation and spherical harmonics. Comput Methods Programs Biomed. 2017;139:197‐207. doi:10.1016/j.cmpb.2016.10.024 [DOI] [PubMed] [Google Scholar]
- 22. Yilmaz E, Kayikcioglu T, Kayipmaz S. Computer‐aided diagnosis of periapical cyst and keratocystic odontogenic tumor on cone beam computed tomography. Comput Methods Programs Biomed. 2017;146:91‐100. doi:10.1016/j.cmpb.2017.05.012 [DOI] [PubMed] [Google Scholar]
- 23. Krois J, Ekert T, Meinhold L, et al. Deep learning for the radiographic detection of periodontal bone loss. Sci Rep. 2019;9:8495. doi:10.1038/s41598‐019‐44839‐3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Lee JH, Kim DH, Jeong SN, Choi SH. Diagnosis and prediction of periodontally compromised teeth using a deep learning‐based convolutional neural network algorithm. J Periodontal Implant Sci. 2018;48:114‐123. doi:10.5051/jpis.2018.48.2.114 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Mao YC, Huang YC, Chen TY, et al. Deep learning for dental diagnosis: a novel approach to furcation involvement detection on periapical radiographs. Bioengineering (Basel). 2023;10:802. doi:10.3390/bioengineering10070802 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. LeCun Y, Bengio Y, Hinton G, Deep learning. Nature 2015;521:436‐444. doi:10.1038/nature14539 [DOI] [PubMed] [Google Scholar]
- 27. Deek JJ, Bossuyt PM, Leeflang MM, Takwoingi Y. Cochrane handbook for systematic reviews of diagnostic test accuracy. Version 2.0 (updated July 2023). In: Cochrane. 2023. Available from https://training.cochrane.org/handbook‐diagnostic‐test‐accuracy/current [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. McInnes MDF, Moher D, Thombs BD, et al. Preferred reporting items for a systematic review and meta‐analysis of diagnostic test accuracy studies: the PRISMA‐DTA statement. JAMA 2018;319:388‐396. doi:10.1001/jama.2017.19163 [DOI] [PubMed] [Google Scholar]
- 29. Faggion CM Jr., Atieh MA, Park S. Search strategies in systematic reviews in periodontology and implant dentistry. J Clin Periodontol. 2013;40:883‐888. doi:10.1111/jcpe.12132 [DOI] [PubMed] [Google Scholar]
- 30. Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. CACM. 2017;60:84‐90. doi:10.1145/3065386 [Google Scholar]
- 31. Whiting PF, Rutjes AW, Westwood ME, et al. QUADAS‐2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155:529‐536. doi:10.7326/0003‐4819‐155‐8‐201110180‐00009 [DOI] [PubMed] [Google Scholar]
- 32. Dinnes J, Deeks J, Kirby J, Roderick P. A methodological review of how heterogeneity has been examined in systematic reviews of diagnostic test accuracy. Health Technol Assess. 2005;9:1‐113, iii. doi:10.3310/hta9120 [DOI] [PubMed] [Google Scholar]
- 33. Jaeschke R, Guyatt GH, Sackett DL. Users' guides to the medical literature. III. How to use an article about a diagnostic test. B. What are the results and will they help me in caring for my patients? The evidence‐based medicine working group. Jama. 1994;271:703‐707. doi:10.1001/jama.1994.03510330081039 [DOI] [PubMed] [Google Scholar]
- 34. Higgins JP, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta‐analyses. Bmj. 2003;327:557‐560. doi:10.1136/bmj.327.7414.557 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Akobeng AK. Understanding diagnostic tests 3: receiver operating characteristic curves. Acta Paediatr. 2007;96:644‐647. doi:10.1111/j.1651‐2227.2006.00178.x [DOI] [PubMed] [Google Scholar]
- 36. Deeks JJ, Macaskill P, Irwig L. The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J Clin Epidemiol. 2005;58:882‐893. doi:10.1016/j.jclinepi.2005.01.016 [DOI] [PubMed] [Google Scholar]
- 37. Glas AS, Lijmer JG, Prins MH, Bonsel GJ, Bossuyt PM. The diagnostic odds ratio: a single indicator of test performance. J Clin Epidemiol. 2003;56:1129‐1135. doi:10.1016/S0895‐4356(03)00177‐X [DOI] [PubMed] [Google Scholar]
- 38. Hoss P, Meyer O, Wolfle UC, et al. Detection of periodontal bone loss on periapical radiographs‐a diagnostic study using different convolutional neural networks. J Clin Med. 2023;12:7189. doi:10.3390/jcm12227189 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Karadsheh O, Zabadi S. Digital aid in diagnosis of periodontitis by deep learning algorithm. Oral presentation at IADR Jordanian Section Meeting, Amman, Jordan 2023 ;2023.
- 40. Khan HA, Haider MA, Ansari HA, et al. Automated feature detection in dental periapical radiographs by using deep learning. Oral Surg Oral Med Oral Pathol Oral Radiol. 2021;131:711‐720. doi:10.1016/j.oooo.2020.08.024 [DOI] [PubMed] [Google Scholar]
- 41. Kurt‐Bayrakdar S, Bayrakdar IS, Kuran A, Celik O, Orhan K, Jagtap R. Advancing periodontal diagnosis: harnessing advanced artificial intelligence for patterns of periodontal bone loss in cone‐beam computed tomography. Dentomaxillofac Radiol. 2025;54:268‐278. doi:10.1093/dmfr/twaf011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Kurt‐Bayrakdar S, Bayrakdar IS, Yavuz MB, et al. Detection of periodontal bone loss patterns and furcation defects from panoramic radiographs using deep learning algorithm: a retrospective study. BMC Oral Health. 2024;24:155. doi:10.1186/s12903‐024‐03896‐5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Kurt‐Bayrakdar S, Celik O, Bayrakdar IS, et al. Success of artificial intelligence system in determining alveolar bone loss from dental panoramic radiography images. CDJ. 2020;23:1‐7. doi:10.7126/cumudj.777057 [Google Scholar]
- 44. Shetty S, Talaat W, AlKawas S, et al. Application of artificial intelligence‐based detection of furcation involvement in mandibular first molar using cone beam tomography images‐ a preliminary study. BMC Oral Health. 2024;24:1476. doi:10.1186/s12903‐024‐05268‐5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Tajima S, Sonoda C, Kobayashi T. Development of an artificial intelligence model using an automatic detection of furcation involvement through panoramic radiography. JCP. 2021;63:119‐128. [Google Scholar]
- 46. Tian E, Hong J, Tang Z, et al. Development and validation of a polyfit approach for assessing alveolar bone loss using panoramic radiography. BMC Oral Health. 2025;25:417. doi:10.1186/s12903‐025‐05714‐y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Vilkomir K, Phen C, Baldwin F, Cole J, Herndon N, Zhang W. Classification of mandibular molar furcation involvement in periapical radiographs by deep learning. Imaging Sci Dent. 2024;54:257‐263. doi:10.5624/isd.20240020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Zhang X, Guo E, Liu X, et al. Enhancing furcation involvement classification on panoramic radiographs with vision transformers. BMC Oral Health. 2025;25:153. doi:10.1186/s12903‐025‐05431‐6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Kwon O, Yong TH, Kang SR, et al. Automatic diagnosis for cysts and tumors of both jaws on panoramic radiographs using a deep convolution neural network. Dentomaxillofac Radiol. 2020;49:20200185. doi:10.1259/dmfr.20200185 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Bragger U. Radiographic parameters: biological significance and clinical use. Periodontol 2000. 2005;39:73‐90. doi:10.1111/j.1600‐0757.2005.00128.x [DOI] [PubMed] [Google Scholar]
- 51. Gurgan C, Grondahl K, Wennstrom JL. Radiographic detectability of bone loss in the bifurcation of mandibular molars: an experimental study. Dentomaxillofac Radiol. 1994;23:143‐148. doi:10.1259/dmfr.23.3.7835514 [DOI] [PubMed] [Google Scholar]
- 52. Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol Artif Intell. 2024;6:e240300. doi:10.1148/ryai.240300 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Khubrani YH, Thomas D, Slator PJ, White RD, Farnell DJJ. Detection of periodontal bone loss and periodontitis from 2D dental radiographs via machine learning and deep learning: systematic review employing APPRAISE‐AI and meta‐analysis. Dentomaxillofac Radiol. 2025;54:89‐108. doi:10.1093/dmfr/twae070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Patil S, Joda T, Soffe B, et al. Efficacy of artificial intelligence in the detection of periodontal bone loss and classification of periodontal diseases: a systematic review. J Am Dent Assoc. 2023;154:795‐804 e791. doi:10.1016/j.adaj.2023.05.010 [DOI] [PubMed] [Google Scholar]
- 55. Li X, Zhao D, Xie J, et al. Deep learning for classifying the stages of periodontitis on dental images: a systematic review and meta‐analysis. BMC Oral Health. 2023;23:1017. doi:10.1186/s12903‐023‐03751‐z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Owler J, Rockett P. Influence of background preprocessing on the performance of deep learning retinal vessel detection. J Med Imaging (Bellingham). 2021;8:064001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Supriyadi MR, Samah ABA, Muliadi J, et al. A systematic literature review: exploring the challenges of ensemble model for medical imaging. BMC Med Imaging. 2025;25:128. doi:10.1186/s12880‐025‐01667‐4 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting Information
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
