ABSTRACT
Aim
To construct predictive models of periodontitis progression by applying Machine Learning (ML) to baseline data from a study of periodontitis progression.
Materials and Methods
Logistic regression (LR), multi‐layer perceptron (MLP) and probabilistic graphic model (PGM) were utilised on data from a multi‐centre longitudinal study in which periodontally healthy (n = 113) and periodontitis participants (n = 302) were examined bi‐monthly for 12 months without treatment. Periodontal examination was performed, and salivary levels of 10 analytes were determined. Clinical and demographic parameters and analytes levels were input into the model. The performance of 14 models was compared using the area under the receiver operating characteristic curve (AUROC), and feature importance was assessed using SHapley Additive exPlanations (SHAP).
Results
The PGM model (Clinical measures, saliva IL‐1β, age, sex) demonstrated the best overall performance (AUROC = 0.88), compared to LR (AUROC = 0.72) and MLP (AUROC = 0.58). Although MLP had a lower Brier score (0.12), its sensitivity was 0, limiting its clinical utility. In contrast, PGM achieved a balanced sensitivity (0.55) and specificity (0.81). Feature importance analyses highlighted the number of deep periodontal pockets as a key driver of model predictions in both PGM and MLP.
Conclusions
ML models can predict periodontitis progression, supporting early detection strategies. Our integrative approach, combining clinical data with salivary biomarkers such as IL‐1β, improved predictive accuracy.
Keywords: artificial intelligence, disease progression, machine learning, periodontitis
1. Introduction
Periodontitis affects 47% of US adults and, in its severe form, affects 1.1 billion people worldwide (Chen et al. 2021; Jain et al. 2023). It can lead to bone and tooth loss, tooth migration and mobility. The burden of untreated periodontitis also represents a considerable economic challenge, particularly in the context of an aging global population (Jain et al. 2023). Thus, early identification of patients at greatest risk of periodontitis progression and timely intervention are crucial (Bumm et al. 2024; Fardal et al. 2023). The few strategies proposed to predict periodontitis progression often using linear regression (LR) (Leite et al. 2017; Lindskog et al. 2010; Martinez‐Canut 2015; Morelli et al. 2018; Serroni et al. 2025). They focus on clinical parameters and patient‐specific risk factors such as diabetes and smoking, but have limited prediction performance and clinical utility (Du et al. 2018).
Artificial intelligence (AI) and machine learning (ML) have revolutionised data analysis, particularly for complex and non‐linear datasets (Obermeyer and Emanuel 2016; Theodosiou and Read 2023). Yet, most ML models for periodontitis are retrospective or cross‐sectional, focus on diagnosis and relying on imaging or electronic health records (EHR). They utilise unstructured data and variable case definitions, limiting their clinical applicability Despite demonstrating high accuracy levels, these models lack external validation and replication in diverse populations (Patel et al. 2022, 2023; Enevold et al. 2023; Swinckels et al. 2025). Notably, most models fail to incorporate biological data, missing the opportunity to leverage high‐throughput technologies to explore the biological drivers of periodontal destruction (Cho et al. 2021; Kang et al. 2022; Martorell‐Marugán et al. 2019). Integrating clinical, immunological, microbiological and genetic data into predictive models requires well‐designed longitudinal studies and collaboration between data scientists and periodontal experts (Adeoye and Su 2024; Bashir et al. 2022; Scott et al. 2023). This approach can help predicting periodontitis progression, an underexplored area (Ossowska et al. 2022; Patel et al. 2022, 2023).
Traditional LR models are favoured for simplicity but often miss complex, non‐linear relationships in biological data. In contrast, ML models, including multi‐layer perceptron (MLP) and probabilistic graphical models (PGM), handle non‐linear data effectively. MLP's computational power enables deeper insights into disease progression, supporting more personalised treatments (Bikku 2020; Bokhare et al. 2023; Larrañaga and Moral 2011), while PGM's probabilistic relationships and data dependencies enhance predictive accuracy (A. Gupta et al. 2019).
The present study seeks to identify the most precise and sensitive method for predicting periodontitis progression by utilising minimal, easy to obtain, input information and employing LR, MLP and PGM in predicting periodontitis progression based on baseline clinical, sociodemographic and immunological data from a longitudinal multi‐centre study. The underlying hypothesis is that integrating clinical and biological information could enhance model performance compared to more simplistic approaches and potentially contribute to the development of more effective, personalised strategies for periodontitis management.
2. Materials and Methods
2.1. Study Design and Population
Participants were recruited from a multi‐center clinical study on periodontal disease biomarkers (ClinicalTrials.gov ID: NCT01489839). Participants had bi‐monthly clinical assessments over one year to monitor periodontitis progression and were classified according to the 2018 criteria (Papapanou et al. 2018). Study protocols were approved by the Institutional Review Boards of The Forsyth Institute and affiliates (R. Teles et al. 2018).
Eligibility criteria are detailed in Supporting Information. In brief, it included age ≥ 25, at least 20 natural teeth, of which 12 were premolars or molars. Periodontitis participants had at ≥ 4 teeth with pocket depth (PD) ≥ 5 mm and clinical attachment level (CAL) ≥ 2 mm, as well as radiographic evidence of bone loss, thus focusing on patients with generalised periodontitis. Periodontally healthy participants presented PD ≤ 3 mm, or PD ≥ 4 mm without CAL and no evidence of radiographic bone loss. Exclusions were pregnancy, lactation and recent antibiotic or periodontal treatment and chronic NSAIDs use. Although smokers and diabetes are at higher risk for periodontitis progression, they were excluded this study aimed at assessing the natural course of periodontitis in the absence of disease modifiers (Figure S1).
2.2. Clinical Examination of Subjects and Sites
Participants had periodontal parameters measured at up to 168 sites each; six sites per tooth for up to 28 teeth, excluding third molars. Measurements included PD, CALand presence of plaque (PLAQ), bleeding on probing (BOP) and suppuration. Measures were recorded using calibrated North Carolina manual probes (PCPUNC 15 Hu‐Friedy Co, Chicago, IL), rounded to the nearest millimetre. The intra‐examiner accuracy was high, with exact agreement (±SD) at 66.2% (±6.6%) and agreement within 1 mm at 95.9% (±2.3%) (R. Teles et al. 2016) (Supporting Information). Only individuals who completed the monitoring phase (bi‐monthly exams over 12 months in the absence of treatment) with the same examiner were included in this analysis.
2.3. Definition of Progression of Periodontitis
Periodontitis progression was defined based on CAL (R. Teles et al. 2016), measured at baseline, 2, 4, 6, 8, 10 and 12 months to track changes over time. A linear mixed model (LMM) accounted for fixed effects (age, gender, time) and random effects (subjects and sites) to capture individual variability. The LMM generated subject‐specific regression curves to predict CAL at each time point, classifying sites as progressing or regressing based on CAL increase thresholds (1, 2 and 3 mm). Participants were categorised by progression: P0 (no progressing sites), P1 (1–2 sites progressing) and P2 (≥ 3 sites progressing) (F. R. F. Teles, Chandrasekaran, et al. 2024; F. Teles, Martin, et al. 2024; R. Teles et al. 2016).
2.4. Saliva Sample Collection
Saliva samples were collected as described earlier (F. Teles, Chandrasekaran, et al. 2024; F. Teles, Martin, et al. 2024). Participants refrained from brushing, chewing gum, eating or drinking for 1½ h before each visit, scheduled at a consistent time (e.g., morning or afternoon) to control for diurnal analyte variations. Subjects accumulated saliva for 60 s without swallowing and deposited it into a tube on ice during the 10‐min collection. Protease inhibitors (20 μL of 1 mg/mL Aprotinin and 10 μL of 200 mM PMSF) were added to 2 mL of saliva. Samples were aliquoted, snap‐frozen using CoolRack/CoolBox and stored at −80°C.
2.5. Measurement of Analytes in Using Bead‐Based Immunoassay (Luminex)
Salivary immunoassays were performed as described earlier (F. Teles, Chandrasekaran, et al. 2024; F. Teles, Martin, et al. 2024). Briefly, whole saliva samples were analysed for IFN‐g, IL‐6, VEGF, IL‐1b, IL‐10, OPG, MMP‐8, MCP‐1, IL‐8 and MMP‐9 using Luminex multiplex assay kits (R&D Systems) on a Bio‐Rad Bio‐Plex 200. Undiluted samples were thawed, centrifuged to remove debris and assayed across three panels. Samples were thawed at 4°C overnight after selection with Freezer Works. The Bio‐Plex 200 was calibrated each morning, and assay kits were warmed to room temperature. Standard curves and controls were prepared in duplicate, and plates washed with a Bio‐Rad plate washer. Data was saved in Excel and securely stored for later analysis. The detection range of the assay is presented under Supporting Information.
2.6. Models Design and Predictors
Model design adhered to the TRIPOD‐AI guidelines for reporting machine learning models (Collins et al. 2015). Feature selection prioritised routinely collected clinical data and salivary analytes to enable accessible, cost‐effective prediction of periodontitis progression, maximising clinical utility and return on investment (Figure S4). Features were selected based on prior clinical and biological relevance (Ossowska et al. 2023) (Deng et al. 2023; Matuliene et al. 2008) and included age, sex, disease stage (Health/Gingivitis, Stage II, Stage III), number of missing teeth (TMISSN), percentage of sites with BOP and plaque (PLAQ), number of deep pockets (PD ≥ 5 mm) and mean log‐transformed levels of 10 salivary analytes.
Fourteen models were tested using LR, MLP and PGM (Table 1), aiming to predict 12‐month periodontitis progression in the absence of treatment. Outcomes were classified as Stable (0–2 progressing sites; P0/P1) or Progressing (≥ 3 sites; P2). Due to MLP's sensitivity to missing data, participants with incomplete input data were excluded. Models were developed in Python (3.12.4), leveraging TensorFlow (2.18.0) and Keras (3.8.0) for MLP and Scikit‐Learn (1.6.0) for PGM. LR was implemented in R (4.4.1) using the glm function. Given the small sample size, the MLP architecture was kept shallow with dropout regularisation to mitigate overfitting (Figure S2). All code and prediction tools are publicly available on GitHub (Caruth and Verma 2025).
TABLE 1.
Features inputted in each tested model.
| Feature type | Feature | Feature description (assessed at baseline) | Feature type | Model 1 | Model 2 | Model 3 | Model 4 | Model 5 | Model 6 | Model 7 | Model 8 | Model 9 | Model 10 | Model 11 | Model 12 | Model 13 | Model 14 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Demographic | Age | Age in years | Continuous | ☑ | ☑ | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ |
| Sex | Male or female represented numerically (M = 1 or F = 2) | Categorical | ☑ | ☑ | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | |
| Clinical features | TMISSN | Number of missing teeth | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ |
| PD ≥ 5 mm | Number of sites with deep pockets greater than 5 | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ | |
| PLAQ | Percentage of sites with plaque | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ | |
| BOP | Percentage of sites with bleeding on probing | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ | |
| Disease classification | Patient disease status | Categorical | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ | |
| Salivary analytes | VEGF | Log‐transformed values | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☐ | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ |
| MMP‐8 | Log‐transformed values | Continuous | ☑ | ☑ | ☑ | ☑ | ☑ | ☐ | ☑ | ☐ | ☐ | ☑ | ☐ | ☐ | ☐ | ☐ | |
| IL‐1β | Log‐transformed values | Continuous | ☑ | ☑ | ☐ | ☑ | ☐ | ☐ | ☐ | ☑ | ☐ | ☐ | ☑ | ☐ | ☐ | ☐ | |
| IFN‐γ | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| IL‐6 | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| IL‐10 | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| OPG | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| MCP‐1 | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| IL‐8 | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | |
| MMP‐9 | Log‐transformed values | Continuous | ☑ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ | ☐ |
2.7. Performance Measures and Measure of Feature Importance
Model performance was primarily evaluated using the area under the receiver operating characteristic curve (AUROC), which reflects the model's ability to assign higher probabilities to true positive cases. An AUROC of 0.5 indicates random performance. Area under the precision‐recall curve (AUPRC) was also used to capture the balance between sensitivity and precision. Additional metrics included sensitivity, specificity and accuracy to provide a comprehensive performance profile. Calibration was assessed using the Brier score, which measures the mean squared difference between predicted probabilities and observed outcomes (range: 0–1; lower = better).
Model interpretability was enhanced using SHAP (SHapley Additive exPlanations), which quantifies each feature's contribution to predictions based on game theory (Lundberg and Lee 2017). While SHAP improves understanding of complex models, it assumes feature independence—a potential limitation. Feature correlations are provided in Figure S3.
3. Results
3.1. Clinical and Demographic Characteristics of the Study Population
Because MLP is sensitive to data missingness, meticulous data cleaning was necessary to create a complete dataset for all features (Model 1). Thus, the cohort of 415 patients resulted in a final dataset of 273 participants, comprising 164 females and 109 males. They were classified into periodontal health/gingivitis (n = 72), stage II (n = 24) and stage III (n = 177) periodontitis. Eighty‐six percent of the patients in the study's original dataset had little to no progression, with the remaining 14% experiencing substantial progression. Once patients with missing features were excluded, we saw a similar distribution: 87% and 13%, respectively. The presence and severity of disease were evident by the magnitude of baseline mean PD and CAL. Stage III patients had the highest mean numbers of missing teeth and pockets ≥ 5 mm and mean percentages of sites presenting plaque, redness, BOP and suppuration (Table 2). Disease progression is evidenced by the increase in mean PD and CAL as well as the mean number of pockets ≥ 5 mm from baseline to 12 months. Progressing patients had higher mean % of plaque, redness and BOP. Progression occurred primarily among stage III periodontitis participants (Table 3).
TABLE 2.
Clinical and demographic characteristics of the population inputted into the model by periodontal status.
| Healthy | Stage II | Stage III | |
|---|---|---|---|
| Number of subjects | 72 | 24 | 177 |
| Baseline clinical groups (healthy/mild/severe) | 72/0/0 | 0/14/10 | 0/77/100 |
| Progression class (P0/P1/P2) | 55/14/3 | 18/4/2 | 83/64/30 |
| Number of male/female | 19/53 | 5/19 | 85/92 |
| AA/White/Oth/Unk | 13/47/12/0 | 4/16/4/0 | 48/120/9/0 |
| Age (years; median, IQR) | 39 (29–49) | 50.5 (40–55) | 52 (45–59) |
| Periodontal parameters | Baseline | 12 M | Baseline | 12 M | Baseline | 12 M (n = 174) |
|---|---|---|---|---|---|---|
| # missing teeth (median, IQR) | 0 (0–1) | 0 (0–1) | 1 (0–1.5) | 1 (0–1.5) | 1 (0–2) | 1 (0–3) |
| PD (mm; median, IQR) | 1.8 (1.6–2.0) | 1.8 (1.6–2.0) | 2.2 (2.0–2.4) | 2.2 (2.0–2.4) | 2.6 (2.3–2.8) | 2.6 (2.3–2.9) |
| CAL (mm; median, IQR) | 1.2 (0.9–1.5) | 1.2 (0.9–1.6) | 1.7 (1.6–1.9) | 1.6 (1.4–2.1) | 2.3 (1.9–2.7) | 2.2 (1.9–2.9) |
| % sites per subject with | ||||||
| Plaque (median, IQR) | 45.2 (31.7–65.2) | 43.2 (24.4–68.7) | 66.1 (48.4–83.6) | 79.8 (39.6–96.1) | 73.5 (53.2–85.4) | 74.7 (47.3–88.9) |
| Gingival redness (median, IQR) | 21.1 (11.5–34.2) | 26.2 (16.1–43.5) | 47.1 (24.7–77.4) | 50.0 (31.8–93.8) | 60.7 (39.3–80.6) | 62.7 (41.4–84.6) |
| BOP (median, IQR) | 10.2 (3.6–30.9) | 10.8 (2.4–26.9) | 40.0 (16.0–54.2) | 24.7 (20.7–48.7) | 41.4 (27.4–64.2) | 38.8 (23.5–61.7) |
| Suppuration (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) |
| # of sites/subject (median, IQR) | ||||||
| PD < 4 mm (median, IQR) | 168 (162–168) | 166 (156–168) | 148.5 (137.5–157.5) | 152 (144–160) | 133 (117–144) | 134 (119–147) |
| PD 4–6 mm (median, IQR) | 0 (0–0.5) | 0 (0–2) | 12 (6.5–23.5) | 9 (4–17.5) | 24 (16–37) | 22 (13–39) |
| PD > 6 mm (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–3) | 0 (0–2) |
| PD < 5 mm (median, IQR) | 168 (162–168) | 168 (162–168) | 160 (154–167.5) | 160 (156–165) | 145 (134–154) | 147 (138–157) |
| PD 5+ mm (median, IQR) | 0 (0–0) | 0 (0–0) | 2 (0–4) | 1 (0–3) | 12 (8–22) | 10 (4–20) |
| CAL < 4 mm (median, IQR) | 168 (161–168) | 167 (156.5–168) | 160 (153–162) | 157.5 (150.5–162) | 135 (117–147) | 136.5 (118–151) |
| CAL 4–6 mm (median, IQR) | 0 (0–0.5) | 0 (0–1) | 5 (1.5–8) | 1 (0–8.5) | 21 (14–36) | 19 (10–33) |
| CAL > 6 mm (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 1 (0–4) | 0 (0–4) |
TABLE 3.
Clinical and demographic characteristics of the population inputted into the model by periodontitis progression.
| Healthy/gingivitis | Periodontitis | |||
|---|---|---|---|---|
| Stable | Progressing | Stable | Progressing | |
| Number of subjects | 69 | 3 | 169 | 32 |
| Baseline clinical groups (healthy/mild/severe) | 69/0/0 | 3/0/0 | 0/76/93 | 0/15/17 |
| Progression class (P0/P1/P2) | 55/14/0 | 0/0/3 | 101/68/0 | 0/0/32 |
| Number of male/female | 18/51 | 1/2 | 73/96 | 17/15 |
| AA/White/Oth/Unk | 12/46/11/0 | 1/1/1/0 | 45/116/8/0 | 7/20/5/0 |
| Age (years; median, IQR) | 39 (29–49) | 42 (25–48) | 52 (47–59) | 48 (39–58) |
| Periodontal parameters | Baseline | 12 M | Baseline | 12 M | Baseline | 12 M (n = 168) | Baseline | 12 M (n = 30) |
|---|---|---|---|---|---|---|---|---|
| # missing teeth (median, IQR) | 0 (0–1) | 0 (0–1) | 0 (0–0) | 0 (0–0) | 1 (0–2) | 1 (0–3) | 1 (0–2) | 1 (0–2) |
| PD (mm; median, IQR) | 1.8 (1.6–1.9) | 1.8 (1.6–2.0) | 1.9 (1.5–1.9) | 2.4 (1.5–2.6) | 2.5 (2.3–2.8) | 2.4 (2.2–2.8) | 2.5 (2.0–2.7) | 2.7 (2.5–3.0) |
| CAL (mm; median, IQR) | 1.2 (0.9–1.5) | 1.2 (0.9–1.5) | 1.5 (0.3–1.9) | 2.4 (0.7–2.7) | 2.2 (1.8–2.6) | 2.1 (1.7–2.6) | 2.5 (1.8–2.9) | 2.8 (2.2–3.2) |
| % sites per subject with | ||||||||
| Plaque (median, IQR) | 43.5 (30.1–64.9) | 42.4 (23.8–69.4) | 58.3 (45.2–66.1) | 44.0 (27.4–63.7) | 70.8 (50.0–82.7) | 72.5 (45.8–91.5) | 82.4 (67.9–89.4) | 80.3 (62.5–87.0) |
| Gingival redness (median, IQR) | 22.6 (11.1–34.5) | 26.2 (15.5–44.0) | 17.9 (16.1–21.4) | 30.4 (22.6–32.1) | 59.1 (36.8–79.2) | 60.9 (38.8–86.3) | 63.8 (49.4–82.7) | 65.8 (45.2–91.7) |
| BOP (median, IQR) | 9.5 (3.6–29.2) | 10.3 (2.4–23.2) | 36.9 (21.4–61.3) | 44.6 (3.6–46.4) | 39.7 (25.0–62.3) | 36.3 (22.2–57.7) | 48.4 (32.1–72.7) | 52.1 (32.7–70.4) |
| Suppuration (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) |
| # of sites/subject (median, IQR) | ||||||||
| PD < 4 mm (median, IQR) | 168 (162–168) | 166 (158–168) | 166 (164–168) | 155 (153–168) | 135 (120–146) | 138 (123–149.5) | 134.5 (120–147) | 129 (111–137) |
| PD 4–6 mm (median, IQR) | 0 (0–0) | 0 (0–1) | 2 (0–4) | 13 (0–15) | 24 (14–36) | 18 (11–33) | 21 (16–34.5) | 31 (18–46) |
| PD > 6 mm (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–2) | 0 (0–1) | 0 (0–3.5) | 1.5 (0–5) |
| PD < 5 mm (median, IQR) | 168 (162–168) | 168 (162–168) | 168 (168–168) | 167 (165–168) | 147 (135–156) | 150 (140–160) | 143.5 (134.5–156) | 144.5 (131–153) |
| PD 5+ mm (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 1 (0–3) | 11 (6–21) | 7 (2.5–15.5) | 14.5 (8–21.5) | 15.5 (11–25) |
| CAL < 4 mm (median, IQR) | 168 (160–168) | 167 (159–168) | 166 (166–168) | 155 (153–168) | 140 (122–152) | 144 (126–154.5) | 127.5 (112.5–144) | 118.5 (107–136) |
| CAL 4–6 mm (median, IQR) | 0 (0–0) | 0 (0–1) | 2 (0–2) | 13 (0–15) | 19 (9–33) | 14 (6–27) | 22 (18–39.5) | 36.5 (23–47) |
| CAL > 6 mm (median, IQR) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–0) | 0 (0–2) | 0 (0–2) | 1.5 (0–7) | 2.5 (0–8) |
3.2. Levels of Salivary Analytes in Periodontitis Progression and Stability
Saliva analysis revealed distinct patterns between progression and stability (Figure 1). Albeit not evident in periodontal health/gingivitis (Figure 1a), they were clear among periodontitis patients (Figure 1b,c). Progressing Stage II and III participants exhibited consistently higher levels of IL‐8, OPG, IFN‐γ, IL‐1β, IL‐10, IL‐6, MMP‐8 and VEGF than their stable counterparts, which were often observed as early as baseline (Figure 1b,c).
FIGURE 1.

(a) Salivary profile of 10 analytes analysed over one year of oral health monitoring, categorised by progression class; (b) salivary profile of 10 analytes analysed over one year for periodontitis participants at stage III, categorised by progression class. Note that there is only a sample size of 2 for stage II progressing individuals month‐18 profiles, resulting in a sample size of 1 at month 18 for stage II progressing. Consequently, error bars are not displayed for stage II progressing saliva at month 18; (c) salivary profile of 10 analytes analysed over one year and 6 months post‐treatment of periodontitis participants stage III, categorised by progression class.
3.3. Assessment of Algorithms and Models
Seventeen input features were evaluated, and 14 feature combinations were tested across three algorithms (LR, MLP and PGM) (Table 1). Training and testing for the MLP model spanned 50 epochs, using an 80% training and 20% testing split. Cross‐validation included five dataset splits, with averaged results reported. Effective regularisation was indicated by the validation curve outperforming the training curve in all models (Figure S5). The PGM, as a naïve Bayesian model, did not follow an iterative epoch‐based approach. AUROC and confidence intervals for all models tested are presented in Figure 2.
FIGURE 2.

The plot illustrates the Mean ROAUCs for each model and method (MLP, PGM, LR) with 95% confidence intervals. The plot shows PGM outperforming the other methods across all feature combinations. MLP shows the lowest AUROC across many of the combinations. Age + Sex have the lowest performance across the models but still achieve notable AUROC with PGM.
Across all 17 feature sets evaluated, PGM demonstrated the most balanced and robust overall performance. With all features included (Model 1), LR showed the highest AUROC (0.78 ± 0.09), followed closely by PGM (AUROC 0.77 ± 0.25), though PGM offered a better trade‐off between sensitivity (1.0) and specificity (0.75) compared to LR (sensitivity: 0.99, specificity: 0.15). The MLP had lower performance in this setting (AUROC = 0.60 ± 0.24) with 0 sensitivity and a Brier score of 0.19. Model 11 (Clinical + IL‐1β + Age + Sex) achieved the highest AUROC overall with PGM (0.88 ± 0.24), along with a reasonable sensitivity (0.55) and specificity (0.81), and a Brier score of 0.14. LR and MLP with the same features reached AUROC values of 0.72 ± 0.10 and 0.64 ± 0.24, respectively. Although MLP had a slightly lower Brier score (0.12), its sensitivity remained 0, suggesting poor utility for detecting progression cases. Similarly, Model 8 (Clinical + IL‐1β) achieved high AUROC with PGM (0.87 ± 0.24), while LR and MLP again trailed behind at 0.72 and 0.66, respectively. Across all models, MLP consistently yielded 0 sensitivity with perfect specificity, reflecting poor discrimination despite calibration improvements (lower Brier scores). LR, while performing adequately in terms of AUROC and accuracy, demonstrated unbalanced classification behaviour, often predicting all samples as positive (sensitivity = 1.0, specificity = 0). In contrast, PGM achieved the best trade‐off between discrimination and calibration, with strong AUROC, reasonable sensitivity/specificity balance and competitive Brier scores (range: 0.13–0.21 across models), establishing it as the most clinically meaningful and interpretable approach for predicting periodontitis progression. Discrimination and calibration statistics are presented in Table 4, including AUCROC, AUPRC, sensitivity, specificity, Brier score and accuracy.
TABLE 4.
AUC model results for all combinations using three different approaches (PGM, MLP and LR).
| Model | Probabilistic graph model | Logistic regression | Multi‐layer perceptron with Keras Turner | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy | AUROC | AUPRC | Sensitivity | Specificity | Brier score | Accuracy | AUROC | AUPRC | Sensitivity | Specificity | Accuracy | AUROC | AUPRC | Sensitivity | Specificity | Brier score | |
| Clinical + IL‐1β + age + sex | 0.79 ± 0.80 | 0.88 ± 0.24 | 0.20 ± 0.18 | 0.55 | 0.81 | 0.14 | 0.87 ± 0.04 | 0.72 ± 0.10 | 0.35 ± 0.07 | 1 | 0 | 0.87 ± 0.68 | 0.58 ± 0.24 | 0.20 ± 0.14 | 0.0 | 1.0 | 0.12 |
| Clinical + IL‐1β | 0.79 ± 0.80 | 0.87 ± 0.24 | 0.19 ± 0.17 | 0.60 | 0.80 | 0.15 | 0.87 ± 0.04 | 0.72 ± 0.10 | 0.23 ± 0.07 | 1 | 0 | 0.87 ± 0.68 | 0.66 ± 0.23 | 0.37 ± 0.20 | 0.0 | 1.0 | 0.11 |
| Clinical + VEGF +MMP‐8 + IL‐1β | 0.79 ± 0.81 | 0.82 ± 0.24 | 0.20 ± 0.18 | 0.83 | 0.78 | 0.17 | 0.87 ± 0.04 | 0.74 ± 0.10 | 0.24 ± 0.07 | 1 | 0 | 0.87 ± 0.68 | 0.61 ± 0.24 | 0.43 ± 0.22 | 0.0 | 1.0 | 0.12 |
| Clinical + VEGF +MMP‐8 + IL‐1β +age + sex | 0.80 ± 0.78 | 0.82 ± 0.25 | 0.19 ± 0.17 | 0.51 | 0.82 | 0.14 | 0.87 ± 0.04 | 0.73 ± 0.10 | 0.26 ± 0.07 | 1 | 0 | 0.87 ± 0.68 | 0.63 ± 0.23 | 0.40 ± 0.21 | 0.0 |
1.0 |
0.12 |
| Clinical + VEGF | 0.78 ± 0.82 | 0.81 ± 0.25 | 0.19 ± 0.17 | 0.65 | 0.79 | 0.16 | 0.87 ± 0.04 | 0.71 ± 0.10 | 0.22 ± 0.07 | 1 | 0 |
0.87 ± 0.68 |
0.58 ± 0.20 | 0.38 ± 0.20 |
0.0 |
1.0 |
0.14 |
| Clinical + age + sex | 0.79 ± 0.80 | 0.83 ± 0.24 | 0.19 ± 0.17 | 0.52 | 0.81 | 0.14 | 0.87 ± 0.04 | 0.73 ± 0.10 | 0.25 ± 0.07 | 1 | 0 |
0.87 ± 0.68 |
0.66 ± 0.23 | 0.37 ± 0.19 |
0.0 |
1.0 |
0.11 |
| Clinical + VEGF + age + sex | 0.79 ± 0.80 | 0.82 ± 0.25 | 0.19 ± 0.17 | 0.56 | 0.81 | 0.15 | 0.87 ± 0.04 | 0.73 ± 0.10 | 0.26 ± 0.07 | 1 | 0 | 0.87 ± 0.68 | 0.54 ± 0.23 | 0.36 ± 0.20 |
0.0 |
1.0 |
0.12 |
| All features | 0.77 ± 0.83 | 0.77 ± 0.25 | 0.18 ± 0.17 | 1.0 | 0.75 | 0.21 | 0.88 ± 0.03 | 0.78 ± 0.09 | 0.42 ± 0.10 | 0.99 | 0.15 | 0.87 ± 0.68 | 0.60 ± 0.24 | 0.32 ± 0.19 | 0.0 | 1.0 | 0.11 |
| Clinical + MMP‐8 + age + sex | 0.79 ± 0.81 | 0.77 ± 0.25 | 0.19 ± 0.17 | 0.55 | 0.80 | 0.15 | 0.87 ± 0.04 | 0.71 ± 0.10 | 0.25 ± 0.07 | 1 | 0 | 0.86 ± 0.69 | 0.59 ± 0.23 | 0.35 ± 0.19 |
0.0 |
1.0 |
0.12 |
| Clinical + MMP‐8 | 0.79 ± 0.81 | 0.76 ± 0.25 | 0.19 ± 0.17 | 0.62 | 0.79 | 0.16 | 0.87 ± 0.04 | 0.71 ± 0.10 | 0.22 ± 0.07 | 1 | 0 |
0.87 ± 0.68 |
0.59 ± 0.23 | 0.35 ± 0.19 |
0.0 |
1.0 |
0.11 |
| Clinical + VEGF + MMP‐8 + age + sex | 0.79 ± 0.81 | 0.74 ± 0.25 | 0.19 ± 0.17 | 0.75 | 0.77 | 0.18 | 0.87 ± 0.04 | 0.71 ± 0.10 | 0.25 ± 0.07 | 1 | 0 |
0.87 ± 0.68 |
0.58 ± 0.24 | 0.28 ± 0.17 |
0.0 |
1.0 |
0.13 |
| Clinical + VEGF +MMP‐8 | 0.78 ± 0.81 | 0.74 ± 0.24 | 0.20 ± 0.18 | 0.75 | 0.78 | 0.17 | 0.87 ± 0.04 | 0.71 ± 0.10 | 0.22 ± 0.07 | 1 | 0 |
0.87 ± 0.68 |
0.61 ± 0.24 | 0.21 ± 0.15 |
0.0 |
1.0 |
0.11 |
| Age + sex | 0.80 ± 0.78 | 0.72 ± 0.25 | 0.19 ± 0.17 | 0.47 | 0.82 | 0.13 | 0.87 ± 0.04 | 0.56 ± 0.10 | 0.14 ± 0.05 | 1 | 0 |
0.87 ± 0.68 |
0.54 ± 0.24 | 0.49 ± 0.22 |
0.0 |
1.0 |
0.14 |
SHAP was used to interpret the machine learning model predictions by assigning each variable a SHAP value, which quantifies its contribution to predicting disease progression (stable or progressing). For our analysis, we visualised feature importance based on the model that identified the most relevant variables. This model included clinical data, VEGF, MMP‐8, IL‐1β, age and sex. In the context of the PGM, the features that had the most significant impact on progression score predictions were the number of DP ≥ 5 mm, IL‐1β, VEGF and MMP‐8 (Figure 3a). These features were key in distinguishing between stable and progressing cases. For the MLP model (Figure 3b), the features with the highest absolute SHAP values were disease stage, plaque, number of pockets ≥ 5 mm, followed by MMP‐8 and VEGF. These variables played a pivotal role in enhancing the predictive performance of the MLP model.
FIGURE 3.

(a) PGM: The plot illustrates feature importance scores for the PGM, highlighting that number of DP ≥ 5 mm are the most influential features affecting model performance. Cytokine IL‐1β and VEGF are also significant but secondary in importance. (b) MLP: The plot shows feature importance for the MLP model, with Disease stage following by plaque and number of DP ≥ 5 mm being the most critical feature for model accuracy. This is followed by cytokines that collectively substantially impact model performance.
4. Discussion
In this study, we used ML techniques to evaluate the predictive performance of three models for forecasting periodontitis progression. MLP, a neural network architecture capable of modelling complex non‐linear relationships, and the PGM, which encodes conditional dependencies among variables, were assessed alongside LR. Despite their methodological differences, all models demonstrated the potential of leveraging high‐dimensional data for predictive modelling. Our results indicate that incorporating multidimensional inputs—beyond traditional clinical features—substantially improves predictive accuracy. Among all models tested, PGM achieved the most robust performance, with balanced accuracy scores of ~0.79–0.80, AUROC values up to 0.88 and a well‐maintained sensitivity–specificity balance. These findings underscore the model's suitability for early identification of patients at risk of progression. In contrast, MLP performance was suboptimal, with AUROC values ranging from ~0.54 to 0.66 and zero sensitivity across all configurations, despite achieving perfect specificity. This suggests limited clinical applicability in its current form. The reduced performance may stem from the model's complexity—multiple hidden layers increase the risk of overfitting or underfitting, particularly in small datasets. A more parsimonious architecture, such as a single‐ or dual‐layer design, may yield better results. This aligns with previous findings suggesting that model simplification can enhance performance and generalisability in similar clinical settings (Yu et al. 2019). Overall, our findings emphasise the importance of selecting models that align with the size and structure of the dataset. Tailoring model complexity can improve interpretability, explainability and, ultimately, clinical utility in predicting periodontitis progression (Zhong et al. 2023).
When comparing ML algorithms to predict tooth loss, Troiano et al. (2023) reported the highest accuracy using Naïve Bayes, a type of PGM, with an accuracy of 0.97 when all features were included across multiple cohorts. In our cohort, the PGM model performed better than other models, but the highest accuracy was achieved with a subset of features rather than including all features together, supporting the notion that including all features may introduce noise that hinders predictive accuracy rather than improving accuracy. When using only age and sex as input features, PGM achieves an AUROC of 0.72 (±0.25) and shows improvement once clinical features are added (AUROC = 0.83 ± 0.24), which is consistent across all three methods. There is marginal improvement across all methods when IL‐1β is included. While age and sex are easily included in all models, the potentially enhanced predictive ability of subsequent models with greater feature inclusion is important to consider in future validation and implementation efforts. To optimise future models, research should focus on advanced feature selection or dimensionality reduction techniques that minimise the influence of irrelevant variables while preserving the predictive value of clinically significant ones. Such strategies will improve model robustness and accuracy and contribute to the development of streamlined, cost‐effective and clinically applicable tools for the prediction of periodontitis progression (Xiong et al. 2006; Gupta and Gupta 2019).
While the classification of progression relies primarily clinical parameters (i.e., staging and grading) Papapanou et al. (2018) emphasises the need to develop methodologies tailored for longitudinal, multidimensional data, encompassing biological and clinical insights alongside prognostic evaluations (Papapanou et al. 2018). Few prediction models for periodontitis incidence and progression exist (Farina et al. 2023; Lang and Tonetti 2003). Further, they show limited predictive performance and clinical usefulness (Du et al. 2018) and cannot predict the disease in its initial stages, as their use in periodontally healthy patients is limited (Farina et al. 2023). Here we selected characteristics that can be easily obtained by clinicians and incorporated biological data from saliva, leveraging its accessibility. One example integrated into the model is % BOP: it can be easily determined and is demonstrated indicator of bone loss (Matuliene et al. 2008). Similarly, the number of sites PD ≥ 5 mm, also simple to assess, have been identified as an indicator of potential stability or treatment success among patients with periodontitis (Feres et al. 2020). In our study, this feature ranked highest in the feature importance analysis, suggesting its strong predictive value for disease progression.
Our model incorporated 17 features, with the most relevant combination being clinical data, IL‐1β, age and sex. Notably, IL‐1β—a cytokine closely linked to inflammation and disease severity (Balu et al. 2024)—emerged as a key predictor for disease progression, consistent with its established role in inflammatory processes (F. R. F. Teles, Chandrasekaran, et al. 2024; F. Teles, Martin, et al. 2024). In our SHAP plot, IL‐1β emerged as an important feature for model importance. Several studies (Blanco‐Pintos et al. 2023; F. R. F. Teles, Chandrasekaran, et al. 2024; Tomas et al. 2017) have shown differences in cytokine levels and profiles between individuals with and without periodontitis. While cytokine analysis is not yet a routine clinical tool, its predictive value underscores the need for rapid, accessible tests—similar to those used for COVID‐19 or pregnancy detection.
One of the limitations of this study is the exclusion of diabetes and smokers, which results from the fact that the goal of the original was to study the natural periodontitis progression in the absence of those factors. Another limitation is our sample size. Despite being a large study in the context of periodontitis progression, that is not the case for AI approaches, where algorithms require about 1000 samples for effective performance. However, it is rare for periodontitis studies to reach those number of participants, except for those utilising HER and radiographs (Patel et al. 2023, 2022; Swinckels et al. 2025). Yet, they carry other potential issues, such as excluding certain patient populations due to lack of access to care and the use of self‐reported oral health measures, which do not capture the true biological and inflammatory state of participants (Du et al. 2018). Obtaining prospectively collected longitudinal data, in the absence of treatment, on more than 415 patients, including periodontally healthy participants, as done in our study, is extremely difficult, with cost, logistics and patient attrition presenting significant hurdles. However, it creates a unique dataset that needs to be explored. We acknowledge that our sample size falls considerably below the 1000 threshold, making our model more prone to overfitting. Thus, we attempted to address this both in the model architecture, through dropout regularisation, and in the model training and evaluation, through five‐fold cross‐sectional internal validation. We recognise that, for greater generalisability, the external validation of our model on a completely new and distinct dataset would be ideal. However, to our knowledge, no other comparable dataset exists and creating a new one would be time‐consuming and cost‐prohibitive. We also excluded ensemble methods such as Random Forests (RF) and Gradient Boosted Trees (GBT), despite their advantages in handling imbalanced data and modelling complex non‐linear relationships. This decision was guided by our dataset's limited size, which is better suited to models like logistic regression, PGM and MLP, all of which perform more reliably with smaller sample sizes. Additionally, we prioritised model interpretability to support clinical decision‐making and foster trust among healthcare providers. This choice represents a deliberate trade‐off between maximising predictive performance and ensuring practical clinical utility.
Future research should include additional predictive algorithms such as random forests, decision trees and other neural network models. Also, employing data‐driven biologically informed feature selection may reduce selection bias and improve the accuracy and generalisability of the model, so that it performs well across diverse datasets and patient populations. We believe that future iterations of our models will benefit from incorporating data from additional visits (rather than baseline only, as done here) and integrating microbial and metabolic data.
In conclusion, our findings demonstrate that ML algorithms hold promise in predicting periodontitis progression, offering a valuable tool for early detection and personalised disease management. By integrating clinical variables and biological markers—particularly IL‐1β—our approach outperforms models based solely on clinical data, achieving improved predictive accuracy. These results highlight the potential of data‐driven strategies in advancing periodontal care. Nevertheless, further research is warranted to externally validate these models and explore their applicability in real‐world clinical settings, ensuring robustness, scalability and clinical utility.
Author Contributions
Substantial contributions to the conception or design of the work: Camila Furquim, Lannawill Caruth, Ganesh Chandrasekaran, Andrew Cucchiara, Michael J. Kallan, Lynn Martin, Magda Feres, Kyle Bittinger, Kimon Divaris, Joseph Glessner, Alpdogan Kantarci, William Giannobile, Shefali Setia Verma, Flavia Teles. Acquisition and analysis: Renata Tavares, Lina Suarez, Belen Retamal‐Valdez. Interpretation of data for the work: Camila Furquim, Lannawill Caruth, Ganesh Chandrasekaran, Andrew Cucchiara, Michael J. Kallan, Lynn Martin, Magda Feres, Kyle Bittinger, Kimon Divaris, Joseph Glessner, Alpdogan Kantarci, William Giannobile, Shefali Setia Verma, Flavia Teles. Drafting the work or revising it critically for important intellectual content: Camila Furquim, Lannawill Caruth, Ganesh Chandrasekaran, Andrew Cucchiara, Michael J. Kallan, Lynn Martin, Magda Feres, Kyle Bittinger, Kimon Divaris, Joseph Glessner, Alpdogan Kantarci, William Giannobile, Shefali Setia Verma, Flavia Teles. Final approval of the version to be published: Camila Furquim, Lannawill Caruth, Ganesh Chandrasekaran, Andrew Cucchiara, Michael J. Kallan, Lynn Martin, Magda Feres, Kyle Bittinger, Kimon Divaris, Joseph Glessner, Alpdogan Kantarci, William Giannobile, Shefali Setia Verma, Flavia Teles. Agreement to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved: Camila Furquim, Lannawill Caruth, Ganesh Chandrasekaran, Andrew Cucchiara, Michael J. Kallan, Lynn Martin, Magda Feres, Kyle Bittinger, Kimon Divaris, Joseph Glessner, Alpdogan Kantarci, William Giannobile, Shefali Setia Verma, Flavia Teles.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Figure S1. (a) Study design; (b) flow chart of subject recruitment for the study. During the study, all sites present were monitored with comprehensive periodontal examinations; serum and saliva samples were obtained at each visit. (c) Salivary levels of analytes were determined using bead‐based immunoassays.
Figure S2. Example of ML model architecture graph. Created using visual keras to illustrate connections between input, hidden and output layers.
Figure S3. Correlation matrix. Heatmap displaying correlation between all features collected from with a range of (−1.0, 1.0) with ±1.0 representing perfect correlation (dark red, dark blue).
Figure S4. Workflow diagram illustrating flow of data inputs from Teles et al. (2018) dataset to the models tested and reported on.
Figure S5. Accuracy versus epoch diagram for one training iteration (50 epochs) of MLP demonstrating relationship between training and validation set.
Data S1. Supporting Information.
Acknowledgements
This study was supported by research grants U01‐DE021127 (NIH/NIDCR), DE021127 and DE033033 from the National Institute of Dental and Craniofacial Research, the UPenn Schoenleber Pilot Grant, IDEA Finalist Prize (UPenn CiPD 2023), R01DE033033‐01A1 (NIH/NIDCR), Josephine and Joseph Rabinowitz Research Award (Penn Dental Medicine), CiPD‐IBI Artificial Intelligence in Oral Health Innovation Award and the Center for Human Phenomic Science (CHPS) CTSA grant number 2UL1TR001878‐06. First author was financed in part by the Conselho Nacional de Desenvolvimento Científico e Tecnológico—CNPq Number# 200154/2022‐2, and in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior—CAPES. The authors would like to acknowledge, in memoriam, the invaluable contributions of Dr. Ricardo Teles and Dr. Robert Genco. The authors would also like to acknowledge the contributions of John S. Preisser, Kevin Moss, Patricia Corby, Nathalia Garcia, Heather Jared, Gay Torresyap, Elida Salazar, Julie Moya, Cynthia Howard, Robert Schifferle, Karen L. Falkner, Jane Gillespie, Debra Dixon and MaryAnn Cugini in the clinical aspects of the study.
Furquim, C. P. , Caruth L., Chandrasekaran G., et al. 2025. “Developing Predictive Models for Periodontitis Progression Using Artificial Intelligence: A Longitudinal Cohort Study.” Journal of Clinical Periodontology 52, no. 10: 1478–1490. 10.1111/jcpe.14194.
Funding: This study was supported by research grants U01‐DE021127 and R01DE033033‐01A1 from the National Institutes of Health/National Institute of Dental and Craniofacial Research (NIH/NIDCR), the UPenn Schoenleber Pilot Grant, the Josephine and Joseph Rabinowitz Research Award (Penn Dental Medicine), the CiPD‐IBI Artificial Intelligence in Oral Health Innovation Award, the IDEA Finalist Prize (UPenn CiPD) and the Center for Human Phenomic Science (CHPS) CTSA grant number 2UL1TR001878‐06. First author was financed in part by the Conselho Nacional de Desenvolvimento Científico e Tecnológico—CNPq Number# 200154/2022‐2, and in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior—CAPES
Contributor Information
Shefali Setia Verma, Email: shefali.setiaverma@pennmedicine.upenn.edu.
Flavia Teles, Email: fteles@upenn.edu.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
References
- Adeoye, J. , and Su Y. X.. 2024. “Artificial Intelligence in Salivary Biomarker Discovery and Validation for Oral Diseases.” Oral Diseases 30, no. 1: 23–37. [DOI] [PubMed] [Google Scholar]
- Balu, P. , Balakrishna Pillai A. K., Mariappan V., and Ramalingam S.. 2024. “Cytokine Levels in Gingival Tissues as an Indicator to Understand Periodontal Disease Severity.” Current Research in Immunology 5: 100080. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bashir, N. Z. , Rahman Z., and Chen S. L. S.. 2022. “Systematic Comparison of Machine Learning Algorithms to Develop and Validate Predictive Models for Periodontitis.” Journal of Clinical Periodontology 49, no. 10: 958–969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bikku, T. 2020. “Multi‐Layered Deep Learning Perceptron Approach for Health Risk Prediction.” Journal of Big Data 7, no. 1: 50. [Google Scholar]
- Blanco‐Pintos, T. , Regueira‐Iglesias A., Seijo‐Porto I., et al. 2023. “Accuracy of Periodontitis Diagnosis Obtained Using Multiple Molecular Biomarkers in Oral Fluids: A Systematic Review and Meta‐Analysis.” Journal of Clinical Periodontology 50, no. 11: 1420–1443. [DOI] [PubMed] [Google Scholar]
- Bokhare, A. , Bhagat A., and Bhalodia R.. 2023. “Multi‐Layer Perceptron for Heart Failure Detection Using SMOTE Technique.” SN Computer Science 4, no. 2: 1–9. [Google Scholar]
- Bumm, C. V. , Ern C., Folwaczny J., et al. 2024. “Periodontal Grading—Estimation of Responsiveness to Therapy and Progression of Disease.” Clinical Oral Investigations 28, no. 5: 289. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Caruth, L. , and Verma S.. 2025. “Setia‐Verma‐Lab/Periodontitis_Progression_Prediction: Release1_042025 (v.1.0.0‐Beta).” Zenodo. 10.5281/zenodo.15224694. [DOI]
- Chen, M. X. , Zhong Y. J., Dong Q. Q., Wong H. M., and Wen Y. F.. 2021. “Global, Regional, and National Burden of Severe Periodontitis, 1990–2019: An Analysis of the Global Burden of Disease Study 2019.” Journal of Clinical Periodontology 48, no. 9: 1165–1188. [DOI] [PubMed] [Google Scholar]
- Cho, Y. D. , Kim W. J., Ryoo H. M., et al. 2021. “Current Advances of Epigenetics in Periodontology From ENCODE Project: A Review and Future Perspectives.” Clinical Epigenetics 13, no. 1: 1–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Collins, G. S. , Reitsma J. B., Altman D. G., and Moons K. G. M.. 2015. “Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD): The TRIPOD Statement.” BMC Medicine 13, no. 1: 1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Deng, K. , Zonta F., Yang H., Pelekos G., and Tonetti M. S.. 2023. “Development of a Machine Learning Multiclass Screening Tool for Periodontal Health Status Based on Non‐Clinical Parameters and Salivary Biomarkers.” Journal of Clinical Periodontology 51, no. 12: 1–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Du, M. , Bo T., Kapellas K., and Peres M. A.. 2018. “Prediction Models for the Incidence and Progression of Periodontitis: A Systematic Review.” Journal of Clinical Periodontology 45, no. 12: 1408–1420. [DOI] [PubMed] [Google Scholar]
- Enevold, C. , Nielsen C. H., Christensen L. B., et al. 2023. “Suitability of Machine Learning Models for Prediction of Clinically Defined Stage III/IV Periodontitis From Questionnaires and Demographic Data in Danish Cohorts.” Journal of Clinical Periodontology 51, no. 12: 1561–1573. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fardal, Ø. , Skau I., and Grytten J.. 2023. “A 30‐Year Retrospective Cohort Outcome Study of Periodontal Treatment of Stages III and IV Patients in a Private Practice.” Journal of Clinical Periodontology 52: 102–112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Farina, R. , Lopez R., Simonelli A., and Trombelli L.. 2023. “Accuracy and Applicability of Periodontitis Risk Assessment Tools: A Critical Appraisal.” Periodontology 2000: 1–18. [DOI] [PubMed] [Google Scholar]
- Feres, M. , Retamal‐valdes B., Faveri M., et al. 2020. “Proposal of a Clinical Endpoint for Periodontal Trials: The Treat‐To‐ Target Approach.” Journal of the International Academy of Periodontology 22, no. 2: 41–53. [PubMed] [Google Scholar]
- Gupta, A. , Slater J. J., Boyne D., et al. 2019. “Probabilistic Graphical Modeling for Estimating Risk of Coronary Artery Disease: Applications of a Flexible Machine‐Learning Method.” Medical Decision Making: An International Journal of the Society for Medical Decision Making 39, no. 8: 1032–1044. [DOI] [PubMed] [Google Scholar]
- Gupta, S. , and Gupta A.. 2019. “Dealing With Noise Problem in Machine Learning Data‐Sets: A Systematic Review.” Procedia Computer Science 161: 466–474. [Google Scholar]
- Jain, N. , Dutt U., Radenkov I., and Jain S.. 2023. “WHO's Global Oral Health Status Report 2022: Actions, Discussion and Implementation.” Oral Diseases 30, no. 2: 73–79. [DOI] [PubMed] [Google Scholar]
- Kang, M. , Ko E., and Mersha T. B.. 2022. “A Roadmap for Multi‐Omics Data Integration Using Deep Learning.” Briefings in Bioinformatics 23, no. 1: 1–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lang, N. P. , and Tonetti M. S.. 2003. “Periodontal Risk Assessment (PRA) for Patients in Supportive Periodontal Therapy (SPT).” Oral Health & Preventive Dentistry 1, no. 1: 7–16. [PubMed] [Google Scholar]
- Larrañaga, P. , and Moral S.. 2011. “Probabilistic Graphical Models in Artificial Intelligence.” Applied Soft Computing 11, no. 2: 1511–1528. [Google Scholar]
- Leite, F. R. M. , Peres K. G., Do L. G., Demarco F. F., and Peres M. A. A.. 2017. “Prediction of Periodontitis Occurrence: Influence of Classification and Sociodemographic and General Health Information.” Journal of Periodontology 88: 731–743. [DOI] [PubMed] [Google Scholar]
- Lindskog, S. , Blomlof J., Persson I., et al. 2010. “Validation of an Algorithm for Chronic Periodontitis Risk Assessment and Prognostication: Risk Predictors, Explanatory Values, Measures of Quality, and Clinical Use.” Journal of Periodontology 81: 584–593. [DOI] [PubMed] [Google Scholar]
- Lundberg, S. M. , and Lee S. I.. 2017. “A Unified Approach to Interpreting Model Predictions.” Advances in Neural Information Processing Systems 30: 4766–4775. [Google Scholar]
- Martinez‐Canut, P. 2015. “Predictors of Tooth Loss due to Periodontal Disease in Patients Following Long‐Term Periodontal Maintenance.” Journal of Clinical Periodontology 42: 1115–1125. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Martorell‐Marugán, J. , Tabik S., Benhammou Y., et al. 2019. Deep Learning in Omics Data Analysis and Precision Medicine, 37–53. Computational Biology. [PubMed] [Google Scholar]
- Matuliene, G. , Pjetursson B. E., Salvi G. E., et al. 2008. “Influence of Residual Pockets on Progression of Periodontitis and Tooth Loss: Results After 11 Years of Maintenance.” Journal of Clinical Periodontology 35, no. 8: 685–695. [DOI] [PubMed] [Google Scholar]
- Morelli, T. , Moss K. L., Preisser J. S., et al. 2018. “Periodontal Profile Classes Predict Periodontal Disease Progression and Tooth Loss.” Journal of Periodontology 89, no. 2: 148–156. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Obermeyer, Z. , and Emanuel E. J.. 2016. “Predicting the Future—Big Data, Machine Learning, and Clinical Medicine.” New England Journal of Medicine 375, no. 13: 1216–1219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ossowska, A. , Kusiak A., and Świetlik D.. 2022. “Evaluation of the Progression of Periodontitis With the Use of Neural Networks.” Journal of Clinical Medicine 11, no. 16: 4667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ossowska, A. , Kusiak A., and Świetlik D.. 2023. “Progression of Selected Parameters of the Clinical Profile of Patients With Periodontitis Using Kohonen's Self‐Organizing Maps.” Journal of Personalized Medicine 13, no. 2: 1–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Papapanou, P. N. , Sanz M., Buduneli N., et al. 2018. “Periodontitis: Consensus Report of Workgroup 2 of the 2017 World Workshop on the Classification of Periodontal and Peri‐Implant Diseases and Conditions.” Journal of Clinical Periodontology 45, no. December 2017: S162–S170. [DOI] [PubMed] [Google Scholar]
- Patel, J. S. , Su C., Tellez M., et al. 2022. “Developing and Testing a Prediction Model for Periodontal Disease Using Machine Learning and Big Electronic Dental Record Data.” Frontiers in Artificial Intelligence 5: 979525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Patel, J. S. , Kumar K., Zai A., Shin D., Willis L., and Thyvalikakath T. P.. 2023. “Developing Automated Computer Algorithms to Track Periodontal Disease Change From Longitudinal Electronic Dental Records.” Diagnostics 13, no. 6: 1028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Scott, J. , Biancardi A. M., Jones O., and Andrew D.. 2023. “Artificial Intelligence in Periodontology: A Scoping Review.” Dentistry Journal 11, no. 2: 43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Serroni, M. , Ravidà A., Santamaria P., et al. 2025. “Multi‐Centre External Validation of a Nomogram for 10‐Year Periodontal Tooth Loss Prediction.” Journal of Clinical Periodontology 52, no. 7: 1044–1055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Swinckels, L. , de Keijzer A., Loos B. G., et al. 2025. “A Personalized Periodontitis Risk Based on Nonimage Electronic Dental Records by Machine Learning.” Journal of Dentistry 153: 105469. [DOI] [PubMed] [Google Scholar]
- Teles, F. R. F. , Chandrasekaran G., Martin L., et al. 2024. “Salivary and Serum Inflammatory Biomarkers During Periodontitis Progression and After Treatment.” Journal of Clinical Periodontology 51: 1619–1631. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teles, F. , Martin L., Patel M., et al. 2024. “Gingival Crevicular Fluid Biomarkers During Periodontitis Progression and After Periodontal Treatment.” Journal of Clinical Periodontology 52: 40–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teles, R. , Benecha H. K., Preisser J. S., et al. 2016. “Modelling Changes in Clinical Attachment Loss to Classify Periodontal Disease Progression.” Journal of Clinical Periodontology 43, no. 5: 426–434. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Teles, R. , Moss K., Preisser J. S., et al. 2018. “Patterns of Periodontal Disease Progression Based on Linear Mixed Models of Clinical Attachment Loss.” Journal of Clinical Periodontology 45, no. 1: 15–25. [DOI] [PubMed] [Google Scholar]
- Theodosiou, A. A. , and Read R. C.. 2023. “Artificial Intelligence, Machine Learning and Deep Learning: Potential Resources for the Infection Clinician.” Journal of Infection 87, no. 4: 287–294. [DOI] [PubMed] [Google Scholar]
- Tomas, I. , Regueira‐Iglesias A., Lopez M., et al. 2017. “Quantification by qPCR of Pathobionts in Chronic Periodontitis: Development of Predictive Models of Disease Severity at Site‐Specific Level.” Frontiers in Microbiology 8: 1443. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Troiano, G. , Nibali L., Petsos H., et al. 2023. “Development and International Validation of Logistic Regression and Machine‐Learning Models for the Prediction of 10‐Year Molar Loss.” Journal of Clinical Periodontology 50, no. 3: 348–357. [DOI] [PubMed] [Google Scholar]
- Xiong, H. , Pandey G., Steinbach M., and Kumar V.. 2006. “Enhancing Data Analysis With Noise Removal.” IEEE Transactions on Knowledge and Data Engineering 18, no. 3: 304–319. [Google Scholar]
- Yu, H. , Samuels D. C., Zhao Y., and Guo Y.. 2019. “Architectures and Accuracy of Artificial Neural Network for Disease Classification From Omics Data.” BMC Genomics 20, no. 1: 167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhong, Y. , Peng Y., Lin Y., et al. 2023. “MODILM: Towards Better Complex Diseases Classification Using a Novel Multi‐Omics Data Integration Learning Model.” BMC Medical Informatics and Decision Making 23, no. 1: 1–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figure S1. (a) Study design; (b) flow chart of subject recruitment for the study. During the study, all sites present were monitored with comprehensive periodontal examinations; serum and saliva samples were obtained at each visit. (c) Salivary levels of analytes were determined using bead‐based immunoassays.
Figure S2. Example of ML model architecture graph. Created using visual keras to illustrate connections between input, hidden and output layers.
Figure S3. Correlation matrix. Heatmap displaying correlation between all features collected from with a range of (−1.0, 1.0) with ±1.0 representing perfect correlation (dark red, dark blue).
Figure S4. Workflow diagram illustrating flow of data inputs from Teles et al. (2018) dataset to the models tested and reported on.
Figure S5. Accuracy versus epoch diagram for one training iteration (50 epochs) of MLP demonstrating relationship between training and validation set.
Data S1. Supporting Information.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
