Skip to main content
Patient preference and adherence logoLink to Patient preference and adherence
. 2026 Sep 7;20:621510. doi: 10.2147/PPA.S621510

Diagnostic Prediction Models for Oral Frailty in Older Adults: A Systematic Review and Critical Appraisal

Yuzhu Fan 1,*, Shuang Zhang 1,*, Yinuo Wang 1, Zhirun Cheng 1, Xueer Yu 1, Yuanyuan Chen 1, Hui Shi 1,✉, Wangqin Shen 1,✉
PMCID: PMC13564451  PMID: 42730340

Abstract

Background

Oral frailty is a clinically relevant marker of vulnerability in older adults, but the quality, performance, and applicability of multivariable models for its identification remain uncertain.

Objectives

To identify and critically appraise multivariable prediction models for oral frailty and summarize their predictors, performance, validation, and risk of bias.

Methods

Nine English and Chinese language databases were searched from inception to 14 July 2026, supplemented by citation and website searches. Two reviewers independently selected studies, extracted data, and assessed risk of bias and applicability using the Prediction Model Risk of Bias Assessment Tool (PROBAST). Owing to substantial heterogeneity, findings were synthesized narratively.

Results

Twenty-three reports representing 22 independent studies and 22 models were included. Twenty-one studies were conducted in China, 21 used cross-sectional data, and 21 defined oral frailty as an Oral Frailty Index-8 score ≥4. All models were diagnostic and identified prevalent oral frailty at assessment; none predicted incident oral frailty. Logistic regression and nomograms were the most common approaches, although two studies compared multiple machine-learning algorithms. Common predictors included age, nutritional vulnerability, physical frailty or sarcopenia, chronic disease burden, smoking, and oral-function indicators. Reported area under the receiver operating characteristic curve (AUC) values ranged from 0.713 to 0.985. Internal validation was reported for 20 models, while six models underwent external validation (four temporal and two geographical); one model reported apparent performance only. Calibration was incompletely assessed, and no model had an overall low risk of bias.

Conclusions

Existing models show generally moderate-to-high discrimination, but the evidence remains methodologically immature. Predominantly cross-sectional designs, high risk of bias, predictor-outcome overlap, incomplete calibration, and limited external validation, especially across geographical settings, preclude routine clinical use. Current models should be regarded as candidate diagnostic screening aids that complement standardized assessment; rigorous external validation and prospective prognostic modelling are priorities.

PROSPERO Registration

CRD420251229868.

Keywords: oral frailty, prediction models, systematic review, PROBAST, older adults

Introduction

Population ageing is a major global demographic transition accompanied by increasing multimorbidity, functional decline, and complex care needs among older adults. Frailty has been widely conceptualized as a multidimensional geriatric syndrome characterized by reduced physiological reserve and diminished resilience to stressors, which predisposes individuals to disability, hospitalization, and mortality.1,2 Within this broader construct, oral frailty represents a distinct but often underrecognized domain of vulnerability. It reflects progressive declines in oral functions such as mastication, swallowing, articulation, and oral motor performance.3

Evidence links impaired oral function to malnutrition, sarcopenia, physical frailty, reduced quality of life and social participation, and cognitive impairment.4–6 Because oral function is central to nutrition and communication, timely identification of oral frailty may support targeted assessment and intervention and may help mitigate subsequent functional decline.

Several instruments have been developed to identify oral frailty based on indicators such as tooth loss, masticatory performance, swallowing difficulty, or tongue pressure. Multivariable prediction models may serve either diagnostic or prognostic purposes. Consistent with the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) framework, diagnostic models estimate the probability that oral frailty is present at the time of assessment, whereas prognostic models estimate the probability of its future onset or progression over a specified period.7 Clarifying this distinction is important for interpreting the intended clinical role of a model.

Multivariable models for oral frailty have recently been developed using different predictors and modelling approaches. Prediction model studies, however, often have methodological and reporting limitations, including inappropriate predictor selection, insufficient validation, inadequate assessment of calibration, and incomplete reporting.8,9 These limitations may lead to optimistic estimates of model performance and restrict clinical applicability.

A recent scoping review synthesized risk prediction models for oral frailty in older adults and highlighted methodological limitations and the need for stronger validation.10 However, the evidence base has continued to expand, and important issues regarding the intended diagnostic or prognostic purpose of these models, their validation strategies, calibration, risk of bias, and clinical applicability warrant updated critical appraisal. Building on this prior evidence synthesis, the present systematic review aimed to: (1) identify multivariable prediction models for oral frailty and characterize their diagnostic or prognostic purpose; (2) evaluate their risk of bias and applicability using the Prediction Model Risk of Bias Assessment Tool (PROBAST); and (3) summarize their predictors, validation strategies, and model performance, including discrimination and calibration.

Methods

This systematic review was prospectively registered in PROSPERO (CRD420251229868) and is reported in accordance with the Transparent Reporting of Multivariable Prediction Models for Individual Prognosis or Diagnosis: Systematic Reviews and Meta-Analyses (TRIPOD-SRMA) guideline.11

Search Strategy

A comprehensive literature search was conducted in PubMed, Embase, the Cochrane Library, CINAHL, and Web of Science, and four Chinese databases: China National Knowledge Infrastructure (CNKI), WanFang Data, the VIP Database, and the Chinese Biomedical Literature Database (CBM), from database inception to 14 July 2026. The search combined terms related to oral frailty and prediction modelling, using database-specific subject headings and free-text terms as appropriate. Core search terms included “oral frailty”, “oral hypofunction”, “oral dysfunction”, and “oral functional decline”, combined with terms such as “predict*”, “prediction model*”, “predictive model*”, “risk prediction”, “risk model*”, “nomogram*”, “validat*”, “machine learning”, and “artificial intelligence”. Equivalent Chinese-language terms were used for the Chinese databases, with search syntax adapted to the requirements of each database. No restrictions on study design were applied at the search stage. In addition to the database searches, supplementary searches were conducted to identify potentially eligible reports not retrieved through the electronic database search. Citation searching was performed by screening the reference lists of included studies and relevant reviews, and additional website searches were also undertaken.

Eligibility Criteria

Inclusion criteria were as follows: (1) studies that developed and/or validated a multivariable model for diagnosing prevalent oral frailty or predicting its future onset or progression; (2) studies that included adults aged ≥60 years, or provided separable results for participants aged ≥60 years; (3) studies that reported at least one measure of model performance, including discrimination and/or calibration; and (4) studies in which oral frailty was defined using a previously validated assessment instrument or explicit criteria reported in the literature. Only full-text reports published in English or Chinese were eligible.

Exclusion criteria were as follows: (1) study protocols, reviews, conference abstracts, letters, editorials, or comments; (2) studies evaluating a single predictor rather than developing or validating a multivariable prediction model; (3) studies in which oral frailty was used as a predictor of another outcome rather than as the target outcome; (4) reports with insufficient methodological information for model-level data extraction or risk-of-bias assessment; and (5) studies evaluating the diagnostic accuracy of a newly proposed oral frailty instrument without developing or validating a multivariable prediction model.

Study Selection

All identified records were imported into EndNote X9 for reference management and deduplication. After duplicate removal, titles and abstracts were independently screened by two reviewers. Records considered potentially eligible by either reviewer were retrieved for full-text assessment. The same two reviewers independently assessed the full texts against the predefined eligibility criteria. Disagreements were resolved through discussion or, when necessary, consultation with a third reviewer.

Data Extraction

Data were extracted using a standardized form developed a priori. Extracted data included study design, setting, sample size and number of outcome events; outcome definition; candidate and retained predictors; modelling and predictor-selection methods; missing-data handling; validation procedures; and measures of discrimination, calibration, and clinical utility. Two reviewers independently extracted data from the full texts of included studies. Where multiple reports described the same or an overlapping cohort, the reports were linked and treated as a single study; the most complete report was used as the primary source, with companion reports consulted for supplementary information. Disagreements were resolved through discussion. For studies comparing multiple algorithms, the model identified by the study authors as the final or preferred model was used for model-level counting and primary synthesis, whereas the performance of alternative algorithms was summarized descriptively.

Risk of Bias Assessment

Two reviewers independently assessed the risk of bias and applicability of the included models using the Prediction Model Risk of Bias Assessment Tool (PROBAST).9 Before the formal assessment, the two reviewers conducted a calibration exercise on a subset of included studies to standardize the interpretation of the PROBAST signalling questions and domain-level judgement criteria. PROBAST evaluates risk of bias across four domains—participants, predictors, outcome, and analysis—and applicability across the first three domains. Signalling questions were rated as “yes”, “probably yes”, “probably no”, “no”, or “no information”, and domain-level judgements were classified as low, high, or unclear risk of bias or applicability concern. Overall risk of bias was rated as low only when all four domains were judged as low risk, high when at least one domain was judged as high risk, and unclear otherwise. Overall applicability was judged analogously across the participants, predictors, and outcome domains. For studies comparing multiple algorithms, the final or preferred model used in the primary synthesis was used for PROBAST assessment. Disagreements were resolved through discussion.

Data Synthesis

Given substantial clinical and methodological heterogeneity in populations, outcome definitions, predictor measurement and coding, modelling approaches, and performance reporting, quantitative pooling was not undertaken. Study and model characteristics were synthesized narratively, with apparent, internally validated, and externally validated performance distinguished where possible. Internal validation included bootstrap resampling, cross-validation, and random split-sample procedures using the model-development data. External validation was defined as evaluation in a separately collected sample not used for model development, including temporal or geographical validation where applicable. Two descriptive sensitivity analyses were conducted: one restricted the synthesis to studies explicitly recruiting community-dwelling older adults, and the other excluded the single model reporting apparent performance only. Findings from the sensitivity analyses were compared qualitatively with the main synthesis.

Results

Search results

The searches identified 3889 records. After deduplication and screening, 263 full-text reports were assessed for eligibility, and 23 reports representing 22 independent studies and 22 diagnostic prediction models for prevalent oral frailty were included.12–34 Jiang et al19 and Liu et al28 reported an overlapping Anhui cohort and were therefore treated as two reports of one study. The detailed study-selection process and reasons for exclusion are presented in Figure 1.

Figure 1.

A flowchart of study selection process for oral frailty prediction models. The flowchart outlines the study selection process for oral frailty prediction models. It begins with the identification of studies via databases and registers, totaling 3,877 records from sources like Pubmed, Embase and others and 12 records from other methods such as websites and citation searching. Duplicate records removed before screening amount to 1,783. The screening phase shows 2,106 records screened, with 1,836 excluded at title or abstract screening. Reports sought for retrieval are 270, with 7 not retrieved. Reports assessed for eligibility are 263, with 240 excluded for reasons such as not being an oral frailty prediction model, ineligible population, or insufficient model-performance data. Finally, 22 studies are included in the review, with 23 reports of included studies.

PRISMA 2020 flow diagram of study selection.

Characteristics of Included Studies

The 23 reports represented 22 independent studies published between 2022 and 2026 (Table 1). Twenty-one studies were conducted in China,12–23,25–34 and one was conducted in Japan.24 Twenty-one studies used cross-sectional data,12–20,22–34 whereas one used a retrospective cohort design.21 All studies evaluated diagnostic models for oral frailty present at the time of assessment; none evaluated the future onset of oral frailty over a prespecified prediction horizon.

Table 1.

Characteristics of Included Studies and Prediction Models for Oral Frailty

Study (Author,Year) Population/Setting Country Study Design Outcome Definition Model Type Sample Size (Events/Total) Predictors (n) Final Predictors
Cheng et al,32 Hospitalized stroke patients >60 y China Cross-sectional OFI-8 ≥4 Five-algorithm comparison; LR selected 169 / 301 4 OHAT score, denture use, activities of daily living, physical frailty
Feng et al,25 Rural hypertensive adults >60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 232 / 377; NR / 161 8 Age, education, living alone, polypharmacy, smoking, dysphagia, xerostomia, comorbidity
Guo et al,21 Hospitalized PD patients ≥60y China Retrospective cohort OFI-8 ≥4 LASSO + Nomogram (LR) NR / 214 6 Age, education, living situation, tobacco exposure, comorbidity burden, imbalanced diet
Li et al,23 Hospitalized adults ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 323 / 780 7 Age, denture use, number of oral medications, dry mouth, frailty, nutritional status, social support
Liu et al,19,28 Community-dwelling adults ≥60y (overlapping Anhui cohort) China Cross-sectional OFI-8 ≥4 Nomogram (LR) 1433 / 3061 6 Hospitalization, depressive symptoms, social isolation, malnutrition, eHealth literacy, subjective cognitive decline
Liu et al,29 Hospitalized older adults with chronic diseases ≥60y China Multicenter cross-sectional OFI-8 ≥4 Nomogram (LR) 307 / 443 5 Age, work status, appetite loss, frailty, clinical physiological resilience
Liu Q et al,33 Community-dwelling adults ≥60 y China Cross-sectional OFI-8 ≥4 Eight-algorithm comparison; SVM + SHAP 939 / 1,457 6 Chronic disease burden, age, depressive symptoms, smoking, malnutrition, physical frailty
Liu W et al,34 Hospitalized older stroke patients ≥60 y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 179 / 286 7 Age, female sex, smoking, diabetes, physical frailty, malnutrition, depressive symptoms
Lu et al,18 Hospitalized cancer patients ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 190 / 406 7 Age, radiotherapy, oral mucositis, grip strength, oral health status, oral health-related self-efficacy, nutritional status
Lv et al,26 Hospitalized esophageal cancer patients ≥60y China Cross-sectional OFI-8 ≥4 LASSO + Nomogram (LR) 250 / 555 6 Radiotherapy, tumor stage, physical frailty, smoking, age, nutritional status
Lv et al,17 Community-dwelling adults ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 258 / 460 5 Falls, sarcopenia, oral health knowledge, oral health beliefs, oral health behaviours
Ma et al,16 Hospitalized first-time stroke patients≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 317 / 664 7 Age, frailty, comorbidity burden, nutritional status, NIHSS score, oral health status, Barthel index
Qiao et al,15 Hospitalized COPD patients ≥60y China Cross-sectional OFI-8 ≥4 Logistic regression + Nomogram 296 / 320 3 Degree of dyspnea, comorbidity burden, nutritional status
Song et al,30 Hospitalized CHF patients ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 176 / 343 6 Age, smoking, physical frailty, malnutrition, polypharmacy, oral health self-efficacy
Wang et al,14 Community-dwelling adults ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 246 / 548 7 Age, denture use, outdoor activity frequency, oral health behaviours, dry mouth, chewing difficulty, choking
Wu et al,13 Community-dwelling adults ≥60y China Cross-sectional OFI-8 ≥4 LASSO + Nomogram (LR) 318 / 586 6 Tobacco exposure, living situation, comorbidity burden, frailty, denture use, nutritional status
Xiao et al,12 Ischaemic stroke patients ≥65y China Cross-sectional OFI-8 ≥4 LASSO + Nomogram 398 / 633 7 Denture use, oral health behaviours, dry mouth, chewing difficulty, swallowing dysfunction, oral health literacy, oral health status
Yamamoto et al,24 Community-dwelling adults ≥65y Japan Cross-sectional ≥3/6 OF items Logistic regression 215 / 843 4 Age, number of teeth, chewing difficulty, choking
Yang et al,27 Hospital-based older adults with T2DM ≥60y China Multicenter cross-sectional OFI-8 ≥4 Nomogram (LR) 246 / 533 8 Age, chronic disease burden, diabetes duration, HbA1c, periodontitis, natural teeth, chewing difficulty, dysphagia
Yang et al,31 Community-dwelling adults ≥60y China Cross-sectional OFI-8 ≥4 LASSO + Nomogram (LR) 126 / 388 9 Sex, age, education, chronic diseases, smoking, natural teeth, chewing difficulty, frailty, oral health self-efficacy
Zhong et al,22 Hospitalized T2DM patients ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 129 / 370 6 Age, BMI, ASMI, tobacco exposure, monthly income, chewing difficulty
Zou et al,20 Rural adults ≥60y China Cross-sectional OFI-8 ≥4 Nomogram (LR) 224 / 595 6 Age, polypharmacy, oral health-related self-efficacy, swallowing dysfunction, denture use, comorbidity burden

Notes: Sample size is reported as the number of outcome events/total participants. Jiang et al (2025)19 and Liu et al (2026)28 reported an overlapping Anhui cohort and were therefore treated as two reports of one independent study; the more complete 2026 report was used as the primary source for data extraction.

Abbreviations: ASMI, appendicular skeletal muscle mass index; BMI, body mass index; CHF, chronic heart failure; COPD, chronic obstructive pulmonary disease; HbA1c, glycated hemoglobin; LASSO, least absolute shrinkage and selection operator; LR, logistic regression; NIHSS, National Institutes of Health Stroke Scale; NR, not reported; OF, oral frailty; OFI-8, Oral Frailty Index-8; OHAT, Oral Health Assessment Tool; PD, Parkinson’s disease; SHAP, SHapley Additive exPlanations; SVM, support vector machine; T2DM, type 2 diabetes mellitus.

Study settings and participant sources varied. Seven independent studies, represented by eight reports, recruited community-dwelling older adults,13,14,17,19,24,28,31,33 two focused on rural older populations,20,25 and 13 were conducted in hospital-based populations.12,15,16,18,21–23,26,27,29,30,32,34 The hospital-based studies included general hospitalized older adults and patients with Parkinson’s disease, cancer, stroke, chronic obstructive pulmonary disease, type 2 diabetes mellitus, chronic heart failure, or other chronic diseases.

Sample sizes ranged from 214 to 3,061 participants, and the reported prevalence of oral frailty ranged from 32.5% to 92.5%. Twenty-one studies defined oral frailty as an Oral Frailty Index-8 score ≥4,12–23,25–34 whereas one classified oral frailty as the presence of at least three of six predefined components.24

Characteristics of Model Development

Model development characteristics are summarized in Table 1. Logistic regression was the predominant modelling approach. Five studies used the least absolute shrinkage and selection operator (LASSO) for predictor selection,12,13,21,26,31 and two compared multiple machine-learning algorithms.32,33 The remaining studies mainly used univariable pre-screening, stepwise selection, or conventional multivariable logistic regression. The final models retained between three and nine predictors.

Nineteen of the 22 models were presented as nomograms. Of the remaining three, one reported a logistic regression equation without a graphical nomogram,24 one selected logistic regression after comparing five algorithms,32 and one implemented a support vector machine model with SHAP-based interpretation and a web-based calculator.33

Despite heterogeneity in candidate predictors and selection strategies, several domains recurred across studies. Age was the most frequently retained predictor. Other common domains included nutritional vulnerability or appetite loss, physical frailty or sarcopenia, multimorbidity or chronic disease burden, and tobacco exposure. Overall, retained predictors spanned three broad domains: demographic and social factors; general health indicators; and oral health and function-related factors, including dentition, denture use, xerostomia, chewing difficulty, dysphagia, and oral health behaviors. Several models also incorporated disease- or context-specific predictors. A comprehensive list of predictors retained in each final model is provided in Table 1.

Characteristics of Model Validation and Performance

Validation strategies and performance metrics are presented in Table 2. Twenty of the 22 models underwent at least one form of internal validation.12–15,17–33 Six models underwent external validation: four temporal validations12,16,17,21 and two geographical validations.13,26 One of the temporally externally validated models did not report a separate internal-validation procedure,16 whereas one model reported apparent performance only.34 Random split-sample validation and bootstrap resampling were the most commonly used internal-validation approaches.

Table 2.

Predictive Performance and Validation of the Included Oral Frailty Prediction Models

Study Discrimination (AUC, 95% CI) Calibration (Hosmer-Lemeshow p-value) Clinical Utility Validation Type
Cheng et al,32 Train 0.907 (0.844–0.970); validation 0.713 (0.415–0.981) NR Validation sensitivity/specificity 0.636/0.653; marked train–validation decline for tree models Split-sample internal validation (70/30) + reported 5-fold cross-validation
Feng et al,25 Train 0.781 (0.735–0.827); bootstrap 0.769 (0.754–0.779); validation 0.813 (0.748–0.879) H–L P: train 0.089; validation 0.447 DCA reported favorable net benefit Bootstrap internal + hold-out internal validation
Guo et al,21 Train 0.863 (0.823–0.903);
Validation 0.882 (0.842–0.922)
H–L P: 0.727 DCA positive above an 8% threshold Bootstrap internal
+ temporal external validation (same hospital, later cohort)
Li et al,23 Train 0.866 (0.837–0.895); Validation 0.887 (0.846–0.928) H–L P: train 0.337; validation 0.466 DCA reported favorable; proposed for bedside use Random split internal validation (70/30)
Liu et al,19,28 Apparent 0.747 (0.729–0.764); bootstrap C-index 0.747 (0.731–0.764) H–L P: 0.437; calibration slope ≈1 Sensitivity/specificity 0.799/0.609 at cutoff 0.414; DCA net benefit at threshold probabilities 18–65% Bootstrap internal validation
Liu et al,29 Train 0.805 (0.753–0.857); Validation 0.900 (0.843–0.957) H–L P: train 0.465; validation 0.161; Brier score 0.158/0.123 DCA reported favorable Bootstrap internal validation
Liu Q et al,33 SVM train 0.814; SVM validation 0.783; LR validation 0.792 SVM Brier score
0.188/0.180 (train/validation)
SVM validation F1 score 0.659; SHAP interpretation; web-based calculator Random split internal validation (70/30) + repeated 10-fold cross-validation for tuning
Liu W et al,34 Apparent 0.916 (0.891–0.943) H–L P: 0.722 Sensitivity/specificity 0.890/0.822; nomogram None (apparent performance only)
Lu et al,18 Train 0.954 (0.934–0.974); Validation 0.977 (0.957–0.998) H–L P: 0.932 DCA reported favorable Random split internal validation (70/30)
Lv et al,17 Train 0.895 (0.854–0.937); internal-validation
0.880 (0.816–0.944); temporal external-test 0.835 (0.760–0.910)
H–L P: train 0.633; internal validation 0.486; temporal external test 0.692 No DCA reported; sensitivity/specificity: train 0.838/0.831, internal validation 0.909/0.762, temporal external test 0.661/0.851 Random split internal validation + temporal external validation (later test cohort)
Lv et al,26 Train 0.812 (0.771–0.853); external validation 0.796 (0.730–0.862) H–L P: train 0.193; external validation 0.093; Brier score 0.176/0.187 DCA reported favorable Bootstrap internal validation + geographical external validation
Ma et al,16 Train 0.945 (0.925–0.965); temporal external-validation
0.915 (0.878–0.952)
H–L P: train 0.688; external validation 0.384 DCA reported favorable; external-validation sensitivity/specificity 0.882/0.802 Temporal external validation (same hospital, later cohort)
Study Discrimination (AUC, 95% CI) Calibration (Hosmer-Lemeshow p-value) Clinical Utility Validation Type
Qiao et al,15 Train 0.970 (0.94–1.00); Validation 0.960 (0.910–1.000) H–L P: train 0.994; validation 0.540 DCA net benefit higher than treat-all and treat-none strategies Random split internal validation (70/30)
Song et al,30 Apparent 0.857 (0.818–0.896); bootstrap optimism-corrected C-index 0.845 H–L P: 0.790; Brier score 0.153 DCA reported favorable Bootstrap internal validation
Wang et al,14 Train 0.950 (0.930–0.970); Validation 0.980 (0.960–1.000) H–L P (train): 0.932; calibration mean absolute error: train 0.009, validation 0.018 DCA reported favorable; train sensitivity/specificity 0.91/0.85 and PPV/NPV 0.75/0.95 Random split internal validation (70/30) + 1,000-bootstrap internal validation
Wu et al,13 Train 0.952 (0.934–0.970); geographical external validation 0.936 (0.902–0.970); bootstrap 0.878 H–L P: train 0.100; external validation 0.112 DCA favorable at thresholds 5–96% (train) and 5–92% (external validation) Bootstrap internal validation + geographical external validation (separate community)
Xiao et al,12 Train 0.985 (0.976–0.994); bootstrap 0.985 (0.975–0.993); temporal external-validation 0.982 (0.967–0.996) H–L P: train 0.343; external validation 0.398 DCA reported favorable; online dynamic nomogram Bootstrap internal validation + temporal external validation (later-period cohort, same hospital)
Yamamoto et al,24 Train 0.890; test 0.860 (0.806–0.915) NR No DCA reported; test sensitivity/specificity 0.904/0.657, PPV/NPV 0.524/0.943, accuracy 0.730 Random split internal validation (70/30)
Yang et al,27 Train 0.847 (0.808–0.886); Validation 0.831 (0.768–0.894) H–L P (validation): 0.094 DCA reported favorable Random split internal validation (70/30) + 1,000-bootstrap internal validation
Yang et al,31 Train 0.945 (0.919–0.970); Validation 0.910 (0.857–0.962) H–L P: 0.915 DCA reported favorable Random split internal validation (70/30) + 1,000-bootstrap internal validation
Zhong et al,22 Development 0.887 (0.847–0.925); hold-out validation 0.839 (0.755–0.923) H–L P: 0.773 NR Bootstrap internal validation + hold-out internal validation
Zou et al,20 Train 0.883 (0.853–0.914); Validation 0.827 (0.759–0.895) H–L P: 0.61 No DCA reported; proposed for rural screening Hold-out internal validation

Notes: Jiang et al (2025)19 and Liu et al (2026)28 reported an overlapping Anhui cohort and were therefore treated as two reports of one independent study; the more complete 2026 report was used as the primary source for data extraction.

Abbreviations: AUC, area under the receiver operating characteristic curve; CI, confidence interval; C-index, concordance index; DCA, decision curve analysis; F1, F1 score; H–L, Hosmer–Lemeshow; LR, logistic regression; NPV, negative predictive value; NR, not reported; PPV, positive predictive value; SHAP, SHapley Additive exPlanations; SVM, support vector machine.

All studies assessed discrimination using the area under the receiver operating characteristic curve (AUC). Reported AUCs ranged from 0.713 to 0.985 across apparent, internally validated, and externally validated evaluations. Across the six externally validated models, AUCs ranged from 0.796 to 0.982: 0.835–0.982 in temporal validation12,16,17,21 and 0.796–0.936 in geographical validation.13,26 The model reporting apparent performance only had an AUC of 0.916 (95% CI 0.891–0.943) and underwent no internal or external validation.34

Calibration was most commonly assessed using the Hosmer-Lemeshow goodness-of-fit test, with all reported P values exceeding 0.05. More informative calibration measures were uncommon: four models reported Brier scores ranging from 0.123 to 0.188,26,29,30,33 and only one reported a calibration slope approximating 1.28 Fifteen models conducted decision curve analysis and reported favorable net benefit.12–16,18,21,23,25–31 One machine-learning model additionally incorporated SHAP-based interpretation and was implemented as a web-based calculator.33

Risk of Bias and Applicability

Risk of bias and applicability were assessed using the Prediction Model Risk of Bias Assessment Tool (PROBAST). Model-level judgements are presented in Figure 2.

Figure 2.

A table and bar charts assess risk of bias and applicability in studies using PROBAST. The table lists studies by author and year, assessing risk of bias and applicability using PROBAST. Columns include ′Risk of bias′ and ′Applicability′ for participants, predictors, outcome and analysis, with overall judgments. Symbols indicate low, unclear and high risk. Bar chart B shows percentages for risk of bias in participants, predictors, outcome, analysis and overall, with low, unclear and high risk. Bar chart C displays applicability percentages for participants, predictors, outcome and overall, with low and unclear risk. Judgments are categorized as low risk, unclear risk and high risk.

PROBAST risk-of-bias and applicability assessment.

Notes: (A) presents model-level judgements for risk of bias and applicability. “+” indicates low risk of bias or low applicability concern, “-” indicates high risk of bias or high applicability concern, and “?” indicates unclear risk of bias or applicability concern; the same symbols are used for the overall judgements. (B) summarizes domain-level risk-of-bias judgements, and (C) summarizes domain-level applicability concerns.

None of the 22 models was judged to have an overall low risk of bias. In the Participants domain, 19 models were rated as low risk, two as high risk, and one as unclear. The Predictors domain was rated as unclear for all 22 models. In the Outcome domain, 11 models were rated as high risk and 11 as unclear. The Analysis domain was the principal source of bias, with 21 models rated as high risk and only one as low risk.

The risk-of-bias profile was driven primarily by analytical and outcome-related concerns rather than by participant selection. Recurrent analytical limitations included data-driven predictor selection, reliance on random split-sample evaluation, limited use of robust internal validation and optimism correction, and incomplete reporting or handling of missing data. The uniformly unclear Predictors ratings reflected insufficient information to support a low-risk judgement across the included reports. Study-specific concerns included internally inconsistent reporting of the age effect and model equation in one stroke nomogram,34 and a small validation sample, inconsistent performance reporting, absence of calibration assessment, and conceptual overlap between the Oral Health Assessment Tool predictor and the OFI-8 outcome in another study.32

Applicability concerns were generally low or unclear, and no model was judged to have high overall applicability concern. All 22 models had low applicability concerns in the Predictors and Outcome domains. In the Participants domain, eight models were rated as having low concern and 14 as unclear concern; consequently, overall applicability was judged as low concern for eight models and unclear for 14.

Sensitivity Analyses

Sensitivity analyses did not materially alter the main findings. When the synthesis was restricted to community-dwelling populations, seven independent studies, reported in eight publications, contributed seven corresponding models. All seven models had an overall high risk of bias and underwent some form of internal validation. Two also underwent external validation: one temporal17 and one geographical.13 Reported AUCs ranged from 0.747 to 0.980.

After excluding the model that reported apparent performance only,34 21 models remained. Twenty had undergone internal validation, and six had undergone external validation, comprising four temporal12,16,17,21 and two geographical validations.13,26 One model had temporal external validation without a separately reported internal-validation procedure.16 The overall AUC range remained 0.713–0.985. These findings were consistent with the primary synthesis, indicating that the main conclusions were not materially influenced by study setting or by inclusion of the model without validation.

Discussion

This systematic review identified 22 diagnostic prediction models for prevalent oral frailty from 22 independent studies reported in 23 publications. Although reported discrimination was generally moderate to high (AUC 0.713–0.985), confidence in these estimates is limited by a geographically concentrated and predominantly cross-sectional evidence base, uniformly high overall risk of bias, incomplete assessment of calibration, and limited external validation, especially across geographical settings. All models were designed to identify oral frailty already present at the time of assessment rather than to predict its future onset. Current models should therefore be regarded as preliminary diagnostic screening aids rather than established tools for prospective risk prediction or routine clinical decision-making.

These findings broadly accord with the recent scoping review, which similarly identified a predominance of cross-sectional model development, apparently favorable discrimination, limited external validation, and substantial risk of bias.10 The present systematic review incorporates a more recent and expanded evidence base and extends the previous synthesis by appraising the diagnostic versus prognostic purpose of the identified models. It also integrates evidence on model development, validation strategies, discrimination, calibration, clinical utility, and risk of bias to assess readiness for clinical use. All identified models estimate prevalent oral frailty rather than future incident oral frailty. Their reported performance therefore primarily reflects diagnostic classification rather than prognostic prediction. This distinction is clinically and methodologically important because favorable diagnostic discrimination should not be interpreted as evidence of prognostic performance.

This distinction warrants emphasis because the term “prediction model” does not necessarily imply prognostic prediction. Across the included studies, the models estimated the probability of oral frailty present at the time of assessment rather than the risk of its future occurrence. This remained true for all six externally validated models,12,13,16,17,21,26 including the retrospective cohort study;21 none evaluated incident oral frailty over a prespecified prediction horizon. Consequently, the reported discrimination reflects diagnostic classification of prevalent oral frailty and should not be extrapolated to forecasting its onset or progression. Distinguishing diagnostic from prognostic purposes is essential because prognostic modelling requires temporally antecedent predictors, a clearly defined prediction horizon, and evaluation of future outcomes, whereas diagnostic models are intended to identify an existing condition.

Despite heterogeneity in candidate predictors and modelling strategies, several predictor domains recurred across studies. Age was the most frequently retained predictor, while nutritional vulnerability or appetite loss, physical frailty or sarcopenia, multimorbidity or chronic disease burden, and tobacco exposure were also commonly represented. These patterns suggest that prevalent oral frailty is identified in conjunction with broader nutritional, systemic, and functional vulnerability rather than as an isolated oral condition. Because the evidence was predominantly cross-sectional, these variables should be interpreted as contemporaneous correlates or diagnostic markers rather than causal or temporal predictors of future oral frailty.

The recurrence of similar predictor domains across studies should not be interpreted as evidence of reproducible prediction across populations or settings. Thirteen studies were conducted in hospital-based populations, many involving selected disease-specific cohorts, whereas others recruited community-dwelling or rural older adults with different case mixes and baseline prevalence. Such heterogeneity may alter predictor effects, model intercepts, calibration, and clinically relevant decision thresholds, even when similar predictors are retained. In addition, the concentration of 21 of the 22 studies in China substantially limits confidence in geographical, cultural, and healthcare-system transportability. Restricting the synthesis to community-dwelling populations did not materially improve the overall risk-of-bias or validation profile, suggesting that these limitations reflect broader weaknesses in model development and evaluation rather than hospital-based sampling alone.

A further methodological concern is the conceptual overlap between some predictors and the OFI-8 outcome. Variables such as chewing difficulty, swallowing dysfunction, xerostomia, denture-related factors, oral health behaviors, and reduced social participation are closely related to, or partially represented within, the oral-frailty construct itself. The inclusion of the Oral Health Assessment Tool score in one model is particularly illustrative because several of its domains overlap with components assessed by the OFI-8.32 Such predictor–outcome overlap may produce circular diagnostic models in which predictors partly reproduce information contained in the outcome, potentially inflating apparent discrimination while providing limited etiological or anticipatory information. These variables may still be appropriate when the explicit objective is to develop a simplified case-finding or screening tool; however, that purpose should be prespecified and clearly distinguished from prediction based on information that is conceptually independent of the outcome and available at the intended moment of use.

Logistic regression and nomogram-based presentation remained the predominant modelling approaches, whereas only two studies compared multiple machine-learning algorithms.32,33 The available evidence does not indicate a consistent performance advantage of more complex algorithms over conventional regression: model rankings varied according to the performance metric used, and validation performance was less favorable than training performance for some candidate algorithms. These findings underscore that greater algorithmic complexity does not necessarily translate into better generalizability, particularly when model development is constrained by limited samples, data-driven predictor selection, or inadequate validation. Although SHAP-based interpretation and web-based implementation may enhance model interpretability and accessibility, these features should be regarded as adjuncts rather than substitutes for rigorous model development, comprehensive calibration assessment, and independent external validation.

Risk-of-bias and validation findings further limit confidence in the reported model performance. The revised PROBAST assessment indicates that analytical limitations were the dominant source of bias: 21 of 22 models were rated as high risk in the Analysis domain, while the Predictors domain was unclear for all models and the Outcome domain was high risk in half. Recurrent analytical concerns included data-driven predictor selection, incomplete handling of missing data, reliance on random split-sample validation, and insufficient control of optimism. External validation was reported for six models, but four were temporal validations12,16,17,21 and only two evaluated performance across geographical settings.13,26 Thus, evidence for transportability across institutions or communities remains particularly limited. Calibration assessment was also substantially less developed than discrimination assessment: the Hosmer-Lemeshow test predominated, whereas Brier scores were reported for only four models and a calibration slope for only one. A non-significant Hosmer-Lemeshow test alone should not be interpreted as evidence of adequate calibration in the absence of more informative measures such as calibration plots, calibration-in-the-large, and calibration slopes.35 Accordingly, favourable AUCs, decision-curve results, or explainability analyses should be interpreted cautiously when model development and validation remain methodologically limited. These patterns are consistent with prediction-model reviews in other geriatric domains36,37 and support further validation and model refinement before clinical implementation.

Taken together, the current models should be regarded as candidate diagnostic screening aids rather than established tools for prospective risk prediction or routine clinical decision-making. Future research should prioritize independent external validation and comprehensive calibration of existing diagnostic models. Genuine prognostic models should be developed using prospective longitudinal designs, temporally appropriate predictors, and clearly prespecified prediction horizons. Candidate models should also be evaluated against simpler standardized approaches, including the OFI-8 itself, to determine whether multivariable modelling provides meaningful incremental clinical value.

Strengths and Limitations

This review has several strengths. It was conducted according to TRIPOD-SRMA standards, systematically evaluated risk of bias and applicability using PROBAST, used a prespecified dual-reviewer process for study selection and data extraction to minimize errors, and included a bilingual search across English- and Chinese-language databases. Building on the recent scoping review, this review incorporates a more recent and expanded evidence base and provides a focused critical appraisal of the diagnostic versus prognostic purpose of the identified models, together with their development methods, validation strategies, discrimination, calibration, clinical utility, and risk of bias.

Several limitations should be acknowledged. First, substantial heterogeneity in study populations, oral-frailty definitions, predictor measurement and coding, modelling strategies, and performance reporting precluded quantitative pooling and direct comparison of model performance. Second, incomplete or internally inconsistent reporting in some primary studies limited the precision of data extraction and introduced uncertainty into risk-of-bias judgements. Finally, the search was restricted to English- and Chinese-language sources and did not specifically cover Japanese-language literature. Given the Japanese origins of the oral-frailty concept and the OFI-8, some relevant studies may have been missed, and the apparent geographical concentration of the evidence should therefore be interpreted with caution.

Conclusion

Current evidence on oral-frailty prediction comprises 22 diagnostic models for prevalent oral frailty, developed predominantly in cross-sectional Chinese populations. Although reported discrimination was generally moderate to high, confidence in model performance is limited by high overall risk of bias, incomplete calibration assessment, limited external validation, especially across geographical settings, and conceptual overlap between some predictors and the oral-frailty outcome. Accordingly, no existing model can yet be recommended for routine clinical decision-making or broad application across settings.

Future research should prioritize rigorous external validation and, where necessary, recalibration or updating of existing diagnostic models. Genuine prognostic models should be developed in prospective longitudinal cohorts with clearly defined prediction horizons and temporally appropriate predictors. Until robust evidence of transportability and clinical utility is available, current models should be regarded as candidate diagnostic screening aids that complement, rather than replace, standardized oral-frailty assessment and clinical judgement.

Acknowledgments

We sincerely thank Nantong University, Yangjie Cao, Yuying Wu, and all team members for their assistance with this research.

Funding Statement

This work was funded by the Nantong Natural Science Foundation (MS2024059); Jiangsu Province University Philosophy and Social Science Research Project (2023SJYB1683); Provincial College Students’ Innovation and Entrepreneurship Training Program Funding Project (S202510304174).

Author Contributions

All authors made a significant contribution to the work reported, whether that is in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; have agreed on the journal to which the article has been submitted; and agree to be accountable for all aspects of the work.

Disclosure

The authors report no conflicts of interest in this work.

References

  • 1.Gobbens RJ, Van Assen MA, Luijkx KG, Wijnen-Sponselee MT, Schols JM. Determinants of frailty. J Am Med Dir Assoc. 2010;11(5):356–15. doi: 10.1016/j.jamda.2009.11.008 [DOI] [PubMed] [Google Scholar]
  • 2.Fried LP, Tangen CM, Walston J, et al. Frailty in older adults: evidence for a phenotype. J Gerontol A Biol Sci Med Sci. 2001;56(3):M146–156. doi: 10.1093/gerona/56.3.M146 [DOI] [PubMed] [Google Scholar]
  • 3.Tanaka T, Takahashi K, Hirano H, et al. Oral frailty as a risk factor for physical frailty and mortality in community-dwelling elderly. J Gerontol A Biol Sci Med Sci. 2018;73(12):1661–1667. doi: 10.1093/gerona/glx225 [DOI] [PubMed] [Google Scholar]
  • 4.Takahashi T, Hatta K, Ikebe K. Risk factors of cognitive impairment: impact of decline in oral function. Jpn Dent Sci Rev. 2023;59:203–208. doi: 10.1016/j.jdsr.2023.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.de Sire A, Ferrillo M, Lippi L, et al. Sarcopenic dysphagia, malnutrition, and oral frailty in elderly: a comprehensive review. Nutrients. 2022;14(5):982. doi: 10.3390/nu14050982 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Iwasaki M, Yoshihara A, Sato M, et al. Dentition status and frailty in community-dwelling older adults: a 5-year prospective cohort study. Geriatr Gerontol Int. 2018;18(2):256–262. doi: 10.1111/ggi.13170 [DOI] [PubMed] [Google Scholar]
  • 7.Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. 2015;162(1):55–63. doi: 10.7326/M14-0697 [DOI] [PubMed] [Google Scholar]
  • 8.Wynants L, Van Calster B, Collins GS, et al. Prediction models for diagnosis and prognosis of COVID-19: systematic review and critical appraisal. BMJ. 2020;369:m1328. doi: 10.1136/bmj.m1328 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Moons KGM, Wolff RF, Riley RD, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: explanation and elaboration. Ann Intern Med. 2019;170(1):W1–W33. doi: 10.7326/M18-1377 [DOI] [PubMed] [Google Scholar]
  • 10.Chen Y, Tang H, Li B, Yuan L. Risk prediction models for oral frailty in older adults: a scoping review. Front Med Lausanne. 2026;13:1868632. doi: 10.3389/fmed.2026.1868632 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Snell KIE, Levis B, Damen JAA, et al. Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: checklist for systematic reviews and meta-analyses (TRIPOD-SRMA). BMJ. 2023;381:e073538. doi: 10.1136/bmj-2022-073538 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Xiao W, Gu D, Zhang M, et al. Development and validation of a nomogram for predicting oral frailty risk in elderly patients with ischaemic stroke. J Clin Nurs. 2026;35(1):346–362. doi: 10.1111/jocn.17855 [DOI] [PubMed] [Google Scholar]
  • 13.Wu Q, Cheng F, Cai W, Jiang Y, Shi Q. Construction and validation of a risk prediction model for oral frailty among community-dwelling older adults. Chin Nurs Res. 2025;39(18):3041–3047. doi: 10.12102/j.issn.1009-6493.2025.18.003 [DOI] [Google Scholar]
  • 14.Wang M, Yang W, Liao T, et al. Construction and validation of a risk prediction model for oral frailty in the elderly community population. Chin J Nurs. 2025;60(3):274–280. doi: 10.3761/j.issn.0254-1769.2025.03.003 [DOI] [Google Scholar]
  • 15.Qiao X, Gao S, Zhao H, Lu Y, Zhang H. A risk prediction model for oral frailty in elderly patients with COPD was constructed based on the health ecology model. BMC Oral Health. 2025;25(1):1897. doi: 10.1186/s12903-025-07186-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ma RR, Fan XL, Tao XB, Zhang W, Li ZB. Development and validation of risk-predicting model for oral frailty in older adult patients with stroke. BMC Oral Health. 2025;25(1):263. doi: 10.1186/s12903-025-05428-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lv KL, Yu P, Xue Y, Su JH, Ren YJ, Tang J. Construction and validation of an oral frailty risk prediction model for community-dwelling older adults. BMC Geriatr. 2025;25(1):808. doi: 10.1186/s12877-025-06393-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Lu CQ, Lu Q, Jin X, Song J. Construction and validation of a risk prediction model for oral frailty in elderly cancer patients. Chin J Geriatr Dent. 2025;23(3):210–218. doi: 10.19749/j.cn.cjgd.1672-2973.2025.03.009 [DOI] [Google Scholar]
  • 19.Jiang W, Liu H, Tao X, et al. Research on the current status and risk prediction model of oral frailty among the elderly in Anhui Province. J Pract Stomatol. 2025;41(2):261–266. doi: 10.3969/j.issn.1001-3733.2025.02.019 [DOI] [Google Scholar]
  • 20.Zou J, Shao Y, Liu J, et al. Construction and validation of a risk prediction model for oral frailty in rural elderly populations. J Bengbu Med Univ. 2025;50(8):1028–1034. doi: 10.13898/j.cnki.issn.2097-5252.2025.08.002 [DOI] [Google Scholar]
  • 21.Guo R, Xue D. Construction and verification of line graph model for oral frailty in elderly with Parkinson’s disease. Chin J Prev Contr Chron Dis. 2025;33(5):363–368. doi: 10.16386/j.cjpccd.issn.1004-6194.20240816.0614 [DOI] [Google Scholar]
  • 22.Zhong L, Zhang H, Xu J, Lu Y, Xiang X, Wang H. Construction of nomogram prediction model for the risk of oral frailty in elderly patients with type 2 diabetes mellitus. J Clin Med Prac. 2024;28(16):98–103. doi: 10.7619/jcmp.20240638 [DOI] [Google Scholar]
  • 23.Li Z, Pan Y, Zhou H, Zhou H, Wei Z, Li Y. Construction and validation of a risk prediction model for oral frailty in elderly inpatients. J Mudanjiang Med Univ. 2024;45(6):39–44,79. doi: 10.13799/j.cnki.mdjyxyxb.2024.06.015 [DOI] [Google Scholar]
  • 24.Yamamoto T, Tanaka T, Hirano H, Mochida Y, Iijima K. Model to predict oral frailty based on a questionnaire: a cross-sectional study. Int J Environ Res Public Health. 2022;19(20):13244. doi: 10.3390/ijerph192013244 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Feng N, Tian Y, Zhang S, et al. Construction and validation of a risk prediction model for oral frailty in rural hypertensive patients. Front Public Health. 2026;14:1687651. doi: 10.3389/fpubh.2026.1687651 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Lv J, Li J, Wang Y, et al. Construction and validation of a risk prediction model for oral frailty in elderly patients with esophageal cancer. Front Oncol. 2026;15:1736063. doi: 10.3389/fonc.2025.1736063 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yang W, Cai K, Zhang C, et al. Development and validation of a risk prediction model for oral frailty among Chinese older adults with type 2 diabetes mellitus. BMC Geriatr. 2026;26(1):336. doi: 10.1186/s12877-026-07039-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Liu H, Zhang M, Mei G, et al. A nomogram for predicting oral frailty in older adults: a small-sample cross-sectional study in Anhui Province, China. Front Public Health. 2026;14:1698294. doi: 10.3389/fpubh.2026.1698294 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Liu H, Hu X, Liu Q, et al. Development and internal validation of a risk prediction model for oral frailty in hospitalized older adults with chronic diseases. Front Public Health. 2026;14:1717485. doi: 10.3389/fpubh.2026.1717485 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Song X, Chai R, Ye J, Chen X, Lin Y, Xu C. Construction and validation of a risk prediction model for oral frailty in elderly patients with chronic heart failure. Front Med Lausanne. 2026;13:1775356. doi: 10.3389/fmed.2026.1775356 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Yang W, Pan Q, Lin Q, et al. Development and validation of a nomogram prediction model for oral frailty in community-dwelling older adults. Front Public Health. 2026;14:1793603. doi: 10.3389/fpubh.2026.1793603 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Cheng J, Wang T, Du Y, Sun M, Wu J. Construction and validation of an oral frailty risk prediction model for stroke patients based on multiple machine learning methods. Anhui Med J. 2026;47(3):355–362. doi: 10.3969/j.issn.1000-0399.2026.03.015 [DOI] [Google Scholar]
  • 33.Liu Q, Guo L, Liu H, et al. Construction of a risk prediction model for oral frailty in community-dwelling older adults using machine learning and SHAP analysis. J Nurs Sci. 2026;41(7):107–112,123. doi: 10.3870/j.issn.1001-4152.2026.07.107 [DOI] [Google Scholar]
  • 34.Liu W, Wang X, Yang S, Zhang X. Construction of a risk prediction model for oral frailty in older stroke patients. Chin J Gen Pract. 2026;24(4):593–595,600. doi: 10.16766/j.cnki.issn.1674-4152.004447 [DOI] [Google Scholar]
  • 35.Kramer AA, Zimmerman JE. Assessing the calibration of mortality benchmarks in critical care: the Hosmer-Lemeshow test revisited. Crit Care Med. 2007;35(9):2052–2056. doi: 10.1097/01.CCM.0000275267.64078.B0 [DOI] [PubMed] [Google Scholar]
  • 36.Lin T, Liao H, Su L, et al. The risk prediction models for sarcopenia in older adults: a systematic review and critical appraisal. Front Public Health. 2026;14:1751954. doi: 10.3389/fpubh.2026.1751954 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Ou Y, Jiang D, Li P, et al. Prediction of frailty in community older adults based on machine learning: a systematic review and meta-analysis. Front Public Health. 2026;13:1667792. doi: 10.3389/fpubh.2025.1667792 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Patient preference and adherence are provided here courtesy of Dove Press

RESOURCES