Skip to main content
Frontiers in Medicine logoLink to Frontiers in Medicine
. 2026 Jul 15;13:1855638. doi: 10.3389/fmed.2026.1855638

Systematic review and meta-analysis of machine learning-based prediction models for readmission risk after total hip and knee arthroplasty

Jiabin Feng 1,†, Min Ma 2,†, Changliang Ou 1, Kaiwei Zhang 1,*
PMCID: PMC13415355  PMID: 42528854

Abstract

Background

Postoperative readmission is a critical metric after total hip and knee arthroplasty (THA/TKA). While machine learning (ML) models for predicting readmission risk are proliferating, the stability of their performance and the robustness of their methodology remain contentious. This study aimed to systematically review and quantitatively synthesize the evidence base for ML-based readmission prediction after THA/TKA.

Methods

A systematic search was conducted across PubMed, Embase, Cochrane Library, and Web of Science from inception to December 31, 2025. Studies developing or validating ML models for THA/TKA readmission were included. Model performance (C-statistic/AUC) was extracted. Study quality was assessed using the PROBAST + AI tool. A multivariate random-effects meta-analysis was performed to pool C-statistics and quantify heterogeneity, with subgroup analyses stratified by study design, surgery type, prediction timeframe, and algorithmic class.

Results

Fifteen studies (57 distinct models) were included. The pooled C-statistic was 0.76 (95% CI: 0.71–0.81). However, extreme heterogeneity (I2 = 99.9%) rendered this point estimate of limited clinical utility; the 95% prediction interval (0.38–0.94) highlighted profound outcome unpredictability. Subgroup analyses revealed significant moderators: single-center models showed optimistically higher performance (0.86) compared to multicenter models (0.65), and THA-specific models yielded higher estimates than TKA-specific models, although these findings were derived from few studies and should be interpreted as exploratory. Advanced ML algorithms did not demonstrate consistent superiority over traditional logistic regression. Crucially, the PROBAST + AI assessment identified a high risk of bias in the majority of studies, primarily due to analytical shortcomings and a universal lack of model recalibration.

Conclusion

The current body of evidence for ML-based readmission prediction after THA/TKA is characterized by extreme heterogeneity and high methodological bias, severely constraining clinical utility. The inability to pool calibration metrics represents a critical evidence gap. Future research must prioritize multi-institutional validation, stringent adherence to reporting standards (e.g., TRIPOD-AI), and the mandatory transparent reporting of both discrimination and calibration metrics to realize any potential clinical benefit.

Systematic review registration

https://www.crd.york.ac.uk/PROSPERO/view/1305608, identifier CRD420261305608.

Keywords: artificial intelligence, machine learning, Patient Readmission, risk prediction, systematic review, total hip arthroplasty, total knee arthroplasty

1. Introduction

When conservative treatments fail to alleviate severe joint dysfunction caused by conditions such as osteoarthritis or rheumatoid arthritis, joint arthroplasty represents a definitive and crucial therapeutic solution. Among these procedures, total hip arthroplasty (THA) and total knee arthroplasty (TKA) are the most widely performed and representative interventions (1). These surgeries involve implanting customized prostheses, typically fabricated from materials such as metals, high-molecular-weight polyethylene, and ceramics, to replace the diseased joint, with the primary goals of relieving pain, correcting deformity, and restoring function. They play an indispensable role in improving patients’ quality of life (2).

In recent years, machine learning (ML) has gained significant prominence in medical prediction (3, 4). By analyzing structured data, including patient demographics, biomarkers, medical history, and anesthetic records, and employing algorithms to explore relationships between features and outcomes, ML models can identify features most strongly associated with the target event, thereby enhancing prediction accuracy. Postoperative readmission following THA/TKA, a key metric for evaluating surgical efficacy and patient prognosis, has consequently attracted increasing research attention (5–19). Although numerous studies have applied machine learning (ML) for THA/TKA readmission prediction, substantial heterogeneity persists in reported performance and predictors. Existing models predominantly incorporate demographic, comorbidity, and procedural variables, yet largely overlook neurocognitive factors that are increasingly recognized as key drivers of postoperative complications and unplanned readmission. While broad reviews of AI in orthopedics exist (3, 11, 20), they lack the granularity required for arthroplasty-specific risk stratification and fail to address the methodological inconsistencies driving performance variability. Critically, no prior synthesis has quantified the extreme between-study heterogeneity or systematically compared the generalizability of single-center models against multicenter cohorts. Furthermore, the putative advantage of complex algorithms over traditional logistic regression remains unverified in this domain. Therefore, this systematic review and meta-analysis aims to bridge this gap by providing a granular evaluation of ML model performance, with a specific focus on quantifying heterogeneity, comparing algorithmic efficacy, and identifying consistent predictors of readmission. Our findings seek to guide the development of robust, clinically applicable models that avoid common pitfalls such as optimism bias and calibration errors induced by improper class imbalance handling. Therefore, this study aims to conduct a systematic review and meta-analysis of the existing literature on ML-based prediction models for readmission after THA/TKA, to summarize the current evidence and provide a reference for future model development and clinical application.

2. Materials and methods

This systematic review and meta-analysis was conducted and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement (21). Its protocol was prospectively registered with the International Prospective Register of Systematic Reviews (PROSPERO; registration number: CRD420261305608). The study design adhered to the PICOTS framework. The target population comprised patients undergoing total hip or knee arthroplasty (THA/TKA). The predictive intervention of interest was the application of a machine learning (ML)-based model to estimate the risk of postoperative readmission. The primary outcome of interest was the discriminative accuracy of these models for predicting readmission, quantified by the area under the receiver operating characteristic curve (AUC) (22), measured from discharge to a specified postoperative interval. As the focus was on evaluating the models themselves, no direct comparator was required. The intended application setting was the hospital environment.

2.1. Literature search

A comprehensive literature search was conducted in PubMed, the Cochrane Library, Embase, and Web of Science from their inception to December 31, 2025. The search strategy combined Medical Subject Headings (MeSH) and free-text terms covering three key concepts: arthroplasty (“Total Joint Arthroplasty,” “Total Knee Arthroplasty,” “Total Hip Arthroplasty,” “Joint Replacement”), readmission (“Readmission,” “Rehospitalization,” “Re-admission”), and artificial intelligence-based predictive modeling (“Machine Learning,” “Deep Learning,” “Artificial Intelligence,” “Predictive Model,” “Risk Prediction,” “Prediction Model,” “Risk Model,” “Algorithm”). The detailed search strings for each database are provided in Supplementary material.

2.2. Inclusion and exclusion criteria

Studies were included according to the following criteria: (1) The study population consisted of patients who underwent THA or TKA. (2) The primary or secondary outcome included postoperative readmission. (3) The study developed or validated a multivariable prediction model for readmission. For the purposes of this review, “machine learning (ML)-based models” were operationally defined broadly to include the full spectrum of supervised learning algorithms. This encompassed both traditional statistical approaches (e.g., Logistic Regression, LASSO) and advanced ML algorithms (e.g., Random Forest, Support Vector Machines, Artificial Neural Networks, Gradient Boosting Machines). Studies were excluded only if they failed to construct a multivariable prediction model (e.g., only reported univariate associations or descriptive statistics). (4) The publication was an original research study in a peer-reviewed journal. (5) The full text was available in English.

Studies were excluded if they were: (1) conference abstracts, reviews, commentaries, case reports, or dissertations; (2) not accessible in full text or were duplicate publications; (3) studies that failed to develop or validate a multivariable prediction model, such as those involving only univariate risk factor analysis or basic statistical comparisons; (4) studies reporting composite endpoints from which data specific to readmission could not be separately extracted; or (5) studies that did not report key model performance metrics (such as AUC) or the follow-up duration.

2.3. Study selection and data extraction

Two reviewers independently screened the titles, abstracts, and full texts of retrieved records against the pre-defined eligibility criteria. Disagreements were resolved through discussion or, if necessary, by consulting a third reviewer. For each included study, data were extracted independently by two reviewers using a standardized form based on the CHARMS (Checklist for critical Appraisal and data extraction for systematic reviews of prediction modeling studies) checklist (5), which provides a structured framework for extraction. The extracted information included: first author, publication year, study population, predictors, method for handling class imbalance, type of surgery (THA, TKA, or both), postoperative observation period (days), total sample size, number of readmission events, models developed, and the reported performance metrics (e.g., AUC).

2.4. Quality assessment (risk of bias)

The risk of bias and concerns regarding the applicability of the included prediction modeling studies were independently assessed by two reviewers using the PROBAST + AI tool (Prediction model Risk Of Bias ASsessment Tool for studies developing, validating, or updating a multivariable prediction model, with an extension for Artificial Intelligence) (23). This tool, an update and extension of the original PROBAST, evaluates four domains (participants, predictors, outcome, and analysis) for both model development and validation. Prior to formal assessment, the two reviewers, both trained in evidence-based medicine, conducted a calibration exercise on the PROBAST + AI tool items. They then performed a pilot assessment on four randomly selected studies to ensure a consistent understanding and application of the criteria. The inter-rater reliability, measured by Cohen’s kappa, ranged from 0.464 to 0.868 for the risk-of-bias domains and was 1.00 for all applicability domains (see Supplementary Table 1) (24). Any discrepancies in the final assessments were resolved through team consensus.

2.5. Statistical analysis

A meta-analysis of the C-statistic requires an estimate of its variance. Among the 15 included studies, 7 provided confidence intervals. For the remaining 8 studies, 95% confidence intervals were calculated using the standard error of the C-statistic and assuming a normal approximation. The meta-analysis was performed using the R package “metamisc,” designed for the meta-analysis of prediction model performance measures (25). Crucially, to account for the non-independence of multiple models reported within the same study, we employed a multivariate random-effects model (also known as a multilevel model). This model structure explicitly partitions the residual variance into between-study and within-study components, adjusting the weights accordingly to prevent the artificial inflation of precision. A random-effects model was used to pool the C-statistics, with weighting based on the within-study error variance. The 95% confidence interval for the summary C-statistic was estimated using the Hartung-Knapp-Sidik-Jonkman method (26). To verify the robustness of the findings, subgroup analyses were planned based on surgery type (THA only, TKA only, or combined), data source (single-center vs. multicenter), method for handling class imbalance, length of postoperative follow-up (days), and model category. Potential small-study bias was assessed via a funnel plot, visually inspecting the scatter of the reported C-statistic against its standard error for asymmetry. The statistical significance of the asymmetry was tested using the weighted Egger’s regression test (27).

3. Results

3.1. Literature search and selection

The initial database search yielded 439 records. Following the removal of duplicates and the screening of titles and abstracts, the full texts of 54 potentially eligible studies were retrieved and assessed. Ultimately, 15 studies satisfied all inclusion criteria and were incorporated into this systematic review and meta-analysis. The detailed study selection process is illustrated in the PRISMA flow diagram (Figure 1).

FIGURE 1.

PRISMA flow diagram showing systematic review study selection. Out of 439 identified records, 116 duplicates were removed. After screening 323 records, 54 were assessed for eligibility, and 15 studies were included in the review.

PRISMA flow diagram.

3.2. Characteristics of included studies

The 15 included studies all focused on developing machine learning (ML) models to predict readmission risk following total hip arthroplasty (THA) or total knee arthroplasty (TKA). Key characteristics are summarized in Tables 1, 2 (5–19).

TABLE 1.

Study characteristics.

Study Study population Predictors Class imbalance handling
Mohammadi et al. (5) Hip/knee arthroplasty EHR patients, Partners Healthcare, 2006–2016. Comprehensive clinical and demographic variables (e.g., demographics, medical history, vital signs, laboratory results, comorbidities, medications, procedures, diagnoses, admission details). Undersampling
Slezak et al. (6) TJR patients, NSQIP database, 2011–2017. Demographics, comorbidities, functional status, and perioperative indicators (e.g., age, BMI, diabetes, ASA, modified frailty index). Stratified sampling only
Gould et al. (7) Primary TKA patients, single center, Australia, pre-March 2020. Multidimensional variables encompassing biopsychosocial factors, comorbidities, prior healthcare utilization, and perioperative events. Not explicitly mentioned
Klemt et al. (8) Primary TKA, single center, 2016–2019. Demographics, comorbidities, surgical specifics, and behavioral factors (e.g., insurance, anesthesia type, implant fixation, smoking/substance use). Stratified sampling only
Wu et al. (9) THA/hemiarthroplasty patients, single hospital, Taiwan, Sep 2016-Dec 2018. Comorbidities, surgical factors, and laboratory values with a focus on coagulation profiles (e.g., CKD, CAD, INR, PT, aPTT) and history of falls. SMOTE
Kunze et al. (10) Medicare THA/TKA patients, claims data, Oct 2016–Sep 2018. Encompassed patient demographics, hospital/surgeon procedural volumes, prior medical history, and county-level social determinants of health. Downsampling
Shaikh et al. (11) Elective THA/TKA patients, SPARCS database, 2012–2016. Included demographics, socioeconomic factors, surgical context, and comorbidity burden quantified by the Elixhauser index. Random Undersampling
Park et al. (12) Primary THA/TKA, single academic safety-net hospital, 2016–2019. Covered variables across the entire care continuum, from demographics and preoperative status to hospitalization and post-discharge factors, including patient-reported outcomes. Not specified; oversampling; RUSBoost
Buddhiraju et al. (13) Primary TKA patients, ACS-NSQIP database, 2013–2020. Comprised core demographic and clinical variables: age, sex, BMI, comorbidities, preoperative labs, and hospitalization/surgical factors. Not explicitly mentioned
Khan et al. (14) Primary unilateral THA, US tertiary academic center, 2016–2020. Featured a multidimensional set including demographics, Area Deprivation Index (ADI), comorbidities (CCI), patient-reported outcomes (PROMs), and perioperative details. Not explicitly mentioned
Buddhiraju et al. (15) Primary TKA patients, ACS-NSQIP, 2013–2020. Consisted of preoperative demographic and clinical risk factors commonly used in surgical risk assessment (e.g., functional status, major organ system comorbidities). Not explicitly mentioned
Khan et al. (16) Primary unilateral TKA, US tertiary academic center, 2016-2020. Included demographics, socioeconomic factors (e.g., ADI), comorbidities (CCI), patient-reported outcomes (PROMs), and perioperative details. Not explicitly mentioned
Crespi et al. (17) Primary THA patients, MARCQI, single institution, 2012–2023. Comprised demographics, lifestyle factors, key comorbidities, and perioperative variables (e.g., LOS, discharge type, preoperative opioid use). Not explicitly mentioned
Khan et al. (18) Primary TKA patients, MARCQI center, 2016–2024. Included demographics, lifestyle factors, key comorbidities, and perioperative variables (e.g., ASA grade, LOS, discharge disposition). Not explicitly mentioned
Purbasari et al. (19) THA patients, MIMIC-IV database. Covered demographics, insurance type, a broad range of preoperative diagnoses, and implant-specific details (material and type). Random oversampling (applied to training set)

BMI, body mass index; CHF, congestive heart failure; COPD, chronic obstructive pulmonary disease; CKD, chronic kidney disease; ASA, American Society of Anesthesiologist Physical Status score; CCI, Charlson comorbidity index; CAD, Coronary artery disease; PROM, Patient-reported outcome measure phenotype; LOS, Length of stay; NARX, Opioid overdose risk score; DVT/PE, deep vein thrombosis or pulmonary embolism.

TABLE 2.

Study data.

Study Type Length (Day) N N. events Models Auc
Mohammadi et al. (5) THA 30 478 ∼55 ANN 0.846(0.823–0.871)
TKA 628 ∼73 ANN 0.822(0.809–0.862)
Slezak et al. (6) TKA/THA 30 ∼95523 ∼3486 RF 0.602(0.593–0.613)
LR 0.658(0.648–0.668)
XGBoost 0.673(0.663–0.683)
LightGBM 0.672(0.662–0.682)
Gould et al. (7) TKA 30 ∼921 ∼63 RF 0.692(0.621–0.764)
LR 0.589(0.506–0.673)
Klemt et al. (8) TKA 90 2005 ∼129 ANN 0.85(0.81–0.89)
SVM 0.82(0.78–0.86)
KNN 0.82(0.78–0.83)
LR 0.83(0.79–0.87)
Wu et al. (9) THA 30 296 8 LR 0.982
DT 0.933
RF 0.976
ANN 0.989
Kunze et al. (10) TKA 30 87,930 3,517 GAM 0.66
THA 47,878 1,915 GAM 0.65
Shaikh et al. (11) TKA/THA 30 6000 ∼176 AB 0.48
XGBoost 0.62(0.62–0.63)
LR 0.61(0.61–0.62)
RF 0.68(0.67–0.68)
SVM 0.63(0.63–0.64)
1 Layer NN 0.53(0.53–0.54)
5 Layer NN 0.53(0.53–0.54)
90 6000 ∼25 AB 0.72(0.72–0.73)
XGBoost 0.61(0.6–0.62)
LR 0.62(0.62–0.63)
RF 0.69(0.68–0.69)
SVM 0.64(0.63–0.65)
1 Layer NN 0.53(0.53–0.54)
5 Layer NN 0.53(0.53–0.54)
Park et al. (12) TKA/THA 90 ∼397 ∼23 LR 0.845(0.825–0.87)
LASSO 0.862(0.842–0.885)
SVM-polynomial 0.714(0.705–0.789)
SVM-radial 0.789 (0. 771–0.839)
RF 0.836(0.819–0.855)
RF w/oversampling 0.835(0.816–0.85)
RUS boosting 0.835(0.811–0.863)
Buddhiraju et al. (13) TKA 30 73157 2,297 ANN 0.78
RF 0.78
HGB 0.78
NEPLR 0.78
Khan et al. (14) THA 30 8893 310 LR 0.7(0.67–0.72)
90 502 LR 0.71(0.69–0.74)
Buddhiraju et al. (15) TKA 30 2835 88 ANN 0.72
Khan et al. (16) TKA 30 10521 444 LR 0.65(0.57–0.75)
90 704 LR 0.68(0.67–0.7)
Crespi et al. (17) THA 90 402 21 MPNN 0.71
Khan et al. (18) TKA 90 645 26 MPNN 0.704
Purbasari et al. (19) THA 365 346 71 XGB 0.996
SGB 0.994
RF 0.994
SVM 0.991
DT 0.878
MLP 0.95
LR 0.662

RF, random forest; LR, logistic regression; SVM, support vector machine; NN, neural network; KNN, k-nearest neighbors; DT, decision tree; GB, gradient boosting.

The studies were published between 2020 and 2025. Study populations primarily comprised patients undergoing primary THA and/or TKA. Data sources were heterogeneous, including single-center Electronic Health Records (EHRs), multi-center registries [e.g., the American College of Surgeons National Surgical Quality Improvement Program (ACS-NSQIP), Statewide Planning and Research Cooperative System (SPARCS)], and administrative claims databases (e.g., Medicare) (5, 6, 10, 11).

A wide array of ML algorithms was employed. Commonly used models included traditional methods like Logistic Regression (LR) (6–9, 11–14, 16, 19) and ensemble methods like Random Forest (RF) (6, 7, 9, 11–13, 19), as well as more complex algorithms such as Artificial Neural Networks (ANNs) (5, 8, 9, 13, 15), Support Vector Machines (SVMs) (8, 11, 12, 19), gradient boosting machines (e.g., XGBoost (6, 11, 19), LightGBM) (6), and novel architectures like Message Passing Neural Networks (MPNNs) (17, 18). Model performance, evaluated primarily by the Area Under the Receiver Operating Characteristic Curve (AUC), varied substantially, with reported values ranging from approximately 0.48 (11) to over 0.99 (19).

Various strategies were employed to address class imbalance, including undersampling (5, 11), oversampling (12, 19), the Synthetic Minority Over-sampling Technique (SMOTE) (9), and ensemble methods like RUSBoost (13). Several studies did not specify their approach to handling imbalance (7, 13–18).

Predictors spanned multiple domains: demographic factors, comorbidities (using indices like the Charlson Comorbidity Index or specific diagnoses), surgical details, laboratory values, and socioeconomic indicators (5–19). The outcome was unplanned readmission within a defined postoperative window, most commonly 30 days (5–7, 9–11, 13–16), followed by 90 days (8, 11, 12, 14, 16–18); one study used a 365-day window (19). Sample sizes varied from 296 (9) to over 95,000 (6), with corresponding readmission event counts from single digits (9) to several thousand (6, 10).

3.3. Meta-analysis, heterogeneity, and publication bias

The meta-analysis yielded a pooled summary estimate (logit AUC) of 0.76 (95% CI: 0.71–0.81) (Figure 2). However, given the extreme between-model heterogeneity (I2 = 99.9%, Tau2 = 0.6728), this summary estimate is of limited clinical interpretability as it represents a mathematical average rather than a reproducible performance benchmark. More germane to clinical translation, the 95% prediction interval (0.38–0.94) indicates that the true performance of a model developed in a new setting is highly uncertain, potentially ranging from poor to excellent. This wide interval underscores the profound impact of contextual and methodological moderators on model generalizability.

FIGURE 2.

Forest plot showing odds ratio estimates and 95 percent confidence intervals for various machine learning models used to predict joint arthroplasty outcomes, with sources, model types, and summary estimate of 0.76 (0.71 to 0.81).

Meta-analysis.

Visual inspection of the funnel plot (Figure 3) showed marked asymmetry. High-precision studies clustered near the summary estimate, while lower-precision studies exhibited a wider, more scattered distribution of effect sizes, with several falling below the pooled estimate. This pattern, alongside the extreme statistical heterogeneity, indicates substantial genuine differences between studies beyond sampling error, likely due to unaccounted moderating factors explored in subgroup analyses. While publication bias cannot be ruled out, the observed asymmetry is more strongly attributable to this high heterogeneity.

FIGURE 3.

Scatter plot with effect size (Logit AUC) on the x-axis and standard error on the y-axis, showing black data points, a solid vertical reference line at zero, a dashed regression line, and a blue confidence interval triangle.

Funnel plot.

3.4. Subgroup analyses

Study Design (Single-Center vs. Multicenter): Models from single-center studies (10 studies; k = 32) showed a pooled estimate of 0.86 (95% CI: 0.79–0.90), while those from multicenter studies (4 studies; k = 25) showed a significantly lower estimate of 0.65 (95% CI: 0.61–0.69) (Figure 4). Heterogeneity remained extreme within both subgroups (I2 = 99.0 and 99.8%, respectively).

FIGURE 4.

Forest plot displaying odds ratios (OR) with ninety-five percent confidence intervals for various predictive models in single-center and multicenter studies of knee and hip arthroplasty outcomes, grouped by study and model type, with summary estimates and prediction intervals included at the end of each group.

Subgroup analyses: study design.

Surgical procedure: THA-specific models (from 6 independent studies; k = 16) yielded a pooled estimate of 0.93 (95% CI: 0.84–0.97). TKA-specific models (7 studies; k = 16) showed 0.76 (95% CI: 0.71–0.79), while mixed-cohort models (3 studies; k = 25) yielded 0.68 (95% CI: 0.63–0.73) (Figure 5). Heterogeneity remained extreme (I2 > 98%).

FIGURE 5.

Forest plot comparing odds ratios for THA-specific (pooled estimate 0.93, 95% CI: 0.84 -0.97), TKA-specific (0.76, 0.71 -0.79), and mixed-cohort models (0.68, 0.63 -0.73).

Subgroup analyses: surgical procedure.

Prediction timeframe: Models for 30-day prediction (10 studies; k = 28) yielded a pooled estimate of 0.69 (95% CI: 0.64–0.74). Performance appeared higher for 90-day models (7 studies; k = 22; 0.75, 95% CI: 0.70–0.79) and 365-day models (1 study; k = 7; 0.97, 95% CI: 0.86–1.00) (Figure 6) (Exploratory; based on a single study). Extreme heterogeneity persisted across all timeframes (I2 from 96.3 to 99.8%). Crucially, given the reliance on a single study, the 365-day estimate represents a hypothesis-generating observation rather than a stable performance benchmark and should be interpreted with extreme caution.

FIGURE 6.

Forest plot showing predictive models for 30-day (OR 0.69), 90-day (OR 0.75), and 365-day outcomes (OR 0.97). Note: 365-day estimate is exploratory and based on a single study.

Subgroup analyses: prediction timeframe.

Model class (algorithm type): Support Vector Machines (4 studies; k = 6) and Tree-based/Ensemble Methods (7 studies; k = 21) had the highest pooled point estimates (0.81 and 0.82, respectively), while Traditional Statistical Models (11 studies; k = 16) had a lower estimate (0.71, 95% CI: 0.65–0.77) (Figure 7). Neural Networks (9 studies; k = 13) showed an estimate of 0.74 (95% CI: 0.62–0.84). Heterogeneity was extreme across all algorithm subgroups (I2 from 99.6 to 100.0%).

FIGURE 7.

Forest plot displaying estimates for Support Vector Machines (OR 0.81), Tree-based methods (OR 0.82), Traditional Statistical Models (OR 0.71), and Neural Networks (OR 0.74).

Subgroup analyses: model class.

Class imbalance handling method: Models from studies using oversampling (3 studies; k = 12) had the highest pooled estimate (OR: 0.96, 95% CI: 0.91–0.99), while those using undersampling (3 studies; k = 18) had the lowest (OR: 0.64, 95% CI: 0.59–0.69) (Figure 8). Models with no balancing (2 studies; k = 8) or unreported methods (8 studies; k = 18) showed intermediate performance (OR: 0.75 for both). High heterogeneity persisted within each subgroup.

FIGURE 8.

Forest plot graphic comparing odds ratios with 95 percent confidence intervals for various machine learning models predicting outcomes for knee and hip arthroplasty, grouped by model balancing approach: no specific balancing, unspecified, oversampling, undersampling, and hybrid methods, with summary estimates and prediction intervals for each group.

Subgroup analyses: class imbalance handling method.

3.5. Risk of bias and applicability

Assessment using the PROBAST + AI tool (Figure 9) indicated a high overall risk of bias for most studies, primarily in the Analysis domain and the Participants domain during model development. Common concerns included lack of sample size justification, inadequate reporting on handling missing data, absence of model recalibration when class imbalance techniques were used, and incomplete description of resampling procedures. The risk of bias was generally low for the Predictors and Outcome domains. Concerns regarding the applicability of the studies to the review question were low across all domains.

FIGURE 9.

Heatmap evaluating fourteen studies on model development and model evaluation across domains including quality, applicability, and risk of bias for participants, predictors, outcomes, and analyses; colors indicate high quality, unclear, and low quality according to the provided key.

Risk of bias and applicability.

4. Discussion

This systematic review and meta-analysis synthesizes evidence from 15 studies investigating ML-based prediction models for readmission after THA/TKA. While the pooled C-statistic was 0.76 (95% CI: 0.71–0.81), the extreme statistical heterogeneity (I2 = 99.9%) renders this point estimate of limited clinical utility. The clinically actionable insight is encapsulated by the 95% prediction interval (0.38–0.94), which highlights the profound unpredictability of model performance. Consequently, our subgroup analyses were not undertaken to derive precise performance benchmarks—which would be implausible given the limited number of independent studies (e.g., n = 1 for 365-day models, n = 6 for THA-specific models)—but rather to identify potential methodological and clinical moderators that drive this wide variation.

Our subgroup analyses provide crucial insights into the sources of this variability. First, study design emerged as a strong moderator. Despite the limited number of multicenter studies (n = 4), models derived from single-center data reported significantly higher—and likely optimistic—performance estimates (pooled AUC 0.86) compared to those developed from multicenter data (pooled AUC 0.65). This discrepancy aligns with the known challenge of overfitting to local data patterns and underscores the greater potential for generalization, albeit with more conservative performance metrics, of models trained on diverse, multi-institutional cohorts (11, 13). Second, the target surgical procedure significantly influenced the results. THA-specific models (6 studies) demonstrated notably higher discriminative ability than TKA-specific models (7 studies) or those built on mixed cohorts (3 studies). However, given the small number of contributing studies and overlapping confidence intervals, this difference should be interpreted as hypothesis-generating rather than definitive evidence of superior THA model performance. This may reflect inherent differences in the pathophysiology leading to readmission, variability in population homogeneity, or differences in the predictive utility of available variables for each procedure (17, 28), supporting the argument for procedure-specific model development. Third, while a trend toward higher AUCs was observed with longer prediction timeframes, this finding is severely limited by the fact that the 365-day estimate derives entirely from a single study. Although this seemingly counterintuitive trend may hypothetically reflect that longer-term readmissions are more strongly driven by stable patient factors (e.g., chronic comorbidities), this observation currently lacks the empirical robustness to support definitive conclusions regarding temporal model performance. In contrast, short-term readmissions may be more influenced by acute, peri-operative, or surgical factors that are harder to predict consistently (29). Emerging evidence highlights neurocognitive complications—including postoperative delirium and perioperative neurocognitive disorders—as prominent acute peri-operative drivers of early readmission after arthroplasty (30). Randomized trial protocols indicate that interventions such as repetitive transcranial magnetic stimulation to mitigate postoperative delirium (31) and opioid-free anesthetic regimens to reduce neurocognitive deficits in elderly hip fracture surgery (32) may alter readmission risk profiles. Incorporating neurocognitive risk factors or intervention exposure as candidate predictors in future ML models could therefore improve the capture of acute peri-operative contributors to short-term readmission, addressing a key gap in current model performance.

Regarding modeling techniques, this study observed that advanced ML algorithms (e.g., SVM, tree-based ensembles) did not demonstrate a consistent or decisive superiority over traditional statistical models like logistic regression. In fact, logistic regression models showed more consistent performance (lower Tau2) across studies, albeit at a modest average level. This finding echoes a growing consensus in predictive analytics that increased model complexity does not automatically translate to superior clinical utility, particularly with limited sample sizes or high-dimensional, noisy data (33, 34). Furthermore, the method for handling class imbalance was identified as a critical and frequently under-reported methodological factor. The use of oversampling techniques was associated with optimistically high performance estimates (9, 19). A concomitant and concerning finding was the near-universal lack of model recalibration following the application of such techniques. This oversight can lead to poorly calibrated probability estimates—where predicted risks do not match observed event rates—severely undermining the models’ practical utility for individualized clinical risk stratification (35).

The PROBAST + AI assessment confirmed a high risk of bias in the majority of included studies, with concerns concentrated primarily in the Analysis domain. Common shortcomings—including inadequate handling of missing data, lack of sample size justification, and poor reporting of resampling—represent critical threats to internal validity. Crucially, these pervasive methodological weaknesses are not merely background noise; they are a primary driver of the extreme heterogeneity (I2 = 99.9%) observed in our synthesis. While the predictors and outcomes were largely applicable, the prevailing biases significantly curb confidence in the reported performance estimates. Consequently, the current evidence base is too unreliable to support the clinical translation of these specific models, regardless of their theoretical potential.

Several limitations of this review should be acknowledged. First, while 15 included studies and 57 distinct ML models provided sufficient data points to explore sources of heterogeneity, the small number of independent publications limited the statistical power for definitive subgroup comparisons. Formal interaction tests were not feasible due to extreme heterogeneity. Crucially, several key subgroups were represented by very few studies (e.g., only 1 study contributed to the 365-day models, and only 3–6 studies defined procedure-specific cohorts). Therefore, all subgroup findings—including the apparently high AUC for THA or 365-day models—must be interpreted strictly as exploratory and hypothesis-generating. The observed trends offer insights into potential moderators but do not constitute conclusive evidence of performance differences. Nevertheless, the observed trends—such as the significant performance gap between single-center and multicenter models—offer critical insights into model generalizability. Second, the persistently high heterogeneity, even within our defined subgroups, suggests the influence of important unmeasured or unexamined moderators. These could include specific data quality and curation practices, granular details of feature engineering, or hyperparameter optimization strategies. Third, the observed asymmetry in the funnel plot raises the possibility of publication bias, where studies reporting lower model performance may be under-represented in the published literature. This could lead to an overestimation of the true average discriminative ability in the field. Fourth, our quantitative synthesis focused solely on the C-statistic (discrimination) because a meaningful pooled analysis of calibration metrics was precluded by inconsistent and incomplete reporting across studies. This represents a critical evidence gap in the field of orthopedic ML research. While discrimination (AUC) quantifies a model’s ability to rank patients, it fails to assess whether predicted risks match observed event rates. Without transparent reporting of calibration plots, intercept, slope, and Brier scores, even a model with high AUC may produce systematically biased probability estimates (e.g., predicting a 5% risk when the true risk is 20%), thereby undermining safe clinical decision-making. We strongly advocate that future studies adhere to the TRIPOD-AI reporting guideline to ensure both discrimination and calibration are rigorously evaluated and transparently reported.

5. Conclusion

In conclusion, while machine learning models for THA/TKA readmission suggest preliminary promise, this potential is currently severely constrained by prevalent methodological limitations and a high risk of bias. The pooled analysis indicates a moderate average discriminative ability; however, the performance of any individual model is highly contingent on specific contextual and methodological factors, including study design (single-center vs. multicenter), target surgical procedure, prediction timeframe, and approaches to handling class imbalance. There is no consistent evidence that complex ML algorithms confer a definitive advantage over traditional statistical models like logistic regression in this context. To advance the field, future research should prioritize the development and validation of models using robust, multi-institutional datasets, adhere to stringent methodological and reporting standards (e.g., TRIPOD-AI), ensuring the transparent and mandatory reporting of both discrimination and calibration metrics (e.g., calibration slopes, intercepts, and Brier scores), and proactively address class imbalance with subsequent model recalibration. Closing this calibration reporting gap is paramount before these tools can be safely translated to bedside risk stratification. Before any consideration for widespread clinical implementation, promising prediction models require rigorous external validation in diverse, real-world settings to demonstrate robust, generalizable, and clinically useful performance.

Acknowledgments

The authors thank Kaiwei Zhang for his guidance on this study.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This research was funded by the National Natural Science Foundation of China (82160914), the Guizhou University of Traditional Chinese Medicine Postgraduate Scientific Research Fund Project (YCXKYB2025025), and the Guizhou Provincial Health Commission Science and Technology Fund Project (gzwkj2026-159).

Footnotes

Edited by: Jun Li, Second Hospital of Anhui Medical University, China

Reviewed by: Huadong Ni, Affiliated Hospital of Jiaxing University, China

M. W. Geda, The Hong Kong Polytechnic University, Hong Kong SAR, China

Data availability statement

The original contributions presented in this study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

JF: Data curation, Formal analysis, Investigation, Project administration, Writing – original draft. MM: Data curation, Formal analysis, Project administration, Writing – original draft. CO: Data curation, Project administration, Writing – original draft. KZ: Data curation, Formal analysis, Investigation, Project administration, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmed.2026.1855638/full#supplementary-material

Table_1.docx (14.8KB, docx)

References

  • 1.Adelani MAA, Marx CMM, Humble S. Are neighborhood characteristics associated with outcomes after Tha and Tka? Findings from a large healthcare system database. Clin Orthop Relat Res. (2023) 481:226–35. 10.1097/corr.0000000000002222 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Bell JA, Emara AK, Barsoum WK, Bloomfield M, Briskin I, Higuera C, et al. Should an age cutoff be considered for elective total knee arthroplasty patients? An analysis of operative success based on patient-reported outcomes. J Knee Surg. (2023) 36:1001–11. 10.1055/s-0042-1748821 [DOI] [PubMed] [Google Scholar]
  • 3.Fares MY, Liu HH, da Silva Etges APB, Zhang B, Warner JJP, Olson JJ, et al. Utility of machine learning, natural language processing, and artificial intelligence in predicting hospital readmissions after orthopaedic surgery: a systematic review and meta-analysis. JBJS Rev. (2024) 12:00075. 10.2106/jbjs.Rvw.24.00075 [DOI] [PubMed] [Google Scholar]
  • 4.Lopez CD, Gazgalis A, Boddapati V, Shah RP, Cooper HJ, Geller JA. Artificial learning and machine learning decision guidance applications in total hip and knee arthroplasty: a systematic review. Arthroplasty Today. (2021) 11:103–12. 10.1016/j.artd.2021.07.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Mohammadi R, Jain S, Namin AT, Heller MS, Palacholla R, Kamarthi S, et al. Predicting unplanned readmissions following a hip or knee arthroplasty: retrospective observational study. JMIR Med Informat. (2020) 8:e19761. 10.2196/19761 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Slezak J, Butler L, Akbilgic O. The role of frailty index in predicting readmission risk following total joint replacement using light gradient boosting machines. Informat Med Unlocked. (2021) 25:100657. 10.1016/j.imu.2021.100657 [DOI] [Google Scholar]
  • 7.Gould DJ, Bailey JA, Spelman T, Bunzli S, Dowsey MM, Choong PFM. Predicting 30-day readmission following total knee arthroplasty using machine learning and clinical expertise applied to clinical administrative and research registry data in an Australian cohort. Arthroplasty (London, England). (2023) 5:30. 10.1186/s42836-023-00186-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Klemt C, Tirumala V, Habibi Y, Buddhiraju A, Chen TL-W, Kwon Y-M. The utilization of artificial neural networks for the prediction of 90-day unplanned readmissions following total knee arthroplasty. Arch Orthop Trauma Surg. (2023) 143:3279–89. 10.1007/s00402-022-04566-3 [DOI] [PubMed] [Google Scholar]
  • 9.Wu JM, Cheng BW, Ou CY, Chiu JE, Tsou SS. Applying machine learning methods to predict the hospital re-admission within 30 days of total hip arthroplasty and hemiarthroplasty. J Healthcare Qual Res. (2023) 38:197–205. 10.1016/j.jhqr.2022.11.009 [DOI] [PubMed] [Google Scholar]
  • 10.Kunze KN, So MM, Padgett DE, Lyman S, MacLean CH, Fontana MA. Machine learning on medicare claims poorly predicts the individual risk of 30-day unplanned readmission after total joint arthroplasty, yet uncovers interesting population-level associations with annual procedure volumes. Clin Orthop Relat Res. (2023) 481:1745–59. 10.1097/CORR.0000000000002705 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Shaikh HJF, Botros M, Ramirez G, Thirukumaran CP, Ricciardi B, Myers TG. Comparable performance of machine learning algorithms in predicting readmission and complications following total joint arthroplasty with external validation. Arthroplasty (London, England). (2023) 5:58. 10.1186/s42836-023-00208-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Park J, Zhong X, Miley EN, Rutledge RS, Kakalecik J, Johnson MC, et al. Machine learning-based predictive models for 90-day readmission of total joint arthroplasty using comprehensive electronic health records and patient-reported outcome measures. Arthroplasty Today. (2024) 25:101308. 10.1016/j.artd.2023.101308 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Buddhiraju A, Shimizu MR, Seo HH, Chen TLW, RezazadehSaatlou M, Huang Z, et al. Generalizability of machine learning models predicting 30-day unplanned readmission after primary total knee arthroplasty using a nationally representative database. Med Biol Eng Comput. (2024) 62:2333–41. 10.1007/s11517-024-03075-2 [DOI] [PubMed] [Google Scholar]
  • 14.Khan ST, Pasqualini I, Rullán PJ, Tidd J, Jin Y, Klika AK, et al. Predictive modeling of medical- and orthopaedic-related 90-day readmissions following primary total hip arthroplasty. J Arthroplasty. (2024) 39:2812–9.e2. 10.1016/j.arth.2024.05.058 [DOI] [PubMed] [Google Scholar]
  • 15.Buddhiraju A, Shimizu MR, Chen TL-W, Seo HH, Bacevich BM, Xiao P, et al. Comparing prediction accuracy for 30-day readmission following primary total knee arthroplasty: the Acs-Nsqip risk calculator versus a novel artificial neural network model. Knee Surg Relat Res. (2025) 37:3. 10.1186/s43019-024-00256-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Khan ST, Pasqualini I, Rullan PJ, Cleveland Clinic Arthroplasty Group, Piuzzi NS, Tidd J. Predictive modeling of medical and orthopaedic-related 90-day-readmissions following primary total knee arthroplasty. J Arthroplasty. (2025) 40:286–93.e2. 10.1016/j.arth.2024.07.041 [DOI] [PubMed] [Google Scholar]
  • 17.Crespi Z, Khan U, Shafau AL, Nham F, Chen C, Little B, et al. Machine-learning prediction of 90-day readmission after primary total hip arthroplasty: analysis of 1,340 cases from the Michigan arthroplasty registry (Marcqi). J Orthop. (2025) 65:270–5. 10.1016/j.jor.2025.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Khan U, Crespi Z, Nham F, Chen C, Little B, Darwiche H. Machine-learning prediction of 90-day readmission following primary Tka: insights from 2,123 cases in the Michigan arthroplasty registry. J Clin Orthop Trauma. (2025) 71:103250. 10.1016/j.jcot.2025.103250 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Purbasari IY, Bayuseno AP, Isnanto RR, Winarni TI. Readmission Risk Prediction after Total Hip Arthroplasty Using Machine Learning and Hyperparameters Optimized with Bayesian Optimization. SSRN; (2025). 10.2139/ssrn.5071220 [DOI] [Google Scholar]
  • 20.Geda MW, Tang YM, Lee CKM. Applications of artificial intelligence in orthopaedic surgery: a systematic review and meta-analysis. Eng Appl Artif Intel. (2024) 133:108326. 10.1016/j.engappai.2024.108326 [DOI] [Google Scholar]
  • 21.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. (2021) 372:n71. 10.1136/bmj.n71 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. (2010) 21:128–38. 10.1097/EDE.0b013e3181c30fb2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Moons KGM, Damen JAA, Kaul T, Hooft L, Andaur Navarro C, Dhiman P, et al. Probast+Ai: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. (2025) 388:e082505. 10.1136/bmj-2024-082505 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. (1977) 33:159–74. 10.2307/2529310 [DOI] [PubMed] [Google Scholar]
  • 25.Debray TP, Damen JA, Riley RD, Snell K, Reitsma JB, Hooft L, et al. A framework for meta-analysis of prediction model studies with binary and time-to-event outcomes. Stat Methods Med Res. (2019) 28:2768–86. 10.1177/0962280218785504 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.IntHout J, Ioannidis JP, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard Dersimonian-Laird method. BMC Med Res Methodol. (2014) 14:25. 10.1186/1471-2288-14-25 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. (1997) 315:629–34. 10.1136/bmj.315.7109.629 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Korvink M, Hung CW, Wong PK, Martin J, Halawi MJ. Development of a novel prospective model to predict unplanned 90-day readmissions after total hip arthroplasty. J Arthroplasty. (2023) 38:124–8. 10.1016/j.arth.2022.07.017 [DOI] [PubMed] [Google Scholar]
  • 29.Goltz DE, Ryan SP, Hopkins TJ, Howell CB, Attarian DE, Bolognesi MP, et al. A novel risk calculator predicts 90-day readmission following total joint arthroplasty. J Bone Joint Surg Am Vol. (2019) 101:547–56. 10.2106/jbjs.18.00843 [DOI] [PubMed] [Google Scholar]
  • 30.Yu S, Garvin KL, Healy WL, Pellegrini VD, Jr., Iorio R. Preventing hospital readmissions and limiting the complications associated with total joint arthroplasty. J Am Acad Orthop Surg. (2015) 23:60–71. 10.5435/JAAOS-D-15-00044 [DOI] [PubMed] [Google Scholar]
  • 31.Zhao ZJ, Yang Y, Wei SR, Zhang ZQ, Yao M, Ni HD. Repetitive transcranial magnetic stimulation to prevent postoperative delirium in elderly arthroplasty patients: study protocol for a single-centre, prospective, randomized controlled trial. BMC Geriatr. (2026). 10.1186/s12877-026-07579-4 [Epub ahead of print]. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zhi T, Wei S, Kuang J, Zhou S, Yu D, Gao T, et al. Effects of opioid-free anesthesia combined with iliofascial nerve block on perioperative neurocognitive deficits in elderly patients undergoing hip fracture surgery: study protocol for a prospective, multicenter, parallel-group, randomized controlled trial. Trials. (2025) 26:122. 10.1186/s13063-025-08828-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Chabot PJ, Cirino CM, Gulotta LV. Machine learning: the what, why, and how. Semi Arthroplasty JSES. (2023) 33:857–61. 10.1053/j.sart.2023.06.018 [DOI] [Google Scholar]
  • 34.Karlin EA, Lin CC, Meftah M, Slover JD, Schwarzkopf R. The impact of machine learning on total joint arthroplasty patient outcomes: a systemic review. J Arthroplasty. (2023) 38:2085–95. 10.1016/j.arth.2022.10.039 [DOI] [PubMed] [Google Scholar]
  • 35.Staartjes VE, Kernbach JM. Importance of calibration assessment in machine learning–based predictive analytics. J Neurosurg Spine. (2020) 32:985–6. 10.3171/2019.12.SPINE191503 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table_1.docx (14.8KB, docx)

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.


Articles from Frontiers in Medicine are provided here courtesy of Frontiers Media SA

RESOURCES