Abstract
Backgroud: No universally accepted model exists for predicting bleeding risk in patients receiving low-molecular-weight heparin or fondaparinux. Objective: This study leveraged seven machine learning algorithms to build a short-term bleeding risk prediction platform for this population. Methods: This retrospective real-world observational study included hospitalized patients who received low-molecular-weight heparin or fondaparinux between January 2022 and December 2023. After applying predefined criteria, the cohort were randomly split into training (70%) and validation (30%) sets. Predictors were identified using LASSO regression. Seven machine learning models, including Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), Neural Network (NN), extreme gradient boosting (XGBoost), adaptive boosting (AdaBoost), and CatBoost, were developed and evaluated. The best-performing model was implemented as an internal web-based bleeding risk prediction tool. Results: Among 1,691 hospitalized patients receiving low-molecular-weight heparin or fondaparinux, 126 (7.5%) experienced bleeding events. The cohort was randomly split into training (n = 1,184) and validation (n = 507) sets. LASSO regression identified 12 predictors, including surgical site, pre-medication INR, hemoglobin, platelet count, renal function, body mass index (BMI), indication, and comorbidities. Seven machine learning models were developed and evaluated. In the validation cohort, CatBoost achieved the best discrimination (AUC = 0.659), followed by XGBoost (AUC = 0.651) and LR (AUC = 0.622). CatBoost also demonstrated the highest accuracy (86.0%) and F1 score (0.297), with strong specificity (89.2%) but limited sensitivity (42.9%). Although all models showed robust negative predictive performance (PR-AUC > 0.93), positive predictive capacity was modest (PR-AUC < 0.20) in validation. Based on its overall performance, CatBoost was deployed as an internal web-based bleeding risk calculator. Conclusions: CatBoost emerged as the optimal model among those tested for predicting bleeding risk in patients receiving low-molecular-weight heparin or fondaparinux, demonstrating modest but superior discrimination, acceptable calibration, and favorable clinical utility. However, the model had limited ability to correctly identify patients who experienced bleeding, as indicated by low positive predictive performance. Given its high negative predictive value, it was better suited for ruling out rather than confirming bleeding risk. A web-based risk calculator based on CatBoost has been developed for internal use. Nevertheless, prospective multicenter validation is required before clinical implementation.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-52178-3.
Keywords: Low-molecular-weight heparin, Fondaparinux, Bleeding risk, Machine learning, Prediction model, Risk calculator
Subject terms: Medical research, Risk factors
Introduction
Low-molecular-weight heparins (LMWH) and fondaparinux remained the preferred anticoagulants of choice in specific clinical settings, such as pregnancy, end-stage renal disease, or when rapid and reversible anticoagulation was required, despite their associated bleeding risk1. A large cohort study of 12,934 patients reported an absolute incidence of 2.5 major bleeding events per 1,000 patients treated with LMWH (95% CI, 1.7–3.5)2. Although this incidence appeared low, the widespread use of these agents meant that even rare bleeding complications could lead to significant morbidity or irreversible harm in vulnerable individuals. Thus, striking a balance between effective anticoagulation and minimal bleeding risk remained a persistent challenge in clinical practice.
To guide decision-making, numerous bleeding risk assessment tools were developed. In atrial fibrillation (AF), scores such as HAS-BLED3, ABC-Bleeding4, ATRIA5, DOAC6, GARFIELD-AF7, HEMORR2HAGES8, and ORBIT9 integrated clinical variables to stratify bleeding risk. Among them, HAS-BLED was widely used due to its simplicity and reliance on routinely available data3, while ABC-Bleeding improved predictive performance by incorporating biomarkers alongside clinical factors4. Similarly, in venous thromboembolism (VTE), models like VTE-BLEED10, RIETE11, IMPROVE12, ACCP13, EINSTEIN14, and VTE-PREDICT15 were proposed, often including comorbidities, age, sex, and prior bleeding history. For instance, VTE-BLEED demonstrated good accuracy using six objective parameters10, and RIETE showed strong discriminative ability for in-hospital major bleeding in acute pulmonary embolism11.
However, these traditional tools were constrained by linear assumptions, limited variable interactions, and methodological limitations16–20. Their generalizability was further hampered by heterogeneous derivation populations and insufficient external validation21. Reflecting these concerns, the 2024 European Society of Cardiology (ESC) guidelines explicitly advised against routine use of existing bleeding risk scores due to uncertainties about their accuracy and the risk of withholding necessary anticoagulation22.
Critically, none of the current models were specifically developed or validated for patients receiving LMWH or fondaparinux. Traditional approaches often failed to capture the complex, nonlinear interplay of demographic, laboratory, and comorbidity-related factors that modulated bleeding risk in this context23.
Machine learning (ML) offered a promising alternative by enabling data-driven modeling of high-dimensional, real-world clinical data. Supervised ML algorithms, including decision trees, random forests, and gradient boosting machines, demonstrated superior performance in risk prediction, phenotypic classification, and biomarker discovery across diverse medical domains23,24. By learning intricate patterns from heterogeneous inputs (e.g., lab values, medication history, comorbidities), ML models generated individualized predictions beyond the scope of conventional scoring systems.
Given the paucity of validated, dedicated tools for patients receiving LMWH or fondaparinux, this study aimed to develop and internally validate an interpretable predictive model based on machine learning using real-world clinical data. The model would be further implemented as a web-based risk calculator to facilitate individualized bleeding risk assessment and support timely, targeted clinical interventions in this high-risk population.
Methods
Study design
This retrospective, real-world observational study was conducted in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines25. The analytical workflow comprised four sequential phases: (1) patient enrollment, (2) data collection and preprocessing, (3) model development and internal validation, and (4) deployment of a web-based prediction tool. We identified hospitalized patients who received LMWH or fondaparinux between January 2022 and December 2023 at Dongyang People’s Hospital and applied predefined exclusion criteria to establish the final cohort. Clinical variables were extracted from electronic medical records, independently preprocessed, and then randomly partitioned into training (70%) and validation (30%) sets. Seven machine learning models, including Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), Neural Network (NN), extreme gradient boosting (XGBoost), adaptive boosting (AdaBoost), and CatBoost, were developed and validated. The best-performing model was used to develop an internal web-based bleeding risk calculator. The overall workflow was illustrated in Fig. 1.
Fig. 1.
Overall workflow chart. Abbreviations: LASSO, Least Absolute Shrinkage and Selection Operator; SVM, Support Vector Machine; GBM, Gradient Boosting Machine; ROC, receiver operating characteristic; DCA, decision curve analysis.
This study was performed in accordance with the ethical principles of the Declaration of Helsinki and approved by the Ethics Committee of Dongyang People’s Hospital (Approval No. Dong Ren Yi 2025-YX-096, May 2025). Given its observational and retrospective design, all data were anonymized and de-identified, and informed consent was waived by the Ethics Committee of Dongyang People’s Hospital. All methods were performed in accordance with the relevant guidelines and regulations.
Patient enrollment
The study population included hospitalized patients aged ≥ 16 years who received LMWH or fondaparinux at Dongyang People’s Hospital (affiliated with Wenzhou Medical University) between January 2022 and December 2023. Patients were excluded if they met any of the following criteria: (1) missing key clinical information, (2) duplicated records, (3) history of bleeding within one month prior to drug administration, (4) concomitant or subsequent use of other anticoagulants within seven days after discontinuation, (5) receipt of continuous blood purification therapy, hemoperfusion, or hemodialysis, (6) catheter locking or flushing, or (7) surgery with intraoperative use of other anticoagulants during treatment.
Data collection
This study collected demographic and clinical data from patients receiving LMWH or fondaparinux. To avoid duplication, only the first administration was included for patients with multiple treatments. Variables included age, sex, body mass index (BMI), alcohol and smoking history, clinical and laboratory data, and comorbidity profiles.
Clinical and laboratory data were organized into nine domains: (1) LMWH or fondaparinux medication details (drug name, medication indication, and medication start and end times); (2) concomitant medication affecting bleeding risk (NSAIDs, antiplatelet drugs and additional hemostasis-altering agents); (3) surgical procedures (any form of surgery, and the “non-surgery” group exclusively referring to medically managed internal medicine patients without any surgical or unrecorded minor invasive procedures); (4) coagulation parameters (international normalized ratio (INR) and activated partial thromboplastin time (APTT)); (5) routine blood test results (pre-medication hemoglobin (Hb) and platelet counts (PLT)); (6) hepatic and renal function makers (creatinine, total bilirubin, alanine aminotransferase (ALT), and aspartate aminotransferase (AST)); (7) fall risk assessment; (8) imaging findings (computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, and other radiological reports); and (9) adverse events (onset time, preferred term, severity, outcome, and causality). All predictor variables, including laboratory test results (e.g., Hb, INR, APTT, creatinine), were restricted to the most recent measurement obtained prior to the first administration of LMWH or fondaparinux. Measurements collected more than 7 days before the initiation of LMWH or fondaparinux were considered missing.
The estimated glomerular filtration rate (eGFR) was calculated using the chronic kidney disease epidemiology collaboration (CKD-EPI) formula, and renal function was categorized according to the criteria outlined in Supplementary Table 126,27. Hepatic impairment was graded based on the national cancer institute organ dysfunction working group (NCI-ODWG) criteria (Supplementary Table 2)28. Fall risk was assessed using the morse fall scale and categorized as low, moderate, or high risk based on thresholds in Supplementary Table 329. This standardized categorization ensured consistent and clinically meaningful interpretation of key baseline variables.
Endpoints and follow-up
The primary outcome was bleeding-related adverse events occurring from the the first administration of LMWH or fondaparinux through 7 days after discontinuation. Bleeding events included organ hemorrhage (e.g., intracranial, renal, hepatic), gingival bleeding, epistaxis, skin ecchymosis, gastrointestinal bleeding, occult blood in stool or urine, and a Hb decline > 20 g/L. Patients were followed via inpatient records and outpatient visit documentation. For patients discharged before completing the follow-up window, outpatient records were reviewed to capture any bleeding events. If no visit occurred within 7 days post-discharge, the patient was classified as event-free. In this analysis, occurrence of any bleeding event was defined as a positive outcome, and absence as negative.
Data pre-processing
Clinical data were manually verified and rigorously preprocessed, including outlier removal and missing data handing. Variables with less than 20% missing values were imputed using a training-set–only nested multiple imputation approach. Specifically, the complete cohort (defined as having no missing values in the primary outcome variable “Adverse”) was first randomly split into a training set (70%) and a held-out validation set (30%) using a fixed random seed (seed = 52) to ensure reproducibility. Multiple imputation was then performed exclusively on the training set, generating five imputed datasets via group-stratified median/mode imputation (stratified by Indication and Surgery_site). No model fitting or parameter tuning was conducted on the validation set. The validation set was imputed using only the reference statistics derived from the imputed training set: median for continuous variables and the mode for categorical variables. This approach prevented data leakage and preserved the independence of the validation phase. Variables with more than 20% missing data were excluded from all subsequent analyses. Finally, the imputed training and validation datasets were formatted (including recoding factor variable and numeric conversion for model compatibility) and saved separately, with no further data modifications applied to the validation set thereafter.
Model development
All model development procedures were confined to the imputed training set, where patients were stratified into bleeding (Adverse = Yes) and non-bleeding (Adverse = No) groups. Initially, univariable analyses were performed to compare baseline characteristics between the two groups using appropriate statistical tests, as detailed in Statistical Methods. Subsequently, feature selection was conducted via Least Absolute Shrinkage and Selection Operator (LASSO) regression for binary classification (family = “binomial”, alpha = 1) using the R package “glmnet”. A 10-fold cross-validation (CV) on the training set identified the optimal penalty parameter (λmin) that minimized the cross-validated mean squared error (MSE). Coefficient trajectories were visualized to confirm feature shrinkage, and features with non-zero coefficients were retained as the initial candidate set. A 10-fold nested CV was further implemented to assess feature stability; features selected in ≥ 70% of the folds were defined as stable predictive features for subsequent modeling. The entire feature selection process was strictly restricted to the training set, with no involvement of the validation set.
Hyperparameter configuration followed, during which seven machine learning models, LR, SVM, GBM, NN, XGBoost, AdaBoost, and CatBoost, were trained on the stable features using fixed manual hyperparameter settings (no iterative tuning). The area under the receiver operating characteristic curve (AUC) was used as the optimization metric to select hyperparameter combinations, which remained unchanged after training. Thereafter, the optimal classification threshold for each model was determined on the training set and fixed for all subsequent evaluations, without re-optimization on the validation set.
Finally, the final training-phase models were fitted on the complete imputed training set using the stable features, fixed manual hyperparameters, and the predetermined optimal thresholds. All model objects, including feature importance scores, were saved for validation, and no further retraining or parameter modification was performed after this stage.
Model evaluation and display
An independent hold-out validation set was used for a one-time, non-invasive final evaluation of the trained models. To prevent overfitting and ensure unbiased performance estimation, no model refitting, hyperparameter tuning, or threshold re-selection was performed. Strict prohibitions were enforced against any data leakage from the validation set to the training phase throughout the entire workflow, and the validation set served as a standalone independent cohort for final performance assessment.
To quantify model performance on the validation set, a comprehensive panel of key predictive metrics was computed, including discrimination metrics (e.g., AUC, sensitivity, specificity), calibration metrics (e.g., calibration slope, intercept), clinical utility metrics (e.g., net benefit from decision curve analysis (DCA)), and diagnostic accuracy metrics (e.g., positive/negative predictive values, F1 score).
The optimal model was identified through a holistic assessment encompassing predictive accuracy, calibration, and clinical utility. For model interpretability, SHapley Additive exPlanations (SHAP) was applied to the optimal model using only the features selected during training, with the assistance of the R packages “shapviz” and “kernelshap”. This generated two types of visualizations: beeswarm plots to depict feature importance and the direction of each feature’s effect on predicted risk, and waterfall plots to enable decomposition of individual patient-level bleeding risk predictions. These visualizations enhanced the model’s transparency and clinical interpretability.
Finally, an internal web-based risk calculator was developed using the R package “shiny”, based on the optimal model, the finalized feature set, and the fixed optimal classification threshold. The calculator enabled real-time input of patient clinical characteristics and output the predicted probability of bleeding.
Statistical Methods
Categorical variables were analyzed using the chi-square (χ²) test or Fisher’s exact test, as appropriate. Continuous variables were evaluated for normality and homogeneity of variance. Normally distributed data were analyzed using Student’s t-test, and non-normally distributed data were evaluated with the Wilcoxon rank-sum test.
All statistical analyses were performed in R (version 4.4.1), with key packages assigned to specific procedures: “tableone” was used for baseline table summarization; “glmnet” and “caret” were used for LASSO regression and feature selection; “glm”, “e1071”, “gbm”, “nnet”, “xgboost”, “adabag”, and “catboost” were used to fit LR, SVM, GBM, NN, XGBoost, AdaBoost, and CatBoost models, respectively; “pROC”, “ROCR”, and “plotROC” were used for AUC calculation and ROC curve generation; “rms” and “riskRegression” were used for calibration analyses; “rmda” and “dcurves” were used for DCA; and “epiR” was used for diagnostic metric calculation during performance evaluation. For model interpretability, “shapviz” and “kernelshap” were employed for SHAP-based analysis. Finally, “shiny” and “shinythemes” were used to develop the web-based risk calculator.
A two-tailed P value < 0.05 was defined as statistically significant. All random processes were controlled using fixed random seeds to ensure reproducibility. The versions of all R packages and the detailed hyperparameter settings used during the training phase are provided in Supplementary Tables 4 and 5, respectively.
Results
Demographic information and clinical characteristics
A total of 1691 patients were included in the study, of whom 126 experienced bleeding. Among 22 variables, all had less than 20% missingness (Supplementary Fig. 1). Liver function grade had the highest missing rate (7.87%), followed by body mass index (BMI, 7.63%), INR (7.21%), APTT (7.21%), and eGFR (4.73%). The cohort was randomly partitioned into a training cohort (n = 1,184) and a validation cohort (n = 507) in a 7:3 ratio. Table 1 summarizes the baseline characteristics between the training and validation cohorts.
Table 1.
Demographics and clinical characteristics of study in the training and validation cohorts.
| Variables | Total, n = 1691 | Training, n = 1184 | Validation, n = 507 |
|---|---|---|---|
| Bleeding, n(%) | |||
| No | 1565 (92.55%) | 1093 (92.31%) | 472 (93.10%) |
| Yes | 126 (7.45%) | 91 (7.69%) | 35 (6.90%) |
| Age, years | 49.00 (33.00–67.00) | 49.00 (32.00–67.00) | 50.00 (33.00–67.00) |
| Sex, n(%) | |||
| Female | 1172 (69.31%) | 810 (68.41%) | 362 (71.40%) |
| Male | 519 (30.69%) | 374 (31.59%) | 145 (28.60%) |
| BMI, kg/m2 | 24.76 (22.27–27.77) | 24.67 (22.13–27.73) | 24.97 (22.51–28.15) |
| INR pre | 0.98 (0.94–1.03) | 0.97 (0.94–1.03) | 0.98 (0.94–1.02) |
| APTT pre, s | 34.20 (32.00–36.70) | 34.10 (32.10–36.60) | 34.30 (31.80–36.85) |
| Hb pre, g/L | 118.00 (106.00–131.00) | 118.00 (105.75–130.00) | 120.00 (109.00–133.00) |
| PLT pre, ×10^9/L | 203.00 (166.00–244.00) | 202.00 (166.00–242.00) | 207.00 (166.50–246.50) |
| Fall risk, n(%) | |||
| Low | 67 (3.96%) | 45 (3.80%) | 22 (4.34%) |
| Moderate | 1129 (66.77%) | 788 (66.55%) | 341 (67.26%) |
| High | 495 (29.27%) | 351 (29.65%) | 144 (28.40%) |
| Alcohol frequency, n(%) | |||
| Never | 1403 (82.97%) | 974 (82.26%) | 429 (84.62%) |
| Occasional | 103 (6.09%) | 70 (5.91%) | 33 (6.51%) |
| Regular | 117 (6.92%) | 88 (7.43%) | 29 (5.72%) |
| Cessation | 68 (4.02%) | 52 (4.39%) | 16 (3.16%) |
| Smoke frequency, n(%) | |||
| Never | 1400 (82.79%) | 964 (81.42%) | 436 (86.00%) |
| Occasional | 14 (0.83%) | 12 (1.01%) | 2 (0.39%) |
| Regular | 129 (7.63%) | 102 (8.61%) | 27 (5.33%) |
| Cessation | 148 (8.75%) | 106 (8.95%) | 42 (8.28%) |
| Pregnancy status, n(%) | |||
| No | 1079 (63.81%) | 756 (63.85%) | 323 (63.71%) |
| Yes | 612 (36.19%) | 428 (36.15%) | 184 (36.29%) |
| Hypertension, n(%) | |||
| No | 1189 (70.31%) | 831 (70.19%) | 358 (70.61%) |
| Yes | 502 (29.69%) | 353 (29.81%) | 149 (29.39%) |
| Diabetes, n(%) | |||
| No | 1458 (86.22%) | 1031 (87.08%) | 427 (84.22%) |
| Yes | 233 (13.78%) | 153 (12.92%) | 80 (15.78%) |
| Cancer, n(%) | |||
| No | 1296 (76.64%) | 924 (78.04%) | 372 (73.37%) |
| Yes | 395 (23.36%) | 260 (21.96%) | 135 (26.63%) |
| Cardiac insufficiency, n(%) | |||
| No | 1625 (96.10%) | 1131 (95.52%) | 494 (97.44%) |
| Yes | 66 (3.90%) | 53 (4.48%) | 13 (2.56%) |
| No. of concomitant drugs, n(%) | |||
| 0 | 1204 (71.20%) | 852 (71.96%) | 352 (69.43%) |
| 1 | 386 (22.83%) | 256 (21.62%) | 130 (25.64%) |
| 2 | 56 (3.31%) | 42 (3.55%) | 14 (2.76%) |
| 3 | 21 (1.24%) | 14 (1.18%) | 7 (1.38%) |
| >=4 | 24 (1.42%) | 20 (1.69%) | 4 (0.79%) |
| Surgery site, n(%) | |||
| No | 307 (18.15%) | 214 (18.07%) | 93 (18.34%) |
| Pelvic cavity | 803 (47.49%) | 569 (48.06%) | 234 (46.15%) |
| Abdomen | 34 (2.01%) | 22 (1.86%) | 12 (2.37%) |
| Blood vessel | 244 (14.43%) | 169 (14.27%) | 75 (14.79%) |
| Chest | 265 (15.67%) | 183 (15.46%) | 82 (16.17%) |
| Others | 38 (2.25%) | 27 (2.28%) | 11 (2.17%) |
| Heparin type, n(%) | |||
| Dalteparin | 297 (17.56%) | 208 (17.57%) | 89 (17.55%) |
| Nadroparin | 1330 (78.65%) | 927 (78.29%) | 403 (79.49%) |
| Others | 64 (3.78%) | 49 (4.14%) | 15 (2.96%) |
| Duration, n(%) | |||
| 1 ~ 3 days | 499 (29.51%) | 347 (29.31%) | 152 (29.98%) |
| 4 ~ 7 days | 1138 (67.30%) | 793 (66.98%) | 345 (68.05%) |
| 8 ~ 14 days | 44 (2.60%) | 37 (3.12%) | 7 (1.38%) |
| >=15 days | 10 (0.59%) | 7 (0.59%) | 3 (0.59%) |
| eGFR pre, n(%) | |||
| G1 | 1346 (79.60%) | 949 (80.15%) | 397 (78.30%) |
| G2 | 270 (15.97%) | 180 (15.20%) | 90 (17.75%) |
| G3a | 44 (2.60%) | 30 (2.53%) | 14 (2.76%) |
| G3b | 15 (0.89%) | 12 (1.01%) | 3 (0.59%) |
| >=G4 | 16 (0.95%) | 13 (1.10%) | 3 (0.59%) |
| Liver function grade, n(%) | |||
| Normal | 1475 (87.23%) | 1030 (86.99%) | 445 (87.77%) |
| Mild | 195 (11.53%) | 137 (11.57%) | 58 (11.44%) |
| Moderate | 18 (1.06%) | 15 (1.27%) | 3 (0.59%) |
| Severe | 3 (0.18%) | 2 (0.17%) | 1 (0.20%) |
| Indication, n(%) | |||
| ATE Prevention | 45 (2.66%) | 36 (3.04%) | 9 (1.78%) |
| ATE Treatment | 6 (0.35%) | 2 (0.17%) | 4 (0.79%) |
| Atrial Fibrillation | 25 (1.48%) | 16 (1.35%) | 9 (1.78%) |
| VTE Prevention | 1539 (91.01%) | 1073 (90.62%) | 466 (91.91%) |
| VTE Treatment | 76 (4.49%) | 57 (4.81%) | 19 (3.75%) |
Notes: Fall risk was assessed using the Morse Fall Scale. eGFR was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) formula. Liver function impairment was graded according to the NCI-ODWG criteria.
Abbreviations: BMI, Body Mass Index; INR, International Normalized Ratio; APTT, Activated Partial Thromboplastin Time; Hb, Hemoglobin; PLT, Platelet count; eGFR, Estimated Glomerular Filtration Rate; ATE, arterial thromboembolism; VTE, venous thromboembolism.
The training cohort was stratified by bleeding occurrence and compared using univariable analysis (Table 2). Among 1184 patients, 91 experienced bleeding events. Significant differences were observed in age (p < 0.001), sex (p = 0.004), BMI (p < 0.001), pre-medication INR (p < 0.001) and APTT (p = 0.002), baseline Hb (p = 0.009), baseline PLT (p = 0.041), smoking and alcohol history (p = 0.002 and p = 0.048), pregnancy status (p < 0.001), comorbid hypertension (p = 0.005), or cardiac insufficiency (p < 0.001), surgical site (p < 0.001), type of heparin administered (p < 0.001), baseline eGFR (p < 0.001), and indication (p < 0.001). Specifically, compared with the non-bleeding group, the bleeding group was characterized by a higher proportion of males (45.05% vs. 30.47%), older age (65 years vs. 48 years), and greater prevalence of smoking and alcohol use (24.18% vs. 18.12%, 24.18% vs. 17.20%), and comorbid hypertension (42.86% vs. 28.73%) or cardiac insufficiency (13.19% vs. 3.75%). Pre-medication INR (1.01 vs. 0.97) and APTT (35.46 s vs. 34.10 s) levels were elevated, and renal function (G1: 62.64% vs. 81.61%) was more impaired in bleeding group. A higher proportion of patients in the bleeding group received LMWH or fondaparinux for VTE treatment (17.58% vs. 3.75%). In contrast, the non-bleeding group had a higher mean BMI (24.77 kg/m2 vs. 23.61 kg/m2) and a greater proportion of patients receiving dalteparin or nadroparin (96.43% vs. 89.01%). No significant differences were observed between the two groups in history of diabetes, malignancy, fall risk, concomitant medication use, or duration of drug therapy (P > 0.05).
Table 2.
Univariate differences in the training cohort.
| Variables | Total, n = 1184 | No Bleeding, n = 1093 | Bleeding, n = 91 | P |
|---|---|---|---|---|
| Age, years | 49.00 (32.00–67.00) | 48.00 (32.00–66.00) | 65.00 (41.50–75.00) | < 0.001 |
| Sex, n(%) | 0.004 | |||
| Male | 374 (31.59%) | 333 (30.47%) | 41 (45.05%) | |
| Female | 810 (68.41%) | 760 (69.53%) | 50 (54.95%) | |
| BMI, kg/m2 | 24.67 (22.13–27.73) | 24.77 (22.27–27.93) | 23.61 (20.94–25.53) | < 0.001 |
| INR pre | 0.97 (0.94–1.03) | 0.97 (0.94–1.03) | 1.01 (0.95–1.08) | < 0.001 |
| APTT pre, s | 34.10 (32.10–36.60) | 34.10 (32.00–36.50) | 35.46 (32.75–38.20) | 0.002 |
| Hb pre, g/L | 118.00 (105.75–130.00) | 117.00 (105.00–130.00) | 124.00 (106.50–141.00) | 0.009 |
| PLT pre, ×10^9/L | 202.00 (166.00–242.00) | 203.00 (167.00–242.00) | 188.00 (141.00–236.50) | 0.041 |
| Fall risk, n(%) | 0.245 | |||
| Low | 45 (3.80%) | 42 (3.84%) | 3 (3.30%) | |
| Moderate | 788 (66.55%) | 734 (67.15%) | 54 (59.34%) | |
| High | 351 (29.65%) | 317 (29.00%) | 34 (37.36%) | |
| Alcohol frequency, n(%) | 0.048 | |||
| Never | 974 (82.26%) | 905 (82.80%) | 69 (75.82%) | |
| Occasional | 70 (5.91%) | 63 (5.76%) | 7 (7.69%) | |
| Regular | 88 (7.43%) | 82 (7.50%) | 6 (6.59%) | |
| Cessation | 52 (4.39%) | 43 (3.93%) | 9 (9.89%) | |
| Smoke frequency, n(%) | 0.002 | |||
| Never | 964 (81.42%) | 895 (81.88%) | 69 (75.82%) | |
| Occasional | 12 (1.01%) | 10 (0.91%) | 2 (2.20%) | |
| Regular | 102 (8.61%) | 99 (9.06%) | 3 (3.30%) | |
| Cessation | 106 (8.95%) | 89 (8.14%) | 17 (18.68%) | |
| Pregnancy status, n(%) | < 0.001 | |||
| No | 756 (63.85%) | 683 (62.49%) | 73 (80.22%) | |
| Yes | 428 (36.15%) | 410 (37.51%) | 18 (19.78%) | |
| Hypertension, n(%) | 0.005 | |||
| No | 831 (70.19%) | 779 (71.27%) | 52 (57.14%) | |
| Yes | 353 (29.81%) | 314 (28.73%) | 39 (42.86%) | |
| Diabetes, n(%) | 0.369 | |||
| No | 1031 (87.08%) | 949 (86.83%) | 82 (90.11%) | |
| Yes | 153 (12.92%) | 144 (13.17%) | 9 (9.89%) | |
| Cancer, n(%) | 0.595 | |||
| No | 924 (78.04%) | 855 (78.23%) | 69 (75.82%) | |
| Yes | 260 (21.96%) | 238 (21.77%) | 22 (24.18%) | |
| Cardiac insufficiency, n(%) | < 0.001 | |||
| No | 1131 (95.52%) | 1052 (96.25%) | 79 (86.81%) | |
| Yes | 53 (4.48%) | 41 (3.75%) | 12 (13.19%) | |
| No. of concomitant drugs, n(%) | 0.071 | |||
| 0 | 852 (71.96%) | 795 (72.74%) | 57 (62.64%) | |
| 1 | 256 (21.62%) | 234 (21.41%) | 22 (24.18%) | |
| 2 | 42 (3.55%) | 35 (3.20%) | 7 (7.69%) | |
| 3 | 14 (1.18%) | 12 (1.10%) | 2 (2.20%) | |
| >=4 | 20 (1.69%) | 17 (1.56%) | 3 (3.30%) | |
| Surgery site, n(%) | < 0.001 | |||
| No | 214 (18.07%) | 173 (15.83%) | 41 (45.05%) | |
| Pelvic cavity | 569 (48.06%) | 546 (49.95%) | 23 (25.27%) | |
| Abdomen | 22 (1.86%) | 19 (1.74%) | 3 (3.30%) | |
| Blood vessel | 169 (14.27%) | 165 (15.10%) | 4 (4.40%) | |
| Chest | 183 (15.46%) | 169 (15.46%) | 14 (15.38%) | |
| Others | 27 (2.28%) | 21 (1.92%) | 6 (6.59%) | |
| Heparin type, n(%) | < 0.001 | |||
| Dalteparin | 927 (78.29%) | 855 (78.23%) | 72 (79.12%) | |
| Nadroparin | 208 (17.57%) | 199 (18.21%) | 9 (9.89%) | |
| Others | 49 (4.14%) | 39 (3.57%) | 10 (10.99%) | |
| Duration, n(%) | 0.212 | |||
| 1 ~ 3 days | 347 (29.31%) | 320 (29.28%) | 27 (29.67%) | |
| 4 ~ 7 days | 793 (66.98%) | 736 (67.34%) | 57 (62.64%) | |
| 8 ~ 14 days | 37 (3.12%) | 31 (2.84%) | 6 (6.59%) | |
| >=15 days | 7 (0.59%) | 6 (0.55%) | 1 (1.10%) | |
| eGFR pre, n(%) | < 0.001 | |||
| G1 | 949 (80.15%) | 892 (81.61%) | 57 (62.64%) | |
| G2 | 180 (15.20%) | 161 (14.73%) | 19 (20.88%) | |
| G3a | 30 (2.53%) | 24 (2.20%) | 6 (6.59%) | |
| G3b | 12 (1.01%) | 5 (0.46%) | 7 (7.69%) | |
| >=G4 | 13 (1.10%) | 11 (1.01%) | 2 (2.20%) | |
| Liver function grade, n(%) | 0.053 | |||
| Normal | 1030 (86.99%) | 957 (87.56%) | 73 (80.22%) | |
| Mild | 137 (11.57%) | 119 (10.89%) | 18 (19.78%) | |
| Moderate | 15 (1.27%) | 15 (1.37%) | 0 (0.00%) | |
| Severe | 2 (0.17%) | 2 (0.18%) | 0 (0.00%) | |
| Indication, n(%) | < 0.001 | |||
| VTE Treatment | 57 (4.81%) | 41 (3.75%) | 16 (17.58%) | |
| VTE Prevention | 1073 (90.62%) | 1003 (91.77%) | 70 (76.92%) | |
| Atrial Fibrillation | 16 (1.35%) | 13 (1.19%) | 3 (3.30%) | |
| ATE Treatment | 2 (0.17%) | 2 (0.18%) | 0 (0.00%) | |
| ATE Prevention | 36 (3.04%) | 34 (3.11%) | 2 (2.20%) |
Notes: Fall risk was assessed using the Morse Fall Scale. eGFR was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) formula. Liver function impairment was graded according to the NCI-ODWG criteria.
Abbreviations: BMI, Body Mass Index; INR, International Normalized Ratio; APTT, Activated Partial Thromboplastin Time; Hb, Hemoglobin; PLT, Platelet count; eGFR, Estimated Glomerular Filtration Rate; ATE, arterial thromboembolism; VTE, venous thromboembolism.
Selection of predictors
LASSO regression was employed to select key predictors of bleeding events. As illustrated in Fig. 2A, the coefficient profiles across log(λ) values demonstrated progressive shrinkage of non-informative variables to zero, with 12 predictors retaining non-zero coefficients at the optimal regularization parameter (λ min = 0.00638, lambda.1se = 0.045). The corresponding binomial deviance plot (Fig. 2B) indicated the optimal model fit based on ten-fold cross-validation. The stability and magnitude of the coefficient trajectories suggested that these 12 variables (fall risk, number of concomitant drugs, surgical site, comorbid cardiac insufficiency, hypertension or diabetes, pre-medication INR, baseline Hb, baseline PLT, renal function, BMI, and indication) contributed meaningfully to bleeding risk prediction.
Fig. 2.
Results of variable selection via LASSO binary logistic regression. (A) Coefficient profile plots across the log(λ) sequence; (B) Optimal λ values indicated by vertical dotted lines based on the 1-SE criterion.
Multiple machine learning model performance
The performance of seven machine learning models was evaluated. In the training cohort, GBM achieved the highest discrimination (AUC = 0.852), followed by the CatBoost (AUC = 0.834), NN (AUC = 0.788), and SVM (AUC = 0.757), while Adaboost showed relatively lower performance (AUC = 0.656) (Fig. 3A). In the validation cohort, all models demonstrated slightly reduced AUCs, with CatBoost maintaining the best discrimination (AUC = 0.659), followed by XGBoost (AUC = 0.651), LR (AUC = 0.622), NN (AUC = 0.581), and GBM (AUC = 0.563) (Fig. 3B).
Fig. 3.
Comparative Model Performance Assessment. (A) ROC curve analysis in training cohort; (B) ROC curve analysis in validation cohort; (C) Decision curve analysis in training cohort; (D) Decision curve analysis in validation cohort; (E) Calibration Plot in training cohort; (F) Calibration Plot in validation cohort.
Decision curves showed the standardized net benefit of various machine learning models across high-risk thresholds on the training cohort (Fig. 3C) and validation cohort (Fig. 3D). Ensemble methods (CatBoost, XGBoost) and NN consistently outperformed traditional models such as LR. LR and XGBoost demonstrated minimal performance gap between sets, indicating strong generalization, whereas GBM and SVM showed substantially higher net benefit on the training cohort than on the validation cohort, indicating evident overfitting.
Calibration curves (Fig. 3E and F, Supplementary Table 6) assessed how well each model’s predicted probabilities (represented by bin midpoints on the x-axis) aligned with the observed event rates (y-axis). In the training set (Fig. 3E), the GBM model (green solid line) demonstrated the best calibration, with its curve closely following the ideal 45° reference line. LR, SVM, XGBoost, and NN also exhibited reasonably good calibration, whereas AdaBoost and CatBoost showed greater deviations. However, in the validation set (Fig. 3F), calibration performance deteriorated across most models. Notably, GBM, which performed best in training, displayed severe miscalibration, characterized by a sharply rising and then precipitously falling curve, suggesting poor generalization. Other models, including LR and AdaBoost, also deviated more substantially from the 45° line compared to their training performance. CatBoost, meanwhile, was represented by only a single data point, precluding assessment of its calibration trend. Due to this limitation, the calibration performance of the models in the validation set was difficult to interpret fully.
We analyzed both precision-recall (PR) curves, confusion matrices, matthews correlation coefficient, and balanced accuracy across the training and validation cohorts (Figs. 4 and 5, Supplementary Tables 7, and Supplementary Table 8). The PR curves (Fig. 4) revealed that in the training cohort, CatBoost achieved the highest positive PR-AUC of 0.502, while all models exhibited excellent negative predictive performance with PR-AUC values exceeding 0.94. However, in the validation cohort, the positive PR-AUC of all models dropped sharply to below 0.2, indicating significant attenuation in generalizability for positive sample detection, whereas negative predictive performance remained robust with PR-AUC values above 0.93. The confusion matrices (Fig. 5) further quantified these observations: AdaBoost, XGBoost, GBM and CatBoost demonstrated the strongest true negative (TN) recognition, with TN rates consistently above 79% in both cohorts; notably, AdaBoost achieved the highest TN rate of 85.1% in the training cohort, which remained stable at 85.0% in the validation cohort. In contrast, LR, NN, and SVM exhibited relatively higher true positive (TP) rates in the training cohort (5.2%–6.1%), though these rates declined to 3.2–3.6% in the validation cohort.
Fig. 4.
Comparison of Precision-Recall Curves Across Models. (A) positive predictive in training cohort; (B) negative predictive in training cohort; (C) positive predictive in validation cohort; (D) negative predictive in validation cohort.
Fig. 5.
Confusion Matrices for Seven Predictive Models on Training and Validation cohort Sets. (A) AdaBoost; (B) Gradient Boosting Machine (GBM); (C) Logistic Regression; (D) Neural Network; (E) Support Vector Machine (SVM); (F) XGBoost; (G) CatBoost. Notes: Target Yes (occurrence of the target bleeding adverse event) & Prediction Yes = TP (True Positive), Target Yes & Prediction No = FN (False Negative), Target No (non-occurrence of the target bleeding adverse event) & Prediction Yes = FP (False Positive), and Target No & Prediction No = TN (True Negative).
To evaluate the generalization performance of the seven predictive models, we compared their key diagnostic metrics across the training and validation cohorts (Tables 3 and 4). In the training cohort (Table 3), NN (0.791), SVM (0.725) and GBM (0.714) achieved the highest sensitivity, demonstrating superior ability to identify true positive samples, whereas AdaBoost (0.385) exhibited the lowest sensitivity, leading to the highest rate of false negatives. Conversely, AdaBoost (0.922), CatBoost (0.887), and XGBoost (0.878) showed the highest specificity, indicating strong performance in identifying true negative samples, while NN (0.683) had the lowest specificity, resulting in more false positives. For positive predictive performance, CatBoost (precision = 0.322, F1 = 0.431) and GBM (precision = 0.301, F1 = 0.423) balanced precision and recall most effectively, though all models displayed limited positive predictive power, with precision < 0.33 and F1 < 0.44. In terms of overall performance, CatBoost (accuracy = 0.868, Kappa = 0.366) and GBM (accuracy = 0.851, Kappa = 0.354) achieved the highest accuracy and kappa values, whereas LR (accuracy = 0.717, Kappa = 0.164) and NN (accuracy = 0.692, Kappa = 0.179) performed poorest. All models had Kappa values < 0.4, reflecting moderate-to-fair agreement between predictions and actual outcomes.
Table 3.
Comparative analysis of the performance outcomes across various machine learning models in the training cohorts.
| Model | Threshold | Sensitivity | Specificity | Precision | NPV | F1 | Accuracy | Kappa |
|---|---|---|---|---|---|---|---|---|
| Logistic | 0.077 | 0.670(0.569–0.758) | 0.721(0.694–0.747) | 0.167(0.132–0.208) | 0.963(0.948–0.974) | 0.267(0.228–0.309) | 0.717(0.691–0.742) | 0.164 |
| SVM | 0.077 | 0.725(0.626–0.806) | 0.726(0.698–0.751) | 0.180(0.144–0.223) | 0.969(0.955–0.979) | 0.289(0.249–0.332) | 0.726(0.699–0.750) | 0.189 |
| GBM | 0.097 | 0.714(0.614–0.797) | 0.862(0.840–0.881) | 0.301(0.244–0.365) | 0.973(0.961–0.982) | 0.423(0.369–0.479) | 0.851(0.829–0.870) | 0.354 |
| Neural Network | 0.064 | 0.791(0.697–0.862) | 0.683(0.655–0.710) | 0.172(0.139–0.211) | 0.975(0.962–0.984) | 0.283(0.246–0.324) | 0.692(0.665–0.717) | 0.179 |
| Xgboost | 0.463 | 0.560(0.458–0.658) | 0.878(0.858–0.896) | 0.277(0.218–0.346) | 0.960(0.946–0.970) | 0.371(0.316–0.429) | 0.854(0.833–0.873) | 0.299 |
| Adaboost | 0.114 | 0.385(0.291–0.487) | 0.922(0.905–0.937) | 0.292(0.218–0.378) | 0.947(0.932–0.959) | 0.332(0.272–0.398) | 0.881(0.861–0.898) | 0.268 |
| CatBoost | 0.529 | 0.648(0.546–0.739) | 0.887(0.866–0.904) | 0.322(0.259–0.393) | 0.968(0.955–0.977) | 0.431(0.373–0.490) | 0.868(0.848–0.886) | 0.366 |
Abbreviations: SVM, Support Vector Machine; GBM, Gradient Boosting Machine.
Table 4.
Comparative analysis of the performance outcomes across various machine learning models in the validation cohorts.
| Model | Threshold | Sensitivity | Specificity | Precision | NPV | F1 | Accuracy | Kappa |
|---|---|---|---|---|---|---|---|---|
| Logistic | 0.077 | 0.486(0.330–0.644) | 0.697(0.654–0.737) | 0.106(0.067–0.164) | 0.948(0.920–0.967) | 0.174(0.128–0.234) | 0.682(0.641–0.721) | 0.069 |
| SVM | 0.077 | 0.457(0.305–0.618) | 0.731(0.689–0.769) | 0.112(0.070–0.174) | 0.948(0.920–0.966) | 0.180(0.130–0.243) | 0.712(0.671–0.750) | 0.077 |
| GBM | 0.097 | 0.429(0.280–0.591) | 0.854(0.819–0.883) | 0.179(0.111–0.274) | 0.953(0.928–0.969) | 0.252(0.183–0.337) | 0.824(0.789–0.855) | 0.171 |
| Neural Network | 0.064 | 0.514(0.356–0.670) | 0.657(0.613–0.698) | 0.100(0.064–0.153) | 0.948(0.918–0.967) | 0.167(0.123–0.223) | 0.647(0.604–0.687) | 0.059 |
| Xgboost | 0.463 | 0.343(0.208–0.508) | 0.883(0.851–0.909) | 0.179(0.106–0.287) | 0.948(0.923–0.965) | 0.235(0.164–0.326) | 0.846(0.812–0.875) | 0.159 |
| Adaboost | 0.114 | 0.200(0.100–0.359) | 0.913(0.884–0.935) | 0.146(0.072–0.272) | 0.939(0.913–0.957) | 0.169(0.103–0.263) | 0.864(0.831–0.891) | 0.097 |
| CatBoost | 0.529 | 0.429(0.280–0.591) | 0.892(0.861–0.917) | 0.227(0.143–0.342) | 0.955(0.931–0.970) | 0.297(0.217–0.392) | 0.860(0.827–0.887) | 0.227 |
Abbreviations: SVM, Support Vector Machine; GBM, Gradient Boosting Machine.
In the validation cohort (Table 4), LR (0.486) and NN (0.514) achieved the highest sensitivity, while AdaBoost (0.200) again had the lowest sensitivity, indicating persistent challenges in positive sample detection. Specificity remained strong and stable, with AdaBoost (0.913), CatBoost (0.892), and XGBoost (0.883) maintaining the highest values, and NN (0.657) again showing the lowest specificity. Positive predictive performance further deteriorated in the validation cohort, with CatBoost (precision = 0.227, F1 = 0.297) and GBM (precision = 0.179, F1 = 0.252) still performing best, but all models exhibited precision < 0.23 and F1 < 0.30, meaning over three-quarters of positive predictions were false positives. For overall performance, CatBoost (accuracy = 0.860, Kappa = 0.227) and XGBoost (accuracy = 0.846, Kappa = 0.159) remained the top performers, while NN (accuracy = 0.647, Kappa = 0.059) again performed poorest.
Variable importance and variable interpretation
Figure 6 and Supplementary Fig. 2 presented the feature importance rankings derived from seven machine learning models. AdaBoost identified Hb_pre (pre-medication Hb), indication, and PLT_pre (pre-medication PLT) as the most influential features, with Hb_pre having the highest importance score. GBM prioritized PLT_pre, INR_pre (pre-medication INR), and Hb_pre as its top three features, reflecting the critical role of preoperative hematological parameters. LR, NN, SVM, and CatBoost all ranked Surgery_site as the most important feature, underscoring its consistent predictive value across diverse modeling approaches. XGBoost uniquely identified INR_pre as the most impactful feature, followed by Surgery_site and BMI, highlighting the importance of coagulation status in this model. Across all models, preoperative hematological parameters (Hb_pre, PLT_pre, INR_pre), surgery_site, indication, and patient demographics (BMI) emerged as the most recurrently important features. This consistency suggested that these variables play a central role in driving the model predictions, while clinical comorbidities (e.g., Diabetes, Hypertension) and functional assessments (e.g., Fall_risk) were assigned relatively lower importance.
Fig. 6.
Integrated feature importance network Across Models. Abbreviations: BMI, Body Mass Index; INR, International Normalized Ratio; Hb, Hemoglobin; eGFR, Estimated Glomerular Filtration Rate; GBM, Gradient Boosting Machine; SVM, Support Vector Machine.
To further interpret the CatBoost model’s predictions, SHAP plots were employed (Fig. 7). The SHAP summary plot (Fig. 7A) revealed the overall contribution of each feature to model predictions. Surgery_site and Hb_pre were the most impactful features, as indicated by their wide distribution of SHAP values. A higher feature value was associated with more positive SHAP values, increasing the predicted risk. Other notable features included indication, BMI, and Fall_risk, which also exhibited substantial SHAP value ranges, highlighting their role in driving model outputs.
Fig. 7.
SHAP analyses of the CatBoost model. (A) Global Feature Importance via SHAP Values; (B) SHAP-Based Local Explanation for an Individual Prediction.
The SHAP force plot (Fig. 7B) illustrated the contribution of each feature to a single patient’s prediction, where the baseline prediction (average model output) was 0.075. For this individual, the most significant feature was Surgery_site = Pelvic cavity, which decreased the predicted risk by 0.0147. This was followed by Hb_pre = 102 g/L and BMI = 26.7 kg/m², which also reduced the prediction. Conversely, features such as Hypertension and No history of diabetes had smaller positive contributions, slightly increasing the predicted risk. This granular breakdown demonstrates how individual features collectively shift the prediction from the baseline, enhancing the interpretability of the model.
Implementation of web calculator
Among all candidates, CatBoost demonstrated superior performance across model evaluations and was therefore selected as the optimal model. Leveraging the key predictors identified by CatBoost, we developed an internal web based risk calculator (available at http://127.0.0.1:7149) to enable individualized bleeding risk prediction for patients receiving LMWH or fondaparinux (Fig. 8).
Fig. 8.
A web-based calculator for predicting bleeding risk in patients with heparin or derivatives.
Discussion
Selection of Predictors and Machine Learning Algorithms
Existing bleeding risk models had commonly incorporated variables such as anemia, age, history of cancer, prior bleeding, renal insufficiency, and the use of antiplatelet agents21,22. Some models also included sex, PLT, use of nonsteroidal anti-inflammatory drugs (NSAIDs), uncontrolled hypertension, diabetes, and alcohol consumption21,22. Building on established predictors, we expanded our feature set to include variables with emerging or mechanistic relevance. Hepatic dysfunction and surgical procedures were also associated with an increased risk of bleeding30–32, prompting us to incorporate surgical site, history of liver disease, and baseline liver function. Furthermore, in addition to conventional prohemorrhagic agents such as NSAIDs and antiplatelets, we also considered additional drugs reported to impair hemostasis, including Cefoperazone-Sulbactam, Piperacillin, Diosmin, Tetramethylpyrazine, Urapidil, Edaravone and Linezolid33,34. While modeling each drug individually could theoretically improve clinical specificity, most of these agents were used infrequently, particularly in the bleeding group, where exposure counts for many drugs were zero or nearly zero (Supplementary Fig. 3). Such sparsity would compromise model stability, reduce statistical power, and increase the risk of biased estimates. By aggregating these agents into a single count variable, namely the number of concomitant medications, these challenges were effectively mitigated. In this retrospective analysis, we employed LASSO regression for automatic feature selection and identified twelve key clinical variables with strong predictive power. By integrating these diverse predictors, we aimed to enhance the accuracy of bleeding risk prediction and provide evidence-based support for clinical decision-making.
To balance performance and practicality, we rigorously evaluated seven machine learning algorithms, including Adaboost, GBM, LR, NN, SVM, XGBoost, and CatBoost. Each algorithm had been selected based on its ability to address specific modeling challenges. LR was chosen for its strong clinical interpretability and its well-established role as a benchmark in risk prediction35, as evidenced by its recent application in predicting osteoporosis risk among older adults at high cardiovascular risk36. SVM and NN were included due to their proven capacity to model high-dimensional, non-linear relationships in biomedical data37, a strength highlighted in studies on echinococcosis diagnosis38. Among ensemble methods, AdaBoost and GBM have consistently shown robustness on heterogeneous clinical datasets by iteratively correcting misclassifications39,40. Notably, XGBoost and CatBoost, two advanced gradient boosting implementations, had consistently ranked among the top-performing models in recent clinical predictive modeling competitions and real-world studies23,41,42. These empirical successes across diverse clinical domains provided strong justification for their inclusion in our comparative framework.
Optimal model selection
Among the seven machine learning models evaluated in this study, the CatBoost model emerged as the most reliable and robust predictive tool, supported by its superior comprehensive performance across key dimensions, including discriminative ability, generalization stability, balanced predictive power for positive and negative samples, and clinical applicability.
Discrimination (AUC)
In the training cohort, the GBM model achieved a marginally higher AUC (0.852) than CatBoost (0.834). However, all models exhibited a reduction in discriminative capacity when applied to the validation cohort, and CatBoost maintained the highest AUC (0.659) in this independent dataset, markedly outperforming GBM, whose AUC declined precipitously to 0.563. This observation underscores that CatBoost’s ability to distinguish between true positive and true negative cases remained more stable when generalized to unseen data, a critical attribute for clinical predictive models that must perform consistently in real-world settings.
Balanced predictive power for positive and negative samples
For positive sample prediction, CatBoost demonstrated the highest positive PR-AUC (0.502) in the training cohort, with an F1 score (0.431). In the validation cohort, although positive predictive performance deteriorated across all models (with PR-AUC values falling below 0.2), CatBoost still retained the optimal precision (0.227) and F1 score (0.297), exhibiting the least attenuation in positive predictive capability. For negative sample prediction, CatBoost maintained high specificity (training: 0.887; validation: 0.892) and true negative (TN) rates in both cohorts, consistent with AdaBoost and XGBoost as top performers in negative case recognition. Notably, its negative PR-AUC remained above 0.93 in the validation cohort. The models’ ability to correctly identify patients who will bleed remained limited, and that it was best suited for ruling out bleeding risk given its high negative predictive value rather than ruling it in. However, the near-perfect negative predictive value observed in this study was an expected consequence of the low event prevalence in the cohort, not as a standalone indicator of superior clinical performance.
Overall performance (Accuracy and Kappa Value)
In the training cohort, CatBoost achieved the highest accuracy (0.868) and Kappa value (0.366), reflecting superior overall prediction correctness and strong agreement between predicted and actual outcomes. In the validation cohort, it continued to lead with an accuracy of 0.860 and the highest Kappa value (0.227), whereas other models (e.g., GBM) showed more pronounced declines in Kappa values, indicating diminished consistency between predictions and real-world outcomes. This highlights CatBoost’s superior overall generalization ability compared to its counterparts.
Overfitting risk
The GBM model performed exceptionally well in the training cohort (e.g., optimal calibration and high AUC) but suffered from severe overfitting in the validation cohort, as evidenced by a sharp drop in AUC and severe miscalibration. The NN model, despite achieving the highest sensitivity (0.791) in the training cohort, exhibited far greater decline in sensitivity (0.514) and F1 score (0.167) in the validation cohort. In contrast, CatBoost showed minimal fluctuations in key performance metrics (AUC, precision, accuracy) between the training and validation cohorts, indicating a lower risk of overfitting and enhanced robustness.
Clinical applicability
Beyond its superior statistical performance, CatBoost effectively identified key predictive variables that are clinically meaningful (e.g., surgical site, pre-medication Hb), which directly supported the development of a user-friendly web-based risk calculator for individualized bleeding risk assessment in patients receiving LMWH or fondaparinux. This integration of robust performance with practical clinical translatability further solidifies CatBoost’s status as the optimal model for the study’s objectives.
In contrast, other models had notable limitations: GBM lacked generalization stability; XGBoost showed weak positive predictive power; LR and SVM had low overall accuracy and significant performance reduction; and AdaBoost, while strong in negative recognition, exhibited the poorest ability to identify positive samples. Collectively, these findings confirmed that CatBoost was the most balanced and reliable predictive model for the study’s objectives.
Interpretation and clinical relevance of key risk variables
In this study, the CatBoost model (identified as optimal) ranked risk variables by importance: Surgery_site emerged as the most critical factor, followed by Hb_pre, a key hematological indicator. Variables such as Indication, Fall_risk, and PLT_pre showed moderate importance, while comorbidities (e.g., Hypertension, Diabetes) had the least impact.
Across the seven models, notable differences in variable importance emerged. Surgery_site was a consensus core predictor—ranked first in LR, NN, SVM, and CatBoost. However, discrepancies existed: AdaBoost, GBM, and XGBoost placed greater emphasis on hematological indicators (e.g., Hb_pre, PLT_pre, INR_pre) than on Surgery_site. Although LR also ranked Surgery_site first, it assigned substantially higher weight to Indication compared to tree-based models. The importance of comorbidities also varied: tree-based models (e.g., CatBoost) treated them as secondary factors, whereas traditional models (LR, SVM) assigned them slightly higher weight. These differences reflected inherent algorithmic “learning preferences”. Tree-based models excelled at capturing nonlinear interactions and prioritized high-discrimination features; linear models effectively balanced interpretable linear associations; and neural networks learned complex, high-dimensional patterns through multi-layer architectures.
In conclusion, Surgery_site and hematological indicators (Hb_pre, PLT_pre, INR_pre) consistently served as core predictors across all models. CatBoost’s strength lay in its ability to balance clinical and hematological features while avoiding overbias, which was a key reason for its superior performance.
A notable finding of our study was an inverse association between surgical intervention and bleeding risk, contrary to the prevailing literature, which generally reported a positive relationship43–47. While major surgery was biologically plausible as a risk factor for post-discharge bleeding, our real-world observational cohort exhibited substantial baseline differences between surgical and non-surgical patients across multiple clinical and demographic characteristics. Specifically, patients undergoing pelvic (Supplementary Table 9), thoracic (Supplementary Table 10), vascular (Supplementary Table 11), abdominal (Supplementary Table 12), or other surgeries (Supplementary Table 13) were significantly younger, less likely to use concomitant medications, and more frequently reported no history of alcohol consumption or smoking. They also had lower prevalences of diabetes, hypertension, and cardiac insufficiency, along with better renal and hepatic function and lower INR and APTT values. These patterns reflected a pronounced selection bias inherent in observational studies: only individuals in relatively better health were typically deemed eligible for surgery.
Moreover, surgical patients received lower doses of LMWH or fondaparinux, because a larger proportion were prescribed these agents for VTE prophylaxis rather than therapeutic anticoagulation. Together, these factors partially explained the unexpected inverse association between surgery and bleeding risk.
In our predictive modeling framework, “surgery” thus served as a proxy for a cluster of unmeasured or imperfectly captured health-related attributes, such as functional status, physiological reserve, and frailty, that collectively indicated a lower bleeding propensity in this population. Although this contributed to the model’s strong internal validity and predictive performance, it may have limited external generalizability, particularly in settings with less stringent surgical eligibility criteria. We therefore cautioned against causal interpretation of this association and emphasized that cohort-specific selection biases might affect the model’s transportability to other populations.
Clinical translation and implementation challenges
While our study demonstrated the feasibility of developing a machine learning–based bleeding risk prediction tool for patients receiving LMWH or fondaparinux, several translational considerations had to be addressed before real-world deployment. Notably, the model’s modest discrimination (AUC = 0.659) and particularly low positive predictive performance (PR-AUC < 0.20) in the validation cohort reflected the inherent difficulty of predicting rare adverse events, a challenge frequently encountered in clinical risk modeling when substantial class imbalance was present. These limitations underscored that any clinical implementation needed to prioritize clinical utility over statistical optimality alone. Moreover, prospective external validation across diverse healthcare settings was recognized as essential to assess generalizability, followed by recalibration to local patient demographics and practice patterns48.
Strengths and limitations
The strength of this study was the comprehensive evaluation and comparison of multiple machine learning algorithms for predicting bleeding risk in patients receiving LMWH or fondaparinux. Among the models assessed, CatBoost consistently demonstrated superior performance in terms of discrimination, calibration, and clinical utility. To enhance real-world applicability, we developed an intuitive, internal web-based prediction platform that enables clinicians to rapidly estimate individualized bleeding risk at the point of care, thereby supporting more informed and safer anticoagulation decisions. However, the model’s ability to correctly identify patients who experienced bleeding was limited, as evidenced by a low positive predictive value and a precision-recall AUC below 0.20 in the validation cohort; given its high negative predictive value, it was better suited for ruling out rather than confirming bleeding risk.
Additionally, this study has several limitations. As a single-center retrospective analysis, this study is inherently limited in generalizability and causal inference. The definition of bleeding, based on the proxy assumption that the absence of an outpatient visit within seven days post-discharge indicates no event, may have failed to capture true bleeding episodes that did not lead to healthcare utilization. This potential outcome misclassification could bias the model toward predicting the negative class and may adversely affect sensitivity and negative predictive value. Additional limitations include substantial class imbalance due to the low incidence of bleeding, which constrained model sensitivity and precluded severity-based stratification; high sparsity in concomitant medication data, necessitating aggregation into a count variable and sacrificing clinical granularity; use of default hyperparameters without systematic tuning; residual missingness in key covariates despite exclusion of patients with > 20% missing data; and uncertain model performance in patients with hepatic or renal impairment, given their underrepresentation in the cohort. Importantly, the decision thresholds embedded in the prediction tool have not been empirically validated in clinical practice. Therefore, the predicted probabilities should be interpreted as one component of a holistic clinical assessment rather than as standalone directives for action. Future iterations of the tool may incorporate dynamic variables (e.g., laboratory values obtained after anticoagulation initiation) to support longitudinal monitoring and risk refinement.
Conclusion
In conclusion, this study demonstrated that machine learning models can effectively predict bleeding risk in patients receiving LMWH or fondaparinux, offering a promising approach to support individualized anticoagulation management. The model integrated readily available clinical variables and exhibited acceptable performance in internal validation. To facilitate clinical translation, we have implemented a user-friendly web-based prediction platform. However, this calculator was intended for internal use only and should not be deployed externally without rigorous external validation. Given that the model has not yet undergone external validation and the platform remains unpublished, prospective multicenter validation and implementation studies are essential before broader clinical adoption can be recommended.
Electronic Supplementary Material
Below is the link to the electronic supplementary material.
Author contributions
All authors discussed the results and contributed to the manuscript. Xiang Zheng conceived and designed the study, which was carried out by Haiyan Chen, Qiwei Ran, Fangli Hu, and Xiang Zheng. The same team was responsible for data collection. Xiang Zheng performed the data analysis and interpretation. Haiyan Chen prepared the tables and figures and drafted the manuscript, which was subsequently revised by Xiang Zheng. All authors reviewed and approved the final version for publication.
Funding
This manuscript was funded by the Jinhua Municipal Science and Technology Bureau (grant number 2024-3-106).
Data availability
The datasets, analysis code, and software packages in this study are available from the corresponding author upon reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Ethics approval
This study was performed in accordance with the ethical principles of the Declaration of Helsinki and approved by the Ethics Committee of Dongyang People’s Hospital (Approval No. Dong Ren Yi 2025-YX-096, May 2025). Given its observational and retrospective design, all data were anonymized and de-identified, and informed consent was waived by the Ethics Committee of Dongyang People’s Hospital. All methods were performed in accordance with the relevant guidelines and regulations.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Ceresetto, J. M. et al. Anticoagulantes parenterales. Actualización en uso y monitoreo de la heparina y sus derivados [Parenteral anticoagulants. Update on use and monitoring of heparin and its derivatives]. Med. (B Aires). 85 (Suppl 2), 1–45 (2025). Spanish. [PubMed] [Google Scholar]
- 2.van Rein, N. et al. Major bleeding risks of different low-molecular-weight heparin agents: A cohort study in 12 934 patients treated for acute venous thrombosis. J. Thromb. Haemost.15(7), 1386–1391. 10.1111/jth.13715 (2017). [DOI] [PubMed] [Google Scholar]
- 3.Pisters, R. et al. A novel user-friendly score (HAS-BLED) to assess 1-year risk of major bleeding in patients with atrial fibrillation: the Euro Heart Survey. Chest138 (5), 1093–1100. 10.1378/chest.10-0134 (2010). [DOI] [PubMed] [Google Scholar]
- 4.Hijazi, Z. et al. The ABC (age, biomarkers, clinical history) stroke risk score: A biomarker-based risk score for predicting stroke in atrial fibrillation. Eur. Heart J.37(20), 1582–90. 10.1093/eurheartj/ehw054 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Fang, M. C. et al. A new risk scheme to predict warfarin-associated hemorrhage: The ATRIA (Anticoagulation and Risk Factors in Atrial Fibrillation) Study. J. Am. Coll. Cardiol.58 (4), 395–401. 10.1016/j.jacc.2011.03.031 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Aggarwal, R. et al. Development and validation of the DOAC Score: A novel bleeding risk prediction tool for patients with atrial fibrillation on direct-acting oral anticoagulants. Circulation148(12), 936–946. 10.1161/CIRCULATIONAHA.123.064556 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Fox, K. A. A. et al. Improved risk stratification of patients with atrial fibrillation: An integrated GARFIELD-AF tool for the prediction of mortality, stroke and bleed in patients with and without anticoagulation. BMJ Open7(12), e017157. 10.1136/bmjopen-2017-017157 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Gage, B. F. et al. Clinical classification schemes for predicting hemorrhage: Results from the National Registry of Atrial Fibrillation (NRAF). Am. Heart J.151(3), 713–9. 10.1016/j.ahj.2005.04.017 (2006). [DOI] [PubMed] [Google Scholar]
- 9.O’Brien, E. C. et al. The ORBIT bleeding score: A simple bedside score to assess bleeding risk in atrial fibrillation. Eur. Heart J.36(46), 3258–64. 10.1093/eurheartj/ehv476 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Badescu, M. C. et al. Prediction of bleeding events using the VTE-BLEED risk score in patients with venous thromboembolism receiving anticoagulant therapy (Review). Exp. Ther. Med.22(5), 1344. 10.3892/etm.2021.10779 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ruíz-Giménez, N. et al. Predictive variables for major bleeding events in patients presenting with documented acute venous thromboembolism. findings from the RIETE Registry. Thromb. Haemost.100(1), 26–31. 10.1160/TH08-03-0193 (2008). [DOI] [PubMed] [Google Scholar]
- 12.Decousus, H. et al. Factors at admission associated with bleeding risk in medical patients: findings from the IMPROVE investigators. Chest139 (1), 69–79. 10.1378/chest.09-3081 (2011). [DOI] [PubMed] [Google Scholar]
- 13.Palareti, G. et al. The American College of Chest Physician score to assess the risk of bleeding during anticoagulation in patients with venous thromboembolism. J. Thromb. Haemost.16(10), 1994–2002. 10.1111/jth.14253 (2018). [DOI] [PubMed] [Google Scholar]
- 14.Di Nisio, M. et al. Risk of major bleeding in patients with venous thromboembolism treated with rivaroxaban or with heparin and vitamin K antagonists. Thromb. Haemost.115(2), 424–32. 10.1160/TH15-06-0474 (2016). [DOI] [PubMed] [Google Scholar]
- 15.de Winter, M. A. et al. Recurrent venous thromboembolism and bleeding with extended anticoagulation: The VTE-PREDICT risk score. Eur. Heart J.44(14), 1231–1244. 10.1093/eurheartj/ehac776 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Gao, X. et al. Diagnostic accuracy of the HAS-BLED bleeding score in VKA- or DOAC-treated patients with atrial fibrillation: A systematic review and meta-analysis. Front. Cardiovasc. Med.8, 757087. 10.3389/fcvm.2021.757087 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Caldeira, D. et al. Performance of the HAS-BLED high bleeding-risk category, compared to ATRIA and HEMORR2HAGES in patients with atrial fibrillation: A systematic review and meta-analysis. J. Interv. Card. Electrophysiol.40(3), 277–84. 10.1007/s10840-014-9930-y (2014). [DOI] [PubMed] [Google Scholar]
- 18.Zeng, J. et al. Comparison of HAS-BLED with other risk models for predicting the bleeding risk in anticoagulated patients with atrial fibrillation: A PRISMA-compliant article. Medicine (Baltimore)99(25), e20782. 10.1097/MD.0000000000020782 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Dalgaard, F. et al. GARFIELD-AF model for prediction of stroke and major bleeding in atrial fibrillation: A Danish nationwide validation study. BMJ Open9(11), e033283. 10.1136/bmjopen-2019-033283 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Mori, N. et al. External validation of the ORBIT bleeding score and the HAS-BLED score in nonvalvular atrial fibrillation patients using direct oral anticoagulants (Asian data from the DIRECT registry). Am. J. Cardiol.124(7), 1044–1048. 10.1016/j.amjcard.2019.07.005 (2019). [DOI] [PubMed] [Google Scholar]
- 21.de Winter, M. A. et al. Prediction models for recurrence and bleeding in patients with venous thromboembolism: A systematic review and critical appraisal. Thromb. Res.199, 85–96. 10.1016/j.thromres.2020.12.031 (2021). [DOI] [PubMed] [Google Scholar]
- 22.Van Gelder, I. C. et al. 2024 ESC guidelines for the management of atrial fibrillation developed in collaboration with the European Association for Cardio-Thoracic Surgery (EACTS). Eur. Heart J.45(36), 3314–3414. 10.1093/eurheartj/ehae176 (2024). [DOI] [PubMed] [Google Scholar]
- 23.Rafie, Z. et al. Leveraging XGBoost and explainable AI for accurate prediction of type 2 diabetes. BMC Public Health25(1), 3688. 10.1186/s12889-025-24953-w (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Ghaderzadeh, M., Garavand, A. & Salehnasab, C. Artificial intelligence in polycystic ovary syndrome: A systematic review of diagnostic and predictive applications. BMC Med. Inform. Decis. Mak.25(1), 427. 10.1186/s12911-025-03255-6 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Von Elm, E. et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: Guidelines for reporting observational studies. Ann. Intern. Med.147, 573–577 (2007). [DOI] [PubMed] [Google Scholar]
- 26.Kumar, S. et al. Anticoagulation in Concomitant Chronic Kidney Disease and Atrial Fibrillation: JACC Review Topic of the Week. J. Am. Coll. Cardiol.74 (17), 2204–2215. https://doi.org/10.1016/j.jacc.2019.08.1031 (2019). [DOI] [PubMed] [Google Scholar]
- 27.Stevens, P. E., Levin, A. & Kidney Disease: Improving Global Outcomes Chronic Kidney Disease Guideline Development Work Group Members. Evaluation and management of chronic kidney disease: synopsis of the kidney disease: improving global outcomes 2012 clinical practice guideline. Ann. Intern. Med.158 (11), 825–830. https://doi.org/10.7326/0003-4819-158-11-201306040-00007 (2013). [DOI] [PubMed] [Google Scholar]
- 28.Zahir, H. et al. Effect of mild and moderate hepatic impairment (defined by Child-Pugh Classification and National Cancer Institute Organ Dysfunction Working Group Criteria) on Pexidartinib pharmacokinetics. J. Clin. Pharmacol.62(8), 992–1005. https://doi.org/10.1002/jcph.2042 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Jiang, H. et al. A real-world study on the Morse Fall Scale and clinical judgment method for fall risk in adult inpatients. Int. Nurs. Rev.72(4), e70110. https://doi.org/10.1111/inr.70110(2025). [DOI] [PubMed] [Google Scholar]
- 30.Khoury, T. et al. The complex role of anticoagulation in cirrhosis: An updated review of where we are and where we are going. Digestion93(2), 149–59. 10.1159/000442877 (2016). [DOI] [PubMed] [Google Scholar]
- 31.Bongiovanni, T. et al. Systematic review and meta-analysis of the association between non-steroidal anti-inflammatory drugs and operative bleeding in the perioperative period. J. Am. Coll. Surg.232(5), 765-790.e1. 10.1016/j.jamcollsurg.2021.01.005 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Lu, C. & Zhang, Y. Gastrointestinal bleeding during the transcatheter aortic valve replacement perioperative period: A review. Medicine (Baltimore)101(48), e31953. 10.1097/MD.0000000000031953 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Chen, J., Li, X. & Xiong, X. Severe coagulation dysfunction and active bleeding induced by cefoperazone/sulbactam in a patient with severe renal insufficiency: a case report. Eur. J. Hosp. Pharm. 2025 Apr18:ejhpharm–2025. 10.1136/ejhpharm-2025-004475 [DOI] [PubMed]
- 34.Zhao, H., Chen, J. & Ou, G. A case report of severe drug-induced immune hemolytic anemia caused by piperacillin. Front. Immunol.15, 1478545. 10.3389/fimmu.2024.1478545 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Bui, L. N. & Ding, Q. Logistic regression modeling: Methodological insights and roadmap. Curr. Pharm. Teach. Learn.17(12), 102460. 10.1016/j.cptl.2025.102460 (2025). [DOI] [PubMed] [Google Scholar]
- 36.Peng, Y., Zhang, C. & Zhou, B. A cross-sectional study comparing machine learning and logistic regression techniques for predicting osteoporosis in a group at high risk of cardiovascular disease among old adults. BMC Geriatr.25(1), 209. 10.1186/s12877-025-05840-w (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Yang, X. et al. Deep siamese residual support vector machine with applications to disease prediction. Comput. Biol. Med.196(Pt A), 110693. 10.1016/j.compbiomed.2025.110693 (2025). [DOI] [PubMed] [Google Scholar]
- 38.Huang, Y. et al. Improving the performance of the echinococcosis diagnosis model based on serum Raman spectroscopy via the integration of convolutional neural network and support vector machine. Spectrochim. Acta A Mol. Biomol. Spectrosc.346, 126945. 10.1016/j.saa.2025.126945 (2026). [DOI] [PubMed] [Google Scholar]
- 39.Geng, Z. et al. Development and validation of a machine learning-based predictive model for assessing the 90-day prognostic outcome of patients with spontaneous intracerebral hemorrhage. J. Transl. Med.22(1), 236. 10.1186/s12967-024-04896-3 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Shao, L. et al. Development and external validation of a machine learning-based fall prediction model for nursing home residents: A prospective cohort study. J. Am. Med. Dir. Assoc.25(9), 105169. 10.1016/j.jamda.2024.105169 (2024). [DOI] [PubMed] [Google Scholar]
- 41.Ghaderzadeh, M., Rafie, Z. & Salehnasab, C. Explainable extratreeclassifier model for early detection of type 2 diabetes: Evidence from the PERSIAN Dena Cohort. BMC Med. Inform. Decis. Mak.26(1), 36. 10.1186/s12911-025-03333-9 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zhong, X. et al. Development and validation of a machine learning-based risk prediction model for post-stroke cognitive impairment. Sci. Rep.15(1), 32942. 10.1038/s41598-025-98054-4 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wadhwa, H. et al. Preoperative and postoperative therapeutic anticoagulation in orthopaedic surgery increases the risk of bleeding: A systematic review and meta-analysis. J. Am. Acad. Orthop. Surg.32(24), e1270–e1279. 10.5435/JAAOS-D-24-00161 (2024). [DOI] [PubMed] [Google Scholar]
- 44.Lavikainen, L. I. et al. Risk of thrombosis and bleeding in gynecologic cancer surgery: Systematic review and meta-analysis. Am. J. Obstet. Gynecol.230(4), 403–416. 10.1016/j.ajog.2023.10.006 (2024). [DOI] [PubMed] [Google Scholar]
- 45.Lavikainen, L. I. et al. Risk of thrombosis and bleeding in gynecologic noncancer surgery: Systematic review and meta-analysis. Am. J. Obstet. Gynecol.230(4), 390–402. 10.1016/j.ajog.2023.11.1255 (2024). [DOI] [PubMed] [Google Scholar]
- 46.Avvedimento, M. et al. Bleeding Events After Transcatheter Aortic Valve Replacement: JACC State-of-the-Art Review. J. Am. Coll. Cardiol.81 (7), 684–702. 10.1016/j.jacc.2022.11.050 (2023). [DOI] [PubMed] [Google Scholar]
- 47.Galli, M. & D’Amario, D. High bleeding risk in patients undergoing coronary and structural heart interventions. Interv. Cardiol. Clin.13(4), 483–491. 10.1016/j.iccl.2024.06.003 (2024). [DOI] [PubMed] [Google Scholar]
- 48.Safaie, A. et al. MOCRA: A multi-algorithm clinical decision support system for the early detection of ovarian cancer. J. Ovarian Res.19(1), 33. 10.1186/s13048-025-01929-3 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets, analysis code, and software packages in this study are available from the corresponding author upon reasonable request.








