Abstract
Background
Post-thrombotic syndrome (PTS) is a common and debilitating complication after lower extremity deep vein thrombosis (LEDVT). Early risk assessment of PTS patients is still a clinical challenge.
Methods
The retrospective, multicenter cohort study included 265 patients with unprovoked LEDVT. Baseline clinical and biochemical data including serum uric acid, body mass index (BMI) and treatment delay were noted. Ten machine learning (ML) models (e.g., support vector machine [SVM], XGBoost and LightGBM) were constructed for predicting PTS occurrence in a 70:30 training-validation split. Discrimination of the models was evaluated by F1-score, area under the receiver operating characteristic curve (AUC), calibration curve, and decision curve analysis. Model interpretability was performed through SHapley Additive exPlanations (SHAP).
Results
PTS occurred in 92 patients (34.7%). In multivariate logistic regression, iliofemoral DVT, elevated uric acid, prolonged treatment delay, higher BMI, and lack of statin use were independent predictors of PTS. Among ML models, SVM achieved the best test performance (AUC = 0.985, F1 = 0.926). SHAP analysis identified BMI, treatment delay, and DVT location as top contributors to prediction, while uric acid showed moderate influence. A web-based risk calculator was deployed for clinical use.
Conclusions
The model of machine learning shows a great capacity in prediction of PTS, and had high prediction accuracy, the SVM algorithm predicted better than other algorithms. Addition of uric acid and timing of treatment refines risk stratification. The tool developed might assist clinicians in early identification and risk stratification of patients for individualized care.
Keywords: Post-thrombotic syndrome, Deep vein thrombosis, Machine learning, Uric acid, Risk prediction
Introduction
One of the most frequent and distressing long-term complications of lower extremity deep vein thrombosis (LEDVT) is post-thrombotic syndrome (PTS), which affects up to 50% of patients over the 2 years subsequent to the triggering thrombotic event [1, 2]. In clinical terms, PTS presents as persistent limb pain, swelling, heaviness, hyperpigmentation, and in severe cases it also includes venous ulceration. These complaints may result in chronic physical discomfort, limitations in mobility, and lower quality of life, as well as greater consumption of healthcare resources [3–5]. From an economic perspective, PTS involves significant direct and indirect costs, such as extended use of medication, absence from work, repeated hospital admissions. Although the long-term burden of PTS is well-established, early prediction and prevention of PTS are still challenging [6, 7]. The Villalta scale has been commonly used for the diagnosis but is retrospective and is not designed for prospective risk assessment [8, 9]. There is, therefore, an unmet clinical need for accurate predictive tools to recognize high-risk patients when a DVT is diagnosed.
Conventional clinical prediction models of PTS are frequently based on logistic regression or rule-based scores, including variables, such as thrombus location and extent, and patient characteristics [10]. Although these models are informative, the models are fundamentally constrained by linear assumptions, difficulty in capturing the complexity of interaction among variable and poor possibility to capture the natural world heterogeneity. Their weak discriminator performance out-of-distribution highlights that more flexible and powerful methods are needed. Machine learning (ML) techniques have evolved as a promising tool in medical prediction modelling, which can use high-dimensional data and model non-linear relationships [11–13]. Various ML algorithms have their own strengths: tree-based ensemble methods such as random forest, XGBoost, and gradient boosting machine (GBM) are suitable for well capturing feature interactions and missing data; SVM is good at high-dimensional, limited-sample scenarios; neural network is able to explore potential nonlinearity [14, 15]. Utilising the variety of ML algorithms enables benchmarking of performance and identification of best-fit models for individual clinical tasks. However, few studies have systematically applied ML to PTS prediction, and most fail to incorporate metabolic or inflammatory biomarkers.
Recent artificial intelligence (AI) applications in venous thromboembolism (VTE) have largely focused on short-term outcomes in acute pulmonary embolism (PE) rather than long-term complications after deep vein thrombosis. For example, Cicek et al. developed a deep learning approach to predict short-term mortality in acute PE [16], and another study evaluated the pan-immuno-inflammation value for in-hospital mortality risk stratification in acute PE [17]. These studies highlight the promise of AI/ML for acute prognostication, but they do not address prospective risk stratification for PTS after unprovoked lower extremity deep vein thrombosis.
There is increasing evidence that hyperuricemia which is often observed in metabolic syndrome, obesity, and cardiovascular disease, may be associated with thrombotic and inflammatory outcomes [18, 19]. High levels of serum uric acid are known to induce oxidative stress, endothelial dysfunction and activation of pro-inflammatory programs, such as interleukin-1β and NLRP3 inflammasome signaling [20–22]. These biological actions can result in more severe venous damage and impaired dissolution of thrombus, thus promoting PTS. Moreover, recent epidemiologic studies have demonstrated that hyperuricemia is associated with the increased risk of VTE, especially in East Asian population [23]. Furthermore, inflammatory comorbidities such as gout and rheumatoid arthritis have both been found to be associated with PTS after DVT [24, 25]. Although serum uric acid is a commonly used biomarker, its inclusion in a formal PTS prediction has not yet been incorporated into any prognostic scoring systems. Furthermore, studying uric acid within a machine learning framework may provide new insights into inflammation-driven venous remodeling and the pathogenesis of PTS.
To fill these gaps, the objective of the current study was to develop and internally validate a series of machine learning models to predict PTS in patients with unprovoked LEDVT. The models included traditional clinical covariates and novel risk factors, such as serum uric acid. Ten different ML techniques, including linear and nonlinear classifiers, were employed and evaluated with several assessment metrics. However, interpretability was favoured with the use of SHapley Additive exPlanations (SHAP) that provided insights into feature-level contributions and candidate clinical interactions. In addition, the optimal model was implemented as an online prediction model for bedside risk prediction. Integrating traditional risk factors with metabolic and inflammatory markers, this study also fills an evidence-based structure for the early diagnosis of PTS which will aid in the management of individual patients and in preventive interventions of PTS.
Methods
Study design and participants
This is a retrospective cohort study of patients with their first episode of LEDVT who were hospitalized between January 2016 and January 2023 at Fuxing Hospital Affiliated to Capital Medical University and Xuanwu Hospital Affiliated to Capital Medical University. All the cases were diagnosed by duplex ultrasonography or contrast venography non-compression and collapse of affected vein. Patients were recruited consecutively and were followed for PTS development.
Inclusion criteria participants were included in this study if they met the following criteria:
Age ≥ 18 years;
First-ever episode of LEDVT;
Unprovoked aetiology (i.e., not associated transient risk factors, such as recent surgery, trauma and pregnancy); and.
Serum uric acid levels measured within 48 h of hospital admission.
Exclusion criteria was as below:
Active cancer, or expected survival < 1 year;
Known post-thrombotic syndrome (PTS) or chronic venous disease (CVD);
Active infection, autoimmune disease or other inflammatory conditions that are known to change systemic inflammation markers or uric acid metabolism;
A history of receiving drugs for lowering serum uric acid (e.g., allopurinol, febuxostat); and.
Inability to attend follow-up due to reasons, such as severe cognitive impairment, loss to follow-up, or living outside the catchment area for long periods.
The study was performed in accordance with the Declaration of Helsinki, and approved by the Ethics Committee of Fu Xing Hospital, Capital Medical University (Ethics Number: 2025FXHEC-KSP043). Due to the retrospective nature of the study and use of anonymized data, the requirement for informed consent was waived.
Outcome definition and variable collection
The main response of the study was PTS during follow-up, which was evaluated according to the Villalta scale, a validated clinical scoring system recommended by international guidelines. Patients who had Villalta scores at follow-up were classified into PTS and non-PTS groups.
Demographic, clinical and laboratory characteristics at baseline were recorded from electronic health records. 19 candidate predictors were analyzed in total:
Demographic and behavior: sex (male/female), age (60–75/ ≥ 75), smoking (yes/no), drink (yes/no). Comorbid conditions: hypertension, diabetes mellitus, coronary heart disease (CHD), chronic kidney disease (CKD), and hyperlipidemia—all entered in the model as binary variables (yes/no). Characteristics of thrombus: DVT_Location (femoral vein thrombosis or iliofemoral vein thrombosis). Treatment factors: compression hosiery (yes/no). Type of anticoagulant: low molecular weight heparin (LMWH), direct oral anticoagulants (DOACs), or vitamin K antagonists (VKAs). Duration of anticoagulation: less than 3 months, 3–6 months, or more than 6 months. Statin use (yes/no). Antiplatelet therapy (yes/no). Thrombectomy (yes/no). Clinical and laboratory variables: Uric_Acid (µmol/L); Body Mass Index (BMI, kg/m2); and Treatment_Delay (days between symptom onset and start of anticoagulation).
Categorical variables were transformed into binary and ordinal features when applicable. All numerical variables were kept as continuous features and were standardized when building our models. All data were dual checked by two investigators to guarantee the accuracy and consistency of the data entry. Before modeling, variables were checked for plausibility and standardized, and records with missing values in the predictors required for a given analysis were excluded.
Data set splitting and statistical analysis
All eligible patients were randomly divided into a training set (70%) and test set (30%) with stratified sampling according to PTS status to keep balance of outcome. The training data set was divided into the model development data set and the internal validation data set, the training data set was used to generate models and select features, and the test data set was used to evaluate the model fitting and predictive performance of machine learning.
Descriptive statistics for summary of baseline characteristics of the whole studied population are reported. Normally distributed continuous variables were expressed as mean ± SD, while categorical variables was expressed as counts and percentages. Between-group comparisons of continuous variables were conducted by the independent samples t test, and of categorical variables by chi-square or Fisher’s exact test as appropriate. Results A p value < 0.05 was considered as significant.
Univariate logistic regression analysis was initially used to screen the potential factors associated with PTS in the training set. Further analysis was conducted using a multivariate logistic regression model to discover independent predictors (p < 0.05), in which variables with p values of < 0.05 in univariate analysis were entered. The odds ratios (ORs) with 95% confidence intervals (95% CIs) were calculated and reported.
All analyses were carried out using R software (version 4.5.0).
Machine learning model development and evaluation
We developed a robust prediction model of post-thrombotic syndrome using ten supervised machine learning algorithms in our training data set. Those machine learning algorithms were logistic regression, support vector machine, k-nearest neighbors, random forest, gradient boosting machine, and extreme gradient boosting, categorical boosting, Adaboost, and multilayer perceptron neural network. We performed tenfold cross-validation during the training data sets. For the SVM, we used a non-linear radial basis function (RBF) kernel, with kernel width and regularization tuned as sigma = 0.001 and C = 0.09. To enhance transparency across models and facilitate assessment of overfitting risk, we report the primary hyperparameter settings used for each algorithm: GBM (n.trees = 100, interaction.depth = 3, shrinkage = 0.1, n.minobsinnode = 5); neural network (size = 6, decay = 0.6); random forest (mtry = 11, numRandomCuts = 3); XGBoost (nrounds = 10, max_depth = 3, eta = 0.001, gamma = 0.5, colsample_bytree = 0.5, min_child_weight = 1, subsample = 0.6); k-nearest neighbors (KNN) (kmax = 12, distance = 1, kernel = “optimal”); and Adaboost (mfinal = 2, maxdepth = 2, coeflearn = “Zhu”). Logistic regression used default settings (glm.tune.grid = NULL). The implementation was conducted built on R version 4.5.0 using packages, including caret, xgboost, lightgbm, and e1071. Performance of each model was in both training and test data sets. The model’s discrimination ability was calculated as the area under the receiver operating characteristic curve. In addition, the overall accuracy, sensitivity, specificity, precision, and F1-score was performed. We compared the discrimination metric and calibration curves, which represent the agreement between probabilities predicted and the observed probability. Clinical utility was analyzed utilizing the decision curve analysis, which assessed the net benefit across models. We visualized the test data set for all methods to make confusion matrices with all possible texts to be classified.
Model interpretation using SHAP analysis
For more transparency and interpretability, we also employed SHapley Additive exPlanations (SHAP) on the best models to investigate feature contributions from the global and per-sample perspectives. SHAP values were computed using the iml, shapr, and DALEX packages in R (version 4.5.0).
Global feature importance and direction in contribution to the model output on the entire data set were visualized with SHAP summary plots. Beeswarm plots also showed the distribution and the value of each single SHAP value corresponding to different features. For a few specific cases, we produced force plots and waterfall plots that demonstrate how different sets of features tilted the prediction either toward or away from the PTS class. Dependency plots were used to examine possible variable interactions. Each SHAP plot produced a mechanistic description of the contribution of individual clinical features to the risk estimate, which facilitated both clinical interpretability and translating machine-learning predictions into practice.
Web-based model deployment
To support use of the machine learning models in the current study for clinical predictions in real time, an interactive web-based prediction tool was built through Shiny framework in R version 4.5.0. The application was designed to allow users to input clinical variables that are relevant and identified during the model development process to obtain a predicted risk of PTS based on the trained models.
The interface was built to minimize complication and mimic clinical practice. It receives user entered patient data and produces a prompt, automated estimation of the risk based on an internal calculation of the machine learning algorithm that was used. This back-end includes the trained model architecture, preprocessing steps and classification thresholds.
Results
Study population baseline characteristics
Overall, 265 patients with unprovoked LEDVT were enrolled, 92 (34.7%) with PTS occurring during follow-up and 173 (65.3%) without PTS. Characteristics Baseline demographic and clinical characteristics by PTS status are summarized in Table 1.
Table 1.
Characteristics of included patients
| Demographic characteristics | levels | No post-thrombotic syndrome (N = 173) | Post-thrombotic syndrome (N = 92) | p value |
|---|---|---|---|---|
| Gender | Female | 87 (50.3%) | 47 (51.1%) | 1 |
| Male | 86 (49.7%) | 45 (48.9%) | ||
| Alcohol | No | 107 (61.8%) | 42 (45.7%) | 0.016 |
| Yes | 66 (38.2%) | 50 (54.3%) | ||
| Smoking | No | 104 (60.1%) | 36 (39.1%) | 0.002 |
| Yes | 69 (39.9%) | 56 (60.9%) | ||
| Stocking | No | 115 (66.5%) | 75 (81.5%) | 0.014 |
| Yes | 58 (33.5%) | 17 (18.5%) | ||
| DVT_Location | Femoral_vein_thrombosis | 146 (84.4%) | 43 (46.7%) | < 0.001 |
| Iliofemoral_vein_thrombosis | 27 (15.6%) | 49 (53.3%) | ||
| Hypertension | No | 93 (53.8%) | 55 (59.8%) | 0.418 |
| Yes | 80 (46.2%) | 37 (40.2%) | ||
| Diabetes | No | 105 (60.7%) | 59 (64.1%) | 0.678 |
| Yes | 68 (39.3%) | 33 (35.9%) | ||
| CHD | No | 112 (64.7%) | 53 (57.6%) | 0.314 |
| Yes | 61 (35.3%) | 39 (42.4%) | ||
| CKD | No | 148 (85.5%) | 59 (64.1%) | < 0.001 |
| Yes | 25 (14.5%) | 33 (35.9%) | ||
| Anticoag_Duration | < 3 m | 29 (16.8%) | 34 (37%) | < 0.001 |
| > 6 m | 115 (66.5%) | 35 (38%) | ||
| 3–6 m | 29 (16.8%) | 23 (25%) | ||
| Thrombectomy | No | 119 (68.8%) | 81 (88%) | < 0.001 |
| Yes | 54 (31.2%) | 11 (12%) | ||
| Hyperlipidemia | No | 86 (49.7%) | 51 (55.4%) | 0.448 |
| Yes | 87 (50.3%) | 41 (44.6%) | ||
| Statin | No | 61 (35.3%) | 48 (52.2%) | 0.011 |
| Yes | 112 (64.7%) | 44 (47.8%) | ||
| Antiplatelet | No | 110 (63.6%) | 55 (59.8%) | 0.635 |
| Yes | 63 (36.4%) | 37 (40.2%) | ||
| Anticoagulant_Type | DOACs | 74 (42.8%) | 41 (44.6%) | 0.027 |
| LMWH | 13 (7.5%) | 16 (17.4%) | ||
| VKA | 86 (49.7%) | 35 (38%) | ||
| Treatment_Delay | Mean ± SD | 5.87 ± 2.90 | 10.35 ± 3.72 | < 0.001 |
| Age | Mean ± SD | 61.57 ± 5.39 | 63.14 ± 5.80 | 0.029 |
| Uric_Acid | Mean ± SD | 376.80 ± 46.19 | 423.88 ± 39.81 | < 0.001 |
| BMI | Mean ± SD | 22.40 ± 1.94 | 25.24 ± 1.81 | < 0.001 |
There was no between-group difference with respect to sex (p = 1.000). However, there was substantial variation in clinical and behavioral characteristics. Patients who had developed PTS were more likely to report a history of alcohol use (54.3% vs. 38.2%, p = 0.016), smoking (60.9% vs. 39.9%, p = 0.002), and less likely to have history of use of compression stockings (18.5% vs. 33.5%, p = 0.014). Iliofemoral vein thrombosis was significantly more frequent in PTS than in femoral-only thrombosis (53.3% vs. 15.6%, p < 0.001).
A higher percentage of CKD (35.9% vs. 14.5%, p < 0.001) and shorter duration of anticoagulant treatment (< 3 months) were also common in the PTS group (37.0% vs. 16.8%, p < 0.001). There are fewer patients underwent thrombectomy in PTS group (12.0% vs. 31.2%, p < 0.001).
With respect to anticoagulant agents, LMWH was used more commonly in PTS patients than in non-PTS patients (17.4% vs. 7.5%, p = 0.027), whereas VKA was used more commonly in non-PTS patients than in PTS patients (49.7% vs. 38.0%). There was no statistical difference in anti-platelet drugs and statins applied between the two groups and in the percentage of hypertension, diabetes, coronary heart disease and hyperlipidemia.
Patients with PTS had significantly higher levels of serum uric acid (423.88 ± 39.81 µmol/L vs. 376.80 ± 46.19 µmol/L, p < 0.001), longer delay in treatment to anticoagulation (10.35 ± 3.72 vs. 5.87 ± 2.90 days, p < 0.001), older age (63.14 ± 5.80 vs. 61.57 ± 5.39 years, p = 0.029), and higher BMI (25.24 ± 1.81 vs. 22.40 ± 1.94 kg/m2, p < 0.001).
Logistic regression analysis for predicting PTS
To identify independent predictors of PTS, both univariate and multivariate logistic regression analyses were performed using the training cohort of 187 patients (122 without PTS, 65 with PTS), as shown in Table 2.
Table 2.
Training set univariate and multivariate logistics regression results
| Demographic characteristics | levels | No post-thrombotic syndrome (N = 122) | Post-thrombotic syndrome (N = 65) | OR (univariable) | OR (multivariable) |
|---|---|---|---|---|---|
| Gender | Female | 60 (49.2%) | 33 (50.8%) | ||
| Male | 62 (50.8%) | 32 (49.2%) | 0.94 (0.51–1.71, p = 0.836) | ||
| Alcohol | No | 74 (60.7%) | 27 (41.5%) | ||
| Yes | 48 (39.3%) | 38 (58.5%) | 2.17 (1.18–4.00, p = 0.013) | 2.41 (0.74–7.90, p = 0.145) | |
| Smoking | No | 73 (59.8%) | 25 (38.5%) | ||
| Yes | 49 (40.2%) | 40 (61.5%) | 2.38 (1.29–4.42, p = 0.006) | 1.96 (0.59–6.54, p = 0.272) | |
| Stocking | No | 84 (68.9%) | 52 (80%) | ||
| Yes | 38 (31.1%) | 13 (20%) | 0.55 (0.27–1.13, p = 0.106) | ||
| DVT_Location | Femoral_vein_thrombosis | 102 (83.6%) | 30 (46.2%) | ||
| Iliofemoral_vein_thrombosis | 20 (16.4%) | 35 (53.8%) | 5.95 (3.00–11.79, p < 0.001) | 10.02 (2.50–40.16, p = 0.001) | |
| Hypertension | No | 61 (50%) | 39 (60%) | ||
| Yes | 61 (50%) | 26 (40%) | 0.67 (0.36–1.23, p = .193) | ||
| Diabetes | No | 72 (59%) | 45 (69.2%) | ||
| Yes | 50 (41%) | 20 (30.8%) | 0.64 (0.34–1.21, p = 0.171) | ||
| CHD | No | 79 (64.8%) | 39 (60%) | ||
| Yes | 43 (35.2%) | 26 (40%) | 1.22 (0.66–2.28, p = 0.521) | ||
| CKD | No | 104 (85.2%) | 41 (63.1%) | ||
| Yes | 18 (14.8%) | 24 (36.9%) | 3.38 (1.66–6.88, p < 0.001) | 2.60 (0.56–12.05, p = 0.222) | |
| Anticoag_Duration | < 3 m | 18 (14.8%) | 20 (30.8%) | ||
| 3–6 m | 24 (19.7%) | 18 (27.7%) | 0.68 (0.28–1.63, p = 0.383) | 1.90 (0.31–11.77, p = 0.491) | |
| > 6 m | 80 (65.6%) | 27 (41.5%) | 0.30 (0.14–0.66, p = 0.002) | 0.21 (0.04–1.04, p = .055) | |
| Thrombectomy | No | 85 (69.7%) | 57 (87.7%) | ||
| Yes | 37 (30.3%) | 8 (12.3%) | 0.32 (0.14–0.74, p = .008) | 0.31 (0.07–1.42, p = 0.133) | |
| Hyperlipidemia | No | 64 (52.5%) | 34 (52.3%) | ||
| Yes | 58 (47.5%) | 31 (47.7%) | 1.01 (0.55–1.84, p = 0.984) | ||
| Statin | No | 45 (36.9%) | 35 (53.8%) | ||
| Yes | 77 (63.1%) | 30 (46.2%) | 0.50 (0.27–0.92, p = .026) | 0.27 (0.08–0.95, p = 0.042) | |
| Antiplatelet | No | 75 (61.5%) | 39 (60%) | ||
| Yes | 47 (38.5%) | 26 (40%) | 1.06 (0.57–1.97, p = 0.844) | ||
| Anticoagulant_Type | LMWH | 9 (7.4%) | 12 (18.5%) | ||
| DOACs | 57 (46.7%) | 29 (44.6%) | 0.38 (0.14–1.01, p = 0.052) | 0.34 (0.05–2.41, p = 0.281) | |
| VKA | 56 (45.9%) | 24 (36.9%) | 0.32 (0.12–0.86, p = 0.024) | 0.23 (0.03–1.70, p = 0.150) | |
| Treatment_Delay | Mean ± SD | 6.0 ± 2.9 | 10.3 ± 3.6 | 1.45 (1.29–1.62, p < 0.001) | 1.60 (1.30–1.98, p < 0.001) |
| Age | Mean ± SD | 61.3 ± 5.2 | 62.5 ± 6.2 | 1.04 (0.98–1.10, p = 0.163) | |
| Uric_Acid | Mean ± SD | 377.5 ± 45.0 | 420.9 ± 39.2 | 1.03 (1.02–1.04, p < 0.001) | 1.02 (1.00–1.03, p = 0.029) |
| BMI | Mean ± SD | 22.5 ± 2.0 | 25.2 ± 1.9 | 2.07 (1.66–2.57, p < 0.001) | 2.11 (1.46–3.07, p < 0.001) |
In univariate analysis, several variables were significantly associated with the development of PTS, including alcohol use (OR = 2.17, 95% CI 1.18–4.00, p = 0.013), smoking (OR = 2.38, p = 0.006), iliofemoral vein thrombosis (OR = 5.95, p < 0.001), CKD (OR = 3.38, p < 0.001), shorter anticoagulation duration (< 3 months), absence of thrombectomy (OR = 0.32, p = 0.008), statin nonuse (OR = 0.50, p = 0.026), LMWH use, VKA use (OR = 0.32, p = 0.024), longer treatment delay (OR per day = 1.45, p < 0.001), higher BMI (OR = 2.07 per unit increase, p < 0.001), and serum uric acid level (OR = 1.03 per µmol/L, p < 0.001).
In the multivariate model, five variables remained independently associated with PTS:
Iliofemoral vein thrombosis (adjusted OR = 10.02, 95% CI 2.50–40.16, p = 0.001); Treatment delay (OR = 1.60 per day, 95% CI 1.30–1.98, p < 0.001); BMI (OR = 2.11 per unit, 95% CI 1.46–3.07, p < 0.001); Elevated uric acid (OR = 1.02 per µmol/L, 95% CI 1.00–1.03, p = 0.029); Statin use (protective factor, OR = 0.27, 95% CI 0.08–0.95, p = 0.042).
Other variables such as smoking, CKD, thrombectomy, and anticoagulant type lost statistical significance in the adjusted model.
Performance of machine learning models
We developed and tested ten machine-learning algorithms for the prediction of post-thrombotic syndrome (PTS). The results are presented in Tables 3 (training) and 4 (test) summarizing evaluation metrics, and in Figs. 1 and 2 showing model performance curves and confusion matrices, respectively.
Table 3.
Training set evaluation metrics
| Model | Threshold | Accuracy | Sensitivity | Specificity | Precision | F1 |
|---|---|---|---|---|---|---|
| Logistic | 0.36174677821296 | 0.888 | 0.862 | 0.902 | 0.824 | 0.842 |
| SVM | 0.635896813922154 | 0.877 | 0.846 | 0.893 | 0.809 | 0.827 |
| GBM | 0.419324802700434 | 0.979 | 0.969 | 0.984 | 0.969 | 0.969 |
| NeuralNetwork | 0.381582312654666 | 0.818 | 0.815 | 0.82 | 0.707 | 0.757 |
| RandomForest | 0.5 | 1 | 1 | 1 | 1 | 1 |
| XGboost | 0.49836054444313 | 0.834 | 0.969 | 0.762 | 0.685 | 0.803 |
| KNN | 0.271509508442251 | 0.904 | 1 | 0.852 | 0.783 | 0.878 |
| Adaboost | 0.186531342943985 | 0.802 | 0.892 | 0.754 | 0.659 | 0.758 |
| LightGBM | 0.50155309902651 | 1 | 1 | 1 | 1 | 1 |
| CatBoost | 0.585665096707379 | 0.85 | 0.877 | 0.836 | 0.74 | 0.803 |
Fig. 1.
Discrimination, calibration, and clinical utility of machine learning models in the training and test sets. A–C Receiver operating characteristic (ROC) curves, calibration plots, and decision curve analyses (DCA) for the training set. D–F Corresponding evaluation curves for the test set
Fig. 2.
Confusion matrices of the ten machine learning models in the test set. A–J Adaboost, CatBoost, GBM, KNN, LightGBM, logistic regression, neural network, random forest, SVM, and XGBoost, respectively
Training set performance
Near-perfect classification models were created in the training cohort. The Random Forest and LightGBM recorded perfect performance scores across the multivariable statistics (accuracy, sensitivity, specificity, precision, F1-score = 1.000). Gradient Boosting Machine (GBM) was the second best (F1 = 0.969) followed by K-nearest neighbors (KNN) (F1 = 0.878) and XGBoost (F1 = 0.803). The F1 scores of more complex models (i.e., logistic regression and SVM) was reasonably good (0.842, 0.827) but they slightly had lower training accuracy compared to Ensembles.
Test set performance
On the test set, Support Vector Machine (SVM) performed best overall with the highest F1-score (0.926) and substantial sensitivity (0.926) and specificity (0.961). GBM and CatBoost appeared closely in the next positions (F1 = 0.912 and 0.877), which is a sign that they generalize well. Although Random Forest and LightGBM obtained perfect training performance, the F1 scores decreased to 0.839 and 0.857 on the test set, demonstrating overfitting. XGBoost showed strong generalization (F1 = 0.877), confirming its reputation of being well-balanced.
ROC, calibration, and DCA
Receiver operating characteristic (ROC) curves illustrated that in the test set, SVM had the highest AUC (AUC = 0.985, 95% CI 0.967–1.000), followed by GBM (0.973), KNN (0.963) and XGBoost (0.960) (Fig. 1A–F). The calibration plots showed close calibration among the predicted and observed risks for SVM, logistic regression, and GBM. The decision curve analyses (DCA) curves showed that the GBM, LightGBM, and Random Forest models presented the maximum net clinical benefit in the training set, and the GBM and XGBoost models had better predictive acuity in the test set than other models. SVM also performed well and consistently better on both data sets.
Confusion matrix analysis
The confusion matrices of the two models are shown in Fig. 2A–J on the test set. The SVM (Fig. 2I) model had an outstanding discrimination capability with a low false positive and false negative values, indicating high sensitivity and specificity of the model. XGBoost (Fig. 2J) and GBM (Fig. 2C) also showed balanced classification, low misclassification. Adaboost (Fig. 2A) and neural network (Fig. 2G), in contradiction, had relatively higher relative error in terms of misclassification, mainly in non-PTS cases, leading to their relatively low precision.
Feature importance across machine learning models
The top predictors of PTS across the ten machine learning models are illustrated in Fig. 3. Although the relative ranking of variables varied by algorithm, several features consistently emerged as dominant predictors.
Fig. 3.
Feature importance rankings across 10 machine learning models for predicting post-thrombotic syndrome
Notably, Uric Acid appeared among the top 3 most important variables in several models, including SVM, XGBoost, GBM, and random forest, confirming its robust and consistent predictive power. This finding aligns with logistic regression analysis, where elevated serum uric acid was an independent risk factor for PTS.
BMI and Treatment Delay were also consistently ranked highly across all ensemble-based models (e.g., LightGBM, CatBoost), suggesting that metabolic status and therapeutic timeliness are key determinants in PTS development. The location of thrombosis (iliofemoral) further reinforced its role as a high-impact anatomical variable.
SHAP-based interpretability of machine learning models
To better understand how individual features contributed to model predictions, SHapley Additive exPlanations (SHAP) were applied to the best-performing model (SVM) and a representative tree-based model (XGBoost), as visualized in Figs. 4 and 5, respectively.
Fig. 4.
SHAP interpretability visualizations for the support vector machine (SVM) model
Fig. 5.
SHAP interpretability visualizations for the XGBoost model
In the SVM model (Fig. 4), SHAP summary and beeswarm plots indicated that Body Mass Index (BMI), Treatment Delay, and DVT Location were the top contributors to PTS prediction. Although Uric Acid was not the most important variable, it still exhibited meaningful influence in certain individuals. Notably, the SHAP dependence plot revealed an interaction between Uric Acid and Treatment Delay, suggesting that the effect of uric acid on PTS risk may be modulated by how soon treatment is initiated. The waterfall and force plots illustrated how a combination of elevated BMI, delayed treatment, and in some cases, increased uric acid collectively raised predicted PTS risk above decision thresholds.
In the XGBoost model (Fig. 5), similar patterns were observed. Treatment delay, uric acid, and BMI ranked highest in SHAP importance. While uric acid was not among the leading predictors globally, it appeared in several individual-level explanations with positive SHAP contributions, reinforcing its potential relevance in specific clinical profiles. The interaction between uric acid and treatment delay was also seen in XGBoost’s dependence plot, further supporting this nuanced relationship.
Deployment of a web-based clinical prediction tool
To facilitate clinical translation and real-time risk assessment, the final Support Vector Machine (SVM) model was deployed as an interactive Shiny web application, publicly accessible at: https://cacsriskmodel.shinyapps.io/make_web/.
As illustrated in Fig. 6, the web interface enables users to input five key clinical variables derived from the model: Uric_Acid (serum uric acid level, µmol/L), BMI (body mass index, kg/m2), Treatment_Delay (number of days from symptom onset to anticoagulation), DVT_Location (femoral vs. iliofemoral vein thrombosis), and Oral statins (yes/no). Once entered, the tool returns a predicted probability of PTS, along with a binary classification result.
Fig. 6.
Screenshot of the deployed Shiny web application for predicting post-thrombotic syndrome
This web-based tool allows clinicians to rapidly estimate PTS risk at the bedside or in outpatient settings, enabling early identification of high-risk patients and individualized management planning.
Discussion
We constructed and compared ten supervised machine learning models for predicting PTS in patients with unprovoked LEDVT. The external validity of the final model both in training and test data sets was evident by the good discrimination and calibration of the model. Applying a multivariable combination of demographic, clinical and laboratory features, as well as metabolic and treatment variables, we demonstrated excellent predictive value of our strategy with a better performance compared to generic regression analysis in the identification of patients at risk of developing PTS.
Multivariable logistic regression analysis indicated that the location of thrombosis in proximal veins, delayed treatment, higher BMI and higher serum uric acid level were independent risk factors for PTS. These observations were reinforced in the machine learning models which identified the same variables as having the strongest predictive impact. Importantly, SHAP-based interpretability suggests complex interaction links to underlying pathophysiological dependencies beyond their additive influences.
Finally, for further clinical translation, the predictive model was implemented as a web-based application, allowing individual risk estimation via a user-interface with user-friendly access. Taken together, these findings highlight the power of machine learning for the early diagnosis of high-risk patients and for assisting in the development of targeted interventions that can prevent the long-term complications of LEDVT.
Some independent predictors for PTS in this study, such as proximal DVT location, higher BMI, and delayed treatment, are consistent with the literature from other studies [26–28]. As an example, iliofemoral vein thrombosis is often related to a higher PTS hazard risk because of severe damage to venous valves and impaired outflow from proximal vessel segments [27, 28]. The above is also exacerbated with increasing BMI which promotes venous stasis and systemic inflammation and risk of chronic venous insufficiency [26]. In the same line, a delay in the start of anticoagulation therapy is associated with a persistent thrombus burden, ongoing inflammation, and poor endothelial repair that favor the onset of PTS [27]. One possible interpretation is that hyperuricemia may amplify the adverse effect of delayed anticoagulation on thrombus persistence and vein-wall inflammation/remodeling, which could explain the observed uric acid–treatment–delay interaction in the SHAP dependence plot.
Another important variable we investigated was statin therapy that seemed to be a protective factor for PTS. In addition to the lipid-lowering effect, statins have anti-inflammatory, endothelial-stabilizing, and antithrombotic properties. Previous studies have demonstrated that statins could decrease thrombus burden, suppress vein wall remodeling, and stimulate endothelial healing [29–31]. Providing a possible explanation for the protective effects against PTS. Nevertheless, the substantiation in this field is sparse and heterogeneous and should be further confirmed in prospective interventional trials.
Furthermore, our findings underscore the importance of serum uric acid as a less-studied but emerging relevant biomarker. Historically considered a byproduct of gout, uric acid is understood to be a signaler of systemic inflammation. A recent East Asian cohort and Mendelian randomization study found that genetically determined higher uric acid levels was strongly associated with higher venous thromboembolism risk [23]. Mechanistic investigations have also demonstrated that urate can stimulate NLRP3 inflammasome and endothelial dysfunction, which might lead to acute thrombus formation and chronic venous damage [24]. In addition, a prospective cohort study conducted by Iding et al. found that chronic inflammatory diseases, many overlapping the condition of hyperuricemia, were accompanied by a significant increase in the incidence of PTS, especially in patients with residual vein obstruction [25].
Collectively, these results highlight that the interaction between thrombotic burden, metabolic state, systemic inflammation and vascular repair is complex in PTS pathogenesis. The addition of uric acid and statin use into predictive models may enhance risk stratification and help identify modifiable secondary prevention factors.
The implications of these results are significant for early identification of at-risk patients and may inform future clinical management. A predictive tool with good discriminatory power was designed by incorporating routine clinically and laboratory available features in machine learning architecture. Unlike traditional regression-based models that often rely on linear assumptions and limited interactions, machine learning approaches are capable of capturing complex, nonlinear relationships between predictors, thereby enhancing predictive performance and individual-level risk estimation.
From a clinical perspective the possibility of stratifying patients according to their PTS risk at the time of DVT diagnosis would be able to influence several management decisions. High-risk individuals may be candidates for closer surveillance in follow-up, routine referral to a vascular specialist, or adjunctive intervention strategies, including increased adherence to vascular risk factor control or the addition of more aggressive strategies (e.g., greater use of compression, longer term anticoagulation, behavior modification to target modifiable risk factors, such as obesity and hyperuricemia-sensitive gout). The protective association of statin use observed may also suggest a potential protection of using these agents in high-risk patients, which needs to be confirmed with prospective trials.
Critically, the final model we presented here, through web-based application, can provide easy and real-time decision support, potentially conveying to clinical practice. This is consistent with current trends to practice precision medicine in the care of patients with venous thromboembolism, where risk-stratification scoring systems have been integrated into clinical processes to help decide on management for a given patient. Highlighting metabolic and inflammatory contributors to PTS rather than just anatomical or procedural risk factors, our work offers a roadmap for predicting and ultimately preventing this debilitating long-term complication.
Several methodological strengths of this work increase the trustworthiness and applicability of the results. We first directly compared predictive ability across our ten machine learning algorithms in an exhaustive comparative analysis to discern the most generalizable and clinically applicable model for final deployment. By employing simple models (e.g., logistic regression) as well as advanced ensemble-based methods (e.g., XGBoost, CatBoost, and LightGBM), the analysis was less likely to be skewed toward one methodological paradigm, thereby enhancing its robustness [32].
Second, the model was ensured that training was explicitly separated from testing with a 7:3 split and that the performance was assessed by various standard criteria, including sensitivity, specificity, F1-score, and AUC, along with calibration and decision curve analyses. This comprehensive validation approach enhances the model credibility, not only in terms of discrimination ability, but also in regard to clinical usefulness and prediction-consistency relation.
SHAP also facilitated interpretability of the prediction models as the contributions of features could be interpreted both globally and for individual instances. This method allowed us to uncover subtle relationships (for example, between uric acid and treatment delay) that would probably be lost with more classical linear-model approaches [33, 34]. The ability to interpret individual predictions facilitates possible inclusion into clinical decision support, where decision credibility (cf. physician adherence to algorithm) is crucial.
In our benchmarking, the SVM achieved the best generalization on the independent test set, whereas several tree-based ensemble models showed near-perfect training performance but a noticeable drop in test performance, consistent with overfitting. This pattern is plausible in a moderate-sized, standardized tabular data set such as ours: after feature scaling, class separation may become more “smooth” and better captured by a regularized kernel method. Notably, our radial basis function SVM imposes a controlled bias–variance trade-off by prioritizing a maximal-margin solution with regularization, which can be advantageous when the effective signal is low-dimensional relative to model flexibility. In contrast, high-capacity tree ensembles can fit idiosyncratic threshold-based splits that do not replicate well in held-out data unless the sample size is sufficiently large and/or strong regularization is applied.
Recent cardiovascular and venous thromboembolism artificial intelligence studies have increasingly emphasized two translational pillars: improving predictive performance beyond conventional risk scores and providing clinically trustworthy explanations. For example, deep learning has been applied to short-term risk prediction in acute pulmonary embolism [16], and inflammation-integrated indices have been evaluated for in-hospital mortality stratification [17]. Although these studies address acute outcomes rather than long-term PTS, they highlight a common requirement for clinical uptake: interpretability that can justify why a specific patient is classified as high risk. In our work, we prioritized SHAP to provide both global explanations and patient-level explanations, enabling clinicians to see how metabolic burden and care-process variables jointly shape risk estimates. As an alternative, local interpretable model-agnostic explanations (LIME; a local surrogate-model approach that explains an individual prediction) could be explored in future work as a complementary perspective; however, SHAP was preferred here, because it yields consistent additive attributions across patients and facilitates visualization of feature interactions relevant to PTS risk (Table 4).
Table 4.
Test set evaluation metrics
| Model | Threshold | Accuracy | Sensitivity | Specificity | Precision | F1 |
|---|---|---|---|---|---|---|
| Logistic | 0.102041790672602 | 0.897 | 1 | 0.843 | 0.771 | 0.871 |
| SVM | 0.723373657378884 | 0.949 | 0.926 | 0.961 | 0.926 | 0.926 |
| GBM | 0.264862480728418 | 0.936 | 0.963 | 0.922 | 0.867 | 0.912 |
| NeuralNetwork | 0.33942072721623 | 0.872 | 0.852 | 0.882 | 0.793 | 0.821 |
| RandomForest | 0.222 | 0.872 | 0.963 | 0.824 | 0.743 | 0.839 |
| XGboost | 0.499032497406006 | 0.91 | 0.926 | 0.902 | 0.833 | 0.877 |
| KNN | 0.522654439315569 | 0.923 | 0.852 | 0.961 | 0.92 | 0.885 |
| Adaboost | 0.186531342943985 | 0.833 | 0.926 | 0.784 | 0.694 | 0.794 |
| LightGBM | 0.0913931406508006 | 0.897 | 0.889 | 0.902 | 0.828 | 0.857 |
| CatBoost | 0.591590292097385 | 0.91 | 0.926 | 0.902 | 0.833 | 0.877 |
Our modeling framework has great potential to serve as a bridge between statistical rigor and clinical utility by integrating predictability, interpretability, and operational deployment. These methodological characteristics make the tool a potential candidate for incorporation in the management algorithm of DVT patients, provided that external validation is successfully completed.
There are, however, some limitations of this study that need to be taken into account. First, despite the two-center design, the sample size (n = 265) is modest and may limit statistical power and generalizability; therefore, external validation in larger multicenter cohorts is warranted. Because our primary goal was to develop a bedside-deployable model and to minimize overfitting, we used a parsimonious predictor set for ML training; future multicenter studies should evaluate broader feature engineering and embedded feature selection/regularization strategies to capture potentially informative non-linear effects and interactions.
Second, thrombus extent, and residual vein obstruction (RVO; persistent venous obstruction on follow-up imaging) were not included, because these imaging-derived variables were not consistently recorded in a structured manner or uniformly assessed during follow-up across both centers, and their absence may have influenced model performance. For example, the factor “Stocking” was binary and did not take into consideration timing or adherence to therapy, which might have underestimated its true influence on the outcomes.
Third, the application of SHAP necessarily revealed only the relationship between each feature to the model without inferring causality. For instance, the apparent interaction between serum uric acid and interval to treatment implies a possible interactive effect on PTS risk but needs confirmation in experimental or longitudinal settings. In addition, the uric acid as a modifiable Biomarker should be interpreted with caution until confirmed by prospective interventions.
Finally, although developing a web-based prediction tool would facilitate translation of our findings into clinical practice, such a tool would need to be tested in real-world settings, optimized to its user interface, and potentially reviewed by regulators. The reliable and safe interpretation and clinical validation of these measures are still pending requirements.
Conclusion
We created and internally validated machine learning models to predict PTS in patients with unprovoked LEDVT in this multicenter study. Significant risk factors were iliofemoral thrombosis, high BMI, treatment delay, high uric acid, and not use of statin. The resulting model showcased good performance and interpretability and was implemented as a user-friendly web-based application. These results indicate the feasibility of machine-learned personalized PTS risk stratification. Future studies are needed for external validation, real-world clinical applicability, and prospective testing of modifiable risk factors, such as uric acid and statins.
Author contributions
All authors (YJL, HRD, YQG) made a significant contribution to the work reported, whether in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; agreed on the journal to which the article has been submitted; and agreed to be accountable for all aspects of the work.
Funding
This study was supported by the National Key Research and Development Program of China [2021YFC2500500].
Data availability
The data analyzed and the codes used during the current study are available from the corresponding author on reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Kahn SR. The post-thrombotic syndrome. Hematology Am Soc Hematol Educ Program. 2016;2016(1):413–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Kahn SR, Shapiro S, Wells PS, Rodger MA, Kovacs MJ, Anderson DR, et al. Compression stockings to prevent post-thrombotic syndrome: a randomised placebo-controlled trial. Lancet. 2014;383(9920):880–8. [DOI] [PubMed] [Google Scholar]
- 3.Engeseth M, Enden T, Sandset PM, Wik HS. Predictors of long-term post-thrombotic syndrome following high proximal deep vein thrombosis: a cross-sectional study. Thromb J. 2021;19(1):3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Vedantham S, Goldhaber SZ, Julian JA, Kahn SR, Jaff MR, Cohen DJ, et al. Pharmacomechanical catheter-directed thrombolysis for deep-vein thrombosis. N Engl J Med. 2017;377(23):2240–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Zhang Y, Lin X, Chen T, Gong S. Association between abnormal systemic coagulation inflammation index and recurrence of deep venous thrombosis as well as quality of life: a retrospective study. Phlebology. 2025. 10.1177/02683555241313240. [DOI] [PubMed] [Google Scholar]
- 6.Makedonov I, Kahn SR, Galanaud JP. Prevention and management of the post-thrombotic syndrome. J Clin Med. 2020. 10.3390/jcm9040923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Wang J, Smeath E, Lim HY, Nandurkar H, Kok HK, Ho P. Current challenges in the prevention and management of post-thrombotic syndrome-towards improved prevention. Int J Hematol. 2023;118(5):547–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Engeseth M, Enden T, Sandset PM, Wik HS. Limitations of the Villalta scale in diagnosing post-thrombotic syndrome. Thromb Res. 2019;184:62–6. [DOI] [PubMed] [Google Scholar]
- 9.Utne KK, Ghanima W, Foyn S, Kahn S, Sandset PM, Wik HS. Development and validation of a tool for patient reporting of symptoms and signs of the post-thrombotic syndrome. Thromb Haemost. 2016;115(2):361–7. [DOI] [PubMed] [Google Scholar]
- 10.Yu T, Song J, Yu L, Deng W. A systematic evaluation and meta-analysis of early prediction of post-thrombotic syndrome. Front Cardiovasc Med. 2023;10:1250480. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Fries AH, Choi E, Han SS. Penalized landmark supermodels (penLM) for dynamic prediction for time-to-event outcomes in high-dimensional data. BMC Med Res Methodol. 2025;25(1):22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu X, Wang M, Wen R, Zhu H, Xiao Y, He Q, et al. Following intravenous thrombolysis, the outcome of diabetes mellitus associated with acute ischemic stroke was predicted via machine learning. Front Pharmacol. 2025;16:1506771. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Sakib S, Bajaj K, Sen P, Li W, Gu J, Li Y, et al. Comparative analysis of machine learning algorithms used for translating aptamer-antigen binding kinetic profiles to diagnostic decisions. ACS Sens. 2025;10(2):907–20. [DOI] [PubMed] [Google Scholar]
- 14.Chen H, Song H, Huang H, Fang X, Chen H, Yang Q, et al. Machine learning prediction and interpretability analysis of high-risk chest pain: a study from the MIMIC-IV database. Front Physiol. 2025;16:1594277. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Jinbo Z, Yufu L, Haitao M. Handling missing data of using the XGBoost-based multiple imputation by chained equations regression method. Front Artif Intell. 2025;8:1553220. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Cicek V, Orhan AL, Saylik F, Sharma V, Tur Y, Erdem A, et al. Predicting short-term mortality in patients with acute pulmonary embolism with deep learning. Circ J. 2025;89(5):602–11. [DOI] [PubMed] [Google Scholar]
- 17.Çiçek V, Yavuz S, Şaylık F, Taşlıçukur Ş, Öz A, Babaoğlu M, et al. Evaluation of pan-Immuno-Inflammation value for In-hospital mortality in acute pulmonary embolism patients. Rev Invest Clin. 2024;76(2):065–79. [DOI] [PubMed] [Google Scholar]
- 18.Pigeot I, Ahrens W. Epidemiology of metabolic syndrome. Pflugers Arch. 2025;477(5):669–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhang S, Tang S, Liu Y, Xue B, Xie Q, Zhao L, et al. Protein-bound uremic toxins as therapeutic targets for cardiovascular, kidney, and metabolic disorders. Front Endocrinol (Lausanne). 2025;16:1500336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Du L, Zong Y, Li H, Wang Q, Xie L, Yang B, et al. Hyperuricemia and its related diseases: mechanisms and advances in therapy. Signal Transduct Target Ther. 2024;9(1):212. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wei X, Zhang M, Huang S, Lan X, Zheng J, Luo H, et al. Hyperuricemia: a key contributor to endothelial dysfunction in cardiovascular diseases. FASEB J. 2023;37(7):e23012. [DOI] [PubMed] [Google Scholar]
- 22.Yin W, Zhou QL, OuYang SX, Chen Y, Gong YT, Liang YM. Uric acid regulates NLRP3/IL-1β signaling pathway and further induces vascular endothelial cells injury in early CKD through ROS activation and K(+) efflux. BMC Nephrol. 2019;20(1):319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Weng H, Li H, Zhang Z, Zhang Y, Xi L, Zhang D, et al. Association between uric acid and risk of venous thromboembolism in East Asian populations: a cohort and Mendelian randomization study. Lancet Reg Health West Pac. 2023;39:100848. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Alturky S, Ashfaq Y, Elhance A, Barney M, Wadiwala I, Hunter AK, et al. Association of post-thrombotic syndrome with metabolic syndrome and inflammation - a systematic review. Front Immunol. 2025;16:1519534. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Iding AFJ, Limpens TMP, Ten Cate H, Ten Cate-Hoek AJ. Chronic inflammatory diseases increase the risk of post-thrombotic syndrome: A prospective cohort study. Eur J Intern Med. 2024;120:85–91. [DOI] [PubMed] [Google Scholar]
- 26.Fraile-Martinez O, García-Montero C, Gomez-Lahoz AM, Sainz F, Bujan J, Barrena-Blázquez S, et al. Evidence of inflammatory network disruption in chronic venous disease: an analysis of circulating cytokines and chemokines. Biomedicines. 2025. 10.3390/biomedicines13010150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Li X, Huang Y, Liang H, Zhong C, Ming Z. Predictive factors for post-thrombotic syndrome in patients with deep vein thrombosis treated with AngioJet pharmacomechanical thrombectomy: A retrospective single-center study. Med Sci Monit. 2025;31:e944805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Zheng Y, Cao C, Chen G, Li S, Ye M, Deng L, et al. Analysis of risk factors for post-thrombotic syndrome after thrombolysis therapy for acute deep venous thrombosis of lower extremities. Internat J Cardiol Cardiovasc Risk Prevent. 2024;22:200319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Jiao J, Zhang D, Peng J, Li Y. MDM2 interacts with PTEN to inhibit endothelial cell development and promote deep vein thrombosis via the JAK/STAT signaling pathway. Mol Med Rep. 2025;31(2):31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Pu H, Lei J, Du G, Huang Q, Qiu P, Liu J, et al. Antiproliferative agent attenuates postthrombotic vein wall remodeling in murine and human subjects. J Thromb Haemost. 2025;23(1):325–40. [DOI] [PubMed] [Google Scholar]
- 31.Wang Z, Zhang P, Tian J, Zhang P, Yang K, Li L. Statins for the primary prevention of venous thromboembolism. Cochrane Database Syst Rev. 2024;11(11):Cd014769. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Gui Q, Wang X, Wu D, Guo Y. Constructing and validating models for predicting gleason grade group upgrading following radical prostatectomy in localized prostate cancer: a comparison between machine learning algorithms and conventional logistic regression. Oncology. 2025. 10.1159/000543492. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Fang C, Zhang L, Xu L, He Y, Zhang X, Xing X. Leveraging machine learning for precision medicine: a predictive model for cognitive impairment in cholestasis patients. BMC Gastroenterol. 2025;25(1):185. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Qiu W, Chen H, Dincer AB, Lundberg S, Kaeberlein M, Lee SI. Interpretable machine learning prediction of all-cause mortality. Commun Med. 2022;2:125. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data analyzed and the codes used during the current study are available from the corresponding author on reasonable request.






