Skip to main content
[Preprint]. 2025 Oct 31:rs.3.rs-6585192. [Version 1] doi: 10.21203/rs.3.rs-6585192/v1

Table 2.

Model performances for prediction of depressive symptoms (CESD and MFQ) at 3 months follow-up.

Data type Model CESD MFQ
R2 RMSE MAE R2 RMSE MAE
Univariable Linear Regression 0.645 6.750 4.908 0.671 8.054 5.769
Multivariable- Raw Linear Regression 0.502 8.589 6.631 0.469 10.812 8.210
Elastic Net 0.570 7.098 5.217 0.521 9.019 6.002
Random Forest 0.614 6.723 4.808 0.562 8.616 5.764
XGBoost 0.545 7.476 5.209 0.543 9.039 5.966
SVM 0.510 8.134 5.937 0.459 9.985 7.101
Ensemble 0.616 6.738 4.819 0.549 8.755 5.697
Multivariable-Residuals Linear Regression 0.323 8.891 6.574 0.268 11.11 7.941
Elastic Net 0.540 7.327 5.451 0.524 8.969 6.117
Random Forest 0.503 7.621 5.447 0.558 8.635 5.893
XGBoost 0.332 8.833 6.359 0.464 9.510 6.459
SVM 0.396 8.395 6.215 0.378 10.249 7.331
Ensemble 0.544 7.301 5.302 0.541 8.805 5.849
Multivariable-PC 5 Linear Regression 0.688 6.501 4.873 0.647 8.128 5.931
Elastic Net 0.591 6.912 4.919 0.511 9.093 5.937
Random Forest 0.606 6.785 4.932 0.535 8.866 5.961
XGBoost 0.463 7.891 5.651 0.343 10.523 6.753
SVM 0.602 6.823 4.806 0.509 9.106 5.817
Ensemble 0.600 6.836 4.846 0.528 8.933 5.931
Multivariable-PC 10 Linear Regression 0.588 7.204 4.823 0.521 9.356 5.812
Elastic Net 0.584 6.976 5.014 0.534 8.869 5.792
Random Forest 0.604 6.807 5.022 0.541 8.801 5.883
XGBoost 0.453 7.982 5.672 0.426 9.792 6.404
SVM 0.616 6.705 4.708 0.557 8.649 5.450
Ensemble 0.623 6.638 4.685 0.547 8.752 5.683
Multivariable-PC 20 Linear Regression 0.546 7.468 5.078 0.477 9.248 6.289
Elastic Net 0.562 7.156 5.312 0.529 8.922 5.811
Random Forest 0.580 7.004 5.254 0.506 9.132 6.234
XGBoost 0.425 8.190 5.911 0.364 10.341 6.779
SVM 0.589 6.929 5.007 0.550 8.722 5.662
Ensemble 0.575 7.046 5.096 0.522 8.982 5.959

Results are summarized over 30 repetitions. RMSE: root mean squared error, R2: R squared, MAE: mean absolute error. Univariable: model with respective baseline depression as the sole predictor, Raw: all normalized baseline variables as predictors, Residuals: residualized baseline variables as predictors, PC5: first 5 principal component scores, PC10: first 10 principal component scores, PC20: first 20 principal component scores. The best performance (highest R2, lowest RMSE and MAE) across all data pre-processing procedures are bolded, and the best performance with “Raw” features are underlined.