Table 2.
Model performances for prediction of depressive symptoms (CESD and MFQ) at 3 months follow-up.
| CESD | MFQ | ||||||
|---|---|---|---|---|---|---|---|
| R2 | RMSE | MAE | R2 | RMSE | MAE | ||
| Univariable | Linear Regression | 0.645 | 6.750 | 4.908 | 0.671 | 8.054 | 5.769 |
| Multivariable- Raw | Linear Regression | 0.502 | 8.589 | 6.631 | 0.469 | 10.812 | 8.210 |
| Elastic Net | 0.570 | 7.098 | 5.217 | 0.521 | 9.019 | 6.002 | |
| Random Forest | 0.614 | 6.723 | 4.808 | 0.562 | 8.616 | 5.764 | |
| XGBoost | 0.545 | 7.476 | 5.209 | 0.543 | 9.039 | 5.966 | |
| SVM | 0.510 | 8.134 | 5.937 | 0.459 | 9.985 | 7.101 | |
| Ensemble | 0.616 | 6.738 | 4.819 | 0.549 | 8.755 | 5.697 | |
| Multivariable-Residuals | Linear Regression | 0.323 | 8.891 | 6.574 | 0.268 | 11.11 | 7.941 |
| Elastic Net | 0.540 | 7.327 | 5.451 | 0.524 | 8.969 | 6.117 | |
| Random Forest | 0.503 | 7.621 | 5.447 | 0.558 | 8.635 | 5.893 | |
| XGBoost | 0.332 | 8.833 | 6.359 | 0.464 | 9.510 | 6.459 | |
| SVM | 0.396 | 8.395 | 6.215 | 0.378 | 10.249 | 7.331 | |
| Ensemble | 0.544 | 7.301 | 5.302 | 0.541 | 8.805 | 5.849 | |
| Multivariable-PC 5 | Linear Regression | 0.688 | 6.501 | 4.873 | 0.647 | 8.128 | 5.931 |
| Elastic Net | 0.591 | 6.912 | 4.919 | 0.511 | 9.093 | 5.937 | |
| Random Forest | 0.606 | 6.785 | 4.932 | 0.535 | 8.866 | 5.961 | |
| XGBoost | 0.463 | 7.891 | 5.651 | 0.343 | 10.523 | 6.753 | |
| SVM | 0.602 | 6.823 | 4.806 | 0.509 | 9.106 | 5.817 | |
| Ensemble | 0.600 | 6.836 | 4.846 | 0.528 | 8.933 | 5.931 | |
| Multivariable-PC 10 | Linear Regression | 0.588 | 7.204 | 4.823 | 0.521 | 9.356 | 5.812 |
| Elastic Net | 0.584 | 6.976 | 5.014 | 0.534 | 8.869 | 5.792 | |
| Random Forest | 0.604 | 6.807 | 5.022 | 0.541 | 8.801 | 5.883 | |
| XGBoost | 0.453 | 7.982 | 5.672 | 0.426 | 9.792 | 6.404 | |
| SVM | 0.616 | 6.705 | 4.708 | 0.557 | 8.649 | 5.450 | |
| Ensemble | 0.623 | 6.638 | 4.685 | 0.547 | 8.752 | 5.683 | |
| Multivariable-PC 20 | Linear Regression | 0.546 | 7.468 | 5.078 | 0.477 | 9.248 | 6.289 |
| Elastic Net | 0.562 | 7.156 | 5.312 | 0.529 | 8.922 | 5.811 | |
| Random Forest | 0.580 | 7.004 | 5.254 | 0.506 | 9.132 | 6.234 | |
| XGBoost | 0.425 | 8.190 | 5.911 | 0.364 | 10.341 | 6.779 | |
| SVM | 0.589 | 6.929 | 5.007 | 0.550 | 8.722 | 5.662 | |
| Ensemble | 0.575 | 7.046 | 5.096 | 0.522 | 8.982 | 5.959 | |
Results are summarized over 30 repetitions. RMSE: root mean squared error, R2: R squared, MAE: mean absolute error. Univariable: model with respective baseline depression as the sole predictor, Raw: all normalized baseline variables as predictors, Residuals: residualized baseline variables as predictors, PC5: first 5 principal component scores, PC10: first 10 principal component scores, PC20: first 20 principal component scores. The best performance (highest R2, lowest RMSE and MAE) across all data pre-processing procedures are bolded, and the best performance with “Raw” features are underlined.