Table 2.
Model performance comparison between use of sequential inputs (LSTM) and temporally flattened inputs (Dense, LR and XGBoost)
| Model | Optimized HPs | Best DT | F1 (95% CI) | Precision (95% CI) | Recall (95% CI) | AUPRC (95% CI) | P-value |
|---|---|---|---|---|---|---|---|
| Best LSTM | Dense units: 88, Learning rate: 0.000308, # LSTM layers: 1, LSTM units layer 1: 112, Dropout layer 1: 0.5 | 0.7 | 0.301 (0.296–0.305) | 0.227 (0.223–0.231) | 0.444 (0.439–0.452) | 0.232 (0.228–0.237) | — |
| Dense | Learning rate: 0.000939, # LSTM layers: 2, LSTM units layer 1: 128, Dropout layer 1: 0.6, LSTM units layer 2: 104, Dropout layer 2: 0.5 | 0.5 | 0.175 (0.173–0.178) | 0.107 (0.105–0.108) | 0.490 (0.484–0.496) | 0.111 (0.108–0.113) | < 0.01 |
| LR | max_iter = 10, class_weight = ‘balanced’, solver = ‘saga' | 0.5 | 0.151 (0.149–0.152) | 0.090 (0.089–0.091) | 0.464 (0.459–0.466) | 0.088 (0.087–0.089) | < 0.01 |
| XGBoost | n_estimators=100, max_depth=6, Learning rate=0.05 | 0.7 | 0.264 (0.260–0.267) | 0.236 (0.233–0.238) | 0.300 (0.197–0.303) | 0.198 (0.195–0.201) | < 0.01 |
The LSTM model was trained on full longitudinal time-series data, whereas baseline models (Dense neural network, logistic regression (LR), and XGBoost) used temporally flattened two-dimensional inputs that ignore sequential structure. For each model, the best decision threshold (DT) was selected based on F1 score. Reported values represent the mean performance on the held-out test set with 95% confidence intervals (CIs) estimated via bootstrap resampling. Model performance is summarized using F1 score, precision, recall, and area under the precision-recall curve (AUPRC). Statistical comparisons between the LSTM and each baseline model were performed using the Mann-Whitney U test applied to the bootstrapped F1 score distributions at the selected DT, with resulting p-values reported in the final column.