Skip to main content
. 2026 Jun 11;9:728. doi: 10.1038/s41746-026-02874-1

Table 2.

Model performance comparison between use of sequential inputs (LSTM) and temporally flattened inputs (Dense, LR and XGBoost)

Model Optimized HPs Best DT F1 (95% CI) Precision (95% CI) Recall (95% CI) AUPRC (95% CI) P-value
Best LSTM Dense units: 88, Learning rate: 0.000308, # LSTM layers: 1, LSTM units layer 1: 112, Dropout layer 1: 0.5 0.7 0.301 (0.296–0.305) 0.227 (0.223–0.231) 0.444 (0.439–0.452) 0.232 (0.228–0.237) —
Dense Learning rate: 0.000939, # LSTM layers: 2, LSTM units layer 1: 128, Dropout layer 1: 0.6, LSTM units layer 2: 104, Dropout layer 2: 0.5 0.5 0.175 (0.173–0.178) 0.107 (0.105–0.108) 0.490 (0.484–0.496) 0.111 (0.108–0.113) < 0.01
LR max_iter = 10, class_weight = ‘balanced’, solver = ‘saga' 0.5 0.151 (0.149–0.152) 0.090 (0.089–0.091) 0.464 (0.459–0.466) 0.088 (0.087–0.089) < 0.01
XGBoost n_estimators=100, max_depth=6, Learning rate=0.05 0.7 0.264 (0.260–0.267) 0.236 (0.233–0.238) 0.300 (0.197–0.303) 0.198 (0.195–0.201) < 0.01

The LSTM model was trained on full longitudinal time-series data, whereas baseline models (Dense neural network, logistic regression (LR), and XGBoost) used temporally flattened two-dimensional inputs that ignore sequential structure. For each model, the best decision threshold (DT) was selected based on F1 score. Reported values represent the mean performance on the held-out test set with 95% confidence intervals (CIs) estimated via bootstrap resampling. Model performance is summarized using F1 score, precision, recall, and area under the precision-recall curve (AUPRC). Statistical comparisons between the LSTM and each baseline model were performed using the Mann-Whitney U test applied to the bootstrapped F1 score distributions at the selected DT, with resulting p-values reported in the final column.