Skip to main content
. 2025 Aug 8;15:1623109. doi: 10.3389/fcimb.2025.1623109

Figure 1.

Flowchart showing process for sample selection and multi-model development in a study. Left side details excluding patients based on criteria, resulting in 1,389. Right side involves data analysis, variable selection, and machine learning model testing with emphasis on XGBoost. Models are split for training and testing, highlighting XGBoost's predictive ability.

Study cohort selection and model development workflow. From 58,078 urinary tract infection patients in MIMIC-IV, 1,389 met inclusion criteria. Twelve key features were selected using four machine learning methods (RF, Lasso, Boruta, XGBoost), refined to nine variables through clinical review. The cohort was split 7:3 (training: testing). Seven ML models were trained; XGBoost demonstrated optimal performance and was validated.