Skip to main content
. 2025 Aug 12;13:1537098. doi: 10.3389/fped.2025.1537098

Figure 2.

Flowchart illustrating a machine learning pipeline for NAFLD data. It starts with data acquisition then proceeds through steps like recalibration, feature engineering, class imbalance correction, model training, and evaluation. Techniques include missing data imputation, feature encoding, SMOTE for oversampling, stratified sampling, and algorithms like CatBoost. The chart notes specific metric outcomes, hyperparameter tuning, and performance validation, ending in clinical deployment suggestions for screening NAFLD and guiding biopsy decisions. Performance drift is monitored, with a feedback loop for continuous improvement.

Machine learning pipeline for pediatric NAFLD prediction. Diagram created with MermaidChart (https://www.mermaidchart.com/). Illustrates the end-to-end workflow including: (1) Internal data processing (n = 659, 120 NAFLD+) with median/mode imputation and SMOTE-based class balancing; (2) CatBoost/AdaBoost model training (5-fold CV); (3) External validation with Jensen-Shannon divergence checks. The feedback loop enables automatic recalibration when performance drift >5% is detected.