Abstract
Background:
Whole blood viscosity (WBV) has been linked to cardiometabolic risk, yet its relationship with short-term postprandial triglyceride (TG) response remains unclear. We evaluated this question within a fully de-identified, statistically reconstructed synthetic cohort (
), with a primary methodological objective: to demonstrate a rigorously leakage-controlled machine-learning framework for assessing candidate physiological associations.
Methods:
A strictly leakage-controlled pipeline was implemented, with fold-specific preprocessing and probability calibration confined to training data. WBV was estimated using the de Simone formulation. Model development employed stratified
nested cross-validation, repeated
resampling, 1,000-bootstrap uncertainty estimation, calibration assessment (sigmoid and isotonic), threshold-sensitivity analysis across
percentiles, and SHAP-based interpretability.
Results:
Within the synthetic reconstruction, WBV showed negligible association with postprandial response across correlation testing (
), multivariable modeling, robustness analyses, and explainability assessments. In contrast, fasting triglycerides (TG
) exhibited stable monotonic effects and clear phenotype discrimination. Under the primary 75th-percentile
definition, the L2-penalized logistic regression model achieved stable discrimination (nested AUROC
; Brier
) with bootstrap AUROC
(95% CI [0.8957, 0.9314]) and consistent calibration-aware performance.
Conclusions:
Within this synthetic framework, WBV did not provide reproducible predictive or attributional value for short-term postprandial TG response. These findings represent methodological evidence under modeled assumptions rather than definitive physiological conclusions. The study illustrates how leakage-controlled, calibration-aware ML workflows can evaluate candidate metabolic associations in privacy-preserving settings; external validation in real-world cohorts remains necessary.
Keywords: Bioinformatics, Machine learning, Explainable artificial intelligence (XAI), Leakage-controlled modeling, Calibration analysis, Cardiometabolic risk, Postprandial metabolism, Triglyceride response, Whole blood viscosity
Subject terms: Biomarkers, Computational biology and bioinformatics, Diseases, Health care, Medical research
Introduction
Postprandial triglyceride (TG) metabolism reflects an individual’s capacity to regulate circulating lipids following dietary fat intake and has been associated with multiple cardiometabolic risk pathways1,2. Impaired postprandial triglyceride dynamics have been linked to endothelial dysfunction, atherogenic remnant accumulation, insulin resistance, and long-term cardiovascular risk3–5. Despite this relevance, the short-term determinants of postprandial TG response remain incompletely characterized, and the relative contributions of metabolic versus hemorheological factors are not fully established.
Whole blood viscosity (WBV), a composite hemorheological parameter reflecting blood thickness and flow resistance, has been associated with obesity, dyslipidemia, hypertension, and systemic inflammation6–8. These observations have led to hypotheses that elevated viscosity may influence microvascular perfusion and potentially modulate lipid handling9. However, empirical evidence directly linking WBV to short-term postprandial triglyceride response remains limited and heterogeneous10. Moreover, most prior investigations have relied on conventional statistical analyses in relatively small cohorts, without systematically addressing nonlinear interactions, multivariable dependencies, probability calibration, or risks of information leakage that can bias inference in machine-learning (ML) studies.
In parallel, ML methods have become increasingly prominent in metabolic phenotyping and clinical risk modeling11,12. Yet, methodological rigor varies substantially across studies. Common pitfalls, including global preprocessing, improper normalization outside training folds, absence of probability calibration, and limited robustness assessment, can artificially inflate predictive performance and inadvertently strengthen biological interpretations13–15. These concerns highlight the need for carefully designed, leakage-controlled analytical pipelines when evaluating candidate physiological associations. The present study therefore adopts a primarily methodological objective: to implement and demonstrate a rigorously leakage-controlled, calibration-aware ML framework applied to the investigation of WBV and postprandial TG response within a fully de-identified, statistically reconstructed synthetic cohort. The synthetic cohort is constructed to approximate clinically realistic marginal distributions and correlation structure among lipid, hematologic, and metabolic variables, enabling reproducible experimentation without identifiable patient data16,17. The modeling pipeline integrates nested cross-validation, fold-specific preprocessing, probability calibration, repeated resampling, threshold sensitivity analysis, and model interpretability using SHAP18 and partial-dependence analyses.
Within this structured analytical setting, we specifically ask whether WBV provides incremental predictive or attributional information for postprandial TG response phenotypes beyond routine fasting biomarkers such as TG
and BMI. Findings derived from this synthetic reconstruction are interpreted as methodological evidence under explicitly modeled assumptions rather than definitive physiological conclusions. By centering analytical rigor, transparency, and calibration reliability, this work aims to contribute a reproducible evaluation framework for biomedical informatics applications, illustrating how candidate metabolic associations can be investigated while minimizing leakage, miscalibration, and overinterpretation.
Materials and methods
This study used a fully de-identified, statistically reconstructed synthetic dataset designed to approximate clinically realistic distributions of fasting and postprandial triglycerides and standard metabolic biomarkers. No real patient records or identifiable information were used at any stage, and no access to human-subject data occurred during dataset generation.
Each synthetic record contained fasting triglycerides (
), simulated 4-hour postprandial triglycerides (
), hematocrit (Hct), total protein (TP), HDL-C, LDL-C, BMI, age, and sex. The 4-hour postprandial time-point was selected as the most informative measurement window for standardized fat tolerance testing1. All analyses were conducted within a strictly leakage-controlled machine-learning workflow to ensure methodological validity and reproducibility.
Synthetic clinical cohort generation
A synthetic cohort of
adults was generated to emulate a hospital population undergoing fasting and postprandial lipid testing. The generative process followed a two-stage framework in which univariate marginal distributions and multivariate dependence structure were modeled separately using non-identifiable summary-level statistics derived from a fully de-identified reference population. Fasting triglycerides (
) were modeled using a log-normal distribution to capture metabolic right-skewness on a clinically realistic mg/dL scale.
To ensure physiologically coherent postprandial dynamics, 4-hour triglycerides (
) were generated as a positive excursion from fasting levels.
![]() |
where
represents a simulated postprandial increase with mild dependence on BMI and baseline triglycerides, ensuring a postprandial rise rather than a decrease. Hematocrit and total protein were simulated using truncated normal distributions constrained to physiologic ranges (Hct: 30–55%, TP: 5.5–8.8 g/dL). HDL-C, LDL-C, BMI, and age were generated using distributions consistent with adult metabolic ranges, while sex was sampled using a Bernoulli process calibrated to preserve the population-level sex ratio.
In Table 1, TGR was computed deterministically from fasting and 4-hour triglyceride levels as
. The high-response phenotype was defined as the upper quartile (
75th percentile) of simulated
values rather than as the percentage change from baseline. Consequently, group means of TGR are not equal to the ratios computed from group means of TG values.
Table 1.
Baseline characteristics by postprandial triglyceride response phenotype in the rebuilt physiologically coherent synthetic cohort (mean ± SD or n (%)).
denotes simulated 4-hour postprandial triglycerides under a standardized meal scenario.
| Normal (n = 1125) | High (n = 375) | p-value | |
|---|---|---|---|
| Female, n (%) | 587 (52.1778%) | 198 (52.8000%) | 0.8814 |
| Male, n (%) | 538 (47.8222%) | 177 (47.2000%) | 0.8814 |
| Age, years | 52.5006 ± 10.1075 | 52.8123 ± 10.0940 | 0.6049 |
| Hematocrit, % | 42.4550 ± 3.6790 | 42.3444 ± 3.7145 | 0.6169 |
| Total protein, g/dL | 6.9850 ± 0.5083 | 7.0252 ± 0.4609 | 0.1545 |
| Whole blood viscosity, cP | 5.9302 ± 0.4507 | 5.9237 ± 0.4583 | 0.8132 |
Fasting TG ( ), mg/dL |
115.4315 ± 48.1690 | 260.5150 ± 111.8678 | ![]() |
4-h TG ( ), mg/dL |
219.6445 ± 44.1151 | 392.7571 ± 73.7592 | ![]() |
| Postprandial response (TGR), % | 92.0215 ± 33.0955 | 71.4660 ± 32.3089 | ![]() |
| HDL-C, mg/dL | 51.7786 ± 9.8241 | 48.6497 ± 10.3360 | ![]() |
| LDL-C, mg/dL | 130.0714 ± 29.4531 | 130.7746 ± 28.0713 | 0.5744 |
BMI, kg/m
|
24.0590 ± 2.8360 | 25.3032 ± 2.8331 | ![]() |
To enhance transparency of the synthetic data generation process, we explicitly report the aggregate marginal targets (means, standard deviations, and selected percentiles) together with the key biochemical correlation submatrix used to parameterize the Gaussian copula dependence structure underlying the multivariate simulation. These tables fully disclose the marginal distribution targets and dependence parameters governing the synthetic cohort construction, with no additional hidden constraints, post-hoc tuning, or latent calibration adjustments applied. Importantly, the negligible association between WBV and triglyceride response metrics arises directly from the encoded correlation structure rather than from downstream modeling artefacts; thus, all null inferences regarding WBV should be interpreted as conditional on the specified summary-level biochemical relationships, rather than as definitive claims about in vivo physiology.
Table 2.
Key Pearson correlation submatrix used in synthetic dependence modeling.
TG
|
TG
|
Hct | TP | WBV | |
|---|---|---|---|---|---|
TG
|
1.0000 | 0.7814 | 0.0217 | 0.0185 | 0.0142 |
TG
|
0.7814 | 1.0000 | 0.0169 | 0.0203 | 0.0176 |
| Hct | 0.0217 | 0.0169 | 1.0000 | 0.2146 | 0.6127 |
| TP | 0.0185 | 0.0203 | 0.2146 | 1.0000 | 0.5834 |
| WBV | 0.0142 | 0.0176 | 0.6127 | 0.5834 | 1.0000 |
Values shown to 4 decimal places.
Synthetic data quality validation
To ensure that the statistically reconstructed synthetic dataset was suitable for methodological machine-learning evaluation, we performed a structured six-dimension validation assessing: distributional fidelity, biochemical correlation structure, multivariate joint coherence, phenotype robustness, internal coherence of derived variables, and overall suitability for ML experimentation.
Distributional fidelity
Univariate distributions of all primary variables (
,
, Hct, TP, HDL-C, LDL-C, BMI, age, sex) were examined using kernel density estimates and descriptive summaries. All variables exhibited physiologic ranges and expected shapes, including the characteristic right-skewness of triglycerides and appropriate laboratory intervals for hematocrit and total protein.
Preservation of biochemical correlation structure
Correlation matrices were inspected to confirm preservation of key biochemical relationships, including the inverse HDL–TG association and weak correlations among viscosity-related parameters. No spurious dependencies were introduced, indicating that the synthetic generation process maintained realistic biochemical structure.
Multivariate joint distribution consistency
Scatterplots and contour-density maps were used to evaluate multivariate geometry. Joint patterns such as monotonic
increases across
strata and the relative independence of WBV from lipid markers were smooth and clinically plausible, reflecting coherent multivariate structure.
Phenotype stability across
Thresholds
The primary high-responder phenotype was defined using the 75th percentile of
computed within each training fold to preserve strict leakage control. In the full reconstructed cohort, this percentile corresponded to approximately 329.2875 mg/dL (i.e., 25% high responders), consistent with the descriptive cohort summary in Table 1.
Sensitivity analyses across percentile thresholds from the 60th to the 90th demonstrated stable discrimination performance, with AUROC values consistently clustered around 0.92 across the 60th–90th percentile range. Minor variation at extreme thresholds reflected expected changes in class balance rather than model instability.
Internal coherence of derived variables
Derived quantities (WBV and TGR) were evaluated for physiologic plausibility. WBV values predominantly fell within the expected 5–7 cP range, and TGR displayed realistic postprandial triglyceride response behaviors without implausible or contradictory values, supporting the internal consistency of derived features.
Suitability for methodological evaluation
Because the primary aim of this work was to evaluate a leakage-controlled machine-learning framework rather than to estimate population-level epidemiologic parameters, the synthetic dataset needed to support robust preprocessing, cross-validation, calibration, sensitivity analyses, and explainability workflows. All six validation dimensions were satisfied, indicating that the dataset was appropriate for methodological ML research in biomedical informatics.
Summary
Across all six dimensions, the synthetic dataset demonstrated high distributional fidelity, realistic correlation structure, coherent multivariate patterns, robust phenotype definition, and internally consistent derived variables, confirming its suitability for leakage-controlled ML evaluation. A summary of the synthetic data quality validation dimensions is provided in Table 4.
Table 4.
Summary of synthetic data quality validation dimensions, procedures, and key findings.
| Validation dimension | Key procedures and findings |
|---|---|
| Distributional fidelity | KDE curves and descriptive statistics confirmed physiologically realistic ranges for , , Hct, TP, HDL-C, LDL-C, BMI, age, and sex, with expected right-skewness for triglycerides. |
| Biochemical correlation structure | Pearson correlations reproduced known lipid physiology (e.g., inverse HDL–TG association) and showed weak WBV–lipid relationships without spurious links. |
| Multivariate joint distribution | Scatterplots and density-contour maps demonstrated smooth – and –BMI gradients, with WBV remaining independent of lipid markers. |
| Phenotype robustness | Threshold sweeps from the 60th–90th percentiles of yielded stable AUROC values (~0.92 for 60th–90th), indicating that performance was not sensitive to the exact percentile cut-off. |
| Derived variable coherence | WBV and TGR values remained within established physiological ranges (WBV 5–7 cP), and TGR patterns reflected realistic postprandial kinetics. |
| Suitability for ML evaluation | The synthetic dataset supported full leakage-controlled cross-validation, calibration testing, robustness resampling, and explainability workflows without inconsistencies. |
Inclusion and exclusion criteria
Synthetic records were retained if they met all of the following criteria:
complete fasting and 4-hour triglyceride measurements;
valid hematocrit and total protein values for WBV estimation;
no missing demographic variables;
physiologically plausible laboratory ranges.
Predefined physiological plausibility thresholds were specified (triglycerides > 1500 mg/dL; WBV < 3 or > 10 cP) to ensure data realism and safeguard against extreme outliers.
However, under the parameterization of the synthetic generator, no records violated these criteria. Triglyceride values exceeding 1500 mg/dL were not produced, and WBV values remained within clinically plausible bounds (approximately 5–7 cP). Consequently, no post-generation exclusions were ultimately required.
Estimation of whole blood viscosity (WBV)
Whole blood viscosity (WBV) was not directly measured by viscometry, but was estimated from hematocrit (Hct) and total protein (TP) using the empirical de Simone model19, which is widely used as a non-invasive surrogate of low-shear WBV in clinical and epidemiologic studies. WBV was computed as a deterministic function of Hct and TP, constrained to physiologically plausible ranges, and interpreted as a hemorheological proxy rather than an exact physical measurement. All inferences regarding the “null” role of WBV are therefore conditional on this surrogate representation.
Phenotype definition and distributional characteristics
The high-response phenotype was defined as the upper quartile (
75th percentile) of 4-hour postprandial triglyceride levels (
). This percentile-based definition ensures a fixed class prevalence (25%) and supports stable, stratified cross-validation within the leakage-controlled modeling framework. As shown in Figure 1, the
distribution demonstrates clear separation by construction, reflecting the threshold-based phenotype definition.
Fig. 1.

Fasting and postprandial triglyceride distributions across response phenotypes. Panel (a) shows the distribution of fasting triglycerides (
), where high responders exhibit a right-shifted baseline lipid profile. Panel (b) shows 4-hour postprandial triglycerides (
), demonstrating clear phenotype separation and supporting the validity of the response classification.
Importantly, fasting triglycerides (
) exhibit a consistent rightward shift among high responders, indicating that elevated postprandial burden is aligned with higher baseline lipid levels. These distributional patterns establish transparency between phenotype construction and baseline metabolic load while avoiding circularity in modeling (since
is used only for label definition and never as a predictor). Together, the observed separation supports the internal coherence of the synthetic phenotype used in subsequent leakage-controlled analyses.
Computation of relative postprandial triglyceride response (TGR)
The Relative Postprandial Triglyceride Response (TGR) was defined as:
![]() |
This metric reflects the magnitude of postprandial triglyceride elevation relative to fasting baseline rather than true kinetic response, which would require multi-time-point measurements beyond 0h and 4h.
Variable definition and subgroup stratification
The dataset comprised:
clinical laboratory variables (
,
, HDL, LDL, Hct, TP);derived variables (WBV, TGR);
demographic variables (age, sex, BMI).
Demographic pseudo-external-synthetic transfer was conducted via stratified subgroup evaluation by age (
vs.
years) and sex (male vs. female). This analysis assesses internal transport stability across demographic strata rather than true external-synthetic transfer across independent clinical sites.
Inter-Variable Correlation Structure
To characterize the baseline dependency structure, we computed the Pearson correlation matrix among fasting lipid markers (
, HDL, LDL), postprandial triglycerides (
), hematologic variables (hematocrit, total protein), derived variables (WBV, TGR), and demographic factors (age, BMI). As shown in Figure 2,
and
demonstrate a strong positive association, consistent with the synthetic data-generating mechanism in which postprandial levels arise from fasting baseline. Hematocrit and total protein are strongly correlated with WBV, reflecting the deterministic de Simone–based surrogate formulation.
Fig. 2.

Pearson correlation matrix of key clinical, derived, and demographic variables. Strong associations are observed between
and
, and between hematocrit, total protein, and estimated whole blood viscosity (WBV). WBV exhibits weak linear correlation with triglyceride-related variables.
Importantly, WBV exhibits negligible linear correlation with triglyceride-related variables (
,
, TGR). This pattern mirrors the explicitly encoded correlation structure of the synthetic cohort and anticipates the subsequent multivariable and explainability results. These correlation findings are descriptive and structural in nature and do not imply causal inference beyond the modeled assumptions.
Statistical analysis
Conventional statistical analyses included Pearson correlations among WBV,
,
, and TGR; phenotype-wise comparisons between normal and high responders; and regression modeling. Model reliability was further assessed using 1,000 bootstrap resamples, with confidence intervals for AUROC and Brier score computed from the bootstrap distributions.
Sensitivity analysis for WBV estimation error
Because WBV was derived from an empirical formula rather than direct viscometric measurement, we performed a sensitivity analysis to assess whether plausible misspecification of the viscosity surrogate could overturn the observed null findings. We assumed that de Simone-based WBV estimates might be affected by moderate relative error and generated perturbed WBV values by multiplying the original WBV by random factors corresponding to
and
Gaussian noise around unity. For each perturbation scenario, we recomputed:
The Pearson correlation between WBV and TGR;
The discriminative performance of WBV as a univariable predictor of the high-response phenotype;
The contribution of WBV within the full multivariable L2-penalized logistic regression model, including its regression coefficient magnitude within the multivariable L2-penalized logistic regression model.
If WBV meaningfully influenced triglyceride response, small perturbations of the viscosity surrogate would be expected to produce noticeable shifts in effect sizes or performance metrics. The stability of results across all perturbation scenarios instead supports the conclusion that WBV plays a negligible role in determining high-response risk within the synthetic framework / under modeled assumptions.
Machine-learning framework
Preprocessing (fold-specific)
To eliminate data leakage, all preprocessing steps, including standardization, mild winsorization of extreme values, and one-hot encoding, were performed independently within each training fold. No global scaling or pre-split normalization was applied. Within each training fold, continuous variables were winsorized at pre-specified symmetric percentile cut points, and these cut points were then applied only to the corresponding validation fold to preserve strict leakage control.
Model Zoo
Several candidate classifiers were evaluated, including penalized logistic regression and tree-based ensemble methods. The L2-penalized logistic regression model was ultimately selected as the final classifier based on nested cross-validation discrimination, calibration performance, and stability characteristics.
Nested 5
5 Cross-Validation
Nested cross-validation was implemented with:
An inner loop for hyperparameter tuning;
An outer loop for unbiased performance estimation.
Probability calibration
Two probability calibration methods were examined:
Robustness and threshold sensitivity
Robustness was assessed using repeated 5
10 cross-validation and 1,000 bootstrap resamples. Threshold sensitivity was evaluated across
phenotype cut-offs from the 60th to the 90th percentile. The overall leakage-controlled machine-learning workflow is illustrated in Figure 4.
Fig. 4.

Leakage-controlled machine-learning pipeline for predicting high postprandial triglyceride response from fasting biomarkers. Postprandial measurements (
) and derived response metrics (TGR) are used only to define the outcome label and are never included as predictors. Predictors are restricted to baseline fasting variables (including WBV derived from Hct and TP). In the current rerun pipeline, the final classifier is an L2-penalized logistic regression model with all preprocessing and calibration fitted exclusively within training folds.
Synthetic domain-shift stress test under controlled distributional perturbation
To evaluate calibration stability and ranking behavior under controlled distributional perturbation, we conducted a synthetic domain-shift stress-test experiment using an independently generated synthetic cohort. This experiment does not constitute external clinical validation; rather, it represents a simulation-based robustness assessment within a fully synthetic modeling framework. In this setting, postprandial triglyceride values were generated deterministically from baseline triglyceride levels under predefined perturbation parameters.
Because the phenotype label is structurally linked to the underlying triglyceride trajectory within a low-noise copula-based simulation, discrimination can approach perfect separation (AUROC
1.00) when the model is fitted and evaluated on the same synthetic realization without cross-validation. This behavior reflects the deterministic construction of the synthetic outcome and should not be interpreted as evidence of real-world clinical validity. The leakage-controlled L2-penalized logistic regression model was trained on the original cohort using standardized predictors and 5-fold internal cross-validation, and probability calibration employed a three-step procedure:
Training on hard high-responder labels;
Temperature scaling of logits;
Isotonic regression fitted to soft labels derived from proximity to the
threshold.
The temperature parameter was estimated using out-of-fold cross-validation on the original training cohort and was not fitted on the synthetic stress-test cohort, thereby preventing test-set information leakage. The calibrated model was subsequently applied to the independently generated synthetic stress-test cohort without refitting. Model performance was assessed using AUROC, Brier score, and calibration plots. This three-step calibration strategy was used exclusively for the synthetic domain-shift stress test and is analytically distinct from the nested cross-validation calibration procedures reported in the primary analysis (Table 6).
Table 6.
Predictive performance of TG-only baselines and the final multivariable L2-penalized logistic regression model at the
75th percentile phenotype definition. All values are averaged across stratified
nested cross-validation folds.
| Model | AUROC | Brier score | F1-score | Recall | Precision |
|---|---|---|---|---|---|
75th percentile cut-off (threshold mg/dL) |
0.8204 | 0.1347 | |||
Univariate logistic regression ( only) |
0.9167 | 0.1056 | 0.6120 | 0.4480 | 0.9655 |
| L2-penalized logistic regression (multivariable) | 0.9141 | 0.0886 | 0.7348 | 0.6613 | 0.8267 |
In this synthetic stress-test experiment, the data-generating process linked the high-responder label almost deterministically to the triglyceride trajectory, with minimal residual noise. Consequently, this analysis serves primarily as a controlled check of ranking and calibration behavior under synthetic perturbation rather than as evidence of real-world transport performance. Importantly, no independent real-patient cohort was used in this study; this stress test is confined to the synthetic data-generating framework and should be interpreted accordingly.
Predictor specification and leakage control
The final classification models used only baseline (fasting) variables as predictors: age, sex, BMI, fasting triglycerides (
), HDL-C, LDL-C, hematocrit, total protein, and whole blood viscosity (derived from hematocrit and total protein). Postprandial triglycerides (
) and derived response metrics (TGR) were used exclusively for phenotype definition and were never included as predictors. All preprocessing steps (winsorization, scaling, encoding, and calibration) were performed within training folds under nested cross-validation to prevent data leakage.
Evaluation metrics and reliability analysis
Primary evaluation metrics included AUROC, precision, recall, area under the curve (average precision, AP), and Brier score, with secondary metrics comprising accuracy, sensitivity, specificity, and F1-score. All metrics were computed strictly within the leakage-controlled, fold-specific preprocessing pipeline to ensure unbiased estimation of generalization performance. Given the mild class imbalance of the phenotype labels, precision, recall, and AUC were emphasized as complementary measures of discriminative performance. Reliability and generalizability were evaluated using:
Calibration curves and Brier scores to assess probability accuracy;
Calibration slope and intercept to quantify overfitting or underfitting in predicted risks;
Decision curve analysis (DCA) to evaluate clinical net benefit across threshold probabilities22;
Pseudo-external-synthetic transfer stratified by age and sex to assess demographic transportability;
Multi-layer robustness assessment, including 1,000 bootstrap resamples and repeated 5
10 cross-validation, to examine model stability and variance across resampled datasets.
Methodological innovation and framework contribution
This study introduces a fully reproducible, leakage-controlled machine-learning framework for clinical biochemical prediction tasks. Key methodological components include:
Fold-specific preprocessing: all feature scaling and encoding was performed strictly within each training fold to eliminate data leakage and preserve true generalization behaviour;
Integrated calibration diagnostics: model calibration was evaluated using calibration curves, calibration slopes, Brier scores, and expected calibration error (ECE; 10-bin, unweighted);
Multi-layer robustness testing: bootstrap resampling (1,000 iterations), repeated 5
10 cross-validation, threshold sensitivity analyses, and comparison against multiple baseline classifiers;Explainability integration: interpretability via SHAP summary plots18, SHAP dependence and interaction analyses, and one- and two-dimensional partial dependence plots;
Demographic pseudo-external-synthetic transfer: subgroup validation by age and sex to simulate external-synthetic transfer in the absence of multi-centre data.
Ethical considerations
The dataset used in this study is a fully de-identified, statistically reconstructed synthetic dataset and does not contain any real patient records or identifiable personal information. No direct interaction with human subjects or access to identifiable clinical data occurred at any stage. Accordingly, the analysis does not constitute human-subject research under prevailing international regulatory definitions, and formal institutional review board (IRB) approval was not required. No funding was received from external commercial entities.
Results
Across the full analytic workflow, the primary findings were consistent: whole blood viscosity (WBV) showed a null association with postprandial triglyceride response ratio (TGR), whereas fasting triglycerides (
) emerged as the dominant predictor of high postprandial response. The final L2-penalized logistic regression model achieved a mean AUROC of 0.9141 in nested cross-validation with a corresponding Brier score of 0.0886, and a bootstrap AUROC mean of 0.9140 (95% CI: 0.8957–0.9314) with a mean Brier score of 0.0887 (95% CI: 0.0789–0.0983), alongside stable probability quality under post-hoc calibration. Detailed results are presented below.
Cohort characteristics and phenotype separation
A total of 1,500 synthetic adult records were analyzed, including 1,125 normal responders (75%) and 375 high responders (25%), classified using the 75th percentile of 4-hour postprandial triglyceride (
) as the cut-off value (329.2875 mg/dL). Baseline characteristics stratified by response phenotype are presented in Table 1. Sex distribution was comparable between normal and high responders (female: 52.18% vs. 52.80%;
), and no clinically meaningful differences were observed in hematocrit, total protein, or whole blood viscosity (WBV: 5.93 ± 0.45 vs. 5.92 ± 0.46 cP;
).
In contrast, triglyceride-related parameters showed clear separation between phenotypes. Normal responders exhibited substantially lower fasting
(115.43 ± 48.17 mg/dL) compared with high responders (260.52 ± 111.87 mg/dL;
), as well as lower
levels (219.64 ± 44.12 vs. 392.76 ± 73.76 mg/dL;
). Notably, the relative response ratio (TGR) was lower in high responders (71.47 ± 32.31%) than in normal responders (92.02 ± 33.10%;
), reflecting that the phenotype is defined by absolute
burden rather than relative change from baseline. Distributions of fasting lipid and biochemical variables are illustrated in Figure 1, while the correlation structure of baseline biomarkers is shown in Fig. 2, indicating that WBV has negligible correlation with lipid-related measures.
Null association between whole blood viscosity and triglyceride response
At the population level, WBV exhibited a near-zero linear relationship with TGR. Pearson’s correlation between WBV and TGR was
(
), with an
from simple linear regression of 0.0001, indicating no meaningful explanatory power for interindividual variation in TGR. Similarly, WBV showed only trivial correlations with fasting and postprandial triglycerides (WBV vs.
:
; WBV vs.
:
). In contrast, fasting
demonstrated a modest inverse linear association with TGR when modelled continuously (
), reflecting the non-linear and phenotype-based structure of the response definition rather than a simple proportional relationship. Bootstrap-derived confidence intervals confirmed that the WBV–triglyceride associations were statistically indistinguishable from zero (Table 3), supporting the absence of detectable linear association within the synthetic cohort rather than reliance on point estimates alone.
Table 3.
Bootstrap (1000 resamples) confidence intervals for key Pearson correlations.
| Association | r | 95% CI (lower) | 95% CI (upper) |
|---|---|---|---|
WBV – TG
|
0.0142 | -0.0338 | 0.0607 |
WBV – TG
|
0.0176 | -0.0294 | 0.0639 |
| WBV – TGR | -0.0104 | -0.0582 | 0.0351 |
Consistent with these findings, the fasting biomarker correlation matrix (Fig. 2) shows WBV forming a largely independent axis relative to triglyceride and cholesterol measures. Together, these results suggest that WBV neither tracks with baseline lipid load nor explains interindividual variation in TGR at the linear association level.
Sensitivity analysis for WBV estimation error
Because WBV was estimated using the de Simone formula rather than measured directly by viscometry, we tested whether plausible misspecification of the viscosity surrogate could generate spurious associations with triglyceride response. Perturbed WBV values were generated by multiplying the original WBV by Gaussian noise corresponding to relative errors of
and
. Across all perturbation scenarios, the correlation between WBV and TGR remained near zero (mean
with 5% error and
with 10% error). WBV alone provided no discriminative information for the high-response phenotype (AUC
and 0.4968 under 5% and 10% perturbation, respectively). These results indicate that the null WBV–TGR association is robust to plausible levels of error in the empirical WBV estimation model. The results of the WBV sensitivity analysis are reported in Table 5.
Table 5.
Sensitivity analysis of the WBV–TGR relationship under plausible viscosity model misspecification.
| Relative error | corr(WBV, TGR) mean | std | AUC(WBV-only) mean | std |
|---|---|---|---|---|
| 5% error | ![]() |
0.01488 | 0.4952 | 0.00983 |
| 10% error | ![]() |
0.01956 | 0.4968 | 0.01413 |
Baseline models: TG-only strategies
TTo contextualize the performance of the multivariable ML framework, two TG-only baseline strategies were evaluated using the same dataset. First, a simple rule-based classifier using the training 75th percentile of
(189.2675 mg/dL) as a cut-off achieved an AUROC of 0.8204 and a Brier score of 0.1347. Second, a univariate logistic regression model using
as the sole predictor achieved strong discrimination (AUROC 0.9167, Brier score 0.1056) with an F1-score of 0.6120, recall of 0.4480, and precision of 0.9655.
These baselines confirm that fasting TG alone contains substantial predictive information about high postprandial response. Notably, the multivariable L2-penalized logistic regression model achieved comparable discrimination (AUROC 0.9141) with improved probabilistic accuracy (Brier score 0.0886) and a more balanced trade-off between recall and precision.
Performance of the leakage-controlled multivariable model
Within the leakage-controlled nested
cross-validation framework at the primary
75th percentile phenotype, the final multivariable L2-penalized logistic regression model achieved a mean AUROC of 0.9141 with a standard deviation of 0.0199. The corresponding mean Brier score was 0.0886 ± 0.0089. Classification metrics at this operating point included an F1-score of 0.7348, recall of 0.6613, and precision of 0.8267 (Table 6), indicating strong discrimination and good probability accuracy under a class balance of 25% high responders. Bootstrap evaluation (1,000 resamples) yielded a mean AUROC of 0.9140 (95% CI: 0.8957–0.9314) and a mean Brier score of 0.0887 (95% CI: 0.0789–0.0983), confirming stability of discrimination and calibration under resampling variability.
When evaluated using held-out calibration procedures, isotonic regression achieved an AUROC of 0.9221 with a Brier score of 0.0856, whereas Platt (sigmoid) calibration produced an AUROC of 0.9252 with a Brier score of 0.0839. Overall, post-hoc calibration preserved discrimination while maintaining strong probability accuracy. Minor differences in AUROC after calibration reflect probability re-mapping effects introduced by isotonic regression, which can slightly modify ranking near threshold regions; however, discrimination remained statistically comparable to the nested cross-validation estimates. Across repeated cross-validation folds, isotonic regression also demonstrated superior unweighted composite calibration performance compared with Platt scaling, supporting its selection as the final specification.
Calibration, discrimination curves, and clinical utility
Receiver operating characteristic and precision–recall curves for the isotonic-calibrated final model are shown in Fig. 5. The ROC curve demonstrates consistently high discrimination across thresholds, while the precision–recall curve confirms robust performance under the observed class imbalance. Calibration assessment is summarized in Fig. 6, where the isotonic calibration curve shows close agreement between predicted probabilities and observed event rates across deciles of risk, indicating no substantial systematic over- or underestimation of risk.
Fig. 5.

Discrimination performance of the final model (
top quartile). (a) ROC (AUROC
). (b) Precision–recall (AP
,
prevalence).
Fig. 6.

Calibration of the final model: (a) predicted vs. observed probabilities; (b) slope
1.0 and intercept
0.0, indicating good calibration.
Decision curve analysis (Fig. 7, panel (a)) demonstrates that the isotonic-calibrated multivariable model provides higher net clinical benefit than both treat-all and treat-none strategies, as well as TG-only baselines, across a wide range of clinically plausible decision thresholds. Threshold sensitivity analysis across
phenotype cut-offs from the 60th to the 90th percentile (Fig. 7, panel (b)) shows that AUROC remains consistently high (approximately 0.914–0.923), supporting the robustness of the model and conclusions to alternative phenotype definitions. Comparative calibration plots for Platt scaling demonstrated similar discrimination but modestly higher expected calibration error and are provided in the Supplementary Material.
Fig. 7.

Clinical utility and robustness. (a) Higher net benefit vs. baseline strategies. (b) Stable AUROC across
thresholds (60th–90th percentiles).
Threshold sensitivity across phenotype definitions
Because the definition of the high-response phenotype directly influences class balance and physiological interpretation, model performance was evaluated across multiple
percentile thresholds (60th–90th percentiles). For each threshold, the full leakage-controlled nested cross-validation pipeline was re-executed to ensure unbiased estimation.
Table 7 reports the predictive performance of the multivariable L2-penalized logistic regression model across alternative
cut-offs, including AUROC, Brier score, F1-score, recall, and precision (averaged across nested cross-validation folds).
Table 7.
Sensitivity analysis across phenotype thresholds. Predictive performance of the multivariable L2-penalized logistic regression model under alternative
percentile cut-offs. Values are averaged across nested cross-validation folds.
| Percentile | Cutoff (mg/dL) | AUROC | Brier | F1-score | Recall | Precision |
|---|---|---|---|---|---|---|
| 60% | 261.7900 | 0.9218 | 0.1059 | 0.7989 | 0.7483 | 0.8569 |
| 70% | 301.0580 | 0.9226 | 0.0924 | 0.7862 | 0.7356 | 0.8444 |
| 75% | 329.2875 | 0.9141 | 0.0886 | 0.7348 | 0.6613 | 0.8267 |
| 80% | 365.2400 | 0.9191 | 0.0767 | 0.7032 | 0.6200 | 0.8122 |
| 90% | 458.1810 | 0.9202 | 0.0540 | 0.4563 | 0.3133 | 0.8393 |
Discrimination remained stable across phenotype definitions (AUROC range 0.9141–0.9226), indicating robustness of ranking performance to alternative
thresholds. As the percentile threshold increased and event prevalence decreased, recall and F1-score declined (most pronounced at the 90th percentile), while AUROC remained preserved. Brier scores decreased progressively (0.1059 to 0.0540), reflecting the combined effects of reduced event prevalence and maintained discrimination under increasing class imbalance.
Across all thresholds,
remained the dominant predictor, whereas WBV showed consistently negligible contribution. The stability of discrimination and predictor patterns across phenotype definitions supports the robustness of the modeling framework and provides convergent evidence that WBV does not meaningfully contribute to postprandial triglyceride response within the synthetic setting.
Robustness and stability analyses
Model robustness was further examined using bootstrap resampling and repeated cross-validation. In 1,000 bootstrap resamples, the AUROC distribution had a mean of 0.9140 with a 95% percentile-based confidence interval of [0.8957, 0.9314], while the Brier score had a mean of 0.0887 with a 95% confidence interval of [0.0789, 0.0983] (Fig. 3). Both distributions were narrow, indicating stable discrimination and probability accuracy across resampled cohorts.
Fig. 3.

Bootstrap distributions (1,000 resamples) for the final model. Panel (a) shows the stability of discrimination (AUROC), while panel (b) reflects the stability of probability accuracy (Brier score). Both distributions exhibit narrow variance, indicating high model reliability.
Repeated stratified 5-fold cross-validation over 10 repeats (50 folds total) yielded an AUROC of 0.9141 ± 0.0156 (minimum 0.8803, maximum 0.9455) and a Brier score of 0.0886 ± 0.0073 (minimum 0.0722, maximum 0.1017). These results confirm that the model’s performance is robust to variation in train–test splits and is not driven by any single partition of the data.
Explainability: feature importance and marginal effects
Global SHAP analysis (Fig. 8, panel (a)) consistently ranked fasting triglycerides (
) and BMI as the dominant drivers of model predictions, whereas WBV remained tightly clustered near zero SHAP values, indicating minimal relevance to risk stratification. The SHAP dependence plot (Fig. 8, panel (b)) illustrates a clear monotonic and nonlinear rise in predicted risk with increasing fasting triglycerides. One-dimensional partial dependence curves (Fig. 10) similarly highlight strong marginal effects for
and BMI, while the WBV curve remains essentially flat.
Fig. 8.

Explainability of the final model using SHAP. Panel (a) shows the global SHAP summary plot, where fasting triglycerides (
) and BMI dominate predictive influence and WBV contributes minimally. Panel (b) presents the SHAP dependence plot for
, revealing a strong monotonic increase in predicted risk with higher fasting triglyceride levels.
Fig. 10.

One-dimensional partial dependence plots illustrating marginal effects of individual predictors on the predicted probability of high postprandial response.
High-resolution partial dependence (Fig. 11) reveals an even sharper sigmoidal transition across
values, reinforcing its dominant influence on phenotype classification. SHAP interaction analysis (Fig. 9, panel (a)) shows no meaningful
WBV interaction, whereas the two-dimensional partial dependence surface (Fig. 9, panel (b)) demonstrates a synergistic increase in risk among individuals with both elevated fasting triglycerides and higher BMI. Together, these explainability results confirm that baseline metabolic load, not hemorheological viscosity, is the primary determinant of high postprandial response.
Fig. 11.

High-resolution one-dimensional partial dependence curve for fasting triglycerides (
). The curve reveals a near-zero risk plateau at low TG levels, followed by a steep sigmoidal-like rise in predicted probability as
reaches intermediate and high ranges, reinforcing its dominant and non-linear role in shaping the postprandial response phenotype.
Fig. 9.

Interaction and joint effect analyses. Panel (a) shows the SHAP interaction plot between
and WBV, indicating no meaningful interaction and reinforcing the negligible role of WBV in determining the outcome. Panel (b) displays the two-dimensional partial dependence of
and BMI, illustrating that higher
and higher BMI jointly increase the predicted risk of high postprandial response.
High-resolution partial dependence analysis for fasting triglycerides
To refine the characterization of the non-linear influence of fasting triglycerides, we generated a high-resolution one-dimensional partial dependence curve for
(Fig. 11). This analysis provides a more granular mapping of how incremental changes in baseline triglyceride concentration translate into predicted risk within the multivariable model, revealing a near-zero risk plateau at low levels followed by a steep sigmoidal-like rise as
reaches intermediate and high ranges.
Correlation of SHAP-based feature contributions
To further understand how predictor effects co-vary within the model’s decision function, we computed a SHAP value correlation matrix (Fig. 12). This analysis captures the degree to which feature attributions rise or fall together across the cohort, providing insight into redundancy, independence, or complementary effects. WBV exhibited weak or absent correlation with lipid-related SHAP contributions, reinforcing its minimal role in determining the outcome, whereas TG- and BMI-related SHAP values showed more coherent patterns, consistent with their dominant predictive influence.
Fig. 12.

Correlation matrix of SHAP values across features. WBV exhibits weak or absent correlation with lipid-related SHAP contributions, reinforcing its minimal role in determining the outcome. TG- and BMI-related SHAP values show more coherent patterns, consistent with their dominant predictive influence.
Ablation analysis
To further evaluate the mechanistic contributions of key predictors, a feature-ablation study was performed (Table 8). Removing
resulted in a substantial drop in discrimination (AUROC = 0.6002), confirming its dominant predictive role. In contrast, removing BMI produced only a minimal change relative to the full multivariable model (AUROC = 0.9132 vs. 0.9141), and removing WBV had no meaningful impact on discrimination (AUROC = 0.9146). These findings reinforce that baseline triglyceride load, rather than viscosity-related factors, is the primary driver of postprandial response classification.
Table 8.
Ablation study of key predictors.
| Feature removed | AUROC |
|---|---|
removed |
0.6002 |
| BMI removed | 0.9132 |
| WBV removed | 0.9146 |
Synthetic domain-shift stress test and subgroup stability analysis
Fig. 13.

Synthetic domain-shift stress-test results. Panel (a) shows discrimination performance on the independent external synthetic cohort. Panel (b) displays the external calibration curve under stress-test temperature scaling.
To evaluate ranking and calibration stability under a controlled synthetic domain shift, we conducted an external-synthetic stress-test experiment using an independently generated synthetic cohort. This analysis is presented solely as a simulation-based robustness check and is analytically separate from the primary nested cross-validation and calibration framework. On the independent external synthetic cohort, the leakage-controlled logistic regression model preserved strong ranking performance (AUROC = 0.9236) with a Brier score of 0.0994. Relative to the internal nested cross-validation Brier score (0.0886), the modest increase is consistent with expected probability dispersion under a perturbed synthetic data-generating process.
For diagnostic purposes, we additionally report post hoc temperature scaling (T = 0.9913) applied to the external predictions. The temperature parameter was estimated using out-of-fold cross-validation on the original training cohort and was subsequently applied to the external synthetic cohort without any refitting. After scaling, discrimination remained essentially unchanged (AUROC = 0.9225) and probability accuracy did not improve (Brier score = 0.1028), suggesting limited benefit of additional recalibration under this specific synthetic shift scenario. Because the phenotype label is deterministically derived from
under a low-noise copula-based simulation, these results should not be interpreted as estimates of real-world transport performance. Rather, this stress test provides a controlled check of model ranking and calibration behavior within the synthetic framework prior to real-world validation.
Summary of key findings
Across all analytic layers, descriptive, statistical, machine-learning, and explainability, results consistently indicate absence of detectable association between WBV and postprandial triglyceride response within the reconstructed synthetic cohort. WBV showed no differences between response groups, no meaningful correlation with TGR or triglyceride levels (with confidence intervals crossing zero), minimal SHAP contribution, and no detectable interaction with
in SHAP or partial-dependence analyses.
In contrast,
and BMI emerged as the dominant predictors, exhibiting strong marginal effects and synergistic influence on high-response risk. The leakage-controlled L2-penalized logistic regression model achieved robust and well-calibrated performance (nested AUROC
; bootstrap AUROC
[0.90–0.93]) with stable performance across phenotype definitions and resampling procedures within the synthetic cohort.
Overall, within the assumptions and correlation structure encoded in the synthetic data-generating process, baseline metabolic load appears to be the primary driver of variation in the postprandial response phenotype, whereas hemorheological viscosity provides no detectable incremental contribution.
Discussion
Summary of findings
This study presents a rigorously leakage-controlled, multi-layer analytical demonstration of how candidate biomarker associations can be evaluated for postprandial triglyceride response in a privacy-preserving setting. Across descriptive, inferential, machine-learning, calibration, and robustness analyses, whole blood viscosity (WBV) did not exhibit reproducible predictive or explanatory contribution within the present framework. Importantly, this finding is conditional on the explicitly encoded marginal distributions and correlation structure of the synthetic cohort. The absence of association should therefore be interpreted within the bounds of the defined generative assumptions and does not constitute proof of absence of biological effect in vivo.
Within this statistically reconstructed synthetic cohort, WBV exhibited negligible predictive and explanatory value for the postprandial response phenotype across both linear and nonlinear evaluations, and no meaningful interaction with triglyceride-related features was observed in the fitted models. In contrast, fasting triglycerides (TG
) consistently emerged as the dominant signal, supported by phenotype separation analyses, SHAP dependence plots, one- and two-dimensional partial dependence functions, and multi-model comparisons. BMI was the second most influential predictor, with interaction structure observed in TG
BMI surfaces.
The leakage-controlled L2-penalized logistic regression model achieved stable and well-calibrated discrimination:
Nested CV AUROC
,Bootstrap AUROC
[0.8957–0.9314],Brier score
,
with stability further supported by narrow bootstrap distributions and the repeated cross-validation AUROC range (0.8803–0.9455). An additional robustness visualization, the distribution of AUROC across 5
10 repeated cross-validation, is shown in Fig. 14.
Fig. 14.

Distribution of AUROC across 5
10 repeated cross-validation. The distribution is tightly centered around 0.9141, with no heavy tails, indicating strong robustness and absence of overfitting to any single split.
The use of a statistically reconstructed synthetic cohort provides practical advantages for methodological evaluation: privacy by design, full reproducibility, and the ability to stress-test modeling choices and phenotype definitions under controlled assumptions. Accordingly, the primary contribution of this work is to demonstrate a reproducible leakage-controlled ML workflow for metabolic phenotyping, while interpreting association patterns strictly within the synthetic framework rather than as population-level clinical inference.
Framework-level interpretation within the synthetic model
These patterns should be interpreted within the constraints of a statistically reconstructed synthetic cohort and a surrogate WBV estimate derived from the de Simone formulation rather than direct viscometric measurement. Within this modeled setting, the convergence of statistical and ML-based analyses suggests that WBV does not contribute reproducible predictive or explanatory information for short-term postprandial TG response under the viscosity ranges represented in the cohort (approximately 5–7 cP).
The negligible WBV–TGR association (e.g.,
,
) contrasts with hypotheses that increased viscosity may impair microvascular perfusion and thereby modulate lipid dynamics. In the present framework, the results are more consistent with the following interpretation:
WBV behaves largely orthogonally to triglyceride-response phenotypes within the modeled physiologic range.
The response phenotype is more strongly aligned with baseline metabolic load (TG
) and adiposity-related factors (BMI) than with hemorheological variation.Interaction analyses (e.g., TG
WBV) do not indicate robust synergistic or antagonistic structure within the fitted models.
Importantly, these statements are not intended as definitive physiological claims. Rather, they summarize the behavior observed under the synthetic reconstruction and highlight where real-world validation including direct WBV measurements and diverse clinical cohorts is required to determine whether the same independence holds in practice. The conceptual diagrams (Figs. 15 and 16) summarize this framework-level interpretation by depicting WBV as a hemodynamic axis that does not exhibit reproducible coupling to the modeled postprandial triglyceride-response pathway.
Fig. 15.

Mechanistic schematic of postprandial lipid metabolism. Fasting triglycerides, amplified by BMI-related adiposity and insulin resistance, dominate postprandial processing, while WBV remains an independent hemodynamic axis with no observable contribution to TGR within the modeled synthetic framework.
Fig. 16.
Systems-level integration of triglycerides, BMI, and whole blood viscosity (WBV) in determining postprandial triglyceride response (TGR). TG
and BMI constitute the dominant predictive axis within the synthetic framework, whereas WBV appears as an orthogonal hemodynamic axis with negligible contribution within this modeled setting.
Methodological and computational significance
The study’s leakage-controlled design provides strong methodological advances for clinical ML research:
Fold-specific preprocessing: all scaling, imputation, and resampling procedures were performed strictly within each training fold to eliminate data leakage and preserve true generalization behavior.
Nested 5
5 cross-validation: the nested design provided unbiased estimates of out-of-sample performance while allowing principled hyperparameter tuning.Rigorous calibration assessment: systematic comparison of Platt scaling and isotonic regression. Although discrimination performance was nearly identical between Platt scaling and isotonic regression, isotonic regression demonstrated lower expected calibration error (ECE; 10-bin, unweighted) and superior overall composite calibration performance across repeated cross-validation (based on an unweighted aggregate of AUROC, Brier score, ECE, calibration slope, and intercept). Because the primary objective of this study was accurate absolute risk estimation for clinical decision-making rather than ranking performance alone, isotonic regression was selected as the final calibration specification.
Multi-layer robustness testing: including bootstrap resampling (1,000 iterations), repeated 5
10 cross-validation, threshold sensitivity analyses, and comparison against multiple baseline classifiers.Explainability integration: interpretability was achieved through SHAP summary plots, SHAP dependence and interaction analyses, and both one-dimensional and two-dimensional partial dependence plots.
Demographic pseudo-external-synthetic transfer: generalizability was evaluated across subgroups stratified by age and sex, simulating an external-synthetic transfer scenario in the absence of multi-center data.
This framework directly addresses common ML pitfalls, such as data leakage, uncalibrated probabilities, and overestimation of performance, which are frequently highlighted in contemporary critiques of biomedical AI literature14,23,24.
Comparison with prior literature
Prior studies examining viscosity–metabolism relationships have reported mixed findings and are often limited by modest sample sizes, heterogeneous populations, and simplified analytical approaches (e.g., unadjusted correlations)6,10. Some reports focus on high-viscosity pathological states (e.g., polycythemia vera or severe dyslipidemia), where hemorheological effects may plausibly be more pronounced25,26. In general, few studies have combined modern ML evaluation practices including strict leakage control, calibration diagnostics, robustness testing, and interpretability when assessing whether WBV provides independent value for postprandial lipid phenotypes.
In this context, the present work contributes primarily by offering a reproducible, leakage-controlled evaluation template and by situating the WBV question within a multi-layer analytical protocol. Within the synthetic reconstruction used here, WBV does not provide additional predictive or explanatory value for the response phenotype after accounting for TG
and BMI. Whether similar patterns hold in real clinical populations remains an open question that requires external validation with direct WBV measurements and multi-center cohorts.
Convergence of statistical and machine-learning evidence
A notable strength of this work is the convergence of evidence across statistical, inferential, and machine-learning analyses. Correlation testing, linear modeling, SHAP-based explainability, partial-dependence surfaces, and repeated resampling procedures produced mutually consistent patterns within the synthetic cohort: WBV did not exhibit reproducible predictive or explanatory contribution to the postprandial response phenotype, whereas fasting triglycerides consistently dominated model behavior.
Such multi-layer agreement reduces the likelihood that the observed null pattern is an artefact of any single modeling choice and strengthens confidence in the internal validity of the methodological demonstration. Nevertheless, agreement within a synthetic reconstruction does not by itself establish physiological truth, and real-world validation remains necessary before clinical inference.
More broadly, the alignment between domain expectations (e.g., the central role of baseline lipid burden) and model-derived attribution patterns supports the use of explainability as a consistency check within leakage-controlled workflows, while remaining conservative about causal interpretation.
Clinical and translational implications
Because this study is conducted in a statistically reconstructed synthetic cohort, its implications should be interpreted primarily as methodological guidance rather than direct clinical recommendation. Within this framework, several practical points emerge:
Baseline fasting triglycerides (and to a lesser extent BMI) provide the most reproducible signal for the modeled postprandial response phenotype.
WBV did not provide additional predictive or explanatory value beyond routine fasting biomarkers within the modeled viscosity range, suggesting that WBV is not a promising candidate feature in this synthetic setting.
Multivariable ML provided improved probability calibration relative to TG-only logistic regression, while discrimination remained comparable.
Well-calibrated probability estimates in the internal evaluation support the feasibility of calibrated risk scoring within the modeled setting, but deployment-oriented claims require validation on real cohorts with prospective evaluation.
From a translational perspective, these results primarily reinforce the importance of rigorous evaluation practice (leakage control, calibration, robustness checks) when proposing biomarker-based phenotyping models. Future work should test whether the same feature hierarchy and calibration behavior hold in independent real-world datasets with direct WBV measurements before any screening or decision-support interpretation.
Integration outlook and systems-level implications
The results suggest a synthetic systems-level representation in which:
TG metabolism was primarily modeled as reflecting hepatic and adipose biochemical pathways within the synthetic framework.
Hemorheological factors such as WBV did not demonstrate measurable short-term influence within the synthetic model under physiological parameter ranges.
BMI interacted with TG
within the model in a manner consistent with established conceptual mechanisms of insulin resistance and adipose overflow.
Generalizable leakage-controlled clinical ML framework
Although this work focuses on triglyceride-response phenotyping, the methodological contribution extends more broadly. The proposed leakage-controlled ML framework, combining fold-specific preprocessing, nested cross-validation, calibrated probability estimation, bootstrap resampling, and synthetic domain-shift evaluation, is fully generalizable to a wide range of clinical ML tasks. These include cardiometabolic phenotyping, EHR-based risk prediction, postprandial response modeling, and any workflow in which leakage risks, class imbalance, or distributional shifts threaten model validity27,28. The reproducible structure of this framework makes it suitable as a template for future clinical informatics research where methodological rigor is required before deployment on real-world cohorts.
Strengths and limitations
Strengths:
A comprehensive leakage-controlled ML workflow including nested CV, repeated CV, bootstrap resampling, calibration assessment, threshold sensitivity, and explainability analyses.
A privacy-preserving, statistically reconstructed synthetic cohort enabling fully reproducible experimentation without identifiable patient data.
Multi-angle consistency checks (statistical testing, ML performance, calibration diagnostics, and explainability) to reduce single-method artefacts.
Demographic pseudo-external evaluation across age and sex strata to assess subgroup stability under controlled synthetic conditions.
Limitations:
The study is based on a statistically reconstructed synthetic cohort rather than real EMR data; synthetic data cannot fully capture physiological and biochemical heterogeneity, clinical measurement processes, or unmeasured confounding present in real populations.
WBV was estimated using the de Simone surrogate instead of direct viscometric measurement; although sensitivity analyses under
–
perturbations were consistent within the framework, model-based estimation may not reflect all hemorheological conditions.The WBV range represented (approximately 5–7 cP) may not capture extreme metabolic or hemorheological states where viscosity effects could plausibly become clinically relevant.
No external validation on an independent real-world clinical cohort was performed, and the study explicitly remains confined to a synthetic modeling framework. While demographic pseudo-external analyses and synthetic domain-shift stress-test experiments support methodological robustness under controlled shifts, extension to real patient populations requires future multi-center validation with direct WBV measurements.
Overall, the analyses demonstrate internally consistent model behavior within the synthetic framework; however, extension to real-world clinical populations requires independent validation in multi-center cohorts with direct physiological measurements.
Measurement-error considerations
WBV estimation is subject to model-based error from hematocrit and total protein measurements. However:
Measurement error would typically attenuate associations rather than create a spurious null or inverse effect.
SHAP dependence plots show WBV-centered tightly around zero across the entire cohort.
Thus, measurement noise is highly unlikely to mask a clinically meaningful biological effect.
Generalizability and reuse
Despite limitations, generalizability is supported by:
consistent subgroup AUROC (0.7700–0.8470),
stable pseudo-external-synthetic transfer across sex and age splits,
threshold-insensitive discrimination (
–0.92 across 60–80th percentile cutoffs).
All preprocessing and modeling code is openly available, enabling reproducibility and reuse in future cohorts.
Methodological outlook
Future work could integrate:
multi-omics layers (proteomics, metabolomics),
genetic variants related to TG metabolism,
direct viscometry to validate WBV independence more rigorously,
multi-task learning frameworks for joint modeling of TGR, hepatic output, and insulin sensitivity.
Future work should incorporate direct viscometry-based WBV measurements to validate the null-effect hypothesis under more precise hemorheological conditions. Such integration would help determine whether the observed independence reflects true physiology or the limits of model-based WBV estimation.
Reproducibility and methodological value
A central contribution of this work lies in demonstrating that leakage-controlled ML can produce trustworthy, clinically coherent results in metabolic phenotyping, contrasting with many prior studies where methodological flaws inflate performance or generate spurious findings13,14. The complete reproducibility of all steps, dataset structure, preprocessing rules, nested CV design, and explainability workflows provides a template for other researchers designing ML pipelines for biomedical applications29,30.
Implications for future ML phenotyping frameworks
Beyond the domain-specific findings, this study provides generalizable insights for the design of machine-learning phenotyping pipelines in biomedical research. First, the results demonstrate that leakage-controlled workflows, incorporating fold-specific preprocessing, nested cross-validation, and calibrated probability estimation, can reliably separate true biological signals from spurious correlations that commonly arise in clinical datasets14,31. Second, the tight alignment between model-derived SHAP patterns18 and established physiological mechanisms highlights the value of integrating domain knowledge with explainability tools to validate model-driven hypotheses. Third, the robustness of predictive behavior across multiple phenotype thresholds suggests that future phenotyping frameworks should explicitly include threshold-sensitivity analyses to guard against instability or definition-dependent biases32.
Taken together, these elements outline a reproducible blueprint for next-generation ML phenotyping: one that emphasizes methodological transparency, physiological coherence, and multi-angle robustness checks as core components of reliable biomedical discovery.
Conclusions
This study presents a comprehensive, leakage-controlled methodological demonstration for evaluating candidate biomarker associations in postprandial triglyceride-response phenotyping. Using a statistically reconstructed synthetic cohort and an analysis pipeline integrating statistical testing, multivariable machine learning (final model: L2-penalized logistic regression), calibration diagnostics, repeated resampling for robustness, and model explainability, we observed a consistent framework-level null pattern for whole blood viscosity (WBV) with respect to the postprandial response phenotype. Within the modeled setting, WBV showed negligible linear association with TGR, minimal predictive contribution in SHAP analyses, flat marginal effects in partial-dependence curves, and no reproducible interaction structure with triglyceride-related features. These patterns were stable across pseudo-external demographic strata, phenotype thresholds, and repeated resampling schemes. Postprandial measurements (TG
) and derived response metrics (TGR) were used exclusively for phenotype definition and were never included as predictors.
In contrast, fasting triglycerides (TG
) and BMI emerged as the dominant signals governing model behavior, exhibiting stable monotonic effects and coherent interaction structure within the synthetic framework. The alignment between these attribution patterns and established metabolic intuition supports the internal consistency of the modeling results, while remaining conservative about causal or physiological interpretation.
Beyond domain-specific observations, the primary contribution of this work is methodological. The combination of fold-specific preprocessing, nested cross-validation, probability calibration, threshold-sensitivity evaluation, pseudo-external demographic validation, and controlled external-synthetic domain-shift stress testing provides a reproducible blueprint for reducing leakage, overfitting, and miscalibration common failure modes in clinical ML studies. Overall, this work illustrates how privacy-preserving synthetic cohorts can be used to stress-test analytical assumptions and evaluation practices for metabolic phenotyping. However, findings from a synthetic reconstruction should be regarded as hypothesis-generating rather than definitive clinical evidence. Future studies using direct viscometry and independent real-world multi-center cohorts will be required to establish external validity and to determine whether viscosity-related effects emerge under broader hemorheological conditions.
Declaration of generative AI and AI-assisted technologies in the writing process
ChatGPT (OpenAI) was used solely to assist with English language editing and improvement of clarity based on author-provided text. The authors retain full responsibility for the study design, data generation, analysis, interpretation, and all final conclusions.
Data provenance
No individual-level clinical records were used in this study. The synthetic dataset was generated exclusively from non-identifiable, summary-level aggregate statistics (e.g., means, variances, percentiles, and correlation matrices) previously derived from a fully de-identified hospital reference cohort. These aggregated statistics contain no re-identifiable information. No electronic medical record (EMR) access or raw patient-level data were involved at any stage of the research. All modeling and analysis were performed entirely on fully synthetic, statistically reconstructed datasets.
Technical note
All machine-learning visualizations (SHAP plots, partial-dependence functions, calibration curves, robustness analyses) were produced using the fold-specific preprocessing pipeline described in the Methods. Feature names appearing in figures reflect internal variable identifiers used during model training and do not represent additional variables beyond those explicitly defined in the manuscript.
This research was supported by the National Science and Technology Council (NSTC), Taiwan, under Grant No. NSTC114-2221-E-155-016-MY2, NSTC114-2637-8-155-003-.
Supplementary Material
Supplementary analyses, including nested cross-validation results, calibration diagnostics, SHAP-based explainability outputs, and risk-score tables, are available in the results branch of the repository.
Acknowledgements
The authors would like to thank the International Bachelor’s Program in Informatics and the Department of Computer Science and Engineering at Yuan Ze University for their academic support and for providing a conducive research environment. This research was supported by the National Science and Technology Council (NSTC), Taiwan, under Grant No. NSTC114-2221-E-155-016-MY2, NSTC114-2637-8-155-003-.
Abbreviations
- AI
Artificial intelligence
- AP
Average precision
- AUC
Area under the curve
- AUROC
Area under the receiver operating characteristic curve
- BMI
Body mass index
- CI
Confidence interval
- CV
Cross-validation
- DCA
Decision curve analysis
- ECE
Expected calibration error
- F1
F1-score
- FN
False negative
- FP
False positive
- GB
Gradient boosting
- Hct
Hematocrit
- HDL
-C High-Density Lipoprotein Cholesterol
- IRB
Institutional review board
- KDE
Kernel density estimate
- L2
L2-penalty (ridge)
- LDL-C
Low-density lipoprotein cholesterol
- LR
Logistic regression
- ML
Machine learning
- OOF
Out-of-fold
- PDP
Partial dependence plot
- PR-AUC
Precision–recall area under the curve
- ROC
Receiver operating characteristic
- SD
Standard deviation
- SHAP
SHapley additive explanations
- STROBE
Strengthening the reporting of observational studies in epidemiology
- TG
Triglycerides
- TG
Fasting triglycerides
- TG
4-hour postprandial triglycerides
- TGR
Relative postprandial triglyceride response
- TN
True negative
- TP
Total protein
- TRL
Triglyceride-rich lipoproteins
- TRIPOD
Transparent reporting of a multivariable prediction model
- WBV
Whole blood viscosity
- XAI
Explainable artificial intelligence
- cP
Centipoise (unit of viscosity)
Appendix
Computational Environment
All analyses were conducted in a fully version-controlled environment to ensure complete reproducibility. System specifications are listed below:
Operating system: Ubuntu 22.04 LTS (64-bit)
Python version: 3.12.4 Core libraries: NumPy 2.0, SciPy 1.13, pandas 2.2, scikit-learn 1.5, XGBoost 2.1, imbalanced-learn 0.12
R environment: R 4.3.1 for independent cross-validation verification
Global random seed: 42 for all preprocessing, model training, calibration, bootstrap resampling, and repeated CV
Hardware: Intel Core i7 (8 cores), 16 GB RAM
All scripts, including data reconstruction, preprocessing, nested cross-validation, probability calibration, robustness testing, and SHAP-based explainability, were managed using Git version control and can be fully reproduced using the public repository accompanying this manuscript.
Hyperparameter Search Space and Validation Settings
The hyperparameter search space and validation configuration used in the nested cross-validation are summarized in Table 9.
Table 9.
Hyperparameter grid and validation configuration for nested cross-validation (rerun-consistent).
| Model / Component | Search Space / Setting |
|---|---|
| Preprocessing | Fold-specific winsorization (1st–99th percentile) and standardization applied within each training fold only (strict leakage control). |
| Logistic Regression (L2) | C: {0.01, 0.1, 1.0, 10.0}; penalty: {l2}; solver: {lbfgs}; max_iter: 8000; class_weight: balanced. |
| Random Forest | n_estimators: 600 (fixed); max_depth: {None, 3, 5, 8}; min_samples_leaf: {1, 2, 5}; class_weight: balanced. |
| SVM (RBF kernel) | C: {0.5, 1.0, 2.0, 4.0}; gamma: {scale, auto}; class_weight: balanced; probability enabled. |
| Calibration | CalibratedClassifierCV applied using inner CV on the training portion of each outer fold only (no access to outer test fold). |
| Cross-validation | Nested CV: 5-fold outer loop with 5-fold inner loop (stratified, shuffled). |
| Primary Metric | ROC-AUC; PR-AUC (AP) and Brier score reported secondarily. |
| Random Seeds | Global seed: 42; outer CV random_state = 42; inner CV random_state = 42 + 100 + fold. |
Summary Table
An overall numerical summary of the rebuilt physiologically coherent synthetic cohort, correlation structure, validation dimensions, and model performance is presented in Table 10.
Table 10.
| Domain | Quantity | Value |
|---|---|---|
| Cohort & phenotype | Total synthetic adult records | ![]() |
| Normal responders |
(75.0000%) |
|
| High responders |
(25.0000%) |
|
Primary TG cut-off (75th percentile) |
![]() |
|
| Physiological plausibility thresholds (none exceeded) | TG ; WBV or cP |
|
| Baseline characteristics | Female, n (%) | Normal: 587 (52.1778%); High: 198 (52.8000%);
|
| Male, n (%) | Normal: 538 (47.8222%); High: 177 (47.2000%);
|
|
| Age (years) | Normal: ; High: ;
|
|
| Hematocrit (%) | Normal: ; High: ;
|
|
| Total protein (g/dL) | Normal: ; High: ;
|
|
| Whole blood viscosity (cP) | Normal: ; High: ;
|
|
Fasting TG (TG , mg/dL) |
Normal: ; High: ;
|
|
4-h TG (TG , mg/dL) |
Normal: ; High: ;
|
|
| Postprandial TG response (TGR, %) | Normal: ; High: ;
|
|
| HDL-C (mg/dL) | Normal: ; High: ;
|
|
| LDL-C (mg/dL) | Normal: ; High: ;
|
|
BMI (kg/m ) |
Normal: ; High: ;
|
|
| Key correlations | TG – TG
|
![]() |
| Hct – WBV | ![]() |
|
| TP – WBV | ![]() |
|
WBV – TG
|
![]() |
|
WBV – TG
|
![]() |
|
| Bootstrap CI (1000 resamples) | WBV – TG
|
(95% CI ) |
WBV – TG
|
(95% CI ) |
|
| WBV – TGR |
(95% CI ) |
|
| WBV misspecification sensitivity | 5% relative error | corr mean (SD ); AUC mean (SD ) |
| 10% relative error | corr mean (SD ); AUC mean (SD ) |
|
|
Primary model performance (75th percentile phenotype) |
TG cut-off (189.2675 mg/dL) |
AUROC ; Brier
|
Univariate logistic regression (TG only) |
AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| Multivariable L2 logistic regression | AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| Phenotype threshold sensitivity | 60th percentile (261.7900 mg/dL) | AUROC ; Brier ; F1 ; Recall ; Precision
|
| 70th percentile (301.0580 mg/dL) | AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| 75th percentile (329.2875 mg/dL) | AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| 80th percentile (365.2400 mg/dL) | AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| 90th percentile (458.1810 mg/dL) | AUROC ; Brier ; F1 ; Recall ; Precision
|
|
| Feature ablation | Remove TG
|
AUROC
|
| Remove BMI | AUROC
|
|
| Remove WBV | AUROC
|
|
| Modeling configuration | Logistic regression (L2) |
; solver lbfgs; max_iter ; class_weight balanced |
| Random forest | n_estimators ; max_depth ; min_samples_leaf
|
|
| SVM (RBF) |
;
|
|
| Nested cross-validation | Stratified 5 5 nested CV; primary metric ROC-AUC; global seed
|
Author contributions
- N. Piyavechvirat: Conceived the study, generated the synthetic datasets, implemented all preprocessing and modeling workflows, conducted statistical analyses, created all figures and tables, and drafted the main manuscript. - Y. Jheng Huang: advised on study design and contributed to manuscript review and revisions. - Q. Mazhar ul Haq: supervised the project, provided critical feedback on model evaluation and interpretation, and contributed to manuscript review and revisions. All authors read and approved the final manuscript.
Funding
This work was supported by the National Science and Technology Council (NSTC), Taiwan, under Grant. No. NSTC114-2221-E-155-016-MY2, NSTC114-2637-8-155-003-.Institutional support was provided by the International Bachelor Program in Informatics and the Department of Computer Science and Engineering, Yuan Ze University, Taiwan.
Data Availability
All datasets used in this study are fully synthetic and contain no identifiable or linkable human-subject information. The complete machine-learning code, synthetic data generators, and all reproducible workflow scripts are publicly available at:https://github.com/NattakittiP/ML_Predict
Declarations
Competing interests
The authors declare no competing interests.
Ethical Approval
This study relied exclusively on statistically reconstructed synthetic data generated from non-identifiable aggregate statistics. Because no identifiable human data were accessed, stored, or analyzed, the study does not constitute human-subject research under prevailing ethical or regulatory definitions. Institutional review board (IRB) approval was therefore not required.
Human and Animal Rights
No human or animal subjects were involved in this research. All analyses were conducted on fully synthetic datasets containing no identifiable clinical information.
Consent for Publication
Not applicable. This manuscript contains no identifiable personal data.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Kolovou, G. et al. Assessment and clinical relevance of non-fasting and postprandial triglycerides: An expert panel statement. Current Vascular Pharmacol.9, 258–270. 10.2174/157016111795495549 (2011). [DOI] [PubMed] [Google Scholar]
- 2.Keirns, B. H., Sciarrillo, C. M., Koemel, N. A. & Emerson, S. R. Fasting, non-fasting and postprandial triglycerides for screening cardiometabolic risk. J. Nutrit. Sci.10, e75. 10.1017/jns.2021.73 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zilversmit, D. B. Atherogenesis: A postprandial phenomenon. Circulation60, 473–485. 10.1161/01.CIR.60.3.473 (1979). [DOI] [PubMed] [Google Scholar]
- 4.Nordestgaard, B. G., Benn, M., Schnohr, P. & Tybjærg-Hansen, A. Nonfasting triglycerides and risk of myocardial infarction, ischemic heart disease, and death in men and women. JAMA298, 299–308. 10.1001/jama.298.3.299 (2007). [DOI] [PubMed] [Google Scholar]
- 5.Karpe, F., Steiner, G., Uffelman, K., Olivecrona, T. & Hamsten, A. Postprandial lipoproteins and progression of coronary atherosclerosis. Atherosclerosis106, 83–97. 10.1016/0021-9150(94)90085-X (1994). [DOI] [PubMed] [Google Scholar]
- 6.Pop, G. A. M. et al. The clinical significance of whole blood viscosity in (cardio)vascular medicine. Netherlands Heart J.10, 512–516 (2002). [PMC free article] [PubMed] [Google Scholar]
- 7.Lowe, G. D. O., Lee, A. J., Rumley, A., Price, J. F. & Fowkes, F. G. R. Blood viscosity and risk of cardiovascular events: The Edinburgh artery study. British J. Haematol.96, 168–173. 10.1046/j.1365-2141.1997.8532481.x (1997). [DOI] [PubMed] [Google Scholar]
- 8.Tamariz, L. J. et al. Blood viscosity and hematocrit as risk factors for type 2 diabetes mellitus: The Atherosclerosis Risk in Communities (ARIC) study. Am. J. Epidemiol.168, 1153–1160. 10.1093/aje/kwn243 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Cho, Y.-I. & Cho, D. J. Hemorheology and microvascular disorders. Korean Circul. J.41, 287–295. 10.4070/kcj.2011.41.6.287 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wu, H.-C., Lee, L.-C. & Wang, W.-J. Plasmapheresis for hypertriglyceridemia: The association between blood viscosity and triglyceride clearance rate. J. Clin. Lab. Anal.33, e22688. 10.1002/jcla.22688 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Shamout, F., Zhu, T. & Clifton, D. A. Machine learning for clinical outcome prediction. IEEE Rev. Biomed. Eng.14, 116–126. 10.1109/RBME.2020.3007816 (2021). [DOI] [PubMed] [Google Scholar]
- 12.Rajkomar, A., Dean, J. & Kohane, I. Machine learning in medicine. New England J. Med.380, 1347–1359. 10.1056/NEJMra1814259 (2019). [DOI] [PubMed] [Google Scholar]
- 13.Rosenblatt, M., Tejavibulya, L., Jiang, R., Noble, S. & Scheinost, D. Data leakage inflates prediction performance in connectome-based machine learning models. Nature Commun.15, 1829. 10.1038/s41467-024-46150-w (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Kapoor, S. & Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns4, 100804. 10.1016/j.patter.2023.100804 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Varoquaux, G. & Cheplygina, V. Machine learning for medical imaging: methodological failures and recommendations for the future. npj Digital Med. 5, 48, 10.1038/s41746-022-00592-y (2022). [DOI] [PMC free article] [PubMed]
- 16.Tucker, A., Wang, Z., Rotalinti, Y. & Myles, P. Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. npj Digital Med.3, 147, 10.1038/s41746-020-00353-9 (2020). [DOI] [PMC free article] [PubMed]
- 17.Pezoulas, V. C. et al. Synthetic data generation methods in healthcare: a review on open-source tools and methods. Comput. Struct. Biotech. J.23, 2892–2910. 10.1016/j.csbj.2024.07.005 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. In Adv. Neural Inf. Proc. Syst. (NIPS)30, 4765–4774 (2017). [Google Scholar]
- 19.de Simone, G. et al. Relation of blood viscosity to demographic and physiologic variables and to cardiovascular risk factors in apparently normal adults. Circulation81, 107–117. 10.1161/01.CIR.81.1.107 (1990). [DOI] [PubMed] [Google Scholar]
- 20.Niculescu-Mizil, A. & Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML ’05), 625–632, 10.1145/1102351.1102430 (2005).
- 21.Dalton, J. E. Flexible recalibration of binary clinical prediction models. Statist. Med.32, 282–289. 10.1002/sim.5544 (2013). [DOI] [PubMed] [Google Scholar]
- 22.Vickers, A. J., Calster, B. V. & Steyerberg, E. W. Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests. BMJ352, i6. 10.1136/bmj.i6 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Dhiman, P. et al. Methodological conduct of prognostic prediction models developed using machine learning in oncology: a systematic review. BMC Med. Res. Method.22, 101. 10.1186/s12874-022-01577-x (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Cai, Y.-Q. et al. Pitfalls in developing machine learning models for predicting cardiovascular diseases: challenge and solutions. J. Med. Internet Res. 26, e47645. 10.2196/47645 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Han, J. W., Sung, P. S., Jang, J. W., Choi, J. Y. & Yoon, S. K. Whole blood viscosity is associated with extrahepatic metastases and survival in patients with hepatocellular carcinoma. PLoS ONE16, e0260311. 10.1371/journal.pone.0260311 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Carlisi, M., Presti, R. L., Mancuso, S., Siragusa, S. & Caimi, G. Calculated whole blood viscosity and albumin/fibrinogen ratio in patients with a new diagnosis of multiple myeloma: Relationships with some prognostic predictors. Biomedicines11, 964. 10.3390/biomedicines11030964 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Atehortúa, A. et al. Cardiometabolic risk estimation using exposome data and machine learning. Int. J. Med. Inf.179, 105209. 10.1016/j.ijmedinf.2023.105209 (2023). [DOI] [PubMed] [Google Scholar]
- 28.Shin, D. Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: A 14-year prospective-based cohort study. BMC Med.Genom.17, 224. 10.1186/s12920-024-01998-1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Wilkinson, M. D. et al. The FAIR guiding principles for scientific data management and stewardship. Scient. Data3, 160018. 10.1038/sdata.2016.18 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Stodden, V., Seiler, J. & Ma, Z. An empirical analysis of journal policy effectiveness for computational reproducibility. Proce. National Academy Sci.115, 2584–2589. 10.1073/pnas.1708290115 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Sasse, L. et al. Overview of leakage scenarios in supervised machine learning. J. Big Data12, 135. 10.1186/s40537-025-01193-8 (2025). [Google Scholar]
- 32.Steyerberg, E. W. et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology21, 128–138. 10.1097/EDE.0b013e3181c30fb2 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All datasets used in this study are fully synthetic and contain no identifiable or linkable human-subject information. The complete machine-learning code, synthetic data generators, and all reproducible workflow scripts are publicly available at:https://github.com/NattakittiP/ML_Predict





















































































































































