Skip to main content
Annals of Medicine logoLink to Annals of Medicine
. 2026 Jun 6;58(1):2682583. doi: 10.1080/07853890.2026.2682583

A body roundness index (BRI)-based predictive model for metabolic syndrome in perimenopausal and postmenopausal women—from a cross-sectional machine learning study to a longitudinal dynamic assessment

Yue Xi a, Qiyue Sun a, Yining Han a, Pengxiang Zhu a, Jiaxin Guo a, Beining Zhang a, Jiacheng Fan a, Zhijun Hong b,#,, Xiaofeng Li a,#,
PMCID: PMC13244513  PMID: 42250232

Abstract

Background and Aims

Metabolic syndrome (MetS) is highly prevalent among perimenopausal and postmenopausal women and poses a major public health challenge because of its association with cardiovascular disease, type 2 diabetes, and premature mortality. However, prediction tools for this population remain limited. Therefore, this study aimed to develop a Body Roundness Index (BRI)-based prediction model for MetS by integrating cross-sectional machine learning and longitudinal assessment.

Methods and Results

Cross-sectional models were trained using NHANES 2007–2020 and validated in the Affiliated Hospital of Dalian University (2023–2024). Sixteen predictors were selected via LASSO and Boruta, and eight models were evaluated using AUC, calibration, and decision curve analysis. SHAP ranked high-contribution factors. Longitudinal analysis used a 10-year cohort. Annualized change rates and cumulative exposure metrics of five key predictors were combined with baseline values to build Cox models, compared by C-index and time-dependent ROC.

The artificial neural network (ANN) demonstrated optimal cross-sectional performance (internal AUC: 0.854; external AUC: 0.878) with good calibration and clinical benefit. SHAP identified BRI, WBC, ALT, MCV, and AST as top contributors, with BRI showing the strongest impact. Longitudinal analysis revealed that integrating annual change rates and annual cumulative exposure of these five predictors achieved optimal discriminative ability (C-index: 0.847), with time-dependent AUCs of 0.853, 0.859, and 0.847 at 1, 3, and 5 years, respectively.

Conclusion

BRI is significantly associated with MetS in perimenopausal and postmenopausal women. The ANN model provides an efficient cross-sectional screening tool, while incorporating longitudinal trajectories of BRI and key laboratory indicators enhances long-term MetS risk prediction.

Keywords: Metabolic syndrome, machine learning, women during the perimenopausal and postmenopausal periods, body roundness index, multicenter

1. Introduction

Metabolic syndrome (MetS) is a group of interrelated metabolic abnormalities, usually manifested as a specific combination of cardiovascular risk factors such as abdominal obesity, elevated blood pressure, abnormal blood sugar metabolism, and blood lipid disorders. In recent years, the prevalence of MetS has continued to rise, further aggravating the overall burden of diseases such as cardiovascular disease, diabetes, chronic kidney disease and polycystic ovary syndrome [1]. Studies have suggested that the incidence of MetS in women during perimenopause and postmenopause increases significantly, up to 3.3 times that of premenopause women [2]. The transition from perimenopausal period to postmenopausal period is accompanied by significant endocrine remodelling, which is mainly manifested by the gradual decline of ovarian function and reduced oestrogen secretion. Against this background, lipid and glucose metabolism are more prone to imbalance, so that women at this stage become high-risk groups of MetS.

Based on the above epidemiological and mechanism clues, the research on risk prediction and stratified assessment for high-risk groups is gradually increasing. Traditional prediction models often take the body mass index (BMI) as the core index [3], but BMI is difficult to effectively reflect the characteristics of fat distribution, especially the insufficient characterisation of the accumulation of abdominal fat/visceral fat, which is one of the key pathological links in the development of MetS. As a new anthropometric indicator proposed in recent years, the Body Roundness Index (BRI) integrates waist circumference and height information to depict the distribution of trunk fat and visceral fat levels, which can be closer to the real state of central obesity to a certain extent. Studies have shown that BRI is superior to some traditional indicators (WC, WHtR, and BMI) in predicting in predicting prediabetes [4], hyperuricemia [5], and cardiovascular diseases [6], and exhibits high predictive value in the risk assessment of MetS [7].

This study takes BRI as the core to build a MetS prediction model for perimenopause and postmenopause women. Although there have been MetS prediction studies based on cross-sectional data in the past, and its useability has been verified in the women and elderly [8,9], most of the relevant work adopts static modelling ideas and does not pay enough attention to the dynamic information of individual indicators over time [10], so there is still a certain situation in long-term risk estimation and individualisation. In order to be closer to the real clinical process, this study introduces longitudinal follow-up data on the basis of cross-sectional machine learning modelling to construct a fusion model: on the one hand, it incorporates BRI-related measures, and on the other hand, it combines the high-contribution variables screened by the SHAP algorithm, and integrates its dynamic changes into the prediction framework. By introducing time-dimensional information, the strategy is expected to improve the model’s ability to accurately assess individual risks, and enhance its interpretability and applicability in early screening and clinical application scenarios.

2. Materials and methods

2.1. Source of data

The cross-sectional data section of this study uses the National Health and Nutrition Examination Survey (NHANES) database. This survey is a project designed by the National Center for Health Statistics (NCHS), which has been conducting annual surveys of 5,000 citizens nationwide since 1999 to assess the health and nutrition status of the U.S. population [11]. We selected data from seven cycles between 2007 and 2020 from the database as an internal validation set. NHANES is an open-source database

(https://www.cdc.gov/nchs/nhanes/index.htm), and the patient data collection process has been approved by the ethics review committee of the National Center for Health Statistics. All participants in the study voluntarily joined the survey and signed informed consent forms. For the external validation phase, samples were collected from the Health Management Center of the First Affiliated Hospital of Dalian Medical University from January 2023 to December 2024. To ensure accurate patient analysis, these individuals were selected based on the same inclusion criteria used in the NHANES dataset and underwent additional validation steps.

The inclusion and exclusion criteria were as follows: Inclusion criteria: (1) Age between 40 and 65 years; (2) Female participants; (3) Meeting the data requirements for MetS diagnosis. Exclusion criteria: (1) Pregnant or recent delivery; (2) Missing BRI data (height and waist circumference); (3) Presence of severe diseases such as malignant tumors. The final internal dataset included 2,706 people and the external dataset included 7,651 people.

The longitudinal data analysis in this study utilized a ten-year cohort from the Health Management Center of the First Affiliated Hospital of Dalian Medical University, covering the period 2013–2022. The inclusion criteria were identical to those previously established, with three additional exclusion criteria: (1) Patients with incomplete follow-up data or unclear follow-up duration; (2) Those with fewer than one follow-up visits; (3) Individuals diagnosed with MetS at baseline (2011). The final cohort included 4,908 people.

2.2. Variable selection

Based on existing literature [12–14] and research objectives, the following four categories of variables were included: a. Demographic information: age, race, education level, marital status; b. Physical examination data: height, weight, waist circumference, BRI. c. Laboratory data: blood cell analysis indicators (Red blood cell count, White blood cell count, platelet count, hemoglobin, red blood cell distribution width, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, percentages and absolute counts of lymphocytes, monocytes, eosinophils, and basophils); electrolytes and metabolites (sodium, potassium, calcium, phosphorus, uric acid); renal function indicators (estimated glomerular filtration rate (eGFR), albumin, globulin, total protein); liver function indicators: total bilirubin, alkaline phosphatase, alanine aminotransferase (ALT), aspartate aminotransferase (AST), gamma-glutamyl transferase (GGT), lactate dehydrogenase (LDH); lipid profile (cholesterol, low-density lipoprotein cholesterol, total cholesterol). d. Gynecological questionnaire: number of pregnancies, age of last delivery, use of contraceptive pills or hormonal medications. e. Dietary intake: daily vitamin D intake, daily sodium intake, daily carbohydrate intake, daily cholesterol intake.

2.3. Definition of MetS

MetS is a standard diagnosis established by the National Cholesterol Education Program Adult Treatment Group III (NCEP-ATPIII) [15]. Individuals who meet three or more of the following five criteria are considered to have MetS: (a) waist circumference exceeding 102 cm in men or 88 cm in women; China waist circumference exceeding 90 cm in men or 85 cm in women; (b) triglycerides ≥ 1.7 mmol/L (150 mg/dL); (c) High density lipoprotein cholesterol(HDL-C) <1.03 mmol/L (40 mg/dL) in men or <1.29 mmol/L (50 mg/dL) in women; (d) systolic blood pressure (SBP) of 130 mmHg or higher, or diastolic blood pressure (DBP) of 85 mmHg or higher, or currently receiving antihypertensive therapy; (e) fasting blood glucose> 5.6 mmol/L (100 mg/dL) or higher, or currently receiving hypoglycemic therapy. Considering ethnic variations in waist circumference criteria, analyses of Asian populations, including the external validation dataset from the First Affiliated Hospital of Dalian Medical University and the longitudinal cohort from the same center, used the cut-off values recommended by the Joint Committee for Developing Chinese Guideline [16], namely waist circumference ≥ 90 cm for men and ≥ 85 cm for women, while the other four diagnostic criteria remained unchanged.

2.4. Statistical analysis

2.4.1. Data cleaning

Given the presence of substantial missing and outliers in the cross-sectional data, data cleaning was performed. Responses labeled as ‘7, 9, 77, 99’ were coded as ‘NA’. Variables with over 30% missing values were removed, while those with less than 30% missing values underwent multiple imputation using R’s ‘mice’ package. Descriptive analysis of the cleaned data revealed that normally distributed or nearly normally distributed quantitative data were presented as mean ± standard deviation, with inter-group comparisons conducted via independent samples t-test. For non-normally distributed quantitative data, medians or quartiles were used, and inter-group comparisons employed non-parametric tests. Categorical data were expressed as frequency (%), with chi-square tests applied for inter-group comparisons.

2.4.2. Feature selection

In this study, to select key features, we first conducted univariate regression analyses on all variables, removing those with p-values greater than 0.05. Then, we used two feature selection methods, Lasso regression (Least Absolute Shrinkage and Selection Operator) and the Boruta algorithm, to improve the accuracy and interpretability of the model. We performed Lasso regression using the ‘glmnet’ package in R [17]. To select the optimal λ value, we used 10-fold cross-validation to evaluate the performance of the model under different λ values. In 10-fold cross-validation, the dataset was divided into ten subsets, with nine subsets used for training and the remaining one for testing. This process was repeated ten times to ensure each subset appeared in the validation set. By calculating the error for each validation, we obtained the average error for each λ value, selecting the λ value that minimized the cross-validation error as the final regularization parameter of the model. We determined the optimal λ value using the lambda.min function in the ‘glmnet’ package and trained the final Lasso model with this λ value. Through Lasso regression, we screened out the features that significantly influenced MetS.

To validate the features selected by Lasso regression and further capture potential nonlinear relationships, we employed the Boruta algorithm [18]. This feature selection method based on random forests evaluates each feature’s importance against randomly generated features to identify meaningful variables. When applying Boruta, we considered the complete feature set and used this approach to further filter out features relevant to clinical outcomes. The variables screened by Boruta included features that highly overlapped with Lasso regression results, demonstrating significant clinical diagnostic value. By combining Lasso regression with Boruta, we determined the final feature set. To visually demonstrate the feature selection results of both methods, we plotted a Venn diagram showing the intersection and differences between variables selected by Lasso regression and Boruta, further validating their complementary nature and consistency.

2.4.3. ML Model construction

The cross-sectional data were randomly divided into two proportions: a training set (70%) and an internal validation set (30%). Supervised machine learning models were constructed using logistic regression (LR), support vector machines (SVM), gradient boosting machines (GBM), artificial neural networks (ANN), extreme gradient boosting trees (XGBoost), k-nearest neighbors (KNN), Adaptive Boosting (AdaBoost), and random forests (RF). Model training and parameter tuning were performed in the training set using repeated 10-fold cross-validation with 5 repeats in the caret framework, with the area under the receiver operating characteristic curve (AUC) used as the primary optimization metric. A limited set of prespecified parameter combinations was evaluated for each algorithm, and the combination with the best cross-validated AUC was selected. Additionally, data from the First Affiliated Hospital of Dalian Medical University (2023–2024) were used for external validation. Model performance evaluation was conducted from three dimensions: discrimination, calibration, and clinical application value. The area under the receiver operating characteristic (ROC) curve (AUC), accuracy, sensitivity, specificity, precision, and F1 score were used to assess model discrimination. Calibration curves evaluated model calibration, while decision curve analysis (DCA) and visualization of confusion matrices provided intuitive insights into the model’s classification capabilities. After identifying the optimal model, a supplementary comparative model was constructed by substituting BRI with BMI, WC, and WHtR within the optimal model, thereby validating the advantages of BRI over traditional obesity indicators.

2.4.4. Variable importance evaluation

SHAP (SHapley Additive Explanatory) is a game-theoretic framework for post-hoc model interpretation in machine learning. It quantifies feature importance by calculating the shapley value of each feature’s contribution to the prediction outcome. This study employs SHAP to enhance model interpretability and transparency.

2.4.5. Longitudinal data analysis and model construction

In this study, we constructed a Cox proportional hazards model using the top 5 variables with the highest contribution degree in the SHAP algorithm. In addition to these five key features, we also incorporated time-varying characteristics, calculated as ACR (annualized change rate) and ACE (annualized cumulative exposure). The calculation formulas are as follows:

ACRi=β^i
xijiitij+εij
ACEi=j=1nj1xij+xij+12×ΔtijTi

Note: xij is the value of the biomarker for individual i at visit j; tij is the corresponding visit time (number of years from the baseline); Δtij is the time interval between two consecutive physical examinations; Ti is the total follow-up duration; β^i is the individual-specific annual variation rate, namely ACR.

We constructed four Cox proportional hazards models (A crude model including only baseline characteristics, Model 1 (baseline characteristics plus ACR), Model 2 (baseline characteristics plus ACE), and Model 3 (baseline characteristics plus both ACR and ACE)) using baseline characteristics combined with two dynamic variables (ACR and ACE) as predictors, with MetS status (1 for MetS, 0 for no MetS) and corresponding follow-up duration (years) as dependent variables. The Schoenfeld residual method was employed to test the proportional hazards assumption in Cox models. Discrimination validation was conducted using the timeROC package to plot time-dependent ROC curves at 1, 3, and 5 years, with AUC values calculated. A model with AUC > 0.7 was considered to have good discrimination. The C-index (C-index) was used to evaluate model performance, where values between 0.7–0.8 indicated moderate precision and above 0.8 indicated high precision. Bootstrap replication (500 iterations) was performed to reduce bias. To visually demonstrate predictive power, we constructed nomograms that convert multiple predictors into survival probabilities, enabling clinicians to rapidly assess MetS risk at 1, 3, and 5 years based on patient characteristics. Finally, a sensitivity analysis was conducted to reassess the predictive accuracy of the optimal model among participants with at least two follow-ups versus those with extended observation periods.

All analyses were performed with a statistical significance level of p < 0.05. The statistical software used in this study was Excel and R version 4.4.0.

3. Result

3.1. Cross section analysis

3.1.1. Baseline characteristics

According to the preset inclusion and exclusion criteria, the internal verification set of this study finally included 2,706 women who met the perimenopause and postmenopausal diagnosis and complete MetS diagnosis information (Figure 1). The baseline characteristics of the study subjects are summarised in Table 1, including 1,278 cases in the MetS group and 1,428 cases in the healthy control group.

Figure 1.

Flowchart depicting the study participant selection process with eligibility criteria and model evaluation steps from NHANES and Dalian Medical University studies. This flowchart outlines participant selection and data processing from various studies, including NHANES 2007-2020 (N=75,402) and Dalian Medical University 2013-2024 (N=268,681). It shows eligibility criteria focusing on female participants aged 40-65, exclusions for conditions like pregnancy and missing data, resulting in eligible samples: NHANES (N=2,706), Dalian 2023-2024 (N=7,651), and Dalian 2013-2022 (N=4,908). It details data cleaning, feature selection, model building, and evaluation using metrics like AUC and Cox models.

Flow chart of this study. LR,logistic regression; SVM,support vector machine; GBM, gradient boosting machine; ANN,artificial neural network; XGBoost, eXtreme Gradient Boosting; k-nearest neighbors; AdaBoost, Adaptive Boosting; RF,random forest; Lasso,Least Absolute Shrinkage and Selection Operator; SHAP, SHapley Additive exPlanation.

Table 1.

demographic and clinical characteristics of MetS and No-MetS group.

Variable Overall No-MetS MetS P
N = 2706 N = 1428 N = 1278
Characteristics (mean(SD)andn(%))
Age(year) 52.40 (7.57) 50.92 (7.53) 54.05 (7.26) <0.001
Race       <0.001
Mexican American 453 (16.7) 208 (14.6) 245 (19.2)
Other Hispanic 348 (12.9) 169 (11.8) 179 (14.0)
Non-Hispanic White 961 (35.5) 535 (37.5) 426 (33.3)
Non-Hispanic Black 620 (22.9) 302 (21.1) 318 (24.9)
Other Race 324 (12.0) 214 (15.0) 110 (8.6)
History of taking contraceptive pills       0.425
Yes 2021(74.7) 1402(74.0) 1252(75.4)
No 685 (25.3) 371 (26.0) 314 (24.6)
Gestation-parturition       <0.001
≤3 1618(59.8) 899(63.0) 719(56.3)
4–7 1017(37.6) 504(35.3) 513(40.1)
≥8 71(2.6) 25(1.8) 46(3.6)
Age at last live birth(year) 33.52 (61.99) 33.35 (57.57) 33.72 (66.61) 0.876
24-hour total nutrient intake(mean(SD))
Vitamin D (D2 + D3) (mcg) 3.90 (4.77) 4.03 (5.01) 3.75 (4.48) 0.132
Sodium (mg) 2988.27 (1442.17) 2984.29 (1434.85) 2992.73 (1450.85) 0.879
Sugars (gm) 97.87 (63.40) 97.58 (61.24) 98.20 (65.74) 0.798
Cholesterol (mg) 251.64 (191.41) 246.22 (184.46) 257.69 (198.78) 0.120
Vital signs (mean(SD))
Height (cm) 160.69 (7.03) 161.11 (7.05) 160.23 (6.98) 0.001
Weight (kg) 78.61 (20.51) 70.78 (16.79) 87.36 (20.76) <0.001
Waist (cm) 99.64 (16.41) 91.86 (13.64) 108.33 (14.80) <0.001
BRI (mean (SD)) 6.12 (2.54) 4.93 (1.97) 7.44 (2.45) <0.001
Laboratory Data(mean(SD))
Sodium (mmol/L) 139.47 (2.36) 139.56 (2.26) 139.38 (2.46) 0.048
Potassium (mmol/L) 3.98 (0.33) 3.97 (0.30) 3.98 (0.36) 0.47
Total Calcium (mg/dL) 9.33 (0.36) 9.32 (0.36) 9.33 (0.36) 0.454
Phosphorus (mg/dL) 3.74 (0.51) 3.75 (0.49) 3.74 (0.53) 0.623
Uric acid (mg/dL) 4.91 (1.28) 4.53 (1.03) 5.34 (1.39) <0.001
Albumin, refrigerated serum (g/dL) 4.11 (0.32) 4.15 (0.30) 4.06 (0.33) <0.001
Globulin (g/dL) 3.03 (0.48) 2.95 (0.45) 3.12 (0.49) <0.001
Total Protein (g/dL) 7.13 (0.46) 7.10 (0.45) 7.18 (0.47) <0.001
Total Bilirubin (mg/dL) 0.60 (0.26) 0.63 (0.26) 0.58 (0.25) <0.001
Alkaline Phosphatase (ALP) (IU/L) 74.15 (26.64) 68.31 (24.04) 80.67 (27.88) <0.001
Alanine Aminotransferase (ALT) (U/L) 22.76 (15.88) 20.52 (13.45) 25.27 (17.89) <0.001
Aspartate Aminotransferase (AST) (U/L) 24.34 (21.48) 23.60 (24.60) 25.17 (17.31) 0.058
Gamma Glutamyl Transferase (GGT) (IU/L) 29.98 (45.79) 24.14 (34.55) 36.50 (55.02) <0.001
Lactate Dehydrogenase (LDH) (IU/L) 136.11 (30.34) 133.78 (28.94) 138.72 (31.64) <0.001
Cholesterol, refrigerated serum (mg/dL) 204.45 (41.14) 203.52 (37.53) 205.49 (44.82) 0.214
LDL-Cholesterol, Friedewald (mg/dL) 120.54 (35.03) 119.13 (32.93) 122.12 (37.18) 0.026
Total Cholesterol (mmol/L) 5.27 (1.05) 5.23 (0.96) 5.30 (1.16) 0.075
Red blood cell count (1000 cells/uL) 4.51 (0.41) 4.45 (0.38) 4.59 (0.43) <0.001
White blood cell count (1000 cells/uL) 6.74 (2.15) 6.21 (1.97) 7.32 (2.19) <0.001
Platelet count (1000 cells/uL) 259.10 (67.72) 254.98 (68.07) 263.72 (67.06) 0.001
Hemoglobin (g/dL) 13.37 (1.34) 13.28 (1.34) 13.48 (1.32) <0.001
Red cell distribution width (%) 13.56 (1.69) 13.46 (1.77) 13.67 (1.60) 0.002
Hematocrit (%) 39.70 (3.63) 39.38 (3.60) 40.07 (3.63) <0.001
Mean cell volume (fL) 88.24 (6.59) 88.79 (6.77) 87.63 (6.33) <0.001
Lymphocyte percent (%) 32.07 (8.43) 32.39 (8.42) 31.72 (8.44) 0.04
Monocyte percent (%) 7.45 (2.09) 7.62 (2.05) 7.26 (2.12) <0.001
Eosinophils percent (%) 2.84 (1.99) 2.85 (2.09) 2.82 (1.89) 0.721
Basophils percent (%) 0.75 (0.44) 0.76 (0.47) 0.74 (0.40) 0.087
Lymphocyte number (1000 cells/uL) 2.09 (0.71) 1.95 (0.65) 2.25 (0.74) <0.001
Monocyte number (1000 cells/uL) 0.49 (0.16) 0.46 (0.16) 0.51 (0.17) <0.001
Eosinophils number (1000 cells/uL) 0.19 (0.15) 0.17 (0.14) 0.20 (0.16) <0.001
Basophils number (1000 cells/uL) 0.04 (0.06) 0.04 (0.05) 0.05 (0.06) <0.001
Mean cell hemoglobin (pg) 29.72 (2.72) 29.94 (2.79) 29.48 (2.62) <0.001
eGFR 92.45 (19.93) 94.10 (18.16) 90.61 (21.60) <0.001

BRI: A Body Roundness Index; eGFR: glomerular filtration rate.

The average age of the healthy control group was 50.92 ± 7.53 years old, and the MetS group was 54.05 ± 7.26 years old. The two groups showed statistical differences in physical measurement indicators (height, weight, waist circumference); the number of pregnancies also differed between groups. In terms of blood biochemical indicators, the relevant parameters varied significantly between the two groups, especially the differences in cholesterol level, bilirubin level and liver function-related indicators.

3.1.2. Feature extraction

After eliminating the variables with a missing rate of more than 30%, a total of 44 candidate variables are reserved for subsequent feature screening. First, LASSO regression is used for variable initial screening (Figure 2A and 2B), and the penalty parameter λ is determined by ten-fold cross-verification, and λ_min is taken as the optimal value; after screening, the regression coefficient of a total of 28 variables is non-zero. Given that incorporating all the above variables into the final model may cause the model structure to be too complex and increase the risk of fitting, this study further adopts the Boruta feature selection method based on random forests to retain potential nonlinear information as much as possible while controlling the number of variables. By comparing the importance of real variables with the randomly generated ‘shadow characteristics’, the algorithm finally identifies 23 variables suitable for model fitting (Figure 2(C)). Combining the results from both Lasso regression and Boruta algorithms, we extracted the intersection of features, ultimately retaining 16 variables (Figure 2(D)). These variables are: ‘Age’ ‘Uric acid’ ‘Globulin’ ‘ALP’ ‘ALT’ ‘AST’ ‘GGT’ ‘Total Cholesterol’ ‘WBC’ ‘Hemoglobin’ ‘RDW’ ‘MCV’ ‘Monocyte percent’ ‘Lymphocyte number’ ‘eGFR’ ‘BRI’.

Figure 2.

Line graphs depict LASSO coefficients and mean squared error versus Log Lambda. A boxplot shows distributions of attributes, while a Venn diagram compares LASSO and Boruta features. The image includes four panels: Panel A displays a line graph of LASSO coefficients against Log Lambda, indicating values decreasing toward zero. Panel B shows mean squared error against Log Lambda with a minimum point. Panel C presents a boxplot illustrating distributions of attributes such as sodium and albumin, with some outliers. Panel D features a Venn diagram comparing feature selections from LASSO (12 unique features) and Boruta (7 unique features), with 16 attributes shared between them.

(A)Lasso coefficient path plots for 44 variables. (B)Cross-validation curves (ten-fold cross validation). The left dashed line represents lambda. min and the right dashed line represents lambda.1se. (C)Boruta-based feature selection results. (D)Venn Diagram of Feature Sets Selected by LASSO Regression and the Boruta Algorithm.

3.1.3. Model selection and evaluation

This study developed eight machine learning models: LR, SVM, GBM, ANN, XGBoost, KNN, AdaBoost, RF. Figure 3(A and B) displays the ROC curves of these models on both training and validation datasets. The Random Forest (RF) and Decision Tree (KNN) models exhibited the most significant performance degradation (AUC from 1.000 to 0.855; from 0.947 to 0.815), indicating a high risk of overfitting, whereas the ANN maintained stable discriminative performance across both training and test sets (AUC from 0.847 to 0.854). On the validation set, LR achieved the highest AUC value (0.858), followed by RF(0.855) and ANN (0.854). However, Table 2 indicates that the RF model exhibits severe overfitting and cannot be directly applied to practical predictions, thus precluding its comparison with other models. To evaluate the model’s performance, the AUC values were compared using the Delong test. As shown in Supplementary Table 1, the ANN and LR models demonstrated statistically insignificant differences in AUC values, yet both exhibited excellent classification performance. Table 3 reveals that ANN outperformed other models in accuracy (0.788), Specificity (0.750), Precisionand (0.745) and F1 score (0.786). XGBoost demonstrated the highest sensitivity and specificity, while GBM achieved the best precision. The confusion matrix in Supplementary Figure (1 and 2) further confirm ANN’s superior ability to distinguish between MetS and NO-MetS among the six models. Figures 3(C and D) indicate that ANN’s calibration curve achieved the best fit in the internal validation set, demonstrating strong alignment between predicted probabilities and observed incidence rates. Decision Curve Analysis (DCA) results for training and test sets (Figures 3(E and F)) reveal that ANN maximize net gain at a specific prediction probability threshold, highlighting their clinical applicability. Comprehensive analysis confirms ANN as the most effective model.

Figure 3.

ROC curves and calibration graphs for multiple models showing sensitivity, specificity, and net benefits at various thresholds. The image consists of six panels illustrating performance metrics for various models. Panel A and B present ROC curves for different models, highlighting sensitivity vs. 1-specificity with the RF model showing the best performance. Panels C and D are calibration curves displaying observed event percentages against bin midpoints for the models. Panels E and F show decision curves of standardized net benefits versus high-risk thresholds, comparing the net benefits of all models.

ROC curves for the eight models in the training set (A) and test set (B). Calibration curve for the eight models in the training set (C) and test set (D). DCA curves for the eight models in the training set (E) and test set (F).

Table 2.

Comparison of the predictive ability of several models in the training set.

Model AUC Accuracy Sensitivity Specificity Precision F1
LR 0.836 0.765 0.833 0.703 0.716 0.77
SVM 0.815 0.742 0.824 0.669 0.692 0.752
GBM 0.915 0.832 0.878 0.790 0.790 0.832
ANN 0.847 0.770 0.821 0.725 0.729 0.772
XGBoost 0.860 0.774 0.853 0.703 0.721 0.782
KNN 0.947 0.856 0.938 0.783 0.796 0.861
Adaboost 0.790 0.739 0.833 0.655 0.685 0.752
RF 1 1 1 1 1 1
Table 3.

Comparison of the predictive ability of several models in the test set.

Model AUC Accuracy Sensitivity Specificity Precision F1
LR 0.858 0.784 0.845 0.731 0.735 0.786
SVM 0.846 0.773 0.882 0.678 0.707 0.785
GBM 0.847 0.776 0.853 0.708 0.72 0.781
ANN 0.854 0.788 0.832 0.750 0.745 0.786
XGBoost 0.848 0.766 0.787 0.748 0.733 0.759
KNN 0.815 0.741 0.782 0.706 0.700 0.739
Adaboost 0.782 0.746 0.855 0.650 0.683 0.759
RF 0.855 0.784 0.845 0.731 0.735 0.786

Additionally, the other three traditional indicators (BRI, WC, and WHtR) were each incorporated into the ANN for 20 replicate analyses. The mean performance evaluation metrics (Supplementary Table 2 and Supplementary Figure 3) showed that BRI exhibited the highest average AUC value (0.8334), slightly higher than WC (0.8332) and WHtR (0.8329), while BMI had the lowest value (0.8209). Overall, BRI, WC, and WHtR demonstrated comparable predictive performance, with BRI showing a slight advantage in discrimination capability, whereas BMI performed relatively weaker overall.

3.1.4. External validation

The external validation dataset comprised 7,651 samples, including 1,524 patients with MetS and 6,127 healthy controls (Supplementary Table 4). Eight models were constructed, with the ANN model demonstrating an AUC of 0.878 on its ROC curve (Figure 4(A)), showing improved discrimination compared to the internal validation set. The calibration curve closely matched the reference line, indicating good calibration (Figure 4(B)). Decision curves revealed significant clinical net benefits when probability thresholds ranged between 0.1–0.8, demonstrating practical value (Figure 4(C)). The confusion matrix in Figure6D clearly demonstrated the ANN model’s superior discrimination. Beyond the AUC, all evaluation metrics showed balanced performance with strong predictive capacity, particularly the high sensitivity (0.844) that highlighted its effectiveness in screening potential MetS risk populations (Supplementary Table 3).

Figure 4.

ROC curve, calibration curve, decision curve analysis, and confusion matrix for ANN model evaluation. Panel A: ROC curve with AUC = 0.878 showing sensitivity vs 1 - specificity. Panel B: Calibration curve of observed vs predicted probability follows the ideal line closely. Panel C: Decision curve analysis graph with net benefit vs high-risk threshold, showing ANN, All, and None lines. Panel D: Confusion matrix displays true/false positives and negatives for predictions (No, Yes).

External validation of the ANN model: (A) ROC curve, (B) Calibration curve, (C) Decision curve, and (D) Confusion matrix.

3.1.5. Explanation of risk factors

Using the SHAP algorithm, we analyzed the importance of predictive variables in the ANN model with the best performance. The SHAP value reflects a variable’s contribution to the model (Figure 5(B)). A higher SHAP value indicates greater contribution. As shown in Figure 5(A), the top-down variable ranking indicates their contribution to patient mortality from hospitalization decreases in order. With SHAP value as the vertical axis, yellow variables on the right represent positive contributions to the outcome, while purple variables on the right indicate negative contributions. The top five variables for predicting MetS in this population are: BRI > WBC > ALT > MCV > AST. Notably, BRI positively contributes to MetS, meaning higher BRI values increase the likelihood of MetS in Perimenopausal and Postmenopausal women. The SHAP dependence plot (Figure 5(D)) further reveals that as BRI increases, the SHAP value exhibits a distinct nonlinear upward trend. After approximately 5, its contribution to predicting risk begins to increase significantly. Additionally, SHAP dependence plots colored according to age, MCV, and WBC were plotted (Figure 5(E–G)), showing no significant interactions. To further investigate the long-term effects of these characteristics, we ultimately included the top five contributing features as predictors in subsequent longitudinal cohort studies.

Figure 5.

Vertical strip plot of SHAP values for various features; bar chart of mean SHAP values; additive force plot of contributions; scatter plots showing SHAP values vs. BRI colored by age, WBC, and MCV. The figure includes multiple panels presenting SHAP values from an ANN model. Panel A is a vertical strip plot showing SHAP values for features such as BR and WBC. Panel B is a horizontal bar chart indicating mean SHAP values. Panel C features an additive force plot detailing contributions to a prediction, f(x)=1, from an expected base E[f(x)]=0.52. Panel D shows a scatter plot of SHAP values against BRI, exhibiting a positive trend. Panels E to G display SHAP dependence plots for BRI, color-coded by age, WBC, and MCV values, illustrating their interrelationships.

(A)Beeswarm plots of the ANN Model.(B)Importance ranking plot of variables for ANN model.(C)Force plots of the ANN Model.(D) SHAP dependence plot for BRI. (E–G) SHAP dependence plots for BRI colored by age, WBC, and MCV.

3.2. Longitudinal cohort analysis

3.2.1. Baseline characteristics

As shown in Table 4, the study analyzed 4,908 participants in the valid cohort. Those with MetS had a median baseline age of 54 (49–60) years and a median follow-up duration of 3 (2–5) years. During the follow-up period, 537 cases of MetS were identified, with a cumulative incidence rate of 10.94% and an incidence density of 3.23 cases per 100 person-years. The baseline age of the MetS group was higher (54.0 vs 50.0 years old, p < 0.001), and the baseline BRI (47.00 vs 38.54), WBC and liver function index were all significantly increased (p < 0.001). Dynamic change analysis shows that the annual change rate of BRI in the MetS group accelerated significantly (0.74 vs 0.10/year, p < 0.001), ALT also showed an upward trend (0.57 vs 0.14 U/L/year, p = 0.001), while MCV showed a downward trend (p < 0.05).

Table 4.

Demographic and clinical characteristics of the MetS and non-MetS groups in the longitudinal cohort.

Variable No-MetS
MetS
P
N = 4,371
N = 537
M P25 P75 M P25 P75
Age 50.00 44.00 55.00 54.00 48.00 59.00 <0.001
BRI_base 38.54 33.89 44.18 47.00 41.77 53.19 <0.001
WBC_base 5.32 4.50 6.24 5.73 4.90 6.84 <0.001
ALT_base 14.00 11.00 19.00 17.00 12.00 24.00 <0.001
MCV_base 89.40 86.70 91.90 89.00 86.50 91.30 0.061
AST_base 18.00 15.00 21.00 18.00 16.00 22.00 <0.001
BRI_ACR 0.24 −0.90 1.33 1.36 0.00 3.22 <0.001
WBC_ACR 0.26 −0.10 6.19 0.08 −0.16 0.45 <0.001
ALT_ACR 0.25 −1.00 1.50 0.90 −1.00 3.00 <0.001
MCV_ACR 0.11 −0.52 0.70 0.03 −0.67 0.68 0.138
AST_ACR 0.17 −0.80 1.02 0.33 −1.00 1.75 0.087
BRI_ACE 39.09 34.77 44.28 49.70 44.15 54.70 <0.001
WBC_ACE 6.22 5.00 15.35 6.04 5.13 7.24 <0.001
ALT_ACE 15.00 12.00 19.67 19.25 15.00 26.00 <0.001
MCV_ACE 89.55 87.00 91.70 88.85 86.78 90.93 <0.001
AST_ACE 18.33 16.00 21.43 19.50 17.00 23.00 <0.001

This study systematically evaluated the relationship between different dimensions of BRI and the risk of MetS using Kaplan-Meier survival analysis. As shown in Supplementary Figure 4, the cumulative incidence of MetS in the high baseline BRI group (>67th) is significantly higher than that in the medium-low baseline group (p < 0.001). At the same time, the BRI annual change rate increased significantly (ACR > 0.5) group has the highest risk, while the declining group has the lowest risk (p < 0.001). Cumulative exposure analysis further shows that there is a significant long-term risk accumulation effect (p < 0.001) in the high BRI cumulative exposure group (>67th). The risk stratification model of comprehensive multi-dimensional information successfully distinguished five risk groups, among which group T1 (ACR > 0.5 and cumulative exposure > 67th) has the highest risk, group T5 (continuous low load) has the lowest risk, and the separation of survival curves between groups is obvious (p < 0.001).

3.2.2. Model selection and evaluation

This study developed four Cox proportional hazards models to evaluate the associations of different metabolic marker exposure patterns with the risk of MetS. The crude model included baseline values only, Model 1 included baseline values plus ACR, Model 2 included baseline values plus ACE, and Model 3 combined baseline values, ACR, and ACE. After collinearity analysis (Supplementary Table 5), variables with VIF > 10 were removed and the models were refitted. As shown in Table 5, all three dynamic models outperformed the crude model. The C-index increased from 0.770 in the crude model to 0.836 in Model 1, 0.839 in Model 2, and 0.847 in Model 3, indicating that longitudinal information substantially improved discrimination. Among them, Model 3 showed the best overall performance, with the highest likelihood ratio test value (χ2 = 875.3, p < 0.001) and Wald test value (χ2 = 1123.0, p < 0.001). In Model 1, both BRI_base (HR = 1.089, 95% CI: 1.081–1.097, p < 0.001) and BRI_ACR (HR = 1.221, 95% CI: 1.198–1.245, p < 0.001) were significant risk factors. In Model 2, BRI_ACE was the main dynamic predictor (HR = 1.113, 95% CI: 1.097–1.130, p < 0.001). In Model 3, BRI_base, BRI_ACR, BRI_ACE, and WBC_ACE remained statistically significant, suggesting that baseline levels, dynamic changes, and cumulative exposure each contributed independently to MetS risk prediction. Schoenfeld residual tests indicated time-varying effects for some dynamic variables (Supplementary Figure 5). Time-dependent ROC analysis (Figure 6(D–F)) further showed that, at 1, 3, and 5 years, model discrimination generally followed the order of Model 3 > Model 2 > Model 1. Taken together, Model 3 demonstrated the best overall predictive performance and was therefore selected as the final model. A nomogram based on Model 3 was subsequently constructed to estimate the 1-, 3-, and 5-year risk of MetS (Figure 7).

Table 5.

Comprehensive comparison of cox models for MetS risk prediction.

Variable Crude Model
Model 1
Model 2
Model 3
HR P HR P HR P HR P
(95% CI) (95% CI) (95% CI) (95% CI)
Baseline Values                
BRI_base 1.077 (1.070–1.084) <0.001 1.089 (1.081–1.097) <0.001 0.977 (0.962–0.993) 0.004 1.055 (1.028–1.083) <0.001
WBC_base 1.146 (1.083–1.214) <0.001 1.148 (1.084–1.217) <0.001 1.309 (1.233–1.389) <0.001 1.298 (1.212–1.390) <0.001
AST_base 1.001 (0.989–1.012) 0.931 0.999 (0.986–1.012) 0.85 1.001 (0.990–1.012) 0.813 1.002 (0.990–1.014) 0.738
MCV_base 1.020 (1.003–1.037) 0.019 1.001 (0.982–1.020) 0.915 1.025 (0.999–1.052) 0.059 1.017 (0.985–1.050) 0.311
ALT_base 1.007 (1.002–1.011) 0.006 1.008 (1.002–1.014) 0.009 1.006 (1.001–1.010) 0.013 1.005 (1.000–1.010) 0.041
Annual Change Rate (ACR)                
BRI_ACR 1.221 (1.198–1.245) <0.001 1.190 (1.157–1.224) <0.001
WBC_ACR 0.876 (0.848–0.905) <0.001 1.017 (0.975–1.061) 0.439
AST_ACR 0.984 (0.957–1.012) 0.248
MCV_ACR 0.948 (0.910–0.988) 0.011 0.971 (0.920–1.024) 0.271
ALT_ACR 1.012 (0.997–1.027) 0.113
Annual Cumulative Exposure (ACE)                
BRI_ACE 1.113 (1.097–1.130) <0.001 1.032 (1.005–1.059) 0.018
WBC_ACE 0.887 (0.867–0.908) <0.001 0.877 (0.841–0.915) <0.001
MCV_ACE 0.977 (0.952–1.002) 0.074 0.982 (0.952–1.013) 0.243
Model Performance                
C-index (SE) 0.770 (0.011) 0.836 (0.011) 0.839 (0.009) 0.847 (0.010)
Likelihood Ratio Test χ²=429.3, p < 0.001 χ²=833.5, p < 0.001 χ²=752.3, p < 0.001 χ²=875.3, p < 0.001
Wald Test χ²=607.7, p < 0.001 χ²=1180.0, p < 0.001 χ²=905.7, p < 0.001 χ²=1123, p < 0.001

Crude Model: BRI_base, WBC_base, MCV_base, AST_base, ALT_base.

Model 1: BRI_base, WBC_base, MCV_base, AST_base,ALT_base, BRI_ACR, WBC_ACR, AST_ACR MCV_ACR, ALT_ACR.

Model 2: BRI_base, WBC_base, MCV_base, AST_base,ALT_base, BRI_ACE, WBC_ACE, MCV_ACE.

Model 3: BRI_base, WBC_base, MCV_base, AST_base,ALT_base, BRI_ACR, WBC_ACR, MCV_ACR, BRI_ACE, WBC_ACE, MCV_ACE.

Figure 6.

Multi-panel figure with forest plots (A-D) depicting various baseline variables and ROC curves (E-H) showing sensitivity and AUC values over time. The figure includes eight panels (A-H). Panels A-D are forest plots showing hazard ratios (HR) and 95% confidence intervals for clinical variables across different measurements. Panel A features baseline variables, while panels B-D incorporate ACR and ACE variables. Panels E-H present ROC curves plotting sensitivity against 1-specificity for 1-year, 3-year, and 5-year intervals, including AUC values for each duration. A dashed line marks random classification. AUC values vary from 0.775 (1-year, E) to 0.859 (3-year, H).

Forest plot of the Crude model (A), model 1 (B), model 2(C) and model 3(D). Time-dependent ROC curves of the Crude model (E), model 1 (F), model 2 (G) and model 3 (H) at 1, 3, and 5 years.

Figure 7.

A multi-panel figure showing health metrics on vertical axes and prediction outcomes for 1, 3, and 5 years. This figure displays three panels: Panel A presents a bar chart with various health metrics (BRI_base, WBC_base, ALT_base, BRI_ACR, BRI_ACE, WBC_ACE). Panels B and C show total points (0-100) and a linear predictor, respectively, indicating outcome probabilities for 1-year, 3-year, and 5-year predictions with slight value decline over time.

Nomogram for Predicting 1, 3, and 5Year Risk of MetS based on the Model 3.

3.2.3. Sensitivity analyses

In the sensitivity analysis, two groups were selected; participants with at least two follow-up visits (minimum two follow-up subgroup) and those with extended observation periods(extended follow-up subgroup). Supplementary Table 6 shows that The model demonstrated acceptable predictive performance in both the minimum two follow-up subgroup (C-index = 0.813; 3-year AUC = 0.824; 5-year AUC = 0.795) and the extended follow-up subgroup (C-index = 0.879; 3-year AUC = 0.833; 5-year AUC = 0.858). Baseline BRI and BRI_ACR remained significant predictors in both strata, the robustness of the results over the follow-up period was demonstrated.

4. Discussion

In order to identify the high-risk population of MetS, there have been more and more studies on the risk prediction model of the disease in recent years [19–21]. However, judging from the existing evidence, there are still some common shortcomings in this kind of model. For example, the research of Ibrahim et al. [21] is mainly based on a single cohort construction model, which lacks external verification support, and its generalisation ability and extraportion validity still need to be further tested. Lim et al. [22] proposed that convolutional ANN can be used for electrocardiogram signal classification, which can distinguish between MetS patients and healthy people to a certain extent. However, many machine learning models are still ‘traditional’ in variable systems and often rely heavily on conventional anthropometric indicators such as BMI. Although such indicators are easy to obtain, it is difficult to accurately depict the distribution of visceral fat, which is one of the key pathological characteristics of the development of MetS [23]. In comparison, Li et al. [3] suggest that BRI is superior to a variety of traditional anthropometric indicators in predicting cardiovascular risk factors, and can more finely reflect the differences in fat distribution. In addition, the previous MetS prediction models mainly target middle-aged and elderly people [24,25], and the research on specific female populations such as perimenopause is relatively limited. Studies have compared four machine learning models in perimenopause women [10], but the samples only cover the perimenopause population aged 45–55 and are not included in the postmenopause stage, so it is difficult to observe further changes in the risk of postmenopause disease. On this basis, this study expands the age range and covers Perimenopausal and Postmenopausal women, aiming to improve the applicability and popularisation value of the model in a wider range of people. Therefore, this study adopts the integrated idea of ‘cross-sectional machine learning + longitudinal dynamic evaluation’ in Perimenopausal and Postmenopausal women to construct and form a MetS prediction model with BRI as the core. Cross-sectional analysis shows that the ANN model has high prediction efficiency (AUC = 0.854) and maintains stable performance in external verification (AUC = 0.878). Additionally, when models incorporating traditional indicators (BMI, WC, and WHtR) were constructed based on the ANN architecture, BRI, WC, and WHtR demonstrated comparable discriminative performance, with BRI showing only marginally higher AUC values than the other two indicators, while all three indicators outperformed BMI. Based on the interpretability analysis of the SHAP algorithm, this study identifies high-contribution characteristics such as BRI, WBC, ALT, MCV and AST, which were strongly associated with MetS risk in the prediction model; its direction is generally consistent with the previous research results [26–28], thus enhancing the reliability of the model interpretation to a certain extent. It should be noted that SHAP values reflect feature importance for model prediction and do not directly imply biological causation. Furthermore, the three Cox proportional risk models based on the longitudinal cohort are used to evaluate the predictive value of different metabolic biomarkers exposure patterns. Among these, the model incorporating the ACR and ACE dynamic variables (Model 3) demonstrated superior overall discriminative power (C-index = 0.847) and exhibited good clinical interpretability; additionally, it maintained relatively stable predictive performance at the 1-year, 3-year, and 5-year follow-up time points.

This study fully verifies the advantages of BRI as the core indicator from the cross-sectional and longitudinal levels. The results show that BRI has the highest contribution (the largest SHAP value) in the cross-sectional model, and in the longitudinal model, it shows significant predictive value in both the baseline level and the rate of change. This finding is consistent with the conclusion of Sergio et al. that BRI is better than traditional anthropometric indicators in metabolic risk assessment [29]. Further longitudinal cohort analysis also strengthened the above understanding: in the Kaplan–Meier single-factor survival analysis, the multidimensional characteristics of BRI are clearly related to the risk of MetS - higher baseline BRI, faster upward trend and higher cumulative exposure all correspond to a significantly increased MetS risk. The composite stratification model (T1–T5) based on the three-dimensional information of ‘baseline state + change rate + cumulative load’ further suggests that the comprehensive assessment framework can identify the continuous risk spectrum from the highest risk (T1) to the lowest risk (T5) [30], and more intuitively reflect the risk related to visceral fat from The pathological trajectory of the evolution of ‘static accumulation’ to ‘dynamic progression’. It should be noted that KM analysis, as a single-factor description method, cannot control the mixed effects of other metabolic indicators at the same time, and it is difficult to quantify the independent contributions of each dimension. There have been studies such as Sun and other MetS prediction models based on the multi-centre health management population in Shandong Province, using multi-factor Cox regression to complete variable screening and model construction [31]. Based on this idea, this study further establishes a multi-factor Cox proportional risk model; after adjusting key covariables such as WBC and liver function, BRI_base (HR = 1.055) and BRI_ACR (HR = 1.190) still maintain independent and stable prediction effects. This result has two meanings: first, the predictive value of BRI dynamic indicators (ACR) does not come from colinearity with other metabolic indicators, but contains relatively independent risk information; second, the HR of BRI_ACR is close to BRI_base, suggesting ‘growth rate’ and ‘existing level’ The weight in risk assessment may be comparable, which to a certain extent supplements and corrects the traditional paradigm of over-reliance on single measurement values in clinical practice. What is more noteworthy is that the BRI dynamic effect continues to be significant in the model, which is consistent with the ‘metabolic memory’ hypothesis [32]. This study suggests that metabolic memory may be reflected not only through long-term cumulative load (ACE) but also in the form of continuously deteriorating change trajectory (ACR); even when baseline levels are similar, different change trends are associated with different metabolic outcomes. This finding may help account for why the T1 group in the complex stratification (rapid rise + high exposure) had a significantly higher risk than other groups. In terms of clinical translation, this finding suggests a potential shift in risk assessment from ‘current state’ to ‘change trajectory’ and offers a quantifiable tool for identifying high-risk individuals (e.g. T1 and T2 groups) during the critical stage of ‘metabolic decompensation’. At the same time, the model has better time-dependent prediction performance in 1–5 years of follow-up (AUC > 0.8), which further supports its clinical useability, and the overall prediction accuracy is better than that of Liu and other research models [33]. In addition, most models demonstrate superior performance in short-term (1-year) prediction, which may be related to the more sensitive capture of recent metabolic deterioration signals at the rate of change, so it is more suitable for clinical situations that require rapid decision-making. In summary, this study not only verifies the application value of BRI as a core biomarker, but also demonstrates the necessity and feasibility of incorporating dynamic change indicators into routine metabolic risk assessment through the evidence chain of ‘descriptive stratification (KM) - effect quantification (Cox) - predictive verification (AUC)’, so as to The accurate prevention of metabolic diseases provides methodological ideas and empirical support.

In addition to BRI, WBC, liver function indicators, red blood cell parameters and kidney function markers also have significant predictive significance, which constitute a multi-dimensional assessment framework for the risk of MetS in Perimenopausal and Postmenopausal women. The importance of WBC is second only to BRI: the MetS group WBC increased significantly (5.73 vs 5.32 × 109/L), which is consistent with the theory of chronic low-grade inflammation [34,35], and supports Pei and other studies [36]. The differences in white blood cell subtypes further suggest more detailed pathological information, in which the lymphocyte count increased significantly in the MetS group (2.25 vs 1.91 × 109/L), which may be associated with adaptive immune-mediated metabolic inflammation [37]. It is worth noting that WBC annual change trend index WBC_ACR has an independent protective effect, suggesting that metabolic risk is not only related to the ‘level’ of inflammation, but also to the reversibility of inflammation [38]; as a dynamic marker, WBC_ACR may be used to evaluate the improvement of inflammation after intervention and its risk return., provide a basis for the accurate prevention of ‘trend-oriented’. In terms of liver function, ALT and AST have high contribution characteristics, suggesting the key role of the liver in metabolic regulation. The elevation of ALT_base in the MetS group (17.00 vs 14.00 U/L) is consistent with previous studies [39]; however, ALT/AST dynamic indicators (ALT_ACR, AST_ACR) are not significant. The reasons may include individual variation caused by perimenopausal hormone fluctuations [40], liver enzyme abnormalities are more partially accompanied/lagging in this population, and the cross-sectional association is difficult to extrapor to the longitudinal stability effect.

Advantages and Limitations: This research reflects the outstanding advantages at the methodological level, and also presents a systematic promotion from traditional statistical analysis to modern predictive modelling. First, the multi-stage data integration strategy makes up for the shortcomings of a single data source to a certain extent. This study connects and complements the NHANES cross-sectional data of the United States and the single-centre longitudinal cohort in China: the former provides a large sample and strong population representation, and the latter supplements the dynamic information brought by long-term follow-up, so that the study achieves a relative balance between ‘breadth’ and ‘depth’. This kind of cross-country and cross-data type design is not common in the field of MetS prediction, which can provide a reference path for subsequent international comparison and cross-group verification. Second, the construction of the dynamic risk assessment framework has certain theoretical innovation significance. The study not only verified the predictive value of BRI as a static indicator, but also expanded risk assessment from a single ‘state judgement’ to ‘process monitoring’ by introducing two types of dynamic measures of change (ACR) and cumulative exposure (ACE). The framework is closer to the essence of MetS chronic progressive diseases, that is, the risk is not only determined by the current level, but also affected by the trajectory of change. Furthermore, the interactive analysis of dynamic indicators and baseline indicators suggests that there may be a metabolic adaptation threshold, which provides a new perspective for explaining individual differences. Nevertheless, there are also limitations in this study that require objective understanding. First, the heterogeneity of data sources may have an impact on the generalisation ability of the model. Cross-sectional data comes from multi-ethnic populations (NHANES) in the United States, and longitudinal data comes from single-centre cohorts in China. Even if quality control and standardised processing have been carried out, differences in genetic background, lifestyle and medical systems of different groups may still bring potential bias. It should be especially noted that there may be systematic differences in the distribution of BRI in the United States and China, which will affect the applicable performance of the American data training model in the Chinese population. Second, the duration of the follow-up limits the robustness of long-term forecasts. The median follow-up time of this study is 3 years (2–5 years), and the lifetime risk assessment of chronic diseases such as MetS is still insufficient; the metabolic changes in perimenopausal women may last for 10–20 years, and it may be difficult for a shorter observation window to fully capture the whole trajectory of risk evolution.

5. Conclusion

This study is aimed at Perimenopausal and Postmenopausal women, builds a MetS prediction model based on BRI, and integrates cross-sectional machine learning and longitudinal dynamic evaluation. Cross-sectional analysis shows that the ANN model has a strong distinguishing ability (AUC = 0.854); combined with SHAP interpretation results, BRI, WBC, ALT, MCV and AST are identified as core prediction factors that contribute highly to model output. The longitudinal cohort analysis further shows that the Cox model 3 (Base + ACR+ACE), which combines the baseline level and the annual change rate, performs best in terms of prediction accuracy (C-index = 0.847), statistical stability and clinical useability, and maintains a relatively stable discrimination in the 1–5-year prediction window. Force. Overall, this study has formed a comprehensive assessment system extending from static feature identification to longitudinal risk monitoring, which provides a basis for the integration of BRI and its dynamic changes into routine clinical evaluation, and can provide relatively reliable tool support for early warning and individualised intervention of MetS in Perimenopausal and Postmenopausal women.

Supplementary Material

Supplementary material revision.docx

Funding Statement

No funding was received.

Disclosure statement

No potential conflict of interest was reported by the author(s).

Ethical approval and consent to participate

Ethical approval for this study was conducted by the Ethics Committee of the Research Department atthe First Affiliated Hospital of Dalian Medical University (approval number PJ-KS-KY-2026-39, Jan. 2026). The study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki. All participants were informed about the purpose of the study, assured of confidentiality, and provided written consent prior to participation. Participation was voluntary, and respondents could withdraw at any time without consequence.

Data availability statement

The original contributions presented in the study are included in the article and the supplementary material. Further inquiries can be directed to the corresponding author. Detailed information and data from NHANES can be downloaded from https://www.cdc.gov/nchs/nhanes/

References

  • 1.Grundy SM. Metabolic syndrome update. Trends Cardiovasc Med. 2016;26(4):364–373. doi: 10.1016/j.tcm.2015.10.004. [DOI] [PubMed] [Google Scholar]
  • 2.Pu D, Tan R, Yu Q, et al. Metabolic syndrome in menopause and associated factors: a meta-analysis. Climacteric. 2017;20(6):583–591. doi: 10.1080/13697137.2017.1386649. [DOI] [PubMed] [Google Scholar]
  • 3.Li Y, He Y, Yang L, et al. Body Roundness Index and Waist–Hip Ratio Result in Better Cardiovascular Disease Risk Stratification: results From a Large Chinese Cross-Sectional Study. Front Nutr. 2022;9:801582. doi: 10.3389/fnut.2022.801582. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Pramono A, Nursari EN, Dieny FF, et al. Assessing the predictive accuracy of the body roundness index for prediabetes in indonesian adults: analisis prediksi body roundness index untuk Prediabetes pada Orang Dewasa di Indonesia. AMNT. 2025;9(4):689–697. doi: 10.20473/amnt.v9i4.2025.689-697. [DOI] [Google Scholar]
  • 5.Zhang N, Chang Y, Guo X, et al. A body shape index and body roundness index: two new body indices for detecting association between obesity and hyperuricemia in rural area of China. Eur J Intern Med. 2016;29:32–36. doi: 10.1016/j.ejim.2016.01.019. [DOI] [PubMed] [Google Scholar]
  • 6.Xu J, Zhang L, Wu Q, et al. Body roundness index is a superior indicator to associate with the cardio-metabolic risk: evidence from a cross-sectional study with 17,000 Eastern-China adults. BMC Cardiovasc Disord. 2021;21(1):97. doi: 10.1186/s12872-021-01905-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Chen Z, Cheang I, Zhu X, et al. Associations of body roundness index with cardiovascular disease and mortality among patients with metabolic syndrome. Diabetes Obes Metab. 2025;27(6):3285–3298. doi: 10.1111/dom.16346. [DOI] [PubMed] [Google Scholar]
  • 8.Cybulska AM, Rachubińska K, Grochans E, et al. Systemic inflammation indices, chemokines, and metabolic markers in perimenopausal women. Nutrients. 2025;17(17):2885. doi: 10.3390/nu17172885. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Pu F, He R, Wei Y, et al. Sex differences in the optimal cut-off values of visceral fat area for predicting metabolic syndrome among Chinese middle-aged and elderly populations: a cross-sectional study. Br J Nutr. 2026;135(3):250–260. doi: 10.1017/S0007114525105801. [DOI] [PubMed] [Google Scholar]
  • 10.Eyvazlou M, Hosseinpouri M, Mokarami H, et al. Prediction of metabolic syndrome based on sleep and work-related risk factors using an artificial neural network. BMC Endocr Disord. 2020;20(1):169. doi: 10.1186/s12902-020-00645-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ahluwalia N, Dwyer J, Terry A, et al. Update on NHANES dietary data: focus on collection, release, analytical considerations, and uses to inform public policy. Adv Nutr. 2016;7(1):121–134. doi: 10.3945/an.115.009258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Xiaoxue W, Zijun W, Shichen C, et al. Risk prediction model of metabolic syndrome in perimenopausal women based on machine learning. Int J Med Inform. 2024;188:105480. doi: 10.1016/j.ijmedinf.2024.105480. [DOI] [PubMed] [Google Scholar]
  • 13.Shin H, Shim S, Oh S.. Machine learning-based predictive model for prevention of metabolic syndrome. PLoS One. 2023;18(6):e0286635. doi: 10.1371/journal.pone.0286635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Zhang Y, Razbek J, Li D, et al. Construction of Xinjiang metabolic syndrome risk prediction model based on interpretable models. BMC Public Health. 2022;22(1):251. doi: 10.1186/s12889-022-12617-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Grundy SM, Cleeman JI, Daniels SR, et al. Diagnosis and management of the metabolic syndrome: an American Heart Association/National Heart, Lung, and Blood Institute Scientific Statement. Circulation. 2005;112(17):2735–2752. doi: 10.1161/CIRCULATIONAHA.105.169404. [DOI] [PubMed] [Google Scholar]
  • 16.Lu J, Wang L, Li M, et al. Metabolic syndrome among adults in China: the 2010 China noncommunicable disease surveillance. J Clin Endocrinol Metab. 2017;102(2):507–515. [DOI] [PubMed] [Google Scholar]
  • 17.Friedman JH, Hastie T, Tibshirani R.. Regularization paths for generalized linear models via coordinate descent. J Stat Softw. 2010;33(1):1–22. doi: 10.18637/jss.v033.i01. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zhou H, Xin Y, Li S.. A diabetes prediction model based on Boruta feature selection and ensemble learning. BMC Bioinformatics. 2023;24(1):224. doi: 10.1186/s12859-023-05300-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hua YX, Shan TQ, Sun F[, et al. Setting up a risk prediction model on metabolic syndrome among 35-74 year-olds based on the Taiwan MJ Health-checkup Database]. Zhonghua Liu Xing Bing Xue Za Zhi. 2013;34(9):874–878. [PubMed] [Google Scholar]
  • 20.Chen S, Xu Y, Jiang Y, et al. Development and validation of a predictive model for metabolic syndrome in a large cohort of people living with HIV. Virol J. 2024;21(1):321. doi: 10.1186/s12985-024-02592-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ibrahim MS, Pang D, Randhawa G, et al. Development and validation of a simple risk model for predicting metabolic syndrome (MetS) in midlife: a cohort study. Diabetes Metab Syndr Obes. 2022;15:1051–1075. doi: 10.2147/DMSO.S336384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Lim C, Kim JY, Nam Y. ECG signal analysis for patient with metabolic syndrome based on 1d-convolution neural network. In 2020 International Conference on Computational Science and Computational Intelligence (CSCI). 2020. p. 731–3. https://ieeexplore.ieee.org/document/9457885. doi: 10.1109/CSCI51800.2020.00134. [DOI] [Google Scholar]
  • 23.Guo T, Zheng S, Chen T, et al. The association of long-term trajectories of BMI, its variability, and metabolic syndrome: a 30-year prospective cohort study. EClinicalMedicine. 2024;69:102486. doi: 10.1016/j.eclinm.2024.102486. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Pang Y, Wang Y, Hao H, et al. Associations of multiple serum metals with the risk of metabolic syndrome among the older population in China based on a community study: a mediation role of peripheral blood cells. Ecotoxicol Environ Saf. 2024;284:116981. doi: 10.1016/j.ecoenv.2024.116981. [DOI] [PubMed] [Google Scholar]
  • 25.Jhuang YH, Kao TW, Peng TC, et al. Serum phosphorus as a risk factor of metabolic syndrome in the elderly in Taiwan: a large-population cohort study. Nutrients. 2019;11(10):2340. doi: 10.3390/nu11102340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Tanaka M, Okada H, Hashimoto Y, et al. Combined effect of hemoglobin and mean corpuscular volume levels on incident metabolic syndrome: a population-based cohort study. Clin Nutr ESPEN. 2020;40:314–319. doi: 10.1016/j.clnesp.2020.08.010. [DOI] [PubMed] [Google Scholar]
  • 27.Wang M, Ma G, Tao Z.. The association of neutrophil-to-lymphocyte ratio with cardiovascular and all-cause mortality among the metabolic syndrome population. BMC Cardiovasc Disord. 2024;24(1):594. doi: 10.1186/s12872-024-04284-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Oda E, Kawai R.. Comparison between high-sensitivity C-reactive protein (hs-CRP) and white blood cell count (WBC) as an inflammatory component of metabolic syndrome in Japanese. Intern Med. 2010;49(2):117–124. doi: 10.2169/internalmedicine.49.2670. [DOI] [PubMed] [Google Scholar]
  • 29.Rico-Martín S, Calderón-García JF, Sánchez-Rey P, et al. Effectiveness of body roundness index in predicting metabolic syndrome: a systematic review and meta-analysis. Obes Rev. 2020;21(7):e13023. doi: 10.1111/obr.13023. [DOI] [PubMed] [Google Scholar]
  • 30.Ross R, Neeland IJ, Yamashita S, et al. Waist circumference as a vital sign in clinical practice: a Consensus Statement from the IAS and ICCR Working Group on Visceral Obesity. Nat Rev Endocrinol. 2020;16(3):177–189. doi: 10.1038/s41574-019-0310-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Yu G. Research on the test method of proportional hazard hypothesis in Cox model. Interl J Soc Sci Educ Res. 2023;6(4):23–34. [Google Scholar]
  • 32.Dong H, Sun Y, Nie L, et al. Metabolic memory: mechanisms and diseases. Signal Transduct Target Ther. 2024;9(1):38. doi: 10.1038/s41392-024-01755-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Liu W, Tang X, Cui T, et al. Development and visualization of a risk prediction model for metabolic syndrome: a longitudinal cohort study based on health check-up data in China. Front Nutr. 2023;10:1286654. doi: 10.3389/fnut.2023.1286654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Gluvic Z, Zaric B, Resanovic I, et al. Link between metabolic syndrome and insulin resistance. Curr Vasc Pharmacol. 2017;15(1):30–39. doi: 10.2174/1570161114666161007164510. [DOI] [PubMed] [Google Scholar]
  • 35.Esser N, Legrand-Poels S, Piette J, et al. Inflammation as a link between obesity, metabolic syndrome and type 2 diabetes. Diabetes Res Clin Pract. 2014;105(2):141–150. doi: 10.1016/j.diabres.2014.04.006. [DOI] [PubMed] [Google Scholar]
  • 36.Pei C, Chang J-B, Hsieh C-H, et al. Using white blood cell counts to predict metabolic syndrome in the elderly: A combined cross-sectional and longitudinal study. Eur J Intern Med. 2015;26(5):324–329. doi: 10.1016/j.ejim.2015.04.009. [DOI] [PubMed] [Google Scholar]
  • 37.Li N, Liu C, Luo Q, et al. Correlation of white blood cell, neutrophils, and hemoglobin with metabolic syndrome and its components. Diabetes Metab Syndr Obes. 2023;16:1347–1355. doi: 10.2147/DMSO.S408081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Hotamisligil GS. Inflammation, metaflammation and immunometabolic disorders. Nature. 2017;542(7640):177–185. doi: 10.1038/nature21363. [DOI] [PubMed] [Google Scholar]
  • 39.Md D, B L, Sl F, et al. Severity of metabolic syndrome is greater among nonalcoholic adults with elevated ALT and advanced fibrosis. Nutrition research (New York, NY). 2021. [DOI] [PMC free article] [PubMed]
  • 40.Ruhl CE, Everhart JE.. Determinants of the association of overweight with elevated serum alanine aminotransferase activity in the United States. Gastroenterology. 2003;124(1):71–79. doi: 10.1053/gast.2003.50004. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material revision.docx

Data Availability Statement

The original contributions presented in the study are included in the article and the supplementary material. Further inquiries can be directed to the corresponding author. Detailed information and data from NHANES can be downloaded from https://www.cdc.gov/nchs/nhanes/


Articles from Annals of Medicine are provided here courtesy of Taylor & Francis

RESOURCES