| Abdel-Aal and Mangoud [26], 1997 |
Models for WHRb as a continuous variable predict the actual values within an error rate of 7.5% at the 90% confidence limits.
Categorical models predict the correct logical value of WHR with an error in only 2 of the 300 evaluation cases.
Analytical relationships derived from simple categorical models explain global observations on the total survey population to an accuracy rate as high as 99%.
Simple continuous models represented as analytical functions highlight global relationships and trends.
There is a strong correlation between WHR and diastolic blood pressure, cholesterol level, and family history of obesity.
|
|
| Positano et al [71], 2008 |
CVd values in VATe, SATf, and VAT/SAT ratio assessment by the standard algorithm without image inhomogeneities correction were 10.7%, 11.9%, and 17.3%, respectively. Correlation coefficients were r=0.97, r=0.93, and r=0.95, respectively (all P<.001).
When correction for field inhomogeneities was applied, VAT, SAT, and VAT/SAT ratio CVs became 9.8%, 6.7%, and 13.1%, respectively. Correlation coefficients became r=0.97, P<.001 for VAT; r=0.99, P<.001 for SAT; and r=0.97, P<.001 for VAT/SAT ratio.
|
The CV between manual and unsupervised analyses was significantly improved by inhomogeneities correction in SAT evaluation. Systematic underestimation of SAT was also corrected. A less critical performance improvement was found in VAT measurement.
The compensation of signal inhomogeneities improves the effectiveness of the unsupervised assessment of abdominal fat.
Correction of intensity distortions is necessary for SAT evaluation but less significant in VAT measurement.
|
| Ergün [70], 2009 |
The classification rate of neural networks in obesity is 90.2%, and the classification rate of logistic regression in obesity is 87.8%.
After these classifications, in obesity, the BMI is more affected than the divergent arteries.
|
|
| Yang et al [69], 2009 |
|
|
| Zhang et al [68], 2009 |
Prediction at 8 months’ accuracy is improved very slightly, in this case by using neural networks, whereas for prediction at 2 years, the obtained accuracy is enhanced by >10%, in this case by using Bayesian methods.
|
SVMg and Bayesian algorithms seem to be the best algorithms for predicting overweight and obesity from the Wirral database.
The incorporation of nonlinear interactions could be important in childhood obesity prediction. Data mining techniques are becoming sufficiently well established to offer the medical research community a valid alternative to logistic regression.
|
| Heydari et al [67], 2012 |
Regarding logistic regression and neural networks, the respective values were 80.2% and 81.2% for correct classification 80.2% and 79.7% for sensitivity, and 81.9% and 83.7% for specificity; the values for the area under the receiver operating characteristic curve were 0.888 and 0.884, respectively, and the values for the kappa statistic were 0.600 and 0.629, respectively.
Abdominal thickness, weight, BMI, and HCh were significantly associated with obesity.
|
|
| Kupusinac et al [66], 2014 |
|
|
| Shao [65], 2014 |
|
Compared with traditional single-stage approaches, the proposed hybrid models—multiple regression, ANN, multivariate adaptive regression splines, and support vector regression techniques—can effectively predict BFP.
|
| Chen et al [64], 2015 |
The most important correlated indexes are creatinine, hemoglobin, hematocrit, uric acid, red blood cells, high-density lipoprotein, alanine transaminase, triglyceride, and γ-glutamyl transpeptidase.
|
The ELMk performs much more efficiently than the SVM and BPNNl and with higher recognition rates.
The proposed ELM-based approach for overweight detection in biomedical applications holds promise as a new, accurate method for identifying participants’ overweight status. It provides a viable alternative to traditional overweight modeling tools by offering excellent predictive ability.
|
| Dugan et al [63], 2015 |
The ID3m model trained on the CHICAn data set demonstrated the best overall performance with an accuracy of 85% and sensitivity of 89%. In addition, the ID3 model had a positive predictive value of 84% and a negative predictive value of 88%.
Being overweight between the ages of 12 and 24 months is a key risk factor for obesity after the second birthday. Furthermore, it is more of a risk factor if the child was not overweight before 12 months.
|
|
| Nau et al [62], 2015 |
After examining 44 community characteristics, the researchers identified 13 features of the social, food, and physical activity environment that, in combination, correctly classified 67% of communities as obesoprotective or obesogenic using the mean BMI z score as a surrogate. Social environment characteristics emerged as the most critical classifiers and might leverage intervention.
|
|
| Almeida et al [61], 2016 |
All BFP-grade predictive models presented a good global accuracy (≥91.3%) for obesity discrimination. Both overfat and obese as well as obese prediction models showed, respectively, good sensitivity (78.6% and 71%), specificity (98% and 99.2%), and reliability for positive or negative test results (≥82% and ≥96%).
For boys, the order of parameters, by relative weight in the predictive model, was BMI z score, height, WHtRq squared variable (_Q), age, weight, CCr_Q, and HCs_Q (adjusted R2=0.847 and RMSEt=2.852); for girls, it was BMI z score, WHtR_Q, height, age, HC_Q, and CC_Q (adjusted R2=0.872 and RMSE=2.171).
|
|
| Lingren et al [60], 2016 |
|
The rule-based exclusion algorithm performed better than the ML algorithm. The best feature set for ML used Unified Medical Language System concept unique identifiers; International Classification of Diseases, Ninth Revision, codes; and RxNorm codes.
|
| Seyednasrollah et al [59], 2017 |
|
WGRS19 improves the prediction of adulthood obesity. The model helps screen children with a high risk of developing obesity. Predictive accuracy is highest among young children (aged 3-6 years), whereas among older children (aged 9-18 years), the risk can be identified using childhood clinical factors.
|
| Hinojosa et al [58], 2018 |
Violent crime, English learners, socioeconomic disadvantage, fewer physical education and fully credentialed teachers, and diversity index were positively associated with obesity. By contrast, the academic performance index, physical education participation, mean educational attainment, and per capita income were negatively associated with obesity. The most highly ranked built or physical environment variables were distance to the nearest highway and green spaces, 10th and 11th most important, respectively.
|
|
| Maharana and Nsoesie [57], 2018 |
Features of the built environment explained 64.8% (RMSE=4.3) of the variation in obesity prevalence across all US census tracts. Individually, the variation explained was 55.8% (RMSE=3.2) for Seattle, Washington (213 census tracts); 56.1% (RMSE=4.2) for Los Angeles, California (993 census tracts); 73.3% (RMSE=4.5) for Memphis, Tennessee (178 census tracts); and 61.5% (RMSE=3.5) for San Antonio, Texas (311 census tracts).
|
|
| Wang et al [56], 2018 |
The SVM model significantly outperformed other classifiers based on the same training features. The SVM model exhibits 70.77% accuracy, 80.09% sensitivity, and 63.02% specificity.
The selected SNPsaa were effective in the detection of obesity risk.
|
|
| Duran et al [55], 2018 |
In female participants, the sensitivity of the BMI, WCbb, and ANN approaches to predict excess body fat was 0.751 (95% CI 0.730‐0.771), 0.523 (95% CI 0.487‐0.559), and 0.782 (95% CI 0.754‐0.810), respectively.
In male participants, the sensitivity of the BMI, WC, and ANN approaches to predict excess body fat was 0.721 (95% CI 0.699‐0.743), 0.572 (95% CI 0.549‐0.594), and 0.795 (95% CI 0.768‐0.821).
|
The diagnostic performance in identifying excess body fat was better in male participants when an ANN approach was used than when BMI and WC z scores were applied.
The ANN and BMI z scores performed comparably and significantly better, respectively, than WC z scores in female participants.
|
| Gerl et al [54], 2019 |
The lipidome, based on a LASSOcc model, predicted BFP the best (R2=0.73). In this model, the strongest positive predictor and strongest negative predictor were sphingomyelin molecules, which differ by only 1 double bond, implying the involvement of an unknown desaturase in obesity-related aberrations of lipid metabolism.
The regression was used to probe the clinically relevant information in the plasma lipidome and found that the plasma lipidome also includes information on body fat distribution because WHR (R2=0.65) was predicted more accurately than BMI (R2=0.47).
|
|
| Hammond et al [53], 2019 |
LASSO regression predicted obesity with an area under the receiver operating characteristic curve of 81.7% for girls and 76.1% for boys.
In each of the separate models for boys and girls, the researchers found that the weight-for-length z score, BMI between 19 and 24 months, and the last BMI measure recorded before the age of 2 years were the most important features for prediction.
|
Comparable to cohort-based studies, EHRdd data with area under the receiver operating characteristic curve values could be used to predict obesity at the age of 5 years, reducing the need for investment in additional data collection.
|
| Hong et al [52], 2019 |
As the results of the 4 ML classifiers showed, the RF algorithm performed the best with micro F1-score 0.9466 and macro F1-score 0.7887 and micro F1-score 0.9536 and macro F1-score 0.6524 for intuitive classification (reflecting medical professionals’ judgments) and textual classification (reflecting the decisions based on explicitly reported information of diseases), respectively.
The MIMICee-III obesity data set was successfully integrated for prediction with minimal configuration of the NLPff2FHIRgg pipeline and ML models.
|
|
| Ramyaa et al [51], 2019 |
SVM, neural network, and KNNhh algorithms performed modestly for the numerical predictions, with mean approximate errors of 6.70 kg, 6.98 kg, and 6.90 kg, respectively.
K-means cluster analysis improved prediction using numerical data and identified 10 clusters suggestive of phenotypes, with a minimum mean approximate error of approximately 1.1 kg. A classifier was used to phenotype participants into the identified clusters, with mean approximate errors of <5 kg for 15% of the test set (approximately, n=2000). SVM performed the best (54.5% accuracy), followed closely by the bagged tree ensemble and KNN algorithms.
|
SVM regression was the best-suited predictive and inferential tool for this task, closely followed by neural network and KNN algorithms. Although the overall data model showed a good fit and predictive ability, clustering produced relatively superior fit statistics.
|
| Scheinker et al [50], 2019 |
Multivariate linear regression and gradient boosting machine regression (the best-performing ML model) of obesity prevalence using all county-level demographic, socioeconomic, health care, and environmental factors had R2 values of 0.58 and 0.66, respectively (P<.001).
|
|
| Shin et al [49], 2019 |
The performance of the proposed system was compared with those of 2 commercial systems that were designed to measure body composition using either a whole body or upper body impedance value. The results showed that the correlation coefficient (R2) value was improved by approximately 9%, and the SE of the estimate was reduced by 28%.
|
|
| Stephens et al [48], 2019 |
|
|
| Blanes-Selva et al [47], 2020 |
|
The implementation of the PU learning methodology in identifying obesity produced results that were satisfactory, providing high sensitivity, and consistent with the World Health Organization’s obesity report.
|
| Dunstan et al [46], 2020 |
Using only 5 categories, RF could predict obesity prevalence with absolute error <10% for approximately 60% of the countries considered and absolute error <20% for 87%.
The most relevant food category with regard to predicting obesity consists of baked goods and flours, followed by cheese and carbonated drinks.
|
|
| Fu et al [45], 2020 |
The 2 most important features—trajectory of infant BMI z score change and maternal BMI at enrollment—were identified from the ML algorithm.
The aforementioned features showed similar predictive capacity compared with all features (area under the curve=0.68 vs 0.68; P=.83; DeLong test). The sensitivity analyses identified the same 2 features (ie, trajectory of infant BMI z score change and maternal BMI at enrollment), and the ranking of these features’ Shapley additive explanations value was unchanged.
In the independent test cohort, the area under the curve for childhood overweight and obesity classification using the aforementioned 2 features was 0.71 (95% CI 0.66 to 0.76), which was comparable to that based on all features (0.72, 95% CI 0.67 to 0.76).
|
An ML algorithm is applied to identify risk factors contributing to childhood overweight or obesity based on a large longitudinal study and addresses the relationships between all collected features and outcomes without any assumption.
A novel unified framework, Shapley additive explanations, is used to interpret predictions, and the identified predictive factors are robust.
|
| Kibble et al [44], 2020 |
New potential links between cytokines and weight gain are identified, as well as associations among dietary, inflammatory, and epigenetic factors.
|
|
| Park et al [43], 2020 |
The actual and predicted ΔBMI showed a significant intraclass correlation value with a low RMSE, and classification between people with increased BMI and those with nonincreased BMI resulted in a high area under the receiver operating characteristic curve value using only the degree centrality values obtained at the baseline visit.
|
|
| Phan et al [42], 2020 |
|
|
| Taghiyev [41], 2020 |
The proposed hybrid system demonstrated 91.4% accuracy, which is higher than that of other classifiers (ie, 4.6% higher than the performance of logistic regression and 2.3% higher than the performance of DTmm).
|
|
| Xiao et al [40], 2020 |
All aspects of horizontal greenery, vertical greenery, and proximity of green levels affected body weight; however, only the VGInn consistently had an adverse effect on weight and obesity.
|
|
| Yao et al [39], 2020 |
|
|
| Alkutbe et al [27], 2021 |
For the gradient boosting models, the predicted fat percentage values were more aligned with the actual value than those in regression models. Gradient boosting achieved better performance than the regression equation because it combined multiple simple models into a single composite model to take advantage of this weak classifier.
The developed predictive model archived RMSE values of 3.12 for girls and 2.48 for boys.
|
|
| Bhanu et al [38], 2021 |
The accuracy of segmentation was superficial SAT: 0.92, deep SAT: 0.88, and VAT: 0.9. The average Hausdorf distance was <5 mm. Automated segmentation significantly correlated R2>0.99 (P<.001) with ground truth for all 3-fat compartments. Predicted volumes were within 1.96 SD from Bland-Altman analysis.
|
DL-based, comprehensive superficial SAT, deep SAT, and VAT analysis tools showed high accuracy and reproducibility and provided a comprehensive fat compartment composition analysis and visualization in <10 seconds.
|
| Cheng et al [37], 2021 |
Physical activity was an important factor in predicting weight status, with gender, age, and race or ethnicity being less important factors associated with weight outcomes.
The durations of vigorous-intensity activity in 1 week and moderate-intensity activity in 1 week were essential attributes.
|
With physical activity and basic demographic information of all methods analyzed, the random subspace classifier algorithm achieved the highest overall accuracy and area under the receiver operating characteristic curve value.
In general, most algorithms showed similar performance.
Logistic regression was middle ranking in terms of overall accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve value among all methods.
|
| Delnevo et al [36], 2021 |
|
Certain psychological variables such as depression are highly predictive of BMI.
ML has several advantages over traditional statistics and can be used to compare the impact of many variables on predicting a chosen outcome and can handle various types of variables.
|
| Lee et al [35], 2021 |
For predicting a newborn’s BMI, linear regression (2.0744) and RF (2.1610) were better than ANN with 1, 2, and 3 hidden layers (150.7100, 154.7198, and 152.5843, respectively) in the mean squared error.
On the basis of variable importance from the RF, the major predictors of a newborn’s BMI were the first abdominal circumference value and estimated fetal weight in week 36 or later, gestational age at delivery, the first abdominal circumference value during week 21 to week 35, maternal BMI at delivery, maternal weight at delivery, and the first biparietal diameter value in week 36 or later.
|
|
| Lin et al [34], 2021 |
Metabolic healthy obesity (44% of the patients) was characterized by a relatively healthy metabolic status with the lowest incidents of comorbidity.
Hypermetabolic obesity–hyperuricemia (33% of the patients) was characterized by extremely high uric acid and an increased incidence of hyperuricemia (adjusted odds ratio 73.67 to metabolic healthy obesity, 95% CI 35.46-153.06).
Hypermetabolic obesity–hyperinsulinemia (8% of the patients) was distinguished by overcompensated insulin secretion and an increased incidence of polycystic ovary syndrome (adjusted odds ratio 14.44 to metabolic healthy obesity, 95% CI 1.75-118.99).
Hypometabolic obesity (15% of the patients) was characterized by extremely high glucose levels, decompensated insulin secretion, and the worst glucolipid metabolism (diabetes: adjusted odds ratio 105.85 to metabolic healthy obesity, 95% CI 42.00-266.74; metabolic syndrome: adjusted odds ratio 13.50 to metabolic healthy obesity, 95% CI 7.34-24.83).
|
|
| Pang et al [33], 2021 |
XGB yielded a mean area under the curve value of 0.81 (SD 0.001), which outperformed all other models. It also achieved a statistically significant better performance than all other models on standard classifier metrics (sensitivity fixed at 80%): precision, mean 30.9% (SD 0.22%); F1-score, mean 44.6% (SD 0.26%); accuracy, mean 66.14% (SD 0.41%); and specificity, mean 63.27% (SD 0.41%).
|
|
| Park et al [32], 2021 |
ML algorithms were used to determine the stances of tweets on Black Lives Matter. ML models showed better performance than lexicon-based sentiment analysis (accuracy: 61%). The NBoo model had an overall accuracy of 85%, slightly higher than that of the CNN model (83.8%); both had higher accuracy than the other models.
However, NB had the highest recall and F1-score for predicting the against stance, whereas CNN performed poorly on identifying the against stance.
|
|
| Rashmi et al [31], 2021 |
|
|
| Snekhalatha and Sangamithirai [30], 2021 |
Among the region of interest studied, the abdomen region exhibited a high temperature difference of 4.703% between normal participants and participants who were obese compared with other regions. The proposed custom network-2 provided an overall accuracy of 92%, with an area under the curve value of 0.948. By contrast, the pretrained model VGG16 produced an accuracy of 79% and an area under the curve value of 0.90 for discrimination into obese and normal thermograms.
|
The DL system based on custom CNN provided a reliable classification performance to identify the occurrence of obesity in test participants.
Custom CNN network-2 provided a commendable accuracy in classifying normal participants and participants who were obese from the thermal images.
The trained custom-2 CNN model can be used for computer-aided screening of test participants for obesity detection.
|
| Thamrin et al [29], 2021 |
Location, marital status, age group, education, sweet drinks, fatty or oily foods, grilled foods, preserved foods, seasoning powders, soft drinks or carbonated beverages, alcoholic beverages, mental or emotional disorders, diagnosed hypertension, physical activity, smoking, and fruit and vegetable consumption are significant in predicting obesity status in adults.
The classification prediction using the logistic regression method achieves the best performance based on the accuracy metric (72%), specificity (71%), precision (69%), kappa (44%), and Fβ-score (70%). Classification prediction by the classification and regression tree method achieves the highest sensitivity (82%) and the highest F1-score (72%).
With regard to the area under the receiver operating characteristic curve performance of the respective classification methods with 10-fold cross-validation, the logistic regression classifier has the highest average area under the receiver operating characteristic curve value (0.798).
|
Logistic regression has a better performance than the classification and regression tree and NB methods.
Kappa coefficients show only moderate concordance between predicted and measured obesity.
The constructed obesity classification model can evaluate and predict the risk of obesity using ML methods for the population of Indonesia, which can then be applied to publicly available open data.
|
| Zare et al [28], 2021 |
The kindergarten BMI z score is the most important predictor of obesity by grade 4.
Including the kindergarten BMI z score of students in the model meaningfully increases the prediction accuracy.
Logistic regression, RF, and neural network algorithms performed similarly in terms of accuracy, sensitivity, specificity, and area under the curve values. The 95% CIs around the area under the curve overlap among these 3 algorithms.
The DT showed lower performance with an area under the curve value that was statistically lower than the area under the curve values from each of the other algorithms. Nevertheless, the performance of the DT algorithm was close to that of the others.
|
Data from the Arkansas, United States, BMI screening program significantly improve the ability to identify children at a high risk of obesity to the extent that better prediction can be translated into more effective policy and better health outcomes.
The ability to predict obesity by grade 4 was robust across the ML algorithms and logistic regression with these data.
|