Skip to main content
China CDC Weekly logoLink to China CDC Weekly
. 2026 Sep 4;8(36):1135–1140. doi: 10.46234/ccdcw2026.184

A Hierarchical Multi-Stage Machine-Learning Study on Predicting Oil and Salt in Chinese Restaurant Dishes for Nutrition Surveillance — 6 Survey Sites, China, 2019–2020

Ying Xing 1, Jiguo Zhang 1, Xiaofang Jia 1, Chang Su 1, Wenwen Du 1,*, Huijun Wang 1, Aidong Liu 1
PMCID: PMC13586539  PMID: 42761454

Abstract

Introduction

The rise in dining out has created a demand for scalable methods for estimating the oil, salt, and sodium content in restaurant meals. These preparation-dependent components are challenging to measure consistently.

Methods

A hierarchical multi-stage prediction model was developed to estimate five dish-level targets: added cooking oil, salt, salt-containing seasonings, sodium from salt-containing seasonings, and total sodium from salt and seasonings in 8,455 Chinese dishes from 192 restaurants across six survey sites in China. The model performance was evaluated using standard regression metrics and task-specific indicators, including zero-use identification, tolerance-based accuracy, multiplicative consistency, symmetric mean absolute percentage error (SMAPE), Spearman's rank correlation, and discrete intake-level classification.

Results

The stage-wise light gradient-boosting machine (LightGBM) model outperformed the linear regression, random forest, and multi-task LightGBM models for all five targets. The log-scale coefficient of determination (R2) values ranged from 0.224 for salt to 0.570 for salt-containing seasonings. On the original scale, strict tolerance-based accuracy was limited, with 22%–37% of the predictions meeting the tolerance band. Multiplicative and rank-based performances were stronger: 40% to 73% of the predictions were within two-fold of the observed value, and the Spearman correlations ranged from 0.486 to 0.771. The predicted means were lower than the observed means, indicating a conservative estimation and limited suitability for uncalibrated mean estimation.

Conclusion

The model can support the preliminary estimation, relative comparison, and screening of restaurant dishes for nutrition surveillance. Therefore, it should not be interpreted as a precise dish-level measurement tool without further calibration.

Keywords: dietary oil and salt estimation, nutrient intake prediction, food composition inference, machine learning, population nutrition monitoring


Oil, salt, and sodium are key targets for dietary surveillance because their excess intake is associated with obesity, hypertension, and cardiovascular disease (1–3). Surveillance increasingly requires scalable estimates from restaurant dishes; however, direct weighing and laboratory analysis are difficult to scale (4). Existing nutrition estimation methods use food images, labels, recipes, or ingredient data to infer intrinsic nutrients (5–8). The added cooking oil, added salt, and salt-containing seasonings are different; they are preparation-dependent quantities that vary with cooking methods, chef habits, customer requests, recipe standardization, pre-prepared sauces, condiment brands, batch preparation, and kitchen practices. This study developed a hierarchical multistage machine learning framework to evaluate whether routinely collected dish information can support the preliminary estimation and comparison of these targets for nutrition surveillance.

METHODS

Data Sources and Study Targets

Data were obtained from 8,455 restaurant dishes collected during 2019–2020 from 192 restaurants across 6 survey sites selected to cover geographic and culinary diversity: Shijiazhuang, Hebei Province, North China; Xining, Qinghai Province, Northwest China; Daqing, Heilongjiang Province, Northeast China; Changsha, Hunan Province, Central China; Chengdu, Sichuan Province, Southwest China; and Nanchang, Jiangxi Province, East/Southeast China. A total of 2 districts/counties and 32 restaurants were surveyed in each province.

Trained staff collected dish names, ingredients, ingredient weights, cooking methods, and restaurant characteristics through interviews with chefs and on-site investigations. Identically named dishes from different restaurants were retained as separate records. Five targets were predicted: added cooking oil, salt-containing seasonings, sodium in salt-containing seasonings, salt, and total sodium from added salt and salt-containing seasonings

Data Preprocessing and Feature Description

All targets were log1p-transformed before model fitting, and observations above the 99.5th percentile for each target were excluded. Predictors included three feature blocks: ingredient composition, including total ingredient weight, main ingredient weight, and 65 food-category weight variables; cooking method, including 6 major categories and 22 subtypes; and restaurant-level characteristics, including survey site, north–south region, and restaurant size. The feature definitions are shown in Table 1, and the food category and cooking method classifications are shown in Supplementary Table S1 (available at https://weekly.chinacdc.cn/).

Table 1. Feature dictionary for model input variables.

Feature block Variable name(s) Definition Coding Source/derivation
Ingredient composition Total ingredient weight Total recorded weight of the dish Continuous (g) Calculated by summing all recorded ingredient weights
Main ingredient weight Weight of the dominant ingredient in the dish Continuous (g) Calculated from raw ingredient-level records
65 food-category weight variables Dish-level total weight assigned to each mapped food category after manual ingredient-to-category mapping Continuous (g), 0 if absent Raw item-level ingredients were mapped to the 65-category system and aggregated at the dish level; complete list shown in Table S1
Cooking method 6 major cooking-category indicators Major cooking category of the dish Binary indicator (0/1) Derived from the original coarse cooking-method records and standardized according to the study classification
22 cooking-subtype indicators Finer-grained cooking subtype of the dish Binary indicator (0/1) Manually refined from original cooking-method information according to the study classification scheme
Restaurant/site level 6 survey-site indicator variables Survey-site location of the restaurant Binary indicator (0/1) Directly derived from recorded restaurant location
North–south regional indicator Whether the restaurant was located in southern China Binary indicator (0/1) Derived from survey-site location
Restaurant size indicators Restaurant scale category (large, medium, small) Binary indicator (0/1) Directly recorded in field survey

Model Development and Design

The primary aim of this study was to predict five dish-level targets: added cooking oil, salt-containing seasonings, sodium in salt-containing seasonings, salt, and total sodium from seasonings and salt. These outcomes are related in both culinary practice and empirical data, and the pairwise Spearman correlations among them suggested that upstream information could be informative for downstream prediction (Figure 1A). We therefore used a stage-wise framework with five sequential stages (Figure 1B). Stage 1 predicted added cooking oil from the original dish features. Stage 2 predicted salt-containing seasonings using the same features plus the stage 1 output. Stage 3 predicted sodium in salt-containing seasonings using the updated feature set. Stage 4 predicted salt using the original features plus upstream stage outputs. Stage 5 predicted total sodium from seasonings and salt using all previously accumulated stage outputs.

Figure 1.

Figure 1

Target variable correlations, model framework, and distribution characteristics for predicting oil, salt, and sodium in Chinese restaurant dishes. (A) Spearman correlation heatmap of the five target variables. (B) Hierarchical multi-stage prediction framework. (C) Distribution characteristics of target variables.

Note: For (A), correlations were calculated among added cooking oil, salt-containing seasonings, sodium in salt-containing seasonings, added salt, and total sodium from seasonings and salt. For (B), original dish features from ingredient composition, cooking method, and restaurant/site-level feature blocks were used as initial inputs. The model predicted the five targets sequentially, and each upstream prediction was added to the input of the next stage. Solid arrows indicate forward prediction flow, and dashed arrows indicate out-of-fold predictions used during training to prevent information leakage. For (C), distributions of the five target variables are shown.

Four models were evaluated to compare different modeling strategies. Linear regression and random forest were implemented as alternative learners using the same stage-wise prediction logic. The multi-task light gradient-boosting machine (LightGBM) model predicted all five targets simultaneously from the same original input features without stage-to-stage feedback. The stage-wise LightGBM model predicted the five targets sequentially and passed predictions forward from one stage to the next.

Model Evaluation and Performance Metrics

Model performance was evaluated using the coefficient of determination (R2), mean absolute error (MAE), and root mean squared error (RMSE) on a log-transformed scale. As the primary goal of this study was to support the rapid estimation of oil, salt, and sodium intake in large-scale nutrition surveillance or surveys, task-specific metrics were also calculated on the original scale, including zero-use identification, tolerance-based accuracy, multiplicative consistency, symmetric mean absolute percentage error (SMAPE), Spearman rank correlation, and discrete intake-level classification. Definitions of these metrics are provided in Supplementary Table S2 (available at https://weekly.chinacdc.cn/).

Statistical Analysis

The dataset was split 80/20 into training and test sets. Five-fold cross-validation was applied to the training set. To prevent information leakage, out-of-fold upstream predictions were used as downstream training inputs and the test-set upstream predictions were passed forward during the final evaluation. Subgroup performance was evaluated according to survey site, north-south region, and major cooking-method category. Feature-group contributions were assessed using grouped Shapley additive explanations (SHAP) values for ingredient composition, cooking method, and restaurant-level features.

All analyses were conducted in Python 3.11 using pandas, Numpy, scikit-learn, and LightGBM.

RESULTS

Data Description

The full dataset covered six major cooking categories (Supplementary Table S1). Mean values per dish were 24.72 g for cooking oil, 21.79 g for salt-containing seasonings, 3.46 g for salt, 1,275.57 mg for sodium in salt-containing seasonings, and 2,663.71 mg for total sodium. The targets were right-skewed and heavy-tailed, consistent with Figure 1C. Of the 8,455 records, 3,729 (44.1%) shared dish names, including 987 repeated dish names. Same-name dishes still varied substantially: median within-name ranges were 27.0 g for cooking oil, 13.5 g for salt-containing seasonings, and 3.0 g for added salt.

Model Results

Statistical metrics on log scale

On a log-transformed scale, the stage-wise LightGBM model performed the best for all five targets (Table 2). The R2 values were 0.550 for cooking oil, 0.570 for salt-containing seasonings, 0.454 for sodium in salt-containing seasonings, 0.224 for added salt, and 0.394 for total sodium from seasonings and salt. Subgroup analyses by survey site, north-south region, and major cooking-method category generally supported the stability of the results, although added salt remained the weakest target (Supplementary Table S3, available at https://weekly.chinacdc.cn/).

Table 2. Test-set performance of four regression models*.
Target Model R 2 MAE RMSE
Abbreviation: R²=coefficient of determination; MAE=mean absolute error; RMSE=root mean squared error.
* All outcomes were log-transformed.
† indicates the primary model used in this study.
Cooking oil Linear regression 0.432 0.878 1.226
Random forest 0.438 0.989 1.225
LightGBM multi-task 0.527 0.860 1.124
LightGBM stage-wise (proposed) † 0.550 0.831 1.069
Salt-containing seasonings Linear regression 0.527 0.664 0.870
Random forest 0.393 0.762 0.994
LightGBM multi-task 0.544 0.656 0.862
LightGBM stage-wise (proposed)† 0.570 0.632 0.813
Sodium in salt-containing seasonings Linear regression 0.310 0.662 0.925
Random forest 0.323 0.727 0.931
LightGBM multi-task 0.426 0.655 0.858
LightGBM stage-wise (proposed)† 0.454 0.624 0.806
Salt Linear regression 0.142 0.557 0.711
Random forest 0.175 0.547 0.704
LightGBM multi-task 0.165 0.539 0.708
LightGBM stage-wise (proposed)† 0.224 0.518 0.666
Total sodium in seasonings and salt Linear regression 0.303 0.528 0.746
Random forest 0.288 0.553 0.790
LightGBM multi-task 0.327 0.545 0.768
LightGBM stage-wise (proposed)† 0.394 0.506 0.707

Grouped SHAP analysis showed that ingredient-composition, cooking-method, and restaurant-level features contributed 49.8%, 37.3%, and 12.9% for cooking oil; 52.2%, 37.1%, and 10.7% for salt-containing seasonings; 58.8%, 17.9%, and 23.3% for sodium in salt-containing seasonings; 40.8%, 14.8%, and 44.4% for added salt; and 69.8%, 16.4%, and 13.8% for total sodium, respectively.

Task-specific metrics on original scale

The predicted means and standard deviations were lower than the observed values for all five targets (Table 3), indicating a conservative shrinkage of high values. Thus, raw predictions should be interpreted as preliminary reference values rather than exact dish-level measurements, and calibration is required for absolute mean intake estimation. On the original scale, BandAcc ranged from 0.217 to 0.374, WIF@2× from 0.403 to 0.729, and Spearman’s ρ from 0.486 to 0.771, indicating better relative ranking than precise gram-level agreement (Supplementary Table S2).

Table 3. Observed and predicted mean values and standard deviations of the five target variables in the test set.
Target Observed mean Observed SD Predicted mean Predicted SD Mean difference
(Pred - Obs)
Abbreviation: SD=standard deviation; Pred=predicted; Obs=observed.
Cooking oil (g) 24.84 34.84 17.13 20.18 −7.70
Salt-containing seasonings (g) 21.63 41.35 14.31 16.01 −7.32
Sodium in salt-containing seasonings (mg) 1,301.05 2,487.01 888.15 886.21 −412.90
Salt (g) 3.55 4.59 2.65 1.35 −0.89
Total sodium in seasonings and salt (mg) 2,745.68 3,633.57 2,162.92 1,267.24 −582.75

DISCUSSION

This study evaluated whether routinely collected restaurant dish information could support the estimation of added oil, salt, salt-containing seasonings, and sodium-related outcomes for nutrition surveillance. Stage-wise LightGBM outperformed the comparison models and showed generally stable performance across survey sites, regions, and cooking-method categories. However, the model should be interpreted as a surveillance support tool for preliminary estimation and comparison, and not as a substitute for direct measurement.

This task differs from conventional nutrient estimation. Intrinsic nutrients can often be estimated from ingredient weights and food composition tables (5–10), whereas added oil, salt, and salt-containing seasonings depend on preparation practices. Chef habits, customer requests, recipe standardization, pre-prepared sauces, condiment brands, batch preparation, and kitchen practices are difficult to measure consistently. The repeated-dish analysis supports this point: identical dish names do not necessarily represent standardized recipes or stable seasoning amounts.

The feature-contribution results also support the use of multiple information sources. Ingredient composition captured much of the dish-level structure, cooking methods captured preparation-related variation, and restaurant-level variables captured part of geographic or operational variation, especially for added salt. Restaurant size was included as a predictor; however, it is only a coarse indicator and cannot directly measure recipe standardization or chef-level practices.

The model's public health value lies in its screening and prioritization. In large-scale monitoring, direct measurement of restaurant dishes is often impractical. Approximate model outputs can help field investigators identify dishes, cooking methods, regions, or restaurant types that are likely to have high oil or sodium content and prioritize them for further assessment or intervention. Raw predictions should not be used for menu labeling, regulatory enforcement, individual dietary assessment, or uncalibrated estimation of the population mean intake.

The findings in this report are subject to some limitations. First, although the dataset covers diverse survey sites, it was not designed to generate province-specific estimates. Second, the target values derived from chef interviews and on-site investigations may contain measurement errors. Third, routinely collected predictors did not include condiment brands, chef-specific habits, customer requests, or detailed cooking process information. Fourth, external validation and calibration are required before broader application.

In conclusion, routinely collected dish information can support the preliminary estimation and relative comparison of oil, salt, and sodium in Chinese restaurant dishes, while revealing clear limits for precise gram-level prediction.

SUPPLEMENTARY DATA

Supplementary data to this article can be found online.

ccdcw-8-36-1135-S1.pdf (246.4KB, pdf)

Funding Statement

Supported by Public Health Emergency Project Nutrition Health and Healthy Diet Campaign (No. 102393220020070000012)

Conflicts of interest

No conflicts of interest.

References

  • 1.Allison SJ High salt intake as a driver of obesity. Nat Rev Nephrol. 2018;14(5):285. doi: 10.1038/nrneph.2018.23. [DOI] [PubMed] [Google Scholar]
  • 2.Watso JC, Fancher IS, Gomez DH, Hutchison ZJ, Gutiérrez OM, Robinson AT The damaging duo: obesity and excess dietary salt contribute to hypertension and cardiovascular disease. Obes Rev. 2023;24(8):e13589. doi: 10.1111/obr.13589. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zhou L, Stamler J, Chan Q, Van Horn L, Daviglus ML, Dyer AR, et al Salt intake and prevalence of overweight/obesity in Japan, China, the United Kingdom, and the United States: the INTERMAP Study. Am J Clin Nutr. 2019;110(1):34–40. doi: 10.1093/ajcn/nqz067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.McLean RM Measuring population sodium intake: a review of methods. Nutrients. 2014;6(11):4651–62. doi: 10.3390/nu6114651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Yunus R, Arif O, Afzal H, Amjad MF, Abbas H, Bokhari HN, et al A framework to estimate the nutritional value of food in real time using deep learning techniques. IEEE Access. 2019;7:2643–52. doi: 10.1109/ACCESS.2018.2879117. [DOI] [Google Scholar]
  • 6.Lei ZF, ul Haq A, Dorraki M, Zhang DF, Abbott D Composing recipes based on nutrients in food in a machine learning context. Neurocomputing. 2020;415:382–96. doi: 10.1016/j.neucom.2020.08.071. [DOI] [Google Scholar]
  • 7.Shen ZD, Shehzad A, Chen S, Sun H, Liu J Machine learning based approach on food recognition and nutrition estimation. Procedia Comput Sci. 2020;174:448–53. doi: 10.1016/j.procs.2020.06.113. [DOI] [Google Scholar]
  • 8.Al-Saffar M, Baiee W Nutrition information estimation from food photos using machine learning based on multiple datasets. Bull Electr Eng Inform. 2022;11(5):2922–9. doi: 10.11591/eei.v11i5.4007. [DOI] [Google Scholar]
  • 9.Ispirova G, Eftimov T, Koroušić Seljak B P-NUT: predicting nutrient content from short text descriptions. Mathematics. 2020;8(10):1811. doi: 10.3390/math8101811. [DOI] [Google Scholar]
  • 10.Li JT, Han FD, Guerrero R, Pavlovic V. Picture-to-amount (PITA): predicting relative ingredient amounts from food images. In: Proceedings of the 2020 25th international conference on pattern recognition. Milan, Italy: IEEE. 2021; p. 10343-50. http://dx.doi.org/10.1109/icpr48806.2021.9412828.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary data to this article can be found online.

ccdcw-8-36-1135-S1.pdf (246.4KB, pdf)

Articles from China CDC Weekly are provided here courtesy of Chinese Center for Disease Control and Prevention

RESOURCES