Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 18.
Published in final edited form as: Ultrasound Med Biol. 2025 Oct 16;52(1):207–215. doi: 10.1016/j.ultrasmedbio.2025.09.016

Enhancing Newborn Health Assessment: Ultrasound-based Body Composition Prediction Using Deep Learning Techniques

Keshi He a, Julia Hohenberg b, Yi Li b, Audrey Xiao a, Hayoung Cho a, Emily Nagel c, Sara Ramel c, Katherine A Bell d,e, Donglai Wei b, Jinhee Park f, Bryan J Ranger a,f,*
PMCID: PMC12614439  NIHMSID: NIHMS2112232  PMID: 41102035

Abstract

Objective:

This study investigates the feasibility of deep learning to predict body composition with ultrasound, specifically fat mass (FM) and fat free mass (FFM), to improve newborn health assessments.

Methods:

We analyzed 721 ultrasound images of the biceps, quadriceps, and abdomen from 65 preterm infants. A deep learning model incorporating a modified U-Net architecture was developed to predict FM and FFM, using air displacement plethysmography (ADP) as ground-truth labels for training. Model performance was assessed using mean absolute error (MAE), mean squared error (MSE), root mean square error (RMSE), and mean absolute percentage error (MAPE), along with Bland-Altman plots to evaluate mean bias and limits of agreement. We tested different image combinations to determine the contribution of anatomical regions. Grad-CAM was applied to identify image regions with the strongest influence on predictions.

Results:

Combining biceps, quadriceps, and abdomen ultrasound images to predict whole-body composition showed strong agreement with ground truth values, with low MAE (FM: 0.0145 kg, FFM: 0.0794 kg), MSE (FM: 0.0003 kg2, FFM: 0.0073 kg2), RMSE (FM: 0.0183 kg, FFM: 0.0854 kg), and MAPE (FM: 2.65%, FFM: 8.40%). Using only abdomen images for prediction improved FFM performance (MAPE=4.62%, MSE=0.0041 kg2, RMSE=0.0486 kg, MAE=0.0378 kg). Grad-CAM revealed muscle regions as key contributors to FM and FFM predictions.

Conclusion:

Deep learning provides a promising approach for predicting body composition with ultrasound and a valuable tool for assessing nutritional status in neonatal care.

Keywords: Body composition prediction, Deep learning, Ultrasound imaging, Malnutrition, Newborn and child health

Introduction

Malnutrition represents a significant public health challenge, affecting over 45 million children under five years of age globally.1 Approximately one in eight children does not receive adequate nutrition to support healthy growth and development.2 At the same time, the rising prevalence of obesity is also influenced by early-life growth patterns, with rapid weight gain in infancy increasing the risk of obesity later in childhood and adulthood.3 Accurate monitoring of children’s growth patterns is essential for identifying nutritional risk and timely intervention. However, current methods primarily rely on simple anthropometric measurements such as height, weight, and body mass index (BMI), which do not provide detailed insights into whether weight gain is due to fat accumulation or healthy muscle development.4,5

Assessing body composition, including fat mass (FM) and fat-free mass (FFM), is valuable for evaluating nutritional status, health outcomes, and responses to interventions. However, accurate measurement typically requires specialized equipment, such as dual-energy X-ray absorptiometry (DXA), computed tomography (CT), or air displacement plethysmography (ADP). These methods are costly, require trained personnel, and are often inaccessible in many clinical and resource-limited settings.

Ultrasound offers a portable, affordable, noninvasive, and radiation-free approach to body composition assessment.6 It has been applied in various populations, including older adults, children with nephrotic syndrome, and individuals with disease-related malnutrition, to assess fat and muscle.7–9 Compared to DXA and CT, ultrasound is safer and more suitable for repeated use. However, its use in automated body composition analysis remains limited. Most existing methods rely on manual measurement of single-plane fat or muscle thickness, which limits reproducibility and does not utilize the full potential of image data. Automated methods leveraging multidimensional features may enable more robust and scalable assessments.

The advantages of ultrasound are particularly relevant for vulnerable populations, such as preterm infants, yet its application in this population remains limited. Prior work has shown ultrasound can reliably measure muscle and fat thickness in premature infants,10 but efforts to utilize these measurements to predict whole-body composition have shown poor accuracy and low predictive ability.11 Importantly, prior models relied on human operators to extract single-plane fat or muscle thickness from ultrasound images, limiting reproducibility and data utility. Integrating ultrasound with deep learning-based automated prediction methods may better leverage multidimensional imaging data and improve predictive accuracy.

Deep learning has emerged as a powerful tool for medical image analysis, including ultrasound image segmentation and classification.12,13 It has also been applied to body composition prediction using CT for image acquisition and segmentation.14–17 However, its use with ultrasound for this purpose remains unexplored. Given ultrasound’s ability to distinguish adipose from muscle tissue and the growing availability of low-cost portable systems, deep learning integrated with ultrasound presents a promising method to assess nutritional and growth metrics.18,19

The objective of this study is to develop a deep learning-based end-to-end method for automated prediction of body composition with ultrasound. Specifically, we aim to validate the feasibility of a deep learning model trained with ultrasound images obtained from preterm infants and ADP-measured ground-truth body composition data, and to evaluate specific model components that contribute to whole-body composition prediction. This approach has the potential to enhance body composition assessment, providing a practical, accurate, and scalable solution for clinical implementation.

Materials and Methods

Subjects

We conducted a secondary analysis of ultrasound images and clinical data from an observational cohort study at the University of Minnesota Medical Center.11 The original study enrolled 68 preterm infants born between 25- and 34-weeks’ gestation who were clinically stable (i.e., not requiring respiratory support or intravenous fluids). Ultrasound imaging of the arm, leg, and abdomen was performed alongside body composition measurement using ADP. Postmenstrual age (PMA) at the time of measurement ranged from 32.60 to 38.70 weeks. For this secondary analysis, we included 65 infants with available ultrasound data. For clinical data, descriptive characteristics of the cohort are summarized in Table 1. Only FM and FFM from ADP were used as ground truth for model training.

Table 1.

Descriptive characteristics infant cohort (N=65)

Variable Mean (SD) or Number (%) Range
Gestational age at birth (weeks) 32.04 (2.05) 25.00–34.90
Birthweight (kg) 1.73 (0.49) 0.75-2.94
Sex, female 36 (55.38%) NA
Length of hospital stay (days) a 38.70 (22.73) 10.0–96.0
PMA at discharge (weeks) a 37.45 (1.79) 34.74–43.81
PMA at study measurement (weeks) 35.13 (1.20) 32.60–38.70
Weight at study measurement (kg) 2.16 (0.36) 1.38–3.09
Length at study measurement (cm) 43.97 (2.30) 38.90–49.00
Body fat percentage (%) 8.81 (4.11) 1.70–19.70
Fat-free body percentage (%) 91.19 (4.11) 80.30–98.30
Fat mass (kg) 0.19 (0.10) 0.03–0.45
Fat-free mass (kg) 1.96 (0.33) 1.27–2.83

Abbreviations: PMA, postmenstrual age.

a

n=61 due to missing data.

Data curation

A dataset of 721 ultrasound images from the 65 infants was curated. Images were collected at three anatomic sites – the biceps (brachii and brachialis), abdomen (rectus abdominis), and quadriceps (rectus femoris and vastus intermedius) – using a B-mode ultrasound imaging system (NextGen LOGIQ e R7, GE Medical Systems, Chicago, IL).11 At least three images were acquired per region (Figure 1).

Figure 1.

Figure 1.

Clinical ultrasound images of a) biceps, b) abdomen, and c) quadriceps.

In the ultrasound images, subcutaneous fat appears as bright layers just beneath the skin surface, while muscle layers are visible beneath the fat and distinguished by darker coloration. In biceps and quadriceps images, deeper layers reveal the underlying humerus or femur. Gold standard body composition values (FM, FFM) were obtained using ADP (PeaPod, Cosmed, Ltd, Concord, CA). Figure 2 presents a flowchart detailing subject selection and the data processing workflow.

Figure 2.

Figure 2.

Study flow chart.

Data pre-processing

Ultrasound images in DICOM format were converted to JPEG to reduce file size and ensure compatibility with image processing workflows. The dataset was cleaned through a systematic quality control process. Exclusion criteria included: (i) presence of strong acoustic shadowing or reverberation artifacts, (ii) insufficient contrast between tissue layers, (iii) truncated or inappropriate fields of view, including imaging depth set too shallow to capture the full neonatal region of interest or too deep such that superficial structures were inadequately resolved, and (iv) visible tissue compression or distortion of anatomical boundaries caused by excessive probe pressure. These criteria were applied systematically to ensure consistency across the dataset, and each excluded image was carefully reviewed by at least two researchers to ensure that only valid scans were retained.

A series of preprocessing steps, illustrated in Figure 3, was subsequently applied to the remaining images. Denoising was performed using a median filter, with a pixel size of 5, which we determined through empirical testing to provide the most effective balance between noise reduction and preservation of anatomical detail. The images were then resized to 256×256 pixels to meet model input requirements. Finally, standard normalization with a mean of 0.5 and a standard deviation of 0.5 was applied and standardization were performed to ensure that all images were on the same scale.

Figure 3.

Figure 3.

Ultrasound image preprocessing. The DICOM files of biceps, abdomen, and quadriceps were converted to JPG. Images were then cropped, resized, normalized, and denoised prior to model training.

Data augmentation was used to enhance generalizability and reduce overfitting. Augmentations included horizontal and vertical flipping, small-angle rotations (±10°), and minor adjustments to brightness and contrast to simulate variability while preserving anatomical integrity. Five augmentations per image provided optimal balance between computational efficiency and predictive performance. This approach is commonly used in deep learning applications to expand limited datasets without requiring additional labeled data.20 Augmented images were qualitatively reviewed to ensure anatomical fidelity, and stable model performance across augmentation levels demonstrated that the approach increased data diversity without causing lossy reduction of critical anatomical details.

Workflow for deep learning-based body composition prediction

Figure 4 illustrates the workflow for deep learning-based automated body composition prediction. Preprocessed ultrasound images, along with ground truth (i.e., ADP-measured FM and FFM), were used to train the model. The trained model, upon deployment, can process ultrasound images and automatically generate end-to-end predictions of FM and FFM.

Figure 4.

Figure 4.

Workflow for ultrasound-based automated body composition prediction.

Deep learning model description

As part of model development, we explored several commonly used architectures, including U-Net21, EfficientNet22, ResNet23, and Attention U-Net24. U-Net was selected based on its balance of computational efficiency, architectural simplicity, and predictive performance. The architecture of the deep learning model, based on the U-Net framework, is illustrated in Figure 5. U-Net employs an encoder-decoder structure that captures local and global image features to predict FM and FFM from ultrasound images.21 For this study, we employed a modified U-Net architecture tailored for the ultrasound regression prediction task. U-Net is frequently used in image segmentation due to its ability to capture spatial hierarchies and contextual relationships within images. This characteristic makes it well-suited for predicting body composition, as structural features play a key role. The encoder-decoder structure facilitates effective feature extraction, while skip connections preserve spatial details essential for accurate prediction.

Figure 5.

Figure 5.

The architecture of our U-Net model. Input from left to right: bicep, abdomen, and quadriceps images, with two samples provided for each region.

Each training batch consisted of a PyTorch tensor containing six ultrasound images per infant – two images each from the abdomen, biceps, and quadriceps. This configuration balanced complexity and computational efficiency. The encoder pathway progressively downsampled the input, reducing the spatial dimensions (from 256×256 to 16×16), while increasing the number of channels (from 64 to 512) to enhance feature capture. The decoder pathway then upsampled the feature maps, restoring the spatial dimensions to 256×256, while decreasing the number of channels back to 64. To mitigate overfitting, dropout layers with a rate of 0.1 were incorporated throughout the network. The final feature map was flattened to (1, 256×256) and passed through linear layers to aggregate information across anatomical regions to generate predictions of FM and FFM. A sigmoid activation function was applied to the final output to normalize predictions and improve training stability.

Model training

A 5-fold cross-validation approach was used to improve training robustness and reduce overfitting, particularly given the limited dataset. Ten percent of the dataset was held out as an independent test set (Figure 2), while the remaining data were used to train and validate the model. In each fold, 72% of the dataset was used for training and 18% used for validation. Model training, hyperparameter tuning using the validation set, and performance evaluation on the test set were performed iteratively to optimize model performance. The model was trained using stochastic gradient descent with a momentum of 0.9, a mini-batch batch size of 32, and a fixed learning rate schedule of 0.001. The Adam optimization algorithm was employed to fine-tune the model.25 A regression sequence duration of 1.0 seconds was selected to balance temporal information and computational efficiency. The loss function integrated mean square error (MSE) and L1 loss to improve accuracy and robustness, while penalizing negative predictions. The number of U-Net layers was optimized based on runtime and performance, and further tuning was not pursued, as deeper architectures yielded minimal gains and longer training times.

Model evaluation

The model’s performance was evaluated using a validation dataset and quantified with standard regression metrics, including root mean square error (RMSE), mean absolute percentage error (MAPE), mean absolute error (MAE), and MSE.

Model implementation

The deep learning models were implemented using Python 3.9 and the PyTorch framework. The computational environment consisted of a Dell G15 Gaming Laptop with a NVIDIA® GeForce RTX™ 3060 GPU and a 11th Generation Intel® Core™ i7-11800H processor. Additionally, a T4 GPU hosted on Google Colab was utilized to accelerate the training process.

Model computation

During model computation, the training and testing loss curves exhibited an expected decline in the initial stages, followed by a gradual decrease and stabilization with minimal variation after 90 epochs. The total duration of the model computation was approximately 20 minutes, depending on the specific computational configuration.

Model validation

Model performance was evaluated using: (a) Bland-Altman plots to assess agreement between predicted and actual values, highlighting any systematic bias or outliers; (b) model errors to quantify the difference between predicted and actual values, providing a gauge for prediction accuracy; and (c) scatterplots for visual inspection of model predictions.

Comparison to classical feature-based regression methods

To compare performance with traditional approaches, classical feature extraction methods including Histogram of Oriented Gradients (HOG) and Scale-Invariant Feature Transform (SIFT) were applied to ultrasound images. Features were encoded using a Bag of Visual Words (BoVW) model and subsequently used as input to a Random Forest regression model for predicting FM and FFM.

Evaluation of anatomical region contributions to prediction

To assess the contribution of each anatomical region to body composition prediction, we tested image combinations derived from the biceps (B), abdomen (A), and quadriceps (Q). Initially, we trained the model with ultrasound images from all three anatomical regions (BAQ) to establish baseline performance. Subsequently, we iteratively excluded one region at a time to quantitatively evaluate how its omission affected the model’s predictive accuracy for FM and FFM, using MAPE as the evaluation metric. This analysis helped identify which regions contributed most to body composition prediction, providing insights into which regions should be prioritized for scanning in future imaging protocols.

Visual identification of image region contributions to prediction

Model interpretability and transparency are essential in healthcare applications. To elucidate how the model makes predictions, we employed feature importance analysis and visualization techniques, including Grad-CAM, to identify regions within the images that contributed to the model’s prediction.26,27 Grad-CAM computes the gradient of the trained model with respect to the input data and processes it into heatmaps that highlight the areas most significant to the final prediction.26 When applied to the three anatomic sites (abdomen, biceps, and quadriceps) at the final layer of the U-Net architecture, Grad-CAM identifies key regions in the ultrasound images that influence FM and FFM prediction, with warmer colored (red) regions indicating greater importance.

Results

Body composition prediction

Figure 6 presents comparisons between the predicted and actual values of FM. The Bland-Altman plot in Figure 6(a) illustrates strong consistency between the measured values and the model predictions for FM, with minimal mean bias of −0.0052 kg and limits of agreement ranging from −0.0516 to 0.0412 kg. The deep learning model yielded the following prediction results: MAE = 0.0145 kg, MSE = 0.0003 kg2, RMSE = 0.0183 kg, and MAPE = 2.65%, as shown in Table 2 and Figure 6(b) and Figure 6(c). Overall, there is good agreement between ground truth and predictions.

Figure 6.

Figure 6.

Comparison of prediction and real value of fat mass (UNet FM BAQ). a) Bland-Altman plot; b) model error; c) model prediction; and d) ground truth vs prediction.

Table 2.

Results of image combinations for prediction of FM

Metrics B A Q BA BQ AQ BAQ
MAE (kg) 0.0301 0.0167 0.0167 0.0492 0.0179 0.0179 0.0145
MSE (kg2) 0.0011 0.0004 0.0039 0.0051 0.0009 0.0005 0.0003
RMSE (kg) 0.0338 0.0207 0.0197 0.0599 0.0312 0.0224 0.0183
MAPE 5.48% 3.06% 3.11% 6.37% 4.35% 3.29% 2.65%

Figure 7 presents a systematic comparison between predicted and actual values for FFM. The Bland-Altman plot in Figure 7(a) shows strong consistency between the measured values and model predictions, with a low mean bias of −0.0613 kg. However, the limits of agreement were wider than those for FM, ranging from −0.2172 to 0.0947 kg. The FFM prediction showed MAPE of 8.40%, while the MAE (0.0749 kg), MSE (0.0073 kg2), and RMSE (0.0854 kg) were larger compared to the FM prediction, as shown in Table 3 and Figure 7(b) and Figure 7(c). Despite this, the ground truth and predictions demonstrate acceptable differences.

Figure 7.

Figure 7.

Comparison of prediction and real value of fat-free mass (UNet FFM BAQ). a) Bland-Altman plot; b) model error; c) model prediction; and d) ground truth vs prediction.

Table 3.

Results of image combinations for prediction of FFM

Metrics B A Q BA BQ AQ BAQ
MAE (kg) 0.046 0.0378 0.0627 0.0551 0.0905 0.0843 0.0749
MSE (kg2) 0.0031 0.0041 0.0064 0.0048 0.0098 0.0084 0.0073
RMSE (kg) 0.0561 0.0486 0.0801 0.0693 0.0993 0.0918 0.0854
MAPE 5.26% 4.62% 7.33% 6.41% 10.09% 9.39% 8.40%

As shown in Table 4, our model outperformed traditional feature-based regression approaches, which used HOG and SIFT features with Random Forest regression.

Table 4.

Performance comparison of models

Mean FM prediction performance Mean FFM prediction performance
MAE (kg) MSE (kg2) RMSE (kg) MAPE MAE (kg) MSE (kg2) RMSE (kg) MAPE
SIFT-based model 0.0879 0.0120 0.1083 73.68% 0.2722 0.1111 0.3301 13.23%
HOG-based model 0.0858 0.0108 0.1021 70.20% 0.2752 0.1114 0.3312 14.19%
Our model 0.0233 0.0017 0.0294 4.04% 0.0645 0.0063 0.0758 7.36%

Abbreviations: FM, fat mass. FFM, fat-free mass. B, biceps image only. A, abdominal image only. Q, quadriceps image only. AB, abdominal and biceps images. BQ, biceps and quadriceps images. AQ, abdominal and quadriceps images. BAQ, biceps, abdominal, and quadriceps images.

Anatomical region contributions to prediction

Table 2 and Table 3 provide the performance of predicted FM and FFM for each body region. For FM prediction, the combination of images from all three anatomical regions (BAQ) resulted in the lowest MAPE (2.65%), MSE (0.0003 kg2), RMSE (0.0183 kg) and MAE (0.0145 kg). For FFM prediction, we found the lowest MAPE (4.62%), MSE (0.0041 kg2), RMSE (0.0486 kg) and MAE (0.0378 kg) when using only abdomen images as input (A) to the model. These findings suggest that abdomen images alone provide a strong contribution to accurate body composition prediction, particularly for FFM.

Image regions that contribute to predictions

Figure 8 presents visualization results from a representative infant, identifying key regions within the ultrasound images of the abdomen, biceps, and quadriceps that contribute to model predictions, as identified by Grad-CAM. Warmer colors (reds) indicate regions of higher importance in the model’s prediction, while cooler colors (blues) indicate lower importance. Separate heatmaps were generated for FM and FFM predictions at each anatomic site (Figure 8). Both FM and FFM predictions emphasized similar areas, highlighting fat regions in blue and nonfat regions in red. This suggests that the model primarily utilizes areas with more muscle tissue for prediction, while subcutaneous fat and bone regions (femur and humerus) contribute less. Notably, the abdomen shows the greatest distinction between fat and muscle in its heatmap, indicating that subcutaneous abdominal fat plays a minimal role in predictions. In contrast, fat overlying the biceps and quadriceps contributed more significantly, although muscle tissue still took precedence in the overall predictions.

Figure 8.

Figure 8.

The contributory regions in the ultrasound images of abdomen, biceps and quadriceps identified by Grad-CAM. Warmer colors (reds) indicate higher importance in the model’s prediction, while cooler colors (blue) indicate lower importance in the model’s prediction.

Discussion

Accurately predicting body composition from ultrasound images is a complex task influenced by several factors, including the quality and quantity of imaging data, image preprocessing, and the selection of modeling techniques. Clinically, bedside estimation of body composition could improve neonatal care by enabling more frequent, non-invasive monitoring of growth and nutrition in premature infants. Ultrasound is portable and repeatable, making it a useful complement to existing methods such as ADP and extending assessments to community clinics and resource-limited settings. This complexity makes it well suited to a deep learning approach, which can quickly perform sophisticated analyses and fully leverage rich imaging data. In practice, manual analysis by human operators often requires a priori selection of regions in the image, along with simplified measures, such as single plane tissue thickness. In contrast, deep learning can independently identify important imaging regions for prediction and effectively utilize all available data within a region.

In this study, we validated the feasibility of applying deep learning techniques to B-mode ultrasound images for predicting body composition in premature infants. Our preliminary results demonstrate that combining ultrasound imaging with deep learning shows promise as a non-invasive method for predicting body composition. The model demonstrated a low mean bias of −0.0052 kg for FM and −0.0613 for FFM, indicating good agreement with ground truth measures. The primary innovation lies in the first demonstration of ultrasound-based automatic prediction of FM and FFM in preterm infants by using deep learning. We also incorporated multiple anatomical regions, evaluated their individual and combined predictive value, and used Grad-CAM to support interpretability – distinguishing this work from previous approaches.

While gold standard methods achieve a low error of 3-5%, our model demonstrated MAPE results ranging from 2.65-10.09%, as shown in Table 2 and Table 3.28,29 Deep learning-based models are heavily dependent on the dataset used for training. This initial analysis was based on a limited sample size, suggesting that larger and more diverse datasets may improve performance. Many body composition prediction methods, such as ADP, incorporate clinical data such as infant sex, weight, length, and age. Our model, however, demonstrated reasonable agreement using imaging data alone, which is advantageous since certain clinical data such as infant length can be challenging to measure accurately.30 Compared with traditional models, our model offers stronger overall predictive value and promising applicability.

Analysis of the contributions of different anatomical regions to body composition prediction revealed that abdomen images were particularly useful for accurate prediction. Using only abdomen images for FFM prediction resulted in the lowest error values. For FM, the abdomen image alone had only marginally higher error than using images from all three anatomical regions. Though preliminary, this finding is important for optimizing ultrasound scanning protocols for body composition assessments. Given that abdomen-only images performed well for FFM prediction, it may be worth exploring in future studies whether scanning just the abdomen is sufficient for accurate body composition prediction as compared to scanning all three anatomical regions. Streamlining the scanning protocol to focus on fewer regions could reduce scanning time, lower costs, and simplify the process for clinicians, enhancing the feasibility of this approach in diverse clinical settings. These results provide valuable insights to guide the development of future studies and clinical applications.

The Grad-CAM visualization results indicate that muscle regions play a significant role in predicting body composition in preterm infants. FFM heatmaps aligned with clinical expectations, demonstrating that non-fat, predominantly muscle tissue, is most important for predicting FFM. In contrast, FM heatmaps were more complex; the model does not focus primarily on subcutaneous fat for its prediction, but emphasizes surrounding tissue regions. This suggests that the prediction of FM may involve other fat deposits, such as visceral fat near the organs, which might not be easily detected by surface ultrasound images. This is consistent with existing literature, which shows that preterm infants have greater visceral fat depots at term-corrected age compared to full-term infants.31 Furthermore, the limited subcutaneous fat in preterm infants could explain why subcutaneous fat did not emerge as the primary determinant for FM prediction. Additionally, we found that subcutaneous fat over the biceps and quadriceps is a stronger predictor of total body composition than abdominal subcutaneous fat, as indicated by the warmer colors in these regions. This suggests that the model’s emphasis on these areas reflects their greater predictive value, rather than a limitation in ultrasound’s ability to differentiate between tissue layers. These results highlight the importance of considering anatomical context when interpreting ultrasound images for body composition analysis.

This study provides a foundation for future multi-center, longitudinal studies to improve generalizability with more diverse datasets. To address the limited diversity of infants in the present cohort, we have initiated data collection at additional sites, including high- and low-resource settings. Nevertheless, the proposed deep learning pipeline shows strong potential for assessing infant nutrition with portable imaging systems outside traditional settings. Compared to manual methods, it offers a more objective, efficient, and less operator-dependent approach. Ultrasound can also be applied in critically ill infants or those on respiratory support who cannot undergo ADP. With mobile integration, this tool could be deployed in underserved settings for immediate feedback and point-of-care monitoring. A next step toward clinical translation will be incorporating poor-quality images into training, enabling the model to automatically identify and exclude them and reduce the need for manual quality control.

Conclusion

This study introduces an ultrasound-based deep learning approach for automated body composition prediction, showing good agreement with reference measures. With further refinement, this approach may provide a cost-effective tool for assessing infant growth and nutrition across diverse care settings.

Acknowledgements

We express our gratitude to our funding sources: Google Research Scholar Program (Ranger), the Schiller Institute for Integrated Science and Society Grant for Exploratory Collaborative Scholarship (SI-GECS; Ranger, Park, Wei), the Research Expense Grant (REG; Ranger, Park), the Boston College Undergraduate Research Fellowship (URF) Program (Li, Cho, Hohenberg), the Gerber Foundation (Nagel, Ramel), the University of Minnesota Healthy Foods Healthy Lives Institute (Nagel, Ramel), and the National Institute of Child Health and Human Development (K23HD104000) (Bell). We also thank other members of our broader research team for their insights and support on this work, including Marisa Albert, Ji In Kim, Dr. Judy Estroff, Dr. Tsinuel Girma, and Dr. Melkamu Berhane.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Conflict of Interest Statement

The authors declare no competing interests.

Data Availability Statement

The data used during the study are available from the corresponding author on reasonable request.

References

  • 1.Katoch OR. Determinants of malnutrition among children: A systematic review. Nutrition 2022; 96: 111565. [DOI] [PubMed] [Google Scholar]
  • 2.Saavedra JM, Prentice AM. Nutrition in school-age children: a rationale for revisiting priorities. Nutrition Reviews. 2023. Jul 1;81(7):823–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Haire-Joshu D, Tabak R. Preventing obesity across generations: evidence for early life intervention. Annual Review of Public Health. 2016. Mar 18;37(1):253–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ramel SE, Zhang L, Misra S, Anderson CG, Demerath EW. Do anthropometric measures accurately reflect body composition in preterm infants?. Pediatric Obesity. 2017. Aug;12:72–7. [DOI] [PubMed] [Google Scholar]
  • 5.Bell KA, Wagner CL, Perng W, Feldman HA, Shypailo RJ, Belfort MB. Validity of body mass index as a measure of adiposity in infancy. The Journal of Pediatrics. 2018. May 1;196:168–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ranger BJ, Lombardi A, Kwon S, Loeb M, Cho H, He K, Wei D, Park J. Ultrasound for assessing paediatric body composition and nutritional status: Scoping review and future directions. Acta Paediatrica. 2025. Jan;114(1):14–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Mateos-Angulo A, Galán-Mercant A, Cuesta-Vargas AI. Ultrasound muscle assessment and nutritional status in institutionalized older adults: a pilot study. Nutrients. 2019: 11(6): 1247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Gehad MH, Yousif YM, Metwally MI, AbdAllah AM, Elhawy LL, El-Shal AS, Abdellatif GM. Utility of muscle ultrasound in nutritional assessment of children with nephrotic syndrome. Pediatr. Nephrol 2023: 38(6): 1821–1829. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.López-Gómez JJ, García-Beneitez D, Jiménez-Sahagún R, Izaola-Jauregui O, Primo-Martín D, Ramos-Bachiller B, Gómez-Hoyos E, Delgado-García E, Pérez-López P, De Luis-Román DA. Nutritional ultrasonography, a method to evaluate muscle mass and quality in morphofunctional assessment of disease related malnutrition. Nutrients. 2023: 15(18): 3923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.McLeod G, Geddes D, Nathan E, Sherriff J, Simmer K, Hartmann P. Feasibility of using ultrasound to measure preterm body composition and to assess macronutrient influences on tissue accretion rates. Early Human Development. 2013. Aug 1;89(8):577–82. [DOI] [PubMed] [Google Scholar]
  • 11.Nagel E, Hickey M, Teigen L, Kuchnia A, Holm T, Earthman C, Demerath E, Ramel S. Can ultrasound measures of muscle and adipose tissue thickness predict body composition of premature infants in the neonatal intensive care unit?. Journal of Parenteral and Enteral Nutrition. 2021. Feb;45(2):323–30. [DOI] [PubMed] [Google Scholar]
  • 12.LeCun Y, Bengio Y, Hinton G. Deep learning. nature. 2015. May 28;521(7553):436–44. [DOI] [PubMed] [Google Scholar]
  • 13.Liu S, Wang Y, Yang X, Lei B, Liu L, Li SX, Ni D, Wang T. Deep learning in medical ultrasound analysis: a review. Engineering. 2019. Apr 1;5(2):261–75. [Google Scholar]
  • 14.Weston AD, Korfiatis P, Kline TL, Philbrick KA, Kostandy P, Sakinis T, Sugimoto M, Takahashi N, Erickson BJ. Automated abdominal segmentation of CT scans for body composition analysis using deep learning. Radiology. 2019. Mar;290(3):669–79. [DOI] [PubMed] [Google Scholar]
  • 15.Koitka S, Kroll L, Malamutmann E, Oezcelik A, Nensa F. Fully automated body composition analysis in routine CT imaging using 3D semantic segmentation convolutional neural networks. European Radiology. 2021. Apr;31:1795–804. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Magudia K, Bridge CP, Bay CP, Babic A, Fintelmann FJ, Troschel FM, Miskin N, Wrobel WC, Brais LK, Andriole KP, Wolpin BM. Population-scale CT-based body composition analysis of a large outpatient population using deep learning to derive age-, sex-, and race-specific reference curves. Radiology. 2021. Feb;298(2):319–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Xu K, Li T, Khan MS, Gao R, Antic SL, Huo Y, Sandler KL, Maldonado F, Landman BA. Body composition assessment with limited field-of-view computed tomography: A semantic image extension perspective. Medical Image Analysis. 2023. Aug 1;88:102852. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Casey P, Alasmar M, McLaughlin J, Ang Y, McPhee J, Heire P, Sultan J. The current use of ultrasound to measure skeletal muscle and its ability to predict clinical outcomes: a systematic review. Journal of cachexia, sarcopenia and muscle. 2022. Oct;13(5):2298–309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ranger BJ, Bradburn E, Chen Q, Kim M, Noble JA, Papageorghiou AT. Portable ultrasound devices for obstetric care in resource-constrained environments: mapping the landscape. Gates Open Research. 2023. Dec 6;7:133. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chlap P, Min H, Vandenberg N, Dowling J, Holloway L, Haworth A. A review of medical image data augmentation techniques for deep learning applications. Journal of medical imaging and radiation oncology. 2021. Aug;65(5):545–63. [DOI] [PubMed] [Google Scholar]
  • 21.Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 2015. (pp. 234–241). Springer International Publishing. [Google Scholar]
  • 22.Tan M, and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. International conference on machine learning. PMLR, 2019. [Google Scholar]
  • 23.He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition 2016. (pp. 770–778). [Google Scholar]
  • 24.Oktay O, Schlemper J, Folgoc LL, Lee M, Heinrich M, Misawa K, Mori K, McDonagh S, Hammerla NY, Kainz B, Glocker B. Attention u-net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999. 2018. Apr 11. [Google Scholar]
  • 25.Kingma DP, Ba J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. 2014. Dec 22. [Google Scholar]
  • 26.Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A. Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016. (pp. 2921–2929). [Google Scholar]
  • 27.Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision 2017. (pp. 618–626). [Google Scholar]
  • 28.Ellis KJ. Body Composition Assessment in Early Infancy: A Review. Baylor Coll Med USDA/ARS Child Nutition Research Cent Houst. 2002. Nov 18. [Google Scholar]
  • 29.Ellis KJ, Yao M, Shypailo RJ, Urlando A, Wong WW, Heird WC. Body-composition assessment in infancy: air-displacement plethysmography compared with a reference 4-compartment model. The American Journal of Clinical Nutrition. 2007. Jan 1;85(1):90–5. [DOI] [PubMed] [Google Scholar]
  • 30.Wood AJ, Raynes-Greenow CH, Carberry AE, Jeffery HE. Neonatal length inaccuracies in clinical practice and related percentile discrepancies detected by a simple length-board. Journal of paediatrics and child health. 2013. Mar;49(3):199–203. [DOI] [PubMed] [Google Scholar]
  • 31.Casirati A, Somaschini A, Perrone M, Vandoni G, Sebastiani F, Montagna E, Somaschini M, Caccialanza R. Preterm birth and metabolic implications on later life: A narrative review focused on body composition. Frontiers in Nutrition. 2022. Sep 15;9:978271. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data used during the study are available from the corresponding author on reasonable request.

RESOURCES