Abstract
The image quality in clinical PET scan can be severely degraded due to high noise levels in extremely obese patients. Our work aimed to reduce the noise in clinical PET images of extremely obese subjects to the noise level of lean subject images, to ensure consistent imaging quality. The noise level was measured by normalized standard deviation (NSTD) derived from a liver region of interest. A deep learning-based noise reduction method with a fully 3D patch-based U-Net was used. Two U-Nets, U-Nets A and B, were trained on datasets with 40% and 10% count levels derived from 100 lean subjects, respectively. The clinical PET images of 10 extremely obese subjects were denoised using the two U-Nets. The results showed the noise levels of the images with 40% counts of lean subjects were consistent with those of the extremely obese subjects. U-Net A effectively reduced the noise in the images of the extremely obese patients while preserving the fine structures. The liver NSTD improved from 0.13±0.04 to 0.08±0.03 after noise reduction (p = 0.01). After denoising, the image noise level of extremely obese subjects was similar to that of lean subjects, in terms of liver NSTD (0.08±0.03 vs. 0.08±0.02, p = 0.74). In contrast, U-Net B over-smoothed the images of extremely obese patients, resulting in blurred fine structures. In a pilot reader study comparing extremely obese patients without and with U-Net A, the difference was not significant. In conclusion, the U-Net trained by datasets from lean subjects with matched count level can provide promising denoising performance for extremely obese subjects while maintaining image resolution, though further clinical evaluation is needed.
Index Terms—: noise reduction, extremely obese patient, FDG PET, deep learning
I. Introduction
IMAGE quality of clinical PET-CT scans is frequently degraded due to large body habitus, especially for patients with high body mass index (BMI) [1]. The large body habitus causes substantial photon attenuation and greater scatter fraction, resulting in significant image noise that can impair the visual detection of lesions, and may render SUV measurements inaccurate [2, 3]. Our retrospective analysis of 2,916 PET-CT studies performed at Yale-New Haven Hospital revealed high prevalence of obesity (39.1%), similar to widespread national data [4, 5]: Class 1 and 2 obesities with BMI between 30 and 40 kg/m2 (987 subjects, 33.8%), Class 3 Extreme Obesity with BMI > 40 kg/m2 (153 subjects, 5.2%). PET-CT images of patients with extreme obesity demonstrated much higher noise levels than scans on patients with normal BMI, leading to difficulties in interpretation, and potentially higher risks of misdiagnosis [6–10]. Deep-learning methods have been well established for reducing the noise and recovering the image quality for low-dose PET images [11–18]. However, the performance of similar noise reduction methods for the clinical PET images of extremely obese subjects were largely unknown and the optimal noise level of PET images to be used in network training was unexplored.
The goal of this work was to recover image quality by reducing the noise level of clinical PET-CT images on extremely obese subjects (BMI > 40 kg/m2) to a noise level similar to that of lean subjects by using a deep learning-based method. The noise levels were gauged by normalized standard deviation (NSTD) derived from liver regions of interest (ROIs) [19]. A 3D patch-based U-Net was used for noise reduction. The training dataset, includes the low- and full- count PET images from 100 randomly selected subjects with normal BMI. The low-count PET images were generated with 40% counts to match the noise levels that are typical in extremely obese subjects in terms of liver NSTD.
II. Materials and methods
A. Study datasets
The subjects were put into three groups as shown in Table 1. Groups 1 and 2 are not extremely obese patients (BMI < 40 kg/m2), and Group 3 are extremely obese patients (BMI > 40 kg/m2). All datasets of the three groups were acquired on a Siemens Biograph mCT at Yale-New Haven Hospital. The subjects were injected with 18F-FDG and a whole-body protocol with continue-bed-motion scanning was used. The images were reconstructed using ordered-subsets expectation maximization (OSEM) algorithm with 2 iterations and 21 subsets, provided by the vendor. A post-reconstruction Gaussian filter with 5 mm full width at half maximum (FWHM) was used. The voxel size of the reconstructed image was 4.07 × 4.07 × 3 mm3. The image size was 200 × 200 in the transverse plane and varied in the axial direction depending on patient height. For the dataset of Group 2, we obtained the low-count PET images for deep convolution neural network (CNN) training after independent uniform down-sampling of the patient list-mode data with Siemens E7 tool. The down-sampling ratio was controlled so that the noise level of the low-count images used in training matched that of the clinical PET images of the extremely obese subjects in Group 3 using liver NSTD as described below.
TABLE I.
The population characteristic of the subjects
| Group 1 | Group 2 | Group 3 | |
|---|---|---|---|
| Description | Not extremely obese | Not extremely obese | Extreme obesity |
| For count level estimation | For training | For testing | |
| Subject Number | 10 | 100 | 10 |
| BMI (kg/m2) | 25.6±3.8 | 26.8±5.2 | 44.9±2.7 |
| Height (m) | 1.70±0.10 | 1.66±0.10 | 1.73±0.09 |
| Weight (kg) | 74.4±15.8 | 73.8±18.3 | 134.0±15.6 |
| Age (years) | 61.0±21.7 | 58.5±22.4 | 64.0±10.9 |
| Gender (male) | 7 | 42 | 4 |
| Dose (mCi) | 9.97±0.81 | 9.95±0.83 | 10.32±0.86 |
| Scan duration (min) | 15.1±2.5 | 16.4±3.9 | 17.1±3.1 |
| Timing post injection (min) | 70.9±11.3 | 69.8±10.0 | 74.1±13.8 |
B. Noise level estimation for extremely obese patients
We used the NSTD inside a ROI with 10 × 10 × 5 voxels in relatively unform liver parenchyma as the surrogate to measure the image-based noise level. The NSTD for the extremely obese subjects in Group 3 was calculated to verify the target noise level of the down-sampled low-count PET images used in training. We down-sampled the listmode data of the 10 patient PET dataset in Group 1 with different down-sampling ratios, ranging from 10% to 100% with an 10% increment. The NSTDs for the PET images obtained from those 10 different count levels were measured. Then the down-sampling ratio with NSTD matching that of the extremely obese patients was selected to guide the generation of the low-count PET images of Group 2 for network training.
C. Network architecture and training
A fully 3D patch-based U-Net [20], including three parts of layers: contracting path, bottleneck and expanding path, as shown in Fig. 1, was used for noise reduction in this study. In the contracting path, each contracting layer consists of two convolution operations, each of which is a 3 × 3 × 3 convolution followed by a rectified linear unit (ReLU) and a 2 × 2 × 2 max pooling operation. The bottleneck includes two convolution operations. In the expanding path, each expanding layer includes a 2 × 2 × 2 up-convolution operation and two convolution operations. The feature maps extracted from the expanding path were concatenated with copied feature maps from the contracting path. The feature number in the first layer is 64.
Fig. 1.

The U-Net architecture.
The network was trained using the pairs of low-count PET images (the input) and full-count PET images (the label) from the 100 subjects in Group 2. The count level of the low-count PET images was determined according to Section B above. As shown in result section below, the noise level of the images with 40% counts matches with that of the extremely obese dataset in Group 3. During training, the patch size was 32 × 32 × 32. Each patch was randomly chosen from the training images. L2 loss function was used and Adam optimizer was applied with an initial learning rate as 0.0001 and exponential decay rate as 0.999. Each epoch contained 16,000 batches with size of 64 while 2,400 epochs were used. For testing process, a patch size of 256 × 256 × 128 was implemented to avoid the patching artifact. All the experiments were carried out on a Dell workstation with NVIDIA Titan Xp GPU.
D. Quantitative Evaluation
The clinical PET images of the 10 extremely obese subjects in Group 3 were denoised using the above-trained network (U-Net A), which was trained by the dataset with consistent count level. For comparison, we obtained the denoised images for these subjects using another network (U-Net B) trained by a dataset with inconsistent count level, consisting of the pairs of 10%-count image and full-count image. The noise of the denoised images was quantitatively calculated using the NSTD inside the liver ROI. The paired t-test was used to compare the noise between the denoised images and the full-count images for the extremely obese subjects. The un-paired t-test was used to compare the noise between the denoised images of the extremely obese subjects and the full-count images of the lean subjects in Group 1.
To quantitively evaluate the resolution preservation, we manually drawn the lesion and lung ROIs to estimate the lesion contrast, which was the ratio of the lesion SUV to lung SUV. The boundary of lesion ROIs were determined by 10% of the maximum value inside the lesion while the lung ROIs were large but away from the lung boundary to reduce the partial volume effect. A total of 10 lesions were obtained for the evaluation.
E. Reader Study Evaluation
To evaluate the denoising performance in a more clinically relevant setting, 49 faculty nuclear radiologists from multiple academic hospitals throughout the USA were invited via email to participate in an online survey (https://www.radiology-universe.org/PET-CT-algorithm-research/), and were also asked to invite other faculty and trainees from their institutions. This online survey presented multiple image slices from various patients; one side displayed images reconstructed using the conventional technique, while the other side displayed corresponding slice denoised using the deep-learning method using U-Net A. Axial, coronal, and sagittal images were displayed. Identity of the algorithms was concealed, labelling these as A and B. Radiologists indicated their overall preference for an algorithm, and selected their training level as either an Attending, Fellow, or Resident. A sample survey screen is shown in Fig 2. A total of 26 respondents completed the survey: 12 Attendings, 2 fellows, and 12 residents. Proportions of respondents preferring one algorithm vs. the other were computed. P value was obtained using Fischer’s Exact test. Binomial confidence intervals were computed using R software.
Fig. 2.

Sample screen capture of the online survey.
III. Results
A. Count level estimation
The NSTD for the extremely obese subjects was 0.13±0.04 and Fig. 3 is the plot showing NSTD with different down-sampling ratios (percentage count) for Group 1 (lean subjects). Apparently, the NSTD increased while the down-sampling ratios decreased. From the results, the NSTD of 40% count PET images of lean subjects is the closest to that of the extremely obese subjects without significant difference (p = 0.62). So, in this work, we used 40% dose PET images as the input images to train the U-Net A.
Fig. 3.

The NSTD of the image with different counts for the lean patients of Group 1. The red line is the desired NSTD range derived from the extremely obese patients of Group 3.
B. Denoised images
As shown in Fig. 4, the U-Net A trained by the dataset with matched count level can effectively reduce the noise for the extremely obese subjects while largely maintaining image resolution, even for the subjects with an extremely large BMI (Subject #3). In contrast, although the U-Net B trained by the dataset with inconsistent count level reduces the noise, the denoised images by U-Net B are over-smoothed with losing fine structure details indicated by while solid arrows in Fig. 4.
Fig. 4.

Sample slices of the raw image, denoised images with U-Nets A and B from two extremely obese subjects. The top row, subject #1 with 46.6 kg/m2 BMI, the middle row, subject #2 with 45.7 kg/m2 BMI, and the bottom row, subject #3 with 49.2 kg/m2 BMI. The dashed arrows indicate the line profile shown in Fig. 5. The display colorbar for the zoomed-in part is from 0 to 10.
The line profiles along the three arrows highlight the fine structures (Fig. 5) where U-Net A is more consistent than U-Net B with the raw images. Specifically, the jawbone after denoising using U-Net A is much sharper than that with U-Net B. The nodule peaks of the U-Net A are also much higher in the intensity and consistent with the raw images than that of the U-Net B.
Fig. 5.

The profiles of the three subjects, pointed out by the arrows.
C. Quantitative measurement
As shown in Fig. 6, the NSTDs of the denoised images with both U-Net A and U-Net B are significantly decreased compared to the raw images (0.08±0.03, 0.05±0.02 vs. 0.13±0.04, for U-Net A, U-Net B and raw images, respectively). Furthermore, the NSTDs of the raw images of the extremely obese subjects in Group 3 are significantly different from those of the full dose images of the subjects in Group 1 (0.13±0.04 vs. 0.08±0.02, p=0.001). After denoising using the deep learning method, there is no significant difference between the NSTDs from U-Net A and those from Group 1 (0.08±0.03 vs. 0.08±0.02, p =0.74) whereas there is a significant difference between the denoised images using U-Net B and the full dose images in Group 1 (0.05±0.02 vs. 0.08±0.02, p=0.002).
Fig. 6.

The NSTD of the raw image, denoised images with U-Net A and B, for the ten extremely obese subjects in Group 3.
As shown in Fig. 7, the contrasts of the denoised images with U-Net B are decreased compared to the raw images while those with U-Net A are close to the raw images (2.18±1.55, 2.18±1.57, and 1.99±1.33 for U-Net A, U-Net B and raw images, respectively).
Fig. 7.

The contrast of the raw image, denoised images with U-Net A and B, for the ten extremely obese subjects in Group 3.
D. Reader study
Radiologist’s preferences for a reconstruction algorithm varied, without a statistically-significant result. Within the attending group, 6 preferred the conventional method, 6 preferred the denoised method (Confidence interval, CI 21.1% to 78.9%). Within the trainees group, 4 preferred the denoised method, 10 preferred the conventional method (95% CI 8.4% to 58.1%). P-value between these groups was not significant, 0.4.
IV. Discussion
We have developed a deep convolutional neural network (U-Net) to reduce the noise of clinical PET images for extremely obese subjects to the noise level of lean subjects. The U-Net was trained by the count-level-controlled dataset, meaning the noise level of the input images in training was matched to that of the PET images of extremely obese subjects. The denoised images were compared with the images denoised by another network trained using inconsistent count level dataset. The results showed that the U-Net trained by the images with consistent count level is superior to that trained with inconsistent count level and can effectively reduce the noise for the extremely obese patients while preserving the fine structures without over-smoothing. After denoising, the noise level is similar to that of the clinical full dose images of lean subjects. On the other hand, the U-Net trained by the images with inconsistent count level over-smoothed the images for extremely obese patients and showed a loss in fine structure. In this study, based on our population, the 40% value was recommended for the matched count level. However, this value might not be applicable for all the other obese patients. To denoise the image for the obese patient in clinic, a more comprehensive study based on a large population will be needed. Furthermore, a fixed value and network may not be suitable and then a self-adaptive solution based on the NSTD of each subject could be used.
In the pilot reader study, 50% of the participants in the more experienced attending group preferred denoised images. While in the less experienced trainee group, only 4 out of 14 preferred denoise images, likely due to the fact that the trainees were trained with noisy images and are not used to images with lower noise. Nevertheless, further evaluations with comprehensive human observer study designs are needed to investigate the clinical benefit of deep learning-based PET denoising methods.
V. Conclusion
We developed a U-Net to reduce the noise of clinical PET images for extremely obese subjects to be similar to the noise level of lean subjects. To avoid the over-smoothing effect caused by denoising, we recommended using the dataset with consistent count level to train the U-Net.
Acknowledgment
All authors declare that they have no known conflicts of interest in terms of competing financial interests or personal relationships that could have an influence or are relevant to the work reported in this paper.
This work was supported by NIH grant R01EB025468.
Footnotes
This work involved human subjects or animals in its research. The authors confirm that all human/animal subject research procedures and protocols are exempt from review board approval.
Contributor Information
Hui Liu, Department of Engineering Physics, Tsinghua University, and Key Laboratory of Particle & Radiation Imaging, Ministry of Education (Tsinghua University), Beijing, China, on leave from the Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, 06511, USA..
Hamed Yousefi, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Niloufar Mirian, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
MingDe Lin, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA.; Visage Imaging, Inc., San Diego, CA, USA.
David Menard, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Matthew Gregory, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Mariam Aboian, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Annemarie Boustani, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Ming-Kai Chen, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Lawrence Saperstein, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Darko Pucar, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Michal Kulon, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
Chi Liu, Department of Radiology and Biomedical Imaging, Yale University, New Haven, CT, USA..
References
- [1].Pi-Sunyer X, “The medical risks of obesity,” Postgrad Med, vol. 121, no. 6, pp. 21–33, Nov 2009, doi: 10.3810/pgm.2009.11.2074. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Chang T, Chang G, Kohlmyer S, Clark JW, Rohren E, and Mawlawi OR, “Effects of injected dose, BMI and scanner type on NECR and image noise in PET imaging,” Phys Med Biol, vol. 56, no. 16, pp. 5275–85, Aug 21 2011, doi: 10.1088/0031-9155/56/16/013. [DOI] [PubMed] [Google Scholar]
- [3].Tatsumi M, Clark PA, Nakamoto Y, and Wahl RL, “Impact of body habitus on quantitative and qualitative image quality in whole-body FDG-PET,” Eur J Nucl Med Mol Imaging, vol. 30, no. 1, pp. 40–5, Jan 2003, doi: 10.1007/s00259-002-0980-5. [DOI] [PubMed] [Google Scholar]
- [4].Flegal KM, Kruszon-Moran D, Carroll MD, Fryar CD, and Ogden CL, “Trends in Obesity Among Adults in the United States, 2005 to 2014,” JAMA, vol. 315, no. 21, pp. 2284–91, Jun 7 2016, doi: 10.1001/jama.2016.6458. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Seidell JC and Halberstadt J, “Obesity: The obesity epidemic in the USA - no end in sight?,” Nat Rev Endocrinol, vol. 12, no. 9, pp. 499–500, Sep 2016, doi: 10.1038/nrendo.2016.121. [DOI] [PubMed] [Google Scholar]
- [6].Harnett DT et al. , “Clinical performance of Rb-82 myocardial perfusion PET and Tc-99m-based SPECT in patients with extreme obesity,” J Nucl Cardiol, vol. 26, no. 1, pp. 275–283, Feb 2019, doi: 10.1007/s12350-017-0855-6. [DOI] [PubMed] [Google Scholar]
- [7].Mc Ardle BA, Dowsley TF, deKemp RA, Wells GA, and Beanlands RS, “Does rubidium-82 PET have superior accuracy to SPECT perfusion imaging for the diagnosis of obstructive coronary disease?: A systematic review and meta-analysis,” J Am Coll Cardiol, vol. 60, no. 18, pp. 1828–37, Oct 30 2012, doi: 10.1016/j.jacc.2012.07.038. [DOI] [PubMed] [Google Scholar]
- [8].Botkin CD and Osman MM, “Prevalence, challenges, and solutions for 18F-FDG PET studies of obese patients: a technologist’s perspective,” J Nucl Med Technol, vol. 35, no. 2, pp. 80–3, Jun 2007, doi: 10.2967/jnmt.106.034918. [DOI] [PubMed] [Google Scholar]
- [9].Tatsumi M, Clark PA, Nakamoto Y, and Wahl RL, “Impact of body habitus on quantitative and qualitative image quality in whole-body FDG-PET,” Eur J Nucl Med Mol Imaging, vol. 30, no. 1, pp. 40–45, 2003/01/01 2003, doi: 10.1007/s00259-002-0980-5. [DOI] [PubMed] [Google Scholar]
- [10].El Fakhri G, Santos PA, Badawi RD, Holdsworth CH, Van Den Abbeele AD, and Kijewski MF, “Impact of acquisition geometry, image processing, and patient size on lesion detection in whole-body 18F-FDG PET,” J Nucl Med, vol. 48, no. 12, pp. 1951–1960, 2007. [DOI] [PubMed] [Google Scholar]
- [11].Liu H, Wu J, Lu W, Onofrey JA, Liu YH, and Liu C, “Noise reduction with cross-tracer and cross-protocol deep transfer learning for low-dose PET,” Phys Med Biol, vol. 65, no. 18, p. 185006, Sep 14 2020, doi: 10.1088/1361-6560/abae08. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Lu W et al. , “An investigation of quantitative accuracy for deep learning based denoising in oncological PET,” Eur J Nucl Med Mol Imaging, vol. 64, no. 16, p. 165019, 2019, doi: 10.1088/1361-6560/ab3242. [DOI] [PubMed] [Google Scholar]
- [13].Gong K, Guan J, Liu C-C, and Qi J, “Pet image denoising using a deep neural network through fine tuning,” IEEE Trans Radiat Plasma Med Sci, vol. 3, no. 2, pp. 153–161, 2018, doi: 10.1109/TRPMS.2018.2877644. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Xu J, Gong E, Pauly J, and Zaharchuk G, “200x low-dose PET reconstruction using deep learning,” arXiv preprint arXiv:1712.04119, 2017. [Google Scholar]
- [15].Xiang L et al. , “Deep auto-context convolutional neural networks for standard-dose PET image estimation from low-dose PET/MRI,” Neurocomputing, vol. 267, pp. 406–416, 2017, doi: 10.1016/j.neucom.2017.06.048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Kang J, Gao Y, Shi F, Lalush DS, Lin W, and Shen D, “Prediction of standard-dose brain PET image by using MRI and low-dose brain 18F-FDG PET images,” Med Phys, vol. 42, no. 9, pp. 5301–5309, 2015, doi: 10.1118/1.4928400. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Liu H, Viswanath V, Karp J, Liu C, and Surti S, “Investigation of lesion detectability using deep learning based denoising methods in oncology PET: a cross-center phantom study,” J Nucl Med, vol. 61, no. supplement 1, p. 430, May 1 2020. [Google Scholar]
- [18].Gong Y et al. , “Parameter-Transferred Wasserstein Generative Adversarial Network (PT-WGAN) for Low-Dose PET Image Denoising,” IEEE Trans Radiat Plasma Med Sci, vol. 5, no. 2, pp. 213–223, 2021, doi: 10.1109/trpms.2020.3025071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Turkington TG, Wilson JM, Bowsher JE, and Gilland DR, “Limiting iterations vs. post smoothing for noise control in PET,” in 2007 IEEE Nuclear Science Symposium Conference Record, 2007, vol. 4: IEEE, pp. 2772–2775. [Google Scholar]
- [20].Ronneberger O, Fischer P, and Brox T, “U-net: Convolutional networks for biomedical image segmentation,” presented at the International Conference on Medical image computing and computer-assisted intervention, 2015. [Google Scholar]
