Abstract
Objective.
Noise magnitude is one of image quality indicators in computed tomography (CT) for assessing imaging performance, and the CMS (Centers for Medical and Medicaid Service) recently included a measure of noise magnitude in terms of “global noise”. Despite its importance, no standard method currently exists to measure noise magnitude in patient images. Theoretically, the most accurate approach is to assess voxel value variation across repeated images, the so-called ensemble noise. Obviously, such a method is not ethically feasible in actual patients. To surmount this impasse, we deployed virtual imaging techniques to benchmark three noise magnitude calculation methods against two gold standard ensemble noise measures across 36 imaging conditions.
Methods.
Over 1800 virtual image datasets were generated from imaging an American College of Radiology (ACR) phantom and Extended Cardiac-Torso (XCAT) human models using a validated, scanner-specific CT simulator (DukeSim). The ACR phantom was imaged under 36 different imaging conditions defined by combinations of chest and abdominopelvic protocols, three dose levels, three reconstruction kernels, and both Filtered Back Projection and Iterative Reconstruction algorithms. At the same conditions, XCAT models were repeatedly imaged 50 times. Noise magnitudes in the ACR phantom were calculated in 5 circular ROIs. In patients, noise was measured in air surrounding the body and in soft tissues by applying HU<−900 and −300≤HU≤100 thresholds, respectively. Per each imaging condition, measured noise magnitudes were compared against ensemble noise in soft tissue, liver, and lungs.
Results.
Across all imaging conditions, noise measurements in the ACR phantom and in air surrounding the patient underestimated ensemble noise by approximately 60% and 50%, respectively. In contrast, soft tissue-based noise measurements were closer to the gold standard with median differences between −8% and +4%.
Conclusions.
This study introduced a virtual imaging-based framework to benchmark clinical CT noise metrics against ensemble noise measures. Virtual imaging enabled objective comparison of different noise magnitude calculation methods in large, realistic populations simulating clinical conditions. Noise measured in air cannot represent soft tissues noise. The results validated soft tissue-based noise measurements as a reliable surrogate to inform protocol design, technology assessment, and equitable healthcare reimbursement.
Keywords: CT image quality, CT noise magnitude methods, virtual imaging techniques
3. Introduction
In diagnostic X-ray imaging, the only reason why patients are exposed to minimal stochastic cancerogenic risks induced by ionizing radiation is to generate images from which valuable diagnostic information may be obtained (1). The quality of such information is qualitatively assessed by the diagnosticians when reading the images. Although a qualitative image evaluation can be timely at the point of care, it does not provide an objective evidence of the actual diagnostic suitability of the image. In computed tomography (CT), such an ephemeral evaluation of the image quality conflicts with the measurement of the radiation risks which can rely on a myriad of well-established metrics and methods, all with different degrees of appropriateness (2, 3). As a result, although the assessment of CT procedures and technologies, the justification of the exam, and the design of optimization actions should rely on an appropriate balance between radiation dose and image quality, they are largely dominated by the urge to reduce radiation burden, often overlooking the final procedure outcome evaluation in terms of the quality of the images (4).
To close this gap, and inspired by international guidelines from the International Commission on Radiological Protection (ICRP) and the International Atomic Energy Agency (IAEA) (5, 6), several research groups developed different methods to quantitatively measure image quality in CT images both in phantoms (in phantasma) and in patients (in vivo) (7–10). These image quality measurements ensure reproducibility and robustness and have been deployed to evaluate attributes such as noise magnitude, spatial resolution, and signal in terms of Hounsfield Unit (HU). In this milieu, noise magnitude measurements recently hit the headlines after a ruling of the Centers for Medical and Medicaid Service (CMS) proposing an ambiguous “global noise” figure as a metric of noise magnitude for diagnostic computed tomography exams in adults (11). The mandate implied that such a metric can be used as a reliable gauge of “inadequate image quality” as a basis of healthcare reimbursement.
Despite the implication of the CMS mandate and the ambiguity of its definition of noise, currently there exists no fully reliable, singular method to measure noise magnitude in CT images. Indeed, there has been a myriad of noise magnitude measures from those in single size uniform phantoms, in multi-sized phantoms (12), in patient images considering the air surrounding the patients (8), in patient images considering only soft tissue (7), and in individual organs of patients (13). Also, noise is not uniform across a CT dataset – aggregating a value across that diversity is a task of its own. This heterogeneity can lead to drastically different results, even when analyzing same datasets, biasing the evaluation of radiological procedures, protocol design, and healthcare reimbursement.
Therefore, it is necessary to objectively compare different approaches to estimate noise magnitude against a gold-standard method. The most accurate approach to measure noise magnitude is to ascertain the so-called ensemble noise (Nensemble) by scanning a patient multiple times, and sampling noise of each pixel across the ensemble of images. This approach requires repeated imaging, an ethically unfeasible prospect for patient imaging (8). Moreover, current quality assurance guidelines for CT table positioning reproducibility allow a tolerance of ±1 mm which may exceed the image pixel size, making it impossible to accurately sample the same pixel across repeated acquisitions (24). As a result, it is unknown which method provides the closest representation of noise magnitude to the ensemble noise gold standard.
To overcome this impasse, in this study, we estimated ensemble noise deploying Virtual Imaging (in silico) techniques and used it as a gold standard to compare three noise magnitude calculation methods, including measuring noise from phantom data, and from patient data in the air outside patients and in the soft tissue. The analysis determined to what extent different noise measurements can accurately represent the true noise magnitude in patient images.
4. Materials and methods
4.1. Image acquisition
A total of 1800 virtual patient and 18 phantom virtual CT image datasets, one for each imaging condition described below, were included in this study. The patient data was used to assess in silico noise measurement techniques from patient images whereas the phantom data was used to evaluate in phantasma noise measurements.
The patient cases included Chest and Abdominopelvic human models representing a typical adult with BMI at the 50th percentile of the US population (XCAT; male; 67 years old; BMI: 28.2 kg/m2, water equivalent diameter: 29.5 cm) generated by segmenting patient images using nonuniform rational B-spline as described by Segars et al. (14, 15); the phantom data were based on a computational ACR phantom (16). The virtual scans were done using a validated CT simulator (DukeSim) (14, 15, 17, 18). DukeSim encompasses a GPU-accelerated approach that combines ray-tracing and Monte Carlo methods to generate a final sinogram incorporating both primary and scatter signals. It accurately replicates various scanner-specific characteristics, including geometry, bowtie-filtered source spectrum, automatic tube current modulation, flying focal spot wobbling, anti-scatter grid, electronic noise, and poly-energetic detector response. Additionally, DukeSim has been integrated with manufacturer-specific reconstruction software and has been validated across various scanner models and imaging conditions (17, 19–22).
The virtual ACR phantom was imaged with DukeSim simulating a clinical CT scanner (Somatom Definition Flash, Siemens Healthineers, Erlangen, Germany) at 120 kV, pitch of 1, and three different dose levels (50, 100 and 150 mAs, corresponding to CTDIvol of 2.9, 5.7, and 8.6 mGy, respectively). The obtained datasets were reconstructed with a 5.0 mm slice thickness and a 500 mm field of view, using an offline reconstruction prototype (ReconCT 15.0.35098.0, Siemens Healthineers), with three different reconstruction kernels (Br32f, Br46f, and Br62f), with both weighted filtered back projection (FBP) and iterative reconstruction (IR) algorithms (ADMIRE, level 3 across all imaging conditions), to generate a total of 18 datasets. Previous studies showed that DukeSim can simulate CT images with image quality attributes similar to real CT images. In particular, noise magnitude was compared in ACR and Mercury Phantom (Duke University, Durham, NC) both in energy-integrating detector and photon-counting detector CT for different doses and reconstruction kernels (17, 23).
4.2. Noise measurement methods
Following the AAPM TG 233 methodology, a medical physicist with 4 years of experience measured the noise for each dataset from five ROIs, with size covering approximating 1% of the phantom area, placed at the center, and at 12, 3, 6, and 9 o’clock of each slice over the uniform section of the imaged phantom, as shown in Figure 1 (24). Noise magnitude (NACR) was obtained by averaging the standard deviation of pixel values across each ROI location and for 19 slices. NACR standard deviation and coefficient of variation were analogously calculated across the 19 slices.
Figure 1:

ROI placements (in red) for noise magnitude measurement on the virtual ACR phantom.
The two virtual patients were also imaged using DukeSim with the same conditions used to acquire the ACR phantom images. Each scan included 39 image slices and was repeated 50 times to enable stable ensemble noise estimations. The sinogram images were then reconstructed with a 5.0 mm slice thickness, a matrix size of 1024 × 1024 pixels, and a 500 mm field of view using ReconCT. Each sinogram data was reconstructed six times, utilizing three different kernels (Br32f, Br46f, and Br62f) and both FBP and IR algorithms, generating a total of 1800 studies (Figure 2.a). Although FBP and IR exhibit different relationships between the input projection data and the reconstructed images, both were included to reflect clinically relevant reconstruction conditions.
Figure 2.

Abdominopelvic (top) and Chest (bottom) image reconstructed with FBP (a); with GNI threshold in red (b); and with NAIR threshold in red (c).
Using a previously published automated methods (7), for each imaging condition and in each repeated scan, noise magnitudes were calculated in soft tissues as global noise index (GNI), and in the air surrounding the patient (NAIR) by applying thresholds of −300≤HU≤100 and HU<−900, respectively (Figure 2.b–c). In particular, for each slice, an ROI (30 pixels × 30 pixels) is identified around each pixel, generating a standard deviation and a histogram of standard deviation map across the slice. Noise of that slice is computed as the value corresponding to the peak of the histogram. Finally, the noise values for the study were computed as the mean of the values calculated in each slice. Standard deviation and coefficient of variation across the 50 repetitions were also calculated for each noise magnitude metric and imaging condition.
4.3. Enslemble noise measurement
The ensemble noise was calculated in the same GNI soft tissue areas (Nensemble), and in the liver and in the lungs (Norgan) for abdominopelvic and chest studies, respectively (Figure 3). In particular, for each slice and for every pixel within the segmented area, the HU of that pixel was measured across the 50 repeated images. The standard deviation of these 50 measurements defined the ensemble noise of that pixel. The slice ensemble noise was obtained by averaging the pixel-level ensemble noise across all the pixels in the slice. Lastly, the final ensemble noise (Nensemble and Norgan) was then calculated by averaging the slice ensemble noise across all the slices. Because the HU variation for a given pixel across the 50 repetitions cannot be attributed to anatomical or morphological changes in the patient or phantom, this variation reflects only photon statistics and electronic noise within the simulation framework. Therefore, Nensemble and Norgan can be regarded as gold standard for quantifying noise magnitude. Standard deviation and coefficient of variation values were also calculated across the 50 repetitions for each noise magnitude metric and imaging condition.
Figure 3.

Human model (XCAT) liver (left) and lungs (right) organ masks applied to calculate Norgan.
The analysis was performed separately for each set of dose level, kernel, and reconstruction algorithm. Median noise magnitude values from different methods were compared in terms of percentage difference. NACR, NAIR and GNI were compared to the ensemble noise calculated in soft tissue (Nensemble) and to the ensemble noise calculated in the liver and in the lungs (Norgan) for abdominopelvic and chest studies. Lastly, for each metric, the relative deviation from Nensemble and Norgan was calculated across variations in mAs, reconstruction algorithm, and kernel, and the maximum absolute percentage error with respect to Nensemble and Norgan was reported across all parameter combinations.
5. Results
Table 1 and Table 2 summarize the median NACR, NAIR, GNI, Nensemble, and Norgan for chest and abdominopelvic images per each kernel, mAs value, and reconstruction algorithm involved in the study. As expected, the noise values decreased with the increase in radiation dose and increased with the kernel sharpness. Moreover, because the ACR phantom is smaller than the virtual patients, the noise magnitude in the phantom is consistently lower than the values measured in patients for constant tube current.
Table 1.
Median NACR, NAIR, GNI, Nensemble, and Norgan for chest images per each kernel, mAs value, and reconstruction algorithm involved in the study. Units are HU.
| Chest | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | NAIR | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 4.9 | 4.1 | 7.4 | 7.7 | 7.0 | 3.2 | 3.2 | 6.0 | 6.7 | 6.7 | |
| 100 | 3.5 | ||||||||||
| Br32f | mAs | 3.0 | 5.4 | 5.4 | 5.0 | 2.3 | 2.4 | 4.4 | 4.7 | 4.7 | |
| 150 | 2.9 | ||||||||||
| mAs | 2.5 | 4.5 | 4.4 | 4.0 | 1.9 | 2.1 | 3.7 | 3.9 | 3.9 | ||
| 50 mAs | 13.6 | 14.3 | 21.5 | 21.9 | 19.6 | 8.4 | 9.5 | 14.4 | 16.4 | 15.6 | |
| 100 | 9.6 | ||||||||||
| Br46f | mAs | 10.0 | 15.2 | 15.4 | 13.8 | 5.9 | 6.7 | 10.2 | 11.5 | 11.1 | |
| 150 | 7.8 | ||||||||||
| mAs | 8.1 | 12.5 | 12.6 | 11.6 | 4.7 | 5.5 | 8.4 | 9.4 | 9.2 | ||
| 50 mAs | 45.4 | 44.0 | 102.8 | 98.0 | 86.4 | 24.6 | 30.2 | 57.7 | 57.5 | 52.6 | |
| Br62f | 100 | 31.8 | |||||||||
| mAs | 35.9 | 74.4 | 69.1 | 61.3 | 17.1 | 25.4 | 41.0 | 40.7 | 37.2 | ||
| 150 | 25.8 | ||||||||||
| mAs | 32.0 | 61.6 | 56.4 | 50.1 | 13.7 | 22.8 | 34.0 | 33.2 | 30.5 | ||
Table 2.
Median NACR, NAIR, GNI, Nensemble, and Norgan for abdominopelvic images per each kernel, mAs value, and reconstruction algorithm involved in the study. Units are HU.
| Abdomen | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | Nair | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 4.9 | 4.8 | 10.8 | 10.8 | 12.1 | 3.2 | 3.8 | 8.9 | 9.5 | 10.3 | |
| 100 | |||||||||||
| Br32f | mAs | 3.5 | 3.4 | 7.8 | 7.6 | 8.5 | 2.3 | 2.8 | 6.3 | 6.7 | 7.2 |
| 150 | |||||||||||
| mAs | 2.9 | 2.9 | 6.4 | 6.2 | 7.0 | 1.9 | 2.3 | 5.2 | 5.4 | 5.8 | |
| 50 mAs | 13.6 | 17.1 | 32.2 | 30.1 | 33.9 | 8.4 | 11.5 | 22.0 | 22.5 | 24.2 | |
| 100 | |||||||||||
| Br46f | mAs | 9.6 | 12.0 | 22.6 | 21.3 | 23.8 | 5.9 | 7.9 | 15.4 | 15.9 | 16.9 |
| 150 | |||||||||||
| mAs | 7.8 | 9.6 | 18.5 | 17.4 | 19.4 | 4.7 | 6.5 | 12.5 | 13.0 | 13.7 | |
| 50 mAs | 45.4 | 51.5 | 143.5 | 131.4 | 149.3 | 24.6 | 33.8 | 84.5 | 77.7 | 86.6 | |
| 100 | |||||||||||
| Br62f | mAs | 31.8 | 40.2 | 103.6 | 92.5 | 105.0 | 17.1 | 27.4 | 60.0 | 54.9 | 60.8 |
| 150 | |||||||||||
| mAs | 25.8 | 35.8 | 86.1 | 75.4 | 85.5 | 13.7 | 24.8 | 48.9 | 44.8 | 49.4 | |
NACR and NAIR both underestimated the gold standards Nensemble and Norgan across all scanning conditions for both chest and abdominopelvic exams (Figure 4 and 5). The difference between NACR and Nensemble was −54% and −61% for FBP and IR studies, respectively; and the difference between NACR and Norgan was −54% and −60% for FBP and IR. The difference between NAIR and Nensemble was −46% and −49% for FBP and IR studies; the difference between NAIR and Norgan was −49% and −51% for FBP and IR, respectively.
Figure 4.

Boxplots of the relative percentage differences between NACR, NAIR, GNI and Nensemble, and Norgan for FBP and IR studies considering all kernels and mAs values. The red line represents the median value, and the blue boxes represent the 1st and 3rd quartiles.
Figure 5.

Heatmap of the relative percentage differences across mA values between NACR, NAIR, GNI and Nensemble, and Norgan for each protocol, kernel, and reconstruction algorithm.
In contrast, Global noise index (GNI) was closer to gold standard Nensemble: +4% in FBP and −3% in IR images. Moreover, GNI values showed a difference of +3% and −8% when compared with Norgan measured in FBP and IR images, respectively. Tables 3 and 4 report standard deviation and coefficient of variation for NACR, NAIR, GNI, Nensemble, and Norgan for chest and abdominopelvic images per each kernel, mAs value, and reconstruction algorithm. Differences in standard deviation and coefficient of variation magnitude reflect the different methods used to calculate the associated noise values.
Table 3a.
Standard deviation values for NACR, NAIR, GNI, Nensemble, Norgan for chest images per each kernel, mAs value, and reconstruction algorithm involved in the study. Units are HU.
| Chest | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | NAIR | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 0.141 | 0.009 | 0.026 | 1.27 | 2.09 | 0.092 | 0.006 | 0.016 | 1.14 | 1.73 | |
| 100 | 0.095 | ||||||||||
| Br32f | mAs | 0.006 | 0.015 | 0.896 | 1.45 | 0.063 | 0.005 | 0.026 | 0.801 | 1.20 | |
| 150 | 0.060 | ||||||||||
| mAs | 0.004 | 0.007 | 0.732 | 1.18 | 0.041 | 0.005 | 0.015 | 0.655 | 0.970 | ||
| 50 mAs | 0.324 | 0.043 | 0.058 | 3.51 | 5.46 | 0.178 | 0.033 | 0.042 | 2.76 | 4.23 | |
| Br46f | 100 | 0.208 | |||||||||
| mAs | 0.024 | 0.042 | 2.49 | 3.80 | 0.140 | 0.012 | 0.033 | 1.95 | 2.90 | ||
| 150 | 0.157 | ||||||||||
| mAs | 0.027 | 0.037 | 2.06 | 3.15 | 0.083 | 0.010 | 0.023 | 1.60 | 2.34 | ||
| 50 mAs | 1.336 | 0.086 | 0.240 | 14.3 | 24.3 | 0.708 | 0.062 | 0.158 | 8.80 | 14.9 | |
| 100 | 0.812 | ||||||||||
| Br62f | mAs | 0.076 | 0.243 | 10.0 | 16.9 | 0.437 | 0.058 | 0.110 | 6.20 | 10.4 | |
| 150 | 0.615 | ||||||||||
| mAs | 0.069 | 0.185 | 8.19 | 13.7 | 0.275 | 0.033 | 0.102 | 5.08 | 8.49 | ||
Table 4a.
Coefficient of variation values for NACR, NAIR, GNI, Nensemble, Norgan for chest images per each kernel, mAs value, and reconstruction algorithm involved in the study.
| Chest | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | NAIR | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 0.029 | 0.002 | 0.004 | 0.156 | 0.267 | 0.029 | 0.002 | 0.003 | 0.160 | 0.236 | |
| Br32f | 100 mAs | 0.027 | 0.002 | 0.003 | 0.156 | 0.263 | 0.028 | 0.002 | 0.006 | 0.160 | 0.231 |
| 150 mAs | 0.021 | 0.002 | 0.002 | 0.156 | 0.261 | 0.022 | 0.003 | 0.004 | 0.160 | 0.229 | |
| Br46f | 50 mAs | 0.024 | 0.003 | 0.003 | 0.152 | 0.251 | 0.021 | 0.003 | 0.003 | 0.159 | 0.246 |
| 100 mAs | 0.022 | 0.002 | 0.003 | 0.152 | 0.248 | 0.024 | 0.002 | 0.003 | 0.160 | 0.237 | |
| 150 mAs | 0.020 | 0.003 | 0.003 | 0.154 | 0.249 | 0.018 | 0.002 | 0.003 | 0.161 | 0.233 | |
| 50 mAs | 0.029 | 0.002 | 0.002 | 0.140 | 0.254 | 0.029 | 0.002 | 0.003 | 0.146 | 0.258 | |
| Br62f | 100 mAs | 0.026 | 0.002 | 0.003 | 0.139 | 0.250 | 0.026 | 0.002 | 0.003 | 0.145 | 0.255 |
| 150 mAs | 0.024 | 0.002 | 0.003 | 0.139 | 0.249 | 0.020 | 0.001 | 0.003 | 0.145 | 0.254 | |
The maximum error compared to Nensemble across differentials due to mAs, reconstruction type, and kernel was 30% for NACR, and NAIR, 16% for GNI, and 5% for Norgan. The maximum error compared to Norgan across differentials due to mAs, reconstruction type, and kernel was 30% for NACR and NAIR, 13% for Nensemble, and 5% for Norgan.
6. Discussion
Through the application of virtual imaging techniques, we performed an objective and unbiased comparison of different noise magnitude estimation methods against the gold standard ensemble noise. The noise measured in the ACR phantom and in the air surrounding the patient largely underestimated the ensemble noise measured in soft tissues and in the organ of interest. Such methods can be instrumental in the comparison of different scanning techniques or in the evaluation of scanner performance, as there is a linear correlation between NAIR and GNI, but they were not introduced to represent noise magnitude in patient tissues, as confirmed in this study. Conversely, the global noise index (GNI) in soft tissue represented the closest surrogate to ensemble noise, affirming the validity of soft tissue-based noise measurements to inform protocol design and technology assessments (7).
Regarding the comparison between ACR phantom and patient-based measurements, it is important to consider differences in object attenuation. The ACR phantom has a diameter of 20 cm, whereas the water-equivalent diameter of the virtual patient model is approximately 29.5 cm. Since all simulations were performed under a fixed tube current, absolute noise values are not directly comparable between these two configurations without accounting for differences in attenuation. To this end, Jadick et al. showed strong agreement between real and images simulated with DukeSim in terms of noise magnitude (20). Moreover, they calculated noise magnitude in the 30 cm section of the Mercury Phantom, corresponding to a 29 cm water equivalent diameter, using DukeSim to simulate same scanner and dose levels considered in the present study (20). Their simulation employed a slice thickness of 0.6 mm and the Br40d reconstruction kernel. Under the simplifying assumption that noise magnitude scales inversely with the square root of slice thickness, and without accounting for potential effects of reconstruction kernel or inter-slice noise correlation, the noise values reported by Jadick et al. can be adjusted to the 5 mm slice thickness used in the present study. Specifically, the corresponding noise levels were equivalent to circa 18 HU, 13.8 HU, and 9.7 HU for the three dose levels. These values fall between the noise levels we measured in our virtual patient using the Br32f and the Br46f kernels, supporting the consistency of the simulation results.
Among other strategies to measure image quality features, namely contrast and resolution (9, 10), methods to measure noise are more technologically and clinically mature. Indeed, a few applications to assess noise magnitude have been already implemented in clinical practice (7, 25–27). For this reason, when it came to choosing an image quality measurand to assess the performance of diagnostic computed tomography exams in adults, the Centers for Medical and Medicaid Service proposed the noise magnitude (11). This CMS measure applies to payments for hospitals and clinicians: the associated data collection and reporting begun in January 2025, impacting the reimbursements starting in January 2027. Because of its critical role in the implementation of this new provision, it is essential that noise measurements are consistent, robust, reproducible, and standardized across institutions.
To that end, a commissioned panel from the American Association of Physicists in Medicine (AAPM) recently published a list of issues and ambiguities related to the CMS measure (28). In particular, the AAPM panel highlighted that: “It is unclear how the measurement of calculated CT global noise is performed given the lack of an official standard and limited documentation in the specification”. Specifically, the experts emphasized how peer-reviewed literature currently recognized two main methods to measure noise magnitude in patient CT images: the noise measured in soft tissues and the noise measured in the air surrounding the patient. Our study showed that the noise in air (NAIR) can underestimate the gold standard Nensemble by as much as 62%, whereas the GNI absolute average difference with Nensemble is 5%. Therefore, the new CMS rule cannot be implemented without referencing a single method to calculate noise magnitude or without resolving how values obtained with different methods can be unbiasedly compared.
This issue underscores the growing importance of image quality in assessing the diagnostic value of CT examinations. Historically, CT performance metrics focused almost exclusively on radiation dose, as before the implementation of automated current and kV modulation techniques, dose once served as a proxy for image quality. Moreover, in vivo image quality assessment required unavailable computational resources. However, both the International Commission on Radiological Protection (ICRP) and the International Atomic Energy Agency (IAEA) now emphasize that optimization must integrate radiation dose and image quality (6, 29). Because image quality is directly correlated with diagnostic detection (30, 31) and to the patient benefit, it is important to assess it in clinically relevant image areas, as shown in this study.
This study has some limitations that should be acknowledged. First, it considered virtual images generated with the simulator configured to replicate a clinical CT scanner. It is known that different vendors implement different strategies to balance radiation output and image quality (12, 24). Future studies can apply the presented methodology to include different CT models. Moreover, concerning the choice of simulating only adult chest and abdominopelvic clinical scenarios with three radiation outputs and images reconstructed with FBP and IR and three kernels at 5 mm slice thickness, it is reasonable to suppose that the specific anatomies and scanning parameters do not affect the trends reported in the study. Future comparisons including different patient anatomical regions for pediatric and adult patients, and different imaging techniques, as well as different range of tissues within a patient CT scan (i.e.: bones), can be performed to confirm this assumption. Finally, we note that there is more to image quality than noise alone. Caution is in order not to presume a singular assessment of noise magnitude can reflect image quality which includes a myriad of other features including noise texture, resolution, and contrast (28, 32).
7. Conclusion
Soft tissue-based noise measurements, particularly the global noise index (GNI), provide more accurate, robust, and reproducible representation of the gold standard ensemble noise in CT imaging, compared to measuring noise in air or phantom datasets. As healthcare policies increasingly mandate image quality evaluation alongside radiation dose, establishing a consistent methodology for noise assessment is essential to ensure reliable procedure optimization and equitable reimbursement practices.
Table 3b.
Standard deviation values for NACR, NAIR, GNI, Nensemble, Norgan for abdominopelvic images per each kernel, mAs value, and reconstruction algorithm involved in the study. Units are HU.
| Abdomen | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | NAIR | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 0.141 | 0.010 | 0.026 | 1.54 | 1.80 | 0.092 | 0.008 | 0.026 | 1.16 | 1.37 | |
| 100 | 0.095 | ||||||||||
| Br32f | mAs | 0.006 | 0.032 | 1.07 | 1.26 | 0.063 | 0.005 | 0.028 | 0.793 | 0.942 | |
| 150 | 0.060 | ||||||||||
| mAs | 0.005 | 0.022 | 0.869 | 1.02 | 0.041 | 0.004 | 0.001 | 0.635 | 0.751 | ||
| 50 mAs | 0.324 | 0.043 | 0.067 | 4.29 | 4.89 | 0.178 | 0.038 | 0.048 | 2.76 | 3.08 | |
| 100 | 0.208 | ||||||||||
| Br46f | mAs | 0.038 | 0.038 | 2.97 | 3.41 | 0.140 | 0.016 | 0.028 | 1.89 | 2.10 | |
| 150 | 0.157 | ||||||||||
| mAs | 0.023 | 0.040 | 2.41 | 2.77 | 0.083 | 0.013 | 0.030 | 1.52 | 1.69 | ||
| 50 mAs | 1.34 | 0.070 | 0.244 | 19.3 | 21.6 | 0.708 | 0.052 | 0.126 | 10.6 | 11.2 | |
| 100 | 0.812 | ||||||||||
| Br62f | mAs | 0.064 | 0.132 | 13.4 | 15.1 | 0.437 | 0.055 | 0.094 | 7.25 | 7.72 | |
| 150 | 0.615 | ||||||||||
| mAs | 0.050 | 0.106 | 10.9 | 12.3 | 0.275 | 0.044 | 0.079 | 5.84 | 6.25 | ||
Table 4b.
Coefficient of variation values for NACR, NAIR, GNI, Nensemble, Norgan for abdominopelvic images per each kernel, mAs value, and reconstruction algorithm involved in the study.
| Abdomen | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| |||||||||||
| FBP | IR | ||||||||||
|
| |||||||||||
| NACR | NAIR | GNI | Nensemble | Norgan | NACR | NAIR | GNI | Nensemble | Norgan | ||
|
| |||||||||||
| 50 mAs | 0.029 | 0.002 | 0.002 | 0.143 | 0.151 | 0.029 | 0.002 | 0.003 | 0.124 | 0.136 | |
| 100 | 0.027 | ||||||||||
| Br32f | mAs | 0.002 | 0.004 | 0.141 | 0.151 | 0.028 | 0.002 | 0.004 | 0.121 | 0.134 | |
| 150 | 0.021 | ||||||||||
| mAs | 0.002 | 0.003 | 0.140 | 0.150 | 0.022 | 0.002 | 0.001 | 0.119 | 0.132 | ||
| 50 mAs | 0.024 | 0.003 | 0.002 | 0.143 | 0.148 | 0.021 | 0.003 | 0.002 | 0.124 | 0.130 | |
| 100 | 0.022 | ||||||||||
| Br46f | mAs | 0.003 | 0.002 | 0.140 | 0.147 | 0.024 | 0.002 | 0.002 | 0.120 | 0.127 | |
| 150 | 0.020 | ||||||||||
| mAs | 0.002 | 0.002 | 0.139 | 0.147 | 0.018 | 0.002 | 0.002 | 0.119 | 0.125 | ||
| 50 mAs | 0.029 | 0.001 | 0.002 | 0.147 | 0.148 | 0.029 | 0.002 | 0.002 | 0.137 | 0.133 | |
| 100 | 0.026 | ||||||||||
| Br62f | mAs | 0.002 | 0.001 | 0.145 | 0.148 | 0.026 | 0.002 | 0.002 | 0.133 | 0.130 | |
| 150 | 0.024 | ||||||||||
| mAs | 0.001 | 0.001 | 0.144 | 0.147 | 0.020 | 0.002 | 0.002 | 0.132 | 0.130 | ||
Take home points.
Soft tissue-based noise measurements, particularly the global noise index (GNI), provide more accurate, robust, and reproducible representation of the gold standard ensemble noise in CT imaging, compared to measuring noise in air or phantom datasets. Soft tissue-based noise measurements is a close surrogate to inform protocol design and technology assessment.
Virtual imaging techniques enabled an unbiased comparison of different noise magnitude calculation methods in large and realistic populations simulating clinical conditions. As healthcare policies increasingly mandate image quality evaluation alongside radiation dose, establishing a consistent methodology for noise assessment is essential to ensure reliable procedure optimization and equitable reimbursement practices.
Funding
This work was funded in part by National Institutes of Health (P41EB028744, R44EB031658, R01HL155293).
Footnotes
Conflict of interest statement
Authors do not list relationships related to the present publication.
Declaration of interests
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:
Francesco Ria reports financial support was provided by National Institutes of Health. Ehsan Abadi reports financial support was provided by National Institutes of Health. Ehsan Samei reports financial support was provided by National Institutes of Health. Francesco Ria reports a relationship with Metis Health Analytics that includes: board membership. Ehsan Samei reports a relationship with Metis Health Analytics that includes: board membership. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Contributor Information
Francesco Ria, Carl E. Ravin Advanced Imaging Labs, Clinical Imaging Physics Group, Center for Virtual Imaging Trials, Departments of Radiology, Duke University Health System, 2424 Erwin Road, Suite 302, Durham, NC 27710, USA.
Martina Talarico, Carl E. Ravin Advanced Imaging Labs, Center for Virtual Imaging Trials, Departments of Radiology, Duke University Health System, 2424 Erwin Road, Suite 302, Durham, NC 27710, USA; Department of Medical Physics, SABES-ASDAA Bolzano Hospital, Via Lorenz Böhler 5 39100 Bolzano (BZ), ITALY.
Em Harkness, Carl E. Ravin Advanced Imaging Labs, Center for Virtual Imaging Trials, Departments of Radiology, Duke University Health System, 2424 Erwin Road, Suite 302, Durham, NC 27710, USA.
Ehsan Abadi, Carl E. Ravin Advanced Imaging Labs, Center for Virtual Imaging Trials, Departments of Radiology, Duke University Health System, 2424 Erwin Road, Suite 302, Durham, NC 27710, USA.
Ehsan Samei, Carl E. Ravin Advanced Imaging Labs, Clinical Imaging Physics Group, Center for Virtual Imaging Trials, Departments of Radiology, Duke University Health System, 2424 Erwin Road, Suite 302, Durham, NC 27710, USA.
Data statement
The authors declare that they had full access to all the data in this study and the authors take complete responsibility for the integrity of the data and the accuracy of the data.
8. References
- 1.Icrp. The 2007 Recommendations of the International Commission on Radiological Protection. ICRP Publication 103. Annals of the ICRP. 2007;37(2–4):9–34. [DOI] [PubMed] [Google Scholar]
- 2.Ria F, Fu W, Hoye J, Segars WP, Kapadia AJ, Samei E. Comparison of 12 surrogates to characterize CT radiation risk across a clinical population. European Radiology. 2021;31(9):7022–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ria F, Rehani MM, Samei E. Characterizing imaging radiation risk in a population of 8918 patients with recurrent imaging for a better effective dose. Scientific Reports. 2024;14(1):6240. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ria F, Zhang AR, Lerebours R, Erkanli A, Abadi E, Marin D, et al. Optimization of abdominal CT based on a model of total risk minimization by putting radiation risk in perspective with imaging benefit. Communications Medicine. 2024;4(1):272. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Samei E, Järvinen H, Kortesniemi M, Simantirakis G, Goh C, Wallace A, et al. Medical imaging dose optimisation from ground up: Expert opinion of an international summit. Journal of Radiological Protection. 2018;38(3):967–89. [DOI] [PubMed] [Google Scholar]
- 6.Vañó E, Miller DL, Martin CJ, Rehani MM, Kang K, Rosenstein M, et al. ICRP Publication 135: Diagnostic Reference Levels in Medical Imaging. Annals of the ICRP. 2017;46(1):1–144. [DOI] [PubMed] [Google Scholar]
- 7.Christianson O, Winslow J, Frush DP, Samei E. Automated Technique to Measure Noise in Clinical CT Examinations. American Journal of Roentgenology. 2015;205(1):W93–W9. [DOI] [PubMed] [Google Scholar]
- 8.Malkus A, Szczykutowicz TP. A method to extract image noise level from patient images in CT. Medical Physics. 2017;44(6):2173–84. [DOI] [PubMed] [Google Scholar]
- 9.Sanders J, Hurwitz L, Samei E. Patient-specific quantification of image quality: An automated method for measuring spatial resolution in clinical CT images. Medical Physics. 2016;43(10):5330–8. [DOI] [PubMed] [Google Scholar]
- 10.Abadi E, Sanders J, Samei E. Patient-specific quantification of image quality: An automated technique for measuring the distribution of organ Hounsfield units in clinical chest CT images. Medical Physics. 2017;44(9):4736–46. [DOI] [PubMed] [Google Scholar]
- 11.Excessive Radiation Dose or Inadequate Image Quality for Diagnostic Computed Tomography (CT) in Adults (Facility IQR), CMS1074v1 (2023). [Google Scholar]
- 12.Ria F, Solomon JB, Wilson JM, Samei E. Technical Note: Validation of TG 233 phantom methodology to characterize noise and dose in patient CT data. Medical Physics. 2020;47(4):1633–9. [DOI] [PubMed] [Google Scholar]
- 13.Fu W, Sharma S, Solomon J, Ria F, Setiawan H, Ding A, et al. Patient-specific organ dose and in-vivo image quality assessment in clinical CT. Physica Medica. 2025;136:105017. [DOI] [PubMed] [Google Scholar]
- 14.Segars WP, Bond J, Frush J, Hon S, Eckersley C, Williams CH, et al. Population of anatomically variable 4D XCAT adult phantoms for imaging research and optimization. Medical Physics. 2013;40(4):043701–. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Segars WP, Sturgeon G, Mendonca S, Grimes J, Tsui BMW. 4D XCAT phantom for multimodality imaging research. Medical Physics. 2010;37(9):4902–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.McCollough CH, Bruesewitz MR, McNitt-Gray MF, Bush K, Ruckdeschel T, Payne JT, et al. The phantom portion of the American College of Radiology (ACR) Computed Tomography (CT) accreditation program: Practical tips, artifact examples, and pitfalls to avoid. Medical Physics. 2004;31(9):2423–42. [DOI] [PubMed] [Google Scholar]
- 17.Abadi E, Harrawood B, Sharma S, Kapadia A, Segars WP, Samei E. DukeSim: A Realistic, Rapid, and Scanner-Specific Simulation Framework in Computed Tomography. IEEE Transactions on Medical Imaging. 2019;38(6):1457–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Abadi E, Segars WP, Tsui BMW, Kinahan PE, Bottenus N, Frangi AF, et al. Virtual clinical trials in medical imaging: a review. Journal of Medical Imaging. 2020;7(04):1–. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Abadi E, Harrawood B, Rajagopal JR, Sharma S, Kapadia A, Segars WP, et al. Development of a scanner-specific simulation framework for photon-counting computed tomography. Biomed Phys Eng Express. 2019;5(5). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Jadick G, Abadi E, Harrawood B, Sharma S, Segars WP, Samei E. A scanner-specific framework for simulating CT images with tube current modulation. Phys Med Biol. 2021;66(18). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Shankar SS, Felice N, Hoffman EA, Atha J, Sieren JC, Samei E, et al. Task-based validation and application of a scanner-specific CT simulator using an anthropomorphic phantom. Medical Physics. 2022;49(12):7447–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Sharma S, Abadi E, Kapadia A, Segars WP, Samei E. A GPU-accelerated framework for rapid estimation of scanner-specific scatter in CT for virtual imaging trials. Phys Med Biol. 2021;66(7). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.McCabe C, Harrawood B, Samei E, Abadi E. In silico modeling of a clinical photon-counting CT system: Verification and validation. Medical Physics. 2025;52(6):3840–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Samei E, Bakalyar D, Boedeker KL, Brady S, Fan J, Leng S, et al. Performance evaluation of computed tomography systems: Summary of AAPM Task Group 233. Medical Physics. 2019;46(11). [DOI] [PubMed] [Google Scholar]
- 25.Kuo HC, Mahmood U, Kirov AS, Mechalakos J, Della Biancia C, Cerviño LI, et al. An automated technique for global noise level measurement in CT image with a conjunction of image gradient. Phys Med Biol. 2024;69(9). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Satu II, Teemu M, Touko K, Juha P, Marko K, Mika K. Automatic head computed tomography image noise quantification with deep learning. Physica Medica. 2022;99:102–12. [DOI] [PubMed] [Google Scholar]
- 27.Ria F, Davis JT, Solomon JB, Wilson JM, Smith TB, Frush DP, et al. Expanding the Concept of Diagnostic Reference Levels to Noise and Dose Reference Levels in CT. American Journal of Roentgenology. 2019;213(4):889–94. [DOI] [PubMed] [Google Scholar]
- 28.Wells JR, Christianson O, Gress D, Gingold E, Jacobs J, Boedeker K, et al. The New CMS Measure of Excessive Radiation Dose or Inadequate Image Quality in CT: Issues and Ambiguities—Perspectives from an AAPM-Commissioned Panel. American Journal of Roentgenology. 2025. [DOI] [PubMed] [Google Scholar]
- 29.Iaea. Radiation Protection of Patients (RPOP) – Diagnostic Reference Levels (DRLs), International Atomic Energy Agency, 2017. [Available from: https://www.iaea.org/resources/rpop/health-professionals/radiology/diagnostic-reference-levels. [Google Scholar]
- 30.Smith TB, Abadi E, Solomon J, Samei E. Development, validation, and relevance of in vivo low-contrast task transfer function to estimate detectability in clinical CT images. Medical Physics. 2021;48(12):7698–711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zarei M, Ria F, Jensen CT, Liu X, Abbey CK, Samei E. Correlation of Automated in Vivo Image Quality With Radiologist’s Performance in Abdomen Computed Tomography Across Conventional and Deep Learning Reconstructions. Journal of Computer Assisted Tomography. 2026: 10.1097/RCT.0000000000001845. [DOI] [PubMed] [Google Scholar]
- 32.Gress DA, Samei E, Frush DP, Pelzl CE, Fletcher JG, Mahesh M, et al. Ranking the Relative Importance of Image Quality Features in CT by Consensus Survey. Journal of the American College of Radiology. 2025;22(1):66–75. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The authors declare that they had full access to all the data in this study and the authors take complete responsibility for the integrity of the data and the accuracy of the data.
