Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 12.
Published in final edited form as: J Expo Sci Environ Epidemiol. 2019 Sep 2;30(3):420–429. doi: 10.1038/s41370-019-0164-z

Influence of exposure measurement errors on results from epidemiologic studies of different designs

Jennifer Richmond-Bryant 1,2, Thomas C Long 1
PMCID: PMC13355284  NIHMSID: NIHMS2187219  PMID: 31477780

Abstract

In epidemiologic studies of health effects of air pollution, measurements or models are used to estimate exposure. Exposure estimates have errors that propagate to effect estimates in exposure-response models. We critically evaluate how types of exposure measurement error influenced bias and precision of effect estimates to understand conditions affecting interpretation of exposure-response models for epidemiologic studies of exposure to PM2.5, NO2, and SO2. We reviewed available literature on exposure measurement error for time-series and long-term exposure epidemiology studies. For time-series studies, time–activity error (daily exposure concentration did not account for variation in exposure due to time–activity during a day) and nonambient (indoor) sources negatively biased the effect estimates and increased standard error, so uncertainty grew with increasing bias while underestimating the true health effect in these studies. Spatial error (deviation between true exposure concentration at an individual’s location and concentration at a receptor) was ascribed to negatively biased effect estimates in most cases. Positive bias occurred for spatially variable pollutants when the variance of error correlated with the exposure estimate. For long-term exposure studies, most spatial errors did not bias the effect estimate. For both time-series and long-term exposure studies reviewed, large uncertainties were observed when exposure concentration was modeled with low spatial and temporal resolution for a spatially variable pollutant.

Keywords: Epidemiology, Exposure modeling, Criteria pollutants

Introduction

The US Environmental Protection Agency (EPA) is mandated to review the National Ambient Air Quality Standards (NAAQS) for criteria air pollutants (CAPs) [particulate matter (PM), nitrogen dioxide (NO2), sulfur dioxide (SO2), carbon monoxide (CO), ozone (O3), and lead (Pb)] under the Clean Air Act [1]. The NAAQS review process involves an evaluation of the scientific evidence base pertaining to each CAP in the Integrated Science Assessment (ISA), which draws from atmospheric science, exposure assessment, dosimetry, epidemiology, toxicology, controlled human exposure, and ecology literature, to evaluate whether CAP exposure results in health or ecological effects [24]. A relationship is deemed causal if it is possible to “rule out chance, confounding, and other biases.” To establish a causal determination, there must be consistency among high-quality studies within each discipline (epidemiology, toxicology, and controlled human exposure), coherence of conclusions among disciplines, and evidence of biological plausibility [5]. Epidemiology studies use estimates of exposure in their statistical models, because true exposure is unknown, especially for large populations. In the ISA, evaluation of exposure assessment methodologies used in the epidemiologic studies is provided to consider whether epidemiologic study results used to support a causal determination may be influenced by confounding, other biases, or uncertainties.

For the purpose of this paper, we refer to exposure measurement error as the bias and uncertainty associated with using concentration metrics to represent the actual exposure of an individual or population [6]. Exposure measurement error can result in bias and incorrect estimates of SEs of the effect estimate. Bias refers to the deviation between the true health effect averaged over a study population and the effect estimate for that population derived from a statistical model [7]. When the effect estimate is positively biased, the true health effect is smaller than the estimate (Fig. 1). A positively biased effect estimate can potentially undermine a causal determination, as it is unknown whether the true effect is different from the null. When the effect estimate is negatively biased, the true effect is larger than the estimate to produce a conservative effect estimate. When the effect estimate is unbiased, the true effect is well represented. Precision, the uncertainty of an effect estimate for a population, can be measured by the standard error (SE) of the effect estimate.

Fig. 1.

Fig. 1

Illustration of the influence of bias on health effect estimates

Exposure measurement error has two components: (1) exposure measurement error derived from uncertainty in the metric being used to represent exposure and (2) error due to use of a surrogate parameter of interest in the epidemiologic study in lieu of the true exposure, which may be unobservable. Classical exposure measurement error is defined as exposure measurement error scattered around the true personal exposure and independent of the level of the measured exposure. Classical exposure measurement error may occur when a fixed-site monitor measuring ambient concentration is imprecise, even if it is accurate. It is also independent of time and space [8]. Classical exposure measurement error can result in bias of the epidemiologic effect estimate. When variation in the exposure measurements is greater than variation in the true exposures, classical exposure measurement error typically biases the effect estimate negatively (indicating no or lesser effect of the exposure compared with the true effect). This would cause the effect to be underestimated. Classical exposure measurement error can also cause inflation or reduction of the SE of the effect estimate. Berkson exposure measurement error is defined as error scattered around the measured exposure surrogate (in most cases, the measured ambient concentration) and is independent of the true exposure [9, 10]. Berkson exposure measurement error may occur when the time series of ambient air pollutant concentrations measured at a monitor differs from the time series of a person’s true exposure such that the true variability in the person’s ambient air pollutant exposure goes unmeasured. Berkson exposure measurement error is not expected to bias the effect estimate. In reality, exposure measurement error has characteristics of both Berkson and classical error, which could introduce both bias and changes in the precision of the effect estimate.

Estimates of PM, NO2, and SO2 exposures are subject to errors from various sources and the point of reference (i.e., “true” exposure) varies with the error type (Table 1). Three exposure measurement error types are based on a reference of ambient concentration measured at a fixed location. Spatial error occurs when concentration is measured or modeled to provide an exposure surrogate at a location that differs from the location of the population centroid. In this case, the reference exposure concentration is the concentration at the location of the centroid of the exposed population. Instrument error occurs when the concentration reported by the monitor differs from the true concentration at that location due to instrument measurement interference or other artifacts. The reference in this case would be the concentration at the monitor location. Model misspecification occurs when important prediction variables are omitted from land-use regression (LUR) or geostatistical models used to estimate exposure concentrations. The reference exposure concentration in this case is the concentration surface over the modeled area of interest. For this review, we do not consider misspecification of the health effect model. Two exposure measurement error types consider personal exposure to the ambient air pollutant to be a reference. Time–activity error occurs when changes in exposure concentration occurring over time are not captured, because exposure concentration is measured or modeled at a single point. In this case, the reference exposure concentration is the concentration at the location of the exposed individual, integrated over time and over the population. Nonambient sources add to exposure measurement error by obscuring the exposure contribution of the ambient sources of primary interest for air pollution epidemiology. In this case, the reference exposure concentration is the exposure concentration due to ambient sources. Bias can be computed by comparing the effect estimates calculated using the exposure estimates with the effect estimates calculated using the reference concentrations.

Table 1.

Sources of exposure measurement error evaluated in the literature review

Error source Description

Spatial error Deviation between the true exposure concentration at the location of the individual and the concentration measured at the location of the monitor or modeled at the location of the receptor. The reference exposure concentration is the concentration at a fixed location of the exposed individual, such as a residence or place of work.
Time–activity error Assignment of a single daily exposure concentration does not account for variations in exposure throughout the day. The reference exposure concentration is the concentration at the location of the exposed individual, integrated over time.
Nonambient sources Because ambient air pollutant exposure is the metric of interest, nonambient exposures can be considered to bias the exposure estimate. A reference exposure concentration can be obtained from a personal monitor taking measurements when no nonambient sources exist or using tracer measurements (e.g., sulfate) to estimate the ambient portion of personal exposure.
Model misspecification The omission of prediction variables, when using land-use regression or geostatistical models to obtain exposure concentration estimates, can add error to the exposure concentration estimate and health effect estimate. A reference exposure concentration could, in theory, be derived from a correctly specified exposure model. Practically speaking, the correctly specified exposure model is not known. The reference exposure concentration is the concentration at the location of a receptor.
Instrument error Deviation of the measured concentration from the true concentration at the location of the monitor due to instrument measurement interference or other artifacts. The reference exposure concentration is the concentration at a collocated Federal Reference Method or Federal Equivalent Method monitor.

The objective of this study is to review the exposure assessment literature for three CAPs (PM, NO2, SO2), to identify patterns in biases and precision in the effect estimates related to sources of exposure measurement error. These CAPs were selected given differences in spatial variability related to their sources and atmospheric chemistry. Specifically, studies reporting bias and/or SE were evaluated for exposure measurement error type. Bias and SE data were analyzed by exposure measurement error type across studies, and emerging patterns were identified and considered for each study design.

Methods

Separate literature reviews for each CAP were conducted to follow the EPA’s systematic review process for the ISAs [5]. Briefly, each literature search began with a broad keyword search of PubMed and Web of Science for papers potentially related to PM, NO2, and SO2. See Supplementary Table S1 for search terms. The PM, NO2, and SO2 literature searches initially produced 305,098, 215,468, and 36,327 references, respectively. References were stored on the US EPA Health and Environmental Research Online database (http://hero.epa.gov). References were assigned to a discipline such as exposure assessment, epidemiology, and experimental studies using a topic classification algorithm based on a set of discipline-specific seed references. These seed references were selected to be representative of the type of literature relevant to each discipline in evaluating the health and environmental effects of CAPs. References from the literature searches with similar terminology to the seed references from a particular discipline were placed into the corresponding bin for further analysis. Applying this algorithm reduced the reference pool to 28,225 exposure assessment references for PM, 10,395 for NO2, and 180 for SO2. References were next screened by title to determine whether they were likely to present information on the impact of exposure measurement error on epidemiologic effect estimates. To be included in this analysis, references had to provide quantitative information on bias or a change in the SE of the effect estimate due to exposure measurement error of various types. Those that seemed likely to include such information were deemed “considered”; the remainder were discarded at this stage. Abstracts and possibly full text were then read to determine whether the references would be “included.” Overall, 8, 7, and 4 references related to exposure assessment were identified for PM, NO2, and SO2, respectively. All size fractions of PM were included, as were species related to PM, NO2, or SO2, such as elemental carbon (EC) in PM2.5 or oxides of nitrogen (NOX). Some studies included multiple pollutants; a total of 14 references were reviewed. Studies evaluating short-term (days to weeks) and long-term exposures (months to years) were analyzed separately by each author of this review. Upon comparison of our evaluations, we found that we determined error type differently for one long-term exposure study and two time-series studies. These differences were resolved through discussion and consensus building.

Data for bias and SE were extracted from the selected references. Several studies reported bias and SE data, which were extracted directly. Some studies presented bias and SE data graphically, in which case we estimated data from the figures. These distinctions are noted in Supplementary Table S2. Some studies presented the “true” health effect (βtrue) from a simulation or measurement and effect estimates (βest) from different scenarios where exposure measurement errors were modeled. If relative bias was not presented explicitly, it was estimated as the difference between the true effect (as defined by the study authors) and the effect estimate normalized by the true effect:

ϵ=βtrue-βestβtrue. (1)

In some cases, but not all, reported effect estimates were standardized by some increment. All effect estimates were adjusted to unit standardization for comparability. If the statistical model was not linear (e.g., for a logistic or Poisson health model), then the model was transformed to a linear form before the bias was extracted, so bias could be comparable across studies once unit scaling was applied.

Exposure measurement error type was determined based on expert judgment for each data point. Each author reviewed all studies and made an independent determination of error type. The authors discussed any differences along with their rationale for selecting a given error type and made a final determination during this discussion. If a study incorporated multiple error types in their analysis of bias in a way that the error due to one type could not be distinguished [11], then the study data were presented in the Supplementary Information (Table S2) but not included in the figures. Boxplots for bias stratified by exposure measurement error type were plotted using the R Statistical Programming Software (v.3.1.2).

Results

This analysis includes studies of the impact of exposure measurement error on effect estimates from time-series and long-term average exposure epidemiologic studies. For time-series studies, six studies evaluated spatial error [1217], two evaluated error due to mischaracterization of time–activity [15, 16], one evaluated error from nonambient sources [16], one evaluated model misspecification [18], and one evaluated instrument error [12]. For long-term average studies, two studies evaluated spatial error [18, 19], one evaluated error due to mischaracterization of time–activity [20], and four evaluated model misspecification [8, 2123]. The majority of these studies developed exposure surfaces from models, such as spatiotemporal, LUR, or kriging models, using monitoring data to train and/or validate the models. Simulation studies such as these are devised to compare the effect estimate from some predetermined error scenario with the true effect. The distributions of the input concentration data are specifically designed to embody the characteristics of the error type of interest.

Boxplots of bias allow for comparison of the relative magnitude and direction of bias among the different types of exposure measurement error studied for time-series and long-term exposure epidemiologic studies. This comparison is less focused on the exact values of the bias as the general pattern of data. Bias is presented by type of exposure measurement error in Fig. 2 for time-series epidemiologic studies and in Fig. 3 for long-term exposure epidemiologic studies. All data are provided in Supplementary Table S2.

Fig. 2.

Fig. 2

Relative bias across studies for each exposure measurement error type examined in the time-series epidemiologic literature reviewed. The solid vertical line indicates zero bias

Fig. 3.

Fig. 3

Relative bias across studies for each exposure measurement error type examined in the long-term exposure epidemiologic literature reviewed. The solid vertical line indicates zero bias

For time-series epidemiologic studies, median bias in the effect estimate was negative for time–activity error and spatial error, and it was near-zero for error due to nonambient sources, model misspecification, and instrument error. Although median bias due to spatial error was below zero, several positive and negative values beyond the interquartile range were observed for spatial errors. SE was shown to increase for spatial error when NO2 exposure was modeled for rural areas, spatial resolution decreased, and error was proportional to the exposure [17]. SE was near-zero for model misspecification. No data for SE were available for time–activity error.

For long-term exposure epidemiologic studies, median bias in the effect estimate was near-zero for spatial error and model misspecification. The bias data were evenly split between positive and negative values for spatial error and model misspecification. Four positive outliers were also observed for bias due to spatial error. Bias was uniformly negative for time–activity error. Nonambient sources and instrument error were not considered in the long-term exposure epidemiology studies. Unlike for time-series simulations, SE was low in most cases when spatial error was simulated, but a few exceptions were observed when coarse grid simulations were employed [19].

Discussion

Time-series epidemiologic studies

Spatial error

The mix of positive and negative biases observed for spatial errors within time-series studies implies a potential for overestimation or underestimation of the health effect. Several studies tested the influence of spatial error on effect estimates in time-series studies [1217]. Biases were negative or zero for the majority (83 of 121) of data points. Although median bias was just under zero, biases ranged from −0.88 to 0.77. The majority of negative bias outliers were computed for two studies of the metropolitan Atlanta area [12, 14]. In Goldman et al. [12], concentrations measured at Atlanta-area monitors were used to develop a spatial semivariogram function of concentration time series, which served as input to an assumed true exposure function when added to an assumed baseline concentration time series. Error was assumed to be the difference between the semivariogram-derived concentration estimate and the baseline. This study was designed to assess the impact of error from using a fixed-site monitor rather than exposure at the study participant’s location [12]. Strickland et al. [14] compared effect estimates for metropolitan Atlanta obtained when exposures were estimated by the nearest fixed-site monitor concentrations, unweighted average concentrations, and population-weighted average concentrations, with health effects calculated using an assumed “true” exposure from a spatially dense model based on concentration data distribution in each sampling grid. Berkson errors were estimated by comparing effect estimates obtained by simulating health effects from a Poisson distribution using the “true” exposure with health effects data, and classical errors were estimated by adding errors from collocated monitors from the Goldman et al. study [12]. The Berkson and classical errors were superimposed on the census tract centroid estimates of concentration prior to comparison of the effect estimates for each of the three exposure estimation schemes with that for the “true” exposure. The Berkson component of error produced positive bias in the effect estimate for the most spatially variable pollutants where the variance of the error was correlated with the exposure estimate. However, overall errors produced negative bias in the effect estimate, which were more pronounced for the fixed-site monitor and less so for the population-weighted average. Bias outliers were observed for spatially variable SO2, NO2, NOX, and EC in both studies, with outliers coming from the exposure assignments from the nearest fixed-site monitors.

In a follow-up study, Goldman et al. [13] extended the true exposure model to include temporal autocorrelation for each pollutant studied. That model was compared with exposure estimates based just on monitored concentrations, from which data were aggregated as a centrally located monitor and using unweighted, population-weighted, and area-weighted averages of concentrations as exposure surrogates. Goldman et al. [13] computed the largest negative bias for NOX exposure estimated by an area-weighted average. All positive bias outliers were published in that study of the impact of exposure measurement errors resulting from the use of fixed-site monitors for SO2, NOX, NO2, and EC or, in one case, from the use of an unweighted average for SO2. With the exception of OC, pollutants associated with bias outliers tended to be spatially variable, and the single monitor and unweighted average approaches assigned a uniform concentration across a large radius. Similarly, Dionisio et al. [15] simulated spatial error by taking the difference between concentrations measured at a fixed-site monitor and concentration calculated using the AERMOD dispersion model [24] at the location of simulated individuals assumed to live at the centroid of 193 Zip Codes in the metropolitan Atlanta area [15]. They obtained a negative bias outlier for EC. Most studies producing a large underestimation of health effects involved using a single fixed-site monitor to estimate exposure to spatially variable air pollutants among study participants in a large urban area.

SE data were available for fewer studies compared with bias, with only 67 data points for SE compared with 121 for bias. SE was estimated for the Sheppard et al. [16], Butland et al. [17], Strickland et al. [14], and Goldman et al. [12] studies. Butland et al. [17] simulated error across a domain subdivided into 1, 2, 3, 10, or 25 grids, to examine how the average error changes with increasing spatial resolution of the model for urban and rural NO2 exposures when additive or proportional error models were used and concentration provided the exposure surrogate. Relative bias decreased with increasing grid resolution for both error models and for urban and rural settings. However, SE was markedly higher for both proportional error models, with the rural model having the highest levels of imprecision. Relative bias decreased similarly with increasing grid resolution for each error model-setting combination, but the range of SE was distinct for each combination in Butland et al. [17]. Otherwise, SE was low for all other studies. Similarly, Sheppard et al. [16] simulated the bias and variability associated with using 1, 3, or 10 concentration monitors to represent average population exposure in a time-series model of the health effects from exposure to ambient PM2.5. Bias was larger in magnitude and negative when using one monitor to represent ambient exposure, but the magnitude of bias was the same for the use of three or ten monitors to represent average population exposure. SE did not change appreciably as the number of monitors increased. The spatial variability of PM2.5 modeled by Sheppard et al. [16] accounted for a small fraction of the total variability in PM2.5, based on monitoring data used to fit the exposure distributions. In comparison, spatial variability of NO2 modeled in Butland et al. [17] contributed a larger fraction of total variability of the concentration.

Time–activity error

Time–activity error appeared to produce negative bias and increased SE in time-series studies, resulting in underestimation of the true health effect in two studies. Dionisio et al. [15] calculated exposure measurement errors and biases by comparing exposure estimates and effect estimates for a model with no time–activity data with those calculated using the Air Pollutants Exposure population exposure model, which accounts for time–activity patterns and was considered to produce a set of “true” exposures. This study was conducted in 193 Zip Codes in the metropolitan Atlanta area. The effect estimates were all negatively biased for NOX, PM2.5, PM2.5-EC, and PM2.5-SO42− with the greatest bias estimated for NOX (90%). Sheppard et al. [16] conducted a time-series simulation to evaluate the impact of personal monitoring data availability on population-level effect estimates from PM2.5 exposure. Sheppard et al. [16] simulated increasing numbers of personal exposure monitors in the dataset and then compared across simulations, with the reference effect estimate derived from 100,000 simulated personal monitors to represent exposures of 100,000 simulated individuals. For this simulated population, using a single personal monitor to represent the average population exposure resulted in high negative bias (90%). Bias decreased with increasing number of personal monitors, with bias almost eliminated when using 100 personal monitors to represent the exposures of 100,000 people. However, SE increased with increasing number of personal monitors, indicating that increasing the amount of individual-level temporal variation in exposure decreases the amount of precision in the effect estimate.

Nonambient sources

One study in this review examined nonambient sources [16]. The influence of nonambient sources on effect estimates for ambient concentrations is usually small. However, temporal variation in the fraction of ambient concentrations to which people are exposed can bias effect estimates when the variations are temporally correlated with changes in ambient concentration. Sheppard et al. [16] conducted simulations adding within-individual variance in nonambient exposure and calculated near-zero bias. SE was virtually unchanged across nonambient exposure scenarios despite different assumptions regarding the influence of ambient and nonambient exposures, suggesting that nonambient sources do not substantially change the variability of effect estimates based on ambient concentrations. Adding between-individual variance to the case assuming normally distributed within-individual variance resulted in even less bias, indicating that the nonambient exposure contribution only added Berkson error to the effect estimate. Variation in ambient exposure fraction (i.e., fraction of ambient concentration to which an individual is exposed, which varies for location and activity) across the simulated population also produced essentially no bias, whether the variation was normally or uniformly distributed. However, adding a temporal component to the variation resulted in a negatively biased effect estimate when the trend in ambient exposure fraction was positively correlated with the annual trend in PM2.5. It resulted in positive bias when the trends were negatively correlated. The magnitude of bias increased with increasing correlation.

Model misspecification

Alexeef et al. [18] examined model misspecification for time-series studies. Model misspecification produced mostly small positive biases in effect estimates, but in a few cases the effect estimate was negatively biased. Alexeef et al. [18] compared kriging simulations with results from LUR with different combinations of inputs. One set of LUR simulations included spatial variables such as distance-to-road, density of major roads within 1 km, and vegetation, along with temporally varying meteorological variables for humidity, wind speed, and planetary boundary layer. A second set of LUR simulations only included spatial variables. A third set of LUR simulations used a two-step process to adjust the spatial model for daily variation. The LUR simulations with spatial variables and temporally variable meteorology produced larger negative biases with high SE relative to the other simulations. The authors attributed these modeling errors to oversimplification of the terms resulting in an inaccurate representation of complex spatial and temporal phenomena [18]. In contrast, the spatial-only simulations and two-step models both had very small positive biases with lower SE that were comparable to those produced from the kriging model. Rather than being attributed to resolution itself, inclusion of variables that capture spatial and temporal contrasts was thought to result in a decrease in the magnitude of bias in the effect estimate [18].

Instrument error

One study in this review examined instrument error [12]. Small biases due to instrument error are unlikely to produce a meaningful change in effect estimates. Goldman et al. [12] studied the influence of instrument error on bias for NO2, NOX, SO2, PM10, PM2.5, and PM2.5 species (EC, NH4, NO3, OC, and SO42−) measured in metropolitan Atlanta. Biases ranged from −0.057 (for PM10) to 0.021 (for NO3), with median bias of −0.040. In most cases, one would expect health effects to be slightly overestimated due to instrument error, which should not impact interpretation of the health effect. If an underestimate were to occur, it would be small enough to not impose a material change in a causal determination. No studies investigated the influence of instrument error on variability.

Long-term exposure epidemiologic studies

Spatial error

Positive biases related to spatial errors present the potential for overestimation of health effects, but in most cases the literature on spatial error reported positive biases within 5% of null. Four positive outliers were observed for linear exposure-response models [18]. In that study, different spatial exposure assignment approaches were used to represent PM2.5 exposure in a linear health effects model over a period of 32 days: two types of kriging and LUR. Higher positive biases were computed when a single value was assigned to each monitor over the study period and kriging was used for spatial interpolation of these period-level averages. The positive bias depended on the degree of smoothness of the concentration surface and the number of monitors, with more monitors resulting in lower bias. Biases were either positive but within the 5–95% whiskers, or negative for LUR or the other kriging approach, where daily concentration was assigned at each monitor then averaged over the study period. Hence, coarse temporal resolution coupled with coarse spatial resolution resulted in substantial overprediction of the health effect, whereas increased temporal resolution in the kriging method and inclusion of covariates in the LUR reduced this overprediction, in some cases resulting in negative bias [18]. Comparison of several approaches for spatial exposure estimation for a health effect model by Gryparis et al. [19], including Bayesian approaches and regression calibration, indicated that most approaches resulted in relatively little bias in the effect estimate. The largest negative outliers in bias occurred when the authors used exposure estimates generated randomly from a prior inverse Gamma data distribution to model concentration profiles with high and moderate spatial variability, respectively [19]. Biases were much smaller and still negative when a simplified covariance model was used. The relatively good performance of multiple approaches provides flexibility in model selection based on sample size and computational intensity. Although exposure predictions were inaccurate when the variance structure was not accurately modeled, they underestimated the health effects and so were conservative.

Coarse spatial and temporal resolution generally produced the highest SE values. SE was highest for the coarse temporal kriging approach in the Alexeef et al. [18] study, especially when fewer virtual monitors were used. The other approaches produced SE nearzero, including for the spatial LUR simulations. Rough concentration surfaces, combined with Bayesian exposure and health models, resulted in elevated SE in the Gryparis et al. [19] study. The approaches producing the highest negative bias had somewhat elevated SE values, although not the highest observed. Lower SE values were found for the simulated true exposure compared with the other tested exposure estimation methods.

Time–activity error

Setton et al. [20] examined time–activity error, which produced slight to moderate negative bias, resulting in underestimation of the effect estimate. This simulation study used annual average NO2 concentration surfaces for Vancouver, BC, and the California South Coast Air Basin, to evaluate the influence of time spent away from home on effect estimates. The authors constructed concentration surfaces using either LUR and IDW (Vancouver) or CAMx (California). Exposure estimates were generated for two situations: individuals spent all their time within 5 km of their residence and individuals spent time further away from home and accordingly experienced different NO2 concentrations. Bias in the effect estimate was evaluated by comparing the residential exposure estimate with that of the mobility-based exposure estimate. The authors used a bias factor estimator designed to adjust the classical error model to account for covariance between the true and surrogate exposure, considering the mobility-based exposure estimate to be the true exposure and the residential estimate as the surrogate exposure. Both time spent away from home and distance from home resulted in negatively biased effect estimates, with time having a stronger effect. The authors attributed the negative bias to the lower variance of the mobility exposure estimate, which was due to a spatial averaging effect resulting from time-weighted exposure concentrations in different parts of the city. The residential exposure estimate better reflected the spatial variability in NO2 concentrations and the spatial averaging resulted in classical error and negative bias. This suggests that not accounting for time–location information results in underestimation of long-term effects of air pollution.

Model misspecification

The effect of model misspecification on bias and SE has been tested in several simulation papers. Szpiro et al. [8] developed an LUR with a matrix of covariates (fully specified exposure model) and then omitted one of the geographic covariates for a nonspecific air pollutant (misspecified exposure model). Exposure estimates from the fully specified model had better agreement with the designated true exposure surface compared with the misspecified model. However, the misspecified model produced nearly unbiased effect estimates, whereas negative biases in the effect estimates were estimated for the fully specified model. SE was higher for the misspecified model than for the fully specified model in the case where the variability in the true exposure surface was σ = 1 but was lower for the misspecified model when σ = 0.1. The findings of Szpiro et al. [8] suggested that model misspecification does not necessarily lead to a biased or imprecise effect estimate. Szpiro and Paciorek [21] used a two-stage health effects model and compared effect estimates derived from using a misspecified exposure model with that derived using a bias correction-bootstrap variance fitting model robust to model misspecification. In this study, small-magnitude negative bias in the effect estimate was observed. Szpiro and Paciorek [21] noted that bias correction adjusted the model for bias due to imprecision of the effect estimate, due in part to model misspecification. They further noted that model misspecification was simulated via an estimator of the exposure mean and variance, but the error stemmed from an overly smoothed exposure surface due to an inadequate number of samplers in the study design. Cefalu and Dominici [25] performed simulations of misspecified PM2.5 exposure assessment models with correctly specified and misspecified health effects models. Results of that study are not presented here, because the authors did not examine direction of bias. However, their results supported Szpiro and Paciorek [21] by elucidating conditions that would prevent effect estimate bias due to exposure model misspecification: correct specification of the exposure model, lack of correlation between the exposure covariates and the health effect confounders, and inclusion of the health effect confounders in the exposure model. SE was elevated when the spatial variability of the pollutant was high and roughly a factor of ten lower when spatial variability of the pollutant was low, regardless of the model specification. Model variance was not improved through application of bias correction and bootstrap estimation of variance, indicating unaddressed error.

Bergen et al. [22] fit multivariable partial least-squares models to data for PM2.5 components (EC, OC, S, Si) in a two-stage model and found that the variance was either unchanged or increased after application of bias correction to address error due to model covariate selection. The authors note that the exposure prediction model fit moderately well but that the exposure model would have been improved with inclusion of geographic covariates for wood-burning sources, and that model fit may have been inflated by overfitting. The majority of data points in this analysis came from Bergen and Szpiro [23] for PM2.5, in which they modeled exposure with and without a set of spatial covariates in a two-stage model. This study design provides data on the influence of model misspecification. Bias correction was applied to some points in this study, which were not included here, because the present study evaluated the influence of measurement error type on bias. Values of bias were negative and generally well within 5% for PM2.5, but SE was generally lower when no geographic covariates were included in the exposure model. This result is similar to Szpiro et al. [8] for a pollutant with less spatial variability. Substantial bias was not noted for more spatially variable PM2.5 components (EC, OC, S, Si) [22] and NOX [21]. The two larger values of bias were found when pollutant type was not specified in simulations [8].

Study strengths and limitations

To our knowledge, this is the first review that surveys the air quality exposure assessment literature to extract relationships between different types of exposure measurement errors with biases of health effect estimates. Our review followed a systematic process, beginning with screening hundreds of thousands of references until 14 remained to form the core of this analysis. In addition, our review incorporated two epidemiologic study designs, five types of errors, and three air pollutants plus several PM2.5 species. Including data from different pollutants enabled consideration of the chemistry, transport, and dispersion characteristics of the pollutants when analyzing data extrema in the bias distribution. The majority of studies used simulated exposure error scenarios to contrast with “true” exposure conditions, which were either simulated or taken from monitored concentrations. Although contrived, simulating exposure error allows the study authors to have greater control over the type and magnitude of exposure errors introduced so that the resulting biases and variations in the effect estimates can be ascribed to a specific source of error, which is more difficult to achieve with field data. As a result, simulation studies become a valuable tool in understanding the impact of exposure measurement error on effect estimates. However, when considering different types of errors within different epidemiologic study designs, a small number of studies were available per error type. For this reason, we explore the direction and general magnitude (large or small) of bias and variance of the effect estimate rather than focusing on specific values. Furthermore, several spatial error studies were performed by a single research group for the city of Atlanta. A greater diversity of locations and researchers would add confidence to the body of literature.

Conclusions

We identified several types of exposure measurement errors and their associated biases and influence on variability for effect estimates in short-term and long-term air-pollutant exposure studies. Despite study design differences, the direction of biases was similar within most types of error. Biases and SE typically had larger magnitude for times-eries studies compared with long-term studies, likely because the data are not as smooth in time-series studies as they are in long-term studies. An exception was for model misspecification. Given findings by Szpiro et al. [8] that model misspecification does not always lead to greater bias, these findings are not surprising.

Positive bias indicates the effect estimate is larger than the true health effect. Positive bias occurred in cases where the exposure data were not suitable for the problem studied, such as when spatial error resulted from a single fixed-site monitor being used as an exposure surrogate for a spatially variable air pollutant in time-series studies. Identification of these characteristics in the exposure models used for epidemiologic studies may undermine causal determinations, because the true effect is thought to be smaller than observed and may even be null. These results are informative for study design, because sampling strategies can be developed that are appropriate for the spatial characteristics of the pollutant being studied to avoid producing positively biased effect estimates.

Negative bias indicates the effect estimate is smaller than the true health effect. Hence, existence of the effect is not in doubt. The magnitude of effect is uncertain but known to be larger than shown by available data. Time–activity error, instrument error, many cases of spatial error, and nonambient sources (in time-series studies) were observed to negatively bias effect estimates with some imprecision. Negatively biased effect estimates also tended to occur when exposure model inputs were misspecified, e.g., when using the wrong distribution of exposure data in spatial error studies, or when time–activity patterns were highly correlated with ambient concentrations in nonambient source studies. These types of exposure measurement errors and related biases are unlikely to undermine causal determinations, even with greater uncertainty, because the true effect is larger than observed.

Often, the conditions leading to exposure measurement error can be anticipated before a study is conducted. Based on the results of this review, exposure measurement error more often than not led to bias towards the null, meaning that health effects were likely even larger than predicted in most cases. In general, when conditions leading to biased and uncertain effect estimates are recognized, they may be addressed to some extent during study design.

Supplementary Material

JESEE supplement

Supplementary information The online version of this article (https://doi.org/10.1038/s41370-019-0164-z) contains Supplementary Material, which is available to authorized users.

Acknowledgements

We thank Dr Kathie Dionisio, Dr Rebecca Nachman, Dr Andrew Hotchkiss, and Dr John Vandenberg for their insightful comments.

Footnotes

Conflict of interest The authors declare that they have no conflict of interest.

Compliance with ethical standards

Disclaimer The study was reviewed by the EPA-NCEA and approved for publication. Mention of trade names or commercial products does not constitute endorsement or recommendation for use. Views expressed here are those of the authors and do not necessarily reflect EPA’s views or policies.

References

  • 1.Clean Air Act, as amended by Pub. L. No. 101–549 (1990). [Google Scholar]
  • 2.U.S. EPA. Integrated science assessment for particulate matter. EPA Report. EPA/600/R-08/139F. Research Triangle Park, NC: U.S. Environmental Protection Agency, Office of Research and Development, National Center for Environmental Assessment-RTP Division, 2009. [Google Scholar]
  • 3.EPA U.S. Integrated science assessment for oxides of nitrogen (final report). EPA Report. EPA/600/R-15/068. Research Triangle Park, NC: U.S. Environmental Protection Agency, National Center for Environmental Assessment, 2016. [Google Scholar]
  • 4.U.S. EPA. Integrated science assessment for sulfur oxides: Health criteria. EPA Report. EPA/600/R-17/451. Research Triangle Park, NC: U.S. Environmental Protection Agency, Office of Research and Development, National Center for Environmental Assessment-RTP, 2017. [Google Scholar]
  • 5.EPA U.S. Preamble to the Integrated Science Assessments. EPA Report. EPA/600/R-15/067. Research Triangle Park, NC: National Center for EnvironmentalAssessment, Office of Research and Development, 2015. [Google Scholar]
  • 6.Lipfert FW, Wyzga RE. The effects of exposure error on environmental epidemiology. In Proceedings of the 2nd Colloquium on Particulate Air Pollution and Health, Park City, UT, 1996. [Google Scholar]
  • 7.Armstrong BK, White E, Saracci R. Principles of exposure measurement in epidemiology. New York, NY: Oxford Univ. Press; 1992. [Google Scholar]
  • 8.Szpiro AA, Paciorek CJ, Sheppard L. Does more accurate exposure prediction necessarily improve health effect estimates? Epidemiology. 2011;22:680–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Goldman GT, Mulholland JA, Russell AG, Strickland MJ, Klein M, Waller LA, et al. Impact of exposure measurement error in air pollution epidemiology: effect of error type in time-series studies. Environ Health. 2011;10:61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Reeves GK, Cox DR, Darby SC, Whitley E. Some aspects of measurement error in explanatory variables for continuous and binary regression models. Stat Med. 1998;17:2157–77. [DOI] [PubMed] [Google Scholar]
  • 11.Basagaña X, Aguilera I, Rivera M, Agis D, Foraster M, Marrugat J, et al. Measurement error in epidemiologic studies of air pollution based on land-use regression models. Am J Epidemiol. 2013;178:1342–6. [DOI] [PubMed] [Google Scholar]
  • 12.Goldman GT, Mulholland JA, Russell AG, Srivastava A, Strickland MJ, Klein M, et al. Ambient air pollutant measurement error: characterization and impacts in a time-series epidemiologic study in Atlanta. Environ Sci Technol. 2010;44:7692–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Goldman GT, Mulholland JA, Russell AG, Gass K, Strickland MJ, Tolbert PE. Characterization of ambient air pollution measurement error in a time-series health study using a geostatistical simulation approach. Atmos Environ. 2012;57:101–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Strickland MJ, Gass KM, Goldman GT, Mulholland JA. Effects of ambient air pollution measurement error on health effect estimates in time-series studies: a simulation-based analysis. J Expo Sci Environ Epidemiol. 2013;25:160–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dionisio KL, Baxter LK, Chang HH. An empirical assessment of exposure measurement error and effect attenuation in bipollutant epidemiologic models. Environ Health Perspect. 2014;122:1216–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Sheppard L, Slaughter JC, Schildcrout J, Liu JS, Lumley T. Exposure and measurement contributions to estimates of acute air pollution effects. J Expo Anal Environ Epidemiol. 2005;15:366–76. [DOI] [PubMed] [Google Scholar]
  • 17.Butland BK, Armstrong B, Atkinson RW, Wilkinson P, Heal MR, Doherty RM, et al. Measurement error in time-series analysis: a simulation study comparing modelled and monitored data. BMC Med Res Methodol. 2013;13:136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Alexeeff SE, Schwartz J, Kloog I, Chudnovsky A, Koutrakis P, Coull BA. Consequences of kriging and land use regression for PM2.5 predictions in epidemiologic analyses: insights into spatial variability using high-resolution satellite data. J Expo Sci Environ Epidemiol. 2015;25:138–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Gryparis A, Paciorek CJ, Zeka A, Schwartz J, Coull BA. Measurement error caused by spatial misalignment in environmental epidemiology. Biostatistics. 2009;10:258–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Setton E, Marshall JD, Brauer M, Lundquist KR, Hystad P, Keller P, et al. The impact of daily mobility on exposure to traffic-related air pollution and health effect estimates. J Expo Sci Environ Epidemiol. 2011;21:42–8. [DOI] [PubMed] [Google Scholar]
  • 21.Szpiro AA, Paciorek CJ. Measurement error in two-stage analyses, with application to air pollution epidemiology. Environmetrics. 2013;24:501–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Bergen S, Sheppard L, Sampson PD, Kim SY, Richards M, Vedal S, et al. A national prediction model for PM2.5 component exposures and measurement error-corrected health effect inference. Environ Health Perspect. 2013;121:1017–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Bergen S, Szpiro AA. Mitigating the impact of measurement error when using penalized regression to model exposure in two-stage air pollution epidemiology studies. Environ Ecol Stat. 2015;22:601–31. [Google Scholar]
  • 24.Cimorelli AJ, Perry SG, Venkatram A, Weil JC, Paine R, Wilson RB, et al. AERMOD: a dispersion model for industrial source applications. Part I: general model formulation and boundary layer characterization. J Appl Meteorol. 2005;44:682–93. [Google Scholar]
  • 25.Cefalu M, Dominici F. Does exposure prediction bias health-effect estimation? The relationship between confounding adjustment and exposure prediction. Epidemiology. 2014;25:583–90. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

JESEE supplement

RESOURCES