Skip to main content
HHS Author Manuscripts logoLink to HHS Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 7.
Published in final edited form as: Health Place. 2020 Mar 13;63:102324. doi: 10.1016/j.healthplace.2020.102324

Assessing county-level determinants of diabetes in the United States (2003–2012)

Justin M Feldman a,*, David C Lee a, Priscilla Lopez a, Pasquale E Rummo a, Annemarie G Hirsch b, April P Carson c, Leslie A McClure d, Brian Elbel a, Lorna E Thorpe a
PMCID: PMC12965709  NIHMSID: NIHMS2136776  PMID: 32217279

Abstract

Using data from the United States Behavioral Risk Factor Surveillance System (2003–2012; N = 3,397,124 adults), we estimated associations between prevalent diabetes and four county-level exposures (fast food restaurant density, convenience store density, unemployment, active commuting). All associations confirmed our a priori hypotheses in conventional multilevel analyses that pooled across years. In contrast, using a random-effects within-between model, we found weak, ambiguous evidence that within-county changes in exposures were associated with within-county change in odds of diabetes. Decomposition revealed that the pooled associations were largely driven by time-invariant, between-county factors that may be more susceptible to confounding versus within-county associations.

Keywords: Diabetes, Built environment, Multilevel modeling

1. Introduction

Beginning in 1990, diabetes prevalence rapidly increased among adults throughout the United States (Geiss et al., 2012, 2017; 2014; Menke et al., 2015). While prevalence plateaued for the overall population by 2008, the burden of diabetes continues to rise among non-Hispanic black and Hispanic populations as well as individuals of low socioeconomic position (Geiss et al., 2014). This increase reflects a global epidemic of type-2 diabetes that has begun in different countries at different times over the past several decades and, while largely linked to changes in the global food system, has had heterogeneous impacts across places and social groups (Popkin, 2015).

Experimental and quasi-experimental research examining the long-term effects of exposure to neighborhood socioeconomic deprivation provides evidence of a causal relationship between geographic areas and the risk of diabetes among their residents. One such study was Moving to Opportunity, which assigned randomly selected families living in public housing in high-poverty neighborhoods to receive vouchers for subsidized housing in low-poverty neighborhoods in the United States during the period 1994–1998; an analysis of glycosylated hemoglobin (HbA1C) levels – a marker of uncontrolled diabetes – measured after a mean of 13 years of follow-up, found a reduction in diabetes of 4.3 percentage points (95% CI: 0.8, 7.8) among the treatment group relative to controls (Ludwig et al., 2011). A second study leveraged the quasi-random assignment of new refugees to neighborhoods of varying levels of economic deprivation throughout Sweden during 1987–1991, ascertained annual diabetes status using national health registry data during 2002–2008, and found an cumulative increase in risk over the 9-year follow-up period and an overall odds ratio of 1.2 (95% CI: 1.1, 1.4) for the effect of assignment to a high-deprivation versus low-deprivation neighborhood (Ludwig et al., 2011; White et al., 2016). While these two studies did not analyze specific environmental exposures beyond neighborhood socioeconomic composition, additional factors hypothesized to increase risk of diabetes include air pollution, noise, unhealthy food environments, and poor walkability; evidence from various observational study designs supports these as risk factors (den Braver et al., 2018; Dendup et al., 2018).

The present study builds upon this growing body of research by applying a novel epidemiologic analytic approach to perform temporal, multilevel analysis of nationally representative, repeated cross-sectional data in order to assess the relationship between diabetes and relevant county-level environmental measures that are available for the entire US over our study period (2003–2012) involving economic context, active commuting, and the food environment.

1.1. Geographic determinants of diabetes and confounding

Confounding is a key concern in analyses of hypothesized geographic determinants of diabetes as with other health outcomes. A person’s place of residence is strongly influenced by social processes (e.g. segregation by socioeconomic position, race/ethnicity, and prior health status), making it difficult to differentiate between causes of diabetes attributable to unmeasured, individual-level characteristics versus place-level characteristics (Macintyre and Ellaway, 2000). Additionally, multiple environmental hazards often cluster in the same places (Honold et al., 2012), meaning that attributing diabetes risk to one specific place-level exposure is difficult – there is likely to be confounding by other, unmeasured place-level factors. Despite the well-known limitations of cross-sectional designs in controlling for confounding, including confounding by prior disease status, cross-sectional research is very common in studies of geographic determinants of diabetes, accounting for 24 of the 50 included non-ecologic studies in a recent literature review (Dendup et al., 2018).

Data from repeated cross-sectional surveys that include diabetes status are widely available and have been used in multilevel analyses of geographic determinants of diabetes. Repeated cross-sectional surveys sample a population over multiple waves but, unlike longitudinal designs, they do not follow the same respondents over time. The surveys are typically meant to provide a sample that is representative of a unit (such as a geographic area) during a given period. While their primary purpose within epidemiology is disease surveillance, researchers often analyze cross-sectional surveys to identify potential causes of disease using either individual-level covariates collected in the survey or contextual exposures determined by respondents’ geographic place of residence (or nesting within another type of higher-level unit, such as a school). When researchers analyze repeated cross-sectional surveys for this purpose, they generally ignore temporality by treating the data as a single, cross-sectional sample or by pooling across survey years, losing informative data about change over time in the exposures and outcomes (Jia et al., 2009; Mehta and Chang, 2008).

In econometrics, it has long been practice to treat repeated cross-sectional data as a “pseudo-panel” and apply the same longitudinal methods that one would employ for data collected from a true cohort design (Verbeek, 1996). One such longitudinal method especially well-suited to multilevel datasets is the random-effects within-between (REWB) model (Bell et al., 2019; Bell and Jones, 2015; Fairbrother, 2011), although it has rarely been employed in epidemiology (Astell--Burt and Feng, 2018; Dieleman and Templin, 2014) and, to our knowledge, never for repeated cross-sectional health surveys. While repeated cross-sectional surveys cannot account for reverse causation because they lack data on the timing of disease development, each wave is in essence a longitudinal measure of a geographic area (Fairbrother, 2014). REWB models can take advantage of the survey design by simultaneously estimating cross-sectional (between-group) and temporal (within-group) associations for group-level predictors. They do so by decomposing each group-level exposure into two variables: (1) the overall, time-invariant group mean over the entire study period (the “between” coefficient); and (2) the time-varying, group-mean-centered value for each wave, e.g. the group value at time t minus the group mean (the “within” coefficient). The longitudinal associations estimated by REWB are equivalent to the coefficients estimated by the more-commonly employed econometric fixed-effects model, as both use group-mean-centering to estimate the change in an independent variable associated with a change in the dependent variable. In the case of a dichotomous outcome, REWB logistic regression is equivalent to a conditional logistic fixed-effects model. When the identification assumptions are met for an econometric fixed-effects model (Table 1), within-group associations are not subject to bias by group-level, unmeasured, time-invariant confounding (Angrist and Pischke, 2009; Imai and Kim, 2019).

Table 1.

Identifying assumptions of pooled multilevel models versus random-effects within-between (REWB) models.

Pooled models REWB models (within-unit estimators)
Assumes no unmeasured time-stable confounding and no unmeasured time-varying confounding Assumes no unmeasured time-varying confounding
Assumes no selection bias Assumes no selection bias
Does not require defining an etiologic period (i.e. time between exposure and disease development) Requires, and is potentially sensitive to, specification of etiologic period
No explicit assumptions regarding causal dynamics Assumes no causal dynamics (past exposures cannot affect future outcome, past outcome cannot affect future exposure)
More robust to measurement error in independent variables Less robust to measurement error in dependent variables
Higher statistical power/higher precision of effect estimates Lower statistical power/lower precision of effect estimates

Although still subject to unmeasured time-varying confounding, these within-group estimators may provide stronger evidence of causality than the standard approach of ignoring temporality and simply pooling across survey waves. However, temporal analysis of diabetes is complicated by ambiguity regarding etiologic period – existing research has supported a view in which long-term cumulative exposure over the lifecourse, beginning with fetal development and persisting into adulthood, matters for developing diabetes (Best et al., 2005; Kanaka--Gantenbein, 2010; Maty et al., 2010). This leads to a lack of clarity about the best approach constructing lags or moving averages to account for latency between community exposures and measured diabetes status.

In this study, we employed a REWB model to analyze ten years of nationally representative, repeated cross-sectional survey data and compare it to a more traditional multilevel model that pools exposures across survey waves. Our study’s primary aims were to (1) estimate the associations between four county-level diabetes exposures and individual-level diabetes, and (2) to compare results from a popular analytic approach (a standard multilevel model that pools across time periods) to those from an alternative approach (a REWB model) that provides arguably stronger evidence that a hypothetical intervention that changes a county’s environment would, or would not, affect diabetes risk. We hypothesized that, consistent with prior research, the odds of diabetes within a county would be positively associated with three county-level exposures (unemployment rates, fast food restaurant availability, and convenience store availability) and inversely associated with active commuting rates. We hypothesized that the directionality of the associations to be the same for both the cross-sectional analysis (county characteristics associated with odds of diabetes) and for the temporal analysis (change in county factors associated with change in odds of diabetes).

2. Methods

We obtained data from the 2003–2012 Behavioral Risk Factor Surveillance System (BRFSS), an ongoing, nationally representative, cross-sectional telephone survey of the US noninstitutionalized civilian population ages ≥ 18 years (CDC, 2018; Mokdad, 2009). The BRFSS is co-ordinated by the Centers for Disease Control and Prevention and administered by states; participants are selected via random-digit dial. For 2012 and all years prior, identifiers indicating respondents’ counties of residence were publicly available for counties with ≥50 respondents in a given year. Because county sample sizes vary at random between years, a county represented in the public-use data in one wave may not appear in the next. We created a multilevel dataset by merging county-level variables pertaining to the economic, food, and built environment to individual survey participants.

2.1. Diabetes and individual-level covariates

We classified respondents as having diabetes if they reported having ever been told by a health care provider that they had diabetes. The BRFSS does not differentiate between subtypes of the disease, however >90% of adults with diabetes in the US have type 2 diabetes (Imperatore et al., 2018). Consistent with CDC analyses, we excluded respondents reporting pregnancy-related diabetes only. Additionally, we obtained data on potential confounders including respondent age (years), gender (male or female, with no data available about whether respondents were transgender), race/ethnicity (Hispanic; or non-Hispanic white, black, Asian/Pacific Islander, American Indian/Alaska Native, or multiracial), marital status (currently married or unmarried), income (continuous, in 2012 dollars), education (<high school, high school, some college, college or more), number of adults residing in the household, any health insurance coverage (yes or no), and smoking history (current, never, or ever smoker). To account for the inclusion of cell phones in the BRFSS sampling frame starting in 2011, we included data on whether the interview took place via cell phone or landline. We did not consider body mass index as a confounder because adiposity is hypothesized mediator through which our exposures of interest can cause diabetes (Chadt et al., 2000).

2.2. County-level exposures

Given prior research on the life course epidemiology of diabetes, we posited that its etiologic process involved long-term, cumulative exposure to environmental risk factors (Ludwig et al., 2011; Smith et al., 2011; Stringhini et al., 2013; White et al., 2016). Accordingly, we calculated county-level exposure variables based on the mean value for the 5 years preceding the BRFSS interview (e.g. a person whose diabetes status was measured in 2003 were assigned exposure values based on the mean of 1998–2002). A 5-year exposure period is shorter than those used for experimental and quasi-experimental research on the impact of community exposures on diabetes risk (Ludwig et al., 2011; White et al., 2016) but data on all exposures of interest were unavailable for earlier periods. Additionally, since BRFSS provides no data on residential histories, there is a tradeoff between longer exposure periods and accurately characterizing participants’ exposure periods (e.g., they were less likely to have lived in the county 10 years ago versus 5 years ago).

As an indicator of county economic context, we used mean annual unemployment rate from the US Bureau of Labor Statistics, the same data used to calculate the official unemployment rate based on the Current Population Survey, a monthly survey of approximately 60,000 US households (Bureau of Labor Statistics, 2018). Unlike other area economic measures (e.g. the poverty rate), unemployment data were available annually throughout the entire study period.

We used two measures related to restaurants and retail food stores to characterize county food environments. Prior research suggests that local retail environments in which there is a preponderance of unhealthy food retailers specializing in selling low-cost, energy-dense products (i.e. ‘food swamps’) may be more responsible for encouraging unhealthy dietary composition than areas in which there are simply few grocery stores (‘food deserts’) (Bridle-Fitzpatrick, 2015; Cooksey-Stowers et al., 2017; Mezuk et al., 2016). We chose to use existing food environment measures that quantify the relative balance of unhealthy and healthy food options and differentiate between food stores (e.g. grocery and convenience stores) and restaurants (Rummo et al., 2017, 2015), as both components of the retail food environment would have different policy implications for intervention. We obtained food environment data from County Business Patterns, a US Census Bureau series that provides annual tabulations on the number of businesses in counties (United States Census Bureau, 2018). Since 1998, the businesses have been categorized according to the North American Industry Classification System (NAICS).

Following prior research, we calculated a convenience store availability measure as proportion of the total Food and Beverage Stores category (NAICS: 445) that were Convenience Stores (NAICS: 44512; convenience stores are establishments offering a limited selection of pre-packaged foods and few if any fresh foods). We calculated our fast food availability measure as the proportion of total businesses in the Restaurants categories (NAICS: 7221 and 7222) that were Limited Service Restaurants (NAICS: 7222). Limited service restaurants are ones in which patrons typically pay at a counter before receiving their food, including ‘fast food’ restaurants.

We measured rates of active commuting in each county by calculating the proportion of total workers age ≥16 years who commute to work by walking, cycling, or using public transportation using decennial census data (1990 and 2000) and the American Community Survey (2009–2013 5-year estimates, whose mean is centered around 2011). Employees who are active commuters are defined as those who travel to work using forms of transportation that involve physical activity (Besser and Dannenberg, 2005). Active commuting rates are also a proxy for the walkability of the area in general, as they are more closely associated with various aspects of the built environment than with area socio-demographic composition (Fan et al., 2014; Glazier et al., 2014). We used linear interpolation to estimate active commuting for intercensal years.

2.3. Statistical analyses

We first calculated descriptive tabulations of the study sample along with estimates of age-adjusted diabetes prevalence across values of person-level covariates. To describe the distribution of county-level exposures and their longitudinal change, we used modified Bland-Altman plots: we created scatterplots for each measure with the mean for 1998–2002 on the x-axis and the mean for 2007–2011 on the y-axis.

We then employed a REWB approach, fitting a logistic mixed-effects model to estimate associations between the community exposures at the county level and odds of prevalent diabetes at the individual level. We standardized the county-level exposures and then decomposed each into two covariates: (1) the overall county mean value over the entire study period, whose coefficient is interpreted as the between-county, cross-sectional association, and (2) the county-mean-centered annual value (county-year value minus county mean); this coefficient is interpreted as the within-county, longitudinal association. We also included a linear predictor to control for national time trends in the exposure and outcome.

We estimated the model using Laplace approximation with the glmer function of the lme4 package in R (Bates et al., 2018) using a logit link function. Our model included three levels of nested random intercepts for county-years, within counties, within states. We treated census region (Northeast, Midwest, South, West) as an indicator variable. We controlled for the individual-level covariates described above both to limit confounding and as an alternative to non-response weights (BRFSS provides national-level weights, but they are not valid for county-level inferences). We controlled for all the covariates used by BRFSS in 2012 to construct their non-response weights and included interactions when cross-tabulated weights were used (age by gender, race/ethnicity, education, marital status, gender by race/ethnicity, age by race/ethnicity, region, region by age, region by gender, region by race/ethnicity), with the exception of household tenure (rental vs. home owner, as this variable was not collected for the full sample prior to 2011) (CDC, 2013). Additionally, we included a linear predictor for year.

For methodologic comparison, we also estimated a model that pooled across survey years. The specification was identical to that of the REWB model, with three levels of random intercepts (county-year, county, state), except there was only a single county-year-level variable for each exposure (the standardized value prior to being decomposed into separate within and between coefficients). The pooled model is the default approach in analyses of repeated cross-sectional survey data.

3. Results

Our study sample included data on the 3,397,124 adult BRFSS respondents with an identifiable county of residence surveyed over 2003–2012. Of the 3141 counties in the United States, 2392 were identified in the dataset for at least one year. There were 893 counties that appeared in the dataset for all ten years; these counties together represent >80% of the total US population.

Age-adjusted diabetes prevalence for the total sample across the study period was 7.90% (95% CI: 7.87%, 7.93%; Table 2). The burden of diabetes was higher among men (vs. women), respondents aged ≥55 years (vs. those ages 18–54), and all categories for people of color (vs. non-Hispanic white). Prevalence decreased with increasing socioeconomic position, whether measured by household income or educational attainment.

Table 2.

Characteristics of the behavioral risk factor surveillance system (BRFSS) sample and temporal changes in diabetes (DM) prevalence for adults age ≥ 18 years – United States, 2003–2012.

Sample characteristics
DM prevalence (2003–2012)a
N % % 95% CI
Total sample 3,397,124 100.0% 7.90% 7.87, 7.93
Demographic
Age
 18-24 141,003 4.2% 1.05% 1.04, 1.05
 25-34 349,428 10.3% 2.04% 2.02, 2.06
 35-44 521,997 15.4% 4.28% 4.26, 4.30
 45-54 680,452 20.0% 8.47% 8.44, 8.50
 55-64 708,820 20.9% 14.90% 14.86, 14.94
 65-74 543,535 16.0% 19.43% 19.39, 19.47
 75-84 351,725 10.4% 18.75% 18.71, 18.79
 ≥85 100,161 3.0% 13.61% 13.57, 13.65
  (missing)b 3 0.0% – –
Gender
 Female 2,079,996 61.2% 8.47% 8.43, 8.52
 Male 1,317,128 38.8% 7.57% 7.54, 7.61
  (missing) – –
Race/ethnicity
 Non-Hispanic White 2,689,777 80.0% 6.91% 6.88, 6.93
 Non-Hispanic Black 278,516 8.3% 13.96% 13.84, 14.09
 Hispanic 213,049 6.3% 11.82% 11.68, 11.97
 Non-Hispanic Asian 38,666 1.2% 14.19% 13.83, 14.55
 Non-Hispanic Alaskan Native/Pacific Islander 65,299 1.9% 8.01% 7.80, 8.21
  Multiracial 78,709 2.3% 10.16% 9.95, 10.37
  (missing) 33,108 1.0% – –
Socio-economic
Education
 Less than high school 315,978 9.3% 13.06% 12.94, 13.19
 High school 1,000,571 29.5% 9.00% 8.95, 9.07
 Some college 907,480 26.8% 8.19% 8.14, 8.25
 College/graduate school 1,166,940 34.4% 5.49% 5.45, 5.53
Household income
 <$15,000 754,188 22.2% 10.81% 10.73, 10.88
 $15,000 to <$25,000 421,232 12.4% 11.40% 11.30, 11.51
 $25,000 to <$35,000 395,942 11.7% 8.99% 8.90, 9.08
 $35,000 to <$50,000 443,031 13.0% 7.54% 7.46, 7.61
 $50,000 or more 1,382,731 40.7% 5.48% 5.45, 5.52
Marital status
 Married 1,853,621 54.7% 7.10% 7.06, 7.14
 Not married 1,533,361 45.3% 9.18% 9.13, 9.22
Health and Health Behaviors
Smoking
 Current 599,045 17.7% 7.78% 7.71, 7.85
 Ever 993,990 29.4% 8.56% 8.50, 8.61
 Never 1,784,057 52.8% 7.40% 7.36, 7.43
Health insurance
 Insured 3,011,713 88.9% 7.84% 7.81, 7.87
 Uninsured 377,290 11.1% 8.48% 8.35, 8.60
BMI
 Under or normal weight 1,337,796 40.0% 4.04% 4.01, 4.07
 Overweight 1,185,243 34.9% 6.49% 6.45, 6.53
 Obese 874,085 25.7% 15.45% 15.37, 15.53
Survey design
Number of adults in household
 1-2 adults 2,842,245 87.2% 7.74% 7.71, 7.77
 3+ adults 418,064 12.8% 9.41% 9.31, 9.52
Sample selection
 Landline 3,119,599 92.0% 7.84% 7.81, 7.87
 Cell phone 272,565 8.0% 8.66% 8.55, 8.76
a

DM prevalence is weighted to the 2000 US Standard Population age distribution, except for prevalence by age group. Excludes respondents (N = 4670, 0.14%) with missing diabetes status.

b

The percent distributions are based on the observed (non-missing) values, and the percent missing is based on the total number of BRFSS respondents.

Relative to the distribution of their means over the study period, county rates of active commuting and fast food availability changed little between the first five and last five years that we analyzed (Fig. 1). There was a greater degree of longitudinal change for unemployment (which rose over the study period in 94% of counties) and convenience store availability (for which the number of counties experiencing increases and decreases were approximately equal). Only a small number of counties, most with large populations, had rates of active commuting that exceeded 10%. Although two measures quantified aspects of the food environment, the correlation between mean county fast food availability and convenience store availability was relatively low (r = 0.27).

Fig. 1.

Fig. 1.

Mean and longitudinal change for county-level environment variables.

Coefficients from the pooled model (Table 3) showed evidence of positive associations between odds of diabetes and county unemployment, fast food availability, and convenience store availability; and an inverse association for rates of active commuting. Odds ratios (ORs) ranged from 0.95 for a one standard deviation (SD) increment in active commuting, meaning a one SD increase was associated with a 5% reduction in odds of prevalent diabetes (OR 95% CI: 0.94, 0.96) to 1.05 (i.e. a 5% increase in odds of prevalent diabetes; OR 95% CI: 1.04, 1.06) for a one standard deviation increment in convenience store availability.

Table 3.

Odds ratios and 95% confidence intervals from logistic multilevel models comparing associations derived from pooled and random-effects within-between models (Outcome: self-report individual-level prevalent diabetes).

County Measures
Pooled county-years model Random effects within-between model
(Continuous; 1 unit = 1 standard deviation) Within-County Between-County
Active commuting 0.95 (0.94, 0.96) 0.92 (0.82, 1.02) 0.95 (0.94, 0.96)
Convenience store density 1.05 (1.04, 1.06) 1.01 (1.00, 1.02) 1.05 (1.05, 1.07)
Fast food density 1.01 (1.01, 1.02) 1.00 (0.99, 1.02) 1.01 (1.00, 1.02)
Unemployment 1.02 (1.01, 1.03) –c 1.06 (1.05, 1.07)
c

The coefficient for within-county unemployment not estimable due to its collinearity with the coefficient for year.

In the REWB model, all within-county associations were estimated with lower precision than between-county associations. Compared to between-county associations, within-county effect sizes were attenuated for convenience stores and fast food restaurants, but the point estimate for within-county change in active commuting (OR: 0.92; 95% CI: 0.82, 1.02; with a point estimate suggesting an 8% within-county decrease in odds of prevalent diabetes) was more extreme. The within-county association for unemployment was not estimable due to collinearity between this variable and the linear effect of time, as assessed by the variance inflation factor, and the same issue arose when we substituted county poverty as an alternative measure. In both cases, county economic deprivation showed a strong, national increasing trend and it was not possible to estimate the within-county effects of economic deprivation independent of the linear effect of time.

4. Discussion

Our hypotheses regarding associations between diabetes and county-level exposures were supported by a multilevel model that pooled across county-years and by the cross-sectional, between-county estimates from a REWB model. However, the within-county longitudinal associations estimated by the REWB model, which are less vulnerable to time-invariant confounding, provided more ambiguous results. Temporal increases in county active commuting were associated with decreasing odds of diabetes, but precision for this estimate was low. While change in the food environment variables predicted change in odds of diabetes in the expected direction, the effect sizes were very small and confidence intervals were wide. Lower precision for the within-county estimates is to be expected, as statistical power is generally lower for within-unit estimators. It is not clear why precision was particularly low for the within-county active commuting coefficient, but it may reflect the fact that few counties experienced meaningful levels of change in active commuting rates over the study period. The association of changes in unemployment, independent of a national time trend, was inestimable due to the nearly ubiquitous growth in county unemployment over the study period (the last four years of which correspond to the Great Recession) and this problem also arose for county poverty when attempted as an alternative measure. Results from the above analysis, which measured exposure at only one geographic level (county) and assumed no spatial variability in the effect of the exposures, suggest that the estimated associations for the food environment measures may have resulted from time-invariant confounding.

To our knowledge, there have been no prior comparisons of regression coefficient for pooled versus REWB multilevel models. Our finding of an inverse cross-sectional (between-county) association of county unemployment with odds of diabetes, even after adjusting for individual socioeconomic position, is consistent with existing research on the contextual effects of economic deprivation and diabetes (Ludwig et al., 2011; Sundquist et al., 2015; White et al., 2016). Prior cohort studies using relative food environment measures calculated at smaller spatial scales have identified positive associations between the proportion of food outlets that are unhealthy and risk of diabetes; these studies used similar definitions of ‘healthy’ and ‘unhealthy’ food establishments as our analysis, but incorporated spatial buffers ranging from 720 m to 1 km (Mezuk et al., 2016; Paquet et al., 2014; Polsky et al., 2016). Cohort studies of neighborhood walkability have found inverse associations with diabetes, and these also measured walkability smaller scales, such as 1 km buffers from individual homes or within Canadian census dissemination areas – units whose populations generally range from 400 to 700 residents (Booth et al., 2013; Creatore et al., 2016; Paquet et al., 2014).

4.1. Strengths and limitations

Our study is subject to a number of important limitations. While county was the smallest geographic identifier available in the BRFSS, this is likely not the ideal level on which to measure all of the community exposures as it misses important intra-county variability, potentially biasing associations towards the null of no effect (Krieger et al., 2017). Considering the food environment, survey data shows that the mean distance travelled to a supermarket in the United States is 6.1 km (Morrison and Mancino, 2015) whereas the distance across a typical US county is many times farther. For walkability, neighborhood is a more relevant scale, as people tend to walk in the areas near their homes (Duncan et al., 2011). Additionally, the BRFSS relies on self-report diabetes and therefore misclassifies cases that are not diagnosed or not recalled; validity of self-report diabetes varies by social group (Mendola et al., 2018). Also, since repeated cross-sectional surveys do not follow the same individuals over time, we do not have data on whether respondents were diabetes-free prior to exposure. Additionally, BRFSS provides no data on respondents’ duration of county residence and – since inter-county migration may be related to health – this raises the potential for selection bias (Findley, 1988). Finally, it is possible that there are threshold effects for certain county-level measures and the relatively limited ranges of within-county change for these measures did not exceed the thresholds. Our study is strengthened by its use of a large, nationally representative sample, its consideration of a biologically plausible etiologic period, and its use of multiple analytic approaches, each with different assumptions with regard to confounding.

5. Conclusions

Analyses of repeated cross-sectional survey data to identify associations between diseases and hypothesized geographic determinants are common; considering those that analyze BRFSS data alone, they number in the dozens. While analyses of BRFSS and similar repeated cross-sectional surveys often benefit from large, representative samples, concerns regarding internal validity arise given the cross-sectional nature of the data and the unavailability of subcounty geographic scales that may be more relevant for the exposures of interest. However, such studies often estimate statistically significant associations in the hypothesized direction between the exposure and disease of interest. It is likely that these associations are biased by unmeasured confounding, given the high degree of geographic sorting that occurs by social characteristics, including by baseline health status; people who are socially deprived in various respects are likely to be sorted into neighborhoods with multiple health hazards (Honold et al., 2012).

We have demonstrated that using REWB models for repeated cross-sectional surveys can reveal different associations for the between vs. within-group coefficients, each of which have distinct assumptions with regard to internal validity. In the analyses above, while results from a commonly employed pooled analysis supported our hypotheses, the REWB within-county estimates were more ambiguous. We found evidence of a meaningful effect size for active commuting, in which within-county growth predicted lower within-county odds of diabetes, although precision for this estimate was low. The REWB model results also revealed that the coefficients from the pooled model were primarily driven by cross-sectional, between-county associations. For researchers who are set on analyzing multiple years of geocoded, repeated cross-sectional survey data from BRFSS and similar data sources, REWB models may be a preferable option to pooled models as they offer a way to analyze change over time. REWB models allow for a comparison between cross-sectional associations and temporal associations, each of which have distinct threats to internal validity regarding confounding and etiologic period as described above and, when taken together, they may be used to help triangulate results if one is able to determine the likely magnitude and direction of bias in each (Lawlor et al., 2016). REWB models also allow the effects of exposures to vary by space and time (Bell et al., 2019) and future research on community determinants of diabetes could explore this heterogeneity.

Funding

This work was supported by the Centers for Disease Control and Prevention [Grant number 1U01DP006299-01].

References

  1. Angrist JD, Pischke J-S, 2009. Parallel worlds: fixed effects, difference-in-differences, and panel data. In: Mostly Harmless Econometrics: an Empiricist’s Companion. Princeton University Press, Princeton, NJ. 10.1057/be.2009.37. [DOI] [Google Scholar]
  2. Astell-Burt T, Feng X, 2018. Geographic variation in the impact of a type 2 diabetes diagnosis on behavioural change: a longitudinal study using random effects within-between (REWB) models. Health Place 54, 164–169. 10.1016/j.healthplace.2018.07.007. [DOI] [PubMed] [Google Scholar]
  3. Bates D, Maechler M, Bolker B, Walker S, Christensen RHB, Singmann H, et al. , 2018. Package “lme4.” R Found. Stat. Comput, Vienna, Austria: 10.18637/jss.v067.i01. [DOI] [Google Scholar]
  4. Bell A, Fairbrother M, Jones K, 2019. Fixed and random effects models: making an informed choice. Qual. Quantity 53, 1051–1074. 10.1007/s11135-018-0802-x. [DOI] [Google Scholar]
  5. Bell A, Jones K, 2015. Explaining fixed effects: random effects modeling of time-series cross-sectional and panel data. Polit. Sci. Res. Methods 3, 133–153. 10.1017/psrm.2014.7. [DOI] [Google Scholar]
  6. Besser LM, Dannenberg AL, 2005. Walking to public transit. Am. J. Prev. Med 29, 273–280. 10.1016/j.ampre.2005.06.010. [DOI] [PubMed] [Google Scholar]
  7. Best LE, Hayward MD, Hidajat MM, 2005. Life course pathways to adult-onset diabetes. Biodemogr. Soc. Biol 52, 94–111. 10.1080/19485565.2005.9989104. [DOI] [PubMed] [Google Scholar]
  8. Booth GL, Creatore MI, Moineddin R, Gozdyra P, Weyman JT, Matheson FI, Glazier RH, 2013. Unwalkable neighborhoods, poverty, and the risk of diabetes among recent immigrants to Canada compared with long-term residents. Diabetes Care 36. 10.2337/dc12-0777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Bridle-Fitzpatrick S, 2015. Food deserts or food swamps?: a mixed-methods study of local food environments in a Mexican city. Soc. Sci. Med 142, 202–213. 10.1016/j.socscimed.2015.08.010. [DOI] [PubMed] [Google Scholar]
  10. Bureau of Labor Statistics, 2018. Local area unemployment Statistics [WWW Document]. https://www.bls.gov/lau/. accessed 12.27.18.
  11. CDC, 2018. Behavioral risk factor surveillance system [WWW Document]. https://www.cdc.gov/brfss/index.html. accessed 12.27.18.
  12. CDC, 2013. 2012 behavioral risk factor surveillence system: weighting the data [WWW Document]. https://www.cdc.gov/brfss/annual_data/2012/pdf/Weighting-the-Data_webpage-content-20130709.pdf. accessed 12.27.18.
  13. Chadt A, Scherneck S, Joost H-G, Al-Hasani H, 2000. Molecular Links between Obesity and Diabetes: “Diabesity,” Endotext. MDText.com, Inc. [Google Scholar]
  14. Cooksey-Stowers K, Schwartz M, Brownell K, 2017. Food swamps predict obesity rates better than food deserts in the United States. Int. J. Environ. Res. Publ. Health 14, 1366. 10.3390/ijerph14111366. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Creatore MI, Glazier RH, Moineddin R, Fazli GS, Johns A, Gozdyra P, Matheson FI, Kaufman-Shriqui V, Rosella LC, Manuel DG, Booth GL, 2016. Association of neighborhood walkability with change in overweight, obesity, and diabetes. JAMA, J. Am. Med. Assoc 10.1001/jama.2016.5898. [DOI] [PubMed] [Google Scholar]
  16. den Braver NR, Lakerveld J, Rutters F, Schoonmade LJ, Brug J, Beulens JWJ, 2018. Built environmental characteristics and diabetes: a systematic review and meta-analysis. BMC Med. 16 10.1186/s12916-017-0997-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Dendup T, Feng X, Clingan S, Astell-Burt T, 2018. Environmental risk factors for developing type 2 diabetes mellitus: a systematic review. Int. J. Environ. Res. Publ. Health 15. 10.3390/ijerph15010078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Dieleman JL, Templin T, 2014. Random-effects, fixed-effects and the within-between specification for clustered data in observational health studies: a simulation study. PloS One 9. 10.1371/journal.pone.0110257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Duncan DT, Aldstadt J, Whalen J, Melly SJ, Gortmaker SL, 2011. Validation of Walk Score® for estimating neighborhood walkability: an analysis of four US metropolitan areas. Int. J. Environ. Res. Publ. Health 10.3390/ijerph8114160. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Fairbrother M, 2014. Two multilevel modeling techniques for analyzing comparative longitudinal survey datasets. Polit. Sci. Res. Methods 2, 119–140. 10.1017/psrm.2013.24. [DOI] [Google Scholar]
  21. Fairbrother M, 2011. Explaining social change: the application of multilevel models to repeated cross-sectional survey data. In: European Consortium for Political Research General Conference. [Google Scholar]
  22. Fan JX, Wen M, Kowaleski-Jones L, 2014. An ecological analysis of environmental correlates of active commuting in urban. U.S. Heal. Place 30, 242–250. 10.1016/j.healthplace.2014.09.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Findley SE, 1988. The directionality and age selectivity of the health-migration relation: evidence from sequences of disability and mobility in the United States. Int. Migr. Rev 10.2307/2546583. [DOI] [PubMed] [Google Scholar]
  24. Geiss L, Li Y, Kirtland K, Barker L, Burrows NR, Gregg EW, 2012. Increasing prevalence of diagnosed diabetes — United States and Puerto Rico, 1995–2010. MMWR Morb. Mortal. Wkly. Rep 61, 918–921 mm6145a4 [pii]. [PubMed] [Google Scholar]
  25. Geiss LS, Kirtland K, Lin J, Shrestha S, Thompson T, Albright A, Gregg EW, 2017. Changes in diagnosed diabetes, obesity, and physical inactivity prevalence in US counties, 2004-2012. PloS One 12, e0173428. 10.1371/journal.pone.0173428. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Geiss LS, Wang J, Cheng YJ, Thompson TJ, Barker L, Li Y, Albright AL, Gregg EW, 2014. Prevalence and incidence trends for diagnosed diabetes among adults aged 20 to 79 years, United States, 1980-2012. JAMA, J. Am. Med. Assoc 312, 1218–1226. 10.1001/jama.2014.11494. [DOI] [PubMed] [Google Scholar]
  27. Glazier RH, Creatore MI, Weyman JT, Fazli G, Matheson FI, Gozdyra P, Moineddin R, Shriqui VK, Booth GL, 2014. Density, destinations or both? A comparison of measures of walkability in relation to transportation behaviors, obesity and diabetes in Toronto, Canada. PloS One. 10.1371/journal.pone.0085295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Honold J, Beyer R, Lakes T, van der Meer E, 2012. Multiple environmental burdens and neighborhood-related health of city residents. J. Environ. Psychol 32, 305–317. 10.1016/J.JENVP.2012.05.002. [DOI] [Google Scholar]
  29. Imai K, Kim IS, 2019. When should we use fixed effects regression models for causal inference with longitudinal data?. Am. J. Polym. Sci 63, 467–490. 10.1186/s13014-015-0471-z. [DOI] [Google Scholar]
  30. Imperatore K, Bullard M, Cowie CC, Lessem SE, Saydah SH, Menke A, Geiss LS, Orchard TJ, Rolka DB, 2018. Prevalence of diagnosed diabetes in adults by diabetes type — United States, 2016. MMWR Morb. Mortal. Wkly. Rep 67, 359–361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Jia H, Moriarty DG, Kanarek N, 2009. County-level social environment determinants of health-related quality of life among us adults: a multilevel analysis. J. Community Health 35, 430–439. 10.1007/s10900-009-9173-5. [DOI] [PubMed] [Google Scholar]
  32. Kanaka-Gantenbein C, 2010. Fetal origins of adult diabetes. Ann. N. Y. Acad. Sci 1205, 99–105. 10.1111/j.1749-6632.2010.05683.x. [DOI] [PubMed] [Google Scholar]
  33. Krieger N, Feldman JM, Waterman PD, Chen JT, Coull BA, Hemenway D, 2017. Local residential segregation matters: stronger association of census tract compared to conventional city-level measures with fatal and non-fatal assaults (total and firearm related), using the index of concentration at the extremes (ICE) for racial, economic, and racialized economic segregation, Massachusetts (US), 1995–2010. J. Urban Health 94, 244–258. 10.1007/s11524-016-0116-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Lawlor DA, Tilling K, Davey Smith G, 2016. Triangulation in aetiological epidemiology. Int. J. Epidemiol 45, 1866–1886. 10.1093/ije/dyw314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Ludwig J, Sanbonmatsu L, Gennetian L, Adam E, Duncan GJ, Katz LF, Kessler RC, Kling JR, Lindau ST, Whitaker RC, McDade TW, 2011. Neighborhoods, obesity, and diabetes — a randomized social experiment. N. Engl. J. Med 365, 1509–1519. 10.1056/NEJMsa1103216. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Macintyre S, Ellaway A, 2000. Ecological approaches: rediscovering the role of the physical and social environment. In: Kawachi I, Berkman L (Eds.), Social Epidemiology. Oxford University Press, New York, pp. 332–348. [Google Scholar]
  37. Maty SC, James SA, Kaplan GA, 2010. Life-course socioeconomic position and incidence of diabetes mellitus among blacks and whites: the Alameda County Study, 1965-1999. Am. J. Publ. Health 100, 137–145. 10.2105/AJPH.2008.133892. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Mehta NK, Chang VW, 2008. Weight status and restaurant availability. A multilevel analysis. Am. J. Prev. Med. 34, 127–133. 10.1016/j.amepre.2007.09.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Mendola ND, Chen T-C, Gu Q, Eberhardt MS, Saydah S, 2018. Prevalence of total, diagnosed, and undiagnosed diabetes among adults: United States, 2013-2016 key findings data from the national health and nutrition examination survey (NHANES). NCHS Data Brief 319. 10.5194/acp-11-4725-2011. [DOI] [PubMed] [Google Scholar]
  40. Menke A, Casagrande S, Geiss L, Cowie CC, 2015. Prevalence of and trends in diabetes among adults in the United States, 1988-2012. JAMA, J. Am. Med. Assoc 10.1001/jama.2015.10029. [DOI] [PubMed] [Google Scholar]
  41. Mezuk B, Li X, Cederin K, Rice K, Sundquist J, Sundquist K, 2016. Beyond access: characteristics of the food environment and risk of diabetes. Am. J. Epidemiol 183, 1129–1137. 10.1093/aje/kwv318. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Mokdad AH, 2009. The behavioral risk factors surveillance system: past, present, and future. Annu. Rev. Publ. Health 30, 43–54. 10.1146/annurev.publhealth.031308.100226. [DOI] [PubMed] [Google Scholar]
  43. Morrison RM, Mancino L, 2015. Most U.S. Households do their main grocery shopping at supermarkets and supercenters regardless of income. Amber Waves. 10.1016/j.yofte.2009.07.003. [DOI] [Google Scholar]
  44. Paquet C, Coffee NT, Haren MT, Howard NJ, Adams RJ, Taylor AW, Daniel M, 2014. Food environment, walkability, and public open spaces are associated with incident development of cardio-metabolic risk factors in a biomedical cohort. Health Place 28, 173–176. 10.1016/j.healthplace.2014.05.001. [DOI] [PubMed] [Google Scholar]
  45. Polsky JY, Moineddin R, Glazier RH, Dunn JR, Booth GL, 2016. Relative and absolute availability of fast-food restaurants in relation to the development of diabetes: a population-based cohort study. Can. J. Public Health 10.17269/CJPH.107.5312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Popkin BM, 2015. Nutrition transition and the global diabetes epidemic. Curr. Diabetes Rep 10.1007/s11892-015-0631-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Rummo PE, Guilkey DK, Ng SW, Meyer KA, Popkin BM, Reis JP, Shikany JM, Gordon-Larsen P, 2017. Does unmeasured confounding influence associations between the retail food environment and body mass index over time? The Coronary Artery Risk Development in Young Adults (CARDIA) study. Int. J. Epidemiol 10.1093/ije/dyx070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Rummo PE, Meyer KA, Boone-Heinonen J, Jacobs DR, Kiefe CI, Lewis CE, Steffen LM, Gordon-Larsen P, 2015. Neighborhood availability of convenience stores and diet quality: findings from 20 years of follow-up in the coronary artery risk development in young adults study. Am. J. Publ. Health 10.2105/AJPH.2014.302435. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Smith BT, Lynch JW, Fox CS, Harper S, Abrahamowicz M, Almeida ND, Loucks EB, 2011. Life-course socioeconomic position and type 2 diabetes mellitus the Framingham offspring study. Am. J. Epidemiol 173, 438–447. 10.1093/aje/kwq379. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Stringhini S, da Batty GD, Bovet P, Shipley MJ, Marmot MG, Kumari M, Tabak AG, Kivimäki M, 2013. Association of lifecourse socioeconomic status with chronic inflammation and type 2 diabetes risk: the whitehall II prospective cohort study. PLoS Med. 10, 1–15. 10.1371/journal.pmed.1001479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Sundquist K, Eriksson U, Mezuk B, Ohlsson H, 2015. Neighborhood Walkability, Deprivation and Incidence of Type 2 Diabetes: A Population-Based Study on 512,061 Swedish Adults. Heal. Place 10.1016/j.healthplace.2014.10.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. United States Census Bureau, 2018. County business Patterns [WWW Document]. https://www.census.gov/programs-surveys/cbp.html. accessed 12.27.18.
  53. Verbeek M, 1996. Pseudo panels and repeated cross-sections. In: The Econometrics of Panel Data. Springer, Dordrecht, pp. 280–292. [Google Scholar]
  54. White JS, Hamad R, Li X, Basu S, Ohlsson H, Sundquist J, Sundquist K, 2016. Long-term effects of neighbourhood deprivation on diabetes risk: quasi-experimental evidence from a refugee dispersal policy in Sweden. Lancet Diabetes Endocrinol 4, 517–524. 10.1016/j.compmedimag.2015.12.001.Uncertainty. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES