Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2022 Jun 1.
Published in final edited form as: Soc Sci Med. 2021 Apr 19;278:113952. doi: 10.1016/j.socscimed.2021.113952

Type 1 Diabetes Incidence Among Youth in Utah: A Geographical Analysis

Matthew L McCullough 1, Neng Wan 2, Marcus G Pezzolesi 3,4,5, Timothy W Collins 6, Sara Elizbeth Grineski 7, Dennis Wei 8, Jose Lazaro-Guevara 9,10, Scott G Frodsham 11, James A Vanderslice 12, John R Holmen 13, Titte R Srinivas 14, Scott A Clements 15
PMCID: PMC8686266  NIHMSID: NIHMS1761668  PMID: 33933801

Abstract

Type 1 Diabetes (T1D) poses an increasing threat to public health, as incidence rates continue to rise globally. However, the etiology of T1D is still poorly understood, especially from the perspective of geography. The objective of this research is to examine the incidence of T1D among youth and to identify high-risk clusters and their association with socio-demographic and geographic variables. The study area was the entire state of Utah and included youth with T1D from birth to 19 years of age from 1998 to 2015 (n=4,161). Spatial clustering was measured both globally and locally using the Moran’s I statistic and spatial scan statistic. Ordinary least squares (OLS) regression was used to measure the association of high-risk clusters with certain risk factors at the Census Block Group (CBG) level. The mean age at diagnosis was 9.3 years old. The mean incidence rate was 25.67 per 100,000 person-years (95% CI, 24.57 – 26.75). The incidence rate increased by 14%, from 23.94 per100,000 person-years in 1998 to 27.98 per 100,000 person-years in 2015, with an annual increase of 0.80%. The results of the spatial scan statistic found 42 high-risk clusters throughout the state. OLS regression analysis found a significant association with median household income, population density, and latitude. This study provides evidence that incidence rates of T1D are increasing annually in the state of Utah and that significant geographic high-risk clusters are associated with socio-demographic and geographic factors.

Keywords: Type 1 Diabetes, diabetes epidemiology, spatial analysis, spatial scan statistic, Utah, geography

Introduction

Type 1 Diabetes (T1D) is a unique and complex autoimmune disease that is heavily influenced by non-genetic factors such as socioeconomic status and environmental triggers (Knipp et al., 2005). T1D is differentiated from Type 2 Diabetes (T2D) in that it is not caused primarily by dietary or lifestyle behaviors. In individuals with T1D, the body does not produce insulin and cannot, therefore, break down the sugars and starches consumed (Rønningen, 2015). A great deal of research has been done on the clinical aspects of T1D and the complications that result from it. Still, much less has aimed to identify the socio-demographic and environmental triggers associated with the spatial pattern of T1D incidence, as we do here. This research is unique in that it leveraged individual-level medical records from the Utah Population Database and analyzed the spatial patterns of T1D at a very fine geographic scale, the Census Block Group (CBG) level, in the state of Utah. One of the primary limitations of others papers has been the use of much larger scale geographic areas and the aggregation of data, which tends to hide spatial patterns (Ball et al., 2014). The primary contribution of this research was to identify hotspots or clusters at the CBG level, where the risk of developing T1D is significantly higher than the general population, and to clarify certain predictors of those hotspots.

T1D most often affects youth between the ages of 0–18. The highest rates of T1D occur between the ages of 5–7 and at or near puberty (Harjutsalo et al., 2010). T1D has been termed juvenile-onset diabetes, but it can occur at any age. The rate of diagnosed T1D has risen substantially over the past 30 years (Patterson et al., 2009). The incidence trend worldwide is increasing at a rate of 3% annually (Moltchanova et al., 2009). Finland has the highest incidence rate of T1D in the world, with 60 cases per 100,000 people each year, whereas China has a very low incidence rate of 1.01 per 100,000 (Weng et al., 2018). The incidence rate of T1D among U.S. non-Hispanic white youth is 23.6 cases per 100,000 person-years (Bell et al., 2009). In Colorado, a state with a similar elevation as our study site Utah, the incidence rate reported in 2002–2004 was 23.9 per 100,000 person-years (Vehik et al., 2007).

The cause of T1D is unknown but is believed to be a combination of genetic and environmental factors (Atkinson and Eisenbarth, 2001). Astonishingly, 85% of new cases of T1D have no family history of the disease (Hamalainen and Knip, 2002).

T1D and T2D have some common characteristics such as high blood glucose levels that can have life-threatening complications, but they also differ in many ways. In T1D, the body’s immune system is responsible for attacking the beta cells of the pancreas and destroying them so that they no longer produce insulin to moderate blood sugar levels (Van Belle et al., 2011). In T2D, the body continues to produce insulin but does not use it properly (termed insulin resistance). Because of the body’s inability to produce insulin, T1D patients must receive insulin injections multiple times per day for the remainder of their life following clinical diagnosis.

A number of dietary and lifestyle factors have been considered to be triggers of T1D, including the early consumption of gluten and cow’s milk (Knip et al., 2010); vitamin D deficiency, as a result of higher latitude and reduced exposure to ultraviolet radiation (Cooper et al., 2011); viral infections (Yeung et al., 2011); and excessive hygiene (Bach and Chatenoud, 2012). Research has shown that the incidence of T1D is positively related to distance north of the equator (Karvonen et al., 2000).

Seasonal and temporal changes have also been found to influence the incidence of T1D, with more cases being diagnosed in the fall and winter months (Moltchanova et al., 2009). Seasonal variations in the diagnosis of T1D have been linked to seasonal variations of certain viral infections, and have been documented as far back as 1926 (Adams, 1926). These seasonal changes, along with the fact that very few cases have any family history of the disease, lead to the conclusion that certain unknown environmental factors are a crucial contributor to the disease process (Atkinson et al., 2014). This paper contributes to the body of research done on T1D by seeking to identify these unknown environmental factors that contribute to the incidence of T1D at the CBG level, with a focus on socio-demographic and geographic risk factors.

Socioeconomic status (SES) plays a vital role in regard to health behaviors and health outcomes. Many causes of mortality and morbidity are influenced by one’s income, education, and occupation. Research has shown that the association between SES and health outcomes typically shows a gradient pattern with improved health and decreased mortality with each increase in occupational grade (Adler and Ostrove, 1999). SES is often treated as a potential confounder between other variables and health outcomes (Braveman et al., 2005). Variations in SES across geographies often lead to differences in lifestyle, maternal age, diet, and exposure to infections, all of which may play a role in similar patterns of incidence of T1D among populations (Patterson and Waugh, 1992). Comparisons of the impact of SES on T1D are difficult, due to the fact that different definitions of SES are used as well as at different geographic scales. Little research has been done on the association between T1D incidence and SES at a fine spatial scale. To the best of our knowledge, only one study (Haynes et al. 2006) has assessed the relationship between T1D and SES for small area units (i.e., Australian census collection districts with approximately 250 dwelling per unit). The current study seeks to build on this knowledge gap by examining the relationship of T1D incidence and SES at the CBG level in Utah.

The primary purposes of this study are to calculate incidence rates of T1D among youth in Utah, examine the spatial pattern of cases, and measure the socio-demographic and geographic associations with the spatial patterns.

Materials and Methods

Disease mapping is an effective approach to understanding the etiological origination of disease (Wakefield, 2007). In order to develop accurate maps of the geographic variation in incidence rates of T1D, we first present the study setting and data sources used in our analysis, followed by an introduction to the spatial methods used to measure both global spatial autocorrelation and local spatial clustering. Finally, we use OLS regression to measure the association between T1D and socio-demographic and geographic risk factors.

Study Setting

This study covers the entire geographic area of the state of Utah, which had a population of 3,161,105 in 2018. The majority of the population (approximately 80%) reside in four urban counties along the Wasatch Front (Weber, Davis, Salt Lake, and Utah). These four urban counties cover roughly 10% of the land area of the state, which totals approximately 84,000 square miles. The other 20% of the population is spread out among 25 rural counties.

Data Sources

The number of new cases of youth diagnosed with T1D from 1998 to 2015 was collected from medical records from Intermountain Healthcare and the University of Utah Healthcare (Frodsham et al., 2019). These two healthcare systems serve the majority of newly diagnosed patients with T1D in the state and are the major source for pediatric endocrinology specialists. The data for this study was provided by the Utah Population Database (UPDB) located at the University of Utah (https://uofuhealth.utah.edu/huntsman/utah-population-database/). The UPDB is one of the world’s richest sources of linked population-based information for demographic, genetic, and epidemiological studies. The study was approved by the Institutional Review Board (IRB) of the University of Utah (#IRB_00096551).

For the purposes of this research, patients with T1D were identified using the following International Classification of Diseases (ICD) codes, 9th and 10th revision codes (ICD-9 and ICD-10, respectively): ICD-9 250.01, 250.03, 250.11, 250.13, 250.21, 250.23, 250.31, 250.33, 250.41, 250.43, 250.51, 250.53, 250.41, 250.43, 250.51, 250.53, 250.61, 250.63, 250.71, 250.73, 250.81, 250.83, 250.91, and 250.93 and ICD-10 E10. In total, a cohort of 4,161 patients with T1D from birth to 19 years of age between 1998 and 2015 was identified for the statistical analysis.

When patients in the database had ICD diagnosis codes for both T1D and T2D we determined individuals to have T1D if two-thirds of the codes indicated that they had a diagnosis of T1D, and they were treated with insulin within one year of their first ICD code for T1D.

Socioeconomic factors for each CBG were were compiled from tables from the 2010–2014 ACS 5-year estimates and included population density, number of housing units, median household income, percent of population with no health insurance, and the percent of the population over age 25 with a bachelor’s degree. Our intent in choosing these specific factors was to include key categories of socioeconomic information such as population, housing, education, income, and health insurance.

This study used CBG identifiers based on the residential address of T1D patients provided by the UPDB to perform a join with polygon data from the Census TIGER/Line Shapefiles of Block Groups. There are 1,690 CBGs in the state of Utah, and each contains approximately 600 to 3,000 people. The geographic size of CBGs ranges from 0.04 sq. miles in urban areas to 5,467 sq. miles in rural areas. The CBG is the smallest Census area unit that provides socio-demographic data, which makes it useful for locating small and irregularly shaped clusters without the need for individual-level address information.

Spatial Methods

T1D incidence rates were mapped at the CBG level in order to perform spatial cluster analysis to detect high and low-risk clusters. Incidence rates were averaged for each year of data and spatial analysis was performed on the aggregate rate. We used GeoDa software (Anselin, 2006) to calculate the global Moran’s I statistic and test for spatial autocorrelation on the residuals of the OLS regression. We generated a spatial weights matrix file with Row standardization and contiguity edges only as the conceptualization method of spatial relationships, in order to account for spatial dependence. We also used the Spatial Scan Statistic (SaTScan) to identify local clusters (Ozdenerol et al., 2005). SaTScan settings were purely spatial, scanning for clusters with high and low rates, using the discrete Poisson model with a maximum spatial cluster size of 50% of the population at risk and a circular window shape. SaTScan results were compared with results from a local Moran’s I test to verify consistency of different spatial methods.

A spatial scan statistic is a useful tool used in disease surveillance for spatial cluster detection (Glaz et al., 1999). Numerous studies have used the spatial scan statistic to study a diverse array of rare health outcomes, including lateral sclerosis, breast cancer, congenital anomalies, giardiasis intestinal parasite, listeriosis, hospital emergency visits, and West Nile Virus (Kulldorff et al., 2006). Cluster areas are detected by a moving window of user-defined shape and size across a map and calculating the likelihood of clusters by comparing the observed and expected number of cases both within and outside of the area covered by the window. The most common shape used in the analysis is a circle, but it is also possible to use an elliptical window, which may have higher power if the cluster shape is non-circular (Kulldorff, 1997). Results can be used to create isoplethic maps that show the estimated probability of clusters being highly significant (Ozdenerol et al., 2005). One of the benefits of using this method is that it tests the null hypothesis against an alternative hypothesis that there is a higher rate of a health outcome within the window compared to outside the window. It can be used for aggregated data such as Census Tract or ZIP codes as well as point level data. It has been recommended that the use of a Poisson model is effective when analyzing spatial clusters on polygon data. The spatial scan statistic can also be effectively used to detect seasonal patterns and trends in spatial data (Azage et al., 2015).

This method of identifying spatial clusters relies on a null hypothesis of risk being completely random throughout the study area. As a tool used for probability mapping, it produces results with an associated p-value (α = 0.05), which is used to determine the statistical significance of the risk within the window as well as without the window. Significantly high and low areas can then be mapped inside of GIS software to visualize the clusters at the local scale.

Population-level Incidence

Incidence rates were calculated annually by dividing the number of newly diagnosed youth from age 0 to 19 by the total number of youth in the population (person years). When it was not already recorded, the date of diagnosis was determined by a formula including the first date seen by an endocrinologist, the first prescription of insulin, and the first date of a matching ICD code in the medical records. The overall incidence rate for the study period was calculated by dividing the total number of youth with T1D in the study period by the total number of person-years for youth age 0 to 19 during the same period. The numerator included the number of youth age 0 to 19 who were newly diagnosed with T1D each year. The denominator included the total number of youth in the population age 0 to 19 in the state. Denominator data for years 1998–1999 were from the Utah Governor’s Office of Planning and Budget (GOPB). For years 2000 and later, the population estimates were provided by the National Center for Health Statistics (NCHS) through a collaborative agreement with the U.S. Census Bureau. The denominator data were compared to the 2000 and 2010 U.S. census data for accuracy, and in both cases, the difference was less than 0.45% of the total population.

Ordinary Least Squares (OLS) Regression

We relied on an OLS regression model (α = 0.05) for this study with the aggregate incidence rate per CBG as the outcome variable and various CBG variables as covariates. OLS is a linear regression model that minimizes the sum of squared difference between observed and predicted values. OLS regression provides a Koenker Statistic to assess the stationarity of the model coefficient. A spatial autoregressive model was not used because the global Moran’s I test on OLS residuals was not significant (p = 0.371). Latitude and elevation were modeled using the centroid of each CBG. Elevation was calculated from the USGS 30-meter 3DEP product. STATA/MP 13.0 was used for regression analysis. Standardized variables included population density and the number of housing units per CBG. Median household income was measured per $10,000 U.S. dollars annually. Elevation was measured per 1,000 ft above sea level and latitude was measured per degree north of the equator. Variables not standardized included the percent of the population with no insurance and the percent of the population age 25+ with a bachelor’s degree.

Results

Incidence rates

In the state of Utah between 1998 and 2015, there were 4,161 new cases of T1D among youth age 0 to 19 from a population of more than 16 million person-years. The mean incidence rate was 25.67 per100,000 person-years (95% CI, 24.57 – 26.75). The overall increase in the incidence rate from 1998 to 2015 was 14%, and the average annual increase was 0.80%. Table 1 lists the descriptive statistics of those with T1D (n=4,161) by age, sex, race, and ethnicity. A higher incidence was noted in males than in females. Youth in the 10 to 14 years age group had the highest incidence rate compared to the other age groups. In terms of race/ethnicity, the incidence rate was highest for individuals who reported Non-Hispanic White (NHW) ethnicity followed by Non-Hispanic Black, Hispanic or Latino, American Indian and Alaska Native, and Asian/Native Hawaiian or Pacific Islander.

Table 1 –

Incidence of T1D among youth from 1998–2015 (n=4,161)

Variable n Rate per 100,000 person years

Age
0–4 807 18.68
5–9 1,281 31.59
10–14 1,503 39.06
15–19 570 14.70
Sex
Male 2,242 27.23
Female 1,919 24.39
Race/Ethnicity
Hispanic or Latino 367 14.64
Non-Hispanic White 3,696 28.88
Non-Hispanic Black 43 19.07
Asian/Native Hawaiian or Pacific Islander 24 5.57
American Indian/Alaska Native 31 11.14

The annual incidence rate of T1D among youth in Utah has grown steadily from 1998 to 2015, as shown in Figure 1. This trend is slightly lower than in a comparable study of youth with T1D in Colorado (Vehik et al., 2007).

Figure 1 -.

Figure 1 -

Incidence of Type 1 Diabetes in Utah by year

A strong seasonal variation in incidence was noted in the fall and winter months (i.e., December – February) compared to the summer (i.e., May – July), which is consistent with the results of other research (Knipp et al. 2005). Figure 2 shows the average percent of new cases by month throughout the study period.

Figure 2 -.

Figure 2 -

Seasonal variation in the percentage of new cases by month

Spatial Scan Results

There are 1,690 CBGs in the state of Utah. The number of cases per CBG ranged from 1 to 81, with an average of 3.6 cases per CBG. The total number of CBGs with at least one case that joined with the CBG geography was 1,406 (83%).

The global Moran’s I statistic for the incidence of T1D in each CBG statewide was not significant (p = 0.124) with an index of −0.007. This means that the distribution of T1D is less spatially clustered than would be expected if underlying spatial processes were not random.

The results of the spatial scan statistic found 42 significant clusters with a p-value of less than 0.05. Relative risks ranged from 1.8 to 36.2, with a mean of 7.8. The results found that 37 of the 42 clusters consisted of a single CBG, meaning that clusters are isolated to small geographic areas. The spatial scan statistic is effective at leveraging the Poisson model when analyzing spatial clusters on polygon data and controlling for population and small numbers. Correlations are likely to be reliable, given the strength, direction, and significance of the results. Results of the spatial scan statistic were compared with the results of a local Moran’s test to verify consistency. The local Moran’s test found 37 of the 42 clusters identified by the spatial scan statistic.

The results of the spatial scan are shown in Figure 3. Only CBGs with significant risk ratios are shown to make it easier to read. The 42 significant clusters are located mainly along the Wasatch Front with a few outliers in Carbon, Washington, Iron, and Summit Counties. Figure 4 shows the results of the scan statistic along the Wasatch Front in greater detail.

Figure 3 -.

Figure 3 -

Results of spatial scan statistic cluster analysis

Figure 4 -.

Figure 4 -

Results of spatial scan statistic cluster analysis for the Wasatch Front

Statistical results

The results of the OLS regression are shown in Table 2. The Pseudo R2 for the model was 1.23%. Regression analysis was statistically significant for median household income, the number of housing units, and latitude when controlling for all other variables in the model. T1D incidence increased an average of 3.3% for each degree of latitude north of the equator. Both population density and annual median household income (per $10,000 U.S. dollars) had a negative association with incidence of T1D. The incidence of T1D decreased with increasing population density and decreased 5.3% per $10,000 increase in annual income.

Table 2 –

Linear regression result of predictor variables

Predictor Coefficient Standard Error Confidence Interval
Population density (standardized) −3.734** 1.207 (−6.102, −1.366)
Number of housing units (standardized) 0.101 1.120 (−2.096, 2.295)
Median household income (per $10,000) −1.37** 0.479 (2.31, −0.43)
Percent with no health insurance −0.207 0.143 (−0.488, 0.075)
Percent with Bachelor’s degree (age 25+) −0.024 0.204 (−0.424, 0.376)
Elevation (per 1,000 ft) 0.436 0.001 (−2.042, 2.913)
Latitude (per degree) 0.859* 1.142 (0.156, 4.884)
*

p < 0.05

**

p < 0.01

Discussion

This study used individual-level medical records data to determine incidence rates of T1D in the state of Utah. It also used spatial methods to identify high-risk clusters of T1D at the CBG level, and built upon the results of previous studies that measured the association of T1D incidence rates with risk factors at a larger geographic scale. Although we did not find a global clustering trend of T1D in the state, we found 42 local clusters with an elevated risk of T1D. The mean incidence rate in Utah (25.67 per 100,000 person-years) is slightly higher than the mean rate for South Carolina, Ohio, Colorado, California, and Washington (23.6 per 100,000 person-years), all of which were part of the SEARCH for Diabetes in Youth study (Bell et al., 2009).

The state of Utah has 29 counties, only five of which are considered urban areas, with a population density of more than 100 people per square mile. 36 of the 42 high-risk clusters identified (86%) were located in these five urban counties. Spatial autocorrelation results were not significant, meaning that overall, there was no detectable spatial pattern between CBGs in TID incidence. At the same time, the presence of a large number of single-CBG high-risk clusters (evident from the spatial scan analysis) suggests that clustering of TID is isolated to very small geographic areas. This finding suggests that there may be unique factors within each CBG that contribute to a higher risk of developing T1D.

Predictor variables used in this study to measure the impact of SES on incidence of T1D included median household income, the percent of the population with no health insurance, and the percent of the population (age 25+) with a bachelor’s degree. All three variables had a negative relationship with incidence of T1D. Incidence of T1D decreased 5.3% per $10,000 increase in annual income. Many studies since the 1980s have reported a greater incidence rate of T1D among individuals with high SES (Tarn et al., 1983). In a study by Haynes et al., (2006) on the effect of SES on T1D in Western Australia, there was a 9.9% increase in the incidence of T1D for each increase in the socioeconomic group. It was also shown that the incidence of T1D in the least disadvantaged group was 56% higher than in the most disadvantaged group. Similar results were found by Torres-Aviles et al. (2010), who used the Human Development Index (HDI) to suggest that children living in higher SES areas are at greater risk for T1D. A research article published by Liese et al. (2012) examined neighborhood-level risk factors for T1D in the U.S., and the results indicated that higher level of SES was associated with greater risks of T1D. In the same article, a higher percentage of minority populations, income from social security, the proportion of crowded households, and poverty were all associated with lower odds of T1D.

The hygiene hypothesis postulates that exposure to viral infections and bacteria at a young age may have a protective effect against autoimmune diseases in general. This hypothesis is consistent with the found association between higher SES status and increased risk of T1D (Boettler and Von Herrath, 2011). The incidence of infectious diseases has decreased over time due to better hygiene standards, the use of antibiotics, and vaccination programs, but at the same time, the number of autoimmune diseases has sharply increased (Bach, 2002).

In our study, the incidence of T1D decreased with increasing population density, revealing that more densely populated CBGs have a lower risk of T1D than less densely populated CBGs. Population density is generally a good predictor of urban/rural classification because CBGs tend to be much larger and less densely populated in more rural areas. It has been assumed that higher population density would increase the risk for T1D by increasing exposure to pathogens (Patterson et at., 1996), but other studies have found population density not to be a significant factor (Ball et al., 2014). In the state of Utah, 79.2% of CBGs are located in the five urban counties. Our analysis revealed that 86% of the high-risk clusters identified in this study are located in urban areas, which suggests that CBGs in urban areas with low population density are particularly at-risk. Factors to consider in future analysis could include specific land use, proximity to open spaces, and environmental exposures, among others.

Latitude north of the equator had a significant and positive association with risk of T1D in this study, when controlling for all other variables. The state of Utah spans 5 degrees of latitude (37°N to 42°N) with nearly 90% of the population living in the northern half of the state. In our analysis, the incidence of T1D increased an average of 3.3% for each degree of latitude north of the equator. In a similar study by Ball et al. (2014), the risk of T1D increased 3.5% for each degree of latitude away from the equator.

Elevation has not typically been included as a variable in regression analysis of T1D risk, but was included in this study because of the positive relationship with latitude and because accurate elevation data for each CBG was available statewide. The average elevation in the state of Utah is 6,100 feet above sea level, which is the third-highest in the United States. Our analysis indicated that T1D incidence increased 1.7% for each 1,000 feet of elevation above sea level.

While the etiology of T1D is still not completely understood, studies continue to show that the increasing incidence rates among young children are occurring too rapidly to be caused by genetic alterations alone, and are likely the result of environmental factors (Gillespie, 2006). Research has identified the most popular environmental risk factor to be viruses and that early exposure to microbes and other pathogens early in life promotes immune responses and protect against the risk of autoimmunity (Hyoty, 2002).

The results of this study found that the percent of new cases reported in the month of January (11%) were almost double those reported in the month of June (6%). Maternal intake of Vitamin D during pregnancy, as well as high doses of Vitamin D early in life, have been shown to have a protective factor against autoimmunity and T1D (Fronczak et al., 2003). Season of birth and maternal exposure to UVB has been found to significantly influence the risk of many different immune-mediated diseases such as multiple sclerosis and T1D. The risk of disease is inversely correlated with second trimester UVB exposure and third trimester vitamin D status (Disanto, 2012).

This study has a number of limitations in regard to the data and methodology. First, it relied on electronic health records from two different healthcare systems, from which ICD diagnosis codes for T1D and T2D were not consistent and required a formula to impute the correct diagnosis. In order to ensure that only patients with T1D were included in the analysis, we verified that they were treated with insulin and that two-thirds of the ICD codes were for T1D only. These two factors, coupled with the age of diagnosis of the cohort from birth to 19, helped to ensure that only patients with T1D were included.

Second, the use of aggregated data from defined geographical units as the base of any ecological model runs the risk of introducing ecological fallacy, which is the idea that relationships observed for groups of people hold true for the individuals within the group (Freedman, 1999). The fundamental problem with ecological fallacy is the loss of individual-level information due to aggregation. In certain circumstances, the ecological associations between the outcome and an explanatory variable differ from the individual-level data, and may even reverse direction (Wakefield et al., 2010). Using predefined geographical units for data collection and modeling purposes can also introduce the Modifiable Areal Unit Problem (MAUP). This problem arises because aggregated neighborhood-level data are analyzed in predefined shapes or units, which are not homogenous and cohesive real-world units. Individual point level data aggregated into predefined geographical units of different shapes and sizes can lead to opposite conclusions (Fotheringham, 1991). The CBG was the base unit of this study. Results based on these spatial units should not be assumed to be applicable to individuals within the CBGs. In addition to the spatial properties of the dataset used in this paper, this study aggregates incidence rates from a temporal perspective as well. Changes in incidence rates over time and space can introduce bias in the results. This temporal effect was identified by Cheung and Adepeju in 2014 as the Modifiable Temporal Unit Problem (MTUP). Although the design of this study does not include space-time analysis, it does provide evidence about the etiology of T1D and the influence of risk factors.

Future analysis should consider the correlation of T1D incidence with proximity to major freeways and highways as well as air pollution levels in commercial and industrial areas. The association between air pollutant emissions and T1D incidence has been observed in European countries (Di Ciaula, 2014). This study does not include environmental risk factors unique to each cluster area, such as air pollution, water quality, ultraviolet radiation (UVB), and climate variables. Future research will explore each of these risk factors.

The intersection of medical research and spatial methods is improving our understanding of complex human-environmental interactions. The increasing incidence rate of T1D among youth is due in part to socioeconomic, demographic, and geographic factors. There is no effective means at this time to prevent T1D, which should lead researchers to pursue the identification of the causal factors of the disease. Understanding the complex genetic and environmental interactions should improve efforts to prevent and cure T1D in the future.

Acknowledgements:

We thank the Pedigree and Population Resource of Huntsman Cancer Institute, University of Utah (funded in part by the Huntsman Cancer Foundation) for its role in the ongoing collection, maintenance and support of the Utah Population Database (UPDB). We also acknowledge partial support for the UPDB through grant P30 CA2014 from the National Cancer Institute, University of Utah and from the University of Utah’s program in Personalized Health and Center for Clinical and Translational Science.

Footnotes

Declaration of Interest: None

Contributor Information

Matthew L. McCullough, Department of Geography, University of Utah, Salt Lake City, UT, USA

Neng Wan, Department of Geography, University of Utah, Salt Lake City, UT, USA.

Marcus G. Pezzolesi, Diabetes and Metabolism Research Center, University of Utah School of Medicine, Salt Lake City, UT, USA Division of Nephrology and Hypertension, Department of Internal Medicine, University of Utah School of Medicine, Salt Lake City, UT, USA; Department of Human Genetics, University of Utah School of Medicine, Salt Lake City, UT, USA.

Timothy W. Collins, Department of Geography, University of Utah, Salt Lake City, UT, USA

Sara Elizbeth Grineski, Department of Sociology, University of Utah, Salt Lake City, UT, USA.

Dennis Wei, Department of Geography, University of Utah, Salt Lake City, UT, USA.

Jose Lazaro-Guevara, Division of Nephrology and Hypertension, Department of Internal Medicine, University of Utah School of Medicine, Salt Lake City, UT, USA; Department of Human Genetics, University of Utah School of Medicine, Salt Lake City, UT, USA.

Scott G. Frodsham, Division of Nephrology and Hypertension, Department of Internal Medicine, University of Utah School of Medicine, Salt Lake City, UT, USA

James A. Vanderslice, Division of Public Health, University of Utah School of Medicine, Salt Lake City, UT, USA

John R. Holmen, Medical Informatics Department, Intermountain Healthcare, Salt Lake City, UT, USA

Titte R. Srinivas, Division of Nephrology, Intermountain Healthcare, Salt Lake City, UT

Scott A. Clements, Department of Pediatrics, University of Utah, Salt Lake City, UT, USA

References

  1. Adams SF (1926). The seasonal variation in the onset of acute diabetes: the age and sex factors in 1,000 diabetic patients. Archives of Internal Medicine, 37(6), 861–864. [Google Scholar]
  2. Atkinson MA, & Eisenbarth GS (2001). Type 1 diabetes: new perspectives on disease pathogenesis and treatment. The Lancet, 358(9277), 221–229. [DOI] [PubMed] [Google Scholar]
  3. Atkinson MA, Eisenbarth GS, & Michels AW (2014). Type 1 diabetes. The Lancet, 383(9911), 69–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Azage M, Kumie A, Worku A, & Bagtzoglou AC (2015). Childhood diarrhea exhibits spatiotemporal variation in northwest Ethiopia: a SaTScan spatial statistical analysis. PLoS One, 10(12), e0144690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Bach JF, 2002. The effect of infections on susceptibility to autoimmune and allergic diseases. New England Journal of Medicine, 347(12), pp.911–920. [DOI] [PubMed] [Google Scholar]
  6. Bach JF, & Chatenoud L (2012). The hygiene hypothesis: an explanation for the increased frequency of insulin-dependent diabetes. Cold Spring Harbor Perspectives in Medicine, 2(2), a007799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Ball SJ, Haynes A, Jacoby P, Pereira G, Miller LJ, Bower C, & Davis EA (2014). Spatial and temporal variation in type 1 diabetes incidence in Western Australia from 1991 to 2010: increased risk at higher latitudes and over time. Health & Place, 28, 194–204. [DOI] [PubMed] [Google Scholar]
  8. Bell RA, Mayer-Davis EJ, Beyer JW, D’Agostino RB, Lawrence JM, Linder B, & Dabelea D (2009). Diabetes in non-Hispanic white youth: prevalence, incidence, and clinical characteristics: the SEARCH for Diabetes in Youth Study. Diabetes Care, 32 (Supplement 2), S102–S111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Boettler T, & Von Herrath M (2011). Protection against or triggering of Type 1 diabetes? Different roles for viral infections. Expert Review of Clinical Immunology, 7(1), 45–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cooper JD, Smyth DJ, Walker NM, Stevens H, Burren OS, Wallace C, & Spector TD (2011). Inherited variation in vitamin D genes is associated with predisposition to autoimmune disease type 1 diabetes. Diabetes, 60(5), 1624–1631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Disanto G, Chaplin G, Morahan JM, Giovannoni G, Hyppoenen E, Ebers GC, & Ramagopalan SV (2012). Month of birth, vitamin D and risk of immune-mediated disease: a case control study. BMC Medicine, 10(1), 69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Di Ciaula A (2014). Association between air pollutant emissions and type 1 diabetes incidence in European countries. Advances in Research, 409–425. [Google Scholar]
  13. Fotheringham AS, & Wong DW (1991). The modifiable areal unit problem in multivariate statistical analysis. Environment and Planning A, 23(7), 1025–1044. [Google Scholar]
  14. Freedman DA (1999). Ecological inference and the ecological fallacy. International Encyclopedia of the Social & Behavioral Sciences, 6(4027–4030), 1–7. [Google Scholar]
  15. Frodsham SG, Yu Z, Lyons AM, Agarwal A, Pezzolesi MH, Dong L, & Smith KR (2019). The familiality of rapid renal decline in diabetes. Diabetes, 68(2), 420–429. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Fronczak CM, Barón AE, Chase HP, Ross C, Brady HL, Hoffman M, & Norris JM (2003). In utero dietary exposures and risk of islet autoimmunity in children. Diabetes Care, 26(12), 3237–3242. [DOI] [PubMed] [Google Scholar]
  17. Gillespie KM (2006). Type 1 diabetes: pathogenesis and prevention. Canadian Medical Association Journal, 175(2), 165–170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Glaz J, & Balakrishnan N (1999). Introduction to scan statistics. In Scan Statistics and Applications (pp. 3–24). Birkhäuser, Boston, MA. [Google Scholar]
  19. Hämäläinen AM, & Knip M (2002). Autoimmunity and familial risk of type 1 diabetes. Current Diabetes Reports, 2(4), 347–353. [DOI] [PubMed] [Google Scholar]
  20. Harjutsalo V, Lammi N, Karvonen M, & Groop PH (2010). Age at onset of type 1 diabetes in parents and recurrence risk in offspring. Diabetes, 59(1), 210–214. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Hyöty H (2002). Enterovirus infections and type 1 diabetes. Annals of Medicine, 34(3), 138–147. [PubMed] [Google Scholar]
  22. Karvonen M, Viik-Kajander M, Moltchanova E, Libman I, LaPorte RONALD, & Tuomilehto J (2000). Incidence of childhood type 1 diabetes worldwide. Diabetes Mondiale (DiaMond) Project Group. Diabetes Care, 23(10), 1516–1526. [DOI] [PubMed] [Google Scholar]
  23. Knip M, Veijola R, Virtanen SM, Hyöty H, Vaarala O, & Åkerblom HK (2005). Environmental triggers and determinants of type 1 diabetes. Diabetes, 54 (suppl 2), S125–S136. [DOI] [PubMed] [Google Scholar]
  24. Knip M, Virtanen SM, & Åkerblom HK (2010). Infant feeding and the risk of type 1 diabetes. The American Journal of Clinical Nutrition, 91(5), 1506S–1513S. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Kulldorff M (1997). A spatial scan statistic. Communications in Statistics-Theory and Methods, 26(6), 1481–1496. [Google Scholar]
  26. Kulldorff M, Huang L, Pickle L, & Duczmal L (2006). An elliptic spatial scan statistic. Statistics in Medicine, 25(22), 3929–3943. [DOI] [PubMed] [Google Scholar]
  27. Moltchanova EV, Schreier N, Lammi N, & Karvonen M (2009). Seasonal variation of diagnosis of Type 1 diabetes mellitus in children worldwide. Diabetic Medicine, 26(7), 673–678. [DOI] [PubMed] [Google Scholar]
  28. Ozdenerol E, Williams BL, Kang SY, & Magsumbol MS (2005). Comparison of spatial scan statistic and spatial filtering in estimating low birth weight clusters. International Journal of Health Geographics, 4(1), 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Patterson CC, Carson DJ, Hadden DR, & Northern Ireland Diabetes Study Group. (1996). Epidemiology of childhood IDDM in Northern Ireland 1989–1994: low incidence in areas with highest population density and most household crowding. Diabetologia, 39(9), 1063–1069. [DOI] [PubMed] [Google Scholar]
  30. Patterson CC, Dahlquist GG, Gyürüs E, Green A, Soltész G, & EURODIAB Study Group. (2009). Incidence trends for childhood type 1 diabetes in Europe during 1989–2003 and predicted new cases 2005–20: a multicentre prospective registration study. The Lancet, 373(9680), 2027–2033. [DOI] [PubMed] [Google Scholar]
  31. Rønningen KS (2015). Environmental trigger (s) of type 1 diabetes: why so difficult to identify? BioMed Research International, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Van Belle TL, Coppieters KT, & Von Herrath MG (2011). Type 1 diabetes: etiology, immunology, and therapeutic strategies. Physiological Reviews, 91(1), 79–118. [DOI] [PubMed] [Google Scholar]
  33. Vehik K, Hamman RF, Lezotte D, Norris JM, Klingensmith G, Bloch C, & Dabelea D (2007). Increasing incidence of type 1 diabetes in 0-to 17-year-old Colorado youth. Diabetes Care, 30(3), 503–509. [DOI] [PubMed] [Google Scholar]
  34. Wakefield J (2007). Disease mapping and spatial regression with count data. Biostatistics, 8(2), 158–183. [DOI] [PubMed] [Google Scholar]
  35. Wakefield J, & Lyons H (2010). Spatial aggregation and the ecological fallacy. In Handbook of Spatial Statistics (pp. 537–554). CRC Press. [Google Scholar]
  36. Weng J, Zhou Z, Guo L, Zhu D, Ji L, Luo X, & Jia W (2018). Incidence of type 1 diabetes in China, 2010–13: population-based study. British Medical Journal, 360, j5295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Yeung WCG, Rawlinson WD, & Craig ME (2011). Enterovirus infection and type 1 diabetes mellitus: systematic review and meta-analysis of observational molecular studies. British Medical Journal, 342, d35. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES