Abstract
This dataset presents the most comprehensive estimate of residential air conditioning (AC) prevalence across the continental United States. Using property-level data for over 103 million housing units from the Dewey database, we imputed and classified four AC types: central, other, evaporative cooler, and none, using XGBoost models optimized for performance. Housing characteristics, socioeconomic indicators, and environmental conditions, such as Cooling Degree Days and elevation, informed predictions. The final product offers national coverage with spatial resolution at the census tract, ZIP code, and metropolitan levels. Model validation was conducted using American Housing Survey data, with strong alignment observed for the central and no air conditioning (AC) categories. This dataset addresses longstanding gaps in understanding the geographic and demographic disparities in AC access, critical for public health, climate adaptation, and energy equity research. Users may integrate these data into epidemiological modeling, resilience planning, and policy analysis to support heat vulnerability assessments and infrastructure interventions.
Subject terms: Geography, Climate change
Background & Summary
Multiple regions around the world experienced unprecedented heat in the warm seasons of 2023 and 20241. Record-breaking temperature extremes can cause severe impacts on society and the environment. Extreme heat is often considered a silent killer, but better preparation can reduce negative health outcomes2,3. Scientists and decision-makers have been making efforts to increase awareness of extreme heat and emphasize the urgency of preparedness. For instance, since 2022, scientists have started to name heat waves4. In 2024, the European Court declared that having a safe climate is a human right5.
The impacts of extreme heat vary according to demographic, economic, and environmental factors. The risk of mortality and mobility increases among people experiencing financial instability, inadequate housing, limited access to public spaces, safe and connected communities, and water sources6–8. For vulnerable populations such as the elderly, pregnant women, children, and people with pre-existing conditions, researchers have identified the use of air conditioning (AC) as one of the most straightforward measures to combat extreme heat9–12.
However, limited data on AC prevalence have hampered further investigation on this matter. The most recent nationwide household-level AC prevalence survey was conducted in 198013. While the American Housing Survey (AHS) tracks AC ownership, this survey focuses on selected populations, and geographical information is only available at the regional or metropolitan level. Moreover, the lack of data on the trajectory of AC ownership at the national level poses a challenge to researchers in understanding its impact.
Some scholars have tried to understand the prevalence of AC and its effects on health. Gronlund & Berrocal14, Sera et al.15, and Romitti et al.16 developed recent AC estimates using AHS data. However, these studies focus on specific cities, lack nationwide information, and predict probabilities based on the number of households with AC. Ahn & Uejio17 also estimated AC prevalence using real estate data. However, their analysis was limited to California and classified AC systems into only three categories: none, other, and central. Although the Energy Information Administration (EIA)18 reports that 90% of Americans possessed an AC unit encompassing central, window, and wall units as of 2022, the effectiveness of these units varies significantly by type and regional climate. Some units perform better in certain climates than others19,20. For example, a survey found that 83% of respondents owned a Central AC unit, and 40% answered that their home still felt hot during the summer of 2011 in Houston, Texas21. Quinn et al.22 observed that the average indoor temperature differed by 3 degrees Celsius between households with Central AC and those with portable AC units in New York City, New York.
Another concern with lacking AC information is that there are significant spatial disparities of AC ownership due to socioeconomic and infrastructural determinants, especially among the population at risk23–26. Various studies have also indicated that low-income households have less access to AC23,27,28. Rural areas that lack infrastructure or have less access to cooling centers27. Additionally, evidence showed that within urban areas, AC prevalence varies, and this differs across various climates23,27,28.
To address the limitations of prior research, this study developed AC prevalence maps of various types (central, other, evaporative coolers, and none) in the continental United States (U.S.). These maps will be invaluable for health and indoor environment research, protecting vulnerable populations, and sustainable energy planning. Furthermore, this data can contribute to climate change adaptation and mitigation strategies12.
Methods
This study utilized real estate data from Dewey to create a nationwide AC prevalence map. Next, we validated the results against those of the American Housing Survey and previously published studies’ estimates. In the following sections, we detail the datasets, their available attributes, and geographical scope. We then outline the data processing and estimation procedures.
Data sources
Real estate data (Warren Group)
The Warren Group provides proprietary real estate data, which was distributed by Dewey and contained property characteristics (100 variables) for more than 155 million properties across the U.S. for the year 202129. Previous studies have found that several property characteristics are highly related to AC ownership, such as geographical location (state), year built, occupancy status, housing types, number of stories, heating type, total number of rooms, renovated year, and building quality14,16,17 Studies have also found that newer houses are more likely to have AC units than older ones, indicating that the year built is strongly associated with AC ownership27. For instance, the prevalence of different types of AC units has varied over time. Window unit ACs have been available since the 1930s, but were not widely used initially due to their high cost. For example, in 1947, about 43,000 units had been sold. By the late 1960s, most new homes had Central AC, and window units became more affordable30.
The geographical location of the property is a critical part of the modeling since cooler climate regions tend to have Others (window, wall, or portable) rather than Central AC16. Additionally, housing types, heating systems, and building quality are also important indicators of AC ownership17. Housing types are also considered as a strong predictor, such as mobile homes, which are less likely to have an AC unit14,31. Poor housing quality, including low-quality insulation and poor construction, is likely related to AC ownership.
Therefore, this study examined several variables: geographical location (latitude, longitude), year built, housing types, occupancy status, number of stories, heating type, AC types, total number of rooms, and building quality. The year built refers to the original construction year, while land use specifies the county’s land use code (e.g., commercial, residential, office, industrial). Housing types indicate types such as condo, duplex, quadplex, townhouse, or multi-family. Occupancy status reflects whether the unit is owner-occupied or rented. The original AC types include central, evaporative cooler, packaged unit, window, none, others, partial, chilled water, refrigeration, ventilation, wall, yes, and geothermal. Heating types indicate central, heat pump, forced, electric, solar radiant, geothermal, etc. We recategorized the AC types into four categories: Central (central, geothermal, packaged unit), Others (partial, chilled water, wall, window unit), Evaporative Cooler, and No AC (refrigeration, ventilation).
The dataset includes 103,022,912 housing unit records categorized by AC type, which were 73.6% of the total number of housing units in the U.S. CONUS (140,056,552)32. Central AC was the most commonly reported system, present in 30.73% of housing units. Other types of AC systems included miscellaneous systems (0.75%) and Evaporative Coolers (0.39%), which were much less common. Approximately 4.08% of units report having No AC, while 10.63% fall under a general Yes AC category, which may indicate unspecified AC systems. Notably, 54.52% of records have missing or unreported data for the AC type (Table 1).
Table 1.
Distribution of AC types among residential properties.
| AC Type | Count | Percentage |
|---|---|---|
| Central | 31,656,924 | 30.73% |
| Others | 775,698 | 0.75% |
| Evaporative Cooler | 396,962 | 0.39% |
| No AC | 4,205,172 | 4.08% |
| Yes AC | 10,948,801 | 10.63% |
| NA | 56,170,915 | 54.52% |
| Total | 103,022,912 | 100.00% |
The completeness of real estate data, including AC, varied across states due to non-standardized jurisdictional recording systems. The East and Midwest regions had relatively fewer missing records and higher completeness rates (ranging between 60–100%) compared to the West and West Coast (Fig. 1). Fig. 1a illustrates how data missingness varies substantially by state. Louisiana exhibited the highest level of missingness, with 93.8% of its counties (60 out of 64) lacking any valid AC information, and North Dakota followed closely with 81.1% (43 out of 53 counties). These high percentages suggest potential gaps in data collection or reporting practices in specific regions. In contrast, several states showed minimal geographic missingness. Georgia, New York, and Ohio each had only one county with completely missing AC data, representing a small share of their total number of counties (0.6%, 1.6%, and 1.1%, respectively).
Fig. 1.
(a) Number of properties missing AC information at the county level, (b) Completeness of AC information at the county level (completeness = (Number of properties with AC information / Total number of properties)*100).
Fig. 1b shows the percentage of completeness at the county level. For example, Virginia showed the highest rate of completeness among the listed states, with 8 out of 95 counties (8.4%) having more than 90% records with AC data. Michigan followed with 3 out of 83 counties (3.6%), and Georgia ranked third with 3 out of 159 counties (1.9%). By contrast, the remaining states, including New York, Florida, and Illinois, each have only one county with more than 90% of AC data.
Real estate property data imputation
The Warren Group property data (https://www.deweydata.io) contains a substantial amount of property information; however, the completeness of each attribute value varies across variables. Ahn and Uejio17 discovered that real estate data is originally collected at the municipality level with non-standardized jurisdiction recording systems, and the missing values do not occur completely at random, which can cause some bias in the data. Consequently, this study conducted an imputation of missing variables using a model based on the Random Forest algorithm. This algorithm automatically identifies columns with missing values as targets and uses the remaining available variables as predictors during the iterative imputation process33.
The advantages of imputation based on Random Forest include the ability to impute categorical variables and a demonstrated high accuracy of results33,34. To address missing values in the dataset, we implemented a non-parametric imputation method using the missForest algorithm, which employs random forests to iteratively impute missing values in both categorical and numerical variables33. This approach was selected due to its robustness to complex interactions and mixed data types, which are common in real estate and housing datasets. Initially, missForest imputes the missing values with the mean/mode before models that minimize the sum of squared differences between the current and previous imputation are iteratively fitted. Of the input variables (e.g., renovation year, heating type, total number of rooms, housing condition, housing type, and ownership), housing type and ownership had the highest missingness, each exceeding 50%. We eliminated these variables from the imputation and analysis33.
To simulate, evaluate, and optimize model performance, we introduced artificial missingness by randomly masking 20% of the values. This procedure simulates a Missing Completely at Random (MCAR) scenario, where the probability of missingness is unrelated to the observed or unobserved values. We conducted sensitivity testing for the missForest algorithm, running multiple configurations to optimize performance. This included varying the number of randomly selected variables (mtry) from 1 to 15, setting the number of trees (ntree) to 10 and 1000 in different tests, allowing up to 7 iterations to ensure convergence (Supplementary Fig. 1). Model performance was evaluated using the built-in out-of-bag (OOB) error estimates provided by missForest, reporting both Normalized Root Mean Squared Error (NRMSE) for continuous variables and proportion of falsely classified (PFC) values for categorical variables. This approach serves as a robust internal validation method in the absence of an external gold-standard dataset. A tuning process was performed to identify the optimal mtry value that minimized both error metrics. Based on these diagnostics, the final imputation model was run using the optimal parameters (mtry = 10, ntree = 100), and the imputed dataset was used for all subsequent analyses.
Environmental data
Several environmental factors are associated with AC ownership. Cooling Degree Days (CDDs) are defined as a measure of the number of annual days when the air temperature exceeds 18 °C and are often used as a measure of climate. CDDs are a critical predictor of AC ownership14. Cooler climate areas tend to have less AC, and urban areas in cooler climates show greater variation in AC prevalence between urban cores and suburban areas16. The National Centers for Environmental Information (NCEI) Climate Divisional Database provided long-term average CDDs at the county level from 1970 to 2020 (https://www.ncei.noaa.gov/access/monitoring/climate-at-a-glance/county/mapping/110/tavg/197001/12/value)35.
For comparison analysis and data visualization, we used Core Based Statistical Areas (CBSAs) to distinguish urban and rural areas (https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html). Following U.S. Census definitions, we classified urban areas as Metropolitan Statistical Areas (population ≥ 50,000) and Micropolitan Statistical Areas (population between 10,000 and 49,999), while all remaining areas were classified as rural36.
Demographic and socioeconomic
Studies have suggested that socioeconomic factors are highly associated with AC ownership. For instance, technology and economic development are related to the introduction of AC30. Various studies applied nighttime satellite data or Gross Domestic Product (GDP) as proxies to measure the technology and economic development37–39. Commonly, metrics of urban growth, land use, and economic development have been linked to the adoption of AC. Several studies have provided evidence of a close link between urbanization growth and GDP per capita40,41. Therefore, this study applied Historical Settlement Data Compilation for the U.S. (HISDAC-US) data, which describes settlement development from 1810 to 2020 in the U.S. at a 250 m spatial resolution42,43. HISDAC-US was developed by combining various data related to settlement development, such as nationwide parcel data, real estate data (Zillow’s Transaction and Assessment Database; ZTRAX), and the Open City Model based on Microsoft building footprint data. HISDAC-US contains various characteristics that describe urban development, such as building intensity (BUI), which represents the sum of indoor areas of each property within a spatial unit, the number of unique building locations (BPUL), and the number of unique property records (BUPR) (10.7910/DVN/45B8IU)44. BUI, BUPL, and BUPR were strongly correlated to each other. We selected BUPR since it had the most completeness and illustrates the overall development of the surrounding areas. We conducted a zonal statistic of BUPR at the census tract level with 2020 census boundaries (The Warren Group boundaries were based on 2020 census geographies) and utilized them as explanatory variables (Table 2).
Table 2.
Input data for AC prevalence estimation.
| Category | Variables | Spatial resolution | Data source |
|---|---|---|---|
| Property | Air conditioning types | Latitude Longitude | Dewey29 |
| Renovation year | Latitude Longitude | Dewey29 | |
| Heating Type | Latitude Longitude | Dewey29 | |
| Total number of rooms | Latitude Longitude | Dewey29 | |
| Housing condition | Latitude Longitude | Dewey29 | |
| Climate | Cooling Degree Days | County | National Centers for Environmental Information (NCEI)58 |
| Socioeconomic | % Black or African American | Census tract | NHGIS13 |
| % Hispanic and Latino | Census tract | NHGIS13 | |
| Education attainment | Census tract | NHGIS13 | |
| Median household income | Census tract | NHGIS13 | |
| BUPR | Census tract | Ahn et al.44 | |
| Historical housing policy | Census tract | Meier & Mitchell59 |
Various studies also indicated that household income, age, education attainment, ethnicity, and historical ethnicity of neighborhoods have an association with the AC ownership16,23,27,28. For example, the low-income and low-education attainment population is least likely to have an AC16,23. The neighborhood ethnicity showed a strong relation with AC ownership23. The neighborhoods have a high proportion of Historical Black/African American populations, which are less likely to have AC23,27,45. Additionally, Hispanic, Latino, and Black/African American populations tend not to have AC with some variations in different climate regions23,28. Therefore, we included the percentage of historical Black/African American, the percentage of Black/African American, the percentage of Hispanic and Latino, and the median household income from the American Community Survey at the census tract level (https://www.census.gov/programs-surveys/acs.html)32.
Data integration
The property data were matched with sociodemographic data derived from the American Community Survey. These data were further linked to the CDDs and HISDAC data. All datasets were integrated using the 2020 census tract boundaries, as both the property and CDDs data were already organized according to this geographic framework.
Estimating air conditioning prevalence
The Warren Group data contains property records for various land use categories, including governmental, residential, vacant, commercial, industrial, office, vacant, and agricultural with 9 categories and 341 subcategories. The original dataset contained approximately 150,856,800 property records. Due to the absence of validation data for certain territories, we excluded records from Hawaii, Alaska, and Puerto Rico, accounting for 2% of the dataset (2,478,277 records), to ensure consistency and reliability in subsequent analyses. Following this, we filtered for residential properties, resulting in a final analytic sample of 103,022,912 records representing 68.3% of the original dataset.
We applied a sequential two-step AC classification procedure. The first series of procedures estimates the prevalence of properties with Yes AC to Others, Evaporative Coolers, and Central AC (Fig. 2, step 1). The second classification procedure focused on records with more specific information (central, evaporative cooler, other, or none) (Fig. 2, step 2). The procedures are followed as described below.
Fig. 2.
Workflow for estimating residential air conditioning prevalence.
We partitioned the data into training and test sets, assigning 80% of households to the training set and 20% to the hold-out test set46. We used stratified random splitting by AC category and CDDs to mitigate class imbalances and ensure that evaluation metrics reflect performance across all AC types47,48. We used the XGBoost (Extreme Gradient Boosting) algorithm to model the relationship between property features, demographic characteristics, and household AC type. XGBoost is a scalable, high-performance implementation of gradient-boosted decision trees that can accurately model structured datasets49. Its success in spatial and geodemographic contexts has been well documented50,51, making it a suitable choice for this multi-class classification task. We configured the model using the softmax objective, training an ensemble of decision trees in which each tree improves upon the errors of the previous one (boosting).
To optimize model performance, we tuned key hyperparameters related to model complexity and regularization, including the number of boosting iterations, maximum tree depth, minimum loss reduction required for further partitioning, L1 regularization term on weights, minimum sum of instance weight needed in a child node, and the subsample ratio of features used per tree. Tuning was guided by cross-validation within the training data to reduce variance in model evaluation and ensure generalizable parameter estimates46. This approach is commonly recommended for structured data when high predictive stability is desired. We aimed to maximize the macro-averaged F1-score, which balances precision and recall across all classes. We used 10 cross validations, acknowledging this as a computationally efficient choice for a preliminary tuning pass. Despite the limited search, the model yielded strong performance, highlighting the robustness of the XGBoost framework.
After identifying the best hyperparameters, we retrained the XGBoost model on the full training set and evaluated it on the test set. We used standard classification metrics to assess model performance, including accuracy (the overall correct classifications), precision (the correctness of positive predictions), recall (sensitivity to actual positives), and F1-score (the harmonic mean of precision and recall). These metrics were calculated for each AC class, and a macro-averaged F1-score was reported to reflect overall model balance.
The optimized XGBoost model achieved high overall accuracy on the test set, with a classification accuracy of 98.87%, indicating strong predictive performance across all AC categories. However, performance varied by class, with especially high metrics for the majority class. The results showed particularly strong performance for the majority class, Central AC (F1-score: 0.99). Performance was lower for less common categories. Evaporative Coolers had a precision of 0.93 and a recall of 0.78, indicating some under-identification. The Others category showed the lowest recall (0.68) and F1-score (0.78), indicating a higher frequency of misclassification. The macro-averaged (averaging metrics across all classes equally) F1-score was 0.87, highlighting a drop in performance when treating all classes equally, while the weighted F1-score (0.9) closely reflected the model’s strong performance on dominant classes. These results underscore the model’s high overall accuracy but also suggest the need for refinement in identifying minority classes. Similar results were obtained for the prediction using all datasets, with a slightly lower overall accuracy of 0.97 (Table 3). Supplementary Table 1 provides a confusion matrix.
Table 3.
Accuracy results from Yes AC prediction and all AC type prediction.
| Class | Precision | Recall | F1-score | Support | |
|---|---|---|---|---|---|
| Yes AC prediction | Central | 0.99 | 1.00 | 0.99 | 15645210 |
| Evaporative Cooler | 0.93 | 0.78 | 0.85 | 201057 | |
| Others | 0.90 | 0.68 | 0.78 | 339684 | |
| macro avg | 0.94 | 0.82 | 0.87 | 16185951 | |
| weighted avg | 0.99 | 0.99 | 0.99 | 16185951 | |
| Overall accuracy | 0.99 | ||||
| All AC type prediction | Central | 0.97 | 0.99 | 0.98 | 30012013 |
| Evaporative Cooler | 0.81 | 0.68 | 0.74 | 391363 | |
| No AC | 0.9 | 0.71 | 0.79 | 1975044 | |
| Others | 0.87 | 0.65 | 0.74 | 642140 | |
| macro avg | 0.89 | 0.76 | 0.82 | 33020560 | |
| weighted avg | 0.97 | 0.97 | 0.97 | 33020560 | |
| Overall accuracy | 0.97 |
To further investigate the accuracy of the model, we compared the performance based on urbanicity. Table 4 shows the model’s classification performance by urbanicity, revealing distinct differences between urban and rural tracts. Overall, the model demonstrated high predictive accuracy in both contexts, with weighted F1-scores of 0.95 for urban areas and 0.92 for rural areas. However, class-specific results show nuanced patterns that underscore regional prediction challenges.
Table 4.
Accuracy results from all AC type predictions by urban and rural.
| Class | Precision | Recall | F1-score | Support | |
|---|---|---|---|---|---|
| Urban | Central AC | 0.96 | 1.00 | 0.98 | 15,333,619 |
| Evaporative Cooler | 0.88 | 0.75 | 0.81 | 198,914 | |
| No AC | 0.93 | 0.60 | 0.73 | 1,294,641 | |
| Other AC | 0.84 | 0.66 | 0.74 | 315,500 | |
| Macro avg | 0.90 | 0.75 | 0.82 | 17,142,674 | |
| Weighted avg | 0.95 | 0.96 | 0.95 | 17,142,674 | |
| Overall accuracy | 0.96 | ||||
| Rural | Central AC | 0.93 | 0.99 | 0.96 | 273,645 |
| Evaporative Cooler | 0.91 | 0.79 | 0.85 | 1,895 | |
| No AC | 0.95 | 0.67 | 0.79 | 55,310 | |
| Other AC | 0.79 | 0.69 | 0.74 | 22,959 | |
| Macro avg | 0.90 | 0.79 | 0.84 | 353,809 | |
| Weighted avg | 0.92 | 0.92 | 0.92 | 353,809 | |
| Overall accuracy | 0.92 |
In urban areas, the model achieved exceptionally strong performance for the dominant class, Central AC, with near-perfect recall (1.00) and an F1-score of 0.98. Predictions for Evaporative Coolers and Others were somewhat less accurate, with F1-scores of 0.81 and 0.74, respectively. The No AC category showed a recall of 0.60, indicating under-identification in urban settings despite reasonable precision (0.93).
In rural areas, the model maintained strong performance for Central AC (F1-score: 0.96), but class balance was more even, and sample sizes were smaller. The recall for Evaporative Coolers improved to 0.79, yielding a higher F1-score (0.85) than in urban areas. However, No AC predictions remained challenging, with a recall of 0.67 and an F1-score of 0.79. The Others category continued to show moderate performance with a rural F1-score of 0.74.
Figs. 3, 4 visualize SHAP feature importance values from the two-step classification process outlined in Fig. 2. Fig.3 reflects the first model step, which classifies AC types only among properties identified as having some form of AC. It includes three categories: Central, Evaporative Cooler, and Other. Fig. 4, in contrast, displays the results of the second model that includes the No AC category and classifies all four groups across the full dataset.
Fig. 3.
SHAP feature importance plots from the first-step AC type classifier by AC types (a) Central AC, (b) Evaporative Cooler, and (c) Others.
Fig. 4.
SHAP feature importance plots from the second-step full-sample classifier by AC types (a) Central AC, (b) Evaporative Cooler, (c) Others, and (d) No AC.
Fig. 3 presents SHAP summary plots that illustrate the relative importance and direction of influence of various predictors across three AC types. SHAP values show how much each feature in a model pushes a prediction higher or lower. Positive values mean the feature increases the prediction; negative values mean it decreases it52. The plots show both the magnitude and polarity of each’eature’s impact on the model output, with color gradients representing high (red) and low (blue) feature values.
Across all AC types, Heating Type, CDDs, and Renovation Year emerge as dominant predictors, though their effects differ by category. For Central AC (Fig. 3a), higher CDDs values and more recent renovation years increase the likelihood of Central AC installation, suggesting alignment with climate demand and housing modernization. Demographic variables, such as the proportion of Hispanic and Black/African American residents, also show a moderate influence, though less pronounced than structural variables.
In contrast, Evaporative Coolers (Fig. 3b) are strongly associated with high CDDs and a higher proportion of Hispanic residents. Still, they are less associated with newer renovations, indicating a more climate-specific and potentially cost-sensitive pattern of adoption. For Others (Fig. 3c), the model reveals a broader distribution of influential factors. While Heating Type and CDDs remain key, socioeconomic indicators such as Median Income, Education, and Historical Housing Policy Score show more pronounced effects compared to Others, reflecting the diversity and potential inequities in this residual category.
Fig. 4 presents SHAP summary plots illustrating the relative importance and directional influence of key features in predicting each AC type. Consistent with patterns observed in Fig. 3, Heating Type, CDDs, and Renovation Year remain among the most influential predictors across all AC types, though their specific impacts vary. For Central AC (Fig. 4a), the combination of high CDDs and recent renovation years is associated with a greater likelihood of adoption, reinforcing climate-driven needs and the role of newer, modernized housing. This pattern is also evident in Fig. 3a.
In Evaporative Coolers (Fig. 4b), high CDDs are again a prominent driver. Still, there is a clearer demographic signal compared to Central AC, with stronger contributions from the proportion of Hispanic residents. A trend was also found in the prediction of Yes AC (Fig. 3b), suggesting regionally or culturally specific usage of cooling technologies.
For Others (Fig. 4c), the model reveals a broader mix of contributing variables. Although heating type and CDDs remain significant, the importance of median income, education, and historical housing policy score is less important, suggesting a heterogeneous category shaped by socioeconomic and structural constraints similar to the more dispersed influences seen in Fig. 3c.
Finally, the No AC (Fig. 4d) shows a distinctive result, with the absence of cooling most strongly influenced by heating type and demographic indicators such as the proportion Hispanic, the proportion Black/African American, lower median income, and poor housing conditions. Notably, the historical housing policy score plays a more substantial role here compared to other categories, echoing insights from Fig. 3 that point to the long-term impacts of structural inequities and underinvestment in housing infrastructure.
Fig. 5 provides an overview of the dataset’s geographic coverage, resolution, and variable structure. Fig. 5a displays county-level choropleths for the conterminous U.S. showing the prevalence (%) of four air-conditioning (AC) categories: Central, Evaporative Cooler, Other, and No AC computed directly from the county table. Fig. 5b presents census-tract choropleths for ten major U.S. metropolitan areas to illustrate the intra-urban detail available in the tract table.
Fig. 5.
AC prevalence by each type (a) nationwide at the county level (b) 10 metropolitan cities at the census tract level (black dots mark the downtown core of each city).
Fig. 6 illustrates the prevalence of AC types according to urbanicity. Fig. 6 shows the distribution of urban areas (census tracts within the U.S. Office of Management and Budget Metropolitan Statistical Areas) and rural (non-metropolitan) areas, and compares the AC prevalence.
Fig. 6.
State-level summary of AC prevalence by urbanicity (Note: each state is located approximately where it is located in a map of the U.S.).
Data Records
The full dataset is archived in Harvard Dataverse53–56. It includes AC prevalence estimates for each AC type (Central, Other, Evaporative Cooler, and No AC) at multiple spatial scales and time points, and probabilities for residential air conditioning (AC) types as an uncertainty layer. The following files are provided:
| File names | Descriptions |
|---|---|
| summary_zcta.zip | The archive contains ZCTA-level air-conditioning (AC) prevalence (count and percentage) by AC type for the 2010, 2015, and 2020 boundary years, provided as CSV files: summary_zcta_adjusted_2010.csv, summary_zcta_adjusted_2015.csv, summary_zcta_adjusted_2020.csv |
| summary_tract.zip | The archive contains census tract-level air-conditioning (AC) prevalence (count and percentage) by AC type for the 2010, 2015, and 2020 boundary years, provided as CSV files: summary_tract_adjusted_2010.csv, summary_tract_adjusted_2015.csv, summary_tract_adjusted_2020.csv |
| summary_metro.zip | The archive contains Census Metro level air-conditioning (AC) prevalence (count and percentage) by AC type for the 2010, 2015, and 2020 boundary years, provided as CSV files: summary_metro_adjusted_2010.csv, summary_metro_adjusted_2015.csv, summary_metro_adjusted_2020.csv |
| Uncertainty.zip | The archive contains modeled probabilities for air conditioning (AC) types at the census tract level for the 2010, 2015, and 2020 boundary years, provided as CSV files: summary_ACprob_tract_2010.csv, summary_ACprob_tract_2015.csv, summary_ACprob_tract_2020.csv |
Technical Validation
Comparison with the american housing survey (AHS)
Data on the prevalence of air conditioning is limited, and existing sources such as surveys and prior research often differ in how they categorize AC types and present results (e.g., some report probabilities, while others provide percentages). These inconsistencies make direct comparisons challenging. Nevertheless, we have incorporated several external sources to help validate and contextualize our results.
We utilized national and metropolitan-level data from the AHS. We incorporated the most recent three surveys conducted in 2019, 2021, and 2023 from both the metropolitan and national levels. For metropolitan areas, we analyzed surveys from 20 unique cities (Raleigh-Cary, NC, Birmingham-Hoover, AL, San Jose-Sunnyvale-Santa Clara, CA; Portland-Vancouver-Hillsboro, OR-WA, Las Vegas-Henderson-Paradise, NV, Oklahoma City, OK; Milwaukee-Waukesha, WI, Cleveland-Elyria, OH, New Orleans-Metairie, LA, Rochester, NY, Memphis, TN-MS-AR, Baltimore-Columbia-Towson, MD, Kansas City, MO-KS, Minneapolis-St. Paul-Bloomington, MN-WI, Pittsburgh, PA, Tampa-St. Petersburg-Clearwater, FL; Denver-Aurora-Lakewood, CO; Richmond, VA; Cincinnati, OH-KY-IN; and San Antonio-New Braunfels, TX). Among these cities, five (Milwaukee-Waukesha, Cleveland-Elyria, New Orleans-Metairie, Denver-Aurora-Lakewood, and Cincinnati) had data from both 2019 and 2023. We excluded the 2019 survey data from these cities in our analysis. National-level data from the AHS provides geographical information only at a regional level, which is divided into eight regions: New England, Middle Atlantic, South Atlantic, East South Central, West South Central, Mountain, West, and Pacific.
The AHS survey includes detailed questions about air conditioning ownership, types, and energy sources. To facilitate the comparison, we reclassified the AHS information into three categories: Central AC (electric-powered, piped gas-powered, liquefied petroleum gas-powered, and other fuel source-powered AC), Others (air conditioning units in 1 to 7 rooms), and No AC (no air conditioning). Likewise, our XGBoost prediction results were also reclassified into Central, Others (including Others and Evaporative Cooler), and No AC.
In Fig. 7a, the predicted values for Central AC align closely with the observed AHS data in warmer cities, such as Las Vegas, NV, Miami, FL, and San Antonio, TX, suggesting good model performance in regions with high AC penetration. Cities like Cleveland, OH, and Portland, OR show higher discrepancies, especially in the “Others” and “No AC” categories. The predictions for No AC exhibited a strong correlation with AHS data (Pearson’s r = 0.8, p < 0.05), while Central AC showed a moderate but statistically significant correlation (r = 0.5, p < 0.05). The correlation for Others was weak and not statistically significant. Fig. 7b reveals regional patterns at the division level, including metropolitan areas. Predictions for Central AC and No AC again align better with AHS data compared to Others. Notably, divisions such as the South Atlantic and West South Central showed close agreement for Central AC, while the East North Central and Pacific divisions matched better for No AC. Fig. 7c illustrates validation for rural/non-metropolitan areas, where greater prediction uncertainty is expected due to limited AHS coverage. Nevertheless, a statistically significant correlation was found for No AC (r = 0.7, p < 0.05), particularly in regions such as the East South Central and Mountain divisions, where observed and predicted values align more closely. In contrast, no significant relationships were observed for Central AC types in these non-metro areas.
Fig. 7.
Comparison between predicted AC prevalence and AHS data across spatial scales and system types. (a) Metropolitan-level AHS data; (b) Regional-level national data including metropolitan areas; and (c) Regional-level national data excluding metropolitan areas. Each panel shows predictions versus AHS estimates for Central AC, No AC, and Others (including Evaporative Cooler).
Comparison with the previous studies’ outcomes
In our study, we compared our data with findings from previous literature, including Sera et al.15 and Romitt et al.16. Both studies applied AHS data. Sera et al.15 utilized the AHS data in conjunction with other local-level data sources, and Romitt et al.16 applied micro AHS data to predict the AC prevalence. Romitt et al.16 assessed the probability of prevalence for any type (central and others) of air conditioning. We recategorized our prediction results to encompass any type of AC (Central, Others, and Evaporative Cooler) and calculated the percentage of any type of AC, facilitating a direct comparison. While there are some variances, particularly noted in California and Arizona, these discrepancies are within an expected range when juxtaposing methodologies and regional assessments.
Similarly, Sera et al.15 focused on the probability of Central AC prevalence. Our results revealed some deviations in states such as Colorado, Oregon, and California. Fig. 8 compares the predicted AC prevalence (any type) from the present study with that of Sera et al.15 and Romitti et al.16. We suspect our higher AC prevalence estimates may be more realistic in high-income Western U.S. counties (e.g., Salinas, CA; San Francisco-Oakland-Berkeley, CA; Boulder, CO; Klamath Falls, OR). On the other hand, the existing databases may be more accurate in places where evaporative cooling is common, such as Arizona and central California. Despite these discrepancies, our findings generally align with those reported in the existing literature in terms of trends and magnitudes, illustrating that regional variations can arise from differing input datasets and study methodologies. Nonetheless, our study presents the first rural and continental U.S. AC estimates since the 1980 decennial census, including predictions for various AC system types.
Fig. 8.
Comparison between predicted results and previous studies’ outcomes (a) data from Sera et al.15 (b) data from Romitti et al.16.
Usage Notes
This dataset provides estimates of AC prevalence across multiple geographic levels. Although the predictive model is based on 2021 data, we have generated outputs using boundary definitions from ZCTA level (2000, 2010, and 2020), tract level (2010,2015, and 2020), and metro level (2020, 2015, and 2010) based on the available boundary data to support a wider range of user needs and enable historical comparisons. The prediction was conducted based on the Dewey, which included both occupied and vacant units. To ensure consistency with observed housing conditions, we adjusted the predicted counts of AC types to align with the number of occupied housing units reported in the Decennial Census57. This adjustment was performed using proportional scaling, where the total number of predicted AC-equipped units in each geographic unit was scaled to match the actual occupied housing count.
For the 2010 and 2015 tract-based datasets, we employed geographic crosswalks from the National Historical Geographic Information System (NHGIS)57 to reconcile changes in tract boundaries over time. Specifically, we mapped housing unit counts from 2020 tract definitions back to 2010 boundaries, enabling longitudinal consistency in spatial units while preserving alignment with contemporary AC prevalence patterns.
This dataset covers approximately 73.6% of U.S. households. While it offers valuable insights, we strongly encourage users to integrate these predictions with additional sources, such as local-level survey data or results from prior research, to enhance accuracy and reduce uncertainty.
We encourage users to review Supplementary Fig. 1, which shows variable-level missingness by state, when interpreting model outputs, particularly in states with high rates of missing data. While the missForest algorithm was tuned and tested for robustness, imputed variables inherently introduce a degree of uncertainty that may affect predictions, especially for rare AC types or atypical housing stock. In regions with substantial imputation, users should exercise caution and consider aggregating estimates to higher geographic levels or validating findings with supplementary data sources. For instance, users are encouraged to use mean or average values across geographic units, or to combine categories with lower prediction accuracy, such as grouping AC types into “available” (Central, Evaporative Cooler, Others) and “not available” (No AC) when integrating this dataset with other sources or conducting comparative studies. These strategies can help mitigate variability and reduce uncertainty in the predictions.
We note that while CDDs are a common proxy for cooling demand, our model does not include relative humidity (RH), which can influence AC system choice, especially Evaporative Cooler. RH was excluded due to its correlation with CDDs and the low prevalence of Evaporative Cooler ( < 0.4%) in the dataset. Since RH typically decreases with rising temperature in arid regions but remains high in humid areas, omitting it may underestimate latent cooling needs in humid zones and overestimate Evaporative Cooler use in dry ones. Users should be aware of these potential regional biases. Researchers can analyze this dataset using standard data processing tools in R, Python, or GIS platforms. Suggested downstream steps may include normalization by population or housing counts or comparison against observed survey benchmarks.
Supplementary information
Acknowledgements
Funding for this work was provided through the National Academy of Sciences Gulf Research Program’s Understanding the Effects of Climate Change on Environmental Hazards in Overburdened Communities grant (SCON-10000677) and the University of Kansas General Research Fund allocation 2302047. We gratefully acknowledge access to the Warren Group property data, provided under a data use agreement between the University of Kansas, Florida State University, and Dewey. The findings and interpretations presented here are solely those of the authors and do not necessarily reflect the views of Dewey or Warren Group. .
Author contributions
Y.A. designed the research study. Y.A. conducted the data analysis, visualization, and validation. Y.A. drafted the manuscript, and C.U. provided critical revisions.
Data availability
The air-conditioning prevalence datasets generated in this study are available in the Harvard Dataverse: Metro-Level Summary of Residential Air Conditioning Prevalence (10.7910/DVN/BEKYUN)53, ZCTA-Level Summary of Residential Air Conditioning Prevalence (10.7910/DVN/A0CGTX)54, census tract-Level Summary of Air Conditioning Prevalence (10.7910/DVN/7GLPD7)55 and Uncertainty in Predicted Air Conditioning Types at the census tract Level (10.7910/DVN/DSGOTC)56.
Code availability
The code we generated for this research is available at https://github.com/YoonjungAhn/NationwideAC/.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s41597-025-06104-3.
References
- 1.Blanchard-Wrigglesworth, E., Bilbao, R., Donohoe, A. & Materia, S. Record Warmth of 2023 and 2024 was Highly Predictable and Resulted From ENSO Transition and Northern Hemisphere Absorbed Shortwave Anomalies. Geophys Res Lett52, e2025GL115614 (2025). [Google Scholar]
- 2.Bittner, M. I., Matthies, E. F., Dalbokova, D. & Menne, B. Are European countries prepared for the next big heat-wave? Eur J Public Health24, 615–619 (2014). [DOI] [PubMed] [Google Scholar]
- 3.Thompson, V. et al. The most at-risk regions in the world for high-impact heatwaves. Nature Communications14, 1–8 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Pappas, S. For the first time, scientists have named a heat wave | Live Science. Planet Earth (2022).
- 5.Blattner, C. E. European ruling linking climate change to human rights could be a game changer — here’s how. Nature628, 691 (2024). [DOI] [PubMed] [Google Scholar]
- 6.Mackres, E., Wong, T., Null, S., Campos, R., & Mehrotra, S. The Future of Extreme Heat in Cities: What We Know — and What We Don’t. https://www.wri.org/insights/future-extreme-heat-cities-data (2023).
- 7.Madrigano, J., Ito, K., Johnson, S., Kinney, P. L. & Matte, T. A case-only study of vulnerability to heat wave–related mortality in New York City (2000–2011). Environ Health Perspect123, 672–678 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Gabbe, C. J. & Pierce, G. Extreme Heat Vulnerability of Subsidized Housing Residents in California. Hous Policy Debate30, 843–860 (2020). [Google Scholar]
- 9.Barreca, A., Clay, K., Deschenes, O., Greenstone, M. & Shapiro, J. S. Adapting to climate change: The remarkable decline in the US temperature-mortality relationship over the Twentieth Century. Journal of Political Economy124, 105–159 (2016). [Google Scholar]
- 10.Sera, F. et al. Air Conditioning and Heat-related Mortality. Epidemiology31, 779–787 (2020). [DOI] [PubMed] [Google Scholar]
- 11.Ostro, B., Rauch, S., Green, R., Malig, B. & Basu, R. The effects of temperature and use of air conditioning on hospitalizations. Am J Epidemiol172, 1053–1061 (2010). [DOI] [PubMed] [Google Scholar]
- 12.De Cian, E., Pavanello, F., Randazzo, T., Mistry, M. N. & Davide, M. Households’ adaptation in a warming climate. Air conditioning and thermal insulation choices. Environ Sci Policy100, 136–157 (2019). [Google Scholar]
- 13.IPUMS NHGIS. Decennial Census of Population and Housing Data. https://data2.nhgis.org/main (2025).
- 14.Gronlund, C. & Berrocal, V. Modeling and comparing central and room air conditioning ownership and cold-season in-home thermal comfort using the American Housing Survey. J Expo Sci Environ Epidemiol30, 814–823 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Sera, F. et al. Air Conditioning and Heat-related Mortality: A Multi-country Longitudinal Study. Epidemiology31, 779–787 (2020). [DOI] [PubMed] [Google Scholar]
- 16.Romitti, Y., Sue Wing, I., Spangler, K. R. & Wellenius, G. A. Inequality in the availability of residential air conditioning across 115 US metropolitan areas. PNAS Nexus1, 1–12 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Ahn, Y. & Uejio, C. K. Modeling air conditioning ownership and availability. Urban Clim46, 101322 (2022). [Google Scholar]
- 18.Energy Information Administration. Nearly 90% of U.S. households used air conditioning in 2020 - U.S. Energy Information Administration (EIA). https://www.eia.gov/todayinenergy/detail.php?id=52558 (2022).
- 19.Cardoza, J. E. et al. Heat-Related Illness Is Associated with Lack of Air Conditioning and Pre-Existing Health Problems in Detroit, Michigan, USA: A Community-Based Participatory Co-Analysis of Survey Data. Int J Environ Res Public Health17, 5704 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Bell, M. L., Ebisu, K., Peng, R. D. & Dominici, F. Adverse Health Effects of Particulate Air Pollution: Modification by Air conditioning. Epidemiology20, 682–686 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Hayden, M. H. et al. Adaptive Capacity to Extreme Heat: Results from a Household Survey in Houston, Texas. Weather, Climate, and Society9, 787–799 (2017). [Google Scholar]
- 22.Quinn, A., Kinney, P. & Shaman, J. Predictors of summertime heat index levels in New York City apartments. Indoor Air27, 840–851 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Guirguis, K. et al. Heat, Disparities, and Health Outcomes in San Diego County’s Diverse Climate Zones. Geohealth2, 212–223 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.O’Neill, M. S., Zanobetti, A. & Schwartz, J. Disparities by race in heat-related mortality in four US cities: The role of air conditioning prevalence. Journal of Urban Health82, 191–197 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Biddle, J. E. Making consumers comfortable: the early decades of air conditioning in the United States. (2011).
- 26.Pavanello, F. et al. Air-conditioning and the adaptation cooling deficit in emerging economies. Nat Commun10.1038/s41467-021-26592-2 (2021). [DOI] [PMC free article] [PubMed]
- 27.Ahn, Y., Uejio, C. K., Wong, S., Powell, E. & Holmes, T. Spatial Disparities in Air Conditioning Ownership in Florida, United States. J Maps19 (2023).
- 28.Vant-Hull, B. et al. The harlem heat project A unique media-community collaboration to study indoor heat waves. Bull Am Meteorol Soc99, 2491–2506 (2018). [Google Scholar]
- 29.Dewey (Warren Group). Real Estate Data. https://www.deweydata.io/ (2021).
- 30.Department of Energy. History of Air Conditioning. https://www.energy.gov/articles/history-air-conditioning (2015).
- 31.Peterson, K. Extreme heat is killing people in Arizona’s mobile homes - The Washington Post. The Washington Posthttps://www.washingtonpost.com/climate-environment/2021/07/02/arizona-mobile-home-deaths/ (2021).
- 32.United States Census Bureau. American Community Survey (ACS). https://www.census.gov/programs-surveys/acs (2021).
- 33.Stekhoven, D. J. & Bühlmann, P. Missforest-Non-parametric missing value imputation for mixed-type data. Bioinformatics28, 112–118 (2012). [DOI] [PubMed] [Google Scholar]
- 34.Tang, F. & Ishwaran, H. Random forest missing data algorithms. Stat Anal Data Min10, 363–377 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.National Centers for Environmental Information (NCEI). Climate at a Glance. https://www.ncei.noaa.gov/access/monitoring/climate-at-a-glance/county/mapping/110/tavg/197001/12/value (2024).
- 36.US Census Bureau. Metropolitan and Micropolitan Statistical Areas Map (2020).
- 37.Ghosh, T., Anderson, S. J., Elvidge, C. D. & Sutton, P. C. Using Nighttime Satellite Imagery as a Proxy Measure of Human Well-Being. Sustainability5, 4988–5019 (2013). [Google Scholar]
- 38.van den Bergh, J. C. J. M. The GDP paradox. J Econ Psychol30, 117–135 (2009). [Google Scholar]
- 39.Bloom, D. E., Canning, D. & Fink, G. Urbanization and the wealth of nations. Science (1979)319, 772–775 (2008). [DOI] [PubMed] [Google Scholar]
- 40.Wilhelmi, O. V. & Hayden, M. H. Connecting people and place: a new framework for reducing urban vulnerability to extreme heat. Environmental Research Letters5, 014021 (2010). [Google Scholar]
- 41.Henderson, V. The urbanization process and economic growth: The so-what question. Journal of Economic Growth8, 47–71 (2003). [Google Scholar]
- 42.Ahn, Y., Leyk, S., Uhl, J. H. & McShane, C. M. An Integrated Multi-Source Dataset for Measuring Settlement Evolution in the United States from 1810 to 2020. Scientific Data11, 1–16 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Ahn, Y., Leyk, S., Uhl, J. H. & Mc Shane, C. M. HISDAC-US Version II. https://dataverse.harvard.edu/dataverse/hisdac-usv2 (2024).
- 44.Ahn, Y., Leyk, S. & Uhl, J. H. Historical built-up records (BUPR) Version II - gridded surfaces for the U.S. from 1810 to 2020. https://dataverse.harvard.edu/dataset.xhtml?persistentId=10.7910/DVN/45B8IU (2024).
- 45.Hoffman, J. S., Shandas, V. & Pendleton, N. The Effects of Historical Housing Policies on Resident Exposure to Intra-Urban Heat: A Study of 108 US Urban Areas. Climate8, 12 (2020). [Google Scholar]
- 46.Kohavi, R. A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection. (1995).
- 47.Kotsiantis, S. B. Supervised Machine Learning: A Review of Classification Techniques. Informatica vol. 31 (2007).
- 48.Kuhn, M. & Johnson, K. Applied predictive modeling. Applied Predictive Modeling 1–600 10.1007/978-1-4614-6849-3/COVER (2013).
- 49.Chen, T. & Guestrin, C. XGBoost: A scalable tree boosting system. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining13-17-August-2016, 785–794 (2016).
- 50.Hebryn-Baidy, L. & Rees, G. Machine Learning Algorithms Evaluated for Urban Land Use and Land Cover Classification Using Sentinel 2 Data. (2024).
- 51.Shao, Z., Ahmad, M. N. & Javed, A. Comparison of Random Forest and XGBoost Classifiers Using Integrated Optical and SAR Features for Mapping Urban Impervious Surface. Remote Sensing16, 665 (2024). [Google Scholar]
- 52.Lundberg, S. M., Allen, P. G. & Lee, S.-I. A Unified Approach to Interpreting Model Predictions (2017).
- 53.Ahn, Y. Metro-Level Summary of Residential Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/BEKYUN (2025). [DOI] [PubMed]
- 54.Ahn, Y. ZCTA-Level Summary of Residential Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/A0CGTX (2025). [DOI] [PubMed]
- 55.Ahn, Y. Census Tract-Level Summary of Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/7GLPD7 (2025).
- 56.Ahn, Y. Uncertainty in Predicted Air Conditioning Types at the Census Tract Level - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/DSGOTC (2025).
- 57.Manson, S. et al. IPUMS National Historical Geographic Information System: Version 18.0 [dataset]. IPUMShttps://www.nhgis.org/frequently-asked-questions-faq#GISJOIN_changes, 10.18128/D050.V18.0 (2024).
- 58.Climate at a Glance | County Mapping | National Centers for Environmental Information (NCEI). https://www.ncei.noaa.gov/access/monitoring/climate-at-a-glance/county/mapping/110/tavg/197001/12/value.
- 59.Meier, H. & Mitchell, B. Historic Redlining Indicator for 2010 and 2020 US Census Tracts. nter-university Consortium for Political and Social Researchhttps://www.openicpsr.org/openicpsr/project/141121/version/V3/view?path=/openicpsr/141121/fcr:versions/V3&type=project (2023).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Citations
- Ahn, Y. Metro-Level Summary of Residential Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/BEKYUN (2025). [DOI] [PubMed]
- Ahn, Y. ZCTA-Level Summary of Residential Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/A0CGTX (2025). [DOI] [PubMed]
- Ahn, Y. Census Tract-Level Summary of Air Conditioning Prevalence - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/7GLPD7 (2025).
- Ahn, Y. Uncertainty in Predicted Air Conditioning Types at the Census Tract Level - Nationwide Residential Air Conditioning Prevalence in the United States [Dataset]. Harvard Dataverse10.7910/DVN/DSGOTC (2025).
- Meier, H. & Mitchell, B. Historic Redlining Indicator for 2010 and 2020 US Census Tracts. nter-university Consortium for Political and Social Researchhttps://www.openicpsr.org/openicpsr/project/141121/version/V3/view?path=/openicpsr/141121/fcr:versions/V3&type=project (2023).
Supplementary Materials
Data Availability Statement
The air-conditioning prevalence datasets generated in this study are available in the Harvard Dataverse: Metro-Level Summary of Residential Air Conditioning Prevalence (10.7910/DVN/BEKYUN)53, ZCTA-Level Summary of Residential Air Conditioning Prevalence (10.7910/DVN/A0CGTX)54, census tract-Level Summary of Air Conditioning Prevalence (10.7910/DVN/7GLPD7)55 and Uncertainty in Predicted Air Conditioning Types at the census tract Level (10.7910/DVN/DSGOTC)56.
The code we generated for this research is available at https://github.com/YoonjungAhn/NationwideAC/.








