Skip to main content
Springer logoLink to Springer
. 2026 Mar 5;198(3):278. doi: 10.1007/s10661-026-15102-w

Spatial analysis of air pollution and cancer prevalence in Texas: a machine learning approach using satellite remote sensing data

Fangchao Dong 1,, Muhammad Tauhidur Rahman 1,, Hao Chen 1
PMCID: PMC12963168  PMID: 41784679

Abstract

Cancer prevalence in the world has been attributed to exposure to air pollutants. However, spatial analyses utilizing remote sensing data have been limited. This study combined satellite-derived measurements of eight air pollutants (SO2, NO2, CO, HCHO, CH4, PM1, PM10, and AER AI) with machine learning algorithms to analyze the relationship between air pollution and cancer prevalence across all 254 Texas counties in 2022. Satellite-based air pollution data were obtained from Sentinel-5 Precursor and the Copernicus Atmosphere Monitoring Service, while cancer prevalence data were sourced from the U.S. Centers for Disease Control and Prevention. The spatial clustering analysis revealed distinct regional patterns. The West Texas counties exhibited lower pollution except for elevated AER AI, the East Texas counties showed uniformly high pollutant concentrations, and the Central Texas counties displayed moderate levels. There was high spatial autocorrelation (Moran’s I = 0.69), suggesting cancer prevalence was geographically clustered. Among the four machine learning algorithms tested (Random Forest, Support Vector Machine, Multi-Layer Perceptron, and Ordinary Least Squares), SVM showed the best performance (RMSE = 0.347, R2 = 0.791). ANOVA analysis confirmed that air pollutant variables significantly improved model fit (p < 0.05). The permutation importance analysis identified SO2 as the most influential environmental predictor, followed by AER AI, CH4, and particulate matter, each contributing 30–60% of the predictive weight relative to the top socio-demographic predictor (liquor consumption). These findings reveal significant associations between air pollution and cancer prevalence in Texas, with important implications for targeted public health interventions in high-risk regions. This study shows the value of integrating remote sensing data with machine learning algorithms for regional environmental health assessment where ground-based monitoring is unavailable.

Keywords: Remote sensing, Machine learning, Cancer prevalence, Air pollutants, Permutation importance, Spatial clustering

Introduction

Air pollution is now recognized as a major cancer risk factor worldwide (Chen et al., 2025; Kim et al., 2025; Li & Wang, 2024), with exposure to ambient pollutants such as particulate matter (PMx) and nitrogen oxides (NOx) associated with increased risks of various cancers and other health conditions (Chen et al., 2023; Liu et al., 2022; Nicolaou et al., 2025; Rammah et al., 2020; Tokuda et al., 2023; Zhuang et al., 2022). Globally, elevated concentrations of PM2.5 and CH4 contribute to premature mortality from lung cancer (Fang et al., 2013; Ma et al., 2023) . However, the relationships between air pollution and other cancer types (e.g., breast and bladder) remain unclear as observational studies have produced complex and often conflicting results (Li & Wang, 2024).

Texas presents a unique case study for examining air pollution-cancer relationships due to its distinctive industrial profile and cancer burden patterns. As the center of the US petroleum industry, Texas accounts for approximately 30% of national oil refining capacity and houses North America’s largest concentration of petrochemical plants (U.S. Energy Information Administration, 2025). This industrial infrastructure exposes residents to complex mixtures of hazardous air pollutants, including elevated concentrations of benzene, formaldehyde, 1,3-butadiene, and other known or suspected carcinogens (Cicalese et al., 2017). Consequently, Texas exhibits cancer incidence patterns that differ markedly from national averages, with particularly elevated rates of leukemia, myeloma, lymphoma, lung, and bladder cancers in counties with high industrial activity (Mungi et al., 2019). The spatial distribution of cancer incidence within Texas reflects these industrial exposure patterns. Hepatocellular carcinoma rates exceed the national average by 20–30%, with the highest incidence concentrated in petrochemical-intensive regions along the Gulf Coast (Cicalese et al., 2017). Breast cancer incidence shows geographic clustering near industrial facilities, with distance to emission sources significantly affecting risk levels (Madrigal et al., 2024). These patterns suggest that localized industrial emissions create cancer hotspots that could be identified and addressed through targeted public health interventions.

Traditional epidemiological approaches to studying air pollution-cancer relationships face significant methodological limitations at regional scales. Ground-based monitoring networks provide limited spatial coverage. Individual personal exposure measurement is resource-intensive and cannot be used for population-level studies. These constraints have historically limited research to either small-scale studies with detailed exposure data or large-scale studies with imprecise exposure proxies. It has also created a critical gap in understanding spatial pollution-health relationships. Remote sensing technology offers the potential for addressing these methodological limitations. Using observations from satellite-based platforms such as Sentinel-5 Precursor, researchers are able to observe important atmospheric constituents, such as ozone (O3), nitrogen dioxide (NO2), carbon monoxide (CO), sulfur dioxide (SO2), methane (CH4), formaldehyde (HCHO), aerosols, and clouds at various temporal and spatial resolutions (Amoroso et al., 2022; Bibault et al., 2020; Liu & Song, 2024; Veefkind et al., 2012). Recent applications have successfully used satellite data to model air pollution exposure for cardiovascular and respiratory health outcomes, though applications to cancer research remain limited (Amoroso et al., 2022; Elshorbany et al., 2021).

Machine learning provides effective methods for analyzing large and complex datasets acquired from remote sensing. Although traditional epidemiological models face difficulties with high-dimensional data, machine learning algorithms including Random Forest (RF), Support Vector Machines (SVM), Gaussian Naive Bayes, Artificial Neural Networks (ANN), gradient boosting, and k-nearest neighbors have been successfully applied to predict the effects of air pollution (Bellinger et al., 2017; Kumar & Pande, 2023; Liu et al., 2019; Masih, 2019; Vachon et al., 2024; Wang et al., 2020; Yarragunta et al., 2021). These approaches can quantify effects of specific pollutants on health outcomes and enhance spatial resolution in environmental and epidemiological analyses (Amoroso et al., 2022). Machine learning has also been applied extensively in cancer research, allowing for the prediction, diagnosis, and modeling of disease progression, decision-making, and treatment outcomes (Kourou et al., 2015). High-resolution satellite imagery and deep learning approaches are able to account for much of the spatial variation in cancer prevalence to support targeted interventions (Bibault et al., 2020).

Several critical gaps exist in understanding the relationship between air pollution and cancer prevalence in Texas. First, most studies have focused on specific metropolitan areas (i.e., Houston and Dallas) or individual pollutants rather than comprehensive statewide assessments. Second, traditional ground-based monitoring provides limited spatial coverage, creating gaps in exposure assessment for rural and suburban populations. Third, while numerous studies have examined individual cancer types (particularly lung cancer), few have investigated overall cancer prevalence patterns across diverse geographic and demographic contexts. Fourth, the application of satellite-based remote sensing data to cancer epidemiology remains underutilized despite its advantages for large-scale spatial analysis. Finally, most existing studies rely on linear statistical approaches that may not capture complex, non-linear interactions between multiple pollutants and socio-demographic factors.

This study aims to address these gaps by providing the first comprehensive, statewide assessment of associations between satellite-derived air pollutant concentrations and cancer prevalence across all 254 counties of Texas. Specifically, it has four major objectives. First, it analyzes the spatial distribution of air pollutant concentrations and cancer prevalence across Texas using remote sensing and geographic information systems. Second, it examines the statistical associations between satellite-derived air pollutants and county-level cancer prevalence using multiple machine learning approaches. Third, it identifies the most effective predictive modeling approach for capturing pollution-cancer relationships. Finally, it determines which air pollutants show the strongest associations with cancer prevalence to inform public health priorities. The remaining sections are organized as follows: the “Study area” section describes the study area, the “Methodology” section details the methodology, the “Results” and “Discussion” sections present results and discuss their implications, and the “Conclusions” section provides concluding remarks.

Study area

The entire state of Texas was chosen as the area of interest in this study (Fig. 1). The state is subdivided into 254 counties housing a population of approximately 30 million residents with the cities of Austin, Dallas-Fort Worth, and Houston having the nation’s highest annual population growth. The state has a very diversified economy, with large industries including energy (e.g., oil and natural gas), agriculture and mining, manufacturing technology, aerospace engineering, and healthcare. It consistently ranks among the top three states in the USA for non-fuel mineral production (Kyle & Elliott, 2019). Oil and gas extraction and petrochemical production activities are nearly three times higher in the state compared to the rest of the country (Assanie & Yücel, 2007). Consequently, the industrial activities in Texas are responsible for the emission of pollutants including CH4 (primarily from fracking), SO2, NOx, CO, PM1, and ground-level ozone, which are all considered major air pollutants in the world (Ridlington & Rumpler, 2013; Weinhold, 2012).

Fig. 1.

Fig. 1

Map of population density (people/km2) across the state of Texas

Methodology

Data collection

Three primary categories of data were collected for this study (Tables 1 and 2). First, mean air quality data for the entire 2022 were gathered from a suite of remote sensing satellite data products covering the entire state of Texas. Satellite-based datasets allowed for consistent, high-resolution spatial coverage and allowed for detailed regional analysis of atmospheric pollutant distributions. Specifically, tropospheric vertical column densities of NO2, CO, HCHO, and SO2 were extracted from Sentinel-5 Near Real-Time Instrument (NRTI) data of the COPERNICUS Program to represent concentrations of important gaseous pollutants.

Table 1.

Cancer prevalence and air pollutants variables description and source. All the data were collected during 2022

Variables Source and description Resolution/scale
Cancer prevalence

▪ Data from PLACES dataset, Centers for Disease Control and Prevention (CDC)

▪ This dataset contains model-based county-level estimates in GIS-friendly format

▪ Represent model-based estimate for age-adjusted prevalence of cancer (non-skin) or melanoma among adults

County level
NO2 (Mol/M2)

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/NRTI/L3)

▪ Tropospheric nitrogen dioxide (NO2) column number density

▪ Measure total nitrogen dioxide (NO2) in a vertical column of the atmosphere above each pixel

3.5 km × 7 km
CO (Mol/M2)

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/NRTI/L3)

▪ Tropospheric CO column number density

▪ Measure total carbon monoxide (CO) in a vertical column of the atmosphere above each pixel

3.5 km × 7 km
HCHO (mol/m2)

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/NRTI/L3)

▪ Tropospheric HCHO column number density

▪ Measure total formaldehyde (HCHO) in a vertical column of the atmosphere above each pixel

3.5 km × 7 km
SO2 (mol/m2)

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/NRTI/L3)

▪ Tropospheric vertical column density of SO2

▪ Measure total sulfur dioxide (SO2) in a vertical column of the atmosphere above each pixel

3.5 km × 7 km
AER AI

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/NRTI/L3)

▪ A dimensionless indicator of the presence of UV-absorbing aerosols

▪ Measure prevalence of coarse aerosols in the atmosphere by absorbing index;

3.5 km × 7 km
CH4 (mol/m2)

▪ Data from The European Union’s Earth observation program/Sentinel-5 Precursor/Level-3 product (COPERNICUS/S5P/OFFL/L3)

▪ Column volume mixing ratio

▪ Measure total methane (CH4) in a vertical column of the atmosphere above each pixel, corrected for water vapor

3.5 km × 7 km
PM1 (µg/m3)

▪ Data from European Centre for Medium-Range Weather Forecasts/Copernicus Atmosphere Monitoring Service/Near Real-Time data (ECMWF/CAMS/NRT)

▪ Particulate matter ≤ 1 µm; Measure surface concentration of PM1

40 km × 40 km
PM10 (µg/m3)

▪ Data from European Centre for Medium-Range Weather Forecasts/Copernicus Atmosphere Monitoring Service/Near Real-Time data (ECMWF/CAMS/NRT)

▪ Particulate matter ≤ 10 µm; measure surface concentration of PM10

40 km × 40 km

Table 2.

Social-demographic variables description and sources

Variables Source/description Resolution/scale Time
Senior population

▪ Data from Esri Enrichment

▪ Household with population 65+ (ACS 5 year)

County level 2022
Household income

▪ Data from Esri Enrichment

▪ Median Household Income (ACS 5 year)

County level 2022
Population

▪ Data from Esri Enrichment

▪ Population (Esri)

County level 2022
Housing affordability

▪ Data from Esri Enrichment

▪ Percentage of Households/Gross Rent 50% + of Income (ACS 5 year)

County level 2022
Insurance

▪ Data from Esri Enrichment

▪ Civilian noninstitutional population (ACS 5 year)

County level 2022
Consumed liquor

▪ Data from Esri Enrichment

▪ Percentage of consumed liquor at home/30 days (Esri)

County level 2024
Smoke

▪ Data from Esri Enrichment

▪ Percentage of smoked cigar in last 12 month

County level 2024

Aerosol prevalence was recorded by the Aerosol Absorbing Index (AERAI) from the Tropospheric Monitoring Instrument (TROPOMI) instrument on board Sentinel-5P satellite sensor product. This index measures the abundance and distribution of absorbing aerosols such as black carbon or desert dust. The state’s CH4 concentration was tracked with the Level 3 Offline Methane (CH4) data from Sentinel-5P dataset. Particulate matter concentrations were collected from the Near Real-Time (NRT) global atmospheric composition service provided by the Copernicus Atmosphere Monitoring Service (CAMS) from the European Centre for Medium-Range Weather Forecasts (ECMWF), which provide global, near-real-time indicators of PM1 and PM10 based on advanced atmospheric composition modeling.

The second category of data collected was related to the socio-demographic characteristics of the residents within the state. Detailed information including the percent of elderly population, household income, population density, housing affordability, insurance coverage, alcohol consumption, and smoking behavior were gathered and organized. These socio-demographic characteristics are known to be important factors for cancer risk and potential modifiers of environmental exposure (Li et al., 2019). Table 2 summarizes the socio-demographic variables included in the analysis together with descriptions and reference timeframes.

Finally, the county-level cancer prevalence data were taken from the PLACES: County Data (released in 2024). This dataset contains model-based, geospatial (GIS) compatible health estimates generated by the Centers for Disease Control and Prevention (CDC) and provides comprehensive and consistent information on health across county and census tract levels.

Data analysis

Once all data were collected, they underwent a series of processing steps. Remote sensing data were first acquired as raster layers using Google Earth Engine (GEE). The mean values of each raster layer at the county level were then calculated using the zonal statistics function in ArcGIS Pro v.3.3.0. These raster-derived values were compared and spatially analyzed alongside socio-demographic and cancer data for each county in Texas. Spatial autocorrelation of county-level cancer rates was assessed using Moran’s I. Next, multivariate cluster analysis was conducted to investigate air pollutant distribution patterns and identify factors with the greatest impact (Cicalese et al., 2017). Cancer prevalence statistics were subsequently summarized for each cluster.

Pearson’s correlation analysis was performed to detect multicollinearity among predictor variables. Any pairs of variables with a correlation coefficient greater than 0.7 were removed to reduce redundancy. The modeling algorithms RF, SVM, Multi-Layer Perceptron (MLP), and Ordinary Least Squares (OLS) were then applied. These four modeling approaches were selected to capture both linear and non-linear relationships. Specifically, the OLS regression was used as a baseline linear model. Random Forest was utilized for handling high-dimensional data and non-linear relationships. The SVM with radial kernel was used for modeling complex patterns with regularization. Finally, the MLP was performed to assess neural network–based performance. These algorithms span parametric to non-parametric to neural network approaches, enabling comprehensive performance assessment. Hyperparameters for all models were optimized using repeated tenfold cross-validation within the caret framework, with model performance evaluated by root mean squared error (RMSE). For the Random Forest model, a grid search was conducted over candidate values of the mtry parameter, while the number of trees was fixed at 1000 to ensure stable ensemble performance; the optimal configuration minimized the cross-validated RMSE. For the Support Vector Machine, candidate values of the cost (C) and radial kernel width (sigma) parameters were automatically generated (tuneLength = 10) and selected based on RMSE minimization. Similarly, the Multi-Layer Perceptron model was tuned using a grid search over the number of hidden neurons (size), with the optimal architecture determined by the lowest cross-validated RMSE. Feature selection was performed for RF and SVM models, while OLS regression, which does not require prior regularization, was used as a baseline. An ANOVA test compared OLS models with and without air pollutant variables to assess their contribution to cancer outcomes.

This study did not apply deep learning algorithms for model comparison. Although deep learning approaches have demonstrated strong predictive capabilities in air pollution studies (Gholami et al., 2021; Rabie et al., 2024), their effectiveness is often limited when sample sizes are small or data are spatially and temporally aggregated. Given that the data used in this study are aggregated at the county level, deep learning models were not considered appropriate for comparison with traditional machine learning algorithms. Moreover, deep learning models typically function as black boxes, offering limited interpretability and making it difficult to gain insight into variable relationships. Because a key objective of this study is to examine the impacts of air pollutants and identify the major contributing factors to health outcomes, deep learning approaches are less suitable for achieving these interpretive goals.

For modeling, the dataset was split into a training set (70%) and a testing set (30%). The performance of each model was evaluated and compared using root mean square error (RMSE) and mean absolute error (MAE). For data preprocessing, the zonal air pollutant data was joined to the shapefile that had already been merged with CDC cancer prevalence data and socio-demographic variables. The resulting shapefile was then imported into RStudio for processing. After dropping the geometry, the dataset’s outlier—Loving County—was removed. A Box–Cox transformation was applied to reduce skewness and improve normality. The data was then centered by subtracting the mean of each numeric variable and scaled so that each variable had a standard deviation of 1. Finally, near-zero variance predictors were removed from the dataset.

Permutation importance analyses were then conducted to identify the most influential predictors (including both socio-demographic and environmental factors) in the selected SVM model. Additional analyses were performed using models trained separately with only remote sensing variables or only socio-demographic variables to assess their relative contributions. All analyses were conducted in R using the caret package, with a 10-time repeated tenfold cross-validation strategy applied across all models to ensure robust evaluation.

Results

Spatial patterns of air quality

The distribution of the individual pollutants across the state is shown in Fig. 2. The NO2, CO, and SO2 are concentrated in the eastern and southeastern parts of Texas near oil refining and power generation facilities. The CH4 levels are highest in the west and southeastern parts of the state and around the cities of Austin, San Antonio, Dallas, and Houston (due to oil and gas extraction and processing sites). The HCHO increases in central and east Texas, reflecting combined industrial emissions and photochemical activities. AER AI is high in the western region whereas PM1 and PM10 are high in the southern, eastern, and central regions, indicative of higher particulate pollution.

Fig. 2.

Fig. 2

Spatial distribution of various air pollutants across the Texas counties

The results from the clustering analysis of CO, SO2, NO2, HCHO, CH4, AER AI, PM1, and PM10 using k-means with 3 clustering groups are shown in Figs. 3 and 4. These clusters highlight areas with similar air quality characteristics which in turn reflect underlying population density, industrial activity, and geographic conditions. The green cluster is characterized by lower levels of CO, HCHO, PM1, PM10, and NO2, and fluctuating levels for CH4. This region has sparse population and few industries. In the orange cluster, uniformly elevated values for most of the pollutants except for CH4 and AER AI indicate high pollution levels for this region due to urbanization and the region’s industrial activity. Moderate levels (in-between the other two clusters) of pollutants are found in the central region of Texas and are shown as the purple cluster. Interestingly, areas of high CH4 and AER AI are in places with fewer other pollutants, suggesting air quality is affected by different sources and processes in the region.

Fig. 3.

Fig. 3

Spatial distribution of air quality clusters across the Texas counties

Fig. 4.

Fig. 4

Boxplots showing standardized values of atmospheric variables (AER AI, CH4, CO, HCHO, NO2, PM1, PM10, and SO2), overlaid with cluster-specific mean trends for groups 1–3

Spatial analysis of cancer prevalence

The prevalence of cancer throughout the state is shown in Fig. 5. The highest rates (7.2–7.7%) occurred in northwestern counties (Carson, Roberts, and Randall) and northeastern counties (Hopkins, Wood, and Grayson), with moderate rates (6.6–7.1%) in surrounding areas. The lowest rates (4.6–5.9%) were clustered in southern and western Texas. The Moran’s I value of approximately 0.69 with Z score of 21.58 and p value of 0 indicated a strong positive spatial autocorrelation. The Z score was far beyond common significance thresholds; the p value indicated that the observed spatial pattern is not random and is statistically significant. Together, they confirm that counties with high cancer prevalence are geographically clustered rather than randomly distributed across the state.

Fig. 5.

Fig. 5

Map of cancer prevalence (in percentage of adult population) in Texas counties

Spatial overlap between pollution and cancer patterns was evident. In northwest and northeast Texas (areas already identified as having high levels of SO2, CO, and particulate matter), higher cancer rates were also observed. The Gulf Coast region exhibits lower cancer prevalence (may be associated with reduced levels of AER AI and PM1). It is also evident that larger urban centers, such as Dallas and Houston, display clear cancer clusters, indicating that additional factors (e.g., urbanization and behavior) may influence cancer prevalence.

The statistical summaries by cluster are shown in Table 3. When combined with the map in Fig. 3, they illustrate the regional differences in cancer prevalence. The East Texas cluster (high pollution, low AER AI) showed notably higher cancer prevalence (mean = 6.78%, median = 6.80%) compared to the moderate-pollution Central Texas cluster (mean = 6.40%, median = 6.60%) and the West Texas cluster with high AER AI (mean = 6.28%, median = 6.30%).

Table 3.

Statistical mean of cancer prevalence of clusters in Texas

Cluster Cancer prevalence mean Cancer prevalence median
West Texas with low pollution but high AER AI 6.28 6.30
East Texas with high pollution but low AER AI 6.78 6.80
Central Texas with moderate pollution 6.40 6.60

Correlation analysis

The correlation values between prevalence of cancer, pollutants, and socioeconomic variables are given in Fig. 6. The incidence of cancer is weakly correlated with most individual pollutants. The cancer prevalence and CO have a weak negative correlation (−0.16), weak positive correlation (0.01) with NO2, CH4 (0.11), and AER AI (0.21). Particulate matter variables (PM10 and PM1) are highly correlated (0.63) with each other. The CO with PM10, HCHO, and AER AI are also highly correlated with a possibility of high degree of multicollinearity. These correlation values indicate that while pollutants may work to increase cancer risk, their effects are subtle and potentially mediated by other variables. Among the socio-demographic variables, population density and insurance coverage are highly correlated (r = 0.94). In order to avoid multicollinearity and distortion in modeling, CO, HCHO, PM10, and population density were excluded from modeling considerations.

Fig. 6.

Fig. 6

Correlation matrix of cancer prevalence with environmental pollutants (CO, NO2, SO2, CH4, HCHO, AER AI, PM1, PM10) and socioeconomic variables

Machine learning modeling comparison

ANOVA testing compared OLS models with and without air pollutant variables with resulting p < 0.05 and confirming that air pollutants significantly improve model fit and are associated with cancer prevalence. Prediction accuracy and explanatory power were compared via testing metrics, and the results are shown in Table 4. The SVM outperformed the other models in terms of RMSE (RMSE = 0.347). The MLP and OLS also achieved low RMSE (0.348 and 0.349) values. In contrast, RF showed the weakest performance (RMSE = 0.393, R2 = 0.748). In terms of explanatory power, SVM achieved the highest R2 (0.791), followed by MLP (0.772) and OLS (0.748). RF regression performed the weakest overall, with both the highest RMSE and the lowest R2 (0.748). These findings suggest that while both OLS, MLP, and SVM excel in prediction accuracy, only SVM demonstrates strong capability in capturing variations, making it the most balanced and effective model among those evaluated.

Table 4.

Comparison of model performance metrics, including RMSE and R2, for the four predictive modeling approaches used in the study

Model Hyper parameter Test RMSE Test R-square
Random Forest Regression mtry = 6; ntree = 1000 0.393 0.748
Ordinary Least Square 0.349 0.748
Support Vector Machine sigma = 0.07; C = 1 0.347 0.791
Multilayer Perceptron size = 3 0.348 0.772

Models were also statistically evaluated using cross-validation metrics to assess their accuracy and stability. The combined MAE and RMSE summary boxplot is shown in Figs. 7 and 8. They highlight significant differences in predictive performance among the four machine learning models. The SVM with radial consistently achieved the lowest median MAE (0.25) and second lowest median RMSE, indicating the best predictive performance among all models in cross-validation. Both RF and OLS also demonstrated competitive performance in terms of cross-validated prediction accuracy. In contrast, the Multi-Layer Perceptron (MLP) exhibited higher median RMSE (0.39) and median MAE (0.30), along with greater variance, suggesting reduced accuracy and stability. Overall, the SVM emerged as the most effective model, offering the best balance between precision and reliability. MLP’s elevated variance further underscored its limited stability. In contrast, RF showed competitive cross-validation performance but notable discrepancy with test results, indicating potential overfitting.

Fig. 7.

Fig. 7

RMSE of the four different machine learning models examined in the study

Fig. 8.

Fig. 8

MAE of the four different machine learning models examined in the study

Figure 9 presents scatterplots comparing predicted versus observed cancer prevalence across four regression models: RF Regression, OLS, SVM, and MLP. Among these, the SVM model demonstrates the closest alignment with the 1:1 reference line. Most data points are tightly clustered around it, indicating high predictive accuracy and minimal bias. The OLS model also shows strong alignment.

Fig. 9.

Fig. 9

Predicted vs. observed from different machine learning models

The RF model generally follows the reference line but exhibits greater dispersion, particularly in the mid-to-high observed values (6.0–7.5). This suggests moderate predictive performance with instances of both overestimation and underestimation. In contrast, the MLP model displays the greatest deviation from the reference line. It tends to underpredict lower values and cluster excessively at the upper end, which points to reduced accuracy and signs of overfitting or instability.

The boxplots of MAE and RMSE from cross-validation (Figs. 10 and 11) provide a clear comparison of model performance, illustrating differences in predictive accuracy and stability across modeling approaches. Among all models, the SVM models (especially those using the full set of predictors) achieved the lowest prediction errors with minimal variance (median MAE values around 0.25–0.27), indicating high accuracy and robustness. The RF models also performed strongly (especially when trained on the full dataset with median MAE approximately 0.27–0.28). However, they demonstrated comparatively better performance than other models when restricted to socio-demographic variables alone (median MAE around 0.30–0.32).

Fig. 10.

Fig. 10

RMSE comparison between models (full model vs. air pollutant only vs. social features only)

Fig. 11.

Fig. 11

MAE comparison between models (full model vs. air pollutant only vs. social features only)

In contrast, the MLP models consistently underperformed, yielding the highest RMSE and MAE values along with substantial variance (primarily when trained on the full or social-only datasets with median MAE values ranging from 0.30 to 0.35 and notable outliers extending to 0.60–0.90). While the MLP slightly outperformed OLS in the air-only configuration for MAE (median MAE around 0.35 compared to 0.38 for OLS), its higher variability suggests instability in predictions. OLS performed competitively when using the full predictor set (median MAE approximately 0.28–0.30), confirming that even simpler linear models benefit from incorporating diverse data inputs. Models using air-only features generally showed the poorest performance, with median MAE values ranging from 0.32 to 0.38 across different approaches.

Overall, SVM emerged as the most reliable and accurate modeling approach for this dataset. It offers the best balance between predictive accuracy and stability. Furthermore, models trained exclusively on air pollutant variables were markedly less effective than those incorporating both environmental and socio-demographic factors. These results underscore the importance of adopting a multidimensional modeling framework when analyzing health outcomes, as integrating multiple determinants yields more accurate and insightful predictions.

Permutation importance analysis

The results from the permutation importance analysis identifying the top ten key drivers of cancer prevalence are shown in Fig. 12. Socioeconomic factors emerged as the most influential, with liquor consumption ranking highest (importance value of approximately 0.21) and household income ranking third (importance of approximately 0.08), suggesting that socioeconomic status is a meaningful contributor to health predictions among all the factors. Among the environmental factors, SO2 demonstrated moderate importance (importance value of approximately 0.13). The AER AI and CH4 (importance value of about 0.08) and PM10 (importance value of approximately 0.06) showed moderate importance. Social-demographic variables such as insurance coverage (importance value almost 0.08) and the presence of seniors in households (importance value around 0.07) also demonstrated moderately important factors. Finally, smoking rate and NO2 showed lowest importance (importance value of between 0.04 and 0.05). Overall, the results reveal that behavioral factors like alcohol consumption, followed by a combination of air quality indicators and socioeconomic variables play a major role in having cancer. Interestingly, traditional risk factors like smoking were found to have lower importance in this analysis.

Fig. 12.

Fig. 12

Permutation importance of top ten predictors for cancer prevalence, showing relative contribution of socio-demographic and environmental variables

Discussion

This study represents one of the first comprehensive assessments of associations between satellite-derived air pollution data and cancer prevalence across Texas. Three key findings emerged. First, spatial analysis revealed distinct regional pollution patterns, with East Texas experiencing elevated concentrations of most pollutants while West Texas showed high AER AI levels despite otherwise cleaner air. Second, machine learning models, particularly SVM, effectively captured complex associations between environmental and socio-demographic factors and cancer prevalence. It significantly outperformed models using only air pollutants or only socio-demographic variables. Third, SO2 emerged as the most influential environmental predictor among air pollutants examined, followed by AER AI and CH4, and PM10, though socio-demographic factors (especially alcohol consumption and household income) demonstrated stronger overall predictive importance.

Spatial distribution patterns of air pollutants and cancer prevalence

The spatial distribution of air pollution revealed three distinct clusters across Texas. West Texas exhibited relatively low pollution levels except for elevated AER AI, a region with sparse population density. High AER AI values indicate strong concentrations of UV-absorbing aerosols, typically linked to desert dust storms, biomass burning, and industrial emissions. This suggests that UV-absorbing aerosols should be a primary focus of environmental protection policies for western Texas cities. East Texas demonstrated relatively poor air quality compared to the state average, with elevated levels of most pollutants except AER AI. This region is highly populated and industrialized, with parts reporting higher prevalence of cancer. The overlap between poorer air quality and elevated cancer prevalence aligns with previous research highlighting the role of urban air pollution in increasing cancer risk (Hemminki & Pershagen, 1994).

However, cancer prevalence is inherently complex and influenced by multiple factors. The high Moran’s I value (0.69) for cancer prevalence indicated strong spatial autocorrelation, confirming that counties with similar cancer prevalence rates cluster geographically rather than being randomly distributed. The relationships among pollutants further complicate the interpretations. Particulate matter showed a negative correlation with AER AI, while CO demonstrated a strong positive correlation with both AER AI and particulate matter. HCHO was negatively correlated with AER AI. These interdependencies suggest complex pollutant interaction structures that may influence cancer prevalence through non-additive pathways. Given these correlations and their modeling significance, CO and HCHO require further investigation.

SO2 was identified as the most influential air pollutant affecting cancer prevalence from permutation importance analysis. This result reinforces various previous studies on the impact of SO2 on cancer, especially lung cancer (Johns & Linn, 2011; Tseng et al., 2019). However, it should be emphasized that this statistical association does not establish causation. The elevated SO2 concentrations in east Texas (where oil refining and power generation are concentrated) may serve as a marker for broader industrial pollution exposure rather than indicating SO2 specific carcinogenic effects. This result also reflects the spatial distribution pattern in northwest Texas, where both cancer prevalence and SO2 levels are elevated. Notably, AER AI and SO2 were the only two air pollutants with markedly high concentrations in this region. Counties with high SO2 levels likely experience elevated concentrations of multiple hazardous established carcinogenic air pollutants not captured in this analysis (i.e., benzene, 1,3-butadiene, and polycyclic aromatic hydrocarbons) (Cicalese et al., 2017). The co-occurrence of multiple pollutants in industrial areas makes it difficult to isolate the specific contribution of any single compound.

Particulate matter and AER AI showed moderate influence on cancer prevalence. Both have been widely demonstrated to pose cancer risks (Corrêa et al., 2007; Deng et al., 2017; Santibáñez-Andrade et al., 2019). While lung cancer has been extensively studied in relation to air pollutants, other cancer types have also been linked to these compounds. For example, particulate matter has been associated with increased risk for liver cancer (Deng et al., 2017). Beyond cancer, well-established epidemiological evidence identifies particulate matter as the predominant contributor to adverse health outcomes (Bae & Hong, 2018; Schwarze et al., 2006). Although this study found SO2 to be the most influential air pollutant rather than particulate matter, this result is not contradictory. Epidemiological studies typically examine direct, individual-level relationships between specific pollutants and health outcomes across thousands of cases. In contrast, our spatial analysis evaluates broader associations using county-level aggregated data. As such, the results reflect regional averages rather than individual exposure pathways. Spatial patterns can also be shaped by population distribution, geographic features, and administrative boundaries and may be further influenced by spatial autocorrelation. Consequently, SO2 may emerge as a more prominent predictor at the regional scale–informing policy decisions aimed at addressing area-wide environmental health issues, while particulate matter may remain the more critical factor for individual-level health risks.

The moderate predictive importance of CH4 requires careful interpretation. Limited evidence directly links methane to carcinogenesis, suggesting that CH4 concentrations likely serve as a proxy for oil and gas extraction activities, which involve exposure to multiple carcinogenic compounds. The finding that CH4 shows moderate importance while demonstrating low correlation with other pollutants suggests it captures unique aspects of industrial activity not represented by other variables. As noted by Özlü and Yalçin (2024), associations between methane production and cancer mortality may reflect co-pollutant exposures rather than methane-specific effects. This interpretation is particularly relevant in Texas, where oil and gas extraction is a dominant industry and methane leakage from these operations is significant. Alvarez et al. (2018) asserted that 1.2% of methane escapes during extraction and use. Increasing development of fracking and pipeline transport of natural gas is highly associated with methane increases (Russo & Carpenter, 2019). More specific spatial analysis of the oil and gas industry locations and cancer prevalence is necessary to uncover the association with CH4.

Machine learning performance and non-linear relationships

The comparative performance of machine learning models revealed important insights into the nature of associations between air pollution, socio-demographic factors, and cancer prevalence. The ANOVA test confirmed that air pollution variables significantly improve model fit (F = 3.84, p = 0.005), demonstrating that environmental exposures provide explanatory power beyond socio-demographic factors alone. This finding has important policy implications: while addressing socioeconomic determinants of health remains critical, environmental interventions targeting air quality can provide complementary pathways for cancer risk reduction. The statistical significance of air pollutant contributions reinforces conclusions from previous regional studies (Pinakana et al., 2023; Sexton et al., 2007) while extending evidence to the statewide scale.

Another key finding from this study was that the spatial associations between cancer prevalence, air pollution, and socioeconomic factors involve a mix of linear and non-linear patterns. The overall performance of OLS regression with socio-demographic variables and air pollutants was comparable to the best non-linear model (SVM with radial kernel) and better than RF, across both cross-validation and test datasets. However, closer examination revealed that OLS performed substantially worse when only air pollutant variables were included (MAE = 0.38 vs. 0.35 for SVM, 0.35 for MLP), suggesting that the spatial associations between air pollution and cancer prevalence are more non-linear than those of socioeconomic variables.

This study identified SVM as the best-performing model both in cross-validation and testing. However, it differs from findings by Amoroso et al. (2022) who found RF performed best when examining associations between NO2 and COVID-19 casualties. This discrepancy may be attributable to several factors. First, the outcome variables are different (cancer prevalence vs. COVID-19 mortality) for each study. Second, different sample sizes and spatial scales along with different temporal dynamics (cancer with long latency vs. acute COVID-19 outcomes) may affect the model performances. Finally, since RF requires larger datasets to demonstrate its full potential, our relatively small sample of 253 counties may reduce RF’s ability to predict accurately in the testing dataset. Still, both studies selected non-linear machine learning models (SVM and RF) as top performers. Similarly, Madukpe et al. (2025) showed that linear models inadequately represent non-linear air pollution dynamics, while non-linear methods, including LSTM and ANNs, achieved superior performance. These demonstrate the capability of non-linear approaches to capture complexity in the associations between air pollutants and health outcomes.

Both deep learning and machine learning models have been applied in air pollution study (Gholami et al., 2021; Rabie et al., 2024; Tao et al., 2023; Wood, 2022). Although deep learning models have demonstrated strong predictive capabilities for air pollution (Rabie et al., 2024) and dust emission (Gholami et al., 2021), they typically require very large training datasets. For example, 37,044 records were used to train a CNN–Bi-LSTM model that achieved 97% accuracy in Rabie et al. (2024). These findings suggest that deep learning is most effective for pure air pollution prediction tasks when large volumes of station-based data are available. In contrast, the present study relies on annually aggregated cancer prevalence data and satellite-derived air pollution variables, for which sample sizes are insufficient to support robust deep learning training. Therefore, machine learning algorithms represent a more appropriate and reliable modeling framework for this study.

Remote sensing for air quality assessment

Remote sensing provides a cost-effective approach to studying air pollution across large geographic regions where ground-based monitoring is sparse or absent. Compared to ground-based point measurements, Sentinel-5P (3.5 km × 7 km) and ECMWF (40 km × 40 km) data has an admittedly coarse resolution. Nevertheless, they were appropriate at the county-level analysis in Texas (especially where there are no extensive ground-based monitoring networks). Existing validation studies have shown that satellite-based air pollutant concentrations have high correlations with ground-based measurements (Kloog et al., 2012; Van Donkelaar et al., 2010), supporting the reliability of this approach for regional-scale analysis.

Nevertheless, satellite-derived measurements have several sources of uncertainties that must be taken into consideration when interpreting the results. First, column measurements (which integrate pollutant concentrations throughout the atmospheric column) may not accurately reflect ground-level exposures (especially in areas with complex topography or vertical stratification of pollutants). This limitation is especially relevant for pollutants like SO2 and NO2 that can exhibit substantial vertical gradients. Second, cloud coverage creates temporal gaps in optical sensor data. Finally, temporal coverage is limited to satellite overpass times (usually once or twice daily). This could fail to represent diurnal pollution trends or acute exposure to pollution. It also might fail to record irregular events that might be toxicologically significant.

The spatial and temporal resolution constraints currently limit investigations at census tract or neighborhood level. The exposure at these spatial levels may be significant since urban residents within a county may experience markedly different pollution exposures than the county’s rural residents. However, county-average values cannot capture these important spatial variations. Due to these limitations, satellite remote sensing applications to cancer epidemiology remain underutilized (Amoroso et al., 2022; Elshorbany et al., 2021). This study demonstrates that satellite-derived pollution data can provide comprehensive spatial coverage for statewide cancer prevalence analysis. Thus, it addresses critical gaps in understanding regional exposure patterns that ground-based monitoring alone cannot achieve.

Future research would benefit from several methodological advances. First, incorporating cancer type–specific incidence data would help identify pollutant-specific effects (usually hidden by aggregate cancer prevalence outcomes). Different cancer types have varying latency periods and susceptibility to specific pollutants. Disaggregated analysis could reveal associations obscured in overall prevalence measures. Second, validating satellite-derived estimates against ground-based monitoring data (especially in high-pollution industrial corridors and urban centers) would help quantify measurement uncertainties and improve exposure assessment accuracy. Third, exploring ensemble machine learning approaches and spatial regression models that take into consideration geographic autocorrelation could improve prediction accuracy. They can also capture spatial dependencies in both pollution and health outcomes. Fourth, as higher-resolution satellite data becomes available in the future, it would be possible to analyze the data at census tract or neighborhood scale at finer scale to investigate within-county exposure gradients. Finally, time-series analysis incorporating multi-year satellite observations could offer valuable insights into temporal dynamics of pollution-health associations.

In many pollutions and health analyses, the variations in indoor air quality standards across different regions of the state could influence individual exposure estimates. Texas does not have region-specific indoor air quality (IAQ) standards. The IAQ standard is largely unregulated. The state only has statewide guidelines for government buildings (Texas Department of Health, 2003). Thus, in this study area, the variation of standards does not impact our statistical analysis and interpretation.

Limitations of the study

This study has a few limitations. First, it is a cross-sectional, ecological study that examines the associations on the county level. It cannot determine causality or determine exposure-disease relationships on an individual level. The county-level analysis also masks important within-county variation in both pollution exposure and cancer risk. Urban residents within a county may experience markedly different exposures than rural residents, but county-average pollution values cannot capture these gradients. Second, the temporal mismatch between 2022 air quality data and cancer prevalence (which reflects exposures over preceding decades) limits our ability to capture long-term exposure effects. Third, cancer prevalence is measured as an aggregate outcome across all cancer types, potentially obscuring pollutant-specific effects on cancers. Despite these limitations, this study provides valuable insights into spatial patterns and can guide more targeted investigations.

Conclusions

The study is the first statewide assessment of relationships among satellite-derived air pollutants and cancer prevalence across 254 Texas counties. It also demonstrates the usefulness of remote sensing combined with machine learning in environmental health research at the regional level. The spatial analysis results showed specific patterns of pollution. West Texas showed elevated UV-absorbing aerosols but otherwise clean air, East Texas exhibited uniformly high concentrations of SO2, NO2, CO, and particulate matter with correspondingly higher cancer prevalence (mean = 6.78%), and Central Texas displayed moderate levels. Strong spatial autocorrelation of cancer prevalence (Moran’s I = 0.69) confirmed geographic clustering rather than random distribution and confirms that cancer risk is dependent upon the local environment.

The use of machine learning analysis showed that air pollutants significantly improve predictive models beyond socio-demographic factors alone (ANOVA: F = 3.84, p = 0.005). Support Vector Machine achieved superior performance (RMSE = 0.347, R2 = 0.791), effectively capturing non-linear pollution-health relationships. Permutation importance identified SO2 as the most influential environmental predictor and likely serving as a marker for broader industrial pollution exposure, followed by AER AI, CH4, and particulate matter.

These findings have important policy implications. The significant contribution of air pollutants confirms that environmental quality represents a modifiable cancer risk factor warranting targeted interventions (particularly in East Texas). The importance of both environmental and socio-demographic factors demonstrates that effective prevention requires integrated approaches addressing air quality and healthcare access simultaneously. Regional pollution clusters provide evidence-based guidance for resource prioritization, while high spatial autocorrelation suggests place-based interventions targeting geographic hotspots may be more efficient than uniform statewide programs.

This study advances environmental health methodology by showing that satellite remote sensing can provide valuable exposure data for regional cancer epidemiology in spite of resolution constraints. Nevertheless, the study has few limitations including the cross-sectional design precluding causal inference, county-level analysis masking within-county gradients, and temporal mismatch between 2022 pollution data and cancer prevalence reflecting decades of exposure. Future studies should incorporate cancer type–specific data, analyze multi-year satellite time series, validate against ground-based measurements, and conduct finer-scale analyses as higher-resolution products become available. In conclusion, satellite-derived pollution data analyzed with machine learning can effectively identify regional cancer risk patterns and inform targeted interventions. The spatial clustering of cancer prevalence in Texas, combined with significant pollution-health associations, underscores the need for comprehensive environmental health policies addressing both pollution sources and vulnerable populations.

Acknowledgements

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. The authors have no competing interests.

Author contribution

All authors contributed to the study conception and design. Material preparation, data collection, and analysis were performed by Fangchao Dong and Hao Chen. Fangchao Dong and Muhammad Tauhidur Rahman drafted and finalized the manuscript with input from all authors. All authors reviewed, commented on, and approved the final manuscript.

Funding

The authors declare that no funding was received to assist with the preparation of this manuscript.

Data availability

All data generated or analyzed during this study will be available upon request.

Declarations

Ethical responsibilities of authors

All authors have read, understood, and have complied as applicable with the statement on “Ethical responsibilities of Authors” as found in the Instructions for Authors.

Ethics approval

This is an observational study. Therefore, no ethical approval is required.

Competing interests

The authors declare no competing interests.

Clinical trial number

Not applicable.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Fangchao Dong, Email: fangchao.dong@utdallas.edu.

Muhammad Tauhidur Rahman, Email: mtr@utdallas.edu.

References

  1. Alvarez, R. A., Zavala-Araiza, D., Lyon, D. R., Allen, D. T., Barkley, Z. R., Brandt, A. R., et al. (2018). Assessment of methane emissions from the U.S. oil and gas supply chain. Science,361(6398), 186–188. 10.1126/science.aar7204 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Amoroso, N., Cilli, R., Maggipinto, T., Monaco, A., Tangaro, S., & Bellotti, R. (2022). Satellite data and machine learning reveal a significant correlation between NO2 and COVID-19 mortality. Environ Res,204, Article 111970. 10.1016/j.envres.2021.111970 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Assanie, L., & Yücel, M. (2007). Industry clusters shape Texas economy. Federal Reserve Bank of Dallas. https://www.dallasfed.org/~/media/documents/research/swe/2007/swe0705b.pdf. Accessed 20 Feb 2026.
  4. Bae, S., & Hong, Y.-C. (2018). Health effects of particulate matter. Journal of the Korean Medical Association,61(12), Article 749. 10.5124/jkma.2018.61.12.749 [Google Scholar]
  5. Bellinger, C., Mohomed Jabbar, M. S., Zaïane, O., & Osornio-Vargas, A. (2017). A systematic review of data mining and machine learning for air pollution epidemiology. BMC Public Health,17(1), Article 907. 10.1186/s12889-017-4914-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Bibault, J.-E., Bassenne, M., Ren, H., & Xing, L. (2020). Deep learning prediction of cancer prevalence from satellite imagery. Cancers,12(12), Article 3844. 10.3390/cancers12123844 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Chen, M.-J., Leon Guo, Y., Lin, P., Chiang, H.-C., Chen, P.-C., & Chen, Y.-C. (2023). Air quality health index (AQHI) based on multiple air pollutants and mortality risks in Taiwan: Construction and validation. Environmental Research,231, Article 116214. 10.1016/j.envres.2023.116214 [DOI] [PubMed] [Google Scholar]
  8. Chen, Y., Yang, S., Lin, J., Gu, S., Wu, L., Huang, W., et al. (2025). Long-term exposure to ambient air pollutants and risk of prostate cancer: A prospective cohort study. Environmental Research,270, Article 121020. 10.1016/j.envres.2025.121020 [DOI] [PubMed] [Google Scholar]
  9. Cicalese, L., Curcuru, G., Montalbano, M., Shirafkan, A., Georgiadis, J., & Rastellini, C. (2017). Hazardous air pollutants and primary liver cancer in Texas. PLoS ONE,12(10), Article e0185610. 10.1371/journal.pone.0185610 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Corrêa, M. P., Dubuisson, P., & Plana-Fattori, A. (2007). An overview of the ultraviolet index and the skin cancer cases in Brazil¶. Photochemistry and Photobiology,78(1), 49–54. 10.1562/0031-8655(2003)0780049AOOTUI2.0.CO2 [Google Scholar]
  11. Deng, H., Eckel, S. P., Liu, L., Lurmann, F. W., Cockburn, M. G., & Gilliland, F. D. (2017). Particulate matter air pollution and liver cancer survival. International Journal of Cancer,141(4), 744–749. 10.1002/ijc.30779 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Elshorbany, Y. F., Kapper, H. C., Ziemke, J. R., & Parr, S. A. (2021). The status of air quality in the United States during the COVID-19 pandemic: A remote sensing perspective. Remote Sensing,13(3), Article 369. 10.3390/rs13030369 [Google Scholar]
  13. Fang, Y., Naik, V., Horowitz, L. W., & Mauzerall, D. L. (2013). Air pollution and associated human mortality: the role of air pollutant emissions, climate change and methane concentration increases from the preindustrial period to present. Atmospheric Chemistry and Physics,13(3), 1377–1394. 10.5194/acp-13-1377-2013
  14. Gholami, H., Mohamadifar, A., Rahimi, S., Kaskaoutis, D. G., & Collins, A. L. (2021). Predicting land susceptibility to atmospheric dust emissions in central Iran by combining integrated data mining and a regional climate model. Atmospheric Pollution Research,12(4), 172–187. 10.1016/j.apr.2021.03.005 [Google Scholar]
  15. Hemminki, K., & Pershagen, G. (1994). Cancer risk of air pollution: Epidemiological evidence. Environmental Health Perspectives,102(4), 187–192. 10.1289/ehp.94102s4187 [Google Scholar]
  16. Johns, D. O., & Linn, W. S. (2011). A review of controlled human SO2 exposure studies contributing to the US EPA integrated science assessment for sulfur oxides. Inhalation Toxicology,23(1), 33–43. 10.3109/08958378.2010.539290 [DOI] [PubMed] [Google Scholar]
  17. Kim, C. H., Kim, S., Kang, Y.-H., Kim, S., Kim, B., & Park, B. (2025). Association between long-term exposure to a mixture of ambient air pollutants and the incidences of bladder and kidney cancers. Environmental Research,285, Article 122667. 10.1016/j.envres.2025.122667 [DOI] [PubMed] [Google Scholar]
  18. Kloog, I., Nordio, F., Coull, B. A., & Schwartz, J. (2012). Incorporating local land use regression and satellite aerosol optical depth in a hybrid model of spatiotemporal PM2.5 exposures in the Mid-Atlantic States. Environmental Science & Technology,46(21), 11913–11921. 10.1021/es302673e [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Kourou, K., Exarchos, T. P., Exarchos, K. P., Karamouzis, M. V., & Fotiadis, D. I. (2015). Machine learning applications in cancer prognosis and prediction. Computational and Structural Biotechnology Journal,13, 8–17. 10.1016/j.csbj.2014.11.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Kumar, K., & Pande, B. P. (2023). Air pollution prediction with machine learning: A case study of Indian cities. International Journal of Environmental Science and Technology,20(5), 5333–5348. 10.1007/s13762-022-04241-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Kyle, J. R., & Elliott, B. A. (2019). Past, present, and future of Texas industrial minerals. Mining, Metallurgy & Exploration,36(2), 475–486. 10.1007/s42461-019-0050-1 [Google Scholar]
  22. Li, W., & Wang, W. (2024). Causal effects of exposure to ambient air pollution on cancer risk: Insights from genetic evidence. Science of the Total Environment,912, Article 168843. 10.1016/j.scitotenv.2023.168843 [DOI] [PubMed] [Google Scholar]
  23. Li, Z., Konisky, D. M., & Zirogiannis, N. (2019). Racial, ethnic, and income disparities in air pollution: A study of excess emissions in Texas. PLoS ONE,14(8), Article e0220696. 10.1371/journal.pone.0220696 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Liu, C., Cai, J., Chen, R., Sera, F., Guo, Y., Tong, S., et al. (2022). Coarse particulate air pollution and daily mortality: A global study in 205 cities. American Journal of Respiratory and Critical Care Medicine,206(8), 999–1007. 10.1164/rccm.202111-2657OC [DOI] [PubMed] [Google Scholar]
  25. Liu, C., & Song, W. (2024). Mapping property redevelopment via GeoAI: Integrating computer vision and socioenvironmental patterns and processes. Cities,144, Article 104644. 10.1016/j.cities.2023.104644 [Google Scholar]
  26. Liu, H., Li, Q., Yu, D., & Gu, Y. (2019). Air quality index and air pollutant concentration prediction based on machine learning algorithms. Applied Sciences,9(19), Article 4069. 10.3390/app9194069 [Google Scholar]
  27. Ma, C., Jung, C.-R., Nakayama, S. F., Tabuchi, T., Nishihama, Y., Kudo, H., et al. (2023). Short-term association of air pollution with lung cancer mortality in Osaka, Japan. Environmental Research,224, 115503. 10.1016/j.envres.2023.115503
  28. Madrigal, J. M., Pruitt, C. N., Fisher, J. A., Liao, L. M., Graubard, B. I., Gierach, G. L., et al. (2024). Carcinogenic industrial air pollution and postmenopausal breast cancer risk in the National Institutes of Health AARP Diet and Health Study. Environment International,191, Article 108985. 10.1016/j.envint.2024.108985 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Madukpe, V. N., Ugoala, B. C., & Zulkepli, N. F. S. (2025). Topological approach and kernel principal component analysis for air pollution source apportionment. International Journal of Environmental Research,19(6), Article 260. 10.1007/s41742-025-00912-6 [Google Scholar]
  30. Masih, A. (2019). Machine learning algorithms in air quality modeling. Global Journal of Environmental Science and Management, 5(4). 10.22034/GJESM.2019.04.10
  31. Mungi, C., Lai, D., & Du, X. L. (2019). Spatial analysis of industrial benzene emissions and cancer incidence rates in Texas. International Journal of Environmental Research and Public Health,16(15), Article 2627. 10.3390/ijerph16152627 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Nicolaou, L., Rowell, E., Gaviola, C., Chandyo, R. K., Sharma, A. K., Shrestha, L. P., et al. (2025). Personal exposures to air pollutants and respiratory health among brick kiln workers and household members. Environmental Research,279, Article 121760. 10.1016/j.envres.2025.121760 [DOI] [PubMed] [Google Scholar]
  33. Özlü, C., & Yalçin, C. (2024). Effects of methane emissions on multiple myeloma-related mortality rates: A World Health Organization perspective. Medicine,103(15), Article e37580. 10.1097/MD.0000000000037580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Pinakana, S. D., Mendez, E., Ibrahim, I., Majumder, Md. S., & Raysoni, A. U. (2023). Air pollution in South Texas: A short communication of health risks and implications. Air,1(2), 94–103. 10.3390/air1020008 [Google Scholar]
  35. Rabie, R., Asghari, M., Nosrati, H., Emami Niri, M., & Karimi, S. (2024). Spatially resolved air quality index prediction in megacities with a CNN-Bi-LSTM hybrid framework. Sustainable Cities and Society,109, Article 105537. 10.1016/j.scs.2024.105537 [Google Scholar]
  36. Rammah, A., Whitworth, K. W., & Symanski, E. (2020). Particle air pollution and gestational diabetes mellitus in Houston. Texas. Environmental Research,190, 109988. 10.1016/j.envres.2020.109988 [DOI] [PubMed] [Google Scholar]
  37. Ridlington, E., & Rumpler, J. (2013). Fracking by the numbers: Key impacts of dirty drilling at the state and national level. Environment Colorado Research & Policy Center.
  38. Russo, P. N., & Carpenter, D. O. (2019). Air emissions from natural gas facilities in New York State. International Journal of Environmental Research and Public Health,16(9), Article 1591. 10.3390/ijerph16091591 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Santibáñez-Andrade, M., Chirino, Y. I., González-Ramírez, I., Sánchez-Pérez, Y., & García-Cuellar, C. M. (2019). Deciphering the code between air pollution and disease: The effect of particulate matter on cancer hallmarks. International Journal of Molecular Sciences,21(1), Article 136. 10.3390/ijms21010136 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Schwarze, P. E., Øvrevik, J., Låg, M., Refsnes, M., Nafstad, P., Hetland, R. B., & Dybing, E. (2006). Particulate matter properties and health effects: Consistency of epidemiological and toxicological studies. Human & Experimental Toxicology,25(10), 559–579. 10.1177/096032706072520 [DOI] [PubMed] [Google Scholar]
  41. Sexton, K., Linder, S. H., Marko, D., Bethel, H., & Lupo, P. J. (2007). Comparative assessment of air pollution-related health risks in Houston. Environmental Health Perspectives,115(10), 1388–1393. 10.1289/ehp.10043 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Tao, H., Jawad, A. H., Shather, A. H., Al-Khafaji, Z., Rashid, T. A., Ali, M., et al. (2023). Machine learning algorithms for high-resolution prediction of spatiotemporal distribution of air pollution from meteorological and soil parameters. Environment International,175, Article 107931. 10.1016/j.envint.2023.107931 [DOI] [PubMed] [Google Scholar]
  43. Texas Department of Health. (2003). Texas voluntary indoor air quality guidelines for government buildings. https://gato-docs.its.txst.edu/jcr:afe38366-c8fe-4e39-8e6f-9eab96d729ba/Gov_Bld_Gd.pdf. Accessed 20 Feb 2026.
  44. Tokuda, N., Ishikawa, R., Yoda, Y., Araki, S., Shimadera, H., & Shima, M. (2023). Association of air pollution exposure during pregnancy and early childhood with children’s cognitive performance and behavior at age six. Environmental Research,236, Article 116733. 10.1016/j.envres.2023.116733 [DOI] [PubMed] [Google Scholar]
  45. Tseng, C.-H., Tsuang, B.-J., Chiang, C.-J., Ku, K.-C., Tseng, J.-S., Yang, T.-Y., et al. (2019). The relationship between air pollution and lung cancer in nonsmokers in Taiwan. Journal of Thoracic Oncology,14(5), 784–792. 10.1016/j.jtho.2018.12.033 [DOI] [PubMed] [Google Scholar]
  46. U.S. Energy Information Administration. (2025). Texas State Energy Profile. https://www.eia.gov/state/print.php?sid=TX. Accessed 20 Feb 2026.
  47. Vachon, J., Kerckhoffs, J., Buteau, S., & Smargiassi, A. (2024). Do machine learning methods improve prediction of ambient air pollutants with high spatial contrast? A systematic review. Environmental Research,262, 119751. 10.1016/j.envres.2024.119751 [DOI] [PubMed] [Google Scholar]
  48. Van Donkelaar, A., Martin, R. V., Brauer, M., Kahn, R., Levy, R., Verduzco, C., & Villeneuve, P. J. (2010). Global estimates of ambient fine particulate matter concentrations from satellite-based aerosol optical depth: Development and application. Environmental Health Perspectives,118(6), 847–855. 10.1289/ehp.0901623 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Veefkind, J. P., Aben, I., McMullan, K., Förster, H., De Vries, J., Otter, G., et al. (2012). TROPOMI on the ESA Sentinel-5 precursor: A GMES mission for global observations of the atmospheric composition for climate, air quality and ozone layer applications. Remote Sensing of Environment,120, 70–83. 10.1016/j.rse.2011.09.027 [Google Scholar]
  50. Wang, A., Xu, J., Tu, R., Saleh, M., & Hatzopoulou, M. (2020). Potential of machine learning for prediction of traffic related air pollution. Transportation Research Part D: Transport and Environment,88, Article 102599. 10.1016/j.trd.2020.102599 [Google Scholar]
  51. Weinhold, B. (2012). The future of fracking: New rules target air emissions for cleaner natural gas production. Environmental Health Perspectives, 120(7). 10.1289/ehp.120-a272
  52. Wood, D. A. (2022). Local integrated air quality predictions from meteorology (2015 to 2020) with machine and deep learning assisted by data mining. Sustainability Analytics and Modeling,2, Article 100002. 10.1016/j.samod.2021.100002 [Google Scholar]
  53. Yarragunta, S., Nabi, M. A., Jeyanthi, P., & Revathy, S. (2021). Prediction of air pollutants using supervised machine learning. In 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS) (pp. 1633–1640). Presented at the 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS), Madurai, India: IEEE. 10.1109/ICICCS51141.2021.9432078
  54. Zhuang, J., Hu, J., Bei, F., Huang, J., Wang, L., Zhao, J., et al. (2022). Exposure to air pollutants during pregnancy and after birth increases the risk of neonatal hyperbilirubinemia. Environmental Research,206, Article 112523. 10.1016/j.envres.2021.112523 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All data generated or analyzed during this study will be available upon request.


Articles from Environmental Monitoring and Assessment are provided here courtesy of Springer

RESOURCES