Skip to main content
PLOS One logoLink to PLOS One
. 2026 Aug 20;21(8):e0343606. doi: 10.1371/journal.pone.0343606

Use of machine learning to detect Escherichia coli in drinking water in Bangladesh

Iqramul Haq 1,2, Md Yusuf Hossain Ador 3, Diego Nobrega 1,*
Editor: Sohail Saif4
PMCID: PMC13492749  PMID: 42623399

Abstract

Escherichia coli (E. coli) is a key indicator of fecal contamination in freshwater and can signal the presence of other harmful bacteria and viruses. The aim of the study is to evaluate the performance of machine learning (ML) tools to detect E. coli in drinking water in Bangladesh using surveillance data under two scenarios: an imbalanced dataset and a balanced dataset. We utilized data from the 2019 Bangladesh Multiple Indicator Cluster Survey, which included a total of 6,069 household drinking water samples. We used agglomerative hierarchical clustering with Ward’s linkage to identify district-level hotspots. Extreme Gradient Boosting with SHapley Additive exPlanations values were used for feature selection, and the Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance in the classification task. We applied nine classical ML models in this study: Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), k-Nearest Neighbors (KNN), Light Gradient-Boosting Machine (LightGBM), Logistic Regression (LR), Naïve Bayes (NB), Random Forest (RF), and Support Vector Machine (SVM), along with a Deep Learning Multi-Layer Perceptron (DL-MLP) model to predict the risk of E. coli contamination (REcC) in water. Model performance was evaluated using accuracy, precision, recall, F1 score, Cohen Kappa, area under the curve (AUC), and a violin plot. E. coli contamination in drinking water was detected in 39.2% (95% CI: 37.4–41.2) of households. Bandarban district had the highest REcC. After applying SMOTE and 10-fold cross-validation with hyperparameter tuning, model performance was more consistent across algorithms. In terms of model evaluation, AdaBoost slightly outperformed the others with an accuracy of 81.6%, Cohen kappa statistic of 19.4%, precision of 82.2%, recall of 99%, F1-score of 89.8%, and an AUC of 68.6%. Ensemble model for example AdaBoost and GBA models had ability to accurately classify drinking water samples with respect to the presence of E. coli using surveillance data than others selected model in this study.

Introduction

Enterobacteriaceae, particularly Escherichia coli (E. coli), are the dominant commensal bacteria found in the gastrointestinal tracts of humans and other warm-blooded animals, and are also considered important pathogens [1]. Contamination of household drinking water with E. coli is a significant public health concern, particularly amongst children [2]. Neonatal meningitis, urinary tract infections, enteritis, septicaemia, and other clinical illnesses are commonly caused by E. coli [1].

Over half of the global population uses groundwater as their primary drinking water source [3]. Conflicting demands have led to over-exploitation of 20% of aquifers, with contamination from chemicals, radionuclides, and microorganisms, as well as improper waste and wastewater management, posing significant risks to groundwater supplies [3]. Acute gastrointestinal illnesses (AGIs) brought on by drinking contaminated water spread irregularly throughout location and time as a result of microbiological groundwater contamination events. Contaminated groundwater is the leading cause of waterborne illness in the US, accounting for 64% of drinking water outbreaks between 1989 and 2002, with up to 150 million North American residents relying on groundwater systems [4].

The burden of waterborne illness is increased by the possibility of exposure to enteric pathogens [5]. Enteric infections are more common among sensitive subpopulations, such as the young, old, pregnant, and malnourished. Both low-and-middle income and high-income countries have high rates of infections caused by enteric pathogens. Groundwater contamination can occur in different ways, such as exposure to septic tanks, sewage, municipal wastewater treatment facilities, wildlife, and agricultural residues [5]. Nearly 60% of private well owners in Ireland do not use an appropriate domestic water treatment system [6].

In many countries, storing drinking water is a common practice due to intermittent or lack of direct access to drinking water within homes. Water storage has been identified as an important risk factor for diarrheal diseases [7]. Unhygienic handling during storage can lead to microbial contamination of drinking water [8]. The type of water source also affects contamination levels [9]. Additionally, water supplies may become contaminated in case of open defecation or subpar toilets are used. Social characteristics such as income will also be associated with contamination levels. Households categorised as middle-class or lower-class in Bangladesh were shown to be at a higher risk of drinking water contamination when they obtained their water from sources with elevated levels of E. coli [2]. The possibility of increased contamination levels was greatly decreased by treating drinking water, although pet ownership was strongly linked to recontamination.

From a statistical perspective, different approaches were used to identify the presence of E. coli contamination among household drinking water. A study in North America used pooled analysis to investigate groundwater contamination in the US and Canada [5]. Another researcher used association rule learning, regression analysis, and variable discretization techniques to identify the relationship between contamination and other predictors [3].

Machine learning (ML) has emerged as a significant solution for various applications such as bioinformatics [10,11], and activity recognition in recent decades. Deep learning (DL) models are powerful tools for computation on large datasets [12]. Since the early 2000s, DL has greatly improved predictive models by automatically extracting and analyzing valuable insights from raw data [13]. A hybrid model is created for general classification tasks, showing better performance than individual ML or DL models [14].

ML techniques were used to investigate the growth of E. coli O157:H7 in spiked raw ground beef and evaluated prediction performance using coefficient of determination (R²) and mean squared error [15]. A second study employed an impedimetric electrochemical aptasensor with gold interdigitated electrodes to measure E. coli O157:H7 in surface water for hydroponic lettuce irrigation. A ML framework was used to evaluate the effectiveness of existing methods for contamination detection and model performance was assessed using the root mean squared error [16]. Many ML algorithms were used to predict E. coli levels such as stochastic gradient boosting, random forest (RF), support vector machines (SVM), and k-nearest neighbor (KNN) [17]. Predictive models for E. coli were developed using regression techniques, artificial neural networks (ANN), and two adaptive neuro-fuzzy inference system (ANFIS) structures and their performance was compared using R² and the root mean squared log error [18]. RF, Categorical Boosting, and Naïve Bayes (NB) were used to classify surgery-ward E. coli isolates from susceptibility profiles, and performance was evaluated using accuracy, precision, recall, F1-score, and area under the curve (AUC) [19]. RF, logistic regression (LR), and SVM analyzed flow cytometry 2D histograms to detect E. coli associated microbial patterns, evaluated by accuracy, sensitivity, precision, and F1-score [20].

Although previous studies identified risk factors for E. coli contamination using traditional statistical methods [2,21], this study advances beyond risk factor identification by developing predictive models that classify contamination status in settings where routine laboratory surveillance is absent or infeasible. ML algorithms offer two critical advantages over conventional statistical approaches: (1) higher predictive accuracy through detection of complex non-linear relationships and high-order interactions among environmental, socioeconomic, and household-level predictors, and (2) the capacity to prioritize high-risk locations for targeted sampling when universal testing is not financially or logistically viable. To achieve this, we test the ability of classical ML [Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), KNN, Light Gradient-Boosting Machine (LightGBM), LR, NB, RF, and SVM] and the DL Multi-layer Perceptron (DL-MLP) algorithms using data from the 2019 Multiple Indicator Cluster Survey (MICS) to estimate the risk of E. coli contamination (REcC) in the drinking water under two scenarios: (1) an imbalanced dataset assessed using 10-fold stratified cross-validation (CV) without hyperparameter tuning, and (2) Synthetic Minority Over-sampling Technique (SMOTE)-balanced training data assessed using 10-fold stratified CV with hyperparameter tuning. We also identified the relevant predictors. The novelty of this work lies in transitioning from explanatory risk factor analysis to operational prediction tools that enable resource-constrained health systems to allocate limited diagnostic capacity based on model-predicted contamination probability rather than reactive testing. The Bangladeshi government is implementing advanced analytical approaches to predict E. coli contamination risk, aiming to improve water quality and protect public health. The 2019 MICS data can help identify hotspots and vulnerable populations, enhancing policymakers’ understanding of socio-demographic factors influencing contamination risk. This evidence-based approach can inform targeted interventions and resource allocation.

Methods

Data source and study design

The study analyzed secondary dataset from the 2019 MICS in Bangladesh to assess E. coli concentrations in drinking water. A two-stage, stratified cluster sampling method was implemented. Primary sampling units were selected from enumeration areas (EAs) defined by the 2011 Bangladesh Population and Housing Census in the first stage. Subsequently, 20 households (HHs) were selected from each EA via systematic random sample, with four households chosen for arsenic level assessment in their drinking water. Additionally, two HHs were randomly selected to measure E. coli concentrations in both household drinking water and its source. The study targeted a sample size of 6,440 households, with 6,069 (98.7%) successfully providing data for E. coli testing [22].

Outcome variable

In this paper, REcC in the HH drinking water is the response variable. The most recommended indicator for fecal contamination is the number of E. coli bacteria counts in a 100 mL sample. The outcome variable is E. coli contamination in household drinking water, which is categorized using WHO criteria into low (<1 colony forming units [CFU]/100mL), medium (1–10 CFU/100mL), and high risk (11–100 or more CFU/100 mL) categories [23,24]. To this study, medium and high-risk categories were grouped as increased risk category. Specifically, households with 1 or more CFU per 100 mL of drinking water were categorized as ‘Yes’ for the REcC, while those with less than 1 CFU per 100 mL were categorized as ‘No’ for the REcC. Mathematically, the outcome variable can be defined as:

REcC= {Yes,          If  CFU≥1 per 100 mL of drinking waterNo,           If CFU<1 per 100 mL of drinking water

Independent variables

In addition to the outcome variable, division, place of residence, age of household head, household head education, household size, wealth index, livestock ownership, types of toilet facility, sources of drinking water, location of the water source, treatment of drinking water, place of handwashing, mass media exposure, and women functional difficulties were selected as potential factors influencing REcC among drinking water in households (Table 1).

Table 1. Summary of the selected covariates.

Variables Categorization
Division Barishal, Chattogram, Dhaka, Khulna, Mymensingh, Rangpur, Rajshahi, Sylhet
Place of residence Urban or Rural
Age of household head 15-64, 65+
Household head education No education, Primary, Secondary+
Household size <4, 4-5, 6+
Wealth index Poor, Middle, Rich
Livestock ownership Yes or No
Types of toilet facility Flush, Pit latrine, Others
Sources of drinking water Piped water, Tube wall, Others
Location of the water source Own dwelling, Own yard/plot, Elsewhere
Treatment of drinking water Yes or No
Place of handwashing Dwelling, Yard/plot, Mobile object, No place
Mass media exposure Yes or No
Women functional difficulty Yes or No

Statistical analysis

We first used a traditional approach based on Chi-squares tests to explore associations between REcC and explanatory variables. Agglomerative hierarchical clustering using Ward’s linkage was used to identify district-level hotspots for the REcC in drinking water. This method automatically groups data points into clusters based on their similarities, starting with each data point in its own cluster and merging the closest two until a single cluster remains.

XGBoost SHapley additive explanations (SHAP)

Prior to the ML analysis, a predictive model for imbalanced classification was developed using eXtreme Gradient Boosting (XGBoost), with the goal of reducing overfitting and improving accuracy through high speed [25,26]. We then used SHapley Additive exPlanations (SHAP) to interpret XGBoost model predictions, providing a comprehensive understanding of each feature’s contribution to the model’s predictions [27]. The SHAP summary plot was used to rank predictors by their mean absolute SHAP values (average absolute impact on the XGBoost output). We then dropped predictors with near-zero SHAP contribution and refit the model using the remaining variables to reduce noise and overfitting. Next, we reviewed SHAP dependence plots to identify clear non-linear patterns and strong interactions among key predictors. Based on these patterns, we updated feature coding where needed (for example, collapsing sparse categories and using more appropriate transformations) and refit the final model. Model performance was re-assessed using the same validation procedure to confirm that these SHAP-guided changes improved generalization and produced more stable predictions.

Additionally, Cramer’s V correlation was used to identify potential collinearity among features, ranging from 0 (no association) to 1 (perfect association) [28]. Values between 0.36 and 0.49 indicated a substantial correlation, while values of 0.50 or higher were considered strong correlation [29].

ML classifiers

This study employed a comprehensive approach, integrating nine (09) classical ML models and one (01) DL techniques, to predict the REcC in household drinking water in Bangladesh. In the training dataset, we applied nine classical ML models utilized in this study include AdaBoost, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM. Additionally, we incorporated DL-MLP model into our analysis. ML models and evaluation metrics are extensively utilized in the literature for classification tasks [25,30–37].

Data preparation and model evaluation

For the ML approach, the MICS dataset was split into training (70%) and test (30%) sets using stratification to preserve the REcC class distribution. The dataset was checked for missing values, which were handled using KNN imputation [38] before training the selected ML and DL algorithms was implemented within a scikit-learn Pipeline so that the same steps were applied consistently during training and validation. Data preprocessing was performed using Python [35], Scikit-learn [39], pandas [40] and Keras [41]. Imbalanced data refers to unequal representation of classes in classification problems, including binary and multi-class tasks [42]. The training set was subjected to 10-fold CV to identify optimal hyperparameters. We used the test dataset to evaluate the performance of models measured by various evaluation parameters, including accuracy, precision, recall, F1 score, and AUC [33,34,43–46]. Results were presented in both tabulated and graphical formats, with violin plots [33] being used to illustrate the distribution of model performances. The overall procedure of the ML approach is displayed in Fig 1.

Fig 1. Flow chart of the procedure of the nine ML classifiers and one DL classifier.

Fig 1

Classification outputs

After training, each household sample was classified using two model outputs. Predicted class labels were obtained using predict(X) and used to generate the confusion matrix and compute accuracy, precision, recall, F1-score, and Cohen’s kappa. When available, predicted probabilities were obtained using predict_proba(X)[:,1]; for SVM, probability estimation was enabled (probability = True) to obtain probability outputs. Receiver operating characteristic (ROC)-AUC was computed using predicted probabilities, whereas the other metrics were computed using predicted class labels.

Hyperparameter

Hyperparameters control the learning process and help determine the best model parameters [47]. A ML model’s performance depends on optimal hyperparameters, which can be tuned using methods like grid search, random search, and Bayesian optimization [48]. Grid search divides the hyperparameter space into a grid and evaluates all combinations using CV to find the optimal values [49]. We used grid search with CV to optimize each model’s hyperparameters. We optimized model hyperparameters using GridSearchCV applied to the training set only. For each algorithm, we evaluated the parameter grids listed in S1 Table in S1 File. Model selection was performed using stratified 10-fold CV (StratifiedKFold, n = 10, shuffle = True, random_state = 42) and the mean F1-score was used as the tuning criterion. The best model (best_estimator_) was refitted on the full training set and evaluated once on the held-out test set.

SMOTE

Class imbalance arises when one class has far more samples than another, causing models to overfit the majority class and perform poorly on the minority class [50–52]. SMOTE generates synthetic minority-class samples to balance the dataset, reducing majority-class bias and helping the classifier learn the minority class more effectively [53–55]. Since the outcome variable (REcC in drinking water) was imbalanced (39.2% “Yes” and 60.8% “No”), we used the SMOTE to address class imbalance in this binary outcome [56]. After SMOTE, synthetic samples are generated for the minority class so that the class distribution becomes more balanced, with roughly equal numbers of instances in each class [57]. Several studies have shown that SMOTE usually boosts minority-class precision, recall, and F1-score, giving a more balanced and fair assessment of model performance [52,57]. For our analysis, we evaluated two training scenarios for each model: (i) the original imbalanced training data, assessed using 10-fold stratified CV without hyperparameter tuning, and (ii) SMOTE-balanced training data, assessed using 10-fold stratified CV with hyperparameter tuning. When SMOTE was used, it was applied only to the training split within each CV fold; the validation folds and the final held-out test set remained non-synthetic.

Results

Prevalence of E. coli in household drinking water

The overall prevalence of E. coli contamination in drinking water was 39.2% (95% CI: 37.4 to 41.2).

Association between sociodemographic factors and E. coli contamination

In univariate analysis, the REcC in drinking water in Bangladesh was significantly associated with several factors, including division, residence, household head education, household size, wealth index, livestock ownership, types of toilet facilities, sources of drinking water, treatment methods for drinking water, handwashing locations, mass media exposure, and women functional difficulties (p < 0.05) (Table 2).

Table 2. Characteristics of enrolled households according to REcC in drinking water.

Variables REcC χ2 (p-value)
Yes (%) No (%)
Division
Barishal 15.8 84.2 353.92 (<0.001)
Chattogram 49.7 50.3
Dhaka 51.4 48.6
Khulna 37.1 62.9
Mymensingh 43.4 56.6
Rangpur 23.4 76.6
Rajshahi 27.5 72.5
Sylhet 37.1 62.9
Place of residence
Urban 46.9 53.1 43.10 (<0.001)
Rural 37.0 63.0
Age of household head (in years)
15-64 39.0 61.0 0.577 (0.447)
65+ 40.4 59.6
Household head education
No education 42.3 57.7 13.89 (0.001)
Primary 37.9 62.1
Secondary+ 37.0 63.0
Household size
<4 36.2 63.8 13.78 (0.001)
4-5 40.2 59.8
6+ 42.2 57.8
Wealth index
Poor 37.6 62.4 19.69 (<0.001)
Middle 35.9 64.1
Rich 42.6 57.4
Livestock ownership
Yes 39.5 60.5 24.81 (<0.001)
No 42.8 57.2
Types of toilet facility
Flush 40.0 60.0 3.554 (0.045)
Pit latrine 38.6 61.4
Others 36.1 63.9
Sources of drinking water
Piped water 54.4 45.6 186.70 (<0.001)
Tube wall 36.0 64.0
Others 77.6 22.4
Location of the water source
Own dwelling 37.9 62.1 0.242 (0.886)
Own yard/plot 37.3 62.7
Elsewhere 38.0 62.0
Treatment of drinking water
Yes 58.0 42.0 102.58 (<0.001)
No 37.1 62.9
Place of handwashing
Dwelling 45.1 54.9 19.56 (<0.001)
Yard/plot 38.0 62.0
Mobile object 39.7 60.3
No place 36.6 63.4
Mass media exposure
Yes 40.7 59.3 7.41 (0.007)
No 37.3 62.7
Women functional difficulty
Yes 31.5 68.5 4.889 (0.027)
No 40.1 59.9

The highest percentage of positive households was observed in Dhaka division (51.4%), followed by Chattogram division (49.7%), Mymensingh division (43.4%), Khulna and Sylhet divisions (both at 37.1%), Rajshahi division (27.5%), Rangpur division (23.4%), and Barishal division, which had the lowest risk (15.8%). The proportion of positive households was higher in urban areas (46.9%) than rural areas (37%). There was a significant association between education level of household head and REcC, with 42.3% of households in Bangladesh having heads with no formal education. Approximately 42.2% of households with more than six members were positive for E. coli contamination in their drinking water. With respect to social-economic characteristics, most houses positive for E. coli contamination (42.6%) were of high-income families. Livestock ownership and toilet facilities were significant factors for E. coli contamination risk among households in Bangladesh. Households with flush toilet facilities had increased prevalence of E. coli contamination compared to those with other facilities. A larger proportion (58%) of households that treated their drinking water were at increased REcC. Additionally, a higher percentage of households where hands were washed inside the dwelling (45.1%) were at REcC, followed by those washing hands in the yard/plot (38%), on a mobile object (39.7%), and those with no designated place for hand washing (36.6%). Households with exposure to mass media had a higher REcC (40.7%) compared to those without exposure to mass media (37.3%) (Table 2).

District-wise REcC

Certain districts in Bangladesh had higher REcC compared to others. We detected geographical clustering in contamination results using agglomerative hierarchical clustering, as shown in the dendrogram in S1 Fig in S1 File. The clustering was based on data from all sixty-four districts affected by E. coli in drinking water. The analysis resulted in four distinct clusters, with sizes of 10, 15, 12, and 27 districts, respectively, which were sorted according to the level of contamination of drinking water. Cluster C1 contained districts with relatively high REcC (up to 60%), whereas districts in cluster C4 had low REcC (around 20%). Cluster 1 (highest prevalence of contamination) comprises Bandarban, Dhaka, Habigonj, Narsingdi, Feni, Tangail, Rangamati, Rajbari, Lakshmipur, and Netrokona. Cluster 2 includes Khagrachari, Kushtia, Khulna, Madaripur, Pirojpur, Noakhali, Chattogram, Satkhira, Lalmonirhat, Meherpur, Cumilla, Kurigram, Cox’s Bazar, Bagerhat, and Sirajganj. Cluster 3 covers Jessore, Joypurhat, Rangpur, Thakurgaon, Naogaon, Panchagarh, Bhola, Patuakhali, Barisal, Barguna, Bogura, and Gaibandha. Cluster 4 consists of all remaining districts not categorized in the first three clusters (Fig 2).

Fig 2. District-wise map of the proportion of E. coli contamination in household drinking water across Bangladesh, generated using agglomerative hierarchical clustering with Ward linkage.

Fig 2

Administrative district boundaries were obtained from the GADM database (Bangladesh ADM2) and the map was generated by the authors in R using sf and ggplot2 without any proprietary basemap tiles. Data: MICS 2019, Bangladesh. Map: GADM (ADM2).

Fig 2 shows the district level variation of REcC in the drinking water using agglomerative hierarchical cluster approach. High REcC was observed in Bandarban, Dhaka, Habigonj, Narshingdi, Feni, Rangamati, Tangail, Rajbari, Lakshmipur, and Netrokona districts, which incorporated the first cluster with > 58% of households being positive for E. coli.

Feature selection

Fig 3 displays SHAP values and top fourteen features with highest impact, with yellow indicating high and dark blue indicating low, with horizontal axis representing each feature’s contribution to prediction. Positive SHAP values indicate higher contamination risk, while negative values indicate lower risk; for example, higher division differences were associated with increased contamination. Regarding SHAP, geographical division was found to be the most significant predictor, followed by the household size, wealth index, household head’s education level, place of handwashing, and livestock ownership.

Fig 3. Feature importance based on SHAP values.

Fig 3

The beeswarm summary plot displays the distribution of SHAP values for the top predictors of REcC in household drinking water, with features ordered by their overall contribution to the model.

Cramér’s V correlation

Cramér’s V correlation analysis revealed strong associations between various factors, particularly the place of handwashing and the location of the water source. Consequently, we excluded the place of handwashing from further ML modeling to predict E. coli levels in water (S2 Fig in S1 File).

The distribution of drinking-water sources across divisions is presented in S3 Fig in S1 File. There was a significant association between division and drinking-water source (χ² = 497.52, p < 0.001). Tubewell water was the predominant source in every division, whereas the use of piped water varied by region. Piped water use was highest in Dhaka (20.5%) and substantially lower in the other divisions (2.5%−9.1%). This geographic variation in water-source profiles is consistent with the SHAP results, which identified both division and drinking-water source as important predictors of E. coli contamination risk.

REcC in the drinking water using ML models

For the ML models, we tested nine algorithms RF, LR, DT, GBA, AdaBoost, LightGBM, KNN, NB, and SVM and as well as a DL approach based on an DL-MLP. Model performance after applying SMOTE and CV with hyperparameter tuning is presented in Fig 4. The corresponding results for imbalanced dataset including 10-fold CV (before SMOTE and before hyperparameter tuning) are provided in the Supplementary Materials (S4 and S5 Figs in S1 File). The predictive performance of each algorithm was compared based on accuracy, Cohen’s Kappa (κ), precision, recall, and F1-score. In both scenarios, scenario 2 (SMOTE-balanced training with CV and hyperparameter tuning) performed better than scenario 1 (original imbalanced training data).

Fig 4. Classification performance measure of the algorithms and comparison.

Fig 4

Performance indicators (accuracy, Cohen’s kappa coefficient, precision, recall, and F1 score) for ML algorithms (SVM, RF, NB, LR, LightGBM, KNN, GBA, DT, AdaBoost and DL-MLP).

For scenario 2 (after SMOTE), AdaBoost demonstrated the highest accuracy at 81.6%, correctly predicting the REcC in household drinking water (Fig 4). It also achieved the highest Cohen’s Kappa (κ = 0.194), recall (0.990), and F1-score (0.898). GBA and LightGBM performed similarly well, with accuracies of 0.799 and 0.791, respectively, and strong balance between sensitivity and precision (GBA: recall 0.965, F1 0.887; LightGBM: recall 0.943, F1 0.881). RF and DT showed moderate-to-strong performance (RF: accuracy 0.759, F1 0.858; DT: accuracy 0.739, F1 0.842), while KNN achieved reasonable accuracy (0.717) and F1 (0.830) but had the lowest κ (0.077), indicating weaker improvement over chance agreement.

In contrast, LR and NB had lower accuracies (0.556 and 0.578) and lower F1-scores (0.676 and 0.696), despite relatively high precision (both ≥ 0.841). SVM also showed low accuracy (0.587) and a comparatively low F1-score (0.707). The DL-MLP model showed intermediate performance (accuracy 0.625, recall 0.682, F1 0.749). Among the ten classifiers, the AdaBoost ML classifier slightly outperformed the others (Fig 4 and S6 Fig in S1 File).

To predict the REcC among households in Bangladesh, the AUC measure was estimated for the LR, RF, GBA, DT, DL-MLP, KNN, AdaBoost, LightGBM, SVM and NB model models (Fig 5). The estimated AUC values were 0.591 (LR), 0.634 (NB), 0.670 (DT), 0.678 (RF), 0.651 (SVM), 0.619 (KNN), 0.686 (AdaBoost), 0.685 (GBA), 0.676 (LightGBM), and 0.626 (DL-MLP). Overall, the boosting and tree-based models showed the best discrimination for predicting REcC in household drinking water. AdaBoost achieved the highest AUC (0.686), followed closely by GBA (0.685), RF (0.678), and LightGBM (0.676). In contrast, LR had the lowest AUC (0.591), indicating the weakest discriminative performance among the evaluated models.

Fig 5. ROC curves and AUC for ten classifiers (AdaBoost, DL-MLP, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM) using 10-fold CV to predict the REcC in household drinking water.

Fig 5

AdaBoost model achieved the highest and most consistent median accuracy across folds, followed by GBA, LightGBM, and RF. Unlike a box plot, the violin plot in Fig 6 displays the full distribution of the 10-fold CV accuracy results.

Fig 6. Violin plots showing the distribution of accuracy across 10-fold CV after SMOTE and hyperparameter tuning for ten classifiers (AdaBoost, DL-MLP, DT, GBA, KNN, LightGBM, LR, NB, RF, and SVM) used to predict the REcC in household drinking water.

Fig 6

The central point indicates the median and the inner band shows the interquartile range.

Feature importance

Fig 7 illustrates the feature importance based on the AdaBoost model, providing valuable insights into the factors influencing E. coli levels in drinking water in Bangladesh. The analysis identifies five key variables division, wealth index, types of toilet facility, household size, and household head education as significant predictors of E. coli contamination in household water.

Fig 7. Feature importance based on AdaBoost.

Fig 7

Discussion

This cross-sectional study used 2019 MICS dataset and main purpose of this study is the assess the performance of ML algorithms to detect REcC in drinking water and identified the hotspot of highly contaminated district of risk of E. coli using the cluster dendrogram of the hierarchical clustering based on Wald’s criterion. To the best of our knowledge, there was no available study which have conducted in Bangladesh to focused on REcC based on the different ML classifiers and clustering approaches we have used in this study. A comparison table between our results and previous research is provided in S2 Table in S1 File.

This study also evaluated that prevalence of REcC in drinking water in Bangladesh was 39.2%. About 39% and 65% of pathogens were detected in point-of-drinking water samples and public domain water source samples, respectively, and classified as low-risk using PCR testing [58]. In a similar study researcher was done using API and conventional methods and concluded that the prevalence of E. coli from different water sources was 49.48% [59]. In our study REcC in drinking water in Bangladesh was significantly associated with several factors, including division, residence, household head education, household size, wealth index, livestock ownership, types of toilet facilities, sources of drinking water, treatment methods for drinking water [8], handwashing location, mass media exposure and maternal functional difficulties. Moreover, a univariate and multivariate analysis conducted in a study [4] which was found that demographic factors such as household heads, room occupancy rate and wealth status was significantly correlated with the occurrence of E. coli in stored drinking water [60].

We evaluated the performance of nine ML algorithms and one DL algorithm to predict REcC in drinking water across two scenarios: original (imbalanced) data and balanced data (SMOTE, with CV and hyperparameter tuning). The ML models included RF, LR, DT, GBA, AdaBoost, LightGBM, KNN, NB, and SVM; DL-MLP was used for DL. Across all ML and DL models, Scenario 2 (balanced data) outperformed Scenario 1 (imbalanced data). For Scenario 2, AdaBoost demonstrated the highest accuracy (81.6%) for correctly predicting REcC in drinking water, along with the highest recall (0.990) and F1-score (0.898). AdaBoost is often highly effective for classification because it combines weak learners to refine decision boundaries, capture feature interactions, and preserve good generalization, which frequently leads to superior accuracy and F1 scores compared with many other models [61,62]. From a practical perspective, our results suggest that boosting and other tree-based ensemble models (AdaBoost, GBA, and LightGBM) are strong first-choice methods for similar survey-based contamination studies, where predictors are largely categorical and relationships may be non-linear. In this dataset, LR showed weaker discrimination (lower AUC) than the ensemble approaches, which may reflect the importance of non-linear effects and interactions across household, environmental, and geographic factors.

Additionally, in terms of estimated AUC values, both the AdaBoost and GBA models outperformed LR, RF, GBA, DT, KNN, LightGBM, SVM, NB and DL-MLP models in predicting the REcC among households in Bangladesh. AdaBoost is versatile because it can use different types of weak learners (for example, decision stumps, SVMs, or neural networks) and be tuned to specific problems, which often improves its performance [63,64]. In the Milwaukee River system, physicochemical water-quality variables explained only a limited proportion of the variation in E. coli (R² = 0.29–0.42). The hybrid ANN–GBM model performed best, but prediction accuracy remained low (about 42%) [65]. RF achieved the highest R² value (0.933) when combining water quality and RGB data in the USA [66], as well as in predicting E. coli concentrations in the Marne River (Paris Area, France) [67]. KNN was the most accurate model for predicting E. coli concentrations in Midmar Dam, with XGBoost and SVM performing next best, while RF and ANN had the largest errors [68].

Moreover, RF has been more effective in predicting E. coli concentrations in agricultural pond waters, as it provided the smallest RMSE in almost all cases [17]. Furthermore, both RF and TPOT exhibited the best performance in predicting E. coli concentrations in the Göta älv river in Gothenburg, Sweden [69], and in Ethiopia [70]. On the other hand, SVM outperforms LR in 26 cities with an accuracy of 70%, except in Kathmandu, where it achieves 79%. In contrast, LR demonstrates a significantly lower accuracy of 61% across all cities [71].

For feature importance in the AdaBoost model, division was the strongest predictor of REcC in drinking water, followed by wealth index, types of toilet facility, household size, household head education, and source of drinking water. Recent study showed that climate, land use, and population density were key predictors of E. coli contamination in shallow tubewell water in Bangladesh [72]. A recent study using a Bayesian censored model on a similar dataset found that household E. coli levels were significantly associated with division, water source/location, treatment, handwashing, toilet type, education and wealth, and livestock ownership [73].

Despite all efforts to optimize ML algorithms even after class balancing for predicting REcC using the 2019 MICS dataset, models achieved a maximum accuracy of 81.6% but only a modest AUC of 0.686. This discrepancy reveals ML’s limited discriminatory ability to reliably predict contamination in household drinking water, even at moderate prevalence levels.

Furthermore, a study conducted in Ireland explored spatial mapping of E. coli occurrence in the community [74] to examine local geographic patterns of E. coli contamination. This study identified hotspot areas with high levels of contamination by using a cluster dendrogram from hierarchical clustering based on Wald’s criterion. In our study, we concluded that the highest REcC in household drinking water was observed in districts such as Bandarban, Dhaka, Habiganj, Narsingdi, Feni, Rangamati, Tangail, Rajbari, Lakshmipur, and Netrokona. These districts were marked as the first cluster, where more than 58% of cases were identified, with red spots indicating the highest levels of contamination. In contrast, districts represented in sky blue were grouped into the fourth cluster, which was characterized as low risk for E. coli contamination in household drinking water across Bangladesh.

This study has limitations. In terms of limitations, it is a cross-sectional study, so no cause-and-effect relationships can be established. Additionally, the analysis relied on socio-demographic variables available in the MICS dataset, excluding environmental factors such as distance from sewage treatment plants and population density because these variables were absent from the dataset and could impact contamination risks.

Conclusion

In this study, we applied nine ML classifiers and one DL-MLP classifier to predict the REcC in household drinking water in Bangladesh using the 2019 MICS dataset. We also used agglomerative hierarchical clustering with Ward’s linkage to identify district-level hotspots. More than one-third of households were classified as having increased REcC in their drinking water. We compared model performance using accuracy, precision, recall, F1-score, Cohen’s kappa, ROC-AUC, and violin plots to summarize the distribution of performance across models. Across all models, SMOTE combined with CV and hyperparameter tuning substantially improved performance compared to the original imbalanced data. Based on ROC-AUC, the models showed limited discrimination between contaminated and non-contaminated household drinking-water samples, even though the best-performing model achieved 81.6% accuracy. Feature-importance results indicated that division, wealth index, type of toilet facility, household size, education, and source of drinking water were the most influential predictors of REcC. Future work should focus on improving predictive performance by incorporating additional environmental and spatial exposure variables where possible and by validating models on independent data.

Supporting information

S1 File. Supporting information figures and tables (S1-S6 Figs; S1-S2 Tables).

(DOCX)

Data Availability

The minimal derived datasets required to reproduce the results reported in this manuscript are available on Figshare (https://doi.org/10.6084/m9.figshare.31063324). The underlying Bangladesh MICS 2019 household microdata are third-party controlled-access data provided by UNICEF and cannot be redistributed by the authors. Researchers may request access via https://mics.unicef.org/surveys.

Funding Statement

This study was financially supported by the Canada Research Chairs Program in the form of an award received by DN (2022-00168). No additional external funding was received for this study. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Allocati N, Masulli M, Alexeyev MF, Di Ilio C. Escherichia coli in Europe: An overview. Int J Environ Res Public Health. 2013;10(12):6235–54. doi: 10.3390/ijerph10126235 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Khan JR, Bakar KS. Spatial risk distribution and determinants of E. coli contamination in household drinking water: A case study of Bangladesh. Int J Environ Health Res. 2020;30(3):268–83. doi: 10.1080/09603123.2019.1593328 [DOI] [PubMed] [Google Scholar]
  • 3.White K, Dickson-Anderson S, Majury A, McDermott K, Hynds P, Brown RS, et al. Exploration of E. coli contamination drivers in private drinking water wells: An application of machine learning to a large, multivariable, geo-spatio-temporal dataset. Water Res. 2021;197:117089. doi: 10.1016/j.watres.2021.117089 [DOI] [PubMed] [Google Scholar]
  • 4.Fong T-T, Mansfield LS, Wilson DL, Schwab DJ, Molloy SL, Rose JB. Massive microbiological groundwater contamination associated with a waterborne outbreak in Lake Erie, South Bass Island, Ohio. Environ Health Perspect. 2007;115(6):856–64. doi: 10.1289/ehp.9430 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Hynds PD, Thomas MK, Pintar KDM. Contamination of groundwater systems in the US and Canada by enteric pathogens, 1990–2013: A review and pooled-analysis. PLoS One. 2014;9(5):e93301. doi: 10.1371/journal.pone.0093301 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.O’Dwyer J, Hynds PD, Byrne KA, Ryan MP, Adley CC. Development of a hierarchical model for predicting microbiological contamination of private groundwater supplies in a geologically heterogeneous region. Environ Pollut. 2018;237:329–38. doi: 10.1016/j.envpol.2018.02.052 [DOI] [PubMed] [Google Scholar]
  • 7.Roberts L, Chartier Y, Chartier O, Malenga G, Toole M, Rodka H. Keeping clean water clean in a Malawi refugee camp: A randomized intervention trial. Bull World Health Organ. 2001;79(4):280–7. [PMC free article] [PubMed] [Google Scholar]
  • 8.Dada N, Vannavong N, Seidu R, Lenhart A, Stenström TA, Chareonviriyaphap T, et al. Relationship between Aedes aegypti production and occurrence of Escherichia coli in domestic water storage containers in rural and sub-urban villages in Thailand and Laos. Acta Trop. 2013;126(3):177–85. doi: 10.1016/j.actatropica.2013.02.023 [DOI] [PubMed] [Google Scholar]
  • 9.Bain R, Cronk R, Wright J, Yang H, Slaymaker T, Bartram J. Fecal contamination of drinking-water in low- and middle-income countries: A systematic review and meta-analysis. PLoS Med. 2014;11(5):e1001644. doi: 10.1371/journal.pmed.1001644 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Holder LB, Haque MM, Skinner MK. Machine learning for epigenetics and future medical applications. Epigenetics. 2017;12(7):505–14. doi: 10.1080/15592294.2017.1329068 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Mavaie P, Holder L, Beck D, Skinner MK. Predicting environmentally responsive transgenerational differential DNA methylated regions (epimutations) in the genome using a hybrid deep-machine learning approach. BMC Bioinformatics. 2021;22(1):575. doi: 10.1186/s12859-021-04491-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Deng J, Dong W, Socher R, Li L-J, Kai Li, Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009. 248–55. 10.1109/cvpr.2009.5206848 [DOI]
  • 13.Mamoshina P, Vieira A, Lane E, Zhavoronkov A. Applications of deep learning in biomedicine. Mol Pharm. 2016;13(5):1445–54. doi: 10.1021/acs.molpharmaceut.5b00982 [DOI] [PubMed] [Google Scholar]
  • 14.Mavaie P, Holder L, Skinner MK. Hybrid deep learning approach to improve classification of low-volume high-dimensional data. BMC Bioinformatics. 2023;24(1):419. doi: 10.1186/s12859-023-05557-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Al S, Uysal Ciloglu F, Akcay A, Koluman A. Machine learning models for prediction of Escherichia coli O157:H7 growth in raw ground beef at different storage temperatures. Meat Sci. 2024;210:109421. doi: 10.1016/j.meatsci.2023.109421 [DOI] [PubMed] [Google Scholar]
  • 16.Qian H, McLamore E, Bliznyuk N. Machine learning for improved detection of pathogenic E. coli in hydroponic irrigation water using impedimetric aptasensors: A comparative study. ACS Omega. 2023;8(37):34171–9. doi: 10.1021/acsomega.3c05797 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Stocker MD, Pachepsky YA, Hill RL. Prediction of E. coli concentrations in agricultural pond waters: Application and comparison of machine learning algorithms. Front Artif Intell. 2022;4:768650. doi: 10.3389/frai.2021.768650 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Abimbola OP, Mittelstet AR, Messer TL, Berry ED, Bartelt-Hunt SL, Hansen SP. Predicting Escherichia coli loads in cascading dams with machine learning: An integration of hydrometeorology, animal density and grazing pattern. Sci Total Environ. 2020;722:137894. doi: 10.1016/j.scitotenv.2020.137894 [DOI] [PubMed] [Google Scholar]
  • 19.Tolan HK, Aydın İ, Tanyildizi-Kokkulunk H, Karakuş M, Akkaya Y, Kaya O, et al. Machine learning model for predicting multidrug resistance in clinical Escherichia coli isolates: A retrospective general surgery study. Antibiotics (Basel). 2025;14(10):969. doi: 10.3390/antibiotics14100969 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Erb IK, Gador N, Jinbäck M, Lindberg E, Paul CJ. A data-driven early warning system for Escherichia coli in water based on microbial community analysis using flow cytometry 2D histograms. Water Research X. 2025;29:100404. doi: 10.1016/j.wroa.2025.100404 [DOI] [Google Scholar]
  • 21.Hasan MM, Hoque Z, Kabir E, Hossain S. Differences in levels of E. coli contamination of point of use drinking water in Bangladesh. PLoS One. 2022;17(5):e0267386. doi: 10.1371/journal.pone.0267386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Bangladesh Bureau of Statistics BBS, UNICEF Bangladesh. Progotir Pathey Bangladesh: Multiple Indicator Cluster Survey 2019. Dhaka (Bangladesh): Bangladesh Bureau of Statistics (BBS). 2019. [Google Scholar]
  • 23.World Health Organization. Guidelines for drinking-water quality. Geneva: World Health Organization. 2011. [Google Scholar]
  • 24.Stauber C, Miller C, Cantrell B, Kroell K. Evaluation of the compartment bag test for the detection of Escherichia coli in water. J Microbiol Methods. 2014;99:66–70. doi: 10.1016/j.mimet.2014.02.008 [DOI] [PubMed] [Google Scholar]
  • 25.Chen T, Guestrin C. XGBoost: a scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD’16), 2016. 785–94. 10.1145/2939672.2939785 [DOI]
  • 26.Zhang C, Jia D, Wang L, Wang W, Liu F, Yang A. Comparative research on network intrusion detection methods based on machine learning. Comput Secur. 2022;121:102861. doi: 10.1016/j.cose.2022.102861 [DOI] [Google Scholar]
  • 27.Mangalathu S, Hwang S-H, Jeon J-S. Failure mode and effects analysis of RC members based on machine-learning-based SHapley Additive exPlanations (SHAP) approach. Engineering Structures. 2020;219:110927. doi: 10.1016/j.engstruct.2020.110927 [DOI] [Google Scholar]
  • 28.Zelko E, Švab I, Rotar Pavlič D. Quality of life and patient satisfaction with family practice care in a roma population with chronic conditions in Northeast Slovenia. Zdr Varst. 2014;54(1):18–26. doi: 10.1515/sjph-2015-0003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Fan C, Oh DS, Wessels L, Weigelt B, Nuyten DSA, Nobel AB, et al. Concordance among gene-expression-based predictors for breast cancer. N Engl J Med. 2006;355(6):560–9. doi: 10.1056/NEJMoa052933 [DOI] [PubMed] [Google Scholar]
  • 30.Breiman L. Random forests. Machine Learning. 2001;45(1):5–32. doi: 10.1023/a:1010933404324 [DOI] [Google Scholar]
  • 31.Gupta V, Mishra VK, Singhal P, Kumar A. An Overview of Supervised Machine Learning Algorithm. In: 2022 11th International Conference on System Modeling & Advancement in Research Trends (SMART), 2022. 87–92. 10.1109/smart55829.2022.10047618 [DOI]
  • 32.Jiang T, Gradus JL, Rosellini AJ. Supervised machine learning: A brief primer. Behav Ther. 2020;51(5):675–87. doi: 10.1016/j.beth.2020.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Nayan MdIH, Uddin MSG, Hossain MdI, Alam MdM, Zinnia MA, Haq I, et al. Comparison of the Performance of machine learning-based algorithms for predicting depression and anxiety among university students in Bangladesh. Asian Journal of Social Health and Behavior. 2022;5(2):75–84. doi: 10.4103/shb.shb_38_22 [DOI] [Google Scholar]
  • 34.Rahman R, Talukder A, Das S, Saha J, Sarma H. Understanding and predicting pregnancy termination in Bangladesh: A comprehensive analysis using a hybrid machine learning approach. Medicine (Baltimore). 2024;103(26):e38709. doi: 10.1097/MD.0000000000038709 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Raschka S, Patterson J, Nolet C. Machine learning in python: main developments and technology trends in data science, machine learning, and artificial intelligence. Information. 2020;11(4):193. doi: 10.3390/info11040193 [DOI] [Google Scholar]
  • 36.Sarker IH. Machine learning: Algorithms, real-world applications and research directions. SN Comput Sci. 2021;2(3):160. doi: 10.1007/s42979-021-00592-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Wang R, Kim J-H, Li M-H. Predicting stream water quality under different urban development pattern scenarios with an interpretable machine learning approach. Sci Total Environ. 2021;761:144057. doi: 10.1016/j.scitotenv.2020.144057 [DOI] [PubMed] [Google Scholar]
  • 38.Rizvi STH, Latif MY, Amin MS, Telmoudi AJ, Shah NA. Analysis of machine learning based imputation of missing data. Cybern Syst. 2023;54(8):1–15. doi: 10.1080/01969722.2023.2247257 [DOI] [Google Scholar]
  • 39.Abraham A, Pedregosa F, Eickenberg M, Gervais P, Mueller A, Kossaifi J, et al. Machine learning for neuroimaging with scikit-learn. Front Neuroinform. 2014;8:14. doi: 10.3389/fninf.2014.00014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Gupta P, Bagchi A. Introduction to pandas. Essentials of Python for artificial intelligence and machine learning. Cham (Switzerland): Springer Nature Switzerland. 2024. 161–96. [Google Scholar]
  • 41.Schneider P, Xhafa F. Anomaly detection and complex event processing over IoT data streams: with application to eHealth and patient data monitoring. Cambridge (MA): Academic Press. 2022. [Google Scholar]
  • 42.Han Q, Gui C, Xu J, Lacidogna G. A generalized method to predict the compressive strength of high-performance concrete by improved random forest algorithm. Constr Build Mater. 2019;226:734–42. doi: 10.1016/j.conbuildmat.2019.07.315 [DOI] [Google Scholar]
  • 43.Hicks SA, Strümke I, Thambawita V, Hammou M, Riegler MA, Halvorsen P, et al. On evaluation metrics for medical applications of artificial intelligence. Sci Rep. 2022;12(1):5979. doi: 10.1038/s41598-022-09954-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Khan MU, Aziz S, Iqtidar K, Zaher GF, Alghamdi S, Gull M. A two-stage classification model integrating feature fusion for coronary artery disease detection and classification. Multimed Tools Appl. 2021;81(10):13661–90. doi: 10.1007/s11042-021-10805-3 [DOI] [Google Scholar]
  • 45.Roth GA, Mensah GA, Johnson CO, Addolorato G, Ammirati E, Baddour LM, et al. Global burden of cardiovascular diseases and risk factors, 1990–2019. J Am Coll Cardiol. 2020;76(25):2982–3021. doi: 10.1016/j.jacc.2020.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Rainio O, Teuho J, Klén R. Evaluation metrics and statistical tests for machine learning. Sci Rep. 2024;14(1):6086. doi: 10.1038/s41598-024-56706-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Pravin PS, Tan JZM, Yap KS, Wu Z. Hyperparameter optimization strategies for machine learning-based stochastic energy efficient scheduling in cyber-physical production systems. Digit Chem Eng. 2022;4:100047. doi: 10.1016/j.dche.2022.100047 [DOI] [Google Scholar]
  • 48.Yasir M, Karim AM, Malik SK, Bajaffer AA, Azhar EI. Prediction of antimicrobial minimal inhibitory concentrations for Neisseria gonorrhoeae using machine learning models. Saudi J Biol Sci. 2022;29(5):3687–93. doi: 10.1016/j.sjbs.2022.02.047 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Wu Z, Tran A, Rincon D, Christofides PD. Machine learning‐based predictive control of nonlinear processes. Part I: Theory. AIChE Journal. 2019;65(11). doi: 10.1002/aic.16729 [DOI] [Google Scholar]
  • 50.Niaz NU, Shahariar KMN, Patwary MJA. Class Imbalance Problems in Machine Learning: A Review of Methods And Future Challenges. In: Proceedings of the 2nd International Conference on Computing Advancements, 2022. 485–90. 10.1145/3542954.3543024 [DOI]
  • 51.Brishti F, Zhang F, Mohammed S, Bai L, Wu F, Chen B. Imbalanced classification with label noise: A systematic review and comparative analysis. ICT Express. 2025;11(6):1127–45. doi: 10.1016/j.icte.2025.09.011 [DOI] [Google Scholar]
  • 52.Wibowo P, Fatichah C. An in-depth performance analysis of the oversampling techniques for high-class imbalanced dataset. regist j ilm teknol sist inf. 2021;7(1):63. doi: 10.26594/register.v7i1.2206 [DOI] [Google Scholar]
  • 53.Heroza RI, Gan JQ, Raza H. Sia-SMOTE: a SMOTE-based oversampling method with better interpolation on high-dimensional data by using a Siamese network. In: International Work-Conference on Artificial Neural Networks (IWANN). Cham (Switzerland: ): Springer Nature Switzerland; 2023. 448–460. [Google Scholar]
  • 54.Hairani H, Widiyaningtyas T, Dwi Prasetya D. Addressing class imbalance of health data: A systematic literature review on modified synthetic minority oversampling technique (SMOTE) Strategies. JOIV : Int J Inform Visualization. 2024;8(3):1310. doi: 10.62527/joiv.8.3.2283 [DOI] [Google Scholar]
  • 55.Fernández A, Garcia S, Herrera F, Chawla NV. SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary. J Artif Intell Res. 2018;61:863–905. doi: 10.1613/jair.1.11192 [DOI] [Google Scholar]
  • 56.Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-sampling Technique. J Artif Intell Res. 2002;16:321–57. doi: 10.1613/jair.953 [DOI] [Google Scholar]
  • 57.Coşkuner A, Rençber ÖF. The imbalanced data problem: investigating factors affecting financial freedom using data mining techniques with SMOTE method. Machine learning in finance: trends, developments and business practices in the financial sector. Cham (Switzerland): Springer Nature Switzerland. 2025. 87–100. [Google Scholar]
  • 58.Saima S, Ferdous J, Sultana R, Rashid RB, Almeida S, Begum A, et al. Detecting Enteric Pathogens in Low-Risk Drinking Water in Dhaka, Bangladesh: An Assessment of the WHO Water Safety Categories. Trop Med Infect Dis. 2023;8(6):321. doi: 10.3390/tropicalmed8060321 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Odonkor ST, Addo KK. Prevalence of multidrug-resistant escherichia coli isolated from drinking water sources. Int J Microbiol. 2018;2018:7204013. doi: 10.1155/2018/7204013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Vannavong N, Overgaard HJ, Chareonviriyaphap T, Dada N, Rangsin R, Sibounhom A, et al. Assessing factors of E. coli contamination of household drinking water in suburban and rural Laos and Thailand. Water Supply. 2017;18(3):886–900. doi: 10.2166/ws.2017.133 [DOI] [Google Scholar]
  • 61.Martinez W, Gray JB. Noise peeling methods to improve boosting algorithms. Computational Statistics & Data Analysis. 2016;93:483–97. doi: 10.1016/j.csda.2015.06.010 [DOI] [Google Scholar]
  • 62.Maji K, Gupta S, Dutta PK. Enhancing heart disease prediction accuracy: Comprehensive analysis of XGBoost and AdaBoost. IET Conf Proc. 2025;2024(37):130–6. doi: 10.1049/icp.2025.0833 [DOI] [Google Scholar]
  • 63.Nabi RM, Saeed SAB, Haron H. Enhanced AdaBoostM1 with multilayer perceptron for stock price prediction. KJAR. 2023;8(1):73–85. doi: 10.24017/science.2023.1.7 [DOI] [Google Scholar]
  • 64.Gai F, Li Z, Jiang X, Guo H. In: 2016. 27–37.
  • 65.Nafsin N, Li J. Prediction of total organic carbon and E. coli in rivers within the Milwaukee River basin using machine learning methods. Environ Sci: Adv. 2023;2:278–93. doi: 10.1039/D2VA00285J [DOI] [Google Scholar]
  • 66.Hong SM, Morgan BJ, Stocker MD, Smith JE, Kim MS, Cho KH, et al. Using machine learning models to estimate Escherichia coli concentration in an irrigation pond from water quality and drone-based RGB imagery data. Water Res. 2024;260:121861. doi: 10.1016/j.watres.2024.121861 [DOI] [PubMed] [Google Scholar]
  • 67.Naloufi M, Lucas FS, Souihi S, Servais P, Janne A, Wanderley Matos De Abreu T. Evaluating the performance of machine learning approaches to predict the microbial quality of surface waters and to optimize the sampling effort. Water. 2021;13(18):2457. doi: 10.3390/w13182457 [DOI] [Google Scholar]
  • 68.Ibrahim AAMS, Nkonyane M, Ngcobo M, Walingo T, Tapamo J-R. Data-Driven machine learning models for E. coli concentration prediction. Sustainability. 2025;18(1):179. doi: 10.3390/su18010179 [DOI] [Google Scholar]
  • 69.Sokolova E, Ivarsson O, Lillieström A, Speicher NK, Rydberg H, Bondelind M. Data-driven models for predicting microbial water quality in the drinking water source using E. coli monitoring and hydrometeorological data. Sci Total Environ. 2022;802:149798. doi: 10.1016/j.scitotenv.2021.149798 [DOI] [PubMed] [Google Scholar]
  • 70.Ambel AA, Bain R, Degefu TB, Donmez A, Johnston R, Slaymaker T. Addressing gaps in data on drinking water quality through data integration and machine learning: Evidence from Ethiopia. npj Clean Water. 2023;6(1). doi: 10.1038/s41545-023-00272-8 [DOI] [Google Scholar]
  • 71.Kuroki S, Ogata R, Sakamoto M. Predicting the presence of E. coli in tap water using machine learning in Nepal. Water Environ J. 2023;37(3):402–11. doi: 10.1111/wej.12844 [DOI] [Google Scholar]
  • 72.Wu J, Cao Y, Islam MdS, Emch M. Application of machine learning to identify influential factors for fecal contamination of shallow groundwater. Water. 2025;17(2):160. doi: 10.3390/w17020160 [DOI] [Google Scholar]
  • 73.Haq I, Rahman A, Akter MR, Hossain D, Nobrega D. Bayesian modeling of Escherichia coli contamination in household drinking water in Bangladesh: evidence from the Multiple Indicator Cluster Survey 2019. Int Health. 2025;ihaf138. doi: 10.1093/inthealth/ihaf138 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Samarasundera E, Walsh T, Cheng T, Koenig A, Jattansingh K, Dawe A, et al. Methods and tools for geographical mapping and analysis in primary health care. Prim Health Care Res Dev. 2012;13(1):10–21. doi: 10.1017/S1463423611000417 [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Sohail Saif

1 Dec 2025

-->PONE-D-25-52920-->-->Use of machine learning to detect Escherichia coli in drinking water in Bangladesh-->-->PLOS ONE

Dear Dr. Haq,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Jan 15 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Sohail Saif, Ph.D

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS One has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. PLOS requires an ORCID iD for the corresponding author in Editorial Manager on papers submitted after December 6th, 2016. Please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field. This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager.

4. We note that Figure 2 in your submission contain map images which may be copyrighted. All PLOS content is published under the Creative Commons Attribution License (CC BY 4.0), which means that the manuscript, images, and Supporting Information files will be freely available online, and any third party is permitted to access, download, copy, distribute, and use these materials in any way, even commercially, with proper attribution. For these reasons, we cannot publish previously copyrighted maps or satellite images created using proprietary data, such as Google software (Google Maps, Street View, and Earth). For more information, see our copyright guidelines: http://journals.plos.org/plosone/s/licenses-and-copyright.

We require you to either (1) present written permission from the copyright holder to publish these figures specifically under the CC BY 4.0 license, or (2) remove the figures from your submission:

A.  You may seek permission from the original copyright holder of Figures 1 and 2 to publish the content specifically under the CC BY 4.0 license.

We recommend that you contact the original copyright holder with the Content Permission Form (http://journals.plos.org/plosone/s/file?id=7c09/content-permission-form.pdf) and the following text:

“I request permission for the open-access journal PLOS ONE to publish XXX under the Creative Commons Attribution License (CCAL) CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). Please be aware that this license allows unrestricted use and distribution, even commercially, by third parties. Please reply and provide explicit written permission to publish XXX under a CC BY license and complete the attached form.”

Please upload the completed Content Permission Form or other proof of granted permissions as an "Other" file with your submission.

In the figure caption of the copyrighted figure, please include the following text: “Reprinted from [ref] under a CC BY license, with permission from [name of publisher], original copyright [original copyright year].”

B. If you are unable to obtain permission from the original copyright holder to publish these figures under the CC BY 4.0 license or if the copyright holder’s requirements are incompatible with the CC BY 4.0 license, please either i) remove the figure or ii) supply a replacement figure that complies with the CC BY 4.0 license. Please check copyright information on all replacement figures and update the figure caption with source information. If applicable, please specify in the figure caption text when a figure is similar but not identical to the original image and is therefore for illustrative purposes only.

The following resources for replacing copyrighted map figures may be helpful:

USGS National Map Viewer (public domain): http://viewer.nationalmap.gov/viewer/

The Gateway to Astronaut Photography of Earth (public domain): http://eol.jsc.nasa.gov/sseop/clickmap/

Maps at the CIA (public domain): https://www.cia.gov/library/publications/the-world-factbook/index.html and https://www.cia.gov/library/publications/cia-maps-publications/index.html

NASA Earth Observatory (public domain): http://earthobservatory.nasa.gov/

Landsat: http://landsat.visibleearth.nasa.gov/

USGS EROS (Earth Resources Observatory and Science (EROS) Center) (public domain): http://eros.usgs.gov/#

Natural Earth (public domain): http://www.naturalearthdata.com/

5. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information.

6. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

Reviewer #4: Yes

**********

-->2. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: No

Reviewer #4: Yes

**********

-->3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

Reviewer #4: Yes

**********

-->4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: No

Reviewer #4: Yes

**********

-->5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: The paper presents an interesting application of machine learning to the selected dataset. The topic is relevant and timely, and the manuscript is generally well organized. However, several methodological and reporting aspects require clarification or further development to enhance the rigor, interpretability, and overall contribution of the study.

Literature Review:

The introduction provides a useful overview of related work, but it would benefit from including the accuracy or performance metrics reported in previous studies. This would help contextualize the authors’ results and allow for direct comparison of findings. Additionally, if any machine learning analyses have been previously conducted on the same dataset, these should be reviewed and referenced. The authors could then discuss their results in comparison with those prior findings in the discussion section.

Use of SHAP and Feature Engineering:

The paper correctly highlights the use of SHAP for model interpretability. However, it would be beneficial to explain how SHAP can be leveraged to improve model performance, not just interpretation. Similarly, since the paper mentions feature engineering, the authors could explore how SHAP and feature engineering together might be used to exclude unimportant features, which could potentially improve model accuracy or stability.

Use of SMOTE:

The authors apply SMOTE but do not clearly explain why it was required. The manuscript should describe the class imbalance in the dataset and justify the need for oversampling. Providing before-and-after class distributions or comparing model performance with and without SMOTE would make this methodological choice more transparent.

Hyperparameter Tuning:

The manuscript mentions the use of Grid Search for hyperparameter tuning but does not provide sufficient detail about the process. The authors should specify which parameters were optimized, the ranges or values tested, and how the tuning process was validated (for example, cross-validation approach or performance metric used for selection). It would also strengthen the study to include performance metrics before and after tuning to show the effect of this optimization.

Variable Analysis and Visualization:

Table 2 presents variable distributions and percentages, which is informative. However, this section could be enhanced by exploring interactions between key variables. For example, examining the relationship between “Division” and “Source of Drinking Water”—identified as important features by SHAP—could reveal deeper insights. Adding visualizations such as heatmaps or grouped bar charts based on two-variable analyses would add another dimension to the findings and improve interpretability.

Figure Captions:

Several figure captions are overly long. It would be preferable to move part of the explanation to the main text, keeping the captions concise and focused on describing what the figures show rather than providing detailed interpretations.

Model Performance and Discussion:

The authors report multiple performance metrics, including F1-score, precision, recall, and AUC, which is commendable. However, the discussion should further address whether the achieved accuracy (approximately 0.7) is acceptable or competitive within their specific research field. Comparing this performance to similar studies would provide important context for evaluating the model’s effectiveness and practical applicability.

Reviewer #2: The manuscript presents an important application of classical and deep learning models for predicting E. coli contamination in household drinking water in Bangladesh using a nationally representative dataset.

Two Minor, Yet Scientifically Important Improvements

1. Clarify the rationale behind SMOTE for an already balanced dataset segment

- The reported prevalence of contamination is ~39% vs 61%. This is moderately imbalanced, not severely.

- Please briefly justify the use of SMOTE and discuss whether oversampling may have artificially inflated model performance or introduced bias, especially since several models showed modest AUC values (<0.70).

2. Add a short explanation of why AdaBoost outperformed other models

- AdaBoost achieved the highest accuracy and F1 score, but the manuscript does not interpret why it may be more suitable for this problem (e.g., handling weak learners, sensitivity to decision boundaries, feature interactions).

- A 2–3 sentence interpretation would strengthen the Discussion and provide insight for practitioners choosing ML techniques for similar epidemiological datasets.

Reviewer #3: This research represents an important scientific contribution in the field of public health and drinking water quality .This is achieved by applying advanced machine learning algorithms to nationally representative population data, a successful experiment in Bangladesh for predicting the risk of drinking water contamination ,but i have Nom. of comments

1-Some paragraphs are repetitive and need academic rewriting.

2-Add a comparison table between your results and the results of previous research.

3-Weaknesses in the wording of research contributions. What's new?

4- The statistical performance of the models was weak, and I suggest to improve by adding environmental variables related to potential pollution, such as:

• Distance from sewage treatment plants

• Population density

5- The accuracy of the results is important, and this observation is crucial. These values are considered very weak for sensitive health applications. Had the researcher used appropriate deep learning algorithms for their data, the results would have been better, as these are important findings.

Reviewer #4: The research was good in terms of the sources and samples used, but the practical aspect was not clear enough. The practical side needs further explanation; there should be a brief explanation of the functions used to classify the samples, as well as an explanation of how the system used in the sample analysis was evaluated.

**********

-->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

Reviewer #4: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

PLoS One. 2026 Aug 20;21(8):e0343606. doi: 10.1371/journal.pone.0343606.r002

Author response to Decision Letter 1


14 Jan 2026

Rebuttal Letter

Date: January 11, 2026

Dear Dr. Saif,

On behalf of all the authors, I wish to convey our gratitude for the critical and constructive feedback that we received from reviewers in our recent submission (Manuscript number: PONE-D-25-52920) entitled “Use of machine learning to detect Escherichia coli in drinking water in Bangladesh.” The manuscript has been revised in accordance with the feedback provided. All changes are highlighted using track changes in the revised manuscript. I hope our efforts satisfy the requirements of the journal. I will be looking forward to your positive response. Thank you for your time and consideration.

Iqramul Haq (on the behalf of authors)

Corresponding author

## The following is a point-by-point response to the reviewer(s) comments:

Reviewer #1: The paper presents an interesting application of machine learning to the selected dataset. The topic is relevant and timely, and the manuscript is generally well organized. However, several methodological and reporting aspects require clarification or further development to enhance the rigor, interpretability, and overall contribution of the study.

Author Response: Thank you for your comments. We have revised the manuscript following your suggestions.

Literature Review:

The introduction provides a useful overview of related work, but it would benefit from including the accuracy or performance metrics reported in previous studies. This would help contextualize the authors’ results and allow for direct comparison of findings. Additionally, if any machine learning analyses have been previously conducted on the same dataset, these should be reviewed and referenced. The authors could then discuss their results in comparison with those prior findings in the discussion section.

Author Response: We have updated the Introduction section. To the best of our knowledge, no previous study has used a machine-learning approach with this dataset, but alternative approaches were indeed used to analyze the same dataset. We added that to the introduction (pages 4-5).

Use of SHAP and Feature Engineering:

The paper correctly highlights the use of SHAP for model interpretability. However, it would be beneficial to explain how SHAP can be leveraged to improve model performance, not just interpretation. Similarly, since the paper mentions feature engineering, the authors could explore how SHAP and feature engineering together might be used to exclude unimportant features, which could potentially improve model accuracy or stability.

Author Response: Thank you for your feedback. We have incorporated this information into the SHAP subsection within the Statistical Analysis section (pages 7-8).

Use of SMOTE:

The authors apply SMOTE but do not clearly explain why it was required. The manuscript should describe the class imbalance in the dataset and justify the need for oversampling. Providing before-and-after class distributions or comparing model performance with and without SMOTE would make this methodological choice more transparent.

Author Response: We have made the necessary changes to the manuscript. We also evaluated model performance metrics before and after SMOTE in the revised version (with SMOTE: pages 10, and 16-19, without SMOTE: see the supplemental pages 3-4).

Hyperparameter Tuning:

The manuscript mentions the use of Grid Search for hyperparameter tuning but does not provide sufficient detail about the process. The authors should specify which parameters were optimized, the ranges or values tested, and how the tuning process was validated (for example, cross-validation approach or performance metric used for selection). It would also strengthen the study to include performance metrics before and after tuning to show the effect of this optimization.

Author Response: Thank you for your comment. We addressed this in the Hyperparameter Tuning subsection of the Statistical Analysis section (page10). We also compared model performance before and after hyperparameter tuning (with hyperparameter tuning: pages 16-19, without hyperparameter tuning: see the supplemental pages 3-4)

Variable Analysis and Visualization:

Table 2 presents variable distributions and percentages, which is informative. However, this section could be enhanced by exploring interactions between key variables. For example, examining the relationship between “Division” and “Source of Drinking Water”—identified as important features by SHAP—could reveal deeper insights. Adding visualizations such as heatmaps or grouped bar charts based on two-variable analyses would add another dimension to the findings and improve interpretability.

Author Response: We have revised and added this information to the Cramer's V correlation subsection in the Results section (page 16).

Figure Captions:

Several figure captions are overly long. It would be preferable to move part of the explanation to the main text, keeping the captions concise and focused on describing what the figures show rather than providing detailed interpretations.

Author Response: That has now been adjusted for all figures. Thank you.

Model Performance and Discussion:

The authors report multiple performance metrics, including F1-score, precision, recall, and AUC, which is commendable. However, the discussion should further address whether the achieved accuracy (approximately 0.7) is acceptable or competitive within their specific research field. Comparing this performance to similar studies would provide important context for evaluating the model’s effectiveness and practical applicability.

Author Response: We have revised the Discussion section accordingly (pages 21-22). Note that, after models were revisited following suggestions from reviewers, model performance improved slightly after applying SMOTE, 10-fold cross-validation, and hyperparameter tuning.

Reviewer #2: The manuscript presents an important application of classical and deep learning models for predicting E. coli contamination in household drinking water in Bangladesh using a nationally representative dataset.

Two Minor, Yet Scientifically Important Improvements

1. Clarify the rationale behind SMOTE for an already balanced dataset segment

- The reported prevalence of contamination is ~39% vs 61%. This is moderately imbalanced, not severely.

Author Response: Thank you. The outcome was indeed moderately imbalanced. Our rationale for using SMOTE was to test its ability to increase accuracy, especially for the contaminated subset. We added the rationale for SMOTE in page 10. Note that we now compare this approach with the original imbalanced dataset (with SMOTE: pages 10, and 16-19, without SMOTE: see the supplemental pages 3-4).

- Please briefly justify the use of SMOTE and discuss whether oversampling may have artificially inflated model performance or introduced bias, especially since several models showed modest AUC values (<0.70).

Author Response: As explained above, we now compare the SMOTE approach with the original imbalanced dataset (with SMOTE: pages 10, and 16-19, without SMOTE: see the supplemental pages 3-4). Additionally, please note that SMOTE was used only in the training dataset -the validation folds and the final held-out test set remained non-synthetic.

2. Add a short explanation of why AdaBoost outperformed other models

- AdaBoost achieved the highest accuracy and F1 score, but the manuscript does not interpret why it may be more suitable for this problem (e.g., handling weak learners, sensitivity to decision boundaries, feature interactions).

Author Response: This has now been added to the Discussion section (page 21).

- A 2–3 sentence interpretation would strengthen the Discussion and provide insight for practitioners choosing ML techniques for similar epidemiological datasets.

Author Response: This has now been added to the Discussion section (page 21).

Reviewer #3: This research represents an important scientific contribution in the field of public health and drinking water quality .This is achieved by applying advanced machine learning algorithms to nationally representative population data, a successful experiment in Bangladesh for predicting the risk of drinking water contamination ,but i have Nom. of comments

1-Some paragraphs are repetitive and need academic rewriting.

Author Response: Thank you for your comments. We have the entire manuscript to the best of our ability.

2-Add a comparison table between your results and the results of previous research.

Author Response: Thank you. That has now been added to supplemental files (Table 2).

3-Weaknesses in the wording of research contributions. What's new?

Author Response: We have added this information to the last paragraph of the Introduction section (page 5).

4- The statistical performance of the models was weak, and I suggest to improve by adding environmental variables related to potential pollution, such as:

• Distance from sewage treatment plants

• Population density

Author Response: We agree that the performance was weak and that is discussed in the last paragraph of the Discussion section (page 23). Additionally, please note that we now have optimized our hyperparameter turning approach based on reviewers feedback, which has increased our overall accuracy from 70.2% to 81.6% (pages 16-19).

5- The accuracy of the results is important, and this observation is crucial. These values are considered very weak for sensitive health applications. Had the researcher used appropriate deep learning algorithms for their data, the results would have been better, as these are important findings.

Author Response: Author Response: Thank you for this observation. We agree that the reported performance is insufficient for sensitive, individual-level health decision-making, and we have revised the manuscript to clarify that these models are not intended for clinical use. We included a commonly used deep learning multilayer perceptron (DL-MLP), but its performance was not substantially better than the classical ML models (page 16-19. Additionally, we now have optimized our hyperparameter turning approach based on reviewers feedback, which has increased our overall accuracy from 70.2% to 81.6% (pages 16-19).

Reviewer #4: The research was good in terms of the sources and samples used, but the practical aspect was not clear enough. The practical side needs further explanation; there should be a brief explanation of the functions used to classify the samples, as well as an explanation of how the system used in the sample analysis was evaluated.

Author Response: Thank you for your comments. We added a paragraph discussing the practical aspects of the study. Additionally, we have revised the Methods section accordingly to include the functions used (pages 8-10).

Attachment

Submitted filename: Rebuttal Letter.docx

pone.0343606.s003.docx (25.1KB, docx)

Decision Letter 1

Sohail Saif

16 Apr 2026

-->PONE-D-25-52920R1-->-->Use of machine learning to detect Escherichia coli  in drinking water in Bangladesh-->-->PLOS One

Dear Dr. Nobrega,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================-->-->As discussed via email, we have rescinded the acceptance decision and returned the manuscript to you, so that you can make any necessary changes to the text and figures. Specifically, please add your updates to figure 3, figure 7, and the text of the manuscript. Once you resubmit your revised manuscript, it will be reviewed by the Academic Editor.

==============================

Please submit your revised manuscript by May 31 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only  the individual author can complete the verification step; PLOS staff cannot  verify ORCID iDs on behalf of authors.

We look forward to receiving your revised manuscript.

Kind regards,

Katherine Kokkinias, Ph.D.

Staff Editor, PLOS One

on behalf of

Sohail Saif, Ph.D

Academic Editor

PLOS One

Journal Requirements:

If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

[Note: HTML markup is below. Please do not edit.]

Reviewer's Responses to Questions

-->Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.-->

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

Reviewer #3: All comments have been addressed

**********

-->2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->3. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: Thanks for addressing my previous comments which I think improved the paper so much. I don’t have any further comments.

Reviewer #2: The authors have adequately addressed the major and minor concerns raised during the previous round of review. Key methodological clarifications, including the justification and controlled use of SMOTE, expanded hyperparameter tuning details, improved SHAP-based feature selection explanation, and strengthened discussion on model performance, have been satisfactorily incorporated. The manuscript is now substantially improved in clarity, rigor, and transparency.

Reviewer #3: (No Response)

**********

-->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review?  For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

Decision Letter 2

Sohail Saif

10 May 2026

Use of machine learning to detect Escherichia coli  in drinking water in Bangladesh

PONE-D-25-52920R2

Dear Dr. Nobrega,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Sohail Saif, Ph.D

Academic Editor

PLOS One

Additional Editor Comments (optional):

Response is satisfactory.

Reviewers' comments:

Acceptance letter

Sohail Saif

PONE-D-25-52920R1

PLOS One

Dear Dr. Nobrega,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Sohail Saif

Academic Editor

PLOS One

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 File. Supporting information figures and tables (S1-S6 Figs; S1-S2 Tables).

    (DOCX)

    Attachment

    Submitted filename: Rebuttal Letter.docx

    pone.0343606.s003.docx (25.1KB, docx)
    Attachment

    Submitted filename: response.pdf

    pone.0343606.s004.pdf (86.8KB, pdf)

    Data Availability Statement

    The minimal derived datasets required to reproduce the results reported in this manuscript are available on Figshare (https://doi.org/10.6084/m9.figshare.31063324). The underlying Bangladesh MICS 2019 household microdata are third-party controlled-access data provided by UNICEF and cannot be redistributed by the authors. Researchers may request access via https://mics.unicef.org/surveys.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES