Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 14.
Published in final edited form as: J Environ Manage. 2025 Sep 12;394:127210. doi: 10.1016/j.jenvman.2025.127210

Prediction and Optimization of Electrochemical Oxidation Efficiency Using Machine Learning: Insights from Carbon-Based Anodes for Water Treatment

Nima Sakhaee 1, Stephanie Sarrouf 1, Paul Kim 2, Jong-Gook Kim 3,*, Akram N Alshawabkeh 1
PMCID: PMC12867521  NIHMSID: NIHMS2135834  PMID: 40945418

Abstract

Electrochemical oxidation (EO) has emerged as a promising and environmentally sustainable strategy for the treatment of contaminated water, enabling the degradation of refractory pollutants through the in-situ generation of reactive oxidative species. Carbon-based anodes—such as graphite plates, carbon felt, and carbon fibers—are increasingly adopted in EO systems due to their cost-effectiveness, electrochemical stability, and environmental compatibility. However, despite their widespread application, a systematic understanding based on data-driven modeling of how different types of unmodified carbon-based anodes influence pollutant removal under varying operational conditions remains lacking . In this study, we applied nine machine learning (ML) algorithms to evaluate the effects of both material and operational parameters on EO performance using carbon-based anodes. A comprehensive dataset of over 1,400 experimental entries was analyzed, encompassing variables such as anode type, current density, pH, electrolyte concentration, pollutant characteristics, and reaction time. Among all models tested, Light Gradient Boosting Machine (LightGBM) exhibited the highest predictive accuracy (R² = 0.926; RMSE = 8.846). SHAP-based feature interpretation revealed that operational factors—particularly reaction time, pollutant type, and current density—had the greatest influence on removal efficiency. In contrast, the specific type of unmodified carbon-based anode had minimal impact under the conditions studied, likely due to the inherent similarity in their electrochemical behavior. This trend may yield different results in systems employing surface-modified electrodes or direct electron transfer mechanisms. These findings emphasize that optimized operational conditions play a more decisive role than material selection in enhancing EO performance. This study presents a data-driven framework to guide EO system optimization, reduce experimental trial-and-error, and support the development of scalable, high-efficiency water and wastewater treatment technologies.

Keywords: Electrochemical oxidation, Machine learning, Carbon-based anode, Light Gradient Boosting Machine (LightGBM), Water treatment

1. Introduction

Water contamination by organic and inorganic pollutants is one of the great concerns, necessitating sustainable treatment technologies. Various methods have been paid attention for water treatment including coagulation, adsorption, advanced oxidation processes, and photocatalysis. Recently, the electrochemical degradation of contaminants has garnered significant attention due to its high efficiency in pollutant removal and in-situ generation of powerful oxidants (Cerrillo-Gonzalez et al., 2025; Sirés et al., 2014; Soares et al., 2024). Various reactive oxygen species, such as hydrogen peroxide (H2O2), hydroxyl radicals (•OH), superoxide radical (O2•−), and singlet oxygen (1O2), are generated within the reactor (Jasmann et al., 2016; Kim et al., 2023), which can effectively mineralize organic pollutants (Jasmann et al., 2016; Kim et al., 2023). The oxidized species will accept an electron and become reduced. For instance, the oxygen evolution reaction (OER) happens at anode. Moreover, the oxygen can undergo electrochemical reduction to produce hydrogen peroxide through a two-electron pathway at cathode. The half reactions for both of them are mentioned at Eq (1) and Eq (2), respectively (Jung et al., 2020). Figure S1 demonstrates a simple electrochemical system, including anode and cathode, and the corresponding half reactions happening at each electrode are also included. Additionally, this process can be utilized to transform inorganic pollutants and heavy metals into less toxic forms, further enhancing its applicability in water purification (Kim et al., 2024b; Wang et al., 2023). Also, they have versatility and adaptability when integrated with other technologies (Chaplin, 2019), energy efficiency and sustainability (Kim et al., 2018), as well as simplicity in control and ease of setup (Valica and Hostin, 2016).

H2O→12O2+2H++2e- (1)
O2+2H++2e-→H2O2 (2)

This method is praised for its energy efficiency, environmental compatibility, and adaptability with other systems. Its performance in electrochemical systems can be influenced by the choice of electrode material. Nevertheless, in this study, where only unmodified carbon-based anodes were considered, the electrode type showed minimal influence on degradation efficiency. In particular, the anode materials are crucial for determining removal efficiency and mechanism (Kim et al., 2024a). Traditional anodes (e.g., PbO2, and SnO2), which are composed of metals, exhibit excellent oxidation capabilities but suffer from high costs, scalability challenges, and metal leaching into the water environment, raising environmental concerns (Kim et al., 2024a). To overcome these limitations, carbon-based anodes have emerged as considerable alternatives due to their high surface area, stability, and cost-effectiveness (Kim et al., 2022). According to several studies, the anodic reactions of carbon-based anodes involve not only the generation of hydroxyl radicals but also direct electron transfer (DET) and reactions mediated by oxygen-containing functional groups (Abdalrhman et al., 2019; Periyasamy and Muthuchamy, 2018). Carbon-based anodes, such as graphite plates, graphite felt, and carbon fiber, have demonstrated effectiveness in removal efficiency compared to metal-based anodes (Chen et al., 2021; Kim et al., 2024a). The intrinsic properties of carbon-based anodes such as the tunable surface chemistry, electrical conductivity, and the promotion of both direct electron transfer (DET) and the generation of reactive oxidative species (ROS), make them emergent for environmental applications (Mauter and Elimelech, 2008; Xu et al., 2013). While carbon-based anodes offer tunable surface chemistry and electrochemical properties, this study focused on unmodified carbon materials to reduce variability and allow a controlled comparison of operational parameters. Different carbon-based anodes exhibit distinct physicochemical properties—such as surface area, surface functional groups, and electrical conductivity—which critically influence the generation of reactive oxygen species, pollutant adsorption capacity, and overall electron transfer efficiency during electrochemical oxidation (Aji et al., 2025). Nevertheless, there is limited research on which carbon-based anodes and operating conditions most significantly affect removal performance. In particular, it is unclear which types of carbon-based anodes perform best under which conditions.

In addition to the anode material, various experimental factors influence the removal mechanisms and efficiency (Coria et al., 2016). The electrochemical degradation of water encompasses a broad spectrum of physical and chemical properties, each influencing the overall efficiency of the process (Tarpeh et al., 2018). Although substantial effort has been directed at examining the effect of individual parameters on system efficiency, there is a notable gap in understanding how multiple parameters interact concurrently, and the impact of simultaneous multi-factor analysis remains underexplored. Additionally, experimentally determining the relative contributions of each parameter and assessing the influence of one parameter on others can be highly time-consuming and challenging (Franceschini and Macchietto, 2008). Especially, given the numerous input variables in electrochemical water treatment systems, including current density, electrode spacing, reaction time, electrode type, electrolyte concentration and pH, identifying an optimal configuration through conventional methods can be exceedingly challenging (Ditta et al., 2023; Körbahti and Erdem, 2017; Narkevica et al., 2014). To handle complex systems, the Design of Experiments (DoE) method is used, which statistically evaluates the effects of multiple inputs on outcomes. A subtype, Definitive Screening Design (DSD), identifies key variables and interactions with fewer experimental runs (Tai et al., 2015). However, DSD relies on the sparsity assumption—that only a few parameters significantly affect the outcome—which may limit its usefulness in complex real-world systems (Aguirre, 2016). Also, implementing and analyzing DSD experiments often depends on commercial software or high coding skills, which can pose a barrier for many researchers (Hayasaka et al., 2024). Response Surface Methodology (RSM), another subset of DoE, combines statistical and experimental techniques to model the relationship between input and output variables using fitted surfaces (Youssefi et al., 2009). Although widely used, RSM has limitations such as limited accuracy and suitability for only a small number of input features (Al-Ani, 2023). Moreover, RSM methods are imprecise in highly complex and non-linear systems (such as electrochemical degradation) (Ranjan et al., 2011), and since RSM uses second-order polynomial models, it is highly sensitive to outliers, which can significantly reduce prediction accuracy (Widodo et al., 2015).

In this context, machine learning (ML) models can be valuable tools to address the aforementioned challenges. Data-driven models have become increasingly popular across various scientific disciplines due to their ease of use, high reliability, and precision in making predictions (Alsolai and Roper, 2020; Hu et al., 2023). These attributes have positioned data-driven modeling as a valuable tool in many engineering domains, particularly in water treatment systems (Gupta et al., 2021; Kim et al., 2025). ML is one of these powerful approaches that leverage available data to make accurate predictions (Zhou et al., 2017). In ML techniques, unlike RSM and second order polynomial expression, various computational techniques are used for prediction, where they can capture complex, non-linear relationships without requiring predefined equations (Zhang and Wu, 2021). Many studies have found that ML models outperform RSM models in terms of accuracy and fit to experimental data, suggesting that ML models offer superior predictive capabilities compared to traditional approaches (Hashemi et al., 2022; Ram Talib et al., 2019; Swain et al., 2024).

The existing literature has explored the application of ML techniques to predict and optimize the performance of various materials and processes for the removal of pollutants from wastewater (Taoufik et al., 2022; Ye et al., 2025). Several studies have focused on using ML models to predict the adsorption capacity and removal efficiency of pollutants, such as heavy metals (including arsenic) and organic contaminants on materials like biochar, bentonite, and cobalt-based oxide catalysts (Guo et al., 2024; Jaffari et al., 2023; Jiang et al., 2024; Liu et al., 2023; Zhang et al., 2023). Data-driven ML models have also been employed to predict the removal efficiency of various pollutants, such as perfluorooctanoic acid, in electrochemical degradation processes, demonstrating high accuracy and identifying key operational parameters (Alnaimat et al., 2024). In addition to predicting pollutant removal, ML models have also been applied for multivariate optimization—such as combining neural networks with genetic algorithms—to maximize electrochemical removal of COD and TN from wastewater with high accuracy and stability (Yang et al., 2013). While prior studies have used ML to model adsorption processes, few have applied ML to electrochemical oxidation systems, especially focusing on carbon-based anodes.

This study aims to bridge this gap by applying multiple ML algorithms to a dataset of over 1,400 experimental data from electrochemical remediation systems. And to evaluate the relationship with carbon-based anodes and contaminants removal in electrochemical reaction. The objectives are to (i) identify the most accurate predictive model, (ii) determine the key features influencing removal efficiency, and (iii) provide insight into the performance of various carbon-based anodes under different operational conditions. Through this approach, the authors aim to support the development of more efficient and sustainable electrochemical treatment technologies. This approach holds the potential to significantly advance the field, contributing to the development of more efficient, sustainable, and environmentally friendly technologies for pollutant removal.

2. Materials and Methods

The methodology of this study starts with the collection of data, followed by the necessary preprocessing steps prior to training the ML models. After this initial phase, hyperparameter tuning is conducted to optimize each model’s performance. After optimization, SHAP importance analysis and partial dependency plots are generated to interpret the model outputs and understand the influence of individual features. Figure 1 illustrates a process diagram that encapsulates the methodology, providing a concise overview of the sequential steps from start to finish. Detailed discussions of each step are presented in subsequent sections.

Figure 1.

Figure 1.

Process flow diagram illustrating the methodological steps

2.1. Data collection

Input data were collected from a review of several relevant studies published between 2016 and 2024, except one paper from 2008, highlighting that the data used in this study are both recent and reliable. In total, 19 scientific papers were meticulously analyzed to extract the necessary data. A structured literature search was conducted using Web of Science and Scopus databases, with the keywords “electrochemical oxidation + carbon anode + pollutant removal,” and a search window spanning 2000–2024. During data cleaning, duplicate and inconsistent entries were removed, and typographical errors were corrected through cross-checking with the original articles. Outliers exceeding three standard deviations from the mean were excluded. Additional details about the collected data and their corresponding references are provided in Table S1. For visual data presented in graphs, primarily depicting degradation over time, data points were extracted using WebPlotDigitizer, an online tool designed to digitize information from figures (Dashti et al., 2024).

During a comprehensive review of these papers, the key parameters of interest were categorized into two groups: categorical parameters, including anode and cathode type, reactor type, type of electrolyte, pollutant type, and numerical parameters, including electrode area (cm2), effective area (cm2), electrode spacing (cm), reactor volume (mL), electrolyte concentration (mM), initial pH, current density (mA/cm2), pollutant concentration (mg/L), reaction time (min), and the output which is removal content of the pollutant (%). All electrochemical oxidation experiments in this study operated under constant current conditions, where the applied voltage is automatically adjusted by the DC power source to maintain the set current. As a result, exact voltages were rarely reported in the source studies and were excluded from our analysis. Various anode and cathode types were used in our dataset. Common carbon-based anodes, including graphite felt, graphite plate, carbon fiber, and activated carbon, were selected based on their unmodified form without surface functionalization, metal doping, or catalyst loading. Boron-doped diamond (BDD) and reduced graphene oxide (rGO) were excluded due to their limited practicality; BDD requires expensive boron doping, and rGO involves strong acids and complex chemical processes (He et al., 2019; Lowe et al., 2019). However, the graphite emerged as the common utilized anode material in these experiments (Periyasamy and Muthuchamy, 2018). For cathodes, a greater diversity of materials was observed, including stainless steel, which accounted for a significant proportion of the choices. More than 70% of the reactors used were undivided. The types of electrolytes and pollutants were also diverse. We included a variety of pollutants, such as methylene blue, sulfanilamide, and ibuprofen, among others, to investigate the significance of pollutant type and explore potential valuable trends related to this factor.

In our dataset, the reactor type refers to whether the reactor is divided or undivided. A divided reactor features an anode and cathode separated by a membrane, whereas in undivided reactors, the anode and cathode share the same electrolyte. Undivided systems offer simpler operation and cost-efficient experimentation. On the other hand, divided reactors provide better control over reaction conditions, enhancing selectivity and reaction management. They often face challenges such as reduced current density and increased design complexity. In contrast, undivided reactors simplify reactor design and operation, improve mass transfer rates, and reduce costs, making them a favorable choice for many researchers (Folgueiras-Amador et al., 2022).

2.2. Pearson correlation coefficient matrix

To quantitively measure the linear correlation between each pair of data, the Pearson correlation coefficient (PCC) analysis was performed (Zhu et al., 2020). It consists of a square matrix, and each element represents the Pearson correlation coefficient between the two variables. The mathematical expression for PCC is presented in Eq (3):

ρx,y=∑i=1N(xi-x-)(yi-y-)∑i=1N(xi-x-)2∑i=1N(yi-y-)2 (3)

Where ρx,y is the correlation coefficient between x and y. The elements of this matrix have values between 1 and −1. The diagonal elements of this matrix are obviously 1, and a correlation close to 1 implies a strong positive correlation between the two variables, and an increase in one will result in an increase in the other one. On the other hand, values close to −1 imply a negative correlation opposite to the previous condition. Conversely, ρx,y=0 indicated no correlation between the two parameters.

2.3. Pre-processing

2.3.1. Imputation of Missing Data (K-nearest neighbor)

While working with large datasets, encountering missing data is inevitable. For instance, in one dataset from Song et al., all necessary experimental parameters were recorded except for the initial pH values (Song et al., 2018). This dataset comprises approximately 50 subsets; including these would likely enhance the accuracy of our machine-learning model and improve our predictions. However, the absence of pH values renders these subsets initially unusable for our analysis. While the easiest solution might be removing the incomplete data sets from the system, this would result in removing valuable data that could be useful in the final prediction of the ML model, despite having some parts. Imputation is a strong and reliable method to add the missing data to the system based on the available data. The total number of missing data points are 2023 data points, from a total of 22710 points, which makes it less than 9% of total data (refer to table S2). Several methods have been proposed to address the problem of missing data, including mean or median imputation. However, these approaches often fail to accurately reflect the underlying data distribution, as they tend to underestimate variability and can introduce bias. This is particularly problematic in cases with substantial missing data, where a large portion of values may be replaced with a single constant, leading to distorted and unreliable estimations (Ren et al., 2023). On the other hand, K-nearest neighbor imputation has emerged as a compelling alternative to methods like mean and median, and it consistently outperforms these traditional methods (Sania et al., 2021) and has been applied in environmental engineering studies to address missing data (Quinteros et al., 2019). In this algorithm, the value of missing data is determined based on the k nearest neighbors in the data set, where k is adjustable by the user. Choosing a very small value for k (e.g., k=1) will result in less generalization, and potentially overlooking some valuable information, and choosing a large value (e.g., k=20) risks underfitting by diluting the influence of closer, more relevant data points with distant, less relevant ones(Thomas and Rajabi, 2021). In this study, we selected k=5, following values commonly used in the literature.

2.3.2. Normalization

While working with large data sets in ML analysis, it is crucial to normalize the input values to eliminate the possibility of errors due to the presence of huge differences in the scale of inputs, both numerical and categorical ones (Apicella et al., 2023). For instance, pH values typically range between 0 and 14, representing a relatively narrow numerical scale, whereas electrolyte concentrations in our study can reach up to 500 mM. In a linear regression model utilizing only these two features, the coefficient for electrolyte concentration might be considerably smaller than that for pH. This discrepancy complicates interpretability, as it is unclear whether the smaller coefficient indicates a lesser impact or merely reflects the difference in scales. Additionally, many machine learning algorithms, such as gradient descent—which aims to minimize the loss function—perform more effectively when the features are normalized. Ensuring that data values are on comparable scales enhances the model’s ability to make more reliable and accurate predictions. Tree-based models are insensitive to the scale of features because they make decisions by splitting data based on threshold values. These models inherently identify the optimal threshold for each feature during training, regardless of its magnitude or unit. But models like linear regression, KNN or bagging (If the weak learner is not decision tree) need normalization. In our study, for models that were not tree-based, the input data were normalized to ensure that all features were on comparable scales, thereby enhancing the models’ ability to produce more reliable and accurate predictions. The normalization was performed using Eq (4):

zi=xi-x-σ (4)

Where xi is the initial value of the input parameter, zi is the normalized value, x- is the mean and σ is the standard deviation of all values of x. This will linearly transform all the input values to a new data set with a mean of zero and does not alter the relative distribution of the points, ensuring that the relative distance is preserved.

2.3.3. Performance metrics

To quantitatively assess the performance of each model, two numerical indexes were used, including the coefficient of determination (R2) and root mean square error (RMSE) (Garosi et al., 2019). The ideal value for R2 is 1, and for RMSE is 0. A model is considered more accurate in predicting the outcome if its values are closer to these ideal metrics in both training and testing data sets. The formulas for calculating these metrics are provided as Eq (5) and Eq (6):

R2=1-∑i=1N(Ei-Pi)2∑i=1N(Ei-E-)2 (5)
RMSE=∑i=1N(Ei-Pi)2N (6)

Where Ei is the experimental value of output collected from the dataset, Pi is the predicted value generated from the model, P- indicates the mean of all experimental data, and N is the number of data points.

2.4. Development of ML models

9 different ML models, including Extreme Gradient Boosting (XGboost), Categorical Boosting (Catboost), Random Forest, Light Gradient Boosting Machine (LightGBM), Linear regression, K-nearest neighbors (KNN), Bagging, Gradient boosting (Gradboost), and Extra tree were chosen for this study. A brief description of each ML model is provided in the Supplementary Information file.

While training ML models, the data set will randomly split into training and testing sets, where we chose 80% of the data to be used for training each model and the rest 20% to be used for testing and validation. 3 models in our study, including Catboost, LightGBM and Gradboost can handle both numerical and categorical inputs internally. For the rest of the models, one-hot encoding technique was used. In this approach, each category within a feature is transformed into a binary column. Thus, for any given categorical feature (e.g., anode type such as graphite), a corresponding data entry is marked with a ‘1’ if it matches the category and ‘0’ otherwise. This method proves more effective than ordinal labeling for managing mixed datasets. Ordinal encoding is prone to give a false bias to higher values, for instance, for two random pollutants in the model, none of them have any priority or advantage over the other one. But assigning numbers to them will lead to sorting bias — a false advantage is given to the one with higher values (Almajid, 2021)

In this study, the grid search method was employed to determine the optimized hyperparameters for each model. Although there are several hyperparameter tuning methods available, such as random search and gradient-based search, we opted for grid search. Given the high dimensionality of our feature space and the nonlinear relationships among the features, a grid search seemed advantageous for comprehensively exploring the range of parameters to identify the optimal set of hyperparameters for each model. While a random search may offer benefits in terms of computational efficiency, it risks overlooking the most effective hyperparameter combination. Furthermore, gradient-based methods, which navigate towards the optimal value (e.g., highest R2 value), may become trapped in local optima, thus failing to yield the best overall set of hyperparameters. The model that gives parameter combination with the best performance is selected as the optimized model (Ghawi and Pfeffer, 2019). Additional details on this process are provided in the supplementary information file.

2.5. Feature importance analysis

One of the key applications of ML models is to quantitatively assess the impact of individual features on the outcome (Wang et al., 2024). While the PCC matrix provides a correlation value between each feature and the result, it is inherently linear and, therefore, fails to account for the combined effects of features on the outcome. This limitation is particularly significant in complex electrochemical systems, where linearity and linear relationships are not valid assumptions.

A promising tool for this process is SHAP (SHapley Additive exPlanations), a method that explains predictions by attributing contributions of individual features to the outcome, accounting for both positive and negative impacts (Yang et al., 2024). SHAP assigns each feature an importance value for a specific prediction, known as the SHAP value (Lundberg and Lee, 2017). This method assumes the actual ML model, f(x) can be approximated with a linear function g(x′):

fx=gx′=∅0+∑i=1n∅ix′i (7)

Where x′={0,1}n is a n×2 binary matrix and n is the number of features in the data. ∅i is feature attribution value, also known as SHAP value, which can be calculated using Eq (8):

∅i=∑S∈F\{i}S!F-S-1!F![fS∪ixS∪i-fS(xS)] (8)

where F is the set of all features, S is a subset of all features excluding i and f(S) is the prediction of ML model using the features available in S. As mentioned, the difference fS∪ixS∪i-fS(xS) will be calculated for all possible subsets without i and then normalized (Wang et al., 2022). SHAP value represents the contribution of each feature in the final prediction of the output (Refer to Eq (7)), and thus, by relatively comparing the magnitudes and the sign of these values for each individual feature, we can gain a quantitative insight about their importance. The magnitude indicates the strength of a feature’s influence, while the sign reveals whether the feature drives the prediction upward or downward.

2.6. Partial dependency plot analysis

Partial dependency plot (PDP) analysis is a method used to visualize and explain the dependency between one or several features and the target variable in a trained ML model (Park et al., 2022). In this method, a single feature is varied across its range while keeping all other features constant, and the model’s output is calculated. This process will generate a single plot, showing the relationship between one instance and the target parameter. To create a Partial Dependency Plot (PDP), the process is repeated for all instances in the dataset, and the results are averaged. The resulting plot represents the average dependency of the feature (e.g., current density) on the target variable (e.g., removal efficiency percentage) while holding all other features constant. This provides a clear and interpretable visualization of how changes in a feature influence the target value. The partial function, fS, is determined using Eq (8):

fSxS=1n∑i=1nf(xS,xCi) (9)

Where xS is the feature of interest for which the relationship with the target parameter is required varied over its range, and xCi is the set of features all of the features except xS which are held constant (Goldstein et al., 2015).

3. Results

3.1. Data analysis

Both numerical and categorical data were utilized to train the ML models. Figure 2 presents a bar chart illustrating the categorical input features, highlighting the distribution of each feature type within the dataset. Figure 3 demonstrates a violin plot for each numerical data, which is commonly used to illustrate the distribution of data. These plots provide valuable insights into the data by visualizing its density and variability. The box in the center of the violin plot represents the interquartile range (IQR), spanning from the first quartile (Q1), which exceeds 25% of the data values, to the third quartile (Q3), which is higher than 75% of the data. The line within the box indicates the median value, representing the midpoint of the dataset. Efficiency, considered the output variable in our ML models, ranges from 0 to 100%. The spacing between electrodes varies from 1 to 14 cm, with most data points being less than 6 cm. Current density is predominantly below 10 mA/cm2, likely because exceeding the limiting current threshold does not enhance the electrochemical reaction rate. Reactor volume ranges from 100 mL to approximately 1L. Electrolyte concentrations are typically below 100 mM. Pollutant concentrations also show variability, with over 75% of the data having concentrations below 50 mg/L, though some values reach as high as 700 mg/L. This diversity in input variables ensures that the ML models are not biased toward a narrow range of data, enabling robust predictions across a wide range of data.

Figure 2.

Figure 2.

Different types and distributions of categorical data used in the models for a) Anode, b) Cathode, c) Reactor type, d) Electrolyte, and e) Pollutant

Figure 3.

Figure 3.

Violin plots for the distribution of a) Effective area, b) Reactor volume, c) Spacing between electrodes, d) Current density, e) Electrolyte concentration, f) Pollutant concentration, g) initial pH, and h) Contamination removal efficiency

Figure 4 illustrates the Pearson correlation coefficient matrix for our data set. As previously mentioned, we utilized one-hot encoding for our ML models to eliminate any inherent ordinality that could potentially introduce bias. However, for the Pearson Correlation Coefficient matrix, which assesses linear relationships, we employed label encoding. This approach allowed us to gain preliminary insights into the linear correlations among the data. The current density exhibits a strong correlation with removal efficiency, underscoring its critical role in generating oxidizing agents that facilitate the degradation of contaminants (Niu et al., 2016). This highlights the significance of optimizing current density to enhance the overall system performance. Similarly, the initial pollutant concentration shows a high correlation factor, as anticipated. A higher initial contamination level poses greater challenges for the system due to the increased demand on degradation rate and extended reaction time required within the reactor. A low negative value of correlation can be seen for the spacing value. This can be attributed to the constrained space between the two electrodes. This limitation potentially affects the generation of oxidizing agents or restricts the space available for the degradation of pollutants. The anode type exerts minimal influence on the final removal efficiency, suggesting that carbon-based anodes have a negligible impact on degradation outcomes. Instead, other factors, particularly experimental parameters such as current density and pollutant concentration, play more significant roles. A negligible correlation coefficient is also observed for the cathode type, suggesting that optimizing other experimental conditions is more effective for achieving higher removal percentages than focusing on the physical characteristics of the electrode, such as its material.

Figure 4.

Figure 4.

Pearson correlation coefficient matrix for each combination of variables

While these factors are all linear and serve as useful indicators of the water treatment system’s characteristics, the fundamental assumption of input independence—where each input is considered to have no effect on the others—overlooks the combined influence of features on the final prediction. This simplification can undermine the model’s accuracy by neglecting the complex interplay among variables. This limitation underscores the importance of employing ML methods. ML approaches account for the inherent dependencies between inputs and their collective impact on the final prediction, thereby enhancing the model’s ability to capture the intricate relationships within the dataset.

3.2. Model performance evaluation

9 ML models which described earlier were employed to predict the removal efficiency percentage based on the input features. The R2 value is calculated for the predicted and actual value of each data set in testing data set to show a quantitative evaluation of the performance of each model. Figure 5 shows regression plots for predicted removal content and actual value which is gathered from the data set. The straight line represents the ideal line (y=x) signifying an absolutely accurate prediction. The closer the data points are to the y=x line on a regression plot, the better and more accurate the model is. In such cases, the R2 value approaches 1, indicating a high level of predictive accuracy.

Figure 5.

Figure 5.

Regression plots and data distribution of actual vs predicted removal content equipped with optimal hyperparameters for a) XGboost, b) Catboost, c) Random Forest, d) LightGBM, e) Linear Regression, f) KNN, g) Bagging, h) Gradboost, and i) Extra Tree

The accuracy of the models varies with each other, reflecting the differences in their methodologies for predicting contamination removal. Table 1 summarizes the performance of each ML model. Typically, the R² value for the training dataset is higher than for the testing dataset, as the model is directly fitted to the training data. For this reason, the RMSE is also lower for the train compared to the test.

Table 1.

Summary of performance metrics for each ML model

Model R2 Train RMSE Train R2 Test RMSE Test
XGboost 0.996 2.185 0.900 10.292
Catboost 0.976 5.367 0.898 10.359
Random Forest 0.985 4.157 0.891 10.721
LightGBM 0.994 2.754 0.926 8.846
Linear Regression 0.273 29.250 0.243 28.262
KNN 0.997 1.807 0.719 17.230
Bagging 0.981 4.700 0.848 12.652
Gradboost 0.993 2.901 0.904 10.065
Extra Tree 0.997 1.748 0.899 10.336

LightGBM has the highest prediction accuracy, reaching to a test R2 value of 0.92. After that, we have Gradboost with R2 value of 0.90. This superior performance may be attributed to the ability of these models to internally handle categorical features, whereas other models require one-hot encoded datasets. Although one-hot encoding is a widely accepted method, models like LightGBM and Gradient Boosting are especially well-optimized to simultaneously handle both categorical and numerical features. Other tree-based methods, such as Random Forest, Extra Trees, and XGBoost, also demonstrate strong predictive performance, indicating that tree models are well-suited for our dataset. On the other hand, KNN, an instance-based learning method, did not exhibit strong performance, despite optimization across 20 different hyperparameters (Refer to Table S3). This suggests that such ML models may not be ideally suited for our project. Overall, most models achieve high regression coefficients, with R2 values exceeding 0.85. Linear regression, however, demonstrates poor performance and clear underfitting. This is expected, as the relationship between electrochemical characteristics and contamination removal is inherently nonlinear. Consequently, linear regression produced the lowest R2 value (0.243) among all the models evaluated. This confirms the discussion about the PCC part, where the values are not so reliable, due to the fact that linear assumption is not so accurate in a complex project like this. The slight variation in the results among different models can be attributed to their underlying prediction mechanisms. For example, tree-based models split the data based on feature values to minimize an error function, while KNN relies on a distance metric for prediction. Boosting models iteratively correct the errors of previous learners in a sequential manner (Schratz et al., 2019). Additionally, during hyperparameter optimization, different parameter configurations can result in varied predictions across models. Ensemble models combine the predictions of multiple weak learners, which can further contribute to differences in the final output (Ilemobayo et al., 2024). Further details regarding the implementation of each model are provided in the supplementary information.

3.3. Feature importance analysis (SHAP)

Figure 6 demonstrates the feature importance analysis results based on the results of our best model, LightGBM. Absolute SHAP values illustrate the average impact on the final output but do not indicate the direction—whether negative or positive—of each feature’s effect. As anticipated, reaction time exerts the most significant influence on the final removal efficiency of the system. Over time, a greater proportion of contaminants is removed, resulting in a progressively lower concentration of contaminants. Both pollutant type and pollutant concentration exhibit high SHAP values, indicating their substantial impact on the model. Each pollutant possesses unique chemical properties, and higher concentrations of pollutants are generally more challenging to degrade. Current density also has a relatively high impact on the final removal efficiency. The SHAP analysis revealed a lower impact of anode type, which may be attributed either to the similarity in physicochemical properties among the selected anodes or to the fact that all anodes used in this study were unmodified carbon-based materials. This observation does not inherently suggest that anode type has a negligible influence on removal efficiency in general. Rather, it reflects the specific scope of our study, wherein the findings and interpretations are applicable within the context of unmodified carbon-based anodes.

Figure 6.

Figure 6.

a) Mean absolute SHAP values, and b) net SHAP values for each feature.

The bar chart only shows the magnitude of values without the sign, and the highest one represents the feature with the highest impact on the final contamination removal, which here is reaction time with a positive impact. More comprehensive and detailed results for feature importance are presented in Figure 7. Categorical features, which do not exhibit higher or lower values, are depicted in grey. In contrast, the effects of increasing or decreasing numerical features are represented by their respective colors.

Figure 7.

Figure 7.

PDP plots for a) Electrode area, b) Current density, c) Effective area, d) Electrolyte concentration, e) pH, f) Pollutant concentration, g) reaction time, h) spacing, and i) reactor volume

3.4. Partial dependence plots (PDP)

Figure 7 presents the partial dependence plots for numerical features, by the best-performing ML model, LightGBM. By isolating individual features, the plots illustrate the average behavior of the final efficiency across the dataset, providing a general trend of how the final efficiency varies with changes in each feature

As expected, an increase in both electrode area and effective area generally enhances efficiency, up to a limiting point, beyond which further increases do not significantly affect efficiency. This behavior could be attributed to saturation effects or diminishing returns in the reaction kinetics, where there may not be any reactant molecules in the electrolyte solution to maintain the reaction, this leads to a mass transport limitation. Also, increasing the surface area beyond the limiting point could lead to overpotentials, where increasing the electrode area can also increase the current density to a level where resistive losses may happen. Pollutant concentration exhibits a strong negative yet non-linear relationship. This could be due to the fact that higher pollutant concentrations require more time and a greater quantity of reducing agents for effective degradation. Reactor volume shows only minor changes regarding the final contaminant removal, indicating that this feature is not a very important factor compared to others. By examining the PDP plot for initial pH, it is recommended to maintain a slightly acidic condition (pH between 5 and 6), where the dependency will increase by up to 10% compared to the neutral region. Increasing the spacing between the two electrodes, as expected, will have a considerable negative effect on the efficiency of the system. This should be highly considered when designing relevant experiments. Current density also exhibits a positive relationship in the PDP plot, eventually stabilizing at a plateau. This trend can be attributed to the initial increase in reactive oxygen species (ROS) generation, which enhances kinetics and mass transfer within the boundary layer. However, efficiency ceases to increase beyond a certain point, typically when mass transfer limits are reached, or excessive bubble formation impedes electrochemical reactions at the electrodes. Reaction time, by a considerable margin, proves to be the most effective factor in degradation. Since many degradation reactions follow first-order or second-order kinetics, signifying that the rate of degradation greatly depends on the concentration of contaminants. In addition, in some degradation processes, the contaminants may need time to reach a saturation point or equilibrium where the degradation process reaches its maximum efficiency. This underscores the importance of allowing sufficient time for the electrochemical system to degrade contaminants. The PDP plot displays a significant increase, indicating a direct relationship, which is a notable finding. Its substantial impact is highlighted when compared to other features discussed in this paper.

While this study focused on identifying key predictors of removal efficiency and comparing the performance of carbon-based anodes across a diverse dataset, it does not capture detailed chemical interactions due to variability in experimental conditions. Future work will address this gap by conducting systematic investigations into pollutant–electrolyte compatibility, pH effects, and mechanistic relationships using well-controlled experimental datasets.

4. Conclusion

In this study, ML models were applied to predict electrochemical contaminant removal using carbon-based anodes to identify the most influential factors affecting removal efficiency. A total of nine ML models, including LightGBM, Random Forest, and XGboost were employed to analyze key parameters such as current density, reaction time, and different types of carbon-based anodes. Among the evaluated models, LightGBM demonstrated superior predictive performance, achieving the lowest RMSE (8.846) and highest R² (0.926). These results underscore the model’s accuracy and reliability, making it an excellent candidate for further investigation and feature importance analysis. Based on the SHAP feature importance results, we identified valuable insights. They indicate that experimental conditions—specifically reaction time, pollutant type, and current density—have a more significant impact on removal efficiency than the specific type of anode or cathode material used. Our findings demonstrate that current density and electrode spacing, which are two experimental variables adjustable by operators, have a significant impact on contaminant removal efficiencies due to their relatively high SHAP values. By modifying these parameters, experiments can achieve higher pollutant degradation. While anode material is generally considered an important factor in electrochemical water treatment systems, the effect of carbon-based anodes on removal efficiency was found to be minimal. Conversely, key experimental factors such as current density and reaction time exhibited a higher impact. This could be attributed to the inherent similarity in the electrochemical properties of carbon-based anodes, which show a negligible effect on performance variations compared to other factors in the system. Moreover, in electrochemical degradation using carbon-based anodes, the choice of graphite, carbon felt, or carbon fiber had minimal influence on performance. Regardless of the anode material, optimized experimental conditions were the primary determinant of removal efficiency. These findings suggest that the inherently high removal capability of carbon-based anodes makes their intrinsic properties less critical compared to operational parameters. Intrinsic descriptors such as catalyst composition or adsorbent pore volume were not included in the model due to limited availability in the literature. Only a few studies reported these parameters quantitatively, and incorporating such sparse features would likely introduce bias and reduce model reliability. This limitation highlights the need for more comprehensive and standardized reporting of material properties in future studies to enable more robust, descriptor-rich machine learning models. Our study provides a scientifically robust guideline for researchers utilizing carbon-based anodes in water treatment systems, enabling the optimization of experimental conditions to enhance electrochemical removal efficiency.

Supplementary Material

SI

Acknowledgment

This study was supported by the Superfund Research Program of the National Institute of Environmental Health Sciences (NIH; grant number P42ES017198)

References

  1. Abdalrhman AS, Ganiyu SO, El-Din MG, 2019. Degradation kinetics and structure-reactivity relation of naphthenic acids during anodic oxidation on graphite electrodes. Chem. Eng. J 370, 997–1007. 10.1016/j.cej.2019.03.281 [DOI] [Google Scholar]
  2. Aguirre VM, 2016. Bayesian analysis of definitive screening designs when the response is nonnormal. Appl. Stochastic Models Bus. Ind 32, 440–452. 10.1002/asmb.2160 [DOI] [Google Scholar]
  3. Aji A, Sidik F, Lin J-L, 2025. Molecular-level insights into the degradation of dissolved organic matter from cyanobacteria-impacted water by electro-oxidation and electro-Fenton with carbon-based electrodes. J. Environ. Manage 373, 123539. 10.1016/j.jenvman.2024.123539 [DOI] [PubMed] [Google Scholar]
  4. Al-Ani H, 2023. Artificial neural network in the prediction of surface roughness: A comparative study. Sustain. Eng. Innov 5, 141–150. 10.37868/sei.v5i2.id216 [DOI] [Google Scholar]
  5. Almajid AS, 2021. Multilayer perceptron optimization on imbalanced data using svm-smote and one-hot encoding for credit card default prediction. J. Adv. Inf. Technol 3, 67–74. 10.15294/jaist.v3i2.57061 [DOI] [Google Scholar]
  6. Alnaimat S, Mohsen O, Elnakar H, 2024. Perfluorooctanoic Acids (PFOA) removal using electrochemical oxidation: A machine learning approach. J. Environ. Manage 370, 122857. 10.1016/j.jenvman.2024.122857 [DOI] [PubMed] [Google Scholar]
  7. Alsolai H, Roper M, 2020. A systematic literature review of machine learning techniques for software maintainability prediction. Inf. Software Technol 119, 106214. 10.1016/j.infsof.2019.106214 [DOI] [Google Scholar]
  8. Apicella A, Isgrò F, Pollastro A, Prevete R, 2023. On the effects of data normalization for domain adaptation on EEG data. Eng. Appl. Artif. Intell 123, 106205. 10.1016/j.engappai.2023.106205 [DOI] [Google Scholar]
  9. Cerrillo-Gonzalez M.d.M., Taqieddin A., Sarrouf S., Sakhaee N., Paz-García JM., Alshawabkeh AN., Ehsan MF., 2025. Enhancing H2O2 Generation Using Activated Carbon Electrocatalyst Cathode: Experimental and Computational Insights on Current, Cathode Design, and Reactor Configuration. Catalysts 15, 189. 10.3390/catal15020189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Chaplin BP, 2019. The prospect of electrochemical technologies advancing worldwide water treatment. Acc. Chem. Res 52, 596–604. 10.1021/acs.accounts.8b00611 [DOI] [PubMed] [Google Scholar]
  11. Chen Z, Lai W, Xu Y, Xie G, Hou W, Zhanchang P, Kuang C, Li Y, 2021. Anodic oxidation of ciprofloxacin using different graphite felt anodes: kinetics and degradation pathways. J. Hazard. Mater 405, 124262. 10.1016/j.jhazmat.2020.124262 [DOI] [PubMed] [Google Scholar]
  12. Coria G, Sirés I, Brillas E, Nava JL, 2016. Influence of the anode material on the degradation of naproxen by Fenton-based electrochemical processes. Chem. Eng. J 304, 817–825. 10.1016/j.cej.2017.06.038 [DOI] [Google Scholar]
  13. Dashti A, Navidpour AH, Amirkhani F, Zhou JL, Altaee A, 2024. Application of machine learning models to improve the prediction of pesticide photodegradation in water by ZnO-based photocatalysts. Chemosphere 362, 142792. 10.1016/j.chemosphere.2024.142792 [DOI] [PubMed] [Google Scholar]
  14. Ditta A, Tabish AN, Farhat I, Razzaq L, Fouad Y, Miran S, Mujtaba MA, Kalam MA, 2023. The Optimization of Operational Variables of Electrochemical Water Disinfection Using Response Surface Methodology. Sustainability 15, 4390. 10.3390/su15054390 [DOI] [Google Scholar]
  15. Folgueiras-Amador AA, Teuten AE, Salam-Perez M, Pearce JE, Denuault G, Pletcher D, Parsons PJ, Harrowven DC, Brown RC, 2022. Cathodic Radical Cyclisation of Aryl Halides Using a Strongly-Reducing Catalytic Mediator in Flow. Angew. Chem. Int. Ed 61, e202203694. 10.1002/anie.202203694 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Franceschini G, Macchietto S, 2008. Model-based design of experiments for parameter precision: State of the art. Chem. Eng. Sci 63, 4846–4872. 10.1016/j.ces.2007.11.034 [DOI] [Google Scholar]
  17. Garosi Y, Sheklabadi M, Conoscenti C, Pourghasemi HR, Van Oost K, 2019. Assessing the performance of GIS-based machine learning models with different accuracy measures for determining susceptibility to gully erosion. Sci. Total Environ 664, 1117–1132. 10.1016/j.scitotenv.2019.02.093 [DOI] [PubMed] [Google Scholar]
  18. Ghawi R, Pfeffer J, 2019. Efficient hyperparameter tuning with grid search for text categorization using kNN approach with BM25 similarity. Open Comput. Sci 9, 160–180. 10.1515/comp-2019-0011 [DOI] [Google Scholar]
  19. Goldstein A, Kapelner A, Bleich J, Pitkin E, 2015. Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation. J. Comput. Graphical Stat 24, 44–65. 10.1080/10618600.2014.907095 [DOI] [Google Scholar]
  20. Guo L, Xu X, Niu C, Wang Q, Park J, Zhou L, Lei H, Wang X, Yuan X, 2024. Machine learning-based prediction and experimental validation of heavy metal adsorption capacity of bentonite. Sci. Total Environ 926, 171986. 10.1016/j.scitotenv.2023.166678 [DOI] [PubMed] [Google Scholar]
  21. Gupta S, Aga D, Pruden A, Zhang L, Vikesland P, 2021. Data analytics for environmental science and engineering research. Environ. Sci. Technol 55, 10895–10907. 10.1021/acs.est.1c01026 [DOI] [PubMed] [Google Scholar]
  22. Hashemi A, Basafa M, Behravan A, 2022. Machine learning modeling for solubility prediction of recombinant antibody fragment in four different E. coli strains. Sci. Rep 12, 5463. 10.1038/s41598-022-09500-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Hayasaka R, Hänisch J, Cayado P, 2024. DSDApp: An Open-Access Tool for Definitive Screening Design. J. Open Res. Software 12. 10.5334/jors.462 [DOI] [Google Scholar]
  24. He Y, Lin H, Guo Z, Zhang W, Li H, Huang W, 2019. Recent developments and advances in boron-doped diamond electrodes for electrochemical oxidation of organic pollutants. Sep. Purif. Technol 212, 802–821. 10.1016/j.seppur.2018.11.056 [DOI] [Google Scholar]
  25. Hu M, Stephen B, Browell J, Haben S, Wallom DC, 2023. Impacts of building load dispersion level on its load forecasting accuracy: Data or algorithms? Importance of reliability and interpretability in machine learning. Energy Build. 285, 112896. 10.1016/j.enbuild.2023.112896 [DOI] [Google Scholar]
  26. Ilemobayo JA, Durodola O, Alade O, Awotunde OJ, Olanrewaju AT, Falana O, Ogungbire A, Osinuga A, Ogunbiyi D, Ifeanyi A, 2024. Hyperparameter tuning in machine learning: A comprehensive review. J. Eng. Res. Rep 26, 388–395. 10.9734/jerr/2024/v26i61188 [DOI] [Google Scholar]
  27. Jaffari ZH, Abbas A, Lam S-M, Park S, Chon K, Kim E-S, Cho KH, 2023. Machine learning approaches to predict the photocatalytic performance of bismuth ferrite-based materials in the removal of malachite green. J. Hazard. Mater 442, 130031. 10.1016/j.jhazmat.2022.130031 [DOI] [PubMed] [Google Scholar]
  28. Jasmann JR, Borch T, Sale TC, Blotevogel J, 2016. Advanced electrochemical oxidation of 1, 4-dioxane via dark catalysis by novel titanium dioxide (TiO2) pellets. Environ. Sci. Technol 50, 8817–8826. 10.1021/acs.est.6b02183 [DOI] [PubMed] [Google Scholar]
  29. Jiang S, Xu W, Xia Q, Yi M, Zhou Y, Shang J, Cheng X, 2024. Application of machine learning in the study of cobalt-based oxide catalysts for antibiotic degradation: an innovative reverse synthesis strategy. J. Hazard. Mater 471, 134309. 10.1016/j.jhazmat.2024.134309 [DOI] [PubMed] [Google Scholar]
  30. Jung E, Shin H, Hooch Antink W, Sung Y-E, Hyeon T, 2020. Recent advances in electrochemical oxygen reduction to H2O2: catalyst and cell design. ACS Energy Lett. 5, 1881–1892. 10.1021/acsenergylett.0c00812 [DOI] [Google Scholar]
  31. Kim H-B, Ehsan MF, Alshawabkeh AN, Kim J-G, 2025. Electrochemical activation of alum sludge for the adsorption of lead (Pb (II)) and arsenic (As): Mechanistic insights and machine learning (ML) analysis. Bioresour. Technol, 132563. 10.1016/j.biortech.2025.132563 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kim J-G, Kim H-B, Jeong W-G, Lee K-H, Baek K, 2024a. Electrochemical oxidation and mechanism of sulfanilamide from groundwater in a flow-through system using carbon fiber (CF) anode. Chemosphere 349, 140817. 10.1016/j.chemosphere.2023.140817 [DOI] [PubMed] [Google Scholar]
  33. Kim J-G, Kim H-B, Lee S, Kwon EE, Baek K, 2022. Mechanistic investigation into flow-through electrochemical oxidation of sulfanilamide for groundwater using a graphite anode. Chemosphere 307, 136106. 10.1016/j.chemosphere.2022.136106 [DOI] [PubMed] [Google Scholar]
  34. Kim J-G, Sarrouf S, Ehsan MF, Alshawabkeh AN, Baek K, 2024b. In-situ groundwater remediation of contaminant mixture of As (III), Cr (VI), and sulfanilamide via electrochemical degradation/transformation using pyrite. J. Hazard. Mater 473, 134648. 10.1016/j.jhazmat.2024.134648 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Kim J-G, Sarrouf S, Ehsan MF, Baek K, Alshawabkeh AN, 2023. In-situ hydrogen peroxide formation and persulfate activation over banana peel-derived biochar cathode for electrochemical water treatment in a flow reactor. Chemosphere 331, 138849. 10.1016/j.chemosphere.2023.138849 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Kim S, Kim C, Lee J, Kim S, Lee J, Kim J, Yoon J, 2018. Hybrid electrochemical desalination system combined with an oxidation process. ACS Sustainable Chem. Eng 6, 1620–1626. 10.1021/acssuschemeng.7b02789 [DOI] [Google Scholar]
  37. Körbahti BK, Erdem MC, 2017. Ph change in electrochemical oxidation of imidacloprid pesticide using boron-doped diamond electrodes. Turk. J. Eng 1, 32–36. 10.31127/tuje.316683 [DOI] [Google Scholar]
  38. Liu J, Xu Z, Zhang W, 2023. Unraveling the role of Fe in As (III & V) removal by biochar via machine learning exploration. Sep. Purif. Technol 311, 123245. 10.1016/j.seppur.2023.123245 [DOI] [Google Scholar]
  39. Lowe SE, Shi G, Zhang Y, Qin J, Wang S, Uijtendaal A, Sun J, Jiang L, Jiang S, Qi D, 2019. Scalable production of graphene oxide using a 3D-printed packed-bed electrochemical reactor with a boron-doped diamond electrode. ACS Appl. Nano Mater 2, 867–878. 10.1021/acsanm.8b02126 [DOI] [Google Scholar]
  40. Lundberg SM, Lee S-I, 2017. A unified approach to interpreting model predictions. Adv. Neural Inf. Process 30. [Google Scholar]
  41. Mauter MS, Elimelech M, 2008. Environmental applications of carbon-based nanomaterials. Environ. Sci. Technol 42, 5843–5859. 10.1021/es8006904 [DOI] [PubMed] [Google Scholar]
  42. Narkevica I, Reimanis M, Kleperis J, Ozolins J, Berzina-Cimdina L, 2014. Electrochemical Studies of Nonstoichiometric TiO2-x Ceramic. Key Eng. Mater 604, 254–257. 10.4028/www.scientific.net/KEM.604.254 [DOI] [Google Scholar]
  43. Niu J, Li Y, Shang E, Xu Z, Liu J, 2016. Electrochemical oxidation of perfluorinated compounds in water. Chemosphere 146, 526–538. 10.1016/j.chemosphere.2015.11.115 [DOI] [PubMed] [Google Scholar]
  44. Park J, Lee WH, Kim KT, Park CY, Lee S, Heo T-Y, 2022. Interpretation of ensemble learning to predict water quality using explainable artificial intelligence. Sci. Total Environ 832, 155070. 10.1016/j.scitotenv.2022.155070 [DOI] [PubMed] [Google Scholar]
  45. Periyasamy S, Muthuchamy M, 2018. Electrochemical oxidation of paracetamol in water by graphite anode: effect of pH, electrolyte concentration and current density. J. Environ. Chem. Eng 6, 7358–7367. 10.1016/j.jece.2018.08.036 [DOI] [Google Scholar]
  46. Quinteros ME, Lu S, Blazquez C, Cárdenas-R JP, Ossa X, Delgado-Saborit J-M, Harrison RM, Ruiz-Rudolph P, 2019. Use of data imputation tools to reconstruct incomplete air quality datasets: A case-study in Temuco, Chile. Atmos. Environ 200, 40–49. 10.1016/j.atmosenv.2018.11.053 [DOI] [Google Scholar]
  47. Ram Talib NS, Halmi MIE, Abd Ghani SS, Zaidan UH, Shukor MYA, 2019. Artificial Neural Networks (ANNs) and Response Surface Methodology (RSM) Approach for Modelling the Optimization of Chromium (VI) Reduction by Newly Isolated Acinetobacter radioresistens Strain NS-MIE from Agricultural Soil. Biomed Res. Int 2019, 5785387. 10.1155/2019/5785387 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Ranjan D, Mishra D, Hasan S, 2011. Bioadsorption of arsenic: an artificial neural networks and response surface methodological approach. Ind. Eng. Chem. Res 50, 9852–9863. 10.1021/ie200612f [DOI] [Google Scholar]
  49. Ren L, Wang T, Seklouli AS, Zhang H, Bouras A, 2023. A review on missing values for main challenges and methods. Inf. Syst 119, 102268. 10.1016/j.is.2023.102268 [DOI] [Google Scholar]
  50. Sania A, Pini N, Nelson ME, Myers MM, Shuffrey LC, Lucchini M, Elliott AJ, Odendaal HJ, Fifer WP, 2021. The K nearest neighbor algorithm for imputation of missing longitudinal prenatal alcohol data. 10.21203/rs.3.rs-32456/v3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Schratz P, Muenchow J, Iturritxa E, Richter J, Brenning A, 2019. Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data. Ecol. Modell 406, 109–120. 10.1016/j.ecolmodel.2019.06.002 [DOI] [Google Scholar]
  52. Sirés I, Brillas E, Oturan MA, Rodrigo MA, Panizza M, 2014. Electrochemical advanced oxidation processes: today and tomorrow. A review. Environ. Sci. Pollut. Res 21, 8336–8367. 10.1007/s11356-014-2783-1 [DOI] [PubMed] [Google Scholar]
  53. Soares BS, de Mello R, Motheo AJ, 2024. Groundwater treatment and disinfection by electrochemical advanced oxidation process: Influence of the supporting electrolyte and the nature of the contaminant. Appl. Res 3, e202300008. 10.1002/appl.202300008 [DOI] [Google Scholar]
  54. Song H, Yan L, Jiang J, Ma J, Pang S, Zhai X, Zhang W, Li D, 2018. Enhanced degradation of antibiotic sulfamethoxazole by electrochemical activation of PDS using carbon anodes. Chem. Eng. J 344, 12–20. 10.1016/j.cej.2018.03.050 [DOI] [Google Scholar]
  55. Swain SS, Khura TK, Sahoo PK, Chobhe KA, Al-Ansari N, Kushwaha HL, Kushwaha NL, Panda KC, Lande SD, Singh C, 2024. Proportional impact prediction model of coating material on nitrate leaching of slow-release Urea Super Granules (USG) using machine learning and RSM technique. Sci. Rep 14, 3053. 10.1038/s41598-024-53410-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Tai M, Ly A, Leung I, Nayar G, 2015. Efficient high-throughput biological process characterization: Definitive screening design with the Ambr250 bioreactor system. Biotechnol. Progr 31, 1388–1395. 10.1002/btpr.2142 [DOI] [PubMed] [Google Scholar]
  57. Taoufik N, Boumya W, Achak M, Chennouk H, Dewil R, Barka N, 2022. The state of art on the prediction of efficiency and modeling of the processes of pollutants removal based on machine learning. Sci. Total Environ 807, 150554. 10.1016/j.scitotenv.2021.150554 [DOI] [PubMed] [Google Scholar]
  58. Tarpeh WA, Barazesh JM, Cath TY, Nelson KL, 2018. Electrochemical stripping to recover nitrogen from source-separated urine. Environ. Sci. Technol 52, 1453–1460. 10.1021/acs.est.7b05488 [DOI] [PubMed] [Google Scholar]
  59. Thomas T, Rajabi E, 2021. A systematic review of machine learning-based missing value imputation techniques. Data Technol. Appl 55, 558–585. 10.1108/DTA-12-2020-0298 [DOI] [Google Scholar]
  60. Valica M, Hostin S, 2016. Electrochemical treatment of water contaminated with methylorange. 10.1515/nbec-2016-0006 [DOI] [Google Scholar]
  61. Wang D, Thunéll S, Lindberg U, Jiang L, Trygg J, Tysklind M, 2022. Towards better process management in wastewater treatment plants: Process analytics based on SHAP values for tree-based machine learning methods. J. Environ. Manage 301, 113941. 10.1016/j.jenvman.2021.113941 [DOI] [PubMed] [Google Scholar]
  62. Wang H, Zhai P, Long X, Ma J, Li Y, Liu B, Xu Z, 2023. Research progress on using biological cathodes in microbial fuel cells for the treatment of wastewater containing heavy metals. Front. Microbiol 14, 1270431. 10.3389/fmicb.2023.1270431 [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Wang W, Li D, Zhou S, Wang Y, Yu L, 2024. Exploring the key influencing factors of low-carbon innovation from urban characteristics in China using interpretable machine learning. Environ. Impact Assess. Rev 107, 107573. 10.1016/j.eiar.2024.107573 [DOI] [Google Scholar]
  64. Widodo E, Guritno S, Haryatmi S, 2015. Response Surface Models with Data Outliers through a Case Study. Appl. Math. Sci 9, 1803–1812. 10.12988/ams.2015.5179 [DOI] [Google Scholar]
  65. Xu W, Pignatello JJ, Mitch WA, 2013. Role of black carbon electrical conductivity in mediating hexahydro-1, 3, 5-trinitro-1, 3, 5-triazine (RDX) transformation on carbon surfaces by sulfides. Environ. Sci. Technol 47, 7129–7136. 10.1021/es4012367 [DOI] [PubMed] [Google Scholar]
  66. Yang C, Guan X, Xu Q, Xing W, Chen X, Chen J, Jia P, 2024. How can SHAP (SHapley Additive exPlanations) interpretations improve deep learning based urban cellular automata model? Comput. Environ. Urban Syst 111, 102133. 10.1016/j.compenvurbsys.2024.102133 [DOI] [Google Scholar]
  67. Yang YY, Ren XJ, Li SC, Zhu LM, Tian SY, 2013. Quantitative investigation of Flocs flotation removal and Hydrocyclonic separation for the Washing water treatment. Applied Mechanics and Materials 333, 1857–1861. 10.4028/www.scientific.net/AMM.333-335.1857 [DOI] [Google Scholar]
  68. Ye C, Tran TTT, Yang Y, 2025. Machine learning and genetic algorithm for effluent quality optimization in wastewater treatment. J. Water Process Eng 71, 107294. 10.1016/j.jwpe.2025.107294 [DOI] [Google Scholar]
  69. Youssefi S, Emam-Djomeh Z, Mousavi S, 2009. Comparison of artificial neural network (ANN) and response surface methodology (RSM) in the prediction of quality parameters of spray-dried pomegranate juice. Drying Technol. 27, 910–917. 10.1080/07373930902988247 [DOI] [Google Scholar]
  70. Zhang W, Ashraf WM, Senadheera SS, Alessi DS, Tack FM, Ok YS, 2023. Machine learning based prediction and experimental validation of arsenite and arsenate sorption on biochars. Sci. Total Environ 904, 166678. 10.1016/j.scitotenv.2023.166678 [DOI] [PubMed] [Google Scholar]
  71. Zhang Y, Wu Y, 2021. Introducing machine learning models to response surface methodologies, Response Surface Methodology in Engineering Science. IntechOpen. 10.5772/intechopen.98191 [DOI] [Google Scholar]
  72. Zhou L, Pan S, Wang J, Vasilakos AV, 2017. Machine learning on big data: Opportunities and challenges. Neurocomputing 237, 350–361. 10.1016/j.neucom.2017.01.026 [DOI] [Google Scholar]
  73. Zhu X, Tsang DC, Wang L, Su Z, Hou D, Li L, Shang J, 2020. Machine learning exploration of the critical factors for CO2 adsorption capacity on porous carbon materials at different pressures. J. Cleaner Prod 273, 122915. 10.1016/j.jclepro.2020.122915 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

SI

RESOURCES