Abstract
Accurate water quality forecasting—specifically tracking phytoplankton dynamics like chlorophyll a—is crucial for safeguarding aquatic ecosystems and global water security. Deep temporal architectures offer powerful non-linear modeling capabilities, promising to overcome the parameterization bottlenecks of traditional process-based models. However, current applications treat these models as black boxes, leaving it obscure how specific architectural modules interact with the intrinsic ecological time scales governing aquatic environments. Here, we reveal a fundamental architectural–ecological duality in water quality forecasting by systematically dissecting nine prediction architectures—spanning classic baselines to modern deep models—across two ecologically contrasting reservoirs. Using chlorophyll a dynamics to represent integrated system responses, we show that patch embedding serves as a universal temporal operator, mitigating noise and enhancing forecasting accuracy by up to 10.1% when integrated into standard baselines. Crucially, inter-variable modeling strategies dictate the effective forecast horizon: channel-independent architectures dominate short-term (1–7 days) chlorophyll a forecasting (Nash–Sutcliffe efficiency up to 0.88) by capturing biomass self-persistence, whereas cross-variable attention architectures excel in medium-term (8–15 days) prediction (efficiency up to 0.83) by uncoupling delayed nutrient–temperature interactions. Explainable AI further confirms a temporal shift in feature reliance from current biomass to lagged drivers over extended horizons. Beyond water quality, this mechanistic alignment between deep learning modules and ecological time scales establishes a scalable blueprint for building interpretable, horizon-adaptive AI frameworks across complex climate and environmental systems.
Keywords: Architectural–ecological duality, Water quality forecasting, Chlorophyll a, Deep forecasting architectures, Ecological time scales, Horizon-adaptive AI
Graphical abstract

Highlights
-
•
Dissecting nine architectures uncovers architectural–ecological duality in water quality forecasting.
-
•
Patch embedding acts as a universal operator, boosting forecasting accuracy across baselines by up to 10.1%.
-
•
Channel independence governs short-term persistence, while cross-variable attention resolves lagged drivers.
-
•
Feature reliance shifts from current chlorophyll a biomass to nutrient-temperature dynamics as lead time expands.
-
•
Establishes a scalable blueprint for building horizon-adaptive AI across complex environmental systems.
1. Introduction
Clean water is a fundamental need for both human survival and the health of aquatic ecosystems [1]. Accurate forecasting of water quality dynamics is essential for obtaining reliable information on future water quality conditions, formulating effective countermeasures, and ultimately safeguarding water security [[2], [3], [4]]. Water quality forecasting models are broadly categorized into two main types: process-based mechanistic models and data-driven approaches. Mechanistic models operate by simulating the complex biogeochemical interactions among various water quality parameters [5]. They offer a high degree of physical representation, but their practical application is often time-intensive, and extensive hydrological, meteorological, and biogeochemical data are required for accurate parameterization and calibration. Their inherent complexity, significant computational intensity, and substantial resource requirements frequently impede their implementation in real-world scenarios. Data-driven models derive patterns and relationships directly from historical monitoring data, thereby eliminating the need for explicit mechanistic assumptions and offering enhanced operational flexibility [6,7]. However, traditional data-driven methods, such as multivariable linear regression and autoregressive integrated moving average, often struggle to capture the complex nonlinearity, nonstationarity, and intricate medium-term dependencies intrinsic to environmental time series. These limitations are particularly evident when handling high-frequency, multivariate water quality datasets and frequently lead to less reliable forecasts [8]. Deep learning models offer compelling solutions to these challenges due to their capacity for hierarchical feature extraction. They can learn complex temporal dependencies directly from raw data, minimizing the need for laborious manual feature engineering [9]. This ability to model complex nonlinear systems has propelled architectures such as multilayer perceptrons (MLPs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and various Transformer variants to the forefront of water quality forecasting [[10], [11], [12], [13]].
Despite these advancements, a gap exists between the rapid development of advanced deep learning architectures for time series forecasting (i.e., deep forecasting models) and their comprehensive application to water quality prediction. The broad time series forecasting domain has demonstrated strong performance of deep forecasting models in fields such as energy consumption, traffic flow, and meteorology [[14], [15], [16], [17], [18], [19], [20]]. These architectures incorporate innovative modules, including patch embedding, decomposition structures, channel-independent learning, and cross-variable attention, to capture local temporal patterns and long-range dependencies. Such capabilities are also highly relevant for environmental time series, as forecast performance is often limited by rapidly shifting environmental conditions and complex, lagged interactions among variables.
Water quality forecasting studies have largely focused on comparing model-level performance, whereas the functional roles of the core architectural modules remain less understood. It is therefore unclear which deep forecasting structures are best suited for different ecological forecasting scenarios, why certain modules improve predictive skill, and whether these advantages are robust across reservoirs with contrasting trophic characteristics and bloom dynamics. This uncertainty limits informed model selection and the development of interpretable, horizon-adaptive forecasting frameworks for aquatic ecosystems.
Here, we compared two conventional baseline models and seven advanced deep learning architectures for multivariate water quality time series. Beyond model-level comparisons, we used ablation analysis to examine how core structural modules embedded within deep forecasting architectures contribute to predictive skill. We selected two reservoirs with different ecological characteristics and data conditions to assess model performance and generalization capability: Königshütte Reservoir in Germany and Yazidang Reservoir in China. Chlorophyll a (Chl a), a key indicator of eutrophication with complex dynamics [21,22], was selected as the target variable. Using high-frequency monitoring data from the aforementioned reservoirs, we evaluated model performance over short-term (1–7 days) and medium-term (8–15 days) forecast horizons. We aimed to benchmark advanced deep forecasting architectures against established approaches under realistic ecological forecasting scenarios and to elucidate how core modeling modules, including patch embedding, decomposition structures, and cross-variable attention, influence forecasting skill across ecological contexts. This systematic module-level analysis provides information for model selection and for developing more interpretable, horizon-adaptive water quality forecasting frameworks.
2. Materials and methods
2.1. Study sites and datasets
Königshütte Reservoir (KS; 51.74° N, 10.80° E) is one of the six reservoirs within the Rappbode reservoir system in the Harz Mountains, Germany [23]. Monitoring data for KS were obtained from the Rappbode Reservoir Observatory [24] of the Terrestrial Environmental Observatories Network (TERENO) project (http://www.tereno.net) and covered December 2011 to March 2016. The dataset included water temperature, nitrate nitrogen (NO3-N), and Chl a. Yazidang Reservoir (YZD; 106.53° E, 38.15° N) is located in Ningxia Province, China. Monitoring data for YZD were obtained from the China National Environmental Monitoring Center (https://www.cnemc.cn/) and covered April 2021 to November 2025. This dataset included 10 water quality indicators: water temperature, pH, dissolved oxygen (DO), electrical conductivity (EC), turbidity, permanganate index (CODMn), ammonia nitrogen (NH4-N), total phosphorus (TP), total nitrogen (TN), and Chl a. The initial sampling intervals were 15 s for KS and 4 h for YZD. Missing values accounted for 5.8% of the KS dataset and 2.6% of the YZD dataset. We imputed gaps shorter than five consecutive time steps by linear interpolation and then extracted the daily median values to facilitate multi-day-ahead forecasting and to harmonize the different sampling frequencies between reservoirs. To assess the sensitivity of our results to the aggregation method, we conducted a supplementary analysis using the 90th percentile instead of the median.
The contrasting trophic conditions of the two reservoirs allowed us to assess model performance across distinct ecological contexts (Supplementary Table S1–S2). KS was predominantly mesotrophic, with a mean Chl a concentration of 6.60 μg L−1; it was classified as eutrophic on 25.6% of the days considered and exceeded the hypereutrophic threshold, defined as Chl a > 25 μg L−1, on 1.7% of days (Fig. 1a). YZD had a higher mean Chl a concentration (8.98 μg L−1). The Chl a observations were classified as eutrophic for 26.0% of the days considered and as hypereutrophic for 7.7% of days (Fig. 1b). Maximum Chl a concentration was substantially higher in YZD than in KS (83.5 versus 47.2 μg L−1). Seasonality was stronger in KS, where the growing-season to non-growing-season Chl a ratio was 2.69, and the September mean peaked at 16.60 μg L−1.YZD showed weaker seasonal contrast, with a ratio of 1.50, but higher year-round variability, with a Chl a coefficient of variation of 1.02 compared with 0.86 in KS, reflecting more episodic bloom dynamics. The two reservoirs also differed in predictor availability: KS represented a predictor-limited setting, whereas YZD provided a predictor-rich dataset with broader nutrient coverage.
Fig. 1.

Daily and monthly chlorophyll a concentrations in the two reservoirs. Time series of chlorophyll a (Chl a) concentrations in the Königshütte (a) and Yazidang (b) reservoirs. Red lines show daily observations, and blue bars show monthly means.
Chl a is widely used as an indicator of phytoplankton biomass and reflects an ecosystem's response to nutrient enrichment, particularly excess nitrogen and phosphorus inputs, the key drivers of eutrophication. The relationship between nutrients and phytoplankton is often nonlinear due to complex biological and physical interactions. Therefore, Chl a was selected as the target variable in this study to represent the system's integrated ecological response and to evaluate the model's capability to capture nonlinear water quality dynamics. All water quality indicators (including the current Chl a concentration) were included as input variables (predictors for future Chl a concentration) in the models, based on their established roles in influencing phytoplankton growth and Chl a dynamics.
2.2. Machine learning models for Chl a concentration forecasting
We forecast Chl a concentration over 1–15-day horizons. Forecasts at 1–7 days were defined as short-term forecasts, which are relevant to early-bloom warning and operational decision-making [25,26]. During this period, Chl a dynamics are associated with rapid phytoplankton growth and short-term physicochemical variability, such as changes in water temperature and DO. Forecasts at 8–15 days were defined as medium-term forecasts, during which Chl a dynamics are influenced by delayed ecological responses to antecedent nutrient loading and broader water quality changes [27].
We evaluated nine machine learning models for Chl a forecasting. Long short-term memory (LSTM) [28] and random forests (RF) [29] are two of the most widely used machine learning models for water quality forecasting [[30], [31], [32]]. However, both models exhibit inherent limitations when applied to Chl a forecasting. LSTM may struggle to capture long-range dependencies and can suffer from information loss over extended sequences, whereas RF lacks an explicit mechanism for modeling temporal dependencies. We also evaluated seven recent deep-learning architectures for time series forecasting [33,34]: Crossformer [19], decomposition linear (DLinear) [18], Informer [20], nonstationary Transformer (NSTransformer) [15], patch time series Transformer (PatchTST) [16], segment recurrent neural network (SegRNN) [14], and TimesNet [17]. Together, these models covered tree-based machine learning, MLPs, RNNs, CNNs, and various Transformer-based models (Fig. 2). A detailed comparison of the temporal representation-learning modules, multivariate interaction-modeling strategies, and backbones of these models is provided in Supplementary Fig. S1. Brief introductions to the nine models in this study are provided in Supplementary Texts S1–S9.
Fig. 2.

Hierarchical classification of the nine forecasting models. The models are organized according to model paradigm, architecture, core computational module, and individual model. Ribbons indicate the corresponding classification relationships; ribbon widths are schematic and do not represent quantitative values. RNN, recurrent neural network; MLP, multilayer perceptron; CNN, convolutional neural network; PatchTST, patch time series Transformer; SegRNN, segment recurrent neural network; DLinear, decomposition linear; NSTransformer, non-stationary Transformer; LSTM, long short-term memory; RF, random forests.
Briefly, Crossformer is a deep learning model based on the Transformer architecture that incorporates patch embedding and cross-variable attention modules [35]. PatchTST also consists of a Transformer architecture with patch embedding and channel independence, whereas NSTransformer combines the Transformer backbone with a stationarization module. DLinear comprises an MLP architecture with trend decomposition and channel independence. SegRNN is an RNN-based model with patch embedding and channel independence, whereas TimesNet relies on a CNN architecture with frequency-domain analysis. Informer is a Transformer-based model; LSTM is an RNN-based model; and RF is a tree-based machine learning model.
2.3. Model construction and explanation
The model construction and explanation workflow comprised six steps: data partitioning, sliding window sampling, model training and validation, hyperparameter tuning, model evaluation, and explanation (Fig. 3). We split the dataset into training (70%), validation (10%), and test (20%) sets. All variables were standardized using z-scores, with the mean and standard deviation estimated from the training set only and then applied to the validation and test sets to avoid data leakage [36]:
| (1) |
where represents the observation order, represents the raw data, is the mean of the training set, is the standard deviation of the training set, and represents the standardized data. We generated time series samples with corresponding features and labels using the sliding-window scheme [31,34]. Samples with missing values (whether in features or labels) were excluded from the dataset to maintain data integrity. We then trained the models on the prepared training sets.
Fig. 3.

Workflow for data preparation, model training, optimization, evaluation and interpretation. The workflow comprises six steps: (1) partitioning the data into training, validation and test sets at a ratio of 70:10:20; (2) generating samples using a sliding window for forecast horizons of 1–15 days; (3) training and validating two baseline models and seven deep-learning forecasting models; (4) optimizing hyperparameters using the tree-structured Parzen estimator (TPE); (5) evaluating model performance over short-term (1–7 days), medium-term (8–15 days), and overall (1–15 days) forecast horizons; and (6) interpreting predictor importance using DeepSHAP. RF, random forests; LSTM, long short-term memory; DLinear, decomposition linear; NSTransformer, non-stationary Transformer; PatchTST, patch time series Transformer; SegRNN, segment recurrent neural network; , Nash–Sutcliffe efficiency; , root mean square error; , mean absolute percentage error; , Kling–Gupta efficiency; DeepSHAP, Deep Shapley additive explanations; Val, validation; Eval, evaluation; N/A, not available.
Hyperparameters were optimized using the tree-structured Parzen estimator (TPE) [37] on the validation set. TPE is a Bayesian optimization method that models the distributions of high-performing hyperparameter configurations as the density function and of low-performing ones as the density , where represents a hyperparameter configuration, and iteratively selects configurations that maximize [38,39]. The optimized hyperparameters for each model are detailed in Table 1.
Table 1.
TPE-optimized hyperparameters for forecasting models at the Königshütte and Yazidang reservoirs.
| Station | Model | Input sizea | Number of layersb | Hidden sizec | Learning rated |
|---|---|---|---|---|---|
| Königshütte reservoirs |
RFe | 7 | - | - | - |
| LSTM | 15 | 2 | 32 | 10−4 | |
| Crossformer | 15 | (2, 2) | 64 | 10−4 | |
| DLinear | 15 | 3 | 64 | 10−3 | |
| Informer | 7 | (3, 1) | 128 | 10−2 | |
| NSTransformer | 7 | (2, 2) | 64 | 10−2 | |
| PatchTST | 15 | (2, 1) | 32 | 10−3 | |
| SegRNN | 15 | 2 | 64 | 10−3 | |
| TimesNet |
30 |
2 |
32 |
10−4 |
|
| Yazidang reservoirs | RFe | 15 | - | - | - |
| LSTM | 15 | 3 | 32 | 10−2 | |
| Crossformer | 15 | (3, 2) | 64 | 10−4 | |
| DLinear | 30 | 2 | 128 | 10−4 | |
| Informer | 15 | (2, 3) | 64 | 10−3 | |
| NSTransformer | 7 | (1, 2) | 128 | 10−2 | |
| PatchTST | 15 | (2, 2) | 64 | 10−3 | |
| SegRNN | 15 | 2 | 32 | 10−3 | |
| TimesNet | 15 | 2 | 64 | 10−3 |
For a pair of (x, y) entry, x and y denote the number of encoder and decoder layers, respectively.
Abbreviations: TPE, tree-structured Parzen estimator; RF, random forests; LSTM, long short-term memory; Dlinear, decomposition linear; NSTransformer, non-stationary Transformer; PatchTST, patch time series Transformer; SegRNN, segment recurrent neural network.
Candidate input sizes were 7, 15, and 30.
Candidate numbers of layers were 1, 2, and 3. For paired values, the first and second numbers denote the numbers of encoder and decoder layers, respectively.
Candidate hidden dimensions were 32, 64, and 128.
Candidate learning rates were 10−2,10−3,and 10−4.
RF hyperparameters were optimized separately, including n_estimators (100, 300, or 500), min_samples_leaf (1, 8, or 16), and max_depth (5, 10, or unrestricted). The optimal combination was n_estimators = 100, min_samples_leaf = 8, and max_depth = unrestricted for both the Königshütte and Yazidang reservoirs.
We used the test set and the Nash-Sutcliffe efficiency () metric to evaluate model performance. measures the goodness-of-fit between the observed and predicted time series, with values closer to 1 indicating a stronger agreement [40]:
| (2) |
where denotes the original observed values, denotes the predicted values, denotes the average of the observed values, and is the number of test set samples.
To provide complementary perspectives on model accuracy, we computed root mean square error (), mean absolute percentage error (), and Kling–Gupta efficiency (). These metrics capture absolute error magnitude, relative error, and combined agreement in correlation, bias, and variability, respectively. Their formulas are provided below.
| (3) |
| (4) |
| (5) |
In equation (5), is the Pearson correlation coefficient between the observed and predicted values, is the ratio of predicted-to-observed standard deviations, and is the ratio of predicted-to-observed means.
was used as the primary metric for model comparison across forecast horizons. , , and were used as complementary metrics to assess absolute, relative, and multiple-dimensional error characteristics, thereby verifying the robustness of the comparative conclusions across multiple evaluation perspectives.
Model interpretation was performed using Shapley additive explanations (SHAP), which estimate the contribution of each input variable to model predictions. For deep-learning models, to address the computational impracticality of exact Shapley values for deep learning models, we used DeepSHAP [41], an approximation of SHAP values based on DeepLIFT multipliers, to quantify variable contributions relative to a reference baseline [42,43]. The Shapley value for the -th feature () is calculated as:
| (6) |
where represents the full feature set, is a subset of features excluding feature , and is the model's output when only features in are considered. We used the mean of the training set as the reference baseline and averaged the resulting SHAP values across all samples and forecast horizons to obtain a global feature importance ranking for each variable.
2.4. Ablation study
Current deep-learning forecasting models increasingly use modular architectures [44], in which different structural modules can be combined to improve temporal representation learning and multivariate interaction modelling [[45], [46], [47]] (Fig. 2). To investigate the contributions and importance of the core architectural modules, we conducted ablation studies [44] on the best-performing models. The objective was to isolate the impact of individual modules and assess the extent of improvements across model architectures, providing insights into which architectural components support predictive performance [45,46].
We used the best-performing model as the baseline and constructed simplified variants by removing or disabling a distinct core module from the baseline model at a time [47]. Hyperparameters, training procedure, data splits (using both YZD and KS reservoirs), and evaluation metrics remained identical to ensure that any performance differences were attributed solely to the absence of the ablated module. For PatchTST, we ablated patch embedding and channel independence, generating PatchTST-nP (no patch embedding), PatchTST-nCI (no channel independence), and PatchTST-nP-nCI (neither patch embedding nor channel independence). For Crossformer, we ablated patch embedding and cross-variable attention in the two-stage attention layer, yielding the Crossformer-nP (no patch embedding), Crossformer-nCV (no cross-variable attention), and Crossformer-nP-nCV (neither patch embedding nor cross-variable attention). Each ablated model's performance was compared with the baseline model to quantify the degradation caused by each module's removal. This comparison allowed for the assessment of the relative importance of each architectural element [8,48].
To further confirm the effectiveness and generalizability of these core modules, we integrated patch embedding, channel independence, and cross-variable attention into two baseline architectures: MLP and LSTM. The resulting variants were Patch-MLP, CI-MLP, CV-MLP, Patch-LSTM, CI-LSTM, and CV-LSTM. The performance of these augmented models was compared with their standard counterparts to determine whether the modules’ benefits extend to other model architectures.
All analyses were performed in Python 3.8. We utilized scikit-learn 1.3.2 to develop RF and PyTorch 2.0.1 to develop the remaining models, and used Optuna 3.4.0 for hyperparameter optimization. Computations were run on Windows 11 with an NVIDIA RTX 4080 GPU and an Intel Core i7-13700K CPU.
3. Results
3.1. Performance of different models across varied forecast horizons
Forecasting accuracy generally decreased for all models as the number of forecasting steps increased (Fig. 4b–f). Over the entire forecasting period, the average across all models for the KS reservoir was 0.58; the models exhibited the best performance for short-term forecasting ( = 0.67) and poor performance for medium-term forecasting (0.47). Models performed better for YZD, with mean values of 0.82 over the full period, 0.85 for short-term forecasts, and 0.78 for medium-term forecasts.
Fig. 4.

NSE-based comparison of nine forecasting models across forecast horizons. a, Locations of the Königshütte (KS) and Yazidang (YZD) reservoirs. b,f, Nash–Sutcliffe efficiency () at forecast horizons of 1–15 days for KS (b) and YZD (f). Symbols and colors denote the nine models, and lines connect values from successive forecast horizons. The vertical dashed line separates short-term (1–7 days) and medium-term (8–15 days) forecasts. c–e, Mean for KS over short-term horizons (c), medium-term horizons (d), and overall (1–15 days) horizons (e). g–i, Corresponding mean values for YZD. Red dashed lines indicate the mean across the nine models in each panel. DLinear, decomposition linear; LSTM, long short-term memory; NSTransformer, non-stationary Transformer; PatchTST, patch time series Transformer; RF, random forests; SegRNN, segment recurrent neural network.
PatchTST and Crossformer showed the strongest performance across forecast horizons (Fig. 4b–f). PatchTST achieved the highest short-term forecasting values (0.75 for KS and 0.88 for YZD), whereas Crossformer performed best over medium-term forecasting, with values of 0.60 for KS and 0.83 for YZD. The differences in forecasting performance across the models increased as the forecast horizon extended. The two baseline models, RF and LSTM, delivered average forecasting performance among all the models and were comparable to DLinear, NSTransformer, and TimesNet for short-term forecasting. SegRNN performed slightly better than the baseline models, whereas Informer and TimesNet (used for medium-term forecasting) underperformed the baseline models. The supplementary analysis using the 90th percentile for daily aggregation yielded consistent results: PatchTST and Crossformer remained the best-performing models for short-term and medium-term forecasting, respectively, and the relative ranking of the other models also remained consistent with the median-based results (Supplementary Fig. S4). Notably, the aforementioned findings regarding model performance and comparisons based on were consistently supported by the , , and evaluations (Fig. 5), indicating that the observed performance patterns are robust across evaluation metrics with different sensitivities to prediction errors.
Fig. 5.

-, -, and -based comparison of nine forecasting models across forecast horizons. a–c, Root-mean-square error (; a), mean absolute percentage error (; b), and Kling–Gupta efficiency (; c) for the Königshütte reservoir. d–f, Corresponding (d), (e), and (f) values for the Yazidang reservoir. Points show metric values at forecast horizons of 1–15 days; colors and symbols denote the nine forecasting models. Vertical dashed lines separate short-term (1–7 days) and medium-term (8–15 days) forecasts. DLinear, decomposition linear; LSTM, long short-term memory; NSTransformer, non-stationary Transformer; PatchTST, patch time series Transformer; RF, random forests; SegRNN, segment recurrent neural network.
3.2. Importance of the predictors for Chl a concentration forecasting
We investigated the importance of the predictors for Chl a concentration forecasting by leveraging PatchTST and Crossformer, the two best-performing models with both reservoirs. Predictor-importance rankings showed broadly similar patterns across the two reservoirs and models (Fig. 6). Current Chl a concentration was the dominant predictor, although its relative importance decreased as the forecast steps increased. Both water temperature and nutrient variables (NO3-N for KS and NH4-N for YZD) contributed to predicting future Chl a concentration with increasing relative importance at longer lead times. For YZD, the collective contribution of predictors excluding Chl a, water temperature, and NH4-N showed an increasing trend.
Fig. 6.

Forecast-horizon dependence of predictor importance in PatchTST and Crossformer. a,b, Relative predictor importance derived from PatchTST (a) and Crossformer (b) for the Königshütte reservoir. c,d, Corresponding results for the Yazidang reservoir. Stacked bars show the normalized contribution of each predictor at forecast steps of 1–15 days and sum to 100% within each forecast step. Relative importance was calculated by dividing the mean absolute Shapley value of each predictor by the sum across all predictors. “Others” includes pH, dissolved oxygen, electrical conductivity, turbidity, permanganate index, total phosphorus, and total nitrogen. PatchTST, patch time series Transformer; Chl a, chlorophyll a; NO3-N, nitrate nitrogen; NH4-N, ammonia nitrogen.
The exact values of relative predictor importance derived from the two models differed. PatchTST placed greater emphasis on the current Chl a concentration, as evidenced by its high relative importance values when compared to the Crossformer model. In contrast, Crossformer placed greater weight on the other predictors. PatchTST thus maintained a relatively high level of importance for the current Chl a concentration throughout all horizons, whereas Crossformer distributed importance more evenly across all input predictors in longer forecasting steps, particularly when the lead time exceeded 10 days.
3.3. Contributions of core modules to model performance
Ablating the core modules of PatchTST and Crossformer impaired the models' predictive accuracies (Fig. 7). For PatchTST, removing the patch embedding resulted in an 11.8% reduction in accuracy for KS and a 6.9% reduction for YZD. This systematic performance drop confirms the critical role of patch embedding in capturing local temporal patterns and enhancing the robustness of feature representations. Channel-independent processing ablation resulted in smaller reductions of 4.0% with KS data and 2.4% with YZD data, indicating its secondary yet complementary contribution. The combined removal of both modules caused the largest performance drop, with declining by 12.6% and 7.4% for KS and YZD, respectively, which highlights the dominant role of patch embedding. With Crossformer, ablating the patch embedding module reduced by 10.2% and 6.8% for KS and YZD, respectively, while removing cross-variable attention reduced by 8.1% and 5.3%, respectively. The simultaneous removal of both modules resulted in the most substantial performance degradation, with reductions of 16.9% for KS and 11.0% for YZD. These results confirm that both modules are critical to Crossformer's performance, with patch embedding playing a slightly more influential role.
Fig. 7.

Effects of removing key architectural modules from PatchTST and Crossformer. a,b, Ablation results for PatchTST (a) and Crossformer (b) at the Königshütte reservoir. c,d, Corresponding results at the Yazidang reservoir. Colored points show Nash–Sutcliffe efficiency () at individual forecast steps (1–15 days) for all model variants (complete and ablated), with colors indicating the forecast step. Bars denote the corresponding mean averaged over all steps for each variant. The red dashed line indicates the full-model overall mean, which serves as the benchmark against which ablated versions are assessed. PatchTST-nP lacks patch embedding; PatchTST-nCI lacks channel independence; PatchTST-nP-nCI lacks both modules. Crossformer-nP lacks patch embedding; Crossformer-nCV lacks cross-variable attention; and Crossformer-nP-nCV lacks both modules. PatchTST, patch time series Transformer.
The enhancement experiments further confirmed these findings (Fig. 8). Patch embedding yielded the most pronounced and consistent improvements. For KS, Patch-MLP increased the average by 9.0% to 0.67, outperforming both the baseline MLP (0.58) and the original PatchTST (0.66). Patch-LSTM achieved a 10.1% gain over its baseline. For YZD, Patch-MLP and Patch-LSTM improved by 5.8% and 1.9%, respectively, demonstrating the robustness of the patch mechanism across different architectures and datasets. Channel independence and cross-variable attention yielded smaller but positive effects. For KS, MLP and LSTM performances improved by 4.8% and 5.6%, respectively, with the channel independence module, and by 3.2% and 3.4%, respectively, with the cross-variable attention module. At YZD, channel independence enhanced MLP and LSTM by 3.5% and 2.3%, respectively, and cross-variable attention by 2.1% and 1.7%, respectively. These findings suggest that channel independence and cross-variable attention offer complementary benefits by refining variable independence and intervariable relationships, whereas patch embedding is the dominant factor that drives predictive improvement.
Fig. 8.

Effects of integrating advanced modules into MLP and LSTM baseline models. Patch embedding (Patch), channel independence (CI), and cross-variable attention (CV) were individually incorporated into multilayer perceptron (MLP) and long short-term memory (LSTM) models. a,b, Nash–Sutcliffe efficiency () of the baseline and modified models at the Königshütte reservoir (a) and Yazidang reservoir (b). Colored points represent the at individual forecast steps (1–15 days) for all model variants (baseline MLP and LSTM and their modified counterparts), with colors indicating the forecast step. Bars denote the corresponding mean averaged over all steps for each variant. The red dashed line indicates the baseline model overall mean, which serves as the benchmark against which the modified versions are assessed.
4. Discussion
4.1. Performance differences among forecasting models
Although the two reservoirs differed in their trophic characteristics, predictor availability, and ecological temporal dynamics, the overall model ranking patterns remained highly consistent across the forecast horizons, supporting the robustness of the comparative evaluations. Across the comparisons, the deep forecasting architectures showed advantages over conventional models such as LSTM and RF, indicating their greater suitability for complex water quality forecasting tasks. Notably, the performance advantages of the advanced models were greater for the predictor-limited KS reservoir than for the predictor-rich YZD reservoir, indicating that modern architectures are particularly effective in extracting predictive signals from sparse variable sets. Although LSTM and RF are widely used in water quality simulation [49,50], our results suggest that model selection should depend on data structure, forecast horizon and ecological context, rather than on model popularity alone.
TimesNet, a TCN-based model, demonstrated strong short-term forecasting capability, particularly in the first forecasting step, indicating its effectiveness in capturing localized temporal patterns [17]. However, its performance deteriorated in the medium-term forecasting scenarios, indicating a potential limitation in its ability to capture global dependencies over extended horizons. DLinear, a relatively simple MLP-based model, outperformed many more complex deep learning models. This finding emphasizes that fundamental MLP architectures, when coupled with appropriate architectural modules, can achieve competitive performance.
PatchTST and Crossformer, both Transformer-based models, showed advantages in short-term and medium-term forecasting, respectively, which suggests that these models are well-suited to capturing the complex temporal dependencies inherent in water quality time series [16,19]. In contrast, Informer showed weaker performance in this study, despite its reported use in other long-sequence forecasting applications. These results indicate that Transformer-based architecture alone does not ensure improved forecasting performance; model design should match the temporal structure, predictor availability and forecast horizon of the target water quality dataset.
4.2. Effectiveness of core modules for forecasting
Ablation experiments showed that patch embedding is the most influential module in both PatchTST and Crossformer, as it substantially enhanced temporal representation learning. The channel independence module in PatchTST and the cross-variable attention module in Crossformer provided auxiliary benefits, further refining feature interactions among variables. These results indicate that core architectural modules can influence forecasting skill, but their effects depend on the temporal structure of the data, predictor availability, and forecast horizon.
Patch embedding represents a fundamental innovation for water quality forecasting. The consistently strong performance of the patch-based models (PatchTST, Crossformer, and SegRNN), together with the results of the ablation and combination experiments, confirms the effectiveness of patches in improving time series prediction. Patch embedding operates by dividing raw time series into several patches (Supplementary Fig. S5). Its utility stems from the following key aspects.
First, patch embedding significantly enhances regional feature extraction. Dividing the time series into discrete patches allows each block to effectively capture localized information and the associated model to focus on the characteristics within different temporal segments, leading to a more nuanced understanding of local patterns [51]. In this study, the Chl a forecasting task used a 15-day input window and a patch length of 3 days, allowing the model to extract the short-term trends and micro-fluctuations crucial for accurate forecasting.
Second, patch embedding may improve model robustness to noise. Localized aggregation inherent in patching can act as a smoothing mechanism that mitigates random fluctuations and highlights dominant patterns. This property is useful for real-world water quality datasets, which often contain noise and short-term perturbations [52]. Since patches primarily capture local information, the presence of some noise or disturbances in the input data does not disproportionately affect the overall network performance.
Third, patch embedding can reduce the parameters required for modeling. By representing time series as aggregated blocks rather than individual time points, patching lowers the dimensional burden of temporal representation and can improve computational efficiency. However, whether this also reduces the number of model parameters depends on the specific implementation.
The incorporation of channel independence or cross-variable attention modules affects forecasting capability across different lead times. The former module type is better suited to short-term forecasting, while the latter excels at medium-term forecasting. PatchTST employs a channel-independent strategy, processing each time series channel in separate pipelines (Supplementary Fig. S6). By extracting robust features from individual channels, PatchTST efficiently captures the patterns and trends associated with the target variable, showing competitive performance for short-term forecast horizons. The ablation study conducted on PatchTST further showed that channel independence contributed to short-term forecasting, and its removal left medium-term performance stable or slightly improved. This suggests that while channel independence is beneficial for capturing short-term dynamics, it may limit a model's capacity to exploit crucial intervariable dependencies [53], which is a prominent, vital requirement for accurate medium-term forecasts. In contrast, Crossformer employs a cross-variable modeling strategy designed to capture intricate intervariable dependencies across multiple time series channels (Supplementary Fig. S7). This architecture integrates information from heterogeneous data dimensions, enabling comprehensive modeling of system dynamics and yielding improved medium-term forecasting accuracy.
4.3. Ecological interpretation of predictor importance
Despite the differences across the models and reservoirs, the predictors’ relative importance for Chl a forecasting showed a consistent pattern. The primary importance of current Chl a for short-term forecasts highlighted its dominant role in determining near-future Chl a concentrations. As the lead time increased, the relative importance of current Chl a declined, whereas water temperature and the nutrient variables gradually became dominant predictors of future Chl a concentrations. These findings align with the literature on aquatic ecology [30]. Future phytoplankton biomass can be determined by combining the current biomass level and the increment over the lead time. The relative contribution of the current biomass level to this increment decreases as the lead time increases. Water temperature and nutrients are two major factors that affect phytoplankton growth and accumulation [49]. Their relative importance increases with extended lead time, while the relative importance of current Chl a gradually decreases [30]. The consistency of these patterns of feature importance across both reservoirs suggests that the identified relationships are not dataset-specific but reflect the general ecological mechanisms governing phytoplankton dynamics under different trophic conditions. This supports the broader applicability of the identified forecasting behaviors across contrasting ecological scenarios.
The relative importance values of current Chl a deduced from PatchTST were always larger than those obtained from Crossformer, mainly due to modular differences: the channel independence module in PatchTST captured temporal dependencies, whereas the cross-variable attention module in Crossformer emphasized interactions among multiple predictors. Based on domain knowledge of aquatic ecology, current Chl a contributes more to short-term forecasting, whereas other predictors are required for medium-term forecasting. PatchTST extracted information from intra-series temporal patterns in Chl a, thereby improving forecasting accuracy and highlighting the importance of current Chl a.
4.4. Implications for water quality forecasting
Machine learning models are increasingly used for water quality forecasting, but model choice should depend on data structure, forecast horizon and the ecological processes represented by the predictors. Here, model comparison, ablation analysis, and DeepSHAP interpretation showed that forecasting performance was linked to specific architectural modules, including patch embedding, channel independence, and cross-variable attention. These results provide three implications for water quality forecasting.
Prioritize deep forecasting models when appropriate. Our study of the two reservoir cases showed that PatchTST and Crossformer achieved high forecasting accuracy, robust feature importance, and process-aligned explainability. These findings suggest the potential of these deep forecasting models for Chl a forecasting in similar reservoir systems. When appropriately validated, such models may provide decision-makers with useful forecasts of future water quality and interpretable information on influential predictors.
Adapt core modules to forecast horizons. Although advanced machine learning methods are constantly emerging and evolving, their structures often consist of core innovative modules. Patch embedding excels at full-period forecasting, channel independence works well for short-term tasks, and cross-variable modeling is effective for longer lead times. This understanding can help modelers determine which core modules to use and then align deep forecasting models with the demands of the forecast horizon.
Use domain knowledge to guide modeling. Integrating process-based understanding with data-driven models has become an increasingly important approach in environmental modeling [50,54]. Based on the common understanding of phytoplankton growth processes, short-term dynamics are primarily controlled by current biomass and immediate physiological responses, while medium-term dynamics increasingly reflect the delayed effects of nutrient enrichment and temperature. This study provides a possible explanation for why PatchTST and Crossformer exhibited better performance for short-term and longer-lead forecasting, respectively. PatchTST's channel-independent design effectively captured the temporal autocorrelations in Chl a, aligning with the strong persistence of phytoplankton biomass over short horizons. Crossformer's cross-variable attention mechanism was well suited to modeling lagged dependencies between nutrients and temperature, consistent with the ecological mechanisms that govern medium-term phytoplankton variability. This consistency between existing domain knowledge and our findings demonstrates that process understandings can inform the development of interpretable and accurate forecasting models.
4.5. Limitations and future work
This study has several limitations. First, we did not investigate long-term forecasting performance beyond 15 days. Forecasting accuracy declined substantially with longer lead times, particularly for KS. Improving long-term ecological forecasting is essential for providing sufficient lead times for water quality management and bloom mitigation strategies. Future work could incorporate additional environmental predictors, such as meteorological forecasting outputs and hydrodynamic information, to improve longer-term ecological forecasting.
Second, although the present study included two reservoirs with contrasting trophic characteristics, predictor availability, and bloom temporal dynamics, the datasets could not fully represent the diversity of aquatic ecosystems and extreme environmental conditions encountered in real-world applications. Severe cyanobacterial outbreaks, abrupt ecological regime shifts, extreme hydrological disturbances, and highly eutrophic systems were not evaluated. The findings should therefore be extrapolated cautiously to broader lake types, climatic regions, and water quality variables. Future studies should validate deep forecasting architectures across more diverse trophic states, disturbance regimes, and ecological transition events to establish generalizable forecasting frameworks.
Third, although lower generalization error indicates more accurate point predictions on the test set, it does not fully characterize forecast reliability or predictive risk. Future studies should incorporate uncertainty quantification to support risk-based water quality management and to develop trustworthy machine learning models. Process knowledge could also be integrated into deep forecasting models, for example through process-informed predictors or loss-function constraints, to improve ecological consistency and interpretability.
5. Conclusion
We evaluated two conventional baseline models and seven advanced deep forecasting architectures for multivariate water quality forecasting under different ecological scenarios in two reservoirs with contrasting trophic characteristics, predictor availability, and bloom temporal dynamics. PatchTST performed best for short-term Chl a forecasting, whereas Crossformer performed best at medium-term horizons. Ablation analyses identified patch embedding as the most influential architectural module, while channel-independent learning and cross-variable attention contributed differently across forecast horizons. Channel independence enhanced short-term forecasting by strengthening the ability to capture temporal persistence, whereas cross-variable attention improved medium-term forecasting by capturing delayed ecological interactions among the environmental variables. Model interpretation revealed a consistent shift in predictor importance from current Chl a concentration to water temperature and nutrient variables as the forecast horizon increased, which aligns with established ecological understanding of phytoplankton dynamics. These findings demonstrate that forecast performance depends on model complexity and the compatibility between architectural design and ecological temporal structures. Overall, our module-level analyses and ecological interpretation provide insights into how core architectural modules influence forecast performance across reservoirs with contrasting trophic and bloom conditions. These findings inform the development of interpretable and horizon-adaptive water quality forecasting models to support accurate prediction and interpretation of aquatic dynamics. Future studies should evaluate these architectures across broader trophic states and extreme environmental conditions, while incorporating uncertainty quantification and additional environmental drivers.
CRediT authorship contribution statement
Yiqi Yu: Writing – original draft, Visualization, Validation, Methodology, Formal analysis, Data curation. Mingzhen Zhang: Writing – review & editing, Visualization, Validation, Funding acquisition, Formal analysis. Zhongyao Liang: Writing – review & editing, Investigation, Funding acquisition, Formal analysis, Data curation. Fan Qu: Validation, Formal analysis. Yao Wang: Visualization, Validation. Xiangzhen Kong: Writing – review & editing, Validation. Karsten Rinke: Writing – review & editing, Data curation. Nengwang Chen: Writing – review & editing, Supervision.
Data availability
Data and codes are available from the corresponding author upon reasonable request (Z. Liang, liangzhongyao@xmu.edu.cn).
Declaration of competing interests
We declare that (1) this manuscript contains original scientific research that has not been published elsewhere previously and (2) all authors listed on the above-referenced manuscript are aware of the submission to this journal.
Acknowledgments
This research was supported by the National Key Research and Development Program of China (2023YFC3209900), National Natural Science Foundation Youth Basic Research Project (424B2052), and State Key Laboratory of Lake and Watershed Science for Water Security.
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.ese.2026.100755.
Appendix A. Supplementary data
The following is the Supplementary data to this article:
References
- 1.Vörösmarty C.J., McIntyre P.B., Gessner M.O., Dudgeon D., Prusevich A., Green P., Glidden S., Bunn S.E., Sullivan C.A., Liermann C.R. Global threats to human water security and river biodiversity. Nature. 2010;467:555–561. doi: 10.1038/nature09440. [DOI] [PubMed] [Google Scholar]
- 2.Heydari S., Nikoo M.R., Mohammadi A., Barzegar R. Two-stage meta-ensembling machine learning model for enhanced water quality forecasting. J. Hydrol. 2024;641 [Google Scholar]
- 3.Yan T., Zhou A.N., Shen S.L. Prediction of long-term water quality using machine learning enhanced by Bayesian optimisation. Environ. Pollut. 2023;318 doi: 10.1016/j.envpol.2022.120870. [DOI] [PubMed] [Google Scholar]
- 4.Zhi W., Appling A.P., Golden H.E., Podgorski J., Li L. Deep learning for water quality. Nat. Water. 2024;2:228–241. doi: 10.1038/s44221-024-00202-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Libera D.A., Sankarasubramanian A. Multivariate bias corrections of mechanistic water quality model predictions. J. Hydrol. 2018;564:529–541. [Google Scholar]
- 6.Nong X.Z., He Y., Chen L.H., Wei J.H. Machine learning-based evolution of water quality prediction model: an integrated robust framework for comparative application on periodic return and jitter data. Environ. Pollut. 2025;369 doi: 10.1016/j.envpol.2025.125834. [DOI] [PubMed] [Google Scholar]
- 7.Zamani M.G., Nikoo M.R., Rastad D., Nematollahi B. A comparative study of data-driven models for runoff, sediment, and nitrate forecasting. J. Environ. Manag. 2023;341 doi: 10.1016/j.jenvman.2023.118006. [DOI] [PubMed] [Google Scholar]
- 8.Xu R., Hu S.R., Wan H., Xie Y.L., Cai Y.P., Wen J.H. A unified deep learning framework for water quality prediction based on time-frequency feature extraction and data feature enhancement. J. Environ. Manag. 2024;351 doi: 10.1016/j.jenvman.2023.119894. [DOI] [PubMed] [Google Scholar]
- 9.Fu G.T., Jin Y.W., Sun S., Yuan Z.G., Butler D. The role of deep learning in urban water management: a critical review. Water Res. 2022;223 doi: 10.1016/j.watres.2022.118973. [DOI] [PubMed] [Google Scholar]
- 10.Li Y., Shi K., Zhu M.Y., Li H.Y., Guo Y.L., Miao S., Ou W., Zheng Z.B. Data-driven models for forecasting algal biomass in a large and deep reservoir. Water Res. 2025;270 doi: 10.1016/j.watres.2024.122832. [DOI] [PubMed] [Google Scholar]
- 11.Peng L., Wu H., Gao M., Yi H.L., Xiong Q.Y., Yang L.D., Cheng S.P. TLT: recurrent fine-tuning transfer learning for water quality long-term prediction. Water Res. 2022;225 doi: 10.1016/j.watres.2022.119171. [DOI] [PubMed] [Google Scholar]
- 12.Pyo J., Park L.J., Pachepsky Y., Baek S.-S., Kim K., Cho K.H. Using convolutional neural network for predicting cyanobacteria concentrations in river water. Water Res. 2020;186 doi: 10.1016/j.watres.2020.116349. 116349–116349. [DOI] [PubMed] [Google Scholar]
- 13.Xu T.T., Coco G., Neale M. A predictive model of recreational water quality based on adaptive synthetic sampling algorithms and machine learning. Water Res. 2020;177 doi: 10.1016/j.watres.2020.115788. [DOI] [PubMed] [Google Scholar]
- 14.Lin S.S., Lin W.W., Wu W.T., Zhao F.Y., Mo R.C., Zhang H.T. SegRNN: segment recurrent neural network for long-term time series forecasting. IEEE Internet Things J. 2026;13:9861–9871. [Google Scholar]
- 15.Liu Y., Wu H.X., Wang J.M., Long M.S. Non-stationary transformers: exploring the stationarity in time series forecasting. Adv. Neural Inf. Process. Syst. 2022;35:9881–9893. [Google Scholar]
- 16.Nie Y.Q., Nguyen N.H., Sinthong P., Kalagnanam J. The Eleventh International Conference on Learning Representations. 2023. A time series is worth 64 words: long-term forecasting with transformers. [Google Scholar]
- 17.Wu H.X., Hu T.G., Liu Y., Zhou H., Wang J.M., Long M.S. The Eleventh International Conference on Learning Representations. 2023. TimesNet: temporal 2d-variation modeling for general time series analysis. [Google Scholar]
- 18.Zeng A.L., Chen M.X., Zhang L., Xu Q. Are transformers effective for time series forecasting? Proc. AAAI Conf. Artif. Intell. 2023:11121–11128. [Google Scholar]
- 19.Zhang Y.H., Yan J.C. The Eleventh International Conference on Learning Representations. 2023. Crossformer: transformer utilizing cross-dimension dependency for multivariate time series forecasting. [Google Scholar]
- 20.Zhou H.Y., Zhang S.H., Peng J.Q., Zhang S., Li J.X., Xiong H., Zhang W.C. Proceedings of the AAAI Conference on Artificial Intelligence. 2021. Informer: beyond efficient transformer for long sequence time-series forecasting; pp. 11106–11115. [Google Scholar]
- 21.Jin L., Chen H.H., Matsuzaki S.I.S., Shinohara R., Wilkinson D.M., Yang J. Tipping points of nitrogen use efficiency in freshwater phytoplankton along trophic state gradient. Water Res. 2023;245 doi: 10.1016/j.watres.2023.120639. [DOI] [PubMed] [Google Scholar]
- 22.Naderian D., Noori R., Kim D., Jun C., Bateni S.M., Woolway R.I., Sharma S., Shi K., Qin B., Zhang Y. Others, environmental controls on the conversion of nutrients to chlorophyll in lakes. Water Res. 2025;274 doi: 10.1016/j.watres.2025.123094. 123094–123094. [DOI] [PubMed] [Google Scholar]
- 23.Kong X.Z., Zhan Q., Boehrer B., Rinke K. High frequency data provide new insights into evaluating and modeling nitrogen retention in reservoirs. Water Res. 2019;166 doi: 10.1016/j.watres.2019.115017. [DOI] [PubMed] [Google Scholar]
- 24.Rinke K., Kuehn B., Bocaniov S., Wendt-Potthoff K., Büttner O., Tittel J., Schultze M., Herzsprung P., Rönicke H., Rink K., Rinke K., Dietze M., Matthes M., Paul L., Friese K. Reservoirs as sentinels of catchments: the rappbode reservoir observatory (Harz Mountains, Germany) Environ. Earth Sci. 2013;69:523–536. [Google Scholar]
- 25.Chen C., Chen Q.W., Yao S.Y., He M.N., Zhang J.Y., Li G., Lin Y.Q. Combining physical-based model and machine learning to forecast chlorophyll-a concentration in freshwater lakes. Environ. Sci. Technol. 2024;907 doi: 10.1016/j.scitotenv.2023.168097. [DOI] [PubMed] [Google Scholar]
- 26.Meng H.B., Zhang J., Chang Y., Zheng Z. A new method for predicting chlorophyll-a concentration in a reservoir: coupling EFDC hydrodynamic and water quality model with ConvLSTM-MLP network. J. Hydrol. 2025;660 [Google Scholar]
- 27.Jia Q.M., Xu C.Q., Jia H.F., Velazquez C., Leng L.Y., Yin D.K. Mining spatiotemporal information for harmful algal bloom forecasting and mechanism interpreting. ACS ES&T Water. 2024;4:2608–2618. [Google Scholar]
- 28.Hochreiter S., Schmidhuber J. Long short-term memory. Neural Comput. 1997;9:1735–1780. doi: 10.1162/neco.1997.9.8.1735. [DOI] [PubMed] [Google Scholar]
- 29.Breiman L. Random forests. Mach. Learn. 2001;45:5–32. [Google Scholar]
- 30.Liang Z.Y., Zou R., Chen X., Ren T.Y., Su H., Liu Y. Simulate the forecast capacity of a complicated water quality model using the long short-term memory approach. J. Hydrol. 2020;581 [Google Scholar]
- 31.Regier P., Duggan M., Myers-Pigg A., Ward N. Effects of random forest modeling decisions on biogeochemical time series predictions. Limnol Oceanogr. Methods. 2022;21:40–52. [Google Scholar]
- 32.Pyo J.C., Pachepsky Y., Kim S., Abbas A., Kim M., Kwon Y.S., Ligaray M., Cho K.H. Long short-term memory models of water quality in inland water environments. Water Res. X. 2023;21 doi: 10.1016/j.wroa.2023.100207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Benidis K., Rangapuram S.S., Flunkert V., Wang Y.Y., Maddix D., Turkmen C., Gasthaus J., Bohlke-Schneider M., Salinas D., Stella L., Aubet F.-X., Callot L., Januschowski T. Deep learning for time series forecasting: tutorial and literature survey. ACM Comput. Surv. 2022;55:1–36. [Google Scholar]
- 34.Song X.B., Deng L.W., Wang H., Zhang Y.A., He Y.X., Cao W.M. Deep learning-based time series forecasting. Artif. Intell. Rev. 2025;58 [Google Scholar]
- 35.Wang S., Wang J.Y., Xin K.L., Yan H.X., Li S.P., Tao T. Enhancing real-time urban drainage network modeling through crossformer algorithm and online continual learning. Water Res. 2025;268 doi: 10.1016/j.watres.2024.122614. [DOI] [PubMed] [Google Scholar]
- 36.Zhu J.J., Yang M.Q., Ren Z.J. Machine learning in environmental research: common pitfalls and best practices. Environ. Sci. Technol. 2023;57:17671–17689. doi: 10.1021/acs.est.3c00026. [DOI] [PubMed] [Google Scholar]
- 37.Bergstra J., Bardenet R., Bengio Y., Kégl B. Algorithms for hyper-parameter optimization. Adv. Neural Inf. Process. Syst. 2011;24:2546–2554. [Google Scholar]
- 38.Yu J.X., Zheng W.G., Xu L.L., Meng F.Y., Li J., Zhangzhong L. TPE-CatBoost: an adaptive model for soil moisture spatial estimation in the main maize-producing areas of China with multiple environment covariates. J. Hydrol. 2022;613 [Google Scholar]
- 39.Amini A., Dolatshahi M., Kerachian R. Effects of automatic hyperparameter tuning on the performance of multi-variate deep learning-based rainfall nowcasting. Water Resour. Res. 2023;59 e2022WR032789. [Google Scholar]
- 40.Zhi W., Feng D.P., Tsai W.P., Sterle G., Harpold A., Shen C.P., Li L. From hydrometeorology to river water quality: can a deep learning model predict dissolved oxygen at the continental scale? Environ. Sci. Technol. 2021;55:2357–2368. doi: 10.1021/acs.est.0c06783. [DOI] [PubMed] [Google Scholar]
- 41.Shrikumar A., Greenside P., Kundaje A. Learning important features through propagating activation differences. Proceedings of the 34th International Conference on Machine Learning. 2017;70:3145–3153. [Google Scholar]
- 42.Lundberg S.M., Lee S.-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017;30:4764–4774. [Google Scholar]
- 43.Li L.B., Qiao J.D., Yu G., Wang L.Z., Li H.Y., Liao C., Zhu Z.D. Interpretable tree-based ensemble model for predicting beach water quality. Water Res. 2022;211 doi: 10.1016/j.watres.2022.118078. [DOI] [PubMed] [Google Scholar]
- 44.Vishnusai Y., Kulakarni T.R., Sowmya Nag K. Springer International Publishing; 2020. Ablation of Artificial Neural Networks; pp. 453–460. [Google Scholar]
- 45.Dhang S., Idogho A., Zhang M., Dev S. Towards accurate billboard detection: an ablation and benchmarking study of deep learning models. IET Conference Proceedings. 2024;2024:178–185. [Google Scholar]
- 46.Liu Y.X., Yang B., Xie K.T., Sun J.L., Zhu S.M. Dongting Lake algal bloom forecasting: robustness and accuracy analysis of deep learning models. J. Hazard Mater. 2025;485 doi: 10.1016/j.jhazmat.2024.136804. [DOI] [PubMed] [Google Scholar]
- 47.Sheikholeslami S., Meister M., Wang T., Payberah A.H., Vlassov V., Dowling J. Proceedings of the 1st Workshop on Machine Learning and Systems. 2021. AutoAblation: automated parallel ablation studies for deep learning; pp. 55–61. [Google Scholar]
- 48.Wan H., Xiang L., Cai Y.P., Xie Y.L., Xu R. Temporal and spatial feature extraction using graph neural networks for multi-point water quality prediction in river network areas. Water Res. 2025;281 doi: 10.1016/j.watres.2025.123561. [DOI] [PubMed] [Google Scholar]
- 49.Anderson S.I., Franzè G., Kling J.D., Wilburn P., Kremer C.T., Menden-Deuer S., Litchman E., Hutchins D.A., Rynearson T.A. The interactive effects of temperature and nutrients on a spring phytoplankton community. Limnol. Oceanogr. 2022;67:634–645. [Google Scholar]
- 50.Tripathy K.P., Mishra A.K. Deep learning in hydrology and water resources disciplines: concepts, methods, applications, and research directions. J. Hydrol. 2024;628 [Google Scholar]
- 51.Kong X.J., Chen Z.H., Liu W.Y., Ning K.L., Zhang L.C., Muhammad Marier S., Liu Y.C., Chen Y.H., Xia F. Deep learning for time series forecasting: a survey. Int. J. Mach. Learn. Cybern. 2025;16:5079–5112. [Google Scholar]
- 52.Uddin M.G., Rahman A., Taghikhah F.R., Olbert A.I. Data-driven evolution of water quality models: an in-depth investigation of innovative outlier detection approaches-A case study of Irish water quality index (IEWQI) model. Water Res. 2024;255 doi: 10.1016/j.watres.2024.121499. [DOI] [PubMed] [Google Scholar]
- 53.Han L., Ye H.J., Zhan D.C. The capacity and robustness trade-off: revisiting the channel independent strategy for multivariate time series forecasting. IEEE Trans. Knowl. Data Eng. 2024;36:7129–7142. [Google Scholar]
- 54.Wang Y., Wang W.K., Ma Z.T., Zhao M., Li W.X., Hou X.Y., Li J., Ye F., Ma W.J. A deep learning approach based on physical constraints for predicting soil moisture in unsaturated zones. Water Resour. Res. 2023;59 e2023WR035194. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data and codes are available from the corresponding author upon reasonable request (Z. Liang, liangzhongyao@xmu.edu.cn).
