Abstract
The modelling of cascade reactions is currently playing a major role in the scale-up of industrial processes, particularly in biomass valorization. This study investigates machine learning (ML) techniques to optimize the reaction conditions of one-pot synthesis of 2,5-furandicarboxylic acid (FDCA), a biobased platform chemical, from sugarcane bagasse via Fe-Mn zeolite catalyst. The objective of the work is to evaluate various ML regressor models for predicting FDCA yield and selectivity, particularly in the context of limited data sets generated through box-Behnken design of experiments. Three different ML models, such as ridge regression, support vector regressor (SVR), and gradient boosting regression (GBR), were compared to identifying the most suitable model for accurate prediction. Among the models, the ridge regression approach demonstrated superior performance to the lowest mean absolute error (MAE) of 0.595 and the highest coefficient of determination (R2) of 95.6% for FDCA yield. The SVR model optimized reaction conditions as 165.65 °C, 5.41 h, and 0.80 g of catalyst dosage to yield 66.65% of FDCA. The proposed ML-based regressor provides new insights into effectively handling small data sets and highlights the potential of machine learning for reliable prediction and process optimization in biomass conversion to FDCA.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-54149-0.
Keywords: Artificial intelligence; Machine learning; 2,5-furandicarboxylic acid; Supported vector regressor; Ridge regression; Gradient boosting regression
Subject terms: Chemistry, Energy science and technology, Engineering, Environmental sciences
Introduction
The sustainable chemical industry has recently focused more on the conversion of biomass into high-value products, including platform chemicals and alternative fuels. This shift responds to the rapid decline of fossil resources and growing concern of environmental challenges1. Among the various biomass derived value-added chemicals, 2,5-furandicarboxylic acid garnered specific interest, due to its use as a platform chemical for making various polymers and other chemicals, with specific interest directed towards its use as plastic precursor as an alternative to petroleum derived terephthalic acid (TPA) for polyester production. FDCA has received significant attention owing to its wide applications in developing various polymers, including nylons, polyesters, polyamides, and plasticizers2. FDCA can be derived from renewable biomass sources such as glucose or fructose through multi-step catalytic process involving hydrolysis, isomerization, dehydration and oxidation. Researchers have explored several approaches for the synthesis of FDCA from various feedstocks, including 5-hydroxymethyl furfural (HMF), furfural, and diglycolic acid3. Among these, the chemo-catalytic oxidative conversion of HMF to FDCA remains the most widely studied and commonly employed pathway. HMF itself can be synthesized from biomass-derived feedstocks such as glucose and fructose or directly from lignocellulosic biomass through acid catalyzed dehydration in the presence of Lewis’s and Bronsted acid sites on the catalyst surface4. However, this FDCA synthesis route is complex, as it involves the sequential conversion of biomass or fructose to HMF, followed by oxidative transformation through various intermediates, viz., 2,5-diformylfuran (DFF), hydroxymethyl furancarboxylic acid (HMFCA), and 5-formylfurancarboxylic acid (FFCA) finally to produce FDCA5.
In this context, the concept of one-pot FDCA synthesis represents a sustainable alternative as it eliminates the need for separating and isolating the major intermediate HMF. This integrated method of synthesis offers significant advantages that align with the principles of green chemistry6,7. Designing a multifunctional heterogeneous catalyst capable of facilitating both reduction and oxidation reactions provides a promising approach for enabling efficient one-pot synthesis of FDCA. This heterogeneous catalysis addresses the limitations associated with the secondary environment pollution typically occurred in homogenous catalysis8,9. Despite advances in catalyst development, the efficient direct conversion of biomass to FDCA remains challenging, as yield and selectivity are highly dependent on reaction parameters including temperature, catalyst type and loading, substrate dosage etc10. Conventional optimization approaches including response surface methodology (RSM), provide as effective statistical tool to explore the influence of process variables on FDCA yield and selectivity11. In particular, the box-Behnken design offers advantages in reducing the number of required experiments while capturing second-order interactions. Although these designs are statistically robust, the limited number of experimental data points can restricts the predictive accuracy of model, particularly when non-linear interactions among variables are significant12,13. While, RSM offers valuable insights into factor-response relationship, its predictive accuracy may be limited when used with small experimental datasets14.
The modelling of industrial processes through intelligent manufacturing has emerged as a powerful tool to improve process efficiency and product quality. The fourth industrial revolution, characterized by emerging technologies, specifically artificial intelligence (AI), is playing a pivotal role in modern industrial development15. This modelling technique can predict the desired output parameters when there is sufficient experimental data. Machine Learning (ML) is a subset of AI and employs specific algorithms to develop models with the minimum number of experimental tests16. ML has been extensively studied by many researchers in the field of heterogeneous catalysis, specifically in applications such as CO2 reduction17, water splitting18 and O2 reaction19 etc. Various ML methods such as Linear Regression (LR), Artificial Neural Network (ANN), K-nearest Neighbor (KNN) regression, Random Forest (RF), Decision Trees and Support Vector Regression (SVR) etc., have been employed in heterogenous catalysis reaction optimization20,21. Ridge regression, is a linear model with L2 regularization, it can manage multicollinearity among the input process parameters and provide interpretable relationships. It is used for develop ML models for catalyst feature designs20,22. While, gradient boosting regression (GBR), a powerful non-linear ensemble method, it has ability to capture complex non-linear interactions between factors and reduce prediction errors. For instance, Ross-Veitía and coworkers analysed the emissions from steam boilers using ML techniques. GBR, showed better performance in the predictions made on the test data, and outperformed all other model such as DNN, MLR, RFR23. Additionally, the support vector regressor (SVR) model is effective in creating a strong and general predictive model that can handle smooth nonlinear patterns in the smaller data sets. Sultana et al. employed three mathematical models such as SVR, ANN, and RSM have been to predict the yield of biodiesel synthesis from papaya seed waste oil24. They reported that, SVR models fit the experimental data compared to other models with lower relative error and higher correlation coefficient24. However, only few works are reported on the application of ML to predict and optimize the synthesis of biomass derived platform chemicals. Recently, Li et al. reported a machine learning-driven framework for the optimization of catalytic oxidation of HMF to FDCA. The independent influence of input factors and their interactive effect on FDCA yield was successfully analyzed25. Similarly, our research team has reported that the FDCA yield, and selectivity can be predicted using artificial neural network- Levenberg-Marquardt (ANN-LM) model, achieving an R2 of 99.3% for both responses11. Inspired by these findings, distinct ML regression models can be employed to learn from existing experimental datasets to predict the FDCA yield within the experimental range.
In this work, a 15-point box–Behnken dataset from our previous research is utilized to evaluate the influence of key process factors on FDCA yield and selectivity. Three different ML regressor models such as ridge regression, GBR, and SVR were analyzed for their ability to accurately predict FDCA yield and selectivity. By analyzing various approaches, viz., linear, ensemble, and kernel based, this study aims to identify the best predictive model for FDCA yield and selectivity using a limited box-Behnken dataset. The optimization of process parameters to maximize the FDCA production was carried out with each model. The model’s predictions were validated with confirmation experiments with the optimized process parameters.
Materials and methods
Materials
The metal precursor like iron nitrate nonahydrate (Fe (NO3)3·9H2O) (CAS7782-61-8, 98%) and manganese nitrate Mn (NO3)2·xH2O (CAS 13446-34-9, 98%) were purchased from Loba chemie international Pvt. Ltd. Molecular sieves zeolite-5 A was procured from Sigma Aldrich. Dimethyl sulfoxide (DMSO, CAS No. 67–68-5) and potassium carbonate (K₂CO₃, CAS No. 584-08-7) were obtained from SRL Chemicals Pvt. Ltd., India, and local suppliers. Analytical grade of 2,5- furandicarboxylic acid (FDCA) (CAS 3238-40-2, 99.5%) was purchased from TCI chemical industry (India) Pvt. Ltd.
Methods
Catalyst synthesis
The Fe-Mn zeolite catalyst was synthesized via simple wetness impregnation method as per the procedure reported in the literature9. The bare zeolite was first heated at 80 °C for about 2 h to drive off any moisture. Earlier studies by Chai et al. and Pandey et al. have shown that both metal loading and the ratio between metal oxides can significantly affect FDCA yield, so those factors were taken into account here8,26. Based on this, a co-impregnation strategy was followed. The precursor solutions for Fe and Mn precursor were prepared by dissolving the required metal salts in water to achieve a total metal oxide loading of 6 wt% with a 1:1 molar ratio (pH of solution is 2–4). The calculated amounts of each precursor were dissolved in distilled water and then added slowly, drop by drop, to the preheated zeolite. The volume of solution used was guided by the pore volume obtained from BET analysis. After impregnation, the material was left to dry overnight at 80 °C. It was then calcined in air at 400 °C for 4 h using a Nastnas muffle furnace (India). The calcination conditions were selected in line with earlier reports8. Once prepared, the catalysts were stored in airtight containers until further use.
Catalyst characterization
The calcined catalyst was subjected to various characterization methods like X-ray diffractogram (Rigaku Ultima Ⅳ, X-ray diffractometer) and FTIR (Thermo Scientific Nicolet iS10 mid frequency Fourier Transform spectrometer) to confirm the structural changes of the catalyst. FE-SEM (ZEISS Gemini SEM 300 Field emission scanning electron microscope) was done for morphology analysis. UVDRS (JASCO V-750 UV-visible spectrophotometer) and X-ray photoelectron spectroscopy (XPS) (Thermofisher Nexsa) to confirm the presence of different metal oxides on the surface was also done, which is indeed necessary for the selective oxidation and reduction process in a one-pot synthesis. The surface area was determined by the method of Brunauer, Emmett, and Teller (BET) and the acidity was calculated from ammonia temperature programme desorption (NH3-TPD) using BELCAT Ⅱ instrument, Japan.
One-pot synthesis & analysis of FDCA
The validation synthesis for the optimized process parameters by the specific model was done with sugarcane bagasse as the feedstock. The collected feedstock was dried at 80 °C and ground to a fine powder of a particle size below 300 μm. The one-pot synthesis of FDCA was carried out in a teflon-lined autoclave reactor. Ground sugarcane bagasse (1 g) was mixed with a water-DMSO solution (50 mL) at a 3:1 (v/v) ratio and introduced into the reactor. The mixture was sonicated at room temperature for a specified duration, after which the required amount of catalyst was added. The pH of the reaction medium was adjusted to 8 using a base additive, and the reaction was conducted under optimized conditions. Upon completion, the mixture was cooled and filtered through a vacuum suction filter, and the filtrate was collected for analysis.
The FDCA yield was calculated from HPLC analysis (LC 2030 C plus Shimadzu). Quantification of FDCA was performed by HPLC using a Shimadzu LC-2030 C Plus system equipped with a C18 reverse-phase column (250 mm × 4.6 mm, 5 μm particle size), UV detector, and low-pressure gradient pump. The mobile phase consisted of 40% methanol and 60% 0.01% orthophosphoric acid (v/v), operated at 35 °C with an injection volume of 20 µL and a flow rate of 1 mL min⁻¹. Data acquisition and quantification of FDCA concentration were performed using LC solution software. The FDCA yield was calculated based on the cellulose content of the biomass, considering the molecular weight of glucose units (Eq. 1). Catalyst selectivity toward FDCA (CoCrZ) was determined using Eq. 2.
![]() |
1 |
![]() |
2 |
The FDCA yield and selectivity data were modelled from the optimization of the process parameters designed by box-Behnken design of experiments for the one-pot synthesis of FDCA from sugarcane bagasse (Table 1). The independent or input variables of the model were reaction time and temperature, and catalyst dosage. The lower limit for catalyst dosage was chosen as 0.6 g, while the higher limit was set at 1 g. The range of reaction time was selected from 4 to 6 h and the temperature of the reaction was from 130 to 170 °C. The dependent or output variables were FDCA yield and selectivity. Three different regression models such as ridge regression, support vector regressor (SVR), and gradient boosting regression (GBR) were compared to identify the most suitable model for accurate prediction of FDCA yield and selectivity. The overall methodology of the work is depicted in Fig. 1.
Table 1.
Data sets used from box-Behnken experimental design11.
| Std | Run | Input variable 1 time h | Input variable 2 temp °C | Input variable 3 catalyst dosage g | FDCA yield % | Selectivity % | RSM BBD predicted values | |
|---|---|---|---|---|---|---|---|---|
| FDCA yield % | Selectivity % | |||||||
| 8 | 1 | 6 | 150 | 1.0 | 61.29 | 74.03 | 61.28 | 74.82 |
| 4 | 2 | 6 | 170 | 0.8 | 67.54 | 80.47 | 67.46 | 80.37 |
| 6 | 3 | 6 | 150 | 0.6 | 59.06 | 71.03 | 59.39 | 70.78 |
| 3 | 4 | 4 | 170 | 0.8 | 64.54 | 78.91 | 64.79 | 79.36 |
| 11 | 5 | 5 | 130 | 1.0 | 55.94 | 65.52 | 56.19 | 65.18 |
| 7 | 6 | 4 | 150 | 1.0 | 58.82 | 68.13 | 58.48 | 68.37 |
| 2 | 7 | 6 | 130 | 0.8 | 59.56 | 71.03 | 59.30 | 70.57 |
| 10 | 8 | 5 | 170 | 0.6 | 62.84 | 76.04 | 62.58 | 76.37 |
| 9 | 9 | 5 | 130 | 0.6 | 54.03 | 65.85 | 53.95 | 66.54 |
| 12 | 10 | 5 | 170 | 1.0 | 62.80 | 73.48 | 62.87 | 72.78 |
| 14 | 11 | 5 | 150 | 0.8 | 64.01 | 87.84 | 63.66 | 86.86 |
| 5 | 12 | 4 | 150 | 0.6 | 57.85 | 78.16 | 57.85 | 77.37 |
| 13 | 13 | 5 | 150 | 0.8 | 63.38 | 86.82 | 63.66 | 86.86 |
| 1 | 14 | 4 | 130 | 0.8 | 57.56 | 71.64 | 57.36 | 71.73 |
| 15 | 15 | 5 | 150 | 0.8 | 63.59 | 85.92 | 63.66 | 86.86 |
Fig. 1.
Graphical representation of the methodology for FDCA production and modelling.
In supervised learning method, the algorithm receives input parameters in the form of labelled data and is trained to predict associated output parameters that are initially unknown. As the data set in this case is small, 100% of the data was used as feed for training the model. However, the performance of the model may be sensitive to the random train-test split if certain experimental points are chosen as the test set. So, the model was validated with leave-one-out cross-validation (LOOCV)27,28. In LOOCV, a single experimental point is used to test the model and the remaining 14 points are used to fit the model. This process is repeated 15 times until each point has been selected as the test set.
Regression models
Model selection was guided by two constraints: (1) dataset size (15 box–Behnken runs) and (2) the need to assess both interpretability and predictive power across linear and non-linear hypothesis classes. The three chosen models serve distinct roles:
Ridge regression was employed as a baseline linear model to account for potential multicollinearity among process variables. The model adds an L2-penalty term to the ordinary least squares (OLS) cost function. This limits the sizes of coefficient magnitudes and improves generalization. By reducing coefficients, ridge decreases variance when predictors are correlated, while still allowing for easy interpretation of the model parameters29. This interpretability facilitates mechanistic insight and the identification of linear trends between process variables and FDCA metrics. The penalty hyperparameter (α) was optimized using grid search with cross-validation over the range of 0.001to 10030.
Support vector regressor (SVR) was considered to capture potential non-linear relationships between process parameters, FDCA yield, and selectivity. SVR utilizes the structural risk minimization and employs ε- insensitive loss function by ignoring minor deviations from the true responses while penalizing large errors. Kernel functions allow SVR to project input variables into higher dimensional feature space. This makes it effective for modelling smooth non-linear relationships in experimental datasets with smaller size29. In this study, a radial basis function (RBF) kernel was applied. The models hyperparameters like the regularization constant (c), kernel coefficient (γ), and loss margin (ε) were optimized using grid over 0.1–100, 0.001–1.001, and 0.01–0.5 respectively.
Gradient Boosting Regression (GBR) was adopted for its ability to model complex interactions between variables and handle mixed feature scales without requiring detailed feature engineering31. GBR builds a series of shallow decision tress in a consecutive manner to reduce the prediction residuals. This model is robust in capturing non-linear relationship between variables and their effect on output variable that often occur in catalytic reaction system32. Moreover, it also provides feature importance, which is helpful for understanding mechanism. This three different modelling on regularized linear, kernel-based nonlinear, and boosting ensemble enables the identification of the best trade-off between interpretability, robustness to small samples, and predictive accuracy. The models hyperparameters like the learning rate (ɳ) and number of estimators(n_estimators) were optimized using grid over 0.01–0.2 and 50–200 respectively.
Model training and evaluation
All models were implemented in Python 3.11 using the scikit-learn (v1.5) library. Hyperparameter tuning was performed via GridSearchCV. The objective metric for tuning was root-mean-square error (RMSE). Model performance was evaluated on the independent test dataset using standard regression metrics as follows:
Mean Squared Error (MSE) and Root Mean Squared Error (RMSE).
It is a common accuracy measure for regression problems that quantifies the average squared difference between true values and model predictions. The square root of MSE gives the RMSE. For a dataset with i observations with true targets Xei and predictions Xpi the sample MSE is.
![]() |
3 |
![]() |
4 |
Mean Absolute Error (MAE).
MAE is a widely used metric for evaluating the accuracy of predictive models. It calculates the average of the absolute difference between predicted values and observed values, offering a straightforward and interpretable measure of model performance. A lower MAE signifies higher predictive accuracy.
![]() |
5 |
Coefficient of Determination (R²).
The coefficient of determination (R2) indicates the proportion of variance in the dependent variable that is explained by the independent variables in a regression model. It ranges from 0 to 1, with higher values indicating a stronger model fit. An R2 of 1 implies perfect prediction, while 0 indicates that the model explains none of the variability in the response data.
![]() |
6 |
![]() |
7 |
Where, Xe is the experimental value, Xp is the predicted value and
is the sample mean. After fitting the training data the best ML model, the trained model has imported to NSGA Ⅱ (Genetic algorithm) to maximize the FDCA yield and selectivity. Validation of the predicted maximum FDCA yield was also done as per the synthesis procedure at optimized reaction conditions.
Model interpretability
Interpreting the model with explanatory techniques adds practical value for decision-making, as it helps uncover both overall trends and individual feature contributions33. The SHapley Additive exPlanations (SHAP) and partial dependence plot (PDP) were employed to enhance the interpretability of the model. The PDP highlights overall trends and SHAP gives a more detailed view by showing how each feature contributes to a prediction.
SHAP is a model-agnostic method based on concepts from game theory. It assigns a contribution value to each feature, indicating how much it influences the prediction for a given instance. These contributions are derived from Shapley values, which consider the marginal effect of each feature in combination with others. In this way, the model’s prediction can be understood as the sum of all individual feature contributions, providing a clear and additive explanation of the output. Whereas, PDP vary one feature at a time while keeping the others fixed, allowing us to see how that particular variable influences the prediction34. This makes it easier to understand the relationship between input features and model outputs, which is a key advantage over many other interpretability approaches35.
Results and discussion
Physicochemical analysis of catalyst
From Fig. 2a, represents the FTIR spectrum of catalyst. The major peaks observed at 1007 cm− 1, 552 cm− 1, and 467 cm− 1 correspond to the Si-Al-O asymmetric stretching, double ring in the framework structure of zeolite, and bending of Si-Al-O linkages respectively36. Thus, the presence of characteristic absorption band for zeolite-5 A as reported in literature even after the metal oxide impregnation, indicating that Fe and Mn impregnation did not significantly alter the zeolite framework structure37.
Fig. 2.
(a) FTIR spectrum, (b) XRD spectrum, (c) FESEM image, and (d) UV-DRS spectrum of catalyst.
The XRD spectrum also reveals similar to the highly crystallized zeolite XRD spectrum as shown in Fig. 2b. For raw zeolite-5 A, the observed peaks at 2θ at 7.1°, 10.21°, 12.3°, 21.68°, 23.84°, 27.1°, 29.96°, and 34.1° can be assigned to the (100), (110), (222), (300), (311), (321), (410), and (421) crystal planes of aluminosilicate respectively (ICDD no. 04–025-2961)38. Whereas for the catalyst, the metal oxide impregnated zeolites displayed peaks that were present in raw zeolite. The reduction in peak intensity of metal oxide-based zeolite compared to raw zeolite can be attributed to the presence of metals like Mn and Fe on the surface38. Additionally, the XRD pattern of Fe oxide at 24.18°, 33.26°, 35.6°, 49.72°, 57.34°, 62.66°, and 64.34° can be assigned to the (012), (104), (110), (024), (018), (214), and (300) crystal planes was consistent with the values in the standard card (JCPDS 86–0550)39. Whereas Mn oxide diffraction at 29.68°, 52.66°, 66.58°, and 74.72° attributed to the plane (110), (211), (112), and (533) respectively (ICDD card no: 00–001-1127)40. However, some of the peaks corresponding to each metal oxide coincide with those of zeolite. This suggests that, Fe and Mn, species are well dispersed within the zeolite framework without accumulation or sintering, allowing the zeolite diffraction pattern to dominate. The crystallite gain size was calculated using the Debye-Scherrer equation, and it is found that 19.83 nm and 15.9 nm for bare zeolite and catalyst respectively. The zeolite framework peaks exhibited a slight positive peak shift after the metal oxide impregnation, due to minor contraction of the d spacing41. Furthermore, the higher full width at half maxima (FWHM) for the dominant peak in the metal oxide impregnated catalyst (0.498°) was slightly higher than the pure zeolite framework (0.495°). This peak broadening indicates a reduction in crystallite size, confirming the successful incorporation of metal oxides within the zeolite framework.
Figure 2c represents the SEM image of catalyst showing the uniform distribution of both metal oxides on the catalyst without sintering or aggregation42. Additionally, the elemental composition of catalysts was also analyzed using SEM-EDS (Fig. 3a). The EDS layered image is displayed in Fig. 3b. The elemental mapping (Fig. 3c) of catalysts such as Si, Al, O, Fe, and Mn, validates the uniform distribution of impregnated metal oxide on the catalyst surface. Figure 2d represents the UV DRS spectrum of catalyst. The UVDRS spectrum of the catalyst shows the typical framework absorption of the zeolite around 270 nm43. Bands below 300 nm can be linked to isolated Fe³⁺ species in tetrahedral or octahedral coordination. A feature around 315 nm likely indicates the presence of small Fe–oxo clusters. In the visible region, the bands near 451 nm, 576 nm, and 684 nm may be associated with iron oxide aggregates and overlapping of weak d–d transitions involving both Fe and Mn species44,45. However, because of the ligand-to-metal charge transfer (LMCT) from both ions, it is not possible to conclude the presence of each metal oxides. So, to confirm the presence of oxidation states of metal species XPS analysis was also done. From the XPS analysis (Fig. 4), it can be concluded that the various oxidation states like Fe2+/Fe3+ and Mn3+/Mn4+ are present on the catalyst surface. The peaks for Mn2P and Fe2P spectrum are summarized in Table 2. Furthermore, the quantitative analysis of both Fe and Mn ions were done, and it is found that 2.9 and 2.44 atomic percentage for Fe and Mn respectively.
Fig. 3.
(a) EDs spectrum, (b) EDS layered image, and (c) elemental mapping of catalyst.
Fig. 4.
XPS (a) survey scan, (b) Fe2p, (c) Mn2p, and (d) O1s spectrum of catalyst.
Table 2.
XPS peak positions and corresponding chemical states.
| Element | Core level | Binding energy (eV) | Assignment |
|---|---|---|---|
| Fe | Fe2p1/2 | 724.01 | Fe2+ |
| Fe | Fe2p1/2 | 725.66 | Fe3+ |
| Fe | Fe2P3/2 | 710.70 | Fe2+ |
| Fe | Fe2P3/2 | 712.49 | Fe3+ |
| Mn | Mn2P3/2 | 642.20 | Mn3+ |
| Mn | Mn2P3/2 | 643.51 | Mn4+ |
| Mn | Mn2p1/2 | 654.56 | Mn4+ |
| Mn | Mn2p1/2 | 653.81 | Mn3+ |
| O | O1s | 530.02 | Lattice oxygen |
| O | O1s | 532.81 | Surface oxygen |
The nitrogen sorption isotherms plotting the adsorbed volume (in cc/g) vs. the relative pressure (P/P0) at 77 K is displayed in Fig. 5. As seen in Fig. 5a, zeolite and catalyst exhibited a sharp increment in volume of adsorption at low relative pressure (P/P0), this behavior confirms the typical type 1 Langmuir isotherm characteristic to micropore adsorption. The bare zeolite shows a surface area of 300 m2/g with a pore volume of 0.21 cc/g. Compared to bare zeolite, the catalyst’s isotherm displayed a hysteresis loop, validating the formation of certain mesopores by the metal oxides. The Fe-Mn zeolite catalyst possessed a surface area of 58.40 m²/g with an average pore radius of 26.73 Å. This characteristic was further validated by the modest pore volume and limited nitrogen uptake observed in the adsorption isotherm giving a pore volume of 0.101 cc/g (approximately 50% reduction in pore volume compared to zeolite). This decrease in surface area and pore volume is due to the impregnation of low surface a rea Fe and Mn oxides on the zeolite pores and channels. Figure 5b displays the differential pore volume (dV/(r)) vs. on pore radius (Å), representing the pore size distribution of various catalysts. For the bare zeolite, the majority of area falls under 4.4 A°, indicating microporous nature of surface. The catalyst has a broad distribution with major area falls in the range of 18–25 Å, attributed to the presence of larger mesopores46. This mesoporous nature of catalysts is advantageous for catalytic applications that require a balance between accessibility and surface exposure47. Moreover, the catalyst surface acidity was deduced by performing NH3-TPD analysis and the total acidity was observed of 2.97 mmol/g. In conclusion, the uniform dispersion of metal oxides on the zeolite framework structure, their oxidation states, and the optimal amount of acidity of the Fe-Mn zeolite catalyst make it an ideal choice for one-pot cascade reaction to synthesize FDCA from biomass.
Fig. 5.
(a) BET isotherm and (b) pore size distribution of catalyst.
Experimental design
The FDCA yield and selectivity of each combination of process parameters as per the Box-Behnken design was calculated and tabulated (Table 1). Initially, ridge, SVR and GBR models were trained using default parameters. Following preliminary evaluations, hyperparameter optimization was performed to enhance predictive accuracy. Grid Search and Random Search techniques were employed to identify suitable configurations, and the optimized combinations for Ridge regression, SVR, and GBR are summarized in Table 3.
Table 3.
Hyperparameters for predicting FDCA yield and selectivity.
| Model | Hyperparameters search range | Best FDCA yield | Best FDCA selectivity |
|---|---|---|---|
| Ridge regression |
degree=[2] α = [0.001, 0.01, 0.1, 1, 10, 100] |
degree=[2] α = 0.1 |
degree=[2] α = 0.1 |
| SVR |
Degree=[2] C=[0.1,1,10,100] γ=[0.001,0.01,0.1,1] ε=[0.01,0.1,0.5] kernel = rbf |
[2] C=100 γ = 0.01 ε = 0.01 kernel = rbf |
[2] C=100 γ = 0.01 ε = 0.01 kernel = rbf |
| GBR |
n_estimators= [50,100,200] learning_rate= [0.01,0.05,0.1, 0.2] max_depth = [2, 3, 4] |
n_estimators = 100 learning_rate = 0.1 max_depth = 2 |
n_estimators = 200 learning_rate = 0.2 max_depth = 2 |
Evaluation metrics such as coefficient of determination (R2), mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), were used to assess how well the selected models performed, and the obtained results are shown in Table 4.
Table 4.
Performance metrics of the implemented models.
| Model | Response | Train R2 | LOOCV R2 | Train RMSE | LOOCV MSE | MAE | LOOCV MAE |
|---|---|---|---|---|---|---|---|
| Ridge | FDCA yield | 0.996 | 0.956 | 0.218 | 0.744 | 0.179 | 0.595 |
| Selectivity | 0.993 | 0.931 | 0.571 | 1.837 | 0.480 | 1.565 | |
| SVR | FDCA yield | 0.998 | 0.916 | 0.122 | 1.027 | 0.050 | 0.811 |
| Selectivity | 0.994 | 0.918 | 0.544 | 2.003 | 0.195 | 1.655 | |
| GBR | FDCA yield | 0.998 | 0.768 | 0.124 | 1.711 | 0.068 | 1.329 |
| Selectivity | 0.997 | 0.638 | 0.350 | 4.224 | 0.130 | 3.312 |
The comparative evaluation of ridge regression, SVR, and GBR provided insights into the suitability of different ML models for predicting FDCA yield and selectivity from a limited box-Behnken dataset. Given the limited number of data points (15 experimental runs), the final performance of the models was computed using LOOCV to give a better indication of the model’s predictive performance rather than their fitting performance. As a result, ridge regression produced the most consistent final model performance with final R² values of 0.956 and 0.931 for FDCA yield and selectivity, respectively. RMSE values were 0.744 and 1.837, and the MAE values of 0.595 and 1.565 for the FDCA yield and selectivity, respectively. Compared to other models, ridge regression demonstrated superior performance consistent with its L2 regularization advantage for small experimental data sets. In this work, the convergent final score of ridge regression suggests that the major relationships between reaction time, temperature, catalyst amount, and FDCA yields and selectivity can be modelled without a highly complex model.
The SVR model provided a middle ground between the ridge and GBR. Operating by mapping input features into a high dimensional space and identifying an optimal hyperplane to minimize prediction errors within a defined margin, SVR is extensively employed in scientific and engineering applications because it effectively handles the high dimensional inputs and non-linear patterns through kernel functions48. This kernel-based formulation reduces sensitivity to input dimensionality and often achieves lower generalization error compared to purely linear models. In our experiments, the SVR also demonstrated good final predictive capability with R² values of 0.916 and 0.918 for FDCA yield and selectivity, respectively. The corresponding RMSE values were 1.027 and 2.003, while the MAE values were 0.811 and 1.655 for FDCA yield and selectivity respectively. SVR projects the raw variables onto a higher-dimensional space, and seeks an optimal hyperplane with a certain allowed error. The use of kernel functions allows SVR to deal with smooth non-linear correlations between process parameters and property responses. But the final error of the SVR model was comparatively higher than ridge regression, which suggests improvements in stability with the regularized linear model for the current small dataset.
GBR is a versatile ensemble learning method that builds models sequentially by adding weak learners, usually decision trees. Each tree is trained to minimize the residual errors left by the previous model. By estimating the gradient of the residuals related to the model parameters, GBR improves the loss function and steadily improves the predictive accuracy48. However, in the current LOOCV test, the final predictive power of GBR using LOOCV was weaker than those of ridge regression and SVR, with R² values of 0.768 and 0.638 for FDCA yield and selectivity, respectively. The corresponding RMSE values were 1.711 and 4.224, while the MAE values were 1.329 and 3.312 for FDCA yield and selectivity respectively. This has shown that while GBR may be able to capture non-linear relationships, final predictive accuracy and stability was compromised by small data set sizes.
The correlation between actual and predicted values of FDCA yield and selectivity using three regression models is depicted in Fig. 6. Across all models, the predicted values align closely with the experimental data, as evidenced by the proximity of data points to the ideal fit line (blue dashed line). For ridge regression (Fig. 6a,b), strong correlations were observed with R² values of 0.956 (yield) and 0.931 (selectivity), confirming the ability of this linear model to capture the underlying data patterns. This indicates nearly perfect prediction accuracy. The close clustering of predicted values compared to experimental values along the ideal line shows that the model can capture nonlinear interaction and account for residual variations. However, minor deviations from the ideal line suggest some limitation in handling nonlinearity within the dataset49. Figure 6c and d represent the SVR models predictability with R2 values of 0.916 and 0.918 for FDCA yield and selectivity respectively. The comparable predictive accuracy confirms the prediction robustness across non-linear approaches, though ridge excels in generalization of small datasets. The GBR predictability is shown in Fig. 6e,f. The model exhibited a moderate performance with an R2 value of 0.768 for FDCA yield and 0.638 for selectivity is consistent with the tree-based models in similar catalytic modelling studies50. Overall, these results show that while all three regression models had strong predictive capability, ridge regression provided the most accurate and reliable predictions of FDCA yield and selectivity.
Fig. 6.
Actual versus predicted values of FDCA yield (a, c, e) and selectivity (b, d, f) using ridge regression (a, b), SVR (c, d), and GBR (e, f).
SHAP and PDP analysis
The SHAP analysis were conducted using the ridge model to analyze the effect of each process variable on the FDCA yield and selectivity (Fig. 7). For FDCA yield, the SHAP feature-importance plot indicated reaction temperature had the highest contribution (2.23), followed by catalyst loading (1.92) and reaction time (0.69). The relatively higher influence of temperature shows FDCA yield is sensitive to the heat-enabled conversion process via intermediate products51. The SHAP beeswarm plot also indicated that higher reaction temperatures have a positive influence on FDCA yield, while lower temperatures have a negative influence. Whereas, for FDCA selectivity, mean |SHAP| value of catalyst dosage (4.767) is observed being the most influential factor, followed by temperature and reaction time. This suggests that selectivity is also highly dependent on the active site exposure for the multiple conversion and oxidation step. However, it has exhibited predominantly negative value in SHAP beeswarm plot. This may be due to the increased number of active sites/excessive amounts can cause side reactions, mass transfer limitations, or catalyst deactivation52,. Thus, temperature is the prominent factor for ridge model to govern both FDCA yield and selectivity. Additionally, SHAP summary plots for SVR and GBR also presented in Fig.S1 and S2. For SVR (Fig.S1), the temperature has identified as the most influential parameter with a value of mean |SHAP| 2.21, followed by catalyst dosage and reaction time, indicating thermodynamic control as primary yield driver. The SHAP beeswarm plot further validates that temperatures predominantly result in +ve SHAP value. Additionally, catalyst dosage also shows a positive effect, although with some variability, suggesting non-linear interactions. In contrast, reaction time exhibits a weaker and more scattered influence, indicating limited sensitivity within the range. For the response FDCA selectivity, the catalyst dosage emerges as the dominant factor with a mean |SHAP| value of 4.43, similar to ridge model. A similar pattern is exhibited by GBR model also as summarized in Fig.S2. In summary, all the model results show that temperature has the strongest influence on FDCA yield, while selectivity is more sensitive to how much catalyst is used. Reaction time, on the other hand, plays a smaller and less consistent role in both cases.
Fig. 7.
SHAP summary plot for ridge regression (a) and (b) FDCA yield and (c) and (d) FDCA selectivity.
The PDP for ridge model is illustrated in Fig. 8. An increase in FDCA yield in the range of investigated reaction temperatures, indicating the positive effect of temperature on biomass conversion and oxidation of intermediate compounds to FDCA. But the influence should be considered only in the studied region between 130 and 170 °C. Catalyst loading exhibited a higher impact at 0.8 g, which led to a constant contribution over 0.8 g. This can be attributed to the fact that there is an optimum amount of catalyst needed to provide enough active sites, and that a higher catalyst loading might not be beneficial for FDCA yield10. The effect of reaction time was lower compared to temperature and catalyst loading, suggesting that temperature and catalyst amount were the key parameters in improving the FDCA yield.
Fig. 8.
PDP for ridge regression (a) FDCA yield and (b) FDCA selectivity.
The PDP showed that selectivity was favored in the intermediate region around 5 h reaction time, 150 °C reaction temperature and 0.8 g catalyst loading. The lower estimated selectivity at higher temperature or excessive catalyst loading may be due to side reactions or over-oxidation and degradation of intermediates11. Therefore, SHAP and PDP analyses suggest that the improved condition should be chosen from a balanced region in the design space, rather than from the top of the design space. The PDP summary plots for SVR and GBR presented in Fig. S3 and S4.
From Fig. S3, the SVR results suggest that both reaction time and temperature tend to increase FDCA yield. The yield rises steadily as the temperature goes from 130 to 170 °C, and longer reaction times also help, though the effect is less strong. Catalyst loading behaves differently, it shows a curved trend, with the best yield around 0.8 g, after which the yield starts to drop. For FDCA selectivity, all three parameters show a clear peak, indicating that moderate conditions work best. The model points to optimal values of about 5 h, 150 °C, and 0.8 g. This means that while pushing temperature and time higher can improve yield, selectivity is maximized only within a narrower, balanced range of conditions. Similarly, the PDP summary plot for GBR model also provide further insight into the individual effects of reaction parameters on FDCA yield and selectivity. For yield, reaction time shows only a slight upward trend, suggesting it doesn’t play a major role within the range studied. Temperature, on the other hand, has a much stronger effect, with yield steadily increasing as it rises from 130 to 170 °C, highlighting its importance in driving the reaction. Catalyst loading follows a non-linear pattern: the yield improves up to about 0.8 g and then drops slightly, pointing to an optimal amount beyond which additional catalyst doesn’t help. A similar behaviour is seen for selectivity. Reaction time again has a limited impact, with a small improvement at mid-range values followed by a decline at longer durations. Temperature continues to have a positive influence, although its effect is not as strong as it is for yield. Catalyst loading stands out as a key factor, with selectivity peaking around 0.8 g before decreasing at higher amounts, possibly due to side reactions or transport limitations. Overall, the results suggest that temperature mainly drives FDCA yield, while selectivity depends more strongly on catalyst loading, with both showing clear optimal ranges and reaction time playing a less significant role.
Optimization of process parameters
The optimization of process parameters was also done for maximizing the FDCA yield and selectivity using the GA algorithm, and the results are summarized as follows in Table 5.
Table 5.
The optimized process parameters maximize the FDCA yield and selectivity.
| Model | Reaction time (h) | Reaction temperature (°C) | Catalyst dosage (g) | Predicted FDCA yield (%) | Predicted FDCA selectivity (%) |
|---|---|---|---|---|---|
| Ridge | 5.31 | 166.45 | 0.80 | 66.54 | 85.34 |
| SVR | 5.41 | 165.65 | 0.80 | 66.65 | 85.06 |
| GBR | 5.05 | 165.11 | 0.84 | 65.46 | 84.96 |
| Regression model from BBD | 5.36 | 166.8 | 0.80 | 66.69 | 85.13 |
The multi-objective optimization via-NSGA-Ⅱ results derived from different predictive models to achieve maximum FDCA yield and selectivity is summarized in Table 5. Among the machine learning approaches, the ridge regression model predicted the optimum conditions at a reaction time of 5.31 h, temperature of 166.45 °C, and for a catalyst dosage of 0.80 g, while the SVR model outperformed (66.65%) at a longer reaction time of 5.41 h and GBR offered fastest conversion at 5.05 h at a reduced selectivity of 84.96%. Ridge’s superior LOOCV R2 of 0.956 and interpretable linear coefficient of 2.23 for temperature, establish ridge as the primary model, with SVR confirming prediction robustness. This, robustness and consistency of these models in modelling chemical processes has also been reported in previous studies53. Whereas in BBD, the model suggested a relatively higher reaction time and temperature (5.36 h and 166.8 °C) with a similar catalyst dosage (0.80 g) to 66.69% and 85.13% of FDCA yield and selectivity respectively. The slightly elevated values in BBD could be attributed to its response surface methodology-based design, which tends to capture broader curvature effects in the experimental space, whereas the machine learning models provide refined predictions through data-driven pattern recognition.
Experimental validation
Based on the predicted outcomes and considering the practical feasibility of the one-pot synthesis of FDCA, the following parameter settings near the optimal setting suggested by various models were employed for experimental validation: catalyst dosage of 0.8 g, reaction temperature of 165 °C, and reaction times of 5.25 h and 5.5 h. A confirmatory experiment was carried out at the estimated optimal operating conditions to validate the predictions. With 5.25 h reaction time, the experimental FDCA yield and selectivity is observed as 66.36% and 86.02% respectively. While with 5.5 h, the test results indicate that the average yield of FDCA 67.54% and selectivity of 80.47% is closely aligned with the predicted result of 66.65% FDCA yield and 85.06% of selectivity from the model. This suggests that optimizing process parameters for maximizing FDCA yield through one-pot biomass conversion using a regressor model is both accurate and reliable. It also offers valuable insights for designing subsequent processes. In summary, this study proves that data driven models offer precise optimization pathways for efficient FDCA synthesis, thereby reducing energy input and maximizing production without compromising product yield and selectivity.
Moreover, the ML models achieved a high value R2, using the process parameters such as reaction time, temperature, and catalyst dosage to maximize the FDCA yield, without considering the catalysts textural properties such as acidity (2.97 mmol/g) or BET surface area (58.4 m²/g). This is mainly because the same catalyst was used in all experiments, so those properties never changed. As a result, the models aren’t really set up to generalize across different catalyst systems. In future studies, this could be addressed by including a range of catalyst formulations to broaden the input features.
Conclusion
In this study, the application of three regression models ridge regression, supported vector regressor (SVR), and gradient boosting regression GBR were analyzed in consideration of the one-pot conversion of biomass to FDCA with Fe-Mn zeolite catalyst. Each model gave insightful predictability for the responses viz. FDCA yield and selectivity with limited number of experimental data generated by box-Behnken design. The ridge regression outperformed among all the models and provided a valuable insight for evaluating the linear correlation between the variables and responses with an R2 value of 0.956 with a MSE value of 0.744. While SVR added multiple learners to achieve a predictive accuracy in FDCA yield with a R2 of 0.916 and MSE of 1.027. The GBR model defined evaluation matrices such as MSE of 1.711 with R2 0.768 for FDCA yield. The SHAP and PDP summary plot for the various models has been analyzed and temperature has been identified as the most prominent factor affecting both FDCA yield and selectivity. Additionally, the multi-objective optimization via-NSGA-Ⅱ of process parameters was done using GA algorithm and it is observed that the SVR model outperformed (66.65%) at a reaction time of 5.41 h and 166.65 °C temperature with a catalyst dosage of 0.80 g. In summary, the integration of experimental and AI-ML strategy provides a more efficient way to explore the feasibility of scale up of one-pot conversion of biomass to FDCA and can be applied to the sustainable production of value-added chemicals from renewable sources.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
The authors gratefully acknowledge the support of MHRD. (Govt. of India) extended to Amrita Vishwa Vidyapeetham through their FAST grant (F. No. 5-6/2013-TS.VII), DST (Govt. of India) extended to Amrita Vishwa Vidyapeetham, through their FIST grant (SR/FST/ETI- 416/2016) for the conduct of this research.
Author contributions
Fathima Safeeda N V : Conceptualization; Data curation; Formal analysis; Investigation; Methodology; Validation; Visualization; Roles/Writing—original draft. Thirumurugan G: Methodology; formal analysis; validation Suvvada Shankara Narayana Rao: Methodology, formal analysis, validation, Meera Balachandran : Conceptualization; Methodology; Resources; Supervision; and Writing—review & editing.
Funding
Open access funding provided by Amrita Vishwa Vidyapeetham.
Data availability
The datasets used and/or analyzed during the current study available from the corresponding author on reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Anandaram, H. et al. Co-pyrolysis characteristics and synergistic interaction of waste polyethylene terephthalate and woody biomass towards bio-oil production. J. Chem.2022, 1–9 (2022). [Google Scholar]
- 2.Sreehari, H., Gopika, V., Jayan, J. S., Sethulekshmi, A. S. & Saritha, A. A comprehensive review on bio epoxy based IPN: Synthesis, properties and applications. Polymer252, 124950 (2022). [Google Scholar]
- 3.Zhang, Z. & Deng, K. Recent advances in the catalytic synthesis of 2,5-furandicarboxylic acid and its derivatives. ACS Catal.5, 6529–6544 (2015). [Google Scholar]
- 4.Shao, Y., Ding, Y., Dai, J., Long, Y. & Hu, Z. T. Synthesis of 5-hydroxymethylfurfural from dehydration of biomass-derived glucose and fructose using supported metal catalysts. Green Synth Catal2, 187–197 (2021). [Google Scholar]
- 5.Ban, H., Cheng, Y., Wang, L. & Li, X. One-pot method for the synthesis of 2,5-furandicarboxylic acid from fructose: In situ oxidation of 5-hydroxymethylfurfural and 5-acetoxymethylfurfural over Co/Mn/Br catalysts in acetic acid. Ind. Eng. Chem. Res.62, 291–301 (2023). [Google Scholar]
- 6.Fathima Safeeda, N. V. et al. Synergistic Cr–Co metal oxide catalysis on molecular sieve support for cascade one-pot conversion of biomass derived cellulose to FDCA. Sustain. Chem. Pharm.45, 102040 (2025). [Google Scholar]
- 7.Kröger, M., Prüße, U. & Vorlop, K. [No title found]. Top. Catal.13, 237–242 (2000). [Google Scholar]
- 8.Chai, Y. et al. Cr-Mn bimetallic functionalized USY zeolite monolithic catalyst for direct production of 2, 5-furandicarboxylic acid from raw biomass. Chem. Eng. J.429, 132173 (2022). [Google Scholar]
- 9.Fathima Safeeda, N. V., Balachandran, M. & Ragula, U. B. R. Engineering a heterogeneous catalyst with earth-abundant metal oxides for efficient one-pot synthesis of 2,5-furan dicarboxylic acid from agro-waste. J. Environ. Chem. Eng.13, 115653 (2025). [Google Scholar]
- 10.Chai, Y. et al. Direct production of 2, 5-Furandicarboxylicacid from raw biomass by manganese dioxide catalysis cooperated with ultrasonic-assisted diluted acid pretreatment. Bioresour. Technol.337, 125421 (2021). [DOI] [PubMed] [Google Scholar]
- 11.Fathima Safeeda, N. V. & Balachandran, M. Implementation of RSM and ANN optimization approach for the valorization of agro-wastes to renewable plastic precursor FDCA over non-precious bimetal oxide functionalized heterogeneous catalyst. J. Environ. Manage.380, 125026 (2025). [DOI] [PubMed] [Google Scholar]
- 12.Balachandran, M., Devanathan, S., Muraleekrishnan, R. & Bhagawan, S. S. Optimizing properties of nanoclay–nitrile rubber (NBR) composites using face centred central composite design. Mater. Des.35, 854–862 (2012). [Google Scholar]
- 13.Shankar, K. V. et al. Influence of T6 heat treatment analysis on the tribological behaviour of cast Al-12.2Si-0.3 Mg-0.2Sr alloy using response surface methodology. J. Bio. Tribo. Corros.7, 96 (2021). [Google Scholar]
- 14.Bañon, F., Martin, S., Vazquez-Martinez, J. M., Salguero, J. & Trujillo, F. J. Predictive models based on RSM and ANN for roughness and wettability achieved by laser texturing of S275 carbon steel alloy. Opt. Laser Technol.168, 109963 (2024). [Google Scholar]
- 15.Sudarshan, V. & Seider, W. D. Advancing machine learning in Industry 4.0: Benchmark framework for rare-event prediction in chemical processes. Comput. Chem. Eng.194, 108929 (2025). [Google Scholar]
- 16.Nussbaum, M. Machine learning and processing of large data. In Encyclopedia of Soils in the Environment 509–520 (Elsevier, 2023). 10.1016/B978-0-12-822974-3.00065-3. [Google Scholar]
- 17.Yohannes, A. G. et al. Combined high-throughput DFT and ML screening of transition metal nitrides for electrochemical CO2 reduction. ACS Catal.13, 9007–9017 (2023). [Google Scholar]
- 18.Estahbanati, M. R. K., Feilizadeh, M. & Iliuta, M. C. Photocatalytic valorization of glycerol to hydrogen: Optimization of operating parameters by artificial neural network. Appl. Catal. B Environ.209, 483–492 (2017). [Google Scholar]
- 19.Jenewein, K. J. et al. Navigating the unknown with AI: Multiobjective Bayesian optimization of non-noble acidic OER catalysts. J. Mater. Chem. A12, 3072–3083 (2024). [Google Scholar]
- 20.Shambhawi, Mohan, O., Choksi, T. S. & Lapkin, A. A. The design and optimization of heterogeneous catalysts using computational methods. Catal. Sci. Technol.14, 515–532 (2024). [Google Scholar]
- 21.Suzuki, K. et al. Statistical analysis and discovery of heterogeneous catalysts based on machine learning from diverse published data. ChemCatChem11, 4537–4547 (2019). [Google Scholar]
- 22.Noh, J., Back, S., Kim, J. & Jung, Y. Active learning with non- ab initio input features toward efficient CO2 reduction catalysts. Chem. Sci.9, 5152–5159 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ross-Veitía, B. D. et al. Machine learning regression algorithms to predict emissions from steam boilers. Heliyon10, e26892 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Sultana, N. et al. Modeling and optimization of non-edible papaya seed waste oil synthesis using data mining approaches. S. Afr. J. Chem. Eng.33, 151–159 (2020). [Google Scholar]
- 25.Li, M. et al. High-efficiency catalyst screening driven by machine learning for the oxidation of 5-Hydroxymethylfurfural to 2, 5-Furandicarboxylic Acid. ACS Sustain. Chem. Eng.13, 13940–13952 (2025). [Google Scholar]
- 26.Pandey, S., Mottoul, M., Orsat, V., Morin, J.-F. & Dumont, M.-J. Base-free synthesis of renewable furan-2,5-dicarboxylic acid (FDCA) over an earth-abundant magnetic catalyst. J. Environ. Chem. Eng.12, 112763 (2024). [Google Scholar]
- 27.Weese, M. L., Smucker, B. J. & Edwards, D. J. The use of cross validation in the analysis of designed experiments. Preprint at.10.48550/ARXIV.2506.14593 (2025). [Google Scholar]
- 28.Austin, G. I., Pe’er, I. & Korem, T. Distributional bias compromises leave-one-out cross-validation. Sci. Adv.11, eadx6976 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Vanjari, P., Kamesh, R. & Y. Rani, K. Machine learning models representing catalytic activity for direct catalytic CO2 hydrogenation to methanol. Mater. Today: Proc.72, 524–532 (2023). [Google Scholar]
- 30.Hoerl, A. E. & Kennard, R. W. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics12, 55–67 (1970). [Google Scholar]
- 31.Chen, T., Guestrin, C. & XGBoost: A Scalable Tree Boosting System. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794ACM, San Francisco California USA, (2016). 10.1145/2939672.2939785
- 32.Ologunagba, D. & Kattel, S. Machine learning prediction of surface segregation energies on low index bimetallic surfaces. Energies13, 2182 (2020). [Google Scholar]
- 33.Sun, D. et al. SHAP-PDP hybrid interpretation of decision-making mechanism of machine learning-based landslide susceptibility mapping: A case study at Wushan District, China. Egypt. J. Remote Sens. Space Sci.27, 508–523 (2024). [Google Scholar]
- 34.Friedman, J. H. Greedy function approximation: A gradient boosting machine. Ann Statist29, (2001).
- 35.Petch, J., Di, S. & Nelson, W. Opening the Black Box: The Promise and Limitations of Explainable Machine Learning in Cardiology. Can. J. Cardiol.38, 204–213 (2022). [DOI] [PubMed] [Google Scholar]
- 36.Zheng, X. et al. Adsorption Evaluation with Zeolites for Gases in the Confined Space. IOP Conf. Ser. : Earth Environ. Sci.453, 012086 (2020). [Google Scholar]
- 37.Khalil, A. K. A. et al. Preparation of iron doped zeolite-coated porous clay ceramic membrane (Fe/ZSM − 5) for heavy metal filtration: Electrochemical study of the rejection mechanism. J. Hazard. Mater. Adv.18, 100720 (2025). [Google Scholar]
- 38.Liang, T., Wang, B., Fan, Z. & Liu, Q. A facile fabrication of superhydrophobic and superoleophilic adsorption material 5A zeolite for oil–water separation with potential use in floating oil. Open Phys.19, 486–493 (2021). [Google Scholar]
- 39.Qayoom, M., Shah, K. A., Pandit, A. H., Firdous, A. & Dar, G. N. Dielectric and electrical studies on iron oxide (α-Fe2O3) nanoparticles synthesized by modified solution combustion reaction for microwave applications. J. Electroceram.45, 7–14 (2020). [Google Scholar]
- 40.Atique Ullah, A. K. M. et al. Oxidative degradation of methylene blue using Mn3O4 nanoparticles. Water Conserv. Sci. Eng.1, 249–256 (2017). [Google Scholar]
- 41.Sadeghi, M., Farhadi, S. & Zabardasti, A. Construction of magnetic MgFe2O4/CdS/MoS2 ternary nanocomposite supported on NaY zeolite and highly efficient sonocatalytic degradation of organic pollutants. RSC Adv.10, 44034–44049 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Xu, S. et al. Highly efficient Cr/β zeolite catalyst for conversion of carbohydrates into 5‑hydroxymethylfurfural: Characterization and performance. Fuel Process. Technol.190, 38–46 (2019). [Google Scholar]
- 43.Yue, Y. et al. Direct synthesis of hierarchical FeCu-ZSM‐5 zeolite with wide temperature window in selective catalytic reduction of NO by NH3. ChemCatChem11, 4744–4754 (2019). [Google Scholar]
- 44.Liu, X., Gao, S., Yang, F., Zhou, S. & Kong, Y. High promoting of selective oxidation of ethylbenzene by Mn-ZSM-5 synthesized without organic template and calcination. Res. Chem. Intermed.46, 2817–2832 (2020). [Google Scholar]
- 45.Liu, H. et al. Efficient aerobic oxidation of 5-hydroxymethylfurfural to 2,5-diformylfuran over Fe2O3-promoted MnO2 catalyst. ACS Sustain. Chem. Eng.7, 7812–7822 (2019). [Google Scholar]
- 46.Gawade, A. B., Nakhate, A. V. & Yadav, G. D. Selective synthesis of 2, 5-furandicarboxylic acid by oxidation of 5-hydroxymethylfurfural over MnFe 2 O 4 catalyst. Catal. Today. 309, 119–125 (2018). [Google Scholar]
- 47.Mondal, S., Ruidas, S., Chongdar, S., Saha, B. & Bhaumik, A. Sustainable porous heterogeneous catalysts for the conversion of biomass into renewable energy products. ACS Sustainable Resour. Manage.1, 1672–1704 (2024). [Google Scholar]
- 48.Dandekar, P., Ambesh, A. S., Khan, T. S. & Gupta, S. Machine learning assisted approximation of descriptors (CO and OH) binding energy on Cu-based bimetallic alloys. Phys. Chem. Chem. Phys.27, 7151–7168 (2025). [DOI] [PubMed] [Google Scholar]
- 49.Šinkovec, H., Heinze, G., Blagus, R. & Geroldinger, A. To tune or not to tune, a case study of ridge logistic regression in small or sparse datasets. Preprint at.10.48550/arXiv.2101.11230 (2021). [DOI] [PMC free article] [PubMed]
- 50.Sheikhmohammadi, A., Khakzad, P., Rasolevandi, T. & Azarpira, H. Leveraging artificial intelligence models (GBR, SVR, and GA) for efficient chromium reduction via UV/trichlorophenol/sulfite reaction. Results Eng.26, 104599 (2025). [Google Scholar]
- 51.Wadaugsorn, K., Lin, K.-Y., Kaewchada, A. & Jaree, A. Production of 2,5-furandicarboxylic acid via oxidation of 5-hydroxymethylfurfural over Pt/C in a continuous packed bed reactor. RSC Adv.12, 18084–18092 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Xing, L., Xu, M., Liu, R., Jing, F. & He, J. Synergistic effect of the fourth metal component on Co/Mn/Br catalyst in 2, 5-furanediformic acid preparation. Catal. Commun.171, 106505 (2022). [Google Scholar]
- 53.Rezaei, I., Amirshahi, S. H. & Mahbadi, A. A. Utilizing support vector and kernel ridge regression methods in spectral reconstruction. Results Opt.11, 100405 (2023). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets used and/or analyzed during the current study available from the corresponding author on reasonable request.















