Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Jan 29;16:6767. doi: 10.1038/s41598-026-37751-0

Optimization of the extraction process of Sanhuang Qingre Formula by integrating response surface methodology, grey correlation analysis, and machine learning

Qisong Chen 1, Pan Meng 1, Xinyue Hu 1, Yang Liu 2, Jiaozhang Tan 2, Zhirong Peng 2,✉, Jianye Yan 1,3,✉
PMCID: PMC12913784  PMID: 41611957

Abstract

This study aimed to optimize the ethanol reflux extraction process of the Sanhuang Qingre Formula (SHQRF) using response surface methodology (RSM), grey correlation analysis (GCA), and machine learning. A combination of single-factor experiments and a Box-Behnken design (BBD) was applied to optimize three key variables: ethanol concentration, reflux time, and liquid-solid ratio. During the experiments, the comprehensive score—calculated from the contents of 11 bioactive components (coptisine, epiberberine, berberine hydrochloride, palmatine, baicalin, chrysin-7-O-β-glucuronide, wogonoside, baicalein, wogonin, oroxylin A, atractylodin) and extraction yield—was used as the evaluation index. The combined weighting method of the fuzzy analytic hierarchy process and entropy weight method was employed to determine the comprehensive score. Results showed that the process optimized by RSM and the support vector machine (SVM) model achieved a higher comprehensive score of 56.08 compared with that of 52.67 obtained by GCA optimization. Consequently, the optimal extraction parameters for SHQRF were determined as follows: ethanol concentration of 55%, reflux time of 2 h per cycle, and a liquid-solid ratio of 12 mL/g. The optimized process significantly increased the individual and total contents of the 11 target components compared with the original process (the pre-optimization process). Furthermore, the original process, the GCA-optimized process, and the optimal processes derived from RSM and SVM were distinctly classified by hierarchical cluster analysis (HCA) and principal component analysis (PCA). This study provides an effective strategy to enhance the extraction efficiency and quality of SHQRF and offers new insights and approaches for optimizing the extraction of active components in traditional Chinese medicine.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-37751-0.

Keywords: Sanhuang qingre formula, Reflux extraction, Box-Behnken design, Response surface methodology, Grey correlation analysis, Machine learning

Subject terms: Chemical biology, Drug discovery, Plant sciences, Chemistry

Introduction

Sanhuang Qingre Formula (SHQRF) is a clinical prescription developed by Professor Wang Dahai of the First Hospital Affiliated with Hunan Traditional Chinese Medical College, based on extensive departmental research and clinical experience. The formula comprises Coptidis Rhizoma, Scutellariae Radix, Astragali Radix, Atractylodis Rhizoma, Magnoliae Flos, Poria, Spirodelae Herba, and Glycyrrhizae Radix Et Rhizoma. SHQRF exerts effects of consolidating vitality, promoting tissue regeneration, clearing heat, and resolving dampness, and is clinically applied in the treatment of chronic purulent sinusitis and allergic sinusitis. Since 1998, SHQRF has been formulated as Sanhuang Qingre Nasal Drops, which were approved by the Zhuzhou Health Bureau as a hospital Class II preparation (Approval No. Xiangyao Zhizi Z20070343). Although the nasal drops have demonstrated notable clinical efficacy, limitations such as short drug retention time and poor stability have restricted their wider clinical application. To overcome these shortcomings, it is necessary to develop Sanhuang Qingre Nasal Drops into alternative dosage forms. The extraction process plays a crucial role in the preparation of traditional Chinese medicine (TCM), as it directly affects pharmaceutical quality and therapeutic consistency. Therefore, optimizing the extraction process of SHQRF is essential for improving its quality and facilitating subsequent formulation development.

In SHQRF, the total alkaloids derived from Coptidis Rhizoma exert significant anti-inflammatory effects1, with coptisine, epiberberine, berberine hydrochloride, and palmatine identified as the principal alkaloids2. Modern pharmacological studies have demonstrated that berberine hydrochloride possesses potent antibacterial activity3, inhibiting bacterial DNA synthesis by reducing DNA topoisomerase activity4,5. Baicalin, a major constituent of Scutellariae Radix, exhibits strong antioxidant properties by directly scavenging oxygen-free radicals and inhibiting xanthine oxidase6. Baicalein has been reported to display antibacterial activity against methicillin-resistant Staphylococcus aureus7. Other bioactive compounds in Scutellariae Radix, including wogonoside8–10, wogonin11,12, and chrysin-7-O-β-glucuronide13, have demonstrated confirmed anti-inflammatory effects, while oroxylin A shows antiviral activity14. Atractylodin, the main active component of Atractylodis Rhizoma, alleviates airway inflammation by suppressing IL-8 and MUC5AC expression and inhibiting NF-κB pathway activation15,16. Due to the complex active components in SHQRF, which may respond differently to extraction variables, relying solely on a single indicator for evaluating extraction efficiency could lead to biased or incomplete assessment results. Therefore, in this study, the contents of the above 11 index components (Fig. 1) together with the extraction yield were selected as evaluation indicators to assess the extraction process of SHQRF.

Fig. 1.

Fig. 1

Structures of 11 active components in SHQRF.

In addition, it is necessary to obtain a comprehensive score by assigning appropriate weights to each indicator during multi-index evaluation. The fuzzy analytic hierarchy process (FAHP) and entropy weight method (EWM) are widely applied in the comprehensive evaluation of multiple indicators17–19. FAHP subjectively assigns weights based on the relative importance of evaluation indicators17, while EWM, grounded in information entropy theory, objectively reflects the variability among indicators18,19. When subjective and objective weighting methods are combined, the resulting weighting coefficients become more scientific, rational, and systematic20,21.

The Box-Behnken design-response surface methodology (BBD-RSM) is widely used in food and pharmaceutical processing due to its advantages of rapidity, simplicity, and reliability22,23. By analyzing the effects of multiple variables on a response variable and exploring their interactions, BBD-RSM provides a rational experimental design for arranging and combining influencing factors, thereby obtaining optimal predictive results21,24. However, as the complexity of extraction processes increases, the limitations of traditional RSM in addressing high-dimensional nonlinear problems have become increasingly evident. Grey correlation analysis (GCA), a fuzzy system analysis method, evaluates the influence of each factor on a system by measuring the correlation degree between different factor sequences based on the similarity or difference in their developmental trends25–28. Meanwhile, machine learning techniques have been increasingly applied in the field of TCM for their strong capability in nonlinear modeling and adaptive optimization29–32. The backpropagation neural network (BPNN), inspired by the structure of the human brain, utilizes a backpropagation algorithm for training. When combined with a genetic algorithm (GA) to optimize network weights and thresholds, the GA-BPNN model demonstrates excellent performance in solving complex nonlinear relationships33–36. Similarly, the support vector machine (SVM), a supervised learning algorithm, performs effectively in complex nonlinear regression with small sample sizes by mapping data into high-dimensional spaces using kernel functions to construct nonlinear regression models37–39. By integrating these approaches, data complexity and uncertainty can be better managed, thereby enhancing the comprehensiveness and accuracy of the extraction process optimization studies.

In this study, BBD-RSM and GCA were employed to optimize the extraction process of SHQRF, using the contents of 11 bioactive components and extraction yield as evaluation indices. A comprehensive score was calculated through the FAHP-EWM to balance the subjective importance and objective variability of the indicators. In parallel, two machine learning models, GA-BPNN and SVM, were developed to identify the superior algorithm for process optimization by comparing model performance parameters. Multivariate statistical analyses were then applied to verify the consistency and robustness of the optimization results across different methodologies. To the best of our knowledge, this is the first study to integrate RSM, GCA, and machine learning (GA-BPNN and SVM) for the multi-index optimization of SHQRF extraction. This hybrid approach addresses the limitations of traditional single-method optimization. The findings not only identify the optimal extraction process but also provide a theoretical basis for future formulation development.

Materials and methods

Materials and chemicals

Plant materials

Coptidis Rhizoma (batch no. 231003), Scutellariae Radix (batch no. 240301), Astragali Radix (batch no. 240202), Atractylodis Rhizoma (batch no. 220701), Magnoliae Flos (batch no. 231201), Poria (batch no. 231205), Spirodelae Herba (batch no. 230302), and Glycyrrhizae Radix Et Rhizoma (batch no. 240104) were purchased from the Nanguo Yaodu TCM Decoction Pieces Co., Ltd. (Shaoyang, China). All materials were authenticated by Associate Professor Limin Gong, Hunan University of Chinese Medicine, China, based on morphological and microscopic characteristics, in accordance with the Chinese Pharmacopoeia (2025 edition).

Chemicals and reagents

Coptisine (batch no. HR8236W16, CAS: 3486-66-6), epiberberine (batch no. HR2731W7, CAS: 6873-09-2), berberine hydrochloride (batch no. HR1571S1, CAS: 633-65-8), palmatine (batch no. HC034213, CAS: 3486-67-7), baicalin (batch no. HR2182W7, CAS: 21967-41-9), chrysin-7-O-β-glucuronide (batch no. HR2182W7, CAS: 35775-49-6), wogonoside (batch no. HR15125B1, CAS: 51059-44-0), baicalein (batch no. HR16319S1, CAS:491-67-8), wogonin (batch no. HR21029B1, CAS:632-85-9), oroxylin A (batch no. HS21102B1, CAS:480-11-5), and atractylodin (batch no. HR3333W16, CAS:55290-63-6) were purchased from Chenguang Biotech Co., Ltd (Baoji, China). The purity of all reference standards was ≥ 98%. Anhydrous ethanol was purchased from China National Pharmaceutical Group Chemical Reagent Co., Ltd. (Shanghai, China). Phosphoric acid was obtained from Kemio Chemical Reagent Co., Ltd. (Tianjin, China). Chromatographic-grade methanol and acetonitrile were supplied by Thermo Fisher Scientific Inc. (Waltham, MA, USA). Deionized water (resistivity ≥ 18.2 MΩ·cm) was prepared in-house using an Option R7 ultra AN water purification system (ELGA LabWater, High Wycombe, UK).

Preparation of test solution

The SHQRF was composed of Coptidis Rhizoma (5 g), Scutellariae Radix (15 g), Astragali Radix (5 g), Atractylodis Rhizoma (2.5 g), Magnoliae Flos (15 g), Poria (5 g), Spirodelae Herba (7.5 g), and Glycyrrhizae Radix et Rhizoma (2.5 g). The mixture was extracted twice by reflux using 10 times (v/w) the volume of 60% ethanol (575 mL per cycle), with each extraction lasting 1 h. The filtrates from both extractions were combined and concentrated to a final volume of 1000 mL. After thorough mixing, 2.5 mL of the concentrated extract was transferred into a 10 mL volumetric flask and diluted to volume with ethanol. The resulting solution was filtered through a 0.22 μm microporous membrane to obtain the test solution.

The original extraction process corresponded to the method used in the preparation of Sanhuang Qingre Nasal Drops. The same procedure was followed as described above, except that water was used as the extraction solvent instead of ethanol.

Establishment of HPLC methods

Chromatographic conditions

HPLC analysis was performed using a Waters ACQUITY HPLC system (Waters Corporation, Milford, MA, USA) equipped with a Symmetry® C18 column (250 mm × 4.6 mm, 5 μm). The mobile phase consisted of (A) aqueous phosphoric acid (0.1%, v/v) and (B) acetonitrile, with gradient elution as follows (0–85 min): 0–25 min, 85%A; 25–40 min, 85%-81%A; 40–45 min, 81%-75%A; 45–60 min, 75%-56%A; 60–70 min, 56%-10%A; 70–75 min, 10%-5%A; 75–85 min, 5%-85%A. The flow rate was maintained at 0.8 mL/min, and the column temperature was set at 35 °C. The injection volume was 2 µL, and detection was performed at a wavelength of 348 nm.

Preparation of standard solutions

Appropriate amounts of each reference standard were accurately weighed and dissolved in 60% ethanol in a 10 mL volumetric flask. The resulting stock solutions contained the following mass concentrations: coptisine (1.892 mg/mL), epiberberine (1.849 mg/mL), berberine hydrochloride (1.700 mg/mL), palmatine (1.826 mg/mL), baicalin (1.919 mg/mL), chrysin-7-O-β-glucuronide (1.190 mg/mL), wogonoside (2.434 mg/mL), baicalein (1.060 mg/mL), wogonin (1.046 mg/mL), oroxylin A (1.696 mg/mL), and atractylodin (1.761 mg/mL). The mixed standard solution was subsequently diluted with 60% ethanol to prepare a series of working solutions at different concentrations for establishing calibration curves and assessing linearity.

Analytical method validation

Method validation was conducted in terms of precision, stability, repeatability, and recovery. Precision was evaluated by consecutively injecting the same sample six times under identical chromatographic conditions. Stability was assessed by analyzing the same sample at time intervals of 0, 2, 4, 6, 12, and 24 h. Repeatability was determined by preparing and analyzing six independent SHQRF sample solutions in parallel. Recovery was assessed by adding accurately known amounts of each of the eleven reference standards to pre-analyzed samples, followed by analysis to calculate the recovery rates.

Single-factor experiment

To optimize the extraction efficiency of SHQRF, three factors—ethanol concentration, liquid-solid ratio, and reflux time—were investigated through single-factor experiments. When studying the effect of ethanol concentration, the liquid-solid ratio was fixed at 10 mL/g and the reflux time at 1 h, while the ethanol concentration was set at five levels (40%, 50%, 60%, 70%, and 80%). For the liquid-solid ratio experiments, the ethanol concentration was fixed at 60% and the reflux time at 1 h, with the liquid-solid ratio set at 6, 8, 10, 12, and 14 mL/g. In the investigation of reflux time, the ethanol concentration was fixed at 60% and the liquid-solid ratio at 10 mL/g, while the reflux time was varied at 0.5, 1.0, 1.5, 2.0, and 2.5 h.

Calculation of comprehensive weights

FAHP is a multi-criteria evaluation method that integrates hierarchical analysis with fuzzy comprehensive evaluation. A fuzzy judgment matrix was constructed based on the compatible regularity and the contributions of the monarch, minister, assistant, and guide. The fuzzy complementary judgment matrix—characterized by reduced computational complexity and improved practicality17—was employed to derive the subjective weight coefficients (wjp) for each indicator. The consistency ratio (CR) was subsequently calculated to verify matrix consistency. EWM quantifies indicator weights according to the amount of information entropy contained in the data. The entropy value reflects the degree of system complexity: the smaller the entropy, the greater the variability among indicators and, consequently, the higher the assigned weight40. Data were first normalized and analyzed using IBM SPSS Statistics version 25.0 (SPSS Inc., Chicago, IL, USA) to generate the raw data matrix, which was then converted into a probability matrix for calculating the objective weight coefficients (wjq). Finally, the multiplication synthesis normalization method was applied to integrate the subjective and objective weights, yielding the comprehensive weights (wjs) of the evaluation indicators, calculated according to Eq. (1).

graphic file with name d33e430.gif 1

Optimization of extraction conditions

Experimental design and statistical analysis of RSM

Based on the results of the single-factor experiments, three key variables—ethanol concentration (50–70%), reflux time (1–2 h), and liquid-solid ratio (8–12 mL/g)—were selected as influencing factors. Each factor was coded at three levels: low (-1), medium (0), and high (+ 1). A Box-Behnken design (BBD) comprising three factors and three levels was implemented using Design-Expert software (Stat-Ease Inc., Minneapolis, MN, USA). The experimental design included a total of 17 runs, consisting of 12 factorial points and 5 center points.

GCA

The gray pattern recognition dataset was established using the contents of 11 index components and the extraction yield as evaluation indicators. Due to the inconsistency of measurement units among the evaluation indicators, the data were normalized according to Eq. (2) to achieve dimensionless comparability.

graphic file with name d33e449.gif 2

Where Yik was the normalized data, Xik was the original data, and Xk was the mean of the k th index of samples (k = 1, 2, …, 12). The optimal reference sequence {Xsk} and the worst reference sequence {Xtk} were the maximum and minimum values of each indicator among the 17 samples.

Firstly, the correlation coefficients of each evaluation unit with respect to both the optimal and worst reference sequences were calculated using Eq. (3).

graphic file with name d33e473.gif 3

Where Inline graphic denotes the resolution coefficient (commonly set to 0.5), ∆´min is the minimum absolute difference between the reference and comparison sequences for each indicator, and ∆´max is the maximum absolute difference. Yk represents the reference sequence.

Secondly, Eq. (4) was applied to calculate the correlation degree of each evaluation unit. The relative correlation degree (rst) was calculated using Eq. (5), and the results were ranked according to their magnitudes to identify the optimal combination of process parameters.

graphic file with name d33e498.gif 4
graphic file with name d33e505.gif 5

Where ri(s) represents the correlation degree relative to the optimal reference sequence, ri(t) corresponds to the worst reference sequence, and rst denotes the relative correlation of the evaluation unit sequence.

Building machine learning models

To optimize the extraction parameters of SHQRF, the experimental data obtained from the RSM experiments were used to construct a BPNN model optimized by a GA. The model architecture consisted of an input layer with three neurons (ethanol concentration, reflux time, and liquid-solid ratio), a hidden layer containing seven neurons, and an output layer with one neuron representing the comprehensive score. Model construction and training were performed using MATLAB software. The core principle of the SVM algorithm is to identify an optimal regression hyperplane that encompasses all training samples with minimal error. The quantitative relationship between the extraction parameters and the comprehensive score was analyzed, and a nonlinear regression model of the extraction process was established in a high-dimensional feature space.

Validation experiment and comparison between different models

The extraction process was performed according to the optimal parameters obtained from the GCA, RSM, and SVM models. Three parallel experiments were conducted for each optimized condition to determine the contents of the 11 bioactive components and the extraction yields. The comprehensive scores were then calculated to evaluate and compare model performance.

Comparison of extraction processes before and after optimization

Test solutions were prepared following both the optimized extraction processes and the original extraction process. The contents of the 11 index components were determined for comparative analysis. Additionally, the evaluation indicators obtained from each optimized model and the original extraction process were used as variables and imported into TBtools software (CJ-Chen, China) for HCA. PCA was further conducted using SIMCA software (Umetrics, Umeå, Sweden).

Results

Method validation of HPLC

Representative HPLC chromatograms of the mixed standard solutions and SHQRF samples are shown in Fig. 2. As illustrated in Fig. 2, the chromatographic peaks of the standard solution corresponded precisely to those in the test solution, and no interfering peaks were observed in the negative control, demonstrating good method specificity. The limits of detection (LOD) and limits of quantification (LOQ) were determined at signal-to-noise (S/N) ratios of 3:1 and 10:1, respectively. The results for linear relationships, linear ranges, LOD, LOQ, precision, stability, and repeatability are summarized in Table 1. Each analyte exhibited excellent linearity within its respective concentration range, with correlation coefficients (R2) exceeding 0.9990. The relative standard deviations (RSDs) for precision, stability, and repeatability were all below 3.0%, confirming the high precision of the instrument, the repeatability of the analytical method, and the stability of the samples over 24 h. The average recovery rates ranged from 96.86% to 103.33%, with RSDs below 3.0%, indicating that the developed HPLC method is highly reliable and accurate for the quantitative determination of the 11 index components in SHQRF.

Fig. 2.

Fig. 2

HPLC chromatograms of SHQRF. (A) Blank solution; (B) Standard solution; (C) SHQRF test solution; (D) Without Coptidis Rhizoma negative test solution; (E) Without Scutellariae Radix negative test solution; (F) Without Atractylodis Rhizoma negative test solution. (1) coptisine; (2) epiberberine; (3) berberine hydrochloride; (4) palmatine; (5) baicalin; (6) chrysin-7-O-β-glucuronide; (7) wogonoside; (8) baicalein; (9) wogonin; (10) oroxylin A; 11. atractylodin.

Table 1.

Linear relation, linear range, LOD, LOQ, precision, stability, repeatability and sample recovery results of 11 components.

Name Linear relation Linear range
(µg/mL)
LOQ
(µg/mL)
LOD
(µg/mL)
RSD% Sample recovery
(RSD%)
Precision stability repeatability
Coptisine

Y=9535X-26,648

(R2 = 0.9991)

2.84–114 1.4 0.43 1.3 2.1 1.0

96.9

(1.3)

Epiberberine

Y = 8428.1X-26,001

(R2 = 0.9990)

2.77–111 1.4 0.42 2.5 2.9 2.8

99.7

(2.6)

Berberine hydrochloride

Y = 9821.5X-28,674

(R2 = 0.9995)

8.50–340 0.76 0.26 2.6 2.6 2.5

102

(2.6)

Palmatine

Y = 8616.2X-10,831

(R2 = 0.9996)

2.74–110 0.91 0.27 2.5 1.2 1.5

101

(2.6)

Baicalin

Y = 1656.4X + 1691

(R2 = 0.9998)

48.0-1919 1.4 0.41 0.57 1.2 1.5

97.5

(1.4)

Chrysin-7-O-β-glucuronide

Y=2344X + 5071.2

(R2 = 0.9993)

4.76–190 0.95 0.30 0.52 1.2 1.2

99.4

(2.5)

Wogonoside

Y = 2188.3X-6057.8

(R2 = 0.9994)

9.74–389 0.97 0.25 0.59 1.3 1.7

103

(1.0)

Baicalein

Y = 4128.2X-6704.2

(R2 = 0.9996)

3.18–127 0.80 0.22 0.73 1.7 2.5

101

(2.5)

Wogonin

Y = 3317.5X + 412.9

(R2 = 0.9998)

1.05–41.8 0.52 0.16 0.94 1.1 2.6

102

(1.8)

Oroxylin A

Y=3787X-1045

(R2 = 0.9993)

0.848–33.9 0.42 0.13 1.9 1.2 2.6

102

(2.8)

Atractylodin

Y=16011X + 2238.3

(R2 = 0.9991)

0.880–35.2 0.32 0.091 0.96 1.1 2.2

97.8

(1.6)

Single-factor experiments

The results of the single-factor experiments provided a foundation for the subsequent Box-Behnken experimental design, as illustrated in Fig. 3. As shown in Fig. 3A and D, the total contents of the indicator components and the extraction yield reached their respective peaks at ethanol concentrations of 60% and 50%. Therefore, ethanol concentrations of 50%, 60%, and 70% were selected for further investigation.

Fig. 3.

Fig. 3

Single-factor experimental results of SHQRF extraction process. (A–C) are effect of factors on the content of active components. (D–F) are effect of factors on the extraction yield.

The effects of reflux time on the indicator contents and extraction yield are presented in Fig. 3B and E. The total content of the indicators exhibited an increasing trend with prolonged reflux time, with no significant difference observed between 1.5 h and 2.0 h. The extraction yield increased initially and then declined, reaching its maximum at 2.0 h. Accordingly, reflux times of 1.0 h, 1.5 h, and 2.0 h were chosen for subsequent experiments.

The influence of the liquid-solid ratio on the indicator contents and extraction yield is shown in Fig. 3C and F. Both the total content of the indicators and the extraction yield increased initially and then decreased with an increasing liquid-solid ratio. The maximum values were observed at 10 mL/g and 12 mL/g, respectively. Hence, liquid-solid ratios of 8 mL/g, 10 mL/g, and 12 mL/g were selected for further analysis.

Calculation of comprehensive weights

FAHP weight

According to the pharmacological activity and relative importance of each indicator in the experiment, the priority order of the indicators was determined as follows: berberine hydrochloride = baicalin > coptisine = epiberberine = palmatine > chrysin-7-O-β-glucuronide = wogonoside = baicalein = wogonin > oroxylin A > atractylodin > extraction yield. According to the criteria for constructing a fuzzy judgment matrix using the FAHP method41, the corresponding judgment matrix is presented in Supplementary Table S1. The relative importance among indicators at the same hierarchical level was scored, and the resulting feature matrix is shown in Supplementary Table S2. The calculated weight coefficients for berberine hydrochloride, baicalin, coptisine, epiberberine, palmatine, chrysin-7-O-β-glucuronide, wogonoside, baicalein, wogonin, oroxylin A, atractylodin, and extraction yield were 9.848%, 9.848%, 8.939%, 8.939%, 8.939%, 8.030%, 8.030%, 8.030%, 8.030%, 8.030%, 7.121%, and 6.212%, respectively. The CR was calculated as 0.08, which is below the acceptable threshold of 0.1, indicating that the consistency test was successfully passed and the constructed judgment matrix was reliable.

EWM weight

The weighting coefficients of coptisine, epiberberine, berberine hydrochloride, palmatine, baicalin, chrysin-7-O-β-glucuronide, wogonoside, baicalein, wogonin, oroxylin A, atractylodin, and extraction yield were calculated as 11.92%, 6.18%, 8.46%, 11.35%, 6.60%, 5.61%, 5.57%, 6.16%, 8.74%, 9.68%, 9.87%, and 9.80%, respectively.

FAHP-EWM hybrid weighting method

The FAHP method determines subjective weight coefficients based on the relative importance of active components in TCM, whereas the EWM method objectively assigns weights according to the internal variability of the indicators. By integrating the subjective evaluation of FAHP with the objective weighting of EWM, comprehensive weights were obtained to better reflect both empirical importance and statistical differentiation. The resulting comprehensive weight coefficients for coptisine, epiberberine, berberine hydrochloride, palmatine, baicalin, chrysin-7-O-β-glucuronide, wogonoside, baicalein, wogonin, oroxylin A, atractylodin, and extraction yield were calculated as 25.55%, 6.35%, 9.09%, 14.03%, 5.81%, 4.04%, 3.92%, 4.62%, 7.21%, 7.31%, and 6.18%, respectively (Table 2).

Table 2.

The comprehensive weight coefficients of indicator components. wjp was the subjective weight coefficients, Wjq was the objective weight coefficients, wjs was the comprehensive weights.

Name wjp wjq wjs
Coptisine 8.94 24.14 25.55
Epiberberine 8.94 6.00 6.35
Berberine hydrochloride 9.85 7.80 9.09
Palmatine 8.94 13.25 14.03
Baicalin 9.85 4.98 5.81
Chrysin-7-O-β-glucuronide 8.03 4.24 4.04
Wogonoside 8.03 4.12 3.92
Baicalein 8.03 4.86 4.62
Wogonin 8.03 7.58 7.21
Oroxylin A 8.03 7.69 7.31
Atractylodin 7.12 7.32 6.18
Extraction yield 6.21 8.01 5.89

Model fitting

RSM modeling

The comprehensive scores obtained from the BBD experiments were calculated using the weight coefficients derived from the FAHP-EWM method and are presented in Table 3. The comprehensive score data were analyzed using Design-Expert software (Stat-Ease Inc., Minneapolis, MN, USA) to evaluate model fitting and establish a polynomial regression equation. The resulting quadratic regression model was expressed as follows:

Table 3.

The results of Box-Behnken design. Coptisine (Y1), epiberberine (Y2), Berberine hydrochloride (Y3), palmatine (Y4), Baicalin (Y5), chrysin-7-O-β-glucuronide (Y6), Wogonoside (Y7), Baicalein (Y8), Wogonin (Y9), oroxylin A (Y10), atractylodin (Y11), extraction yield (Y12), comprehensive score (Y).

No. A
(%)
B
(h)
C
(mL/g)
Concentration(µg/mL) (%) Y
Y1 Y2 Y3 Y4 Y5 Y6 Y7 Y8 Y9 Y10 Y11 Y12
1 50 1 10 16.26 16.28 64.92 17.59 444.29 36.99 97.23 17.31 6.57 4.49 3.21 27.85 48.12
2 70 1 10 16.72 17.53 65.03 18.86 354.90 32.99 83.78 19.22 6.12 4.17 3.36 26.42 42.58
3 50 2 10 17.87 19.69 84.31 23.01 454.72 39.85 98.25 21.24 8.16 5.57 2.86 28.94 52.45
4 70 2 10 18.85 16.29 71.02 17.31 406.53 34.12 93.42 22.00 6.57 3.77 4.07 27.60 47.04
5 50 1.5 8 16.23 17.60 67.58 17.29 409.44 34.45 95.34 16.02 6.07 3.83 2.56 28.32 46.04
6 70 1.5 8 15.73 16.58 64.73 16.94 399.13 32.61 88.60 17.75 5.60 3.28 3.75 26.77 44.59
7 50 1.5 12 16.45 18.40 74.54 21.17 446.81 35.31 98.04 18.95 8.25 4.58 2.25 30.88 50.11
8 70 1.5 12 15.82 15.19 60.74 16.55 336.24 28.35 75.86 14.34 5.33 3.20 3.83 29.82 39.78
9 60 1 8 19.65 20.00 72.08 17.57 448.41 38.29 98.57 19.84 7.68 5.27 3.56 28.33 50.52
10 60 2 8 22.91 18.63 75.91 19.61 453.73 35.25 98.28 19.58 6.77 4.07 4.31 30.17 52.06
11 60 1 12 18.42 17.83 74.10 19.30 423.63 33.74 95.31 21.71 7.19 4.44 3.33 29.16 48.76
12 60 2 12 22.10 19.07 85.95 20.82 474.82 37.46 105.64 20.03 8.18 4.98 2.00 31.63 54.70
13 60 1.5 10 15.85 17.68 74.94 22.79 435.10 36.97 100.06 18.99 7.78 4.56 4.44 29.34 49.65
14 60 1.5 10 17.37 18.88 76.99 23.37 435.79 36.98 100.10 18.97 7.79 4.56 4.45 29.87 50.46
15 60 1.5 10 17.30 18.32 76.50 23.16 435.04 36.86 99.85 19.08 7.72 4.51 4.46 28.11 50.17
16 60 1.5 10 16.00 17.78 74.89 22.79 434.46 37.05 100.22 18.93 7.81 4.64 4.41 29.35 49.67
17 60 1.5 10 16.39 18.12 74.83 21.09 439.75 37.65 100.37 19.66 7.90 4.54 4.63 28.27 49.87
graphic file with name d33e1213.gif 6

The analysis of variance (ANOVA) for the quadratic model is summarized in Table 4. The ANOVA results demonstrated that the quadratic model was highly significant (F-value = 152.61, P < 0.0001), with a non-significant lack-of-fit term (F-value = 1.74, P = 0.2971), confirming the model’s adequacy. The coefficient of determination (R² = 0.9949) indicated an excellent correlation between the predicted and experimental values. Among the model terms, ethanol concentration (A) and reflux time (B) exerted significant effects on the comprehensive score (P < 0.0001). Significant interaction effects were observed between ethanol concentration and liquid-solid ratio (AC, P < 0.001), as well as between reflux time and liquid-solid ratio (BC, P < 0.001). The contour plots and three-dimensional response surface diagrams (Fig. 4), generated using the Model Graphs tool, visually illustrated these interaction effects. Steeper slopes in the 3D response surfaces indicated stronger synergistic or antagonistic interactions among variables, whereas flatter surfaces represented weaker interactions. Based on regression coefficients and graphical analyses, the influence of individual factors on the comprehensive score followed the order: ethanol concentration (A) > reflux time (B) > liquid-solid ratio (C).

Table 4.

Analysis of quadratic model variance of response surface. * represents significant at p < 0.05; ** represents highly significant at p < 0.01; *** represents extremely significant at p < 0.001.

Source Sum of Squares DF Mean Square F Value P Value Significance
Modal 217.99 9 24.22 152.61 < 0.0001*** Significant
A 64.58 1 64.58 406.91 < 0.0001***
B 33.09 1 33.09 208.48 < 0.0001***
C 0.0025 1 0.0025 0.0154 0.9046
AB 0.0042 1 0.0042 0.0266 0.875
AC 19.71 1 19.71 124.21 < 0.0001***
BC 4.84 1 4.84 30.5 0.0009**
A² 81.45 1 81.45 513.19 < 0.0001***
B² 16.54 1 16.54 104.19 < 0.0001***
C² 0.7995 1 0.7995 5.04 0.0597
Residual error 1.11 7 0.1587
Lack of fit 0.6287 3 0.2096 1.74 0.2971 Not significant
Pure error 0.4823 4 0.1206

Total

deviation

219.1 16
Fig. 4.

Fig. 4

Interaction among various factors. (A) Three-dimensional response surface map of reflux time and ethanol concentration. (B) Three-dimensional response surface map of liquid-solid ratio and ethanol concentration. (C) Three-dimensional response surface map of liquid-solid ratio and reflux time. (D) Contour map of reflux time and ethanol concentration. (E) Contour map of liquid-solid ratio and ethanol concentration. (F) Contour map of liquid-solid ratio and reflux time.

According to the BBD-RSM, the optimal extraction conditions were determined as follows: ethanol concentration of 54.3%, reflux time of 2 h, and liquid-solid ratio of 12 mL/g, yielding a predicted comprehensive score of 56.10. Considering practical production feasibility, the parameters were adjusted to an ethanol concentration of 55%, a reflux time of 2 h, and a liquid-solid ratio of 12 mL/g.

GCA

During the GCA, the data presented in Table 3 were normalized, and the results are provided in Supplementary Table S3. The relative correlation degrees of each sample are illustrated in Fig. 5. In general, a larger ri(s) and smaller ri(t) result in a higher rst, indicating that the corresponding process is closer to the ideal state. Therefore, the evaluation units can be ranked according to the magnitude of their relative correlation degrees—the higher the relative correlation, the better the extraction performance. Consequently, the rst value serves as an indicator of the quality of each sample’s extraction process. As shown in Fig. 5, sample No. 3 exhibited the highest relative correlation degree, suggesting that it had the most favorable extraction process among the 17 experimental runs. Accordingly, the optimal extraction conditions determined by GCA were an ethanol concentration of 50%, a reflux time of 2 h, and a liquid-solid ratio of 10 mL/g.

Fig. 5.

Fig. 5

Bubble chart of the grey correlation degree.

GA-BPNN modeling

Artificial neural networks (ANNs) are powerful tools for modeling complex and nonlinear systems and have been widely applied to multivariate process control42. In this study, the GA-BPNN model was established using MATLAB R2022b. A BPNN with a 3-7-1 topology was constructed, using the 17 sets of experimental data from the BBD as input samples. Among these, 14 datasets were used for training and 3 for validation. The BPNN model consisted of one hidden layer, and the Levenberg-Marquardt algorithm was employed as the training algorithm. The number of training iterations was set to 1000, and the learning rate was fixed at 0.01. The network’s training accuracy was evaluated using different numbers of neurons in the hidden layer. It was observed that when the hidden layer contained seven neurons, the model achieved higher training accuracy after multiple iterations, and this structure was therefore adopted. To further enhance model performance, the thresholds and connection weights of the BPNN were optimized using a GA. The GA parameters were set as follows: 25 generations and an initial population size of 5, resulting in the final GA-BPNN model. The comparative results of model performance are presented in Table 5. For the training set, the mean squared error (MSE) was 0.5504 and the R2 was 0.9558. For the test set, the MSE and R2 were 2.2620 and 0.7828, respectively. These results indicated that the predicted values of the GA-BPNN model deviated considerably from the experimental values, suggesting inadequate model fitting and limited applicability for optimizing the SHQRF extraction process.

Table 5.

Comparison of different model effects.

Modal MSE RMSE MAE R 2
Training Test Training Test Training Test Training Test
GA-BPNN 0.5504 2.2620 0.7419 1.5040 0.5072 1.4096 0.9558 0.7828
SVM 0.0583 0.0043 0.2414 0.0659 0.2013 0.0653 0.9959 0.9991

SVM modeling

The SVM model is based on the principle of minimizing structural risk and aims to identify an algorithm that maximizes the margin of separation between datasets, ensuring that as many data points as possible lie within the boundaries defined by the optimal hyperplane function. For nonlinear regression problems, a kernel function is applied to map the input data into a high-dimensional feature space, allowing the model to capture complex nonlinear relationships. The SVM model was developed using MATLAB R2022b, with 17 sets of experimental data from the BBD serving as sample data for network training. Among these, 14 sets were used for training and 3 for validation. The radial basis function (RBF) kernel was selected for the model, with a penalty factor of 4.0 and an RBF parameter of 0.8.

As shown in Table 5, the MSEs of the training and test sets were 0.0583 and 0.0043, respectively, indicating that the predicted values of the SVM model were in close agreement with the experimental results. These errors were notably smaller than those obtained using the GA-BPNN model. The R2 for both the training and test sets exceeded 0.99, which was also higher than that of the GA-BPNN model. These findings demonstrated that the SVM model exhibited reliable predictive performance, strong fitting ability, and superior accuracy. The comparison between experimental and predicted values, as well as the corresponding error analysis, is presented in Fig. 6. Overall, the SVM model showed better predictive capability for practical extraction processes and was more suitable for modeling the SHQRF extraction process.

Fig. 6.

Fig. 6

Comparison of experimental data with the predicted value obtained by model (A); analysis of error (B).

The SVM model was subsequently applied to predict the optimal extraction conditions. Ethanol concentration levels were set at 50%, 55%, 60%, 65%, and 70%; reflux times were set at 1.0, 1.5, and 2.0 h; and liquid-solid ratios were set at 8, 9, 10, 11, and 12 mL/g. A total of 75 combinations were used as input data, and the fully trained SVM model was employed for simulation tests to identify the optimal process parameters. The optimal extraction conditions predicted by the SVM model were an ethanol concentration of 55%, a reflux time of 2 h, and a liquid-solid ratio of 12 mL/g. Under these conditions, the predicted comprehensive score was 55.71. The optimized parameters obtained from the SVM model were consistent with those derived from the RSM analysis.

Validation experiment and comparison between different models

The results of the validation experiments are summarized in Table 6. As shown in the table, the optimal extraction conditions derived from the RSM and SVM models were consistent. The comprehensive scores of the optimal processes obtained through the Box-Behnken response surface and the SVM model were both higher than those obtained using the GCA optimization method. The experimentally validated comprehensive score for the RSM and SVM models was 56.08, which was in close agreement with the predicted value of 55.71, confirming the accuracy and reliability of the established models. Therefore, the optimal extraction parameters for SHQRF were finally determined as follows: ethanol concentration of 55%, reflux time of 2 h, and liquid-solid ratio of 12 mL/g.

Table 6.

Verification of experiment results.

Method Concentration(µg/mL) (%) Y Mean ± SD RSD(%)
Y 1 Y 2 Y 3 Y 4 Y 5 Y 6 Y 7 Y 8 Y 9 Y 10 Y 11 Y 12
GCA 19.58 20.90 72.58 21.38 468.88 38.06 99.19 17.51 6.85 3.86 2.86 30.05 52.13 52.67 ± 0.72 1.37
19.34 21.16 73.75 21.71 469.56 38.57 99.53 17.71 6.98 4.81 2.91 30.15 52.41
19.67 21.60 75.71 22.06 479.70 39.33 101.65 18.22 7.12 4.85 3.02 30.10 53.49
RSM or SVM 24.50 20.84 89.30 22.26 478.99 37.75 106.06 20.59 8.45 5.33 2.04 30.17 56.19 56.08 ± 0.21 0.37
24.32 20.60 89.20 22.21 480.66 37.38 105.87 20.63 8.53 5.47 2.02 30.10 56.20
24.43 20.77 88.71 21.67 476.52 37.26 105.71 20.45 8.54 5.31 1.99 29.92 55.84

Comparison of extraction processes before and after optimization

The extraction processes before and after optimization were carried out, and the contents of the 11 bioactive components were determined. The results of the chromatographic mirror comparison are shown in Fig. 7. As illustrated, the retention times of the main peaks were identical between the two processes, indicating consistent component identities. However, the peak areas and overall chemical composition were markedly enhanced in the optimized extraction process. The mass fractions of the 11 components are presented in Fig. 8. Compared with the original extraction process, the mass fractions of all 11 index components increased significantly under the optimized conditions, resulting in a total content increase of 131.90%. These findings demonstrate that the optimized extraction process markedly improved extraction efficiency and overall yield compared with the original method.

Fig. 7.

Fig. 7

Mirror comparison of the post-optimization process (A) and the pre-optimization process (B). (1) coptisine; (2) epiberberine; (3) berberine hydrochloride; (4) palmatine; (5) baicalin; (6) chrysin-7-O-β-glucuronide; (7) wogonoside; (8) baicalein; (9) wogonin; (10) oroxylin A; 11. atractylodin.

Fig. 8.

Fig. 8

The mass fractions of the indicator components post-optimization process and the pre-optimization process.

Multivariate statistical analysis

HCA

HCA is an unsupervised pattern recognition technique that groups data based on the similarity or distance between samples, progressively merging or splitting them in a hierarchical structure43,44. To visually compare the differences between the optimal extraction processes obtained through various analytical methods and the original extraction process, HCA was performed for statistical evaluation. The data presented in Table 6, along with the 11 evaluation indices obtained from the test solution prepared using the original extraction process, were used as variables and imported into TBtools software for analysis. The between-group linkage method and Euclidean distance were employed as the clustering criteria, and the resulting heatmap is shown in Fig. 9A. The samples were initially clustered into two groups: the original extraction process formed one cluster, while all optimized extraction processes were grouped into the other. The contents of the 11 evaluation indices in the original extraction process were significantly lower than those in the optimized processes, indicating substantial improvement after optimization. When the Euclidean distance threshold was further reduced, the samples were divided into three distinct categories: (1) the original extraction process, (2) the optimal process derived from GCA, and (3) the optimal processes obtained from RSM and SVM. This clustering pattern demonstrated clear differences among the extraction methods.

Fig. 9.

Fig. 9

HCA scatter plot (A) and PCA scatter plot (B).

PCA

PCA is a commonly used dimensionality reduction technique that simplifies complex datasets by transforming correlated variables into a smaller number of uncorrelated principal components. This method has been widely applied in the comprehensive evaluation and classification of TCM45. In this study, the experimental data were imported into SIMCA 14.1 for PCA analysis, and the results are presented in Fig. 9B. The samples were clearly classified into three distinct groups. The optimal process obtained by GCA was located in quadrant I, while those optimized by RSM and SVM were distributed in quadrant IV. In contrast, the original extraction process was positioned in quadrants II and III. This clear spatial separation indicated significant differences in the overall chemical composition of the samples before and after process optimization, as well as among the optimization results derived from different analytical methods—findings consistent with the HCA results. Moreover, the tight clustering of sample points within each group suggested good homogeneity and reproducibility among replicates.

Discussion

In this study, the ethanol reflux extraction method was employed to extract the bioactive components of SHQRF. The extraction time, ethanol concentration, reflux duration, and liquid-solid ratio were identified as the major factors influencing ethanol reflux extraction46–48. In our previous work, the number of extraction cycles was investigated, and it was observed that both the total content of the 11 index components and the extraction yield increased with additional extraction cycles. However, when the extraction was performed more than twice, the dissolution of the total components tended to approach saturation, and the contents became stable despite further increases in extraction times49. Considering practical production requirements and cost efficiency, the number of extraction cycles was fixed at two in this study. During the single-factor experiments, the total contents of the index components and extraction yields increased initially and then decreased with rising ethanol concentration. According to the principle of “like dissolves like,” the solubility of SHQRF components improved as the polarity of the solvent approached that of the solutes22. The solubility increased up to 60% ethanol and declined beyond this concentration, likely due to excessive reduction in solvent polarity. Similar findings were reported for the extraction of Yujin Powder20. In the reflux time experiments, both the total content of the 11 components and the extraction yield exhibited an upward trend with increasing reflux time but plateaued after 2 h. This phenomenon may be attributed to the establishment of a dynamic equilibrium between the extraction solvent and the herbal matrix50. Once the dissolution of active components reached equilibrium, further prolongation of the reflux time no longer improved extraction efficiency51. A similar trend was observed in the extraction of Arctic Chlorella sp.52. Likewise, the total content of the indicator components and extraction yield increased initially and then decreased as the liquid-solid ratio increased. This can be explained by the fact that a higher solvent volume enhances the contact area between the herbal material and solvent, thereby improving mass transfer. However, when the liquid-solid ratio exceeded an optimal threshold, most soluble components had already been extracted, and the extraction efficiency declined due to solvent saturation53.

The RSM, GCA, and SVM models were constructed to predict and optimize the extraction process of SHQRF. According to the prediction results, the optimal extraction conditions obtained using the RSM and SVM models were highly consistent, showing strong agreement between the predicted and experimentally measured values. Among the models, the SVM approach demonstrated significant advantages in rapidly predicting extraction outcomes under various parameter combinations, thereby reducing the number of required experimental runs and saving both time and resources. Due to its strong capacity for modeling complex nonlinear relationships, the SVM method effectively captured the interactive and nonlinear dependencies among multiple process parameters, resulting in more accurate and reliable optimization. Therefore, the SVM model can serve as a powerful alternative to traditional RSM regression for process prediction and optimization in multivariate systems. In the validation experiments, the comprehensive scores of the optimal extraction processes predicted by the RSM and SVM models were higher than those obtained using the GCA method. The GCA technique can evaluate the effects of each factor combination comprehensively by calculating the degree of correlation between different parameter levels, thereby minimizing the impact of experimental randomness on the final conclusion54. In this study, GCA further clarified the specific influence of each parameter on extraction performance across 13 level combinations, identifying combination No. 3 as superior. However, as a statistical correlation-based method, GCA lacks predictive capability for unseen samples and can only optimize within the range of known data, which limits its broader applicability.

Between the two machine learning models, the SVM exhibited lower error values and higher fitting accuracy compared with the GA-BPNN model, demonstrating superior reliability and precision in optimizing the SHQRF extraction process. The BPNN is one of the most widely used neural network architectures, characterized by strong multidimensional function-mapping capability and a flexible network structure. It features a simple and efficient training algorithm with a fast learning rate. When combined with a GA, the BPNN gains enhanced global optimization capability and improved resistance to overfitting55. However, in this study, the GA-BPNN model may have been limited by the small sample size, resulting in insufficient data for effective generalization. Under such conditions, inter-sample variability and component complexity can be amplified, making it difficult for the model to accurately capture key variable relationships, thereby leading to suboptimal fitting performance. In contrast, the SVM model proved more suitable for small-sample datasets due to its structural risk minimization principle and flexible kernel function. By efficiently mapping data into a high-dimensional feature space, the SVM effectively handles nonlinear regression problems, achieving robust and accurate predictions even under limited data conditions.

Conclusions

In this study, a combined approach integrating BBD-RSM, GCA, and machine learning models was applied to optimize the ethanol reflux extraction process of SHQRF. The comprehensive score (Y), calculated based on the contents of 11 bioactive components and extraction yields, served as the evaluation index. The optimized extraction parameters were determined as follows: extraction performed twice, ethanol concentration of 55%, reflux time of 2 h per cycle, and liquid-solid ratio of 12 mL/g. Under these conditions, the contents of the 11 index components and the total yield were significantly higher than those obtained from the original extraction process. The results of HCA and PCA demonstrated clear distinctions among the original extraction process, the optimal process obtained by GCA, and those optimized through RSM and SVM, confirming the consistency and validity of the optimization results. This study successfully demonstrated the effectiveness of multi-index evaluation combined with machine learning for process optimization in TCM. Subsequent research should explore other advanced machine learning algorithms or larger datasets to enhance model robustness and generalize the integrated optimization strategy to other complex TCM compound formulations.

Supplementary Information

Below is the link to the electronic supplementary material.

Author contributions

Q.C. : Conceptualization, Methodology, Software, Writing—original draft preparation. P.M. : Writing—review and editing, Supervision. X.H. : Methodology, Validation. Y .L. : Methodology, Data curation. J. T. : Validation. Z.P. : Resources, Writing—review and editing, Project administration, Funding acquisition. J. Y. : Conceptualization, Resources, Project administration, Funding acquisition. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Hunan Provincial Natural Science Foundation (no. 2023JJ60133) and Key Discipline Project on Chinese Pharmacology of Hunan University of Chinese Medicine [202302].

Data availability

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Zhirong Peng, Email: 13272052757@163.com.

Jianye Yan, Email: yanjy@hnucm.edu.cn.

References

  • 1.Wang, A. Q., Yuan, Q. J., Guo, N., Yang, B. & Sun, Y. Research progress on medicinal resources of Coptis and its isoquinoline alkaloids. Zhongguo Zhong Yao Za Zhi. 46, 3504–3513. 10.19540/j.cnki.cjcmm.20210430.103 (2021). [DOI] [PubMed] [Google Scholar]
  • 2.Wang, J. et al. Coptidis rhizoma: a comprehensive review of its traditional uses, botany, phytochemistry, pharmacology and toxicology. Pharm. Biol.57, 193–225. 10.1080/13880209.2019.1577466 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Huang, X. et al. Inhibition of berberine hydrochloride on Candida albicans biofilm formation. Biotechnol. Lett.42, 2263–2269. 10.1007/s10529-020-02938-6 (2020). [DOI] [PubMed] [Google Scholar]
  • 4.Rui, Z., Chang-Pei, X., Jing-Jing, Z. & Hong-Jun, Y. Research progress on chemical compositions of Coptidis rhizoma and pharmacological effects of berberine. Zhongguo Zhong Yao Za Zhi. 45, 4561–4573. 10.19540/j.cnki.cjcmm.20200527.202 (2020). [DOI] [PubMed] [Google Scholar]
  • 5.Li, T. K. et al. Human topoisomerase I poisoning by protoberberines: potential roles for both drug-DNA and drug-enzyme interactions. Biochemistry39, 7107–7116. 10.1021/bi000171g (2000). [DOI] [PubMed] [Google Scholar]
  • 6.Wen, Y., Wang, Y., Zhao, C., Zhao, B. & Wang, J. The pharmacological efficacy of Baicalin in inflammatory diseases. Int. J. Mol. Sci.2410.3390/ijms24119317 (2023). [DOI] [PMC free article] [PubMed]
  • 7.Liu, T. et al. Antibacterial synergy between linezolid and baicalein against methicillin-resistant Staphylococcus aureus biofilm in vivo. Microb. Pathog. 147, 104411. 10.1016/j.micpath.2020.104411 (2020). [DOI] [PubMed] [Google Scholar]
  • 8.Liao, H., Ye, J., Gao, L. & Liu, Y. The main bioactive compounds of Scutellaria baicalensis Georgi. for alleviation of inflammatory cytokines: A comprehensive review. Biomed. Pharmacother. 133, 110917. 10.1016/j.biopha.2020.110917 (2021). [DOI] [PubMed] [Google Scholar]
  • 9.Yu, X. et al. Wogonoside ameliorates airway inflammation and mucus hypersecretion via NF-κB/STAT6 signaling in ovalbumin-induced murine acute asthma. J. Agric. Food Chem.72, 7033–7042. 10.1021/acs.jafc.3c04082 (2024). [DOI] [PubMed] [Google Scholar]
  • 10.Yu, X. et al. Wogonoside inhibits inflammatory cytokine production in lipopolysaccharide-stimulated macrophage by suppressing the activation of the JNK/c-Jun signaling pathway. Ann. Transl Med.8, 532. 10.21037/atm.2020.04.22 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Cho, J. & Lee, H. K. Wogonin inhibits excitotoxic and oxidative neuronal damage in primary cultured rat cortical cells. Eur. J. Pharmacol.485, 105–110. 10.1016/j.ejphar.2003.11.064 (2004). [DOI] [PubMed] [Google Scholar]
  • 12.Lee, H. et al. Flavonoid wogonin from medicinal herb is neuroprotective by inhibiting inflammatory activation of microglia. Faseb J. 17, 1943– 1944. 10.1096/fj.03-0057fje (2003). [DOI] [PubMed]
  • 13.Yi, Y. et al. Chrysin 7-O-β-D-glucuronide, a dual inhibitor of SARS-CoV-2 3CL(pro) and PL(pro), for the prevention and treatment of COVID-19. Int. J. Antimicrob. Agents. 63, 107039. 10.1016/j.ijantimicag.2023.107039 (2024). [DOI] [PubMed] [Google Scholar]
  • 14.Gao, J. et al. Oroxylin A is a severe acute respiratory syndrome coronavirus 2-spiked pseudotyped virus blocker obtained from radix scutellariae using angiotensin-converting enzyme II/cell membrane chromatography. Phytother Res.35, 3194–3204. 10.1002/ptr.7030 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dong, Y., Zhang, X., Yao, C., Xu, R. & Tian, X. Atractylodin attenuates the expression of MUC5AC and extracellular matrix in lipopolysaccharide-induced airway inflammation by inhibiting the NF-κB pathway. Environ. Toxicol.36, 1911–1922. 10.1002/tox.23311 (2021). [DOI] [PubMed] [Google Scholar]
  • 16.Lin, Y. C. et al. Atractylodin ameliorates ovalbumin–induced asthma in a mouse model and exerts immunomodulatory effects on Th2 immunity and dendritic cell function. Mol. Med. Rep.22, 4909–4918. 10.3892/mmr.2020.11569 (2020). [DOI] [PubMed] [Google Scholar]
  • 17.Jiang, M. et al. Optimization of the extraction process for Shenshou Taiyi powder based on Box-Behnken experimental design, standard relation, and FAHP-CRITIC methods. BMC Complement. Med. Ther.24, 251. 10.1186/s12906-024-04554-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zang, Z. et al. Evaluation of the effect of ultrasonic pretreatment on vacuum far-infrared drying characteristics and quality of Angelica sinensis based on entropy weight-coefficient of variation method. J. Food Sci.88, 1905–1923. 10.1111/1750-3841.16566 (2023). [DOI] [PubMed] [Google Scholar]
  • 19.Du, Y. et al. Optimization of extraction or purification process of multiple components from natural products: entropy weight method combined with Plackett-Burman design and central composite design. Molecules. 2610.3390/molecules26185572 (2021). [DOI] [PMC free article] [PubMed]
  • 20.Jiang, L. et al. Optimization of ethanol extraction technology for Yujin powder using response surface methodology with a Box-Behnken design based on analytic hierarchy Process-Criteria importance through intercriteria correlation weight analysis and its safety evaluation. Molecules. 2810.3390/molecules28248124 (2023). [DOI] [PMC free article] [PubMed]
  • 21.Xie, R. F. et al. Optimization of high pressure machine decocting process for Dachengqi Tang using HPLC fingerprints combined with the Box-Behnken experimental design. J. Pharm. Anal.5, 110–119. 10.1016/j.jpha.2014.07.001 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Cao, Q. et al. Simultaneous optimization of ultrasound-assisted extraction for total flavonoid content and antioxidant activity of the tender stem of Triarrhena lutarioriparia using response surface methodology. Food Sci. Biotechnol.30, 37–45. 10.1007/s10068-020-00851-2 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Ferreira, S. L. et al. Box-Behnken design: an alternative for the optimization of analytical methods. Anal. Chim. Acta. 597, 179–186. 10.1016/j.aca.2007.07.011 (2007). [DOI] [PubMed] [Google Scholar]
  • 24.Zhang, L. et al. Simultaneous optimization of ultrasound-assisted extraction for flavonoids and antioxidant activity of Angelica keiskei using response surface methodology (RSM). Molecules. 2410.3390/molecules24193461 (2019). [DOI] [PMC free article] [PubMed]
  • 25.Zhang, Y. et al. Evaluation of acetic acid treatment of fresh-cut water chestnuts using gray-correlation analysis based on the variation-coefficient weight. Front. Nutr.11, 1370611. 10.3389/fnut.2024.1370611 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Lai, P. et al. Grey correlation analysis of drying characteristics and quality of Hypsizygus marmoreus (crab-flavoured mushroom) by-products. Molecules. 2810.3390/molecules28217394 (2023). [DOI] [PMC free article] [PubMed]
  • 27.Deng, X. et al. A novel strategy for active compound efficacy status identification in multi-tropism Chinese herbal medicine (Scutellaria baicalensis Georgi) based on multi-indexes spectrum-effect gray correlation analysis. J. Ethnopharmacol.300, 115677. 10.1016/j.jep.2022.115677 (2023). [DOI] [PubMed] [Google Scholar]
  • 28.Lan, Q. et al. Comprehensive application of AHP-CRITIC hybrid weighting method, grey correlation analysis and BP-ANN in optimization of extraction process of Qizhi prescription. Chin. J. Exp. Tradit Med. Form. 1–11. 10.13422/j.cnki.syfjx.20250667 (2024). [DOI] [PubMed]
  • 29.Peng, W. et al. Network pharmacology combines machine learning, molecular simulation dynamics and experimental validation to explore the mechanism of acetylbinankadsurin A in the treatment of liver fibrosis. J. Ethnopharmacol.323, 117682. 10.1016/j.jep.2023.117682 (2024). [DOI] [PubMed] [Google Scholar]
  • 30.Chen, H. & He, Y. Machine learning approaches in traditional Chinese medicine: A systematic review. Am. J. Chin. Med.50, 91–131. 10.1142/s0192415x22500045 (2022). [DOI] [PubMed] [Google Scholar]
  • 31.Wang, Y., Shi, X., Li, L., Efferth, T. & Shang, D. The impact of artificial intelligence on traditional Chinese medicine. Am. J. Chin. Med.49, 1297–1314. 10.1142/s0192415x21500622 (2021). [DOI] [PubMed] [Google Scholar]
  • 32.Ma, S. et al. Machine learning in TCM with natural products and molecules: current status and future perspectives. Chin. Med.18, 43. 10.1186/s13020-023-00741-9 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Gao, H., Wu, C., Huang, D., Zha, D. & Zhou, C. Prediction of fetal weight based on back propagation neural network optimized by genetic algorithm. Math. Biosci. Eng.18, 4402–4410. 10.3934/mbe.2021222 (2021). [DOI] [PubMed] [Google Scholar]
  • 34.Shu, J. et al. Optimization of tetrastigma Hemsleyanum extraction process based on GA-BPNN model and analysis of its antioxidant effect. Heliyon9, e20200. 10.1016/j.heliyon.2023.e20200 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Lu, M. et al. Optimization of adsorption performance of cerium-loaded intercalated bentonite by CCD-RSM and GA-BPNN and its application in simultaneous removal of phosphorus and ammonia nitrogen. Chemosphere336, 139241. 10.1016/j.chemosphere.2023.139241 (2023). [DOI] [PubMed] [Google Scholar]
  • 36.Wu, J., Yang, F., Guo, L. & Sheng, Z. Modeling and optimization of ellagic acid from chebulae fructus using response surface methodology coupled with artificial neural network. Molecules. 2910.3390/molecules29163953 (2024). [DOI] [PMC free article] [PubMed]
  • 37.Heikamp, K. & Bajorath, J. Support vector machines for drug discovery. Expert Opin. Drug Discov. 9, 93–104. 10.1517/17460441.2014.866943 (2014). [DOI] [PubMed] [Google Scholar]
  • 38.Silva, G. F. S., Fagundes, T. P., Teixeira, B. C. & Chiavegatto Filho, A. D. P. Machine learning for hypertension prediction: a systematic review. Curr. Hypertens. Rep.24, 523–533. 10.1007/s11906-022-01212-6 (2022). [DOI] [PubMed] [Google Scholar]
  • 39.Li, X. et al. Optimization of ultrasonic-assisted extraction of active components and antioxidant activity from Polygala tenuifolia: A comparative study of the response surface methodology and least squares support vector machine. Molecules. 2710.3390/molecules27103069 (2022). [DOI] [PMC free article] [PubMed]
  • 40.Wang, W., Li, J., Fang, Y., Zheng, Y. & You, F. An effective hybrid feature selection using entropy weight method for automatic sleep staging. Physiol. Meas.4410.1088/1361-6579/acff35 (2023). [DOI] [PubMed]
  • 41.Liu, W., Zheng, Q., Pang, L., Dou, W. & Meng, X. Study of roof water inrush forecasting based on EM-FAHP two-factor model. Math. Biosci. Eng.18, 4987–5005. 10.3934/mbe.2021254 (2021). [DOI] [PubMed] [Google Scholar]
  • 42.Ranjan, D., Mishra, D. & Hasan, S. H. Bioadsorption of arsenic: an artificial neural networks and response surface methodological approach. Ind. Eng. Chem. Res.50, 9852–9863. 10.1021/ie200612f (2011). [Google Scholar]
  • 43.Cao, X. et al. Comparative investigation for rotten xylem (kuqin) and strip types (tiaoqin) of Scutellaria baicalensis Georgi based on fingerprinting and chemical pattern recognition. Molecules. 2410.3390/molecules24132431 (2019). [DOI] [PMC free article] [PubMed]
  • 44.Li, S., Huang, Y., Zhang, F., Ao, H. & Chen, L. Comparison of volatile oil between the Ligusticum sinese Oliv. and Ligusticum jeholense Nakai et Kitag. based on GC-MS and chemical pattern recognition analysis. Molecules. 2710.3390/molecules27165325 (2022). [DOI] [PMC free article] [PubMed]
  • 45.Zheng, C., Li, W., Yao, Y. & Zhou, Y. Quality evaluation of atractylodis macrocephalae rhizoma based on combinative method of HPLC fingerprint, quantitative analysis of multi-components and chemical pattern recognition analysis. Molecules. 2610.3390/molecules26237124 (2021). [DOI] [PMC free article] [PubMed]
  • 46.Hu, Y. et al. Optimisation of Ethanol-Reflux extraction of saponins from steamed Panax Notoginseng by response surface methodology and evaluation of hematopoiesis effect. Molecules. 2310.3390/molecules23051206 (2018). [DOI] [PMC free article] [PubMed]
  • 47.Chen, Y., Zhang, C., Zhang, M. & Fu, X. Three statistical experimental designs for enhancing yield of active compounds from herbal medicines and anti-motion sickness bioactivity. Pharmacogn Mag. 11, 435–443. 10.4103/0973-1296.160444 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Jiang, Y. et al. Ultrasonic-assisted ionic liquid extraction of two biflavonoids from Selaginella tamariscina. ACS Omega. 5, 33113–33124. 10.1021/acsomega.0c04723 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zhang, Y. et al. Extraction process optimization and link evaluation of Reyanning mixture based on quality by design Idea. Zhongguo Zhong Yao Za Zhi. 49, 3229–3241. 10.19540/j.cnki.cjcmm.20240313.302 (2024). [DOI] [PubMed] [Google Scholar]
  • 50.Meng, Z., Zhao, J., Duan, H., Guan, Y. & Zhao, L. Green and efficient extraction of four bioactive flavonoids from Pollen Typhae by ultrasound-assisted deep eutectic solvents extraction. J. Pharm. Biomed. Anal.161, 246–253. 10.1016/j.jpba.2018.08.048 (2018). [DOI] [PubMed] [Google Scholar]
  • 51.Chen, Y. et al. Extraction optimization of polysaccharides from wet red microalga Porphyridium purpureum using response surface methodology. Mar. Drugs. 2210.3390/md22110498 (2024). [DOI] [PMC free article] [PubMed]
  • 52.Song, H. et al. Extraction optimization, purification, antioxidant activity, and preliminary structural characterization of crude polysaccharide from an Arctic chlorella sp. Polymers. 10. 10.3390/polym10030292 (2018). [DOI] [PMC free article] [PubMed]
  • 53.Ma, Y. et al. Reflux extraction optimization and antioxidant activity of phenolic compounds from Pleioblastus amarus (Keng) shell. Molecules. 2710.3390/molecules27020362 (2022). [DOI] [PMC free article] [PubMed]
  • 54.Yin, S. et al. Simultaneous determination of multiple bioactive constituents in Abelmoschi corolla by UFLC-QTRAP-MS/MS. Chin. Mater. Med.46, 2527–2536. 10.19540/j.cnki.cjcmm.20201225.302 (2021). [DOI] [PubMed] [Google Scholar]
  • 55.Fernandez, M., Caballero, J., Fernandez, L. & Sarai, A. Genetic algorithm optimization in drug design QSAR: Bayesian-regularized genetic neural networks (BRGNN) and genetic algorithm-optimized support vectors machines (GA-SVM). Mol. Divers.15, 269–289. 10.1007/s11030-010-9234-9 (2011). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES