Abstract
The quality variation of Goji berry across cultivated zones leaded to an urgent requirement of geographical discrimination. Classical untargeted to targeted strategy faced with the problem of reducing prediction capability. Here, a new untargeted to targeted transition strategy was proposed. At first, the AntDAS-GCMS was used for automatically resolving compounds in untargeted metabolic profiling data to initialize a biomarker set. Compounds with similar chemical structures and/or homologous metabolic pathways were introduced to enhance the coverage of biomarker set, which were used to construct an MRM-based targeted metabolomics method. Quantitative analysis revealed that, proline was dramatically higher in Gansu (252.31 ± 310.68 μg/g) than Ningxia (1.85 ± 4.56 μg/g) and Xinjiang (0.08 ± 0.18 μg/g); cellobiose showed 4.2-fold higher concentration in Xinjiang (84.32 ± 28.54 μg/g) than Ningxia (20.32 ± 3.90 μg/g); glucose was increased in Ningxia (5255.56 ± 425.48 μg/g). The developed strategy achieved >96% prediction accuracy for Goji berries from different zones, outperforming classical untargeted-to-targeted approaches (>91%).
Keywords: Goji berry, Untargeted to targeted transition strategy, GC–MS, AntDAS-GCMS, Chemometrics
Graphical abstract

Highlights
-
•
An untargeted to targeted strategy was developed for discriminating Goji berry.
-
•
AntDAS-GCMS enables automatic marker screening and identification in untargeted way.
-
•
Valuable compounds beyond the screened ones can be obtained with the new strategy.
-
•
GC–MS/MS-based targeted method can greatly enhance geographical discrimination.
-
•
The developed strategy outperformed the traditional untargeted to targeted method.
1. Introduction
Goji berry, a traditional health food and medicinal herb, has attracted significant attention due to its rich nutritional composition and various pharmacological properties. Goji berries contain high levels of sugars, amino acids, vitamins, minerals, and bioactive substances, offering benefits such as anti-fatigue, antioxidant, immune-regulating, and anti-aging effects (Miranda et al., 2024; Zheng et al., 2021). In China, the primary production regions for Goji berries include Ningxia, Gansu, Qinghai, and Xinjiang provinces, with Ningxia Goji berries recognized as authentic medicinal materials (Zhou et al., 2023). Goji berries from different regions can command different prices in the local market, with the price of Ningxia Goji berries 2–3 times higher than those from other regions, creating more opportunities for fraudulent behavior among sellers and posing serious challenges for market regulation (Ju et al., 2025). For example, Li and Nirere et al. (Li et al., 2017; Nirere et al., 2023) discovered that there was a phenomenon of adulteration of Goji berry. As a result, there is an urgent need for accurate geographical traceability of Goji berries to ensure product authenticity and quality.
A comprehensive understanding of the chemical composition of Goji berries is essential for developing an accurate classification model for geographical discrimination. Currently, most research has relied on chromatography-mass spectrometry techniques. For example, Li et al. combined GC–MS and electronic nose technology with multivariate statistical analysis to identify differences in the aroma of Goji berry samples, enabling differentiation based on geographical origins (Li et al., 2017). Similarly, Bertoldi et al. developed a method using HPLC-DAD-MS to distinguish the geographical origins of Goji berries based on carotenoid content, successfully differentiating Italian Goji berries from Asian varieties (Bertoldi et al., 2019). In addition, He et al. developed an intelligent recognition method based on color space transformation and textural morphological features to classify Goji berries by origin (He et al., 2023). Given the complex composition of Goji berries, classical GC–MS-based untargeted metabolomics faced with false-positive compound identification results and GC–MS/MS-based targeted metabolomics restricted to metabolite coverage. There is an urgent requirement of developing untargeted-to-targeted workflow for constructing reliable geographical discrimination model for Goji berries.
The use of gas chromatography with mass spectrometry, such as gas chromatography–mass spectrometry (GC–MS) and GC–MS/MS, has been extensively used in metabolomics for the geographical determination of analyzed samples (De Lima Morais da Silva et al., 2017). In practical applications, GC–MS is frequently used in an untargeted manner to screen and identify metabolites that vary across different zones or groups (Rasekh et al., 2023), while GC–MS/MS is used for the precise quantification of specific compounds (Kim et al., 2024). Each technique offers distinct advantages. GC–MS provides extensive compound information in an untargeted approach, allowing for the characterization of unknown compounds based on collected mass spectra (Fan et al., 2021). By contrast, GC–MS/MS offers higher sensitivity than GC–MS and delivers significantly more accurate quantification for hundreds of targeted compounds (Wang et al., 2023). Both techniques have been successfully employed for the geographical discrimination of analyzed samples (Gamal et al., 2024). Moreover, GC–MS provides high chromatographic resolution and rich mass spectral information, enabling the identification and quantification of a wide range of volatile and semi-volatile metabolites. These metabolites—including sugars, organic acids, amino acids, and nitrogen-containing compounds—are closely associated with environmental factors such as climate, soil nutrients, and irrigation conditions, making them highly relevant for geographical origin discrimination (Mourão et al., 2023).
Currently, researchers may resort to an untargeted to targeted strategy for the purpose of screening valuable compounds to construct a reliable geographical discrimination model. In practical applications, however, analysts frequently faced with the problems like i) an efficient data analysis tool for accurate compound resolution, identification, and valuable screening; ii) reduction of model prediction capability when the constructed untargeted model was transformed into a targeted one. For instance, advanced data analysis tools (Heuckeroth et al., 2024; Huan et al., 2017; Stein, 1999; Tsugawa et al., 2015) faced with numerous false-positive results in untargeted data analysis that cannot be confirmed with targeted method. Our research group developed a new GC–MS data analysis tool, AntDAS-GCMS, for addressing false-positive results and screening valuable components (Zhang et al., 2020), which may provide another solution for classical untargeted to targeted methodology. With respect to the geographical discrimination of natural products, the screened metabolites come from different pathways, implying that a number of underlying valuable compounds that distributed in the same pathway or obtain similar chemical structure may be used to improve prediction capability of discrimination model from GC–MS/MS (Mourão et al., 2023).
In this study, we proposed a new untargeted to targeted transition strategy for Goji berries, which was achieved based on our recently proposed GC–MS data analysis software, AntDAS-GCMS (Zhang et al., 2020). First, volatile and semi-volatile compounds in Goji berries were analyzed using GC–MS in an untargeted manner. The acquired raw GC–MS data files were then imported into AntDAS-GCMS to automatically perform component resolution, registration, and chemometric analysis to identify underlying biomarkers. The screened compounds were then used in combination with a number of known compounds to establish a targeted quantification method using GC–MS/MS. Finally, a chemometric model was established for the geographical discrimination of hundreds of Goji berries from different areas in China.
2. Experimental
2.1. Chemicals and reagents
High-performance liquid chromatography (HPLC) grade isopropanol, acetonitrile, dichloromethane, and methanol were obtained from Merck (Darmstadt, Germany). Watsons water (Guangzhou, China), methoxyamine pyridine standard solution (20 mg/mL, CHEMISCI, China), and N,O-bis(trimethylsilyl)trifluoro-acetamide (BSTFA) were obtained from Sigma-Aldrich (St. Louis, USA). C6–C32 n-alkanes (GC quality), which were used to calculate the retention index (RI) (Zhu et al., 2022), were purchased from Accustandard Inc. (USA).
The following 44 standards, namely, alanine, 3-aminobutyric acid, valine, leucine, isoleucine, proline, glycine, serine, threonine, erythritol, threitol, methionine, aspartic acid, pyroglutamic acid, hydroxyproline, 3-hydroxy-3-methylglutaric acid, arginine, glutamic acid, phenylalanine, xylose, asparagine, 1,6-anhydro-β-d-glucose, xylitol, arabitol, tyrosine methyl ester, glutamine, ornithine, fructose, citrulline, glucose, galactose, mannose, histidine, lysine, mannitol, tyrosine, gluconic acid, tryptophan, cysteine, glucose-6-phosphate, sucrose, cellobiose, raffinose, and melezitose (purity of ≥98%), were purchased from Sigma-Aldrich (St. Louis, USA).
2.2. Sample collection
In this work, we collected 186 Goji berry samples, including 50 samples from Ningxia province, 70 samples from Gansu province, and 66 samples from Xinjiang province. All samples were collected on the year of 2023. The geographical zone of each Goji berry sample was provided in Table S1. Samples were stored in a dark place at room temperature. Experiment was performed immediately when the sample collection procedure was finished.
2.3. Sample preparation
The Goji berry samples used in this study were sourced from Ningxia, Gansu, and Xinjiang, China. In this experiment, 10.0 mg of powder was weighed and placed into a 2.0 mL centrifuge tube. Then, 1.5 mL of extraction solvent (isopropanol:acetonitrile:water = 3:3:2, v/v/v) was added and vortexed for 2 min, followed by ultrasound treatment on ice for 30 min. Subsequently, 500 μL of the supernatant was transferred into a chromatography vial, which was dried in a vacuum centrifuge for 4 h. The derivatization treatment consisted of two steps. In the first step, 100 μL of methoxyamine pyridine standard solution was added for reconstitution, which was then incubated on a thermostatic shaker at 37 °C for 120 min. In the second step, 100 μL of BSTFA reagent was added, and the mixture was reacted on a thermostatic shaker at 60 °C for 90 min (Li et al., 2022; Matsueda et al., 2023). After cooling to room temperature, a 1 μL sample was used for GC–MS analysis.
2.4. GC–MS conditions
2.4.1. Untargeted metabolomics based on full scan mode of GC–MS
A Thermo TSQ 9000-TRACE 1310 GC–MS system was used for untargeted metabolomics analysis. Split injection mode was adopted, with a split ratio of 50:1, and an inlet temperature of 300 °C. High-purity helium gas (99.999%) was used as the carrier gas, which was performed at a flow rate of 1.0 mL/min. Compounds were separated with a DB-5MS capillary column (30 m × 0.25 mm × 0.25 μm). The oven temperature was initiated at 70 °C for 4 min, then gradually increased to 310 °C at a rate of 5 °C/min, where it was held for 15 min (Ma et al., 2023; Zhang et al., 2023).
Full scan mode was used for mass spectrum collection to comprehensively capture all extracted metabolites. An EI ion source was employed for ionization, with the temperature set at 300 °C and the transfer line temperature at 280 °C. A mass range was set between 50 and 550 Da, with a solvent delay time of 7.2 min. The scan speed was 0.2 s per spectrum.
2.4.2. Targeted metabolomics based on the multiple reaction monitoring mode of GC–MS/MS
The same Thermo TSQ 9000-TRACE 1310 GC–MS was used for targeted metabolomics based on multiple reaction monitoring (MRM) analysis. The oven programming conditions and chromatographic column were identical to those used in the untargeted analysis. In this work, ionization was employed under the same conditions. By contrast, the mass spectrometry parameters were specifically optimized for each targeted compound, where the solvent delay time was 7.2 min and the scan speed was 0.2 s per spectrum. Specifically, we first screened high-intensity precursor ions within defined retention time windows, followed by automated selection of product ions and systematic optimization of collision energies in AutoSRM mode. Detailed parameter settings were shown in Table S2, and the parameter settings for Chameleon data processing software was shown in Table S3.
3. Methodology
The workflow of the developed untargeted to targeted transition strategy is illustrated in Fig. 1, which mainly consists of three sequence steps: i) compound resolution, time-shift correction, and registration, as well as underlying biomarker screening and identification; ii) construction of the GC–MS/MS-based targeted analysis method; and iii) development of the geographical discrimination model of Goji berries using chemometric methods. The success of the developed strategy relays on two factors, an accurate components resolution in the first step and the extension of identified compounds in the second step.
Fig. 1.
Workflow of the developed untargeted to targeted transition strategy.
3.1. Compound resolution, time-shift correction and registration, and underlying biomarker screening and identification (Step 1 in Fig. 1)
The raw GC–MS data were imported into AntDAS-GCMS software for TIC peak detection, and a Gaussian-based smoothing algorithm was used to detect peaks in the TIC (Fu et al., 2016). The peaks with a signal-to-noise (s/n) ratio above 10 were used for the following component resolution. Extracted ion chromatogram (EIC) peaks were also extracted using the Gaussian smoothing algorithm, and component resolution was performed separately for each TIC peak. First, an EIC peak clustering strategy was used to classify the EICs into a series of clusters, and a respective chromatogram was calculated for each cluster. Then, a multivariate curve resolution-alternative least squares (MCR-ALS) approach was introduced for component resolution. The retrieved mass spectra of each component were then imported into the National Institute of Standards and Technology (NIST) database for compound identification.
An example of component resolution and identification is shown in Fig. 2, with the TIC peaks from AntDAS-GCMS shown in Fig. 2a, indicating that 243 TIC peaks were detected. The inset plots in Fig. 2a present the peak detection results in detail, revealing that TIC peaks could be reasonably detected, and the range of eluted components was properly estimated. Component resolution was illustrated with three TIC peaks, as shown in Fig. 2b, namely, 131#, 132#, and 133#. Component resolution from AntDAS-GCMS suggested that five components could be retrieved, corresponding to Com.205 to Com.209. The resolved mass spectra were imported into the NIST database for compound identification, and finally, four components were identified. The resolved mass spectra of Com.206, Com.207, Com.208, and Com.209 are shown in Fig. 2c.
Fig. 2.
GC–MS data analysis in AntDAS. (a) The detected peaks in TIC of a Goji berry sample by GC–MS. (b) TIC peak resolution in AntDAS-GCMS. The left plot shows three TIC peaks; the middle plot shows five compounds resolved from the three TIC peaks, and the right plot shows the identification list of five compounds. (c) Compound identification results for the resolved compounds. (d) Results of the manual identification of compounds.
Traditional compound identification was implemented through a manual operation, as shown in Fig. 2d, and the original mass spectra at the peak apexes were imported into the NIST database for compound identification. In this case, only three components were resolved. Comparison with AntDAS-GCMS clearly indicated that two components were missing. In addition, the coeluted components reduced the quality of compound identification. For example, the mass spectra at the third peak apex were influenced by the coeluted components and provided a match factor (MF) and revised match factor (RMF) of 729 and 755, respectively. By contrast, the two resolved components from AntDAS-GCMS achieved acceptable identification results, with Com.208 obtaining MF and RMF values of 759 and 908, respectively. These results suggest that AntDAS-GCMS is suitable for data analysis to screen underlying components.
A time shift correction was performed to align the retention times of the resolved components that corresponded to the same compound at almost identical positions (Ma et al., 2024). The correction procedure was first performed by screening compounds present in the analyzed samples, with the mass spectra used to construct a spectral similarity matrix. The time shift correction was then transformed into an optimization problem that extracted a maximum path from the similarity matrix. A modified dynamic programming algorithm was subsequently introduced to extract the path, and the nodes in the maximum path served as benchmarks for aligning components. Component registration was performed by employing the mass spectra and modified retention times of the resolved components. Finally, a component registration table was created with each row corresponding to a component that was distributed in the analyzed samples, with each column corresponding to all resolved components within the sample.
Biomarker screening was performed according to the registered component list table, and analysis of variance (ANOVA) was used to screen metabolites that showed significant differences among various geographical zones. Compound identification was subsequently performed by importing the mass spectra of the screened metabolites into the NIST database. The identified components were subsequently used to construct targeted analytical methods to accurately quantify targeted compounds.
3.2. Construction of GC–MS/MS-based targeted analysis method (Step 2 in Fig. 1)
Unlike the classic methodology which employed the identified valuable compounds from untargeted metabolomics as targeted ones for constructing GC–MS/MS method, the developed strategy will treat the identified compound from untargeted metabolomics as initialized seeds to search for the other underlying compounds that may be useful for geographical discrimination. There were different methods of selecting valuable compounds, like finding the other metabolites in the same pathway of the identified ones and/or employing molecular network analysis based on mass spectra. In this work, we employed the ideology behind both methods was employed, which was performed by empirically selecting compounds with similar structures. To accomplish this goal, we fist construct a compound library with standard compounds by using component identification from AntDAS-GCMS (Fig. 1). Compounds which obtained similar chemical structures with the identified ones was selected as additional targeted ones for GC–MS/MS method.
Finally, the screened compounds from untargeted metabolomics and the ones from additional selection were used for constructing a GC–MS/MS method to accurately quantify their contents in Goji berry samples. Parameters including limit of detection (LOD), the limit of quantification (LOQ), and linearity of each compound were obtained with standards. The contents of all targeted compounds were used for constructing geographical discrimination model.
3.3. Geographical discrimination model of Goji berries using chemometric methods (Step 3 in Fig. 1)
Geographical discrimination was performed using supervised chemometric methods such as partial least squares discriminant analysis (PLS-DA) and unsupervised chemometric methods such as principal component analysis (PCA). The Monte-Carlo simulation was then used to validate the constructed model, and the best model was selected.
4. Results and discussion
We first performed untargeted metabolomics with a limited number of samples to screen the underlying components, followed by targeted metabolomics for large-scale samples. In this case, the analytical task was efficiently accomplished and the underlying metabolites were reasonably screened for accurate quantification. Thus, eight samples from each geographical zone were selected to obtain a small number of samples for untargeted GC–MS data analysis.
4.1. Biomarker screening based on untargeted metabolomics
Initially, untargeted metabolomics analysis was conducted using GC–MS to acquire the data of all metabolites present in Goji berries. After analyzing the Goji berry GC–MS data using AntDAS-GCMS software, a 2248 × 24 compound registration list was obtained, with 2248 representing the number of semi-volatile compounds and 24 representing the number of samples. The large number of registered components indicated an extreme task for screening valuable metabolites. To ensure the stability and reliability of the GC–MS system during data acquisition, eight quality control (QC) samples (QC1–QC8) were inserted throughout the analytical sequence. A control chart was constructed based on the TIC peak areas of the QC samples. As shown in Fig. S1a, all QC measurements remained within ±2 standard deviations, indicating that the GC–MS system maintained stable signal output during the entire run. This level of instrumental stability ensures the reliability of the metabolite quantification and supports the robustness of the subsequent chemometric modeling. Because the composition of metabolites in the same plant matrix, the Goji berry, was almost identical, except for the intensity variation across different zones, the distribution characteristics of the registered components were studied, with the results shown in Fig. S1b. Most components were only distributed within a limited number of samples. For example, 82% components were detected in 20–60% of samples, while 100% components were found in at least 80% samples. In this work, compounds that could be detected by at least 80% of samples were selected. ANOVA was then applied to identify significantly different metabolites, using a confidence level of p < 0.05. Finally, 219 different compounds were identified.
The quality of the screened metabolites was then evaluated using chemometric clustering methods. First, unsupervised hierarchical clustering analysis (HCA) was performed, with the results shown in Fig. S1c. The HCA results suggested that samples from different zones could be clearly distinguished. Among the three geographical zones, samples from Ningxia and Gansu provinces exhibited greater similarity to each other but were distinctly different from those from Xinjiang province. This resemblance between Ningxia and Gansu may be attributed not only to their geographical proximity but also to their comparable environmental conditions in 2023. Specifically, both regions experienced similar annual mean temperatures (approximately 8–12 °C) and relatively low precipitation levels, with Ningxia receiving around 20 mm of annual rainfall and Gansu showing representative values of 50–300 mm depending on local terrain. Both provinces also rely heavily on irrigation from the Yellow River, resulting in comparable water quality and soil moisture conditions during the growing season. These shared climatic and irrigation characteristics, together with similar cultivation practices, likely contributed to the similarity observed in their metabolite profiles. In contrast, Xinjiang experienced record-breaking high temperatures and historically low precipitation in 2023, with typical annual rainfall ranging from 50 to 150 mm across most areas, reflecting its more diverse and arid ecological zones. These distinct environmental conditions may explain the clear separation of Xinjiang samples in the HCA analysis.
Sample clustering results from PCA are presented in Fig. S1d, focusing on the sample distribution of the first two principal components. As shown in Fig. S1d, the first and second principal components explained 50.5% and 16.9% of the data variance, respectively. Samples distributed within the first two principal components could be clearly separated into three clusters within the 95% confidence ellipses. Specifically, samples from Xinjiang province were distributed on the right side of the first principal component, while the other samples were located on the opposite side. These results suggested that Goji berries from Xinjiang province were significantly different from the two remaining zones, consistent with the HCA results. Focusing on the sample distribution results of Ningxia and Gansu provinces, we found that samples could be approximately separated along the second principal component. Specifically, samples from Ningxia province were distributed on the positive axis, while those from Gansu province were distributed on the negative axis, suggesting a clear separation from PCA. The sample clustering results from HCA and PCA suggested that Goji berries from different zones could be separated with the screened components. Further investigation of the PCA results indicated that removing a portion of variance from the raw data did not affect sample clustering, suggesting that there was no need for selecting all screened metabolites for sample clustering. Thus, we focused on the screened metabolites that could be identified using the NIST database.
The 219 screened metabolites were imported into the NIST database for compound identification. In this work, the compound identification criteria were as follows: i) a retention index between the calculated value and NIST value of less than 40; ii) an MF and RMF value from the NIST library above 700. Finally, 55 reliable compounds were identified, and detailed information including the retention time, MF, RMF, and CAS was summarized, as shown in Table S4. The 55 compounds were briefly classified according to their chemical structures, including 26 sugars, 10 amino acids, 10 organic acids, 4 esters, 1 pyridine, 1 pyrrole, 1 pyrone, 1 alcohol, and 1 quinoline (Fig. 3a). Notably, many of these metabolites have been reported as key biomarkers associated with plant stress responses, carbon–nitrogen metabolism, and fruit quality formation, suggesting their potential relevance to geographical origin discrimination (Liu et al., 2023; Ramírez-Meraz et al., 2024; Zhang et al., 2022).
Fig. 3.
(a) The number of compounds in each class. (b) The area proportion of compounds in each class. (c) Principal component analysis (PCA) results based on the identified 55 compounds. (d) Partial least squares discriminant analysis (PLS-DA) results based on the identified 55 compounds.
Among the identified metabolites, amino acids such as proline, threonine, aspartic acid, glutamic acid, alanine, and asparagine are well-known indicators of environmental stress and nitrogen assimilation pathways (Dan et al., 2021). Previous studies have shown that variations in amino acid pools can reflect differences in soil nitrogen availability, temperature fluctuations, and water stress, all of which differ significantly among Ningxia, Gansu, and Xinjiang. Sugars and sugar derivatives constituted the largest group of differential metabolites, including glucose, fructose, galactose, mannitol, ribose, xylose, arabinose, lactose, maltose, turanose, palatinose, and several sugar alcohols and lactones. Carbohydrate metabolism is highly sensitive to climatic conditions, and previous reports have demonstrated that sugar composition can serve as a robust indicator for geographical traceability in fruits and medicinal plants (Lima et al., 2022; Wu et al., 2025). The presence of characteristic sugars such as allofuranose, tagotofuranose, and galactinol further highlights the metabolic specificity associated with regional ecological factors.
In addition, organic acids including malic acid, lactic acid, glycolic acid, glyceric acid, and butanoic acid are central intermediates in primary metabolism and have been widely used as chemotaxonomic markers in fruit authentication (Luo et al., 2025). Their distinct abundance patterns across regions may reflect differences in sunlight exposure, diurnal temperature variation, and irrigation water sources. In summary, the diversity and functional relevance of these 55 differential metabolites indicate strong discriminatory potential for distinguishing the geographical origins of Goji berries. Their metabolic roles and region-specific abundance patterns provide a biochemical basis for origin traceability and support the robustness of the chemometric classification results.
The content of each group was measured by the peak area, with the results shown in Fig. 3b. Sugar occupied the largest proportion of all compounds, with a value of 79.2%. The second largest content group was organic acid, with a proportion of 18.4%, which was one-quarter of the sugar content. The remaining compound groups had much lower abundances, collectively accounting for 2.4% of the total. For example, amino acids were the most abundant, comprising only 0.7% of the total.
The quality of sample clustering results based on the identified 55 compounds is shown in Fig. 3c and d, corresponding to PCA and PLS-DA. The PCA results indicated a serious overlap of analyzed samples, and the samples from the three zones could not be separated. We subsequently investigated the performance of the supervised clustering method, PLS-DA, and the results in Fig. 3d suggested that samples from the three cultivation zones could be clearly separated. The results also indicated that the identified samples could not be satisfactorily separated from the identified 55 compounds in an unsupervised manner, implying a number of key compounds that could not be identified. Therefore, combined supervised chemometric methods such as PLS-DA could provide an efficient strategy for classifying Goji berries; thus, PLS-DA was employed for the following large-scale data analysis.
4.2. Targeted analysis using identified compounds
MRM-based targeted metabolomics was performed using the identified compounds. Among the preliminary screened 55 compounds, 18 compounds were confirmed with standards and selected as targeted compounds for MRM analysis. Given that most of the identified compounds were sugars and amino acids, an additional 26 compounds were included to expand the targeted metabolomics, comprising 6 sugars and 20 amino acids. Finally, the performance of the developed MRM-based targeted metabolomics was investigated using a routine analysis conducted at the Ningxia Food Testing Institute (NXFTI) to accurately classify Goji berry samples from different geographical zones. A total of 186 samples were collected in this dataset, including 50 samples from Ningxia, 70 samples from Gansu, and 66 samples from Xinjiang. The following section details the development and validation of the MRM method.
4.2.1. Construction and validation method for MRM-based targeted metabolomics
The precursor and product ions for each compound were optimized by a Thermo GC–MS/MS instrument, as summarized in Table S5, with the chromatogram profiles for each compound shown in Fig. S2. The linearity of each compound was first investigated, and the concentration of each compound was designed based on the peak area in the analyzed samples, as shown in Table S6. The linearity of each compound was calculated using the least squares method using designed concentration levels and peak areas. As shown in Table S6, all 44 compounds obtained high linearity values, with the Pearson coefficients exceeding 0.99.
To assess the accuracy of the constructed MRM method, a recovery study was performed by spiking actual samples with known concentrations of standard compounds. The recovery of each compound was calculated by comparing the predicted concentrations from the linear prediction model in the table with the actual values. As shown in Table S6, the recovery values for all compounds were located in the range of 80–120%, indicating that acceptable accuracy performance could be obtained (Tan et al., 2019).
The precision was evaluated by determining the intra-day and inter-day variations of the targeted compounds. Intra-day precision was assessed using three replicate analyses within a single day, while inter-day precision was assessed over three consecutive days, with the precision expressed as relative standard deviation (RSD%). We observed that all analyzed compounds exhibited both intra-day and inter-day RSD values below 6%, indicating acceptable precision for routine analysis. The limit of detection and the limit of quantification were determined according to the signal-to-noise (S/N) ratios, using values of 3 and 10, respectively. Table S6 summarizes the calculated LOD and LOQ values for each compound, which covered the concentration range of the analyzed samples (Xie et al., 2023).
4.2.2. Quantification of the targeted compounds using MRM-based targeted metabolomics
The content variations of the 44 compounds within each cultivated zone are provided in Table S7. The detailed information was shown in Supplementary data 1, along with a scatter plot and bar chart for each compound, as shown in Fig. 4. All quantitative results are reported as mean ± standard deviation (SD), where the SD represents the intra-regional variability among Goji berry samples collected from Ningxia (n = 50), Gansu (n = 70), and Xinjiang (n = 66). The SD values listed for each metabolite correspond to the dispersion of its content across all samples within the same production region. The relatively large SD values observed for certain metabolites reflect the natural biological variability within each region, which may be influenced by cultivation practices, microclimatic conditions, and post-harvest handling. ANOVA was employed to provide a statistical analysis for the targeted compounds, and the results suggested that 40 compounds showed significant differences at p < 0.05 (except for 1,6-anhydro-β-glucose, arginine, gluconic acid, and sucrose). An additional investigation indicated that four compounds contained the largest content values from Ningxia, 26 compounds contained the largest values from Gansu, and 10 compounds contained the largest values from Xinjiang province.
Fig. 4.
The scatter plot and bar plot of the content distribution of 44 compounds in Goji berry samples from Ningxia, Gansu and Xinjiang provinces (NX: Ningxia, GS: Gansu, XJ: Xinjiang).
To explore the impact of geographical environments on the distribution of sugars and amino acids in Goji berries and identify key metabolic markers for origin authentication, a comprehensive correlation network analysis was conducted on 44 target metabolites across different production regions (Fig. 5). The results revealed high metabolic differences among regions at p < 0.05, reflecting the influence of soil nutrients, climate, and ecological factors on plant metabolism.
Fig. 5.
Comprehensive correlation network analysis of 44 target metabolites across different production regions.
In Goji berries from Ningxia, nitrogen metabolism-related amino acids (e.g., alanine, aspartic acid, and glutamic acid) and osmolytes (e.g., mannitol, hydroxyproline) were highly enriched, suggesting that metabolic adaptation in this region was primarily influenced by soil nitrogen availability and drought stress, with a Pearson coefficient exceeding 0.7. By contrast, Goji berries from Gansu exhibited the most diverse metabolic characteristics, including multiple essential amino acids and sugars, indicating a complex metabolic regulation strategy to address varying ecological conditions. Meanwhile, in Xinjiang, sugar metabolism emerged as the predominant metabolic feature, with a high accumulation of fructose (r = 0.9), galactose (r = 0.7), and other carbohydrates. This pattern was likely influenced by intense solar radiation, large diurnal temperature variations, and an arid climate, which collectively affected carbon metabolism and sugar storage (Chá et al., 2024; Feng et al., 2022; Sun et al., 2024).
Furthermore, metabolites such as valine, mannitol, hydroxyproline, glutamic acid, and aspartic acid were found to be correlated across multiple regions at p < 0.05, suggesting their potential as stable geographic markers for metabolic fingerprint-based origin discrimination. The distinct regional variations in sugar and amino acid metabolism highlighted their value as key metabolic markers for Goji berry authentication, providing a scientific basis for metabolite-based geographical traceability and a robust approach for origin verification.
Because the target components mostly consisted of sugars and amino acids, the total amino acid (TAA) and total sugar (TS) contents of the Goji berry samples from Ningxia, Gansu, and Xinjiang provinces were quantitatively analyzed (Fig. S3), revealing significant differences in their accumulation across different geographical origins. The TAA content was the highest in the Gansu samples (588.27 μg/g), which was significantly greater than in the Ningxia (110.79 μg/g) and Xinjiang (116.40 μg/g) samples at p < 0.05.
Additional analysis indicated that the TAA composition in Gansu Goji berries was predominantly contributed by proline and pyroglutamic acid, accounting for 64.5% of the total amino acid content. By contrast, arginine and pyroglutamic acid were the primary contributors to TAA in the Ningxia and Xinjiang Goji berries, comprising 82.3% and 88.4% of their total amino acid content, respectively. The amino acid and sugar metabolism in the Goji berries showed distinct regional characteristics. The Gansu samples exhibited significantly higher TAA content, approximately five times that of Ningxia and Xinjiang, likely due to fertile, nitrogen-rich soil, moderate precipitation, and substantial diurnal temperature variations that enhanced amino acid biosynthesis.
Regarding TS content, the Xinjiang samples exhibited the highest accumulation (41,347.86 μg/g), followed by those from Gansu (36,493.39 μg/g), while the Ningxia samples had the lowest content (5508.37 μg/g). Notably, sucrose, galactose, glucose, fructose, and mannose were the predominant contributors to TS accumulation across all three regions, accounting for over 90% in Ningxia, Gansu, and Xinjiang. By contrast, the Xinjiang samples had the highest total sugar (TS) content (41,347.86 μg/g), surpassing the Gansu (36,493.39 μg/g) and Ningxia (5508.37 μg/g) (p < 0.05) samples. This was possibly due to intense solar radiation and large diurnal temperature fluctuations, which enhanced carbon assimilation while reducing nighttime respiratory consumption. These results were consistent with the findings obtained by Ryu, Glew, and Gong et al. (Glew et al., 2003; Gong et al., 2022; Ryu et al., 2020).
Variability analysis further demonstrated that amino acid content exhibited the highest fluctuations in the Xinjiang samples (CV = 123.20%), followed by Gansu (98.89%) and Ningxia (54.94%), while the sugar content variability was most pronounced in the Ningxia (100.00%) samples, with the Gansu (27.38%) and Xinjiang (26.35%) samples exhibiting relatively stable sugar levels. These distinct metabolic patterns across different regions highlighted the potential of TAA and TS profiles as biochemical markers for geographical traceability. The pronounced variability in amino acid and sugar accumulation, influenced by soil nutrient composition, climatic conditions, and genetic backgrounds, suggested that integrating metabolic fingerprinting with environmental and genetic data could enhance the accuracy of origin authentication models (Lu et al., 2025). These findings provided a scientific basis for the establishment of a robust traceability system for wolfberries, facilitating industry quality control and geographical authentication.
4.2.3. Geographical discrimination using chemometric models
We subsequently investigated the performance of several classic chemometric models on the geographical discrimination of Goji berries. Four algorithms were selected, including SVM, k-means, Fisher, and PLS-DA (T. Wang et al., 2009). Monte-Carlo simulations were conducted with 10,000 iterations. In each simulation, 80% of the samples were randomly selected as the validation dataset for model construction, while the remaining 20% of the samples served as the test dataset. Fig. 6a presents the accuracy of the four methods in terms of validation and prediction. The parameter settings for chemometric model were shown in Table S8. The accuracy values of the validation set were 99.0%, 100.0%, 99.9%, and 99.4% for SVM, k-means, Fisher, and PLS-DA, respectively, showing consistent performance across methods. However, when evaluating the prediction accuracy, SVM, k-means, Fisher, and PLS-DA were 97.4%, 75.7%, 96.0%, and 97.8%, respectively. These results suggested that SVM and PLS-DA provided the best performance among the four algorithms. Thus, PLS-DA was selected for this work. The superior performance of PLS-DA can be attributed to its algorithmic characteristics and the structure of the metabolomics dataset. The targeted metabolite profiles contain a large number of correlated variables, and PLS-DA is particularly effective in handling multicollinearity by projecting predictors into latent variables that maximize covariance between metabolite levels and class labels. Unlike PCA, which maximizes overall variance without considering group information, PLS-DA incorporates class membership during model construction, enabling it to capture subtle but highly discriminative metabolic differences among geographical origins. Furthermore, Monte-Carlo cross-validation demonstrated that PLS-DA provided the most stable prediction accuracy, confirming its suitability for datasets where class separation relies on multivariate patterns rather than single high-variance features.
Fig. 6.
(a) The accuracy of the four methods in terms of both validation and prediction. (b) Partial least squares discriminant analysis (PLS-DA) results based on 44 target metabolites. The top-left panel presents the score plot of the first and second latent variables, while the top-right panel displays the score plot of the first and third latent variables. The bottom-left panel shows the score plot of the second and third latent variables, and the bottom-right panel illustrates the 3D PLS-DA score plot. (c) The confusion matrix of 44 target metabolites across different production regions. (d) The confusion matrix of 18 target metabolites screened in untargeted metabolomics.
Geographical discrimination models were investigated using the collected samples, with the samples distributed on the first two latent variables of PLS-DA provided in Fig. 6b. Samples from different zones could be briefly separated with calculated ellipses. Goji berries from Ningxia province could be clearly separated from those from Gansu and Xinjiang provinces, as the ellipses corresponding to Ningxia province were distributed farther away from the other two zones. The calculated ellipses of Gansu and Xinjiang provinces were relatively close to each other. When examining sample clustering along the second and third latent variables of PLS-DA, samples from Xinjiang and Gansu partially overlapped. By contrast, the samples from Ningxia remained distinctly separated from Gansu and Xinjiang provinces. In addition, samples from Ningxia province were scattered within the smallest zone in the clustering zone, likely due to the province's limited cultivation region, primarily centered in Zhongning city. By contrast, Goji berries cultivated in Gansu and Xinjiang provinces were across much larger areas, resulting in larger ellipses in Fig. 6b. These results indicated that a geographical model could be constructed to accurately discriminate Goji berry samples from different cultivated zones.
A PLS-DA model was constructed using custom-developed MATLAB code, and Monte-Carlo simulations were used to investigate the performance of PLS-DA (Hu et al., 2025). Specifically, 106 samples were divided into two subsets: 80% samples from each geographical zone were selected as the validation set for constructing the PLS-DA model, and the remaining 20% of the samples served as test samples. A mean center treatment was employed, and the constructed geographical model was used for predicting the origins of the test samples. This procedure was repeated 10,000 times. The confusion matrix for all repeats is statistically shown in Fig. 6c, indicating that the prediction accuracy values of Ningxia, Gansu, and Xinjiang provinces were 96.3%, 97.8%, and 96.6%, respectively. The results suggested that most of the samples could be precisely predicted using the PLS-DA model. A further investigation of the confusion matrix suggested that a very small fraction of Ningxia Goji berries (3.4%) was incorrectly identified as Gansu berries, and only 0.3% were misclassified as Xinjiang berries. Notably, even though most of the samples (96.6%) from Xinjiang province were accurately classified, a small portion (3.4%) was misclassified as originating from Gansu province, and none of the samples were identified as from Ningxia. These results suggested that Goji berries from Ningxia and Xinjiang provinces were relatively different from each other, consistent with the sample clustering results in Fig. 6b.
Metabolites with VIP (Variable Importance in Projection) values greater than 1.0 were considered key contributors to the PLS-DA model and used to interpret origin discrimination (Table S9). These variables were mainly carbohydrates, polyols, and a limited number of amino acids–related metabolites, indicating that primary metabolism dominated the classification. Cellobiose showed the highest VIP value (VIP = 1.88), suggesting a strong association with origin-dependent metabolic variation. As a cellulose-derived disaccharide, its variation likely reflects differences in structural carbohydrate turnover during fruit development. Glucose (VIP = 1.83) and fructose (VIP = 1.42) were the most influential soluble sugars, consistent with differences in carbon assimilation and allocation under distinct environmental conditions. Among polyols, threitol (VIP = 1.77) and mannitol (VIP = 1.36) contributed substantially to sample separation, indicating the involvement of osmotic and stress-related regulation. In addition, several amino acids–related metabolites, including tyrosine (VIP = 1.19), histidine (VIP = 1.07), and citrulline (VIP = 1.03), exceeded the VIP threshold. Overall, origin discrimination was driven by coordinated variation in carbohydrate, polyol, and nitrogen-related metabolism rather than by a single marker.
4.2.4. A comparison with classical untargeted-to-targeted method
The key difference between our MRM-based targeted metabolomics and the classic untargeted-to-targeted method was that the latter focused on compounds that were identified compounds from the results of untargeted metabolomics. To compare these approaches, we constructed another PLS-DA model on the basis of the 18 compounds that were screened in the untargeted metabolomics analysis, with the results shown in Fig. 6d. Similarly, a majority of samples from each zone could be accurately classified, as Goji berry samples from Ningxia, Gansu, and Xinjiang provinces achieved accuracy values of 91.2%, 94.6%, and 92.1%, respectively. A comparison of the results in Fig. 6c implied that the accuracy values of all cultivated zones, i.e., Ningxia, Gansu, and Xinjiang provinces, decreased. In this case, about 7.9% of the samples from Xinjiang province were inaccurately classified as from Gansu province, and 4.0% of the samples from Gansu province were recognized as from Xinjiang province. In targeted metabolomics, advanced data analysis tools are often used to retrieve and screen key metabolites; however, this approach can overlook compounds that cannot be resolved or identified using NIST libraries or reference standards. This comparison indicated that the introduction of additional compounds could greatly improve the total performance of the PLS-DA model, especially for Gansu and Xinjiang provinces. However, the introduction of compounds with similar chemical structures or the same biological pathway was valuable for improving metabolite coverage and enhancing the accuracy of geographical discrimination models.
4.3. Discussion
Untargeted and targeted metabolomics have been extensively employed for the geographical discrimination of food samples. The employment of untargeted metabolomics has frequently utilized data analysis tools such as AMDIS (Acosta-García et al., 2025), AntDAS-GCMS (Y. Fan et al., 2023), and MS-DIAL (Xiaqiong Fan et al., 2022). However, one issue is that the majority of screened compounds from untargeted metabolomic cannot be identified using NIST, Wiley, or in-house libraries. This may be due to i) the limited coverage of compounds in these libraries and ii) the low quality of resolved mass spectra of compounds. In this case, identified compounds must be accurately quantified to further confirm their contribution to geographical discrimination. Unfortunately, only a small fraction of studies has employed targeted methods such as MRM for confirmation. The findings from this work suggested that identified compounds may not always yield optimal geographical discrimination results compared with using the entire set of screened metabolites. Instead, it would be beneficial to incorporate additional valuable compounds to construct a geographical model for Goji berries. Since identifying valuable compounds for targeted metabolomics is somewhat of a black box, it may be beneficial to find the compounds with similar chemical structures or those involved in the same biological pathways as the identified compounds, whose efficiency was confirmed with this work.
Although untargeted metabolomics is often promoted as providing a comprehensive overview of metabolites in analyzed samples, its scope remains limited by the analytical technique employed. In this study, the sample treatment method provided a comprehensive analysis of semi-volatile compounds in the analyzed samples, whose volatility could be improved after derivatization, making it suitable for analysis using GC–MS. According to the identified compounds, most of the compounds consisted of sugars and amino acids. A limitation of the derivatization strategy was that sugars and amino acids produced similar mass spectra, leading to challenges in compound identification. Analysts frequently encounter issues where several candidates match in the NIST database with high MFs and RMFs. To address this, we confirmed targeted compounds using standards before performing downstream analyses such as MRM.
In conclusion, substantial differences in the sugar and amino acid content were found among Goji berries from different regions. These differential markers were likely influenced by the specific climate and soil conditions of each area, which, in turn, affected the overall quality and functional properties of the berries. Our findings not only promoted the understanding of the distribution of chemical components in Goji berries but also provided crucial evidence for geographical origin analysis. Based on the quantitative data obtained for amino acids and sugars, origin discrimination analysis was performed. The results highlighted distinct characteristics in the amino acid and sugar profiles of Goji berries from different regions, making them effective indicators for origin identification. In addition, the strategy presented in this work could provide a new solution for the quality control of food samples.
5. Conclusions
This study proposed an untargeted to targeted transition strategy, which was used to achieve an efficient conversion from untargeted metabolomics to targeted metabolomics. The AntDAS-GCMS method was employed to perform biomarker screening on a small-scale set of samples based on untargeted metabolomics, and compounds were screened and identified. Subsequently, compounds with similar chemical structures or homologous metabolic pathways identified were introduced, and an MRM-based GC–MS/MS quantification method covering 44 target compounds was developed. The targeted metabolomic method was then used to analyze the large-scale set of samples, which revealed that compared with the PLS-DA model constructed with 18 screened compounds from the untargeted approach, the prediction accuracy improved from 91.2%, 94.6%, and 92.1% for the Ningxia, Gansu, and Xinjiang samples to 96.3%, 97.8%, and 96.6%, respectively, significantly enhancing the accuracy and reliability of the constructed geographical discrimination model. This strategy provided a new approach for the transition from untargeted to targeted metabolomics, offering strong support for food traceability and quality control research.
CRediT authorship contribution statement
Long-He Wang: Writing – original draft, Validation, Investigation, Data curation. Yan-Jin Wen: Writing – review & editing, Supervision. Wen-Xin Wang: Writing – review & editing, Supervision. Meng Zhai: Writing – review & editing, Supervision. Shu-Fang Li: Writing – review & editing, Supervision. Qing-Xia Zheng: Writing – review & editing. Ping-Ping Liu: Writing – review & editing. Yao Zhang: Writing – review & editing. Yi Lv: Writing – review & editing. Hui-Na Zhou: Writing – review & editing. Yong-Jie Yu: Writing – review & editing, Methodology, Funding acquisition, Conceptualization.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgments
This study was supported by the National Natural Science Foundation of China (22378214), the Key Scientific and Technological Project of Henan Province (232102310334), the Chief Scientist Innovation Project of State Tobacco Monopoly Administration/China National Tobacco Corporation (932025CK0500), the Research project of State Administration for Market Regulation (2023MK125), the Key Research and Development Program of Ningxia Province (2025BEG02027, 2022BEG03170), the Natural Science Foundation of Ningxia Province (2023AAC05038 and 2023AAC03253), and the Foundation of Ningxia Medical University (XJKF230202 and XM2020021).
Footnotes
Supplementary data to this article can be found online at https://doi.org/10.1016/j.fochx.2026.103549.
Contributor Information
Ping-Ping Liu, Email: Liu_pingping2012@163.com.
Yong-Jie Yu, Email: yongjie.yu@163.com.
Appendix A. Supplementary data
Data availability
No data was used for the research described in the article.
References
- Acosta-García E.D., Páez-Lerma J.B., Martínez-Prado M.A., Soto-Cruz N.O. Volatile compound analysis in mezcal based on multiple extraction/concentration methods, deconvolution software, and multivariate analysis. Food Control. 2025;168 doi: 10.1016/j.foodcont.2024.110852. [DOI] [Google Scholar]
- Bertoldi D., Cossignani L., Blasi F., Perini M., Barbero A., Pianezze S., Montesano D. Characterisation and geographical traceability of Italian goji berries. Food Chemistry. 2019;275:585–593. doi: 10.1016/j.foodchem.2018.09.098. [DOI] [PubMed] [Google Scholar]
- Chá L.C., Ressurreição S., Oliveira L., Santos S., Nunes M., Vidal M.…Gomes F. Sugar content in Arbutus unedo L. fruit and its relationship with climatic and edaphic characteristics. Plants. 2024;13(23):3383. doi: 10.3390/plants13233383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dan Z., Chen Y., Li H., Zeng Y., Xu W., Zhao W., He R., Huang W. The metabolomic landscape of rice heterosis highlights pathway biomarkers for predicting complex phenotypes. Plant Physiology. 2021;187(2):1011–1025. doi: 10.1093/plphys/kiab273. [DOI] [PMC free article] [PubMed] [Google Scholar]
- De Lima Morais da Silva P., De Lima L.S., Caetano Í.K., Torres Y.R. Comparative analysis of the volatile composition of honeys from Brazilian stingless bees by static headspace GC-MS. Food Research International. 2017;102:536–543. doi: 10.1016/j.foodres.2017.09.036. [DOI] [PubMed] [Google Scholar]
- Fan X., Jiao X., Liu J., Jia M., Blanchard C., Zhou Z. Characterizing the volatile compounds of different sorghum cultivars by both GC-MS and HS-GC-IMS. Food Research International. 2021;140 doi: 10.1016/j.foodres.2020.109975. [DOI] [PubMed] [Google Scholar]
- Fan X., Xu Z., Zhang H., Liu D., Yang Q., Tao Q., Wen M., Kang X., Zhang Z., Lu H. Fully automatic resolution of untargeted GC-MS data with deep learning assistance. Talanta. 2022;244 doi: 10.1016/j.talanta.2022.123415. [DOI] [PubMed] [Google Scholar]
- Fan Y., Bai X., Chen H., Yang X., Yang J., She Y., Fu H. A novel simultaneous quantitative method for differential volatile components in herbs based on combined near-infrared and mid-infrared spectroscopy. Food Chemistry. 2023;407 doi: 10.1016/j.foodchem.2022.135096. [DOI] [PubMed] [Google Scholar]
- Feng Y., Fan X., Zhang S., Wu T., Bai L., Wang H., Ma Y., Guan X., Wang C., Yang H. Effects of variety and origin on the metabolic and texture characteristics of quinoa seeds based on ultrahigh-performance liquid chromatography coupled with high-field quadrupole-orbitrap high-resolution mass spectrometry. Food Research International. 2022;162 doi: 10.1016/j.foodres.2022.111693. [DOI] [PubMed] [Google Scholar]
- Fu H.-Y., Guo J.-W., Yu Y.-J., Li H.-D., Cui H.-P., Liu P.-P.…Lu P. A simple multi-scale Gaussian smoothing-based strategy for automatic chromatographic peak extraction. Journal of Chromatography A. 2016;1452:1–9. doi: 10.1016/j.chroma.2016.05.018. [DOI] [PubMed] [Google Scholar]
- Gamal A., Soliman M., Al-Anany M.S., Eissa F. Optimization and validation of high throughput methods for the determination of 132 organic contaminants in green and roasted coffee using GC-QqQ-MS/MS and LC-QqQ-MS/MS. Food Chemistry. 2024;449 doi: 10.1016/j.foodchem.2024.139223. [DOI] [PubMed] [Google Scholar]
- Glew R.H., Ayaz F.A., Sanz C., VanderJagt D.J., Huang H.S., Chuang L.T., Strnad M. Changes in sugars, organic acids and amino acids in medlar (Mespilus germanica L.) during fruit development and maturation. Food Chemistry. 2003;83(3):363–369. doi: 10.1016/S0308-8146(03)00097-9. [DOI] [Google Scholar]
- Gong H., Rehman F., Li Z., Liu J., Yang T., Liu J.…Wang Y. Discrimination of geographical origins of wolfberry (Lycium barbarum L.) fruits using stable isotopes, earth elements, free amino acids, and saccharides. Journal of Agricultural and Food Chemistry. 2022;70(9):2984–2997. doi: 10.1021/acs.jafc.1c06207. [DOI] [PubMed] [Google Scholar]
- He J., Wang T., Yan H., Guo S., Hu K., Yang X.…Duan J. Intelligent identification method of geographic origin for Chinese wolfberries based on color space transformation and texture morphological features. Foods. 2023;12(13):2541. doi: 10.3390/foods12132541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Heuckeroth S., Damiani T., Smirnov A., Mokshyna O., Brungs C., Korf A.…Pluskal T. Reproducible mass spectrometry data processing and compound annotation in MZmine 3. Nature Protocols. 2024;19(9):2597–2641. doi: 10.1038/s41596-024-00996-y. [DOI] [PubMed] [Google Scholar]
- Hu L., Wang Y., Wu X., Shan Y., Zhu F., Zhang F., Yang Q., Liu M. Geographic origin discrimination and quantification of phenolic compounds and moisture in Artemisia argyi folium using NIRS and chemometrics. Journal of Agriculture and Food Research. 2025;23 doi: 10.1016/j.jafr.2025.102295. [DOI] [Google Scholar]
- Huan T., Forsberg E.M., Rinehart D., Johnson C.H., Ivanisevic J., Benton H.P.…Siuzdak G. Systems biology guided by XCMS online metabolomics. Nature Methods. 2017;14(5):461–462. doi: 10.1038/nmeth.4260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ju Y., Liu H., Niu S., Kang L., Ma L., Li A., Zhao Y., Yuan Y., Zhao D. Optimizing geographical traceability models of Chinese Lycium barbarum: Investigating effects of region, cultivar, and harvest year on nutrients, bioactives, elements and stable isotope composition. Food Chemistry. 2025;467 doi: 10.1016/j.foodchem.2024.142286. [DOI] [PubMed] [Google Scholar]
- Kim Y.-K., Baek E.J., Na T.W., Sim K.S., Kim H., Kim H.J. LC–MS/MS and GC–MS/MS cross-checking analysis method for 426 pesticide residues in agricultural products: A method validation and measurement of uncertainty. Journal of Agricultural and Food Chemistry. 2024;72(41):22814–22821. doi: 10.1021/acs.jafc.4c01992. [DOI] [PubMed] [Google Scholar]
- Li Q., Yu X., Xu L., Gao J.-M. Novel method for the producing area identification of Zhongning Goji berries by electronic nose. Food Chemistry. 2017;221:1113–1119. doi: 10.1016/j.foodchem.2016.11.049. [DOI] [PubMed] [Google Scholar]
- Li S.-F., Guo X.-M., Hao X.-F., Feng S.-H., Hu Y.-J., Yang Y.-Q., Wang H.-F., Yu Y.-J. Untargeted metabolomics study of Lonicerae japonicae flos processed with different drying methods via GC-MS and UHPLC-HRMS in combination with chemometrics. Industrial Crops and Products. 2022;186 doi: 10.1016/j.indcrop.2022.115179. [DOI] [Google Scholar]
- Lima K.R.P., Cavalcante F.L.P., Paula-Marinho S.d.O., Pereira I.M.C., Lopes L.D.S., Nunes J.V.S.…Carvalho H.H.D. Metabolomic profiles exhibit the influence of endoplasmic reticulum stress on sorghum seedling growth over time. Plant Physiology and Biochemistry. 2022;170:192–205. doi: 10.1016/j.plaphy.2021.11.041. [DOI] [PubMed] [Google Scholar]
- Liu T., Qiao N., Ning F., Huang X., Luo L. Identification and characterization of plant-derived biomarkers and physicochemical variations in the maturation process of Triadica cochinchinensis honey based on UPLC-QTOF-MS metabolomics analysis. Food Chemistry. 2023;408 doi: 10.1016/j.foodchem.2022.135197. [DOI] [PubMed] [Google Scholar]
- Lu Y., Zhai R., Chu Z., Zhu M., Li J., Jiang Y., Ye Z. LC-MS/MS-based quantitative method and metrological traceability technology for measuring components of animal origin in beef and lamb and their products. Food Chemistry. 2025;464 doi: 10.1016/j.foodchem.2024.141600. [DOI] [PubMed] [Google Scholar]
- Luo Y., Zhang M., Zhu Z., Zhang W., Ding S., Lin X., Sun X., Ma T. Early harvesting is undesirable in kiwifruit production: Insights from metabolomics and physiological evidence. Food Chemistry: X. 2025;30 doi: 10.1016/j.fochx.2025.103003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma G.-M., Wang J.-N., Wang X.-C., Ma F.-L., Wang W.-X., Li S.-F.…She Y. AntDAS-GCMS: A new comprehensive data analysis platform for GC–MS-based untargeted metabolomics with the advantage of addressing the time shift problem. Analytical Chemistry. 2024;96(23):9379–9389. doi: 10.1021/acs.analchem.4c00100. [DOI] [PubMed] [Google Scholar]
- Ma X.-L., Wang X.-C., Zhang J.-N., Liu J.-N., Ma M.-H., Ma F.-L., Lv Y., Yu Y.-J., She Y. A study of flavor variations during the flaxseed roasting procedure by developed real-time SPME GC–MS coupled with chemometrics. Food Chemistry. 2023;410 doi: 10.1016/j.foodchem.2023.135453. [DOI] [PubMed] [Google Scholar]
- Matsueda M., Asai S., Watanabe A., Pipkin W., Teramae N., Ohtani H., Watanabe C. Quantitative analysis of acrylic acid in acrylic pressure-sensitive adhesives by reactive pyrolysis-GC/MS using N,O-bis(trimethylsilyl)trifluoroacetamide as a trimethylsilylation reagent with two-step heating. Journal of Analytical and Applied Pyrolysis. 2023;175 doi: 10.1016/j.jaap.2023.106170. [DOI] [Google Scholar]
- Miranda M.R., Basilicata M.G., Vestuto V., Aquino G., Marino P., Salviati E.…Manfra M. Anticancer therapies based on oxidative damage: Lycium barbarum inhibits the proliferation of MCF-7 cells by activating pyroptosis through endoplasmic reticulum stress. Antioxidants. 2024;13(6):708. doi: 10.3390/antiox13060708. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mourão R.S., Sanson A.L., Poleti Martucci M.E. HS-SPME-GC-MS combined with metabolomic approach to discriminate volatile compounds of Brazilian coffee from different geographic origins. Food Bioscience. 2023;56 doi: 10.1016/j.fbio.2023.103395. [DOI] [Google Scholar]
- Nirere A., Sun J., Kama R., Atindana V.A., Nikubwimana F.D., Dusabe K.D., Zhong Y. Nondestructive detection of adulterated wolfberry (Lycium Chinense) fruits based on hyperspectral imaging technology. Journal of Food Process Engineering. 2023;46(4) doi: 10.1111/jfpe.14293. [DOI] [Google Scholar]
- Ramírez-Meraz M., Méndez-Aguilar R., Zepeda-Vallejo L.G., Hernández-Guerrero C.J., Hidalgo-Martínez D., Becerra-Martínez E. Exploring the chemical diversity of Capsicum chinense cultivars using NMR-based metabolomics and machine learning methods. Food Research International. 2024;178 doi: 10.1016/j.foodres.2023.113796. [DOI] [PubMed] [Google Scholar]
- Rasekh M., Karami H., Kamruzzaman M., Azizi V., Gancarz M. Impact of different drying approaches on VOCs and chemical composition of Mentha spicata L. essential oil: A combined analysis of GC/MS and E-nose with chemometrics methods. Industrial Crops and Products. 2023;206 doi: 10.1016/j.indcrop.2023.117595. [DOI] [Google Scholar]
- Ryu M.-J., Kim M., Ji M., Lee C., Yang I., Hong S.-B.…Nam S.-J. Discrimination of Lycium chinense and L. barbarum based on metabolite analysis and hepatoprotective activity. Molecules. 2020;25(24):5835. doi: 10.3390/molecules25245835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stein S.E. An integrated method for spectrum extraction and compound identification from gas chromatography/mass spectrometry data. Journal of the American Society for Mass Spectrometry. 1999;10(8):770–781. doi: 10.1016/S1044-0305(99)00047-1. [DOI] [Google Scholar]
- Sun Q., Wu F., Wu W., Yu W., Zhang G., Huang X., Hao Y., Luo L. Identification and quality evaluation of Lushan Yunwu tea from different geographical origins based on metabolomics. Food Research International. 2024;186 doi: 10.1016/j.foodres.2024.114379. [DOI] [PubMed] [Google Scholar]
- Tan S., Niu Y., Liu L., Su A., Hu C., Meng Y. Development of a GC–MS/SIM method for the determination of phytosteryl esters. Food Chemistry. 2019;281:236–241. doi: 10.1016/j.foodchem.2018.12.092. [DOI] [PubMed] [Google Scholar]
- Tsugawa H., Cajka T., Kind T., Ma Y., Higgins B., Ikeda K.…Arita M. MS-DIAL: Data-independent MS/MS deconvolution for comprehensive metabolome analysis. Nature Methods. 2015;12(6):523–526. doi: 10.1038/nmeth.3393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang T., Shao K., Chu Q., Ren Y., Mu Y., Qu L.…Xia B. Automics: An integrated platform for NMR-based metabonomics spectral processing and data analysis. BMC Bioinformatics. 2009;10(1):83. doi: 10.1186/1471-2105-10-83. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang X., Yan S., Zhao W., Wu L., Tian W., Xue X. Comprehensive study of volatile compounds of rare Leucosceptrum canum Smith honey: Aroma profiling and characteristic compound screening via GC–MS and GC–MS/MS. Food Research International. 2023;169 doi: 10.1016/j.foodres.2023.112799. [DOI] [PubMed] [Google Scholar]
- Wu L., Lou R., Zhang Q., Li K., Hou T. Identification and quality evaluation of Fuyun 6 and Zhongcha 108 tea fresh leaves at different altitudes using non-targeted metabolomics combined with machine learning. Food Chemistry: X. 2025;32 doi: 10.1016/j.fochx.2025.103271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xie Z., Zeng D., Wang J., Zhao M., Feng Y. Dispersive liquid-liquid microextraction coupled with gas chromatography-mass spectrometry (GC-MS) for the determination of soy sauce aroma compounds. Food Control. 2023;152 doi: 10.1016/j.foodcont.2023.109838. [DOI] [Google Scholar]
- Zhang H., Zhou F., Li F., Zhao C., Wang H., Yu H.…Wang C. Quality differentiation method of similar phytomedicines with high sugar content based on the sugar-marker: Taking Schisandrae Chinensis Fructus and Schisandrae Sphenantherae Fructus as an example. Arabian Journal of Chemistry. 2022;15(4) doi: 10.1016/j.arabjc.2022.103727. [DOI] [Google Scholar]
- Zhang J.-N., Ma M.-H., Ma X.-L., Ma F.-L., Du Q.-Y., Liu J.-N., Wang X.-C., Zhao Q.-P., Yu Y.-J., She Y. A comprehensive study of the effect of drying methods on compounds in Elaeagnus angustifolia L. flower by GC-MS and UHPLC-HRMS based untargeted metabolomics combined with chemometrics. Industrial Crops and Products. 2023;195 doi: 10.1016/j.indcrop.2023.116452. [DOI] [Google Scholar]
- Zhang Y.-Y., Zhang Q., Zhang Y.-M., Wang W.-W., Zhang L., Yu Y.-J., Bai C.-C., Guo J.-Z., Fu H.-Y., She Y. A comprehensive automatic data analysis strategy for gas chromatography-mass spectrometry based untargeted metabolomics. Journal of Chromatography A. 2020;1616 doi: 10.1016/j.chroma.2019.460787. [DOI] [PubMed] [Google Scholar]
- Zheng Y., Pang X., Zhu X., Meng Z., Chen X., Zhang J., Ding Q., Li Q., Dou G., Ma B. Lycium barbarum mitigates radiation injury via regulation of the immune function, gut microbiota, and related metabolites. Biomedicine & Pharmacotherapy. 2021;139 doi: 10.1016/j.biopha.2021.111654. [DOI] [PubMed] [Google Scholar]
- Zhou Y., Wang D., Duan H., Zhou S., Guo J., Yan W. Detection and analysis of volatile flavor compounds in different varieties and origins of goji berries using HS-GC-IMS. LWT. 2023;187 doi: 10.1016/j.lwt.2023.115322. [DOI] [Google Scholar]
- Zhu M., Sun J., Zhao H., Wu F., Xue X., Wu L., Cao W. Volatile compounds of five types of unifloral honey in Northwest China: Correlation with aroma and floral origin based on HS-SPME/GC–MS combined with chemometrics. Food Chemistry. 2022;384 doi: 10.1016/j.foodchem.2022.132461. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
No data was used for the research described in the article.






