Skip to main content
The Plant Genome logoLink to The Plant Genome
. 2026 Jun 22;19(2):e70266. doi: 10.1002/tpg2.70266

Haplotype‑resolved comparison of transcription factor superfamilies between wild and cultivated autotetraploid green jujube and prioritization of candidate transcription factors via machine learning

Xudong Zhu 1, Pengyan Chang 1, Yanzhen Liao 2, Huini Wu 1, Fan Jiang 1,✉
PMCID: PMC13287549  PMID: 42333035

Abstract

A framework beyond single‐reference genomes is needed to understand transcription factor evolution. This study employed an integrated haplotype‑resolved genomes–transcriptome atlas–machine learning to characterize the transcription factors of autotetraploid green jujube (Ziziphus mauritiana). The first haplotype‑resolved comparison of transcription factor superfamilies from eight haplotype genomes (HapGenome) representing a specific wild and a specific cultivated green jujube accession, encompassing 42 superfamilies and 12,123 gene copies. Evolutionary analyses revealed high structural conservation with minimal copy number variation, gene presence/absence variations, and strong purifying selection (Ka/Ks < 1). Dispersed duplication (47.24%), not whole‐genome duplication (36.10%), was the most frequently observed duplication event in the expansion of transcription factor superfamily. A haplotype‑resolved transcriptome atlas demonstrated that tissue‐specific expression divergence occurred at the superfamily level and between the core/dispensable genes. Integrating transcriptomic and metabolomic data, support vector machine classification with leave‑one‑out cross‑validated distinguished three wild fruits from six cultivated fruits with the accuracy of 89% using orthologous gene groups (OGGs) expression profiles. The eXtreme gradient boosting was employed as an exploratory tool to prioritize OGGs related to metabolite changes. Finally, OGG‐95, a Lesion Simulating Disease Zn finger transcription factor, was screened out, which was significantly upregulated in cultivated fruits, and its expression was significantly correlated with differential accumulation of nucleotides and organic acids that need further functional validation. This integrative study provided novel insights into the genomic architecture and regulatory evolution of transcription factors in a polyploid fruit crop, highlighting the power of multi‐dimensional analyses for gene discovery.

Plain Language Summary

The genomes of wild plants experienced numerous changes when they transited into cultivated crops. We studied how a key group of regulatory genes, called transcription factors, changed between the genomes of wild and cultivated green jujube, a tetraploid fruit crop. By comparing the complete set of these genes between the two specific wild and cultivated accessions, we found they are highly conserved, meaning most genes stayed the same. Surprisingly, small‐scale gene duplications created more new gene variants than the large‐scale whole‐genome duplication event that shaped this species. We also mapped where and when these genes are active in different plant parts, revealing family‐specific patterns. Using machine learning models as an exploratory tool, we screened out one specific gene family that shows promise for future studies on fruit quality. This research shows how combining large‐scale genetic data with advanced computing can reveal the key regulators of desirable crop traits.


Abbreviations

AUC

area under the curve

CNV

copy number variation

CV

coefficient of variation

HapGenome

haplotype genome

LOO‐CV

leave‑one‑out cross‑validation

LSD

Lesion Simulating Disease

OGG

orthologous gene group

PAV

presence/absence variations

ROC

receiver operating characteristic

SVM

support vector machine

TAU

tissue specificity index

WGD

whole‐genome duplication

XGBoost

eXtreme gradient boosting

1. INTRODUCTION

The process from wild plants to cultivated crops led to the emergence of human‐favored agronomic traits in fruit crops, which has been demonstrated to bring in numerous changes in plant genome (Cao et al., 2025; Cheng et al., 2025; Roy et al., 2025). These genomic changes, including markedly shifts in gene content, such as the gain and loss of genes in cultivated crops compared with their wild relatives, have been widely documented in fruit crops (Long et al., 2024; Su et al., 2024; Y. Wang, Xia, et al., 2025; Y. Wang, Zeng, et al., 2025; X. Zhang, Song, et al., 2025). Polyploidy, or whole‐genome duplication (WGD), represented another pivotal evolutionary mechanism in fruit crops, which has been revealed in kiwifruit (X. Li, Kuhl, et al., 2025), Chinese cherry (W. Zhang, Wang, et al., 2025), citrus plants (X. Song, Zhang, et al., 2025), strawberry (H. Han, Salinas, et al., 2025), and Malus plants (W. Li, Chu, et al., 2025). By generating genetic redundancy, WGD fueled trait improvement through gene duplication, loss, and regulatory rewiring, ultimately enhancing phenotypic diversity and adaptability, which were key genetic resources for breeding. Compared to their diploid relatives, polyploids often exhibited increased organism size, enhanced competitive advantages, and improved adaption levels (H. Han, Salinas, et al., 2025; X. Li, Kuhl, et al., 2025; X. Song, Zhang, et al., 2025).

Green jujube (Ziziphus mauritiana), also commonly known as “Indian jujube” or “ber,” is an economically important fruit crop in the Rhamnaceae family (M. Guo et al., 2024). Its cultivation on the Indian subcontinent began around 11,000 years ago, making it one of the world's oldest cultivated fruit crops. Hence, the process from wild to cultivated has resulted in marked phenotypic variations, including enlarged leaves, reduced stipular spines, larger fruit size, and accelerated growth rates compared to its wild progenitor. Notably, green jujube is an autotetraploid, and its genome has undergone substantial restructuring following polyploidization, a process that has driven the expansion, contraction, and neofunctionalization of specific gene families, thereby generating novel genetic variation (X. Zhu et al., 2026). A comparative genomic analysis between wild and cultivated green jujube thus offered a powerful system to dissect the molecular basis of the process from wild to cultivate, particularly changes in gene family evolution that may underlie key phenotypic transitions.

Transcription factors are regulatory factors that modulate gene expression by binding to specific DNA sequences, governing essential biological processes in plants (M. Sun et al., 2025; Tong et al., 2025). In the past two decades, the genome‐wide analysis of transcription factors gene families has become a standard research methodology with the increasing availability of reference genomes. However, as high‐fidelity, haplotype‐resolved genome are continually being released, studies that rely solely on a single reference genome will fail to capture critical structural variations, like gene presence/absence variations (PAVs) and copy number variations (CNVs) that are prevalent across individuals and populations. In contrast, the haplotype‐resolved genome framework, widely adopted in large‐scale genomic studies, comprehensively accounted for these variable genomic contents. Recent haplotype‐resolved genome analyses in plants have demonstrated that PAVs and CNVs are widespread across species (Long et al., 2024; Su et al., 2024; Y. Wang, Xia, et al., 2025; Y. Wang, Zeng, et al., 2025; X. Zhang, Song, et al., 2025). Nevertheless, most haplotype‐resolved genome analyses to date have focused on specific functional gene families, such as resistance genes (Su et al., 2024), UDP‐dependent glycosyltransferases (M. Huang, Zheng, et al., 2025), terpene synthase gene (Lu et al., 2025), and aroma‐related genes (Long et al., 2024). Systematic haplotype‐resolved genome‐wide explorations of transcription factor superfamilies remain infrequent, with only limited reports (Fan et al., 2025; L. Guo et al., 2025; X. Han, Qiu, et al., 2025; Javed et al., 2026; R. Li, Lei, et al., 2025; T. Li, Liao, et al., 2025; Liu et al., 2025; X. Wang, Cheng, et al., 2025; T. Wang, Zheng, et al., 2025; F. Zhao, Li, et al., 2025; W. Zhao, Wang, et al., 2025; Z. Zhu & Stein, 2025), such as the MYB and bZIP family in Cucurbitaceae (M. Sun et al., 2025; W. Zhao, Wang, et al., 2025) and the bHLH family in barley (Tong et al., 2025). A comprehensive analysis encompassing multiple transcription factor superfamilies across a haplotype‐resolved genome is therefore lacking.

Meanwhile, machine learning has emerged as a powerful computational tool for gene discovery, enabling the prioritization of candidate genes from large‐scale biological datasets through the analysis of feature importance scores (Farhan et al., 2025; Zhou et al., 2026). Algorithms such as support vector machine (SVM) and eXtreme gradient boosting (XGBoost) were particularly effective for identifying significant features within high‐dimensional gene expression data (W. Song, Yang, et al., 2025; Z. Wang, Li, et al., 2025). Although extensively applied in biomedicine to identify molecular signatures and biomarkers for diseases (W. Song, Yang, et al., 2025; X. Wang, Cheng, et al., 2025; Xu et al., 2025), the use of these advanced machine learning approaches in plant sciences remains limited. Most studies that employed machine learning in crop field predominantly concentrated on a limited scope, like classifying disease‐resistant genotypes (Ahn et al., 2025; Farhan et al., 2025). As high‐throughput sequencing has become common, the systematic mining integrated with machine learning to screen genes that regulate complex traits is a significant opportunity.

To address these gaps, this study performed haplotype‐resolved genome‐wide identification of nearly all known plant transcription factor superfamilies (42 in total) for autotetraploid green jujube, encompassing eight haplotype‐resolved genomes from one wild and one cultivated green jujube accession. Then the study comprehensively characterized the accession‐specific evolutionary dynamics and tissue‐specific expression profiles of the two accessions, including CNV, PAV, duplication types, and selection pressures, providing high‐resolution insights into genomic differences between the two specific accessions. Finally, the study employed machine learning approaches to prioritize candidate regulatory OGGs through the transcriptomic and metabolomic variation in fruits from three wild and six cultivated green jujube accessions. Specifically, SVM with leave‑one‑out cross‑validated (LOO‐CV) was used to distinguish wild and cultivated genotypes based on fruit transcriptome data, and XGBoost was further applied to exploratory screen out OGG‐95 as high‐priority candidates for future functional studies to reveal its possible relations with the changes in nucleotides and organic acid.

This integrated haplotype‐resolved genome–transcriptome atlas–machine learning framework provided novel understandings into the genomic architecture, evolutionary conservation, and functional roles of transcription factor superfamilies between the two wild and cultivated green jujube accessions. It suggested that the clear advantage of such a multi‐dimensional approach over conventional single‐genome, single‐family analyses for accurately characterizing gene family evolution and prioritizing key regulators, thereby laying a foundation for future functional studies and molecular breeding in green jujube and other polyploid crops.

2. MATERIALS AND METHODS

2.1. Identification of transcription factor superfamilies

First, we respectively retrieved the four haplotype genomes (hereafter referred to as HapGenome) and annotation files (fasta and gff3 format) of one specific wild (XSYS) and one specific cultivated (Misi) green jujube released by M. Guo et al. (2024) at figshare (https://doi.org/10.6084/m9.figshare.23530068), and following this, the protein sequences were extracted using GFFREAD (https://github.com/gpertea/gffread) (Pertea & Pertea, 2020). Then, three tools, iTAK (https://github.com/FeiLab/iTAK) (Zheng et al., 2016), PlantTFDB (https://planttfdb.gao‐lab.org/prediction.php) (Tian et al., 2020), and PlantTFcat (https://www.zhaolab.org/PlantTFcat/) (Dai et al., 2013), were used to search the protein sequences and generated three independent predictive datasets. When one gene was predicted by at least two of three datasets to be one specific transcription factor, this gene will be recognized as the member of a specific transcription factor superfamily.

For each transcription factor gene superfamily, key descriptive statistics were calculated across the eight HapGenomes: the mean, standard deviation (Std), and range. The coefficient of variation (CV), defined as the ratio of the Std to the mean (CV = Std/mean), was employed as the primary metric to quantify relative fluctuation. The R package RIdeogram (https://github.com/TickingClock1992/RIdeogram) was used to visualize the distribution of transcription factor across haplotype chromosomes (Hao et al., 2020). GO (gene ontology) and KEGG (Kyoto Encyclopedia of Genes and Genomes) enrichment assessments of transcription factor genes were performed using eggnog‐mapper (http://eggnog‐mapper.embl.de/) (Cantalapiedra et al., 2021).

2.2. Analysis of orthologous gene groups (OGGs)

OrthoFinder (https://github.com/davidemms/OrthoFinder) (Emms & Kelly, 2019) was used to identify the orthologous gene groups (OGGs) across eight HapGenomes. OrthoFinder was applied with an inflation parameter of 1.5 and E‐value cutoff of 1e‐5 to group protein sequences into orthologous gene groups. Thereafter, OGGs present in all eight HapGenomes were defined as core OGGs, while those that were shared by two to seven HapGenomes were defined as dispensable OGGs. In addition, some transcription factor gene copies could not be assigned to any OGG, that is, they had no orthologs in any other haplotype. These unassigned singletons were defined as private OGGs.

MUSCLE (https://github.com/rcedgar/muscle) (Edgar, 2004) was used to perform sequence alignment, and TrimAL (https://vicfero.github.io/trimal/) (Capella‐Gutiérrez et al., 2009) was used to extract and trim the conserved sequences. FastTree2 (https://github.com/morgannprice/fasttree) (Price et al., 2010) was used to construct phylogenetic evolutionary tree by maximum likelihood method based on highly conserved single‐copy gene, the LG + CAT evolutionary models and SH test were used. The time divergence between different HapGenomes was obtained from TimeTree5 (https://timetree.org/). CAFE5 (https://github.com/hahnlab/CAFE5) (Mendes et al., 2021) was used to estimate the expansion or contraction of OGGs based on a model of copy number evolution across the species phylogeny.

2.3. Analysis of phylogenetic relationships

To determine whether genes are colinear between different HapGenomes, the MCScanX (https://github.com/wyp1125/MCScanX) (Y. Wang et al., 2012) software was used to investigate the synteny, collinearity, and gene duplication patterns of transcription factor genes with default parameters, and further, the regions comprising these colinear genes are referred to as colinear blocks to facilitate understanding of chromosome rearrangement and evolution.

KaKs_Calculator 3.0 (https://github.com/Chenglin20170390/KaKs_Calculator‐3.0) (Z. Zhang, 2022) was used to calculate the synonymous (Ks) and nonsynonymous (Ka) mutation rates of the duplicated gene pairs, which measure the selection pressure during evolution.

2.4. Analysis of expression profiles through transcriptome datasets

The raw RNA‐seq data of different tissues from one wild (XSYS) and one cultivated (Misi) green jujube accession were downloaded from the Sequence Read Archive (SRA) database of the National Center for Biotechnology Information (NCBI) under accession numbers listed in Table S1 using the SRA toolkit.

The FastQC tool (http://www.bioinformatics.babraham.ac.uk/projects/fastqc/) was used to check the quality of the raw sequencing data, and fastP (https://github.com/OpenGene/fastp, Chen et al., 2018) was used to trim the adapter and low‐quality base. To enable direct comparison of expression levels between the one wild and one cultivated green jujube accession, the study built a combined Kallisto (https://github.com/pachterlab/kallisto, Bray et al., 2016) index containing all coding sequences from the eight HapGenomes (cI–cIV and wI–wIV). Reads from each sample were quantified against this combined index, and transcript per million (TPM) values were obtained. The TPM value was processed with log2 (TPM + 1) to obtain the expression heat map using TBtools‐II. The log2 (TPM + 1) value < 1 was determined as the gene was not expressed.

Finally, genes were assigned to their respective OGGs based on OrthoFinder analysis. To obtain a single representative expression value for each OGG across the HapGenomes, the median TPM of all its member genes is the expression value of individual OGG, and therefore copy/allelic‐specific variation would be ignored. The R package ClusterGVis (https://github.com/junjunlab/ClusterGVis) was used to conduct the K‐means cluster analysis based on the expression profiles of OGGs.

Additionally, tissue specificity index (TAU) of genes was calculated using the online tool tspex (https://tspex.lge.ibi.unicamp.br/). The TAU indicates how specific or broadly expressed a gene or transcript is within studied tissues. Genes with a TAU close to 1 are more specifically expressed in one tissue, while genes with a TAU closer to 0 are equally expressed across all tissues studied (Mantica et al., 2024).

2.5. Selection of candidate OGGs discriminating between wild and cultivated fruits

The raw RNA‐seq data of nine fruit samples from three wild (M58, M60, and M74) and six cultivated (M18, M25, M28, M31, M46, and M47) green jujube accessions were downloaded from the NCBI SRA database (Table S1).

This study performed SVM with LOO‐CV to discriminate wild from cultivated fruits based on the expression profiles of OGGs. First, expression data were aggregated to the accession level by taking the median expression of each OGG across biological replicates for each of the nine accessions (three wild and six cultivated). This yielded a dataset of nine samples, each representing an independent genotype. In each fold of the LOO‐CV, one accession was held out as the test set, and the remaining eight accessions were used to train an SVM with default cost parameter C = 1 and feature scaling. No pre‐selection of features was applied and all OGGs were used as predictors. The trained model was then used to predict the class (wild or cultivated) of the held‐out accession. This procedure was repeated until every accession had served as the test set once.

This study performed a permutation test with 1000 iterations to assess the classification accuracy. In each iteration, the class labels of the nine accessions were randomly shuffled, and the entire LOO‐CV procedure was repeated. The p‐value was calculated as the proportion of permutations where the accuracy was equal to or greater than the observed accuracy.

All analyses were implemented in the R package “e1071” (version 1.7‐17) (https://github.com/cran/e1071) for SVM and base functions for cross‐validation and permutation.

2.6. Selection of candidate OGGs related to altered metabolites between wild and cultivated fruit

The metabolome data of nine fruit samples from six cultivated green jujube accessions (M18, M25, M28, M31, M46, and M47) and three wild green jujube accessions (M58, M60, and M74) were retrieved from the MetaboLights database with the identifier MTBLS10827 (https://www.ebi.ac.uk/metabolights/editor/study/MTBLS10827) released by M. Guo et al. (2024). First, the significance of total abundance of alkaloids, amino acids and derivatives, flavonoids, lipids, nucleotides and derivatives, organic acids, saccharides and alcohols, vitamin, phenolic acids, and terpenoids between wild and cultivated green jujube was calculated by the Wilcoxon rank sum test (p < 0.05). The result was drawn by box plots. Finally, the metabolites with significant differences were further used for the selection of key OGGs using the R package “xgboost” (version 3.1.3.1) (https://cran.r‐project.org/web/packages/xgboost/index.html).

Besides, the receiver operating characteristic (ROC) curves were constructed using the R package pROC (version 1.19.0.1) (https://cran.r‐project.org/web/packages/pROC/index.html), and the area under the curve (AUC) was calculated. OGGs with an AUC > 0.5 were considered to have acceptable diagnostic performance, where a value closer to 1 indicates a stronger ability to discriminate.

Due to the small sample size (nine accessions), the XGBoost results were intended for feature prioritization rather than definitive inference. All identified associations were further examined by Pearson correlation analysis to filter robust relationships.

2.7. Statistical analysis

Statistical analyses were performed using the SPSS software (versi on 21.0). Statistical assumptions of normality and homogeneity of variance were evaluated using the Shapiro–Wilk and Levene's tests, respectively. Based on these results, parametric or nonparametric tests were applied. Parametric data were analyzed with a one‐way analysis of variance and Tukey's honestly significant difference post hoc test. Nonparametric data were analyzed with the Wilcoxon rank sum test (two groups) and Kruskal–Wallis H test (three or more groups), and significant results were adjusted by Bonferroni for multiple comparisons.

3. RESULTS

3.1. Identified transcription factor superfamilies across eight HapGenomes from the two wild and cultivated green jujube accessions

To accurately and comprehensively identify all transcription factor superfamilies across eight haplotype green jujube genome, this study performed a combined criterion to select the transcription factor genes based on three unique tools, iTAK, PlantTFDB, and PlantTFcat. As shown in Figure S1, the number of transcription factor genes identified was completely different from these unique tools. Generally, PlantTFcat was the highest, followed by iTAK, and PlantTFDB is the lowest. The number of shared transcription factor genes using three tools was close among eight HapGenomes (ranged from 1247 to 1320). When one gene was confirmed by two or three tools to be a transcription factor, this gene could be recognized as a specific transcription factor (detailed prediction results can be found in Table S2). After selection on the basis of this criterion, the total number of identified transcription factor genes was 1486 in cI HapGenome, 1532 in cII HapGenome, 1522 in cIII, 1509 in cIV, 1541 in wI, 1508 in wII, 1518 in wIII, and 1507 in wIV (Figure S1).

Finally, 42 transcription factor gene superfamilies, that is, Alfin‐like, AP2, AUX/IAA, B3, BBR‐BPC, BES, bHLH, bZIP, CAMTA, CPP, E2F/DP, EIL, FAR, GeBP, GRAS, GRF, HMG‐box, Homeo, Homobox, HRT, HSF, Jumonji, LFY, LOB, MADS, MYB, NAC, NF‐X1, NF‐Y, Nin‐like, NZZ/SPL, S1Fa‐like, SBP, SET, SNF2, SRS, STAT, Trihelix, TUBBY, Whirly, WRKY, and Zn finger, were identified across the eight HapGenomes (Figure S2). And the subclass compositions of these transcription factor superfamilies are shown in Table S3. The copy number of each transcription factor gene superfamily in each HapGenome was highly divergent across the eight HapGenomes (as shown in Figure S2) and thus divided into four subgroups. For example, the copy number of Zn finger, MYB, bHLH, and AP2 superfamily was about 100 in most of HapGenomes; the copy number of Homobox, NAC, B3, WRKY, MADS, and FAR ranged from 46 to 85 in all HapGenomes; and the copy number of bZIP, GRAS, LOB, Homeo, NF‐Y, SET, Trihelix, HMG‐box, HSF, SBP, AUX/IAA, SNF2, and Jumonji ranged from 6 to 53, and the remaining 19 transcription factor gene superfamilies were lower than 10.

3.2. CNV of transcription factor superfamilies between the two wild and cultivated accessions

CNV refers to a class of structural variation involving an abnormal number of copies of a specific genomic region, resulting from gain or loss of gene copies. The CNV of transcription factor superfamilies between the two specific wild and cultivated green jujube accessions was evaluated at multiple levels.

3.2.1. Comparison of copy number in different superfamilies between the two wild and cultivated accessions

The total copy number of 42 transcription factor gene superfamilies was obviously differential. The highest number of transcription factor gene superfamilies was Zn finger, MYB, bHLH, and AP2, with a number of 1835, 1579, 1114, and 948, respectively. And the lowest number of transcription factor gene superfamilies was LFY, NZZ/SPL, S1Fa‐like, and STAT, with a number of 8, 10, 8, and 8, respectively (Figure 1a).

FIGURE 1.

FIGURE 1

Copy number variation of identified transcription factor superfamilies. (a) Total number of all transcription factor superfamilies in all HapGenomes. (b) Coefficient of variation in copy number of all transcription factor superfamilies. Based on the mean copy number, transcription factor superfamilies were classed into four groups: >100, 100–50, 50–10, and <10. Red arrow indicated the families with CV value were 0. The cI, cII, cIII, and cIV were the four HapGenomes of the cultivated green jujube accession. The wI, wII, wIII, and wIV were the four HapGenomes of the wild green jujube accession.

To quantitatively measure the variation of copy number of each transcription factor superfamily among eight HapGenomes, the CV was calculated, which proved to be a robust metric for such comparisons, effectively normalizing for differences in average family size to reveal true relative fluctuations (Figure 1b). Several superfamilies, such as Jumonji (CV was 0.27, ranged from 7 to 14 copies), NAC (CV was 0.233, ranged from 48 to 82 copies), and AUX/IAA (CV was 0.20, ranged from 11 to 21 copies), displayed considerable variation in copy number. In contrast, a distinct set of superfamilies, like Zn finger (CV was 0.02, ranged from 221 to 233 copies), bHLH (CV was 0.01, ranged from 137 to 143 copies), and WRKY (CV was 0.01, ranged from 60 to 62 copies), implied exceptional conservation in number of gene copies across all HapGenomes of these two specific accessions. In addition, E2F/DP (six copies each), EIL (five copies each), SRS (five copies each), STAT, LFY, S1Fa‐like (one copy each), HRT, NF‐X1, and Whirly (two copies each) families maintained an identical copy number in all eight HapGenomes, highlighting their fundamental and nonredundant biological functions.

3.2.2. Comparison of copy number between the two wild and cultivated accessions

To explore differences in copy numbers of transcription factor superfamily between the two wild and cultivated accessions, the study compared the median copy number across the four cultivated HapGenomes with the median across the four wild HapGenomes for each superfamily (Figure 2 and Figure S3). Because each accession was represented by a single biological sample, no statistical tests were performed, and only descriptive differences were reported. The total copy number of transcription factors across all 42 superfamilies was similar between the specific cultivated (median across four HapGenomes = 1516) and the specific wild (median = 1513) accession (Figure 2). For most superfamilies, the median copy numbers differed by <10% between the two accessions (Figure 2 and Figure S3). Notable exceptions were AP2, NAC, and B3, which showed considerable differences between the two wild and cultivated accessions. Specifically, the median copy number of AP2 superfamily was lower in the cultivated accession (median = 114) than in the wild accession (median = 124). A similarly lower copy number from the B3 superfamily in the cultivated accession was found. In contrast, a higher copy number of the NAC superfamily was reported in the cultivated accession (median = 79) than in the wild (median = 66).

FIGURE 2.

FIGURE 2

Comparison of the copy numbers of 23 transcription factor superfamilies between the two cultivated and wild green jujube accessions. Each dot represented one HapGenome, and bars indicated medians.

3.2.3. CNV among different chromosomes between the two wild and cultivated accessions

The distribution of transcription factors in 12 chromosomes harbored a highly similar trend across the two wild and cultivated green jujube accessions (Figure S4). The median copy number of transcription factors in Chr01, Chr06, and Chr08 was relatively high, while the median copy number of transcription factors in Chr05 and Chr07 was relatively low (Figure S4a). Further analysis indicated that only in Chr01, the variation of median copy number of transcription factor was observed between the two wild and cultivated green jujube accessions (cultivated = 172 and wild = 179), and there was no observed variation in other 11 chromosomes (Figure S4b). The distribution map of transcription factors in Chr01 illustrated the observed difference (Figure S4c).

3.3. PAVs of OGGs between the two wild and cultivated accessions

To investigate the conservation of transcription factors in these two specific accessions, the transcription factors identified in the eight HapGenomes from the two wild and cultivated accessions were assigned to different OGGs. In total, of the 12,123 transcription factor gene copies, 12,022 could be assigned to 1652 OGGs, the remaining 101 gene copies lacked orthologs in any other HapGenome and were therefore classified as private OGGs (unassigned singletons) (Table S4). The number of OGGs was higher than the average number of transcription factors per HapGenome (1515), indicating the presence of PAV for transcription factors in the two specific wild and cultivated accessions. Based on the PAV of each OGG in the eight HapGenomes, the 1652 OGGs were classified into 1167 “core” OGGs (present in all eight HapGenomes) and 485 “dispensable” OGGs (absent in at least one HapGenome), including 137 OGGs (conserved in seven HapGenomes), 119 OGGs (conserved in six HapGenomes), 73 OGGs (conserved in five HapGenomes), 48 OGGs (conserved in four HapGenomes), 31 OGGs (conserved in three HapGenomes), and 77 OGGs (conserved in two HapGenomes) (Figure 3a). The number of OGGs identified in each HapGenome varied dramatically from 1448 (cI) to 1508 (wI) (Figure 3b). The number of genes included in OGGs of a single HapGenome also varied from 0 to 4, and most were 1 (Figure 3c). The remaining 101 genes were defined as “private” OGGs (singletons without orthologs), which varied from 6 to 32 across eight HapGenomes (Figure 3d). Additionally, the number of common OGGs in the wild accession (wI, wII, wIII, and wIV) and the cultivated accession (cI, cII, cIII, and cIV) green jujube genome was 41 and 32, respectively (Figure 3d).

FIGURE 3.

FIGURE 3

Identification of OGGs across eight HapGenomes of the two wild and cultivated green jujube accessions. (a) Number of core and dispensable OGGs. (b) Number of OGGs in eight HapGenomes. (c) More/less information of transcription factor genes included in OGGs in each HapGenome. The color represented the number of transcription factor genes that a specific OGG contained in each HapGenome. (d) Number of OGGs shared in wild and cultivated green jujube genomes, respectively, and private OGGs, which are unassigned gene copies that do not belong to any orthogroup.

In detail, core OGGs occupied 9503 (78.39%) distinct transcription factor gene copies, dispensable OGGs were composed with 2519 (20.78%) transcription factor gene copies, and private OGGs were 101 (0.83%) (Figure 4a). Regarding the number of gene copies in OGGs, the core OGGs of the cultivated green jujube accession were 4753, dispensable OGGs were 1258, and private OGGs were 38. Similarly, the core OGGs in the wild green jujube accession were 4750, dispensable OGGs were 1261, and private OGGs were 63 (Figure 4a). Distribution of core, dispensable, and private OGGs across eight HapGenomes was also evaluated (as shown in Figure 4b), the core OGGs’ number was nearly identical (from 1186 to 1189), and the dispensable OGGs’ number varied from 284 to 347 in all HapGenomes. The count of private OGGs in eight HapGenomes was also exhibited in Figure 3d. This study further illustrated the PAV of OGGs across the eight HapGenomes from the two specific wild and cultivated accessions (Figure 4c). As expected, most of OGGs exited in all eight HapGenomes, and a proportion of OGGs was only occurred in limited number of HapGenomes (Figure 4c).

FIGURE 4.

FIGURE 4

Distribution of core, dispensable, and private OGGs in the two wild and cultivated green jujube accessions. (a) Number of core, dispensable, and private OGGs (singletons without orthologs) in wild and cultivated green jujube. (b) Distribution of core, dispensable, and private OGGs (singletons without orthologs) in eight HapGenomes. (c) Presence/absence information of OGGs across eight HapGenomes.

However, all the above observations were based on a single wild and a single cultivated accession and might reflect accession‐specific rather than general patterns; therefore, they need to be further investigated in a larger panel including multiple wild and cultivated accessions.

3.4. Pathway enrichment analysis of OGGs between the two wild and cultivated accessions

The further pathway enrichment analyses of OGGs revealed that core OGGs are involved in broader pathways than dispensable and private OGGs (Figure S5a). However, the pathways involved in core OGGs between the two wild and cultivated green jujube accessions nearly are the same except for the difference in the number of OGGs (Figure S5b). Similarly, the pathways associated with dispensable OGGs were nearly the same in the two wild and cultivated green jujube accessions, differing primarily in the number of OGGs involved (Figure S5c). In contrast, the pathways of private OGGs of the wild accession are highly enriched in genetic information processing, chromosome‐associated proteins, and transcription machinery than in the cultivated accession, which only possesses pathways related to transcription factors and circadian rhythm‐plant (Figure S5d).

The number of core, dispensable, and private genes in distinct transcription factor superfamilies was highly separated (Figure S6). The number of core genes in most of transcription factor superfamily was undoubtedly higher than the dispensable and private genes, whereas NAC, AUX‐IAA, and Jumonji superfamilies displayed a completely opposite pattern. In detail, the core copies of NAC numbered 174, while the dispensable copies numbered 359, which is approximately twice the core copies. More obviously, nearly all the AUX‐IAA genes were categorized as dispensable, while only eight core and two private copies, respectively. Interestingly, Jumonji superfamily had no core copies at all, and almost all were dispensable copies, with only 14 copies classified as private.

3.5. Evolution of OGGs across eight HapGenomes between the two wild and cultivated accessions

To assess whether the conservation level of transcription factors in the green jujube genomes might be related to gene duplication class, the duplication types (WGD/segmental, dispersed, tandem, proximal, dispersed, and singleton) for each conservation category (core, dispensable, and private) were analyzed across eight HapGenomes of the two wild and cultivated accessions (Figure 5). In general, dispersed and WGD/segmental duplication were the dominant gene‐duplication types (occupying 47.24% and 36.10%, respectively). Tandem, singleton, and proximal duplicates accounted for 9.14%, 4.30%, and 3.22% of the genes, respectively (Figure 5a). Distribution of gene duplication across the eight HapGenomes was unquestionably similar to the overall distribution pattern (Figure 5b).

FIGURE 5.

FIGURE 5

Collinearity analysis of OGGs in the two wild and cultivated green jujube accessions. (a) General distribution of different gene‐duplication classes. (b) Chromosome distribution of different gene‐duplication classes. (c) Distribution of different gene‐duplication classes in core, dispensable, and private OGGs. (d) Distribution of different gene‐duplication classes in 42 different transcription factor families.

In terms of gene duplication type, core and dispensable had a similar composition of duplication, with dispersed and WGD/segmental occurring predominantly (Figure 5c). By contrast, among private genes, proximal duplications were more dominant than WGD/segmental duplications, even though dispersed duplications were frequently observed (Figure 5c). Additionally, the distribution of gene duplication events was completely unequaled among identified 42 transcription factor superfamilies (Figure 5d). In detail, only 12 superfamilies, including Zn finger, MYB, bHLH, and AP2 covered all five gene duplication classes, NZZ/SPL, S1Fa‐like, and STAT superfamilies were both singleton duplicates, and NF‐X1 and Whirly superfamilies were solely dispersed duplicates (Figure 5d). Transcription factor superfamily for Homobox had no proximal duplicates, FAR had no singleton duplicates, GRAS had no proximal duplicates, Homeo had no tandem duplicates, and SNF2 had no WGD/segmental duplicates. A visualization of syntenic block comprised by the transcription factor genes across eight HapGenomes evidently showed their syntenic relationship (Figure S7), consistent with the haplotype‐resolved genome analyses.

The ratio of non‐synonymous (Ka) to synonymous (Ks) substitutions (Ka/Ks) indicated selection pressure, where Ka/Ks < 1, = 1, and >1 means purifying, neutral, and positive selection, respectively. To evaluate the selection pressures on the transcription factors of the two wild and cultivated accessions, this study calculated Ka/Ks ratios for each gene (Figure 6; Table S5). The general comparison of all Ka/Ks value revealed no significant difference between the two wild and cultivated green jujube accessions. Even though 499 gene pairs between the two wild and cultivated accessions had the Ka/Ks value of >1, only 182 and 137 gene pairs were >1 inside the two cultivated and wild accessions, respectively (Figure 6a). Specifically, when selection pressures on core, dispensable, and private genes were compared, dispensable genes possessed significantly (p < 2.22e‐16) higher medium Ka/Ks values than core genes, and the Ka/Ks values of dispensable genes were also more divergent than those of core genes (Figure 6b). In contrast, the Ka/Ks ratios of core and dispensable genes did not differ significantly between the specific two cultivated and wild green jujube accessions (Figure 6c,d).

FIGURE 6.

FIGURE 6

Comparison of Ka/Ks values in the transcription factors between the two wild and cultivated green jujube accessions. (a) Overall comparison of Ka/Ks values between the two wild and cultivated green jujube accessions. From top to bottom, half violin plot indicates the data density, box plot indicates the data centration and dispersion (including minimum, lower quartile, median, upper quartile, and maximum value), scatter plot indicates the data distribution, red line indicates mean values, blue line indicates the Ka/Ks = 1, and blue number indicates the number of gene pairs with Ka/Ks > 1. Black number indicates the total number of gene pairs. (b) Comparison of Ka/Ks values among core, dispensable, and private genes. (c) Comparison of Ka/Ks values of core genes in the two wild and cultivated green jujube accessions. (d) Comparison of Ka/Ks values of dispensable genes in the two wild and cultivated green jujube accessions. Kruskal–Wallis H test was conducted for significant analysis at p < 0.05, and ns: no significant.

Based on the copy number of each superfamily, the comparison of Ka/Ks values from different transcription factor superfamilies classified into four groups was also estimated (Figure S8). The Ka/Ks values for the Zn finger, MYB, bHLH, and AP2 superfamilies differed significantly, with average value of 0.39, 0.42, 0.36, and 0.32, respectively (Figure S8a). Likely, the Ka/Ks values between most transcription factor superfamilies from second group showed significant differences, except for the comparison between B3 and FAR, NAC and FAR, B3 and MADS, and MADS and FAR (Figure S8b). The other two groups also discovered significantly differential Ka/Ks values between some superfamilies, accompanied by a proportion of superfamilies with no difference (Figure S8c,d).

3.6. Expansion and contraction of OGGs between the two wild and cultivated accessions

To analyze the contraction and expansion of OGGs in an evolutionary context, this study constructed a phylogenetic tree based on the eight HapGenomes of the two wild and cultivated green jujube accessions (Figure S9). It is crucial to note that the divergence time of approximately 4 Mya between the two wild and cultivated accessions, as estimated from the TimeTree database, represents the lineage divergence of their respective wild ancestors. This event occurred long before the onset of the human‐directed process from wild to cultivated, which is estimated to have begun around 11,000 years ago. As expected, the HapGenomes for cI and cII, cII and cIV, wI and wII, and wIII and wIV were clustered into one clade in terms of green jujube as autotetraploid plant (Figure S9). The OGGs’ contraction and expansion analysis presented that the divergence of wild and cultivated green jujube (approximately 4 Mya) had a minor role in the contraction and expansion of OGGs, that is, only one OGG gained in the cultivated accession and one OGG lost in the wild accession (Figure S9). By contrast, the more recent WGD event (estimated at <1 Mya) was followed by a net loss of OGGs, which may reflect an obvious reduction in OGG number. In detail, the two OGGs gained and 30 OGGs lost during the duplication of wIII–wIV HapGenome, one OGG gained and 21 OGGs lost for wII–wI, one OGG gained and 30 lost for cII–cI, two OGGs gained and 25 OGGs lost for cIV–cIII (Figure S9).

Furthermore, after the chromosome duplications, the long evolution of eight HapGenomes also influenced substantially the OGG number, with an obvious reduction in OGG number observed (Figure S9). For example, five OGGs gained and 64 OGGs lost in the wIV HapGenome when it evolved, while two OGGs gained and 59 OGGs lost in the cI and cIV HapGenomes.

3.7. Haplotype‐resolved transcriptome atlas of transcription factors across multiple tissues of the two wild and cultivated accessions

The expression‐pattern analysis by transcription factor superfamilies unveiled clear family‐specific and tissue‐specific preferences (Figure 7). Although the majority members from most of superfamilies were expressed in all tissues, a considerable number of members from some superfamilies were not expressed in some tissues. For instance, only 60%–80% members of AP2 and B3 superfamilies were expressed in the four tissues of the cultivated green jujube accession, and nearly 50% members of LOB superfamily were expressed (Figure 7a). Additionally, the expression levels of top four transcription factor superfamilies, including Zn finger, AP2, MYB, and bHLH, showed imbalance between these four tissues of the cultivated accession (Figure 7b). Moreover, only a proportion of AP2 and B3 superfamily members was expressed in root, leaf, branch, and enlarged fruit of the wild green jujube accession, while all were expressed in mature fruit. Interestingly, no members of LFY family expressed in branch and enlargement fruit (Figure 7c). Likely, the top four transcription factor superfamilies (Zn finger, AP2, MYB, and bHLH) presented unequal expression between these five tissues of the wild accession (Figure 7d).

FIGURE 7.

FIGURE 7

Expression profiles of transcription factor superfamilies in different tissues of the two wild and cultivated green jujube accessions. (a) Proportions of 42 transcription factor superfamilies expressed in different tissues of the cultivated accession. (b) Expression patterns of Zn finger, MYB, bHLH, and AP2 superfamilies expressed in different tissues of the cultivated accession. (c) Proportions of 42 transcription factor superfamilies expressed in different tissues of the wild accession. (d) Expression patterns of Zn finger, MYB, bHLH, and AP2 superfamilies expressed in different tissues of the wild accession.

3.8. Transcriptome atlas of OGGs across three shared tissues of the two wild and cultivated accessions

Furthermore, the expression profiles of OGGs were examined in three tissues (leaf, branch, and mature fruit) shared by the two wild and cultivated green jujube accessions (Figure S10; Tables S6 and S7). Obviously, the core, dispensable, and private OGGs are expressed in the three tissues of the cultivated (Figure S10a) and wild (Figure S10b) green jujube accession. However, the proportion of expressed core OGGs was the highest, followed by the dispensable, and the private OGGs were the lowest in the cultivated green jujube accession (Figure S10a). Likely, the more core OGGs were expressed in the three tissues of the wild accession than the dispensable and private OGGs, which both were nearly the same (Figure S10b).

The TAU indicated the tissue‐specific expression of one gene. In this study, the TAU of all the OGGs in three tissues revealed the slight shift in tissue‐specific expression of OGGs between the two wild and cultivated accessions (Figure S10c). Compared to the wild green jujube accession, the OGGs of the cultivated accession expressed broadly in branch, and narrowly in leaf and mature fruit (Figure S10c).

Furthermore, the K‐means clustering of 1167 core OGGs was conducted to uncover their expression patterns (Figure S10d; Tables S8 and S9). Altogether, 1091 core OGGs were grouped into eight clusters (C1–C8) across the leaf, branch, and mature fruit of the two wild and cultivated green jujube accessions (Figure S10d). Notably, these results further confirmed the tissue‐specific expression patterns of core OGGs (Figure S10d). For instance, the core OGGs in C3 subclade were highly expressed in mature fruits of the cultivated accession, whereas showed low expression in mature fruits of the wild accession (Figure S10d). In contrast, C5 subclade expressed highly in both branches of the two cultivated and wild accessions (Figure S10d).

3.9. Transcriptome atlas of OGGs across nine fruits from three wild and six cultivated accessions

Based on the transcriptome data of nine different mature fruit samples from three wild green jujube accessions and six cultivated accessions, 1078 core OGGs were divided into eight clusters (C1–C8) (Figure 8a). Lineage‐specific expression patterns were then observed among the core OGGs (Figure 8a). For instance, C6 subclade was expressed highly in six cultivated fruits, whereas C5 subclade was expressed highly in three wild fruits.

FIGURE 8.

FIGURE 8

Expression profiles of core OGGs in nine fruits comprised of three wild and six cultivated green jujube accessions and their support vector machine (SVM) with leave‑one‑out cross‑validation (LOO‐CV) analysis. (a) K‐means cluster analysis of core OGGs in nine fruits from three wild and six cultivated green jujube accessions. (b) Confusion matrix of LOO‐CV predictions using SVM. Rows represented predicted classes, and columns represented true classes. Numbers indicated the count of accessions. (c) Distribution of accuracy from 1000 permutation tests with randomly shuffled class labels. The red dashed line marked the observed accuracy. The annotated p‑value indicated the proportion of permutations achieving accuracy ≥ the observed value. (d) Receiver operating characteristic (ROC) curve derived from the predicted probabilities of the SVM with LOO‐CV. (e) Top 20 OGGs ranked by mean absolute SVM weight across the nine LOO‐CV folds. Higher weight indicated stronger contribution to the classification model. OGG‑95 (highlighted in red box) showed the high importance. (f) Boxplot of OGG‑95 expression (log2(TPM + 1)) between three wild and six cultivated fruits. The thick horizontal line indicated the median, boxes represented the interquartile range, and points showed individual accession values. The Wilcoxon rank‑sum test p‑value was shown.

In summary, the haplotype‐resolved transcriptome atlas revealed that tissue‐ and lineage‐specific expression divergence was a prominent feature of the green jujube transcription factor superfamilies, even under the background of high genomic structural conservation between the two specific accessions. These expression differences between these wild and cultivated accessions, particularly in fruits, provided a functional basis for discriminating the two germplasm types and for screening out transcription factors related to the expression changes between wild and cultivated accessions. Therefore, the observed expression divergence served as a direct bridge from evolutionary analysis to the machine learning‐based prioritization of candidate regulatory gene.

3.10. Classification of wild and cultivated fruits using OGG expression profiles

To determine whether the overall expression of transcription factor OGGs can discriminate between nine green jujube accessions comprised of the fruits of three wild and six cultivated fruit samples, this study performed an SVM with LOO‐CV on the nine accession‐level samples (three wild and six cultivated). The model achieved an overall accuracy of 8/9 (88.9%, 95% confidence interval: 51.8%–99.7%), with a sensitivity of 66.7% (two out of three wild accessions correctly classified) and a specificity of 100% (all six cultivated accessions correctly classified) (Figure 8b; Table S10). The single misclassified wild accession was consistently predicted as cultivated.

A permutation test with 1000 random label shuffles yielded a p‐value of 0.04 (Figure 8c), indicating that the observed classification performance was significantly better than chance. The area under the curve (AUC) derived from the leave‐one‐out predicted probabilities was 0.604 (Figure 8d), further indicating the modest discriminative power of the OGG expression profiles, consistent with the exploratory nature of this analysis. However, given the limited sample size, the permutation p‐value should be interpreted cautiously.

Among individual OGGs, OGG‐95 showed the highest mean absolute SVM weight across the nine cross‐validation folds (Figure 8e; Table S11), suggesting it was the feature most heavily weighted in the classification model. Consistently, OGG‐95 expression was significantly higher in cultivated fruits than in wild fruits with a p value of 0.024 (Figure 8f; Table S11). These results implied OGG‐95 as a top candidate associated with the transcriptional shift in green jujube fruit.

3.11. Prioritization of OGGs associated with altered metabolites between wild and cultivated fruits using XGBoost and correlation validation

Fruit metabolites between three wild and six cultivated green jujube accessions had been shown to be significantly differential (Jiang et al., 2025). Herein, the study performed a XGBoost machine learning to screen the potential OGGs, which might be associated with the discrepancy in metabolites.

First, the significance of the difference in metabolite profiles from three wild and six cultivated green jujube accessions was evaluated using the Wilcoxon rank sum test (Figure S11; Table S12). In total, among 10 metabolite classes, including alkaloids, amino acids and derivatives, flavonoids, lipids, nucleotides and derivatives, organic acids, phenolic acids, saccharides and alcohols, terpenoids, and vitamin, only five metabolite classes (alkaloids, flavonoids, nucleotides and derivatives, organic acids, and phenolic acids) possessed the significant difference (p < 0.01) between these wild and cultivated green jujube accessions (Figure S11). Furthermore, XGBoost was applied to screen the potential candidate OGGs related to these differential accumulated metabolites. The XGBoost analysis screened out OGG‐4, OGG‐178, OGG‐713, OGG‐95, OGG‐110, OGG‐24, and OGG‐3 as candidate features potentially associated with differential accumulation of alkaloids, flavonoids, nucleotides, organic acids, and phenolic acids (Figure 9). Specifically, OGG‐4 was found to be related to the alkaloids differential with a gain value of 0.53 and AUC value of 0.604 (Figure 9a–c). OGG‐178 and OGG‐713 have a significant association with flavonoids, evidenced by gain values of 0.23 and 0.21 and AUC values of 0.625 and 0.965, respectively (Figure 9d–g). OGG‐95 was associated with the differential accumulation of nucleotides, achieving a gain value of 0.75 (Figure 9h–j) and an AUC value of 1.000 (Figure 9j). The classification model highlighted OGG‐110 and OGG‐95 as two candidate variables for distinguishing changes of organic acids, showing moderate predictive power with gain values of 0.34 and 0.25 and both AUC of 1.000 (Figure 9j,k–m). In addition, the XGBoost screened out OGG‐24 and OGG‐3 as the candidate features, showing a clear relationship with changes of phenolic acids. Its predictive importance was validated by gain values of 0.31 and 0.24 and AUC values of 0.646 (Figure 9n–q).

FIGURE 9.

FIGURE 9

Screening of candidate OGGs related to differential accumulated metabolites in nine fruits comprised of three wild and six cultivated green jujube accessions through eXtreme gradient boosting (XGBoost). (a) XGBoost learning curve plot related to alkaloids. (b) Histogram in XGBoost importance of feature gene related to alkaloids. (c) Receiver operating characteristic (ROC) curves of OGG‐4. (d) XGBoost learning curve plot related to flavonoids. (e) Histogram in XGBoost importance of feature gene related to flavonoids. (f) ROC curves of OGG‐178. (g) ROC curves of OGG‐713. (h) XGBoost learning curve plot related to nucleotides. (i) Histogram in XGBoost importance of feature gene related to nucleotides. (j) ROC curves of OGG‐95. (k) XGBoost learning curve plot related to organic acids. (l) Histogram in XGBoost importance of feature gene related to organic acids. (m) ROC curves of OGG‐110. (n) XGBoost learning curve plot related to phenolic acids. (o) Histogram in XGBoost importance of feature gene related to phenolic acids. (p) ROC curves of OGG‐24. (q) ROC curves of OGG‐3. Regarding the histogram, the higher the score, the higher the importance ranking and the greater the contribution to the prediction result. The red bar inside the histogram indicates the OGG with the highest performance and blue bar indicates the performance of other OGGs. ROC curves of OGG‐95 is shown in (j). Cox‐nloglik, negative log‐likelihood of Cox proportional hazards regression. Iter, iteration. Gain: the importance of feature, the average gain of a feature across all trees, indicating the contribution of that feature to the model's predictions.

The subsequent correlation between OGG expression and metabolite abundance was explored to validate their prediction (Figure 10). The results presented that most of the correlation coefficients (r) between metabolites and OGGs were nonsignificant, with r < 0.50 and/or p > 0.05. Only the correlations between alkaloids and OGG‐4 (r = 0.74) and between nucleotides and OGG‐95 (r = 0.79) were strong and statistically significant (both p < 0.01) (Figure 10a,d). Additionally, the correlation coefficients between flavonoids and OGG‐178 were 0.41 with p < 0.05 (Figure 10b), while the correlation coefficients between flavonoids and OGG‐713 were −0.02, with p = 0.94 (Figure 10c). And the similar results were also found between phenolic acids and OGG‐24/OGG‐3 (Figure 10 g,h). These results suggested that their high XGBoost importance scores may be overfitted due to the small sample size.

FIGURE 10.

FIGURE 10

Correlation of metabolites’ accumulated abundances and OGGs’ expression levels in nine fruits comprised by three wild and six cultivated green jujube accessions. (a) Alkaloids and OGG‐4. (b) Flavonoids and OGG‐178. (c) Flavonoids and OGG‐713. (d) Nucleotides and OGG‐95. (e) Organic acids and OGG‐110. (f) Organic acids and OGG‐95. (g) Phenolic acids and OGG‐24. (h) Phenolic acids and OGG‐3. The purple bar in the left panel indicated the expression levels of OGGs, and the green bar in the right panel indicated the accumulation abundances of metabolites. The dashed line separated cultivated samples (left) from wild samples (right).

4. DISCUSSION

Understanding the structure variations of genome during the process from wild to cultivated crop is fundamental for uncovering crop development and growth mechanism (Cao et al., 2025; Cheng et al., 2025; Long et al., 2024; Su et al., 2024; Y. Wang, Xia, et al., 2025). Recent investigations have primarily focused on the global variants encompassing all gene families of several or a dozen genomes, especially for the plant immune genes (Long et al., 2024; Su et al., 2024). However, a comprehensive study on large‐scale transcription factors based on the plant haplotype‐resolved genome was generally lacking, even though there were several limited reports on single transcription factor, such as the bZIP superfamily in Cucurbitaceae (M. Sun et al., 2025) and the bHLH superfamily in barley (Tong et al., 2025). In green jujube, the genomic evolution was characterized, revealing thousands of structure variations between wild and cultivated accessions, and many fruit astringency‐related genes were lost during the process from wild to cultivated (M. Guo et al., 2024). Nonetheless, the changing trajectory of transcription factor genes of green jujube has not yet been comprehensively analyzed. In this study, a haplotype‐resolved genome and transcriptome atlas framework was applied to characterize the variation of the gene content and expression profile of the entire transcription factor superfamily, which revealed the high structural conservation of the transcription factor superfamily between two specific wild and cultivated green jujube HapGenomes, even though there exists a certain level of CNV and PAV.

4.1. Conservation of transcription factors between the two wild and cultivated accessions

This study found high conservation in gene contents of 42 transcription factor superfamilies, including a total of 12,123 gene copies between the eight HapGenomes of the one wild and one cultivated green jujube accession (Figures S1 and S2; Tables S2 and S3). These 42 superfamilies nearly covered all the known plant transcription factors, and the average number was around 1500 in a single HapGenome, aligning to the previous investigation in kiwifruit (Brian et al., 2021) and Arabidopsis (Zenker et al., 2025).

Within 42 transcription factor superfamilies, families involved in signaling and transcriptional regulation, such as Jumonji, NAC, and AUX/IAA, displayed high fluctuation, likely reflecting environmental‐adaptive evolution in the specific genome. Conversely, the extreme stability of superfamilies like STAT, LFY, and E2F/DP suggested they represent core regulatory modules indispensable for basic cellular functions (Figure 1). Further orthologous analysis indicated that the high variation in the Jumonji, NAC, and AUX/IAA superfamilies could be attributed to their greater proportion of dispensable OGGs compared to other superfamilies, the proportion of which was 85.57%, 65.63%, and 91.38%, respectively (Figure S6). It suggested these superfamilies experienced relaxed evolution pressure between the two wild and cultivated green jujube accessions, implying their functional plasticity in green jujube. In contrast, a haplotype‐resolved genome‐wide study across 20 barley genomes found 21.39% NACs were classified into dispensable OGGs (Liu et al., 2025), possibly resulting from the difference in sample sizes.

More importantly, a key finding was that the total copy number of transcription factors showed no obvious difference between the two wild and cultivated green jujube accessions. The observed differences were limited to just four superfamilies (AP2, NAC, and B3) (Figure 2 and Figure S3), implying conservation between these two specific accessions, a conclusion further supported by observed variations in chromosomal distribution (Figure S4). Similarly, several haplotype‐resolved genome analyses based on a specific transcription factor superfamily also reported consistent findings (Liu et al., 2025; M. Sun et al., 2025; Tong et al., 2025; T. Wang, Zheng, et al., 2025). Specifically, the total number of bZIP was reported from 48 to 59 in nine haplotype‐resolved genomes of Cucurbitaceae crop with the CV value of 0.08 (M. Sun et al., 2025), and the total number of bHLHs ranged from 161 to 176 in 20 haplotype‐resolved genomes of barley (Tong et al., 2025). Nevertheless, the CNV and PAV analysis was solely based on a single wild accession and a single cultivated accession in this study, and therefore, future studies with large‐scale wild and cultivated populations were needed to determine whether these CNV and PAV differences were general features or accession‐specific.

Furthermore, this study observed that 66.57% of OGGs were shared among all the eight HapGenomes, while 5.76% of them were specific to a single HapGenome, and the presence of PAV in the haplotype‐resolved genomes (Figure 3). In addition, a total of 78% gene copies were present in all eight HapGenomes between the two wild and green jujube accessions (Figure 4). Correspondingly, most research showed the similar percentage in count of core, dispensable, and private genes (Liu et al., 2025; M. Sun et al., 2025; Tong et al., 2025; T. Wang, Zheng, et al., 2025), further supporting the high conservation in the HapGenomes of the two specific green jujube accessions. Most of OGGs and gene copies were enriched in similar biological processes, including genetic information processing, plant hormone signal transduction, and environmental information processing (Figure S5). Overall, these findings might suggest a high degree of conservation in gene content of HapGenomes in the two specific wild and cultivated green jujube accessions and underscore the importance of haplotype‐resolved genome‐wide analysis for understanding changes in transcription factor gene superfamily between the two wild and cultivated green jujube accessions.

4.2. Natural selection of transcription factors between the two wild and cultivated accessions

Gene duplication is a principal driver of plant diversification, occurring primarily through WGD and small‐scale events (tandem, dispersed, and proximal). WGD has repeated throughout plant evolution, generating genetic redundancy, which, against the genomic instability of polyploidization, enables neofunctionalization, and ultimately contributes to genetic robustness and species survival. However, the haplotype‐resolved genomic analysis of green jujube implied that dispersed duplication, not WGD/segmental duplication, maybe the predominant force in the expansion of transcription factor genes in these two specific green jujube accessions. Dispersed duplication comprised 47.24% of events compared to 36.10% for WGD/segmental duplication, while other duplication types were minimal (Figure 5 and Figure S7). Given that green jujube is an autotetraploid fruit crop species, the predominance of dispersed duplication over WGD/segmental duplication was initially unexpected.

Moreover, the evolutionary analysis of transcription factor genes suggested a strong signal of purifying selection, with the vast majority exhibiting Ka/Ks ratios below 1 (Figure 6a). This indicated strong functional constraint, a typical mark of essential regulatory genes. Consistent with this, a comparative analysis of OGGs showed that dispensable OGGs possessed significantly greater evolutionary flexibility than core OGGs, as shown by their higher Ka/Ks values (Figure 6b). These results suggested that core OGGs tend to experience stronger purifying selection than dispensable OGGs, which was consistent with their essential functions. Whether dispensable OGGs were more likely to contribute to environmental‐adaptive divergence remained a hypothesis that requires functional validation.

Further supporting the assumptions of conservation in the two specific accessions, this study found no significant difference in selective pressure between the two wild and cultivated green jujube accessions, either within or between them (Figure 6a,c,d). These unexpected results indicated that the changes between the two specific wild and cultivated green jujube accessions have not imposed a strong, genome‐wide selective pressure on transcription factor genes, highlighting the high degree of conservation in this critical regulatory network. However, considering Ka/Ks reflect long‐term evolutionary constraints rather than recent selection, and only one wild and one cultivated accession were selected in this study, the possibility that the process from wild to cultivated affected regulatory network was not ruled out.

This study observed significant divergence in evolutionary rates among different transcription factor superfamilies (Figure S8), implying distinct evolutionary dynamics and potential differences in functional conservation. Additionally, the phylogenetic tree showed that many OGGs lost between the two wild and cultivated green jujube accessions (Figure S9). These variations among superfamilies presented a promising field for future research to elucidate the specific drivers of specific transcription factor family evolution within the genus Ziziphus.

Collectively, this analysis revealed conservative patterns among duplication types between the two wild and cultivated green jujube accessions. The dispersed duplication accounted for the largest proportion (47.24%) of transcription factor gene copies, whereas WGD/segmental duplication accounted for 36.10%. This observation, together with the predominantly purifying selection (Ka/Ks < 1) across most gene pairs, suggested that the two duplication types may have contributed differently to the evolution of transcription factor families. One hypothesis is that after polyploidization, different duplication modes have experienced different evolutionary fates, with dispersed duplications contributing more to recent transcription factor family expansion (Bird et al., 2018; Dhatterwal et al., 2024; Golicz et al., 2020; Panchy et al., 2016). However, phylogenomic analyses across a broader population of green jujube accessions, along with formal duplication‐dating and retention‐bias analyses, would be required to validate this model.

4.3. Functional implications of transcription factor genes between wild and cultivated green jujube

This study established the first global transcriptome atlas for green jujube, characterizing the expression landscape of transcription factor genes across multiple tissues and diverse accessions (Figures 7 and 8 and Figure S10). The majority of core transcription factor OGGs were expressed in at least one tissue, whereas the proportion of dispensable and private transcription factor OGGs that were expressed in at least one tissue tended to be lower (Figure S10). This pattern aligned with the predicted roles of core OGGs in fundamental biological processes, which required stable, constitutive expression. Notably, the expression profiles of core, dispensable, and private OGGs themselves showed no tissue specificity (Figure S10), suggesting that OGG PAV in the haplotype‐resolved genome was not a primary determinant of tissue‐specific transcriptional profiles. Instead, tissue‐specific regulation emerged at the level of transcription factor superfamilies. Different superfamilies exhibited unequal expression proportions across tissues (Figure 7) with distinct tissue‐specific expression profiles for individual OGGs (Figure S10d). This indicated a critical functional divergence among transcription factor superfamilies, where specific superfamilies tended to regulate developmental or physiological processes unique to certain tissues.

Furthermore, a comparative haplotype‐resolved transcriptome atlas of core OGGs analysis between three wild and six cultivated accessions revealed significant species‐specific expression patterns (Figure 8a). These divergent transcript levels suggested that the regulatory networks of green jujube have changed from wild to cultivated accessions, potentially underlying key phenotypic differences in fruit traits, stress tolerance, and plant architecture between the two wild and cultivated accessions. Similarly, the specific expansion of dispensable barley bHLH may be directly related to environmental adaptation and CNV/PAV may contribute to variation in stress tolerance (Tong et al., 2025). Together, these findings implied that the functional diversification of the specific transcription factor superfamilies between the two wild and cultivated green jujube accessions was influenced by the genomic context (core/dispensable) of genes themselves. Nevertheless, this OGG‐level aggregation might mask intra‐OGG expression heterogeneity, and that future copy/allelic‐specific expression analyses were needed to resolve such finer regulatory dynamics.

While the genomic structural analyses were limited to a single wild and a single cultivated accession, the transcriptome and metabolome datasets include multiple independent accessions. Consequently, the changes in gene expression and metabolite accumulation were supported by broader sampling, whereas the structural differences should be interpreted as differences between the two specific genomes studied. Future studies with additional large‐scale wild and cultivated accessions will be required to determine whether the observed CNV and PAV patterns were general features of green jujube transition.

4.4. Screen of potential candidate regulatory genes via machine learning

Machine learning provided a powerful analytical framework for discerning complex patterns within high‐dimensional datasets, significantly enhancing feature discrimination and classification accuracy. Previous studies have successfully harnessed machine learning models across various fields of plant research, such as geographical origin and botanical traceability (Y. S. Huang, An, et al., 2025; Lou et al., 2026; Ratnasekhar et al., 2025; Xue et al., 2026); identified transcriptional predictors of plant responding to abiotic and biotic stress (Ahn et al., 2025; El‐Khadir et al., 2026; Farhan et al., 2025; Ikram et al., 2025; Karlström et al., 2025); and selected gene markers for classifying disease‐resistant genotypes (Dublino et al., 2025) and for predicting complex quantitative traits (Mishra et al., 2026; Zhou et al., 2026).

Using a stringent SVM with LOO‐CV, this study found that the global OGG expression profile could discriminate wild from cultivated fruits with an accuracy of 89% (p = 0.04). Although the sample size is limited (nine accessions), the statistically significant performance suggested that the process from wild to cultivated has left a detectable transcriptional footprint on the fruit regulatory network. Among individual OGGs, OGG‐95 was significantly upregulated in cultivated fruits (p < 0.05) and was prioritized by XGBoost as a top feature associated with nucleotide (r = 0.74, p < 0.01) and organic acid accumulation (r = 0.79, p < 0.01) (Figures 9 and 10). Functional annotation revealed OGG‐95 as a member of the Lesion Simulating Disease (LSD) transcription factor family, a class of small C2C2 Zn finger proteins known to be central regulators of programmed cell death, oxidative stress homeostasis, and biotic/abiotic stress responses (Akbar et al., 2025; S. Sun et al., 2023; Zeng et al., 2022). The recent studies reported that LSD can modulate seed vigor in alfalfa (Medicago sativa) (S. Sun et al., 2023), coordinate the oxidative stress response in cassava (Zeng et al., 2022), and show a significant correlation with sucrose transporter activity in sugarcane (Akbar et al., 2025). Therefore, the upregulation of OGG‐95 (LSD) in fruits of the six cultivated green jujube accessions and its association with differential accumulation of nucleotide and organic acid suggested a possible role that requires experimental validation in modulating fruit development and ripening processes of green jujube.

Overall, OGG‐95 (LSD) emerged not merely as a statistical classifier but as a potential candidate that related to metabolic phenotypes between these wild and cultivated green jujube accessions. Considering the loss of green jujube fruit astringency in the process from wild to cultivated (M. Guo et al., 2024), the specific correlation with organic acids was especially noteworthy, as these compounds were key determinants of fruit flavor, acidity, and post‐harvest physiology. These results prioritized the LSD family as a promising potential candidate for understanding the difference of the regulatory networks related to fruit characteristics. While these integrative bioinformatic and modeling results provided preliminary in silico clues, targeted functional validation through techniques such as gene silencing/overexpression, coupled with detailed metabolomic profiling in engineered plants, was essential to definitively establish the causal role of OGG‐95 (LSD) in governing these fruit quality traits.

However, the SVM classification and XGBoost association analyses in this study were both limited by the small sample size (nine accessions). The high accuracy (89%) and perfect AUC values observed in some XGBoost models likely reflect overfitting rather than generalizable performance. For instance, OGG‐713 was identified by XGBoost as an important feature for flavonoid accumulation (gain = 0.21 and AUC = 0.965), yet its direct Pearson correlation with total flavonoid content was negligible (r = −0.02, p = 0.94). Therefore, these models were not recognized as validated predictors but as exploratory tools to prioritize candidates. Future studies with larger sample sizes and functional experiments will be required to validate these candidates.

5. CONCLUSIONS

This study provided a comprehensive, multi‐dimensional analysis of the evolution and regulation of transcription factor superfamilies identified in one wild and one cultivated autotetraploid green jujube accession. By moving beyond a single reference genome, the haplotype‐resolved genome revealed a landscape of high conservation in transcription factor gene content of two specific accessions, countered by the unexpected predominance of dispersed duplication as a key evolutionary force. The integrated haplotype‐resolved transcriptome atlas proposed how functional diversity operates at the level of transcription factor superfamilies, orchestrating tissue‐specific and lineage‐specific expression profiles. Importantly, the application of an SVM with LOO‐CV suggested that OGG expression profiles can significantly discriminate wild from cultivated fruit. Combined with XGBoost prioritization and correlation validation, this study screened out the OGG‐95 (LSD) gene family as a high‐priority candidate gene family for future functional validation to determine its potential role in metabolic divergence between the wild and cultivated fruits. Our findings underscored that in the two specific green jujube accessions, the evolutionary trajectory of transcription factor gene superfamilies was shaped by an interplay of duplication mechanisms and selective pressures, not merely by their polyploid. This work established a powerful integrative model, haplotype‐resolved genome, haplotype‐resolved transcriptome atlas, and machine learning for dissecting the genetic basis of complex traits, offering a robust foundation for future functional validation and precision breeding in green jujube and other polyploid species.

AUTHOR CONTRIBUTIONS

Xudong Zhu: Conceptualization; formal analysis; methodology; resources; writing—original draft; writing—review and editing. Pengyan Chang: Visualization; formal analysis; software; validation. Yanzhen Liao: Resources. Huini Wu: Formal analysis; software. Fan Jiang: Project administration; supervision; writing—review and editing.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflicts of interest.

Supporting information

Figure S1. Venn diagram of transcription factors identified across eight HapGenomes from the two wild and cultivated green jujube accessions using three unique tools. Four HapGenomes of the cultivated green jujube accession, (a) cI, (b) cII, (c) cIII and (d) cIV, and four HapGenomes of the wild green jujube accession, (e) wI, (f) wII, (g) wIII and (h) wIV.

Figure S2. Copy number of 42 transcription factor superfamilies identified across eight HapGenomes from the two wild and cultivated green jujube accessions. The cI, cII, cIII and cIV was the four HapGenomes of the cultivated green jujube accession. The wI, wII, wIII and wIV was the four HapGenomes of the wild green jujube accession.

Figure S3. Comparison in the copy numbers of remained 19 transcription factor families between wild and cultivated green jujube.

Figure S4. Comparison of chromosome distribution of transcription factor superfamilies between the two cultivated and wild green jujube accessions. (a) Stacked line chart of copy number of transcription factor genes in 12 chromosomes of 8 HapGenomes. (b) Per‑chromosome comparison of copy number in 12 chromosomes between the two wild and cultivated green jujube accessions. Each dot represented one HapGenome and bars indicated median of copy number. (c) Distribution of transcription factor gene copies in Chr01 in the two wild and cultivated accessions.

Figure S5. KEGG pathway enrichment analysis of core, dispensable and private OGGs in the two wild and cultivated green jujube accessions. (a) Comparison of pathway enrichment between core, dispensable and private OGGs. (b) Comparison of pathway enrichment of core OGGs between the two wild and cultivated accessions. (c) Comparison of pathway enrichment of dispensable OGGs between the two wild and cultivated accessions. (d) Comparison of pathway enrichment of private OGGs between the two wild and cultivated accessions. Red to blue bar indicated the p value and the circle size indicated the number of OGGs.

Figure S6. Distribution of core, dispensable and private gene copies in 42 transcription factor superfamilies.

Figure S7. Syntenic relationships of OGGs across 8 HapGenomes in the two wild and cultivated green jujube accessions.

Figure S8. Comparison of Ka/Ks values of transcription factor superfamilies. (a) Comparison of Ka/Ks values between Zn finger, MYB, bHLH and AP2 family (copy number > 100). (b) Comparison of Ka/Ks values between 6 superfamilies for Homobox, NAC, B3, WRKY, MADS and FAR with copy number ranged from 50 to 100. (c) Comparison of Ka/Ks values between 13 superfamilies with copy number ranged from 10 to 50. (d) Comparison of Ka/Ks values between 15 superfamilies with copy number ranged < 10. Kruskal‐Wallis H test was conducted for significant analysis at p < 0.05, ns: no significant. The LFY, S1Fa‐like and STAT was found no collinearity relationships, and no Ka/Ks value was calculated in NZZ/SPL superfamily.

Figure S9. Phylogenetic relationship and OGG contraction and expansion results among 8 HapGenomes. The blue minus sign indicated the number of contracted OGGs, and the red plus sign indicated the number of expanded OGGs. The 4 Mya divergence time represented the lineage divergence of the wild and cultivated ancestors.

Figure S10. Expression profiles of core OGGs in leaf, branch and mature fruit of the two wild and cultivated green jujube accessions. (a) Proportions of core OGGs expressed in leaf, branch and mature fruit of the cultivated green jujube accession. (b) Proportions of OGGs expressed in leaf, branch and mature fruit of the wild green jujube accession. (c) Rader map of number of tissue‐specific OGGs between the two wild and cultivated green jujube accessions based on their TAU. (d) K‐means cluster analysis of core OGGs in leaf, branch and mature fruit of the two wild and cultivated green jujube accessions.

Figure S11. The significance of difference in different metabolite profiles from three wild and six cultivated green jujube accessions that were evaluated using Wilcoxon rank sum test.

TPG2-19-e70266-s001.docx (20.8MB, docx)

Table S1. The SRA lists of all tissues and accessions used in this study.

Table S2. Identified of transcription factor superfamilies through iTAK, PlantTFcat and PlantTFDB database.

Table S3. Summary of 42 transcription factor superfamilies and their subclasses.

Table S4. Composition of gene member in each OGGs.

Table S5. Distribution of all Ka/Ks values between transcription factor gene pairs among wild and cultivated green jujube.

Table S6. Expression profiles of all transcription factor genes in different tissues of cultivated green jujube.

Table S7. Expression profiles of all transcription factor genes in different tissues of wild green jujube.

Table S8. Haplotype‐resolved transcriptome atlas of all OGGs in different tissues of wild and cultivated green jujube.

Table S9. Haplotype‐resolved transcriptome atlas expression profiles of all OGGs in different fruits of wild and cultivated green jujube.

Table S10. Detailed results of leave‑one‑out cross‑validation (LOO‑CV).

Table S11. Differential expression analysis of all OGGs between wild and cultivated fruits and their SVM feature weights across the nine leave‑one‑out folds.

Table S12. Metabolome profiles in fruits from wild and cultivated green jujube.

TPG2-19-e70266-s002.xlsx (2.8MB, xlsx)

ACKNOWLEDGMENTS

This research was financially supported by a grant from the Natural Science Foundation Project of Zhangzhou City (ZZ2025JH09), Basic Scientific Research Project of Fujian Public Welfare Scientific Research Institute (2025R1028002 and 2025R1028003), National Tropical Plants Germplasm Resource Center (NTPGRC2025‐032), and National Germplasm Repository for Fujian‐Taiwan Characteristic Crops (Zhangzhou).

Zhu, X. , Chang, P. , Liao, Y. , Wu, H. , & Jiang, F. (2026). Haplotype‑resolved comparison of transcription factor superfamilies between wild and cultivated autotetraploid green jujube and prioritization of candidate transcription factors via machine learning. The Plant Genome, 19, e70266. 10.1002/tpg2.70266

Assigned to Associate Editor Abdulqader Jighly.

DATA AVAILABILITY STATEMENT

The raw transcriptome data of different tissues used in this work have been accessed in the NCBI Sequence Read Archive under accession number PRJNA979989. The genomic sequences and annotation datasets of wild and cultivated green jujube genome are available at figshare (https://doi.org/10.6084/m9.figshare.23530068). The metabolomic data of all samples have been deposited to the MetaboLights database with the identifier MTBLS10827 (https://www.ebi.ac.uk/metabolights/editor/study/MTBLS10827). All other relevant data are included in the manuscript and its supplementary materials.

REFERENCES

  1. Ahn, E. , Baek, I. , Prom, L. K. , Park, S. , Kim, M. S. , Meinhardt, L. W. , & Magill, C. (2025). Candidate genes for anthracnose resistance in Senegalese sorghum: A machine learning‐based exploration. Functional & Integrative Genomics, 26(1), Article 5. 10.1007/s10142-025-01797-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Akbar, S. , Hua, X. , Zhang, Y. , Liu, G. , Wang, T. , Shi, H. , Li, Z. , Qi, Y. , Habiba, H. , Yao, W. , Zhang, M. Q. , & Zhang, J. (2025). Genome‐wide analysis of sugar transporter gene family in Erianthus rufipilus and Saccharum officinarum, expression profiling and identification of transcription factors. Frontiers in Plant Science, 15, 1502649. 10.3389/fpls.2024.1502649 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bird, K. A. , VanBuren, R. , Puzey, J. R. , & Edger, P. P. (2018). The causes and consequences of subgenome dominance in hybrids and recent polyploids. New Phytologist, 220(1), 87–93. 10.1111/nph.15256 [DOI] [PubMed] [Google Scholar]
  4. Bray, N. L. , Pimentel, H. , Melsted, P. , & Pachter, L. (2016). Near‐optimal probabilistic RNA‐seq quantification. Nature Biotechnology, 34(5), 525–527. 10.1038/nbt.3519 [DOI] [PubMed] [Google Scholar]
  5. Brian, L. , Warren, B. , McAtee, P. , Rodrigues, J. , Nieuwenhuizen, N. , Pasha, A. , David, K. M. , Richardson, A. , Provart, N. J. , Allan, A. C. , Varkonyi‐Gasic, E. , & Schaffer, R. J. (2021). A gene expression atlas for kiwifruit (Actinidia chinensis) and network analysis of transcription factors. BMC Plant Biology, 21(1), Article 121. 10.1186/s12870-021-02894-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Cantalapiedra, C. P. , Hernández‐Plaza, A. , Letunic, I. , Bork, P. , & Huerta‐Cepas, J. (2021). eggNOG‐mapper v2: Functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Molecular Biology and Evolution, 38(12), 5825–5829. 10.1093/molbev/msab293 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Cao, Y. , Yan, J. , Ross‐Ibarra, J. , & Yang, N. (2025). Plant domestication revisited: Genomic insights into origins, mechanisms, and convergent evolution. iScience, 29(1), 114062. 10.1016/j.isci.2025.114062 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Capella‐Gutiérrez, S. , Silla‐Martínez, J. M. , & Gabaldón, T. (2009). trimAl: A tool for automated alignment trimming in large‐scale phylogenetic analyses. Bioinformatics, 25(15), 1972–1973. 10.1093/bioinformatics/btp348 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Chen, S. , Zhou, Y. , Chen, Y. , & Gu, J. (2018). fastp: An ultra‐fast all‐in‐one FASTQ preprocessor. Bioinformatics, 34(17), i884–i890. 10.1093/bioinformatics/bty560 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cheng, Y. J. , Zhang, Y. , Yu, J. , Meng, X. R. , Chang, Z. Y. , Jiang, F. L. , Wu, Z. , Zhang, Y. Y. , Xue, J. Y. , & Zhang, T. Q. (2025). Chromatin accessibility analysis reveals functional cis‐regulatory regions related to fruit development and domestication in tomato. The Plant Journal, 123(3), e70416. 10.1111/tpj.70416 [DOI] [PubMed] [Google Scholar]
  11. Dai, X. , Sinharoy, S. , Udvardi, M. , & Zhao, P. X. (2013). PlantTFcat: An online plant transcription factor and transcriptional regulator categorization and analysis tool. BMC Bioinformatics, 14, Article 321. 10.1186/1471-2105-14-321 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Dhatterwal, P. , Sharma, N. , & Prasad, M. (2024). Decoding the functionality of plant transcription factors. Journal of Experimental Botany, 75(16), 4745–4759. 10.1093/jxb/erae231 [DOI] [PubMed] [Google Scholar]
  13. Dublino, R. , D'Esposito, D. , Guadagno, A. , Capuozzo, C. , Crinò, P. , Formisano, G. , & Ercolano, M. R. (2025). AI‐integrated omics analysis reveals cultivar‐specific resistance mechanisms to powdery mildew in Cucurbita pepo . International Journal of Molecular Sciences, 26(23), 11488. 10.3390/ijms262311488 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Edgar, R. C. (2004). MUSCLE: Multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32(5), 1792–1797. 10.1093/nar/gkh340 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. El‐Khadir, I. , Mouniane, Y. , El Bakkali, M. , Chriqui, A. , & Hmouni, D. (2026). Improvement of salt stress tolerance in Salvia officinalis L. by co‐culture with halophytic plants: Comparative study with Aptenia cordifolia and Spergularia salina . Journal of the Science of Food and Agriculture. 10.1002/jsfa.70495 [DOI] [PubMed] [Google Scholar]
  16. Emms, D. M. , & Kelly, S. (2019). OrthoFinder: Phylogenetic orthology inference for comparative genomics. Genome Biology, 20(1), Article 238. 10.1186/s13059-019-1832-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Fan, H. , Li, J. , Huang, W. , Liang, A. , Jing, L. , Li, J. , Yang, Q. Y. , Liu, K. , & Yang, Z. (2025). Pan‐genome analysis of the R2R3‐MYB genes family in Brassica napus unveils phylogenetic divergence and expression profiles under hormone and abiotic stress treatments. Frontiers in Plant Science, 16, 1588362. 10.3389/fpls.2025.1588362 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Farhan, M. , Ikram, M. , Tung, S. A. , Ren, M. J. , & Wang, Y. (2025). Identification of stem rust‐associated candidate genes in wheat (Triticum aestivum L.) via RNA‐sequencing, cluster analysis, weighted gene co‐expression networks, and machine learning models. Plant Physiology and Biochemistry, 229(Pt D), 110701. 10.1016/j.plaphy.2025.110701 [DOI] [PubMed] [Google Scholar]
  19. Golicz, A. A. , Bayer, P. E. , Bhalla, P. L. , Batley, J. , & Edwards, D. (2020). Pangenomics comes of age: From bacteria to plant and animal applications. Trends in Genetics, 36(2), 132–145. 10.1016/j.tig.2019.11.006 [DOI] [PubMed] [Google Scholar]
  20. Guo, L. , Song, Q. , Zhang, X. , Tian, T. , Zhang, Y. , Wu, Y. , Zhang, P. , Ma, J. , Chen, T. , & Yang, D. (2025). Pangenome identification and functional characterization of AHL genes in wheat (Triticum aestivum L.) reveal the role of TaAHL67 in grain weight regulation. BMC Plant Biology, 26(1), Article 25. 10.1186/s12870-025-07817-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Guo, M. , Bi, G. , Wang, H. , Ren, H. , Chen, J. , Lian, Q. , Wang, X. , Fang, W. , Zhang, J. , Dong, Z. , Pang, Y. , Zhang, Q. , Huang, S. , Yan, J. , & Zhao, X. (2024). Genomes of autotetraploid wild and cultivated Ziziphus mauritiana reveal polyploid evolution and crop domestication. Plant Physiology, 196(4), 2701–2720. 10.1093/plphys/kiae512 [DOI] [PubMed] [Google Scholar]
  22. Han, H. , Salinas, N. , Barbey, C. R. , Jang, Y. J. , Fan, Z. , Verma, S. , Whitaker, V. M. , & Lee, S. (2025). A telomere‐to‐telomere phased genome of an octoploid strawberry reveals a receptor kinase conferring anthracnose resistance. GigaScience, 14, giaf005. 10.1093/gigascience/giaf005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Han, X. , Qiu, C. , Gai, Z. , Zhai, J. , Song, J. , Sun, J. , & Li, Z. (2025). Pan‐Genome‐based characterization of the PYL transcription factor family in Populus. Plants, 14(16), 2541. 10.3390/plants14162541 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Hao, Z. , Lv, D. , Ge, Y. , Shi, J. , Weijers, D. , Yu, G. , & Chen, J. (2020). RIdeogram: Drawing SVG graphics to visualize and map genome‐wide data on the idiograms. PeerJ Computer Science, 6, e251. 10.7717/peerj-cs.251 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Huang, M. , Zheng, P. , Li, N. , Chen, Q. , Liu, Y. , Huang, B. , Tao, X. , Yu, J. , & Xu, S. (2025). Evolutionary dynamics and functional divergence of the UDP‐glycosyltransferases gene family revealed by a pangenome‐wide analysis in tomato. Horticulture Research, 12(11), uhaf204. 10.1093/hr/uhaf204 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Huang, Y. S. , An, Y. L. , Zheng, Y. Y. , Zhao, W. J. , Song, C. Q. , Zhang, L. J. , Chen, J. T. , Tang, Z. J. , Feng, L. , Li, Z. W. , Liu, X. K. , Zhang, D. D. , & Guo, D. A. (2025). A holistic strategy for the in‐depth discrimination and authentication of 16 citrus herbs and associated commercial products based on machine learning techniques and non‐targeted metabolomics. Journal of Chromatography A, 1745, 465747. 10.1016/j.chroma.2025.465747 [DOI] [PubMed] [Google Scholar]
  27. Ikram, M. , Farhan, M. , Derakhshani, B. , Kumar, S. , Khan, N. , Gupta, R. , Usman, B. , & Liu, P. (2025). Machine learning and CRISPR‐based validation elucidate OsWOX13 involvement in rice heat stress tolerance and flowering. Physiologia Plantarum, 177(6), e70714. 10.1111/ppl.70714 [DOI] [PubMed] [Google Scholar]
  28. Javed, J. , Wang, G. , Sun, S. , Saleem, H. , Baloch, F. S. , Xue, X. , & Gou, J. Y. (2026). Pan‐genomic diversity of the EPF/EPFL gene family across wild and modern wheat species. BMC Plant Biology, 26(1), Article 402. 10.1186/s12870-025-08046-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Jiang, F. , Zhu, X. , Wu, M. , Chang, P. , Wu, H. , & Li, H. (2025). Domestication has reshaped gene families, gene expressions and flavonoid metabolites in green jujube (Ziziphus mauritiana Lam.) fruit. Horticulturae, 11(8), 974. 10.3390/horticulturae11080974 [DOI] [Google Scholar]
  30. Karlström, A. , Gómez‐Cortecero, A. , Connell, J. , Nellist, C. F. , Ordidge, M. , Dunwell, J. M. , & Harrison, R. J. (2025). Identifying key genes for European canker resistance in apple: Machine learning and gene expression profiling of quantitative disease resistance. Scientific Reports, 16(1), Article 3419. 10.1038/s41598-025-33478-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Li, R. , Lei, C. , Zhang, Q. , Guo, X. , Cui, X. , Wang, X. , Li, X. , & Gao, J. (2025). Pan‐genome‐based characterization of the SRS transcription factor family in foxtail millet. Plants, 14(8), 1257. 10.3390/plants14081257 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Li, T. , Liao, W. , Wang, H. , Wang, Z. , Li, J. , Zhou, X. , Cai, Y. , Zhang, J. , Feng, F. , Wang, Y. , Wang, W. , Hu, J. , & Sun, Y. (2025). Integrating ATAC‐seq and pan‐genomics identifies stress‐memory AP2/ERF hubs in foxtail millet. Food Science & Nutrition, 13(12), e71109. 10.1002/fsn3.71109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Li, W. , Chu, C. , Zhang, T. , Sun, H. , Wang, S. , Liu, Z. , Wang, Z. , Li, H. , Li, Y. , Zhang, X. , Geng, Z. , Wang, Y. , Li, Y. , Zhang, H. , Fan, W. , Wang, Y. , Xu, X. , Cheng, L. , Zhang, D. , … Han, Z. (2025). Pan‐genome analysis reveals the evolution and diversity of Malus. Nature Genetics, 57(5), 1274–1286. 10.1038/s41588-025-02166-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Li, X. , Kuhl, H. , Zhu, S. , Xie, S. , Zhang, C. , Huo, L. , Hu, X. , Gong, J. , Wang, M. , Yin, X. , & Sun, X. (2025). The haplotype‐resolved genome assembly of the hexaploid kiwifruit Actinidia deliciosa reveals its hybrid origin and polysomic inheritance. Plant Communications, 6(10), 101436. 10.1016/j.xplc.2025.101436 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Liu, X. , Zhang, M. , Su, J. , Wu, L. , Shen, M. , Zhuang, Y. , Wang, Q. , & Chen, G. (2025). The evolution, variation, and expression patterns under development and stress responses of the NAC gene family in the barley pan‐genome. Frontiers in Plant Science, 16, 1635416. 10.3389/fpls.2025.1635416 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Long, Q. , Cao, S. , Huang, G. , Wang, X. , Liu, Z. , Liu, W. , Wang, Y. , Xiao, H. , Peng, Y. , & Zhou, Y. (2024). Population comparative genomics discovers gene gain and loss during grapevine domestication. Plant Physiology, 195(2), 1401–1413. 10.1093/plphys/kiae039 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Lou, R. , Wu, L. , Zhang, Q. , Li, K. , & Hou, T. (2026). Identification and quality evaluation of Camellia sinensis cv. Longjing 43 fresh leaves from Enshi with different altitudes based on non‐targeted metabolomics combined with chemometrics. Food Research International, 227, 118276. 10.1016/j.foodres.2025.118276 [DOI] [PubMed] [Google Scholar]
  38. Lu, T. , Jiang, S. , Liu, X. , Lu, Z. , Sadaqat, M. , Chen, C. , Zuo, X. , & Tahir Ul Qamar, M. (2025). Pan‐genome analysis and functional characterisation of the terpene synthase (TPS) gene family in five varieties of yellowhorn (Xanthoceras sorbifolium). Functional Plant Biology, 52(10), FP24349. 10.1071/FP24349 [DOI] [PubMed] [Google Scholar]
  39. Mantica, F. , Iñiguez, L. P. , Marquez, Y. , Permanyer, J. , Torres‐Mendez, A. , Cruz, J. , Franch‐Marro, X. , Tulenko, F. , Burguera, D. , Bertrand, S. , Doyle, T. , Nouzova, M. , Currie, P. D. , Noriega, F. G. , Escriva, H. , Arnone, M. I. , Albertin, C. B. , Wotton, K. R. , Almudi, I. , … Irimia, M. (2024). Evolution of tissue‐specific expression of ancestral genes across vertebrates and insects. Nature Ecology & Evolution, 8(6), 1140–1153. 10.1038/s41559-024-02398-5 [DOI] [PubMed] [Google Scholar]
  40. Mendes, F. K. , Vanderpool, D. , Fulton, B. , & Hahn, M. W. (2021). CAFE 5 models variation in evolutionary rates among gene families. Bioinformatics, 36(22–23), 5516–5518. 10.1093/bioinformatics/btaa1022 [DOI] [PubMed] [Google Scholar]
  41. Mishra, D. S. , Yadav, V. , Berwal, M. K. , Ravat, P. , Singh, A. , Rane, J. , Kaushik, P. , Patel, P. J. , Rao, V. V. A. , Kumar, P. , Goswami, A. K. , Chaturvedi, K. , & Kumar, P. (2026). Integrative multivariate and machine learning analyses reveals trait architecture and selection potential in guava (Psidium Guajava L.). BMC Plant Biology, 26, 1–59. 10.1186/s12870-026-08216-3 [DOI] [PubMed] [Google Scholar]
  42. Panchy, N. , Lehti‐Shiu, M. , & Shiu, S. H. (2016). Evolution of gene duplication in plants. Plant Physiology, 171(4), 2294–2316. 10.1104/pp.16.00523 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Pertea, G. , & Pertea, M. (2020). GFF utilities: GffRead and GffCompare. F1000Research, 9, 304. 10.12688/f1000research.23297.2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Price, M. N. , Dehal, P. S. , & Arkin, A. P. (2010). FastTree 2–approximately maximum‐likelihood trees for large alignments. PLoS ONE, 5(3), e9490. 10.1371/journal.pone.0009490 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Ratnasekhar, C. H. , Rai, A. K. , Rakwal, P. , Khan, S. , Verma, A. K. , Mukhopadhyay, P. , Rathor, P. , Hinghrani, L. , Birse, N. , Trivedi, R. , & Trivedi, P. K. (2025). Machine learning‐guided Orbitrap‐HRAMS‐based metabolomic fingerprinting for geographical origin, variety and tissue specific authentication, and adulteration detection of turmeric and ashwagandha. Food Chemistry, 482, 144078. 10.1016/j.foodchem.2025.144078 [DOI] [PubMed] [Google Scholar]
  46. Roy, A. , Chandra, T. , Mondal, R. , Molla, J. , Jaiswal, S. , Srivastava, M. , Kumar, D. , Molla, K. A. , & Iquebal, M. A. (2025). Genomic resources and genetic improvement of vital tropical and subtropical fruit crops: Current status and prospects. AoB Plants, 17(6), plaf059. 10.1093/aobpla/plaf059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Song, W. , Yang, Q. , Du, M. , Wei, J. , Zhou, B. , Shi, J. , Liang, L. , Liu, Z. , Liang, M. , Li, M. , & Gao, Y. (2025). Identification of MUC5B as a lymph node metastasis‐associated gene in lung adenocarcinoma through integrated transcriptomic and machine learning approaches. Frontiers in Immunology, 16, 1666240. 10.3389/fimmu.2025.1666240 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Song, X. , Zhang, M. , Wang, T. T. , Duan, Y. Y. , Ren, J. , Gao, H. , Fan, Y. J. , Xia, Q. M. , Cao, H. X. , Xie, K. D. , Wu, X. M. , Zhang, F. , Zhang, S. Q. , Huang, Y. , Boualem, A. , Bendahmane, A. , Tan, F. Q. , & Guo, W. W. (2025). Polyploidization leads to salt stress resilience via ethylene signaling in citrus plants. New Phytologist, 246(1), 176–191. 10.1111/nph.20428 [DOI] [PubMed] [Google Scholar]
  49. Su, Y. , Yang, X. , Wang, Y. , Li, J. , Long, Q. , Cao, S. , Wang, X. , Liu, Z. , Huang, S. , Chen, Z. , Peng, Y. , Zhang, F. , Xue, H. , Cao, X. , Zhang, M. , Yisilam, G. , Chu, Z. , Gao, Y. , Zhou, Y. , … Tian, X. (2024). Phased telomere‐to‐telomere reference genome and pangenome reveal an expansion of resistance genes during apple domestication. Plant Physiology, 195(4), 2799–2814. 10.1093/plphys/kiae258 [DOI] [PubMed] [Google Scholar]
  50. Sun, M. , Jiang, Q. , Zhang, E. , Zhu, Z. , Yang, Y. , Tan, C. L. , Wang, Z. , Li, R. , Tao, Y. , & Zhao, Q. (2025). Evolutionary dynamics and functional diversification of bZIP transcription factors in Cucurbitaceae: A pan‐genome approach. Journal of Agricultural and Food Chemistry, 73(50), 32363–32378. 10.1021/acs.jafc.5c10450 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Sun, S. , Ma, W. , Jia, Z. , Ou, C. , Li, M. , & Mao, P. (2023). Genomic identification and expression profiling of Lesion Simulating Disease genes in alfalfa (Medicago sativa) elucidate their responsiveness to seed vigor. Antioxidants, 12(9), 1768. 10.3390/antiox12091768 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Tian, F. , Yang, D. C. , Meng, Y. Q. , Jin, J. , & Gao, G. (2020). PlantRegMap: Charting functional regulatory maps in plants. Nucleic Acids Research, 48(D1), D1104–D1113. 10.1093/nar/gkz1020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Tong, C. , Jia, Y. , Hu, H. , Zeng, Z. , Chapman, B. , & Li, C. (2025). Pangenome and pantranscriptome as the new reference for gene‐family characterization: A case study of basic helix‐loop‐helix (bHLH) genes in barley. Plant Communications, 6(1), 101190. 10.1016/j.xplc.2024.101190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Wang, T. , Zheng, Y. , Sun, L. , Guan, M. , Hu, Y. , Yu, H. , Wu, D. , & Du, J. (2025). Analysis of maize PAL pan gene family and expression pattern under lepidopteran insect stress. Frontiers in Plant Science, 16, 1651563. 10.3389/fpls.2025.1651563 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Wang, X. , Cheng, S. , & Du, B. (2025). Pan‐genomic identification and analysis of the maize BBX family. Genes, 17(1), 46. 10.3390/genes17010046 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Wang, Y. , Tang, H. , Debarry, J. D. , Tan, X. , Li, J. , Wang, X. , Lee, T. H. , Jin, H. , Marler, B. , Guo, H. , Kissinger, J. C. , & Paterson, A. H. (2012). MCScanX: A toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Research, 40(7), e49. 10.1093/nar/gkr1293 [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Wang, Y. , Xia, B. , Lin, Q. , Wang, H. , Wu, Z. , Zhang, H. , Zhou, Z. , Yan, Z. , Gao, Q. , Zhang, X. , Wang, S. , Liu, Z. , Meng, X. , Zhang, Y. , Gleave, A. P. , Zhang, H. , & Yao, J. L. (2025). Convergent domestication of bitter apples and pears by selecting mutations of MYB transcription factors to reduce proanthocyanidin levels. Molecular Horticulture, 5(1), Article 51. 10.1186/s43897-025-00173-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Wang, Y. , Zeng, Y. , Wang, Y. , Chen, H. , Xiao, W. , Fang, T. , Zhu, J. , Li, C. , Gao, L. , & Liu, J. H. (2025). Haplotype‐resolved genome and multi‐omics landscape reveal an epigenetic regulation of citric acid accumulation in lemon. Plant Physiology, 199(1), kiaf400. 10.1093/plphys/kiaf400 [DOI] [PubMed] [Google Scholar]
  59. Wang, Z. , Li, X. , Wang, J. , Yu, H. , Zhao, D. , Xu, Y. , Zhou, S. , & Men, W. (2025). Comprehensive molecular characterization of high‐stemness gastric cancer cells using single‐cell transcriptomics, spatial mapping, and machine learning. npj Precision Oncology, 9(1), Article 400. 10.1038/s41698-025-01177-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Xu, L. , Jiang, P. , Xiao, Y. , Wang, H. , Liu, H. , & Chen, H. (2025). CD48, CD69, and TIGIT as diagnostic biomarkers for primary Sjögren's syndrome: An integrated machine learning and multi‐disease discrimination validation study. Frontiers in Immunology, 16, 1700831. 10.3389/fimmu.2025.1700831 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Xue, Q. , Zhu, B. , Li, Z. , Tang, Y. , & Wang, Y. (2026). Machine learning‐enabled metabolomics for geographical authentication of Lonicera japonica via UHPLC‐Q‐TOF‐MS/MS and SHAP interpretation. Food Chemistry, 505, 148062. 10.1016/j.foodchem.2026.148062 [DOI] [PubMed] [Google Scholar]
  62. Zeng, H. , Xu, H. , Wang, H. , Chen, H. , Wang, G. , Bai, Y. , Wei, Y. , & Shi, H. (2022). LSD3 mediates the oxidative stress response through fine‐tuning APX2 activity and the NF‐YC15‐GSTs module in cassava. The Plant Journal, 110(5), 1447–1461. 10.1111/tpj.15749 [DOI] [PubMed] [Google Scholar]
  63. Zenker, S. , Wulf, D. , Meierhenrich, A. , Viehöver, P. , Becker, S. , Eisenhut, M. , Stracke, R. , Weisshaar, B. , & Bräutigam, A. (2025). Many transcription factor families have evolutionarily conserved binding motifs in plants. Plant Physiology, 198(2), kiaf205. 10.1093/plphys/kiaf205 [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Zhang, W. , Wang, J. , Wang, X. , Duan, X. , Peng, J. , Zhang, X. , Xing, Q. , Zhang, K. , & Yan, J. (2025). Chromosome‐level genome assembly of tetraploid Chinese cherry (Prunus pseudocerasus). Scientific Data, 12(1), Article 136. 10.1038/s41597-025-04462-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Zhang, X. , Song, B. , Du, S. , Zhang, S. , Ren, Y. , Xue, C. , Xu, S. , Zheng, P. , Chen, S. , Qiao, Z. , Liu, J. , Wei, W. , & Wu, J. (2025). Genomic insights into deleterious mutations and their impact on agronomic traits during pear domestication. Horticulture Research, 12(9), uhaf140. 10.1093/hr/uhaf140 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Zhang, Z. (2022). KaKs_Calculator 3.0: Calculating selective pressure on coding and non‐coding sequences. Genomics, Proteomics & Bioinformatics, 20(3), 536–540. 10.1016/j.gpb.2021.12.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Zhao, F. , Li, X. , Chen, Z. , & Guo, C. (2025). Pan‐genome‐wide investigation and expression analysis of GATA gene family in maize. Plants, 14(11), 1693. 10.3390/plants14111693 [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Zhao, W. , Wang, J. , Yang, H. , Hou, X. , Zhang, Z. , Chen, J. , Wang, H. , & Yan, C. (2025). Large‐scale analysis of MYB genes in Cucurbitaceae identifies a novel gene regulating plant height. Horticulture Research, 12(11), uhaf210. 10.1093/hr/uhaf210 [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Zheng, Y. , Jiao, C. , Sun, H. , Rosli, H. G. , Pombo, M. A. , Zhang, P. , Banf, M. , Dai, X. , Martin, G. B. , Giovannoni, J. J. , Zhao, P. X. , Rhee, S. Y. , & Fei, Z. (2016). iTAK: A program for genome‐wide prediction and classification of plant transcription factors, transcriptional regulators, and protein kinases. Molecular Plant, 9(12), 1667–1670. 10.1016/j.molp.2016.09.014 [DOI] [PubMed] [Google Scholar]
  70. Zhou, W. , Ouyang, H. , Yan, Z. , Song, J. , Li, Y. , Tang, Z. , & Zhao, C. (2026). Integrative machine learning approach for identifying genes associated with quantitative traits: A soybean (Glycine max) yield case study. The Plant Genome, 19(1), e70178. 10.1002/tpg2.70178 [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Zhu, X. , Wu, M. , Chang, P. , Li, H. , Wu, S. , Huang, H. , & Jiang, F. (2026). Comparative genomics of light harvesting chlorophyll a/b binding protein (LHC) family in green jujube (Ziziphus mauritiana) reveal the effect of domestication based on haplotype‐resolved genome. BMC Plant Biology, 26, Article 393. 10.1186/s12870-026-08189-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Zhu, Z. , & Stein, N. (2025). Pangenome insights into structural variation and functional diversification of barley CCT motif genes. The Plant Genome, 18(3), e70098. 10.1002/tpg2.70098 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1. Venn diagram of transcription factors identified across eight HapGenomes from the two wild and cultivated green jujube accessions using three unique tools. Four HapGenomes of the cultivated green jujube accession, (a) cI, (b) cII, (c) cIII and (d) cIV, and four HapGenomes of the wild green jujube accession, (e) wI, (f) wII, (g) wIII and (h) wIV.

Figure S2. Copy number of 42 transcription factor superfamilies identified across eight HapGenomes from the two wild and cultivated green jujube accessions. The cI, cII, cIII and cIV was the four HapGenomes of the cultivated green jujube accession. The wI, wII, wIII and wIV was the four HapGenomes of the wild green jujube accession.

Figure S3. Comparison in the copy numbers of remained 19 transcription factor families between wild and cultivated green jujube.

Figure S4. Comparison of chromosome distribution of transcription factor superfamilies between the two cultivated and wild green jujube accessions. (a) Stacked line chart of copy number of transcription factor genes in 12 chromosomes of 8 HapGenomes. (b) Per‑chromosome comparison of copy number in 12 chromosomes between the two wild and cultivated green jujube accessions. Each dot represented one HapGenome and bars indicated median of copy number. (c) Distribution of transcription factor gene copies in Chr01 in the two wild and cultivated accessions.

Figure S5. KEGG pathway enrichment analysis of core, dispensable and private OGGs in the two wild and cultivated green jujube accessions. (a) Comparison of pathway enrichment between core, dispensable and private OGGs. (b) Comparison of pathway enrichment of core OGGs between the two wild and cultivated accessions. (c) Comparison of pathway enrichment of dispensable OGGs between the two wild and cultivated accessions. (d) Comparison of pathway enrichment of private OGGs between the two wild and cultivated accessions. Red to blue bar indicated the p value and the circle size indicated the number of OGGs.

Figure S6. Distribution of core, dispensable and private gene copies in 42 transcription factor superfamilies.

Figure S7. Syntenic relationships of OGGs across 8 HapGenomes in the two wild and cultivated green jujube accessions.

Figure S8. Comparison of Ka/Ks values of transcription factor superfamilies. (a) Comparison of Ka/Ks values between Zn finger, MYB, bHLH and AP2 family (copy number > 100). (b) Comparison of Ka/Ks values between 6 superfamilies for Homobox, NAC, B3, WRKY, MADS and FAR with copy number ranged from 50 to 100. (c) Comparison of Ka/Ks values between 13 superfamilies with copy number ranged from 10 to 50. (d) Comparison of Ka/Ks values between 15 superfamilies with copy number ranged < 10. Kruskal‐Wallis H test was conducted for significant analysis at p < 0.05, ns: no significant. The LFY, S1Fa‐like and STAT was found no collinearity relationships, and no Ka/Ks value was calculated in NZZ/SPL superfamily.

Figure S9. Phylogenetic relationship and OGG contraction and expansion results among 8 HapGenomes. The blue minus sign indicated the number of contracted OGGs, and the red plus sign indicated the number of expanded OGGs. The 4 Mya divergence time represented the lineage divergence of the wild and cultivated ancestors.

Figure S10. Expression profiles of core OGGs in leaf, branch and mature fruit of the two wild and cultivated green jujube accessions. (a) Proportions of core OGGs expressed in leaf, branch and mature fruit of the cultivated green jujube accession. (b) Proportions of OGGs expressed in leaf, branch and mature fruit of the wild green jujube accession. (c) Rader map of number of tissue‐specific OGGs between the two wild and cultivated green jujube accessions based on their TAU. (d) K‐means cluster analysis of core OGGs in leaf, branch and mature fruit of the two wild and cultivated green jujube accessions.

Figure S11. The significance of difference in different metabolite profiles from three wild and six cultivated green jujube accessions that were evaluated using Wilcoxon rank sum test.

TPG2-19-e70266-s001.docx (20.8MB, docx)

Table S1. The SRA lists of all tissues and accessions used in this study.

Table S2. Identified of transcription factor superfamilies through iTAK, PlantTFcat and PlantTFDB database.

Table S3. Summary of 42 transcription factor superfamilies and their subclasses.

Table S4. Composition of gene member in each OGGs.

Table S5. Distribution of all Ka/Ks values between transcription factor gene pairs among wild and cultivated green jujube.

Table S6. Expression profiles of all transcription factor genes in different tissues of cultivated green jujube.

Table S7. Expression profiles of all transcription factor genes in different tissues of wild green jujube.

Table S8. Haplotype‐resolved transcriptome atlas of all OGGs in different tissues of wild and cultivated green jujube.

Table S9. Haplotype‐resolved transcriptome atlas expression profiles of all OGGs in different fruits of wild and cultivated green jujube.

Table S10. Detailed results of leave‑one‑out cross‑validation (LOO‑CV).

Table S11. Differential expression analysis of all OGGs between wild and cultivated fruits and their SVM feature weights across the nine leave‑one‑out folds.

Table S12. Metabolome profiles in fruits from wild and cultivated green jujube.

TPG2-19-e70266-s002.xlsx (2.8MB, xlsx)

Data Availability Statement

The raw transcriptome data of different tissues used in this work have been accessed in the NCBI Sequence Read Archive under accession number PRJNA979989. The genomic sequences and annotation datasets of wild and cultivated green jujube genome are available at figshare (https://doi.org/10.6084/m9.figshare.23530068). The metabolomic data of all samples have been deposited to the MetaboLights database with the identifier MTBLS10827 (https://www.ebi.ac.uk/metabolights/editor/study/MTBLS10827). All other relevant data are included in the manuscript and its supplementary materials.


Articles from The Plant Genome are provided here courtesy of Wiley Periodicals LLC on behalf of Crop Science Society of America

RESOURCES