Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2024 Jun 3;121(24):e2319679121. doi: 10.1073/pnas.2319679121

Genome evolution of the ancient hexaploid Platanus × acerifolia (London planetree)

Xu Yan a,1, Gehui Shi a,1, Miao Sun a,1, Shengchen Shan b,1, Runzhou Chen a, Runhui Li a, Songlin Wu a, Zheng Zhou a, Yuhan Li a, Zhenhua Liu c, Yonghong Hu d, Zhongjian Liu e, Pamela S Soltis b,f,g, Jiaqi Zhang a,2, Douglas E Soltis b,f,g,h,2, Guogui Ning a,2, Manzhu Bao a,2
PMCID: PMC11181145  PMID: 38830106

Significance

We assembled a high-quality genome of the London planetree, Platanus × acerifolia, a hybrid between Platanus occidentalis and Platanus orientalis. In contrast to the chromosomal rearrangements often observed following WGD (Whole-genome duplication) in other extant angiosperm genomes, the ancient hexaploid P. × acerifolia genome exhibited surprising karyotype stability, with almost no chromosomal rearrangements. The Platanus genome shows widespread sub-/neo-functionalization without a dominant subgenome. WGD and sub-/neo-functionalization events contributed greatly to the evolutionary innovation of flowering-time genes in Platanus. Our results provide insights into the evolution of genome organization and improve our understanding of the evolution of flowering-time regulation in angiosperms.

Keywords: London planetree, comparative genomics, karyotypic stasis, sub-/neo-functionalization, flowering-time

Abstract

Whole-genome duplication (WGD; i.e., polyploidy) and chromosomal rearrangement (i.e., genome shuffling) significantly influence genome structure and organization. Many polyploids show extensive genome shuffling relative to their pre-WGD ancestors. No reference genome is currently available for Platanaceae (Proteales), one of the sister groups to the core eudicots. Moreover, Platanus × acerifolia (London planetree; Platanaceae) is a widely used street tree. Given the pivotal phylogenetic position of Platanus and its 2-y flowering transition, understanding its flowering-time regulatory mechanism has significant evolutionary implications; however, the impact of Platanus genome evolution on flowering-time genes remains unknown. Here, we assembled a high-quality, chromosome-level reference genome for P. × acerifolia using a phylogeny-based subgenome phasing method. Comparative genomic analyses revealed that P. × acerifolia (2n = 42) is an ancient hexaploid with three subgenomes resulting from two sequential WGD events; Platanus does not seem to share any WGD with other Proteales or with core eudicots. Each P. × acerifolia subgenome is highly similar in structure and content to the reconstructed pre-WGD ancestral eudicot genome without chromosomal rearrangements. The P. × acerifolia genome exhibits karyotypic stasis and gene sub-/neo-functionalization and lacks subgenome dominance. The copy number of flowering-time genes in P. × acerifolia has undergone an expansion compared to other noncore eudicots, mainly via the WGD events. Sub-/neo-functionalization of duplicated genes provided the genetic basis underlying the unique flowering-time regulation in P. × acerifolia. The P. × acerifolia reference genome will greatly expand understanding of the evolution of genome organization, genetic diversity, and flowering-time regulation in angiosperms.


Ranunculales, Proteales, Buxales, and Trochodendrales comprise a grade of eudicots that are subsequent sisters to the core eudicots, which comprise ~70% of all angiosperm species diversity. Despite representing only a small fraction of angiosperm species diversity, these four orders occupy pivotal phylogenetic positions in the evolution of flowering plants (1), exhibiting enormous diversity in growth habit, floral form, and flowering-time (2). Proteales comprise Sabiaceae, Nelumbonaceae, Proteaceae, and Platanaceae. Platanaceae (plane tree or sycamore family) consist of the single extant genus, Platanus, and also include extensive diversity in the fossil record. Genomic resources including genome assemblies are available for Proteaceae and Nelumbonaceae (36), as well as for the other clades that are sister to the core eudicots (Ranunculales, Buxales, and Trochodendrales) (79). However, no genome assembly is available for Platanaceae, hampering a comparative understanding of genome evolution in eudicots.

The ten extant species of Platanaceae are naturally distributed from southeastern Europe to Pakistan, Indo-China, eastern Canada to Guatemala (Fig. 1A) (1012). Platanus × acerifolia, a hybrid between Platanus orientalis (oriental plane) and Platanus occidentalis (American sycamore) (13), is commonly cultivated. Herein, we use P. × acerifolia rather than its preferred synonym P. × hispanica (11, 12) given its wide usage in horticulture. Platanus × acerifolia is cultivated worldwide and has gained popularity in the urban environment owing to its attractive canopy, fast growth, and high tolerance to infertile soil and air pollution (Fig. 1 A and B, SI Appendix, Fig. S1 and Supplementary Note 1, and Datasets S1 and S2) (14). All members of Platanaceae have a chromosome number of 2n = 42 (15) and are considered to be ancient polyploids (16, 17). Oginuma and Tobe (18) proposed that Platanus species are likely to be ancient hexaploids with a basic chromosome number of x = 7. However, the hypothesized ancient hexaploidy of Platanus has not been confirmed with genetic or genomic data.

Fig. 1.

Fig. 1.

The biological profile and genome of Platanus × acerifolia. (A) The natural distribution of extant Platanus species (green) with an ornamental hybrid P. × acerifolia highlighted (red). The red dots represent occurrence data assembled from GBIF (access on 06 August 2022) and natural distribution range (green) is based on POWO (Plants of the World Online; https://powo.science.kew.org/taxon/urn:lsid:ipni.org:names:32120-1). (B) The outstanding tall trees of P. × acerifolia cultivated among small trees (Osmanthus fragrans, Ligustrum lucidum, etc) on Purple Mountain in Nanjing, China. The whole P. × acerifolia community is designed by horticulturalists forming a palace-like corridor with fall foliage coloration (yellow line of trees in the image) Photographed by Guoxiao Tu. (C) Phylogeny-based genomic phasing of P. × acerifolia. The red, blue, and dark yellow branches correspond to subgenomes A, B, and C, respectively. (D) Genomic features of the representative pseudohaploid genome of P. × acerifolia. The circle layers from outer to inner represent: Each wedge represents a chromosome, the length of wedge indicates the chromosome size (Mb), and each group of 3 wedges of the same color represents 3 ancestral homeologs. (a), repeat density (b), the density of long terminal repeat (LTR) elements (c), the density of long interspersed nuclear elements (LINEs) (d), the density of genes (e), the densities of single nucleotide polymorphisms (SNPs) (navy) and insertions-deletions (InDels) (purple) in P. occidentalis (f), the densities of SNPs (navy) and InDels (purple) in P. orientalis (g), and the syntenic relationships (h).

Polyploidy and subsequent chromosome rearrangements (also known as genome shuffling) can lead to reproductive barriers and are thought to be an important force behind the evolutionary success of plants (19). Eudicots are hypothesized to be derived from an ancestor that contained seven chromosomes and then experienced extensive genome shuffling, reflected in the present-day karyotypes of some members (20). For example, compared to a hypothetical reconstructed ancestral eudicot genome (2123), Aquilegia coerulea (n = 7; Ranunculales) experienced four fissions and 11 fusions after one round of WGD (Whole-genome duplication) (21); Nelumbo nucifera (n = 8; Proteales) experienced 14 fissions and 20 fusions after one round of WGD (21); Vitis vinifera (n = 19; Vitales) experienced two fissions and four fusions after the gamma paleopolyploidization (22); and Arabidopsis thaliana (n = 5; Brassicales) experienced 83 fissions and four fusion events after the gamma paleopolyploidization and two additional lineage-specific WGDs (23). Notably, the karyotypes of Buxus sinica and Buxus austro-yunnanensis (both n = 14; Buxales) showed high similarity to the hypothetical ancestral eudicot karyotype (AEK) after one round of WGD, despite the large number of noncolinear regions between homeologous chromosomes (9, 21).

In addition to its unique phylogenetic position and proposed polyploidization history, Platanus exhibits key evolutionary innovations in the transition from vegetative growth to flowering. Specifically, P. × acerifolia undergoes a 2-y flowering transition, forming flower subpetiolar buds and becoming dormant in the first year, and then flowering in April or May of the following year (24). In contrast, the flowering schedules of annual model plants such as A. thaliana and Oryza sativa are within a single growing season (25, 26). “Flowering-time genes” here are defined as those Platanus genes whose overexpression alters flowering-time in any Arabidopsis or tobacco (Nicotiana tobacum) accession. In P. × acerifolia, eight early flowering homologs (SQUAMOSA PROMOTER BINDING PROTEIN-LIKE3 [SPL3], FLOWERING LOCUS D [FD], FLOWERING LOCUS T [FT], MOTHER of FT and TFL1 [MFT], SUPPRESSOR OF OVEREXPRESSION OF CONSTANS 1 [SOC1], LEAFY [LFY], AGAMOUS-LIKE 6 [AGL6], and FRUITFULL [FUL]) and three late flowering homologs (TERMINAL FLOWER 1 [TFL1], BROTHER OF FT AND TFL1 [BFT], and SHORT VEGETATIVE PHASE [SVP]) have been identified (2732). However, the genomic drivers for the innovation of these flowering-time genes in P. × acerifolia remain unknown.

Given its widespread cultivation, value as a genomic resource, and evolutionary significance, P. × acerifolia was selected for assembly and annotation of a high-quality, chromosome-level reference genome. A comparison of the Platanus genome with those from other eudicots will enable a better understanding of the evolution of ploidy and subsequent changes in genomic structure, as well as the genomic drivers for evolutionary innovation in flowering-time regulation in P. × acerifolia.

Results

The Genome of the Ancient Hexaploid P. × acerifolia.

The genome size (1.37 Gb; SI Appendix, Fig. S2 and Supplementary Note 2) and chromosome number (2n = 42; SI Appendix, Fig. S3) of P. × acerifolia are similar to those of other Platanus species (13). We assembled a high-quality, haplotype-resolved genome for P. × acerifolia with a high BUSCO (Benchmarking Universal Single-Copy Orthologs) completeness (99.6%; SI Appendix, Table S1), a high NGS (Next-generation sequencing) read alignment rate (99.89%; SI Appendix, Table S2), a high Hi-C-anchored rate (95.54%; SI Appendix, Table S3), as well as the longest contig N50 among all sequenced genome in the grade of subsequent sisters to the core eudicots (24.15 Mb; SI Appendix, Tables S4 and S5). The longest 42 scaffolds corresponded to the known chromosome number (2n = 42) of Platanus species (SI Appendix, Fig. S4 and Table S6). Therefore, we assume each of the 42 scaffolds represents one chromosome. Notably, repetitive elements accounted for 69.89% of the P. × acerifolia genome (SI Appendix, Fig. S5 and Table S7 and Dataset S3). We annotated 94,518 gene structures, 92.25% of which had a functional annotation (SI Appendix, Table S8).

Oginuma and Tobe (18) proposed that all Platanus species are hexaploid with a basic chromosome number of x = 7. We conducted colinearity analysis based on all 42 chromosomes (i.e., the 42 longest scaffolds) of P. × acerifolia to identify homologous and homeologous chromosomes. We observed that: 1) Each chromosome was colinear with the other five chromosome copies (SI Appendix, Fig. S6). For each chromosome, there was always one other chromosome with higher similarity (red dots in SI Appendix, Fig. S6) and identical colinearity, indicating that the assembled genome consists of the two haploid genomes contributed by the 21 pairs of homologous chromosomes (2n = 42). 2) Each haploid genome could be further subdivided into seven sets of three homeologous chromosomes based on structural features (e.g., inversions and/or translocations), thus representing the three subgenomes within each haploid genome (n = 3x = 21; SI Appendix, Fig. S6).

To differentiate among the three subgenomes, we performed a phylogenetic analysis of 1,675 orthogroups (OGs) of low-copy-number (LCN) protein-coding genes (LCN OGs) from the seven chromosome groups (SI Appendix, Fig. S6). Within each group, the chromosomes were clustered into three subgenomes based on phylogenetic relationships (Fig. 1C and SI Appendix, Table S9 and Supplementary Note 4). The topology of the homeologous chromosomes became more stable as the number of LCN OGs increased; for the majority of the homeologous chromosome groups, the improvement of the phylogeny stability gradually reached a plateau when 80% or more LCN OGs were used to build the phylogeny (SI Appendix, Fig. S7). These results indicated that sufficient genes were used to build the subgenome phasing phylogeny, and the topology remains unchanged as gene numbers increased. Importantly, our phasing method also successfully phased allotetraploid Gossypium hirsutum and the allohexaploid Triticum aestivum subgenomes as test cases (SI Appendix, Fig. S8 and Supplementary Note 4), suggesting the effectiveness of our method. However, using the standard K-mer-based approaches (such as SubPhaser), we could not distinguish the subgenomes of P. × acerifolia (SI Appendix, Fig. S9; also see Discussion).

Once the three subgenomes within Platanus had been distinguished, the longer chromosome from each homologous chromosome pair (derived from the two parental gametes) was selected to constitute the pseudohaploid genome, which retained most of the genetic information (Fig. 1C). This resulted in a total of 21 chromosomes across the three subgenomes in the pseudohaploid genome of P. × acerifolia (SI Appendix, Table S10). Although the pseudohaploid genome might be a mosaic of parental chromosomes rather than the true biological haploid genome, there was a high degree of colinearity between homologous chromosomes of the two haplotypes, suggesting that the pseudohaploid genome well represented the haplotype of the P. × acerifolia genome. Indeed, approximately 93% of the transcripts were mapped onto the 21 representative chromosomes, which was similar to the transcriptome alignment rate (ca. 95%) when all 42 chromosomes were used (SI Appendix, Fig. S10 and Dataset S4). Accordingly, we considered the 21 chromosomes (1n = 3x = 21), along with the remaining unanchored contigs, as representative of the P. × acerifolia pseudohaploid genome, and these were used for subsequent analyses (Fig. 1D and SI Appendix, Table S11).

Platanus × acerifolia subgenome A (589.12 Mb) consists of seven chromosomes with lengths ranging from 71.02 to 110.33 Mb; subgenome B (593.20 Mb) includes seven chromosomes with lengths ranging from 69.43 to 99.59 Mb; subgenome C (555.54 Mb) comprises seven chromosomes with lengths ranging between 61.01 and 92.65 Mb (SI Appendix, Table S10). We annotated 13,969, 13,882, and 13,179 protein-coding genes for subgenomes A, B, and C, respectively (SI Appendix, Table S10).

Together with genomic data from 106 representative angiosperm species, the dated species tree shows that the phylogenetic placement and divergence of Platanus (Platanaceae) are consistent with the current understanding of the angiosperm tree of life (SI Appendix, Figs. S11–S18, Table S12, and Supplementary Note 5 and Dataset S5) (33). To gain insight into gene loss and gain during the hexaploidization of Platanus, we compared the expansion and contraction of orthologous gene families in P. × acerifolia and 16 other phylogenetically representative gymnosperm and angiosperm species (SI Appendix, Figs. S19–S21 and Supplementary Note 6). We detected 5,864 gene families that experienced expansion and 3,155 that underwent contraction in P. × acerifolia compared with the most recent common ancestor of Platanaceae and Proteaceae (SI Appendix, Fig. S19). The genes that exhibited significant expansion in P. × acerifolia were associated with hormone regulation, plant–pathogen interaction, as well as flower and leaf development, which may contribute to the rapid growth and strong disease resistance in Platanaceae (SI Appendix, Fig. S20). In addition, the trichome differentiation-related genes TCP4 and TCP15 (34) also underwent an expansion and were found to be highly expressed in branch, subpetiolar buds, and male/female inflorescences, which may explain the dense trichomes of P. × acerifolia (SI Appendix, Fig. S21).

Karyotypic Stasis after Independent WGD Events.

WGD events within Proteales were inferred by genome colinearity analysis and the distribution of KS values for duplicated genes (Fig. 2 A and B and SI Appendix, Figs. S22–S25). One significant KS peak was observed for Platanaceae at approximately 0.42, corresponding to a polyploidization event approximately 49 Mya (Fig. 2A). Similarly, one independent KS peak was observed for each of N. nucifera, Macadamia integrifolia, and Telopea speciosissima, indicating that the families within Proteales (Nelumbonaceae, Platanaceae, and Proteaceae) might all have undergone independent WGD events. Moreover, the estimated date of the WGD event and the age of the crown group of Platanaceae (approximately 56 Mya; see SI Appendix, Supplementary Note 7) were very close.

Fig. 2.

Fig. 2.

P. × acerifolia has retained the karyotype of the ancestral eudicot after two sequential Platanaceae-specific WGD events. (A) WGD analysis based on KS distribution. The KS distribution of paralogs (solid lines) in colinear regions in Proteales and orthologs (dashed lines) between P. × acerifolia (red) and T. speciosissima (orange), M. integrifolia (green), N. nucifera (blue), or V. vinifera (purple). (B) Macrosyntenic alignments of Proteales. The seven colors indicate the syntenic blocks of seven sets of chromosomes in P. × acerifolia. (C) The KS distribution of orthologs in colinear regions between pairs of the three subgenomes. (D) Sister-to-core eudicot karyotype evolution was predicted based on the reconstructed AEK model. The seven protochromosomes comprising the hypothetical AEK are illustrated using distinct colors. Circles label putative duplication events and stars highlight the two rounds of WGD where the third genome was donated to the original tetraploid (green star) by an extinct lineage, yielding the hexaploid (purple star).

The results of our colinearity analysis further showed that there was a 3:3 syntenic depth ratio between P. × acerifolia and V. vinifera; in contrast, other Proteales families (N. nucifera, M. integrifolia, and T. speciosissima) each exhibited a syntenic depth of 2:3 with P. × acerifolia (Fig. 2B and SI Appendix, Figs. S22–S25). Both colinearity and phylogenetic analyses of the duplicated genes indicated that WGDs occurred independently in each family of the Proteales after they diverged from each other (SI Appendix, Fig. S26 and Supplementary Note 7) and that Platanaceae independently underwent two rounds of genome duplications that were not shared by other Proteales families (Fig. 2C and SI Appendix, Fig. S26).

To verify the occurrence of the two sequential ancient WGDs in P. × acerifolia, the KS values between paralogs of any two subgenomes were calculated. The peak of the KS distribution between paralogs of subgenomes A and C was almost identical to that between paralogs of subgenomes B and C (KS = 0.421; Fig. 2C). The peak of the KS distribution between paralogs of subgenomes A and B was different from those of subgenomes A and C and subgenomes B and C (KS = 0.419; Fig. 2C). This finding indicated that within a short time, the diploid ancestors of subgenomes A and B formed a tetraploid, followed quickly by an ancient hexaploidy formation that involved the diploid ancestor of subgenome C; this explained the above presence of only one KS peak in hexaploid P. × acerifolia.

Murat et al. (35) inferred that the hypothetical AEK comprised seven protochromosomes (i.e., the ancestral preduplication chromosomes). We used the hypothetical AEK model consisting of 23,374 protogenes (the ancestral preduplication genes) developed by Wang et al. (21) to examine possible karyotypic changes in P. × acerifolia and other eudicot species (Fig. 2D and SI Appendix, Supplementary Note 8). Our analysis of karyotype evolution revealed that each subgenome in P. × acerifolia has maintained strong colinearity with the reconstructed AEK (Fig. 2D). The colinearity blocks between P. × acerifolia and AEK gene pairs accounted for 60.01% (14,041) of the progenitors in the AEK model (Dataset S6). Colinearity analysis using a 1-Mb sliding window showed that 90.51% (1,583) of the P. × acerifolia genome shared colinearity blocks with the AEK (Dataset S6). We identified 413 colinear blocks (47,306 syntenic gene pairs) between P. × acerifolia and AEK (Dataset S7). In addition, genome shuffling including 39 translocated inversions (3,250 translocated inversion gene pairs), 32 translocations (2,761 gene pairs), and 103 inversions (9,025 gene pairs) were identified. Notably, the three subgenomes of P. × acerifolia showed almost no exchange between nonhomeologous chromosomes relative to their eudicot ancestor (Fig. 2D and SI Appendix, Fig. S32). We found that P. × acerifolia had only seven potential genome shuffles between nonhomeologous chromosomes (four rearrangements between LG05 and LG03, two rearrangements between LG07 and LG03, and one rearrangement between LG07 and LG01). Only 0.27% (132) of syntenic gene pairs in P. × acerifolia are involved in the potential genome shuffles between nonhomeologous chromosomes (Dataset S7).

These findings suggested that almost no genome shuffling occurred between nonhomeologous chromosomes in the P. × acerifolia genome after the formation of the ancient hexaploid. Because the karyotype of Platanus was nearly identical to the inferred AEK, the most recent ancestor of Proteales (the ancestor of Platanaceae) likely possessed a very similar karyotype to the putative AEK (Fig. 2D). Compared with other eudicot species in Ranunculales, Proteales, Trochodendrales, and Buxales (Aquilegia coerulea, Coptis chinensis, Papaver somniferum, N. nucifera, M. integrifolia, T. speciosissima, Tetracentron sinense, and B. sinica) as well as V. vinifera, sister to other rosids, Platanus retained a karyotype that is most similar to the reconstructed AEK (Fig. 2D and SI Appendix, Figs. S27–S37).

No Significant Dominance was Detected among the Three Subgenomes.

Subgenome dominance in P. × acerifolia was analyzed by comparing gene content and expression among its three subgenomes (SI Appendix, Supplementary Note 9). The number of subgenome-specific differentially expressed genes (DEGs) was approximately equal among the three subgenomes (Dataset S8), and there was no significant difference in the expression of homeologous genes among subgenome colinearity blocks (Fig. 3A). Additionally, there were 4,096 Pfam domains in the genes of the P. × acerifolia genome, 96.19% (3,940) of which did not differ significantly among the three subgenomes (SI Appendix, Fig. S38A). This result suggested that there was no apparent subgenome-biased loss of genes. When A. coerulea (Ranunculaceae, Ranunculales), N. nucifera (Nelumbonaceae, Proteales), V. vinifera (Vitaceae, Vitales), and Populus trichocarpa (Salicaceae, Malpighiales) were used as outgroups, subgenomes A, B, and C showed an average of 10,939, 10,728.5, and 10,872 homeologous genes in colinear blocks, respectively (SI Appendix, Fig. S38B). The coefficient of variation (CV) of the number of homologs in the colinearity blocks among three subgenomes (0.81%) was significantly lower than CV of subgenome sizes among the three subgenomes (2.91%), suggesting that there was no significant difference in the number of homeologs among subgenomes. To further investigate subgenome dominance in P. × acerifolia, we compared the density of the six most abundant types of transposable elements (TEs) [hAT-Tag1, MuLE-MuDR, L1-LINE, RTE-BovB-LINE, Copia LTR (long terminal repeat), and Gypsy LTR] among the three subgenomes (SI Appendix, Fig. S38C). We found that the TE density did not differ significantly among subgenomes A, B, and C (SI Appendix, Fig. S38C).

Fig. 3.

Fig. 3.

The P. × acerifolia genome lacks subgenome dominance and exhibits unbiased fractionation and gene sub-/neo-functionalization. (A) Homologous gene expression percentage in syntenic blocks among the three subgenomes. The nested blue lines show 90%, 95%, and 99% CI. (B) Heterozygosity frequency among the subgenomes (per 500-k window). The red, blue, and green dotted lines delineate the locations of seven chromosomes of subgenomes A, B, and C, respectively. (C) FractBias plots comparing B. sinica (target genome) with P. × acerifolia (query genome). The figure depicts pair-wise comparisons between each B. sinica chromosome (PGA1 to PGA14) and the corresponding syntenic regions of the three subgenomes of each P. × acerifolia chromosome (LG01 to LG07). The x-axis corresponds to iterated sliding windows of 50 genes on the target chromosome. The y-axis indicates gene retention percentages at syntenic locations on query chromosomes. Each chromosomal item in the figure can be regenerated at https://genomevolution.org/r/1qojf. (D) A heatmap of duplicated genes with subfunctionalization and/or neo-functionalization events. Each row represents genome duplication-derived copies of a specific gene in the three subgenomes. Different columns represent different subpetiolar bud developmental stages (shown as Area 4 of Fig. 4A). Gene expression (FPKM) was z-normalized.

Differences in heterozygosity between pairs of subgenomes can also serve as a measure of subgenome dominance. Recessive subgenomes are expected to experience lowered purifying selection pressure compared with dominant ones. We estimated the per-500-kb window distribution of nucleotide diversity (π) between P. × acerifolia subgenomes and found that the π curves for the three P. × acerifolia subgenomes were entangled, and none of the subgenomes displayed a consistent lower/higher nucleotide diversity compared to the other subgenomes (Fig. 3B). Based on the Wilcoxon rank-sum test, there were no significant differences in heterozygosity between P. × acerifolia subgenomes A and B (0.0092 vs. 0.0085, P = 0.858), subgenomes A and C (0.0092 vs. 0.0091, P = 0.232), or subgenomes B and C (0.0085 vs. 0.0091, P = 0.291) (Dataset S9). The analysis of heterozygosity is consistent with there being no subgenome dominance among the subgenomes of P. × acerifolia.

In addition, based on the Wilcoxon nonparametric test (SI Appendix, Figs. S39 and S40), no significant differences in the distribution of Ka/Ks were observed among the three subgenomes. However, we found a significant difference in Ka between subgenomes A and B, and in Ks between subgenomes A and B and between subgenomes B and C (SI Appendix, Figs. S41 and S42). The similar Ka/Ks distribution among the three subgenomes, together with the nucleotide diversity results (Fig. 3B), indicated that selection pressure seems most likely unbiased among the subgenomes of P. × acerifolia.

In summary, based on the results from analyses of gene content and expression, the presence of TEs, and selection pressure among the three subgenomes, our data suggest that there is no subgenome dominance in P. × acerifolia.

Fractionation and Sub-/Neo-Functionalization in P. × acerifolia.

A comparative analysis of the composition of homeologous genes among the three subgenomes (SI Appendix, Supplementary Note 10) showed that over one-third of the genes (15,163/41,030) were shared among all three subgenomes (SI Appendix, Figs. S43 and S44); these shared genes were enriched in functions related to plant growth, stress responses, and hormone and signal transduction (SI Appendix, Fig. S45). Furthermore, approximately one-third of the genes (13,765/41,030) in the pseudohaploid subgenomes did not have homeologs in the other two subgenomes (SI Appendix, Fig. S44), suggesting that a large number of genes lost their redundant copies following the WGD events (i.e., fractionation). Genes specific to subgenome A (4,800) are enriched in RNA synthesis, metabolism, and modification (SI Appendix, Fig. S46); genes specific to subgenome B (4,823) were enriched in cell wall formation and pollen recognition (SI Appendix, Fig. S47); genes specific to subgenome C (4,142) were enriched in pollen tube guidance, leaf abscission, regulation of ncRNAs, and modification of tRNAs (SI Appendix, Fig. S48). These results suggested that all three subgenomes have undergone extensive fractionation.

SynMap’s FractBias tool was used for quantitative comparisons of gene loss across the P. × acerifolia subgenomes using B. sinica (Buxaceae) as a reference (https://genomevolution.org/r/1qojf). Overall, no clear fractionation bias pattern was observed with FractBias among the three subgenomes of P. × acerifolia (relative to B. sinica), as shown in SI Appendix, Fig. S56. Although a small portion of the FractBias curve showed that subgenome B had a lower gene retention rate than subgenomes A and C of P. × acerifolia on chromosome LG07 (relative to B. sinica chromosome PGA 12), the curves of the three subgenomes in P. × acerifolia were tangled in most genomic regions (Fig. 3C). These results, together with the fact that a similar number of homeologs was retained among the three subgenomes following fractionation (SI Appendix, Fig. S38B), indicated no fractionation bias among the three subgenomes in P. × acerifolia.

We also analyzed the colinearity between P. × acerifolia and A. coerulea, N. nucifera, V. vinifera, and P. trichocarpa to identify the locations of homeologs retained after fractionation (SI Appendix, Supplementary Note 10). The results showed that the distribution of retained homeologs after fractionation was not significantly biased among the three subgenomes of P. × acerifolia (SI Appendix, Fig. S49). A microcolinearity analysis of P. × acerifolia showed that the three subgenomes lost genes at different loci, but retained a similar number of overall genes, leading to unbiased fractionation (SI Appendix, Fig. S50).

To detect whether genes that were retained in triplicate (without fractionation) may have undergone sub-/neo-functionalization (36, 37), we examined the expression patterns of genes in P. × acerifolia using the transcriptome data of flowers at different development stages (SI Appendix, Fig. S51A and Supplementary Note 11). Among the 2,099 gene sets retained in triplicate (6,297 genes), 1,220 (58.12%) exhibited significant shifts in expression patterns among gene copies (Fig. 3D and SI Appendix, Fig. S51B), implying these genes are subjected to sub-/neo-functionalization. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis showed that the gene copies displaying sub-/neo-functionalization were significantly enriched (adjusted P ≤ 0.05) in many key biological processes, including plant hormone signal transduction- and circadian rhythm-related pathways (SI Appendix, Fig. S52). These results suggested that the replicated genes in the P. × acerifolia genome underwent extensive sub-/neo-functionalization and may have contributed to shaping the phenotype of P. × acerifolia.

WGD and Sub-/Neo-Functionalization Drove Evolutionary Innovation in Flowering-Time Genes.

Platanus species show innovation in the regulation of flowering-time, characterized by a 2-y flowering transition (30). However, the major genomic drivers underlying this innovation remain unknown. Thus, we analyzed the phylogeny, colinearity, sequence differences, and expression trends of 12 P. × acerifolia flowering-time gene duplicates whose functions have been confirmed (SI Appendix, Figs. S53–S73 and Supplementary Note 12 and Dataset S10). The copy numbers of SOC1 (15) and FT (10) were markedly expanded in P. × acerifolia compared with that in A. coerulea (4 and 3, respectively), N. nucifera (2 and 9, respectively), T. sinense (2 and 3, respectively), and B. sinica (2 and 1, respectively) (Fig. 4A). Phylogenetic and colinearity analyses of the FT genes and their genomic neighborhoods indicated that two ancient FT duplicates existed before the WGD events in Platanaceae, resulting in two paralogous clades (clade 1 and clade 2; SI Appendix, Figs. S59 and S61). In addition, owing to the WGD events in Platanaceae, P. × acerifolia had paralogs from each of the three subgenomes in both clade 1 and clade 2 (SI Appendix, Fig. S61).

Fig. 4.

Fig. 4.

WGD and sub-/neo-functionalization events of flowering-time genes in P. × acerifolia. (A) Phylogenetic tree and distribution of flowering-time genes across the representative flowering plants. Circles label putative duplication events and stars highlight the two rounds of WGD where the third genome was donated by an extinct lineage to the original tetraploid (green star) to form the hexaploid (purple star) in the phylogenetic tree. The background colors of the heatmap denote the extent of copy number variation in each species. (B) A phylogenetic tree of SUPPRESSOR OF OVEREXPRESSION OF CONSTANS 1 (SOC1), which was extracted from the phylogenetic tree of the MADS-box gene family (SI Appendix, Fig. S54). (C) Colinearity analysis of the SOC1 genes and their genomic neighborhood. (D) Area 1: The colored boxes represent the seasons (blue, winter; red, spring; yellow, summer; green, autumn); the division of seasons is based on data provided by the Wuhan Meteorological Bureau Administration. Area 2: Eight developmental stages divided according to histological sections of inflorescence development. Stage 1 (S1), subpetiolar bud differentiation phase, which is the critical stage of flowering transition, May; Stage 2 (S2), inflorescence primordium differentiation phase, late May/early June; Stage 3 (S3), single flower, June; Stage 4 (S4), floral organ differentiation phase, late June/late July; Stage 5 (S5), the first phase of floral organ development, August; Stage 6 (S6), endodormant phase, September/November; Stage 7 (S7), maturation phase of stamen and carpel development, indicating that flower bud differentiation is complete, December until the following early April; and Stage 8 (S8), flowering phase, April. Area 3: Heatmap of the expression of early-flowering (red) and late-flowering (blue) gene duplicates. Area 4: The developmental status of the subpetiolar bud. Area 5: Images of the canopy showing the developmental status of P. × acerifolia. (E) Phenotypes of PaBBX28/29-1-overexpressing Arabidopsis plants grown under short-day (SD) conditions. (Scale bar, 5 cm.) (F) The phenotypes of PaBBX28/29-2-expressing Arabidopsis plants grown under SD conditions. (Scale bar, 5 cm.)

Phylogenetic results of the SOC1 gene showed that the cluster of P. × acerifolia SOC1 duplicates further clustered into three clades owing to the occurrence of WGDs in Platanaceae (Fig. 4B and SI Appendix, Fig. S53). All P. × acerifolia SOC1 duplicates originated from chromosome LG07 in each of three subgenomes, including four duplicates in subgenome A (clade 2), nine in subgenome B (one in clade 3 and eight in clade 1), and two in subgenome C (clade 3; Fig. 4B). Colinearity analysis of SOC1 and its genomic regions showed that 10 SOC1 duplicates (PaA7G298710, PaA7G298740, PaA7G298750, PaA7G298760, PaB7G381950, PaB7G381980, PaB7G382000, PaB7G382030, PaC7G398240, and PaC7G398260) were derived from WGDs in Platanaceae, while five SOC1 duplicates (PaB7G381970, PaB7G382020, PaB7G530940, PaB7G530950, and PaB7G530960) were derived from post-WGD tandem duplications (TDs) in subgenome B (Fig. 4C). We also found that SOC1 duplicates formed a 5 to 6-Mb gene cluster with AGL6 and BFT (Fig. 4C), both of which also play a role in the regulation of flowering-time. The expansion of AGL6 and BFT was also driven by WGDs in Platanaceae (SI Appendix, Figs. S53–S55 and S59). Of the 27 duplicates of the other eight validated flowering-time genes (excluding the already-mentioned SOC1 and FT as well as LFY, with only one copy), 22 (81.5%) were derived from WGD events, two (7.4%) from TDs, and three (11.1%) via transpositions or other mechanisms (SI Appendix, Fig. S53).

Subfunctionalization and neo-functionalization are key facets of the long-term evolution of WGD-mediated gene expansion and play a role in trait innovation (36). We analyzed the role of sub-/ neo-functionalization in the evolutionary innovation of flowering-time genes based on the expression patterns of 12 validated flowering- time gene duplicates (Fig. 4D). Notably, the expression patterns of some genes differ from their homeologs, such as PaBBX28/29. The BBX28/29 genes PaA2G260070 (PaBBX28/29-2), PaB2G067580 (PaBBX28/29-1), and PaC2G369370 (PaBBX28/29-3) were continuously expressed during floral organ differentiation, while only PaBBX28/29-1 was also expressed during the subpetiolar bud differentiation phase (Stage1; Fig. 4D and SI Appendix, Fig. S72). The significantly different expression patterns of PaBBX28/29 duplicates indicated the possibility of sub-/neo-functionalization (Fig. 4D and SI Appendix, Fig. S73). To verify whether there were functional differences among PaBBX28/29 copies, we independently transferred two of the three copies (PaBBX28/29-1 and PaBBX28/29-2) into A. thaliana (see cloning and functional analysis of PaBBX28/29 in the Materials and Methods and SI Appendix, Figs. S74–S79). Both the 35S::PaBBX28/29-1 and 35S::PaBBX28/29-2 transgenic lines exhibited a strong late-flowering phenotype under long-day (LD) conditions, suggesting that no significant functional difference existed between these two gene copies under the LD condition (SI Appendix, Fig. S76). However, significant functional divergence was observed between the 35S::PaBBX28/29-1 and 35S::PaBBX28/29-2 transgenic lines under short-day conditions; specifically, the 35S::PaBBX28/29-2 transgenic line exhibited a late-flowering phenotype, which was not observed in the 35S::PaBBX28/29-1 transgenic A. thaliana (Fig. 4 E and F). Subsequently, a dual-luciferase reporter assay showed that both PaBBX28/29-1 and PaBBX28/29-2 inhibited the transcription of AtFT, while only PaBBX28/29-2 suppressed the transcription of PaFT (SI Appendix, Fig. S78 A and B). The amino acid sequences of the two gene copies exhibited a 77% similarity value (SI Appendix, Supplementary Note 13). Results based on phylogeny and multiple sequence comparisons showed that the structure and phylogeny of the alleles were very similar between the parents (SI Appendix, Fig. S74). These findings indicated that the duplicated genes of P. × acerifolia with shifted expression patterns might have undergone sub-/neo-functionalization after genome duplication; however, we could not exclude the possibility that the functional divergence between the two copies was inherited from their now-extinct ancestral counterparts.

Discussion

We assembled and annotated high-quality genome for Platanaceae, filling an important phylogenetic gap in the lineages that are sisters to core eudicots. Although the pseudohaplotype subgenome assembled here may comprise a mosaic of chromosomes originating from different gametes (and thus the parental species P. occidentalis and P. orientalis), we confirmed that this possible mosaicism did not affect the analysis of the triplicated subgenomes based on the results of genome colinearity and Hi-C interaction heatmaps analysis (SI Appendix, Fig. S4). The hexaploid P. × acerifolia genome provides a valuable resource for studying genome evolution not only in Proteales but also in eudicots in general. Genome assemblies of P. occidentalis and P. orientalis would enable the separation of the homologous chromosomes (contributed by the gametes) in P. × acerifolia into true subgenomes (rather than the pseudosubgenomes identified in the current work). Classifying the haploid genomes of P. × acerifolia would reveal whether there 1) has been any crossing-over between the two haploid subgenomes; 2) is any subgenome dominance at the level of the parental species-specific subgenomes; and 3) has been divergence between the two haploid subgenomes subsequent to the formation of P. × acerifolia. Answers to these questions will promote a more comprehensive understanding of polyploid genome evolution in P. × acerifolia.

The sequencing and assembly of haplotype-resolved polyploid plant genomes are no longer formidable tasks (38, 39). However, phasing subgenomes in ancient polyploids of unknown progenitors remains extremely challenging (40). Current methods for assigning subgenomes without ancestry-related information rely on genomic features such as K-mer distribution, GC content, and the phylogeny of centromeric repeats (41). In our initial pipeline, both phylogenetic and K-mer-based approaches were used for subgenome phasing. The phylogeny-based phasing method successfully phased the genome of P. × acerifolia into three putative subgenomes (Fig. 1C). However, the K-mer-based approach could not distinguish the subgenomes of P. × acerifolia. We used SubPhaser v.1.2.6 (42), a widely used algorithm for subgenome phasing, with a comprehensive combination (including default values: K-mer length = 13, minimum K-mer frequency = 200) of 10 K-mer values and eight minimum frequency thresholds for each K-mer (a total of 80 parameter combinations). However, none of them could phase the P. × acerifolia genome into three subgenomes (SI Appendix, Fig. S9; data and logfiles are available at https://doi.org/10.5061/dryad.j6q573nm9). We suspect that one cause could be likely due to the evolutionary similarity of the ancestral parents of Platanus. Indeed, Jia et al. (42) indicated that SubPhaser may not be able to successfully phase subgenomes in autopolyploids or allopolyploids derived from hybridization events between closely related species. In addition, SubPhaser used subgenome-specific repetitive DNA sequences for subgenome assignment (42); in contrast, our phylogenetic phasing approach used the sequences from LCN genes. Phasing subgenomes by using divergent genomic regions (i.e., repetitive DNAs vs. genes) between the two approaches may also explain the observed results. Our approach expands the subgenome phasing methods available for allopolyploid plants of unknown progenitors and provides a valuable supplementary method for future studies of polyploid genome evolution; it should be used with caution considering some of the assumptions we made here (e.g., no large amount of gene exchange between subgenomes).

The results of our KS distribution, colinearity, and phylogenetic analyses suggested that independent genome duplication events occurred in Platanaceae with no shared WGD event being detected among the ancestors of four Proteales families (i.e., Sabiaceae, Nelumbonaceae, Platanaceae, and Proteaceae). The WGD events in Platanaceae are also independent of those in Vitaceae (SI Appendix, Fig. S26). Moreover, previous studies using the B. sinica (Buxales) and T. sinense (Trochodendrales) genomes had shown that there were no shared WGD events among Ranunculales, Nelumbonaceae, Buxales, Trochodendrales, and Vitales (9). Our results further supported the hypothesis that the gamma polyploidy event only involved the stem lineage of core eudicot ancestors (9). The timing of many WGD events often coincides with major periods of the Earth's climatic and/or geologic changes or periods of mass extinction (43). The genome duplication events in Platanaceae occurred approximately 49 Mya, which is close to the time of the Paleocene-Eocene Thermal Maximum (56 Mya) (44). Therefore, the genome duplication events in Platanus might have occurred as a stress response to major paleoclimatic changes that took place at that time (44).

Comparing the karyotypes of extant species with the hypothetical AEK provides a perspective on eudicot chromosomal evolution (9, 19, 45). A growing number of studies of eudicot karyotypes based on whole-genome sequencing have shown that extensive nonhomeologous chromosome rearrangements occurred in flowering plants following WGD events (9, 21, 45, 46). Among the eudicots, the extant chromosomes of both P. × acerifolia and B. sinica are very similar to the inferred hypothetical AEK (9). However, the colinearity blocks shared between P. × acerifolia and AEK contained 60.01% (14,041) of the genes of the reconstructed AEK model (Dataset S6), whereas those shared between B. sinica and AEK contained only 37.49% (8,763) of the genes of the reconstructed AEK model (Dataset S11). Of these genes, 7,910 AEK model progenitors were colinear with the chromosomes of both P. × acerifolia and B. sinica, and 6,131 AEK progenitors were only colinear with P. × acerifolia (Dataset S6), while 853 AEK progenitors were only colinear with B. sinica (Dataset S11). Using sliding windows of 1 Mb, 90.51% (1,583) of the windows from the P. × acerifolia genome contained colinearity blocks with AEK. Only 42.47% (327) of the windows in B. sinica contained colinearity blocks with AEK (Datasets S6 and S11). Additionally, although B. sinica had only one WGD event and P. × acerifolia had two WGD events, B. sinica had 33 potential genome shuffling regions between nonhomeologous chromosomes (Dataset S12), which was 4.71 times higher than P. × acerifolia (Dataset S7). B. sinica had 498 potential alignments in genome shuffling regions between nonhomeologous chromosomes, 3.77 times more than P. × acerifolia (Datasets S7 and S12). These data and comparison with other sisters to core eudicots also showed that the P. × acerifolia retained a karyotype that was most similar to the reconstructed AEK (SI Appendix, Figs. S27–S36). This "freezing" (i.e., retention of the reconstructed AEK) of the P. × acerifolia genome represents an example that genome shuffling between nonhomeologous chromosomes may not always happen following WGD.

Polyploids frequently display subgenome dominance and biased fractionation (47, 48). Examples of this include the allotetraploid Brassica napus, the allopolyploid G. hirsutum (49, 50). However, the P. × acerifolia genome exhibited extensive unbiased fractionation and no apparent subgenome dominance as also seen in a limited number of allotetraploid species, such as Tragopogon mirus (near codominance), Cucurbita maxima, Cucurbita moschata, and soybean (Glycine max) (5153). Although the lack of subgenome dominance in C. maxima has been suggested to be related to the similar number of TE-containing genes between its subgenomes (53), the mechanisms behind the absence of subgenome dominance in polyploids remain unclear and require further investigation. Sub-/neo-functionalization has long been recognized as evolutionary forces driving the retention and evolution of duplicated genes in plants (5456). The sub-/neo-functionalization of triplicated genes may have played an important role in the formation of key P. × acerifolia traits (such as flowering-time). The elucidation of the mechanism underlying P. × acerifolia diploidization requires further analysis of the epigenetic regulation and functional identification of polyploid-derived copies in additional plant samples.

The expansion of flowering-time genes in P. × acerifolia is primarily attributed to WGD compared to other no-core eudicots, with similar patterns being reported in many flowering plants, such as Nymphaea colorata (57). Our data suggested that TD was also a significant factor in the expansion of genes involved in flowering-time regulation in P. × acerifolia, such as SOC1 and FT (Fig. 4 AC and SI Appendix, Fig. S61). Both WGD and TD contribute to gene functional innovation (58). Indeed, it has been shown that TD-derived duplicates exhibit rapid gene sequence divergence (59). Nonetheless, the relative contributions of WGD and TD to flowering-time innovation in P. × acerifolia were found to be unequal. Future investigations of whether there is a significant bias in functional categorization between retained duplicated genes derived from WGD and post-WGD TD events in P. × acerifolia may provide perspectives regarding the preference and consequence of TD retention in eudicots.

We also observed that the copy numbers of some flowering-time genes, such as MFT (belonging to the PEBP family) (60), were low in most species of angiosperms. Phylogenetic analysis indicated that the two MFT homologs of P. × acerifolia originated from independent evolutionary branches and were located on different chromosomes (Fig. 4A and SI Appendix, Fig. S53). In closely related species, the number of copies of MFT also did not increase, even following multiple WGD events, which contrasts with the pattern of expansion observed in other PEBP family members (Fig. 4A and SI Appendix, Fig. S53). This finding implied that there is a widespread deletion of MFT homologs in angiosperms after WGD, which may be explained by the dominant-negative mutation hypothesis (61). That is, the duplications of MFT increase the mutational target; when the mutant variant is coexpressed with the wild-type MFT, the expression of the mutant MFT disrupts wild-type function, producing dominant-negative phenotypes. Therefore, there may be selection for removal of the extra MFT copy such that it cannot acquire dominant-negative mutations.

We found that most of the duplicates of flowering-time genes exhibited sub-/neo-functionalization. For instance, PaBBX28/29 duplicates showed different expression patterns (Fig. 4D and SI Appendix, Fig. S72), and exhibited different functions in promoting the floral transition (Fig. 4 E and F and SI Appendix, Fig. S78). Further sub-/neo-functionalization analysis of additional genes involved in flowering-time regulation will shed light on the mechanism of the 2-y flowering transition in P. × acerifolia.

In conclusion, this study revealed Platanaceae-specific WGD events, karyotypic stasis, subgenome gene diversity, and WGD and sub-/neo-functionalization as drivers of flowering-time gene evolution in P. × acerifolia. We found that the three P. × acerifolia subgenomes have remained remarkably conserved following ancient WGDs, making this species a recent, stabilized, homoploid hybrid that could be a model for studying homoploid hybridization. Accordingly, the P. × acerifolia genome serves as a valuable research subject for gene discovery in woody plants.

Materials and Methods

Materials.

To assemble the P. × acerifolia reference genome, genomic DNA was extracted from a mixture of young leaves and subpetiolar buds obtained from a single P. × acerifolia plant (voucher accession number CCAU0011690) cultivated in a nursery garden at Huazhong Agricultural University (HZAU, Wuhan, China). Additionally, in April 2020, RNA was extracted from several tissues of P. × acerifolia plants (root, branch, male inflorescences, and female inflorescences) for transcriptome sequencing, genome annotation, and tissue-specific expression analysis, which were performed with three technical replicates per tissue. To explore the multiyear flowering regulation in P. × acerifolia, subpetiolar buds were collected from April 2020 to March 2021, and subpetiolar buds at six different stages of development were collected from four few-flower varieties of P. × acerifolia between April and July 2020 (SI Appendix, Supplementary Note 2).

Genome and Transcriptome Sequencing.

The extracted genomic DNA met the requirements for both Illumina (San Diego, CA, USA) sequencing and PacBio High Fidelity (HiFi; Covaris, MA) sequencing (SI Appendix, Supplementary Note 2). Illumina™ (150-bp paired-end reads) and PacBio SMRTbell™ (10 kb) libraries were constructed for the Illumina NovaSeq and PacBio Sequel II platforms, respectively. Hi-C libraries (62) were also constructed for the MGI-SEQ 2000 platform (Wuhan, Hubei, China) to anchor genome scaffolds onto chromosomes (for details, see SI Appendix, Supplementary Note 2 and Tables S13 and S14). All the transcriptome libraries were also sequenced on the MGI-SEQ 2000 platform (150-bp paired-end reads) to obtain high-quality paired-end reads for subsequent analysis (for details, see SI Appendix, Supplementary Note 2).

De Novo Assembly and Chromosome Anchoring.

A total of 98 Gb of trimmed genomic data were generated using fastp (63). The genome size was then estimated from 17- to 31-bp K-mers using Jellyfish (64) and GenomeScope (65) (SI Appendix, Table S15). HiFi Circular Consensus Sequence (CCS) reads were generated using CCS software (github.com/PacificBiosciences/ccs) with the parameters "--min-passes 1 --min-rq 0.99 --min-length 100" Hifiasm v.0.16.0 (45) was employed to generate contigs and resolve segmental duplications. Subsequently, contigs annotated to mitochondrial DNA, plastid DNA, plasmids, and metazoa via searches against the NT (Nucleotide Sequence Database) (66) using BLASTN (67) were removed. BUSCO (68) was used to assess the completeness of the assembled genome by comparison against embryophyta_odb10 in genome mode. Sequence identity was also assessed by aligning the paired-end reads to the assembled genome using BWA (69).

For chromosome anchoring, high-quality paired-end reads were mapped to the assembled contigs using Bowtie2 (70) with the parameters "-end-to-end --very-sensitive -L 30". A total of 194,928,461 valid read pairs were generated by HiC-Pro (71) (SI Appendix, Table S16). The contigs were further clustered, ordered, and oriented onto chromosomes through valid interaction pairs using LACHESIS (72).

Genome Annotation.

Repeat elements were identified using RepeatModeler (73) and RepeatMasker (74). LTRs were detected by LTR_retriever (75) and LTR_finder (76). INFERNAL (77) was used to predict noncoding RNA based on the Rfam database (78) and tRNAscan-SE (79) was employed to identify tRNAs. The structure of protein-coding genes was annotated by RNA-seq-based prediction, ab initio prediction, and similarity-based prediction. Trinity (80), and Cufflinks [81; following alignment using TopHat2 (82)] were used to generate de novo transcript assemblies. CD-HIT (83) was employed to filter the combined transcripts with the parameters "-aL -c 0.8". Subsequently, PASA (84) was used to align transcripts to the P. × acerifolia genome and predict the complete structure with high confidence. For ab initio annotation, high-quality genic models were predicted using PASA, SNAP (85), GeneMark (86), and AUGUSTUS (87). The protein profiles of four species [N. nucifera (3), T. speciosissima (4), M. integrifolia (5), and P. occidentalis (PRJNA79937)] with the closest phylogenetic relationship to P. × acerifolia according to APG IV (33) were used for similarity-based prediction of gene structures. Finally, gene structures were predicted using MAKER (88) by integrating the above-mentioned prediction methods. The functions of protein-coding genes were identified by aligning protein sequences against the NR, GO (89), eggnog (90), InterPro (91), Uniprot (92), TrEMBL (93), and KEGG (94) databases.

Genome Phasing.

First, chromosomes colinearity in the P. × acerifolia genome was identified using WGDI (95) with the parameters "E-value ≤10−5; score ≥100". Then OGs within all P. × acerifolia chromosomes were identified using OrthoFinder (96), with Nelumbo nucifera (all chromosomes) serving as the outgroup. An LCN OG was defined as a single-copy gene in each homologous chromosome of P. × acerifolia and at least one copy in the N. nucifera genome. The protein sequences of each OG were aligned using MAFFT (97) and trimmed by trimAL (98) using default parameters. ModelTest-NG (99) was used to identify the best-fitting model and 100 bootstrap replicates were performed for each LCN OG tree using RAxML-NG (100). Phylogenetic relationships among the six homologous chromosomes in each of the seven homologous chromosome groups were inferred from LCN OGs using ASTRAL (101). Finally, RAxML-NG was used to calculate the optimal branch length with the amino acid supermatrix concatenated by FASconCAT (102).

Then homeologous chromosomes were distinguished using ASTRAL topology, following which the two chromosome pairs with the highest similarity among the three pairs were distinguished by branch length. Once the three subgenomes had been separated, each subgenome was further divided into a long group and a short group based on the assembly chromosome length.The workflow for subgenome phasing is described in SI Appendix, Fig. S6B. To capture more genetic information, a representative haploid genome was selected by only including the longer chromosomes from each homologous pair as the pseudohaploid genome of Platanus (SI Appendix, Table S10). We also used the phylogeny-based subgenome phasing method to successfully phase the subgenomes of the allotetraploid G. hirsutum and allohexaploid T. aestivum as validation tests (SI Appendix, Fig. S8 and Supplementary Note 4).

For subgenome identification using the K-mer-based approach, SubPhaser (42) was used with a comprehensive combination of 10 K-mer values (13, 15, 17, 19, 21, 23, 25, 27, 29, and 31) and eight minimum frequency thresholds for each K-mer (25, 50, 75, 100, 150, 200, 250, and 300) (a total of 80 parameter combinations including the default values).

WGD Analysis.

To identify WGD events in P. × acerifolia and Proteales, the Ks distribution of duplicate genes and synteny were investigated. For the Ks distribution analysis, the Ks values of orthologous gene pairs in the syntenic blocks were estimated by the Nei-Gojobori method using WGDI (95) with a maximum gene gap of 40. The Ks distribution of paralogs and orthologs of P. × acerifolia, N. nucifera, M. integrifolia, and T. speciosissima were estimated by WGDI. The Ks values of P. × acerifolia/V. vinifera orthologs (with a mean value of 1.07) were used to calculate the number of substitutions per synonymous site per year (r) using the formula of divergence date = Ks/(2r) following the methods reported by Badouin et al. (103). The date of divergence for P. × acerifolia/V. vinifera was 124 mya according to the Timetree (104) service (http://www.timetree.org). The dates of WGD events for P. × acerifolia and other Proteales species were estimated based on the same r value and the respective mean Ks values. Then, the sizer plot of the Ks distributions was plotted by WGDI. Macrosynteny between P. × acerifolia and other Proteales species was investigated using JCVI (105) with the parameters “E-value = 10−5 and n = 20” (at least 20 genes within one synteny block) (SI Appendix, Supplementary Note 7).

Karyotype Evolution Analysis.

To characterize the karyotype changes in Platanus and compare them with other Proteales species after WGDs, an analysis of karyotype evolution was performed. First, paralogous and orthologous relationships between the published AEK model (21) and the genomes of 10 representative eudicots (P. × acerifolia, eight noncore eudicots [A. coerulea, Coptis chinensis, Papaver somniferum, N. nucifera, M. integrifolia, T. speciosissima, T. sinense, and B. sinica], and the core eudicot V. vinifera) were identified following previously reported methods (106). Pair-wise sequence alignments were performed by LAST (107). The syntenic relationships within each genome, between genomes, and between the AEK model and all sampled genomes were analyzed by JCVI (SI Appendix, Supplementary Note 8). Finally, all the species were mapped to the AEK model to identify the karyotypic differences among these modern species, and the ancestors of Proteales were inferred following the shortest evolutionary pathway to the AEK.

Fractionation and Sub-/Neo-Functionalization Analyses.

To analyze the fractionation bias among the subgenomes of P. × acerifolia, we used the FractBias tool of SynMap to obtain quantitative comparisons of gene loss between P. × acerifolia and B. sinica with per-50-gene windows (108). The distribution curve of fractionation was plotted in GraphPad Prism 10 (www.graphpad.com). Colinearity between each assembled subgenome of P. × acerifolia and the genomes of A. coerulea, N. nucifera, V. vinifera, and P. trichocarpa was analyzed by JCVI. Fractionations were detected from gene losses in colinear blocks among the three subgenomes. Fractionation distribution and retained gene numbers were visualized using Hiplot (109).

The gene copies present in triplicate were identified, following the DupGen_finder pipeline, as previously described (110). To compare gene expression across the subgenomes of the haploid genome, we calculated the annual expression profile, which refers to the gene expression of genes during subpetiolar bud development from April to March (11 periods; see Dataset S13). Our analysis included the following: expression averaged over all rhythm periods (G effect), differences in mean expression for each of the subpetiolar bud development rhythm periods averaged over the gene copies (R effect), and whether the mean expression within a rhythm period varied by different gene copies (G × R effect), as previously described (111). Statistics for the different expression patterns were performed by split-plot ANOVA of duplicate pairs using R (112). The existence of the G × R effect means that there is a difference in the expression levels of genes from the subgenomes A, B, and C in one period but not another, thereby showing a complementary expression pattern that is referred to here as sub-/neo-functionalization (111). Subfunctionalization could not be differentiated from neo-functionalization in this study because of lacking close ancestors of Platanus. Therefore, subfunctionalization and neo-functionalization were collectively considered as a whole as suggested in previous studies (111).

Subgenome Dominance.

Colinearity among the three subgenomes was analyzed and plotted using JCVI (105), while subgenomic dominance was analyzed by using the method as described for Pogostemon cablin (113). Significant differences in the counts of genes with the same Pfam (114) domain among the three subgenomes were determined by Fisher’s exact test (P ≤ 0.05). Then, the paired t-test was employed to reveal the differences in the expression of the gene copies between each two of the three subgenomes, respectively. The DEG pairs based on transcriptomes from April to the following March were identified using |Fold Change| ≥ 2 and P ≤ 0.05 as thresholds.

To determine whether there was molecular divergence among subgenomes in P. × acerifolia, Ks, Ka, and Ka/Ks values were calculated based on orthologues in syntenic blocks in the three subgenomes of P. × acerifolia and the genome of N. nucifera. Putative ortholog pairs were first identified using all-against-all protein similarity alignment between all genes in the genomes of P. × acerifolia and N. nucifera using BLASTP with a cut-off E-value of 10−5 and a minimum score of 100. Subsequently, the Ks, Ka, and Ka/Ks values of paralogs in the syntenic blocks were calculated based on the putative ortholog pairs as described above in the WGD Analysis section. Finally, GraphPad Prism 10 was employed to reveal the differences in the expression of the gene copies among the three subgenomes.

Paired-end reads were aligned to the genome using BWA (69). Then, single nucleotide polymorphisms were identified by GATK v.4.0.0 employing the parameters “QD < 2.0 || MQ <40.0 || FS >60.0 || SOR >3.0 || MQRankSum <−12.5 || ReadPosRankSum <−8.0” (115). The heterozygosities (π) among P. × acerifolia subgenomes A, B, and C were calculated by VCFtools with the parameters “--window-pi 500000, --window-pi-step 500000” (116). Finally, the distribution curve was plotted in GraphPad Prism 10.

Identification and Expression Quantification of Flowering-Time Gene Duplicates.

To identify orthologs and paralogs of the verified flowering-time genes in P. × acerifolia, the predicted proteomes of P. × acerifolia and nine other angiosperm species [N. colorata, O. sativa, A. coerulea, N. nucifera, T. sinense, B. sinica, V. vinifera, Solanum lycopersicum, and A. thaliana], were searched using BLASTP (1e−30) and InterProScan (117) with Chlamydomonas reinhardtii (118) as the outgroup. The domains used for the InterProScan search are listed in SI Appendix, Supplementary Note 12. For phylogenetic analysis, sequences were aligned using MAFFT with the parameter E-INS-I (97). The phylogenetic tree of MADS-box genes was constructed using IQ-TREE2 with untrimmed aligned sequences (119). Aligned sequences of other gene families were trimmed by trimAL (98), and the phylogenetic tree was constructed using RAxML-NG (100). The models used to construct the phylogenetic tree were identified using ModelTest-NG (101); the specific models used for each gene family are listed in SI Appendix, Supplementary Note 12. Phylogenetic trees were visualized using ITOL (120). The identification of orthologs and paralogs was based on genes that formed a clade with known A. thaliana genes, the sequences of conserved domains, and gene family information in published articles. Details regarding the identification of orthology/paralogy relationships are given in SI Appendix, Supplementary Note 12. The genome-wide syntenic genes in P. × acerifolia were identified using JCVI with default parameter settings. The putative genes expanded by post-WGD TD (duplicate gene clusters in one subgenome of Platanus but not in other subgenomes or other angiosperms) or WGD were manually verified based on genomic location information.

The transcriptomes were mapped using Bowtie2, and the fragments were counted and normalized to fragments per kilobase per million reads using RSEM (121). Gene expression was visualized using the R package pheatmap (122) and GraphPad Prism 10.

The primers used for cloning the copies of PaBBX28/29 were designed using Primer5 (123) based on the genome of P. × acerifolia (Dataset S7). Mixed-buds cDNA was used as the template for PCR. The amplification products were cloned into the PMD18-T vector (TaKaRa, Otsu, Japan) for sequencing. Conserved domain analysis was performed using the CCD (124). Details of the methods used for the functional verification of sub-/neo-functionalization of PaBBX28/29 are described in SI Appendix, Supplementary Note 13.

Supplementary Material

Appendix 01 (PDF)

pnas.2319679121.sapp.pdf (19.1MB, pdf)

Dataset S01 (XLSX)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2319679121.sd03.xlsx (30.8KB, xlsx)

Dataset S04 (XLSX)

Dataset S05 (XLSX)

Dataset S06 (XLSX)

Dataset S07 (XLSX)

pnas.2319679121.sd07.xlsx (30.6KB, xlsx)

Dataset S08 (XLSX)

Dataset S09 (XLSX)

pnas.2319679121.sd09.xlsx (132.9KB, xlsx)

Dataset S10 (XLSX)

Dataset S11 (XLSX)

pnas.2319679121.sd11.xlsx (755.8KB, xlsx)

Dataset S12 (XLSX)

pnas.2319679121.sd12.xlsx (32.6KB, xlsx)

Dataset S13 (XLSX)

pnas.2319679121.sd13.xlsx (24.7MB, xlsx)

Dataset S14 (XLSX)

pnas.2319679121.sd14.xlsx (12.3KB, xlsx)

Acknowledgments

We are thankful to Dr. Chai-Shian Kua, Dr. Charles Cannon, and The Morton Arboretum for providing the sample collections, and Guoxiao Tu for his courtesy of the Platanus image used in Fig. 1B and the National Key Laboratory for Germplasm Innovation Utilization of Horticultural Crops, HZAU for the help with this study. This work was supported by the National Natural Science Foundation of China (grant number 32272762).

Author contributions

X.Y., G.S., M.S., J.Z., D.E.S., G.N., and M.B. designed research; X.Y., G.S., Z.Z., and Y.L. performed research; Y.H. contributed new reagents/analytic tools; X.Y., G.S., M.S., R.C., R.L., S.W., and Zhenhua Liu analyzed data; M.B. provide materials; and X.Y., G.S., M.S., S.S., Zhongjian Liu, P.S.S., J.Z., D.E.S., G.N., and M.B. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

Although PNAS asks authors to adhere to United Nations naming conventions for maps (https://www.un.org/geospatial/mapsgeo), our policy is to publish maps as provided by the authors.

Contributor Information

Jiaqi Zhang, Email: jiaqizhang@mail.hzau.edu.cn.

Douglas E. Soltis, Email: dsoltis@ufl.edu.

Guogui Ning, Email: ggning@mail.hzau.edu.cn.

Manzhu Bao, Email: mzbao@mail.hzau.edu.cn.

Data, Materials, and Software Availability

The transcriptome data generated during the current study are available in the NCBI Sequence Read Archive (SRA) repository under project accession: PRJNA1010979 (125). The assembly and annotation of the P. × acerifolia genome have been deposited in CoGe (Genome id: 67387). The underlying data and scripts from this study have been deposited in the Dryad Digital Repository (https://doi.org/10.5061/dryad.j6q573nm9) (126).

Supporting Information

References

  • 1.Soltis D. E., et al. , “Phylogeny of angiosperms: An overview” in Phylogeny and Evolution of the Angiosperms (University of Chicago Press, 2018), pp. 35–62. [Google Scholar]
  • 2.Kramer E. M., Zimmer E. A., Gene duplication and floral developmental genetics of basal eudicots. Adv. Bot. Res. 44, 353–384 (2006). [Google Scholar]
  • 3.Ming R., et al. , Genome of the long-living sacred lotus (Nelumbo nucifera Gaertn.). Genome Biol. 14, 1–11 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Chen S. H., et al. , Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C. Mol. Ecol. Resour. 22, 1836–1854 (2022). [DOI] [PubMed] [Google Scholar]
  • 5.Nock C. J., et al. , Genome and transcriptome sequencing characterises the gene space of Macadamia integrifolia (Proteaceae). BMC Genom. 17, 1–12 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chang J., et al. , The genome of the king protea, Protea cynaroides. Plant J. 113, 262–276 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Filiault D. L., et al. , The Aquilegia genome provides insight into adaptive radiation and reveals an extraordinarily polymorphic chromosome with a unique history. elife 7, e36426 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Guo L., et al. , The opium poppy genome and morphinan production. Science 362, 343–347 (2018). [DOI] [PubMed] [Google Scholar]
  • 9.Chanderbali A. S., et al. , Buxus and Tetracentron genomes help resolve eudicot genome history. Nat. Commun. 13, 643 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Stevens P. F., Dara form “Angiosperm phylogeny website version 14”. Angiosperm Phylogeny Website. Available at: http://www.mobot.org/MOBOT/research/APweb. Deposited 03 March 2022.
  • 11.Tropicos, Data from "Facilitated by the missouri botanical garden". Tropicos. Available at: https://www.tropicos.org/. Deposited 02 August.
  • 12.POWO, Data from "Plants of the world online”. Royal Botanic Gardens, Kew. Available at: http://www.plantsoftheworldonline.org. Deposited 02 May 2023.
  • 13.Henry A., Flood M. G., "The history of the london plane, Platanus acerifolia, with notes on the genus Platanus" in Proceedings of the Royal Irish Academy Section B: Biological, Geological, and Chemical Science (Royal Irish Academy, Dublin, Ireland, 1919), pp. 9–28. [Google Scholar]
  • 14.Willis K. J., Petrokofsky G., The natural capital of city trees. Science 356, 374–376 (2017). [DOI] [PubMed] [Google Scholar]
  • 15.Rice A., et al. , The Chromosome Counts Database (CCDB)–A community resource of plant chromosome numbers. New Phytol. 206, 19–26 (2015). [DOI] [PubMed] [Google Scholar]
  • 16.Stebbins G. L., "Variation and evolution in plants" in Variation and Evolution in Plants, Dunn L. C., Ed. (Columbia University Press, 1950), pp. 423–572. [Google Scholar]
  • 17.Grant V., "Polyploidy: Range and frequency" in Plant Speciation (Columbia University Press, 1981), pp. 283–295. [Google Scholar]
  • 18.Oginuma K., Tobe H., Karyomorphology and evolution in some Hamamelidaceae and Platanaceae (Hamamelididae; Hamamelidales). Bot. Mag. Tokyo 104, 115–135 (1991). [Google Scholar]
  • 19.Salse J., Ancestors of modern plant crops. Curr. Opin. Plant Biol. 30, 134–142 (2016). [DOI] [PubMed] [Google Scholar]
  • 20.Amborella Genome Project et al. , The Amborella genome and the evolution of flowering plants. Science 342, 1241089 (2013). [DOI] [PubMed] [Google Scholar]
  • 21.Wang Z., et al. , A high-quality Buxus austro-yunnanensis (Buxales) genome provides new insights into karyotype evolution in early eudicots. BMC Biol. 20, 216 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Gui S., et al. , Improving Nelumbo nucifera genome assemblies using high-resolution genetic maps and BioNano genome mapping reveals ancient chromosome rearrangements. Plant J. 94, 721–734 (2018). [DOI] [PubMed] [Google Scholar]
  • 23.Murat F., Van de Peer Y., Salse J., Decoding plant and animal genome plasticity from differential paleo-evolutionary patterns and processes. Genome Biol. Evol. 4, 917–928 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Nixon K. C., Poole J. M., Revision of the Mexican and Guatemalan species of Platanus (Platanaceae). Lundellia 6, 103–137 (2003). [Google Scholar]
  • 25.Bouché F., Lobet G., Tocquin P., Périlleux C., FLOR-ID: An interactive database of flowering-time gene networks in Arabidopsis thaliana. Nucleic Acids Res. 44, D1167–D1171 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Zhou S., et al. , Transcriptional and post-transcriptional regulation of heading date in rice. New Phytol. 230, 943–956 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cai F., et al. , Identification and characterisation of a novel FT orthologous gene in London plane with a distinct expression response to environmental stimuli compared to PaFT. Plant Biol. 21, 1039–1051 (2019). [DOI] [PubMed] [Google Scholar]
  • 28.Han H., et al. , Four SQUAMOSA PROMOTER BINDING PROTEIN-LIKE homologs from a basal eudicot tree (Platanus acerifolia) show diverse expression pattern and ability of inducing early flowering in Arabidopsis. Trees 30, 1417–1428 (2016). [Google Scholar]
  • 29.Zhang J., et al. , The FLOWERING LOCUS T orthologous gene of Platanus acerifolia is expressed as alternatively spliced forms with distinct spatial and temporal patterns. Plant Biol. 13, 809–820 (2011). [DOI] [PubMed] [Google Scholar]
  • 30.Zhang S., "Functional and evolutional analysis of flowering regulatory genes in Platanus acerifolia [in Chinese], " PhD dissertation, Huazhong Agricultural University, Wuhan: (2016). [Google Scholar]
  • 31.Zhang S., et al. , Identification and characterization of FRUITFULL-like genes from Platanus acerifolia, a basal eudicot tree. Plant Sci. 280, 206–218 (2019). [DOI] [PubMed] [Google Scholar]
  • 32.Zhang S., et al. , Functional characterization of three TERMINAL FLOWER 1-like genes from Platanus acerifolia. Plant Cell Rep. 42, 1071–1088 (2023). [DOI] [PubMed] [Google Scholar]
  • 33.Angiosperm Phylogeny Group et al. , An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG IV. Bot. J. Linn. Soc. 181, 1–20 (2016). [Google Scholar]
  • 34.Shao C., et al. , A class II TCP transcription factor PaTCP4 from Platanus acerifolia regulates trichome formation in Arabidopsis. DNA Cell Biol. 40, 1235–1250 (2021). [DOI] [PubMed] [Google Scholar]
  • 35.Murat F., Armero A., Pont C., Klopp C., Salse J., Reconstructing the genome of the most recent common ancestor of flowering plants. Nat. Genet. 49, 490–496 (2017). [DOI] [PubMed] [Google Scholar]
  • 36.Birchler J. A., Yang H., The multiple fates of gene duplications: Deletion, hypofunctionalization, subfunctionalization, neofunctionalization, dosage balance constraints, and neutral variation. Plant Cell 34, 2466–2474 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Cheng F., et al. , Gene retention, fractionation and subgenome differences in polyploid plants. Nat. Plants 4, 258–268 (2018). [DOI] [PubMed] [Google Scholar]
  • 38.Cheng H., Concepcion G. T., Feng X., Zhang H., Li H., Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Kyriakidou M., Tai H. H., Anglin N. L., Ellis D., Strömvik M. V., Current strategies of polyploid plant genome sequence assembly. Front. Plant Sci. 9, 385637 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Edger P. P., McKain M. R., Bird K. A., VanBuren R., Subgenome assignment in allopolyploids: Challenges and future directions. Curr. Opin. Plant Biol. 42, 76–80 (2018). [DOI] [PubMed] [Google Scholar]
  • 41.Freyman W. A., Johnson M. G., Rothfels C. J., homologizer: Phylogenetic phasing of gene copies into polyploid subgenomes. Methods. Ecol. Evol. 14, 1230–1244 (2023). [DOI] [PubMed] [Google Scholar]
  • 42.Jia K. H., et al. , SubPhaser: A robust allopolyploid subgenome phasing method based on subgenome-specific k-mers. New Phytol. 235, 801–809 (2022). [DOI] [PubMed] [Google Scholar]
  • 43.Novikova P. Y., Hohmann N., Van de Peer Y., Polyploid Arabidopsis species originated around recent glaciation maxima. Curr. Opin. Plant Biol. 42, 8–15 (2018). [DOI] [PubMed] [Google Scholar]
  • 44.Van de Peer Y., Ashman T. L., Soltis P. S., Soltis D. E., Polyploidy: An evolutionary and ecological force in stressful times. Plant Cell 33, 11–26 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Murat F., et al. , Ancestral grass karyotype reconstruction unravels new mechanisms of genome shuffling as a source of plant evolution. Genome Res. 20, 1545–1557 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Zheng C., Chen E., Albert V. A., Lyons E., Sankoff D., Ancient eudicot hexaploidy meets ancestral eurosid gene order. BMC Genom. 14, 1–13 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Doyle J. J., et al. , Evolutionary genetics of genome merger and doubling in plants. Annu. Rev. Genet. 42, 443–461 (2008). [DOI] [PubMed] [Google Scholar]
  • 48.Soltis D. E., Visger C. J., Soltis P. S., The polyploidy revolution then.And now: Stebbins revisited. Am. J. Bot. 101, 1057–1078 (2014). [DOI] [PubMed] [Google Scholar]
  • 49.Cheng F., et al. , Biased gene fractionation and dominant gene expression among the subgenomes of Brassica rapa. PloS One 7, e36442 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Zheng D., et al. , Histone modifications define expression bias of homoeologous genomes in allotetraploid cotton. Plant Physiol. 172, 1760–1771 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Koh J., Soltis P. S., Soltis D. E., Homeolog loss and expression changes in natural populations of the recently and repeatedly formed allotetraploid Tragopogon mirus (Asteraceae). BMC Genom. 11, 1–16 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Gill N., et al. , Molecular and chromosomal evidence for allopolyploidy in soybean. Plant Physiol. 151, 1167–1174 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Sun H., et al. , Karyotype stability and unbiased fractionation in the paleo-allotetraploid Cucurbita genomes. Mol. Plant. 10, 1293–1306 (2017). [DOI] [PubMed] [Google Scholar]
  • 54.Lynch M., Conery J. S., The evolutionary fate and consequences of duplicate genes. Science 290, 1151–1155 (2000). [DOI] [PubMed] [Google Scholar]
  • 55.Lynch M., Conery J. S., The evolutionary demography of duplicate genes. Genome Evol. 3, 35–44 (2003). [PubMed] [Google Scholar]
  • 56.Soltis P. S., Soltis D. E., The role of hybridization in plant speciation. Annu. Rev. Plant Biol. 60, 561–588 (2009). [DOI] [PubMed] [Google Scholar]
  • 57.Zhang L., et al. , The water lily genome and the early evolution of flowering plants. Nature 577, 79–84 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Qiao X., et al. , Gene duplication and evolution in recurring polyploidization–diploidization cycles in plants. Genome Biol. 20, 1–23 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Liang Z., Schnable J. C., Functional divergence between subgenomes and gene pairs after whole genome duplications. Mol. Plant. 11, 388–397 (2018). [DOI] [PubMed] [Google Scholar]
  • 60.Yoo Y. S., et al. , Acceleration of flowering by overexpression of MFT (MOTHER OF FT AND TFL1). Mol. cells 17, 95–101 (2004). [PubMed] [Google Scholar]
  • 61.De Smet R., et al. , Convergent gene loss following gene and genome duplications creates single-copy families in flowering plants. Proc. Natl. Acad. Sci. U.S.A. 110, 2898–2903 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Zhu W., et al. , Altered chromatin compaction and histone methylation drive non-additive gene expression in an interspecific Arabidopsis hybrid. Genome Biol. 18, 1–16 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Chen S., Zhou Y., Chen Y., Gu J., fastp: An ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Marçais G., Kingsford C., A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764–770 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Vurture G. W., et al. , GenomeScope: Fast reference-free genome profiling from short reads. Bioinformatics 33, 2202–2204 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Sayers E. W., et al. , Database resources of the national center for biotechnology information. Nucleic Acids Res. 50, D20–D26 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.McGinnis S., Madden T. L., BLAST: At the core of a powerful and diverse set of sequence analysis tools. Nucleic Acids Res. 32, W20–W25 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Manni M., Berkeley M. R., Seppey M., Simão F. A., Zdobnov E. M., BUSCO update: Novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol. Biol. Evol. 38, 4647–4654 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Li H., Durbin R., Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 26, 589–595 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Langmead B., Wilks C., Antonescu V., Charles R., Scaling read aligners to hundreds of threads on general-purpose processors. Bioinformatics 35, 421–432 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Servant N., et al. , HiC-Pro: An optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 1–11 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Burton J. N., et al. , Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119–1125 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Flynn J. M., et al. , RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. U.S.A. 117, 9451–9457 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Chen N., Using repeat masker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 5, 4.10.1–4.10.14 (2009). [DOI] [PubMed] [Google Scholar]
  • 75.Ou S., Jiang N., LTR_retriever: A highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Xu Z., Wang H., LTR_FINDER: An efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265–W268 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Nawrocki E. P., Eddy S. R., Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29, 2933–2935 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Kalvari I., et al. , Rfam 14: Expanded coverage of metagenomic, viral and microRNA families. Nucleic Acids Res. 49, D192–D200 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Chan P. P., Lowe T. M., tRNAscan-SE: Searching for tRNA genes in genomic sequences. Methods Mol. Biol. 1962, 1–14 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Grabherr M. G., et al. , Trinity: Reconstructing a full-length transcriptome without a genome from RNA-Seq data. Nat. Biotechnol. 29, 644–652 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Trapnell C., et al. , Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat. Biotechnol. 28, 511–515 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Kim D., et al. , TopHat2: Accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome Biol. 14, 1–13 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Fu L., Niu B., Zhu Z., Wu S., Li W., CD-HIT: Accelerated for clustering the next-generation sequencing data. Bioinformatics 28, 3150–3152 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Haas B. J., et al. , Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 9, 1–22 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Korf I., Gene finding in novel genomes. BMC Bioinformatics 5, 1–9 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Brůna T., Lomsadze A., Borodovsky M., GeneMark-EP+: Eukaryotic gene prediction with self-training in the space of genes and proteins NAR Genom. Bioinform. 2, lqaa026 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Stanke M., Diekhans M., Baertsch R., Haussler D., Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24, 637–644 (2008). [DOI] [PubMed] [Google Scholar]
  • 88.Holt C., Yandell M., MAKER2: An annotation pipeline and genome-database management tool for second-generation genome projects. BMC Bioinformatics 12, 1–14 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Ashburner M., et al. , Gene ontology: Tool for the unification of biology. Nat. Genet. 25, 25–29 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Huerta-Cepas J., et al. , eggNOG 5.0: A hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 47, D309–D314 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Blum M., et al. , The InterPro protein families and domains database: 20 years on. Nucleic Acids Res. 49, D344–D354 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.UniProt Consortium, UniProt: The universal protein knowledgebase in 2021. Nucleic Acids Res. 49, D480–D489 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Boeckmann B., et al. , The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 31, 365–370 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Kanehisa M., Furumichi M., Sato Y., Ishiguro-Watanabe M., Tanabe M., KEGG: Integrating viruses and cellular organisms. Nucleic Acids Res. 49, D545–D551 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Sun P., et al. , WGDI: A user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol. Plant. 15, 1841–1851 (2022). [DOI] [PubMed] [Google Scholar]
  • 96.Emms D. M., Kelly S., OrthoFinder: Solving fundamental biases in whole genome comparisons dramatically improves orthogroup inference accuracy. Genome Biol. 16, 1–14 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Katoh K., Standley D. M., MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. Mol. Biol. Evol. 30, 772–780 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Capella-Gutiérrez S., Silla-Martínez J. M., Gabaldón T., trimAl: A tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25, 1972–1973 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Darriba D., et al. , ModelTest-NG: A new and scalable tool for the selection of DNA and protein evolutionary models. Mol. Biol. Evol. 37, 291–294 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Kozlov A. M., Darriba D., Flouri T., Morel B., Stamatakis A., RAxML-NG: A fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics 35, 4453–4455 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Zhang C., Rabiee M., Sayyari E., Mirarab S., ASTRAL-III: Polynomial time species tree reconstruction from partially resolved gene trees. BMC Bioinformatics 19, 15–30 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Kück P., Meusemann K., FASconCAT: Convenient handling of data matrices. Mol. Phylogenet. Evol. 56, 1115–1118 (2010). [DOI] [PubMed] [Google Scholar]
  • 103.Badouin H., et al. , The sunflower genome provides insights into oil metabolism, flowering and Asterid evolution. Nature 546, 148–152 (2017). [DOI] [PubMed] [Google Scholar]
  • 104.Kumar S., Stecher G., Suleski M., Hedges S. B., TimeTree: A resource for timelines, timetrees, and divergence times. Mol. Biol. Evol. 34, 1812–1819 (2017). [DOI] [PubMed] [Google Scholar]
  • 105.Tang H., et al. , Synteny and collinearity in plant genomes. Science 320, 486–488 (2008). [DOI] [PubMed] [Google Scholar]
  • 106.Sun P., et al. , Early diversification and karyotype evolution of flowering plants. Research Square [Preprint] (2022). 10.21203/rs.3.rs-1410884/v1. Accessed 6 November 2022. [DOI]
  • 107.Kiełbasa S. M., Wan R., Sato K., Horton P., Frith M. C., Adaptive seeds tame genomic sequence comparison. Genome Res. 21, 487–493 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Joyce B. L., et al. , FractBias: A graphical tool for assessing fractionation bias following polyploidy. Bioinformatics 33, 552–554 (2017). [DOI] [PubMed] [Google Scholar]
  • 109.Li J., et al. , Hiplot: A comprehensive and easy-to-use web service for boosting publication-ready biomedical data visualization. Brief. Bioinform. 23, bbac261 (2022). [DOI] [PubMed] [Google Scholar]
  • 110.Qiao X., et al. , Gene duplication and evolution in recurring polyploidization-diploidization cycles in plants. Genome Biol. 20, 1–23 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Duarte J. M., et al. , Expression pattern shifts following duplication indicative of subfunctionalization and neofunctionalization in regulatory genes of Arabidopsis. Mol. Biol. Evol. 23, 469–478 (2006). [DOI] [PubMed] [Google Scholar]
  • 112.R-Core-team, R. R-Core-team. https://www.r-project.org. Deposited 10 March 2022.
  • 113.Shen Y., et al. , Chromosome-level and haplotype-resolved genome provides insight into the tetraploid hybrid origin of patchouli. Nat. Commun. 13, 3511 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Mistry J., et al. , Pfam: The protein families database in 2021. Nucleic Acids Res. 49, D412–D419 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.DePristo M. A., et al. , A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat. Genet. 43, 491–498 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Danecek P., et al. , The variant call format and VCFtools. Bioinformatics 27, 2156–2158 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Quevillon E., et al. , InterProScan: Protein domains identifier. Nucleic Acids Res. 33, W116–W120 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Merchant S. S., et al. , The Chlamydomonas genome reveals the evolution of key animal and plant functions. Science 318, 245–250 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Minh B. Q., et al. , IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Mol. Biol. Evol. 37, 1530–1534 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Letunic I., Bork P., Interactive Tree Of Life (iTOL) v5: An online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 49, W293–W296 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Li B., Dewey C. N., RSEM: Accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics 12, 1–16 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Malcolm P., Heatmaps. CRAN. https://CRAN.R-project.org/package=pheatmap. Deposited 11 August 2022.
  • 123.Lalitha S., Primer premier 5. Biotech. Softw. Intern. Rep. 1, 270–272 (2000). [Google Scholar]
  • 124.Lu S., et al. , CDD/SPARCLE: The conserved domain database in 2020. Nucleic Acids Res. 48, D265–D268 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Yan X., Data from “Platanus × acerifolia transcriptome of the floral transition.” National Center for Biotechnology Information. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA1010979. Deposited 31 August 2023.
  • 126.Yan X., Data from “Assemblies, associated annotation files, and analysis source data of Platanus x acerifloia genome.” Dryad. 10.5061/dryad.j6q573nm9. Deposited 07 September 2023. [DOI]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

pnas.2319679121.sapp.pdf (19.1MB, pdf)

Dataset S01 (XLSX)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2319679121.sd03.xlsx (30.8KB, xlsx)

Dataset S04 (XLSX)

Dataset S05 (XLSX)

Dataset S06 (XLSX)

Dataset S07 (XLSX)

pnas.2319679121.sd07.xlsx (30.6KB, xlsx)

Dataset S08 (XLSX)

Dataset S09 (XLSX)

pnas.2319679121.sd09.xlsx (132.9KB, xlsx)

Dataset S10 (XLSX)

Dataset S11 (XLSX)

pnas.2319679121.sd11.xlsx (755.8KB, xlsx)

Dataset S12 (XLSX)

pnas.2319679121.sd12.xlsx (32.6KB, xlsx)

Dataset S13 (XLSX)

pnas.2319679121.sd13.xlsx (24.7MB, xlsx)

Dataset S14 (XLSX)

pnas.2319679121.sd14.xlsx (12.3KB, xlsx)

Data Availability Statement

The transcriptome data generated during the current study are available in the NCBI Sequence Read Archive (SRA) repository under project accession: PRJNA1010979 (125). The assembly and annotation of the P. × acerifolia genome have been deposited in CoGe (Genome id: 67387). The underlying data and scripts from this study have been deposited in the Dryad Digital Repository (https://doi.org/10.5061/dryad.j6q573nm9) (126).


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES