Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 May 26;17:6866. doi: 10.1038/s41467-026-73233-7

An X-linked sex determination mechanism in cannabis and hop

Sarah B Carey 1, Philip C Bentz 1, John T Lovell 1, Laramie M Akozbek 1,2, Zachary A Myers 1, Walid Korani 1, Joshua S Havill 3, Lillian Padgitt-Cobb 4, Ryan C Lynch 4, Nicholas Allsing 4, Jack Mangels 5, Zachary Stansell 6,7, George M Stack 6,7, Tyler Gordon 6, Austin Osmanski 1, Katherine A Easterling 8,15, Leonardo R Orozco 9, Zach E Marcus 9, Haley Hale 1, Hannah McCoy 1, Zachary Meharg 1,2, Jane Grimwood 1, Lawrence B Smart 7, Daniela Vergara 9, Rafael F Guerrero 10, Nolan C Kane 9, Rich Fletcher 5, John K McKay 5,11, Todd P Michael 4,12,13,14, Gary J Muehlbauer 3, Josh Clevenger 1, Alex Harkess 1,
PMCID: PMC13389243  PMID: 42192122

Abstract

Sex chromosomes in cannabis and hop were identified a century ago because of their obvious visible differences in size (heteromorphy). However, we know little about the genes they contain that control the development of the inflorescences. Here we assembled genomes, with phased sex chromosomes, for hop and cannabis. The XY chromosomes share an origin prior to the divergence between the genera around 36 MYA. Due to the inheritance patterns of the XYs, the male-specific region of the Y is highly-degenerated, with substantial gene loss, while the X shows faster rates of molecular evolution. Consistent with the hypothesis that these species lack an active-Y system, no clear sex-determining genes reside on the Y. Instead, an X-linked homolog of aminocyclopropane-1-carboxylate synthase (ACS), that is involved in the ethylene biosynthesis pathway, determines the fate of the female inflorescence. Beyond sex determination, the sex chromosomes contribute to the sexual dimorphism in ecology and physiology and have played a role in the domestication and breeding of these species.

Subject terms: Evolutionary genetics, Plant reproduction, Plant genetics, Plant hormones


The ancient X and Y chromosomes of cannabis and hop display interesting patterns of molecular evolution and harbor key floral genes. Here the authors show that the X chromosome, not the Y, determines sex, and identify an X-linked ethylene biosynthesis gene as a candidate for sex-determination.

Introduction

The inflorescences in the Cannabaceae family have a deep history of human application. Cannabis sativa L. was cultivated by humans around 12,000 years ago1 and is still grown globally for medicinal compounds (cannabinoids) concentrated in glandular trichomes on female inflorescences, as well as for multi-purpose oil in seeds, and bast fiber in stems2. The female inflorescences of hop (Humulus lupulus L.) also concentrate secondary compounds in glandular trichomes and have been used for brewing beer3 since as far back as 1200 years ago, where they provide the characteristic bitterness and aroma. Thus, females in both genera have been selected over centuries of human cultivation. While there is a chromosomal basis that determines sex in these species, a major complication is that dioecy is not strictly expressed. Typical XX females have been shown to produce male flowers4, which can cause unintended fertilization and changes in secondary metabolite production5,6. Curiously, reversions to monoecy have also been observed in XX and XY individuals, particularly in cannabis7,8. Identifying the genes involved with sex-determination is critical to better control sex expression and enhance the production of fiber, medicinal compounds, and seed oil.

Since their discovery9, sex chromosomes have been of keen interest to study because of their unique inheritance patterns that drive differences in their size, structure, and gene content, as well as their role in sex determination. A century ago, the sex chromosomes of cannabis and hop were among the first to be identified in flowering plants because of their heteromorphic XY pairs10,11. Humulus lupulus var. lupulus, the European hop, has one of the only known XY pairs in plants where the Y chromosome is cytologically smaller than the X. Cannabis sativa, instead, has a Y chromosome that is larger than the X, but this form of heteromorphy is also rare across angiosperms; most examined plant sex chromosomes are homomorphic12, including the other botanical varieties of H. lupulus13. Although there is extreme variation in cytotype, the shared origin of the XY pairs14 suggests that cannabis and hop share the same sex determination mechanism.

Here we assemble fully-phased, chromosome-scale genomes with both X and Y chromosomes for cannabis and hop, and use these to examine their sex chromosome structure and evolution. Counter to the expectations from other XY systems, like in mammals, where the non-recombining region of the Y contains the sex-determining gene15, we instead identify a candidate sex-determination gene on the X, consistent with the leading hypothesis that hop and cannabis do not have an active-Y system16. This result, along with the additional floral genes identified, underlines the importance of sex chromosomes in the development of the economically important cannabis and hop flowers.

Results

Phased sex chromosome assemblies

Until recently, assembling Cannabaceae genomes was a challenge because species in the family are highly heterozygous with repeat-rich genomes17,18. Further, sex chromosomes are particularly arduous to contiguously assemble19. For example, the Y chromosome was the last to be included in the human “T2T” genome20, and sex chromosomes are incomplete in many mammalian genomes. While progress has been made to identify sex-linked sequences and to assemble sex chromosome pairs (e.g.,2127), to our knowledge, no approach has been designed to assemble and validate perfectly phased sex chromosomes across the immense range of sizes and levels of divergence that exist28. Therefore, we developed a custom pipeline to phase the sex chromosomes that relies on the generation of male-specific k-mers (henceforth, Y-mers) and a haplotype-combined scaffolding approach to identify X- versus Y-linked contigs (Supplementary Note 1).

To examine the structure and gene content of the Cannabaceae sex chromosomes, we built genomes that span the range of XY cytotypes. This includes males of three botanical varieties of hop used across the breeding pedigree29: a line of H. lupulus var. lupulus (USDA “21110M”) and two wild, American relatives, H. lupulus var. lupuloides (“MN-1421”) and var. neomexicanus (“MN-586”). We also built genomes for two male genotypes of cannabis (“Carmagnola” and “Otto II”) and two European fiber-grain genotypes that have derived monoecy from dioecy (“USO 31” and “Futura 75”). We used hifiasm with PacBio HiFi long reads and Omni-C incorporation to build contig assemblies30 (Supplementary Figs. 1, 2, and Supplementary Data 1). Prior to scaffolding, we identified X- and Y-linked contigs and manually placed them into separate haplotypes (Supplementary Data 2 and Supplementary Note 1). After necessary sex chromosome phase correction, in cannabis, the haplotype 1 (HAP1) and 2 (HAP2) assemblies are 760.6–828.8 Mb, while the hop genomes are three times larger at 2.4–2.7 Gb (initial contig N50s: 22–167.6 Mb; Supplementary Data 3). These genomes contain very few gaps (Supplementary Figs. 3 and 4) and exhibit Merqury31 k-mer completeness values of 93–98.3% and QV values of 38.7–58 (Supplementary Data 3). Importantly, the sex chromosomes of these assemblies meet the expected cytological patterns. The Y is the largest chromosome in the cannabis genome at 110.6–114 Mb, while the X is 23% smaller at 84.4–87.4 Mb (Fig. 1). Similar to the increased overall genome size, the X chromosomes of the three varieties of hop are three times larger than in cannabis, at 267.4–272 Mb. The hop Y chromosomes show the most drastic difference overall. The American hop Y chromosomes are cytologically homomorphic, where they are 89.3–92% the size of the X. The Y in H. lupulus var. lupulus is the smallest chromosome in its genome, and the smallest hop Y, at 168.7 Mb, fitting the approximate 1.5:1 ratio observed in the cytotype (62.9% the size of the X)13. These near-perfect, phased sex chromosomes mark a methodological advancement in the field of comparative genomics, enabling the generation of chromosome-scale assemblies from a single heterogametic sex individual without parental trio-binning.

Fig. 1. Genome architecture of Cannabis and Humulus.

Fig. 1

a Identification of the male-specific region of the Y chromosomes (MSY). The heatmap within each ideogram shows coverage of male-specific k-mers (Y-mers) within 1 Mb windows and a 100 kb jump on the X, Y, and Xm (monoecious X) chromosomes. b GENESPACE gene and repeat landscapes, and syntenic relationships, for cannabis and hop. The chromosomes are represented with a single haplotype for the autosomes, but both sex chromosomes. Due to the genome size differences, the hop genome was scaled to cannabis. c GENESPACE gene and repeat landscapes, and syntenic relationships, between the cannabis X and Y chromosomes, and the hop X and Y chromosomes (d). Syntenic relationships are evident in the pseudoautosomal region (PAR) of both species, but very little synteny remains between the MSY and the homologous region of the X (HXR). In the hop PAR (d), the dark gray highlights two inversions based on synteny.

The origin of cannabis and hop sex chromosomes

In an XY system, the pseudoautosomal region (PAR) recombines relatively freely32, while the male-specific region of the Y (MSY) is typically non-recombining and only inherited by males. Understanding the timing of gene capture into the MSY is important, as the sex-determining genes are likely those that ceased recombining first. Reconstructions of sexual systems suggest that dioecy is the ancestral state in Cannabaceae and its sister families Moraceae and Urticaceae that contain figs, mulberries, and nettles33. Example species within these families have been shown to switch sex when ethylene concentration or perception is modified34,35, suggesting the possibility of a shared sex-determining mechanism. Currently, it is unknown if the two families share an ancestral XY pair; however, recent assemblies in Moraceae36,37 and our phased XY chromosomes provide an opportunity to refine our understanding of the timing of the sex chromosome evolution and uncover which genes, or region of the sex chromosomes, have the oldest origin.

We used a rigorous Y-mer mapping and gene tree approach to delineate the boundary between the MSY and the PAR (Supplementary Data 4). We used a combination of RNAseq and protein homology, and annotated an average of 34.9k gene models per haplotype using BRAKER38,39, with BUSCOs ranging between 93–97.2% (Supplementary Data 3 and 5). The hop and cannabis Y chromosomes contain a single PAR with a majority of the sequence contained within the MSY (72.4–88.7%; Supplementary Data 6), also matching cytological observations (Fig. 1)40. We additionally used the gene trees to identify the homologous region of the X (HXR) to the MSY (i.e., the region of the X that does not recombine with the Y in males). Genes on the Cannabaceae sex chromosomes largely maintain a linear order between XY pairs in the PARs, in utter contrast to the MSYs that have very few genes and little synteny remaining (Fig. 1).

To investigate the origin of the sex chromosomes, we used evidence from complementary analyses. Syntenic relationships show the sex chromosomes evolved from the same ancestral autosomes in cannabis and hop, but different chromosomes in mulberries and figs (Fig. 2). Likewise, using gene trees we found support for a shared origin between cannabis and hop, where sex-linkage of a gene prior to the divergence of the genera shows all Y-linked copies forming a clade that is sister to the X-linked clade (Supplementary Fig. 5). No gene tree topology suggests that Moraceace shares this origin. To estimate the age of the MSY, we used synonymous protein changes (Ks) between one-to-one XY orthologs (hereafter, gametologs). Using Ks in hop (max = 0.4; mean top 10% = 0.33; Fig. 2), we estimate the sex chromosomes evolved 38.6–48.3 MYA. We next used a phylogeny based on whole-plastome sequences and fossil records to calibrate and estimate divergence times for comparison across Rosales. Resulting age estimates were transposed onto a species tree, inferred using a multispecies coalescent summary approach41, with 825 nuclear gene trees at concordant nodes (Fig. 2 and Supplementary Fig. 6). The ancestors of cannabis and hop began diverging between 33.9 and 41.4 MYA (95% CI; mean = 36.38), whereas Cannabaceae diverged from Moraceae around 66–84.4 MYA (95% CI; mean = 74.2). These results firmly place the sex chromosome origin on the stem branch leading to the split between cannabis and hop, consistent with our Ks-based age estimate and previous comparative analyses14, but nonetheless support that their shared XY evolution is among the oldest identified in plants, and independent from the species in Moraceae.

Fig. 2. Origin of the cannabis and hop sex chromosomes.

Fig. 2

a GENESPACE synteny plot between hop, cannabis, fig, and mulberry. The chromosomes were ordered based on syntenic relationships to the hop reference. Green highlights synteny from the X chromosomes in hop and cannabis, while purple and salmon highlight synteny from the sex chromosomes of mulberry and fig, respectively. b Synonymous protein changes (Ks) for one-to-one XY gametologs for hop and cannabis highlight the homologous region of the X (HXR) to the male-specific region of the Y (MSY). Ks values above one were removed. Syntenic relationships between the hop and cannabis X chromosomes were identified using GENESPACE. The hop reference is shown in its reverse complement, so that the HXR has the same orientation as cannabis. c Cladogram showing dates as the mean age in millions of years for species across Rosales. The species tree was estimated using a coalescent-based summary approach with ASTRAL based on 825 nuclear genes.

Genes captured into the MSY in the same event are expected to have similar levels of Ks, where the older captures will have higher Ks between gametologs compared to younger captures (i.e., evolutionary strata42). While the max Ks in cannabis (0.81) is twice as high as hop (0.4), Ks is largely consistent when plotted across the X in both genera (Fig. 2), and we observed no evidence for a change point within this region that would indicate strata (Supplementary Fig. 7), suggesting the vast majority of the MSY was captured in one event. In addition to the region of shared capture into the MSY, there is evidence for continual, but independent, gene captures in cannabis and hop. In hop, we found evidence for nearly 50 Y-linked genes that were captured in the American hops, that are clearly in the PAR in H. lupulus var. lupulus (Supplementary Fig. 5). In cannabis, Ks between gametologs shows a clear, continuous pattern of addition from the PAR boundary at 29.3–35.7 Mb, suggesting that genes in this region have been gradually added into the MSY over time, rather than through a large structural change (Fig. 2). Thus we found the Cannabaceae sex chromosomes have two regions of gene capture into the MSY: a region of independent captures near the PAR boundary, and a large, shared region that was captured prior to the divergence of the two genera that likely contains the sex-determining gene.

Y degeneration and faster-X evolution

The sex chromosomes have unique inheritance patterns that, coupled with the suppressed recombination of the MSY, can drive divergent molecular evolutionary dynamics. A classic observation in the MSY of sex chromosomes is degeneration. The suppressed recombination of the MSY reduces the efficacy of natural selection and can drive the accumulation of slightly deleterious mutations43. As a consequence, the MSY tends to accumulate repetitive sequences, which can quickly increase the size of the Y, and over a greater time can lead to gene loss and reduced size44,45. It has, however, been established that sex chromosome degeneration is not linear with time12, and consistent with this, the Cannabaceae sex chromosomes have a shared origin but exhibit extreme cytological variation (Fig. 1). Like the Y, the X chromosome has unusual inheritance patterns. The X spends more evolutionary time in females, but the hemizygous state of the X chromosome in males is predicted to cause elevated rates of molecular evolution, a phenomenon referred to as the faster-X effect that has mostly been investigated in animal systems thus far46 (but see47).

The Cannabaceae sex chromosomes show clear signs of degeneration. In cannabis, which has the largest Y relative to its X counterpart, there is evidence for substantial gene loss (overall MSY models were 46% the number found in the HXR). In H. lupulus var. lupulus, the MSY contains 17% of the total number of genes in the HXR, with similarly low numbers in the Y chromosomes of the hop varieties lupuloides and neomexicanus (both 20%; Supplementary Data 7). One potential cause for the variation in the physical sizes of the Y chromosomes is through the differential loss of genes from the MSY, where the smaller Y is expected to have fewer overall genes and, more conservatively, fewer detectable 1:1 gametologs. However, in hop, which shows the most extreme variation in size within a species, the number of XY gametologs are nearly the same, suggesting instead that the difference in size of the Y is driven by transposon accumulation. To test this, we annotated transposable elements (TE) using EDTA48 and found the cannabis and hop genomes to be on average 74.3% and 86.8% repetitive, respectively (Supplementary Data 5). TE content sharply increases at the beginning of the MSY (Fig. 1) and is 15.8% and 6.8% greater compared with the genome-wide averages for cannabis and hop, and 13.3% and 5.3% compared to the HXRs, respectively (Supplementary Data 8). Recent bursts in TE activity are evident on the Y chromosomes compared to their X counterparts, consistent with the increased TE abundance (Fig. 3). Importantly, the H. lupulus var. lupulus Y shows reduced TE expansion relative to the other hop Y chromosomes, which could explain its smaller size. Together, these results support that the MSYs are highly degenerated and that the cytological variation in Y chromosome sizes is a complex, ongoing interplay between gene loss and repeat expansion.

Fig. 3. Molecular evolution of cannabis and hop XY chromosomes.

Fig. 3

a Kimura substitution plot showing the repeat landscapes of the sex chromosomes in two hop genotypes. Values closer to 0 represent more recent events, and higher values approaching 50 represent older events. Time was approximated by dividing the Kimura distance by twice the mutation rate (secondary x-axis), and the black lines indicate the estimated time of the initial sex chromosome evolution. b Boxplot of the ratio of nonsynonymous to synonymous protein changes (dN/dS). The boxplots represent the 25th, 50th (median), and 75th percentiles of values, while the top and bottom whiskers extend to 1.5 times the interquartile range. Asterisks indicate the level of significance from a Wilcoxon rank sum test, with a Benjamini–Hochberg correction for multiple tests, where “***” indicates p ≤ 0.001, “*” indicates p ≤ 0.05, and “ns” indicates not significant. Values above a dN/dS of three were removed for the comparisons between the cannabis homologous region of the X (HXR; n = 597), pseudoautosomal region (PAR; n = 782), and autosomes (n = 8671), and between the hop HXR (n = 497), PAR (n = 226), and autosomes (n = 7932). c Fst outliers between high-THC and high-CBD genotypes in cannabis. Fst was calculated within 100 kb windows and with a 10 kb jump. Outliers were defined as values that were three standard deviations above the mean (0.206), and are indicated with black points. Locations for the FLOWERING LOCUS T (FT), CONSTANS (CO), GIGANTEA (GI), and prenyltransferases (PT) in the PAR are indicated.

The X chromosomes in cannabis and hop also show interesting patterns of molecular evolution. We used codeml in PAML49 to calculate dN/dS ratios on 10,056 genes genome-wide. In cannabis, the HXR (n = 597) has significantly higher dN/dS overall than the autosomes (n = 8671; Wilcoxon rank sum test with Benjamini and Hochberg correction, p = 8.3–08) and PAR (n = 782; p = 8.2−08) and in hop, we find the same pattern with the HXR (n = 497) showing significantly greater dN/dS than the autosomes (n = 7932; p = 0.00018) and PAR (n = 226; p = 0.02; Fig. 3). This supports that in both genera the X chromosomes show signatures of the faster-X effect. Given the ongoing evidence that sex chromosomes play a critical role in speciation50, these molecular patterns could contribute to difficulty with making wide crosses (e.g., between distinct botanical varieties).

Key floral genes on the sex chromosomes

Flowering time is one of the several critical differences in the development of the sexes in cannabis and hop, wherein males flower earlier than females13,51. The overall inflorescences are also architecturally distinct between the sexes. The male inflorescences consist of clusters of flowers (cymose panicles), while the female inflorescences are compound racemes13. Within the female inflorescences are the glandular trichomes that produce the majority of the secondary metabolites, many of which are prenylated compounds (e.g., bitter acids and cannabinoids). It is currently unknown how much the sex chromosomes contribute to these sexually dimorphic traits.

Understanding flowering time is a major goal of many breeding programs in order to develop cultivars suitable for different latitudes and to control flowering between the sexes. In cannabis, the appearance of male flowering typically occurs after producing four true leaf pairs (L4), 10–14 days prior to the appearance of female inflorescences (at nine true leaf pairs, L952). The photoperiod-dependent pathway for flowering includes genes, such as FLOWERING LOCUS T (FT), CONSTANS (CO), and GIGANTEA (GI), and their conservation across plants could suggest similar functions in Cannabaceae53,54. We found homologs for these three genes on the XYs of cannabis and hop (Supplementary Data 9 and 10). Consistent with these findings, GWAS and QTL analyses in cannabis identified peaks for flowering time across many loci, including on the X chromosome55. FT in particular stands out as a potential candidate for the difference in flowering time between the sexes because, in addition to a PAR homolog and two autosomal homologs, there are additional copies that are found within the MSY and HXR. However, two observations might rule out FT. First, the origin of the duplicates to the MSY are recent and do not predate the divergence with hop, therefore they are not likely to be driving the sex-specific differences in flowering given males also flower earlier in hop (Supplementary Fig. 8). The second observation is related to gene expression of FT, which is expressed in leaf phloem cells and transported to the shoot apical meristem to initiate floral transition56. Using RNAseq data of leaf tissue before and after flowering in cannabis57, the PAR FT is significantly differentially expressed, as well as the X-linked homolog in the HXR and the autosomal homolog on Chr07 (Supplementary Data 11). However, none of the FT genes are significantly differentially expressed in the leaf tissues when contrasted between the sexes at flowering. Another strong candidate is FLOWERING LOCUS D (FD), which interacts with FT to promote flowering58. Gene tree topology supports that FD became sex-linked prior to the split between cannabis and hop (Supplementary Fig. 8). Using RNAseq of apical meristems during flowering (L9 in female, L4 in male)59, where FD is primarily expressed, the X copy is not significantly different between the sexes (adjusted p = 0.53). However, male expression is ten times higher on the Y-linked ortholog (Supplementary Data 11). These results could suggest that FD is involved with the earlier timing of floral initiation found in male plants. Many additional floral regulators are located on the XYs, and a fine-stage developmental time series of gene expression, coupled with functional analyses, will help to better disentangle the roles of these genes.

The female inflorescences of hop and cannabis are economically important, since they produce high densities of glandular trichomes. In hop, the trichomes (called lupulin glands) produce the bitter acids, commonly referred to as alpha and beta acids, and essential oils used in brewing beer. Similarly, in cannabis, the female inflorescences produce the majority of the cannabinoids, such as cannabidiol (CBD) and tetrahydrocannabinol (THC). A critical step in the bitter acid and cannabinoid pathways is prenylation via aromatic prenyltransferases (PT)60,61. In hop, two PT genes have been identified62 and are a tandem duplication located in the PAR of the XY chromosomes. The hop PT1 has three to six homologs in cannabis, and several other PT genes are found in cannabis that vary both in function and copy number62, but are also located in the PAR. Bitter acids and cannabinoids are major targets of breeding programs, so we tested for signs of selection among Cannabaceae PT genes. We first used an Fst outlier approach using HiFi data for 49 female cannabis genotypes63,64, where we contrasted high-THC (n = 30) and high-CBD (n = 19) genotypes. Our analysis showed a local peak of higher Fst surrounding the PT genes, which could in part be driven by fiber hemp samples in the high-CBD class being bred for lower cannabinoid production (Fig. 3). While Cannabis is monotypic, Humulus contains two species, which gives us the power to test for selection using the McDonald-Kreitman test (MKT). Using WGS for 23 hops, we found PT1 and PT2 show signatures of adaptive variation (alpha = 0.76 and 0.1, respectively), of which PT1 shows the strongest signal under Fisher’s exact test (p = 0.03, FDR = not significant). Given that PT genes prenylate CBD, THC, and the bitter acids, robust comparisons that contrast the abundance of metabolite production may be more revealing.

A curious pattern is evident when contrasting cannabis and hop for these genes. Cannabis has evolved many more copies of PT, with diverse functions62, and shows copy-number variation among genotypes, while hop is more static. Moreover, the gradual expansion of the MSY in cannabis leaves the closest PT gene ~1.1 Mb away from the boundary. This gradual expansion of the MSY is potentially related to centromere proximity (Supplementary Fig. 9) as the existing reduced recombination could facilitate sex-linkage65,66. However, given that females are strongly selected in cannabis, instead the MSY boundary may have expanded due to sexually antagonistic selection32,67. Indeed, recent analyses show the boundary of the MSY is variable in cannabis, with some genotypes supporting this boundary expanding further into the PAR63, closer to the PT and FT genes. The MSY expansion in the wild hops is also towards the PT genes, but a notable, different pattern is found in the H. lupulus var. lupulus “21110M” genotype. In the PAR, there are two large inversions (each ~9 Mb) and a translocation (~4.8 Mb) (Fig. 1), suggesting only ~⅓ of the PAR in this genotype can freely recombine; these SVs co-locate at the PAR/MSY boundary, which may lead to the eventual expansion or contraction of the MSY over time. FT is also located in the inversion with PT in the “21110M” genotype, which could have implications for breeding that involves these two targets and other linked genes. It is unclear, based on the genome references available, whether the PT genes in hop and cannabis originated on the sex chromosomes or if they perhaps translocated there. Nevertheless, given that the PAR is physically linked to the Y (or X), the selection on females for products that involve PT could be driving some of these patterns of expansion of the MSY or structural variation at the boundary with the PAR. Altogether, these results suggest the sex chromosomes played a role in the domestication of these species, but it is also possible that the speciation process and subsequent adaptation has also helped to shape the sex chromosomes.

Evidence for a background, not active, Y system

It is well established that ethylene plays a demonstrable role in sex determination in cannabis and hop. Ethephon treatment, when applied to XY plants, which releases ethylene, will initiate female inflorescence development, rather than male35,68. The opposite pattern has also been demonstrated, wherein ethylene inhibitors promote male development on XX plants69. In addition to likely involving ethylene, the sex determination genes are also expected to act early in flower development, because the dimorphic architectures between the sexes are apparent shortly after floral meristem initiation13 and neither sex produces vestigial structures of the other sex52. Analyses to determine the sex-determining genes59,7072 have made promising discoveries, but so far, no clear candidate has been identified.

Given the shared origin of sex chromosomes, we searched for highly conserved genes across the cannabis and hop MSYs. Sixty-eight genes are shared on the Y chromosomes between the genera, 21 of which are suspected to be involved with floral development based on orthologous functions (Supplementary Data 10). None of these appear to be obvious sex-determination genes based on their developmental timing and function. For instance, the maintenance of SIGNAL PEPTIDE PEPTIDASE (AtSPP) and ROP ENHANCER 1 (REN1), which in Arabidopsis are involved in pollen germination73 and pollen tube growth74, respectively, suggests a role in male fertility in Cannabaceae, but not overall stamen development. Moreover, when contrasting male gene expression at the L2 stage (pre-flowering) to the L4 stage (flowering), only three MSY genes were significantly differentially expressed, but none of these are shared with hop nor have clear roles in reproductive organ development (Supplementary Data 12). The lack of an obvious sex-determining gene on the Y tracks with crossing experiments using polyploids that vary in X copy number. In an active-Y sex determination system, inheritance of the Y is expected to lead to male individuals, regardless of the number of X copies75. In hop and cannabis diploids, XY individuals typically develop strictly male inflorescences, whereas polyploid individuals with XXY or XXXY karyotypes commonly produce female flowers or are monoecious76,77. These lines of evidence suggest that the sex-determining gene instead resides on the X. The maintained pollen and flowering timing genes, like FD, in the MSY suggest that the Y is not inactive, but plays other critical fertility roles, which we term a “background-Y”.

A candidate X-linked sex-determination gene

Natural variation in sexual systems can be used to overcome limitations in functional genomics in order to identify sex-determining genes78. Monoecy, wherein a plant can produce both functional female and male flowers, has been observed in XX and XY diploids in cannabis7,8. While sex-expression can be “leaky” in typical cannabis and hop plants, these monoecious genotypes reliably produce fully-functional inflorescences for both sexes. One possible route to transitioning from dioecy to monoecy in an otherwise XX karyotype individual involves the movement of the sequence from the MSY to the X chromosome. To test this hypothesis, we examined the genomes for the monoecious “Futura 75” and “Uso 31” genotypes. There are no signatures of the Y chromosome on the Xs of these two monoecious genotypes (henceforth denoted as Xm), including no dense peaks of Y-mers (Fig. 1 and Supplementary Data 2), and assessing the gene trees showed no evidence of a topology suggesting an MSY gene translocated to the X. These results suggest that in these individuals, monoecy evolved through a modification of a gene on the X chromosome, likely involving the sex-determining gene. Moreover, these results highlight that fully-functional male inflorescences can develop in the absence of a Y chromosome, consistent with the background-Y hypothesis.

To uncover the genetic basis for monoecy in cannabis, we generated a reciprocal backcross (BC1) between a female and a monoecious line and genotyped 950 BC1 offspring with low-pass Illumina shotgun sequencing (n = 887 genotypes after processing; see “Methods”). We mapped reads to a pangenome graph using the two male and two monoecious cannabis references and used KhufuPAN79 to identify loci associated with monoecy. All 19 hits are located on the X chromosome, with 73.7% of these landing between 77.6 and 81.7 Mb (Fig. 4). We generated PacBio HiFi data for 10 additional monoecious genotypes and used existing HiFi reads from 51 cannabis genotypes63,64 outside of the BC1 and performed an Fst outlier analysis (n = 49 female, n = 14 monoecious). 66.8% of the Fst outliers are found on the X chromosome, and particularly surrounding the monoecy locus identified with the BC1 (Fig. 4), strongly suggesting this is the locus associated with the monoecious trait.

Fig. 4. Identifying the sex-determining locus in cannabis.

Fig. 4

a Karyoplot showing Khufu hits from a backcross segregating for monoecy and Fst outliers between female and monoecious cannabis genotypes. Fst was calculated within 100 kb windows and with a 10 kb jump. Outliers were defined as values that were three standard deviations above the mean (0.438), and are indicated with black points. b The region identified as the monoecy locus is highlighted in light gray, with the same Khufu hits and Fst outliers as in (a). Also shown are the locations of genes that are significantly differentially expressed between the sexes in cannabis during flowering (adjusted p-value < 0.05). The positive axis, highlighted in green, are genes with higher expression in females, while the negative axis, and in yellow, are higher in males. Homologs of two ethylene biosynthesis genes are located in this locus, ACC synthase (ACS) and ACC oxidase (ACO). The ETHYLENE RESPONSE (ETR) gene shown was recently proposed as a candidate for sex-determination, but is outside the monoecy locus. The reproductive meristem (REM) and KANADI (KAN) genes are located in the monoecy locus, but these, and ETR, are not significantly differentially expressed in this contrast.

We next analyzed gene expression of apical meristems during flowering between the sexes (L9 in female, L4 in male), to narrow in on candidate sex-determining genes. Forty-six genes were significantly differentially expressed in the monoecy region (Fig. 4); 11 are related to either flowering or ethylene (Supplementary Data 13). None of the floral genes are obvious candidates for sex-determination based on their orthologous function; however, because of the demonstrated role of ethylene in promoting female inflorescence development, the ACO and ACS genes are notable candidates. The ethylene biosynthesis pathway involves two enzymatic steps. First, S-adenosyl-L-methionine is converted into 1-aminocyclopropane-1-carboxylic acid (ACC) by ACC synthase (ACS), then ACC is converted to ethylene by ACC oxidase (ACO)8084. The ACO gene located in the monoecy locus is orthologous to ACO5 (Supplementary Fig. 8), and expression of this gene was greater in males, which is counter to the expectation that female flowers are initiated with ethylene perception (Supplementary Data 13). Moreover, the ACO expression is most abundant in root tissue in cannabis (Supplementary Fig. 10). The ACS gene, however, is in a clade with other type II ACS genes (Supplementary Fig. 8), the expression was significantly greater in females than males (Fig. 4), and was exclusively expressed in females and monoecious plants within floral tissues and stems during flowering, with one weakly expressed male floral sample being the exception (Supplementary Fig. 10). These patterns are consistent with the ACS genes that have been demonstrated to determine sex in cucurbits85,86.

Other sex-determination gene candidates for cannabis and hop have recently been proposed71,72. The reproductive meristem (REM) and KANADI (KAN) genes in Toscani et al.72 fall within the monoecy locus proposed here, but the ethylene-response (ETR) gene in Akagi et al.71 falls well outside of the monoecy locus (Fig. 4). Importantly, the expression of ETR, REM, and KAN are not significantly different between the sexes in apical meristems during flowering, (Fig. 4) and they are ubiquitously expressed across analyzed tissue types rather than specifically in floral tissues (Supplementary Fig. 10). Therefore, these genes do not follow the patterns expected for sex-determination; that is, there is not a narrow window of expression before or during floral meristem initiation when the sexes developmentally differentiate13. These are instead patterns we have found with the ACS gene.

Additionally, in some cannabis crosses, monoecy is shown as a recessive trait7,87. We searched for variants that were homozygous in monoecious genotypes that could be causative. We used a combination of Sawfish, DeepVariant, and mpileup to call SNPs, INDELS, and structural variants (SV). No large SVs were identified within the ACO5 and ACS gene models, and no SNP or INDEL segregated in ACO5 for monoecy. However, we found two INDELS, each three base pairs in length, in exon 4 of ACS just outside an aminotransferase domain, that are homozygous in 13/14 monoecious genotypes (one “Futura 75” was heterozygous, but another “Futura 75” was homozygous). Of the 49 female genotypes scanned, only two had the INDELs, but they were both heterozygous. These results lead us to the hypothesis that sex is determined in cannabis and hop using putatively the same ethylene-based sex determination pathway that involves ACS as in cucurbits, which are more than 100 million years diverged.

These results parallel and differ from other sex-determining systems in important ways. While some plants have support for a single gene that determines sex88, others are two-gene systems, where one gene impacts female and another male development89. Our ACS results suggest a single-gene sex-determination system, which in plants is theorized to be a viable pathway to dioecy when the ancestral state is monoecy90. And while cannabis and hop have an XY system that mirrors the mammalian XYs in some aspects, like the substantial degeneration of the Y, the X-linked basis for sex-determining more closely resembles avian systems. In chickens, females are the heterogametic sex (i.e., a ZW sex chromosome system) and inherit the sex-specific chromosome. However, the sex-determining gene, DMRT1, resides on the Z chromosome91. Like DMRT1, ACS in hop and cannabis could function by a dose-related mechanism, in which a threshold amount of ACS expression is needed to produce females. Altogether, while cannabis and hop were among the first sex chromosomes to be identified in flowering plants, these results highlight that there is more these sex chromosomes can help us to better understand about the evolution of these complex regions of a genome.

Discussion

Understanding the genetic basis for sex has been a topic of interest for over a century. Cytologists, like Nettie Stevens, uncovered the inheritance of the heteromorphic sex chromosome that determined maleness9. While both cannabis and hop sex chromosomes were identified a century ago due to their heteromorphic XYs, uncovering the genes that determine sex has remained elusive. Our near-complete genome assemblies have shown the evolution of sex chromosomes in cannabis and hop in unprecedented detail. The ancient MSY contains male-fertility genes, but given the lack of clear genes associated with sex determination, it leads us to hypothesize that it is a background-Y rather than an active-Y system. The X, however, has demonstrably been shown to be involved with sex determination through crosses and the reversion to monoecy. Here, we identified a candidate gene in the ethylene pathway that promotes female inflorescence development, ACS. While some of these results are in line with patterns found in mammals, like gene loss in the MSY and faster-X, an open question remains on how background-Y and an X chromosome that contains the sex determining gene may evolve differently. Altogether, the diverse sex chromosomes and their associated reproductive genes in both cannabis and hop played a pivotal role in the domestication, breeding, and ongoing production of these species.

Methods

Genome assemblies

All sequencing for this study was performed at HudsonAlpha Institute for Biotechnology (Huntsville, AL, USA), unless otherwise noted. To obtain high molecular weight (HMW) DNA for the genome assemblies, we flash-froze young leaf tissue in liquid nitrogen, then extracted the HMW genomic DNA using a Takara NucleoBond HMW kit (San Jose, CA). Libraries were constructed using a SMRTbell Template Prep Kit 3.0, sheared using a Megaruptor (Diagenode), and sized on a Blue Pippin instrument (Sage Science, Beverly, MA, USA). Long-read sequencing was performed using circular consensus sequencing (CCS) mode on a PacBio Sequel II (Pacific Biosciences of California). Sequencing was performed using a 30 h movie time with 2 h pre-extension, and the resulting raw data were processed using the CCS6.2 algorithm. Next, we generated Omni-C data to phase the haplotypes during the assembly process. On flash frozen young leaf tissue, libraries were constructed using standard protocols (Cantata Bio, Dovetail Omni-C Kit Catalog #21005) using the Dovetail Omni-C protocol V2 and were run on an Illumina NovaSeq 6000 using paired-end reads and a read length of 150 base pairs.

For cannabis, HiFi reads were assembled into contigs using hifiasm v0.16.030, invoking the mode that incorporates Omni-C reads. For hop, we instead used hifiasm v0.19.5 (with Omni-C incorporation). For “21110M”, we also changed the “-s” parameter from 0.55 (default) to 0.2 to better balance the haplotypes out of the assembly. Across all assemblies, we polished our contigs with the HiFi data using Racon v1.5.092, removed contigs below 50 Kb using BBTools v38.18-0 (https://sourceforge.net/projects/bbmap/), and removed contigs identified as contamination by FCS-GX v0.4.093.

Prior to scaffolding, we manually placed contigs from each of the sex chromosomes into a single haplotype. To accomplish this, we first identified male-specific k-mers (Y-mers). We generated Illumina whole-genome sequencing (WGS) for 22 H. lupulus individuals (Supplementary Data 14). DNA was extracted using a CTAB approach on flash-frozen leaf tissue, and sequencing libraries were constructed using an Illumina TruSeq DNA PCR-free library kit (Catalog #20015962) using standard protocols. Libraries were sequenced on an Illumina NovaSeq 6000 instrument using paired ends and a read length of 150 base pairs. For two isolates (“64035 M” and “Bullion”), DNA was instead extracted using the BioEcho EchoLUTION Plant DNA 96 Kit (product no. 010-103-002), and sequencing libraries were constructed using a seqWell purePlex DNA Library Preparation Kit. Sequencing was performed on an Illumina NovaSeq X Plus instrument using paired-end reads and a read length of 150 base pairs at NovoGene (Beijing, China). For hop, we built the Y-mer list using 12 of the WGS genotypes and the HiFi for two of the genome assembly genotypes (H. lupulus var. lupulus “21110M” and H. lupulus var. lupuloides “MN-1421”). For cannabis, we used 14 existing WGS libraries (Supplementary Data 14). All paired-end Illumina data had adapters removed and were quality filtered using TRIMMOMATIC v0.3994 with leading and trailing values of three, a sliding window of 30, a jump of 10, and a minimum remaining read length of 40. We next found all canonical 21-mers in each isolate using Jellyfish v2.3.095 and used the bash comm command to identify a list of k-mers found across all males, but never found in any female, for hop and cannabis separately. We mapped the Y-mers to both haplotypes of each assembly using BWA-MEM v0.7.17, with parameters -k 21 -T 21 -a -c 10, and generated a table of coverage per contig using SAMtools v1.10 coverage96 (Supplementary Data 2). Additional details for manually phasing the sex chromosomes are described in the Supplementary Note 1.

The sex chromosome-corrected assembly was then scaffolded by first mapping the Omni-C reads using BWA-MEM, with parameters -5SP and samtools v1.10 with parameters view -S -h -b -F 2316. Next, we scaffolded the assembly using Yet Another Hi-C Scaffolding tool (YaHS) v1.1 with default parameters97. We edited contact maps manually with Juicebox Assembly Tools v2.1598 to produce the expected 10 chromosomes per haplotype. Scaffolds containing assembled telomeres were checked for correct orientation to the terminal ends by searching for the (TTTAGGG)n repeat using GENESPACE v1.3.199. The cannabis assemblies were reordered and oriented to the C. sativa “cs10” reference100, and the hop was reordered and oriented to H. lupulus var. lupulus “Cascade”18. Final plots showing the number of contigs per chromosome and telomeric regions (density of the window for the telomeric repeat set to 0.75) were generated using GENESPACE. We assessed the quality of the genome assemblies using compleasm v0.2.6101 using the Eudicots odb10 database and Merqury v1.331. To delineate the boundary of the MSY with the PAR, we remapped the Y-mers to the final assembly using the approach described above. We coupled the MSY/PAR boundary call based on Y-mer coverage with gene trees (described below).

Gene annotations

To generate gene model annotations, we first annotated and masked repetitive sequences. For cannabis, this was accomplished by using the existing repeat library produced in Lynch et al.63. Repeats were then masked with RepeatMasker v4.1.2. We annotated gene models using BRAKER v2.1.6 and TSEBRA102, using RNAseq and Full-length Oxford Nanopore cDNA (Supplementary Data 14). The short-read RNA-seq libraries were first trimmed with fastp103 and then aligned with hisat2 v2.2.1104. The full-length cDNA were instead mapped with minimap2 v2.24105. Protein evidence was also included from Arabidopsis thaliana, Theobroma cacao, Glycine max, Rhamnella rubrinervis, Ziziphus jujuba, Trema orientale, Vitis vinifera, Prunus persica, Morus notabilis, C. sativa, and H. lupulus (following the same protocol from Lynch et al.63).

For the hop assemblies, we annotated repeats by building a de novo library for each haplotype separately using RepeatModeler v2.0.5106, invoking the -LTRStruct parameter. We then masked repeats using RepeatMasker v4.1.5107. We annotated gene models with BRAKER3 v3.0.639 using both proteins and RNAseq data as extrinsic evidence. To generate RNAseq data, we flash-froze leaf tissue and inflorescences for H. lupulus var. lupulus “21110M” and “Comet male”, respectively. We extracted RNA using a CTAB approach108. Sequencing libraries were constructed using an Illumina TruSeq Stranded mRNA Library Prep Kit (Catalog #20020595) using standard protocols and were sequenced on an Illumina NovaSeq 6000 instrument using paired-ends and a read length of 150 bp. We additionally used existing RNAseq data for different tissue types (Supplementary Data 14). The RNAseq data were trimmed as described above and were mapped using the defaults in BRAKER3 when supplying fastq files as the input. For the protein support, we used the OrthoDB v11109 for Viridiplantae, combined with proteins from genome references for Parasponia110, Trema110, H. lupulus “Cascade”, and C. sativa “cs10”. We assessed the gene annotations using compleasm v0.2.6101 using the Eudicots odb10 database.

Repeat analyses

To examine transposable elements (TEs), we additionally annotated repeats using EDTA v2.0.048 in the sensitive mode for each haplotype separately. To compare repeat landscapes across genotypes, we generated a panEDTA library using EDTA v2.2.0111. We next used this pan-library to identify the repeats on the sex chromosomes using RepeatMasker, and Kimura substitution values were calculated using the script createRepeatLandscape.pl.

To identify putative centromere locations, we used StainedGlass v0.5112 to visualize the massive tandem repeat arrays. To identify tandem repeats, we used Tandem Repeats Finder v4.09.1113 (parameters 2 7 7 80 10 50 500 -f -d -m -h), and to identify CRM elements, we used BLAST v2.15.0114 using previously identified elements115. To plot the landscape of repeats and genes in the context of the putative centromere locations, we used sliding windows of 100 and 200 Kb for the Y and X, respectively.

Synteny analyses

For cannabis and hop, visualization and analysis of local synteny was conducted with DEEPSPACE, which aligns genomic windows between assemblies and uses GENESPACE machinery to infer syntenic blocks. Source code for the riparian plots was modified where needed to visually compare two genomes with such different total sizes. To examine syntenic relationships between Cannabaceae and Moraceae, we used GENESPACE using default parameters, using the H. lupulus var. lupulus “21110M” HAP2, C. sativa “Otto II” HAP2, Morus alba116, and Ficus hispida117.

Gene tree analyses

To compare the boundary of the MSY with the PAR when using Y-mer mapping, and examine the timing of gene capture into the MSY, we built 2080 gene trees that contained a Y-linked gene model. We first ran OrthoFinder v2.5.2118,119 in ultra-sensitive mode, with the cannabis and hop proteins and 43 additional angiosperms (Supplementary Data 15) to identify OrthoGroups. We aligned the genes within OrthoGroups using MAFFT v7.471120 with the parameter maxiterate set to 1000 and using genafpair. We built gene trees using RAxML v8.2.12121 with 100 bootstrap replicates and invoking the model PROTGAMMAWAG. We visually examined the topology to assess whether sex-linkage of the gene occurred prior to, or after, the divergence of the genera. To examine phylogenetic evidence of monoecy evolving from a Y-linked gene translocating to the X, we used the ETE3 v3.1.3122 check_monophyly function. We narrowed this analysis to trees that contained genes from the MSY in “Otto II” and “Carmagnola”, an X-linked gene in either haplotype of “Futura 75” and “Uso 31”, and a homolog on the “Otto II” and “Carmagnola” X chromosomes, leaving 110 trees. This OrthoFinder run was also used to examine gene homology, and due to copy number variation, gene numbers given in the main text are from the C. sativa “Otto II” reference.

Ks-based analyses and age estimation for the sex chromosomes

To identify one-to-one XY orthologs on the sex chromosomes (gametologs), we ran OrthoFinder v2.5.2 within a reference’s haplotypes. We aligned coding sequences of the gametologs using MAFFT and calculated synonymous (Ks) changes in codons using Ka/Ks calculator v2.0123. To test for evidence of multiple capture events into the MSY (i.e., evolutionary strata), we used the R package mcp v0.3.4124 on Ks values. In cannabis, we excluded one outlier (Ks of 3.7) identified using PMCMRplus v1.9.12125, which runs Rosner’s generalized extreme studentized deviate many-outlier test126. For mcp, we used the model, “y ~ 1, ~1” to identify the change point between two plateaus, and we used 10,000 iterations, 3 chains, and a burn-in of 10,000 (i.e., “adapt”). We also tested the model “y ~ 1, ~ 1, ~1” and “y ~ 1, ~ 1, ~1, ~1” to search for a third and fourth plateau, respectively.

To estimate the age of the sex chromosomes, we used the following formula and values from hop: time (T) = (Ks/π − 1) × (2Netg). For Ks, we used both the max Ks (0.402), as well as the average of the top 10% of Ks values (0.325). For pi (π), we used 0.016, which we calculated by first using BWA-MEM v0.7.17127 to map the WGS data to the H. lupulus var. lupulus “21110M” HAP1 reference, where we removed the Y chromosome, but included the X chromosome from HAP2. We used bcftools v1.9 mpileup and call128 functions to call variants, where we applied the -G parameter for call. We filtered the VCF file using “QUAL > 20 & DP > 5 & MQ > 30”, minor allele frequency of 0.05, and dropped sites with > 25% missing data. We identified SNPs using all 22 H. lupulus genotypes generated here, and one existing H. scandens genotype; however, we restricted the estimation of π to wild-collected H. lupulus hop genotypes (n = 7). π was estimated using Pixy v2.0.0.beta8129 within 10,000-bp windows, and the average taken across the autosomes only. To estimate the effective population size (Ne), we used the following formula, where Ne = π/4μ. As a mutation rate has not been estimated in hop, we assumed a μ of 8 × 10−9. Finally, we used a generation time (tg) of two.

Species tree inference

We used OrthoFinder v3.0.1118,119 to identify conserved, single-copy genes for species tree inference with weighted ASTRAL (wASTRAL) v1.22.3.741,130, which employs a coalescent-based summary method that is statistically consistent under the multi-species coalescent model. To identify orthologs conserved as single copy across the Rosales order, we used gene annotations from the following Rosales taxa: H. lupulus var. lupuloides, H. lupulus var. lupulus, H. lupulus var. neomexicanus, C. sativa, Ficus hispida36, Malus domestica131, Morus alba116, Pyrus commune132, and Ziziphus jujuba133. All single-copy genes conserved across those samples were used for gene and species tree inference, after removing all orthologs of sex chromosome-linked genes (found in Cannabaceae species), since any of those may exhibit unique evolutionary histories in different lineages. Coalescent-based, rather than concatenation, methods were used to analyze nuclear genes because concatenation implicitly assumes an absence of recombination and ILS, or incomplete lineage sorting (i.e., a single history is shared amongst all genes), and can result in statistical inconsistencies in multi-locus datasets with high levels of gene tree discordance134,135; whereas coalescent approaches assume free recombination between genes, accounting for gene tree–species tree discordance due to ILS.

Along with our Humulus dataset, publicly available Illumina short reads were downloaded from NCBI SRA for additional lineages spanning the Rosales (Supplementary Data 16), to produce gene assemblies for phylogenomic analysis. We employed HybPiper v2.3.3136 to (i) map reads to target nucleotide sequences with BWA-MEM, (ii) assemble mapped reads into contiguous sequences with SPAdes v3.15.2137, (iii) extend gene assemblies into flanking intronic regions, and (iv) test for paralogs with default HybPiper parameters. The target sequence file used for read mapping with HybPiper was manually curated with nucleotide coding sequences for 825 one-to-one orthologs from genome annotations for: (1) H. lupulus var. neomexicanus, (2) M. domestica, (3) C. sativa, (4) M. alba, and (5) F. hispida.

Multiple sequence alignments for each target gene were produced using MAFFT v7.520 with the --auto option, which automatically selects the best alignment strategy based on the data120. Poorly-aligned sequences were trimmed using trimAl v1.4.1 with the flag -automated1138. Gene tree inference was performed using maximum likelihood (ML) with IQ-TREE v1.6.2139, 1000 ultrafast bootstraps (BS)140, and model selection based on Bayesian information criterion with ModelFinder141. The resulting species tree was rooted with Rosaceae taxa (i.e., P. communis, M. domestica).

Plastome divergence time estimations

Whole (circularized) plastome assemblies were used to estimate relationships and divergence times among the same, or similar, Rosales taxa included in the species tree analysis (Supplementary Data 16). To assemble chloroplast genomes, we used oatk v0.2142 when HiFi data was available. For all others, where we used Novoplasty v4.3.5143. Limited taxon sampling was used to optimize time calibrations across the Rosales tree and obtain robust divergence time estimates for ingroup branches (i.e., Cannabaceae and Humulus). Whole plastomes were aligned using MAFFT v7.520 with the --auto option and used as input for ML tree inference with IQ-TREE v1.6.2 (as performed for individual nuclear gene trees) and Bayesian divergence time estimation with BEAST v2.7.7144. BEAST was executed using a GTR + G + I substitution and site heterogeneity model with four rate categories; a fossilized birth-death (FBD) tree prior to model speciation and extinction145,146; and optimized relaxed clock model with a lognormal distribution147. The plastome tree topology was constrained according to the ML tree from IQ-TREE and rooted with Rosaceae taxa. Two independent BEAST runs, totaling 48,858,000 MCMC generations, were combined to obtain >200 effective sample size (ESS) for all priors, with MCMC sampling every 1000 generations and 10% burn-in. ESS was assessed with Tracer v1.7.2148. To calibrate the plastome tree to time (i.e., million years ago = Ma) and estimate the tree root height, we estimated the FBD origin under a normal distribution with a mean of 107 Ma and a liberal standard deviation of 4.0. Those parameters virtually constrained the total tree height to approximately 100–114 Ma, which is in accordance with crown group age estimates for Rosales in previous studies33,149. Under the FBD model, two fossil calibrations were applied to further calibrate the plastome tree to time: the first applied a 66 Ma minimum age for the Cannabaceae stem group33,150,151, and the second applied a 33.9 Ma minimum age for the stem group of Humulus33,152.

Tests for selection

To examine protein evolution on the X chromosomes, and test for evidence of faster-X evolution, we first identified OrthoGroups using OrthoFinder v2.5.2 with HAP2 of H. lupulus var. lupulus “21110M”, H. lupulus var. neomexicanus “MN-586”, and C. sativa “Otto II”, as well as Morus notibilis37, and Trema orientale110. We found 10,056 single-copy OrthoGroups where all five genomes were represented. To calculate the ratio of nonsynonymous to synonymous changes in proteins (dN/dS), we used codeml in PAML v4.10.749. To generate the alignments required as input to PAML, we first aligned the genes using MAFFT (same parameters as above), then we converted the alignments to coding sequences using pal2nal v14153. For the input trees for PAML, we used the gene trees generated from OrthoFinder. We assessed significance between autosomal, PAR, and HXR genes, where we excluded dN/dS values greater than three for hop or cannabis (1401 and six genes, respectively) by converting them to NA, and ran a Wilcoxon rank sum test with a Benjamini–Hochberg correction for multiple tests.

To calculate Fst and identify outlier regions between high-THC and high-CBD cannabis, and between female and monoecious genotypes, we generated PacBio HiFi long-read sequencing for 12 additional monoecious genotypes. For two samples (2023_mono1 and 2023_mono2), the methods for generating HiFi were the same as described above. For the 10 remaining genotypes, seeds were sown in greenhouse conditions before being transplanted to an outdoor (in-ground) plot in Geneva, New York (USA). Young leaf tissue was collected and flash frozen in liquid nitrogen. We extracted HMW DNA, prepared and pooled PacBio HiFi libraries, and sequenced using the PacBio Revio platform as described by ref. 154. Plant sex phenotypes were recorded for each individual on a weekly basis throughout the growing season. We additionally used 51 existing HiFi datasets (Supplementary Data 14)63. We aligned the reads to the C. sativa “Otto II” HAP2 reference using pbmm2 v.1.17.0 (https://github.com/PacificBiosciences/pbmm2). To generate an all-sites VCF, we called SNPs using bcftools mpileup and call. We filtered the vcf file using “QUAL > 20 & DP > 3 & MQ > 30”, a minor allele frequency of 0.05, and dropped sites with > 25% missing data. We used Pixy v2.0.0.beta8129 using a 100,000-bp window with a 10,000-bp jump to calculate Fst. We defined outliers as three standard deviations above the mean, making the outliers above 0.438 for female versus monoecious comparisons and 0.206 for between high-THC and high-CBD.

To estimate adaptive substitutions, or alpha (α)155, in hop, we used the McDonald–Kreitman test (MKT). We first identified variants as described above for hop, with the exception that we used the H. lupulus var. lupulus “21110M” HAP1 reference. Next, we used the GATK v4.5156 FastaAlternateReferenceMaker tool to convert the VCF for each genotype to fasta format and used gffread v0.12.7 (https://github.com/gpertea/gffread) to output coding sequences. We used the R package iMKT v0.1.1157, where we applied a Fay, Wyckoff, and Wu correction158 of 0.15. Because the MKT can be underpowered when there are low counts in the 2-by-2 contingency table, we only compared results for which the row and column totals were six or greater159. After calculating the MKT, we used Fisher's exact test to test for significance on each gene and applied a false discovery rate adjustment to the p-values.

Backcross generation and analysis 

To generate the backcross population (BC1), a cross was made between New West Genetics’ breeding lines using NWG 7619 as the female and NWG 7384 as the sire. NWG 7619 is an early-flowering dioecious variety, and NWG 8232 is a later-flowering monoecious variety. Crossing was performed in a greenhouse in Fort Collins, CO, where plants were grown in 3.5” pots using Pro-Mix HP potting soil with supplemental lighting set for 16-h daylength. A monoecious NWG 7384 was paired with a NWG 7619 female under a Midco (St. Louis, MO) pollination bag. F1 seed was then planted outdoors in a small plot in Winsor, CO, where XY males were rogued so that only NWG 8232 pollen was available to sire BC1 seed on F1 females. Approximately 20 grams of BC1 seed were harvested from two F1 plants. BC1 seed was planted in a greenhouse under conditions similar to those already described and used for phenotyping the degree of monoecy using the following 0–4 scale: 0 = no staminate flowers, 1 = more pistillate than staminate, 2 = roughly equal pistillate and staminate flowers, 3 =  more staminate than pistillate, 4 = only staminate flowers. Floral phenotypes were scored as BC1 plants flowered over the course of an approximately 4-week period.

DNA was extracted using the BioEcho EchoLUTION Plant DNA 96 Kit (Product No. 010-103-002), and sequencing libraries were constructed using a Twist 96-Plex Library Prep kit. Sequencing was performed on an Illumina NovaSeq 6000 at Discovery Life Sciences. To ensure that no Y chromosomes were present in the BC1 data, which may impact the results, we tested for the presence of the Y chromosome. To accomplish this, we used the bbduk function of BBmap v38.86 (Bushnell, sourceforge.net/projects/bbmap/) to identify reads that contained perfect matches to the list of cannabis Y-mers. We mapped these reads to the Y-containing HAP1 of the “Otto II” reference, using BWA v0.7.17, and we used Samtools v.1.1096 to calculate coverage of the non-recombining region. We plotted these coverage values using ggplot2160 and found no evidence of a Y chromosome (Supplementary Fig. 11).

KhufuPAN (https://github.com/w-korani/KhufuPAN) was used to produce a genome graph and analyze low-coverage whole-genome sequencing. A graph was constructed using Minigraph-cactus161 with each haplotype individually in the graph of “Carmagnola”, “Futura 75”, “Otto II”, and “Uso 31”. The reference genome used was HAP2 of “Otto II” with the X. The graph consists of 8,303,397 SNPs and 2,610,350 SVs. A total of 950 lines (average depth 1.01X) were mapped to the pangraph, and variants were genotyped. After filtering for a minimum depth of 2, missing data across lines less than 75%, and minor allele frequency of 0.01, 906,075 variants were retained in the panmap. Filtering for lowly sequenced lines resulted in retaining 887 lines for the final analysis. Identification of significant allele frequency differences between male and female classes was done using KhufuEnv79. The genotypes from the female (n = 367) and monoecious (n = 520) samples were extracted, group-specific allele frequency was calculated, and the difference in allele frequency was calculated. Significant sites with allele frequency differences of greater than 0.4 were identified.

Gene expression analyses

To examine gene expression in cannabis, we used existing RNAseq data for leaf, stem, and apical meristem tissues (n = 72; Supplementary Data 14)59,87. We first trimmed the data as described above, and aligned the data using STAR v.2.7.9a162 to the C. sativa “Otto II” HAP2 reference, but also included the Y from HAP1. To inhibit multi-mapping issues with the PAR being represented twice on the sex chromosomes, we hard-masked the Y PAR using bedtools v2.31.0 maskfasta163. We assembled transcripts using StringTie v2.1.7164 and ran differential gene expression using DESeq2 v1.46165. To generate the heatmap of expression, in addition to the samples used for gene expression, we included RNAseq libraries across diverse tissue types (n = 100; Supplementary Data 14)59,87,166,167. We generated the heatmap using pheatmap v1.0.12168 in R v4.4.2.

Additional variant calling

In addition to identifying SNPs with bcftools mpileup for the cannabis HiFi data, we also called SNPs and small INDELs using DeepVariant v1.8.0169 and larger SVs using Sawfish v0.12.10 (https://github.com/PacificBiosciences/sawfish).

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

41467_2026_73233_MOESM2_ESM.pdf (203.1KB, pdf)

Description of Additional Supplementary Information

Supplementary Data 1 (50.5KB, xlsx)
Supplementary Data 2 (269.4KB, xlsx)
Supplementary Data 3 (54.4KB, xlsx)
Supplementary Data 4 (52.2KB, xlsx)
Supplementary Data 5 (51.3KB, xlsx)
Supplementary Data 6 (46KB, xlsx)
Supplementary Data 7 (50.9KB, xlsx)
Supplementary Data 8 (51.8KB, xlsx)
Supplementary Data 9 (306.9KB, xlsx)
Supplementary Data 10 (194.7KB, xlsx)
Supplementary Data 11 (49.7KB, xlsx)
Supplementary Data 12 (45.8KB, xlsx)
Supplementary Data 13 (58.9KB, xlsx)
Supplementary Data 14 (104.1KB, xlsx)
Supplementary Data 15 (52.7KB, xlsx)
Supplementary Data 16 (59.8KB, xlsx)
Supplementary Data 17 (50.5KB, xlsx)
Supplementary Data 18 (52.5KB, xlsx)
Reporting Summary (119.1KB, pdf)

Source data

Source Data (21.7MB, xlsx)

Acknowledgements

We would like to thank the US Department of Agriculture Hemp Germplasm Laboratory (8060-21000-034-000-D) for generously providing materials for this manuscript. We would also like to thank the HudsonAlpha IT Department for their research computing support. Support for this work was provided by the US Department of Agriculture National Institute of Food and Agriculture Postdoctoral Fellowship (USDA NIFA) no. 2022-67012-38987 (S.B.C.), USDA NIFA no. 2023-67013-39620 (A.H.), National Science Foundation (NSF) IOS-PGRP CAREER no. 2239530 (A.H.), USDA Non-Assistance Cooperative Agreement (NACA) no. 58-8060-4-002 (A.H., Z.S.), NIH R35 GM147107 (R.F.G.), and Institute of Cannabis Research at Colorado State University Pueblo ICR-FY22-Kane (R.F.G., D.V., and N.C.K.).

Author contributions

Conceptualization: S.B.C., J.S.H., J.K.M., G.J.M., and A.H. Funding acquisition: S.B.C., Z.S., R.F.G., D.V., N.C.K., and A.H. Formal analysis: S.B.C., P.C.B., Z.A.M., W.K., L.P.-C., R.C.L., N.A., and A.O. Investigation: S.B.C., J.M., Z.S., G.S., T.G., R.F., H.H., H.M., Z.M., J.G., and J.K.M. Methodology: S.B.C, A.H. Resources: J.S.H., Z.S., G.S., T.G., K.A.E., Z.E.M., L.R.O., L.B.S., N.C.K., G.J.M., J.C., and A.H. Visualization: S.B.C., P.C.B., J.T.L., L.M.A., and Z.A.M. Writing-original draft: S.B.C., A.H. Writing-review and editing: S.B.C., P.C.B., J.T.L., L.M.A., J.S.H., L.P.-C., R.C.L., Z.S., G.S., K.A.E., L.B.S., D.V., R.F.G., N.C.K., J.K.M., T.P.M., and A.H.

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.

Data availability

The genome assemblies have been deposited on NCBI under the BioProjects listed in Supplementary Data 3. The genome assemblies have additionally been deposited on Zenodo 10.5281/zenodo.16326991, as well as their gene and repeat annotations. Sequencing libraries used for the genome assemblies and annotations, and the remaining sequencing generated in support of this manuscript, are publicly available on NCBI under BioProject PRJNA1437418, and NCBI SRA accessions are provided in Supplementary Data 14. Additional files in support of this manuscript have also been deposited on Zenodo 10.5281/zenodo.16326991. Source data are provided with this paper.

Code availability

An R script that supports this manuscript can be found at https://github.com/sarahcarey/Cannabaceae_sexChromosomes. A static release can be found at Zenodo 10.5281/zenodo.19376146.

Competing interests

A.H., J.C. are co-founders and board members of Veil Genomics, a long-read genotyping company. D.V. is a board member of the 501(C)3 non-profit Agricultural Genomics Foundation and the sole owner of the company CGRI, LLC. J.K.M., R.F. are founders of New West Genetics, a breeding program that produces hybrid hemp seed. All other authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains Supplementary material available at 10.1038/s41467-026-73233-7.

References

  • 1.Ren, G. et al. Large-scale whole-genome resequencing unravels the domestication history of Cannabis sativa. Sci. Adv.7, eabg2286 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Small, E. & Marcus, D. Hemp: a new crop with new uses for North America. Trends New Crops New Uses24, 284–326 (2002). [Google Scholar]
  • 3.Hornsey, I. S. Brewing (Royal Society of Chemistry, 2013).
  • 4.Käfer, J., Méndez, M. & Mousset, S. Labile sex expression in angiosperm species with sex chromosomes. Philos. Trans. R. Soc. Lond. B Biol. Sci.377, 20210216 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Harrison, J. Effect of hop seeds on beer quality. J. Inst. Brew.77, 350–352 (1971). [Google Scholar]
  • 6.Lipson Feder, C. et al. Fertilization following pollination predominantly decreases phytocannabinoids accumulation and alters the accumulation of terpenoids in Cannabis inflorescences. Front. Plant Sci.12, 753847 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Razumova, O. V., Alexandrov, O. S., Divashuk, M. G., Sukhorada, T. I. & Karlov, G. I. Molecular cytogenetic analysis of monoecious hemp (Cannabis sativa L.) cultivars reveals its karyotype variations and sex chromosomes constitution. Protoplasma253, 895–901 (2016). [DOI] [PubMed] [Google Scholar]
  • 8.Garcia-de Heer, L., Mieog, J., Burn, A. & Kretzschmar, T. Why not XY? Male monoecious sexual phenotypes challenge the female monoecious paradigm in Cannabis sativa L. Front. Plant Sci.15, 1412079 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Stevens, N. M. Studies in Spermatogenesis (Carnegie Institution of Washington, 1905).
  • 10.Winge, O. On sex chromosomes, sex determination and preponderance of females in some dioecious plants. Compt. Rend. Trav. Lab. Carlsberg15, 1–26 (1923). [Google Scholar]
  • 11.Hirata, K. Sex reversal in hemp. J. Soc. Agric. For. Sapporo16, 145–168 (1924).
  • 12.Renner, S. S. & Müller, N. A. Plant sex chromosomes defy evolutionary models of expanding recombination suppression and genetic degeneration. Nat. Plants7, 392–402 (2021). [DOI] [PubMed] [Google Scholar]
  • 13.Shephard, H. L., Parker, J. S., Darby, P. & Ainsworth, C. C. Sexual development and sex chromosomes in hop. New Phytol.148, 397–411 (2000). [DOI] [PubMed] [Google Scholar]
  • 14.Prentout, D. et al. Plant genera Cannabis and Humulus share the same pair of well-differentiated sex chromosomes. New Phytol.231, 1599–1611 (2021). [DOI] [PubMed] [Google Scholar]
  • 15.Sinclair, A. H. et al. A gene from the human sex-determining region encodes a protein with homology to a conserved DNA-binding motif. Nature346, 240–244 (1990). [DOI] [PubMed] [Google Scholar]
  • 16.Parker, J. S. & Clark, M. S. Dosage sex-chromosome systems in plants. Plant Sci.80, 79–92 (1991). [Google Scholar]
  • 17.Gao, S. et al. A high-quality reference genome of wild Cannabis sativa. Hortic. Res. 7, 73 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Padgitt-Cobb, L. K., Pitra, N. J., Matthews, P. D., Henning, J. A. & Hendrix, D. A. An improved assembly of the ‘cascade’ hop (Humulus lupulus) genome uncovers signatures of molecular evolution and refines time of divergence estimates for the Cannabaceae family. Hortic. Res.10, uhac281 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Rhie, A. et al. Towards complete and error-free genome assemblies of all vertebrate species. Nature592, 737–746 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Rhie, A. et al. The complete sequence of a human Y chromosome. Nature621, 344–354 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Muyle, A. et al. SEX-DETector: a probabilistic approach to study sex chromosomes in non-model organisms. Genome Biol. Evol.8, 2530–2543 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Käfer, J., Lartillot, N., Marais, G. A. B. & Picard, F. Detecting sex-linked genes using genotyped individuals sampled in natural populations. Genetics218, iyab053 (2021). [DOI] [PMC free article] [PubMed]
  • 23.Xu, X.-W., Sun, P., Gao, C., Zheng, W. & Chen, S. Assembly of the poorly differentiated Verasper variegatus W chromosome by different sequencing technologies. Sci. Data10, 893 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Makova, K. D. et al. The complete sequence and comparative analysis of ape sex chromosomes. Nature630, 401–411 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Morris, J., Darolti, I., Bloch, N. I., Wright, A. E. & Mank, J. E. Shared and species-specific patterns of nascent Y chromosome evolution in two guppy species. Genes9, 238 (2018). [DOI] [PMC free article] [PubMed]
  • 26.Akagi, T., Henry, I. M., Tao, R. & Comai, L. A. Y-chromosome–encoded small RNA acts as a sex determinant in persimmons. Science346, 646–650 (2014). [DOI] [PubMed] [Google Scholar]
  • 27.Tennessen, J. A. et al. Repeated translocation of a gene cassette drives sex-chromosome turnover in strawberries. PLoS Biol.16, e2006062 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Carey, S. B. et al. Representing sex chromosomes in genome assemblies. Cell Genom.2, 100132 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Haunold, A. Hop production, breeding, and variety development in various countries. J. Am. Soc. Brew. Chem.39, 27–34 (1981). [Google Scholar]
  • 30.Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods18, 170–175 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol.21, 245 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Otto, S. P. et al. About PAR: the distinct evolutionary dynamics of the pseudoautosomal region. Trends Genet.27, 358–367 (2011). [DOI] [PubMed] [Google Scholar]
  • 33.Zhang, Q., Onstein, R. E., Little, S. A. & Sauquet, H. Estimating divergence times and ancestral breeding systems in Ficus and Moraceae. Ann. Bot.123, 191–204 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Dennis Thomas, T. In vitro modification of sex expression in mulberry (Morus Alba) by ethrel and silver nitrate. Plant Cell Tissue Organ Cult.77, 277–281 (2004). [Google Scholar]
  • 35.Mohan Ram, H. Y. & Sett, R. Modification of growth and sex expression in Cannabis sativa by aminoethoxyvinylglycine and ethephon. Z. Pflanzenphysiol.105, 165–172 (1982). [Google Scholar]
  • 36.Zhang, X. et al. Genomes of the banyan tree and pollinator wasp provide insights into fig-wasp coevolution. Cell183, 875–889 (2020). [DOI] [PubMed] [Google Scholar]
  • 37.Xia, Z. et al. Chromosome-level genomes reveal the genetic basis of descending dysploidy and sex determination in Morus plants. Genom. Proteom. Bioinform.10.1016/j.gpb.2022.08.005 (2022). [DOI] [PMC free article] [PubMed]
  • 38.Bruna, T., Hoff, K. J., Lomsadze, A., Stanke, M., & Borodovsky, M. BRAKER2: automatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS supported by a protein database. NAR genomics and bioinformatics3, 10.1093/nargab/lqaa108 (2021). [DOI] [PMC free article] [PubMed]
  • 39.Gabriel, L. et al. BRAKER3: fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA. Genome Res. 10.1101/gr.278090.123 (2024). [DOI] [PMC free article] [PubMed]
  • 40.Razumova, O. V., Divashuk, M. G., Alexandrov, O. S. & Karlov, G. I. GISH painting of the Y chromosomes suggests advanced phases of sex chromosome evolution in three dioecious Cannabaceae species (Humulus lupulus, H. japonicus, and Cannabis sativa). Protoplasma260, 249–256 (2023). [DOI] [PubMed] [Google Scholar]
  • 41.Zhang, C., Rabiee, M., Sayyari, E. & Mirarab, S. ASTRAL-III: polynomial time species tree reconstruction from partially resolved gene trees. BMC Bioinform.19, 153 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Lahn, B. T. & Page, D. C. Four evolutionary strata on the human X chromosome. Science286, 964–967 (1999). [DOI] [PubMed] [Google Scholar]
  • 43.Charlesworth, D., Charlesworth, B. & Marais, G. Steps in the evolution of heteromorphic sex chromosomes. Heredity95, 118–128 (2005). [DOI] [PubMed] [Google Scholar]
  • 44.Papadopulos, A. S. T., Chester, M., Ridout, K. & Filatov, D. A. Rapid Y degeneration and dosage compensation in plant sex chromosomes. Proc. Natl. Acad. Sci. USA112, 13021–13026 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Sacchi, B. et al. Phased assembly of neo-sex chromosomes reveals extensive Y degeneration and rapid genome evolution in Rumex hastatulus. Mol. Biol. Evol.41, msae074 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Mank, J. E., Vicoso, B., Berlin, S. & Charlesworth, B. Effective population size and the faster-X effect: empirical results and their interpretation. Evolution64, 663–674 (2010). [DOI] [PubMed] [Google Scholar]
  • 47.Krasovec, M., Nevado, B. & Filatov, D. A. A comparison of selective pressures in plant X-linked and autosomal genes. Genes9, 234 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Ou, S. et al. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol.20, 275 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Yang, Z. PAML 4: phylogenetic analysis by maximum likelihood. Mol. Biol. Evol.24, 1586–1591 (2007). [DOI] [PubMed] [Google Scholar]
  • 50.Payseur, B. A., Presgraves, D. C. & Filatov, D. A. Introduction: sex chromosomes and speciation. Mol. Ecol.27, 3745–3748 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Mediavilla, V., Jonquera, M. & Schmid-Slembrouck, I. Decimal code for growth stages of hemp (Cannabis sativa L.). J. Int. Hemp Assoc.5, 65 (1998).
  • 52.Shi, J., Schilling, S. & Melzer, R. Morphological and genetic analysis of inflorescence and flower development in hemp (Cannabis sativa L.). Preprint at bioRxiv10.1101/2024.01.25.577276 (2024).
  • 53.Brandoli, C., Petri, C., Egea-Cortines, M. & Weiss, J. Gigantea: uncovering new functions in flower development. Genes11, 1142 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Romero, J. M. et al. CONSTANS, a HUB for all seasons: how photoperiod pervades plant physiology regulatory circuits. Plant Cell36, 2086–2102 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Steel, L., Welling, M., Ristevski, N., Johnson, K. & Gendall, A. Comparative genomics of flowering behavior in Cannabis sativa. Front. Plant Sci. 14, 1227898 (2023). [DOI] [PMC free article] [PubMed]
  • 56.Corbesier, L. et al. FT protein movement contributes to long-distance signaling in floral induction of Arabidopsis. Science316, 1030–1033 (2007). [DOI] [PubMed] [Google Scholar]
  • 57.Dowling, C. A., Michael, T. P., McCabe, P. F., Schilling, S. & Melzer, R. FT-like genes in Cannabis and hops: sex specific expression and copy-number variation may explain flowering time variation. BMC genomics26, 930 (2025). [DOI] [PMC free article] [PubMed]
  • 58.Abe, M. et al. FD, a bZIP protein mediating signals from the floral pathway integrator FT at the shoot apex. Science309, 1052–1056 (2005). [DOI] [PubMed] [Google Scholar]
  • 59.Shi, J., Toscani, M., Dowling, C. A., Schilling, S. & Melzer, R. Identification of genes associated with sex expression and sex determination in hemp (Cannabis sativa L.). J. Exp. Bot. 10.1093/jxb/erae429 (2024). [DOI] [PMC free article] [PubMed]
  • 60.Li, H. et al. A heteromeric membrane-bound prenyltransferase complex from hop catalyzes three sequential aromatic prenylations in the bitter acid pathway. Plant Physiol.167, 650–659 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Luo, X. et al. Complete biosynthesis of cannabinoids and their unnatural analogues in yeast. Nature567, 123–126 (2019). [DOI] [PubMed] [Google Scholar]
  • 62.Rea, K. A. et al. Biosynthesis of cannflavins A and B from Cannabis sativa L. Phytochemistry164, 162–171 (2019). [DOI] [PubMed] [Google Scholar]
  • 63.Lynch, R. C. et al. Domesticated cannabinoid synthases amid a wild mosaic cannabis pangenome. Nature10.1038/s41586-025-09065-0 (2025). [DOI] [PMC free article] [PubMed]
  • 64.Stack, G. M. et al. Comparison of recombination rate, reference bias, and unique pangenomic haplotypes in Cannabis sativa using seven DE novo genome assemblies. Int. J. Mol. Sci. 26, 1165 (2025). [DOI] [PMC free article] [PubMed]
  • 65.Yu, Q. et al. Chromosomal location and gene paucity of the male specific region on papaya Y chromosome. Mol. Genet. Genom.278, 177–185 (2007). [DOI] [PubMed] [Google Scholar]
  • 66.She, H. et al. Evolution of the spinach sex-linked region within a rarely recombining pericentromeric region. Plant Physiol.193, 1263–1280 (2023). [DOI] [PubMed] [Google Scholar]
  • 67.Rice, W. Sex chromosomes and the evolution of sexual dimorphism. Evolution38, 735–742 (1984). [DOI] [PubMed]
  • 68.Galoch, E. The hormonal control of sex differentiation in dioecious plants of hemp (Cannabis sativa). The influence of plant growth regulators on sex expression in male and female plants. Acta Soc. Bot. Pol.47, 153–162 (2015). [Google Scholar]
  • 69.Lubell, J. D. & Brand, M. H. Foliar sprays of silver thiosulfate produce male flowers on female hemp plants. HortTechnology28, 743–747 (2018). [Google Scholar]
  • 70.Faux, A.-M., Draye, X., Flamand, M.-C., Occre, A. & Bertin, P. Identification of QTLs for sex expression in dioecious and monoecious hemp (Cannabis sativa L.). Euphytica209, 357–376 (2016). [Google Scholar]
  • 71.Akagi, T. et al. Evolution and functioning of an X-A balance sex-determining system in hops. Nat. Plants10.1038/s41477-025-02017-6 (2025). [DOI] [PubMed]
  • 72.Toscani, M. et al. An ancient X chromosomal region harbours three genes potentially controlling sex determination in Cannabis sativa. Preprint at bioRxiv10.1101/2025.07.03.663031 (2025).
  • 73.Han, S., Green, L. & Schnell, D. J. The signal peptide peptidase is required for pollen function in Arabidopsis. Plant Physiol.149, 1289–1301 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Hwang, J.-U., Vernoud, V., Szumlanski, A., Nielsen, E. & Yang, Z. A tip-localized RhoGAP controls cell polarity by globally inhibiting Rho GTPase at the cell apex. Curr. Biol.18, 1907–1916 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Westergaard, M. The mechanism of sex determination in dioecious flowering plants. Adv. Genet.9, 217–281 (1958). [DOI] [PubMed] [Google Scholar]
  • 76.Warmke, H. E. & Davidson, H. Polyploidy investigations. Yearb. Carnegie Inst. Wash.43, 135–139 (1944). [Google Scholar]
  • 77.Haunold, A. Cytology, sex expression, and growth of a tetraploid × diploid cross in hop (Humulus lupulus L.). Crop Sci.11, 868–871 (1971). [Google Scholar]
  • 78.Hyden, B. et al. Structural variation of a sex-linked region confers monoecy and implicates GATA15 as a master regulator of sex in Salix purpurea. New Phytol.238, 2512–2523 (2023). [DOI] [PubMed] [Google Scholar]
  • 79.Wright, H. C., Davis, C. E. M., Clevenger, J. & Korani, W. KhufuEnv, an auxiliary toolkit for building computational pipelines for plant and animal breeding. Preprint at bioRxiv10.1101/2025.03.28.645917 (2025). [DOI] [PubMed]
  • 80.Adams, D. O. & Yang, S. F. Methionine metabolism in apple tissue: implication of S-adenosylmethionine as an intermediate in the conversion of methionine to ethylene: implication of S-adenosylmethionine as an intermediate in the conversion of methionine to ethylene. Plant Physiol.60, 892–896 (1977). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Adams, D. O. & Yang, S. F. Ethylene biosynthesis: identification of 1-aminocyclopropane-1-carboxylic acid as an intermediate in the conversion of methionine to ethylene. Proc. Natl. Acad. Sci. USA76, 170–174 (1979). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Boller, T., Herner, R. C. & Kende, H. Assay for and enzymatic formation of an ethylene precursor, 1-aminocyclopropane-1-carboxylic acid. Planta145, 293–303 (1979). [DOI] [PubMed] [Google Scholar]
  • 83.Hamilton, A. J., Bouzayen, M. & Grierson, D. Identification of a tomato gene for the ethylene-forming enzyme by expression in yeast. Proc. Natl. Acad. Sci. USA88, 7434–7437 (1991). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Ververidis, P. & John, P. Complete recovery in vitro of ethylene-forming enzyme activity. Phytochemistry30, 725–727 (1991). [Google Scholar]
  • 85.Boualem, A. et al. A cucurbit androecy gene reveals how unisexual flowers develop and dioecy emerges. Science350, 688–691 (2015). [DOI] [PubMed] [Google Scholar]
  • 86.Zhang, S. et al. The control of carpel determinacy pathway leads to sex determination in cucurbits. Science378, 543–549 (2022). [DOI] [PubMed] [Google Scholar]
  • 87.Dowling, C. A. et al. A FLOWERING LOCUS T ortholog is associated with photoperiod-insensitive flowering in hemp (Cannabis sativa L.). Plant J.119, 383–403 (2024). [DOI] [PubMed] [Google Scholar]
  • 88.Müller, N. A. et al. A single gene underlies the dynamic evolution of poplar sex determination. Nat. Plants6, 630–637 (2020). [DOI] [PubMed] [Google Scholar]
  • 89.Harkess, A. et al. Sex determination by two Y-linked genes in garden asparagus. Plant Cell32, 1790–1796 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Cronk, Q. The distribution of sexual function in the flowering plant: from monoecy to dioecy. Philos. Trans. R. Soc. Lond. B Biol. Sci.377, 20210486 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Smith, C. A. et al. The avian Z-linked gene DMRT1 is required for male sex determination in the chicken. Nature461, 267–271 (2009). [DOI] [PubMed] [Google Scholar]
  • 92.Vaser, R., Sović, I., Nagarajan, N. & Šikić, M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res.27, 737–746 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Astashyn, A. et al. Rapid and sensitive detection of genome contamination at scale with FCS-GX. Genome Biol.25, 60 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics30, 2114–2120 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Marcais, G. & Kingsford, C. Jellyfish: a fast k-mer counter. Tutor. Manuais1, 1038 (2012).
  • 96.Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics25, 2078–2079 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Zhou, C., McCarthy, S. A. & Durbin, R. YaHS: yet another Hi-C scaffolding tool. Bioinformatics39, btac808 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Dudchenko, O. et al. The Juicebox Assembly Tools module facilitates de novo assembly of mammalian genomes with chromosome-length scaffolds for under $1000. Preprint at bioRxiv10.1101/254797 (2018).
  • 99.Lovell, J. T. et al. GENESPACE tracks regions of interest and gene copy number variation across multiple genomes. Elife11, e78526 (2022). [DOI] [PMC free article] [PubMed]
  • 100.Grassa, C. J. et al. A new Cannabis genome assembly associates elevated cannabidiol (CBD) with hemp introgressed into marijuana. New Phytol.230, 1665–1679 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Huang, N. & Li, H. compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics39, btad595 (2023). [DOI] [PMC free article] [PubMed]
  • 102.Gabriel, L., Hoff, K. J., Brůna, T., Borodovsky, M. & Stanke, M. TSEBRA: transcript selector for BRAKER. BMC Bioinform.22, 566 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics34, i884–i890 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol.37, 907–915 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics34, 3094–3100 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA117, 9451–9457 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Smit, A. F. A., Hubley, R. & Green, P. RepeatMasker Open-4.0. http://www.repeatmasker.org (2015).
  • 108.Gasic, K., Hernandez, A. & Korban, S. S. RNA extraction from different apple tissues rich in polyphenols and polysaccharides for cDNA library construction. Plant Mol. Biol. Rep.22, 437–438 (2004). [Google Scholar]
  • 109.Kuznetsov, D. et al. OrthoDB v11: annotation of orthologs in the widest sampling of organismal diversity. Nucleic Acids Res.51, D445–D451 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.van Velzen, R. et al. Comparative genomics of the nonlegume parasponia reveals insights into evolution of nitrogen-fixing rhizobium symbioses. Proc. Natl. Acad. Sci. USA115, E4700–E4709 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Ou, S. et al. Differences in activity and stability drive transposable element variation in tropical and temperate maize. Genome Res.34, 1140–1153 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Vollger, M. R., Kerpedjiev, P., Phillippy, A. M. & Eichler, E. E. StainedGlass: interactive visualization of massive tandem repeat structures with identity heatmaps. Bioinformatics38, 2049–2051 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res.27, 573–580 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Madden, T. L. The BLAST Sequence Analysis Tool. in The NCBI Handbook (National Center for Biotechnology Information, 2013).
  • 115.Neumann, P. et al. Plant centromeric retrotransposons: a structural and cytogenetic perspective. Mob. DNA2, 4 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Xia, Z. et al. Haplotype-resolved chromosomal-level genome assembly reveals regulatory variations in mulberry fruit anthocyanin content. Hortic. Res.11, uhae120 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Liao, Z. et al. A telomere-to-telomere reference genome of ficus (Ficus hispida) provides new insights into sex determination. Hortic. Res.11, uhad257 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Emms, D. M. & Kelly, S. OrthoFinder: solving fundamental biases in whole genome comparisons dramatically improves orthogroup inference accuracy. Genome Biol.16, 157 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Emms, D. M. & Kelly, S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol.20, 238 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol.30, 772–780 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Stamatakis, A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics30, 1312–1313 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Huerta-Cepas, J., Serra, F. & Bork, P. ETE 3: reconstruction, analysis, and visualization of phylogenomic data. Mol. Biol. Evol.33, 1635–1638 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Zhang, Z. et al. KaKs_Calculator: calculating Ka and Ks through model selection and model averaging. Genom. Proteom. Bioinform.4, 259–263 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Lindeløv, J. K. mcp: An R Package for Regression With Multiple Change Points. 10.31219/osf.io/fzqxv (2020).
  • 125.Pohlert, T. & Pohlert, M. T. Package ‘pmcmr’. R Package Version1, (2018).
  • 126.Rosner, B. Percentage points for a generalized ESD many-outlier procedure. Technometrics25, 165–172 (1983). [Google Scholar]
  • 127.Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Preprint at 10.48550/arXiv.1303.3997 (2013).
  • 128.Li, H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics27, 2987–2993 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Korunes, K. L. & Samuk, K. Pixy: unbiased estimation of nucleotide diversity and divergence in the presence of missing data. Mol. Ecol. Resour.21, 1359–1368 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Zhang, C. & Mirarab, S. Weighting by gene tree uncertainty improves accuracy of quartet-based species trees. Mol. Biol. Evol. 39, msac215 (2022). [DOI] [PMC free article] [PubMed]
  • 131.Khan, A. et al. A phased, chromosome-scale genome of ‘honeycrisp’ apple (Malus domestica). GigaByte2022, 1–15 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Yocca, A. et al. A chromosome-scale assembly for ‘d’Anjou’pear. G3 Genes Genomes Genet.14, jkae003 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Yang, M. et al. Insights into the evolution and spatial chromosome architecture of jujube from an updated gapless genome assembly. Plant Commun.4, 100662 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Kubatko, L. S. & Degnan, J. H. Inconsistency of phylogenetic estimates from concatenated data under coalescence. Syst. Biol.56, 17–24 (2007). [DOI] [PubMed] [Google Scholar]
  • 135.Edwards, S. V. et al. Implementing and testing the multispecies coalescent model: a valuable paradigm for phylogenomics. Mol. Phylogenet. Evol.94, 447–462 (2016). [DOI] [PubMed] [Google Scholar]
  • 136.Johnson, M. G. et al. HybPiper: extracting coding sequence and introns for phylogenetics from high-throughput sequencing reads using target enrichment. Appl. Plant Sci.4, 1600016 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Bankevich, A. et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J. Comput. Biol.19, 455–477 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Capella-Gutiérrez, S., Silla-Martínez, J. M. & Gabaldón, T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics25, 1972–1973 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Nguyen, L.-T., Schmidt, H. A., von Haeseler, A. & Minh, B. Q. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol. Biol. Evol.32, 268–274 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Hoang, D. T., Chernomor, O., von Haeseler, A., Minh, B. Q. & Vinh, L. S. UFBoot2: improving the ultrafast bootstrap approximation. Mol. Biol. Evol.35, 518–522 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Kalyaanamoorthy, S., Minh, B. Q., Wong, T. K. F., von Haeseler, A. & Jermiin, L. S. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat. Methods14, 587–589 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Zhou, C. et al. Oatk: a de novo assembly tool for complex plant organelle genomes. Genome Biology26, 235 (2025). [DOI] [PMC free article] [PubMed]
  • 143.Dierckxsens, N., Mardulyn, P. & Smits, G. NOVOPlasty: de novo assembly of organelle genomes from whole genome data. Nucleic Acids Res.45, e18 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Bouckaert, R. et al. BEAST 2.5: an advanced software platform for Bayesian evolutionary analysis. PLoS Comput. Biol.15, e1006650 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Stadler, T. Sampling-through-time in birth-death trees. J. Theor. Biol.267, 396–404 (2010). [DOI] [PubMed] [Google Scholar]
  • 146.Heath, T. A., Huelsenbeck, J. P. & Stadler, T. The fossilized birth-death process for coherent calibration of divergence-time estimates. Proc. Natl. Acad. Sci. USA111, E2957–E2966 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Drummond, A. J., Ho, S. Y. W., Phillips, M. J. & Rambaut, A. Relaxed phylogenetics and dating with confidence. PLoS Biol.4, e88 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Rambaut, A., Drummond, A. J., Xie, D., Baele, G. & Suchard, M. A. Posterior summarization in Bayesian phylogenetics using tracer 1.7. Syst. Biol.67, 901–904 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Magallón, S., Gómez-Acevedo, S., Sánchez-Reyes, L. L. & Hernández-Hernández, T. A metacalibrated time-tree documents the early rise of flowering plant phylogenetic diversity. New Phytol.207, 437–453 (2015). [DOI] [PubMed] [Google Scholar]
  • 150.Knobloch, E., & Mai, D. H. Monographie der Früchte und Samen in der Kreide von Mitteleuropa (Ústřední Ústav Geologický v Academii, Nakl. Československé Akademie Věd, 1986).
  • 151.Friis, E. M., Crane, P. R. & Pedersen, K. R. Early Flowers and Angiosperm Evolution (Cambridge University Press, 2011).
  • 152.Collinson, M. E. The fossil history of the Moraceae, Urticaceae (including Cecropiaceae), and Cannabaceae. in Evolution, Systematics, and Fossil History of the Hamamelidae, Vol. 2, 319–339 (Oxford University Press, 1989).
  • 153.Suyama, M., Torrents, D. & Bork, P. PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res.34, W609–W612 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.Lee, K. C. et al. Long-read low-pass sequencing for high-resolution trait mapping. Preprint at bioRxiv10.1101/2025.01.09.632261 (2025).
  • 155.Smith, N. G. C. & Eyre-Walker, A. Adaptive protein evolution in Drosophila. Nature415, 1022–1024 (2002). [DOI] [PubMed] [Google Scholar]
  • 156.McKenna, A. et al. The genome analysis toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res.20, 1297–1303 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157.Murga-Moreno, J., Coronado-Zamora, M., Hervas, S., Casillas, S. & Barbadilla, A. iMKT: the integrative McDonald and Kreitman test. Nucleic Acids Res.47, W283–W288 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 158.Fay, J. C., Wyckoff, G. J. & Wu, C. I. Positive and negative selection on the human genome. Genetics158, 1227–1234 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 159.Begun, D. J. et al. Population genomics: whole-genome analysis of polymorphism and divergence in Drosophila simulans. PLoS Biol.5, e310 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160.Wickham, H. ggplot2: Elegant Graphics for Data Analysis (Springer, 2016).
  • 161.Hickey, G. et al. Pangenome graph construction from genome alignments with minigraph-cactus. Nat. Biotechnol.42, 663–673 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 162.Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics29, 15–21 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 163.Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics26, 841–842 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 164.Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat. Biotechnol.33, 290–295 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 165.Love, M., Anders, S. & Huber, W. Differential analysis of count data–the DESeq2 package. Genome Biol.15, 10–1186 (2014). [Google Scholar]
  • 166.Mi, Y. et al. Characterization and co-expression analysis of ATP-binding cassette transporters provide insight into genes related to cannabinoid transport in Cannabis sativa L. Int. J. Biol. Macromol.242, 124934 (2023). [DOI] [PubMed] [Google Scholar]
  • 167.Braich, S., Baillie, R., Jewell, L. S., Spangenberg, G. & Cogan, N. Generation of a comprehensive transcriptome atlas and transcriptome dynamics in medicinal cannabis. Sci. Rep. 9, 16583 (2019). [DOI] [PMC free article] [PubMed]
  • 168.Kolde, R. & Kolde, M. R. Package ‘pheatmap’. R package (2015).
  • 169.Poplin, R. et al. A universal SNP and small-indel variant caller using deep neural networks. Nat. Biotechnol.36, 983–987 (2018). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

41467_2026_73233_MOESM2_ESM.pdf (203.1KB, pdf)

Description of Additional Supplementary Information

Supplementary Data 1 (50.5KB, xlsx)
Supplementary Data 2 (269.4KB, xlsx)
Supplementary Data 3 (54.4KB, xlsx)
Supplementary Data 4 (52.2KB, xlsx)
Supplementary Data 5 (51.3KB, xlsx)
Supplementary Data 6 (46KB, xlsx)
Supplementary Data 7 (50.9KB, xlsx)
Supplementary Data 8 (51.8KB, xlsx)
Supplementary Data 9 (306.9KB, xlsx)
Supplementary Data 10 (194.7KB, xlsx)
Supplementary Data 11 (49.7KB, xlsx)
Supplementary Data 12 (45.8KB, xlsx)
Supplementary Data 13 (58.9KB, xlsx)
Supplementary Data 14 (104.1KB, xlsx)
Supplementary Data 15 (52.7KB, xlsx)
Supplementary Data 16 (59.8KB, xlsx)
Supplementary Data 17 (50.5KB, xlsx)
Supplementary Data 18 (52.5KB, xlsx)
Reporting Summary (119.1KB, pdf)
Source Data (21.7MB, xlsx)

Data Availability Statement

The genome assemblies have been deposited on NCBI under the BioProjects listed in Supplementary Data 3. The genome assemblies have additionally been deposited on Zenodo 10.5281/zenodo.16326991, as well as their gene and repeat annotations. Sequencing libraries used for the genome assemblies and annotations, and the remaining sequencing generated in support of this manuscript, are publicly available on NCBI under BioProject PRJNA1437418, and NCBI SRA accessions are provided in Supplementary Data 14. Additional files in support of this manuscript have also been deposited on Zenodo 10.5281/zenodo.16326991. Source data are provided with this paper.

An R script that supports this manuscript can be found at https://github.com/sarahcarey/Cannabaceae_sexChromosomes. A static release can be found at Zenodo 10.5281/zenodo.19376146.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES