Significance
Using a population-genomic approach, we identified copy number variants (CNVs)—stretches of DNA that can be either present, absent, or in multiple copies—displaying parallel signatures of local adaptation across the native and introduced ranges of the invasive weed Ambrosia artemisiifolia. We further identified 16 large CNVs, some associated with ecologically important traits including sex allocation and height, that show strong signatures of selection over space, along with dramatic temporal changes over the past several decades. These results highlight the importance of an often-overlooked form of genomic variation for both local adaptation and rapid adaptation of invasive species.
Keywords: copy number variation, local adaptation, invasive species
Abstract
Adaptation is a critical determinant of the diversification, persistence, and geographic range limits of species. Yet the genetic basis of adaptation is often unknown and potentially underpinned by a wide range of mutational types—from single nucleotide changes to large-scale alterations of chromosome structure. Copy number variation (CNV) is thought to be an important source of adaptive genetic variation, as indicated by decades of candidate gene studies that point to CNVs underlying rapid adaptation to strong selective pressures. Nevertheless, population-genomic studies of CNVs face unique logistical challenges not encountered by other forms of genetic variation. Consequently, few studies have systematically investigated the contributions of CNVs to adaptation at a genome-wide scale. We present a genome-wide analysis of CNV contributing to the adaptation of an invasive weed, Ambrosia artemisiifolia. CNVs show clear signatures of parallel local adaptation between North American (native) and European (invaded) ranges, implying widespread reuse of CNVs during adaptation to shared heterogeneous patterns of selection. We used a local principal component analysis (PCA) to genotype CNV regions in whole-genome sequences of samples collected over the last two centuries. We identified 16 large CNV regions of up to 11.85 megabases in length, eight of which show signals of rapid evolutionary change, with pronounced frequency shifts between historic and modern populations. Our results provide compelling genome-wide evidence that CNV underlies rapid adaptation over contemporary timescales of natural populations.
Understanding how populations adapt and persist in the face of rapid environmental change is one of the most pressing challenges of our time. Fundamental to this goal is determining the genetic basis of adaptive evolution. But despite considerable empirical and theoretical work in this area, many questions remain unresolved. For example, does adaptation typically rely on new and beneficial mutations or on standing genetic variation? Does adaptation generally result in the removal or maintenance of genetic variation affecting fitness? Do mutations contributing to adaptation have uniformly small phenotypic effects, or are large-effect mutations important as well? Do populations exposed to similar environments evolve using the same or different genetic variants?
The evolution of quantitative traits was traditionally thought to almost exclusively depend on evolutionary changes at many polymorphic loci with individually small phenotypic effects (1, 2). However, comparatively recent theoretical and empirical studies demonstrate that large-effect variants can also play important roles in adaptation (3–5). Large-effect mutations are particularly likely to contribute to the initial stages of a population’s evolutionary response to a sudden shift in the environment (6) and to facilitate stable adaptive genetic differentiation among populations connected by migration. Such large-effect variants promote local adaptation by resisting the swamping effects of gene flow (7), including cases in which the alleles carry substantial pleiotropic costs (8).
Genomic structural variants, which include inversions, translocations, duplications, and deletions, are predicted to have both large phenotypic effects and strong potential to contribute to adaptation (9, 10). Chromosomal inversions have a long history of evolutionary study, initially facilitated by classical cytogenetics [e.g., polytene chromosome studies (11, 12)], and more recently through advances in genome sequencing and analysis, which have produced compelling new evidence that inversions often underpin adaptation to environmental change (13–15). These findings have reinvigorated interest in the role of inversions in adaptation, yet other types of structural variation have not garnered the same level of attention.
Several case studies show that copy number variants (CNVs)—structural changes that include gene duplications, deletions, and variation in transposable element abundance—have facilitated adaptation in well-characterized systems such as Drosophila melanogaster (16), Anopheles gambiae (17), and Arabidopsis thaliana (18). Prominent examples include the repeated contribution of CNVs to pesticide resistance (19–22), evident in the parallel evolution of CNVs in the agricultural weed Amaranthus tuberculatus as a response to glyphosate exposure (23), and amplification of cytochrome P450 family genes in Antarctic killifish and fall armyworm populations exposed to toxins (24, 25). These studies demonstrate the immense adaptive potential of CNVs, yet most are candidate-driven analyses that cannot resolve the broader contributions of CNVs to adaptation. Few studies have systematically characterized the genome-wide contributions of CNVs to adaptive divergence across a range of environmental conditions and stresses (13, 26, 27).
Invasive species have several unique features that make them powerful systems for studying adaptation in nature and uncovering its genetic basis. First, because recently introduced populations are likely to be initially maladapted to local conditions in the new range, there is strong scope for rapid adaptive evolution that is observable within decades (28, 29). Second, some plant invasions have extensive documentation in georeferenced herbarium collections that can be phenotyped and sequenced to identify and track evolutionary changes over time—an approach rarely possible in natural populations (15, 30, 31). Third, invasive species typically occupy climatically diverse native and invasive ranges, promoting adaptive evolution to local environmental conditions (28, 29, 32). In particular, those with broad ranges further enable tests of the predictability of evolution, since local adaptation across native and invasive ranges may stem from parallel or unique genetic solutions to similar environmental challenges (15, 32, 33). Moreover, they are well suited for evaluating contributions of CNVs, which have been predicted to be important in invasive species adaptation (34, 35). Yet to our knowledge, this hypothesis has never been tested at a genome-wide scale.
Over the last 200 y, the North American native plant Ambrosia artemisiifolia (common ragweed) has become a widespread pest on all continents except Antarctica (36). This wind-pollinated, outcrossing species produces highly allergenic pollen that accounts for ∼50% of hay-fever cases in Europe (37). It is also an agricultural pest (38), with glyphosate resistance—a phenotype associated with CNVs in other species (19)—reported in some A. artemisiifolia populations (39). Furthermore, climate change is predicted to exacerbate this weed’s impacts, with increased pollen production due to an elongated flowering season (40), as well as reducing the future geographic overlap with key biocontrol agents (41). Previous studies show that A. artemisiifolia has established strong signals of local adaptation to climate across its native range, and in introduced ranges in Europe, Asia, and Australia, with rapid local adaptation following each introduction (15, 33, 41–45). We have previously demonstrated a significant contribution of both SNPs and large-effect structural variants—in the form of chromosomal inversions—to climate adaptation in Europe and North America (15). This raises the question of whether other types of genome structural variants (i.e., CNVs) have similarly contributed to adaptive divergence following A. artemisiifolia’s range expansion.
In this study, we leveraged a large temporally and spatially resolved dataset to investigate the genome-wide contributions of CNVs to local adaptation in A. artemisiifolia. Over 600 whole-genome sequences from individuals collected across the native North American range and the introduced European range, including herbarium samples dating back to 1830, provided a unique opportunity to detect signals of adaptation over space and time. We first investigated whether putative CNVs display signatures of divergent selection in both Europe and North America. By comparing the selection signatures of putative CNVs in each range, we then assessed the degree to which these shared variants evolve in parallel between them. Second, using a subset of 121 phenotyped individuals, we tested for associations between putative CNVs and ecologically important traits, such as flowering time and size. The typically low coverage and error-prone nature of herbarium sequences renders many existing CNV detection methods unsuitable for these data. As such, we developed an approach that combines read depth and local principal component analysis (PCA) methods to genotype large CNV regions in both modern and low-coverage historic samples, enabling us to identify CNVs, estimate temporal changes in their frequencies, and thus directly track CNV evolution in these populations.
Results
CNV Identification.
To identify CNVs, we calculated read depth in nonoverlapping 10 kbp windows (normalized by the average coverage of each individual sample) across 311 modern resequenced genomes spanning both the North American and European ranges of A. artemisiifolia. We defined a window as a CNV if at least 5% of samples had average standardized read depths either greater, or less, than two SD from the window mean across samples. Out of the total 105,175 genomic windows, this resulted in 17,855 candidate CNVs retained for subsequent analyses.
CNV Selection Analysis.
Local adaptation to spatially heterogeneous environments is expected to leave a signal of extreme allele frequency differentiation among populations (46). As our measures of read depth for identifying candidate CNV windows were continuous, we used a QST−FST outlier test to identify windows with population differentiation in excess of neutral expectations. We first calculated an FST distribution of neutral SNPs (Fig. 1 B and D) using the method described by Weir & Cockerham (47). Under neutrality, FST and QST values should have similar distributions, whereas an enrichment of QST values within the upper tail of this null FST distribution provides evidence for local adaptation, with QST outliers representing local adaptation candidates (48).
Fig. 1.

CNV QST values indicate regions of selection in the Ambrosia artemisiifolia genome. (A) QST values of filtered coverage windows in North American populations. Pink windows indicate those above 1% FST threshold shown in B, blue windows indicate top 1% of QST values. (C) QST values of filtered coverage windows in European populations. Pink windows indicate those above 1% FST threshold shown in D, blue windows indicate top 1% of QST values. Bars above manhattan plots indicate merged windows > 300 kbp.
In North America, 1,382 CNV windows exhibited QST values at or above the top 1% threshold of neutral FST values: a 7.7-fold enrichment relative to the neutral expectation that 1% of QST windows will fall within this tail (P < 2.2e−16; binomial test; Fig. 1 A and B). In Europe, 339 CNV windows exhibited QST values exceeding the top 1% of the FST distribution: a 1.9-fold enrichment (Fig. 1 C and D; P < 2.2e−16; binomial test). Using an equation adapted from the McDonald–Kreitman test (49), this excess of outliers is consistent with a true positive rate of 87% for the North American CNV candidate windows, and 47% true positives for European outlier windows (see methods). Of the CNV windows displaying differentiation, 111 were outliers in both ranges (32% of European outliers; P = 1.17e−40, hypergeometric test; Fig. 2A), a highly significant excess indicative of parallelism in the same CNVs subject to divergent selection in both ranges. In contrast, there was no overlap between the top 1% of neutral SNP FST values for the two ranges, suggesting that neutral processes cannot explain the parallelism observed in CNVs.
Fig. 2.
Overlapping QST outliers in both ranges and their potential biological functions. (A) Distribution of QST values of all 10 kbp coverage windows in both North America and Europe. Overlapping windows exceeding the neutral FST threshold of 1% in each respective range are colored in pink. (B) Gene ontology enrichment plot of genes within overlapping outlier QST windows, with biological pathways only retained if represented by 2 or more genes (pink in A).
Variation in recombination rate across the genome may interfere with the identification of signatures of selection (50). To account for potential effects of local recombination rate on patterns of CNV differentiation, we separated QST windows into three recombination rate bins based on the genetic map described in ref. 51: low (< 0.5 cM/Mbp), medium (0.5 to 2 cM/Mbp), and high (> 2 cM/Mbp). When QST−FST analyses were repeated within each recombination rate bin separately, 98.9% and 98.3% of the original QST outliers remained significant in North America and Europe, respectively. This demonstrates the minimal impact of recombination rate on the divergent patterns of read depth within CNV windows in this dataset (SI Appendix, Fig. S2). We also investigated the possibility that nonindependence between 10 kbp windows drives the observed patterns of divergence and repeatability. To do so, we merged windows with correlated variation in read depth (R2 > 0.6) within 1 Mbp of one another. After merging windows, the number of candidate CNVs was reduced from 17,855 to 11,877, with the largest window measuring 11.85 Mbp on chromosome 4. We then repeated the QST−FST analysis on these merged windows. QST values remained elevated relative to neutral FST distributions in both North America and Europe (sixfold and 1.3-fold respectively), with 41 outlier windows shared between ranges: far more than expected by chance (hypergeometric test P = 2.736e−16).
To identify candidate CNVs associated with climate, we estimated correlations between each candidate CNV window and the six bioclimatic variables (SI Appendix, Figs. S5 and S6) that were selected after filtering highly correlated variables from the original 19 WorldClim variables (52). Of the 1,382 significant QST windows in North America, 315 (22.7%) were associated with at least one of the six climate variables, whereas only 12 of 339 significant windows in Europe (3%) correlated with climate variables (SI Appendix, Figs. S5 and S6), suggesting the primary selective forces driving differentiation of CNVs in Europe are likely not climate-related.
To determine the putative biological functions of the adaptation candidates, we used gene ontology (GO) enrichment analyses of annotated genes residing within the outlier QST windows. Candidate CNVs in North America are enriched for biotic and abiotic stress response genes, with significant GO terms including “systemic acquired resistance,” “response to oomycetes,” and “response to freezing” (SI Appendix, Fig. S3A and Table S1). In Europe, significant GO enrichments include the hormonal stress response pathways “response to abscisic acid” and “response to jasmonic acid” (SI Appendix, Fig. S3B and Table S2). Overlapping QST candidates between the ranges exhibited GO term enrichment for the defense response terms “defense response to virus” and “response to abscisic acid” (Fig. 2B and SI Appendix, Table S3).
CNV–Trait Associations.
We tested for relationships between genome-wide CNVs and 29 ecologically important traits phenotyped in 121 of our samples, each reared in a common garden experiment (phenotype data were previously reported in refs. 43 and 53). Eighteen traits (SI Appendix, Table S5) were significantly associated (using a Bonferroni-corrected 0.05 threshold) with normalized read depth in at least one of the 17,855 candidate CNV windows (SI Appendix, Fig. S4). Of these trait-associated windows, 17 and 4 overlapped with QST outlier windows in North America and Europe, respectively. With a more relaxed significance threshold of FDR = 0.05 using the Benjamini–Hochberg method, these overlaps were increased to 76 in North America and 10 in Europe. Of particular interest, two nearby windows on chromosome 14 (h1s14:17180001-17190000, h1s14:18470001-18480000) were associated with flowering onset, dichogamy, and sex allocation (defined as female reproductive biomass/male reproductive biomass; SI Appendix, Fig. S4). These traits display strong latitudinal clines in A. artemisiifolia, with overall earlier flowering, much earlier male flowering compared to female flowering, and female-biased sex allocation occurring at high absolute latitudes (43). These two windows flank an annotated gene, AGL-104, that is linked to pollen production in Arabidopsis (54). Moreover, one of these windows (h1s14:18470001-18480000) is a QST outlier in North America and Europe, suggestive of parallel divergent selection.
Large CNV Region Identification (CNVr).
The continuous measures of read depth in 10 kbp windows, while reliable in modern samples, were inaccurate when using low-coverage historic data. Since we had previously obtained accurate measures of genotypes for large inversions in these historic samples (15), we implemented a similar approach to identify large CNV regions (CNVr) in order to perform temporal comparisons between the historic and contemporary samples. Furthermore, we would expect many CNVs to be larger than 10 kbp. We therefore used the same linkage disequilibrium–based approach as stated above to identify and merge adjacent windows which appeared to be components of a larger CNV. To corroborate the presence of large segregating CNVs within our dataset, we performed a local PCA of genotype likelihoods in 100 kbp windows across the genome. We determined potential segregating structural variants as regions with at least three adjacent windows that were outliers for distortions in local population structure relative to the rest of that chromosome. This resulted in a minimal size cutoff of 300 kbp for CNVrs. As such, we defined CNV regions as those in which merged read-depth windows (greater than 300 kbp) overlapped with at least three adjacent windows exhibiting variation in local population structure. This approach identified 52 candidate CNVrs.
We then genotyped individuals for CNVrs in our population-genomic data, including low-coverage historic samples. To do this, we used a combination of normalized read depth across the genomic location of each candidate CNVr alongside a PCA calculated within that same region to cluster samples into genotypes differing in both read depth and PC1 (Fig. 3). With this approach, we were able to identify distinct clusters corresponding to genotypes in both modern and historic samples for 16 out of 52 CNVrs. These 16 CNVrs, which ranged in size from 0.3 to 11.85 Mbp, accounted for 8.1% of the 17,885 10 kbp windows identified in the contemporary samples, including 22.6% of outlier QST windows in North America and 7.2% of outlier QST windows in Europe. Fifteen of these CNVrs exist as heterozygotes within the diploid reference genome, with haplotype 1 containing presence variants, meaning they can be corroborated with alignments between each haplotype of the diploid assembly (SI Appendix, Fig. S9). The high heterozygosity of CNVrs in the reference was likely due to our genotyping method favoring loci with absence alleles that are common in our samples, yet the presence alleles needed to be found in the reference haplotype in order for the CNV to be identified. This is consistent with the low frequencies (mean = 0.196; range: 0.027 to 0.479) of all CNVr presence variants (SI Appendix, Table S6). GO analysis of annotated genes indicates that the 16 CNVrs were enriched for several biological processes, including “methylglyoxal catabolism,” “peroxisome fission,” “pollen tube adhesion,” and “glyphosate metabolism” (SI Appendix, Table S4).
Fig. 3.

CNVr genotyping of cnv-chr4a. (A) MDS coordinates from local PCA in 100 kb windows across chromosome 4. Values diverging from zero indicate local structure compared with the rest of the chromosome. (B) Distribution of merged window sizes along chromosome 4, with boxes indicating window size on the Y axis and chromosomal position on the X axis. (C) Alignment of chromosome 4 Hap1 and Hap2 of the diploid reference genome. A clear gap is present corresponding to local PCA and merged window locations. (D and E) Regional PC1 against regional read-depth corresponding to merged window identified in B with clustering computed using k-means and manual annotation. (D) represents modern samples and (E) represents historic samples.
The largest CNVr that we identified was cnv-chr4a, which we estimated to be 11.85 Mbp in length. Closer inspection of this region within the diploid reference reveals that large regions on haplotype 1 are absent on haplotype 2 but are interspersed by three smaller inversions (Fig. 3 and SI Appendix, Figs. S8 and S9). Analysis of coverage depth across chromosome 4 of three closely related Ambrosia species sequences mapped to the A. artemisiifolia reference revealed the absence haplotype as the likely ancestral state (SI Appendix, Fig. S10). The derived insertion variant contains an excess of transposable elements (TEs) relative to the remainder of chromosome 4 (87.21% versus 70.48%). TE family Ty1/Copia accounts for 32.45% of the TEs within cnv-chr4a, compared to just 14.87% throughout the rest of this chromosome. Repetitive elements display greater density toward the beginning of this region (SI Appendix, Fig. S11), where gene density is very low. The region toward the end of the cnv-chr4a, which exhibits greater gene density, shows strong synteny with chromosome 2 and aligns with inversions present in the reference alignment (SI Appendix, Figs. S8 and S11). Most of chromosome 4 displays synteny with chromosome 2, suggesting they are homoeologous chromosomes. The largest gap in syntenic blocks corresponds to the gene-depleted and TY1/Copia-enriched region of cnv-chr4a (SI Appendix, Fig. S11). It is therefore likely that this complex structural variant consists of a series of inversions which have subsequently been separated by a large TE expansion. Recombination is likely strongly suppressed within this structural variant, as the coverage windows exhibit strong linkage disequilibrium across the region. Candidate CNV windows within cnv-chr4a are QST outliers (Fig. 1 A and C) and associated with the bioclimatic variable mean diurnal range (SI Appendix, Fig. S5), consistent with the CNVr contributing to local adaptation.
Another noteworthy CNVr, cnv-chr8a, contains an ortholog of the A. thaliana EPSPS locus. This CNVr lies within a large inversion, hb-chr8, previously described in ref. 15. While the frequency of cnv-chr8a is not strongly correlated with the frequency of this inversion (R2 = 0.006), cnv-chr8a presence alleles occur exclusively on the common, and presumably ancestral, orientation of the inverted region.
Spatiotemporal CNVr Modeling.
We used whole genome sequences derived from > 284 historical herbarium samples (dating back to 1830) and generalized linear models (GLMs) to uncover how the 16 CNVr alleles may have changed in frequency over both space and time. These GLMs predicted genotype as a function of range (native or introduced), latitude, and year of collection, and each model was reduced to remove any nonsignificant interactions between these variables. To correct for population structure, we also included the first principal component of genetic variation (calculated from 10,000 neutral SNPs) as a covariate in each model. Twelve of the 16 CNVrs displayed at least one significant predictor variable (time, range, latitude, or interactions). Nine CNVrs exhibited significant associations with latitude (cnv-chr4a, cnv-chr4b, cnv-chr5a, cnv-chr9a, cnv-chr14a, cnv-chr17b, cnv-chr17c, and cnv-chr18a) (SI Appendix, Table S7 and Fig. S12). Models of three CNVrs (cnv-chr10a, cnv-chr14a, and cnv-chr17a) contained significant three-way interactions (Fig. 4 and SI Appendix, Tables S7 and S8 and Fig. S12). Of note, both cnv-chr14a and cnv-chr17a displayed clinal patterns in North America regardless of year, whereas this same latitudinal pattern was present only in modern European samples — strong evidence of clinal reformation following an initial period of maladaptation in the introduced range (Fig. 4 and SI Appendix, Tables S7 and S8). The large cnv-chr4a insertion displays latitudinal clines in modern populations across both ranges, with the insertion at higher frequencies at lower latitudes. In the native range, this appears to be driven by increasing frequencies over time in more central and southern populations of North America (SI Appendix, Table S8). Correspondingly, this CNV overlaps with a previous SNP-based result consistent with a selective sweep in the St. Louis population (15).
Fig. 4.

Spatio-temporal CNV frequency shifts. (A and B) Logistic regression models for (A) cnv-chr14a and (B) cnv-chr17a with error bands representing 95% CI of least-squares regressions of CNV frequency. Time was binned into five categories, ranging from historic (purple) to modern (green). Model information is found in SI Appendix, Table S6. (C and D) Frequency of (C) cnv-chr14a and (D) cnv-chr17a in modern A. artemisiifolia populations.
To further analyze the potential adaptive significance of these CNVrs, we assessed associations between each variant and the same 29 phenotypes analyzed above (43). The cnv-chr10a variant displayed Bonferroni-significant associations with total biomass, root weight, and total number of male inflorescences, while cnv-chr14a was significantly associated with sex allocation (SI Appendix, Table S9).
Discussion
CNVs are increasingly recognized as important in local adaptation (13, 24, 26, 27, 55). However, previous empirical evidence is predominantly limited to examples of pesticide resistance (18, 25) and candidate gene studies (56, 57). Genome-wide analyses of CNVs at the population scale are rare beyond model organisms (55). While many methods exist for identifying CNVs from resequenced genomes (58, 59), these often rely on long reads, or short reads with higher coverage than available for many population genomic datasets, including our historic specimens. We therefore combined local read depth, linkage disequilibrium, and deviations in population structure along the chromosome, to identify CNVs in modest-coverage samples collected over the past 190 y.
We implemented a genome-wide approach to identify CNVs and examined their potential contributions to local adaptation across the native and an invaded range of A. artemisiifolia. CNV windows were enriched for signatures of local adaptation, which occurred disproportionately in parallel between native and invasive ranges. As this signal was not replicated in putatively neutral SNP loci, this implies CNV-driven local adaptation across Europe takes place via standing variation inherited from North America. Furthermore, large CNV regions identified in modern and historical samples, such as cnv-chr14a and cnv-chr17a (Fig. 4), show evidence of rapid local adaptation across the short timescale (ca. 150 y) of A. artemisiifolia’s invasion in Europe.
A. artemisiifolia CNV windows exhibited extensive signals of local adaptation, including elevated geographic divergence in fold coverage, relative to neutral expectations, in both North America and Europe (Fig. 1). These windows contain an overrepresentation of genes involved in abiotic and biotic stress responses in North America (SI Appendix, Fig. S3A and Table S1) and pathogen defense in Europe (SI Appendix, Fig. S3B and Table S2)—consistent with SNP-based FST outlier windows in Europe being enriched for defense-related functions (30). The overall weaker patterns of CNV differentiation in Europe relative to North America are also consistent with previous SNP-based analyses showing fewer differentiated outlier SNP windows in Europe (15). Some CNVs were associated with traits important to local adaptation, including a CNV on chromosome 16 associated with mature plant height (SI Appendix, Fig. S4), which also overlaps an annotated NB-ARC domain. NB-ARC domains, occurring in most plant resistance (R) genes, are involved in nucleotide binding and recognition (60), and duplications of such genes underlie the evolution of resistance to pathogens (61, 62). CNVs also appear to influence phenological traits. For example, the CNV windows on chromosome 14 are associated with flowering time (SI Appendix, Fig. S4). Overall, our data point to important roles of CNVs in the local adaptation of A. artemisiifolia, which aligns with evidence from other species that structural variants widely contribute to adaptation (63–65).
One-third of adaptive CNV windows in Europe were also candidates for local adaptation across North America. Previous work in A. artemisiifolia has revealed similar patterns of repeatability, or parallelism, with respect to SNPs (15, 33), inversions (15), and genes affecting locally adapted traits like flowering time and sex allocation (43). Invasive species are expected to evolve in parallel when responding to analogous selection pressures, as observed in our system and in others, such as Drosophila suzukii and European starling (66, 67). Such parallelism is promoted when standing genetic variation involved in local adaptation in the native range is recruited as a source of adaptive variation within the invasive range (68, 69). Multiple introductions from several genetically diverse source populations from North America to Europe presumably facilitated repeatability by ensuring that most of the important standing variants successfully made the journey to the new range (30, 33). That all CNVrs were present in historical European populations indicates that they were introduced into Europe during the early stages of invasion (SI Appendix, Table S6).
The 16 large CNV regions that we detected using a combination of read depth, linkage disequilibrium, and local population structure analyses (Fig. 3 A and B) contained 22.6% of QST outlier windows from North America and 7.2% of the outlier windows from Europe. Yet these CNVrs comprised only 0.23% of the genome, demonstrating their disproportionate contribution to these signals of local adaptation. Remarkably, 15 of the CNVrs colocalize with segregating presence/absence variants in the highly heterozygous diploid reference, supporting our detection method. Closer analysis of the largest CNVr (cnv-chr4a) within the reference reveals that our detection method may lack sensitivity in fully revealing structural complexity within CNVrs (Fig. 3). Multiple abutting inversions and CNVs within this region appear to segregate together as a single, complex structural variant. Nevertheless, the regions we identified demonstrate the general adaptive potential of structural variants, in which CNVs are predominant features. Since our method of detection for CNVrs was biased toward identifying large CNVrs with presence variants on haplotype 1 of the reference, and our genotyping method was biased toward identifying loci with common absence alleles, our results represent a lower bound on the prevalence and adaptive significance of CNVrs in A. artemisiifolia. Investigations using pangenomics (70) to elucidate a more complete picture of the contribution of CNVs to adaptation are therefore warranted.
Our study goes beyond the traditional population-genomic approach of detecting signals of local adaptation using contemporary samples alone. Our use of historical sequence data also allowed us to track temporal change in CNV frequencies across nearly two centuries—a period that spans the establishment and spread of A. artemisiifolia within Europe (36, 71), and significant environmental upheaval in both ranges, owing to industrialization, agriculture, and climate change (72). CNVs are rarely genotyped in historic genomes because sequence quality is poor (73). However, by focusing only on large CNVs identified using modern data, we were able to confidently assign genotypes in historic samples. This use of modern data to validate historic sequence variant calls is common in temporal genomic studies (30, 74). Eight CNVrs display clear frequency shifts over time, consistent with rapid adaptation over its recent evolutionary history. For example, while cnv-chr14a and cnv-chr17a both exhibit a consistent latitudinal cline in historical and modern North American populations, these clines are only evident in modern European populations (Fig. 4A and SI Appendix, Tables S7 and S8), which is consistent with the rapid cline reformation in Europe following an initial period of postintroduction maladaptation. The cnv-chr14a CNVr is associated with sex allocation (SI Appendix, Table S9), and candidate CNV windows within the region exhibit associations with flowering time, sex allocation, and dichogamy (SI Appendix, Fig. S4), traits which show parallel clines in Europe and North America (43). Furthermore, cnv-chr14a lies within 20 kbp of the AGL-104 gene, which is involved in pollen development in Arabidopsis (54). Flowering time is a complex trait that is affected by diverse forms of genomic variation (75)—our results indicate that SNPs, inversions (15), and CNVs each play important roles in the rapid adaptation of this important trait in A. artemisiifolia populations.
Many well-characterized CNVs in other species are associated with the evolution of pesticide resistance (19, 20, 25). We identified a CNVr (cnv-chr8a) potentially involved in herbicide resistance. An ortholog of the A. thaliana EPSPS locus, the molecular target of glyphosate herbicides, lies within the cnv-chr8a region. CNVs confer resistance to glyphosate in numerous other weed species, where the increased gene expression caused by EPSPS gene amplification ameliorates the herbicide’s toxic effects (19, 56). Glyphosate resistance has been documented in some A. artemisiifolia populations (39), and while we do not know which populations in our study might be glyphosate resistant, this CNVr is a strong candidate for future study.
Previous population-genomic analyses of A. artemisiifolia provide strong evidence that SNPs and putative chromosomal inversions contribute to local adaptation; here, we provide evidence of a similar role for CNVs. Although CNVs are known to have large effects on traits (17), we cannot be sure that the variants we have identified are the direct targets of selection—they may instead be in linkage disequilibrium with other variants that are the actual targets. Assessing relationships between CNVs and nearby SNPs is fraught because CNVs disrupt SNP calling (76). Two CNVrs overlap inversions identified in ref. 15 but are not in strong LD with the inversion genotypes (R2 = 0.006 to 0.01). However, smaller CNVs may exhibit stronger associations with other SVs as part of coadapted gene complexes (77) or neutral hitchhikers. Large insertion–deletion variants can result in hemizygous regions of reduced recombination which may collect and bind together multiple variants (78). Functional assays such as RNA-seq experiments are required to understand the mechanistic effect of these CNVs on traits and fitness (79). Further efforts to determine the functional effects of CNVs, alongside greater sample sizes, would also help uncover the likely epistatic interactions between CNVs and other adaptive variants. The potential existence and nature of these interactions are pertinent questions in evolutionary biology lacking empirical investigations on genome-wide scales (80). Future work should also consider the role of other forms of genomic variation, for example, transposable element abundance and genome size which have been linked with local adaptation and aggressive range expansion (75), alongside investigating the roles of CNVs in biotic interactions, such as pathogen response.
Our study highlights the importance of copy number variation (CNV) in the evolution of a widely distributed and rapidly adapting invasive weed. While CNVs have previously been implicated in adaptation in response to specific selection pressures in other species (23, 24), our genome-wide discovery approach was able to identify candidate genomic regions that are more broadly representative of the contribution of CNVs to local adaptation. We have linked several of these candidates with traits ranging from flowering time to pathogen resistance. Along with previous studies showing that SNPs and chromosomal inversions underlie local adaptation during A. artemisiifolia’s expansion across vast environmental gradients, these new findings make it clear that CNVs account for a significant and previously unrecognized component of this plant’s past success and are consequential for its invasive capacity wherever it may be introduced in the future.
Methods
Samples and Alignments.
Analyses were conducted on 613 whole-genome A. artemisiifolia sequences described in refs. 30, 81–84 and a chromosome-level, phased, diploid A. artemisiifolia genome assembly (15, 85, 86). Alignments and SNP calls against the primary haplotype of the diploid reference (haplotype 1) were generated by Battlay et al. (15, 85, 86), using the Paleomix v1.2.13.4 (87) pipeline and GATK UnifiedGenotyper v3.7 (88). Modern and historic samples were sequenced from across the species’ native North American (155 modern and 92 historic samples) and introduced European (156 modern and 192 historic samples) ranges. Modern samples were collected between 2014 and 2018 and sequenced to a mean coverage of 2.9×. Historic samples were sequenced from herbarium samples collected between 1830 and 1973 with a median collection date of 1905 (SI Appendix, Fig. S1) and sequenced to a mean coverage of 1.4×. Present-day population samples, whose geographic coordinates were recorded during sampling, included between n = 1 to n = 10 individuals; however, populations where n = 1 were removed from analyses requiring population-level information. Historic individuals were grouped into populations according to their proximity (15, 30). Cases where only one sample was obtained from a geographic location were excluded from analyses in which population information was a requirement. Additionally, we aligned sequences of three other Ambrosia species (30) to the primary haplotype of the diploid reference using the Paleomix pipeline, as described in refs. 15, 85, 86.
Depth of Coverage Analysis.
In order to identify CNV within our resequenced common ragweed individuals, we analyzed depth of coverage in nonoverlapping 10 kbp windows using Samtools v1.9 depth (89) on alignment bam files. In the initial analyses of modern samples we only used reads with mapping quality > 30 (-Q 30). Subsequent read-depth analyses of historic samples used reads with mapping quality > 5 (-Q 5) in order to accommodate their poorer mapping quality. Each window was then normalized by dividing window depth by the genome-wide coverage for the sample. To apply a population frequency-based filter to this dataset, we kept only windows which had at least 5% of samples greater than or less than 2 SD from the population mean. This filtering procedure resulted in 17,855 of 105,175 windows (5.8%) being classified as CNVs.
QST−FST Analysis.
Genomic loci that have responded to spatially heterogeneous selection are expected to show elevated differentiation among populations, relative to neutrally evolving loci. SNPs associated with local adaptation can be detected as outliers of genome-wide scans of FST or similar statistics (47, 90, 91). However, unlike SNP data, candidate CNVs have been identified by depth of coverage represented on a continuous scale. As such, coverage can be viewed as a phenotypic measurement that is analogous to measurements of continuous traits. Tests for trait responses to divergent selection are often inferred using a QST−FST analysis (92, 93). Theory predicts that QST values for neutrally evolving traits should follow the same distribution as FST for neutrally evolving loci (48). We used coverage data to measure QST values for each window (analyses were conducted separately for the European and North American ranges), adapting the relationship between QST and population variation from Whitlock (48):
| [1] |
where VA, among represents the among-population variation in coverage for a given window, and VA, within represents within-population coverage variation.
We performed linear mixed models, using the lme4 package in R (94) on populations from each range to extract variance components attributed to within- and among-population variation for each coverage window. Unlike other analyses performed in this study, we used unnormalized coverage for each window in the model, and accounted for variation in individual sample depth by including individual median coverage as a covariate in the model, and population was included as a random effect. The variance among populations (VA, among) was extracted as the variance component attributed to the population, and within-population variance (VA, within) corresponded to the model’s residual variance component.
To identify QST values with divergence in excess of neutral expectations, the distribution of QST values in each range was compared to the distribution of neutral FST values from the same populations. FST distributions were calculated in VCFtools (95) using 10,000 putatively neutral and independently segregating LD-pruned SNPs, randomly sampled from outside both genic regions and known structural variants (15). Outliers were classified as windows with QST values greater than 99% of neutral FST values. Under neutrality, 1% of QST values are expected to fall above this 99% threshold, and we therefore tested whether there was a significant excess of windows above this null expectation. To identify the rate of true positives, we used the following calculation:
| [2] |
where n represents the total number of windows analyzed, M represents the total number of outliers, m represents the number of true positives, and therefore, M−m represents the number of false positives. Analyses were performed using modern data from North America and European ranges separately. We used a hypergeometric test to identify significant overlaps in both QST and FST loci between both ranges.
Environmental–CNV Associations.
Correlations between genetic variation and environment provide further evidence for local adaptation on the genomic level. CNV–environment associations were identified by measuring Spearman’s correlation coefficients (Rho) between normalized coverage of each 10 kbp window and climate variables. Climatic variables were extracted from WorldClim (52) for each of the geographic coordinates of our sample of individuals. We excluded highly correlated variables (R2 > 0.7) from the analysis, resulting in six variables: BIO1: annual mean temperature, BIO2: mean diurnal range, BIO8: mean temperature of wettest quarter, BIO9: mean temperature of driest quarter, BIO12: annual precipitation, and BIO15: precipitation seasonality. Outliers were identified as the top 1% of the empirical P-value of the Rho distribution. Coverage windows that were in the 1% tails of the CNV–environment distribution and outliers in our QST analysis were considered CNV–climate adaptation candidates.
Identifying Large CNV Regions.
Since larger CNVs will affect coverage in multiple adjacent windows, we used a linkage disequilibrium–based approach to merge nearby windows. We merged windows within 1 Mbp (and all windows within this region) that had sample depths that were correlated at R2 > 0.6. By generating heatmaps of LD between CNV windows, we visualized the presence of correlated read depth, indicative of larger CNVs (SI Appendix, Fig. S7). This resulted in 11,877 windows ranging in size from 10 kbp to 11.85 Mbp.
Local PCA has previously been used to detect population-genomic signatures of chromosomal inversions (15, 63, 96). Here, we employed this method to identify distortions in local relatedness due to CNVs using two modifications. Due to low-coverage and bias in historic samples, we performed these steps initially on modern samples only, including historic samples for genotyping later (see below). First, we calculated local covariance matrices for each window using ANGSD (97) and PCAngsd (98). Second, we did not include a filter for missing data in our ANGSD command, which meant that missingness would cause distortion in the local PCAs. The analysis was performed in nonoverlapping windows of 100 kbp, and multidimensional scaling axes 1-4 (calculated from local PCAs across each chromosome) were examined for blocks of outliers indicative of large structural variants. While individual MDS outlier windows are not necessarily due to structural variations, SVs are the most likely cause of signals that are consistent across adjacent windows. Therefore, we retained candidates that included at least three adjacent windows that were outliers, so that the lower limit of CNVr size was 300 kbp. As such, we identified candidate CNVr as those where MDS candidates overlapped the merged read depth–based CNV windows which were also greater than 300 kbp in length. These 52 candidate CNVrs corresponded to 68% of the merged CNV windows greater than 300 kbp in length.
CNV Region Genotyping.
We attempted to genotype modern and historical samples, separately, for each of the 52 candidate CNVr. To do so, we performed another local PCA across the length of each candidate CNVr using PCAngsd (98). We then compared PC1 of the candidate CNVr against the average normalized read depth across that CNVr to identify whether individuals clustered by read depth and PC1. To genotype these samples according to these clusters, we used k-means clustering from the ClusterR package in R (99), followed by manual annotation of genotypes. Samples appeared to either cluster into groups of two or three genotypes. CNVr candidates displaying two clear genotypes were likely to have one rare homozygote, or the heterozygote and one homozygote class were indistinguishable. As such, we assigned k-values of either 2 or 3 depending on whether visual inspection indicated the presence of two or three segregating genotypes. To test for association between CNVrs overlapping or neighboring chromosomal inversions, we used the cor function in R (R team) to calculate correlations between CNVr and the genotypes of overlapping inversions that were previously identified by Battlay et al. (15). To visualize CNVrs that were heterozygous between haplotypes of the diploid reference genome, we aligned both reference haplotypes using minimap2 v2.1.8 (-k19 -w19 -m200) (100) and generated dotplots of the alignments. To call segregating SVs in the diploid reference, we aligned both haplotypes using nucmer (-maxmatch -c 100 -b 500 -l 50) within the mummer v3.23 software package (101). We then used SyRI (102) to identify and plotsr (-s 300000) (103) to visualize structural variants greater than 300 kbp in length. This method of identifying large segregating CNVs is constrained in that it can only identify CNVrs that contain a presence variant on haplotype 1 of the reference. Since haplotype 1 is larger than haplotype 2, it is more likely to contain more presence variants. Second, to create a strong signature of divergence within the local PCA, a large number of individuals must contain the absence variant, meaning that presence variants will also be rare. Furthermore, closer inspection of segregating CNVrs within the reference indicated that some CNVrs may consist of complex SVs, containing inversions and translocations, which our identification method using WGS failed to identify.
Since cnv-chr4a was surprisingly large and segregating in our diploid reference, we conducted further analyses to understand its genomic makeup. First, we analyzed TE content within this region. We identified TEs using EDTA (104) and used RepeatMasker v4.1.1 (105) to obtain a summary of various TE families within this region relative to the rest of chromosome 4. We then performed a synteny analysis to determine whether the small number of genes in that region were collinear with other genomic regions and were therefore an ancestral arrangement or whether they were novel combinations of genes. We used McScanX v97e74f4 (106) to determine syntenic gene groups resulting from a self-alignment of protein sequences on haplotype 1 using blastp (-evalue 1e−10) in BLAST v2.7.1 (107). We calculated read depth on chromosome 4 for three aligned Ambrosia outgroup species (30). Calculating the mean read depth exclusively within the cnv-chr4a region allowed us to ascertain the likely ancestral state of this structural variant.
Temporal CNV Changes.
To analyze shifts in CNV frequency over both time and space, we used GLMs with the glm function in R (108) to predict presence/absence counts in populations from range (Europe or North America), sample year, and sample latitude, including significant interactions between the three variables. PC1, which was calculated from a covariance matrix of 10,000 SNPs randomly sampled from outside of genes and putative inversions identified by Battlay et al. (15), was included in each model to control for population structure. Model significance was tested with the anova function using a type-3 test [car v3.1-2 package (109)]. The models were reduced in a stepwise fashion, removing nonsignificant interactions until all remaining interactions (if any) were significant (P < 0.05). The emtrends function [emmeans v1.10.2 package (110)] was used to test directionality and obtain CI within interacting predictors.
CNV–Trait Associations.
To investigate associations between CNVs and traits, we measured associations between 29 traits [in 121 samples for which trait data were available (43, 53)] and the coverage windows described above. For each window, we fit linear models, using lm in R (108), between each sample’s normalized window depth and each trait value. To account for population structure, the first principal component of neutral covariance was added to the model. As previously, this was obtained by calculating a covariance matrix on 10,000 putatively neutral sites that were outside gene and inversion regions using PCangsd. A PCA was performed on this covariance matrix using prcomp in R (108). We assessed significance using a Bonferroni-corrected significance threshold: 0.05 divided by the number of windows tested (17,855). We performed CNVr-trait associations using EMMAX [v.beta-7Mar2010 (111)] to identify associations between the CNV regions, using a covariance matrix consisting of the same 10,000 putatively neutral SNPs in previous analyses in this study to correct for population structure.
Gene Ontology Analyses.
Gene ontology enrichment analyses were performed using the R package topGO (112), using GO terms from A. thaliana TAIR 10 BLAST results. Fisher’s exact test, using a significance threshold of P < 0.05, was used to identify GO terms enriched within candidate gene lists relative to QST outliers, as well as genes within CNVr.
Supplementary Material
Appendix 01 (PDF)
Acknowledgments
We thank Stephen Wright for his advice in the early stages of this project, and we thank Keyne Monro for her advice and expertise in developing the QST−FST analysis. We also thank Claire Mérot, an anonymous reviewer, and the editor for their thoughtful comments which greatly improved the manuscript.
Author contributions
J.W., T.C., M.D.M., P.B., and K.A.H. designed research; J.W., V.C.B., L.v.B., T.C., P.B., and K.A.H. performed research; J.W. contributed new reagents/analytic tools; J.W. and P.B. analyzed data; with advice from T.C. and K.A.H.; and J.W. wrote the paper with contributions from all authors.
Competing interests
The authors declare no competing interest.
Footnotes
This article is a PNAS Direct Submission. D.O.-B. is a guest editor invited by the Editorial Board.
Data, Materials, and Software Availability
All code for analyses performed in this work have been deposited in GitHub (113). The haplotype 1 gene annotation GFF file is available from Figshare (114). Phenotype data is available from Figshare (53). Previously published data were used for this work (NCBI: PRJNA819156 (85); NCBI: PRJNA929657 (86); NCBI: PRJNA929658 (81); ENA: PRJEB48563 (82); ENA: PRJNA339123 (83); ENA: PRJEB34825 (84)).
Supporting Information
References
- 1.Fisher R. A., XV.—The correlation between relatives on the supposition of mendelian inheritance. Earth Environ. Sci. Trans. R. Soc. Edinb. 52, 399–433 (1919). [Google Scholar]
- 2.Pritchard J. K., Pickrell J. K., Coop G., The genetics of human adaptation: Hard sweeps, soft sweeps, and polygenic adaptation. Curr. Biol. 20, R208–R215 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Van’t Hof A. E., et al. , The industrial melanism mutation in British peppered moths is a transposable element. Nature 534, 102–105 (2016). [DOI] [PubMed] [Google Scholar]
- 4.Jensen A. J., et al. , Large-effect loci mediate rapid adaptation of salmon body size after river regulation. Proc. Natl. Acad. Sci. U.S.A. 119, e2207634119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Connallon T., Hodgins K. A., Allen Orr and the genetics of adaptation. Evolution 75, 2624–2640 (2021). [DOI] [PubMed] [Google Scholar]
- 6.Orr H. A., The population genetics of adaptation: The distribution of factors fixed during adaptive evolution. Evolution 52, 935–949 (1998). [DOI] [PubMed] [Google Scholar]
- 7.Yeaman S., Whitlock M. C., The genetic architecture of adaptation under migration-selection balance. Evol. Int. J. Org. Evol. 65, 1897–1911 (2011). [DOI] [PubMed] [Google Scholar]
- 8.Battlay P., Yeaman S., Hodgins K. A., Impacts of pleiotropy and migration on repeated genetic adaptation. bioRxiv [Preprint] (2023). Available at: https://www.biorxiv.org/content/10.1101/2021.09.13.459985v5. Accessed 2 June 2024. [DOI] [PMC free article] [PubMed]
- 9.Wellenreuther M., Bernatchez L., Eco-evolutionary genomics of chromosomal inversions. Trends Ecol. Evol. 33, 427–440 (2018). [DOI] [PubMed] [Google Scholar]
- 10.Chakraborty M., Emerson J. J., Macdonald S. J., Long A. D., Structural variants exhibit widespread allelic heterogeneity and shape variation in complex traits. Nat. Commun. 10, 4872 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sturtevant A. H., A case of rearrangement of genes in drosophila. Proc. Natl. Acad. Sci. U.S.A. 7, 235–237 (1921). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Dobzhansky T., Genetics of the Evolutionary Process (Columbia University Press, 1970). [Google Scholar]
- 13.Cayuela H., et al. , Thermal adaptation rather than demographic history drives genetic structure inferred by copy number variants in a marine fish. Mol. Ecol. 30, 1624–1641 (2021). [DOI] [PubMed] [Google Scholar]
- 14.Tepolt C. K., Grosholz E. D., de Rivera C. E., Ruiz G. M., Balanced polymorphism fuels rapid selection in an invasive crab despite high gene flow and low genetic diversity. Mol. Ecol. 31, 55–69 (2022). [DOI] [PubMed] [Google Scholar]
- 15.Battlay P., et al. , Large haploblocks underlie rapid adaptation in the invasive weed Ambrosia artemisiifolia. Nat. Commun. 14, 1717 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Magwire M. M., Bayer F., Webster C. L., Cao C., Jiggins F. M., Successive increases in the resistance of drosophila to viral infection through a transposon insertion followed by a duplication. PLOS Genet. 7, e1002337 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lucas E. R., et al. , Whole-genome sequencing reveals high complexity of copy number variation at insecticide resistance loci in malaria mosquitoes. Genome Res. 29, 1250–1261 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.DeBolt S., Copy number variation shapes genome diversity in Arabidopsis over immediate family generational scales. Genome Biol. Evol. 2, 441–453 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Gaines T. A., et al. , Gene amplification confers glyphosate resistance in Amaranthus palmeri. Proc. Natl. Acad. Sci. U.S.A. 107, 1029–1034 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Daborn P. J., et al. , A single p450 allele associated with insecticide resistance in Drosophila. Science 297, 2253–2256 (2002). [DOI] [PubMed] [Google Scholar]
- 21.Ffrench-Constant R. H., The molecular genetics of insecticide resistance. Genetics 194, 807–815 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Zhang C., et al. , Investigating the mechanisms of glyphosate resistance in goosegrass (Eleusine indica) population from South China. J. Integr. Agric. 14, 909–918 (2015). [Google Scholar]
- 23.Kreiner J. M., et al. , Multiple modes of convergent adaptation in the spread of glyphosate-resistant Amaranthus tuberculatus. Proc. Natl. Acad. Sci. U.S.A. 116, 21076–21084 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Reid N. M., et al. , The genomic landscape of rapid repeated evolutionary adaptation to toxic pollution in wild fish. Science 354, 1305–1308 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Gimenez S., et al. , Adaptation by copy number variation increases insecticide resistance in the fall armyworm. Commun. Biol. 3, 664 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Prunier J., et al. , Gene copy number variations involved in balsam poplar (Populus balsamifera L.) adaptive variations. Mol. Ecol. 28, 1476–1490 (2019). [DOI] [PubMed] [Google Scholar]
- 27.Dorant Y., et al. , Copy number variants outperform SNPs to reveal genotype–temperature association in a marine species. Mol. Ecol. 29, 4765–4782 (2020). [DOI] [PubMed] [Google Scholar]
- 28.Hodgins K. A., Bock D. G., Rieseberg L. H., “Trait evolution in invasive species” in Annual Plant Reviews Online, (John Wiley & Sons Ltd, 2018), pp. 459–496. [Google Scholar]
- 29.Colautti R. I., Barrett S. C. H., Rapid adaptation to climate facilitates range expansion of an invasive plant. Science 342, 364–366 (2013). [DOI] [PubMed] [Google Scholar]
- 30.Bieker V. C., et al. , Uncovering the genomic basis of an extraordinary plant invasion. Sci. Adv. 8, eabo5115 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.McGaughran A., et al. , Genomic tools in biological invasions: current state and future frontiers. Genome Biol. Evol. 16, evad230 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Stern D. L., The genetic causes of convergent evolution. Nat. Rev. Genet. 14, 751–764 (2013). [DOI] [PubMed] [Google Scholar]
- 33.van Boheemen L. A., Hodgins K. A., Rapid repeatable phenotypic and genomic adaptation following multiple introductions. Mol. Ecol. 29, 4102–4117 (2020). [DOI] [PubMed] [Google Scholar]
- 34.Dlugosch K. M., Anderson S. R., Braasch J., Cang F. A., Gillette H. D., The devil is in the details: Genetic variation in introduced populations and its contributions to invasion. Mol. Ecol. 24, 2095–2111 (2015). [DOI] [PubMed] [Google Scholar]
- 35.Estoup A., et al. , Is there a genetic paradox of biological invasion?. Annu. Rev. Ecol. Evol. Syst. 47, 51–72 (2016). [Google Scholar]
- 36.Essl F., et al. , Biological Flora of the British Isles: Ambrosia artemisiifolia. J. Ecol. 103, 1069–1098 (2015). [Google Scholar]
- 37.Schaffner U., et al. , Biological weed control to relieve millions from Ambrosia allergies in Europe. Nat. Commun. 11, 1745 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Cowbrough M. J., Brown R. B., Tardif F. J., Impact of common ragweed (Ambrosia artemisiifolia) aggregation on economic thresholds in soybean. Weed Sci. 51, 947–954 (2003). [Google Scholar]
- 39.Brewer C. E., Oliver L. R., Confirmation and resistance mechanisms in glyphosate-resistant common ragweed (Ambrosia artemisiifolia) in Arkansas. Weed Sci. 57, 567–573 (2009). [Google Scholar]
- 40.Hamaoui-Laguel L., et al. , Effects of climate change and seed dispersal on airborne ragweed pollen loads in Europe. Nat. Clim. Change 5, 766–771 (2015). [Google Scholar]
- 41.Sun Y., Roderick G. K., Rapid evolution of invasive traits facilitates the invasion of common ragweed, Ambrosia artemisiifolia. J. Ecol. 107, 2673–2687 (2019). [Google Scholar]
- 42.van Boheemen L. A., et al. , Multiple introductions, admixture and bridgehead invasion characterize the introduction history of Ambrosia artemisiifolia in Europe and Australia. Mol. Ecol. 26, 5421–5434 (2017). [DOI] [PubMed] [Google Scholar]
- 43.van Boheemen L. A., Atwater D. Z., Hodgins K. A., Rapid and repeated local adaptation to climate in an invasive plant. New Phytol. 222, 614–627 (2019). [DOI] [PubMed] [Google Scholar]
- 44.Gorton A. J., Moeller D. A., Tiffin P., Little plant, big city: A test of adaptation to urban environments in common ragweed (Ambrosia artemisiifolia). Proc. R. Soc. B Biol. Sci. 285, 20180968 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.McGoey B. V., Stinchcombe J. R., Introduced populations of ragweed show as much evolutionary potential as native populations. Evol. Appl. 14, 1436–1449 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Lewontin R. C., Krakauer J., Distribution of gene frequency as a test of the theory of the selective neutrality of polymorphisms. Genetics 74, 175–195 (1973). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Weir B. S., Cockerham C. C., Estimating F-statistics for the analysis of population structure. Evolution 38, 1358–1370 (1984). [DOI] [PubMed] [Google Scholar]
- 48.Whitlock M. C., Evolutionary inference from QST. Mol. Ecol. 17, 1885–1896 (2008). [DOI] [PubMed] [Google Scholar]
- 49.McDonald J. H., Kreitman M., Adaptive protein evolution at the Adh locus in Drosophila. Nature 351, 652–654 (1991). [DOI] [PubMed] [Google Scholar]
- 50.Booker T. R., Yeaman S., Whitlock M. C., Variation in recombination rate affects detection of outliers in genome scans under neutrality. Mol. Ecol. 29, 4274–4279 (2020). [DOI] [PubMed] [Google Scholar]
- 51.Prapas D., et al. , Quantitative trait loci mapping reveals an oligogenic architecture of a rapidly adapting trait during the European invasion of common ragweed. Evol. Appl. 15, 1249–1263 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Fick S. E., Hijmans R. J., WorldClim 2: New 1-km spatial resolution climate surfaces for global land areas. Int. J. Climatol. 37, 4302–4315 (2017). [Google Scholar]
- 53.van Boheemen L., Atwater D. Z., Hodgins K. A., phenotype Raw, heterozygosity, population structure and climate data. Figshare. 10.6084/m9.figshare.7217744.v1. Deposited 10 February 2025. [DOI]
- 54.Adamczyk B. J., Fernandez D. E., MIKC* MADS domain heterodimers are required for pollen maturation and tube growth in arabidopsis. Plant Physiol. 149, 1713–1723 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Schrider D. R., Hahn M. W., Begun D. J., Parallel evolution of copy-number variation across continents in drosophila melanogaster. Mol. Biol. Evol. 33, 1308–1316 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Patterson E. L., Pettinga D. J., Ravet K., Neve P., Gaines T. A., Glyphosate resistance and EPSPS gene duplication: Convergent evolution in multiple plant species. J. Hered. 109, 117–125 (2018). [DOI] [PubMed] [Google Scholar]
- 57.Good R. T., et al. , The molecular evolution of cytochrome P450 genes within and between drosophila species. Genome Biol. Evol. 6, 1118–1134 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Abyzov A., Urban A. E., Snyder M., Gerstein M., CNVnator: An approach to discover, genotype, and characterize typical and atypical CNVs from family and population genome sequencing. Genome Res. 21, 974–984 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Zarate S., et al. , Parliament2: Accurate structural variant calling at scale. GigaScience 9, giaa145 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.van Ooijen G., et al. , Structure-function analysis of the NB-ARC domain of plant disease resistance proteins. J. Exp. Bot. 59, 1383–1397 (2008). [DOI] [PubMed] [Google Scholar]
- 61.Żmieńko A., Samelak A., Kozłowski P., Figlerowicz M., Copy number polymorphism in plant genomes. Theor. Appl. Genet. 127, 1–18 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Dolatabadian A., Patel D. A., Edwards D., Batley J., Copy number variation and disease resistance in plants. Theor. Appl. Genet. 130, 2479–2490 (2017). [DOI] [PubMed] [Google Scholar]
- 63.Todesco M., et al. , Massive haplotypes underlie ecotypic differentiation in sunflowers. Nature 584, 602–607 (2020). [DOI] [PubMed] [Google Scholar]
- 64.Mérot C., et al. , Genome assembly, structural variants, and genetic differentiation between lake whitefish young species pairs (Coregonus sp.) with long and short reads. Mol. Ecol. 32, 1458–1477 (2023). [DOI] [PubMed] [Google Scholar]
- 65.Hämälä T., et al. , Genomic structural variants constrain and facilitate adaptation in natural populations of Theobroma cacao, the chocolate tree. Proc. Natl. Acad. Sci. U.S.A. 118, e2102914118 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Koch J. B., et al. , Population genomic and phenotype diversity of invasive Drosophila suzukii in Hawai‘i. Biol. Invasions 22, 1753–1770 (2020). [Google Scholar]
- 67.Hofmeister N. R., Werner S. J., Lovette I. J., Environmental correlates of genetic variation in the invasive European starling in North America. Mol. Ecol. 30, 1251–1263 (2021). [DOI] [PubMed] [Google Scholar]
- 68.Lin T., Klinkhamer P. G. L., Vrieling K., Parallel evolution in an invasive plant: Effect of herbivores on competitive ability and regrowth of Jacobaea vulgaris. Ecol. Lett. 18, 668–676 (2015). [DOI] [PubMed] [Google Scholar]
- 69.Stuart K. C., Sherwin W. B., Edwards R. J., Rollins L. A., Evolutionary genomics: Insights from the invasive European starlings. Front. Genet. 13, 1010456 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Ebler J., et al. , Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes. Nat. Genet. 54, 518–525 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Chauvel B., Dessaint F., Cardinal-Legrand C., Bretagnolle F., The historical spread of Ambrosia artemisiifolia L. France from herbarium records. J. Biogeogr. 33, 665–673 (2006). [Google Scholar]
- 72.Case M. J., Stinson K. A., Climate change impacts on the distribution of the allergenic plant, common ragweed (Ambrosia artemisiifolia) in the eastern United States. PLOS ONE 13, e0205677 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Kim A. S., et al. , Temporal collections to study invasion biology. Mol. Ecol. 32, 6729–6742 (2023). [DOI] [PubMed] [Google Scholar]
- 74.Alves J. M., et al. , Parallel adaptation of rabbit populations to myxoma virus. Science 363, 1319–1326 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Kreiner J. M., Hnatovska S., Stinchcombe J. R., Wright S. I., Quantifying the role of genome size and repeat content in adaptive variation and the architecture of flowering time in Amaranthus tuberculatus. PLOS Genet. 19, e1010865 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Stranger B. E., et al. , Relative impact of nucleotide and copy number variation on gene expression phenotypes. Science 315, 848–853 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Kirkpatrick M., Barrett B., Chromosome inversions, adaptive cassettes and the evolution of species’ ranges. Mol. Ecol. 24, 2046–2055 (2015). [DOI] [PubMed] [Google Scholar]
- 78.Potente G., et al. , Comparative genomics elucidates the origin of a supergene controlling floral heteromorphism. Mol. Biol. Evol. 39, msac035 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Battlay P., et al. , Structural variants underlie parallel adaptation following global invasion. bioXriv [Preprint] (2024). Available at: https://www.biorxiv.org/content/10.1101/2024.07.09.602765v1. Accessed 12 December 2024.
- 80.Wellenreuther M., Mérot C., Berdan E., Bernatchez L., Going beyond SNPs: The role of structural genomic variants in adaptive evolution and species diversification. Mol. Ecol. 28, 1203–1209 (2019). [DOI] [PubMed] [Google Scholar]
- 81.National Center for Biotechnology Information, PRJNA929658. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 82.National Center for Biotechnology Information, PRJEB48563. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 83.National Center for Biotechnology Information, PRJNA339123. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 84.National Center for Biotechnology Information, PRJEB34825. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 85.National Center for Biotechnology Information, PRJNA819156. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 86.National Center for Biotechnology Information, PRJNA929657. NCBI BioProject. https://www.ncbi.nlm.nih.gov/bioproject?term=PRJNA819156&cmd=DetailsSearch. Accessed 2 June 2023.
- 87.Schubert M., et al. , Characterization of ancient and modern genomes by SNP detection and phylogenomic and metagenomic analysis using PALEOMIX. Nat. Protoc. 9, 1056–1082 (2014). [DOI] [PubMed] [Google Scholar]
- 88.DePristo M. A., et al. , A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat. Genet. 43, 491–498 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Danecek P., et al. , Twelve years of SAMtools and BCFtools. GigaScience 10, giab008 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Coop G., Witonsky D., Di Rienzo A., Pritchard J. K., Using environmental correlations to identify loci underlying local adaptation. Genetics 185, 1411–1423 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Gautier M., Genome-wide scan for adaptive divergence and association with population-specific covariates. Genetics 201, 1555–1579 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Spitze K., Population structure in Daphnia obtusa: Quantitative genetic and allozymic variation. Genetics 135, 367–374 (1993). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Prout T., Barker J. S., F statistics in Drosophila buzzatii: Selection, population size and inbreeding. Genetics 134, 369–375 (1993). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Bates D., Mächler M., Bolker B., Walker S., Fitting Linear Mixed-Effects Models Using lme4. J. Stat. Softw. 67, 1–48 (2015). [Google Scholar]
- 95.Danecek P., et al. , The variant call format and VCFtools. Bioinformatics 27, 2156–2158 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Li H., Ralph P., Local PCA shows how the effect of population structure differs along the genome. Genetics 211, 289–304 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Korneliussen T. S., Albrechtsen A., Nielsen R., ANGSD: Analysis of next generation sequencing data. BMC Bioinformatics 15, 356 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Meisner J., Albrechtsen A., Inferring population structure and admixture proportions in low-depth NGS data. Genetics 210, 719–731 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Mouselimis L., et al. , ClusterR: Gaussian mixture models, K-means, mini-batch-Kmeans, K-medoids and affinity propagation clustering (2023), Deposited 4 December 2023.
- 100.Li H., Minimap2: Pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Kurtz S., et al. , Versatile and open software for comparing large genomes. Genome Biol. 5, R12 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Goel M., Sun H., Jiao W.-B., Schneeberger K., SyRI: Finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol. 20, 277 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Goel M., Schneeberger K., plotsr: Visualizing structural similarities and rearrangements between multiple genomes. Bioinformatics 38, 2922–2926 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Ou S., et al. , Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol. 20, 275 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Smit A. F. A.Hubley R.Green P., RepeatMasker Open-4.0. 2013–2015 http://www.repeatmasker.org.
- 106.Wang Y., et al. , MCScanX: A toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 40, e49 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Camacho C., et al. , BLAST+: Architecture and applications. BMC Bioinformatics 10, 421 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.R Foundation for Statistical Computing, Vienna, Austria, R: A language and environment for statistical computing. (2024), Deposited 2024.
- 109.Fox J., Weisberg S., An R Companion to Applied Regression, (Sage, 3 ed., 2019). [Google Scholar]
- 110.Lenth R. V., et al. , emmeans: Estimated Marginal Means, aka Least-Squares Means. (2024), Deposited 23 January 2024.
- 111.Kang H. M., et al. , Variance component model to account for sample structure in genome-wide association studies. Nat. Genet. 42, 348–354 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Alexa A., Rahnenführer J., Gene set enrichment analysis with topGO. Bioconductor Improv. 27, 1–26 (2009). [Google Scholar]
- 113.Wilson J., ragweed_cnv. GitHub. https://github.com/jonrobwil/ragweed_cnv. Deposited 9 February 2025).
- 114.Battlay P., Ambrosia artemisiifolia annotations GFF. Figshare. 10.6084/m9.figshare.19672710.v1. Deposited 9 February 2025. [DOI]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Data Availability Statement
All code for analyses performed in this work have been deposited in GitHub (113). The haplotype 1 gene annotation GFF file is available from Figshare (114). Phenotype data is available from Figshare (53). Previously published data were used for this work (NCBI: PRJNA819156 (85); NCBI: PRJNA929657 (86); NCBI: PRJNA929658 (81); ENA: PRJEB48563 (82); ENA: PRJNA339123 (83); ENA: PRJEB34825 (84)).

