Abstract
Premise
Current phylogenies of Amaranthaceae sensu stricto (s.s.) are inadequately sampled and resolved to reflect the entire evolutionary history of the lineage, which is likely complex due to at least three whole‐genome duplication events, occasionally followed by subsequent additional polyploidization events and rapid diversification of individual sublineages. We designed a new target enrichment bait set to overcome these challenges when reconstructing a phylogeny and demonstrated its applicability to the entire Amaranthaceae s.s. lineage.
Methods
We analyzed 12,775 orthologous and low‐copy genes from a previous comprehensive transcriptomic study for marker selection. Following a newly developed approach that allows the selection of long exons and thus avoids the assembly of chimeric loci, we selected 1000 orthologous exons for phylogenomic analyses.
Results
Our in vivo application showed a high locus recovery rate across all major clades of Amaranthaceae s.s., generated a robust phylogenetic tree, and clarified previously ambiguous relationships of the genera Bosea and Charpentiera. Gene tree conflict analysis revealed mainly high levels of gene tree concordance within the lineage, with a few notable exceptions.
Discussion
The Amaranthaceae1000 kit will provide the basis for a phylogenetic tree across the Amaranthaceae s.s., facilitating future studies on systematics, diversification, and genome evolution within this economically important lineage.
Keywords: Amaranthaceae, gene tree conflict, orthology inference, target enrichment, taxon‐specific bait set, whole‐genome duplications
The advent of next‐generation sequencing (NGS) technologies revolutionized the field of phylogenomics by enabling millions of individual sequencing reactions to be performed simultaneously, dramatically increasing throughput (Shendure and Ji, 2008). Although NGS made whole‐genome sequencing more affordable and efficient, it also introduced data analysis and storage challenges (Batley and Edwards, 2009). For phylogenomic studies, especially of plant groups with large and polyploid genomes, sequencing entire genomes often remains unfeasible, and focusing on a set of preselected genomic regions of interest is a more cost‐effective and convenient approach. Therefore, target‐enriched or hybridization‐based sequencing methods have become more popular in botanical research in recent decades (Mamanova et al., 2010). Another major advantage of target‐enriched approaches is their suitability for fragmented DNA, as they allow for the use of herbarium material, a valuable resource, especially for taxa that are difficult to collect because they are rare or grow in remote regions (Hale et al., 2020). While “universal” bait sets like Angiosperms353 have been widely applied in angiosperms (Johnson et al., 2019; Zuntini et al., 2024), numerous studies have illustrated their limitations (e.g., Lee et al., 2021; Yardeni et al., 2022; Haigh et al., 2023; Helmstetter et al., 2025) and underscored the need for more specific bait sets tailored to particular taxonomic groups. The level of taxonomic specificity varies, ranging from the family level (Nikolov et al., 2019; Christe et al., 2021; Eserman et al., 2021) to the genus or even species level (Bogarín et al., 2018; Villaverde et al., 2018).
Selecting loci for target enrichment is a crucial step in phylogenomic studies. Many studies aim to select single‐copy nuclear genes, such as ribosomal or housekeeping genes (Eserman et al., 2021; Acha and Majure, 2022; Timilsena et al., 2022), to avoid conflicts with paralogous copies in downstream analyses. However, gene duplications are widespread in plants, and many angiosperm families tend to retain duplicated genes that evolve independently (Li et al., 2016). In such cases, the history of these genes may not accurately reflect the species’ history, leading to either false or biased topologies. It is therefore essential to identify orthologous genes that are the product of speciation events and share a common ancestor. In contrast, paralogous genes arising from gene duplication events can introduce misleading signals into phylogenetic analyses (Fitch, 1970). Hence, high‐copy and paralogous loci should be excluded when selecting loci for bait design.
Orthology can be inferred using two main approaches: tree‐based methods and graph‐based methods. The tree‐based approach uses phylogenetic tree topologies to infer orthology from transcriptomic or genomic data (Gabaldón, 2008). In contrast, the graph‐based approach relies on pairwise sequence comparisons, typically performed using all‐against‐all BLAST searches. In doing so, orthology is inferred based on sequence similarity, as orthologs tend to have the highest sequence similarity between species (Gabaldón and Koonin, 2013). Similarity‐based approaches are the most commonly used in various pipelines for designing baits. MarkerMiner (Chamala et al., 2015), for example, performs a reciprocal BLAST search against a selected reference to identify putative orthologs. CAPTUS (Ortiz et al., 2024) employs a clustering approach using MMseqs2 (Many‐against‐Many sequence searching) (Steinegger and Söding, 2017). Other methods, such as Hyb‐Seq (Weitemier et al., 2014) and the approach described by Folk et al. (2015), compare putative orthologs to databases of known orthologs or single‐copy loci. These approaches limit the number of putative loci to those in the reference database, which may overlook lineage‐specific orthologs not included in the reference. Furthermore, sequence similarity–based methods can be prone to errors due to incomplete and heterogeneous datasets caused by lineage‐specific gene loss or duplication events, leading to misinterpretations (Koonin, 2005). Therefore, tree‐based methods are considered more reliable for determining orthologs, as they adhere strictly to the formal definition of orthology (Gabaldón, 2008).
Several tools for tree‐based orthology inference exist. Most pipelines first construct homolog gene trees and then detect subtrees with only one single sequence per taxon in a second step (Chiu et al., 2006; Dunn et al., 2013). A comparable tree‐based approach from Yang and Smith (2014), however, provides different algorithms to prune orthologous subtrees from homologous gene trees: maximum inclusion (MI), rooted ingroups (RT), and monophyletic outgroups (MO). The advantage of these methods is that they do not require a reference database and can therefore be applied to non‐model organisms (Yang and Smith, 2014).
Amaranthaceae sensu stricto (s.s.) is a well‐supported clade within the broader circumscription of Amaranthaceae sensu lato (s.l.) (Huang et al., 2020; Morales‐Briones et al., 2021; Xu et al., 2024). This group comprises approximately 900 species in 81 genera, which are further organized into five well‐supported monophyletic tribes: Achyranthoids, Aervoids, Amaranthoids, Celosioids, and Gomphrenoids (Hernández‐Ledesma et al., 2015; Hammer et al., 2019; Di Vincenzo et al., 2025). The origin of the family remains unknown, and a recent study unveiled that the five tribes probably originated simultaneously via a rapid radiation event (Morales‐Briones et al., 2021). Bosea L. and Charpentiera Gaudich., which have been placed as sister taxa to the rest of the clade (Müller and Borsch, 2005), have a widely disjunct distribution. Bosea can be found in Macaronesia, the eastern Mediterranean, and the Himalayas, while Charpentiera occurs in Hawaii and French Polynesia (Kadereit et al., 2003; Di Vincenzo et al., 2018). Previous phylogenies of Amaranthaceae s.s. relied mostly on a few plastid (trnL‐F, rpl16, trnK, matK) or nuclear loci (ITS, A36, G3PDH, waxy), and therefore displayed numerous poorly resolved relationships, especially on the genus level (Müller and Borsch, 2005; Sánchez‐Del Pino et al., 2009, 2012; Hammer et al., 2015; Bena et al., 2017; Waselkov et al., 2018; Limarino and Borsch, 2020; Bena et al., 2024). To date, there are only a few studies that are based on numerous loci; these include a plastome‐based phylogeny for the Australian genus Ptilotus R. Br. (Hammer et al., 2019) and two phylogenies encompassing the whole Amaranthaceae s.l. family (with an incomplete sampling of Amaranthaceae s.s.) based on plastome (Xu et al., 2024) and transcriptomic data (Morales‐Briones et al., 2021). Given the lack of a robust phylogenomic framework, combined with the important role of Amaranthaceae s.s. as both crop plants (Joshi and Verma, 2020; Aderibigbe et al., 2022) and noxious weeds (Bayón, 2022; Roberts and Florentine, 2022), there is a clear need for a deeper understanding of the evolution and systematics of this lineage.
The genetic complexity of Amaranthaceae s.s. presents significant challenges for phylogenomic studies. Genome sizes of taxa in the clade vary from 1 C = 0.48 pg (Amaranthus palmeri S. Watson; Bennett and Smith, 1976) to 1 C = 4.9 pg (Celosia whitei W. F. Grant; Nath et al., 1992), and the polyploidy level can reach up to 12× in C. whitei (Nath et al., 1992). Furthermore, three whole‐genome duplication (WGD) events have been identified in the backbone of the clade (Yang et al., 2018), adding to the complexity of its evolutionary history. Given that a target enrichment approach is particularly effective for taxa with large and/or polyploid genomes and has been successfully applied before (e.g., Morales‐Briones et al., 2021; Mendez‐Reneau et al., 2023; Walden et al., 2024), we employed the same method for Amaranthaceae s.s.
Here, we present a new workflow for the design of a taxon‐specific bait set for Amaranthaceae s.s., called Amaranthaceae1000. This workflow utilizes a tree‐based orthology inference approach, ensuring accurate identification of orthologous loci without relying on reference databases. In addition, our method prioritizes the selection of long loci, reducing the risk of chimeric gene assembly. We show that these newly developed nuclear markers enable a lineage‐wide application, successfully overcoming the technical challenges posed by the genomic complexity of Amaranthaceae s.s. With this approach, we aim to provide a robust phylogenomic framework that will serve as a prerequisite to gain deeper insights into the spatial and temporal evolution of this economically and ecologically important lineage.
METHODS
Selection of target loci
We extracted 12,775 FASTA files from the ortholog gene trees generated by Morales‐Briones et al. (2021) and filtered out all but the Amaranthaceae s.s. sequences. The orthologs were inferred using the tree‐based MO approach developed by Yang and Smith (2014). To identify exon–intron boundaries, we downloaded the Amaranthus hypochondriacus L. v2.1 genome available on Phytozome v13 (Goodstein et al., 2012). We then masked the introns by extracting all genes and exons from the annotated genome and subtracted them from their respective genes using BEDTools v2.30.0 (Quinlan and Hall, 2010). From the 12,775 orthologs, we proceeded with a subset of 10,612 orthologs that comprised the sequence of A. hypochondriacus. We then aligned the 10,612 orthologs of the Amaranthaceae s.s. taxa using MAFFT v7.490 (Katoh and Standley, 2013) and, in a second step, added the intron‐masked genome of A. hypochondriacus into the existing alignment using the option ‐‐add (Katoh and Frith, 2012).
Exons were split under two conditions: (i) if the intron was present in the genomic reference and (ii) if gaps were observed in the transcriptomes. These criteria ensured the retrieval of complete exons, resulting in 73,262 exons with an average sequence length of 285.67 bp. The scripts used for this process are available at https://github.com/tinakiedaisch/bait_design_from_orthologs (see Data Availability Statement). The exons were filtered by length, and 6168 exons longer than 700 bp were retained. Next, we aimed to remove highly conserved loci by targeting exons with more than 2% parsimony informative sites. Informativeness was assessed using PhyKIT (Steenwyk et al., 2021), resulting in the retention of 5947 exons that exceeded this threshold. Finally, sequences with a gap proportion greater than 0.5 were identified using the get_sequences_gaps_ratio.py script of trimAL v.3.29 (Capella‐Gutiérrez et al., 2009) and removed from the alignments.
We realigned the exons using the OMM_MACSE v12.01 pipeline (Ranwez et al., 2021) to preserve their amino acid translation, remove non‐homologous sequence fragments, and trim extremities. The resulting alignments were again filtered with trimAL under the conditions mentioned above. The remaining exons were filtered in Geneious Prime 2023.2.1 (Biomatters Ltd., Auckland, New Zealand; http://www.geneious.com/), retaining only those with a minimum length of 800 bp, a minimum pairwise identity of 75%, and present in at least 10 sequences per alignment. After the filtering, we manually reviewed the 1894 exons and removed those with excessive gaps and stop codons, resulting in 1216 carefully selected exons for downstream analyses. We ensured that all loci included representatives from both major clades (clade 1: Aervoids, Achyranthoids, and Gomphrenoids; clade 2: Celosioids and Amaranthoids) to guarantee lineage‐wide application of the baits.
Because three WGD events in Amaranthaceae s.s. have been detected (Yang et al., 2018), finding single‐copy loci was unrealistic. To get a general impression of the duplication rates, we thus counted the occurrence of individual taxa in each of the 14,549 isoform‐masked homologous gene trees, which were accessed through the supplementary material of Morales‐Briones et al. (2021).
Deeringia amaranthoides (Lam.) Merr. and Hermbstaedtia glauca (J. C. Wendl.) Rchb. ex Steud. did not undergo a WGD (Morales‐Briones, unpublished work) and were therefore used to identify whether the selected loci belonged to large gene families. Loci with more than three duplications in either species were removed, leaving 1200 exons (in 1163 genes) for the final step of the bait design. We randomly selected 1000 exons to comply with the size restrictions of the 40,000‐bait kit. Summary statistics of the selected exons were performed with AMAS (Borowiec, 2017).
Finally, we wanted to determine whether one sequence (representing either clade 1 with Gomphrenoids, Achyranthoids, and Aervoids or clade 2 with Amaranthoids and Celosioids) or two sequences (representing both clades 1 and 2) should be selected for bait design. To assess how the choice of representative sequences could potentially affect the success of locus recovery, we performed an in‐silico capture analysis using CAPTUS v1.0.1 (Ortiz et al., 2024). For this, we generated three different target files: (i) containing the targeted 1000 loci from species of clade 1 (see above), (ii) containing the targeted 1000 loci from species of clade 2 (see above), and (iii) containing the targeted 1000 loci from species of both clades 1 and 2 representing all major groups in Amaranthaceae s.s. These three target files were used for exon extraction from the 29 Amaranthaceae s.s. transcriptomes of Morales‐Briones et al. (2021). We observed that target files constructed following strategies i and ii resulted in a biased capture efficiency, with reduced efficiency in the clade absent from the target file. In contrast, the target file following strategy iii resulted in equally high capture efficiency in both clades. We therefore followed strategy iii by randomly selecting two sequences with the fewest gaps in each clade for the bait design. Bait synthesis was performed using myBaits at Daicel Arbor BioSciences (Ann Arbor, Michigan, USA). A total of 39,091 biotinylated 120‐nucleotide RNA probes were designed with a 2× tiling strategy.
Taxon sampling
To demonstrate the functionality of the designed baits in vitro, we selected 24 species from Amaranthaceae s.s. (Appendix 1), covering Achyranthoids, Aervoids, Amaranthoids, Celosioids, and Gomphrenoids (Hernández‐Ledesma et al., 2015), as well as Bosea and Charpentiera. The samples were collected from the herbaria CANB, M, MSB, and PERTH (herbarium acronyms per Index Herbariorum [Thiers, 2025]).
From the original Amaranthaceae s.l. dataset of Morales‐Briones et al. (2021), we kept the following 16 samples for the phylogenetic reconstruction to represent the eight major clades: Suaeda divaricata Moq. and Suaeda maritima (L.) Dumort. from the Suaedoideae; Salicornia pacifica Standl. and Tecticornia pergranulata (J. M. Black) K. A. Sheph. & Paul G. Wilson from the Salicornioideae; Haloxylon ammodendron (C. A. Mey.) Bunge ex Fenzl and Kali collinum (Pall.) Akhani & Roalson (syn. Salsola collina Pall.) from the Salsoloideae; Bassia scoparia (L.) A. J. Scott and Eokochia saxicola (Guss.) Freitag & G. Kadereit from the Camphorosmoideae; Chenopodium quinoa Willd. and C. amaranticolor H. J. Coste & A. Reyn. from the Chenopodioideae; Agriophyllum squarrosum (L.) Moq. and Corispermum hyssopifolium L. from the Corispermoideae; Beta macrocarpa Guss. and Hablitzia tamnoides M. Bieb. from the Betoideae; and Nitrophila occidentalis (Moq.) S. Watson and Polycnemum majus A. Braun ex Bogenh. from the Polycnemoideae. In addition, we included 29 transcriptomes of Amaranthaceae s.s., also derived from Morales‐Briones et al. (2021), to assess the accuracy of our sample placements within the phylogenetic framework.
Moreover, transcriptomes from two Achatocarpaceae species, Phaulothamnus spinescens A. Gray and Achatocarpus gracilis H. Walter, which are sister to Amaranthaceae s.l. (Morales‐Briones et al., 2021), along with those of the more distantly related Spergularia media (L.) C. Presl, Dianthus caryophyllus L., Illecebrum verticillatum L., Herniaria latifolia Lapeyr., Corrigiola litoralis L., Mesembryanthemum crystallinum L., Delosperma echinatum (Lam.) Schwantes, Commicarpus scandens (L.) Standl., Mollugo pentaphylla L., and Microtea debilis Sw. from the Caryophyllales were chosen as outgroups.
DNA extraction, library preparation, and sequencing
DNA was extracted from 20–30 mg of leaf material using the DNeasy Plant Mini Kit (QIAGEN, Hilden, Germany) and eluted in 75 µL of 10 mM Tris. DNA quantity was determined with a high‐sensitivity assay kit using the Qubit 4 Fluorometer (Thermo Fisher Scientific, Waltham, Massachusetts, USA), and fragment size was assessed visually via electrophoresis on a 0.8% agarose gel. For long DNA fragments (>1000 bp), 25 µL was sheared to an average size of 350 bp with an M220 Focused‐ultrasonicator with the M220 Holder XTU Insert microTUBE 50 µL (Covaris, Woburn, Massachusetts, USA) prior to library preparation.
One hundred nanograms of DNA was aliquoted in 25 µL of 10 mM Tris and used directly for library preparation. We used the NEBNext Ultra II Kit for Illumina (New England Biolabs, Ipswich, Massachusetts, USA) following the manufacturer's protocol (for half‐volume reactions). After the adapter ligation step, size selection was performed for samples with fragment sizes of 500–1000 bp. The DNA quantity was then measured with the Qubit 4 Fluorometer, and fragment size was assessed with the 4150 TapeStation System using the High Sensitivity D1000 ScreenTape Assay (Agilent Technologies, Santa Clara, California, USA). All libraries were dual‐indexed using NEBNext Multiplex Oligos for Illumina (NEB #E6444, New England Biolabs).
Finished libraries were normalized to 10 nM, and up to 24 samples of similar sizes were pooled together, with the final pool containing 623.99 ng of DNA in 240 µL. The DNA was concentrated to a final volume of 7 µL ddH2O using a vacuum pump. The hybridization followed the myBaits Hybridization Capture for Targeted NGS protocol version 5.02 (Daicel Arbor Biosciences) using the Q5 polymerase. The hybridization took place for 17 h at 62°C. Captured DNA was enriched with a 12‐cycle PCR and a bead clean‐up at a 0.9× ratio. The final DNA concentration and size were measured with the High Sensitivity D1000 TapeStation (Agilent Technologies). Sequencing was carried out at Novogene (Munich, Germany), with a targeted amount of 2.3 Gbp per sample, resulting in 900× coverage of the 2,571,494 bp total reference length.
Sequence assembly
Read quality was assessed with FastQC v0.11.7 (Andrews, 2010), and results were compiled with MultiQC v1.22 (Ewels et al., 2016). We deduplicated the reads using ParDRe v2.2.5 (González‐Domínguez and Schmidt, 2016) and used Trimmomatic v0.39 (Bolger et al., 2014) with the parameters SLIDINGWINDOW:4:5 LEADING:5 TRAILING:5 MINLEN:25 to remove sequencing adapters and low‐quality reads. For the in silico capture in HybPiper2 v2.1.6, we chose BLASTX as a mapping tool (Johnson et al., 2016). For loci extraction, we created two separate target files, each representing one of the two large clades in Amaranthaceae (Clade 1: Aervoids, Achyranthoids, Gomphrenoids; Clade 2: Celosioids, Amaranthoids). The similarity threshold to map the reads was adjusted for each sample according to its proximity to the species in the reference file (‐‐thresh option) and ranged between 75% and 85%. An individual value (‐‐cov_cutoff option) in SPAdes (Prjibelski et al., 2020) was used to adjust the coverage during loci assembly for each sample. The coverage was calculated based on a previous HybPiper run under default coverage settings using a custom R script. This custom assembly ensured a higher quality of the extracted loci and fewer misassemblies.
For de novo chloroplast assembly, we applied Fast‐Plast v1.2.6 (McKain, 2017) using the untrimmed, de‐duplicated reads. As a reference, we used the whole chloroplast genomes of all members of the order Caryophyllales integrated in the program.
Orthology inference and phylogenomic analysis
To include the transcriptomic data of Morales‐Briones et al. (2021), we split the 1000 targeted exons from the orthologous genes described above. Therefore, all major clades of Amaranthaceae s.l., representatives of Amaranthaceae s.s., and their respective outgroups from the Caryophyllales were incorporated.
We inferred orthology from the paralog_no_chimeras output from HybPiper2 following the phylogenomic dataset construction pipeline (Yang and Smith, 2014) with modifications from Morales‐Briones et al. (2021). Because we had included a comprehensive list of outgroup taxa, we followed the MO approach with at least 15 ingroup taxa (out of 53 total) (Yang and Smith, 2014). The individual 972 inferred orthologous loci were aligned using MACSE v2.07 (Ranwez et al., 2018), and columns with more than 70% missing data were trimmed using pxclsq from phyx v.1.3 (Brown et al., 2017). Gene trees were inferred using a maximum likelihood (ML) approach in IQ‐TREE2 v2.3.1 (Minh et al., 2020) with an ultrafast bootstrap approximation (UFBoot) of 1000 (Hoang et al., 2018). Branches with UFBoot support values below 70 were collapsed. The species tree was reconstructed under a coalescent‐based model with ASTRAL IV v1.20.4.6 (Zhang and Mirarab, 2022a) using the ML gene trees. Node support was expressed in local posterior probability values (LPP). In addition, we employed Astral‐Pro3 (ASTRAL for PaRalogs and Orthologs) v1.20.3.6 (Zhang and Mirarab, 2022b) to reconstruct a species tree using all cleaned homologous trees, as it allows for multi‐copy genes.
Gene tree discordance analysis
For gene tree discordance analysis, we used PhyParts v0.0.1 (Smith et al., 2015), a tool that classifies the nodes of gene trees into four categories: (i) supporting the species tree topology, (ii) supporting the main alternative topology, (iii) supporting all other alternative topologies, and (iv) uninformative (including missing data). The conflict can be quantified by calculating each node's internode certainty all (ICA) values. An ICA close to 1 indicates strong concordance, close to 0 indicates equal support for one or more conflicting bipartitions, and a negative ICA indicates discordance.
We rooted the uncollapsed gene trees and the species tree on the outgroups consisting of Commicarpus scandens, Corrigiola litoralis, Delosperma echinatum, Dianthus caryophyllus, Herniaria latifolia, Illecebrum verticillatum, Mesembryanthemum crystallinum, Microtea debilis, Mollugo pentaphylla, and Spergularia media (all Caryophyllales). Because each node could have a different number of gene trees, which may skew the proportion of uninformative gene trees, the PhyParts results were plotted, including informativeness and missingness. The node was treated as uninformative if the UFBoot was <70%. We plotted the results using the Python script available at phypartspiecharts_missing_uninformative.py (https://bitbucket.org/dfmoralesb/target_enrichment_orthology/src/master/).
To further test for confidence, consistency, and informativeness of the phylogeny, we used Quartet Sampling v1.3.1 (QS) (Pease et al., 2018). By resampling the quartet tree counts, three scores were calculated for each internal branch of the phylogeny: (i) the Quartet Concordance (QC) score, which is similar to ICA and quantifies the concordant quartet; (ii) the Quartet Differential (QD) score, which measures the disparity between the sampled proportions of the two discordant topologies and can therefore indicate when one alternative relationship is sampled more frequently than the other; and (iii) the Quartet Informativeness (QI) score, which quantifies the proportion of informative replicates.
As input, we concatenated the trimmed orthologous sequences with AMAS (Borowiec, 2017) and converted them from FASTA to PHYLIP format using pxs2phy from phyx. Before running the QS analysis using IQ‐TREE with 1000 replicates, we rerooted the species tree using pxrr from phyx.
Mapping whole‐genome duplications
Lastly, we wanted to explore whether the already known WGD in Amaranthaceae s.s. (Yang et al., 2018) could be detected in our target‐enriched dataset. To exclude the potential bias from the transcriptomic data, we performed this analysis only on the 24 newly sequenced samples enriched with the Amaranthaceae1000 custom baits, using available scripts from Yang et al. (2018) (https://bitbucket.org/blackrim/clustering).
We extracted the rooted orthogroups (ingroup lineages with genes descended from a single ancestor) from our cleaned homologous trees, setting the MIN_INGROUP_TAXA feature to 10 (extract_clades.py). We then mapped the duplications from each orthogroup onto the species tree (map_dups_concordant.py) and plotted the calculated duplication percentages onto the individual branches (plot_branch_labels.py).
RESULTS
Bait design
The final bait set included 1000 exons from 989 genes and targeted 1,294,760 bp. A minimum of 10 taxa (out of 29 taxa) were represented in all loci. The length of the targeted loci ranged between 800 and 2500 bp, with an average locus length of 1295 bp (Figure 1A). The proportion of parsimony informative sites ranged between 0.21 and 0.45 (average 0.31) and correlated with sequence length (Figure 1B).
Figure 1.

Summary statistics for 1000 loci selected for bait design. (A) Length distribution of the included loci (in base pairs). (B) Scatterplot showing the relationship between locus length (in base pairs) and the number of parsimony‐informative sites. (C) Frequency of species sequence representation per clade in the bait design for the 1000 loci.
The in‐silico capture in CAPTUS was used to determine whether one or two sequences should be used for bait design, and it showed that two sequences resulted in better locus capture across the entire Amaranthaceae s.s. When only sequences from either clade were used, we observed reduced capture efficiency in the non‐sampled clade (Appendix S1, see Supporting Information with this article). Therefore, for each locus, one sequence with the fewest gaps from each clade was randomly selected for bait design (Figure 1C). Here, we should briefly mention that the incomplete and fragmented transcriptome of Nelsia quadrangula (Engl.) Schinz was never selected due to its low quality and poor assembly, but all 28 other species are represented in the sequences for bait design (Figure 1C). Most represented are Deeringia amaranthoides and Hermbstaedtia glauca from the Celosioids, which were newly generated in Morales‐Briones et al. (2021) and had good‐quality transcriptomes, while Amaranthus retroflexus L. was less represented due to its incomplete transcriptome (Figure 1C).
Sequence data and orthology inference
We generated 3.8 to 13 million raw reads per sample, about half of which consisted of PCR duplicates (on average). Between 1.2 and 5.7 million reads were mapped to the target file (9.4–49.2%; average: 36.6% of reads on target). Between 960 and 995 genes with more than 50% of the targeted locus length (average: 982 genes, 98.2%) and between 927 and 983 recovered genes with more than 75% of the targeted locus length (average: 966 genes, 96.6%) were recovered (Figure 2). All 1000 loci contained sequences for which more than 50% of the length was recovered, with only one locus (AH017606_E2) recovered for less than 50% of the length in any of the sampled species (Figure 2A). All clades of Amaranthaceae s.s. showed comparable recovery efficiency at both 50% and 75% of the target locus length: Achyranthoids (988 and 973 genes), Aervoids (982 and 965 genes), Amaranthoids (982 and 963 genes), Celosioids (984 and 973 genes), and Gomphrenoids (985 and 965 genes), respectively (Figure 2B). In addition to the nuclear data, the percentage of chloroplast genes recovered ranged from 3.7% (Arthraerua leubnitziae (Kuntze) Schinz) to 82.7% (Ptilotus mollis Benl), with an average of 43.4% (Appendix S2).
Figure 2.

Locus recovery success. (A) Percentage of the reference length recovered for 24 species of Amaranthaceae s.s. for 1000 loci, ranging from 0–100%. (B) Number of recovered genes reaching 50% (solid line) and 75% (dotted line) of the reference length across the five major clades of Amaranthaceae s.s., as well as the two genera Bosea and Charpentiera.
Out of 972 pruned orthologs, between 606 (in Ptilotus mollis) and 937 (in Bosea yervamora L.) were retrieved per sample (with an average of 729). Trimmed MO orthologs ranged from 756 to 2490 characters (with an average of 1290) and about 2.7% missing data (from 0% to 32%).
Phylogenetic reconstruction
The coalescent ASTRAL tree confirmed the monophyly of Amaranthaceae s.s. within Amaranthaceae s.l. with the former Chenopodiaceae, Betoideae, and Polycnemoideae (LPP = 1.00; Figure 3). Within Amaranthaceae s.s., we recovered five clades, and the relationships between them were resolved with maximum support (LPP = 1.00; Figure 3). These five clades clustered into two major lineages: one containing the Gomphrenoids, the Achyranthoids, and the Aervoids, and the other containing the Amaranthoids and the Celosioids, with the genera Bosea and Charpentiera forming a sister grade to the latter major lineage. The placement of the 24 newly sequenced samples with the Amaranthaceae1000 baits was consistent with the phylogenetic framework of the transcriptomic data (Figure 3). The genera Quaternella Pedersen and Gomphrena L. were recovered as non‐monophyletic, whereas all other genera, albeit not densely sampled, were found to be monophyletic. The newly sequenced Australian species Gomphrena arida J. Palmer was nested within a well‐supported clade (LPP = 1.00) comprising G. vermicularis L. (syn. Blutaparon vermiculare (L.) Mears), G. arida, G. celosioides Mart., and Gossypianthus lanuginosus (Poir.) Moq. (syn. Gomphrena lanuparonychioides T. Ortuño & Borsch). For the first time, the genus Neocentema Schinz could be placed phylogenetically within the Amaranthoids with maximum support (LPP = 1; Figure 3). All results were congruent with the Astral‐Pro3 inference that was performed on all homologs (Appendix S3).
Figure 3.

Coalescent phylogenetic inference based on 57 transcriptomes and 24 samples sequenced with the Amaranthaceae1000 baits. The species tree resulted from an ASTRAL IV analysis of “monophyletic outgroup” (MO) ortholog trees. Species sequenced with the bait set are in bold, followed by their respective laboratory accession numbers. All branches have maximum support, except those marked with an asterisk, with local posterior probabilities (LPP) of 0.9 and 0.98 (from top to bottom). The tree is rooted in members of the Caryophyllales. Pie charts represent results from a PhyParts analysis, with blue indicating concordance, red indicating discordance, green showing the main alternative topology, dark gray representing informativeness, and light gray denoting missing data. Branch values correspond to quartet sampling analysis, showing the quartet concordance (QC) score, quartet differential (QD) score, and quartet informativeness (QI) score, respectively.
Gene tree discordance
The Gomphrenoids and the Achyranthoids showed a high degree of gene tree concordance (ICA = 0.61 and ICA = 0.50, respectively) and a strong QS score (0.95/0.22/1 and 1/–/1, respectively). However, their sister group relationship remained controversial with moderate discordance (ICA = 0.29, QS score 0.41/0.35/0.99), indicating a main alternative topology. The Aervoids had the maximum QS score but suffered from a high level of missing data (ICA = 0.26), whereas the Amaranthoids had high levels of gene tree concordance (ICA = 0.89) and a maximum QS score. The Celosioids showed some gene tree discordance but had a good QS score (ICA = 0.63, QS score 0.95/0/1). The highest levels of gene tree discordance in Amaranthaceae s.s. were found at the base of Amaranthaceae s.s., with Bosea placed as sister to the Amaranthoids and Celosioids, and Charpentiera recovered as sister to both (QS score 0.34/0.1/1 and −0.24/0/0.99, respectively). The analysis yielded a main alternative at both nodes, with Bosea and Charpentiera resolved as sister clades. Nevertheless, the relationship between the Amaranthoids and the Celosioids was highly congruent, as was the monophyly of the two clades.
Whole‐genome duplications
We detected the highest percentage of duplicated genes at the base of clade 1 (Gomphrenoids, Achyranthoids, and Aervoids) with 37.79% (Appendix S4). On the subsequent branches, the proportion of duplicated genes was still elevated, with 14.77% for the Achyranthoids and Gomphrenoids and 10.49% for the Aervoids. A second WGD event was detected at the base of the Amaranthoids with 8.52% duplicated genes. The known WGD event in Alternanthera Forssk. was undetected (Yang et al., 2018).
DISCUSSION
Target enrichment strategies have transformed molecular systematics (Soltis et al., 2013), and the subsequent development of “universal” angiosperm baits (Johnson et al., 2019) has played an important role in resolving the phylogenies of many taxa (Antonelli et al., 2021; Thomas et al., 2021; Giaretta et al., 2022). This success has contributed immensely to the angiosperm tree of life, which includes over 8000 genera and demonstrates the potential of universal baits (Zuntini et al., 2024). However, the relationship cannot be disentangled for complex groups using this universal bait set alone (Lee et al., 2021). To address this drawback, custom baits have become increasingly popular in recent years and have been shown to function at different taxonomic scales (Bogarín et al., 2018; Villaverde et al., 2018; Nikolov et al., 2019; Christe et al., 2021; Eserman et al., 2021). Here, we introduce the Amaranthaceae1000 bait kit, which retrieves up to 1000 low‐copy orthologous nuclear exons for phylogenomic analyses of Amaranthaceae s.s. We show that almost all (98.2%) of the targeted genes can be recovered across Amaranthaceae s.s. and a robust phylogenetic framework can be established.
Marker selection and capture efficiency
The efficiency of the probe set depends highly on the evolutionary distance between the taxa used for bait design and the studied taxa in question (Carlsen et al., 2018; Andermann et al., 2020; Veltman et al., 2024). Thereby, closely related taxa are expected to perform better, and the bait efficiency decreases as sequence similarity decreases in other, less closely related samples. For this reason, we used representatives of all major Amaranthaceae s.s. clades to design the baits (Figure 1C). We observed no difference in the captured genes between the five subclades (Figure 2B), concluding that the bait kit is suitable across the whole lineage. While not explicitly tested, our taxon sampling illustrates that the baits may have enough resolution power to resolve even lower generic levels, as with Amaranthus L. (Figure 3). However, denser sampling is needed to corroborate this claim, and further studies are currently underway in our lab.
Another important factor to consider when selecting markers for target enrichment is their length. Longer loci tend to contain more single‐nucleotide polymorphisms (SNPs) (e.g., Figure 1B), making them more informative for phylogenomic studies. However, it is crucial to recognize that in a species tree inference based on gene trees, each locus contributes one gene tree, regardless of its length. This means that while longer loci may capture more genetic variation, they require more sequencing resources but ultimately carry the same weight in downstream analyses as shorter loci. Nevertheless, shorter loci tend to have more gene tree errors (Zhang et al., 2017), affecting two‐step species tree inference methods like ASTRAL (Mirarab and Warnow, 2015). Furthermore, erroneous gene trees can affect orthology inference and downstream analyses (Morales‐Briones et al., 2022). For this reason, we targeted longer loci between 800 and 2500 bp (with an average length of 1295 bp) (Figure 1A) to provide sufficient SNP density without unnecessarily increasing sequencing costs or data complexity.
In the context of genome duplications, the risk of assembling chimeric genes increases, especially when selected markers include multi‐exon genes (Morales‐Briones et al., 2018). Such errors may obscure the evolutionary history of the locus in question, making orthology inference difficult. For instance, the universal Angiosperms353 bait set contains 2304 exons from 353 genes, with an average exon length of 165 bp (Johnson et al., 2019). While this bait set has proven effective in many applications, its use in polyploid species carries the risk of assembling genes containing exons from different gene copies. To address this issue, we prioritized the selection of longer exons (Figure 1A) to minimize the potential for compound loci and improve the reliability of orthology inference, as suggested in Morales‐Briones et al. (2022).
We found a relatively low percentage of on‐target reads (36.6%) compared with other custom baits: 42% in Zingiberaceae (Carlsen et al., 2018), 42.47% in Arachis L. (Peng et al., 2017), 48.6% in Euphorbia L. (Villaverde et al., 2018), 52% in Ochnaceae (Shah et al., 2021), and 73% in Bignoniaceae (Fonseca et al., 2023). However, these differences are strongly influenced by the choice of mapping tool. BLASTX, which we used, typically maps fewer reads than alignment‐based tools like BWA, because it performs protein‐level searches with stringent statistical filtering. In turn, BLASTX typically retrieves more genes, despite being computationally more demanding (https://github.com/mossmatters/HybPiper/wiki/Troubleshooting,-common-issues,-and-recommendations) (e.g., Maurin et al., 2021). This is highlighted in our data, where we show that, even if a low percentage of reads are mapped to the target (e.g., Deeringia amaranthoides A18: 9.4% on target), most of the loci are recovered (e.g., Deeringia amaranthoides A18: 975 genes at 75% length) (Appendix S5). Furthermore, we set strict requirements for locus assembly. Using a customized coverage cutoff for SPAdes (‐‐cov_cutoff 15 to 60) or the high percent identity threshold for retaining exon hits, we avoid assembling potential contaminations or chimeric exons consisting of contigs with very different coverage values.
Paralog handling
While many bait design studies aim to select single‐copy genes to minimize downstream conflict with paralogs, we selected genes from tree‐inferred orthologs. Single‐copy loci often represent highly conserved genes involved in essential biological processes such as photosynthesis or the cell cycle, which often limits their use when addressing other evolutionary questions (De Smet et al., 2013; Li et al., 2016). In contrast, using non‐housekeeping orthologous genes allows us to potentially broaden the scope of the bait kit by targeting more variable genes that can be used to infer the evolutionary histories of lower taxonomic units. By directly using the orthologs for bait design, we ensure that we capture the copies that reflect the evolutionary history of the species in question. Nevertheless, despite numerous bioinformatic tools (Emms and Kelly, 2019; Grau‐Bové and Sebé‐Pedrós, 2021; Persson and Sonnhammer, 2022), determining orthology remains challenging, as duplicated regions may undergo re‐diploidization over time, resulting in hidden paralogy (Xiang et al., 2017; Bomblies, 2020).
Paralogs have gained attention in phylogenomic studies due to their potential to provide valuable evolutionary insights (Gardner et al., 2021; Morales‐Briones et al., 2022; Walden et al., 2024). In our study, we retained all paralogous copies extracted with HybPiper2 and inferred orthology using an automated tree‐based approach (Yang and Smith, 2014) rather than discarding paralogs at the outset and thus losing potentially important information (Ufimov et al., 2021). This method was developed for genomic or transcriptomic datasets but has also proven effective for paralogous loci from target enrichment datasets (Morales‐Briones et al., 2022). Using the MO approach, we retained 972 out of 1000 loci, while 853 genes were paralog‐flagged after extraction with HybPiper2 and, therefore, theoretically had to be excluded (Appendix S5). Hence, we concur with Ufimov et al. (2022) and Morales‐Briones et al. (2022) that, in the case of WGD, paralogous flagged loci should not be removed from the analysis as this will result in substantial data loss. Furthermore, by discarding paralogous flagged loci at the outset, valuable information about biological processes such as hybridization is lost (Joyce et al., 2025). Instead, all copies should be retrieved and orthologous alignments should be generated (Yang and Smith, 2014).
A notable drawback of the MO approach is its potential to introduce bias in cases of allopolyploidy (Morales‐Briones et al., 2022). Yang et al. (2018) identified two allopolyploidy events (AMAR1 and AMAR2) within Amaranthaceae s.s.: one at the base of the Aervoids and Gomphrenoids, and another within the Aervoids, between the species Aerva javanica (Burm. f.) Juss. ex Schult. and Ouret lanata (L.) Kuntze. These two nodes exhibit a high percentage of missing data in PhyParts (Figure 3), likely reflecting these challenges associated with allopolyploidy. In such cases, the MO algorithm can create imbalances between subtrees. This occurs because the algorithm prunes the subtree with fewer taxa, which may disproportionately affect one subgenome due to unequal gene loss following allopolyploidization. Additionally, the design of baits on allopolyploid taxa can introduce bias by preferentially targeting one subgenome over the other (Morales‐Briones et al., 2022). To address the heterogeneity of gene trees due to gene duplication and loss, other tree‐based orthology inference methods, such as the RT algorithm that keeps all subtrees (Yang and Smith, 2014), DISCO (Willson et al., 2022), or ASTRAL‐Pro (Zhang and Mirarab, 2022b), can be used instead (Appendix S3).
Phylogenomics of Amaranthaceae s.s
The phylogenomic tree of Amaranthaceae s.s. was well‐supported and overall in line with existing phylogenies (Figure 3) (Müller and Borsch, 2005; Sánchez‐Del Pino et al., 2009, 2012; Hammer et al., 2015, 2019; Bena et al., 2017, 2024; Di Vincenzo et al., 2018, 2025; Waselkov et al., 2018; Huang et al., 2020; Limarino and Borsch, 2020; Morales‐Briones et al., 2021; Xu et al., 2024). We confirmed the topology among the five main clades—Gomphrenoids, Achyranthoids, Aervoids, Amaranthoids, and Celosioids. While many studies placed the genera Bosea and Charpentiera as a sister grade to the rest of Amaranthaceae s.s. (Kadereit et al., 2003; Müller and Borsch, 2005; Sage et al., 2007; Ogundipe and Chase, 2009; Bena et al., 2017; Di Vincenzo et al., 2018; Huang et al., 2020), we were able to resolve this relationship for the first time by using an NGS‐based approach and found that they are sister genera to the Amaranthoids and Celosioids. However, we discovered high levels of gene tree discordance between them (Figure 3), with the main alternative being that both lineages are sisters to each other (data not shown). In the past, the two genera Bosea and Charpentiera had been placed together with the genera Amaranthus and Chamissoa Kunth in the subtribe Amaranthineae of Amarantheae (Townsend, 1993), and proximity to Deeringia R. Br. had also been suggested (Kadereit et al., 2003). Although both claims are based on limited morphology only, they are substantiated by our results.
Within the Gomphrenoids, the genus Alternanthera was found to be monophyletic and sister to Tidestromia Standl., as found in previous studies (Sánchez‐Del Pino et al., 2012; Bena et al., 2017; Morales‐Briones et al., 2021). In contrast, the genera Gomphrena and Quaternella were retrieved as polyphyletic. While the genus Quaternella has never been included in a phylogenetic reconstruction, except for one species (Quaternella ephedroides Pedersen) (Morales‐Briones et al., 2021), Gomphrena has previously been shown to be polyphyletic (Sánchez‐Del Pino et al., 2009; Bena et al., 2017). The placement of Achyranthoids largely aligns with previous studies by Di Vincenzo et al. (2018, 2025) and Ogundipe and Chase (2009). Only the monotypic genus Arthraerua (Kuntze) Schinz, here recovered as a sister to Achyranthes bidentata Blume, Nototrichium humile Hillebr., Nelsia quadrangula, and Pandiaka involucrata (Moq.) B. D. Jacks. (Figure 3), was in a polytomy with Calicorema Hook. f. in the backbone of Achyranthoids in Di Vincenzo et al. (2018). Within the Aervoids, Ouret lanata (syn. Aerva lanata (L.) Juss. ex Schult.) and Aerva javanica were recovered as sister species, contrary to the findings of Hammer et al. (2019), who found Aerva javanica as sister to Paraerva T. Hammer and Ouret Adans. as sister to Ptilotus.
The topology of the Amaranthoids is consistent with previous studies, with the largest genus, Amaranthus, being monophyletic (Di Vincenzo et al., 2018; Waselkov et al., 2018; Xu et al., 2024). Only the genus Neocentema has never been phylogenetically analyzed. Morphologically, it was placed in the subtribe Achyranthinae by Schinz (1934) but was later revised by Suessenguth (1949), who placed the genus in the Amaranthinae due to the root tip similarities of the embryo and suggested a close relationship with Digera Forssk. Here, we provide the first molecular confirmation of this morphological classification. Within the Celosioids, a sister relationship of the genera Celosia L. and Hermbstaedtia Rchb. was found, contrary to the Sanger‐based topology in Di Vincenzo et al. (2018), where Celosia and Deeringia were retrieved as sister genera, albeit not well supported. Although Celosioids comprise important crop plants like Celosia argentea L. (Ayodele, 2023), they are still severely understudied phylogenetically, and further molecular analyses are needed to resolve their evolutionary history.
Conclusions
Our newly developed bait kit, Amaranthaceae1000, successfully targeted the intended loci. With a customized pipeline, we generated a well‐resolved phylogeny of Amaranthaceae s.s. While the phylogeny itself suffers from incomplete taxon sampling, it provides valuable insights, including the placement of previously recalcitrant groups, such as Bosea and Charpentiera, or previously non‐sampled genera, such as Neocentema. While further sampling is needed to address clade‐specific questions, we here illustrate that the Amaranthaceae1000 bait kit is a powerful and promising tool to address questions beyond pure phylogenomics, including spatio‐temporal reconstructions, population and comparative genomics, and trait evolution.
AUTHOR CONTRIBUTIONS
D.F.M.B., T.K., A.Z.C., and G.K. designed the research. T.K. conducted the sampling in M and MSB, while A.Z.C. carried out the sampling in PERTH and CANB. T.K. performed the laboratory work. D.F.M.B. and T.K. analyzed the data. G.K. provided the funding. All authors interpreted the results. T.K. led the writing under the supervision of D.F.M.B., A.Z.C., and G.K., and all authors contributed to and approved the final version of the manuscript.
Supporting information
Appendix S1. Loci extraction report from 29 Amaranthaceae s.s. transcriptomes using the CAPTUS pipeline. The completeness of recovered loci is color‐coded in a gradient from black (0%) to red (100%). (Top) Locus extraction using a target file containing only sequences of clade 1 (Gomphrenoids, Achyranthoids, and Aervoids) in blue. (Bottom) Locus extraction using a target file containing only sequences of clade 2 (Amaranthoids and Celosioids) in yellow.
Appendix S2. Fast‐Plast results showing the percentage of known angiosperm chloroplast genes recovered in 24 samples sequenced with the Amaranthaceae1000 baits.
Appendix S3. Astral‐Pro3 phylogenetic inference with all cleaned homologous trees from 57 transcriptomes and 24 samples sequenced with the Amaranthaceae1000 baits. The tree is rooted on members of the Caryophyllales, and support values on the branches correspond to local posterior probabilities (LPPs).
Appendix S4. Gene duplication mapping results. (Left) Histogram showing the percentage of gene duplications per branch. (Right) Phylogenetic inference from 24 species using ASTRAL IV with “monophyletic outgroup” (MO) ortholog trees, rooted on members of the Caryophyllales. Branch values indicate the proportion of duplicated genes with above 6 in bold. The stars mark the known WGD from Yang et al. (2018).
Appendix S5. Summary statistics of sequencing success obtained from HybPiper2. Metrics include species name, clade, number of reads, mapped reads, percentage mapped to targets, number of mapped genes, and additional gene recovery statistics (e.g., genes with contigs, sequences, or specific coverage thresholds). The table also reports paralog warnings (by length and depth), genes without or with stitched contigs, skipped contigs, chimera warnings, and the total bases recovered.
ACKNOWLEDGMENTS
The authors thank the Elfriede und Franz Jakob Foundation for their financial support. We also thank the curators and staff of the herbaria M, MSB, CANB, and PERTH for allowing us to study their specimens and to perform limited destructive sampling. We would also like to express our gratitude to Alina Höwener (Ludwig‐Maximilians‐Universität München) for her support and advice in the laboratory process. Open Access funding enabled and organized by Projekt DEAL.
Appendix 1. Voucher information for the newly sequenced Amaranthaceae s.s. species.
| Species | Collector and collection no. | Collection datea | Collection locality | Herbarium barcode | DNA ID | Laboratory accession no. | SRA accession no. |
|---|---|---|---|---|---|---|---|
| Aerva javanica (Burm. f.) Juss. ex Schult. | D. Podlech 36191 | 02.10.1981 | Yemen | MSB‐122362 | 4753 | A9 | SAMN48338600 |
| Alternanthera bettzickiana (Regel) G. Nicholson | G.T. Prance et al. 15574 | 23.10.1971 | Brazil | M‐0342909 | 4765 | A16 | SAMN48338601 |
| Amaranthus albus L. | Richard R. Halse 8987 | 11.09.2013 | USA | MJG015317 | 4729 | A32 | SAMN48338602 |
| Amaranthus macrocarpus Benth. | SJ 9245 | 01.18.2005 | Australia | — | 840 | A49 | SAMN48338603 |
| Arthraerua leubnitziae (Kuntze) Schinz | S.W. Breckle 10308 | 14.12.1986 | Namibia | M‐0331726 | 4719 | A25 | SAMN48338604 |
| Bosea yervamora L. | W. Hdhg. DT 2294 | 08.03.1995 | Spain | M‐0331728 | 4730 | A35 | SAMN48338605 |
| Celosia elegantissima Hauman | Reekmanns 10648 | 12.06.1981 | Burundi | MSB‐122608 | 4731 | A36 | SAMN48338606 |
| Charpentiera elliptica (Hillebr.) A. Heller | Paul C. Hutchison & John Obata 2824 | 15.01.1967 | Hawaii, USA | M‐0331740 | 4757 | A12 | SAMN48338607 |
| Cyphocarpa angustifolia Lopr. | K. Balkwill et al. 5344 | 13.01.1990 | South Africa | M‐0331747 | 4720 | A26 | SAMN48338608 |
| Deeringia amaranthoides (Lam.) Merr. | H. Y. Liang 63925 | 10.1933 | China | M‐0331749 | 4767 | A18 | SAMN48338609 |
| Digera muricata (L.) Mart. | Thulin 11452 | 2006 | Oman | — | 3731 | A50 | SAMN48338610 |
| Gomphrena arida J. Palmer | A. Fraser 348 | 13.03.2001 | Australia | CANB 636242.1 | 4734 | A37 | SAMN48338611 |
| Hebanthe erianthos (Poir.) Pedersen | Strang 1057 | 23.07.1967 | Brazil | M‐0335298 | 4768 | A19 | SAMN48338612 |
| Hermbstaedtia glauca (J. C. Wendl.) Rchb. ex Steud. | H. Müller 772 | 01.08.1977 | Namibia | M‐0331756 | 4763 | A15 | SAMN48338613 |
| Neocentema robecchii (Lopr.) Schinz | Kazmi et al. 58 | 21.12.1977 | Somalia | M‐0331769 | 4745 | A8 | SAMN48338614 |
| Nototrichium humile Hillebr. | Otto Degener 20640 | 07.05.1950 | Hawaii, USA | M‐0335291 | 4771 | A20 | SAMN48338615 |
| Ouret lanata (L.) Kuntze | I. Hageman & R. Claßen 3224 | 01.10.1986 | Kenya | MJG012475 | 4725 | A30 | SAMN48338616 |
| Pandiaka involucrata (Moq.) B. D. Jacks. | R. Bartha 1/12 | 1967 | Nigeria | M‐0335293 | 4726 | A31 | SAMN48338617 |
| Pfaffia glabrata Mart. | A.P. Duarte 8305 | 19.07.1964 | Brazil | M‐0335294 | 4772 | A33 | SAMN48338618 |
| Pleuropterantha revoilii Franch. | J.J.F.E. De Wilde 5953 | 30.06.1969 | Ethiopia | M‐0335302 | 4774 | A21 | SAMN48338619 |
| Ptilotus mollis Benl | J. Hruban & D. Coultas SBJ‐01 | 10.02.2022 | Australia | PERTH 09468528 | 4750 | A4 | SAMN48338620 |
| Ptilotus sericostachyus (Nees) F. Muell. | PAGL 1/34 | 16.10.2005 | Australia | PERTH 07508271 | 4751 | A43 | SAMN48338621 |
| Quaternella glabratoides (Suess.) Pedersen | Pereira 7917 | 1963 | Brazil | M‐0335295 | 4752 | A44 | SAMN48338622 |
| Sericorema remotiflora (Hook.) Lopr. | Venter S. 12878 | 30.03.1988 | South Africa | M‐0335307 | 4739 | A5 | SAMN48338623 |
Note: SRA = Sequence Read Archive.
Collection date is formatted as day.month.year.
Kiedaisch, T. , Kadereit G., Žerdoner Čalasan A., and Morales‐Briones D. F.. 2025. Advancing phylogenomics in Amaranthaceae sensu stricto: Development and application of a new nuclear target enrichment bait set. Applications in Plant Sciences 13(5): e70019. 10.1002/aps3.70019
DATA AVAILABILITY STATEMENT
All generated data are available from the National Center for Biotechnology Information (NCBI; Bioproject PRJNA1209683). The workflow for the bait design is available on Github (https://github.com/tinakiedaisch/bait_design_from_orthologs), and the bait design files are available on Dryad (https://doi.org/10.5061/dryad.k3j9kd5m6; Kiedaisch et al., 2025).
REFERENCES
- Acha, S. , and Majure L. C.. 2022. A new approach using targeted sequence capture for phylogenomic studies across Cactaceae. Genes 13: 350. 10.3390/genes13020350 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aderibigbe, O. R. , Ezekiel O. O., Owolade S. O., Korese J. K., Sturm B., and Hensel O.. 2022. Exploring the potentials of underutilized grain amaranth (Amaranthus spp.) along the value chain for food and nutrition security: A review. Critical Reviews in Food Science and Nutrition 62: 656–669. 10.1080/10408398.2020.1825323 [DOI] [PubMed] [Google Scholar]
- Andermann, T. , Torres Jiménez M. F., Matos‐Maraví P., Batista R., Blanco‐Pastor J. L., Gustafsson A. L. S., Kistler L., et al. 2020. A guide to carrying out a phylogenomic target sequence capture project. Frontiers in Genetics 10: 1407. 10.3389/fgene.2019.01407 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Andrews, S. 2010. FastQC: A quality control tool for high throughput sequence data. Website: https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ [accessed 1 July 2025].
- Antonelli, A. , Clarkson J. J., Kainulainen K., Maurin O., Brewer G. E., Davis A. P., Epitawalage N., et al. 2021. Settling a family feud: A high‐level phylogenomic framework for the Gentianales based on 353 nuclear genes and partial plastomes. American Journal of Botany 108: 1143–1165. 10.1002/ajb2.1697 [DOI] [PubMed] [Google Scholar]
- Ayodele, O. P. 2023. Growth, yield and nutritional quality of Lagos spinach (Celosia argentea L.) as influenced by the density of goat weed (Ageratum conyzoides L.). Journal of Plant Protection Research 61: 20–27. 10.24425/jppr.2021.136265 [DOI] [Google Scholar]
- Batley, J. , and Edwards D.. 2009. Genome sequence data: Management, storage, and visualization. BioTechniques 46: 333–336. 10.2144/000113134 [DOI] [PubMed] [Google Scholar]
- Bayón, N. D. 2022. Identifying the weedy amaranths (Amaranthus, Amaranthaceae) of South America. Advances in Weed Science 40: e0202200013. 10.51694/AdvWeedSci/2022;40:Amaranthus007 [DOI] [Google Scholar]
- Bena, M. J. , Acosta J. M., and Aagesen L.. 2017. Macroclimatic niche limits and the evolution of C4 photosynthesis in Gomphrenoideae (Amaranthaceae). Botanical Journal of the Linnean Society 184: 283–297. 10.1093/botlinnean/box031 [DOI] [Google Scholar]
- Bena, M. J. , Baranzelli M. C., Costas S. M., Cosacov A., Acosta M. C., Moreira‐Muñoz A., and Sérsic A. N.. 2024. Linking South American dry regions by the Gran Chaco: Insights from the evolutionary history and ecological diversification of Gomphrena s.str. (Gomphrenoideae, Amaranthaceae). Journal of Systematics and Evolution 62: 758–774. 10.1111/jse.13023 [DOI] [Google Scholar]
- Bennett, M. D. , and Smith J. B.. 1976. Nuclear DNA amounts in angiosperms. Philosophical Transactions of the Royal Society of London, B: Biological Sciences 274: 227–274. 10.1098/rstb.1976.0044 [DOI] [PubMed] [Google Scholar]
- Bogarín, D. , Pérez‐Escobar O. A., Groenenberg D., Holland S. D., Karremans A. P., Lemmon E. M., Lemmon A. R., et al. 2018. Anchored hybrid enrichment generated nuclear, plastid and mitochondrial markers resolve the Lepanthes horrida (Orchidaceae: Pleurothallidinae) species complex. Molecular Phylogenetics and Evolution 129: 27–47. 10.1016/j.ympev.2018.07.014 [DOI] [PubMed] [Google Scholar]
- Bolger, A. M. , Lohse M., and Usadel B.. 2014. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 30: 2114–2120. 10.1093/bioinformatics/btu170 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bomblies, K. 2020. When everything changes at once: Finding a new normal after genome duplication. Proceedings of the Royal Society B, Biological Sciences 287: 20202154. 10.1098/rspb.2020.2154 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Borowiec, M. L. 2017. AMAS: A fast tool for large alignment manipulation and computing of summary statistics. PeerJ 4: e1660. 10.7717/peerj.1660 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brown, J. W. , Walker J. F., and Smith S. A.. 2017. Phyx: Phylogenetic tools for unix. Bioinformatics 33: 1886–1888. 10.1093/bioinformatics/btx063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Capella‐Gutiérrez, S. , Silla‐Martínez J. M., and Gabaldón T.. 2009. trimAl: A tool for automated alignment trimming in large‐scale phylogenetic analyses. Bioinformatics 25: 1972–1973. 10.1093/bioinformatics/btp348 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Carlsen, M. M. , Fér T., Schmickl R., Leong‐Škorničková J., Newman M., and Kress W. J.. 2018. Resolving the rapid plant radiation of early diverging lineages in the tropical Zingiberales: Pushing the limits of genomic data. Molecular Phylogenetics and Evolution 128: 55–68. 10.1016/j.ympev.2018.07.020 [DOI] [PubMed] [Google Scholar]
- Chamala, S. , García N., Godden G. T., Krishnakumar V., Jordon‐Thaden I. E., De Smet R., Barbazuk W. B., et al. 2015. MarkerMiner 1.0: A new application for phylogenetic marker development using angiosperm transcriptomes. Applications in Plant Sciences 3: e1400115. 10.3732/apps.1400115 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chiu, J. C. , Lee E. K., Egan M. G., Sarkar I. N., Coruzzi G. M., and DeSalle R.. 2006. OrthologID: Automation of genome‐scale ortholog identification within a parsimony framework. Bioinformatics 22: 699–707. 10.1093/bioinformatics/btk040 [DOI] [PubMed] [Google Scholar]
- Christe, C. , Boluda C. G., Koubínová D., Gautier L., and Naciri Y.. 2021. New genetic markers for Sapotaceae phylogenomics: More than 600 nuclear genes applicable from family to population levels. Molecular Phylogenetics and Evolution 160: 107123. 10.1016/j.ympev.2021.107123 [DOI] [PubMed] [Google Scholar]
- De Smet, R. , Adams K. L., Vandepoele K., Van Montagu M. C. E., Maere S., and Van De Peer Y.. 2013. Convergent gene loss following gene and genome duplications creates single‐copy families in flowering plants. Proceedings of the National Academy of Sciences, USA 110: 2898–2903. 10.1073/pnas.1300127110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Di Vincenzo, V. , Gruenstaeudl M., Nauheimer L., Wondafrash M., Kamau P., Demissew S., and Borsch T.. 2018. Evolutionary diversification of the African achyranthoid clade (Amaranthaceae) in the context of sterile flower evolution and epizoochory. Annals of Botany 122: 69–85. 10.1093/aob/mcy055 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Di Vincenzo, V. , Berendsohn W., Wondafrash M., and Borsch T.. 2025. Phylogenetics and morphological character evolution in the achyranthoid clade (Amaranthaceae): Evidence to re‐circumscribe the genera Achyranthes and Cyathula and to resurrect a third species of the former genus Sericocomopsis in East Africa. Taxon 74: 66–100. 10.1002/tax.13284 [DOI] [Google Scholar]
- Dunn, C. W. , Howison M., and Zapata F.. 2013. Agalma: An automated phylogenomics workflow. BMC Bioinformatics 14: 330. 10.1186/1471-2105-14-330 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Emms, D. M. , and Kelly S.. 2019. OrthoFinder: Phylogenetic orthology inference for comparative genomics. Genome Biology 20: 238. 10.1186/s13059-019-1832-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eserman, L. A. , Thomas S. K., Coffey E. E. D., and Leebens‐Mack J. H.. 2021. Target sequence capture in orchids: Developing a kit to sequence hundreds of single‐copy loci. Applications in Plant Sciences 9: e11416. 10.1002/aps3.11416 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ewels, P. , Magnusson M., Lundin S., and Käller M.. 2016. MultiQC: Summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32: 3047–3048. 10.1093/bioinformatics/btw354 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fitch, W. M. 1970. Distinguishing homologous from analogous proteins. Systematic Biology 19: 99–113. 10.2307/2412448 [DOI] [PubMed] [Google Scholar]
- Folk, R. A. , Mandel J. R., and Freudenstein J. V.. 2015. A protocol for targeted enrichment of intron‐containing sequence markers for recent radiations: A phylogenomic example from Heuchera (Saxifragaceae). Applications in Plant Sciences 3: e1500039. 10.3732/apps.1500039 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fonseca, L. H. M. , Carlsen M. M., Fine P. V. A., and L. G. Lohmann. 2023. A nuclear target sequence capture probe set for phylogeny reconstruction of the charismatic plant family Bignoniaceae. Frontiers in Genetics 13: 1085692. 10.3389/fgene.2022.1085692 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gabaldón, T. 2008. Large‐scale assignment of orthology: Back to phylogenetics? Genome Biology 9: 235. 10.1186/gb-2008-9-10-235 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gabaldón, T. , and Koonin E. V.. 2013. Functional and evolutionary implications of gene orthology. Nature Reviews Genetics 14: 360–366. 10.1038/nrg3456 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gardner, E. M. , Johnson M. G., Pereira J. T., Puad A. S. A., Arifiani D., Sahromi, Wickett N. J., and Zerega N. J. C.. 2021. Paralogs and off‐target sequences improve phylogenetic resolution in a densely sampled study of the breadfruit genus (Artocarpus, Moraceae). Systematic Biology 70: 558–575. 10.1093/sysbio/syaa073 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giaretta, A. , Murphy B., Maurin O., Mazine F. F., Sano P., and Lucas E.. 2022. Phylogenetic relationships within the hyper‐diverse genus Eugenia (Myrtaceae: Myrteae) based on target enrichment sequencing. Frontiers in Plant Science 12: 759460. 10.3389/fpls.2021.759460 [DOI] [PMC free article] [PubMed] [Google Scholar]
- González‐Domínguez, J. , and Schmidt B.. 2016. ParDRe: Faster parallel duplicated reads removal tool for sequencing studies. Bioinformatics 32: 1562–1564. 10.1093/bioinformatics/btw038 [DOI] [PubMed] [Google Scholar]
- Goodstein, D. M. , Shu S., Howson R., Neupane R., Hayes R. D., Fazo J., Mitros T., et al. 2012. Phytozome: A comparative platform for green plant genomics. Nucleic Acids Research 40: D1178–D1186. 10.1093/nar/gkr944 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Grau‐Bové, X. , and Sebé‐Pedrós A.. 2021. Orthology clusters from gene trees with Possvm . Molecular Biology and Evolution 38: 5204–5208. 10.1093/molbev/msab234 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haigh, A. L. , Gibernau M., Maurin O., Bailey P., Carlsen M. M., Hay A., Leempoel K., et al. 2023. Target sequence data shed new light on the infrafamilial classification of Araceae. American Journal of Botany 110: e16117. 10.1002/ajb2.16117 [DOI] [PubMed] [Google Scholar]
- Hale, H. , Gardner E. M., Viruel J., Pokorny L., and M. G. Johnson. 2020. Strategies for reducing per‐sample costs in target capture sequencing for phylogenomics and population genomics in plants. Applications in Plant Sciences 8: e11337. 10.1002/aps3.11337 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hammer, T. , Davis R., and Thiele K.. 2015. A molecular framework phylogeny for Ptilotus (Amaranthaceae): Evidence for the rapid diversification of an arid Australian genus. Taxon 64: 272–285. 10.12705/642.6 [DOI] [Google Scholar]
- Hammer, T. A. , Zhong X., Colas Des Francs‐Small C., Nevill P. G., Small I. D., and Thiele K. R.. 2019. Resolving intergeneric relationships in the aervoid clade and the backbone of Ptilotus (Amaranthaceae): Evidence from whole plastid genomes and morphology. Taxon 68: 297–314. 10.1002/tax.12054 [DOI] [Google Scholar]
- Helmstetter, A. J. , Ezedin Z., De Lírio E. J., De Oliveira S. M., Chatrou L. W., Erkens R. H. J., Larridon I., et al. 2025. Toward a phylogenomic classification of magnoliids. American Journal of Botany 112: e16451. 10.1002/ajb2.16451 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hernández‐Ledesma, P. , Berendsohn W. G., Borsch T., Mering S. V., Akhani H., Arias S., Castañeda‐Noa I., et al. 2015. A taxonomic backbone for the global synthesis of species diversity in the angiosperm order Caryophyllales . Willdenowia 45: 281. 10.3372/wi.45.45301 [DOI] [Google Scholar]
- Hoang, D. T. , Chernomor O., Von Haeseler A., Minh B. Q., and Vinh L. S.. 2018. UFBoot2: Improving the ultrafast bootstrap approximation. Molecular Biology and Evolution 35: 518–522. 10.1093/molbev/msx281 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huang, J. , Chen W., Li Y., and Yao G.. 2020. Phylogenetic study of Amaranthaceae sensu lato based on multiple plastid DNA fragments. Chinese Bulletin of Botany 55: 457. 10.11983/CBB19228 [DOI] [Google Scholar]
- Johnson, M. G. , Gardner E. M., Liu Y., Medina R., Goffinet B., Shaw A. J., Zerega N. J. C., and Wickett N. J.. 2016. HybPiper: Extracting coding sequence and introns for phylogenetics from high‐throughput sequencing reads using target enrichment. Applications in Plant Sciences 4: e1600016. 10.3732/apps.1600016 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Johnson, M. G. , Pokorny L., Dodsworth S., Botigué L. R., Cowan R. S., Devault A., Eiserhardt W. L., et al. 2019. A universal probe set for targeted sequencing of 353 nuclear genes from any flowering plant designed using k‐medoids clustering. Systematic Biology 68: 594–606. 10.1093/sysbio/syy086 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Joshi, N. , and Verma K. C.. 2020. A review on nutrition value of Amaranth (Amaranthus caudatus L.): The crop of future. Journal of Pharmacognosy and Phytochemistry 9: 1111–1113. [Google Scholar]
- Joyce, E. M. , Schmidt‐Lebuhn A. N., Orel H. K., Nge F. J., Anderson B. M., Hammer T. A., and McLay T. G. B.. 2025. Navigating phylogenetic conflict and evolutionary inference in plants with target‐capture data. Australian Systematic Botany 38: SB24011. 10.1071/SB24011 [DOI] [Google Scholar]
- Kadereit, G. , Borsch T., Weising K., and Freitag H.. 2003. Phylogeny of Amaranthaceae and Chenopodiaceae and the evolution of C4 photosynthesis. International Journal of Plant Sciences 164: 959–986. 10.1086/378649 [DOI] [Google Scholar]
- Katoh, K. , and Frith M. C.. 2012. Adding unaligned sequences into an existing alignment using MAFFT and LAST. Bioinformatics 28: 3144–3146. 10.1093/bioinformatics/bts578 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Katoh, K. , and Standley D. M.. 2013. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. Molecular Biology and Evolution 30: 772–780. 10.1093/molbev/mst010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kiedaisch, T. , Kadereit G., Žerdoner Čalasan A., and Morales‐Briones D. F.. 2025. Data from: Advancing phylogenomics in Amaranthaceae sensu stricto: Development and application of a new nuclear target enrichment bait set. Dryad Dataset. 10.5061/dryad.k3j9kd5m6 [accessed 21 July 2025]. [DOI]
- Koonin, E. V. 2005. Orthologs, paralogs, and evolutionary genomics. Annual Review of Genetics 39: 309–338. 10.1146/annurev.genet.39.073003.114725 [DOI] [PubMed] [Google Scholar]
- Lee, A. K. , Gilman I. S., Srivastav M., Lerner A. D., Donoghue M. J., and Clement W. L.. 2021. Reconstructing Dipsacales phylogeny using Angiosperms353: Issues and insights. American Journal of Botany 108: 1122–1142. 10.1002/ajb2.1695 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, Z. , Defoort J., Tasdighian S., Maere S., Van De Peer Y., and De Smet R.. 2016. Gene duplicability of core genes is highly consistent across all angiosperms. Plant Cell 28: 326–344. 10.1105/tpc.15.00877 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Limarino, T. O. , and Borsch T.. 2020. Gomphrena (Amaranthaceae, Gomphrenoideae) diversified as a C4 lineage in the New World tropics with specializations in floral and inflorescence morphology, and an escape to Australia. Willdenowia 50: 345–381. 10.3372/wi.50.50301 [DOI] [Google Scholar]
- Mamanova, L. , Coffey A. J., Scott C. E., Kozarewa I., Turner E. H., Kumar A., Howard E., et al. 2010. Target‐enrichment strategies for next‐generation sequencing. Nature Methods 7: 111–118. 10.1038/nmeth.1419 [DOI] [PubMed] [Google Scholar]
- Maurin, O. , Anest A., Bellot S., Biffin E., Brewer G., Charles‐Dominique T., Cowan R. S., et al. 2021. A nuclear phylogenomic study of the angiosperm order Myrtales, exploring the potential and limitations of the universal Angiosperms353 probe set. American Journal of Botany 108: 1087–1111. 10.1002/ajb2.1699 [DOI] [PubMed] [Google Scholar]
- McKain, M. A. 2017. Mrmckain/Fast‐Plast: Fast‐Plast V.1.2.6. Available at Zenodo repository: 10.5281/ZENODO.973887 [posted 23 September 2017; accessed 1 July 2025]. [DOI]
- Mendez‐Reneau, J. , Burleigh J. G., and Sigel E. M.. 2023. Target capture methods offer insight into the evolution of rapidly diverged taxa and resolve allopolyploid homeologs in the fern genus Polypodium s.s. Systematic Botany 48: 96–109. 10.1600/036364423X16758873924135 [DOI] [Google Scholar]
- Minh, B. Q. , Schmidt H. A., Chernomor O., Schrempf D., Woodhams M. D., Von Haeseler A., and Lanfear R.. 2020. IQ‐TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution 37: 2461. 10.1093/molbev/msaa131 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mirarab, S. , and Warnow T.. 2015. ASTRAL‐II: Coalescent‐based species tree estimation with many hundreds of taxa and thousands of genes. Bioinformatics 31: i44–i52. 10.1093/bioinformatics/btv234 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morales‐Briones, D. F. , Liston A., and Tank D. C.. 2018. Phylogenomic analyses reveal a deep history of hybridization and polyploidy in the Neotropical genus Lachemilla (Rosaceae). New Phytologist 218: 1668–1684. 10.1111/nph.15099 [DOI] [PubMed] [Google Scholar]
- Morales‐Briones, D. F. , Kadereit G., Tefarikis D. T., Moore M. J., Smith S. A., Brockington S. F., Timoneda A., et al. 2021. Disentangling sources of gene tree discordance in phylogenomic data sets: Testing ancient hybridizations in Amaranthaceae s.l. Systematic Biology 70: 219–235. 10.1093/sysbio/syaa066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morales‐Briones, D. F. , Gehrke B., Huang C.‐H., Liston A., Ma H., Marx H. E., Tank D. C., and Yang Y.. 2022. Analysis of paralogs in target enrichment data pinpoints multiple ancient polyploidy events in Alchemilla s.l. (Rosaceae). Systematic Biology 71: 190–207. 10.1093/sysbio/syab032 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Müller, K. , and Borsch T.. 2005. Phylogenetics of Amaranthaceae based on matK/trnK sequence data: Evidence from parsimony, likelihood, and Bayesian analyses. Annals of the Missouri Botanical Garden 92: 66–102. [Google Scholar]
- Nath, P. , Ohri D., and Pal M.. 1992. Nuclear DNA content in Celosia (Amaranthaceae). Plant Systematics and Evolution 182: 253–257. 10.1007/BF00939191 [DOI] [Google Scholar]
- Nikolov, L. A. , Shushkov P., Nevado B., Gan X., Al‐Shehbaz I. A., Filatov D., Bailey C. D., and Tsiantis M.. 2019. Resolving the backbone of the Brassicaceae phylogeny for investigating trait diversity. New Phytologist 222: 1638–1651. 10.1111/nph.15732 [DOI] [PubMed] [Google Scholar]
- Ogundipe, O. T. , and Chase M.. 2009. Phylogenetic analyses of Amaranthaceae based on matK DNA sequence data with emphasis on west African species. Turkish Journal of Botany 33: 153–161. 10.3906/bot-0707-15 [DOI] [Google Scholar]
- Ortiz, E. M. , Höwener A., Shigita G., Raza M., Maurin O., Zuntini A., Forest F., et al. 2024. A novel phylogenomics pipeline reveals complex pattern of reticulate evolution in Cucurbitales. bioRxiv 564367 [preprint]. Available at: 10.1101/2023.10.27.564367 [posted 26 September 2024; accessed 1 July 2025]. [DOI]
- Pease, J. B. , Brown J. W., Walker J. F., Hinchliff C. E., and Smith S. A.. 2018. Quartet Sampling distinguishes lack of support from conflicting support in the green plant tree of life. American Journal of Botany 105: 385–403. 10.1002/ajb2.1016 [DOI] [PubMed] [Google Scholar]
- Peng, Z. , Fan W., Wang L., Paudel D., Leventini D., Tillman B. L., and Wang J.. 2017. Target enrichment sequencing in cultivated peanut (Arachis hypogaea L.) using probes designed from transcript sequences. Molecular Genetics and Genomics 292: 955–965. 10.1007/s00438-017-1327-z [DOI] [PubMed] [Google Scholar]
- Persson, E. , and Sonnhammer E. L. L.. 2022. InParanoid‐DIAMOND: Faster orthology analysis with the InParanoid algorithm. Bioinformatics 38: 2918–2919. 10.1093/bioinformatics/btac194 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prjibelski, A. , Antipov D., Meleshko D., Lapidus A., and Korobeynikov A.. 2020. Using SPAdes de novo assembler. Current Protocols in Bioinformatics 70: e102. 10.1002/cpbi.102 [DOI] [PubMed] [Google Scholar]
- Quinlan, A. R. , and Hall I. M.. 2010. BEDTools: A flexible suite of utilities for comparing genomic features. Bioinformatics 26: 841–842. 10.1093/bioinformatics/btq033 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ranwez, V. , Douzery E. J. P., Cambon C., Chantret N., and Delsuc F.. 2018. MACSE v2: Toolkit for the alignment of coding sequences accounting for frameshifts and stop codons. Molecular Biology and Evolution 35: 2582–2584. 10.1093/molbev/msy159 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ranwez, V. , Chantret N., and Delsuc F.. 2021. Aligning protein‐coding nucleotide sequences with MACSE. In Katoh K. [ed.], Multiple Sequence Alignment: Methods and protocols, 51–70. Humana Press, New York, New York, USA. 10.1007/978-1-0716-1036-74 [DOI] [PubMed] [Google Scholar]
- Roberts, J. , and Florentine S.. 2022. A review of the biology, distribution patterns and management of the invasive species Amaranthus palmeri S. Watson (Palmer amaranth): Current and future management challenges. Weed Research 62: 113–122. 10.1111/wre.12520 [DOI] [Google Scholar]
- Sage, R. F. , Sage T. L., Pearcy R. W., and Borsch T.. 2007. The taxonomic distribution of C4 photosynthesis in Amaranthaceae sensu stricto. American Journal of Botany 94: 1992–2003. 10.3732/ajb.94.12.1992 [DOI] [PubMed] [Google Scholar]
- Sánchez‐Del Pino, I. S. , Borsch T., and Motley T. J.. 2009. trnL‐F and rpl16 sequence data and dense taxon sampling reveal monophyly of unilocular anthered Gomphrenoideae (Amaranthaceae) and an improved picture of their internal relationships. Systematic Botany 34: 57–67. 10.1600/036364409787602401 [DOI] [Google Scholar]
- Sánchez‐Del Pino, I. , Motley T. J., and Borsch T.. 2012. Molecular phylogenetics of Alternanthera (Gomphrenoideae, Amaranthaceae): Resolving a complex taxonomic history caused by different interpretations of morphological characters in a lineage with C4 and C3–C4 intermediate species. Botanical Journal of the Linnean Society 169: 493–517. 10.1111/j.1095-8339.2012.01248.x [DOI] [Google Scholar]
- Schinz, H. 1934. Amaranthaceae. In Engler A. and Prantl K. [eds.], Die natürlichen Pflanzenfamilien, ed. 2, 16c, pp. 7–85. W. Engelmann, Leipzig, Germany. [Google Scholar]
- Shah, T. , Schneider J. V., Zizka G., Maurin O., Baker W., Forest F., Brewer G. E., et al. 2021. Joining forces in Ochnaceae phylogenomics: A tale of two targeted sequencing probe kits. American Journal of Botany 108: 1201–1216. 10.1002/ajb2.1682 [DOI] [PubMed] [Google Scholar]
- Shendure, J. , and Ji H.. 2008. Next‐generation DNA sequencing. Nature Biotechnology 26: 1135–1145. 10.1038/nbt1486 [DOI] [PubMed] [Google Scholar]
- Smith, S. A. , Moore M. J., Brown J. W., and Yang Y.. 2015. Analysis of phylogenomic datasets reveals conflict, concordance, and gene duplications with examples from animals and plants. BMC Evolutionary Biology 15: 150. 10.1186/s12862-015-0423-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Soltis, D. E. , Gitzendanner M. A., Stull G., Chester M., Chanderbali A., Chamala S., Jordon‐Thaden I., et al. 2013. The potential of genomics in plant systematics. Taxon 62: 886–898. 10.12705/625.13 [DOI] [Google Scholar]
- Steenwyk, J. L. , Buida T. J., Labella A. L., Li Y., Shen X.‐X., and Rokas A.. 2021. PhyKIT: A broadly applicable UNIX shell toolkit for processing and analyzing phylogenomic data. Bioinformatics 37: 2325–2331. 10.1093/bioinformatics/btab096 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Steinegger, M. , and Söding J.. 2017. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology 35: 1026–1028. 10.1038/nbt.3988 [DOI] [PubMed] [Google Scholar]
- Suessenguth, K. 1949. Some new or noteworthy Amaranthaceae from East Africa. Kew Bulletin 1949: 475–480. [Google Scholar]
- Thiers, B. 2025. (continuously updated). Index Herbariorum. Website http://sweetgum.nybg.org/science/ih/ [accessed 3 July 2025].
- Thomas, S. K. , Liu X., Du Z., Dong Y., Cummings A., Pokorny L., Xiang Q.‐Y., and Leebens‐Mack J. H.. 2021. Comprehending Cornales: Phylogenetic reconstruction of the order using the Angiosperms353 probe set. American Journal of Botany 108: 1112–1121. 10.1002/ajb2.1696 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Timilsena, P. R. , Wafula E. K., Barrett C. F., Ayyampalayam S., McNeal J. R., Rentsch J. D., McKain M. R., et al. 2022. Phylogenomic resolution of order‐ and family‐level monocot relationships using 602 single‐copy nuclear genes and 1375 BUSCO genes. Frontiers in Plant Science 13: 876779. 10.3389/fpls.2022.876779 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Townsend, C. C. 1993. Amaranthaceae. In Kubitzki K., Rohwer J. G., and Bittrich V. [eds.], The families and genera of vascular plants, Vol. 2: Flowering plants: Dicotyledons, 70–91. Springer, Berlin, Germany. 10.1007/978-3-662-02899-5_7 [DOI] [Google Scholar]
- Ufimov, R. , Zeisek V., Píšová S., Baker W. J., Fér T., Van Loo M., Dobeš C., and Schmickl R.. 2021. Relative performance of customized and universal probe sets in target enrichment: A case study in subtribe Malinae. Applications in Plant Sciences 9: e11442. 10.1002/aps3.11442 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ufimov, R. , Gorospe J. M., Fér T., Kandziora M., Salomon L., Van Loo M., and Schmickl R.. 2022. Utilizing paralogues for phylogenetic reconstruction has the potential to increase species tree support and reduce gene tree discordance in target enrichment data. Molecular Ecology Resources 22: 3018–3034. 10.1111/1755-0998.13684 [DOI] [PubMed] [Google Scholar]
- Veltman, M. A. , Anthoons B., Schrøder‐Nielsen A., Gravendeel B., and De Boer H. J.. 2024. Orchidinae‐205: A new genome‐wide custom bait set for studying the evolution, systematics, and trade of terrestrial orchids. Molecular Ecology Resources 24: e13986. 10.1111/1755-0998.13986 [DOI] [PubMed] [Google Scholar]
- Villaverde, T. , Pokorny L., Olsson S., Rincón‐Barrado M., Johnson M. G., Gardner E. M., Wickett N. J., et al. 2018. Bridging the micro‐ and macroevolutionary levels in phylogenomics: Hyb‐Seq solves relationships from populations to species and above. New Phytologist 220: 636–650. 10.1111/nph.15312 [DOI] [PubMed] [Google Scholar]
- Walden, N. , Kiefer C., and Koch M. A.. 2024. Unravelling complex hybrid and polyploid evolutionary relationships using phylogenetic placement of paralogs from target enrichment data. bioRxiv 601132 [preprint]. Available at: 10.1101/2024.06.28.601132 [posted 2 July 2024; accessed 1 July 2025]. [DOI]
- Waselkov, K. E. , Boleda A. S., and Olsen K. M.. 2018. A phylogeny of the genus Amaranthus (Amaranthaceae) based on several low‐copy nuclear loci and chloroplast regions. Systematic Botany 43: 439–458. 10.1600/036364418X697193 [DOI] [Google Scholar]
- Weitemier, K. , Straub S. C. K., Cronn R. C., Fishbein M., Schmickl R., McDonnell A., and Liston A.. 2014. Hyb‐Seq: Combining target enrichment and genome skimming for plant phylogenomics. Applications in Plant Sciences 2: e1400042. 10.3732/apps.1400042 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Willson, J. , Roddur M. S., Liu B., Zaharias P., and Warnow T.. 2022. DISCO: Species tree inference using multicopy gene family tree decomposition. Systematic Biology 71: 610–629. 10.1093/sysbio/syab070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xiang, Y. , Huang C.‐H., Hu Y., Wen J., Li S., Yi T., Chen H., et al. 2017. Evolution of Rosaceae fruit types based on nuclear phylogeny in the context of geological times and genome duplication. Molecular Biology and Evolution 34: 263–281. 10.1093/molbev/msw242 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu, H. , Guo Y., Xia M., Yu J., Chi X., Han Y., Li X., and Zhang F.. 2024. An updated phylogeny and adaptive evolution within Amaranthaceae s.l. inferred from multiple phylogenomic datasets. Ecology and Evolution 14: e70013. 10.1002/ece3.70013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang, Y. , and Smith S. A.. 2014. Orthology inference in nonmodel organisms using transcriptomes and low‐coverage genomes: Improving accuracy and matrix occupancy for phylogenomics. Molecular Biology and Evolution 31: 3081–3092. 10.1093/molbev/msu245 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang, Y. , Moore M. J., Brockington S. F., Mikenas J., Olivieri J., Walker J. F., and Smith S. A.. 2018. Improved transcriptome sampling pinpoints 26 ancient and more recent polyploidy events in Caryophyllales, including two allopolyploidy events. New Phytologist 217: 855–870. 10.1111/nph.14812 [DOI] [PubMed] [Google Scholar]
- Yardeni, G. , Viruel J., Paris M., Hess J., Groot Crego C., De La Harpe M., Rivera N., et al. 2022. Taxon‐specific or universal? Using target capture to study the evolutionary history of rapid radiations. Molecular Ecology Resources 22: 927–945. 10.1111/1755-0998.13523 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang, C. , and Mirarab S.. 2022a. Weighting by gene tree uncertainty improves accuracy of quartet‐based species trees. Molecular Biology and Evolution 39(12): msac215. 10.1093/molbev/msac215 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang, C. , and Mirarab S.. 2022b. ASTRAL‐Pro 2: Ultrafast species tree reconstruction from multi‐copy gene family trees. Bioinformatics 38: 4949–4950. 10.1093/bioinformatics/btac620 [DOI] [PubMed] [Google Scholar]
- Zhang, C. , Sayyari E., and Mirarab S.. 2017. ASTRAL‐III: Increased scalability and impacts of contracting low support branches. In Meidanis J. and Nakhleh L. [eds.], Comparative Genomics, Lecture Notes in Computer Science, 53–75. Springer International Publishing, Cham, Switzerland. 10.1007/978-3-319-67979-2_4 [DOI] [Google Scholar]
- Zuntini, A. R. , Carruthers T., Maurin O., Bailey P. C., Leempoel K., Brewer G. E., Epitawalage N., et al. 2024. Phylogenomics and the rise of the angiosperms. Nature 629: 843–850. 10.1038/s41586-024-07324-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix S1. Loci extraction report from 29 Amaranthaceae s.s. transcriptomes using the CAPTUS pipeline. The completeness of recovered loci is color‐coded in a gradient from black (0%) to red (100%). (Top) Locus extraction using a target file containing only sequences of clade 1 (Gomphrenoids, Achyranthoids, and Aervoids) in blue. (Bottom) Locus extraction using a target file containing only sequences of clade 2 (Amaranthoids and Celosioids) in yellow.
Appendix S2. Fast‐Plast results showing the percentage of known angiosperm chloroplast genes recovered in 24 samples sequenced with the Amaranthaceae1000 baits.
Appendix S3. Astral‐Pro3 phylogenetic inference with all cleaned homologous trees from 57 transcriptomes and 24 samples sequenced with the Amaranthaceae1000 baits. The tree is rooted on members of the Caryophyllales, and support values on the branches correspond to local posterior probabilities (LPPs).
Appendix S4. Gene duplication mapping results. (Left) Histogram showing the percentage of gene duplications per branch. (Right) Phylogenetic inference from 24 species using ASTRAL IV with “monophyletic outgroup” (MO) ortholog trees, rooted on members of the Caryophyllales. Branch values indicate the proportion of duplicated genes with above 6 in bold. The stars mark the known WGD from Yang et al. (2018).
Appendix S5. Summary statistics of sequencing success obtained from HybPiper2. Metrics include species name, clade, number of reads, mapped reads, percentage mapped to targets, number of mapped genes, and additional gene recovery statistics (e.g., genes with contigs, sequences, or specific coverage thresholds). The table also reports paralog warnings (by length and depth), genes without or with stitched contigs, skipped contigs, chimera warnings, and the total bases recovered.
Data Availability Statement
All generated data are available from the National Center for Biotechnology Information (NCBI; Bioproject PRJNA1209683). The workflow for the bait design is available on Github (https://github.com/tinakiedaisch/bait_design_from_orthologs), and the bait design files are available on Dryad (https://doi.org/10.5061/dryad.k3j9kd5m6; Kiedaisch et al., 2025).
