Abstract
Mapping genotypes to phenotypes is a fundamental goal in biology. Phylogenetic Genotype to Phenotype mapping methods are a relatively new set of tools that aim to identify genomic regions associated with trait variation between species. Here, we review recent developments in Phylogenetic Genotype to Phenotype mapping methods, focusing on three key areas: methods based on replicated substitutions at individual amino acid sites; methods detecting changes in evolutionary rates; and methods analyzing gene duplication and loss. We discuss how trait definition and measurement can impact these methods, as well as the genetic mechanisms that can give rise to trait variation between lineages. We examine the strengths and limitations of different approaches, highlighting the importance of explicit modeling of evolutionary processes. Finally, we outline promising future directions, including the integration of within-species variation, as well as epigenetic and environmental information. Since no single method is likely to identify all genomic regions of interest, we encourage users to apply a representative range of methods that are capable of detecting different types of associations. Overall, this review provides practitioners a roadmap for understanding and applying Phylogenetic Genotype to Phenotype mapping approaches.
Keywords: phylogenomics, forward genomics, convergent evolution, replicated evolution
Significance.
Phylogenetic Genotype to Phenotype mapping (PhyloG2P) methods have developed rapidly over recent years, offering an exciting opportunity to expand genotype to phenotype mapping to trait variation between species. However, PhyloG2P studies have often not engaged with the full complexity of replicated phenotypic and genotypic evolution, and common challenges have not been synthesised into a clear direction for future research. In this review, we address both of these points by providing a detailed overview of the mechanisms of replicated phenotypic and genotypic evolution, and discuss the outstanding challenges facing PhyloG2P research.
The Evolutionary Biology Behind PhyloG2P
Introduction to PhyloG2P and the Relevance of Replicated Evolution
Understanding how genomic changes produce phenotypic differences is a core pursuit in biology. In recent years there have been a growing number of methods that explore this relationship from a phylogenetic perspective, broadly referred to as PhyloG2P (Smith et al. 2020). PhyloG2P methods use reconstructed evolutionary relationships to link genotypic variation between lineages to their variation in phenotype. This evolutionary perspective enables genotype to phenotype mapping in situations where typical population genetics approaches would be challenging, if not impossible to implement (Nagy et al. 2020; Smith et al. 2020).
PhyloG2P research follows two main approaches. The first approach identifies phenotypes that have arisen only once in the evolutionary history of a group of lineages. Phylogenetic methods are used to reconstruct the ancestral genomic and phenotypic states of the group of lineages, identifying genetic changes that are uniquely correlated with the transition to the target phenotype. However, since most lineages will have a substantial number of unique phenotypic features, the identified genetic changes may be associated with any of these features rather than the phenotype of interest (a phenomenon similar to “Darwin’s Scenario”, see Maddison and FitzJohn 2014). To address this challenge, single lineage PhyloG2P studies tend to utilize prior information about the identified genetic elements (genes or noncoding regulatory regions), such as annotations from related model organisms. While successfully applied across a range of organisms, this PhyloG2P approach primarily deepens our understanding of existing genotype to phenotype maps, rather than developing novel ones.
The second approach to PhyloG2P research focuses on phenotypic changes that have evolved independently in multiple lineages (known as replicated evolution, see Box 1). Again, phylogenetic methods are used to reconstruct the evolutionary history of the lineages being studied, now identifying multiple independent phenotypic and genotypic transitions. Utilizing evolutionary replication provides greater power to differentiate lineage-specific genetic changes from those that are shared between some or all of the lineages with the phenotype of interest. Like single lineage studies, this multilineage approach can be complemented with prior knowledge of relevant genetic elements, though such information is not essential for uncovering likely genotype to phenotype associations. While this PhyloG2P approach has some limitations (discussed in the following subsections), it has the significant advantage of being able to identify novel genotype to phenotype associations without requiring detailed prior knowledge about the genetic elements involved.
Box 1 The Language of Replication.
The independent evolution of similar traits or phenotypes in response to similar environmental pressures is commonly known as parallel or convergent evolution, depending on the similarity of the ancestral species or lineages. On a phylogenetic scale, parallel evolution typically refers to independent evolution of similar phenotypes in “closely” related species, while convergent evolution occurs in more “distantly” related species. At the genetic sequence level, parallel evolution is often used to describe identical nucleotide or amino acid substitutions from a shared ancestral state, while convergent evolution involves different ancestral states. At the genetic network scale, parallel evolution involves changes in homologous genetic elements, whereas convergent evolution involves changes in distinct network elements (for more details, see Arendt and Reznick 2008; Cerca 2023; Pereira and Kohlsdorf 2023 and James et al. 2023). However, it is not always clear where the boundaries between these categories of close/similar and distant/distinct lie. Additionally, some systems may contain both parallelism and convergence at different levels, such as distantly related species (convergent on the phylogenetic level) which both have had identical substitutions at a particular nucleotide position from the same ancestral state (parallel on the genetic sequence level). We wish to avoid potential confusion caused by the language of parallelism and convergence. Following James et al. (2023), we use the term “replicated evolution” to refer to all forms of independent evolution of similar phenotypes in response to shared environmental pressures, regardless of the underlying genetic mechanisms.
Both of these approaches facilitate the mapping of genotype to phenotype for traits that only vary between species, which is not possible with existing methods that depend upon the ability to cross individuals with contrasting phenotypes (Smith et al. 2020). However, many available PhyloG2P methods depend on replicated evolution, as these independent evolutionary transitions provide the statistical power to distinguish phenotype-associated genetic changes from background genomic variation. It has also been noted that utilizing replicated evolution provides a key advantage for discovering genotype to phenotype relationships, even in genomic regions with no prior functional characterization (Nagy et al. 2020). We therefore focus this review on PhyloG2P methods applied to phenotypic changes that have independently evolved across multiple lineages, with reference to single lineage studies where relevant.
In this review, we begin by covering the essential aspects of replicated evolution that underpin PhyloG2P analysis involving multiple independent phenotypic transitions. Grossnickle et al. (2024) and Thomas et al. (2024) suggested that PhyloG2P methods would benefit from a more comprehensive consideration of the different biological processes that can cause replication—to this end we first discuss phenotypic replication and the genomic processes that drive it. We describe existing PhyloG2P methods categorized by the type of genetic change they detect: namely, specific amino acid substitutions, molecular evolutionary rate variation, and duplication and loss of genetic elements. We build on the earlier reviews of Rey et al. (2019), Smith et al. (2020), Nagy et al. (2020) and Pereira and Kohlsdorf (2023) and focus on recent developments, particularly in the first two categories of methods described above. We conclude the review by identifying key challenges and areas where current analysis could be improved and highlight promising future directions for PhyloG2P development.
Measuring and Comparing Traits
One of the most critical decisions in any PhyloG2P study is how to define and measure the trait of interest (Speed and Arbuckle 2017; Grossnickle et al. 2024; Redlich et al. 2024). Many PhyloG2P studies analyze traits as a binary presence/absence pattern across lineages, as they can be relatively simple to measure for a large number of individuals or species (Stefen et al. 2022). While presence/absence based studies have successfully identified links between genetic elements and traits (Chikina et al. 2016; Eliason et al. 2023; Zhang et al. 2023), especially in the context of trait loss (Prudent et al. 2016; Partha et al. 2019; Sackton et al. 2019), they may also oversimplify the biological processes under investigation, limiting the insights gained (Lamichhaney et al. 2019). For example, expanding the phylogenetic breadth of a study may reveal additional biological variation, such as more phenotypic categories or intermediate phenotypes between “presence” and “absence” (Redlich et al. 2024). Similarly, measuring traits continuously rather than categorically can enhance statistical power, as for each lineage there is often more information than just the presence/absence of a trait. PhyloG2P studies based on continuous traits can also provide a more detailed understanding of what aspect of the trait the genetic elements control (e.g. length, developmental timing, colour, etc) (Kowalczyk et al. 2020; Baker et al. 2021). Practically, trait definition and measurement are constrained by data availability and sample size, and they are further informed by prior knowledge of the genetic pathways associated with the trait. Overall, research groups using PhyloG2P methods should carefully consider and communicate potential alternative ways of defining and measuring the trait.
The replicated evolution of marine adaptations in mammals illustrates the importance of trait definitions. PhyloG2P studies of these marine adaptations typically focus on three lineages, namely Cetacea (whales and dolphins), Pinnipedia (seals and walruses), and Sirenia (dugongs and manatees), which have all independently transitioned from terrestrial ancestors to a predominantly or entirely marine environment. Many studies have attempted to identify genetic changes associated with this phenotypic transition (e.g. Zhou et al. 2015; Chikina et al. 2016; Hu et al. 2019 and Treaster et al. 2021), highlighting their numerous shared trait changes. However, not all of the traits associated with marine mammals are exclusive to their marine adaptation. For example, the evolved resistance to the stress of hypoxia (low blood oxygen) is important to marine mammals, as it facilitates extended dive duration (Chikina et al. 2016). However, evolved resistance to hypoxia has also been essential for mammals that have transitioned to subterranean or high altitude environments, where oxygen availability is also limited (Wang et al. 2024) (Fig. 1). As such, searching for genetic changes that are exclusively found in marine mammals may fail to detect some genetic elements underpinning hypoxia tolerance, as this adaptation is shared with other nonmarine lineages. Therefore, to fully characterize the genetic changes associated with evolutionary adaptations involving multiple traits, such as those seen in marine mammals, the traits may need to be investigated individually in addition to studying their collective replicated evolution (Speed and Arbuckle 2017; Lamichhaney et al. 2019).
Fig. 1.
A schematic diagram representing a simplified mammalian phylogeny and a table indicating the presence of particular traits. The phylogeny includes three independent mammalian transitions to marine environments (manatee, seal, and whale). The dots in the table correspond to the presence of the trait. The loss of rear limbs is associated with the transition to a marine environment, but only for a subset of marine lineages. Similarly, hypoxia tolerance is also associated with the transition to a marine environment, but it is also shared with multiple other mammalian lineages.
We suggest that the results of PhyloG2P studies will depend in part on where traits lie along a spectrum of complexity. On one end, we can conceptualize “simple” traits with both easily quantifiable phenotypes and a simple genetic basis. Here PhyloG2P studies will have a chance to identify specific genotype to phenotype maps (Lamichhaney et al. 2019; Nagy et al. 2020; Okuyama et al. 2025). On the other end of the spectrum are “compound” traits, that consist of multiple simple traits and resultantly have a complex genetic basis (for example, being a marine mammal). For these more complex traits, PhyloG2P studies may not be able to provide the same level of specificity, but they will still be useful for identifying broader changes in the genetic pathways associated with an evolutionary transition (Chikina et al. 2016; Hu et al. 2019; Zhang et al. 2023; Eastment et al. 2024).
The Genetic Underpinnings of Replicated Changes in Phenotype
Replicated phenotypic changes in independent lineages can be caused by genetic changes which may be replicated at different levels of organization (Fig. 2) (Losos 2011). For instance, identical substitutions (or insertions/deletions) at the same nucleotide position or amino acid may cause identical phenotypic changes (Fig. 2a), such as resistances to a common toxin (Zhen et al. 2012). Changes in different sites within the same gene or regulatory element may also produce a replicated phenotypic change (Fig. 2b). This result could occur when there are multiple ways to achieve a similar change in protein structure (Sykes et al. 2022). Changes in phenotype can also be produced by duplication or loss of entire genes or noncoding regulatory elements (Nagy et al. 2014, 2017). For example, an increase in the number of copies of a genetic element can lead to increased expression levels, which in turn can produce a substantial change in a trait (Klure et al. 2025) (Fig. 2c). Alternatively, lineages may achieve a similar phenotypic change through substitutions in, or duplications of, distinct genetic elements. These distinct genetic elements may be part of the same genetic pathway or network where changes have similar impacts on network function (Láruson et al. 2020; Shakya et al. 2025) (Fig. 2d). Finally, genetic changes in different genetic networks, may produce similar changes in a phenotype (Fig. 2e). If a replicated change in phenotype has been achieved through changes in different genetic elements, then PhyloG2P methods will fail to identify them unless considerable prior knowledge of the genetic elements involved can be leveraged to identify genetic replication at the network or pathway level (Lamichhaney et al. 2019; Wheeler et al. 2022; Eastment et al. 2024; Shakya et al. 2025). In any group of lineages that have independently achieved similar phenotypic changes, some of them may have undergone similar or identical genetic changes, where as others may have achieved the same phenotype through unique genetic changes (Yusuf et al. 2023).
Fig. 2.
Replicated phenotypic changes can result from genetic variation at different organizational levels: a) Nucleotide substitutions at the same site; b) Substitutions at different sites within the same genetic element; c) Copy number variations affecting gene expression; d) Mutations in different genetic elements within the same genetic pathway or network; e) Changes in distinct genetic networks controlling the same trait.
The likelihood of similar phenotypic changes being achieved through similar genetic changes in distinct lineages depends on the degree of pleiotropy (Zhang 2023) (where a single genetic element can affect multiple traits) and genetic redundancy (Láruson et al. 2020) (where multiple different mutations can have similar impacts on a trait) in the genetic elements and networks that control the trait. For genetic networks under a high degree of pleiotropic constraint, only a limited number of genetic changes may produce a particular trait change without causing deleterious changes in other traits controlled by the network (Stern 2013). Traits controlled by pleiotropically constrained networks are ideal for PhyloG2P research, as there is a greater chance that changes in the trait will have been caused by the same genetic change in multiple lineages, improving our power to identify them. Alternatively, if a genetic network controlling a trait has a low pleiotropic constraint, or high genetic redundancy, then genetic changes in different genetic pathways may create similar trait changes (Griesmann et al. 2018). PhyloG2P methods that utilize replicated evolution may struggle in such cases, as there is a low chance of any of the genetic changes being replicated across lineages (Zou and Zhang 2015). Successfully identifying the genetic elements controlling such traits from a PhyloG2P perspective would require prior knowledge of the genetic elements likely to be involved.
Advantages of Replicated Trait Loss in PhyloG2P Studies
Replicated loss of homologous traits has been the focus of many PhyloG2P studies, as it has a high chance of producing replicated genetic changes (Prudent et al. 2016; Sackton et al. 2019). Trait loss may occur when a change in the environment of a lineage makes a trait redundant (Partha et al. 2019), or when a change in development has the side effect of causing a trait to no longer form (Womack et al. 2018; Pereira and Kohlsdorf 2023; Womack and Hoke 2023). Either of these phenomena can lead to a decrease in selective pressure on the genetic elements controlling the trait, provided they are not pleiotropically linked to other traits. While changes in a trait may be caused by a small number of genetic changes as described above, over time the loss of selective pressure associated with trait loss will lead to the accumulation of various mutational changes across all impacted homologous genetic elements. As a result, when a homologous trait has been lost multiple times in independent lineages, there is a relatively high chance of replicated increases in substitution rate, or reductions in copy number of genetic elements associated with the trait (Hiller et al. 2012; Nagy et al. 2017; Zhang et al. 2023).
The concepts we have outlined in this section provide the background upon which PhyloG2P research has been developed. The various PhyloG2P methods each account for or model these concepts in different ways. Later in this review, we explore the detailed implications of these different PhyloG2P approaches with respect to the concepts described in this section. First, we provide a brief history and broad overview of currently available PhyloG2P methods.
An Overview of Current PhyloG2P Methods
Origins of PhyloG2P Methods
PhyloG2P methods have seen a surge of development and application over the past decade. Phylogenetic methods for identifying correlations between phenotypic and genotypic changes were first developed almost 30 years ago (Zhang and Kumar 1997; Yang 1998), and several more were developed over the following 15 years (Castoe et al. 2009; O’Connor and Mundy 2009; Pollard et al. 2010; Lartillot and Poujol 2011; Hiller et al. 2012). Many of these early methods saw limited application, largely because of their poor statistical properties, such as simplistic null models leading to poor false positive rates. However, as sequenced genomes rapidly increased in availability over the past 20 years, it became clear that there was a need for statistical analysis methods that could be applied genome-wide while maintaining low error rates (Nagy et al. 2020). This need prompted the development of several new and updated phylogenetic methods that could be applied to this genomic data (Zou and Zhang 2015; Prudent et al. 2016; Rey et al. 2018; Hu et al. 2019; Kowalczyk et al. 2019; Treaster et al. 2021; Duchemin et al. 2022; Fukushima and Pollock 2023; Morel et al. 2024) in the field that would later come to be called PhyloG2P (Smith et al. 2020). Most of the recently developed PhyloG2P methods identify correlations between replicated phenotypic transitions and replicated genetic changes. Each method uses a different statistical framework to determine the significance of identified correlations, based on their defined criteria for what constitutes replicated evolution.
We argue that PhyloG2P methods fall into three broad categories based on the type of genetic changes they identify. The first group search for replication at individual amino acid sites. There are a variety of approaches taken to determine which amino acid changes constitute replication, and when observed replication is significantly greater than expected by chance. The second group of approaches search for either individual or replicated changes in the rate of genetic evolution. Decreases in rate can be associated with novel constraints on genetic elements, while increases in rate can be indicative of directional or relaxed selection. The third group search for duplications or losses of genetic elements, encompassing a variety of genetic phenomena and phenotypic transitions associated with these changes. Here, we provide a summary of the current PhyloG2P methods for each of these categories (Table 1), building on the earlier reviews of Rey et al. (2019), Smith et al. (2020), Nagy et al. (2020) and Pereira and Kohlsdorf (2023).
Table 1.
Current PhyloG2P methods grouped into three categories
| Individual Amino Acids | ||||
|---|---|---|---|---|
| Model | Description | Noncoding Regions | Relaxed Match | Example Applications |
| Zou2015, WGCCRR | Identify genes with a greater number of replicated amino acid substitutions than would be expected by chance (Zou and Zhang 2015; Dong et al. 2024) | ✕ | ✕ | Wang et al. (2024) |
| Castoe2009 | Compares the number of replicated substitutions to the number of divergent substitutions between two lineages (Castoe et al. 2009) | ✕ | ✕ | Thomas and Hahn (2015), Xu et al. (2021) and Yusuf et al. (2023) |
| CAAStools | Similar to Zou2015, but without the strict requirement that all foreground lineages must have the same amino acid (Barteri et al. 2023) | ✕ | ✓ | |
| CONDOR | Web application similar Zou2015, but with the option to only require some foreground lineages to have replicated substitutions (Morel et al. 2024) | ✕ | ✓ | |
| Chabrol2018 | Determines the likelihood of replicated amino acid substitutions having occurred (Chabrol et al. 2018) | ✕ | ✓ | |
| Pelican | Identifies amino acid positions where a change in amino acid profile has occurred for foreground lineages (Duchemin et al. 2022) | ✕ | ✓ | Vargas-Chavez et al. (2025) |
| PCOC | Similar to Pelican, but with added requirement for a change in amino acid to have occurred in foreground lineages (Rey et al. 2018) | ✕ | ✓ | Eliason et al. (2023) and Zhang et al. (2023) |
| CSUBST | Compares replicated synonymous and nonsynonymous substitutions to determine significance (Fukushima and Pollock 2023) | ✕ | ✓ | Sadanandan et al. (2023), Matsuda and Makino (2024) and Morales et al. (2024) |
| Evolutionary Rates | ||||
|---|---|---|---|---|
| Model | Description | Noncoding Regions | Cont. Traits | Example Applications |
| PAML | Compares ratio of synonymous to nonsynonymous substitutions in foreground and background lineages (Yang 2007) | ✕ | ✕ | Eliason et al. (2023), Zhang et al. (2023) and Wang et al. (2024) |
| phyloP | Generalized framework for testing hypotheses about the rate of sequence evolution (Pollard et al. 2010) | ✓ | ✕ | Feigin et al. (2019) and Sakamoto et al. (2024) |
| RELAX, aBSREL, TraitRELAX | Detailed methods for determine the type of selective pressure a gene has evolved under (see text for citations) | ✕ | ✕ | Sadanandan et al. (2023), Zhang et al. (2023) and Eliason et al. (2023) |
| Forward Genomics | Identifies genetic elements with a high level of sequence divergence in lineages that have lost a trait (Prudent et al. 2016) | ✓ | ✕ | |
| RERconverge | Identifies genetic elements with significantly different relative evolutionary rate in lineages with a replicated trait (Kowalczyk et al. 2019) | ✓ | ✓ | Eliason et al. (2023), Morales et al. (2024) and Zhang et al. (2023) |
| TRACCER | Similar to RERconverge, but based on pairwise comparisons with more closely related pairs weighted higher (Treaster et al. 2021) | ✓ | ✕ | |
| PhyloAcc | Nested likelihood model for evolutionary rates of noncoding regions (Thomas et al. 2024) | ✓ | ✓ | Sackton et al. (2019) and Eastment et al. (2024) |
| Duplication and Loss of Genetic Elements | ||||
|---|---|---|---|---|
| Model | Description | Noncoding Regions | Dup. and/or Loss | Example Applications |
| COMPARE, CAFE | Identifies gene families that undergo duplications or losses correlated with phenotypic changes (Nagy et al. 2014), (Mendes et al. 2020) | ✓ | Both | Nagy et al. (2017) and Zhang et al. (2023) |
| REforge | Uses a model of transcription factor binding to determine when a binding site has deteriorated in lineages that have lost a trait (Langer et al. 2018) | ✓ | Loss | |
| Evolink | Associated presence/absence of genetic elements with presence/absence of traits in microbial lineages (Yang and Jiang 2023) | ✓ | Loss | |
Noncoding Regions: method can be applied to noncoding regions. Relaxed Match: Match between trait and amino acid does not have to be perfect. Continuous Traits: method can be applied to continuous traits. Duplication and/or Loss: if method can detect both phenomena or just loss. Note, some methods do not have example applications apart from the original paper.
In our discussion, we use the language of “foreground” lineages to refer to groups of independent lineages which have repeatedly evolved a particular phenotype of interest and “background” lineages to refer to all lineages without the replicated phenotype. In the context of a PhyloG2P study based on a phenotypic transition that has only occurred once, the foreground will typically be a single species or clade, and the background will be all other species included. This language can also be extended to foreground and background branches to refer to internal branches of the phylogeny for which an estimate of the trait value is made.
Replicated Amino Acid Substitutions
PhyloG2P methods based on replication at the level of single amino acids (Fig. 2a) search for sites with repeated amino acid substitutions in lineages with a replicated binary phenotype of interest. These methods aim to identify coding regions which show an excess of repeated amino acid substitutions (Zhang and Kumar 1997; Chabrol et al. 2018; Fukushima and Pollock 2023; Morel et al. 2024), with each approach differing in how they define replicated changes in amino acids. The earliest statistical method of identifying this form of genetic replication was introduced by Zhang and Kumar (1997), and subsequently revised by Zou and Zhang (2015). These approaches use reconstructed ancestral sequences to identify sites that have the same amino acid in the foreground lineages, which is different to the amino acid of the most recent ancestor of each foreground lineage. The significance of observed amino acid replication is based on the expected number of replicated substitutions given the amino acids likely to be observed at that site (see Zou and Zhang 2015 for a revised approach to this process). Recently, Whole Genome Comparative Coding Region Read (WGCCRR, Dong et al. 2024) has been developed as a web-based tool that can facilitate this type of PhyloG2P analysis. However, the WGCCRR framework has the additional condition of requiring background lineages to have different amino acids to those present in the foreground lineages (see further discussion in the following section).
Castoe et al. (2009) propose a method with a similar definition of replication to Zhang and Kumar (1997), but with a different approach to identify statistical significance. In addition to counting the number of replicated changes of the same amino acid, they also count the number of divergent changes, where the foreground lineages have different amino acids at a particular site. They take the ratio of these values as their test statistic and construct a null distribution by computing it many times for random combinations of lineages. They argue that if the ratio calculated for the foreground lineages is significantly greater than the ratios calculated for other groups of species, then the observed replication in the foreground lineages is significant. While the original study by Castoe et al. (2009) focused on replication across the whole mitochondrial genome of reptiles, subsequent studies have used the same approach to identify specific genes associated with repeated phenotypic evolution (Thomas and Hahn 2015; Xu et al. 2021).
The methods of both Zou and Zhang (2015) and Castoe et al. (2009) require that all of the foreground lineages share the same amino acid at a site for it to be identified as a replicated substitution. However, there are cases where replicated changes in protein function arise from different amino acids at the same site (Zhen et al. 2012; Mohammadi et al. 2022). An approach to expand PhyloG2P methods to detect this form of replication is to relax the requirement for each foreground lineage to have evolved the same amino acid. For example, this relaxation can be achieved by running the above two methods multiple times for different subsets of the foreground lineages. However this approach is computationally intensive for studies with a large number of foreground lineages (Fukushima and Pollock 2023). Convergent Amino Acid Substitutions tools (CAAStools, Barteri et al. 2023) provides an expansion of the method of Zou and Zhang (2015) and can allow for the foreground amino acids to not all be identical. While the foreground lineages are not required to all have the same amino acid substitution, CAAStools has the requirement that the amino acid present for all foreground lineages must be different to the amino acid present in background lineages at that site. The ConDor workflow (Morel et al. 2024, available as a web-based tool) can also facilitate some foreground lineages not sharing a particular amino acid by allowing the user to specify a minimum number of lineages required to share the same amino acid substitution. Their approach was designed to account for potential uncertainty in phenotypic classification, which can be common in viral studies, and hence, unalike CAAStools, allows for some foreground lineages to have the same amino acid as the background lineages. In contrast, Chabrol et al. (2018) developed a far more relaxed method that inherently accommodates situations where some foreground species have different amino acids. Their approach steps away from explicit ancestral sequence reconstruction and instead uses a Continuous Time Markov Model of sequence evolution to determine the likelihood of replicated amino acid transitions occurring in the phylogeny branches leading to the foreground lineages.
Another way to relax the requirement for foreground lineages to have the same amino acid is found in amino acid profile methods (Rey et al. 2019). These methods identify sites where the foreground lineages have a different amino acid preference (e.g. a preference for a hydrophobic versus hydrophilic amino-acid) compared to the background lineages. While methods like ConDor focus solely on the most common amino acid in the foreground lineages, profile methods use all the amino acids in the foreground lineage to determine changes in amino acid profile. This form of analysis is efficiently implemented in Pelican (Duchemin et al. 2022). The Profile Change with One Change (PCOC) method (Rey et al. 2018) expands on this approach by simultaneously implementing a requirement of all foreground lineages undergoing at least one change in amino acid at the site (the “OC” component).
The recently developed CSUBST model (Fukushima and Pollock 2023) offers a novel approach to analyze replicated amino acid substitutions. CSUBST uniquely examines replicated codon changes by independently analyzing synonymous and nonsynonymous substitutions. For each gene in the alignment, the method divides the observed rate by the expected rate of replicated synonymous and nonsynonymous substitutions respectively (). In a manner inspired by the ω metric of selective pressure (Yang 1998), the method takes the ratio of the relative excess of nonsynonymous and synonymous replicated substitutions, producing the metric . By accounting for synonymous substitutions, this metric is robust to errors in topology which can confound the other methods described above (Mendes and Hahn 2016; Mendes et al. 2019). Additionally, as the metric accounts for the expected number of both synonymous and nonsynonymous replicated substitutions, it is also robust to poor estimations of those expected numbers, which can pose a challenge for several of the other methods (Zou and Zhang 2015).
Changes in Evolutionary Rates
PhyloG2P methods based on changes in evolutionary rates search for genetic elements which evolve at different rates in foreground and background lineages. These methods can identify genetic elements that either experience decreased evolutionary rates due to new genetic constraints, or increased rates following relaxation of constraint (as can occur with trait loss). Additionally, in the context of replicated phenotypic transitions, evolutionary rates-based methods can identify genetic elements experiencing positive selection in which multiple different mutations could cause a similar change in phenotype (Fig. 2b), which may not be identified by methods based on replicated changes at single sites (Treaster et al. 2021). These methods vary in how they measure evolutionary rate, what types of rate changes they identify, and which phylogenetic branches they consider for these changes. One of the earliest implementations of this type of method in a PhyloG2P context was the branch model of Yang (1998) (available in PAML, Yang 2007). The branch model works by analyzing variation in the ratio of rates of nonsynonymous to synonymous substitutions in coding sequences (ω) between foreground and background branches of the phylogeny. As synonymous substitutions are expected to arise in a selectively neutral manner, an ω value of 1 indicates neutral evolution of the gene in a particular branch. indicates that the gene is under purifying selection, while is a sign of potential positive selection. A likelihood test can compare whether a model with distinct ω values for foreground and background branches fits the genetic data better than a single ω value for the whole phylogeny. Genes which have a higher likelihood under the two rates model likely experienced a change in selective regime following or during the phenotypic transition, and hence may be associated with the trait. The phyloP program (Pollard et al. 2010) extends the basic principles of the likelihood test to general substitution rates. Using substitution rates instead of ω allows for the analysis of noncoding regulatory elements, albeit without the same information about the selective pressure acting on the genetic element.
There are a number of other models which utilize ω as a measure of selective pressure. These methods are particularly useful for providing a detailed characterization of the evolutionary history of candidate coding regions identified by other PhyloG2P methods, potentially uncovering evidence of adaptation. One of the motivators for these methods is that not all codons in a gene will be under the same level of constraint, and hence will likely have different values of ω. Branch-site models (Zhang et al. 2005, also available in PAML) account for this variation in constraint by allowing each codon to have its own branch model. However, the branch and branch-site models impose specific permissible ω values for the background and foreground lineages. To address this limitation, both the RELAX (Wertheim et al. 2015) and adaptive branch-site random effects likelihood (aBSREL, Smith et al. 2015) models assign each branch a discrete distribution of ω parameters and each codon can evolve under one of their rates. RELAX allows three ω parameters per branch, centered around a value of . A narrow distribution around indicates that most sites are evolving under approximately neutral selection, while a broad distribution indicates both constraint and directional selection are present. aBSREL assigns a variable number of ω parameters to each branch dependent on the branch length, and places no constraint on the ω values taken. Replicated changes in the ω distributions of either RELAX or aBSREL can then be associated with replicated phenotypic transitions, with the particular distributions providing more detail as to the selective pressures acting on the genes.
Another motivation for more complex models of selective pressure is that trait changes do not necessarily occur at the point of speciation. The models discussed above all implicitly make this assumption as they assign whole branches as having a particular trait, which can cause erroneous inference of the selective pressures acting on a genetic element (Halabi et al. 2021). One way to avoid making that assumption is to jointly model both trait and sequence evolution. In this framework, a continuous time Markov Model of trait evolution is used to infer when changes in phenotype may have occurred (potentially part way along a phylogenetic branch), with the evolutionary rate dependent on the trait value. The joint modeling approach was implemented by O’Connor and Mundy (2009) for binary traits, and in CoEvol (Lartillot and Poujol 2011) for continuous traits. This approach was combined with the consideration for between-site variation discussed above in TraitRepProp (Levy Karin et al. 2017), and further extended to account for uncertainty in phenotypic reconstructions in TraitRELAX (Halabi et al. 2021). Jointly modeling trait and sequence evolution requires substantial computation, which has largely prevented these models from being applied in a genome-wide context.
The first rates-based method to be explicitly developed for whole genome applications without requiring prior knowledge of the genetic elements associated with the trait was Forward Genomics (Hiller et al. 2012; Prudent et al. 2016). Forward Genomics specifically identifies genetic elements which have evolved neutrally following replicated loss of a trait. The method works by reconstructing ancestral sequences and searching for genetic elements where foreground lineages have accumulated a relatively large number of changes from the ancestral sequence. The method was revised to account for the background evolutionary rate of lineages, and to apply a more rigorous test of statistical significance (Prudent et al. 2016). The Forward Genomics approach was flipped in the “Reverse Genomics” approach of Marcovitz et al. (2016), which searches for genetic elements that had repeatedly accumulated large numbers of substitutions in independent lineages. A trait database is then used to search for traits with a presence/absence pattern that matches the pattern of the presence/absence of divergence for a given genetic element.
RERconverge (Chikina et al. 2016; Kowalczyk et al. 2019) is a widely used method that was also developed for whole genome scans without the requirement for prior knowledge of the genetic elements associated with the trait. RERconverge works by constructing a model of the evolutionary rate of each branch accounting for a range of statistical and evolutionary properties (Partha et al. 2019). This model is then compared to the branch specific evolutionary rates inferred from the alignment for each genetic element, with the difference in value being termed the relative evolutionary rate (RER). Ancestral state reconstruction is used to determine the foreground branches with the trait, and the RERs of the foreground branches are compared to the RERs of the background branches for each genetic element. If there is a strong correlation, with the foreground branch RERs being either greater than or less than the background RERs, then the genetic element is inferred to be associated with the trait. The statistical significance of this correlation is determined through extensive simulations of permuted phylogenies (called “permulations”) which aim to capture the distribution of RERs expected to be observed in the context of replicated evolution (Saputra et al. 2021). This approach can be applied to both coding and noncoding genetic elements (Kowalczyk et al. 2022), as well as continuous (Kowalczyk et al. 2020) and categorical traits (Redlich et al. 2024) (see further discussion in the following section).
While RERconverge has been applied widely, Treaster et al. (2021) identified two potential issues with the framework which they attempt to address with their model TRACCER. The first issue is that RERconverge requires branches to be assigned as foreground or background through ancestral state reconstruction, which can be challenging for many traits (see further discussion in the following section). TRACCER instead compares extant lineages with and without the phenotype in a pairwise manner from their most recent common ancestor, removing the requirement to determine which sections of their evolutionary history did and did not have the replicated phenotype. The second issue is that under the RERconverge framework (and many other rates-based methods) the RERs of all background and foreground lineages are weighted equally when determining correlations. Treaster et al. (2021) argue that closely related pairs of foreground and background lineages will have less sequence divergence than distantly related ones, and hence the differences observed between them are more likely to be associated with the replicated trait. They further argue that if a background clade that is not closely related to any foreground lineages is highly sampled, then providing equal weighting to the RERs of that clade can mislead correlations of RER values. To address this issue, TRACCER weights RER significance inversely to the branch length separating the two lineages being compared. Even with these two changes, TRACCER is still likely not entirely robust to the issue of trait change timing discussed above (Halabi et al. 2021), but it may be less vulnerable to issues with ancestral state reconstruction than RERconverge and some of the other rates-based methods.
Along with RERconverge, PhyloAcc (Hu et al. 2019; Thomas et al. 2024) is one of the more thoroughly developed frameworks for PhyloG2P analysis, with a focus on noncoding genetic elements. PhyloAcc uses a Bayesian approach with three nested models to detect replicated acceleration of genetic evolution in foreground lineages (as may be observed following trait loss Sackton et al. 2019). The simplest model (M0) assumes that all branches of the phylogeny evolve under neutral or constrained evolution, and does not permit any acceleration. The M0 model is nested within the M1 model, which allows for acceleration only on the foreground branches. The M1 model is again nested within the M2 model which allows for acceleration on any combination of branches. All three of these models are fit to a genetic element, and the Bayes factors are compared. If M1 is preferred over M0, then that is evidence that acceleration has occurred, and if M1 is preferred over M2, then it is evidence that it has only occurred in foreground lineages. The second comparison with M2 prevents potential false positives resulting from acceleration associated with a nonforeground trait, which are often not accounted for in other rates-based methods. This approach has been expanded to account for gene tree discordance (Yan et al. 2023), and to model continuous traits (Gemmell et al. 2024), which are both discussed further in the following section.
Duplication and Loss of Genetic Elements
The final main group of PhyloG2P methods search for genetic elements which are repeatedly duplicated or lost in lineages with a replicated phenotype (Fig. 2c). Increases in gene family size through duplication may be associated with trait-gain (Nagy et al. 2017; Griesmann et al. 2018; Zhang et al. 2023), while decreases may be associated with trait-loss (Nagy et al. 2017; Eastment et al. 2024). Models such as Computation Analysis of gene Family Evolution (CAFE, Mendes et al. 2020) reconstruct the phylogenetic history of gene family expansion and contraction. Replicated changes in gene families occurring in foreground branches can indicate that the gene family plays a role in controlling the trait of interest. The comparative phylogenomic analysis of trait evolution (COMPARE, Nagy et al. 2014) framework is specifically designed to facilitate this type of PhyloG2P analysis. For example, it can be used to identify a gene family which underwent significant expansion in an ancestral lineage that gained a new trait, and subsequently contracted in descendant lineages which have lost the trait (Nagy et al. 2017).
The loss of transcription factor binding sites (TFBSs) is a special case of genetic element loss. TFBSs are often contained within larger conserved noncoding regions, and can be difficult to anayze using DNA alignment based methods. If a TFBS acquires several substitutions that change its binding properties in a particular lineage, then a transcription factor that targets it may no longer bind to the regulatory region. However, if substitutions in another section of the regulatory region create a new TFBS that has similar binding properties to the original TFBS, then the transcription factor may still be able to bind to the regulatory region following mutations in the original binding site. Analysis of a DNA alignment of the region would suggest that the TFBS has been lost in this particular lineage, but the properties of the regulatory element have not actually changed. To facilitate analysis of TFBS loss, Regulatory Element forward genomics (REforge Langer et al. 2018) calculates a binding affinity score between a set of transcription factors and an alignment of conserved regulatory elements. These scores are also calculated for reconstructed ancestral sequences, providing an estimate for the phylogenetic history of binding affinity. Replicated decreases in binding affinity can then be associated with replicated loss of a trait. REforge was demonstrated to perform better at identifying regulatory elements associated with trait loss than methods based on evolutionary rates. However, it also requires prior knowledge of transcription factors and their binding motifs, which makes it challenging to apply REforge to traits that have not been studied previously.
Finally, we would like to briefly discuss phylogenetic profiling methods (Dembech et al. 2023). These methods are similar to PhyloG2P methods, except that they search for associations between genetic elements rather than between a genetic element and a phenotype. These methods identify genetic elements which are repeatedly lost or conserved together, which indicates that they may have a common function. They are often best suited for analysis of genetic elements associated with fundamental metabolic processes, and can be applied to multiple-kingdom alignments. The basic idea of phylogenetic profiling can be expanded to PhyloG2P analysis by associating the presence and absence of a trait with the presence and absence of genetic elements. This is the idea behind Evolink (Yang and Jiang 2023), which is specifically designed for PhyloG2P analysis of alignments of tens of thousands of microbial lineages. While the simplistic presence/absence approach would likely fail to detect many of the genetic elements associated with phenotypic change in multicellular organisms with complex genomes, they demonstrated it’s effectiveness on microbial lineages with comparatively simpler genomes. The simplified approach also leads to a substantial reduction in computational intensity, which is a barrier to analysis of such large alignments.
Additional Considerations for Applying PhyloG2P Methods
Alternatives to Trait Presence/Absence
The majority of PhyloG2P methods outlined above focus on binary trait definitions such as presence/absence or retained/lost. This focus is not surprising, as such trait data is relatively easy to collect and model, and can still provide valuable insights into the genetic processes that control the trait. However, as discussed above, binary representations of traits often simplify the underlying biology, potentially limiting the insights that can be gained by studying them (Lamichhaney et al. 2019).
To address the limitations posed by binary trait PhyloG2P methods, recent developments have produced methods capable of analyzing continuous valued traits on a genome-wide scale (Kowalczyk et al. 2020; Gemmell et al. 2024). RERconverge has introduced an approach to correlate branch RER values with the change in trait value on those branches (Kowalczyk et al. 2020). This framework identifies correlations between any combination of high/low RER scores and increases/decreases in trait values. Kowalczyk et al. (2020) argued that distinct sets of genetic elements would likely be involved in increases and decreases in trait value, and hence chose to consider increases and decreases separately. A similar directional approach was taken by Tejada-Martinez et al. (2022) in their study of longevity in great apes, where they looked for correlation between evolutionary rate and maximum lifespan. In contrast, Baker et al. (2021) searched for correlations between evolutionary rate and both trait value and absolute change in trait value. They argued that some genetic elements may be able to drive both increases and decreases in trait value, hence using the absolute change in trait value when searching for associations. For example, a high evolutionary rate could cause both large increases and decreases in the trait value for different lineages. A similar approach for analyzing continuous traits was taken by Gemmell et al. (2024) in the PhyloAcc framework (called PhyloAcc-C), where they model changes in continuous trait value as a form of Brownian motion. The rate of this Brownian motion is set by which of the three PhyloAcc rate classes the lineage is evolving under.
PhyloAcc-C will potentially identify similar correlations between evolutionary rate and absolute change in trait value as Baker et al. (2021), as the increased Brownian rate is more likely to produce both large increases and decreases in trait value. Even though the directional and nondirectional approaches aim to identify distinct sets of genetic elements associated with the continuous trait, Gemmell et al. (2024) found a considerable overlap between their results and the results of Kowalczyk et al. (2020). While more studies employing these new continuous trait methods will be required to fully understand how they may complement each other, they are a promising start to increase the range of traits that PhyloG2P analysis can be applied to. It is also notable that PhyloAcc-C is one of the few PhyloG2P methods that incorporates an explicit model of trait evolution. While there is evidence that this approach may provide some benefits over methods that depend upon predefined specification of trait values (Halabi et al. 2021), the full impacts of this modeling choice remain relatively unexplored.
Compared to continuous traits, development of PhyloG2P methods for nonbinary categorical traits has seen less focus. The δ (delta) statistic of Ribeiro et al. (2023) is one approach that is able to identify associations between genetic elements and variation in a categorical trait. The δ statistic measures the uncertainty in ancestral state reconstruction of internal nodes based on a categorical trait for a given phylogeny (such as a gene tree). A low level of uncertainty in ancestral state reconstruction suggests a strong phylogenetic signal, indicating that the genetic element may be associated with the trait of interest. However, it is important to note that the δ statistic was not developed explicitly for PhyloG2P applications, and a high confidence in reconstructed ancestral states could be produced by phylogenies of genes unrelated to the trait. The RERconverge framework (which is explicitly designed for PhyloG2P applications) has also been expanded to work with categorical traits by using a Continuous Time Markov Model of trait evolution to reconstruct ancestral states (Redlich et al. 2024). This approach conducts pairwise RER comparisons between all combinations of the traits. The differences in evolutionary rate are assessed using a revised version of their “permulation” approach to determine statistical significance, and then combined in an omnibus test (similar to a traditional ANOVA) to identify genetic elements that are correlated with the trait.
The δ statistic detects patterns of evolutionary rate change that are different to the patterns identified by RERconverge. The δ statistic will be large if a genetic element has an increased rate of evolution during transitions between trait values, as the ancestral states of the trait will be inferred with greater certainty. By comparison, RERconverge explicitly identifies genetic elements that evolve at different rates depending on the trait value of the lineage. That is, the δ statistic identifies higher rates of genetic change associated with transitions between trait values, while RERconverge identifies different rates of genetic change associated with different trait values. It has previously been found that these two approaches tend to identify the same genetic elements for a binary trait with few transitions in phenotype (Kowalczyk et al. 2022), but it is not yet clear if the same will be true for nonbinary categorical traits with many ancestral transitions in phenotype (Treaster et al. 2021). Redlich et al. (2024) found that the δ statistic and RERconverge generally identified distinct sets of genetic elements, suggesting that further research into the impacts of these two approaches to modeling changes in evolutionary rate may be required.
In addition to suitable PhyloG2P methods, expanding PhyloG2P analysis beyond presence/absence studies requires the availability of continuous or categorical trait data (Lamichhaney et al. 2019). Accessing high quality phenotypic data is a common challenge for genomics studies, as methods of recording phenotypic data often vary between species and are rarely compiled in a computer-accessible manner (Stefen et al. 2022). This challenge is beginning to be alleviated through large-scale digitization efforts (Lamichhaney et al. 2019). One example of this digitization process is the MaTrics database of mammalian trait data, which has been specifically created for PhyloG2P applications (Stefen et al. 2022). The creators of MaTrics specifically chose to not include continuous traits in their database, arguing that PhyloG2P methods have largely focussed on presence/absence data. However, as is now becoming clear with the recent studies discussed above, there is use for both discrete and continuous trait data in PhyloG2P research, and the two approaches will likely complement each other. We encourage any groups that are digitizing phenotypic data to include both discrete and continuous measurements in their databases where possible, as exemplified in the AusTraits database (Falster et al. 2021).
Cautionary Tales of Tree Topology
An issue that has arisen in some PhyloG2P studies is using phylogenetic models that do not accurately represent the type of genetic replication being searched for. One of the most widely identified and discussed occurrences of this issue is the association of genes with the replicated evolution of echolocation in bats and toothed whales by Parker et al. (2013) (see (Thomas and Hahn 2015; Rey et al. 2018) for subsequent discussion). Parker et al. (2013) argued that echolocation related genes in bats and toothed whales would likely share more similarities than expected based on the species tree. To attempt to model this evolved similarity, they searched for genes that had higher likelihood values under a phylogeny that grouped the foreground lineages into a clade sister to other mammalian lineages than their likelihood under the species phylogeny. While Parker et al. (2013) intend for this model to represent the biological process of echo-locating lineages independently evolving similar genetic changes, it is instead more likely to represent an alternative hypothesis of the echolocation related genes having a distinct evolutionary history to that of the species tree. It is unclear what biological process this distinct evolutionary history would represent, and as such it is difficult to interpret the evolutionary significance of any genes which are identified with this method. The alternate topology approach also assumes that the entirety of a genetic element is expected to be more similar between foreground lineages, when phenotypic replication may only require substitutions in a subset of sites (Levy Karin et al. 2017). We argue that attempts to identify evolved genetic similarity should use approaches that model it explicitly, such as models of replicated amino acid substitution.
A related issue can occur when incomplete lineage sorting (ILS) or introgression has resulted in apparent replication in phenotypes (Smith et al. 2020). When genetic elements have a history of ILS (such as selection on standing variation) or introgression, the phylogeny for the genetic element will differ from the species tree (referred to as gene tree discordance). Forcing such genetic elements onto the species tree can result in apparent increases in substitution rate on branches that differ between the two trees (Mendes and Hahn 2016; Mendes et al. 2019). Evolutionary rates based PhyloG2P methods can be used to identify genetic elements under these circumstances. However, this approach would again make it challenging to interpret the evolutionary significance of any genetic elements identified, as these models assume that changes in the genetic element have arisen independently when they have not. When it is possible that ILS or introgression have been involved in phenotypic replication, it may be more appropriate to model those phenomena explicitly. A recent development in PhyloAcc, named PhyloAcc-GT (Yan et al. 2023), analyses replicated acceleration in evolutionary rates while accounting for the impacts of potential gene tree discordance using a multispecies coalescent model, allowing these two phenomena to be disentangled. More broadly, the capacity of PhyloG2P methods to handle introgression may be improved by incorporating phylogenetic networks, which can explicitly model a range of phenomena associated with hybridization (Bastide et al. 2018).
The Nuances of Replicated Amino Acid Substitutions
There has been much discussion as to the nature of replicated amino acid substitutions and how they should be modeled in the context of replicated phenotypic evolution (Thomas and Hahn 2015; Zhou et al. 2015; Zou and Zhang 2015; Thomas et al. 2017; Rey et al. 2018; Morel et al. 2024). Central to this discussion is the notion of necessity and sufficiency of particular amino acids underlying the replicated phenotype. Necessity describes how a phenotype cannot arise without a specific amino acid substitution, requiring all foreground lineages to share this substitution. On the other hand, sufficiency indicates that any lineage with the amino acid will have the phenotype and hence no background lineages should have this substitution. While these assumptions may hold in specific contexts (such as small viral genomes under strong constraint, Morel et al. 2024), we argue that they are likely too restrictive to apply to most biological contexts. Necessity does not hold in cases where multiple amino acids at the same site or substitutions at different sites can produce a similar change in phenotype (Rey et al. 2018; Treaster et al. 2021; Sykes et al. 2022). This restrictive assumption also undermines the replicated/divergent distinction of Castoe et al. (2009) and Thomas and Hahn (2015), as they implicitly assume that a replicated trait change cannot be caused by two different amino acids in the same position. Likewise, sufficiency does not hold when multiple substitutions are required to produce the replicated phenotype, or where the phenotypic effect of particular amino acid substitutions depends on the genetic and environmental context. Model choices regarding necessity and sufficiency will largely come down to balancing false positives and negatives in the identified genotype to phenotype associations. Methods that implement strict assumptions (such as Castoe et al. 2009, CSUBST Fukushima and Pollock 2023 or WGCCRR Dong et al. 2024) will likely be less susceptible to false positives, while methods that do not rely on as strict assumptions (such as Chabrol et al. 2018, PCOC Rey et al. 2018 or CONDOR Morel et al. 2024) are less likely to produce false negatives. Regardless of the model assumptions, it has also been identified that false positives can be reduced by including a large number of lineages in both the foreground and background (Thomas et al. 2017).
Challenges of Ancestral State Reconstruction and Alignment of Genetic Elements
A major caveat that applies to most PhyloG2P methods is that they rely, explicitly or implicitly, on accurate ancestral state reconstruction of the trait of interest to differentiate foreground and background lineages. However, most reconstruction methods assume that the trait of interest is evolving neutrally, which can lead to inaccuracies when this assumption does not hold (Holland et al. 2020). For example, if a trait has undergone replicated changes in many lineages then the ancestor of those lineages may be incorrectly inferred to have the replicated trait (Sackton et al. 2019; Thomas et al. 2024).
The challenges associated with ancestral state reconstruction do not impact all PhyloG2P methods equally. For instance, ambiguity about the ancestral state is unlikely for replicated degradation or loss of a complex trait. In such a case, once a lineage accumulates multiple degrading substitutions in a genetic element that underlies the lost trait, there is very little chance of those substitutions being reversed so as to re-evolve the trait. Loss of complex traits is likely to follow Dollo’s law of irreversibility, where we can be confident that the most recent common ancestor of lineages with and without the trait possessed the trait. Conversely, it may be challenging to reconstruct ancestral states of replicated transitions between various states of a categorically valued complex trait, such as ecomorphs (Corbett-Detig et al. 2020; Morales et al. 2024). Without reliable estimation of ancestral states, it is unclear where particular phenotypic transitions (and hence genotypic changes) occurred on the phylogeny. Some studies have attempted to avoid this issue by focusing on the branches leading to extant lineages (Corbett-Detig et al. 2020; Morales et al. 2024). However, this approach risks overlooking genotypic changes that occurred on internal branches of the tree, potentially creating false negatives which would significantly reduce the power of the analysis (Chabrol et al. 2018; Halabi et al. 2021; Treaster et al. 2021). Individually analyzing the component traits involved in complex phenotypic adaptations may improve the capacity to infer ancestral states, and hence the power of PhyloG2P analysis. This is especially true for continuous valued traits, where uncertainty can be accounted for by using a phylogenetic generalized least squares approach (Baker et al. 2021; Tejada-Martinez et al. 2022). Additionally, incorporating fossil records (Lamichhaney et al. 2019) or methods that do not heavily rely on ancestral state reconstruction (Treaster et al. 2021) may also help. However, uncertainty in ancestral state reconstruction still represents an important limitation, and it’s potential impacts should be considered carefully when drawing conclusions from PhyloG2P analysis.
Another important caveat for the application of PhyloG2P methods based on replicated amino acid substitutions or changes in evolutionary rate is that they depend on alignments of orthologous genetic elements. If replicated phenotypic changes have been caused by substitutions in lineage-specific paralogues, then only searching for genetic changes in orthologous genes will likely fail to detect the genetic replication. Methods based on gene duplication may be able to work in these scenarios, but in many cases the rate of duplication may be too low to be identified as significant. Paralogues could potentially be included in PhyloG2P analysis through adjustments of the mathematical frameworks used, similar to recent advances in gene-to-gene mapping methods (Forsythe et al. 2025).
More generally, identifying and aligning orthologous genetic elements becomes challenging when studying genetic variation across deep evolutionary time scales (Steenwyk et al. 2023). It may be unlikely that replicated phenotypic changes would be caused by changes in the same genetic element if it has diverged to the extent that it is unalignable between the lineages. However, even after such divergence, there may still be conserved functions on higher levels, such as the functional role of protein domains, and their involvement in replicated phenotypic changes would not be detectable using current PhyloG2P approaches. This challenge creates an upper limit for the application of presently available PhyloG2P methods. Even so, there is still a great abundance of unexplored cases of replicated evolution within more closely related lineages that can be investigated through PhyloG2P approaches (aided by the growing availability of alignments Christmas et al. 2023).
Future Directions
Phenotypic replication can be caused by replicated changes at several genetic levels in both coding and noncoding genetic elements. As we have discussed, there are a range of PhyloG2P methods available to identify replication on these different genetic levels. The diversity of methods available means that there is no one single approach to determine the replicated genetic changes underlying phenotypic replication. This diversity has been actively embraced with several recent studies applying multiple complementary PhyloG2P methods (Eliason et al. 2023; Zhang et al. 2023; Eastment et al. 2024; Matsuda and Makino 2024) (Box 2). However, there have been other recent studies that only use a single PhyloG2P method (Xu et al. 2021; Sakamoto et al. 2024). Such studies limit the scope of the conclusions they can draw, and risk missing important genetic elements associated with the trait (Wheeler et al. 2022). Similarly, PhyloG2P studies have revealed the presence of genetic replication in noncoding regions when it was not found in coding regions (Langer et al. 2018; Feigin et al. 2019; Thomas et al. 2024; Shakya et al. 2025). As such, we strongly encourage future PhyloG2P studies to use a representative range of methods and analyze both coding and noncoding genetic elements whenever possible, while providing clear justifications for their methodological choices.
Box 2 Exemplar PhyloG2P Applications.
Viviparity (giving birth to live young) has independently evolved many times in vertebrates such as squamate reptiles and fish, as well as once in mammals. Eastment et al. (2024) utilized several PhyloG2P methods to identify the genetic changes associated with evolutionary transitions to viviparity. Few previous studies had used genome-wide approaches to explore whether the same genes were involved in independent transitions to viviparity across deep evolutionary time scales. Eastment et al. constructed a set of gene alignments from vertebrate genomes spanning over 400 million years of evolution that included 17 independent transitions to viviparity, a far larger number than earlier PhyloG2P studies (Yusuf et al. 2023). They found that replicated expansion and contraction of gene family size was the most widespread form of replication, although no family changed significantly in more than six lineages. Notably, they identified a contraction of a gene family related to yolk development in a subset of viviparous species. While yolk is an essential component of eggs in oviparous species, it is also present within the uterus during pregnancy for most viviparous species. However, the subset of lineage that Eastment et al. (2024) detected the replicate gene family contraction in have all previously been identified as no longer producing yolk (including mammals). This finding highlights the potential for gene family contraction associated with trait loss, as well as the importance of having a detailed understanding of trait variation in PhyloG2P studies. In addition, they identified diverse molecular changes associated with viviparity, but only some of these genetic changes were replicated, and even then only across a subset of lineages. Together, the findings of Yusuf et al. (2023) and Eastment et al. (2024) suggests that complex traits evolving over large timescales show limited genetic repeatability, likely reflecting redundancy in the genetic routes that can lead to viviparity.
Many lineages of angiosperm plants have independently adapted to harsh alpine environments. Zhang et al. (2023) used PhyloG2P methods from all three categories (replicated amino acid substitutions, changes in evolutionary rates, and duplication/loss of genetic elements) to identify genetic replication across seven independent alpine transitions in the angiosperms. Through analysis of gene duplication and loss, they identified expansions of gene families related to hypoxia tolerance, which are likely required to survive in low oxygen alpine environments. Additionally, they found plant immune response gene families had repeatedly contracted, potentially as a result of reduced microorganism presence. Evolutionary rates analysis identified replicated changes in genes across different biological pathways. Genes related to basic biological process were found to be under positive selection, which the authors hypothesized was related to the core physiological adaptations required to survive in harsh conditions. They also identified the accelerated evolution of genes relating to self-incompatibility (the inability to self fertilize) using RERconverge. Follow up investigations using RELAX and PCOC revealed the presence of both relaxed selection and a small number of replicated amino acid profile changes, likely due to the low abundance of pollinators. Zhang et al. (2023)’s study demonstrates how the different categories of PhyloG2P methods can identify distinct sets of genetic elements related to changes in component traits involved in complex phenotypic adaptations.
PhyloG2P methods can also be used to compliment other molecular and genetics approaches. Okuyama et al. (2025) used a range of transcriptomic, molecular, and transgenic approaches to understand the genetic mechanisms underlying replicated changes in floral scent compounds. After identifying the responsible gene, they used CSUBST (Fukushima and Pollock 2023) to identify replicated amino acid substitutions in the foreground lineages. In contrast to Zhang et al. (2023), who used PCOC due to substantial genetic divergence between the alpine plant species, Okuyama et al. (2025) chose CSUBST because the gene was highly conserved between species and they wanted to identify any identical substitutions driving the replicated phenotypic change. Their analysis identified three substitutions within the gene that had occurred in all three independent transitions studied. Their follow up experiments manipulating those three sites confirmed they were necessary and sufficient to cause the phenotypic change. This study highlights how PhyloG2P methods can compliment other molecular genetics methods.
There are potential barriers that can prevent the widespread application of PhyloG2P methods. One is the accessibility of PhyloG2P methods from a technical perspective. Methods which lack neat software implementation or basic documentation may prove inaccessible for researchers without the required background knowledge and skills, limiting their potential usage. To improve accessibility, researchers creating PhyloG2P methods should include detailed walk-throughs and user-guides (such as Kowalczyk et al. 2019 and Thomas et al. 2024), and where possible make easy to use software implementations (such as PCOC being available as a container (Rey et al. 2018) or RERconverge being available as an R package Kowalczyk et al. 2019). Another barrier is computational intensity. While high computational demands are common in genomic analysis, they can limit the capacity for researchers to analyze genomic data. This limitation not only restricts researchers’ ability to apply PhyloG2P methods, it also inhibits the capacity for reproduction and verification of previous findings, eroding an important component of rigorous scientific research (Kumar 2022). Research groups developing PhyloG2P methods should explore ways to reduce the computational intensity of their methods (e.g. Yan et al. 2023 found that tuning model complexity on a gene by gene basis reduced computation time while no impacting statistical performance). Computational accessibility can also be provided through publicly available web applications (Ribeiro et al. 2023; Dong et al. 2024; Morel et al. 2024). Ensuring that PhyloG2P methods are accessible will enhance the capacity for researchers to apply comprehensive sets of methods in their analyses, improving the capacity for PhyloG2P research to uncover novel genotype to phenotype maps.
The PhyloG2P methods discussed in this review either focus on changes in molecular sequences, or gene duplication and loss. However, there are many other epigenetic and transcriptomic features of genomes which can impact phenotypes. These features include DNA methylation, chromatin structure, structural rearrangements (Vargas-Chavez et al. 2025), and the interactions between RNA, proteins, and their environment. Including variation in these features can improve phenotype to genotype maps (Ritchie et al. 2015), and incorporating them into PhyloG2P analysis may prove fruitful. One approach to including these other forms of molecular variation is to use machine learning models, such as the Evolutionary Sparse Learning (ESL) model of Kumar and Sharma (2021). ESL uses binary encoding to enable joint analysis of any genetic feature, such as SNPs or presence/absence of a genetic element, as well as epigenetic or environmental variables. Machine learning processes are then used to narrow down the features to those which are most important in explaining a particular phylogenetic hypothesis, such as the presence of a particular phenotype. As well as its capacity to integrate a variety of variables relevant to phenotypic prediction, ESL is also orders of magnitude less computationally intensive than the PhyloG2P models discussed in this review. However, one shortcoming of using machine learning is that it does not explicitly model evolutionary processes, complicating the interpretation of the features it identifies. As a result, it may be most appropriate to combine applications of ESL with other methods that use explicit models of evolution. This approach may prove beneficial, as the more computationally intensive models will only need to be applied to the subset of genetic elements identified by the computationally efficient ESL (provided ESL is not overly inaccurate in its initial identification of candidate genetic elements) (Allard et al. 2025). As such, the machine learning approach of ESL provides a first step into modeling a broader range of biological features that may impact phenotypes, and also provides a more computationally accessible option for PhyloG2P adjacent analysis.
PhyloG2P methods identify a set of genetic elements correlated with a trait of interest, and follow up studies are required to determine the nature of that relationship. Although infeasible for many species, most PhyloG2P studies use existing genetic annotations to verify the relevance of the candidate genetic elements they identify. This verification can be done through gene ontology enrichment tests, which have often identified significant enrichment in functions related to the trait of interest (Chikina et al. 2016; Prudent et al. 2016; Nagy et al. 2017; Hu et al. 2019; Zhang et al. 2023; Redlich et al. 2024). However, relatively few applications of PhyloG2P methods have been followed up by transgenic or knockout experiments to verify the role of genetic elements identified in their analysis (Booker et al. 2016; Sackton et al. 2019; Xu et al. 2021; Okuyama et al. 2025). These experiments are important, as they provide evidence of causal relationships between a genetic element and a trait (Losos 2011). More studies such as these will be required in future to build confidence in the capacity for PhyloG2P methods to produce novel genotype to phenotype maps. Although most current applications of PhyloG2P methods focus on deepening our understanding of previously established genotype-phenotype maps, developing novel genotype-phenotype relationships should remain at the forefront of PhyloG2P research.
An important limitation of current PhyloG2P methods is that they rely on a single “representative” trait value and consensus genome for each lineage, which does not account for the phenotypic and genotypic variation within lineages. Trait values can vary widely within lineages, and this variance will likely not be consistent through time or between lineages. These forms of variance are particularly important for the analysis of traits with continuous values, as allelic variation between individuals may contribute to significantly different values of the trait (such as height or lifespan). Extending PhyloG2P methods to account for within-lineage variance, such as by modeling changes in trait distributions, may improve their accuracy. It is also likely that allelic variation within populations that have undergone replicated phenotypic evolution will show corresponding signs of adaptation in the genetic elements controlling the trait (Barghi et al. 2020). Finding ways to include this information in PhyloG2P analysis may improve the capacity to identify genetic elements associated with trait variation between species. More broadly, a model that unifies both population genetics and PhyloG2P analysis will pave the way for understanding the connection between evolution on short and long timescales. This gap has been partially bridged by the PhyloGWAS approach of (Pease et al. 2016), and may be more comprehensively closed with a unified mathematical framework such as the one proposed by Schraiber et al. (2024).
In summary, recent years have seen many developments in the PhyloG2P research program. With the growing abundance of genomic sequences, well-resolved phylogenies, and detailed trait databases, PhyloG2P methods are well positioned to expand genotype to phenotype mapping to include variation between species. A detailed understanding of both phenotypic and genotypic replication will reveal a largely untapped wealth of potential applications of these methods, further enabled by the development of PhyloG2P methods that can be applied to continuously valued traits. While some challenges remain for the integration of epigenetic information and within population trait variation in these studies, recent developments provide promising approaches to addressing them. There are many potential next steps that the PhyloG2P research program could take, and all of them are likely to reveal exciting new opportunities for expanding the mapping of genotype to phenotype.
Acknowledgments
We thank the anonymous reviewers for their helpful and constructive feedback.
Contributor Information
Arlie R Macdonald, School of Natural Sciences, University of Tasmania, Hobart, TS, Australia; Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, University of Tasmania, Sandy Bay, TS, Australia; School of the Environment, The University of Queensland, Brisbane, QL, Australia.
Maddie E James, Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, University of Tasmania, Sandy Bay, TS, Australia; School of the Environment, The University of Queensland, Brisbane, QL, Australia.
Jonathan D Mitchell, School of Natural Sciences, University of Tasmania, Hobart, TS, Australia; Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, University of Tasmania, Sandy Bay, TS, Australia.
Barbara R Holland, School of Natural Sciences, University of Tasmania, Hobart, TS, Australia; Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture, University of Tasmania, Sandy Bay, TS, Australia.
Funding
This work was partly funded by The Australian Research Council Centre of Excellence for Plant Success in Nature and Agriculture (CE200100015).
Data Availability
No new data were generated or analyzed in support of this research.
Literature Cited
- Allard JB et al. Evolutionary sparse learning reveals the shared genetic basis of convergent traits. Nat Commun. 2025:16:3217. 10.1038/s41467-025-58428-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Arendt J, Reznick D. Convergence and parallelism reconsidered: what have we learned about the genetics of adaptation? Trends Ecol Evol. 2008:23:26–32. 10.1016/j.tree.2007.09.011. [DOI] [PubMed] [Google Scholar]
- Baker J, Meade A, Venditti C. Genes underlying the evolution of tetrapod testes size. BMC Biol. 2021:19:1–10. 10.1186/s12915-021-01107-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barghi N, Hermisson J, Schlötterer C. Polygenic adaptation: a unifying framework to understand positive selection. Nat Rev Genet. 2020:21:769–781. 10.1038/s41576-020-0250-z. [DOI] [PubMed] [Google Scholar]
- Barteri F et al. Caastools: a toolbox to identify and test convergent amino acid substitutions. Bioinformatics. 2023:39:btad623. 10.1093/bioinformatics/btad623. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bastide P, Solís-Lemus C, Kriebel R, William Sparks K, Ane C. Phylogenetic comparative methods on phylogenetic networks with reticulations. Syst Biol. 2018:67:800–820. 10.1093/sysbio/syy033. [DOI] [PubMed] [Google Scholar]
- Booker BM et al. Bat accelerated regions identify a bat forelimb specific enhancer in the HoxD locus. PLoS Genet. 2016:12:e1005738. 10.1371/journal.pgen.1005738. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Castoe TA et al. Evidence for an ancient adaptive episode of convergent molecular evolution. Proc Natl Acad Sci U S A. 2009:106:8986–8991. 10.1073/pnas.0900233106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cerca J. Understanding natural selection and similarity: convergent, parallel and repeated evolution. Mol Ecol. 2023:32:5451–5462. 10.1111/mec.v32.20. [DOI] [PubMed] [Google Scholar]
- Chabrol O, Royer-Carenzi M, Pontarotti P, Didier G. Detecting the molecular basis of phenotypic convergence. Methods Ecol Evol. 2018:9:2170–2180. 10.1111/mee3.2018.9.issue-11. [DOI] [Google Scholar]
- Chikina M, Robinson JD, Clark NL. Hundreds of genes experienced convergent shifts in selective pressure in marine mammals. Mol Biol Evol. 2016:33:2182–2192. 10.1093/molbev/msw112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Christmas MJ et al. Evolutionary constraint and innovation across hundreds of placental mammals. Science. 2023:380:eabn3943. 10.1126/science.abn3943. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Corbett-Detig RB, Russell SL, Nielsen R, Losos J. Phenotypic convergence is not mirrored at the protein level in a lizard adaptive radiation. Mol Biol Evol. 2020:37:1604–1614. 10.1093/molbev/msaa028. [DOI] [PubMed] [Google Scholar]
- Dembech E et al. Identification of hidden associations among eukaryotic genes through statistical analysis of coevolutionary transitions. Proc Natl Acad Sci U S A. 2023:120:e2218329120. 10.1073/pnas.2218329120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dong Z, Wang C, Qu Q. WGCCRR: a web-based tool for genome-wide screening of convergent indels and substitutions of amino-acids. Bioinform Adv. 2024:4:vbae070. 10.1093/bioadv/vbae070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Duchemin L, Lanore V, Veber P, Boussau B. Evaluation of methods to detect shifts in directional selection at the genome scale. Mol Biol Evol. 2022:40:msac247. 10.1093/molbev/msac247. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eastment RV, Wong BBM, McGee MD. Convergent genomic signatures associated with vertebrate viviparity. BMC Biol. 2024:22:34. 10.1186/s12915-024-01837-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eliason CM et al. Genomic signatures of convergent shifts to plunge-diving behavior in birds. Commun Biol. 2023:6:1011. 10.1038/s42003-023-05359-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Falster D et al. AusTraits, a curated plant trait database for the Australian flora. Sci Data. 2021:8:254. 10.1038/s41597-021-01006-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Feigin CY, Newton AH, Pask AJ. Widespread cis-regulatory convergence between the extinct Tasmanian tiger and gray wolf. Genome Res. 2019:29:1648–1658. 10.1101/gr.244251.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Forsythe ES et al. Ercnet: phylogenomic prediction of interaction networks in the presence of gene duplication. Mol Biol Evol. 2025:42:msaf089. 10.1093/molbev/msaf089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fukushima K, Pollock DD. Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence. Nat Ecol Evol. 2023:7:155–170. 10.1038/s41559-022-01932-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gemmell P, Sackton TB, Edwards SV, Liu JS. A phylogenetic method linking nucleotide substitution rates to rates of continuous trait evolution. PLoS Comput Biol. 2024:20:e1011995. 10.1371/journal.pcbi.1011995. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griesmann M et al. Phylogenomics reveals multiple losses of nitrogen-fixing root nodule symbiosis. Science. 2018:361:eaat1743. 10.1126/science.aat1743. [DOI] [PubMed] [Google Scholar]
- Grossnickle DM et al. Challenges and advances in measuring phenotypic convergence. Evolution. 2024:78:1355–1371. 10.1093/evolut/qpae081. [DOI] [PubMed] [Google Scholar]
- Halabi K, Karin EL, Guéguen L, Mayrose I. A codon model for associating phenotypic traits with altered selective patterns of sequence evolution. Syst Biol. 2021:70:608–622. 10.1093/sysbio/syaa087. [DOI] [PubMed] [Google Scholar]
- Hiller M et al. A “forward genomics” approach links genotype to phenotype using independent phenotypic losses among related species. Cell Rep. 2012:2:817–823. 10.1016/j.celrep.2012.08.032. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Holland BR, Ketelaar-Jones S, o’Mara AR, Woodhams MD, Jordan GJ. Accuracy of ancestral state reconstruction for non-neutral traits. Sci Rep. 2020:10:7644. 10.1038/s41598-020-64647-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hu Z, Sackton TB, Edwards SV, Liu JS. Bayesian detection of convergent rate changes of conserved noncoding elements on phylogenetic trees. Mol Biol Evol. 2019:36:1086–1100. 10.1093/molbev/msz049. [DOI] [PMC free article] [PubMed] [Google Scholar]
- James ME, Brodribb T, Wright IJ, Rieseberg LH, Ortiz-Barrientos D. Replicated evolution in plants. Annu Rev Plant Biol. 2023:74:697–725. 10.1146/arplant.2023.74.issue-1. [DOI] [PubMed] [Google Scholar]
- Klure DM, Greenhalgh R, Orr TJ, Shapiro MD, Dearing MD. Parallel gene expansions drive rapid dietary adaptation in herbivorous woodrats. Science. 2025:387:156–162. 10.1126/science.adp7978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kowalczyk A, Chikina M, Clark N. Complementary evolution of coding and noncoding sequence underlies mammalian hairlessness. Elife. 2022:11:e76911. 10.7554/eLife.76911. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kowalczyk A et al. RERconverge: an R package for associating evolutionary rates with convergent traits. Bioinformatics. 2019:35:4815–4817. 10.1093/bioinformatics/btz468. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kowalczyk A, Partha R, Clark NL, Chikina M. Pan-mammalian analysis of molecular constraints underlying extended lifespan. Elife. 2020:9:e51089. 10.7554/eLife.51089. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar S. Embracing green computing in molecular phylogenetics. Mol Biol Evol. 2022:39:msac043. 10.1093/molbev/msac043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar S, Sharma S. Evolutionary sparse learning for phylogenomics. Mol Biol Evol. 2021:38:4674–4682. 10.1093/molbev/msab227. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lamichhaney S et al. Integrating natural history collections and comparative genomics to study the genetic architecture of convergent evolution. Philos Trans R Soc Lond B Biol Sci. 2019:374:20180248. 10.1098/rstb.2018.0248. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Langer BE, Roscito JG, Hiller M. REforge associates transcription factor binding site divergence in regulatory elements with phenotypic differences between species. Mol Biol Evol. 2018:35:3027–3040. 10.1093/molbev/msy187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lartillot N, Poujol R. A phylogenetic model for investigating correlated evolution of substitution rates and continuous phenotypic characters. Mol Biol Evol. 2011:28:729–744. 10.1093/molbev/msq244. [DOI] [PubMed] [Google Scholar]
- Láruson ÁJ, Yeaman S, Lotterhos KE. The importance of genetic redundancy in evolution. Trends Ecol Evol. 2020:35:809–822. 10.1016/j.tree.2020.04.009. [DOI] [PubMed] [Google Scholar]
- Levy Karin E, Wicke S, Pupko T, Mayrose I. An integrated model of phenotypic trait changes and site-specific sequence evolution. Syst Biol. 2017:66:917–933. 10.1093/sysbio/syx032. [DOI] [PubMed] [Google Scholar]
- Losos JB. Convergence, adaptation, and constraint. Evolution. 2011:65:1827–1840. 10.1111/evo.2011.65.issue-7. [DOI] [PubMed] [Google Scholar]
- Maddison WP, FitzJohn RG. The unsolved challenge to phylogenetic correlation tests for categorical characters. Syst Biol. 2014:64:127–136. 10.1093/sysbio/syu070. [DOI] [PubMed] [Google Scholar]
- Marcovitz A, Jia R, Bejerano G. “Reverse genomics”’ predicts function of human conserved noncoding elements. Mol Biol Evol. 2016:33:1358–1369. 10.1093/molbev/msw001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Matsuda Y, Makino T. Comparative genomics reveals convergent signals associated with the high metabolism and longevity in birds and bats. Proc Biol Sci. 2024:291:20241068. 10.1098/rspb.2024.1068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mendes FK, Hahn MW. Gene tree discordance causes apparent substitution rate variation. Syst Biol. 2016:65:711–721. 10.1093/sysbio/syw018. [DOI] [PubMed] [Google Scholar]
- Mendes FK, Livera AP, Hahn MW. The perils of intralocus recombination for inferences of molecular convergence. Philos Trans R Soc Lond B Biol Sci. 2019:374:20180244. 10.1098/rstb.2018.0244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mendes FK, Vanderpool D, Fulton B, Hahn MW. CAFE 5 models variation in evolutionary rates among gene families. Bioinformatics. 2020:36:5516–5518. 10.1093/bioinformatics/btaa1022. [DOI] [PubMed] [Google Scholar]
- Mohammadi S et al. Epistatic effects between amino acid insertions and substitutions mediate toxin resistance of vertebrate Na+, K+-ATPases. Mol Biol Evol. 2022:39:msac258. 10.1093/molbev/msac258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morales AE et al. Distinct genes with similar functions underlie convergent evolution in Myotis bat ecomorphs. Mol Biol Evol. 2024:41:msae165. 10.1093/molbev/msae165. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morel M, Zhukova A, Lemoine F, Gascuel O. Accurate detection of convergent mutations in large protein alignments with ConDor. Genome Biol Evol. 2024:16:evae040. 10.1093/gbe/evae040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nagy LG et al. Latent homology and convergent regulatory evolution underlies the repeated emergence of yeasts. Nat Commun. 2014:5:4471. 10.1038/ncomms5471. [DOI] [PubMed] [Google Scholar]
- Nagy LG et al. Genetic bases of fungal white rot wood decay predicted by phylogenomic analysis of correlated gene-phenotype evolution. Mol Biol Evol. 2017:34:35–44. 10.1093/molbev/msw238. [DOI] [PubMed] [Google Scholar]
- Nagy LG, Merényi Z, Hegedüs B, Bálint B. Novel phylogenetic methods are needed for understanding gene function in the era of mega-scale genome sequencing. Nucleic Acids Res. 2020:48:2209–2219. 10.1093/nar/gkz1241. [DOI] [PMC free article] [PubMed] [Google Scholar]
- O’Connor TD, Mundy NI. Genotype–phenotype associations: substitution models to detect evolutionary associations between phenotypic variables and genotypic evolutionary rate. Bioinformatics. 2009:25:i94–i100. 10.1093/bioinformatics/btp231. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Okuyama Y et al. Convergent acquisition of disulfide-forming enzymes in malodorous flowers. Science. 2025:388:656–661. 10.1126/science.adu8988. [DOI] [PubMed] [Google Scholar]
- Parker J et al. Genome-wide signatures of convergent evolution in echolocating mammals. Nature. 2013:502:228–231. 10.1038/nature12511. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Partha R, Kowalczyk A, Clark NL, Chikina M. Robust method for detecting convergent shifts in evolutionary rates. Mol Biol Evol. 2019:36:1817–1830. 10.1093/molbev/msz107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pease JB, Haak DC, Hahn MW, Moyle LC. Phylogenomics reveals three sources of adaptive variation during a rapid radiation. PLoS Biol. 2016:14:e1002379. 10.1371/journal.pbio.1002379. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pereira AG, Kohlsdorf T. Repeated evolution of similar phenotypes: integrating comparative methods with developmental pathways. Genet Mol Biol. 2023:46:e20220384. 10.1590/1678-4685-gmb-2022-0384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A. Detection of nonneutral substitution rates on mammalian phylogenies. Genome Res. 2010:20:110–121. 10.1101/gr.097857.109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prudent X, Parra G, Schwede P, Roscito JG, Hiller M. Controlling for phylogenetic relatedness and evolutionary rates improves the discovery of associations between species’ phenotypic and genomic differences. Mol Biol Evol. 2016:33:2135–2150. 10.1093/molbev/msw098. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Redlich R et al. RERconverge expansion: using relative evolutionary rates to study complex categorical trait evolution. Mol Biol Evol. 2024:41:msae210. 10.1093/molbev/msae210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rey C, Guéguen L, Sémon M, Boussau B. Accurate detection of convergent amino-acid evolution with PCOC. Mol Biol Evol. 2018:35:2296–2306. 10.1093/molbev/msy114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rey C et al. Detecting adaptive convergent amino acid evolution. Philos Trans R Soc Lond B Biol Sci. 2019:374:20180234. 10.1098/rstb.2018.0234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ribeiro D, Borges R, Rocha AP, Antunes A. Testing phylogenetic signal with categorical traits and tree uncertainty. Bioinformatics. 2023:39:btad433. 10.1093/bioinformatics/btad433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ritchie MD, Holzinger ER, Li R, Pendergrass SA, Kim D. Methods of integrating data to uncover genotype–phenotype interactions. Nat Rev Genet. 2015:16:85–97. 10.1038/nrg3868. [DOI] [PubMed] [Google Scholar]
- Sackton TB et al. Convergent regulatory evolution and loss of flight in paleognathous birds. Science. 2019:364:74–78. 10.1126/science.aat7244. [DOI] [PubMed] [Google Scholar]
- Sadanandan KR et al. Convergence in hearing-related genes between echolocating birds and mammals. Proc Natl Acad Sci U S A. 2023:120:e2307340120. 10.1073/pnas.2307340120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sakamoto F et al. Detection of evolutionary conserved and accelerated genomic regions related to adaptation to thermal niches in Anolis lizards. Ecol Evol. 2024:14:e11117. 10.1002/ece3.v14.3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saputra E, Kowalczyk A, Cusick L, Clark N, Chikina M. Phylogenetic permulations: a statistically rigorous approach to measure confidence in associations in a phylogenetic context. Mol Biol Evol. 2021:38:3004–3021. 10.1093/molbev/msab068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schraiber JG, Edge MD, Pennell M. Unifying approaches from statistical genetics and phylogenetics for mapping phenotypes in structured populations. PLoS Biol. 2024:22:e3002847. 10.1371/journal.pbio.3002847. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shakya SB, Edwards SV, Sackton TB. Convergent evolution of noncoding elements associated with short tarsus length in birds. BMC Biol. 2025:23:52. 10.1186/s12915-025-02156-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith MD et al. Less is more: an adaptive branch-site random effects model for efficient detection of episodic diversifying selection. Mol Biol Evol. 2015:32:1342–1353. 10.1093/molbev/msv022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith SD, Pennell MW, Dunn CW, Edwards SV. Phylogenetics is the new genetics (for most of biodiversity). Trends Ecol Evol. 2020:35:415–425. 10.1016/j.tree.2020.01.005. [DOI] [PubMed] [Google Scholar]
- Speed MP, Arbuckle K. Quantification provides a conceptual basis for convergent evolution. Biol Rev Camb Philos Soc. 2017:92:815–829. 10.1111/brv.2017.92.issue-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Steenwyk JL, Li Y, Zhou X, Shen XX, Rokas A. Incongruence in the phylogenomics era. Nat Rev Genet. 2023:24:834–850. 10.1038/s41576-023-00620-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stefen C et al. Phenotyping in the era of genomics: MaTrics—a digital character matrix to document mammalian phenotypic traits. Mamm Biol. 2022:102:235–249. 10.1007/s42991-021-00192-5. [DOI] [Google Scholar]
- Stern DL. The genetic causes of convergent evolution. Nat Rev Genet. 2013:14:751–764. 10.1038/nrg3483. [DOI] [PubMed] [Google Scholar]
- Sykes J, Holland BR, Charleston MA. A review of visualisations of protein fold networks and their relationship with sequence and function. Biol Rev Camb Philos Soc. 2022:98:243–262. 10.1111/brv.v98.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tejada-Martinez D et al. Positive selection and enhancer evolution shaped lifespan and body mass in great apes. Mol Biol Evol. 2022:39:msab369. 10.1093/molbev/msab369. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thomas GW et al. Practical guidance and workflows for identifying fast evolving non-coding genomic elements using PhyloAcc. Integr Comp Biol. 2024:64:1513–1525. 10.1093/icb/icae056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thomas GW, Hahn MW. Determining the null model for detecting adaptive convergence from genomic data: a case study using echolocating mammals. Mol Biol Evol. 2015:32:1232–1236. 10.1093/molbev/msv013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thomas GW, Hahn MW, Hahn Y. The effects of increasing the number of taxa on inferences of molecular convergence. Genome Biol Evol. 2017:9:213–221. 10.1093/gbe/evw306. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Treaster S, Daane JM, Harris MP. Refining convergent rate analysis with topology in mammalian longevity and marine transitions. Mol Biol Evol. 2021:38:5190–5203. 10.1093/molbev/msab226. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vargas-Chavez C et al. An episodic burst of massive genomic rearrangements and the origin of non-marine annelids. Nat Ecol Evol. 2025:9:1263–1279. 10.1038/s41559-025-02728-1. [DOI] [PubMed] [Google Scholar]
- Wang QP et al. Adaptive evolution of antioxidase-related genes in hypoxia-tolerant mammals. Front Genet. 2024:15:1315677. 10.3389/fgene.2024.1315677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wertheim JO, Murrell B, Smith MD, Kosakovsky Pond SL, Scheffler K. RELAX: detecting relaxed selection in a phylogenetic framework. Mol Biol Evol. 2015:32:820–832. 10.1093/molbev/msu400. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wheeler LC et al. Transcription factors evolve faster than their structural gene targets in the flavonoid pigment pathway. Mol Biol Evol. 2022:39:msac044. 10.1093/molbev/msac044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Womack MC, Fiero TS, Hoke KL. Trait independence primes convergent trait loss. Evolution. 2018:72:679–687. 10.1111/evo.2018.72.issue-3. [DOI] [PubMed] [Google Scholar]
- Womack MC, Hoke KL. Convergent anuran middle ear loss lacks a universal, adaptive explanation. Brain Behav Evol. 2023:98:290–301. 10.1159/000534936. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu D et al. A single mutation underlying phenotypic convergence for hypoxia adaptation on the Qinghai-Tibetan Plateau. Cell Res. 2021:31:1032–1035. 10.1038/s41422-021-00517-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yan H et al. PhyloAcc-GT: a Bayesian method for inferring patterns of substitution rate shifts on targeted lineages accounting for gene tree discordance. Mol Biol Evol. 2023:40:msad195. 10.1093/molbev/msad195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang Y, Jiang X. Evolink: a phylogenetic approach for rapid identification of genotype–phenotype associations in large-scale microbial multispecies data. Bioinformatics. 2023:39:btad215. 10.1093/bioinformatics/btad215. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang Z. Likelihood ratio tests for detecting positive selection and application to primate lysozyme evolution. Mol Biol Evol. 1998:15:568–573. 10.1093/oxfordjournals.molbev.a025957. [DOI] [PubMed] [Google Scholar]
- Yang Z. PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 2007:24:1586–1591. 10.1093/molbev/msm088. [DOI] [PubMed] [Google Scholar]
- Yusuf LH, Saldívar Lemus Y, Thorpe P, Macias Garcia C, Ritchie MG. Genomic signatures associated with transitions to viviparity in cyprinodontiformes. Mol Biol Evol. 2023:40:msad208. 10.1093/molbev/msad208. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang J. Patterns and evolutionary consequences of pleiotropy. Annu Rev Ecol Evol Syst. 2023:54:1–19. 10.1146/ecolsys.2023.54.issue-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang J, Kumar S. Detection of convergent and parallel evolution at the amino acid sequence level. Mol Biol Evol. 1997:14:527–536. 10.1093/oxfordjournals.molbev.a025789. [DOI] [PubMed] [Google Scholar]
- Zhang J, Nielsen R, Yang Z. Evaluation of an improved branch-site likelihood method for detecting positive selection at the molecular level. Mol Biol Evol. 2005:22:2472–2479. 10.1093/molbev/msi237. [DOI] [PubMed] [Google Scholar]
- Zhang X et al. Genomic convergence underlying high-altitude adaptation in alpine plants. J Integr Plant Biol. 2023:65:1620–1635. 10.1111/jipb.v65.7. [DOI] [PubMed] [Google Scholar]
- Zhen Y, Aardema ML, Medina EM, Schumer M, Andolfatto P. Parallel molecular evolution in an herbivore community. Science. 2012:337:1634–1637. 10.1126/science.1226630. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou X, Seim I, Gladyshev VN. Convergent evolution of marine mammals is associated with distinct substitutions in common genes. Sci Rep. 2015:5:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zou Z, Zhang J. Are convergent and parallel amino acid substitutions in protein evolution more prevalent than neutral expectations? Mol Biol Evol. 2015:32:2085–2096. 10.1093/molbev/msv091. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No new data were generated or analyzed in support of this research.


