Abstract
In the tree reconciliation approach for species tree inference, a tree that has the minimum reconciliation score for given gene trees is taken as an estimate of the species tree. The scoring models used in existing tree reconciliation methods include the duplication, mutation, and deep coalescence costs. Since existing inference methods all are heuristic, their performances are often evaluated by using the Robinson-Foulds (RF) distance between the true species trees and the estimates output on simulated multi-locus datasets. To better understand these methods, we study the relationships between the duplication cost and the RF distance. We prove that the gap between the duplication cost and the RF distance is unbounded, but the symmetric duplication cost is logarithmically equivalent to the RF distance. The relationships between other reconciliation costs and the RF distance are also investigated.
Key words: : gene duplication and loss, incomplete lineage sorting, metric equivalence, Ronbinson-Foulds distance, species tree inference, tree reconciliation
1. Introduction
With more and more genomes being fully sequenced, species tree inferences from multilocus datasets have become one of the basic bioinformatics tasks in comparative and evolutionary biology (Dunn et al., 2008). Reconstructing a species tree from a multi locus dataset requires reconciling the phylogenetic information contained in the genes. The challenge for such an analysis is that the genes sampled from the same set of species often contain conflicting signals and thus have discordant gene trees. Although some of these conflicting signals can be due to systematic errors occurring in phylogenetic analysis, they often reflect gene-specific mutational events. During evolution, recombination, gene duplication, gene loss, incomplete lineage sorting, or horizontal gene transfer, all might occur in a gene family (Fitch, 1970; Goodman et al., 1979; Maddison, 1997; Pamilo and Nei, 1988).
Existing methods for inferring species trees from gene trees can be divided into two broad categories: parametric methods (Ané et al., 2007; Kubatko et al., 2009; Liu and Pearl, 2007) and tree reconciliation methods (Warnow, 2013). Parametric methods adopt gene duplication or incomplete lineage-sorting models and infer species trees using either likelihood or Bayesian framework. These approaches have strong statistical foundation, but they can be computationally expensive. Therefore, tree reconciliation methods are often used in phylogenetic studies on a genomic scale (Burleigh et al., 2011; Katz et al., 2012; Sanderson and McMahon, 2007).
Gene tree reconciliation methods take a collection of gene trees and output the tree that has the smallest reconciliation score for the gene trees as an estimate of the species tree. The scoring models used in these methods include the duplication cost (that is, the number of duplications), the mutation cost (that is, the number of gene duplication and loss events), the deep coalescence cost (that is, the number of incomplete lineage sorting events), and the Robinson-Foulds (RF) distance (Bansal et al., 2010; Bayzid et al., 2012; Chang et al., 2013; Chaudhary et al., 2013; Than and Nakhleh, 2009; Wehe et al., 2008; Yu et al., 2011). Since finding the parsimony tree from gene trees is NP-hard for the reconciliation scoring models (Bryant, 1997; Ma et al., 2000; Zhang, 2011), existing methods all compute a species tree by navigating through the part of the species tree space guided by the reconciliation score. The performances of these heuristic methods are usually evaluated using the RF distances between true species trees and their estimate output on simulation datasets (Chaudhary et al., 2013; Than and Nakhleh, 2009; Yang and Warnow, 2011). Hence, uncovering the relationships between the RF distance and the reconciliation costs is important to design good gene tree reconciliation methods. Unfortunately, many of these relationships are incompletely known. In particular, the relationship between the RF distance and the duplication cost is rather elusive, as both are bounded above by a linear function in the number of taxa. In this work, we mainly investigate this relationship. We consider all the reconciliation costs as different metrics in species tree spaces.
The rest of this article is divided into three sections. The basic concepts, notions, and the various reconciliation costs are introduced in Section 2. Several results are presented in Section 3. We first show that the gap between the duplication cost and the RF distance is unbounded. On the other hand, the symmetric duplication cost and the RF distance are proved to be logarithmically equivalent. A similar relationship between the gene loss cost and the RF distance is obtained. We also examine the distribution of the sizes of terraces for these metrics. We conclude with suggestions for future work in Section 4.
2. Concepts and Notions
2.1. Species trees
A species tree T over a set of species X is a rooted tree in which one node is designated as the root, all edges are oriented away from the root, and the species in X are one-to-one mapped to leaves. We use V(T) and E(T) to denote the sets of nodes and (directed) edges in T, respectively. For
, v is the parent of u and, equivalently, u is a child of v if
. A node is called a leaf if it does not have any child. In this article, we focus on binary species trees in which every nonleaf node has two children. We also use the following notation:
• p(u) denotes the parent of a nonroot node u in a tree;
• Ch(u) denotes the set of two children of a nonleaf u in a tree;
• Vlf(T) denotes the set of leaves in T;
• Vit(T) denotes the set of internal (i.e., nonleaf) nodes in T;
• T(u) denotes the subtree rooted at
, which consists of u and its descendants of u;• PT(−, u) denotes the unique path from the root to u in T;
• T|U denotes the subtrees induced by a subset
, whose sets of nodes and branches are
and
, respectively.
For
, v is an ancestor of u and, equivalently, u is a descendant of v, denoted by
, if v is in PT(−, u). We write
if the relationship is uncertain, but we know that v = u or
. Therefore, it is not hard to see that v is a leaf (terminal node) if and only if
for any
. For
, lca(U) denotes the latest common ancestor (lca) of the nodes in U.
2.2. The RF distance
Let X be a set of species, and let
denote the set of all species tree over X. Consider
. A species
is called the label of the corresponding leaf
, written l(u). Conversely, we define l−1(x) = u. For
, we define:
![]() |
called the (species) cluster of u. Define
![]() |
Consider two species trees S′ and S″ over X. They are identical if and only
. Therefore, the dissimilarity of S′ and S″ can be measured by the number of clusters that are found in one but not in the other, that is,
![]() |
called the RF distance (Semple and Steel, 2003). Since both S′ and S″ have |X| − 1 internal nodes and their roots' cluster is X, The following inequalities hold:
![]() |
2.3. Tree reconciliation costs
Consider
. Since each leaf in these trees has a unique label in X, there is a one-to-one map between Vlf(S′) and Vlf(S″). Such a correspondence can be extended into a map from V(S′) to V(S″), λS′,S″, as:
![]() |
The duplication cost (DPC), Cdup(S′, S″), is defined as |Dup(S′, S″)|, where
![]() |
The nodes in Dup(S′, S″) are called duplication nodes.
Proposition 1. For
.
Proof. Since S′ ≠ S″, there exists
having the following properties (Fig. 1):
FIG. 1.
Illustration of the proof of Proposition 1. The cluster of
is not found in S′, but the clusters of the descendants of u are all in S′. If
such that
, then lca(t1, t2) is a duplication node under λS′,S″.
(i) CS″(u) ≠ CS′(v) for any
;(ii) for each descendant t of u, that is,
, CS″(t) = CS′(v) for some
.
Let Ch(u) = {u1, u2}. By the above assumption, CS″(ui) = CS′(ti) for some
for i = 1, 2. Since CS″(u) is not a cluster in S′, t1 and t2 must not be siblings in S′, that is, p(t1) ≠ p(t2). Let lca(t1, t2) = v and Ch(v) = {v1, v2} such that
and
. Since t1 and t2 are not siblings, either t1 ≠ v1 or t2 ≠ v2. If t2 ≠ v2, λS′,S″(v2) is in PS″(−, u). If t1 ≠ v1, λS′,S″(v1) is in PS″(−, u). Thus, λS′,S″(v1), λS′,S″(v2), or both are in PS″(−, u), the path from r(T″) to u.
If
. If
. Therefore, v is a duplication node and hence Cdup(S′,S″) ≥ 1.
Let
have the largest distance from the root of S′. Then, x must have two leaf children, x1 and x2. By definition, λS′,S″(x1) ≠ λS′,S″(x) ≠ λS′,S″(t2). Hence, x is not a duplication node when S′ is mapped onto S″. Since S′ has |X| − 1 internal nodes, S′ has at most |X| − 1 − 1 duplication nodes. Therefore, Cdup(S′,S″) ≤ |X| − 2. ■
For
, we use λS′,S″(e′) to denote the path from λS′,S″(u) to λS′,S″(v). We further use |λS′,S″(u, v)| to denote the number of the non-end nodes in λS′,S″(e′). |λS′,S″(u, v)| = 0 if λS′,S″(u) = λS′,S″(v). The gene loss cost, Closs(S′, S″), is defined as:
![]() |
where χu = 1 if u is a duplication node such that |λS′,S″(u, û)| ≠ 0 for some
and χu = 0 otherwise.
If S′ is a gene tree and S″ is a species tree, Cdup(S′, S″) and Closs(S′, S″) are the numbers of gene duplications and losses occurring in the duplication history of the genes represented by the leaves in S′ defined by the lca map (Goodman et al., 1979).
For
, define:
![]() |
Clearly,
is in DC(e″) if and only if e″ is an edge in λS′,S″(e). The deep coalescence cost, Cdc(S′,S″), is defined as:
![]() |
The deep coalescence cost Cdc(S′, S″) is introduced to measure the dissimilarity of S′ and S″ that are caused by incomplete lineage sorting for a gene tree S′ and a species tree S″ (Maddison, 1997).
Finally, we note that all the reconciliation costs introduced above are not a symmetric metric in general. Hence, the symmetric duplication cost (SDPC) of S′ and S″, defined as Cdup(S′, S″) + Cdup(S″, S′), is used to study species tree inference (Ma et al., 2000).
3. Results
3.1. Unbounded gap between the DPC and RF distance
For two different trees over the same set of species, both the duplication cost and half of the RF distance are in the same range from 1 to n − 2, where n is the number of species involved in the trees. Therefore, the following question naturally arises:
Are the DPC and RF distance metrically equivalent?
and, equivalently, does there exist c1 ≥ c2 > 0 such that
![]() |
for any
?
Consider
such that Crf(S1, S2) = 2(|X| − 2), that is,
. By definition, the parent of two sibling leaves in S1 is not a duplication node under λS1, S2. Since such an internal node always exists in S1, we have:
![]() |
However, surprisingly, there are many pairs of trees S1 and S2 in
satisfying that Crf(S1, S2) = 2(|X| − 2) but Cdup(S1, S2) = 1, where |X| ≥ 3. For example, we consider a class of complete binary species trees, as illustrated in Figure 2. Let n = 2k for some k ≥ 1 and
. We set T to be the rooted complete binary tree with n unlabeled leaves. An denotes the species tree over X obtained from T by labeling the i-th leaf (counting from left to right) with i (Fig. 2). Bn denotes the species tree obtained from T by labeling the i-th leaf with φ(i) that is defined as:
![]() |
FIG. 2.
Complete binary species trees An and Bn for n = 23. Cdup(An, Bn) = 1 but Crf(An, Bn) = 2(n − 2).
for
.
Proposition 2. (1). Crf(An, Bn) = 2(n − 2).
(2). Cdup(An, Bn) = 1.
(3). The DPC and the RF distance are not equivalent.
Proof.
(1). For any nonroot
, either
, or
. In Bn, the (2i − 1)th and (2i)th leaves are siblings, labeled with i and n/2 + i, respectively, where i ≤ n/2. Therefore, for
,
![]() |
Combining these two facts together, we conclude that only the roots of An and Bn have the same cluster X. Hence, Crf(An, Bn) = 2(n − 2), as both trees have n − 1 internal nodes.
(2). Consider
. Notice that
for some i and k. For the left and right children, u1 and u2, of u,
![]() |
![]() |
If u is in the left subtree of An, then
and
![]() |
If u is in the right subtree of An,
![]() |
Hence, the root r of An and its two children are mapped to the root of Bn, implying that r is a duplication node.
If u is in the left subtree of An such that
, then,
![]() |
![]() |
![]() |
This implies that u is not a duplication node. Similarly, all the nodes in the right subtree are not duplication nodes. Therefore, Cdup(An, Bn) = 1.
(3). This statement follows directly from (1) and (2). ■
Although Cdup(An, Bn) = 1, we have that Cdup(Bn, An) = n/2 − 2. However, for any fixed large k > 0, there exist G and S such that
when X is large enough. For example, we take an integer m and define
![]() |
Consider the two trees G and S over X defined in Figure 3. It is easy to see that Crf(G, S) = 2(m2 − 2). The lca map λG,S maps different nodes in Gi to different nodes in S for each i. Furthermore, it maps the roots of
and their ancestors to the root of S. Thus, Cdup(G, S) = m − 1. Similarly, we can also show that Cdup(S, G) = m − 1. Hence, for any k < m, we have
.
FIG. 3.
Trees G and S over m2 species satisfying that Crf(G, S) = 2(m2 − 2), but Cdup(S, G) = Cdup(G, S) = m − 1.
This example shows that even the symmetric duplication cost (SDPC) is not equivalent to the RF distance in the species tree space. However, they are weakly equivalent.
3.2. Logarithmic equivalence of the SDPC and RF distance
Consider
. For the sake of simplicity, we use the following notations:
• μ denotes the lca map from S′ to S″, that is, μ = λS′,S″;
• ψ = λS″,S′;
• Dupμ(S′, S″) denotes the set of duplication nodes in S′ under μ.
Lemma 1.
For
such that Crf(S′, S″) = 2(|X| − 2),
and
.
Proof.
We prove the first fact by mathematical induction. By symmetry, the second fact holds.
Consider
. It has two children, u1 and u2. There are three possible cases.
Case 1.
u1 and u2 are leaves.
Since S′ and S″ do not have any common node cluster except X, the leaves, w1 and w2, that have the same labels as u1 and u2 are not siblings in S′. By definition, μ(wi) = ui and ψ(ui) = wi for i = 1, 2. Let v = lca(w1, w2). By assumption, ψ(u) = lca(ψ(u1), ψ(u2)) = lca(w1, w2) = v. Assume that the children of v are v1 and v2 such that
. Since w1 and w2 are not siblings, we have that v1 ≠ w1, v2 ≠ w2, or both. Without loss of generality, we may assume that v1 ≠ w1. Since μ(w1) = u1, which is a leaf, and μ(v1) ≠ u1, μ(v1) is either equal to u or an ancestor of u, that is, μ(v1) is in the path PS″(−, u). If v2 = w2, μ(v) = lca{μ(v1), μ(v2)} = μ(v1) (Fig. 4A), as μ(v2) ( = u2) is a child of u and hence a descendant of μ(v1). If v2 ≠ w2, μ(v2) is in PS″(−, u) and hence μ(v1) and μ(v2) are comparable under
, as they are in the path from the root to u (Fig. 4B). If
, μ(v) = μ(v2). If
, μ(v) = μ(v1). Therefore,
.
FIG. 4.
Illustration of the proof of Lemma 1 in Case 1. Here, u has two leaf children, u1 and u2, ψ(u) = v, and ψ(ui) = wi for i = 1, 2. (A) Here, w2 is a child of lca(w1, w2), but w1 is not. In this case, μ(v) = μ(v1). (B) Both w1 and w2 are not the children of v. The children, v1 and v2, are mapped into the path from the root to u under μ.
Case 2.
Only one of u1 and u2 is a leaf.
This case is similar to Case 1. Let
and
. Assume w1 is the leaf corresponding to u1 in S′, that is,
and assume that
. If w1 is a descendant of ψ(u2), then
. If w1 is not a descendant of ψ(u2), that is,
, we let v = lca(w1, ψ(u2)). Then, v = ψ(u) and μ(v) is in PS″(−, u). Since S′ and S″ do not have any common cluster except X, μ(ψ(u2)) is also in PS″(−, u). If w1 is a child of v, then μ(v) = lca(w1, μ(ψ(u2))) = μ(ψ(u2)). If w1 is not a child of v, we assume that v1 is the child of v such that
. Then μ(v1) is in the path PS″(−, u) and is comparable with μ(ψ(u2)). We obtain that μ(v) = μ(v1) if
and μ(v) = μ(ψ(u2)) otherwise. Therefore, ψ(u) is v, a duplication node under μ.
Case 3.
Both u1 and u2 are not leaves.
Assume
for i = 1, 2. If ψ(u) = ψ(ui) for some i, it is obvious that
.
If ψ(u) ≠ t1 and ψ(u) ≠ t2, we have that ψ(u) = lca(t1, t2). Let lca(t1, t2) = v and Ch(v) = {v1, v2} (Fig. 1). We prove that v is a duplication node under μ by considering the following two cases.
Case 3.1.
ti ≠ vi for i = 1, or 2. Assume t1 ≠ v1. Since t1 and u1 have different node clusters, μ(t1) is an ancestor of u1 and hence μ(v1) is in the path PS″(−, u). Since μ(v) = lca(μ(v1), μ(v2)), μ(v) must be in PS″(−, u) and thus is equal to μ(v1) if
. It is equal to μ(v2) if
, as μ(v2) is also in PS″(−, u) in this case.
Case 3.2.
t1 = v1 and t2 = v2. Since Crf(S′, S″) = 2(|X| − 2) and the roots of S′ and S″ have the same species cluster,
but CS″(u1) ≠ CS′(v1). Hence, μ(v1) is in the path PS″(−, u). Similarly, μ(v2) is also in PS″(−, u). This implies that they are comparable under
. Hence, μ(v) = lca(μ(v1), μ(v2)), equal to μ(v1) if
and equal to μ(v2) otherwise. Hence,
. ■
For
, we use μ−1(v) to denote the set of nodes that are mapped to v under μ. If v is the root of
is a subtree rooted at the root of S′. If v is a nonroot internal node in
is a union of disjoint subtrees of S′.
Lemma 2.
Let
be a nonroot node and m duplications nodes be mapped into PS″(−, p(v)). Then, μ−1(v) is a union of m + 1 subtrees of S′ at most.
Proof.
Let v be a nonroot internal node in S″. Assume that μ−1(v) is a union of
. Let Ti be rooted at
. Then, xi's are incomparable, that is,
and
for i ≠ j. This implies that
is a tree with k leaves in which k − 1 internal nodes are binary and other internal nodes have only one child.
Let y be a binary internal node that has two children in
. Clearly, y = lca(x′, x″) for some
. Assume that y1 and y2 are the children of y in S′ such that
and
. Since μ(x′) = μ(x″) = v, both μ(y1) and μ(y2) are in the path PS″(−, p(v)). This implies that μ(y1) and μ(y2) are comparable under
. Therefore, we have either
or
. If the former is true, μ(y) = μ(y1). If the latter is true, μ(y) = μ(y2). Thus, y is a duplication node under μ.
Since all the k − 1 binary internal nodes in
are duplication nodes under μ, k − 1 ≤ m. Therefore, k ≤ m + 1. ■
Lemma 3.
Let
. If there exist nonroot internal nodes
and
such that CS′(u) = CS″(v). Let
denote the tree obtained from S′ by replacing S′(u) with a single leaf and
the tree obtained from S″ by replacing S″(v) with a single leaf with the same label. Then,
![]() |
Theorem 1.
For
.
Proof.
By Lemma 3, we may just assume that only the root clusters of these two trees are identical, that is, Crf(S′, S″) = 2(|X| − 2). Let μ and ψ be defined as above. Assume that
![]() |
Since each vi is a duplication node with respect to the lca map ψ from S″ to S′,
![]() |
We further assume that μ−1(vi) contains qi nodes and di out of them are duplication nodes, for each 1 ≤ i ≤ k. Note that
is a union of subtrees of S′ and every internal node in such a subtree is a duplication node with respect to μ.
Assuming that
is a union of ti subtrees of
. Let Tj have bj internal nodes, each of which is a duplication node. Since Tj is binary, it has exactly bj + 1 leaves, each of which is not a duplication node. Therefore, we have:
![]() |
and
![]() |
Since there are at most Cdup(S′, S″) − di duplications nodes that are mapped in the path PS″(′, p(vi)), by Lemma 2, we have that ti ≤ Cdup(S′, S″) − di + 1 if vi is a nonroot node in S″, and it is trivial for the root of S″ for which t1 = 1. This implies that
![]() |
Hence, by Inequalities (8) and (9),
![]() |
By Proposition 1,
![]() |
and thus
![]() |
Using Inequality (10), we obtain the following fact:
![]() |
■
Theorem 2.
For arbitrary species trees S′, S″ over the same set of species,
![]() |
Therefore, the SDPC and the RF distance are logarithmically equivalent
Proof.
By Theorem 1,
![]() |
and thus
![]() |
On the other hand, by Inequality (7) and Lemma 3,
![]() |
This implies the following inequality:
![]() |
■
3.3. The relationships between the RF distance and the gene loss cost
Consider
. Let
such that
. Assume
and Ch(v) = {v1, v2}. We have that
is not empty for i = 1, 2. Since C(u) ≠ C(v), C(u) is a proper subset of C(v). This implies that λS′,S″(v) is in P(−, u). We consider the following two cases.
Case 1. Only one of v1 and v2 is mapped into P(−, u). We assume that v1 is mapped to a descendant of u, that is,
. Then, (p(u), u) is an edge in λS′,S″(v, v1). Since λS′,S″(v2) is in P(−, u), v2 has a descendant
such that (p(u), u) is an edge in
. Thus, there is one extra lineage in (p(u), u) at least.
Case 2. Both v1 and v2 are mapped into P(−.u). Similar to Case 1, vi has a descendant
such that (p(u), u) is an edge in
for i = 1, 2. Thus, there is one extra lineage in (p(u), u) at least. This proves:
![]() |
Conversely, if |X| ≥ 3 and Crf(S′,S″) = 2(|X| − 2), by theorem 6 in Than and Rosenberg (2013),
![]() |
It is easy to see that the result in Lemma 3 is also true for the deep coalescence cost. Combining all these facts together, we have:
![]() |
This implies the logarithmic equivalence of the RF distance and the deep coalescence cost.
For any
, Closs(S′, S″) = Cdc(S′, S″) + 2Cdup(S′, S″), proved in Zhang (2011). By Inequality (11),
![]() |
On the other hand, by Inequalities (7) and (11) and Lemma 3,
![]() |
Therefore, the RF distance and the gene loss cost are also logarithmically equivalent.
3.4. The size distribution of terraces in species tree space
Since species tree space grows exponentially with the number of taxa, understanding the landscape of the tree space is important in designing good heuristic algorithms that navigate through parts of this space guided by the parsimony score of a tree to infer species trees. Given a metric defined in the species tree space, the set of species trees with the same metric score from a species tree is called a terrace (Sanderson et al., 2011).
Here, we examine the size distribution of the DPC and RF terraces. Formally, a terrace defined by a species tree S for the DPC score c is
. Similarly, we can define the concept for the RF distance. We adopt the Furnas rank (Furnas, 1984) to display the size distributions of these terraces. It is a complete order
in the species tree space defined as follows. For
if one of the following three conditions holds:
1. The left subtree
of S′ has fewer leaves than the left subtree
of S′.2.
, but
.3.
and
are identical, but
, where
and
are the right subtrees of S and S′, respectively.
Figure 5 displays the size distributions of terraces with respect to the DPC (left panel) and RF distance (right panel) for the species tree space over 10 species. We observe the following facts:
FIG. 5.

The size distribution of terraces defined by the DPC and the RF distance. All 98 topologies of the species tree space over 10 species are listed in increasing Furnas rank along the X-axis. Each color bar indicates the portion of species trees in a terrace that have the same score (given inside a circle in the corresponding band) with the corresponding species tree.
• The two terraces defined by the two largest values contain over 95% of the trees for the RF distance, independent of the Furnas rank of the reference tree.
• In contrast, the terraces defined by the middle values of the DPC are large, whereas the terraces defined by the extreme DPC cost are small. The size of the terrace of the largest DPC value increases slightly as the Furnas rank of the reference tree increases.
Figure 6 displays the size distributions of terraces with respect to the SDPC and gene loss costs. Interestingly, these size distributions are quite similar to each other. They indicate that SDPC and gene loss costs are more sensitive to the topology of the trees than the DPC and the RF distance.
FIG. 6.

The size distribution of the terraces defined by the symmetric duplication cost (left panel) and the gene loss cost (right panel). Ninety-eight topologies of the species trees over 10 species are listed in increasing Furnas rank along the X-axis. Each color bar indicates the portion of species trees in a terrace.
4. Conclusion
All the tree reconciliation methods have high accuracy in inferring true species trees from real gene trees data. However, recently, Chaudhary et al. (2013) demonstrated by simulation that the RF-based parsimony approach often outperforms the reconciliation methods that use the gene duplication or mutation costs. A recent study of Yang and Warnow (2011) also suggests that the duplication-cost-based approach has poorer performance at large. One of our results is that the DPC and the RF distances are not equivalent. This provides a plausible explanation for the facts just mentioned. The weak equivalence of the RF distance and the gene loss cost also gives a clue to why the mutation cost is desirable for inferring species trees from a collection of gene trees (Chaudhary et al., 2013; Yang and Warnow, 2011).
Mathematical properties of the tree reconciliation costs have been studied in recent years. The deep coalescence cost has the Pareto property (Lin et al., 2012), meaning that a node cluster that appears in every input gene tree must also appear in the parsimony species tree. Than and Rosenberg (2013) report the formulas for the number of trees that achieve the maximum deep coalescence cost for a fixed gene or species tree. Górecki and Eulenstein (2014) presented a linear time algorithm for identifying a tree that maximizes the deep coalescence cost for a fixed gene or species tree. Zhang (2011) studied the relationship among the duplication cost, gene loss cost, and deep coalescence cost. The relationships between the nearest neighbor interchange distance and the reconciliation costs are also investigated (Bansal et al., 2009; Luo et al., 2011; Than and Rosenberg, 2013). However, many of mathematical properties of the reconciliation costs are incompletely known. For example, the distributions of these costs remain unclear for the uniform and Yule-Harding distribution (Harding, 1971) of species trees. In contrast, the distribution of the RF distance is much more clear. It is asymptotically Poisson with different means for several underlying species tree distributions (Steel and Penny, 1993). Our preliminary studies on terraces for the DPC and RF distances indicate that the distribution of the DPC seems to be different from that of the RF distance. Therefore, the following problems are open for future research:
1. Is the distribution of the DPC (respectively, SDPC) asymptotically normal?
2. Is the distribution of the gene loss cost asymptotically normal?
Solutions to these problems are definitely informative for studying species tree inference.
Acknowledgments
L.X.Z would like to thank Jian Chen for helpful discussions. The work was supported by the Singapore MOE (Tier-1 FRC grant R146-000-177-112).
Author Disclosure Statement
No competing financial interests exist.
References
- Ané C., Larget B., Baum D.A., et al. 2007. Bayesian estimation of concordance among gene trees. Mol. Biol. Evol. 24, 412–426 [DOI] [PubMed] [Google Scholar]
- Bansal M.S., Burleigh J.G., Eulenstein O., et al. 2010. Robinson-Foulds supertrees. Alg. Mol. Biol. 5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bansal M.S., Eulenstein O., and Wehe A.2009. The gene-duplication problem: Near-linear time algorithms for NNI-based local searches. IEEE-ACM Trans. Comput. Biol. Bioinform. 6, 221–231 [DOI] [PubMed] [Google Scholar]
- Bayzid M., Mirarab S., and Warnow T.2012. Inferring optimal species trees under gene duplication and loss, 250–261. In Pacific Symposium on Biocomputing. World Scientific, Singapore: [DOI] [PubMed] [Google Scholar]
- Bryant D.1997. Building trees, hunting for trees, and comparing trees [PhD Thesis]. Department of Mathematics, University of Canterbury, New Zealand [Google Scholar]
- Burleigh J.G., Bansal M.S., Eulenstein O., et al. 2011. Genome-scale phylogenetics: inferring the plant tree of life from 18,896 gene trees. Syst. Biol. 60, 117–125 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang W.-C., Górecki P., and Eulenstein O.2013. Exact solutions for species tree inference from discordant gene trees. J. Bioinform. Comput. Biol. 11, 5. [DOI] [PubMed] [Google Scholar]
- Chaudhary R., Burleigh J., and Fernandez-Baca D.2013. Inferring species trees from incongruent multi-copy gene trees using the robinson-foulds distance. Alg. Mol. Biol. 8, 28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dunn C.W., Hejnol A., Matus D.Q., et al. 2008. Broad phylogenomic sampling improves resolution of the animal tree of life. Nature 452, 745–749 [DOI] [PubMed] [Google Scholar]
- Fitch W.M.1970. Distinguishing homologous from analogous proteins. Syst. Biol. 19, 99–113 [PubMed] [Google Scholar]
- Furnas G.W.1984. The generation of random, binary unordered trees. J. Classification 1, 187–233 [Google Scholar]
- Goodman M., Czelusniak J., Moore G.W., et al. 1979. Fitting the gene lineage into its species lineage, a parsimony strategy illustrated by cladograms constructed from globin sequences. Syst. Zool. 28, 132–163 [Google Scholar]
- Górecki P., and Eulenstein O.2014. Maximizing deep coalescence cost. IEEE/ACM Trans. Comput. Biol. Bioinform. 11, 231–242 [DOI] [PubMed] [Google Scholar]
- Harding E.1971. The probabilities of rooted tree-shapes generated by random bifurcation. Adv. Applied Prob. 3, 44–77 [Google Scholar]
- Katz L.A., Grant J.R., Parfrey L.W., et al. 2012. Turning the crown upside down: gene tree parsimony roots the eukaryotic tree of life. Syst. Biol. 61, 653–660 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kubatko L.S., Carstens B.C., and Knowles L.L.2009. Stem: species tree estimation using maximum likelihood for gene trees under coalescence. Bioinformatics 25, 971–973 [DOI] [PubMed] [Google Scholar]
- Lin H.T., Burleigh J.G., and Eulenstein O.2012. Consensus properties for the deep coalescence problem and their application for scalable tree search. BMC Bioinformatics 13 (Suppl. 10), S12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu L., and Pearl D.K.2007. Species trees from gene trees: reconstructing Bayesian posterior distributions of a species phylogeny using estimated gene tree distributions. Syst. Biol. 56, 504–514 [DOI] [PubMed] [Google Scholar]
- Luo C.-W., Chen M.-C., Chen Y.-C., et al. 2011. Linear-time algorithms for the multiple gene duplication problems. IEEE-ACM Trans. Comput. Biol. Bioinform. 8, 260–265 [DOI] [PubMed] [Google Scholar]
- Ma B., Li M., and Zhang L.X.2000. From gene trees to species trees. SIAM J. Comput. 30, 729–752 [Google Scholar]
- Maddison W.P.1997. Gene trees in species trees. Syst. Biol. 46, 523–536 [Google Scholar]
- Pamilo P., and Nei M.1988. Relationships between gene trees and species trees. Mol. Biol. Evol. 5, 568–583 [DOI] [PubMed] [Google Scholar]
- Sanderson M.J., and McMahon M.M.2007. Inferring angiosperm phylogeny from est data with widespread gene duplication. BMC Evol. Biol. 7 (Suppl. 1), S3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sanderson M.J., McMahon M.M., and Steel M.2011. Terraces in phylogenetic tree space. Science 333, 448–450 [DOI] [PubMed] [Google Scholar]
- Semple C., and Steel M.A.2003. Phylogenetics. Oxford University Press, Oxford, U.K [Google Scholar]
- Steel M.A., and Penny D.1993. Distributions of tree comparison metrics: some new results. Syst. Biol. 42, 126–141 [Google Scholar]
- Than C., and Nakhleh L.2009. Species tree inference by minimizing deep coalescences. PLoS Comput. Biol. 5, e1000501. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Than C.V., and Rosenberg N.A.2013. Mathematical properties of the deep coalescence cost. IEEE-ACM Trans. Comput. Biol. Bioinformatics 10, 61–72 [DOI] [PubMed] [Google Scholar]
- Warnow T.2013. Large-scale multiple sequence alignment and phylogeny estimation, 85–146. InChauve C., El-Mabrouk N., and Tannier E., eds. Models and Algorithms for Genome Evolution. Springer, New York [Google Scholar]
- Wehe A., Bansal M.S., Burleigh J.G., et al. 2008. Duptree: a program for large-scale phylogenetic analyses using gene tree parsimony. Bioinformatics 24, 1540–1541 [DOI] [PubMed] [Google Scholar]
- Yang J., and Warnow T.2011. Fast and accurate methods for phylogenomic analyses. BMC Biol. 12 (Suppl. 9), S4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu Y., Warnow T., and Nakhleh L.2011. Algorithms for MDC-based multi-locus phylogeny inference: beyond rooted binary gene trees on single alleles. J. Comput. Biol. 18, 1543–1559 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang L.2011. From gene trees to species trees II: Species tree inference by minimizing deep coalescence events. IEEE-ACM Trans. Comput. Biol. Bioinform. 8, 1685–1691 [DOI] [PubMed] [Google Scholar]













































