Skip to main content
Journal of Computational Biology logoLink to Journal of Computational Biology
. 2014 Aug 1;21(8):578–590. doi: 10.1089/cmb.2014.0021

Are the Duplication Cost and Robinson-Foulds Distance Equivalent?

Yu Zheng 1, Louxin Zhang 1,,2,
PMCID: PMC4116105  PMID: 24988427

Abstract

In the tree reconciliation approach for species tree inference, a tree that has the minimum reconciliation score for given gene trees is taken as an estimate of the species tree. The scoring models used in existing tree reconciliation methods include the duplication, mutation, and deep coalescence costs. Since existing inference methods all are heuristic, their performances are often evaluated by using the Robinson-Foulds (RF) distance between the true species trees and the estimates output on simulated multi-locus datasets. To better understand these methods, we study the relationships between the duplication cost and the RF distance. We prove that the gap between the duplication cost and the RF distance is unbounded, but the symmetric duplication cost is logarithmically equivalent to the RF distance. The relationships between other reconciliation costs and the RF distance are also investigated.

Key words: : gene duplication and loss, incomplete lineage sorting, metric equivalence, Ronbinson-Foulds distance, species tree inference, tree reconciliation

1. Introduction

With more and more genomes being fully sequenced, species tree inferences from multilocus datasets have become one of the basic bioinformatics tasks in comparative and evolutionary biology (Dunn et al., 2008). Reconstructing a species tree from a multi locus dataset requires reconciling the phylogenetic information contained in the genes. The challenge for such an analysis is that the genes sampled from the same set of species often contain conflicting signals and thus have discordant gene trees. Although some of these conflicting signals can be due to systematic errors occurring in phylogenetic analysis, they often reflect gene-specific mutational events. During evolution, recombination, gene duplication, gene loss, incomplete lineage sorting, or horizontal gene transfer, all might occur in a gene family (Fitch, 1970; Goodman et al., 1979; Maddison, 1997; Pamilo and Nei, 1988).

Existing methods for inferring species trees from gene trees can be divided into two broad categories: parametric methods (Ané et al., 2007; Kubatko et al., 2009; Liu and Pearl, 2007) and tree reconciliation methods (Warnow, 2013). Parametric methods adopt gene duplication or incomplete lineage-sorting models and infer species trees using either likelihood or Bayesian framework. These approaches have strong statistical foundation, but they can be computationally expensive. Therefore, tree reconciliation methods are often used in phylogenetic studies on a genomic scale (Burleigh et al., 2011; Katz et al., 2012; Sanderson and McMahon, 2007).

Gene tree reconciliation methods take a collection of gene trees and output the tree that has the smallest reconciliation score for the gene trees as an estimate of the species tree. The scoring models used in these methods include the duplication cost (that is, the number of duplications), the mutation cost (that is, the number of gene duplication and loss events), the deep coalescence cost (that is, the number of incomplete lineage sorting events), and the Robinson-Foulds (RF) distance (Bansal et al., 2010; Bayzid et al., 2012; Chang et al., 2013; Chaudhary et al., 2013; Than and Nakhleh, 2009; Wehe et al., 2008; Yu et al., 2011). Since finding the parsimony tree from gene trees is NP-hard for the reconciliation scoring models (Bryant, 1997; Ma et al., 2000; Zhang, 2011), existing methods all compute a species tree by navigating through the part of the species tree space guided by the reconciliation score. The performances of these heuristic methods are usually evaluated using the RF distances between true species trees and their estimate output on simulation datasets (Chaudhary et al., 2013; Than and Nakhleh, 2009; Yang and Warnow, 2011). Hence, uncovering the relationships between the RF distance and the reconciliation costs is important to design good gene tree reconciliation methods. Unfortunately, many of these relationships are incompletely known. In particular, the relationship between the RF distance and the duplication cost is rather elusive, as both are bounded above by a linear function in the number of taxa. In this work, we mainly investigate this relationship. We consider all the reconciliation costs as different metrics in species tree spaces.

The rest of this article is divided into three sections. The basic concepts, notions, and the various reconciliation costs are introduced in Section 2. Several results are presented in Section 3. We first show that the gap between the duplication cost and the RF distance is unbounded. On the other hand, the symmetric duplication cost and the RF distance are proved to be logarithmically equivalent. A similar relationship between the gene loss cost and the RF distance is obtained. We also examine the distribution of the sizes of terraces for these metrics. We conclude with suggestions for future work in Section 4.

2. Concepts and Notions

2.1. Species trees

A species tree T over a set of species X is a rooted tree in which one node is designated as the root, all edges are oriented away from the root, and the species in X are one-to-one mapped to leaves. We use V(T) and E(T) to denote the sets of nodes and (directed) edges in T, respectively. For Inline graphic, v is the parent of u and, equivalently, u is a child of v if Inline graphic. A node is called a leaf if it does not have any child. In this article, we focus on binary species trees in which every nonleaf node has two children. We also use the following notation:

  • • p(u) denotes the parent of a nonroot node u in a tree;

  • • Ch(u) denotes the set of two children of a nonleaf u in a tree;

  • • Vlf(T) denotes the set of leaves in T;

  • • Vit(T) denotes the set of internal (i.e., nonleaf) nodes in T;

  • • T(u) denotes the subtree rooted at Inline graphic, which consists of u and its descendants of u;

  • • PT(−, u) denotes the unique path from the root to u in T;

  • • T|U denotes the subtrees induced by a subset Inline graphic, whose sets of nodes and branches are Inline graphic and Inline graphic, respectively.

For Inline graphic, v is an ancestor of u and, equivalently, u is a descendant of v, denoted by Inline graphic, if v is in PT(−, u). We write Inline graphic if the relationship is uncertain, but we know that v = u or Inline graphic. Therefore, it is not hard to see that v is a leaf (terminal node) if and only if Inline graphic for any Inline graphic. For Inline graphic, lca(U) denotes the latest common ancestor (lca) of the nodes in U.

2.2. The RF distance

Let X be a set of species, and let Inline graphic denote the set of all species tree over X. Consider Inline graphic. A species Inline graphic is called the label of the corresponding leaf Inline graphic, written l(u). Conversely, we define l−1(x) = u. For Inline graphic, we define:

graphic file with name eq19.gif

called the (species) cluster of u. Define

graphic file with name eq20.gif

Consider two species trees S′ and S″ over X. They are identical if and only Inline graphic. Therefore, the dissimilarity of S′ and S″ can be measured by the number of clusters that are found in one but not in the other, that is,

graphic file with name eq22.gif

called the RF distance (Semple and Steel, 2003). Since both S′ and S″ have |X| − 1 internal nodes and their roots' cluster is X, The following inequalities hold:

graphic file with name eq23.gif

2.3. Tree reconciliation costs

Consider Inline graphic. Since each leaf in these trees has a unique label in X, there is a one-to-one map between Vlf(S′) and Vlf(S″). Such a correspondence can be extended into a map from V(S′) to V(S″), λS′,S, as:

graphic file with name eq25.gif

The duplication cost (DPC), Cdup(S′, S″), is defined as |Dup(S′, S″)|, where

graphic file with name eq26.gif

The nodes in Dup(S′, S″) are called duplication nodes.

Proposition 1.For Inline graphic.

Proof. Since S ≠ S″, there exists Inline graphic having the following properties (Fig. 1):

FIG. 1.

FIG. 1.

Illustration of the proof of Proposition 1. The cluster of Inline graphic is not found in S′, but the clusters of the descendants of u are all in S′. If Inline graphic such that Inline graphic, then lca(t1, t2) is a duplication node under λS′,S.

  • (i) CS(u) ≠ CS(v) for any Inline graphic;

  • (ii) for each descendant t of u, that is, Inline graphic, CS(t) = CS(v) for some Inline graphic.

Let Ch(u) = {u1, u2}. By the above assumption, CS(ui) = CS(ti) for some Inline graphic for i = 1, 2. Since CS(u) is not a cluster in S′, t1 and t2 must not be siblings in S′, that is, p(t1) ≠ p(t2). Let lca(t1, t2) = v and Ch(v) = {v1, v2} such that Inline graphic and Inline graphic. Since t1 and t2 are not siblings, either t1 ≠ v1 or t2 ≠ v2. If t2 ≠ v2, λS′,S(v2) is in PS(−, u). If t1 ≠ v1, λS,S(v1) is in PS(−, u). Thus, λS′,S(v1), λS′,S(v2), or both are in PS(−, u), the path from r(T″) to u.

If Inline graphic. If Inline graphic. Therefore, v is a duplication node and hence Cdup(S′,S″) ≥ 1.

Let Inline graphic have the largest distance from the root of S′. Then, x must have two leaf children, x1 and x2. By definition, λS,S(x1) ≠ λS′,S(x) ≠ λS′,S(t2). Hence, x is not a duplication node when S′ is mapped onto S″. Since S′ has |X| − 1 internal nodes, S′ has at most |X| − 1 − 1 duplication nodes. Therefore, Cdup(S′,S″) ≤ |X| − 2.   ■

For Inline graphic, we use λS,S(e′) to denote the path from λS,S(u) to λS′,S″(v). We further use |λS′,S(u, v)| to denote the number of the non-end nodes in λS′,S(e′). |λS′,S(u, v)| = 0 if λS′,S(u) = λS′,S(v). The gene loss cost, Closs(S′, S″), is defined as:

graphic file with name eq42.gif

where χu = 1 if u is a duplication node such that |λS′,S(u, û)| ≠ 0 for some Inline graphic and χu = 0 otherwise.

If S′ is a gene tree and S″ is a species tree, Cdup(S′, S″) and Closs(S′, S″) are the numbers of gene duplications and losses occurring in the duplication history of the genes represented by the leaves in S′ defined by the lca map (Goodman et al., 1979).

For Inline graphic, define:

graphic file with name eq45.gif

Clearly, Inline graphic is in DC(e″) if and only if e″ is an edge in λS′,S(e). The deep coalescence cost, Cdc(S′,S″), is defined as:

graphic file with name eq47.gif

The deep coalescence cost Cdc(S′, S″) is introduced to measure the dissimilarity of S′ and S″ that are caused by incomplete lineage sorting for a gene tree S′ and a species tree S″ (Maddison, 1997).

Finally, we note that all the reconciliation costs introduced above are not a symmetric metric in general. Hence, the symmetric duplication cost (SDPC) of S′ and S″, defined as Cdup(S′, S″) + Cdup(S″, S′), is used to study species tree inference (Ma et al., 2000).

3. Results

3.1. Unbounded gap between the DPC and RF distance

For two different trees over the same set of species, both the duplication cost and half of the RF distance are in the same range from 1 to n − 2, where n is the number of species involved in the trees. Therefore, the following question naturally arises:

Are the DPC and RF distance metrically equivalent?

and, equivalently, does there exist c1 ≥ c2 > 0 such that

graphic file with name eq48.gif

for any Inline graphic?

Consider Inline graphic such that Crf(S1, S2) = 2(|X| − 2), that is, Inline graphic. By definition, the parent of two sibling leaves in S1 is not a duplication node under λS1, S2. Since such an internal node always exists in S1, we have:

graphic file with name eq52.gif

However, surprisingly, there are many pairs of trees S1 and S2 in Inline graphic satisfying that Crf(S1, S2) = 2(|X| − 2) but Cdup(S1, S2) = 1, where |X| ≥ 3. For example, we consider a class of complete binary species trees, as illustrated in Figure 2. Let n = 2k for some k ≥ 1 and Inline graphic. We set T to be the rooted complete binary tree with n unlabeled leaves. An denotes the species tree over X obtained from T by labeling the i-th leaf (counting from left to right) with i (Fig. 2). Bn denotes the species tree obtained from T by labeling the i-th leaf with φ(i) that is defined as:

graphic file with name eq55.gif

FIG. 2.

FIG. 2.

Complete binary species trees An and Bn for n = 23. Cdup(An, Bn) = 1 but Crf(An, Bn) = 2(n − 2).

for Inline graphic.

Proposition 2.(1). Crf(An, Bn) = 2(n − 2).

  • (2). Cdup(An, Bn) = 1.

  • (3). The DPC and the RF distance are not equivalent.

Proof.

(1). For any nonroot Inline graphic, either Inline graphic, or Inline graphic. In Bn, the (2i − 1)th and (2i)th leaves are siblings, labeled with i and n/2 + i, respectively, where i ≤ n/2. Therefore, for Inline graphic,

graphic file with name eq61.gif

Combining these two facts together, we conclude that only the roots of An and Bn have the same cluster X. Hence, Crf(An, Bn) = 2(n − 2), as both trees have n − 1 internal nodes.

(2). Consider Inline graphic. Notice that Inline graphic for some i and k. For the left and right children, u1 and u2, of u,

graphic file with name eq64.gif
graphic file with name eq65.gif

If u is in the left subtree of An, then Inline graphic and

graphic file with name eq67.gif

If u is in the right subtree of An,

graphic file with name eq68.gif

Hence, the root r of An and its two children are mapped to the root of Bn, implying that r is a duplication node.

If u is in the left subtree of An such that Inline graphic, then,

graphic file with name eq70.gif
graphic file with name eq71.gif
graphic file with name eq72.gif

This implies that u is not a duplication node. Similarly, all the nodes in the right subtree are not duplication nodes. Therefore, Cdup(An, Bn) = 1.

(3). This statement follows directly from (1) and (2).   ■

Although Cdup(An, Bn) = 1, we have that Cdup(Bn, An) = n/2 − 2. However, for any fixed large k > 0, there exist G and S such that Inline graphic when X is large enough. For example, we take an integer m and define

graphic file with name eq74.gif

Consider the two trees G and S over X defined in Figure 3. It is easy to see that Crf(G, S) = 2(m2 − 2). The lca map λG,S maps different nodes in Gi to different nodes in S for each i. Furthermore, it maps the roots of Inline graphic and their ancestors to the root of S. Thus, Cdup(G, S) = m − 1. Similarly, we can also show that Cdup(S, G) = m − 1. Hence, for any k < m, we have Inline graphic.

FIG. 3.

FIG. 3.

Trees G and S over m2 species satisfying that Crf(G, S) = 2(m2 − 2), but Cdup(S, G) = Cdup(G, S) = m − 1.

This example shows that even the symmetric duplication cost (SDPC) is not equivalent to the RF distance in the species tree space. However, they are weakly equivalent.

3.2. Logarithmic equivalence of the SDPC and RF distance

Consider Inline graphic. For the sake of simplicity, we use the following notations:

  • • μ denotes the lca map from S′ to S″, that is, μ = λS′,S;

  • • ψ = λS″,S;

  • • Dupμ(S′, S″) denotes the set of duplication nodes in S′ under μ.

Lemma 1.

For Inline graphic such that Crf(S′, S″) = 2(|X| − 2), Inline graphic and Inline graphic.

Proof.

We prove the first fact by mathematical induction. By symmetry, the second fact holds.

Consider Inline graphic. It has two children, u1 and u2. There are three possible cases.

Case 1.

u1 and u2 are leaves.

Since S′ and S″ do not have any common node cluster except X, the leaves, w1 and w2, that have the same labels as u1 and u2 are not siblings in S′. By definition, μ(wi) = ui and ψ(ui) = wi for i = 1, 2. Let v = lca(w1, w2). By assumption, ψ(u) = lca(ψ(u1), ψ(u2)) = lca(w1, w2) = v. Assume that the children of v are v1 and v2 such that Inline graphic. Since w1 and w2 are not siblings, we have that v1 ≠ w1, v2 ≠ w2, or both. Without loss of generality, we may assume that v1 ≠ w1. Since μ(w1) = u1, which is a leaf, and μ(v1) ≠ u1, μ(v1) is either equal to u or an ancestor of u, that is, μ(v1) is in the path PS(−, u). If v2 = w2, μ(v) = lca{μ(v1), μ(v2)} = μ(v1) (Fig. 4A), as μ(v2) ( = u2) is a child of u and hence a descendant of μ(v1). If v2 ≠ w2, μ(v2) is in PS(−, u) and hence μ(v1) and μ(v2) are comparable under Inline graphic, as they are in the path from the root to u (Fig. 4B). If Inline graphic, μ(v) = μ(v2). If Inline graphic, μ(v) = μ(v1). Therefore, Inline graphic.

FIG. 4.

FIG. 4.

Illustration of the proof of Lemma 1 in Case 1. Here, u has two leaf children, u1 and u2, ψ(u) = v, and ψ(ui) = wi for i = 1, 2. (A) Here, w2 is a child of lca(w1, w2), but w1 is not. In this case, μ(v) = μ(v1). (B) Both w1 and w2 are not the children of v. The children, v1 and v2, are mapped into the path from the root to u under μ.

Case 2.

Only one of u1 and u2 is a leaf.

This case is similar to Case 1. Let Inline graphic and Inline graphic. Assume w1 is the leaf corresponding to u1 in S′, that is, Inline graphic and assume that Inline graphic. If w1 is a descendant of ψ(u2), then Inline graphic. If w1 is not a descendant of ψ(u2), that is, Inline graphic, we let v = lca(w1, ψ(u2)). Then, v = ψ(u) and μ(v) is in PS(−, u). Since S′ and S″ do not have any common cluster except X, μ(ψ(u2)) is also in PS(−, u). If w1 is a child of v, then μ(v) = lca(w1, μ(ψ(u2))) = μ(ψ(u2)). If w1 is not a child of v, we assume that v1 is the child of v such that Inline graphic. Then μ(v1) is in the path PS(−, u) and is comparable with μ(ψ(u2)). We obtain that μ(v) = μ(v1) if Inline graphic and μ(v) = μ(ψ(u2)) otherwise. Therefore, ψ(u) is v, a duplication node under μ.

Case 3.

Both u1 and u2 are not leaves.

Assume Inline graphic for i = 1, 2. If ψ(u) = ψ(ui) for some i, it is obvious that Inline graphic.

If ψ(u) ≠ t1 and ψ(u) ≠ t2, we have that ψ(u) = lca(t1, t2). Let lca(t1, t2) = v and Ch(v) = {v1, v2} (Fig. 1). We prove that v is a duplication node under μ by considering the following two cases.

Case 3.1.

ti ≠ vi for i = 1, or 2. Assume t1 ≠ v1. Since t1 and u1 have different node clusters, μ(t1) is an ancestor of u1 and hence μ(v1) is in the path PS(−, u). Since μ(v) = lca(μ(v1), μ(v2)), μ(v) must be in PS(−, u) and thus is equal to μ(v1) if Inline graphic. It is equal to μ(v2) if Inline graphic, as μ(v2) is also in PS(−, u) in this case.

Case 3.2.

t1 = v1 and t2 = v2. Since Crf(S′, S″) = 2(|X| − 2) and the roots of S′ and S″ have the same species cluster, Inline graphic but CS(u1) ≠ CS(v1). Hence, μ(v1) is in the path PS(−, u). Similarly, μ(v2) is also in PS(−, u). This implies that they are comparable under Inline graphic. Hence, μ(v) = lca(μ(v1), μ(v2)), equal to μ(v1) if Inline graphic and equal to μ(v2) otherwise. Hence, Inline graphic.   ■

For Inline graphic, we use μ−1(v) to denote the set of nodes that are mapped to v under μ. If v is the root of Inline graphic is a subtree rooted at the root of S′. If v is a nonroot internal node in Inline graphic is a union of disjoint subtrees of S′.

Lemma 2.

Let Inline graphic be a nonroot node and m duplications nodes be mapped into PS(−, p(v)). Then, μ−1(v) is a union of m + 1 subtrees of Sat most.

Proof.

Let v be a nonroot internal node in S″. Assume that μ−1(v) is a union of Inline graphic. Let Ti be rooted at Inline graphic. Then, xi's are incomparable, that is, Inline graphic and Inline graphic for i ≠ j. This implies that Inline graphic is a tree with k leaves in which k − 1 internal nodes are binary and other internal nodes have only one child.

Let y be a binary internal node that has two children in Inline graphic. Clearly, y = lca(x′, x″) for some Inline graphic. Assume that y1 and y2 are the children of y in S′ such that Inline graphic and Inline graphic. Since μ(x′) = μ(x″) = v, both μ(y1) and μ(y2) are in the path PS(−, p(v)). This implies that μ(y1) and μ(y2) are comparable under Inline graphic. Therefore, we have either Inline graphic or Inline graphic. If the former is true, μ(y) = μ(y1). If the latter is true, μ(y) = μ(y2). Thus, y is a duplication node under μ.

Since all the k − 1 binary internal nodes in Inline graphic are duplication nodes under μ, k − 1 ≤ m. Therefore, k ≤ m + 1.   ■

Lemma 3.

Let Inline graphic. If there exist nonroot internal nodes Inline graphic and Inline graphic such that CS(u) = CS(v). Let Inline graphic denote the tree obtained from Sby replacing S′(u) with a single leaf and Inline graphic the tree obtained from Sby replacing S″(v) with a single leaf with the same label. Then,

graphic file with name eq125.gif

Theorem 1.

For Inline graphic.

Proof.

By Lemma 3, we may just assume that only the root clusters of these two trees are identical, that is, Crf(S′, S″) = 2(|X| − 2). Let μ and ψ be defined as above. Assume that

graphic file with name eq127.gif

Since each vi is a duplication node with respect to the lca map ψ from S″ to S′,

graphic file with name eq128.gif

We further assume that μ−1(vi) contains qi nodes and di out of them are duplication nodes, for each 1 ≤ i ≤ k. Note that Inline graphic is a union of subtrees of S′ and every internal node in such a subtree is a duplication node with respect to μ.

Assuming that Inline graphic is a union of ti subtrees of Inline graphic. Let Tj have bj internal nodes, each of which is a duplication node. Since Tj is binary, it has exactly bj + 1 leaves, each of which is not a duplication node. Therefore, we have:

graphic file with name eq132.gif

and

graphic file with name eq133.gif

Since there are at most Cdup(S′, S″) − di duplications nodes that are mapped in the path PS(′, p(vi)), by Lemma 2, we have that ti ≤ Cdup(S′, S″) − di + 1 if vi is a nonroot node in S″, and it is trivial for the root of S″ for which t1 = 1. This implies that

graphic file with name eq134.gif

Hence, by Inequalities (8) and (9),

graphic file with name eq135.gif

By Proposition 1,

graphic file with name eq136.gif

and thus

graphic file with name eq137.gif

Using Inequality (10), we obtain the following fact:

graphic file with name eq138.gif

   ■

Theorem 2.

For arbitrary species trees S′, Sover the same set of species,

graphic file with name eq139.gif

Therefore, the SDPC and the RF distance are logarithmically equivalent

Proof.

By Theorem 1,

graphic file with name eq140.gif

and thus

graphic file with name eq141.gif

On the other hand, by Inequality (7) and Lemma 3,

graphic file with name eq142.gif

This implies the following inequality:

graphic file with name eq143.gif

   ■

3.3. The relationships between the RF distance and the gene loss cost

Consider Inline graphic. Let Inline graphic such that Inline graphic. Assume Inline graphic and Ch(v) = {v1, v2}. We have that Inline graphic is not empty for i = 1, 2. Since C(u) ≠ C(v), C(u) is a proper subset of C(v). This implies that λS′,S(v) is in P(−, u). We consider the following two cases.

Case 1. Only one of v1 and v2 is mapped into P(−, u). We assume that v1 is mapped to a descendant of u, that is, Inline graphic. Then, (p(u), u) is an edge in λS′,S(v, v1). Since λS′,S(v2) is in P(−, u), v2 has a descendant Inline graphic such that (p(u), u) is an edge in Inline graphic. Thus, there is one extra lineage in (p(u), u) at least.

Case 2. Both v1 and v2 are mapped into P(−.u). Similar to Case 1, vi has a descendant Inline graphic such that (p(u), u) is an edge in Inline graphic for i = 1, 2. Thus, there is one extra lineage in (p(u), u) at least. This proves:

graphic file with name eq154.gif

Conversely, if |X| ≥ 3 and Crf(S′,S″) = 2(|X| − 2), by theorem 6 in Than and Rosenberg (2013),

graphic file with name eq155.gif

It is easy to see that the result in Lemma 3 is also true for the deep coalescence cost. Combining all these facts together, we have:

graphic file with name eq156.gif

This implies the logarithmic equivalence of the RF distance and the deep coalescence cost.

For any Inline graphic, Closs(S′, S″) = Cdc(S′, S″) + 2Cdup(S′, S″), proved in Zhang (2011). By Inequality (11),

graphic file with name eq158.gif

On the other hand, by Inequalities (7) and (11) and Lemma 3,

graphic file with name eq159.gif

Therefore, the RF distance and the gene loss cost are also logarithmically equivalent.

3.4. The size distribution of terraces in species tree space

Since species tree space grows exponentially with the number of taxa, understanding the landscape of the tree space is important in designing good heuristic algorithms that navigate through parts of this space guided by the parsimony score of a tree to infer species trees. Given a metric defined in the species tree space, the set of species trees with the same metric score from a species tree is called a terrace (Sanderson et al., 2011).

Here, we examine the size distribution of the DPC and RF terraces. Formally, a terrace defined by a species tree S for the DPC score c is Inline graphic. Similarly, we can define the concept for the RF distance. We adopt the Furnas rank (Furnas, 1984) to display the size distributions of these terraces. It is a complete order Inline graphic in the species tree space defined as follows. For Inline graphic if one of the following three conditions holds:

  • 1. The left subtree Inline graphic of S′ has fewer leaves than the left subtree Inline graphic of S′.

  • 2. Inline graphic, but Inline graphic.

  • 3. Inline graphic and Inline graphic are identical, but Inline graphic, where Inline graphic and Inline graphic are the right subtrees of S and S′, respectively.

Figure 5 displays the size distributions of terraces with respect to the DPC (left panel) and RF distance (right panel) for the species tree space over 10 species. We observe the following facts:

FIG. 5.

FIG. 5.

The size distribution of terraces defined by the DPC and the RF distance. All 98 topologies of the species tree space over 10 species are listed in increasing Furnas rank along the X-axis. Each color bar indicates the portion of species trees in a terrace that have the same score (given inside a circle in the corresponding band) with the corresponding species tree.

  • • The two terraces defined by the two largest values contain over 95% of the trees for the RF distance, independent of the Furnas rank of the reference tree.

  • • In contrast, the terraces defined by the middle values of the DPC are large, whereas the terraces defined by the extreme DPC cost are small. The size of the terrace of the largest DPC value increases slightly as the Furnas rank of the reference tree increases.

Figure 6 displays the size distributions of terraces with respect to the SDPC and gene loss costs. Interestingly, these size distributions are quite similar to each other. They indicate that SDPC and gene loss costs are more sensitive to the topology of the trees than the DPC and the RF distance.

FIG. 6.

FIG. 6.

The size distribution of the terraces defined by the symmetric duplication cost (left panel) and the gene loss cost (right panel). Ninety-eight topologies of the species trees over 10 species are listed in increasing Furnas rank along the X-axis. Each color bar indicates the portion of species trees in a terrace.

4. Conclusion

All the tree reconciliation methods have high accuracy in inferring true species trees from real gene trees data. However, recently, Chaudhary et al. (2013) demonstrated by simulation that the RF-based parsimony approach often outperforms the reconciliation methods that use the gene duplication or mutation costs. A recent study of Yang and Warnow (2011) also suggests that the duplication-cost-based approach has poorer performance at large. One of our results is that the DPC and the RF distances are not equivalent. This provides a plausible explanation for the facts just mentioned. The weak equivalence of the RF distance and the gene loss cost also gives a clue to why the mutation cost is desirable for inferring species trees from a collection of gene trees (Chaudhary et al., 2013; Yang and Warnow, 2011).

Mathematical properties of the tree reconciliation costs have been studied in recent years. The deep coalescence cost has the Pareto property (Lin et al., 2012), meaning that a node cluster that appears in every input gene tree must also appear in the parsimony species tree. Than and Rosenberg (2013) report the formulas for the number of trees that achieve the maximum deep coalescence cost for a fixed gene or species tree. Górecki and Eulenstein (2014) presented a linear time algorithm for identifying a tree that maximizes the deep coalescence cost for a fixed gene or species tree. Zhang (2011) studied the relationship among the duplication cost, gene loss cost, and deep coalescence cost. The relationships between the nearest neighbor interchange distance and the reconciliation costs are also investigated (Bansal et al., 2009; Luo et al., 2011; Than and Rosenberg, 2013). However, many of mathematical properties of the reconciliation costs are incompletely known. For example, the distributions of these costs remain unclear for the uniform and Yule-Harding distribution (Harding, 1971) of species trees. In contrast, the distribution of the RF distance is much more clear. It is asymptotically Poisson with different means for several underlying species tree distributions (Steel and Penny, 1993). Our preliminary studies on terraces for the DPC and RF distances indicate that the distribution of the DPC seems to be different from that of the RF distance. Therefore, the following problems are open for future research:

  • 1. Is the distribution of the DPC (respectively, SDPC) asymptotically normal?

  • 2. Is the distribution of the gene loss cost asymptotically normal?

Solutions to these problems are definitely informative for studying species tree inference.

Acknowledgments

L.X.Z would like to thank Jian Chen for helpful discussions. The work was supported by the Singapore MOE (Tier-1 FRC grant R146-000-177-112).

Author Disclosure Statement

No competing financial interests exist.

References

  1. Ané C., Larget B., Baum D.A., et al. 2007. Bayesian estimation of concordance among gene trees. Mol. Biol. Evol. 24, 412–426 [DOI] [PubMed] [Google Scholar]
  2. Bansal M.S., Burleigh J.G., Eulenstein O., et al. 2010. Robinson-Foulds supertrees. Alg. Mol. Biol. 5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bansal M.S., Eulenstein O., and Wehe A.2009. The gene-duplication problem: Near-linear time algorithms for NNI-based local searches. IEEE-ACM Trans. Comput. Biol. Bioinform. 6, 221–231 [DOI] [PubMed] [Google Scholar]
  4. Bayzid M., Mirarab S., and Warnow T.2012. Inferring optimal species trees under gene duplication and loss, 250–261. In Pacific Symposium on Biocomputing. World Scientific, Singapore: [DOI] [PubMed] [Google Scholar]
  5. Bryant D.1997. Building trees, hunting for trees, and comparing trees [PhD Thesis]. Department of Mathematics, University of Canterbury, New Zealand [Google Scholar]
  6. Burleigh J.G., Bansal M.S., Eulenstein O., et al. 2011. Genome-scale phylogenetics: inferring the plant tree of life from 18,896 gene trees. Syst. Biol. 60, 117–125 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Chang W.-C., Górecki P., and Eulenstein O.2013. Exact solutions for species tree inference from discordant gene trees. J. Bioinform. Comput. Biol. 11, 5. [DOI] [PubMed] [Google Scholar]
  8. Chaudhary R., Burleigh J., and Fernandez-Baca D.2013. Inferring species trees from incongruent multi-copy gene trees using the robinson-foulds distance. Alg. Mol. Biol. 8, 28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Dunn C.W., Hejnol A., Matus D.Q., et al. 2008. Broad phylogenomic sampling improves resolution of the animal tree of life. Nature 452, 745–749 [DOI] [PubMed] [Google Scholar]
  10. Fitch W.M.1970. Distinguishing homologous from analogous proteins. Syst. Biol. 19, 99–113 [PubMed] [Google Scholar]
  11. Furnas G.W.1984. The generation of random, binary unordered trees. J. Classification 1, 187–233 [Google Scholar]
  12. Goodman M., Czelusniak J., Moore G.W., et al. 1979. Fitting the gene lineage into its species lineage, a parsimony strategy illustrated by cladograms constructed from globin sequences. Syst. Zool. 28, 132–163 [Google Scholar]
  13. Górecki P., and Eulenstein O.2014. Maximizing deep coalescence cost. IEEE/ACM Trans. Comput. Biol. Bioinform. 11, 231–242 [DOI] [PubMed] [Google Scholar]
  14. Harding E.1971. The probabilities of rooted tree-shapes generated by random bifurcation. Adv. Applied Prob. 3, 44–77 [Google Scholar]
  15. Katz L.A., Grant J.R., Parfrey L.W., et al. 2012. Turning the crown upside down: gene tree parsimony roots the eukaryotic tree of life. Syst. Biol. 61, 653–660 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Kubatko L.S., Carstens B.C., and Knowles L.L.2009. Stem: species tree estimation using maximum likelihood for gene trees under coalescence. Bioinformatics 25, 971–973 [DOI] [PubMed] [Google Scholar]
  17. Lin H.T., Burleigh J.G., and Eulenstein O.2012. Consensus properties for the deep coalescence problem and their application for scalable tree search. BMC Bioinformatics 13 (Suppl. 10), S12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Liu L., and Pearl D.K.2007. Species trees from gene trees: reconstructing Bayesian posterior distributions of a species phylogeny using estimated gene tree distributions. Syst. Biol. 56, 504–514 [DOI] [PubMed] [Google Scholar]
  19. Luo C.-W., Chen M.-C., Chen Y.-C., et al. 2011. Linear-time algorithms for the multiple gene duplication problems. IEEE-ACM Trans. Comput. Biol. Bioinform. 8, 260–265 [DOI] [PubMed] [Google Scholar]
  20. Ma B., Li M., and Zhang L.X.2000. From gene trees to species trees. SIAM J. Comput. 30, 729–752 [Google Scholar]
  21. Maddison W.P.1997. Gene trees in species trees. Syst. Biol. 46, 523–536 [Google Scholar]
  22. Pamilo P., and Nei M.1988. Relationships between gene trees and species trees. Mol. Biol. Evol. 5, 568–583 [DOI] [PubMed] [Google Scholar]
  23. Sanderson M.J., and McMahon M.M.2007. Inferring angiosperm phylogeny from est data with widespread gene duplication. BMC Evol. Biol. 7 (Suppl. 1), S3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Sanderson M.J., McMahon M.M., and Steel M.2011. Terraces in phylogenetic tree space. Science 333, 448–450 [DOI] [PubMed] [Google Scholar]
  25. Semple C., and Steel M.A.2003. Phylogenetics. Oxford University Press, Oxford, U.K [Google Scholar]
  26. Steel M.A., and Penny D.1993. Distributions of tree comparison metrics: some new results. Syst. Biol. 42, 126–141 [Google Scholar]
  27. Than C., and Nakhleh L.2009. Species tree inference by minimizing deep coalescences. PLoS Comput. Biol. 5, e1000501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Than C.V., and Rosenberg N.A.2013. Mathematical properties of the deep coalescence cost. IEEE-ACM Trans. Comput. Biol. Bioinformatics 10, 61–72 [DOI] [PubMed] [Google Scholar]
  29. Warnow T.2013. Large-scale multiple sequence alignment and phylogeny estimation, 85–146. InChauve C., El-Mabrouk N., and Tannier E., eds. Models and Algorithms for Genome Evolution. Springer, New York [Google Scholar]
  30. Wehe A., Bansal M.S., Burleigh J.G., et al. 2008. Duptree: a program for large-scale phylogenetic analyses using gene tree parsimony. Bioinformatics 24, 1540–1541 [DOI] [PubMed] [Google Scholar]
  31. Yang J., and Warnow T.2011. Fast and accurate methods for phylogenomic analyses. BMC Biol. 12 (Suppl. 9), S4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Yu Y., Warnow T., and Nakhleh L.2011. Algorithms for MDC-based multi-locus phylogeny inference: beyond rooted binary gene trees on single alleles. J. Comput. Biol. 18, 1543–1559 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Zhang L.2011. From gene trees to species trees II: Species tree inference by minimizing deep coalescence events. IEEE-ACM Trans. Comput. Biol. Bioinform. 8, 1685–1691 [DOI] [PubMed] [Google Scholar]

Articles from Journal of Computational Biology are provided here courtesy of SAGE Publications

RESOURCES