Skip to main content
Molecular Biology and Evolution logoLink to Molecular Biology and Evolution
. 2025 Jun 6;42(6):msaf124. doi: 10.1093/molbev/msaf124

A General Substitution Matrix for Structural Phylogenetics

Sriram G Garg 1,, Georg K A Hochberg 2,3,4,
Editor: Tal Pupko
PMCID: PMC12198762  PMID: 40476610

Abstract

Sequence-based maximum likelihood phylogenetics is a widely used method for inferring evolutionary relationships, which has illuminated the evolutionary histories of proteins and the organisms that harbor them. However, modern implementations with sophisticated models of sequence evolution struggle to resolve deep evolutionary relationships, which can be obscured by excessive sequence divergence and substitution saturation. Structural phylogenetics has emerged as a promising alternative because protein structure evolves much more slowly than protein sequences. Recent developments in protein structure prediction using AI have made it possible to predict protein structures for entire protein families and then to translate these structures into a sequence representation—the 3Di structural alphabet—that can in theory be directly fed into existing sequence-based phylogenetic software. To unlock the full potential of this idea, however, requires the inference of a general substitution matrix for structural phylogenetics, which has so far been missing. Here, we infer this matrix from large datasets of protein structures and show that it results in a better fit to empirical datasets than previous approaches. We then use this matrix to re-visit the question of the root of the tree of life. Using structural phylogenies of universal paralogs, we provide the first unambiguous evidence for a root between archaea and bacteria. Finally, we discuss some practical and conceptual limitations of structural phylogenetics. Our 3Di substitution matrix provides a starting point for revisiting many deep phylogenetic problems that have so far been extremely difficult to solve.

Keywords: phylogenetics, maximum likelihood, structural phylogenetics, evolution, substitution models

Introduction

The field of phylogenetics has evolved from relying on morphological comparisons to sophisticated sequence-based analyses (Whelan et al. 2001). The advent of computational methods marked a turning point, introducing a range of algorithms from neighbor joining (NJ) (Saitou and Nei 1987) and maximum parsimony (MP) (Farris 1970; Fitch 1971) to maximum likelihood (ML) (Felsenstein 1981) and Bayesian inferences (Rannala and Yang 1996 ; Mau and Newton 1997) on nucleotide and amino acid (AA) sequences. Each methodological leap has brought with it a deeper understanding of evolutionary history through better trees. Among the various phylogenetic methods, ML approaches have emerged as particularly powerful tools for modeling evolutionary processes (Posada and Crandall 2021). The flexibility and robustness of ML techniques have made them indispensable for contemporary phylogenetic studies, especially those tackling large datasets or seeking to resolve deep evolutionary relationships. But especially deep, sequenced-based phylogenetics remains difficult. Substitution saturation is a particular challenge, in which each site in the alignment has accumulated multiple substitutions over a branch of interest (Philippe and Forterre 1999). Depending on the accuracy of the substitution model of sequence evolution, saturation can lead to spurious phylogenetic signals and artifacts in phylogenetic trees (Felsenstein 2003). The problem of saturation cannot always be solved by adding more sequences (Philippe et al. 2011) or by better models of sequence evolution.

Saturation is a relevant problem for the identification of the root of the tree of life. It is traditionally placed on the branch between bacteria and archaea (Gouy et al. 2015), which has important implications for the nature of the last universal common ancestor (LUCA). This inference is based on paralog rooting with universally duplicated genes, where the paralogs reciprocally root each other (Iwabe et al. 1989). Although this root is tacitly accepted by the majority of biologists, the paralog trees are riddled with potential problems. In all previous attempts, the branch between universal paralogs remains so long as to be probably saturated (Brown and Doolittle 1995; Philippe and Forterre 1999; Gouy et al. 2015; Mahendrarajah et al. 2023). This means that the root position within each paralog might be mostly determined by the preferences of the substitution model, rather than real phylogenetic signal, which has been erased almost entirely. Some phylogeneticists, therefore, still consider the root of the tree of life an unsolved problem (Gouy et al. 2015).

Structural phylogenetics offers a potentially powerful alternative to traditional sequence-based approaches. Structures evolve much more slowly than sequences, and if a model for structural evolution could be inferred, this could help resolve phylogenies that are beyond the reach of sequenced-based methods. Early attempts at this idea were limited by the lack of high-quality protein structures or reliable methods of scoring multiple sequence alignments (MSAs) of protein structures (Johnson et al. 1990a, 1990b; Balaji et al. 2001; Balaji and Srinivasan 2001). This changed with the advent of artificial intelligence models that can predict protein structures with good accuracy (Jumper et al. 2021; Varadi et al. 2024). The availability of a large database of structures has prompted researchers to mold this novel source of information for the identification of structural homologs in a process similar to the Basic Local Alignment Search Tool (BLAST). Chief among these tools is FoldSeek which translates the 3D information in predicted and experimentally determined structures into 20 unique characters the authors call the 3Di alphabet (van Kempen et al. 2024). Briefly, each 3Di character represents local geometric environments derived from the relative positions and orientation of neighboring AAs residues both spatially and sequentially. Each character's definition, therefore, depends upon its immediate structural context and consequently encodes information about multiple AAs (sites in the context of phylogenetics). The advantage of using an alphabet of 20 characters is that it enables the direct use of these 3Di characters in conventional implementations of AA-based likelihood methods.

The conversion of a large dataset of 3D structures into the 3Di alphabet allows the computation of a scoring matrix like the BLOSUM scoring matrix commonly employed by MSA programs (van Kempen et al. 2024). This scoring matrix enables the quick identification of structural homologs of proteins which has been very successful in the identification of divergent orthologs. Such a scoring system also allows to compute a similarity score (fident in case of FoldTree) which can then be used to compute NJ trees as demonstrated by FoldTree (Moi et al. 2023). Furthermore, one could also calculate a substitution matrix from this BLOSUM style scoring matrix which can be directly implemented in ML approaches such as in the case of 3DiPhy (Puente-Lelievre et al. 2024). Neither of these approaches correspond to standard ML phylogenetics for AAs: FoldTree's NJ method is fast and simple but inherits all limitations of classical NJ in that it relies on the true distance between sequences being close to their observed distance, an assumption that is often violated in realistic AA datasets, and may equally be violated in structure trees if branches are either very long or very short (Atteson 1999; Elias and Lagergren 2009). NJ also does not account for site rate variation (Mihaescu et al. 2009). 3DiPhy does use a full likelihood model, which can account for these phenomena. However, its substitution matrix is derived from a BLOSUM-like alignment scoring matrix. Such matrices are constructed by counting co-occurrences of particular characters in sequence pairs, rather than inferring their contents using ML (Le and Gascuel 2008). In standard sequence phylogenetics, the BLOSUM matrix has long been superseded by empirical models which are inferred in a full phylogenetic likelihood framework, and generally result in a much better fit to empirical data (Le and Gascuel 2008).

These features of existing structural phylogenetics frameworks motivated us to infer a new substitution model using a phylogenetic ML framework. This substitution model can in theory be directly inferred from each alignment in the form of a general time reversible (GTR) model but inferring a substitution matrix for a 20-letter alphabet from a single MSA is difficult and prone to overfitting. For conventional protein models, this problem is solved by combining large numbers of protein alignments and inferring from them one substitution model that best describes all the data. Once computed, this general model, also denoted as Q, can then be used for individual protein families, which avoids overfitting using GTR. Here, we make use of AlphaFold and a recently developed protein large language model to infer a general substitution matrix for structural phylogenetics. We show that this Q-matrix is chosen as the best matrix by modelfinder across previous methods to use 3Di characters to infer ML phylogenies. Finally, we use our Q-matrix to re-infer the phylogenies of universal paralogs and photosystems to settle long-standing questions in deep evolution that previously suffered from saturation.

Results and Discussion

Estimation of the 3Di Q-Matrix

We set out to compute a general Q-matrix for structural phylogenetics. Given a large enough dataset, this is straightforward to achieve using the QMaker routine of IQ-Tree (Minh et al. 2021). We used two strategies to gather a large dataset of protein families and their predicted structures. Our first goal was to use the set of 6,653 protein families that were used to infer a Q-matrix in the initial study by Minh et al. To avoid having to predict AlphaFold structures for every sequence in this large database, we opted to use a recently developed bilingual large language model (LLM), ProtT5. This model was trained to directly translate between an AA sequence and its corresponding 3Di sequence, without having to infer an AlphaFold model (Heinzinger et al. 2024). We used this method to translate all sequences in the PFAM dataset from AA sequences to 3Di. The ProtT5 model is not perfect, as it introduces some randomness into the 3Di translation, meaning that translating to a 3Di sequence from the same input AA sequence results in a slightly different prediction (supplementary fig. S1a to c, Supplementary Material online). In addition, when comparing 3Di sequences extracted from AlphFold2 structures to the same 3Di sequence predicted with the LLM model, we found large numbers of sequences in which the AlphaFold and LLM predictions had low pairwise identities (supplementary figs. S1d to f and S2, Supplementary Material online). In order to safeguard against potential errors in estimating the substitution model using incorrectly translated 3Di sequences, we also estimated a separate Q-matrix using 3Di sequences extracted from AlphaFold predictions. We employed FoldSeek to cluster the SwissProt AlphaFold Database. These 1,660 AF clusters (hereafter AF-db) were used along with 3Di translation of the 6,653 protein families (hereafter PFAM-db) for the QMaker pipeline. Crucially, both sets of 3Di sequences were then aligned using the alignment program mafft using the 3Di scoring matrix from van Kempen et al. (2024) instead of the standard BLOSUM62 matrix used for AA alignments.

We then estimated a tree for each of 3Di MSAs in our two datasets. For this initial tree, we employed GT20 as the starting model, despite the concern of model overfitting given the unique nature of the 3Di alphabet and the lack of other models that could serve as the initial starting model. These initial trees were then used to estimate a single Q-matrix that best explains the respective sets of MSAs as described in the QMaker pipeline (Minh et al. 2021). This resulted in two Q-matrices hereafter denoted as Q.3Di.AF and Q.3Di.LLM. The two Q-matrices estimated were very similar with minor differences in exchangeability (Fig. 1c), with a Pearson’s correlation of 1 (Fig. 1d). We then checked if these matrices are preferred by IQ-Tree's modelfinder over the GTR20 or the previously published 3DiPhy model using a test set of 6,653 3Di MSAs from PFAM that were not used for estimating the Q-matrix (Table 1). Indeed, the 3DiPhy model is only preferred in 278 MSAs over 6,267 MSAs that prefer either the Q.3Di.AF or the Q.3Di.LLM model, which are practically the same (Fig. 1b to d). This increased our confidence that we had successfully captured the mechanism of change describing the mutability in the structural alphabet across a wide range of proteins. In the analyses of specific protein families that follow, IQ-Tree's modelfinder predominantly chose Q.3Di.AF over Q.3Di.LLM or GTR20 according to the corrected Akaike information criterion (AICc). Generally, we encourage future users of these matrices to always test if they are using Q.3Di.AF changes any conclusions in cases where Q.3Di.LLM is the better fit model. This is because the AF matrix is much less affected by the misprediction issues than the LLM (which we discuss further below).

Fig. 1.

Fig. 1.

a) Overview of the pipeline employed in the manuscript. Briefly, AlphaFold structures or AA protein sequences were translated into 3Di characters using FoldSeek or the bilingual ProtT5 model, respectively. These 3Di characters are aligned with MAFFT using the 3Di scoring matrix before being used to estimate the general substitution Q-matrix using QMaker which were subsequently used to estimate 3Di-ML trees using IQ-Tree. b) Lower triangular portion is a representation of the Q-matrix estimated from 1,660 AF clusters, while the upper triangular section denotes the Q-matrix estimated from 6,653 PFAM clusters translated to 3Di alphabet using the ProtT5 bilingual language model. In both cases, values higher than 2 are colored orange. c) Ratio of exchangeabilities between the Q.3Di.AF and the Q.3Di.LLM matrix. Each square represents the value (m1ijm2ij)/(m1ij+m2ij), where m1 and m2 represent Q.3Di.AF and Q.3Di.LLM, respectively. d) Pearson’s correlation between the exchangeabilities of the two matrices indicates very little differences between the two matrices.

Table 1.

Number of trees that preferred each model/Q-matrix as identified using modelfinder from IQ-Tree according to AICc, Akaike information criterion (AIC), and Bayesian information criterion (BIC)

Model AICc AIC BIC
Q.3Di.AF 3,925 2,958 3,697
Q.3Di.LLM 2,342 2,065 2,309
3DiPhy 278 322 267
GTR20 108 1,308 380
Total 6,653 6,653 6,653

Rooting the ToL Using Structural Phylogenetics

Rooting the tree of life is a particularly challenging problem owing to the lack of outgroups that can reliably root phylogenetic trees. Paralog rooting is a powerful method that uses phylogenetic trees with duplicated genes that reciprocally root each other. In most cases, the paralogs root each other along the same branch recovering an unambiguous root for the species tree containing the paralogs. However, in cases of highly divergent paralogs, the two paralogs sometimes do not agree on the same root (Fig. 2a). We tested if our new matrix can help improve trees used to root the tree of life using two universal paralogs that have been previously used for this purpose: Elongation factors and catalytic and noncatalytic subunits of the rotary ATPase. We begin with the elongation factor phylogeny. Elongation factor EF-Tu/EF-2 delivers aminoacyl-tRNAs to the A-site of the ribosome while the elongation factor EF-G/EF-1A catalyzes the translocation of the peptidyl-tRNA (Miller 1972). Both paralogs are conserved across the tree of life, making them an ideal candidate for paralog rooting (Baldauf et al. 1996; Philippe and Forterre 1999; Gouy et al. 2015). In all previous attempts to root the tree of life using EF-G and EF-Tu, the branch separating the paralogs is extremely long and potentially completely saturated, which implies that the position of the root within each paralog might be determined entirely by the substitution model and not by any synapomorphies between the paralogs. In addition, the two paralogs do not root each other consistently, increasing the uncertainty.

Fig. 2.

Fig. 2.

a) A schematic representation of paralog rooting. Three possible root positions are shown with the “true” root depicted with a green star and two other possible roots with circles. In the scenario where paralogous rooting is successful, both paralog subtrees reciprocally root each other (right). Other possible scenarios are also shown where the paralog subtrees are ambiguously rooted (left). b) AA ML tree containing 1,076 EF-Tu and EF-G homologs from eukaryotes, bacteria, and archaea. The mitochondrial and plastid-encoded copies are not included. Note that the branch separating the EF-Tu and EF-G is broken for illustration. c) 3Di structural ML tree estimated 3Di sequences and the Q.3Di.AF model from the predicted AlphaFold structures of 1,069 EF-Tu and EF-G homologs. In both cases, blue, red, and gray clades represent bacteria, archaea, and eukaryotes, respectively. Numbers in red and black indicate branch lengths and ultrafast bootstrap supports, respectively.

To test if our new matrix can help solve this problem, we first assembled a dataset of 1,076 homologs of EF-Tu and EF-G. In an AA-based ML tree, we also recover a very long branch (branch length (BL) = 3.284) between the two paralogs albeit still separating the bacteria and archaea (Fig. 2a). In line with previous phylogenies, this tree recovers different roots for the tree of life in the two paralogs: between bacteria and archaea plus eukaryotes, and between archaea and bacteria plus eukaryotes (Fig. 2b). We then extracted 3Di sequences from 1,076 AlphaFold predictions using FoldSeek (see methods) and utilized our new Q.3Di.AF Q-matrix as the substitution model, to estimate a new tree of the EF-G and EF-Tu paralogs. This recovered a phylogenetic tree with the length of the branch separating the paralogs far below 1 (0.186). Crucially, the root position is now consistent in both the paralogs and indicates a root between archaea and bacteria for life (Fig. 2c). The archaea in both paralogs remain paraphyletic, which is consistent with the two-domain tree of life.

Another universally conserved paralogous gene family used to root the tree of life is the catalytic and noncatalytic subunits of the rotary ATPase. The head group of the rotary ATPase is a hexamer consisting of two subunits, only one of which is catalytic (Fig. 3a). The bacterial and mitochondrial ATPases are called the F0F1-ATPases, and their subunits are called F1-alpha and F1-beta for the noncatalytic and catalytic subunits, respectively (Grüber et al. 2001). The archaeal ATPase is called the V-ATPase and shares a similar architecture with a noncatalytic and a catalytic subunit in its headgroup (Fig. 3a). Owing to the endosymbiotic event between archaea and bacteria at eukaryogenesis, the eukaryotes and archaea also share this ATPase which in eukaryotes is in the vacuole, where it functions to acidify lysosomes (Gogarten et al. 1989). The archaeal/eukaryotic subunits are named V1-beta and V1-alpha for the noncatalytic and catalytic subunits, respectively (Grüber et al. 2001; Cross and Müller 2004). A recent analysis on rooting the ToL using the ATPase subunits (Mahendrarajah et al. 2023) recovers a tree that separates the four major subunits with extremely long basal branches (Fig. 3b). This tree is consistent with the idea that the catalytic and noncatalytic subunits originated before the divergence of archaea and bacteria, and roots the tree of life between these two domains. The same study also identified an early transfer of the archaeal noncatalytic subunit into bacteria; however, the catalytic counterpart to this transfer was not recovered in the catalytic sub-tree suggesting multiple transfer events (Fig. 3b).

Fig. 3.

Fig. 3.

a) A schematic representation of the bacterial and archaeal ATPase highlighting the subunits under investigation. They are represented using the same colors in the phylogenetic trees. b) AA ML tree of 1,520 sequences across the ToL reproduced from Mahendrarajah et al. (2023) of the catalytic and noncatalytic subunits of bacterial, archaeal, and eukaryotic rotary ATPase. The early branching transfer from bacteria and archaea in the noncatalytic V1 clade is highlighted in white with a black outline. The corresponding clade in the V1 catalytic clade branches deep inside the archaeal sequences and is highlighted similarly. c) 3Di structural tree estimated using the Q.3Di.AF model. Sequences assigned to the early transfer from the archaeal clade to bacteria are highlighted as in (b), but now this transfer is inferred for both the catalytic and noncatalytic subunits. Numbers in red and black indicate branch lengths and ultrafast bootstrap supports, respectively. In both cases, gray clades represent eukaryotes. The green circles and orange squares indicate cyanobacterial and proteobacterial contributions in eukaryotes representing the plastid and mitochondrial ATPases.

As before, we predicted AlphaFold structures for all 1,520 sequences and extracted the 3Di sequences using FoldSeek, and calculated a 3Di (structural) ML tree with the Q.3Di.AF as the substitution model. While the tree in this case looks remarkably like the AA ML tree, the 3Di structural ML tree has significantly shorter branches (Fig. 3c). This new topology also reconfirms the root of ToL as between the archaea and bacteria. Furthermore, in the 3Di tree, the early transfer of the archaeal ATPase subunits is recovered basal in both catalytic and the noncatalytic subtrees suggesting a single early transfer from archaea to bacteria. Together with the elongation factors, our results bolster support for the two-domain tree of life with the eukaryotes branching within archaea. In both these cases, it is evident that structural phylogenetics can resolve deep phylogenies and recover consistent groupings within the paralogs despite large divergences in AA sequences.

Evolution of Photosystems RCI and RCII

The issue of saturation is not exclusive to tree-of-life problems but to all evolutionarily divergent proteins that share remote homology in sequence. The origin of oxygenic photosynthesis is another event that impacted the overall geochemistry of the planet and has been the subject of contentious debate. Photosynthesis can be classified into two major types: anoxygenic photosynthesis, which uses either reaction center II (RCII) or reaction center I (RCI), but never both together and oxygenic photosynthesis which uses both reaction centers I and II (RCI and RCII) coupled to a water splitting reaction that leads to the formation of oxygen (Hohmann-Marriott and Blankenship 2011). One set of theories suggests that anoxygenic photosynthesis evolved first and later developed into oxygenic photosynthesis (Martin et al. 2018). An alternative view favors oxygenic photosynthesis to have evolved first, with anoxygenic phototrophs having lost either RCI or RCII. One piece of evidence for the latter view is the lack of any bacterial group that harbors the anoxic versions of both RCI and RCII, which is thought to be a necessary precursor to oxygenic photosynthesis (Sánchez-Baracaldo and Cardona 2020). Until recently, members of the chloroflexota phylum have only been known to harbor anoxic RCII. This changed when a chloroflexota group, Ca. Chloroheliales, was identified that contain RC1 (Tsuji et al. 2024). This still falls short of proving that anoxic RCI and RCII have existed together in the same genome however, one possible interpretation of these data is that an ancestral Chloroflexus might have contained both, leading to differential losses in extant lineages of Chloroflexota. This would support the idea that anoxic photosynthesis may have come first if these photosystems are close relatives of the photosystems that were eventually transferred into cyanobacteria.

The phylogenetic tree based on AAs of RCI containing Chloroflexota does not place their RCI sequences as close relatives to those of cyanobacteria (4A, re-inferred for this study). But this tree suffers from extremely long branches, and we wondered whether this placement is the result of long branch attraction. We therefore set out to re-infer this tree using 3Di characters and our structural substitution matrix (Fig. 4b). This shortened all relevant internal branches to lengths well below one but yielded the same topology as the AA tree. This confirms the author's original inferences and leaves the evolution of oxygenic photosynthesis an unsolved problem for now.

Fig. 4.

Fig. 4.

a) AA ML tree of 321 RC1 protein sequences. Note that long branches are broken as indicated for illustration. b) 3Di structural ML tree of 297 3Di sequences from AlphaFold structures using the Q.3Di.AF model. Numbers in red and black indicate branch lengths and ultrafast bootstrap supports, respectively.

Our work in this manuscript and that of others (Moi et al. 2023; Puente-Lelievre et al. 2024) clearly points to the utility of structural phylogenetics in cases where structures can be predicted reliably and with one possible structure per sequence. There are several practical and conceptual caveats that come with using this method, which we will briefly elaborate on. We present these caveats in the spirit of critical optimism about the utility and impact of this new method.

Prediction Accuracy of LLMs

Structural phylogenies can only ever be as good as the predicted models that are used to derive 3Di sequences. Predicting large numbers of sequences with AlphaFold is computationally costly and potentially prohibitive for many interested users. Using bilingual Protein LLMs like ProtT5 may seem like an obvious solution because it removes the computationally expensive requirement of predicting the AF structures of a large number of protein clusters not only in the Q-matrix estimation but also for tree inference of single protein families with a lot of members. Encouragingly, the Q-matrix estimated from 3Di sequences derived from AlphaFold structures (Q.3di.AF) is very similar to the one from PFAM clusters translated using ProtT5 (Q.3Di.LLM) despite their low accuracy compared with AlphaFold predictions (Fig. 1c and d, supplementary fig. S1, Supplementary Material online). This could be due to the fact that the LLM has issues when dealing with long repeated stretches which in some cases leads to possible register shifts of structural motifs (supplementary fig. S2, Supplementary Material online). These register shifts can be dealt with during a 3Di sequence alignment using the 3Di scoring matrix, which we did not perform for our pairwise identity calculations (as the two sequences are the same length). It is also possible that prediction errors average out when inferring a Q-matrix for thousands of protein families, even if there are substantial errors in the alignments of any one family. It is clear, however, that ProtT5 translations are not reliable for inferring individual trees. We tested this by using ProtT5-derived 3Di alignments for the three protein families we investigated here. In two out of three cases, we recovered phylogenies that either were biologically improbable (for the ATPase tree, supplementary fig. S4a, Supplementary Material online) and/or erroneous with nonsensical topologies (for the RC1 tree, supplementary fig. S5a, Supplementary Material online). Most of these issues stem from the faulty prediction of 3Di sequences. While we did not observe this problem here when using AlphaFold structures, we expect similar issues when using structures that are not confidently predicted by AlphaFold. For now, reliable tree inference only seems possible using AlphaFold-generated structures and therefore comes with a significant computational overhead. Better language and structure prediction models are certain to be available in the future and they should make structural phylogenetics more widely accessible.

We also compared the performance of our substitution matrix to that of the 3DiPhy matrix on these three trees, using the AlphaFold predictions. In all three cases, 3DiPhy was not the preferred model by IQ-Tree's modelfinder. For the ATPase phylogeny, this yielded the same biologically implausible tree as when using ProtT5-derived 3Di alignments and our own substitution matrix (supplementary fig. S4b, Supplementary Material online). For RC1 and elongation factor trees, this resulted in topologies similar to that resulting using the best-fit matrix was Q.3di for all three trees (supplementary figs. S3b and S5b, Supplementary Material online, Table 2).

Table 2.

Results of the ModelFinder (IQ-Tree version 2, default settings) performed to choose the best-suited model for the structural phylogenies indicated

Model LogL AIC AICc BIC
Elongation factor Q.3Di.AF* −58,723.393 121,716.786 9,242,436.79 129,454.064
Q.3Di.LLM −59,584.159 123,438.317 9,244,158.32 131,175.595
3Diphy −60,995.253 126,260.506 9,246,980.51 133,997.783
Q.3Di.AF + F −58,703.863 121,715.726 9,405,455.73 129,521.86
Q.3Di.LLM + F −58,985.445 122,278.89 9,406,018.89 130,085.024
3Diphy + F −60,109.338 124,526.676 9,408,266.68 132,332.81
GTR20 + F −56,650.338 117,986.675 11,101,970.7 126,477.748
GTR20 + F −56,650.338 117,986.675 11,101,970.7 126,477.748
GTR20 + F + G4 −54,496.053 113,680.105 11,107,040.1 122,174.802
GTR20 + F + I −56,311.927 117,311.853 11,110,671.9 125,806.55
GTR20 + F + I + G4 −54,468.409 113,626.817 11,116,366.8 122,125.138
GTR20 + F + R2 −55,042.585 114,775.17 11,117,515.2 123,273.491
GTR20 + F + I + R2 −54,945.953 114,583.907 11,126,707.9 123,085.852
GTR20 + F + R3 −54,632.119 113,958.238 11,135,470.2 122,463.808
GTR20 + F + I + R3 −54,526.097 113,748.195 11,144,652.2 122,257.388
ATPase Q.3Di.AF + G4* −169,038.61 344,153.226 18,809,117.2 355,829.914
Q.3Di.LLM + G4 −171,848.18 349,772.35 18,814,736.4 361,449.038
3Diphy + G4 −176,533 359,141.997 18,824,106 370,818.685
Q.3Di.AF + F + G4 −168,793.62 343,701.231 19,040,313.2 355,450.946
Q.3Di.LLM + F + G4 −169,434.28 344,982.55 19,041,594.6 356,732.265
3Diphy + F + G4 −174,385.18 354,884.364 19,051,496.4 366,634.079
GTR20 + F + G4 −163,357.01 333,206.01 21,412,730 345,682.155
GTR20 + F + G4 −163,357.01 333,206.01 21,412,730 345,682.155
GTR20 + F −173,735.87 353,961.735 21,420,501.7 366,434.037
GTR20 + F + I + G4 −163,273.73 333,041.469 21,425,553.5 345,521.457
GTR20 + F + R2 −165,922.62 338,339.242 21,430,851.2 350,819.231
GTR20 + F + I −172,814.7 352,121.392 21,431,645.4 364,597.537
GTR20 + F + I + R2 −165,685.28 337,866.563 21,443,370.6 350,350.395
GTR20 + F + R3 −163,752.68 334,003.361 21,452,503.4 346,491.037
GTR20 + F + I + R3 −163,648.3 333,796.599 21,465,296.6 346,288.118
RC1 Q.3Di.AF + G4* −33,466.467 68,184.934 853,188.934 70,926.891
Q.3Di.LLM + G4 −33,586.172 68,424.344 853,428.344 71,166.301
3Diphy + G4 −34,508.202 70,268.403 855,272.403 73,010.36
Q.3Di.AF + F + G4 −33,810.473 68,910.947 902,250.947 71,736.126
Q.3Di.LLM + F + G4 −33,988.575 69,267.15 902,607.15 72,092.329
3Diphy + F + G4 −34,087.819 69,465.637 902,805.637 72,290.816
GTR20 + F + G4 −37,321.28 76,310.56 1,469,090.56 79,963.582
GTR20 + F + G4 −37,321.28 76,310.56 1,469,090.56 79,963.582
GTR20 + F −40,522.62 82,711.24 1,472,155.24 86,359.882
GTR20 + F + I + G4 −37,295.336 76,260.671 1,472,380.67 79,918.074
GTR20 + F + R2 −37,772.143 77,214.287 1,473,334.29 80,871.689
GTR20 + F + I −39,788.729 81,245.458 1,474,025.46 84,898.48
GTR20 + F + I + R2 −37,671.319 77,014.637 1,476,478.64 80,676.419
GTR20 + F + R3 −37,355.927 76,385.854 1,479,197.85 80,052.016
GTR20 + F + I + R3 −37,330.311 76,336.623 1,482,500.62 80,007.165

Models tested were Q.3Di.AF, Q.3Di.LLM, 3DiPhy, and GTR20.

Fold-Switching and Conformational Variability

Many proteins undergo conformational changes and some even switch their folds entirely as part of their functions (Chang et al. 2015). Previous analyses using AlphaFold suggest that it can sometimes predict structures in different conformations despite having a strong bias toward one dominant conformation (Chakravarty and Porter 2022; Sala et al. 2023; Wayment-Steele et al. 2024). Since this can lead to different 3Di sequences for the same protein, depending on which conformation it is predicted in, we wondered if this could lead to spurious grouping according to conformation rather than genealogical relationships on 3Di phylogenies. We examined two proteins for this purpose. One is KaiB, which is known to fold-switch as part of its catalytic cycle, involving a drastic change of a helix to a beta-sheet (Chang et al. 2015; Zhang et al. 2023). The other is the RNA Polymerase III subunit Rpc10, which undergoes a conformational change during its function in gene transcription (Girbig et al. 2021).

To test how much this issue can affect 3Di trees, we constructed a worst-case scenario for both proteins. In both cases, we modeled each sequence on the two distinct conformations using homology modeling and inferred their 3Di sequences using FoldSeek. For tree inference, we then randomly chose the 3Di sequence of one of the two possible conformations for each protein, such that approximately half of our sequences were predicted in one conformation, and the other half in the other conformation. For both KaiB and Rpc10, we found that the phylogenetic tree splits the two conformational states with long branches (Fig. 5b and d) as opposed to a 3Di structural tree which was inferred from 3Di sequences reflecting a single conformation (Fig. 5a and c). This highlights a severe limitation of structural phylogenetics where the presence of multiple predicted conformations can generate spurious branches and relationships. Here, we concocted an extreme case by forcing sequences randomly into distinct conformations. However, in cases where only a small minority of proteins within the analysis share a different conformation, these artifacts can lead to false conclusions. It is therefore very important to assess the conformational homogeneity of the predicted sequences before inferring a 3Di tree.

Fig. 5.

Fig. 5.

a) 3Di structural ML tree constructed from KaiB proteins modeled in the ground state. b) 3Di structural ML tree constructed from approximately 50% of the KaiB proteins modeled in the ground state (blue) and the other 50% modeled in the fold-switched state (green). c) 3Di structural ML tree constructed from RPC10 proteins modeled in the IN conformation. d) 3Di structural ML tree constructed from approximately 50% of the RPC10 proteins modeled in the OUT conformation (blue) and the other 50% modeled in the IN conformation (green). In both cases, the distinct conformations form monophyletic groups in contrast to their placements in (a) and (c), respectively.

Site Independence in Structural Alignments

One of the main assumptions of ML is site independence, which allows the likelihood to be computed independently for all sites in the alignment (Liò and Goldman 1998). It has been obvious for a long time that this is not a realistic assumption. Epistasis between AAs is a well-demonstrated phenomenon and quite extensive among proteins (Starr and Thornton 2016). This violates the site-independence assumption of ML phylogenetics, even though it has been argued that increasing the number of sites normally associated with a protein sequence or increasing the number of proteins used for a concatenated alignment averages out the signal in most cases (Starr and Thornton 2016; Magee et al. 2021). In the case of the structural phylogenetics and 3Di alphabet, however, this assumption is explicitly violated since each letter corresponds to at least five other AA positions in 3D space. It is for example not clear to us that it is even possible for a single substitution to occur at the level of 3Di characters, because of the structural dependence between sites. In a sense, structural phylogenetics makes the ugly truth of model violation explicit in its alphabet. Whether or not this approach becomes widely accepted in evolutionary biology will depend on investigating the consequences of this violation, which is beyond the scope of this manuscript.

Information Loss

The 3Di alphabet compresses information that would be present in AAs. This is the very reason for its utility in deep phylogenetics, because it overcomes the saturation problem. But it also makes evolution on short time-scales harder to resolve using these characters and relationships at the very tips of 3Di trees are probably much less reliable than in an AA or DNA tree (Mutti et al. 2024). A potential solution is to use partitioned models, in which a tree is inferred from both 3Di and AA alignments simultaneously, using different substitution models for the partitions (Puente-Lelievre et al. 2024). So far, this has been done using edge-linked models, which allow the partitions to have different evolutionary rates (slow for 3Di, fast for AAs) (Puente-Lelievre et al. 2024), with one set of relative branch lengths for both partitions. It may in the future be productive to investigate using edge-unlinked models, where the 3Di and AA partitions would have their own set of branch lengths (Lopez et al. 2002). This may be necessary because 3Di is not a simple recoding of AAs, but a fundamentally different type of character. Branches that are long on 3Di trees could in principle be short on AA trees and vice versa. This may for example happen when the gain of a novel tertiary structure element through an insertion changes the 3Di state at many sites that are in contact with this new tertiary structure element, even if the AA states do not change at all at the same sites. This kind of heterotachy is a serious violation of the assumptions of ML phylogenetics and can make the method statistically inconsistent (Kolaczkowski and Thornton 2004). Edge-unlinked models can account for this problem and are implemented in standard phylogenetics software such as IQ-Tree (Chernomor et al. 2016), but they present a difficult optimization problem and were not tested here. Whether such complex models are justified statistically will depend on the protein family under investigation. Another question is the size of the alphabet. 3Di uses 20 characters because this allows simple integration with existing phylogenetic software. It is not yet clear whether this is even close to the optimal number of characters for structural phylogenetics. Larger alphabets could perhaps retain more short-term information. They would, however, make the inference of substitution matrices much harder.

Structural Phylogenetics and the Future of Deep History

Our work complements and builds on other recent tools that utilize the 3Di alphabet for structural phylogenetics (Moi et al. 2023; Puente-Lelievre et al. 2024). Our structural Q-matrices should make it much easier to infer structure-based trees for anyone familiar with ML phylogenetics. Newly developed online tools for the generation of 3Di alignments should further lower the bar for adoption (Gilchrist et al. 2024). As with every new method, it is difficult to know exactly what impact structural phylogenetics will have. For now, we see its main utility in solving difficult rooting problems involving distant outgroups that AA phylogenies cannot solve with any degree of confidence. This will help polarize the direction of evolutionary change in the emergence of many important functions. Better resolved deep protein phylogenies will also improve our reconstructions of the gene content of ancient organisms (the LUCA, the Last Eukaryotic Common Ancestor, and the Last Archaeal Common Ancestor, for example).

For now, these methods will not be useful for ancestral sequence reconstruction, because 3Di sequences cannot be back translated into a unique AA sequence (Heinzinger et al. 2024). Even though our matrix allows us to infer 3Di sequences at internal nodes of structural phylogenies, it is at present not possible to then turn these reconstructed 3Di sequences into resurrected proteins composed of AAs. It may, however, be possible to restrict a set of plausible AA reconstructions at one particular node on an AA phylogeny to a subset that agrees with the reconstructed ancestral 3Di sequence at the corresponding node on a structural phylogeny.

The true impact of viewing the past through the glacial change in the structure of proteins will only emerge when this method is robustly tested and becomes widely adopted in evolutionary biology. We hope the matrix inferred here will be a first step in this process.

Methods

Datasets for QMaker

The SwissProt AlphaFold database (https://alphafold.ebi.ac.uk/download) was downloaded and then clustered with FoldSeek (https://github.com/steineggerlab/foldseek) easy-cluster with default settings and a coverage of 80%. This yielded 1,660 clusters which contained at least 50 members and a maximum of 2,500 members. Databases of structures in Protein Data Bank (PDB) format were then created and 3Di sequences were subsequently extracted from these 1,660 clusters using FoldSeek as previously described. The PFAM sets were taken from Minh et al. (2021) which contained 6,655 protein families used for training the Q-matrix and a further 6,653 families were used for testing. In the case of PFAM families, the AA FASTA files were directly translated to the 3Di alphabet using the scripts provided by Heinzinger et al. (2024) (https://github.com/mheinzinger/ProstT5).

Q-Matrix Estimation

Both the AF-db and PFAM-db sets of 3Di sequences were aligned using ginsi method within Mafft (v7.515) and the 3Di scoring matrix from FoldSeek using the −aamatrix flag implemented within mafft. The 3Di MSAs thus generated were then used in the QMaker routine as described in Minh et al. (2021) (http://www.iqtree.org/doc/Estimating-amino-acid-substitution-models). Briefly, for each MSA, the best-fit substitution model was initialized with GTR20 along with the best fit model for rate-heterogeneity among sites using ModelFinder (Kalyaanamoorthy et al. 2017). In the next step, we estimate a joint reversible Q-matrix for all the 3Di MSAs as described.

Individual Protein/3Di Sequences and Phylogenetic Tree Reconstructions

Elongation factors and Reaction Center I homologs were identified using BLAST against the NCBI nonredundant (nr) database, and then filtered with a minimum similarity threshold of 50% and an e-value cutoff of 1E−5. For the ATPase, phylogeny was exactly reproduced from Mahendrarajah et al. (2023), and the same sequences were used for the 3Di sequences. The AA sequences were aligned using linsi and then subsequently trimmed using TrimAl (v1.4) (Capella-Gutiérrez et al. 2009) with the -automated1 setting. Trimmed AA alignments were then used for ML tree estimation using IQ-Tree with the best-fit model suggested by ModelFinder. 3Di sequences for individual protein trees were extracted from PDB files from individual AlphaFold (v2.2.0) predictions. The best-ranked AlphaFold models were used to create a database using FoldSeek which allowed us to extract 3Di sequences from the PDB structures. For 3Di sequences translated from ProtT5, the model was queried as described in Heinzinger et al. (2024) using AA sequences as input. All 3Di sequences were aligned with Mafft (ginsi) using the −aamatrix option specifying the 3Di scoring matrix provided by van Kempen et al., as part of FoldSeek. 3Di MSAs were then used to estimate the structural ML tree as described above. For individual 3Di tree reconstructions, IQ-Tree (v2.3.0) was used to identify the best-fit model (Q.3Di.AF, Q.3Di.LLM or GTR20) according to AICc, along with rate-heterogeneity using ModelFinder. Both AA and 3Di trees were estimated with 10,000 Ultrafast bootstraps (-bb) and 10,000 iterations for SH-test (-alrt).

Homology Modeling

For the KaiB and RPC10 proteins, homologs were identified via BLAST as described above. Then PDB structures or AlphaFold structures of the two conformations in question were used as a template in SWISS-MODEL (Waterhouse et al. 2018). KaiB was modeled using the PDB structure 2QKE in the ground state and 5JYT in the fold-switched state from Thermosynechococcus elongatus. The RPC10 was homology modeled on the PdB structure 7AE1 in the OUT conformation and 7AE3 in the IN conformation as described in Girbig et al. (2021). 3Di sequences were extracted from both sets of states/conformations and then randomly sampled to generate a set composed approximately 50% of 3Di sequences from PDB of KaiB and RPC10 in one of the two states/conformation. ML trees were then estimated using these protein sequences as described above.

Supplementary Material

msaf124_Supplementary_Data

Acknowledgments

We would like to thank Mathias Girbrig for suggesting RPC10 for studying the impact of conformational heterogeneity.

Contributor Information

Sriram G Garg, Evolutionary Biochemistry Group, Max Planck Institute for Terrestrial Microbiology, Marburg 35043, Germany.

Georg K A Hochberg, Evolutionary Biochemistry Group, Max Planck Institute for Terrestrial Microbiology, Marburg 35043, Germany; Center for Synthetic Microbiology (SYNMIKRO), Philipps-University Marburikg, Marburg 35043, Germany; Department of Biology, Philipps-University Marburg, Marburg 35043, Germany.

Supplementary Material

Supplementary material is available at Molecular Biology and Evolution online.

Author Contributions

S.G.G. and G.K.A.H. conceptualized, designed, and wrote the manuscript. S.G.G. performed the computations.

Funding

The study was funded by the Human Frontiers Science Program Grant (RGP0028) awarded to G.K.A.H.

Data Availability

All datasets, trees, and alignments are available at https://doi.org/10.17617/3.1MJJBH.

References

  1. Atteson  K. The performance of neighbor-joining algorithms of phylogeny reconstruction. Algorithmica. 1999:25:251–278. 10.1007/PL00008277. [DOI] [Google Scholar]
  2. Balaji  S, Srinivasan  N. Use of a database of structural alignments and phylogenetic trees in investigating the relationship between sequence and structural variability among homologous proteins. Protein Eng Des Select. 2001:14(4):219–226. 10.1093/protein/14.4.219. [DOI] [PubMed] [Google Scholar]
  3. Balaji  S, Sujatha  S, Kumar  SSC, Srinivasan  N. PALI—a database of Phylogeny and ALIgnment of homologous protein structures. Nucleic Acids Res. 2001:29(1):61–65. 10.1093/nar/29.1.61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Baldauf  SL, Palmer  JD, Doolittle  WF. The root of the universal tree and the origin of eukaryotes based on elongation factor phylogeny. Proc Natl Acad Sci U S A.  1996:93(15):7749–7754. 10.1073/pnas.93.15.7749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Brown  JR, Doolittle  WF. Root of the universal tree of life based on ancient aminoacyl-tRNA synthetase gene duplications. Proc Natl Acad Sci U S A.  1995:92(7):2441–2445. 10.1073/pnas.92.7.2441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Capella-Gutiérrez  S, Silla-Martínez  JM, Gabaldón  T. Trimal: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009:25(15):1972–1973. 10.1093/bioinformatics/btp348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Chakravarty  D, Porter  LL. AlphaFold2 fails to predict protein fold switching. Protein Sci. 2022:31(6):e4353. 10.1002/pro.4353. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Chang  Y-G, Cohen  SE, Phong  C, Myers  WK, Kim  Y-I, Tseng  R, Lin  J, Zhang  L, Boyd  JS, Lee  Y, et al.  A protein fold switch joins the circadian oscillator to clock output in cyanobacteria. Science. 2015:349(6245):324–328. 10.1126/science.1260031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Chernomor  O, von Haeseler  A, Minh  BQ. Terrace aware data structure for phylogenomic inference from supermatrices. Syst Biol. 2016:65(6):997–1008. 10.1093/sysbio/syw037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cross  RL, Müller  V. The evolution of A-, F-, and V-type ATP synthases and ATPases: reversals in function and changes in the H+/ATP coupling ratio. FEBS Lett. 2004:576(1-2):1–4. 10.1016/j.febslet.2004.08.065. [DOI] [PubMed] [Google Scholar]
  11. Elias  I, Lagergren  J. Fast neighbor joining. Theor Comput Sci.  2009:410(21-23):1993–2000. 10.1016/j.tcs.2008.12.040. [DOI] [Google Scholar]
  12. Farris  JS. Methods for computing Wagner trees. Syst Biol. 1970:19(1):83–92. 10.1093/sysbio/19.1.83. [DOI] [Google Scholar]
  13. Felsenstein  J. Evolutionary trees from DNA sequences: a maximum likelihood approach. J Mol Evol. 1981:17(6):368–376. 10.1007/BF01734359. [DOI] [PubMed] [Google Scholar]
  14. Felsenstein  J.  Inferring phylogenies. Sunderland (MA): Sinauer Associates; 2003. [Google Scholar]
  15. Fitch  WM. Toward defining the course of evolution: minimum change for a specific tree topology. Syst Zoöl. 1971:20(4):406. 10.2307/2412116. [DOI] [Google Scholar]
  16. Gilchrist  CLM, Mirdita  M, Steinegger  M. Multiple protein structure alignment at scale with FoldMason. bioRxiv 606130. 10.1101/2024.08.01.606130, 01 August 2024, preprint: not peer reviewed. [DOI]
  17. Girbig  M, Misiaszek  AD, Vorländer  MK, Lafita  A, Grötsch  H, Baudin  F, Bateman  A, Müller  CW. Cryo-EM structures of human RNA polymerase III in its unbound and transcribing states. Nat Struct Mol Biol. 2021:28(2):210–219. 10.1038/s41594-020-00555-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Gogarten  JP, Kibak  H, Dittrich  P, Taiz  L, Bowman  EJ, Bowman  BJ, Manolson  MF, Poole  RJ, Date  T, Oshima  T, et al.  Evolution of the vacuolar H+-ATPase: implications for the origin of eukaryotes. Proc Natl Acad Sci U S A.  1989:86(17):6661–6665. 10.1073/pnas.86.17.6661. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Gouy  R, Baurain  D, Philippe  H. Rooting the tree of life: the phylogenetic jury is still out. Philos Trans R Soc B Biol Sci. 2015:370(1678):20140329. 10.1098/rstb.2014.0329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Grüber  G, Wieczorek  H, Harvey  WR, Müller  V. Structure–function relationships of A-, F- and V-ATPases. J Exp Biol.  2001:204(15):2597–2605.. 10.1242/jeb.204.15.2597. [DOI] [PubMed] [Google Scholar]
  21. Heinzinger  M, Weissenow  K, Sanchez  JG, Henkel  A, Steinegger  M, Rost  B. Bilingual language model for protein sequence and structure. NAR Genom Bioinform. 2024:6(4):lqae150. 10.1093/nargab/lqae150. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Hohmann-Marriott  MF, Blankenship  RE. Evolution of photosynthesis. Ann Rev Plant Biol. 2011:62(1):515–548. 10.1146/annurev-arplant-042110-103811. [DOI] [PubMed] [Google Scholar]
  23. Iwabe  N, Kuma  K, Hasegawa  M, Osawa  S, Miyata  T. Evolutionary relationship of archaebacteria, eubacteria, and eukaryotes inferred from phylogenetic trees of duplicated genes. Proc Natl Acad Sci U S A.  1989:86(23):9355–9359. 10.1073/pnas.86.23.9355. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Johnson  MS, Šali  A, Blundell  TL. Phylogenetic relationships from three-dimensional protein structures. Methods Enzym. 1990a:183:670–690. 10.1016/0076-6879(90)83044-A. [DOI] [PubMed] [Google Scholar]
  25. Johnson  MS, Sutcliffe  MJ, Blundell  TL. Molecular anatomy: phyletic relationships derived from three-dimensional structures of proteins. J Mol Evol. 1990b:30(1):43–59. 10.1007/BF02102452. [DOI] [PubMed] [Google Scholar]
  26. Jumper  J, Evans  R, Pritzel  A, Green  T, Figurnov  M, Ronneberger  O, Tunyasuvunakool  K, Bates  R, Žídek  A, Potapenko  A, et al.  Highly accurate protein structure prediction with AlphaFold. Nature. 2021:596(7873):583–589. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Kalyaanamoorthy  S, Minh  BQ, Wong  TKF, von Haeseler  A, Jermiin  LS. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017:14(6):587–589. 10.1038/nmeth.4285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kolaczkowski  B, Thornton  JW. Performance of maximum parsimony and likelihood phylogenetics when evolution is heterogeneous. Nature. 2004:431:980–984. 10.1038/nature02917. [DOI] [PubMed] [Google Scholar]
  29. Le  SQ, Gascuel  O. An improved general amino acid replacement matrix. Mol Biol Evol.  2008:25(7):1307–1320. 10.1093/molbev/msn067. [DOI] [PubMed] [Google Scholar]
  30. Liò  P, Goldman  N. Models of molecular evolution and phylogeny. Genome Res. 1998:8(12):1233–1244. 10.1101/gr.8.12.1233. [DOI] [PubMed] [Google Scholar]
  31. Lopez  P, Casane  D, Philippe  H. Heterotachy, an important process of protein evolution. Mol Biol Evol. 2002:19(1):1–7. 10.1093/oxfordjournals.molbev.a003973. [DOI] [PubMed] [Google Scholar]
  32. Magee  AF, Hilton  SK, DeWitt  WS. Robustness of phylogenetic inference to model misspecification caused by pairwise epistasis. Mol Biol Evol. 2021:38(10):4603–4615. 10.1093/molbev/msab163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Mahendrarajah  TA, Moody  ERR, Schrempf  D, Szánthó  LL, Dombrowski  N, Davín  AA, Pisani  D, Donoghue  PCJ, Szöllősi  GJ, Williams  TA, et al.  ATP synthase evolution on a cross-braced dated tree of life. Nat Commun. 2023:14(1):7456. 10.1038/s41467-023-42924-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Martin  WF, Bryant  DA, Beatty  JT. A physiological perspective on the origin and evolution of photosynthesis. FEMS Microbiol Rev.  2018:42(2):205–231. 10.1093/femsre/fux056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Mau  B, Newton  MA. Phylogenetic inference for binary data on dendograms using Markov chain Monte Carlo. J Comput Graph Stat. 1997:6(1):122–131. 10.1080/10618600.1997.10474731. [DOI] [Google Scholar]
  36. Mihaescu  R, Levy  D, Pachter  L. Why neighbor-joining works. Algorithmica. 2009:54(1):1–24. 10.1007/s00453-007-9116-4. [DOI] [Google Scholar]
  37. Miller  DL. Elongation factors EF Tu and EF G interact at related sites on ribosomes. Proc Natl Acad Sci U S A.  1972:69(3):752–755. 10.1073/pnas.69.3.752. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Minh  BQ, Dang  CC, Vinh  LS, Lanfear  R. QMaker: fast and accurate method to estimate empirical models of protein evolution. Syst Biol.  2021:70(5):1046–1060 .   10.1093/sysbio/syab010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Moi  D, Bernard  C, Steinegger  M, Nevers  Y, Langleib  M, Dessimoz  C. Structural phylogenetics unravels the evolutionary diversification of communication systems in gram-positive bacteria and their viruses. bioRxiv 558401. 10.1101/2023.09.19.558401, 23 September 2023, preprint: not peer reviewed. [DOI]
  40. Mutti  G, Ocaña-Pallarés  E, Gabaldón  T. Newly developed structure-based methods do not outperform standard sequence-based methods for large-scale phylogenomics. bioRxiv 606352. 10.1101/2024.08.02.606352, 6 August 2024, preprint: not peer reviewed. [DOI]
  41. Philippe  H, Brinkmann  H, Lavrov  DV, Littlewood  DTJ, Manuel  M, Wörheide  G, Baurain  D. Resolving difficult phylogenetic questions: why more sequences are not enough. PLoS Biol. 2011:9(3):e1000602. 10.1371/journal.pbio.1000602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Philippe  H, Forterre  P. The rooting of the universal tree of life is not reliable. J Mol Evol. 1999:49(4):509–523. 10.1007/PL00006573. [DOI] [PubMed] [Google Scholar]
  43. Posada  D, Crandall  KA. Felsenstein phylogenetic likelihood. J Mol Evol. 2021:89(3):134–145. 10.1007/s00239-020-09982-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Puente-Lelievre  C, Malik  AJ, Douglas  J, Ascher  D, Baker  M, Allison  J, Poole  A, Lundin  D, Fullmer  M, Bouckert  R, et al. Tertiary-interaction characters enable fast, model-based structural phylogenetics beyond the twilight zone. bioRxiv 571181. 10.1101/2023.12.12.571181, 13 December 2023, preprint: not peer reviewed. [DOI]
  45. Rannala  B, Yang  Z. Probability distribution of molecular evolutionary trees: a new method of phylogenetic inference. J Mol Evol. 1996:43(3):304–311. 10.1007/BF02338839. [DOI] [PubMed] [Google Scholar]
  46. Saitou  N, Nei  M. The neighbor-joining method: a new method for reconstructing phylogenetic trees. Mol Biol Evol. 1987:4(4):406–425. 10.1093/oxfordjournals.molbev.a040454. [DOI] [PubMed] [Google Scholar]
  47. Sala  D, Engelberger  F, Mchaourab  HS, Meiler  J. Modeling conformational states of proteins with AlphaFold. Curr Opin Struct Biol. 2023:81:102645. 10.1016/j.sbi.2023.102645. [DOI] [PubMed] [Google Scholar]
  48. Sánchez-Baracaldo  P, Cardona  T. On the origin of oxygenic photosynthesis and cyanobacteria. New Phytol. 2020:225(4):1440–1446. 10.1111/nph.16249. [DOI] [PubMed] [Google Scholar]
  49. Starr  TN, Thornton  JW. Epistasis in protein evolution. Protein Sci. 2016:25(7):1204–1218. 10.1002/pro.2897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Tsuji  JM, Shaw  NA, Nagashima  S, Venkiteswaran  JJ, Schiff  SL, Watanabe  T, Fukui  M, Hanada  S, Tank  M, Neufeld  JD. Anoxygenic phototroph of the Chloroflexota uses a type I reaction centre. Nature. 2024:627(8005):915–922. 10.1038/s41586-024-07180-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. van Kempen  M, Kim  SS, Tumescheit  C, Mirdita  M, Lee  J, Gilchrist  CLM, Söding  J, Steinegger  M. Fast and accurate protein structure search with FoldSeek. Nat Biotechnol. 2024:42(2):243–246. 10.1038/s41587-023-01773-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Varadi  M, Bertoni  D, Magana  P, Paramval  U, Pidruchna  I, Radhakrishnan  M, Tsenkov  M, Nair  S, Mirdita  M, Yeo  J, et al.  AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024:52(D1):D368–D375. 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Waterhouse  A, Bertoni  M, Bienert  S, Studer  G, Tauriello  G, Gumienny  R, Heer  FT, de Beer  TAP, Rempfer  C, Bordoli  L, et al.  SWISS-MODEL: homology modelling of protein structures and complexes. Nucleic Acids Res. 2018:46(W1):W296–W303. 10.1093/nar/gky427. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Wayment-Steele  HK, Ojoawo  A, Otten  R, Apitz  JM, Pitsawong  W, Hömberger  M, Ovchinnikov  S, Colwell  L, Kern  D. Predicting multiple conformations via sequence clustering and AlphaFold2. Nature. 2024;625:832–839. 10.1038/s41586-023-06832-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Whelan  S, Liò  P, Goldman  N. Molecular phylogenetics: state-of-the-art methods for looking into the past. Trends Genet. 2001:17(5):262–272. 10.1016/S0168-9525(01)02272-7. [DOI] [PubMed] [Google Scholar]
  56. Zhang  N, Guan  W, Cui  S, Ai  N. Crowded environments tune the fold-switching in metamorphic proteins. Commun Chem.  2023:6(1):117. 10.1038/s42004-023-00909-2. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

msaf124_Supplementary_Data

Data Availability Statement

All datasets, trees, and alignments are available at https://doi.org/10.17617/3.1MJJBH.


Articles from Molecular Biology and Evolution are provided here courtesy of Oxford University Press

RESOURCES