Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Jan 8.
Published in final edited form as: Methods Mol Biol. 2008;413:125–146. doi: 10.1007/978-1-59745-574-9_5

Algorithms for Multiple Protein Structure Alignment and Structure-Derived Multiple Sequence Alignment

Maxim Shatsky, Ruth Nussinov, Haim J Wolfson
PMCID: PMC10773980  NIHMSID: NIHMS1948696  PMID: 18075164

Summary

Primary amino acid content and the geometry of the folded protein 3D structure are major parameters of protein function. During the course of evolution the protein 3D structure is more preserved than its primary sequence. Thus, analysis of protein structures is expected to lead to a deep insight into protein function. Recognition of a structural core common to a set of protein structures serves as a basic tool for the studies of protein evolution and classification, analysis of similar structural motifs and functional binding sites, and for homology modeling and threading.

In this chapter, we discuss several biologically related computational aspects of the multiple structure alignment and propose a method that provides solutions to these problems. Finally, we address the problem of structure-based multiple sequence alignment and propose an optimization method that unifies primary sequence and 3D structure information.

Keywords: Multiple structure alignment, partial alignment, structure base sequence alignment, structure-sequence conservation

1. Introduction

The increasing number of determined protein structures opens new horizons for studies of protein function. There are numerous examples of similar functioning proteins, for example, isomerases, cytokines, myoglobins, immunoglobulins, and transferases, with similar 3D structure but less than 25% sequence identity. Therefore, in order to study relationships between such proteins sequence analysis alone is not sufficient. While methods for sequence analysis have significantly advanced in the past years, methods for structural analysis are still at an earlier, exploratory stage. Here, we address one of the most basic structure-related problems, the problem of multiple structure alignment and structure-derived multiple sequence alignment.

A number of methods have been proposed to solve the problem of structural alignment between a pair of proteins, for example, VAST (1), Geometric Hashing (GH) (2), CE (3), DALI (4), FlexProt (5), and others (6). Obviously, multiple structure alignment can provide much more information. Recognition of a structural core common to a set of protein structures has many applications in the studies of protein evolution and classification (4,7), analysis of similar functional binding sites and protein–protein interfaces (810), homology modeling and threading (11,12), and so on. However, despite this need, the multiple structure alignment problem has not been extensively studied, and, consequently, there are very few available methods that solve this task.

Let us formulate a list of some principal requirements for a multiple structure alignment method. Subjects like protein structure representation and structural similarity scoring functions are extensively discussed in several reviews (6,13,14); therefore, our aim is to emphasize the topics that are specifically related to the multiple alignment problem and topics that were paid less attention in previous reviews and they are the following:

  • Partial alignment.

  • Subset alignment.

  • Flexible alignment.

  • Sequential alignment.

  • Sequence order independent alignment.

  • Time efficiency.

1.1. Partial Alignment

There might be only a sub-structure (motif and domain) that is similar between a set of molecules, for example, in a set of multi-domain proteins having one or several common domains. In case of protein domain swapping (15), a protein chain of a monomer can be partially aligned with two protein chains of a dimer. Another example is alignment between multi-protein complexes having some structurally similar combination of molecules. Thus, a detection of all common motifs, domains, or multi-protein combinations may be required for a multiple structure alignment method. We consider a local alignment as a special case of partial alignment. For example, a partial alignment may consist of several locally matched structural elements that can be aligned under the same Euclidean transformation.

1.2. Subset Alignment

An important aspect of any multiple, sequence, or structure alignment is a detection of a subset of molecules that are more similar than the whole input set. For example, consider an input set of 10 proteins from one family and 5 proteins from another family. Assume that the proteins in each family are structurally similar, but there is little similarity between any two proteins from the first and second family. A multiple alignment between these 15 molecules would probably detect at most one common secondary structure element. Therefore, it is very important for a multiple alignment method to be able to automatically distinguish between two such subsets.

To demonstrate the significance of the partial and subset alignment ability, consider a schematic example in Fig. 1A. Three proteins share a common small pattern, and each pair of the proteins share additional, larger, patterns. The desired goal of a multiple structure alignment method is to detect all four patterns. An additional example with real protein structures is given below in Subheading 4.3. It should be clear that the number of all possible solutions that may be also biologically meaningful could be exponential in the number of input molecules. For example, consider proteins that contain a large number of α-helices. Each pair of α-helices could be structurally matched (at least partially, if they are different in their lengths). Any multiple combination of α-helices from different proteins results in some multiple alignment. Obviously, the number of such multiple alignments is exponential. Therefore, even if an algorithm is capable of detecting all such combinations, it is not practical to report them.

Fig. 1.

Fig. 1.

(A) A schematic example of three proteins that share a common pattern X. Applying a pairwise alignment method that detects the most similar common pattern will result in pattern A for proteins P1 and P2, pattern B for P1 and P3, and pattern D for P2 and P3. Therefore, no common pattern can be derived from the patterns A, B, and D. One possible solution is to store two (or more) high scoring solutions for each pairwise comparison. However, in this case, the number of iterations to compare all pairwise results to detect the best combination of multiple alignments becomes exponential. (B) The MultiProt method aims to compute a large number of different local multiple alignments. The depicted alignments are detected while selecting each structure as a pivot, for example, patterns X, A, and B are detected when protein P1 is selected as a pivot. Finally, pattern X will appear in the multiple alignment of three molecules. Patterns A, B, and D will appear in the set of alignments consisting of two molecules.

1.3. Flexible Alignment

Proteins are flexible molecules, which may appear in different conformations. Hinge motion may divide a protein structure into several almost rigid parts, which move one relative to the other. In such a case, a rigid structural alignment may detect only separate partially matched regions. Therefore, only a manual inspection of the final solution may reveal the whole picture as in the example given below in Subheading 4.3. Obviously, a method that is able to automatically detect a multiple flexible alignment is more beneficial.

1.4. Sequential and Sequence Order Independent Alignment

Sequence alignment methods naturally produce alignments that follow the protein sequence order, that is, aligned amino acids indices are always in increasing order. However, protein evolution imposes less constraints on the sequential order than on the structural properties. Consequently, proteins may have a similar function but topologically different 3D structure. One such example is the calcium/phospholipid-binding domain (CaLB, C2 domain) which consists of a β-sandwich (eight strands in two sheets). The proteins synaptogamin I (pdb:1rsy) and cytosolic phospholipase A2 (pdb:1rlw) both have the C2 domain but with different topology (16). The phenomenon of similar secondary structure arrangements but with different topology has been the case of a previous study [(17) and references therein]. Therefore, in order to fully discover the structural similarities, sequence order independent alignments should be considered. In Subheading 4.3., we consider one such example.

1.5. Time Efficiency

Optimal pairwise structural alignment can be solved in polynomial time; however, it is still computationally expensive and currently not practical for an implementation (18). Approximation techniques can significantly reduce time complexity with relatively small degradation in solution accuracy (19,20). However, even for three structures the multiple alignment problem is NP-hard (21) (i.e., practically solvable only in exponential time). While the worst case scenario may be computationally infeasible for the detection of an exact solution, considering specific geometrical properties of the protein molecules can, in practice, significantly reduce the computational cost. Examples of such properties, which are far from resembling a random point distribution, include sequentiality of the protein backbone, secondary structure element composition, and protein compactness. Therefore, “smart” heuristic methods that utilize such properties, in practice, may give results that are sufficient for a biological research. An excellent example of a heuristic method for multiple sequence alignment is MUSCLE (22). Still further research, theoretical and practical, is required for the multiple structural alignment problem.

Below we briefly review available methods for the multiple structure alignment task and try to correlate them with the list of requirements defined above.

A center-star approach is one of the efficient ways to compute a multiple sequence alignment. Analogously, it can be applied for multiple structure alignment. A center structure is selected which is most similar to the rest of the molecules. Then, iteratively, all other structures are joined into a multiple alignment based on their pairwise alignments with the center structure (11,23). Alternatively, one can apply a tree-progressive approach, where a multiple alignment is created according to some distance tree (24,25). Therefore, a tree-progressive alignment first aligns similar proteins, then proceeds to more distant relationships. An advantage of such an approach is its ability to detect sub-set alignments of structurally different families.

In order to tackle the flexible alignment problem, the POSA method (26) utilizes a partial order graph representation of multiple alignments. The advantage of this method is in automatic detection of larger structurally similar regions that cannot be detected without considering hinge motions of the protein backbone. The multiple alignments are computed from pairwise alignments using a tree-progressive approach.

The center-star and the tree-progressive approaches are essentially based on some pairwise alignment method that is iteratively applied for the construction of a multiple alignment. Therefore, such technique is less suitable for detection of small structurally similar motifs as at each stage of the iterative alignment only one, the best, solution is selected. Figure 1A shows a simple example where a straightforward application of a pairwise alignment method will fail to recognize a pattern common to more than two sequences/structures.

The MALECON (27) and MUSTANG (28) methods aim to avoid the shortcomings of the iterative pairwise approaches using two different techniques. MALECON (27) considers all possible combinations of the input molecules. When the number of input proteins is large, such a combinatorial approach becomes exponential; therefore, the method considers at least all possible protein triplets while other proteins are progressively added to the aligned triplets. MUSTANG (28) uses a tree-progressive approach; however, it reduces possible artifacts of the iterative multiple alignment by applying a refinement of the residue correspondence scores based on transitive relations between the aligned structure pairs. Consequently, the advantage of both methods is in detection of subset alignments. However, because only one solution is considered for any given combination of proteins, some smaller local alignments can be missed. Both methods produce sequential alignments.

The MUSTA algorithm (29,30) computes a common geometric core that appears simultaneously in all the input molecules, thus avoiding the shortcomings of the iterative pairwise approaches. The method applies the Geometric Hashing technique (27) which allows detection of the sequence order independent alignments. This technique was successfully applied in a number of pairwise structure alignment methods (32,33). Because the method requires that all input molecules participate in the multiple structural alignment, the drawback of this method is inability to distinguish outliers. It is sufficient that one structure is very distinct from the others to result in an empty alignment. Consequently, this method cannot detect subset alignments. Second, its efficiency limits practical application for only 10–15 molecules.

Another approach, SPratt2 (34), aims to detect small common, local structural motifs of size 3–20 amino acids. The method describes each residue as a short string of its spatial neighbors. Then, an efficient sequence pattern discovery technique is applied to detect sets of residues with common environmental descriptors. The computed alignments are sequential. The method is efficient and allows subset alignments.

The MASS (16,35) method utilizes the secondary structure information (SSE) to reduce the computational cost of initial common core detection. Therefore, it requires that at least two pairs of SSE be multiply aligned. The method is capable of detecting partial, subset, sequential, and non-sequential alignments.

In order to produce structure-based multiple sequence alignment, the recently developed method 3DCoffee (36) incorporates spatial weights into the multiple sequence alignment method TCoffee. The spatial weight of an amino acid pair is defined as a positive large constant number, when this pair is structurally aligned according to some pairwise structural alignment method. Therefore, the method does not distinguish between amino acids that are structurally aligned at different distances. Because the method applies information only from the best pairwise structure alignment, the weighting decision may be inaccurate for the multiple alignment problem.

Here, we discuss a method that aims to solve the multiple structural alignment problem with the support of the most above-defined requirements. One of the advantages of our method is its ability of subset and partial alignment. This is illustrated in Fig. 1A and B. Consider the set of proteins from Fig. 1A. The goal of our method is to detect local multiple alignments of all four patterns. This is achieved by performing all possible local multiple alignments of ungapped fragments. The final solutions are constructed from these locally aligned multiple fragments. This makes our approach different from most existing methods, which generally derive a multiple alignment from the high scoring pairwise superpositions. If a pattern appears more than once in some protein, our method recognizes only one combination of this pattern from all possible appearances; however, all sets of possible combinations are reported by the program. Our method is extremely efficient and is suitable for simultaneous comparison of up to tens of proteins.

Below, we start with the brief description of the multiple structural alignment method, MultiProt (37). Then, we discuss the problem of structure-based multiple sequence alignment and propose an optimization method that unifies primary sequence and 3D structure information. Finally, we present some experimental results that include comparison with the HOMSTRAD (38) benchmark of manually curated multiple structure-based sequence alignments. We argue that our automated approach produces slightly more accurate alignments.

2. MultiProt—an Algorithm for Multiple Protein Structure Alignment

The input is k protein structures, {Pi}i=1k, each represented as a sequence of the centers of the Cα atoms. In addition, the input contains a parameter ε, which is the distance threshold between the matched Cα atoms. The goal of the algorithm is to compute, for each r = 2,…, k, the largest multiple alignments consisting of exactly r structures. Practically, the number of multiple alignment solutions computed for each r is a user-defined parameter.

Here, we briefly explain the main idea of the method (37). First, we pick a pivot structure and require that it is included in all multiple alignments. In order to prevent dependency on a pivot structure, all input structures are iteratively selected to be a pivot one. We call two sequential (without gaps) fragments of the same length to be ε-congruent if there exists a Euclidean 3D transformation that super-imposes both fragments with root mean square deviation (rmsd) less than ε.

The MultiProt algorithm consists of three major stages. In the first stage, all ε-congruent fragment pairs are efficiently detected between the pivot and all other structures. Secondly, we compute all possible combinations of ε-congruent multiple (sub)-fragments. This stage is analogous to the detection of all non-gaped local multiple alignments. To prevent an exponential number of multiple local alignments, we do not compute them explicitly but rather store all possible alignments by means of combination sets. Such set consists of one fragment from the pivot structure and itsε-congruent fragments from other structures; therefore, such set may include several fragments from some molecule. Thirdly, for each local multiple alignment set, we heuristically select [the problem of selecting the optimal combination is NP-hard (39)] a combination of fragments, one fragment from each structure. Once a unique combination is selected, we compute a global multiple correspondence between the Cα atoms. At this stage, we have a choice (user defined parameter) whether to compute a sequential alignment or a non-sequential one.

The main idea of the MultiProt approach is its ability to efficiently compute a large number of local non-gapped multiple structure alignments. Essentially, a local multiple alignment is computed for each possible fragment of the input molecules. Such local alignments serve as a basis for the extension to the larger partial multiple alignments. In addition, in order to detect subset alignments, the solutions are scored separately according to protein composition, that is, a scoring of alignment between proteins {a, b, c} does not effect a ranking of an alignment between {a, b, d} (the application for this requirement is demonstrated in Subheading 4.3).

To compute a significance of a multiple structural alignment, we apply a simple estimation by means of p-value. Naturally, the p-value depends on the number of input proteins, k, and their sizes. Therefore, we computed the multiple alignment size distribution for different values of k (practically only for k = 2,…, 10). We selected a representative set of 5674 protein structures from the SCOP database (1.65) (40) which have less than 40% of pairwise sequence identity [this data set is provided by ASTRAL (41)]. This resulted in 2304 protein domains (according to the SCOP classification). For each domain we arbitrarily selected only one structure. For each k, the number of structures, we applied MultiProt on k randomly selected structural domains. In total, we performed 10,000 such random alignments for each k. We computed these distributions separately for sequential and non-sequential alignments. Clearly, larger structures will likely produce larger alignments. Therefore, given some multiple alignment size and the minimal structure size, smin, of the aligned molecules, we estimate its significance only from distributions of multiple alignments with minimal molecule size within 20% of smin.

MultiProt is time efficient. On a standard PC with Pentium(R) 4, 2.00 GHz, on proteins with average size of 179 amino acids (from the above data set), the average running time for 2,…, 10 structures is 0.5 s, 4 s, 10.5 s, 21 s, 38 s, 1 min, 1 min 30 s, 2 min 10 s and 3 min 3 s, respectively.

3. Structure-Derived Multiple Sequence Alignment

Sequence alignment methods may produce inaccurate alignments due to low sequence identity. For proteins with solved 3D structure, a structural superposition provides a basis for a more robust assessment of evolutionary relationships between amino acids. Yet, a structural 3D superposition does not uniquely define an alignment between protein sequences. Consider the following scenario, where are given a multiple structure superposition between three proteins {ai},{bj}, and {ck}. Assume that the following groups of amino acids have been superimposed close in 3D space: (a1,b1), (a2,b2,c1), (a3,b3,c2), and (a3,b3,c3). Therefore, there can be several multiple sequence alignments that are consistent with the structural superposition, for example, the following three combinations (for the purpose of sequence alignment we assume that all sequence and structural alignments are according to protein sequence order):

I. II. III.
a1 a2 a3 a1 a2 a3 a1 a2 a3
b1 b2 b3 b1 b2 b3 b1 b2 b3
c1 c2 c3 c1 c2 c3 c1 c2 c3

Which alignment is preferable? All three satisfy the geometrical constraints. Obviously, in this case, we would prefer an alignment with less gaps that also places more similar amino acid types, according to some substitution matrix, in the same column. Therefore, we face an optimization problem that is similar to the multiple sequence alignment problem but has additional spatial constraints. Here, we propose to perform a multiple sequence alignment that unifies structural information, derived from a multiple structure alignment, with amino acid substitution matrices (42). Because protein structure is generally more conserved than sequence, we propose to perform an optimization of the multiple alignment, first, according to structure and then according to amino acid types combined with 3D information. Namely, we propose the following scheme: Given a set of protein structures, first, perform a multiple structural alignment [e.g., apply the MultiProt (37) method]. Second, based on the multiple structure superposition, perform a multiple sequence alignment optimizing a sequence structure unified scoring function.

The scoring function is the likelihood of amino acid a to substitute b assuming that a(b) is located in a secondary structure of type SSE(a), SSE(b), and the 3D distance, according to a multiple structural alignment, between Cα atoms of a and b is d (a,b). SSE(a) is either a helix, a strand or unspecified.

LikelihoodRatio (a,b)=P(a,b)P3D[d(a,b)SSE(a),SSE(b)]P(a,b)P3D[d(a,b)SSE(a),SSE(b)]

where P(a, b)/P′(a, b) is a commonly used likelihood ratio of amino acid substitutions, P(a, b) is the probability for a to substitute b, and P′(a, b) is the randomly expected probability [our default values are taken from the Blosum62 matrix (43)], P3D [d (a, b)] is the observed probability of distance d (a, b) in a set of structural alignments of closely related proteins (conditioned by the type of secondary structures), and P′3D [d (a, b)] is the randomly expected probability of 3D distances (42).

Finally, the score for the 3D substitution matrix is defined as log-odds,

Score (a,b)=2log2[P(a,b)P(a,b)]+2log2{P3D[d(a,b)SSE(a),SSE(b)]P3D     [d(a,b)SSE(a),SSE(b)]}.

Therefore, the values of the first term are taken from a standard substitution matrix and the newly computed values of the second term are given in Table 1. These scores are applied in the multiple sequence alignment algorithm. To solve the multiple alignment, we apply an iterative profile-profile alignment procedure. First, each structure is initialized as a singleton profile. At each step of iteration, the two most similar profiles are joined into one until only one profile, which includes all the input structures, is left.

Table 1.

3D Substitution Scores

Distance [0,1] [1,2] [2,3] [3,4] [4,5] [5,6] [6,7] [7,8] [8,9] [9, ∞]
H H 5.23 4.23 3.82 2.62 −0.19 −1.93 −3.06 −3.37 −3.28 −3.13
H S −5.01 −5.45 −4.90 −5.23 −5.79 −6.26 −6.40 −6.17 −6.09 −5.94
HU 2.36 1.07 0.92 0.15 −1.17 −2.26 −3.15 −3.58 −3.98 −4.23
S S 9.09 5.80 3.99 3.11 −2.07 −3.67 −4.33 −4.61 −4.97 −4.11
S U 4.73 2.37 1.63 0.81 −2.28 −3.54 −4.26 −4.56 −4.56 −4.72
UU 8.66 5.26 3.97 3.04 1.19 −0.30 −1.58 −2.42 −3.21 −3.63
a

H, helix, S, strand, U, undefined.

Optionally, the method allows to produce distance constrained alignments. This is achieved by requiring that all pairwise distances between the Cα atoms of amino acids from the same column are less than some predefined parameter. Such requirement is trivially incorporated into the scoring function of the profile-profile alignment procedure. The constrained multiple alignments allow identification and clustering of structurally similar regions. One such application is shown in Subheading 4.6.

The program output format of multiple alignments is ClustalX or PIR. Optionally, for each column of multiple alignment a structure-sequence conservation score can be reported. It combines the amino acid types and the structural superposition of these amino acids. For reasons of practical convenience, the scores are scaled into the range [0,9] and are displayed under each multiple alignment column. Therefore, a visual examination of these scores reveals whether a region is conserved or not (see examples in Subheading 4.6). In addition, for each input protein file, the amino acid temperature factor field can be set to the corresponding conservation score. This allows convenient 3D visualization in color of the amino acid conservation scores.

4. Experimental Examples

4.1. Pairwise Alignment Cases

First, we test the MultiProt method with non-trivial pairwise alignment cases. We repeated the experiment presented by Shindyalov and Bourne (3). The experiment presents a set of 10 protein pairs and pairwise alignments performed by different (pairwise) methods. The results are presented in Table 2. Two kinds of MultiProt results are given: alignments that preserve protein backbone order and sequence order independent alignments. As can be observed from the table, our pairwise results are very competitive. The maximal running time (pair 1crl:534, 1ede:310) is less than 4 s.

Table 2.

Pairwise Structural Alignment test

Molecule 1 (size) Molecule 2 (size) VAST Sal/rms Dali Sal/rms CE Sal/rms GH Sal/rms MultiProt1 Sal/rms MultiProt2 Sal/rms
1fxi:A(96) 1ubq(76) 48/2.1 51/1.6 44/1.7 50/1.8
1 ten(89) 3hhr:B(195) 78/1.6 86/1.9 87/1.9 81/1.7 81/1.3 82/1.3
3hla:B(99) 2rhe(114) 63/2.5 85/3.5 62/1.8 60/1.8 67/1.9
2aza:A(129) 1paz(120) 4/2.2 - 85/2.9 74/1.9 75/2.0 85/2.5
1cew:I(108) 1mol:A(94) 71/1.9 81/2.3 69/1.9 66/1.6 76/1.8 75/1.9
1cid(177) 2rhe(114) 85/2.2 95/3.3 94/2.7 70/1.5 84/1.8 88/1.9
1crl(534) 1ede(310) 211/3.4 187/3.2 80/1.9 161/2.3 232/2.4
2sim(381) 1nsb:A(390) 284/3.8 286/3.8 264/3.0 197/2.0 233/2.3 268/2.3
1bge:B(159) 2gmf:A(121) 74/2.5 98/3.5 94/4.1 72/1.8 78/2.5 88/2.2
1tie(166) 4fgf(124) 82/1.7 108/2.0 116/2.9 87/1.7 95/2.1 99/2.3
a

RMS, root mean square; Sal, number of aligned atoms. The protein pairs are classified as ‘difficult’ for structural anaysis (48). The alignments are performed by VAST (1), Dali (44), CE (3), Geometric Hashing (GH) method (2) (http://bioinfo3d.cs.ac.il/c_alpha_match/), and MultiProt. The information in this table, except for the GH method and MultiProt results, is taken from Shindyalov and Bourne (3) MultiProt1 results do preserve the sequence order, while MultiProt2 are sequence order independent.

4.2. Sequence Order Independent Structure Alignment

A four-helix arrangement appears in a large number of proteins. SCOP includes at least 40 folds with a four-helix bundle. Holm and Sander (44) show an alignment of the Rop protein (1rop) with cytochrome b56 (256b). Both proteins have a four-helix bundle, but the topological arrangement is different, that is, when the two structures are aligned, at least one helix pair is aligned in an opposite sequential order. Here, we show a multiple structural alignment of four proteins (1f4n, 2cbl:A, 1b3q, and 1rhg:A) which share a four-helix bundle (see Fig. 2). Figure 2C shows the direction of the protein sequences according to a structural alignment when all four helices are aligned. As one can see the direction is different for the last two helices. Thus, none of the commonly used sequence alignment methods can align simultaneously the four α-helices. Figure 2B shows a multiple structural alignment with the four helices aligned. The running time is 14 s.

Fig. 2.

Fig. 2.

Sequence order independent structure alignment. (A) Four proteins, 1f4n, 2cbl:A, 1b3q, and 1rhg:A, containing a four-helix bundle. (B) Multiple alignment of four-helix bundle produced by MultiProt. (C) Schematic representation of the sequence alignment derived from the multiple structural superposition. This common structural motif cannot be detected by standard sequence alignment methods due to different topological arrangement and different chain composition.

4.3. Detection of Partial and Subset Alignments

Here, we demonstrate the ability of MultiProt to detect partial and subset multiple alignments. We consider five multi-domain molecules 1adj:A, 1hc7:A, 1qf6:A, 12as:A, and 1v95:A. Some domains are structurally similar. Our goal in studying such a set of proteins is to identify the two common domains (see Fig. 3): Class II aminoacyl-tRNA synthetase (aaRS)-like, catalyic domain and Anticodon-binding domain of Class II aaRS [the classification is according to SCOP (40)]. The multiple alignment for all five structures resulted in a common structural core of size 39 amino acids, consisting mainly of β-sheet and α-helix. Despite the fact that these two domains are differently classified, there is some partial non-random (p-value < 0 001) structural similarity. The solutions containing four structures revealed two high scoring multiple alignments with different protein composition. These multiple alignments are alignments of the first domain (1adj:A, 1hc7:A, 1qf6:A, and 12as:A) and of the second domain (1adj:A, 1hc7:A, 1qf6:A, and 1v95:A). Therefore, in this example, MultiProt successfully carries out the task of subset and partial multiple alignment. The running time is 2 min.

Fig. 3.

Fig. 3.

Partial and subset alignments. (A) Simplified schematic view of protein domains of 1adj:A, 1hc7:A, 1qf6:A, 12as:A, and 1v95:A. (B) All five proteins are aligned. The common core is 39 amino acids (p-value < 0 001). (C) Domain A is aligned between the first four proteins. The common core has 125 amino acids, this is highest ranked solution for four structures. (D) Domain B is aligned between 1adj:A, 1hc7:A, 1qf6:A, and 1v95:A. The common core has 76 amino acids, this is the second ranked solution with molecule ID composition different from the larger alignment of domain A.

4.4. Comparison Against the HOMSTRAD Database

HOMSTRAD is a benchmark database of manually curated multiple alignments of 1032 homologous protein families with available 3D structures (38). Our aim is to compare the quality of multiple alignments produced by MultiProt and STACCATO against the HOMSTRAD database. However, an objective comparison is not trivial as we do not know neither (1) the ultimate scoring function for multiple alignment nor (2) the correct alignment (which will require to know, for instance, the exact phylogenetic tree of the protein family). To overcome this problem, we decided to select two kinds of measures. The first measure is a commonly used sequence similarity score, namely, a sum-of-pairs score according to the BLOSUM62 matrix, which we denote by Seq score. The second measure computes the fitness of structural alignment according to a multiple sequence alignment. We denote the Str score by the number of columns in a multiple alignment that contains at least half non-gap positions and their 3D rmsd is less than 3 Å. Both scores, Seq and Str, are normalized by the length of the multiple alignment. Given two multiple alignments, we conjecture that an alignment with higher values of both scores is closer to the optimum. However, if only one score is higher at the expense of the other, it is not clear which alignment is better.

Table 3 demonstrates a comparison of MultiProt and STACCATO with the HOMSTRAD data set. We applied different gap opening penalties in order to detect the optimal value. The values in the table represent the average (over 1032 alignments) of the improvement/degradation of the Seq and Str scores. For example, for gap opening penalty −7, the number of multiple alignments where our approach improved at least one score while the other score was at least as good as HOMSTRAD score is 587 (57% of the 1032 cases). The number of alignments where our approach degraded both Seq score and Str score is only 16. There are 1000 cases where our approach improved either Seq score or Str score (at the expense of the other).

Table 3.

Comparison Against the HOMSTRAD Database

Gap opening penalty Seq % Str % Number of gaps
−3 +9.5 (+9.4) +0.4 (+8.3) +24
−7 +7.5 (+7.4) +0.3 (+8.2) −7
−10 +6.4 (+6.1) −0.04 (+8.0) −16
−13 +5.4 (+5.1) −0.4 (+7.6) −20
−15 +4.8 (+4.6) −0.7 (+7.3) −23
−30 +1.3 (+1) −3 (+5.4) −33
a

Two experiments have been conducted. In the first one STACCATO has been applied on the multiple structure alignments as found in HOMSTRAD. In the second experiment, STACCATO has been applied on the multiple structure alignments computer by the MultiProt method. The numeric values represent the difference in scores of STACCATO and HOMSTRAD alignments measured in percents relative to the HOMSTRAD score (positive values mean an improvement over the HOMSTRAD alignments). The results of the second experiment are presented in the parentheses. The last column represents a relative difference in the number of gap openings (negative values mean less gaps are opened in the STACCATO alignments).

The default value of gap opening penalty is selected to be −10, for which MultiProt and STACCATO arguably give more accurate alignments than HOMSTRAD. Outside the range of [−7,−10], either the number of gap openings is increased or the Seq and Str scores are decreased.

4.5. Low Sequence Identity with High Structural Similarity

Here, we give a simple example that demonstrates that in order to achieve a correct sequence alignment it is essential to use structural information (if available). We selected three proteins from the Glutathione S-Transferase family with less than 15% of pairwise sequence identity, which is extremely low. These are 1gnw:A:86–211 (Class phi GST from Arabidopsis thaliana), 1g7o:A:76–215 (Glutaredoxin 2 from Escherichia coli), and 1gwc:A:87–224 (Class tau GST from Triticum tauschii l.).

Due to low sequence similarity, multiple sequence alignment methods may produce an inaccurate alignment. However, these proteins come from the same family and share high structural similarity; therefore, an accurate alignment can be computed from a multiple structure superposition (see Fig. 4).

Fig. 4.

Fig. 4.

(A) A fragment of alignment produced by MultiProt and STACCATO. The structural superposition conservation score, which is displayed under each column, shows that most of the fragment is structurally aligned within 2–3 Å. (B) Alignment computed by ClustalW (47). Some of the discrepancies are marked with boxes.

The active site residues of the Glutathione S-Transferase family were analyzed by Zhang et al. (45). They proposed a method, SAPS, to compute a more accurate sequence alignment based on multiple structure information.

However, SAPS has a restriction that only non-gap fragments are aligned. Therefore, our proposed method has an advantage of computing complete alignments including optimization of gap regions (see Fig. 5).

Fig. 5.

Fig. 5.

Alignment of Glutathione S-Transferase proteins as computed by STACCATO. Active site residues are correctly aligned and are marked with boxes. The correct alignment of these regions was shown by the SAPS method (45). Bold box marks structurally more variable region which contains additional active site residues which were not aligned by SAPS.

4.6. Loop Movement in Tyrosine Kinase

Tyrosine kinase represents a large family of evolutionarily conserved enzymes that play a critical role in cellular signaling pathways (46). Here, our aim is to analyze a multiple alignment of the activation loop. We selected two kinase proteins in the active state (1ir3:A and 1cdk:A) and two in the inactive state (1irk and 1iep). These are insulin receptors from human (1ir3:A and 1irk), a cAMP-dependent protein kinase catalytic subunit from pig (1cdk:A), and Abelsone tyrosine kinase from mouse (1iep).

At the beginning of the activation loop there is a well-conserved DFG motif, which is involved in Mg-ATP binding. During the activation, the loop undergoes a significant conformational change, when some amino acids change their position by as much as 31 Å (see Fig. 6A). What kind of analysis should be performed in order to detect and distinguish between two conformational states and detect residues participating in this reorganization? Clearly, the multiple sequence alignment does not recognize active and inactive states, as it multiply aligns the sequence of the whole activation loop. In our approach, as discussed above in Subheading 3., we are able to apply distance constrained alignment which aligns only amino acids closely located in space. Applying a distance threshold of 5 Å, we are able to distinguish between two states of the activation loop as shown in Fig. 6C. Only the aspartic residue from the DFG motif is spatially conserved (within 5 Å distance threshold).

Fig. 6.

Fig. 6.

(A) Activation loop of tyrosine kinase in the active and the inactive state. The DFG motif is located at the beginning of the loop. Structural disposition of some residues may be as large as 31 Å. (B) Alignment produced by STACCATO. Visual inspection of the structural superposition conservation score, which is displayed under each column, suggests a significant structural variability of the region. (C) Alignment produced with distance constraint of 5 Å. Two separate structurally similar clusters are clearly revealed. Only aspartic acid from the DFG motif is aligned.

5. Conclusions

Here, we have discussed a powerful method, MultiProt, for multiple protein structure alignment. The main advantages of our method are (1) simultaneous structure superposition (no side effects of pairwise alignment methods), (2) solutions are detected for any number of molecules (subset alignments), (3) proteins can consist of several domains or even several chains (partial alignments), (4) the final alignments can optionally preserve the sequence order or be sequence order independent, and (v) time efficiency (tens and even hundreds of molecules). In case that protein structures differ modulo hinge motion, a complete alignment can be detected by a manual examination of several largest partial solutions. In addition, we have discussed the problem of structure-based multiple sequence alignment, which in the case of protein structure availability overcomes the problems inherent to sequence alignment methods. The discussed method, STACCATO, with the combination of MultiProt produces multiple alignments as good (and arguably slightly better) as multiple alignments from the manually curated HOMSTRAD data base.

Acknowledgments

The research of M. Shatsky is supported by a PhD fellowship in “Complexity Science” from the Yeshaya Horowitz association. This research was supported by the Israel Science Foundation (grant no. 281/05), the Binational US-Israel Science Foundation (BSF) and by the Hermann Minkowski-Minerva Center for Geometry at Tel Aviv University. The research of R. Nussinov has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under contract number NO1-CO-12400. This research was supported (in part) by the Intramural Research Program of the NIH, National Cancer Institute, Center for Cancer Research. The content of this publication does not necessarily reflect the view or policies of the Department of Health and Human Services, nor does mention of trade names, commercial products, or organization imply endorsement by the U.S. Government.

References

  • 1.Madej T, Gibrat J, and Bryant S Threading a database of protein cores. Proteins 23:356–369, 1995. Online available on http://www.ncbi.nlm.nih.gov/Structure/VAST/vast.shtml. [DOI] [PubMed] [Google Scholar]
  • 2.Bachar O, Fischer D, Nussinov R, and Wolfson H A computer vision based technique for 3-D sequence independent structural comparison. Protein Eng 6:279–288, 1993. [DOI] [PubMed] [Google Scholar]
  • 3.Shindyalov I and Bourne P Protein structure alignment by incremental combinatorical extension (ce) of the optimal path. Protein Eng 11(9):739–747, 1998. Online available on http://cl.sdsc.edu/ce.html. [DOI] [PubMed] [Google Scholar]
  • 4.Dietmann S, Park J, Notredame C, Heger A, Lappe M, and Holm L A fully automatic evolutionary classification of protein folds: dali domain dictionary version 3. Nucleic Acids Res 29(1):55–57, 2001. Online available on http://www.embl-ebi.ac.uk/dali/. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Shatsky M, Nussinov R, and Wolfson H FlexProt: alignment of flexible protein structures without a pre-definition of hinge regions. Journal of Computational Biology 11(1):83–106, 2004. [DOI] [PubMed] [Google Scholar]
  • 6.Eidhammer I, Jonassen I, and Taylor W Structure comparison and structure patterns. J Comput Biol 7:685–716, 2000. [DOI] [PubMed] [Google Scholar]
  • 7.Orengo CA, Michie AD, Jones S, Jones DT, Swindells MB, and Thornton JM CATH – a hierarchic classification of protein domain structure. Structure 5(8):1093–1108, 1997. [DOI] [PubMed] [Google Scholar]
  • 8.Ma B, Elkayam T, Wolfson H, and Nussinov R Protein-protein interactions: structurally conserved residues distinguish between binding sites and exposed protein surfaces. Proc Natl Acad Sci USA 100(10):5772–5777, 2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chung J, Wang W, and Bourne P Exploiting sequence and structure homologs to identify protein-protein binding sites. Proteins 62(3):630–640, 2006. [DOI] [PubMed] [Google Scholar]
  • 10.Aytuna A, Gursoy A, and Keskin O Prediction of protein-protein interactions by combining structure and sequence conservation in protein interfaces. Bioinformatics 21(12):2850–2855, 2005. [DOI] [PubMed] [Google Scholar]
  • 11.Akutsu T and Sim K Protein threading based on multiple protein structure alignment. In Genome Informatics (GIW’99), Asai K and Miyano S and Takagi T (eds). Universal Academy Press, Tokyo, 23–29, 1999. [PubMed] [Google Scholar]
  • 12.Goldsmith-Fischman S and Honig B Structural genomics: computational methods for structure analysis. Protein Sci 12(9):1813–1821, 2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Koehl P Protein structure similarities. Curr Opin Struct Biol 11:348–353, 2001. [DOI] [PubMed] [Google Scholar]
  • 14.Kolodny R, Koehl P, and Levitt M Comprehensive evaluation of protein structure alignment methods: scoring by geometric measures. J Mol Biol 346(4): 1173–88, 2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Bennett M, Schlunegger M, and Eisenberg D 3d domain swapping: a mechanism for oligomer assembly. Protein Sci 4:2455–2468, 1995. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Dror O, Benyamini H, Nussinov R, and Wolfson H MASS: multiple structural alignment by secondary structures. Bioinformatics 19 Suppl. 1:i95–i104, 2003. [DOI] [PubMed] [Google Scholar]
  • 17.Yuan X and Bystroff C Non-sequential structure-based alignments reveal topology-independent core packing arrangements in proteins. Bioinformatics 21(7):1010–1019, 2005. [DOI] [PubMed] [Google Scholar]
  • 18.Ambuhl C, Chakraborty S, and Gartner B Computing largest common point sets under approximate congruence. In Proceedings of the 8th Annual European Symposium on Algorithms, 52–63, Springer-Verlag, Springer, Berlin, 2000. [Google Scholar]
  • 19.Akutsu T Protein structure alignment using dynamic programming and iterative improvement. IEICE Trans Inf Syst E79-D:1629–1636, Springer; Berlin, 1996. [Google Scholar]
  • 20.Kolodny R and Linial N Approximate protein structural alignment in polynomial time. Proc Natl Acad Sci USA 101(33):12201–12206, 2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Shatsky M, Shulman-Peleg A, Nussinov R, and Wolfson H The multiple common point set problem and its application to molecule binding pattern detection. J Comput Biol 13(2):407–428, 2006. [DOI] [PubMed] [Google Scholar]
  • 22.Edgar R Muscle: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res 32(5):1792–1797, 2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Gerstein M and Levitt M Using iterative dynamic programming to obtain accurate pairwise and multiple alignments of protein structures. In Proceedings of the Fourth International Conference on Intelligent Systems in Molecular Biology, 59–67, Menlo Park, CA, AAAI Press, Heidleberg, Germany, 1996. [PubMed] [Google Scholar]
  • 24.Russell R and Barton G Multiple protein sequence alignment from tertiary structure comparison: assignment of global and residue confidence levels. Proteins 14:309–323, 1992. [DOI] [PubMed] [Google Scholar]
  • 25.Taylor WR, Flores T, and Orengo C Multiple protein structure alignment. Protein Sci 3:1858–1870, 1994. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Ye Y and Godzik A Multiple flexible structure alignment using partial order graphs. Bioinformatics 21(10):2362–2369, 2005. [DOI] [PubMed] [Google Scholar]
  • 27.Ochagavia ME and Wodak S Progressive combinatorial algorithm for multiple structural alignments: application to distantly related proteins. Proteins 55(2):436–454, 2004. [DOI] [PubMed] [Google Scholar]
  • 28.Konagurthu A, Whisstock J, Stuckey P, and Lesk A Mustang: a multiple structural alignment algorithm. Proteins 64(3):559–574, 2006. [DOI] [PubMed] [Google Scholar]
  • 29.Leibowitz N, Nussinov R, and Wolfson H MUSTA-a general, efficient, automated method for multiple structure alignment and detection of common motifs: application to proteins. J Comput Biol 8:93–121, 2001. [DOI] [PubMed] [Google Scholar]
  • 30.Leibowitz N, Fligelman Z, Nussinov R, and Wolfson H Automated multiple structure alignment and detection of a common substructural motif. Proteins 43:235–245, 2001. [DOI] [PubMed] [Google Scholar]
  • 31.Wolfson HJ and Rigoutsos I Geometric hashing: an overview. IEEE Comput Sci Eng 4(4):10–21, 1997. [Google Scholar]
  • 32.Nussinov R and Wolfson H Efficient detection of three-dimensional structural motifs in biological macromolecules by computer vision techniques. Proc Natl Acad Sci USA 88:10495–10499, 1991. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Shatsky M, Fligelman Z, Nussinov R, and Wolfson H Alignment of flexible protein structures. In 8th International Conference on Intelligent Systems for Molecular Biology, 329–343, AAAI press, Heidleberg, Germany, 2000. [PubMed] [Google Scholar]
  • 34.Jonassen I, Eidhammer I, Conklin D, and Taylor W Structure motif discovery and mining the pdb. Bioinformatics 18(2):362–367, 2002. [DOI] [PubMed] [Google Scholar]
  • 35.Dror O, Benyamini H, Nussinov R, and Wolfson H Multiple structural alignment by secondary structures: – algorithm and applications. Protein Sci 12:2492–2507, 2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.O’Sullivan O, Suhre K, Abergel C, Higgins D, and Notredame C 3Dcoffee: combining protein sequences and structures within multiple sequence alignments. J Mol Biol 340(2):385–395, 2004. [DOI] [PubMed] [Google Scholar]
  • 37.Shatsky M, Nussinov R, and Wolfson H A method for simultaneous alignment of multiple protein structures. Proteins 56(1):143–156, 2004. [DOI] [PubMed] [Google Scholar]
  • 38.Mizuguchi K, Deane C, Blundell T, and Overington J Homstrad: a database of protein structure alignments for homologous families. Protein Sci 7:2469–2471, 1998. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Akutsu T and Halldorson MM On the approximation of largest common subtrees and largest common point sets. Theor Comput Sci 233:33–50, 2000. [Google Scholar]
  • 40.Murzin A, Brenner S, Hubbard T, and Chothia C SCOP: a structural classification of proteins database for the investigation of sequences and structures. J Mol Biol 247:536–540, 1995. [DOI] [PubMed] [Google Scholar]
  • 41.Chandonia J, Hon G, Walker N, Lo Conte L, Koehl P, Levitt M, and Brenner S The astral compendium in 2004. Nucleic Acids Res 32:D189–D192, 2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Shatsky M, Nussinov R, and Wolfson HT Optimization of multiple sequence alignment based on multiple structure alignment. Proteins 62(1):209–217, 2006. [DOI] [PubMed] [Google Scholar]
  • 43.Henikoff S and Henikoff J Amino acid substitution matrices from protein blocks. Proc Natl Acad Sci USA 89(22):10915–10919, 1992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Holm L and Sander C Protein structure comparison by alignment of distance matrices. J Mol Biol 233:123–138, 1993. [DOI] [PubMed] [Google Scholar]
  • 45.Zhang Z, Lindstam M, Unge J, Peterson C, and Lu G Potential for dramatic improvement in sequence alignment against structures of remote homologous proteins by extracting structural information from multiple structure alignment. J Mol Biol 332(1):127–142, 2003. [DOI] [PubMed] [Google Scholar]
  • 46.Hubbard S and Till JH Protein tyrosine kinase structure and function. Ann Rev Biochem 69:373–398, 2000. [DOI] [PubMed] [Google Scholar]
  • 47.Higgins D, Thompson J, and Gibson T CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res 22:4673–4680, 1994. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Fischer D, Elofsson A, Rice D, and Eisenberg D Assessing the performance of fold recognition methods by means of a comprehensive benchmark. In Proceedings of Pacific Symposium on Biocomputing (Hunter L and Klein T, editors), World Scientific Press, Singapore, 300–318, 1996. [PubMed] [Google Scholar]

RESOURCES