Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 31.
Published in final edited form as: J Biomed Inform. 2015 Jan 13;54:10–38. doi: 10.1016/j.jbi.2015.01.002

Adapting simultaneous analysis phylogenomic techniques to study complex disease gene relationships

Joseph D Romano a,1, William G Tharp b, Indra Neil Sarkar a,c,d,*
PMCID: PMC13422749  NIHMSID: NIHMS2187726  PMID: 25592479

Abstract

The characterization of complex diseases remains a great challenge for biomedical researchers due to the myriad interactions of genetic and environmental factors. Network medicine approaches strive to accommodate these factors holistically. Phylogenomic techniques that can leverage available genomic data may provide an evolutionary perspective that may elucidate knowledge for gene networks of complex diseases and provide another source of information for network medicine approaches. Here, an automated method is presented that leverages publicly available genomic data and phylogenomic techniques, resulting in a gene network. The potential of approach is demonstrated based on a case study of nine genes associated with Alzheimer Disease, a complex neurodegenerative syndrome.

The developed technique, which is incorporated into an update to a previously described Perl script called “ASAP,” was implemented through a suite of Ruby scripts entitled “ASAP2,” first compiles a list of sequence-similarity based orthologues using PSI-BLAST and a recursive NCBI BLAST+ search strategy, then constructs maximum parsimony phylogenetic trees for each set of nucleotide and protein sequences, and calculates phylogenetic metrics (Incongruence Length Difference between orthologue sets, partitioned Bremer support values, combined branch scores, and Robinson–Foulds distance) to provide an empirical assessment of evolutionary conservation within a given genetic network. In addition to the individual phylogenetic metrics, ASAP2 provides results in a way that can be used to generate a gene network that represents evolutionary similarity based on topological similarity (the Robinson–Foulds distance).

The results of this study demonstrate the potential for using phylogenomic approaches that enable the study of multiple genes simultaneously to provide insights about potential gene relationships that can be studied within a network medicine framework that may not have been apparent using traditional, single-gene methods. Furthermore, the results provide an initial integrated evolutionary history of an Alzheimer Disease gene network and identify potentially important co-evolutionary clustering that may warrant further investigation.

Keywords: Bioinformatics, Phylogenetics, Simultaneous analysis, Alzheimer Disease, Comparative genomics

1. Introduction

Classical genetic diseases typically arise due to isolated genetic changes within a single gene or allele [1]. Many of these “simple” or “monogenic” diseases follow Mendelian patterns of inheritance. The responsible genetic lesion is often the result of an insertion or deletion event, or the transversion/transposition of a nucleotide. The probability for transmission of simple genetic disorders may thus be easily predicted and generally follow sex-linked or autosomal patterns of heredity. Classic examples of monogenic disorders include cystic fibrosis, sickle cell anemia, and achondroplasia [24]. By contrast, complex diseases or disorders may not follow clear hereditary patterns or be diagnosed based on isolated genetic lesions. However, many complex diseases such as cardiovascular disease, type 2 diabetes mellitus, and Alzheimer Disease occur with higher frequency among families and close genetic relatives – suggesting that the interaction of genetic elements may play a central role in their pathogenesis, beyond environmental or behavioral factors [5]. Identifying risks for complex diseases and developing new approaches for treating or preventing them may benefit from high-throughput, computational, or bioinformatics based approaches. Related advances in biotechnology have facilitated the identification of genotypes that may be factors involved in the heritability of complex genetic diseases [6]. For example, specific genotypes can be associated with a probabilistic value of susceptibility relative to the gene(s) they influence and thus correlated with a disease phenotype [1,79].

Due to limited knowledge about the specific mechanisms by which multiple genetic factors may influence complex diseases, pharmacotherapies are often aimed at managing symptoms or laboratory values, and are therefore reactionary and not preventative. Thus, the approach to complex disease management necessarily extends beyond pharmacotherapy, attempting environmental and behavioral changes through patient education or lifestyle modification [2,10]. A major current goal of biomedical research is therefore to better characterize complex relationships between contributing factors associated with complex diseases for identifying possible targets for therapeutic intervention. Genetic background influences the susceptibility to complex disease, which is an artifact of the structural or functional relationships between some or all members of a disease gene network [3,7]. These relationships may include direct physical interaction between the protein products of the genes, parallel functionality in metabolic pathways, or co-localization of protein products in a certain cell or tissue type [4,7]. These data are not easily elucidated using approaches that are focused on a single gene or pathway, and instead require a broader systems-based methodology. Understanding the shared history of multiple genes may provide guidance in developing approaches that target multiple genes that have evolved to work together through evolutionary time.

Complicating the assessment of such systems-based methodologies is the lack of reference standards for benchmarking approaches for discovering how myriad complex disease genes work in the context of disease phenotypes. It should therefore be noted that methodologies for interpreting multiple genes associated with complex disease may not quantifiably be benchmarked against previously used methods that focus on single-gene analysis. Instead, methodologies for multiple gene analysis can be seen as identifying potential relationships between genes, necessitating the development of benchmarks and controls that can be performed internally against known genetic interactions (and cases where one can be fairly certain that no interaction is taking place).

Within the context of translational bioinformatics, there have been a limited number of attempts to address understanding complex diseases that accommodate multiple disease genes simultaneously. One method has described the use of Mendelian genetic traits that occur in coincidence with complex genetic diseases to predict underlying mechanisms of those diseases, primarily by linking diverse biomedical databases [11]. Another approach involves the linking of complex disease genes to other diseases based on molecular similarity, which adapts a vector space model approach [12].

Perhaps the most significant approach developed to date for studying the potential impact of multiple disease genes is that of “network medicine,” which systematically describes the rise of disease phenotypes as perturbations in the normal interactions of molecular, environmental, and population networks rather than a single macromolecule or biological pathway [13]. Network medicine postulates relationships between genes based on observed interacting phenomena, accounting for numerous genetic events or environmental factors may contribute to similar phenotypic results in one class of organism but not in another. Notably missing from network medicine approaches to data is the inclusion of models that reflect evolutionary similarities between genes that comprise a disease network. Evolutionary models may provide a perspective to the network medicine approach and enable the inclusion of ancient environmental and population data into the construction of a disease network and may identify both previously unknown factors contributing to disease development and new model organisms for understanding disease pathology.

Phylogenetic analyses infer potential evolutionary relationships based on similarities implying common descent from shared ancestry and are performed on data sets consisting of physical, functional, or molecular representations [14]. Genomic analyses typically construct the analytic matrix using nucleotide or amino-acid sequences from different individuals or species (termed “taxa”; singular “taxon”). Classically, the resulting data are presented as trees where the branching points (termed “nodes”) give rise to hierarchical groupings of more similar taxa (akin to leaves on a branch). These trees can be used to explore potential patterns of divergence from a common ancestor as well as the degree of difference among taxa included in the tree. This degree of difference is usually described as an evolutionary “distance” that can be inferred multiple ways, but typically represents a measure of evolutionary change (based upon sequence differences) or an amount of time since divergence likely occurred [15,16].

However, like experiments focused on a single gene or pathway, an isolated phylogenetic analysis may not capture important features of co-evolution or conservation of gene clusters impacting complex disease processes. Additionally, reliance on phylogenetic trees of individual genes may not fully address the potential for genetic changes such as lateral gene transfer, reversion of mutations, or recombination events [17,18].

To account for multiple evolutionary patterns represented by multiple genes, data matrices can be combined into a single phylogenetic analysis through a “simultaneous analysis” (SA) approach [1921]. In SA, individual data blocks (e.g., a sequence matrix for a particular gene; referred to as a “partition”) are systematically combined to enable higher-order analyses that transcend data derived from analysis of an individual partition. Frequently, SA values are derived by applying arithmetic operations on other (already determined) SA values, so the workflow tends to follow a stepwise pattern. SA techniques have been shown to strengthen the overall support for the evolutionary patterns represented by trees determined by single partition phylogenetic analyses [22]. However, while the aforementioned network medicine approach aims to examine relationships between multiple genes simultaneously, there has been no direct incorporation of phylogenetic metrics, such as those that result from a SA, in the study of disease genes. There is thus a significant opportunity to explore the potential utility of SA techniques in light of network medicine approaches towards the ultimate goal of developing a multi-faceted perspective of the evolution and interaction of disease genes.

This study aimed to explore the potential of leveraging SA techniques for providing additional information that can be used within a network medicine context. The approach builds on the notion that evolutionarily significant types of relationships may be identified using phylogenetic techniques within the genetic context of a given complex disease [5,39,40]. In particular, this study sought to explore the relative evolutionary conservation of nine genes associated with Alzheimer Disease (AD). The potential meaningfulness of the suggested evolutionarily similar relationships is gauged relative to both sets of genes that are known to be highly correlated (mitochondrial genes) and non-correlated (randomly selected genes from different, unrelated diseases). Furthermore, the relative co-evolution of AD disease genes was gauged relative to those genes that have been previously identified as being uniquely associated with AD versus those that are more polygenic. The promising results of this feasibility study suggest that phylogenetic approaches that can incorporate information from multiple genes may offer evolutionary insights into complex disease gene relationships. The results further suggest that disease genes that are unique to a single disease may have a relatively conserved evolution across orthologues across evolutionary time compared to those that are shared with other diseases.

2. Materials and methods

2.1. Overview

The overall goal of this study was to develop an automated method for an SA phylogenetic analysis and to use it to construct an evolutionary history for a set of genes that may be correlated with a given complex genetic disorder within a network medicine context. AD was chosen as a case study because it has that distinct genes have been identified previously to correlate strongly with formation of the disease, which is a likely artifact of the fact that AD is a popular complex disease to study. Co-evolutionary patterns were examined for nine of these genes using a SA technique mediated by a series of scripts written in the Ruby scripting language. The resulting phylogenetic tree provides the first description of the co-evolution of genes within a network medicine framework that may reveal the development and pathogenesis patterns of AD. The developed technique provides a framework for an automated approach to study the co-evolution of gene sets associated with a complex disease using a robust phylogenetic methodology.

In this study, a previous automated SA approach (Automated Simultaneous Analysis Phylogenetics; ASAP [41]) was both refined and expanded upon to collect and analyze disease genes based on: (1) the degree of corroboration between partitions; and (2) the support for an overall consensus tree modeling a putative evolutionary relationship common to all partitions, using maximum parsimony analysis [42]. The final phase then generates a gene network based on the Robinson–Foulds tree similarity metric [43]. The original ASAP began with a set of given completed data partitions, which would then be aligned, subjected to maximum parsimony phylogenetic analyses that resulted in the reporting of accepted simultaneous analysis support values. ASAP2 includes many of the features of ASAP, but also is able to automatically construct data partition from a set of query genes, perform data partition homogeneity tests (the Incongruence Length Difference [ILD], described below) and calculate Robinson–Foulds topology distances between trees. An added benefit of ASAP2 is that it utilizes freely available and modern phylogenetic tools (e.g., TNT in place of PAUP*) and provides a more general implementation of SA that could be used in the context of traditional multi-gene analyses or to specifically enable the analysis of complex disease genes, such as presented here.

The automated collection and simultaneous phylogenetic analysis process was developed using a sequential set of Ruby [44] scripts (entitled and referred to henceforth as “ASAP2” – DOI: http://dx.doi.org/10.5281/zenodo.11075) that made use of the Bioruby gem [45] as well as the following freely available genomic or phylogenetic analysis tools: BLAST+ (Basic Local Alignment Search Tool [46]), MUSCLE (MUltiple Sequence Comparison by Log-Expectation [47]), and TNT (Tree analysis using New Technology [48]). The overall workflow for ASAP2 is shown in Fig. 1. From a technical perspective, ASAP2 is an improvement over an earlier version of the implemented approach (ASAP, which was written as a single Perl script) both in terms of modularity and use of contemporary phylogenomic tools.

Fig. 1.

Fig. 1.

Overview of ASAP2 workflow. The process, as implemented in the study, begins with the providing of GenBank IDs for protein sequences, which may originate from reference resources like OMIM or other user chosen sequences. A combination of a highly specific recursive BLAST+ approach and PSI-BLAST is used to identify sequence-similarity based orthologues (using a stringent E-value cutoff of 0.0). For each orthologue protein sequence identified, its corresponding nucleotide GenBank entry is retrieved based on metadata within the protein GenBank sequence. The remaining workflow follows the standard process for Simultaneous Analysis (SA) for both the protein and nucleotide sequence sets (called “partitions” in SA): Sequence Alignment (e.g., using MUSCLE) and phylogenetic tree building (e.g., using TNT). The resulting trees are then compared for each protein and nucleotide partition as well as for the overall protein or nucleotide SA tree. MUSCLE: Multiple Sequence Comparison by Log-Expectation (Multiple sequence alignment software). TNT: Tree analysis using New Technology (Maximum Parsimony phylogenetic analysis software).

2.2. Alzheimer Disease gene sequences

Nucleotide sequences for genes implicated as contributing to a higher risk for Alzheimer Disease (AD) in humans were manually identified using the data associated with the “Alzheimer Disease; AD” entry in Online Mendelian Inheritance in Man (OMIM) (OMIM ID #104300) [49]. This yielded ten discrete genes shown to be related to AD, which were loaded into ASAP2. For the purposes of this study, nonspecific chromosomal regions that encompass numerous genes or the noncoding regions between them were not included.

2.3. Identifying potentially related disease genes based on sequence similarity

From the initial set of human AD disease gene sequences, ASAP2 performed two types of sequence searches from within the non-redundant (nr) protein database maintained at the United States National Library of Medicine’s National Center for Biotechnology Information (NCBI) using NCBI BLAST+ [46]. First, a PSI-BLAST (Position-Specific Iterated BLAST) search (which searches for similar sequences using an iterative profiling approach [50]) was done for each gene sequence. Second, a recursive search was initiated with a BLASTx (which uses a translated nucleotide sequence query to perform a protein search) search of the nr protein database for the gene of interest and the best results used to iteratively search for additional protein sequences using BLASTp (which searches for protein sequences using an amino-acid query) until no additional sequences were found. An expect value (E-value) of 1.0 × 10−256 was used as the criterion for inclusion of results for both the PSI-BLAST and the recursive BLAST algorithm. Candidate data partition matrices for each gene were then constructed based on the combination of PSI-BLAST and recursive BLAST results. Corresponding nucleotide sequences were determined based on information in the DBSOURCE metadata field that links a given protein sequence to its coding nucleotide sequence.

After candidate data partitions were assembled, ASAP2 culled taxa and sequences from each data partition that were not uniformly represented for each gene (i.e., a sequence for a given species must be present in each data partition for that species to be retained for further analysis). Additionally, if any species was represented more than once in any partition, ASAP2 only kept the first (most similar according to BLAST) protein sequence for that species. ASAP2 then assembled the resulting data partitions into FASTA files, aligned them using the default parameters of MUSCLE, and formatted them into TNT-compatible data matrix files. ASAP2 also includes the ability to align sequences using MAFFT or Clustal Omega, which are packaged alongside MUSCLE and may be specified using arguments at runtime. Additionally, the user has the ability to provide sequences aligned through other means (including manually). For the purpose of this study, and to demonstrate the automated capabilities of ASAP2, all sequences were aligned using the default MUSCLE strategy. The generated FASTA files and TNT data matrix files are available as Supplementary Data.

In order to confirm that only relevant sequences were identified by the BLAST search strategy, the description of each identified sequence was manually inspected to ensure that it was appropriate to compare to its corresponding human gene. If the description line of the GenBank entry was not informative, further manual investigation of the sequence was performed (e.g., inspection of the sequence alignment).

Sequence length was manually verified following the selection of sequences, prior to conducting simultaneous analysis. As the SA techniques implanted in ASAP2 do not require a distinction between orthologues and paralogues, this was done only with the intent of preventing the use of large chromosomal assemblies or very short gene fragments. Additionally, the BLAST2 expect value between each pair of sequences in each partition was computed to verify that all sequences are statistically similar to one another; No maximum expect value was used as a cutoff, but all values were smaller than 1 × 10−100.

2.4. Phylogenetic analyses

ASAP2 used TNT to conduct the maximum parsimony phylogenetic analyses of each data partition. Trees were constructed using tree bisection and reconnection (TBR) rearrangement, finding optimal scores 20 times followed by 10 cycles of tree-drifting. Subsequently, group support values were determined by counting the minimum number of steps needed to lose each group by TBR rearrangement [51]. The TNT analysis included individual plotting of apomorphies and synapomorphies, bootstrap resampling, and calculation of both the relative and absolute Bremer support values at each branch.

ASAP2 then generated a SA consensus tree using TNT by creating an interleaved matrix of the data partitions. The interleaved matrix was built by concatenating each aligned data partition (minus headers and metadata) sequentially, separated by line breaks, into a single TNT data file. This data file was then interpreted by TNT as if the sequences for each species were concatenated in the order of the data partitions in the interleaved matrix. The tree building routine was the same as used for analyzing the individual data partitions, except 30 cycles of tree-drifting were used.

The Partitioned Bremer Support (PBS; also known as Partitioned Branch Support) at each node in the SA consensus tree was used as the primary criterion for the evaluation of each data partition. The PBS value is defined as “the minimum number of character steps for [a] partition on the shortest topologies for the combined data set that do not contain that node, minus the minimum number of character steps for that partition on the shortest topologies for the combined data set that do contain that node” [20]. Therefore, a specific PBS value can be interpreted as a measurement of how well the data from a particular partition either support (represented by positive values) or refute (negative values) a particular node on the consensus tree. Branch Support (BS) values, defined as “the minimum number of character steps for that data set on the shortest topologies that do not contain that node, minus the minimum number of character steps for that data set on the shortest topologies that contain that node” were used as the second criterion for evaluation of the SA consensus tree [20]. After determining PBS values across all tree nodes on the consensus tree for each data partition, the BS was determined for each node on the consensus tree by the sum of all PBS values for that particular node. A positive BS score indicates that the overall combined set of data partitions support the topology at that node rather than refute it. ASAP2 uses a slightly modified version (for the purpose of automation) of a previously developed TNT script to calculate the PBS values [52]. Modifications were made to the original TNT script were to facilitate automated data input and processing of output as required by ASAP2 without altering the tree building routines, and minimizing the text-based front end displayed to the user.

The Hidden Branch Support (HBS) for a particular node on the consensus tree was computed as the difference between the BS value at that node in the consensus tree and the sum of the BS values for that node from each data partition. The magnitude of an HBS value of a given node in the consensus tree was used as the final criterion for determining the overall strength of supporting or refuting the topology at the node.

Finally, a gene network was generated from the consensus analyses for each data partition using the Robinson–Foulds (RF) metric to quantify the distance between each pair of trees [43]. This was implemented using a previously written TNT script [53] that was modified to fit within the automated workflow of ASAP2. All calculations and parameters in the script were unchanged from the original version. To transform RF values onto a scale where larger values corresponded to more similarity (conventionally, higher RF values indicate greater dissimilarity based on a normalized count of symmetric differences between trees), the following calculation was used:

RF=1eRF

Cytoscape [54] was then used to visualize the gene network that was based on the RF values as normalized edge weights using a force-directed layout.

2.5. Compatibility and homogeneity testing of partitions

The Incongruence Length Difference (ILD) test [55] was used to quantify the combinability of the data partitions, which is a statistical measure that is commonly employed in Maximum Parsimony analyses. The ILD test compares the difference in length of the best tree for each individual partition with the best tree for all data partitions concatenated into a single block, and then performs the same calculation for a large number of random partitionings of the same data.

2.6. Establishing a benchmark for ASAP2

ASAP2 was run using the mitochondrial genes for the 34 species included in the AD study, where the data partitions were constructed from the protein and translated nucleotide sequences for mtDNA genes. The gene clustering diagrams based on the RF values for the mtDNA analysis were viewed individually, and again when ASAP2 was run using the gene partitions from both the AD study and the mtDNA genomes. Comparison of the relative RF scores was facilitated through a visual comparison of the diagrams to qualitatively interpret the strength of the conclusions made by the AD analysis alone.

Following this benchmark of mtDNA genes (which were presumed to be conserved over evolutionary time), ASAP2 was run again using a set of genes that each affect susceptibility to different complex genetic disorders, aiming to explore any potential relationship of distantly-related genomic elements. As with the mtDNA analysis, the results were interpreted based on RF scores relative to the nine AD genes and also facilitated by visual comparison of the generated diagrams.

Finally, to calibrate the results of the analysis of eight random genes against the set of nine AD genes initially used, ASAP2 was run using the union of the two sets of data partitions, which totaled 16 partitions. Since ASAP2 requires all partitions to contain the same set of taxa, the intersection of species across both sets of partitions was used.

3. Results

3.1. ASAP2

ASAP2. was developed as a set of Ruby scripts and is available at GitHub under the GNU General Public License (https://github.com/UVM-BIRD/asap2 [DOI: http://dx.doi.org/10.5281/zenodo.11075]). The script guides the process of performing a SA from an initial set of Genbank identifiers. By the end of the analysis, ASAP2 produces files containing the data partitions, E-value tables, FASTA files of the final data partitions (both unaligned and aligned), TNT data matrices, and all TNT output, including log files and parenthetically-notated tree files. The ASAP2 data workflow is illustrated in Fig. 1.

3.2. Gathering uniform taxonomic distribution of AD genes

Ten genes associated with Alzheimer Disease susceptibility were initially selected through OMIM for analysis using ASAP2. Due to incompatibility issues with BLASTx, one gene (PAX-interacting protein 1 [PAXIP1]; GI:530387259) was removed from the analysis. In brief, because PAXIP1 contains six BRCT (BRCA C terminus) domains that are homologous to many sequences in GenBank, BLASTx quit at each attempt due to memory overflow. The nine remaining genes used for the remainder of the study are listed in Table 1.

Table 1.

The 10 genes originally selected for the Alzheimer Disease (AD) ASAP2 analysis.

Gene Common name GenBank GI GenPept GI

apbb2 Amyloid-B Precursor Protein-Binding Family B Member 2 225007611 50083291
nos3 Nitric Oxide Synthase 3 231571328 231571329
plau Urokinase-Type Plasminogen Activator 222537757 4505863
sorl1 Sortilin-Related Receptor L 307611954 4507157
a2m α-2-Macroglobulin 66932946 62088808
blmh Bleomycin Hydrolase 530411126 194378004
mpo Myeloperoxidase 4557758 4557759
ace Angiotensin I-Converting Enzyme 295844836 295844837
app Amyloid-β Precursor Protein 324021737 324021738
paxip1 PAX Transcription Activation Domain-Interacting Protein 1

Omitted from the study due to issues with recursive BLAST and PSI-BLAST search strategies.

The combined PSI-BLAST and recursive BLAST results for each gene included in this study resulted in nine data partitions representing 34 unique species (including Homo sapiens; Table 2). If the BLAST analyses resulted in any species being represented more than once in a data partition, only the first sequence (the one most similar to the query sequence) was kept. The protein sequences identified – along with the corresponding source nucleotide sequences – using this process are provided in Supplementary Table 1.

Table 2.

The 34 species identified by ASAP2 using Alzheimer Disease gene queries.

Scientific name Vernacular name NCBI Taxon ID

Homo sapiens Human 9606
Pan troglodytes Chimpanzee 9598
Gorilla gorilla gorilla Western lowland gorilla 9595
Nomascus leucogenys Northern white-cheeked gibbon 61853
Macaca mulatta Rhesus monkey 9544
Pan paniscus Bonobo 9597
Papio Anubis Olive baboon 9555
Callithrix jacchus White-tufted-ear marmoset 9483
Saimiri boliviensis boliviensis Bolivian squirrel monkey 39432
Mus musculus House mouse 10090
Otolemur garnettii Small-eared galago 30611
Pteropus alecto Black flying fox 9402
Ovis aries Sheep 9940
Cavia porcellus Domestic guinea pig 10141
Sus scrofa Pig 9823
Rattus norvegicus Norway rat 10116
Bos Taurus Cattle 9913
Oryctolagus cuniculus Rabbit 9986
Ailuropoda melanoleuca Giant panda 9646
Tupaia chinensis Chinese tree shrew 246437
Felis catus Domestic cat 9685
Cricetulus griseus Chinese hamster 10029
Heterocephalus glaber Naked mole-rat 10181
Ceratotherium simum simum Southern white rhinoceros 73337
Orcinus orca Killer whale 9733
Odobenus rosmarus divergens Pacific walrus 9708
Dasypus novemcinctus Nine-banded armadillo 9361
Chinchilla lanigera Long-tailed chinchilla 34839
Ictidomys tridecemlineatus Thirteen-lined ground squirrel 43179
Trichechus manatus latirostris Florida manatee 127582
Mustela putorius furo Domestic ferret 9669
Condylura cristata Star-nosed mole 143302
Octodon degus Degu 10160
Jaculus jaculus Lesser Egyptian jerboa 51337

3.3. Incongruity of AD data partitions

The p-values for the ILD test between each of the nucleotide and protein data partitions are listed in Tables 5 and 6, respectively. All p-values were 0.01, with the exception of the a2m/plau nucleotide data comparison, indicating very high degrees of incongruence. Likewise, performing the ILD test on the complete concatenation of all nine nucleotide and all nine protein data partitions yielded p-values of 0.01. Such highly incongruent data partitions were expected, since the goal of the study was to highlight and inspect variations between MP analyses over the same set of species.

Table 5.

ILD and RF′ distance between each pair of AD trees for nucleotide sequence data partitions.

A2M ACE APBB2 APP BLMH MPO NOS3 PLAU SORL1

A2M 0.379 0.405 0.417 0.417 0.444 0.472 0.430 0.404
ACE 0.01 0.380 0.430 0.379 0.379 0.379 0.391 0.391
APBB2 0.01 0.01 0.380 0.392 0.405 0.380 0.392 0.405
APP 0.01 0.01 0.01 0.391 0.379 0.417 0.379 0.430
BLMH 0.01 0.01 0.01 0.01 0.391 0.430 0.404 0.391
MPO 0.01 0.01 0.01 0.01 0.01 0.404 0.444 0.404
NOS3 0.01 0.01 0.01 0.01 0.01 0.01 0.417 0.417
PLAU 0.01 0.01 0.01 0.01 0.01 0.01 0.01 0.404
SORL1 0.01 0.01 0.01 0.01 0.01 0.01 0.01 0.01

Unshaded cells show ILD confidence interval (p-value: lower numbers indicate higher degree of incongruity), and shaded cells show RF′ distances. Values are unitless. Lower RF′ values are indicative of higher similarity between trees. Using this implementation of the Robinson–Foulds algorithm, values should only be quantitatively compared within the same table – scale varies based on the specific analysis.

Table 6.

ILD and RF′ distance between each pair of AD trees for protein sequence data partitions.

A2M ACE APBB2 APP BLMH MPO NOS3 PLAU SORL1

A2M 0.492 0.438 0.421 0.472 0.634 0.553 0.745 0.590
ACE 0.01 0.401 0.402 0.423 0.456 0.449 0.463 0.424
APBB2 0.01 0.01 0.435 0.400 0.418 0.426 0.443 0.436
APP 0.01 0.01 0.01 0.384 0.402 0.408 0.386 0.402
BLMH 0.01 0.01 0.01 0.01 0.454 0.465 0.461 0.454
MPO 0.01 0.01 0.01 0.01 0.01 0.468 0.562 0.526
NOS3 0.01 0.01 0.01 0.01 0.01 0.01 0.499 0.487
PLAU 0.01 0.01 0.01 0.01 0.01 0.01 0.01 0.584
SORL1 0.01 0.01 0.01 0.01 0.01 0.01 0.01 0.01

Unshaded cells show ILD confidence interval (p-value: lower numbers indicate higher degree of incongruity), and shaded cells show RF′ distances. Values are unitless. Lower RF′ values are indicative of higher similarity between trees. Using this implementation of the Robinson–Foulds algorithm, values should only be quantitatively compared within the same table – scale varies based on the specific analysis.

3.4. Simultaneous analysis

All phylogenetic analyses were rooted to Dasypus novemcinctus (nine-banded armadillo), which was determined to be the furthest diverged from humans using TimeTree [16]. Individual maximum parsimony trees for each nucleotide and protein data partition are shown in Figs. 2 and 3, respectively. Consensus SA trees based on the combination of the nine data partitions are shown in Figs. 4 and 5 (nucleotide and protein tree, respectively). Computed Branch Score (BS) values are shown on the consensus trees, and corresponding Partitioned Bremer Support (PBS) values are listed in Tables 3 and 4, respectively.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Fig. 2.

Individual nucleotide trees. Nucleotide trees are shown for each Alzheimer Disease gene included in this study, which show the drastic topological differences between individual AD nucleotide phylogenies within the AD gene network. Multiple taxa from the same taxonomic order are accordingly colored (Red = Primates, Blue = Rodentia, Green = Artiodactyla, Pink = Carnivora). Each tree is presented as a sub-figure, with the name of the human gene in the upper-left corner. Yellow highlighted region in (c) emphasizes the only instance of polytomy in the AD nucleotide trees. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Fig. 3.

Individual protein trees. Protein trees are shown for each Alzheimer Disease gene included in this study, which show the drastic differences between the individual AD protein phylogenies. Additionally, the high incidence of polytomies demonstrates the difficulty in resolving meaningful phylogenies for highly-conserved protein sequences. Multiple taxa from the same taxonomic order are accordingly colored (Red = Primates, Blue = Rodentia, Green = Artiodactyla, Pink = Carnivora). Each tree is presented as a sub-figure, with the name of the human gene in the upper-left corner. Red arrow in (d) indicates the node with the highest degree of polytomy in this study (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.).

Fig. 4.

Fig. 4.

Simultaneous analysis tree based on nucleotide data for included Alzheimer Disease genes. The tree is constructed by concatenating all nucleotide data partitions, and SA support values are labeled at individual nodes to demonstrate how well each data partition corroborates the tree topology at that node. It should also be noted that this tree follows commonly accepted evolutionary patterns more closely than most single-gene trees. Values at nodes are in the format x[y], where x is the Branch Score (BS) for that node, and y is a unique identifier (“node number”) corresponding to a row in Table 3, containing Partitioned Bremer Scores (PBS) for each nucleotide data partition corresponding to the given node on the simultaneous analysis tree. Yellow highlighted region emphasizes paraphyly of primates (which are shown in red font) (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.).

Fig. 5.

Fig. 5.

Simultaneous analysis tree based on protein data for included Alzheimer Disease genes. The tree is constructed by concatenating all protein data partitions, and SA support values are labeled at individual nodes to demonstrate how well each data partition corroborates the tree topology at that node. It should also be noted that this tree follows commonly accepted evolutionary patterns more closely than most single-gene trees, and polytomies are likewise eliminated. Values at nodes are in the format x[y], where x is the Branch Score (BS) for that node, and y is a unique identifier (“node number”) corresponding to a row in Table 4, containing Partitioned Bremer Scores (PBS) for each protein data partition corresponding to the given node on the simultaneous analysis tree. Yellow highlighted regions emphasize division of primates into two clades, and green highlighted region shows two species that might be considered possible model organisms for AD. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Table 3.

Nucleotide PBS values for each internal node on nucleotide simultaneous analysis tree for each data partition.

Node apbb2 nos3 plau sorl1 a2m blmh mpo ace app

  1 1929 −188 788 −2136 −111 775 54 134 −305
  2 −457 −173 −21 −604 −197 493 −49 1634 −68
  3 360 −175 −46 −739 −203 874 −45 128 −142
  4 360 −175 −46 −739 −203 874 −45 128 −142
  5 −2 −52 844 −505 181 −435 −101 368 −205
  6 −2 −52 844 −505 181 −435 −101 368 −205
  7 771 −9 −5 −1414 −231 897 −34 −25 67
  8 −1748 −209 1283 1781 31 −538 −200 347 −267
  9 360 −175 −46 −739 −203 874 −45 128 −142
10 771 −9 −5 −1414 −231 897 −34 −25 67
11 −46 −197 1889 −772 103 −529 −141 370 −436
12 830 −42 27 −1570 −138 −1040 510 339 1607
13 993 −110 152 −343 179 −1005 57 274 986
14 133 −14 21 76 38 −1942 366 223 1116
15 950 627 −1229 1429 −568 554 −365 452 −191
16 133 −14 21 76 38 −1942 366 223 1116
17 133 −14 21 76 38 −1942 366 223 1116
18 138 1203 23 80 39 −1989 368 211 1094
19 360 −175 −46 −739 −203 874 −45 128 −142
20 771 −9 −5 −1414 −231 897 −34 −25 67
21 771 −9 −5 −1414 −231 897 −34 −25 67
22 771 −9 −5 −1414 −231 897 −34 −25 67
23 1216 −158 757 −1448 −176 −69 19 21 −47
24 268 300 21 −552 −35 −1 5 −40 223
25 1086 −348 2094 −781 −267 −622 −173 183 −107
26 360 −175 −46 −739 −203 874 −45 128 −142
27 360 −175 −46 −739 −203 874 −45 128 −142
28 36 −7 −35 −63 10 22 1 22 35
29 360 −175 −46 −739 −203 874 −45 128 −142
30 360 −175 −46 −739 −203 874 −45 128 −142
31 −109 1 45 −502 −60 522 152 140 −126

Node numbers are ASAP2 assigned labels, reported next to BS score on SA tree (Fig. 4). Values for apbb2 and app gene partitions at node 13 are emphasized in bold due to their values being significantly higher than other partitions at that node.

Table 4.

Protein PBS values given for each internal node on protein simultaneous analysis tree for each data partition.

Node apbb2 nos3 plau sorl1 a2m blmh mpo ace app

1 6 10 12.5 1 4.5 4 3.5 −13.5 3
2 265 0 −37 11 −131 −6 −10 −69 0
3 263 6 −28 16 −112 −8 −7 −56 19
4 0 13 3 −4 −1 0 3 2 0
5 261 −1 −38 11 −131 −7 −12 −62 0
6 273 6 −15 18 −89 1 −14 −49 5
7 2 1 0 3 −5 3 −1 1 1
8 0 3 10 2 14 2 7 −31 39
9 −3 −1 −25 −21 −61 −583 −13 581 186
10 27 1 5 6 16 2 8 15 −16
11 0 7 0 0 1 3 12 1 0
12 0 548 0 0 2 0 12 0 −18
13 1 3 0 13 4 0 −3 2 0
14 4 1 0 9 17 1 −1 12 2
15 2 3 2 5 7 −1 6 9 1
16 13 6 3 11 24 5 13 14 1
17 4 3 8 9 18 −2 7 6 0
18 3 1 −1 −6 −4 −2 −1 2 17
19 3 1 −1 −6 −4 −2 −1 2 17
20 −16 10 7 20 40 7 2 −27 −5
21 −16 10 7 20 40 7 2 −27 −5
22 −3 −1 −25 −21 −61 −583 −13 581 186
23 −16 10 7 20 40 7 2 −27 −5
24 −4 2 1 0 4 2 0 5 2
25 −21 5 5 16 12 1 −8 10 −2
26 30 4 14 −1 24 −909 160 571 176
27 −2 6 2 3 5 −3 −1 7 1
28 2 13 1 3 −4 8 15 −33 40
29 12 27 31 42 120 7 90 68 −334
30 2 13 1 3 −4 8 15 −33 40
31 −3 −1 −25 −21 −61 −583 −13 581 186

Node numbers are ASAP2 assigned labels, reported next to BS score on SA tree (Fig. 5).

Each of the SA trees showed a distinct branching pattern, with all parent nodes on all trees strictly bifurcating. Furthermore, while some PBS values were negative (indicating that the data in a specific partition was not congruent with the consensus tree at that branch), all the BS values on the protein SA tree were positive. The nucleotide SA tree had positive BS scores at each node with no polytomies, suggesting that the genes selected for this study supported all the internal branches in the protein simultaneous analysis tree.

While the topologic organization of the SA trees generally followed canonical patterns of mammalian evolution there were some notable exceptions that received high levels of statistical support. In the SA nucleotide tree, most primates were grouped together into the monophyletic clade rooted at node 13, with the exception of Macaca mulatta (rhesus macaque), Callithrix jacchus (common marmoset) and Otolemur garnettii (northern greater galago) that each occurred distally from all other primates (Fig. 4). In the SA protein tree, primates were divided into two dis-tinct clades: (1) a monophyletic clade rooted at node 27, or (2) a paraphyletic clade rooted at node 8 that also included Sus scrofa (wild boar) and Jaculus jaculus (lesser Egyptian jerboa) (Fig. 5).

3.5. Comparison of trees using the Robinson–Foulds metric

The Robinson–Foulds metric was used to quantify the similarity between the generated trees. The pairwise comparisons between each of the nucleotide and protein data partitions are shown in Tables 5 and 6, respectively. Additionally, the RF distances for respective nucleotide and protein trees for a given partition as well as for the SA trees are shown in Table 7. The RF distances were used as input into CytoScape to visualize the relative relationship between both the nucleotide and the protein sequences for a given gene based on shared evolutionary history, shown in Fig. 6A (nucleotide sequences) and Fig. 6B (protein sequences). The resulting gene networks showed a tight clustering of MPO, A2M, NOS3, SORL1, and PLAU evolutionary patterns.

Table 7.

RF and RF′ values between corresponding nucleotide and protein trees for each gene.

Gene RF RF

APBB2 0.880 0.415
NOS3 0.593 0.553
PLAU 0.714 0.490
SORL1 0.900 0.407
A2M 0.559 0.572
BLMH 0.770 0.463
MPO 0.367 0.693
ACE 0.900 0.407
APP 0.837 0.433

Values are unitless. Higher RF values (and lower RF′ values) are indicative of higher similarity between trees. Using this implementation of the Robinson–Foulds algorithm, values should only be quantitatively compared within the same table – scale varies based on the specific analysis.

Fig. 6.

Fig. 6.

Clustering of Alzheimer Disease gene sequences based on similarity of tree topology. Based on clustering based on tree topology similarity for nucleotide (a) and protein (b), the graphs reveal a tight clustering of the A2M, MPO, PLAU, and NOS3 genes. This suggests that these genes may have a shared evolutionary history and thus may infer a yet-uncharacterized functional association related to the AD phenotype. Clustering of genes with most common evolutionary patterns are identified using gray ellipses. Nodes on the graph represent individual nucleotide gene partitions, and connecting edges are relative distances between trees, represented numerically as RF values (shorter edges correspond to more highly similar trees). RF values are listed in Table 5. RF is unitless.

3.6. Comparison of AD gene clustering to mtDNA genes

A benchmark for ASAP2 was accomplished by comparing the results for the set of AD genes to the corresponding results using sequences for mtDNA genes for the same set of 34 taxa. The RF gene clustering network diagrams (Fig. S1a for nucleotide sequences and Fig. S1b for protein sequences) for the mtDNA genes alone show relatively tight clustering, with the slight exception of the ND4L gene in the protein analysis and the ND6 gene in the nucleotide analysis. Overall, this finding recapitulates that mtDNA genes may be suitable as taxonomic markers or for species identification. When both mtDNA and AD gene partitions were used simultaneously in ASAP2 (Fig. S1c and S1d, for nucleotide and protein sequences, respectively), the network diagrams reveal that the mtDNA genes cluster closely with a number of AD genes that themselves form a cluster – namely MPO, A2M, NOS3, SORL1, and PLAU.

Additionally, the SA trees (Supplementary Data) for the mtDNA analysis demonstrated a striking departure from the topology shown in the AD SA trees. In agreement with mtDNA genes being highly conserved throughout the tree of life, trees constructed from mtDNA gene sequences closely resemble the actual evolutionary histories of the species in the trees [56], while the AD SA trees exhibited striking departures from the expected topology. This strengthens the proposition that abnormal patterns of evolutionary conservation may help distinguish and characterize the molecular basis of complex genetic diseases. These promising findings suggest that future work is needed to include the use of ASAP2 to study additional complex diseases, which will help further validate the methodology proposed here as well as better support the proposed findings relative the AD genes included in this study.

3.7. Benchmarking ASAP2 using randomly selected query genes

A second benchmark was established by running an ASAP2 analysis against eight randomly selected genes, each of which is implicated in susceptibility to a different complex genetic disorder (Table 8). In searching for non-human sequences corresponding to these eight genes, a different set of species was identified, each member of which possesses at least one putative orthologue to each of the human query genes. The RF values generated using both the protein and translated nucleotide sequences demonstrated far less conservation of a common evolutionary pattern than with the AD analysis. Whereas the AD nucleotide analysis resulted in RF values ranging from 0.750 to 0.969 (mean: 0.907 ± 0.055), the analysis of the benchmark of eight randomly selected genes resulted in RF values (Table 9) ranging 0.877–0.970 (mean: 0.952 ± 0.026). Likewise, the AD protein analysis resulted in RF values ranging from 0.294 to 0.957 (mean: 0.770 ± 0.149), and the analysis of the benchmark of eight randomly selected genes resulted in RF values (Table 10) ranging 0.895–0.968 (mean: 0.944 ± 0.025). Larger RF distances indicate more highly dissimilar trees, corresponding to less evolutionary conservation. Overall, this demonstrates a higher degree of variability in the AD analysis, which corroborates the claim that certain members of the AD gene family have been conserved together over evolutionary time.

Table 8.

Randomly selected genes for negative control benchmark, with corresponding complex genetic disorder(s) included. GenBank/GenPept query GIs for each gene are available in Supplementary Data.

Gene Common name Complex disease

abcc8 Sulfonylurea Receptor Type 2 Diabetes mellitus
adrb2 Beta-2-Adrenergic Receptor Obesity/Asthma
app Amyloid-β Precursor Protein Alzheimer Disease
brca2 Breast Cancer Type 2 Susceptibility Protein Various cancers
park2 Parkin Parkinson Disease
ldlr Low Density Lipoprotein Receptor Hypercholesterolemia
calcr Calcitonin Receptor Osteoporosis
nod2 Nucleotide-Binding Oligomerization Domain Protein 2 Inflammatory Bowel Disease

Also in AD analysis.

Table 9.

RF and RF′ distance between each pair of second benchmark trees for nucleotide sequence data partitions.

ABCC8 ADRB2 APP BRCA2 CALCR LDLR NOD2 PARK2

ABCC8 0.381 0.380 0.380 0.380 0.380 0.379 0.379
ADRB2 0.966 0.381 0.380 0.380 0.381 0.393 0.393
APP 0.968 0.966 0.380 0.380 0.380 0.379 0.379
BRCA2 0.968 0.967 0.968 0.404 0.392 0.416 0.379
CALCR 0.968 0.967 0.968 0.906 0.392 0.403 0.391
LDLR 0.968 0.966 0.968 0.937 0.937 0.404 0.379
NOD2 0.969 0.934 0.969 0.877 0.908 0.906 0.379
PARK2 0.969 0.934 0.969 0.969 0.938 0.969 0.970

Unshaded cells show RF distances, and shaded cells show RF′ distances.

Table 10.

RF and RF′ distance between each pair of second benchmark trees for protein sequence data partitions.

ABCC8 ADRB2 APP BRCA2 CALCR LDLR NOD2 PARK2

ABCC8 0.380 0.385 0.380 0.382 0.380 0.380 0.381
ADRB2 0.967 0.384 0.391 0.396 0.380 0.404 0.393
APP 0.955 0.957 0.384 0.387 0.384 0.400 0.385
BRCA2 0.968 0.938 0.958 0.409 0.391 0.403 0.380
CALCR 0.962 0.927 0.949 0.895 0.395 0.409 0.397
LDLR 0.967 0.968 0.957 0.938 0.929 0.403 0.380
NOD2 0.968 0.906 0.917 0.909 0.895 0.908 0.380
PARK2 0.966 0.933 0.955 0.968 0.925 0.967 0.968

Unshaded cells show RF distances, and shaded cells show RF′ distances.

As with the AD analysis, force-directed graphs of the RF results were generated for both the nucleotide (Fig. S2a) and protein (Fig. S2b) data partitions in the benchmark. However, the graphs for the benchmark analysis do not adequately demonstrate a lack of clustering as would be expected by the numerical data. This is due to the graph being projected onto a 2-dimensional plane, resulting in an apparent distortion of some of the edge lengths. Regardless, the numerical data show a distinct pattern of large RF distances between trees, as opposed to the formation of tight “clusters” within the AD and mtDNA analyses, which strongly supports the hypothesis that these “clusters” predict a close functional relationship between certain genes in the given framework.

When ASAP2 was run using the union of all AD data partitions and the eight randomly-selected data partitions (full data available in Supplementary Data), the resultant force-directed graphs of RF distances (Fig. S2c and S2d) depicted the AD genes as a distinct set segregated within the graph, forming another tight “cluster” among the sparsely located random genes. These observations suggest that the results of the AD analysis cannot be quantitatively compared against the random gene analysis when they are run separately (due to support values being relative). Additionally, this suggests that ASAP2 may be used as a potential as a tool for identifying complex disease genes that have a shared history, in addition to being useful for analyzing relationships within an already known network of shared phenotypic functionality.

4. Discussion

Alzheimer Disease (AD), a complex neurodegenerative disorder, is the most common form of dementia, accounting for between 70% and 90% of diagnosed cases of dementia [23,24]. Worldwide, more than 24 million people are estimated to have AD, with estimates exceeding 80 million to be affected over the next 30 years [25]. The pathognomonic histological finding that originally defined AD is the presence amyloid plaques in cortical brain tissues [2628]. The plaques arise from the production, and eventual extracellular precipitation, of fibrillar aggregates of amyloid-β peptides causing disruption of the neural architecture and induction of inflammation [26]. Clinically, AD is characterized by progressive memory loss and cognitive decline, leading to general functional impairment [2732]. However, the precise etiology of AD remains elusive. Amyloid plaques have been shown to differ widely in manifestation with regards to protein composition, structural characteristics, and prevalence [33,34]. Despite advances in predicting the presence of the disease based on symptoms and diagnostic imaging, a definitive diagnosis of AD can only be made after an autopsy of the affected brain after death [29,30]. As such, researchers are faced with significant difficulties both in studying the progression of the disease as well as identifying potential therapeutic techniques to prevent or treat the disease.

Further compounding the difficulty in identifying causes of AD is the fact that the majority of cases do not follow a well-characterized pattern of heritability, even though susceptibility to AD is widely considered to have a genetic basis [33,3537]. Familial forms of AD have been identified resulting from single or double gene lesions leading to increased amyloid plaque burden, but these account for less than 5% of the total cases of AD [38]. A wide range of human genes have been linked to susceptibility at various stages of life, suggesting sporadic AD has a multifactorial, complex genetic component [38]. The complex, multi-gene aspect of AD, taken in combination with the potential impact of shared evolution of its associated genes offers an excellent candidate disease to explore the potential utility of SA techniques.

The use of ASAP2 enabled the generation of a first view of a gene network based on an integrated phylogeny of AD-associated genes. The results suggest that SA techniques may have utility in the development of network perspectives of large-scale studies that aim to model the evolutionary development, transmission, and interaction of disease associated gene sets. As a new paradigm for studying complex disease genes, the network approach presented here relies on the assumption that individual genes may indeed diverge substantially from what one might expect at the species level. Indeed, the Incongruence Length Difference (ILD) test indicates that, in a conventional phylogenetic analysis, the AD genes used here should not be used to impute traditional hypotheses of species-level evolution. However, individual gene phylogenies may provide unique information about the functionality of the genes from which they are constructed relative to the associated disease within a network medicine context. Thus, a combined analysis of complex disease genes, which may be counter-intuitive from traditional phylogenetic practices (where genes are generally only be combined when reflecting common topologies), can be used to identify potential shared evolutionarily shared relationships that would otherwise not be discovered. The divergences from conventional phylogenetic analyses make the methodology described in this study particularly amenable to network analyses, where the correlations and relations are based on overall topological similarity that reflect concordance of evolutionary patterns across multiple genes for a common set of taxa.

4.1. ASAP2 function

ASAP2 consolidates a set of phylogenetic techniques into a single pipeline of Ruby scripts designed to expose higher-order quantitative relationships between genes not visible through more traditional single-gene based analyses. Implementing phylogenetic techniques often requires a significant amount of manual data curation that is both labor- and time-intensive, especially in the context of disease genes. ASAP2 was designed as a flexible automated tool that performs these tasks with minimal required intervention beyond entering an initial set of GenBank identifiers. ASAP2 thus supports the ability to do large-scale phylogenetic analyses in a tractable manner. ASAP2 execution time is generally Θ(n2) with respect to both the number of data partitions and the average length of sequences, but overall runtime can be considerably reduced based on available computational resources for individual BLAST queries, sequence alignments, and phylogenetic tree search routines. The data structures produced by ASAP2 were designed to be user-readable and manually editable during a given analysis. This supports the ability to adjust subsequent analyses based on results generated at any point along the analysis pipeline. It is important to note that orthology determination based on sequence similarity alone may result in unexpected sequences being included in an analysis (e.g., it has been shown sequence similarity methods such as BLAST do not always return the closest related sequences [57]. To address this potential issue, ASAP2 enables the user to create a set of data partitions using their own preferred search strategy. Nonetheless, in this feasibility study that focused on AD genes, it was felt that it would be of utility to demonstrate the potential of an automated search strategy, as both a proof of concept and as a means to standardize the selected orthologues. The selected putative orthologue sequences were manually verified to ensure that no irrelevant sequences were introduced into a given data partition. Uniformly across the AD analysis, no sequences identified using the described BLAST algorithm were found to be irrelevant in the scope of their respective query sequences.

The original Perl version of ASAP [41] required a prior file containing sequences that was then aligned using MUSCLE and the phylogenetic analysis was subsequently executed using PAUP* [58]. ASAP also allowed for the inclusion of pre-aligned or morphological data. By contrast, ASAP2 was developed in Ruby, and uses MUSCLE based alignment with the subsequent phylogenetic analyses done in TNT (which is freely licensable, unlike PAUP*). Additionally, ASAP2 was specifically designed to work exclusively with molecular data available from GenBank/GenPept, requiring only that the user provide an initial set of Accession numbers.

In this study, the utility of ASAP2 was demonstrated by performing analyses on a discrete set of pre-identified sequences within the context of developing a gene network to enhance network medicine knowledge for disease genes. However, the script may also be used for a number of large-scale multi-gene phylogenetic investigations. For example, one could use ASAP2 to study whole genomes with the goal to identify essential, evolutionarily conserved genes in groups of species [9,21]. Conceptually, by adjusting inclusion thresholds for the BLAST search mechanism and by specifying different options for the sequence alignment procedure, analyses could be expanded for constructing SA networks with respect to entire gene families.

Although this study made specific use of MUSCLE for multiple sequence alignment, ASAP2 includes the ability to alternatively use MAFFT or Clustal Omega. It should be noted that it might be possible for future versions of ASAP2 to include the ability to perform phylogenetic analyses within a maximum likelihood framework, although the majority of simultaneous analysis methods to date have relied upon maximum parsimony, and new methods of computing SA values would need to first be devised that are analogous to the ones used in maximum parsimony (e.g., for assessing deviations from actual evolution).

4.2. Putative orthologue sequence identification

Based on an initial OMIM query for Alzheimer Disease, orthologues for 34 species were identified across nine disease genes. In addition to the recursive BLAST based approaches implemented by ASAP2, there are specific orthologue resources that could have also been searched to identify orthologous sequences for each of the nine disease genes. For example, inParanoid [59] and OrthoMCL [60] had eight and 13 species spanning the nine genes, respectively. Interestingly, in identifying the set of species that contains putative orthologues for each of the AD genes through each of the three identification methods, ASAP2 and inParanoid identified only mammalian species, while OrthoMCL identified a set of organisms that included several non-mammalian species, including Danio rerio (zebrafish), Takifugu rubripes (tiger blowfish), Tetraodon nigroviridis (spotted green pufferfish), and Gallus gallus (chicken). Additionally, both inParanoid and OrthoMCL identified the species Canis familiaris (dog) and Equus caballus (horse), while ASAP2 did not. The differences in orthologue identification may be due to the conservative filtering parameters used for BLAST queries in ASAP2 that were tuned to ensure a high degree of similarity between sequences and to minimize the possibility of random homologies (as implicated by using an E-value of 1.0 × 10−256). Neither inParanoid nor OrthoMCL identified the same set of additional species across all nine genes that were the focus of this study. ASAP2 does allow for the inclusion or removal of sequences to increase or reduce the taxonomic diversity of a given analysis immediately following the BLAST analyses; however, since no additional taxa were identified uniformly across the nine genes of interest by either OrthoMCL or inParanoid, no such modification of taxon diversity was performed in this study. Additionally, future studies may benefit from starting with a wider empirical set of genes or with parameters for the recursive BLAST strategy that are tuned to higher E-values that could lead to greater taxonomic diversity.

The ASAP2 sequence identification method does not have the ability to definitively distinguish between orthologues and paralogues. Distinguishing between the two is an ongoing challenge. However, due to the low expect value currently used in the ASAP2 BLAST search strategy, only highly similar sequences are preserved (and, if more than one sequence is found for a given species, only the ‘best’ sequence will be kept). Therefore, regardless of whether an identified sequence is an orthologue or an in-paralogue to a human sequence, the phylogenetic analysis of the sequences should reveal how that sequence has evolved across the set of species. In this way, ASAP2 provides a systematic mechanism to leverage both similarity-based approaches for sequence identification alongside phylogenetic approaches for evolutionary study of those similar sequences.

4.3. Phylogenetic analysis

The individual gene genealogies returned by the phylogenetic analyses of individual gene partitions were not expected to imply speciation or specific transmission of genetic elements. “Evolutionary conservation” of certain groups of genes in this study refers to the conservation of gene function within a constrained group of other genes. Effectively, these individual phylogenies highlight constrained variations in the genome that are at a number of magnitudes higher resolution than easily discernable traits, with the exception of disease phenotypes arising due to a malfunction within the given set of genes.

The TNT analyses used by ASAP2 were optimized to only include the most unambiguous groupings. As such, the TNT scripts produce fewer trees, but the likelihood of the trees reflecting evolutionary history is correspondingly more reliable. The final consensus tree represents a likely model of evolutionary transmission of the group of Alzheimer Disease genes studied, and the partitioned Bremer support values indicate the degree to which each gene fits the predicted pattern of evolution. The partitioned Bremer values may also be used to identify genes or species in a study that did not (for one reason or another) follow a similar pattern of transmission as the others. Topologically, the SA protein tree in this study exhibited a small number of groupings that differ from the accepted model of mammalian evolution, notably the separation of primates into two distinct clades (Fig. 5, highlighted in yellow). On the SA nucleotide tree, the paraphyletic grouping of some primates (Fig. 4, highlighted in yellow) also merits scrutiny since this suggests that the genes included in this study deviate from taxonomically accepted evolution. Since only the most similar GenBank-catalogued sequence was retained for each species included in this study, there is only a modest risk of accidental selection of a paralogue instead of an orthologue. It is important to note that the aggregate PBS values for these different nodes were low and may be subject to topologic changes with the addition of more partitions. However, these “alternative” placements of certain primate species in the SA tree might also be explained by a reversion to an ancestral state for a particular disease gene. In this instance, the “state” is the pattern of interaction between the disease genes included in the study – the SA trees can be thought of as a phylogenetic analysis of the possible network in which some of the AD genes may function, and placement on the tree represents nonspecific alterations to that network. Likewise, an “ancestral state” is the structure and genetic landscape of this possible network in a common ancestor to the organisms on the tree. Therefore, this type of deviation from taxonomic evolution represents potential evolutionary divergence of this theorized Alzheimer Disease gene network within isolated species. The presence of these types of alternate evolutionary patterns suggests a potential differential susceptibility of species in the development of AD. For example, the APBB2 and APP PBS values at node 13 in the nucleotide SA tree (Table 3, emphasized in bold font) are significantly higher than for other partitions: 993 and 986, respectively (compared to average values of 131.4 ± 620.1). These values suggest a potential interaction (based on a strongly corroborated evolutionary history) between the protein products of APBB2 and APP in primates. Building on the known interaction between APBB2 and APP in H. sapiens, exploration of the polymorphisms in these genes in M. mulatta and O. garnettii may elucidate the potential for differences in functional interactions. Such further exploration of these types of findings, especially relative to critical synapomorphic characters, could therefore yield valuable data regarding the evolutionarily important functional or potentially interacting sites for a given disease gene.

The individual data partition protein trees had a high incidence of polytomy, which is when more than two species branch off of a single node. This is generally considered uninformative in determining ancestry, as there are not enough data to determine whether species branching off of the same node are more or less closely related. However, these observations highlight the evolutionary conservation of fundamental protein sequences over many related organisms [6165]. APP, one of the central genes in Alzheimer Disease research, displays the most drastic examples of polytomy, with 17 branches underneath one node alone (Fig. 3d, node emphasized with red arrow). This reinforces previous studies showing a high degree of conservation of the APP gene family over time [6670].

While the protein phylogenies demonstrate conservation of structure across multiple species, the nucleotide sequences generate trees allowing a more precise elucidation of ancestry. Since nucleotide sequences can have differences that do not affect protein structure or function due to the degeneracy of the genetic code, rates of change in nucleotide sequences are more closely tied to evolutionary time [49,7173]). Among the individual partition nucleotide trees, only the APBB2 tree has an occurrence of more than two branches rooted at a single parent node. The branch generated at this node contains four species of very closely related great apes (Nomascus leucogenys, Gorilla gorilla gorilla, Pan troglodytes, and Pan paniscusFig. 2c, highlighted in yellow). This suggests that the nucleotide sequences corresponding to APBB2 in each of these species are so similar that a more descriptive phylogenetic relationship between them cannot be determined, which underscores the fact that APBB2 is highly conserved among closely related species.

Determination of the distance between individual trees prior to constructing a consensus tree can help to preliminarily identify clustering patterns among specific genes prior to constructing a consensus tree [74]. Additionally, once a consensus is reached, these distances can be used to explain the strength of the support for the SA tree and generate representations of the gene network [75]. While multiple methods may be used to evaluate the distance between trees consisting of the same set of taxa, this study used the Robinson–Foulds (RF) distance [43]. The RF distance between two trees is defined by the sum of the number of data partitions implied by one, but not both, of the trees. A variety of algorithms exist for computing RF distance [76,77], and an optimal method is usually selected on the basis of algorithmic complexity and worst-case running time [51,78,79]. In this study, a gene network was constructed based on tree topology similarity using RF (transformed to RF, which converts RF values onto a scale where higher values correspond to less similarity).

On examination of the gene network for the Alzheimer Disease genes used in this study (Figs. 2 and 3), a tight clustering of oxidative stress genes was observed with the gene for plasminogen activator (PLAU) and a member of the sortilin related receptor gene family (SORL1). While SORL1 has been found to have an important association with Alzheimer Disease and oxidative stress genes are involved in the unfolded protein response associated with increased amyloid formation, a relationship between these genes has not been shown before [80,81]. This type of association is not observable using single pathway experiments or phylogenetic methods that do not incorporate an SA approach. Further investigation will be needed to understand the nature of this clustering within the gene network.

In interpreting the SA values and the gene-clustering network diagrams, two important assumptions regarding network medicine and complex genetic disorders should be noted: (1) Although many genetic components may be identified for a given disease, it should not be assumed that all of these factors are involved in the same pathway – there may be multiple pathways involved; and (2) Co-evolution of genes does not necessarily imply conservation due to related function, and vice versa – co-evolution may be a result of factors as simple as two genes being in close proximity on the same chromosome. These assumptions underscore how the information provided by SA techniques and the broader disease implications may be potentially misleading or incomplete. Nevertheless, an SA methodology might still provide insights to potential targets for clinical therapies that would not have been highlighted using more traditional single-gene based analytic approaches.

A final aspect of this study is that it further highlights the fact that choice of model organism is paramount for the study of complex disease. The relatively short lifespan of M. musculus and malleability of the murine genome has led to an explosion of experimental approaches centered on manipulation of genes thought to be involved in human disease [82,83]. However, especially with relation to complex diseases, alternative model organisms need to be considered [84]. The recent increase in biological systems data and continued growth in bioinformatics methodologies for analyzing these data may allow for the development of more data driven choices of model organisms for complex diseases. For example, based on the preliminary findings of this study of the shared evolution of a limited set of genes thought to influence AD susceptibility in humans, the SA consensus trees suggest that S. scrofa (pig) and J. jaculus (jerboa) may be more suitable model organisms than murine species (Fig. 5, highlighted in green).

4.4. Verification and benchmarking of ASAP2 results

Similar to how phylogenetic analyses can be difficult to fully validate, since a given phylogenetic analysis is often perceived as a hypothesis, the evaluation of SA techniques can be difficult. Further challenging the validation of the approach presented here is the paradigm shift of identifying potential co-evolutionary relationships relative to a complex disease. The validation of the approach developed in this study was thus done by examining the characteristic of a gene network developed from two types of genes: (1) those that have been known to have co-evolved (a positive control; should cluster); and, (2) a random selection of disease genes that have not been described in a co-evolutionary context (negative control; should not cluster).

Based on the premise that mitochondrial genes evolve at a distinguishable rate between different species (implying that phylogenetic analyses of mtDNA genes often recapitulate taxonomy) [85,86], this study explored whether the AD gene clustering might be meaningful relative to taxonomy (thus implying evolutionary conservation of clustered genes). This strategy for benchmarking the ASAP2 analysis of AD genes using mitochondrial DNA sequences reinforces the principle that SA techniques may be used to study deviations in evolutionary conservation among isolated genes. In this study, ASAP2 revealed that the included mitochondrial genes showed a pattern of clustering (based on RF values) that was highly similar to the clustering of a subset of AD genes that could potentially be involved in previously unpredicted relationships via metabolic stress pathways. While this does not definitively prove that a causal relationship exists between the aforementioned AD genes, it does demonstrate that ability of ASAP2 to highlight possible relationships of interest for complex disease genes that might warrant further investigation.

By contrast, genes that operate in different metabolic and structural protein pathways are not likely to have many strong functional links between the genes associated with other disease related pathways [87]. The SA analysis and resultant gene network of the eight randomly selected disease genes did not show any apparent clustering, especially relative to either the AD genes or the mtDNA genes. The results of this second benchmark support the implied corollary to the hypothesis of this study: Strongly dissimilar patterns of evolutionary conservation emerge when comparing phylogenetic analyses of genes that are not closely related or not likely to interact with one another. In sequence, this strengthens the premise that simultaneous analysis may yield a new type of information that is rooted in tangible, evolutionarily significant (and potentially yet undiscovered) pathophysiological events.

The results of this study suggest that SA techniques can provide metrics that complement network medicine approaches that have been used to explore disease genes. For example, the resulting gene network from this study can be used to further add structure to the Diseasome [88], which is a network of disease genes that are ascertained from inferred relationships as reported in OMIM. To this point, the tightest cluster of genes observed in the SA-based gene network developed in this study (A2M, PSEN2, APBB2, PLAU, BLMH, NOS3, and PAXIP1) are all uniquely connected to Alzheimer Disease (i.e., they have no other disease connections in the Diseasome); the two genes that are observed as having slightly less correlation (APP and ACE) are linked with other diseases (APP with Schizophrenia and Amyloidosis; and, ACE with Renal Tubular Dysgenesis, Progression of SARS, and Myocardial Infarction). The results of this study therefore suggest that, in addition to the uniqueness that has been shown by the Diseasome, there is a strong correlative history of these genes. ASAP2 thus offers an evolutionary framework by which one might study disease genes that can further warrant concurrent investigation of putatively associated genes relative to a shared disease phenotype.

4.5. Identifying new complex disease genes relationships using ASAP2

The AD genes selected for use throughout this study were identified from previous investigations. ASAP2 was designed to take a known set of related genes and compare their gene genealogies so that previously unidentified interactions might make themselves apparent. The strong segregation of AD genes in the context of including a random set of complex disease genes in ASAP2 suggests that if a number of genes for a given complex disease cluster together, other genes within that cluster may be related to the given complex disease. Although this does not in any way absolutely confirm that relationships between putatively related disease genes exist, it provides a target for future studies that can aim to resolve potential relationships.

The originally described method provides a forward genetics approach to complex diseases (i.e., identifying genotypic functionality from phenotype data), while the retooling of ASAP2 could be described as a reverse genetics approach – starting with widely distributed genomic data and linking a phenotype to a subset of those data. In consort, these two approaches to using ASAP2 (and similar such tools) could define a highly effective means of resolving the array of incongruities that are often observed between genotypic (genomic) and phenotypic (complex disease) data sources, eventually defining new therapies and treatments for diseases.

The orthologous genes and their associated taxa were chosen using a strictly bioinformatics approach; there was no a priori selection of taxa that represent specific phenotypes (e.g., those that exhibit characteristics of AD similar to human manifestation of the condition) or to ensure complete taxonomic representation. Instead, the approach was chosen to include all available data based solely on a stringent sequence similarity criterion. For translational utility, it would be essential for future studies to include some selection of taxa that represent different observable phenotypes and then seek to identify patterns of evolution that may connote protective or causative genetic characteristics. Additionally, further validation of the approach presented in this study will undoubtedly require the verification of its applicability for other complex diseases. This will be essential to resolve the potential issue that the gene network characteristics observed in this study may be unique to Alzheimer Disease. Future work will thus be focused on the application of ASAP2 in the context of other complex diseases, including those that have been characterized in the Diseasome. It is important to consider that the identification of orthologues across multiple genes for multiple diseases for the same set of taxa is a formidable challenge. Because of an artifact that AD disease genes are heavily studied, it was possible for ASAP2 to easily identify orthologues (based on a strict sequence similarity cut-off) for 34 taxa. However, it was not possible to readily identify another set of orthologues for another disease for the same 34 taxa. A significant future enhancement of the SA approach, therefore, would require either the accommodation of different sets of taxa across different diseases or identification of a core set of taxonomically diverse taxa that can be used across all the diseases of interest. The identification of taxa that are missing orthologous gene sequences could also be used to guide future molecular sequencing efforts.

5. Conclusion

Phylogenomic studies using simultaneous analysis techniques are positioned to become more commonplace as increasing amounts of genomic data are available across the spectrum of life and systematically available through resources such as GenBank. Here, an automated tool (ASAP2) is presented with the intent of enabling researchers to leverage these data to support studies that aim to unveil potentially evolutionarily significant relationships that can complement other network approaches. The application of ASAP2 to a set of nine genes associated with Alzheimer Disease demonstrated a potentially important clustering of genes that corroborates other network approaches, and also suggests that there may be an evolutionarily meaningful reason for their correlation with the Alzheimer Disease phenotype. The results thus suggest that the methodology presented here may be used to add additional, evolutionarily informed structure to gene network studies.

Supplementary Material

supplementary information 2
supplementary information

Appendix A. Supplementary material

Supplementary data associated with this article can be found, in the online version, at http://dx.doi.org/10.1016/j.jbi.2015.01.002.

References

  • [1].Badano JL, Katsanis N. Beyond Mendel: an evolving view of human genetic disease transmission. Nat Rev Genet 2002;3:779–89. [DOI] [PubMed] [Google Scholar]
  • [2].Velinov M, Slaugenhaupt SA, Stoilov I, Scott CI, Gusella JF, Tsipouras P. The gene for achondroplasia maps to the telomeric region of chromosome 4p. Nat Genet 1994;6:314–7. [DOI] [PubMed] [Google Scholar]
  • [3].Kerem B, Rommens JM, Buchanan JA, Markiewicz D, Cox TK, Chakravarti A, et al. Identification of the cystic fibrosis gene: genetic analysis. Science 1989;245:1073–80. [DOI] [PubMed] [Google Scholar]
  • [4].Rees DC, Williams TN, Gladwin MT. Sickle-cell disease. Lancet 2010;376:2018–31. [DOI] [PubMed] [Google Scholar]
  • [5].Sillén A, Forsell C, Lilius L, Axelman K, Björk BF, Onkamo P, et al. Genome scan on Swedish Alzheimer’s disease families. Mol Psychiatry 2006;11:182–6. [DOI] [PubMed] [Google Scholar]
  • [6].Yonan AL, Alarcón M, Cheng R, Magnusson PKE, Spence SJ, Palmer AA, et al. A genomewide screen of 345 families for autism-susceptibility loci. Am J Hum Genet 2003;73:886–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Li Y, Hollingworth P, Moore P, Foy C, Archer N, Powell J, et al. Genetic association of the APP binding protein 2 gene (APBB2) with late onset Alzheimer disease. Hum Mutat 2005;25:270–7. [DOI] [PubMed] [Google Scholar]
  • [8].Newton-Cheh C, Johnson T, Gateva V, Tobin MD, Bochud M, Coin L, et al. Genome-wide association study identifies eight loci associated with blood pressure. Nat Genet 2009;41:666–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Klein BA, Tenorio EL, Lazinski DW, Camilli A, Duncan MJ, Hu LT. Identification of essential genes of the periodontal pathogen Porphyromonas gingivalis. BMC Genomics 2012;13:578. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Estruch R, Ros E, Salas-Salvadó J, Covas M-I, Corella D, Arós F, et al. Primary prevention of cardiovascular disease with a Mediterranean diet. N Engl J Med 2013;368:1279–90. [DOI] [PubMed] [Google Scholar]
  • [11].Regan K, Wang K, Doughty E, Li H, Li J, Lee Y, et al. Translating Mendelian and complex inheritance of Alzheimer’s disease genes for predicting unique personal genome variants. J Am Med Inform Assoc 2012;19:306–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Sarkar IN. A vector space model approach to identify genetically related diseases. J Am Med Inform Assoc 2012;19:249–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Barabási A-L, Gulbahce N, Loscalzo J. Network medicine: a network-based approach to human disease. Nat Rev Genet 2011;12:56–68. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Swiderski DL, Zelditch ML, Fink WL. Why morphometrics is not special: coding quantitative data for phylogenetic analysis. Syst Biol 1998. [PubMed] [Google Scholar]
  • [15].Zharkikh A Estimation of evolutionary distances between nucleotide sequences. J Mol Evol 1994;39:315–29. [DOI] [PubMed] [Google Scholar]
  • [16].Hedges SB, Dudley J, Kumar S. TimeTree: a public knowledge-base of divergence times among organisms. Bioinformatics 2006;22:2971–2. [DOI] [PubMed] [Google Scholar]
  • [17].Dagan T Phylogenomic networks. Trends Microbiol 2011;19:483–91. [DOI] [PubMed] [Google Scholar]
  • [18].Layeghifard M, Peres-Neto PR, Makarenkov V. Inferring explicit weighted consensus networks to represent alternative evolutionary histories. BMC Evol Biol 2013;13:274. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Nixon KC, Carpenter JM. On simultaneous analysis. Cladistics 1996;12:221–41. [DOI] [PubMed] [Google Scholar]
  • [20].Gatesy J, O’Grady P, Baker RH. Corroboration among data sets in simultaneous analysis: hidden support for phylogenetic relationships among higher level artiodactyl taxa. Cladistics 1999;15:271–313. [DOI] [PubMed] [Google Scholar]
  • [21].Rokas A, Williams BL, King N, Carroll SB. Genome-scale approaches to resolving incongruence in molecular phylogenies. Nature 2003;425:798–804. [DOI] [PubMed] [Google Scholar]
  • [22].Baker RH, Yu X, DeSalle R. Assessing the relative contribution of molecular and morphological characters in simultaneous analysis trees. Mol Phylogenet Evol 1998;9:427–36. [DOI] [PubMed] [Google Scholar]
  • [23].Ferri CP, Prince M, Brayne C, Brodaty H, Fratiglioni L, Ganguli M, et al. Global prevalence of dementia: a Delphi consensus study. Lancet 2005;366:2112–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Mayeux R, Stern Y. Epidemiology of Alzheimer disease. Cold Spring Harb Perspect Med 2012;2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Ballard C, Gauthier S, Corbett A, Brayne C, Aarsland D, Jones E. Alzheimer’s disease. Lancet 2011;377:1019–31. [DOI] [PubMed] [Google Scholar]
  • [26].Hardy J, Selkoe DJ. The amyloid hypothesis of Alzheimer’s disease: progress and problems on the road to therapeutics. Science 2002;297:353–6. [DOI] [PubMed] [Google Scholar]
  • [27].Nelson PT, Braak H, Markesbery WR. Neuropathology and cognitive impairment in Alzheimer disease: a complex but coherent relationship. J Neuropathol Exp Neurol 2009;68:1–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Nelson PT, Alafuzoff I, Bigio EH, Bouras C, Braak H, Cairns NJ, et al. Correlation of Alzheimer disease neuropathologic changes with cognitive status: a review of the literature. J Neuropathol Exp Neurol 2012;71:362–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Blacker D, Albert MS, Bassett SS, Go R. Reliability and validity of NINCDS-ADRDA criteria for Alzheimer’s disease: the National Institute of Mental Health Genetics Initiative. Arch Neurol 1994. [DOI] [PubMed] [Google Scholar]
  • [30].Dubois B, Feldman HH, Jacova C, DeKosky ST. Research criteria for the diagnosis of Alzheimer’s disease: revising the NINCDS–ADRDA criteria. Lancet Neurol 2007. [DOI] [PubMed] [Google Scholar]
  • [31].Querfurth HW, LaFerla FM. Alzheimer’s disease. N Engl J Med 2010;362:329–44. [DOI] [PubMed] [Google Scholar]
  • [32].Karran E, Mercken M, De Strooper B. The amyloid cascade hypothesis for Alzheimer’s disease: an appraisal for the development of therapeutics. Nat Rev Drug Discov 2011;10:698–712. [DOI] [PubMed] [Google Scholar]
  • [33].Silverman JM, Smith CJ, Marin DB, Birstein S, Mare M, Mohs RC, et al. Identifying families with likely genetic protective factors against Alzheimer disease. Am J Hum Genet 1999;64:832–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Lu J-X, Qiang W, Yau W-M, Schwieters CD, Meredith SC, Tycko R. Molecular structure of β-amyloid fibrils in Alzheimer’s disease brain tissue. Cell 2013;154:1257–68. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Sjögren T, Sjogren H, Lindgren AG. Morbus Alzheimer and morbus Pick; a genetic, clinical and patho-anatomical study. Acta Psychiatr Neurol Scand Suppl 1952;82:1–152. [PubMed] [Google Scholar]
  • [36].McIlroy SP, Crawford VL, Dynan KB, McGleenon BM, Vahidassr MD, Lawson JT, et al. Butyrylcholinesterase K variant is genetically associated with late onset Alzheimer’s disease in Northern Ireland. J Med Genet 2000;37:182–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Ramanan VK, Risacher SL, Nho K, Kim S, Swaminathan S, Shen L, et al. APOE and BCHE as modulators of cerebral amyloid deposition: a florbetapir PET genome-wide association study. Mol Psychiatry 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Tanzi RE. The genetics of Alzheimer disease. Cold Spring Harb Perspect Med 2012;2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39].Thornton JW, DeSalle R. Gene family evolution and homology: genomics meets phylogenetics. Annu Rev Genomics Hum Genet 2000;1:41–73. [DOI] [PubMed] [Google Scholar]
  • [40].Watson JD, Baker TA, Bell SP, Gann A, Levine M, Losick R. Molecular biology of the gene. Benjamin Cummings; 2014. [Google Scholar]
  • [41].Sarkar IN, Egan MG, Coruzzi G, Lee EK, DeSalle R. Automated simultaneous analysis phylogenetics (ASAP): an enabling tool for phlyogenomics. BMC Bioinformatics 2008;9:103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Fitch WM. Toward defining the course of evolution: minimum change for a specific tree topology. Syst Zool 1971;20:406–16. [Google Scholar]
  • [43].Robinson DF, Foulds LR. Comparison of phylogenetic trees. Math Biosci 1981. [Google Scholar]
  • [44].Matsumoto Y, editor. Ruby Programming Language; n.d. [Google Scholar]
  • [45].Goto N, Prins P, Nakao M, Bonnal R, Aerts J, Katayama T. BioRuby: bioinformatics software for the Ruby programming language. Bioinformatics 2010;26:2617–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [46].Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, et al. BLAST+: architecture and applications. BMC Bioinformatics 2009;10:421. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Edgar RC. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics 2004;5:113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [48].Goloboff P, Farris S, Nixon K. TNT (Tree analysis using New Technology); 2000. [Google Scholar]
  • [49].McKusick VA. Mendelian inheritance in man and its online version, OMIM. Am J Hum Genet 2007;80:588–604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res 1997;25:3389–402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [51].Goloboff PA, Farris JS. Methods for quick consensus estimation. Cladistics 2001;17:S26–34. [Google Scholar]
  • [52].Peña C, Wahlberg N, Weingartner E, Kodandaramaiah U, Nylin S, Freitas AVL, et al. Higher level phylogeny of Satyrinae butterflies (Lepidoptera: Nymphalidae) based on DNA sequence data. Mol Phylogenet Evol 2006;40:29–49. [DOI] [PubMed] [Google Scholar]
  • [53].Goloboff PA, editor. TNT wiki; n.d.
  • [54].Smoot ME, Ono K, Ruscheinski J, Wang P-L, Ideker T. Cytoscape 2.8: new features for data integration and network visualization. Bioinformatics 2011;27:431–2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [55].Farris JS, Källersjö M, Kluge AG. Constructing a significance test for incongruence. Syst Biol 1995. [Google Scholar]
  • [56].van Oven M, Kayser M. Updated comprehensive phylogenetic tree of global human mitochondrial DNA variation. Hum Mutat 2009;30:E386–94. [DOI] [PubMed] [Google Scholar]
  • [57].Koski LB, Golding GB. The closest BLAST hit is often not the nearest neighbor. J Mol Evol 2001;52:540–2. [DOI] [PubMed] [Google Scholar]
  • [58].Wilgenbusch JC, Swofford D. Inferring evolutionary trees with PAUP⁄. Curr Protoc Bioinformatics 2003. [chap. 6:Unit6.4]. [DOI] [PubMed] [Google Scholar]
  • [59].Ostlund G, Schmitt T, Forslund K, Köstler T, Messina DN, Roopra S, et al. InParanoid 7: new algorithms and tools for eukaryotic orthology analysis. Nucleic Acids Res 2010;38:D196–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [60].Li L, Stoeckert CJ, Roos DS. OrthoMCL: identification of ortholog groups for eukaryotic genomes. Genome Res 2003;13:2178–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [61].Alexander PA, He Y, Chen Y, Orban J, Bryan PN. The design and characterization of two proteins with 88% sequence identity but different structure and function. Proc Natl Acad Sci USA 2007;104:11963–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Kaneko I, Yamada N, Sakuraba Y, Kamenosono M, Tutumi S. Suppression of mitochondrial succinate dehydrogenase, a primary target of beta-amyloid, and its derivative racemized at Ser residue. J Neurochem 1995;65:2585–93. [DOI] [PubMed] [Google Scholar]
  • [63].Pardossi-Piquard R, Petit A, Kawarai T, Sunyach C, Alves da Costa C, Vincent B, et al. Presenilin-dependent transcriptional control of the Abeta-degrading enzyme neprilysin by intracellular domains of betaAPP and APLP. Neuron 2005;46:541–54. [DOI] [PubMed] [Google Scholar]
  • [64].Liu Q, Zerbinatti CV, Zhang J, Hoe H-S, Wang B, Cole SL, et al. Amyloid precursor protein regulates brain apolipoprotein E and cholesterol metabolism through lipoprotein receptor LRP1. Neuron 2007;56:66–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [65].Nikolaev A, McLaughlin T, O’Leary DDM, Tessier-Lavigne M. APP binds DR6 to trigger axon pruning and neuron death via distinct caspases. Nature 2009;457:981–9. [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  • [66].Freir DB, Nicoll AJ, Klyubin I, Panico S, Mc Donald JM, Risse E, et al. Interaction between prion protein and toxic amyloid b assemblies can be therapeutically targeted at multiple sites. Nat Commun 2011;2:336. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [67].Yang J, Ji Y, Mehta P, Bates KA, Sun Y, Wisniewski T. Blocking the apolipoprotein E/amyloid-β interaction reduces fibrillar vascular amyloid deposition and cerebral microhemorrhages in TgSwDI mice. J Alzheimers Dis 2011;24:269–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [68].Manczak M, Reddy PH. Abnormal interaction of oligomeric amyloid-β with phosphorylated tau: implications to synaptic dysfunction and neuronal damage. J Alzheimers Dis 2013;36:285–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [69].Coulson EJ, Paliga K, Beyreuther K, Masters CL. What the evolution of the amyloid protein precursor supergene family tells us about its function. Neurochem Int 2000;36:175–84. [DOI] [PubMed] [Google Scholar]
  • [70].Tharp WG, Sarkar IN. Origins of amyloid-β. BMC Genomics 2013;14:290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [71].Brown TA. Genomes 2. Garland Publishing; 2002. [Google Scholar]
  • [72].Bejerano G, Pheasant M, Makunin I, Stephen S, Kent WJ, Mattick JS, et al. Ultraconserved elements in the human genome. Science 2004;304:1321–5. [DOI] [PubMed] [Google Scholar]
  • [73].Lehmann J, Libchaber A. Degeneracy of the genetic code and stability of the base pair at the second position of the anticodon. RNA 2008;14:1264–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [74].Vilella AJ, Severin J, Ureta-Vidal A, Heng L, Durbin R, Birney E. EnsemblCompara GeneTrees: complete, duplication-aware phylogenetic trees in vertebrates. Genome Res 2009;19:327–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [75].Degnan JH, DeGiorgio M, Bryant D, Rosenberg NA. Properties of consensus methods for inferring species trees from gene trees. Syst Biol 2009;58:35–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [76].Bansal MS, Burleigh JG, Eulenstein O, Fernández-Baca D. Robinson–Foulds supertrees. Algorithms Mol Biol 2010;5:18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [77].Chaudhary R, Burleigh JG, Fernández-Baca D. Inferring species trees from incongruent multi-copy gene trees using the Robinson–Foulds distance. Algorithms Mol Biol 2013;8:28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [78].Pattengale ND, Gottlieb EJ, Moret BME. Efficiently computing the Robinson–Foulds metric. J Comput Biol 2007;14:724–35. [DOI] [PubMed] [Google Scholar]
  • [79].Chaudhary R, Burleigh JG, Fernández-Baca D. Fast local search for unrooted Robinson–Foulds supertrees. IEEE/ACM Trans Comput Biol Bioinform 2012;9:1004–13. [DOI] [PubMed] [Google Scholar]
  • [80].Rogaeva E, Meng Y, Lee JH, Gu Y, Kawarai T, Zou F, et al. The neuronal sortilin-related receptor SORL1 is genetically associated with Alzheimer disease. Nat Genet 2007;39:168–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [81].Haataja L, Gurlo T, Huang CJ, Butler PC. Islet amyloid in type 2 diabetes, and the toxic oligomer hypothesis. Endocr Rev 2008;29:303–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [82].Bedell MA, Jenkins NA, Copeland NG. Mouse models of human disease. Part I: techniques and resources for genetic analysis in mice. Genes Dev 1997;11:1–10. [DOI] [PubMed] [Google Scholar]
  • [83].Bedell MA, Largaespada DA, Jenkins NA, Copeland NG. Mouse models of human disease. Part II: recent progress and future directions. Genes Dev 1997;11:11–43. [DOI] [PubMed] [Google Scholar]
  • [84].Ostrander EA, Franklin H. Epstein Lecture. Both ends of the leash – the human links to good dogs with bad genes. N Engl J Med 2012;367:636–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [85].Brown WM, George M. Rapid evolution of animal mitochondrial DNA; 1979. [DOI] [PMC free article] [PubMed]
  • [86].Kocher TD, Thomas WK, Meyer A, Edwards SV, Pääbo S, Villablanca FX, et al. Dynamics of mitochondrial DNA evolution in animals: amplification and sequencing with conserved primers. Proc Natl Acad Sci USA 1989;86:6196–200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [87].Soh D, Dong D, Guo Y, Wong L. Finding consistent disease subnetworks across microarray datasets. BMC Bioinformatics 2011;12(Suppl. 13):S15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [88].Goh K-I, Cusick ME, Valle D, Childs B, Vidal M, Barabási A-L. The human disease network. Proc Natl Acad Sci USA 2007;104:8685–90. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

supplementary information 2
supplementary information

RESOURCES