Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 10.
Published in final edited form as: Nat Methods. 2025 Dec 12;23(2):312–327. doi: 10.1038/s41592-025-02931-9

Computational strategies for cross-species knowledge transfer and translational biomedicine

Hao Yuan 1, Christopher A Mancuso 2, Kayla Johnson 3, Ingo Braasch 4,✉, Arjun Krishnan 5,✉
PMCID: PMC12970299  NIHMSID: NIHMS2139894  PMID: 41388108

Abstract

Research organisms provide invaluable insights into human biology and diseases, serving as essential tools for functional experiments, disease modeling, and drug testing. However, evolutionary divergence between humans and research organisms hinders effective knowledge transfer across species. Here, we review state-of-the-art methods for computationally transferring knowledge across species, primarily focusing on methods that utilize transcriptome data and/or molecular networks. Our review addresses four key areas: (1) transferring disease and gene annotation knowledge across species, (2) identifying functionally equivalent molecular components, (3) inferring equivalent perturbed genes or gene sets, and (4) identifying equivalent cell types. We conclude with an outlook on future directions and several key challenges that remain in cross-species knowledge transfer, including introducing the concept of “agnology” to describe functional equivalence of biological entities, regardless of their evolutionary origins. This concept is becoming pervasive in integrative data-driven models where evolutionary origins of functions can remain unresolved.

Introduction

The use of research organisms, also known as model species or model systems, is essential to biomedical research [1]. Leveraging their resemblance in anatomy, physiology, behavior and genetics to corresponding human conditions, research organisms help scientists investigate a wide range of biomedical phenomena and therapeutic treatments before they are applied to humans. For instance, the zebrafish (Danio rerio) is a well-established vertebrate research organism that has external and fast development in large offspring numbers, transparent embryos, and allows for easy genetic manipulation and drug administration through compound exposure [2].

There are two main reasons that research organisms are critical to studying human phenotypes, processes, and diseases. First, it is most often unethical to study biomedical processes or apply novel therapeutic treatments directly in humans. Although there are many unanswered questions on the ethics of using research organisms [3,4,5], the knowledge derived from these organisms has helped save countless human lives [2,6]. Secondly, genetic studies in humans typically have very high variability due to confounding effects such as population genetics and environmental factors [7,8,9,10], and these confounders can be controlled much more under defined laboratory conditions with research organisms.

However, using research organism data to explore human biology presents its own challenges. Evolutionary divergence often gives rise to similar yet different underlying biological processes between human and research organisms [11], thereby impeding the transfer of knowledge across species. Orthologs, i.e., homologous genes from different species that evolved from a shared ancestral gene by speciation, have taken a privileged role in transferring functional annotations across species. Nevertheless, orthologs may have experienced significant, lineage-specific functional changes across species [12,13,14,15,16,17]. Additionally, the complex genetics underlying phenotypes/processes shared across species may differ. This is because interactions among genes responsible for a phenotype could be rewired during the evolutionary process, potentially encompassing some de novo genes [18]. Consequently, the effective and precise transfer of knowledge considering de novo genes remains challenging. Finally, since no single research organism can fully recapitulate the entirety of a complex biological condition in humans [1], how do we predict which research organisms capture the different facets of the human biology of interest?

To address these issues, researchers have begun developing computational models that can robustly transfer knowledge between a variety of research organisms and humans to identify functionally similar genes or groups of genes across species regardless of their evolutionary origin. Recent discoveries of such cross-species gene pairs were evidenced by large-scale computational competitions such as sbvIMPROVER Species Translation Challenge [19] and the Critical Assessment of Protein Function Annotation Challenge (CAFA) [20,21,22,23].

In this review, we present a suite of cross-species knowledge transfer approaches with a significantly broader scope than previous reviews [24,25,26,27]. We comprehensively lay out the landscape of recent and state-of-the-art data-driven strategies, including those that leverage AI and machine learning, to answer four classes of important questions that frequently arise when using research organisms to study biomedical questions and translating findings to humans (Figure 1). (1) How to predict disease-gene or function-gene relationships across species (Figure 1a)? (2) How to identify functionally equivalent molecular components across species (Figure 1b)? (3) How to infer perturbed molecular profiles across species (Figure 1c)? (4) How to map equivalent cell types and cell states across species (Figure 1d)?

Figure 1: Schematic diagram of each section of this review article.

Figure 1:

(a) How to predict disease-gene or function-gene relationships across species? Diagram depicts associating genes in each species to specific functions, diseases, and phenotypes. (b) How to identify functionally equivalent molecular components across species? Diagram depicts finding the most equivalent gene, pathway/expression module, or phenotype between species. (c) How to infer perturbed molecular profiles across species? Diagram depicts gene expression in each species as a result of taking a particular perturbation like drug. (d) How to map equivalent cell types and cell states across species? Diagram depicts alignment of cell types across species.

Instead of discussing methods through an algorithmic lens, we are taking a question-first perspective to provide inspiration and background both for computational researchers interested in developing new methods in this area and for experimental/wet-lab researchers interested in finding and using the best methods in this area. Although we have separated the methods into four sections, many of the approaches described across these sections share similarities in algorithmic design, data types, and output formats. Detailed information about the methods underlying each of the main sections can be found in Supplementary File 1. In addition, we have curated numerous datasets and databases used in cross-species methodologies, summarized in Supplementary File 2 and Supplementary Note 1. These resources will enable researchers to test and improve upon existing computational methods in the field.

How to predict disease-gene or function-gene relationships across species

Comprehensive knowledge of the roles genes play in molecular functions, phenotypes, and diseases is fundamental to our understanding of the molecular underpinnings of biological, physiological, and pathological processes. However, only a limited number of genes in the human genome have been experimentally characterized. For example, ~33% of protein-coding genes have no functional annotation to any Gene Ontology term [28]. Knowledge of gene-function, gene-phenotype, and gene-disease relationships is significantly richer, though far from complete, in research organisms due to the availability of a variety of experimental techniques such as gene editing, knock-in, and knockout experiments [29,30]. How do we leverage this knowledge available across species to close massive annotation gaps, especially in light of the potential functional divergence of homologous genes and the presence of gene duplications, de novo genes, and gene losses? Moreover, how do we use existing knowledge in human and traditional research organisms (e.g. mouse, frog, zebrafish, fly, worm) to fill annotation gaps in non-traditional research organisms (e.g., dog [31], python [32], gar [33], planaria [34]) that still lack sufficient data. These needs have spurred the development of a number of computational methods for predicting disease-gene or function-gene relationships across multiple species (Figure 2).

Figure 2: How to predict disease-gene or function-gene relationships across species.

Figure 2:

(a) Network-based methods can annotate diseases or functions for genes across species by aligning networks or by embedding networks into a shared, low-dimensional feature space where a supervised learning model is then trained to propagate annotations. (b) Phylogenetic profile-based methods identify co-presence or co-absence of genes throughout evolutionary history, implying closely related functions among these genes. This relationship is then used to propagate annotations across species. (c) Disease and function annotations can be transferred by combining ontologies across species. Genes with similar functions or disease associations across species can then be identified through the ontology structure.

Utilizing molecular networks to predict functional and disease gene annotations across species

Molecular interaction networks, which map experimentally or computationally derived relationships between genes or their products, enable the prediction of gene functions and disease associations through the “guilt-by-association” principle, which posits that genes interacting within shared network neighborhoods are likely involved in similar biological processes [35,36,37]. Interaction neighborhoods within these networks capture the “functional context” of genes, which is highly complementary to sequence homology information. Further, these networks typically contain and therefore enable making inferences about nearly all genes in the genome, including non-homologous and un(-der)studied genes. Therefore, a number of network-based computational approaches have been developed to transfer and predict function, phenotype, and disease annotations across species, especially using machine learning (ML) approaches [38,39,40,41] (Figure 2a).

For instance, the Functional Knowledge Transfer (FKT) method [42] first finds “functionally-similar homologous gene pairs” that are in the same gene family and in similar network neighborhoods [40], and uses these pairs to propagate functional annotation across species. FKT successfully predicted zebrafish genes involved in mitotic exit regulation, such as cdh2, which was later experimentally validated and has then been experimentally confirmed related to cell cycle progression in zebrafish retina cells [43]. Similar to the idea in FKT, GenePlexusZoo [44] enhances cross-species predictions by simultaneously linking networks of human and five research organisms (mouse, zebrafish, fly, worm, yeast) through gene homology to generate a “functional representation” of genes from all these species, which improved prediction performance for pathway, phenotype and disease annotation tasks. Beside considering network similarity alone, NetQuilt [45] integrates multi-species networks using IsoRank [41], which aligns networks considering both sequence similarity and network similarity, and applies deep learning to predict annotations across species.

Beyond gene function prediction, network-based methods also aid in disease-gene association discovery. Katz measure [46], a well-established technique in social network link prediction [47], quantifies gene-disease proximity in heterogeneous networks combining cross-species gene orthologs and disease associations [38]. Further, “Katz features” that represent the number of paths of a certain length between a gene and a disease node, when incorporated into a ML framework, were shown to significantly enhance the performance of predicting gene-disease relationships [38]. Instead of using curated disease annotation, DiseaseQUEST [39] predicts disease-gene candidates in research organisms by combining human genome-wide association studies (GWAS) [48] with gene networks that reflect pathways underlying tissue-specific physiology [49] and disease [50,51]. Using the homologous and functionally equivalent genes of human GWAS disease genes as positive examples in a target research organism, and gene interactions in an appropriate tissue network as features, diseaseQUEST trains an ML model to prioritize novel disease-gene candidates. Strikingly, the authors demonstrated this approach by prioritizing and then experimentally validating genes related to Parkinson’s disease in the nematode Caenorhabditis elegans [39].

In addition, numerous methods have been developed for predicting disease-gene associations in single species, typically in humans. These methods can be adapted to predict disease-gene associations across species [36,52,53,54,55].

Discovering disease-related genes using phylogenetic profiles

Genes with similar functions tend to appear and disappear together during evolutionary history. Most human disease genes have ancient origins [56,57,58,59], with many traceable to eukaryotic ancestors and some dating back to the evolution of multicellularity [56]. By identifying genes that co-evolve with known disease-related genes, we therefore could propagate the disease annotations from the known genes to the co-evolved genes, thereby helping predict functions for genes that are not well characterized [60,61] (Figure 2b).

Maxwell and colleagues [62] used evolutionary profiles to examine the evolutionary distribution of human disease genes from the Online Mendelian Inheritance in Man (OMIM) database [63]. They revealed heterogeneity underlying the evolutionary origins of different classes of human disease genes. For example, genes related to inflammatory and immune diseases are of vertebrate origin, while genes associated with cardiovascular and hematological diseases originate from early metazoans.

Recently published methods [64,65] enabled systematic screens of genome-wide co-occurring functional modules, leading to functional predictions for numerous previously uncharacterized genes. For example, using phylogenetic profiling, Dey et al. [65] identified understudied candidate genes linked to the actin-nucleating WASH complex and ciliary and centrosomal defects by scanning co-occurring gene modules across 177 species from the eukaryotic tree of life. They further experimentally evaluated candidate functions of genes by identifying gene product co-localization and gene knockdown in cell models.

Bridging disease-gene association through phenotypic similarity

Extensive gene-phenotype associations in research organisms have been discovered through hypothesis-driven experiments and large-scale genetics screens [66,67,68]. These links can lead to the identification of genes in a research organism whose perturbation results in phenotypes similar to those observed in human patients with a particular disease, thereby pointing to a viable model for the disease under observation.

However, comparing phenotypes across species is challenging due to evolutionary divergence and non-standard descriptions of phenotypes in different species. The continuous efforts in phenotypic ontology curation and the development of cross-species phenotype ontologies such as uPheno [69] and PhenomeNET [70] have now significantly mitigated this challenge, enabling researchers to use ontology-based semantic similarity measurements between phenotypes to explore suitable research organisms and identify new disease-gene and function-gene relationships (Figure 2c). For instance, Exomiser [71] leveraged disease-to-gene relationships discovered through semantic similarity measures among integrated ontologies to assist in disease diagnosis.

Recent studies improve on semantic similarity by creating a latent embedding space based on the phenotype ontology combined with known disease-phenotype and disease-/phenotype-gene associations [72,73]. “Node embeddings” created this way contain a low-dimensional numerical representation of each entity, such as gene, phenotype, and disease that captures that entity’s relationships. Such embeddings naturally lend themselves as inputs into ML algorithms. In this study [72], the authors trained a supervised neural network model to predict gene-disease associations based on the node embeddings of genes and diseases, which performed better than an unsupervised approach based on the similarities between gene and disease embeddings.

How to identify functionally equivalent molecular components across species

Complementary to inferring genes in each species that are associated with a particular function or disease, another critical task is to infer molecular components that are functionally equivalent counterparts of each other across species. A number of computational approaches have been developed to identify such functionally equivalent components at the gene, pathway (gene set), and genomic level across species (Figure 3). While the concept of identifying functionally equivalent genes partially overlaps with gene-level predictions introduced in the previous section, this section expands on describing functional equivalency between higher-order functional elements.

Figure 3: How to identify functionally equivalent molecular components across species.

Figure 3:

(a) Functionally equivalent gene pairs can be identified through genes that connect to similar sets of genes in aligned cross-species gene networks. (b) When prior knowledge about gene sets is known, functionally equivalent gene sets are discovered through large overlaps between gene sets across species. When prior knowledge is not available, functionally equivalent gene sets can be discovered by finding sub-networks with similar topology and expression patterns across species. (c) Resemblance at the organismal level, i.e., functionally equivalent profiles, can be examined by comparing genomic profiles across species.

Identifying functionally equivalent gene pairs across species

An essential ingredient of finding functionally equivalent genes is to go beyond solely using sequence similarity and homology because, in many organisms, proteins performing similar functions (e.g., playing similar roles in the same biological pathway) may not be the most similar in sequence or share a common evolution origin [74]. Instead, equivalent molecular functions may have convergently, i.e., independently evolved. Gene Analogue Finder implemented this notion by measuring functional similarity between a pair of genes based on the overlap between gene ontology (GO) terms associated with them [75]. Han et al. [15] used a comparable approach to calculate functional similarity between orthologous human-mouse gene pairs based on the average pairwise similarity between human and mouse phenotype ontology terms annotated to those genes. This study identified several cases of functional divergence of orthologs that could be traced to changes in noncoding regulatory sequences of gene pairs with high protein sequence similarity. Nevertheless, the performance of such methods depends heavily on the completeness of gene annotations to terms in function and phenotype ontologies, with term overlap estimates becoming less meaningful for particular genes or entire species with sparse (i.e., highly incomplete) annotations [15].

Genome-scale gene interaction networks help overcome this limitation by providing a powerful alternative way to capture the “functional context” of each gene in terms of its local network neighborhood. Thus, two genes compared across species interacting with similar sets of genes in their respective molecular network neighborhoods are likely functionally equivalent (Figure 3a). Chikina et al. [40] realized this idea by first grouping the network neighbors of individual genes into meta-genes that correspond to Treefam families [76] and measuring the hypergeometric overlap between sets of meta-genes that are neighbors of a pair of genes across species in the respective species network. Intersecting meta-genes identified through this approach revealed specific biological processes underlying network similarities.

Other recent network-based approaches for finding functionally equivalent genes take advantage of the idea of node embedding. Methods such as MUNK [77], MUNDO [78], and ETNA [79] create a joint network-based representation of genes by embedding the molecular networks in a pair of species individually and then align the two embeddings based on sequence orthologs. MUNK produces different joint embeddings based on the “source” species and the “target” species, where the size of joint embedding set is as large as the number of genes in the smaller network. ETNA uses cross-training to align the two embeddings into a bidirectional compressed (low-dimensional) space. MUNK, MUNDO and ETNA use the similarity of embedding vectors of genes to estimate their functional similarity or to inform function or disease gene prediction.

Another promising approach to identify functionally equivalent genes involves comparing protein embeddings generated by protein language models (PLMs), such as ESM2 [80]. These embeddings implicitly capture protein structure, molecular characteristics, as well as evolutionary relationships between proteins [81]. Protein embeddings also enable convenient comparison of sequences across diverse species using simple cosine similarity, independent of sequence similarity. This strategy has led to recent methods, including those by Hamamsy et al. [82] and Hong et al. [83], which utilize PLMs to efficiently identify deeply homologous proteins across diverse species, even in cases of limited sequence similarity. Although these methods primarily focus on identifying homologous protein domains, they provide an important foundation for identifying structurally and functionally equivalent or similar genes across species. By encoding complex DNA sequence patterns into embeddings, DNA foundation models such as Evo2 [84] also offer complementary approaches for identifying functionally equivalent genes across species.

Discovering functionally equivalent gene sets across species

Going beyond individual gene pairs, it is also of interest to find functionally conserved gene sets across species because sets of genes could represent concepts like pathways and molecular mechanisms underlying conserved phenotypes. McGary et al. [85] introduced the concept of orthologous phenotypes, i.e., “phenologs”, defined as phenotypes that are associated with orthologs. Phenologs could include phenotype counterparts that are observably very different from each other while being influenced by conserved molecular functions (e.g., significantly overlapping sets of orthologous genes). Examples include a yeast model for angiogenesis defects, a worm model for breast cancer, mouse models of autism, and a plant model for human neural crest defects [85]. Phenologs serve as a valuable tool to quantitatively identify non-obvious phenotypes in research organisms for studying human diseases, along with disease gene candidates such as genes annotated to non-obvious phenotypes and which are not orthologous to any known disease genes. MUNK [77] removes the restriction of comparing only the overrepresentation of orthologous genes. Instead, it identifies phenologs by detecting functionally equivalent genes with similar protein embeddings, uncovering more phenologs than McGary et al.

Discovering functionally equivalent gene modules across species

Phenolog approaches rely on prior knowledge of functional and phenotypic annotations of genes, which is often highly incomplete, especially in non-traditional model species. Network-based methods have proven valuable in overcoming this limitation by helping to find conserved gene sets, usually called gene modules, while filling in knowledge gaps. This is because, in addition to capturing relationships between pairs of genes, networks also capture higher-level organization between groups of genes in the form of tight sub-networks. Consequently, similar sub-networks across species correspond to homologous functional modules or pathways (Figure 3b). CoCoCoNet [86] takes such a network-based approach to test whether a given gene set in one species is involved in similar functions as homologous gene sets in another species. If a subset of genes accurately predicts the remaining genes in the gene set using neighbor voting in both species, that is taken as an indication that the gene set corresponds to a conserved functional module.

Molecular networks can further refine comparisons of differential gene expression across species to help identify functionally conserved “active modules” that comprise tightly connected sub-networks of conserved genes that respond similarly to a given condition (e.g., disease, perturbation). However, finding such modules is difficult because active modules in different species are often not conserved across species, while conserved modules are not necessarily active [87]. So, algorithms for finding conserved and active modules need to consider activity plus conservation at the same time.

The neXus algorithm [88] was developed to meet this need. neXus identifies modules using a greedy seed-and-extend algorithm. It begins with a pair of orthologous nodes as seeds and iteratively extends both sub-networks by incorporating neighboring orthologous gene pairs. ModuleBlast [89] works similarly to neXus with the added capability of distinguishing the resulting modules based on the direction of the expression change. This separation allows ModuleBlast to evaluate whether conserved active modules display expression patterns in the same or opposite way. Both neXus and ModuleBlast limit the identification of modules to fully conserved genes. xHeinz [87] relaxes this constraint by allowing users to define the proportion of conserved nodes, offering a more flexible approach that leads to including functionally conserved but non-homologous genes in the final modules. xHeinz was further applied to understand the regulatory mechanism of interleukin-17-producing helper T (Th17) cell differentiation in humans and mice. The study revealed that key regulators of Th17 cells are conserved across species [87].

Measuring equivalent profiles between human and research organisms

Beyond genes and gene modules, evaluating the ability of a research organism to recapitulate specific human phenotypes or disease conditions is crucial for functional studies and experimental design. However, such evaluations are subjective, with researchers often using vague terms such as “poorly” or “greatly” to indicate resemblances. Therefore, some recent methods have leveraged functional genomics data to draw quantitative conclusions about the biological resemblance between human and research organisms (Figure 3c).

Congruence Analysis for Model Organisms (CAMO) [90] quantitatively measures the congruence between human and research organisms by comparing the distribution of differentially expressed gene (DEG) profiles under matching conditions. To improve the accuracy of identifying DEGs, CAMO infers differential posterior probabilities of genes based on p-values from conventional pipelines. The concordance level of differential gene expression across species was summarized as concordance (c-scores) and discordance scores (d-scores), which served as quantitative measures of congruence across species. CAMO was applied to reconcile studies [91,92] that reached contradicting conclusions about whether mice are a suitable model for studying human inflammatory diseases. By reanalyzing and comparing inflammatory expression data within and across species, the authors concluded that burn- and infection-induced inflammation in mice resembles human inflammation [90].

Comparing phenotypes provides another avenue to discover biological resemblance in research organisms. Cross-species phenotypic ontologies can link phenotypes related to human diseases to research organism phenotypes. Leveraging such connections, the Monarch Initiative [93] provides tools, such as the Phenotype Explorer, to prioritize research organisms based on phenotypic similarities between phenotype terms from human and research organisms, aiding in the discovery of relevant research organisms and phenotypes.

How to infer perturbed transcriptomes across species

Bulk and single-cell transcriptome profiling have emerged as preeminent technologies for nearly any organism to capture genome-wide molecular responses to a variety of developmental and physiological states as well as treatments, perturbations, and other conditions. Transcriptome profiling in research organisms has especially been valuable in studying perturbations that may be impractical or ethically infeasible in humans. Further, comparing transcriptomic profiles across species sheds light on conserved and distinct cellular states and gene responses, and makes way for context-specific knowledge transfer. However, directly comparing expression changes across species based on gene homology is challenging due to evolutionary divergence in gene expression programs. This section describes several methods that have been developed to utilize computational techniques to enable accurate comparisons of transcriptomic profiles across species (Figure 4).

Figure 4: How to infer perturbed transcriptomes across species.

Figure 4:

(a) Some methods use linear models to relate DEG effect sizes utilizing matched samples between humans and research organisms. Other approaches develop ML models based solely on research organisms to predict conditions in humans, which is then used to identify DEGs between conditions. (b) To infer enriched gene sets across species, gene set representations are aligned, then models are built that can transfer enrichment scores between species.

Determining differentially expressed genes across species

Identifying differentially expressed genes (DEGs) from transcriptomic profiles has led to insights into the impact of numerous diseases, perturbations, or experimental conditions. To systematically capture the relationship of DEGs across species, a class of methods train ML models using carefully constructed cross-species dataset pairs (CSDPs) with matching conditions, offering curated examples of how expression changes correspond to phenotypic outcomes across species, which can be further used for model training [94,95] (Figure 4a).

Found In Translation (FIT) [94] is a linear regression model that is used to understand relationships of DEGs between human and research organisms for each orthologous gene pair. The authors built models based on 170 mouse-human CSDPs across 28 conditions, which were pairs of mouse and human experiments associated with the same disease or condition. The model was then used to predict DEGs in humans under conditions analogous to those captured in novel mouse data. FIT identified more human DEGs compared to simple transfers of DEGs from research organisms to human based on orthology. For instance, using protein immunostaining, the authors confirmed their prediction that the gene ILF3 is upregulated in the colon of patients with inflammatory bowel diseases (IBD) even though the gene was not differentially expressed in either human or mouse data [94].

The FIT approach is likely to be limited to diseases or drug perturbations with sufficient training data from both human and research organisms. In contrast, Brubaker et al. [95] proposed a semi-supervised method to predict DEGs across species that only requires phenotype information in research organisms. Initial supervised models were built to predict mouse phenotypes using mouse expression data, and then these models were iteratively augmented with high-confidence human samples to predict the phenotypes of the remaining human expression data. Finally, differential gene expression analyses were performed on human samples with distinct predicted phenotypes.

Transcomp-R [96] was developed to predict DEGs across species when phenotype labels are available for only one species, which applied to identify mouse genes predictive of infliximab responses in IBD patients, despite the absence of corresponding phenotype labels in mice. By applying principal component analysis (PCA) to murine expression data and projecting human expression data into the murine PC space, TransComp-R identified murine PCs and genes associated with infliximab response. This approach highlighted integrin signaling activation in resistant patients, with single-cell sequencing revealing ITGA1 overexpression in immune cells. The role of ITGA1 in resistance was further validated through anti-ITGA1 treatment on patient immune cells. Beyond infliximab response in IBD, TransComp-R has been used to identify other translatable signatures, such as MK2 kinase inhibition in IBD [97] and common disease signatures among mouse and human in Alzheimer’s disease [98]. Similar projection techniques have also been applied to cross-species pathway comparisons [99].

Finally, other recent work [100,101] have treated biological conditions—such as tissue type or disease state—as a “style” and applied style transfer techniques to predict transcriptomic changes within a species. AutoTransOP [102] extends this concept to cross-species comparisons by treating species as a style. Using an autoencoder structure, AutoTransOP compresses gene expression profiles into a latent space, maximizing mutual information and minimizing the distance between same-condition embeddings to create a meaningful representation. When a new sample from one species is introduced, the model can predict its corresponding expression profile in another species. Interestingly, AutoTransOP also modified TransComp-R for cross-species profile prediction by using PCA as a latent variable and decoding it to the translated profile. In translating lung fibrosis profiles between humans and mice, TransComp-R achieves performance comparable to AutoTransOP when orthologous genes are sufficiently available, suggesting that linear transformations may be adequate in such contexts. However, when orthologs are limited, AutoTransOP excels.

Though promising, it is important to note that genetic differences across species are significantly larger and more complex than the condition changes observed within a single species. As a result, applying style transfer techniques directly to cross-species comparisons may face additional challenges and limitations that need to be carefully considered and addressed.

Determining functionally enriched terms across species

Gene set analyses (GSA) have been widely used to identify predefined gene sets (e.g. Gene Ontology biological processes) that are significantly enriched in a gene list of interest [103]. GSA provides a powerful approach to gaining insights into enriched functional terms associated with diseases, phenotypes, or perturbations. However, the discrepancy of genes across species makes the definition of functionally equivalent gene set across species difficult, so directly transferring enriched terms across species is impractical (Figure 4b).

XGSEA [104] tackles the challenge of transferring enriched terms across species by using affine mapping, which projects genes from humans and research organisms into the same space. In this joint space, gene sets sharing more homologous genes will be closer to each other in the transformed space. With the transformed gene set representations, logistic regression models were built to characterize the relationship between the representation and enrichment metrics from the research organism, including p-values and enrichment scores of gene sets. Finally, the trained models are used to predict enrichment metrics for gene sets from humans. XGSEA successfully estimates enrichment metrics in human gene sets based on metrics of homologous gene sets in research organisms. For instance, XGSEA was able to train computational models based on zebrafish data and predict enriched pathways associated with melanomas in human patients [104].

How to map equivalent cell types and cell states across species

The advent of single-cell and single-nucleus sequencing techniques has opened up new avenues in biological research. This emerging frontier focuses on identifying and comparing cell types and states across different species to better understand cellular diversity and cellular innovations that arose through cell type evolution [105,106]. By transferring insights from extensively studied organisms to less-explored ones, these cross-species comparisons offer a powerful insight for cell type identification and discovery, thereby advancing our understanding of cellular biology across species.

Cell mapping problems may be simply conceptualized as mapping transcriptionally similar cells across species (i.e., based on homologous genes having similar expression levels). However, cells of a particular type within a species are more alike than those of corresponding cell types across different species [106,107]. Therefore, effectively mapping cells across species requires correcting for species-specific expression differences so that cells of homologous types are mixed in irrespective of species origin, while cells of different types remain distinct from each other. Additionally, methods must account for experiment-induced batch effects, which arise from technical variations rather than true biological signals. Some tools have been developed to reconcile heterogeneous single-cell RNA-seq (scRNA-seq) data from multiple species [108,109]. These methods typically seek to project cells from multiple species into a unified space, facilitating cross-species comparison. As a result, the cell mapping problem can be viewed as a domain adaptation problem, i.e., aligning cells from different species in a unified space (Figure 5).

Figure 5: How to map equivalent cell types and cell states across species.

Figure 5:

(a) The majority of methods align cells with similar expression patterns of homologous genes across a pair of species. (b) SATURN stands out as the only non-transformer-based method capable of aligning single-cell datasets from multiple species simultaneously. It first aggregates genes into synthesized macrogenes using protein language model. These synthesized macrogenes are then utilized to align cells across multiple species. (c) Newer methods are employing transformer architectures for self-supervised alignment of cells across species. Genes are first subsampled – either selecting top-expressed genes or random subsets – and then embedded to represent cells. These embeddings pass through transformer blocks followed by a feed-forward layer to predict the expression level of masked genes from the input or whether the masked genes are expressed in the cells.

Aligning cells across a pair of species based on homologs

Shafer [106] summarized various methods for cross-species scRNA-seq integration. Originally designed to mitigate the batch effects that were introduced by experimental variation, these methods may not sufficiently correct for species differences, which are more pronounced than batch effects [107]. When comparing human cells to those of anciently polyploid species like the teleost fishes (e.g. zebrafish), the abundance of gene duplicates results in fewer one-to-one gene matches, leading to a reduced signal for alignment. The reliance on orthology stems from the orthology conjecture, which posits that orthologs (i.e., homologous genes in different species that evolved from a common ancestral gene by speciation) are generally the most functionally similar genes across species, while gene duplicates (paralogs) tend to functionally diverge more at equivalent sequence distances. However, orthology does not necessarily imply similar functions across species, and paralogs might exhibit more functional similarity than orthologs. For instance, when an ortholog acquires a loss-of-function mutation, its function might be compensated by the upregulation of a paralog, resulting in the paralog having a more similar function to the ortholog in the other species [110,111]. Therefore, incorporating one-to-many, many-to-one and many-to-many homology into the analyses could improve cross-species alignment performance by accounting for functional similarity between paralogs (Figure 5a).

In a benchmark study conducted by Song et al. [112], SAMap [113] was identified as the only method that is capable of mapping divergent cell types. To account for the effect that gene expression patterns may be divergent in homologous cell types, SAMap uses one-to-one, one-to-many, many-to-one and many-to-many homology information and incorporates neighboring cells within species into the calculation of similarity between cells across species. Consequently, cells with differing expression profiles remain closely associated if they are among the nearest cross-species neighbors to each other. Based on this analysis, SAMap identified a general alignment of gene expression patterns and developmental relationships during embryogenesis in frogs and zebrafish. Interestingly, they also detected a group of secretory cell types that have similar expression patterns while having different developmental origins in the two different species including cells arising from different germ layers [113].

Aligning cells among multiple species using all genes

Instead of using precalculated gene homology, the new framework SATURN [114] integrates cell atlases from divergent species by harnessing the power of protein language models (PLMs) to relate genes across species to each other and project all cells from multiple species into shared cell embedding space (Figure 5b). Using a neural network, SATURN first projects scRNA-seq datasets to a joint space composed of “macrogenes” representing groups of genes across species that have similar protein embeddings inferred from the PLM. The neural network weights are used to define how a gene contributes to a macrogene. For each macrogene, SATURN learns the cross-species cell embedding as a non-linear combination of macrogenes, guided by an unsupervised objective function that simultaneously maximizes the distance between distinct cells within the same species and minimizes the distance between similar cells across species. This approach enables SATURN to integrate data from multiple species (and not just perform pairwise comparisons), including divergent species where gene homology relationships are hard to unravel. The output of SATURN also enables easy cross-species comparison, such as finding differentially-expressed macrogenes across species, which can then be traced back to actual genes based on the weight of connections to the macrogenes in the neural network. SATURN’s application to a multi-species dataset of frog and zebrafish embryogenesis shows success in aligning evolutionarily related cell types and revealing differentially expressed genes in macrophage/myeloid progenitors and ionocytes across these two distantly related species [114].

Advancing cross-species cell type alignment with transformer models

A new wave of approaches for aligning cell types across species is based on transformer models, which take large-scale single-cell data from multiple species and tissues as input and are trained in a self-supervised manner (Figure 5c). A key challenge in these transformer models is encoding cells with different numbers of genes across species into a fixed-length token sequence due to token limitations. Another issue is determining the gene ordering within a cell. Following solutions from single-species single-cell transformer models [115,116], cross-species models like Precious3GPT [117] and GeneCompass [118] encode cells as sequences of gene embeddings ordered by expression level. They also add special species tokens to differentiate cells from different species. However, as the typical maximum token length is limited, lower-expressed genes, which may have important functions, are often ignored. The Universal Cell Embeddings (UCE) [119] model addresses this limitation by subsampling a fixed number of genes based on their expression level, allowing it to capture expression patterns from genes with low-expression. UCE also applies an alternative gene ordering strategy, sorting genes by genomic location to capture gene regulatory and gene order (synteny) relationships that are often evolutionarily conserved, aiding cross-species integration.

Precious3GPT and GeneCompass models include all human genes but only homologs from other species. This reliance on homology limits these models to a few closely related species. On the other hand, instead of learning gene embeddings from scratch, UCE [119] uses protein sequence embeddings as input, similar to strategies applied in SATURN [114]. While this approach restricts the model to protein-coding genes, it bypasses homology constraints and enables easier and broader cross-species comparisons. Some models integrate prior knowledge with the expression data to improve performance. For example, GeneCompass [118] incorporates information on gene regulatory networks, gene promoters, gene family annotations, and gene co-expression relationships. Precious3GPT [117] enhances predictions using knowledge graphs and text description of genes from National Center for Biotechnology Information (NCBI).

These models provide a way to leverage insights from model species to better understand non-model species through predictions without requiring prior knowledge (i.e., zero-shot predictions) or relying on homology. With the number of cross-species single-cell models steadily growing, it is critical to perform comprehensive benchmarking to evaluate their ability to learn complex expression patterns across species, assess integration quality, and ensure accurate cross-species cell-type predictions.

Future perspective

The explosion of large-scale human biological data and accompanying analysis methods, while seen as a prospect of a golden age for human genetics, has also raised the question of whether it will diminish the significance of research organisms [7,120]. From the perspective of this review, the answer is certainly “No”. In fact, the ocean of new data will lead to an increasing number of candidate disease-related genes, dysregulated pathways, drugs, etc., that will require validation and functional characterization, which can be carried out comprehensively only in the in vivo settings of research organisms [7]. Therefore, the goal of increasingly better computational strategies to effectively gain biological and biomedical insights from research organisms will continue to remain significant and grow further. The landscape of current methods, summarized in this review, demonstrates great promise towards this goal. Instead of solely relying on sequence similarity and gene homology, these approaches utilize sophisticated inferential techniques and diverse datasets, paving a promising path for effectively transferring knowledge between human and research organisms. The next frontier in unlocking the full potential of research organisms and the ever-growing datasets involves addressing some key challenges that still remain.

Capturing the specific facets of complex diseases

Complex human diseases exhibit staggering genetic and phenotypic heterogeneity. Despite this complexity, much of the current focus is on identifying a single research organism that can fully recapitulate a disease condition. Consequently, a major need in the field is to develop methods that can identify optimal phenotypes, conditions, and genes for studying the many facets of complex diseases.

Several approaches have been proposed to dissect diseases into molecular components. Given that disrupted biological processes are often shared among diseases [121,122,123], even those that are seemingly unrelated [124], it is promising to dissect diseases into dysregulated processes [125] or modules [126,127]. Additionally, phenotype ontologies, such as Human Phenotype Ontology (HPO) [128], are widely used tools for phenotype-driven disease analyses, enabling the breakdown of diseases into related phenotypes [129,130].

By leveraging dissected processes, modules, or phenotypes, we can identify combinations of research organisms that maximally capture the multifaceted nature of complex diseases. The methods discussed in this review, particularly in the section “How to identify functionally equivalent molecular components across species”, shed light on approaches for relating these dissected components to suitable research organisms. However, there remains a critical need for future approaches that can utilize functionally equivalent molecular components of complex diseases to guide the strategic selection of the most appropriate research organisms for specific disease aspects.

The concept of agnology: sometimes we just don’t know

Homology is widely used as the bridge to find functional similarities across species. However, the definition of homology, i.e., shared evolutionary origin of a gene or phenotype, in itself does not include function nor does homology guarantee functional similarity. Many factors can lead to functional divergence between homologous genes. Examples include changes in non-coding regulatory sequences [15], reciprocal gene loss [131], and developmental system drift [132].

Moreover, solely focusing on homology restrains the power of predictive models to explore comprehensive functional relationships across divergent species. For instance, research organisms like the nematode C. elegans may not share a substantial number of clear and direct gene orthologs with humans, resulting in a scarcity of genes available for model training and testing. This limited orthologous gene repertoire can hinder the ability of computational models to learn comprehensive functional relationships across species. Furthermore, orthologs alone can only explain a fraction of observed biological variations [106]. Therefore, it is crucial to also incorporate non-homologous genes to construct more comprehensive models. Relying solely on homologous genes is also limiting from the perspective of network biology, where species evolve through gain or loss of interactions, indicating that non-homologous genes can play roles in similar functional and regulatory programs [133].

Some methods discussed in this review, like FIT [94] for cross-species profile prediction or CoCoCoNet [86] for network module comparison, focus mainly on homologous genes. However, there is a growing shift toward broader comparisons beyond homology. Some approaches, like MUNK [77] and GenePlexusZoo [44], use homology as anchors to expand comparisons to non-homologous and species-specific genes. Others, like SATURN [114] and UCE [119], directly use protein embedding similarity to create a more expansive map between genes across species without being restricted to a priori inferred gene homologies. To formally describe computationally discovered fully or partially functionally equivalent genes that may or may not be homologous, i.e., when “we just do not know”, we introduce the term “agnolog” (Figure 6). Agnologs are biological entities, processes, or responses—such as genes, gene sets, or biological systems—that are functionally equivalent across species regardless of their evolutionary origin. The prefix “agno-” means “unknown” or “not known”, reflecting a data-driven observation, independent of evolutionary assumptions.

Figure 6: Definition of agnology.

Figure 6:

Traditional methods rely on homology to determine functional equivalence. In contrast, modern data-driven approaches introduce agnology, identifying fully or partially functionally equivalent genes regardless of evolutionary origin, allowing for ambiguity in homology or convergent evolution/analogy.

In this context, agnology bridges the gap between homology and analogy. Analogs are functionally similar entities that evolved independently, while homologs share a common ancestry, though some may have diverged functionally. Agnologs encompass both analogs and homologs with similar functions. Identifying agnologs allows for further evolutionary analysis to determine whether they are in fact analogs or functionally equivalent homologs.

Gene rescue experiments provide a potential way to determine agnologs, where a gene is functionally disabled in one species and replaced by a gene from another species to assess whether the original phenotype can be restored. These assays, commonly performed in research organisms like Drosophila [134] and yeast [135], have traditionally focused on orthologous genes to reduce the search space. Methods discussed in this review can help to broaden candidates to non-orthologous genes. For example, MUNK predicted human gene sets associated with muscular dystrophy to be functionally related to mouse gene sets linked to abnormal muscle fiber morphology, despite minimal shared homology. Specifically, MUNK identified 21 functionally similar protein pairs between these phenotypes, with only one being homologous [77]. Another representative example is GenePlexusZoo, which predicted genes associated with Bardet-Biedl Syndrome 1 (BBS) across five research organisms. Notably, among the predicted genes annotated to top BBS-related biological processes in these organisms, none of the genes turned out to be orthologous to the human BBS genes [44].

Similar to homology being defined at different levels [136], agnologs can be defined and analyzed at multiple biological levels, including genes, gene sets, and cell types. At the gene level, agnologous pairs are genes with similar functions across species, typically identified by first anchoring a subset of genes through homology, followed by exploration of functional similarities to all genes beyond homologous relationships. These agnologous genes are often recognized using embedding representations, enabling systematic comparisons. At the gene-set level, agnologous gene sets are groups of genes responsible for similar functions or phenotypes, identified through overrepresentation analyses of agnologous genes or through predictive approaches, relying less on homology.

Identifying agnologous cell types across species represents a challenging yet exciting frontier in cell type evolution studies. Methods leveraging homologous genes include SAMap [113], Precious3GPT [117], and GeneCompass [118], whereas SATURN [114] and UCE [119] utilize agnologous genes identified through protein language models. For instance, SATURN identified early-stage zebrafish macrophages and frog myeloid progenitors as agnologous cell types, as both share the capability of differentiating into macrophages regulated by conserved mechanisms.

These concepts – agnologous genes, gene sets, and cell types – are deeply interconnected. Agnologous genes serve as a backbone to guide predictions of agnologous gene sets, as demonstrated by methods such as MUNK and GenePlexusZoo. Similarly, embedding representations derived from agnologous genes facilitate the identification of agnologous cell types across diverse species, such as in SATURN or UCE.

Networks in more species and more contexts

Molecular networks are one of the most widely applied data types in cross-species knowledge transfer. Regardless of their representation (e.g., edgelists, adjacency matrices, or node embeddings), networks capture interactions among neighboring genes and provide mechanistic representations of genes to computational models. However, network-based methods also have limitations.

Most networks generated using experimental approaches are for humans or popular research organisms like mice. There is an urgent need to expand the repertoire of networks beyond these species to include non-traditional research organisms. Current methods, such as STRING [137] and those described in recent literature [138], use orthology information to infer “interologs”, i.e., conserved interactions between pairs of proteins that have interacting homologs in another organism [139,140]. For example, STRING utilizes high-confidence networks in humans and data-rich research organisms to derive protein-protein interaction (PPI) networks for over 1,000 species. However, the assumption that paired orthologous genes have conserved interactions across species may not always hold true due to interaction rewiring over the course of evolution. Additionally, limiting the scope to orthologous genes precludes inferring interactions involving taxon-specific genes. Other previously introduced methods like FKT/IMP [42,141] use homologous genes with similar network neighborhoods to transfer knowledge across species, but these methods require experimental functional genomics data as prior to generating networks, limiting their application to popular research organisms.

Another important future step in generating species-specific networks is the development of context-specific networks, particularly those with tissue or cell-type specificity. Context specificity plays a crucial role in biomedicine since disease-gene associations frequently arise from disrupted interactions among tissue-specific and cell lineage-specific processes under particular environmental conditions [51,142]. Unfortunately, the nuanced interactions that vary across tissues may not be fully captured by experimentally generated large-scale networks such as PPI. Coexpression networks have the potential to capture tissue specificity more effectively [143], but they tend to be noisy [144]. To obtain robust signals, current studies often extract only a small fraction of information for downstream analysis, such as the top 0.5% of co-expressed gene pairs [145]. This approach, however, results in significant loss of information. Limited efforts have been made to build tissue-specific networks by integrating multi-modal functional genomic data [39,51] or to contextualize network representations by incorporating tissue-specific expression [146,147,148].

Recent advances in sequence-based deep learning models could help expand the repertoire of species-specific and context-specific networks. For example, AlphaFold-Multimer [149] can predict protein interactions based on sequences. ExPecto [150] and ExPectoSC [151] can predict tissue-/cell-specific regulatory landscapes based on DNA sequences alone. Thanks to the advancement of genome projects for non-model species, such as the Vertebrate Genomes Project [152], genome sequence resources are much richer than other genomics types such as transcriptomes or epigenomes. These sequence-based models and transfer learning techniques pave a promising way to expand networks beyond popular research organisms. However, we need to integrate phylogenetic information into transfer learning to consider the taxonomic-specific logic of gene regulation or protein interactions.

Automated construction of ontologies and knowledge graphs

In this review, we have described the power of ontologies as a framework for transferring knowledge across species. Furthermore, combining various ontologies and annotations into knowledge graphs can provide new insights for translational biomedicine. The Monarch Knowledge Graph exemplifies this approach, integrating knowledge from 33 biomedical resources, including information from all major research organism databases [93]. Leveraging such knowledge graphs, we can employ advanced techniques like graph deep learning [153] to utilize information from research organisms and assist biomedical inquiries such as rare disease variant prioritization [154]. Integrating knowledge graphs with large language models (LLMs) can also help reduce “hallucinations” in AI-powered question-answering systems within the biomedical field [155].

Despite the power of ontologies and knowledge graphs in cross-species knowledge transfer, the generation of ontologies requires laborious curation by specialists. Manual curation may also introduce biases towards popular research areas. For instance, since zebrafish is a widely used research organism for developmental and embryogenic studies, GO terms annotated to zebrafish might be skewed towards early life stages. Such differences in terms of annotation frequency could in turn bias downstream genomic analyses [156], especially when terms are then transferred to other, less-studied species.

The capability of current LLMs to automatically extract entities and relationships from the literature offers an efficient alternative for generating ontologies and knowledge graphs. For example, SPIRE [157] can generate ontologies by processing text input and a user-provided schema that describes the desired structure of the ontology. While more effort is needed to maintain the quality and consistency of automatically generated ontologies, this approach holds great potential for significantly enriching ontologies across a wide spectrum of organisms, from well-studied to understudied research organisms and even non-model organisms.

Better benchmarking studies for cross-species scRNAseq analyses

Mapping single-cell and single-nuclei transcriptomics (scRNA-seq/snRNA-seq) data across different species provides valuable insights into cell type evolution and facilitates cell type and cell state annotation [105,106]. However, most existing cross-species cell mapping algorithms are limited to closely related species and do not account for one-to-many, many-to-one, and many-to-many homology relationships [106], while considering all homologous genes is crucial when comparing species involving gene duplication derived from common ancestors. For example, due to a whole genome duplication in the ancestor of teleost fishes, considering all homologous gene pairs is essential for cross-species integration of cell types between human and biomedical fish models such as the teleost zebrafish [158]. To overcome such limitations, algorithms such as SAMap [113] and SATURN [114] have been developed. While recent benchmarking studies highlight properties of these methods that are important for effective cross-species integration of scRNA-seq data [112], most of these methods were originally designed for batch corrections. Furthermore, since SAMap applied additional restrictions on within-species manifolds, SAMap was not systematically compared to other methods in current benchmarking studies. Transformer-based alignment models further expand the landscape of cross-species scRNA-seq alignment methods, differing in their strategies for encoding cell embeddings, model architectures, species divergence, and the proportions of species included in the training data. However, there is currently no comprehensive benchmarking study that systematically evaluates these models. Thus, it is essential to conduct more comprehensive benchmarking of these methods to evaluate their performance on different types of scRNA-seq datasets and species with varying degrees of divergence and data curation. Additionally, the development of new methods for downstream analyses is crucial to extract meaningful biological insights from the growing mountain of single-cell/nuclei data.

Extracting and curating knowledge from non-traditional research organisms

The majority of the organisms noted in this review are classic research organisms such as mouse, zebrafish, and nematode, with non-traditional and emerging research organisms often being overlooked in “mainstream” biomedical studies. However, non-traditional research organisms are significant for translational research due to their unique biological characteristics [159]. For instance, the axolotl (Ambystoma mexicanum), a neotenic salamander, shows remarkable regenerative abilities, making it a crucial research organism for regenerative medicine and developmental biology [160]. The spotted gar (Lepisosteus oculatus), a ray-finned fish with a uniquely slow evolutionary rate, has proven valuable in facilitating genomic comparison between teleost biomedical research organisms and the human genome [33]. The tunicate, Ciona intestinalis, is a close chordate relative of vertebrates. Its unique evolutionary position, simple body plan, and ease of manipulation make it highly suitable for studying chordate embryonic development and morphogenesis [161]. However, challenges such as polyploidy, large genome sizes, genomic rearrangements, and taxon-specific cell types can make analyzing the genomic data from non-traditional species more difficult than classic research organisms. It is thus necessary to explore additional methods that can effectively transfer knowledge and data from and to non-traditional research organisms.

Conclusions

In this review, we have discussed the current state of computational methods for cross-species knowledge transfer. We emphasize that these methods surpass simple comparisons of molecular profiles across species and highlight the utilization of orthogonal information sources such as phenotypic ontologies and molecular networks to facilitate cross-species knowledge transfer.

Our review provides resources and insights into the advancements and challenges of methods for cross-species knowledge transfer among various research organisms and human. The resources summarized in this review will facilitate biomedical studies using research organisms, including traditional and non-traditional research organisms, by leveraging knowledge from human or extensively studied model systems.

There are still critical unmet needs and plenty of room for further improvements and refinements of existing approaches. More advanced computational approaches are needed to identify a range of research organisms for studying different aspects of diseases. Instead of relying solely on homology as a bridge, expanding analyses to agnologs can enhance the effectiveness of cross-species knowledge transfer. To gain a more comprehensive understanding of diseases, we need networks from more species, and also need to incorporate context specificity, especially cell- and tissue-specificity, into networks, despite the inherent difficulties involved. To fully unlock the power of ontology-based knowledge transfer, we need methods to automatically generate high-quality and robust ontologies as well as knowledge graphs. As a rapidly growing field, it is essential to establish more benchmarks for state-of-the-art cross-species cell mapping between humans and evolutionary divergent research organisms. Currently, most methods are focused on the few well-studied research organisms, but greater attention and analytical methods should be employed to extract biomedical insights from non-traditional, emerging research organisms.

We propose that with the application of comprehensive computational approaches, the field will gain more exciting insights from big data of research organisms, ultimately enhancing our understanding of human biology and diseases.

Supplementary Material

Supplementary File 2
Supplementary File 1
Supplementary Note 1

Acknowledgements

This work is supported by NIH R35 GM128765 and Simons Foundation 1017799 (to A.K.). (Zebra)fish-human research transfer in the Braasch Lab has been supported by NIH R01 OD011116. We thank all members in Krishnan and Braasch Labs for helpful discussion and feedback on the manuscript.

Footnotes

Competing interests

The authors declare no competing interests.

Contributor Information

Hao Yuan, Genetics and Genome Science Program; Ecology, Evolution, and Behavior Program, Michigan State University.

Christopher A. Mancuso, Department of Biostatistics & Informatics, University of Colorado Anschutz Medical Campus

Kayla Johnson, Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus.

Ingo Braasch, Department of Integrative Biology; Genetics and Genome Science Program; Ecology, Evolution, and Behavior Program, Michigan State University.

Arjun Krishnan, Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus.

References

  • 1.Animal Models are Essential to Biological Research: Issues and Perspectives; Françoise Barré-Sinoussi, Xavier Montagutelli Future Science OA (2015-July-31) https://doi.org/gftwtr DOI: 10.4155/fso.15.63 [DOI] [Google Scholar]
  • 2.Model organisms contribute to diagnosis and discovery in the undiagnosed diseases network: current state and a future vision,; Baldridge Dustin, Wangler Michael F, Bowman Angela N, Yamamoto Shinya, Schedl Tim, Pak Stephen C, Postlethwait John H, Shin Jimann, Lilianna Solnica-Krezel, … Monte Westerfield Orphanet Journal of Rare Diseases (2021-May-07) https://doi.org/gkjnhw DOI: 10.1186/s13023-021-01839-9 [DOI] [Google Scholar]
  • 3.Use of animals in experimental research: an ethical dilemma?; Baumans V Gene Therapy (2004-September-29) https://doi.org/dsrjkb DOI: 10.1038/sj.gt.3302371 [DOI] [PubMed] [Google Scholar]
  • 4.The ethics of animal research; Festing Simon, Wilkinson Robin EMBO reports (2007-May-18) https://doi.org/bzr78f DOI: 10.1038/sj.embor.7400993 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Harmonization of Animal Care and Use Guidance; Demers Gilles, Griffin Gilly, Vroey Guy De, Haywood Joseph R, Zurlo Joanne, Bédard Marie Science (2006-May-05) https://doi.org/bh2jbk DOI: 10.1126/science.1124036 [DOI] [PubMed] [Google Scholar]
  • 6.Model Organisms Facilitate Rare Disease Diagnosis and Therapeutic Research; Michael F Wangler Shinya Yamamoto, Chao Hsiao-Tuan, Jennifer E Posey Monte Westerfield, Postlethwait John, Hieter Philip, Boycott Kym M, Campeau Philippe M, Bellen Hugo J Genetics (2017-August-31) https://doi.org/gbwb63 DOI: 10.1534/genetics.117.203067 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.The future of model organisms in human disease research; Aitman Timothy J, Boone Charles, Churchill Gary A, Hengartner Michael O, Mackay Trudy FC, Stemple Derek L Nature Reviews Genetics (2011-July-18) https://doi.org/b3ht5z DOI: 10.1038/nrg3047 [DOI] [PubMed] [Google Scholar]
  • 8.Genetic Heterogeneity in Human Disease; McClellan Jon, King Mary-Claire Cell (2010-April) https://doi.org/b9rhz5 DOI: 10.1016/j.cell.2010.03.032 [DOI] [Google Scholar]
  • 9.Molecular Characterization Reveals Genetic Uniformity in Experimental Chicken Resources; TADANO Ryo, KINOSHITA Keiji, MIZUTANI Makoto, ATSUMI Yuusuke, FUJIWARA Akira, SAITOU Toshiki, NAMIKAWA Takao, TSUDZUKI Masaoki Experimental Animals (2010) https://doi.org/fg9q89 DOI: 10.1538/expanim.59.511 [DOI] [PubMed] [Google Scholar]
  • 10.Isogenic lines in fish – a critical review; Franěk Roman, Baloch Abdul Rasheed, Kašpar Vojtěch, Saito Taiju, Fujimoto Takafumi, Arai Katsutoshi, Pšenička Martin Reviews in Aquaculture (2019-October-21) https://doi.org/ggddr9 DOI: 10.1111/raq.12389 [DOI] [Google Scholar]
  • 11.Defining the functional divergence of orthologous genes between human and mouse in the context of miRNA regulation; Cui Chunmei, Zhou Yuan, Cui Qinghua Briefings in Bioinformatics (2021-July-05) https://doi.org/gr86f6 DOI: 10.1093/bib/bbab253 [DOI] [PubMed] [Google Scholar]
  • 12.Testing the Ortholog Conjecture with Comparative Functional Genomic Data from Mammals; Nehrt Nathan L, Clark Wyatt T, Radivojac Predrag, Hahn Matthew W PLoS Computational Biology (2011-June-09) https://doi.org/cmg5m4 DOI: 10.1371/journal.pcbi.1002073 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.The ortholog conjecture revisited: the value of orthologs and paralogs in function prediction; Stamboulian Moses, Guerrero Rafael F, Hahn Matthew W, Radivojac Predrag Bioinformatics (2020-July-01) https://doi.org/gr96nw DOI: 10.1093/bioinformatics/btaa468 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Pervasive Variation of Transcription Factor Orthologs Contributes to Regulatory Network Evolution; Nadimpalli Shilpa, Persikov Anton V, Singh Mona PLOS Genetics (2015-March-06) https://doi.org/f68jmg DOI: 10.1371/journal.pgen.1005011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Divergence of Noncoding Regulatory Elements Explains Gene–Phenotype Differences between Human and Mouse Orthologous Genes; Han Seong Kyu, Kim Donghyo, Lee Heetak, Kim Inhae, Kim Sanguk Molecular Biology and Evolution (2018-April-24) https://doi.org/gdd4g2 DOI: 10.1093/molbev/msy056 [DOI] [Google Scholar]
  • 16.The origin and evolution of cell types; Arendt Detlev, Musser Jacob M, Baker Clare VH, Bergman Aviv, Cepko Connie, Erwin Douglas H, Pavlicev Mihaela, Schlosser Gerhard, Widder Stefanie, Laubichler Manfred D, Wagner Günter P Nature Reviews Genetics (2016-November-07) https://doi.org/f9b62x DOI: 10.1038/nrg.2016.127 [DOI] [PubMed] [Google Scholar]
  • 17.Tissue evolution: mechanical interplay of adhesion, pressure, and heterogeneity; Büscher Tobias, Ganai Nirmalendu, Gompper Gerhard, Elgeti Jens New Journal of Physics (2020-March-01) https://doi.org/gqpwh8 DOI: 10.1088/1367-2630/ab74a5 [DOI] [Google Scholar]
  • 18.Evolutionary rewiring of regulatory networks contributes to phenotypic differences between human and mouse orthologous genes; Ha Doyeon, Kim Donghyo, Kim Inhae, Oh Youngchul, Kong JungHo, Han Seong Kyu, Kim Sanguk Nucleic Acids Research (2022-February-07) https://doi.org/gtsx5v DOI: 10.1093/nar/gkac050 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Understanding the limits of animal models as predictors of human biology: lessons learned from the sbv IMPROVER Species Translation Challenge; Rhrissorrakrai Kahn, Belcastro Vincenzo, Bilal Erhan, Norel Raquel, Poussin Carine, Mathis Carole, Dulize Rémi HJ, Ivanov Nikolai V, Alexopoulos Leonidas, Rice J Jeremy, … Hoeng Julia Bioinformatics (2014-September-17) https://doi.org/f63qxc DOI: 10.1093/bioinformatics/btu611 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens; Zhou Naihui, Jiang Yuxiang, Bergquist Timothy R, Lee Alexandra J, Kacsoh Balint Z, Crocker Alex W, Lewis Kimberley A, Georghiou George, Nguyen Huy N, Hamid Md Nafiz, … Friedberg Iddo Genome Biology (2019-November-19) https://doi.org/ggnxpz DOI: 10.1186/s13059-019-1835-8 [DOI] [Google Scholar]
  • 21.An expanded evaluation of protein function prediction methods shows an improvement in accuracy; Jiang Yuxiang, Oron Tal Ronnen, Clark Wyatt T, Bankapur Asma R, D’Andrea Daniel, Lepore Rosalba, Funk Christopher S, Kahanda Indika, Verspoor Karin M, Ben-Hur Asa, … Radivojac Predrag Genome Biology (2016-September-07) https://doi.org/gg7qsc DOI: 10.1186/s13059-016-1037-6 [DOI] [Google Scholar]
  • 22.A large-scale evaluation of computational protein function prediction; Radivojac Predrag, T Clark Wyatt, Oron Tal Ronnen, Schnoes Alexandra M, Wittkop Tobias, Sokolov Artem, Graim Kiley, Funk Christopher, Verspoor Karin, Ben-Hur Asa, … Friedberg Iddo Nature Methods (2013-January-27) https://doi.org/gjh8q2 DOI: 10.1038/nmeth.2340 [DOI] [Google Scholar]
  • 23.Community-Wide Evaluation of Computational Function Prediction; Friedberg Iddo, Radivojac Predrag arXiv (2016) https://doi.org/gsc3dr DOI: 10.48550/arxiv.1601.01048 [DOI] [Google Scholar]
  • 24.Translating preclinical models to humans; Brubaker Douglas K, Lauffenburger Douglas A Science (2020-February-14) https://doi.org/ggk6dt DOI: 10.1126/science.aay8086 [DOI] [PubMed] [Google Scholar]
  • 25.Applications of comparative evolution to human disease genetics; McWhite Claire D, Liebeskind Benjamin J, Marcotte Edward M Current Opinion in Genetics & Development (2015-12) https://doi.org/gr86f5 DOI: 10.1016/j.gde.2015.08.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Transfer learning of clinical outcomes from preclinical molecular data, principles and perspectives; Kowald Axel, Barrantes Israel, Möller Steffen, Palmer Daniel, Escobar Hugo Murua, Schwerk Anne, Fuellen Georg Briefings in Bioinformatics (2022-April-23) https://doi.org/gr86f7 DOI: 10.1093/bib/bbac133 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Systems biology approaches help to facilitate interpretation of cross-species comparisons; Dougherty Bonnie V, Papin Jason A Current Opinion in Toxicology (2020-10) https://doi.org/gr86f4 DOI: 10.1016/j.cotox.2020.06.002 [DOI] [Google Scholar]
  • 28.A compendium of human gene functions derived from evolutionary modelling; Feuermann Marc, Mi Huaiyu, Gaudet Pascale, Muruganujan Anushya, Lewis Suzanna E, Ebert Dustin, Mushayahama Tremayne, Aleksander Suzanne A, Balhoff James, … Thomas Paul D Nature (2025-February-26) https://doi.org/g86gqc DOI: 10.1038/s41586-025-08592-0 [DOI] [Google Scholar]
  • 29.A Rapid Method for Directed Gene Knockout for Screening in G0 Zebrafish; Wu Roland S, Lam Ian I, Clay Hilary, Duong Daniel N, Deo Rahul C, Coughlin Shaun R Developmental Cell (2018-July) https://doi.org/gdv6c4 DOI: 10.1016/j.devcel.2018.06.003 [DOI] [PubMed] [Google Scholar]
  • 30.Loss of circadian rhythmicity in bdnf knockout zebrafish larvae; D’Agostino Ylenia, Frigato Elena, Noviello Teresa MR, Toni Mattia, Frabetti Flavia, Cigliano Luisa, Ceccarelli Michele, Sordino Paolo, Cerulo Luigi, Bertolucci Cristiano, D’Aniello Salvatore iScience (2022-04) https://doi.org/gr8cpc DOI: 10.1016/j.isci.2022.104054 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Leading the way: canine models of genomics and disease; Shearin Abigail L, Ostrander Elaine A Disease Models & Mechanisms (2010-January-14) https://doi.org/dqcgd5 DOI: 10.1242/dmm.004358 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.The Burmese python genome reveals the molecular basis for extreme adaptation in snakes; Castoe Todd A, de Koning APJason, Hall Kathryn T, Card Daren C, Schield Drew R, Fujita Matthew K, Ruggiero Robert P, Degner Jack F, Daza Juan M, Gu Wanjun, … Pollock David D Proceedings of the National Academy of Sciences (2013-December-02) https://doi.org/f5kfpx DOI: 10.1073/pnas.1314475110 [DOI] [Google Scholar]
  • 33.The spotted gar genome illuminates vertebrate evolution and facilitates human-teleost comparisons; Braasch Ingo, Gehrke Andrew R, Smith Jeramiah J, Kawasaki Kazuhiko, Manousaki Tereza, Pasquier Jeremy, Amores Angel, Desvignes Thomas, Batzel Peter, Catchen Julian, … Postlethwait John H Nature Genetics (2016-March-07) https://doi.org/f3rndn DOI: 10.1038/ng.3526 [DOI] [Google Scholar]
  • 34.The genome of Schmidtea mediterranea and the evolution of core cellular mechanisms; Grohme Markus Alexander, Schloissnig Siegfried, Rozanski Andrei, Pippel Martin, Young George Robert, Winkler Sylke, Brandl Holger, Henry Ian, Dahl Andreas, Powell Sean, … Rink Jochen Christian Nature (2018-January-24) https://doi.org/gcv62s DOI: 10.1038/nature25473 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function; Warde-Farley David, Donaldson Sylva L, Comes Ovi, Zuberi Khalid, Badrawi Rashad, Chao Pauline, Franz Max, Grouios Chris, Kazi Farzana, Lopes Christian Tannus, … Morris Quaid Nucleic Acids Research (2010-June-21) https://doi.org/bkds9d DOI: 10.1093/nar/gkq537 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Supervised learning is an accurate method for network-based gene classification; Liu Renming, Mancuso Christopher A, Yannakopoulos Anna, Johnson Kayla A, Krishnan Arjun Bioinformatics (2020-April-14) https://doi.org/gmvnfc DOI: 10.1093/bioinformatics/btaa150 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Network-based methods for human disease gene prediction; Wang X, Gulbahce N, Yu H Briefings in Functional Genomics (2011-July-15) https://doi.org/ctnj62 DOI: 10.1093/bfgp/elr024 [DOI] [PubMed] [Google Scholar]
  • 38.Prediction and Validation of Gene-Disease Associations Using Methods Inspired by Social Network Analyses; Singh-Blom UMartin, Natarajan Nagarajan, Tewari Ambuj, Woods John O, Dhillon Inderjit S, Marcotte Edward M PLoS ONE (2013-May-01) https://doi.org/f24vnt DOI: 10.1371/journal.pone.0058977 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.An integrative tissue-network approach to identify and test human disease genes; Yao Victoria, Kaletsky Rachel, Keyes William, Mor Danielle E, Wong Aaron K, Sohrabi Salman, Murphy Coleen T, Troyanskaya Olga G Nature Biotechnology (2018-October-22) https://doi.org/gfg4bd DOI: 10.1038/nbt.4246 · This study introduced DiseaseQUEST, a method to combine human GWAS-derived gene-disease associations and research organism functional networks to predict candidate disease-related genes in research organisms. [DOI] [Google Scholar]
  • 40.Accurate Quantification of Functional Analogy among Close Homologs; Chikina Maria D, Troyanskaya Olga G PLoS Computational Biology (2011-February-03) https://doi.org/bh988k DOI: 10.1371/journal.pcbi.1001074 · This method uses functional gene networks and metagenes of neighboring genes in these networks to find agnologs across species. [DOI] [Google Scholar]
  • 41.Global alignment of multiple protein interaction networks with application to functional orthology detection; Singh Rohit, Xu Jinbo, Berger Bonnie Proceedings of the National Academy of Sciences (2008-September-02) https://doi.org/cn7rgd DOI: 10.1073/pnas.0806627105 [DOI] [Google Scholar]
  • 42.Functional Knowledge Transfer for High-accuracy Prediction of Under-studied Biological Processes; Park Christopher Y, Wong Aaron K, Greene Casey S, Rowland Jessica, Guan Yuanfang, Bongo Lars A, Burdine Rebecca D, Troyanskaya Olga G PLoS Computational Biology (2013-March-14) https://doi.org/f4qtp9 DOI: 10.1371/journal.pcbi.1002957 · This study introduced the Functional Knowledge Transfer (FKT) method, a network-based approach that utilizes functional equivalents (which we call “agnologs”) to enhance gene annotation across species. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Mutations in N-cadherin and a Stardust homolog, Nagie oko, affect cell-cycle exit in zebrafish retina; Yamaguchi Masahiro, Imai Fumiyasu, Tonou-Fujimori Noriko, Masai Ichiro Mechanisms of Development (2010-05) https://doi.org/djzmf9 DOI: 10.1016/j.mod.2010.03.004 [DOI] [PubMed] [Google Scholar]
  • 44.Joint representation of molecular networks from multiple species improves gene classification; Mancuso Christopher A, Johnson Kayla A, Liu Renming, Krishnan Arjun PLOS Computational Biology (2024-January-10) https://doi.org/gtsx8m DOI: 10.1371/journal.pcbi.1011773 · This study introduced GenePlexusZoo, a method to simultaneously integrate network information across more than two species to improve gene classification. [DOI] [Google Scholar]
  • 45.NetQuilt: deep multispecies network-based protein function prediction using homology-informed network similarity; Barot Meet, Gligorijević Vladimir, Cho Kyunghyun, Bonneau Richard Bioinformatics (2021-February-12) https://doi.org/gk2rt6 DOI: 10.1093/bioinformatics/btab098 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.A New Status Index Derived from Sociometric Analysis; Katz Leo Psychometrika (1953-03) https://doi.org/c6g25t DOI: 10.1007/bf02289026 [DOI] [Google Scholar]
  • 47.The link prediction problem for social networks; Liben-Nowell David, Kleinberg Jon Proceedings of the twelfth international conference on Information and knowledge management (2003-November-03) https://doi.org/c7ttqw DOI: 10.1145/956863.956972 [DOI] [Google Scholar]
  • 48.Large-Scale Discovery of Disease-Disease and Disease-Gene Associations; Gligorijevic Djordje, Stojanovic Jelena, Djuric Nemanja, Radosavljevic Vladan, Grbovic Mihajlo, Kulathinal Rob J, Obradovic Zoran Scientific Reports (2016-August-31) https://doi.org/f8znpb DOI: 10.1038/srep32404 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Understanding Tissue-Specific Gene Regulation; Sonawane Abhijeet Rajendra, Platig John, Fagny Maud, Chen Cho-Yi, Paulson Joseph Nathaniel, Lopes-Ramos Camila Miranda, DeMeo Dawn Lisa, Quackenbush John, Glass Kimberly, Kuijjer Marieke Lydia Cell Reports (2017-10) https://doi.org/ggcrpf DOI: 10.1016/j.celrep.2017.10.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Mechanisms of tissue and cell-type specificity in heritable traits and diseases; Hekselman Idan, Yeger-Lotem Esti Nature Reviews Genetics (2020-January-08) https://doi.org/ggkx9v DOI: 10.1038/s41576-019-0200-9 [DOI] [PubMed] [Google Scholar]
  • 51.Understanding multicellular function and disease with human tissue-specific networks; Greene Casey S, Krishnan Arjun, Wong Aaron K, Ricciotti Emanuela, Zelaya Rene A, Himmelstein Daniel S, Zhang Ran, Hartmann Boris M, Zaslavsky Elena, Sealfon Stuart C, … Troyanskaya Olga G Nature Genetics (2015-April-27) https://doi.org/f7dvkv DOI: 10.1038/ng.3259 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Network propagation: a universal amplifier of genetic associations; Cowen Lenore, Ideker Trey, Raphael Benjamin J, Sharan Roded Nature Reviews Genetics (2017-June-12) https://doi.org/gbhkwn DOI: 10.1038/nrg.2017.38 [DOI] [PubMed] [Google Scholar]
  • 53.Benchmarking network propagation methods for disease gene identification; Picart-Armada Sergio, Barrett Steven J, Willé David R, Perera-Lluna Alexandre, Gutteridge Alex, Dessailly Benoit H PLOS Computational Biology (2019-September-03) https://doi.org/gtsx8k DOI: 10.1371/journal.pcbi.1007276 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Genome-wide prediction and functional characterization of the genetic basis of autism spectrum disorder; Krishnan Arjun, Zhang Ran, Yao Victoria, Theesfeld Chandra L, Wong Aaron K, Tadych Alicja, Volfovsky Natalia, Packer Alan, Lash Alex, Troyanskaya Olga G Nature Neuroscience (2016-August-01) https://doi.org/f889pv DOI: 10.1038/nn.4353 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.XGDAG: explainable gene–disease associations via graph neural networks; Mastropietro Andrea, De Carlo Gianluca, Anagnostopoulos Aris Bioinformatics (2023-August-01) https://doi.org/gtsx8h DOI: 10.1093/bioinformatics/btad482 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.An Ancient Evolutionary Origin of Genes Associated with Human Genetic Diseases; Domazet-Loso T, Tautz D Molecular Biology and Evolution (2008-August-05) https://doi.org/bpbtgp DOI: 10.1093/molbev/msn214 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Similarly Strong Purifying Selection Acts on Human Disease Genes of All Evolutionary Ages; Cai James J, Borenstein Elhanan, Chen Rong, Petrov Dmitri A Genome Biology and Evolution (2009-January-01) https://doi.org/bgbqrk DOI: 10.1093/gbe/evp013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Genome-wide identification of genes likely to be involved in human genetic disease; Lopez-Bigas N Nucleic Acids Research (2004-June-02) https://doi.org/bbzrfz DOI: 10.1093/nar/gkh605 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.On the Origins of Mendelian Disease Genes in Man: The Impact of Gene Duplication; Dickerson JE, Robertson DL Molecular Biology and Evolution (2011-June-24) https://doi.org/czx744 DOI: 10.1093/molbev/msr111 [DOI] [Google Scholar]
  • 60.Assigning protein functions by comparative genome analysis: Protein phylogenetic profiles; Pellegrini Matteo, Marcotte Edward M, Thompson Michael J, Eisenberg David, Yeates Todd O Proceedings of the National Academy of Sciences (1999-April-13) https://doi.org/br9862 DOI: 10.1073/pnas.96.8.4285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Detecting Protein Function and Protein-Protein Interactions from Genome Sequences; Edward M Marcotte, Pellegrini Matteo, Ng Ho-Leung, Rice Danny W, Yeates Todd O, Eisenberg David Science (1999-July-30) https://doi.org/fxg2p3 DOI: 10.1126/science.285.5428.751 [DOI] [PubMed] [Google Scholar]
  • 62.Evolutionary profiling reveals the heterogeneous origins of classes of human disease genes: implications for modeling disease genetics in animals; Maxwell Evan K, Schnitzler Christine E, Havlak Paul, Putnam Nicholas H, Nguyen Anh-Dao, Moreland RTravis, Baxevanis Andreas D BMC Evolutionary Biology (2014-October-04) https://doi.org/f6qffg DOI: 10.1186/s12862-014-0212-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.OMIM.org: leveraging knowledge across phenotype–gene relationships; Amberger Joanna S, Bocchini Carol A, Scott Alan F, Hamosh Ada Nucleic Acids Research (2018-November-16) https://doi.org/gjxh3j DOI: 10.1093/nar/gky1151 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Human disease locus discovery and mapping to molecular pathways through phylogenetic profiling; Tabach Yuval, Golan Tamar, Hernández-Hernández Abrahan, Messer Arielle R, Fukuda Tomoyuki Kouznetsova Anna, Liu Jian-Guo, Lilienthal Ingrid, Levy Carmit, Ruvkun Gary Molecular Systems Biology (2013-01) https://doi.org/f2pc78 DOI: 10.1038/msb.2013.50 · This study uses phylogenetic profiles to enable systematic identifications of genome-wide functional modules that co-occur across the eukaryotic phylogeny. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Systematic Discovery of Human Gene Function and Principles of Modular Organization through Phylogenetic Profiling; Dey Gautam, Jaimovich Ariel, Collins Sean R, Seki Akiko, Meyer Tobias Cell Reports (2015-02) https://doi.org/ggvqsm DOI: 10.1016/j.celrep.2015.01.025 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.High-throughput mouse phenomics for characterizing mammalian gene function; Brown Steve DM, Holmes Chris C, Mallon Ann-Marie, Meehan Terrence F, Smedley Damian, Wells Sara Nature Reviews Genetics (2018-April-06) https://doi.org/gr8fcd DOI: 10.1038/s41576-018-0005-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Progress towards completing the mutant mouse null resource; Peterson Kevin A, Murray Stephen A Mammalian Genome (2021-October-26) https://doi.org/gr8fcc DOI: 10.1007/s00335-021-09905-0 [DOI] [Google Scholar]
  • 68.Harmonizing model organism data in the Alliance of Genome Resources,; Agapite Julie, Albou Laurent-Philippe, Aleksander Suzanne A, Alexander Micheal, Anagnostopoulos Anna V, Antonazzo Giulia, Argasinska Joanna, Arnaboldi Valerio, Attrill Helen, … Zytkovicz Mark Genetics (2022-February-25) https://doi.org/gr8fcg DOI: 10.1093/genetics/iyac022 [DOI] [Google Scholar]
  • 69.The Unified Phenotype Ontology : a framework for cross-species integrative phenomics; Matentzoglu Nicolas, Bello Susan M, Stefancsik Ray, Alghamdi Sarah M, Anagnostopoulos Anna V, Balhoff James P, Balk Meghan A, Bradford Yvonne M, Bridges Yasemin, Callahan Tiffany J, … Osumi-Sutherland David GENETICS (2025-03) https://doi.org/g9ttnm DOI: 10.1093/genetics/iyaf027 [DOI] [Google Scholar]
  • 70.PhenomeNET: a whole-phenome approach to disease gene discovery; Hoehndorf R, Schofield PN, Gkoutos GV Nucleic Acids Research (2011-July-06) https://doi.org/cmvr4t DOI: 10.1093/nar/gkr538 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Next-generation diagnostics and disease-gene discovery with the Exomiser; Smedley Damian, Jacobsen Julius OB, Jäger Marten, Köhler Sebastian, Holtgrewe Manuel, Schubach Max, Siragusa Enrico, Zemojtel Tomasz, Buske Orion J, Washington Nicole L, … Robinson Peter N Nature Protocols (2015-November-12) https://doi.org/f73qtp DOI: 10.1038/nprot.2015.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Contribution of model organism phenotypes to the computational identification of human disease genes; Alghamdi Sarah M, Schofield Paul N, Hoehndorf Robert Disease Models & Mechanisms (2022-July-01) https://doi.org/gr8fch DOI: 10.1242/dmm.049441 · This study uses phenotype ontologies to relate research organisms-specific and human-specific ontologies. By treating connected ontologies as graphs, they also proposed using graph embedding-based machine learning approaches to predict disease-gene associations across species. [DOI] [Google Scholar]
  • 73.Prioritizing genomic variants through neuro-symbolic, knowledge-enhanced learning; Althagafi Azza, Zhapa-Camacho Fernando, Hoehndorf Robert Bioinformatics (2024-May-01) https://doi.org/gt6n4x DOI: 10.1093/bioinformatics/btae301 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Convergent Evolution of Enzyme Active Sites Is not a Rare Phenomenon; Gherardini Pier Federico, Wass Mark N, Helmer-Citterich Manuela, Sternberg Michael JE Journal of Molecular Biology (2007-09) https://doi.org/cc7qn6 DOI: 10.1016/j.jmb.2007.06.017 [DOI] [PubMed] [Google Scholar]
  • 75.Gene analogue finder: a GRID solution for finding functionally analogous gene products; Tulipano Angelica, Donvito Giacinto, Licciulli Flavio, Maggi Giorgio, Gisel Andreas BMC Bioinformatics (2007-September-03) https://doi.org/dtn246 DOI: 10.1186/1471-2105-8-329 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.TreeFam: a curated database of phylogenetic trees of animal gene families; Li H Nucleic Acids Research (2006-January-01) https://doi.org/fss2j9 DOI: 10.1093/nar/gkj118 [DOI] [Google Scholar]
  • 77.Functional protein representations from biological networks enable diverse cross-species inference; Fan Jason, Cannistra Anthony, Fried Inbar, Lim Tim, Schaffner Thomas, Crovella Mark, Hescott Benjamin, Leiserson Mark DM Nucleic Acids Research (2019-March-08) https://doi.org/gmdn7p DOI: 10.1093/nar/gkz132 · This study introduces MUNK, a method that utilizes aligned protein-protein interaction network representations to identify agnologous gene pairs. Additionally, this work expands the concept of Phenologs, characterizing them by a significant overlap of agnologous gene pairs. [DOI] [Google Scholar]
  • 78.MUNDO: protein function prediction embedded in a multispecies world; Arsenescu Victor, Devkota Kapil, Erden Mert, Shpilker Polina, Werenski Matthew, Cowen Lenore J Bioinformatics Advances (2021-September-29) https://doi.org/gtsx8g DOI: 10.1093/bioadv/vbab025 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Joint embedding of biological networks for cross-species functional alignment; Li Lechuan, Dannenfelser Ruth, Zhu Yu, Hejduk Nathaniel, Segarra Santiago, Yao Vicky Bioinformatics (2023-August-26) https://doi.org/gtsx8j DOI: 10.1093/bioinformatics/btad529 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Evolutionary-scale prediction of atomic-level protein structure with a language model; Lin Zeming, Akin Halil, Rao Roshan, Hie Brian, Zhu Zhongkai, Lu Wenting, Smetanin Nikita, Verkuil Robert, Kabeli Ori, Shmueli Yaniv, … Rives Alexander Science (2023-March-17) https://doi.org/grzgts DOI: 10.1126/science.ade2574 [DOI] [PubMed] [Google Scholar]
  • 81.Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences; Rives Alexander, Meier Joshua, Sercu Tom, Goyal Siddharth, Lin Zeming, Liu Jason, Guo Demi, Ott Myle, Zitnick CLawrence, Ma Jerry, Fergus Rob Proceedings of the National Academy of Sciences (2021-April-05) https://doi.org/gmmft5 DOI: 10.1073/pnas.2016239118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Protein remote homology detection and structural alignment using deep learning; Hamamsy Tymor, Morton James T, Blackwell Robert, Berenberg Daniel, Carriero Nicholas, Gligorijevic Vladimir, Strauss Charlie EM, Leman Julia Koehler, Cho Kyunghyun, Richard Bonneau Nature Biotechnology (2023-September-07) https://doi.org/gss86j DOI: 10.1038/s41587-023-01917-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Fast, sensitive detection of protein homologs using deep dense retrieval; Hong Liang, Hu Zhihang, Sun Siqi, Tang Xiangru, Wang Jiuming, Tan Qingxiong, Zheng Liangzhen, Wang Sheng, Xu Sheng, King Irwin, … Li Yu Nature Biotechnology (2024-August-09) https://doi.org/g7b5zf DOI: 10.1038/s41587-024-02353-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Genome modeling and design across all domains of life with Evo 2; Brixi Garyk, Durrant Matthew G, Ku Jerome, Poli Michael, Brockman Greg, Chang Daniel, Gonzalez Gabriel A, King Samuel H, Li David B, Merchant Aditi T, … Hie Brian L Cold Spring Harbor Laboratory (2025-February-21) https://doi.org/g85t7g DOI: 10.1101/2025.02.18.638918 [DOI] [Google Scholar]
  • 85.Systematic discovery of nonobvious human disease models through orthologous phenotypes; McGary Kriston L, Park Tae Joo, Woods John O, Cha Hye Ji, Wallingford John B, Marcotte Edward M Proceedings of the National Academy of Sciences (2010-March-22) https://doi.org/b8hnkn DOI: 10.1073/pnas.0910200107 · This study introduced the concept of Phenologs, which facilitates finding phenotypic counterparts across species by looking at ortholog overlap between genes annotated to phenotypes in different species. [DOI] [Google Scholar]
  • 86.CoCoCoNet: conserved and comparative co-expression across a diverse set of species; Lee John, Shah Manthan, Ballouz Sara, Crow Megan, Gillis Jesse Nucleic Acids Research (2020-May-11) https://doi.org/ghjzdz DOI: 10.1093/nar/gkaa348 · This study introduced CoCoCoNet, a method that helps to find agnologous gene sets across species by comparing co-expression networks generated in different species. [DOI] [Google Scholar]
  • 87.xHeinz: an algorithm for mining cross-species network modules under a flexible conservation model; El-Kebir Mohammed, Soueidan Hayssam, Hume Thomas, Beisser Daniela, Dittrich Marcus, Müller Tobias, Blin Guillaume, Heringa Jaap, Nikolski Macha, Wessels Lodewyk FA, Klau Gunnar W Bioinformatics (2015-May-27) https://doi.org/f7vck8 DOI: 10.1093/bioinformatics/btv316 [DOI] [PubMed] [Google Scholar]
  • 88.A Scalable Approach for Discovering Conserved Active Subnetworks across Species; Deshpande Raamesh, Sharma Shikha, Catherine M Verfaillie Wei-Shou Hu, Myers Chad L PLoS Computational Biology (2010-December-09) https://doi.org/bfn9rk DOI: 10.1371/journal.pcbi.1001028 · This study introduced neXus, a method that integrates differentially expressed gene lists and network topology to find agnologous gene modules across species. [DOI] [Google Scholar]
  • 89.ModuleBlast: identifying activated sub-networks within and across species; Zinman Guy E, Naiman Shoshana, O’Dee Dawn M, Kumar Nishant, Nau Gerard J, Cohen Haim Y, Bar-Joseph Ziv Nucleic Acids Research (2014-November-26) https://doi.org/f66rt7 DOI: 10.1093/nar/gku1224 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Transcriptomic congruence analysis for evaluating model organisms; Zong Wei, Rahman Tanbin, Zhu Li, Zeng Xiangrui, Zhang Yingjin, Zou Jian, Liu Song, Ren Zhao, Li Jingyi Jessica, Sibille Etienne, … Tseng George C Proceedings of the National Academy of Sciences (2023-February-02) https://doi.org/grqm9w DOI: 10.1073/pnas.2202584120 [DOI] [PMC free article] [PubMed] [Google Scholar]; This study introduced CAMO, a method to systematically quantify the similarity between perturbed transcriptomic profiles across species.
  • 91.Genomic responses in mouse models poorly mimic human inflammatory diseases; Seok Junhee, Warren HShaw, Cuenca Alex G, Mindrinos Michael N, Baker Henry V, Xu Weihong, Richards Daniel R, McDonald-Smith Grace P, Gao Hong, Hennessy Laura, … Wong Wing H Proceedings of the National Academy of Sciences (2013-February-11) https://doi.org/f4p74d DOI: 10.1073/pnas.1222878110 [DOI] [Google Scholar]
  • 92.Genomic responses in mouse models greatly mimic human inflammatory diseases; Takao Keizo, Miyakawa Tsuyoshi Proceedings of the National Academy of Sciences (2014-August-04) https://doi.org/f6zqzv DOI: 10.1073/pnas.1401965111 [DOI] [Google Scholar]
  • 93.The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species; Putman Tim E, Schaper Kevin, Matentzoglu Nicolas, Rubinetti Vincent P, Alquaddoomi Faisal S, Cox Corey, Caufield JHarry, Elsarboukh Glass, Gehrke Sarah, Hegde Harshad, … Munoz-Torres Monica C Nucleic Acids Research (2023-November-24) https://doi.org/gs6kmr DOI: 10.1093/nar/gkad1082 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Found In Translation: a machine learning model for mouse-to-human inference; Normand Rachelly, Du Wenfei, Briller Mayan, Gaujoux Renaud, Starosvetsky Elina, Ziv-Kenet Amit, Shalev-Malul Gali, Tibshirani Robert J, Shen-Orr Shai S Nature Methods (2018-November-26) https://doi.org/gfkfvg DOI: 10.1038/s41592-018-0214-9 This study introduced FIT, a method that investigated the possibility of transferring DEG results across species using linear models. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Computational translation of genomic responses from experimental model systems to humans; Brubaker Douglas K, Proctor Elizabeth A, Haigis Kevin M, Lauffenburger Douglas A PLOS Computational Biology (2019-January-10) https://doi.org/gr8cpf DOI: 10.1371/journal.pcbi.1006286 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.An interspecies translation model implicates integrin signaling in infliximab-resistant inflammatory bowel disease; Brubaker Douglas K, Kumar Manu P, Chiswick Evan L, Gregg Cecil, Starchenko Alina, Vega Paige N, Southard-Smith Austin N, Simmons Alan J, Scoville Elizabeth A, Coburn Lori A, … Lauffenburger Douglas A Science Signaling (2020-August-04) https://doi.org/gqbd3g DOI: 10.1126/scisignal.aay3258 · This study introduced Transcomp-R, a method that allows transferring knowledge across species when the data from each species was not measured using the same omics modality. [DOI] [Google Scholar]
  • 97.Cross-species transcriptomic signatures predict response to MK2 inhibition in mouse models of chronic inflammation; Suarez-Lopez Lucia, Shui Bing, Brubaker Douglas K, Hill Marza, Bergendorf Alexander, Changelian Paul S, Laguna Aisha, Starchenko Alina, Lauffenburger Douglas A, Haigis Kevin M iScience (2021-12) https://doi.org/g879jq DOI: 10.1016/j.isci.2021.103406 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Computational Interspecies Translation Between Alzheimer’s Disease Mouse Models and Human Subjects Identifies Innate Immune Complement, TYROBP, and TAM Receptor Agonist Signatures, Distinct From Influences of Aging; Lee Meelim J, Wang Chuangqi, Carroll Molly J, Brubaker Douglas K, Hyman Bradley T, Lauffenburger Douglas A Frontiers in Neuroscience (2021-September-30) https://doi.org/g879js DOI: 10.3389/fnins.2021.727784 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Translatable pathways classification (TransPath-C) for inferring processes germane to human biology from animal studies data: example application in neurobiology; Carroll Molly J, Garcia-Reyero Natàlia, Perkins Edward J, Lauffenburger Douglas A Integrative Biology (2021-10) https://doi.org/gr8hv2 DOI: 10.1093/intbio/zyab016 [DOI] [PubMed] [Google Scholar]
  • 100.Style transfer with variational autoencoders is a promising approach to RNA-Seq data harmonization and analysis; Russkikh Nikolai, Antonets Denis, Shtokalo Dmitry, Makarov Alexander, Vyatkin Yuri, Zakharov Alexey, Terentyev Evgeny Bioinformatics (2020-July-16) https://doi.org/gjrszt DOI: 10.1093/bioinformatics/btaa624 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.scGen predicts single-cell perturbation responses; Lotfollahi Mohammad, Wolf FAlexander, Theis Fabian J Nature Methods (2019-July-29) https://doi.org/gf9t5z DOI: 10.1038/s41592-019-0494-8 [DOI] [PubMed] [Google Scholar]
  • 102.AutoTransOP: translating omics signatures without orthologue requirements using deep learning; Meimetis Nikolaos, Pullen Krista M, Zhu Daniel Y, Nilsson Avlant, Hoang Trong Nghia, Magliacane Sara, Lauffenburger Douglas A npj Systems Biology and Applications (2024-January-29) https://doi.org/gtg885 DOI: 10.1038/s41540-024-00341-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles; Subramanian Aravind, Tamayo Pablo, Mootha Vamsi K, Mukherjee Sayan, Ebert Benjamin L, Gillette Michael A, Paulovich Amanda, Pomeroy Scott L, Golub Todd R, Lander Eric S, Mesirov Jill P Proceedings of the National Academy of Sciences (2005-September-30) https://doi.org/d4qbh8 DOI: 10.1073/pnas.0506580102 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.XGSEA: CROSS-species gene set enrichment analysis via domain adaptation; Cai Menglan, Nguyen Canh Hao, Mamitsuka Hiroshi, Li Limin Briefings in Bioinformatics (2021-January-30) https://doi.org/gr8hvz DOI: 10.1093/bib/bbaa406 [DOI] [PubMed] [Google Scholar]
  • 105.The ancestral gene repertoire of animal stem cells; Alexandre Alié, Hayashi Tetsutaro, Sugimura Itsuro, Manuel Michaël, Sugano Wakana, Mano Akira, Satoh Nori, Agata Kiyokazu, Funayama Noriko Proceedings of the National Academy of Sciences (2015-December-07) https://doi.org/gf3xjn DOI: 10.1073/pnas.1514789112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Cross-Species Analysis of Single-Cell Transcriptomic Data; Shafer Maxwell ER Frontiers in Cell and Developmental Biology (2019-September-02) https://doi.org/gg2rpw DOI: 10.3389/fcell.2019.00175 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Benchmarking atlas-level data integration in single-cell genomics; Luecken Malte D, Büttner M, Chaichoompu K, Danese A, Interlandi M, Mueller MF, Strobl DC, Zappia L, Dugas M, Colomé-Tatché M, Theis Fabian J Nature Methods (2021-December-23) https://doi.org/gnvvd5 DOI: 10.1038/s41592-021-01336-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Integrating single-cell transcriptomic data across different conditions, technologies, and species; Butler Andrew, Hoffman Paul, Smibert Peter, Papalexi Efthymia, Satija Rahul Nature Biotechnology (2018-April-02) https://doi.org/gc87v6 DOI: 10.1038/nbt.4096 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Single-Cell Multi-omic Integration Compares and Contrasts Features of Brain Cell Identity; Welch Joshua D, Kozareva Velina, Ferreira Ashley, Vanderburg Charles, Martin Carly, Macosko Evan Z Cell (2019-06) https://doi.org/gf3m3v DOI: 10.1016/j.cell.2019.05.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Variable paralog expression underlies phenotype variation; Bailon-Zambrano Raisa, Sucharov Juliana, Mumme-Monheit Abigail, Murry Matthew, Stenzel Amanda, Pulvino Anthony T, Mitchell Jennyfer M, Colborn Kathryn L, Nichols James T eLife (2022-September-22) https://doi.org/gt6n5b DOI: 10.7554/elife.79247 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.OrthoDisease: tracking disease gene orthologs across 100 species; Forslund K, Schreiber F, Thanintorn N, Sonnhammer ELL Briefings in Bioinformatics (2011-May-12) https://doi.org/fhwbjk DOI: 10.1093/bib/bbr024 [DOI] [PubMed] [Google Scholar]
  • 112.Benchmarking strategies for cross-species integration of single-cell RNA sequencing data; Song Yuyao, Miao Zhichao, Brazma Alvis, Papatheodorou Irene Nature Communications (2023-October-14) https://doi.org/gsv2bj DOI: 10.1038/s41467-023-41855-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Mapping single-cell atlases throughout Metazoa unravels cell type evolution; Tarashansky Alexander J, Musser Jacob M, Khariton Margarita, Li Pengyang, Arendt Detlev, Quake Stephen R, Wang Bo eLife (2021-May-04) https://doi.org/gkczvh DOI: 10.7554/elife.66747 · This study introduced SAMap, a method that fully utilizes complex homology patterns to align single-cell transcriptomes across a pair of species. [DOI] [Google Scholar]
  • 114.Toward universal cell embeddings: integrating single-cell RNA-seq datasets across species with SATURN; Rosen Yanay, Maria Brbić Yusuf Roohani, Swanson Kyle, Li Ziang, Leskovec Jure Nature Methods (2024-February-16) https://doi.org/gtsx8f DOI: 10.1038/s41592-024-02191-z · This study introduced SATURN, a method that uses all genes, regardless of their evolutionary relationships, to align single-cell transcriptomes across species. SATURN is also able to simultaneously integrate single-cell transcriptomes from more than two species. [DOI] [Google Scholar]
  • 115.Transfer learning enables predictions in network biology; Theodoris Christina V, Xiao Ling, Chopra Anant, Chaffin Mark D, Al Sayed Zeina R, Hill Matthew C, Mantineo Helene, Brydon Elizabeth M, Zeng Zexian, Liu XShirley, Ellinor Patrick T Nature (2023-May-31) https://doi.org/gr9x63 DOI: 10.1038/s41586-023-06139-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.scGPT: toward building a foundation model for single-cell multi-omics using generative AI; Cui Haotian, Wang Chloe, Maan Hassaan, Pang Kuan, Luo Fengning, Duan Nan, Wang Bo Nature Methods (2024-February-26) https://doi.org/gtkxpk DOI: 10.1038/s41592-024-02201-0 [DOI] [PubMed] [Google Scholar]
  • 117.Precious3GPT: Multimodal Multi-Species Multi-Omics Multi-Tissue Transformer for Aging Research and Drug Discovery; Galkin Fedor, Naumov Vladimir, Pushkov Stefan, Sidorenko Denis, Urban Anatoly, Zagirova Diana, Alawi Khadija M, Aliper Alex, Gumerov Ruslan, Kalashnikov Aleksandr, … Zhavoronkov Alex Cold Spring Harbor Laboratory (2024-July-25) https://doi.org/g879jr DOI: 10.1101/2024.07.25.605062 [DOI] [Google Scholar]
  • 118.GeneCompass: deciphering universal gene regulatory mechanisms with a knowledge-informed cross-species foundation model; Yang Xiaodong, Liu Guole, Feng Guihai, Bu Dechao, Wang Pengfei, Jiang Jie, Chen Shubai, Yang Qinmeng, Miao Hefan, Zhang Yiyang, … Li Xin Cell Research (2024-October-08) https://doi.org/g7r673 DOI: 10.1038/s41422-024-01034-y · This method employs a transformer-based architecture to align agnoglogous cells across species while integrating auxiliary information such as gene regulatory networks, gene promoters, gene families, and co-expression networks. [DOI] [Google Scholar]
  • 119.Universal Cell Embeddings: A Foundation Model for Cell Biology; Rosen Yanay, Roohani Yusuf, Agrawal Ayush, Samotorcan Leon, Tabula Sapiens Consortium, Quake Stephen R, Leskovec Jure Cold Spring Harbor Laboratory (2023-November-29) https://doi.org/gs7tcx DOI: 10.1101/2023.11.28.568918 [DOI] [Google Scholar]
  • 120.Modeling Human Disease Phenotype in Model Organisms; Marian Ali J Circulation Research (2011-August-05) https://doi.org/b2mhvn DOI: 10.1161/circresaha.111.249409 [DOI] [Google Scholar]
  • 121.A Pathway-Based View of Human Diseases and Disease Relationships; Li Yong, Agarwal Pankaj PLoS ONE (2009-February-04) https://doi.org/d7483w DOI: 10.1371/journal.pone.0004346 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Shared Biological Pathways Between Alzheimer’s Disease and Ischemic Stroke; Cui Pan, Ma Xiaofeng, Li He, Lang Wenjing, Hao Junwei Frontiers in Neuroscience (2018-September-07) https://doi.org/gfbdpb DOI: 10.3389/fnins.2018.00605 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Unravelling the Shared Genetic Mechanisms Underlying 18 Autoimmune Diseases Using a Systems Approach; Gokuladhas Sreemol, Schierding William, Golovina Evgeniia, Fadason Tayaza, O’Sullivan Justin Frontiers in Immunology (2021-August-13) https://doi.org/gt6n46 DOI: 10.3389/fimmu.2021.693142 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Molecular Processes Involved in the Shared Pathways between Cardiovascular Diseases and Diabetes; Tokarek Julita, Budny Emilian, Saar Maciej, Stańczak Kamila, Wojtanowska Ewa, Młynarska Ewelina, Rysz Jacek, Franczyk Beata Biomedicines (2023-September-23) https://doi.org/gs3jxw DOI: 10.3390/biomedicines11102611 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Mapping biological process relationships and disease perturbations within a pathway network; Stoney Ruth, Robertson David L, Nenadic Goran, Schwartz Jean-Marc npj Systems Biology and Applications (2018-June-11) https://doi.org/gt6n4v DOI: 10.1038/s41540-018-0055-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.A network-based approach for isolating the chronic inflammation gene signatures underlying complex diseases towards finding new treatment opportunities; Hickey Stephanie L, McKim Alexander, Mancuso Christopher A, Krishnan Arjun Frontiers in Pharmacology (2022-October-12) https://doi.org/gt6n47 DOI: 10.3389/fphar.2022.995459 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Network analysis reveals rare disease signatures across multiple levels of biological organization; Buphamalai Pisanu, Kokotovic Tomislav, Nagy Vanja, Menche Jörg Nature Communications (2021-November-09) https://doi.org/gnpnsc DOI: 10.1038/s41467-021-26674-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.The Human Phenotype Ontology in 2024: phenotypes around the world; Gargano Michael A, Matentzoglu Nicolas, Coleman Ben, Addo-Lartey Eunice B, Anagnostopoulos Anna V, Anderton Joel, Avillach Paul, Bagley Anita M, Bakštein Eduard, Balhoff James P, … Robinson Peter N Nucleic Acids Research (2023-November-11) https://doi.org/gt6qm8 DOI: 10.1093/nar/gkad1005 [DOI] [Google Scholar]
  • 129.Common genetic variation associated with Mendelian disease severity revealed through cryptic phenotype analysis; Blair David R, Hoffmann Thomas J, Shieh Joseph T Nature Communications (2022-June-27) https://doi.org/gt6qm7 DOI: 10.1038/s41467-022-31030-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Curation and expansion of the Human Phenotype Ontology for systemic autoinflammatory diseases improves phenotype-driven disease-matching; Maassen Willem, Legger Geertje, Cinar Ovgu Kul, van Daele Paul, Gattorno Marco, Bader-Meunier Brigitte, Wouters Carine, Briggs Tracy, Johansson Lennart, van der Velde Joeri, … van Gijn Marielle Frontiers in Immunology (2023-September-12) https://doi.org/gt6qm9 DOI: 10.3389/fimmu.2023.1215869 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Gene Loss and Evolutionary Rates Following Whole-Genome Duplication in Teleost Fishes; Brunet Frédéric G, Crollius Hugues Roest, Paris Mathilde, Aury Jean-Marc, Gibert Patricia, Jaillon Olivier, Laudet Vincent, Robinson-Rechavi Marc Molecular Biology and Evolution (2006-June-29) https://doi.org/c4tn7b DOI: 10.1093/molbev/msl049 [DOI] [PubMed] [Google Scholar]
  • 132.Developmental system drift and flexibility in evolutionary trajectories; True John R, Haag Eric S Evolution & Development (2001-03) https://doi.org/fd5m8w DOI: 10.1046/j.1525-142x.2001.003002109.x [DOI] [PubMed] [Google Scholar]
  • 133.Interactome evolution: insights from genome-wide analyses of protein–protein interactions; Ghadie Mohamed A, Coulombe-Huntington Jasmin, Xia Yu Current Opinion in Structural Biology (2018-06) https://doi.org/gd7tzw DOI: 10.1016/j.sbi.2017.10.012 [DOI] [PubMed] [Google Scholar]
  • 134.Inter-Species Rescue of Mutant Phenotype—The Standard for Genetic Analysis of Human Genetic Disorders in Drosophila melanogaster Model; Ecovoiu Alexandru Al, Ratiu Attila Cristian, Micheu Miruna Mihaela, Chifiriuc Mariana Carmen International Journal of Molecular Sciences (2022-February-27) https://doi.org/g9q9bt DOI: 10.3390/ijms23052613 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Systematic humanization of yeast genes reveals conserved functions and genetic modularity; Kachroo Aashiq H, Laurent Jon M, Yellman Christopher M, Meyer Austin G, Wilke Claus O, Marcotte Edward M Science (2015-May-22) https://doi.org/f7czv3 DOI: 10.1126/science.aaa0769 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Evolution; Futuyma Douglas J Sinauer Associates, Inc. Publishers; (2013) ISBN: 9781605351155 [Google Scholar]
  • 137.The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest; Szklarczyk Damian, Kirsch Rebecca, Koutrouli Mikaela, Nastou Katerina, Mehryary Farrokh, Hachilif Radja, Gable Annika L, Fang Tao, Doncheva Nadezhda T, Pyysalo Sampo, … von Mering Christian Nucleic Acids Research (2022-November-12) https://doi.org/gs7sn3 DOI: 10.1093/nar/gkac1000 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction; Szymborski Joseph, Emad Amin Cold Spring Harbor Laboratory (2024-February-16) https://doi.org/gt6n44 DOI: 10.1101/2024.02.13.580150 [DOI] [Google Scholar]
  • 139.Protein Interaction Mapping in C. elegans Using Proteins Involved in Vulval Development; Walhout Albertha JM, Sordella Raffaella, Lu Xiaowei, Hartley James L, Temple Gary F, Brasch Michael A, Thierry-Mieg Nicolas, Vidal Marc Science (2000-January-07) https://doi.org/db3c99 DOI: 10.1126/science.287.5450.116 [DOI] [PubMed] [Google Scholar]
  • 140.Annotation Transfer Between Genomes: Protein–Protein Interologs and Protein–DNA Regulogs; Yu Haiyuan, Luscombe Nicholas M, Lu Hao Xin, Zhu Xiaowei, Xia Yu, Han Jing-Dong J, Bertin Nicolas, Chung Sambath, Vidal Marc, Gerstein Mark Genome Research (2004-06) https://doi.org/fn7gwg DOI: 10.1101/gr.1774904 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.IMP 2.0: a multi-species functional genomics portal for integration, visualization and prediction of protein functions and networks; Wong Aaron K, Krishnan Arjun, Yao Victoria, Tadych Alicja, Troyanskaya Olga G Nucleic Acids Research (2015-May-12) https://doi.org/f7nxbp DOI: 10.1093/nar/gkv486 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Machine learning methods to model multicellular complexity and tissue specificity; Sealfon Rachel SG, Wong Aaron K, Troyanskaya Olga G Nature Reviews Materials (2021-July-15) https://doi.org/gr83px DOI: 10.1038/s41578-021-00339-3 [DOI] [Google Scholar]
  • 143.Co-expression networks reveal the tissue-specific regulation of transcription and splicing; Saha Ashis, Kim Yungil, Gewirtz Ariel DH, Jo Brian, Gao Chuan, McDowell Ian C, Engelhardt Barbara E, Battle Alexis Genome Research (2017-October-11) https://doi.org/gb2qx6 DOI: 10.1101/gr.216721.116 [DOI] [Google Scholar]
  • 144.Addressing noise in co-expression network construction; Burns Joshua JR, Shealy Benjamin T, Greer Mitchell S, Hadish John A, McGowan Matthew T, Biggs Tyler, Smith Melissa C, Feltus FAlex, Ficklin Stephen P Briefings in Bioinformatics (2021-November-30) https://doi.org/gr83p2 DOI: 10.1093/bib/bbab495 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Applying differential network analysis to longitudinal gene expression in response to perturbations; Xue Shuyue, Rogers Lavida RK, Zheng Minzhang, He Jin, Piermarocchi Carlo, Mias George I Frontiers in Genetics (2022-October-17) https://doi.org/gr83p3 DOI: 10.3389/fgene.2022.1026487 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.CONE: COntext-specific Network Embedding via Contextualized Graph Attention; Liu Renming, Yuan Hao, Johnson Kayla A, Krishnan Arjun Cold Spring Harbor Laboratory (2023-October-24) https://doi.org/gt6n42 DOI: 10.1101/2023.10.21.563390 [DOI] [Google Scholar]
  • 147.Contextual AI models for single-cell protein biology; Li Michelle M, Huang Yepeng, Sumathipala Marissa, Liang Man Qing, Valdeolivas Alberto, Ananthakrishnan Ashwin N, Liao Katherine, Marbach Daniel, Zitnik Marinka Nature Methods (2024-July-22) https://doi.org/gt6n4w DOI: 10.1038/s41592-024-02341-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Predicting multicellular function through multi-layer tissue networks; Zitnik Marinka, Leskovec Jure Bioinformatics (2017-July-12) https://doi.org/gbnk7p DOI: 10.1093/bioinformatics/btx252 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Protein complex prediction with AlphaFold-Multimer; Evans Richard, O’Neill Michael, Pritzel Alexander, Antropova Natasha, Senior Andrew, Green Tim, Žídek Augustin, Bates Russ, Blackwell Sam, Yim Jason, … Hassabis Demis Cold Spring Harbor Laboratory (2021-October-04) https://doi.org/gm2vcp DOI: 10.1101/2021.10.04.463034 [DOI] [Google Scholar]
  • 150.Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk; Zhou Jian, Theesfeld Chandra L, Yao Kevin, Chen Kathleen M, Wong Aaron K, Troyanskaya Olga G Nature Genetics (2018-July-16) https://doi.org/gdvmqw DOI: 10.1038/s41588-018-0160-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Atlas of primary cell-type-specific sequence models of gene expression and variant effects; Sokolova Ksenia, Theesfeld Chandra L, Wong Aaron K, Zhang Zijun, Dolinski Kara, Troyanskaya Olga G Cell Reports Methods (2023-09) https://doi.org/gt6n4t DOI: 10.1016/j.crmeth.2023.100580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Towards complete and error-free genome assemblies of all vertebrate species; Rhie Arang, Shane A McCarthy Olivier Fedrigo, Damas Joana, Formenti Giulio, Koren Sergey, Uliano-Silva Marcela, Chow William, Fungtammasan Arkarachai, Kim Juwan, … Jarvis Erich D Nature (2021-April-28) https://doi.org/gjtrn9 DOI: 10.1038/s41586-021-03451-0 [DOI] [Google Scholar]
  • 153.Graph representation learning in biomedicine and healthcare; Li Michelle M, Huang Kexin, Zitnik Marinka Nature Biomedical Engineering (2022-October-31) https://doi.org/gq533s DOI: 10.1038/s41551-022-00942-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.The effects of biological knowledge graph topology on embedding-based link prediction; Bradshaw Michael S, Gaskell Alisa, Layer Ryan M Cold Spring Harbor Laboratory (2024-June-11) https://doi.org/gt6n45 DOI: 10.1101/2024.06.10.598277 [DOI] [Google Scholar]
  • 155.Phenomics Assistant: An Interface for LLM-based Biomedical Knowledge Graph Exploration; O’Neil Shawn T, Schaper Kevin, Elsarboukh Glass, Reese Justin T, Moxon Sierra AT, Harris Nomi L, Munoz-Torres Monica C, Robinson Peter N, Haendel Melissa A, Mungall Christopher J Cold Spring Harbor Laboratory (2024-February-02) https://doi.org/gt6n43 DOI: 10.1101/2024.01.31.578275 [DOI] [Google Scholar]
  • 156.Gene Ontology: Pitfalls, Biases, and Remedies; Gaudet Pascale, Dessimoz Christophe Methods in Molecular Biology (2016-November-04) https://doi.org/gf4d5v DOI: 10.1007/978-1-4939-3743-1_14 [DOI] [PubMed] [Google Scholar]
  • 157.Structured prompt interrogation and recursive extraction of semantics (SPIRES): A method for populating knowledge bases using zero-shot learning; Caufield JHarry, Hegde Harshad, Emonet Vincent, Harris Nomi L, Joachimiak Marcin P, Matentzoglu Nicolas, Kim HyeongSik, Moxon Sierra AT, Reese Justin T, Haendel Melissa A, … Mungall Christopher J arXiv; (2023) https://doi.org/gt6n48 DOI: 10.48550/arxiv.2304.02711 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 158.The zebrafish reference genome sequence and its relationship to the human genome; Howe Kerstin, Clark Matthew D, Torroja Carlos F, Torrance James, Berthelot Camille, Muffato Matthieu, Collins John E, Humphray Sean, McLaren Karen, Matthews Lucy, … Stemple Derek L Nature (2013-April-17) https://doi.org/k9c DOI: 10.1038/nature12111 [DOI] [Google Scholar]
  • 159.Non-model model organisms; Russell James J, Theriot Julie A, Sood Pranidhi, Marshall Wallace F, Landweber Laura F, Fritz-Laylin Lillian, Polka Jessica K, Oliferenko Snezhana, Gerbich Therese, Gladfelter Amy, … Brunet Anne BMC Biology (2017-June-29) https://doi.org/gfc837 DOI: 10.1186/s12915-017-0391-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160.A chromosome-scale assembly of the axolotl genome; Smith Jeramiah J, Timoshevskaya Nataliya, Timoshevskiy Vladimir A, Keinath Melissa C, Hardy Drew, Voss SRandal Genome Research (2019-January-24) https://doi.org/gf5xcq DOI: 10.1101/gr.241901.118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 161.An organismal perspective on C. intestinalis development, origins and diversification; Kourakis Matthew J, Smith William C eLife (2015-March-25) https://doi.org/gt6n49 DOI: 10.7554/elife.06024 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary File 2
Supplementary File 1
Supplementary Note 1

RESOURCES