Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Aug 30.
Published in final edited form as: Nat Methods. 2018 Nov 26;15(12):1067–1073. doi: 10.1038/s41592-018-0214-9

Found in Translation: A machine learning model for improving mouse to human inference

Rachelly Normand 1, Wenfei Du 2, Mayan Briller 1, Renaud Gaujoux 1, Elina Starosvetsky 1, Amit Ziv-Kenet 1, Gali Shalev-Malul 1, Robert J Tibshirani 2,3, Shai S Shen-Orr 1,*
PMCID: PMC12396630  NIHMSID: NIHMS1509549  PMID: 30478323

Abstract

Cross-species differences form barriers to translational research that ultimately hinder the success of clinical trials. Yet systematic incorporation of the wealth of knowledge on species differences in the interpretation of animal model data is lacking. Here we present the Found In Translation (FIT) model, a data-driven statistical methodology which leverages public domain gene expression data to predict from the results of a new mouse experiment genes expected to be altered in the equivalent human condition. We applied FIT to mouse model data from 28 different human diseases and identified experimental conditions in which FIT predictions outperform direct cross-species extrapolation from mouse results, increasing the overlap of differentially expressed genes by 20–50%. FIT predictions suggested novel disease-associated genes, an example of which we experimentally validated. FIT highlights signals that may otherwise be missed and reduces the time spent on false leads, with zero experimental cost. FIT is available at www.mouse2man.org.


Mice are the most widely used and cost-effective model to study human diseases. Despite their essentiality to clinical research and global expression similarity1,2, mice have substantial differences from humans3,4, likely due to evolutionary distance, differences in lifespan5, environment6 and the artificiality of many disease models. The utility of mice for human translational research is a much debated issue, dividing camps with fierce intensity7,8,9. From a clinical perspective, successful therapeutic experiments in mice often fail in human clinical trials10,11,12.

Therefore, there is an urgent need to develop methodologies for improving cross-species translational research. Experimental efforts to develop mice that more closely mimic human biology have become a norm for improved modeling of human disease13. Despite these processes being both long and expensive, there have been scarce attempts to develop in-silico methodologies for overcoming translational challenges14,15,16,17. Successful cross-species inference requires a rich set of observations to generalize from in a robust manner. Publicly available gene expression repositories have amassed over two million samples18,19 and have been used for example to identify and validate novel biomarkers and understand gene regulation20,21,22. Given these successes and the broad availability of mouse and human gene expression datasets, we reasoned these data can be used to bridge the ‘cross-species gap’.

Here we present FIT (Found in Translation), a novel data-driven statistical model that given a mouse gene expression experiment predicts the genes relevant to the parallel human condition. A previous attempt to improve mouse to human translation was restricted to predicting outcomes when a highly similar experiment profiling mice and humans was observed17. In contrast, FIT leverages the comprehensive collection of public mouse and human gene expression experiments to allow a more informative mapping from a new mouse experiment to human without having seen any prior human data for the condition.

Results

A mouse-human gene expression compendium

We manually annotated 170 microarray and RNA-seq gene expression datasets from the Gene Expression Omnibus (GEO)19 consisting of disease versus healthy control samples in mouse and human. We then paired human and mouse model datasets with a human disease dataset of comparable conditions to Cross Species Pairings (CSPs) (Supplementary Fig. 1, Methods). We differentiated between CSPs that were compared by the authors who generated at least one of the datasets (reference (RF) CSPs) and those we found separately in GEO and paired to each other (standard (ST) CSPs, Methods). For each CSP in each species, we computed the within-dataset effect-size capturing the difference between disease and control samples (Methods). This effort yielded a total compendium of 170 CSPs (62 RF and 108 ST) spanning 28 different diseases, which we use to study how human and mouse orthologous genes express under similar conditions and as training data for FIT (Supplementary Table 1).

A model predicting human disease genes from mouse data

A basic assumption when analyzing mouse gene expression data is that the genes with the highest absolute fold-change and low significance values are the most relevant to the condition being tested. We sought to develop a method that incorporates mouse-human biological differences when prioritizing genes, rather than solely base on the mouse data and ortholog gene information. We designed FIT, a linear model fit using lasso regression23, to predict human effect-size per gene by combining the measured mouse gene expression with prior knowledge on mouse-human similarities and differences The model outputs per gene a new effect-size value (i.e. alternative to fold-change), which can then be used for any downstream analyses that would commonly be run on mouse gene expression data.

Given a new mouse disease vs. control gene expression experiment, FIT predicts the expected human effect-size values per gene in three steps (Fig. 1a, Methods): (1) Effect-size computation: for each dataset in the training compendium, FIT computes an effect-size (an estimated fold-change for RNA-seq datasets and a z-test for microarray datasets). (2) Construction of a gene level model of mouse to human relation: Using both training and new mouse data FIT learns a regularized linear model (LASSO) whereby the new mouse data dictates the estimated effect-size, unless there is evidence to prefer a shift of the estimated effect-size based on the training data. To compute confidence intervals (CI), we bootstrap the training data 100 times. (3) Prediction of the human effect-size: FIT predicts the expected human effect-size based on the relationship learned from the training data by computing the mean of the estimated effect-sizes resulting from the bootstrapped models. The genes FIT predicts to have high absolute effect-sizes are more likely associated with the human condition of interest. We include a pre-step to FIT, in which an accurate classifier predicts FIT applicability to a new mouse input data. Applying FIT on mouse gene expression data may refocus researchers in directions other than those suggested from the original mouse results.

Figure 1: Cross-species similarity of disease versus control gene expression is markedly low across multiple disease models.

Figure 1:

(a) The FIT model pipeline: Given the new mouse gene expression experiment and a reference compendium consisting of paired human and mouse-model disease versus control datasets (Cross-Species Pairs, CSP) used for training the model. FIT performs three steps per each gene: (1) For each CSP in the compendium, FIT computes an effect-size between disease and control samples; (2) FIT then learns the relation between the mouse and human disease effect-sizes using a set of regularized linear models obtained from repeated sampling; (3) FIT estimates the human effect-size per gene as the mean of the regularized predictions. Confidence intervals are obtained from the sampling procedure. We evaluate the model by comparing the predictions to human gene expression data from a parallel condition to that of the input mouse. Due to the complexity of translating mouse to human biology, as a pre-step for a FIT analysis we use an SVM classifier that predicts with high accuracy whether FIT is applicable to the new mouse data. (b) The mouse-human disease reference compendium is comprised of CSPs organized by diseases. For each CSP, we quantify the cross-species gap by comparing mouse differentially expressed genes (DEGs) and human DEGs at set thresholds of fold-change and q-value (117 combinations tested in total). We consider the cross-species overlapping DEGs as true positive (TP). (c) TP fractions (TP/total amount of mouse DEGs) are illustrated (mean+/-SE) for each CSP in the compendium for fold-change threshold = 0.1 (genes with the top 10% absolute fold-change values) and a q-value threshold = 0.05, and compared to intra-human and intra-mouse pairings across 28 diseases. Cross-species pairings have significantly lower TP fractions compared to intra-species pairings (P-value<=10−308 for intra-mouse and P-value<=0.0029 for intra-human by Fisher’s combined probabilities test), emphasizing the cross-species gap. Disease names are colored by comparison type, number of pairs per disease per type appear on right.

FIT predictions better resemble human data

To set a quantitative reference point of the cross-species gap and allow evaluation of FIT’s performance, we performed a systematic assessment of human-mouse similarity across all CSPs in the compendium. Limiting our analysis to one-to-one orthologues genes, we identified differentially expressed genes (DEGs) in each species by setting varying thresholds for fold-change and q-values (Methods). We defined the overlap between mouse and human DEGs as true positive (TP) (Fig. 1b) and the ratio of TP genes out of all mouse DEGs as the TP fraction. For q-value=0.05 and fold-change=0.1 (10% of highest fold-change genes in absolute values), the maximal TP fraction in mouse was 34% – implying that at best, when directly translating the mouse results to human, only one out of three genes was shared, with a mean of one out of 20 genes (Fig. 1c). As control, we computed the TP fraction between intra-species pairings, (all possible pairwise comparisons within each disease of the same species). We observed significantly higher TP fractions within intra-species pairings compared to CSPs (Fisher’s combined P-value<=10−308 for intra-mouse and P-value<=0.003 for intra-human on one-tailed Mann-Whitney). Extending this analysis to all 117 thresholds, we observed that cross-species TP fractions were on average a third of the size of intra-species, (Supplementary Fig. 23, P-value<=10−16 by one-tailed Mann-Whitney on mean values, n=117 as the number of threshold combinations). Taken together, we concluded that identification of genes associated with human disease by direct inference from mouse gene expression fails to capture the majority of human signal.

To evaluate FIT, we used a leave-one-disease-out methodology: for each disease dataset, we removed all CSPs associated with the disease from the training compendium. We treated the mouse disease dataset as a test experiment and used it, jointly with the training compendium, as input for FIT to predict the human effect-size per gene. We then tested these estimates against the left-out disease human dataset (Fig. 1a, evaluation step). Building intuition by example, Sdf2 encodes an ER-stress associated protein not previously associated with disease24. Presenting FIT with the training compendium consisting of all prior measurements for Sdf2 except for the measurements from sepsis CSPs (Fig. 2a, top). FIT estimated negative slopes in the majority of the bootstrapped regression models, implying an inverse cross-species relation. This resulted in FIT predicting opposite effect-sizes compared to the original mouse values (Fig. 2a, bottom). To evaluate FIT globally, we defined FIT DEGs per dataset for a specific threshold combination, as a set of genes equal in size to the mouse DEGs at the same thresholds whose FIT estimated effect-sizes were ranked highest in absolute value. Comparing the performance of mouse-direct analysis against the FIT-derived analysis for one of the sepsis CSPs, we contrasted mouse TP genes with FIT TP genes and observed a total of 184 human DEGs ‘rescued’ by FIT (Fig. 2b, for q-value=1, fold-change= 0.1).

Figure 2: FIT is able to increase the true positive fractions compared to the mouse model.

Figure 2:

(a) An example of FIT prediction for the Sdf2 gene in sepsis. Top: Sdf2 effect-sizes in mouse vs. human across all CSPs in the training data (black) from which slopes of the 100 bootstrapped regularized linear regression models are computed (grey lines). Applying the leave-one-disease-out methodology, sepsis CSPs were held out from the model (green). Bottom: for every CSP in the sepsis dataset (colored lines), shown are the mouse (top), FIT (middle) and human (bottom) effect-size percentiles, computed separately for positive (0 – 100) and negative (-100 – 0) effect-size values. FIT predicts an effect-size with opposite signs compared to the mouse, due to the negative slopes estimated from the training data. (b) Human vs. mouse fold-change values of an example sepsis CSP. FIT rescued 184 human DEGs undetected by the mouse (DEGs defined at fold-change=0.25 and q-value =1). (c) Diagrams showing the evaluation procedure for FIT. Left: Venn diagrams for mouse and FIT compared at a given fold-change and q-value threshold to the silver-standard human DEGs each yield a full confusion matrix: True-positive (TP), true-negative (TN), false-positive (FP) and false-negative (FN). Right: The resulting direct-to-mouse and FIT derived confusion matrices can be compared with one another across all thresholds either by computing the TP/ TN/ FP/ FN ratios (top) or by computing sensitivity and specificity (bottom). (d) FIT performance varies between CSPs and can be split into five classes ranging from major signal gain to major signal loss based on the confusion matrix values at q-value=1 and across fold-changes thresholds. For each class, the sensitivity and specificity values are illustrated in standard boxplots (median, 25% and 75% percentiles, number of CSPs per class are 28, 51, 27, 63, 1 for major signal gain to major signal loss). (e) An SVM classifier, based on the first 50 PCs of the mouse gene expression data, predicts which datasets are improved by FIT. The results are shown for each pair of thresholds (x-axis), averaged over 100 classifier runs per threshold pair, each with random splits of the data to training and test (80%-20%). Inset shows zoomed-in results for q-value=0.1 and fold-change=0.15 with an overall accuracy of 84%.

Next, we performed the leave-one-disease-out validation procedure across all 28 diseases and 170 CSPs, comparing TP fractions between FIT and the mouse data (Fig. 2c left, Methods). We extended the evaluation of FIT to a complete confusion matrix by computing the false positive (FP), false negative (FN) and true negative (TN) values for both FIT and the original mouse analysis. Then, we compared between the mouse performance and FIT’s performance by calculating the sensitivity and specificity values for FIT and the mouse and by computing ratios of TP, TN, FP and FN values between FIT and mouse (Fig. 2c right, Methods). We repeated this procedure across 117 fold-change and q-value threshold combinations. An improvement in translational insight will occur when sensitivity and/ or specificity increases in FIT compared to the original mouse analysis, and ideally when TP ratio>1, TN ratio>1, FP ratio<1 and FN<1.

Classifying the 170 CSPs by the relative performance FIT compared to the mouse, we observed five classes of performance which we categorized based on the behavior of true and false ratios across multiple fold-change thresholds (Fig. 2d q-value=1, Supplementary Fig. 4a, Methods): These included major and minor signal gain, that is, cases where FIT clearly captured more human relevant signal than the original mouse with or without reduced false leads; to major and minor signal loss: where a FIT-derived analysis captured less human relevant signal and possibly introduced more false leads than the direct-from-mouse analysis. Between these two extremes, we observed an ‘equal signal’ class exhibiting intermediate performance.

To enable a robust use of FIT we built a classifier aimed at predicting when FIT will improve similarity to human data based on dimensionality reduced mouse training data (Methods). Across all threshold pairs, the classifier correctly predicted in 80% of tested cases on average whether FIT will provide benefit (i.e. an increase of TP rate compared to conventional analysis of the mouse data, Figure 3e, Supplementary Fig. 4b). Taken together, our evaluation suggests FIT is able to markedly increase the overlap with human DEGs and decrease the amount of false results in a predictable subset of mouse experiments.

Figure 3: FIT rescues human disease relevant genes that could not have been found based on mouse data alone.

Figure 3:

(a) True positive fractions (mean+/-SE across all 170 CSPs) for mouse DEGs (blue), FIT DEGs (pink), a subset of FIT DEGs with the smallest 30% confidence intervals (CI, red) and intra-species DEGs (black). There is a significantly higher TP-fraction in FIT DEGs compared to mouse DEGs, and in small-CI genes compared to all predicted genes (P-value<=10−16 by one-tailed Mann-Whitney across all threshold means for both comparisons, n=117 as the number of threshold combinations). Intra-species TP fraction (black) may set an theoretical upper limit on translation improvement. (b) A histogram of percentages of CSPs in which each gene is considered a DEG (q-value threshold=0.1, fold-change threshold=0.25) highlights that FIT does not repeatedly boost a specific group of genes across different CSPs. (c) GO-term enrichment analysis on the putative genes of each FIT-improved disease shows they have diverse biological functions. Circle size represents the percent of putative genes that are tagged with the GO annotation. P-values (by hyper-geometric) are represented by color, p-values>0.05 are colored in grey. (d) The association of ILF3 with IBD was predicted by FIT despite not being observed in gene expression data of mouse models of IBD, nor in human IBD disease samples. DAB staining area quantification of the protein encoded by ILF3 shows significant differences between the expression in colons of healthy (top) and IBD patients (bottom, P-value<=0.005 by one-tailed Mann-Whitney, average of total area (mean±SD): 3.3±1.05% and 9.1±3.7% for 5 healthy and 7 IBD patients, respectively). Values are illustrated in standard boxplots (median, 25% and 75% percentiles). Scale bars: 100µm. (e) Enrichment of three pathways in FIT but not in the original mouse IBD data demonstrate that FIT boosts the signal of disease-relevant pathways. Shown are GSEA results of an IBD dataset that were enriched in FIT but not in the mouse (nominal GSEA p-value<0.05 and GSEA FDR<0.25 in FIT but not in the mouse). In these pathways, known to be associated with IBD, FIT’s leading edge has a higher enrichment score and in the two top examples the leading edge is reached at a smaller rank.

FIT predicts human disease genes masked when using mouse data alone

To understand when FIT performs best, we characterized the features that correlated with FIT’s performance. We noted that with the present training data, FIT increased true positive fractions predominantly in infectious diseases and inflammatory conditions yet performed poorly in cancer. Tissue wise, FIT performed well when the mouse tissue was from spleen, blood, lung and gut whereas in human tissues were blood and gut (Supplementary Fig. 5). We focused on the 10 diseases for which the majority of CSPs showed signal-improvement by FIT (from here on FIT-improved diseases, covering 81 CSPs) (Supplementary Fig. 4a, including: Sepsis, Duchene Muscular Dystrophy (DMD), Staphylococcus Infection, Injury, Burns, Inflammatory Bowel Diseases (IBD), Gaucher, Cystic Fibrosis, E.coli, Infection and RA). In these diseases, across all 117 thresholds, restricting the analysis to only the genes with the 30% smallest CIs (small-CI genes), we observed significantly higher TP fractions compared to when considering all FIT DEGs (Fig. 3a, P-value<=10−16 by one-tailed Mann-Whitney on mean values), a difference not seen in randomly chosen gene set (data not shown). Intra-species TP fraction may set the upper limit of translation improvement; we noted that on average FIT increased the TP fraction by 20% whereas focusing on small-CI genes increased the TP fraction by 50%, and in some thresholds was close to the intra-species asymptote, suggesting that CIs can be used as reliability scores for FIT predictions (Fig. 3a).

Next, we characterized the genes whose estimates FIT alters: We defined ‘putative genes’ as genes FIT suggested to have high relevance to human yet were undetected by the mouse. On average, 28% of the putative genes were ‘rescued’ in that we observed they were highly expressed in the parallel human condition (Methods, Supplementary Fig. 6a, Supplementary Table 2). The percent of rescued genes out of the putative gene set increased when focusing on small-CI genes (Supplementary Fig. 6b). To identify whether FIT achieves improved performance by boosting a specific set of genes across all mouse models or via re-ranking of a diverse set of genes, we computed the number of CSPs FIT boosted each gene in. Only 16% of genes were identified as a DEG by FIT in over 10% of CSPs, and less than 7% were identified in over 20% of CSPs (Fig. 3b, Supplementary Table 3, counting genes measured in at least 40 CSPs for q-value=0.1, fold-change=0.25). This trend was evident across thresholds (data not shown), illustrating that FIT’s predictions were predominantly tailored to the input mouse data. Furthermore, looking at functional enrichment in putative genes showed that FIT’s predictions were not limited to a specific functional group (Fig. 3c, Methods).

To experimentally showcase the strength of FIT predictions, we chose as an example one FIT putative gene, ILF3, for which FIT predicted upregulation in the colon of IBD patients compared to healthy controls despite it being undetected in either IBD mouse model or human datasets. To the best of our knowledge ILF3 has not been associated with IBD, yet genetic variants of it have been associated with Rheumatoid arthritis25, Psoriasis26 and several cancers27. Validating by an orthogonal methodology to gene expression, we observed a significant increase in the immuno-staining of ILF3 protein in colons of IBD patients versus healthy controls (Fig. 3d, p<0.005 by one-tail Mann-Whitney, Methods). Finally, to showcase FIT utility for downstream analyses, we identified pathways significantly enriched in FIT only in the IBD dataset (nominal P-value<0.05 and FDR <0.25 using FIT but not direct-from-mouse data). Amongst the top were pathways associated with IBD, including: Complement and coagulation cascades28, Antigen processing and presentation29, Neurotrophin signaling pathway30. Thus, FIT is able to boost the signal of disease-relevant pathways that are weaker in the mouse data (Fig. 3e). Taken together, these results suggest that by leveraging prior data a FIT analysis may rescue genes and pathways masked by cross-species differences and predict novel human-disease associations.

Discussion

The cross-species gap is a major hurdle for translational research. Here we presented FIT (Found In Translation), a statistical model that enables a more informative translation from a mouse gene expression dataset to human by leveraging public human-mouse gene expression data. FIT predicted effect-sizes allow any sequential analysis that is currently performed on mouse data. We show that FIT outperforms the direct-from-mouse approach for a large number of tested datasets. To allow researchers to assess FIT’s applicability for their own experiment we provide a classifier that predicts for which mouse datasets FIT is able to predict values resembling human data. FIT is highly beneficial for the datasets it is predicted to improve, producing 20–50% more human-relevant genes compared to the mouse.

FIT’s performance is likely a complex function of the size of the training compendium, the complexity and variability of the disease studied and whether the mouse model emulates it well, the mouse and human tissues of interest and experiment similarity to those in the training compendium. As a data-driven methodology, FIT performance assessment relies on the human data to which FIT predictions are compared to. We consider the human data a “silver-standard” because on the one hand, it is the most intuitive and logical data to compare FIT’s predictions to, but on the other hand it reflects only one angle of human disease biology and has its own internal inconsistences. For example, as we show and is well known, different human gene expression datasets in the same disease may vary widely in the genes they report as differentially expressed. Considering the human data as a silver standard also implies that incorporation of a large training set provides strength to the model much greater than the findings from any single experiment – similar to how meta-analyses usually capture a more accurate signal than that of any single study they contain.

FIT can be augmented in a number of ways: First, the training compendium can be expanded with increased data growth and with it FIT’s performance and the breadth of conditions it is applicable for will rise. Second, with more data, the model can incorporate additional information such as explicit modeling of gene relations (e.g. pathways), disease relations, or a more sophisticated weighting scheme for the training data. FIT presents a novel paradigm of data reuse which can be extended to other data types and species, as well as to other biological questions involving mappings.

As translational research is difficult and drug development a costly and long process, reducing false leads from animal models is critical. The FIT application is fast and does not entail high costs or efforts. We believe that FIT should be used in conjunction with standard analyses to flush out more valid conclusions regarding human biology. Born out of the realization that with twenty years of whole genome expression profiling data behind us and still growing, no new experiment should be analyzed without taking prior data into account; we see FIT as a method leveraging this data to account for the evolutionary and environmental differences in gene function and regulation between mouse and man. Future data methodology enhancements and continued data accumulation will ultimately lead to a Rosetta Stone for cross-species translation.

Online Methods

Assembly and pairing of matched human and mouse datasets from the public domain

We manually assembled and annotated human and mouse microarray and RNA-seq gene expression experiments from the Gene Expression Omnibus (GEO)19. For an experiment to be included in our compendia, we required the following criteria: (1) For every human dataset/s there must be a corresponding mouse model dataset/s; (2) Each condition in each species must contain at least 3 disease samples and at least 3 control samples; (3) the microarray platform used must be a 1-color array.

In each GEO dataset, we identified sub-datasets (referred in the paper as ‘datasets’ for simplicity) consisting of a subset of disease and control samples measured under similar conditions.

These datasets are either disease and control samples originally compared to each other by the researchers, or the most comparable conditions, if such a comparison was not made in the original study (Supplementary Fig. 1a, Supplementary Table 1). We identified two types of CSPs: Standard (ST), which consisted of human datasets and corresponding datasets of a mouse model of the human disease in a separate study; and Reference (RF), where the human and mouse datasets were directly contrasted to one another in a publication authored by the researchers who had generated at least one of the datasets. We made this distinction as we assumed that for RF-CSPs researchers who are experts in the disease found the data from the mouse experiment as adequate for gaining insight of the human disease. In both RF and ST-CSPs, if one condition contains samples in more than one microarray platform, then the set of samples from each platform were considered as a separate dataset.

Following collection of all the relevant mouse and human datasets, we paired human and mouse datasets: In RF-CSPs, the cross-species comparisons in each dataset were chosen as closely as possible to the comparisons done in the original study. In ST-CSPs, or in RF-CSPs in which the original study did not contain pair-wise comparisons of the species (e.g. when all samples were simply clustered), the most similar conditions in both species were chosen for comparison. Specifically, we favored CSPs that were assayed in parallel tissue in both species or if that was not available, the CSP whose similarity of global fold-change profile as computed by Spearman correlation, was highest. We allowed a one-to-many pairing scheme, meaning one mouse dataset can be paired with a few human datasets and vice versa. Many-to-many pairings were not permitted in order to avoid dependencies in the reference compendium. RNA-seq and microarray datasets were not mixed, that is a CSP included either only microarray datasets or only RNA-seq datasets from both species.

We then filtered the reference compendium and omitted datasets with a low amount of differentially expressed genes (DEGs) in either the mouse or the human dataset, to avoid using datasets with no signal. Specifically, for a q-value threshold of 0.3 and a fold-change cutoff of 15% of the top genes in absolute value, datasets that had less than 1% DEGs out of the total number of genes were omitted from the reference compendium. In addition, datasets with less than 3,000 genes in total were not included.

This manual assembly effort resulted in an annotated gene expression compendium containing 3,033 human and 1,181 mouse samples that were sub-divided into a total of 90 human datasets and 137 mouse datasets (Supplementary Table 1). The compendium includes datasets assayed across 20 and 13 different human and mouse microarray platforms respectively, as well as high throughput sequencing.

Pre-processing of gene expression datasets

Gene expression data was obtained from GEO using the R package GEOquery31. All the analyses performed on the gene expression datasets were limited to 16,491 genes that are 1-to-1 human-mouse orthologs as identified in MGD32.

Pre-processing of microarray datasets

We systematically verified that each dataset is logged transformed, by checking that all the values in the dataset are lower than 100. In case the dataset was found to be not log-transformed, a log-2 transformation was applied. Additionally, we systematically verified that each dataset is normalized by checking that the difference between the sample means < 10−5. In case the difference was greater, the samples in the dataset were quantile-normalized.

Gene-level fold-change (FC) was calculated first at the probe level:

FCprobe=meansample  Disease(log2(Intensitysample)) meansample  Control(log2(Intensitysample))

The gene fold-change was then taken as the highest probe-level value:

FCgene=max(FCprobe|probegene)

In order to get comparable dynamic range for all datasets we computed Z-test as previously suggested33:

Ztest (g)=mean(ZsampDisease(g))mean(ZsampControl(g))SDControlNControl+SDDiseaseNDisease

Where Z(g) is the Z-score of gene g for either the control or disease samples; SDDisease and SDControl are the standard deviation of the disease and control samples respectively; Ncontrol and NDisease are the number of samples in the control and the disease conditions, respectively.

P-values were computed from the Z-test distribution, which is normally distributed, and were adjusted for multiple hypotheses by Benjamini-Hochberg adjustment, to produce q-values. The Z-test values, which were referred to as effect-size throughout the paper, were given to FIT as input, whereas the fold-change and q-values were used to define DEGs.

Pre-processing of RNA-seq datasets

Raw reads were downloaded using R package SRAdb34 V1.28.1 and then pseudo-mapped and quantified using Kallisto35 version 0.43.0 with the following parameters “kallisto quant –b 100”. Sleuth36 version 0.28.1 was used in order to compute the effect-size (beta) and q-values per gene.

The effect-size values were given to FIT as input, and were used, along with the q-values, to define DEGs.

Definition of DEGs per species and cross-species comparison

DEGs are defined by setting fold-change and q-value thresholds. A range of q-values were checked (0.01, 0.05, 0.1, 0.15, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1), and a range of fold-change values were checked (0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4), consisting of a total of 117 threshold combinations. The fold-change thresholds represent the fraction of genes that have the highest absolute fold-change.

When comparing DEGs between the species the fold-change sign was taken into account: for a gene to be a true positive result, it should be below the q-value threshold in both species, above the absolute fold-change threshold in both species and have the same fold-change sign.

The FIT model

We assume a linear model for each gene g:

ZpH(g)=αg+βgZpM(g)

Where ZpX(g) correspond to the effect-sizes of gene g in CSP p and species X (H= human, M= Mouse).

We introduce regularization with penalties that shrink the slope to 1 and the intercept to 0. In other words, in order to change the mouse expression value, there should be a strong cross-species signal that suggests otherwise. The regularization is done by adding the penalty terms |αg| and |1-βg| to the least squares optimization. Thus, for each gene g, we solve

minαgβg d(ZpH(g)αgβgZpM(g))2+λ(|αg|+|1βg|)

Substituting βg'=1βg the objective becomes:

minαgβg d(ZpH(g)ZpMαgβgZpM(g))2+λ(|αg|+|β'g|)

We use 10-fold cross validation at the comparison level to choose the λ parameter. We chose λ by taking the smallest value giving a mean square error that is one standard error measure away from the minimum. Given the optimal λ, we find the optimal solution α^g, β^g.

The prediction for gene g for a new CSP p' is then given by:

Zp'H(g)=α^g+β^gZp'M(g)

Zp'H(g) is referred to as FIT estimated effect-size.

Classification of FIT results into five classes

CSPs were manually clustered into one of five classes based on the true and false ratios measured in tests with q-value=1 and across fold-change values (0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4). This choice of thresholds allowed a sufficient number of DEGs in mouse, FIT and human in order to avoid noise. Each CSP was represented by four vectors (TP ratio, TN ratio, FP ratio and FN ratio) containing 7 values each, that were manually clustered using the following criteria based on the number of values in each vector:

# TP ratio>1 # TN ratio>1 #FP ratio<1 #FN ratio<1
Class 1 (Major signal gain) >=5 >=5 >=5 >=5
Class 2 (Minor signal gain) >=5 >=5
# TP ratio>1 # TN ratio>1 #FP ratio<1 #FN ratio<1
Class 4 (Minor signal loss) >=5 >=5
Class 5 (Major signal loss) >=5 >=5 >=5 >=5

Any CSP that did not fit any of the above criteria was classified as class 3 (Equal performance).

SVM classifier

We sought to train an SVM classifier to predict FIT’s ability to improve mouse-human translation. Specifically, we performed PCA on the mouse gene expression data and observed a notable separation of good versus poor FIT performance in the PC space, which could be seen even in the first principle component, suggesting that similarity of expression of the new mouse results to the mouse data in the training compendium may be predictive of FIT performance (Supplementary Fig. 4b). The PCA was performed on all 137 unique mouse datasets (mouse datasets may be paired more than once to different human datasets).

In order to run PCA on the training data we looked for a common subset of genes available in most of the mouse datasets. We found a subset of 4,957 genes that were missing in less than 15% of the CSPs. The missing expression data for those genes (9% out of the whole gene expression matrix) was imputed by averaging the expression of the gene in the remaining datasets. PCA was run on the imputed gene expression data of the 4,957 genes and the first 50 PCs were given as input to an SVM classifier, capturing more than 80% of the variation. This classifier is available in the FIT website and R package.

In order to test the classifier, we randomly split the data to training (20%) and test (80%). The training data was used to create the 50 PCs and train the SVM classifier, which was run on the test data in order to predict whether FIT is able of improve it. This splitting of the data to training and test was performed 100 times per q-value and threshold pair (summarized in Figure 2e). In each iteration, we tuned the classifier’s parameters using the “tune” function from the “e1071” R package 37 V1.6–8.

FIT putative and rescued genes

We considered predictions to be putative genes in a disease if they met the following criteria for at least one CSPs of the disease:

  1. FIT estimated effect-size >= 80th percentile

  2. Mouse effect-size <= 60th percentile.

A putative gene that had a human effect-size >= 60th percentile was considered as a FIT rescued gene.

The criteria for putative and rescued genes is checked for positive and negative values separately.

Functional analysis

The functional annotations (Fig. 3c) were downloaded from Panther38 on August 2017. We used a hyper-geometric test to assess the significance of functional enrichment of FIT putative genes in each disease.

Experimental validation of ILF3

The estimated effect-sizes of ILF3 by FIT were above the 60th percentile for 2 out of 3 IBD mouse datasets it was measured in. In the mouse and human datasets the effect-sizes in all datasets were below the 50th percentile.

We performed analyses by immunostaining of archival slides of 7 IBD patients (5 females, 2 males, age (mean±sd) 32±10) and 5 healthy individuals (3 females, 2 males, age (mean±sd) 50±21), treated at the gastroenterology department of the Rambam Health Care Campus (RHCC). This study was performed in accordance with the Declaration of Helsinki and was approved by the institutional ethics committee of RHCC and TASMC (August 2016, Israel).

Formalin-fixed slides of paraffin-embedded colon tissues (5 healthy patients, 7 IBD patients), sectioned at 4µm, slices were deparaffinized in xylene (three times for 5 minutes each), rehydrated through a graded series of treatment with ethanol (100%, 95%, 85%, and 70% for 5 minutes each), and rinsed in tap water. To quench endogenous peroxidase activity, the sections were incubated in 0.3% H2O2 in Methanol for 15 minutes. For all samples, antigen retrieval was performed by boiling the sections in 0.01M Sodium Citrate Buffer (Sigma) for 20 minutes in the microwave, followed by incubation at room temperature for 30 minutes. A 10% solution of normal goat serum (Sigma) in PBS was used as a blocking buffer (2 hours at room temperature). Sections were incubated with anti-ILF3 primary antibody (1:400, LS-BIO, LS-B11424) for an hour at room temperature. Sections were washed in PBS-Tween (0.1%) and secondary staining with DAB were performed using a Polink universal kit (GBI-Labs Polink-1 HRP Detection System for Broad Spectrum) and counterstained with Mayer’s hematoxylin (Sigma). Stained sections were mounted in Immu-Mount (Thermo Fisher).

Slides were scanned on an automatic slide scanner (3DHistech Pannoramic 250 Flash III) at x20 magnification, and images were taken using Adimec Q12A180 camera. Quantification of DAB staining was performed using ImagePro Premier 9.3.2 software. Minor gamma corrections were made to three of the samples. Each slide contained two slices from each patient. In order to calculate the amount of staining in each slice, the area of stained cells was divided by the total area of the slice and averaged to every patient.

GSEA analysis

To find enriched pathways we ran pre-ranked GSEA39 on the mouse effect size values and on the FIT predicted values of the IBD comparison “Microarrays_ST_IBD_13” (as detailed in Supp. Table. 1). We used GSEA software V3.0 build 0160 with the default parameters and used KEGG gene set list (c2.cp.kegg.v6.1.symbols.gmt).

Statistics

Significance in the difference between cross- to intra-species TP fractions across diseases was computed by Fisher combined P-values on Mann-Whitney tests per disease.

Significance in the difference between TP-fraction in FIT DEGs and mouse DEGs and difference between small-CI genes and all predicted genes was computed by one-tailed Mann-Whitney test across all threshold means.

Significance in the difference of % of total areas in immune-staining between healthy controls and IBD patients was computed by a one-tailed Mann-Whitney test.

Code availability

FIT is available as a web-tool at www.mouse2man.org and as an R package at https://github.com/shenorrLab/FIT.mouse2man. The code of the package at the time of publication is available under Supplementary Software linked to this article.

Reporting Summary

Further information on experimental design is available in the Nature Research Reporting Summary linked to this article.

Supplementary Material

1
2
1509549_Sup_tab_1
1509549_Sup_tab_3
1509549_Sup_tab_2

Acknowledgments

We wish to thank Mark Davis, Atul Butte and Larry Steinman, discussions with which inspired this work. Yuval Shaked, Tomer Shlomi, Zohar Yakhini, Yoram Reiter and Elhanan Borenstein for fruitful discussions. Yehuda Chowers for his assistance with IBD-related experiment and analysis. David Cohen for computing help. This research was supported by the Israel Science Foundation (ISF) grants 1365/12 and the Rappaport Institute of Biomedical Research (both to S.S.S.-O.). R.J.T was supported by NSF grant DMS1208164 and NIH grant 5R01 EB001988-16.

Footnotes

Competing Financial Interests Statement

S.S.S.-O. and E.S. are scientific advisers and hold equity in CytoReason. R.G. is an employee of CytoReason and holds equity. All other authors declare no competing interests.

Data availability

All gene expression data was acquired from the public domain and is listed in Supplementary Table 1. The training data used by FIT can be downloaded from the site and package mentioned under “Code availability” section.

References

  • 1.Stuart JM, Segal E, Koller D & Kim SK A gene-coexpression network for global discovery of conserved genetic modules. Science 302, 249–55 (2003). [DOI] [PubMed] [Google Scholar]
  • 2.Zheng-Bradley X, Rung J, Parkinson H & Brazma A Large scale comparison of global gene expression patterns in human and mouse. Genome Biol 11, R124 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Liao B-Y & Zhang J Null mutations in human and mouse orthologs frequently result in different phenotypes. Proc. Natl. Acad. Sci. U. S. A 105, 6987–92 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Mestas J & Hughes CCW Of mice and not men: differences between mouse and human immunology. J. Immunol 172, 2731–8 (2004). [DOI] [PubMed] [Google Scholar]
  • 5.Geifman N & Rubin E The mouse age phenome knowledgebase and disease-specific inter-species age mapping. PLoS One 8, e81114 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Beura LK et al. Normalizing the environment recapitulates adult human immune traits in laboratory mice. Nature 532, 512–516 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Shay T et al. Conservation and divergence in the transcriptional programs of the human and mouse immune systems. Proc. Natl. Acad. Sci. U. S. A 110, 2946–51 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Seok J et al. Genomic responses in mouse models poorly mimic human inflammatory diseases. Proc. Natl. Acad. Sci. U. S. A 110, 3507–12 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.KOLATA G Mice Fall Short as Test Subjects for Some of Humans’ Deadly Ills. New York Times (2013). [Google Scholar]
  • 10.Bugelski PJ & Martin PL Concordance of preclinical and clinical pharmacology and toxicology of therapeutic monoclonal antibodies and fusion proteins: cell surface targets. Br. J. Pharmacol 166, 823–46 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Wilkins HM, Bouchard RJ, Lorenzon NM & Linseman DA in Horizons in Neuroscience Research Vol.5 (eds. Costa A & Villalba E) 5, 67–72 (Nova Science Publishers, Inc., 2011). [Google Scholar]
  • 12.Hünig T The storm has cleared: lessons from the CD28 superagonist TGN1412 trial. Nat. Rev. Immunol 12, 317–8 (2012). [DOI] [PubMed] [Google Scholar]
  • 13.Brehm M. a, Wiles MV, Greiner DL & Shultz LD Generation of improved humanized mouse models for human infectious diseases. J. Immunol. Methods 1–15 (2014). doi: 10.1016/j.jim.2014.02.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Hwang S, Kim E, Yang S, Marcotte EM & Lee I MORPHIN: a web tool for human disease research by projecting model organism biology onto a human integrated gene network. Nucleic Acids Res 42, W147–53 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zinman GE et al. ModuleBlast: identifying activated sub-networks within and across species. Nucleic Acids Res 43, e20–e20 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Djordjevic D, Kusumi K & Ho JWK XGSA: A statistical method for cross-species gene set analysis. Bioinformatics 32, i620–i628 (2016). [DOI] [PubMed] [Google Scholar]
  • 17.Seok J Evidence-Based Translation for the Genomic Responses of Murine Models for the Study of Human Immunity. PLoS One 10, e0118017 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kolesnikov N et al. ArrayExpress update--simplifying data submissions. Nucleic Acids Res 43, D1113–D1116 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Barrett T et al. NCBI GEO: Archive for functional genomics data sets - Update. Nucleic Acids Res 41, 991–995 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sweeney TE, Braviak L, Tato CM & Khatri P Genome-wide expression for diagnosis of pulmonary tuberculosis: a multicohort analysis. Lancet Respir. Med 4, 213–224 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Segal E et al. Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data. Nat. Genet 34, 166–76 (2003). [DOI] [PubMed] [Google Scholar]
  • 22.Szász AM et al. Cross-validation of survival associated biomarkers in gastric cancer using transcriptomic data of 1,065 patients. Oncotarget 5, (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Tibshirani R Regression Shrinkage and Selection via the Lasso. J. Stat. Soc 58, 267–288 (1996). [Google Scholar]
  • 24.Lorenzon-Ojea AR et al. Stromal cell derived factor-2 (Sdf2): A novel protein expressed in mouse. Int. J. Biochem. Cell Biol 53, 262–270 (2014). [DOI] [PubMed] [Google Scholar]
  • 25.Izumi T et al. Activation of synoviolin promoter in rheumatoid synovial cells by a novel transcription complex of interleukin enhancer binding factor 3 and GA binding protein alpha. Arthritis Rheum 60, 63–72 (2009). [DOI] [PubMed] [Google Scholar]
  • 26.O’Rielly DD & Rahman P Genetic, Epigenetic and Pharmacogenetic Aspects of Psoriasis and Psoriatic Arthritis. Rheum. Dis. Clin. North Am 41, 623–642 (2015). [DOI] [PubMed] [Google Scholar]
  • 27.Hou Q, Chen K & Shan Z The construction of cDNA library and the screening of related antigen of ascitic tumor cells of ovarian cancer. Eur. J. Gynaecol. Oncol 36, 590–4 (2015). [PubMed] [Google Scholar]
  • 28.Senchenkova E, Seifert H & Granger D Hypercoagulability and Platelet Abnormalities in Inflammatory Bowel Disease. Semin. Thromb. Hemost 41, 582–589 (2015). [DOI] [PubMed] [Google Scholar]
  • 29.Stagg AJ, Hart AL, Knight SC & Kamm MA The dendritic cell: its role in intestinal inflammation and relationship with gut bacteria. Gut 52, 1522–9 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.di Mola FF et al. Nerve growth factor and Trk high affinity receptor (TrkA) gene expression in inflammatory bowel disease. Gut 46, 670–9 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]

Methods-only references

  • 31.Davis S & Meltzer PS GEOquery: a bridge between the Gene Expression Omnibus (GEO) and BioConductor. Bioinformatics 23, 1846–7 (2007). [DOI] [PubMed] [Google Scholar]
  • 32.Eppig JT, Blake J. a., Bult CJ, Kadin J. a. & Richardson JE The Mouse Genome Database (MGD): facilitating mouse as a model for human biology and disease. Nucleic Acids Res 43, D726–D736 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Cheadle C, Vawter MP, Freed WJ & Becker KG Analysis of microarray data using Z score transformation. J. Mol. Diagn 5, 73–81 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhu Y, Stephens RM, Meltzer PS & Davis SR SRAdb: query and use public next-generation sequencing data from within R. BMC Bioinformatics 14, 19 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Bray NL, Pimentel H, Melsted P & Pachter L Near-optimal probabilistic RNA-seq quantification. Nat. Biotechnol 34, 525–527 (2016). [DOI] [PubMed] [Google Scholar]
  • 36.Pimentel H, Bray NL, Puente S, Melsted P & Pachter L Differential analysis of RNA-seq incorporating quantification uncertainty. Nat. Methods 14, 687–690 (2017). [DOI] [PubMed] [Google Scholar]
  • 37.Meyer D, Dimitriadou E, Hornik K, Weingessel A & Leisch F e1071: Misc Functions of the Department of Statistics, Probability Theory Group (Formerly: E1071) (2017). [Google Scholar]
  • 38.Mi H et al. PANTHER version 11: expanded annotation data from Gene Ontology and Reactome pathways, and data analysis tool enhancements. Nucleic Acids Res 45, D183–D189 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Subramanian A, Tamayo P, Mootha VK, Mukherjee S & Ebert BL Gene set enrichment analysis : A knowledge-based approach for interpreting genome-wide (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
1509549_Sup_tab_1
1509549_Sup_tab_3
1509549_Sup_tab_2

Data Availability Statement

All gene expression data was acquired from the public domain and is listed in Supplementary Table 1. The training data used by FIT can be downloaded from the site and package mentioned under “Code availability” section.

RESOURCES