Skip to main content
Journal of Computational Biology logoLink to Journal of Computational Biology
. 2023 Jun 14;30(6):726–735. doi: 10.1089/cmb.2022.0243

EnsMOD: A Software Program for Omics Sample Outlier Detection

Nathan P Manes 1,*, Jian Song 1,*, Aleksandra Nita-lazar 1,
PMCID: PMC10282819  PMID: 37042708

Abstract

Detection of omics sample outliers is important for preventing erroneous biological conclusions, developing robust experimental protocols, and discovering rare biological states. Two recent publications describe robust algorithms for detecting transcriptomic sample outliers, but neither algorithm had been incorporated into a software tool for scientists. Here we describe Ensemble Methods for Outlier Detection (EnsMOD) which incorporates both algorithms. EnsMOD calculates how closely the quantitation variation follows a normal distribution, plots the density curves of each sample to visualize anomalies, performs hierarchical cluster analyses to calculate how closely the samples cluster with each other, and performs robust principal component analyses to statistically test if any sample is an outlier. The probabilistic threshold parameters can be easily adjusted to tighten or loosen the outlier detection stringency. EnsMOD can be used to analyze any omics dataset with normally distributed variance. Here it was used to analyze a simulated proteomics dataset, a multiomic (proteome and transcriptome) dataset, a single-cell proteomics dataset, and a phosphoproteomics dataset. EnsMOD successfully identified all of the simulated outliers, and subsequent removal of a detected outlier improved data quality for downstream statistical analyses.

Keywords: hierarchical cluster analysis, multivariate, omics, outlier detection, proteomics, robust principal component analysis

1. INTRODUCTION

Outlier detection has been the subject of extensive research for over a century (Aggarwal, 2017; Pimentel et al., 2014). An often-cited definition is “An outlier is an observation that deviates so much from other observations as to arouse suspicion that it was generated by a different mechanism” (Hawkins, 1980). An outlier is a statistical anomaly, and in general it is reasonable to expect that the percentage of outliers is between 0% and 5%. In biomedical research, a biological outlier is the result of a biological mechanism, and a technical outlier is caused by a technical (nonbiological) mechanism.

Regardless of the mechanism, outliers should be avoided if they impair a statistical analysis. For example, consider a dataset consisting of 10 replicate measurements. Nine measurements were sampled from a standard normal distribution (mean = 0 and standard deviation = 1) resulting from Mechanism A. The other measurement is an outlier, and it was sampled from a different normal distribution (mean = 20 and standard deviation = 1) resulting from Mechanism B. In general, including the outlier will reduce the accuracy of the estimated mean resulting from Mechanism A, and it will impair other descriptive statistics and statistical modeling, and this can result in an incorrect biological conclusion. Of course, biological outliers might contain important information about a previously unobserved biological mechanism (e.g., a rare signaling event). Likewise, technical outliers can point researchers toward how to improve their experimental methods. Therefore, detection of outliers is important not only for avoiding erroneous biological conclusions, but also for discovering rare biological phenomena and improving research protocols.

Strategies for outlier mitigation include repeating the experiment from the beginning, increasing the number of samples to wash out the effect of the tentative outlier(s), improving the experimental protocol to better avoid outliers, taking no action [i.e., retaining the tentative outlier(s)], and detecting and removing (or winsorizing) the outlier(s). Numerous tests for detecting univariate outliers have been developed, including the Grubbs test (i.e., the extreme studentized deviate test), the Tietjen–Moore test, the generalized extreme studentized deviate test, and the ROUT method (Aggarwal, 2017; Manoj and Senthamarai Kannan, 2013; Motulsky and Brown, 2006). In general, these tests are unsuitable for detecting outliers within omics datasets because omics datasets are highly multivariate. Numerous algorithms have been developed for multivariate sample outlier detection, and these have been reviewed in detail by others (Aggarwal, 2017; Filzmoser and Todorov, 2013; Filzmoser and Todorov, 2011; Møller et al., 2005; Rousseeuw and Hubert, 2018). Some of these algorithms were designed for specific nonomics applications and are unsuitable for omics sample outlier detection.

RNA-seq sample outlier detection and removal has been shown to statistically benefit downstream data analyses (Chen et al., 2020; Merino et al., 2016), and this is likely true for omics data in general. RNA-seq quality control guidelines recommend detecting outliers by manually inspecting a classical (i.e., not robust) principal component analysis (PCA) biplot (Conesa et al., 2016; Merino et al., 2016). However, this approach is subjective and not statistically robust, especially because outliers can skew classical PCA analyses, causing the results to be unreliable (Hubert et al., 2005b).

Robust PCA (rPCA) algorithms use nonclassical PCA and statistical modeling to confidently detect multivariate sample outliers that exceed a preset probabilistic threshold (Rousseeuw and Hubert, 2018). rPCA algorithms include PcaCov (Todorov and Filzmoser, 2009), PcaGrid (Croux et al., 2007), PcaHubert (ROBPCA) (Hubert et al., 2005a, 2005b), and PcaLocantore (Locantore et al., 1999). These four rPCA algorithms are implemented in the R package rrcov (Todorov and Filzmoser, 2009). In a comparison of six PCA algorithms (classical PCA, PcaProj, PcaGrid, ROBPCA, PcaLocantore, and WCPCA), PcaGrid had the fewest false positives and ROBPCA was the most sensitive (Pascoal et al., 2010). In a comparison of three PCA algorithms [classical PCA, PcaGrid, and PcaHubert (ROBPCA)], PcaGrid had the best sensitivity and specificity (Chen et al., 2020).

An alternative algorithm “Ensemble outlier detection” was developed to detect transcriptomics sample outliers using logistic regression with elastic net regularization, sparse partial least squares discriminant analysis, and sparse generalized partial least squares analysis (Lopes et al., 2018). Ensemble outlier detection was compared to two transcriptomics sample misclassification algorithms using an RNA-seq dataset, and it was found to be relatively accurate (Sun et al., 2020). However, Ensemble outlier detection has only been used to analyze datasets containing ≥100 samples. Many omics datasets contain <100 samples, and it is currently unclear if Ensemble Outlier Detection would accurately detect outliers within them.

The rPCA outlier detection algorithms described above assume that the overall measurement variation of the input table follows a normal distribution. An rPCA-based algorithm for non-normal variation has been developed, but it has not yet been tested using omics datasets (Hubert et al., 2009). Most biological outcomes are the result of aggregation of numerous semi-independent non-Gaussian processes, each small in scale relative to the aggregate, and if no process dominates, the distribution typically converges to the normal distribution (Frank, 2009). In contrast, some biological processes are dominated by one or two non-Gaussian processes.

Also, many experimental measurements use technical mechanisms that result in variation that does not follow a normal distribution. For example, we and others have reported that quantitative proteomics using liquid chromatography–mass spectrometry (LC-MS) results in variation that follows a nearly log-normal distribution (Boehm et al., 2007; Ji et al., 2015; Manes et al., 2011). Log-transformation of these data result in variation that is nearly normally distributed. If the variation of a dataset deviates from a normal distribution, then a transformation to achieve normality should be considered (Huber et al., 2002; Kelmansky et al., 2013; Raymaekers and Rousseeuw, 2021; Rocke and Durbin, 2003).

Recently, our laboratory used quantitative LC-MS proteomics to study dissected regions of clinical esophageal excisional biopsy samples. So far, we have analyzed 20 samples, and this project is ongoing. These samples are proteomically diverse, and they might originate from slightly different esophageal regions, resulting in different populations of cells and proteins. Consequently, we decided to use two previously published strategies to robustly test for proteomics sample outliers within our dataset.

For the first strategy, PcaGrid was used to detect transcriptomics sample outliers within simulated and genuine RNA-seq datasets (Chen et al., 2020). The outlier detection stringency parameter was set to 97.5% (for both the score distance and the orthogonal distance). For data with normally distributed variation, this value is the estimated fraction of the samples that are not falsely classified as outliers. After outlier removal, the authors performed statistical testing and discovered novel differentially expressed genes (DEGs). They performed quantitative reverse transcription PCR and validated these novel DEGs. Therefore, the authors demonstrated the utility of sample outlier detection and removal for improving RNA-seq data analysis and DEG discovery.

For the second strategy, hierarchical cluster analysis (HCA) and ROBPCA were integrated to analyze simulated and genuine gene microarray datasets (Selicato et al., 2021). The HCA used the Pearson correlation coefficient as the distance function, and five linkage functions (average, complete, single, centroid, and Ward.D2) were compared. For each HCA, the cophenetic correlation coefficient (CCC) was calculated. The CCC is a measure of how well a dendrogram preserves the pairwise distances of the original dataset. The CCC can range from zero (the clustering is uninformative) to one (the clusters perfectly represent the original distances). The linkage method resulting in the largest CCC was used for the downstream analyses, and a CCC ≥0.8 was required.

The gap statistics algorithm was used to calculate the optimal number of clusters, and the Silhouette coefficient (SC) was calculated for each sample. The SC can range from −1 (the sample would fit much better in a different cluster) through zero (the sample does not fit in any cluster) to one (the sample fits perfectly in its cluster). A sample with SC <0.25 was considered a potential outlier (“No substantial structures have been found.”). A sample with 0.25 ≤ SC <0.5 was borderline (“The structure is weak and may be artificial.”). ROBPCA (specifically, robpca of the rospca R package) was used with the outlier detection stringency parameter set to 97.5% (for both the score distance and the orthogonal distance).

Unfortunately, neither of these two strategies had been incorporated into a software program accessible to scientists without a strong background in bioinformatics or biostatistics. Consequently, we implemented both of these strategies within Ensemble Methods for Outlier Detection (EnsMOD). EnsMOD can be used to analyze any omics dataset with normally distributed variance. EnsMOD has been made freely available to the public (https://github.com/niaid/EnsMOD ), it is open-source, and users do not need a strong background in bioinformatics or biostatistics to use it to analyze their own datasets.

2. WORKFLOW AND USAGE

EnsMOD is provided both as a stand-alone Rmarkdown and as a R Shiny application with a graphical user interface. Both versions have exactly the same functionality. To use the EnsMOD application or Rmarkdown, users needed to install R (https://www.r-project.org/ ) and RStudio (https://www.rstudio.com/ ). To use the application version, download it as a ZIP file (freely available at https://github.com/niaid/EnsMOD), decompress the ZIP file into a directory, and open the/app/app.R file in RStudio. To use the Rmarkdown version, simply download “Ensemble_Methods_for_Outlier_Detection_v2_0_stand_alone.Rmd” from the same GitHub location, and open it using RStudio.

The Rmarkdown and application versions automatically install and update all of their required R packages. Alternatively, to manually install the R packages required by the script version, open RStudio and select “Tools” and “Install Packages…” from the drop-down menu, and install the required packages (including dependencies): BiocManager, cluster, factoextra, fitdistrplus, ggraph, gplots, RColorBrewer, rospca, rrcov, and tidyverse; use the Console to run “BiocManager::install(limma)” to install the limma package. It is possible that antivirus software will need to be paused during these steps. To confirm that the “cluster” package was successfully installed, try to load it by running “library(cluster)” using the Console, and similarly for the other packages. In addition to the above R packages, the application version requires: shiny, shinyjs, xfun, DT, readr, dplyr, data.table, reshape2, htmltools, and readxl.

The input table needs to be in an XLSX file named “Gene_Expression_Table.xlsx,” and it needs to be in the same directory as the EnsMOD Rmarkdown. The application version allows users to open their input data files using the GUI. The XLSX file should contain only one worksheet and only one table. In this table of gene/protein/etc. abundance data, each column corresponds to a sample, and each row corresponds to a gene (or a protein, or a metabolite, etc., depending on the omics dataset type). The first row (the header) should contain unique identifiers of the samples. The first column is not for gene/protein/etc. identifiers (they are not needed); all of the columns correspond to samples and contain abundance values. Rows can contain missing values (such as “NaN”), but these rows will be excluded from the analyses (note that “0” is not treated as a missing value). Data imputation might be necessary if most of the rows contain one or more missing values. The overall quantitation variation is assumed to be normally distributed. If this requires a transformation of the data (e.g., log-transformation), it must be performed separately before the EnsMOD analysis.

Using RStudio, click the “Knit” button to run the Rmarkdown version, or click “Run App” to run the application version. The output files will be in the current working directory (Rmarkdown version) and in/app/www/EnsMODoutpupts/(application version). For the input datasets that we describe below, the runtimes ranged from ∼1 minute to 5 hours (the SCoPE2 dataset took 5 hours; the input table dimensions were 1490 samples by 731 proteins).

Running EnsMOD will produce an html file (in the same directory as the Rmarkdown) that can be reviewed using a web browser. Each step of the Rmarkdown is described in the output html file. First, the four outlier detection stringency parameters are set (the minimum CCC threshold, the maximum SC threshold, and the robpca and PcaGrid probabilistic thresholds). These four threshold values can easily be loosened or tightened by the user in RStudio. The input data are imported, and rows containing one or more missing values are excluded from the analyses. A density curve for each sample is plotted. A sample with a density curve different from the others might be an outlier and should be carefully investigated.

The overall quantitation variation is assumed to follow a normal distribution. No genuine experimental data exactly follow a normal distribution, and unfortunately it is unknown how non-normal the variation can be before the rPCA outlier detection algorithms become erroneous. Therefore, EnsMOD includes multiple tools to inspect the normality of the variation. An empirical histogram and density curve is plotted against a standard normal distribution (mean = 0, standard deviation = 1), and the corresponding coefficient of determination (R2) is calculated (using the empirical density curve and the standard normal distribution). The empirical and fitted cumulative distributions are plotted, as are the Quantile–Quantile and Probability–Probability plots. The skewness versus kurtosis is plotted, and the Shapiro–Wilk and Kolmogorov–Smirnov tests for normality are performed. If the variation deviates from a normal distribution, then an upstream data transformation to achieve normality should be considered (Huber et al., 2002; Kelmansky et al., 2013; Raymaekers and Rousseeuw, 2021; Rocke and Durbin, 2003).

In Sections 2 and 3 of the Rmarkdown, 15 HCAs are performed using three distance functions (Euclidean, Manhattan, and Pearson) and five linkage functions (average, complete, single, centroid, and Ward.D2). Nonparametric distance functions (e.g., Spearman's rank correlation coefficient) were not used because they are not recommended for data with normally distributed variation. The HCA distance–linkage pair that resulted in the highest CCC was selected for the downstream analyses. The optimal number of clusters was calculated using gap statistics, and if this failed, the maximum average SC method was used (Charrad et al., 2014). The SC for each sample was calculated using the eclust function of the factoextra R package. Note that eclust is limited to a maximum of 20 clusters. eclust might work poorly for datasets with more than ∼20 experimental conditions. Although not ideal, large omics datasets with more than ∼20 experimental conditions could be analyzed by partitioning the samples into subsets (each having <20 experimental conditions) and using EnsMOD to analyze each subset separately (this would just be for the SC values; EnsMOD would be used normally for all of the other values).

In Section 4, a classical PCA is performed to visualize the data on a biplot of principal component 1 versus principal component 2. The robpca and PcaGrid rPCA algorithms are performed, and distance–distance plots are used to robustly detect outliers. Note that the density plot and the HCA analyses are only descriptive, and do not rely on a probability threshold to detect outliers. In contrast, robpca and PcaGrid both use statistical modeling to detect multivariate outliers that exceed their preset probability threshold. In Section 5 the results are summarized.

We first used EnsMOD to analyze a simulated proteomics dataset. There were 20 samples in total, 19 were partitioned into 4 groups (each contained 4 or 5 samples), and there was 1 outlier. For each gene and sample group, the abundance values were drawn from a normal distribution (coefficient of variation = 15%). EnsMOD easily identified the outlier. The simulated proteomics dataset and the corresponding EnsMOD output are available at https://github.com/niaid/EnsMOD/tree/main/EnsMOD_Examples. We used EnsMOD to analyze our esophageal proteomics dataset (n = 20), and we detected three outliers. After careful consideration, we decided to double our sample set (n = 40), and this work is ongoing.

We used EnsMOD to analyze our multiomic dataset (proteomics and transcriptomics) from analyses of terminal ileum tissue from conventional and germ-free mice (n = 39; female and male; BALB/c and C57BL/10A strains; 3088 proteins; 35,958 transcript probes) (Manes et al., 2017). The proteomic and transcriptomic datasets were analyzed separately. Analysis of the proteomic dataset resulted in the detection of zero outliers. Analysis of the transcriptomic dataset resulted in the detection of one outlier (“OME02” satisfied the HCA, robpca, and PcaGrid criteria). Analyses of multimodal datasets for sample outlier detection might be expected to result in the detection of the same set of outliers from the multiple data types if they result from a biological mechanism.

In this study, sample OME02 was a transcriptomic outlier and not a proteomic outlier, and we would expect that a biological outlier would typically be both. An exception would be if a mechanism altered the transcriptome, but this had not yet translated into an alteration of the proteome. OME02 might be a technical outlier, it might be a false positive, or it may be a biological outlier resulting from mRNA–protein discordance. The mouse ileum multiomic dataset and the corresponding EnsMOD output are available at https://github.com/niaid/EnsMOD/tree/main/EnsMOD_Examples.

Single-cell transcriptomic and proteomic datasets are often highly sparse and require specialized algorithms for sample outlier detection. In addition to correctly detecting technical outliers, these algorithms need to avoid misclassification of small populations of rare cell types that might resemble biological outliers. Numerous single-cell RNA-seq outlier detection algorithms have been developed, including EnsembleKQC (Ma et al., 2019), FEATS (Vans et al., 2021), Granatum (Zhu et al., 2017), MetaCell (Baran et al., 2019), RaceID3 (Sagar and Grun, 2019), SC3 (Kiselev et al., 2017), Scater (McCarthy et al., 2017), scImpute (Li and Li, 2018), scPipe (Tian et al., 2018), and scRNABatchQC (Liu et al., 2019). We used EnsMOD to analyze a published single-cell proteomics dataset (Specht et al., 2021). The “Proteins-processed.csv” table described in the article has dimensions 1490 samples (i.e., cells) by 3042 proteins. It contains no missing values because data imputation was performed by the authors. We imposed an additional requirement that each protein have a peptide that was quantified within ≥20% of the samples (298 of 1490).

The resulting table had dimensions 1490 samples by 731 proteins. EnsMOD was run and the HCA CCC equaled 0.698, which indicated that the clustering was of little value (informative HCA clustering results in CCC ≥0.8). Three samples were identified as outliers by both robpca and PcaGrid (i300, i518, and i1640). Further research is needed before single-cell proteomic outliers can be confidently detected. We, therefore, do not recommend using EnsMOD to analyze sparse datasets (missing measured [i.e., nonimputed] values >50%), and we caution that data imputation might negatively affect omics outlier detection in general. The single-cell proteomics dataset and the corresponding EnsMOD output are available at https://github.com/niaid/EnsMOD/tree/main/EnsMOD_Examples.

Lastly, we used EnsMOD to analyze our quantitative LC-MS phosphoproteomic time-course dataset of spleens from mice with anthrax (Manes et al., 2011). Mice were not injected at all, or were injected (intraperitoneal) with vehicle alone, or with either the Sterne strain (lethal in 2–3 days) or the ΔSterne strain (asymptomatic) of Bacillus anthracis. Fifteen mice were prepared for the Sterne 72 hours experimental condition, but only one reached 72 hours. Spleens were isolated at 24, 48, and 72 hours postinjection for LC-MS phosphoproteomics (Fig. 1A). The one Sterne-infected mouse that survived 72 hours functioned here as a “pseudo-outlier” because it was the sole member of its experimental group. EnsMOD was used to analyze this dataset (specifically, the globally normalized dataset from the article), and the variance was found to closely follow a normal distribution (Fig. 1B). The density curves were all similar to each other (Fig. 1C). The HCA that used Euclidean distance and average linkage resulted in the largest CCC, which was equal to 0.899, indicating that the clustering was informative (Fig. 1D).

FIG. 1.

FIG. 1.

FIG. 1.

Outlier detection from analyzing the anthrax phosphoproteomics dataset. Mice were not injected at all, or were injected with vehicle alone, or with either the Sterne or the ΔSterne strain of Bacillus anthracis (A). Fifteen mice were prepared for the Sterne 72 hours experimental condition, but only one reached 72 hours. Spleens were analyzed using LC-MS phosphoproteomics (Manes et al., 2011). Empirical histogram and density curve (black) plotted against the standard normal distribution (red) (B). The R2 value was calculated between the empirical density curve and the standard normal distribution. Density curve of the phosphopeptide abundance values for each sample (C). Dendrogram of the samples from the HCA with the largest CCC (D). Silhouette coefficients for each sample (black line is the cutoff, red dashed line is the mean calculated across all of the samples) (E). Distance–Distance plots from the robpca and PcaGrid analyses (F, G). Venn diagram of the differentially abundant phosphopeptides discovered using one-way ANOVAs (q-value ≤0.05), with (left circle) and without (right circle) including the genuine sample outlier (SplnC0_0003) (H). CCC, cophenetic correlation coefficient; HCA, hierarchical cluster analysis; LC-MS, liquid chromatography–mass spectrometry.

Four of the 46 samples had SC <0.25, indicating that these four samples (SplnC0_0003, SplnS2_0001, SplnS2_0004, and SplnS3_0001) were potential outliers (Fig. 1E). The robpca algorithm detected 16 outliers (Fig. 1F), which is an implausibly high number (16/46 = 35%, and we required 0%–5%). The PcaGrid algorithm detected eight outliers (Fig. 1G), which is also implausibly high (8/46 = 17%). There was not an obvious pattern to the implausible rPCA results (the Stern mice appeared to be overrepresented). We noted that requiring the CCC+SC+ROBPCA criteria (Selicato et al., 2021) resulted in the detection of just four outliers (SplnC0_0003, SplnS2_0001, SplnS2_0004, and SplnS3_0001) (4/46 = 8.7%).

Next, we increased the robpca, and PcaGrid probability threshold parameters from 0.975 (the default values) to 0.999 to make outlier classification much stricter. The robpca algorithm still detected the same sixteen outliers, but the PcaGrid algorithm now detected just three outliers (SplnC0_0003, SplnS2_0002, and SplnS3_0001), which was far more reasonable (3/46 = 6.5%). As mentioned in the introduction, an outlier is a statistical anomaly, and in general it is reasonable to expect that the percentage of outliers is between 0% and 5%. Therefore, we recommend using strict rPCA thresholds, and we recommend against using robpca alone.

Furthermore, we examined which outliers were common across the three analyses (HCA SC, robpca, and PcaGrid) when we used the default robpca and PcaGrid probability thresholds (0.975). Four outliers were common to both rPCA analyses (SplnC0_0003, SplnS2_0002, SplnS2_0003, and SplnS3_0001). If all three criteria were required, two outliers were detected (SplnC0_0003 and SplnS3_0001). The Sterne 72 hours pseudo-outlier SplnS3_0001 was successfully detected, suggesting that the use of all three criteria is a reasonable strategy for outlier detection. One of the noninjected mice (SplnC0_0003) was detected as a genuine outlier. The anthrax spleen phosphoproteomics dataset and the corresponding EnsMOD output are available at https://github.com/niaid/EnsMOD/tree/main/EnsMOD_Examples.

This dataset included 20 uninfected mice, and we questioned if a single outlier among them could impact downstream statistical analyses. To test this, we used InfernoRDN v. 1.1.7995 (2021-OCT-12) (https://github.com/PNNL-Comp-Mass-Spec/InfernoRDN) (Polpitiya et al., 2008) to perform one-way ANOVAs both with and without including sample SplnC0_0003 (q-value ≤0.05 was required; the pseudo-outlier SplnS3_0001 was not included in either set of ANOVA analyses; all of the uninfected mice were grouped into a single experimental condition for all of the ANOVAs). Retaining sample SplnC0_0003 resulted in the discovery of 574 differentially abundant phosphopeptides, but removing this sample resulted in 617 (Fig. 1H). Apparently, this one outlier among the twenty uninfected mice caused a negative impact on the ANOVAs by increasing the number of false negative results. Thus, omics sample outlier detection and removal (or the use of another outlier mitigation strategy) is important for improving discovery of differentially abundant analytes within omics datasets.

3. CONCLUSION

EnsMOD enables scientists to confidently detect sample outliers within their datasets. A strong background in bioinformatics or biostatistics is not required to use EnsMOD. We used EnsMOD to analyze our anthrax phosphoproteomics dataset and discovered novel differentially abundant phosphopeptides. We also used EnsMOD to analyze our esophageal proteomics dataset and EnsMOD detected three outliers, which prompted us to double our sample set. In this way, EnsMOD enables scientists to mitigate the effects of statistical anomalies.

4. AVAILABILITY

EnsMOD (both application and stand-alone versions) is open source and freely available at https://github.com/niaid/EnsMOD.

ACKNOWLEDGMENTS

The authors thank Mark Rochman and Marc E. Rothenberg at Cincinnati Children's Hospital Medical Center for their contributions.

AUTHORs' CONTRIBUTIONS

Software: J.S., N.P.M.; Supervision: A.N.-L.; Writing—original draft: N.P.M., Writing—review and editing: N.P.M., J.S., A.N.-L.

AUTHOR DISCLOSURE STATEMENT

The authors declare they have no conflicting financial interests.

FUNDING INFORMATION

This work was supported by the Intramural Research Program of NIAID, NIH.

REFERENCES

  1. Aggarwal CC. Outlier Analysis, 2nd ed. Springer International Publishing AG: Cham, Switzerland; 2017; doi: 10.1007/978-3-319-47578-3 [DOI] [Google Scholar]
  2. Baran Y, Bercovich A, Sebe-Pedros A, et al. MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions. Genome Biol 2019;20:206; doi: 10.1186/s13059-019-1812-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Boehm AM, Putz S, Altenhofer D, et al. Precise protein quantification based on peptide quantification using iTRAQ. BMC Bioinformatics 2007;8:214; doi: 10.1186/1471-2105-8-214 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Charrad M, Ghazzali N, Boiteau V, et al. NbClust: An R package for determining the relevant number of clusters in a data set. J Stat Softw 2014;61:1–36; doi: 10.18637/jss.v061.i06 [DOI] [Google Scholar]
  5. Chen X, Zhang B, Wang T, et al. Robust principal component analysis for accurate outlier sample detection in RNA-Seq data. BMC Bioinformatics 2020;21:269; doi: 10.1186/s12859-020-03608-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Conesa A, Madrigal P, Tarazona S, et al. A survey of best practices for RNA-seq data analysis. Genome Biol 2016;17:13; doi: 10.1186/s13059-016-0881-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Croux C, Filzmoser P, Oliveira MR. Algorithms for Projection–Pursuit robust principal component analysis. Chemometr Intell Lab Syst 2007;87:218–225; doi: 10.1016/j.chemolab.2007.01.004 [DOI] [Google Scholar]
  8. Filzmoser P, Todorov V. Review of robust multivariate statistical methods in high dimension. Anal Chim Acta 2011;705:2–14; doi: 10.1016/j.aca.2011.03.055 [DOI] [PubMed] [Google Scholar]
  9. Filzmoser P, Todorov V. Robust tools for the imperfect world. Inf Sci 2013;245:4–20; doi: 10.1016/j.ins.2012.10.017 [DOI] [Google Scholar]
  10. Frank SA. The common patterns of nature. J Evol Biol 2009;22:1563–1585; doi: 10.1111/j.1420-9101.2009.01775.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Hawkins DM. Identification of Outliers. Springer International Publishing AG: Cham, Switzerland; 1980; doi: 10.1007/978-94-015-3994-4 [DOI] [Google Scholar]
  12. Huber W, von Heydebreck A, Sultmann H, et al. Variance stabilization applied to microarray data calibration and to the quantification of differential expression. Bioinformatics 2002;18 Suppl 1:S96–S104; doi: 10.1093/bioinformatics/18.suppl_1.s96 [DOI] [PubMed] [Google Scholar]
  13. Hubert M, Rousseeuw P, Verdonck T. Robust PCA for skewed data and its outlier map. Comput Stat Data Anal 2009;53:2264–2274; doi: 10.1016/j.csda.2008.05.027 [DOI] [Google Scholar]
  14. Hubert M, Rousseeuw PJ, Van Aelst S. 10—Multivariate outlier detection and robustness. In: Handbook of Statistics. vol. 24 (Rao CR, Wegman EJ, Solka JL. eds.) Elsevier: Amsterdam, Netherlands; 2005a; pp. 263–302; doi: 10.1016/S0169-7161(04)24010-X [DOI] [Google Scholar]
  15. Hubert M, Rousseeuw PJ, Vanden Branden K. ROBPCA: A new approach to robust principal component analysis. Technometrics 2005b;47:64–79; doi: 10.1198/004017004000000563 [DOI] [Google Scholar]
  16. Ji C, Li YF, Bellinger EP, et al. A maximum-likelihood approach to absolute protein quantification in mass spectrometry. In: Proceedings of the 6th ACM Conference on Bioinformatics, Computational Biology and Health Informatics. 2015; pp. 296–305; doi: 10.1145/2808719.2808750 [DOI] [Google Scholar]
  17. Kelmansky DM, Martinez EJ, Leiva V. A new variance stabilizing transformation for gene expression data analysis. Stat Appl Genet Mol Biol 2013;12:653–666; doi: 10.1515/sagmb-2012-0030 [DOI] [PubMed] [Google Scholar]
  18. Kiselev VY, Kirschner K, Schaub MT, et al. SC3: Consensus clustering of single-cell RNA-seq data. Nat Methods 2017;14:483–486; doi: 10.1038/nmeth.4236 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Li WV, Li JJ. An accurate and robust imputation method scImpute for single-cell RNA-seq data. Nat Commun 2018;9:997; doi: 10.1038/s41467-018-03405-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Liu Q, Sheng Q, Ping J, et al. scRNABatchQC: Multi-samples quality control for single cell RNA-seq data. Bioinformatics 2019;35:5306–5308; doi: 10.1093/bioinformatics/btz601 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Locantore N, Marron JS, Simpson DG, et al. Robust principal component analysis for functional data. Test 1999;8:1–73; doi: 10.1007/BF02595862 [DOI] [Google Scholar]
  22. Lopes MB, Verissimo A, Carrasquinha E, et al. Ensemble outlier detection and gene selection in triple-negative breast cancer data. BMC Bioinformatics 2018;19:168; doi: 10.1186/s12859-018-2149-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Ma A, Zhu Z, Ye M, et al. EnsembleKQC: An unsupervised ensemble learning method for quality control of single cell RNA-seq sequencing data. In: Intelligent Computing Theories and Application. (Huang DS, Jo KH, Huang ZK. eds.) Springer International Publishing: New York, NY; 2019; pp. 493–504; doi: 10.1007/978-3-030-26969-2_47 [DOI] [Google Scholar]
  24. Manes NP, Dong L, Zhou W, et al. Discovery of mouse spleen signaling responses to anthrax using label-free quantitative phosphoproteomics via mass spectrometry. Mol Cell Proteomics 2011;10:M110.000927; doi: 10.1074/mcp.M110.000927 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Manes NP, Shulzhenko N, Nuccio AG, et al. Multi-omics comparative analysis reveals multiple layers of host signaling pathway regulation by the gut microbiota. mSystems 2017;2:e00107-17; doi: 10.1128/mSystems.00107-17 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Manoj K, Senthamarai Kannan K. Comparison of methods for detecting outliers. International J Sci Eng Res 2013;4:709–714; ISSN: 2229–5518 [Google Scholar]
  27. McCarthy DJ, Campbell KR, Lun AT, et al. Scater: Pre-processing, quality control, normalization and visualization of single-cell RNA-seq data in R. Bioinformatics 2017;33:1179–1186; doi: 10.1093/bioinformatics/btw777 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Merino GA, Fresno C, Netto F, et al. The impact of quality control in RNA-seq experiments. J Phys Conf Series 2016;705:012003; doi: 10.1088/1742-6596/705/1/012003 [DOI] [Google Scholar]
  29. Møller SF, von Frese J, Bro R. Robust methods for multivariate data analysis. J Chemometr 2005;19:549–563; doi: 10.1002/cem.962 [DOI] [Google Scholar]
  30. Motulsky HJ, Brown RE. Detecting outliers when fitting data with nonlinear regression—a new method based on robust nonlinear regression and the false discovery rate. BMC Bioinformatics 2006;7:123; doi: 10.1186/1471-2105-7-123 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Pascoal C, Oliveira MR, Pacheco A, et al. Detection of outliers using robust principal component analysis: A simulation study. In: Combining Soft Computing and Statistical Methods in Data Analysis. Springer: Berlin; 2010; pp. 499–507; doi: 10.1007/978-3-642-14746-3_62 [DOI] [Google Scholar]
  32. Pimentel MAF, Clifton DA, Clifton L, et al. A review of novelty detection. Signal Process 2014;99:215–249; doi: 10.1016/j.sigpro.2013.12.026 [DOI] [Google Scholar]
  33. Polpitiya AD, Qian WJ, Jaitly N, et al. DAnTE: A statistical tool for quantitative analysis of -omics data. Bioinformatics 2008;24:1556–1558; doi: 10.1093/bioinformatics/btn217 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Raymaekers J, Rousseeuw PJ. Transforming variables to central normality. Mach Learn 2021; doi: 10.1007/s10994-021-05960-5 [DOI] [Google Scholar]
  35. Rocke DM, Durbin B. Approximate variance-stabilizing transformations for gene-expression microarray data. Bioinformatics 2003;19:966–972; doi: 10.1093/bioinformatics/btg107 [DOI] [PubMed] [Google Scholar]
  36. Rousseeuw PJ, Hubert M. Anomaly detection by robust statistics. WIREs Data Mining Knowl Discov 2018;8:e1236; doi: 10.1002/widm.1236 [DOI] [Google Scholar]
  37. Sagar, Grun D.. Lineage inference and stem cell identity prediction using single-cell RNA-sequencing data. Methods Mol Biol 2019;1975:277–301; doi: 10.1007/978-1-4939-9224-9_13 [DOI] [PubMed] [Google Scholar]
  38. Selicato L, Esposito F, Gargano G, et al. A new ensemble method for detecting anomalies in gene expression matrices. Mathematics 2021;9:882; doi: 10.3390/math9080882 [DOI] [Google Scholar]
  39. Specht H, Emmott E, Petelski AA, et al. Single-cell proteomic and transcriptomic analysis of macrophage heterogeneity using SCoPE2. Genome Biol 2021;22:50; doi: 10.1186/s13059-021-02267-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Sun H, Cui Y, Wang H, et al. Comparison of methods for the detection of outliers and associated biomarkers in mislabeled omics data. BMC Bioinformatics 2020;21:357; doi: 10.1186/s12859-020-03653-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Tian L, Su S, Dong X, et al. scPipe: A flexible R/Bioconductor preprocessing pipeline for single-cell RNA-sequencing data. PLoS Comput Biol 2018;14:e1006361; doi: 10.1371/journal.pcbi.1006361 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Todorov V, Filzmoser P. An object-oriented framework for robust multivariate analysis. J Stat Softw 2009;32:1–47; doi: 10.18637/jss.v032.i03 [DOI] [Google Scholar]
  43. Vans E, Patil A, Sharma A. FEATS: Feature selection-based clustering of single-cell RNA-seq data. Brief Bioinform 2021;22:bbaa306; doi: 10.1093/bib/bbaa306 [DOI] [PubMed] [Google Scholar]
  44. Zhu X, Wolfgruber TK, Tasato A, et al. Granatum: A graphical single-cell RNA-Seq analysis pipeline for genomics scientists. Genome Med 2017;9:108; doi: 10.1186/s13073-017-0492-3 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Journal of Computational Biology are provided here courtesy of SAGE Publications

RESOURCES