Abstract
Existing genomic sequencing data can be used to study host-microbiome ecosystems, however distinguishing signals originating from truly present microbes versus contaminating species and artifacts is a substantial and often prohibitive challenge. Here we show that emerging sequencing technologies definitely capture reads from present microbes. We developed SAHMI, a computational resource to identify truly present microbial nucleic acids and filter contaminants and spurious false-positive taxonomic assignments from standard transcriptomic sequencing of mammalian tissues. In benchmark studies, SAHMI correctly identifies known microbial infections present in diverse tissues, and we validate SAHMI’s enrichment for correctly classified, truly present species using multiple orthogonal computational experiments. The application of SAHMI to single-cell and spatial genomic data thus enables co-detection of somatic cells and microorganisms and joint analysis of host-microbiome ecosystems.
Introduction
Retrospective genomics analysis of host-microbiome ecosystems is challenging because true signal is often obfuscated by abundant contamination and taxonomic classification noise (see Supplementary Information). While in silico decontamination approaches have been developed for specific purposes1,2, these methods generally rely on sample type or batch effect analysis of large-scale datasets, or they require extra experimental data. There are few methods3 that tackle the question of denoising microbial genomic data solely from individual samples or in the context of small studies. Accurate in silico decontamination of small-scale studies would greatly increase the utility of existing clinical genomic data for microbiome analysis.
Results
We thus developed SAHMI (Single-cell Analysis of Host-Microbiome Interactions), a computational pipeline to identify microbial reads and remove false positives and contaminants from emerging genomic data of mammalian tissues (Fig. 1a). SAHMI denoises and decontaminates the output of taxonomic classifiers (e.g. Kraken24). It truly present taxa by examining the relationship between the number of total and unique k-mers for each taxon, and it removes contaminants by comparing taxa profiles to a negative control reference dataset.
Figure 1. Filtering false positives with k-mer correlation tests.

a, Schematic of the SAHMI pipeline. SAHMI identifies taxa that are truly present in tissues using a k-mer correlation test and identifies false positives and contaminants by comparing taxa distributions to an extensive negative control reference. b, Scatter plot showing the total number of sequencing reads and species detected in each study. Blue, experimentally introduced pathogen; red, natural infection involving a human tissue. c, Boxplots showing significantly increased reads assigned to bacteria when the human genome is not included as a reference during taxonomic classification. Boxplots show median (line), 25th and 75th percentiles (box) and 1.5 x IQR (whiskers). Points represent outliers; 2-sided t-tests; **** p < 2e-16 (native R statistical test limit); *** p=0.005; ns, not significant. Left panel, n=23; middle, n=38; right, n=147. d, Histogram of the number of unique k-mers per sample assigned to the known pathogens and to all detected species in the benchmark studies. e, k-mer correlation test results for the truly present Salmonella enterica in Saliba et al., 2016 (top panels) and for Fusarium venenatum, an example false positive or contaminant from the same samples (bottom panels). The left three panels show correlations across samples, and the rightmost shows the k-mer test across barcodes. The plot titles indicate the Spearman correlation value. See also Figure S1 for correlations for all truly present taxa in each benchmark dataset. f, Scatter plot of the k-mer correlation tests for species in the benchmark studies. Each point represents an individual species. Correlations are run across samples within a study. The x-axis represents the Spearman correlation values between #k-mers vs. #unique k-mers. The y-axis represents the correlation value between #k-mers vs. #reads; colors represent correlation values between #reads vs. #unique k-mers. Lines represent contour densities. g, 2-sided Fisher test result for taxa in all benchmark studies that passed or failed the k-mer correlation tests. h, Boxplots showing the fraction of reads mapped per taxon with BLAST for taxa that passed or failed the k-mer correlation tests in a subset of the skin leprosy data. Boxplots are as in (b), 2-sided Wilcoxon rank sum testing. Failed, n=162; Passed, n=64. i, Boxplots showing the distribution of BLAST mapping scores for taxa that passed or failed the k-mer correlation tests for a subset of the skin leprosy data. Boxplots are as in (b), 2-sided Wilcoxon rank sum testing. Failed, n=162; Passed, n=64. j, Similar to (e) but for taxa detected in cell line experiments.
Defining false positive, contaminant, and truly present taxa
First, we confirmed that mRNA from truly present microbes is captured in single-cell RNA-sequencing (scRNA-seq) by using Kraken2 to human samples from 8 studies with clinically verified infection5–7 or in which a pathogen was experimentally introduced8–11 (see Supplementary Information, Supplementary Figure 1, Supplementary Data 1).
While the known pathogen was identified in each study, a mean of 4735 other species (range 801–7084) were also classified in the benchmark datasets, and total species correlated with total reads across studies (r=0.7, p=0.08, Fig. 1b). We also observed that reported microbial profiles differed substantially depending on the mapping parameters used. Mapping reads to the microbiome without including the host genome led to significantly increased reads that mapped to bacteria in general but to only negligible differences in the number of reads that mapped to the verified pathogen (t-test, p<1e-4; Fig. 1c). This underscores the importance of including all relevant reference genomes during taxonomic classification, and especially the host genome.
We reasoned that the putative microbial reads could either originate from (1) correctly classified microbes present in the original tissue (truly present), (2) correctly classified microbes introduced during sample handling or sequencing (contaminant), or (3) be misclassified reads (false positive). False positives may arise from multiple sources, including sequence homology, sequencing errors, off-target amplification, mapping errors, or reference genome errors, and contaminating reads may be introduced at various stages of sample handling. Both severely obfuscate analysis of tissues with unknown microbial burden and without matched negative controls.
The k-mer correlation tests
The number of unique k-mer sequences assigned to a taxon can be used to filter false positives3. However, we observed a wide range in these values in the benchmark studies, and unique k-mer counts alone did not distinguish the known pathogens (Fig. 1d). While high counts of unique k-mers may reflect wider genome coverage, they could also be a result of contamination in the reference genomes or artifacts in Kraken’s least common ancestor database. To mitigate these possibilities, we posited that species that are truly present the degree of genome coverage (as estimated by unique k-mers) would be proportional to the total reads sampled, consistent with the established principle in genomics that wider genome coverage is achieved with deeper sequencing depth13.
To test this, we first we examined the pairwise Spearman correlations across samples between the numbers of reads, k-mers, and unique k-mers for the human reads and indeed found strong relationships (Spearman r=0.99, r=0.93, r=0.94, all p<2e-16). We then calculated these correlation values for all species detected in the benchmark studies. Similar to the human data, the true pathogens had extremely high correlation values in all metrics (median r=0.99, p<2e-16, Fig. 1e, Supplementary Figure 2a, Supplementary Data 1). The relationships between total and unique k-mers were also strong for the known pathogens when tested across scRNA-seq barcodes (median r=0.92, p<2e-16, Fig. 1e, Supplementary Figure 2a). However, when we examined these correlation tests for all initially classified taxa, we found a wide spectrum of values (Fig. 1f), with only a minority of taxa (871/7084) having significant correlations (p<0.05) in all four tests.
Among the taxa with high correlation values were cutaneous species and common contaminants, such as C. acnes, S. epidermidis, M. restricta, and Mycoplasma spp., whereas more surprising species like Proteus virus Isfahan had nonsignificant values (Supplementary Data 2). This suggested that the correlation tests were selecting for correctly classified taxa (truly present and contaminants) and against misclassified false positives. To test this, for the lepromatic skin data6, we mapped reads from taxa that passed or failed the correlation tests to the human genome using STAR14 and to their respective microbial genomes using BLAST15. Despite explicitly including the human genome in the Kraken database, 23% of the putative microbial reads could be mapped to the human genome using STAR, and that reads from taxa that failed the k-mer tests were significantly more likely to map to the human genome and comprised the majority (92%) of human mappable reads (Fisher test, p<2e-16, Fig. 1g). Taxa that passed the k-mer correlation tests had significantly higher mapped read fractions to their genomes with BLAST (Wilcoxon p=2e-4, Fig. 1h) and had higher BLAST mapping scores (Wilcoxon p=4e-7, Fig. 1i). These results indicate that the k-mer correlation tests remove human mappable reads missed by Kraken and that they select for correctly classified microbial taxa.
The cell-line test
We next sought to identify the contaminants and remaining artifacts that passed the k-mer correlation tests based on the observation that contaminants appear in higher frequencies in negative control samples1. In the absence of matched controls, we posited that sterile cell line data could serve as a substitute. We profiled the microbiome from publicly available human cell-line RNA-seq data for 1,057 samples from >75 sources from around the world (Supplementary Data 3), thereby creating a negative control resource (Supplementary Data 4). From these studies, we identified a mean of 982 species per sample (range: 124–5918). K-mer correlation tests found a range of values (Fig. 1j) which were significantly weaker than those from the benchmark studies (Wilcoxon, p<2e-16) due to the cell line data being enriched in false positives. The most ubiquitous species included cutaneous microbiota and common contaminants (Supplementary Data 4), as has previously been identified16.
Comparing taxa reads counts from the benchmark studies to their distribution in the cell line data clearly distinguished true positive signal rom background noise in all studies (Fig. 2a). The known pathogens had higher reads per million microbiome reads compared to the 90th percentile of their respective distribution in the cell line data (or 99th percentile if zero-count samples are included), whereas most taxa did not. Only a minority of all reported species per sample (median: 2.8%, range: 0.26–22%) and initially classified microbial reads per sample (median: 12%, range: 3.2–81%) passed both the k-mer correlation and cell line quantile tests (Fig. 2b).
Figure 2. Filtering contaminants and false positives with the cell-line test.

a, Overlaid histograms of specified taxon reads per million microbiome reads assigned to the known pathogens (top row) and to example contaminants (bottom row) in the benchmark studies as well as in the cell line data. Bars are colored by study of origin. b, Bar plots of relative proportions of reads (left panel) and species (right panel) that passed or failed the k-mer correlation and cell-line quantile tests. c, Final SAHMI filtering results for all sample types. The 5 left columns indicate the fraction of samples that detected the truly present pathogen. The 3 right columns indicate the mean number of species detected per sample after SAHMI, before SAHMI, and the total number species detected per sample type after SAHMI filtering. See also Table S2. d, Table showing the number of taxa that passed or failed the cell-line quantile test and the k-mer correlation tests across all benchmark studies. e, Boxplots showing the number of reads per taxon mapped by STAR for microbial reads from the lepromatous leprosy study grouped by k-mer correlation and cell line quantile test results. Boxplots show median (line), 25th and 75th percentiles (box) and 1.5 x IQR (whiskers). Points represent outliers; p-values indicate 2-sided Wilcoxon rank sum testing. Sample sizes from left to right: n=11, n=42, n=4, n=51. f, Boxplots indicating the body site enrichment scores for taxa that passed SAHMI filtering in the skin leprosy study. Enrichment scores are calculated using subsampled and body site label shuffled data. Boxplots are as in (e). **** p < 1e-3; ns, not significant. 2-sided Wilcoxon rank sum testing performed. N=100 subsampled and n=100 shuffled replicates for each body site.
Examining the specific taxa that passed in each sample type indicates that SAHMI selected for truly present microbes (Fig. 2c, Supplementary Data 2). In the stomach samples, SAHMI identified H. pylori in 8/8 infected samples17 and only in 1/15 control samples. No other microorganisms were identified in the infected samples. In the HSV-1 cell-line infections, SAHMI identified HSV-1 in 34/38 infected samples compared to 7/18 uninfected, which indicates substantial cross-sample contamination greater than expected in general experiments; this could be mitigated by setting stricter thresholds during the cell-line test. Nonetheless, in these samples only 1–2 microorganisms in total were reported by SAHMI. In the Salmonella infection samples, SAHMI identified Salmonella in 8/14 samples with growing bacteria and reported no Salmonella in 46/46 samples without growing bacteria. In the samples from COVID-19 patients, SAHMI identified SARS-CoV-2 in 6/7 severe cases and none in mild cases or healthy controls. These were all samples in which only one microorganism was expected, and only 1–4 species were detected by SAHMI out of 121–6338 species/sample type (Fig. 2c) that were initially reported. However, the skin hosts a rich microbiome18, and the lepromatic skin samples SAHMI detected a total of 125 unique species, including M. leprae in 6/8 samples with infection. Examining the specific k-mer correlation and cell-line test results in all samples indicates that most taxa failed both tests, but that both tests were still necessary to filter out hundreds of microbes that were not truly present (Fisher test, p<2e-16, Fig. 2d, Supplementary Data 2).
Validation of SAHMI results
We validated our results from the skin leprosy study using two approaches. First, we used STAR14 to do full-length read mapping for reads that passed or failed the SAHMI tests. Species that passed both k-mer correlation and cell line quantile tests had significantly more mapped reads per species than all other initially reported taxa (Wilcoxon p<5e-4, Fig. 2e). Second, we used mBodyMap19 to assess the likely ecological source(s) of taxa from the skin samples that passed the SAHMI tests. The purpose of this experiment was to determine whether taxa that passed the SAHMI tests were predominantly the expected skin flora, which would indicate that SAHMI is enriching for ecologically expected microbes. We computed body site enrichment scores (see Methods, Body site enrichment analysis) for the SAHMI-reported taxa and found that these species were most dominantly enriched from the skin microbiome. The enrichment score for “skin” was both the highest of all body sites and was significantly greater than expected by chance when body site labels were randomized (Wilcoxon p<2e-16, Fig. 2j).
Lastly, we assessed the performance of SAHMI in the context of pancreatic cancer, where many tumors harbor gastrointestinal microbiota20,21. We analyzed pancreatic cancer samples that were dually sequenced by either scRNA-seq and whole genome sequencing or by total RNA-seq and 16S-rRNAgene-seq20. We posited that truly present genera and species would be more likely to be detected using two sequencing technologies compared to taxa that are false positives – as was demonstrated in a recent analysis of false positive detection using different sequencing platforms22. In both sets of experiments, taxa that passed SAHMI denoising criteria were significantly more likely to be detected in the same sample with the two separate technologies compared to taxa that failed SAMHI (p < 2e-16; Fisher test for both clinical sample cohorts Supplementary Figure 3). These results collectively validate that SAHMI enriches for truly present taxa.
Discussion
This study has a number of limitations. First, we only benchmark its performance on scRNA-seq and not on other sequencing technologies. Second, there is subjectivity in setting the SAHMI test thresholds, and while the same thresholds were used throughout this manuscript, adjusting them may alter the final results. Third, SAHMI requires either >3 barcodes or samples to run the correlation tests, limiting its use in very small studies. Fourth, while SAHMI removes human reads in the initial taxonomic classification step with Kraken and further removes human reads misclassified as microbes via the k-mer and cell-line tests, some users may wish to add an explicit subtractive host-alignment step. Fifth, in some samples that we analyzed in which only one microbe was expected, 1–4 species still passed all filters; these are misclassifications or contaminants that were more prevalent than in the cell-line data. Researchers may still wish to add additional filters, such as prevalence cutoffs, species intersections across samples or cohorts, or filters based on the strength of species associations with host gene expression. Sixth, we demonstrated that microbial-host barcode pairing could localize microorganisms to specific host cells. While in a previous analysis of the pancreatic cancer microbiome20 we found substantial bacteria-cell-type specific pairing, we did not find such substantial pairing in the benchmark datasets analyzed in this manuscript. This may be due to (1) true lack of cell-type specific enrichment, (2) dilution of observed enrichments due to substantial presence of microbes in inter-cellular spaces, or (3) ambient RNA contamination amongst barcodes. Future work can elucidate the extent to which SAHMI captures true cell-microbe enrichment.
Methods
SAHMI pipeline for microbiome detection and denoising from transcriptomic data
We developed SAHMI, a statistical pipeline for detection of microbes and analysis of host-microbiome interactions from scRNA-seq and other transcriptomic data. After sequencing and taxonomic classifications, SAHMI has two primary functions: (1) it identifies true microbial signal by running correlation analyses across barcodes and samples (k-mer correlation tests), and (2) it filters contaminants and false positives by comparing metagenomic counts to distributions of profiles in negative control samples (cell line quantile test). This enables systematic retrospective identification of microbes in host tissues. Downstream analysis of associations can be done at the sample level or at the level of individual cells for somatic cells and microbes that are tagged with the same cell barcodes.
Taxonomic classification: metagenomic classification of paired-end reads from bulk RNAseq, scRNA-seq, or spatial transcriptomic sequencing fastq files can be performed using a k-mer based mapper that identifies a taxonomic ID for each k-mer and read. While SAHMI can work with any k-mer-mapper, we reported the results for SAHMI with Kraken23,4, a popular and benchmarked tool which finds exact matches of candidate 35-mer genomic substrings to the lowest common ancestor of genomes in a reference metagenomic database. It is essential that all realistically possible genomes are included as mapping references at this stage, or that host mappable reads are excluded. The required outputs from this step are: a Kraken summary report with sample level metagenomic counts, a Kraken output file with read and k-mer level taxonomic classifications, and raw sequencing fastq files with taxonomic classification for each read, or the equivalent data.
Barcode level signal denoising (barcode k-mer correlation test): SAHMI first extracts microbiome reads from the raw data given their taxonomic IDs and removes reads that contain k-mers that map to the host genome. Next, for each taxon in the sample, SAHMI identifies the corresponding reads and removes reads with less than a default of 50% of the k-mers mapped directly to the taxon or to a parent taxon in its lineage. While analyzing scRNA-seq data, the cell barcode ID is used to identify reads originating from the same droplet. The number of total and unique k-mers mapping to the taxon or its lineage is then tabulated. For computational efficiency, a default of 1000 barcodes per taxon are randomly sampled. The Spearman correlation between the number of total and unique k-mers across barcodes for each taxon is computed. SAHMI reports the correlation and p-value and recommends removing taxa with non-significant correlations. This enables identification of true microbes in an individual sample.
Sample-level signal denoising (sample k-mer correlation test): the correlation analysis is also conducted across samples when possible. The Kraken report tabulates the total number of reads, minimizers (k-mers), and an estimate of unique k-mer counts for each taxon, or the equivalent data can be obtained from mappers. For each taxon, SAHMI correlates the number of k-mers vs. unique k-mers, reads vs. k-mers, and reads vs. unique k-mers across all samples in a study. True taxa are identified as those having significant positive Spearman correlation values and p-values for all three tests.
Identifying contaminants and false positives (cell line quantile test): these can be identified in the SAHMI workflow based on the widely observed pattern that contaminants appear at higher frequencies in low concentration or negative control samples1. We observed that this pattern also extends to false positive assignments. In the absence of experimentally matched negative controls, we provide a negative control resource comprised of microbiome profiles from 2,491 sterile cell experiments from around the world. For each taxon in a test sample, SAHMI compares the fraction of microbiome reads assigned to the taxon [i.e. taxon counts/sum(all bacterial, fungal, viral counts), in reads per million] to the microbiome fraction assigned to the taxon in all cell line experiments. Using the microbiome fraction comparison normalizes for experiments having a varying number of total sequencing reads or varying underlying contamination. SAHMI tests whether the taxon microbial fraction in the test sample is >99th percentile (by default, or >90th percentile of only non-zero cell-line counts are included) of the taxon’s microbiome fraction distribution in cell line data. Taxa whose counts fall below the cell line distribution’s prespecified quantile are identified as below the cell-line noise threshold. Users may choose how stringently to select the quantile threshold for testing.
Quantitation of microbes and creating the barcode-metagenome counts matrix: after identifying true taxa, reads assigned to those taxa are extracted and passed though a series of filters. ShortRead is used to remove low complexity reads (< 20 non-sequentially repeated nucleotides), low quality reads (PHRED score < 20), and PCR duplicates tagged with the same unique molecular identifier and cellular barcode. Non-sparse cellular barcodes can be selected by using an elbowplot of barcode rank vs. total reads, smoothed with a moving average of 25, and using a cutoff at a change in slope < 10−3, in a manner analogous to how cellular barcodes are typically selected in single-cell sequencing data (CellRanger (10x Genomics), Drop-seq Core Computational Protocol v2.0.0 (McCarroll laboratory)). Lastly, the full taxonomic classification of all resulting reads and the number of reads assigned to each clade are tabulated.
Assembling the negative control cell lines microbiome data
The Sequence Read Archive (SRA) was queried using the following search: ((((“public”[Access]) AND “rna seq”[Strategy]) AND “transcriptomic”[Source]) AND cell line) AND “Homo sapiens”[orgn:__txid9606] to identify sequencing runs with human cell lines. This resulted in 52,397 sample entries. We then selected for samples with “library selection = cDNA OR library selection = PolyA”, and we removed experiments with mouse strain information, experiments involving infection, runs without a submitting center name, and runs with “cell line” designated as “none”. From each remaining submitting center, we randomly selected 5 runs. We downloaded raw RNA-seq fastq files for these samples and profiled their microbiome using Kraken23,4. Samples were retained in the data base only if >90% of the reads were classified as human (taxid=9606). We then manually checked each sample’s metadata to ensure they did not involve infection or stimulation with microbial agents. The resulting samples and their metadata are tabulated in Supplementary Data 3 and cell line microbiome data are available in Supplementary Data 4.
True positive datasets selection and metagenomic mapping
We analyzed scRNA-seq data from patient samples with the following clinically verified infections: Mycobacterium leprae (skin)6, Helicobacter pylori (stomach)7, and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2; bronchoalveolar lavage fluid)5. In addition data were analyzed from scRNA-seq experiments for tissues in which the following pathogens were experimentally introduced: Candida albicans [Peripheral Blood Mononuclear Cells (PBMCs)]9, Salmonella enterica (PBMCs)8, Mycobacterium tuberculosis (lung)10, human alpha herpesvirus 1 (HSV-1) (PBMCs)12, and human immunodeficiency virus 1 (HIV-1) (PBMCs)11.
Taxonomic classification was done using Kraken23,4 using a combined reference database including the human genome and all bacterial, fungal, and viral reference genomes recorded in RefSeq as of April 2022. Taxa were only included if they were detected in >5 samples. Cutoffs of r>0 and p<0.01 were used for all k-mer correlation tests. The cell-line quantile test cutoff for all analyses was set to reads per million > 99th percentile [or 90th percentile of nonzero entries] in the cell line dataset. Microbial read alignment was done for a subset of samples using the STAR14 RNA-seq aligner with the following parameters: alignIntronMax=1 and outFilterScoreMinOverLread=0.05. Alignment of reads to the human genome version hg19 using STAR was done using default parameters. Full-length alignment of microbial reads to their respective genomes was done using BLAST15 with default parameters and using the nt reference genomes. The “bitscore” output from BLAST was used to compare mapping scores for taxa that passed or failed the k-mer correlation tests.
scRNA-seq data processing
For the clinical datasets, scRNA-seq data were processed using the standard Seurat23 pipeline. Cell types were identified by comparing cluster marker genes to the PanglaoDB24 reference database. Cell-types in the spatial transcriptomic map of a lesion from a patient with lepromatous leprosy were identified as follows. Marker genes for each cell-type were identified from scRNA-seq data of the lepromatous leprosy lesions using the FindAllMarkers function with logfc.threshold=1 and min.pct.diff=0.25. Cell-type annotations were taken from the original study. Cell-type scores for each spatial spot were computed as the mean expression of cell-type marker genes, and these were scaled per cell-type. More specifically, each spot had a score for each cell-type, which was the mean gene expression of marker genes for that cell-type. These scores were z-score scaled per cell-type across all spatial spots. The predominant cell-type at a spatial spot was identified as the cell-type with the highest score at that spot.
Body site enrichment analysis
The mBodyMap19 was used to calculate body site enrichment scores for the taxa that passed the SAHMI tests from the leprosy skin study6. The enrichment score reflects both the relative abundance of the SAHMI-reported taxa at each body site and the number of SAHMI-reported taxa detected at each body site. For each body site, the enrichment score was calculated as follows:
To compute enrichment score significance, the scores were calculated 100 times each for subsampled and label-shuffled data. For subsampled data, 80% of samples from each body site were randomly sampled and the enrichment score was calculated. These were compared to 100 calculations of enrichment scores for all samples with randomized body site labels. The subsampled and label-shuffled enrichment scores for each body site were compared using Wilcoxon rank sum tests.
Statistical analyses
All statistical analyses were performed using R version 3.6.1. All p-values were corrected for false-discovery rate (fdr) for multiple hypotheses using the p.adjust function with method= “fdr”, unless otherwise stated. The tidyverse package was used heavily for data processing (https://www.tidyverse.org/). The ggpubr package (https://github.com/kassambara/ggpubr) was used to compare group means with nonparametric tests and to perform multiple hypothesis correction for statistics that are noted in the figures. P-values reported as <2.2×10−16 result from reaching the calculation limit for the native R statistical test functions and indicate values below this number, not a range of values. Barcode level analyses were done for the benchmark studies in which reads had identifiable cell barcodes, unique molecular identifiers, and poly-A sequences. Comparisons to the human cell line negative control dataset were done for the benchmark studies that used human cells.
Supplementary Material
Acknowledgments
We acknowledge the Office of Advanced Research Computing (OARC) at Rutgers, The State University of New Jersey for providing access to the Amarel cluster URL: https://it.rutgers.edu/oarc. We also acknowledge grant support from National Institutes of Health grant R01GM129066, R21CA248122, R35GM149224 (SD), National Institutes of Health grant U01AI22285 (MJB), Sergei Zlinkoff Foundation (MJB), Canadian Institute for Advanced Research (MJB), National Institutes of Health, National Center for Advancing Translational Sciences, Rutgers Clinical and Translational Science Award TL1TR003019 (BG).
Footnotes
Code Availability
The SAHMI pipeline is available on our Github: https://github.com/sjdlabgroup/SAHMI and at Zenodo: https://doi.org/10.5281/zenodo.7017103.25
Competing interests
MJB declares that he serves on the Scientific Advisory Board of Micronoma, Inc. BG and SD have jointly filed PCT patent applications PCT/US2022/025829 and PCT/US2022/025832.
Data Availability
The cell lines microbiome negative control dataset is available in Supplementary Data 4. Other data are available on our Github: https://github.com/sjdlabgroup/SAHMI and at Zenodo: https://doi.org/10.5281/zenodo.7017103.25 The following infection datasets were analyzed in this manuscript: COVID-19 (GSE145926), M. leprae (GSE151528 and GSE167889), gastric samples (GSE134520), Salmonella (GSE79363), Candida (GSE111731), M. tuberculosis (GSE167232), HIV (GSE111727). The following human reference genome was used: hg19 (PRJNA31257). Source Data for Figures 1–2 are available with this manuscript.
References
- 1.Davis NM, Proctor DM, Holmes SP, Relman DA & Callahan BJ Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data. Microbiome 6, 226 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Poore GD et al. Microbiome analyses of blood and tissues suggest cancer diagnostic approach. Nature 579, 567–574 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
- 3.Breitwieser FP, Baker DN & Salzberg SL KrakenUniq: confident and fast metagenomics classification using unique k-mer counts. Genome Biol. 19, 198 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Wood DE, Lu J & Langmead B Improved metagenomic analysis with Kraken 2. Genome Biol. 20, 257 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Liao M et al. Single-cell landscape of bronchoalveolar immune cells in patients with COVID-19. Nat. Med 26, 842–844 (2020). [DOI] [PubMed] [Google Scholar]
- 6.Ma F et al. The cellular architecture of the antimicrobial response network in human leprosy granulomas. Nat. Immunol 22, 839–850 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Zhang P et al. Dissecting the Single-Cell Transcriptome Network Underlying Gastric Premalignant Lesions and Early Gastric Cancer. Cell Rep. 27, 1934–1947.e5 (2019). [DOI] [PubMed] [Google Scholar]
- 8.Saliba AE et al. Single-cell RNA-seq ties macrophage polarization to growth rate of intracellular Salmonella. Nat. Microbiol 2, 1–8 (2016). [DOI] [PubMed] [Google Scholar]
- 9.Muñoz JF et al. Coordinated host-pathogen transcriptional dynamics revealed using sorted subpopulations and single macrophages infected with Candida albicans. Nat. Commun 10, (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Pisu D et al. Single cell analysis of M. tuberculosis phenotype and macrophage lineages in the infected lung. J. Exp. Med 218, e20210615 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Golumbeanu M et al. Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity in Latent and Reactivated HIV-Infected Cells. Cell Rep. 23, 942–950 (2018). [DOI] [PubMed] [Google Scholar]
- 12.Wyler E et al. Single-cell RNA-sequencing of herpes simplex virus 1-infected cells connects NRF2 activation to an antiviral program. Nat. Commun 10, 4878 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Sims D, Sudbery I, Ilott NE, Heger A & Ponting CP Sequencing depth and coverage: key considerations in genomic analyses. Nat. Rev. Genet 15, 121–132 (2014). [DOI] [PubMed] [Google Scholar]
- 14.Dobin A et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Camacho C et al. BLAST+: architecture and applications. BMC Bioinformatics 10, 421 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Salter SJ et al. Reagent and laboratory contamination can critically impact sequence-based microbiome analyses. BMC Biol. 12, 87 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Zhang P et al. Dissecting the Single-Cell Transcriptome Network Underlying Gastric Premalignant Lesions and Early Gastric Cancer. Cell Rep. 27, 1934–1947.e5 (2019). [DOI] [PubMed] [Google Scholar]
- 18.Byrd AL, Belkaid Y & Segre JA The human skin microbiome. Nat. Rev. Microbiol 16, 143–155 (2018). [DOI] [PubMed] [Google Scholar]
- 19.Jin H et al. mBodyMap: a curated database for microbes across human body and their associations with health and diseases. Nucleic Acids Res. 50, D808–D816 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Ghaddar B et al. Tumor microbiome links cellular programs and immunity in pancreatic cancer. Cancer Cell 40, 1240–1253.e5 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Riquelme E et al. Tumor Microbiome Diversity and Composition Influence Pancreatic Cancer Outcomes. Cell 178, 795–806.e12 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Jia Y et al. Sequencing introduced false positive rare taxa lead to biased microbial community diversity, assembly, and interaction interpretation in amplicon studies. Environ. Microbiome 17, 43 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Stuart T et al. Comprehensive Integration of Single-Cell Data. Cell 177, 1888–1902.e21 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Franzén O, Gan L-M & Björkegren JLM PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data. Database 2019, baz046 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.sjdlabgroup. (2022). sjdlabgroup/SAHMI: SAHMI v1.0 (v1.0) Zenodo. 10.5281/zenodo.7017103 [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The cell lines microbiome negative control dataset is available in Supplementary Data 4. Other data are available on our Github: https://github.com/sjdlabgroup/SAHMI and at Zenodo: https://doi.org/10.5281/zenodo.7017103.25 The following infection datasets were analyzed in this manuscript: COVID-19 (GSE145926), M. leprae (GSE151528 and GSE167889), gastric samples (GSE134520), Salmonella (GSE79363), Candida (GSE111731), M. tuberculosis (GSE167232), HIV (GSE111727). The following human reference genome was used: hg19 (PRJNA31257). Source Data for Figures 1–2 are available with this manuscript.
