Skip to main content
Communications Biology logoLink to Communications Biology
. 2026 Jun 24;9:864. doi: 10.1038/s42003-026-10382-x

ScQCenrich enables multi-metric quality control for single-cell RNA sequencing

Yuanyuan Liu 1,2,✉,#, Cheng Yang 3,#, Chenghui Wang 4, Mingwang Zhang 5, Kai Luo 6, Lihua Wu 6, Xiufeng Xie 6
PMCID: PMC13315592  PMID: 42342867

Abstract

Single-cell RNA sequencing (scRNA-seq) is highly susceptible to dissociation stress, partial lysis, and nuclear–cytoplasmic imbalance, yet quality control still often relies on fixed thresholds for mitochondrial RNA, gene counts, and UMIs. Here we present scQCenrich, an interpretable multi-metric QC framework for post-cell-calling whole-cell scRNA-seq that integrates canonical metrics with intronic fraction, MALAT1 enrichment, dissociation-stress features, and optional splice-aware information. Across mouse brain, mouse heart and lung cancer datasets, scQCenrich reduces over-filtering relative to conventional and model-based comparators while preserving coherent neuronal, erythroid, cardiomyocyte and malignant-cell populations. In high-quality peripheral blood mononuclear cell data, the method remains conservative. Automated reports link quality-control calls to cluster-level metrics, marker genes and functional enrichment. scQCenrich therefore provides a transparent and reproducible framework for quality-control decisions in whole-cell single-cell RNA sequencing analyses.

Subject terms: Data processing, RNA sequencing


scQCenrich integrates canonical, nuclear-enrichment, and stress-related metrics to improve quality control in single-cell RNA sequencing while preserving biologically coherent cell populations.

Introduction

Single-cell RNA sequencing (scRNA-seq) has transformed cellular genomics by enabling transcriptome-wide profiling of heterogeneous tissues at single-cell resolution. However, the interval between tissue dissociation and droplet capture is intrinsically vulnerable to technical distortion. Warm enzymatic dissociation can induce acute stress programs, dying or ruptured cells can release ambient RNA into suspension, and partial lysis can selectively deplete cytoplasmic transcripts while leaving behind a nucleus-biased library composition1–4. Although ambient RNA contamination is mechanistically distinct and often addressed by dedicated decontamination methods, it arises from the same damaged-cell landscape and reinforces the need for rigorous post–cell-calling quality control (QC)3. Because these artifacts can masquerade as biology, QC is not merely a preprocessing convenience but a major determinant of which cell states remain available for downstream interpretation.

Most routine scRNA-seq QC workflows are anchored to the canonical triad of total unique molecular identifiers (UMIs), number of detected genes, and mitochondrial RNA fraction (mt%)5–9. These metrics are biologically informative, but each is only an imperfect proxy for cell integrity. Total UMIs reflect overall RNA capture and library size, such that severely damaged or inefficiently captured cells typically yield fewer molecules. The number of detected genes reflects transcriptome complexity and often declines when RNA is degraded or incompletely recovered, although naturally small or transcriptionally quiescent cells may also exhibit low complexity. mt% measures the relative abundance of mitochondrially encoded transcripts and often rises when cytoplasmic RNA is lost; however, it can also be intrinsically elevated in metabolically active or metabolically rewired cell states8,9. Consequently, hard filtering on these axes can remove bona fide biology, whereas reliance on these axes alone may still fail to identify subtly compromised profiles.

More adaptive frameworks have improved on rigid cutoffs, but an important gap remains. scater formalizes data-driven QC through robust outlier detection across per-cell metrics, and miQC advances further by using a probabilistic framework that jointly models detected genes and mitochondrial RNA fraction to infer low-quality cells6,7. The key contribution of scQCenrich is motivated by the observation that damage in scRNA-seq is multi-axial: partial lysis generates nuclear–cytoplasmic imbalance that is not fully captured by mt% or library complexity, and dissociation stress can distort transcriptional programs even when conventional QC metrics remain borderline1,2,10,11. A QC framework that relies primarily on canonical axes may therefore conflate metabolically unusual but intact cells with genuinely damaged libraries.

These considerations motivate the feature set used in scQCenrich. When cytoplasmic RNA is lost, nuclear and pre-mRNA signals become relatively enriched, making intronic fraction an informative proxy for nuclear enrichment and damaged droplets10,11. MALAT1, a nuclear-retained lncRNA, provides an additional orthogonal indicator of nucleus-biased libraries12. Stress-response signatures induced during tissue processing capture ex vivo perturbation, whereas the canonical metrics continue to summarize global RNA abundance, complexity, and mitochondrial enrichment1,2,5. Building on this logic, scQCenrich integrates UMIs, detected genes, mt%, intronic fraction, MALAT1 enrichment, and dissociation-stress signals within an interpretable multi-metric framework to estimate low-quality risk while allowing context-aware rescue of borderline cells when orthogonal evidence supports transcriptome integrity. Its main advance is the explicit incorporation of biologically motivated damage signals such as intronic fraction and MALAT1 enrichment that extend beyond mt%-gene models.

Critically, scQCenrich performs context-aware rescue of borderline cells when orthogonal evidence supports intact transcriptomes and when cell-type identity remains biologically coherent. Finally, it generates an automated HTML report that couples QC diagnostics with cell-type and cluster-level functional validation, turning QC from a black box into a reviewable and reproducible decision process.

Results

scQCenrich preserves biologically coherent, metabolically active cells lost to conventional QC

In a 10x Genomics mouse brain dataset, scQCenrich removed 1185 cells, substantially fewer than Seurat (2520), scater (1801), miQC (2770), and DropletQC (1968). Cells uniquely retained by scQCenrich tended to show borderline-elevated mitochondrial RNA fractions but normal intronic fractions, consistent with intact transcriptomes rather than damaged libraries (Figs. 1, 2A, S1A, S1B). Differences between methods were concentrated in neuronal, astrocytic, and erythrocytes (Figs. 2B–D, S1C, S1D). These identifications were supported by canonical marker expression and concordant Gene Ontology Biological Process (GO BP) enrichment, supporting that they represented bona fide biological cell states rather than technical artifacts (Table 1 and Figs. 2E, S1E, S1F, S2A, S2B). Notably, both scQCenrich and scQCenrich without splice information retained more functional neurons than Seurat, scater, or miQC, and QC metrics and GO BP analysis of rescued neurons further supported their biological validity (Figs. 2B, S2C–F). Although DropletQC retained the largest fraction of neurons, this came at the expense of disproportionate loss of functional astrocytes (Figs. 2C, S2G). Similarly, Seurat, scater, and DropletQC depleted functional erythrocytes that were preserved by both miQC and scQCenrich (Fig. 2D). Collectively, these findings indicate that scQCenrich reduces over-filtering of metabolically active yet biologically coherent cell states while maintaining separation from damaged profiles.

Fig. 1. scQCenrich integrates orthogonal QC signals to make context-aware filtering decisions.

Fig. 1

Schematic of the framework quantifying canonical metrics, including UMIs, detected genes and mitochondrial RNA fraction (mt%), together with nuclear-enrichment proxies, including intronic fraction and MALAT1 enrichment, and a compact dissociation-stress score. These features are integrated to prioritize removal of damaged or stressed profiles while retaining borderline but biologically credible cells when orthogonal evidence supports transcriptome integrity.

Fig. 2. scQCenrich preserves biologically coherent brain cell states that are over-filtered by conventional QC.

Fig. 2

A Shows mitochondrial RNA fraction (mt%) versus intronic fraction in the mouse brain dataset; blue points indicate cells retained by scQCenrich but removed by scater. B–D Show the proportions of neurons, astrocytes and erythrocytes retained by each QC method. E Shows GO BP enrichment for the neuron cluster (n = 1261 cells); bars represent enriched GO BP terms with FDR-adjusted q < 0.05. Abbreviation: GO BP, Gene Ontology Biological Process.

Table 1.

Gene markers

Stress markers FOS,FOSB,JUN,JUNB,JUND,ATF3,EGR1,DUSP1,DUSP2,IER2,IER3,BTG1,BTG2,ZFP36,ZFP36L1,HSPA8,HSP90AB1,HSPA1A,HSPA1B,HSPB1,HSP90AA1,HSP90B1,HSPH1,HSPE1,HSPD1,XBP1,DDIT3
Cardiomyocyte score markers Ttn,Myh6,Myh7,Tnnt2,Tnni3,Actn2,Mybpc3,Ldb3,Csrp3,Tcap,Des,Ryr2,Casq2,Jph2,Pln,Cacna1c,Slc8a1,Kcnq1,Kcne1,Cav3,Popdc2,Bves,Popdc3
Neuron markers Snap25,Rbfox3,Stmn2,Syn1,Syp,Map2,Dcx,Tubb3,Gap43,Calb1
erythroid markers Hba-a1,Hba-a2,Hbb-bt,Hbb-bs,Alas2,Fech,Gypa,Slc4a1,Slc25a37,Tspo2,Rhd

scQCenrich safeguards functional cardiomyocytes disproportionately removed by standard filters

We next evaluated performance in mouse heart, a tissue in which bona fide cardiomyocytes frequently occupy conventional QC failure zones because of their high metabolic demand. In this dataset, scQCenrich removed 1194 cells, again fewer than Seurat (3170), scater (1612), miQC (2441), and DropletQC (2147) (Figs. 3A, S3A, S3B). As in brain, cells uniquely retained by scQCenrich typically showed modestly elevated mitochondrial fractions without concomitant increases in intronic fraction, consistent with intact rather than damaged transcriptomes. Inspection of the retained and removed populations showed that Seurat, scater, and miQC each removed a substantial portion of the functional cardiomyocyte compartment, whereas scQCenrich preserved most of these cells (Figs. 3B, 3C, S3C). Cardiomyocyte identity was supported by strong expression of canonical markers and by GO BP enrichment for cardiomyocyte-associated biological programs, supporting that the cluster represented a functional cardiomyocyte population (Table 1 and Figs. 3D, S3D, S3E). Moreover, QC metrics and enrichment analyses of cells rescued by scQCenrich from Seurat, scater, and miQC consistently supported a cardiomyocyte functional identity (Fig. S3F–I). Of note, DropletQC again nearly eliminated functional erythrocytes in this dataset (Fig. S3J–L). Together, these results show that scQCenrich better preserves biologically meaningful cardiomyocytes.

Fig. 3. scQCenrich preserves functional cardiomyocytes that are preferentially lost under conventional QC.

Fig. 3

A Shows mt% versus intronic fraction in the mouse heart dataset; blue points indicate cells retained by scQCenrich but removed by Seurat. B Shows the cardiomyocyte compartment on UMAP in green and the corresponding overlay of cells removed by each QC method, with red points indicating removed cells. C Shows the proportion of cardiomyocytes retained by each QC method. D Shows GO BP enrichment for the cardiomyocyte cluster (n = 190 cells); bars represent enriched GO BP terms with FDR-adjusted q < 0.05.

scQCenrich retains malignant cells of interest while removing damaged profiles

We further tested scQCenrich in a cancer dataset, where the key challenge was to preserve malignant cells of interest while removing severely compromised profiles. Cell identities were assigned on the basis of marker expression, and damaged clusters were identified by extremely high mitochondrial transcript fractions ( > 85%) together with GO BP enrichment overwhelmingly dominated by mitochondrial processes (Fig. S4A–C). GO BP analysis confirmed that the malignant cluster was enriched for cancer-associated functional programs (Fig. 4B). Across methods, performance diverged substantially in the trade-off between malignant-cell retention and damaged-cell removal (Figs. 4C, 4D, S4D). Both scQCenrich and scQCenrich without splice information rescued malignant cells that were removed by Seurat, scater, and miQC, and enrichment analysis of these rescued cells supported their functional malignant identity (Fig. S4E–G). By contrast, DropletQC preserved the highest proportion of malignant cells but failed to eliminate damaged cells in this dataset (Fig. 4C, D).

Fig. 4. scQCenrich balances malignant-cell retention with removal of damaged profiles in lung cancer.

Fig. 4

A Shows the lung cancer dataset on UMAP colored by cell annotation. B Shows GO BP enrichment for the malignant-cell cluster (n = 1224 cells); asterisks mark cancer-relevant terms, and enriched terms are shown at FDR-adjusted q < 0.05. C Shows the proportion of malignant cells retained by each QC method. D Shows the proportion of damaged cells removed by each QC method. E Shows the proportion of malignant cells retained by each QC method in the external validation dataset.

To evaluate performance in a more practical multi-sample setting and enable comparison with SampleQC, we further benchmarked scQCenrich on a published lung adenocarcinoma dataset lacking splicing information, in which malignant cells had been annotated by the original authors (Fig. S5A)13. Even without splice-aware input, scQCenrich retained a larger fraction of malignant cells of interest than Seurat, scater, miQC, and SampleQC (Fig. 4E, Fig. S5B–F). Together, these results indicate that scQCenrich achieves a favorable balance between preserving biologically relevant malignant cells and excluding clearly compromised profiles.

Comparable performance on high-quality datasets

We also tested scQCenrich on a high-quality peripheral blood mononuclear cell (PBMC) dataset to ensure that it was not overly stringent. In this case, all modern QC tools performed similarly: scQCenrich, scater, and miQC each removed only a small, scattered set of outlier cells, whereas Seurat’s rigid thresholds culled a substantially larger region of cells (Figure S6A). The overlap of cells removed by scQCenrich and by the other probabilistic methods was large (Figure S6B). Quantitatively, scQCenrich discarded 416 cells versus 437 (miQC) and 555 (scater) in this dataset (versus 1484 by Seurat). Seurat disproportionately discarded more monocytes and T cells than other QC methods (Figs. S6C, S7A, S7B). QC metrics and enrichment analysis verified that additional cells removed by Seurat are functional monocytes and T cells (Fig. S7C–F). These results confirm that on well-behaved data scQCenrich acts conservatively: it performs on par with current probabilistic QC methods and does not indiscriminately cull cells. In other words, our extra nuclear- and stress-based checks do not reduce sensitivity on high-quality data but rather prevent undue over-filtering.

Automated reporting facilitates detection of low-quality clusters

A major advantage of scQCenrich is its comprehensive QC report. Cells were provisionally annotated using cluster-level DEGs to facilitate review in the HTML report. In an in-house low-quality dataset, the report quickly flagged clusters 0 and 1 as having heavy cell removal (Fig. 5A–B). Crucially, cells in these clusters had predicted cell-type labels that did not match the most dominant GO BP enrichment for those clusters (Fig. 5C). For fibroblast and keratinocyte annotations, the dominant GO BP terms were expected to align with ECM organization and skin development, respectively; deviation from this pattern prompted further review. By surfacing such discrepancies between annotation and functional profiles, the report guided targeted investigation of suspect clusters. In summary, the automated report transforms QC into an interactive diagnostics tool: it makes quality issues apparent at the cluster level, facilitating tuning of parameters and increasing confidence in the final cell set

Fig. 5. Automated cluster-level QC reporting reveals low-quality structure and annotation-function mismatches.

Fig. 5

A Shows UMAP positions of all cells in the low-quality dataset, with cells removed by scQCenrich shown in red. B Shows UMAP cluster identities. C Shows per-cluster GO BP profiles highlighting discordance between predicted cell-type labels and dominant biological processes, thereby guiding targeted review and parameter tuning.

Discussion

Quality control in droplet-based scRNA-seq is challenging because low-quality profiles do not arise from a single mechanism1,2,5. Tissue dissociation can induce acute stress programs, dying or ruptured cells can release ambient RNA, and partial lysis can preferentially deplete cytoplasmic transcripts while leaving behind a nucleus-biased library composition9. These processes distort transcriptome structure in different ways, yet routine filtering often compresses this complexity into a small set of canonical metrics. Among these, mitochondrial RNA fraction is particularly problematic: although it often increases when cytoplasmic RNA is lost or cells are stressed, it can also be intrinsically elevated in metabolically active or metabolically rewired states. Consequently, mt%-centered QC can remove biologically meaningful cells together with truly compromised profiles1,2,5,9.

scQCenrich provided the most consistent balance in the datasets examined. Across brain, heart, cancer, and high-quality PBMC datasets, it consistently reduced over-filtering of biologically coherent populations while remaining conservative on well-behaved data. This pattern was most evident in metabolically demanding or transcriptionally atypical compartments, including neurons, erythroid cells, cardiomyocytes, and malignant cells, where reliance on conventional QC axes alone can be especially punitive. These results support scQCenrich as a robust post-cell-calling QC framework for droplet-based whole-cell scRNA-seq5,9.

The Seurat comparison should be interpreted as a benchmark against a conventional fixed-cutoff workflow rather than an optimized per-dataset configuration. We selected prespecified literature-informed thresholds to represent common practice, while recognizing that suitable hard cutoffs vary across tissues and study designs8,14. Our results therefore do not argue that all Seurat parameterizations are uniformly suboptimal; instead, they highlight the sensitivity of fixed-threshold QC to parameter choice and the relative stability of scQCenrich across diverse datasets. The practical advantage of scQCenrich is thus reduced dependence on manual threshold tuning, especially in datasets containing metabolically active or otherwise atypical but biologically valid cell states.

These findings also clarify how scQCenrich relates to prior QC methods. miQC models mitochondrial fraction jointly with detected genes in a probabilistic framework, whereas scater provides a widely used QC utility workflow rather than an end-to-end probabilistic damage model6. DropletQC introduced nuclear fraction as an informative axis for distinguishing empty droplets, damaged cells, and intact cells in droplet-based scRNA-seq10. SampleQC addresses a different but complementary problem by modeling multivariate QC structure across samples and cell types in multi-sample studies15. In this context, scQCenrich is best understood as an interpretable post-cell-calling whole-cell scRNA-seq framework that combines canonical metrics, nuclear-enrichment signals, MALAT1, dissociation-stress features, and optional splice-aware information with a biologically reviewable rescue step.

A key strength of scQCenrich is that the incorporated signals are related but not redundant. Total counts and detected genes summarize RNA recovery and transcriptome complexity. Mitochondrial RNA fraction is sensitive to damage and stress, but biologically ambiguous. By contrast, intronic fraction and nuclear-retained transcripts such as MALAT1 more directly reflect nuclear-cytoplasmic imbalance expected after cytoplasmic leakage, while dissociation-stress programs capture ex vivo transcriptional perturbation that may occur even when counts and complexity remain within conventional ranges. Optional splice-aware features add further evidence when such data are available. The value of the framework therefore lies in joint interpretation across partially independent axes, rather than in treating any single metric as decisive.

The comparison with DropletQC is particularly informative in this regard. Both methods recognize that nuclear enrichment contains information not captured by mt%-gene-count space alone. However, our benchmarks suggest that nuclear-fraction-oriented damage signals alone do not fully resolve the tradeoff between sensitivity and specificity across all relevant cellular states. In the datasets examined here, DropletQC performed favorably in some compartments, such as neuronal retention, but less favorably in others, including astrocytes, erythroid cells, and damaged-cell exclusion in the cancer dataset. This pattern supports the rationale for combining nuclear-enrichment evidence with additional orthogonal features and biological review, rather than treating any single QC axis as sufficient on its own10. Importantly, scQCenrich does not require splice-aware matrices for routine use: splice-aware features improve resolution when available, but the framework remains applicable to standard expression matrices using canonical QC metrics, nuclear-enrichment signals, and stress-related features.

At the same time, the scope of the present benchmarks should be stated plainly. The current study evaluates scQCenrich on public droplet-based whole-cell scRNA-seq datasets that include both 10x Genomics reference datasets and an external published lung adenocarcinoma cohort. These results therefore support the utility of scQCenrich beyond a single-source benchmark set, including in settings where splice-aware information is unavailable. However, the present analyses do not establish performance across all assay types or study designs, particularly single-nucleus RNA-seq, where high intronic content and elevated nuclear-retained transcripts are expected features of intact nuclei rather than damage signatures. In such settings, assay-aware recalibration would be necessary, and nucleus-focused methods such as DIEM and QClus remain important comparators2,15–17.

Beyond retention alone, scQCenrich emphasizes interpretability. The automated HTML report links QC-feature diagnostics, low-dimensional embeddings, marker expression, and functional enrichment into a reviewable output that allows analysts to inspect why cells were removed, retained, or flagged as borderline. This is particularly useful in datasets containing low-quality substructure, where technical artifacts can resemble plausible biological clusters. By coupling QC decisions to biological coherence rather than leaving them as opaque binary calls, the framework turns QC from a black box into a transparent analytical decision layer. This interpretability is likely to be especially valuable in studies where downstream conclusions are sensitive to whether stressed, partially lysed, or metabolically unusual populations are retained or discarded.

In summary, scQCenrich is best viewed as an interpretable multi-metric QC framework for post-cell-calling whole-cell scRNA-seq. Within the assay scope tested here, it reduces over-filtering of biologically coherent cell states, remains conservative on high-quality datasets, and offers a more flexible and transparent alternative to mt%-centered or single-axis QC strategies.

Methods

scQCenrich workflow

scQCenrich is an R-based QC framework for scRNA-seq data stored in Seurat objects. The workflow calculates per-cell QC metrics, identifies low-quality cells using threshold-based or model-based detection, optionally performs cluster-level preliminary annotation, applies a coherence-based rescue step for biologically plausible borderline profiles and renders an HTML report. In the analysis reported here, scQCenrich was run as a post-cell-calling QC method on filtered Cell Ranger count matrices. Unless otherwise stated, the published benchmarks used run_qc_pipeline with method = “gmm”, qc_strength = “auto”, rescue_mode = “moderate”, assay = “RNA”, report_html=TRUE and enrichment_plots=TRUE. The wrapper default output directory is qc_outputs and the default report file is qc_report.html.

Preliminary annotation and clustering

Preliminary annotation provides biological context for report interpretation and rescue, but it is not used as a stand-alone removal criterion. For marker-based annotation, clusters are generated using the standard Seurat workflow. In the benchmark scripts, the Seurat object is normalized with NormalizeData, variable features are selected with FindVariableFeatures, data are scaled with ScaleData, principal components are computed with RunPCA, UMAP is computed using dimensions 1-20, neighbors are computed with FindNeighbors and clusters are called with FindClusters at resolution 0.4. Cluster labels can be assigned using PanglaoDB-derived marker sets, SingleR or ScType18,19.

QC metrics calculation

For each cell, scQCenrich computes total UMI counts, the number of detected genes, mitochondrial transcript fraction, MALAT1 fraction and a dissociation-stress score. The dissociation-stress score is the mean normalized expression of curated immediate-early and heat-shock stress genes (Table 1)1,2,4,12. When splice-aware data are supplied, spliced and unspliced count layers are aligned to the main Seurat object. scQCenrich then calculates spliced counts, unspliced counts, spliced fraction, unspliced fraction, unspliced-to-spliced ratio and intronic fraction, defined as unspliced counts divided by the sum of spliced and unspliced counts. If splice-aware layers are absent, intronic-fraction fields remain unavailable and MALAT1 two-tailed deviation is used as the nuclear-enrichment proxy. These features provide orthogonal information on nuclear enrichment and cell integrity beyond canonical QC metrics10,11.

Model-based low-quality detection

Model-based detection first embeds cells in QC feature space. The feature matrix includes pctMT, nFeature, stress_score and u2s_ratio when available. If intronic fraction is available, scQCenrich adds the absolute robust z-score of intronic fraction as a two-tailed nuclear-deviation axis and includes MALAT1 fraction as a secondary signal. If intronic fraction is unavailable but MALAT1 is available, the absolute robust z-score of MALAT1 fraction is used as the nuclear-deviation proxy. A PCA embedding with up to five components is computed. In Gaussian mixture mode, mclust::Mclust is fit to this embedding with G = 1:6 components. The covariance model is selected by the default Mclust model search, and the number of mixture components and covariance parameterization are selected by Bayesian information criterion. If Mclust fails because of a model-fitting error, the detector falls back to dbscan::hdbscan. In HDBSCAN mode, minPts is set adaptively to max(20, min(50, round(0.01 × N))), where N is the number of cells. No separate fixed density threshold was manually imposed. HDBSCAN density clusters, including the noise cluster, are then passed to the same cluster-level QC scoring procedure.

QC burden scoring and removal rules

QC burden is summarized at the QC-cluster level. For each QC-defined cluster, scQCenrich calculates the mean pctMT, mean negative nFeature, mean stress score, mean intronic or MALAT1 two-tailed deviation, mean unspliced-to-spliced ratio and mean MALAT1 fraction when available. These cluster summaries are converted to robust z-scores using median and median absolute-deviation scaling. The default cluster QC burden score is the sum of robust z-scores for mitochondrial burden, low gene complexity and stress, plus 1.5 times the intronic or MALAT1 nuclear-deviation z-score, 1.0 times the unspliced-to-spliced z-score and 0.5 times the MALAT1 z-score when true intronic fraction is already available. The 0.5 MALAT1 term is suppressed when MALAT1 is being used as the no-splice nuclear proxy to avoid double counting. Clusters with burden scores at or above the remove_quantile threshold, together with the single highest-scoring cluster, are initially marked for removal. In qc_strength = “auto” mode, remove_quantile is learned from a preliminary burden score based on pctMT, negative nFeature, stress and intronic or MALAT1 deviation, then constrained to the interval 0.85-0.97. Cell-level borderline calls are assigned to cells outside removal clusters with mitochondrial or low-gene z-scores above border_z or nuclear-deviation scores above intr_border_z. The border_z defaults are 1.2 for auto/default, 2.0 for lenient and 0.8 for strict; intr_border_z defaults to 2.0.

Default safeguards and final QC states

Several explicit safeguards are applied after the model-based call. Cells with pctMT greater than or equal to 50% are force removed by the extreme mitochondrial hard cap. Borderline cells are capped at 25% of the pre-rescue or final removed-cell count, whichever is larger; the least severe borderline cells are demoted to keep if this cap is exceeded. The final mutually exclusive QC states are kept, borderline and remove. Cells with keep or borderline states are retained for downstream analysis, whereas removed cells are excluded.

Coherence-based rescue

Coherence-based rescue is applied only to cells that were initially assigned to remove within eligible QC-defined clusters. The default minimum QC-cluster size for rescue is 50 cells. For each QC cluster, scQCenrich calculates cluster size, removed-cell fraction, median pctMT, median nFeature, median stress score, median intronic fraction, dominant-label coherence and the fraction of available metric gates passed. A cluster passes the mitochondrial gate when its median pctMT is no higher than the global median, the gene-complexity gate when its median nFeature is no lower than the global median and the stress gate when its median stress score is no higher than the global median. The metric pass fraction is the number of passed gates divided by the number of available gates. The rescue score is the mean of dominant-label coherence, metric pass fraction and 1 minus the removed-cell fraction. Rescue thresholds are 0.90 for strict mode, 0.70 for moderate mode and 0.50 for lenient mode; rescue is disabled in none mode. Intronic gating provides a hard veto when splice-aware data are available: a cluster is considered intronically healthy only when its median intronic fraction is below max(1.20 times the global median intronic fraction, 1.20 times 0.35, 0.65). In all modes, rescued cells are reassigned from remove to borderline rather than directly to keep. Cells removed by extreme mitochondrial rules are not rescued.

Datasets and benchmarking design

We benchmarked scQCenrich on five public droplet-based whole-cell scRNA-seq datasets selected to represent complementary QC scenarios: 10k E18 mouse brain20, 10k E18 mouse heart21, 10k human PBMCs22, a 10x Genomics non-small-cell lung cancer dataset23 and an external published multi-sample lung adenocarcinoma dataset13. Brain and heart datasets were used to test whether the method preserves metabolically active yet biologically coherent cell states that frequently occupy conventional QC failure zones. PBMCs provided a relatively high-quality reference to assess whether additional QC features introduce excess stringency. The cancer datasets provided disease settings in which malignant-cell retention had to be balanced against removal of severely damaged profiles. The external multi-sample dataset further allowed evaluation without splice-aware input and enabled comparison with SampleQC.

Splice-aware input generation and comparator methods

To incorporate RNA nuclear/cytoplasmic balance into QC, we computed spliced and unspliced count layers using the velocyto v0.17.16 counting pipeline, which parses aligned reads to generate loom files with spliced and unspliced partitions; logic settings followed velocyto recommendations for 10x Genomics data24. Loom files were imported into Seurat using SeuratWrappers v0.4.0. We compared scQCenrich with four benchmark baselines representing distinct categories of current QC practice. Seurat v5.3.0 and scater v1.37.0 were treated as QC frameworks/utilities rather than standalone statistical QC models6,18. For the Seurat baseline, we applied a prespecified fixed-threshold workflow intended to represent common Seurat-style filtering practice rather than an optimized dataset-specific configuration: minimum 200 detected genes per cell, maximum mitochondrial transcript fraction 10% and upper cap 10,000 detected genes8. These thresholds were not tuned separately for each dataset. scater was run using addPerCellQCMetrics followed by 3-median-absolute-deviation lower-tail outlier detection for library size and detected genes and higher-tail outlier detection for mitochondrial percentage. As model-based comparators, we included miQC v1.17.0, which jointly models mitochondrial fraction and detected genes, and DropletQC, which uses nuclear-fraction-based damage signals to distinguish damaged from intact droplets. We additionally evaluated SampleQC on the external published multi-sample lung adenocarcinoma dataset, as this method is specifically designed for multivariate QC across samples and cell types in multi-sample studies 15.

Automated reporting

scQCenrich automatically generates an HTML report summarizing QC outcomes, including low-dimensional visualizations, QC metric distributions and summaries of kept, borderline and removed cells. The report includes feature maps for total UMIs, detected genes, mt%, MALAT1 fraction, stress score, intronic fraction and splice-aware layers when present; cluster-level and cell-type-level marker plots; and GO BP over-representation analysis. Biological coherence is assessed using marker-gene and Gene Ontology enrichment analyses implemented with clusterProfiler25. The report is intended to make the basis for QC decisions reviewable rather than to provide an additional hidden filtering criterion.

Statistics and reproducibility

No animals or human participants were newly recruited for this computational study. All benchmark analyses used publicly available scRNA-seq datasets or previously generated data, and cells were analyzed as observations within each dataset. Cell numbers retained or removed by each method are reported as direct counts or proportions; no inferential statistical test was used to compare QC methods unless explicitly stated. GO BP enrichment analyses were performed with clusterProfiler, and terms were considered enriched at Benjamini-Hochberg false-discovery-rate-adjusted q < 0.05. For analyses involving random sampling or stochastic dimension reduction, the benchmarking script records the random seeds used. No data points were excluded except through the explicitly documented QC rules. Benchmarking scripts and package code are provided with the source code to support independent reproduction.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

42003_2026_10382_MOESM3_ESM.pdf (35.9KB, pdf)

Description of Additional Supplementary Files

Reporting Summary (21.2MB, xlsx)

Acknowledgements

The authors thank the developers and maintainers of the public single-cell datasets and open-source software used in this study.

Author contributions

Y.L. and C.Y. conceived the study. C.Y. developed scQCenrich and performed computational analyses. Y.L. supervised the study and interpreted results. M.Z., K.L., L.W. and X.X. contributed to study design, biological interpretation and manuscript revision. Y.L. and C.Y. wrote the manuscript with input from all authors. C.W. contributed to conceptualization, project administration and funding acquisition.

Peer review

Peer review information

Communications Biology thanks Josephine Yates, Isar Nassiri, and Tomàs Montserrat-Ayuso for their contribution to the peer review of this work. Primary Handling Editors: Massimo Andreatta and Laura Rodríguez Pérez. A peer review file is available.

Funding

This work was supported by the Young Scientists Fund of the National Natural Science Foundation of China (82301991).

Data availability

All datasets analyzed in this study are publicly available. The 10x Genomics 10k E18 mouse brain, 10k E18 mouse heart, 10k human PBMC and non-small-cell lung cancer datasets are available from the 10x Genomics datasets portal. The external lung adenocarcinoma dataset from Kim et al. is available as processed data from the NCBI Gene Expression Omnibus under accession GSE13190713. Source data underlying the graphs in the main and Supplementary Figs. are provided as Supplementary Data 1. All other analysis-derived data are available from the corresponding author on reasonable request.

Code availability

The scQCenrich source code and benchmarking scripts are available at https://github.com/lemonlyy755/scQCenrich10.5281/zenodo.2005079826. The benchmark reproduction script is provided in inst/benchmarking.R within the repository. The analyses described in this manuscript used R with Seurat v5.3.0, scater v1.37.0, miQC v1.17.0, velocyto v0.17.16, SeuratWrappers v0.4.0 and clusterProfiler v4.18.4.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yuanyuan Liu, Cheng Yang.

Supplementary information

The online version contains supplementary material available at 10.1038/s42003-026-10382-x.

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

42003_2026_10382_MOESM3_ESM.pdf (35.9KB, pdf)

Description of Additional Supplementary Files

Reporting Summary (21.2MB, xlsx)

Data Availability Statement

All datasets analyzed in this study are publicly available. The 10x Genomics 10k E18 mouse brain, 10k E18 mouse heart, 10k human PBMC and non-small-cell lung cancer datasets are available from the 10x Genomics datasets portal. The external lung adenocarcinoma dataset from Kim et al. is available as processed data from the NCBI Gene Expression Omnibus under accession GSE13190713. Source data underlying the graphs in the main and Supplementary Figs. are provided as Supplementary Data 1. All other analysis-derived data are available from the corresponding author on reasonable request.

The scQCenrich source code and benchmarking scripts are available at https://github.com/lemonlyy755/scQCenrich10.5281/zenodo.2005079826. The benchmark reproduction script is provided in inst/benchmarking.R within the repository. The analyses described in this manuscript used R with Seurat v5.3.0, scater v1.37.0, miQC v1.17.0, velocyto v0.17.16, SeuratWrappers v0.4.0 and clusterProfiler v4.18.4.


Articles from Communications Biology are provided here courtesy of Nature Publishing Group

RESOURCES