Skip to main content
The Journal of Molecular Diagnostics : JMD logoLink to The Journal of Molecular Diagnostics : JMD
. 2025 Jun 25;27(9):869–881. doi: 10.1016/j.jmoldx.2025.05.008

Development of an End-to-End Total RNA Sequencing Quality Control Framework for Blood-Based Biomarker Discovery

Cheryl L Sesler ∗, Guzel I Shaginurova ∗, Lukasz S Wylezinski ∗,†, Elena V Grigorenko ∗, Franklin R Cockerill III ∗,‡,§, Charles F Spurlock III ∗,†,¶,‖,∗
PMCID: PMC12489346  PMID: 40578551

Abstract

Next-generation RNA sequencing (RNA-seq) enables comprehensive transcriptomic profiling for disease characterization, biomarker discovery, and precision medicine. Despite its potential, RNA-seq has not yet been widely adopted for clinical applications, and a key barrier to its adoption is the variability introduced during processing and analysis. Quality controls (QCs) must be considered through all stages of biomarker discovery. This study describes a comprehensive QC framework for effective RNA-seq biomarker discovery. Multilayered quality metrics were established across preanalytical, analytical, and postanalytical processes. Total RNA-seq was performed by using RNA isolated from whole blood (PAXgene Blood RNA tubes). Bulk RNA controls were incorporated to monitor sequencing batches. This framework was applied to a catalog of prospectively collected or biobanked clinical specimens spanning multiple disease indications. Among all QCs, preanalytical metrics (specimen collection, RNA integrity, and genomic DNA contamination) exhibited the highest failure rates and resulted in the addition of a secondary DNase treatment, which reduced genomic DNA levels. The additional DNase treatment significantly lowered intergenic read alignment and provided sufficient RNA for downstream sequencing and analysis. This end-to-end QC framework for RNA-seq biomarker discovery was developed and implemented to enhance the confidence and reliability of results. To advance the clinical adoption of RNA-seq, developing and implementing standards will improve reliability, accelerate biomarker discovery, and facilitate its translation into clinically actionable diagnostics and therapeutics.


Next-generation RNA sequencing (RNA-seq) has emerged as a powerful tool in molecular diagnostics, enabling comprehensive transcriptomic profiling for disease characterization, biomarker discovery, and precision medicine applications.1, 2, 3, 4 By quantifying gene expression at an unprecedented scale, RNA-seq facilitates the identification of molecular signatures associated with disease pathophysiology, treatment response, and prognosis.5, 6, 7, 8, 9 However, despite its potential, RNA-seq has not yet been widely adopted for clinical applications. Unlike DNA sequencing, which provides a stable representation of the genome, RNA-seq presents unique technical and biological challenges that hinder clinical translation.2,7,10,11

A key barrier to the widespread clinical adoption of RNA-seq is the variability introduced across preanalytical, analytical, and postanalytical stages.12,13 RNA is inherently labile and susceptible to degradation, and its expression levels are influenced by intrinsic biological factors and external conditions, including specimen collection, handling, storage, and processing protocols.11,13, 14, 15 Technical variability in RNA isolation, library preparation, sequencing platforms, and bioinformatic pipelines further compounds the challenges of standardizing gene expression analyses.16,17 The difficulties associated with RNA biomarker discovery, especially the reproducibility of initial findings in independent studies, highlight the critical need for implementing strong quality control (QC) measures throughout the RNA-seq workflow.10,18

Although analytical QC tools are widely implemented to assess sequencing quality after sequencing is performed, these approaches do not sufficiently address preanalytical sources of variability that can significantly affect gene expression measurements.11, 12, 13,15,19 Furthermore, there is currently no standardized, end-to-end QC framework guiding the integration of RNA-seq into biomarker discovery and molecular diagnostics. The absence of such standardization limits reproducibility and poses challenges for regulatory acceptance, as no formal guidelines exist for RNA-seq–based clinical testing. A comprehensive, multilayered QC approach that spans the assessment of specimen collection, RNA integrity, sequencing quality, and post-sequencing validation is essential to improving the reliability of transcriptomic biomarker studies.20, 21, 22

The current study describes a systematic QC framework to enhance the reproducibility and robustness of RNA-seq–based biomarker discovery. This framework incorporates preanalytical, analytical, and postanalytical QC checkpoints, leveraging best practices to mitigate technical variability. By implementing a rigorous, end-to-end QC strategy, a reproducible and scalable workflow was established to enhance the reliability of RNA-seq for biomarker discovery and broader clinical adoption. Implementing rigorous QC measures at every stage, from specimen collection to data interpretation, aims to enhance RNA-seq efficiency, minimize technical variability, and promote clinical translation.

Materials and Methods

Clinical Specimens

Peripheral whole blood was collected according to the manufacturer’s recommendations into PAXgene Blood RNA tubes (PreAnalytiX GmbH, Hombrechtikon, Switzerland) from individuals enrolled in metabolic or neurodegenerative disease studies and from healthy control subjects (Supplemental Table S1). Specimens were frozen at –70°C or lower by the specimen provider and shipped on dry ice. Upon receipt, each frozen specimen tube was weighed and stored at –70°C or lower before processing. For prospective studies, repeat blood collections were requested for patients with underfilled blood specimens, which include specimens filled with less than the manufacturer’s recommended fill volume of 2.5 mL of blood.

Specimens were obtained from Quest Diagnostics (San Juan Capistrano, CA), Discovery Life Sciences (Huntsville, AL), and the Accelerated Cure Project (Waltham, MA). Relevant institutional review board approval and written informed consent were obtained at each site. Specimens were de-identified prior to receipt and use in this study.

RNA Controls

Bulk total RNA samples isolated from cervical (HeLa) and breast (MCF7) adenocarcinoma cell lines were obtained from BioChain Institute Inc. (Newark, CA). Total RNA was treated with DNase I, PCR verified as DNA free, diluted to 50 ng/μL, and aliquoted by the manufacturer. Single-use aliquots were shipped on dry ice and frozen at –70°C or lower before processing.

Whole-Blood Total RNA Isolation and Sequencing

Total RNA was isolated from PAXgene Blood RNA tubes on the QIAcube using the PAXgene Blood RNA Kit (Qiagen, Hilden, Germany), incorporating an initial DNase treatment per the manufacturer’s protocol. RNA and DNA concentrations were determined on the Qubit 4.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA). RNA integrity characteristics, including the RNA Integrity Number (RIN) and percentage of RNA fragments >200 nucleotides [DV200 (RNA)], were assessed by using the TapeStation 4200 or Bioanalyzer 2100 (Agilent Technologies, Santa Clara, CA). Secondary DNase treatment was performed by using the TURBO DNA-free Kit (Thermo Fisher Scientific) followed by RNA sample clean-up with the RNA MagClean DX kit (Aline Biosciences, Nashua, NH). RNA was subjected to random fragmentation and converted to double-stranded cDNA. Libraries were constructed with the NEBNext Ultra II Directional RNA Kit with rRNA depletion (New England Biolabs, Ipswich, MA). Globin depletion was performed with the QIAseq FastSelect –Globin Kit (Qiagen). Libraries were pooled and assessed on the D1000 ScreenTape System (Agilent), and 150-bp paired-end sequencing was performed on the NovaSeq 6000 (Illumina, San Diego, CA).

Sequence Analysis

Sequencing quality parameters, such as sequencing quality score, sequencing depth, and percentage of GC content, were assessed with FastQC version 0.11.9.23 Sequences were aligned to the human genome (GRCh38.104/hg38) using Spliced Transcripts Alignment to a Reference (STAR version 2.7.9a),24 and gene expression levels were quantified by using featureCounts version 2.0.0.25 Genomic origins of uniquely mapped reads were determined with Qualimap RNA-seq QC version 2.3.26,27 Genomic mapping was visualized with Integrative Genomics Viewer version 2.13.2.28

Batch Analysis

Batches were numbered chronologically (batch 1, oldest; batch 6, newest). Sequencing batch analysis was assessed with HeLa and MCF7 RNA controls. Quantified gene expression values for each QC sample were used for cluster analysis and principal component analysis (PCA). Unsupervised cluster analysis was performed with the Leiden algorithm29 and visualized through the Uniform Manifold Approximation and Projection dimensionality reduction technique.30 PCA was performed in R version 4.2.2 using the prcomp() function and visualized with the ggplot() function (ggplot2 package version 3.4.1).31,32

Comparative Analysis

Detailed pipelines for differential expression (DE) and machine learning (ML) analyses have been described previously.33 In brief, DE was performed with DESeq2 version 1.32.034; genes were retained if adjusted P < 0.05 (Benjamini-Hochberg false discovery rate35), log2 fold change ≥|1|, and normalized mean expression counts ≥10. For ML, genes were filtered by using three feature selection strategies (recursive feature elimination, highly variable gene binning, and the χ2 test).36, 37, 38 Models were trained by using a training (80%) and validation (20%) partitioning strategy. Mean feature-impact scores from the three best performing models per feature selection were calculated, and the top 200 features across all nine models were selected. Consensus non-negative matrix factorization was applied to normalized counts to generate consensus factors across multiple iterations.39 The optimal rank (k) was chosen by maximizing solution stability and minimizing Frobenius error. Resulting factors were embedded in a k-nearest neighbor graph, partitioned with Leiden clustering, and visualized with the Uniform Manifold Approximation and Projection technique. Genes with spectra scores >3 SDs above the mean within each Leiden community were carried forward. Functional enrichment of these gene sets was assessed by using the Genomic Regions Enrichment of Annotations Tool version 4.0.4.40

Statistical Analysis

Quality metrics were summarized with the number and percentage of samples that failed to meet the defined acceptance criteria. Pearson correlation analysis was performed using the cor.test() function in R version 4.2.2 or GraphPad Prism version 9.0.31 Two-tailed t-tests assuming equal variance were performed to assess the statistical significance of categorical variables.

Results

Multilayered QC Framework Increases Efficiency and Application Flexibility

Gene expression analysis relies on accurate clinical annotation and requires QC checkpoints at each process step to ensure accurate results and reliable interpretation, particularly for large RNA-seq data sets. This multilayered QC framework integrates established internal practices with previously validated best practices.20, 21, 22,41, 42, 43, 44 This framework applies to both prospectively collected and biobanked whole-blood specimens processed using PAXgene Blood RNA tubes (Supplemental Table S1) and incorporates QC measures across the preanalytical, analytical, and postanalytical phases of RNA biomarker discovery (Figure 1). In addition to evaluating the quality of individual samples, external RNA controls were included in standard processing to monitor sequencing batch consistency.

Figure 1.

Figure 1

Embedded quality control framework. Workflow diagram of preanalytical, analytical, and postanalytical quality control (QC) stages assessed throughout the RNA sequencing biomarker discovery platform. External RNA controls are integrated for batch and analytical evaluation. Orange QC checkpoints indicate assessment of individual samples, and gold QC checkpoints indicate batch assessment.

Preanalytical QC assessed specimen collection, RNA isolation, DNase treatment, and library preparation, ensuring RNA integrity before sequencing. Analytical QC focused on sequencing data quality and downstream computational analyses, whereas postanalytical QC included result verification, pathway enrichment analysis, and literature review. Both sample- and batch-level QC metrics enabled the early detection of low-quality samples, batch inconsistencies, or process errors, reducing resource waste and preventing unreliable data from affecting interpretation.

The framework and acceptance criteria were developed using whole-blood specimens from patients with neurologic and metabolic diseases, as well as from healthy control subjects. Special consideration was given to specimen handling, storage conditions, and disease-related variability to account for both technical and biological factors. The framework is now being applied in ongoing biomarker discovery studies, supporting efforts to improve RNA-seq reproducibility and clinical utility.

Preanalytical Quality Assessment

The primary objective of preanalytical QC steps is to assess sample quality and suitability for RNA-seq. These preanalytical procedures evaluate whole-blood specimen volume and RNA integrity to ensure successful downstream processing. By identifying samples with insufficient RNA concentration or compromised integrity early, these assessments improve efficiency, reduce costs, and prevent unreliable data from affecting downstream analyses. Only samples meeting all predefined acceptance criteria are used in biomarker discovery efforts.

Acceptance criteria for preanalytical QC (Table 1) were established based on empirical testing, manufacturer guidelines, and published literature.15,45, 46, 47, 48, 49 Key QC metrics include specimen weight, RNA concentration, RNA integrity (RIN and DV200), and residual genomic DNA (gDNA) contamination. Proper specimen collection, handling, and storage are essential to obtaining high-quality RNA for sequencing.

Table 1.

Summary of Preanalytical QC Metrics Assessed for PAXgene Blood RNA Specimens and External RNA Controls

QC metrics Acceptance criteria Clinical specimens
RNA controls
No. assessed No. failed (%) No. assessed No. failed (%)
Specimen collection
 Specimen weight ≥18.0 g 788∗ 29 (3.7) NA NA
RNA isolation and initial DNase treatment
 RNA concentration ≥25 ng/μL 297 4 (1.4) 12 0 (0)
 RIN value ≥6.0 297 12 (4.0) 12 0 (0)
 DV200 (RNA) ≥70% 272† 5 (1.8) 10 0 (0)
 DNA concentration ≤10% RNA concentration 234‡ 139 (59.4) 8 5 (62.5)

DV200 (RNA), percentage of RNA fragments >200 nucleotides; QC, quality control; RIN, RNA Integrity Number.

∗

A total of 297 specimens meeting acceptance criteria have been processed, and 462 specimens have been banked for future processing.

†

DV200 (RNA) was not measured on one batch (25 clinical specimens and two RNA controls).

‡

DNA concentration was spot checked across the first two sequencing batches (not measured for 63 clinical specimens and four RNA controls).

Upon receipt, all specimens were weighed to confirm adequate blood volume for RNA isolation. Underfilled PAXgene Blood RNA tubes yield insufficient RNA, making them unsuitable for processing. RNA integrity was assessed by using RIN and the DV200. RNA purity was evaluated by measuring residual gDNA, which can artificially inflate transcript counts in RNA-seq data.50 Although minimizing residual gDNA is ideal, no universally accepted threshold exists for total RNA-seq experiments.

Among all QC metrics, specimen weight, RINs, and gDNA contamination had the highest failure rates (Table 1). Failure rates were <5% for all metrics except gDNA contamination, which exceeded the concentration threshold in 60% of RNA samples and control samples. Total RNA-seq processes are prone to residual gDNA contamination compared with targeted RNA-seq methods that selectively isolate a subset of RNA transcripts from a sample. Excessive gDNA can distort transcript quantification and affect analysis. Despite following manufacturer-recommended automated RNA isolation protocols, the high failure rate suggested that the initial DNase treatment was insufficient, warranting consideration of an additional DNase treatment.

Presence of Excess gDNA Affects Sequencing Results

A secondary DNase treatment was implemented and evaluated to further reduce residual gDNA before library preparation and RNA-seq. In a preliminary test, four RNA samples underwent secondary DNase treatment and RNA clean-up to improve RNA purity. This treatment reduced gDNA concentration by >90% (Figure 2A), and all samples met acceptable gDNA concentration thresholds after secondary DNase treatment (Figure 2B).

Figure 2.

Figure 2

Secondary DNase treatment reduces residual genomic DNA (gDNA) for downstream analysis. Total RNA was isolated from whole blood PAXgene Blood RNA specimens. A and B: gDNA concentration (A) and percentage of gDNA relative to RNA concentration (B) were measured before [orange; untreated (Unt)] and after [blue; treated (Trt)] secondary DNase treatment. The dotted line indicates the acceptable upper threshold for residual gDNA. C: Stacked bar graphs (left) illustrate the percentage of sequenced reads by genomic origins for a representative sample set before (Trt) and after (Unt) secondary DNase treatment. Pie charts (right) illustrate the average percentage of uniquely mapped reads mapping to intergenic (orange), exonic (blue), and intronic (teal) genomic regions (n = 4).

With gDNA significantly reduced, sequencing results were further assessed to identify the impact of the secondary DNase treatment. Although intronic read mapping remained unchanged, secondary DNase treatment led to a significant increase in exonic read alignment (P = 0.0013) and a decrease in intergenic read mapping (P = 0.0007) across all samples (Figure 2C). Compared with poly(A) libraries, total RNA-seq libraries often exhibit higher proportions of intergenic and intronic reads because noncoding RNAs, pre-mRNAs, and other RNA species are retained during processing.51,52 gDNA contamination increases background expression in RNA-seq data, particularly for intergenic regions and low-expression genes. This can lead to inaccuracies in quantification and result in false-positive discoveries, particularly for total RNA-seq studies.50, 51, 52, 53 Although DNase-treated samples can exhibit lower mapping percentages in targeted poly(A) capture, the inclusion and optimization of DNase treatment protocols in total RNA-seq studies may be necessary to minimize overestimation of gene expression.54, 55, 56 Comparative analysis using Integrative Genomics Viewer showed that samples treated with secondary DNase exhibited lower background sequence counts, particularly in low-expression genes (Supplemental Figure S1). These findings indicate that removing residual gDNA improves transcript quantification and overall sequencing accuracy.

Because DNase treatment can affect both RNA and DNA yield, its effects on RNA integrity and sequencing performance were examined. As expected, secondary DNase treatment lowered RNA concentrations by an average of 23% (data not shown), but yields remained sufficient for RNA-seq. RINs before and after treatment were compared across 84 RNA samples to assess potential impacts on RNA integrity. Although a significant decrease in RINs was observed, the majority of samples remained within acceptable analytical thresholds (Figure 3A). Only 3.6% of samples (n = 3) fell below the RIN threshold after treatment. Interestingly, pretreatment and posttreatment RINs showed no correlation (R2 = 2 × 10–5), suggesting that the decrease in RIN was independent of initial RNA quality (Figure 3B). These results indicate that although secondary DNase treatment affects RNA integrity, most samples remained suitable for RNA-seq analysis.

Figure 3.

Figure 3

RNA integrity after secondary DNase treatment is sufficient for downstream processing. Total RNA was isolated from whole blood PAXgene Blood RNA specimens. RNA Integrity Number (RIN) for RNA samples were measured before (untreated) and after (treated) secondary DNase treatment. A: Scatter plot of RINs before (untreated) and after (treated) secondary DNase treatment. Gray circles indicate samples below the acceptable RIN value of 6.0. B: Correlation plot of RINs measured before (untreated) and after (treated) secondary DNase treatment. Dotted lines indicate the linear trend lines for the coefficient of determination (R2). n = 84. ∗∗∗∗P = 0.0001.

Analytical Quality Assessment

The purpose of analytical quality assessment is to evaluate sequencing data quality and alignment accuracy against a reference genome. These assessments ensure that only high-quality data are used in downstream analyses. Similar to preanalytical QC, multiple metrics were defined and evaluated at each stage of the analytical process. Key parameters included sequencing depth, base quality scores, and GC content, which were assessed using FastQC, a widely used tool for evaluating Illumina sequencing reads.

After alignment, mapping rates were analyzed to determine sequencing accuracy. Although multiple alignment tools, including STAR,24 HISAT2,57 Salmon,58 and RSEM,59 are available, the mapping rates of uniquely and multimapped reads were assessed in this study using STAR.24,41 Additional genomic region mapping (exonic, intronic, and intergenic reads) was evaluated using Qualimap 2.27 Acceptance criteria were based on tool developer recommendations and prior empirical testing. In biomarker discovery, only samples meeting all analytical QC criteria were used for DE analysis and other advanced computational methods.

Of the initial 297 processed samples, 209 RNA-seq samples were assessed and passed all preanalytical QC metrics (PAXgene specimen weight, RNA concentration, RIN, DV200, and gDNA concentration) before proceeding to sequencing. A total of 92.3% of samples met all analytical QC thresholds (Table 2). The highest failure rates were observed in alignment metrics, with 2.9% failing uniquely mapped read thresholds and 5.3% failing multimapped read thresholds. No samples failed intergenic read mapping, with intergenic mapping rates ranging from 2.9% to 6.5%. These results indicate that the majority of samples passing preanalytical QC produced high-quality sequencing data, making them suitable for DE, ML, and other advanced analyses.

Table 2.

Summary of Analytical QC Metrics Assessed for PAXgene Blood RNA Specimens and External RNA Controls

QC metrics Acceptance criteria Clinical specimens
RNA controls
No. assessed∗ No. failed (%) No. assessed No. failed (%)
Sequencing
 Sequencing quality score ≥80% of bases ≥Q30 209 0 (0) 12 0 (0)
 Sequencing depth ≥40 million reads 209 5 (2.4) 12 1 (8.3)
 % GC content 40%–60% 209 0 (0) 12 0 (0)
Alignment
 % Uniquely mapped reads ≥70% 209 6 (2.9) 12 0 (0)
 % Multimapped reads ≤15% 209 11 (5.3) 12 0 (0)
 % Intergenic reads ≤10% 209 0 (0) 12 0 (0)
∗

A total of 209 samples met all preanalytical quality control (QC) metrics (PAXgene specimen weight, RNA concentration, RNA Integrity Number, percentage of RNA fragments >200 nucleotides, and genomic DNA concentration).

Postanalytical Quality Assessment

The goal of postanalytical quality assessment is to evaluate the biological and clinical relevance of biomarkers identified through advanced analytics, ensuring their successful translation into diagnostic or therapeutic applications. The approach used in postanalytical QC depends on factors such as disease context, study stage, and data availability. RNA-seq data analysis often yields hundreds to thousands of candidate biomarkers, making validation and prioritization challenging. A thorough postanalytical QC process refines this list to the most robust, biologically meaningful, and clinically actionable biomarkers.

Unlike earlier QC stages, postanalytical assessment does not rely on rigid acceptance criteria. Instead, it compares results from multiple analytical methods and validates those findings in biological and clinical contexts (Figure 4A). For instance, DE analysis, which is central to biomarker discovery, can be performed by using tools such as DESeq2,34 NOIseq,60 or EdgeR,61 each employing different statistical models and normalization algorithms that may yield varying sets of differentially expressed genes. Comparative analyses can also encompass multiple approaches (eg, DE, ML, unsupervised clustering), and identifying overlapping candidates across these pipelines strengthens confidence in the results. Further validation involves examining gene networks, functional pathways, and existing clinical data, which support explainability and show that the biomarkers are biologically plausible and clinically relevant to researchers and clinicians. Ultimately, the strongest evidence for reproducible biomarkers emerges when multiple methods, data layers, and independent cohorts converge on the same candidates (Figure 4B). Such consistency across diverse analyses and data sets underscores the reliability of these candidate biomarkers and enhances their potential for clinical application.

Figure 4.

Figure 4

Postanalytical assessments emphasize the discovery of reproducible and biologically relevant biomarkers. A: The Venn diagram (left) represents a result comparison across four independent biomarker discovery analyses to identify shared results (candidate biomarkers). Rationale and explainability (right) of identified candidate biomarkers are primarily conducted through reviewing published literature, assessing pathway enrichment and network interactions, and verifying results with publicly available data sets. B: Illustration depicts the funneling of biomarkers through multilayered comparative analyses, including various tools and methods, omics data sets, and independent cohorts, to deliver a narrowed and robust list of candidate biomarkers for independent validation.

To illustrate how multistage QC strategies can enhance biomarker discovery, candidate RNA biomarkers were identified that distinguish treatment-naive relapsing-remitting multiple sclerosis (RRMS) from treatment-naive neuromyelitis optica (NMO), two chronic demyelinating inflammatory diseases with similar clinical presentations.33 DE analysis identified 208 genes exhibiting a ≥2-fold change (adjusted P value <0.05) and an average normalized expression ≥10 (Supplemental Figure S2A). Of these, 77 (41.2%) have been previously associated with multiple sclerosis, NMO, neurodegenerative diseases, or autoimmune diseases. ML methods incorporating recursive feature elimination, highly variable gene binning, and χ2 test feature selection strategies produced models with classification accuracies of 78% to 94% (Supplemental Figure S2B). The top 200 features with highest feature-impact scores across all nine models were identified. In parallel, unsupervised Leiden clustering performed with consensus non-negative matrix factorization revealed distinct expression profiles for RRMS and NMO (Supplemental Figure S2C).

Comparative analysis of the DE, ML, and consensus non-negative matrix factorization results converged on 16 blood-based candidate biomarkers (Supplemental Figure S2D). Pathway enrichment analysis of these 16 genes (Supplemental Figure S2E) highlighted processes involved in the microbial host response (implicated in NMO’s pronounced inflammatory episodes),62, 63, 64 iron and heme metabolism (reported in demyelination and neurodegeneration in RRMS),65, 66, 67 translational regulation, and redox balance; these mechanisms are repeatedly associated with neurodegenerative autoimmune pathophysiology in both conditions.68, 69, 70, 71

External RNA Controls Monitor Batch-to-Batch Performance

RNA-seq data are highly susceptible to technical variability, which can introduce biases that affect data interpretation.16,17 One major concern is batch-to-batch variation, collectively known as batch effects, which commonly affect next-generation sequencing technologies.72, 73, 74 Although computational tools include batch effect correction methods, biomarker discovery programs do not always monitor these effects. Technical variations arise during preanalytical processing and sequencing, and, in many studies, samples are sequenced in a single batch to minimize cost and complexity. As a result, batch effects may go undetected when comparing data across studies or integrating data sets.

To monitor batch effects, two bulk-prepared and aliquoted external controls (HeLa-QC and MCF7-QC) were processed through a secondary DNase treatment, library construction, and sequencing in parallel with PAXgene Blood RNA samples across six independent runs. Both controls were sequenced at the same depth as clinical specimens and showed comparable rates of unique and intergenic mapping (Figure 5A).

Figure 5.

Figure 5

External RNA controls show minimal batch-to-batch variance in the RNA biomarker discovery program. Bulk control RNA was isolated from HeLa and MCF7 cell lines, DNase treated, and stored in 50 ng/μL single-use aliquots. Total RNA was isolated from whole-blood PAXgene Blood RNA specimens. Secondary DNase-treatment was performed on all RNA samples with subsequent stranded library preparation with rRNA and globin depletion. Raw sequences were aligned and quantified. A: Bar graphs illustrate analytical quality metrics: sequencing depth (left), percentage of uniquely mapped reads (mapping rate, middle), and percentage of uniquely mapped reads with intergenic genomic origin (right) for RNA controls. B: Uniform Manifold Approximation and Projection (UMAP) embeddings depict unsupervised Leiden clustering populations across batches (B1, B2, B3, B4, B5, and B6) for each HeLa and MCF7 RNA control. C: Correlations (r) were calculated across batch-to-batch principal component analysis (PCA) for each HeLa and MCF7 RNA control. HeLa, n = 6 and MCF7, n = 6; and PAXgene Blood RNA, n = 288 (A). QC, quality control.

Batch-to-batch variation was negligible. Unsupervised Leiden clustering effectively separated the HeLa-QC and MCF7-QC control libraries sequenced across six independent runs, confirming stable identity (Figure 5B). PCA supported these findings: pairwise correlations ranged from 0.81 to 0.99 for HeLa-QC and 0.84 to 0.99 for MCF7-QC, with no significant longitudinal trends (Figure 5C). CVs followed the expected abundance-dependent pattern: highly expressed transcripts (mean approximately 7700 counts) showed tight dispersion (CV 17% to 18%), moderately expressed genes (approximately 200 counts) exhibited intermediate variability (CV 34% to 36%), and low-abundance transcripts (<10 counts) were dominated by stochastic noise (CV approximately 130%) (Supplemental Table S2). This mean–variance relationship mirrors findings from large benchmarking studies, indicating that control performance is consistent with established RNA-seq reproducibility metrics.16,21,75 Filtering out genes with <10 counts, the principal contributors to the noise tier, did not significantly change the PCA correlations (data not shown), indicating that inter-batch consistency is driven by the moderate- and high-expression genes retained for downstream analyses.

Although PCA is a standard tool for detecting batch effects,76,77 establishing definitive acceptance thresholds requires extended monitoring over many runs. Here, the combination of high PCA correlations, clear clustering, and a preliminary correlation threshold (0.80) suggests that no batch failed QC and supports the overall robustness of the sequencing workflow.

Discussion

Next-generation sequencing has significantly advanced molecular diagnostics, with DNA sequencing becoming an established tool for evaluating inherited diseases. Its clinical adoption was facilitated by collaborative efforts among manufacturers, clinical providers, and regulatory agencies, leading to standardized development and validation guidelines.42, 43, 44,78, 79, 80 However, RNA-seq lacks similar regulatory oversight, making its integration into clinical diagnostics more challenging. The FDA’s 2018 guidance for next-generation sequencing–based in vitro diagnostics specifically excluded RNA-seq,43 leaving gaps in regulatory recommendations for its clinical application. However, many principles from next-generation sequencing–based guidance, such as specimen integrity, sequencing depth, and quality controls, are relevant to RNA-seq studies. Clinical adoption of RNA-seq remains limited due to technical variability, lack of standardization, and challenges in reproducibility.

One of the major obstacles in RNA biomarker discovery is the variability in RNA-seq results, particularly when using liquid biopsy samples such as whole blood.11,13,19,81 Study design and cohort inclusion and exclusion criteria, especially the selection of healthy or disease controls, can influence the reproducibility of biomarker discovery studies.82,83 In addition to biological variability, differences in sample collection, handling, and processing, including varying reagent lots and controls or PCR bias, can affect gene expression measurements and reproducibility.11, 12, 13, 14, 15,19,46,47,81 The current study did not aim to address all RNA-seq challenges but rather highlights critical steps and process controls in laboratory and informatics workflows that affect data variability. Ensuring the appropriate selection of disease controls and accurate clinical annotation, along with reducing technical variability early in transcriptomic studies, is essential for improving research efficiency and translating RNA biomarkers into actionable diagnostic or therapeutic applications.

RNA biomarker discovery spans multiple stages, from specimen collection to data analysis. Although RNA-seq QC is often emphasized at the analytical stage, significant risks arise during preanalytical processing, including sample collection, RNA isolation, DNase treatment, and library preparation. Relying on a single QC metric, such as RIN, may overlook quality deficiencies that affect biomarker selection and interpretation. A multilayered QC approach is therefore necessary to detect and mitigate quality issues at multiple stages, ensuring accurate and reproducible gene expression analysis.

Preanalytical QC measures provide a more comprehensive assessment of RNA sample quality and can identify process deficiencies early. Peripheral whole blood is a widely accepted specimen type, and FDA-approved collection devices maintain RNA stability. However, variations in blood collection volume can affect results. The current study found that 3.6% of PAXgene Blood RNA tubes were underfilled (<18.0 g), leading to exclusion from downstream processing. Implementing simple quality assessments, such as tube weight measurements, allows for early specimen rejection or re-collection, minimizing wasted resources and suboptimal sequencing outcomes.

Beyond simply measuring RNA quantity, assessing RNA purity and integrity is essential for robust QC. RIN remains the most widely used metric; however, because these values rely heavily on ribosomal RNA measurements, RIN may not fully capture mRNA integrity. The DV200 metric, which quantifies the proportion of RNA fragments >200 nucleotides, can better reflect the quality of functional mRNA in certain applications.48 In this study, RIN values had the highest failure rate (4.0%) compared with those of DV200 (1.8%), reinforcing the importance of using both metrics to comprehensively evaluate RNA integrity.

Another key concern in total RNA-seq is the presence of residual gDNA, which can distort gene quantification more significantly compared with poly(A) selection protocols.50,52, 53, 54, 55 Secondary DNase treatments can help reduce gDNA contamination but may also lower yields for less abundant transcripts.50,54,84,85 The optimal DNase strategy varies based on the sample collection system and the reagents used. For instance, PAXgene Blood RNA tubes and their associated protocols may result in more residual gDNA, necessitating additional off-column DNase digestion.86,87 gDNA measurements revealed that 59.4% of clinical RNA samples and 62.5% of batch controls exceeded acceptable contamination thresholds after the first DNase treatment, prompting a second DNase step. This secondary treatment significantly reduced gDNA levels and improved sequencing accuracy.

Analytical QC further ensures high-quality sequencing data by examining read depth, base quality, and alignment metrics. This workflow targeted a minimum of 40 million reads per sample to support robust transcriptome profiling. Base quality and GC content were evaluated by using FastQC, and STAR and Qualimap were used to assess alignment accuracy. Importantly, optimized DNase protocols reduced intergenic mapping rates but had a relatively minor impact on intronic reads, which often reflect legitimate biological signals such as unspliced pre-mRNA, nascent transcripts, or retained introns that can have regulatory functions.52 Due to the potential biological relevance of these transcripts, care was taken to avoid excessive DNase treatment that might degrade them. Ultimately, 92.3% of samples that passed preanalytical QC also met all analytical QC criteria, confirming their suitability for downstream DE and ML analyses.

Postanalytical QC focuses on biomarker validation and reproducibility, incorporating comparative analyses, pathway enrichment, literature review, and clinical relevance. Because the various DE tools (eg, DESeq2, EdgeR, NOIseq) use different normalization strategies, cross-validation of results across multiple tools improves confidence in biomarker selection. In addition, comparing identified biomarkers across independent data sets of patients with the same disease or disease state (eg, treatment, remission, progression) enhances their clinical relevance and likelihood of successful validation. ML and network-based analyses can further refine top candidate biomarkers, ensuring that only the most robust and biologically meaningful biomarkers advance to clinical testing.

A comprehensive QC framework for RNA biomarker discovery must also incorporate batch-to-batch performance monitoring. Batch effects, caused by variations in reagents, sequencing lanes, operators, and time of processing, can significantly affect RNA-seq reproducibility.72, 73, 74,77 Although computational tools such as limma88 and ComBat-seq89 correct for batch effects, monitoring batch-to-batch variation using external RNA controls provides an additional layer of quality assurance. In this study, external HeLa and MCF7 RNA controls exhibited minimal batch-to-batch variability (correlation coefficients, 0.81 to 0.99 for HeLa-QC and 0.84 to 0.99 for MCF7-QC), supporting consistent sequencing performance across multiple batches.

In summary, RNA-seq biomarker discovery demands a multitiered QC framework that encompasses preanalytical, analytical, and postanalytical stages. The absence of standardized QC measures in publicly available data sets restricts reproducibility and limits broader applications. By implementing rigorous QC strategies at each step, from specimen collection to data analysis, researchers can ensure reliable biomarker identification and enhance the potential of RNA-seq for clinical translation. The current study introduces a structured approach to RNA-seq QC, promoting more effective biomarker discovery and the future integration of transcriptomic data into clinical diagnostics.

Several important considerations arise when adopting an end-to-end QC framework for RNA-seq biomarker discovery. Although originally designed for whole blood processed by total RNA-seq, the core principles here can be adapted to a variety of disease indications, sample types, and library preparation methods. One example is the optional secondary DNase treatment: although it helps reduce gDNA contamination, it may impair the detection of low-abundance transcripts and should be carefully evaluated based on specific study goals and the specimens used.56 Similarly, acceptance criteria for evaluating RNA integrity, such as combining RIN and DV200 metrics, may need adjustment for different biospecimen sources (eg, formalin-fixed, paraffin-embedded tissue, liquid biopsy specimens) and experimental designs.56 In some situations, only a subset of the QC steps (eg, analytical QC alone) may be sufficient, depending on the application and available sample volume.

This adaptability can be illustrated with cell-free plasma, in which limited sample input is entirely allocated to a low-input library protocol. In such instances, the absence of extra material shifts much of the quality verification to library preparation and downstream analytical QC, which may not generalize to other specimen types. Although the framework has these inherent limitations, it offers a practical roadmap for multitiered quality assurance at the preanalytical, analytical, and postanalytical stages. Implementing comprehensive QC procedures mitigates the current lack of standardized measures, thus enhancing reproducibility and expanding the utility of RNA-seq in clinical applications. Adopting rigorous QC checkpoints from specimen collection through data interpretation advances the reliability of RNA-seq–based biomarker discovery and strengthens its potential for clinical translation.

To advance the clinical adoption of RNA-seq, a standardized QC framework must be established across all data generation and analysis stages. Preanalytical QC should consider specimen collection, RNA integrity, and gDNA contamination to ensure sample quality before sequencing. Analytical QC should define metrics to assess sequencing, alignment, and batch effects. Bioinformatics pipelines should incorporate validated analytical tools and transparent reporting of computational parameters to ensure data consistency across studies. Postanalytical QC should require independent validation of RNA biomarkers across multiple analytical methods, data sets, and functional characterization. In addition, regulatory agencies should establish clear classification and approval pathways for RNA-seq–based diagnostics, including analytical and clinical validation recommendations. Implementing these guidelines will improve RNA-seq data reliability, accelerate biomarker discovery, and facilitate its translation into clinically actionable diagnostics and therapeutics.

Disclosure Statement

C.L.S., G.I.S., L.S.W., and C.F.S. are employed by and own stock in Decode Health, Inc. E.V.G. and F.R.C. are consultants and own stock in Decode Health, Inc.

Acknowledgments

We thank the patients who have participated in continuing research, without whom this work would not be possible, and Jamieson Gray for his help with the publication of this manuscript.

Author Contributions

C.L.S., G.I.S., L.S.W., E.V.G., F.R.C., and C.F.S. conceived the study; C.L.S., G.I.S., L.S.W., E.V.G., and C.F.S. collected and analyzed the data; C.L.S., G.I.S., L.S.W., E.V.G., and C.F.S. wrote the manuscript; all authors revised the manuscript; and C.F.S. supervised the study. C.F.S. had full access to all study data and takes responsibility for the integrity of the data analysis.

Footnotes

Supported by Decode Health, Inc., Quest Diagnostics, Inc., and the National Institutes of Health (NIH/NIAID) R43 AI157674-01.

Supplemental material for this article can be found at http://doi.org/10.1016/j.jmoldx.2025.05.008.

Supplemental Data

Supplemental Table S1
mmc1.docx (29KB, docx)
Supplemental Table S2
mmc2.docx (28.6KB, docx)
Supplemental Figure S1

Secondary DNase treatment reduces background sequence counts of representative low-expression genes. Total RNA was isolated from a representative set of whole blood PAXgene Blood RNA specimens. An aliquot of each RNA sample was subjected to a secondary DNase treatment before library preparation and sequencing. Integrative Genomics Viewer images depict sequence read counts for two low-expression genes of interest, KIR3DL1 and SERINC2, for samples with (blue) and without (orange) a secondary DNase treatment. Gray rectangles below the graphs indicate exonic regions.

mmc3.pdf (306.3KB, pdf)
Supplemental Figure S2

Identification and postanalytical assessment of candidate biomarkers that distinguish patients with relapsing-remitting multiple sclerosis (RRMS) and neuromyelitis optica (NMO). A: Volcano plot shows statistical significance [–log10 (adjusted P value)] by the magnitude of change [log2 fold change (FC)] for differentially expressed genes between RRMS and NMO. Blue dots indicate statistically differentially expressed genes (SDEGs; adjusted P value <0.05) that exhibited log2 FC ≥1.0 and mean expression counts ≥10. The table indicates summary statistics for genes indicated in blue. B: Validation accuracies of the top three machine learning (ML) models from each feature selection technique [recursive feature elimination (RFE), χ2 test (Chi2), and highly variable gene binning (HVG)] with confusion matrix from the overall top model. Expression patterns depict the Z-score normalized differential expression (DE) values of the top 15 genes with the highest feature-impact scores across all nine models. Coloration denotes SDs from the mean computed across all samples (mean, 0). C: Uniform Manifold Approximation and Projection embedding depicts consensus non-negative matrix factorization (cNMF) usage similarity. The Leiden clustering algorithm was used to partition samples according to community gene expression profiles. Representative usages indicate unique gene expression patterns for sample subsets (indicated with ovals) with an optimal k value (k = 3). D: A Venn diagram illustrates the overlap of top genes (candidate biomarkers) identified from DE, ML, and cNMF for NMO versus RRMS. The table indicates the top five candidate biomarkers shared across all methods by the magnitude of DE (log2 FC). E: Genomic Regions Enrichment of Annotations Tool analysis of perturbed biological networks using the candidate biomarkers identified through triangulation of advanced analytics approaches (DE, ML, and cNMF) comparing whole blood samples of the NMO and RRMS cohorts. Orange circles signify P value (–log10) denoted on the top x-axis. Blue bars represent the pathway fold enrichment denoted on the lower x-axis. ∗Average of normalized count values across samples. snoRNA, small nucleolar RNA.

mmc4.pdf (707.7KB, pdf)

References

  • 1.Wang Z., Gerstein M., Snyder M. RNA-seq: a revolutionary tool for transcriptomics. Nat Rev Genet. 2009;10:57–63. doi: 10.1038/nrg2484. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Byron S.A., Van Keuren-Jensen K.R., Engelthaler D.M., Carpten J.D., Craig D.W. Translating RNA sequencing into clinical diagnostics: opportunities and challenges. Nat Rev Genet. 2016;17:257–271. doi: 10.1038/nrg.2016.10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Smail C., Montgomery S.B. RNA sequencing in disease diagnosis. Annu Rev Genomics Hum Genet. 2024;25:353–367. doi: 10.1146/annurev-genom-021623-121812. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.National Academies of Sciences, Engineering, and Medicine . The National Academies Press; Washington, DC: 2024. Charting a Future for Sequencing RNA and Its Modifications: A New Era for Biology and Medicine. [PubMed] [Google Scholar]
  • 5.Costa V., Aprile M., Esposito R., Ciccodicola A. RNA-Seq and human complex diseases: recent accomplishments and future perspectives. Eur J Hum Genet. 2013;21:134–142. doi: 10.1038/ejhg.2012.129. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Gao M., Zhong A., Patel N., Alur C., Vyas D. High throughput RNA sequencing utility for diagnosis and prognosis in colon diseases. World J Gastroenterol. 2017;23:2819–2825. doi: 10.3748/wjg.v23.i16.2819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Marco-Puche G., Lois S., Benítez J., Trivino J.C. RNA-seq perspectives to improve clinical diagnosis. Front Genet. 2019;10:1152. doi: 10.3389/fgene.2019.01152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Buzdin A., Sorokin M., Garazha A., Glusker A., Aleshin A., Poddubskaya E., Sekacheva M., Kim E., Gaifullin N., Giese A., Seryakov A., Rumiantsev P., Moshkovskii S., Moiseev A. RNA sequencing for research and diagnostics in clinical oncology. Semin Cancer Biol. 2020;60:311–323. doi: 10.1016/j.semcancer.2019.07.010. [DOI] [PubMed] [Google Scholar]
  • 9.Boyle T.A., Bossler A.D. RNA sequencing steps toward the first line. Cancer. 2023;129:2294–2296. doi: 10.1002/cncr.34801. [DOI] [PubMed] [Google Scholar]
  • 10.Cabús L., Lagarde J., Curado J., Lizano E., Pérez-Boza J. Current challenges and best practices for cell-free long RNA biomarker discovery. Biomark Res. 2022;10:62. doi: 10.1186/s40364-022-00409-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sprang M., Möllmann J., Andrade-Navarro M.A., Fontaine J.F. Overlooked poor-quality patient samples in sequencing data impair reproducibility of published clinically relevant datasets. Genome Biol. 2024;25:222. doi: 10.1186/s13059-024-03331-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lippi G., Banfi G., Church S., Cornes M., De Carli G., Grankvist K., Kristensen G.B., Ibarz M., Panteghini M., Plebani M., Nybo M., Smellie S., Zaninotto M., Simundic A.M., European Federation for Clinical Chemistry and Laboratory Medicine Working Group for Preanalytical Phase Preanalytical quality improvement. In pursuit of harmony, on behalf of European Federation for Clinical Chemistry and Laboratory Medicine (EFLM) Working group for Preanalytical Phase (WG-PRE) Clin Chem Lab Med. 2015;53:357–370. doi: 10.1515/cclm-2014-1051. [DOI] [PubMed] [Google Scholar]
  • 13.Peplies J., Günther K., Bammann K., Fraterman A., Russo P., Veidebaum T., Tornaritis M., Vanaelst B., Mårild S., Molnár D., Moreno L.A., Ahrens W., IDEFICS Consortium Influence of sample collection and preanalytical sample processing on the analyses of biological markers in the European multicentre study IDEFICS. Int J Obes (Lond) 2011;35(Suppl 1):S104–S112. doi: 10.1038/ijo.2011.41. [DOI] [PubMed] [Google Scholar]
  • 14.Hansen K.D., Wu Z., Irizarry R.A., Leek J.T. Sequencing technology does not eliminate biological variability. Nat Biotechnol. 2011;29:572–573. doi: 10.1038/nbt.1910. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Shen Y., Li R., Tian F., Chen Z., Lu N., Bai Y., Ge Q., Lu Z. Impact of RNA integrity and blood sample storage conditions on the gene expression analysis. Onco Targets Ther. 2018;11:3573–3581. doi: 10.2147/OTT.S158868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.McIntyre L.M., Lopiano K.K., Morse A.M., Amin V., Oberg A.L., Young L.J., Nuzhdin S.V. RNA-seq: technical variability and sampling. BMC Genomics. 2011;12:293. doi: 10.1186/1471-2164-12-293. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Arora S., Pattwell S.S., Holland E.C., Bolouri H. Variability in estimated gene expression among commonly used RNA-seq pipelines. Sci Rep. 2020;10:2734. doi: 10.1038/s41598-020-59516-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kalesinskas L., Gupta S., Khatri P. Increasing reproducibility, robustness, and generalizability of biomarker selection from meta-analysis using Bayesian methodology. PLoS Comput Biol. 2022;18 doi: 10.1371/journal.pcbi.1010260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Cadamuro J., Baird G., Baumann G., Bolenius K., Cornes M., Ibarz M., Lewis T., Oliveira G.L., Lippi G., Plebani M., Simundic A.-M., von Meyer A. Preanalytical quality improvement—an interdisciplinary journey, on behalf of the European Federation for Clinical Chemistry and Laboratory Medicine (EFLM) Working Group for Preanalytical Phase (WG-PRE) Clin Chem Lab Med. 2022;60:662–668. doi: 10.1515/cclm-2022-0117. [DOI] [PubMed] [Google Scholar]
  • 20.Conesa A., Madrigal P., Tarazona S., Gomez-Cabrero D., Cervera A., McPherson A., Szcześniak M.W., Gaffney D.J., Elo L.L., Zhang X., Mortazavi A. A survey of best practices for RNA-seq data analysis. Genome Biol. 2016;17:13. doi: 10.1186/s13059-016-0881-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.SEQC/MAQC-III Consortium A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the sequencing quality control consortium. Nat Biotechnol. 2014;32:903–914. doi: 10.1038/nbt.2957. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Sheng Q., Vickers K., Zhao S., Wang J., Samuels D.C., Koues O., Shyr Y., Guo Y. Multi-perspective quality control of Illumina RNA sequencing data analysis. Brief Funct Genomics. 2017;16:194–204. doi: 10.1093/bfgp/elw035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Brown J., Pirrung M., McCue L.A. FQC dashboard: integrates FastQC results into a web-based, interactive, and extensible FASTQ quality control tool. Bioinformatics. 2017;33:3137–3139. doi: 10.1093/bioinformatics/btx373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Dobin A., Davis C.A., Schlesinger F., Drenkow J., Zaleski C., Jha S., Batut P., Chaisson M., Gingeras T.R. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. doi: 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Liao Y., Smyth G.K., Shi W. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30:923–930. doi: 10.1093/bioinformatics/btt656. [DOI] [PubMed] [Google Scholar]
  • 26.García-Alcalde F., Okonechnikov K., Carbonell J., Cruz L.M., Götz S., Tarazona S., Dopazo J., Meyer T.F., Conesa A. Qualimap: evaluating next-generation sequencing alignment data. Bioinformatics. 2012;28:2678–2679. doi: 10.1093/bioinformatics/bts503. [DOI] [PubMed] [Google Scholar]
  • 27.Okonechnikov K., Conesa A., García-Alcalde F. Qualimap 2: advanced multi-sample quality control for high-throughput sequencing data. Bioinformatics. 2016;32:292–294. doi: 10.1093/bioinformatics/btv566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Robinson J.T., Thorvaldsdóttir H., Winckler W., Guttman M., Lander E.S., Getz G., Mesirov J.P. Integrative genomics viewer. Nat Biotechnol. 2011;29:24–26. doi: 10.1038/nbt.1754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Traag V.A., Waltman L., van Eck N.J. From Louvain to Leiden: guaranteeing well-connected communities. Sci Rep. 2019;9:5233. doi: 10.1038/s41598-019-41695-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.McInnes L., Healy J., Saul N., Groβberger L. UMAP: uniform manifold approximation and projection. J Open Source Softw. 2018;3:861. [Google Scholar]
  • 31.R Core Team . R Foundation for Statistical Computing; Vienna, Austria: 2022. R: A language and environment for statistical computing. [Google Scholar]
  • 32.Wickham H. Springer; New York, NY: 2016. ggplot2: Elegant Graphics for Data Analysis. [Google Scholar]
  • 33.Wylezinski L.S., Sesler C.L., Shaginurova G.I., Grigorenko E.V., Wohlgemuth J.G., Cockerill F.R., III, Racke M.K., Spurlock C.F., III Machine learning analysis using RNA sequencing to distinguish neuromyelitis optica from multiple sclerosis and identify therapeutic candidates. J Mol Diagn. 2024;26:520–529. doi: 10.1016/j.jmoldx.2024.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Love M.I., Huber W., Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550. doi: 10.1186/s13059-014-0550-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Benjamini Y., Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat B Methodol. 1995;57:289–300. [Google Scholar]
  • 36.Sanz H., Valim C., Vegas E., Oller J.M., Reverter F. SVM-RFE: selection and visualization of the most relevant features through non-linear kernels. BMC Bioinformatics. 2018;19:432. doi: 10.1186/s12859-018-2451-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Setiono R., Liu H. Proceedings of the Seventh IEEE International Conference on Tools with Artificial Intelligence. IEEE Computer Society; Los Alamitos, CA: 1995. Chi2: feature selection and discretization of numeric attributes; pp. 388–391. [Google Scholar]
  • 38.Yip S.H., Sham P.C., Wang J. Evaluation of tools for highly variable gene discovery from single-cell RNA-seq data. Brief Bioinform. 2019;20:1583–1589. doi: 10.1093/bib/bby011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Kotliar D., Veres A., Nagy M.A., Tabrizi S., Hodis E., Melton D.A., Sabeti P.C. Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-Seq. Elife. 2019;8 doi: 10.7554/eLife.43803. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.McLean C.Y., Bristor D., Hiller M., Clarke S.L., Schaar B.T., Lowe C.B., Wenger A.M., Bejerano G. GREAT improves functional interpretation of cis-regulatory regions. Nat Biotechnol. 2010;28:495–501. doi: 10.1038/nbt.1630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Dobin A., Gingeras T.R. Mapping RNA-seq reads with STAR. Curr Protoc Bioinformatics. 2015;51:11.14.1–11.14.19. doi: 10.1002/0471250953.bi1114s51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Endrullat C., Glökler J., Franke P., Frohme M. Standardization and quality management in next-generation sequencing. Appl Transl Genom. 2016;10:2–9. doi: 10.1016/j.atg.2016.06.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.US Food and Drug Administration . US Food and Drug Administration; Silver Spring, MD: 2018. Considerations for Design, Development, and Analytical Validation of Next Generation Sequencing (NGS)–Based In Vitro Diagnostics (IVDs) Intended to Aid in the Diagnosis of Suspected Germline Diseases. [Google Scholar]
  • 44.Clinical and Laboratory Standards Institute . Human Genetic and Genomic Testing Using Traditional and High-Throughput Nucleic Acid Sequencing Methods, 3rd ed., CLSI guideline MM09. Clinical and Laboratory Standards Institute; Wayne, PA: 2023. [Google Scholar]
  • 45.Schroeder A., Mueller O., Stocker S., Salowsky R., Leiber M., Gassmann M., Lightfoot S., Menzel W., Granzow M., Ragg T. The RIN: an RNA integrity number for assigning integrity values to RNA measurements. BMC Mol Biol. 2006;7:3. doi: 10.1186/1471-2199-7-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Gallego Romero I., Pai A.A., Tung J., Gilad Y. RNA-seq: impact of RNA degradation on transcript quantification. BMC Biol. 2014;12:42. doi: 10.1186/1741-7007-12-42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Schuierer S., Carbone W., Knehr J., Petitjean V., Fernandez A., Sultan M., Roma G. A comprehensive assessment of RNA-seq protocols for degraded and low-quantity samples. BMC Genomics. 2017;18:442. doi: 10.1186/s12864-017-3827-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Matsubara T., Soh J., Morita M., Uwabo T., Tomida S., Fujiwara T., Kanazawa S., Toyooka S., Hirasawa A. DV200 index for assessing RNA integrity in next-generation sequencing. Biomed Res Int. 2020;2020 doi: 10.1155/2020/9349132. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Lu W., Zhou Q., Chen Y. Impact of RNA degradation on next-generation sequencing transcriptome data. Genomics. 2022;114 doi: 10.1016/j.ygeno.2022.110429. [DOI] [PubMed] [Google Scholar]
  • 50.Li X., Zhang P., Wang H., Yu Y. Genes expressed at low levels raise false discovery rates in RNA samples contaminated with genomic DNA. BMC Genomics. 2022;23:554. doi: 10.1186/s12864-022-08785-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Zhao W., He X., Hoadley K.A., Parker J.S., Hayes D.N., Perou C.M. Comparison of RNA-Seq by poly (A) capture, ribosomal RNA depletion, and DNA microarray for expression profiling. BMC Genomics. 2014;15:419. doi: 10.1186/1471-2164-15-419. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Liu H., Hu K., O’Connor K., Kelliher M.A., Zhu L.J. CleanUpRNAseq: an R/Bioconductor package for detecting and correcting DNA contamination in RNA-seq data. BioTech (Basel) 2024;13:30. doi: 10.3390/biotech13030030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Sultan M., Amstislavskiy V., Risch T., Schuette M., Dökel S., Ralser M., Balzereit D., Lehrach H., Yaspo M.L. Influence of RNA extraction methods and library selection schemes on RNA-seq data. BMC Genomics. 2014;15:675. doi: 10.1186/1471-2164-15-675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Markou A.N., Smilkou S., Tsaroucha E., Lianidou E. The effect of genomic DNA contamination on the detection of circulating long non-coding RNAs: the paradigm of MALAT1. Diagnostics (Basel) 2021;11:1160. doi: 10.3390/diagnostics11071160. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Ura H., Togi S., Niida Y. Poly(A) capture full length cDNA sequencing improves the accuracy and detection ability of transcript quantification and alternative splicing events. Sci Rep. 2022;12 doi: 10.1038/s41598-022-14902-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Liu Y., Bhagwate A., Winham S.J., Stephens M.T., Harker B.W., McDonough S.J., Stallings-Mann M.L., Heinzen E.P., Vierkant R.A., Hoskin T.L., Frost M.H., Carter J.M., Pfrender M.E., Littlepage L., Radisky D.C., Cunningham J.M., Degnim A.C., Wang C. Quality control recommendations for RNASeq using FFPE samples based on pre-sequencing lab metrics and post-sequencing bioinformatics metrics. BMC Med Genomics. 2022;15:195. doi: 10.1186/s12920-022-01355-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Kim D., Paggi J.M., Park C., Bennett C., Salzberg S.L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019;37:907–915. doi: 10.1038/s41587-019-0201-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Patro R., Duggal G., Love M.I., Irizarry R.A., Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14:417–419. doi: 10.1038/nmeth.4197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Li B., Dewey C.N. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011;12:323. doi: 10.1186/1471-2105-12-323. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Tarazona S., Furió-Tarí P., Turrà D., Di Pietro A., Nueda M.J., Ferrer A., Conesa A. Data quality aware analysis of differential expression in RNA-seq with NOISeq R/Bioc package. Nucleic Acids Res. 2015;43 doi: 10.1093/nar/gkv711. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Robinson M.D., McCarthy D.J., Smyth G.K. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26:139–140. doi: 10.1093/bioinformatics/btp616. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Xie Q., Sun J., Sun M., Wang Q., Wang M. Perturbed microbial ecology in neuromyelitis optica spectrum disorder: evidence from the gut microbiome and fecal metabolome. Mult Scler Relat Disord. 2024;92 doi: 10.1016/j.msard.2024.105936. [DOI] [PubMed] [Google Scholar]
  • 63.Bar-Or A., Steinman L., Behne J.M., Benitez-Ribas D., Chin P.S., Clare-Salzler M., Healey D., Kim J.I., Kranz D.M., Lutterotti A., Martin R., Schippling S., Villoslada P., Wei C.-H., Weiner H.L., Zamvil S.S., Smith T.J., Yeaman M.R. Restoring immune tolerance in neuromyelitis optica: Part II. Neurol Neuroimmunol Neuroinflamm. 2016;3 doi: 10.1212/NXI.0000000000000276. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Zamvil S.S., Spencer C.M., Baranzini S.E., Cree B.A.C. The gut microbiome in neuromyelitis optica. Neurotherapeutics. 2018;15:92–101. doi: 10.1007/s13311-017-0594-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Schipper H.M., Song W., Zukor H., Hascalovici J.R., Zeligman D. Heme oxygenase-1 and neurodegeneration: expanding frontiers of engagement. J Neurochem. 2009;110:469–485. doi: 10.1111/j.1471-4159.2009.06160.x. [DOI] [PubMed] [Google Scholar]
  • 66.Stephenson E., Nathoo N., Mahjoub Y., Dunn J.F., Yong V.W. Iron in multiple sclerosis: roles in neurodegeneration and repair. Nat Rev Neurol. 2014;10:459–468. doi: 10.1038/nrneurol.2014.118. [DOI] [PubMed] [Google Scholar]
  • 67.Levi S., Ripamonti M., Moro A.S., Cozzi A. Iron imbalance in neurodegeneration. Mol Psychiatry. 2024;29:1139–1152. doi: 10.1038/s41380-023-02399-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Junker A. Pathophysiology of translational regulation by microRNAs in multiple sclerosis. FEBS Lett. 2011;585:3738–3746. doi: 10.1016/j.febslet.2011.03.052. [DOI] [PubMed] [Google Scholar]
  • 69.Harsij A., Gharebaghi A., Ghiasian M., Eslami S., Ghafouri-Fard S., Taheri M., Sayad A. Expression analysis of Treg-related lncRNAs in neuromyelitis optica spectrum disorder. Mult Scler Relat Disord. 2024;81 doi: 10.1016/j.msard.2023.105350. [DOI] [PubMed] [Google Scholar]
  • 70.Ohl K., Tenbrock K., Kipp M. Oxidative stress in multiple sclerosis: central and peripheral mode of action. Exp Neurol. 2016;277:58–67. doi: 10.1016/j.expneurol.2015.11.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Arslan B., Arslan G.A., Tuncer A., Karabudak R., Dinçel A.S. Evaluation of thiol homeostasis in multiple sclerosis and neuromyelitis optica spectrum disorders. Front Neurol. 2021;12 doi: 10.3389/fneur.2021.716195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Leek J.T., Scharpf R.B., Bravo H.C., Simcha D., Langmead B., Johnson W.E., Geman D., Baggerly K., Irizarry R.A. Tackling the widespread and critical impact of batch effects in high-throughput data. Nat Rev Genet. 2010;11:733–739. doi: 10.1038/nrg2825. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Goh W.W.B., Wang W., Wong L. Why batch effects matter in omics data, and how to avoid them. Trends Biotechnol. 2017;35:498–507. doi: 10.1016/j.tibtech.2017.02.012. [DOI] [PubMed] [Google Scholar]
  • 74.Goh W.W.B., Yong C.H., Wong L. Are batch effects still relevant in the age of big data? Trends Biotechnol. 2022;40:1029–1040. doi: 10.1016/j.tibtech.2022.02.005. [DOI] [PubMed] [Google Scholar]
  • 75.Law C.W., Chen Y., Shi W., Smyth G.K. voom: precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biol. 2014;15:R29. doi: 10.1186/gb-2014-15-2-r29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Holmes S., Alekseyenko A., Timme A., Nelson T., Pasricha P.J., Spormann A. Visualization and statistical comparisons of microbial communities using R packages on Phylochip data. Pac Symp Biocomput. 2011:142–153. doi: 10.1142/9789814335058_0016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Tom J.A., Reeder J., Forrest W.F., Graham R.R., Hunkapiller J., Behrens T.W., Bhangale T.R. Identifying and mitigating batch effects in whole genome sequencing data. BMC Bioinformatics. 2017;18:351. doi: 10.1186/s12859-017-1756-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Aziz N., Zhao Q., Bry L., Driscoll D.K., Funke B., Gibson J.S., Grody W.W., Hegde M.R., Hoeltge G.A., Leonard D.G.B., Merker J.D., Nagarajan R., Palicki L.A., Robetorye R.S., Schrijver I., Weck K.E., Voelkerding K.V. College of American Pathologists’ laboratory standards for next-generation sequencing clinical tests. Arch Pathol Lab Med. 2015;139:481–493. doi: 10.5858/arpa.2014-0250-CP. [DOI] [PubMed] [Google Scholar]
  • 79.Gargis A.S., Kalman L., Bick D.P., da Silva C., Dimmock D.P., Funke B.H., et al. Good laboratory practice for clinical next-generation sequencing informatics pipelines. Nat Biotechnol. 2015;33:689–693. doi: 10.1038/nbt.3237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Roy S., Coldren C., Karunamurthy A., Kip N.S., Klee E.W., Lincoln S.E., Leon A., Pullambhatla M., Temple-Smolkin R.L., Voelkerding K.V., Wang C., Carter A.B. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20:4–27. doi: 10.1016/j.jmoldx.2017.11.003. [DOI] [PubMed] [Google Scholar]
  • 81.Simeon-Dubach D., Burt A.D., Hall P.A. Quality really matters: the need to improve specimen quality in biomedical research. J Pathol. 2012;228:431–433. doi: 10.1002/path.4117. [DOI] [PubMed] [Google Scholar]
  • 82.Rundle A., Ahsan H., Vineis P. Better cancer biomarker discovery through better study design. Eur J Clin Invest. 2012;42:1350–1359. doi: 10.1111/j.1365-2362.2012.02727.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Zheng Y. Study design considerations for cancer biomarker discoveries. J Appl Lab Med. 2018;3:282–289. doi: 10.1373/jalm.2017.025809. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Heera R., Sivachandran P., Chinni S.V., Mason J., Croft L., Ravichandran M., Yin L.S. Efficient extraction of small and large RNAs in bacteria for excellent total RNA sequencing and comprehensive transcriptome analysis. BMC Res Notes. 2015;8:754. doi: 10.1186/s13104-015-1726-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Loontiens S., Depestel L., Vanhauwaert S., Dewyn G., Gistelinck C., Verboom K., Van Loocke W., Matthijssens F., Willaert A., Vandesompele J., Speleman F., Durinck K. Purification of high-quality RNA from a small number of fluorescence activated cell sorted zebrafish cells for RNA sequencing purposes. BMC Genomics. 2019;20:228. doi: 10.1186/s12864-019-5608-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Chai V., Vassilakos A., Lee Y., Wright J.A., Young A.H. Optimization of the PAXgene blood RNA extraction system for gene expression analysis of clinical samples. J Clin Lab Anal. 2005;19:182–188. doi: 10.1002/jcla.20075. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Nguyen L.T., Pollock C.A., Saad S. Extraction of high quality and high yield RNA from frozen EDTA blood. Sci Rep. 2024;14:8628. doi: 10.1038/s41598-024-58576-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Ritchie M.E., Phipson B., Wu D., Hu Y., Law C.W., Shi W., Smyth G.K. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43 doi: 10.1093/nar/gkv007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Zhang Y., Parmigiani G., Johnson W.E. ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom Bioinform. 2020;2 doi: 10.1093/nargab/lqaa078. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Table S1
mmc1.docx (29KB, docx)
Supplemental Table S2
mmc2.docx (28.6KB, docx)
Supplemental Figure S1

Secondary DNase treatment reduces background sequence counts of representative low-expression genes. Total RNA was isolated from a representative set of whole blood PAXgene Blood RNA specimens. An aliquot of each RNA sample was subjected to a secondary DNase treatment before library preparation and sequencing. Integrative Genomics Viewer images depict sequence read counts for two low-expression genes of interest, KIR3DL1 and SERINC2, for samples with (blue) and without (orange) a secondary DNase treatment. Gray rectangles below the graphs indicate exonic regions.

mmc3.pdf (306.3KB, pdf)
Supplemental Figure S2

Identification and postanalytical assessment of candidate biomarkers that distinguish patients with relapsing-remitting multiple sclerosis (RRMS) and neuromyelitis optica (NMO). A: Volcano plot shows statistical significance [–log10 (adjusted P value)] by the magnitude of change [log2 fold change (FC)] for differentially expressed genes between RRMS and NMO. Blue dots indicate statistically differentially expressed genes (SDEGs; adjusted P value <0.05) that exhibited log2 FC ≥1.0 and mean expression counts ≥10. The table indicates summary statistics for genes indicated in blue. B: Validation accuracies of the top three machine learning (ML) models from each feature selection technique [recursive feature elimination (RFE), χ2 test (Chi2), and highly variable gene binning (HVG)] with confusion matrix from the overall top model. Expression patterns depict the Z-score normalized differential expression (DE) values of the top 15 genes with the highest feature-impact scores across all nine models. Coloration denotes SDs from the mean computed across all samples (mean, 0). C: Uniform Manifold Approximation and Projection embedding depicts consensus non-negative matrix factorization (cNMF) usage similarity. The Leiden clustering algorithm was used to partition samples according to community gene expression profiles. Representative usages indicate unique gene expression patterns for sample subsets (indicated with ovals) with an optimal k value (k = 3). D: A Venn diagram illustrates the overlap of top genes (candidate biomarkers) identified from DE, ML, and cNMF for NMO versus RRMS. The table indicates the top five candidate biomarkers shared across all methods by the magnitude of DE (log2 FC). E: Genomic Regions Enrichment of Annotations Tool analysis of perturbed biological networks using the candidate biomarkers identified through triangulation of advanced analytics approaches (DE, ML, and cNMF) comparing whole blood samples of the NMO and RRMS cohorts. Orange circles signify P value (–log10) denoted on the top x-axis. Blue bars represent the pathway fold enrichment denoted on the lower x-axis. ∗Average of normalized count values across samples. snoRNA, small nucleolar RNA.

mmc4.pdf (707.7KB, pdf)

Articles from The Journal of Molecular Diagnostics : JMD are provided here courtesy of American Society for Investigative Pathology

RESOURCES