Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Apr 24;17:5654. doi: 10.1038/s41467-026-72089-1

TUSCO: benchmarking transcriptome reconstruction with endogenous single-isoform controls

Tianyuan Liu 1,2, Alejandro Paniagua 1,2, Fabian Jetzinger 1,2,3, Luis Ferrández-Peral 1,5, Adam Frankish 4, Ana Conesa 1,✉
PMCID: PMC13315946  PMID: 42026080

Abstract

Long-read sequencing (LRS) platforms, such as Oxford Nanopore and Pacific Biosciences, enable comprehensive transcriptome analysis but face challenges such as sequencing errors, sample quality variability, and library preparation biases. Current benchmarking approaches address these issues insufficiently: BUSCO assesses transcriptome completeness using conserved single-copy orthologous genes but can misinterpret alternative splicing as gene duplications, while SIRV spike-ins and ERCCs oversimplify real sample complexity, neglecting RNA degradation and RNA-extraction artifacts, thus inflating performance metrics. Simulation algorithms are limited in their ability to recapitulate the complexity of real samples. To overcome these limitations, we introduce the Transcriptome Universal Single-isoform COntrol (TUSCO) benchmarking framework, centered on a curated TUSCO gene set of genes lacking alternative isoforms that can be confidently treated as an internal ground truth. The TUSCO evaluation quantifies precision by identifying reconstructed transcripts that deviate from reference annotations and quantifies sensitivity by verifying detection completeness in human and mouse samples. Masking TUSCO gene set transcripts and replacing them with modified splice variants in the annotation creates a TUSCO-novel challenge that assesses reconstruction of the true, now-unannotated isoforms. Our validation demonstrates that TUSCO metrics provide accurate and reliable benchmarking without external controls, significantly improving quality control standards for transcriptome reconstruction using LRS.

Subject terms: Data processing, Computational platforms and environments


Long-read sequencing enables comprehensive transcriptome characterization but remains challenging to benchmark due to sequencing errors, sample variability, and the limited scope of existing evaluation strategies. This study introduces TUSCO, a benchmarking framework based on a curated set of single-isoform genes that act as an internal ground truth, enabling robust assessment of precision and sensitivity, including recovery of masked isoforms, without requiring external controls in human and mouse transcriptomes.

Introduction

Advances in long-read sequencing (LRS) technologies, such as Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have revolutionized transcriptomic studies by enabling the comprehensive profiling of full-length RNA molecules1. These emerging LRS platforms allow precise quantification of isoform diversity and in-depth characterization of alternative splicing at bulk2,3 and single-cell4,5 levels, tasks that short-read methods struggle to accomplish6 (reviewed in ref. 7). Despite these breakthroughs, the Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium and other benchmarking efforts8–14 have shown that transcript identification, particularly the accurate detection of novel isoforms absent from current annotations, remains challenging. Factors such as sample quality, library preparation, RNA capture by the instrument, sequencing errors, and processing choices of analysis method may introduce biases and compromise the accuracy of transcript calls. This has motivated the development of tools and strategies for the quality evaluation of long-read transcriptomics data.

Comprehensive quality control (QC) is critical for identifying potential biases and technical artifacts in LRS data. Read-level QC tools, such as LongQC15, MinKNOW (https://github.com/nanoporetech/minknow_api), SQANTI-reads16, and PycoQC17, assess sequencing quality using metrics including read counts, length distribution, basecalling accuracy, adapter contamination, GC content, and coverage, among others. At the transcript level, SQANTI3 is widely used for evaluating the structural novelty of transcripts, integrating complementary data such as short-read alignments, CAGE peaks, poly(A) motifs and poly(A) sites, and machine learning to improve confidence and curation of the long-read-defined transcriptome18. Although these tools yield valuable insights into data quality, they do not offer an absolute ground truth and therefore cannot provide formal benchmarking capabilities.

Current long-read benchmarking frameworks have adopted and adapted existing short-read methods. For example, BUSCO, which measures completeness by comparing reconstructed transcriptomes against a database of conserved single-copy orthologous genes19, has been used by several LRS assessment studies8,20,21. However, BUSCO is limited in its capacity to evaluate isoform diversity, often misinterpreting alternatively spliced transcripts as gene duplications22,23. Another strategy is the utilization of spiked-in RNAs mimicking alternatively spliced transcripts, such as ERCCs24, SIRVs (https://lexogen.com/sirvs/), and Sequins25. These products offer LRS-tailored benchmarking capabilities but fail to capture the complexity of actual RNA samples, including RNA degradation patterns. As such, spike-in controls tend to overestimate the performance of LRS methods. Finally, several data simulation algorithms, including computational strategies to simulate novel transcripts, have been proposed26–28. While these tools can generate comprehensive ground truth datasets for method evaluation, they face challenges in faithfully simulating the complexity of the different sources of bias present in the data, and their results need to be complemented with additional benchmarking solutions. More recently, short-read-based junction-level benchmarking approaches such as MAJIQ-L have been proposed, using matched short-read data as a reference proxy to evaluate junction detection29. While such methods avoid synthetic spike-ins and can be applied to any expressed gene, they cannot assess transcript-level features such as transcription start and termination sites, and require the availability of matched short-read sequencing from the same samples.

Given these limitations, we developed the TUSCO benchmarking framework, an internal ground truth for LRS transcript identification that evaluates both real-sample artifacts and technological artifacts. The TUSCO benchmarking framework is centered on a curated TUSCO gene set of widely expressed genes characterized by highly conserved splice sites and the absence of evidence for alternative isoforms in the recount330 database. The TUSCO benchmarking framework offers several advantages. First, it detects false positives by identifying reconstructed transcripts mapped to the TUSCO gene set that deviate from established annotations, providing a measure of precision in transcript identification and revealing partial transcripts that may result from low sample quality. Second, it uncovers false negatives by leveraging the widespread expression of transcripts in the TUSCO gene set across human or mouse samples, thereby assessing the sensitivity and completeness of transcript detection methods. Third, via TUSCO-novel, the TUSCO evaluation assesses performance in detecting novel transcripts under conditions of incomplete annotation. Moreover, the TUSCO evaluation provides direct and effortless assessment, as it eliminates the need to manage spike-in control reagents or synthetic data. Our development satisfies a requirement in the LRS transcriptomics field by providing an internal standard for QC.

Results

Selection and validation of the TUSCO gene set for transcript benchmarking

We built the TUSCO gene set by enforcing cross-annotation single-isoform agreement and population evidence of splice/transcription start site (TSS) invariance (Fig. 1a). Candidate genes had to have one annotated isoform with identical exon-intron structure and strand across GENCODE and RefSeq (plus Matched Annotation from NCBI and EMBL-EBI [MANE] Select for human)31–33. We then screened for unannotated alternative splicing using large Illumina-based splice junction compendia: IntroVerse (human)34 and recount3 junctions (human and mouse)30, removing loci with intragenic novel junction support above predefined thresholds. TSS consistency was evaluated using refTSS (integrating FANTOM5, DBTSS, EPDnew, ENCODE)35; in human, we applied both exon-overlap and  ± 300 bp CAGE window checks, and in mouse, we required the exon-overlap rule for single-exon genes. A parameter sensitivity analysis is provided in Supplementary Fig. 1.

Fig. 1. Overview of TUSCO gene set selection and the TUSCO benchmarking framework.

Fig. 1

a Schematic illustration of the four-step pipeline for identifying single-isoform genes suitable for benchmarking. Candidate genes are cross-referenced across multiple annotation databases (MANE, RefSeq, and Ensembl for human; RefSeq and GENCODE for mouse), evaluated for potential alternative splice sites, and checked for broad, consistent expression in multiple tissues. b The TUSCO benchmarking framework leverages SQANTI3 structural categories to define benchmarking labels. True Positives (TP) are transcripts in the TUSCO gene set correctly identified as Reference Match (RM); Partial True Positives (PTP) are those that align to alternative categories (FSM_non-RM or ISM). False Negatives (FN) indicate undetected transcripts in the TUSCO gene set, whereas False Positives (FP) are transcripts with Novel in Catalog, Novel not in Catalog, Genic Intron, and Genic Genomic categories. c High-level workflow from sample preparation to computational analysis, highlighting how different biases (sample, protocol, platform, and transcript reconstruction) can affect the user-defined transcriptome and how different benchmarking approaches assess these biases.

Expression prevalence was defined using Bgee present/absent calls with stringent prevalence cutoffs (≥95% of retained tissues for human and ≥90% for mouse), reinforced by housekeeping genes from HRT Atlas36,37. Bgee integrates multiple expression data types (RNA-Seq, Affymetrix, in situ hybridization, ESTs) curated from healthy wild-type samples, including large resources such as GTEx38 alongside many smaller datasets, to provide a comparable reference of normal expression across species. To further favor widely expressed, splice-invariant loci, we applied AlphaGenome-based filters39 on two metrics: Expression, defined as the tissue-median reads per kilobase per million mapped reads (RPKM) across reference RNA-seq panels, and Splice ratio, defined as the ratio of AlphaGenome-predicted counts for unannotated versus annotated splice junctions. Because the AlphaGenome expression reference panels and absolute RPKM distributions differ between species, we used species-specific RPKM cutoffs rather than a single universal value. In practice, these thresholds still select the high-expression tail: based on AlphaGenome tissue-median RPKM across all available expression tracks, the median gene in the TUSCO gene set ranks at the 90.3rd percentile in human and the 96.5th percentile in mouse, supporting consistent detectability in our benchmarking datasets. Occasional edge cases underwent manual review by GENCODE curators.

This process yielded tissue-resolved TUSCO gene sets (human: 36–113 per tissue; mouse: 34–461) and universal cores (human n = 46; mouse n = 32) consistently expressed across all tested tissues (Fig. 2a, b). Compared with reference annotations and spike-ins, the TUSCO gene set spans endogenous size and structure ranges relevant for full-length detection: transcripts are comparable to RefSeq transcripts and longer than SIRVs/ERCCs; sets include both single- and multi-exon loci (Fig. 2c; Supplementary Fig. 2). Genes in the TUSCO gene set also show expression distributions similar to those of medium to highly expressed genes across tissues (Fig. 2d) and demonstrate, by AlphaGenome-predicted analysis, higher expression and lower novel-junction usage than all single-isoform genes annotated in both GENCODE and RefSeq (Supplementary Fig. 3), indicating that our filtering criteria were effective in selecting consistently expressed genes with minimal chances of containing hidden alternative isoforms. Together, these properties support the TUSCO gene set as an endogenous, broadly expressed, splice-invariant ground truth for benchmarking transcript identification and completeness in long-read data.

Fig. 2. Overview of the TUSCO gene set in human and mouse.

Fig. 2

a, b Tissue distribution of genes in the TUSCO gene set in human and mouse. Values denote retained genes per tissue; universal cores are expressed in ≥95% of human and ≥90% of mouse tissues. c Transcript length distributions in discrete bins (<600 bp, 600–1 kb, 1–2 kb, 2–4 kb, >4kb) for transcripts in the TUSCO gene set in human and mouse compared with all GENCODE transcripts from the corresponding species; ERCC and SIRV spike-ins are overlaid for reference. Transcript length is defined as the sum of non-overlapping exonic bases per transcript. d Cross-tissue median expression in log10 transcripts per million (TPM) for genes in the TUSCO gene set shown in green versus a background of the top 10,000 most highly expressed genes per tissue shown in gray, computed from GTEx for human and ENCODE for mouse. Violin plots show the distribution; embedded boxplots indicate the median as the center line, the interquartile range as the box, and whiskers extending to 1.5 × IQR; outliers are omitted. A small offset of 1 × 10−4 was added before log transformation. GENCODE versions: human v49 and mouse vM38. Source data are provided as a Source Data file.

TUSCO benchmarking framework in SQANTI3

The TUSCO benchmarking framework aims to evaluate four key aspects of long-read transcript reconstruction: (1) sensitivity: whether expressed transcripts are detected; (2) precision: whether detected transcripts match expected structures; (3) robustness: performance under real-world sample quality variations; and (4) novel isoform discovery capacity: the ability to identify true transcripts when annotations are incomplete.

The TUSCO benchmarking framework leverages the SQANTI3 framework to assign LRS transcript models to benchmarking labels (Fig. 1b). We applied consistent classification criteria for both the TUSCO gene set and SIRV spike-ins. True Positives (TP) include: (i) transcripts classified by SQANTI3 as Reference Match (RM); (ii) mono-exonic full splice match (FSM) transcripts with TSS and transcription termination site (TTS) within ±50 bp of annotated boundaries; and (iii) for transcripts longer than 3000 bp, FSM with TSS and TTS within ±100 bp. Partial True Positives (PTP) comprise remaining FSM, incomplete-splice matches (ISM), and mono-exon by intron retention calls, indicating deviations from reference boundaries that may reflect RNA degradation or incomplete sequencing. False Negatives (FN) represent reference units (genes in the TUSCO gene set or SIRV transcripts) that remain undetected, possibly due to insufficient sequencing depth or biases in RNA capture. False Positives (FP) are transcripts mapped to genes in the TUSCO gene set classified as Novel in Catalog (NIC), Novel not in Catalog (NNC), Genic Intron, Genic Genomic, or Fusion events, which may arise from sequencing artifacts, incorrect transcript reconstruction, or mapping errors.

These benchmarking labels allow for the calculation of TUSCO metrics (sensitivity, precision, and redundancy), which are provided in a dedicated TUSCO evaluation report. Seamlessly integrated with SQANTI3, the TUSCO benchmarking framework operates without requiring additional user input, enabling users to execute the TUSCO evaluation directly after completing SQANTI3 processing. The incorporation of this TUSCO evaluation report within SQANTI3 results in an enhanced, comprehensive, and unified QC resource that both describes the properties of the long-read data and provides rigorous and standard TUSCO metrics. Additionally, the TUSCO benchmarking framework enables a more complete assessment of the RNA sequencing workflow by capturing biases at every stage—from sample preparation to RNA extraction, library preparation, sequencing, and transcript reconstruction—unlike synthetic spike-ins such as SIRVs, which only gauge library preparation, sequencing, and transcript reconstruction (Fig. 1c). Note that the TUSCO benchmarking framework is agnostic to the strategy used by any transcript reconstruction algorithm. In principle, the TUSCO evaluation can assess all three major classes of transcriptome reconstruction tools: genome-guided de novo (no annotation provided), genome-guided discovery (annotation provided, with the goal of identifying novel transcripts), and genome reference-free (no genome or transcriptome prior information). The TUSCO evaluation can assess any such method as long as it is applied to human or mouse data, the two species for which the TUSCO gene set is currently defined.

Validation and comparison between the TUSCO benchmarking framework and SIRVs

To assess the TUSCO benchmarking framework as an endogenous benchmark, we first compared its performance to SIRVs using comprehensive datasets from the LRGASP Consortium. Three major library preparation workflows—direct RNA (dRNA) ONT, cDNA ONT, and cDNA PacBio—were applied to both human WTC11 induced pluripotent stem cells (iPSCs) and mouse embryonic stem (ES) cells. Transcript models were generated using the FLAIR40 pipeline and evaluated against either the TUSCO gene set or SIRVs (Fig. 3a).

Fig. 3. Performance comparison of the TUSCO gene set and SIRVs across library preparation workflows, including RNA degradation and false-negative analyses.

Fig. 3

a Dumbbell plots of sensitivity (Sn), redundant precision (rPre), positive detection rate (PDR), 1 − FDR, inverse redundancy (Redundancy−1), and non-redundant precision (nrPre) for human WTC11 and mouse ES processed with dRNA ONT, cDNA ONT, or cDNA PacBio libraries. The TUSCO gene set (green) and SIRVs (purple) lines summarize TUSCO metrics per dataset. b Bar charts showing the proportion of true positives (TP), partial true positives (PTP), false positives (FP), and false negatives (FN) for long-read-only (left) and hybrid long-plus-short-read (right) library protocols; bars show means ± 1 s.d., points indicate individual pipelines (n = 6 paired pipelines per species), and two-sided paired sign-flip permutation tests evaluate differences between TUSCO metrics and SIRVs: human TP p = 0.046, PTP p = 0.046, FP p = 0.077, FN p = 0.046; mouse TP p = 0.015, PTP p = 0.046, FP p = 0.877, FN p = 0.046. c RNA degradation experiment schematic and scatter plots relating RIN to fully reconstructed transcripts for the TUSCO gene set (top) and Sequins (bottom), with dataset-specific linear regression fits and shaded bands representing 95% confidence intervals. Pearson correlation: TUSCO gene set R = 0.810, p = 4.60 × 10−5; Sequins R = 0.075, p = 0.77; n = 17 samples. TP correspond to reference matches (RM); PTP include FSM or ISM transcripts that retain annotated junctions but shift 5'/3' ends. d False-negative (FN) genes in the TUSCO gene set across transcript reconstruction pipelines for human WTC11 and mouse ES; bars count FN genes present in exactly k of six long-read protocols plus one Illumina dataset, highlighting that none are absent from all datasets. Source data are provided as a Source Data file.

Across all library types and samples, TUSCO metrics closely mirrored those obtained with SIRVs in sensitivity, precision, and redundancy (Fig. 3a). To quantify this concordance, we calculated cosine similarity (cosim) scores between the two benchmarking approaches (Supplementary Table 1). Values approaching 1 indicate strong agreement in evaluating transcript reconstruction performance but do not necessarily reflect absolute performance quality. As a practical guide, values >0.95 indicate strong concordance, while values <0.90 would suggest meaningful benchmark divergence. The cDNA PacBio libraries, which were the best-performing experimental option according to LRGASP results, consistently achieved cosim values ranging from 0.9969 to 0.9992 in both human and mouse samples, demonstrating that the TUSCO benchmarking framework and SIRV benchmarks equally recognize the performance of high-quality data. Likewise, cDNA ONT, dRNA ONT, and hybrid long + short read library protocols (LS) also exhibited high agreement (cosim: 0.9493–0.9977), reinforcing the TUSCO benchmarking framework as a reliable internal benchmark comparable to established spike-in controls, while also revealing small differences between the two QC datasets. These findings underscore that TUSCO metrics should be interpreted as serving two main purposes: providing relative performance rankings of reconstruction methods within a dataset, and serving as indicators of sample-specific quality, where expected variance reflects real technical and biological heterogeneity.

The TUSCO benchmarking framework more stringently reflects sample preparation and sequencing platform

We asked whether endogenous controls (the TUSCO gene set) behave differently from synthetic spike-ins (SIRVs/Sequins) across pipelines and species. To answer this, we compared the fractions of true positives (TP), partial true positives (PTP), false positives (FP), and false negatives (FN) under benchmarking by the TUSCO gene set versus SIRVs for human and mouse datasets (Fig. 3b).

Across both species, spike-ins showed higher TP rates than TUSCO evaluation (two-sided paired sign-flip permutation test across six pipelines, n = 6 pairs; human TP p = 0.046, mouse TP p = 0.015). Closer inspection reveals that the lower TP rate for TUSCO evaluation (Fig. 3b) is driven by increases in PTP and FN rather than an explosion of FP: PTP and FN are higher under TUSCO evaluation (human PTP p = 0.046, FN p = 0.046; mouse PTP p = 0.046, FN p = 0.046), whereas FP are not elevated (human FP p = 0.077, mouse FP p = 0.877). We hypothesized that PTP arises from RNA degradation causing truncated molecules, while FN reflects protocol limitations.

Because TUSCO gene set transcripts are endogenous, they are exposed to RNA integrity, extraction and library biases, and depth limitations that intact spike-ins do not experience. We tested the PTP hypothesis directly by applying TUSCO evaluation to ONT cDNA samples spanning a broad range of RNA integrity number (RIN) values (Fig. 3c)41. The TUSCO evaluation revealed increased PTP at lower RIN, and the fraction of fully reconstructed transcripts in the TUSCO gene set (TPTP+PTP) strongly correlated with RIN (Pearson correlation, n = 17, R = 0.810, p = 4.60 × 10−5), whereas Sequins showed no correlation (R = 0.075, p = 0.77). As secondary validation, we manually inspected representative FP calls and consistently found technical artifacts unsupported by long- and short-read evidence (Supplementary Notes 1–4).

Together, these results show that TUSCO evaluation complements spike-in benchmarks by revealing quality- and protocol-linked failure modes (PTP/FN) that spike-ins tend to mask, while maintaining comparable robustness against FP.

We further examined false negative (FN) genes in the TUSCO gene set across diverse reconstruction pipelines and sequencing technologies within the LRGASP Consortium dataset. We investigated whether evidence of expression existed for all genes in the TUSCO gene set in any of the LRGASP biological samples. We calculated the number of datasets where genes in the TUSCO gene set were detected by at least one mapped read. We found that all genes in the TUSCO gene set were detected in at least one dataset (Fig. 3d), confirming their expression in the biological samples and their true false-negative status when not detected. Consistent with our hypothesis that FN reflects protocol limitations, FN rates across these same LRGASP datasets are primarily determined by platform and protocol choice (PacBio cDNA lowest; ONT higher; Supplementary Fig. 4).

We concluded that by reflecting real-sample pitfalls, including RNA degradation, library preparation biases, and variable read depth, the TUSCO benchmarking framework provides a realistic, endogenously anchored benchmark strategy that surpasses synthetic controls such as SIRVs and Sequins.

We compared TUSCO evaluation with MAJIQ-L, a junction-level benchmarking approach using matched short-read data as a proxy reference29. Among LRGASP library preparation workflows, cDNA-PacBio consistently ranked best under both frameworks, while cDNA-ONT and dRNA-ONT rankings diverged (Supplementary Fig. 5; Supplementary Table 2). This divergence reflects complementary quality dimensions: TUSCO evaluation is sensitive to RNA integrity (detecting partial transcripts), whereas MAJIQ-L captures each method’s junction-level sensitivity against a common short-read-derived splice graph. For example, dRNA-ONT achieves high TUSCO evaluation sensitivity (75.6%) despite only 760K reads, but shows high MAJIQ-L false-negative rate (73.5%) due to limited depth. Conversely, cDNA-ONT has 16.6M reads and better MAJIQ-L junction sensitivity (37.0%) but lower TUSCO evaluation sensitivity (59.5%) due to degraded RNA. These results demonstrate that TUSCO evaluation and MAJIQ-L provide complementary assessments of long-read transcriptome reconstruction quality. Overall, pipeline rankings were similar between TUSCO evaluation and MAJIQ-L benchmarks, although MAJIQ-L results indicated lower overall performance for sensitivity and FN rate, while TUSCO evaluation highlighted FDR limitations unseen by MAJIQ-L.

TUSCO-novel evaluates the capacity of reconstruction tools to detect unannotated isoforms

Since genes in the TUSCO gene set are well-annotated, they assess the capacity of long-read methods to detect known transcripts but do not reveal the accuracy of these methods in detecting novel transcripts, which often correspond to alternative isoforms of annotated genes12,42,43. To address this shortcoming, we developed a strategy to use the TUSCO benchmarking framework to assess novel transcript calls. For each single-isoform gene in the TUSCO gene set, we modified its transcript model in the reference annotation by creating an entirely different splice pattern with structurally plausible artificial donor and acceptor sites at all junctions and supplying the artificially modified transcript as the sole reference. This effectively updates the annotation of the single-isoform genes to reflect a nonexistent transcript while omitting the real, expressed transcript (Fig. 4a). Isoform reconstruction tools with the capacity to detect novel transcripts should call the true isoform despite a misleading annotation. We applied TUSCO-novel to Bambu (v3.8.3), StringTie2 (v2.2.3), FLAIR, and the Iso-Seq + SQANTI3 machine learning (ML) pipeline across PacBio and ONT CapTrap and cDNA libraries in human and mouse (LRGASP Consortium). Performance with the native reference annotation was contrasted with that under TUSCO-novel (Fig. 4b).

Fig. 4. TUSCO-novel benchmark for evaluating multi-exon transcript discovery.

Fig. 4

a Schematic of the TUSCO-novel approach. The native multi-exon isoform in the TUSCO gene set is removed from the annotation and replaced with a simulated transcript model carrying structurally plausible artificial donor and acceptor splice sites. The modified annotation is used as input to reconstruction pipelines to assess their ability to reconstruct the true (but now unannotated) isoform. b Dumbbell plots of six TUSCO metrics (Sn, nrPre, Redundancy−1, 1 − FDR, PDR, rPre) comparing native-reference TUSCO evaluation (circles; solid) versus TUSCO-novel (triangles; dashed) across Bambu, StringTie2, FLAIR, and Iso-Seq + SQANTI3 machine learning (ML) for PacBio and ONT CapTrap/cDNA libraries in human (WTC11) and mouse (ES). Bambu and StringTie2 perform well with the true reference annotation but drop sharply under TUSCO-novel, indicating strong reliance on annotation. FLAIR partially reconstructs true isoforms in the novel context but with elevated false positives and false negatives. Iso-Seq + SQANTI3 ML best controls false positives under TUSCO-novel (0 FP across human and mouse), consistent with reference-free discovery followed by junction-support filtering. Lower sensitivity and precision reflect the lack of TSS/TTS adjustment in the Iso-Seq pipeline, yielding partial true positives. Per-read coverage analysis (Supplementary Fig. 6) supports that these PTP calls are enriched for truncated reads, consistent with boundary shifts rather than incorrect splice junction chains. Source data are provided as a Source Data file.

The results show a clear split between reference-guided and more reference-agnostic behavior (Fig. 4b). Under native annotation, Bambu achieved near-perfect performance (F1 score: human 99.5%, range 97.9–100.0; mouse 98.0%). Under TUSCO-novel, F1 dropped to 37.3% (range 30.5–53.1) in human and 41.0% (range 36.9–49.3) in mouse, indicating strong annotation dependence. StringTie2 showed similar declines: F1 fell from 69.8% (range 57.1–95.7) to 38.0% (range 31.8–46.8) in human and from 76.2% (range 70.4–92.0) to 24.4% (range 20.3–26.2) in mouse. FLAIR exhibited moderate decreases (F1: human 72.0–44.2%; mouse 69.3–42.1%) but high protocol-dependent variability (range 19.5–68.2 under TUSCO-novel). By contrast, Iso-Seq + SQANTI3 ML (n = 1 per species) showed a stable F1 regardless of annotation (human: 54.5–57.1%; mouse: 51.0–48.9%) and reported 0 FP under TUSCO-novel (Supplementary Table 3).

We attribute the stability of Iso-Seq + SQANTI3 ML to combining reference-free Iso-Seq transcript discovery with SQANTI3 filtering that removes transcript models lacking junction support from complementary data, improving precision for novel transcripts. However, the pipeline lacks TSS/TTS refinement against the reference, yielding more PTP even when the internal junction chain is correct (Fig. 4). Because TP and PTP largely differ at transcript boundaries (TSS/TTS) rather than in internal junctions, we asked whether PTP calls reflect truncated molecules that do not span the full exonic body of the true isoform. We quantified per-read coverage f=LreadLexonic for each read supporting TP or PTP calls in the TUSCO gene set, where Lread is the observed Iso-Seq full-length non-chimeric (FLNC) read length (PacBio cDNA) and Lexonic is the total exonic length of the true reference isoform. Thus, f ≈ 1 indicates near full-length support, whereas low f indicates 5’/3’ truncation expected under RNA degradation or incomplete capture (and f can exceed 1 when a read is longer than the reference isoform). In human, TP coverage averaged 84.30% ± 26.52% (median 98.3%, n = 4892), while PTP averaged 54.68% ± 43.09% (median 31.7%, n = 633); in mouse, TP averaged 79.73% ± 25.02% (median 94.7%, n = 8920) and PTP 62.10% ± 30.50% (median 53.7%, n = 2564). These patterns are consistent with partial transcript coverage effects (Supplementary Fig. 6).

Collectively, these findings indicate that TUSCO-novel serves as a controlled stress test for unannotated isoform reconstruction in the presence of misleading reference models. It complements TUSCO evaluation under a correct reference annotation by distinguishing tools that strongly base their transcript calls on the supplied annotation from those capable of reference-agnostic, data-driven transcript reconstruction, a key requirement for transcriptomics in emerging models, under-characterized tissues, or disease contexts.

The TUSCO benchmarking framework helps assess replication in PacBio long-read transcriptomics

With increasing throughput of LRS technologies, we asked whether the TUSCO benchmarking framework could assist in experimental design decisions. Specifically, we used the TUSCO evaluation to assess how biological replicates affect sensitivity and precision. Five cDNA libraries were prepared from mouse brain and kidney and sequenced on the PacBio Sequel IIe (approximately 4–5 million FLNC reads per sample; Supplementary Table 4). TUSCO metrics were computed in smart intersection (consensus) mode, i.e., considering the transcripts consistently detected across an increasing number of replicates to detect consistently-expressed transcripts. All possible combinations for a given replicate number were computed, and the mean and 95% confidence intervals were computed (Fig. 5a). FLAIR was used as the transcript reconstruction method.

Fig. 5. Replication gains and agreement between universal and tissue-specific TUSCO gene sets in mouse PacBio cDNA data.

Fig. 5

a Multi-sample (“consensus”) TUSCO evaluation: five biological replicates per tissue (brain, kidney) were sequenced; SQANTI3 assigns transcript models; TUSCO metrics are computed on the transcript set common to the selected replicates. b Performance versus replication (1 → 5 replicates) for brain and kidney. Bars show Sensitivity (Sn), non-redundant Precision (nrPre), inverse Redundancy (Redundancy−1), 1 − FDR, PDR, and redundant Precision (rPre). Red callouts mark percentage-point changes relative to one replicate; note the large FDR reduction from one to two replicates, especially in kidney. c Benchmark agreement between the universal mouse set (∣U∣ = 32) and tissue-specific panels requiring only tissue-consistent expression (brain: 63 genes; kidney: 44 genes). Values shown in the dumbbell plots are computed per sample and then averaged across biological replicates within each tissue and panel type. Source data are provided as a Source Data file.

With a single replicate of approximately 5 million reads, most transcripts in the TUSCO gene set were reconstructed, but precision was modest: brain sensitivity was 81.3% (95% CI, 74.5–88.0) with PDR 93.1% (95% CI, 89.9–96.4) and non-redundant precision 67.1% (95% CI, 61.4–72.8; FDR 32.9%); kidney sensitivity was 71.9% (95% CI, 66.4–77.4) with PDR 90.6% (95% CI, 85.9–95.4) and non-redundant precision 65.2% (95% CI, 61.0–69.4; FDR 34.8%). Requiring concordant detection in at least two replicates markedly reduced false positives: non-redundant precision increased to 87.8% (95% CI, 86.5–89.1) in brain and 81.0% (95% CI, 79.2–82.8) in kidney, while sensitivity was 81.3% (95% CI, 79.4–83.1) and 70.6% (95% CI, 67.8–73.5), respectively. With three to five replicates, brain precision stabilized near 88.7–90.0%, at sensitivity of 83.1–84.4%. In kidney, sensitivity increased with replicate number—from 74.7% (three replicates) to 78.1% with five—while precision plateaued near 79.7–80.6%. We attribute the larger gain in kidney sensitivity at three versus two replicates to boundary refinement through replicate consensus, whereby additional replicates help refine TSS/TTS and junction boundaries, converting borderline partial matches into exact TUSCO gene set-concordant structures.

Collectively, these results indicate that TUSCO evaluation effectively reveals the effect of replication on transcript detection accuracy. The largest precision gain occurred between one and two replicates, with additional replicates (3–5) yielding incremental improvements at this sequencing depth (Supplementary Table 4). Thus, for this experimental setup (mouse brain/kidney, PacBio cDNA, FLAIR), at least two biological replicates substantially improve precision, with diminishing returns thereafter; additional replicates can still provide modest sensitivity gains (notably in kidney) via consensus boundary refinement (Fig. 5b). By contrast, analyses of SIRV spike-in controls yielded highly consistent concordance metrics across 1–5 replicates, providing little guidance on replicate number optimization (Supplementary Fig. 7).

Universal and tissue-specific TUSCO gene sets give similar results

Previous analyses were conducted with the universal TUSCO gene set, which is rather small (n = 32). We built tissue-specific TUSCO gene sets to expand locus coverage and tested whether conclusions persisted (Fig. 2a, b). Tissue-specific TUSCO gene sets were constructed using the same single-isoform and invariance filters, requiring only consistent expression in the focal tissue; additionally, tissue-specific TUSCO gene sets used a stricter tissue-level AlphaGenome-predicted screen (human RPKM > 2.0 and splice ratio < 0.001; mouse RPKM > 1.5 and splice ratio < 0.001), as detailed in Methods. Using the same data, the brain (63 genes) and kidney (44 genes) panels produced TUSCO metrics that were minimally lower than those of the universal TUSCO gene set in all cases, with cosine similarity between universal and tissue results of 0.9999 for brain and 0.9992 for kidney (Fig. 5c). We concluded that the universal TUSCO gene set is a suitable benchmarking option for human and mouse samples, while tissue-specific TUSCO gene sets provide slightly higher coverage to fine-tune results.

Discussion

The reconstruction of full-length transcriptomes using LRS technologies has revolutionized our understanding of transcript diversity. Yet, this advancement brings with it a pressing need for accurate and realistic benchmarking strategies. Our study demonstrates that synthetic spike-in controls, while useful, present significant limitations when used as stand-alone benchmarks. Tools like SIRVs and Sequins offer a controlled environment to assess transcript identification, but they fail to replicate key biological and technical variables that are intrinsic in real samples—such as RNA degradation, extraction bias, and natural variability across tissues. These shortcomings may result in artificially inflated performance estimates and obscure technology-specific limitations that are critical in challenging experimental contexts, such as clinical or disease-focused transcriptomic studies.

To address these gaps, we developed the TUSCO benchmarking framework, built around a highly curated TUSCO gene set of single-isoform endogenous genes. These genes are robustly and consistently expressed across tissues, offering a more biologically relevant and nuanced ground truth to evaluate transcriptome reconstruction performance. The TUSCO benchmarking framework surpasses synthetic benchmarks by reflecting the same degradation, extraction biases, and length distribution characteristics of native RNAs, particularly in the medium-to-long transcript length range, which is underrepresented in many synthetic datasets. This results in a more representative assessment of sequencing performance.

When compared with short-read-based benchmarking approaches, we note that the TUSCO evaluation, which relies on a true annotation-based ground truth derived from full-length transcript data rather than short-read approximations, enables transcript-level assessment of both PTP and FN, requires no additional sequencing costs (as no matched short-read data are needed), and uniquely captures TSS and TTS accuracy through the PTP category. In contrast, evaluation using short-read-based approaches such as MAJIQ-L captures junction-level information that complements transcript-level assessment. In cDNA-based long-read protocols, reverse transcription proceeds from the poly-A tail, so 5’-proximal splice junctions in longer, multi-exon transcripts are less likely to be covered, and short-read junction data can expose these coverage gaps. However, such approaches require matched short-read data, entailing additional sequencing effort. We anticipate that, as LRS increases in throughput and becomes more affordable, the need for matched short-read data will diminish.

Here we demonstrate that the TUSCO benchmarking framework effectively reveals performance differences among LRS technologies and associated analysis methods, highlighting its utility as a benchmarking framework. In the datasets evaluated in this study, the observed performance differences can largely be attributed to variation in sequencing depth, read accuracy, and sample quality. However, this work is not intended to provide a systematic comparison of LRS technologies, which would require a more balanced and controlled experimental design. Rather, our results illustrate that TUSCO can reliably detect performance differences and help interpret them in the context of the intrinsic properties of the analyzed datasets.

Moreover, because the TUSCO gene set is endogenous and tissue-resolved, it enables benchmarking across a wide variety of experimental settings and sample qualities. This versatility makes the TUSCO benchmarking framework particularly valuable for several practical applications. For example, we show that the TUSCO benchmarking framework can inform trade-offs between sequencing depth and number of replicates, a common dilemma in experimental design. The TUSCO benchmarking framework also enables structured evaluation of novel transcript detection via TUSCO-novel. While the current implementation focuses on novel splice junctions, the method can be easily adapted to assess alternative TSS, TTS, or transcript disambiguation by modifying the underlying GTF annotations. We demonstrated the utility of the TUSCO benchmarking framework in both human and mouse samples, with TUSCO gene sets adaptable to tissue-specific or cross-tissue applications. This flexibility ensures that researchers working on tissues not directly included in the resource can still leverage the TUSCO benchmarking framework. If the community widely adopts this approach, it could be extended to additional species where sufficient transcriptomic data exist to define a reliable TUSCO gene set.

Despite its advantages, the TUSCO benchmarking framework also has limitations. It is inherently restricted to single-isoform loci and thus cannot evaluate transcript reconstruction accuracy within highly complex multi-isoform genes, a domain where synthetic spike-ins, which include alternatively spliced genes, remain useful. However, our comparison with SIRVs shows that true positive (TP) rates were similar across both approaches, suggesting similar capacity for assessing the identification of the right junction chain. Synthetic controls failed to capture sample-specific variation in pre-sequencing biases, as reflected by greater discrepancies in precision and false negative rates, which are better illuminated by the TUSCO gene set. As sequencing error rates continue to drop, biases introduced during RNA extraction and library preparation are likely to become the dominant limiting factors. In this context, the endogenous nature of the TUSCO gene set offers a crucial advantage.

Moreover, while genes in the TUSCO gene set are selected for consistent expression across healthy tissues, their expression profiles may be altered in unusual biological conditions such as severe disease states. In these contexts, specific genes in the TUSCO gene set might be downregulated or absent, potentially inflating the apparent false negative rate. Consequently, users should interpret TUSCO metrics in disease samples with appropriate caution, using them primarily as relative indicators for comparing pipeline performance on the same dataset rather than as absolute measures of sensitivity.

Another important consideration is the potential for the future discovery of alternative isoforms in currently annotated single-isoform genes. As GENCODE annotations evolve with the likely addition of new transcripts across the genome, some current genes in the TUSCO gene set may eventually reveal alternative isoforms. While the current TUSCO gene set is based on extensive sequencing evidence from multiple large-scale databases, we anticipate that regular updates and cross-referencing with evolving annotations will be necessary to maintain its accuracy.

Looking forward, the TUSCO benchmarking framework could be extended beyond strictly single-isoform genes. As our understanding of tissue-specific and condition-specific isoform expression patterns matures, a complementary approach could leverage dominantly expressed transcripts at multi-isoform loci, where one isoform represents >95% of expression in specific tissues or conditions. While implementing such an approach would present additional complexity in defining expression thresholds and handling tissue-specificity, it would substantially expand the number of benchmark loci and could provide continuity even as annotations evolve. The modular design of the TUSCO benchmarking framework within SQANTI3 would facilitate such extensions, allowing the community to adapt the framework to emerging transcriptomic knowledge while maintaining the core principle of using endogenous, well-characterized transcripts as internal controls.

In summary, the TUSCO benchmarking framework provides a powerful, biologically grounded framework to evaluate the performance of long-read transcriptomics across real-world conditions. By bridging the gap between synthetic controls and complex biological samples, it offers a more realistic benchmark to guide method development, protocol optimization, and experimental design.

Methods

TUSCO gene set selection pipeline

The TUSCO gene set was constructed with a single, reproducible pipeline that integrates multi-source annotations, population splicing and TSS evidence, and broad expression compendia, and then enforces strict “no evidence for alternative isoforms” criteria. For human, we retrieved and harmonized GENCODE44 v49, RefSeq45 GRCh38.p14 (GCF_000001405.40), and MANE33 v1.4 annotations together with an Ensembl BioMart Ensembl ↔ RefSeq/NCBI mapping46; for mouse, we used GENCODE vM38 and RefSeq GRCm39 with the corresponding Ensembl BioMart mapping. After normalizing chromosome labels where needed, we identified genes that are single-isoform in each source and required exact identity of exon-intron structure and strand across all GTFs; only those cross-annotation matches were retained and propagated with synchronized mapping subsets.

We next screened for population-level alternative splicing and TSS inconsistencies. Splice evidence was derived from recount3 junction BED files30. For each multi-exon gene, we computed the mean coverage of its annotated junctions, μ, and treated an intragenic, strand-consistent novel junction as evidence of an alternative isoform if its read support exceeded T = αμ, where α is species-specific (human: 0.01; mouse: 0.05); these values were calibrated so that >95% of final TUSCO gene set genes fall below the threshold. We additionally required at least one supporting read to avoid undefined ratios when μ = 0 (extremely low-coverage loci); in practice, αμ dominates because μ is typically large (e.g., among multi-exon genes in the final TUSCO gene sets, min(μ)=5.1×105 in human and 9.8 × 105 in mouse). Novel junctions shorter than 80 bp were ignored. TSS evidence was assessed with refTSS v4.135. Our TSS filtering primarily focused on single-exon genes: in human, we required single-exon loci to pass both an exon-overlap check (exactly one refTSS interval must overlap the TSS exon and contain the transcript’s TSS coordinate) and a CAGE window check centered at the transcript TSS that removed loci with more than one peak in a ±300 bp window. In mouse, TSS screening was restricted to single-exon genes using the exon-overlap rule. For human only, we further removed genes flagged by IntroVerse_80 as having widespread novel isoform usage in population data (>80% of samples)34.

Expression-based universality and tissue sets were then established using a hybrid approach that leverages multiple expression data sources. For human, we employed a dual-source strategy: Bgee36 present/absent calls for identifying universal genes, and GTEx38 for tissue-specific gene validation. To achieve universal gene identification, we analyzed Bgee data by retaining calls of “gold quality” and “silver quality” and restricting analysis to tissues with at least 25,000 distinct expressed genes to ensure robust tissue representation. Genes were called universal if marked as “present” in at least 95% of these retained tissues; sensitivity analysis confirmed that relaxing this threshold sharply increases false-negative rates. For tissue-specific sets, we used GTEx v10 data across all 54 available GTEx tissues. Within each tissue, genes were required to satisfy dual expression criteria: (i) prevalence ≥95% of samples with TPM  > 0.1, and (ii) median expression ≥1.0 TPM across all tissue samples.

For mouse, we used Bgee with the same quality filters (“gold quality” and “silver quality”), minimum tissue gene coverage (25,000), and called universal genes at a 90% prevalence threshold across retained tissues. Anatomical entity IDs were mapped to tissue names, and per-tissue gene sets were constructed using those IDs. HRT Atlas housekeeping genes meeting our single-isoform criteria were included to ensure universal expression37. For mouse tissue-specific sets, we optionally augmented expression data with ENCODE RNA-seq when available, requiring median RPKM ≥ 1.0 and prevalence ≥ 95% at RPKM > 0.1.

To ensure consistency at the population level beyond presence/absence, we applied an AlphaGenome-based screen with species- and class-specific thresholds39. AlphaGenome provides two key metrics for each gene across tissues: (i) Expression, defined as the tissue-median RPKM across reference RNA-seq panels, and (ii) Splice ratio, defined as the ratio of AlphaGenome-predicted counts for unannotated versus annotated splice junctions. We defined splice purity as the fraction of tissues where the splice ratio falls below a specified threshold:

SP=t∈T:rt<τt∈T:rtobserved 1

where rt denotes the splice ratio in tissue t, τ is the species- and class-specific splice ratio threshold, and T the set of tissues with available data. This metric quantifies how consistently a gene maintains its annotated splicing pattern across tissues, with higher values indicating lower evidence of alternative splicing.

For human universal genes, we required: (i) single-exon genes to have median RPKM ≥ 1.0 and ≥95% of tissues with RPKM > 0.05; (ii) multi-exon genes to have median RPKM ≥ 1.0 and ≥95% of tissues with RPKM > 0.05; and (iii) splice purity ≥ 95% with a splice ratio threshold < 0.001. For mouse universal genes, the criteria were: (i) single-exon genes with median RPKM ≥ 0.5 and ≥90% of tissues with RPKM > 0.01; (ii) multi-exon genes with median RPKM ≥ 0.25 and ≥90% of tissues with RPKM > 0.005; and (iii) splice purity ≥ 90% with a splice ratio threshold < 0.01. When generating tissue-specific sets, we applied stricter tissue-level AlphaGenome-predicted screens: human required RPKM > 2.0 and splice ratio < 0.001 in the tissue of interest, while mouse required RPKM > 1.5 and splice ratio < 0.001. A small, curated exclusion list was applied at the end to remove remaining edge cases. The pipeline writes synchronized mapping subsets alongside each GTF subset and records, for every removed gene, the first filter responsible for the removal.

Parameter sensitivity analysis

To validate the robustness of the TUSCO gene set selection pipeline, we conducted systematic sensitivity analyses across four key parameter categories. For TSS/CAGE window size, we evaluated windows from 100 bp to 500 bp using refTSS peak data, measuring sensitivity (fraction of annotated TSSs captured) against inter-gene collision rate. For expression thresholds, we swept four configurations (strict to lenient) using LRGASP evaluation data and benchmarked false-negative rates against SIRV controls, selecting parameters that balance gene count against acceptable FN increase. Species-specific thresholds were calibrated to target comparably high-expression percentiles (90th–97th percentile) based on AlphaGenome-predicted tissue-median RPKM in both human and mouse. For splice junction filtering, we used AlphaGenome-predicted junction read counts to verify that the relative threshold T = αμ (where μ is mean annotated-junction coverage and α is 0.01 for human, 0.05 for mouse) effectively discriminates genes in the TUSCO gene set from excluded single-isoform genes.

Sample preparation, library construction, and computational pipeline parameters

Long-read RNA-seq datasets used in this study were derived from multiple sources to enable a robust evaluation of transcriptome reconstruction using the TUSCO benchmarking framework. Data were collected from both public repositories and our laboratory to provide diverse experimental conditions.

The LRGASP Consortium dataset comprises LRS data generated using ONT and PacBio platforms. These data, which include samples from human WTC11 iPSC and mouse ES cells, were retrieved from the Synapse repository (https://www.synapse.org/Synapse:syn25007472/wiki/608702). The RNA degradation experiment data were obtained to assess the impact of sample quality on transcript reconstruction. FAST5 and FASTQ files for these experiments are available via the European Nucleotide Archive (ENA) under accession PRJEB53210, and additional FASTQ files from a post-mortem brain sample are accessible from the European Genome-Phenome Archive (EGA) under accession EGAS00001006542.

Mouse kidney dataset (PacBio Iso-Seq, 3-month-old cohort). Kidneys were dissected from five male C57BL/6J wild-type mice (3-month-old). Brain samples used in this study were obtained from the same cohort of 3-month-old male C57BL/6J mice. Tissues were homogenized using FastPrep, and total RNA was extracted with the Maxwell 16 LEV simplyRNA Purification Kit. cDNA synthesis and sample barcoding followed PacBio Iso-Seq recommendations using the NEBNext Single Cell/Low Input cDNA Synthesis & Amplification kit. Five barcoded Iso-Seq libraries (K31–K35) were prepared with the Iso-Seq SMRTbell prep kit 3.0 and sequenced on a PacBio Sequel IIe. This dataset is deposited in the European Nucleotide Archive under accessions PRJEB85167 (released 2025-01-23, brain) and PRJEB94912 (released 2025-09-30, kidney).

Mouse-specific laboratory procedures. RNA quantity and integrity were assessed prior to library preparation (e.g., Agilent Bioanalyzer to obtain RIN values). Library preparation strictly followed the PacBio Iso-Seq workflow: full-length cDNA amplification and barcoding, SMRTbell construction with the Iso-Seq SMRTbell prep kit 3.0, and sequencing on the Sequel IIe platform. Run and enzyme/incubation parameters followed the manufacturer’s recommendations for Iso-Seq libraries. Lexogen SIRV-Set 1 spike-in controls were added at 3% (E0 mix to brain samples and E1 mix to kidney samples).

For PacBio datasets, raw subreads were processed with the PacBio Iso-Seq analysis pipeline (v4.3.0) to generate FLNC reads.

For all datasets, alignment of reads to the reference genome was conducted using minimap2 (version 2.28). The reference genome FASTA assemblies were GRCh38/hg38 primary assembly for human and GRCm39/mm39 for mouse, and alignments were performed with parameters specifically tuned for long-read RNA-seq data. Presets were stratified by platform: PacBio Iso-Seq (HiFi) used -ax splice:hq, while ONT cDNA or direct RNA used -ax splice (with -k14). Key settings included the use of the -ax splice option for spliced alignment where applicable, –secondary=no to suppress secondary alignments, and -C5 to penalize ambiguous mappings. The flag -uf was enabled to guide canonical splice-site discovery and strand/orientation inference during spliced alignment, and the maximum intron length was set to 2,000,000 nucleotides via the -G 2000000 parameter. Platform-specific options recommended by minimap2 for PacBio and ONT long-read RNA-seq were applied where relevant.

For transcript QC and classification, the SQANTI3 pipeline (version 5.4) was applied. In general, the pipeline uses reference annotation, genome, and CAGE peak data specific to the organism under study. The workflow is divided into three major stages: a QC stage using sqanti3_qc.py, a filtering stage using sqanti3_filter.py that applies criteria defined in a default JSON configuration file to remove low-confidence transcript models, and a rescue stage using sqanti3_rescue.py to rescue valid transcripts that may have been overly filtered18.

We ran the pipeline with common options such as –skipORF to bypass open reading frame prediction and provided coverage files—STAR-derived splice-junction files (SJ.out.tab) generated from Illumina short-read alignments to support splice-junction validation. Illumina reads were mapped with STAR (v2.7.11b) in two-pass mode against GRCh38/hg38 or GRCm39/mm39 using standard RNA-seq settings, and the resulting junctions were supplied to SQANTI3.

TUSCO-novel simulation of structurally plausible isoforms

To construct the TUSCO-novel challenge, we generated misleading transcript models for single-isoform TUSCO gene set genes by modifying splice junctions while preserving canonical splice-site dinucleotides. For multi-exon genes, we shifted each annotated donor and acceptor coordinate within the original intron boundaries while searching the genomic sequence for canonical motifs (plus strand: GT donor and AG acceptor; minus strand: AC donor and CT acceptor in reference orientation). If a canonical motif could be found, shifts were constrained to be non-trivial (minimum 5 bp) and preferentially selected within a local search window (100 bp), expanding to the full intron when needed. For single-exon genes, we introduced one synthetic intron by selecting two positions separated by at least 50 bp and, when possible, snapping each boundary to a nearby canonical motif within 20 bp.

These simulated isoforms are structurally plausible only in the limited sense of matching canonical splice-site motifs; they do not model or guarantee functional splicing. In particular, we do not enforce preservation of splice enhancer/silencer sequences, branch point or polypyrimidine tract features, GC content, or exon-length distributions.

Transcriptome reconstruction with FLAIR

We reconstructed transcriptomes for three datasets—(i) the LRGASP Consortium project (Oxford Nanopore [MinION R10.4.1] and PacBio Sequel II), (ii) the RNA degradation dataset (Oxford Nanopore direct RNA [GridION R9.4.1]), and (iii) the replicated Mouse Brain & Kidney dataset (PacBio Sequel IIe and Oxford Nanopore cDNA [R9.4.1])—using FLAIR v2.2.040. Briefly, raw LRS data were first converted to FASTQ format. Reads were then aligned to the reference genome (GRCh38/hg38 for human; GRCm39/mm39 for mouse) using FLAIR’s align functionality, which internally wraps minimap2 to produce BED-formatted alignments. Alignment files were refined with the FLAIR correct subcommand to adjust splice junctions and transcript boundaries based on an external reference GTF annotation. The collapse step, which groups corrected reads into isoforms, was run with the same reference annotation provided via the --gtf flag, with Illumina short-read support applied only in the LRGASP long+short (LS) dataset.

For the TUSCO-novel evaluation, transcript models were additionally generated using Bambu (v3.8.3) and StringTie2 (v2.2.3) with default parameters.

Junction-level benchmarking with MAJIQ-L

For junction-level benchmarking, we applied MAJIQ-L to compare long-read transcript reconstructions against matched short-read data from the LRGASP WTC11 human samples29. Short-read Illumina data were first processed through MAJIQ v3 (majiq build with the GENCODE v49 annotation and --strandness NONE; all other build parameters at default) to construct splice graphs, which served as the reference for junction detection. Long-read transcript models (GTF files from each reconstruction pipeline) were evaluated with voila lr, and all junctions from all quantified local splice variations were included in the analysis. Default MAJIQ-L parameters were used throughout: PSI threshold of 0.05, exact junction matching (0 bp fuzziness at both 5’ and 3’ ends), and minimum read support of 1.

Junctions were classified as known (present in annotation) or novel (absent from annotation). For each category, we computed: True Positives (TP): junctions detected by both long reads and MAJIQ; False Positives (FP): junctions detected by long reads but absent from MAJIQ; and False Negatives (FN): junctions present in MAJIQ but not detected by long reads. True Negatives cannot be computed because unexpressed genes cannot be distinguished from expressed but undetected genes. From these counts, we calculated Sensitivity as (TP-known + TP-novel)/(TP-known + TP-novel + FN-known + FN-novel), False Negative rate as (FN-known + FN-novel)/(TP-known + TP-novel + FN-known + FN-novel), and False Discovery Rate as 100 × (N − TP)/N where N is the total number of long-read junctions detected.

TUSCO metrics and SIRV metrics calculation

We implemented unified TUSCO metrics and SIRV metrics based on SQANTI3 structural classifications. For the TUSCO gene set, we curated a single-isoform GTF reference (“TUSCO gene set reference set”) comprising genes with no evidence of alternative splicing. For SIRVs, we used the provided SIRV reference annotations. Using the SQANTI3 classification file, we restricted evaluation to reconstructed transcripts mapping to either genes in the TUSCO gene set (matched by the predominant identifier present in the classification: Ensembl, RefSeq, or gene symbol) or SIRV chromosomes.

For benchmarking by the TUSCO gene set and by SIRVs, we applied consistent classification criteria:

True positives (TP)

Transcripts meeting any of the following criteria:

  • SQANTI3 structural category “reference_match”

  • Full-splice_match transcripts with single exons where TSS and TTS are within  ± 50 bp of the annotated boundaries.

  • For transcripts with reference length >3000 bp: full-splice_match transcripts where TSS and TTS are within ±100 bp of the annotated boundaries.

Partial true positives (PTP)

Transcripts classified as:

  • Full-splice_match (excluding those already classified as TP)

  • Incomplete-splice_match.

  • Mono-exon_by_intron_retention (subcategory).

False positives (FP)

For genes in the TUSCO gene set, transcripts labeled as novel_in_catalog, novel_not_in_catalog, genic (including genic_intron and genic_genomic), or fusion. For SIRVs, any transcript mapping to SIRV chromosomes that is neither TP nor PTP.

False negatives (FN)

Genes in the TUSCO gene set (or SIRV transcripts) with no observed transcript in any structural category.

Let G be the number of reference units (genes in the TUSCO gene set or SIRV transcripts), N the number of observed transcripts mapping to reference units, TP the number of TP transcripts, PTP the number of PTP transcripts, FP the number of FP transcripts, TPu the number of reference units with at least one TP, D the number of reference units detected by either TP or PTP, and TFSM+ISM the total number of full-splice_match plus incomplete-splice_match transcripts. From these quantities we calculated:

Sn=100×TPuG, 2
nrPre=100×TPN, 3
rPre=100×TP+PTPN, 4
PDR=100×DG, 5
FDR=100×N−TPN, 6
FDeR=100×FPN,(FalseDetectionRate) 7
Redundancy=TFSM+ISMD. 8

For the TUSCO-novel evaluation, we additionally computed the F1 score as the harmonic mean of sensitivity and non-redundant precision:

F1=2×Sn×nrPreSn+nrPre 9

where Sn and nrPre are defined as above, expressed as percentages.

To quantify the concordance between benchmarking by the TUSCO gene set and by SIRVs, we computed cosine similarity (cosim) scores using the six performance metrics as a vector. For two benchmark approaches yielding metric vectors vTUSCO metrics and vSIRV, cosine similarity is defined as:

cosim=∑i=16vTUSCOmetrics,ivSIRV,i∑i=16vTUSCOmetrics,i2∑i=16vSIRV,i2 10

where the six-dimensional metric vectors comprise Sn, nrPre, Redundancy−1, 1 − FDR, PDR, and rPre, all expressed as percentages.

Statistical analysis

All statistical analyses were performed using R version 4.3.3. Python 3.11 was used for the TUSCO gene set selector and novel simulator pipelines. For the comparison between benchmarking by the TUSCO gene set and by SIRVs, we applied a two-sided sign-flip permutation test on the paired pipeline-level differences (n = 6 paired pipelines per species) for each category (TP, PTP, FP, FN). This test is exact for small n and does not assume normality of paired differences. No correction for multiple testing was applied, as the four comparisons address distinct, biologically motivated hypotheses about the relative behavior of endogenous versus synthetic benchmarks.

Correlation analyses between RIN values and transcript reconstruction rates employed Pearson product-moment correlation with n = 17 samples spanning RIN values from 7.2 to 9.9. The fraction of fully reconstructed transcripts was calculated as TPTP+PTP for benchmarking by the TUSCO gene set and for Sequins benchmarks. For Supplementary Fig. 4, we summarized the percentage of false negatives under benchmarking by the TUSCO gene set across LRGASP datasets (n = 12) grouped by sequencing platform and library method; this analysis was descriptive.

For the multi-replicate analysis, 95% confidence intervals were computed using the exact binomial method for sensitivity, precision, and positive detection rate metrics. When multiple combinations of k replicates were possible, we calculated the mean and confidence interval across all combinations, weighting each combination equally.

Intersection mode across replicates

For replicate analyses, we computed TUSCO metrics on a k-way intersection call set that retains only transcript structures reconstructed in all of the k selected replicates. Within each replicate, SQANTI3-classified transcripts were first assigned to genes in the TUSCO gene set and collapsed by structure (identical junction chain; TSS/TTS considered equivalent within ±50 bp). A structure entered the intersection if an equivalent collapsed transcript appeared in every replicate; one representative was then chosen by (i) minimum summed 5’/3’ end deviation from the TUSCO gene set model, (ii) higher read support, (iii) longer length. Correctness labels (TP, PTP, FP) and all downstream metrics (Sn, nrPre, rPre, PDR, FDR, FDeR, Redundancy) were computed exclusively on this intersection set; FN were genes in the TUSCO gene set lacking TP or PTP in the intersection. This conservative strategy emphasizes reproducibility, sharply reducing spurious calls while preserving consistently detected transcripts.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (1.6MB, pdf)

Source data

Source Data (938.3KB, xlsx)

Acknowledgements

Computational analyses were performed on the high-performance computing cluster Garnatxa at the Institute for Integrative Systems Biology (I2SysBio). I2SysBio is a joint research centre formed by University of Valencia (UV) and Spanish National Research Council (CSIC). T.L., F.J., and A.C. disclose support for the research of this work from the European Union Marie Skłodowska-Curie Actions Doctoral Network project LongTREC [HORIZON-MSCA-2021-DN-01, grant agreement 101072892]. A.P. discloses support for the research of this work from the Spanish MICINN [FPU21/01597]. A.C. and L.F.P. disclose support for the research of this work from the National Human Genome Research Institute of the National Institutes of Health [1R21HG011280-01]. A.C. discloses additional support from MICIU/AEI [PID2023-152976NB-I00]. A.F. discloses support for the research of this work from the National Human Genome Research Institute of the National Institutes of Health [U24HG007234, U24HG011451], Wellcome Trust [WT222155/Z/20/Z], and the European Molecular Biology Laboratory. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

Author contributions

T.L. developed TUSCO, performed analyses and drafted the manuscript. A.P., F.J., and L.F.P. contributed to the analysis of the PacBio replicated mouse dataset. A.F. contributed to the curation of TUSCO genes. A.C. supervised the study and manuscript drafting. All authors approved the manuscript.

Peer review

Peer review information

Nature Communications thanks Liliana Florea, Quentin Gouil, Paweł Łabaj and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Data availability

Publicly available LRS datasets used in this study include the LRGASP Consortium data for human WTC11 iPSC and mouse ES cells, accessible at the Synapse repository under accession syn25007472/wiki/608702. The RNA degradation datasets are deposited in the European Nucleotide Archive (ENA) under accession PRJEB53210, with supplementary FASTQ data for post-mortem brain samples available from the European Genome-Phenome Archive (EGA) under accession EGAS00001006542. The Replicated Mouse Brain and Kidney dataset, generated in-house, has been submitted in its entirety to the ENA under accessions PRJEB85167 (secondary accession ERP168594, “Sequencing of mice brains using Iso-Seq”) and PRJEB94912 (secondary accession ERP177675, “Sequencing of mice kidneys using Iso-Seq”). For benchmarking, the universal and tissue-specific TUSCO gene sets for human and mouse can be freely obtained at https://github.com/ConesaLab/SQANTI3 and at the dedicated TUSCO benchmarking framework portal https://tusco.uv.es. An example of the TUSCO report is available at: https://github.com/ConesaLab/tusco-paper/blob/main/example/tusco_output/tusco_report.html. The example input data and full TUSCO output files are provided in the example/ directory of the same repository. Source data are provided with this paper. All data are also available from the corresponding author upon request.

Code availability

All scripts used to generate the Figures in this study and to implement the TUSCO gene set selector are available at https://github.com/ConesaLab/tusco-paper (10.5281/zenodo.19204457).

Competing interests

A.C. has received in-kind support from PacBio. A.C., T.L., and F.J. are participants in the MSCA-DN project LongTREC, where Oxford Nanopore is also a partner. The remaining authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-72089-1.

References

  • 1.Marx, V. Method of the year: long-read sequencing. Nat. Methods20, 6–11 (2023). [DOI] [PubMed] [Google Scholar]
  • 2.Chen, Y. et al. Context-aware transcript quantification from long-read RNA-seq data with Bambu. Nat. Methods20, 1187–1195 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Prjibelski, A. D. et al. Accurate isoform discovery with IsoQuant using long reads. Nat. Biotechnol.41, 915–918 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Philpott, M. et al. Nanopore sequencing of single-cell transcriptomes with scCOLOR-seq. Nat. Biotechnol.39, 1517–1520 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Kabza, M. et al. Accurate long-read transcript discovery and quantification at single-cell, pseudo-bulk and bulk resolution with Isosceles. Nat. Commun.15, 7316 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Conesa, A. et al. A survey of best practices for RNA-seq data analysis. Genome Biol.17, 13 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Monzó, C., Liu, T. & Conesa, A. Transcriptomics in the era of long-read sequencing. Nat. Rev. Genet26, 681–701 (2025). [DOI] [PubMed] [Google Scholar]
  • 8.Pardo-Palacios, F. J. et al. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification. Nat. Methods21, 1349–1363 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Weirather, J. L. et al. Comprehensive comparison of Pacific Biosciences and Oxford Nanopore Technologies and their applications to transcriptome analysis. F1000Res6, 100 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Soneson, C. et al. A comprehensive examination of Nanopore native RNA sequencing for characterization of complex transcriptomes. Nat. Commun.10, 3359 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Dong, X. et al. Benchmarking long-read RNA-sequencing analysis tools using in silico mixtures. Nat. Methods20, 1810–1821 (2023). [DOI] [PubMed] [Google Scholar]
  • 12.Su, Y. et al. Comprehensive assessment of mRNA isoform detection methods for long-read sequencing data. Nat. Commun.15, 3972 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chen, Y. et al. A systematic benchmark of Nanopore long-read RNA sequencing for transcript-level analysis in human cell lines. Nat. Methods22, 801–812 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Monzó, C., Frankish, A. & Conesa, A. Notable challenges posed by long-read sequencing for the study of transcriptional diversity and genome annotation. Genome Res.35, 583–592 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Fukasawa, Y., Ermini, L., Wang, H., Carty, K. & Cheung, M.-S. LongQC: a quality control tool for third generation sequencing long read data. G3 Genes∣Genomes∣Genet.10, 1193–1196 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Keil, N., Monzó, C., McIntyre, L. & Conesa, A. Quality assessment of long read data in multisample lrRNA-seq experiments using SQANTI-reads. Genome Res.35, 987–998 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Leger, A. & Leonardi, T. pycoQC, interactive quality control for Oxford Nanopore Sequencing. JOSS4, 1236 (2019). [Google Scholar]
  • 18.Pardo-Palacios, F. J. et al. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms. Nat. Methods21, 793–797 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics31, 3210–3212 (2015). [DOI] [PubMed] [Google Scholar]
  • 20.Velasco, V. M. E. et al. A long-read and short-read transcriptomics approach provides the first high-quality reference transcriptome and genome annotation for Pseudotsuga menziesii (Douglas-fir). G3 Genes∣Genomes∣Genet.13, jkac304 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Nip, K. M. et al. Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2. Nat. Commun.14, 2940 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Nenasheva, N. et al. Annotation of protein-coding genes in 49 diatom genomes from the Bacillariophyta clade. Sci. Data12, 985 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Paniagua, A. et al. Evaluation of strategies for evidence-driven genome annotation using long-read RNA-seq. Genome Res.35, 1053–1064 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Jiang, L. et al. Synthetic spike-in standards for RNA-seq experiments. Genome Res.21, 1543–1551 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hardwick, S. A. et al. Spliced synthetic genes as internal controls in RNA sequencing experiments. Nat. Methods13, 792–798 (2016). [DOI] [PubMed] [Google Scholar]
  • 26.Yang, C., Chu, J., Warren, R. L. & Birol, I. NanoSim: Nanopore sequence read simulator based on statistical characterization. GigaScience6, gix010 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ono, Y., Hamada, M. & Asai, K. PBSIM3: a simulator for all types of PacBio and ONT long reads. NAR Genom. Bioinform.4, lqac092 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Mestre-Tomás, J., Liu, T., Pardo-Palacios, F. & Conesa, A. SQANTI-SIM: a simulator of controlled transcript novelty for lrRNA-seq benchmark. Genome Biol.24, 286 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Han, S. W., Jewell, S., Thomas-Tikhonenko, A. & Barash, Y. Contrasting and combining transcriptome complexity captured by short and long RNA sequencing reads. Genome Res.34, 1624–1635 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Wilks, C. et al. recount3: summaries and queries for large-scale RNA-seq expression and splicing. Genome Biol.22, 323 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Mudge, J. et al. GENCODE 2025: reference gene annotation for human and mouse. Nucleic Acids Res.53, D966–D975 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Goldfarb, T. et al. NCBI RefSeq: reference sequence standards through 25 years of curation and annotation. Nucleic Acids Res.53, D243–D257 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Morales, J. et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature604, 310–315 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.García-Ruiz, S. et al. IntroVerse: a comprehensive database of introns across human tissues. Nucleic Acids Res.51, D167–D178 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Abugessaisa, I. et al. refTSS: a reference data set for human and mouse transcription start sites. J. Mol. Biol.431, 2407–2422 (2019). [DOI] [PubMed] [Google Scholar]
  • 36.Bastian, F. B. et al. The Bgee suite: integrated curated expression atlas and comparative transcriptomics in animals. Nucleic Acids Res.49, D831–D847 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Hounkpe, B. W., Chenou, F., de Lima, F. & De Paula, E. HRT Atlas v1.0 database: redefining human and mouse housekeeping genes and candidate reference transcripts by mining massive RNA-seq datasets. Nucleic Acids Res.49, D947–D955 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Lonsdale, J. et al. The Genotype-Tissue Expression (GTEx) project. Nat. Genet.45, 580–585 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Avsec, i et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature649, 1206–1218 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Tang, A. D. et al. Detecting haplotype-specific transcript variation in long reads with FLAIR2. Genome Biol.25, 173 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Prawer, Y. D. J., Gleeson, J., De Paoli-Iseppi, R. & Clark, M. B. Pervasive effects of RNA degradation on Nanopore direct RNA sequencing. NAR Genom. Bioinform.5, lqad060 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wang, E. T. et al. Alternative isoform regulation in human tissue transcriptomes. Nature456, 470–476 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Pan, Q., Shai, O., Lee, L. J., Frey, B. J. & Blencowe, B. J. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing. Nat. Genet.40, 1413–1415 (2008). [DOI] [PubMed] [Google Scholar]
  • 44.Harrow, J. et al. GENCODE: the reference human genome annotation for The ENCODE Project. Genome Res.22, 1760–1774 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.O’Leary, N. A. et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res.44, D733–D745 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Kinsella, R. J. et al. Ensembl BioMarts: a hub for data retrieval across taxonomic space. Database2011, bar030 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (1.6MB, pdf)
Source Data (938.3KB, xlsx)

Data Availability Statement

Publicly available LRS datasets used in this study include the LRGASP Consortium data for human WTC11 iPSC and mouse ES cells, accessible at the Synapse repository under accession syn25007472/wiki/608702. The RNA degradation datasets are deposited in the European Nucleotide Archive (ENA) under accession PRJEB53210, with supplementary FASTQ data for post-mortem brain samples available from the European Genome-Phenome Archive (EGA) under accession EGAS00001006542. The Replicated Mouse Brain and Kidney dataset, generated in-house, has been submitted in its entirety to the ENA under accessions PRJEB85167 (secondary accession ERP168594, “Sequencing of mice brains using Iso-Seq”) and PRJEB94912 (secondary accession ERP177675, “Sequencing of mice kidneys using Iso-Seq”). For benchmarking, the universal and tissue-specific TUSCO gene sets for human and mouse can be freely obtained at https://github.com/ConesaLab/SQANTI3 and at the dedicated TUSCO benchmarking framework portal https://tusco.uv.es. An example of the TUSCO report is available at: https://github.com/ConesaLab/tusco-paper/blob/main/example/tusco_output/tusco_report.html. The example input data and full TUSCO output files are provided in the example/ directory of the same repository. Source data are provided with this paper. All data are also available from the corresponding author upon request.

All scripts used to generate the Figures in this study and to implement the TUSCO gene set selector are available at https://github.com/ConesaLab/tusco-paper (10.5281/zenodo.19204457).


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES