Summary
Existing RNA velocity estimation tools often fail to calculate velocities for a substantial portion of genes due to technical limitations or model assumptions, thereby restricting downstream analyses that rely on velocity estimations. To tackle this problem, we propose NARVI (Neural Network-Assisted RNA Velocity Imputation), a deep learning framework that learns the relationship between the expression patterns and velocities of computable genes to accurately estimate velocities for otherwise incalculable genes. We evaluated the performance of NARVI across multiple single-cell transcriptome datasets and applied it to trajectory inference and marker gene analysis using the reconstituted velocities of dropped genes. This approach recovers velocities for thousands of genes that were previously impossible to estimate, thereby broadening the scope of downstream analyses and providing deeper insights into gene transcriptional dynamics.
Subject areas: biochemistry, biocomputational method, neural networks
Graphical abstract

Highlights
-
•
NARVI imputes RNA velocity for genes dropped by conventional tools
-
•
NARVI-recovered velocities enable trajectory, marker-gene, and differential velocity analyses
-
•
Applicable across developmental, stimulation, and cross-species single-cell contexts
Biochemistry; Biocomputational method; Neural networks
Introduction
The advent of single-cell RNA sequencing (scRNA-seq) technology has enabled the profiling of gene expression in tens of thousands of individual cells.1,2 This technology allows for the detection of compositional changes in heterogeneous cell populations, the discovery of novel cell types, and the identification of cell-type-specific markers.3 Moreover, ordering cells along a pseudotemporal axis based on similarities in gene expression patterns elucidates dynamic cellular processes, such as differentiation and activation, through trajectory inference.4 In addition to the many graph-based and optimal transport-based methods,5,6 RNA velocity offers an approach that traces cell differentiation lineages using transcription dynamics.7,8 Specifically, RNA velocity captures gene transcriptional dynamics by measuring the ratio between unspliced (immature) and spliced (mature) mRNA, thereby estimating both the direction and magnitude of short-term state changes in each cell.8 When projected onto a low-dimensional embedding, these velocity vectors depict the dynamic transitions of cells as trajectories. In addition, applications of RNA velocity extend beyond trajectory inference to other downstream analyses. For example, there are differential kinetics analysis, which extracts genes that dynamically change transcription in lineages, and in silico perturbation and cell fate prediction based on the reconstruction of continuous vector fields using machine learning.9,10 However, these downstream analyses are limited to a small subset of genes, only those for which velocity can be inferred by existing tools.11 For instance, velocity cannot be calculated for genes that do not undergo splicing or do not produce sufficient amounts of immature mRNA.
To address dropout events in scRNA-seq read count values, various imputation methods have been reported. SAVER uses a Bayesian approach to impute zero values based on gene-gene interaction in the same cell,12 while scImpute uses a non-negative least squares method to impute expression for genes with a high probability of dropout as inferred from the expression levels of similar cells.13 More recently, deep neural network methods capable of handling large datasets have emerged for imputing read count values.14,15,16 For example, deep count autoencoder (DCA) is a variational autoencoder-based method that—because single cells are zero inflated sparse data—applies zero-inflated negative binomial (ZINB) regression.15 However, no direct methodology has been developed for imputing the RNA velocity of genes that cannot be calculated by existing tools.
Therefore, we propose here Neural Network-Assisted RNA Velocity Imputation (NARVI). NARVI uses gene expression values to recover the RNA velocity for genes that are excluded from calculations using conventional tools. In this study, we demonstrate the prediction performance of NARVI using public datasets. Additionally, we present trajectory inference and differentiation marker analyses using the reconstructed velocity. NARVI can recover velocity for thousands of genes, making the transcriptional kinetic information available for downstream analyses.
Results
A deep learning framework for the RNA velocity imputation of dropped genes
NARVI can decipher the complex relationships between gene expression value and RNA velocity using a deep neural network. This model learns the relationship between the expression pattern and velocity of specific genes across all cells and attempts to accurately predict velocities for genes that cannot be calculated using conventional tools.
The NARVI velocity estimation workflow is depicted in Figure 1. First, raw RNA velocity calculations are performed for the unspliced and spliced read count matrices using existing tools such as scVelo. This yields RNA velocities for the genes that pass the calculation (hereafter referred to as “passed genes”), as well as the expression values of total 2,000 highly variable genes (HVGs). For the passed genes, we add the observed/current expression value and raw RNA velocity to obtain a projected/future expression value. Genes for which raw RNA velocity could not be calculated are defined as “dropped genes.” Second, we use MAGIC to preprocess the expression values of all genes, including passed genes and dropped genes, as well as the predicted future expression values of passed genes. MAGIC reconstructs the denoised expression matrix based on diffusion between similar cells.17 Because the data matrices for the current and future expression values have the same number of cells, a square Markov matrix of size cells × cells, calculated from the expression values of all genes, can be applied to the predicted future expression values of passed genes. Third, to train NARVI, we use the current expression values of the passed genes as explanatory variables and their velocities as the objective variables. The deep neural network learns to infer the velocity of each gene from the expression pattern in each cell. Essentially, this model can recover the RNA velocities of dropped genes that have expression distributions similar to those of passed genes. Finally, after completing the training, the model predicts the RNA velocities of the dropped genes using their current expression values as input.
Figure 1.
Workflow of the NARVI
Raw RNA velocity is first calculated for passed genes using existing tools (e.g., scVelo or UniTVelo). Current and projected future expression values are then denoised using MAGIC to obtain smoothed matrices. NARVI trains a deep neural network to learn the relationship between expression patterns and RNA velocity from passed genes. After training, the model imputes RNA velocities for dropped genes based on their expression profiles, enabling reconstruction of a more comprehensive velocity landscape for downstream analyses such as trajectory inference and marker gene dynamics.
To validate the performance of NARVI, we used five datasets (Table S1): mouse hematopoietic cell differentiation,18 mouse lung epithelial bleomycin stimulation,19 mouse erythroid differentiation,20 mouse pancreatic endocrine cell differentiation,21 and a human bone marrow dataset.22 We applied a cross-validation (CV) method to evaluate the performance of NARVI using the validation data (Figure S1). We then assessed model performance by examining the correlation between predicted and true velocity values for the genes in the validation data.
NARVI application for hematopoiesis dataset
To verify NARVI’s performance in imputing RNA velocity, we tried to recover the velocities for dropped genes in a hematopoietic cell differentiation dataset obtained by inDrop.18 We used scVelo, a widely adopted Python-based tool for velocity analysis,9 to calculate raw RNA velocity. For this dataset, 916 genes were extracted as passed genes by scVelo, and their RNA velocities were calculated. By visualizing the RNA velocity of the passed genes on an embedding, we confirmed the dynamic transition of each cell (Figure 2A). As expected, we observed the differentiation from undifferentiated cells into various functional cells, including monocytes and neutrophils (Figure 2A). Next, we divided the passed genes into training and validation sets and evaluated the performance of NARVI using CV. For the validation genes with the highest prediction accuracy, the correlation coefficients between the predicted and true values were 0.95 or higher (Figure 2B). We then examined the distribution of the correlation coefficients between the predicted and true values for each gene in both the training and validation data. For the 305 validation genes, the median and mean correlation coefficients were 0.72 and 0.78, respectively (Figure 2C). In addition, across all 3-fold CV models, the average correlation coefficient for the validation data was 0.70 (Figure S2A). Furthermore, we restored the velocities of the 813 dropped genes (41% of total 2,000 HVGs) that were missing and examined whether the cell dynamics were similar to those of the passed genes (Figures 2D–2F). In the case of dropped genes, although a minor flow in megakaryocytes and mast cells was observed in the opposite direction to that of the passed genes, the starting point of the velocity flow and the overall cellular dynamics from undifferentiated to functional cells were similar for both passed and dropped genes (Figures 2A, 2D, and 2F). The streams of velocity by a combination of the passed and dropped genes diminished the partially emergent contradictory flow and reproduced a flow pattern closer to that of only passed genes (Figures 2A and 2E). To further explore the potential of NARVI for downstream analysis in biological contexts, we then restored the velocities of representative differentiation marker genes. Cluster of differentiation 74 (Cd74) is a marker for monocytes and monocyte-derived dendritic cells and is known to play an important role in antigen presentation.23,24 Consistent with this, the NARVI-reconstructed velocity of Cd74 emphasized transcriptional induction during differentiation from progenitor cells to monocytes (Figure 2G). Lipocalin-2 (Lcn2), also known as neutrophil gelatinase-associated lipocalin (Ngal), is a marker that is highly expressed in neutrophils and contributes to protection against bacterial infections and the regulation of oxidative stress.25,26 The recovered velocity of Lcn2 was consistent with its increased expression during neutrophil maturation (Figure 2G).
Figure 2.
NARVI application for hematopoiesis dataset
Validation of NARVI using expression and velocity data obtained with scVelo (A–G) and UniTVelo (H–N) for the hematopoiesis dataset from Weinreb et al.
(A) The streams of the velocity of passed genes calculated by scVelo on SPRING embedding, showing differentiation trajectories from progenitors to blood cells such as monocytes and neutrophils.
(B) Scatterplots of predicted vs. true RNA velocity values for six validation genes with highest correlation in one CV fold; each dot represents a cell. The Pearson correlation coefficient (PCC) between the predicted and true values for each gene is displayed on the subplot.
(C) Boxplots of PCCs between predicted and true velocities for training and validation genes in the best CV fold.
(D) The streams of reconstructed dropped gene’s velocity.
(E) The streams of velocity for all genes (passed and dropped genes) after imputation by NARVI.
(F) Violin plots of cosine similarity between velocity vectors of passed genes vs. dropped genes or all genes for each cell; red line indicates median.
(G) Expression and recovered velocity for two marker genes (Cd74 and Lcn2) among dropped genes, visualized on embedding.
(H–N) Same analyses as (A)–(G)– but using UniTVelo as the base tool.
Moreover, we verified the imputation of RNA velocity calculated using UniTVelo,27 a different tool from scVelo. Through UniTVelo calculations, we obtained RNA velocities for 1,348 passed genes in this dataset. The resulting cellular dynamics of the passed genes exhibited a flow from undifferentiated cells to monocytes and basophils, with slight differences in lymphoid and plasmacytoid dendritic cell (pDCs) peripheral flow compared to the scVelo results (Figures 2A and 2H). We then verified the performance of NARVI using a portion of the passed genes as validation data (Figures 2I and 2J). The gene with the highest prediction accuracy showed a high value of around 0.99 (Figure 2I). In the best CV model, the mean and median correlation coefficients for the validation data were 0.64 and 0.87, respectively (Figure 2J), with an average of 0.62 across all three models (Figure S2A). Additionally, we reconstructed the velocities of the 489 dropped genes (24% of total 2,000 HVGs) using NARVI (Figures 2K, 2L, and 2M). A comparison of the velocity streams on the embedding of dropped and passed genes showed similar velocity flows from undifferentiated to terminal differentiated cells, such as monocytes and basophils (Figures 2H and 2K). The cosine similarity of vectors for each cell in the reduced dimension also showed high values (Figure 2M). Finally, we imputed the velocity of Cd74 and Lcn2, which were also defined as dropped genes by UniTVelo. The velocity of Cd74 and Lcn2 highlighted the transcriptional induction from undifferentiated to monocyte maturation and from undifferentiated to neutrophil, respectively (Figure 2N).
NARVI application for lung epithelial dataset
To test the applicability of NARVI to a broader dataset, we used a dataset capturing state changes of lung epithelial cells following bleomycin stimulation.19 Using scVelo with standard parameters, we identified 490 passed genes in the lung data and reproduced the cell differentiation flow based on the RNA velocities of passed genes (Figure 3A). For the validation data, the most accurate gene group exhibited a correlation coefficient of 0.95 or more (Figure 3B). When applying NARVI to the 327 passed genes and the 163 validation genes, the median correlation coefficients were 0.77 and 0.79, respectively (Figure 3C), and there was no tendency for overfitting in any of the CV (Figure S2B). Furthermore, NARVI reconstructed the velocity of the 1,132 dropped genes (57% of total 2,000 HVGs) (Figures 3D–3F). The flows from goblet cells through club-to-ciliated cells to ciliated cells, those from club activated cells to club cells, and those from aberrant basal cells (ABCs) to AT1 were consistently observed on all embeddings (Figures 3A, 3D, and 3E). Then, we imputed the velocities of the marker genes among the dropped genes, following the same approach used for the verification of the Weinreb’s dataset. Tubulin alpha 1a (Tuba1a) is an expression marker for ciliated cells in the respiratory airways and is necessary for cilia formation and maturation.28,29 The reconstructed velocity of Tuba1a suggests its transcriptional induction along the cluster from club-to-ciliated cells, which is consistent with its increased expression in ciliated cells (Figure 3G). Tissue inhibitor of metalloproteinases 2 (Timp2) is known to be highly expressed in the airway epithelium in animal models of bleomycin-induced lung injury for pulmonary fibrosis research30 and is thought to contribute to fibrosis by regulating the amount of extracellular matrix via modulation of matrix metalloproteinases.31 In this dataset, Timp2 is highly expressed in alveolar type cells such as ABCs, which increase in response to bleomycin stimulation. Consistent with this observation, the recovered velocity of Timp2 was highly elevated in AT2 activated, ABCs, and AT1 cells (Figure 3G).
Figure 3.
NARVI application for lung epithelial dataset
Validation of NARVI using the expression and velocity obtained with scVelo (A–G) and UniTVelo (H–N) for lung epithelial dataset.
(A) The streams of the velocity of passed genes calculated by scVelo on UMAP embedding.
(B) Scatterplots of predicted vs. true RNA velocity values for six validation genes with highest correlation in one CV fold; each dot represents a cell. The PCC between the predicted and true values for each gene is displayed on the subplot.
(C) Boxplots of PCCs between predicted and true velocities for training and validation genes in the best CV fold.
(D) The streams of the reconstructed dropped gene’s velocity.
(E) The streams of the velocity for all genes (passed and dropped genes) after imputation by NARVI.
(F) Violin plots of cosine similarity between velocity vectors of passed genes vs. dropped genes or all genes for each cell; red line indicates median.
(G) Expression and recovered velocity for two marker genes (Tuba1a and Timp2) among dropped genes, visualized on embedding.
(H–N) Same analyses as (A)–(G) but using UniTVelo as the base tool.
Next, we validated this dataset using UniTVelo. The RNA velocities of the 1,078 extracted passed genes were used to visualize the cellular state changes (Figure 3H), and these genes were divided into training and validation sets to verify the performance of NARVI (Figures 3I and 3J). As with the UniTVelo verification of Weinreb’s dataset, the genes with the highest prediction accuracy showed a very high correlation between predicted and true values (Figure 3I). In the best CV model, the mean and median correlation coefficients for the validation data were 0.77 and 0.99, respectively (Figure 3J). Furthermore, we imputed the velocities of 691 dropped genes (35% of total 2,000 HVGs) and confirmed the agreement with the velocities of the passed genes on uniform manifold approximation and projection (UMAP) (Figures 3K, 3L, and 3M). The flow from the progenitor cells of club-to-ciliated cells into ciliated cells and that from AT2 activated cells to AT1 and ABCs were similar in all cases (Figures 3H, 3K, and 3L). Additionally, we also performed NARVI-based imputation of two marker genes included among the dropped genes. As we expected, we observed that the recovered velocity of Tuba1a was enhanced in the ciliated cell lineage and that Timp2 recovered velocity was high in clusters that increased in response to bleomycin stimulation, such as AT2 activated cells, which was concordant with the higher expression of this gene in AT1 and ABCs (Figure 3N).
Application of NARVI to the erythroid dataset to examine differences between base methods
To investigate the effects of different base tools used for calculating raw RNA velocity on velocity imputation by NARVI, we applied NARVI to a dataset of mouse erythroid differentiation.20 In this dataset, applying scVelo with standard parameters leads to distorted lineage inference due to the presence of multiple rate kinetics genes that violate model assumptions.32 As expected, the velocity flow using the 248 passed genes in scVelo emphasized the lineage from mature erythroid to progenitor cells (Figure 4A). Subsequently, we trained NARVI and predicted the velocity of the 1,287 dropped genes (64% of total 2,000 HVGs) (Figure 4B). The velocity streams of the dropped genes were similar to those of the passed genes; in other words, the contradictory flow from mature erythroid to progenitor cells was reproduced (Figures 4A and 4B, left).
Figure 4.
NARVI application for erythroid dataset
Performance evaluation of NARVI using expression and velocity acquired by scVelo (A and B) and velocity acquired by UniTVelo (C, D, and E) for the erythroid differentiation dataset.
(A) The streams of the passed gene’s velocity calculated by scVelo on UMAP embedding.
(B) The streams of the dropped gene’s velocity reconstructed by NARVI (left) and those of all genes (passed and dropped genes) after imputation by NARVI (right).
(C and D) Same analyses as (A) and (B) but using UniTVelo as the base velocity estimation tool.
(E) Expression and recovered RNA velocity for six marker genes (Klf1, Alas2, Prdx2, Hba-x, Hba-a2, and Hbb-bt) among dropped genes, visualized on UMAP embedding. These velocities were imputed by NARVI using UniTVelo as the base tool.
In contrast, the application of UniTVelo to this dataset has previously been shown to recover the expected trajectory erythroid cell maturation,27 and we also reproduced this trend using RNA velocity estimates from 1,015 passed genes (Figure 4C). NARVI reconstructed a flow similar to that of the passed genes on the embedding by restoring the velocities of 739 dropped genes (37% of total 2,000 HVGs) (Figures 4C and 4D). These observations suggest that the velocity flow of the dropped genes, as reconstructed by NARVI using information from the passed genes, is similar to that of the passed genes, regardless of whether the base tool’s predictions align with the biological findings.
Finally, using UniTVelo—which produces a lineage that matches biological knowledge on UMAP—we recovered the RNA velocities of six marker genes found among the dropped genes. All of these genes are required for the maturation of red blood cells,33,34,35 and we found that they are transcriptionally induced in the progression from progenitor to fully mature red blood cells (Figure 4E). Interestingly, the transcription factors Kruppel-like factor 1 (Klf1) and hemoglobin beta adult t chain (Hbb-bt) were detected with particularly low expression values, yet their induction in line with a developmental lineage was appropriately predicted. These results demonstrate NARVI’s potential for enabling downstream velocity-based analyses of previously inaccessible genes.
Discussion
We propose NARVI as a tool for directly imputing RNA velocity for genes that are dropped by existing estimation tools. We evaluated the performance of this workflow using multiple datasets and reconstructed cell trajectories using the imputed velocities of dropped genes. Across all datasets, the recovered velocities resulted in trajectories similar to those derived from passed genes. Furthermore, by employing different raw RNA velocity estimation tools, we evaluated the applicability of NARVI. In addition, we interpreted the transcription kinetics of key marker genes within the dropped genes by reconstructing their velocities in each dataset.
More than a dozen methods have been reported for imputing scRNA-seq data.12,13,14,15,16 These approaches are primarily designed to impute expression values and excel at recovering count matrices or expression data.36 However, those methods are not intended for RNA velocity imputation. This limitation arises because the data types have different characteristics; unlike count data, velocity values can be negative to indicate transcriptional suppression and are very small, following a distribution distinct from that of expression values. Accordingly, given the characteristics of velocity data, NARVI is designed to model RNA velocity directly, and our results verified the feasibility of this approach to imputation.
DeepVelo is a deep-learning-based method that predicts RNA velocity from measured transcriptome profiles.37 It employs a variational autoencoder to learn velocity and then performs iterative calculations using a model based on the Euler method to predict continuous cell transitions. In contrast to NARVI, DeepVelo focuses on long-term cellular dynamics by leveraging information from genes with successfully estimated RNA velocities. NARVI and DeepVelo are complementary: NARVI expands the RNA velocity field by imputing dropped genes, enabling DeepVelo to operate on a more comprehensive velocity landscape. This synergy allows for more accurate in silico perturbation and long-term trajectory prediction, particularly in contexts where key regulatory genes are missing from conventional velocity estimates.
From our analysis of the Weinreb’s dataset, differential velocity analysis identified Cd74 and Lcn2 among genes associated with monocyte and neutrophil differentiation (Figure S3). Interestingly, these genes are not merely lineage markers: Cd74 has been implicated in hematopoietic stem cell maintenance, while Lcn2 contributes to neutrophil maturation.38,39 This observation implies that NARVI can potentially recover dynamic signals for dropped genes, including those with regulatory roles. Beyond Cd74 and Lcn2, other dropped genes may similarly represent regulators of lineage commitment or functional maturation, underscoring NARVI’s potential to reveal hidden aspects of cellular differentiation.
Recent multimodal approaches are reshaping single-cell analysis by integrating transcriptomic, spatial, and epigenomic layers to uncover gene co-expression patterns and regulatory networks at unique resolution. These advances, exemplified by methods such as LEGEND and scMultiomeGRN, highlight a shift toward comprehensive modeling of cellular states.40,41 Building on this trend, MultiVelo extends RNA velocity into the multi-omic space, combining chromatin accessibility and transcriptional dynamics to refine cell fate prediction.42 In this context, NARVI offers a complementary role: NARVI can potentially expand the velocity landscape predicted by MultiVelo, enabling more complete dynamic modeling and enhancing the interpretability of multimodal frameworks.
Currently, there is no established method for directly imputing RNA velocity, which has restricted downstream analyses that rely on velocity calculations to a small number of genes.11 NARVI addresses this issue by recovering the velocity of dropped genes, making it possible to perform downstream analyses on a broader gene set. In this study, we demonstrated trajectory inference that integrates the velocities of passed and dropped genes, as well as velocity imputation for dropped marker genes, but those are just a few potential applications for this model. NARVI will be a particularly useful tool in fields such as developmental biology, immunology, and oncology, where detailed understanding of single-cell gene expression dynamics is critical.43,44 Ultimately, NARVI has the potential to contribute to the discovery of regulatory factors and pathways involved in cellular state transitions and to enrich our understanding of gene expression dynamics.
Limitations of the study
NARVI learns the relationship between the expression pattern and velocity in all cells for specific genes. The prediction performance of this architecture depends on the similarity between the expression patterns of passed and dropped genes. Therefore, it is difficult to accurately impute RNA velocity for out-of-distribution genes, as their expression profiles differ substantially from the training data.
NARVI currently learns velocity patterns using input expression vectors whose dimension is the number of cells in the dataset. Because this architecture is tied to dataset-specific cell composition, the trained model cannot be directly applied to other datasets without retraining. This limits NARVI’s current use to imputing RNA velocity within a single dataset of interest. Future work could overcome this limitation by leveraging atlas-scale training to enable cross-dataset generalization.
The performance of NARVI is sensitive to the choice of preprocessing. RNA velocity is estimated from the unspliced and spliced counts, but they are often noisy due to transcriptional stochasticity, low abundance of unspliced transcripts, and uneven detection across genes and cells. Applying MAGIC-based denoising consistently improved the prediction accuracy of validation genes across all datasets compared to no denoising, confirming that noise reduction is critical for NARVI’s performance (Table S2). Since NARVI can incorporate general denoising using the expression matrix, and given ongoing improvements in experimental and computational methods for generating higher quality scRNA-seq data, its performance is expected to further benefit from these advances.
In our validation of NARVI for erythroid cells, we highlighted the importance of aligning the base tool’s raw RNA velocity predictions with established biological knowledge. This is because there is a concern that the velocity of dropped genes recovered in violation of biological knowledge may lead to misinterpretation in subsequent downstream analyses. We observed similar discrepancy in mouse pancreas during endocrinogenesis and in human bone marrow during hematopoiesis (Figure S4). For example, in the pancreas dataset, scVelo reproduced the cyclic trajectories of ductal cells along the cell cycle, whereas UniTVelo did not. In the human bone marrow dataset, similar to the mouse erythroid results, scVelo observed a contradictory flow from terminally differentiated cells to hematopoietic stem cells, whereas UniTVelo obtained the velocity streams consistent with prior knowledge. Because NARVI learns the relationship between gene expression and velocity for passed genes, the recovered velocities of dropped genes in principle mirror the flow of the passed genes. This observation highlights the importance of verifying that the velocities estimated by the base tool are consistent with prior biological knowledge when using NARVI for velocity reconstruction and downstream analysis. Although combining the imputed velocities of dropped genes with the velocities of passed genes could potentially yield an improved trajectory consistent with biological expectations, careful consideration is needed when applying the current version of NARVI for this purpose.
Resource availability
Lead contact
Requests for further information should be directed to the lead contact, Hideaki Mizuno (mizunohda@chugai-pharm.co.jp).
Materials availability
This study did not generate new unique reagents.
Data and code availability
The tutorial and demonstration code for NARVI are available at https://bitbucket.org/tech-kobo/narvi/. The analysis code used to generate the results in this paper is provided as supplemental information in a separate compressed archive: Data S1. No new raw sequencing data were generated; all third-party single-cell datasets analyzed here are publicly available under the accessions listed in the key resources table. The NARVI algorithm is the subject of a patent application. Commercial use of the patented algorithm requires a separate patent licensing agreement. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Acknowledgments
We thank Jacob Davis of Chugai Pharmaceutical Co., Ltd. for proofreading.
Author contributions
R.E., M.S., and H.M. conceptualized this study. R.E. and M.S. developed the NARVI workflow. The original draft was written by R.E., M.S., and H.M. with assistance from all authors. The manuscript was reviewed and edited by R.E., M.S., T.T., S.O., and H.M. Project administration and supervision were performed by T.T., S.O., and H.M.
Declaration of interests
All authors are the employees and shareholders of Chugai Pharmaceutical Co., Ltd. The authors have a patent application related to this work.
STAR★Methods
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Deposited data | ||
| Mouse hematopoiesis single cell dataset | Weinreb et al., 202018 | GSE140802 |
| Mouse bleomycin-stimulated lung single cell dataset | Strunz et al., 202019 | GSE141259 |
| Mouse gastrulation single cell dataset | Pijuan-Sala et al., 201920 | E-MTAB-6967 |
| Mouse pancreas endocrinogenesis single cell dataset | Bastidas-Ponce et al., 201921 | GSE132188 |
| Human bone marrow single cell dataset | Setty et al., 201922 | S-SUBS8 |
| Software and algorithms | ||
| NARVI demo code | This paper | https://bitbucket.org/tech-kobo/narvi/ |
| NARVI analysis scripts | This paper | Supplemental Information: Code S1 (ZIP file) |
| Scanpy | Wolf et al., 201845 | https://scanpy.readthedocs.io/en/stable/ |
| scVelo | Bergen et al., 20209 | https://scvelo.readthedocs.io/en/stable/ |
| UniTVelo | Gao et al., 202227 | https://unitvelo.readthedocs.io/en/latest/ |
| Pyro-Velocity | Qin et al., 202246 | https://pinellolab.github.io/pyrovelocity/ |
| CellRank | Lange et al., 202247; Weiler et al., 202448 | https://cellrank.readthedocs.io/en/latest/ |
| MAGIC | van Dijk et al., 201817 | https://magic.readthedocs.io/en/stable/ |
| Optuna | Akiba et al., 201949 | https://github.com/optuna/optuna |
| PyTorch | Paszke et al., 201950 | https://pytorch.org/ |
Experimental model and study participant details
This study did not generate new experimental models, primary cell cultures, or study participants. All analyses were performed using previously published, publicly available single-cell datasets, which are listed in the key resources table. We did not perform new animal experiments, recruit human participants, or collect new human specimens; therefore, institutional permissions, oversight, and informed consent procedures are as described in the corresponding original publications. Because this work is a secondary analysis of publicly available datasets, the availability of sex and gender metadata depends on the original studies. Sex information is not available in the human AnnData object, and we did not perform sex-stratified analyses in any of the datasets analyzed.
Method details
Dataset
The mouse hematopoietic lineage dataset is available through Pyro-Velocity,46 and the original data is derived from Weinreb et al.18 The mouse lung epithelial dataset is available through CellRank,47,48 and the original data is derived from Strunz et al.19 The mouse erythroid lineage, mouse pancreas, and human bone marrow single-cell datasets are available through scVelo,9 and the original data are derived from Pijuan-Sala et al.,20 Bastidas-Ponce et al.,21 and Setty et al.,22 respectively. Each dataset, downloaded in the h5ad format, was loaded using the scVelo function and used as AnnData.
Preprocessing of training data
Raw RNA velocity calculations using scVelo were performed following the tutorial (https://scvelo.readthedocs.io/en/stable/about.html). The filtering and normalization functions were run with min_shared_counts = 20 and highly variable genes = 2000, and the moment computation was performed with n_pcs = 30 and n_neighbors = 30. The mode for RNA velocity estimations using scVelo was set to stochastic. Raw RNA velocity calculations using UniTVelo were also performed, referring to the tutorial. As with scVelo, 2,000 highly variable genes were extracted. Most of the other specified parameters were set in accordance with the tutorial, with R2_ADJUST = True, LEARNING_RATE = 0.01, FIT_OPTION = 1 (unified-time mode), and MAX_ITER = 10000.
Then, the passed genes for which the raw RNA velocity could be calculated were extracted from the elements in anndata.var.[“velocity_genes”] that equaled True. The elements that equaled False were defined as dropped genes for which velocity could not be estimated. For the passed genes, the projected future expression value was computed by adding the raw RNA velocity to the current expression value, following the preprocessing protocol in velocyto.8 For the future expression values, values that became negative due to the addition of velocity were clipped to zero.8 These processes were carried out to facilitate the subsequent application of MAGIC, which is described below.
The MAGIC model was first applied to the data matrix of current expression values across all genes, including both passed and dropped genes, to calculate the exponentialized Markov affinity matrix for computing denoised expressions.17 The obtained matrix was then applied to the current and future expression values to get the denoised expression matrix. Then, we defined RNA velocity as the difference between the denoised future and current expression values.
For some datasets, we found that the NARVI’s performance in imputing RNA velocity was suboptimal for dropped genes with low expression values. Based on this observation, only dropped genes with a current expression value greater than or equal to the first quartile are predicted here. Finally, Z score normalization was applied to the data matrices of the current expression values of the passed and dropped genes, as well as to the RNA velocity values of the passed genes.
Description of NARVI architecture
NARVI is composed of multiple blocks, each of which consists of a fully connected layer, an activation function, and a dropout layer. Specifically, there is a fully connected layer that converts from the input layer to the first hidden layer, followed by the activation function, and then the dropout layer. We assume that h(1) is the output of the first layer, and the transformation from the input layer to the first layer is given by the following equation:
where x is the input vector, W(1) is the weight matrix, b(1) is the bias vector, σ is the activation function, and p is the dropout rate. This sequence of structures is considered to be one block, and this block is repeated for each hidden layer. The transformation from the l-th layer to (l+1)-th layer can be represented as follows:
Finally, the network is completed with a fully connected layer that returns to the input size. The final output vector y is as follows:
where L is the index of the final layer.
Evaluation of model performance on cross-validation
True RNA velocity values for dropped genes cannot be obtained using conventional tools. We evaluated NARVI’s performance by treating the subset of genes with calculable RNA velocities as training data and the remaining genes as validation data (Figure S1). We employed nested CV to split the training and validation data, using k-fold CV in the outer loop (k = 3). Additionally, we used two inner k-fold CV procedures for hyperparameter tuning and for evaluating the predictive performance. We used the Pearson product-moment correlation coefficient as the evaluation metric. The optimization of the hyperparameters, loss function, and algorithm for the NARVI model were performed in the first inner k-fold CV using Optuna, searching within the defined range of values (Table S3).49 The searching of hyperparameters by Optuna involved solving the minimization problem using the Tree-Structured Parzen Estimator algorithm over 100 trials, with the objective function being the mean value across all inner folds.51,52 Each trial was pruned based on the median of the trials. Using the best set of parameters identified by Optuna, we trained and validated the model on each fold of the second inner k-fold CV. For each inner fold, inference was performed using the model from the epoch that yielded the minimum validation loss. The final prediction for the outer fold’s validation data was obtained by averaging the inference results across all inner folds. The precise architecture and hyperparameters of NARVI used for each dataset are shown in the table (Table S4). The predicted values were then transformed back to their original scale using the statistics (mean and variance) applied during Z score normalization of the velocity data matrix. The correlation coefficient between predicted and true velocity values was then calculated.
Reconstruction of vector flow on embedding
The velocities of the passed or dropped genes were mapped onto a two-dimensional embedding. On the embedding, each dot represents a cell, and the coordinates of these were calculated based on the gene expression values. The vectors representing the RNA velocity were computed as having a length equal to the number of genes per cell, and stored in anndata.layers[“velocity”]. As a result, the velocities of the passed and dropped genes were drawn as velocity stream plots on the two-dimensional embedding using the scVelo function pl.velocity_embedding_stream. The figure for all genes was created using the velocities of all genes, both passed and dropped, to draw the vector streams. Also, the pre-calculated embeddings included in the respective datasets were reused for the base embeddings. In other words, the Weinreb’s dataset used SPRING, the mouse lung, mouse red blood cell, and mouse pancreas datasets used UMAP, and the human bone marrow dataset used t-SNE for the embedding.53,54,55 Furthermore, we focused on the vectors of each cell reconstructed on the two-dimensional embedding, and calculated the cosine similarity between the vectors of the dropped genes, all genes, and passed genes. These were the indices used to represent the similarity of the dynamics of each cell on the embedding.
Recovery of velocity for marker genes included among dropped genes
For the mouse hematopoiesis, mouse lung, and mouse erythroid maturation datasets, we performed RNA velocity recovery using NARVI for several key marker genes included among the dropped genes. Cd74 and Lcn2 in mouse hematopoiesis and Tuba1a and Timp2 in mouse lung, genes known to play important roles in cell differentiation and state transitions, were selected as marker genes and were among the genes that were dropped both in scVelo and UniTVelo. In the mouse erythrocyte differentiation dataset, we selected six genes, including hemoglobin-related genes involved in the erythrocyte maturation. The expression and recovered velocities of the selected genes were visualized using scVelo’s pl.velocity function.
Differential velocity analysis and visualization
Differential velocity analysis was performed on the Weinreb’s hematopoiesis dataset to identify differential velocity genes (DVGs) in monocyte and neutrophil lineages. For each lineage, we applied the scVelo function tl.rank_velocity_genes with groupby = "cell_type". After execution, velocity gene scores were extracted from adata.uns['rank_velocity_genes']["scores"]. Genes with a score ≥0 were defined as DVGs. These scores were log10-transformed and ranked in descending order for visualization. Rank plots were generated with the x axis representing the rank and the y axis representing the log10-transformed score.
Quantification and statistical analysis
All analyses were performed in Python using Scanpy with scVelo and/or UniTVelo. NARVI was implemented in PyTorch, with MAGIC used for denoising and Optuna for hyperparameter optimization; all scripts are provided as Code S1. Model performance was assessed by 3-fold nested cross-validation on velocity-estimable (passed) genes, using the Pearson correlation coefficient (PCC) between predicted and true velocity values across cells as the primary metric. Distributions of per-gene PCC values for training and validation genes were summarized as boxplots with median and interquartile range. Agreement between velocity fields was evaluated by projecting velocity vectors onto the dataset-provided embeddings and computing per-cell cosine similarity, summarized as distributions with the median indicated. Differential velocity analysis was performed using scVelo rank_velocity_genes (groupby = cell_type). Exact summary statistics are provided in the results and corresponding figure legends.
Published: February 5, 2026
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.isci.2026.114909.
Supplemental information
References
- 1.Kolodziejczyk A.A., Kim J.K., Svensson V., Marioni J.C., Teichmann S.A. The technology and biology of single-cell RNA sequencing. Mol. Cell. 2015;58:610–620. doi: 10.1016/j.molcel.2015.04.005. [DOI] [PubMed] [Google Scholar]
- 2.Tang F., Barbacioru C., Wang Y., Nordman E., Lee C., Xu N., Wang X., Bodeau J., Tuch B.B., Siddiqui A., et al. mRNA-Seq whole-transcriptome analysis of a single cell. Nat. Methods. 2009;6:377–382. doi: 10.1038/nmeth.1315. [DOI] [PubMed] [Google Scholar]
- 3.Luecken M.D., Theis F.J. Current best practices in single-cell RNA-seq analysis: a tutorial. Mol. Syst. Biol. 2019;15 doi: 10.15252/msb.20188746. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Saelens W., Cannoodt R., Todorov H., Saeys Y. A comparison of single-cell trajectory inference methods. Nat. Biotechnol. 2019;37:547–554. doi: 10.1038/s41587-019-0071-9. [DOI] [PubMed] [Google Scholar]
- 5.Schiebinger G., Shu J., Tabaka M., Cleary B., Subramanian V., Solomon A., Gould J., Liu S., Lin S., Berube P., et al. Optimal-Transport Analysis of Single-Cell Gene Expression Identifies Developmental Trajectories in Reprogramming. Cell. 2019;176:1517. doi: 10.1016/j.cell.2019.02.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wolf F.A., Hamey F.K., Plass M., Solana J., Dahlin J.S., Göttgens B., Rajewsky N., Simon L., Theis F.J. PAGA: graph abstraction reconciles clustering with trajectory inference through a topology preserving map of single cells. Genome Biol. 2019;20:59. doi: 10.1186/s13059-019-1663-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Deconinck L., Cannoodt R., Saelens W., Deplancke B., Saeys Y. Recent advances in trajectory inference from single-cell omics data. Curr. Opin. Syst. Biol. 2021;27 doi: 10.1016/j.coisb.2021.05.005. [DOI] [Google Scholar]
- 8.La Manno G., Soldatov R., Zeisel A., Braun E., Hochgerner H., Petukhov V., Lidschreiber K., Kastriti M.E., Lönnerberg P., Furlan A., et al. RNA velocity of single cells. Nature. 2018;560:494–498. doi: 10.1038/s41586-018-0414-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bergen V., Lange M., Peidli S., Wolf F.A., Theis F.J. Generalizing RNA velocity to transient cell states through dynamical modeling. Nat. Biotechnol. 2020;38:1408–1414. doi: 10.1038/s41587-020-0591-3. [DOI] [PubMed] [Google Scholar]
- 10.Qiu X., Zhang Y., Martin-Rufino J.D., Weng C., Hosseinzadeh S., Yang D., Pogson A.N., Hein M.Y., Hoi Joseph Min K., Wang L., et al. Mapping transcriptomic vector fields of single cells. Cell. 2022;185:690–711.e45. doi: 10.1016/j.cell.2021.12.045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Bergen V., Soldatov R.A., Kharchenko P.V., Theis F.J. RNA velocity-current challenges and future perspectives. Mol. Syst. Biol. 2021;17 doi: 10.15252/msb.202110282. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Huang M., Wang J., Torre E., Dueck H., Shaffer S., Bonasio R., Murray J.I., Raj A., Li M., Zhang N.R. SAVER: gene expression recovery for single-cell RNA sequencing. Nat. Methods. 2018;15:539–542. doi: 10.1038/s41592-018-0033-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Li W.V., Li J.J. An accurate and robust imputation method scImpute for single-cell RNA-seq data. Nat. Commun. 2018;9:997. doi: 10.1038/s41467-018-03405-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Arisdakessian C., Poirion O., Yunits B., Zhu X., Garmire L.X. DeepImpute: an accurate, fast, and scalable deep neural network method to impute single-cell RNA-seq data. Genome Biol. 2019;20:211. doi: 10.1186/s13059-019-1837-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Eraslan G., Simon L.M., Mircea M., Mueller N.S., Theis F.J. Single-cell RNA-seq denoising using a deep count autoencoder. Nat. Commun. 2019;10:390. doi: 10.1038/s41467-018-07931-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Xu J., Xu J., Meng Y., Lu C., Cai L., Zeng X., Nussinov R., Cheng F. Graph embedding and Gaussian mixture variational autoencoder network for end-to-end analysis of single-cell RNA sequencing data. Cell Rep. Methods. 2023;3 doi: 10.1016/j.crmeth.2022.100382. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.van Dijk D., Sharma R., Nainys J., Yim K., Kathail P., Carr A.J., Burdziak C., Moon K.R., Chaffer C.L., Pattabiraman D., et al. Recovering Gene Interactions from Single-Cell Data Using Data Diffusion. Cell. 2018;174:716–729.e27. doi: 10.1016/j.cell.2018.05.061. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Weinreb C., Rodriguez-Fraticelli A., Camargo F.D., Klein A.M. Lineage tracing on transcriptional landscapes links state to fate during differentiation. Science. 2020;367 doi: 10.1126/science.aaw3381. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Strunz M., Simon L.M., Ansari M., Kathiriya J.J., Angelidis I., Mayr C.H., Tsidiridis G., Lange M., Mattner L.F., Yee M., et al. Alveolar regeneration through a Krt8+ transitional stem cell state that persists in human lung fibrosis. Nat. Commun. 2020;11 doi: 10.1038/s41467-020-17358-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Pijuan-Sala B., Griffiths J.A., Guibentif C., Hiscock T.W., Jawaid W., Calero-Nieto F.J., Mulas C., Ibarra-Soria X., Tyser R.C.V., Ho D.L.L., et al. A single-cell molecular map of mouse gastrulation and early organogenesis. Nature. 2019;566:490–495. doi: 10.1038/s41586-019-0933-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Bastidas-Ponce A., Tritschler S., Dony L., Scheibner K., Tarquis-Medina M., Salinno C., Schirge S., Burtscher I., Böttcher A., Theis F.J., et al. Comprehensive single cell mRNA profiling reveals a detailed roadmap for pancreatic endocrinogenesis. Development. 2019;146 doi: 10.1242/dev.173849. [DOI] [PubMed] [Google Scholar]
- 22.Setty M., Kiseliovas V., Levine J., Gayoso A., Mazutis L., Pe'er D. Characterization of cell fate probabilities in single-cell data with Palantir. Nat. Biotechnol. 2019;37:451–460. doi: 10.1038/s41587-019-0068-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Schroder B. The multifaceted roles of the invariant chain CD74--More than just a chaperone. Biochim. Biophys. Acta. 2016;1863:1269–1281. doi: 10.1016/j.bbamcr.2016.03.026. [DOI] [PubMed] [Google Scholar]
- 24.Swann J.W., Olson O.C., Passegué E. Made to order: emergency myelopoiesis and demand-adapted innate immune cell production. Nat. Rev. Immunol. 2024;24:596–613. doi: 10.1038/s41577-024-00998-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Chakraborty S., Kaur S., Guha S., Batra S.K. The multifaceted roles of neutrophil gelatinase associated lipocalin (NGAL) in inflammation and cancer. Biochim. Biophys. Acta. 2012;1826:129–169. doi: 10.1016/j.bbcan.2012.03.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kjeldsen L., Johnsen A.H., Sengeløv H., Borregaard N. Isolation and primary structure of NGAL, a novel protein associated with human neutrophil gelatinase. J. Biol. Chem. 1993;268:10425–10432. [PubMed] [Google Scholar]
- 27.Gao M., Qiao C., Huang Y. UniTVelo: temporally unified RNA velocity reinforces single-cell trajectory inference. Nat. Commun. 2022;13 doi: 10.1038/s41467-022-34188-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Choksi S.P., Lauter G., Swoboda P., Roy S. Switching on cilia: transcriptional networks regulating ciliogenesis. Development. 2014;141:1427–1441. doi: 10.1242/dev.074666. [DOI] [PubMed] [Google Scholar]
- 29.Dogan G., Ozturk M., Karakulak D.T., Karagenc L. Altered Expression of Pulmonary Epithelial Cell Markers in Fetal and Adult Mice Generated by in vitro Embryo Culture and Embryo Transfer. Cells Tissues Organs. 2024;213:1–16. doi: 10.1159/000527044. [DOI] [PubMed] [Google Scholar]
- 30.Kunugi S., Fukuda Y., Ishizaki M., Yamanaka N. Role of MMP-2 in alveolar epithelial cell repair after bleomycin administration in rabbits. Lab. Invest. 2001;81:1309–1318. doi: 10.1038/labinvest.3780344. [DOI] [PubMed] [Google Scholar]
- 31.Arpino V., Brock M., Gill S.E. The role of TIMPs in regulation of extracellular matrix proteolysis. Matrix Biol. 2015;44–46:247–254. doi: 10.1016/j.matbio.2015.03.005. [DOI] [PubMed] [Google Scholar]
- 32.Barile M., Imaz-Rosshandler I., Inzani I., Ghazanfar S., Nichols J., Marioni J.C., Guibentif C., Göttgens B. Coordinated changes in gene expression kinetics underlie both mouse and human erythroid maturation. Genome Biol. 2021;22 doi: 10.1186/s13059-021-02414-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Bieker J.J., Philipsen S. Erythroid Kruppel-Like Factor (KLF1): A Surprisingly Versatile Regulator of Erythroid Differentiation. Adv. Exp. Med. Biol. 2024;1459:217–242. doi: 10.1007/978-3-031-62731-6_10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Sadlon T.J., Dell'Oso T., Surinya K.H., May B.K. Regulation of erythroid 5-aminolevulinate synthase expression during erythropoiesis. Int. J. Biochem. Cell Biol. 1999;31:1153–1167. doi: 10.1016/s1357-2725(99)00073-4. [DOI] [PubMed] [Google Scholar]
- 35.Sadowska-Bartosz I., Bartosz G. Peroxiredoxin 2: An Important Element of the Antioxidant Defense of the Erythrocyte. Antioxidants. 2023;12 doi: 10.3390/antiox12051012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Cheng Y., Ma X., Yuan L., Sun Z., Wang P. Evaluating imputation methods for single-cell RNA-seq data. BMC Bioinf. 2023;24:302. doi: 10.1186/s12859-023-05417-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Chen Z., King W.C., Hwang A., Gerstein M., Zhang J. DeepVelo: Single-cell transcriptomic deep velocity field learning with neural ordinary differential equations. Sci. Adv. 2022;8 doi: 10.1126/sciadv.abq3745. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Becker-Herman S., Rozenberg M., Hillel-Karniel C., Gil-Yarom N., Kramer M.P., Barak A., Sever L., David K., Radomir L., Lewinsky H., et al. CD74 is a regulator of hematopoietic stem cell maintenance. PLoS Biol. 2021;19 doi: 10.1371/journal.pbio.3001121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Schroll A., Eller K., Feistritzer C., Nairz M., Sonnweber T., Moser P.A., Rosenkranz A.R., Theurl I., Weiss G. Lipocalin-2 ameliorates granulocyte functionality. Eur. J. Immunol. 2012;42:3346–3357. doi: 10.1002/eji.201142351. [DOI] [PubMed] [Google Scholar]
- 40.Deng T., Huang M., Xu K., Lu Y., Xu Y., Chen S., Xie N., Tao Q., Wu H., Sun X. LEGEND: Identifying Co-expressed Genes in Multimodal Transcriptomic Sequencing Data. Genom. Proteom. Bioinform. 2025;23 doi: 10.1093/gpbjnl/qzaf056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Xu J., Lu C., Jin S., Meng Y., Fu X., Zeng X., Nussinov R., Cheng F. Deep learning-based cell-specific gene regulatory networks inferred from single-cell multiome data. Nucleic Acids Res. 2025;53 doi: 10.1093/nar/gkaf138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Li C., Virgilio M.C., Collins K.L., Welch J.D. Multi-omic single-cell velocity models epigenome-transcriptome interactions and improves cell fate prediction. Nat. Biotechnol. 2023;41:387–398. doi: 10.1038/s41587-022-01476-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Lim B., Lin Y., Navin N. Advancing Cancer Research and Medicine with Single-Cell Genomics. Cancer Cell. 2020;37:456–470. doi: 10.1016/j.ccell.2020.03.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Stubbington M.J.T., Rozenblatt-Rosen O., Regev A., Teichmann S.A. Single-cell transcriptomics to explore the immune system in health and disease. Science. 2017;358:58–63. doi: 10.1126/science.aan6828. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Wolf F.A., Angerer P., Theis F.J. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19:15. doi: 10.1186/s13059-017-1382-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Qin Q., Bingham E., La Manno G., Langenau D.M., Pinello L. Pyro-Velocity: Probabilistic RNA Velocity inference from single-cell data. bioRxiv. 2022 doi: 10.1101/2022.09.12.507691. Preprint at. [DOI] [Google Scholar]
- 47.Lange M., Bergen V., Klein M., Setty M., Reuter B., Bakhti M., Lickert H., Ansari M., Schniering J., Schiller H.B., et al. CellRank for directed single-cell fate mapping. Nat. Methods. 2022;19:159–170. doi: 10.1038/s41592-021-01346-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Weiler P., Lange M., Klein M., Pe'er D., Theis F. CellRank 2: unified fate mapping in multiview single-cell data. Nat. Methods. 2024;21:1196–1205. doi: 10.1038/s41592-024-02303-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Akiba T., Sano S., Yanase T., Ohta T., Koyama M. Kdd'19: Proceedings of the 25th Acm Sigkdd International Conferencce on Knowledge Discovery and Data Mining. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework; pp. 2623–2631. [DOI] [Google Scholar]
- 50.Paszke A., Gross S., Massa F., Lerer A., Bradbury J., Chanan G., Killeen T., Lin Z.M., Gimelshein N., Antiga L., et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. arXiv. 2019 Preprint at. [Google Scholar]
- 51.Watanabe S. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance. arXiv. 2023 Preprint at. [Google Scholar]
- 52.Watanabe S., Hutter F. c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperparameter Optimization. Preprint at arXiv. 2023;2022:4371–4379. doi: 10.24963/ijcai.2023/486. [DOI] [Google Scholar]
- 53.McInnes L., Healy J., Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv. 2018 doi: 10.48550/arXiv.1802.03426. Preprint at. [DOI] [Google Scholar]
- 54.van der Maaten L., Hinton G. Visualizing Data using t-SNE. J. Mach. Learn. Res. 2008;9:2579–2605. [Google Scholar]
- 55.Weinreb C., Wolock S., Klein A.M. SPRING: a kinetic interface for visualizing high dimensional single-cell expression data. Bioinformatics. 2018;34:1246–1248. doi: 10.1093/bioinformatics/btx792. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The tutorial and demonstration code for NARVI are available at https://bitbucket.org/tech-kobo/narvi/. The analysis code used to generate the results in this paper is provided as supplemental information in a separate compressed archive: Data S1. No new raw sequencing data were generated; all third-party single-cell datasets analyzed here are publicly available under the accessions listed in the key resources table. The NARVI algorithm is the subject of a patent application. Commercial use of the patented algorithm requires a separate patent licensing agreement. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.




