Abstract
Background
Distinguishing self from non-self is a major challenge for the immune system. Endogenous cytoplasmic double-stranded RNA (dsRNA) can mimic viral RNA and activate immune sensors like MDA5. ADAR1-mediated adenosine-to-inosine editing disrupts base-pairing to suppress immunogenicity of these endogenous structures. Global editing indices are widely used to probe this crucial ADAR1 function. However, they are dominated by nuclear pre‑mRNA edits with limited immune relevance. Here we present the cytoplasmic editing index (CEI) that quantifies editing specifically within inverted Alu repeats in 3′ untranslated regions of mature cytoplasmic transcripts, which potentially form cytosolic dsRNA structures carrying higher immunological risk.
Results
Analyzing over 25,000 RNA-sequencing samples, we demonstrate CEI captures ADAR1p150 activity and outperforms the global editing index in terms of sensitivity and signal-to-noise, enabling sharper tissue-specific profiling, enhanced detection power of infection‑induced editing changes, and stronger association with cancer prognoses. An open-source, cloud-native pipeline delivers end‑to‑end, reproducible analysis at very low cost, supporting immediate, scalable adoption.
Conclusions
CEI provides a refined metric for quantifying immune-relevant RNA editing, revealing previously obscured tissue- and disease-specific editing landscapes. The accompanying open-source, cloud-native pipeline enables broad adoption of high-quality editing analysis across research settings. This approach offers new opportunities for investigating ADAR1's role in immunity, infection, and cancer, with potential applications in biomarker development and therapeutic intervention strategies.
Graphical abstract

Supplementary Information
The online version contains supplementary material available at 10.1186/s13059-026-04154-3.
Keywords: RNA editing, Innate immunity, DsRNA, Cytoplasm
Background
Adenosine-to-inosine (A-to-I) RNA editing is a widespread post-transcriptional modification carried out by adenosine deaminase acting on RNA (ADAR) family of enzymes that convert adenosines to inosines within double-stranded RNA (dsRNA) [1–3]. Among the ADAR enzymes, ADAR1 (ADAR) plays a central role in innate immune regulation. Endogenous dsRNAs that resemble viral RNA may be recognized by cytoplasmic RNA sensors and aberrantly activate antiviral pathways, most notably through the MDA5–MAVS axis and PKR, leading to type I interferon (IFN) response [4–6]. Editing by ADAR1 disrupts the base-pairing of the immunogenic self-dsRNAs to the extent that they are no longer recognized by MDA5, or marks them otherwise, preventing inappropriate activation of the antiviral cellular immune system [7–10]. Aberrant editing has been associated with several autoimmune disorders [11–22], and complete loss of ADAR1 is embryonic-lethal due to MDA5-mediated auto-immunity [7–9, 23–25].
In humans, the vast majority of editing events occur within primate-specific Alu retrotransposons [26–32]. These elements are abundant in primate genomes [33, 34], and frequently form long intramolecular dsRNA structures by pairing of neighboring inverted repeats [29, 35]. Most Alu elements lie in intronic and intergenic regions [34, 36]. Of the small fraction incorporated into mature mRNAs, Alu elements are especially enriched in noncoding regions, mainly the 3′ untranslated regions (3′UTRs). These elements may result in dsRNA structures that engage the cytoplasmic innate immune sensors such as MDA5 [7–9, 35].
Given the essential function of A-to-I editing in preventing inappropriate immune activation, there is growing interest in accurate quantification of editing activity [32, 37–44]. Current quantitative methods remain limited in interpretability and precision. The simplest method — counting edited sites — fails to account for the uneven distribution of Alu element coverage and the typically low editing frequency at individual sites within them. Accordingly, it is highly sensitive to sequencing depth and does not provide a robust estimate for enzymatic activity [45]. More advanced metrics, such as the Alu editing index (AEI) [39], calculate a global score by averaging Alu editing frequencies across the human genome, weighted by expression levels. While these are useful for gauging overall ADAR activity [14, 16, 17, 20, 21, 46–64], they conflate signals from heterogeneous transcripts and subcellular compartments including intronic or nuclear-retained RNAs. The pre-mRNA structures are rarely seen in the cytoplasm [65] and are therefore invisible to the cytoplasmic innate immune sensors. Previous analysis [66] has found that only ~ 0.15% of putative dsRNA structures in pre-mRNA exist in mature mRNAs. Furthermore, the vast majority of the global editing signal is due to a very large number of weakly edited Alu repeats [39, 67], whose structure is largely insensitive to ADAR1 perturbation. Consistently, Sun et al. [68] demonstrated that only 1–2% of cellular dsRNAs are truly immunogenic when unedited, bringing the number of critical MDA5 substrates to hundreds or thousands rather than millions. Consequently, aggregate genome-wide editing indices that are mostly dominated by weakly-edited nuclear structures lack specificity for immune-relevant substrates and may be insensitive to the subtle dynamics of editing that regulate innate immunity.
To address this gap, we introduce a cytoplasmic editing index (CEI): a biologically informed metric that specifically quantifies A-to-I editing events within putative cytoplasmic dsRNAs – clusters of inverted Alu elements residing in 3′UTRs, which are expected to form long intramolecular dsRNA structures capable of activating cytoplasmic immune sensors [69]. Integrating both structural and spatial constraints, this metric provides focused measurement of editing in immunologically relevant substrates. The full repertoire of immunogenic dsRNAs remains undefined. Current evidence [25, 68, 69] suggests that only a fraction of cytoplasmic Alu–Alu duplexes are immunostimulatory, and that additional immunogenic dsRNAs may arise from non‑Alu regions [19, 69, 70]. Nevertheless, we propose using the editing status of Alu elements likely to encounter cytoplasmic dsRNA sensors as a biologically meaningful proxy for ADAR1p150 activity, which correlates with its immunosuppressive function.
We demonstrate that the CEI provides enhanced sensitivity for detecting immunologically relevant editing changes and serves as a valuable tool for investigating RNA editing across diverse biological contexts. Benchmarking across over 25,000 RNA-seq samples from infectious disease and normal tissue datasets, we show that the CEI robustly captures editing dynamics under diverse conditions. To support broad adoption, we provide an accessible, cost-efficient cloud-computing framework enabling scalable and systematic quantification of immunogenic dsRNA editing across diverse experimental settings.
Results
A targeted and scalable index for cytoplasmic A-to-I RNA editing
To develop a cytoplasmic editing index that captures immunologically relevant dsRNA substrates, we first established quantitative criteria for identifying Alu elements with immunogenic potential (Fig. 1a). We defined inverted 3′UTR clusters as regions containing ≥ 2 Alu elements (each > 200 bp) with at least one element oriented in each direction, enabling intramolecular dsRNA formation (Fig. 1d). The selected regions are both retained in cytoplasmic transcripts and capable of forming stable secondary structures detectable by cellular dsRNA sensors.
Fig. 1.

Cloud-based pipeline and biological rationale for quantifying cytoplasmic versus global A-to-I RNA editing. a Biological rationale of the cytoplasmic editing index. For an RNA molecule to trigger immune activation, it must (1) localize to the cytoplasm and (2) adopt a double-stranded conformation therein. Following transcription and splicing, mRNAs are transported to the cytoplasm. Reversely oriented Alu repeats in the 3′UTRs can fold into intramolecular double-stranded RNA (dsRNA) structures. When these cytoplasmic duplexes are insufficiently edited, they are recognized by cellular innate immune pathways and trigger type-I interferon responses. The cytoplasmic editing index specifically quantifies editing within these immunologically relevant dsRNA substrates. b Cloud-based analysis workflow. RNA-seq samples are (1) retrieved, (2) processed through a modular cloud pipeline, and (3) analyzed for A-to-I editing patterns. The pipeline is designed for high-throughput processing and optimized for efficient deployment across diverse dataset sizes, with compatibility across multiple cloud-computing environments. c Resource optimization and scalable processing. Each pipeline step operates with minimal computational requirements (virtual CPU and RAM allocations shown beside icons), with processing costs of approximately $0.10 per sample (for GCP and AWS platforms). The containerized Nextflow implementation enables portable deployment to any Nextflow-compatible cloud or local computing environment. Complete parameter specifications are detailed in Methods. d Cytoplasmic vs. global editing index computation. The cytoplasmic editing index is calculated exclusively from RNA-seq reads mapped to inverted Alu clusters within 3′UTRs (Methods). These clusters are likely to form long, intramolecular dsRNA structures that remain stable after splicing and nuclear export. By contrast, the global Alu editing index encompasses reads aligned to any genomic Alu element, regardless of orientation, size or genomic context
Applying these criteria genome-wide (hg38), we identified 3,343 Alu elements across 895 genes within structured 3′UTR Alu clusters (see Methods). The cytoplasmic editing index (CEI) is defined as the aggregate A-to-I editing level calculated exclusively across these immunologically relevant Alu elements. For comparison, we computed the global Alu editing index (AEI), which aggregates editing across 1,113,103 genomic Alu elements (merged contiguous elements of any length) irrespective of orientation or transcriptomic context [39]. The computational strategies underlying both indices are summarized in Fig. 1d.
To enable large-scale comparative analysis of cytoplasmic vs. global editing patterns, we developed a cost-efficient, modular, cloud-based computational pipeline that retrieves RNA-seq data, performs preprocessing, alignment, and quantifies expression and editing (Fig. 1b). The input file for this workflow is a list of Sequence Read Archive (SRA) sample accession codes. Resource allocation is optimized for each step, resulting in a per-sample processing cost of ~ US $0.10 (Fig. 1c). We validated our approach by analyzing 36,334 publicly available RNA-seq samples from diverse contexts — infectious disease studies (15,649 samples; Additional file 2: Table S1), GTEx [71] (9,602 samples) and TCGA [72] (11,083 samples) — demonstrating robust scalability across heterogeneous datasets. All pipeline scripts and reference files are open-access to support reproducibility (Methods).
Few Alu clusters contribute to the CEI
First, we characterize the 3,343 Alu elements within 3′UTRs contributing to the CEI, and their genomic architecture. We compared these to the 5,344 elements in uniformly oriented Alu clusters, including singletons (one isolated 3′UTR Alu element) or tandem repeats (Fig. 2a, b). The latter are not expected to form cytoplasmic intramolecular dsRNAs. Furthermore, uniformly oriented clusters are present in ~ five-fold more genes than inverted clusters (Fig. 2c; 4,011 vs. 791 genes). Similarly, analysis of murine B1 and B2 repeat clusters in mouse 3′UTRs show that uniformly oriented clusters dominate and inverted clusters are restricted to far fewer genes (Fig. 2d, e). These results suggest conserved evolutionary constraints against dsRNA-forming configurations in 3′UTRs.
Fig. 2.

Structural constraint on inverted versus tandem Alu clusters in human 3′UTRs. a Schematic of Alu configurations in 3′UTRs. Inverted Alu clusters can fold into intramolecular dsRNA that can be edited by ADAR and can activate cytoplasmic RNA sensors if unedited, whereas tandem Alu clusters lack complementary elements and do not form dsRNA. b Orientation breakdown of Alu elements in human 3′UTRs (3,343 inverted and 5,344 uniformly oriented elements). Each element is assigned to the largest cluster size in which it participates within each orientation (singleton, pair or ≥ 3). Eighty elements appear in both orientation classes (reflecting different transcript isoforms) and are thus counted twice. Uniformly oriented elements predominate and are overwhelmingly singletons, whereas inverted elements, though fewer overall, make the majority of multi-element clusters. c Number of human genes harboring inverted and tandem 3′UTR Alu clusters. There are 104 genes exhibiting both types, each in a different isoform. d Orientation breakdown of B1/B2 elements in mouse 3′UTRs (650 inverted and 3,184 uniformly oriented elements). Three B1 elements and one B2 element appear in both orientation classes (reflecting different transcript isoforms) and are thus counted twice. Uniformly oriented elements predominate, as seen in human (b). e Same as c for mouse B1/B2 clusters (Euler diagram). f Distribution of the largest Alu cluster per gene (cluster found in the transcript whose 3′UTR has the largest number of Alu elements among all gene’s transcripts), stratified by orientation. Right inset: possible tandem and inverted configurations for cluster sizes 2–4, and expected inverted:tandem ratios under a random-orientation model (see Methods). Left inset: observed vs. expected number of genes with inverted Alu clusters. Labels above each bar indicate gene counts per category; empirical inverted:tandem ratios (bottom, sizes 2–4) compare to right inset expectations. Inverted Alu clusters span a broader size range than tandem Alu clusters, reflecting the larger combinatorial space of mixed orientations. Nevertheless, inverted Alu clusters are consistently rarer than expected (panel c), indicating negative selection against dsRNA-forming configurations. g Depletion of inverted clusters relative to the random model. Bars represent the observed proportion of inverted clusters divided by the proportion expected when orientations are assigned at equal probability (for Alu cluster sizes with > 10 genes per orientation; Methods; Additional file 1: Fig. S1). Values < 100% indicate underrepresentation, consistent with selective constraint against dsRNA-forming orientations. The proportion of observed inverted clusters declines with cluster size (χ2 test, no multiple testing correction for the three tests; Additional file 1: Fig. S1)
Examining tandem and inverted cluster abundance as a function of cluster size reveals depletion of inverted clusters compared to tandem clusters (Fig. 2f). For size-2 clusters, the number of tandem configurations exceeds that of inverted ones; for sizes ≥ 3, inverted clusters outnumber tandems (Fig. 2f), but not as much as expected in the absence of evolutionary purification (left inset, Fig. 2f). For a size-N cluster, there are orientation combinations, of which are mixed-strand configurations associated with inverted clusters, compared to only two homogeneous-orientation ones associated with tandem clusters. Thus, a random, equally probable and uncorrelated assignment of orientations should result in an inverted:tandem ratio of for a cluster of size N (right inset, Fig. 2f). The observed ratio is well below this expectation (left inset, Fig. 2f), indicating strong depletion of dsRNA-forming configurations (Fig. 2g; Additional file 1: Fig. S1). This finding is consistent with previous reports on evolutionary selection against dsRNA-forming arrangements in mature transcripts [66].
Similar results are obtained when the analysis is restricted to the longest 3′UTR isoform, as well as for analyses of murine B1 and B2 repeat clusters in mouse 3′UTRs (Additional file 1: Fig. S2; Additional file 1: Fig. S3).
We tested whether cluster architecture might be influenced by Alu evolutionary age or subfamily composition, but found only minor differences in sequence divergence and slight deviations from genome-wide proportions (Additional file 1: Fig. S1).
Cytoplasmic index reflects ADAR1p150-dependent immune activation
ADAR1p150, the interferon-inducible cytoplasmic isoform of ADAR1 [2, 73], plays a critical role in preventing cytoplasmic dsRNA from triggering innate immune responses. We thus asked whether the CEI, probing editing in Alu elements within inverted Alu clusters, exhibits elevated sensitivity to editing by ADAR1p150. We compared the CEI to two additional editing metrics: the conventional global Alu editing index (AEI), and a tandem editing index quantifying editing within tandem-oriented Alus (negative control, see Methods). Analysis of knockout [10] and reconstitution [74] datasets revealed that the CEI was selectively sensitive to ADAR1p150 loss or reconstitution compared to the AEI, while the tandem index remained consistently low across all conditions.
As expected, ADAR1 knockout suppressed all editing indices; however, selective loss of ADAR1p150 reduced the CEI more strongly. Similarly, in the reconstitution system, ADAR1 knockdown reduced editing signals in both indices as well. Subsequent ADAR1p150 expression selectively restored the CEI almost to its wild-type levels, while global editing remained diminished (Fig. 3a).
Fig. 3.

Biological perturbations of A-to-I RNA editing reveal differential sensitivity of CEI and AEI. a Impact of ADAR1p150 loss and reconstitution on editing indices. Left [10]: HEK293 cells with ADAR1 and ADAR1p150 knockout (KO) ± IFN-β treatment; right [74]: TBP-ARPE cells with ADAR1 knockdown followed by ADAR1p150 overexpression (each sample cultured and sequenced in triplicate). Loss of ADAR1 activity reduces global editing in both systems, while cytoplasmic editing is selectively impaired by ADAR1p150 absence and restored upon isoform presence, demonstrating that the cytoplasmic index specifically measures ADAR1p150-mediated activity. The tandem index (computed from exclusively tandem Alu elements, see Methods) remains near-zero across all genotypes, validating the assumption that these regions cannot form dsRNA. b ADAR1p150 is required for IFN-β responsiveness. Bars show fold-change in mean editing (IFN-β treated versus untreated) from the ADAR1 KO experiment in panel a (mean ± s.d., n = 3). While IFN treatment enhances the cytoplasmic index in WT cells, this response is abolished in both ADAR1p150 KO and ADAR1 KO cells, confirming isoform specificity. Global and tandem indices show minimal IFN responsiveness in WT cells. c Temporal dynamics of cytokine-induced editing in EndoC-ꞵH1 cells. Two independent time-course datasets (IFN-α, left [75]; mixed cytokines, right [76]) reveal index trajectories over 2–48 h. The cytoplasmic index rises earlier and more dramatically than the global index, while the tandem index remains consistently low and largely unaffected by cytokine treatment. d Fold-induction for datasets shown in panel c (mean ± s.d., n = 5). The CEI consistently exhibits greater dynamic induction compared to the AEI. Apparent fold-changes in the tandem index do not reflect biologically meaningful induction, as the absolute editing levels remain near baseline (panel c). e Species conservation of ADAR1p150 dependence. Left [78], tissue-specific patterns: brain (predominantly ADAR1p110-expressing) shows marked reduction of CEI and AEI upon ADAR1p110 knockout, with a further decrease in the ADAR1p110/ADAR2 double knockout. In contrast, thymus (predominantly ADAR1p150-expressing) maintains high editing across genotypes. Right [79], rescue experiments: editing-deficient mouse embryonic fibroblasts (MEFs; ADAR1⁻/⁻; ADARB1⁻/⁻; Gria2ᴿ/ᴿ) show no editing when expressing RFP controls, modest CEI restoration with ADAR1p110, and complete rescue with ADAR1p150. In both contexts, AEI and tandem indices remain consistently flat. f Subcellular specificity. Analysis of subcellular fractions from HeLa cells [80] reveals highest CEI in the cytoplasm, decreasing progressively in nuclear and chromatin fractions, whereas AEI and tandem indices remain relatively constant across compartments
In addition, the knockout dataset included IFN-β stimulation, enabling assessment of interferon‑induced editing responses across genotypes. In wild-type cells, IFN-β treatment induced a marked increase in the CEI compared to untreated controls, while the AEI showed no significant change and the tandem index remained largely unresponsive. IFN-induced enhancement of cytoplasmic editing was absent in both complete ADAR1 knockout and selective ADAR1p150 knockout cells (Fig. 3b), confirming that cytoplasmic editing induced by interferon specifically requires ADAR1p150.
Time-course analyses in two independent datasets [75, 76] revealed that CEI rose faster and higher than AEI following IFN-α or cytokine treatment, whereas the tandem index remained largely unaffected (Fig. 3c,d). Similarly, the cytoplasmic index exhibits preferential responsiveness to cytokines in human islets [75, 77] (Additional file 1: Fig. S4). The contrasting behavior of inverted versus tandem Alus, despite both being part of cytoplasmic 3′UTRs, demonstrates that structural architecture, not subcellular context alone, determines immune-responsive editing activity.
Analysis of mouse tissues [78] and embryonic fibroblasts [79] demonstrated evolutionary conservation of isoform-specific cytoplasmic editing. In mouse thymus, predominantly expressing ADAR1p150 [78], cytoplasmic editing does not change much upon knockout of ADAR1p110 and ADAR2. Conversely, in mouse brain tissue, primarily expressing ADAR1p110 [78], ADAR1p110 knockout causes marked CEI reduction; combined ADAR1p110/ADAR2 knockout further diminishes this effect. In editing-deficient fibroblasts (ADAR1⁻/⁻; ADARB1⁻/⁻; Gria2ᴿ/ᴿ), cytoplasmic editing was rescued robustly by ADAR1p150, partially by ADAR1p110, and not at all by an RFP control. Across these conditions, the effect on AEI and tandem index is limited, confirming the isoform specificity of cytoplasmic editing (Fig. 3e).
Direct assessment of subcellular localization further supported the cytoplasmic specificity of the CEI. Fractionation experiments in HeLa cells [80] revealed highest CEI in cytoplasmic fractions, intermediate in nuclear fractions, and lowest in chromatin fractions, whereas the AEI and tandem indices presented negligible compartmental variation (Fig. 3f). Thus, CEI demonstrates isoform specificity absent in conventional editing indices, with particular utility for monitoring immune activation and infection responses.
The CEI is defined as a weighted average of editing levels across individual elements, with weights proportional to their expression levels. As such, it is inherently sensitive to both editing efficiency and the relative expression of each element. In principle, a change in CEI could therefore reflect shifts in the expression landscape rather than genuine changes in editing activity. However, because the CEI integrates signals from thousands of regions – each contributing only marginally to the overall index – the influence of expression is likely limited. To empirically assess this, we examined the change in CEI upon IFN induction as reported in Fig. 3a.
We first characterized the contribution of individual elements to the CEI. As expected, in both wild-type (WT) and IFN-treated samples, the individual contribution of each element is very small; > 99.7% of the elements in each condition account for less than 1% of the total index each (Additional file 1: Fig. S5a). Thus, even large changes in the expression of isolated elements are unlikely to substantially affect the CEI. Next, we examined per-element editing levels, which are unaffected by expression variation, and observed that the majority of elements exhibited increased editing upon IFN treatment (Additional file 1: Fig. S5b). Finally, to disentangle the contributions of expression and editing, we constructed two pseudo-indices: one that applies IFN-induced editing levels to the WT expression profile, and another that combines the IFN-induced expression profile with WT editing levels. This decomposition clearly separates the two effects and demonstrates that the observed increase in CEI upon IFN stimulation is predominantly driven by elevated per-region editing levels, with only a minimal contribution from expression changes (Additional file 1: Fig. S5c).
Collectively, these findings establish CEI as an ADAR1p150-sensitive, biologically grounded, and computationally efficient metric that specifically captures editing activity in immunologically relevant dsRNA substrates, a small but critical fraction of all cellular dsRNAs [68, 69], with particular utility for monitoring immune activation and infection responses.
Cytoplasmic editing index provides superior signal, specificity and computational performance
We next evaluated the technical performance of the CEI relative to the widely used global Alu editing index (AEI) [39]. Leveraging our standardized cloud-based computational pipeline (Fig. 1b,c), we quantified both indices across 15,649 publicly available RNA-seq samples from 55 infectious-disease datasets, enabling consistent large-scale benchmarking (Figs. 1b,c; Fig. 5a; Methods). We compared editing signal, signal-to-noise ratio, and mean coverage per editable site to assess the relative advantages of the two indices.
Fig. 5.

Cytoplasmic editing index enables detection of immune-regulated RNA editing changes. a Overview of RNA-seq sample collection and analysis workflow. A total of 15,649 publicly available RNA-seq samples spanning 55 datasets and over 20 distinct infectious diseases were compiled and analyzed using standardized cloud-based pipelines to quantify the CEI and AEI. b Comparison of effect sizes (Cohen's d, mean difference divided by pooled standard deviation) for case–control analyses between CEI and AEI across infectious disease datasets (for all case–control comparisons with ≥ 4 samples per group, ≥ 10 total samples). The green triangle indicates the region where the CEI identified a greater positive effect size than the AEI, demonstrating enhanced biological sensitivity. Points are colored based on statistical significance detected exclusively by CEI, exclusively by AEI, by both indices, or by neither index (Wilcoxon test, paired samples where applicable, see Methods). Case–control status was assigned from sample metadata and pathogen load annotations (Methods; Additional file 2: Table S1; Additional file 2: Tables S3-S4). c Longitudinal CEI analysis of tuberculosis treatment and recovery [81] in human samples. Boxplots show editing index trajectories at several time points after treatment initiation, grouped by recovery time (either week 4 or weeks 8–12). The CEI decreases significantly after recovery, while the AEI shows no significant changes (Wilcoxon rank-sum test, raw p-values denoted under comparisons). d Variability of RNA editing indices across TCGA cancer samples. Each point corresponds to one tissue type, and presents the normalized standard deviation (s.d./mean editing) of the index across samples in primary tumors and normal tissues. Inter-tumor variability of CEI is significantly greater than AEI in primary tumors (Wilcoxon signed-rank test), whereas normal tissues show no significant difference between indices. e Kaplan–Meier survival analysis across TCGA cancer types with ≥ 100 samples. Adjusted P-values (Benjamini–Hochberg correction) represent the optimal separation of samples into two groups based on the editing index (Methods). The two indices showed complementary patterns of survival association, with CEI performing better in some cancer types and AEI in others. f Linear modeling (CEI ~ AEI) across antemortem donors of GTEx for tissues with ≥ 10 samples. Points represent model slopes (± standard error) per tissue, ranked by slope magnitude. Inset shows linear fits for tissues with the highest (stomach) and lowest (muscle skeletal) slopes, highlighting variable tissue-specific relationships between cytoplasmic and global editing indices. g Postmortem vs. antemortem editing index stability across GTEx tissues with ≥ 10 samples in both cohorts. Bars indicate postmortem-to-antemortem mean editing ratios for CEI and AEI, highlighting the reduced susceptibility of CEI to postmortem RNA degradation. Insets show distributions of postmortem and antemortem CEI and AEI values in adipose visceral tissue, exemplifying the robustness of CEI against tissue-degradation artifacts
For both indices, A-to-G mismatches consistently ranked first in frequency (> 99% of samples), with C-to-T second (> 93%; Additional file 1: Fig. S6). The values reported for the two indices were correlated across samples, but CEI consistently outperformed AEI across multiple metrics. CEI delivered a ~ 2.6-fold higher editing signal than AEI (ordinary least squares slope = 2.57; Fig. 4a) and exhibits a 2.15-fold higher signal-to-noise ratio (Fig. 4b), reflecting improved sensitivity and specificity for A-to-I events. This enhanced performance stems partly from CEI's ~ 16-fold greater average coverage per editable site (slope = 16.5; Fig. 4c), reflecting its restriction to mature, cytoplasmic transcripts. These gains were obtained without increasing variability across datasets (Fig. 4d), underscoring CEI’s efficiency and robustness for large-scale RNA editing analysis. Given that ~ 0.3% of the genomic Alu elements are included in the cytoplasmic index, their coverage is ~ 16.5-fold enriched, and their editing ~ 2.15-fold enhanced, we estimate that approximately 11% of the A-to-G mismatches contributing to the AEI are captured in the CEI.
Fig. 4.

Cytoplasmic editing index demonstrates superior signal, specificity, and efficiency over global editing metrics. a–c Scatter-density plots compare the CEI and AEI in terms of a) the editing signal (A-to-G index, b) signal-to-noise ratio (SNR), and c) mean coverage per editable site. Points are colored by sample density; dashed black lines mark x = y, and red lines are ordinary-least-squares fits through the origin, giving slopes of 2.57, 2.15 and 16.5, respectively. SNR is the A-to-G index divided by the maximum non-A-to-G mismatch index (four samples with infinite SNR were excluded). Average coverage is computed per editable site (Methods). d Dataset-level variability. Each point is the standard deviation of the index values across all samples within a single dataset normalized by its mean editing level, restricted to samples with noise < 0.3 in both indices (dataset sizes 22–1,299 samples, Additional file 2: Table S2). No significant difference is found between indices (Wilcoxon signed-rank). e Empirical cumulative distribution of editing noise (maximum non-A-to-G mismatch index per sample). Although the two curves largely overlap, the global index has a slightly larger fraction of samples under the threshold of 0.3 (vertical dashed line), consistent with slightly lower background mismatch variability. Inset: noise density histogram. f, g, Computational efficiency. Bars show the mean ± s.d. wall-clock runtime per sample (f) and peak resident-set memory per sample (g), both recorded with the GNU time utility. Both metrics are substantially lower for the CEI, underscoring its performance advantage
The distribution of the noise per sample is similar for both indices, with a slight advantage to the global index (Fig. 4e). In addition, the cytoplasmic index remained robust to sub-optimal mapping conditions (Additional file 1: Fig. S6). Finally, CEI required substantially lower compute time and memory usage than AEI (Fig. 4f,g), highlighting its efficiency. These advantages are recapitulated in 9,602 Genotype-Tissue Expression [71] (GTEx) samples (Additional file 1: Fig. S7).
The cytoplasmic editing index increases sensitivity to editing alteration in infection
We next compared use of CEI and AEI to detect immune-related editing alterations at the population level across diverse infections. Using our uniformly processed compilation of 15,649 RNA-seq samples spanning more than 20 pathogens and associated controls, we computed CEI and AEI for each sample (Fig. 5a), and searched for infection-associated editing changes (Fig. 5b).
Across infections, CEI consistently showed larger absolute effect sizes and greater sensitivity, achieving statistical significance in 46 comparisons, compared to only 28 significant changes in AEI (Fig. 5b). This advantage was especially pronounced for editing increases during active infection, consistent with CEI's sensitivity to ADAR1p150 upregulation, and was most evident for viral infections and respiratory conditions (Additional file 1: Fig. S8). In a longitudinal tuberculosis cohort [81], CEI captured a progressive decrease in editing after treatment and recovery (Fig. 5c), whereas AEI showed no significant changes, highlighting CEI’s capacity to track dynamic, immune-driven regulation of RNA editing.
Recent studies have highlighted the importance of the ADAR protein in the development of autoimmune and autoinflammatory diseases, most notably in Aicardi–Goutières syndrome [13] (AGS) and type 1 diabetes [21]. These findings might lead one to expect elevated CEI values in such conditions. However, the role of ADAR-mediated editing has primarily been demonstrated during the initial, critical stages of disease onset, and may not be accurately captured in bulk tissue profiles obtained at later time points.
To investigate whether altered editing patterns could still be detected in autoimmune and autoinflammatory diseases, we analyzed RNA-seq data from a diverse set of conditions, including multiple sclerosis (MS), Crohn’s disease, atopic dermatitis, psoriasis, Sjögren's syndrome, and ulcerative colitis [82–86]. In these datasets, we observed no significant differences in the overall editing index (Additional file 1: Fig. S9). These findings likely reflect the relatively advanced disease stages represented in the cohorts examined, where multiple feedback loops and compensatory pathways have already been activated. Moreover, clinical tissue samples are often composed of heterogeneous cell populations, such as whole blood or lesional tissue with infiltrating immune cells, which may dilute or obscure disease-relevant editing changes occurring in specific cell subsets.
Nevertheless, it is worth noting that our previous work did identify a general reduction in RNA editing levels in patients with atopic dermatitis and psoriasis [14, 20]. This suggests that disease- and context-specific effects on editing can still occur, even when global CEI measurements appear unchanged.
Improved detection of RNA editing variability in cancer and across normal tissues using the CEI
We next examined whether the enhanced sensitivity of CEI extended beyond infectious diseases. We first looked at variability and patient stratification in cancer samples and controls from the TCGA dataset. CEI captured greater editing variability than AEI across primary tumors but not normal tissues, underscoring its capacity to detect heterogeneity in tumor-associated editing changes (Fig. 5d). Furthermore, Kaplan–Meier survival analyses showed that CEI and AEI provided better patient stratification in different cancer types across TCGA, highlighting complementary strengths and underscoring the clinical value of using both metrics (Fig. 5e). CEI showed stronger signal and lower noise relative to AEI in TCGA tumors, and the indices differed in the cancer types they most strongly distinguished, reflecting complementary strengths (Additional file 1: Fig. S10).
To assess the robustness of CEI across tissues and donor cohorts, we analyzed normal samples from both antemortem (living samples: organ donor, surgical) and postmortem donors, taken from the GTEx [71] dataset. The editing index is affected by postmortem processes, such as hypoxia, cellular metabolism and apoptosis [39, 60], making this distinction critical for evaluating tissue-specific measurements. The correlation between CEI and AEI varied substantially across tissues, revealing that different tissues exhibit specific immune-relevant editing landscapes (Fig. 5f; Additional file 1: Fig. S11). As a notable exception, cortex-derived brain tissues uniquely exhibited CEI values lower than the AEI ones, potentially reflecting distinct editing dynamics driven by high expression of ADAR2 and ADAR3 in this neural tissue [47, 87] (Additional file 1: Fig. S11). Importantly, CEI was less susceptible to postmortem changes (Fig. 5g; Additional file 1: Fig. S11). These findings underscore the stability of CEI across donor cohorts and support its use for quantifying RNA editing in both high-quality and partially degraded tissues.
Discussion
Quantifying immunologically relevant A-to-I RNA editing represents a critical challenge in understanding innate immune regulation, as existing metrics fail to capture the biological specificity required for mechanistic insights. The cytoplasmic editing index (CEI) provides a biologically informed complement to existing RNA editing metrics by targeting the subset of editing events most likely to influence cytoplasmic dsRNA sensing.
The qualitative difference between CEI and global editing metrics is manifested in CEI's selective ADAR1p150 dependence. Notably, comparison with recently mapped immunogenic dsRNA regions revealed substantial overlap with CEI targets: 41% overall (559/1,355), increasing to 76% (542/710) for 3′UTR regions and 81% (111/137) for those with the highest cytoplasmic editing levels [68]. Consistently, the significant depletion of inverted Alu clusters relative to random expectations provides an evolutionary validation for CEI's targeting strategy, indicating a strong selective pressure against dsRNA-forming arrangements, consistent with previous findings [66]. This evolutionary constraint, shown here for human and mouse, confirms that the repeats targeted by CEI targets are enriched in functionally critical editing substrates.
The specific double‑stranded RNAs that trigger innate immune activation remain poorly characterized. Potential contributors to the cytoplasmic pool of immunogenic dsRNAs include cis‑natural antisense pairs, Alu‑derived structures not represented in RefSeq transcripts but present in the cytoplasm due to intron retention or 3′UTR extension, and dsRNAs formed from other repetitive elements such as endogenous retroviruses (ERVs) [19, 60, 69]. The CEI does not account for these largely uncharacterized putative substrates of ADAR1p150 or related RNA‑processing enzymes. Moreover, current evidence suggests that some cytoplasmic Alu–Alu duplexes included in the CEI are immunogenically inert [25, 69, 70]. Given that the complete repertoire of immune‑relevant dsRNAs is still unknown, the CEI framework deliberately focuses on a defined subset of ADAR1p150 editing targets to yield a measurable indicator of enzyme activity, while acknowledging that not all of this activity is directly linked to immune modulation.
The CEI is a weighted average signal across numerous Alu clusters and may therefore be influenced by variations in their relative expression, even when the editing rate at individual elements remains constant. Such shifts can arise from changes in the expression of the host gene or from alterations in isoform composition. In particular, alternative polyadenylation, frequent in cancer and immune activation, can lead to 3′UTR shortening, resulting in the loss of Alu‑containing regions or the removal of one member of an editable Alu pair, thereby impacting the CEI. Our analyses indicate that these transcriptomic effects are largely averaged out in the aggregated CEI signal, and the dominant determinant of CEI variation is ADAR1p150 enzymatic activity. Nonetheless, coordinated 3′UTR remodeling, such as widespread shortening, could disrupt this relationship. Future work integrating isoform‑resolved quantification and RNA secondary‑structure profiling may help disentangle changes in editing efficiency from changes in substrate availability.
Our scalable, cost-effective computational pipeline delivers substantial technical advantages, and is particularly suitable for large-scale or heterogeneous datasets. Per-sample processing costs are reduced to ~ $0.10 while maintaining reproducibility through containerized, cloud-native workflows. This proved crucial for our analysis of > 25,000 samples across diverse contexts, demonstrating that CEI consistently outperformed AEI with higher editing signal, improved signal-to-noise ratio and higher coverage per editable site, without increased variability. These advantages translated into enhanced biological detection across infectious diseases, complementary patient stratification for cancer contexts, and reduced susceptibility to postmortem artifacts.
Importantly, our implementation of the CEI is highly accessible and enables widespread adoption. The user is required to supply just a list of SRA accession codes, and all steps from data download to index calculation are fully automated.
Conclusions
CEI's enhanced sensitivity suggests a possible clinical utility for monitoring immune status in infectious diseases and autoimmune conditions where ADAR1 dysfunction drives pathology. The improved cancer stratification hints at potential prognostic applications, though prospective validation would be required. By targeting functionally meaningful substrates rather than aggregate editing, CEI provides a powerful analytical framework for dissecting how A-to-I editing shapes immune regulation in health and disease, complementing existing global metrics.
Methods
Terminology
Alu element
A single Alu element > 200 bp as defined in the RepeatMasker database. It may or may not overlap other Alu elements.
Alu cluster
The set of all Alu elements located within a 3′UTR of a curated RefSeq transcript. Alu elements that overlap the 3′UTR partially and exceed beyond its limits are trimmed to include only the subsequence overlapping with the 3′UTR. Two distinct 3′UTR isoforms may define two distinct Alu clusters.
Inverted Alu cluster
An Alu cluster containing at least one element of each orientation.
Tandem Alu cluster
An Alu cluster that is not an inverted Alu cluster, where all Alu elements have the same orientation.
Infectious disease data collection and annotation
Fifty-five publicly available infection-related RNA-seq datasets, comprising 15,741 RNA-seq samples, were obtained from the NCBI Sequence Read Archive (SRA) [88]. Datasets were annotated via Sequence Read Archive (SRA) metadata, with additional curation from Gene Expression Omnibus (GEO) [89] SOFT files when metadata were incomplete. The datasets are listed in Additional file 2: Table S1. Accession numbers, metadata, and per-sample annotations are summarized in Additional file 2: Table S4.
Each study was classified as case–control or alternative design (e.g., longitudinal, treatment). For longitudinal datasets, two time-points were selected as case and control based on study specifications, as detailed in Additional file 2: Table S1. The time-point with the highest pathogen load was chosen as the case group whenever possible, unless a different selection criterion was noted in the primary publication.
For multi-strain pathogens, strains were analyzed separately only if each strain group contained at least four samples. Otherwise, strains were combined into a single group. Datasets containing multiple libraries from the same donor were treated as independent observations; no donor-level random-effect correction was performed. Case–control comparisons are detailed in Additional file 2: Table S3. Pathogen classification by clinical manifestation and pathogen type is listed in Additional file 2: Table S6.
Of 15,741 samples in this cohort, 92 failed to run due to various issues, most commonly because of mismatched paired-end lengths (e.g., 50 bp + 150 bp) or missing FASTQ files. The failed samples and a detailed breakdown of failure reasons are provided in Additional file 2: Table S5.
ADAR1 perturbation data collection
In addition, we analyzed 147 publicly available RNA-seq samples from datasets profiling genetic or cytokine perturbations of RNA editing regulators in human and mouse along with cell fractions [10, 74–80]. These datasets included experimental conditions such as ADAR knockdown or reconstitution, as well as interferon stimulation time courses along with sequencing of cell fractionation protocol. Accession numbers, metadata, and sample groupings for these experiments are provided in Additional file 2: Table S7.
Autoimmune diseases data collection
We analyzed 384 RNA-seq samples from 4 independent publicly available datasets encompassing 9 autoimmune conditions in human tissues. Accession numbers and disease-control groupings for these datasets are provided in Additional file 2: Table S8.
GTEx data collection and annotation
Raw RNA-seq data for 9,602 samples from the Genotype-Tissue Expression v10 (GTEx) [71] project were obtained via dbGaP [90] (phs000424). Only normal samples sequenced with paired-end, unstranded, short-read protocols with available FASTQ files via SRA were included. Surgical and organ-donor samples were classified as antemortem (tissues collected from living donors), whereas postmortem samples originated from deceased donors. Accession numbers, metadata, and sample group assignments for all analyzed GTEx samples are provided in Additional file 2: Table S9.
TCGA data collection and analysis
Aligned RNA-seq data from The Cancer Genome Atlas [72] (TCGA) were obtained from the Genomic Data Commons (GDC) [91] via dbGaP (phs000178). We excluded multi-mapped and unmapped reads, to retain only uniquely mapped primary alignments compatible with our reference annotations, using the command samtools view -h -b -q 255 -F 2308 -f 2. Of the initial 11,083 samples, 452 samples were excluded due to the absence of reads after filtering. Accession numbers, metadata, filtering outcomes, and sample group assignments for all analyzed TCGA samples are provided in Additional file 2: Table S10. Detailed workflow adaptations for TCGA data are documented and available on GitHub.
Cloud-based computational pipeline and data processing
All computational analyses used our cloud-optimized Nextflow pipeline engineered for cost-efficient processing of high-throughput RNA-seq data. The workflow runs on Google Cloud Platform (GCP) Batch and Amazon Web Services (AWS) Fargate, but is portable to other providers owing to its modular design. All steps are executed using version-pinned Docker images to ensure reproducibility across analyses. Stage-specific resource allocation reduced average end-to-end cost to ~ $0.10 per sample (Fig. 1c).
Initialization: a one-time initialization step retrieves all reference annotations and generates index files required for expression quantification, alignment and RNA quantification. The resources are then deposited in a bucket of the user’s choice. This step is done once per computing platform and organism.
Per-sample processing: this includes per-sample tasks — read retrieval, quality control, expression quantification, alignment to the genome, and RNA-editing quantification. For each sample, these steps are executed sequentially in discrete containers on resource-optimized virtual machines. Samples are processed in parallel, enabling thousands of concurrent workflows and simplifying quota management and monitoring. The workflow accepts a list of SRA accession IDs (and an NGC file, for dbGaP data) and produces sample-level outputs ready for downstream analysis.
Retrieving all reference annotations: reference annotations and expression files are downloaded from UCSC Genome Browser [92] (http://hgdownload.soe.ucsc.edu/goldenPath; https://hgdownload.soe.ucsc.edu/gbdb) canonical data sources and processed using standard genomics tools (awk, bedtools [93] and bigBedToBed [92]). Specific resource URLs are detailed in GitHub.
Generation of index files required for expression quantification: Salmon [94] (version 1.10.2) index was generated with RefSeq [92] (latest available version as of early 2024) as the reference transcriptome. The entire genome was used as the decoy sequence as per Salmon’s comprehensive index generation protocol.
Generation of index files required for alignment: for each read-length group we built a length-matched STAR [95] index (--sjdbOverhang = R, where R ∈ {50, 75, 100, 125, 150}).
Generation of index files required for RNA editing quantification: for each regions file, we built the appropriate JSD index using the RNAEditingIndexer. This index is used during the editing quantification step to avoid per-sample re-generation and reduce resource usage.
Resource deposit: resources are transferred to a bucket supplied by the user via AWS CLI or gcloud CLI, according to the cloud provider.
Read retrieval and quality control: raw sequencing reads were retrieved using NCBI SRA Toolkit (v3.0.7), and preprocessed with fastp [96] (v0.23.4) for quality filtering, with non-default options -q 25 -u 20 -e 30 --dont_eval_duplication -G -A. For each dataset (library), we extracted X, the average read length per mate as annotated in SRA. We defined L as the element in the set of standard lengths (L ∈ {50, 75, 100, 125, 150}), that is closest to X. If X ≥ L-3, reads shorter than L-3 bp were discarded using the flag --length_required L-3. Reads longer than L were trimmed to this size from the 3’ using --max_len1 L. Else (X < L-3), reads shorter than X-3 bp were discarded using the flag --length_required X-3, and reads longer than X were trimmed to this size from the 3’ using --max_len1 X. In both cases, the alignment index for this data was set to L. In some cases, SRA annotations were found to be inconsistent with read data and were manually curated. All read lengths decisions for both X and L are detailed in Additional file 2: Table S4 at sample level. Studies containing multiple read-length groups (e.g., 50 bp and 100 bp libraries) were split, and each group of specific read-length was processed independently. GTEx samples were processed using X = L = 75 bp as per their documented sequencing protocol.
Expression quantification: transcript expression levels were quantified using Salmon [94] (version 1.10.2) with automatic library detection (-l A) and transcript-to-gene mapping (-g) using a relevant resource file.
Alignment to the genome: unique alignments were generated with STAR [95] (v2.7.10b), with the relevant index and the following parameters:
--alignSJoverhangMin 8 --alignIntronMax 1000000 --alignMatesGapMax 600000 --outFilterMultimapNmax 1 --outFilterMatchNminOverLread 0.95. The index length was chosen to match preprocessed read lengths as previously detailed. For indices robustness testing, we selected 1,000 random samples and realigned them with the non-AEI-optimal default STAR parameter --outFilterMatchNminOverLread 0.66 while keeping other flags unchanged. Comparison of noise is presented in Additional file 1: Fig. S6b.
RNA-editing quantification: RNA editing levels within Alu regions (AEI, CEI, tandem index) were quantified from STAR alignments using RNAEditingIndexer [39] (v1.0) with RefSeq curated data as reference. Common variants from dbSNP151 [97] (hg38) for human samples or dbSNP142 (mm10) for mouse samples were filtered to exclude known polymorphisms from the analysis. RNA editing levels of known sites of human [98] and mouse [99] are quantified via pipeline using a custom-built program based on samtools [100] mpileup file format.
Resource usage monitoring: runtime and peak memory usage were measured using the GNU time utility and aggregated across samples to benchmark computational efficiency. Results are presented in Fig. 4f and g.
The complete Nextflow workflow, pipeline initialization and documentation are available at https://github.com/a2iEditing/CEI [101]. All known sites and CEI annotations, including pre-computed inverted Alu cluster definitions for human and mouse reference genomes (hg38 and mm10 accordingly), are provided in the repository.
Cytoplasmic index regions
The cytoplasmic editing index is calculated over inverted Alu clusters. These were identified by scanning 3′UTR exons for ≥ 2 Alu elements (each exon separately), each spanning > 200 bp, with at least one Alu element in each of the two strand orientations. Alu elements ≤ 200 bp were excluded, even when positioned inside a candidate cluster. For mouse regions, B1 and B2 families were analyzed separately, with clusters considered inverted when at least one element within a family showed inverted orientation relative to other elements in the same family. Mouse repeats were required to span > 120 bp.
Elements extending beyond annotated 3′UTR boundaries were trimmed to retain only the sequence overlapping the 3′UTR. Transcriptomic annotations were obtained from UCSC Genome Browser [92] (hg38 and mm10 RefSeq Curated, February 2021). Adjacent or overlapping Alu elements were merged with pybedtools [102] (v0.9.1, wrapping bedtools [93] v2.30). The lists of clusters and associated elements are available in Additional file 2: Table S11, Additional file 2: Table S12 and the project’s GitHub repository.
Tandem index regions
Tandem index regions are Alu elements that belong to tandem Alu clusters, do not belong to any inverted Alu cluster, and in addition do not have an inverted Alu element of any length (even ≤ 200 bp) occurring anywhere in the same 3′UTR, to ensure structural independence from potential dsRNA-forming configurations. The same transcript annotations (hg38 and mm10 RefSeq Curated, February 2021) and 3′UTR-trimming rules were applied. Adjacent elements were merged with pybedtools [102] (v0.9.1; bedtools v2.30) using the same parameters as for inverted clusters. For mouse samples, tandem regions were defined separately for B1 and B2 families. The resulting tandem set represents non-dsRNA-forming repetitive element clusters and serves as a control for evaluating orientation-specific editing activity.
RNA editing index calculations
RNA editing levels were quantified using RNAEditingIndexer [39] (v1.0). Two editing indices were computed for all samples, with a third index calculated for specific analyses:
Global editing index (AEI): quantified editing levels across all annotated Alu elements genome-wide using default hg38 settings for human samples or mm10 settings for mouse samples (B1 and B2 elements).
Cytoplasmic editing index (CEI): quantified editing within inverted Alu clusters (or inverted B1/B2 elements for mouse) as defined above, using the -rb flag. This metric restricts analysis to cytoplasm-localized, dsRNA-forming clusters and serves as the primary biologically-informed alternative to the AEI.
Tandem index: calculated for control analyses (Fig. 3) using non-inverted 3′UTR Alu elements via the -rb flag as a structural control.
All indices used mismatch profiles across specified regions, masking common SNPs (dbSNP151 for hg38; dbSNP142 for mm10). RNAEditingIndexer incorporates gene expression estimates to guide strand assignment and transcript disambiguation during editing quantification (for algorithmic details, see Ref. [39]).
RNA editing metrics for all samples included in this study are available in Additional file 2: Table S8 and Additional file 2: Tables S13–S16.
Inverted cluster count model
To compute the expected ratio of the numbers of inverted and tandem clusters, we noted that of all 2ᴺ possible allocations of orientation to elements in a cluster with N Alu elements, only two configurations result in a tandem cluster — either all forward or all reverse. Assuming each element is independently oriented, with equal probability for forward or reverse orientation, one expects an inverted-to-tandem ratio:
Calculation of average coverage per editable site
Mean coverage per editable site is calculated as the total read depth at all A and T positions in the reference sequence, divided by the number of such positions. We excluded contributions from mismatches other than A-to-G and T-to-C. Both forward strand A positions and reverse strand T positions are included since A-to-I editing can occur on either strand depending on transcript orientation.
Statistical analysis of infectious diseases cohort
Case–control comparisons in infection datasets were assessed with Wilcoxon rank-sum tests on noise-filtered data (maximum non-A-to-G mismatch < 0.3); Wilcoxon signed-rank tests were applied when study design allowed. Comparisons were considered only when filtered case and control groups included ≥ 4 samples each and ≥ 10 samples in total. If a dataset included multiple tissues or read lengths, each tissue and read length were compared separately. Technical replicates were merged after noise filtering by averaging over editing measures across all replicates. For paired comparisons, unpaired samples (due to experimental design or filtering) were excluded.
Calculation of standard deviation for ratio of means
Standard deviations for ratios of group means were calculated using error propagation. For each group, the sample mean and standard deviation were computed using the unbiased estimator (denominator n-1). We estimated the coefficient of variation as for each group.
For the ratio , the variance was calculated using the delta method, assuming independence between groups. The standard deviation of the ratio is:
Linear modeling and comparative analysis of GTEx cohort
To compare the two indices using linear modeling, we first filtered out noisy samples (maximum non-A-to-G mismatch index ≥ 0.3 for either AEI or CEI). Donor replicates within tissues were merged by averaging over editing measures across all replicates. The filtered and merged cohort included 9,097 samples.
Ordinary least squares regression (CEI ~ AEI, no constant term) was applied to each tissue type with ≥ 10 samples, to find the slopes. To compare antemortem and postmortem cohorts, we repeated the analysis for each cohort separately, requiring ≥ 10 samples in both cohorts.
Survival analysis and comparative analysis of TCGA cohort
Following noise filtering and sample merging as described above, the TCGA cohort included 10,488 samples.
For each TCGA cancer type with ≥ 100 samples, we split the data based on index values to perform survival analysis. Samples were ranked by editing index value and divided into high and low groups using nine different thresholds at 10% increments (10th percentile, 20th percentile, …, 90th percentile, see Additional file 2: Table S17). For each threshold, Kaplan–Meier [103] survival analysis was performed in R v4.5.0 using the survival package (v3.8–3) (survfit()) and log-rank testing (survdiff()). P-values were adjusted for multiple testing (for all cancer types and thresholds, CEI and AEI separately) using the Benjamini–Hochberg procedure to control the false discovery rate, and the threshold yielding the lowest log-rank P-value was chosen. This approach ensures fair comparison between indices by allowing each to be evaluated at its optimal discrimination threshold, while controlling for the multiple comparisons inherent in the threshold optimization procedure.
Statistics
All statistical tests presented in the article are two-sided.
Supplementary Information
Additional file 1: Supplementary figures. This file includes supplementary figures S1-S11.
Additional file 2: Supplementary tables. This file includes supplementary tables S1-S18. Supplementary tables description: Table S1: Infectious-disease dataset annotation. Per-dataset metadata for all infectious-disease studies analysed, including study design, pathogen, disease indication and technical descriptors. Table S2: Infectious-disease datasets: normalized s.d. Per-dataset summary statistics for both editing indices. Table S3: Infectious-disease case–control contrasts. All case-control comparisons used to compute effect-size differences between editing indices. Table S4: Infectious-disease sample annotation. Per-sample metadata for the infectious-disease cohort, detailing original SRA/GEO columns and annotation, curated case–control annotations, study design, read length, replicate status, library type, strand specificity and tissue. Table S5: Failed samples. List of infectious-disease samples excluded from analysis, with the specific failure reason for each. Table S6: Infectious disease taxonomy. Mapping of infectious-disease studies to pathogen type and clinical presentation. Table S7: ADAR perturbation sample classification. Sample classification for ADAR1-perturbation and subcellular fractionation experiments used in Fig. 3. Table S8: Autoimmune diseases classification and editing metrics. Classification and editing statistics for autoimmune diseases datasets. Table S9: GTEx sample classification. Classification of GTEx samples included in the analysis. Table S10: TCGA sample classification. Classification of TCGA samples retained in the study and those removed during filtering. Table S11: Human Alu elements in 3′UTRs. Full list of human inverted and tandem Alu elements > 200 bp, with the 3′UTR regions in which they occur. An element may appear in multiple 3′UTRs of the same or different genes and may be classified as inverted or tandem in different isoforms. Table S12: Mouse B1 and B2 elements in 3′UTRs. Full list of mouse inverted and tandem B1 and B2 elements > 120 bp, with the 3′UTR regions in which they occur. An element may appear in multiple 3′UTRs of the same or different genes and may be classified as inverted or tandem in different isoforms. Table S13: ADAR perturbation datasets editing metrics. Editing statistics for ADAR1-perturbation and fractionation datasets. Table S14: Infectious-disease editing metrics. Editing, coverage and runtime statistics for all infectious-disease samples. Table S15: TCGA editing metrics. Editing statistics for TCGA samples. Table S16: GTEx editing metrics. Editing, coverage and runtime statistics for GTEx samples. Table S17: Raw survival-analysis output. Kaplan-Meier log-rank results for all cancer types and all high/low index dichotomies. Table S18: Contribution of element expression to CEI. Coverage and editing metrics for each expressed CEI element in mock and IFN-beta treated wildtype samples from PRJNA386593. This data was used for calculating the contribution of element expression to CEI.
Acknowledgements
We thank the following members of the Levanon research group for their contributions to initial dataset search and curation: Aviya Zimmermann, David Gorelik, Zohar Rossenwasser-Maymon, Michelle Eidelman, Miriam Karmon, Moshe Perez, Netanel Landesman, Kobi Shapira, Rivka Bouhnik-Cohen, and Renana B. Drummer. We are grateful to Yamin Cohen for compiling the clinical manifestation categories. We acknowledge the numerous research groups, including the GTEx Consortium, who deposited RNA-seq data in public repositories and made this large-scale comparative analysis possible. The results shown here are in whole or part based upon data generated by the TCGA Research Network: https://www.cancer.gov/tcga.
Peer review information
Yang Li and Tim Sands were the primary editors of this article and managed its editorial process and peer review in collaboration with the rest of the editorial team. The peer-review history is available in the online version of this article.
Authors’ contributions
EYL and EE conceived the study and supervised the work. EYL, EE and RCF developed the CEI methodology. IT designed and implemented the cloud platform workflow and conducted the sample analysis. Resource initialization and data analysis were performed by RCF. IT and RCF annotated infectious disease. HK collected and curated infectious disease datasets and contributed to the initial analyses. SHR designed and implemented the known-sites editing quantification program and developed the Docker container. EYL, EE and RCF wrote the manuscript with feedback from IT. All authors read and approved the final manuscript.
Funding
This material is based upon work supported by the Google Cloud Research Credits program with the award GCP19980904. E.Y.L. was supported by the Israel Science Foundation grant 2637/23 and by NIDDK (DK133442). E.E was supported by the Israel Science Foundation grant 867/25.
Data availability
All code, pipeline definitions, and configuration files are released under the open-source GNU General Public License v3.0 (GPL-3.0) and are publicly available in GitHub at https://github.com/a2iEditing/CEI [101] and in Zenodo at https://doi.org/10.5281/zenodo.19158873 [104]. Custom Docker images are archived at https://hub.docker.com/u/levanonlab [105]. Raw RNA-seq data were downloaded from the NCBI Sequence Read Archive (SRA) [88]; accession numbers are provided in Additional file 2: Table S4, Additional file 2: Table S7 and Additional file 2: Table S8. Controlled-access data from TCGA [72] (dbGaP accession phs000178) and GTEx [71] (dbGaP accession phs000424) were obtained under approved data-use agreements; accession numbers are provided in Additional file 2: Table S9 and Additional file 2: Table S10.
All computational steps were executed in version-pinned Docker containers orchestrated by Nextflow [106] (v25.03.1-edge) to ensure end-to-end reproducibility. Statistical analyses were conducted in R v4.5.0 using tidyr v1.3.1, dplyr v1.1.4, purrr v1.0.4, data.table v1.17.6 and stringr v1.5.1. Publication-ready figures were generated in R using ggplot2 v3.5.2, cowplot v1.1.3 and patchwork v1.3.1. Venn diagrams were generated with DeepVenn [107]. Illustrative figures were created with BioRender.com.
Data generated or analyzed during this study are included in this published article and its additional files.
Declarations
Ethics approval and consent to participate
Public RNA-seq data were downloaded from the NCBI Sequence Read Archive (SRA) [88]. Controlled-access data from TCGA [72] (dbGaP accession phs000178) and GTEx [71] (dbGaP accession phs000424) were obtained under approved data-use agreements.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Roni Cohen-Fultheim and Itamar Twersky contributed equally to this work.
Erez Y. Levanon and Eli Eisenberg contributed equally to this work, corresponding authors.
Contributor Information
Erez Y. Levanon, Email: erez.levanon@biu.ac.il
Eli Eisenberg, Email: elieis@post.tau.ac.il.
References
- 1.Bass BL. RNA editing by adenosine deaminases that act on RNA. Annu Rev Biochem. 2002;71:817–46. 10.1146/annurev.biochem.71.110601.135501. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Savva YA, Rieder LE, Reenan RA. The ADAR protein family. Genome Biol. 2012;13:252, 252. 10.1186/gb-2012-13-12-252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Nishikura K. A-to-I editing of coding and non-coding RNAs by ADARs. Nat Rev Mol Cell Biol. 2016;17:83–96. 10.1038/nrm.2015.4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Schlee M, Hartmann G. Discriminating self from non-self in nucleic acid sensing. Nat Rev Immunol Nature Publishing Group. 2016;16:566–80. 10.1038/nri.2016.78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Chen YG, Hur S. Cellular origins of dsRNA, their recognition and consequences. Nat Rev Mol Cell Biol. 2022;23:286–301. 10.1038/s41580-021-00430-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hu S-B, Heraud-Farlow J, Sun T, Liang Z, Goradia A, Taylor S, et al. ADAR1p150 prevents MDA5 and PKR activation via distinct mechanisms to avert fatal autoinflammation. Mol Cell. 2023;83:3869–84. 10.1016/j.molcel.2023.09.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Mannion NM, Greenwood SM, Young R, Cox S, Brindle J, Read D, et al. The RNA-editing enzyme ADAR1 controls innate immune responses to RNA. Cell Rep. 2014;9:1482–94. 10.1016/j.celrep.2014.10.041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Liddicoat BJ, Piskol R, Chalk AM, Ramaswami G, Higuchi M, Hartner JC, et al. RNA editing by ADAR1 prevents MDA5 sensing of endogenous dsRNA as nonself. Science. 2015;349:1115–20. 10.1126/science.aac7049. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Pestal K, Funk CC, Snyder JM, Price ND, Treuting PM, Stetson DB. Isoforms of RNA-editing enzyme ADAR1 independently control nucleic acid sensor MDA5-driven autoimmunity and multi-organ development. Immunity. 2015. 10.1016/j.immuni.2015.11.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Chung H, Calis JJA, Wu X, Sun T, Yu Y, Sarbanes SL, et al. Human ADAR1 Prevents Endogenous RNA from Triggering Translational Shutdown. Cell Cell Press. 2018;172:811-824.e14. 10.1016/j.cell.2017.12.038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Miyamura Y, Suzuki T, Kono M, Inagaki K, Ito S, Suzuki N, et al. Mutations of the RNA-specific adenosine deaminase gene (DSRAD) are involved in dyschromatosis symmetrica hereditaria. Am J Hum Genet. 2003;73:693–9. 10.1086/378209. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Nejentsev S, Walker N, Riches D, Egholm M, Todd JA. Rare variants of IFIH1, a gene implicated in antiviral responses, protect against type 1 diabetes. Science Science. 2009;324:387–9. 10.1126/science.1167728. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Rice GI, Kasher PR, Forte GMA, Mannion NM, Greenwood SM, Szynkiewicz M, et al. Mutations in ADAR1 cause Aicardi-Goutières syndrome associated with a type I interferon signature. Nat Genet. 2012;44:1243–8. 10.1038/ng.2414. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Shallev L, Kopel E, Feiglin A, Leichner GS, Avni D, Sidi Y, et al. Decreased A-to-I RNA editing as a source of keratinocytes’ dsRNA in psoriasis. RNA. 2018;24:828–40. 10.1261/rna.064659.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Tossberg JT, Heinrich RM, Farley VM, Crooke PS III, Aune TM. Adenosine-to-inosine RNA editing of Alu double-stranded (ds)RNAs is markedly decreased in multiple sclerosis and unedited Alu dsRNAs are potent activators of proinflammatory transcriptional responses. J Immunol. 2020;205:2606–17. 10.4049/jimmunol.2000384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Roth SH, Danan-Gotthold M, Ben-Izhak M, Rechavi G, Cohen CJ, Louzoun Y, et al. Increased RNA editing may provide a source for autoantigens in systemic lupus erythematosus. Cell Rep. 2018;23:50–7. 10.1016/j.celrep.2018.03.036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Vlachogiannis NI, Gatsiou A, Silvestris DA, Stamatelopoulos K, Tektonidou MG, Gallo A, et al. Increased adenosine-to-inosine RNA editing in rheumatoid arthritis. J Autoimmun. 2020;106:102329, 102329. 10.1016/j.jaut.2019.102329. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Song B, Shiromoto Y, Minakuchi M, Nishikura K. The role of RNA editing enzyme ADAR1 in human disease. WIREs RNA. 2022;13:e1665. 10.1002/wrna.1665. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Li Q, Gloudemans MJ, Geisinger JM, Fan B, Aguet F, Sun T, et al. RNA editing underlies genetic risk of common inflammatory diseases. Nature. 2022;608:569–77. 10.1038/s41586-022-05052-x. [DOI] [PMC free article] [PubMed]
- 20.Karmon M, Kopel E, Barzilai A, Geva P, Eisenberg E, Levanon EY, et al. Altered RNA editing in atopic dermatitis highlights the role of double-stranded RNA for immune surveillance. J Invest Dermatol. 2023;143:933–43. 10.1016/j.jid.2022.11.010. [DOI] [PubMed] [Google Scholar]
- 21.Knebel UE, Peleg S, Dai C, Cohen-Fultheim R, Jonsson S, Poznyak K, et al. Disrupted RNA editing in beta cells mimics early-stage type 1 diabetes. Cell Metab. 2024;36:48–61. 10.1016/j.cmet.2023.11.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Tamizkar KH, Jantsch MF. RNA editing in disease: mechanisms and therapeutic potential. RNA Cold Spring Harbor Lab. 2025;31:359–68. 10.1261/rna.080331.124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Wang QD, Miyakoda M, Yang WD, Khillan J, Stachura DL, Weiss MJ, et al. Stress-induced apoptosis associated with null mutation of ADAR1 RNA editing deaminase gene. J Biol Chem. 2004;279:4952–61. 10.1074/jbc.M310162200. [DOI] [PubMed] [Google Scholar]
- 24.Hartner JC, Schmittwolf C, Kispert A, Muller AM, Higuchi M, Seeburg PH. Liver disintegration in the mouse embryo caused by deficiency in the RNA-editing enzyme ADAR1. J Biol Chem. 2004;279:4894–902. 10.1074/jbc.M311347200. [DOI] [PubMed] [Google Scholar]
- 25.Li JB, Walkley CR. Leveraging genetics to understand ADAR1-mediated RNA editing in health and disease. Nat Rev Genet. 2025;26:532–46. 10.1038/s41576-025-00830-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Athanasiadis A, Rich A, Maas S. Widespread A-to-I RNA editing of Alu-containing mRNAs in the human transcriptome. PLoS Biol. 2004;2:e391. 10.1371/journal.pbio.0020391. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Blow M, Futreal PA, Wooster R, Stratton MR. A survey of RNA editing in human brain. Genome Res. 2004;14:2379–87. 10.1101/gr.2951204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Kim DDY, Kim TTY, Walsh T, Kobayashi Y, Matise TC, Buyske S, et al. Widespread RNA editing of embedded Alu elements in the human transcriptome. Genome Res. 2004;14:1719–25. 10.1101/gr.2855504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Levanon EY, Eisenberg E, Yelin R, Nemzer S, Hallegger M, Shemesh R, et al. Systematic identification of abundant A-to-I editing sites in the human transcriptome. Nat Biotechnol. 2004;22:1001–5. 10.1038/nbt996. [DOI] [PubMed] [Google Scholar]
- 30.Kleinberger Y, Eisenberg E. Large-scale analysis of structural, sequence and thermodynamic characteristics of A-to-I RNA editing sites in human Alu repeats. BMC Genomics. 2010;11:453, 453. 10.1186/1471-2164-11-453. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Bazak L, Haviv A, Barak M, Jacob-hirsch J, Deng P, Zhang R, et al. A-to-I RNA editing occurs at over a hundred million genomic sites, located in a majority of human genes. Genome Res. 2014;24:365–76. 10.1101/gr.164749.113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Porath HT, Carmi S, Levanon EY. A genome-wide map of hyper-edited RNA reveals numerous new sites. Nat Commun. 2014;5:4726, 4726. 10.1038/ncomms5726. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Lander ES, Linton LM, Birren B, Nusbaum C, Zody MC, Baldwin J, et al. Initial sequencing and analysis of the human genome. Nature. 2001;409:860–921. 10.1038/35057062. [DOI] [PubMed] [Google Scholar]
- 34.Batzer MA, Deininger PL. Alu repeats and human genomic diversity. Nat Rev Genet Nature Publishing Group. 2002;3:370–9. 10.1038/nrg798. [DOI] [PubMed] [Google Scholar]
- 35.Ahmad S, Mu X, Yang F, Greenwald E, Park JW, Jacob E, et al. Breaching Self-Tolerance to Alu Duplex RNA Underlies MDA5-Mediated Inflammation. Cell Cell Press. 2018;172:797-810.e13. 10.1016/J.CELL.2017.12.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Sela N, Mersch B, Gal-Mark N, Lev-Maor G, Hotz-Wagenblatt A, Ast G. Comparative analysis of transposed element insertion within human and mouse genomes reveals Alu’s unique role in shaping the human transcriptome. Genome Biol. 2007;8:R127. 10.1186/gb-2007-8-6-r127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Piechotta M, Wyler E, Ohler U, Landthaler M, Dieterich C. JACUSA: site-specific identification of RNA editing events from replicate sequencing data. BMC Bioinformatics. 2017;18:7. 10.1186/s12859-016-1432-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Eisenberg E, Levanon EY. A-to-I RNA editing - immune protector and transcriptome diversifier. Nat Rev Genet. 2018;19:473–90. 10.1038/s41576-018-0006-1. [DOI] [PubMed] [Google Scholar]
- 39.Roth SH, Levanon EY, Eisenberg E. Genome-wide quantification of ADAR adenosine-to-inosine RNA editing activity. Nat Methods. 2019;16:1131–8. 10.1038/s41592-019-0610-9. [DOI] [PubMed] [Google Scholar]
- 40.Flati T, Gioiosa S, Spallanzani N, Tagliaferri I, Diroma MA, Pesole G, et al. HPC-REDItools: a novel HPC-aware tool for improved large scale RNA-editing analysis. BMC Bioinformatics. 2020;21:353. 10.1186/s12859-020-03562-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Lo Giudice C, Silvestris DA, Roth SH, Eisenberg E, Pesole G, Gallo A, et al. Quantifying RNA editing in deep transcriptome datasets. Front Genet. 2020. 10.3389/fgene.2020.00194. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Kluesner MG, Tasakis RN, Lerner T, Arnold A, Wüst S, Binder M, et al. MultiEditR: The first tool for the detection and quantification of RNA editing from Sanger sequencing demonstrates comparable fidelity to RNA-seq. Molecular Therapy Nucleic Acids Elsevier. 2021;25:515–23. 10.1016/j.omtn.2021.07.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Rosenwasser Z, Levanon E, Levitt M, Oren G. Detection of RNA Editing Sites by GPT Fine-tuning. 2024 Cited 2025 Sept 30. https://openreview.net/forum?id=Ty0aKMV7Ow. Accessed 30 Sept 2025
- 44.Torkler P, Sauer M, Schwartz U, Corbacioglu S, Sommer G, Heise T. Lodei: a robust and sensitive tool to detect transcriptome-wide differential A-to-I editing in RNA-seq data. Nat Commun. 2024;15:9121–024. 10.1038/s41467-024-53298-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Eisenberg E. Chapter eleven - bioinformatic approaches for accurate assessment of A-to-I editing in complete transcriptomes. In: Beal P, editor. Methods in Enzymology. Academic Press; 2025. p. 241–65. 10.1016/bs.mie.2024.11.020. [DOI] [PubMed]
- 46.Han L, Diao L, Yu S, Xu X, Li J, Zhang R, et al. The genomic landscape and clinical relevance of A-to-I RNA editing in human cancers. Cancer Cell. 2015;28:515–28. 10.1016/j.ccell.2015.08.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Tan MH, Li Q, Shanmugam R, Piskol R, Kohler J, Young AN, et al. Dynamic landscape and regulation of RNA editing in mammals. Nature. 2017;550:249–54. 10.1038/nature24041 [DOI] [PMC free article] [PubMed]
- 48.Breen MS, Dobbyn A, Li Q, Roussos P, Hoffman GE, Stahl E, et al. Global landscape and genetic regulation of RNA editing in cortical samples from individuals with schizophrenia. Nat Neurosci. 2019;22:1402–12. 10.1038/s41593-019-0463-7. [DOI] [PMC free article] [PubMed]
- 49.Ishizuka JJ, Manguso RT, Cheruiyot CK, Bi K, Panda A, Iracheta-Vellve A, et al. Loss of ADAR1 in tumours overcomes resistance to immune checkpoint blockade. Nature. 2019;565:43–8. 10.1038/s41586-018-0768-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Schaffer AA, Kopel E, Hendel A, Picardi E, Levanon EY, Eisenberg E. The cell line A-to-I RNA editing catalogue. Nucleic acids research NLM (Medline). 2020;48:5849–58. 10.1093/nar/gkaa305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Giorgio SD, Martignano F, Torcia MG, Mattiuz G, Conticello SG. Evidence for host-dependent RNA editing in the transcriptome of SARS-CoV-2. Science Advances. American Association for the Advancement of Science; 2020 Cited 2025 Sept 30; 10.1126/sciadv.abb5813 [DOI] [PMC free article] [PubMed]
- 52.Buchumenski I, Holler K, Appelbaum L, Eisenberg E, Junker JP, Levanon EY. Systematic identification of A-to-I RNA editing in zebrafish development and adult organs. Nucleic Acids Res. 2021;49:4325–37. 10.1093/nar/gkab247. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Szymczak F, Cohen-Fultheim R, Thomaidou S, de Brachène AC, Castela A, Colli M, et al. ADAR1-dependent editing regulates human β cell transcriptome diversity during inflammation. Front Endocrino. Frontiers; 2022 Cited 2025 May 22;13. 10.3389/fendo.2022.1058345 [DOI] [PMC free article] [PubMed]
- 54.Picardi E, Mansi L, Pesole G. Detection of A-to-I RNA editing in SARS-COV-2. Genes (Basel). 2022;13:41, 41. 10.3390/genes13010041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Cuddleston WH, Li J, Fan X, Kozenkov A, Lalli M, Khalique S, et al. Cellular and genetic drivers of RNA editing variation in the human brain. Nat Commun. Nature Publishing Group; 2022;13:2997. 10.1038/s41467-022-30531-0 [DOI] [PMC free article] [PubMed]
- 56.Amweg A, Tusup M, Cheng P, Picardi E, Dummer R, Levesque MP, et al. The A to I editing landscape in melanoma and its relation to clinical outcome. RNA Biol. 2022. 10.1080/15476286.2022.2110390. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Merdler-Rabinowicz R, Gorelik D, Park J, Meydan C, Foox J, Karmon M, et al. Elevated A-to-I RNA editing in COVID-19 infected individuals. NAR Genom Bioinform. 2023. 10.1093/NARGAB/LQAD092. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Avram-Shperling A, Kopel E, Twersky I, Gabay O, Ben-David A, Karako-Lampert S, et al. Identification of exceptionally potent adenosine deaminases RNA editors from high body temperature organisms. PLoS Genet. 2023;19:e1010661. 10.1371/journal.pgen.1010661. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Mann TD, Kopel E, Eisenberg E, Levanon EY. Increased A-to-I RNA editing in atherosclerosis and cardiomyopathies. PLoS Comput Biol. 2023;19:e1010923. 10.1371/journal.pcbi.1010923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Rodriguez de los Santos M, Kopell BH, Buxbaum Grice A, Ganesh G, Yang A, Amini P, et al. Divergent landscapes of A-to-I editing in postmortem and living human brain. Nat Commun. 2024;15:5366, 5366. 10.1038/s41467-024-49268-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Wang Y, Wu J, Zhao J, Xu T, Zhang M, Liu J, et al. Global characterization of RNA editing in genetic regulation of multiple ovarian cancer subtypes. Mol Ther Nucleic Acids. 2024. 10.1016/j.omtn.2024.102127. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Heraud-Farlow JE, Taylor SR, Chalk AM, Escudero A, Hu S-B, Goradia A, et al. GGNBP2 regulates MDA5 sensing triggered by self double-stranded RNA following loss of ADAR1 editing. Sci Immunol. 2024;9:eadk0412. 10.1126/sciimmunol.adk0412. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Peleg S, Zamashanski L, Belin J, Novoselsky R, Cohen-Fultheim R, Knebel UE, et al. RNA editing deficiency models differential immunogenicity of pancreatic α- and β-cells. Mol Metab. 2025;98:102183, 102183. 10.1016/j.molmet.2025.102183. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.D’Addabbo P, Cohen-Fultheim R, Twersky I, Fonzino A, Silvestris DA, Prakash A, et al. REDIportal: toward an integrated view of the A-to-I editing. Nucleic Acids Res. 2025;53:D233-42. 10.1093/nar/gkae1083. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Dong S, Li C, Zenklusen D, Singer RH, Jacobson A, He F. YRA1 autoregulation requires nuclear export and cytoplasmic Edc3p-mediated degradation of its pre-mRNA. Mol Cell. 2007;25:559–73. 10.1016/j.molcel.2007.01.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Barak M, Porath HT, Finkelstein G, Knisbacher BA, Buchumenski I, Roth SH, et al. Purifying selection of long dsRNA is the first line of defense against false activation of innate immunity. Genome Biol. 2020;21:26. 10.1186/s13059-020-1937-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Bazak L, Levanon EY, Eisenberg E. Genome-wide analysis of Alu editability. Nucleic Acids Res. 2014;42:6876–84. 10.1093/nar/gku414. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Sun T, Li Q, Geisinger JM, Hu S-B, Fan B, Su S, et al. ADAR1 editing is necessary for only a small subset of cytosolic dsRNAs to evade MDA5-mediated autoimmunity. Nat Genet Nature Publishing Group. 2025;57:3101–11. 10.1038/s41588-025-02430-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Levanon EY, Cohen-Fultheim R, Eisenberg E. In search of critical dsRNA targets of ADAR1. Trends Genet. 2024;40:250–9. 10.1016/j.tig.2023.12.002. [DOI] [PubMed] [Google Scholar]
- 70.Rosenwasser Z, Cohen-Fultheim R, Shliefer O, Levanon EY, Eisenberg E. Transcriptome-wide systematic search does not detect A-to-I RNA editing in cis-antisense RNA duplexes. Genome Res. Cold Spring Harbor Lab; 2026;gr.280820.125. 10.1101/gr.280820.125 [DOI] [PMC free article] [PubMed]
- 71.Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, et al. The Genotype-Tissue Expression (GTEx) project. Nat Genet. 2013;45:580–5. 10.1038/ng.2653. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Weinstein JN, Collisson EA, Mills GB, Shaw KRM, Ozenberger BA, Ellrott K, et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet Nature Publishing Group. 2013;45:1113–20. 10.1038/ng.2764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Patterson JB, Samuel CE. Expression and regulation by interferon of a double-stranded-RNA-specific adenosine deaminase from human cells: evidence for two forms of the deaminase. Mol Cell Biol. 1995;15:5376–88. 10.1128/MCB.15.10.5376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.van Gemert F, Drakaki A, Lozano IM, de Groot D, Uiterkamp MS, Proost N, et al. ADARp150 counteracts whole genome duplication. Nucleic Acids Res. 2024;52:10370–84. 10.1093/nar/gkae700. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Colli ML, Ramos-Rodríguez M, Nakayasu ES, Alvelos MI, Lopes M, Hill JLE, et al. An integrated multi-omics approach identifies the landscape of interferon-α-mediated responses of human pancreatic beta cells. Nat Commun. 2020;11:2584. 10.1038/s41467-020-16327-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Ramos-Rodríguez M, Raurell-Vila H, Colli ML, Alvelos MI, Subirana-Granés M, Juan-Mateu J, et al. The impact of proinflammatory cytokines on the β-cell regulatory landscape provides insights into the genetics of type 1 diabetes. Nat Genet Nature Publishing Group. 2019;51:1588–95. 10.1038/s41588-019-0524-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Gonzalez-Duque S, Azoury ME, Colli ML, Afonso G, Turatsinze J-V, Nigi L, et al. Conventional and Neo-antigenic Peptides Presented by β Cells Are Targeted by Circulating Naïve CD8+ T Cells in Type 1 Diabetic and Healthy Donors. Cell Metabolism Elsevier. 2018;28:946-960.e6. 10.1016/j.cmet.2018.07.007. [DOI] [PubMed] [Google Scholar]
- 78.Kim JI, Nakahama T, Yamasaki R, Cruz PHC, Vongpipatana T, Inoue M, et al. RNA editing at a limited number of sites is sufficient to prevent MDA5 activation in the mouse brain. PLoS Genet. 2021;17:e1009516. 10.1371/journal.pgen.1009516. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Kleinova R, Rajendra V, Leuchtenberger AF, Lo Giudice C, Vesely C, Kapoor U, et al. The ADAR1 editome reveals drivers of editing-specificity for ADAR1-isoforms. Nucleic Acids Res. 2023;51:4191–207. 10.1093/nar/gkad265. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Zaghlool A, Niazi A, Björklund ÅK, Westholm JO, Ameur A, Feuk L. Characterization of the nuclear and cytosolic transcriptomes in human brain tissue reveals new insights into the subcellular distribution of RNA transcripts. Sci Rep. 2021;11:4076. 10.1038/s41598-021-83541-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Thompson EG, Du Y, Malherbe ST, Shankar S, Braun J, Valvo J, et al. Host blood RNA signatures predict the outcome of tuberculosis treatment. Tuberculosis. 2017;107:48–58. 10.1016/j.tube.2017.08.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Garrido-Trigo A, Corraliza AM, Veny M, Dotti I, Melón-Ardanaz E, Rill A, et al. Macrophage and neutrophil heterogeneity at single-cell spatial resolution in human inflammatory bowel disease. Nat Commun. 2023;14:4506, 4506. 10.1038/s41467-023-40156-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Tsoi LC, Rodriguez E, Degenhardt F, Baurecht H, Wehkamp U, Volks N, et al. Atopic dermatitis is an IL-13–dominant disease with greater molecular heterogeneity compared to psoriasis. J Invest Dermatol. 2019;139:1480–9. 10.1016/j.jid.2018.12.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Merleev A, Ji-Xu A, Toussi A, Tsoi LC, Le ST, Luxardi G, et al. Proprotein convertase subtilisin/kexin type 9 is a psoriasis-susceptibility locus that is negatively related to IL36G. JCI Insight. American Society for Clinical Investigation; 2022 Cited 2026 Mar 20;7. 10.1172/jci.insight.141193 [DOI] [PMC free article] [PubMed]
- 85.Small J, Kamra Y, Herbert T, Tate P, Matthews E, McGarvey A, Wang J, Albornoz Carranza C, Salvucci M, Watcharapichat P, Gurel M, Barcellos Machado C. Mapping the Molecular Signature of Progressive Multiple Sclerosis. GSE290450. Gene Expression Omnibus. 2025 https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE290450
- 86.Xing S. Salivary Extracellular Vesicle RNA Profiling Reveals Biomarkers for Sjögren’s. GSE305188. 2025 https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE305188
- 87.Chen C-X, Cho D-S, Wang Q, Lai F, Carter KC, Nishikura K. A third member of the RNA-specific adenosine deaminase gene family, ADAR3, contains both single- and double-stranded RNA binding domains. RNA. 2000;6:755–67. 10.1017/S1355838200000170. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Leinonen R, Sugawara H, Shumway M. The sequence read archive. Nucleic Acids Res. 2011;39:D19-21. 10.1093/nar/gkq1019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Barrett T, Wilhite SE, Ledoux P, Evangelista C, Kim IF, Tomashevsky M, et al. NCBI GEO: archive for functional genomics data sets—update. Nucleic Acids Res. 2013;41:D991-5. 10.1093/nar/gks1193. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Mailman MD, Feolo M, Jin Y, Kimura M, Tryka K, Bagoutdinov R, et al. The NCBI dbGaP database of genotypes and phenotypes. Nat Genet. 2007;39:1181–6. 10.1038/ng1007-1181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Heath AP, Ferretti V, Agrawal S, An M, Angelakos JC, Arya R, et al. The NCI Genomic Data Commons. Nat Genet Nature Publishing Group. 2021;53:257–62. 10.1038/s41588-021-00791-5. [DOI] [PubMed] [Google Scholar]
- 92.Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, et al. The human genome browser at UCSC. Genome Res. 2002;12:996–1006. 10.1101/gr.229102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Quinlan AR, Hall IM. BEDTools: A flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–2. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods Nature Publishing Group. 2017;14:417–9. 10.1038/nmeth.4197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018;34:i884–90. 10.1093/bioinformatics/bty560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Sherry ST, Ward M-H, Kholodov M, Baker J, Phan L, Smigielski EM, et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001;29:308–11. 10.1093/nar/29.1.308. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Gabay O, Shoshan Y, Kopel E, Ben-Zvi U, Mann TD, Bressler N, et al. Landscape of adenosine-to-inosine RNA recoding across human tissues. Nat Commun. 2022;13:1–17. 10.1038/s41467-022-28841-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Pinto Y, Cohen HY, Levanon EY. Mammalian conserved ADAR targets comprise only a small fragment of the human editosome. Genome Biol. 2014;15:R5. 10.1186/gb-2014-15-1-r5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The sequence alignment/map format and SAMtools. Bioinformatics. 2009;25:2078–9. 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Twersky I, Cohen-Fultheim R. CEI. GitHub. 2026 https://github.com/a2iEditing/CEI
- 102.Dale RK, Pedersen BS, Quinlan AR. Pybedtools: a flexible Python library for manipulating genomic datasets and annotations. Bioinformatics. 2011;27:3423–4. 10.1093/bioinformatics/btr539. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Kaplan EL, Meier P. Nonparametric estimation from incomplete observations. J Am Stat Assoc. 1958;53:457–81. 10.2307/2281868. [Google Scholar]
- 104.Twersky I, Cohen-Fultheim R. Code for A Cytoplasmic Index for Quantifying Immune-Related A-to-I RNA Editing. Zenodo. 2026 10.5281/zenodo.19158873 [DOI] [PMC free article] [PubMed]
- 105.Twersky I. levanonlab. Docker Hub. 2026 https://hub.docker.com/u/levanonlab
- 106.Di Tommaso P, Chatzou M, Floden EW, Barja PP, Palumbo E, Notredame C. Nextflow enables reproducible computational workflows. Nat Biotechnol. 2017;35:316–9. 10.1038/nbt.3820. [DOI] [PubMed] [Google Scholar]
- 107.Hulsen T. DeepVenn -- a web application for the creation of area-proportional Venn diagrams using the deep learning framework Tensorflow.js. arXiv; 2022 Cited 2025 Sept 17. 10.48550/arXiv.2210.04597
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Additional file 1: Supplementary figures. This file includes supplementary figures S1-S11.
Additional file 2: Supplementary tables. This file includes supplementary tables S1-S18. Supplementary tables description: Table S1: Infectious-disease dataset annotation. Per-dataset metadata for all infectious-disease studies analysed, including study design, pathogen, disease indication and technical descriptors. Table S2: Infectious-disease datasets: normalized s.d. Per-dataset summary statistics for both editing indices. Table S3: Infectious-disease case–control contrasts. All case-control comparisons used to compute effect-size differences between editing indices. Table S4: Infectious-disease sample annotation. Per-sample metadata for the infectious-disease cohort, detailing original SRA/GEO columns and annotation, curated case–control annotations, study design, read length, replicate status, library type, strand specificity and tissue. Table S5: Failed samples. List of infectious-disease samples excluded from analysis, with the specific failure reason for each. Table S6: Infectious disease taxonomy. Mapping of infectious-disease studies to pathogen type and clinical presentation. Table S7: ADAR perturbation sample classification. Sample classification for ADAR1-perturbation and subcellular fractionation experiments used in Fig. 3. Table S8: Autoimmune diseases classification and editing metrics. Classification and editing statistics for autoimmune diseases datasets. Table S9: GTEx sample classification. Classification of GTEx samples included in the analysis. Table S10: TCGA sample classification. Classification of TCGA samples retained in the study and those removed during filtering. Table S11: Human Alu elements in 3′UTRs. Full list of human inverted and tandem Alu elements > 200 bp, with the 3′UTR regions in which they occur. An element may appear in multiple 3′UTRs of the same or different genes and may be classified as inverted or tandem in different isoforms. Table S12: Mouse B1 and B2 elements in 3′UTRs. Full list of mouse inverted and tandem B1 and B2 elements > 120 bp, with the 3′UTR regions in which they occur. An element may appear in multiple 3′UTRs of the same or different genes and may be classified as inverted or tandem in different isoforms. Table S13: ADAR perturbation datasets editing metrics. Editing statistics for ADAR1-perturbation and fractionation datasets. Table S14: Infectious-disease editing metrics. Editing, coverage and runtime statistics for all infectious-disease samples. Table S15: TCGA editing metrics. Editing statistics for TCGA samples. Table S16: GTEx editing metrics. Editing, coverage and runtime statistics for GTEx samples. Table S17: Raw survival-analysis output. Kaplan-Meier log-rank results for all cancer types and all high/low index dichotomies. Table S18: Contribution of element expression to CEI. Coverage and editing metrics for each expressed CEI element in mock and IFN-beta treated wildtype samples from PRJNA386593. This data was used for calculating the contribution of element expression to CEI.
Data Availability Statement
All code, pipeline definitions, and configuration files are released under the open-source GNU General Public License v3.0 (GPL-3.0) and are publicly available in GitHub at https://github.com/a2iEditing/CEI [101] and in Zenodo at https://doi.org/10.5281/zenodo.19158873 [104]. Custom Docker images are archived at https://hub.docker.com/u/levanonlab [105]. Raw RNA-seq data were downloaded from the NCBI Sequence Read Archive (SRA) [88]; accession numbers are provided in Additional file 2: Table S4, Additional file 2: Table S7 and Additional file 2: Table S8. Controlled-access data from TCGA [72] (dbGaP accession phs000178) and GTEx [71] (dbGaP accession phs000424) were obtained under approved data-use agreements; accession numbers are provided in Additional file 2: Table S9 and Additional file 2: Table S10.
All computational steps were executed in version-pinned Docker containers orchestrated by Nextflow [106] (v25.03.1-edge) to ensure end-to-end reproducibility. Statistical analyses were conducted in R v4.5.0 using tidyr v1.3.1, dplyr v1.1.4, purrr v1.0.4, data.table v1.17.6 and stringr v1.5.1. Publication-ready figures were generated in R using ggplot2 v3.5.2, cowplot v1.1.3 and patchwork v1.3.1. Venn diagrams were generated with DeepVenn [107]. Illustrative figures were created with BioRender.com.
Data generated or analyzed during this study are included in this published article and its additional files.
