Abstract
We report that in humans, mice, fruit flies, and worms, the ribosomal RNAs and the transcribed spacers of 45S are densely packed with organism-specific sequence motifs that are primarily shared with nervous system genes. The human ribosomal RNAs and 45S spacers contain 1,723 such motifs. Specific combinations of these motifs are predominantly found in 3,430 human nervous system genes, of which 1,046 are genes associated with brain disorders, including autism spectrum disorder and schizophrenia. The sequences of the 1,723 motifs and their locations in the introns and exons of nervous system genes are unique to primates. Experimental evidence indicates that the motifs are contact points for ribosomal-RNA|messenger-RNA and messenger-RNA|messenger-RNA heteroduplexes, present in the binding sites of many RNA-binding proteins, and carried by endogenous small noncoding RNAs. Moreover, several of the motifs’ intergenic copies overlap genome-wide association-study–determined polymorphisms associated with brain disorders and other conditions. Lastly, based on expression data from bulk brain samples from autism spectrum disorder patients and controls, specific combinations of these motifs are enriched only among genes that are downregulated in the patients and only in the cortical areas that are responsible for language, hearing, and vision. This study shows that genomic architecture, the sequences of ribosomal RNAs/spacers, and the sequences of nervous system genes have remained interlocked across 600 million years of evolution through organism-specific motifs. The findings suggest an extensive regulatory layer and can aid in developing new molecular diagnostics and treatments for disorders considered typically human.
Keywords: autism spectrum disorder, bipolar disorder, schizophrenia, epilepsy, Hirschsprung's disease, system development, neurons, axons, dendrites, synapses, ribosomal RNAs, primates, rodents, insects, worms, LINE-1, LINE-2, pyknons
Introduction
Linking nucleotide combinations to organismal identity goes back more than 60 years. One of the earliest observations is in a 1962 study by Noboru Sueoka, who noted that a bacterium's C + G content is sufficiently unique and can help distinguish it from other bacteria (Sueoka 1962). Fifteen years later, Carl Woese discovered the archaeal domain of life (Woese and Fox 1977) by examining combinations of oligonucleotides (“k-mers”). His discovery was soon followed by other pioneering studies that showed genomes exhibiting characteristic dimer and trimer frequencies, which could be used to compare organisms (Sharp and Li 1987; Karlin and Burge 1995; Karlin and Mrazek 1996). Later studies explored larger values of k (Sandberg et al. 2001; Tsirigos and Rigoutsos 2005a, b; Takahashi et al. 2009), including some that used templates spanning dozens of nucleotides (nts) (Déraspe et al. 2017; Bussi et al. 2021).
Several years ago, we embarked on studying a specific type of eukaryotic k-mers. These k-mers have numerous exact intergenic and intronic copies and at least one exonic copy (Rigoutsos et al. 2006). We required that the k-mers have at least one instance in a messenger RNA (mRNA) to ensure a sequence-based association between one or more protein-coding and one or more noncoding regions. We dubbed these k-mers “pyknons” (pronounced “peek-nons”), from the Greek word for “dense, serried,” to capture the fact that in many mRNAs, pyknons are located close to one another (Rigoutsos et al. 2006). The original study focused on pyknons in mRNAs and found that human pyknons have largely organism-specific sequences and are associated with specific biological processes. Follow-up studies found that the pyknons’ intronic copies capture functional conservation between human and mouse without sequence conservation (Tsirigos and Rigoutsos 2008) and that they co-localize with regions that produce piRNAs (Tsirigos and Rigoutsos 2008; Robine et al. 2009). A parallel genome-level analysis of the plant A. thaliana showed that its pyknons exhibit properties similar to their human counterparts (Feng et al. 2009).
Using a custom-designed expression array, we also explored the transcriptional profiles of several hundred pyknons with instances in loss-of-heterozygosity regions. We found many pyknon-containing RNAs with tissue-specific expression (Rigoutsos et al. 2017). Other pyknon-containing RNAs were found to have aberrant expression in cutaneous melanoma (Khaitan et al. 2011) and colon cancer (Evangelista et al. 2020; Pichler et al. 2020) and to be downregulated in peripheral blood leucocytes after splenectomy (Dragomir et al. 2019).
For a pyknon-containing lncRNA, N-BLR, we showed that it facilitates the epithelial to mesenchymal transition (EMT) in colon cancer (Rigoutsos et al. 2017) by decoying a miRNA that targets a pyknon in N-BLR's 3′ region, thereby diminishing the miRNA's effect on its protein-coding targets. N-BLR was also shown to facilitate the EMT in gastric adenocarcinoma (Youn et al. 2019). A second pyknon-containing lncRNA, FLANC, was shown to be associated with poor survival in colorectal cancer (Pichler et al. 2020). A recent large-scale study of multiple cancers extended these findings by showing that many lncRNAs strongly associated with oncogenes and tumor suppressors are highly enriched in pyknons (Jha et al. 2022).
Lastly, we mention 2 results that implicated pyknons in somewhat unexpected settings. A study of DNA methylation showed that the DNA methyltransferase DNMT1 also binds RNA with high affinity, at sites enriched for pyknons (Di Ruscio et al. 2013). Additionally, an analysis of expression data from human and mouse pre-implantation embryos found that the mRNAs whose levels decrease from the zygote to the blastocyst contain different pyknons than those whose levels increase during the same period (Telonis and Rigoutsos 2021).
The release of the first telomere-to-telomere (T2T) human genome assembly (Nurk et al. 2022) added 200 million nts that were not previously available. A prominent feature of the new assembly was its inclusion of 219 copies of the 45S cassette. We asked whether these newly added copies could reveal novel linkages between noncoding and protein-coding regions. Notably, these sequences are unlike those of other transcripts: this is true of rRNAs (Cherlin et al. 2020; Kim et al. 2024) and of transcribed spacers of 45S (Bourbon et al. 1988; Coleman 2013), whose translocation (Enright et al. 1996; Bao et al. 2024) and processing (Gordon et al. 2019; Pillon et al. 2019; Sloan et al. 2019; Ma et al. 2024) are complex steps of ribosome biogenesis. Below, we present our investigation's findings and counterpart analyses of the mouse, fruit fly, and worm genomes.
Materials and Methods
Pyknon Generation: Initial Enumeration of k-mers
We exhaustively enumerated 16-mers that are in at least one of the 96,441 mRNAs (22,511 protein-coding genes) of Rel. 105 of ENSEMBL. We then sought them in the T2T genome (Nurk et al. 2022), keeping only those with ≥30 identical genomic copies. The 16-mers that survived this step formed a list of “candidate pyknons.” This step did not include the Y chromosome of the T2T genome to avoid biasing in favor of male-specific k-mers.
Pyknon Generation: Statistical Significance of the Enumerated k-mers
Using a Monte Carlo simulation, we determined the distribution of 16-mers in randomly shuffled instances of the T2T genome assembly. Specifically, we shuffled the forward strand, then generated the reverse complement of it, and enumerated all distinct 16-mers in the resulting 2-stranded “genome.” We repeated the process 100 times, each time creating a histogram of the frequencies of the 16-mers in the respective shuffled genome. From these histograms, we estimated the probability of a 16-mer occurring by chance 30 or more times to be less than 3.0E-06.
Pyknon Generation: Identifying Pyknons Among the Enumerated k-mers
Using the approach outlined in our original study (Rigoutsos et al. 2006), we processed the list of candidate pyknons (see above) to remove redundant entries among them and identify the final list (2 candidate pyknons are redundant if they have the same genomic counts and overlap/co-occur at the same genomic locations). First, we sorted all candidate pyknons in order of decreasing number of their genomic copies, sorting lexicographically candidate pyknons with the same count. Then, we marked all nt positions in all ENSEMBL mRNAs as “available.” In order of decreasing genomic counts, we sought copies of each candidate pyknon in the mRNA sequences. If the candidate is found in an mRNA and all positions of the match are “available,” we let the candidate “claim” these positions and change the labels to “claimed.” If we could place the candidate pyknon in the current mRNA more than once, we would do so before examining the next mRNA. A candidate that successfully claimed one or more spots in one or more mRNAs during this traversal was reported as a “pyknon”; otherwise, it was discarded.
Reference rRNAs
For this analysis, we used the following human genome reference sequences from GenBank: RNA45SN1 (accession number NR_145819.1) as the reference entry for 45S; RNA5S12 (accession number NR_023374.1) for 5S; MT-RNR1 (accession number NR_137294.1) for 12S; and MT-RNR2 (accession number NR_137295.1) for 16S. These reference sequences have the following lengths: RNA45SN1, 13,351 nts; RNA5S12, 121 nts; MT-RNR1, 954 nts; and MT-RNR2, 1,559 nts, respectively. Within 45S, we distinguished among the 5′ ETS or ETS1, 18S rRNA, ITS1, 5.8S rRNA, ITS2, 28S rRNA, and 3′ ETS or ETS2. We also used the following 19 nonhuman 28S rRNA reference sequences: NR_000055.1 (C. elegans), NR_003279.1 (M. musculus), NR_046246.2 (R. norvegicus), NR_133562.1 (D. melanogaster), XR_002740182.2 (F. catus), XR_002805825.1 (E. caballus), XR_003510171.1 (B. indicus × B. taurus), XR_003725683.1 (M. mulatta), XR_003887306.1 (T. rubripes), XR_004030810.1 (N. leucogenys), XR_004223802.1 (X. tropicalis), XR_004524865.1 (T. truncatus), XR_005382296.1 (C. lupus familiaris), XR_005582245.1 (S. boliviensis boliviensis), XR_005647714.1 (P. giganteus), XR_006938338.1 (G. gallus), XR_007912736.1 (O. cuniculus), XR_008546126.1 (P. troglodytes), and XR_008676494.1 (G. gorilla). Lastly, we used the 45S segment from the following 5 nonhuman precursor 48S rRNA reference sequences: KX061886.1 (P. troglodytes), KX061888.1 (P. abelii), KX061890.1 (M. mulatta), KX061887.1 (G. gorilla), and KX061889.1 (N. leucogenys). The reference sequences for 45S for the M. musculus, D. melanogaster, and C. elegans genomes are NR_046233.2, NR_133558.1, and X03680.1, respectively.
RepeatMasker Data
We downloaded RepeatMasker annotations for the T2T genome from the UCSC Table Browser (table ID: “hub_3671779_t2tRepeatMasker”) in November 2024. We used a simple, exhaustive, brute-force search to find all appearances of the 1,723 pyknons and their reverse complements within the known repetitive sequences. Because multiple entries of this table correspond to overlapping genomic sequences, pyknons or their reverse complements that map to these overlapping regions are counted multiple times, and their total count may exceed the actual genomic count.
Statistical Significance of Pyknons That Are in rRNAs/spacers, mRNAs, and Pre-mRNAs
For each T2T pyknon found in the 10 reference sequences (6 rRNAs and 4 spacers of 45S), we determined whether it is enriched in each of 5S, 45S, 12S, 16S, 5′ untranslated regions (5′-UTRs), coding sequences (CDSs), 3′ untranslated regions (3′-UTRs), mRNAs, introns, and pre-mRNAs, respectively. To this end, we compared a pyknon's observed sense or antisense counts in each region to its expected counts and computed fold enrichments and P-values using a binomial test.
Enrichment Analysis of Gene Sets
We used ShinyGO (Ge et al. 2020) to find enriched GO terms, KEGG pathways, and diseases among those listed in the Disease Alliance ontology (Schriml et al. 2022) and the Disease Rat Genome Database (Kaldunski et al. 2023). In these analyses, we used a false discovery rate (FDR) threshold of 1.0E-03 and a fold enrichment threshold of 1.5 times. To test a pathway enrichment in a given gene list adjusted for genes’ lengths, we adapted a previously proposed approach (Mi et al. 2012). We considered a logistic regression model, where the indicator of pathway membership is a dependent variable, whereas the indicator of gene list membership and gene length are independent variables. We computed a gene's length as the average of pre-mRNA lengths across all transcripts of the gene. We fit the regression model and calculated P-values for independent variables using the statsmodels v0.14.0 Python module.
Searching for Pyknon-Containing Syntenic Sequences and Testing Sequence Conservation
For each pyknon from the expanded G(CGG)5 and (TC)8 groups (described below) with a copy within the 5′-UTRs, CDSs, 3′-UTRs, or introns of a gene, we determined its syntenic region in each of 19 nonhuman genomes (listed under Reference rRNAs above) as follows. First, we flanked the pyknon instance upstream and downstream to include any additional copies of rRNA/spacer pyknons located ≤20 nts from each other. Second, we flanked the extended sequence on the genome with 100 nts on each side (we did not analyze the few pyknon instances that overlap exon-exon junctions). Third, we sought the resulting pyknons + flanks sequence in the genome of each organism in the proximity (±100 kb) of the orthologous gene. We performed the search using the GLSEARCH v36.3.8i (Pearson and Lipman 1988) with the default E-value threshold, enforcing a ≥75% nucleotide identity threshold, and considering only the best alignment (if one existed). We multiple-aligned the resulting sequences using Clustal Omega (Sievers and Higgins 2021). For each position on the multiple alignment, we calculated a conservation score using the formula score = (1-scaled Shannon's entropy) × (1-fraction of gaps) from Valdar (2002). We averaged the conservation scores in the window of positions covered by the pyknon instances in humans and the flanking alignment portions downstream and upstream of the window.
Analysis of Psoralen–Cross-Linked Chimeras
We downloaded SPLASH sequencing reads from poly-A–enriched libraries of lymphoblastoid cell lines, human embryonic stem cells, and retinoic acid–differentiated cells (Aw et al. 2016) from the Sequence Read Archive (SRA accession number: PRJNA318958). To map the reads, we first created a reference FASTA file containing nucleotide sequences spanning full gene lengths (ENSEMBL Rel. 105) and the 10 rRNA/spacer sequences. Then, we mapped the sequenced reads to the reference sequences using the bwa-mem aligner (Li 2013) with the options “-a -c 1000000” to exhaustively report all alignments. According to the original pipeline, we collapsed reads that mapped to identical locations with identical alignment scores to aggressively remove PCR duplicates (Aw et al. 2016). Finally, we extracted and annotated chimeric rRNA|mRNA, spacer|mRNA, and mRNA|mRNA reads requiring exact and correctly oriented matching of the pyknon in each portion of the chimeric read.
Searching for Pyknons in ENCODE's Datasets of RNA-Binding Proteins
We downloaded the raw eCLIP-seq data generated by the ENCODE project (Van Nostrand et al. 2020) for 113 RNA-binding proteins (RBPs) and 2 cell lines (K562, HepG2). For most RBPs, there were 2 replicates and 1 control (IgG-eCLIP) per RBP/cell-line combination. We mapped the reads to the reference human genome using STAR (Dobin et al. 2013), allowing only exact matches and reporting all hits in case of multi-mapping. For each genomic region covered by a sense or antisense instance of the 1,723 rRNA/spacer pyknons, we computed the number of eCLIP-seq reads from each experiment that supported the underlying region using bedtools v2.31.1 (Quinlan and Hall 2010). We sub-selected sites supported by ≥10 RPM and enriched ≥2× compared to the RPM values generated by the matching IgG-eCLIP control experiment.
Searching for Pyknons in Public Small RNA-Seq Datasets
We downloaded 57,597 human short RNA FASTQ files (library_strategy = ncrna-seq or mirna-seq) from the NIH Sequence Read Archive (SRA) on 2025 May 20, using the fasterq-dump tool from the SRA toolkit (https://hpc.nih.gov/apps/sratoolkit.html). We identified sequencing adapters using findAdapt (Chen et al. 2024), FASTQC (Wingett and Andrews 2018), and DNApi (Tsuji and Weng 2016), then removed them using cutadapt (Martin 2011) (error rate = 12%, quality cutoff = 15, minimum length = 16). To profile the molecules and generate normalized expression in reads per million (RPM), we mapped sncRNAs using the process outlined in (Loher et al. 2025), and the mappers (default settings) described in Loher et al. (2021), Loher et al. (2017), and Cherlin et al. (2020). Furthermore, we required that a minimum of 20% of the sequenced reads map to known sncRNA classes, which left us with 39,359 datasets.
Genome-Wide Association Studies Catalog Data
We downloaded the NIH/genome-wide association studies (GWAS) Catalog (“All associations v1.0”) of Sollis et al. (2023) from https://www.ebi.ac.uk/gwas/docs/file-downloads, in April 2024. The downloaded file included 589,865 associations. We sub-selected single-nucleotide changes and filtered them based on P-values using a threshold of 1.0E-08, which left 315,399 single-nucleotide polymorphisms for downstream analyses.
Searching for Pyknons in PsychENCODE Datasets
We downloaded the lists of differentially expressed genes between individuals with ASD and neurotypical controls for 11 cortical areas (Supplementary Data 3 of Gandal et al. (2022)). Separately for each area, we selected statistically significant (FDR ≤ 5.0E-02) upregulated (log2 fold change ≥ 0.2) and downregulated (log2 fold change ≤ −0.2) protein-coding genes. We then used a hypergeometric test to assess enrichment of our gene collection of interest. We further tested the genes that passed these filters for enrichment in the 1,169 ASD risk genes of the SFARI Gene database that we downloaded from the Simons Foundation portal on 2024 June 25.
Results
We sought the 75,955,725 16-mers found in the mRNAs of Rel. 105 of ENSEMBL in the T2T genome (see Materials and Methods) and kept the subset of 2,635,947 16-mers with at least 30 exact copies across the T2T genome. We processed this subset of candidate pyknons to remove redundancies (see Materials and Methods). Then, we intersected the final set with the GenBank reference sequences for the 6 human rRNAs and 4 transcribed spacers of 45S.
The 4 Nuclear rRNAs and the 4 Spacers of 45S Are Densely Packed With Pyknons
We found 1,723 pyknons in the sequences of the 6 rRNAs and 4 45S spacers: 843 pyknons are in sense, 848 in antisense, and 32 in both orientations (Table S1a). All sequences but 12S and 16S are densely packed with pyknons in sense and antisense (Figures S1 and S2). Figure 1 highlights this point with the help of secondary structures for 2 rRNAs (18S, 28S), determined by R2DT (Sweeney et al. 2021), and 2 spacers (ITS1, ETS2), determined by RNAfold (Lorenz et al. 2011)—we will discuss the meaning of the different colors in the next section. In each structure, most pyknon-covered regions span more than 16 nts, the result of distinct pyknons overlapping partially. Table S1b lists the sequence, location, and orientation of pyknons in each rRNA and 45S spacer, whereas Table S1c lists the pyknons’ counts in the T2T genome.
Fig. 1.
Secondary structures of 18S and 28S rRNAs and the ITS1 and ETS2 spacers. The pyknon-covered regions are highlighted. Red: instances of sequences from groupomni-TC. Orange: instances of sequences from groupomni-CGG. Yellow: instances of sequences from grouppunct. The styles of secondary structures were modified in RNAcanvas (Johnson and Simon 2023).
The Pyknons From the rRNAs and 45S Spacers Match Very Short Snippets From Long Interspersed Nuclear Elements
By definition, all 1,723 pyknons found in rRNAs and spacers are sense to mRNAs. Because they also have many additional exact copies in the genome, we examined whether any of them or their reverse complements—a total of 2,810 sequences—are found within RepeatMasker-annotated repeat elements (see Materials and Methods). Almost all sequences, 2,696 out of 2,810 (95.9%), have one or more exact copies in long interspersed nuclear elements (LINEs) (Table S1d). LINE-1 elements harbor 90.1% of the 2,810 sequences, and LINE-2 elements 74.3%. Multiple sequences are in both LINE-1 and LINE-2 elements. We examined the sequence segments within LINEs that are covered by rRNA/spacer pyknons or their reverse complements and found them to form short, isolated islands: 74.6% of these 117,917 islands are ≤22 nts, whereas 89.0% of the islands are ≤30 nts (data not shown).
The Pyknons Found in the rRNAs and the Transcribed Spacers of 45S Form 3 Groups
We heuristically divided the entries of Table S1c and their reverse complements, which have the same genomic counts, into 2 collections: sequences with very high genomic counts and their reverse complements, and all else. Using an iterative procedure (see Materials and Methods), we let the collection with high counts “attract” sequences from the other collection if those sequences were a Levenshtein distance of 1 away. Upon termination, the augmented first collection comprised 2 subgroups of sequences. The first subgroup had 34 sequences that included TCTCTCTCTCTCTCTC or “(TC)8,” its reverse complement, and their sequence variants; we will refer to this group as “groupomni-TC” (Table S1e). The second subgroup had 141 sequences that included GCGGCGGCGGCGGCGG or “G(CGG)5,” its reverse complement, and their sequence variants; we will refer to this group as “groupomni-CGG” (Table S1e). There remained 2,635 sequences in the second collection; we will refer to them as “grouppunct.” The modifier “omni” acknowledges the genomic omni-presence of the sequences in the corresponding 2 subgroups, and their generally high genomic copy numbers. The modifier “punct” highlights its members’ generally lower genomic copies and their punctate presence in protein-coding genes (Table S1e). In Fig. 1, we used “red,” “orange,” and “yellow” to color the instances of the sequences belonging to groupomni-TC, groupomni-CGG, and grouppunct, respectively.
(TC)8 and G(CGG)5 Are Recent Evolutionary Insertions to 28S rRNA
(TC)8 is sense to 28S rRNA at one location (starting at position 3052) and antisense to ITS1 at 3 locations (starting at positions 191, 193, and 195, respectively). G(CGG)5 is sense to 28S at 9 locations (starting at positions 849, 852, 870, 873, 876, 2153, 2156, 2159, and 3489, respectively) and antisense to ITS1 at one location (starting at position 741). Figure 2 shows a multiple sequence alignment of 28S rRNA sequences from 20 organisms, spanning the spectrum from H. sapiens to C. elegans. The 10 instances of (TC)8 and G(CGG)5 are embedded in the “expansion segments” of the 28S rRNA (Ramesh and Woolford 2016; Hariharan et al. 2023). When comparing distant organisms, e.g. human and mouse, the expansion segments are evident as large insertions, whereas the differences among the 6 primates at those sites are smaller. However, Fig. 2 makes evident that the differences between humans, chimps, and gorillas, on the one hand, and all other primates and non-primates, on the other, correspond to the insertions of the 10 instances of (TC)8 and G(CGG)5. At 15 to 20 million years away, only a single copy of G(CGG)5 remains in the 28S rRNA of N. leucogenys (white-cheeked gibbon).
Fig. 2.
Multiple sequence alignment of 28S rRNA sequences from 20 organisms shows the existence and location of the previously reported expansion segments of 28S rRNAs. A position of the alignment is colored green if the corresponding 21-nt symmetric window has ≥75% of its nts matching the human sequence. Red and orange stars show the locations of the (TC)8 and G(CGG)5 pyknons, respectively. Note how in primates, the pairwise differences revolve around the insertions of the (TC)8 and G(CGG)5 pyknons.
The Pyknons of rRNAs and Spacers Have Exact Copies in Half of All Protein-Coding Genes
We sought sense or antisense copies of the sequences in groupomni-TC, groupomni-CGG, and grouppunct in the introns and exons of all 22,511 protein-coding genes in Rel. 105 of ENSEMBL. We found that one or more of these sequences appear in the unspliced transcripts of 11,034 (49.0%) genes (Table S2a). Combined, the sense and antisense copies of the 1,723 pyknons across rRNAs, spacers, and the introns and exons of protein-coding genes span 2,436,551 nts.
The Pyknons of rRNAs and Spacers Are Recent Insertions Into the Genomes of Primates
Given that the instances of (TC)8 and G(CGG)5 are absent from the 28S rRNAs of non-primates (Fig. 2), we hypothesized that the same holds for the other pyknons we find in rRNAs and spacers of 45S and for those pyknons’ copies in the introns and exons of protein-coding genes.
The pyknons of the human 45S cassette are not conserved in the 45S cassettes of other organisms. We focused on the subset of 2,754 sequences from the 45S cassette and sought them in the 45S cassettes of 8 other organisms: 5 nonhuman primates, the mouse, the fruit fly, and the worm. The number of human sequences in the 45S cassettes of these 8 organisms decreased monotonically with their phylogenetic distance from H. sapiens. Only 1,537 of the 2,754 human sequences are in P. troglodytes, 1,387 in G. gorilla, 1,111 in P. abelii, 914 in M. mulatta, 898 in N. leucogenys, 418 in M. musculus, 58 in D. melanogaster, and 48 in C. elegans.
The intronic and exonic instances of human rRNA/spacer pyknons are not conserved in other organisms. The 1,723 rRNA/spacer pyknons have 31,319 sense or antisense copies in the spans of 11,034 protein-coding genes (above and Table S2a). We flanked each copy with 100 nts on each side, identified its syntenic counterparts in 19 other genomes (see Materials and Methods), and created the corresponding multiple sequence alignment. Separately for each flanking region, we computed a conservation score for the pyknon span. Figure 3a and Figure S3a show representative multiple sequence alignments across several organisms of syntenic regions from 7 genes. Note how (TC)8 (Fig. 3a) and G(CGG)5 (Figure S3a) are present in the human sequences but absent from most other orthologs. Figure 3b highlights the difference between the conservation scores of the pyknon span and flanking regions, respectively. In 3′-UTRs and introns, ∼75% of all pyknon instances are less conserved than the nucleotide segments immediately flanking the pyknon. In 5′-UTRs and CDSs, the differences in conservation percentages are still above 50% and highly statistically significant. Figure 3c shows the average number of nts (i.e. non-gaps) corresponding to the pyknon portion of the multiple sequence alignments for the intronic copies and the different genomes. All non-primates show a striking absence of intronic copies for the 1,723 pyknons or their antisense, indicating that most of these copies are recent insertions in the genomes of primates. Figure S3b shows counterpart plots for 5′-UTRs, CDSs, and 3′-UTRs.
Fig. 3.
Conservation of groupomni-TC, groupomni-CGG, and grouppunct instances across organisms. a) Multiple sequence alignments of syntenic regions from 4 representative genes. The human sequence contains an instance of (TC)8 from groupomni-TC. The examples show a lack of conservation in organisms that are phylogenetically distant from humans. b) Boxplots showing the level of conservation of all exonic and intronic instances of the 2,810 sequences from groupomni-TC, groupomni-CGG, and grouppunct instances vs. sequences in their immediate proximity (flanks). The analysis examined 31,319 instances across 11,034 genes. The x axis measures the difference between the conservation scores of a pyknon instance and its immediate upstream or downstream flank. Positive values indicate that the pyknon instance is more conserved compared to its flanks, whereas negative values indicate the opposite. c) Boxplots showing the number of nts (i.e. non-gaps) spanning the regions containing the 26,109 intronic instances of the 2,810 sequences of groupomni-TC, groupomni-CGG, and grouppunct.
Additionally, we examined whether either (TC)8 or G(CGG)5 is part of a wider motif encompassing flanking nts that perhaps differ from the recurring “TC” and “CGG” cores. We generated sequence logos (Crooks et al. 2004) after enumerating the pyknon's combined 180,955 T2T genomic copies and flanking each copy with 20 nts on each side. Figure S4 shows that these 2 pyknons mark variable-width runs of the “TC” and “CGG” cores.
Each rRNA/spacer Contains a Unique Set of Pyknons That It Shares With a Unique Set of mRNAs
We compared the pyknons found in each of the 6 rRNAs and 4 transcribed spacers of 45S (Table S1c) to determine whether they have more than one copy within these 10 sequences. We carried out the analysis separately for the 875 sense and 880 antisense pyknons. Table 1a and b show that any 2 of the 10 reference sequences share no more than a handful of pyknons, independently of the pyknons’ orientation in the rRNAs or spacers.
Table 1.
Shared pyknons
| a | 5S | ETS1 | 18S | ITS1 | 5.8S | ITS2 | 28S | ETS2 | 12S | 16S | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 6 | 250 | 34 | 121 | 2 | 102 | 333 | 25 | 2 | 10 | ||
| 5S | 6 | 6 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| ETS1 | 250 | 250 | 0 | 0 | 0 | 0 | 5 | 2 | 0 | 0 | |
| 18S | 34 | 34 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||
| ITS1 | 121 | 121 | 0 | 2 | 0 | 0 | 0 | 0 | |||
| 5.8S | 2 | 2 | 0 | 0 | 0 | 0 | 0 | ||||
| ITS2 | 102 | 102 | 0 | 0 | 0 | 0 | |||||
| 28S | 333 | 333 | 3 | 0 | 0 | ||||||
| ETS2 | 25 | 25 | 0 | 0 | |||||||
| 12S | 2 | 2 | 0 | ||||||||
| 16S | 10 | 10 |
| b | 5S | ETS1 | 18S | ITS1 | 5.8S | ITS2 | 28S | ETS2 | 12S | 16S | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 10 | 242 | 35 | 116 | 0 | 125 | 332 | 30 | 2 | 1 | ||
| 5S | 10 | 10 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| ETS1 | 242 | 242 | 0 | 0 | 0 | 0 | 5 | 2 | 0 | 0 | |
| 18S | 35 | 35 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||
| ITS1 | 116 | 116 | 0 | 2 | 0 | 0 | 0 | 0 | |||
| 5.8S | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||||
| ITS2 | 125 | 125 | 3 | 0 | 0 | 0 | |||||
| 28S | 332 | 332 | 3 | 0 | 0 | ||||||
| ETS2 | 30 | 30 | 0 | 0 | |||||||
| 12S | 2 | 2 | 0 | ||||||||
| 16S | 1 | 1 |
The second row and second column show how many pyknons are sense (a) or antisense (b) in the 6 rRNA and 4 45S spacers. The remaining rows and columns show how many of the pyknons of a given orientation are shared by any 2 of the 10 reference sequences. Note how the pyknons are unique to each rRNA and 45S spacer. Nonzero table values are highlighted in bold.
We repeated the comparison after associating each of the 10 sequences with the genes whose mRNAs contain the sequence's pyknons, then calculated the pairwise overlap of the 10 groups at the level of the genes while also considering the relative orientation of pyknons in rRNAs/spacers and mRNAs. To avoid multiple counting, we compared shared pyknons at the level of gene identifiers, not mRNA identifiers. Table 2 summarizes our findings. We can see several off-diagonal pairs with high counts, e.g. ETS1-ITS1, ETS1-28S, ITS1-ITS2, ITS1-28S, and ITS2-28S. This crosstalk is expected and results from segments internal to these spacers/rRNAs that are duplicates or reverse complements of one another. For example, the red-colored region in the secondary structure of ITS1 and 28S (Fig. 1) corresponds to (GA)8 and (TC)8, respectively. Examining pyknons that are antisense to mRNAs may seem paradoxical, considering that pyknons must have, by definition, one sense copy in an mRNA. However, note that pyknons that are sense in one mRNA can appear in antisense orientation in other mRNAs.
Table 2.
Number of genes linked to rRNAs and 45S spacers for all possible relative orientations
| rRNA & mRNA same | rRNA & mRNA opposite | spacer & mRNA same | spacer & mRNA opposite | rRNA & mRNA same | rRNA & mRNA opposite | spacer & mRNA same | spacer & mRNA opposite | rRNA & mRNA same | rRNA & mRNA opposite | spacer & mRNA same | spacer & mRNA opposite | rRNA & mRNA same | rRNA & mRNA opposite | spacer & mRNA same | spacer & mRNA opposite | rRNA & mRNA same | rRNA & mRNA opposite | rRNA & mRNA same | rRNA & mRNA opposite | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5S | ETS1 | 18S | ITS1 | 5.8S | ITS2 | 28S | ETS2 | 12S | 16S | |||||||||||||
| 6 | 10 | 413 | 388 | 41 | 42 | 848 | 1,106 | 2 | 0 | 245 | 322 | 1,584 | 1,289 | 94 | 89 | 3 | 2 | 14 | 3 | |||
| rRNA & mRNA same | 5S | 6 | 6 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| rRNA & mRNA opposite | 10 | 10 | 0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||
| spacer & mRNA same | ETS1 | 413 | 413 | 28 | 1 | 1 | 51 | 62 | 0 | 0 | 19 | 18 | 87 | 84 | 13 | 15 | 0 | 0 | 0 | 0 | ||
| spacer & mRNA opposite | 388 | 388 | 0 | 0 | 45 | 61 | 0 | 0 | 18 | 22 | 83 | 83 | 12 | 15 | 0 | 0 | 0 | 0 | ||||
| rRNA & mRNA same | 18S | 41 | 41 | 0 | 1 | 2 | 0 | 0 | 1 | 1 | 3 | 2 | 1 | 0 | 0 | 0 | 0 | 0 | ||||
| rRNA & mRNA opposite | 42 | 42 | 4 | 4 | 0 | 0 | 2 | 2 | 8 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | ||||||
| spacer & mRNA same | ITS1 | 848 | 848 | 117 | 0 | 0 | 74 | 56 | 150 | 654 | 8 | 9 | 0 | 0 | 1 | 0 | ||||||
| spacer & mRNA opposite | 1,106 | 1,106 | 0 | 0 | 41 | 96 | 886 | 174 | 9 | 10 | 0 | 0 | 0 | 0 | ||||||||
| rRNA & mRNA same | 5.8S | 2 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||||||||
| rRNA & mRNA opposite | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ||||||||||
| spacer & mRNA same | ITS2 | 245 | 245 | 14 | 54 | 94 | 5 | 1 | 0 | 0 | 0 | 0 | ||||||||||
| spacer & mRNA opposite | 322 | 322 | 113 | 79 | 4 | 2 | 0 | 0 | 0 | 0 | ||||||||||||
| rRNA & mRNA same | 28S | 1,584 | 1,584 | 233 | 30 | 24 | 0 | 0 | 0 | 0 | ||||||||||||
| rRNA & mRNA opposite | 1,289 | 1,289 | 26 | 28 | 0 | 0 | 1 | 0 | ||||||||||||||
| spacer & mRNA same | ETS2 | 94 | 94 | 16 | 0 | 0 | 0 | 0 | ||||||||||||||
| spacer & mRNA opposite | 89 | 89 | 0 | 0 | 0 | 0 | ||||||||||||||||
| rRNA & mRNA same | 12S | 3 | 3 | 0 | 0 | 0 | ||||||||||||||||
| rRNA & mRNA opposite | 2 | 2 | 0 | 0 | ||||||||||||||||||
| rRNA & mRNA same | 16S | 14 | 14 | 0 | ||||||||||||||||||
| rRNA & mRNA opposite | 3 | 3 | ||||||||||||||||||||
Off-diagonal entries show the level of crosstalk among the various groups of genes. Notation: “same” means same relative orientation; “opposite” means opposite relative orientation. Nonzero table values are highlighted in bold.
The Sequences of Groupomni-TC, Groupomni-CGG, and Grouppunct Appear in Specific mRNA Regions in Specific Genes
The sequences of groupomni-TC, groupomni-CGG, or grouppunct are found in the mRNAs (exons only) of 3,893 genes (Table S2b). We asked whether these sequences’ 5′-UTR instances favor specific gene groups. We then asked the same question for their CDS and their 3′-UTR instances. Figure 4 and Figure S5 show that the rRNA/spacer motifs are prevalent in the mRNAs of nervous system genes and developmental genes that also have established associations with brain and genetic disorders.
Fig. 4.
Enrichments of genes whose 5′-UTRs, CDSs, and 3′-UTRs share pyknons in sense (blue marks) or antisense (red marks) with rRNAs and 45S spacers. Shown are the enrichments for biological processes and cellular components.
We also asked whether the copies of the rRNA/spacer sequences we find in each mRNA region (5′-UTR, CDS, or 3′-UTR) come from a specific group among groupomni-TC, groupomni-CGG, and grouppunct. Figure S6 shows a division of labor among the 9 possible combinations (3 mRNA regions times 3 groups of pyknons). Namely, the sequences of groupomni-TC, groupomni-CGG, and grouppunct appear in different mRNA regions of largely nonoverlapping gene sets that belong to nonoverlapping biological processes.
Inspection of Table S2a and b shows that most instances of the pyrimidine-rich groupomni-TC sequences are in introns. We asked whether the instances correspond to the “polypyrimidine tract” that is typically located immediately upstream of constitutive 3′ splice sites (Corvelo et al. 2010). We found that only 1.9% of the groupomni-TC intronic copies are ≤60 bases of the nearest downstream exon, indicating that groupomni-TC sequences are unrelated to the canonical polypyrimidine signal.
Only the Genes Whose Pre-mRNAs Pair Up a Sequence From Groupomni-TC With a Sequence From Grouppunct Are Linked to the Nervous System and Development
Table S2b shows that approximately one-half of the genes whose pre-mRNAs contain groupomni-TC sequences do not contain grouppunct sequences and vice versa. The 2,941 genes that contain exclusively groupomni-TC sequences are enriched (FDR ≤ 1.0E-03) for ion transport and metabolism (Fig. 5). On the other hand, the 3,191 genes containing exclusively grouppunct sequences are enriched (FDR ≤ 1.0E-03) for diverse processes. Note also that only 576 genes contain exclusively groupomni-CGG sequences in their pre-mRNAs. These genes are involved in transcription and RNA biogenesis, suggesting a unique role for those groupomni-CGG sequences that do not appear together with either groupomni-TC or grouppunct sequences. The finding also captures a novel link between transcription-related genes and the translation-related rRNAs.
Fig. 5.
a) The 3 groups of sequences (groupomni-TC, groupomni-CGG, and grouppunct) appear in the pre-mRNAs of 11,034 genes. Most of these genes contain either a sequence from groupomni-TC or a sequence from grouppunct: only the genes that contain at least one sequence from each of these 2 groups are enriched in the nervous system and multiple brain disorders. We generated the shown GO term enrichments with ShinyGO (Ge et al. 2020) (FDR ≤ 1.0E-03). For clarity of presentation, not more than the top 20 terms of each list are shown.
Of particular interest are the 3,430 genes whose pre-mRNAs simultaneously contain at least one sequence from each of groupomni-TC and grouppunct. These 3,430 genes have many strong links (FDR ≤ 1.0E-03) to development and the nervous system (Fig. 5). The enriched cellular component terms include axon, dendrite, and synapse. The enriched KEGG terms include signaling- and synapse-related pathways. Moreover, the Disease Alliance ontology and the Disease Rat Genome Database show a striking enrichment of disease terms. Examination of the enriched terms shows that 1,046 of the 3,430 genes are linked to multiple brain and developmental disorders, including autism spectrum disorder (ASD), schizophrenia, bipolar disorder, and epilepsy.
The Enrichment of Nervous System Genes Containing Both Groupomni-TC and Groupomni-CGG Sequences Remains After Accounting for Length
The pre-mRNAs of nervous system genes are typically long (McCoy and Fire 2024), raising the possibility that the sequences of groupomni-TC, groupomni-CGG, and grouppunct appear in these genes by chance. We repeated the enrichment analysis of Fig. 5 by including a gene's genomic span as an independent covariate. Virtually all (484 of 494) of the previously enriched GO biological process terms corresponding to the 3,430 (=2,673 + 757) genes remain significant (FDR ≤ 1.0E-03, fold enrichment ≥ 1.5), indicating the nonrandom presence of these sequences in nervous system genes.
Each rRNA/spacer Is Associated With a Different Set of Human Disorders Through Its Pyknons
Figure 5 highlights the importance of sequences from groupomni-TC and grouppunct co-occurring in pre-mRNAs, whereas Table 1 shows that each rRNA/spacer contains a unique set of pyknons. Since the groupomni-TC sequences are omnipresent, we reasoned that it is the sequences of grouppunct that specify the associations of rRNAs/spacers with different disorders. We tested this hypothesis by sub-selecting those grouppunct sequences in the i-th rRNA/spacer (Table S1c and d), identifying which among them are in any of the 3,430 genes of Fig. 5, and computing enrichments for the genes that contain them. We found the following associations (FDR ≤ 1.0E-03):
5S rRNA ↔ ASD
28S rRNA ↔ substance-related disorders, neurodevelopmental disorders, epilepsy, bipolar disorder, ASD, atrial fibrillation, primary ovarian insufficiency, and prostate cancer
ETS1 ↔ substance-related disorders, developmental disabilities, neurodevelopmental disorders, thrombocytosis, atrial fibrillation, and ASD
ITS1 ↔ substance abuse, developmental disabilities, neurodevelopmental disorders, epilepsy, bipolar disorder, ASD, and prostate cancer
ITS2 ↔ substance-related disorders, neurodevelopmental disorders, intellectual disability, epilepsy, and atrial fibrillation
ETS2 ↔ substance abuse, neurodevelopmental disorders, Hirschsprung's disease, thrombocytosis, epilepsy, atrial fibrillation, ASD, and colorectal cancer
It is noteworthy that several disorders considered typically human (e.g. ASD, schizophrenia, bipolar disorder) are associated with the human-specific sequences of the transcribed spacers ETS1, ITS1, ITS2, and ETS2. We found no notable associations for genes containing the grouppunct sequences from each of the remaining 4 rRNAs.
The Association of Nervous System Genes With Disorders Can Be Human-Specific and Depend on a Pyknon's Orientation
We examined further the association between rRNAs/spacers and disorders with the help of a pyknon CGGGCGTGGTGGTGGG from grouppunct, which is in ETS2. Of the genes whose pre-mRNAs contain groupomni-TC sequences, we sub-selected those that also contain CGGGCGTGGTGGTGGG in sense (551 genes), antisense (671 genes), or both orientations simultaneously (171 genes). Figure S7 shows that the relative orientation of this pyknon matters: each gene group is enriched for nervous system genes involved in different biological processes, cellular components, molecular functions, and disorders. Moreover, this ETS2 pyknon is human-specific; therefore, its associations with the nervous system are also human-specific (Figure S7).
Groupomni-CGG Sequences Are Contact Points for rRNA|mRNA and mRNA|mRNA Heteroduplexes
Given the recurrence of exact sense and antisense pyknon copies in rRNAs, spacers, and mRNAs, we also tested the hypothesis that pyknons serve as contact points for long-distance rRNA|mRNA, spacer|mRNA, and mRNA|mRNA interactions by analyzing the psoralen–cross-linking data of Aw et al. (2016) from 3 cell lines. We found multiple rRNA|mRNA and spacer|mRNA chimeras that primarily consisted of transcripts containing sequences from groupomni-CGG paired with their reverse complement; these pairings occurred 10 to 100 times more frequently than expected by chance (FDR ≤ 5.0E-02; Table S3a). The mRNA segments that form these chimeras are mainly in 5′-UTRs and CDSs (Table S3b), a result that Fig. 4 and Figure S5 anticipated. Very few chimeras involve sequences from groupomni-TC or grouppunct. We also found mRNA|mRNA heteroduplexes involving mRNAs from different genes. These, too, comprise mainly 5′-UTR segments containing groupomni-CGG sequences paired with their reverse complement (FDR ≤ 5.0E-02; Table S3c and d).
The Binding Sites of Numerous RBPs Contain Exact Copies of rRNA/spacer Pyknons
We also investigated the possibility that RBPs bind regions containing pyknon copies. We searched all 1,723 rRNA/spacer pyknons in the eCLIP-seq data of the 113 RPBs cataloged by the ENCODE project (Van Nostrand et al. 2020). We sub-selected sites supported by ≥10 RPM whose support was ≥2× higher than observed in the matching control (IgG-eCLIP) experiment. Of the 1,723 pyknons, 941 (54.6%) are fully embedded, in sense or antisense, in the binding sites of all 113 analyzed RBPs in various combinations (Table S4). These binding sites are in rRNAs, mRNAs, or pre-mRNAs. We found poor support for the pyrimidine-rich sequences of groupomni-TC in the binding sites of PTBP1, U2AF1, or U2AF2 all 3 of which have an affinity for pyrimidine-rich regions (Voith von Voithenberg et al. 2016; Ye et al. 2023): at the RPM and fold-enrichment thresholds mentioned above, we found 157 PTBP1 binding sites across K562 and HepG2 cells, 5 U2AF2 binding sites, and none for U2AF1, but only a few of them occurred in introns. This agrees with the above-mentioned finding that the groupomni-TC sequences are unrelated to the canonical polypyrimidine tract.
Endogenous Abundant rRNA-Derived sncRNAs Carry Copies of rRNA/spacer Pyknons
We hypothesized that sncRNAs carrying rRNA/spacer pyknons may be involved in molecular decoying of RBPs or RNA|RNA heteroduplexes, in analogy to microRNAs (Poliseno et al. 2010; Rigoutsos and Furnari 2010; Tay et al. 2011; Zhuang et al. 2013) and tRNA-derived fragments (tRFs) (Goodarzi et al. 2015). For this potential to exist, endogenous sncRNAs, presumably rRFs (Cherlin et al. 2020, 2024), would need to exist that carry rRNA/spacer pyknons, or the latter's reverse complement. To test this, we searched the 39,359 datasets from NIH's SRA that passed quality control (see Materials and Methods and Table S5a). We found that 1,964 (69.9%) of the 2,810 sequences in groupomni-TC, groupomni-CGG, and grouppunct are carried by 116,915 distinct sncRNAs (microRNAs, tRFs, rRFs), with each of the latter being supported by ≥100 reads in at least one analyzed dataset. Note that the same sncRNA can contain instances of 2 or more overlapping pyknons. 114,241 of the 116,915 (97.7%) sncRNAs are rRFs, as expected. Table S5b lists for each of the 1,964 sequences the number of SRA datasets among the 39,359 in which it was found at ≥100 reads, and the number of distinct sncRNAs that carry it.
The Genomic Copies of rRNA/spacer Pyknons Overlap Risk Variants for Brain Disorders and Other Conditions
We also hypothesized that the genomic copies of rRNA/spacer pyknons overlap known risk polymorphisms for the human disorders we already discussed and perhaps other conditions. To investigate this possibility, we searched the NHGRI/EBI GWAS Catalog (Sollis et al. 2023), considering only variants with an associated P-value ≤ 1.0E-08. We found 131 variants that overlap the genomic copies of 82 rRNA/spacer pyknons (Table S5c), corresponding to a ∼42-fold enrichment (P-value < 3.9E-12). Of the 131 variants, 90 are within genomic copies of groupomni-TC sequences. The variants we identified are linked to brain disorders, including ASD, schizophrenia, bipolar disorder, depression, and also to other conditions/traits.
Genes Containing rRNA/spacer Pyknons Are Downregulated in Specific Cortical Regions of ASD Patients and Include ASD Risk Genes
We analyzed RNA-seq data from bulk brain samples that the PsychENCODE consortium (Gandal et al. 2022) generated from the post-mortem brain samples (11 cerebral cortex areas) of 49 ASD patients and 54 controls (see Materials and Methods). For each cortical area, we assessed if the 3,430 genes whose pre-mRNAs simultaneously contain at least one sequence from each of groupomni-TC and grouppunct are enriched among genes differentially expressed in ASD. We analyzed downregulated and upregulated genes separately. We found that the 3,430 genes containing rRNA/spacer pyknons are enriched only among the genes that are downregulated in ASD patients and only in Brodmann areas (BAs) 22/41/42 (primary auditory cortex, language comprehension) and BA 17 (primary visual cortex) (Figure S8). For BA 22/41/42, we found 236 of the 3,430 among the downregulated genes (enrichment = 1.5, P-value = 7.5E-12) with 54 of the 236 being known ASD risk genes and listed in the SFARI Gene collection (enrichment = 2.31, P-value = 7.8E-09). For BA 17, we found 256 of the 3,430 among the downregulated genes (enrichment = 1.4, P-value = 3.1E-10) with 50 of the 256 being known ASD risk genes and listed in the SFARI Gene collection (enrichment = 1.87, P-value = 1.4E-05). See Table S5d for the lists of genes.
The Sharing of Organism-Specific Motifs Between rRNAs/spacers and Nervous System and Developmental Genes Is Conserved Across Evolution
Lastly, we hypothesized that the sharing of organism-specific rRNA/spacer pyknons with nervous system and developmental genes is not unique to humans. Since most human pyknons arise from the 45S cassette (Figures S1 and S2), if such sharing exists in other organisms, it ought to be evident when examining the 45S sequence only.
We selected 3 model organisms for this analysis: mouse, fruit fly, and worm. Because the fruit fly and the worm have small genomes (140 and 100 Mbps, respectively), we reasoned that the counterpart pyknons likely comprise fewer than 16 nts. Using a Monte Carlo simulation, we determined the probability of a random 13-mer appearing by chance 30 or more times in the fruit fly or the worm to be <1.0E-03. For simplicity, we used redundant pyknons in these calculations (see Materials and Methods). We also removed all 16-mers that contained 15 A's or 15 T's.
We heuristically divided each organism's pyknons into 2 groups based on genomic counts to create the counterparts to the human groupomni-TC and grouppunct. We did not seek a counterpart to the groupomni-CGG since we already showed that these sequences are an innovation of primates (Fig. 2). The mouse counterparts to the human groupomni-TC and grouppunct sets contain 24 and 172 sequences, respectively (Table S6a). The fruit fly counterparts contain 272 and 2,540 sequences, respectively (Table S6e). And the worm counterparts contain 34 and 268 sequences, respectively (Table S6i). There are 10,208 mouse genes, 6,802 fruit fly genes, and 6,561 worm genes containing sequences from either groupomni or grouppunct in their pre-mRNAs (Table S6b, c, f, g, j, k).
In complete analogy to what we found in humans, in all 3 model organisms, only those genes whose pre-mRNAs contain at least one sequence from the counterpart of groupomni-TC and at least one sequence from grouppunct are associated with nervous system and developmental genes and brain and genetic disorders. Figures S9 to S11 show the terms and pathway enrichments for mouse, fruit fly, and the worm, respectively. In each case (Table S6d, h, l), the enriched terms mirror those of the human genome. In the worm, the 3 gene sets captured by the Venn diagram of Figure S11 show a modest pathway crosstalk; however, the enrichments and statistical significance of the genes containing at least one pyknon from each of groupomni and grouppunct far outweigh this crosstalk.
We also sought each organism's 45S cassette pyknons (sense and antisense instances) in the 45S cassettes of the other organisms: there are 2,754 human, 196 mouse, 2,812 fruit fly, and 302 worm sequences. Table 3 lists the pairwise overlaps and demonstrates their organism-specific nature.
Table 3.
Searching the pyknons found in the 45S cassette of the human, mouse, fruit fly, and worm genomes in sense or antisense in the 45S cassette of the other 3 genomes reveals organism-specific sequences.
| Target | ||||
|---|---|---|---|---|
| H. sapiens (45S cassette) | M. musculus (45S cassette) | D. melanogaster (45S cassette) | C. elegans (45S cassette) | |
| Source | ||||
| H. sapiens pyknons in 45S (16 nts) | 2,754 | 418 | 56 | 48 |
| M. musculus pyknons in 45S (16 nts) | 12 | 196 | 2 | 0 |
| D. melanogaster pyknons in 45S (13 nts) | 120 | 124 | 2,812 | 54 |
| C. elegans pyknons in 45S (13 nts) | 4 | 12 | 8 | 302 |
The diagonal values are highlighted in bold.
Discussion
We analyzed a specific type of genome-wide-conserved motifs with organism-specific sequences, the pyknons. We showed that the 6 rRNAs (5S, 5.8S, 12S, 16S, 18S, 28S) and the 4 transcribed spacers of 45S (ETS1, ITS1, ITS2, ETS2) are packed with pyknons they share with the pre-mRNAs of 11,034 human genes. Unlike previous work that sought similarities between full-length rRNAs and genes (Mauro and Edelman 1997; Tranque et al. 1998) or used variable-length k-mers (Parker et al. 2015, 2018) but studied only the expansion segments (Fujii et al. 2018) of 18S and 28S, we examined and presented results on the 4 nuclear rRNAs, the 4 transcribed spacers of 45S, the 2 mitochondrial rRNAs, and the links of these rRNAs/spacers to intronic and exonic sequences, intergenic space, experimentally determined RNA-RNA interactions, RBP assays, transcriptomic data from tens of thousands of public datasets, GWAS-derived risk variants, and data from ASD patients. We also examined the conservation of the uncovered relationships across evolution.
We found a specific combination of rRNA and spacer pyknons that occurs primarily in the pre-mRNAs of nervous system and developmental genes, including many risk genes for brain and genetic disorders. Nearly all of these pyknons are carried by rRFs, the recently reported class of sncRNAs that are derived from rRNAs and spacers (Lambert et al. 2019; Cherlin et al. 2020; Pliatsika et al. 2024). This sharing of pyknons between rRNAs/spacers and nervous system genes is not unique to humans but extends to at least 3 more organisms (mouse, fruit fly, worm), with each organism using its own set of pyknons. The recurrence of the finding across 600 million years of evolution suggests that it is an essential component of these organisms. The finding implicates, for the first time, the sequences of human rRNAs and the transcribed spacers of 45S in human nervous system disorders, through primate-specific pyknons shared between rRNAs/spacers and specific nervous system genes.
Each human rRNA and transcribed spacer contains a unique set of pyknons (Table 1 and Table S1a to c). Except for 12S and 16S, the rRNAs and spacers are very rich in pyknons (Figures S1 and S2). Figure 1 shows the secondary structures for 2 rRNAs and 2 spacers with the pyknon-covered portions highlighted. Most rRNA/spacer pyknons have additional copies in transposable elements, specifically LINE-1 and LINE-2 (Table S1d), where they span very short islands: 89.0% of these pyknon-covered islands are ≤30 nts. LINEs are believed to follow a random distribution across chromosomes (Garcia et al. 2024), but these shared segments, their short span, the rRNA/spacer pyknons’ nonrandom placement, and the importance of rRNAs and 45S spacers for an organism's well-being raise a question about how this sharing could have occurred. We are neither aware of nor can we speculate about processes that may underlie these observations.
The rRNA/spacer pyknons could be divided into 3 groups, each with its own properties: groupomni-TC, groupomni-CGG, and grouppunct (Table S1e), with (TC)8 and G(CGG)5, respectively, leading the first 2 groups. These 2 leaders are notable. Among other things, their instances in 28S represent very recent insertions into the human genome (Fig. 2): most of these instances are absent from N. leucogenys, approximately 15 to 20 million years away. Interestingly, the recency of these segments extends to all the rRNA/spacer pyknons: their tens of thousands of copies in the exons and introns of protein-coding genes (Fig. 3a and b; Figure S3a) are unique to primates. We note here that not all pyknon-covered regions correspond to rRNA expansion segments (Fujii et al. 2018), nor are all expansion segments covered by pyknons.
We also note the multiple instances of the trinucleotide CGG and its variants in 28S (Fig. 2) and the 5′-UTRs and CDSs of genes (Table S2b and Figure S3a). CGG runs were previously discussed (Kleiderlein et al. 1998; Boivin and Charlet-Berguerand 2022) in the context of neuropathologies that are typically human (Kleiderlein et al. 1998; Chen et al. 2003; Metsu et al. 2014; Ishiura et al. 2019; Annear et al. 2021). And Fig. 2 shows that the copies of G(CGG)5 in 28S occur only in primates. The only exception is the 28S rRNA of the odontocete T. truncatus (bottlenose dolphin), a non-primate mammal known to develop Alzheimer's disease–like pathology (Vacher et al. 2023): T. truncatus contains a single copy of G(CGG)5 in its 28S rRNA.
Even when we confine our analysis to only the exonic instances of the rRNA/spacer pyknons, a clear “division of labor” emerges. This is evidenced by Table 2, Fig. 4, and Figure S5. According to Table 2, the pyknons link each rRNA and 45S spacer to a unique set of mRNAs that are associated with different biological processes, molecular functions, pathways, or diseases (Fig. 4; Figure S5). When we group genes whose 5′-UTRs, CDSs, and 3′-UTRs contain any of the sequences in groupomni-TC, groupomni-CGG, and grouppunct, the 9 possible combinations of {groupomni-TC, groupomni-CGG, grouppunct} × {5′-UTR, CDS, 3′-UTR} correspond to genes belonging to different processes (Figure S6). Figures S5 and S6 allow several novel observations:
The exonic instances of groupomni-CGG's members are in genes involved in DNA transcription.
The exonic instances of groupomni-TC's members are in genes involved in RNA splicing (but their intronic instances are not involved in splicing, as analysis of RBP data showed).
The instances of groupomni-TC, groupomni-CGG, or grouppunct members in untranslated regions (5′-UTR, 3′-UTR) occur in genes linked to multiple genetic and brain disorders.
We stress that these findings link rRNAs and translation with transcription and splicing. They also link rRNAs and translation with human disorders.
Our study also indicates a “cooperation” of the sequences in groupomni-TC and grouppunct, with genes containing at least one sequence from each of these 2 groups belonging to the nervous system (Fig. 5). This cooperation appears to be elaborate. Indeed, when we pair groupomni-TC pyknons with the subset of grouppunct sequences found in a single rRNA/spacer, we end up with a different set of genes that are linked to different diseases/conditions each time. This suggests novel relationships. Further analysis suggests that pyknon orientation is also important. For example, genes containing the human-specific pyknon CGGGCGTGGTGGTGGG from ETS2 in sense are enriched for different biological processes, cellular components, molecular functions, and diseases than genes that contain it in antisense (Figure S7).
The recurrence of exact sense and antisense pyknon copies from rRNAs/spacers in numerous genes makes the involvement of pyknons in long-range direct coupling (Tam et al. 2008; Watanabe et al. 2008) and indirect molecular decoying (Keene 2007; Poliseno et al. 2010; Rigoutsos and Furnari 2010; Tay et al. 2011; Zhuang et al. 2013; Goodarzi et al. 2015; Rigoutsos et al. 2017), a real possibility and is supported by the available data.
Direct coupling: Psoralen–cross-linked chimeras from several cell lines contain multiple rRNA/spacer-derived sequences paired with mRNA fragments. One member of the pair contains a pyknon and the other the pyknon's reverse complement (Table S3). Other chimeras involve 2 mRNA segments from different genes coupled at sense and antisense instances of rRNA/spacer pyknons (Table S3).
Decoying: (i) In various combinations, 941 rRNA/spacer pyknons are in the binding sites of 113 RBPs reported by the ENCODE project. (ii) When we searched 39,359 public datasets from NIH's SRA repository, we found 116,915 abundant sncRNAs containing 1,964 of the 2,810 sequences in groupomni-TC, groupomni-CGG, and grouppunct. Of these pyknon-carrying sncRNAs, 97.7% are rRFs (Cherlin et al. 2020; Pliatsika et al. 2024).
To put pyknons in a disease context, we analyzed the genomic copies of rRNA/spacer pyknons and found they overlap 131 risk variants listed in the GWAS Catalog, the latter linked to several of the same disorders that emerged from our genome-only analyses (Table S5). This suggests a strong connection between genomic architecture, the transcriptome, and these disorders. Moreover, the presence of risk variants in the copies of pyknons further points to the pyknons’ functional relevance. We also analyzed RNA-seq data from ASD patients and controls (Gandal et al. 2022) and showed that only those genes whose pre-mRNAs contain at least one sequence from each of groupomni-TC and grouppunct are downregulated in the patients and only in the cortical regions that are linked to language, hearing, and vision (Figure S8). These 3 modalities are typically affected in ASD patients.
The presence of the rRNA/spacer pyknons in sncRNAs indicates consistency and suggests a purpose. However, our understanding of this purpose is currently limited. The evidence suggests a model where pyknon-containing sncRNAs use “sponging” or “decoying” to regulate the formation of RNA-RNA heteroduplexes or the binding of RBPs through shared motifs. The model is supported by the finding that the motifs are the contact points of rRNA|mRNA, spacer|mRNA, and mRNA|mRNA heteroduplexes, which we identified through psoralen cross-linking (Aw et al. 2016). Recent work by Barna and colleagues suggests that a similar mechanism is operating in human neuronal cells (Leppek et al. 2020). Similarly, sncRNAs could affect the binding of RBPs to sites of mRNAs/pre-mRNAs that contain pyknon copies, thereby influencing RNA stabilization and degradation, pre-mRNA splicing, helicase activity, localization, transport, and polyadenylation (Keene 2007; Glisovic et al. 2008; Hentze et al. 2018). These possibilities suggest another way in which the intronic and exonic copies of pyknons are functional features of pre-mRNAs and mRNAs.
The “rRNA/spacer–genome architecture–nervous system” axis is not unique to humans. We found that the same axis exists in mice, fruit flies, and worms. Across 600 million years of evolution, the axis leverages organism-specific motifs shared between the organism's rRNAs, 45S spacers, and its nervous system and developmental genes. In all 4 organisms, a specific combination of rRNA/spacer pyknons appears only in nervous system and developmental genes and is linked to brain and genetic disorders (Figures S9 to S11). The sequences underlying these linkages are organism-specific (Table 3).
A particularly intriguing aspect of our findings is that the pyknons establish a novel, sequence-based connection between rRNAs/spacers and mRNAs implicated in ribosomopathies, i.e. diseases characterized by aberrant ribosome biosynthesis or deficient ribosomal function (Farley-Barnes et al. 2019). Although defective ribosome biosynthesis and function are ubiquitous in these diseases, they appear to be deleterious in specific tissues and cell types, often leading to neurodevelopmental impairment, craniofacial deformities, bone marrow and cardiac defects, and other issues (Narla and Ebert 2010; Kulkarni et al. 2017; Farley-Barnes et al. 2019). Several ribosomopathy-linked genes contain sequences from groupomni-TC, groupomni-CGG, and grouppunct, including POLR1C and POLR1D (Treacher Collins syndrome) (Dauwerse et al. 2011), POLR1A (Weaver et al. 2015), UTP4 (Freed and Baserga 2010), DNAJC21 (Tummala et al. 2016), TRPS1, many ribosomal protein genes, and the helicases DDX51, DDX17, and DDX47 RNA that are involved in ribosome biosynthesis. Moreover, many other factors involved in mRNA translation and its regulation are present in our 3 groups, such as eIF2alpha kinases HRI, PKR, and PERK, and translation initiation and elongation factors. Mitochondrial ribosomal protein genes implicated in various cancers are also found in those lists, a notable observation given that ribosomopathies often show a propensity for cancer (Ruggero 2013; Kulkarni et al. 2017; Farley-Barnes et al. 2019). Finally, emerging evidence connects mRNA translation aberrations with many brain disorders that appear in our analyses, including autism and Rett syndrome (Hooshmandi et al. 2020; Rodrigues et al. 2020). The organism-specific linkages between rRNA/spacers, mRNAs, and ribosomopathies constitute a novel molecular link that could serve as the basis for exploring the poorly understood pathogenetic mechanisms of these conditions.
Lastly, we note the presence of the rRNA/spacer pyknons in many genes linked to substance-related disorders. As with ASD, this finding suggests a genetic basis for substance abuse (Kwako et al. 2018; Wang et al. 2019; Blum et al. 2020). Polymorphisms, duplications, micro-insertions, or micro-deletions at the sites of these pyknons, or their immediate neighborhood, may underlie both conditions.
Supplementary Material
Contributor Information
Isidore Rigoutsos, Computational Medicine Center, Thomas Jefferson University, Jefferson Alumni Hall #M81, 1020 Locust Street, Philadelphia, PA 19107, USA.
Stepan Nersisyan, Computational Medicine Center, Thomas Jefferson University, Jefferson Alumni Hall #M81, 1020 Locust Street, Philadelphia, PA 19107, USA.
Eric Londin, Computational Medicine Center, Thomas Jefferson University, Jefferson Alumni Hall #M81, 1020 Locust Street, Philadelphia, PA 19107, USA.
Iliza Nazeraj, Computational Medicine Center, Thomas Jefferson University, Jefferson Alumni Hall #M81, 1020 Locust Street, Philadelphia, PA 19107, USA.
Bonnie Dong, Sidney Kimmel Medical College, Thomas Jefferson University, 1045 Walnut Street, Philadelphia, PA 19107, USA.
Anastasios Vourekas, Department of Biological Sciences, 202 Life Sciences Building, Louisiana State University, Baton Rouge, LA 70803, USA.
Phillipe Loher, Computational Medicine Center, Thomas Jefferson University, Jefferson Alumni Hall #M81, 1020 Locust Street, Philadelphia, PA 19107, USA.
Supplementary Material
Supplementary material is available at Molecular Biology and Evolution online.
Author Contributions
I.R. designed and supervised the study. I.R., S.N., P.L., I.N., B.D., E.L., and A.V. generated and analyzed the data. I.R. wrote the paper with assistance from S.N., P.L., A.V., and E.L.
Funding
The work was supported by Thomas Jefferson University (I.R.) and the Louisiana Board of Regents RCS LEQSF (2023-26)-RD-A-16 (A.V.).
Data Availability
This study did not generate new software or new genomic sequences. The study analyzed existing, publicly available reference sequences and datasets. The reference sequences of rRNAs/spacers were obtained from GenBank, and the sequences of mRNAs from ENSEMBL (see Materials and Methods). The sequence of the T2T human genome assembly is the one reported by Nurk et al. (2022). The RepeatMasker annotations for the T2T assembly were obtained from the UCSC human genome browser. The analyzed psoralen–cross-linked data are from the NIH SRA dataset with accession number PRJNA318958. The analyzed list of polymorphisms was obtained from the NHGRI/EBI GWAS Catalog described in Sollis et al. (2023). The identifiers of the 39,359 NIH SRA sncRNA-seq datasets we analyzed are listed in the supplemental tables accompanying the manuscript. All the derivative data we generated are also included in the supplemental tables.
References
- Annear DJ et al. Abundancy of polymorphic CGG repeats in the human genome suggest a broad involvement in neurological disease. Sci Rep. 2021:11:2515. 10.1038/s41598-021-82050-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Aw JG et al. In vivo mapping of eukaryotic RNA interactomes reveals principles of higher-order organization and regulation. Mol Cell. 2016:62:603–617. 10.1016/j.molcel.2016.04.028. [DOI] [PubMed] [Google Scholar]
- Bao J et al. A UTP3-dependent nucleolar translocation pathway facilitates pre-rRNA 5′ETS processing. Nucleic Acids Res. 2024:52:9671–9694. 10.1093/nar/gkae631. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blum K et al. Biotechnical development of genetic addiction risk score (GARS) and selective evidence for inclusion of polymorphic allelic risk in substance use disorder (SUD). J Syst Integr Neurosci. 2020:6:1000221. 10.15761/JSIN.1000221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boivin M, Charlet-Berguerand N. Trinucleotide CGG repeat diseases: an expanding field of polyglycine proteins? Front Genet. 2022:13:843014. 10.3389/fgene.2022.843014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bourbon H, Michot B, Hassouna N, Feliu J, Bachellerie JP. Sequence and secondary structure of the 5′ external transcribed spacer of mouse pre-rRNA. DNA (Basel). 1988:7:181–191. 10.1089/dna.1988.7.181. [DOI] [PubMed] [Google Scholar]
- Bussi Y, Kapon R, Reich Z. Large-scale k-mer-based analysis of the informational properties of genomes, comparative genomics and taxonomy. PLoS One. 2021:16:e0258693. 10.1371/journal.pone.0258693. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen H-C, Wang J, Shyr Y, Liu Q. FindAdapt: a python package for fast and accurate adapter detection in small RNA sequencing. PLoS Comput Biol. 2024:20:e1011786. 10.1371/journal.pcbi.1011786. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen LS, Tassone F, Sahota P, Hagerman PJ. The (CGG)n repeat element within the 5′ untranslated region of the FMR1 message provides both positive and negative cis effects on in vivo translation of a downstream reporter. Hum Mol Genet. 2003:12:3067–3074. 10.1093/hmg/ddg331. [DOI] [PubMed] [Google Scholar]
- Cherlin T et al. The subcellular distribution of miRNA isoforms, tRNA-derived fragments, and rRNA-derived fragments depends on nucleotide sequence and cell type. BMC Biol. 2024:22:205. 10.1186/s12915-024-01970-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cherlin T et al. Ribosomal RNA fragmentation into short RNAs (rRFs) is modulated in a sex- and population of origin-specific manner. BMC Biol. 2020:18:38. 10.1186/s12915-020-0763-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Coleman AW. Analysis of mammalian rDNA internal transcribed spacers. PLoS One. 2013:8:e79122. 10.1371/journal.pone.0079122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Corvelo A, Hallegger M, Smith CW, Eyras E. Genome-wide association between branch point properties and alternative splicing. PLoS Comput Biol. 2010:6:e1001016. 10.1371/journal.pcbi.1001016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Crooks GE, Hon G, Chandonia JM, Brenner SE. WebLogo: a sequence logo generator. Genome Res. 2004:14:1188–1190. 10.1101/gr.849004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dauwerse JG et al. Mutations in genes encoding subunits of RNA polymerases I and III cause Treacher Collins syndrome. Nat Genet. 2011:43:20–22. 10.1038/ng.724. [DOI] [PubMed] [Google Scholar]
- Déraspe M et al. Phenetic comparison of prokaryotic genomes using k-mers. Mol Biol Evol. 2017:34:2716–2729. 10.1093/molbev/msx200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Di Ruscio A et al. DNMT1-interacting RNAs block gene-specific DNA methylation. Nature. 2013:503:371–376. 10.1038/nature12598. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dobin A et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013:29:15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dragomir MP et al. The non-coding RNome after splenectomy. J Cell Mol Med. 2019:23:7844–7858. 10.1111/jcmm.14664. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Enright CA, Maxwell ES, Eliceiri GL, Sollner-Webb B. 5′ETS rRNA processing facilitated by four small RNAs: U14, E3, U17, and U3. RNA. 1996:2:1094–1099. [PMC free article] [PubMed] [Google Scholar]
- Evangelista AF et al. Pyknon-containing transcripts are downregulated in colorectal cancer tumors, and loss of PYK44 is associated with worse patient outcome. Front Genet. 2020:11:581454. 10.3389/fgene.2020.581454. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Farley-Barnes KI, Ogawa LM, Baserga SJ. Ribosomopathies: old concepts, new controversies. Trends Genet. 2019:35:754–767. 10.1016/j.tig.2019.07.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Feng J, Naiman DQ, Cooper B. Coding DNA repeated throughout intergenic regions of the Arabidopsis thaliana genome: evolutionary footprints of RNA silencing. Mol Biosyst. 2009:5:1679–1687. 10.1039/b903031j. [DOI] [PubMed] [Google Scholar]
- Freed EF, Baserga SJ. The C-terminus of Utp4, mutated in childhood cirrhosis, is essential for ribosome biogenesis. Nucleic Acids Res. 2010:38:4798–4806. 10.1093/nar/gkq185. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fujii K, Susanto TT, Saurabh S, Barna M. Decoding the function of expansion segments in ribosomes. Mol Cell. 2018:72:1013–1020.e6. 10.1016/j.molcel.2018.11.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gandal MJ et al. Broad transcriptomic dysregulation occurs across the cerebral cortex in ASD. Nature. 2022:611:532–539. 10.1038/s41586-022-05377-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Garcia S et al. The dynamic interplay between ribosomal DNA and transposable elements: a perspective from genomics and cytogenetics. Mol Biol Evol. 2024:41:msae025. 10.1093/molbev/msae025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ge SX, Jung D, Yao R. ShinyGO: a graphical gene-set enrichment tool for animals and plants. Bioinformatics. 2020:36:2628–2629. 10.1093/bioinformatics/btz931. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Glisovic T, Bachorik JL, Yong J, Dreyfuss G. RNA-binding proteins and post-transcriptional gene regulation. FEBS Lett. 2008:582:1977–1986. 10.1016/j.febslet.2008.03.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goodarzi H et al. Endogenous tRNA-derived fragments suppress breast cancer progression via YBX1 displacement. Cell. 2015:161:790–802. 10.1016/j.cell.2015.02.053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gordon J, Pillon MC, Stanley RE. Nol9 is a spatial regulator for the human ITS2 Pre-rRNA endonuclease-kinase Complex. J Mol Biol. 2019:431:3771–3786. 10.1016/j.jmb.2019.07.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hariharan N, Ghosh S, Palakodeti D. The story of rRNA expansion segments: finding functionality amidst diversity. Wiley Interdiscip Rev RNA. 2023:14:e1732. 10.1002/wrna.1732. [DOI] [PubMed] [Google Scholar]
- Hentze MW, Castello A, Schwarzl T, Preiss T. A brave new world of RNA-binding proteins. Nat Rev Mol Cell Biol. 2018:19:327–341. 10.1038/nrm.2017.130. [DOI] [PubMed] [Google Scholar]
- Hooshmandi M, Wong C, Khoutorsky A. Dysregulation of translational control signaling in autism spectrum disorders. Cell Signal. 2020:75:109746. 10.1016/j.cellsig.2020.109746. [DOI] [PubMed] [Google Scholar]
- Ishiura H et al. Noncoding CGG repeat expansions in neuronal intranuclear inclusion disease, oculopharyngodistal myopathy and an overlapping disease. Nat Genet. 2019:51:1222–1232. 10.1038/s41588-019-0458-z. [DOI] [PubMed] [Google Scholar]
- Jha A et al. Identifying common transcriptome signatures of cancer by interpreting deep learning models. Genome Biol. 2022:23:117. 10.1186/s13059-022-02681-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Johnson PZ, Simon AE. RNAcanvas: interactive drawing and exploration of nucleic acid structures. Nucleic Acids Res. 2023:51:W501–W508. 10.1093/nar/gkad302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kaldunski ML et al. Rare disease research resources at the Rat Genome Database. Genetics. 2023:224:iyad078. 10.1093/genetics/iyad078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karlin S, Burge C. Dinucleotide relative abundance extremes: a genomic signature. Trends Genet. 1995:11:283–290. 10.1016/S0168-9525(00)89076-9. [DOI] [PubMed] [Google Scholar]
- Karlin S, Mrazek J. What drives codon choices in human genes? J Mol Biol. 1996:262:459–472. 10.1006/jmbi.1996.0528. [DOI] [PubMed] [Google Scholar]
- Keene JD. RNA regulons: coordination of post-transcriptional events. Nat Rev Genet. 2007:8:533–543. 10.1038/nrg2111. [DOI] [PubMed] [Google Scholar]
- Khaitan D et al. The melanoma-upregulated long noncoding RNA SPRY4-IT1 modulates apoptosis and invasion. Cancer Res. 2011:71:3852–3862. 10.1158/0008-5472.CAN-10-4460. [DOI] [PubMed] [Google Scholar]
- Kim JH et al. Comparative analysis and classification of highly divergent mouse rDNA units based on their intergenic spacer (IGS) variability. NAR Genom Bioinform. 2024:6:lqae070. 10.1093/nargab/lqae070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kleiderlein JJ et al. CCG repeats in cDNAs from human brain. Hum Genet. 1998:103:666–673. 10.1007/s004390050889. [DOI] [PubMed] [Google Scholar]
- Kulkarni S et al. Ribosomopathy-like properties of murine and human cancers. PLoS One. 2017:12:e0182705. 10.1371/journal.pone.0182705. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kwako LE, Bickel WK, Goldman D. Addiction biomarkers: dimensional approaches to understanding addiction. Trends Mol Med. 2018:24:121–128. 10.1016/j.molmed.2017.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lambert M, Benmoussa A, Provost P. Small non-coding RNAs derived from eukaryotic ribosomal RNA. Noncoding RNA. 2019:5:16. 10.3390/ncrna5010016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leppek K et al. Gene- and species-specific hox mRNA translation by ribosome expansion segments. Mol Cell. 2020:80:980–995.e13. 10.1016/j.molcel.2020.10.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv 3997. 10.48550/arXiv.1303.3997, 16 March 2013, preprint: not peer reviewed. [DOI]
- Loher P et al. Isomirmap—fast, deterministic, and exhaustive mining of isomiRs from short RNA-Seq datasets. Bioinformatics. 2021:37:1828–1838. 10.1093/bioinformatics/btab016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Loher P, Londin E, Ilieva H, Pasinelli P, Rigoutsos I. Re-analyses of samples from amyotrophic lateral sclerosis patients and controls identify many novel small RNAs with diagnostic and prognostic potential. Mol Neurobiol. 2025:62:8135–8149. 10.1007/s12035-025-04747-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Loher P, Telonis AG, Rigoutsos I. MINTmap: fast and exhaustive profiling of nuclear and mitochondrial tRNA fragments from short RNA-seq data. Sci Rep. 2017:7:41184. 10.1038/srep41184. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lorenz R et al. ViennaRNA package 2.0. Algorithms Mol Biol. 2011:6:26. 10.1186/1748-7188-6-26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma Y et al. Molecular mechanism of human ISG20L2 for the ITS1 cleavage in the processing of 18S precursor ribosomal RNA. Nucleic Acids Res. 2024:52:1878–1895. 10.1093/nar/gkad1210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 2011:17:10–12. 10.14806/ej.17.1.200. [DOI] [Google Scholar]
- Mauro VP, Edelman GM. rRNA-like sequences occur in diverse primary transcripts: implications for the control of gene expression. Proc Natl Acad Sci U S A. 1997:94:422–427. 10.1073/pnas.94.2.422. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McCoy MJ, Fire AZ. Parallel gene size and isoform expansion of ancient neuronal genes. Curr Biol. 2024:34:1635–1645.e3. 10.1016/j.cub.2024.02.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Metsu S et al. A CGG-repeat expansion mutation in ZNF713 causes FRA7A: association with autistic spectrum disorder in two families. Hum Mutat. 2014:35:1295–1300. 10.1002/humu.22683. [DOI] [PubMed] [Google Scholar]
- Mi G, Di Y, Emerson S, Cumbie JS, Chang JH. Length bias correction in gene ontology enrichment analysis using logistic regression. PLoS One. 2012:7:e46128. 10.1371/journal.pone.0046128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Narla A, Ebert BL. Ribosomopathies: human disorders of ribosome dysfunction. Blood. 2010:115:3196–3205. 10.1182/blood-2009-10-178129. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nurk S et al. The complete sequence of a human genome. Science. 2022:376:44–53. 10.1126/science.abj6987. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parker MS, Balasubramaniam A, Sallee FR, Parker SL. The expansion segments of 28S ribosomal RNA extensively match human messenger RNAs. Front Genet. 2018:9:66. 10.3389/fgene.2018.00066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parker MS, Sallee FR, Park EA, Parker SL. Homoiterons and expansion in ribosomal RNAs. FEBS Open Bio. 2015:5:864–876. 10.1016/j.fob.2015.10.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pearson WR, Lipman DJ. Improved tools for biological sequence comparison. Proc Natl Acad Sci U S A. 1988:85:2444–2448. 10.1073/pnas.85.8.2444. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pichler M et al. Therapeutic potential of FLANC, a novel primate-specific long non-coding RNA in colorectal cancer. Gut. 2020:69:1818–1831. 10.1136/gutjnl-2019-318903. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pillon MC, Lo YH, Stanley RE. IT's 2 for the price of 1: multifaceted ITS2 processing machines in RNA and DNA maintenance. DNA Repair (Amst). 2019:81:102653. 10.1016/j.dnarep.2019.102653. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pliatsika V et al. MINRbase: a comprehensive database of nuclear- and mitochondrial-ribosomal-RNA-derived fragments (rRFs). Nucleic Acids Res. 2024:52:D229–D238. 10.1093/nar/gkad833. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poliseno L et al. A coding-independent function of gene and pseudogene mRNAs regulates tumour biology. Nature. 2010:465:1033–1038. 10.1038/nature09144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010:26:841–842. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ramesh M, Woolford JL Jr. Eukaryote-specific rRNA expansion segments function in ribosome biogenesis. RNA. 2016:22:1153–1162. 10.1261/rna.056705.116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rigoutsos I, Furnari F. Gene-expression forum: decoy for microRNAs. Nature. 2010:465:1016–1017. 10.1038/4651016a. [DOI] [PubMed] [Google Scholar]
- Rigoutsos I et al. Short blocks from the noncoding parts of the human genome have instances within nearly all known genes and relate to biological processes. Proc Natl Acad Sci U S A. 2006:103:6605–6610. 10.1073/pnas.0601688103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rigoutsos I et al. N-BLR, a primate-specific non-coding transcript leads to colorectal cancer invasion and migration. Genome Biol. 2017:18:98. 10.1186/s13059-017-1224-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Robine N et al. A broadly conserved pathway generates 3′UTR-directed primary piRNAs. Curr Biol. 2009:19:2066–2076. 10.1016/j.cub.2009.11.064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rodrigues DC et al. Shifts in ribosome engagement impact key gene sets in neurodevelopment and ubiquitination in Rett syndrome. Cell Rep. 2020:30:4179–4196.e11. 10.1016/j.celrep.2020.02.107. [DOI] [PubMed] [Google Scholar]
- Ruggero D. Translational control in cancer etiology. Cold Spring Harb Perspect Biol. 2013:5:a012336. 10.1101/cshperspect.a012336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sandberg R et al. Capturing whole-genome characteristics in short sequences using a naïve Bayesian classifier. Genome Res. 2001:11:1404–1409. 10.1101/gr.186401. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schriml LM et al. The human disease ontology 2022 update. Nucleic Acids Res. 2022:50:D1255–D1261. 10.1093/nar/gkab1063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharp PM, Li WH. The codon adaptation index–a measure of directional synonymous codon usage bias, and its potential applications. Nucleic Acids Res. 1987:15:1281–1295. 10.1093/nar/15.3.1281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sievers F, Higgins DG. The Clustal Omega multiple alignment package. Methods Mol Biol. 2021:2231:3–16. 10.1007/978-1-0716-1036-7_1. [DOI] [PubMed] [Google Scholar]
- Sloan KE, Knox AA, Wells GR, Schneider C, Watkins NJ. Interactions and activities of factors involved in the late stages of human 18S rRNA maturation. RNA Biol. 2019:16:196–210. 10.1080/15476286.2018.1564467. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sollis E et al. The NHGRI-EBI GWAS catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023:51:D977–D985. 10.1093/nar/gkac1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sueoka N. On the genetic basis of variation and heterogeneity of DNA base composition. Proc Natl Acad Sci U S A. 1962:48:582–592. 10.1073/pnas.48.4.582. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sweeney BA et al. R2DT is a framework for predicting and visualising RNA secondary structure using templates. Nat Commun. 2021:12:3494. 10.1038/s41467-021-23555-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Takahashi M, Kryukov K, Saitou N. Estimation of bacterial species phylogeny through oligonucleotide frequency distances. Genomics. 2009:93:525–533. 10.1016/j.ygeno.2009.01.009. [DOI] [PubMed] [Google Scholar]
- Tam OH et al. Pseudogene-derived small interfering RNAs regulate gene expression in mouse oocytes. Nature. 2008:453:534–538. 10.1038/nature06904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tay Y et al. Coding-independent regulation of the tumor suppressor PTEN by competing endogenous mRNAs. Cell. 2011:147:344–357. 10.1016/j.cell.2011.09.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Telonis AG, Rigoutsos I. The transcriptional trajectories of pluripotency and differentiation comprise genes with antithetical architecture and repetitive-element content. BMC Biol. 2021:19:60. 10.1186/s12915-020-00928-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tranque P, Hu MC-Y, Edelman GM, Mauro VP. rRNA complementarity within mRNAs: a possible basis for mRNA-ribosome interactions and translational control. Proc Natl Acad Sci U S A. 1998:95:12238–12243. 10.1073/pnas.95.21.12238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tsirigos A, Rigoutsos I. A new computational method for the detection of horizontal gene transfer events. Nucleic Acids Res. 2005a:33:922–933. 10.1093/nar/gki187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tsirigos A, Rigoutsos I. A sensitive, support-vector-machine method for the detection of horizontal gene transfers in viral, archaeal and bacterial genomes. Nucleic Acids Res. 2005b:33:3699–3707. 10.1093/nar/gki660. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tsirigos A, Rigoutsos I. Human and mouse introns are linked to the same processes and functions through each genome's most frequent non-conserved motifs. Nucleic Acids Res. 2008:36:3484–3493. 10.1093/nar/gkn155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tsuji J, Weng Z. DNApi: a de novo adapter prediction algorithm for small RNA sequencing data. PLoS One. 2016:11:e0164228. 10.1371/journal.pone.0164228. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tummala H et al. DNAJC21 mutations link a cancer-prone bone marrow failure syndrome to corruption in 60S ribosome subunit maturation. Am J Hum Genet. 2016:99:115–124. 10.1016/j.ajhg.2016.05.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vacher MC et al. Alzheimer's disease-like neuropathology in three species of oceanic dolphin. Eur J Neurosci. 2023:57:1161–1179. 10.1111/ejn.15900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Valdar WS. Scoring residue conservation. Proteins. 2002:48:227–241. 10.1002/prot.10146. [DOI] [PubMed] [Google Scholar]
- Van Nostrand EL et al. A large-scale binding and functional map of human RNA-binding proteins. Nature. 2020:583:711–719. 10.1038/s41586-020-2077-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Voith von Voithenberg L et al. Recognition of the 3′ splice site RNA by the U2AF heterodimer involves a dynamic population shift. Proc Natl Acad Sci U S A. 2016:113:E7169–E7175. 10.1073/pnas.1605873113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang SC, Chen YC, Lee CH, Cheng CM. Opioid addiction, genetic susceptibility, and medical treatments: a review. Int J Mol Sci. 2019:20:4294. 10.3390/ijms20174294. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Watanabe T et al. Endogenous siRNAs from naturally formed dsRNAs regulate transcripts in mouse oocytes. Nature. 2008:453:539–543. 10.1038/nature06908. [DOI] [PubMed] [Google Scholar]
- Weaver KN et al. Acrofacial dysostosis, cincinnati type, a mandibulofacial dysostosis syndrome with limb anomalies, is caused by POLR1A dysfunction. Am J Hum Genet. 2015:96:765–774. 10.1016/j.ajhg.2015.03.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wingett SW, Andrews S. Fastq screen: a tool for multi-genome mapping and quality control. F1000Res. 2018:7:1338. 10.12688/f1000research.15931.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Woese CR, Fox GE. Phylogenetic structure of the prokaryotic domain: the primary kingdoms. Proc Natl Acad Sci U S A. 1977:74:5088–5090. 10.1073/pnas.74.11.5088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ye R et al. Capture RIC-seq reveals positional rules of PTBP1-associated RNA loops in splicing regulation. Mol Cell. 2023:83:1311–1327.e7. 10.1016/j.molcel.2023.03.001. [DOI] [PubMed] [Google Scholar]
- Youn YH, Byun HJ, Yoon JH, Park CH, Lee SK. Long noncoding RNA N-BLR upregulates the migration and invasion of gastric adenocarcinoma. Gut Liver. 2019:13:421–429. 10.5009/gnl18408. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhuang R et al. miR-195 competes with HuR to modulate stim1 mRNA stability and regulate cell migration. Nucleic Acids Res. 2013:41:7905–7919. 10.1093/nar/gkt565. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
This study did not generate new software or new genomic sequences. The study analyzed existing, publicly available reference sequences and datasets. The reference sequences of rRNAs/spacers were obtained from GenBank, and the sequences of mRNAs from ENSEMBL (see Materials and Methods). The sequence of the T2T human genome assembly is the one reported by Nurk et al. (2022). The RepeatMasker annotations for the T2T assembly were obtained from the UCSC human genome browser. The analyzed psoralen–cross-linked data are from the NIH SRA dataset with accession number PRJNA318958. The analyzed list of polymorphisms was obtained from the NHGRI/EBI GWAS Catalog described in Sollis et al. (2023). The identifiers of the 39,359 NIH SRA sncRNA-seq datasets we analyzed are listed in the supplemental tables accompanying the manuscript. All the derivative data we generated are also included in the supplemental tables.

![[ALT TEXT: Graphs of the secondary structures of 18S and 28S rRNAs and the ITS1 and ETS2 spacers. Portions of the structures are colored to denote coverage by the pyknons in groupomni-TC, groupomni-CGG, and grouppunct, respectively.]](https://cdn.ncbi.nlm.nih.gov/pmc/blobs/4120/12515603/2e94b408a4ab/msaf237f1.jpg)
![[ALT TEXT: Graph showing a multiple sequence alignment of the 28S rRNA sequences from twenty organisms with increasing evolutionary distance from humans. The locations of the (TC)8 and G(CGG)5 pyknons are highlighted.]](https://cdn.ncbi.nlm.nih.gov/pmc/blobs/4120/12515603/3db01bd5422b/msaf237f2.jpg)
![[ALT TEXT: Graph showing the conservation of groupomni-TC, groupomni-CGG, and grouppunct instances across organisms. Part A shows the multiple sequence alignments for four genes at the location of a select pyknon. Part B shows summary boxplots indicating diminishing conservation with evolutionary distance at the sites of pyknons in untranslated regions and coding sequences. Part C shows summary boxplots indicating diminishing conservation with evolutionary distance at the sites of pyknons in introns.]](https://cdn.ncbi.nlm.nih.gov/pmc/blobs/4120/12515603/19de38e8d71a/msaf237f3.jpg)
![[ALT TEXT: Graph showing term enrichments for genes containing rRNA and 45S spacer pyknons in their untranslated regions and coding sequences, in sense or antisense.]](https://cdn.ncbi.nlm.nih.gov/pmc/blobs/4120/12515603/b01f1e6e7ff4/msaf237f4.jpg)
![[ALT TEXT: Graph showing a Venn diagram for the genes whose pre-mRNAs contain pyknons from the groupomni-TC, groupomni-CGG, or grouppunct, respectively. Also shown are term enrichments for some of the seven regions of the Venn diagram.]](https://cdn.ncbi.nlm.nih.gov/pmc/blobs/4120/12515603/7c19bab1525a/msaf237f5.jpg)