Abstract
The Rfam database, a widely used repository of non-coding RNA families, has undergone significant updates in release 15.0. This paper introduces major improvements, including the expansion of Rfamseq to 26 106 genomes, a 76% increase, incorporating the latest UniProt reference proteomes and additional viral genomes. Sixty-five RNA families were enhanced using experimentally determined 3D structures, improving the accuracy of consensus secondary structures and annotations. R-scape covariation analysis was used to refine structural predictions in 26 families. Gene Ontology (GO) and Sequence Ontology annotations were comprehensively updated, increasing GO term coverage to 75% of families. The release adds 14 new Hepatitis C Virus RNA families and completes microRNA family synchronization with miRBase, resulting in 1603 microRNA families. New data types, including FULL alignments, have been implemented. Integration with APICURON for improved curator attribution and multiple website enhancements further improve user experience. These updates significantly expand Rfam’s coverage and improve annotation quality, reinforcing its critical role in RNA research, genome annotation and the development of machine learning models. Rfam is freely available at https://rfam.org.
Graphical Abstract
Graphical Abstract.
Introduction
The Rfam database was established in 2002 (1) in order to provide a central repository of non-coding RNA (ncRNA) families for genomic annotations. Each family is represented by three key components: (i) a multiple sequence alignment called the SEED alignment that contains aligned sequences of homologous RNA sequences that share a consensus secondary structure, (ii) a covariance model (CM) built using the Infernal software (2) that was trained on the SEED alignment and (iii) a set of matches that were found using the CM, called the set of FULL hits. The families are built manually by Rfam curators using scientific literature or alignments submitted by the community.
Rfam curation is done in the context of a specific sequence database, called Rfamseq. This database is meant to provide curators with a comprehensive collection of whole genomes. It should attempt to represent the known phylogenetic diversity, while ensuring the many searches required for curation remain feasible. All FULL sequences are computed by searching these sequences. For each model the curators search for homologs in Rfamseq, and select a bit score threshold, called the gathering threshold, that separates homologous sequences from unrelated or more distantly related sequences, taking into account phylogenetic distribution of hits. The cutoff can be used by Infernal to report only sequences that score higher than this threshold when annotating genomes or searching in sequence databases.
Each major release of the Rfam database, such as 14.0 or 15.0, corresponds to a new version of Rfamseq. Prior to Rfam 13.0, Rfamseq was based on a set of sequences from the ENA database (3). However, due to the explosive growth in ENA, in Rfam 13.0 Rfamseq transitioned to a representative and reduced redundancy set of genomes produced by the UniProt team (4) for the Reference Proteomes dataset (5). The representative proteomes from UniProt are mapped to their corresponding genomes and the genomic sequence is analysed by Rfam (5). This provides users with a representative collection of genomes which have comprehensive ncRNA and protein annotations.
Rfam has been widely adopted by various scientific communities, with uses ranging from the annotation of small ncRNAs in genomic resources like Ensembl (6), to serving as a dataset for training machine learning models such as AlphaFold 3 (7). Additionally, Rfam is frequently used as a key reference for known ncRNAs. To ensure it remains current and valuable to the scientific community, we continuously review, update and enhance the Rfam families.
In release 15.0, we have updated and expanded Rfamseq, improved families using 3D structures and R-scape (8), improved Gene Ontology (GO) (9) and Sequence Ontology (SO) (10) annotations and completed the synchronization of miRBase (11) microRNAs into Rfam.
New Rfamseq database
Rfam 15 is based on the 2024_03 release of UniProt reference proteomes that includes 23 158 genomes (see Figure 1 showing the taxonomic distribution). We attempted to download all components of the genomes used to build the proteome set. In some cases, the specified genome was replaced by a newer version, which was used instead. In a very few cases, <50, all versions of the genome were either outdated or removed. In those cases, we did not fetch any genome.
Figure 1.
The taxonomic distribution of cellular organisms in Rfamseq 15 organized by the phylogenetic kingdom. Shown on the outside is the number of reference and non-reference proteomes in each lineage. Each coloured section indicates the kingdom with bacteria (green), archaea (pink) and eukaryotes (blue).
In addition to the genomes in the UniProt set, we also fetched all viral genomes from the protein information resource (12) at 75% per cent protein sequence identity for additional 2985 viral genomes. This set was added because during our viral work, described below, we found that the UniProt reference proteome work underreports the diversity of viral genomes. This has led to some viral Rfam families having very few or no FULL hits. Rfam would prefer that all families have at least one match outside of the SEED sequences. While using an additional set of viral genomes leads to a large increase in the number of genomes, the total size of Rfamseq is dominated by cellular organisms as <1% of the nucleotides come from a viral genome. Table 1 shows the comparison of Rfamseq in the last major release Rfam 14.0 versus Rfam 15.0.
Table 1.
The number of genomes in Rfamseq corresponding to release 14.0 and 15.0
| Kingdom | Rfamseq v14 | Rfamseq v15 |
|---|---|---|
| Total sequence length | 3.68e11 | 1.46e12 |
| Number of species | 14 774 | 26 106 |
| Eukaryotes | 1057 | 2613 |
| Archaea | 418 | 370 |
| Bacteria | 7808 | 9601 |
| Viruses | 5491 | 13 552 |
| Number of Rfam hits | 2897 296 | 10 736 534 |
| Size on disc | 170 GB | 772 GB |
The size on disc is the uncompressed size of all sequences in Rfamseq.
Once Rfamseq was updated, all families were searched against the updated database and the matches were collected. Many older or larger families have not been rescanned in years due to technical pipeline limitations, e.g. the 5S ribosomal RNA (rRNA; RF00001) has not had its matches updated since 2014.
Once all families were updated, we compared the matches of families in 14.3 (which was the last published version) and 15.0 to determine how Rfam has changed. Of the 3431 families in common between 14.3 and 15.0, 98 had no FULL hits in 14.3 and when moving to 15.0, 23 of these families gained at least one hit while 21 lost all hits, leading to 96 families in 15.0 without FULL hits. Families without FULL hits occur because either the curator-selected gathering threshold of the family is too high, or if Rfamseq does not contain any similar sequences. Future work will look to see if these families require updating the threshold or if additional genomes need to be added to Rfamseq.
Of the 3302 families with hits in both releases, the average family grew by 166%, with 2335 gaining at least one hit, 547 losing hits and 420 with no change in the number of hits. Examining the families with an increase shows that there are 26 families with a >10× growth. These families are primarily microRNA families matching to plant genomes. While, as described below, these families have been synchronized with miRBase, we are revisiting these families to see if they need additional future updates in the light of these additional hits, or if the genomes are poorly assembled and should be excluded from Rfamseq.
Overall the update to Rfamseq has led to a considerable change in all Rfam families. This will provide the community with a consistent and up-to-date dataset to reuse. Future updates to Rfamseq will seek to maximize the coverage of known organisms.
Improving existing Rfam families
Using RNA 3D structures to revise Rfam secondary structures
Rfam families are based on multiple sequence alignments annotated with consensus secondary structures that indicate base pairing. The accuracy and completeness of the consensus secondary structure is of critical importance as it guides the SEED alignment, informs the Infernal CM and is used for training and benchmarking of software for RNA 2D and 3D structure prediction, including AlphaFold 3 (7) and R-scape (8).
Most of the Rfam families with known 3D structure were created prior to the 3D structure determination using less accurate, predicted secondary structures. For example, while the FMN riboswitch family (RF00050) correctly captured five helices of the FMN riboswitch, it did not include one helix, two pseudo-knots (PKs) and several key base pairings (Figure 2B). As such, one major goal of this project was to include the PKs which are observed in 3D structures into the Rfam families.
Figure 2.
(A) Pipeline for reviewing Rfam families using experimentally determined RNA 3D structures from PDB. (B) Example of the FMN riboswitch (RF00050) before and after improvements with a 3D structure. Improvements included adding two PKs and creating missing helices. One of the newly added PK basepairs has support from R-scape, while the other does not. (C) A summary of how the IS605-orfB-I RNA (RF03065) family was reviewed and improved using R-scape.
To aid the Rfam curators with updating Rfam alignments, an automated pipeline was developed to map RNA sequences and secondary structures of experimentally determined RNA 3D structures to Rfam SEED alignments. Every week Rfam families are mapped to the latest set of RNA chains from the Protein Data Bank (PDB) (13) using the Infernal cmscan program. Rfam uses the experimental sequence not the sequence observed in the 3D structure of each chain to avoid issues with unresolved residues. For each Rfam family with one or more matching structures, the sequences and secondary structures are iteratively added to the alignment in the Stockholm format using Infernal’s cmalign program. The base pairing information, as determined by FR3D (14), is also included in the alignment as an additional per-residue (GR) annotation line (one per 3D structure). The pipeline is automatically executed weekly and the Rfam curators receive notifications of newly mapped structures (Figure 2A).
While the automatic cmalign-based approach works well for many families and structures, there are two failure modes we have observed. First, not all structures match a family. To deal with this, we periodically review all structures without a family and manually assign a family if one exists. The structure then undergoes the standard automated alignment and review process. Second, the automatically generated alignments are not always correct. In some cases, <20 so far, the automatically generated alignment produces a worse alignment. This is often because the sequences used in the experiments are edited for crystallization. In these cases, the curator will undo the changes to the alignment and may manually align the sequence into the alignment. This alignment is then reviewed. Additionally, curators periodically review all RNA chains without a matching Rfam family to determine if the structures ‘should’ match a family.
The resulting alignments undergo manual review to determine whether the family consensus secondary structure is consistent with the basepair information from PDB structures. The structure and annotated basepairs are always kept in the alignment, even in cases where the structure is not used to update the consensus secondary structure, e.g. low-resolution structure or non-canonical basepairs. Additional annotations are also included in the Stockholm files, such as RNA structural elements. For example P1, P2 and P3 domains are included in the Pistol ribozyme RF02679, while the SAM riboswitch (RF00162) now includes a kink-turn annotation.
The manual review process aims to update the alignment, consensus sequence and consensus secondary structure to reflect the additional information from 3D structures. In summary, we validate that the 3D structure is correctly aligned, the metadata about the structure is correct and the structure is of sufficient resolution (<4.0 Å) to annotate basepairs. We then use the annotated basepairs to update the consensus secondary structure to reflect any additional basepairs, including PKs. A detailed description of the process is available at https://docs.rfam.org/en/latest/adding-3d-structures.html.
Our pipeline detects 143 Rfam families which match at least one 3D structure. We began updating families with 3D structures in release 14.5 and since then Rfam has updated 65 families, with 298 chains from 3D structures. This includes well-known families such as the SAM riboswitch (RF00162), the FMN riboswitch (RF00050) and microRNA 16 precursor (RF00254). The updates included the addition of PKs in 42 of the 65 families, of which six had two PK, and one had three. These features include ligand binding sites for riboswitches (28 annotations), locations of well-known motifs such as kink-turns (eight annotations) and structural regions such as helices and domains (60 annotations). A detailed table showing which families were updated with which 3D structures is available and can also be found online at https://rfam.org/3d. Rfam curators work to keep families as up-to-date with the latest structures as possible, which has led to 11 families being updated more than once since release 14.5.
The review of structures is part of the ongoing curation process, and additional families will be improved in future releases. The pipeline is implemented in Python and is available at https://github.com/Rfam/rfam-3d-seed-alignments and the weekly updates are found in the file pdb_full_region.txt.gz located in the preview section of the FTP archive (http://ftp.ebi.ac.uk/pub/databases/Rfam/.preview/).
Using R-scape to improve Rfam consensus secondary structures
Since version 13.0 (5), the Rfam website included the results of the R-scape analysis (8) to help users evaluate consensus secondary structures for each Rfam family, as well as alternative secondary structures generated using the R-scape CaCoFold algorithm. CaCoFold evaluates all possible consensus secondary structures that are consistent with multiple sequence alignment and can propose a structure maximizing the number of basepairs with statistically significant covariation (15).
In release 14.9, the R-scape software was used to systematically review all Rfam families and identify those that could be enhanced using R-scape CaCoFold structures. We examined all families with an increased number of covarying basepairs and selected 26 families for updates. These families are listed in Table 2. Shown in Figure 2C is an example of a family, IS605-orfB-I, before and after improvement.
Table 2.
Summary of the R-scape-based improvements
| Family | Additional covarying basepairs |
|---|---|
| RF02033 (HEARO) | 24 |
| RF03065 (IS605-orfB-I) | 14 |
| RF03068 (RT-3) | 8 |
| RF03072 (raiA) | 5 |
| RF02969 (DUF3800-I) | 4 |
| RF01688 (Actino-pnp) | 3 |
| RF02004 (group-II-D1D4-5) | 3 |
| RF02005 (group-II-D1D4-6) | 3 |
| RF02913 (pemK) | 3 |
| RF03077 (RT-2) | 3 |
| RF03135 (L4-Archaeoglobi) | 3 |
| RF03144 (eL15-Euryarchaeota) | 3 |
| RF00062 (HgcC) | 2 |
| RF01731 (TwoAYGGAY) | 2 |
| RF01794 (sok) | 2 |
| RF02221 (sRNA-Xcc1) | 2 |
| RF02947 (cow-rumen-2) | 2 |
| RF03000 (LOOT) | 2 |
| RF03158 (L31-Actinobacteria) | 2 |
| RF01864 (plasmodium_snoR21) | 1 |
| RF01867 (CC2171) | 1 |
| RF02944 (c4-2) | 1 |
| RF02968 (DUF3800-IX) | 1 |
| RF02987 (GA-cis) | 1 |
| RF03019 (RT-16) | 1 |
| RF03046 (Pseudo-monadales-1) | 1 |
This shows the families which were improved by using R-scape CaCoFold structures.
GO and SO annotations
Rfam is commonly used as a source of functional information for ncRNAs. This takes several forms, from users reading our curated descriptions and Wikipedia articles to understand the role of an ncRNA, to using the GO and SO annotations that Rfam curators provide to understand the role of an ncRNA. Rfam is the largest source of GO annotations with >10 million sequences having an Rfam-based GO annotation. In addition to providing information to human users, Rfam is also used in training large language models (LLMs). LLMs such as ChatGPT (16) and Claude (https://www.anthropic.com/news/claude-3-5-sonnet) have clearly been trained on Rfam data and are able to output entries in the formats, Stockholm and DESC, that Rfam uses; thus, providing a completely new way for scientists to access the information found in Rfam.
The GO is a resource to classify and provide functional information for biomolecules using structured annotations and ontologies (9). Similarly the SO provides an ontology and structured annotations for the types of biomolecules (10). Both resources are continuously updated to better reflect scientific knowledge. For example, since the last publication of Rfam, the SO gained several terms specific to the location of rRNA, e.g. cytosolic_rRNA (SO:0002343) and obsoleted the more generic rRNA terms.
Rfam provides GO and SO annotations for families; however, these annotations are not regularly reviewed to stay up-to-date with the latest changes in the ontologies. As an example, prior to 15.0, the rRNA families used obsoleted SO terms. For 15.0, we manually reviewed all annotations and updated them to better reflect changes in the GO and SO, as well as improved the specificity of annotations where possible.
We ensured that all families had at least one up-to-date SO term and the annotation was as specific as possible. For GO terms, we ensured that the terms were up-to-date and added related terms where possible. A few examples of the changes we made were to (i) ensure all small nucleolar RNAs (snoRNAs) had a snoRNA specific SO term and included an RNA processing GO term; (ii) add the sodium ion binding term to the NA+ riboswitch; and (iii) updated families with the generic mature_transcript SO term to the more specific ncRNA SO term. Of the 3431 common families between 14.3 and 15.0, 2084 had a GO term in 14.3 while 2426 did in 15.0. For all families in 15.0, we now have at least one GO term for 3157, which covers 75% of all families, an increase from 60% in 14.3. These changes increased the total number of GO terms from 3752 to 4446. As these updates propagate through the community, we expect these changes to lead to millions more sequences annotated with functional information.
Viral RNAs
Starting with Rfam 14.3, we collaborated with the Marz group to develop a new workflow for viral RNA families (17). This workflow is based upon whole genome alignments of viruses and led to the creation of Flaviviridae and Coronaviridae families as described previously (17). Since the last publication, we have continued to use this workflow and have added Hepatitis C Virus (HCV) (18) families. We included 14 new families and removed four outdated families. A schematic of the new families can be seen in Figure 3. These families include structures found in both the non-coding and coding regions of the viral alignments. All viral families organized by viral clade can be seen at https://rfam.org/viruses. We plan to continue to use the viral pipeline and improve the representation of viral families in Rfam and expect to focus on Hepatitis Delta Virus (HDV) and pestivirus families in the near future.
Figure 3.
A schematic of the HCV viral families Rfam now contains. These families were built from whole genome alignments, as described in (18). The figure shows their relative locations in the HCV genomic alignment and shown in the inset are two of the new HCV families.
MicroRNAs
MicroRNAs are an important class of ncRNA that regulate gene expression in animals and plants, with many microRNAs being implicated in disease. For example, the mir-17∼92 cluster (19) and mir-155 (20) are amongst a number of microRNAs which act as oncogenes. Understanding the evolutionary and family relationships of microRNAs across species allows the transfer of annotation, e.g. from model organisms to humans and vice versa.
Since early releases, Rfam has included microRNA families, e.g. let-7 (21) was deposited in 2002 for Rfam 1.1. However, they were not subjected to regular review and were not coordinated with miRBase (11), an authoritative resource for microRNA annotation that assigns identifiers for microRNA genes and sequences. As of Rfam 13.0, out of 1983 microRNA families found in miRBase v21, only 28% matched one or more of the 529 Rfam microRNA families.
Starting in Rfam 14, we began systematically reviewing microRNA families with the goal to synchronize Rfam microRNA families with miRBase. To this end, we worked to ensure that the SEED alignments of all microRNA families in Rfam were built from microRNA sequences that are tracked in the miRBase database. Multiple sequence alignments of microRNA sequences were extracted from miRBase and mapped to Rfam accessions. Each sequence was assigned a RNAcentral Unique RNA Sequence (22) identifier to remove sequence redundancy and represent only distinct sequences for each species. For each alignment, a CM was built using Infernal and used to search the Rfamseq database. Bit score thresholds for each model [known as gathering thresholds (17)] were manually curated. These thresholds enable automatic and accurate genome annotation using Rfam microRNA families with Infernal.
Pre-existing Rfam microRNA families were reviewed and replaced with miRBase-derived SEED alignments where possible. New families were created when Rfam did not already represent the microRNA sequences. MicroRNA families found in Rfam that did not match miRBase were reviewed and updated or deleted. In cases where the miRBase alignment matched several Rfam families, we manually inspected the alignment and families and split and merged families and built clans of families as appropriate.
Rfam release 15.0 marks the completion of the synchronization process, as we have processed all microRNA alignments from miRBase. For roughly 200 miRBase-derived alignments, we determined that the family was not suitable for inclusion into Rfam. In most cases this was because there was a single unique sequence in the miRBase alignment. A small fraction of the alignments which have not been added may be removed from miRBase in the future. Table 3 shows the summary of changes since release 14.3, and a list of all microRNA families is found in Supplementary Table S1. All miRBase microRNA sequences that are not represented in updated Rfam families will be periodically and systematically reviewed in a cycle of improvements and curation of both Rfam and miRBase. Rfam and miRBase will continue to synchronize our resources as new microRNAs and new microRNA alignments are available, with miRBase acting as the repository of microRNA sequences, and Rfam as the authority on the grouping of those sequences into families. Rfam microRNA family classifications will be available through both resources.
Table 3.
A summary of the changes in Rfam microRNA families since last publication (17) with release 14.3
| Changes for Rfam 15.0 | |
|---|---|
| Total number of microRNA families | 1603 |
| New families | 722 |
| Updated families | 881 |
| Deleted families | 8 |
New families
While the primary focus of Rfam has been to complete the microRNA project, and improve existing families with 3D structures, we have continued to create new families. Since Rfam 14.3 (17), Rfam has created 16 new non-microRNA or viral RNA families. The families cover a range of phylogenetic and functional types. A few examples include Bacteroidales small SRP (RF04183), the signal recognition particle RNA of Bacteroidetes (23), Hairpin-meta1, a virus-like ribozyme reported in RNA satellites of plant viruses (RF04190) (24), Hovlinc ribozyme (RF04188) (25), a newly discovered type of self-cleaving ribozymes found in human and other hominids, RF04222 PLRV xrRNA, a exoribonuclease-resistant RNA detected in Flavivirus Potato virus (26) and RF04247 bZIP family which is a non-canonical Hac1/Xbp1 intron found in Metazoa (27). Shown in Figure 4 are the secondary structures of selected families.
Figure 4.
A selection of the new families for Rfam 15.0. The five families, namely Hovlinc ribozyme (RF04188), the PLRV xrRNA (RF04222), the Hairpin-meta ribozyme (RF04190), nqrA-II ncRNA motif (RF04310) and Bacteriodales small SRP (RF04183) are several of the new families created for Rfam 15.0.
These families are curated through a combination of literature reviews and community submissions. We actively encourage groups with candidate RNAs to contribute their alignments and secondary structure data to Rfam. However, it is important to note that Rfam requires all sequences in SEED alignments to be traceable to public databases such as GenBank (28) or RNAcentral (22). With the increasing ease of sequencing, many laboratories maintain private sequence databases for constructing alignments. As a result, many of these new alignments cannot be incorporated into Rfam due to the lack of traceable public records. We strongly recommend that the scientific community submit their sequences to publicly accessible databases, ensuring they are available for reuse and analysis by the broader research community. Furthermore, when publishing alignments, using public accessions will facilitate faster and more efficient integration into resources like Rfam.
Other improvements
Updating FULL alignments
Rfam provides several data types for interested users to download from our FTP site. These include the SEED alignments, CMs and FULL sequences. Older versions of Rfam included FULL alignments, which are an alignment of all matching sequences to the CM. These alignments are all sequences which match above the family’s gathering threshold. While these sequences are matches to our models, it is possible that the sequences are pseudo-genes when viewed in the context of the genome. Previously, these were no longer produced because of technical limitations; however, with improvements to Infernal and growth in available computing power, it is now possible to create these alignments again. We now produce full alignments for all Rfam families. These are available for each model in the ‘full_alignments’ section on the FTP site (https://ftp.ebi.ac.uk/pub/databases/Rfam/CURRENT/full_alignments).
APICURON integration
APICURON (https://apicuron.org/) is a database to track and credit biocurators for the work they do (29). The database allows for databases to credit curators for each activity they perform and provides overall statistics for each database including Rfam. In Rfam, it is common for a curator to perform many small updates and fixes to families to produce a release. Currently, this work goes unacknowledged except during publication, which is generally separated by several years. By integrating with APICURON, we are able to display all the changes curators perform as part of their duties on a more regular schedule. Additionally, APICURON updates ORCID records (https://orcid.org/) to help credit curators for their activities. We update our APICURON records with each release and our APICURON page is available at https://apicuron.org/databases/rfam. APICURON requires that each resource tracks the activities of each curator, unfortunately some historical Rfam activity does not have the required specificity, leading to some older curators lacking attributions. This is being worked on and we hope to properly credit all curators for their contributions to Rfam.
Website improvements
The Rfam website (http://rfam.org) has undergone continuous development since our last publication. We have added several new features, such as dedicated landing pages for each project, e.g. viral families can easily be browsed at http://rfam.org/viruses. These pages make it simpler for users interested in a particular project, e.g. the 3D structure alignments, to see our progress and fetch all data related to the project.
We have recently updated some terminology used for riboswitch families in Rfam. We have found that some users are confused that Rfam riboswitch families generally include just the aptamer domains. This is because aptamers align well, while it is difficult to build an alignment of a whole riboswitch. We have added a short note to the header of each Rfam riboswitch page to indicate that this family is only the aptamer domain and a short explanation of what that means. Additionally, the descriptions of families now include the term ‘aptamer’ to clarify this point. We have not updated the identifiers or accessions of any families. Current Rfam families are aptamer domains, thus our decision to clarify the descriptions. However, future families may include expression domains and when this occurs we will note it as such.
Finally, we have integrated the RNAcentral LitScan widget from RNAcentral. RNAcentral has developed a pipeline, LitScan, which identifies and extracts mentions of any ncRNA in all open access literature (30). The result of this pipeline can be reused by embedding a simple HTML widget. The widget allows users to browse and search the related literature in a convenient way. We have embedded this widget on all Rfam family pages under the new ‘Publications’ tab. The results are updated with each RNAcentral release, roughly every 4 months. An example of the widget for the Glutamine riboswitch family page is shown in Figure 5.
Figure 5.
The LitScan widget embedded in the Glutamine riboswitch family page. This widget shows users which open access articles mention the family. Families are searched using their Rfam ID and Rfam accession in all open access literature. It allows users to search the matched articles by several types of metadata such as the article type (research, review, report, etc.), the section of the paper which mentions the family and organisms mentioned in the paper. Additionally, users can perform a text search within the sentences which mention the Rfam family.
Conclusions
After >20 years of work (1), Rfam has grown to >4000 families and continues to be a key resource in RNA science. We continue to provide a large centralized collection of ncRNA families, which are used in many ways. Rfam was originally created to annotate genomes, but we have grown far beyond that use case. We would like to take this opportunity to thank everyone who has worked on or collaborated with Rfam over the past 20 years. Anyone interested in learning about the history and changes in Rfam since its inception is encouraged to visit https://rfam.org/rfam20 to find interviews with many former and current staff and users.
Since our last publication, we have focused on completing our synchronization with miRBase and improving existing families using R-scape and 3D structures. This has led to a large increase in the number and quality of families. Rfam plans to stay synchronized with miRBase and complete the 3D structure improvements by connecting each family to at least one structure where possible. Outside of those projects, we will continue to import viral RNAs and return to capturing more novel families discovered by the community. Finally, we will explore ways to expand Rfamseq to better improve our coverage of the phylogenetic space.
It is essential to note that there has been a major shift in bioinformatics with the publication of AlphaFold and its effect on protein structure prediction. We expect that as the field of RNA structure prediction matures, Rfam will continue to play a key role as a source of ground truth. In order to best serve this new use case, we will begin exploring ways to grow faster than before, while maintaining the high standards the community has come to expect of Rfam alignments. This will be essential to providing the test and training sets required for deep learning-based approaches.
In addition to our ongoing projects, Rfam plans to generate a large collection of automatically generated candidate families, similar to Pfam’s Pfam-B. This will be a separate collection of alignments and serve as a source of new families for Rfam. We believe the RNA 3D structure community will benefit from this collection of alignments and this will speed up the curation process. Second, we plan to create a comprehensive genome annotation pipeline which will include pseudo-gene detection. Rfam is used extensively in genome annotation, but we do not provide a simple comprehensive method of ncRNA genome annotation. Notably, Rfam models may match pseudo-genes creating incorrect annotations. This is a long requested tool we plan to provide to the community.
Finally, we encourage any interested community members to reach out for collaboration or to provide data. As discussed here, much of our major work is carried out with interested community members and can lead to new directions for the resource. We invite new data submissions and feedback at https://docs.rfam.org/page/contact-us.html.
Supplementary Material
Acknowledgements
The authors would like to thank the community for its excellent feedback and the organizers of the Benasque 2022 and 2024 meetings where several collaborations described in this paper were established. We would also like to thank the broader RNA community for their contributions, including providing alignments and feedback for families. In particular, we thank Lars Barquist, Fei Qi, Christina Weinberg, Zasha Weinberg and Quentin Vincens who kindly shared their data with us and worked with us in the creation of new families.
Contributor Information
Nancy Ontiveros-Palacios, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Emma Cooke, SciBite Limited, BioData Innovation Centre, Wellcome Genome Campus, Hinxton, Cambridge CB10 1DR, UK.
Eric P Nawrocki, National Center for Biotechnology Information, U.S. National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA.
Sandra Triebel, RNA Bioinformatics and High-Throughput Analysis, Friedrich Schiller University Jena, 07743 Jena, Germany; European Virus Bioinformatics Center, Friedrich Schiller University Jena, 07743 Jena, Germany.
Manja Marz, RNA Bioinformatics and High-Throughput Analysis, Friedrich Schiller University Jena, 07743 Jena, Germany; European Virus Bioinformatics Center, Friedrich Schiller University Jena, 07743 Jena, Germany.
Elena Rivas, Department of Molecular and Cellular Biology, Harvard University, Cambridge, MA 02138, USA.
Sam Griffiths-Jones, School of Biological Sciences, Faculty of Medicine, Biology and Health, Michael Smith Building, The University of Manchester, Dover St, Manchester M13 9NT, UK.
Anton I Petrov, Riboscope Ltd, Cambridge CB1 1AH, UK.
Alex Bateman, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Blake Sweeney, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Data availability
All Rfam data are released under the Creative Commons Zero (CC0) licence at https://rfam.org. The data can be accessed via an API, a public MySQL database and the FTP archive. The Rfam documentation (https://docs.rfam.org) and (31) contain detailed instructions. All code is available on GitHub under the Apache 2.0 licence at https://github.com/Rfam and Zenodo at 10.5281/zenodo.13919037 and 10.5281/zenodo.13919054.
Supplementary data
Supplementary Data are available at NAR Online.
Funding
Wellcome Trust [218302/Z/19/Z]; Biotechnology and Biological Sciences Research Council [BB/S020462/1]; NIH Intramural Research Program; ELIXIR; and EMBL-EBI Core Funds. Funding for open access charge: Wellcome Trust [218302/Z/19/Z].
Conflict of interest statement. AB is a member of the Nucleic Acids Research Editorial Board.
References
- 1. Griffiths-Jones S., Bateman A., Marshall M., Khanna A., Eddy S.R.. Rfam: an RNA family database. Nucleic Acids Res. 2003; 31:439–441. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Nawrocki E.P., Eddy S.R.. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics. 2013; 29:2933–2935. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Nawrocki E.P., Burge S.W., Bateman A., Daub J., Eberhardt R.Y., Eddy S.R., Floden E.W., Gardner P.P., Jones T.A., Tate J.et al.. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res. 2015; 43:D130–D137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. UniProt Consortium UniProt: the Universal Protein knowledgebase in 2023. Nucleic Acids Res. 2023; 51:D523–D531. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Kalvari I., Argasinska J., Quinones-Olvera N., Nawrocki E.P., Rivas E., Eddy S.R., Bateman A., Finn R.D., Petrov A.I.. Rfam 13.0: shifting to a genome-centric resource for non-coding RNA families. Nucleic Acids Res. 2018; 46:D335–D342. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Harrison P.W., Amode M.R., Austine-Orimoloye O., Azov A.G., Barba M., Barnes I., Becker A., Bennett R., Berry A., Bhai J.et al.. Ensembl 2024. Nucleic Acids Res. 2024; 52:D891–D899. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A.J., Bambrick J.et al.. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024; 630:493–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Rivas E., Clements J., Eddy S.R.. Estimating the power of sequence covariation for detecting conserved RNA structure. Bioinformatics. 2020; 36:3072–3076. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Gene Ontology Consortium Aleksander S.A., Balhoff J., Carbon S., Cherry J.M., Drabkin H.J., Ebert D., Feuermann M., Gaudet P., Harris N.L.et al.. The Gene Ontology knowledgebase in 2023. Genetics. 2023; 224:iyad031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Sant D.W., Sinclair M., Mungall C.J., Schulz S., Zerbino D., Lovering R.C., Logie C., Eilbeck K.. Sequence Ontology terminology for gene regulation. Biochim. Biophys. Acta Gene Regul. Mech. 2021; 1864:194745. [DOI] [PubMed] [Google Scholar]
- 11. Kozomara A., Birgaoanu M., Griffiths-Jones S.. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019; 47:D155–D162. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Wu C.H., Yeh L.-S.L., Huang H., Arminski L., Castro-Alvear J., Chen Y., Hu Z., Kourtesis P., Ledley R.S., Suzek B.E.et al.. The Protein Information Resource. Nucleic Acids Res. 2003; 31:345–347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Armstrong D.R., Berrisford J.M., Conroy M.J., Gutmanas A., Anyango S., Choudhary P., Clark A.R., Dana J.M., Deshpande M., Dunlop R.et al.. PDBe: improved findability of macromolecular structure data in the PDB. Nucleic Acids Res. 2020; 48:D335–D343. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Sarver M., Zirbel C.L., Stombaugh J., Mokdad A., Leontis N.B.. FR3D: finding local and composite recurrent structural motifs in RNA 3D structures. J. Math. Biol. 2008; 56:215–252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Rivas E. RNA structure prediction using positive and negative evolutionary information. PLoS Comput. Biol. 2020; 16:e1008387. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., Neelakantan A., Shyam P., Sastry G., Askell A.et al.. Language models are few-shot learners. 2020; arXiv doi:22 July 2020, preprint: not peer reviewedhttps://arxiv.org/abs/2005.14165.
- 17. Kalvari I., Nawrocki E.P., Ontiveros-Palacios N., Argasinska J., Lamkiewicz K., Marz M., Griffiths-Jones S., Toffano-Nioche C., Gautheret D., Weinberg Z.et al.. Rfam 14: expanded coverage of metagenomic, viral and microRNA families. Nucleic Acids Res. 2021; 49:D192–D200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Triebel S., Lamkiewicz K., Ontiveros N., Sweeney B., Stadler P.F., Petrov A.I., Niepmann M., Marz M.. Comprehensive survey of conserved RNA secondary structures in full-genome alignment of Hepatitis C virus. Sci. Rep. 2024; 14:15145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Zhao W., Gupta A., Krawczyk J., Gupta S.. The miR-17-92 cluster: yin and Yang in human cancers. Cancer Treat. Res. Commun. 2022; 33:100647. [DOI] [PubMed] [Google Scholar]
- 20. Wang F., Wang J., Zhang H., Fu B., Zhang Y., Jia Q., Wang Y.. Diagnostic value of circulating miR-155 for breast cancer: a meta-analysis. Front. Oncol. 2024; 14:1374674. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Rougvie A.E. Control of developmental timing in animals. Nat. Rev. Genet. 2001; 2:690–701. [DOI] [PubMed] [Google Scholar]
- 22. RNAcentral Consortium RNAcentral 2021: secondary structure integration, improved sequence search and new member databases. Nucleic Acids Res. 2021; 49:D212–D220. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Ryan D., Jenniches L., Reichardt S., Barquist L., Westermann A.J.. A high-resolution transcriptome map identifies small RNA regulation of metabolism in the gut microbe Bacteroides thetaiotaomicron. Nat. Commun. 2020; 11:3557. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Weinberg C.E., Olzog V.J., Eckert I., Weinberg Z.. Identification of over 200-fold more hairpin ribozymes than previously known in diverse circular RNAs. Nucleic Acids Res. 2021; 49:6375–6388. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Chen Y., Qi F., Gao F., Cao H., Xu D., Salehi-Ashtiani K., Kapranov P.. Hovlinc is a recently evolved class of ribozyme found in human lncRNA. Nat. Chem. Biol. 2021; 17:601–607. [DOI] [PubMed] [Google Scholar]
- 26. Steckelberg A.-L., Vicens Q., Costantino D.A., Nix J.C., Kieft J.S.. The crystal structure of a Polerovirus exoribonuclease-resistant RNA shows how diverse sequences are integrated into a conserved fold. RNA. 2020; 26:1767–1776. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Hooks K.B., Griffiths-Jones S.. Conserved RNA structures in the non-canonical Hac1/Xbp1 intron. RNA Biol. 2011; 8:552–556. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Benson D.A., Cavanaugh M., Clark K., Karsch-Mizrachi I., Lipman D.J., Ostell J., Sayers E.W.. GenBank. Nucleic Acids Res. 2013; 41:D36–D42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Hatos A., Quaglia F., Piovesan D., Tosatto S.C.E.. APICURON: a database to credit and acknowledge the work of biocurators. Database (Oxford). 2021; 2021:baab019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Green A., Ribas C., Ontiveros-Palacios N., Griffiths-Jones S., Petrov A.I., Bateman A., Sweeney B.. LitSumm: large language models for literature summarisation of non-coding RNAs. 2023; arXiv doi:06 November 2023, preprint: not peer reviewedhttps://arxiv.org/abs/2311.03056.
- 31. Kalvari I., Nawrocki E.P., Argasinska J., Quinones-Olvera N., Finn R.D., Bateman A., Petrov A.I.. Non-coding RNA analysis using the Rfam Database. Curr. Protoc. Bioinformatics. 2018; 62:e51. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All Rfam data are released under the Creative Commons Zero (CC0) licence at https://rfam.org. The data can be accessed via an API, a public MySQL database and the FTP archive. The Rfam documentation (https://docs.rfam.org) and (31) contain detailed instructions. All code is available on GitHub under the Apache 2.0 licence at https://github.com/Rfam and Zenodo at 10.5281/zenodo.13919037 and 10.5281/zenodo.13919054.






