Abstract
Despite the groundbreaking advances in deep learning-enabled methods for biomolecular modeling, predicting accurate three-dimensional (3D) structures of RNA remains challenging owing to the highly flexible nature of RNA molecules combined with the limited availability of evolutionary sequences or structural homology. Here we introduce RNAbpFlow, a sequence- and base pair-conditioned SE(3)-equivariant flow-matching model for generating RNA 3D structural ensembles. Leveraging a nucleobase center representation, RNAbpFlow enables end-to-end generation of all-atom RNA structures without the explicit or implicit use of evolutionary information or homologous structural templates. Experimental results show that base-pairing conditioning leads to broadly generalizable performance improvements over current approaches for RNA topology sampling and predictive modeling in large-scale benchmarking.
Subject terms: Machine learning, Computational models
RNAbpFlow generates all-atom RNA conformational ensembles for single-chain RNA monomers without the explicit or implicit use of evolutionary information or homologous structural templates.
Main
The determination of three-dimensional (3D) structures of RNA has become a crucial challenge in structural biology, driven by the growing interest in RNA-based therapeutics1,2. High-resolution characterization of 3D RNA structure is essential for the design and understanding of RNA molecules with specific therapeutic functions3, thus expanding the scope of RNA-mediated drug discovery4. However, the intrinsic conformational flexibility of RNA presents major challenges for experimental structure determination methods such as X-ray crystallography, nuclear magnetic resonance (NMR) spectroscopy and cryo-electron microscopy. Computational RNA structure prediction, therefore, is emerging as an attractive alternative to fill the gap in RNA structure space and to elucidate RNA conformational dynamics that underpin diverse cellular processes5.
Traditional RNA 3D-structure prediction methods include template-based approaches such as ModeRNA6 and RNAbuilder7, which rely on homologous structural information, as well as physics- and/or knowledge-based methods such as FARFAR28, 3dRNA9, RNAComposer10 and Vfold3D11, which exploit biophysical potentials and prebuilt fragment libraries to assemble full-length RNA structures. However, these approaches are constrained by the scarcity of RNA structural data in the Protein Data Bank (PDB)12 and are often computationally prohibitive, making them less suitable for the prediction of large RNAs with complex topologies13,14. Although physics-based methods combined with expert human intervention have demonstrated success in community-wide blind RNA-Puzzles15 and Critical Assessment of Structure Prediction (CASP) challenges16, there remains a critical need for fully automated, fast and accurate methods for computational modeling of RNA structures.
Inspired by the transformative impact of AlphaFold 217 on protein structure prediction18, a growing number of deep learning-based methods have recently been developed for modeling RNA structures, including DRfold13, trRosettaRNA14, trRosettaRNA219, RoseTTAFoldNA20, RhoFold+21 and NuFold22, by leveraging attention-powered transformer architectures23. However, except for DRfold, most of these methods are highly dependent on explicit evolutionary sequence information derived from multiple sequence alignments (MSA) or implicitly make use of homologous information learned by biological language models21. Obtaining reliable MSAs for RNA sequences poses notable challenges owing to the isosteric nature of base pair interactions, hindering sequence alignment efforts24. Furthermore, many existing methods fail to fully leverage the RNA base pair (2D) information, including canonical and noncanonical base-pairing interactions, key determinants of the final 3D conformation of RNA25,26. Finally, the static structure predictions made by these approaches might be inadequate to capture the inherent conformational flexibility of an RNA molecule that often adopts a distribution of conformational states instead of folding into a static structure5,27. Thus, there is an urgent need to develop improved computational methods that can generate a conformational ensemble of all-atom RNA 3D structures directly from the nucleotide sequence by making use of base-pairing information without explicitly or implicitly using any evolutionary information.
Diffusion-based generative modeling has achieved remarkable success in the image domain and is now attracting considerbale attention in structural bioinformatics, most notably in AlphaFold 3 (AF3)28 for biomolecular interaction prediction and AF3-inspired approaches such as Boltz-229, Chai-130 as well as 3D protein backbone generation with approaches such as RFdiffusion31 and FrameDiff32. More recently, methods such as FrameFlow33 have showcased the power of SE(3)-equivariant flow matching, achieving better precision in designability than their diffusion-based counterpart but with reduced sampling costs. RNA-FrameFlow34, a recent adaptation of FrameFlow for RNA, is the first generative model that specifically targets 3D RNA backbone generation. However, the method uses naive unconditional flow matching, generating backbones without using sequence or base-pairing information. Such a limitation highlights the opportunity to develop improved deep generative models for RNA that can efficiently sample a large conformational ensemble of all-atom RNA 3D structures by explicitly conditioning on the nucleotide sequence and base-pairing information through conditional flow matching, thereby offering an elegant combination of principles-based and data-driven approach to RNA structural ensemble generation free from sequence- and structural-level homology.
Here, we present RNAbpFlow, a sequence- and base pair-conditioned SE(3)-equivariant flow-matching model for generating all-atom RNA conformational ensemble for single-chain RNA monomers. The major contributions of this work lie in the methodology development of a conditional flow matching-based method for RNA 3D structure generation. First, RNAbpFlow incorporates conditions on the nucleotide sequence and base-pairing information from three complementary base pair annotation methods to comprehensively capture canonical and noncanonical interactions. Second, by incorporating a nucleobase center representation that enables the optimization of angles of all rotatable bonds of nucleobases, RNAbpFlow directly outputs all-atom RNA structures in an end-to-end fashion, bypassing the need for a post hoc geometry optimization module, which is impractical in the context of large-scale sample generation. Third, we introduce base pair-centric auxiliary-loss functions to enable maximal realization of the canonical and noncanonical base-pairing interactions. In summary, the central methodological contribution of RNAbpFlow is its integration of a flexible nucleobase representation with a FrameFlow-inspired SE(3)-equivariant flow-matching process. This design choice enables efficient generation of all-atom RNA conformational ensembles while explicitly modeling nucleobase orientation and flexibility. Furthermore, the incorporation of base pair-centric auxiliary-loss functions during training to promote faithful realization of both canonical and noncanonical interactions improves the fidelity of conditional generation during inference. Collectively, these components yield an efficient conditional sampler that unifies nucleobase-level flexibility with SE(3)-equivariant generative modeling.
We empirically observe performance improvements when base-pairing condition is introduced across a wide range of evaluation metrics, outperforming a recent molecular dynamics (MD) simulation-based global topology sampling method RNAJP35, which explicitly considers base-pairing and base-stacking interactions. In addition, RNAbpFlow generalizes well for sequence- and base pair-conditioned RNA 3D-structure prediction compared with several state-of-the-art methods, including top-performing automated servers participating in the recently concluded CASP16 challenge. RNAbpFlow is freely available via GitHub at https://github.com/Bhattacharya-Lab/RNAbpFlow.
Results
Overview of the RNAbpFlow framework
An overview of our method, RNAbpFlow, is shown in Fig. 1a, with a detailed architectural diagram shown in Fig. 1b. Our framework is built on the foundations of FrameFlow33, a flow-matching formulation tailored for fast protein backbone generation on the SE(3) frame representation. To represent each nucleotide in an RNA sequence as a rigid body frame defined by a translation from the global origin and a rotation matrix, as well as for the full atomic RNA 3D structure generation in an end-to-end manner, we follow the nucleotide representation presented in NuFold22. The rotation matrix is constructed using the Cartesian coordinates of C1′ as the origin of the local frame and O4′–C1′–C2′ for orientation. Given an RNA sequence of length N, we begin with N such frames sampled from a Gaussian distribution as the starting point for our iterative sampling process. Our method incorporates conditions on the nucleotide sequence and base-pairing information to guide the conformational sampling process to generate the all-atom RNA structure. The three base-frame atoms, O4′, C1′ and C2′, are derived from the learned frame representation, while the first nitrogen of the base (N1 for pyrimidines or N9 for purines) is imputed using tetrahedral geometry. The remaining atoms are partitioned into ten frames, and the corresponding atomic coordinates are generated on the basis of the iterative update of the frames using the nine predicted torsion angles lying on the bonds that connect these frames, leading to an all-atom RNA structure. The choice of nucleotide representation yields an orientation-preserving local coordinate frame: O4′, C1′ and C2′ define a stable sugar-based reference frame, while the first nitrogenous base (N1/N9) anchors the base attachment point, enabling the model to recover base placement and orientation via the predicted rigid-body frame and ϕ9 angle. In addition, the full atomic reconstruction includes ribose ring torsion angle predictions, enabling the model to efficiently reproduce sugar-puckering conformations (C3′-endo and C2′-endo) in the generated all-atom structures22. Details of the flow-matching formulation, network architecture and training objectives are provided in Methods.
Fig. 1. Overview of RNAbpFlow.
a, A sequence- and base-pair-conditioned SE(3)-equivariant flow-matching model for generating RNA 3D structural ensemble. Using the nucleotide sequence and base-pairing information as conditions, our end-to-end framework enables efficient sampling of all-atom RNA 3D structures based on a nucleobase center representation without explicitly or implicitly using any evolutionary information. b, Overview of the transformer-based invariant point attention (IPA) module used in the backbone of the flow-matching network in RNAbpFlow. Left: the full six-block model stack operating on single and pair representations to update backbone frames under different training objectives. Right: an expanded view of the IPA block shows how sequence features and input base pair maps modulate the attention computation to update the representations. Here, L is sequence length and h is hidden dimension; q, k and v denote per-head query, key and value projections of the single representation; denote their geometric (point) projections; Ti = (Ri, ti) is the backbone frame for nucleotide i, with rotation Ri ∈ SO(3) and translation , and the outputs Oi indicate the per-head attention components combined to update the representations. Repr., representation.
Structural ensemble generation performance
To evaluate the sampling performance of RNAbpFlow, we compare against RNAJP35, a recent coarse-grained MD simulation-based method for RNA 3D-structure sampling with explicit consideration of base pair information, including noncanonical base-pairing and base-stacking interactions as well as long-range loop–loop interactions. For the curated benchmark set of 12 RNA targets containing three-way junctions described in Methods, we generate 1,000 3D structural samples per target with base pairs extracted from the native 3D structures using three different annotation tools, RNAView36, MC-Annotate37 and DSSR38 (detailed in Methods), and compute the mean and maximum scores (template modeling score (TM-score)39 and local distance difference test (lDDT)40) for each target, which are then averaged across the 12 RNAs. To ensure a fair comparison, we randomly select 1,000 decoy structures for each target from the entire ensemble provided by RNAJP, sampled similarly on the basis of experimental base pairs.
The results, shown in Fig. 2a, demonstrate that RNAbpFlow consistently outperforms RNAJP in both metrics. For example, RNAbpFlow achieves an average over the mean lDDT score of 0.66, surpassing that of 0.59 achieved by RNAJP while maintaining the same standard deviation. Similarly, in terms of global topology sampling, RNAbpFlow generates higher-quality structures, achieving an average over the mean TM-score of 0.38 compared with RNAJP’s 0.32. To evaluate sampling efficacy, we calculate the fraction of targets for which at least one correct fold (TM-score >0.45 or lDDT >0.75) appears in the ensemble41 as well as the fraction of all decoys that are correctly folded. RNAbpFlow finds a TM-score-based correct fold in 66.67% of the RNAs and an lDDT-based correct fold in 25%, compared with RNAJP’s 41.67% and 0.0%, respectively. Across all 12,000 decoys generated by RNAbpFlow, 13.4% achieve a TM-score >0.45 and 9.6% achieve an lDDT >0.75. By contrast, only 1.73% of RNAJP’s decoys exceed a TM-score >0.45 and 0.0% exceed an lDDT >0.75. These results demonstrate that RNAbpFlow not only outperforms RNAJP in terms of top-scoring structures but also results in a substantially higher proportion of high-quality decoys, highlighting its efficiency in sampling both global topology and local conformations. Detailed performance comparison for all targets in the benchmark set is presented in Supplementary Table 19. Figure 2b,c shows two representative examples from the RNAJP benchmark set, illustrating the predictive performance of RNAbpFlow compared with the RNAJP (detailed in Supplementary Results 3.3.1).
Fig. 2. Sampling performance comparison with RNAJP.
a, Performance comparison in terms of maximum, mean and standard deviation scores of 1,000 3D structural samples predicted for 12 RNAJP test RNAs. Values in bold indicate the best performance. b, Two representative RNA targets with PDB ID 2HGH and 3PDR, both containing three-way junctions, shown as predicted structural models (blue) superimposed on the experimental structures (green) with the corresponding evaluation metrics reported at the bottom-right of each superposition. c, Corresponding canonical and noncanonical (annotated in blue) base pair annotations extracted using RNApdbee 3.0 (ref. 54) (default parameters with FR3D annotator) from the experimental and predicted 3D structures of the same targets 2HGH and 3PDR; base pair fidelity is evaluated in terms of all interactions (INF-All), noncanonical interactions (INF-NWC) and base-stacking interactions (INF-Stack).
Performance on CASP15 targets
To compare our method RNAbpFlow against various sequence- and base pair-conditioned RNA 3D-structure prediction approaches on CASP15 RNAs, we train RNAbpFlow by curating a training set from the full RNA3DB dataset42 by retaining RNA chains released in the PDB12 before April 2022. We leverage our recently published RNA 3D structure scoring method lociPARSE43 to select the top-scoring structure from the RNAbpFlow-generated structural ensemble on the basis of lociPARSE’s estimated lDDT score (pMoL). For the physics- and/or knowledge-based methods, we provide accurate (native) base-pairing information in the required dot-bracket notation (DBN) format, as extracted from the experimental 3D structures by DSSR software (the only RNA 3D structure annotator that outputs 2D structure in DBN format), whereas the predictive modeling performance of our method RNAbpFlow is compared against several deep learning-based RNA structure prediction methods by locally installing and running with their default parameter settings and 2D-structure prediction pipeline in MSA-independent manner for a fair evaluation (detailed in Methods). As presented in Table 1, the predictive modeling performance of our method RNAbpFlow noticeably improves in the presence of accurate (native) base pairs, achieving an average TM-score of 0.48, all-atom root mean square deviation (RMSD) of 7.77 and interaction network fidelity (INF) for non-Watson–Crick base pairs (NWC) of 0.62, compared with a 0.40, 10.70 and 0.48 TM-score, RMSD and INF-NWC, respectively, when predicted base pairs are used, showing a 20% improvement in TM-score, a 27.4% reduction in RMSD and 29.2% improvement in INF-NWC. By contrast, all other physics- and/or knowledge-based approaches show only moderate performance even with accurate (native) base pairs (for example, Vfold reaches the maximum TM-score of 0.34 and RNAComposer achieves a minimum RMSD of 14.15 among all). The deep learning-based methods, when supplied with accurate (native) base pairs, exhibit somewhat improved performance compared with the physics- and/or knowledge-based approaches but still much worse than our method RNAbpFlow (Supplementary Table 17). This highlights the versatility of RNAbpFlow in its ability to achieve a higher performance ceiling by effectively leveraging accurate base-pairing information as a key condition in deep generative modeling. In the absence of native base-pairing information, we use pseudoknot-aware sequence-based predicted base pair maps from three RNA 2D-structure predictors, namely, IPKnot44, SPOT-RNA45 and RibonanzaNet46, chosen on the basis of their sampling performance when individually conditioned on CASP15 natural targets as presented in Supplementary Table 18. When conditioned on the combination of these three predicted (noisy) base pairs, RNAbpFlow consistently outperforms the deep learning-based RNA structure prediction methods: NuFold, trRosettaRNA, RhoFold+ and DRfold, all run with their default parameter settings along with any postprediction optimization and/or refinement, exhibiting improved accuracy across all reported metrics on natural RNAs (Table 1). The pairwise statistical significance tests (P−values) are presented in Supplementary Table 16, and the scatter plots comparing RNAbpFlow with each competing method across all evaluation metrics are shown in Supplementary Figs. 1 and 2. However, owing to the small sample size (n = 6), caution should be exercised when interpreting the resulting P−values. In terms of modeling challenging motifs involving noncanonical interactions and pseudoknots, RNAbpFlow achieves improved base pair fidelity than competing methods, particularly for INF-NWC (0.62) and stacking interactions (INF-Stack, 0.89) with accurate (native) base pair input, while successfully recovering all pseudoknots across the six targets, consistent with the input base pair fidelity analysis in the following section. Finally, regarding the stereochemical quality of the predicted structures, physics- and/or knowledge-based methods tend to report fewer clashes, bond and angle violations than RNAbpFlow (Supplementary Table 12), probably owing to the lack of any postprediction optimization and/or refinement used in the default pipeline of our method RNAbpFlow. However, additional PyRosetta-based relaxation through iterative energy minimization47 can substantially improve stereochemical correctness for RNAbpFlow predictions with a minimal change in global and local accuracy metrics as presented in Supplementary Table 13.
Table 1.
Benchmarking the predictive modeling performance on all six available natural RNAs from CASP15 with both predicted and native base pairs as input
| Input modality | Predictors | TM-score | lDDT | All RMSD | GDT-TS | INF-All | INF-WC | INF-NWC | INF-Stack |
|---|---|---|---|---|---|---|---|---|---|
| Native base pairs | RNAbpFlowa | 0.48 | 0.75 | 7.77 | 53.66 | 0.89 | 1.00 | 0.62 | 0.89 |
| Vfold-pipelineb | 0.34 | 0.62 | 15.45 | 38.89 | 0.82 | 0.97 | 0.56 | 0.79 | |
| RNAComposerb | 0.33 | 0.58 | 14.15 | 37.25 | 0.79 | 0.99 | 0.48 | 0.77 | |
| 3dRNAb | 0.31 | 0.51 | 17.36 | 33.45 | 0.76 | 0.85 | 0.35 | 0.76 | |
| SimRNAb | 0.30 | 0.60 | 15.30 | 34.10 | 0.79 | 0.93 | 0.35 | 0.80 | |
| GraphaRNAa | 0.18 | 0.47 | 22.17 | 20.27 | 0.66 | 0.90 | 0.30 | 0.63 | |
| Predicted base pairs | RNAbpFlowa | 0.40 | 0.68 | 10.70 | 44.86 | 0.85 | 1.00 | 0.48 | 0.83 |
| DRfolda | 0.33 | 0.58 | 14.02 | 38.70 | 0.75 | 0.81 | 0.32 | 0.77 | |
| NuFolda | 0.33 | 0.56 | 14.02 | 37.40 | 0.73 | 0.67 | 0.27 | 0.79 | |
| trRosettaRNAa | 0.32 | 0.54 | 13.44 | 38.81 | 0.66 | 0.64 | 0.16 | 0.71 | |
| RhoFold+a,c | 0.29 | 0.42 | 16.56 | 34.88 | 0.52 | 0.49 | 0.07 | 0.57 |
Values in bold indicate the best performance.
aDeep learning methods.
bPhysics and/or knowledge-guided methods.
cLanguage model-based method, run in single-sequence mode.
Beyond natural RNAs, when evaluated on the four synthetic CASP15 RNAs, RNAbpFlow conditioned on native base pairs generates high-quality ensembles, with the highest ensemble accuracies reaching a TM-score ≥0.45 for three out of the four targets as presented in Supplementary Table 14. In terms of overall predictive modeling performance, RNAbpFlow conditioned on predicted base pairs attains the highest average TM-score and lDDT among all competing methods, together with superior base pair fidelity across all INF variants (Supplementary Table 15), specially for noncanonical base pair fidelity (INF-NWC, 0.73) and stacking recovery (INF-Stack, 0.88). The performance, however, slightly drops with predicted input base pair conditioning, consistent with the evaluations in the following section. To further characterize the performance of both natural and synthetic RNAs in the presence of pseudoknots, we present a fine-grained analysis in Supplementary Table 21. In summary, our method RNAbpFlow exhibits improved versatility on both natural and synthetic RNAs, with and without pseudoknots, outperforming all other competing methods.
Performance on CASP16 targets
Performance comparison with state-of-the-art methods
AF328 has emerged as a unified framework for deep learning-based modeling of biomolecular 3D structures, including RNA 3D-structure prediction, in addition to recent AlphaFold 2-inspired architectures such as trRosettaRNA219, NuFold22 and DRfold248. In the recently concluded CASP16 challenge, the two automated server methods officially ranked top 10 in RNA structure prediction category, MSA-based AF3-server and Yang-Server, have demonstrated state-of-the-art performance, alongside top-ranked human expert predictors49,50. Multiple participating groups in CASP16, both human and automated, incorporated AF3 predictions to generate RNA structural ensemble51. However, the accuracy of these state-of-the-art methods still remains heavily reliant on the availability of evolutionary information in the form of homologous structural templates and/or deep MSAs51. To evaluate the sampling accuracy of our MSA- and template-free method RNAbpFlow on blind CASP16 targets, we use a training set from the full RNA3DB dataset42, which was curated from the PDB on 26 April 2024 before the start of the CASP16 challenge, and obtain 994 RNA sequences. Following recent literature22, we further perform data augmentation using a cross-distillation set of an additional 2,170 sequences collected from the bpRNA-1m (90) dataset52, along with their corresponding experimental base pair information. The details of the curation of the cross-distillation set are available in Methods. For training, we use base pairs extracted from experimental 3D structures from the PDB or experimental base pair annotations (for bpRNA-1m (90)) as ground truth. During inference, we incorporate predicted base pairs from three RNA 2D-structure predictors: IPKnot44, SPOT-RNA45 and RibonanzaNet46, following CASP15 (detailed in Methods). We use all 28 CASP16 targets whose experimental structures are available for this blind testing. The complete list of 28 publicly available CASP16 targets and their associated metadata are presented in Supplementary Table 1. As most of the experimental structures in our training set are ≤200 nucleotides (~92%) and the vast majority of RNA sequences of interest in the current Rfam release (v15.0) are relatively short as well (~91% of seed sequences have a length of ≤200), we emphasize 14 targets with a length of ≤200 nucleotides for primary analyses, while the results for the full set are provided in Supplementary Table 5.
As shown in Fig. 3a, our method RNAbpFlow without using any MSA information leads to better accuracy in terms of average of the maximum TM-score and lDDT among the models present in the generated structural ensemble than the two top-performing automated CASP16 servers achieve in their best of five CASP16 submissions despite using MSA and/or template information. To analyze the dependency on MSA and template information of these methods, we compute the target-wise TM-score difference between the best of five CASP16 submissions from either of the servers and the maximum TM-score achieved by the structural ensemble generated by RNAbpFlow. Figure 3b shows that, for hard targets (TM-score <0.45)51 with weak evolutionary signal in terms of shallow MSA depth (Neff ≤130)51, RNAbpFlow consistently yields higher TM-scores in the structural ensemble than both the automated servers, demonstrating the effectiveness of base pair-conditioned structure modeling when evolutionary information is scarce. By contrast, for easy targets such as R1263 or R1264, both servers perform comparably or outperform RNAbpFlow owing to the availability of deep MSA, revealing their dependence on MSA information. As the search for homologous sequences to obtain sufficiently deep MSA is a challenging task for RNA24,53 (only 4 out of 14 CASP16 targets in this analysis have sufficiently deep MSA with Neff >130), our method RNAbpFlow can be an attractive alternative for RNA 3D-structure generation conditioned only on base pair information, which is much more conserved across species than sequences53. Two such representative targets from CASP16 are shown in Fig. 3c (detailed in Supplementary Results 3.3.2), where RNAbpFlow correctly recovers the native fold, whereas neither server predicts a correct model. Detailed target-wise results for the two CASP16 servers are presented in Supplementary Table 2. To benchmark the sampling performance of RNAbpFlow against the state-of-the-art methods in the absence of MSA, we compare our method against the stand-alone versions of AF3 (docker version), NuFold, trRosettaRNA2 and DRfold2, all of which are run via local installation. For both AF3 and NuFold, we generate 1,000 structures per target by varying the random seed and running without using any MSA information. As trRosettaRNA2 and DRfold2 are nonstochastic methods with no seeding mechanism available for generating an ensemble for a fixed input, we consider the default output of a single predicted structure per target for these methods. As shown in Fig. 3a, the performance of RNAbpFlow is better than all the competing methods, including AF3, in terms of maximum TM-score and lDDT over the average of 14 targets. At least one correct fold (TM-score >0.45) is present in the generated ensemble of RNAbpFlow for 12 out of 14 targets (85.71%) compared with 8 out of 14 targets (57.13%) for AF3, demonstrating the consistency of RNAbpFlow. Meanwhile, RNAbpFlow yields much better performance compared with trRosettaRNA2, DRfold2 and NuFold both in terms of TM-score and lDDT. Target-wise sampling results for AF3 and NuFold are presented in Supplementary Table 3, while results for trRosettaRNA2 and DRfold2 are presented in Supplementary Table 4. Target-wise results for RNAbpFlow under different training and inference configurations are presented in Supplementary Table 5, with performance comparison grouped by pseudoknot classification presented in Supplementary Table 22.
Fig. 3. RNAbpFlow can sample higher-quality 3D structures than state-of-the-art methods.
a, Performance comparison of RNAbpFlow with top-performing methods, using predicted base pairs as input during inference, on 14 CASP16 targets with a length of ≤200 nucleotides. Table entries report the average maximum TM-score and lDDT across the 14 target-wise predicted ensembles. Values in bold indicate the best performance. aDid not participate in CASP. b, Target-wise differences in maximum TM-score between the best structure in the RNAbpFlow ensemble (with cross-distillation + predicted base pairs) and the best of five submissions from AF3-server (group 304) and Yang-Server (group 052), with targets ordered by increasing sequence length. Positive values indicate targets on which RNAbpFlow outperforms the corresponding server method, highlighting its advantage on harder targets and/or targets with shallow MSAs. c, Representative predictions for two CASP16 targets, R1288 (in vitro ribozyme) having no MSA or template information and R1255 (SL5 SARS-CoV-2), one of the several conserved RNA structural elements in the -UTR of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Predicted structures from RNAbpFlow, AF3-server and Yang-Server (blue) are superimposed on the corresponding experimental structures (green), with TM-score, lDDT, RMSD and GDT-TS shown for each prediction.
For targets larger than 200 nucleotides, RNAbpFlow consistently outperforms NuFold, trRosettaRNA2 and DRfold2 across all metrics, while slightly trailing AF3. This gap is a limitation of conditioning on noisy predicted input base pairs, which are less accurate for larger RNAs. In particular, the average consensus quality of the three predicted input maps for RNAbpFlow inference in terms of INF is 0.51 for larger targets, compared with 0.84 for smaller targets (a 39% relative drop). As presented in Supplementary Table 20, when accurate native base pairs are provided as input instead of noisy predicted input base pairs, the average over the maximum TM-score achieved by RNAbpFlow for the larger targets increases from 0.32 to 0.42 and the average over the maximum lDDT rises from 0.47 to 0.58, indicating the higher performance ceiling of RNAbpFlow in the presence of accurate input conditions, which is consistent with the base pair fidelity analyses presented below, showing RNAbpFlow’s tight conformity to input base pair conditions regardless of their quality. Nonetheless, attaining improved accuracy for these longer targets remains a substantial challenge for all state-of-the-art approaches, including our method RNAbpFlow, and even for the human expert groups participating in CASP16. For 10 out of the 14 large RNA targets in the CASP16 challenge, no submission from any human or server group achieved an RMSD below 15 Å, underscoring the difficulty in modeling large RNAs. Supplementary Fig. 3 shows the predictions by RNAbpFlow on four such large RNA targets from CASP15 and CASP16 sets.
Contribution of data augmentation, fine-tuning and base pair accuracy
To evaluate the contribution of data augmentation via cross-distillation training of RNAbpFlow, fine-tuning with predicted base pairs and the impact of base pair accuracy, we generate 1,000 3D structural models for each CASP16 target using both native and predicted base pairs as inputs with and without distillation and fine-tuning. As presented in Table 2, data augmentation via cross-distillation training yields substantial gains over the ‘no distillation’ variant regardless of the nature of base pair information (that is, predicted versus experimental) used for conditional generation (target-wise results with and without distillation are presented in Supplementary Table 5). With distillation and predicted base pair conditioning, the average over the maximum TM-score increases from 0.50 to 0.57, a relative improvement of 14%, and the average over the maximum lDDT increases from 0.61 to 0.69, a relative improvement of 13.1%. With experimental base pair conditioning, the performance improves substantially, attaining an average over the maximum of 0.68 and an average over the maximum lDDT of 0.77, highlighting a substantial accuracy gap between predicted and experimental base pair-conditioned sampling (Supplementary Table 6). To reduce this gap, in addition to training on the distillation-augmented dataset, we fine-tune RNAbpFlow using predicted base pairs from three pseudoknot-aware RNA 2D-structure predictors: IPKnot44, SPOT-RNA45 and RibonanzaNet46 (detailed in Methods). The fine-tuned model further improves RNAbpFlow’s sampling quality, achieving an average over the maximum TM-score of 0.61 while also markedly increasing the mean TM-score and lDDT, thereby narrowing the accuracy gap between predicted and experimental base pair conditioning. To further quantify the role of base pair quality under predicted conditioning (using inputs from three different predictors), we examine how per-target input base pair accuracy relates to downstream 3D sampling quality. Supplementary Table 7 presents Spearman correlations between INF (agreement of the individual predicted input base pairs to the consensus experimental base pairs) and ensemble TM-score or lDDT statistics. Overall, RNAbpFlow sampling shows moderate-to-strong associations with base pair accuracy, with correlations generally stronger for lDDT than for TM-score, suggesting that accurate base pair constraints may align more closely with local structural correctness than global fold similarity41. Detailed sampling statistics (percentages of correct folds in terms of TM-score and lDDT in the ensemble) are presented in Supplementary Table 9.
Table 2.
Sampling performance in terms of maximum and mean scores of 1,000 3D structures for each target ≤200 in the CASP16 test set based on different variants of training and inference of RNAbpFlow
| Method | Training | Inference | TM-score | lDDT | ||
|---|---|---|---|---|---|---|
| Max | Mean | Max | Mean | |||
| RNAbpFlow | With distillation | Experimental base pair | 0.68 | 0.5 | 0.77 | 0.72 |
| RNAbpFlow | With fine-tuning | Predicted base pair | 0.61 | 0.45 | 0.72 | 0.66 |
| RNAbpFlow | With distillation | Predicted base pair | 0.57 | 0.41 | 0.69 | 0.63 |
| RNAbpFlow | Without distillation | Predicted base pair | 0.50 | 0.36 | 0.61 | 0.57 |
All models are trained using experimental base pair annotations. Values in bold indicate the best performance.
To analyze the fidelity with which the input base pairs are realized by RNAbpFlow predicted structures, we compute the INF score between the input base pairs (native or predicted) and the output base pairs extracted from the resulting 3D structures using RNAView36. Supplementary Table 8 highlights that when native base pairs are provided, RNAbpFlow generates structures with strong agreement to the input with an average INF value of 0.93. When provided with noisy (predicted) base-pairing information that differs from the native input, RNAbpFlow still reproduces these incorrect input base pairs with slightly lower but still high fidelity (average INF of 0.84), demonstrating the ability of RNAbpFlow to tightly conform to the provided input base pairs regardless of their accuracy. To quantify the extent to which RNAbpFlow introduces base pairs beyond those specified in the input, we compute the per-target number of false positives averaged over 1,000 samples generated (Supplementary Table 8). When conditioned on noisy (predicted) base pairs, RNAbpFlow introduces approximately threefold more erroneous base pairs per target on average than when conditioned on accurate (native) base pairs. This is consistent with RNAbpFlow’s strong adherence to the input conditions regardless of its quality, which may facilitate the formation of these additional spurious base pairs. Overall, the results demonstrate that RNAbpFlow can recover correct 3D structures, both in terms of TM-score and lDDT, when supplied with accurate experimental base pairs; moreover, while an accuracy gap remains when predicted base pairs are used compared with experimental inputs, data augmentation via cross-distillation during training and fine-tuning on predicted base pairs notably help to close this gap.
Ablation study
To assess the importance of the base-pairing information incorporated in our conditional flow-matching formulation for RNA 3D-structure generation, we utilize a sequentially and structurally nonredundant train–test split provided by RNA3DB42 (detailed in Methods). During training and inference, we individually condition on base pairs extracted from the native 3D structures using each of the three different 2D annotation tools: RNAView36, MC-Annotate37 and DSSR38 as well as their combination. We independently train each variant on the RNA3DB training set and evaluate on the 48 targets from the corresponding nonredundant test set. We also train a baseline model by conditioning only on the nucleotide sequence but not on the base-pairing information. For each target sequence, we generate 1,000 3D structural samples and calculate the maximum and mean of both TM-score and lDDT. The resulting target-wise score distributions across all 48 targets, shown in Fig. 4a, indicate that RNAbpFlow achieves the best performance when all three base pair maps are incorporated as conditions, with average maximum per-target TM-score and lDDT values of 0.51 and 0.71, respectively. This represents an average increase of 41.7% in TM-score and 54.3% in lDDT compared with the baseline sequence-conditioned variant of our flow-matching formulation, which achieves an average over the maximum per-target TM-score distribution of 0.36 and lDDT of 0.46. These improvements underscore the critical role of incorporating base pair information in RNA 3D generative modeling. The relative contribution of individual base pair maps is also evident from Fig. 4a, which demonstrates that none of the individual maps alone can outperform their combination. Furthermore, the distributions reveal the consistency of the generated sample qualities when all three maps are used, which achieves a relatively smaller interquartile range while maintaining the highest mean values for both TM-score and lDDT distributions, indicating both high-quality samples and reduced variability. This consistency is also reflected in sampling performance, where RNAbpFlow successfully generates at least one 3D structure with a TM-score exceeding 0.45 (a threshold for assessing RNA global fold correctness41) for 31 out of 48 targets, with 64.6% correctly folded targets from the test set, demonstrating the robustness of RNAbpFlow in RNA 3D structure generation when accurate base-pairing information is available. Complementary analyses in Fig. 4b further support these findings: removing any auxiliary losses reduces sampling quality, highlighting the importance of base pair-centric supervision. Details about loss-function ablation and hyperparameter selection results are discussed in Supplementary Results 3.1–3.2.
Fig. 4. Base pair conditioning improves sample generation.
a, Box plots of the distributions of target-wise maximum and mean scores over 1,000 generated 3D structural samples per target, measured by TM-score (left) and lDDT (right), across n = 48 RNA3DB test targets for various base-pairing conditioning (or lack thereof). b, The same evaluation as in a but for models trained with different auxiliary-loss combinations. In a and b, center lines indicate medians, box limits indicate the upper and lower quartiles, whiskers denote 1.5 times the interquartile range and points beyond the whiskers are plotted as outliers. Green triangles indicate arithmetic means, and the numbers overlaid on the boxes report the corresponding mean values.
Discussion
In this work, we developed RNAbpFlow, a sequence- and base pair-conditioned all-atom RNA 3D structure generation method based on SE(3)-equivariant flow matching model. Experimental results demonstrate that the introduction of base-pairing conditioning leads to performance improvements and the accuracy gain is connected to the quality of the base pairs. Free from the confines of sequence- and structural-level homology, RNAbpFlow enables direct generation of all-atom RNA 3D structural models in an end-to-end manner, thereby opening promising avenues for RNA conformational dynamics through large-scale structural ensemble generation in atomic detail. Despite the advantages, a key limitation of our method RNAbpFlow is that the sampling and predictive performance of RNAbpFlow is heavily dependent on the accuracy of input base pair information. Improving performance in this regard will require improvement in base pair prediction accuracy, especially for larger RNA targets. Furthermore, the current conditional sampling pipeline of RNAbpFlow is not specifically optimized for very large RNAs, which may require substantially more long-RNA training data and a model architecture that scales subquadratically with sequence length, such as locality-aware message passing combined with sparse attention to facilitate efficient long-range information flow. For future work, our method can be extended beyond base-pairing information to incorporate homology information (MSA) or additional experimental constraints, such as chemical probing reactivities (for example, SHAPE/DMS) as per-nucleotide signals and proximity ligation or crosslinking data as sparse pairwise restraints, which may further improve accuracy on challenging targets. Furthermore, evaluation of alternative and openly accessible base pair annotation pipelines (for example, FR3D) for experimental 3D-structure labeling during training can be explored to further improve the robustness and accuracy of RNAbpFlow.
Methods
Model input
Our method operates independently of any MSA or template information, relying solely on sequence and base-pairing (2D) information for the conditional generation of RNA 3D structure. Specifically, given an RNA sequence of length L as input, we encode the nucleotide sequence using one-hot encoding, represented as a binary vector of four entries corresponding to the four types of nucleotides (A, U, C and G). For base-pairing information during the training process, we extract 2D annotations from experimental (native) 3D structures using three different software tools, RNAView36, MC-Annotate37 and DSSR38, and represent them as three separate 2D binary maps, each having the shape L × L. However, despite analyzing the same 3D coordinates, each method applies its own geometric criteria (for example, which edges interact, planarity/angle limits and hydrogen-bond cutoffs), which can result in minor variations between base pair annotations55 as presented in Supplementary Table 10. To capture these diverse canonical and noncanonical base-pairing information, we directly provide all three binary maps as input (L × L × 3) for the bias term in our denoiser architecture as three separate channels for edge features without any reconciliation of contradictory annotations. During the sampling of CASP15 and CASP16 test targets, in the absence of native base-pairing information, we use sequence-based predicted base pair maps from three RNA 2D-structure predictors, namely IPKnot44, SPOT-RNA45 and RibonanzaNet46 (Supplementary Table 11). All three predictors provide pseudoknot-aware base pair predictions and are chosen on the basis of the sampling performance on CASP15 natural RNA targets as presented in Supplementary Table 18. However, RNAbpFlow is versatile in that it supports inference using any user-defined set of three maps or even a single custom map, which can be duplicated three times to match the required three input 2D channels.
Network architecture
RNA frame representation
We represent each nucleotide in a geometric abstraction using the concept of rigid body frames. Each nucleotide frame in the form of a tuple is defined as a Euclidean transform T = (r, x), where r ∈ SO(3) is a rotation matrix and is the translation vector that can be applied to transform a position in local coordinates to a position in global coordinates.
SE(3) flow matching on frames
Flow matching56 is a class of deep generative models in which the goal is to learn a velocity field (or flow) that matches the probability flow of the data distribution to transform a simple distribution, such as a Gaussian, to the desired complex data distribution in high-dimensional space. Flow matching directly learns this velocity field that describes how points should move from the simple distribution to the target distribution without completely destroying the data distribution. By integrating an ordinary differential equation over this learned vector field, flow matching offers simpler trajectories toward achieving the target distribution, offering a huge computational speed-up for large-scale sample generation.
In the context of generative modeling, the geodesic path describes a smooth transformation of one probability distribution into another while minimizing distortion. The concept of geodesics generalizes the notion of shortest paths in non-Euclidean spaces, enabling efficient computation and interpretation. Given a noisy frame T0 sampled from a simple prior density p0(T0) and the experimental frame T1 sampled from a target distribution p1(T1), the geodesic path connecting two points T0 and T1 in a combination of simple manifolds such as and SO(3) can be expressed using exponential and logarithmic maps following the generalization of flow matching to Riemannian manifolds57 in the following way:
| 1 |
where is a vector in the tangent space of T0 pointing toward T1, and t ∈ [0, 1] parameterizes the sequence of probability distributions referred to as the probability path pt between the two distributions p0 and p1. The conditional flow Tt constructed in equation (1) can be decomposed into separate individual flows for the simplification of the training procedure in the following way:
| 2 |
where the prior distribution p0(T0) during training takes the form of , where random translation x0 is sampled from the unit Gaussian distribution and random rotation is sampled from IGSO3(σ = 1.5) following refs. 32,33,58 for better performance. During training, the parameter t is sampled from , where ϵ = 0.1 is chosen for training stability.
Velocity field network
The objective of our flow-matching method is to learn the parametrized vector field ut, which represents a smooth, time-dependent (t) map that generates an ordinary differential equation that describes the transformation between two distributions: p0(T0) (noisy frames) and p1(T1) (ground-truth frames). To learn this mapping, we train a parameterized neural network vθ(Tt, t) that predicts the vector field given corrupted ground-truth frames Tt at time t. Following FrameFlow33, we use the structure module from AlphaFold 217 as the backbone architecture of this network.
Sampling strategy
During conditional sampling of RNA 3D structures, our framework takes a random initialization of the backbone frames T = (r, x), where translation x is sampled from a unit Gaussian distribution in and rotation r is sampled from a uniform distribution in SO(3). During inference, instead of the linear scheduler for rotation matrices, we use the exponential scheduler e−ct with c = 10 for better performance. Thus, our SO(3) flow for rotation in equation (2) changes according to the following equation:
| 3 |
On the basis of the specified number of timesteps, we discretize the interval [0, 1] into a sequence of values for t. Starting from the random set of frames, we iteratively update the frame representations using the denoised clean-frame predictions and from our learned model at each timestep t with the specified condition in the following way:
| 4 |
Training of RNAbpFlow
Our training objective contains multiple loss-function components related to base pair conditions and all-atom structure generation. The primary loss function for this framework is the same SE(3) loss formulated in FrameFlow33, namely, the vector field loss in SE(3). To train our denoiser model vθ, two separate components of this loss are calculated for the predicted rotation and translation given the corrupted frames Tt at time t as follows:
| 5 |
| 6 |
where N is the total number of frames. To predict the nine torsion angles , where each angle is represented using its sine–cosine embedding, we use an additional multilayer perceptron head and calculate the torsional loss in the following way by comparing with the experimental angles ϕ:
| 7 |
Our base pair-augmented auxiliary-loss function components are described below. At the 3D level, we optimize the predicted distance between the annotated base pairs (m, n) extracted from all three base pair annotation methods, termed the bp3D loss, as follows:
| 8 |
| 9 |
where is the set of annotated base pairs from the ith annotation method, is the number of annotations in , Dmn denotes the ground-truth Euclidean distance between the C1′ atoms of the nucleotide pair (m, n) and is the corresponding predicted distance. At the 2D level, we utilize an additional head to reduce the dimensionality of the predicted pair features () and compute the BCEWithLogitsLoss against the three experimental 2D maps (SS), termed the bp2D loss, as described below:
| 10 |
Our final combined weighted loss function is as follows:
| 11 |
We train our model in PyTorch-Lightning using the Adam optimizer with a learning rate of 0.0001. The distributed training process runs on eight NVIDIA H100 GPUs for 1,500 epochs, taking approximately 36 h to train on the cross-distillation augmented training set. We further fine-tune our trained model on predicted base pair maps from three pseudoknot-aware RNA 2D-structure predictors: IPKnot44, SPOT-RNA45 and RibonanzaNet46 for 200 epochs with a learning rate of 0.00001, utilizing the same loss functions, to mitigate the performance gap between noisy and accurate base pair information.
Training and benchmark datasets
To develop and benchmark our method RNAbpFlow, we utilize separate training sets and corresponding nonredundant benchmark sets to prevent data leakage between training and evaluation and therefore use dataset-specific checkpoints for each benchmark instead of a single default model. For the primary method development and internal benchmarking, we use the recently published RNA3DB dataset42, which is both sequentially and structurally nonredundant, making it highly suitable for training and benchmarking deep learning models. We use the version of RNA3DB parsed from the PDB12 on 26 April 2024 and select a representative set of RNA sequences from the original train–test split they provide for our purposes. To ensure that the dataset contains only high-quality native structures, we apply several filtering steps such as excluding structures with only one atom per nucleotide, removing protein residues from RNA structures, extracting contiguous sequences from experimental structures to address mismatches between the provided FASTA sequences and the experimental 3D structures for preserving base-pairing integrity and excluding sequences lacking base pairs in their corresponding native structures. To identify RNA chains with potentially truncated or poorly annotated regions, we extract base pairs from the corresponding experimental structures using RNAView36 and filter out chains containing long consecutive unpaired segments (≥20 nucleotides) to avoid base pair inconsistency. Finally, we only retain sequences with a minimum length of 30 and a maximum length of 200 to ensure efficient training for this experiment. This results in a clean training set consisting of 560 RNA sequences (excluding component #0, which comprises sequences with no hit to any Rfam family at an E-value threshold of 1.0, reducing the chance of data leakage) and 48 test sequences for benchmarking our method development. Both training and test sets originate from RNA3DB and follow the nonredundant splitting strategy defined in the original publication42, which ensures no sequence or structural overlap between the training and test sets.
To compare RNAbpFlow against RNAJP35, a recent MD simulation-based method for RNA 3D-structure sampling based on a given 2D structure, we use the dataset of 22 RNAs containing three-way junctions as used in the original RNAJP study. We exclude multimeter structures from the dataset and apply CD-HIT-est-2D between the remaining sequences and our training set of 560 RNA sequences from RNA3DB to avoid redundancy. This reduces the benchmark set to 12 RNAs (Supplementary Table 19), for which we download the predicted all-atom decoy structures generated by RNAJP from their publicly available repository at https://rna.physics.missouri.edu/RNAJP/index.html.
To enforce a strictly blind evaluation on the CASP15 and CASP16 benchmarks and ensure fair comparison with existing methods, we avoid the RNA3DB-provided train–test split to avoid data leakage. Instead, we curate two new, nonoverlapping training sets from RNA3DB following the recent deep learning-based RNA 3D-structure prediction methods13,14,19,22. For CASP15 blind benchmarking, we curate a training set by including RNA chains released before April 2022 in the PDB12 present in the entire RNA3DB dataset. As the release date of the first target in the CASP15 challenge was May 2022, this guarantees that the CASP15 test targets were published more recently than any targets available in the training set. We apply the same rigorous filtering steps described previously, resulting in a training set of 731 sequences with lengths ranging from 30 to 784 nucleotides. We evaluate the performance of RNAbpFlow on all six natural RNAs and four synthetic RNAs in the CASP15 benchmark set. Similarly, to benchmark the performance of RNAbpFlow on the CASP16 targets, we merge all components provided by RNA3DB into a single training set to train RNAbpFlow. The first CASP16 target was released in May 2024, whereas the RNA3DB training set contains PDBs released on or before 6 April 2024, thus emulating a strictly blind experiment as followed by other participating groups in the CASP16 challenge. Following the filtering procedure described above, we obtain a total of 994 RNA sequences and their corresponding experimental 3D structures. For the blind CASP16 evaluation, we collect the 28 available experimental RNA 3D structures, a list of which is presented in Supplementary Table 1.
For the CASP16 blind experiment, we additionally curate a cross-distillation set from bpRNA-1m (90)52 to investigate the effect of data augmentation. To this end, we first obtain the bpRNA-1m(90) set52, which comprises 28,370 sequences (and their corresponding secondary structures), as originally curated by the authors after removing any sequences sharing >90% identity over at least 70% of their length. We then filter this set by sequence length to retain only those sequences between 30 and 200 nucleotides. To further reduce redundancy, we apply MMseqs2 clustering59 with a 20% identity threshold, yielding 11,048 representative sequences. Next, we remove any of these 11,048 sequences that share >20% identity with the combined RNA3DB set of 994 training sequences and all CASP16 targets, again using MMseqs2, resulting in 10,107 nonredundant sequences. From this pool, we exclude sequences whose bpRNA-1m(90)-provided secondary structures contain <10 base pairs, reducing the set to 9,289 sequences. For data augmentation via cross-distillation during training, we predict structures for these 9,289 sequences using AF328 (Docker version) following recent literature22. We do not use any MSA information during AF3 prediction to generate the cross-distillation dataset for augmentation to ensure that the training of our method RNAbpFlow remains free from the influence of any evolutionary information. To maximize the estimated confidence scores of the AF3-predicted structures in the augmented training set, we filter the predictions using the self-assessment confidence measures predicted local distance different test (plDDT) and predicted aligned error (PAE) of AF3. Specifically, we only retain the predictions when plDDT ≥60 and PAE ≤15, resulting in a final cross-distillation set of 2,170 high-confidence RNA structures. Overall, the training set comprises 994 experimental PDB structures and 2,170 high-confidence predicted structural models for data augmentation via cross-distillation (for a combined total of 3,164), maintaining an approximate 1:2.2 ratio of PDB to cross-distillation data within each training batch22.
Performance evaluation and competing methods
To evaluate the performance of our method, RNAbpFlow, compared with other state-of-the-art 3D-structure sampling and prediction methods for RNA, we use a wide range of evaluation metrics. To evaluate global fold similarity, we calculate the TM-score (based on C3′) using US-align39 and global distance test–total score (GDT-TS) (based on C4′) using the LGA program for RNA60. For the assessment of local environment fitness, we compute lDDT40 using OpenStructure61 version 2.8 and the stereochemical quality metrics such as clashscore and bond length or bond angle violations metric using the MolProbity package62. Furthermore, we utilize the CASP-RNA pipeline41 to (1) evaluate the full atomic structural accuracy by calculating the all-atom RMSD and (2) compute INF variants for Watson–Crick, NWC and stacking interactions55, which quantify how well a structural model reproduces the base interactions present in the experimentally resolved reference structure.
As our method RNAbpFlow relies purely on a sequence- and base pair-conditioned SE(3)-equivariant flow-matching model for RNA 3D structure generation without utilizing MSAs or template information, we provide no MSA information to the competing deep learning-based methods for the sake of a fair performance comparison. For the CASP15 simulated blind assessment, we compare against MSA-agnostic deep learning-based methods such as DRfold13, which integrates end-to-end and geometric potentials and GraphaRNA63, a denoising diffusion method that uses graph representation with base pair information as input to predict coarse-grained RNA 3D structures. In terms of MSA-dependent methods, the comparison includes methods such as RhoFold+21, which leverages language model information, NuFold22, which uses a flexible nucleobase center representation with predicted base pair as input and trRosettaRNA14, a transformer-based architecture that predicts 1D/2D geometric restraints and generates full-atom RNA 3D models through restraint-guided energy minimization. All of these five methods are installed and run locally with their default parameters and 2D-structure prediction pipeline without any MSA information. In addition to these deep learning-based approaches, we compare against four physics- and/or knowledge-based methods: RNAComposer10, 3dRNA64, Vfold-Pipeline65 and SimRNA66. RNAComposer, 3dRNA and SimRNA are accessed via their respective web servers, while Vfold-Pipeline is installed and run locally. For all of these four methods and GraphaRNA, base-pairing information is provided in the required DBN format, as extracted from the experimental 3D structures by DSSR software (the only RNA 3D-structure annotator that outputs 2D structure in DBN format). We use any default postprediction optimization or refinement procedures for structure prediction across all competing methods if available. For the CASP16 simulated blind test, we compare RNAbpFlow against recent state-of-the-art diffusion-based biomolecular prediction method AF328 and three other AlphaFold 2-inspired RNA 3D-structure prediction methods DRfold248, NuFold22 and trRosettaRNA219 without any MSA information. DRfold2 is a sequence-only ab initio RNA 3D-structure predictor that combines a pretrained RNA composite language model with an end-to-end structure module, and trRosettaRNA2 is an end-to-end RNA 3D-structure predictor that uses a pretrained secondary-structure module and structure-aware attention. Both methods are installed and run locally with their default parameters. We additionally benchmark RNAbpFlow against two leading MSA-dependent server methods, AF3-server and Yang-Server, on the basis of their CASP16 official submissions.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Online content
Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41592-026-03128-4.
Supplementary information
Supplementary Results 1–3, Tables 1–22 and Figs. 1–3.
Source data
Comparison with RNAJP. Per-target maximum and mean results.
CASP16 benchmarking and case study. Per-target maximum and mean results.
Ablation study. Per-target maximum and mean results.
CASP15 benchmarking. Per-target evaluation metrics.
Effect of fine-tuning and cross-distillation. Per-target maximum and mean results.
Acknowledgements
This work was partially supported by the National Institute of General Medical Sciences (grant no. 2R35GM138146 to D.B.) and the National Science Foundation (grant no. DBI2208679 to D.B.).
Author contributions
D.B. conceived the study. S.T. processed the data, carried out the model training, conducted the experiments and generated the results. D.B. supervised the study. S.T. and D.B. analyzed the results. S.T. and D.B. wrote the initial paper. S.T. generated the initial figures. All authors discussed the results and contributed to the final paper.
Peer review
Peer review information
Nature Methods thanks Marta Szachniukand the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Primary Handling Editor: Arunima Singh, in collaboration with the Nature Methods team.
Data availability
The raw data used in this study, including the datasets for training and testing, are collected from publicly available sources. For training RNAbpFlow and method development, we downloaded RNA 3D structures and corresponding FASTA sequences from RNA3DB via GitHub at https://github.com/marcellszi/rna3db. Base pair annotations were extracted from training-set RNA 3D structures using three tools: RNAView (https://github.com/rcsb/RNAView), MC-Annotate (https://major.iric.ca/MajorLabEn/MC-Tools.html) and DSSR (https://x3dna.org/). For data augmentation with cross-distillation during training, we used the bpRNA-1m (90) dataset (https://bprna.cgrb.oregonstate.edu/download.php#bpRNA-1m(90)), which provides FASTA sequences and experimentally derived base pair information. During fine-tuning and inference, predicted base pair maps were obtained using three public predictors: RibonanzaNet (https://github.com/Shujun-He/RibonanzaNet), SPOT-RNA (https://github.com/jaswindersingh2/SPOT-RNA) and IPKnot (https://github.com/satoken/ipknot). For comparisons with RNAJP, we downloaded the RNAJP decoy set from https://rna.physics.missouri.edu/RNAJP/index.html and the corresponding native RNA 3D structures from the Protein Data Bank (PDB) at https://www.rcsb.org/. The experimental structures for CASP15 and CASP16 targets were obtained from the official Protein Structure Prediction Center website at https://predictioncenter.org/download_area/, with permission from the CASP organizers for targets cleared for public access. Training datasets are available via Zenodo at 10.5281/zenodo.19388910 (ref. 67). Source data are provided with this paper.
Code availability
An open-source software implementation of RNAbpFlow, licensed under the GNU General Public License v3, is freely available via GitHub at https://github.com/Bhattacharya-Lab/RNAbpFlow. Training and inference scripts for RNAbpFlow are available via Zenodo at 10.5281/zenodo.19388910 (ref. 67).
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s41592-026-03128-4.
References
- 1.Kulkarni, J. A. et al. The current landscape of nucleic acid therapeutics. Nat. Nanotechnol.16, 630–643 (2021). [DOI] [PubMed] [Google Scholar]
- 2.Sheridan, C. First small-molecule drug targeting RNA gains momentum. Nat. Biotechnol.39, 6–9 (2021). [DOI] [PubMed] [Google Scholar]
- 3.Liu, D., Thélot, F. A., Piccirilli, J. A., Liao, M. & Yin, P. Sub-3-Å cryo-EM structure of RNA enabled by engineered homomeric self-assembly. Nat. Methods19, 576–585 (2022). [DOI] [PubMed] [Google Scholar]
- 4.Warner, K. D., Hajdin, C. E. & Weeks, K. M. Principles for targeting RNA with drug-like small molecules. Nat. Rev. Drug Discov.17, 547–558 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Li, Y., Arce, A., Lucci, T., Rasmussen, R. A. & Lucks, J. B. Dynamic RNA synthetic biology: new principles, practices and potential. RNA Biol.20, 817–829 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Rother, M., Rother, K., Puton, T. & Bujnicki, J. M. ModeRNA: a tool for comparative modeling of RNA 3D structure. Nucleic Acids Res.39, 4007–4022 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Flores, S. C., Wan, Y., Russell, R. & Altman, R. B. Predicting RNA structure by multiple template homology modeling. Pac. Symp. Biocomput.15, 216–227 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Watkins, A. M., Rangan, R. & Das, R. FARFAR2: improved de novo Rosetta prediction of complex global RNA folds. Structure28, 963–976 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zhang, Y., Wang, J. & Xiao, Y. 3dRNA: building RNA 3D structure with improved template library. Comput. Struct. Biotechnol. J.18, 2416–2423 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Popenda, M. et al. Automated 3D structure composition for large RNAs. Nucleic Acids Res.40, e112 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Cao, S. & Chen, S.-J. Physics-based de novo prediction of RNA 3D structures. J. Phys. Chem. B115, 4216–4226 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Berman, H. M. et al. The Protein Data Bank. Nucleic Acids Res.28, 235–242 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Li, Y. et al. Integrating end-to-end learning with deep geometrical potentials for ab initio RNA structure prediction. Nat. Commun.14, 5745 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Wang, W. et al. trRosettaRNA: automated prediction of RNA 3D structure with transformer network. Nat. Commun.14, 7266 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Miao, Z. et al. RNA-Puzzles Round IV: 3D structure predictions of four ribozymes and two aptamers. RNA26, 982–995 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Sarzynska, J., Popenda, M., Antczak, M. & Szachniuk, M. RNA tertiary structure prediction using RNAComposer in CASP15. Proteins91, 1790–1799 (2023). [DOI] [PubMed] [Google Scholar]
- 17.Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Moussad, B., Roche, R. & Bhattacharya, D. The transformative power of transformers in protein structure prediction. Proc. Natl Acad. Sci. USA120, e2303499120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Wang, W., Peng, Z. & Yang, J. Predicting RNA 3D structure and conformers using a pre-trained secondary structure model and structure-aware attention. Nat. Mach. Intell.8, 722–734 (2026).
- 20.Baek, M. et al. Accurate prediction of protein–nucleic acid complexes using RoseTTAFoldNA. Nat. Methods21, 117–121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Shen, T. et al. Accurate RNA 3D structure prediction using a language model-based deep learning approach. Nat. Methods21, 2287–2298 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Kagaya, Y. et al. NuFold: end-to-end approach for RNA tertiary structure prediction with flexible nucleobase center representation. Nat. Commun.16, 881 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Tarafder, S., Roche, R. & Bhattacharya, D. The landscape of RNA 3D structure modeling with transformer networks. Biol. Methods Protoc.9, bpae047 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Zhang, Y. et al. Multiple sequence alignment-based RNA language model and its application to structural inference. Nucleic Acids Res.52, e3 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Vicens, Q. & Kieft, J. S. Thoughts on how to think (and talk) about RNA structure. Proc. Natl Acad. Sci. USA119, e2112677119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kryshtafovych, A. et al. New prediction categories in CASP15. Proteins91, 1550–1557 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Ganser, L. R., Kelly, M. L., Herschlag, D. & Al-Hashimi, H. M. The roles of structural dynamics in the cellular functions of RNAs. Nat. Rev. Mol. Cell Biol.20, 474–489 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature630, 493–500 (2024). [DOI] [PMC free article] [PubMed]
- 29.Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. Preprint at bioRxiv10.1101/2025.06.14.659707 (2025).
- 30.Chai Discovery team et al. Chai-1: decoding the molecular interactions of life. Preprint at bioRxiv10.1101/2024.10.10.615955 (2024).
- 31.Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature620, 1089–1100 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Yim, J. et al. SE(3) diffusion model with application to protein backbone generation. In Proc. 40th International Conference on Machine Learning (eds Krause, A. et al.) 40001–40039 (PMLR, 2023).
- 33.Yim, J. et al. Fast protein backbone generation with SE(3) flow matching. Preprint at https://arxiv.org/abs/2310.05297 (2023).
- 34.Anand, R. et al. RNA-FrameFlow: flow matching for de novo 3D RNA Backbone Design. Trans. Mach. Learn. Res.2025, 1–27 (2025).
- 35.Li, J. & Chen, S.-J. RNAJP: enhanced RNA 3D structure predictions with non-canonical interactions and global topology sampling. Nucleic Acids Res.51, 3341–3356 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Yang, H. et al. Tools for the automatic identification and classification of RNA base pairs. Nucleic Acids Res.31, 3450–3460 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Gendron, P., Lemieux, S.& Major, F. Quantitative analysis of nucleic acid three-dimensional structures. J. Mol. Biol.308, 919–936 (2001). [DOI] [PubMed] [Google Scholar]
- 38.Lu, X.-J., Bussemaker, H. J. & Olson, W. K. DSSR: an integrated software tool for dissecting the spatial structure of RNA. Nucleic Acids Res.43, e142 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Zhang, C., Shine, M., Pyle, A. M.& Zhang, Y. US-align: universal structure alignments of proteins, nucleic acids, and macromolecular complexes. Nat. Methods19, 1109–1115 (2022). [DOI] [PubMed] [Google Scholar]
- 40.Mariani, V., Biasini, M., Barbato, A. & Schwede, T. lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics29, 2722–2728 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Das, R. et al. Assessment of three-dimensional RNA structure prediction in CASP15. Proteins91, 1747–1770 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Szikszai, M. et al. RNA3DB: a structurally-dissimilar dataset split for training and benchmarking deep learning models for RNA structure prediction. J. Mol. Biol.436, 168552 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Tarafder, S. & Bhattacharya, D. lociPARSE: a locality-aware invariant point attention model for scoring RNA 3D structures. J. Chem. Inf. Model.64, 8655–8664 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Sato, K., Kato, Y., Hamada, M., Akutsu, T. & Asai, K. IPknot: fast and accurate prediction of RNA secondary structures with pseudoknots using integer programming. Bioinformatics27, i85–i93 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Singh, J., Hanson, J., Paliwal, K. & Zhou, Y. RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning. Nat. Commun.10, 5407 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.He, S. et al. Ribonanza: deep learning of RNA structure through dual crowdsourcing. Preprint at bioRxiv10.1101/2024.02.24.581671 (2024).
- 47.Chaudhury, S., Lyskov, S. & Gray, J. J. PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta. Bioinformatics26, 689–691 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Li, Y. et al. DRfold2 is a deep learning-based tool that enables efficient and accurate RNA structure prediction. PLoS Biol.24, e3003659 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Zhang, S., Li, J., Zhou, Y. & Chen, S.-J. Enhancing RNA 3D structure prediction in CASP16: integrating physics-based modeling with machine learning for improved predictions. Proteins94, 239–248 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Kagaya, Y. et al. Structure modeling protocols for protein multimer and RNA in CASP16 with enhanced MSAs, model ranking, and deep learning. Proteins94, 167–182 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Kretsch, R. C. et al. Assessment of nucleic acid structure prediction in CASP16. Proteins94, 192–217 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Danaee, P. et al. bpRNA: large-scale automated annotation and analysis of RNA secondary structure. Nucleic Acids Res.46, 5381–5394 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Menzel, P., Gorodkin, J. A. N. & Stadler, P. F. The tedious task of finding homologous noncoding RNA genes. RNA15, 2075–2082 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Pielesiak, J. et al. RNApdbee 3.0: a unified web server for comprehensive rna secondary structure annotation from 3D coordinates. J. Mol. Biol. 169795 (2026). [DOI] [PubMed]
- 55.Parisien, M., Cruz, J. A., Westhof, É. & Major, F. New metrics for comparing and assessing discrepancies between RNA 3D structures and models. RNA15, 1875–1885 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M. & Le, M. Flow matching for generative modeling. In 11th International Conference on Learning Representations (ICLR, 2023).
- 57.Chen, R. T. Q. & Lipman, Y. Flow matching on general geometries. In 12th International Conference on Learning Representations (ICLR, 2024).
- 58.Nikolayev, D. I. & Savyolov, T. I. Normal distribution on the rotation group SO(3). Texture Stress Microstructure29, 201–233 (1997). [Google Scholar]
- 59.Steinegger, M. & Söding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol.35, 1026–1028 (2017). [DOI] [PubMed] [Google Scholar]
- 60.Zemla, A. LGA: a method for finding 3D similarities in protein structures. Nucleic Acids Res.31, 3370–3374 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Biasini, M. et al. OpenStructure: an integrated software framework for computational structural biology. Acta Crystallogr. D69, 701–709 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Chen, V. B. et al. MolProbity: all-atom structure validation for macromolecular crystallography. Acta Crystallogr. D66, 12–21 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Justyna, M., Zirbel, C., Antczak, M. & Szachniuk, M. Graph neural network and diffusion model for modeling RNA interatomic interactions. Bioinformatics41, btaf515 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Wang, J., Wang, J., Huang, Y. & Xiao, Y. 3dRNA v2.0: an updated web server for RNA 3D structure prediction. Int. J. Mol. Sci.20, 4116 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Li, J., Zhang, S., Zhang, D. & Chen, S.-J. Vfold-Pipeline: a web server for RNA 3D structure prediction from sequences. Bioinformatics38, 4042–4043 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Moafinejad, S. N. et al. SimRNAweb v2.0: a web server for RNA folding simulations and 3D structure modeling, with optional restraints and enhanced analysis of folding trajectories. Nucleic Acids Res.52, W368–W373 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Tarafder, S. & Bhattacharya, D. RNAbpFlow training data and code. Zenodo10.5281/zenodo.19388910 (2026).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Results 1–3, Tables 1–22 and Figs. 1–3.
Comparison with RNAJP. Per-target maximum and mean results.
CASP16 benchmarking and case study. Per-target maximum and mean results.
Ablation study. Per-target maximum and mean results.
CASP15 benchmarking. Per-target evaluation metrics.
Effect of fine-tuning and cross-distillation. Per-target maximum and mean results.
Data Availability Statement
The raw data used in this study, including the datasets for training and testing, are collected from publicly available sources. For training RNAbpFlow and method development, we downloaded RNA 3D structures and corresponding FASTA sequences from RNA3DB via GitHub at https://github.com/marcellszi/rna3db. Base pair annotations were extracted from training-set RNA 3D structures using three tools: RNAView (https://github.com/rcsb/RNAView), MC-Annotate (https://major.iric.ca/MajorLabEn/MC-Tools.html) and DSSR (https://x3dna.org/). For data augmentation with cross-distillation during training, we used the bpRNA-1m (90) dataset (https://bprna.cgrb.oregonstate.edu/download.php#bpRNA-1m(90)), which provides FASTA sequences and experimentally derived base pair information. During fine-tuning and inference, predicted base pair maps were obtained using three public predictors: RibonanzaNet (https://github.com/Shujun-He/RibonanzaNet), SPOT-RNA (https://github.com/jaswindersingh2/SPOT-RNA) and IPKnot (https://github.com/satoken/ipknot). For comparisons with RNAJP, we downloaded the RNAJP decoy set from https://rna.physics.missouri.edu/RNAJP/index.html and the corresponding native RNA 3D structures from the Protein Data Bank (PDB) at https://www.rcsb.org/. The experimental structures for CASP15 and CASP16 targets were obtained from the official Protein Structure Prediction Center website at https://predictioncenter.org/download_area/, with permission from the CASP organizers for targets cleared for public access. Training datasets are available via Zenodo at 10.5281/zenodo.19388910 (ref. 67). Source data are provided with this paper.
An open-source software implementation of RNAbpFlow, licensed under the GNU General Public License v3, is freely available via GitHub at https://github.com/Bhattacharya-Lab/RNAbpFlow. Training and inference scripts for RNAbpFlow are available via Zenodo at 10.5281/zenodo.19388910 (ref. 67).




