Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Jul 9;17:8484. doi: 10.1038/s41467-026-75395-w

Deep Learning Predicts Dissimilar DNA-DNA Binding and Engineers Hyperconnected Networks

Karishma Matange 1,#, Gunavaran Brihadiswaran 2,#, Kyle J Tomek 1, Kevin Volkel 2, Doug Townsend 2, James M Tuck 2,, Albert J Keung 1,
PMCID: PMC13478463  PMID: 42426013

Abstract

Common frameworks in molecular bioengineering and synthetic biology focus on orthogonality, viewing weak or non-specific interactions as problems to avoid. This constrains the usable sequence space, limits scalability, and neglects scenarios where synthetic systems must operate within natural backgrounds of high sequence diversity. Harnessing the full space is difficult because models are lacking that can accurately and quickly predict non-orthogonal interactions and be validated against ground truth data. Here we develop BINND — Binding and Interaction Neural Network for DNA — using DNA-DNA interactions as a testbed. BINND combines an ultra-high throughput platform measuring millions of interactions with a deep learning model attaining accuracies above 80%, generalizing across diverse sequences and running 50 times faster than current models. We demonstrate its value with a searchable DNA network of fictitious storybook characters. BINND enables accurate prediction for diagnostics, bioengineering, and DNA origami, supporting a shift toward exploiting the full sequence space.

Subject terms: DNA computing, Molecular engineering, Computational models, Chemical engineering


Weak or non-specific DNA-DNA interactions can confound system orthogonality in synthetic biology. Here the authors develop BINND, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design.

Introduction

Exclusivity in pairwise molecular interactions, or orthogonality, confers benefits to engineering biological systems. It reduces noise, enhances predictability, and facilitates the design of Boolean logic. It has therefore been a dominant conceptual framework for bioengineering and synthetic biology1,2. However, many native systems, including chromatin regulation and cellular signaling, evolved messier multivalent networks, where each component can interact with several other distinct components3,4. Such systems exploit a greater proportion of the possible molecular sequence space, and if they could be synthetically designed and engineered, they could theoretically be more efficient and scalable as well as yield previously inaccessible emergent functions (Fig. 1A)510. For example, hyperconnected networks are characterized by high connectivity, redundancy, and robustness, and could drive emergent properties like signal amplification and computational processing (Fig. 1B). The growing interest in DNA-based data storage and computation might also benefit from sophisticated search1113 capabilities conferred by control over ‘messier’ molecular interactions.

Fig. 1. Visualization of the currently sparse usage of the DNA–DNA interaction space.

Fig. 1

A DNA sequence space represented in 3D, where the nodes (black) represent DNA sequences that are highly dissimilar to one another. The gray spheres represent the highly constrained local sequence space around each node, where existing models can make accurate predictions of pairwise interactions with the node. Currently, the nodes must be placed far apart to avoid unpredictable interactions. Such orthogonal engineering results in large gaps between nodes and wasted and unused sequence space. B (left) A bipartite plot of a binary system where each outer node connects to only one inner node. (right) A bipartite plot of a multivalent system where all outer nodes connect to all inner ones.

DNA-DNA interactions are an ideal testbed to explore the concept of engineering beyond orthogonality. Predicting interactions of highly similar DNA sequences is simple due to the rules of complementarity; this is often reflected in common statements of its programmability. Yet, off-target interactions, especially in contexts where there is a high background of sequence diversity, such as from genomes or DNA-based data storage databases, are a widely known and unsolved issue13,14. Current state-of-the-art prediction models spanning thermodynamic models1517, Levenshtein distance, Hamming distance, and GC content and melting temperature18,19, function well for highly similar sequences but drop dramatically in accuracy farther from perfect complementarity. While holistic ensemble modeling can account for competitive binding in complex mixtures, it is computationally intractable for large-scale libraries. Conversely, existing pairwise approximations, though more scalable, often fail to capture the subtle off-target landscape. Addressing these challenges would have a broad impact across applications, including in genomics research20, diagnostics21, molecular computation2224, DNA nanotechnology, including DNA origami and data storage25,26, and molecular biology27.

There are four major challenges to developing an accurate model able to predict outside the orthogonal sequence space: (1) Experimental datasets are small and typically sample sequence spaces that deviate only by a few nucleic acids2836. It is actually hard to comprehensively assess the accuracies of any model due to the lack of ground truth data; (2) hybridization is dependent on three-dimensional conformations, with DNA being very flexible and able to form diverse hairpin structures; (3) hybridization is dependent on dynamic processes that are hard to model especially in a computationally efficient manner37; (4) molecules with the same DNA sequence can adopt many different conformations, and thus a solution of identical DNA sequences will exist as a conformational ensemble. The combined challenge of the physical complexity of DNA–DNA hybridization, paired with the intractability of trying to directly model those mechanisms computationally, is well-suited for a machine learning approach to develop an accurate and fast predictive model. A key limitation is the lack of high-throughput ground truth experimental data that spans far beyond the orthogonal sequence space to capture highly dissimilar sequence pairings.

Here, we address these experimental and computational challenges through the generation of wet lab datasets that capture a broad, diverse, and deep experimental sequence space. In concert, we develop a robust and generalizable deep learning model, BINND - Binding and Interaction Neural Network for DNA, with strong performance on both standard and custom evaluation metrics, including a high area under the receiver operating characteristic (ROC) curve. We comprehensively evaluate the performance characteristics of BINND, including its speed, memory usage, error rates, and accuracy, and we benchmark them against state-of-the-art models. We further perform experiments to assess the generalizability and versatility of BINND to varying environmental conditions. Finally, we challenge the accuracy, power, and scalability of BINND to predict well outside of the orthogonal sequence space by designing and experimentally implementing a complex hyperconnected network representing fictitious children’s book characters. The hyperconnected network is searchable by character attributes and leverages densely overlapping but controlled interactions between highly dissimilar DNA sequences.

Results

Ultra-high-throughput platform robustly enriches hybridizing sequences

Our goal is to develop an accurate and computationally tractable predictive model of DNA hybridization that functions far from the orthogonal sequence space (Fig. 1). We focus here on ~20mer DNA sequences as they are commonly used in biotechnology as probes for diagnostics, hybridization sites for PCR and DNA origami or nanotechnology, and file addresses or computational motifs in DNA-based data storage and computation. With a theoretical parameter space comprising 420 distinct sequences, the first need is to develop a wet lab platform capable of detecting a large number of specific hybridization measurements from within an even larger pool of sequences. Conventional approaches to detect DNA hybridization, like spectrophotometry, microarrays, and melt-curves, do not scale adequately, typically limited to hundreds of measurable interactions at most. In contrast, next-generation sequencing (NGS) could provide measurement scales in the millions to tens of billions if two challenges can be solved: (1) NGS of experiments exploring multiple conditions is cost-prohibitive unless the sequencing is performed in a pooled fashion; (2) existing NGS formats cannot directly sequence short 20mer oligomers nor single-stranded DNAs.

We address these challenges in our design of a platform that measures the interactions of a DNA library consisting of all possible 20mer sequences (‘prey’) with a set of fixed 20mer sequences (‘bait’). To enable multiplexed and pooled sequencing readouts, the prey DNA comprises a fixed 40 base-pair (bp) double stranded DNA (dsDNA) domain followed by a 20 nucleotide (nt) randomized single stranded domain (ssDNA) (Fig. 2A). The dsDNA domain contains a 40 bp barcode sequence that indicates which experimental condition a sequencing read is obtained from, while the ssDNA region is comprised of random 20mer sequences created through phosphoramidite synthesis where all four phosphoramidites are present and can be incorporated in every synthesis cycle. Multiple prey libraries can be generated, with all DNAs within each prey library sharing an identical barcode sequence. The workflow is then to mix the prey library with biotinylated 20mer ssDNA baits composed of a single sequence. The mixture is then incubated with streptavidin magnetic beads to enrich for prey that bound to the bait. The enriched prey is then A-tailed to provide a universal ligation site for adapter oligos at the 3′ end. An extension reaction then generates a strand complementary to the template fragment. Adapter oligos are ligated to the 5′ end of the original strand before PCR adds additional indexing adapters to both ends and amplifies the sequences. These steps extend the length of the prey as well as convert the prey into fully dsDNA compatible with NGS.

Fig. 2. Ultra-high-throughput platform robustly enriches hybridizing sequences.

Fig. 2

A Schematic illustrating bait and prey architectures and the ultra-high-throughput experimental workflow to measure binding interactions. Colored boxes and bars indicate the DNA populations corresponding to measurements in the other panels. Created with BioRender.com. B Levenshtein distance between designed biotinylated Bait sequences. C A dot plot displaying the logarithmic ratio of the mass of bound or unbound prey to the mass of the initial prey library for each of the 26 baits. D Percent of prey that are observed to bind to multiple baits. E Levenshtein distance. 3.2 million bound sequences, 3.9 million unbound sequences, and all possible 20-mer sequences comprise the distributions, and the distributions are min-max scaled. F ΔG with bin sizes of 0.1. 3.2 million bound sequences, 3.9 million unbound sequences, and 1 million random 20mer sequences comprise the distributions, and the distributions are min-max scaled. G, H The enrichment experiment was repeated for 26 different biotinylated bait sequences. Heatmaps show the min-max scaled G Levenshtein distance for the bound, unbound, and initial prey sequences. H ΔG distributions for the bound, unbound, and initial prey sequences. Source data are provided as a Source Data file.

We first assess the robustness of this method. Twenty-six different bait sequences labeled A–Z are computationally designed to cover a broad sequence space and to be highly dissimilar from each other (Fig. 2B, Supplementary Table 1). Each bait is separately mixed with copies of the same prey library at equimolar amounts, magnetically enriched, and submitted for NGS. Three results are expected if the method is robust. First, with equimolar prey and bait, the mass of bound prey should be substantially lower than the mass of the total prey, as the majority of random 20mer sequences would be expected to be very dissimilar to, and hence not bind, the bait sequence. Indeed, this is what is observed, with less than 0.2% by mass of the input prey bound by the bait (Fig. 2C). Second, as the bait sequences are highly dissimilar, they should share very few bound prey sequences. Indeed, only 0.07% of the bound sequences appear in more than one dataset, with most of these appearing only in 2 bound datasets (Fig. 2D). This suggests the assay minimizes non-specific binding.

The third expected result is the most difficult to validate; it is whether the method actually enriches for prey that should bind to each bait, or if the assay is dominated by non-specific binding. As there is no existing experimental method or model that operates at high scale and accuracy, the closeness of the method to reflecting ground truth is difficult to validate, aside from demonstrating high accuracy in its end-stage performance. However, an intermediate step using conventional models, while imperfect, could be used to assess if the method is generally enriching prey as expected. For example, Levenshtein distances, Gibbs free energy differences (ΔG), and Hamming distances between each prey and the bait are expected to be lower for stronger binding partners. Simultaneously, this dataset provides the ability to perform a comprehensive assessment of the accuracy of these models. Indeed, the Levenshtein distance, ΔG and Hamming distance distributions of the bound prey shift lower as expected compared to the unbound and total distributions (Fig. 2E, F, Supplementary Fig. 1A). A similar trend in Levenshtein distance, ΔG and Hamming distance is observed when the assay is repeated on all 26 bait sequences (Fig. 2G, H, Supplementary Fig. 1B). A fluorophore-quencher hybridization assay provided further confidence in individual interaction events (Supplementary Fig. 2, Supplementary Data 1) indicating that the assay is enriching for prey with greater similarity to the bait, and that this enrichment is robust across a diverse sequence space.

To validate dataset representativity, we projected the sequence space into two dimensions by bifurcating 2-bit encoded 20-mers into 10-nt prefixes and suffixes. The resulting projections show broad occupancy across the coordinate space, indicating that the library recapitulates global sequence diversity without significant compositional bias (Supplementary Fig. 3A, B). This coverage extends to both bound and unbound libraries (Supplementary Fig. 3C, D). Further supporting this, nucleotide frequency distributions show that while unbound sequences mirror the initial library, bound sequences exhibit a distinct shift reflecting selective enrichment rather than synthesis bias (Supplementary Fig. 4). Finally, the Levenshtein distance and ΔG distributions (Fig. 2G, H) also confirm that the dataset spans the relevant sequence and thermodynamic landscapes required for generalizable learning.

BINND distinctly stratifies binding and non-binding sequences

When plotted as Levenshtein distance or ΔG distributions, there is substantial overlap between the bound and unbound populations (Fig. 2E, H, Supplementary Fig. 1). What is additionally problematic is that these metrics perform most poorly exactly where the majority of the sequence space is located, for example between Levenshtein distances of 8–16 and ΔG of −8 to −1 kcal/mol (Supplementary Fig. 5). This directly illustrates how conventional metrics and models are unable to distinctly stratify binding from non-binding sequences, especially in the regions between the orthogonal sequence spaces, and underscores the complexity of DNA interactions. Convolutional neural networks (CNN) could learn complex patterns and trends driving DNA interactions if provided sufficient data. Here, we develop a CNN-based classifier (Fig. 3A) and train it on our experimentally generated dataset using a high-performance computing cluster38. The dataset is characterized by high sparsity, with 92% of the unique sequences occurring as singletons, making frequency-based labels unreliable for regression (Supplementary Fig. 6). The CNN is trained on a balanced dataset of 144.6 million unique sequence pairs with an equal number of bound and unbound sequences, randomly split into training, validation, and test sets in an 8:1:1 ratio. At a fixed classification threshold of 0.5, the confusion matrix (Fig. 3B) yields a false-positive rate of 2.97% and a false-negative rate of 13.55%, corresponding to true-positive and true-negative rates of 36.45% and 47.03%, respectively. Reassuringly, the model achieves its highest accuracies at probabilities near 0 and 1 and its lowest accuracies around the 0.5 threshold (Fig. 3C). As a distinct metric of classification accuracy independent of any threshold, the area under the receiver operating characteristic curve (AUC) is 0.88, indicating strong overall discriminatory power (Fig. 3D).

Fig. 3. BINND distinctly stratifies binding and non-binding sequences.

Fig. 3

A Architecture of an 8-layer CNN binary classifier comprising five convolutional layers and three fully connected layers. Each convolutional layer is followed by ReLU activation, batch normalization, and dropout. Each sequence pair is one-hot encoded into a 4 × 40 matrix, which serves as the input to the model. B Confusion matrix showing the distribution of true positives, true negatives, false positives, and false negatives, evaluated on ~14.5 million sequence pairs from the test set. A prediction probability threshold of 0.5 is used to classify sequence pairs as bound or unbound. C Prediction probabilities are binned (bin size = 0.1). Prediction accuracy is measured as a ratio of correctly predicted outcomes to total predictions made for each probability bin. D ROC curve illustrates the trade-off between sensitivity and specificity. E Distributions of bound and unbound sequences as functions of BINND prediction probabilities, obtained using one bait. The distribution is plotted min-max scaled (bin size = 0.05). F Distributions of bound and unbound sequences as functions of BINND prediction probabilities. Each row corresponds to reads obtained from experiments using each of 26 baits. The heatmap is min–max scaled. Source data are provided as a Source Data file.

The discriminatory power of the CNN is further illustrated when the bound and unbound sequences are plotted as functions of the CNN prediction probability (Fig. 3E, F, Supplementary Fig. 7). The same sequences plotted in Fig. 3E, F plotted as functions of Levenshtein distance, ΔG, or Hamming distance exhibit considerable overlap between bound and unbound distributions (Fig. 2E–H, Supplementary Figs. 1 and 5). In contrast, the CNN generates substantially improved discrimination between bound and unbound populations with complete non-overlapping bifurcation of the populations for >90% of the sequences. Furthermore, the CNN performs considerably better than Levenshtein distance or ΔG in the region where the majority of the sequence space lies, specifically between prediction probabilities of 0.3–0 (Supplementary Fig. 7). It is important to note that low-probability bound sequences (0.1–0.2) do not exhibit clear positional nucleotide biases or distinct motifs relative to unbound sequences (Supplementary Fig. 8A). Thermodynamically, they occupy intermediate ΔG regimes rather than strongly favorable binding states (Supplementary Fig. 8B). A dedicated model trained on this subset showed near-random performance (AUC ~ 0.54), despite sufficient sequence diversity, indicating a lack of strong, learnable features. Together, these results suggest that the observed bimodality arises from regions of sequence space where binding is governed by subtle or higher-order interactions not fully captured by the current model, rather than clear adversarial patterns. Alternatively, these data arise from random and truly non-specific binding of some sequences during wetlab processing.

Two additional analyses provide confidence in the CNN model. First, the vast majority of the false CNN predictions lie in the region of ΔG overlap between bound and unbound sequences; furthermore, the false positives are shifted to lower ΔG relative to the false negatives, suggesting the CNN captures relevant features of binding from only sequence information (Supplementary Fig. 9). Second, BINND achieved an accuracy of 100% when tested on 1 million perfectly complementary sequence pairs of which none existed in the experimental (and hence training) dataset.

BINND was designed as a foundational training and benchmarking dataset without a focus on any specific sequence distributions. As such, it is not specifically tailored to represent the sequence distributions of native genomes. Designing bait sets to reflect different genomes could also be very expensive, as even highly reduced representations of genomic sequence spaces would take many baits to distinguish them from other sequence spaces, with some arbitrariness in establishing cutoffs also needed. However, the question of how BINND performs in a genomic context remains interesting and important. We therefore tested, in a limited way, how BINND performs in predicting binding to sequences that are more reflective of the human genome, as an example, and found similar performance for those sequences as compared to those derived from a random distribution (Supplementary Fig. 10).

BINND is faster and more accurate than state-of-the-art prediction models

Both natural and modern synthetic systems would benefit from the analysis of large numbers of pairwise interactions. This is not yet possible due to the high computational cost and error rates of current state-of-the-art nucleic acid design and analysis tools. To assess how BINND performs in speed and accuracy, we benchmark it against state-of-the-art predictors, including Levenshtein distance and thermodynamic (ΔG)-based models like NUPACK and Primer3. All models are equivalently tested using the identical experimental dataset comprising ~14.5 million sequence pairs. BINND outperforms all three models by at least 10% in classification accuracy (Fig. 4A, Supplementary Fig. 11). In addition, a CNN with the same architecture as BINND but trained on random sequences achieves an accuracy of 50%, as expected of a random classifier. A CNN trained on in silico data generated by classifying sequence pairs as bound or unbound using a ΔG threshold reaches an accuracy of 66%. In contrast, BINND obtains an 83% accuracy (Fig. 4A).

Fig. 4. BINND is faster and more accurate than state-of-the-art prediction models.

Fig. 4

A (top to bottom) Hamming distance model with a bound-unbound threshold of 14. Accuracy is measured as a ratio of correctly predicted outcomes to total predictions made; Levenshtein distance model with a bound-unbound threshold of 11; a model based on Primer3, with a bound-unbound threshold set at −6.0 kcal/mol; a model based on NUPACK, with a bound-unbound threshold set at −10.5 kcal/mol; BINND-Lite: a computationally efficient variant of BINND, trained on 26 million experimentally generated sequence pairs using the same 8:1:1 split; BINND; CNN trained and tested on 26 million in-silico generated sequence pairs, using an 8:1:1 split for training, validation, and testing. Binding labels (bound/unbound) are assigned based on a ΔG threshold of −10.5 kcal/mol; CNN trained and tested on a dataset of 26 million randomly generated sequence pairs, using an 8:1:1 split for training, validation, and testing, respectively. B Average execution time, average memory usage (GB), and peak memory usage (GB) are measured for each model. This is repeated three times, and the average values are plotted. Error bars indicate 95% confidence intervals (n = 3). C ROC curve illustrates the trade-off between sensitivity and specificity for BINND (AUC = 0.88), NUPACK-based model (AUC = 0.73), Primer3-based model (AUC = 0.79), Levenshtein distance-based model (AUC = 0.74), and a random classifier model (AUC = 0.5). Source data are provided as a Source Data file.

Owing to the inherently parallel nature of CNN computations, BINND exhibits substantial performance gains when run on a GPU, achieving 54.4× and 7.1× faster speeds over NUPACK and Primer3, respectively (Fig. 4B). Even when run on a CPU, BINND achieves comparable performance. Average and peak memory requirements remain at or below those of existing state-of-the-art binding prediction models. To further facilitate the utility of BINND in resource-constrained settings, BINND-Lite is a more compact variant that reduces model complexity. It trains 2.2× faster on a GPU and achieves 10× faster inference on a CPU compared to BINND, with only a small drop in accuracy. Consistent with these performance gains, ROC analysis demonstrates that BINND achieves superior discriminative capability (AUC = 0.88) compared to NUPACK, Primer3, and Levenshtein distance-based models, all of which perform closer to baseline relative to a random classifier (Fig. 4C).

BINND demonstrates data-efficient learning and generalizes well from limited data

The size of the full sequence space of 20mer oligos is just over 1 trillion (420). The dataset used to train BINND is just ~0.01% of that space, predominantly limited by sequencing costs. The accuracy it achieves despite this data sparsity demonstrates data-efficient learning. However, the training and test data are both derived using the same set of bait sequences. Thus, BINND generalizes well within the context of the bait sequences used, but it is unclear how well it generalizes to data obtained using bait sequences distinct from those used in the training dataset. To quantify its ability to generalize from limited data across diverse bait sequences, multiple models are trained on datasets derived from incrementally increasing numbers of distinct bait sequences (N = 1…25). The models are then tested on datasets obtained using the held-out baits (Fig. 5A). The AUC rises steeply from ~0.55 at N = 1 to a plateau of ~0.8 beginning at N = 11. This early plateau demonstrates that BINND rapidly captures the core features governing binding specificity and is able to predict binding to unseen bait sequences with high accuracy.

Fig. 5. BINND demonstrates data-efficient learning and generalizes well from limited data.

Fig. 5

A 1 million bound and 1 million unbound prey sequences are obtained from each pulldown using each of the 26 baits, A–Z, resulting in 52 million total sequences. BINND is trained on three randomly selected subsets of these 26 datasets, with the number of datasets in each subset indicated by the x-axis. BINND is then tested on the remaining data. The mean AUC of the three subsets is plotted with 95% confidence intervals. B BINND is trained on the bound and unbound data pooled from 25 of the datasets (50 million total sequences) and tested on the held-out dataset indicated on the x-axis to obtain AUCs. C BINND is trained on the dataset indicated on the x-axis and tested on the remaining data from the other 25 datasets to obtain AUCs. D Prey was pulled out using Bait A under varied salt concentrations: ‘low salt’ 0.05 M KCl; ‘medium salt’ 0.5 M KCl; ‘high salt’ 4 M KCl. Bound and unbound libraries sequenced from each condition are used to train three different BINND models. Each model is tested on data from all three conditions, and AUC values are plotted as a heatmap. In each case, the training dataset included 0.9 million bound and 0.9 million unbound sequences from the low, medium, or high salt datasets. 0.1 million bound and 0.1 million unbound sequences from the same salt condition as the training data are used for testing. Also included in the testing are the 1 million bound and 1 million unbound sequences obtained in the other two salt conditions. Source data are provided as a Source Data file.

It is possible that some bait datasets contribute to the AUC differently depending on some unique feature of their sequence or on the quality of the data itself. To assess the impact of each bait dataset, BINND is trained on 25 bait datasets and tested on the 26th, with this repeated for all 26 possible combinations of bait datasets. BINND achieves an AUC ≥ 0.80 for the majority of bait dataset combinations (Fig. 5B). This classification performance suggests model robustness across diverse sequence backgrounds. The converse experiment can also be performed where single bait datasets are used to train a model which is then used to predict binding across the remaining 25 datasets (Fig. 5C, Supplementary Fig. 12). The average AUC remains above the random-guess baseline of 0.50, indicating that even with extremely limited training data the model learns generalizable binding rules rather than overfitting to specific sequences (Supplementary Fig. 13).

The parameter space regulating DNA-DNA interactions is larger than just the sequence space; environmental parameters are known to alter these interactions. BINND is able to learn in distinct environmental contexts as well (Fig. 5D). When BINND is trained on data obtained in a low ionic strength buffer, it achieves strong but uniform classification performance for test datasets obtained in low, medium, and high ionic strength conditions, with AUC values of 0.88, 0.89, and 0.89, respectively. When trained on medium ionic strength datasets, the AUC for low ionic strength datasets drops by 2% while the AUCs for medium and high ionic strength datasets increase by 1% and 2%, respectively. When trained on a high ionic strength dataset, the AUC further drops by 1% when testing on low ionic strength datasets, while medium and high ionic strength AUCs remain the same and increase by 1%, respectively. These correlated shifts in performance highlight the ability of BINND to function in diverse assay environments.

BINND can design hyperconnected and multivalent interaction networks

Hyperconnected, multivalent interaction networks are common in natural biological systems. If they could be engineered, they could drive new classes of technologies such as DNA-based information storage and computation. These networks are characterized by: nodes (in this case, DNA molecules) binding to many other nodes, especially through non-redundant interactions; multiple dissimilar molecules interacting with the same target molecule, with this pattern propagating recursively through the network; high dimensionality and dense connections; and molecules participating in multiple diverse interactions. Such networks are highly efficient as they can reduce the total number of components needed to drive complex emergent functions through diverse feedback and feedforward network motifs39. However, creating such networks is hard. It requires the design of multiple interactions that are mutually dependent on each other, which is a formidable scaling challenge. Simultaneously, it requires creating attractive and repulsive forces between molecules that are, in some cases, similar and in other cases, dissimilar from each other, which requires very high accuracy and precision. BINND is accurate, fast, and generalizes well from limited data; these are the three characteristics precisely needed to address the challenge of designing a hyperconnected, multivalent interaction network.

To test the ability of BINND to create a hyperconnected and multivalent network, we use it to predict a matrix of hyperconnected interactions between the 26 bait sequences (A–Z) and 96 new prey sequences. As an illustrative analogy, consider the 96 prey sequences as children’s book characters, each characterized by a distinct combination of attributes. The 26 bait sequences represent these abstract attributes. Just as a character may simultaneously be a dreamer, brightly dressed, and a unicorn believer, a DNA sequence may hybridize with multiple baits designed using BINND. One could therefore search for all characters (prey sequences) that have a certain attribute (bait sequence) (Fig. 6A, Supplementary Data 2, Supplementary Fig. 14). The interaction matrix can also be visualized as comprising a zigzag pattern (Fig. 6B). The resulting interaction network is highly combinatorial and overlapping, reflecting the non-trivial relationships among features in hyperconnected systems. This is a highly challenging hyperconnected network to design for several reasons: it requires three mutually dissimilar baits to interact with the same prey (Fig. 6B inset), and to do so iteratively through all 26 baits (this is represented by one diagonal of the zigzag); it requires this recursive cycle to be repeated multiple times with distinct sets of prey (the 6 diagonals of the zigzag); there are baits that need to be designed to bind between 2 and 18 of the 96 prey (Supplementary Table 2); and there are substantial numbers of interactions that must remain non-binding, meaning that the system cannot simply be highly non-specific.

Fig. 6. BINND can design hyperconnected and multivalent interaction networks.

Fig. 6

A Hyperconnected character-attribute network is designed, where outer nodes represent the characters and inner nodes represent attributes. B The same network displayed in grid form. The inset illustrates the overlapping subsets of interactions that pose the challenge of designing hyperconnected networks. C Wetlab experiments are performed in triplicate for all 26 Bait sequences (A–Z) in equimolar amounts with a defined 96-prey library. BINND and Primer3 predictions are compared to the empirical dataset and plotted as heatmaps with true-positives as green, true-negatives as blue, false-positives as yellow, and false negatives as vermillion. D The results from C plotted as bipartite plots with true-positives as green, false-positives as yellow, true-negatives as blue, and false negatives as vermillion. E Scatter plot displaying data from all 26 experiments performed, relaying the percent frequency of a sequence of each 96-prey sequence in the sequencing data as a function of BINND prediction probability. Source data are provided as a Source Data file.

This hyperconnected network is implemented in the wet lab by using each biotinylated bait to perform 26 separate equimolar pulldowns from a pooled DNA mixture containing all 96 prey. The experimental results recapitulate much of the predicted zigzag pattern (Fig. 6C, Supplementary Fig. 15A, B) and network architecture (Fig. 6D, Supplementary Fig. 15C). The prediction probability of each prey also correlated monotonically with its enrichment (Fig. 6E, Supplementary Fig. 15D). As expected, compared to a randomer prey pool (Fig. 2A), the mass of prey pulled down from this specially designed 96-member pool was substantially greater, ranging from 18.6% to 58.9% (Supplementary Fig. 15E). BINND also outperformed Primer3 by over 100% in predicting bound sequences (Fig. 6C), a rate far greater than when comparing their prediction accuracies for solely binary interactions (Fig. 4A).

Discussion

Hyperconnected networks are ubiquitous. They are common in data architectures40, societal structures including social networks41, and in native biological systems like gene and protein signaling networks. They are an important example of the potential benefits gained by operating beyond constrained orthogonal sequence spaces. However, engineering synthetic hyperconnected networks using biological components is rare and challenging. This is in large part due to the limited interaction space that can be engineered, and therefore prior networks have been constrained to at most a handful of interconnected interactions designed through the use of combinatorial binding interfaces, where each node is composed of multiple distinct DNA sequences or protein domains fused together in order to effect interactions with multiple other nodes4246.

The present work addresses and implements three key features that unlock the engineering of hyperconnected networks. First, it creates a hyperconnected network in which each node is a single sequence, not a combination of multiple sequences, where each sequence is capable of interacting with multiple other sequences or nodes; this mimics what natural evolution has engineered in many natural settings and is a more efficient use of the molecular design space. Second, no single node is a dominant hub with all other nodes exhibiting limited interactions; rather, all nodes act as highly connected hubs. Third, it creates the numerically largest and most interconnected engineered network to date. These three properties are related in that efficient use of the DNA sequence space supports the efficient scaling of systems and networks, and this reflects the efficiency that many native biological systems have naturally evolved.

The rarity of engineered hyperconnected networks arises from the challenge of predicting and designing intermolecular interactions. BINND directly addresses this challenge and is the central advance of this work. At a superficial level, DNA-DNA interactions are often viewed and stated as highly programmable and predictable; yet, it becomes clear upon closer examination that this is not the case. Indeed, the common polymerase chain reaction and biosensing technologies, which both rely on DNA–DNA hybridization, exhibit high levels of non-specific interactions, particularly in the presence of background genomic or environmental DNA with high sequence diversity47,48, despite the use of state-of-the-art models to try and avoid them. This challenge will be exacerbated in even denser synthetic DNA storage and computation systems13,19,49. This has motivated the widespread approach of engineering only within highly orthogonal spaces, where sets of highly dissimilar sequences are used to create biomolecular systems. This approach constrains the complexity of engineered systems. It also cannot account for non-specific interactions with the sequence diversity of natural biological systems in which many engineered tools and sensors must operate. It is perhaps within this context that the accuracy of DNA–DNA prediction models has been viewed. While BINND outperforms thermodynamic models like NUPACK in conventional settings (Fig. 4), it is striking how BINND outperforms NUPACK by a substantially greater margin when applied within the constraints of a hyperconnected system. Greater accuracy is required when predicting interactions between sequences that are highly dissimilar, and the demands of a hyperconnected network best highlight the gain in accuracy BINND achieves. The ability of BINND to handle these cases represents a shift toward scalable, network-level control of nucleic acid hybridization.

A key gap in the development of accurate models has been the lack of experimental data to both develop and assess models. Most predictions of binding and quantifications of accuracy are generally limited to binary comparisons or comparisons within small sets of binding partners18,5053. They also often assess the impact of a few mutations rather than investigating the broader possible sequence space. The very nature of quantifying true accuracy was previously intractable, as the number and diversity of sequence pairs that could be measured were limited to less than 125 sequences24,2836. The dataset herein provided the ability to train BINND with 9 orders of magnitude greater numbers of sequences that also randomly sampled the entire sequence space; it simultaneously provided the ability to rigorously and comprehensively assess the accuracy of state-of-the-art models.

The theoretical complexity of the 20mer prey library is 4²⁰ (~10¹² sequences), while the experimental starting material provides an average molecular coverage of approximately 100 copies of each unique sequence, assuming idealized uniform DNA synthesis. In practice, oligonucleotide synthesis is known to be non-uniform, and some sequences may be underrepresented or absent. Due to the disparity between theoretical library complexity and achievable sequencing depth, the BINND dataset represents a sparse sampling of interaction space rather than exhaustive coverage. As a result, quantitative conclusions should be interpreted at the level of learned interaction features and statistical trends, rather than at the level of individual sequence completeness. BINND does not require exhaustive enumeration of all possible sequences; rather, it relies on large-scale sampling of interaction space to learn generalizable hybridization features. As a result, synthesis bias and partial sequence dropout are expected to have limited impact on model performance, particularly given the scale of the dataset and measurements across multiple bait sequences.

To assess accuracy and specificity, one needs to make and test predictions across a large sequence space. A corollary is that speed is required to carry out the scale of comparative calculations needed to assess accuracy and specificity. We were inspired by the speed and efficiency of machine learning models. For example, prior work trained a convolutional neural network using in-silico data predicted using computationally intensive thermodynamic algorithms54. This essentially converted thermodynamic in-silico predictions into a more efficient algorithm, but it remained unable to improve in accuracy at predicting ground truth measurements. Data were the major gap, and when it is obtained at scale, it improves the accuracy of machine learned models while leveraging their speed.

While Transformer-based architectures have become the standard for sequence modeling, we opted for a CNN-based approach to prioritize computational efficiency and scalability. Given the high-volume inference requirements of downstream applications like primer search, the lower memory footprint and accelerated inference of a CNN are critical for practical deployment. Comparative benchmarks against a RoBERTa-based transformer showed that while the latter achieved a negligible 0.1% increase in accuracy, it incurred a 13.5× higher training latency, totaling 61 h for a subset of the data.

All models in this study were trained and evaluated using in vitro interaction datasets, which provide controlled measurements of intrinsic DNA–DNA hybridization. This approach is consistent with related experimental paradigms, such as hybridization-based pulldown or capture assays used for genomic enrichment, which are likewise performed under controlled in vitro conditions. While the present work focuses on the development and validation of the BINND platform, the resulting models capture biologically meaningful principles of DNA hybridization. As such, they could be directly applied to biologically relevant tasks, including the identification of off-target primer binding, the design of highly specific diagnostic probes, and the evaluation of unintended hybridization in synthetic DNA systems. Furthermore, by enabling accurate prediction of DNA–DNA interactions at scale, BINND provides a foundation for future studies aimed at incorporating additional biological constraints, extending to in vivo conditions, and exploring organism-specific hybridization behavior.

There are two additional interesting and potentially related features of BINND. It was not initially clear whether a model trained on a narrow subset of sequence space (i.e., 26 highly dissimilar Baits) could generalize to the broader space of possible hybridization events. Our results show that BINND succeeds in doing so, suggesting that the constraints governing hybridization are sufficiently structured to be learned from a limited sequence space. This has significant implications for model development, as it enables accurate prediction without requiring prohibitively large training datasets. It also suggests that it may be tractable in the future to develop interpretable, or at least explainable, models.

In addition to its data-efficient learning, BINND maintained high accuracy in varied environmental conditions. Both attributes of BINND suggest its flexibility could be useful across a broad application space, including enrichment of specific loci or sequences from genomic material in preparation for next-generation sequencing, synthetic biology, biosensing, point-of-care diagnostics, pathogen and biomarker detection, designing PCR primers, implementing search and random access in DNA data storage systems, and engineering interactions for DNA computing and DNA origami. BINND-Lite also provides a minimal tradeoff in model size and accuracy that could be useful for real-time diagnostics or automated synthetic biology designs. The experimental parameters employed here (20mer sequences, hybridization-based pullouts) were selected for their broad applicability. Additional datasets could be obtained for specific applications, including with different sequence lengths or with specific molecular contexts, such as PCR, where the 3′ end would be disproportionately of relevance. To accommodate variable-length inputs, the current encoding scheme can be extended via zero-padding and masking up to a defined maximum sequence length. Furthermore, given the performance of BINND-Lite on a reduced dataset, we anticipate that a targeted expansion of approximately one million data points would be sufficient to adapt the model to extended input lengths without incurring prohibitive data overhead. Application to RNA and hybrid nucleic acid systems would broaden its relevance to gene regulation, therapeutics, and programmable molecular devices. In parallel, incorporating uncertainty-based experimental design could accelerate model improvement by focusing new data collection on regions of low confidence. Together, these extensions would further enhance BINND’s role as a generalizable and efficient tool for nucleic acid interaction design across diverse biological and engineering applications.

Methods

Prey DNA library preparation

DNA oligonucleotides were ordered from Eurofins Genomics. The randomer prey library was prepared by combining 183 pmol of 60-mer barcoded randomized ssDNA and 183 pmol of 40-mer ssDNA complementary to the barcode region of the prey library. The 96 prey library was prepared by combining 91.5 pmol of 60-mer ssDNA comprised of equimolar amounts of each sequence, 91.5 pmol of the randomized prey library, and 183 pmol of 40-mer ssDNA complementary to the barcode region of the prey library. 1 µl of 10× annealing buffer (100 mM Tris pH 8.0, 500 mM KCl, 10 mM EDTA, and distilled water) and the DNA mixtures were combined and brought up to 10 µl with distilled water. This mixture was incubated in a thermocycler at 90 °C for 2 min, followed by a gradual cooling to 4 °C at a rate of −1 °C per minute.

DNA hybridization assay

Prey libraries were mixed with 183 pmol of biotinylated Bait, 14 µl of 10× annealing buffer, and the volume was adjusted to 150 µl with distilled water. This mixture was annealed in a thermocycler by heating to 68 °C for 10 min, then the temperature was gradually reduced to 4 °C at a rate of −1 °C per minute.

Following the annealing step, this reaction mixture was incubated with 183 µl of NEB streptavidin magnetic beads for 30 min at room temperature. The beads were then separated using a magnetic stand, and the supernatant, containing the unbound prey, was collected and stored. The beads were washed once with distilled water before being reconstituted in 50 µl water. The mixture was then incubated at 68 °C for 10 min in a thermocycler. After incubation, the beads were placed on a magnetic stand, and the supernatant containing the bound prey was collected and stored at −20 °C.

The bound, unbound, and initial prey DNA were prepared for NGS using the IDT xGen ssDNA & Low-Input DNA prep kit (cat# 10009817) following the manufacturer’s instructions. 7 PCR cycles were performed in the indexing step. Equal masses of DNA (ng) from each sample were pooled prior to sequencing. NGS was performed by MedGenome, Inc., on an Illumina NovaSeq PE150

Bait sequence design

The 26 bait sequences were computationally designed using the following criteria.

  1. A minimum Hamming distance of ten was maintained between all pairs

  2. GC content was restricted to 40–60% (8–12 bases), and homopolymer repeats were limited to ≤3.

  3. Sequences were evaluated at 55 °C to ensure strong specific binding (specific ΔG between −11.8 and −9.5 kcal/mol) while strictly limiting non-specific interactions with the existing set (non-specific ΔG > −3.59 kcal/mol).

Experimental workflow timeline

The dataset comprising approximately 145 million sequences was generated across 28 distinct DNA assay conditions, each performed in triplicate. Each assay required ~6 h of hands-on time and could be partially parallelized; however, parallelization was constrained to preserve uniform incubation times and ensure data quality. In practice, all experiments were completed in ~7 batches over the course of approximately one week by experienced personnel. Subsequent sequencing was performed through commercial providers, with typical turnaround times of 4–6 weeks for large-scale runs. These practical considerations are discussed to guide researchers aiming to implement or build upon this approach.

Fluorescence–Quencher hybridization assay

The Initial prey sequences were prepared in the same way as in the pulldown assay. 7.53 pmol prey were mixed with an equimolar amount of Bait, 14 µl of 10× annealing buffer, and the volume was adjusted to 150 µl with distilled water. This mixture was annealed in a thermocycler by heating to 68 °C for 10 min, then the temperature was gradually reduced to 4 °C at a rate of −1 °C per minute. A Plate reader was used for fluorescence detection.

Sequencing data processing

Samples were sequenced through Illumina NovaSeq55. Sequencing data was processed (Supplementary Fig. 16) to get valid sequences (Supplementary Data 3 and 4)

Sequence alignments against the human reference genome56 were performed using the NCBI BLAST+, interfaced via the Biopython library (v2.16.0). Given the short length of our query sequences, the -task blastn-short parameter was utilized to optimize sensitivity for short-seed matches. To ensure high-confidence unique mapping, searches were restricted to the top alignment per query (-max_target_seqs 1 -max_hsps 1). The specific command utilized was: blastn -task blastn-short -query <query_file > -db <db_name > -outfmt 6 qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore -max_target_seqs 1 -max_hsps 1.

Model training and implementation

Dataset

The sequencing data were preprocessed by filtering for sequences of the desired length and matching barcodes (Supplementary Fig. 16). The dataset was then balanced to include an equal number of bound and unbound pairs, resulting in approximately 144.6 million samples57. These were split into training, validation, and test sets in an 8:1:1 ratio using a stratified split to preserve the class distribution across all subsets. While we utilized unique prey sequences, we further validated the absence of data leakage by calculating the Jaccard similarity between 1 million randomly selected prey sequence pairs from the training and testing sets. Using a 3-mer window (k = 3), we observed a mean Jaccard similarity of 0.16. This low similarity score confirms that the performance metrics reflect true predictive generalization rather than information leakage.

Data encoding

The input sequence pairs were first one-hot encoded. Let the first sequence be S1 = [S1,1,S1,2,…,S1,20] and second sequence be S2 = [S2,1,S2,2,…,S2,20]. To preserve the assumed proximity of nucleotides under perfect binding, the sequences were interleaved as follows: [S1,1,S2,20,S1,2,S2,19 …,S1,20,S2,1]. The current version of BINND accepts only 20 bp sequences, resulting in a 4 × 40 one-hot encoded input matrix.

Training and evaluation

BINND was implemented using the PyTorch (v2.6) framework58. The model architecture is illustrated in Fig. 3A, and it contains 4,467,969 trainable parameters. A detailed layer-wise summary of the architecture and parameter counts is provided in Supplementary Table 3. The hyperparameters were tuned using the Ray Tune implementation of the Asynchronous Successive Halving Algorithm (ASHA)59. The complete hyperparameter search space is described in Supplementary Table 4. The final training configuration is summarized in Supplementary Table 5.

All computations were performed on high-performance computing nodes equipped with dual Intel® Xeon® Gold 6226 R CPUs (32 total cores) and NVIDIA L40S GPUs. The mean of aggregated training and validation losses, computed over every 500 batches, was logged during training. These loss curves are provided in Supplementary Fig. 17. Early stopping with a patience of two epochs was employed to prevent overfitting and improve training efficiency. Accuracy, AUC, precision, recall, and confusion matrix were used to evaluate BINND performance.

BINND-Lite architecture and training

Designed as a compact variant, BINND-Lite features an architecture with fewer layers and a reduced parameter count (1,092,225) compared to the full BINND model. This design choice enables faster training and inference, which is particularly beneficial for experimental iterations or when operating under computational resource constraints. Its detailed layer-wise architecture and parameter summary are presented in Supplementary Table 6. For direct comparison, BINND-Lite was trained and evaluated utilizing the identical experimental setup, hyperparameters, and evaluation metrics employed for BINND.

Implementation details

Gibbs free energy calculation

ΔG values were calculated using the thermodynamic models implemented in NUPACK and Primer3. The implementations of both models are summarized below.

NUPACK Python library v 4.0.1.760

model=nupack.Model(material='dna', celsius=25, sodium=0.05, magnesium=0.0)seq_1 = nupack.Strand(seq1)seq_2 = nupack.Strand(seq2)complex = nupack.Complex([seq_1, seq_2])set_1 = nupack.ComplexSet(strands = [seq_1, seq_2],complexes=nupack.SetSpec(max_size=2))result=nupack.complex_analysis(complexes=set_1,model=model,compute = ['pfunc'])dg = result[complex].free_energy

Primer3 Python library v2.1.017

model = primer3.thermoanalysis.ThermoAnalysis(temp_c = 25,mv_conc=0.05 *1000, dv_conc=0.0*1000)result = model. calc_heterodimer(seq1, seq2)dG =result['dg']

Levenshtein distance calculation

Levenshtein distance was calculated using the python-Levenshtein 0.27.3 library61.

Statistics and reproducibility

No statistical method was used to predetermine sample size. Sample sizes were selected based on established practices in the field and practical considerations associated with experimental throughput and sequencing capacity. No data were excluded from the analyses. The experiments were not randomized. The investigators were not blinded to allocation during experiments and outcome assessment. Statistical analyses were performed as described in the corresponding figure legends and ‘Methods’ sections. All attempts at replication were successful.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

41467_2026_75395_MOESM2_ESM.pdf (175.6KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (10KB, xlsx)
Supplementary Data 2 (11.3KB, xlsx)
Supplementary Data 3 (11.3KB, xlsx)
Supplementary Data 4 (12KB, xlsx)
Reporting Summary (1.5MB, pdf)

Source data

Source Data (589.9KB, xlsx)

Acknowledgements

We acknowledge the computing resources provided by North Carolina State University High Performance Computing Services Core Facility (RRID:SCR_022168). Some artwork was created with BioRender.

Author contributions

K.J.T., K.V., D.T., J.M.T., and A.J.K. conceived the study. J.M.T. and A.J.K. acquired funding and supervised the project. K.M., G.B., J.M.T., and A.J.K. performed the investigation. K.M. planned and performed the wetlab experiments with guidance from A.J.K., G.B., and J.M.T., developed the software, curated the data, and carried out the formal analysis. K.M., G.B., J.M.T., and A.J.K. wrote the paper with input from all.

Peer review

Peer review information

Nature Communications thanks De-Shuang Huang, and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Funding

A.J.K. and J.M.T. disclose support for the research of this work from the National Science Foundation (ECCS-2027655, CSR-1901324, CSR-2403352). A.J.K., J.M.T., and K.J.T. disclose support for the research of this work from the National Institutes of Health (R41HG013877). K.J.T. discloses support for this work from a Department of Education Graduate Assistance in Areas of Need fellowship, P200A160061. A.J.K. discloses support for this work from the Simons Foundation (990252).

Data availability

The raw sequencing data files generated in this study have been deposited in the Zenodo database at 10.5281/zenodo.20494835. The dataset used to train and evaluate BINND has been deposited in the Zenodo database at 10.5281/zenodo.19500645. Source data for all other analyses are provided as a Source Data file. Source data are provided with this paper.

Code availability

The code used to develop the model, perform the analyses, and generate the results in this study is publicly available and has been deposited in the GitHub repository BINND at https://github.com/dna-storage/BINND, under the MIT license. The specific version of the code associated with this publication is archived in Zenodo and is accessible via 10.5281/zenodo.1948879558. The repository includes detailed installation instructions, a Dockerfile for containerized setup, and environment configuration files (BINND.yml and init.sh). External dependencies used for ΔG calculations, NUPACK (v4.0.1.7) and Primer3, are governed by their own licenses; NUPACK is not open-source.

Competing interests

K.J.T., J.M.T., and A.J.K. are co-founders of DNAli Data Technologies. The remaining authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Karishma Matange, Gunavaran Brihadiswaran

Contributor Information

James M. Tuck, Email: jtuck@ncsu.edu

Albert J. Keung, Email: ajkeung@ncsu.edu

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-75395-w.

References

  • 1.Nielsen, A. A. K. et al. Genetic circuit design automation. Science352, aac7341 (2016). [DOI] [PubMed] [Google Scholar]
  • 2.Siuti, P., Yazbek, J. & Lu, T. K. Synthetic circuits integrating logic and memory in living cells. Nat. Biotechnol.31, 448–452 (2013). [DOI] [PubMed] [Google Scholar]
  • 3.Erez, K., Jangid, A., Feldheim, O. N. & Friedlander, T. The role of promiscuous molecular recognition in the evolution of RNase-based self-incompatibility in plants. Nat. Commun.15, 49163 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Liu, X. et al. Robustness and lethality in multilayer biological molecular networks. Nat. Commun.11, 19841 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jencks, W. P. On the attribution and additivity of binding energies (proteins/ligands/entropy/enzymes). Proc. Natl. Acad. Sci. USA78, 4046–4050 (1981). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Schipler, A. & Iliakis, G. DNA double-strand-break complexity levels and their possible contributions to the probability for error-prone processing and repair pathway choice. Nucleic Acids Res. 41, 7589–7605 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Berger, S. L. The complex language of chromatin regulation during transcription. Nature447, 407–412 (2007). [DOI] [PubMed] [Google Scholar]
  • 8.Jeong, H., Tombor, B., Albert, R., Oltvai, Z. N. & Barabási, A.-L. The large-scale organization of metabolic networks. Nature407, 651–654 (2000). [DOI] [PubMed] [Google Scholar]
  • 9.Courtney, A. H., Lo, W.-L. & Weiss, A. TCR signaling: mechanisms of initiation and propagation. Trends Biochem. Sci.43, 108–123 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Barabási, A.-L. & Oltvai, Z. N. Network biology: understanding the cell’s functional organization. Nat. Rev. Genet.5, 101–113 (2004). [DOI] [PubMed] [Google Scholar]
  • 11.Imburgia, C. et al. Random access and semantic search in DNA data storage enabled by Cas9 and machine-guided design. Nat. Commun.16, 61264 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bee, C. et al. Molecular-level similarity search brings computing to DNA data storage. Nat. Commun.12, 24991 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Tomek, K. J., Volkel, K., Indermaur, E. W., Tuck, J. M. & Keung, A. J. Promiscuous molecules for smarter file operations in DNA-based data storage. Nat. Commun.12, 23669 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kalter, N. et al. Off-target effects in CRISPR-Cas genome editing for human therapeutics: progress and challenges. Mol. Ther. Nucleic Acids36, 102636 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Wolfe, B. R., Porubsky, N. J., Zadeh, J. N., Dirks, R. M. & Pierce, N. A. Constrained multistate sequence design for nucleic acid reaction pathway engineering. J. Am. Chem. Soc.139, 3134–3144 (2017). [DOI] [PubMed] [Google Scholar]
  • 16.Zhang, J. X. et al. Predicting DNA hybridization kinetics from sequence. Nat. Chem.10, 91–98 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Untergasser, A. et al. Primer3—new capabilities and interfaces. Nucleic Acids Res. 40, e115 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Xu, Q., Schlabach, M. R., Hannon, G. J., Elledge, S. J. & Watson, B. Design of 240,000 orthogonal 25mer DNA barcode probes. Proc. Natl. Acad. Sci. Usa.106, 2289–2294 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tomek, K. J. et al. Driving the scalability of DNA-based information storage systems. bioRxiv10.1101/591594 (2019). [DOI] [PubMed]
  • 20.Collins, R. L. & Talkowski, M. E. Diversity and consequences of structural variation in the human genome. Nat. Rev. Genet. 10.1038/s41576-024-00808-9 (2025). [DOI] [PubMed]
  • 21.Esmaeilpour, D., Ghomi, M., Zare, E. N. & Sillanpää, M. Recent advances in DNA nanotechnology for cancer detection and therapy: a review. Int. J. Biol. Macromol.292, 142136 (2025). [DOI] [PubMed] [Google Scholar]
  • 22.Zhang, M. & Han, D. Concept, development and applications of DNA computation. Fundam. Res. 10.1016/j.fmre.2023.06.015 (2023). [DOI] [PMC free article] [PubMed]
  • 23.Zhang, S. F., Li, Y. H., Zhang, R. X., Li, B. Z. & Wang, Q. Multi-mode data organization and file retrieval based on a primer library in large-scale digital DNA storage. Engineering10.1016/j.eng.2023.10.021 (2025).
  • 24.Bae, J. H., Fang, J. Z. & Zhang, D. Y. High-throughput methods for measuring DNA thermodynamics. Nucleic Acids Res.48, e89 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Ma, W. et al. The biological applications of DNA nanomaterials: current challenges and future directions. Signal Transduct. Target. Ther.6, 1 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Doricchi, A. et al. Emerging approaches to DNA data storage: challenges and prospects. ACS Nano16, 17552–17571 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Kumar, J., Choudhary, K., Anukriti, Singh, H., Kumar, A. & Bagga, P. Molecular approaches for authentication and identification of medicinal plants. Plant Mol. Biol. Rep. 10.1007/s11105-025-01557-7 (2025).
  • 28.Breslauert, K. J., Franks, R., Blockers, H. & Markyt, L. A. Predicting DNA duplex stability from the base sequence. Proc. Natl. Acad. Sci. USA83, 3746–3750 (1986). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.SantaLucia, J. A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proc. Natl. Acad. Sci. USA95, 1460–1465 (1998). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.SantaLucia, J. & Hicks, D. The thermodynamics of DNA structural motifs. Annu. Rev. Biophys. Biomol. Struct.33, 415–440 (2004). [DOI] [PubMed] [Google Scholar]
  • 31.SantaLucia, J., Allawi, H. T. & Seneviratne, P. A. Improved nearest-neighbor parameters for predicting DNA duplex stability. Biochemistry35, 3555–3562 (1996). [DOI] [PubMed] [Google Scholar]
  • 32.Allawi, H. T. & SantaLucia, J. Thermodynamics and NMR of internal G·T mismatches in DNA. Biochemistry36, 10581–10594 (1997). [DOI] [PubMed] [Google Scholar]
  • 33.Allawi, H. T. & SantaLucia, J. Thermodynamics of internal C·T mismatches in DNA. Nucleic Acids Res.26, 2694–2701 (1998). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Sugimoto, N., Nakano, S.-I., Yoneyama, M. & Honda, K.-I. Improved thermodynamic parameters and helix initiation factor to predict stability of DNA duplexes. Nucleic Acids Res.24, 4501–4505 (1996). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Weckx, S., Carton, E., de Vuyst, L. & van Hummelen, P. Thermodynamic behavior of short oligonucleotides in microarray hybridizations can be described using Gibbs free energy in a nearest-neighbor model. J. Phys. Chem. B111, 13583–13590 (2007). [DOI] [PubMed] [Google Scholar]
  • 36.Hassibi, A., Vikalo, H., Riechmann, J. L. & Hassibi, B. Real-time DNA microarray analysis. Nucleic Acids Res.37, e102 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Boehr, D. D., Nussinov, R. & Wright, P. E. The role of dynamic conformational ensembles in biomolecular recognition. Nat. Chem. Biol.5, 789–796 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.North Carolina State University High Performance Computing Services Core Facility (RRID:SCR_022168). https://hpc.ncsu.edu/main.php (2026).
  • 39.Alon, U. An Introduction to Systems Biology: Design Principles of Biological Circuits (Chapman & Hall/CRC, 2006).
  • 40.Angles, R. & Gutierrez, C. Survey of graph database models. ACM Comput. Surv.40, 1 (2008). [Google Scholar]
  • 41.Brandes, U., Freeman, L. C. & Wagner, D. Social networks. in Handbook of Graph Drawing Visualization (ed R. Tamassia) (CRC Press, 2013), pp 805–839.
  • 42.Liu, L., Huang, W. & Huang, J. D. Synthetic circuits that process multiple light and chemical signal inputs. BMC Syst. Biol.11, 38 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Noman, N., Palafox, L. & Iba, H. Evolving genetic networks for synthetic biology. N. Gener. Comput.31, 299–322 (2013). [Google Scholar]
  • 44.Qian, L. & Winfree, E. Scaling up digital circuit computation with DNA strand displacement cascades. Science332, 1196–1201 (2011). [DOI] [PubMed] [Google Scholar]
  • 45.Tamsir, A., Tabor, J. J. & Voigt, C. A. Robust multicellular computing using genetically encoded NOR gates and chemical “wires. Nature469, 212–215 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Bashor, C. J., Horwitz, A. A., Peisajovich, S. G. & Lim, W. A. Rewiring cells: synthetic biology as a tool to interrogate the organizational principles of living systems. Annu. Rev. Biophys.39, 515–537 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Walker, S. P. et al. Non-specific amplification of human DNA is a major challenge for 16S rRNA gene sequence analysis. Sci. Rep.10, 15156 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Monserud, J. H. & Schwartz, D. K. Mechanisms of surface-mediated DNA hybridization. ACS Nano8, 4488–4499 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Polak, R. E. & Keung, A. J. A molecular assessment of the practical potential of DNA-based computation. Curr. Opin. Biotechnol.81, 102940 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Green, A. A., Silver, P. A., Collins, J. J. & Yin, P. Toehold switches: de-novo-designed regulators of gene expression. Cell159, 925–939 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Thompson, K. E., Bashor, C. J., Lim, W. A. & Keating, A. E. Synzip protein interaction toolbox: in vitro and in vivo specifications of heterospecific coiled-coil interaction domains. ACS Synth. Biol.1, 118–129 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Langan, R. A. et al. De novo design of bioactive protein switches. Nature572, 205–210 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Wu, H. et al. Expanding DNA origami design freedom with de novo synthesized scaffolds. J. Am. Chem. Soc.146, 16076–16084 (2024). [DOI] [PubMed] [Google Scholar]
  • 54.Buterez, D. Scaling up DNA digital data storage by efficiently predicting DNA hybridisation using deep learning. Sci. Rep.11, 17498 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Matange, K., Brihadiswaran, G., Tuck, J. & Albert, K. BINND—DNA Sequences [Data set]. Zenodo. 10.5281/zenodo.20494836 (2026).
  • 56.Genome Reference Consortium. Genome Reference Consortium Human Build 38 (GRCh38), RefSeq assembly accession GCF_000001405.26. https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_000001405.26/ (2013).
  • 57.Brihadiswaran, G., Matange, K., Tuck, J., & Keung, A. BINND Dataset [Data set]. Zenodo10.5281/zenodo.19500646 (2026).
  • 58.Brihadiswaran, G. & Tuck, J. dna-storage/BINND: BINND Nature Communications Submission (v1.0.0). Zenodo10.5281/zenodo.19488795 (2026).
  • 59.Li, L. et al. A system for massively parallel hyperparameter tuning. Proc. Mach. Learn. Syst.2, 230–246 (2020). [Google Scholar]
  • 60.Fornace, M. E., Porubsky, N. J. & Pierce, N. A. A unified dynamic programming framework for the analysis of interacting nucleic acid strands: enhanced models, scalability, and speed. ACS Synth. Biol.9, 2665–2678 (2020). [DOI] [PubMed] [Google Scholar]
  • 61.python-Levenshtein: Python extension for computing string edit distances and similarities. https://pypi.org/project/python-Levenshtein/ (accessed 2026).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

41467_2026_75395_MOESM2_ESM.pdf (175.6KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (10KB, xlsx)
Supplementary Data 2 (11.3KB, xlsx)
Supplementary Data 3 (11.3KB, xlsx)
Supplementary Data 4 (12KB, xlsx)
Reporting Summary (1.5MB, pdf)
Source Data (589.9KB, xlsx)

Data Availability Statement

The raw sequencing data files generated in this study have been deposited in the Zenodo database at 10.5281/zenodo.20494835. The dataset used to train and evaluate BINND has been deposited in the Zenodo database at 10.5281/zenodo.19500645. Source data for all other analyses are provided as a Source Data file. Source data are provided with this paper.

The code used to develop the model, perform the analyses, and generate the results in this study is publicly available and has been deposited in the GitHub repository BINND at https://github.com/dna-storage/BINND, under the MIT license. The specific version of the code associated with this publication is archived in Zenodo and is accessible via 10.5281/zenodo.1948879558. The repository includes detailed installation instructions, a Dockerfile for containerized setup, and environment configuration files (BINND.yml and init.sh). External dependencies used for ΔG calculations, NUPACK (v4.0.1.7) and Primer3, are governed by their own licenses; NUPACK is not open-source.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES