Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 Apr 22;27(2):bbag180. doi: 10.1093/bib/bbag180

Computational nanobody design using graph neural networks and Metropolis Monte Carlo sampling

Lei Wang 1,#, Xiaoming He 2,#, Xinzhou Qian 3, Gaoxing Guo 4, Qiang Huang 5,6,7,✉
PMCID: PMC13099421  PMID: 42015415

Abstract

Nanobodies are promising protein therapeutics due to their high-stability, low immunogenicity, and ease of production. However, experimental screening of high-affinity nanobodies and their post optimization remain costly and time-consuming due to the vast variant space. Here, we developed a computational approach that integrates graph neural networks (GNNs) with Monte Carlo Metropolis algorithm for nanobody design. We constructed a GNN model, AiPPA, to predict the protein–protein binding free energy (BFE) without requiring the complex structure, achieving a Pearson correlation of 0.62 on benchmark. We then combined AiPPA with Metropolis importance sampling to design low-BFE nanobodies from a non-affinity template. We applied this method to the antigen TL1A and generated two affinity nanobodies. This work establishes a physics-informed deep learning method for computational nanobody design, providing a novel development strategy for protein therapeutics.

Keywords: protein therapeutics, antibody design, protein–protein interaction, binding free energy, Metropolis algorithm

Introduction

Antibodies have played an important role in the treatment of many diseases in recent years [1, 2]. Since 2020, approximately 10 antibody-based biologics have been approved each year [3], and most of the top best-selling drugs worldwide were antibodies [4]. The success of antibody therapeutics has driven the development of various novel types of antibodies [5]. Among these, nanobody (Nb), as the smallest antigen-binding fragment [6] (~15 kDa), offers excellent tissue and tumor penetration [7, 8]. The high stability [9], low immunogenicity [10], and ease of production [11] has further solidified its role in therapeutic innovation. To date, three Nb-based therapeutics—Caplacizumab [12], Envafolimab [13], and Ozoralizumab [14]—have been approved for clinical use; and, many more are in clinical trials [15].

Nanobodies are primarily discovered through library screening; however, these screened nanobodies often exhibit low binding affinity to disease-related antigens [16, 17]. Since high affinity is essential to enhance the efficacy and reduce off-target effects [18], affinity optimization is required for the screened nanobodies. Two primary approaches have been developed for affinity optimization in vitro. The first is based on the hypermutagenic properties of specific cells, such as B cells or H1229 cells, and to replicate the in vivo hypermutation process in an in vitro setting [19, 20]. However, this method heavily relies on cells types and requires complicated manipulations [18]. The second involves generation of mutant libraries by introducing random mutations into nanobodies, followed by selection systems to screen for variants with increased affinities [21]. This method often requires large libraries and multiple rounds of screening [22, 23]. In addition, in vitro affinity optimization has the risk to compromise the stability and other properties of the nanobodies [24, 25].

To address these challenges, computational design methods have been developed [21, 22]. These methods typically treat affinity optimization as an energy minimization task, aiming to generate variants with lower binding free energies (i.e. higher binding affinities). This process usually involves two steps: low-energy variant search in sequence space and binding affinity evaluation of the variants. However, existing computational affinity evaluation methods, such as PRODIGY [26], often rely on the accurate antibody–antigen complex structures. Even with advanced prediction tools like AlphaFold-Multimer [27] or AlphaFold3 (AF3) [28], predicting protein–protein interaction (PPI) complex structures remain time-consuming for large numbers of antibody variants, and sometimes the predictions are insufficiently accurate for the free energy calculations [29, 30]. Consequently, protein–protein complex structure-based methods are difficult to deal with large numbers of potential variants, often resulting in limited improvement [31, 32]. Thus, it is still necessary to develop more efficient methods to identify affinity variants of antibodies targeting given antigens.

In this study, we considered nanobody design for a specific antigen as an affinity optimization process starting from a non-affinity template nanobody. First, we constructed a graph neural network (GNN) model called AiPPA (AI-enabled Protein–Protein Affinity predictor) for predicting the protein–protein binding free energy (BFE) without the need for their complex structures. Then, we used the Monte Carlo (MC) Metropolis importance sampling algorithm [33] to explore sequence variants of the template complementarity-determining regions (CDRs). During sampling, AiPPA calculated the BFEs of the variants to the antigen, guiding the exploration of low-BFE sequences. We applied this approach to design nanobodies targeting TL1A (Tumor necrosis factor-like cytokine 1A), an antigen implicated in autoimmune diseases. Bio-layer interferometry (BLI) confirmed that two of the six designs bind TL1A.

Materials and methods

Construction of protein graphs

Each protein in a PPI pair is represented as an undirected graph Inline graphic, where Inline graphic denotes the set of nodes, each representing a residue in the protein; Inline graphic is the adjacency matrix of the graph with dimensions Inline graphic; and Inline graphic is the set of edges in the graph. An edge only exists between node Inline graphic and node Inline graphic if and only if their Cα-Cα distance is less than or equal to the threshold θ:

graphic file with name DmEquation1.gif (1)

The residue type (Inline graphic), residue-level solvent-accessible surface area (Inline graphic), and Cα coordinates (Inline graphic) were used as the initial features of each node. The 20 amino acid types were encoded as integers ranging from 0 to 19 and transformed into 10-dimensional vectors. The Inline graphic values were directly calculated using Biopython (https://biopython.org) based on the protein structure. Instead of directly using the Cα coordinates of the residues, we first translated the protein such that its centroid was positioned at the origin. The transformed Cα coordinates were then used as the positional features (Inline graphic). Finally, these three features were concatenated to form the initial node features:

graphic file with name DmEquation2.gif (2)

where Inline graphic represents concatenation operator, resulting in an initial feature dimension of 14 for each node. The initial edge features were set to an 8-dimension 8 zero vector:

graphic file with name DmEquation3.gif (3)

Model architecture

During the node and edge updates at layer Inline graphic, the edge features were aggregated as “message”:

graphic file with name DmEquation4.gif (4)

where Inline graphic represents all the neighboring nodes of node Inline graphic and Inline graphic is a learnable weight matrix. In the first layer, the hidden states of the nodes and edges are their initial features. The message is then passed to the node Inline graphic, and update its hidden state using the GINE model:

graphic file with name DmEquation5.gif (5)

where Inline graphic is a learnable. Next, BN was performed and then nonlinear activation was applied to update the hidden states:

graphic file with name DmEquation6.gif (6)

After updating the nodes, their hidden states were passed to the adjacent edges to update the hidden states of the edges:

graphic file with name DmEquation7.gif (7)

After stacking two EGINE layers (Inline graphic), “global add” operation was performed to aggregate the features of all the edges in the last layer, producing the final representation of the protein graph:

graphic file with name DmEquation8.gif (8)

where Inline graphic and represents each of the two interacting proteins. Finally, the representations of the two interacting graphs were combined and fed into a three-layer MLP to predict the BFE of A and B:

graphic file with name DmEquation9.gif (9)

MC sampling of low-BFE variants

Three operations were used to generate sequence variants: point mutation, insertion, and deletion. The probabilities of these operations were similar to the study by Yeh et al. [34]: point mutations were assigned a probability of 70%, and insertions and deletions together accounted for 30%. Based on the natural nanobodies, the CDR length sampling ranges are: CDR1 - 7 amino acids, CDR2 - 5 or 6 amino acids, and CDR3 - randomly varying from 5 to 20 amino acids. The transition probabilities for insertion and deletion are related to the length of a given CDR: deletion (equation 10) and insertion (equation 11) probabilities are determined by the sequence proportion with the same CDR length in the natural nanobodies:

graphic file with name DmEquation10.gif (10)
graphic file with name DmEquation11.gif (11)

where Inline graphic, L is the given length of CDR2 or CDR3, and Inline graphic is the number of the corresponding CDRs with length Inline graphic in the INDI database. For insertion and point mutation, we used the amino acid composition at each CDR position of the natural nanobodies as the sampling probability. In addition, for the hypervariable region of CDR3 (100A–100L), we used the amino acid frequencies given in the McMahon library [35] as the sampling probability.

After calculating the BFE (Inline graphic) between a given variant and the protein target using AiPPA, its acceptance or rejection was determined by the Metropolis sampling criterion:

graphic file with name DmEquation12.gif (12)

where Inline graphic is the acceptance probability of step Inline graphic, Inline graphic is the BFE predicted by AiPPA at the stepInline graphic, and Inline graphic is the binding free energy at the previous step t-1; Inline graphic is the ideal gas constant of 8.314 J mol−1 K−1; and T is temperature. T was set to the room temperature times 0.2 (i.e. 59.6 K) in order to slightly increase the acceptance probability of the variants.

Local refinement of binding interface

The binding interface residues of a given nanobody were defined as those with at least one heavy atom within 4.5 Å of the target. To refine the interface residues, all non-interface residues in the given nanobody were fixed, and then ProteinMPNN was run to generate 100 interface sequences at two temperature (0.1 and 0.3):

python protein_mpnn_run.py \

--jsonl_path ../parsed_pdbs_bb.jsonl \

--chain_id_jsonl ../assigned_chains.jsonl \

--fixed_positions_jsonl ../masked_pos.jsonl \

--out_folder $output_dir \

--num_seq_per_target 100 \

--sampling_temp "0.1 0.3" \

--seed 37 \

--batch_size 1

where parsed_pdbs_bb.jsonl contains PDB information of the given structure, assigned_chains.jsonl specifies chain details, and masked_pos.jsonl defines the fixed positions, allowing for targeted site design. The binding affinities of the refined sequences were also evaluated using the method by Yi et al. [36], and then the top 10 sequences were selected for further interface quality analysis.

Protein expression and purification

Nanobody expression and purification followed the established protocols in our previous study [37]. Briefly, the gene encoding the nanobody with an N-terminal pelB signal peptide and a C-terminal 6 × His tag (separated by a TEV protease site) was cloned into pET-26b(+) and expressed in E. coli Rosetta (DE3). The extracted protein was purified by Ni-NTA affinity chromatography, and the eluate was dialyzed against PBS. The His-tag was removed by TEV protease cleavage, followed by reverse Ni-NTA chromatography to eliminate trace tagged contaminants. The purified protein was concentrated, aliquoted, flash-frozen, and stored at −80°C.

BLI measurements

BLI measurements were performed on the Sartorius Octet R8 system using Octet® SA biosensors. BLI experiments were performed at 25°C using PBS (pH 7.4) supplemented with 0.02% (w/v) Tween-20 as the detection buffer. A BLI measurement began with sensor equilibration in detection buffer for 90 seconds. Biotinylated TL1A was then immobilized on the biosensors for 300 seconds until a binding signal of 3.0 nm was achieved. The sensors were then equilibrated in the detection buffer for 180 seconds. The sensors were then exposed to wells containing Br3s95 and Tr3s90 nanobodies for 160 seconds to allow binding. Finally, the sensors were immersed in detection buffer for 100 seconds to monitor dissociation. The BLI data were analyzed using the Octet® Analysis Studio Software (ver. 13.0.1.35).

Results

Architecture and training of AiPPA

Various machine learning-based models have been developed for protein–protein binding affinity prediction [26, 38, 39]. Many methods used the PPI complex structures for the affinity prediction. To avoid this, we used GNNs to model two interacting protein partners (A and B) as two separate graphs (Fig. 1A). Then, the graph representation learning was applied to each protein graph (A or B). These two protein graphs were combined to predict the BFE of A and B (Fig. 1B).

Figure 1.

The overall architecture of the AiPPA model, illustrating the graph-based representation of proteins, the design of the graph neural network (GNN), and the mechanisms of message passing within the graph.

Architecture of the AiPPA model. (A) Graph representation of two interacting proteins (A and B). Residues of the proteins serve as nodes, with their physicochemical properties encoded as node features. Edges are defined by the contact map and represent spatially adjacent amino acid residues based on the 3D structures of the proteins. The initial edge features are set to 0. (B) Architecture and learning process of the AiPPA model. The AiPPA model consists of two serially connected EGINE blocks, which iteratively update the graph representations of proteins based on the training data. The graph representations of the two proteins are then combined using concatenation and addition operations and fed into a MLP to calculate the binding free energy between A and B. (C) Message passing in the EGINE blocks. During the learning process, message is passing between the nodes and edges of the proteins and iterative updates. Blue-1 arrows indicate messages from adjacent edges, green-2 arrows represent messages from neighboring nodes, and red-3 arrows denote self-referential messages from edges or nodes.

As shown in Fig. 1A, residues were represented as nodes in the graph, and the initial values of their features were assigned according to three properties: Cα coordinates, solvent-accessible surface area (SASA), and residue type. To facilitate message passing between the spatially close residues, we also added edges to the protein graphs for any two nodes whose Cα-Cα distance is less than a given cutoff value, θ. After several tests, we found that AiPPA achieved optimal performance at θ = 8 Å (Fig. S1A). Since the edges represent the relationships between residues and support the message passing between the nodes, we ever defined the initial edge features according to the node distances, etc. However, during the model training process, we found that the initial edge features did not contribute to the model performance (Fig. S1B). Thus, the initial edge features were set as zero vectors.

As shown in Fig. 1B, AiPPA consists of two EGINE blocks, which are based on the graph isomorphism networks (GINs) that uses an injective aggregation function and has the graph classification ability approaching the Weisfeiler–Lehman test [40]. Hu et al. introduced edge features into the aggregation process and thereby developed GINE (GIN with edges), which further improved its performance [41]. We further introduced “edge updating” to construct the edge-updating GINE (EGINE), in order to improve the message passing between the nodes and the edges (Fig. 1C). Then, an EGINE layer combines a batch normalization (BN) layer and a ReLU activation function layer to form an EGINE block (Fig. 1B). After the processing by two stacking EGINE blocks, the output edge features of the two graphs formed the final representations of A and B. These representations were then merged and fed into a multilayer perceptron (MLP) that outputs the BFE (Fig. 1B).

To train AiPPA, we screened a dataset of 1858 PPI samples from the PDBbind database (v2020) [42] (see Supplementary Methods and Fig. S2), which was the largest training dataset available at the time of this study. We divided this dataset into training and validation sets in an 8:2 ratio for the model training and validation, respectively (Fig. S3A). The model training was stopped when the validation loss reached its minimum, in order to avoid overfitting (Fig. S3B). And hyperparameters used in training are listed in Fig. S3C.

AiPPA performance

To verify the AiPPA model, we used a popular benchmark for the PPI affinities: the Kastritis benchmark [43], which contains 144 PPI complexes with experimental affinities and high-resolution structures. For a reliable comparison with other representative models, we selected 138 samples from this benchmark as the test set (see Supplementary Methods). The free energy distributions of this test set are similar as those of the training and validation sets (Fig. S3A). During the period of our study, three main prediction models were available for the comparison: PPA_Pred2 [39], PPI-Affinity [3, 4], and PRODIGY [26]. PPA_Pred2 uses protein sequences as input and was trained by 453 dimers selected from PDBbind (v2015); PPI-Affinity was trained by 833 samples selected from PDBbind (v2020). Unlike them, PRODIGY was trained by the Kastritis benchmark.

Based on the test set, we calculated the Pearson correlation coefficient (PCC) and root mean square error (RMSE) values for each model. As shown in Fig. 2A and Table S1A, AiPPA achieved the best PCC of 0.62, which was 11%, 26%, and 15% higher than those of PPA_Pred2, PRODIGY, and PPA-Affinity, respectively; AiPPA also had a relatively low RMSE (Fig. 2A). Thus, the predicted affinities by AiPPA showed a strong correlation with the experimental values (Fig. 2B). Although PPI-Affinity had the lowest RMSE value, its correlation was not good (PCC = 0.54). This suggests that AiPPA has the best ability to rank the binding free energies of the samples, an important property when searching for the lowest energy sequences. Notably, homologous leakage is not appeared in this test (Table S1B). Significantly, AiPPA performed well for those samples with binding free energies around −10 kcal·mol−1, the most frequent region in the energy distributions of the training set (Fig. S3A). By the way, the evaluation on antibody benchmark also shows the better performance of AiPPA (Table S1C).

Figure 2.

A comprehensive performance evaluation using RMSE and PCC benchmark the predictive performance of the AiPPA model against other computational methods on the Kastritis dataset and its structural subsets.

Performance of the AiPPA model. (A) Comparison of AiPPA, PPA_Pred2, PRODIGY, and PPI-Affinity on the Kastritis benchmark (test set). The left panel shows the PCC, with higher values indicating better predictive accuracy, while the right panel displays the RMSE, where lower values correspond to smaller deviations between predicted and experimental values. (B) Correlation plot of predicted versus experimental binding free energies of the test set. (C) Performance metrics (PCC, left; RMSE, right) for AiPPA, PPA_Pred2, PRODIGY, and PPI-Affinity on rigid and flexible subsets of the test set.

As known, proteins usually undergo conformational changes during their binding processes. Therefore, to accurately predict the interactions between two “flexible” proteins are more difficult than those between two relatively rigid proteins. To evaluate the effectiveness of AiPPA for this kind of interactions, we further divided the test set into rigid complexes (Rigid, I_RMSD ≤1.0 Å, n = 67, Fig. S4A) and flexible complexes (Flexible, I_RMSD >1.0 Å, n = 71, Fig. S4B) based on Vangone and Bonvin’s criteria [26], and then assessed the model performance on these two subsets (Fig. 2C and Table S1A). As seen, sequence-based model PPA_Pred2 exhibited similar performance for both subsets, because the input sequence features were the same for both cases. The highest PCC and the lowest RMSE indicated that PRODIGY had the best prediction for the rigid complexes. However, for the flexible complexes, its performance dropped by ~30% in the PCC. In contrast, PPI-Affinity and AiPPA performed better on the flexible complexes. AiPPA achieved the best results on both subsets, with a high accuracy for the flexible complexes (PCC = 0.67).

For the flexible complexes, conformational changes at the binding interface often result in differences between the bound and unbound states of the protein (A or B; Fig. S4B). To evaluate the robustness of AiPPA to such changes, we assessed its predictions using both bound and unbound protein conformations, respectively. As shown in Fig. 2C, AiPPA demonstrated consistent performance compared to Fig. 2A, indicating its ability to accurately predict the BFEs despite conformational changes during binding. This is important for its application in the nanobody design, where interface flexibility is often a critical factor.

MC sampling of CDR variants

A typical nanobody comprises four framework fragments (FR1, FR2, FR3, and FR4) and three CDR loops (Fig. 3A). The nanobody frameworks are very conserved, but the CDR loops are not. Since the CDRs play a crucial role in the binding to the antigens [44], nanobody design and engineering usually focus on the CDR loops, especially the CDR3 loop. To integrate AiPPA, we considered the nanobody design for a specific antigen as an in silico generation of the CDR variants of a non-affinity template nanobody, especially a nanobody with a humanized framework. To this end, we employed MC simulation, a thermodynamically rigorous method widely used for sampling complex energy landscapes, to explore CDR variants of the template. Using AiPPA, we calculated the BFE of each sampled variant with the target and accepted or rejected variants based on the Metropolis criterion, which ensures convergence to a Boltzmann-weighted equilibrium distribution [45]. As illustrated in Fig. 3B, each MC iteration comprises the following five steps: (i) generating a new CDR variant from the previous sequence, (ii) predicting its solubility index, (iii) predicting its 3D structure, (iv) evaluating its BFE to the target using AiPPA, and (v) accepting or rejecting the variant according to the Boltzmann-weighted energy change criterion. This physics-based framework enables efficient exploration of high-dimensional sequence design spaces, guiding the search toward low-BFE, high-affinity variants.

Figure 3.

This figure illustrates the structural framework and sequence patterns of a template nanobody VHH1 alongside a computational workflow that employs iterative Metropolis Monte Carlo sampling to design and evaluate variant sequences through binding free energies.

Design of high-affinity nanobodies by combining AiPPA and MC Metropolis sampling. (A) 3D structure and sequence of the template nanobody VHH1. The framework is shown in yellow, with highlighted CDRs: CDR1 (upper center), CDR2 (upper right), and CDR3 (upper left). During the design process, CDR2 and CDR3 lengths were varied by inserting amino acids at specific positions: position 52A for CDR2 and positions 100A–100L for CDR3, depending on the sequence length. (B) MC sampling for low-BFE nanobody design based on the Metropolis Boltzmann-weighted criterion. Variants are iteratively generated through MC sampling. Each iteration involves: (1) Generating a new sequence (St + 1) via random mutation, followed by solubility prediction; low-solubility sequences are discarded and regenerated; (2) predicting the 3D structure of high-solubility sequences and calculating their binding free energy (ΔGt + 1) with the target protein using AiPPA; and (3) applying the thermodynamic Metropolis criterion to compare ΔGt + 1 with the binding free energy (ΔGt) at last step to determine sequence acceptance or rejection. The variant sequence and binding free energy are updated iteratively. (C) Accepted variants and their binding free energies from 30 000 MC sampling iterations. The top 5% of variants exhibited binding free energies below −12.2 kcal·mol−1, as indicated by the dashed line.

To enhance the sampling efficiency for nature-like variants, we analyzed ~20,000 natural nanobody sequences from the INDI database [46] to extract sequence features in their CDRs (Fig. 3A, right; Table S2). After removing duplicates, we obtained 3108 CDR1, 3812 CDR2, and 10 770 CDR3 sequences for the analysis (Fig. S5). We found that the most frequent lengths for CDR1, CDR2, and CDR3 are 7, 5–6, and 5–20 residues (>90% frequency), respectively. Specific amino acid preferences were observed at certain positions, such as glycine (G) at H26 and tyrosine (Y) at H102, suggesting conservation in CDR loop sequences. To generate nature-like variants during MC sampling, we constrained CDR lengths to these ranges: CDR1 = 7 aa, CDR2 = 5–6 aa, and CDR3 = 5–20 aa. Length transition probabilities were calculated based on current CDR lengths (Methods; Fig. S6). Furthermore, amino acid sampling probabilities were weighted according to their natural distribution frequencies in the CDR loops.

To ensure a soluble expression of a sampled variant, in each MC iteration step we used Protein-Sol [47] to predict the solubility index of the variant and rejected those with a solubility index <0.45. For a variant with an index ≥0.45, its 3D structure was then predicted using IgFold [48], which provides similar accuracy to AlphaFold2 [49] but is almost 60 times faster. The BFE of the variant to the protein target was then calculated using AiPPA to determine its probability of acceptance using the Metropolis algorithm (Methods, Fig. 3B). The MC sampling process was terminated when the iteration steps of 30 000 were reached.

Low-BFE variant design

To validate our method, we applied it to design low-BFE nanobodies targeting TL1A, a key therapeutic target for inflammatory bowel diseases (IBDs) such as ulcerative colitis and Crohn’s disease, which remain among the most challenging diseases of the 21st century [50, 51]. We used the nanobody VHH1 as the template for the design (Fig. S7). VHH1 is an affinity nanobody targeting TNF-α, which is a homologue protein of TL1A with relevance to rheumatoid arthritis (RA) and IBDs [52] (sequence in Fig. S7A). We expressed and then confirmed via BLI that VHH1 does not bind TL1A (Fig. S7B and C). Starting from the VHH1 sequence, we performed 30 000 MC simulation steps, and ~9400 variants were accepted (Fig. 3C). Notably, ~3200 variants exhibited predicted BFEs below −12.0 kcal·mol−1, with 34 variants reaching energies < −13.0 kcal·mol−1 and a minimum of −13.4 kcal·mol−1. For further evaluation, we selected the top 5% of low-energy variants (ΔG < −12.2 kcal·mol−1; Fig. 3C).

To evaluate the sequence uniqueness and diversity of the selected variants, we calculated the Levenshtein distances (LDs) between their CDRs and those of VHH1, as well as the shortest LDs between the generated variants and natural nanobodies. As shown in Fig. 4a, most variants exhibited LDs > 3 for CDR1 and CDR2 relative to VHH1, while CDR3 showed a minimum LD of 10, with most distances ranging between 13 and 15 (Fig. 4A, right), indicating significant divergence from VHH1. Furthermore, the shortest LDs indicated that CDR1 and CDR2 exhibited a high sequence similarity to those of natural nanobodies (LD = 1-2), with some variants matching natural sequences exactly (LD = 0). In contrast, CDR3 loops displayed a minimum LD of 6 and a maximum of 11 compared to the natural CDR3 loops (Fig. 4A, left), highlighting both the critical role of CDR3 in binding affinity and the potential for generating novel nanobody sequences. To further assess diversity, we computed pairwise LDs between the CDR loops of the variants. As shown in Fig. 4B, the majorities of LDs were nonzero, demonstrating the high sequence diversity of the variants.

Figure 4.

Sequence analysis of MC-designed nanobodies targeting TL1A reveals their novelty and diversity relative to the template and natural nanobodies and uncovers intrinsic sequence patterns across the three CDRs.

Sequence characteristics of the MC-sampled affinity nanobodies targeting TL1A. (A) LDs between the CDRs of the nanobodies and the template VHH1 (right) and the minimum LDs between the CDRs of the nanobodies and natural nanobodies (left). Larger edit distances indicate greater sequence divergence. (B) Pairwise LDs among the CDR loops of the sampled nanobodies. (C) Length distribution of the CDR3 loops in the sampled nanobodies. (D) Amino acid distribution logo plots for key positions in the CDR loops of natural nanobodies and the sampled nanobodies. Logo plots were generated using WebLogo3 [53].

To investigate whether the sampled variants exhibited specific sequence features for binding, we analyzed the length distributions of CDR3 loops and the amino acid compositions of all three CDR loops. As shown in Fig. 4C, the CDR3 loops of the variants were predominantly longer (16–18 amino acids) compared to the 14-residue CDR3 of VHH1, suggesting that increased loop length may expand the binding interface with TL1A. Sequence logos revealed that although the dominant amino acids in CDR1 and CDR2 were comparable to those of VHH1, their relative frequencies differed significantly (Fig. 4D). Notably, CDR3 loops displayed a strong preference for specific residues, particularly arginine (R) and serine (S), which may enhance binding through electrostatic interactions and hydrogen bonding.

Site-specific filtering of designs

Because Nb-TL1A complex structures were not used in the process of the MC variant generation, we did not know the exact binding site of a generated variant on TL1A. Although the top 5% variants in above were predicted as potential affinity variants, their actual binding sites on TL1A might be different. To identify those variants that bind to a desired binding site of TL1A (i.e. residues related to DcR3 binding [54]) (Fig. 5A, top), we performed site-specific filtering on these variants. Given the high false-positive rates in protein design [55], a systematic procedure shown in Fig. 5A (middle) was used to filter and refine the top 5% sampled variants, and finally to get the designs binding to the desired site (Fig. 5A, bottom).

Figure 5.

A post-design workflow combining AlphaFold2, docking, interface analysis, and ProteinMPNN is applied to further filter the designed nanobodies for specific binding to TL1A epitopes.

Site-specific filtering of designed nanobodies targeting TL1A. (A) Workflow for screening nanobodies binding to a desired binding site. The homo-trimeric TL1A is depicted with its three chains using different colors (top and bottom panels). The target epitope, spanning two chains, is highlighted by the dashed circle (top and bottom panels). From 1500 predicted low-binding free energy variants, 6 candidate nanobodies (0.4%) binding to the specified TL1A epitope were identified. The nanobody is shown as cartoon, with its CDRs displayed as sticks (bottom panel). (B) Distributions of pLDDT scores for the CDR loops of the sampled variants. The dashed line indicates pLDDT >70, representing high-confidence predicted structures. (C) Relationship between binding free energy (ΔG_separated) and interface area (ΔSASA_int) in interface quality analysis. (D) Interface quality evaluation using three parameters: packstat >0.65, ΔG/ΔSASA × 100 < −1.5, and ΔG_separated < −40. Binding free energy (in REU) is represented by a color gradient, with darker shades indicating lower (more favorable) binding energies.

We first predicted the 3D structures of all the variants using AlphaFold2 and then selected those variants with a pLDDT (predicted Local Distance Difference Test) value >70, based on the confidence classification established by the AlphaFold Protein Structure Database. As shown in Fig. 5B, the pLDDT distributions of the variants indicated that the prediction confidence of the CDR1 and CDR2 loops was high, but that of the CDR3 loops was relatively low, probably due to the greater length and sequence variability. Using the filtering threshold of pLDDT >70 for all the CDR loops, 405 matching variants were then obtained. To identify their binding sites on TL1A, these variants were docked to TL1A using the HDOCK program [56], and the top-ranked docked pose of a given variant was used as its representative binding conformation to TL1A. This yielded 31 variants bound to the desired binding epitope, as illustrated in Fig. S8. We then minimized the folding energies of their complex structures with the Rosetta Relax protocol and then analyzed their Nb-TL1A interfaces using the InterfaceAnalyzer module.

Three interface properties from InterfaceAnalyzer were considered for the further filtering: Rosetta binding free energy (dG_separated) and solvent accessible area buried at the interface (dSASA_int), and packing state (packstat). Indeed, we found that the binding free energies and the solvent accessible areas buried at the interfaces of the 31 variants showed a linear correlation: a larger dSASA_int corresponds to a lower binding free energy (Fig. 5C). Therefore, the interface property criteria (dG_separated < −40, dG/dSASA × 100 < −1.5, and packstat >0.65, top 10% most favorable designs) were used to filter the remaining 31 variants. As illustrated in Fig. 5D, three variants were found to meet these criteria (i.e. 0.20% of the top 5% variants), exhibiting good interface packing and binding energy per unit area.

Based on the docked complex structures of these three variants, we used ProteinMPNN [57] to locally refine their interface residues in the CDR3 loops, with a preference for amino acids “YGSRV” according to the study by Yi et al. [36]. Finally, we performed a self-consistency prediction of the Nb-TL1A complex structures for all the refined variants using AF3, and then selected those variants that bind to the desired epitope and match the interface criteria above. Six designs were finally selected for the experimental validation (Table S3).

Experimental validation

To verify the binding activities of the six designs, we expressed them in Escherichia coli Rosetta DE3 cells. Except for Tr4s18, the other five designs were successfully expressed in the soluble form. We purified the nanobodies and cleaved their 6-His tags with TEV protease. Then, we further purified the cleaved samples by reverse affinity chromatography to obtain the tag-free proteins (Fig. S9).

To verify their binding activities, we used BLI to measure their binding kinetics to TL1A. To this end, biotin-labeled TL1A proteins were immobilized on the SA sensors, and the nanobodies at different concentrations were then introduced as the mobile phase to observe binding and dissociation over a specific concentration range. The measurements showed that two designs, Br3s95 and Tr3s90, exhibited clear binding curves (Fig. 6A). And the best-fitting Kd (equilibrium dissociation constant) values of the Br3s95 and Tr3s90 were 2.56 and 1.86 μM, respectively. This indicates that we have successfully designed two binding nanobodies for TL1A using our approach (Table S4). To the best of our knowledge, this study reports the first designed nanobody specifically targeting TL1A. While its binding affinity is not yet comparable to that of the existing anti-TL1A antibody ABS-101, this work establishes a critical foundation for expanding the therapeutic landscape of TL1A-related diseases.

Figure 6.

Two nanobodies capable of binding TL1A were validated using BLI, and structural analysis revealed that their interaction with TL1A forms a hydrophobic core and further stabilized by a network of hydrogen bonds.

Experimental validation and structural modeling of the designed nanobodies. (A) Binding kinetics of nanobodies Br3s95 and Tr3s90 to TL1A measured by BLI, with corresponding SDS-PAGE analysis on the right. Light-colored traces represent raw BLI data, and dark-colored curves show the best-fit models. (B) Atomic models of Br3s95 and Tr3s90 in complex with TL1A. TL1A is shown as a surface model, while Br3s95 and Tr3s90 are depicted as cartoon models. CDR1, CDR2, and CDR3 loops are indicated in lime, cyan, and magenta, respectively. (C) Hydrogen bonds between the CDR loops of Br3s95 and TL1A, indicated by dashed lines with distance. (D) Hydrogen bonds between the CDR loops of Tr3s90 and TL1A, indicated by dashed lines with distance. (E) Hydrophobic interactions of Br3s95 mediated by residue F43 of TL1A. Dashed lines indicate surrounding residues in the hydrophobic core, and the nearby surface is colored by hydrophobicity (orange-1: hydrophobic; cyan-2: hydrophilic). (F) Hydrophobic interactions of Tr3s90 mediated by residue F43 of TL1A. Dashed lines indicate surrounding residues in the hydrophobic core, and the nearby surface is colored by hydrophobicity (orange-1: hydrophobic; cyan-2: hydrophilic). Panels C and D were generated using Ligplot+ [58].

Since the affinity maturation process can sometimes affect the structural stability of the antibodies [24], we further evaluated the protein stability of the designed nanobodies compared to the starting template VHH1 by determining their melting temperature (Tm) using nanoscale differential scanning fluorimetry (nanoDSF). As shown in Fig. S10, the Tm value of the template VHH1 is about 70.0°C, but the Tm values of Br3s95 and Tr3s90 are about 72.3°C and 77.1°C, respectively, slightly higher than that of VHH1. This indicates that the computational design not only increased the binding affinity but also slightly improved the protein stability.

To elucidate the key interactions, we performed structural analysis based on the complex structures predicted by AF3. As shown in Fig. 6B, Br3s95 and Tr3s90 interact with TL1A primarily through their CDR3 loops; however, their binding poses at the interface differ: Br3s95 adopts an upward-facing conformation, whereas Tr3s90 adopts a downward-facing conformation. In Br3s95, the CDR2 and CDR3 loops form hydrogen bonds with TL1A (Fig. 6C), whereas in Tr3s90 the hydrogen bonds are predominantly formed by the CDR3 loop (Fig. 6D). For both nanobodies, serine and basic amino acids are central to the hydrogen bonding networks (Fig. 6C and D). In addition, both nanobodies interact with the hydrophobic residue F43 on the TL1A surface, which forms the core of the hydrophobic interactions. In Br3s95, F43 interacts with Y104 in the CDR3 loop through π-π stacking (Fig. 6E), whereas in Tr3s90, F43 interacts with V104 of the CDR3 loop with its surrounding residues (Fig. 6F). Besides the hydrogen bonding, these hydrophobic interactions further stabilize the Nb-TL1A binding.

Discussion

We have combined GNNs with MC sampling to develop a deep learning approach for nanobody design. We used this method to design nanobodies targeting TL1A and validated that two designs could bind TL1A. As mentioned, our AiPPA is a structure-based GNN model that uses only the structures of two individual protein partners (A and B) to predict the BFE. In contrast, many previous models require the protein–protein interface features extracted from the complex structures [26, 38]. Meanwhile, for large-scale computational tasks such as those in this study, accurate prediction of the complex structures by molecular docking or AF3 [28, 30] is often not feasible because the prediction even for a single complex structure takes several minutes or tens of minutes on a workstation computer with ordinary graphics processing units (GPUs). Therefore, affinity prediction methods that do not require any complex structures are helpful. Of course, as the prediction accuracy plays an important role in the affinity variant design, AiPPA still needs to be improved in the future.

Recently, various protein design methods have been used to design or optimize antibodies, including nanobodies [32, 59, 60]. In contrast to general protein design, antibody design often focuses on the design of the CDR loop sequences and does not require extensive design of their frameworks. In practice, to minimize development risk, a therapeutic antibody usually prefers to have a humanized framework and, as far as possible, its sequence is not altered during the antibody design process. Therefore, our method is particularly suited to such an antibody design scenario. In nature, our method mainly designed the sequences of the CDR loops of the template nanobody. Unlike many point mutation optimization methods where the CDR loop lengths were fixed [32, 60], our design method allowed changes not only in the amino acid types of the CDR loops but also in their lengths during the sampling process of the affinity variants. Also, given a specific template protein, our method can be easily extended to design other proteins, including the traditional full-length antibodies.

Optimizing VHH1 from a nonbinding template to high-affinity binding with TL1A is not merely an affinity optimization problem, but also an ab initio design task. Compared to certain generative-model-based approaches, our method enables explicit optimization of user-defined properties—such as solubility and binding affinity—and allows additional physicochemical constraints to be directly incorporated into the MC sampling process. However, to adequately explore the vast protein sequence space, our approach may be less efficient than direct generation methods, highlighting computational efficiency as a key priority for future development.

Furthermore, the performance of AiPPA critically affects the design of affinity-improved variants and is currently constrained by the scarcity of high-quality experimental protein–protein binding affinity data. Although training under data-limited conditions remains a substantial challenge, emerging strategies—including transfer learning and large-scale pretraining—offer promising avenues to improve the predictive accuracy and generalization of AiPPA in the future.

In summary, we developed a computational framework that integrates GNNs with MC sampling to enable efficient design of affinity nanobodies to a given antigen. By applying this approach, we successfully designed two nanobodies targeting TL1A, a key therapeutic antigen in autoimmune diseases. Therefore, this study provides a novel approach to antibody design and offers nanobodies with potential for the treatment of autoimmune diseases.

Key Points

  • We developed AiPPA, a graph neural network model for accurate prediction of protein–protein binding affinity without requiring complex structural information.

  • We established a rational computational pipeline for in silico design of nanobodies by integrating AiPPA with Metropolis Monte Carlo sampling.

  • We validated the high-affinity binding of two designed nanobodies targeting TL1A, demonstrating the practical utility of our computational design approach.

Supplementary Material

BIB-25-2456R2_SuppInfo_bbag180

Acknowledgements

We think Dr. Yong Zhao (U-mab Biopharma, Inc.) for his valuable assistance during the initial phase of this project. This work was supported by the National Key Research and Development Program of China (2021YFA0910604) and the National Natural Science Foundation of China (31971377 and 31671386).

Contributor Information

Lei Wang, State Key Laboratory of Genetics and Development of Complex Phenotypes, Shanghai Engineering Research Center of Industrial Microorganisms, MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, 2005 Songhu Road, Shanghai 200438, China.

Xiaoming He, State Key Laboratory of Genetics and Development of Complex Phenotypes, Shanghai Engineering Research Center of Industrial Microorganisms, MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, 2005 Songhu Road, Shanghai 200438, China.

Xinzhou Qian, State Key Laboratory of Genetics and Development of Complex Phenotypes, Shanghai Engineering Research Center of Industrial Microorganisms, MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, 2005 Songhu Road, Shanghai 200438, China.

Gaoxing Guo, State Key Laboratory of Genetics and Development of Complex Phenotypes, Shanghai Engineering Research Center of Industrial Microorganisms, MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, 2005 Songhu Road, Shanghai 200438, China.

Qiang Huang, State Key Laboratory of Genetics and Development of Complex Phenotypes, Shanghai Engineering Research Center of Industrial Microorganisms, MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, 2005 Songhu Road, Shanghai 200438, China; Multiscale Research Institute of Complex Systems, Fudan University, 825 Zhangheng Road, Shanghai 201203, China; Shanghai Key Laboratory of Gene Editing and Cell Therapy for Rare Diseases, Fudan University, 83 Fenyang Road, Shanghai 200031, China.

Author contributions

Q.H. supervised the project. L.W., X.H. and Q.H. conceived the project; L.W. developed the computer programs and performed the computational study; L.W., G.G. and X.Q. collected the data and analyzed the computer codes. X.H. and G.G. performed the experimental validation and analysis; L.W. and Q.H. drafted and revised the manuscript. All authors contributed to the drafting and revision of the paper.

Conceptualization: L.W., X.H. and Q.H.; Data curation & Validation: L.W., G.G. and X.Q.; Methodology: L.W., X.H. and G.G; Software: L.W.; Supervision: Q.H.; Writing—original draft: L.W. and Q.H.; Writing—review & editing: All authors

Funding

None declared.

Data availability

All datasets used for the model training and evaluation are publicly available, with the details described in Methods and Supplementary Methods.

Code availability

The source codes for AiPPA are publicly available via GitHub. And the address link is https://github.com/Fudan-HQLab/AiPPA.

References

  • 1. Lu  RM, Hwang  YC, Liu  IJ  et al. Development of therapeutic antibodies for the treatment of diseases. J Biomed Sci  2020;27:1. 10.1186/s12929-019-0592-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Mullard  A. FDA approves 100th monoclonal antibody product. Nat Rev Drug Discov  2021;20:491–5. [DOI] [PubMed] [Google Scholar]
  • 3. Strohl  WR. Structure and function of therapeutic antibodies approved by the US FDA in 2023. Antib Ther  2024;7:132–56. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Verdin  P. Top companies and drugs by sales in 2023. Nat Rev Drug Discov  2024;23:240–0. [DOI] [PubMed] [Google Scholar]
  • 5. Carter  PJ, Rajpal  A. Designing antibodies as therapeutics. Cell  2022;185:2789–805. [DOI] [PubMed] [Google Scholar]
  • 6. Bao  G, Tang  M, Zhao  J  et al. Nanobody: a promising toolkit for molecular imaging and disease therapy. EJNMMI Res  2021;11:6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Bates  A, Power  CA. David vs. Goliath: the structure, function, and clinical prospects of antibody fragments. Antibodies  2019;8:28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Kijanka  M, Dorresteijn  B, Oliveira  S  et al. Nanobody-based cancer therapy of solid tumors. Nanomed.  2015;10:161–74. [DOI] [PubMed] [Google Scholar]
  • 9. Muyldermans  S. Nanobodies: natural single-domain antibodies. Annu Rev Biochem  2013;82:775–97. [DOI] [PubMed] [Google Scholar]
  • 10. Vu  KB, Ghahroudi  MA, Wyns  L  et al. Comparison of llama VH sequences from conventional and heavy chain antibodies. Mol Immunol  1997;34:1121–31. [DOI] [PubMed] [Google Scholar]
  • 11. Liu  Y, Huang  H. Expression of single-domain antibody in different systems. Appl Microbiol Biotechnol  2018;102:539–51. [DOI] [PubMed] [Google Scholar]
  • 12. Scully  M, Cataland  SR, Peyvandi  F  et al. Caplacizumab treatment for acquired thrombotic thrombocytopenic purpura. N Engl J Med  2019;380:335–46. 10.1056/NEJMoa1806311 [DOI] [PubMed] [Google Scholar]
  • 13. Li  J, Deng  Y, Zhang  W  et al. Subcutaneous envafolimab monotherapy in patients with advanced defective mismatch repair/microsatellite instability high solid tumors. J Hematol Oncol  2021;14:95. 10.1186/s13045-021-01095-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Ishiwatari-Ogata  C, Kyuuma  M, Ogata  H  et al. Ozoralizumab, a humanized anti-TNFα NANOBODY® compound, exhibits efficacy not only at the onset of arthritis in a human TNF transgenic mouse but also during secondary failure of administration of an anti-TNFα IgG. Front Immunol  2022;13:853008. 10.3389/fimmu.2022.853008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Yong Joon Kim  J, Sang  Z, Xiang  Y  et al. Nanobodies: robust miniprotein binders in biomedicine. Adv Drug Deliv Rev  2023;195:114726. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Muyldermans  S. A guide to: generation and design of nanobodies. FEBS J  2021;288:2084–102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Liu  W, Song  H, Chen  Q  et al. Recent advances in the selection and identification of antigen-specific nanobodies. Mol Immunol  2018;96:37–47. 10.1016/j.molimm.2018.02.012 [DOI] [PubMed] [Google Scholar]
  • 18. Zhang  H, Xue  JB, Wang  AZL  et al. EASINESS: E. Coli assisted speedy affINity-maturation evolution SyStem. Front Immunol  2021;12:747267. 10.3389/fimmu.2021.747267 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Cumbers  SJ, Williams  GT, Davies  SL  et al. Generation and iterative affinity maturation of antibodies in vitro using hypermutating B-cell lines. Nat Biotechnol  2002;20:1129–34. 10.1038/nbt752 [DOI] [PubMed] [Google Scholar]
  • 20. Chen  S, Qiu  J, Chen  C  et al. Affinity maturation of anti-TNF-alpha scFv with somatic hypermutation in non-B cells. Protein Cell  2012;3:460–9. 10.1007/s13238-012-2024-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Li  J, Kang  G, Wang  J  et al. Affinity maturation of antibody fragments: a review encompassing the development from random approaches to computational rational optimization. Int J Biol Macromol  2023;247:125733. 10.1016/j.ijbiomac.2023.125733 [DOI] [PubMed] [Google Scholar]
  • 22. Tabasinezhad  M, Talebkhan  Y, Wenzel  W  et al. Trends in therapeutic antibody affinity maturation: from in-vitro towards next-generation sequencing approaches. Immunol Lett  2019;212:106–13. 10.1016/j.imlet.2019.06.009 [DOI] [PubMed] [Google Scholar]
  • 23. Tiller  KE, Chowdhury  R, Li  T  et al. Facile affinity maturation of antibody variable domains using natural diversity mutagenesis. Front Immunol  2017;8:986. 10.3389/fimmu.2017.00986 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Shehata  L, Maurer  DP, Wec  AZ  et al. Affinity maturation enhances antibody specificity but compromises conformational stability. Cell Rep  2019;28:3300–3308.e4. 10.1016/j.celrep.2019.08.056 [DOI] [PubMed] [Google Scholar]
  • 25. Julian  MC, Li  L, Garde  S  et al. Efficient affinity maturation of antibody variable domains requires co-selection of compensatory mutations to maintain thermodynamic stability. Sci Rep  2017;7:45259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Vangone  A, Bonvin  AM. Contacts-based prediction of binding affinity in protein–protein complexes. eLife  2015;4:e07454. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Evans  R, O'Neill  M, Pritzel  A  et al. Protein complex prediction with AlphaFold-Multimer. bioRxiv. 2021. 10.1101/2021.10.04.463034 [DOI] [Google Scholar]
  • 28. Abramson  J, Adler  J, Dunger  J  et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature  2024;630:493–500. 10.1038/s41586-024-07487-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Bryant  P. Deep learning for protein complex structure prediction. Curr Opin Struct Biol  2023;79:102529. [DOI] [PubMed] [Google Scholar]
  • 30. Lin  P, Li  H, Huang  SY. Deep learning in modeling protein complex structures: from contact prediction to end-to-end approaches. Curr Opin Struct Biol  2024;85:102789. [DOI] [PubMed] [Google Scholar]
  • 31. Cheng  X, Wang  J, Kang  G  et al. Homology modeling-based in silico affinity maturation improves the affinity of a nanobody. Int J Mol Sci  2019;20:4187. 10.3390/ijms20174187 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Bai  Z, Wang  J, Li  J  et al. Design of nanobody-based bispecific constructs by in silico affinity maturation and umbrella sampling simulations. Comput Struct Biotechnol J  2023;21:601–13. 10.1016/j.csbj.2022.12.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Metropolis  N, Rosenbluth  AW, Rosenbluth  MN  et al. Equation of state calculations by fast computing machines. J Chem Phys  1953;21:1087–92. [DOI] [PubMed] [Google Scholar]
  • 34. Yeh  AHW, Norn  C, Kipnis  Y  et al. De novo design of luciferases using deep learning. Nature  2023;614:774–80. 10.1038/s41586-023-05696-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. McMahon  C, Baier  AS, Pascolutti  R  et al. Yeast surface display platform for rapid discovery of conformationally selective nanobodies. Nat Struct Mol Biol  2018;25:289–96. 10.1038/s41594-018-0028-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Yi  CH, Taylor  ML, Ziebarth  J  et al. Predictive models and impact of interfacial contacts and amino acids on protein–protein binding affinity. ACS Omega  2024;9:3454–68. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Qin  Q, Jiang  X, Huo  L  et al. Computational design and engineering of self-assembling multivalent microproteins with therapeutic potential against SARS-CoV-2. J Nanobiotechnol  2024;22:58. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Romero-Molina  S, Ruiz-Blanco  YB, Mieres-Perez  J  et al. PPI-affinity: a web tool for the prediction and optimization of protein–peptide and protein–protein binding affinity. J Proteome Res  2022;21:1829–41. 10.1021/acs.jproteome.2c00020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Nikam  R, Yugandhar  K, Michael Gromiha  M. Discrimination and prediction of protein-protein binding affinity using deep learning approach. In: Huang  DS, Jo  KH, Zhang  XL (eds.), Intelligent Computing Theories and Application, pp. 809–15. Switzerland: Springer Cham, 2018. [Google Scholar]
  • 40. Xu  K, Hu  W, Leskovec  J  et al. How powerful are graph neural networks? In: Proc. of International Conference on Learning Representations 2019. 10.48550/arXiv.1810.00826 [DOI]
  • 41. Hu  W, Liu  B, Gomes  J  et al. Strategies for pre-training graph neural networks. In: Proc. of International Conference on Learning Representations 2020.
  • 42. Wang  R, Fang  X, Lu  Y  et al. The PDBbind database: collection of binding affinities for protein−ligand complexes with known three-dimensional structures. J Med Chem  2004;47:2977–80. [DOI] [PubMed] [Google Scholar]
  • 43. Kastritis  PL, Moal  IH, Hwang  H  et al. A structure-based benchmark for protein-protein binding affinity: protein-protein structure-affinity benchmark. Protein Sci  2011;20:482–91. 10.1002/pro.580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Liu  B, Yang  D. Easily established and multifunctional synthetic nanobody libraries as research tools. Int J Mol Sci  2022;23:1482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Landau DP, Binder K. A Guide to Monte Carlo Simulations in Statistical Physics. 5th ed. Cambridge: Cambridge University Press; 2021. 10.1017/9781108780346 [DOI] [Google Scholar]
  • 46. Deszyński  P, Młokosiewicz  J, Volanakis  A  et al. INDI—integrated nanobody database for immunoinformatics. Nucleic Acids Res  2022;50:D1273–81. 10.1093/nar/gkab1021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Hebditch  M, Carballo-Amador  MA, Charonis  S  et al. Protein–sol: a web tool for predicting protein solubility from sequence. Bioinformatics  2017;33:3098–100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Ruffolo  JA, Chu  LS, Mahajan  SP  et al. Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies. Nat Commun  2023;14:2389. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Jumper  J, Evans  R, Pritzel  A  et al. Highly accurate protein structure prediction with AlphaFold. Nature  2021;596:583–9. 10.1038/s41586-021-03819-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Xu  WD, Li  R, Huang  AF. Role of TL1A in inflammatory autoimmune diseases: a comprehensive review. Front Immunol  2022;13:891328. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Wang  S, Wang  P, Wang  D  et al. Postbiotics in inflammatory bowel disease: efficacy, mechanism, and therapeutic implications. J Sci Food Agric  2024;104:1234–45. [DOI] [PubMed] [Google Scholar]
  • 52. Beirnaert  E, Desmyter  A, Spinelli  S  et al. Bivalent llama single-domain antibody fragments against tumor necrosis factor have picomolar potencies due to intramolecular interactions. Front Immunol  2017;8:867. 10.3389/fimmu.2017.00867 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Crooks  GE, Hon  G, Chandonia  JM  et al. WebLogo: a sequence logo generator. Genome Res  2004;14:1188–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Zhan  C, Patskovsky  Y, Yan  Q  et al. Decoy strategies: the structure of TL1A:DcR3 complex. Structure  2011;19:162–71. 10.1016/j.str.2010.12.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Bennett  NR, Coventry  B, Goreshnik  I  et al. Improving de novo protein binder design with deep learning. Nat Commun  2023;14:2625. 10.1038/s41467-023-38328-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Yan  Y, Zhang  D, Zhou  P  et al. HDOCK: a web server for protein–protein and protein–DNA/RNA docking based on a hybrid strategy. Nucleic Acids Res  2017;45:W365–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Dauparas  J, Anishchenko  I, Bennett  N  et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science  2022;378:49–56. 10.1126/science.add2187 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Laskowski  RA, Swindells  MB. LigPlot+: multiple ligand–protein interaction diagrams for drug discovery. J Chem Inf Model  2011;51:2778–86. [DOI] [PubMed] [Google Scholar]
  • 59. Bennett  NR, Watson  JL, Ragotte  RJ  et al. Atomically accurate de novo design of single-domain antibodies. Preprint at  2024. 10.1101/2024.03.14.585103 [DOI] [Google Scholar]
  • 60. Hie  BL, Shanker  VR, Xu  D  et al. Efficient evolution of human antibodies from general protein language models. Nat Biotechnol  2024;42:275–83. 10.1038/s41587-023-01763-2 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

BIB-25-2456R2_SuppInfo_bbag180

Data Availability Statement

All datasets used for the model training and evaluation are publicly available, with the details described in Methods and Supplementary Methods.


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES