Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2024 Oct 3;121(41):e2410529121. doi: 10.1073/pnas.2410529121

Improved deep learning prediction of antigen–antibody interactions

Mu Gao a,b,1, Jeffrey Skolnick a,1
PMCID: PMC11474075  PMID: 39361651

Significance

Accurately predicting antibody–antigen interactions, which are central to the adaptive immune response, is a daunting task. This study explores the potential of a deep learning approach for computationally predicting these interactions. Using the spike protein of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) as a test case, we demonstrate the capability of this computational approach to predict interactions between antibodies and various epitopes on the surface of the antigen. One particularly promising strategy involves using antibody sequences collected from B cell sequencing. The encouraging findings from this approach have significant implications for practical applications to antibody development.

Keywords: antibody–antigen interaction, structure prediction, deep learning, SARS-CoV-2

Abstract

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.


Deciphering how proteins interact with each other to form functional complexes at the atomic level is key to understanding biological processes. In pursuit of this goal, the remarkable success of AlphaFold 2 (AF2) (1), originally created to predict the atomic structure of individual proteins, has spurred new developments to predict protein–protein complexes using deep learning (24). Among these efforts, we introduced AF2Complex, which leverages AF2 deep learning models to predict protein–protein interactions (PPIs) based on the confidence of complex structure modeling (2). To facilitate high-throughput protein–protein interaction prediction on high-performance computing clusters, AF2Complex assembles input features of multiple individual proteins with a variety of options for input tinkering and introduces a confidence metric specifically designed to evaluate the likelihood that a set of multiple proteins interact.

The usefulness of AF2Complex has been demonstrated in practical applications, including predictions subsequently validated in experimental studies. For example, the prediction of a multimeric Escherichia coli CcmI complex has been corroborated by two subsequent experimental studies (5, 6). In a proof-of-concept investigation, we applied an AF2C workflow to search for PPIs within the cell envelope of E. coli (7). This endeavor uncovered not only known interactions but also unveiled high-confidence unanticipated interactions, as well as conformational changes leading to the formation of protein supercomplexes. More recently, we utilized an enhanced version of AF2Complex to detect and model PPIs that are important to T cell regulation. By searching for interaction partners of the tyrosine kinase Lck among 1,000 human proteins involved in adaptive immune responses, we not only generated insightful structural models for known, functionally crucial protein complexes, but also predicted unexpected partners of Lck with profound biological implications (8). Collectively, these two large-scale PPI screening efforts have yielded multiple mechanistic hypotheses ripe for further experimental investigation.

Despite significant progress, notable challenges remain, particularly in predicting antibody–antigen interactions, which are pivotal to adaptive immune responses. While recent advancements in AF2 for multimeric protein structure prediction, AF-Multimer, have notably improved accuracy in predicting antibody–antigen complex structures (9, 10), the success rate falls short in comparison to typical protein complexes that have many evolutionary orthologs (2, 3). These orthologous sequences serve as key input components, compiled in the multiple sequence alignment (MSA) input to AF2 deep learning models. However, for an antigen–antibody target, such orthologous sequences are unavailable, posing a significant obstacle that limits the predictive capabilities for deep learning methods.

Here, we evaluate the capability of AF2Complex to predict antibody–antigen interactions by addressing two questions: i) What is the accuracy of AF2Complex in predicting the antibody–antigen complex structures? ii) Can AF2Complex identify antibodies that bind to a given antigen within a library of antibodies mixed with known binders and arbitrarily chosen antibodies? To address these two questions, we focus on the receptor-binding domain (RBD) of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein because many experimental data are available for benchmarking purpose. The global scientific community’s intensified efforts on understanding immunity against the SARS-CoV-2 virus during the 2020 pandemic led to the discovery of diverse antibody sequences recognizing the virus’s spike protein, including tens of experimentally determined complex structures (11, 12). These valuable data have been cataloged in the database Cov-AbDab (13), from which we curated our benchmarking sets. Furthermore, leveraging these antibody sequences, we devised an MSA strategy that significantly improves the predictive power of the AF2 deep learning models beyond the standard MSA strategy.

Results

We present the findings from two benchmark tests designed to address the two questions outlined above. The first test assesses the accuracy of AF2Complex in predicting the structures of 36 IgG antibodies in complex with the RBD by comparing them with their experimentally determined structures as the ground truth. The second test evaluates the sensitivity and the specificity in a pool of antibodies mixed with 471 known RBD-binders and 500 arbitrarily selected antibodies from B cell sequencing studies of healthy individuals (14).

Predicting Structures of IgG Antibodies Targeting Diverse Epitopes on the Spike RBD.

From the CoV-AbDab database, which archives antibodies that bind to the RBD of the SARS-CoV-2 spike protein (13), we selected 36 paired IgG antibodies (referred to as PDB36) with structures experimentally determined in the presence of the RBD. All of these structures were released in the Protein Data Bank (PDB) after Sept 30, 2021, the cutoff date of structures used for training AF2 “multimer_v3” models (3) that are used in this study. Additionally, we verified that the heavy chains of these IgG domains are unique and not present in PDB235, a dataset composing antibodies that target various coronavirus species with i) unique heavy chain complementarity-determining region 3 (CDR-H3) sequences and ii) with experimental structures potentially utilized in training AF2 models. For each antibody in PDB36, we evaluated the sequence similarity of its CDR-H3 with those in PDB235, taking the highest sequence identity value, denoted as sid*, from these comparisons. The maximum sid* observed in PDB36 is 84% with 29 queries having sid* values below 70%. By comparison, when performing the same evaluation on 500 randomly selected antibodies from B cell sequencing studies of healthy individuals, we found sid* values as high as 82% and a median of 50%. The comparison indicates that the targets in PDB36 are not trivially similar to those that may have been used in training the AF2 models.

In this benchmark, each target consists of the spike RBD and the variable domains of the antibody light and heavy chains, designated as VL and VH, respectively. We use only sequence features from MSAs without any structural template for prediction. We test three different strategies for deriving the input MSAs by querying three different sequence libraries: i) UniProt—employing the UniProt sequence library which is the standard AF2 pipeline; ii) RBD-binding—using 506 antibodies known to bind RBD according to CoV-AbDab; and iii) Arbitrary—using 1,000 antibodies randomly selected from healthy individuals without known COVID-19 infection. For strategies (ii) and (iii), the VH and VL sequences are compiled separately as the sequence libraries for each query light or heavy chains. The resulted MSAs for the light and heavy chains of a target are then paired according to their cognate pairings in the search library. This pairing mode was chosen for structure prediction. It should be noted that the MSAs of RBD remain unpaired from the MSAs of antibody chains in all three scenarios, because the spike protein only hit homologs in viral species but not in species where homologs of Ig folds of antibodies are found. For each MSA strategy, 50 structures are predicted per target, and the top ranked model selected according to the interface score (iScore) is evaluated. It is important to note that we focus only on the protein–protein interface between the antibody and antigen in the iScore calculation, and the interface between the two Ig chains is ignored because it is usually a high-quality prediction. Finally, we also consider a fourth Combined strategy, where the model with the highest iScore from the top-ranked predictions across the three MSA strategies is selected.

Fig. 1A displays the recall and success rate for these four strategies, with detailed statistics for each target given in Dataset S1. A target is considered recalled if its top-ranked model has a statistically significant similarity to its experimental counterpart at the antibody–antigen interface, as defined by the Interface Similarity score (IS-score) whose P-value < 0.01 (15). Furthermore, success also requires a confident iScore > 0.4, which is important for practical applications when the complex structure of a target is unknown. Under these criteria, AF2Complex achieves a 50% (18/36) recall with both UniProt and Arbitrary MSA strategies, improves to 58% (21/36) with the RBD-binding strategy, and reaches the highest recall of 61% (22/36) when the three strategies are combined. A similar pattern is observed with the success rate: 33% (12/36) in both UniProt and Arbitrary strategies, 39% (14/36) in the RBD-binding strategy, and 47% (17/36) when they are combined. Fig. 1B illustrates the relationship between model confidence and quality for each target in the Combined strategy. All 15 (42%) predictions with high confidence (iScore > 0.5) consistently exhibit high model quality (IS-score P-value < 1 × 10−6), with a mean interface backbone Cα RMSD (iRMSD) of 1.1 Å. Moreover, 17 of 18 confident predictions (iScore > 0.4) show significant results in model quality assessment. Interestingly, there are five cases with low iScore scores yet significant model quality, including three cases that are either of high-quality or nearly so. Although it not surprising that targets with high sid* values are relatively easier to predict than those with low sid* values, the Pearson correlation coefficient is modest: 0.48 between sid* and model quality metric IS-score and 0.46 between sid* and the confidence metric iScore for the top-ranked models from the Combined strategy. Among the 22 cases with significant IS-scores, 59% (13/22) of them have an sid* value below 60%.

Fig. 1.

Fig. 1.

Benchmarking of the predictions of the complex structure of 36 IgG antibodies targeting the RBD of the SARS-CoV-2 spike protein. (A) Comparison of different MSA strategies assessed by their recall and success rate. (B) Model quality versus confidence. Each circle in the plot marks a target antibody–antigen complex. The regime where a prediction is both significant in model quality and confidence in model prediction are colored in light pink. Horizontal dash lines indicate the medium and high confidence prediction levels and vertical dash lines mark the significant and high-quality models. (C) Clustering the target by epitope and selected successful predictions for different epitopes. Clusters that share more than 50% of epitopes on average are colored differently, otherwise they are linked by gray lines. Each target is labeled by their PDB code, along with single-letter identifiers for the heavy, light, and antigen chains; symbols marks confident model and its quality. To the Right four representative examples show the computational and experimental models superimposed in a cartoon representation. In each predicted model, the variable antibody domains are colored to match the target label, the RBD is colored in gray, and the experimental structures are depicted in green. An antibody (Ab) is named according to its original reference. Statistics are provided for the interface RMSD (iRMSD), the P-value of the IS-score (P), the interface score (iScore) reported by AF2Complex, and the highest sequence identity (sid*) in the CDR-H3 region with antibodies whose structures may have been used for training AF2 deep learning models.

Given that antibodies binding to the RBD can recognize various epitopes on the antigen, we further examine the top predictions of PDB36 according to their experimentally determined epitopes. A hierarchical clustering of these epitopes confirms that PDB36 encompasses a diverse set of epitopes (Fig. 1C). Four major clusters are clearly distinct, each sharing 10% or fewer epitope residues on average with any other three clusters. At a 50% epitope overlap rate, there are 11 subclusters. Notably, AF2Complex delivers at least one significant prediction in all four major clusters and six subclusters, with successful predictions observed in three major clusters and five subclusters. Four successful examples from very different epitope clusters are illustrated in Fig. 1C. In each case, the predicted structure closely matches the experimental structure, with an interface RMSD of less than 1 Å and highly significant IS-scores. Moreover, these are nontrivial cases whose sid* ranges from 40 to 59%. For instance, Ab159 has a low sid* of 40%, below the median sid* at 50% of arbitrarily chosen antibodies. It has a unique epitope located at an apex of the major loop of RBD, sharing less than 55% of its epitope with those of the other 35 antibodies in PDB36. This uniqueness was noted by the experimentalists who solved the structure (16). Despite the challenge, its top predicted ranked model has a confident iScore of 0.47 and exhibits highly significant similarity to the experimental structure. Overall, this benchmark test demonstrates that AF2Complex can deliver confident and accurate predictions for antibodies that target a broad spectrum of epitopes on the spike RBD, though there is still room to improve the prediction accuracy.

Discriminating True Antibodies for the Spike RBD from Arbitrarily Selected Antibodies.

Next, we explore whether this approach can identify RBD-binding antibodies within a pool of 971 antibodies mixed with 471 known RBD-binders (true positives, RBD471) and 500 arbitrarily selected antibodies (true negatives). In our benchmark analysis, we focus on the VH and VL domains of these antibodies. None of the true positives exhibits a sid* value of 85% or higher, indicating that they are not trivially similar to antibodies whose experimental structures have been potentially used for deep learning modeling training. It is also worth noting that the experimental structures of almost all antibodies in these sets are unknown.

The same four strategies described above are employed in this benchmark test, as shown in Fig. 2 (data provided in Dataset S2). Overall, we observe the same trend in the performance: the strategy using RBD-binders to derive input MSAs delivers outperforms the other two MSA strategies, and combining them enhances the performance in receiver operating characteristic (ROC) curves. Since only the regime of low false positive rate is practically relevant, we focus on the area where the false positive rate is less than 0.1 (Fig. 2A). The normalized area under the curve (AUC) for this threshold, AUC0.1, is 0.28 and 0.27 for Combined and RBD-binding strategies, significantly higher than 0.17 and 0.16 yielded by the UniProt and Arbitrary, respectively. While there is substantial room for further improvement, all strategies clearly beat random predictions, which have a baseline AUC0.1 value of 0.05. At a significant iScore cutoff of 0.4, the standard UniProt strategy recalls 15% (72/471) true positives with a false positive rate (FPR) of 3.6% (18/500). The RBD-binding strategy boosts the recall to 25% (116/471), a 48% increase, while reducing the false positives to 2.6% (13/500), a 33% decrease in FPR. The precision of this strategy also improves from 80 to 90% (Fig. 2B). Combining the three strategies results in the highest recall of 32% with a higher FPR of 6.8%, yielding a precision of 82%. At a high iScore cutoff of 0.5, the RBD-binding strategy delivers a recall of 17% (78/501) and an FPR of 0.6% (3/500), achieving a high precision of 96%. Meanwhile, the Combined strategy delivers a recall of 21% (98/471) and an FPR of 1.4% (7/500), resulting in a precision of 93%. The high precision suggests the potential applicability of this approach in finding antibody binders given an antigen target.

Fig. 2.

Fig. 2.

Detecting antibodies for the spike RBD of SARS-Cov-2 using different strategies. (A) Receiver operating characteristic curve. (B) The precision–recall curve. The random curve is the expected result by random guess. (C) Prediction confidence versus sid*. The iScore of the top 1 ranked model in the Combined scheme is plotted against the sid* of each target antibody in the RBD471 set. Dashed horizontal lines indicate medium (green) and high confident (purple) levels, respectively. Distributions of iScore and sid* are shown on the Top and Left histogram, respectively. (D) Recall obtained with different MSA strategies. The Left panel shows the results of individual runs using specific sequence libraries: full Arbitrary set, three libraries with different numbers of RBD-binding and arbitrary antibodies mixed, and the full RBD-binding set. The Right panel shows the results of the individual runs combined with the two runs using the UniProt and Arbitrary strategies. Medium and high confidence predictions are shaded in light and solid colors, respectively. The number on the top of each bar is the total recall percentage, split into the corresponding values for the medium and high confidence levels within each bar.

It is also worth noting that the 151 detected true positives in the Combined strategy are distinct from the antibodies (PDB235) possibly used in training, as 85% of them have a sid* value of less than 70% and the median sid* is 58% (Fig. 2C). Further analysis of all 471 RBD-binders reveals a weak correlation at 0.17 between the confidence metric iScore and sid*. Notably, 28% of confident predictions are made for difficult targets with sid* values of less than 50%, which is the median for the arbitrarily selected control set, whereas 48% of cases with sid* values of more than 80% were not detected with confidence. This analysis suggests that the deep learning approach can make surprising predictions that are not obvious based on the CDR-H3 sequence similarity with training structures.

To investigate the number of RBD-binding antibody sequences needed to improve predictions, we further tested three MSA strategies, RBD60, RBD125, and RBD250: 60, 125, and 250 RBD-binding antibody sequences from the RBD-binding set with sid* < 70% are randomly drawn and paired with 440, 375, 250 sequences randomly chosen from the Arbitrary set, respectively, to form sequence libraries for constructing the input MSAs. For each case, we repeat the same prediction procedure on RBD471 as described above. The resulting recall obtained by using individual sequence libraries and combined MSA strategies (with the Arbitrary and UniProt strategies) is shown in (Fig. 2D and Dataset S3). Using only 60 RBD-binder and 440 arbitrary antibodies sequences, the RBD60 strategy yields recall values of 16%, slightly higher than the baseline results at 15% using Arbitrary. Moreover, RDB60 confidently identifies targets missed by either Arbitrary or UniProt strategies, leading to an improvement in the Combined strategy with a recall increasing from 21 to 25%. The recall clearly increases when more RBD-binders sequences are utilized: 18% and 21% for RBD125 and RBD250, respectively, corresponding to 28% and 30% in the combined strategies, and approaching the recall of 32% obtained in the Combined strategy when the full RBD-binding set is used. Evidently, including 60 or more RBD-binding antibody sequences in the input MSA library improves the success rate.

Discussion

Using the spike protein of the coronavirus SARS-CoV-2 as a case study, we demonstrated the capability of AF2Complex to identify antibodies targeting the RBD that recognizes angiotensin-converting enzyme 2, the receptor on a host cell surface. This detection is based on structure prediction of antibody–antigen interactions. In the test set of 36 antibody-RBD targets with known experimental structures, and considering only the top one ranked model, the combination of the three MSA strategies resulted in significant structure predictions in over 60% of the targets and successfully predicted complex models with both high confidence and accuracy in nearly half of the testing set. Expanding this approach to larger test sets, approximately ~1,000 antibodies mixed with RBD-binders and arbitrarily selections, at a confident iScore of 0.4, the RBD-binding strategy detects 25% of true antibodies with a precision of 90%, and the Combined strategy finds 32% of true antibodies at a precision of 82%. These results demonstrate that the deep learning approach can distinguish antibodies that specifically recognize the antigen from those that are arbitrarily chosen.

Our study shows that the accuracy of predicting antibody–antigen interactions can be substantially improved by using input MSAs built from a sequence library that contains antibodies targeting the same antigen. Among the three sequence libraries used for three separate MSA strategies, interestingly, the standard UniProt library and a random set of antibodies yield similar performance. In contrast, a sequence library of ~500 RBD-binding antibodies improves the recall by 50% and reduces the FPR by 33% compared to the first two MSA strategies. This finding indicates that the MSAs derived from true RBD-binders enable more accurate inference from the AF2 deep learning models than those assembled from arbitrary sequence libraries, despite the fact that most antibodies in the RBD-binder library are not well-characterized experimentally and likely target diverse epitopes on the antigen. This has significant practical implications, suggesting that compiling sequences of antibodies targeting the same antigen can effectively enhance prediction accuracy for antibody–antigen interactions. Technologies such as single B cell sequencing for monoclonal antibody production could provide the necessary sequences for this strategy (17). Even though such sequence data from B cell sequencing are likely noisy, empirical evidence suggests that deep learning models are capable of extracting valuable information about physical interactions even from noisy inputs.

Despite these notable advances, there remains substantial room for improvement in predicting the complex structures of antibody–antigen interactions. With respect to the spike protein of SARS-CoV-2, two complicating factors may adversely affect the results. First, we only consider the RBD as a monomer, but the spike proteins naturally form trimers under physiological conditions. This can complicate predictions, especially if an antibody binds at the interface between two spike RBD domains (18), which is difficult to predict using only a monomeric RBD. Second, the spike protein undergoes extensive glycosylation modifications on its surface that our current approach cannot explicitly model, potentially reducing the accuracy of our predictions. Although the latter issue is somewhat mitigated for the RBD, where the main epitope is located on a functionally conserved surface patch that is free from posttranslational modifications, it remains a significant challenge for other domains of the spike protein. This is a common issue with viral antigens.

Notwithstanding these challenges, it is evident that AF2 deep learning models have successfully captured the physical representations necessary for accurately predicting at least some antibody–antigen interactions with the standard sequence libraries. If antigen-specific antibody sequences are available, our benchmark tests on the SARS-CoV-2 RBD suggest that having as few as 60 paired antibody sequences could lead to some improvement. The practical effectiveness of this strategy in broader applications warrants further exploration.

Methods

Datasets.

The database CoV-AbDab (13) was queried in December 2022 to obtain IgG antibodies originated from B cell sequencing of human patient infected with SARS-CoV-2 virus. Additionally, we require that the antibodies have been experimentally confirmed to bind to the spike protein of the wild-type SARS-CoV-2 virus, i.e., the virus strain at the outbreak of the pandemic in 2019 (19) and that the epitopes are experimentally determined to be within the RBD. This search yielded us 506 IgG antibodies, and the sequences of their VH and VL domains were employed in the RBD-binding strategy for deriving input MSAs. To exclude entries with experimentally determined structures potentially used for AF2 deep-learning training, we also compiled all antibodies with experimental structures released by the PDB (20) before Sept 30, 2021, the cutoff date for training AF2 “multimer_v3” models. A total of 235 antibodies with unique CDR-H3 sequences were identified, forming the PDB235 set, which includes antibodies targeting various coronavirus species and strains, such as the SARS and MERS, in addition to SARS-CoV-2. We aligned the CDR-H3 between these antibodies and the 506 RDB-binding antibodies and removed RBD-binders if an antibody shares more than 85% CDR-H3 sequence identity with any of the 235 experimentally determined antibodies potentially used for model training. The sequence identity is determined by the number of identical residues normalized by the mean length of two CDR3 regions subjected to comparison. This filtering procedure finally yields 471 RBD-binders as the true positive set, and a small subset of 36 antibodies with experimental structures, i.e., PDB36.

To establish a control group, we randomly select 1,000 antibodies from B cell sequencing of healthy individuals without known COVID-19 infection from the Observed Antibody Space (14). The sequences of the VH and VL domains of these 1,000 antibodies are employed in the Arbitrary MSA strategy. Half of them were randomly drawn as the true negative set in the benchmark test. Note that while we assume that these antibodies are nonbinders, this assumption has not been experimentally verified. For the standard UniProt MSA strategy, the December 2022 UniProt (21) release was utilized.

Antibody–Antigen Prediction with AF2Complex.

The development of AF2Complex has been described in previous publications (2, 7). Here, we utilize AF2Complex to predict potential interactions between antibodies and antigens. One important feature of AF2Complex relevant to this study is the confidence metric known as the interface score (iScore), which conveniently allows us to focus on the interface between the antigen and the paired antibodies and discard the typically strong interaction signals from the antibody heavy and light chains, whose complexation is usually predicted at high accuracy. For structure prediction, the “multimer_v3” set of deep learning models of AF2 (3) was utilized using input features derived from three MSA strategies for the target antibodies. In each strategy, MSAs of VH and VL domains are paired if the MSAs are derived from the libraries of paired antibodies, as detailed earlier, or from the same species when using the standard UniProt library. The MSAs of the spike RBD are derived from the standard UniProt library, and remain unpaired from the MSAs of antibodies, regardless of the MSA strategy adopted. No structural template was allowed in all structure predictions described in this work. Given a target from the PDB36 set, each AF2 neural network model is invoked ten times with different random seeds and up to 20 recycles. For targets from the two larger sets, to reduce computing costs, 471 RBD-binders and 500 control antibodies, each AF2 model is applied five times and up to 12 recycles. For each target, the top ranked model by the iScore is retained for evaluation purpose.

Interface Score.

The confidence metric has been described previously (2). It is based on the interface TM-score introduced in iAlign (22) and the predicted alignment error estimated by AF2 (3). The score is defined as the follows:

iScore=p=1C1ImaxiI\IpjIp11+eij/d0I2, [1]

where I is the set of protein–protein interface residues in the predicted complex model, I is the total number of interface residues, i.e., the cardinality of I, |I|, and C denotes all protein chains. Each chain p has an observed number of interface residues Ip, and I is the union of Ip. The predicted alignment error eij of interface residue j is calculated according to a local reference frame based on interface residue i. The optimal local reference for calculating the maximum score for interface residues of chain p can only be selected by interface residues not belonging to chain p. The normalization factor d0(I) is given by,

d0I=1.24I-153-1.8if I220.02Iif I<22. [2]

By default, the iScore is evaluated for all interfaces of a multimeric complex. If a complex has three or more chains, it can be specified to evaluate the interface between subsets of chains, such as the interface between an antigen and a paired antibody, while ignoring the interface between chains within the same subsets, such as the heavy and light chains of a paired antibody.

The statistical significance of the iScore is based on the results of 7,000 putatively noninteracting protein pairs from E. coli (2). At iScore thresholds of 0.40, 0.50, and 0.70, the empirical P-value is calculated at 1.2%, 0.4%, and <0.01%, respectively. These thresholds correspond to medium, high, and very high levels of confidence, respectively.

Performance Evaluation.

Standard metrics were applied to the benchmark tests on the RBD-binders and the control set. The predictions were labeled using the predefined classification and the numbers of true positives, false positives, true negatives, and false negatives were then designated as TP, FP, TN, and FN, respectively. Performance measures are defined as follows:

True Positive Rate=Recall=TPTP+FNFalse Positive Rate=FPTN+FPPrecision=TPTP+FP. [3]

We also employed the normalized AUC0.1, which is the area under the ROC curve up to an FPR of 0.1, divided by 0.1. ROC curves were plotted using ROCR (23).

Structural Analysis.

The program IS-score was used to evaluate the quality of predicted model using the experimental structure for comparison (15). The statistical significance of the IS-score is modeled by extreme value distributions of the scores from rigid-body docking decoys. For hierarchical clustering of epitopes, an epitope residue is defined on the surface of RBD if an RBD residue is less than 4.5 Å from any heavy atom of the bound target antigen; and the average linkage method is adopted. VMD (24) was used to visualize the protein structural models.

Computational Resources.

The development and PPI predictions were mainly performed on the Perlmutter supercomputer at National Energy Research Scientific Computing Center (NERSC) and the Delta supercomputer at the National Center for Supercomputing Applications. In the first test on the PDB36 set, it took about two Nvidia A100 h per each antibody–antigen target to predict 50 complex structures in 20 AF2 recycles. In the second set of 971 targets, for each, it took approximately 40 A100 min per target to predict 25 complex structures in 12 AF2 recycles.

Supplementary Material

Dataset S01 (XLSX)

pnas.2410529121.sd01.xlsx (32.2KB, xlsx)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2410529121.sd03.xlsx (99.8KB, xlsx)

Acknowledgments

This project was funded in part by the Division of General Medical Sciences of the NIH (grant R35GM118039) and the Department of Energy Office of Science, Office of Biological and Environmental Research (DOE grant DE-SC0021303). The research used resources through allocations from the DOE-funded Advanced Scientific Computing Research Leadership Computing Challenge program and the National Energy Research Scientific Computing Center (NERSC award ERCAP0027209), the NSF-funded Advanced Cyberinfrastructure Coordination Ecosystem Services & Support (ACCESS award BIO230134), and the Partnership for an Advanced Computing Environment at the Georgia Institute of Technology.

Author contributions

M.G. and J.S. designed research; M.G. performed research; M.G. analyzed data; and M.G. and J.S. wrote the paper.

Competing interests

M.G. is a stockholder at AgnistaBio, Inc. J.S. declares no competing interests.

Footnotes

This article is a PNAS Direct Submission.

Contributor Information

Mu Gao, Email: mu.gao@agnistabio.com.

Jeffrey Skolnick, Email: skolnick@gatech.edu.

Data, Materials, and Software Availability

Lists and sequences of the antibodies used in this study are available at Zenodo https://doi.org/10.5281/zenodo.11110249 (25). The source code of AF2Complex is freely available at https://github.com/FreshAirTonight/af2complex (26).

Supporting Information

References

  • 1.Jumper J., et al. , Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Gao M., Nakajima An D., Parks J. M., Skolnick J., AF2Complex predicts direct physical interactions in multimeric proteins with deep learning. Nat. Commun. 13, 1744 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Evans R., et al. , Protein complex prediction with AlphaFold-Multimer. bioRxiv [Preprint] (2021). 10.1101/2021.10.04.463034 (Accessed 10 April 2022). [DOI]
  • 4.Baek M., et al. , Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Li J., et al. , Structures of the CcmABCD heme release complex at multiple states. Nat. Commun. 13, 6422 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ilcu L., Denkhaus L., Brausemann A., Zhang L., Einsle O., Architecture of the Heme-translocating CcmABCD/E complex required for Cytochrome c maturation. Nat. Commun. 14, 5190 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Gao M., Nakajima An D., Skolnick J., Deep learning-driven insights into super protein complexes for outer membrane protein biogenesis in bacteria. Elife 11, e82885 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Gao M., Skolnick J., Predicting protein interactions of the kinase Lck critical to T cell modulation. Structure In press (2024).
  • 9.Yin R., Pierce B. G., Evaluation of AlphaFold antibody–antigen modeling with implications for improving predictive accuracy. Protein Sci. 33, e4865 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.McCoy K. M., Ackerman M. E., Grigoryan G., A comparison of antibody-antigen complex sequence-to-structure prediction methods and their systematic biases. bioRxiv [Preprint] (2024). 10.1101/2024.03.15.585121 (Accessed 25 March 2024). [DOI] [PMC free article] [PubMed]
  • 11.Hastie K. M., et al. , Defining variant-resistant epitopes targeted by SARS-CoV-2 antibodies: A global consortium study. Science 374, 472–478 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Chen Y., et al. , Broadly neutralizing antibodies to SARS-CoV-2 and other human coronaviruses. Nat. Rev. Immunol. 23, 189–199 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Raybould M. I. J., Kovaltsuk A., Marks C., Deane C. M., CoV-AbDab: The coronavirus antibody database. Bioinformatics 37, 734–735 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Olsen T. H., Boyles F., Deane C. M., Observed antibody space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences. Protein Sci. 31, 141–146 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gao M., Skolnick J., New benchmark metrics for protein-protein docking methods. Proteins 79, 1623–1634 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Takeshita M., et al. , Potent SARS-CoV-2 neutralizing antibodies with therapeutic effects in two animal models. iScience 25, 105596 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Pedrioli A., Oxenius A., Single B cell technologies for monoclonal antibody discovery. Trends Immunol. 42, 1143–1158 (2021). [DOI] [PubMed] [Google Scholar]
  • 18.Li R., et al. , Conformational flexibility in neutralization of SARS-CoV-2 by naturally elicited anti-SARS-CoV-2 antibodies. Commun. Biol. 5, 789 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wu F., et al. , A new coronavirus associated with human respiratory disease in China. Nature 579, 265–269 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Berman H. M., et al. , The protein data bank. Nucleic Acids Res. 28, 235–242 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.The UniProt Consortium, UniProt: A worldwide hub of protein knowledge. Nucleic Acids Res. 47, D506–D515 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Gao M., Skolnick J., iAlign: A method for the structural comparison of protein-protein interfaces. Bioinformatics 26, 2259–2265 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Sing T., Sander O., Beerenwinkel N., Lengauer T., ROCR: Visualizing classifier performance in R. Bioinformatics 21, 3940–3941 (2005). [DOI] [PubMed] [Google Scholar]
  • 24.Humphrey W., Dalke A., Schulten K., VMD: Visual molecular dynamics. J. Mol. Graphics 14, 33–38 (1996). [DOI] [PubMed] [Google Scholar]
  • 25.Gao M., Skolnick J., Benchmark data derived from the antibodies for the spike receptor-binding domain of SARS-CoV-2. 10.5281/zenodo.11110249. Deposited 13 August 2024. [DOI]
  • 26.Gao M., Nakajima An D., Parks J. M., Skolnick J., AF2Complex version 1.4.1. 10.5281/zenodo.13732465. Deposited 8 September 2024. [DOI]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Dataset S01 (XLSX)

pnas.2410529121.sd01.xlsx (32.2KB, xlsx)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2410529121.sd03.xlsx (99.8KB, xlsx)

Data Availability Statement

Lists and sequences of the antibodies used in this study are available at Zenodo https://doi.org/10.5281/zenodo.11110249 (25). The source code of AF2Complex is freely available at https://github.com/FreshAirTonight/af2complex (26).


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES