Abstract
Gaining structural insights into antibody-antigen interactions is essential for understanding immune recognition and therapeutics design. Accurately modeling these complexes remains challenging for both physics-based approaches and AI-based, co-folding methods such as AlphaFold3. These methods not only struggle to generate near-native conformations, but, more critically, often fail to rank those correctly, revealing fundamental limitations for antibody-antigen modeling. We present DeepRank-Ab, a geometric deep learning-based scoring function tailored to antibody-antigen interfaces, together with a rigorously curated benchmark of ~2.3 million decoys from 1,442 complexes, providing the diversity required for robust training and unbiased evaluation. We systematically assessed graph representations, structural and energetic features, and sampling strategies. Our analysis identified that atom-level representations coupled with Voronoi-based surface decomposition and antibody-specific features are the most effective formulation for accurate scoring. Across multiple independent test sets, DeepRank-Ab consistently outperforms all evaluated methods, including AlphaFold3, HADDOCK and state-of-the-art scoring functions. It increases AlphaFold3 Top1 success rate by 35.5% and improves the mean Top1 DockQ by more than a factor of two. DeepRank-Ab generalizes beyond its training distribution, achieving 100% Top5 success rate on external antibody-antigen CAPRI targets, surpassing all tested methods. These results establish DeepRank-Ab as effective scoring method substantially improving identification of near-native antibody-antigen conformations.
Subject terms: Software, Structural biology
Geometric deep learning with DeepRank-Ab improves scoring of antibody–antigen complexes, enabling more accurate ranking of near-native conformations compared to state-of-the-art methods including AlphaFold3.
Introduction
Antibodies are protective proteins generated by B lymphocytes as part of the humoral immune response to antigens. Structurally, they are Y-shaped macromolecules composed of two heavy and two light chains, each containing constant and variable regions. The variable regions include the complementarity-determining regions (CDRs), which are hypervariable loops that form the antigen-binding site and determine antibody specificity and affinity. Sequence and structural variations within these loops give rise to the immense diversity of antibodies, estimated at 10¹⁵ to 10¹⁸ distinct specificities1,2. Studying antibody–antigen interactions is therefore fundamental not only for understanding immune recognition but also for advancing vaccine and diagnostic development, as well as the design of therapeutic antibodies3.
Modeling antibody–antigen interactions at atomic resolution remains highly challenging, even in the age of artificial intelligence. Co-folding algorithms4,5 rely heavily on co-evolutionary signals to infer residue-residue contacts, but such signals are largely absent in antibody–antigen pairs. Although AlphaFold3 (AF3)4 reports success rates of 40–60% in modeling such complexes, these results are typically obtained only after extensive sampling and re-ranking using internal confidence metrics. Open-source reproductions of AF3, including OpenFold36, Boltz-27, etc., still underperform compared to AF3 despite being trained on comparable data with similar architectures. Moreover, across all these approaches, a substantial gap exists between the best conformations produced by the models and those favored in their rankings8.
In AlphaFold2, model ranking relies on a weighted combination of ipTM and pTM5, whereas AF3 additionally applies penalties for steric clashes and structural disorders4. Several recent methods aim to improve ranking, including pDockQ29 for chain-specific evaluation, actifpTM10 for handling flexible regions, and ipSAE11 for better separation of correct and incorrect interfaces. However, for antibody–antigen complexes, recent work12 shows that the main source of error lies not in the ranking formula itself but in AlphaFold’s limited ability to estimate alignment error. This suggests that further refinement of AF3’s ranking scheme is insufficient, and that external scoring functions and/or improving AF3’s confidence prediction head13 are needed to improve performance on these systems.
Given these challenges, physics-based modeling remains highly valuable. Accurate modeling of antibody–antigen interactions is largely dependent on obtaining a reliable starting structure—a task made difficult by the vast sequence diversity of antibodies and the intrinsic flexibility of their CDR loops14. In particular, correct conformation of CDR-H3 has been shown to be essential for successful docking15. The scoring problem is also severe in docking: only about 20% of the top-ranked models correspond to correct binding modes, reflecting the limited capability of current scoring functions, particularly in unbound docking scenarios where conformational changes occur during binding16.
Deep learning approaches trained on docking decoys have been widely explored17–21. Despite these advances, the scoring problem remains unsolved. Poor generalizability is a persistent issue. As shown by Stratiichuk et al., scoring functions trained for rigid docking protocols often fail to generalize to decoys generated through flexible sampling22. One contributing factor is the limited diversity of available training data. Constructing large, heterogeneous docking benchmarks is difficult, and most scoring methods were trained on decoys produced by a single protocol and software. Data leakage poses a second challenge: strict separation between training and test sets is not always enforced, and common practices such as sequence-similarity or temporal filtering do not guarantee dataset independence23, leading to inflated performance estimates.
To address this, we developed DeepRank-Ab, a deep-learning scoring function tailored specifically to antibody–antigen interfaces. Because of the lack of benchmark datasets covering both bound and unbound conformations while capturing the structural flexibility of antibody-antigen interactions, we first constructed a dataset of more than 2 million models generated from 1442 antibody–antigen complexes under four docking protocols. Building upon our previous DeepRank-GNN-esm18 framework, we systematically evaluated multiple levels of graphs, feature representations, and neural network architectures to optimize the overall model formulation. The final DeepRank-Ab models achieve performance that surpasses external methods, including AF3 ranking confidence4, HADDOCK score24,25, and one of the leading scoring functions, VoroIF-jury17. DeepRank-Ab is available at https://github.com/haddocking/DeepRank-Ab.
Results
Assessment of AF3’s scoring capabilities
We first evaluated AF3’s ability to rank antibody–antigen complexes before developing DeepRank-Ab. Our curated benchmark contains 1442 antibody–antigen structures from SAbDab26 (see “Strict structural splitting to avoid data leakage” for details). Following Foldseek-Multimer27 structure-based clustering and splitting, 215 complexes were assigned as benchmark test sets, and among these, 59 complexes were released after the AF3 training cutoff date. Refer to “Methods” for details on clustering and splitting. We refer to this subset as AF3 test sets to ensure an unbiased evaluation of AF3’s scoring performance. For each complex in the AF3 test sets, AF3 was run using 100 random seeds, generating a total of 500 models to ensure sufficient sampling.
AF3 achieves a Top 1 success rate of 40.7%, which is similar to values reported in the literature4 (Fig. 1a). When considering the Top 50 ranked models, the success rate increases to 59.3%. Here, the Top K success rate is defined as the fraction of complexes for which at least one model among the top K ranked predictions reaches near-native quality, where near-native is defined as a DockQ ≥ 0.23, corresponding to CAPRI’s “acceptable” threshold28. We adopted this threshold throughout this work as predicting antibody–antigen interaction is quite challenging for both AF3 and docking; stricter thresholds would yield very few successes, making comparisons between scoring methods less informative.
Fig. 1. Performance evaluation of AF3 and graph-based scoring algorithms.

a Evaluation of AF3’s scoring performance on 59 antibody–antigen complexes in our test set. DockQ ≥ 0.23 was used as the cutoff to define near-native models. The Top K success rate, where K represents the number of models considered, and the Top K DockQ values at each K are shown. The “oracle” curve represents the ideal performance obtained when the best possible model is selected for each complex. The top K success rate is defined by counting the number of docking cases in which at least one near-native model is found among the Top K ranking models, divided by the total number of cases. Top K DockQ is computed by selecting the highest DockQ value among Top K-ranked models for each complex and then averaging across complexes. b One example (PDBID: 7PS0) where AF3 produces high-quality models but ranks it 498th, while the Top 1 ranked model is incorrect. c Distribution of top-ranked DockQ scores for AF3 and oracle ranking (n = 59). Error bars represent standard deviation across complexes.
Both the Top 1 success rate and average Top K DockQ28 would be nearly doubled under the oracle setting, which represents an optimal scoring scenario where the best model is always ranked first. The oracle does not reach 100% because, for some complexes, none of the generated models are near-native; in other words, even an ideal ranking cannot recover an acceptable structure if it was not sampled. To better illustrate the spread and variability of DockQ values across the test set, we plotted the distribution of top-ranked DockQ values for AF3 and the oracle (Fig. 1c). These plots reveal a wide range of DockQ values, indicating that the quality of predicted models varies substantially between complexes. This variability underscores that the success of scoring and ranking for AF3 is highly target dependent.
This observation highlights a central limitation of AF3’s scoring: the method fails to recognize acceptable models and fails to prioritize high-quality structures. Figure 1b illustrates this issue with an example (PDBID: 7PS0). AF3 generates a high-quality model (DockQ = 0.6) but ranks it at 498th, while the Top-ranked model fails to form the correct contacts with the antigen.
Design and optimize DeepRank-Ab in terms of graph level, features, and sampling strategies
Having identified the scoring limitations of AF3, we developed DeepRank-Ab guided by two key principles: (i) constructing reliable and diverse training data, and (ii) identifying an effective structural representation for training. To this end, we curated 1442 antibody–antigen complexes from SAbDab after stringent quality and redundancy filtering and generated a new docking benchmark with HADDOCK325. To prevent data leakage, complexes were clustered using Foldseek-Multimer27, and the resulting clusters were split into training/validation splits and test sets. Details on the dataset curation and clustering procedure are provided in Methods under “Strict structural splitting to avoid data leakage”.
We formulate DeepRank-Ab as a regression task aimed at predicting continuous DockQ scores (pDockQ). This allows the model to distinguish near native from incorrect predictions while ranking them accordingly, preserving the full spectrum of model quality instead of collapsing all acceptable models into a binary category.
Docking benchmarks are often redundant, raising a question of how to best leverage such large datasets. We therefore focused on two rarely discussed but critical design choices: (i) the level of structural information encoded in data representation, and (ii) the strategy to sample models from the benchmark dataset for training and testing.
Graph design
Graph-based learning has proven effective across many structural biology applications29,30. For protein–protein interactions, previously we used residue-level graph representations for the scoring task. However, because antibody-antigen interactions rely on CDR loops that are highly flexible and structurally complex31, we therefore examined whether representing structures at the atomic level could improve performance.
We further expanded the feature set to enrich the graph representation, adding CDR region labels for nodes feature; edge geometry and energetics, as well as atom–atom contact areas derived from Voronoi tessellation32 as edge features. Full descriptions of these features are provided in the “Methods” section.
Sampling strategies for training and evaluation sets
A major challenge in training scoring functions is the distribution mismatch between training data and real docking outputs, where near-native decoys are rare and conformational changes are often observed. How to best sample decoys from docking benchmarks for efficient training remains an underexplored area in the field. Here, we considered two complementary motivations: first, balanced sampling, which stabilizes the training process by ensuring a more even representation of decoy qualities. The second is targeted upsampling of low-quality decoys (DockQ < 0.23), based on the observation that near-native models tend to resemble one another, whereas low-quality decoys, although more abundant, can fail in various ways. To systematically evaluate these considerations, we compared two sampling strategies: (i) a balanced sampling approach and (ii) targeted upsampling of low-quality (DockQ < 0.23) decoys. The overview of our DeepRank-Ab workflow is shown in Fig. 2. Refer to the “Methods” section for details of the sampling method.
Fig. 2. The various stages of the DeepRank-Ab workflow.

a Data generation: experimental structures were re-docked in different scenarios, resulting in 2.3 million models, from which representative datasets were sampled. b Feature definition: new antibody-specific and physics-based descriptors were added to the existing features in DeepRank-GNN-esm. c Interfaces were represented as graphs and used as inputs to train equivariant graph neural networks.
Model training and evaluation
Using combinations of the two sampling strategies and the two graph levels as well as feature combinations that we discussed above, we trained six DeepRank-Ab variants and evaluated them using fivefold cross-validation. The splitting between different training/validation folds was defined based on the structural clustering described in “Methods” section.
Although DeepRank-Ab is trained to predict continuous pDockQ scores, we also assessed its classification performance. Both predicted and ground-truth DockQ values were binarized using a threshold of 0.23 to distinguish near-native from non-native models. Standard binary metrics, including AUC, Accuracy, Precision, and Recall, were used to assess model performance. Full definitions of these metrics are provided in the Supplementary Method 5.
We report per-complex statistics because, in antibody–antigen scoring, overall AUC values can appear high even when performance on individual complexes is poor; per-complex evaluation is therefore more informative. Figure 3 summarizes the average per-complex MSE and AUC for each model across the five folds. Complementing the per-complex analysis, Table 1 reports the overall regression and binary classification performance of each DeepRank-Ab variant.
Fig. 3. Comparison of model performance across graph representations and sampling strategies.

Per-complex area under curve (AUC) (a) and per-complex MSE (b) for models trained using residue-level or atom-level graphs under “balanced” or “upsampled” sampling methodology (see “Methods”). “residue_dist” corresponds to residue-level graphs with distance-based edges; “atom_dist” and “atom_area” denote atom-level graphs using Euclidean distance or Voronoi contact area, respectively. Boxplots show the distribution of values across complexes averaged over fivefold cross-validation. Error bars represent standard deviation across folds.
Table 1.
Overall performance of DeepRank-Ab variants across fivefold cross-validation on the benchmark evaluation set
| DeepRank-Ab variants | ||||||
|---|---|---|---|---|---|---|
| residue_dist | atom_dist | atom_area | residue_dist | atom_dist | atom_area | |
| Sampling | balanced | balanced | balanced | upsampled | upsampled | upsampled |
| RMSE (mean ± std) | 0.165 ± 0.002 | 0.165 ± 0.004 | 0.164 ± 0.003 | 0.157 ± 0.003 | 0.163 ± 0.005 | 0.155 ± 0.005 |
| R² (mean ± std) | 0.578 ± 0.011 | 0.575 ± 0.018 | 0.583 ± 0.012 | 0.508 ± 0.018 | 0.468 ± 0.035 | 0.518 ± 0.026 |
| AUC (mean ± std) | 0.784 ± 0.012 | 0.786 ± 0.014 | 0.787 ± 0.010 | 0.803 ± 0.011 | 0.794 ± 0.017 | 0.808 ± 0.016 |
| Accuracy (mean ± std) | 0.705 ± 0.010 | 0.706 ± 0.012 | 0.708 ± 0.009 | 0.777 ± 0.005 | 0.768 ± 0.011 | 0.784 ± 0.007 |
| Precision (mean ± std) | 0.750 ± 0.010 | 0.762 ± 0.013 | 0.757 ± 0.009 | 0.641 ± 0.013 | 0.622 ± 0.023 | 0.671 ± 0.014 |
| Recall (mean ± std) | 0.710 ± 0.013 | 0.690 ± 0.019 | 0.704 ± 0.029 | 0.542 ± 0.026 | 0.534 ± 0.021 | 0.513 ± 0.028 |
This table reports the global regression and binary classification metrics (mean ± standard deviation) for all six DeepRank-Ab variants. Models differ in graph representation (residue- or atom-level), interaction encoding (Euclidean distance vs. Voronoi contact area), and sampling strategy (balanced vs. upsampled). Metrics include RMSE, R², AUC, Accuracy, Precision, and Recall, computed across all complexes in the dataset during evaluation.
Four variants used atom-level graphs and differed in how interatomic interactions are encoded, either by Euclidean distance (“atom_dist”) or by Voronoi contact area (“atom_area”), and in their sampling strategy (“balanced” vs. “upsampled”). Motivated by the strong performance of Voronoi-based scoring approaches17,33, we examined explicitly how the area-based representations compare to distance-based alternatives.
As shown in Fig. 3, the Voronoi-based “atom_area” representation performs slightly better than “atom_distance” representation, achieving lower or equal MSE under both sampling strategies (0.025 vs. 0.027 for “upsampled” sampling; 0.025 vs 0.025 for “balanced” sampling) and higher AUC or equal (0.801 vs. 0.785 for “upsampled” sampling; 0.799 vs 0.799 for “balanced” sampling). Overall, atom-level graphs outperform residue-level ones, indicating that a detailed structural representation improves accuracy. The “upsampled” variants show slightly better results, suggesting that enriching high-quality and diverse low-quality decoys benefits training.
Ablation study and feature importance analysis
To quantify the importance of individual features, we performed ablation studies on the best model (“upsampled atom_area”) using fivefold cross-validation. Additional ablation experiments were also conducted on the full dataset, which refers to performing training and ablation on the entire training set without splitting it into folds. For a complete list of node and edge features used in this baseline model, refer to Supplementary Table 1.
Figure 4a demonstrates the impact of removing each feature in turn, reported as pDockQ ΔRMSE relative to the baseline model (RMSE = 0.159). Removing region labels, atom type, covalent interactions, orientation, or Voronoi contact area caused the largest increases in RMSE and a larger variance across training folds, indicating that these features are essential for both accuracy and training stability. In contrast, removing ESM-2 embeddings, residue type, or charge caused only minor performance changes, suggesting partial redundancy. Ablations on the full dataset yielded similar trends (Supplementary Fig. 1).
Fig. 4. Comparative performance and ablation analysis of docking scoring methods across benchmark and AF3 test sets.

a Ablation study (fivefold cross-validation). We report the mean RMSE ± SD obtained when each feature group is removed, relative to the full baseline model. b Top k success rate of 5 methods on 215 complexes in benchmark test set, including 3 DeepRank-Ab variants (“atom_area”, “atmo_distanc”,” residue_distance”), HADDOCK EMscoring, and VoroIF-jury. c, d Performance comparison of eight methods on 59 complexes from the AF3 test set, including AF3, three DeepRank-Ab variants with electrostatic features (“atom_area”, “atom_distance”, “residue_distance”), two variants without electrostatics (“atom_area_no_elec”, “atom_distance_no_elec”), HADDOCK EMScoring, and VoroIF-jury. Panel c shows the Top K success rate; panel d shows the best Top K DockQ. Refer to S3 Method 5 for a detailed description of the feature set used by each model. e Distribution of Top 1 DockQ across the AF3 test set for all methods. Error bars represent standard deviation across complexes. f Per-complex comparison of the correlation between predicted scores and DockQ, comparing DeepRank-Ab with AF3 ranking score. Each point represents one complex.
As part of an ablation study on the full dataset, we evaluated the impact of incorporating fine-tuned ESM-2 embeddings, compared to standard ESM-2 embeddings, on model performance. These embeddings were generated using our AbTune protocol34, which performs single-sequence fine-tuning of ESM-2 at inference time. A detailed description of our fine-tuning protocol is available in Supplementary Method 6. Interestingly, ablating fine-tuned embeddings led to a modest increase in ΔRMSE, whereas ablating raw ESM-2 embeddings did not, highlighting the potential benefit of antibody-specific fine-tuning (Supplementary Fig. 1). However, this effect was only marginal and did not improve Top k success rate, especially when K is small (See Results). Considering that including the fine-tuning step would substantially increase the complexity of DeepRank-Ab, we chose to use standard ESM-2 embeddings in the final model.
Performance on the test sets: docking benchmark test sets and AF3 test sets
Performance on the docking benchmark test sets
To rigorously assess model performance under realistic docking conditions, we restricted our test set to the most challenging scenario: unbound docking. Including models from the other three docking conditions would greatly inflate the Top 1 success rate because, in our experience, high-quality structures, particularly those produced by the HADDOCK refinement protocol, are very easy to rank.
We retrained the final models incorporating insights from the ablation study and compared them against the HADDOCK score obtained using the HADDOCK3 emscoring module and VoroIF-jury17. In Fig. 4b, all three DeepRank-Ab variants outperform these baselines. The “atom_dist” model yields the highest Top 1 success rate, whereas “atom_area” performs best at Top 10. Classification and regression metrics of these DeepRank-Ab variants are also reported at Table 2.
Table 2.
Overall performance of three DeepRank-Ab variants on benchmark test sets (n = 215)
| Graph representation | Regression metrics | Classification metrics | Success rate (%) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| RMSE | R2 | AUC | Accuracy | Precision | Recall | T1 | T5 | T10 | |
| residue_dist | 0.156 | 0.525 | 0.808 | 0.786 | 0.641 | 0.544 | 23.47 | 38.5 | 44.6 |
| atom_dist | 0.155 | 0.527 | 0.812 | 0.795 | 0.664 | 0.549 | 24.88 | 40.38 | 45.07 |
| atom_area | 0.16 | 0.498 | 0.79 | 0.769 | 0.594 | 0.564 | 19.72 | 40.85 | 46.01 |
This table reports the regression, binary classification metrics, as well as the Top K success rate (%). Variants differ in graph representation (residue- or atom-level), interaction encoding (Euclidean distance vs. Voronoi contact area).
We further investigated whether DeepRank-Ab’s ranking performance was associated with structural characteristics of the complexes, such as interface size or the lengths of key CDR loops (CDR-H1, CDR-H2, CDR-H3). While no significant differences were observed for interface size or CDR-H1/H2 length (p ≥ 0.05; see Supplementary Table 2), complexes with shorter CDR-H3 loops were more likely to have an acceptable model ranked at Top 1 by DeepRank-Ab. This is to be expected, as longer H3 loops are more challenging to model and rank. Moreover, antigen length showed a strong effect: DeepRank-Ab ranked peptide antigens far more accurately than protein antigens in the test set (Top 1 success rate 52.38% vs. 11.56%, Top 5 success rate 88.09% vs 28.9%). Here, peptide antigens are defined as antigens with fewer than 50 residues, following the convention used in SAbDab26. Using this definition, the 215 complex test set contains 42 peptide–antigen complexes.
Several factors likely contribute to the superior performance on peptides. First, antibody–peptide interfaces are known to have greater hydrophobic burial and higher shape complementarity35, which may make them easier for the model to predict. Second, our graph construction uses a relatively small heavy atom distance cutoff of 5 Å, focusing on local atomic interactions at the interface. This design choice is particularly well-suited for peptide antigens, which are short and whose binding is dominated by local contacts rather than long-range interactions. In contrast, protein antigens often involve more extended interfaces, which are more challenging for graph-based learning approaches to capture. Collectively, the combination of favorable interface properties and a graph representation tailored to local interactions likely explains the strong performance of our method on antibody–peptide complexes.
Although protein language-model embeddings showed only modest improvements in RMSE in our ablation study, they provided clear benefits for our primary objective—accurate ranking of top models rather than precise prediction of every DockQ value. However, we did not observe further gains when using fine-tuned embeddings.
Performance on external test sets
We next evaluated whether scoring functions trained on HADDOCK decoys generalize to AF3-generated structures. All tested methods outperform AF3 ranking score on the AF3 test set (Fig. 4c). The best model (“atom_area”) achieves a Top 1 success rate of 49.15% (vs. AF3 40.67% and HADDOCK 45.7%).
During analysis, we also observed that unrelaxed AF3 structures often exhibited energetic features that differed from those in our training set, likely due to local steric clashes or suboptimal side-chain packing. To mitigate this distribution shift, we trained additional models in which the electrostatic energy term was removed (“atom_area_no_elec” and “atom_dist_no_elec”). This adjustment further improved performance, increasing the Top 1 success rate from 49.15% to 54.24% (Fig. 4c), making DeepRank-Ab the best-performing method on this dataset.
Beyond success rates, DeepRank-Ab is also more effective at prioritizing high-quality structures. As shown in Fig. 4d, the mean Top 1 DockQ increased by 35.5%, from 0.321 for the AF3 baseline to 0.435 with our best-performing method. In 99% of cases, when DeepRank-Ab correctly identified a Top 1 structure among the HADDOCK-generated models for a given complex, it also recovered the correct Top 1 model from AF3. This further illustrates the extent to which DeepRank-Ab generalizes across different model-generation protocols. Examining the Top 1 DockQ distribution further highlights DeepRank-Ab’s performance (Fig. 4e): despite the wide variability of DockQ values across targets, DeepRank-Ab not only increases the average DockQ but also raises the median DockQ more than sevenfold, from 0.074 to 0.546 comparing to AF3.
In addition, we evaluated the pairwise correlation between DeepRank-Ab pDockQ, DockQ, and AF3 ranking score. On a per-complex basis (Fig. 4f), DeepRank-Ab outperforms AF3 in most cases: for 43 out of 59 complexes, the correlation between DeepRank-Ab pDockQ and DockQ lies above the diagonal, meaning that in 73% of the cases DeepRank-Ab provides a better scoring solution.
We further examined the global distribution of scores across all complexes and identified systematic differences between the two scoring functions. Specifically, DeepRank-Ab pDockQ exhibits a substantially broader range that could capture variations in DockQ, whereas the AF3 ranking score is markedly skewed toward higher values and rarely assigns low scores, even for complexes with poor DockQ (Supplementary Fig. 2).
Finally, motivated by the success of the consensus scoring approach in previous rounds of CAPRI36, we evaluated two consensus strategies: by averaging the model predictions and or by applying a jury-based voting method as in VoroIF-GNN33, combining predictions from our best model with AF3. However, both strategies reduced performance (Table 3). Consensus scoring is most effective when the combined methods contribute complementary information. In our case, our top-performing model is already highly optimized; combining it with other methods dilutes rather than improves its ranking power.
Table 3.
Comparison of regression metrics, classification metrics, and Top K success rates for DeepRank-Ab variants, consensus scoring strategies, and AF3’s native ranking score on 59 AF3-generated antibody–antigen complexes
| Regression metrics | Binary classification metrics | Success rate (%) | ||||||
|---|---|---|---|---|---|---|---|---|
| Graph representation | RMSE | AUC | Accuracy | Precision | Recall | T1 | T5 | T10 |
| residue_dist | 0.337 | 0.682 | 0.572 | 0.467 | 0.749 | 35.59 | 49.15 | 57.63 |
| atom_dist | 0.477 | 0.839 | 0.403 | 0.399 | 0.999 | 40.68 | 57.63 | 61.02 |
| atom_area | 0.518 | 0.851 | 0.4 | 0.398 | 0.999 | 42.37 | 57.63 | 61.02 |
| atom_dist_no_elec | 0.431 | 0.786 | 0.537 | 0.459 | 0.951 | 45.76 | 59.32 | 62.71 |
| atom_area_no_elec | 0.399 | 0.853 | 0.542 | 0.464 | 0.978 | 54.24 | 61.02 | 64.41 |
| consensus_jury | N/A | 37.28 | 45.76 | 54.24 | ||||
| consensus_average | 50.84 | 61.02 | 61.02 | |||||
| AF3 | 40.68 | 54.24 | 59.32 | |||||
Variants differ in graph representation (residue- or atom-level), interaction encoding (Euclidean distance vs. Voronoi contact area). “no_elec” variant means the model variant trained without electrostatic features. The two consensus methods are: (1) consensus_jury, a jury-based voting scheme combining AF3 with our best model (“atom_area_no_elec”), and (2) consensus_average, which uses the averaged ranks from AF3 and “atom_area_no_elec” as the final ranking.
To further assess DeepRank-Ab, we evaluated it on an external dataset of five MassiveFold37 targets of antibody–antigen complexes taken from recent rounds of CAPRI36. Re-clustering these targets together with our training set using Foldseek-multimer showed that each formed a singleton cluster, indicating that their binding interfaces were not similar to those of any complexes in the training set.
Four of these targets involve peptide antigens. DeepRank-Ab showed very strong performance on this dataset, achieving a 100% success rate at Top 5, while MassiveFold reached only 60% (Supplementary Fig. 3). These results are consistent with our earlier observations that DeepRank-Ab is particularly effective for peptide–antigen complexes.
Methods
Dataset construction and generation of the docking benchmark
We constructed a non-redundant dataset of antibody–antigen complexes from SAbDab26 by first retrieving all eligible structures and then applying several filtering steps. Concatenated CDR sequence filtering using 100% identity cutoff and structural quality thresholds (resolution ≤ 3.5 Å; R-factor ≤ 0.25) was used to remove redundant or low-quality entries, along with additional checks to exclude incomplete antibodies and non-protein antigens. For antigens composed of multiple chains, only those in contact with the antibody were retained. To improve the modeling efficiency, antigens longer than 500 residues were segmented using Merizo38, and only high-confidence domains (score > 0.7) containing interface residues were kept. Entries for which ABodyBuilder239 failed to generate antibody models or for which AF2 yielded poor antigen predictions (RMSD > 8 Å) were removed. After these steps, the final curated dataset comprised 1,442 antibody–antigen complexes, which served as input to HADDOCK325.
When generating the docking benchmark, we had several design objectives. First, we aimed to obtain a diverse set of low-quality (DockQ < 0.23) models while also ensuring that each complex contained at least one high-quality model. Conformational diversity was introduced through a combination of docking and refinement procedures. To achieve this, we applied four HADDOCK3 protocols, including two docking scenarios (bound–bound docking and unbound–unbound docking) and two refinement scenarios (bound refinement and unbound refinement). Details are provided in the Supplementary Method 1.
To construct representative decoy sets for training, we applied two sampling strategies. In refinement-only scenarios, all generated models were retained. In docking scenarios, decoys from rigid-body docking and energy minimization were grouped by DockQ score into bins of width 0.1, clustered, and a fixed number of models was selected from each bin and scenario (“balanced” sampling). In the second strategy (“upsampled”), low-quality decoys (i.e., those in bin 0 and bin 1) were upsampled by a factor of 20 (See Supplementary Method 2, sampling statistics in Supplementary Fig. 4), leveraging HADDOCK’s wide conformational search to introduce additional structural diversity and improve training robustness.
Strict structural splitting to avoid data leakage
To prevent data leakage during training, all complexes were first clustered with Foldseek-Multimer27 with thresholds of 0.65 for multimer TM, 0.5 for chain-level TM, and 0.65 for interface LDDT, defining interfaces with an 8.5 Å cutoff. Results showed that 89% of complexes end up as singletons, indicating adequate diversity in the dataset. Approximately 10% were grouped into clusters of size 2–9, and only 0.3% belonged to larger clusters (see Supplementary Table 3 for clustering statistics). The dataset was then partitioned at the cluster level: clusters were randomly ordered and assigned to the training/validation split until ~85% of all complexes were included, with the remaining ~15% allocated to the test set. All decoys generated from a given complex were assigned to the same split as their parent complex.
Graph representation, new features, and model architecture
We formulated the scoring task as a regression problem in which a docking model is given as input, and the network predicts an interface-quality score (pDockQ). Antibody–antigen interfaces were encoded as graphs at two levels of detail. First, we constructed residue-level graphs, with nodes and edges defined following our previous DeepRank-GNN-esm framework18, a general-purpose scoring function to rank PPI poses that integrates graph representations with protein language model embeddings from ESM-2, eliminating the need for computationally intensive PSSM features.
Second, we built atom-level graphs, in which heavy atoms serve as nodes and edges are defined either by interatomic distances or by contact areas computed using Voronoi tessellation32. To enrich these representations, we introduced several new node- and edge-level features: nodes were annotated with antibody-specific region labels derived from IMGT numbering40, and edges were assigned geometric descriptors capturing local residue orientation, along with energetic and covalent interaction terms adapted from DeepRank241.
For the model architecture, we adopted an E(n)-equivariant GNN42, chosen for its ability to incorporate geometric information while preserving rotational and translational symmetries. Following the design principles of DeepRank-GNN19, we implemented a two-branch architecture that processes interface edges and internal edges separately, with each branch consisting of seven EGNN layers. Edge messages integrate node features, RBF-encoded distances, and the full set of edge attributes, while coordinate updates are computed via learned transformations. Node updates employ residual connections, LayerNorm, and dropout. After message passing, we apply hierarchical pooling followed by a Global Attention readout. The resulting embeddings from both branches are concatenated and passed through a lightweight prediction head with a residual connection to generate the final pDockQ score.
For each model, we performed fivefold cross-validation using the cluster-aware partitioning described above. After completing cross-validation, we trained our final models on the full training dataset using the same training protocol. These fully trained models were then used to evaluate performance on the benchmark test set, AF3 test set, and MassiveFold decoy set.
Full description of graph design, feature implementation, and training details can be found in Supplementary Methods 3 and 4. A description of all metrics we used to evaluate model performance is provided in Supplementary Method 5.
Curation of external test sets
We evaluated our method on three independent test sets: 215 complexes from our docking Benchmark test set, 59 from AF3 test set, and 5 five sets of MassiveFold37 decoys from previous rounds of CAPRI36,43 targets. For the docking Benchmark test set, we included only unbound–unbound docking models with upsampled acceptable, lower-quality models. For the AF3 test set, we generated 500 structures per complex using a local installation of AF34 with 100 random seeds. The five CAPRI targets (T231, T234, T266, T268, and T270) were used as an external test set, with an average of 8046 models generated per target using MassiveFold37.
Statistics and reproducibility
Statistical analyses were performed as described in the main text and figure legends.
Discussion
In this work, we addressed the scoring challenge for antibody–antigen complexes and showed that existing methods, including AF3, have substantial limitations in their ability to rank near-native structures. Results from previous CAPRI36 rounds further indicate that scoring remains a critical bottleneck. For some antibody–antigen targets, while acceptable models do exist in the pool, neither AF nor current scoring methods could identify them, highlighting the need for new methods.
To support the development of DeepRank-Ab, we curated a high-quality and structurally diverse docking benchmark specifically tailored to antibody–antigen interactions. This dataset was generated using multiple HADDOCK3 workflows, including both docking and refinement protocols, combined with a sampling strategy designed to maximize structural diversity.
Using this dataset, we conducted a systematic evaluation of graph representation and feature sets and developed DeepRank-Ab, including a suite of scoring functions optimized for structures produced by both docking pipelines and AF3. Across multiple independent test sets, DeepRank-Ab consistently demonstrated superior performance, achieving the highest Top k success rates and Top k DockQ values compared with AF3, HADDOCK scoring, and VoroIF-jury.
Compared with existing scoring functions, DeepRank-Ab leverages an enriched graph representation that integrates geometric, physicochemical, and evolutionary features to describe antibody–antigen interfaces. Inspired by the top scorers in previous CAPRI rounds, additional graph edge features were defined based on Voronoi tessellation, which captures interfacial packing and contact topology more accurately than distance-based graphs. Combined with training on a high-quality benchmark dataset with EGNN, DeepRank-Ab achieves strong predictive performance.
Despite these advances, we still observed an upper limit to the model’s performance: compared with the oracle success rate, a gap remains. This gap likely arises from multiple factors. Fundamentally, scoring is inherently a sub-problem of binding site prediction. Despite our best efforts in characterizing their binding interfaces, antibody–antigen recognition remains extremely complicated. Multiple interfaces may appear equally favorable, yet only one corresponds to the true native binding mode, influenced by factors such as binding entropy and enthalpy, interface stability, and correct modeling of conformational changes—many of which are challenging to capture computationally. This complexity is also linked to the challenge of antibody binder selection44, as identifying functional binders essentially requires distinguishing the true native interface.
In addition to these intrinsic limitations, several technical considerations could further improve DeepRank-Ab. These include incorporating global features of the entire complex, training with decoys from multiple docking platforms to reduce bias, and implementing additional checks to ensure the structural integrity of antibodies. For example, as sometimes observed in the MassiveFold dataset based on our experience, the two chains of antibodies are not properly packed, highlighting the need for stricter quality control prior to ranking.
We anticipate that supplementing AF3 with DeepRank-Ab should compensate, at least in part, for the absence of explicitly incorrect structures in AF3’s training regime. Although AF3 is indirectly exposed to suboptimal conformations during its recycling procedure and possibly through entries in AlphaFoldDB4, systemic datasets such as docking decoys may still introduce additional types of structural variation that are not represented in its training data.
Supplementary information
Acknowledgements
We acknowledge Marc Lensink and Guillaume Brysbaert, University of Lille, France, for providing the five MassiveFold decoys.
Author contributions
X.X., I.C. and A.M.J.J.B. designed the research. X.X. and I.C. conducted the research and analyzed the results. I.C. generated the HADDOCK3 docking benchmark. X.X. and I.C. contributed to graph design, model training, and evaluation. V.R. contributed to graph featurization. X.X. wrote the original draft and handled the revisions. All authors reviewed the paper.
Peer review
Peer review information
Communications Biology thanks Alissa M. Hummer and Brian Pierce for their contribution to the peer review of this work. Primary Handling Editor:Christina Karlsson Rosenthal. A peer review file is available.
Funding
Financial support from the European High Performance Computing Joint Undertaking, project BioExcel (101093290), is acknowledged. X. Xu acknowledges financial support from the China Scholarship Council (grant No. 202208310024).
Data availability
All data involved in this work are available at https://zenodo.org/records/17911452.
Code availability
DeepRank-Ab is available online at https://github.com/haddocking/DeepRank-Ab.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Xiaotong Xu, Ilaria Coratella.
These authors jointly supervised this work: Xiaotong Xu, Alexandre MJJ Bonvin.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s42003-026-10408-4.
References
- 1.Rees, A. R. Understanding the human antibody repertoire. mAbs12, 1729683 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Peng, H.-P. et al. Antibody CDR amino acids underlying the functionality of antibody repertoires in recognizing diverse protein antigens. Sci. Rep.12, 12555 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Norman, R. A. et al. Computational approaches to therapeutic antibody design: established methods and emerging trends. Brief Bioinform.21, 1549–1567 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.The OpenFold3 Team. OpenFold3-preview. (2025).
- 7.Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. bioRxiv 10.1101/2025.06.14.659707 (2025). [DOI]
- 8.Ünsal, S., Holland, B., Sardag, I. & Timucin, E. Confidence scoring for deep learning-predicted antibody–antigen complexes: AntiConf as a precision-driven metric. Brief. Bioinform.27, bbag137 (2026). [DOI] [PMC free article] [PubMed]
- 9.Zhu, W., Shenoy, A., Kundrotas, P. & Elofsson, A. Evaluation of AlphaFold-Multimer prediction on multi-chain protein complexes. Bioinformatics39, btad424 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Varga, J. K., Ovchinnikov, S. & Schueler-Furman, O. actifpTM: a refined confidence metric of AlphaFold2 predictions involving flexible regions. Bioinformatics41, btaf107 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Dunbrack, R. L. Rēs ipSAE loquunt: what’s wrong with AlphaFold’s ipTM score and how to fix it. bioRxiv 10.1101/2025.02.10.637595 (2025). [DOI]
- 12.Fromm, S., Ludaic, M. & Elofsson, A. Evaluating deep learning based structure prediction methods on antibody–antigen complexes. Bioinformatics42, btag136 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Verburgt, J., Zhang, Z. & Kihara, D. AlphaFold model quality self-assessment improvement via deep graph learning. Protein Sci.34, e70274 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Spoendlin, F. C. et al. Predicting the conformational flexibility of antibody and T cell receptor complementarity-determining regions. Nat. Mach. Intell.7, 1755–1767 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Xu, X., Giulini, M. & Bonvin, A. M. J. J. Improved prediction of antibody and their complexes with clustered generative modelling ensembles. Bioinforma. Adv.5, vbaf161 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Gaudreault, F., Corbeil, C. R. & Sulea, T. Enhanced antibody-antigen structure prediction from molecular docking using AlphaFold2. Sci. Rep.13, 15107 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Olechnovič, K., Banciul, R., Dapkūnas, J. & Venclovas, Č. FTDMP: a framework for protein–protein, protein–DNA, and protein–RNA docking and scoring. Proteins 10.1002/prot.26792 (2026). [DOI] [PubMed]
- 18.Xu, X. & Bonvin, A. M. J. J. DeepRank-GNN-esm: a graph neural network for scoring protein–protein models using protein language model. Bioinform. Adv.4, vbad191 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Réau, M., Renaud, N., Xue, L. C. & Bonvin, A. M. J. J. DeepRank-GNN: a graph neural network framework to learn patterns in protein–protein interfaces. Bioinformatics39, btac759 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Renaud, N. et al. DeepRank: a deep learning framework for data mining 3D protein-protein interfaces. Nat. Commun.12, 7068 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wang, X., Terashi, G., Christoffer, C. W., Zhu, M. & Kihara, D. Protein docking model evaluation by 3D deep convolutional neural networks. Bioinformatics36, 2113–2118 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Stratiichuk, R. et al. Sampling and ranking of protein conformations using machine learning techniques do not improve the quality of rigid protein–protein docking. J. Chem. Inf. Model.65, 10167–10179 (2025). [DOI] [PubMed] [Google Scholar]
- 23.Bushuiev, A. et al. Learning to design protein-protein interactions with enhanced generalization. Preprint at arXiv 10.48550/arXiv.2310.18515 (2024). [DOI]
- 24.Dominguez, C., Boelens, R. & Bonvin, A. M. J. J. HADDOCK: a protein−protein docking approach based on biochemical or biophysical information. J. Am. Chem. Soc.125, 1731–1737 (2003). [DOI] [PubMed] [Google Scholar]
- 25.Giulini, M. et al. HADDOCK3: a modular and versatile platform for integrative modeling of biomolecular complexes. J. Chem. Inf. Model. 10.1021/acs.jcim.5c00969 (2025). [DOI] [PMC free article] [PubMed]
- 26.Dunbar, J. et al. SAbDab: the structural antibody database. Nucl. Acids Res.42, D1140–D1146 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Kim, W. et al. Rapid and sensitive protein complex alignment with Foldseek-Multimer. Nat. Methods22, 469–472 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Basu, S. & Wallner, B. DockQ: a quality measure for protein-protein docking models. PLoS ONE11, e0161879 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Corso, G., Stark, H., Jegelka, S., Jaakkola, T. & Barzilay, R. Graph neural networks. Nat. Rev. Methods Prim.4, 17 (2024). [Google Scholar]
- 30.Zhou, J. et al. Graph neural networks: A review of methods and applications. AI Open1, 57–81 (2020).
- 31.Regep, C., Georges, G., Shi, J., Popovic, B. & Deane, C. M. The H3 loop of antibodies shows unique structural characteristics. Proteins85, 1311–1318 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Olechnovič, K. & Grudinin, S. Voronota-LT: efficient, flexible, and solvent-aware tessellation-based analysis of atomic interactions. J. Comput. Chem.46, e70178 (2025). [DOI] [PubMed] [Google Scholar]
- 33.Olechnovič, K. & Venclovas, Č VoroIF-GNN: voronoi tessellation-derived protein–protein interface assessment using a graph neural network. Proteins91, 1879–1888 (2023). [DOI] [PubMed] [Google Scholar]
- 34.Xu, X. & Bonvin, A. M. J. J. AbTune: layer-wise selective Fine-Tuning of protein language models for Antibodies. bioRxiv 10.1101/2025.10.17.682998 (2025). [DOI] [PMC free article] [PubMed]
- 35.Lee, J. H., Yin, R., Ofek, G. & Pierce, B. G. Structural features of antibody-peptide recognition. Front. Immunol.13, 910367 (2022). [DOI] [PMC free article] [PubMed]
- 36.Lensink, M. F. et al. Biomolecular interaction prediction in the pre- and post-AlphaFold Era: the 8th CAPRI evaluation. Proteins 10.1002/prot.70018 (2025). [DOI] [PubMed]
- 37.Raouraoua, N. et al. MassiveFold: unveiling AlphaFold’s hidden potential with optimized and parallelized massive sampling. Nat. Comput. Sci.4, 824–828 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Lau, A. M., Kandathil, S. M. & Jones, D. T. Merizo: a rapid and accurate protein domain segmentation method using invariant point attention. Nat. Commun.14, 8445 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Abanades, B. et al. ImmuneBuilder: deep-Learning models for predicting the structures of immune proteins. Commun. Biol.6, 575 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Dunbar, J. & Deane, C. M. ANARCI: antigen receptor numbering and receptor classification. Bioinformatics32, 298–300 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Crocioni, G. et al. DeepRank2: mining 3D protein structures with geometric deep learning. J. Open Source Softw.9, 5983 (2024). [Google Scholar]
- 42.Satorras, V. G., Hoogeboom, E. & Welling, M. E(n) equivariant graph neural networks. Preprint at arXiv 10.48550/arXiv.2102.09844 (2022). [DOI]
- 43.Reys, V. et al. Integrative modeling in the age of machine learning: a summary of HADDOCK strategies in CAPRI rounds 47–55. Proteins 10.1002/prot.26789 (2026). [DOI] [PubMed]
- 44.Smorodina, E. et al. Structural plausibility without binding specificity: limits of AI-based antibody-antigen structure prediction confidence scores. bioRxiv 10.64898/2026.03.02.709004 (2026).42051289 [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All data involved in this work are available at https://zenodo.org/records/17911452.
DeepRank-Ab is available online at https://github.com/haddocking/DeepRank-Ab.
