Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Dec 10.
Published in final edited form as: Nat Biotechnol. 2024 Nov;42(11):1637–1642. doi: 10.1038/s41587-024-02469-9

AIntibody: an experimentally-validated in silico antibody discovery design challenge

M Frank Erasmus 1,*, Laura Spector 1, Fortunato Ferrara 1, Roberto DiNiro 1, Thomas J Pohl 1, Katheryn Perea-Schmittle 1, Wei Wang 2, Peter M Tessier 3, Crystal Richardson 4, Laure Turner 4, Sumit Kumar 4, Daniel Bedinger 5, Pietro Sormanni 6, Monica L Fernández-Quintero 7, Andrew B Ward 7, Johannes R Loeffler 7, Olivia M Swanson 7, Charlotte M Deane 8, Matthew I J Raybould 8, Andreas Evers 9, Carolin Sellmann 9, Sharrol Bachas 10, Jeff Ruffolo 11, Horacio G Nastri 12, Karthik Ramesh 12, Jesper Sørensen 13, Rebecca Croasdale-Wood 14, Oliver Hijano 15, Camila Leal-Lopes 15, Yu Qiu 15, Paolo Marcatili 16, Erik Vernet 16, Rahmad Akbar 16, Simon Friedenson 17, Rick Wagner 18, Vinodh babu Kurella 19, Shipra Malhotra 19, Satyendra Kumar 19, Patrick Kidger 20, Juan C Almagro 21, Eric Furfine 22, Marty Stanton 22, Christilyn P Graff 22, Santiago David Villalba 23, Florian Tomszak 23, Andre AR Teixeira 24, Melody Shahsavarian 25, Elizabeth Hopkins 26, Molly Dovner 26, Sara D’Angelo 1, Andrew R M Bradbury 1,*
PMCID: PMC12688036  NIHMSID: NIHMS2126833  PMID: 39496931

To the editor:

Science is frequently subject to the Gartner hype cycle1: emergent technologies spark intense initial enthusiasm with the recruitment of dedicated scientists. As limitations are recognized, disillusionment often sets in: some scientists turn away, disappointed in the inability of the new technology to deliver on initial promise, while others persevere and further develop the technology. While the value (or not) of a new technology usually becomes clear with time, appropriate benchmarks can be invaluable in highlighting strengths and areas for improvement, substantially speeding up technology maturation. A particular challenge in computational engineering and artificial intelligence (AI)/machine learning (ML) is that benchmarks and best practices are uncommon, so it is particularly hard for non-experts to assess the impact and performance of these methods. While multiple papers have highlighted best practices and evaluation guidelines2–4, the true test for such methods is ultimately prospective performance, which requires experimental testing.

In the 1990’s several groups attempted to predict the structure of proteins from amino acid sequences, and the success and value of different models was assessed ad hoc. The Critical Assessment of Structure Prediction (CASP)5 biannual competition was established in 1994 to compare the performance of various algorithms. In this competition, teams used experimental methods (predominantly X-ray crystallography and NMR) to determine protein structure, but only supplied the protein sequences to the modeling community. Expert modeling teams then used a combination of human expertise and computational methods, or fully automated methods to predict the correct structure from the sequence. Teams were allowed to provide up to five models per target. An independent panel compared the predicted and experimentally determined structures. Widely viewed as the “protein structure prediction world championship”, there were incremental improvements in structural predictions until AlphaFold6 won dramatically in 20187 and 20208. Although AlphaFold has not competed since, many methods today are inspired by the AlphaFold architecture, collectively demonstrating the power of deep learning algorithms for protein structure prediction. Other similar competitions aimed at predicting antibody structure from sequence9,10 were initiated but have not been held since 2014.

Thirty years later, the application of AI/ML solutions and other in silico approaches to the development and improvement of proteins11, particularly antibody therapeutics12–14 has led to a desire to understand the capabilities of these technologies. Numerous companies and academic groups claim AI/ML solutions to affinity maturation, antibody developability and de novo antibody and library design. However, significant challenges lie in understanding the value of these algorithms: how they differ from one another in performance, how the quality of predicted antibodies compares to existing experimental practices, how effective they are in generating antibodies recognizing particular epitopes (including against experimentally more challenging targets such as those with membrane, glycan, or flexible components), how generalizable they are, and particularly whether they are able to provide antibodies with the desired properties more rapidly than experimental approaches. Most results are based on retrospective studies (i.e. without conducting new experiments) and without making data accessible. This limits validation of such claims and the real-world impact of these methods. ML methods can perform very differently in prospective studies compared to retrospective ones, for example, in docking onto AlphaFold2 generated structures15. Prospective evaluations of new problems are essential to assess real-world performance. These difficulties emphasize the importance of a public and prospective benchmarking effort to provide a quantifiable and unbiased assessment of these techniques. Here, we propose the launch of an AI/ML benchmarking exercise (named AIntibody and accessed via the eponymous AIntibody.org website) that will present a series of escalating AI/ML challenges in antibody discovery, calibrated according to the success of the outcome of previous competitions.

Launching an AI/ML benchmarking competition

The AIntibody competition aims to validate the performance of computational models, including those based on AI, to generate antibody candidates. Teams are invited to submit antibody sequences in response to the challenges described below. Submitted sequences will be synthesized as proteins and independently benchmarked in a wet lab. The long-term goal is to produce an ongoing challenge series analogous to CASP. As the competition evolves, challenges are anticipated to become progressively more difficult, and each will be followed by a peer-reviewed manuscript to provide the community with clarity on the results, highlighting areas for improvement and satisfactory capabilities.

In addition to CASP, this competition has similarities to Critical Assessment of Computational Hit-finding Experiments (CACHE)16, Drug Design Data Resource (D3R) Grand Challenge17, and Statistical Assessment of Modeling of Proteins and Ligands (SAMPL)18. CACHE is an extensive public-private partnership benchmarking initiative recently established to enable the development of computational methods for drug hit-finding and addressing the problem that no “algorithm can currently select, design or rank potent drug-like small-molecule protein binders consistently”. The D3R Grand Challenge focuses on (small molecule) ligand-protein pose and binding affinity prediction, while SAMPL conducts blind prediction challenges in computational drug discovery, with an emphasis on binding modes, affinities, and physical properties for small molecules. Unlike these small molecule competitions, the AIntibody competition is focused on antibodies and uses different organizing principles.

Given the urgent need for benchmarking in AI/ML mediated antibody discovery, this first competition is an unfunded ad hoc contest based on datasets generated against the receptor binding domain (RBD) of SARS-CoV-2 by some of the authors and made available to participants and the broader scientific community, with participants paying a significantly discounted fee [$120/antibody] to cover the costs of gene synthesis and expression of each submitted antibody sequence. As the responses are designed to test the present immediate value of AI/ML algorithms, participants will have fourteen days from data access to provide sequence solutions. This time frame is designed to prevent experimental validation prior to submission, which will be carried out independently and conducted blindly by third parties. The structure of the supplied data can be found in Table 1–2, allowing participants ample time to prepare in advance. We anticipate the ad hoc nature of this first competition will provide deeper understanding and time to organize future contests by more formal partnerships or societies, involving steering committees comprising academic and industry scientists, ideally with external funding.

Table 1.

Expected NGS data and format for competition #1

field / column description
[TYPE]
value, # or range of non-redundant values:
sort_population indicates population of single domain library where
Diversity is inserted
[STRING]
values:
“phase1_h1_h2_am”
“phase1_l1_l2_am”
“phase1_l3_am”
sequence_aa_heavy amino acid
[STRING]
  • phase1_h1_h2_am: 31,729

  • phase1_l1_l2_am: 1

  • phase1_l3_am: 1

sequence_aa_light amino acid
[STRING]
  • phase1_h1_h2_am: 1

  • phase1_l1_l2_am: 122,988

  • phase1_l3_am: 30,379

cdr1_aa_heavy
[IMGT]
amino acid
[STRING]
  • phase1_h1_h2_am: 2,528

  • phase1_l1_l2_am: 1

  • phase1_l3_am: 1

cdr2_aa_heavy
[IMGT]
amino acid
[STRING]
  • phase1_h1_h2_am: 1,023

  • phase1_l1_l2_am: 1

  • phase1_l3_am: 1

cdr3_aa_heavy
[IMGT]
amino acid
[STRING]
  • phase1_h1_h2_am: 1

  • phase1_l1_l2_am: 1

  • phase1_l3_am: 1

cdr1_aa_light
[IMGT]
amino acid
[STRING]
  • phase1_h1_h2_am: 1

  • phase1_l1_l2_am: 4,770

  • phase1_l3_am: 826

cdr2_aa_light
[KABAT]
amino acid
[STRING]
  • phase1_h1_h2_am: 1

  • phase1_l1_l2_am: 7,047

  • phase1_l3_am: 784

cdr3_aa_light
[IMGT]
amino acid
[STRING]
  • phase1_h1_h2_am: 1

  • phase1_l1_l2_am: 996

  • phase1_l3_am: 5,585

redundancy # of times full chain appears in NGS at the amino acid level
[INTEGER]
  • phase1_h1_h2_am: [1 – 12,099] [min – max]

  • phase1_l1_l2_am: [1 – 494] [min – max]

  • phase1_l3_am: 1,541 [1 – 968] [min – max]

Table 2.

Expected NGS data and format for competition #2–3

field / column description
[TYPE]
value, # or range of non-redundant values:
characterized TRUE if characterized by SPR
[BOOL]
TRUE
  • # VL+VH: 142

  • #L3+H3: 74

  • #H3: 55

  • #clusters: 40

FALSE
  • # VL+VH: 31,917

  • #L3+H3: 4,362

  • #H3: 754

  • #H3 clusters: 300

lsa_bin experimentally determined bin group
[INTEGER]
values = [1, 2, 5, NA]
cluster_cdr3_heavy unique cluster identifier (e.g., 57F)
[STRING]
300
affinity LSA affinity in molarity (M)
[FLOAT]
2.7×10−11 - 2.1×10−3 M
NA = uncharacterized
1.0×10−6 M = characterized, weak affinity
on_rate LSA on-rate in (M−1s−1)
[FLOAT]
1.2×102 - 4.5×105 M−1s−1
NA = uncharacterized OR characterized, slow on-rate
off_rate LSA off-rate in (s−1)
[FLOAT]
1.0×10−5 - 0.6×10−2 s−1
NA = uncharacterized OR characterized, fast off-rate
sequence_aa_light amino acid
[STRING]
102aa to 120aa
sequence_aa_heavy amino acid
[STRING]
113aa to 133aa
cdr1_aa_heavy
[IMGT]
amino acid
[STRING]
7aa to 14aa
cdr2_aa_heavy
[IMGT]
amino acid
[STRING]
6aa to 9aa
cdr3_aa_heavy
[IMGT]
amino acid
[STRING]
6aa to 26aa
cdr1_aa_light
[IMGT]
amino acid
[STRING]
5aa to 12aa
cdr2_aa_light
[KABAT]
amino acid
[STRING]
6aa to 11aa
cdr3_aa_light
[IMGT]
amino acid
[STRING]
5aa to 18aa
relative_abundance_10nM relative abundance of the concatenated CDRs in the 10nM RBD sort round via NGS
[FLOAT]
0.0% - 6.1%
relative_abundance_1nM relative abundance of the concatenated CDRs in the 1nM RBD sort round via NGS
[FLOAT]
0.0% - 9.5%

Testing antibodies for biological activity is complex, however, the challenges presented in this inaugural competition are relatively straightforward, involving only affinity and developability (Table 3), which are the subjects of many AI/ML state-of-the-art performance claims and among the most important determinants of biological activity. Results will be published blinded. Although not obligatory, participants will be invited to be co-authors on published results papers, akin to the approach taken recently when diverse antibodies against SARS-CoV-2 were analyzed19. Participants will be able to see the range of results and their ranking within that range, but individual performances will remain masked. For each challenge described below, the focus lies in engineering only the complementarity determining regions (CDRs) (the IMGT.org definition is used for all CDRs except the LCDR2, for which the definition Kabat is used). For example, frameworks of the antibodies provided in the datasets should not be modified within the context of the challenge. While in vivo affinity maturation is focused on CDRs, framework mutations are often also introduced. The challenges introduced here limit available diversity space to CDRs, providing a direct comparison to experimental methods20. Future challenges may allow the introduction of framework mutations.

Table 3.

Affinity and developability to be experimentally assessed

Assay Metric (units) Company [Instrument]
surface plasmon resonance (SPR) on-rate (M−1s−1),
off-rate (s−1),
affinity (M)
Carterra [LSA] &
Sapidyne [KinExA]
expression expression titers (mg/L) Azenta
[HPLC; SDS-PAGE]
monomer content (if ≥1 mg) % main peak (AUCmain x 100)/(AUCall_peaks)) Azenta
[SEC-HPLC]
self-interaction AC-SINS (Δƛ shift, nm) Mosaic
[ELISA]
thermal stability Tm (°C) Mosaic
[UNCLE]
thermal aggregation Tagg (°C) Mosaic
[UNCLE]
hydrophobicity retention time (min) Mosaic
[HIC HPLC]
polyreactivity BVP (relative response) Mosaic
[ELISA]

This benchmarking exercise will help us understand the power of AI/ML capabilities applied to antibody discovery at this time. The results of this competition will provide insight for establishing realistic timelines for future AI/ML use in antibody discovery.

Competition Details and Guidelines

The challenges proposed in this inaugural competition are based on unpublished SARS-CoV-2 next-generation sequencing (NGS) datasets from which some antibodies were characterized21,22. Given the vast amount of additional public data available for SARS-CoV-2 binding antibodies (e.g., COVIC19 & Cov-AbDab23), the following should provide the best possible scenario for AI/ML task success.

Once designed/identified sequences have been uploaded, Azenta will synthesize genes for antibodies, express and purify them, carry out size exclusion chromatography and provide coded antibodies to other partners. Antibody affinities will be assessed by Carterra using surface plasmon resonance (Carterra LSA-XT) and Mosaic will assess developability (HIC HPLC, BVP ELISA, AC-SINS, Tm and Tagg). Bio-Techne will provide the target. These assays will ensure standardized conditions and unbiased head-to-head comparison of predicted sequences.

Competition #1 – In Silico Antibody Affinity Maturation:

Participants will be provided with experimental datasets derived from the affinity maturation of an antibody recognizing the RBD of SARS-CoV-2 (Table 1). Each dataset comprises three NGS sub-datasets of CDR sequences (LCDR1+2; LCDR3 and HCDR1+2 in which the remaining CDRs are constant) generated during phase 1 of a previously described experimental affinity maturation method20 (Figure 1), in which phase 2 involves combining all phase 1 outputs and experimentally selecting for higher affinity. Each antibody population has diversity only in the indicated CDRs (Figure 1a), the remaining CDRs being parental, and has been displayed on yeast and sorted for target binding (Figure 1b). While each population binds the target more tightly than the parental, individual NGS sequences have not been assessed for their ability to encode antibodies with improved binding activity and may include PCR or sequencing errors. While amino acid sequences of phase 2 characterized antibodies with their affinities have been determined (Figure 1c–d), this data will not be provided so computational methods are given the opportunity to generate the same (or better) sequences.

Figure 1 – Study design of competition #1.

Figure 1 –

a) Phase 1: DNA library diversity is introduced into L1+L2, L3, or H1+H2 of an ant-RBD scFv. b) Selective pressure is applied by FACS using RBD to select for improved binders. c) Phase 2: RBD binding diversity from each of the three arms (L1+L2, L3, H1+H2) is recovered, PCR-amplified and recombined into a yeast display vector. d) The combined diversity is transformed back into yeast and sorted for improved affinity and expression. Sequencing identifies final improved variants, which are reformatted into e) IgG for expression, f) binding affinity and g) developability measurements.

The computational goal of Competition 1 is to design antibodies (Figure 1e) with improved affinity (Figure 1f) for the RBD of SARS-CoV-2 that also exhibit favorable developability properties (Figure 1g) using the NGS datasets. Designs should only be applied to HCDR1–2 and LCDR1–3 and not frameworks or HCDR3. The blinded assessment will determine the affinities of designed antibodies and how well they compare to those obtained experimentally in phase 2. While the experimental affinity maturation was carried out on scFvs displayed on yeast, antibodies were, and will be, tested as full length IgGs. Results will be compared to the affinities of experimentally derived sequences generated by combining the three phase 1 outputs and selecting from the corresponding combinatorial library displayed on yeast.

Competition #2 – In Silico Affinity Rank Prediction for Antibody Discovery:

Participants will be provided with an NGS dataset of a single selection output recognizing the RBD of SARS-CoV-2, clustered by HCDR3 sequence22 using a previously published library24, which comprises natural CDRs embedded within well-behaved therapeutic scaffolds. While not all individual NGS sequences have been assessed for their ability to encode antibodies with binding activity, and therefore may include PCR or sequencing errors, those sequences encoding antibodies (as IgG) demonstrated to bind the target will be identified/provided, with their corresponding affinities. Furthermore, the relative frequency of different sequences within the clusters will be provided, see Table 2 for a representative dataset.

The computational goal of Competition 2 is to identify those sequences within the existing NGS dataset that encode the highest affinity antibodies (that have not already had their affinities determined) in the two largest HCDR3 clusters, by largest number of VL+VH sequences, from experimental bin group #1 [i.e., 28F and 27F] and the largest cluster from experimental bin group #2 [i.e., 47F] (Table 4). Results will be compared to the affinities of those antibodies that were experimentally derived and for which sequences were provided.

Table 4.

Cluster details for competition #2.

Cluster ID LSA bin # unique
VL+VH
# unique
L3+H3
# unique
H3
28F 1 3,554 193 5
27F 1 2,524 152 12
47F 2 400 86 9

The two largest clusters from bin group #1 (28F and 27F) and bin group #2 (47F) with specific cluster details are shown.

Competition #3 – NGS Inspired Computational Antibody Design:

Participants are given the same NGS output as in Competition #2. The AI/ML goal is to generate out-of-library sequences of antibodies binding the same target with as high affinities as possible that also exhibit favorable developability properties, using the provided NGS datasets described here and any other useful publicly available data. Only CDRs should be designed, and frameworks should remain unmodified (Table 2). The out-of-library sequences should not be present within the NGS dataset itself, nor be derived from any previous independent experiments. Results will be compared to the affinities of those antibodies that were experimentally derived and for which sequences were provided.

Each of these challenges has high affinity as an endpoint, but developability will also be assessed to ensure affinity is not generated at the expense of developability. Antibodies will be judged as passing, questionable, or failing on each of the five developability assays described above, with failing antibodies scored 2, questionable antibodies scored 1, and passing antibodies scored 0; any antibody with a total score of 4 or above will be considered failing.

Publication of results and participation requirements

Participants agree to make the winning algorithm (from each of the three competitions) publicly available, in the form of a paper describing the algorithmic details. This may be done anonymously as part of the follow-up manuscript from the competition organizers, or (to avoid autoplagiarism) as part of a separately-authored standalone manuscript. Authors would be required to open-source their code upon journal request. All participants (anonymously, but regardless of winning position) also agree to provide at least a single-sentence description of their methods (e.g., ‘protein language model based on ESM2 and fine-tuned on antibody database XYZ’). By contributing, participants agree to this requirement for inclusion in the follow-on manuscript.

Future AIntibody competitions

We anticipate this to be the first of a series of AIntibody competitions, each comprising a set of challenges of increasing complexity that will provide opportunities for scientists to test, compare and improve their AI/ML models. Future contests are expected to evolve in complexity depending upon the outcome of this and other competitions, in parallel with the development of AI/ML use in antibody discovery and design. In follow-on competitions anticipated below, the term “antibodies” refers to VHH’s or antibodies, each of which will be separate competitions. To maintain standardization between competitions, Challenges 2 & 3 will be retained using a different target, and NGS selection output, with different iterations proceeding as follows:

  1. Given an antibody sequence with known affinity to a target, generate candidates with improved affinities.

  2. Given an antibody sequence with known affinity to a target, with poor developability properties (e.g., thermal stability, polyreactivity), generate candidates with equal or better affinities that lack the poor developability issues.

  3. Given the sequence of a target with known structure, and structurally similar targets in the PDB, generate specific antibodies binding to the target.

  4. Given the sequence of a target with unknown structure, but similar targets in the PDB, generate specific antibodies binding to that target.

  5. Apply Challenge 3 above to a specific predefined epitope.

  6. Given a set of antibody sequences known to bind distinct epitopes on a given target, predict which antibodies bind to different epitopes.

  7. Given a set of antibody sequences known to bind distinct epitopes on a given target, predict which antibodies will allow sandwich binding.

  8. Given the sequence of a target, and a specific epitope within that target, generate de novo antibodies binding to that specific epitope.

  9. Given a set of diverse antibodies, predict their epitopes.

  10. Given the sequence of a target with known structure, generate antibodies binding to that target.

  11. Given an antibody sequence to a known human target epitope, create antibody designs that are cross-reactive to cynomolgus monkey and mouse with epitopes identical or near parental human epitope.

  12. Create a de novo or rationally guided antibody designs which do not bind any known targets (isotype controls).

To participate in AIntibody, please follow registration instructions at AIntibody.org. Participants may compete in one or more challenges. Registration will open on the publication date [TBD publication date] and close thirty days later. The datasets will be released to participants thirty days after publication and to the community upon publication of the results paper.

Acknowledgements

We thank Leonard Wossnig of LabGenius for his generous insights and contributions to the manuscript. We also thank Bio-Techne for their gracious supply of the RBD antigen for this Benchmarking Study.

Footnotes

Competing Interests

LS, FF, RD, TJP, KPS, SD, MFE, and ARMB are employees of Specifica, an IQVIA business. OH, CLL and YQ are employees of Sanofi. HGN and KR are employees of Incyte and stockholders. DB is an employee of Carterra. RW is an employee of Bonito Biosciences. CMD discloses membership of the Scientific Advisory Board of Fusion Antibodies and AI proteins. PM, EV, and RA are employees of Novo Nordisk A/S. JS is an employee of OpenEye, Cadence Molecular Sciences. RCW is an employee of AstraZeneca. AE and CS are employees of Merck Healthcare KGaA. SB is an employee of Evolutionary Scale. JR is an employee of Profluent Bio. SF is an employee of Alloy Therapeutics. VBK, SM, SK are employees of Takeda. PK is an employee of Cradle Bio. JCA is an employee of GlobalBio. EF, MS and CPG are employees of Mosaic Biosciences. MS is an employee of Eli Lilly. PMT is a member of the scientific advisory boards for Nabla Bio, Aureka Biotechnologies and Dualitas Therapeutics. SDV and FT are employees of Bayer AG. All other authors declare no competing interests.

References

  • 1.Linden A & Fenn J Understanding Gartner’s hype cycles. Strategic Analysis Report N° R-20–1971. Gartner, Inc 88, 1423 (2003). [Google Scholar]
  • 2.Lewis Chinery AMH, Mehta Brij Bhushan, Akbar Rahmad, Rawat Puneet, Slabodkin Andrei, Le Quy Khang, Lund-Johansen Fridtjof, Greiff Victor, Jeliazkov Jeliazko R., Deane Charlotte M.. Baselining the Buzz Trastuzumab-HER2 Affinity, and Beyond. bioRxiv (2024). 10.1101/2024.03.26.586756 [DOI] [Google Scholar]
  • 3.Wossnig L, Furtmann N, Buchanan A, Kumar S & Greiff V Best practices for machine learning in antibody discovery and development. Drug Discov Today 29, 104025 (2024). 10.1016/j.drudis.2024.104025 [DOI] [PubMed] [Google Scholar]
  • 4.Bender A et al. Evaluation guidelines for machine learning tools in the chemical sciences. Nat Rev Chem 6, 428–442 (2022). 10.1038/s41570-022-00391-9 [DOI] [PubMed] [Google Scholar]
  • 5.Moult J The current state of the art in protein structure prediction. Curr Opin Biotechnol 7, 422–427 (1996). 10.1016/s0958-1669(96)80118-2 [DOI] [PubMed] [Google Scholar]
  • 6.Senior AW et al. Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710 (2020). 10.1038/s41586-019-1923-7 [DOI] [PubMed] [Google Scholar]
  • 7.Senior AW et al. Protein structure prediction using multiple deep neural networks in the 13th Critical Assessment of Protein Structure Prediction (CASP13). Proteins 87, 1141–1148 (2019). 10.1002/prot.25834 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jumper J et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). 10.1038/s41586-021-03819-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Almagro JC et al. Antibody modeling assessment. Proteins 79, 3050–3066 (2011). 10.1002/prot.23130 [DOI] [PubMed] [Google Scholar]
  • 10.Almagro JC et al. Second antibody modeling assessment (AMA-II). Proteins 82, 1553–1562 (2014). 10.1002/prot.24567 [DOI] [PubMed] [Google Scholar]
  • 11.Watson JL et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023). 10.1038/s41586-023-06415-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Hie BL et al. Efficient evolution of human antibodies from general protein language models. Nat Biotechnol 42, 275–283 (2024). 10.1038/s41587-023-01763-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Harvey EP et al. An in silico method to assess antibody fragment polyreactivity. Nat Commun 13, 7554 (2022). 10.1038/s41467-022-35276-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Li L et al. Machine learning optimization of candidate antibody yields highly diverse sub-nanomolar affinity antibody libraries. Nat Commun 14, 3454 (2023). 10.1038/s41467-023-39022-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lyu J et al. AlphaFold2 structures template ligand discovery. bioRxiv (2024). 10.1101/2023.12.20.572662 [DOI] [Google Scholar]
  • 16.Ackloo S et al. CACHE (Critical Assessment of Computational Hit-finding Experiments): A public-private partnership benchmarking initiative to enable the development of computational methods for hit-finding. Nat Rev Chem 6, 287–295 (2022). 10.1038/s41570-022-00363-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Gathiaka S et al. D3R grand challenge 2015: Evaluation of protein-ligand pose and affinity predictions. J Comput Aided Mol Des 30, 651–668 (2016). 10.1007/s10822-016-9946-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Guthrie JP A blind challenge for computational solvation free energies: introduction and overview. J Phys Chem B 113, 4501–4507 (2009). 10.1021/jp806724u [DOI] [PubMed] [Google Scholar]
  • 19.Hastie KM et al. Defining variant-resistant epitopes targeted by SARS-CoV-2 antibodies: A global consortium study. Science 374, 472–478 (2021). 10.1126/science.abh2315 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Teixeira AAR et al. Simultaneous affinity maturation and developability enhancement using natural liability-free CDRs. MAbs 14, 2115200 (2022). 10.1080/19420862.2022.2115200 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ferrara F et al. A pandemic-enabled comparison of discovery platforms demonstrates a naive antibody library can match the best immune-sourced antibodies. Nat Commun 13, 462 (2022). 10.1038/s41467-021-27799-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Erasmus MF et al. Insights into next generation sequencing guided antibody selection strategies. Sci Rep 13, 18370 (2023). 10.1038/s41598-023-45538-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Raybould MIJ, Kovaltsuk A, Marks C & Deane CM CoV-AbDab: the coronavirus antibody database. Bioinformatics 37, 734–735 (2021). 10.1093/bioinformatics/btaa739 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Azevedo Reis Teixeira A et al. Drug-like antibodies with high affinity, diversity and developability directly from next-generation antibody libraries. MAbs 13, 1980942 (2021). 10.1080/19420862.2021.1980942 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES