Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2025 Jul 22;22(8):4679–4692. doi: 10.1021/acs.molpharmaceut.5c00250

A Pharmacophore-Based Method for Rapid and Accurate Virtual Screening of Antibody Libraries against Antigens

Christopher I Williams , Farbod Mahmoudinobar , David C Thompson , J Wade Davis , Sandeep Kumar ‡,*
PMCID: PMC12326342  PMID: 40694045

Abstract

Antibody-based biotherapeutics make up an important class of biopharmaceuticals. However, their discovery requires resource- and time-consuming laboratory processes. To ameliorate this situation, several computational methods were used to predict the structures of antibody:antigen complexes (Ab:Ag) and identify potential binders, in-silico. However, there is still a general lack of rapid virtual screening methods capable of screening large antibody libraries against a given antigen or group of antigens. In this work, we explore the application of a successful small-molecule drug discovery strategy and adapt pharmacophore-based virtual screening to the world of antibody discovery. Using a nonredundant data set of 874 Ab:Ag complexes, we have developed an automated method to create pharmacophores from the antibody complementarity determining regions. Our method is 98.6% (862 out of 874) successful at reproducing the ground truth, i.e., it can recapitulate the parental antibody:antigen complexes. In a benchmarking comparison with cognate docking, using 33 Ab:Ag complexes of therapeutic interest, the pharmacophore method was not only much faster than cognate docking but also recovered all the native interfacial contacts. In addition, it can also find additional putative antibody binders to a given antigen within clusters of Ab:Ag complexes with similar interfacial structures. Our method has significant implications toward accelerating biotherapeutic drug discovery as well as drug repurposing research. This method was implemented in MOE 2024 and is available to the scientific community.

Keywords: antibody, antigen, pharmacophore, virtual screening, drug design, biotherapeutics


graphic file with name mp5c00250_0007.jpg


graphic file with name mp5c00250_0005.jpg

Introduction

In 2023, the United States Food and Drug Administration approved as many biotherapeutics as small-molecule drugs for the first time. This growth has been mostly fueled by successful approvals of antibody-based biotherapeutics. This is consistent with the emergence of these biotherapeutics as the most successful class of biopharmaceuticals over the past few decades. Antibody-based biotherapeutics are particularly attractive because of high specificity, low nonmechanism toxicity, and long blood circulation times in comparison to small-molecule pharmaceuticals. In 1986, muromonab (OKT3) became the first approved antibody-based biotherapeutic. There are now over 120 approved biotherapeutics in the US market and over 200 globally (https://www.antibodysociety.org/antibody-therapeutics-product-data/). The cornerstones of the industrialization of antibody-based biotherapeutics are improvements in antibody generation methods. Currently, antibodies against a given antigen are mostly generated via animal immunization, hybridomas, B-cell sequencing, and display technologies. In the meantime, antibody sequence data have been accumulating in the public domain thanks to next generation sequencing of B-cell repertoires (https://www.ncbi.nlm.nih.gov/sra), other sources of information such as GenBank (https://www.ncbi.nlm.nih.gov/genbank/), Observed Antibody Space (OAS, https://opig.stats.ox.ac.uk/webapps/oas/), and the Adaptive Immune Receptor Repertoire community (AIRR, https://docs.airr-community.org/en/stable/getting_started.html).

With the availability of large amounts of antibody sequences together with the advent of generative artificial intelligence (AI), it has now become feasible to generate de novo full amino acid sequences of the antibody variable regions (Fv), accurately predict their three-dimensional structures, and estimate their intrinsic physicochemical properties. Recently, we have proposed the concept of DAbI (Discovery of antibodies in-silico) to enable in silico drug discovery of antibody-based biotherapeutics. The roadmap to enable DAbI consists of the following three steps: (a) in silico generation of highly diverse, developable, and human antigen-specific or antigen-agnostic antibody libraries via generative AI; (b) atomistic definition of potential antibody:antigen (Ab:Ag) complexes via virtual screening and molecular simulations; and (c) optimization of the selected Ab:Ag complexes for binding affinity and species cross-reactivity as per needs of the individual projects. In step (a), generation of antigen- or epitope-specific antibodies via machine learning promises early success in a discovery project by addressing a specific therapeutic concept. However, this approach suffers from the key disadvantage of needing to start afresh every time a new biologic drug discovery program is initiated. Conversely, the computational generation of antigen-agnostic libraries with desired developability attributes can be useful in finding antibody binders over several projects. The computationally generated antibody libraries need to be screened against a given target(s) in step (b). This can be done both computationally and experimentally. In the wet laboratory, one can construct yeast/phage display libraries from computational designs, pan them against given target(s), and then sequence the hits via next generation sequencing. In the dry laboratory, several computational methods can also be used to find potential antibody binders to a given target antigen. These methods include protein:protein docking, finding/heavy chain CDR3 loops capable of binding a specific epitope via fragment-based design, , or complementarity determining region (CDR) redesign. However, these methods are not scalable for screening large antibody libraries. DAbI envisages computationally screening large antibody libraries (at least ∼106–109 Ab molecules), comparable to the sizes of phage and yeast display libraries for it to be competitive with a typical experimental setup. Furthermore, this task needs to be performed much faster than the experiments to help shorten drug discovery project timelines. In the context of small-molecule drug discovery, pharmacophore-based virtual screening has been shown to be useful for this task.

The concept of a pharmacophore is well-known in small-molecule drug discovery. The IUPAC definition states that a pharmacophore is “an ensemble of steric and electronic features that is necessary to ensure the optimal supramolecular interactions with a specific biological target and to trigger (or block) its biological response”. A pharmacophore model (or “query”) can be made from the experimental 3D structure of a small-molecule/protein complex by placing pharmacophore feature spheres (or simply “features”) on the small-molecule atoms that make important contacts with the protein. Each feature is assigned a pharmacophore type to reflect the nature of interaction between the small molecule and the protein. In the MOE software, polar interactions are abstracted as generic H-bond donor (Don) or H-bond acceptor (Acc) features, ionic interactions as anion (Ani) or cation (Cat) features, and nonpolar contact interactions as aromatic (Aro) and/or hydrophobic (Hyd) features depending on the nature of the interacting atoms. Each pharmacophore feature is assigned a sphere radius to reflect the tolerated variation in its position. The steric environment of the binding pocket can be represented as an excluded volume object by placing VDW spheres on the protein atoms that make up the binding pocket. The excluded volume spheres indicate the regions in space that small-molecule binders cannot occupy because of VDW clashes with the protein. Pharmacophore queries can be used to search databases of 3D small-molecule structures to find novel compounds (hits) that match the pharmacophore query and, thus, could potentially bind the protein target of interest. The pharmacophore search can also position the hit molecules in 3D onto the pharmacophore query; therefore, it is an alternate approach to traditional docking for placing small molecules in the protein pocket.

Pharmacophore queries are typically used in the context of small-molecule:protein interactions, but there is no theoretical reason why they cannot be used to predict protein:protein interactions as well. For example, given an experimental antibody:antigen (Ab:Ag) complex, if we treat the Ab as the binder to the antigen, a pharmacophore query for the Ab can be created by placing pharmacophore features on the Ab atoms that make important interactions with the Ag. The surface of the antigen can be represented as an excluded volume object. The resulting query could be used to search a database of 3D antibody structures to find novel antibodies that can potentially bind the antigen. Here we present the application of pharmacophore queries and pharmacophore searching to Ab:Ag complexes. This study was undertaken to answer the following questions: (a) Is it feasible to apply pharmacophore-based virtual screening to Ab:Ag complexes? (b) How does pharmacophore-based virtual screening perform in terms of speed and accuracy when compared with the widely used method of protein:protein docking? Results presented in this report answer the first question affirmatively. In response to the second question, our results show that pharmacophore-based virtual screening is both much faster and more accurate than protein:protein docking.

Methods

Databases of Ab:Ag Complexes

A database of high-quality antibody:antigen (Ab:Ag) complexes was extracted from the Antibody.mdb project database in MOE 2022.02 by filtering the crystal structures using the following criteria. First, the crystal structure was solved to a resolution of 3 Å or better. Second, it contains at least one polypeptide antigen that buries ≥30 Å2 of surface area of the antibody’s CDRs. Third, the antibody contains both light and heavy chains with no disorder or missing portions in the CDRs or the frameworks of the antibody variable regions. When multiple copies of the Ab:Ag complex exist in the same PDB file, only the first copy is retained. The above filters result in a database of 1211 high-quality crystal structures of Ab:Ag complexes henceforth referred to as AbAg.mdb. It is important to note the original Antibody.mdb data set from MOE was created by extracting, as of Feb 2022, all structures from the PDB which contain at least one antibody protein chain as detected by the automatic IMGT Ab annotation routine in MOE. No restrictions were placed on either the class of the Ab:Ag complex or the source organism of the antibody.

Some of the structures in the AbAg.mdb database are members of redundant clusters of structures with the same antigen bound to identical or near-identical Ab units. A separate database of nonredundant complexes was generated by first assigning an antigen cluster code to each structure based on 90% sequence identity of the Ag. Structures with the same antigen cluster code contain the same antigen. Within each antigen cluster, an additional cluster code was assigned based on the combined 90% sequence identity of both the full Ab and CDR-only regions. Each unique combination of antigen and Ab-CDR cluster codes was taken to represent a unique Ab:Ag interface. A database of nonredundant Ab:Ag interfaces (AbAg_NR.mdb) was created by keeping the best resolution crystal structure from each Ab-CDR cluster within each antigen cluster.

Automatic Generation of Pharmacophore Queries from Ab:Ag Complexes

The Protein Contacts application in MOE was used to detect ionic, H-bond, arene, and distance contacts at the Ab:Ag interface given a 3D structure of the Ab:Ag complex. A Scientific Vector Language (SVL) function “protein–protein interface pharmacophore query” (“ph4_from_ppi.svl”) was written to automatically create an Ab pharmacophore query based on the contacts between the Ab CDR atoms and the antigen atoms as detected by the Protein Contacts application.

The Ab pharmacophore query is created from pharmacophore annotation points on the CDR atoms as generated by the “Unified” scheme. The Unified scheme is a rule-based system to automatically assign annotation points such as the H-bond donor, H-bond acceptor, hydrophobic, and aromatic based on the heavy atoms in a 3D structure. Two types of annotation points were used in this study:

  • (1)

    Atom-centered annotations: Points located directly on a molecule’s atom, such as “H-bond donor” or “cation”.

  • (2)

    Centroid annotations: Points located at the geometric center of a subset of atoms. For example, the aromatic ring centroid annotation point is at the centroid of the ring’s atoms.

The Ab pharmacophore query is constructed by creating pharmacophore feature spheres (or features) around annotation points of CDR atoms that make interactions with the Ag. Hydrogen-bond donor (Don) and/or acceptor (Acc) features are created from CDR annotation points involved in H-bond donor and/or acceptor contacts with the Ag. If an atom makes both donor and acceptor interactions, the Don and Acc features are combined into a single feature using the “or” operator (Don|Acc). Cation (Cat) and anion (Ani) features are created from annotation points of CDR atoms that make ionic contacts with the Ag. Coincident cation and H-bond donor features are combined into a single feature using the “or” operator (Cat|Don). Coincident anion and H-bond acceptor features are also combined into a single feature using the “or” operator (Ani|Acc). Arene Ab:Ag contacts are converted to aromatic features (Aro) only if the contacting atoms in the CDR are aromatic. The Aro feature is not created if the contacting atoms in the CDR are not aromatic. The radius of a feature sphere indicates the allowed positional deviation for the feature to be considered a match during pharmacophore searching and placement. All the atom-centered features, Don, Acc, Ani, and Cat and combinations thereof, are assigned a feature radius of 1 Å. Aro features are assigned a radius of 1.5 Å. The resulting Ab pharmacophore query can be used to search a database of 3D Ab structures for those that match the spatial arrangement of features in the query.

Volume objects can be used to filter structures already positioned by the pharmacophore query features. An “excluded volume” object specifies a region in space that atoms cannot occupy and filters any structure with any atoms occupying that region of space. The Ag atoms in an Ab:Ag complex represent a region of space that Ab atoms cannot occupy, so an excluded volume object is made by placing 1.5 Å spheres on all the Ag heavy atoms. The excluded volume is added to the Ab pharmacophore query. During a pharmacophore search, a molecule which is matched and positioned by the pharmacophore features but has atoms in the excluded volume region will be filtered out in the search.

The “occupied volume” object specifies a region in space where at least one atom must be present and will filter structures with no atoms in that region. The distance contacts at Ab:Ag interfaces often involve contiguous regions of atoms with a variety of pharmacophore atom types and lack strong directional interactions with the antigen. Therefore, distance-only contacts are not converted to pharmacophore features but instead are made into occupied volumes. Each contiguous set of atoms in a distance-only contact is converted to a separate “occupied volume” object in the Ab pharmacophore query. During the pharmacophore search, a molecule that is matched and positioned by the pharmacophore features but does not have atoms in the occupied volume objects will be filtered out in the search.

Pharmacophore Feature Survey

The “protein–protein interface pharmacophore query” application was developed and used to automatically generate an Ab pharmacophore query from each Ab:Ag complex in the nonredundant database AbAg_NR.mdb. The queries are stored in the database, along with separate data fields reporting the numbers and types of features in each query for the feature survey.

Pharmacophore Self-Placement Test

Each Ab pharmacophore query in the AbAg_NR.mdb database is subjected to a “self-placement test” where the query is used to search a random 3D orientation of the cognate Ab structure used to construct the query. A query passes the self-placement test if the search can successfully match and position the randomly oriented copy of the cognate Ab structure. When a query passes the self-placement test, the root mean square deviation (RMSD) between the native Ab pose and the pharmacophore-placed Ab pose is measured. The number of contacts preserved between the native pose and the pharmacophore-placed Ab pose is also reported.

Pharmacophore Antibody Search Performance

The performance of Ab pharmacophore queries in pharmacophore searches was tested by using each query in the AbAg_NR.mdb database of nonredundant structures to search the AbAg.mdb database of redundant structures. The following results were reported for the hits from each search:

  • 1.

    Same Ab-CDR:Ag cluster: Number of Ab structures matched by the query where the matched Ab is bound to the same antigen (same Ag cluster) and belongs to the same Ab-CDR cluster as the cognate Ab used to generate the query.

  • 2.

    Different Ab-CDR cluster: Number of Ab structures matched by the query where the matched Ab is bound to the same antigen (same Ag cluster) but belongs to a different Ab-CDR cluster than the cognate Ab used to generate the query.

  • 3.

    Different Ag: Number of Ab structures matched by the query where the matched Ab is bound to an antigen that is different from the antigen bound to the cognate Ab structure used to generate the query (different Ag cluster).

Comparison of Pharmacophore Self-Placement with Cognate Protein–Protein Docking

To compare pharmacophore placement of Abs with placement using a traditional protein–protein docking approach, a subset of 33 therapeutically relevant Ab:Ag complexes was taken from the Antibody.mdb database in MOE 2024 and saved into a separate database Ab33_therapeutic.mdb. Six of the 33 structures were already present in the original AbAg.mdb database of 1211 redundant structures. The remaining 27 structures were not present because 10 of the structures were published after Feb 2022 and the remaining 17 were filtered out during creation of the AbAg.mdb database because of either poor resolution or missing residues or both. The “protein–protein interface pharmacophore query” application was used to automatically generate a pharmacophore query for each of the 33 Ab:Ag complexes and each query was used to search and place the native Ab structure. The pharmacophore results were compared with the results produced using the MOE protein–protein docker to self-dock each of the 33 Abs to its native Ag. To ensure a direct comparison between docking and the pharmacophore placement method, the docking search space was restricted to the CDR regions on the Ab and the known epitope region on the antigen. The rigid body refinement option was used in the docking. The docking runs used the default values of 10,000 poses in preplacement, 1000 poses in placement, and 100 poses in refinement. The docking pose in best agreement with the native pose (lowest RMSD) was used for comparison to the pharmacophore self-placement method.

To explore the protein:protein docking results from a contacts perspective, the best RMSD pose was imported into the MOE and annotated using the IMGT scheme. Contacts were determined between Ag and the CDR, and then a fractional count compared to the native pose was determined. Unlike the pharmacophore self-placement analysis, only covalent, arene, metal, hydrogen bond, and ionic contacts were analyzed (distance was not evaluated).

Consensus Pharmacophore of the KcsA Antigen

Consensus pharmacophores were created for a collection of Ab:Ag structures containing the same antigen by first superposing the structures in 3D on their common antigen unit using sequence and structure alignment tools in the MOE. A pharmacophore query was automatically generated for the Ab:Ag interface of each structure. The Database AutoPH4 (db_AutoPH4) application (“Database AutoPH4, SVL source code provided by Chemical Computing Group ULC, 1010 Sherbrooke St. West, Suite #910, Montreal, QC, Canada, H3A 2R7” 2024) was used to create consensus pharmacophore features by clustering the pharmacophore features from all structures based on type and distance. Features with the same pharmacophore type but from different structures in the collection were clustered by using distance-based single-linkage clustering with a cutoff of 1 Å. Each cluster of features was converted to a consensus feature by placing a sphere at the cluster center with a radius that encompasses all of the spheres in the cluster. Each consensus feature was assigned a frequency value between [0,1] that measures the fraction of complexes in the collection that contain that feature in their query. Consensus features with a frequency of 1.0 represent pharmacophore features that occur in all structures in the collection. The consensus features were each assigned a unique integer label, and every structure in the collection was assigned a list of feature labels, indicating the consensus features present in each structure.

Results

Database of Ab:Ag Complexes

A database of 1211 high-quality crystal structures of Ab:Ag complexes available in the Protein Data Bank (PDB) and satisfying the criteria summarized in the Methods section was constructed in March 2024. Henceforth, this database is referred to as AbAg.mdb. A nonredundant subset of 874 Ab:Ag complex structures based on an antigen and antibody clustering as described in the Methods section was extracted from this database and saved to a database henceforth referred to as AbAg_NR.mdb. The 874 structures in AbAg_NR.mdb were taken to represent examples of unique Ab:Ag interfaces as described below.

The results of the Ag and Ab clustering are reported in Table . Each Ag cluster is a collection of database entries with the same antigen. There are 488 Ag clusters in the database representing 488 unique antigens within the data set of 1211 Ab:Ag complexes. Many Ag clusters consist of only one structure, indicating there is only one example of that antigen in the database. Some Ag clusters consist of 10 or more structures indicating multiple examples of the antigen structure existing in the database. The “Ag cluster size” in Table is the number of Ab:Ag complexes in the Ag cluster, and the “# clusters” column reports how many Ag clusters have a given Ag cluster size. There is a wide variation in Ag cluster size. The database contains 316 singleton Ag clusters, (Ag cluster size = 1) indicating only one representative structure of that antigen existing in the database. There are 172 nonsingleton clusters with 2 or more representative antigen structures in the database. Seventy-nine (79) Ag clusters have only two representative Ab:Ag structures in the database (Ag cluster size = 2) and 83 Ag clusters have between 3 and 11 representative structures (Ag cluster sizes = [3,11]). The antigen identity in the ten clusters with 12 or more representative Ab:Ag structures (Ag cluster size ≥12) is listed in Table . The largest Ag cluster in our database (Ag cluster size = 104) is the spike protein SARS-COV2 antigen.

1. Different Ag Clusters in Our Data Set .

Ag cluster size # clusters antigen identity
104 1 SARS-CoV2 spike protein
56 1 KcsA ion channel
51 1 lysozyme
25 1 HIV-1 gp120 core
24 1 HIV-1 V3 LOOP MIMIC
23 1 alpha-Iib-beta3
19 1 hepatitis C virus envelope glycoprotein E2
17 1 PfCSP peptide 20
15 1 PrP (114–234)
12 1 HIV gp41 11mer
11 2 (multiple antigens)
10 5 (multiple)
9 3 (multiple)
8 3 (multiple)
7 4 (multiple)
6 2 (multiple)
5 10 (multiple)
4 16 (multiple)
3 38 (multiple)
2 79 (multiple)
1 316 (multiple)
a

An Ag cluster is a set of database entries with the same antigen identity. The “Ag cluster size” is the number of structures in an Ag cluster. The “# clusters” column reports how many Ag clusters have a given Ag cluster size. The sum of the “# clusters” column488stands for the number of unique antigen clusters in the database. In summary, the 1211 Ab:Ag complex crystal structures contain 488 unique antigens.

Antigen structures in each of the nonsingleton Ag clusters (Ag cluster size ≥2) may be bound either to the same or different Abs. To generate a set of unique Ab:Ag interfaces for the pharmacophore survey, each of the nonsingleton Ag clusters was further clustered into Ab-CDR clusters based on 90% identity of both the full Ab sequence and the CDR-only sequence. Each Ab-CDR cluster within a given Ag cluster represents a set of highly sequence-similar antibodies bound to the same Ag. Many of the nonsingleton Ag clusters subdivided into different Ab-CDR clusters. The antigen identity and number of unique Ab-CDR clusters bound to the antigen are reported in Table for Ag clusters with 12 or more members. Some Ag clusters contain several Ab-CDR clusters, suggesting the Abs binding the Ag are a diverse set which may potentially recognize different epitopes on the antigen. For example, the 104-member SARS-CoV2 spike protein Ag cluster divides into 32 Ab clusters based on Ab-CDR sequence identity, suggesting the Abs bound to this antigen are a diverse set and potentially bind different epitopes. In contrast, other Ag clusters such as the 56-member KcsA ion channel Ag cluster contain only a few Ab-CDR clusters, suggesting the bound Abs are a set of similar antibodies that possibly bind the same epitope.

2. Number of Ab-CDR Clusters in the 10 Largest Antigen Clusters .

Ag cluster size #Ab-CDR clusters antigen description
104 32 SARS-CoV2 spike protein
56 3 KcsA ion channel
51 12 lysozyme
25 21 HIV-1 gp120 core
24 23 HIV-1 V3 LOOP MIMIC
23 3 alpha-Iib-beta3
19 15 hepatitis C virus envelope glycoprotein E2
17 5 PfCSP peptide 20
15 7 PrP (114–234)
12 7 HIV gp41 11mer
a

The “Ag cluster size” is the number of structures in the database with a given antigen. For each collection of structures with the same antigen, “#Ab-CDR clusters” reports how many clusters are produced when the Abs are clustered based on 90% sequence identity in the full Ab and CDR-only regions. Some collections of antigens such as SARS-CoV2 spike protein have many Ab-CDR clusters, suggesting the Abs binding the SARS-CoV2 spike protein are quite diverse. In contrast, antigens such as the KcsA ion channel have only a few Ab-CDR clusters, suggesting there is little diversity in Abs that bind this antigen.

The clustering of the nonsingleton Ag clusters based on the Ab-CDR sequence identity produced 558 Ag:Ab-CDR clusters. The best resolution structure from each of the 588 clusters was combined with the 316 singleton Ab:Ag structures to produce the database of 874 nonredundant Ab:Ag structures (AbAg_NR.mdb) used in the pharmacophore feature survey and pharmacophore search tests.

Pharmacophore Feature Survey: Features Present in Ab:Ag Interfaces

The 874 structures in AbAg_NR.mdb were taken as a representative sampling of Ab:Ag interfaces. This data set was used for the pharmacophore query generation and feature survey study. A pharmacophore query was created for each entry of this database using the antibody paratope residues and known antigen contacts as described in the Methods section. The plots in Figure report the incidence of different pharmacophore feature types at the Ab:Ag interfaces in our data set. The x-axis is the pharmacophore feature count, and the y-axis is the fraction of the database entries that have queries with that feature count. The average total number of features observed at Ab:Ag interfaces is 10.2 ± 3.8, which is greater than the number of features typically observed for the small-molecule ligands. Acceptor (Acc) and donor (Don) features are the most common feature types found at the interface. Two structures6MWN and 6XJWproduced pharmacophore queries with the maximum observed combined Acc/Don count of 21. The anion, cation, and Aro (aromatic) pharmacophore features were observed with much lower frequencies at Ab:Ag interfaces (Figure ). The median value for the total combined number of pharmacophore features (Acc/Don/Cat/Ani/Aro) at the Ab:Ag interface in this data set is 10. Only 8 Ab:Ag interfaces exhibited 20 or more features, with a maximum of 26 features observed at one Ab:Ag interface (6XJW).

1.

1

A survey of different types of Ab:Ag query features is presented. In each plot, the x-axis plots the number of features (feature count) of a given type in a query. The y-axis plots the fraction of Ab:Ag_NR database queries with that feature count. The total number of pharmacophore features found at Ab:Ag interfaces is small enough to facilitate the application of pharmacophore-based virtual screening to the realm of antibody-based biotherapeutic drug discovery.

In addition to pharmacophore features, the Ab:Ag pharmacophore queries contain occupied volume objects representing contiguous groups of heavy atoms forming distance contacts with the antigen. The plot in Figure shows the fraction of database entries as a function of the occupied volume object. The average number of occupied volume objects is (21.2 ± 0.15), and the median value is 21. The maximum number of occupied volume objects observed at an Ab:Ag interface is 36.

One Ab:Ag complex structure in the database (PDB entry 6MG7) produced a pharmacophore with 0 features because the Protein Contacts application detected only distance contacts at the Ab:Ag interface and no acceptor, donor, or aromatic interactions. Four Ab:Ag complexes (PDB entries 4Q0X, 2Y07, 4BH8, and 5YY5) generated queries with only one feature, and another five complexes (PDB entries 4S1Q, 2XZQ, 4HLZ, 6Z3Q, and 7X08) generated queries with only two features. Since a pharmacophore query needs at least 3 features to perform a 3D search, only the 864 database entries with Ab queries containing ≥3 features were used in the pharmacophore searches to validate our methodology by reproducing the ground truth via pharmacophore self-placement (described in the next section).

Overall, this survey of pharmacophoric features suggests that most pharmacophore queries generated at Ab:Ag interfaces will contain fewer than 25 features. The computational feasibility of a pharmacophore search is related to the number of features defining the search: this number is small enough to be considered practical for pharmacophore searching. If the average number of pharmacophore features at Ab:Ag interfaces was significantly largere.g. more than 100then this approach becomes less useful and certainly less competitive with experimental strategies as explored in the Introduction.

Validation of the Pharmacophore Method for Virtual Screening of Antibody Databases

Our method for pharmacophore-based virtual screening of antibody databases was validated by asking the following questions:

  • (1)

    Pharmacophore self-placement: Can a pharmacophore search using an Ab:Ag derived query reproduce the ground truth by finding and placing the cognate Ab:Ag partners in the same orientations as observed in the crystal structures?

  • (2)

    Pharmacophore antibody searches: Can a pharmacophore search find closely related antibodies within the same Ab:Ag interface cluster?

  • (3)

    Pharmacophore searches versus protein:protein docking: How does the pharmacophore search method compare with docking in terms of speed and accuracy?

Pharmacophore Self-Placement

The 874 nonredundant Ab:Ag complexes in the AbAg_NR.mdb were used to answer the first question on reproducing the ground truth and the results are summarized in Table . Ten out of the 874 Ab:Ag complexes were found to have too few (<3) pharmacophore features. These were excluded from further calculations. The remaining 864 Ab:Ag complexes were used to generate the 864 searchable pharmacophore queries (i.e., queries with 3 or more features). These queries were subjected to a self-placement test as described in the Methods section. The self-placement test is passed if and only if the Ab:Ag complex predicted by the pharmacophore query is placed within 1.0 Å RMSD of the Ab structure in the original Ab:Ag complex that was used to generate the query. Each self-search test was performed starting from a random orientation of the original Ab structure.

3. Results of Pharmacophore Self-Placement Tests with Automatically Generated Pharmacophore Queries.

description value
#Ab:Ag complexes used to make queries 874
#Ab queries with 0 features 1
#Ab queries with 1 feature 4
#Ab queries with 2 features 5
#Ab queries that are searchable (>2 features) 864
#Ab queries pass self-placement 862
average self-placement Ab RMSD (Å) 0.00133548
min self-placement Ab RMSD (Å) 0.000007
max self-placement Ab RMSD (Å) 0.1165
average % native Ab:Ag contacts preserved 98.90
maximum run time (s) 1.49
minimum run time (s) 0.57
average run time (s) 0.71
median run time (s) 0.7
a

RMSD stands for root mean square deviation between the pose placed by the pharmacophore and the original one in the crystal structure of an Ab:Ag complex.

In theory, every searchable pharmacophore query should be able to find the cognate Ab structure used to generate the query. In practice, the results in Table show that 862 (99.8%) of the 864 automatically generated searchable queries passed the self-placement test by successfully placing a random orientation of the Ab structure used to generate the query. The average RMSD between the pharmacophore-placed pose and the native pose is reported in Table along with the average fraction of Ab:Ag native contacts (99%) preserved between the two poses. The average Cα RMSD (0.0013 Å) and maximum Cα RMSD (0.1165 Å) values suggest that the pharmacophore-placed Ab structures are positioned remarkably close to the original Ab structure. The self-placement search times reported in Table show that the pharmacophore self-placement is fast, with a median search time of 0.70 s. The slowest self-placement search took 1.49 s. In our experience, these search times are much faster than those required by the cognate docking.

The two automatically generated pharmacophore queries that failed the self-placement test correspond to the PDB entries 6SV2 and 6UDJ. Visual inspection of the failed queries showed the initial failure was due to filtering by the excluded volume because the excluded volumes in these had some overlap with the pharmacophore features. Once the sphere radii of the excluded volumes were manually reduced from 1.5 to 1 Å, both failed queries passed the self-placement test. These manually adjusted queries were used moving forward for these two structures.

In conclusion, the self-placement results demonstrate that pharmacophore searches using queries derived from Ab:Ag complexes can indeed reproduce the ground truth by finding and placing the cognate Ab:Ag partners in the same 3D orientation as observed in the crystal structures.

Pharmacophore Antibody Searches

The pharmacophore self-placement tests described above demonstrate that pharmacophore searches using Ab:Ag-derived queries can reproduce the ground truth and produce high-quality cognate Ab:Ag poses. In our second question, we asked whether pharmacophore searching can be used as a screening tool to identify antibody structures different from the cognate Ab structure used to generate the query. To investigate this, each of the 864 searchable queries from the nonredundant Ab:Ag complexes (AbAg_NR.mdb) was used to screen against the original database of 1211 redundant Ab:Ag complexes (AbAg.mdb).

As described in Tables and , the AbAg_NR.mdb database was used to generate the unique Ab:Ag interface clusters. The search results for the 25 largest Ab:Ag clusters are reported in Table . The identity of the antigen is given for each cluster. The “cognate Ab:Ag PDB” reports the PDB code of the Ab:Ag complex used to generate the pharmacophore query for that cluster. The “Ab-CDR:Ag cluster size” column reports how many unique Ab:Ag complexes exist in each cluster. The “total hits” column reports the total number of hits retrieved by the query. The “same Ab-CDR:Ag cluster” column reports the number of hits where the Ab is from the same Ab-CDR:Ag cluster as the cognate Ab:Ag used to make the query. This number reflects the ability of a query to find Ab structures that are similar to the Ab used to generate the query. The “different Ab-CDR:Ag cluster” column in Table reports the number of hits where the Ab is from a different Ab-CDR:Ag cluster than the Ab used to generate the query. This number reflects the ability of a query to find more diverse Ab structures (≤90% sequence similarity) that can still potentially bind to the same antigen. Finally, the “different Ag” column in Table reports the number of Ab hits where the Ab binds an antigen that is different from the antigen in the Ab:Ag complex used to generate the query.

4. Pharmacophore Antibody Searches Using Automatically Generated Pharmacophore Queries .

      number of hits (#)
antigen cognate Ab:Ag PDB code Ab-CDR: Ag cluster size total same Ab-CDR: Ag cluster different Ab-CDR: Ag cluster different Ag
KcsA ion channel 2IH3 54 19 18 1 0
LY-CoV488 neutralizing antibody against SARS-CoV-2 7KMH 52 1 1 0 0
hen egg lysozyme 3D9A 30 10 10 0 0
integrin alphaIIbbeta3 receptor 3T3P 21 15 14 1 0
HIV-1 BG505 SOSIP.664 6MTJ 9 8 8 0 0
PfCSP peptide 21 7RD4 9 3 3 0 0
SARS-CoV-2 beta variant spike glycoprotein 7Q0G 9 2 2 0 0
VEGF 7KEZ 8 1 1 0 0
hen egg lysozyme 1G7J 8 1 1 0 0
Gi-coupled MRGPRX2 7S8M 7 1 1 0 0
mouse PrPc fragment 120–230 4H88 7 6 5 1 0
HIV-1 gp41 peptide LLELDKWASLW 3D0V 6 3 3 0 0
histone chaperone ASF1 5UEA 6 3 3 0 0
human TRAAK K+ channel FHIEG 7LJ5 6 6 5 1 0
circumsporozoite protein NANP3 6O24 5 3 3 0 0
clade A/E 93TH057 HIV-1 gp120 core 4OLZ 5 2 2 0 0
SARS-CoV-2 receptor binding domain 7LOP 5 1 1 0 0
Der p 1 3RVW 4 5 4 1 0
short form HGFA 2R0L 4 1 1 0 0
CLC-ec1 6ADB 4 5 4 1 0
medaka fish taste receptor T1r2a-T1r3 ligand binding domains 5X2N 4 3 3 0 0
4E10 (delta loop) with epitope bound 5CIN 4 3 3 0 0
influenza hemagglutinin head 6Q0L 4 1 1 0 0
circumsporozoite protein DND3 6VLN 4 1 1 0 0
hepatitis C virus envelope glycoprotein E2 ectodomain 6MEH 4 1 1 0 0
a

The antigen column lists the identity of the antigen in the cluster. The Ab-CDR: Ag cluster size reports how many Ab:Ag complexes are in each Ab_CDR: Ag cluster. The “cognate Ab:Ag PDB code” reports the PDB code of the Ab:Ag complex used to generate the pharmacophore query for the cluster. The “total hits” column reports the total number of hits found when the pharmacophore query is used to search the database of nonredundant Ab:Ag complexes AbAg.mdb. The “same Ab-CDR:Ag cluster” column reports the number of hits that belong to the same Ab-CDR:Ag cluster as the Ab:Ag complex used to make the query. The “different Ab-CDR:Ag” cluster column reports the number of hits that bind the same antigen but belong to a different Ab-CDR cluster than the Ab:Ag complex used to make the query. The “different Ag” column reports the number of hits where the matched Ab is bound to an antigen different from the antigen in the Ab:Ag complex used to generate the query. None of these searches found an Ab bound to antigen different from the cognate antigen.

Table shows that our PH4 queries can find one or more Abs that are similar but not necessarily identical to the cognate Ab in all 25 Ab:Ag clusters. However, only 6 queries could find more dissimilar Ab structures from a different Ab-CDR-Ag cluster. This observation suggests the automatically generated pharmacophores are quite specific to the interactions observed within the Ab-CDR:Ag cluster used to generate the query. The results in the “different Ag” column in Table show that none of the searches found an Ab that is bound to an antigen different from the cognate antigen in the Ab:Ag complex used to make the query. This result is not surprising because the automatically generated pharmacophores have specific geometrical features and because antibodies with similar sequences (same clonotypes) bind the same antigens/epitopes.

Ab Pharmacophores Query Search Time

The scatter plot in Figure a shows the time it took each of the 864 searchable queries to search the 1211 entries in the AbAg.mdb database versus the number of features in the query. The searches were run in parallel on an IBM laptop using 4-CPUs to perform the search. Only 15 of the pharmacophore queries required more than 50 s to search the 1211 Ab structures. The scatter plot in Figure b is the same as Figure a except showing only the searches t < 50 s. Both plots show no direct correlation between search time and the number of features in the query, but some trends can be noted. The four queries with longest search times (>500 s) all have >20 features, which is somewhat expected since increasing the number of features in a query can slow down the search because more features need to be matched. A query with a small number of features can also be slow because the small number of features can match a larger structure in multiple different ways, so the search slows down as the multiple matches are investigated.

2.

2

A scatter plot of the time (s) to search the 1211 Ab structure database vs the number of features in the query. (a) Results for all 864 searchable queries. (b) Results for the 849 queries with search times <50 s. This result shows that pharmacophore-based virtual screening is fast.

The average time for a query to search the 1211 entry AbAg.mdb database is 10.2 ± 39.1 s, with a median search time of 6.8 s. The average and standard deviation are skewed by 15 searches that took longer than 50 s. When the 15 slower searches are removed from the statistics, the average search time becomes 8.5 ± 6.3 s, with a median of 6.6 s. The times are reported for a search that used 4 processors, so we can calculate one processor will, on average, search 1211 structures in 40.8 s or ∼30 structures per seconds. The average search time suggests that high-throughput screening of Ab is possible because pharmacophore searching is a massively parallel process and scales almost linearly with the number of processors used. A pharmacophore search using a 100-processor cluster could potentially search 3000 structures per second or 10.8 million structures in 1 h.

Pharmacophore Searches versus Protein:Protein Docking

Protein:protein docking is an established method to predict structures of the putative Ab:Ag complexes. Therefore, the third question explored in this work is: In the context of biologic systems, how does a pharmacophore search method compare with traditional protein:protein docking in terms of both speed and accuracy? To answer this question, we selected a control data set of 33 unique Ab:Ag complexes from the MOE 2024 Antibody.mdb database such that the antibodies in these complexes are biotherapeutics. We restricted ourselves to pharmacophore self-placement because it is directly comparable with cognate docking, which consists of supplying the bioactive conformation of the ligand (Ab in our case) as an input. Table shows the results of the 33 pharmacophore self-placement experiments and cognate protein:protein docking. The protein:protein docking was performed using MOE 2024.0601 with parameters defined in the Methods section. The calculations were run on a 2021 Apple M1 MacBook Pro.

5. Comparison of Pharmacophore Placement and Localized Protein:Protein Docking .

      pharmacophore self-placement
cognate docking
generic name of the therapeutic antibody PDB code PH4 features (#) RMSD to native (Å) native contacts preserved (%) search time (s) RMSD to native (Å) native contacts preserved (%) search time (s)
latikafusp 1W72 14 2.41 × 10–5 100 46 × 10–3 0.63 11 1.18 × 103
danburstotug 2DD8 9 2.81 × 10–5 100 83 × 10–3 0.57 62 1.08 × 103
eciskafusp 3ZTJ 5 2.21 × 10–5 100 96 × 10–3 0.68 100 1.93 × 103
becotatug 4KRO 13 2.02 × 10–5 100 60 × 10–3 0.04 67 1.45 × 103
dargistotug 4XHJ 6 2.63 × 10–5 100 105 × 10–3 0.63 79 1.67 × 103
tobevibart 5J13 10 1.80 × 10–5 100 167 × 10–3 1.06 55 0.77 × 103
visugromab 5KEL 5 2.59 × 10–5 100 74 × 10–3 0.59 90 1.03 × 103
fepixnebart 5KN5 5 2.52 × 10–5 100 220 × 10–3 0.25 92 0.63 × 103
riltovetbart 5TLJ 6 2.54 × 10–5 100 390 × 10–3 0.31 67 0.70 × 103
prafnosbart 5TRU 8 2.63 × 10–5 100 97 × 10–3 5.04 0 0.85 × 103
eglatoprutug 5TZT 3 2.53 × 10–5 100 172 × 10–3 0.40 75 0.79 × 103
mevonlerbart 5VYF 14 2.71 × 10–5 100 82 × 10–3 0.71 60 0.84 × 103
varokibart 6OAN 15 2.54 × 10–5 100 143 × 10–3 0.37 71 1.28 × 103
cirevetmab 6IAP 9 3.60 × 10–2 100 98 × 10–3 0.07 91 0.86 × 103
freneslerbart 6UTA 9 2.55 × 10–5 100 82 × 10–3 0.24 88 0.67 × 103
burfiralimab 6XLQ 6 2.51 × 10–5 100 69 × 10–3 0.32 56 0.88 × 103
resugosbart 7BBJ 13 2.61 × 10–5 100 121 × 10–3 0.41 46 1.52 × 103
vilamakitug 7F9W 8 2.32 × 10–5 100 106 × 10–3 0.46 93 0.9 × 103
bebtelovimab 7MMO 9 2.57 × 10–5 100 84 × 10–3 0.35 44 0.88 × 103
camoteskimab 7MSQ 13 2.35 × 10–5 100 68 × 10–3 0.13 77 1.06 × 103
gorivitug 7PA9 13 2.57 × 10–5 100 118 × 10–3 0.37 67 1.43 × 103
acrixolimab 7PNM 6 2.41 × 10–5 100 156 × 10–3 0.64 53 2.98 × 103
anzurstobart 7ST5 13 2.66 × 10–5 100 83 × 10–3 0.84 35 0.8 × 103
izenivetmab 8AS0 11 2.67 × 10–5 100 119 × 10–3 0.14 43 0.71 × 103
bempikibart 8C7M 15 2.49 × 10–5 100 70 × 10–3 0.29 78 1.64 × 103
paridiprubart 8CN9 6 2.60 × 10–5 100 110 × 10–3 0.18 60 1.3 × 103
lunaxafusp 8DCM 14 2.61 × 10–5 100 106 × 10–3 0.09 83 1.63 × 103
boserolimab 8DS5 10 2.72 × 10–5 100 103 × 10–3 0.34 58 0.65 × 103
fiztasovimab 8EZ8 8 2.48 × 10–5 100 76 × 10–3 0.03 91 1.13 × 103
anivovetmab 8IVX 6 2.57 × 10–5 100 82 × 10–3 0.0013 93 1.4 × 103
cifurtilimab 8JRU 6 3.93 × 10–5 100 234 × 10–3 0.27 56 0.79 × 103
ralzapastotug 8SZY 8 1.82 × 10–5 100 177 × 10–3 0.16 42 0.8 × 103
perenostobart 8TUI 9 2.54 × 10–5 100 113 × 10–3 0.52 60 0.83 × 103
average ± std   9.2 ± 3.4 (3.0 ± 0.34) × 10–5 100 (118 ± 66) × 10–3 0.34 ± 0.23 65.4 ± 24.0 (1.12 ± 0.4) × 103
a

Comparison of Ab self-placement for 33 therapeutically relevant Ab:Ag complex crystal structures using pharmacophore searching and cognate docking. Self-placement with the pharmacophore is much faster than docking and produces poses closer to the native structure as measured by RMSD and % of preserved native contacts.

b

Native contacts preserved refers to the fraction of the interfacial contacts common between the original Ab:Ag complex structure and structure predicted by PH4 self-placement searches or localized docking. For the assessment of the contacts of the protein:protein docking poses, only contacts of type arene, metal, ionic, hydrogen bond, and covalent were used in the analysis.

c

In these five instances the pose reported is not the top scoring pose (3ZTJpose 2, 5ITJpose 2, 5TRUpose 10, 5TZTpose 3, 8TUIpose 2).

d

Average value taken over the Cα RMSD values for 32 Ab:Ag complexes after excluding the outlier for pharmacophore placement, 6IAP, whose Cα RMSD values is 0.036 Å.

The interfaces in the 33 unique therapeutic Ab:Ag complexes contain on average nine (average = 9.2 ± 3.4) pharmacophore features and the pharmacophore self-placement test reproduced the ground truth (i.e., recapitulate the known Ab:Ag complex) in all the 33 cases with an average Cα RMSD of 3 × 10–5 Å, preserved 100% native contacts, and took an average of 118 ms per pharmacophore search (Table ). The cognate docking protocol was also able to identify the best RMSD pose as the top scoring pose in 28 of the 33 cases (85%). The median Cα RMSD of the top scoring pose was found to be 0.35 Å, while the median Cα RMSD of the best pose across the set of 33 systems was found to be 0.34 Å. The cognate docking application preserved two-thirds of the native contacts (average = 65.4 ± 24%) and took an average of 1120 s (∼19 min) per experiment (Table ).

While the Ab poses generated using pharmacophore self-placement have a consistently smaller RMSD value, given that in both approaches these values are so small as to be well within any kind of experimental (i.e., observational) tolerance, we can conclude that both methods seem able to reproduce the ground truth. However, the pharmacophore self-placement results in poses with a greater number of native contacts preserved than Ab poses generated from cognate docking. Overall, the results summarized in Table demonstrate that pharmacophore searching is significantly faster and more accurate than the docking runs and, therefore, more suitable for the rapid virtual screening of large antibody libraries.

Figure presents visual comparisons between PH4 self-placements and cognate docking poses for six Ab:Ag complexes that were chosen from Table . The PH4 self-placements and the cognate docking poses have been superposed on to the ground truth (original crystal structure) in each case. The complexes were chosen such that their cognate docking predictions are highly accurate, as reflected by their Cα RMSD values ranging from 0.0013 to 0.84 Å. The PDB entries for these six Ab:Ag complexes are 8IVX (cognate docking Cα RMSD = 0.0013 Å), 8SZY (cognate docking Cα RMSD = 0.16 Å), 6XLQ (cognate docking Cα RMSD = 0.32 Å), 7F9W (cognate docking Cα RMSD = 0.46 Å), 7PNM (cognate docking Cα RMSD = 0.64 Å), and 7ST5 (cognate docking Cα RMSD = 0.84 Å). The Cα RMSD values for PH4 self-placements for these structures range from 1.82 × 10–5 Å (8SZY) to 2.66 × 10–5 Å (7ST5). In Figure a (PDB entry 8IVX), it is not feasible to visually distinguish between the ground truth (crystal structure) and PH4 self-placement (Cα RMSD = 2.57 × 10–5 Å) as well as the ground truth and the cognate docking pose (Cα RMSD = 0.0013 Å) because of extremely low Cα RMSD values in both instances. For the remaining 5 Ab:Ag complexes (Figure b–f), one cannot distinguish between the ground truth and the PH4 self-placements, but the cognate docking poses (shown as blue ribbons) can be visually distinguished, despite their high accuracy (Table ). In summary, Figure provides visual evidence that PH4 self-placements are more accurate compared with the cognate docking poses, even among the most accurate cognate docking examples.

3.

3

Superposition of PH4 self-placement and cognate docking poses on the crystal structures for six Ab:Ag complexes contained in the PDB entries (a) 8IVX, (b) 8SZY, (c) 6XLQ, (d) 7F9W, (e) 7PNM, and (f) 7ST5. In each panel, the Ag is represented as molecular surface (dark yellow) and the Ab is represented as a ribbon. The antibody ribbons have been annotated using the CCG scheme, where heavy chain CDR3 (HCDR3) is shown in red. The remaining two heavy chain CDRs (HCDR1 and HCDR2) are shown in brown, while all of the light chain CDRs are shown in magenta. The heavy chain and light chain framework regions are shown in dark and light green colors. The antibody structures in the crystal structure (ground truth), PH4 self-placement, and cognate docking are visually indistinguishable in the case of the Ab:Ag complex is shown in (a). The cognate docking poses are shown as blue ribbons in the (b–f) to highlight their differences with respect to the ground truth and PH4 placement. This figure shows that PH4 self-placements are more accurate than cognate docking.

Consensus Pharmacophore Model for the KcsA Antigen

The results in Table show that a pharmacophore query created using one representative example from an Ab-CDR:Ag cluster of structures can often find antibodies in the same Ab-CDR:Ag cluster and in some instances also find antibodies from a different Ab-CDR:Ag cluster. For example, the 56 antibody:KcsA ion channel complexes in AbAg.mdb divide into three clusters, one cluster with 54 members and two singleton clusters each with one member, based on 90% sequence identity in the Ab and CDR regions. A pharmacophore search using a query generated from the highest resolution structure in the 54-member cluster (PDB code 2IH3) finds 18 hits from the same Ab-CDR:KcsA cluster as 2IH3 and one hit from another Ab-CDR:KcsA cluster different from the 2IH3 cluster. The query does not find all of the KcsA binding antibodies presumably because not all of the antibodies that bind KcsA exhibit the same contacts as 2IH3. However, the results suggest KcsA binding antibodies share some common binding features, and if so, a consensus pharmacophore for KscA binding could be created from the highly conserved features. In theory, a pharmacophore search with a consensus query could find most if not all of the KcsA binding antibodies in a database. However, if the consensus query contains too few features, it can become nonspecific, and pharmacophore searches with the query may return a large portion of any Ab database being searched, rendering the searches less useful. To test if consensus pharmacophore queries can be made from a set of related Ab:Ag complexes and if consensus pharmacophore queries can be used in meaningful searches, we asked the following questions:

  • (1)

    Consensus pharmacophores detection: Can we find a “consensus pharmacophore query”i.e., a set of features common to all structuresfrom a collection of antibodies that bind the same antigen epitope?

  • (2)

    Consensus pharmacophore searching: Does searching a database with a consensus pharmacophore query find all (or most) of the Abs that bind the antigen used to make the query? Is the consensus pharmacophore query so nonspecific that it hits almost any antibody structure in each database, making it useless for screening?

To answer question 1, consensus pharmacophore analysis as described in the Methods section was performed using all 56 Ab:KcsA complexes from the AbAg.mdb database. The complexes were aligned on the common KcsA unit, as shown in Figure . The antigen units are shown in red and the bound Ab units are shown in blue. The consensus pharmacophore analysis detected six (6) features common to all structures, 2 acceptor features (Acc), and four (4) combined acceptor and anion (Acc & Ani) features. The Ab residue(s) responsible for each consensus feature and the corresponding interacting residues on KcsA are given in Table . The consensus features are shown in Figure as cyan-colored spheres. A consensus pharmacophore query was created using the six consensus features in Table . An excluded volume object was created from the antigen heavy atoms in the crystal structure contained in the PDB entry 2IH3. No occupied volumes representing distance contacts were added. The KcsA consensus query is quite simple, with only 6 features and one excluded volume.

4.

4

Crystal structures of 56 KcsA:antibody complexes overlaid using sequence-based alignment and the superposition of the antigen. The antigen backbone ribbons are red, and the Ab ribbons are blue. The consensus pharmacophore features common to all 56 structures are shown as cyan spheres. Features F1 and F2 are acceptor (Acc), and features F3–F6 are combined acceptor and anion (Acc & Ani).

6. KscA Consensus Pharmacophore Features .

pharmacophore feature type Ab residue Ag residues
acceptor N108 Q58
acceptor G109 Y62, T61
acceptor & anion E107 R52
acceptor & anion D38 R64
a

The table lists the consensus pharmacophore feature types along with the residue on the Ab that generates the feature and the residue(s) on the KcsA antigen that the feature interacts with.

To answer question 2, the KcsA consensus query was used to perform pharmacophore searches on the AbAg.mdb database and the Antibody.mdb database from MOE 2024 to test (a) if the consensus query can find all known antibodies that bind the KcsA antigen, (b) if the consensus query can find some antibodies either unbound or bound to different antigens from KcsA, and (c) if the consensus query is so nonspecific that it hits a large portion of antibody structures in the database. The results in Table report the total number of KcsA structures in each search database, the total number of hits retrieved from searches using the KcsA consensus pharmacophore query, the number of hits where the Ab is bound to KcsA, and the number of hits where the Ab is not bound to KcsA. The PDB codes for the hits where the Ab is not bound to KcsA are also reported. The results in Table show the KcsA consensus pharmacophore query finds all 56 KcsA structures in the AbAg.mdb database and finds zero (0) non-KcsA bound hits. This represents a perfect retrieval rate, which is not surprising given the database being searched is small and the KcsA consensus pharmacophore was made from all the structures in this database. It should be noted that even with a small number of features, the consensus Ab is sufficiently specific that it does not produce a long list of Abs not known to bind KcsA.

7. KcsA Consensus Pharmacophore Search Results .

search database #structures in the database #KcsA structures in the database #hits (total) #KcsA hits from the search #non-KcsA hits from the search non-KcsA hit PDB codes
AbAg.mdb 1211 56 56 56 0 (none)
MOE 2024 16,180 88 89 81 8 1A3R, 4DGI, 4R7D, 4YXH, 4YXK, 4YXL, 7LOK, 7WLY
a

The table shows the total number of structures and the number of KcsA binding antibodies in each database. The number of hits retrieved by using the KcsA consensus pharmacophore model to search each database is also reported. The hits are subdivided into Abs known to bind KcsA and Abs not known to bind KcsA. The PDB codes of hit Abs not known to bind KcsA are also given. Only the search of the larger MOE2024 database found Abs not known to bind KcsA. Three of the eight non-KcsA bound hitsPDB codes 4YXH, 4YXK, 4YXLrepresent the same Ab, so in effect the search found 6 unique non-KcsA-bound Abs.

The MOE 2024 Antibody.mdb database contains 88 structures with Abs bound to KcsA. Searching the MOE 2024 Antibody.mdb database with the KcsA consensus pharmacophore query produces 89 hits, 81 of which are Abs bound to KcsA. The search misses 7 of the known 88 KcsA binding antibodies. The search also finds eight Abs not known to bind KcsA. The PDB codes of the non-KcsA-bound Ab hits are reported in Table . Three of the hits (PDB codes 4YXH, 4YXK, 4YXL) are the same antibody, so in effect the search found six unique Abs which are not known to bind KcsA. Note that the query is very specific, and the search does not produce a long list of hits, representing a significant portion of the search database. This observation implies that the probability of finding false positives is low for our method.

In Table the percent sequence identity, percent sequence similarity, and RMSD when aligned and superposed onto 2IH3 are given for each of the unique non-KcsA binding hits. These antibodies are different from the KcsA binding antibodies and bind different antigens. Particularly, the antibodies contained in the PDB entries 4R7D and 7LOK share low sequence identities with the antibody in the PDB entry 2IH3. These observations provide an example of convergent evolution observed rather rarely in the antibody:antigen complexes.

8. Comparison of Non-KcsA-Bound Hits with the KcsA-Bound Ab Structure in PDB Code 2IH3 .

    full Ab
CDR-Only
PDB entry antigen % identity % similarity Cα RMSD (Å) % identity % similarity Cα RMSD (Å)
1A3R HUMAN RHINOVIRUS (SEROTYPE 2) 75.6 86.1 3.063 31.1 53.3 5.68
4DGI human PrPc fragment 120–230 72.2 76.6 1.199 55.6 62.2 1.625
4R7D no antigen 57.5 73.1 2.071 37.8 46.7 2.335
4YXH deer prion protein 88.6 93.3 1.543 62.2 73.3 2.74
7LOK HIV-1 Env trimer a 44.4 60 2.736 34.4 46.6 5.13
7WLY Omicron S 58.5 75.2 2.351 42.2 64.4 4.03
a

PDB entry codes of the six unique non-KcsA-bound Ab hits that were found by searching the MOE 2024 antibody database using the KcsA consensus pharmacophore. The Abs are compared with a KcsA-bound Ab (PDB code 2IH3) using sequence identity, sequence similarity, and RMSD to 2IH3 in an optimally superposed overlay.

Discussion

With the availability of enormous amounts of antibody sequences from next generation sequencing on naive as well as antigen-experienced B-cell repertoires, display library outputs, and other sources of antibody sequences in the publicly available databases , and rise of generative AI, we asked if it is now feasible to discover therapeutic antibody candidates starting from antibody library design, screening for potential antibody binders to given antigen(s), and optimization of the lead candidates without relying heavily on the experiments. Toward this goal, we have recently proposed a three-step conceptual roadmap called DAbI (Discovery of Antibodies In silico) as described in the introduction. This work is focused on the second step, i.e., structural definition of putative Ab:Ag complexes via virtual screening of antibody libraries against given antigen(s). This is also the most difficult part of our roadmap because of the general lack of virtual screening tools for biologic drug discovery. We have attempted to adapt the pharmacophore-based virtual screening method, commonly used in the realm of small-molecule drug discovery, for use in the context of antibody-based drug discovery. Our results suggest that it is indeed feasible to add this method to the computational toolkit for the discovery and design of novel antibody-based biotherapeutics.

To understand limitations and advantages of the pharmacophore method, we benchmarked it against cognate docking and compared the performance of the two methods using 33 therapeutically relevant Ab:Ag complexes whose crystal structures are available in the PDB. Our results suggest that pharmacophore-based virtual screening is much faster and more accurate when compared with traditional cognate protein:protein docking (Figure and Table ). This is a significant result. However, further tests are needed to fully validate this observation. A limitation of the current approach is that it requires an Ab:Ag complex structure to initiate pharmacophore query building and searches. While this could be still useful for drug repurposing computations, in the real-world scenario for novel biologic drug discovery, we should be able to start virtual screening using the structure of the antigen alone because Ab:Ag crystal structures are often not available at the initial stages of the drug discovery projects.

To our knowledge, this is the first reported use of the pharmacophore concept toward virtual screening of antibodies. There have been previous attempts to fingerprint Ab:Ag interfaces and other protein:protein interfaces and predict antibody:antigen complexes. More recently, deep learning-based methods such as Alpha Fold Multimer have also become available. However, the speed of these methods is incompatible with the time scales required for the rapid virtual screening of large antibody libraries, a clear boon of the pharmacophore-based methodology described in the current work. Moreover, the computational resources and cost of these algorithms running on GPUs and AWS clouds also need to be considered when they are applied at an industrial scale. An exhaustive benchmarking exercise involving all of these different methods, though desirable, is beyond the scope of this study.

It is clearly feasible to use pharmacophore-based methods for virtual screening of antibodies against a given antigen, and it is our hope that widespread adoption of this technique, as part of DAbI or otherwise, shall lead to significant contributions to biotherapeutic drug discovery moving forward.

Acknowledgments

The authors acknowledge their discussions with Dr. Alireza Shahneh and with other colleagues in Moderna, Boehringer-Ingelheim, and Chemical Computing Group (CCG) on developing pharmacophore search-based methods for virtual screening of antibodies. The work described in this manuscript has been implemented as a module in MOE 2024 and is available to all users of MOE.

S.K. came up with the idea of using pharmacophore-based virtual screening for antibody drug discovery. S.K., J.W.D., and C.W. designed the study. C.W., F.M., and D.C.T. carried out the calculations. All authors contributed toward manuscript writing and review.

The authors declare the following competing financial interest(s): C.W. and D.C.T. are employees of Chemical Computing Group. F.M. and J.W.D. are employees of Moderna Inc. S.K. is a former employee of Moderna Inc.

References

  1. Senior M.. Fresh from the Biotech Pipeline: Record-Breaking FDA Approvals. Nat. Biotechnol. 2024;42(3):355–361. doi: 10.1038/s41587-024-02166-7. [DOI] [PubMed] [Google Scholar]
  2. Gray A., Bradbury A. R. M., Knappik A., Plückthun A., Borrebaeck C. A. K., Dübel S.. Animal-Free Alternatives and the Antibody Iceberg. Nat. Biotechnol. 2020;38(11):1234–1239. doi: 10.1038/s41587-020-0687-9. [DOI] [PubMed] [Google Scholar]
  3. Olsen T. H., Boyles F., Deane C. M.. Observed Antibody Space: A Diverse Database of Cleaned, Annotated, and Translated Unpaired and Paired Antibody Sequences. Protein Sci. 2022;31(1):141–146. doi: 10.1002/pro.4205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Christley S., Aguiar A., Blanck G., Breden F., Bukhari S. A. C., Busse C. E., Jaglale J., Harikrishnan S. L., Laserson U., Peters B., Rocha A., Schramm C. A., Taylor S., Vander Heiden J. A., Zimonja B., Watson C. T., Corrie B., Cowell L. G.. The ADC API: A Web API for the Programmatic Query of the AIRR Data Commons. Front. Big Data. 2020;3:22. doi: 10.3389/fdata.2020.00022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Rajagopal N., Choudhary U., Tsang K., Martin K. P., Karadag M., Chen H.-T., Kwon N.-Y., Mozdzierz J., Horspool A. M., Li L., Tessier P. M., Marlow M. S., Nixon A. E., Kumar S.. Deep Learning-Based Design and Experimental Validation of a Medicine-like Human Antibody Library. Briefings Bioinf. 2025;26:bbaf023. doi: 10.1093/bib/bbaf023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Bauer J., Rajagopal N., Gupta P., Gupta P., Nixon A. E., Kumar S.. How Can We Discover Developable Antibody-Based Biotherapeutics? Front. Mol. Biosci. 2023;10:1221626. doi: 10.3389/fmolb.2023.1221626. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Yokoo H., Shibata N., Demizu Y.. Protein-Protein Docking, Methods and Protocols. Methods Mol. Biol. 2024;2780:345–359. doi: 10.1007/978-1-0716-3985-6_18. [DOI] [PubMed] [Google Scholar]
  8. Paggi J. M., Pandit A., Dror R. O.. The Art and Science of Molecular Docking. Annu. Rev. Biochem. 2024;93(1):389–410. doi: 10.1146/annurev-biochem-030222-120000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Lin P., Li H., Huang S.-Y.. Deep Learning in Modeling Protein Complex Structures: From Contact Prediction to End-to-End Approaches. Curr. Opin. Struct. Biol. 2024;85:102789. doi: 10.1016/j.sbi.2024.102789. [DOI] [PubMed] [Google Scholar]
  10. Shor B., Schneidman-Duhovny D.. Integrative Modeling Meets Deep Learning: Recent Advances in Modeling Protein Assemblies. Curr. Opin. Struct. Biol. 2024;87:102841. doi: 10.1016/j.sbi.2024.102841. [DOI] [PubMed] [Google Scholar]
  11. Rangel M. A., Bedwell A., Costanzi E., Taylor R., Russo R., Bernardes G. J. L., Ricagno S., Frydman J., Vendruscolo M., Sormanni P.. Fragment-Based Computational Design of Antibodies Targeting Structured Epitopes. bioRxiv. 2022:2021.03.02.433360. doi: 10.1101/2021.03.02.433360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Sormanni P., Aprile F. A., Vendruscolo M.. Rational Design of Antibodies Targeting Specific Epitopes within Intrinsically Disordered Proteins. Proc. Natl. Acad. Sci. U.S.A. 2015;112(32):9902–9907. doi: 10.1073/pnas.1422401112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Nimrod G., Fischman S., Austin M., Herman A., Keyes F., Leiderman O., Hargreaves D., Strajbl M., Breed J., Klompus S., Minton K., Spooner J., Buchanan A., Vaughan T. J., Ofran Y.. Computational Design of Epitope-Specific Functional Antibodies. Cell Rep. 2018;25(8):2121–2131 e5. doi: 10.1016/j.celrep.2018.10.081. [DOI] [PubMed] [Google Scholar]
  14. Muegge I., Bentzien J., Ge Y.. Perspectives on Current Approaches to Virtual Screening in Drug Discovery. Expert Opin. Drug Discov. 2024;19(10):1173–1183. doi: 10.1080/17460441.2024.2390511. [DOI] [PubMed] [Google Scholar]
  15. da Rocha M. N., de Sousa D. S., da Silva Mendes F. R., dos Santos H. S., Marinho G. S., Marinho M. M., Marinho E. S.. Ligand and Structure-Based Virtual Screening Approaches in Drug Discovery: Minireview. Mol. Divers. 2025;29:2799–2809. doi: 10.1007/s11030-024-10979-6. [DOI] [PubMed] [Google Scholar]
  16. Lima A., Penteado A., de Jesus J., de Paula V., Ferraz W., Trossini G.. Structure-Based Virtual Screening: Successes and Pitfalls. J. Braz. Chem. Soc. 2024;35(10):e-20240112. doi: 10.21577/0103-5053.20240112. [DOI] [Google Scholar]
  17. Wermuth C. G., Ganellin C. R., Lindberg P., Mitscher L. A.. Glossary of Terms Used in Medicinal Chemistry (IUPAC Recommendations 1998) Pure Appl. Chem. 1998;70(5):1129–1143. doi: 10.1351/pac199870051129. [DOI] [Google Scholar]
  18. Molecular Operating Environment (MOE), 2022.02; Chemical Computing Group ULC: 910-1010 Sherbrooke St. W., Montreal, QC: H3A 2R7, 2022. [Google Scholar]
  19. Berman H. M., Westbrook J., Feng Z., Gilliland G., Bhat T. N., Weissig H., Shindyalov I. N., Bourne P. E.. The Protein Data Bank. Nucleic Acids Res. 2000;28(1):235–242. doi: 10.1093/nar/28.1.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Williams, C. Protein-Protein Interface Pharmacophore Query; Scientific Vector Language (SVL) Source Code Provided by Chemical Computing Group ULC: 910-1010 Sherbrooke St. W., Montreal, QC: H3A 2R7, 2024, 2024. [Google Scholar]
  21. Mehra R., Kepp K. P.. Structure and Mutations of SARS-CoV-2 Spike Protein: A Focused Overview. ACS Infect. Dis. 2022;8(1):29–58. doi: 10.1021/acsinfecdis.1c00433. [DOI] [PubMed] [Google Scholar]
  22. Wong W. K., Robinson S. A., Bujotzek A., Georges G., Lewis A. P., Shi J., Snowden J., Taddese B., Deane C. M.. Ab-Ligity: Identifying Sequence-Dissimilar Antibodies That Bind to the Same Epitope. mAbs. 2021;13(1):1873478. doi: 10.1080/19420862.2021.1873478. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Porter K. A., Desta I., Kozakov D., Vajda S.. What Method to Use for Protein-Protein Docking? Curr. Opin. Struct. Biol. 2019;55:1–7. doi: 10.1016/j.sbi.2018.12.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Collins K. W., Copeland M. M., Brysbaert G., Wodak S. J., Bonvin A. M. J. J., Kundrotas P. J., Vakser I. A., Lensink M. F.. CAPRI-Q: The CAPRI Resource Evaluating the Quality of Predicted Structures of Protein Complexes. J. Mol. Biol. 2024;436(17):168540. doi: 10.1016/j.jmb.2024.168540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Zhao N., Wu T., Wang W., Zhang L., Gong X.. Review and Comparative Analysis of Methods and Advancements in Predicting Protein Complex Structure. Interdiscip. Sci.: Comput. Life Sci. 2024;16(2):261–288. doi: 10.1007/s12539-024-00626-x. [DOI] [PubMed] [Google Scholar]
  26. Goldstein L. D., Chen Y.-J. J., Wu J., Chaudhuri S., Hsiao Y.-C., Schneider K., Hoi K. H., Lin Z., Guerrero S., Jaiswal B. S., Stinson J., Antony A., Pahuja K. B., Seshasayee D., Modrusan Z., Hötzel I., Seshagiri S.. Massively Parallel Single-Cell B-Cell Receptor Sequencing Enables Rapid Discovery of Diverse Antigen-Reactive Antibodies. Commun. Biol. 2019;2(1):1–10. doi: 10.1038/s42003-019-0551-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Gainza P., Sverrisson F., Monti F., Rodola E., Boscaini D., Bronstein M. M., Correia B. E.. Deciphering Interaction Fingerprints from Protein Molecular Surfaces Using Geometric Deep Learning. Nat Methods. 2020;17(2):184–192. doi: 10.1038/s41592-019-0666-6. [DOI] [PubMed] [Google Scholar]
  28. Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A., Bridgland A., Meyer C., Kohl S. A. A., Ballard A. J., Cowie A., Romera-Paredes B., Nikolov S., Jain R., Adler J., Back T., Petersen S., Reiman D., Clancy E., Zielinski M., Steinegger M., Pacholska M., Berghammer T., Bodenstein S., Silver D., Vinyals O., Senior A. W., Kavukcuoglu K., Kohli P., Hassabis D.. Highly Accurate Protein Structure Prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Molecular Pharmaceutics are provided here courtesy of American Chemical Society

RESOURCES