Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 May 1.
Published in final edited form as: J Mol Liq. 2022 Feb 18;353:118759. doi: 10.1016/j.molliq.2022.118759

Machine learning assessment of the binding region as a tool for more efficient computational receptor-ligand docking

Matjaž Simončič a, Miha Lukšič a, Maksym Druchok b,c,*
PMCID: PMC8903148  NIHMSID: NIHMS1783332  PMID: 35273421

Abstract

We present a combined computational approach to protein-ligand binding, which consists of two steps: (1) a deep neural network is used to locate a binding region on a target protein, and (2) molecular docking of a ligand is performed within the specified region to obtain the best pose using Autodock Vina. Our in-house designed neural network was trained using the PepBDB dataset. Although the training dataset consisted of protein-peptide complexes, we show that the approach is not limited to peptides, but also works remarkably well for a large class of non-peptide ligands. The results are compared with those in which the binding region (first step) was provided by Accluster. In cases where no prior experimental data on the binding region are available, our deep neural network provides a fast and effective alternative to classical software for its localization. Our code is available at https://github.com/mksmd/NNforDocking.

Keywords: molecular docking, machine learning, deep neural network, Accluster, AutoDock Vina

Graphical Abstract

graphic file with name nihms-1783332-f0001.jpg

1. Introduction

One of the cornerstone research directions in pharma is the search for efficient drug compounds. A well-designed drug needs to selectively bind to an active site of a target receptor and make little or no effect on other receptors. Such a design process is often time-consuming and labor-intensive, as the chemical space of small drug-like molecules is ultimately huge [1]. Therefore, it is critical that the search for in silico candidates precedes in vitro testing.

Molecular docking is a powerful computational approach in biochemistry and drug design for candidate suggestion [26]. The main objective of most docking techniques is to predict a binding pose, i.e., the orientation and conformation of a docked molecule (ligand) within a binding site on a target protein of interest (receptor). Thus, the binding site is a fragment of the receptor where the binding occurs and usually consists of a few amino acid residues. However, if the ligand undergoes a chemical change (catalyzed reaction) upon binding, or if the receptor is either inhibited or conformationally changed, the binding site is refered to as the active site. These are particularly important when it comes to drug design. Finding binding sites, let alone active sites, can be a tedious and time-consuming process. Therefore, several approaches towards finding binding sites as well as the position of a ligand molecule within them have been developed. Simulation-based techniques, such as energy optimization methods, while being sensitive to the choice of force field (see for example Refs. 711), offer a great deal of precision and can be successful in this regard [12, 13]. Unfortunately, this comes with high computational cost, which limits the use of simulation methods in high-throughput screening protocols. Another path relies on the utilization of less precise but faster methods. Here, both false positives and false negatives are tolerated as long as true positives are found at a sufficiently high rate to justify the effort. This is reflected in a number of reviews or comparative studies discussing available docking methods and recent advances in the field [1417]. Hybrid approaches, e.g., based on machine learning (ML) [18], can serve as a good alternative with decent accuracy and low computational time. Similar ideas of multistage protocols are also discussed in Refs. 1921.

In general, docking methods can be classified as template-based and structure-based docking methods. Template-based (homology modelling) approaches use available information about atoms and amino acids within the docked fragments of ligands and receptors and search for similar structural sequences in a receptor-ligand pair of interest [22, 23]. Structure-based techniques examine the surface of the receptor for attractive sites and optimize the position (orientation) of the ligand within these sites using scoring functions [24, 25]. Blind structure-based docking searches for a potential binding site over the whole receptor, whereas in target structure-based docking the search is restricted to a specific part of the receptor. The spatial constraint must be reliable, otherwise docking will give an erroneous result. However, if the region is predicted correctly, target docking has more chance of achieving a reasonable level of precision – fewer iterations will be spent on investigating weakly-confident areas, allowing to dedicate more computational power to fine-tuning the ligand position. A number of papers (see, for example, Refs. 15, 26, 27) report a higher level of success with different target docking-based approaches, which highlight the importance of the binding site localization. Several quite successful receptor binding site prediction protocols exist suitable for various classes of ligand molecules, such as eFindSite [28], ACCLUSTER [29], LIBRA [30], as well as protocols intended solely for peptides: PeptiMap [31], PEP-SiteFinder [32], etc. Also, many docking programs provide the position and conformation of ligands within binding sites, as well as the amino-acid residues responsible for binding. There is a number of platforms for ligand-receptor docking, such as ZDOCK [33], Rosetta FlexPepDock [34, 35], Rosetta FlexPepDock ab initio [36], SwissDock [37], GalaxyPepDock [38], BSP-SLIM [39], DOCK 6 [40], DynaDock [41], CABS-Dock pepATTRACT [43], HADDOCK [44], AutoDock Vina [45], CoDockPP [46], PLANTS [47] etc. For a review see Refs. 15, 17, 48, 49. The drawback of most of these approaches is that they are specialized for docking of certain types of ligands (peptides, nucleic acids, etc.) or impose size limitations on receptors/ligands. Therefore, the development of new and generally applicable hybrid approaches is almost a necessity.

Most of the above docking programs are designed to output multiple ligand poses within different binding sites according to internal scoring mechanisms. What makes determining the correct binding pose even more challenging is that a top-ranked pose in the docking calculation is not always predicted within the experimentally determined binding site on the receptor molecule [4951]. Finding wrong binding sites can occur for a variety of reasons: some arise from physicochemical conditions not fully mapped during docking, others from representation uncertainties (resolution of the X-ray crystallographic structure), and some from binding sites being buried within the protein. For this reason, it is even more important to have an assessment of the binding region on the first place – the part of the receptor that contains the actual binding site – and then perform target docking within the defined region.

It is also worth mentioning the scoring rules used to assess and distinguish ligand poses. Scoring rules can be divided into four groups [25]: force-field, knowledge-based, empirical, and machine-learning. Each docking pose can be represented using different descriptors, which usually carry hints about binding-related mechanisms, e.g. H-bond donors and acceptors, charge, polarity, hydrophobicity, etc. They can be thought of as matching rules, e.g.: “H-bond donors of a particular residue of the receptor will most likely bind with H-bond acceptors of the ligand”. Of course, this is a very simplified interpretation and actual approaches consider multiple rules between multiple sites and rate them to mark the binding site or evaluate the binding energy. For example, a study by Chupakhin et al. [52] discusses an approach based on introduced interaction finger-prints, verifying geometric rules and evaluating binding sites of three target receptors. An ML-approach, where a neural network learns key features from a dataset of protein-ligand complexes, therefore, develops its own empirical scoring rules. A study by Ragoza et al. [53] proposes an approach, where a convolutional neural network is trained to learn key features for protein-ligand interactions and, thus, rate different ligand poses with its score. In a study by Taherzadeh et al. [54] proteins and peptides are featurized and fed to a random forest model to locate binding regions on protein surfaces. A genetic algorithm is used in Ref. 55 to dock ligands to the HIV-1 protease. Another approach combining binding affinity and docking stages is used in Ref. 56 to screen ligands for potential anti-tumor inhibitors. Interestingly, the binding affinity assessment there is done in an ensembled manner that combines predictions from various machine learning techniques. Combined pipelines with a docking stage are also applied in a number of recent studies to mention a few [5761]. In a study by Banchi et al. [62] a more exotic approach of gaussian boson sampling is used for the molecular docking problem. In Ref. 62 the protein-ligand complex is represented as a graph and the problem is formulated as a search of maximum weighted clique within the graph.

In this study, we focus on the approach of combining a machine learning stage to localize a binding region on protein with a stage of target-based molecular docking. The performance is compared to the one of an analogous procedure where the binding region, instead of ML, is provided by another software, Accluster. Results for protein-peptide and protein-(non-peptide) docking are presented. Although our neural network was trained on peptide-protein complexes, we show that the approach is not limited to peptide ligands, but can also be used in cases with non-peptides. In addition, we demonstrate that ML-based hybrid docking can outperform classical approaches with respect to time.

2. Methods and materials

The development and tuning of the neural network was performed on the cases from the PepBDB v.2019 dataset [63]. After this stage we focus on the localization of the binding region (ML, Accluster) on the receptor surface followed by target molecular docking of both peptide and non-peptide ligands using the Autodock Vina package [45].

The general scheme (Fig. 1) consists of two stages: (1) for a given receptor-ligand pair, the neural network predicts the coordinates of the ligand relative to the receptor. This approximate ligand position represents the potential binding region on the target molecule. (2) A target cell is defined around the determined region and docking is performed within it. The receptors (proteins) are large molecules posing a large space of possibilities for docking a ligand. The task of the first step (machine learning) is to narrow down the search space, while in the second stage target docking within a determined binding region is performed. The effectiveness of the first step can often improve the overall procedure with respect to time, especially in high-throughput protocols.

Figure 1:

Figure 1:

The scheme of our approach combining neural network (NN) and AutoDock Vina (ADV). Stage 1: the NN takes protein and ligand structures as inputs (left) and predicts the binding region on the protein (center). Stage 2: ADV performs target-docking of the ligand within the predicted binding region (right).

2.1. Basic Assumptions

Before diving into the computational details of the first step (machine learning), we need to highlight the assumptions and the data preparation steps. The first pre-processing step, performed over receptor-ligand complexes, is a linear translation of the coordinates in a way that the geometric centers of receptors coincide with the origin of coordinates (shifts on average being approximately 60 Å). Another pre-processing involves excluding receptor-ligand complexes containing atoms other than C, O, N, P, S, F, Cl, Br, and I. The hydrogen atoms (PepBDB does not include hydrogens) and oxygen atoms associated with hydrating water molecules are either already pre-deleted from PDB files or deleted by hand. Further, the intramolecular structure of the receptors (as taken from a PDB file) is considered frozen. The naming convention in PepBDB follows the pattern AAAA B, where the first part (before the underscore) refers to the four-letter-digit code of the protein-peptide complex in the Protein Data Bank and the second part (after the underscore) refers to the ID of the peptide chain within the complex.

Our neural network design imposes two constraints on the size of receptors and ligands. The vast majority of samples within PepBDB have receptors with the number of heavy atoms below 4000, thus, we chose this number as the upper size limit for receptors. A similar limit of 300 atoms was also imposed for ligands.

2.2. Data and representations

To represent receptor and ligand molecules participating in docking we use three types of data structures: one-hot encodings for receptors and ligands, adjacency matrices for ligands, and matrices of 3D coordinates for receptors.

It is convenient to represent categorical data using one-hot encodings. In the case of chemical compounds this is a N × M matrix of 0s and 1s. N runs over all atoms, and M is the size of the dictionary of valid chemical elements (in our case C, O, N, P, S, F, Cl, Br, and I). We only consider receptors that are not larger than 4000 atoms and ligands that are not larger than 300, applying zero-padding for the shorter ones. An additional element is added to the dictionary of atomic species – “None”, indicating the absence of atoms. Thus, the one-hot matrices have dimensions of 4000 × 10 and 300 × 10 for receptors and ligands, respectively.

The adjacency matrix for ligands is a square symmetric matrix of dimension 300 × 300, gathering interatomic distances, being placed on the corresponding intersections of rows and columns. This matrix consists only of distances less than 2.0 Å, otherwise the elements are set to zero. Such a threshold of 2.0 Å ensures that only chemical bonds between adjacent atoms are considered, implying no constraints on non-bonded atom pairs. None of the absolute 3D coordinates are used as ligand inputs, since they need to be predicted.

The coordinate representation for receptors is a matrix with dimensions 4000 × 3, containing 3D Cartesian coordinates of receptor atoms. If the number of receptor atoms, Nreceptor, is less than 4000, the concluding 4000−Nreceptor elements are filled with zeros.

The above data structures serve as inputs in the machine learning procedure of the neural network, while the task solved by the neural network is the prediction of atomic coordinates of the ligands. Therefore, during the training step, we also use matrices (of dimension 300 × 3) with Cartesian coordinates of ligands and zero-padding.

As mentioned earlier, the original PepBDB v.2019 dataset was filtered to select samples, following constraints on the number of atoms and allowed atomic types. The final number of samples was about 10 000. A random split into six folds with k-fold validation technique was used to assess the performance metrics.

The last aspect of the data processing is data augmentation. In order to enrich the train dataset, we applied two types of augmentations: (1) coordinate translation by a random shift of |Δr|8.6Å and (2) three random rotations around the X, Y, and Z axes. All these augmentations were treated as independent samples along with the original ones, thus constituting more than 50 000 training samples in total. The above extensions help the neural network learn the translation and rotation invariance of docked complexes.

2.3. Neural network architecture

Supervised machine learning techniques, in particular artificial neural networks, are widely used in routine applications such as speech recognition and computer vision. A basic neural network consists of input and output layers, and hidden layers in between, each consisting of multiple neurons. The neurons pass a received signal through a weighted input and an activation function. The outputs are propagated in a feed-forward manner from the input layer to the output layer. The set of weights that parameterize neurons are iteratively optimized to fit a given training set of data to minimize the error of the network. Modern deep learning setups often involve combinations of neuron types and activation functions, branching architectures (such as separate inputs or residual connections), and other tricks that help improve performance. We implemented our deep learning neural network from scratch using PyTorch, a Python-based scientific computing package. PyTorch provides most of the required building blocks and thus offers enough flexibility to customize the functionality of the neural network.

A high-level architecture of the neural network is depicted in Fig. 2. First, it has four input “arms”, that process the data separately. Each arm consists of four 2D convolution layers (3 consequent and 1 residual) and one fully connected layer. Since the ligand adjacency matrix is sparse by definition, we also included two max-pooling layers to reduce the dimensionality along this arm. Then, the outputs of all arms are concatenated, normalized, and further regressed to predict the ligand coordinates, as shown in the right side of the figure. The activation functions are ReLU over the entire network with the sole exception of the last layer producing ligand coordinates where no activation functions are applied.

Figure 2:

Figure 2:

A principal architecture of the neural network.

The whole network was trained from scratch, back-propagating the errors to input layers of each arm. We used the stochastic gradient descent technique implemented in the ADAM optimizer to train the neural network, with learning rates of 10−3 at the beginning of the training and 10−4 at the end. The loss is calculated as a mean square error between actual and predicted coordinates of ligands. During the train routine we tracked the performance on train and test subsets and saved the corresponding models. Further, general-purpose predictions are made using the model that showed the best performance on the test subset.

We also noticed that the predicted atomic coordinates of large ligands might experience a decay towards the origin of coordinate system. One can suppose that such a bias is due to a weaker population of neurons with non-zero values – the neural network learns the fact that many ligands have number of atoms Nligand lower than 300, therefore, the concluding 300 − Nligand neurons are preferably trained to output zeros. To validate this idea we conducted two additional experiments with different zero-padding setups: (1) zero-padding preceded the coordinates and (2) the coordinates were located in the middle of coordinate matrix with zero-padding applied on both sides. Indeed, these experiments helped confirm the above assumption – the coordinate decay appeared on weakly populated and thus “unconfident” neurons. In order to overcome this issue, we utilize two neural networks with identical architectures but opposite zero-paddings: zeros either precede or follow the non-zero data. Both networks are trained separately, while their outputs are merged with a weighting factor k in a form of a sigmoid function, prioritizing the “confident” neurons.

In order to preserve the intramolecular connectivity for the predicted samples we tested an additional term in the loss function. Using the predicted coordinates of a ligand one can reconstruct an adjacency matrix of a ligand and apply an additional penalty for the difference between the reconstructed and the input adjacency matrices (the adjacency matrix of a ligand is one of the inputs of the prediction). However, this approach resulted in a balance between a slight improvement of the intramolecular geometry and worsening of the binding site prediction. Therefore, we decided to keep the loss function in the original form of the mean square error between the predicted and experimentally determined coordinates.

Our neural network-based predictor for ligand-receptor docking is available on GitHub (URL: https://github.com/mksmd/NNforDocking).

2.4. Combination of ML/Accluster and AutoDock Vina

The binding region for a protein-ligand pair is provided by ML, where the initial guess are the coordinates of the ligand with respect to the protein. For comparison, the binding region is also predicted by Accluster [29], a web-based software for the prediction of peptide binding sites on proteins. The latter uses ZDOCK3.1 [33] for the docking of 20 standard amino acids to globally scan the surface of a given protein. If provided with a FASTA file of the peptide only the amino acids of the peptide are used for this procedure, which reduces the time for the prediction. Modes of amino acids forming favourable interactions with the protein surface are clustered, while the clusters are ranked with respect to their size, which usually results in up to three binding regions. To keep the paper concise, we limit our evaluation only to the top-1 ranked cluster within the main part of the paper. The assessment of all three binding regions (cumulative results) can be found in the Supplementary material.

Once the binding region is predicted, docking is performed using AutoDock Vina with default parameters and point charges assigned according to the AMBER14 [64] force field. The setup was performed using the YASARA [65] molecular modeling program. Twenty-five docking runs were performed per receptor-ligand pair, and the resulting ligand poses were clustered if the root-mean-square deviation (RMSD) between any two poses was less than 5 Å. Each cluster is represented by the pose with the highest score according to the scoring function. Within the scope of the paper we consider only the first ranked pose (top-1). In all cases the receptors were considered rigid (frozen coordinates), allowing flexibility for ligands.

3. Results

3.1. Testing ML predictions on the PepBDB dataset

In order to quantify the predictive capability of the neural network, we plot the distribution over the test samples as a function of their ligand RMSDs between experimental and predicted poses (Fig. 3). The initial sharp peak at 2–6 Å is followed by a gradual decrease of the distribution with distance. We visually inspected the predicted configurations and selected some typical examples with different RMSDs for illustration. These examples are presented in Fig. 4, where gray ribbons depict the receptors, the ligands of the experimental structures are shown with black sticks, and the positions predicted by ML are shown with green spheres. The top snapshot shows an example of a prediction with RMSD of 3.0 Å. The backbone of the peptide is well reproduced in contrast to its side groups. The middle snapshot shows a slightly weaker prediction with RMSD of 8.5 Å – the position of the backbone matches the experimental structure of the ligand only partially. The bottom snapshot shows a prediction with RMSD of 11.8 Å – in this particular case the prediction correctly guessed the cavity, but the position of the backbone was poorly reproduced. One might conclude that the proposed approach reconstructs the intramolecular geometry rather weakly, while being successful in pinpointing the binding region where the biding site is located. From our experience, the binding region can be decently identified when the prediction RMSD falls below 10 Å. There are 824 test samples (out of 1492) with RMSDs below 10 Å. If one is interested in localizing the binding region only, these numbers may be more optimistic for two reasons: (1) some of the RMSDs above 10 Å may be a consequence of the reversed orientation of a ligand, although the binding region is still correctly identified (not accounted for in the 824 “correct” samples), and (2) partial interactions of the ligand with the receptor surface may also generate high RMSDs (some parts of a ligand float in the bulk solvent), overestimating the error due to less constrained parts of the ligand.

Figure 3:

Figure 3:

The distribution of samples from the test subset binned by the RMSDs between experimental and ML-predicted atomic positions.

Figure 4:

Figure 4:

Examples of the ML prediction. The gray ribbons depict proteins, the black sticks represent ligands of the experimental structure, and the green spheres show the ML predictions. Top snapshot shows the prediction for the case 4apr_I with RMSD of 3.0 Å, central snapshot – the case 6gk7_A with RMSD of 8.5 Å, bottom – the case 1bcr_C with RMSD of 11.8 Å.

We wanted to compare the predictive power of the ML algorithm with that of Accluster. We therefore used the same criterion for both approaches. The authors of Accluster considered a prediction successful if the geometric center of the amino acid cluster was within 4 Å of any atom of the peptide. By this metric, the success rate of our approach, calculated from the test data set, is about 60% (891 out of 1492). Accluster provides multiple binding regions (clusters). The success rate when considering only the top-1 cluster was 54.6% [29], although the success rate when considering also the predictions ranked 2 and 3 increased to 71.8%. Given the computational cost of the ML algorithm, which provides the prediction in less than a second, and Accluster, which on average takes approximately 20 minutes for the prediction (ranging from 10 to 50 minutes), the cost-effective computational advantage of our approach is evident (2–3 orders of magnitude gain in speed).

At this stage of the study, we concluded that our ML solution can be used as a fast and lightweight tool to locate binding regions and needs to be complemented by a second step to perform target docking within the located regions.

As further elaborated, our ML solution can be also useful in distinguishing between closely scored predictions from other docking services (the usefulness is showcased for ZDOCK in Appendix A).

3.2. Docking of peptides

In this subsection, we show the results of peptide docking with AutoDock Vina based on our ML approach for estimating the binding region as well as based on Accluster estimation. AutoDock Vina is designed to dock small ligand molecules, so its use for peptides is rather limited to short-chain peptides [50]. However, unlike peptide-oriented software that usually requires FASTA sequences or equivalent PDB files with amino acids, it is more flexible and works easily with non-peptide ligands (peptidomimetics, non-peptides). It provides a target-based approach to molecular docking, which requires the construction of a cell around the binding site (the search space), so a general idea of the potential binding region is a necessity. A hybrid approach of ML, which serves as an initial guess for the binding region, and target molecular docking, which optimizes the ligand geometry, has demonstrated its efficiency.

We analyzed 36 protein-peptide complexes from the Protein Data Bank, which were not included in the PepBDB v.2019 dataset. For each of these cases, we performed two docking procedures: (1) target docking based on the ML prediction and (2) target docking based on the prediction of Accluster (top-1 cluster). In the first case, the prediction from ML served as the initial guess. In the second case we provided Accluster with the FASTA file of the peptide whenever possible (conforming to the size limitation), which resulted in up to three binding regions consisting of a cluster of amino acids, however only the first one (top-1) was considered. A rectangular cell with a selected minimum distance (L) to each atom of the ML-guess (neural-network) or cluster of amino acids (Accluster) was defined and docking was performed within this confined space. The cell size is a hyperparameter, which means that we could perform molecular docking with different cell sizes to investigate their influence on the docking accuracy. If the search space (cell size) is too small, a ligand cannot explore the binding region with enough flexibility, whereas if it is too large the computational power is wasted accessing regions of the protein where no binding occurs. Another inspected parameter was the cutoff distance, rcutoff, between the ligand atoms and the receptor: in receptor-ligand interactions, not all ligand atoms contribute to binding; some part of the ligand may freely float in the bulk. These atoms should therefore not be evaluated when comparing the results of molecular docking with the experimental structure. Our analysis showed that rcutoff does not play a major role in the calculation of RMSDs of a given assessment, so this was not pursued further and will not be discussed here.

Fig. 5 summarizes the results of cell size fine-tuning (RMSDs for each case and cell sizes can be found in Table S1 in the Supplementary material file). The histogram compares three different docking approaches (using AutoDock Vina): target molecular docking using the (1) experimental pose of the ligand in the center of the cell (shown in blue), (2) ML prediction as the center of the cell (shown in green), and (3) the top-1 ranked cluster obtained by Accluster in the center of the cell (shown in red). The histogram accounts for the results with different cell sizes, as indicated in the color legend at the right. The reproducibilities for ML- and Accluster-based dockings are comparable (between 10 and 20% depending on the cell size) for the first bin of 0–2.5 Å. Upon considering the second bin (2.5–5 Å) the ML-based approach is slightly better than Accluster. Such low percentages could be due to either an inaccurate initial estimate of ML/Accluster or the fact that AutoDock Vina is not an optimal tool for docking longer peptides. To evaluate the latter claim we repeated target docking by centering the cell around the experimental pose of the ligand (ground truth). The reproducibility in this case (shown in blue) improves only slightly in comparison to dockings where ML provided the initial guess, however, the success rate is still far from ideal. This supports the above statement that AutoDock Vina was not originally designed for peptides and might miss the binding site even when provided with a cell around the exact experimental pose. The cell size influences the reproducibility to a minor extent, however, general conclusions could not be drawn from the histograms.

Figure 5:

Figure 5:

The histogram displaying the effect of the cell size on the percentage of peptide poses reproduced by (1) target molecular docking within cells centered around the ligand of the experimental pose (blue), centered around (2) the ML prediction (green) or (3) the top-1 ranked cluster by Accluster (red) according to their RMSDs with respect to the experimental pose per each bin (2.5 Å intervals). Only the results with RMSDs below 10 Å are shown.

Since Accluster provides up to three different clusters of amino-acids (binding regions), the outcome could be improved by considering all of them and presenting the results as cumulative reproducibilities in the histogram (see Figure S1 in the Supplementary information). The drawback of such an approach is the computational cost, namely one would have to perform additional dockings to achieve a better result. Considering that it takes Accluster on average 20 minutes (less if the FASTA of the peptide is provided) to estimate the binding regions one would additionally have to perform up to two target dockings, whereas when ML provides the binding region, it does so in under a second, after which only one docking needs to be performed. Such an advantage is extremely sought after in high throughput screening protocols.

3.3. Docking of small organic molecules (non-peptides)

In the above subsection, an ML-assisted target molecular docking for peptide ligands was discussed. Since the binding domains on the receptors generally do not change much with respect to ligand type (peptide, non-peptide), we decided to test the same scheme for small organic molecules. The authors of Accluster noticed that, although in their approach the surface of the protein is probed with amino acids, the software gives adequate results also for non-peptide ligands. Consequently, we decided to also compare our ML-assisted approach to Accluster in regard to non-peptide ligand docking. We searched the Protein Data Bank for complexes with protein receptors and small non-peptide ligands. Cases conformed to the size limitation of the ML approach (no more than 300 atoms for ligands and no more than 4000 atoms for receptors) and included only allowed atomic species (C, O, N, P, S, F, Cl, Br, and I). We selected 58 complexes and performed a similar set of docking experiments as in the previous section.

The histogram of the predicted RMSDs is shown in Fig. 6 (RMSDs for each case and cell sizes can be found in Table S2 in the Supplementary material file). Similar to the histogram for peptides (Fig. 5), target molecular docking based on Accluster (red bars), target molecular docking based on ML (green), and based on experimental structure (blue) are compared for different cell sizes. The histogram is qualitatively different from the peptide case. The first band of bars for RMSDs below 2.5 Å is significantly higher than the other bands. In contrast to what can be seen in Fig. 5, a clear trend in regard to the cell size can be drawn. For ML-assisted docking low reproducibilities for smaller cell sizes are a consequence of the binding site not being within the search space. The outcome can be improved by increasing the cell size. If the cell edges are at least 15 Å from each of the heavy atoms of the ML-prediction, the results are comparable to those where docking is performed within the cell centred around the position of the ligand of the experimental structure. Contrary to ML, the trend for Accluster is reversed – the quality of the prediction decreases with increasing cell size. The reason for this is that the potential binding region, i.e. a cluster of amino acids, is on average larger than the ML guess and consequently the potential binding region often already encompasses the correct binding site. The cumulative histogram, considering the best out of three clusters provided by Accluster, is shown on Figure S2 in the Supplementary material.

Figure 6:

Figure 6:

Same as in Fig. 5 but for small organic (non-peptide) ligands.

To conclude this part, we selected three examples to showcase some successful and failed dockings and present them in Fig. 7. The color code is similar to the previous examples: gray ribbons depict proteins, black sticks the ligand positions of experimental structure, green spheres the ML prediction, and green sticks the target docking made by AutoDock Vina with the use the ML prediction as an initial guess. The red and magenta spheres show the clustered predictions made by Accluster (red stands for ranked 1 clusters, and magenta for ranked 2 and 3 ones), the red and magenta sticks are the corresponding target dockings made by AutoDock Vina. The blue sticks depict target dockings by AutoDock Vina with the use of experimental positions as an initial guess. The top snapshot (6o5j from the Protein Data Bank) in Fig. 7 shows the case where three guesses (ML, top-1 by Accluster, and experimental) lead to success in target docking: the corresponding poses (green, red, and blue sticks) coincide with ground truth (black sticks) almost perfectly. The red “cloud” covers the whole binding site. In the middle snapshot (the 6w63 case), one can see that only black and blue sticks coincide, the ML guess (green spheres) is relatively close, but the target docking (green sticks) fails. All clustered predictions by Accluster (red and magenta spheres) miss the binding site, being located on the opposite side of the receptor, which obviously leads to the fails in the corresponding target predictions. The bottom snapshot shows the case of 6lvm: all Accluster predictions failed to find the binding site, while the target dockings using ML and experimental positions as initial guesses reproduced the ground truth almost perfectly.

Figure 7:

Figure 7:

The examples of target dockings using our ML, Accluster, and experimental ligand positions to localize binding regions. The gray ribbons depict proteins, black sticks – experimental positions of ligands, red and magenta spheres – clustered predictions made by Accluster for ligands. Red spheres indicate the highest scored cluster (by Accluster), magenta spheres – second and third scored clusters. The red and magenta sticks are the target dockings made by AutoDock Vina with the use of the corresponding Accluster predictions. Green spheres show ML predictions for ligands, the top-1 scored positions within target dockings by AutoDock Vina are shown by green sticks. The dockings made by AutoDock Vina using experimental positions as initial guess are depicted by blue sticks. The examples are taken from the Protein Data Bank, their PDB identifiers are: top – 6o5j, middle – 6w63, bottom – 6lvm.

For the sake of simplicity, we evaluated only the target docking poses rated as top-1 by AutoDock Vina, although the presented results could be improved by selecting docked poses beyond top-1. Other approaches towards evaluating molecular docking results often consider multiple docking poses and consequently present a cumulative success rate. This is particularly important because top-ranked poses do not necessarily coincide with experimentally determined ligand positions [51, 52]. In particular, Agrawal et al. [49] in a benchmark study of protein-peptide docking software distinguished between the “best” (according to the scoring function) and “top” (according to the experimental structure of the complex) ranked poses. For this reason, we also demonstrate the use of the ML-prediction for the reprioritization of clusters provided by another docking software, the ZDOCK (see Appendix A). Normally, molecular docking is considered successful if one of the top N predictions coincides with the experimentally determined position of the ligand with RMSD within 2–2.5 Å. However, protein-peptide docking is often subjected to a different criterion (other than RMSD) due to the flexibility of the docked peptides and consequently other metrics, such as CAPRI parameters [66], can be used. This was not considered in our analysis in order to unite the approach for both peptide and non-peptide docking. Moreover, docking of flexible and longer ligands (such as peptides) can be quite challenging, especially in cases where the conformational flexibility of the target protein comes into play [16, 67]. In addition, some studies point out the importance of non-rigidity of the receptor [14, 68, 69], although this is still one of the major challenges for the available docking methods. The software used (AutoDock Vina) allows for flexibility of the receptor, but for reasons of comparison with experimental data as well as minimizing the computational cost of molecular docking, the receptor coordinates were frozen.

4. Conclusions

The article describes a machine learning solution for binding region identification that serves as a guide for target molecular docking. The approach consists of two stages: (1) the machine learning algorithm provides the approximate binding region on a receptor and (2) target molecular docking is performed to fine-tune the pose of the ligand within the predicted region. The choice of software for the second stage is fairly arbitrary and allows some degree of freedom with respect to the task (either a ligand is of peptide or non-peptide type).

We show that the binding region determined by the ML algorithm can be used as an initial guess for target molecular docking. The size of the cell should be large enough to ensure that the ligand explores conformations with the highest binding affinity to the receptor. The work includes an analysis that focuses on the size of the cell.

It is shown that the performance of target molecular docking based on ML is comparable to the performance of docking based on the binding region predicted by Accluster when it comes to peptide-protein cases. The performance of molecular docking for peptide ligands remains a difficult task, mainly due to the conformational freedom of long and flexible peptides. More optimistic results have been obtained in cases where small organic molecules have been docked. The performance of the latter, providing a large enough search space around the initial guess, is comparable to target docking, where the initial guess is the experimentally determined position of the ligand.

It is encouraging that the ML model has shown a powerful capability of learning correlations hidden in the structural data to recognize binding sites. Therefore, in the future, it would be beneficial to develop a ML algorithm that takes into account the chemical topology of the ligand in the binding region. Moreover, the training data can be extended with other classes of compounds to generalize the approach and improve the performance.

Supplementary Material

1

Highlights.

  • A two-step protocol enables more efficient computational receptor-ligand docking

  • Machine learning-based pipeline provides an effective way to evaluate binding regions for protein-ligand complexes

  • The algorithm works for both peptide and non-peptide ligands, and also in cases where the binding site is buried

  • When binding region data are not available, the algorithm provides a fast and effective alternative to classical predictors

  • The approach can be generalized by extending the training dataset to include other classes of compounds

Acknowledgement

M.S. and M.L. acknowledge the financial support from the Slovenian Research Agency (research core funding No. P1-0201). M.S. acknowledges support through the “Young Researchers” program of the Slovenian Research Agency. M.L. acknowledges the National Institutes of Health RM1 award “Solvation modeling for next-gen biomolecule simulations” (grant No. RM1-GM135136). M.D. acknowledges financial support from the National Academy of Sciences of Ukraine (grant No. 14.2021.MM).

Appendix A. Docking peptides with ZDOCK and comparison to ML

Here we discuss the combination of our ML approach with the ZDOCK docking software [33]. The web-based ZDOCK platform performs blind docking and does not allow direct integration with other pipelines, so manual data operations limit its use on a larger scale. We performed docking for a dozen cases beyond the PepBDB v.2019 dataset and compared the results to the ML predictions. Some of these examples are shown in Fig. A1, where the receptor is shown with grey ribbons, the experimental position of the ligand is shown with black sticks, ZDOCK’s predictions for the ligands are shown with red and magenta sticks, and the prediction of ML is shown with green spheres. The red sticks indicate the first top (top-1) ranked position according to the ZDOCK scoring function, while the following nine (out of the top-10) predictions are marked with magenta sticks.

Predictions usually form multiple clusters on the surface of the receptor. The top snapshot in Fig. A1 shows the case with two different clusters of ligand positions, neither of which matches the experimental one (black sticks), so the ML prediction does not improve the ranking in this case. The middle snapshot shows a case with three distinct clusters – one cluster correctly guesses the experimental position but does not include the top-1 predicted pose, however, ML (green) guesses the correct binding region. This example shows that the top scored ligand position (red sticks) does not necessarily match the proper site (black sticks). In such cases, ML prediction can help identify the correct site and reprioritize the ranking. In the bottom snapshot, we see the case where both the top ranked position and the ML algorithm guess the same site. With this rather qualitative overview of the results, we wanted to illustrate a scoring problem among a set of predictions typically offered by popular docking services.

Appendix B. Supplementary material

Supplementary data to this article include the calculated RMSDs for the docking of peptide and non-peptide ligands for different cell sizes.

Figure A1:

Figure A1:

Examples of molecular docking by ZDOCK compared to the ML approach. Gray ribbons depict proteins, black sticks the ligands of experimental structures, red and magenta sticks the ZDOCK predictions for ligands, and green spheres the ML predictions for ligands. Ligands in red indicate the top-1 ZDOCK scored positions, the ligands in magenta depict next scored positions. The examples are taken from the Protein Data Bank, their PDB identifiers are: top – 4z8j, middle – 6hsx, bottom – 5djd.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Credit authorship contribution statement

All authors: Conceptualization, Methodology, Validation, Writing - Original Draft, Writing - Review & Editing, Matjaž Simončič: Software, Investigation, Visualization, Miha Lukšič: Resources, Supervision, Maksym Druchok: Resources, Supervision, Project administration, Software, Investigation, Visualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability statement

The data that support the findings of the study are available from the corresponding author upon reasonable request, while the ML code is available at https://github.com/mksmd/NNforDocking.

References

  • [1].Kirkpatrick P, Ellis C, Chemical space, Bioorganic Med. Chem 432 (2004) 823. [Google Scholar]
  • [2].Shoichet B, McGovern S, Wei B, Irwin J, Lead discovery using molecular docking, Curr. Opin. Chem. Biol 6 (2002) 439–446. [DOI] [PubMed] [Google Scholar]
  • [3].Pinzi L, Rastelli G, Molecular docking: Shifting paradigms in drug discovery, Int. J. Mol. Sci 20 (2019) 4331. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].de Azevedo W Jr. (Ed.), Docking screens for drug discovery, Humana Press, New York, NY, 2019. [Google Scholar]
  • [5].Sethi A, Joshi K, Sasikala K, Alvala M, Molecular docking in modern drug discovery: Principles and recent applications, in: Gaitonde V, Karmakar P, Trivedi A(Eds.), Drug Discovery and Development, IntechOpen, 2020. [Google Scholar]
  • [6].Macari G, Toti D, Polticelli F, Computational methods and tools for binding site recognition between proteins and small molecules: from classical geometrical approaches to modern machine learning strategies, J. Comput. Aided Mol. Des 33 (2019) 887–903. [DOI] [PubMed] [Google Scholar]
  • [7].Guvench O, MacKerell A, Comparison of protein force fields for molecular dynamics simulations, Humana Press, Totowa, NJ, 2008, pp. 63–88. [DOI] [PubMed] [Google Scholar]
  • [8].Pissurlenkar R, Shaikh M, Iyer R, Coutinho E, Molecular mechanics force fields and their applications in drug design, Antiinfect. Agents Med. Chem 8 (2009) 128–150. [Google Scholar]
  • [9].Beauchamp K, Lin Y-S, Das R, Pande V, Are protein force fields getting better? a systematic benchmark on 524 diverse NMR measurements, J. Chem. Theory Comput 8 (2012) 1409–1414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Cino E, Choy W-Y, Karttunen M, Comparison of secondary structure formation using 10 different force fields in microsecond molecular dynamics simulations, J. Chem. Theory Comput 8 (2012) 2725–2740. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Wang L, Wu Y, Deng Y, Kim B, Pierce L, Krilov G, Lupyan D, Robinson S, Dahlgren M, Greenwood J, Romero D, Masse C, Knight J, Steinbrecher T, Beuming T, Damm W, Harder E, Sherman W, Brewer M, Wester R, Murcko M, Frye L, Farid R, Lin T, Mobley D, Jorgensen W, Berne B, Friesner R, Abel R, Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field, J. Am. Chem. Soc 137 (2015) 2695–2703. [DOI] [PubMed] [Google Scholar]
  • [12].Lindorff-Larsen K, Maragakis P, Piana S, Eastwood M, Dror R, Shaw D, Systematic validation of protein force fields against experimental data, PLOS ONE 7 (2012) 1–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Salmaso V, Moro S, Bridging molecular docking to molecular dynamics in exploring ligand-protein recognition process: An overview, Front. Pharmacol 9 (2018) 923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Pagadala N, Syed K, Tuszynski J, Software for molecular docking: a review, Biophys. Rev 9 (2017) 91–102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Ciemny M, Kurcinski M, Kamel K, Kolinski A, Alam N, Schueler-Furman O, Kmiecik S, Protein–peptide docking: opportunities and challenges, Drug Discov. Today 23 (2018) 1530–1537. [DOI] [PubMed] [Google Scholar]
  • [16].Lee A, Harris J, Khanna K, Hong J-H, A comprehensive review on current advances in peptide drug development and design, Int. J. Mol. Sci 20 (2019) 2383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Zhao J, Cao Y, Zhang L, Exploring the computational methods for protein-ligand binding site prediction, Comput. Struct. Biotechnol. J 18 (2020) 417–426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Mamoshina P, Vieira A, Putin E, Zhavoronkov A, Applications of deep learning in biomedicine, Mol. Pharm 13 (2016) 1445–1454. [DOI] [PubMed] [Google Scholar]
  • [19].Alam N, Goldstein O, Xia B, Porter K, Kozakov D, Schueler-Furman O, High-resolution global peptide-protein docking using fragments-based PIPER-FlexPepDock, PLoS Comput. Biol 13 (2017) 1–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [20].Iqbal S, Hoque M, PBRpredict-Suite: a suite of models to predict peptide-recognition domain residues from protein sequence, Bioinformatics 34 (2018) 3289–3299. [DOI] [PubMed] [Google Scholar]
  • [21].Johansson-Åkhe I, Mirabello C, Wallner B, Predicting protein-peptide interaction sites using distant protein complexes as structural templates, Sci. Rep 9 (2019) 4267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Roy A, Zhang Y, Recognizing protein-ligand binding sites by global structural alignment and local geometry refinement, Structure 20 (2012) 987–997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].Alekseenko A, Kotelnikov S, Ignatov M, Egbert M, Kholodov Y, Vajda S, Kozakov D, ClusPro LigTBM: Automated template-based small molecule docking, J. Mol. Biol 432 (2020) 3404–3410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].London N, Movshovitz-Attias D, Schueler-Furman O, The structural basis of peptide-protein binding strategies, Structure 18 (2010) 188–199. [DOI] [PubMed] [Google Scholar]
  • [25].Li J, Fu A, Zhang L, An overview of scoring functions used for protein-ligand interactions in molecular docking, Interdiscip. Sci. Comput. Life Sci 11 (2019) 320–328. [DOI] [PubMed] [Google Scholar]
  • [26].Kalyaanamoorthy S, Phoebe Chen Y-P, Structure-based drug design to augment hit discovery, Drug Discov. Today 16 (2011) 831–839. [DOI] [PubMed] [Google Scholar]
  • [27].Santos L, Ferreira R, Caffarena E, Integrating molecular docking and molecular dynamics simulations, Springer New York, New York, NY, 2019, pp. 13–34. [DOI] [PubMed] [Google Scholar]
  • [28].Brylinski M, Feinstein W, efindsite: Improved prediction of ligand binding sites in protein models using meta-threading, machine learning and auxiliary ligands, J. Comput. Aided Mol. Des 27 (2013) 551–567. [DOI] [PubMed] [Google Scholar]
  • [29].Yan C, Zou X, Predicting peptide binding sites on protein surfaces by clustering chemical interactions, J. Computat. Chem 36 (2015) 49–61. [DOI] [PubMed] [Google Scholar]
  • [30].Viet Hung L, Caprari S, Bizai M, Toti D, Polticelli F, Libra: ligand binding site recognition application, Bioinformatics 31 (2015) 4020–4022. [DOI] [PubMed] [Google Scholar]
  • [31].Lavi A, Ngan C, Movshovitz-Attias D, Bohnuud T, Yueh C, Beglov D, Schueler-Furman O, Kozakov D, Detection of peptide-binding sites on protein surfaces: The first step toward the modeling and targeting of peptide-mediated interactions, Proteins 81 (2013) 2096–2105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Saladin A, Rey J, Thévenet P, Zacharias M, Moroy G, Tufféry P, PEP-SiteFinder: a tool for the blind identification of peptide binding sites on protein surfaces, Nucleic Acids Ress 42 (2014) W221–W226. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Pierce B, Wiehe K, Hwang H, Kim B-H, Vreven T, Weng Z, ZDOCK server: interactive docking prediction of protein–protein complexes and symmetric multimers, Bioinformatics 30 (2014) 1771–1773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Raveh B, London N, Schueler-Furman O, Sub-angstrom modeling of complexes between flexible peptides and globular proteins, Proteins 78 (2010) 2029–2040. [DOI] [PubMed] [Google Scholar]
  • [35].London N, Raveh B, Cohen E, Fathi G, Schueler-Furman O, Rosetta FlexPepDock web server-high resolution modeling of peptide-protein interactions, Nucleic Acids Res 39 (2011) W249–W253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Raveh B, London N, Zimmerman L, Schueler-Furman O, Rosetta FlexPepDock ab-initio: Simultaneous folding, docking and refinement of peptides onto their receptors, PLoS One 6 (2011) 1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Grosdidier A, Zoete V, Michielin O, Swissdock, a protein-small molecule docking web service based on eadock dss, Nucleic Acids Res 39 (2011) W270–W277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Lee H, Heo L, Lee M, Seok C, GalaxyPepDock: a protein–peptide docking tool based on interaction similarity and energy optimization, Nucleic Acids Res 43 (2015) W431–W435. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39].Lee H, Zhang Y, Bsp-slim: A blind low-resolution ligand-protein docking approach using predicted protein structures, Proteins 80 (2012) 93–110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Allen W, Balius T, Mukherjee S, Brozell S, Moustakas D, Lang P, Case D, Kuntz I, Rizzo R, Dock 6: Impact of new features and current docking performance, J. Computat. Chem 36 (2015) 1132–1156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [41].Antes I, DynaDock: A new molecular dynamics-based algorithm for protein–peptide docking including receptor flexibility, Proteins 78 (2010) 1084–1104. [DOI] [PubMed] [Google Scholar]
  • [42].Kurcinski M, Ciemny P, Oleniecki T, Kuriata A, Badaczewska-Dawid AE, Kolinski A, Kmiecik S, CABS-dock standalone: a toolbox for flexible protein–peptide docking, Bioinformatics 35 (2019) 4170–4172. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [43].de Vries S, Rey J, Schindler C, Zacharias M, Tuffery P, The pepATTRACT web server for blind, large-scale peptide–protein docking, Nucleic Acids Res 45 (2017) W361–W364. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].van Zundert G, Rodrigues J, Trellet M, Schmitz C, Kastritis P, Karaca E, Melquiond A, van Dijk M, de Vries S, Bonvin A, The HADDOCK2.2 web server: User-friendly integrative modeling of biomolecular complexes, J. Mol. Biol 428 (2016) 720–725. Computation Resources for Molecular Biology. [DOI] [PubMed] [Google Scholar]
  • [45].Trott O, Olson A, AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multi-threading, J. Comput. Chem 31 (2010) 455–461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [46].Kong R, Wang F, Zhang J, Wang F, Chang S, CoDockPP: A multistage approach for global and site-specific protein–protein docking, J. Chem. Inf. Model 59 (2019) 3556–3564. [DOI] [PubMed] [Google Scholar]
  • [47].Korb O, Stützle T, Exner T, PLANTS: Application of ant colony optimization to structure-based drug design, in: Dorigo M, Gambardella L, Birattari M, Martinoli A, Poli R, Stützle T(Eds.), Ant colony optimization and swarm intelligence, Springer Berlin Heidelberg, Berlin, Heidelberg, 2006, pp. 247–258. [Google Scholar]
  • [48].Schueler-Furman O, London N (Eds.), Modeling peptide-protein interactions: Methods and protocols, Humana Press, New York, NY, 2017. [Google Scholar]
  • [49].Agrawal P, Singh H, Srivastava H, Singh S, Kishore G, Raghava G, Benchmarking of different molecular docking methods for protein-peptide docking, BMC Bioinformatics 19 (2019) 426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Rentzsch R, Renard B, Docking small peptides remains a great challenge: an assessment using AutoDock Vina, Brief. Bioinform 16 (2015) 1045–1056. [DOI] [PubMed] [Google Scholar]
  • [51].Hwang S, Lee C, Lee S, Ma S, Kang Y-M, Cho K, Kim S-Y, Kwon O, Yoon C, Kang Y, Yoon J, Nam K-Y, Kim S-G, In Y, Chai H, Acree W, Grant J, Gibson K, Jhon M, Scheraga H, No K, PMFF: Development of a physics-based molecular force field for protein simulation and ligand docking, J. Phys. Chem. B 124 (2020) 974–989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [52].Chupakhin V, Marcou G, Baskin I, Varnek A, Rognan D, Predicting ligand binding modes from neural networks trained on protein–ligand interaction fingerprints, J. Chem. Inf. Model 53 (2013) 763–772. [DOI] [PubMed] [Google Scholar]
  • [53].Ragoza M, Hochuli J, Idrobo E, Sunseri J, Koes D, Protein–ligand scoring with convolutional neural networks, J. Chem. Inf. Model 57 (2017) 942–957. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [54].Taherzadeh G, Zhou Y, Liew A-C, Yang Y, Structure-based prediction of protein–peptide binding regions using Random Forest, Bioinformatics 34 (2017) 477–484. [DOI] [PubMed] [Google Scholar]
  • [55].de Magalhães C, Almeida D, Barbosa H, Dardenne L, A dynamic niching genetic algorithm strategy for docking highly flexible ligands, Inf. Sci 289 (2014) 206–224. [Google Scholar]
  • [56].Chen J-Q, Chen H-Y, Dai W.-j., Lv Q-J, Chen C-C, Artificial intelligence approach to find lead compounds for treating tumors, J. Phys. Chem. Lett 10 (2019) 4382–4400. [DOI] [PubMed] [Google Scholar]
  • [57].Gao K, Nguyen D, Chen J, Wang R, Wei G-W, Repositioning of 8565 existing drugs for COVID-19, J. Phys. Chem. Lett 11 (2020) 5373–5382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [58].Nand M, Maiti P, Joshi T, Chandra S, Pande V, Kuniyal J, Ramakrishnan M, Virtual screening of anti-HIV1 compounds against SARS-CoV-2: machine learning modeling, chemoinformatics and molecular dynamics simulation based analysis, Sci. Rep 10 (2020) 20397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [59].Batra R, Chan H, Kamath G, Ramprasad R, Cherukara M, Sankaranarayanan S, Screening of therapeutic agents for COVID-19 using machine learning and ensemble docking studies, J. Phys. Chem. Lett 11 (2020) 7058–7065. [DOI] [PubMed] [Google Scholar]
  • [60].Santana M, Silva-Jr F, De novo design and bioactivity prediction of SARS-CoV-2 main protease inhibitors using recurrent neural network-based transfer learning, BMC Chemistry 15 (2021) 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [61].Kadioglu O, Saeed M, Greten H, Efferth T, Identification of novel compounds against three targets of SARS CoV-2 coronavirus by combined virtual screening and supervised machine learning, Comput. Biol. Med 133 (2021) 104359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Banchi L, Fingerhuth M, Babej T, Ing C, Arrazola J, Molecular docking with Gaussian Boson Sampling, Sci. Adv 6 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [63].Wen Z, He J, Tao H, Huang S-Y, PepBDB: a comprehensive structural database of biological peptide–protein interactions, Bioinformatics 35 (2019) 175–177. [DOI] [PubMed] [Google Scholar]
  • [64].Hornak V, Abel R, Okur A, Strockbine B, Roitberg A, Simmerling C, Comparison of multiple amber force fields and development of improved protein backbone parameters, Proteins 65 (2006) 712–725. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [65].Krieger E, Vriend G, YASARA View – molecular graphics for all devices – from smartphones to workstations, Bioinformatics 30 (2014) 2981–2982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [66].Janin J, Henrick K, Moult J, Eyck L, Sternberg M, Vajda S, Vakser I, Wodak S, CAPRI: a critical assessment of predicted interactions, Proteins 52 (2003) 2–9. [DOI] [PubMed] [Google Scholar]
  • [67].Buonfiglio R, Recanatini M, Masetti M, Protein flexibility in drug discovery: From theory to computation, ChemMedChem 10 (2015) 1141–1148. [DOI] [PubMed] [Google Scholar]
  • [68].Frimurer T, Peters G, Iversen L, Andersen H, Møller N, Olsen O, Ligand-induced conformational changes: Improved predictions of ligand binding conformations and affinities, Biophys. J 84 (2003) 2273–2281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [69].Mobley D, Dill K, Binding of small-molecule ligands to proteins: “what you see” is not always “what you get”, Structure 17 (2009) 489–498. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

The data that support the findings of the study are available from the corresponding author upon reasonable request, while the ML code is available at https://github.com/mksmd/NNforDocking.

RESOURCES