Abstract
Protein kinases are vital drug targets, yet designing selective inhibitors is challenging, compounded by resistance and kinome complexity. This review explores Quantitative Structure-Activity Relationship (QSAR) modeling for kinase drug discovery, focusing on integrating traditional QSAR with machine learning (ML)—CNNs, RNNs—and structural data. Methods include structural databases, docking, and deep learning QSAR. Key findings show ML-integrated QSAR significantly improves selective inhibitor design for CDKs, JAKs, PIM kinases. The IDG-DREAM challenge exemplifies ML’s potential for accurate kinase-inhibitor interaction prediction, outperforming traditional methods and enabling inhibitors with enhanced selectivity, efficacy, and resistance mitigation. QSAR combined with advanced computation and experimental data accelerates kinase drug discovery, offering transformative precision medicine potential. This review highlights deep learning-enhanced QSAR’s novelty in automating feature extraction and capturing complex relationships, surpassing traditional QSAR, while emphasizing interpretability and experimental validation for clinical translation.
Keywords: QSAR, kinase, FDA approved, catalytic domain, discovery, machine learning, deep QSAR, CNN
ARTICLE HIGHLIGHTS
Kinase biology and therapeutics
Kinases regulate cellular functions and cell cycle progression, with their dysregulation linked to diseases like cancer and autoimmune disorders.
Understanding kinase structure and binding pockets have facilitated drug discovery efforts, leading to over 80 FDA-approved inhibitors.
Quantitative structure-activity relationship (QSAR) in drug design
QSAR modeling provides a framework to predict biological activity based on molecular features.
Traditional QSAR methods, including CoMFA, CoMSIA, and 3D-QSAR, have been pivotal in optimizing kinase inhibitors.
The role of machine learning in QSAR
Machine learning techniques such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) have revolutionized QSAR, uncovering complex molecular relationships.
Deep QSAR enhances predictive power, facilitating the identification of active kinase inhibitors from vast virtual libraries.
Case studies and applications
QSAR has been involved in understanding ligand-target interactions in kinases like CDKs, JAKs, and PIM kinases.
Future directions
Integrating QSAR with machine learning and experimental approaches, alongside collaborative efforts between computational and medicinal chemists, holds great promise for advancing precision medicine and developing next-generation kinase inhibitors.
1. Introduction
Protein kinases play a crucial role in modifying cellular functions at various levels. The process of protein phosphorylation is described as a reversible mechanism that involves the coordinated activity of phosphatases and kinases in a cyclic process known as the phosphorylation-dephosphorylation cycle. This intricate interplay highlights the significance of protein kinases as essential regulatory elements within cellular biochemical pathways. The active state of protein kinases can either enhance or diminish specific enzymatic actions, thereby finely tuning various biological and cellular activities.
In the last two decades, there has been a notable advancement in the molecular and structural understanding of human kinases, leading both academic and pharmaceutical research community to increasingly appreciate the value of research studies on kinase enzymes. Consequently, the efforts of medicinal chemists in the pharmaceutical industry culminated in the endorsement of 80 FDA-approved protein kinase small molecule inhibitors as of November 2023 (please refer to Tables 1S, 2S, and 3S under supplementary material). This progress is facilitated by four main factors: firstly, the intensified gene sequencing efforts that have enriched our knowledge about the human kinome. The human kinase gene family comprises almost 518 members, among which 106 are pseudogenes [1]. The second factor encompasses the implementation of modern drug discovery techniques and the improvement of the artificial intelligence procedures, including the advancement of computer-aided drug activity profiling panels, QSAR studies, machine learning drug discovery, and the molecular databases screening [2]. These advancements enable scientists to identify and eliminate promiscuous compounds during the early stages of the drug design process, thus aiding in minimizing off-target effects in later stages. The third factor relates to the increased availability of high-resolution ligand-kinase structural crystals, which are freely accessible in the Protein Data Bank (PDB). This availability has facilitated the development of computational drug design technology and provided a better understanding of the kinase binding pocket [3]. Finally, the improved understanding of the clinical applications of kinase inhibitors and the mechanisms underlying molecular drug resistance has proven to be pivotal. The challenge of kinase inhibitor resistance in cancer treatment remains a significant concern, compelling the ongoing development of new generations of kinase inhibitors to counteract the effects of mutation-linked drug resistance [3,4].
The application of computer-aided QSAR studies for treating dysregulated protein kinase enzymatic activity was not officially reported until the approval of Imatinib in the early 2000s. Interestingly, we found a review issue published in “Trends in Pharmacological Sciences” Vol. 23 No. 3 in March 2002 named “On the long road to drug discovery” [5,6]. This issue was giving a comprehensive overview on actual trends in drug discovery in the early times of the 21st century. Crucial points were discussed in the chapters of this issue, most importantly Cyclin-Dependent Kinase (CDK) inhibitory molecules. The issue was highlighting the future progress in drug research at that time and its chapters focused on the structure–activity relationships of CDK inhibitory drugs. In 2002, the chapters were not easily comprehended by the biologists and pharmacologists. However, the issue introduced a very clear insight to the importance of these fields of drug research [5]. Moreover, the issue described the new scientific findings that are related to the function of CDKs in cell cycle regulation and roughly proposed clinical applications for the new chemical inhibitors of CDK (e.g., flavones) in cancer therapy.
The increasing reliance on machine learning QSAR in kinase drug discovery is further validated by large-scale community initiatives such as the IDG-DREAM Drug-Kinase Binding Prediction Challenge (detailed in Supplementary Material Section 1), which rigorously benchmarked and assessed the performance of numerous predictive models developed by research groups worldwide.
This review explores the discovery of approved and investigational therapies targeting kinases, with a particular focus on advancements in QSAR studies addressing dysregulated kinase enzymes. Well-structured QSAR algorithms enable researchers to uncover complex relationships between the molecular structures of anti-kinase agents, leveraging extensive experimental datasets to develop predictive models that serve as powerful tools in the drug design process. Additionally, this review emphasizes the critical importance of comparative evaluations of different QSAR models, highlighting methods with superior predictive power, computational efficiency, and practical applicability. These insights provide a comprehensive understanding of QSAR performance in advancing kinase-targeted drug discovery (Figure 3).
Figure 3.
Phases of traditional drug design and development process.
2. Eukaryotic protein kinases (components, binding pocket, catalytic domain, and types of inhibitors)
2.1. Eukaryotic protein kinase component
Eukaryotic kinases typically exist in a non-active state and become activated only upon being stimulated by a regulatory signal [7]. For instance, Receptor Protein-Tyrosine Kinases (RTKs) are triggered by ligand-induced dimerization. Additionally, Extracellular Signal-Regulated Kinase (ERK) is activated by another kinase known as Mitogen-Activated Protein Kinase (MEK), Cyclin-Dependent Kinases (CDKs) are also activated by their respective cyclins, and the Calcium/Calmodulin-Dependent Protein Kinase is stimulated by the Calcium-Calmodulin complex [8]. Consequently, dysregulation in the kinase activation pathway can lead to human cell cycle abnormalities such as malignancies and inflammatory diseases [8].
Moreover, kinases follow a hierarchical organization into groups, families, and sub-families. Their classification primarily hinges on the amino acid sequences within their catalytic domains, alongside their biological functions [8]. Predominantly protein kinases, comprising nine categories of human kinase groups (please refer to Figure 1). These categories are divided into distinct groups. The AGC group (63 members) includes Protein Kinases A, G, and C (PKA, PKG, PKC), Akt kinases (PKB1–3), Aurora kinases (Aur1–3), and others like PDK1 and RSK1–4. The CAMK group (74 members) features Calcium/Calmodulin-Dependent Kinases (CAMK1, CAMK2, CAMK4), Phosphorylase Kinases (PhKγ1, PhKγ2), Mitogen-Activated Kinase Activators (MAPK2, MAPK3, MAPK5), and Myosin Light Chain Kinases. The CK1 group (12 members) includes Casein Kinases (CK1), Tau Tubulin Kinases, and Vaccinia-Related Kinases. The CMGC group (61 members) comprises Cyclin-Dependent Kinases (CDK1–11), MAPKs, ERK1–5, and GSK3. Other groups include the STE group, activating MAPK cascades, Tyrosyl Kinase (90 members including RTKs and NRTKs), TKL group (43 members), and Receptor Guanylyl Cyclase (5 members). Additionally, 83 atypical kinases are distinct from these families (please refer to Figure 1).
Figure 1.
Eukaryotic protein kinase component.
2.2. General structure of protein kinases
In general, the catalytic domain of protein kinases consists of a small N-terminal lobe and a large C-terminal lobe, in addition to the active catalytic site that connects the two lobes [9].
The small N-terminal lobe is composed of five anti-parallel β-sheet (β1–β5). The β1 and β2 strands are connected by the P-loop, also known as the Glycine-Rich Loop (GRL). Following this, the αC-helix area occurs in two forms: the active form and the inactive form [8]. Although the αC-helix occurs within the N-terminal lobe, it holds a special position between the small N-terminal lobe and the large C-terminal lobe. Moreover, the P-loop is situated in the N-terminal vicinity and includes the GxGxΦG motif, where “G” represents the glycine amino acid, and “Φ” stands for the hydrophobic areas.
Additionally, the β1- and β2-strands interact with the adenine part of ATP. A valine amino acid appears two residues after the GRL loop within the β2-strand. The valine residue is essential for engaging in hydrophobic interactions either with the adenine portion of ATP bound to the kinase or with the kinase inhibitor positioned within the ATP binding pocket. Protein kinases also contain a lysine sequence within the β3-sheet and a conserved glutamate near the center of the αC-helix, forming a salt bridge that connects the positively charged lysine from the β3-strand with the negatively charged glutamate from the αC-helix [1]. Also, the lysine in this sequence functions by linking the α- and β-phosphates of ATP to the αC-helix [10]. The presence of the salt bridge linking the β3-lysine and the αC-glutamate is vital for establishing the active state of the protein kinase enzyme, typically associated with the “αC-in” conformation. This conformation is indispensable for achieving optimal kinase activity [10].
As mentioned earlier, the general structure of protein kinases also includes the large C-terminal lobe, which is composed of eight α-helical structures (αD, αE, αF, αG, αH, αI, αEF1, and αEF2), in addition to four extra short β-sheets (β6, β7, β8, and β9). It’s worth noting that the ATP binding pocket is part of this large lobe, and the adenine part of ATP interacts with the second residue of the β7-strand [1]. Moreover, the C-terminal lobe harbors a flexible polypeptide chain named the catalytic loop, which facilitates the transfer of a phosphate group from ATP to the serine, threonine, or tyrosine residue on the enzyme substrate. The C-terminal lobe functions by positioning the enzyme substrate into the correct active site to enable catalysis. The middle portion of the active catalytic site of protein kinase families varies with their amino acid sequence. The active catalytic site of the kinase family has residues that can be modified by adding a phosphate group (phosphorylation process), which is required for the initiation of maximal enzymatic activity. In many protein kinases, this activation segment initiates with DFG (Asp-Phe-Gly) sequence and concludes with APE (Ala-Pro-Glu) sequence. Next to the activation segment occurs the conserved catalytic loop segment (His-Arg-Asp (HRD)) sequence and the amino-terminus of the αC-helix [1,11]. The gatekeeper residue is the amino acid positioned near the N6 amine of ATP in all kinases, and it aids in determining the selectivity of the kinase inhibitor, as it can vary in size from a large amino acid like Phe to a small residue like Thr [12].
The kinase catalytic binding pocket can adopt different conformations, most importantly the DFG-IN conformation, in which the DFG (Asp-Phe-Gly) triplet occupies an “in” conformation capable of complexing a Mg2+ atom during the phosphate transfer step. On the other hand, a DFG-OUT conformation in kinases is one in which the DFG triplet is incapable of complexing a Mg2+ atom during the phosphate transfer step [10,12].
2.3. Kinase catalytic domains
The structural and functional critical residues within the active conformation of protein kinases entail two spines: The R regulatory spine and the C catalytic spine. The R regulatory spine works by positioning the substrate (recipient protein or enzyme), and the C-spine works by positioning the ATP for catalysis. The two spines occur within the small N-terminal and the large C-terminal lobes. Moreover, the R-spine is composed of four amino acids, and the C-spine is composed of eight amino acids, and both spines are important to generate a catalytically flexible active joint. Regarding the C catalytic spine, two of its eight amino acids belong to the small N-terminal lobe, while the other six amino acids belong to the large C-terminal lobe [1,11].
The hinge region connects the N-terminal and C-terminal lobes. When ATP binds to the catalytic domain through special hydrogen bond interactions, it unites the N-terminal and C-terminal lobes. This is illustrated by the binding of the 6-amino N–H group of the adenine group within the ATP structure with the backbone carbonyl group of the first hinge residue, initiating the transfer of the phosphate group from ATP to the recipient protein or enzyme [1,13]. Additionally, the N1 of the adenine group of ATP acts as a hydrogen-bond acceptor and forms a hydrogen bond with the backbone N–H group of the third hinge residue. This pattern of dual hydrogen bonding in the hinge region is important for catalysis and the potent inhibition of kinases by small molecule inhibitors. Additional hydrogen bonds to the kinase hinge seem to affect only kinase selectivity [13–15].
2.4. Types of kinase inhibitors
Kinase inhibitors exhibit diverse binding mechanisms, leading to their classification into several types (I-VI) [12]. These types differ primarily in their binding site (ATP pocket, allosteric site, or both), conformational preferences induced in the kinase, and overall mechanism of inhibition. For instance, Type I inhibitors target the ATP pocket, while Type II inhibitors can occupy both the ATP pocket and an adjacent allosteric site. Other types, like Type III, bind exclusively to allosteric sites, and further subtypes and types with bivalent or covalent mechanisms exist. A more detailed description of these kinase inhibitor types, including specific examples and references, can be found in Supplementary Material Section 2. Please refer to Tables 1S, 2S, 3S, and 4S under supplementary material for a comprehensive list of 80 FDA-approved kinase drugs and their inhibitor types as of November 2023.
3. A Brief overview of the history of FDA-approved kinase enzyme modulators
3.1. The emergence of kinase-directed therapeutics
The exploration of kinases as crucial regulators in cell cycle processes began in the late 1970s. A pivotal discovery was made in 1977, when Japanese researchers first identified protein kinase C (PKC) (Inoue et al.) [16]. By 1980, the foundational role of epidermal growth factor (EGF) in phosphorylating tyrosine residues on plasma membrane receptors was elucidated by Ushiro and Cohen [17]. This phosphorylation event was recognized as the trigger for a cascade of biochemical processes leading to cell proliferation.
During the mid-1980s, the scientific community viewed kinase binding domains as highly conserved, posing challenges for the development of potent and selective kinase inhibitors. Nevertheless, significant progress was achieved during this period, including the isolation of epidermal growth factor receptor (EGFR) and insulin receptor kinase [18,19]. In 1986, Staurosporine’s inhibitory effect on PKC was confirmed [20], and by 1988, the first sub-micromolar competitive inhibitors for EGFR and insulin receptor kinases were developed [21]. These early inhibitors paved the way for structure-activity relationship (SAR) studies, which began targeting EGFR and insulin receptor kinases.
In 1994, the first selective ATP-mimetic EGFR kinase inhibitor was identified [22,23], marking a critical milestone in kinase inhibitor development. These inhibitors exhibited selectivity for insulin receptor kinase while demonstrating minimal activity against serine/threonine kinases, thereby enabling the generation of preliminary SAR data [21]. This era laid the foundation for advancing kinase-focused drug discovery through computational approaches such as quantitative structure-activity relationship (QSAR) studies (Figures 2 and 3).
Figure 2.
Timeline (1968–2024) showing the simultaneous development of both the science of kinase inhibition and modulation and the computer-aided drug design computational studies including the development of QSAR investigational studies. The development of both sciences simultaneously culminated in 80 USFDA-approved small molecule kinase modulators as of November 2023. Pink boxes represent the development of computational sciences and QSAR studies, the orange boxes represent the development of the kinase modulation science, and the green-boxed represent the USFDA-approved small molecules kinase inhibitors.
It is important to note that kinase inhibitors often target signaling pathways, not just single kinases. For example, the RAF-MEK-ERK pathway is a critical cascade in cancer, and inhibitors have been developed targeting RAF, MEK, and ERK kinases within this pathway. Targeting pathways offers opportunities to modulate signaling more effectively, address pathway redundancy, and potentially develop combination therapies that target multiple points in a pathway for enhanced efficacy or to overcome resistance mechanisms. Consideration of these broader pathway contexts is crucial in kinase drug discovery beyond simply targeting individual kinase enzymes.
3.2. Imatinib: a breakthrough in cancer therapy
The approval of Imatinib (STI571, brand name Gleevec®) by the US Food and Drug Administration (FDA) in 2001 heralded the era of small-molecule protein kinase inhibitors. Imatinib, initially developed as a platelet-derived growth factor receptor (PDGFR) kinase inhibitor, was later found to target BCR-Abl and is now employed in treating multiple malignancies, including Philadelphia chromosome-positive (Ph+) acute lymphoblastic leukemia (ALL), chronic myeloid leukemia (CML), and gastrointestinal stromal tumors (GIST) [3,24].
As a polypharmacological drug, Imatinib demonstrates inhibitory activity against multiple kinase targets at clinically achievable concentrations. This versatility underscores its transformative role in oncology.
3.3. Generational advances in BCR-abl inhibitors
Despite the success of Imatinib, resistance caused by mutations in the BCR-Abl kinase domain spurred the development of second-generation inhibitors, including Nilotinib (AMN107, brand name Tasigna®, approved in 2007) and Dasatinib (BMS-354825, brand name Sprycel®, approved in 2006) [11,22]. Nilotinib offers potent activity against most resistant mutations but lacks efficacy against the gatekeeper mutation T315I, where threonine is replaced by isoleucine, preventing binding to the allosteric pocket.
The approval of third-generation inhibitors such as Ponatinib (AP 24534, brand name Iclusig®) in 2012 addressed this limitation, providing effective treatment for the T315I mutation [25]. More recently, Asciminib (brand name Scemblix®), a type IV allosteric inhibitor, has expanded therapeutic options as a novel BCR-Abl-targeted agent [26].
3.4. Non-receptor tyrosine kinase: Janus kinase (JAK) inhibitors
The first oral JAK inhibitor, Ruxolitinib (INCB-018424, brand name Jakafi®), gained FDA approval in 2011 for the treatment of myeloproliferative disorders such as myelofibrosis [27]. Subsequently, Tofacitinib (CP-690550, brand name Xeljanz®), a JAK3-selective inhibitor, was approved in 2012 for active rheumatoid arthritis [11].
While Ruxolitinib advanced treatment options for several conditions, its hematological side effects and suboptimal responses in some cases necessitated further research. This led to the approval of Fedratinib (TG101348, brand name Inrebic®) in 2019 as a second-line treatment for high-risk myelofibrosis [1]. Recent approvals, including Abrocitinib for atopic dermatitis and Momelotinib for myelofibrosis-associated anemia, highlight ongoing innovation in JAK inhibitor development [23,24,28].
3.5. Advances in receptor tyrosine kinase (RTK) inhibitors
The ErbB family of RTKs, including EGFR (ErbB1) and HER2 (ErbB2), has been a major focus of kinase inhibitor development. In 2003, AstraZeneca’s Gefitinib (ZD1839, brand name Iressa®) became the first FDA-approved EGFR inhibitor, followed by Erlotinib (OSI-774, brand name Tarceva®) in 2004 [9,10,16].
Notably, third-generation inhibitors such as Osimertinib (brand name Tagrisso®), approved in 2017, have transformed treatment paradigms for EGFR-mutant lung cancer [29]. With 80 FDA-approved kinase inhibitors as of 2023, the ErbB/EGFR family remains the most extensively targeted category [11,30].
4. General overview of quantitative structure-activity relationship (QSAR)
4.1. The establishment of the traditional QSAR modelling technique
The beginning of the machine learning era can be traced back to 1968 with the publication of the book “Cybernetics and Forecasting” by Ivakhnenko and Lapa, wherein these scientists investigated modern analytical and computational methods in science [16]. A leading-edge moment occurred 10 years later when Ivakhnenko himself published a revolutionary article in 1978 titled “The group method of data handling in long-range forecasting” [17]. The article focused on the introduction of the “principle of self-organization” in computer modeling. Ivakhnenko proposed that, in applying the principle of self-organization; a researcher could enhance a small amount of statistical data to the computer alongside specific selection criteria for the model. The computer, using this approach, would then systematically analyze several models that met the specified statistical criteria. Through this process, the computer would identify a successful model (a mathematical equation) of optimal complexity. Consequently, the researcher would attain the best model capable of delivering the most compelling predictions. This marked the outset of QSAR studies [2].
From 1980 till 2010, a significant advancement occurred in computing technology, cheminformatics, and the expansion of molecular databases. These developments, rooting through the foundational work in the 1960s and 1970s, paved the way for the emergence of computer-aided modeling. The progress during this period laid the groundwork for a transformative shift in the landscape of machine learning computational chemistry, enabling more sophisticated and refined modeling approaches including the QSAR modeling [2].
Mainly, computer-aided QSAR aims toward reveling a mathematical relationship that can describe an observed bioactivity (IC50, Ki, and ligand efficiency values) for a group of active compounds against a specific target as a function of structural features and meaningful descriptors [18].
The ligand-based computer-aided QSAR modeling process begins the collection and depiction of structures from previously known active molecules available in literature databases. The collected compounds and their activities are forming a preliminary dataset, serving as a foundation for an effective QSAR modeling project. Medicinal chemists and pharmacologists who are performing the QSAR modeling explore and calculate the physicochemical properties of each collected molecule. It is imperative for the QSAR modeling procedure to be accompanied with statistical evaluation, achieved by splitting the collected dataset into training and test sets for internal and external statistical evaluation respectively [19]. Please refer to Figure 4 which represent the flow diagram showing the investigational steps during the QSAR model development process.
Figure 4.
Flow diagram showing the investigational steps during the QSAR model development process.
The quantitative portrayal of a molecule or a drug property is termed as a descriptor. A descriptor can be a physicochemical or pharmacophoric property. These descriptors are essential in the context of formulating any new QSAR equation. Currently, wide ranges of computable molecular descriptors that are derived from various resources are available. As such, a descriptor can encompass constitutional, electronic, geometrical, hydrophobic, steric, quantum mechanical, or topological properties (please refer to Table 1). Additionally, different applications and softwares can calculate diverse spatial, electronic, topological, and other descriptors. For calculating electronic and quantum mechanical descriptors, Density Functional Theory (DFT) is a significant method. DFT, a quantum mechanical approach, allows for the computation of molecular electronic properties, such as charge distribution, ionization potential, and electron affinity. These DFT-derived electronic descriptors can be particularly valuable in 2D-QSAR models, capturing crucial aspects of molecular reactivity and interactions [31].
Table 1.
Quantitative structure-activity relationship (QSAR) descriptors.
| Type of descriptor | Meaning | Examples |
|---|---|---|
| Fragment constants | Changes in structure are likely to produce similar changes in reactivity, ionization, and binding. Fragment constants explain the effect of different types of substituents on reactivity, ionization, and binding. |
|
| Conformational energy related descriptors | Energy of selected conformation for the studied compound. |
|
| Electronic descriptors | Used to describe electronic aspects of the molecule or atoms bonds and moleculer fragments, e.g, Dipole moment, HOMO, LUMO energy, etc. |
|
| Receptor Surface Analysis (RSA) descriptors | Receptor Surface Analysis (RSA) is useful in QSAR model building when the receptor surface is known. The RSA approach clearly differs from pharmacophore approaches, as it captures information about the receptor instead of the ligands. From receptor surface models, one can derive descriptors, which provide 3D information about the (steric or electrostatic) interaction energies between each point of the receptor surface and the ligand. RSA descriptors can be combined with other 3D or 2D descriptors for QSAR analysis. |
|
| Quantum mechanical descriptors | MOPAC descriptors are calculated using a semi-empirical method that is likely to generate values that are more accurate. The following are examples of the MOPAC descriptors. |
|
| Graph-theoretic descriptors Topological descriptors |
Topological indices are 2D descriptors that are calculated based on graph theory concepts (Kier and Hall 1976, 1986; Katritzky and Gordeeva 1993). Topological descriptors differentiate various molecules according to their shape, size, degree of branching, and flexibility. The electrotopological state indices are numerical values computed for each atom in a molecule and which encode information about both the topological environment of that atom and the electronic interactions due to all other atoms in the molecule. |
|
| Graph-theoretic descriptors Information-content descriptors |
Descriptors that view molecule graphs as sources of certain probability distributions to which Shannons statistical information theory tools can be applied. |
|
| Molecular Shape Analysis (MSA) | Molecular Shape Analysis (MSA) is a technique in QSAR analysis, which combines the molecular shape similarity and commonality measures to determine the similarities between molecules. Molecular shape similarity is applied to the comparison of 3D molecular shapes, which are represented by atomic properties. On the other hand, molecular shape commonality is using conformational energy and molecular shape together to measure molecular similarity. |
The basic concept of MSA in QSAR analysis is that the shape of the molecule is related to the binding site cavity (or pocket), thus it is related to biological activity as well.
|
| Shape descriptors | Molecular shape descriptors has a large scientific literature and in the past years several shape descriptors were developed. The basic principle is the projection of the molecular surface onto three mutually perpendicular planes XY, XZ, and YZ. The descriptors encode the conformation and also the orientation of the molecule. Rotational invariance is obtained by the previous alignment of the X, Y, and Z axes along the three axes of principal inertia. |
|
| Thermodynamic descriptors | Thermodynamics descriptors are used to relate chemical structure to observed chemical behavior. |
|
| Molecular Field Analysis (MFA) descriptors | A 3D rectangular grid can represent the molecular field. MFA analysis is based on the calculation of interaction energies (steric and electrostatic interactions in the case of CoMFA) between some probes (H+ or CH3) and the molecule, represented by a rectangular grid. Thus the field of molecules can be described by MFA grids, and the energies associated with MFA grid points may serve as inputs (descriptors) for the calculation of QSAR models. |
|
| Structural descriptors | Constitutional descriptors are simple, commonly used descriptors reflecting the molecular composition of a compound without any information about its topology. The most common constitution descriptors are number of atoms, bond count, atom type, ring count, and molecular weight (MW). These descriptors are inert to any conformation change and, thus, do not distinguish among isomers. |
|
| ADME descriptors | ADMET properties, that are, absorption, distribution, metabolism, excretion, and toxicity. Virtual screening is able to find similar molecules from large databases. Finding a “patentable” compound with a desired property value is an important aim to be pointed out. |
|
Detailed breakdown of descriptors used in QSAR modeling, including fragment constants, conformational energy-related descriptors, electronic descriptors, receptor surface analysis, and more, with examples for each type.
Upon conducting the QSAR model development, several multivariate analysis methods are employed. These include, for example, the multiple linear regression method, the principal component analysis method, the neural network method, the support vector machine method, and the K-nearest neighbor method. Subsequently, the successful equation can serve as a template to predict the activity of new hits against the studied target [18].
Generally, the QSAR methods can be classified into different types based on how their descriptors are produced and calculated. QSAR methods are:
Zero-dimensional (0D)-QSAR models: these are the simplest QSAR methods. Upon performing the 0D-QSAR, the descriptors are developed based on the molecular structure and formula, e.g., molecular weight, atom number and type, sum of atomic properties, number of rotatable bonds, number of hydrogen-bond acceptors, number of hydrogen-bond donors.
1D-QSAR models: these models establish a quantitative activity/property relationship based on the descriptors that are related molecular properties, e.g., solubility, hydrophilicity, hydrophobicity, etc.
2D-QSAR models link activity data with structural patterns and indices, e.g., topological state indices, flexibility indices, Kier & Hall subgraph count index (SC), Kier shape indices, Kier alpha-modified shape indices, Molecular Flexibility index (f), and Balaban indices (JX and JY), etc.
3D-QSAR equations use grid-based descriptors that are derived from molecular interaction fields. Moreover, they are modulated based on the position, orientation, and conformation of training molecules in the three-dimensional space, e.g., the Comparative Molecular Field Analysis (CoMFA) and Comparative Molecular Similarity Indices Analysis (CoMSIA) methods [20]. MFA is based on the calculation of electrostatic and steric interaction energies between selected probes (H+ or CH3) and the selected ligand molecules, in a three-dimensional lattice. Thus, the grid points serve as descriptors for generating 3D-QSAR models [6,20].
4D-QSAR models are also grid-based descriptors, but they are not derived from molecular interaction fields. 4D-QSAR represents ensemble sampling or conformational flexibility, which is defined by the conformational ensemble profile (CEP). The CEP is calculated with molecular dynamics (MD) simulations. Upon performing a 4D-QSAR each molecule in the training set is represented in all possible conformational, stereoisomeric, and protonation states.
5D-QSAR models signify a higher level of representation of the 4D-QSAR models and allow the induced-fit representation of them.
6D-QSAR models allow the incorporation of the H2O solvation models in 5D-QSAR.
It is important to note that a QSAR equation represent a multivariate mathematical relationship between the biological activity (IC50, Ki, and ligand efficiency values) and a set of descriptors. Analyzing the statistically successful QSAR equation helps in comprehending the connection between the chemical structure of a particular drug and its biological activity and so the QSAR tool in the drug discovery and computer-aided applications is regarded as an expletory tool, offering insights into the productive structure-activity relationships among active ligands within their binding pockets (Table 2).
Table 2.
USFDA-approved small molecules kinase inhibitors discovered by using the computer-aided drug design (CADD) approach.
Includes a list of kinase inhibitors, their approval years, associated company labels, and USFDA drug label links.
4.2. Introduction to machine and deep learning QSAR modelling technique
Unlike traditional QSAR methods, which rely on manual feature extraction and descriptor calculation, machine learning QSAR leverages the power of neural networks to automate these crucial steps. This automation has led to the development of more robust and scalable models, especially vital for handling the complexities of modern molecular data.
Deep learning QSAR modeling represents a significant evolution in QSAR, aligning with the advancements in artificial intelligence (AI) [21]. A key innovation is the use of molecular embeddings. These are sophisticated mathematical representations of molecules designed to capture intricate and often non-linear relationships within molecular data. Molecular embeddings enable more efficient and nuanced analysis and prediction compared to traditional descriptors used in QSAR. Importantly, deep learning QSAR integrates molecular representation directly into the model training process [2,24], contrasting sharply with traditional QSAR where descriptor generation is a separate, pre-modeling step.
It’s important to reiterate the distinction between machine learning and deep learning. Machine learning is a broad field encompassing techniques for creating algorithms and predictive models from data. Deep learning is a subfield of machine learning that specifically utilizes neural networks with multiple layers [2,22,24]. Traditional machine learning methods include clustering, regression, outlier detection, linear regression, and support vector machines [21].
A fundamental shift in deep learning QSAR is the integrated approach to molecular representation. As mentioned, traditional QSAR treats descriptor calculation and model training as separate stages. In deep QSAR, however, molecular representations, often in the form of embeddings, are directly incorporated into the neural network’s optimization process. Neural networks themselves are computational models inspired by the human brain, capable of learning intricate patterns and knowledge from raw data and making predictions based on this learned understanding [24].
Unlike traditional QSAR, which relies on pre-defined, standard chemical descriptors, machine learning algorithms can autonomously identify and extract salient features directly from data through their multi-layered architecture. This hierarchical feature extraction, often described as capturing a “butterfly effect” where higher-level features are derived from lower-level outputs, is a key advantage. Deep learning architectures like Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) are designed to discern complex patterns and connections with minimal human intervention. This inherent capability makes deep learning exceptionally valuable for drug discovery, particularly when integrated with QSAR modeling techniques [24,32].
4.3. Advantages of machine and deep learning in kinase QSAR modeling
Deep learning’s effectiveness in QSAR modeling, especially within drug discovery and kinase research, stems from several significant advantages. Firstly, deep learning models excel at handling large datasets, which are increasingly common in kinase-related pharmaceutical research. They possess the capacity to analyze complex, high-dimensional data and extract insights from diverse sources, including chemical structure properties alongside cellular and biological activity data [2,33].
A critical advantage is deep learning’s ability to perform automatic feature extraction from raw data. In QSAR modeling, this translates to the direct identification of relevant molecular features from chemical structures, significantly reducing the need for extensive manual feature engineering. This automated feature learning not only streamlines the modeling process but also enhances model accuracy by capturing features that might be overlooked by human-designed descriptors [27,33].
Deep learning’s ability to automatically extract features from raw data is a key advantage. Unlike traditional QSAR requiring pre-defined descriptors, deep learning models, particularly CNNs and GNNs, can learn relevant molecular features directly from representations like graphs or voxel grids. This reduces reliance on expert-defined descriptors, potentially uncovering novel, non-obvious features that are crucial for activity prediction and streamlining the model development process significantly.
Furthermore, biological systems and chemical interactions, particularly within kinase-regulated cellular activities, are inherently complex and non-linear. Deep learning models are specifically designed to capture these complex, non-linear relationships through their multiple layers of non-linear transformations. This capability leads to more precise predictions of biological activity and compound efficacy.
Deep learning offers remarkable flexibility and scalability in QSAR tasks. Architectures like Convolutional Neural Networks (CNNs) are well-suited for analyzing spatial arrangements in molecular graphs, while Recurrent Neural Networks (RNNs) are effective for sequential data analysis. Moreover, deep learning facilitates multitask learning, enabling a single model to simultaneously predict multiple properties or activities. This is particularly beneficial in kinase research, allowing for a more holistic understanding of kinase interactions and activities, ultimately improving predictive accuracy and providing deeper insights into target relationships [2,22,27,33,34].
In essence, machine learning has significantly advanced QSAR modeling by overcoming limitations of traditional approaches. Its capacity for autonomous learning from complex data and making precise predictions positions it as a transformative technology in kinase inhibitor discovery and beyond.
QSAR methods, such as CoMFA and CoMSIA, often rely on static, pre-defined molecular descriptors and are limited by smaller datasets. This often results in moderate predictive accuracy, reduced flexibility, and limited scalability. While computationally less demanding, these approaches lack the adaptability to address the complex challenges of modern drug discovery.
In contrast, modern machine learning-based QSAR models, particularly deep learning models like CNNs and RNNs, leverage large datasets and excel at capturing complex, non-linear relationships between molecular features and biological activity. These methods achieve high predictive accuracy, flexibility, and scalability, making them invaluable tools in contemporary drug discovery, despite their increased data and computational demands.
4.4. Key machine and deep learning techniques in kinase research
The most prominent machine and deep learning models applied in recent kinase research include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Graph Neural Networks (GNNs), and Multitask Learning Approaches [2,24].
4.4.1. Convolutional neural networks (CNNs)
Convolutional Neural Networks (CNNs) are a class of deep learning models exceptionally well-suited for analyzing structured data, such as images. In the context of molecular structures and kinase inhibitor activity prediction, CNNs are used to identify patterns and features within molecular representations that correlate with biological activity [24]. CNNs have demonstrated high effectiveness in molecular data analysis, for example, achieving a ROC-AUC of 0.87 in identifying Aurora-A kinase inhibitors [23].
To elaborate on their mechanism within QSAR, CNNs utilize convolutional layers to automatically extract hierarchical features from molecular representations. These layers act as filters, detecting spatial patterns like functional groups or substructures within representations such as 2D molecular fingerprints, 3D voxel grids, or even molecular graphs converted to image-like formats. Pooling layers then reduce dimensionality, focusing on the most salient features. This automated feature extraction removes the need for manual descriptor engineering, allowing CNNs to learn complex relationships directly from raw molecular data.
CNNs offer versatility in how molecules are represented. For kinase inhibitor modeling, molecules can be depicted as graphs (atoms as nodes, bonds as edges), which can then be converted into grid-like structures or “images” processable by CNNs. Alternatively, molecules can be represented as 2D images (e.g., molecular fingerprints) or, for more detailed analysis, as 3D molecular structures in voxel grids. CNNs automatically extract hierarchical features from these representations. Initial layers capture local structural features, while deeper layers identify more complex, global molecular patterns. For predictive modeling, CNNs are trained on datasets where the input is the molecular structure and the output is the biological activity (e.g., kinase inhibitory activity). The training process optimizes network weights to minimize prediction errors, enabling the CNN to learn features relevant to biological activity [24,27,28].
Data preparation for CNN-based kinase QSAR begins with assembling a dataset of known kinase inhibitors with their corresponding activities. These inhibitors are then transformed into a CNN-compatible format, such as 2D images of fingerprints or 3D voxel grids. This transformation is essential for enabling the CNN to learn and effectively predict the activity of novel kinase inhibitors [28].
Model training involves using this labeled data (molecular structures with known activities) to train the CNN. During training, the network learns to recognize patterns and features in the molecular structures that are indicative of biological activity against specific kinases specific kinases [27,28].
Prediction, validation, and testing follow training. Once trained, the CNN can predict the activity of new, unseen kinase inhibitors by inputting their molecular structures. The CNN then generates an activity prediction, either as a continuous score or a categorical label (e.g., active/inactive). Model performance is rigorously evaluated using separate test datasets to ensure generalization to new data. Common evaluation metrics include accuracy, precision, recall, and ROC-AUC to validate the model’s predictive capability [2,28]. In summary, CNNs are powerful tools for kinase inhibitor discovery, leveraging their ability to automatically learn patterns from complex molecular data to analyze known structures and predict the activity of novel compounds.
4.4.2. Recurrent neural networks (RNNs)
Recurrent Neural Networks (RNNs) are designed for modeling sequential data, making them suitable for tasks involving temporal dependencies, such as the dynamics of kinase activities or interactions. RNNs are equipped with internal memory to process input sequences, enabling them to capture patterns that evolve over time [24,35,36]. RNNs, while less directly applied to static molecular structure QSAR than CNNs and GNNs, excel at processing sequential data. Their internal memory allows them to consider the order of input features, making them suitable for analyzing time-dependent kinase activities or potentially sequential representations of molecular properties. In QSAR, their strength lies in capturing temporal dependencies, though their application to static molecular data is less common compared to CNNs and GNNs.
In kinase research, RNNs can model the temporal progression of kinase activation or inhibition. Kinase signaling pathways often involve sequential activation or inhibition events. RNNs can learn from time-series data to predict future kinase activities based on past observations. This is valuable for understanding cellular signaling dynamics and predicting the time-dependent effects of kinase inhibitors time [24,35,36].
Furthermore, RNNs can model sequential interactions between kinases and their substrates. By analyzing sequences of kinase-substrate interactions, RNNs can identify patterns in kinase-substrate recognition, aiding in predicting new kinase-substrate interactions and understanding complex signaling cascades.
4.4.3. Graph neural networks (GNNs)
Graph Neural Networks (GNNs) are specialized models designed for analyzing graph-structured data. This makes them highly relevant for capturing the intricate connectivity and interactions within molecular graphs, which is crucial for understanding kinases and their interactions [24,29]. GNNs directly process molecules as graphs, where atoms are nodes and bonds are edges. They employ a message-passing mechanism, where information is iteratively exchanged between neighboring nodes, allowing the network to learn representations that capture both node features (atom properties) and the connectivity of the molecule. This makes GNNs particularly adept at understanding how molecular structure and interatomic relationships influence kinase interactions and activity, offering a powerful approach for capturing intricate structural features beyond linear descriptors.
In the context of kinases, molecular structures can be directly represented as graphs, with atoms as nodes and bonds as edges. GNNs can learn directly from these graph representations to predict various properties and activities related to kinases [24,29,35].
GNNs are particularly effective at capturing the complex structural and functional relationships between kinases and their interacting molecules (substrates, inhibitors, activators). By aggregating information from neighboring nodes in the molecular graph, GNNs can learn to predict kinase activities, substrate specificity, and binding affinities [24,29,35].
4.4.4. Multitask learning approaches
Multitask learning, a powerful technique within machine learning, is used to simultaneously predict multiple kinase activities and interactions, enhancing predictive capabilities and learning efficiency. Algorithms and frameworks like multitask neural networks (MTNNs) and multitask graph neural networks (MTGNNs) are employed, leveraging shared representations across related tasks to improve generalization [37].
Platforms like KinomeX exemplify the application of multitask learning in kinase research. KinomeX uses multitask neural networks to predict multiple kinase activities, capturing shared features across kinases for more accurate predictions and exploration of kinase interactions and cross-reactivity. Integrating data from diverse sources (chemical structures, biological assays, protein interaction networks), KinomeX offers a comprehensive approach to kinase behavior understanding. Case studies and publications demonstrate KinomeX’s utility in enhancing predictive performance, modeling complex biological processes, and generating holistic insights for kinase inhibitor discovery [37].
Comparative analyses show that KinomeX and similar multitask approaches often outperform traditional single-task QSAR models and other multitask learning methods. These approaches offer improved prediction accuracy, robustness, and facilitate the discovery of novel kinase interactions and inhibitor candidates. Tools like DeepKinZero and DeepKinNet also utilize multitask learning for kinase inhibitor activity prediction across multiple targets, aiding drug discovery and design processes [21,24,37]. The potential impact of multitask learning in kinase research and drug discovery is significant, promising advancements in computational biology and medicinal chemistry, and potentially transforming disease understanding and treatment.
4.4.5. Machine learning QSAR modeling and validation framework for comparison
To provide a comparative context, machine learning QSAR models are often developed and used as baselines for validation. Ensemble learning methods, specifically Random Forest (RF) and Support Vector Regression (SVR), are commonly employed for this purpose. RF and SVR, while distinct, are both widely used and effective machine learning algorithms.
Random Forest (RF) is an ensemble method suitable for both classification and regression. It combines predictions from multiple decision trees, trained on random data subsets, enhancing overall performance. RF is robust to noise, handles diverse data types, and reduces overfitting, though it can be computationally intensive and less interpretable than single trees.
Support Vector Regression (SVR) extends Support Vector Machines (SVM) for regression, fitting a hyperplane within a tolerance margin. SVR excels with non-linear data using kernels (linear, polynomial, RBF), is effective for small to medium datasets with complex relationships, and resists overfitting when tuned. However, it’s sensitive to hyperparameter choice and computationally demanding for large datasets, especially with non-linear kernels. RF is generally better for larger, noisy datasets, while SVR is suited for smaller datasets with intricate patterns.
Molecular descriptors, including hydrophobicity, molecular weight, topological indices, and electronic properties, are typically calculated using tools like RDKit and KNIME for these baseline ML models. Feature selection, often using Recursive Feature Elimination (RFE), is performed to optimize model performance.
Datasets for QSAR model development are often sourced from curated resources like ChEMBL and NCI databases. These datasets, comprising kinase inhibitors with activity data, are crucial for QSAR applications. Data preprocessing typically involves removing duplicates, normalizing activity values, and applying filters like Lipinski’s Rule of Five. Datasets are then split into training and testing subsets (e.g., 70% training, 30% testing). Databases such as ChEMBL, PubChem, PDB, BindingDB, ZINC, DrugBank, and the NCI Database are invaluable, providing bioactivity data, structural information, and chemical libraries essential for training QSAR models, virtual screening, and broader drug discovery applications [38–41].
Performance metrics used for model evaluation include RMSE, R2, sensitivity, specificity, and ROC-AUC. In comparative validation, the focus is often on qualitative comparisons of model performance rather than solely presenting numerical results. For example, Random Forest and Graph Neural Networks may be shown to demonstrate robust predictive power and generalizability based on consistent performance across validation datasets. Convolutional Neural Networks might excel in capturing complex molecular relationships in datasets with non-linear activity profiles. Support Vector Regression might show slightly less robustness compared to ensemble and deep learning models across diverse datasets. This qualitative approach emphasizes the comparative strengths and weaknesses of different methodologies based on their performance metrics on validation datasets, providing a nuanced understanding beyond just numerical output values.
4.4.6. Perturbation theory machine learning (PTML) for kinase drug discovery
Perturbation Theory Machine Learning (PTML) offers a valuable and distinct approach within kinase-targeted drug discovery, uniquely addressing the prediction of kinase-relevant phenotypic activity. Moving beyond target-level kinase inhibition, PTML models complex biological responses at the cellular level, crucial for translating in-vitro results to in-vivo efficacy of kinase inhibitors. PTML models provide key advantages for kinase research:
Predicts Kinase Phenotypic Activity: PTML directly predicts phenotypic outcomes relevant to kinase modulation, such as cancer cell growth inhibition or signaling pathway changes. This is more clinically relevant than just predicting kinase binding.
Multitask for Kinase Context: PTML excels in multitask modeling, predicting kinase inhibitor effects across multiple kinase targets, cell lines, and assay conditions. This captures the polypharmacology and context-dependent effects crucial in kinase biology.
Interpretable for Kinase Design: PTML models are often interpretable, revealing structure-activity relationships for kinase inhibitors. These insights guide rational design of improved kinase-targeted drugs.
PTML’s utility in kinase and cancer research is evidenced by a range of successful applications. For instance, PTML has been employed to design virtual agents targeting lung cancer cell phenotypes [42], and has aided in the design of inhibitors for pancreatic cancer that target multiple kinases and cell lines simultaneously [43]. Furthermore, PTML has facilitated the design of dual inhibitors, such as those targeting CDK4 and HER2 [44]. Cell-based PTML models have also been developed to design inhibitors specifically for liver cancer cell lines [45]. These diverse examples underscore PTML’s significant value in kinase drug discovery, particularly when considering phenotypic outcomes and polypharmacological aspects. This demonstrated utility positions PTML as a crucial approach within the broader landscape of kinase machine learning.
While deep learning excels in complex kinase data analysis, PTML provides a complementary approach focused on phenotypic relevance and interpretability, critical for translating kinase inhibition to therapy. PTML, often prioritizing mechanistic insight, can be integrated with deep learning for hybrid approaches. The optimal strategy for kinase-targeted drug discovery depends on research goals and data, balancing prediction, interpretability, and phenotypic focus.
A comparative overview of the machine learning and deep learning methods utilized in this research is presented in Table 3. This table outlines their core principles, advantages, and limitations in the context of kinase QSAR.
Table 3.
Comparative analysis of machine learning and deep learning algorithms for kinase QSAR.
| Method | Type | Description/core idea | Strengths in kinase QSAR | Weaknesses in kinase QSAR | Hyperparameters (mentioned in text or common) | Typical applications in kinase QSAR |
|---|---|---|---|---|---|---|
| Random Forest (RF) | Ensemble ML | Ensemble of Decision Trees; predictions based on majority vote (classification) or average (regression). | Robust to noise, handles diverse data types, reduces overfitting, good baseline performance, feature importance estimation. | Can be computationally intensive for very large datasets, less interpretable than single trees, can be outperformed by DL methods. | Splitting Criterion (Information Gain Ratio), Number of Trees (e.g., 100), Tree depth, Minimum node size. | Baseline model for comparison, activity prediction, feature selection, virtual screening. |
| XGBoost | Ensemble ML | Gradient Boosting of Decision Trees; sequential model building to correct errors of previous trees. | High predictive accuracy, handles complex relationships, efficient implementation, feature importance, regularization to prevent overfitting. | Can be sensitive to hyperparameter tuning, potential for overfitting if not carefully tuned, can be less interpretable than simpler models. | Boosting Rounds (e.g., 100), Eta (learning rate, e.g., 0.3), Max Depth (e.g., 6), Gamma, Lambda, Alpha, Subsampling, Colsampling. | High-performance activity prediction, virtual screening, lead optimization, understanding complex SAR. |
| Naïve Bayes (NB) | Traditional ML | Probabilistic classifier based on Bayes’ theorem; assumes feature independence (naïve assumption). | Simple and fast to train, computationally efficient, works well with high-dimensional data, can be useful for initial screening. | Strong feature independence assumption often violated in molecular data, may not capture complex relationships, lower predictive accuracy compared to complex models. | Prior probabilities, feature distributions (e.g., Gaussian for continuous, Multinomial for discrete). | Initial screening, quick activity prediction, feature importance assessment (though limited by independence assumption). |
| Convolutional Neural Networks (CNNs) | Deep Learning | Neural networks designed for structured data (like images); uses convolutional layers to extract features. | Automatic feature extraction from molecular representations (2D fingerprints, 3D voxel grids), captures spatial patterns, high predictive accuracy for complex data. | Requires structured input representations (images, grids), can be computationally intensive for 3D representations, interpretability can be challenging, needs large datasets for optimal performance. | Number of Convolutional Layers, Filter sizes, Pooling layers, Number of Filters, Activation Functions, Regularization. | Activity prediction using image-like molecular representations, pattern recognition in molecular structures, virtual screening. |
| Recurrent Neural Networks (RNNs) | Deep Learning | Neural networks designed for sequential data; uses recurrent connections to process sequences. | Can model temporal dependencies (though less directly applicable to static QSAR), can handle variable-length input sequences, useful for processing sequential molecular representations. | Less naturally suited for static molecular structures compared to CNNs and GNNs, can be challenging to train (vanishing/exploding gradients), interpretability can be complex. | Number of Recurrent Layers, Hidden Units, Activation Functions, Sequence Length handling, Regularization. | Modeling time-dependent kinase activities (if applicable), processing sequential molecular descriptors (e.g., SMILES strings). |
| Probabilistic NN (PNN) | Neural Network | Feedforward NN with exponential activation function; approximates Bayes optimal decision boundaries. | Fast training, robust to noisy data, good for pattern recognition and classification, non-linear decision boundaries. | Can be less flexible than deep neural networks, may not scale as well to very large datasets, default hyperparameters might not be optimal for all datasets. | Smoothing parameter (Sigma - default in KNIME). | Classification tasks, pattern recognition in molecular data, activity classification (active/inactive/intermediate). |
| k-Nearest Neighbors (kNN) | Traditional ML | Instance-based learning; classifies based on majority class of k-nearest neighbors in feature space. | Simple to understand and implement, non-parametric (no assumptions about data distribution), can capture local patterns. | Computationally expensive for large datasets (needs to store all training data), sensitive to feature scaling and distance metric choice, performance depends heavily on “k” value. | Number of Neighbors (k, e.g., 3-5), Distance Metric (e.g., Euclidean, Manhattan), Distance-weighting (distance-dependent/independent). | Activity prediction based on similarity, exploring neighborhood relationships in chemical space. |
| Multilayer Perceptron (MLP) | Neural Network | Feedforward NN with multiple layers (input, hidden, output); learns complex non-linear functions. | Can learn complex relationships, flexible architecture, basis for deeper neural networks, can be used for both classification and regression. | Can be prone to overfitting, requires careful hyperparameter tuning (architecture, learning rate, regularization), can be computationally intensive to train, interpretability can be challenging. | Number of Hidden Layers, Number of Neurons per Layer, Activation Functions, Learning Rate, Regularization (e.g., L1, L2, Dropout). | Activity prediction, modeling complex SAR, non-linear relationship modeling, virtual screening. |
| Locally Weighted Learning (LWL) | Instance-based ML | “Lazy learner”; weights instances based on their proximity to the query point during prediction. | Adapts to local data patterns, can capture non-linear relationships, instance-based (preserves training data information). | Computationally expensive at prediction time (needs to re-weight instances for each prediction), can be sensitive to distance metric and weighting function, might be overfit in noisy regions. | Distance Metric, Weighting Function, Bandwidth/Kernel parameters. | Modeling local SAR, capturing non-linearities in specific regions of chemical space. |
| Graph Neural Networks (GNNs) | Deep Learning | Neural networks designed for graph data; directly processes molecules as graphs (atoms and bonds). | Directly processes molecular graphs, captures intricate connectivity and relationships, automatic feature extraction from graph structure, high predictive accuracy for complex SAR. | Computationally more intensive than simpler ML methods, interpretability can be challenging, requires graph-based molecular representations, needs substantial data for complex models. | Number of GNN Layers, Message Passing functions, Aggregation functions, Node/Edge feature dimensions, Regularization. | Activity prediction based on molecular graphs, modeling complex SAR, understanding structure-activity relationships, virtual screening, target prediction. |
| Support Vector Regression (SVR) | Traditional ML | Extends Support Vector Machines for regression; finds optimal hyperplane within a tolerance margin. | Effective with non-linear data (using kernels), good for small to medium datasets, resists overfitting when tuned properly. | Sensitive to hyperparameter choice (kernel, regularization), computationally demanding for large datasets, especially with non-linear kernels, can be outperformed by DL methods for very complex data. | Kernel type (linear, polynomial, RBF), Regularization parameter (C), Epsilon (tolerance margin). | Regression tasks, activity prediction, modeling non-linear SAR in smaller datasets. |
This table presents a comparison of machine learning (ML) and deep learning (DL) methods in the context of kinase quantitative structure-activity relationship (QSAR) modeling. It highlights their distinct approaches, advantages, limitations, relevant hyperparameters, and common use cases to facilitate a comparative understanding of their suitability for kinase drug discovery.
5. Kinase enzyme QSAR studies
5.1. The evolution and impact of QSAR studies in kinase inhibitor discovery
Until the late 1980s and early 1990s, there were no significant systematic efforts reported to investigate the kinase modulators or to study the kinase inhibitory structure-activity relationship due to the highly conserved ATP-binding domains among the kinases [46,47].
The 1992 Nobel Prize in Medicine and physiology was awarded to Edmond Fischer and Edwin Krebs for their discovery of the reversible protein phosphorylation-dephosphorylation process as a biological regulatory mechanism [25,48,49]. This discovery is considered as the real beginning of the small-molecule kinase modulator research.
In 1996, Hu et al. published one of the first articles focusing on the protein kinase C (PKC) inhibitors and their structure-activity relationship [50]. PKC is a key element in the human transduction pathway, also it belongs to the family of serine/threonine protein kinases which was discovered by the Japanese scientist Nishizuka and coworkers in 1977 [51–54].
The article focused on the advances of some PKC inhibitors based on their chemical class: Staurosporine, Balanoids, Sphingolipids, Isoxazolones, and natural products [50].
Later, as reported in a 1996 publication by Druker et al., Imatinib itself was designed by using a known structure of the ATP binding pocket, a series of chemical derivatives from the 2-phenylaminopyrimidine class were synthesized and screened for their ability to inhibit a several protein kinases [55]. Following, the drug was subjected to numerous SAR studies and was used to guide the design and discovery of new kinase inhibitors [56,57]. Since 1998, concepts such as drug-likeness, lead-likeness, fragment-based drug design, and target-focused drug design approaches have been developed successively.
The primary factor that contributed to the development of the kinase SAR studies was the disclosed information related to the Imatinib discovery. After the approval of Imatinib, the QSAR field has improved intensely, fueled by the escalated availability of the experimental data sets related to kinase active compounds. These changes occurred concurrently with the advances in chemometric and chemical analysis techniques, resulting in an upsurge in the number of chemical and physical descriptors. Additionally, during the early 2000s, there was a growing interest in advanced computer-aided modeling techniques available for QSAR studies.
In 2002, research on QSAR studies regarding candidates for kinase targets was limited. For instance, Manallack et al. utilized a series of neural networks to identify compounds that specifically interact with kinases [58,59]. At that time, the pharmaceutical industry emphasized the early identification of compounds likely to fail in the later stages of drug design and development, following the so-called “fail fast, fail cheap” approach [58,59].
The application of QSAR studies against the kinase enzymes as an evaluative approach, i.e., with the focus on developing an explanatory model of previously published data was accelerated in 2003. It was when Kamath & Buolamwini explored two members of the family of RTK, namely EGFR and HER-2 receptors, by applying the Comparative Molecular Field Analysis (CoMFA) and the Comparative Molecular Similarity Analysis (CoMSIA) QSAR studies on a group of 50 active benzylidene malonitrile tyrphostins derivatives. The study evaluated the binding mode of the studied compounds inside the EGFR and HER-2 ATP binding pocket and explained the selectivity of the dihydroxy compounds toward the EGFR relative to HER-2. The study also concluded that the active compounds are more likely having multiple binding modes at the EGFR binding site and only one binding mode inside the HER-2 receptor [60,61].
In 2005, only four kinase inhibitor compounds had reached commercial use: fasudil for rho kinase which was only approved in Japan in 1995 for treating cerebral vasospasm [51,62], sirolimus for mammalian target of rapamycin (mTOR), Imatinib for BCR-Abl, and Gefitinib for EGFR [47,51]. In that same year, Sprous et al. published article describing a simple QSAR model that is directed to distinguish kinase inhibitors from other drug-like compounds [47]. The method applied 1D and 2D descriptors in a multivariable QSAR model to determine the chemical space within a library of known kinase inhibitors. The model correctly identified 98% of the 258 training set compounds and allowed the recognition of the kinase inhibitor space.
The scientific advancement of the QSAR studies experienced a surge during the second decade of the 21st century. This was propelled by the collective scientific efforts of researchers worldwide. Notably, Alexander Tropsha and his colleagues at North Carolina University made a dramatical contributions to the development of the QSAR modeling approaches. Their research studies have contributed to the transition of the QSAR studies from being an evaluative method to becoming a practical drug discovery approach that can aid in screening and identifying novel biologically active compounds [32,34,63,64].
Accordingly, QSAR and other Computer-aided drug design and machine-learning tools have recently affected the drug discovery efforts in many aspects. For example, fragment-based lead discovery (FBLD) was used efficiently to design and synthesize Vemurafenib (code PLX-4032, brand name Zelboraf®) (please refer to Table 3S under supplementary material: Serine/Threonine Protein Kinase Inhibitors: B-Raf). Vemurafenib was found active against cancer cells with mutated Raf (V600E). During the process of FBLD, the drug is built up from small fragments. Originally, a library of small fragments is screened against the target, then small fragment binding potential values are established using Nuclear Magnetic Resonance (NMR) and X-ray crystallography. Next, the active hits are synthesized and examined against the investigated target. Surprisingly, Vemurafenib took only 6 years from concept to market and was approved in 2011 [65].
Still, the QSAR modeling is considered as one of the promptest methods in computer-aided drug design since it helped drug designers and medicinal chemists to understand the relationship between kinase enzyme activity and the molecular properties of the targeted inhibitors. Henceforth, QSAR studies are considered as promising technique in synthesizing novel, potent and selective kinase enzymes modulators. And so, when combined with other computational techniques, the QSAR tool has proved to provide rational insights to facilitate the discovery of novel hits, leads, and to treat dysregulated protein kinase enzymatic activity [10,18,66].
5.2. Revolutionizing kinase inhibitor discovery: the impact of machine learning on QSAR modeling
CNNs, RNNs and GCNs studies show the significant influence of deep learning on QSAR modeling for kinase inhibitors. Advanced deep learning techniques have led to significant breakthroughs in kinase research. These efforts have not only improved the accuracy of predicting kinase inhibitor activity but have also accelerated the identification and development of new therapeutic agents.
In a study entitled “Integrating QSAR modeling and deep learning in Drug Discovery: the emergence of deep QSAR” published online in 2024, Tropsha et al. [2] have concluded that deep learning methods, particularly deep QSAR, have not yet led to approved drugs, but they have significantly accelerated preclinical research stages for small-molecule drug candidates. For example, Exscientia’s first AI designed drug entered a phase I clinical trial after only 12 months of exploratory research, while In-silico Medicine developed a phase I anti-fibrotic clinical candidate in 30 months starting from target discovery. Both companies utilized deep learning approaches discussed in this perspective. These successes indicate that deep QSAR is entering a productive phase, likely leading to faster discovery of small-molecule drugs, crucial especially in addressing emerging infectious diseases like COVID-19 [24,37,67].
In a recent study by Nippa et al., an ensemble of 100 deep QSAR models was trained using the ELECTRA algorithm (Efficiently Learning an Encoder that Classifies Token Replacements Accurately) [2,22,67]. These models predicted the extent to which compounds would inhibit phosphoinositide 3-kinase-γ (PI3Kγ), each with slightly different results. Testing showed that the number of votes a compound received was generally reflected in how well it performed in lab tests. The top-scoring compound inhibited the target PI3Kγ, with a Ki of 63 nM [22,67].
In the context of multitask learning approaches Liu et al. demonstrated that multi-task learning, which models docking data for both new and known targets, outperforms single-task and active learning approaches. Emerging methods also use deep learning to predict protein-ligand interactions, allowing for extensive target profiling of various chemicals. For example, Li et al. developed a kinome-wide multi-task deep neural network model using around 140,000 data points from 391 kinases. This model mapped a comprehensive kinome interaction network and led to the creation of KinomeX, an online platform predicting kinome-wide polypharmacology based on 2D molecular structures [22,24,37].
Deep learning improves the accuracy of ligand binding affinity and property calculations. For instance, the ANI molecular mechanics scheme has been expanded to predict protein-ligand binding free energies, which are crucial for deep QSAR modeling. This approach combines molecular dynamics simulations of the protein-ligand complex with molecular mechanics and machine learning potential energies [2,24,33]. Studies testing this approach with kinase inhibitors demonstrated significant improvements in predicting binding affinities, reducing absolute binding free energy errors compared to traditional molecular mechanics calculations. Notably, the corrections from molecular mechanics to machine learning-molecular mechanics were positive and helped improve the accuracy of predictions, particularly for aliphatic groups with high conformational flexibility [24,36].
One emerging trend in the application of deep learning to kinase research is the integration of generative models for designing novel inhibitors. Generative models, such as variational autoencoders (VAEs) and generative adversarial networks (GANs), offer the potential to generate novel molecular structures with desired properties. By leveraging large datasets of kinase inhibitors, these models can aid in the discovery of new compounds with specific activity profiles, potentially accelerating drug development processes [2].
Integrating deep learning into kinase research promises to revolutionize drug discovery by enabling accurate predictions of drug activity and personalized treatment design. These advancements could lead to targeted therapies that improve outcomes and reduce adverse effects.
6. Case studies: investigating dysregulated kinase enzymatic activity using QSAR
To illustrate the practical application and impact of QSAR studies in kinase research, this review now examines several case studies focusing on dysregulated kinase targets of significant scientific and therapeutic interest. These case studies, detailed in Supplementary Material Section 3, highlight QSAR investigations on Cyclin-Dependent Kinases (CDKs), Moloney-Murine Leukemia Virus Kinases (PIM-1, PIM-2, and PIM-3), Ca2+/Calmodulin-Dependent Protein Kinase II (CAMKII), Polo Kinase 1 (Plk1), and Tropomyosin Receptor Kinase A (TrkA). These examples showcase how QSAR modeling has been employed to understand inhibitor binding modes, identify key structural features for activity, and guide the design of novel kinase inhibitors for disease management and treatment. While Computer-Aided Structure-Based Drug Design (SBDD) has directly contributed to FDA-approved kinase inhibitors, these case studies underscore the significant evaluative and explanatory role of QSAR in kinase drug discovery, providing valuable insights from previously published data and paving the way for future advancements. For in-depth analysis of each case study, please refer to Supplementary Material Section 3.
7. Limitations and challenges
Since the 1980s, kinase-based drug discovery has overtaken other drug discovery targets as the most investigated cellular targets for disease treatment [11]. Despite the achieved advancement, several limitations and challenges are reflected during the exploration of new kinase-based approved drugs.
As mentioned earlier, the ATP binding domains of kinase enzymes are more conserved than those of the other substrate binding domains [46]. Additionally, the high intracellular ATP concentrations require high intracellular concentrations of the kinase inhibitors. Accordingly, the scientists found it much simpler to search for substrate competitive kinase inhibitors rather than ATP competitive kinase.
Essentially, only a fraction of the human kinases has been the focus of small molecule drugs, leaving a significant portion of the kinases understudied and unexploited [68]. Some references indicate that approximately 70%, or around 400 kinases, within the kinome is still not investigated [69]. It’s acknowledged that most kinase modulation efforts are directed toward the tyrosine kinase group (e.g., EGFR, FGFR, PDGFR) this is clearly illustrated by the fact that as of November 2023, inhibitors of Protein-Tyrosine Kinases account for more than 40 FDA-approved therapeutic agents [1,11].
Second, most of the FDA-approved kinase modulators are targeting cancer treatment although the kinase signaling cascades are related to diverse cellular activities, for example, inflammatory, CNS, cardiovascular, and diabetes activities.
Third, many of the current kinase inhibitors share the same core with some different substituents. Thus, this is reflected by the limited number of FDA-approved kinase inhibitor compounds.
Fourth, the high amino acid sequence similarity inside the ATP-binding pockets of kinases has been proven challenging to develop kinase inhibitors with potent inhibition against selective targets and minimal interactions with off-targets.
Fifth, closely connected with the previous point, many inhibitors interact with more than one target. By contrast, few absolute-selective inhibitors, which might be evaluated as dual- or multiple-target inhibitors if a more comprehensive screening assay was used.
8. Conclusion
Quantitative Structure-Activity Relationship (QSAR) modeling remains a powerful and advanced approach for kinase-targeted drug design, offering significant potential in modern pharmaceutical research. While structure-based design and other methodologies have contributed to the kinase inhibitors currently available, QSAR provides a valuable ligand-based perspective, particularly when enhanced by modern machine learning techniques. Historically, QSAR has served as an evaluative tool and a predictor of drug-likeness and ADME properties in early drug discovery. The aspiration to utilize QSAR models as accurate filters for virtual screening and to identify crucial activity-driving molecular descriptors remains a central goal.
The integration of machine learning, especially deep learning, offers a transformative path forward for QSAR in kinase inhibitor discovery. These methods enhance QSAR by automating feature extraction, handling complex datasets, and capturing non-linear relationships often missed by traditional approaches. However, the interpretability of these complex models is crucial for accelerating drug design. To maximize impact, future efforts must prioritize model interpretability, employing techniques such as feature importance analysis, attention mechanisms, and Explainable AI (XAI) methods to unlock actionable insights from these powerful models. These interpretable insights, revealing key structural determinants of kinase activity, can directly guide medicinal chemistry optimization and accelerate the design of novel inhibitors.
While statistical validation remains important, a renewed emphasis on experimental validation, including preclinical and clinical studies, is essential to translate QSAR findings into tangible drug discovery progress. Moving beyond purely statistical validation, future kinase QSAR studies should focus on prospective experimental validation to confirm model predictions and drive the design and optimization of novel bioactive kinase modulators.
The field of kinase-targeted drug discovery is rapidly advancing through the integration of machine learning QSAR, as evidenced by collaborative validation efforts like the IDG-DREAM Drug-Kinase Binding Prediction Challenge (see Supplementary Material Section 1 for details). Such initiatives underscore the growing confidence in computational methods and their potential to accelerate the development of novel therapeutics.
In conclusion, the synergy of QSAR modeling with advanced machine learning, particularly when coupled with strategies to enhance model interpretability and rigorous experimental validation, holds immense promise for advancing kinase-targeted drug discovery. QSAR, enriched by these modern approaches, can unlock new clinical opportunities, predict undesired side effects, and ultimately accelerate the development of more effective and safer kinase inhibitors. Future research should focus on refining QSAR modeling procedures, integrating interpretability techniques, and prioritizing experimental validation to fully realize the potential of QSAR in the kinase modulator drug discovery process.
9. Future perspectives
The future of QSAR modeling in kinase inhibitor research is exceptionally bright, building upon a rich history of computer-aided drug design efforts targeting these crucial enzymes. QSAR models, now increasingly sophisticated, represent a powerful distillation of extensive biological and chemical knowledge accumulated against kinase targets. The critical next step is to fully leverage this knowledge for tangible advancements in preclinical and clinical kinase inhibitor development.
One key direction lies in the continued refinement and expansion of QSAR methodologies through the deeper integration of advanced machine learning techniques. Deep learning, with architectures like Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Graph Neural Networks (GNNs), will undoubtedly play an increasingly central role. However, future progress must prioritize model interpretability. Developing and applying techniques to understand why these complex models make their predictions will be paramount. This includes incorporating Explainable AI (XAI) methods, attention mechanisms, and feature importance analyses to extract actionable insights into ligand-target interactions and guide rational kinase inhibitor design.
Furthermore, the integration of Perturbation Theory Machine Learning (PTML) represents a highly promising avenue. PTML’s focus on predicting phenotypic activity, coupled with its inherent interpretability, directly addresses the critical need to translate target-level kinase inhibition to cellular and in vivo efficacy. Future research should explore hybrid approaches that combine the strengths of PTML with deep learning, potentially using deep learning for feature extraction within a PTML framework or employing PTML to validate and interpret findings from deep learning QSAR models.
Crucially, the field must move beyond statistical validation and embrace rigorous experimental validation. Future QSAR studies should prioritize prospective validation, designing experiments to specifically test model predictions in preclinical and, where feasible, early clinical settings. This iterative cycle of model building, experimental testing, and model refinement will be essential to build truly robust and predictive QSAR models that can significantly accelerate kinase inhibitor drug discovery.
Looking ahead to the next 5–10 years, the integration of advanced QSAR and machine learning techniques promises a transformative era in kinase inhibitor discovery. We anticipate the development of highly predictive, interpretable, and experimentally validated models capable of significantly streamlining virtual screening to efficiently identify promising candidates, guiding rational lead optimization by providing detailed structure-activity insights, predicting phenotypic efficacy and polypharmacology to enhance clinical translatability, and ultimately accelerating the discovery of novel kinase targets and therapeutic strategies. By prioritizing interpretability, experimental validation, and synergistic integration of diverse machine learning approaches, the future of QSAR in kinase inhibitor research is poised to fully realize its potential in accelerating and enhancing the development of life-saving kinase-targeted therapies.
Supplementary Material
Acknowledgments
The authors would like to thank the Deanships of Scientific Research at The Hashemite University for funding this project.
Funding Statement
The authors would like to thank the Deanships of Scientific Research at The Hashemite University for funding this project. The authors certify that financial support was received for this research and/or the creation of this work, the sources and extent of involvement of which are clearly identified in the manuscript.
Author contributions
Rand Shahin, Associate professor, r.shahin@hu.edu.jo, Main author. Sawsan Jaafreh, Assistant professor, sawsan.jaafreh@hu.edu.jo, analysis, or interpretation and drawing the graphs in addition to revising the manuscript. Yusra Azzam, Undergraduate student, yusraazz@ub.edu, Language editing. Adding scientific information regarding AI, machine learning and deep learning QSAR.
Disclosure statement
The authors have no other relevant affiliations or financial involvement with any organization or entity with a financial interest in or financial conflict with the subject matter or materials discussed in the manuscript apart from those disclosed.
References
Papers of special note have been highlighted as either of interest (*) or of considerable interest (**) to readers.
- 1.Roskoski R. Properties of FDA-approved small molecule protein kinase inhibitors: A 2023 update. Pharmacol Res. 2023;187:106552. [DOI] [PubMed] [Google Scholar]; **This paper provides the latest information on FDA-approved kinase inhibitors, making it crucial for understanding current drug properties and clinical applications.
- 2.Tropsha A, Isayev O, Varnek A, et al. Integrating QSAR modelling and deep learning in drug discovery: the emergence of deep QSAR. Nat Rev Drug Discov. 2024;23(2):141–155. doi: 10.1038/s41573-023-00832-0 [DOI] [PubMed] [Google Scholar]; **This review highlights the integration of QSAR modeling and deep learning, offering insights into modern computational approaches in drug discovery.
- 3.Cohen P, Cross D, Jänne PA.. Kinase drug discovery 20 years after imatinib: progress and future directions. Nat Rev Drug Discov. 2021;20:551–569. [DOI] [PMC free article] [PubMed] [Google Scholar]; **This comprehensive review discusses the advancements and future directions in kinase drug discovery, reflecting on two decades of research since imatinib’s introduction.
- 4.Fabro F, Lamfers MLM, Leenstra S.. Advancements, challenges, and future directions in tackling glioblastoma resistance to small kinase inhibitors. Cancers (Basel). 2022;14:600. [DOI] [PMC free article] [PubMed] [Google Scholar]; **This paper addresses the specific challenge of glioblastoma resistance to kinase inhibitors, providing valuable insights into cancer therapy
- 5.Jucker EVB. On the long road to drug discovery. Trends Pharmacol Sci. 2002;23:147. [Google Scholar]
- 6.Burkert U, Allinger NL, Wold S, et al. Quantitative approaches to drug design. In: Westphal, U. Steroid-protein interactions II. Springer-Verlag; 1988. [Google Scholar]
- 7.Taylor SS, Keshwani MM, Steichen JM, et al. Evolution of the eukaryotic protein kinases as dynamic molecular switches. Philos Trans R Soc Lond B Biolog Sci. 2012:367:2517–2528. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Roskoski R. A historical overview of protein kinases and their targeted small molecule inhibitors. Pharmacol Res. 2015;100:1–23. [DOI] [PubMed] [Google Scholar]
- 9.Hari SB, Merritt EA, Maly DJ.. Sequence determinants of a specific inactive protein kinase conformation. Chem Biol. 2013;20(6):806–815. doi: 10.1016/j.chembiol.2013.05.005. [DOI] [PMC free article] [PubMed] [Google Scholar]; *This study focuses on the structural aspects of kinase conformation, which is critical for understanding inhibitor binding and specificity.
- 10.Roskoski R. Classification of small molecule protein kinase inhibitors based upon the structures of their drug-enzyme complexes. Pharmacol Res. 2016;103:26–48. [DOI] [PubMed] [Google Scholar]; *This reference offers a classification of kinase inhibitors based on their structural interactions, essential for understanding their mechanisms of action.
- 11.Wu P, Nielsen TE, Clausen MH.. FDA-approved small-molecule kinase inhibitors. Trends Pharmacol Sci. 2015;36(7):422–439. doi: 10.1016/j.tips.2015.04.005 [DOI] [PubMed] [Google Scholar]
- 12.Dar AC, Shokat KM.. The evolution of protein kinase inhibitors from antagonists to agonists of cellular signaling. Annu Rev Biochem. 2011;80(1):769–795. doi: 10.1146/annurev-biochem-090308-173656 [DOI] [PubMed] [Google Scholar]; *This review discusses the evolution of kinase inhibitors, providing historical context and future perspectives on their development.
- 13.Xing L, Klug-Mcleod J, Rai B, et al. Kinase hinge binding scaffolds and their hydrogen bond patterns. Bioorg Med Chem. 2015;23(19):6520–6527. doi: 10.1016/j.bmc.2015.08.006 [DOI] [PubMed] [Google Scholar]; **This paper provides detailed insights into the structural chemistry of kinase inhibitors, specifically focusing on hinge binding, a key aspect of kinase inhibitor design.
- 14.Swellmeen L, Shahin R, Al-Hiari Y, et al. Structure based drug design of Pim-1 kinase followed by pharmacophore guided synthesis of Quinolone-based inhibitors. Bioorg Med Chem. 2017;25(17):4855–4875. doi: 10.1016/j.bmc.2017.07.036 [DOI] [PubMed] [Google Scholar]; **This work on pharmacophore modeling for Pim-1 kinase inhibitors is crucial for understanding novel inhibitor discovery methodologies.
- 15.Shahin R, Swellmeen L, Shaheen O, et al. Identification of novel inhibitors for Pim-1 kinase using pharmacophore modeling based on a novel method for selecting pharmacophore generation subsets. J Comput Aided Mol Des. 2016;30(1):39–68. doi: 10.1007/s10822-015-9887-7 [DOI] [PubMed] [Google Scholar]
- 16.Berners-Lee CM. Cybernetics and forecasting. Nature. 1968;219(5150):202–203. doi: 10.1038/219202b0 [DOI] [Google Scholar]
- 17.Ivakhnenko AG. The group method of data handling in long-range forecasting. Technol Forecast Soc Change. 1978;12:213–227. [Google Scholar]
- 18.Abuhammad A, Taha MO.. QSAR studies in the discovery of novel type-II diabetic therapies. Expert Opin Drug Discov. 2016;11(2):197–214. doi: 10.1517/17460441.2016.1118046 [DOI] [PubMed] [Google Scholar]
- 19.Sadeghi F, Afkhami A, Madrakian T, et al. QSAR analysis on a large and diverse set of potent phosphoinositide 3-kinase gamma (PI3Kγ) inhibitors using MLR and ANN methods. Sci Rep. 2022;12(1):6090. doi: 10.1038/s41598-022-09843-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Bajusz D, Rácz A, Héberger K.. Chemical data formats, fingerprints, and other molecular descriptions for database analysis and searching. In: Comprehensive medicinal chemistry III. Elsevier, Inc.; 2017. p. 329–378. [Google Scholar]
- 21.Elbadawi M, Gaisford S, Basit AW.. Advanced machine-learning techniques in drug discovery. Drug Discov Today. 2021;26(3):769–777. doi: 10.1016/j.drudis.2020.12.003 [DOI] [PubMed] [Google Scholar]
- 22.Nippa DF, Atz K, Hohler R, et al. Enabling late-stage drug diversification by high-throughput experimentation with geometric deep learning. Nat Chem. 2024;16(2):239–248. doi: 10.1038/s41557-023-01360-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Banat R, Daoud S, Taha MO.. Ligand-based pharmacophore modeling and machine learning for the discovery of potent aurora A kinase inhibitory leads of novel chemotypes. Mol Divers. 2024;28(6):4241–4257. doi: 10.1007/s11030-024-10814-y [DOI] [PubMed] [Google Scholar]
- 24.Nag S, Baidya ATK, Mandal A, et al. Deep learning tools for advancing drug discovery and development. 3 Biotech. 2022;12(5):110. doi: 10.1007/s13205-022-03165-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Li L, Liu S, Wang B, et al. An updated review on developing small molecule kinase inhibitors using computer-aided drug design approaches. Int J Mol Sci. 2023;24(18):13953. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Min HY, Lee HY.. Molecular targeted therapy for anticancer treatment. Exp Mol Med. 2022;54:1670–1694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Ma J, Sheridan RP, Liaw A, et al. Deep neural nets as a method for quantitative structure-activity relationships. J Chem Inf Model. 2015;55(2):263–274. doi: 10.1021/ci500747n [DOI] [PubMed] [Google Scholar]
- 28.Kanev GK, Yaran ZA, Lbert JK, et al. Predicting the target landscape of kinase inhibitors using 3D convolutional neural networks. PLoS Comput Biol. 2023;19(9):e1011301. doi: 10.1371/journal.pcbi.1011301 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Khemani B, Patil S, Kotecha K, et al. A review of graph neural networks: concepts, architectures, techniques, challenges, datasets, applications, and future directions. J Big Data. 2024;11(1):18. doi: 10.1186/s40537-023-00876-4 [DOI] [Google Scholar]
- 30.Roskoski R. Rule of five violations among the FDA-approved small molecule protein kinase inhibitors. Pharmacol Res. 2023;191:106774. [DOI] [PubMed] [Google Scholar]
- 31.Zerrouk M, Er-Rajy M, Azzaoui K, et al. DFT computation-assisted design and synthesis of trisodium nickel triphosphate: Crystal structure, vibrational study, electronic properties and application in wastewater purification. J Mol Struct. 2025;1329:141450. doi: 10.1016/j.molstruc.2025.141450 [DOI] [Google Scholar]
- 32.Tropsha A, Gramatica P, Gombar VK.. The importance of being earnest: Validation is the absolute essential for successful application and interpretation of QSPR models. QSAR Comb Sci. 2003;22(1):69–77. doi: 10.1002/qsar.200390007 [DOI] [Google Scholar]
- 33.Li X, Li Z, Wu X, et al. Deep learning enhancing kinome-wide polypharmacology profiling: Model construction and experiment validation. J Med Chem. 2020;63(16):8723–8737. doi: 10.1021/acs.jmedchem.9b00855 [DOI] [PubMed] [Google Scholar]
- 34.Popova M, Isayev O, Tropsha A.. Deep reinforcement learning for de novo drug design. Sci Adv. 2018;4(7):eaap7885. doi: 10.1126/sciadv.aap7885 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Alhajeri MS, Luo J, Wu Z, et al. Process structure-based recurrent neural network modeling for predictive control: A comparative study. 2021. Available from: https://www.sciencedirect.com/science/article/pii/S0263876221005724.
- 36.Poso A. Molecular modeling of protein kinases: Current status and challenges. In: Laufer, S. (Ed.), Topics in Medicinal Chemistry. Springer Science and Business Media Deutschland GmbH; 2021. p. 25–41. [Google Scholar]
- 37.Li Z, Li X, Liu X, et al. KinomeX: A web application for predicting kinome-wide polypharmacology effect of small molecules. Bioinformatics. 2019;35(24):5354–5356. doi: 10.1093/bioinformatics/btz519 [DOI] [PubMed] [Google Scholar]
- 38.Gaulton A, Bellis LJ, Bento AP, et al. ChEMBL: A large-scale bioactivity database for drug discovery. Nucleic Acids Res. 2012;40(Database issue):D1100–D1107. doi: 10.1093/nar/gkr777 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Kim S, Chen J, Cheng T, et al. PubChem 2023 update. Nucleic Acids Res. 2023;51(D1):D1373–D1380. doi: 10.1093/nar/gkac956 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Berman HM, Westbrook J, Feng Z, et al. The Protein Data Bank. Nucleic Acids Res. 2000;28(1):235–42. http://www.rcsb.org/pdb/status.html. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wishart DS, Feunang YD, Guo AC, et al. DrugBank 5.0: A major update to the DrugBank database for 2018. Nucleic Acids Res. 2018;46(D1):D1074–D1082. doi: 10.1093/nar/gkx1037 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Kleandrova VV, Cordeiro MNDS.. Speck-Planche A. Perturbation theory machine learning model for phenotypic early antineoplastic drug discovery: design of virtual anti-lung-cancer agents. Appl Sci (Switzerland). 2024;14(20):9344. [Google Scholar]
- 43.Kleandrova VV, Speck-Planche A.. PTML modeling for pancreatic cancer research: In silico design of simultaneous multi-protein and multi-cell inhibitors. Biomedicines. 2022;10(2):491. doi: 10.3390/biomedicines10020491 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Kleandrova VV, Scotti MT, Scotti L, et al. Multi-target drug discovery via ptml modeling: applications to the design of virtual dual inhibitors of cdk4 and her2. Curr Top Med Chem. 2021;21(7):661–675. doi: 10.2174/1568026621666210119112845 [DOI] [PubMed] [Google Scholar]
- 45.Kleandrova VV, Scotti MT, Scotti L, et al. Cell-based multi-target QSAR model for design of virtual versatile inhibitors of liver cancer cell lines. SAR QSAR Environ Res. 2020;31(11):815–836. doi: 10.1080/1062936X.2020.1818617 [DOI] [PubMed] [Google Scholar]
- 46.Levitzki A. Tyrosine kinase inhibitors: Views of selectivity, sensitivity, and clinical performance. Annu Rev Pharmacol Toxicol. 2013;53:161–185. doi: 10.1146/annurev-pharmtox-011112-140341 [DOI] [PubMed] [Google Scholar]
- 47.Sprous DG, Zhang J, Zhang L, et al. Kinase inhibitor recognition by use of a multivariable QSAR model. J Mol Graph Model. 2006;24(4):278–295. doi: 10.1016/j.jmgm.2005.09.004 [DOI] [PubMed] [Google Scholar]
- 48.The Nobel Assembly at the Karolinska Institute . The nobel prize in physiology or medicine 1992. 1992. https://www.nobelprize.org/prizes/medicine/1992/press-release/.
- 49.Fischer E. Nobel-winning biochemist who discovered a ubiquitous cell-regulatory mechanism. Nature. 2021;597. [Google Scholar]
- 50.Hu H. Recent discovery and development of selective protein kinase C inhibitors. Drug Discov Today. 1996;1(10):438–447. [Google Scholar]
- 51.Shahin R, Shaheen O, El-Dahiyat F, et al. Research advances in kinase enzymes and inhibitors for cardiovascular disease treatment. Future Sci OA. 2017;3(4):FSO204. doi: 10.4155/fsoa-2017-0010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Takai Y, Kishimoto A, Inoue M, et al. Studies on a cyclic nucleotide-independent protein kinase and its proenzyme in mammalian tissues I. Purification and characterization of an active enzyme from bovine cerebellum*. J Biol Chem. 1977;252:7603–7609. [PubMed] [Google Scholar]
- 53.Nishiyama K, Katakami H, Yamarzura H, et al. Functional specificity of guanosine 3’:5-monophosphate-dependent and adenosine 3’:5’-monophosphate-dependent protein kinases from silkworm*. J Biol Chem. 1975;250:1297–1300. [PubMed] [Google Scholar]
- 54.Inoue M, Kishimoto A, Takai Y, et al. Studies on a cyclic nucleotide-independent protein kinase and its proenzyme in mammalian tissues. II. Proenzyme and its activation by calcium-dependent protease from rat brain. J Biol Chem. 1977;252(21):7610–7616. [PubMed] [Google Scholar]
- 55.Druker B, Tamura S, Buchdungerz E, et al. Effects of a selective inhibitor of the Abl tyrosine kinase on the growth of Bcr-Abl positive cells. 1996. Available from: http://www.nature.com/naturemedicine. [DOI] [PubMed]
- 56.Panjarian S, Iacob RE, Chen S, et al. Structure and dynamic regulation of Abl kinases. J Biol Chem. 2013;288(8):5443–5450. doi: 10.1074/jbc.R112.438382 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Winter GE, Rix U, Carlson SM, et al. Systems-pharmacology dissection of a drug synergy in imatinib-resistant CML. Nat Chem Biol. 2012;8(11):905–912. doi: 10.1038/nchembio.1085 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Manallack DT, Pitt WR, Gancia E, et al. Selecting screening candidates for kinase and G protein-coupled receptor targets using neural networks. J Chem Inf Comput Sci. 2002;42(5):1256–1262. doi: 10.1021/ci020267c [DOI] [PubMed] [Google Scholar]
- 59.Sprous DG, Palmer RK, Swanson JT, et al. QSAR in the pharmaceutical research setting: QSAR models for broad, large problems. 2010. [DOI] [PubMed] [Google Scholar]
- 60.Kamath S, Buolamwini JK.. Receptor-guided alignment-based comparative 3D-QSAR studies of benzylidene malonitrile tyrphostins as EGFR and HER-2 kinase inhibitors. J Med Chem. 2003;46(22):4657–4668. doi: 10.1021/jm030065n [DOI] [PubMed] [Google Scholar]
- 61.Tropsha A, Golbraikh A.. Predictive QSAR modeling workflow, model applicability domains, and virtual screening. Curr Pharm Des. 2007;13(34):3494–3504. doi: 10.2174/138161207782794257 [DOI] [PubMed] [Google Scholar]
- 62.Shahin R, Alqtaishat S, Taha MO.. Elaborate ligand-based modeling reveal new submicromolar Rho kinase inhibitors. J Comput Aided Mol Des. 2012;26(2):249–266. doi: 10.1007/s10822-011-9509-y [DOI] [PubMed] [Google Scholar]
- 63.Tropsha A. Best practices for QSAR model development, validation, and exploitation. Mol Inform. 2010;29(6-7):476–488. doi: 10.1002/minf.201000061 [DOI] [PubMed] [Google Scholar]
- 64.Golbraikh A, Tropsha A.. Beware of q2!. J Mol Graph Model. 2002;20(4):269–276. doi: 10.1016/s1093-3263(01)00123-1 [DOI] [PubMed] [Google Scholar]
- 65.Sosman JA, Kim KB, Schuchter L, et al. Survival in BRAF V600-mutant advanced melanoma treated with vemurafenib. N Engl J Med. 2012;366:707–714. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Sadybekov AV, Katritch V.. Computational approaches streamlining drug discovery. Nature. 2023;616:673–685. [DOI] [PubMed] [Google Scholar]
- 67.Clark K, Luong M-T, Le QV, et al. ELECTRA: pre-training text encoders as discriminators rather than generators. 2020. arXiv:2003.10555v1 [cs.CL] [Google Scholar]
- 68.Cichońska A, Ravikumar B, Allaway RJ, et al. Crowdsourced mapping of unexplored target space of kinase inhibitors. Nat Commun. 2021;12(1):3307. doi: 10.1038/s41467-021-23165-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Attwood MM, Fabbro D, Sokolov AV, et al. Trends in kinase drug discovery: targets, indications and inhibitor design. Nat Rev Drug Discov. 2021;20:839–861. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




