Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 9;45(8):e70047. doi: 10.1002/minf.70047

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Amit Kumar Srivastav 1,2,, Unnati Modi 3, Rahul Kumar 4, Dhiraj Bhatia 5,, Raghu Solanki 5,
PMCID: PMC13454516  PMID: 42572360

Abstract

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

Keywords: drug design, drug discovery, geometric deep learning, small molecules


Graph deep learning (GDL)‐based framework for structure‐based drug design (SBDD), illustrating how three‐dimensional molecular geometry, topology, and physicochemical features are integrated to accurately model protein–ligand interactions and accelerate drug discovery.

graphic file with name MINF-45-e70047-g001.jpg

1. Introduction

Geometric deep learning (GDL) is an emerging subfield of machine learning that extends deep neural network (DNN) architectures to non‐Euclidean domains such as graphs, manifolds, and point clouds [1]. GDL is fundamentally based on the idea of constructing models that are invariant or equivariant to geometric transformations, including translations, rotations, and reflections, qualities needed to precisely depict molecular interactions. The tiny molecules interact with their biological targets depending on maintaining this geometric information, since biomolecular structures, by nature, exist in three‐dimensional (3D) space [2].

Whereas conventional deep learning (DL) models run on fixed grid‐based or sequence‐based data, GDL models depict molecules as graphs in which atoms are handled as nodes and bonds or interactions as edges. Spatial coordinates and chemical descriptors enrich these graphs so that graph neural networks (GNNs), SE(3) equivariant networks, and geometric attention mechanisms may learn from both topological and 3D geometric properties of protein‐ligand complexes [3] (Figure 1). This DL model could capture intricate interaction patterns depending on the exact molecular geometry. Such designs are well‐suited for a broad spectrum of structure‐based drug design (SBDD), including binding affinity prediction, posture estimation, and lead optimization [4]. The fundamental method in SBDD uses the 3D structure of biological targets, usually proteins or nucleic acids, to direct the creation and optimization of small‐molecule treatments. SBDD helps to identify ligands that might interact with targets with great specificity and potency by using thorough structural knowledge of active or allosteric binding sites [5].

FIGURE 1.

FIGURE 1

Representations of various geometric deep learning (GDL) models for molecular structure analysis.

Traditional SBDD methods typically include molecular docking, scoring systems, and molecular dynamics (MD) simulations. Molecular Docking generally predicts how a ligand fits into the target binding site, while scoring systems estimate how strongly the ligand is likely to bind based on simplified physical and chemical models. Whereas MD simulations provide insights into the conformational flexibility and stability of the protein‐ligand complex [6]. These methods are sometimes hampered, nevertheless, by critical constraints like reliance on heuristic scoring systems, difficulty accounting for molecular flexibility and induced fit, and large computing costs [78]. These limitations have stimulated the development of more scalable, data‐driven, and high‐precision alternatives.

To address these limitations, DL has emerged as a transformative approach within the drug discovery paradigm. Early uses of DL in SBDD analyzed 3D voxelized representations of protein–ligand interactions using convolutional neural networks (CNNs), therefore treating molecules as 3D images [9]. Outstanding on benchmark datasets, these models showed potential, including binding affinity prediction and posture classification. Still, voxel‐based models have a major trade‐off. CNNs lack natural geometric awareness; discretization of 3D space can result in information loss and significant memory demand. Crucially, in biological systems, molecule symmetries and rotations were not handled by them spontaneously. These models thus have an intrinsically limited capacity to generalize across many conformations or binding environments [10].

Previous reviews on SBDD have summarized classical computational strategies and discussed various GDL approaches developed for modeling 3D macromolecular structures in rational drug discovery [11, 12, 13]. Literature has also provided comprehensive analysis of key SBDD tasks, including protein pocket modeling, binding affinity prediction, linker design, de novo molecule generation, binding site identification, and binding‐pose prediction [14]. The present review provides a broader and more application‐oriented perspective by emphasizing the role of GDL across the complete structure‐based drug discovery pipeline. In particular, this review highlights recent advances in lead optimization, geometry‐aware generative models, multimodal protein representations, and the integration of GDL with experimental structural biology and computational workflows. Practical benchmarking strategies, open‐source software resources, and future research directions are discussed to provide a comprehensive reference for both computational scientists and medicinal chemists. Furthermore, we discuss how GDL‐based frameworks can be integrated with experimental workflows, highlight persistent challenges, and delineate future opportunities through which geometry‐aware AI may substantially accelerate small‐molecule drug discovery.

2. Fundamentals of GDL for Molecular Data

The effectiveness of GDL in drug development was largely determined by how chemical structures were represented and processed by neural networks. Unlike conventional DL, which usually manages fixed‐dimensional, grid‐like data (e.g., images or sequences), GDL runs directly on non‐Euclidean structures such as graphs, point clouds, and manifolds, allowing models to learn complex spatial and relational information inherent in molecular systems [12].

2.1. Representations of Molecular Structures for GDL

2.1.1. Graph‐Based Representations

Graph structures, whereby molecules were represented as mathematical graphs G (V, E), are among the most often used and natural forms for molecular depiction in GDL. Under this paradigm, edges ε E show chemical bonds or spatial interactions like hydrogen bonds, π–π stacking, or proximity‐based contact between residues or ligands and proteins; nodes v i ε V correspond to atoms or amino acid residues [15]. In atomistic graphs, every atom is a node; covalent bonds constitute the edges. While edge features record bond type, length, and angle information, node characteristics often contain atomic number, hybridization state, partial charge, and aromaticity. Such graph‐based representations are particularly well‐suited for small molecule modeling [16].

While mostly applied in protein modeling, residue‐level graphs abstract every amino acid residue as a node. Spatial closeness or biochemical interactions, e.g., salt bridges or hydrophobic contacts, may define edges. This form helps to simplify the structural complexity of proteins and supports long–range interaction modeling [17].

Graph‐based models such as GNNs and message passing neural networks (MPNNs) are essential for predicting molecular properties and drug–target interactions. These models demonstrate exceptional capability in learning local chemical environments, interatomic interactions, and topological motifs. Moreover, their ability to flexibly integrate both 2D and 3D molecular information makes them highly suitable for diverse drug discovery applications [18].

2.1.2. Geometric Graph Representations

While conventional graph‐based representations primarily capture molecular topology through atom–bond connectivity. Geometric graph representations extend this paradigm by embedding explicit 3D geometric information such as directionality and orientation directly into node and edge features. These models combine the relational strength of GNNs with the spatial awareness of geometric tensors, allowing learning algorithms to reason over both chemical connectivity and spatial geometry simultaneously.

In geometric graphs, atomic or residue nodes are represented by scalar features (e.g., atomic number, hybridization, charge) and vector‐valued geometric features that encode directional information such as bond orientations or interatomic displacement vectors. The message passing and aggregation operations were designed to be rotation‐ and translation‐equivariant, ensuring that molecular predictions remain consistent regardless of coordinate frame.

Several state‐of‐the‐art protein and ligand encoders were built on this geometric graph principle:

2.1.2.1. Geometric Vector Perceptron (GVP‐GNN)

Introduced by Jing et al. (2021), this model integrates scalar and vector channels to represent atomic environments while maintaining SE(3)‐equivariance. It has been widely adopted for modeling protein structures, residue–residue interactions, and ligand‐protein complexes, outperforming traditional scalar‐only GNNs in structure‐based learning tasks [19].

2.1.2.2. GearNet

Proposed by Zhang et al. (2023), GearNet enhances GNN message passing with multi‐scale geometric features and relational embeddings that capture both chemical bonding and spatial neighborhood geometry. Variants such as GearNet‐ESM combine graph‐based encoders with protein language models (PLMs), enabling multimodal learning that leverages both structural and sequence‐derived context [20].

2.1.2.3. ProNet (Protein GNN)

A recent framework that represents proteins as geometric graphs where edges encode both distance and orientation‐based constraints between residues. ProNet's hierarchical aggregation mechanism improves long‐range dependency modeling and facilitates accurate binding site characterization and protein–ligand interface prediction.

2.1.2.4. Coordinate‐Based Distance Convolution (CDC)

A geometric operator that replaces standard graph convolutions with coordinate‐aware kernels, allowing the model to directly process interatomic distances and angular relationships within local neighborhoods. CDC‐based models can more faithfully reproduce anisotropic interactions such as hydrogen bonding and π–π stacking.

By combining geometric tensors with graph connectivity, these approaches provide a unified framework for representing both molecular topology and 3D geometry. They are increasingly viewed as the default representation for modern GDL pipelines in SBDD, bridging the gap between purely topological graphs and continuous spatial representations like point clouds or meshes.

2.1.3. Point Cloud Representations

Another geometric view models molecules as point clouds, unordered sets of 3D coordinates matching the atomic positions in space. Every point in this representation could have further associations with atomic type, charge, or partial density. Point clouds retain the exact spatial layout of atoms without imposing a specified connectedness, unlike graph representations [18].

All of which are fundamental in determining binding poses and energetics, point cloud‐based models shine in capturing fine‐grained spatial distributions, noncovalent interactions (NCI), and molecule surface characteristics. This representation closely matches the physical reality of molecular structures, so DL architectures can learn straight from the 3D geometry without losing continuity via discretization (as happens in voxel‐based CNNs) [21].

Models such as PointNet, PointNeXt, and SE(3)‐equivariant networks have recently been adapted for molecular point clouds by including rotation and translation equivariance to provide strong modeling of molecular conformations. This is particularly advantageous for studies focused on protein‐ligand docking, pose prediction, and binding pocket characterization, where orientation‐invariant learning is critically important for accurate modeling and interpretation [22].

2.1.4. Mesh‐Based Representations

Though less extensively utilized than graphs and point clouds, mesh‐based representations have become a useful tool for some purposes, especially in representing protein surfaces, binding pockets, or ligand structures. Molecular surfaces are triangulated in this presentation into meshes made of vertices, edges, and faces, making a continuous manifold [23].

In capturing surface topology, concavity/convexity of binding pockets, and electrostatic potential distribution over the protein surface, mesh‐based representations show special benefits. In ligand recognition, specificity, and binding strength, these qualities can be rather important [24]. By learning geometric and biological patterns directly on the surface mesh, tools such as molecular surface interaction fingerprinting (MaSIF) and DeepSurf have analyzed protein–protein and protein–ligand interactions using mesh‐based GDL. This method offers a higher‐level abstraction of molecular structure and interaction readiness, therefore complementing atomistic and point cloud models even if computationally more demanding [25]. Furthermore, Table 1 summarizes the principal molecular representations used across SBDD.

TABLE 1.

Comparison of molecular representations used in geometric deep learning.

Representation type Input format Key models/ architectures Advantages Limitations Applications in SBDD Computational complexity Training dataset size Open‐source availability
Graph‐based Node and edge feature matrices; optional 3D coordinates. GNN, MPNN, SchNet, DimeNet, GemNet, GAT, GVP‐GNN, GearNet, ProNet. Captures topological and chemical connectivity; flexible integration of 2D and 3D info; efficient message passing. Limited handling of global 3D geometry; may lose spatial continuity; sensitive to bond definitions. Molecular property prediction, protein–ligand interaction modeling, binding affinity prediction. Moderate (efficient message passing; scales approximately with number of nodes and edges) Medium to Large (≈103–106 molecules;, e.g., MoleculeNet, PDBbind, QM9, BindingDB) Yes (PyTorch Geometric, DGL, Deep Graph Library, official GitHub implementations)
Geometric Graph Scalars (atom/residue attributes) + vectors (relative orientations, coordinates). GVP‐GNN, GearNet, CDC, ProNet, EGNN, SE(3)‐Transformer. Encodes both chemical topology and 3D geometry; rotation/translation equivariant; improved interpretability. Higher computational cost; requires consistent coordinate data. Protein–ligand interface modeling, structure‐based learning, pose prediction. High (equivariant operations increase computational cost) Large (typically ≥104–106 structures; PDBbind, CrossDocked, Protein Data Bank, QM9) Yes (e3nn, PyTorch Geometric, TorchMD‐Net, EGNN and SE(3)‐Transformer repositories)
Point Cloud 3D coordinates + atom‐level descriptors (charge, type, radius). PointNet, PointNet++, DiffDock, SE(3)‐equivariant point models. Preserves full 3D spatial arrangement; captures noncovalent interactions; naturally rotation‐invariant. Lacks explicit chemical bonding; limited local context; may need large datasets. Binding pose prediction, binding pocket characterization, molecular docking. High (neighborhood search and hierarchical feature aggregation) Medium to Large (≈103–105 protein–ligand complexes; PDBbind, CrossDocked, AlphaFold‐derived datasets) Yes (PointNet/PointNet++, Open3D‐ML, PyTorch implementations, DiffDock repository)
Mesh‐based Mesh of surface vertices with curvature, normal, and electrostatic attributes. MaSIF, DeepSurf, NeurIPS–MeshCNN. Represents surface topology and electrostatics; ideal for studying interaction sites and recognition. Computationally expensive; limited availability of high‐quality mesh data. Protein–protein and protein–ligand interface prediction, shape complementarity analysis. Very High (surface mesh construction and geometric convolutions are computationally intensive) Medium (≈103–104 high‐quality protein surface meshes; MaSIF, Protein Data Bank‐derived datasets) Limited (MaSIF, MeshCNN, selected research repositories)
Voxel / Grid‐based 3D tensor of features (x, y, z, channels). AtomNet, GNINA, 3D‐CNN, DeepDock. Compatible with CNN architectures; captures volumetric interactions. High memory cost; loss of fine‐grained geometry; sensitive to grid resolution. Binding affinity prediction, virtual screening, scoring function learning. Very High (memory‐intensive 3D convolutions) Large (typically ≥104–106 voxelized protein–ligand complexes; PDBbind, DUD‐E, CrossDocked) Yes (GNINA, AtomNet‐related implementations, DeepChem, TensorFlow/PyTorch frameworks)

2.1.5. Protein versus Ligand Encoders and Multimodal Representations

Accurately modeling protein–ligand interactions requires distinct encoding strategies for proteins and ligands, as the two molecular systems differ fundamentally in size, structural hierarchy, and geometric complexity. Ligands are typically small, rigid or semi‐flexible molecules that can be effectively represented using atom‐level molecular graphs, 3D point clouds, or distance‐aware geometric graphs. GNN‐ and MPNN‐based ligand encoders operate primarily on local chemical neighborhoods, capturing atom types, bond orders, hybridization patterns, and fine‐grained 3D conformational features. These representations are well suited for property prediction, conformer scoring, and ligand‐based virtual screening.

In contrast, proteins exhibit complex tertiary and quaternary organization, long‐range residue contacts, and backbone geometry defined by dihedral angle relationships. Representing proteins therefore requires specialized geometric encoders that can integrate structural features beyond simple atomic connectivity. Recent advances have introduced protein‐focused architectures such as GVP‐GNN, GCPNet, CDC, GearNet, and ProNet, which incorporate vector features, local reference frames, residue‐centric graphs, and SE(3)‐equivariance to capture orientation‐dependent interactions crucial for binding site recognition. These protein encoders outperform generic GNNs by explicitly modeling the geometric and biophysical principles that govern protein folding, surface topology, and interaction hotspots.

A major trend in current SBDD is the development of multimodal protein representations, which combine complementary sources of information to create richer, more expressive encoders. One prominent direction integrates structural geometric graphs with PLMs derived from large‐scale sequence data, as demonstrated in GearNet‐ESM and bidirectional hierarchical protein multimodal encoders. These hybrids leverage evolutionary conservation, contextual sequence embeddings, and structural geometry simultaneously, improving generalization in low‐data settings. Another direction combines graph‐based atomistic or residue‐level encodings with surface‐derived features, as seen in AtomSurf, S3F, and other surface‐aware GDL models. By incorporating curvature, electrostatics, and local surface patches, these models better capture the physicochemical determinants of molecular recognition.

Together, these dedicated protein encoders and multimodal frameworks represent a significant evolution beyond traditional graph‐only or sequence‐only representations. They provide a more accurate and biologically grounded foundation for modeling protein–ligand interactions, exploring binding‐pocket geometry, and improving downstream tasks such as docking, affinity prediction, and interface classification.

2.2. Principal GDL Models/Architectures

GDL is a broad spectrum of architectures intended especially to interpret chemical data in non‐Euclidean domains. Geometric transformations, node connections, and spatial interactions are handled differently across these models. Several important architectures have become rather effective tools for predicting molecular characteristics, binding affinities, and optimizing drug candidates in the framework of SBDD [26] (Figure 2).

FIGURE 2.

FIGURE 2

GDL‐based models for 3D structure‐based drug design.

2.2.1. Geometric Symmetries in Molecular DL: Equivariance and Invariance

The remarkable success of GDL in molecular modeling stems from its ability to explicitly incorporate the geometric symmetries inherent in 3D molecular systems. Unlike conventional DL methods, which often depend on coordinate‐specific representations, GDL architecture is designed to respect the fundamental symmetry properties of molecules under spatial transformations. This enables physically meaningful, data‐efficient, and generalizable learning [27]. The geometric transformations relevant to molecular systems are described by the Euclidean group (E(n)) in Equation (1), which consists of all distance‐preserving transformations in an (n)‐dimensional Euclidean space:

E(n)={(R,t)RO(n),;tR{n}} (1)

where (R) represents an orthogonal transformation (rotation or reflection) and (t) denotes a translation vector [28]. In three‐dimensional molecular modeling, the Special Euclidean group (SE(3)) is of particular importance and is defined as

SE(3)=SO(3){3} (2)

where (SO(3)) denotes the group of three‐dimensional rotations and R{3} represents translations. In Equation (2), unlike (E(3)), the (SE(3)) group excludes reflections and therefore preserves the handedness (chirality) of molecular structures [2728].

Two fundamental concepts govern the design of geometry‐aware neural networks, i.e., invariance and equivariance. In Equation (3), A function (f) is said to be invariant if its output remains unchanged under a geometric transformation

f(gx)=f(x)gG (3)

where (g) represents any transformation belonging to the symmetry group. Molecular properties such as binding affinity, molecular energy, toxicity, ADMET characteristics, and molecular class are expected to be invariant because their physical values do not depend on the orientation or position of the molecule in space [28].

In contrast, a function is equivariant when its output transforms consistently with the input,

f(gx)=gf(x)gG (4)

Equivariant learning is particularly important for predicting geometric quantities such as atomic coordinates, force vectors, dipole moments, and molecular conformations, where rotating or translating a molecular structure should produce an equivalent transformation in the predicted output as defined in Equation (4) [28].

Respecting these symmetry principles is essential for physically consistent molecular modeling. For example, rotating a protein–ligand complex by (90^\circ) does not alter its binding affinity or free energy, which are invariant properties. However, the atomic coordinates and force vectors must rotate by the same angle, demonstrating equivariant behavior. Incorporating these constraints into neural network architectures reduces redundant learning, improves sample efficiency, enhances model generalization, and ensures physically meaningful predictions [29].

Modern GDL architecture explicitly incorporates these symmetry principles. GNNs capture molecular connectivity through message passing, whereas E(n)‐equivariant (EGNNs), SE(3)‐Transformers, Tensor Field Networks (TFNs), and GVPs extend this framework by preserving rotational and translational symmetries during feature propagation. These architectures have demonstrated superior performance in molecular property prediction, protein–ligand interaction modeling, binding affinity prediction, molecular docking, and three‐dimensional structure generation because they naturally align with the geometric principles governing molecular systems [30].

The incorporation of geometric symmetries through invariant and equivariant learning constitutes one of the defining characteristics of modern GDL and provides the theoretical foundation for accurately modeling complex molecular structures and interactions in structure‐based drug discovery. To facilitate understanding of symmetry‐aware learning in molecular modeling, Table 2 compares the defining characteristics, mathematical formulations, representative molecular properties, corresponding GDL architectures, and typical SBDD applications of invariant and equivariant representations.

TABLE 2.

Distinction between invariant and equivariant properties in deep molecular learning.

Aspect Invariant properties Equivariant properties
Definition Output remains unchanged under geometric transformations (rotation, translation, or reflection). Output transforms consistently with the input geometric transformation.
Mathematical expression (f(gx) = f(x)) (f(gx) = g f(x))
Output type Scalar quantities Vector‐ or tensor‐valued quantities
Physical interpretation Physical property is independent of the molecule's orientation or position in space. Geometric output changes in the same manner as the molecular coordinates.
Representative molecular properties Binding affinity, molecular energy, ADMET properties, toxicity prediction, solubility, bioactivity, molecular property prediction Atomic coordinates, force vectors, dipole moments, molecular conformations, protein–ligand binding poses, electron density orientation
Representative GDL architectures SchNet, DimeNet, GemNet, Graph Neural Networks (GNNs), Graph Attention Networks (GATs) E(n)‐Equivariant Graph Neural Networks (EGNNs), SE(3)‐Transformers, Tensor Field Networks (TFNs), Geometric Vector Perceptron Networks (GVP‐GNNs)
Typical applications in SBDD Binding affinity prediction, virtual screening, ADMET prediction, lead prioritization, molecular property prediction Protein–ligand docking, binding‐pose prediction, conformational modeling, protein structure refinement, molecular dynamics and force prediction

2.2.2. GNNs

GNNs naturally allow the analysis of data structured as graphs, a natural abstraction for chemical and biomolecular systems. GNNs form the pillar of GDL. Within the framework of SBDD, molecules and protein‐ligand complexes are sometimes shown as graphs whereby chemical bonds or spatial interactions are seen as edges, and atoms or residues correspond to nodes. Fundamental to precisely capturing molecular action, this structure enables the modeling of both topological and spatial relations [31].

The MPNN framework is one very well‐known class of GNNs in drug development. Through local information exchange with their neighbors, MPNNs update node representations, therefore learning how atoms inside a molecule or residues within a binding pocket interact either directly or indirectly [32]. MPNNs use two main actions carried out at every layer of the network in their message‐passing process. First, depending on their current feature vectors and the properties of the connecting edge, a message is generated between adjacent nodes, maybe encoding bond type, interatomic distance, or contact strength. A state update follows in which every node compiles the incoming messages and utilizes them to hone its own representation [33].

For a node vi , the message‐passing mechanism at the t‐th iteration is defined as:

mi(t)=jϵN(i)M(t)(hi(t),hj(t),eij) (5)
hi(t+1)=U(t)(hi(t),mi(t)) (6)

Here, hi(t)ϵRd represents the feature vector of node i at layer t, e ij denotes the edge feature between nodes i and j, N(i) is the neighborhood of node i, and M(t) and U(t) are learnable functions, typically implemented as multilayer perceptrons (MLP). This formulation facilitates the capturing of both short‐ and long‐range dependencies by allowing for the variable recording of local chemical environments and propagating information along the molecular graph [34].

Many architectural developments inside the MPNN framework have greatly enhanced the value of the framework in drug discovery. Bhatti, U. A. et al. (2023) discussed the simplified learning process of graph convolutional networks (GCNs) by sending messages using normalized adjacency matrices, therefore preserving necessary connectivity information [35]. Graph attention networks (GATs) improve this method by adding dynamic weighting of the relevance of surrounding nodes during aggregation, therefore allowing the model to select more powerful interactions. Using continuous‐filter convolutions based on interatomic distances, architectures such as SchNet allow direct learning from 3D molecule geometries to extend the message transmission paradigm into continuous space [36]. Klicpera, J. et al. (2021) have worked on more complex models, including DimeNet and GemNet use angular features θijk between linked atoms (i, j, k) in the message function, therefore enabling the simulation of direction‐dependent interactions, including π‐stacking and hydrogen bonding. These developments raise the expressiveness and accuracy of GNNs in estimating quantum mechanical characteristics, binding affinities, and other biophysically relevant metrics taken together [37].

Wu et al. (2023) proposed a technique called substructure mask explanation (SME), which builds on well‐established molecular segmentation methods to provide interpretations aligned with domain‐specific chemical knowledge [38]. SME is employed to elucidate how GNNs learn to predict key molecular properties such as aqueous solubility, genotoxicity, cardiotoxicity, and blood–brain barrier permeability. By highlighting model inconsistencies, SME assists chemists in optimizing molecular structures for desired properties while offering interpretable insights consistent with their expert understanding. From atomic nodes to higher‐order subgraphs reflecting pieces or domains, such networks run hierarchically, learning representations at several levels of abstraction. Similar to the word embeddings in natural language processing, molecules are broken down in models such as Mol2Vec and GraphVAE into constituent fragments whose embeddings are learnt in a data‐driven way [39].

Harren, T. et al. applied and compared various explainable AI (XAI) techniques to lead optimization datasets with well‐characterized structure–activity relationships (SARs) and available X‐ray crystal structures. The study demonstrates that integrating deep DNN models with powerful interpretation methods provides clear and comprehensive insights [40].

VirtuDockDL, a Python‐based web platform developed by Noor F. et al. leverages DL for drug discovery. The pipeline utilizes GNN to analyze and predict the therapeutic potential of various chemical compounds (Figure 3). In conclusion, these GNN‐based models direct reasonable drug design and scaffold optimization by spotting and stressing the most pharmacologically relevant segments. In lead optimization, this is especially helpful since small changes to molecular substructures can have a major impact on potency, selectivity, and pharmacokinetics [41].

FIGURE 3.

FIGURE 3

Schematic representation of VirtuDockDL pipeline virtual screening for drug discovery. The process begins with identifying active and inactive compounds, generating de novo molecules, applying drug‐likeness filters, and extracting graph‐based features. This GNN model is then trained and evaluated using metrics such as ROC curves, followed by screening of compound libraries. Protein structures are prepared for molecular docking, and the results are compared with experimental data. Figure is Reproduced (adapted) with permission from [41], Copyright 2024, The Author(s), published in Scientific Reports.

2.2.3. 3D Convolutional Neural Networks (3D CNNs)

3D CNNs represent one of the foundational GDL architectures successfully applied to structure‐based drug discovery (Figure 4a). These models expand conventional 2D convolutional networks into three dimensions, enabling their processing of volumetric molecular representations explicitly encoding the spatial arrangement of atoms and their physicochemical characteristics in Cartesian space. Unlike graph‐ or point cloud‐based models running on abstracted or sparse representations, 3D CNNs operate on dense, structured input grids called voxels, which provide a discretized approximation of molecular geometry and interaction contexts [44].

FIGURE 4.

FIGURE 4

Representative examples of GDL models. Schematic representation of convolutional neural networks (a), figure adapted with permission from [42], Copyright 2017 American Chemical Society. Structure of Graphormer‐based graph contrastive learning method (b), Figure is reproduced (adapted) wih permission from [43], Copyright 2024 Elsevier.

In the study by Jiang, H. et al. (2020), a common approach used to represent the protein–ligand combination is embedded in a cubic grid of dimension d×d×d, whereby each voxel v ijk corresponds to a discrete subvolume in 3D space [45]. Where the channels may encode atomic type, partial charge, hydrophobicity, hydrogen bond donor/acceptor status, or other pertinent descriptors, each voxel is linked with a multichannel feature vector f ijk ∈ R c . This tensor V ∈ Rd×d×c serves as the input to the 3D CNN, which then applies a series of learnable convolutional filters W ∈ Rk×k×k×c performing local operations that extract hierarchical spatial patterns [45].

At a given voxel (i, j,k), the 3D convolution process in the output volume is specified mathematically as:

Oi,j,k=u=1kv=1kw=1kc=1CWu,v,w,cVi+u,j+v,k+w,c (7)

where the summation spans the kernel dimensions and feature channels, and W represents the convolutional kernel of size k. As the network depth rises, a set of activation maps encoding ever more abstract representations of molecule shape and interactions results. 3D CNNs learn a spatial hierarchy of interaction features from low‐level atomistic contacts to higher‐order structural motifs and functional binding interfaces by stacking many convolutional layers with pooling and nonlinear activation functions [46].

Further, Markus, B. et al. (2023) demonstrated that volumetric depiction of spatial interactions was crucial. In the study of protein binding pockets, for instance, 3D CNNs can record minute characteristics of the pocket shape and electrostatics supporting ligand complementarity and specificity. The model learns discriminative spatial patterns in ligand posture prediction that distinguish accurately from erroneous binding conformations. In the prediction of protein–ligand binding affinities, 3D CNNs can similarly be trained to transfer spatially dispersed interaction information to a scalar estimate of binding strength, therefore acting as learned scoring functions [47].

Moreover, the development of several noteworthy 3D CNN models for structure‐based drug discovery is in progress. Among the first DL models to directly predict bioactivity from the 3D structure of protein‐ligand complexes by using 3D convolutions was AtomNet [48]. Wallach I. et al. (2015) showed that by learning elements unavailable to hand‐engineered scoring systems that DNNs could outperform conventional docking‐based approaches in virtual screening tasks. To enhance binding mode classification, GNINA combined 3D CNNs with conventional docking software employing convolutional layers to re‐score docked poses [48].

2.2.4. Geometric Transforms

The transformer architecture was originally designed for natural language processing and has revolutionized machine learning through its self‐attention mechanisms, which enable effective modeling of long‐range relationships. Its scalability and adaptability have motivated its use in vision, protein structure prediction, and molecular modeling. Within the framework of GDL, geometric transformers are a strong class of architecture that combines self‐attention with geometric priors, enabling their effective processing of non‐Euclidean data, including molecular graphs, 3D point clouds, and spatially embedded atomic networks [4950] (Figure 4b).

Geometric transformers could simulate global context by attending to all node pairs inside a molecular structure, unlike conventional GNNs that mostly concentrate on local neighborhoods by message transmission. In the study, E P. B. et al. demonstrates that in molecular systems where distant atoms or residues, such as those engaged in allosteric control or long‐range electrostatic interactions, can greatly affect binding affinity or conformational stability. In these models, each node i computes a weighted sum of the features from every other node j in the graph or spatial structure using attention weights calculated by a compatibility function, these models’ fundamental operation is based on:

αij=softmaxj((WQhi)T(WKhj)dk) (8)
zi=jαijWVhj (9)

where h i  and h j  are the input feature vectors for nodes i and j. W Q, W K, W V  are learnable projection matrices for queries, keys, and values; d k  is the dimensionality of the key vectors used for scaling; this formulation allows each node to aggregate information from all others, independent of topological distance [51].

2.2.4.1. Graph Transformers

Graph transformers represent a new class of neural network architecture or geometric transformers designed to extend the expressive power of transformers to graph‐structured data [52]. By integrating global attention mechanisms with inherent graph topology, these models overcome key limitations of traditional GNNs that rely on local neighborhood aggregation and impose strong structural inductive biases [53]. Inspired by the transformers in natural language processing and computer vision, graph transformers leverage attention interpretability in complex relational systems. Recent studies demonstrated their strong performance across a wide range of graph learning tasks, highlighting their versatility and growing utility in domains such as molecular modeling, protein design, and network analysis [54].

2.2.4.2. Geometric (SE(3)‐Equivariant) Transformers

Since physical systems are invariant to the choice of coordinate frame, the development of equivariant transformers, which guarantees consistent model outputs under spatial transformations such as rotations and translations, is a significant advancement in this field. This property, known as SE(3)‐equivariance, was especially important in molecular modeling. Designing attention and message‐passing procedures that are equivariant with regard to the SE(3) Lie group, the group of 3D rotations and translations, the SE(3)‐transformer exhibits this technique. Tensor field networks and spherical harmonics help the model to learn representations that honor the symmetrical features of 3D space [55].

2.2.4.3. 3D Spatial Transformers

By including distance‐aware spatial encodings and topological context into the attention layers, other models such as GeoFormer and Graphormer broaden the transformer architecture for graph‐structured data. Furthermore, Choi, S. R. and M. Lee (2023) discussed that these models strike a mix between structural bias inherent in molecular graphs and the expressiveness of attention [49]. Graphormer, for example, uses spatial encodings reflecting the shortest‐path distances in the molecular graph to simultaneously consider topological and geometric connections. Molformer fits problems demanding fine‐grained modeling of conformational flexibility and reactivity by combining atom type embeddings, torsion angles, and structural priors into a unified attention framework [4956].

In many different molecular applications where long‐range dependencies and spatial configurations are crucial, geometric transformers have shown exceptional performance. These models could learn subtly occurring interactions throughout the whole binding interface in binding affinity prediction. It helps to develop compounds satisfying both pharmacophoric restrictions and geometric complementarity with the target binding site in ligand creation. Geometric transformers also allow one to replicate the conformational changes ligands and proteins undergo upon binding, including induced fit and allosteric effects, in docking and pose prediction [5758].

Geometric transformers have one clear benefit in their adaptability for both generative and discriminative modeling, their ability to process disparate input sources and capture interactions across several spatial and chemical scales, but they are also computationally demanding, especially in relation to big protein–ligand systems since their quadratic complexity in the atom or node count [59].

2.2.5. Point Cloud Networks

Operating directly on unordered sets of 3D coordinates, point cloud networks are a potent and progressively popular family of GDL architectures that can simulate molecular systems as spatially dispersed atomic point clouds. Point cloud models retain the continuous geometry of the molecular system without imposing artificial structural limitations, unlike voxel‐based approaches, which discretize molecular space into regular grids, or graph‐based methods depending on preset connectedness. Each point in the cloud typically corresponds to an atom and is associated with both spatial coordinates x i ∈ R3 and additional atomic features f i ∈ Rd, yielding a set P={(xi,fi)}i=1N  for a molecule of N atoms [60].

For 3D object recognition, Charles, R. Q. et al. (2017) showed that PointNet is the fundamental architecture for this field and has subsequently been modified for molecular uses. Using a common MLP, PointNet independently processes each point independently generating per‐point feature vectors h i = MLP(f i). PointNet uses a symmetric aggregation process, usually max or mean pooling, over the set to generate a global representation invariant of the ordering of points:

hglobal=γ({hi}i=1N)=maxi{hi} (10)

This approach ensures permutation invariance, which is essential for operating on unordered sets of atoms. PointNet may thus overlook local geometric structures and interatomic connections essential to understanding molecular interactions by considering points independently before aggregation [61].

To address this limitation, in 2025, Lu, X. et al. worked on PointNet++, which expands the architecture using a hierarchical feature learning approach with local neighborhood aggregation in order to overcome this restriction. Local characteristics are learned by recursive sampling and grouping operations; each point has a local spatial context acquired by grouping neighboring points based on Euclidean distance within a ball of radius r . PointNet++ can capture multi‐scale geometric aspects by means of these hierarchical layers, therefore enabling more expressive modeling of molecular surfaces and binding site topologies [62].

Point cloud networks have shown great performance in jobs demanding realistic 3D geometry without voxelization or graph encoding in the framework of molecular modeling and drug development. The primary applications involve mimicking protein binding sites, where atoms within the surface of the binding cavity produce complex and unstable structures. Considering these atoms as point sets helps point cloud networks learn to record the form, curvature, and spatial density of interaction surfaces [60]. These characteristics were particularly pertinent in determining interaction hotspots, ligand binding modalities, and ligand binding affinities depending on ligand structural complementarity with the binding pocket.

Their value has been further improved by recent modifications of point cloud networks especially designed for drug discovery. Zhang, Z. et al. worked on ProteinPocketNet to extract feature representations straight from point clouds depicting protein pockets, thereby enabling the exact characterization of their physicochemical and geometric features. Learning from raw atomic coordinates and related information helps the model to deduce binding site specificity and reactivity patterns vital to rational drug design [63].

In processes requiring structural data collected from cryo‐electron microscopy (Cryo‐EM), AlphaFold‐predicted protein models, or docked ensembles, where the correct 3D arrangement of atoms is critical and frequently not well‐suited for voxelization, point cloud networks are very useful. Point cloud models’ versatility lets them maintain geometric integrity while processing structural data with different degrees of completeness and resolution [64]. Furthermore, Table 3 provides a comparative benchmark summary of representative GDL architectures used in structure‐based drug discovery, highlighting their benchmark datasets, evaluation metrics, computational requirements, and practical performance.

TABLE 3.

Comparative benchmark summary of representative GDL architectures used in structure‐based drug discovery.

Architecture Primary task Dataset Training data scale Evaluation metric Representative performance Computational cost Open‐source
SchNet Molecular property prediction QM9, PDBbind  ~130k molecules RMSE Low prediction error Moderate Yes
DimeNet++ Molecular property prediction QM9  ~130k molecules RMSE, Spearman's ρ Improved over SchNet High Yes
GemNet Molecular property prediction QM9, OC20 130k–1.3M structures MAE, RMSE State‐of‐the‐art High Yes
EGNN Molecular property prediction QM9, MD17 130k molecules MAE, RMSE High accuracy Moderate–High Yes
SE(3)‐Transformer Binding affinity prediction PDBbind  ~19k complexes RMSE, Pearson's R Superior affinity prediction High Yes
EquiBind Docking/Pose prediction PDBbind  ~19k complexes Top‐1 Success Rate Fast inference Low–Moderate Yes
DiffDock Docking/Pose prediction PDBbind, PoseBusters  ~19k complexes Top‐1 Success Rate State‐of‐the‐art docking accuracy Moderate Yes

2.2.6. Recent Advances in GDL

Recent advances in GDL have substantially expanded its applications in SBDD. The convergence of high‐accuracy structural prediction, diffusion‐based generative modeling, and PLMs has significantly enhanced molecular representation learning, protein–ligand interaction modeling, and de novo drug design. These developments enable GDL frameworks to integrate sequence, structural, and geometric information, thereby improving predictive accuracy, generalizability, and computational efficiency [1265].

2.2.6.1. AlphaFold3 and Structural Modeling

The introduction of AlphaFold3 represents a major milestone in computational structural biology by extending structure prediction beyond individual proteins to biomolecular complexes involving proteins, ligands, nucleic acids, and ions. Unlike AlphaFold2, which primarily focused on protein tertiary structure prediction, AlphaFold3 accurately models protein–ligand, protein–DNA, protein–RNA, and protein–ion interactions within a unified framework [66]. This advancement provides high‐quality structural templates for downstream GDL applications, including binding‐site identification, binding affinity prediction, molecular docking, and lead optimization. The improved structural accuracy also facilitates the development of geometry‐aware neural networks by providing more reliable 3D molecular representations. Nevertheless, challenges remain in accurately modeling highly flexible proteins, transient molecular interactions, and conformational dynamics, indicating that experimental validation and MD simulations continue to complement AI‐based structural prediction [67].

2.2.6.2. Diffusion‐Based Generative Models

Diffusion‐based generative models have emerged as one of the most significant developments in AI‐driven drug discovery. Recent frameworks such as DiffSBDD, DiffDock, SurfDock, FlexDock, and BInD employ geometric diffusion processes to generate chemically valid molecules and predict protein–ligand binding poses while preserving three‐dimensional structural constraints [68]. Unlike traditional generative approaches, diffusion models iteratively refine molecular structures from noisy initial states, enabling accurate de novo molecular generation, binding‐pose prediction, ligand optimization, and structure‐conditioned molecular design. By integrating geometric information throughout the diffusion process, these methods achieve improved molecular diversity, docking accuracy, and physicochemical consistency, making them promising tools for next‐generation SBDD pipelines [16].

2.2.6.3. PLMs and Multimodal Geometric Learning

Another important trend is the integration of PLMs with GDL architectures. Foundation models such as ESM‐3, Uni‐Mol, and GearNet‐ESM combine sequence‐derived embeddings with three‐dimensional structural representations to generate comprehensive molecular feature representations [69]. While PLMs capture evolutionary conservation and contextual sequence information from millions of protein sequences, GDL architectures simultaneously learn spatial relationships, molecular geometry, and physicochemical interactions. The resulting multimodal encoders leverage sequence embeddings, structural embeddings, and geometric representations to improve protein–ligand interaction prediction, molecular property prediction, virtual screening, and binding affinity estimation. This integration enhances model robustness, particularly for proteins with limited experimentally determined structures, and represents a promising direction toward foundation models for AI‐assisted drug discovery [70].

Collectively, these recent developments demonstrate a clear transition from conventional geometry‐aware neural networks toward integrated multimodal foundation models that combine structural biology, GDL, and generative artificial intelligence (AI). Such approaches are expected to accelerate rational drug design by enabling more accurate prediction of biomolecular interactions, efficient molecular generation, and improved generalization across diverse therapeutic targets.

3. Advantages of GDL in SBDD

GDL provides a transforming benefit in SBDD by analyzing the 3D spatial information of molecular systems. It presents more accurate modeling of atomic interactions, binding affinities, and conformational dynamics. Furthermore, GDL captures intricate geometric relationships and anisotropic interactions, in contrast to conventional methods that rely on 2D or string‐based representations. This drug design model could improve predictive performance, generalizability, and interpretability in drug discovery tasks.

3.1. Directly Processes 3D Structural Information

One of the fundamental advantages of GDL is its inherent ability to process and learn from the three‐dimensional (3D) structure of molecular systems. GDL frameworks consider the real spatial configuration of atoms, unlike conventional machine learning techniques depending on flattened molecular representations such as SMILES strings, 2D graphs, or manually created descriptions. This is accomplished through representations such as atomistic graphs with embedded Cartesian coordinates, point clouds, or surface meshes; each capturing the geometric arrangement of atoms within ligands and protein targets [15].

Molecular recognition, where binding affinity and selectivity are governed by the 3D complementarity between ligand and target, is inherently geometric. Therefore, it necessitates learning algorithms capable of accurately interpreting molecular shape, volume, orientation, and interaction interfaces. As demonstrated by Crampon K. et al. (2022), during both training and inference processes, GDL models could save and utilize this spatial data information. These properties of GDL help them to accomplish important tasks including binding affinity prediction, posture estimation, and identification of interaction sites with more exact, improved accuracy and biological relevance. By modeling the actual 3D conformation of a complex X={(xi,fi)}i=1N, where x i ∈ R3 are coordinates, and f i ∈ Rd are atomic features, GDL can extract patterns that directly reflect real‐world molecular behavior [71].

3.2. Captures Complex Spatial Relationships

GDL has a prominent capability to capture complex and high‐order spatial interactions between atoms, residues, and molecule surfaces. These interactions were critical for biological activity relations. Traditional scoring functions, often reliant on pairwise interaction terms and simplistic distance cutoffs, fail to capture the geometric and physicochemical nature of protein–ligand interactions. In contrast, graph‐based and equivariant networks are designed to model not only pairwise distances but also angular and torsional characteristics, hence simulating the whole geometric environment of molecular assemblies [72].

MPNNs implement context‐sensitive updates through distance‐ and angle‐based weighting in the information flow across nodes. For example, DimeNet uses both bond lengths d ij  and angles θ jik , where the message function is specified as:

mij=e(dij,θjik,hi,hj) (11)

Crucially for modeling hydrogen bonding, π–π stacking, and other anisotropic interactions, this approach lets models grasp directional dependence on interactions. GDL can thereby simulate spatially diffuse and multi–body interactions, which are necessary for precisely describing complicated chemical systems and binding landscapes [73].

3.3. Improved Accuracy and Generalization

GDL models exhibit exceptional performance not only in predictive accuracy but also in their ability to generalize across novel biological targets and diverse chemical scaffolds. Particularly when used on unknown protein families or chemically complex ligands, traditional models, including empirical scoring systems and nongeometric DL methods, often suffer from overfitting or poor extrapolation [74].

In the study of Li, J. et al. (2024) demonstrated that Geometric priorities in GDL models learn more universal and transferable representations. Equivariant designs such as E(n)‐equivariant networks and SE(3)‐Transformers preserve consistency under 3D rotations and translations, therefore guaranteeing that forecasts are invariant to coordinate frame selections. This feature produces models capable of generalizing from smaller or more heterogeneous datasets as well as more data‐efficient ones. GDL models routinely beat conventional docking‐based scoring systems in predicting binding affinities and binding poses and hit identification rates on extensively used benchmarks such as PDBbind and CASF‐2016. Moreover, pretraining GDL models on large unlabeled protein‐ligand structures using self‐supervised tasks, such as denoising, context prediction, or node masking—has enhanced their generalizing ability to new biochemical settings [75].

3.4. End‐to‐End Learning

GDL designs offer end‐to‐end learning pipelines whereby models are trained directly on raw 3D molecule structures to execute sophisticated prediction tasks without the need for manual feature engineering or domain‐specific preprocessing. This approach contrasts markedly with conventional pipelines that need intermediary representations, physicochemical descriptors, pharmacophore fingerprints, or docking score generation [76].

Shen, C. et al. discussed that end‐to‐end GDL frameworks begin with structural inputs, typically 3D atomic coordinates and features such as atom type, hybridization state, and partial charge, and transform them into task‐specific representations through differentiable neural modules. For instance, in an end‐to‐end model designed to predict binding affinity y, the process can be formalized as:

y=fθ({xi,fi)}i=1N (12)

where f θ  represents the GDL model parameterized by θ. Differentiable with regard to both geometric and chemical characteristics, these models enable integrated optimization and multi‐objective training. From a single structural representation, this design not only simplifies model creation but also enables combined learning of several attributes, including affinity, toxicity, and ADMET profiles [77].

3.5. Potential for Enhanced Interpretability

Despite the remarkable predictive power of DL, the interpretability of these models remains a major challenge, particularly in critical drug therapeutics such as small‐molecule applications in drug discovery and development. Whereas GDL provides unique pathways to improve model transparency through various techniques, including but not limited to graph‐based saliency analysis, gradient‐based attribution, and visual attention methodologies. These methods provide practical insights into the factors of binding and activity, therefore clarifying the molecular substructures or geographical areas most affecting model predictions [78]. Nerrise F. et al. (2023) demonstrated in their study that GATs were shown to assign learned attention coefficients (α ij ) to edges during message passing, which can be understood as markers of interaction significance between atoms i and j. Likewise, saliency maps generated utilizing gradients of the prediction about atomic coordinates yxi highlight areas within a protein–ligand complex that generate binding affinity. Techniques such as integrated gradients, layer‐wise relevance propagation, and perturbation‐based attribution are under active investigation to provide GDL models with more biological interpretability and transparency [79].

4. GDL Applications in Structural‐Based Drug Design

The integration of GDL into SBDD has marked a paradigm shift in computational drug discovery. GDL algorithms facilitate end‐to‐end learning from intrinsic geometric and physicochemical properties of molecular systems, enhancing drug discovery. The important applications demonstrate the transformative ability of GDL, including binding affinity prediction, virtual screening, de novo design, posture prediction, ligand–receptor pose prediction, ADMET profiling, and modeling of protein flexibility [80] (Figure 5).

FIGURE 5.

FIGURE 5

Schematic representation of applications of GDL.

4.1. Protein–ligand Binding Affinity Prediction

GDL models have advanced protein–ligand binding affinity prediction and serve as a robust computational framework for quantifying the interaction intensity between small molecules and protein targets. In a recent study, Wang et al. (2024) introduced GDL frameworks that learn complicated interaction patterns directly from molecular geometry using spatially aware designs, unlike conventional scoring systems depending on simplified physicochemical approximations. Capturing minor NCI, shape complementarity, and solvent exposure, these models include structural information including protein binding pockets, ligand conformations, and interatomic distances [67]. Cremer, J. et al. (2023) demonstrated that the incorporation of rotationally and translationally equivariant characteristics into architectures including SchNet, DimeNet, and SE(3)‐Transformers enhanced binding affinity prediction. These models exhibit superior generalizability across several target classes and chemical scaffolds, outperforming conventional docking‐based scoring systems in benchmark datasets such as PDBbind and CASF [81]. For example, another model, FlexDock, designed by Corso G. et al. improves docking performance by increasing the proportion of energetically favorable conformations from 30% to 73% and enhancing accuracy compared to PDBBind benchmark [82]. Cao et al. introduced SurfDock, a geometric diffusion network distinguished by its equivalent architecture and its ability to integrate multiple protein representations, including primary sequences, 3D structural graphs, and surface‐derived features [83] . By performing generative diffusion on a non‐Euclidean manifold, SurfDock enables precise optimization of ligand translation, rotation and torsion, resulting in more reliable and physically consistent binding‐pose generation.

4.2. Structure‐Based Virtual Screening

GDL models are increasingly applied in virtual screening workflows to predict the binding potential of compounds against specific protein targets and enable the ranking of candidate molecules from large chemical libraries. Conventional virtual screening systems follow molecular docking to stimulate ligand binding poses, followed by scoring functions to estimate binding affinity [84]. Lyu J. et al. (2023) reported that these techniques often exhibit limited accuracy in active chemicals and are sometimes computationally expensive. In contrast, GDL‐based approaches learn statistical patterns from known protein–ligand interactions and apply this knowledge to quickly evaluate new compounds. By incorporating structural representations of both the protein target and ligand candidate, GDL models enhance hit identification accuracy and binding probability estimation. This allows for high‐throughput screening of ultra‐large compound libraries with improved confidence and delivers notable speed and accuracy gains over conventional docking‐based strategies [85]. Pinheiro PO et al. introduced VoxBind, a score‐based generative model for 3D molecular design conditioned on protein structures [86]. The framework employs a 3D voxel‐denoising network capable of learning and generating molecular geometries by representing ligands as atomic density grids, enabling accurate and spatially coherent ligand‐protein pose generation. This model outperforms existing approaches across multiple in silico benchmarks, offering substantially faster sampling and being simple to train. The generated molecules exhibit greater structural diversity, reduced steric clashes, and improved pocket‐specific binding affinity compared with state‐of‐the‐art methods.

4.3. De Novo Drug Design and Lead Optimization

In generative modeling, where the goal is to design new molecules with desired biological and physicochemical characteristics, GDL plays a critical role. GDL models have been effectively combined in the framework of SBDD with variational autoencoders, generative adversarial networks (GANs), and transformer‐based generative models to generate drug‐like compounds physically and functionally tailored to bind protein targets [87]. In the study, Ivanenkov, Y. et al. (2023) reviewed that these models can build compounds that not only satisfy synthetic accessibility and pharmacokinetic criteria but also conform to geometric restrictions of the binding site by including them into the molecular generation process [88]. Furthermore, GDL is utilized more to guide molecular optimization toward enhanced potency, selectivity, and favorable ADMET profiles. Its ability to fine‐tune lead compounds by predicting how small changes influence binding interactions significantly reduces experimental load and accelerates the lead optimization cycle [89]. A diffusion‐based molecular generative framework, bond and interaction‐generating diffusion (BInd) model, was introduced by Lee J. et al. [90]. The model simultaneously generates atoms, bonds, and NCI within a target binding pocket. BInD was specifically designed to address key multi‐objective challenges encountered in DL‐based SBDD: (i) achieving precise local geometry, (ii) ensuring desirable drug‐like molecular properties, and (iii) generating target–appropriate interactions. By jointly modeling both structural and interaction features of protein–ligand binding, BInD provides a comprehensive strategy for molecular design. Its ability to generate chemically viable molecules with realistic 3D poses and favorable binding interactions opens new opportunities for computer‐aided drug discovery and optimization.

4.4. Prediction of Protein–Ligand Interactions and Binding Modes

GDL models not only predict binding affinity but also accurately forecast the specific binding mode; that is, how a ligand orients interacts within the active sites of a target protein. These projections indicate the appropriate binding poses and identify key molecular interactions like hydrophobic packing, van der Waals forces, salt bridges, and hydrogen bonds [91]. Xue, F. et al. (2025) depicted that GDL‐based models such as EquiBind and DiffDock can bypass conventional docking engines by directly predicting binding conformations from unbound protein and ligand structures using SE(3)‐equivariant networks. These models achieve high accuracy even involving flexible targets or uncertain binding sites by considering spatial restrictions and interaction specificity. Furthermore, GDL may find important interacting residues and elucidate the chemical basis of binding, thereby enhancing interpretability and providing biological understanding of drug‐target interactions vital for rational drug design [92].

4.5. Prediction of ADMET Properties With Structural Context

Drug safety and efficacy depend on the prediction of ADMET. Traditional ADMET prediction models, which often depend on molecular fingerprints or physicochemical descriptors, frequently overlook spatial configuration or conformational dynamics [93]. De Vries, M. et al. (2025) demonstrated that GDL models address this limitation by learning from 3D molecular structures and their surroundings, therefore allowing more exact modeling of traits including solubility, permeability, metabolic liability, and toxicity. GDL can more effectively represent steric effects, polar surface area, and bioisosterism, all of which greatly affect ADMET behavior by including geometric and atomic‐level properties. Implementing these models during early‐stage screening offers the potential to eliminate candidates with suboptimal pharmacokinetics or toxicity profiles, which reduces downstream attrition in drug development [26].

4.6. Understanding Protein Flexibility and Conformational Changes

Accurate modeling of ligand binding critically depends on accounting for protein flexibility and conformational plasticity; factors that remain one of the persistent challenges in SBDD. Conventional methods either depend on pre‐generated protein ensembles from computationally expensive MD simulations or consider the protein as a rigid body [94]. Rudden et al. (2022) explored the application of GDL models specifically designed to incorporate protein flexibility, either by learning robust latent representations that tolerate structural variability or by training on multiple conformational states. Some techniques use ensemble learning, whereby several conformations of a target are simultaneously stored to capture a larger conformational landscape. Others learn flexible embeddings from structural databases or use normal mode analysis. These developments allow GDL models to better forecast binding events including induced fit or allosteric control, hence increasing their relevance to increasingly challenging and dynamic drug targets [95].

5. Revolutionizing Small‐Molecule Discoveries Through GDL

The discovery of small‐molecule inhibitors (SMIs) signifies a crucial advancement in the progression of targeted therapeutics [96]. These compounds play a significant role in modern pharmacotherapy, particularly in the management of chronic diseases. Due to their favorable pharmacokinetic properties, such as oral bioavailability, low molecular weight, and well‐balanced ADME profiles, these SMIs were considered as prominent inhibitors. The identification and development of low molecular weight compounds like SMIs has traditionally been a labor‐intensive, costly, and time‐consuming process, often requiring over a decade and investments exceeding several billion dollars to bring a single drug to market [97]. However, the advent of AI has initiated a paradigm shift in pharmaceutical research and development (Figure 6).

FIGURE 6.

FIGURE 6

Comparison between traditional and AI‐assisted approaches in drug discovery. The table highlights key differences across various stages of the drug discovery pipeline, including target identification, hit identification, and lead optimization. AI‐assisted methods leverage machine learning and deep learning techniques such as natural language processing (NLP), virtual screening, and generative models to accelerate timelines, reduce costs, and enhance efficiency. Figure created using Biorender.com.

Leveraging its unparalleled capacity to analyze vast datasets, predict molecular properties, and design novel chemical entities, AI accelerates early‐phase discovery and enhances data‐driven decision‐making. Rather than replacing human innovation, AI acts as an augmentation tool, expanding the scope of rational drug design and enabling more efficient hypothesis generation. In recent years, several pioneering examples have illustrated AI's transformative role in small molecule discovery [97]. A notable case is the development of DSP‐1181 by Exscientia and Sumitomo Dainippon Pharma, an AI‐designed molecule that entered Phase I clinical trials in early 2020 for the treatment of anxiety‐related disorders [98].

Molecular design tasks are generally classified into three categories: macromolecule (e.g., protein) design, SBDD, and ligand‐based drug design (LBDD) [99100]. SBDD and LBDD utilize computational strategies to rationally develop small‐molecule ligands that specifically target biologically relevant molecules. For example, Schneuing A. et al. developed DiffSBDD, an SE(3)‐equivariant diffusion model that conditions ligand generation on 3D protein‐binding pockets, effectively framing SBDD as a 3D conditional molecular generation problem [101]. This model can address key challenges in molecular generation, including off‐the‐shelf property optimization, explicit negative design, partial molecular design. The quality of generated drug candidates can be further enhanced by incorporating additional constraints derived from diverse computational metrics. With the rapid advancement of AI technologies, molecular GDL has emerged as a powerful approach to accelerate drug discovery by enabling efficient learning from molecular structures. Various computational methods have been developed to predict binding affinities based on the three‐dimensional structures of ligand–target complexes. While some approaches utilize CNNs or GNNs directly for prediction, others rely on engineered descriptors that capture key ligand–target interactions, which are subsequently input into predictive algorithms [1].

Optimization of ligands—small molecules that bind to target biomolecules to improve their pharmacological properties—is a common challenge in drug discovery. To address this challenge, Powers et al. developed fragment‐based molecular expansion (FRAME), a machine learning framework that uses 3D protein‐ligand structures (Figure 7a) [4]. FRAME models the expansion process as a sequence of three‐dimensional steps, selecting appropriate molecular fragments, determining optimal attachment sites, and defining their spatial orientation based on the input structure of a starting ligand bound within a protein pocket. Instead of relying on manually encoded rules for synthetic feasibility or binding affinity, FRAME uses neural networks trained on existing high‐affinity, drug‐like protein‐ligand complexes to learn relevant patterns. This data‐driven approach offers medicinal chemists a powerful tool for hypothesis generation, potentially accelerating drug development and contributing to improved therapeutic outcomes. In another study, Putin et al. introduced a novel DNN architecture called reinforced adversarial neural computer (RANC) for the de novo design of small organic molecules [102]. RANC integrates reinforcement learning (RL) with the GAN framework (Figure 7b). Comparative analyses demonstrated that RANC, trained on the SMILES representation of molecules, outperforms its earlier DNN‐based counterpart ORGANIC across several drug discovery‐relevant metrics, including higher QED scores, improved compliance with medicinal chemistry filters (MCFs), greater structural uniqueness, and adherence to the Muegge criteria.

FIGURE 7.

FIGURE 7

Schematic view of FRAME (a), figure adapted from [4] Copyright 2023, The authors, published in American Chemical Society and RANC models (b), figure adapted from [102], Copyright 2018 American Chemical Society. Schematic representation of the structure flow of RLBSIF (c), Figure is reprodcued (adapted) with permission from [103], Copyright 2025 Elsevier.

The functional roles of RNA in catalysis and structural folding are controlled by interactions between RNA and small‐molecule ligands. Therefore, accurate prediction of ligand binding sites within RNA structures is crucial. To address this challenge, Zhu et al. introduced RNABind, a GDL framework guided by structural embeddings [104]. RNABind enables the prediction of small molecule binding sites in both single‐chain and multi‐chain RNA structures. This approach, which represents the complete RNA complex as a graph to capture interchain interactions and RNA flexibility, marks the first application of GDL to identify RNA–ligand binding sites. The integration of large RNA language model embeddings enhances the model's ability to generalize and capture both structural and contextual information. Recently, Sang et al. also developed RNA–ligand binding surface interaction fingerprints (RLBSIF), a computational framework based on GDL [103]. Leveraging MaSIF‐derived surface interaction fingerprints, RLBSIF integrates chemical features (e.g., atomic charges) with geometric descriptors (e.g., shape index, distance‐dependent curvature) to comprehensively characterize RNA–ligand interfaces [105]. Trained on 440 binding pockets, RLBSIF achieved 90% classification accuracy at the pocket level and outperformed existing models in two independent benchmarks, demonstrating its utility in accurately identifying binding sites within complex RNA structures. This method shows promise for RNA‐targeted drug design and the development of RNA‐based therapeutics.

Further, to predict the binding conformation of small bioactive molecules to protein targets, Méndez‐Lucio et al. developed DeepDock, a GDL‐based technique [1]. This method learns a statistical potential specific to each ligand–target pair, based on distance likelihood, enabling accurate prediction of binding conformations. Statistical potentials are then used to efficiently sample small‐molecule conformations. Das et al. developed a new model based on GDL to predict drug–virus interactions against COVID‐19 [106]. Their findings demonstrate that the proposed approach achieves 97% accuracy in predicting drug–virus interactions, outperforming existing methods. In summary, all these GDL models are revolutionizing small molecule drug discovery by offering accurate, robust, and rapid predictions of binding sites. While traditional methods like template‐based and consensus approaches remain valuable, integrating them with machine learning can enhance insights by incorporating chemical‐specific information.

6. Challenges and Future Directions

GDL has transforming potential in SBDD, yet several challenges must be addressed to ensure its effective integration into practical drug discovery pipelines. These limitations extend beyond technical concerns and encompass broader scientific, infrastructural, and interpretability issues. Addressing these obstacles will require a multifaceted approach involving methodological innovation, high‐quality data curation, and multidisciplinary collaboration as the field continues to evolve [13].

The availability and quality of large‐scale, well‐annotated 3D structural datasets are a main obstacle in the evolution of advanced GDL models. Deep geometric models require extensive and diverse data to capture the heterogeneity of molecular interactions and structural conformations across several protein‐ligand complexes. However, commonly used datasets such as PDBbind and CASF are relatively limited in size and suffer from problems including inconsistent resolution, biased distributions, and incomplete binding annotations [107]. Many crystallographic structures exhibit poorly resolved binding pockets or missing side chains that can cause noise in training performance and compromise model generalization. Furthermore, the lack of high‐resolution experimental structures for novel targets, combined with reliance on predicted or homology‐modeled complexes, further undermines model reliability. Addressing these limitations will need the extension of carefully selected, varied, and standardized structural datasets, potentially augmented with synthetic data derived from MD simulations or high‐confidence computational predictions from computational frameworks such as AlphaFold [108].

Effective modeling of protein flexibility and conformational dynamics remains a significant challenge. Most existing GDL models are built upon static representations of protein‐ligand complexes typically derived from a single crystallographic structure [109]. However, proteins are inherently dynamic macromolecules that undergo conformational changes in response to ligand interaction and environmental cues. Ignoring this flexibility can lead to inaccurate predictions, especially for targets undergoing notable induced fit or allosteric modulation. Emerging approaches attempt to capture conformational ensembles by including data from MD simulations or by training on several protein configurations. Nonetheless, scalable and data‐efficient integration of protein flexibility into GDL pipelines remains unresolved. environment. Future directions may involve the development of temporally aware geometric models or hybrid frameworks that combine GDL with coarse‐grained or atomistic dynamic simulations, thereby more accurately representing the biophysical landscape of small molecules [110].

Another important issue that requires further investigation is interpretability and explainability. While GDL models demonstrate strong predictive potential, their opaque internal decision‐making systems can limit their use in hypothesis‐driven molecular design. Establishing trust in model predictions and informing experimental validation relies on a comprehensive understanding of the underlying chemical properties and spatial relationships that drive outcomes such as binding affinity and pose ranking [111]. Although several interpretability approaches, such as gradient‐based saliency maps and attention‐based visualization, have been adapted for geometric models, these approaches are still in their early stages and frequently lack biological interpretability. An essential first step in making GDL models more transparent and useful is building strong, domain‐specific interpretability models that can translate predictions back to chemically significant properties or pharmacophores [112].

Furthermore, the current models, such as SE(3)‐Transformers and E(n)‐equivariant GNNs, have shown greater advancement; several difficulties remain in accurately capturing the anisotropic, multi‐scale, and context‐dependent character of protein–ligand binding. Future architectures may incorporate higher‐order tensors, rotational harmonics, or topological elements like persistent homology to better represent the geometry and topological complexity of molecular interactions. Additionally, the development of multi‐resolution models capable of simultaneously processing atomic, residue‐level, and surface‐level data could further enhance predictive accuracy and model interpretability [32]. Furthermore, Table 4 provides an overview of current challenges and emerging research directions in GDL for structure‐based drug discovery.

TABLE 4.

Current challenges and future research directions for GDL in structure‐based drug discovery.

Current Challenge Proposed solution Potential impact
Limited high‐quality structural datasets Self‐supervised learning, contrastive pretraining, transfer learning Improved generalization and reduced dependence on labeled data
Protein flexibility and conformational dynamics Multi‐resolution GDL, hybrid physics‐informed ML, integration with molecular dynamics simulations More realistic modeling of protein–ligand interactions
Limited model interpretability Attention visualization, gradient‐based attribution, explainable AI (XAI), scaffold‐level interpretation Improved confidence in predictions and rational lead optimization
High computational cost Sparse attention mechanisms, efficient equivariant architectures, distributed computing Faster training and improved scalability
Limited experimental validation Closed‐loop integration with cryo‐EM, X‐ray crystallography, high‐throughput screening and medicinal chemistry workflows Faster translation from computational predictions to experimental validation

Future progress in GDL will depend on the convergence of data‐centric AI, physics‐informed modeling, and experimental validation. Self‐supervised pretraining on large structural databases, multimodal foundation models integrating sequence and 3D structural information, and hybrid MD–geometric learning frameworks are expected to substantially improve predictive accuracy for protein flexibility and molecular recognition [113]. In parallel, advances in explainable GDL, including attention‐based visualization and gradient attribution methods, may provide interpretable insights into protein–ligand interactions and support rational scaffold optimization during lead optimization. Ultimately, integrating GDL with automated experimental platforms and medicinal chemistry workflows has the potential to establish closed‐loop AI‐driven drug discovery pipelines, accelerating the identification and optimization of next‐generation therapeutics [113].

7. Conclusion

GDL has emerged as a transformative force in SBDD, enabling unprecedented capacity to learn directly from the intricate 3D geometries that govern molecular recognition and interaction. By inherently incorporating spatial, topological, and physicochemical information, GDL surpasses traditional machine learning models that rely on flattened or heuristic representations. Its physically grounded and computationally scalable framework empowers a new generation of predictive models that deliver enhanced accuracy, generalizability, and biological relevance across key tasks such as binding affinity prediction, pose estimation, virtual screening, and de novo molecular generation.

The integration of GDL into the drug discovery pipeline has significantly accelerated the transition from hit identification to lead optimization, while reducing dependence on manually engineered features and conventional scoring functions. GDL's ability to encode molecular symmetries, capture orientation–invariant interactions, and execute end‐to‐end learning from raw 3D structures offers substantial advantages in modeling real‐world biological systems. Despite these advancements, the limited availability of diverse, high‐resolution structural datasets continues to constrain model robustness and scalability of GDL models. Moreover, incorporating protein flexibility and dynamic conformational changes into learning frameworks remains an unresolved challenge. Additionally, the interpretability of GDL models is an essential component for hypothesis‐driven molecular design, and the computational demands of training complex architectures pose additional barriers to their widespread adoption. Future directions suggest the development of hybrid approaches that integrate GDL with physics‐based methods such as molecular docking, MD simulations, and quantum mechanics to design interpretable, multi‐scale models. Advancements in architectural design, especially in multimodal and multi‐resolution learning, along with emerging techniques such as transfer learning and self‐supervised pretraining, would be instrumental in fully harnessing the potential of GDL and molecular modeling.

In summary, GDL offers a unified, geometry‐aware, and data‐adaptive paradigm capable of addressing the inherent complexity of molecular interactions in drug discovery. Its continued evolution, supported by methodological innovation and interdisciplinary collaboration, could be poised to redefine rational drug design and accelerate the development of safer, more effective therapeutics.

Author Contributions

A.K.S., U.M., and R.K. wrote the main manuscript text and prepared the figures. R.S. and D.B. reviewed and edited the manuscript. All authors reviewed the final version of the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Acknowledgments

The authors sincerely thank IIT Gandhinagar for providing the necessary facilities. R.S. acknowledges the Anusandhan National Research Foundation (ANRF), Government of India and IIT Gandhinagar for financial support through the National Post‐Doctoral Fellowship (NPDF) (PDF/2023/000365) and Achievers Post‐Doctoral Fellowship (APD) (OTH/R&D/13516). U.M. further acknowledges financial support from STARS‐MoES. D.B. thanks SERB‐DST, GUJCOST, GSBTM, and STARS‐MoES for research grants.

Contributor Information

Amit Kumar Srivastav, Email: amitks.kit@gmail.com.

Dhiraj Bhatia, Email: dhiraj.bhatia@iitgn.ac.in.

Raghu Solanki, Email: raghu.solanki@iitgn.ac.in.

Data Availability Statement

No datasets were generated or analyzed during the current study.

References

  • 1. Méndez‐Lucio O., Ahmad M., del Rio‐Chanona E. A., and Wegner J. K., “A Geometric Deep Learning Approach to Predict Binding Conformations of Bioactive Molecules,” Nature Machine Intelligence 3 (2021): 1033–1039. [Google Scholar]
  • 2. Atz K., Grisoni F., and Schneider G., “Geometric Deep Learning on Molecular Representations,” Nature Machine Intelligence 3 (2021): 1023–1032. [Google Scholar]
  • 3. Townshend R. J. L., Eismann S., Watkins A. M., et al., “Geometric Deep Learning of RNA Structure,” Science 373 (2021): 1047–1051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Powers A. S., Yu H. H., Suriana P., et al., “Geometric Deep Learning for Structure‐Based Ligand Design,” ACS Central Science 9 (2023): 2257–2267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Lionta E., Spyrou G., Vassilatis D. K., and Cournia Z., “Structure‐Based Virtual Screening for Drug Discovery: Principles, Applications and Recent Advances,” Current Topics in Medicinal Chemistry 14 (2014): 1923–1938. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Sun Y., Jiao Y., Shi C., and Zhang Y., “Deep Learning‐Based Molecular Dynamics Simulation for Structure‐Based Drug Design against SARS‐CoV‐2,” Computational and Structural Biotechnology Journal 20 (2022): 5014–5027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Glielmo A., Husic B. E., Rodriguez A., Clementi C., Noé F., and Laio A., “Unsupervised Learning Methods for Molecular Simulation Data,” Chemical Reviews 121 (2021): 9722–9758. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Sarkar A., Concilio S., Sessa L., Marrafino F., and Piotto S., “Advancements and Novel Approaches in Modified AutoDock Vina Algorithms for Enhanced Molecular Docking,” Results in Chemistry 7 (2024): 101319. [Google Scholar]
  • 9. Vijayan R. S. K., Kihlberg J., Cross J. B., and Poongavanam V., “Enhancing Preclinical Drug Discovery with Artificial Intelligence,” Drug Discovery Today 27 (2022): 967–984. [DOI] [PubMed] [Google Scholar]
  • 10. Scott O. B., Gu J., and Chan A. W. E., “Classification of Protein‐Binding Sites Using a Spherical Convolutional Neural Network,” Journal of Chemical Information and Modeling 62 (2022): 5383–5396. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Bai Q., Xu T., Huang J., and Pérez‐Sánchez H., “Geometric Deep Learning Methods and Applications in 3D Structure‐Based Drug Design,” Drug Discovery Today 29 (2024): 104024. [DOI] [PubMed] [Google Scholar]
  • 12. Isert C., Atz K., and Schneider G., “Structure‐Based Drug Design with Geometric Deep Learning,” Current Opinion in Structural Biology 79 (2023): 102548. [DOI] [PubMed] [Google Scholar]
  • 13. Weller J. A. and Rohs R., “Structure‐Based Drug Design with a Deep Hierarchical Generative Model,” Journal of Chemical Information and Modeling 64 (2024): 6450–6463. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Zhang Z., Yan J., Huang Y., et al., “Structure‐Based Drug Design with Geometric Deep Learning: A Comprehensive Survey,” ACM Computing Surveys 58 (2025): 1–130. [Google Scholar]
  • 15. Zhang S., Liu Y., and Xie L., “A Universal Framework for Accurate and Efficient Geometric Deep Learning of Molecular Systems,” Scientific Reports 13 (2023): 19171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Alakhdar A., Poczos B., and Washburn N., “Diffusion Models in De Novo Drug Design,” Journal of Chemical Information and Modeling 64 (2024): 7238–7256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Gao W., Mahajan S. P., Sulam J., and Gray J. J., “Deep Learning in Protein Structural Modeling and Design,” Patterns 1 (2020): 100142. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Jo J., Kwak B., Lee B., and Yoon S., “Flexible Dual‐Branched Message‐Passing Neural Network for a Molecular Property Prediction,” ACS Omega 7 (2022): 4234–4244. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Jing B., Eismann S., Suriana P., Townshend R., and Dror R., “Learning from Protein Structure with Geometric Vector Perceptrons,” in ICLR 2021 (2021). [Google Scholar]
  • 20. Zhang Z., Xu M., Jamasb A. R., et al., “Protein Representation Learning by Geometric Structure Pretraining,” in The Eleventh International Conference on Learning Representations (2023). [Google Scholar]
  • 21. D’Hondt S., Oramas J., and De Winter H., “A Beginner's Approach to Deep Learning Applied to VS and MD Techniques,” Journal of Cheminformatics 17 (2025): 47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Chen H., Liu S., Chen W., Li H., and Hill R., “Equivariant Point Network for 3D Point Cloud Analysis,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2021), 14509–14518, 10.1109/CVPR46437.2021.01428. [DOI] [Google Scholar]
  • 23. Yang H., Qureshi R., and Sacan A., ”Protein Surface Representation and Analysis by Dimension Reduction,” Proteome Science 1 (2012): 1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Kudo G., Hirao T., Yoshino R., Shigeta Y., and Hirokawa T., ”Pocket to Concavity: A Tool for the Refinement of Protein–Ligand Binding Site Shape from Alpha Spheres,” Bioinformatics 39 (2023): btad212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Marchand A., Buckley S., Schneuing A., et al., “Targeting Protein–ligand Neosurfaces with a Generalizable Deep Learning Tool,” Nature 639 (2025): 522–531. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. De Vries M., Dent L. G., Curry N., et al., “Geometric Deep Learning and Multiple‐Instance Learning for 3D Cell‐Shape Profiling,” Cell Systems 16 (2025): 101229. [DOI] [PubMed] [Google Scholar]
  • 27. Cohen T. S., and Welling M., Proceedings of the 33rd International Conference on International Conference on Machine Learning ‐ Volume 48 (JMLR.org, 2016), 2990–2999. [Google Scholar]
  • 28. Fuchs F., Worrall D., Fischer V., and Welling M., “SE(3)‐Transformers: 3D Roto‐Translation Equivariant Attention Networks,” arXiv (2020), 10.48550/arXiv.2006.10503. [DOI] [Google Scholar]
  • 29. Bronstein M. M., Bruna J., Cohen T., and Velivckovi’c P., “Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges,” arXiv (2021), 10.48550/arXiv.2104.13478. [DOI] [Google Scholar]
  • 30. Sestak F., Schneckenreiter L., Brandstetter J., Hochreiter S., Mayr A., and Klambauer G., “VN‐EGNN: E(3)‐ and SE(3)‐Equivariant Graph Neural Networks with Virtual Nodes Enhance Protein Binding Site Identification,” Journal of Cheminformatics 18 (2025): 11, 10.1186/s13321-025-01127-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Zhang X. M., Liang L., Liu L., and Tang M. J., “Graph Neural Networks and Their Current Applications in Bioinformatics,” Frontiers in Genetics 12 (2021): 690049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Zhou J., Cui G., Hu S., et al., “Graph Neural Networks: A Review of Methods and Applications,” AI Open 1 (2020): 57–81. [Google Scholar]
  • 33. Liu C., Sun Y., Davis R., Cardona S. T., and Hu P., “ABT‐MPNN: an Atom‐Bond Transformer‐Based Message‐Passing Neural Network for Molecular Property Prediction,” Journal of Cheminformatics 15 (2023): 29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Chen Y., Wang J., Zou Q., et al., “DrugDAGT: A Dual‐Attention Graph Transformer with Contrastive Learning Improves Drug‐Drug Interaction Prediction,” BMC Biology 22 (2024): 233. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Bhatti U. A., Tang H., Wu G., Marjan S., and Hussain A., “Deep Learning with Graph Convolutional Networks: An Overview and Latest Applications in Computational Intelligence,” International Journal of Intelligent Systems 2023 (2023): 8342104. [Google Scholar]
  • 36. Vrahatis A. G., Lazaros K., and Kotsiantis S., “Graph Attention Networks: A Comprehensive Review of Methods and Applications,” Future Internet 16 (2024. [Google Scholar]
  • 37. Klicpera J., Becker F., and Günnemann S., Proceedings of the 35th International Conference on Neural Information Processing Systems (Curran Associates Inc., 2021), Article 520. [Google Scholar]
  • 38. Wu Z., Wang J., Du H., et al., “Chemistry‐Intuitive Explanation of Graph Neural Networks for Molecular Property Prediction with Substructure Masking,” Nature Communications 14 (2023): 2585. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Chen J., Liu Y., Li J., Su B., and Wen J., Atomic and Subgraph‐Aware Bilateral Aggregation for Molecular Representation Learning , 2023.
  • 40. Harren T., Matter H., Hessler G., Rarey M., and Grebner C., “Interpretation of Structure–Activity Relationships in Real‐World Drug Design Data Sets Using Explainable Artificial Intelligence,” Journal of Chemical Information and Modeling 62 (2022): 447–462. [DOI] [PubMed] [Google Scholar]
  • 41. Noor F., Junaid M., Almalki A. H., Almaghrabi M., Ghazanfar S., and Tahir ul Qamar M., “Deep Learning Pipeline for Accelerating Virtual Screening in Drug Discovery,” Scientific Reports 14 (2024): 28321. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Xu Y., Pei J., and Lai L., “Deep Learning Based Regression and Multiclass Models for Acute Oral Toxicity Prediction with Automatic Chemical Feature Extraction,” Journal of Chemical Information and Modeling 57 (2017): 2672–2685. [DOI] [PubMed] [Google Scholar]
  • 43. Wang J. and Ren J., “Graphormer Based Contrastive Learning for Recommendation,” Applied Soft Computing 159 (2024): 111626. [Google Scholar]
  • 44. Francoeur P. G., Masuda T., Sunseri J., et al., “Three‐Dimensional Convolutional Neural Networks and a Cross‐Docked Data Set for Structure‐Based Drug Design,” Journal of Chemical Information and Modeling 60 (2020): 4200–4215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Jiang H., Fan M., Wang J., et al., “Guiding Conventional Protein–Ligand Docking Software with Convolutional Neural Networks,” Journal of Chemical Information and Modeling 60 (2020): 4594–4602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Yamashita R., Nishio M., Do R. K. G., and Togashi K., “Convolutional Neural Networks: an Overview and Application in Radiology,” Insights into Imaging 9 (2018): 611–629. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Markus B., C. G.C., Andreas K., et al., “Accelerating Biocatalysis Discovery with Machine Learning: A Paradigm Shift in Enzyme Engineering, Discovery, and Design,” ACS Catalysis 13 (2023): 14454–14469. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Wallach I., Dzamba M., and Heifets A., “AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure‐based Drug Discovery,” arXiv (2015), 10.48550/arXiv.1510.02855. [DOI] [Google Scholar]
  • 49. Choi S. R. and Lee M., “Transformer Architecture and Attention Mechanisms in Genome Data Analysis: A Comprehensive Review,” Biology 12, no. 7 (2023): 1033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Liu Z., Roberts R. A., Lal‐Nag M., Chen X., Huang R., and Tong W., “AI‐Based Language Models Powering Drug Discovery and Development,” Drug Discovery Today 26 (2021): 2593–2607. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. P. Barros E. D.;lia, Malmstrom R. D., Nourbakhsh K., et al., “Electrostatic Interactions as Mediators in the Allosteric Activation of Protein Kinase a RIα,” Biochemistry 56 (2017): 1536–1545. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Shehzad A., Xia F., Abid S., et al., “Graph Transformers: A Survey,” arXiv (2024), 10.48550/arXiv.2407.09777. [DOI] [PubMed] [Google Scholar]
  • 53. Chen C., Wu Y., Dai Q., et al., “A Survey on Graph Neural Networks and Graph Transformers in Computer Vision: A Task‐Oriented Perspective,” IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (2024): 10297–10318. [DOI] [PubMed] [Google Scholar]
  • 54. Yuan C., Zhao K., Kuruoglu E. E., et al., “A Survey of Graph Transformers: Architectures, Theories and Applications,” arXiv (2025), 10.48550/arXiv.2502.16533. [DOI] [Google Scholar]
  • 55. Fuchs F. B., Worrall D. E., Fischer V., and Welling M., Proceedings of the 34th International Conference on Neural Information Processing Systems (Curran Associates Inc, 2020), Article 166. [Google Scholar]
  • 56. Mu J., Li Z., Zhang B., et al., “Graphormer Supervised De Novo Protein Design Method and Function Validation,” Briefings in Bioinformatics 25 (2024): bbae135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Min Y., Wei Y., Wang P., et al., “From Static to Dynamic Structures: Improving Binding Affinity Prediction with Graph‐Based Deep Learning,” Advanced Science 11 (2024): e2405404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Li J. and Gong X., “Harnessing Pre‐Trained Models for Accurate Prediction of Protein‐Ligand Binding Affinity,” BMC Bioinformatics 26 (2025): 55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Ali R. S. Aal E., Meng J., Khan M. E. I., and Jiang X., “Machine Learning Advancements in Organic Synthesis: A Focused Exploration of Artificial Intelligence Applications in Chemistry,” Artificial Intelligence Chemistry 2 (2024): 100049. [Google Scholar]
  • 60. Saranti A., Pfeifer B., Gollob C., Stampfer K., and Holzinger A., “From 3D Point‐cloud Data to Explainable Geometric Deep Learning: State‐of‐the‐art and Future Challenges,” WIREs Data Mining and Knowledge Discovery 14 (2024): e1554. [Google Scholar]
  • 61. Charles R. Q., Su H., Kaichun M., and Guibas L. J., “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” arXiv, 10.48550/arXiv.1612.00593. [DOI] [Google Scholar]
  • 62. Lu X., Guan Z., Pang D., Dou R., and Zheng X., “PointNet++SAKS: A Point Cloud Model Based on KANs and Attention Mechanism for Objects Classification and Semantic Segmentation,” IEEE Access 13 (2025): 29292–29304. [Google Scholar]
  • 63. Zhang Z., Shen W. X., Liu Q., and Zitnik M., “Efficient Generation of Protein Pockets with PocketGen,” Nature Machine Intelligence 6 (2024): 1382–1395. [Google Scholar]
  • 64. Giri N., Roy R. S., and Cheng J., “Deep learning for Reconstructing Protein Structures from Cryo‐EM Density Maps: Recent Advances and Future Directions,” Current opinion in Structural Biology 79 (2023): 102536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Pineda J., Midtvedt B., Bachimanchi H., et al., “Geometric Deep Learning Reveals the Spatiotemporal Features of Microscopic Motion,” Nature Machine Intelligence 5 (2023): 71–82. [Google Scholar]
  • 66. Peng C., Ni W., Liu Q., Hu G., and Zheng W., “A Comprehensive Benchmarking of the AlphaFold3 for Predicting Biomacromolecules and their Interactions,” Briefings in Bioinformatics 26 (2025): bbaf616. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Wang K., Huang Y., Wang Y., You Q., and Wang L., “Recent Advances from Computer‐Aided Drug Design to Artificial Intelligence Drug Design,” RSC Medicinal Chemistry 15 (2024): 3978–4000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68. Wang J., Zhou P., Wang Z., et al., “Diffusion‐Based Generative Drug‐Like Molecular Editing with Chemical Natural Language,” Journal of Pharmaceutical Analysis 15 (2025): 101137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Cheng R., Liu T., Liao C., Wu X., Zhu L., and Zhang S., “Integrating Protein Language Models with Multimodal Embeddings to Accelerate Function Prediction of Uncharacterized Proteins,” International Journal of Molecular Sciences 27 (2026): 3891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Bjerregaard A., Groth P. M., Hauberg S., Krogh A., and Boomsma W., “Foundation Models of Protein Sequences: A Brief Overview,” Current Opinion in Structural Biology 91 (2025): 103004. [DOI] [PubMed] [Google Scholar]
  • 71. Crampon K., Giorkallos A., Deldossi M., Baud S., and Steffenel L. A., “Machine‐Learning Methods for Ligand–protein Molecular Docking,” Drug Discovery Today 27 (2022): 151–164. [DOI] [PubMed] [Google Scholar]
  • 72. Agoni C., Fernández‐Díaz R., Timmons P. B., Adelfio A., Gómez H., and Shields D. C., “Molecular Modelling in Bioactive Peptide Discovery and Characterisation,” Biomolecules 15 (2025): 15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Hajiashrafi T., Zekriazadeh R., Flanagan K. J., et al., “The Role of π–π Stacking and Hydrogen‐Bonding Interactions in the Assembly of a Series of Isostructural Group IIB Coordination Compounds,” Acta Crystallographica Section C Structural Chemistry 75 (2019): 178–188. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74. Khakzad H., Igashov I., Schneuing A., Goverde C., Bronstein M., and Correia B., “A New Age in Protein Design Empowered by Deep Learning,” Cell Systems 14 (2023): 925–939. [DOI] [PubMed] [Google Scholar]
  • 75. Li J., Cheng C., Ma J., and Liu G., “Geometric Point Attention Transformer for 3D Shape Reassembly,” arXiv (2024), 10.48550/arXiv.2411.17788. [DOI] [Google Scholar]
  • 76. Krikid F., Rositi H., and Vacavant A., “State‐of‐the‐Art Deep Learning Methods for Microscopic Image Segmentation: Applications to Cells, Nuclei, and Tissues,” Journal of Imaging 10 (2024): 131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. Shen C., Luo J., and Xia K., “Molecular Geometric Deep Learning,” Cell Reports Methods 3 (2023): 100621. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78. Brima Y. and Atemkeng M., “Saliency‐Driven Explainable Deep Learning in Medical Imaging: Bridging Visual Explainability and Statistical Quantitative Analysis,” BioData Mining 17 (2024): 18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79. Nerrise F., Zhao Q., Poston K. L., Pohl K. M., and Adeli E., “An Explainable Geometric‐Weighted Graph Attention Network for Identifying Functional Networks Associated with Gait Impairment,” arXiv (2023), 10.48550/arXiv.2307.13108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80. Verburgt J., Jain A., and Kihara D., “Recent Deep Learning Applications to Structure‐Based Drug Design,” Methods in Molecular Biology 2714 (2024): 215–234, 10.1007/978-1-0716-3441-7_13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81. Cremer J., Medrano Sandonas L., Tkatchenko A., Clevert D.‐A., and De Fabritiis G., “Equivariant Graph Neural Networks for Toxicity Prediction,” Chemical Research in Toxicology 36 (2023): 1561–1573. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82. Corso G., Somnath V. R., Getz N., Barzilay R., Jaakkola T., and Krause A., “Composing Unbalanced Flows for Flexible Docking and Relaxation,” in International Conference on Learning Representations 2025 (ICLR 2025, 2025), 27699–27722. [Google Scholar]
  • 83. Cao D., Chen M., Zhang R., et al., “SurfDock Is a Surface‐Informed Diffusion Generative Model for Reliable and Accurate Protein–ligand Complex Prediction,” Nature Methods 22 (2025): 310–322. [DOI] [PubMed] [Google Scholar]
  • 84. Pradeep P., Struble C., Neumann T., Sem D. S., and Merrill S. J., “A Novel Scoring Based Distributed Protein Docking Application to Improve Enrichment,” IEEE/ACM Transactions on Computational Biology and Bioinformatics 12 (2015): 1464–1469. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85. Lyu J., Irwin J. J., and Shoichet B. K., “Modeling the Expansion of Virtual Screening Libraries,” Nature Chemical Biology 19 (2023): 712–718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Pinheiro P. O., Jamasb A. R., Mahmood O., Sresht V., and Saremi S., “Structure‐Based Drug Design by Denoising Voxel Grids,” arXiv (2024), 10.48550/arXiv.2405.03961. [DOI] [Google Scholar]
  • 87. Zeng X., Wang F., Luo Y., et al., “Deep Generative Molecular Design Reshapes Drug Discovery,” Cell Reports Medicine 3 (2022): 100794. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88. Ivanenkov Y., Zagribelnyy B., Malyshev A., et al., “The Hitchhiker's Guide to Deep Learning Driven Generative Chemistry,” ACS Medicinal Chemistry Letters 14 (2023): 901–915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89. Cáceres E. L., Tudor M., and Cheng A. C., “Deep Learning Approaches in Predicting ADMET Properties,” Future Medicinal Chemistry 12 (2020): 1995–1999. [DOI] [PubMed] [Google Scholar]
  • 90. Lee J., Zhung W., Seo J., and Kim W. Y., “BInD: Bond and Interaction‐Generating Diffusion Model for Multi‐Objective Structure‐Based Drug Design,” Advanced Science 12 (2025): e02702. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91. Wang H., “Multi‐View Learning Framework for Predicting Unknown Types of Cancer Markers Via Directed Graph Neural Networks Fitting Regulatory Networks,” Briefings in bioinformatics 25 (2024): bbae546. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92. Xue F., Zhang M., Li S., et al., “SE(3)‐Equivariant Ternary Complex Prediction Towards Target Protein Degradation,” Nature Communications 16, no. 1 (2025): 5514, 10.1038/s41467-025-61272-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93. Jung W., Goo S., Hwang T., et al., ”Absorption Distribution Metabolism Excretion and Toxicity Property Prediction Utilizing a Pre‐Trained Natural Language Processing Model and its Applications in Early‐Stage Drug Development,” Pharmaceuticals 17 (2024): 382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94. Buonfiglio R., Recanatini M., and Masetti M., “Protein Flexibility in Drug Discovery: From Theory to Computation,” ChemMedChem 10 (2015): 1141. [DOI] [PubMed] [Google Scholar]
  • 95. Rudden L. S. P., Hijazi M., and Barth P., “Deep Learning Approaches for Conformational Flexibility and Switching Properties in Protein Design,” Frontiers in Molecular Biosciences 9 (2022): 928534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96. Zhong L., Li Y., Xiong L., et al., “Small Molecules in Targeted Cancer Therapy: Advances, Challenges, and Future Perspectives,” Signal Transduction and Targeted Therapy 6 (2021): 201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97. Martinelli C., Repetto M., and Curigliano G., Artificial Intelligence for Medicine, ed. Ben‐David S., Curigliano G., Koff D., Jereczek‐Fossa B A,, La Torre D., and Pravettoni G. (Academic Press, 2024), 37–45. [Google Scholar]
  • 98. Bienstock R. J., “AI/ML Methodologies and the Future‐Will they be Successful in Designing the Next Generation of New Chemical Entities?, Journal of Cheminformatics 17 (2025): 46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99. Liu M., Li C., Chen R., Cao D., and Zeng X., “Geometric Deep Learning for Drug Discovery,” Expert Systems with Applications 240 (2024): 122498. [Google Scholar]
  • 100. Zheng Y., Koh H. Y., Ju J., et al., “Large Language Models for Drug Discovery and Development,” Patterns 6, 2025, 101346. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101. Schneuing A., Harris C., Du Y., et al., “Structure‐Based Drug Design with Equivariant Diffusion Models,” Nature Computational Science 4 (2024): 899–909. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102. Putin E., Asadulaev A., Ivanenkov Y., et al., “Reinforced Adversarial Neural Computer for De Novo Molecular Design,” Journal of Chemical Information and Modeling 58 (2018): 1194–1204. [DOI] [PubMed] [Google Scholar]
  • 103. Sang C., Shu J., Wang K., et al., “The Prediction of RNA‐Small Molecule Binding Sites in RNA Structures Based on Geometric Deep Learning,” International Journal of Biological Macromolecules 310 (2025): 143308. [DOI] [PubMed] [Google Scholar]
  • 104. Zhu W., Ding X., Shen H.‐B., and Pan X., “Identifying RNA‐Small Molecule Binding Sites Using Geometric Deep Learning with Language Models,” Journal of Molecular Biology 437 (2025): 169010. [DOI] [PubMed] [Google Scholar]
  • 105. Gainza P., Sverrisson F., Monti F., et al., “Deciphering Interaction Fingerprints from Protein Molecular Surfaces Using Geometric Deep Learning,” Nature Methods 17 (2020): 1–192. [DOI] [PubMed] [Google Scholar]
  • 106. Das B., Kutsal M., and Das R., “A Geometric Deep Learning Model for Display and Prediction of Potential Drug‐Virus Interactions Against SARS‐CoV‐2,” Chemometrics and Intelligent Laboratory Systems 229 (2022): 104640. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107. Wu F., Wu L., Radev D., Xu J., and Li S. Z., “Integration of Pre‐Trained Protein Language Models into Geometric Deep Learning Networks,” Communications Biology 6 (2023): 876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108. Taylor R. and Wood P. A., “A Million Crystal Structures: The Whole Is Greater than the Sum of Its Parts,” Chemical Reviews 119 (2019): 9427–9477. [DOI] [PubMed] [Google Scholar]
  • 109. Lill M. A., “Efficient Incorporation of Protein Flexibility and Dynamics into Molecular Docking Simulations,” Biochemistry 50 (2011): 6157–6169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110. Stachowski T. R. and Fischer M., “Large‐Scale Ligand Perturbations of the Protein Conformational Landscape Reveal State‐Specific Interaction Hotspots,” Journal of Medicinal Chemistry 65 (2022): 13692–13704. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111. Huanbutta K., Burapapadh K., Kraisit P., et al., “Artificial Intelligence‐Driven Pharmaceutical Industry: A Paradigm Shift in Drug Discovery, Formulation Development, Manufacturing, Quality Control, and Post‐Market Surveillance,” European Journal of Pharmaceutical Sciences 203 (2024): 106938. [DOI] [PubMed] [Google Scholar]
  • 112. Mohamed E., Sirlantzis K., and Howells G., “A Review of Visualisation‐as‐Explanation Techniques for Convolutional Neural Networks and Their Evaluation,” Displays 73 (2022): 102239. [Google Scholar]
  • 113. Klenam D., “Artificial Intelligence in the Development of Structural and Functional Materials: A Transformative Frontier and Complementary Enabler of the Rational Alloy Design Framework,” Next Research 3 (2026): 101157. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analyzed during the current study.


Articles from Molecular Informatics are provided here courtesy of Wiley

RESOURCES