Abstract
Cancer remains among the most aggressive and treatment-resistant diseases, with a persistent failure of therapeutic strategies. Addressing the bottlenecks in cancer drug discovery, we present a feature-driven AI-integrated pipeline designed for systematic identification of repurposable drug candidates against druggable targets across diverse types of cancers. As proof, we applied this pipeline to glioma. We utilized Gen AI to identify an antiglioma target, alkylglycerone phosphate synthase (AGPS), a key enzyme in tumor metabolism and progression. Using a deep learning model, we screened over 5,76,510 compounds from the life chemicals high-throughput screening database for their potential to inhibit AGPS. ROC analysis of top candidates identified through graph neural network modeling and Glide docking yielded an AUC of 0.89, supporting the model’s ability to discriminate between active and inactive compounds. Top-scoring candidates were subjected to rigorous molecular dynamics (MD) simulations to assess the binding stability. Among them, F2881-0267 emerged with favorable drug-like properties. To evaluate the binding free energy landscape, we developed a hybrid deep learning model combining 3D convolutional neural networks and multilayer perceptrons. This framework integrates spatial features, molecular interaction fingerprints, and physics-based energy descriptors derived from MD trajectories. Our findings showcase the potential of this AI transformative model to streamline drug discovery workflows, which can be applied to other therapeutically relevant targets similar to AGPS.


1. Introduction
Artificial intelligence (AI) is redefining the landscape of drug discovery by accelerating the identification of therapeutic targets, optimizing compound screening, and predicting drug–target interactions with unprecedented precision. − Recent advances in deep learning-augmented virtual screening captures complex drug–target binding, convolutional neural networks (CNNs) trained on three-dimensional protein–ligand complexes (e.g., Pafnucy, DeepFusionNet), extract spatial interaction features, and outperform classical docking in pose and affinity prediction. , Concurrently, graph neural networks (GNNs) generalize CNNs to non-Euclidean molecular graphs, directly encoding the atomic connectivity and 3D structure of proteins and ligands. Ligand-centric deep learning models including sequence-based CNNs such as DeepDTA and multitask neural networks enable large-scale screening of extensive compound libraries, prioritizing potent small molecules and facilitating drug repurposing in oncology. Advanced generative models (e.g., GAN-based frameworks) are part of standard practices aimed at improving the speed and reliability of drug discovery processes. , These developments collectively offer a robust and scalable solution to overcome the time-intensive and failure-prone stages of traditional drug discovery.
However, most models rely on static protein–ligand snapshots rather than dynamic conformational ensembles, leading to incomplete representation of binding thermodynamics. Additionally, target-specific generalizability is constrained by data sparsity and model overfitting, particularly for underexplored cancer targets. Moreover, many frameworks operate in isolated prediction modules, lacking an integrated end-to-end pipeline that spans from target nomination to free energy estimation and lead optimization.
In this study, we introduce an AI-driven framework for drug repurposing and target-specific inhibitor identification with general applicability across various cancer types. To demonstrate its effectiveness, we applied this framework to glioma, one of the most aggressive and treatment-resistant forms of cancer. Glioma accounts approximately 38.7% of primary central nervous system tumors and are the most prevalent malignancies arising from glial cells in the brain and spinal cord. The World Health Organization (WHO) classifies gliomas into four grades based on malignancy, with grades I and II being less aggressive, while grades III and IV including glioblastoma multiforme (GBM) are highly invasive and resistant to conventional therapies. Despite advances in surgical resection, chemotherapy, and radiotherapy, high-grade gliomas remain difficult to treat. Recent shifts toward targeted and personalized therapies underscore the importance of identifying molecular biomarkers and understanding oncogenic signaling networks. Gliomas frequently exhibit mutations in critical regulatory genes such as TP53 (31–38%) and PTEN (24–37%), leading to disrupted apoptosis, genomic instability, and hyperactivation of the PI3K/AKT pathway. − Targeting aberrant protein–protein interactions and their upstream modulators presents a promising therapeutic avenue for glioma. In this context, we focused on AGPS, a metabolic enzyme involved in ether lipid biosynthesis, which is increasingly recognized for its role in cancer metabolism. Inhibition of AGPS has been shown to suppress the expression of protumorigenic lipids (LPA, LPAe, PGE2), reduce glioma cell adhesion and invasion, modulate the G0/G1 cell cycle, and influence protein–protein interactions critical to tumor signaling pathways, including cyclin D1, CD44, E-cadherin, and matrix metalloproteinases (MMP-2/9), highlighting its potential as a multifunctional therapeutic target in glioma. ,
Given its critical role in cancer progression, our study focuses on identifying a potent AGPS inhibitor through an AI-integrated framework with the aim of advancing targeted treatment strategies for glioma. Traditional drug discovery can take many years before a new medication is available on the market. So, drug repurposing is a more efficient alternative for treating various diseases. From this perspective, we delve into generative AI with protein–protein interaction (PPI) for glioma target identification and a deep neural network system to screen potential drug compounds with high precision.
2. Methods
2.1. AI-Assisted Literature Search for Target Identification
To identify glioma-specific therapeutic targets, we implemented a generative AI framework. First, relevant literature was systematically extracted by applying search terms “glioma associated proteins” and “glioma biomarkers” using the PubMed E-utilities API. A tailored Python script automated the retrieval of PubMed articles using disease-specific search terms. Titles and abstracts were parsed from XML-formatted responses, retaining only those with nonempty abstracts. The extracted metadata comprising PubMed IDs (PMIDs), titles, and abstracts was compiled into a structured data set for downstream processing.
In the second phase, a custom-generative AI framework was developed to identify potential protein targets. We employed a dual-model architecture integrating a Sentence-BERT embedder (paraphrase-MiniLM-L6-v2) with a transformer-based generative model (FLAN-T5-XXL). Abstract texts were embedded into a vector space and indexed via FAISS for efficient semantic retrieval. Upon receiving a protein inference query, relevant abstracts were retrieved and context-specific prompts were constructed for generative inference. The model synthesized protein candidates by conditioning on literature-derived contextual embeddings, producing a refined list of glioma-associated proteins for subsequent validation and virtual screening.
2.2. PPI Network Construction and Visualization of Regulated Genes
The protein–protein interaction network was generated using STRING, a search tool for the Retrieval of Interacting Genes/Proteins (https://string-db.org/). The PPI network of upregulated and downregulated genes was built using STRING, considering interactions with a confidence score >0.4 as statistically significant. The data were subsequently imported into Cytoscape version 3.10.1 for visualization of the interaction network. Cytoscape software is specifically designed for the analysis and visualization of large-scale networks, offering enhanced flexibility for importing additional data mapping it onto network visualizations.
2.3. Homology Modeling and Protein Preparation
Homology modeling of AGPS was performed since there was no structure for Human AGPS in the Protein Data Bank (PDB). We utilized the AGPS protein complex with ZINC69435460 (Organism: Guinea Pig, PDB ID: 5AE1, Resolution-2.10 Å) and AGPS wild type (Organism: Guinea Pig, PDB ID: 4BBY, Resolution-1.90 Å) to generate respective homology models using Homology Modeling Wizard in Schrodinger and SWISS-MODEL (https://swissmodel.expasy.org/). The selection of templates with optimal sequence alignment facilitated the identification of a modeled structure that displayed a minimal RMSD, thereby confirming the robustness of our computational predictions. The protein structure was prepared using the Preparation Wizard in Schrodinger software. This wizard identifies and corrects defects in the protein structure by replacing hydrogen atoms, removing water molecules, and adding zero-order bonds to metals and disulfide bonds. Following this, the structure was optimized and minimized using the OPLS_2005 force field for all atoms with a maximum RMSD of 0.30 Å to achieve the most stable conformation with the minimized energy.
2.4. Deep Neural Network for Binding Affinity Prediction
To screen the ligands for binding potential, we previously developed an attention-based deep neural network architecture trained on pharmacophore-derived molecular features extracted from 3D SDF files of 5,76,510 compounds using RDKit. Ligand descriptors, including rotatable bonds, hydrogen bond donors/acceptors, molecular weight, LogP, and estimates of acidic, basic, and hydrophobic properties, were computed and structured for input. In parallel, cavity features of the target AGPS were extracted using BioPython’s PDBParser, capturing atom-level attributes (e.g., element type, hydrogen bonding, hydrophobicity) and spatial centroid coordinates for each residue in the defined cavity. Binding affinity was computed as a weighted dot product between ligand and cavity vectors, normalized, and labeled for binary classification. The model comprised ligand and protein embedding layers, a multihead attention mechanism to capture context-aware interactions, and a fully connected output layer, trained using binary cross-entropy loss and optimized with Adam. Batch training and affinity-based labeling enabled robust classification of successful ligand–protein interactions.
2.4.1. Ligand Preprocessing and Pharmacophore Feature Extraction
Ligand molecules (SDF files) are read using RDKit, sanitized, and hydrogenated before descriptor computation. We compute standard cheminformatics descriptors via RDKit’s Descriptors and rdMolDescriptors modules: molecular weight (MolWt), octanol–water partition coefficient (MolLogP), number of hydrogen bond donors (NumHDonors) and acceptors (NumHAcceptors), number of rotatable bonds (NumRotatableBonds), topological polar surface area (TPSA), as well as Lipinski-style counts of acidic (HBA) and basic (HBD) functional groups, aromatic rings, aliphatic rings, and aliphatic carbocycles. These features, nine in total in our implementation, form a tensorized fingerprint for each ligand. In practice, we collect these values into a PyTorch tensor and associate it with the ligand identifier.
A custom PyTorch data set is used to batch-process all SDFs. All .sdf files in the ligand folder are enumerated, and each molecule index in each file is recorded (e.g., (filename, index) tuples) so that the data set can retrieve any ligand on demand. The __getitem__ method loads the specified molecule, computes its descriptor vector (via extract_pharmacophore_features), and returns a 1 × 9 feature tensor plus feature names and ligand name. In this way, the entire ligand library can be streamed into memory-efficient batches. The resulting descriptors for all ligands can then be saved (e.g., to a pandas DataFrame or Excel) for downstream modeling.
2.4.2. Protein Structure Featurization
The protein 3D structure is processed with ProDy to extract the active site and construct an atomic graph. The target PDB is parsed (chain “A”, residues [616,617] in our case), and the subset of atoms in the active-site residues is selected. From these atoms, we compute the active-site centroid by averaging coordinates
where N is the number of active-site atoms and r i is their coordinates. If the selection fails or contains too few atoms, an exception is raised to ensure data integrity.
Atomic-level features are then assigned to each active-site atom. We define a feature matrix of shape N × F (here F = 10) for residue-specific properties. In our prototype code, these features are placeholders (zeros) for quantities like atomic charge, solvent exposure, or ANM modes, but in general, this could include partial charges, atom types, or normal-mode amplitudes. We also associate per-atom electrostatic potential and hydrophobicity values (here, dummy zeros) and compute a global active-site volume (placeholder). The “surface” of the site is represented by the 3D coordinates of these atoms themselves (treated as surface points). All of these are stored in a feature dictionary:
-
1.
residue_features: an N × 10 matrix of atomic descriptors (e.g., partial charges, atom types).
-
2.
Electrostatic: length-N vector of atomic electrostatic potentials.
-
3.
Hydrophobicity: length-N vector of atomic hydrophobicities.
-
4.
Volume: scalar pocket volume.
-
5.
Anm_modes: an array of atomic displacements (dummy ANM modes of size 3N × 3).
-
6.
Surface_points: N × 3 array of atomic coordinates.
-
7.
Centroid: 3-vector of the site centroid.
Finally, the active site is represented as a graph for the GNN input. We treat each atom as a node with the above features (in practice, concatenated or summarized into a per-node vector). Edges are added between any two atoms whose Euclidean distance is below a threshold (here 5 Å). Concretely, we compute the pairwise distance matrix of the active-site coordinates and add edges (j) whenever ∥r i – r j ∥<5.0 Å. The resulting undirected atomic graph captures the local geometry of the pocket. (The protein data object for the GNN thus contains x = atomic feature matrix, pos = atomic coordinates, and edge_index as the adjacency list.)
2.4.3. Ligand Quantum and Graph Representation
Each ligand is also represented as a 3D graph featuring both atomistic and pharmacophoric properties. First, a single 3D conformation is generated: hydrogens are explicitly added (Chem.AddHs), an initial embedding is computed (AllChem.EmbedMolecule with fixed random seed), and then, a force-field optimization (MMFF94) is run. This ensures a chemically reasonable geometry and allows the computation of physicochemical properties.
From the optimized molecule, we extracted per-atom features and global descriptors. A PyTorch Geometric Data object is constructed with:
-
1.
Node features x: for each atom, we use a vector [a] where Z = atomic number, d = degree (number of bonded neighbors), q = formal charge, h = hybridization index, and a∈{0,1} indicating aromaticity. These features are converted to a floating-point tensor of shape n × 5 (for n atoms).
-
2.
Edge indices: every bond in the molecule gives an undirected edge (v) (encoded in COO format). If multiple bonds exist, they can be represented by parallel edges or a bond order feature, but here we simply record connectivity.
-
3.
Node positions (pos): an n × 3 array of the atomic coordinates from the optimized conformer. Including the 3D positions enables the model to use geometric information.
In parallel, we computed several global “pharmacophoric” 3D descriptors of the ligand. The code extracts (or falls back to) values for electrostatic potential (ESP) at a reference point (set to 0.0 if no explicit calculation is done), the principal moment of inertia PMI1 (shape anisotropy), and the normalized principal moment ratio NPR1 (often interpreted as steric hindrance). If these descriptors are unavailable (e.g., if the RDKit version lacks NPR1/PMI1), we use a default of 0.0. A placeholder “affinity” field is also initialized to 0.0 for a later GNN prediction. In addition, we compute standard 1D/2D molecular descriptors via RDKit (MolWt, MolLogP, NumRotatableBonds, TPSA, HBA, HBD, ring counts, etc.) and Gasteiger partial charges (to obtain maximum and minimum charges, although our pipeline only reads the first atom’s charge due to a coding quirk). These descriptors are stored in a dictionary keyed by name.
Thus, each ligand yields a GNN graph (Data(x, edge_index, pos)), a surface point set (its atomic coordinates), a small set of 3D descriptors (ESP,PMI,NPR), and its RDKit scalar features (molecular weight, logP, rotatable bonds, etc.). Together, these capture both the quantum-chemical and the topological character of the ligand.
2.4.4. Entropy and Geometric Scoring
We estimate the conformational entropy loss from ligand flexibility. Following Tidor and co-workers, the torsional entropy change is approximated as
where n rot is the number of rotatable bonds and R is the gas constant (1.987cal·mol–1·K –1 in our code). This formula reflects the loss of rotational degrees of freedom upon binding. For each ligand, we compute ΔS torsion using its rotatable-bond count and store it as an entropy loss (in energy units at the chosen temperature).
We also quantify the ligand’s spatial relationship to the binding site by distance metrics. Using the ligand’s atomic positions and the protein site centroid c, we compute the Euclidean distance from each ligand atom r i to c. From these, we recorded the mean, minimum, and maximum distances
In practice, we use the average distance D avg as a metric of geometric complementarity (the smaller the D avg, the deeper the ligand sits in the pocket). This distance enters the final scoring (below) as the term D geom = 1/(1 + D avg), so that closer ligands obtain a higher score.
2.4.5. Physics-Informed Binding Score
We define a composite binding score E total that combines surface complementarity, electrostatics, GNN prediction, geometric fit, and entropy. Formally
where each term is normalized to be in a comparable range. The terms are
-
1.
Cosine surface similarity S cos: we compute the cosine similarity between the protein pocket surface points and ligand surface points. Concretely, we build distance matrices of all protein–ligand point pairs and compute the mean cosine distance. Then, S cos = 1 – ⟨d cosine⟩ so that higher S cos means more similar surfaces.
-
2.
Electrostatic correlation ρelec: we compute the Pearson correlation coefficient between the protein active-site electrostatic potential (length-N vector) and the ligand’s ESP value. This captures whether positive/negative regions align. Numerically, ρelec = corrcoef(V protein,V ligand).
-
3.
GNN affinity A GNN: the predicted affinity score from the trained GNN model (see Section 6).
-
4.
Geometric score D geom: defined as 1/(1 + D avg), where D avg is the mean ligand-centroid distance. This ranges in (0,1] and rewards ligands that occupy the pocket interior (larger D geom means closer on average).
-
5.
Entropy score S entropy: proportional to the (negative) entropy loss, here taken as -ΔS torsion/R so that higher entropy loss yields a lower score.
In our implementation, we set α1 = 0.3, α2 = 0.2, α3 = 0.2, α4 = 0.2, α5 = 0.1, and compute
This weighted sum produces a final score reflecting both data-driven (GNN) and physics-based considerations. (Different choices of α i could be calibrated on known binders, but here they serve as illustrative hyperparameters.)
2.4.6. Deep Graph Neural Network Model
The core predictive model (“PharmaGNN”) is a dual-branch graph neural network with attention. One branch encodes the protein site graph, and the other encodes the ligand graph. Each branch has two stacked graph-attention layers (GATv2Conv) followed by readout pooling. Specifically:
Protein branch: input node dimension is 10 (from residue_features). We apply a GATv2Conv(10 → 256) layer with 4 attention heads, producing 256-dimensional features per head (total 1024). A second GATv2Conv(1024 → 512) (single head) then outputs 512 features per node.
Ligand branch: input dimension is 5 (atomic features). We apply GATv2Conv(5 → 256) with 4 heads (output 1024), then GATv2Conv(1024 → 512) to yield 512 features per atom.
After the graph convolutions, we apply global mean pooling separately to each graph. If h i are the node embeddings in the protein graph, the protein’s graph embedding is
and similarly for the ligand . This yields fixed-size 512-dimensional vectors h prot,h lig for the protein pocket and ligand, respectively.
We then fuse protein and ligand representations via a cross-attention mechanism. We treat h prot as the query Q, and h lig as key K and value V, into a standard multihead attention layer with 8 heads (embedding size 512). The scaled dot-product attention is given by
where d k = 64 is the per-head key dimension. Concretely, Q and K,V are 1 × 512 vectors, so the output is a 1 × 512 vector h attn. Multiple heads are concatenated and projected.
Finally, the protein embedding h prot and the attended ligand context h attn are concatenated into a 1024-dimensional vector. A small multilayer perceptron (MLP) decodes this to a single affinity prediction
where the MLP has a hidden width of 512, layer normalization, and LeakyReLU activation, and outputs a scalar. (All componentsGATv2 and multihead attentionare implemented via PyTorch Geometric and PyTorch modules.) The result is the GNN’s affinity score A GNN.
2.4.7. Training and Validation Setup
In this prototype, no experimental binding affinities are available; therefore, we use placeholder labels. We assemble a DataFrame of ligands and assign an arbitrary “Experimental” value (e.g., zero) for each. The pipeline’s PharmaValidator then performs a 5-fold K-fold cross-validation: in each fold, we split the ligands into training and test sets, train (hypothetically) on the train set, and evaluate on the test set. As a metric, we compute the coefficient of determination R 2 between the experimental affinities and the model’s predicted scores. Since all experimental labels are zero, this simply measures variance explained by the constant predictor; in practice, one would replace these placeholders with real data.
Feature standardization is applied using scikit-learn’s StandardScaler. That is, ligand descriptors and predicted scores are rescaled to zero mean and unit variance before training to improve numerical conditioning. In our code, a StandardScaler object is created and could be fit to the training data; in the current stub, it serves to indicate that all inputs should be normalized. After cross-validation, the average R 2 is reported as a rough sanity check of model behavior. In a real deployment, one would withhold a validation set of known affinities and possibly perform model selection on weights α i . Following the prediction phase, we conducted Pearson and Spearman correlation analyses between the model’s output scores and the corresponding docking scores to quantitatively assess the model’s predictive validity.
2.5. Ligand Preparation
We downloaded the three-dimensional structures of 5,76,510 lakh compounds selected from the high-throughput screening (HTS) collection in the Life Chemicals database (https://lifechemicals.com/). The LigPrep tool from Schrodinger Molecular Modeling Suite version 14.4 was used to prepare the ligands. Each structure was ionized at a pH of 7.0 using the Epik ionizer and generated at most 1 specified chirality of each ligand.
2.6. Grid Generation
A receptor grid with coordinates 20 × 20 × 20 Å was generated using the “receptor grid generator” for the active-site residues His616 and His617. Accurate measurements were taken into account, utilizing a partial cutoff of 0.25 and a van der Waals scale factor of 1.00. The highlighted atoms are the active site’s V loop.
2.7. Virtual Screening
Virtual screening is a technique used to dock a large number of ligands to identify the most promising candidates for a specific target. We performed virtual screening of the prepared 5,76,510 compounds and the known inhibitor (zinc 69435460) in the generated grid using the GLIDE module of the Schrodinger Molecular Modeling Suite. Three levels of precision were utilized: high throughput virtual screening (HTVS), standard precision (SP), and extra precision (XP). Specifically, 20% of the compounds were screened through HTVS, 20% through SP, and 50% through XP. After each screening phase, the batch of ligands was refined, focusing on those with the highest docking scores. Using the postprocess with Prime MM GBSA (Molecular Mechanics Generalized Born Surface Area), the binding free energy was predicted.
This comprehensive virtual screening method led to the identification of top-ranking hits based on Glide docking scores. The docking score is further used for validation of our DNN model, delivering an additional layer of insight for AI-assisted screening.
2.8. ADMET
AI Drug Lab (https://ai-druglab.smu.edu/admet), an AI-based prediction, was used to predict the drug-likeness of the compound through pharmacokinetic and pharmacodynamic parameters, including absorption, distribution, metabolism, excretion, and toxicity (ADMET).
2.9. IFD
The induced fit docking (IFD) process is a receptor-based computational method that involves docking the receptor with the ligands that are induced to fit in different conformations. Introducing new ligands into the loop can cause conformational changes, resulting in movement within the histidine-rich loop, which is a challenge to the docking algorithm; IFD helps improve binding accuracy in flexible protein structures. We conducted IFD for the compounds that exhibited a high binding energy and docking scores in virtual screening. This approach helps to eliminate steric conflicts and identify ligand conformations that have low energy.
2.10. Molecular Dynamics
MD is a method used to assess the behavior of a compound in a dynamic environment. The stability of the complex in a dynamic environment is analyzed using the Desmond module of Schrodinger. The compounds with high docking scores were further analyzed by using MD. A solvate model TIP3P was used in an orthorhombic box shape to minimize the structure, and Na and Cl ions were added to neutralize the solvent system. Molecular dynamics simulations were conducted for 200 ns at a temperature of 300 K and 1.01325 bar pressure using the Nose–Hoover chain as the thermostat and the Martyna–Tobias Klein barostat methods with isotropic coupling style in a force field of OPLS 2005. The results of the MD analysis were evaluated using various metrics, including RMSD, RMSF, hydrogen bonds, MM GBSA, Rg, and SASA. Additionally, FEL analysis and principal component analysis (PCA) were performed as part of the post-MD analysis.
2.11. Binding Free Energy Prediction from MD Trajectories Using a Hybrid Deep Learning Model
To predict the ligand–protein binding free energy (ΔG bind) with atomic precision, we developed a hybrid deep learning framework trained on MD trajectories and trajectory-derived structural features.
2.12. Trajectory Processing and Featurization
Protein–ligand complexes were simulated using explicit-solvent MD, and trajectories were postprocessed to exclude solvent and ion atoms. Each frame of the trajectory was featurized using three complementary descriptors. Initially, the 3D voxel grids encode the spatial and physicochemical properties V ∈ R (C × D × D × D) with C = 10 channels representing atomic types, hydrogen bond donors/acceptors, hydrophobicity, and SASA. To compute the interaction fingerprints (f in ∈ R 128), we have used the ProLIF library. Characterizing the physical properties (f phys ∈ R 3) represents the electrostatic and van der Waals energies and the estimated solvation contribution.
The voxel grid is centered on the ligand centroid and spans a cubic region of radius r = 10 Å and resolution ρ = 1 Å. Each atom feature is encoded as
| 1 |
where ∥ c (a) is an indicator function for the atom’s channel and δ ijk (a) is a discrete delta function for its voxel location.
The calculation of electrostatic interaction by Coulomb’s law is as follows
| 2 |
where q i and q j are the partial atomic charges, r ij is the interatomic distance, ϵ is the dielectric constant, and δ is the numerical stability constant. Van der Waals interactions were modeled using the Lennard-Jones potential
| 3 |
Solvation energy estimation is as follows
| 4 |
2.13. Hybrid Model Architecture
To capture the complex biophysical interactions that govern protein–ligand binding, we designed a hybrid neural network architecture that integrates structural, interaction-based, and physical features derived from MD trajectories. The architecture consists of three parallel branches, each specializing in a distinct representation modality.
2.13.1. 3D CNN for Voxelized Spatial Encoding
The 3D CNN ingests voxelized spatial data. The voxel tensor encoded the spatial arrangement and physicochemical context of the protein–ligand interface. The CNN processes V as
| 5 |
The CNN comprises the first convolutional 3D (Conv3D) layer with 32 filters, kernel size 3 × 3 × 3, followed by BatchNormalization 3D (BatchNorm3D) for normalization and rectified linear unit function (ReLU) as the activation function. The second layer is the max pooling layer with a configuration of 2 × 2 × 2. The second Conv3D layer contains 64 filters with a kernel size of 3 × 3 × 3 and a BatchNorm3D and ReLU. The final layer delineates the maximum pooling size with 2 × 2 × 2. The output feature map is flattened to form vector hcnn.
2.13.2. MLP for Interaction Fingerprints
A separate MLP processes the ProLIF-derived interaction fingerprint f int ∈ R 128, encoding residue level interactions such as hydrogen bonds, hydrophobic contacts, and π–π stacking
| 6 |
This branch applies 64- and 32-unit linear layers with a ReLU and dropout (0.3) layer, having a final output h fp.
2.13.3. MLP for Physics-Based Energy Terms
Another MLP processes coarse-grained physicochemical descriptors f phys ∈ R 3 representing three parameters, electrostatic energy, van der Waals energy, and solvation energy. The transformation is
| 7 |
This MLP includes a 16 unit linear layer with ReLU and a dropout layer of 0.3.
2.13.4. Final Regression Head
The final prediction combines all three modalities through concatenation
| 8 |
here, ∥ denotes vector concatenation, and the output ΔG frame ∈ R represents the predicted binding free energy for a single MD frame.
2.13.5. Postprediction Aggregation and Free Energy Estimation
For N trajectory frames, model predictions were aggregated using a Boltzmann-weighted ensemble approach
| 9 |
where R = 1.987 × 10–3 kcal\mol–1 and T = 298 K. This corrects for thermal sampling by emphasizing low-energy conformers.
2.14. Entropy Estimation
To estimate configuration entropy, atomic coordinates from a trajectory framer were flattened into x i ∈ R 3n , and the covariance matrix ∑ ∈ R 3n × 3n was computed using the formula
| 10 |
where λ j are the eigenvalues of sigma and ε is a regularization constant. The final corrected free energy is
| 11 |
2.15. Model Training and Evaluation
The model was trained by using the mean-square error (MSE) loss
| 12 |
Optimization of this model was performed using AdamW with a learning rate of 1 × 10–4 and a batch size of 4. This hybrid model enables the integration of spatially rich CNN features, interaction-driven biochemical patterns, and physically grounded energy terms, capturing binding dynamics on multiple scales. By training on per-frame MD data, the model generalizes well to unseen conformations and allows robust estimation of binding free energies across sampled states.
3. Results
3.1. Generative AI and PPI Framework for Glioma Target Identification
Streamlining the target protein identification for glioma, we implemented generative pretrained transformer google/flan-t5-xxl framework, trained with 33117 glioma-associated protein literature to induce a domain-specific contextual bias. Further, we ranked the proteins by incorporating features such as undercharacterization in the literature and functional relevance to the glioma biology. This integrative strategy yielded AGPS as a candidate with a score of 0.972. To prioritize the potential target from the identified candidates, we performed protein–protein interaction analysis using the STRING database. The AGPS emerged as a key node with a high confidence score of 0.73. Supporting evidence of AGPS laid the foundation of its role in glioma progression and tumorigenesis. Moving beyond traditional heuristic approaches, we deployed a generative AI-driven method to contextualize and rank glioma-associated targets with therapeutic potential (Figure ).
1.
Identification and prioritization of glioma-associated proteins using a large language model framework. A schematic overview of the computational pipeline leveraging a transformer-based large language model (flan-t5-xxl) to systematically extract, screen, and prioritize glioma-associated proteins from published literature. (a) Relevant keywords (e.g., glioma-associated proteins, glioma biomarkers) were used to retrieve articles from PubMed. After manual and automated screening, entities were evaluated based on their functional relevance, extent of characterization, and presence in the existing literature. (b) The outputs were integrated with advanced proteomic visualization tools, including expression analysis and distribution plots. (c) Further validated through protein–protein interaction (PPI) network analysis using the STRING database. This integrative approach enables discovery and prioritization of functionally relevant but underexplored glioma biomarkers for potential downstream therapeutic or diagnostic targeting.
3.2. Homology Modeling of AGPS
Since the crystallized structure of human AGPS is not available in the PDB (https://www.rcsb.org/), a high-confidence structure was generated by homology modeling using the Swiss model. The resulting structure demonstrated 99.85% sequence similarity to the AGPS template with PDB ID 4BCA. The modeled structure of AGPS using PDB ID 4BBY exhibited the lowest root-mean-square deviation (RMSD) value of 0.100 Å, outperforming the 5AE1 structure.
Stereochemical quality assessment, based on deviation from ideal bond angles in the Ramachandran plot, is indicated by an average deviation of 1.62 Å, supporting the structural reliability of the model. Notably, 94.53% of residues were located within the most favorable regions, further validating the model’s conformational stability (Figure ). The final modeled structure comprised chains A and B, along with a flavin adenine dinucleotide (FAD) cofactor, which is crucial for the AGPS functionality. Study of Nenci et al. (2012) has demonstrated that FAD facilitates alkyl–acyl transformations without altering the substrate’s redox state. The alignment analysis indicates a significant degree of compositional conservation, highlighting the structural and evolutionary similarities between the two proteins. Most residues, particularly the predominant hydrophobic (A, L, V, and I) and charged (E, K, and D) amino acids, exhibit minimal variations, suggesting the preservation of critical physicochemical properties necessary for core folding, domain architecture, and functional motifs. These interactions contribute to the interface between the beta strands and the alpha helices. A closer inspection of these contact sites illustrates a recurring pattern where alanine or valine protruding from the β barrel strands engages with leucine or isoleucine residues.
2.
High-confidence template-based model reveals distinct binding-site architecture with rigorous geometric validation: (a) homology-modeled dimeric structure of the target protein rendered as a gray cartoon, with conserved active-site residues highlighted as magenta spheres. The inset shows a close-up of the binding cleft (V-loop), positioned at the interface between the FAD and cap domains where two representative side-chain conformations (teal and salmon sticks) illustrate key interactions and modeling confidence; (b) side-by-side comparison of amino acid frequencies in the target sequence (blue bars) versus the chosen template (red bars), demonstrating high conservation across residue types and supporting the reliability of the sequence alignment used for model building; (c) Ramachandran plot of backbone φ/ψ torsion angles for all modeled residues, with green contours denoting the most favored regions and blue shading the allowed regions; over 90% of residues fall within the core, reflecting excellent stereochemical quality and minimal outliers.
To further validate the modeled structure, we superimposed it with the AlphaFold-predicted human AGPS structure (https://deepmind.google/science/alphafold/). The alignment yielded a Cα RMSD of 0.960 Å and an alignment score of 0.039, indicating a strong structural concordance (Figure ). Notably, the binding pocket exhibited prominent spatial conservation, supporting the reliability of the modeled geometry for subsequent docking and molecular dynamics analyses.
3.
Superimposition of the homology-modeled human AGPS structure (green) with the AlphaFold-predicted AGPS model (magenta), showing a high degree of structural concordance between the two conformations. The inset highlights the active-site region (V site), demonstrating a close spatial overlap of key secondary structural elements.
3.3. Deep Learning Architecture for HTS Compounds Screening
Based on our previous publication on Nipah virus, we retrained the model with the MD trajectories and found that the model was stable and produced consistent prediction results. The previously validated time-efficient deep learning pharmacophore framework was used to screen a library of 5,76,510 ligands against AGPS, using pharmacophore properties.
The predictive pipeline retained its ability to differentiate high-affinity binders, successfully identifying as F3309-0655, F2881-0267 top-scoring compounds with predicted binding affinity values comparable to or exceeding known reference inhibitors. Following the prediction phase, we conducted Pearson and Spearman correlation analyses between the model’s output scores and corresponding docking scores. The resulting coefficients (Pearson r = – 0.47, Spearman ρ = – 0.51) indicate a moderate negative correlation, supporting the predictive validity and ranking consistency of our screening model (Figure ). These findings confirm the robustness and transferability of our approach across distinct protein targets and highlight its applicability in accelerating virtual screening workflows. The binding affinity predictions from the DNN model are given in Table S1.
4.
Correlation analysis between deep learning-based pharmacophore screening scores and docking scores for AGPS inhibitors.
Pearson and Spearman correlation analyses were performed to evaluate the relationship between predicted binding affinity scores generated by the deep learning pharmacophore model and the corresponding molecular docking scores for screened compounds. A moderate negative correlation was observed (Pearson r = – 0.47; Spearman ρ = – 0.51), indicating consistency between model-based ranking and structure-based docking evaluation. These results support the predictive reliability and transferability of the screening framework for identifying candidate AGPS inhibitors.
3.4. Evaluation of DNN Efficiency in Virtual Screening
This study investigated the molecular docking scores in classifying small molecules as biologically active or inactive within a virtual screening pipeline. We utilized a multitiered virtual screening approach to find the potential restraints targeting a specific protein. Multiple docking precision levels, HTVS, SP, and XP were applied to retain the top-scoring compounds after each iteration. From the 576,510 compounds and the control compound (ZINC69435460), a total of 3930 top-ranking hits were identified by the screening conducted based on the GLIDE module Also, the validation of the DNN model yielded an AUC score of 0.89, which indicated a moderate discriminatory power between the compounds. Comparative scoring results for the validation are indicated in Table S1 and Figure S1.
3.5. Binding Affinity of the Selected Ligand Compounds
To identify ideal inhibitors, the filtered compounds were further evaluated based on their docking scores. Among the screened compounds, F3309-0655 shows a docking score of −11.629 kcal/mol and a binding energy of −66.53 kcal/mol (Figure ). The interacting residues include PHE428, GLN425, and TYR580. F5463-0122 and F6472-097 also showed high docking scores (−11.498 and −11.218 kcal/mol, respectively), forming stable interactions with residues, such as PHE428, TRP131, and ARG566. F2881-0267 demonstrated a strong binding profile with a docking score of −11.055 kcal/mol, along with a binding energy of −58.71 kcal/mol, engaging TYR580, TRP131, and GLN425. F1885-1052 shows a docking score of −10.735 kcal/mol and a binding energy of −65.29 kcal/mol forming a hydrogen bond with PHE428. Prominently, F0554-0714 showed a docking score of −10.039 kcal/mol and the highest binding energy of −67.42 kcal/mol. They form interactions with LYS433, ARG503, PHE426, PHE428, and TYR580. In contrast, the control compound Zinc69435460 has a very low docking score of −1.956 kcal/mol and a very low binding energy of −25.48 kcal/mol. It shows interactions with ARG503, LYS433, and ALA496. A summary of docking scores, binding energies, and interacting residues is presented in Table S2.
5.
Molecular interaction landscapes highlight differential residue engagement across ligand complexes. The two-dimensional interaction diagrams present a detailed analysis of the molecular interactions between the protein and six distinct ligands (a) F3309-0655, (b) F5463-0122, (c) F6472-0974, (d) F2881-0267, (e) F1885-1052, (f) F0554-0714, and (g) Zinc69435460, derived from representative frames of docking studies. Residue interactions are visualized using a color-coding scheme: polar interactions are indicated in light blue, hydrophobic contacts in light green, positively charged residues in blue, and negatively charged residues in red or pink, where applicable. Although the ligands occupy overlapping binding pockets, they engage distinct subsets of amino acid residues, which highlights their varying pharmacophoric characteristics. These interaction fingerprints provide a structural foundation for understanding ligand binding specificity and offer valuable insights into the structure–activity relationships that are critical for inhibitor design.
3.6. ADMET Analysis of the Selected Ligands
The six shortlisted compounds were ranked based on ADMET profiles (Table S3). F5463-0122 showed the most favorable properties, followed by F2881-0267 and F6472-0974. F3309-0655 and F1885-1052 ranked fourth and fifth, respectively, while F0554-0714 exhibited the least desirable ADMET characteristics. Compounds F0554-0714, F6472-0974, and F5463-0122 demonstrated favorable Caco-2 permeability (log P > −5.2), and all candidates showed positive human intestinal absorption. Only F5463-0122 and F2881-0267 met oral bioavailability criteria (>50%). All compounds crossed the blood–brain barrier (BBB > 30) and were largely nontoxic based on DILI prediction, except for F2881-0267. hERG inhibition was acceptable (<45) for F6472-0974, F5463-0122, and F1885-1052.
3.7. Active-Site Guided Docking of Lead Candidates
All compounds except for F0554-0714 displayed favorable docking scores, binding energies, and ADMET properties. However, the ligand docking of the compounds failed to show interactions with the key active residues. Henceforth, we decided to use IFD for the remaining five compounds, allowing for receptor flexibility and conformational effects on the ligand interactions (Figure ). F5463-0122 exhibits a docking score of −10.302 kcal/mol with residues HIS616, HIS617, SER301, GLY238, GLY236, and PRO238. This interacts with HIS616 and HIS617, which are the active-site residues. Additionally, F2881-0267 and F6472-0974 also interact with active-site residues. The docking scores for the remaining compounds are detailed in Table S4. The difference in the virtual screening and IFD may be due to the mobility of the amino acids of the protein, which is not seen in XP docking.
6.
3D binding pocket conformational occupancy maps reveal ligand-induced adaptation across glioma target complexes. Representative snapshots of 5 glioma protein–ligand complexes, obtained through IFD, illustrate the dynamic accommodation of structurally diverse ligands (shown in red stick representation) within the protein binding site (gray ribbon and sticks). Panels (a) F3309-0655, (b) F5463-0122, (c) F6472-0974, (d) F2881-0267, and (e) F1885-1052 delineate distinct binding poses across different targets, highlighting the conformational plasticity of critical binding site residues upon ligand engagement. Key noncovalent interactions, including hydrogen bonds (represented by yellow dashed lines) and π–π contacts (illustrated by cyan dashed lines), are annotated to reveal the contributions of specific residues, namely, HIS 616, HIS 617, ARG 419, GLY 236-238, TRP 96, and SER 304, in stabilizing ligand binding. Importantly, these residues exhibit side-chain reorientations or loop shifts in response to ligand binding, demonstrating localized induced fit effects that are absent in rigid docking protocols.
3.8. Molecular Dynamics of AGPS–Ligand Complexes
By evaluating the ligand docking, ADMET, and IFD results, we selected five compounds with high scores for a molecular dynamics simulation period of 200 ns The atomic deviation of binding and unbinding states unveils the occurrence of stable interactions, as reflected by consistently low RMSD values. The unbound protein exhibited an average RMSD of 2.92 Å, with noticeable fluctuations after 50 ns and pronounced deviations between 95–100 ns and 130–145 ns. In contrast, complexes with F5463-0122 and F6472-0974 showed improved stability, with average RMSD values of 2.34 and 2.36 Å, respectively. F2881-0267 also demonstrated stable binding (2.50 Å), while F3309-0655 and F1885-1052 showed slightly higher RMSD values of 2.78 and 3.03 Å, respectively. Specifically, the F6472-0974 complex displayed the most stable RMSD trajectory, with only minor fluctuations between 40 and 120 ns (Figure a). F5463-0122 exhibited moderate fluctuations between 80 and 100 ns, while F3309–0655 initially rose to 3.5 Å but stabilized after 90 ns. The F1885-1052 complex stabilized after 60 ns, with minimal variation between 105–120 ns and 140–150 ns. F2881-0267 showed an initial fluctuation, followed by consistent stability throughout the trajectory. All complexes remained within the acceptable RMSD range (<3 Å) for globular proteins, with F2881-0267 and F6472-0974 exhibiting the most stable profiles.
7.
Dynamic stability and flexibility profiling reveal distinct binding behaviors of antiglioma candidates. (a) The RMSD analysis quantifies the temporal structural deviation of the protein backbone across the 200 ns MD simulations. The protein only exhibits increased global fluctuations, indicating an increased conformational flexibility. In contrast, ligand-bound complexes, particularly those with F2881-0267 and Zinc69435460, exhibit lower RMSD profiles, indicating enhanced structural stabilization upon ligand binding. Conversely, the F1885-1052 and F3309-0655 complexes show elevated RMSD values during intermediate simulation windows, suggesting induced flexibility or transient destabilization in specific binding modes. (b) RMSF profiles measure per-residue flexibility and highlight specific regions of conformational plasticity. The protein only shows pronounced fluctuations, especially near residue indices ∼370 and ∼480, corresponding to loop regions. Ligand binding attenuates these fluctuations to varying degrees, with F5463-0122 exerting a notable dampening effect on flexible segments, consistent with stabilizing ligand–residue interactions. This observation supports the hypothesis that ligand occupancy selectively modulates local flexibility and dynamic communication across the binding pocket and distal residues. Together, the RMSD and RMSF analyses demonstrate that ligand binding can fine-tune protein conformational dynamics, with implications for stability, function, and druggability.
Root-mean-square fluctuation (RMSF) describes the variation in individual amino acids of a protein that occurs due to ligand binding in a dynamic environment. The AGPS protein structure, used as a template for homology modeling, begins at position 81, meaning that the modeled protein has been constructed by starting from this point. RMSF of the unbound protein varies between residues 156 and 165 (PPSIVNEDFL), which includes a loop and a small portion of an α-helix, residues 351–364 (GPRMSTGPDIHHFI); proper orientation is necessary for the catalytic activity of the enzyme, and residues 431–460 (ENNLTAHVEAGITG) correspond to a disordered segment, indicating that it is unstable and highly mobile. The RMSF value of all complexes at residues 156–165 shows a reduction in fluctuations; however, the other two fluctuation regions remain unchanged (Figure b). For the protein–F3309-0655 complex, there is a significant decrease in the RMSF value at residues 156–165 to 1.5 Å, while the other fluctuations observed in the unbound protein persist. The RMSF of the protein–F1885-1052 complex shows fluctuations in the amino acid sequences from 351–364 and from 431–460, but it reduces the RMSF value for sequences 156 and 165 to 3 Å. The protein–F2881-0267 complex exhibits a high RMSF value of 4 Å for sequences 156 and 165 and fluctuations at 351–364 but reduces fluctuations in the 431–460 regions to 3.2 Å. In contrast, the protein–F6472-0974 complex shows the greatest reduction in RMSF, with minimal fluctuations. Protein–Zinc69435460 shows fluctuations at 351–364 and 431–460 up to 3.5 and 4 Å, respectively. The RMSF values of protein complexes with F2881-0267, F1885-1052, F3309-0655, and F5463-0122 show fluctuation in the region of 351–364, which suggests that they prevent the catalytic activity of the AGPS enzyme.
The binding free energy and affinity of protein–ligand complexes, estimated by the MMGBSA method, are influenced by several contributing factors. These include van der Waals interactions, electrostatic forces, and polar solvation energies, as elaborated in Table S5. Notably, the post-MD binding energies of the complexes have increased in comparison to the values obtained after docking, suggesting that all five complexes exhibit enhanced binding affinities and stability under dynamic simulation conditions. The compound exhibiting the most negative ΔG reflects the strongest binding energy. Specifically, F1885-1052 and F2881-0267 show significant binding energies of −136.93 kcal/mol and −129.78 kcal/mol, respectively, indicating consistent interactions in solution. These values indicate improved stability under dynamic conditions, outperforming the standard inhibitor ZINC69435460 (−119.24 kcal/mol).
3.8.1. Protein–Ligand Interaction and Conformational Dynamics
The interaction profile from the MD simulations highlights hydrogen bonds as the most frequent and persistent noncovalent interactions, crucial for ligand stabilization and pharmacophore specificity. Hydrophobic contacts, though weaker, aid in enhancing ligand orientation and minimizing solvent exposure. Among the compounds, F5463-0122 demonstrated sustained hydrogen bonding with P234, G237, and G238 and the catalytically important H616, along with persistent hydrophobic interactions involving H617. F3309-0655 formed hydrogen bonds with Q425 and transient hydrophobic contacts with F428, R566, and Y580. F1885-1052 maintained a stable hydrogen bonding with Q425 and weak hydrophobic interactions with Y580 throughout the simulation. Compared to other ligands, F2881-0267 exhibited the most extensive hydrogen-bonding network, interacting with multiple residues, including P234, G236–238, T239, S240, T260, D303, T309, S315, S319, E368, and I374. In addition, it formed intermittent but functionally significant hydrogen bonds and water bridges with active-site residues H616 and H617, contributing to its favorable binding profile. Notably, F2881-0267 also exhibited the highest number of water-mediated interactions, particularly involving catalytically relevant residues (Figure S2). F6472-0974 also demonstrated persistent hydrogen bonding with H616, H617, G236, G238, T239, and T316, with longer interaction durations compared to those of F2881-0267. Bound water molecules on protein surfaces enhance ligand–protein interaction stability by facilitating bridging interactions, boosting enzyme–ligand affinity, and influencing inhibitor binding.
Rg measures the compactness of protein–ligand complexes, with lower values indicating stability and higher values suggesting expansion. , The unbound protein exhibited an average R g of 24.18 Å, Figure S3a. Among the ligand-bound complexes, F5463-0122 exhibited the highest degree of structural compactness, with an R g value of 24.24 Å, followed by F2881-0267 at 24.42 Å. F3309-0655 and Zinc69435460 maintained moderate levels of compactness, with R g values of 24.56 Å and 24.52 Å, respectively. In contrast, F6472-0974 and F1885-1052 displayed substantially elevated R g values of 26.32 Å and 26.63 Å, respectively, indicating reduced compactness. The relatively lower R g values associated with F5463-0122, F2881-0267, F3309-0655, and Zinc69435460 suggest their ability to maintain structural integrity and potentially impede tumor progression.
SASA measures the solvent-exposed surface area of biomolecules, reflecting the conformational dynamics and interaction potential. Higher SASA indicates greater surface exposure and ligand accessibility, while lower SASA suggests compactness and reduced flexibility. In our analysis, the unbound protein presented an average SASA of 23,049.26 Å2 (Table S6). In the ligand-bound complexes, F5463-0122, F2881-0267, and F3309-0655 exhibited comparatively lower average SASA values of 23,408.91 Å2, 23,688.39 Å2, and 23,940.28 Å2, respectively, suggesting that these ligands promote tightening of the protein conformation and decrease solvent exposure. However, ligands F1885-1052 and F6472-0974 showed higher average SASA values of 24,166.85 Å2 and 24,183.21 Å2, respectively, indicating a greater solvent-exposed conformation that may affect binding kinetics and interaction dynamics (Figure S3b). Additionally, the standard inhibitor, Zinc69435460, displayed a relatively low SASA of 23,464.90 Å2, aligning with a stable and compact binding conformation.
3.9. Protein Dynamics, PCA, and Free Energy Landscape Analysis
The DCCM analysis served as an essential analytical framework for probing the conformational dynamics of proteins by quantitatively assessing the correlated atomic motions observed in MD simulations. , In the DCCM visualizations, positively correlated movements are represented in blue, whereas negatively correlated motions are illustrated in dark red. Analysis of the ligand-bound complexes F2881-0267, F1885-1052, F5463-0122, and F3309-0655 reveals significant positive correlations in the active site reflecting the overall binding stability of the complexes. F6472-0974 shows weaker active-site correlations but notable anticorrelated dynamics in the other regions, suggesting moderate binding affinity (Figure S4). Percentage motion (% motion) reflects protein atomic movement, with lower values indicating greater stability. F1885-1052 exhibited the least motion (26.97%), followed by F6472-0974 (27.51%), F2881-0267 (30.69%), F5463-0122 (33.23%), and F3309-0655 (40.05%). The unbound protein showed substantially higher motion (48.27%), whereas the standard inhibitor Zinc69435460 exhibited the lowest % motion (23.10%), emphasizing its high stability in the binding environment.
Protein flexibility was assessed using PCA of the first two principal components (PC1 and PC2), with conformational distributions visualized as color-coded scatter plots. The clustering of ligand-bound systems together with the protein-only simulation within the central region of the PCA space indicates limited conformational variability and the predominance of stable structural states with only minor excursions toward peripheral regions. Notably, the F2881-0267-protein and F3309-0655-protein complexes exhibit PCA distributions more compact than those of the other ligand–protein systems, suggesting reduced conformational deviations and restricted conformational mobility. Collectively, these results indicate ligand-dependent modulation of protein dynamics, as captured by PCA through deviations from the mean conformational state following ligand binding (Figure ).
8.
PCA outlines distinct conformational states across ligand-bound ensembles. PCA was performed on the Cα atoms derived from MD trajectories to identify the predominant modes of motion. The two-dimensional projection onto the principal components PC1 and PC2 illustrates the conformational landscapes explored by both protein only (yellow) and ligand-bound systems (color-coded as per the accompanying legend). Protein only exhibits a broadly dispersed distribution along the principal component axes, indicative of an extensive sampling of conformational space. In contrast, ligand-bound complexes form more compact and distinct clusters, with each ligand (F5463-0122, F2881-0267, F3309-0655, F6472-0974, F1885-1052, Zinc69435460) inducing unique shifts within the conformational ensemble. Notably, the F2881-0267–protein complex occupies a more restricted region of the PCA space, reflecting reduced conformational fluctuations relative to the other ligand-bound systems.
The FEL provides detailed insight into conformational shifts, protein folding dynamics, and overall stability profiles. By assessing unbound proteins in contrast to their ligand-bound complexes, FEL analysis facilitates an evaluation of thermodynamic variances, with lower free energy states indicating higher conformational stability. In the case of the F2881-0267–protein complex, the FEL describes a significantly stabilized conformational ensemble, marked by a dense low-free-energy basin (depicted in violet) situated near the origin of the PC1–PC2 plane. This suggests a considerable enhancement in stability compared to the unbound protein (Figure ). The observation of multiple accessible low-energy minima spread over an extensive region implies that the complex can explore various energetically favorable conformations with minimal energetic penalties. Moreover, both the ZINC69435460 protein complex and the protein-only system exhibit clearly defined low-free-energy basins, supporting the idea of stable conformational states within these systems.
9.
FEL captures the ligand-driven conformational shifts and energy barriers of protein dynamics. The FELs mapped along the first two principal components (PC1 and PC2) illustrate the conformational heterogeneity and energetic preferences of protein only and ligand-bound protein complexes. In each panel (a) F2881-0267, (b) F1885-1052, (c) F3309-0655, (d) F5463-0122, (e) F6472-0974, (f) protein only, (g) Zinc69435460, distinct energy basins can be observed, corresponding to metastable conformational states sampled during MD simulations. Protein only displays a broader, more rugged energy surface, indicating a higher degree of conformational flexibility. In contrast, the ligand-bound systems exhibit more localized and deeper minima, reflecting ligand-induced stabilization of specific conformational states. Notably, ligands such as F2881-0267 and Zinc69435460 demonstrate particularly sharp and well-defined minima, suggesting strong confinement within energetically favorable conformational substates. These FELs underscore the role of ligand binding in remodeling the protein’s energy landscape, effectively narrowing its conformational ensemble and raising energy barriers between transitions.
3.10. Deep Learning Hybrid Model for Binding Free Energy Prediction of AGPS-F2881_0267
Our results demonstrate the effectiveness of the custom hybrid deep learning framework for predicting the binding energy of protein–ligand complexes. The implementation of parallel convolutional and fully connected layers effectively captured 3D vowelized features, interaction fingerprints (ProLIF), and electrostatic and van der Waals descriptors.
3.11. Capturing AGPS-F2881_0267 Atomics via MD Simulation
The effect of every atom, specifically its force-field dependence with a statistical precision, is necessary for absolute binding free energy calculation. , The resulting spatial coordinates and velocity of atoms in nanometers per picosecond during the interaction simulation from the trajectory file are utilized for framing the energy level variations in the atomic context. In order to confine all of the atoms in the MD-generated synthetic environment, we stripped the atoms for a perfect topology matching. While inputting the trajectories, the spatial challenge is to correctly identify the same atoms across the AGPS-F2881_0267 pdb format, so the learning method can cause consistent information for comparing two molecular environment conformations in its main function.
3.12. Deep Interaction Featurization Using ProLIF and RDKit
To accurately characterize binding free energies across molecular dynamic trajectories, we implemented a modular featurization pipeline that proved to be effective in capturing the complexity of dynamic protein–ligand interactions. Our framework successfully integrated topology-based slicing using MDTraj, interaction fingerprinting via ProLIF, and custom voxelization strategies centered on ligand binding sites. This combination yielded a computationally efficient yet detailed representation of transient binding events.
Using RDKit, we maintained compatibility with diverse ligand chemistries while preserving atomic-level detail. Protein structures were parsed and annotated at high resolution using Bio.PDB, allowing for accurate structural and chemical characterization. The voxel-based features embedded in dense 4D tensors captured spatial, chemical, and energetic information, making them well-suited for 3D CNN processing. Simultaneously, the ProLIF-derived interaction fingerprints effectively represented interpretable biochemical contacts such as hydrogen bonding and π–π-stacking.
In addition to these learned descriptors, we incorporated Lennard–Jones and electrostatic terms, producing a hybrid representation that balances physical accuracy with machine learning adaptability. This architecture enabled high-fidelity prediction of binding free energies, offering a compelling integration of interpretability, physical realism, and computational efficiency (Figure a).
10.
Hybrid deep learning model for binding free energy estimation- (a) The hybrid neural network architecture integrates 3 complementary input modalities to estimate binding free energy: A 3D voxel branch that encodes spatial interaction patterns via 3D convolutional layers, a molecular fingerprint branch that captures topological and physicochemical features through dense layers, and a physics-informed branch that processes handcrafted features derived from interaction energy terms. These parallel branches are concatenated and fed into a fully connected regressor module, which outputs predicted binding affinity (ΔG\Delta GΔG). (b) Frame-wise prediction of binding free energy across a 1000-frame MD trajectory is shown. Each point corresponds to an individual frame’s predicted ΔG\Delta GΔG, with high-frequency fluctuation highlighting conformational variability. The red dashed line represents the Boltzmann-weighted average binding free energy (−4.88 kcal/mol), with the shaded area delineating a confidence interval of ±1 standard deviation (1.08 kcal/mol) around the predicted ensemble average. The model’s capacity to generalize across diverse representations underscores its significant promise in improving drug discovery frameworks.
3.13. Model Evaluation
To assess the efficacy of our hybrid deep learning model for predicting protein–ligand binding free energies, we trained and evaluated it on MD trajectories processed through our custom pipeline. The training was performed over ten epochs using the Hybrid Binding Model, which integrates three distinct feature types: 3D voxelized representations of the binding pocket, ligand–protein interaction fingerprints, and van der Waals, electrostatics, and solvation. A Binding Data set was constructed from MD snapshots, enabling framewise supervision of the model.
The model converged stably, showing a consistent loss reduction across training epochs (e.g., final average MSE loss below 0.01), indicating its ability to learn from the joint feature representation. After training, we applied the model to a held-out trajectory processed via the MD Trajectory Processor, which extracts features per frame, predicts per-frame ΔG values, and calculates ensemble-averaged predictions.
3.14. Binding Free Energy and Dissociation Constant
We extracted 985 frames from a 200 ns molecular dynamics trajectory of the protein–ligand complex, stripping solvent, and ions to retain only relevant binding atoms. Each frame was featurized and evaluated using our pretrained model. The predicted per-frame binding free energies were aggregated using a Boltzmann-weighted averaging scheme, accounting for the statistical mechanics of binding ensemble populations. The resulting Boltzmann-weighted binding free energy (ΔG) was −4.88 kcal/mol, with a standard deviation of 1.08 and the dissociation constant is 1.25 mM where the gas constant K 1.987 × 10–3 kcal/mol and T = 298 K (Figure b, See Methods).
Rather than chasing artificially low binding free energies, our model targets experimentally realistic affinities. This provides a more interpretable and translatable framework for early stage drug design, especially for fragment-based screening and shallow protein pockets. We validate this using deep learning-assisted Boltzmann-weighted sampling, producing ΔG values consistent with what can be measured by ITC or SPR.
4. Discussion
AI-assisted in-silico drug discovery pipelines now enable the simultaneous disruption of multiple signaling pathways, offering promising solutions to therapeutic resistance especially for cancer. In glioma, the dual inhibition of AKT and mTOR has emerged as a more effective approach. The enzyme AGPS has been implicated in glioma progression via the PI3K/Akt and mTOR pathways as well as through regulation by lncRNAs and microRNAs. Given the absence of a crystallized AGPS structure, a homology model was built using SWISS-MODEL and virtually screened against 576,510 compounds from the Life Chemicals library using an attention-based deep learning model. Five top candidates demonstrated favorable ADMET profiles and BBB permeability, with polar interactions involving critical residues GLN 425 and TYR580. Further IFD and RMSD analysis confirmed the structural stability of several complexes, with Rg data supporting protein folding, consistent with AGPS inhibition.
Among the top hits, F2881-0267 exhibited the most promising profile, showing a strong binding energy (below −100 kcal/mol) and disrupting AGPS catalytic activity without engaging 5-deazaFAD, a cofactor essential for enzymatic function. Comparative analysis with ZINC69435460, the only known AGPS inhibitor, emphasized the advantage of using the curated Life Chemicals HTS library for targeted drug discovery. Molecular dynamics simulations validated by SASA, DCCM, PCA, and FEL further supported the stability and efficacy of the identified compounds. These findings position F2881-0267 as a leading AGPS inhibitor candidate. Nonetheless, interpopulation variability in therapeutic response remains a limitation. Future work should incorporate extended MD simulations, deep learning-based structural refinement, and experimental validation through ITC or SPR to confirm binding affinities and optimize therapeutic potential.
5. Conclusion
This AI based framework can be adopted for alleviating cancer progression. Evidencing that, this optimized AI offers scalable and time efficient solutions for generating target specific therapeutic candidates with potent anticancer activity. There is a pressing need to identify molecular targets to support cancer therapeutic treatment outcomes. In this study, we focused on AGPS, a protein that is relatively understudied in glioma progression. By integrating AI-based drug repurposing with molecular modeling, we identified promising HTS-derived compounds with strong potential for therapeutic intervention. Our approach underscores the utility of AI in accelerating target identification and lead optimization for complex cancers like glioma.
Supplementary Material
Acknowledgments
We acknowledge Yenepoya (Deemed to be University), Mangalore, for providing infrastructure and funding for the Centre for Integrative Omics Data Science (CIODS).
Data available through this GitHub Link: https://github.com/naveen-joy-18/Feature-driven-AI-Guided-Drug-Repurposing-Pipeline-for-AGPS-Targeted-Glioma-Therapy-.git
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.5c09368.
Binding affinity predictions from the DNN model (Table S1) (XLSX)
Quantitative docking and comprehensive binding affinity assessment and interaction mapping of glioma-targeted ligands (Table S2); Physicochemical and toxicological properties precited via ADMET modeling for lead optimization (Table S3); Predicted binding energetics and conformational adaptation assessment from induced fit docking coupled with MM-GBSA analysis (Table S4); Postdynamic MM-GBSA binding free energy evaluation of ligand–target interactions (Table S5); Average structural and solvent exposure metrics of ligand–protein complexes from MD simulations (Table S6); ROC curve evaluating the predictive performance of AGPS inhibitor screening mode (Figure S1); Profiling hydrogen bond interactions among ligand-bound states (Figure S2); Structural compactness and solvent exposure reveal ligand-induced modulation of AGPS dynamics (Figure S3); DCCM analysis uncovers ligand-driven modulation of residue dynamics (Figure S4) (PDF)
AT: writingreview and editing, writingoriginal draft, methodology, investigation, formal analysis, conceptualization. SDT: methodology, formal analysis, writingreview and editing. LD: writingreview and editing, writingoriginal draft, methodology, formal analysis, conceptualization. LJ: writingreview and editing, methodology, conceptualization. JV: methodology, writingreview and editing. NJ: methodology, writingreview and editing, RRD: writingreview and editing, RR: writingreview and editing, conceptualization. AJ: writingreview and editing, writingoriginal draft, methodology, conceptualization.
The authors declare that the article publishing charges, computational resources, and software infrastructure were provided by Yenepoya (Deemed to be University). Dr. Abhithaj J reports a relationship with Yenepoya that includes employment and nonfinancial support.
The authors declare no competing financial interest.
References
- You Y., Lai X., Pan Y., Zheng H., Vera J., Liu S., Deng S., Zhang L.. Artificial intelligence in cancer target identification and drug discovery. Signal Transduct Target Ther. 2022;7(1):156. doi: 10.1038/s41392-022-00994-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharma V., Singh A., Chauhan S., Sharma P. K., Chaudhary S., Sharma A., Porwal O., Fuloria N. K.. Role of Artificial Intelligence in Drug Discovery and Target Identification in Cancer. Curr. Drug Deliv. 2024;21(6):870–886. doi: 10.2174/1567201821666230905090621. [DOI] [PubMed] [Google Scholar]
- Pandiyan S., Wang L.. A comprehensive review on recent approaches for cancer drug discovery associated with artificial intelligence. Comput. Biol. Med. 2022;150:106140. doi: 10.1016/j.compbiomed.2022.106140. [DOI] [PubMed] [Google Scholar]
- Issa N. T., Stathias V., Schurer S., Dakshanamurthy S.. Machine and deep learning approaches for cancer drug repurposing. Semin. Cancer Biol. 2021;68:132–142. doi: 10.1016/j.semcancer.2019.12.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jones D., Kim H., Zhang X., Zemla A., Stevenson G., Bennett W. F. D., Kirshner D., Wong S. E., Lightstone F. C., Allen J. E.. Improved Protein-Ligand Binding Affinity Prediction with Structure-Based Deep Fusion Inference. J. Chem. Inf. Model. 2021;61(4):1583–1592. doi: 10.1021/acs.jcim.0c01306. [DOI] [PubMed] [Google Scholar]
- Reiser P., Neubert M., Eberhard A., Torresi L., Zhou C., Shao C., Metni H., van Hoesel C., Schopmans H., Sommer T.. et al. Graph neural networks for materials science and chemistry. Commun. Mater. 2022;3(1):93. doi: 10.1038/s43246-022-00315-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ozturk H., Ozgur A., Ozkirimli E.. DeepDTA: deep drug-target binding affinity prediction. Bioinformatics. 2018;34(17):i821–i829. doi: 10.1093/bioinformatics/bty593. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ren F., Aliper A., Chen J., Zhao H., Rao S., Kuppe C., Ozerov I. V., Zhang M., Witte K., Kruse C.. et al. A small-molecule TNIK inhibitor targets fibrosis in preclinical and clinical models. Nat. Biotechnol. 2025;43(1):63–75. doi: 10.1038/s41587-024-02143-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thaikkad A., Henna F., Thomas S. D., John L., Raju R., Jayanandan A.. Cangrelor and AVN-944 as repurposable candidate drugs for hMPV: analysis entailed by AI-driven in silico approach. Mol. Divers. 2025;29:3587. doi: 10.1007/s11030-025-11206-6. [DOI] [PubMed] [Google Scholar]
- Sadybekov A. V., Katritch V.. Computational approaches streamlining drug discovery. Nature. 2023;616(7958):673–685. doi: 10.1038/s41586-023-05905-z. [DOI] [PubMed] [Google Scholar]
- King A.. Four ways to power-up AI for drug discovery. Nature. 2025:00602-5. doi: 10.1038/d41586-025-00602-5. [DOI] [PubMed] [Google Scholar]
- Yasinjan F., Xing Y., Geng H., Guo R., Yang L., Liu Z., Wang H.. Immunotherapy: a promising approach for glioma treatment. Front. Immunol. 2023;14:1255611. doi: 10.3389/fimmu.2023.1255611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cahill D. P., Sloan A. E., Nahed B. V., Aldape K. D., Louis D. N., Ryken T. C., Kalkanis S. N., Olson J. J.. The role of neuropathology in the management of patients with diffuse low grade glioma: A systematic review and evidence-based clinical practice guideline. J. Neurooncol. 2015;125(3):531–549. doi: 10.1007/s11060-015-1909-8. [DOI] [PubMed] [Google Scholar]
- Xu S., Tang L., Li X., Fan F., Liu Z.. Immunotherapy for glioma: Current management and future application. Cancer Lett. 2020;476:1–12. doi: 10.1016/j.canlet.2020.02.002. [DOI] [PubMed] [Google Scholar]
- Wang L. M., Englander Z. K., Miller M. L., Bruce J. N.. Malignant Glioma. Adv. Exp. Med. Biol. 2023;1405:1–30. doi: 10.1007/978-3-031-23705-8_1. [DOI] [PubMed] [Google Scholar]
- Noor H., Briggs N. E., McDonald K. L., Holst J., Vittorio O.. TP53 Mutation Is a Prognostic Factor in Lower Grade Glioma and May Influence Chemotherapy Efficacy. Cancers. 2021;13(21):5362. doi: 10.3390/cancers13215362. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giotta Lucifero A., Luzzi S.. Immune Landscape in PTEN-Related Glioma Microenvironment: A Bioinformatic Analysis. Brain Sci. 2022;12(4):501. doi: 10.3390/brainsci12040501. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang J. M., Schiapparelli P., Nguyen H. N., Igarashi A., Zhang Q., Abbadi S., Amzel L. M., Sesaki H., Quinones-Hinojosa A., Iijima M.. Characterization of PTEN mutations in brain cancer reveals that pten mono-ubiquitination promotes protein stability and nuclear localization. Oncogene. 2017;36(26):3673–3685. doi: 10.1038/onc.2016.493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou W., Liu Y., Li H., Song Z., Ma Y., Zhu Y.. Mass Spectrometry and Computer Simulation Predict the Interactions of AGPS and HNRNPK in Glioma. BioMed Res. Int. 2021;2021:6181936. doi: 10.1155/2021/6181936. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benjamin D. I., Cozzo A., Ji X., Roberts L. S., Louie S. M., Mulvihill M. M., Luo K., Nomura D. K.. Ether lipid generating enzyme AGPS alters the balance of structural and signaling lipids to fuel cancer pathogenicity. Proc. Natl. Acad. Sci. U. S. A. 2013;110(37):14912–14917. doi: 10.1073/pnas.1310894110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen L., Zhang W., He L., Jin L., Qian L., Zhu Y.. Effect of alkylglycerone phosphate synthase on the expression levels of lncRNAs in glioma cells and its functional prediction. Oncol. Lett. 2020;20(4):66. doi: 10.3892/ol.2020.11927. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Doncheva N. T., Morris J. H., Gorodkin J., Jensen L. J.. Cytoscape StringApp: Network Analysis and Visualization of Proteomics Data. J. Proteome Res. 2019;18(2):623–632. doi: 10.1021/acs.jproteome.8b00702. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Piano V., Benjamin D. I., Valente S., Nenci S., Marrocco B., Mai A., Aliverti A., Nomura D. K., Mattevi A.. Discovery of Inhibitors for the Ether Lipid-Generating Enzyme AGPS as Anti-Cancer Agents. ACS Chem. Biol. 2015;10(11):2589–2597. doi: 10.1021/acschembio.5b00466. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nenci S., Piano V., Rosati S., Aliverti A., Pandini V., Fraaije M. W., Heck A. J., Edmondson D. E., Mattevi A.. Precursor of ether phospholipids is synthesized by a flavoenzyme through covalent catalysis. Proc. Natl. Acad. Sci. U. S. A. 2012;109(46):18791–18796. doi: 10.1073/pnas.1215128109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Manthattil Vysyan S., Suraj Prasanna M., Jayanandan A., Gangadharan A. K., Chittalakkottu S.. Phytocompounds hesperidin, rebaudioside a and rutin as drug leads for the treatment of tuberculosis targeting mycobacterial phosphoribosyl pyrophosphate synthetase. J. Biomol. Struct. Dyn. 2024:1–15. doi: 10.1080/07391102.2024.2438363. [DOI] [PubMed] [Google Scholar]
- Yang Y., Yao K., Repasky M. P., Leswing K., Abel R., Shoichet B. K., Jerome S. V.. Efficient Exploration of Chemical Space with Docking and Deep Learning. J. Chem. Theory Comput. 2021;17(11):7106–7119. doi: 10.1021/acs.jctc.1c00810. [DOI] [PubMed] [Google Scholar]
- Jacobson M. P., Pincus D. L., Rapp C. S., Day T. J., Honig B., Shaw D. E., Friesner R. A.. A hierarchical approach to all-atom protein loop prediction. Proteins. 2004;55(2):351–367. doi: 10.1002/prot.10613. [DOI] [PubMed] [Google Scholar]
- John L., Dcunha L., Ahmed M., Thomas S. D., Raju R., Jayanandan A.. A deep learning and molecular modeling approach to repurposing Cangrelor as a potential inhibitor of Nipah virus. Sci. Rep. 2025;15(1):16440. doi: 10.1038/s41598-025-00024-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tian H., Ketkar R., Tao P.. ADMETboost: a web server for accurate ADMET prediction. J. Mol. Model. 2022;28(12):408. doi: 10.1007/s00894-022-05373-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henna F., Arun Kumar G., Thaikkad A., Varun T., Jayadevi Variyar E., Raju R., Abhithaj J.. In silico and In vitro Profiling of Lariciresinol Against PLA2: A molecular Approach to Regulate Inflammation. Eur. J. Med. Chem. Rep. 2025;15:100290. doi: 10.1016/j.ejmcr.2025.100290. [DOI] [Google Scholar]
- Nabuurs S. B., Wagener M., de Vlieg J.. A flexible approach to induced fit docking. J. Med. Chem. 2007;50(26):6507–6518. doi: 10.1021/jm070593p. [DOI] [PubMed] [Google Scholar]
- Bowers, K. J. ; Chow, E. ; Xu, H. ; Dror, R. O. ; Eastwood, M. P. ; Gregersen, B. A. ; Klepeis, J. L. ; Kolossvary, I. ; Moraes, M. A. ; Sacerdoti, F. D. . Scalable algorithms for molecular dynamics simulations on commodity clusters, Proceedings of the 2006 ACM/IEEE Conference on Supercomputing; IEEE, 2006. [Google Scholar]
- Gangadharan, A. K. ; Kundil, V. T. ; Jayanandan, A. . Computational Tools in Drug-Lead Identification and Development. In Drugs from Nature: Targets, Assay Systems and Leads; Springer, 2024; pp 89–119. [Google Scholar]
- Dubchak I., Holbrook S. R., Kim S. H.. Prediction of protein folding class from amino acid composition. Proteins. 1993;16(1):79–91. doi: 10.1002/prot.340160109. [DOI] [PubMed] [Google Scholar]
- Fatriansyah J. F., Boanerges A. G., Kurnianto S. R., Pradana A. F., Fadilah, Surip S. N.. Molecular Dynamics Simulation of Ligands from Anredera cordifolia (Binahong) to the Main Protease (M (pro)) of SARS-CoV-2. J. Trop Med. 2022;2022:1–13. doi: 10.1155/2022/1178228. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hong S., Kim D.. Interaction between bound water molecules and local protein structures: A statistical analysis of the hydrogen bond structures around bound water molecules. Proteins. 2016;84(1):43–51. doi: 10.1002/prot.24953. [DOI] [PubMed] [Google Scholar]
- Siddiquee N. H., Tanni A. A., Sarker N., Sourav S. H., Islam L., Mili M. A., Akter F., Chandra Roy S., Abdullah-Al-Mamun M., Malek S.. et al. Insights into novel inhibitors intending HCMV protease a computational molecular modelling investigation for antiviral drug repurposing. Inform. Med. Unlocked. 2024;48:101522. doi: 10.1016/j.imu.2024.101522. [DOI] [Google Scholar]
- Khan A., Ali S. S., Khan M. T., Saleem S., Ali A., Suleman M., Babar Z., Shafiq A., Khan M., Wei D. Q.. Combined drug repurposing and virtual screening strategies with molecular dynamics simulation identified potent inhibitors for SARS-CoV-2 main protease (3CLpro) J. Biomol. Struct. Dyn. 2021;39(13):4659–4670. doi: 10.1080/07391102.2020.1779128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parida P. K., Paul D., Chakravorty D.. The natural way forward: Molecular dynamics simulation analysis of phytochemicals from Indian medicinal plants as potential inhibitors of SARS-CoV-2 targets. Phytother. Res. 2020;34(12):3420–3433. doi: 10.1002/ptr.6868. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maisuradze G. G., Liwo A., Scheraga H. A.. Relation between free energy landscapes of proteins and dynamics. J. Chem. Theory Comput. 2010;6(2):583–595. doi: 10.1021/ct9005745. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fu H., Chen H., Blazhynska M., Goulard Coderc de Lacam E., Szczepaniak F., Pavlova A., Shao X., Gumbart J. C., Dehez F., Roux B.. et al. Accurate determination of protein:ligand standard binding free energies from molecular dynamics simulations. Nat. Protoc. 2022;17(4):1114–1141. doi: 10.1038/s41596-021-00676-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Comer J., Gumbart J. C., Henin J., Lelievre T., Pohorille A., Chipot C.. The adaptive biasing force method: everything you always wanted to know but were afraid to ask. J. Phys. Chem. B. 2015;119(3):1129–1151. doi: 10.1021/jp506633n. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chmiela S., Sauceda H. E., Muller K. R., Tkatchenko A.. Towards exact molecular dynamics simulations with machine-learned force fields. Nat. Commun. 2018;9(1):3887. doi: 10.1038/s41467-018-06169-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McGibbon R. T., Beauchamp K. A., Harrigan M. P., Klein C., Swails J. M., Hernandez C. X., Schwantes C. R., Wang L. P., Lane T. J., Pande V. S.. MDTraj: A Modern Open Library for the Analysis of Molecular Dynamics Trajectories. Biophys. J. 2015;109(8):1528–1532. doi: 10.1016/j.bpj.2015.08.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouysset C., Fiorucci S.. ProLIF: a library to encode molecular interactions as fingerprints. J. Cheminf. 2021;13(1):72. doi: 10.1186/s13321-021-00548-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Scalfani V. F., Patel V. D., Fernandez A. M.. Visualizing chemical space networks with RDKit and NetworkX. J. Cheminf. 2022;14(1):87. doi: 10.1186/s13321-022-00664-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data available through this GitHub Link: https://github.com/naveen-joy-18/Feature-driven-AI-Guided-Drug-Repurposing-Pipeline-for-AGPS-Targeted-Glioma-Therapy-.git










