Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jan 30.
Published in final edited form as: Annu Rev Med. 2026 Jan;77(1):381–398. doi: 10.1146/annurev-med-050224-122802

Artificial Intelligence to Guide Repurposing of Drugs

Zhimin Fu 1, Yuxin Yang 2,3, Mina Chung 4,5, Serpil Erzurum 6, Feixiong Cheng 2,3,4,*
PMCID: PMC12854511  NIHMSID: NIHMS2112616  PMID: 41592930

Abstract

With the pharmacokinetics, dosing, safety, and manufacturing of approved or investigational drugs already well-characterized, drug repurposing/repositioning offers an emerging strategy to rapidly develop effective treatments for various challenging diseases. However, the growing mass of genetic and multi-omics data has not been effectively explored by the drug repurposing community due to a lack of accurate approaches. This Review aims to be an authoritative, critical, and accessible review and discussion of general interest to the drug repurposing community concerning the use of Artificial Intelligence (AI) and Machine Learning (ML) tools. Emerging questions include what is achievable with AI in this domain and what will the impact be, what does AI/ML embrace, and how can we, as geneticists, pharmacologists, and computational scientists, contribute to the discovery of new, cheap, and affordable repurposable medicines. The fast growth of genetics and multi-omics data (Genomics, Transcriptomics, Proteomics, Metabolomics, and Radiomics) and Electronic Health Records (EHR) in diverse populations contribute to answering questions, such as how to rapidly identify effective repurposable medicines, what is a clinically meaningful effect size in trials, and what are the potential implications for precision medicine. This review will discuss AI/ML for drug repurposing in the context of genetics, multi-omics, real-world data collection, and crowdsourcing of knowledge. We finally conclude by considering questions on how AI/ML methodologies can unite the diverse aspects of translational medicine for emerging treatment development in human challenging diseases.

1. Introduction

Although investment in biomedical and pharmaceutical research has increased significantly over the past two decades, specifically in understanding disease genetic risk factors, we still need to develop more effective disease-modifying treatments for multiple challenging diseases, such as Alzheimer’s disease (AD), heart disease, cancer, and Coronavirus disease 2019 (COVID-19)13. This is partly because the growing mass of genetic and multi-omics datasets has not been effectively explored for drug development due to a lack of accurate approaches. Drug repurposing/repositioning reduces the time and cost of drug development14. With the pharmacokinetics, dosing, safety, and manufacturing of approved or investigational drugs already well-characterized, the goal of drug repositioning is to identify new indications for drugs46, e.g., an anti-inflammatory agent for arthritis might be repositioned for treatment of AD or AD-related dementia (ADRD). However, how to prioritize drug targets and candidate repurposable medicines for human complex diseases at drugome-wide and genome-wide scales is challenging.

The recent increase in generation of “multi-omics” data, including genetics, genomics, epigenomics, transcriptomics, proteomics, metabolomics, lipidomics, radiomics, and phenomics, as well as data digitalization in patient care and the pharmaceutical sector, present both challenges and opportunities (Figure 1). For example, the growing availability of big data prompts challenges for personalized clinical diagnosis and treatment of human disease. These challenges have motivated the application of advanced Artificial Intelligence (AI) and Machine Learning (ML) tools to help scientists identify repurposable drugs and accelerate progress toward effective treatments for a variety of challenging complex diseases.

Figure 1.

Figure 1.

A proposed novel AI agent-based drug discovery pipeline for personalized medicine. a, An overview pipeline in AI-based drug repurposing. b, Constructing AI-ready datasets. These databases consist of three pivotal datasets, biomedical knowledge graph (KG) dataset, multi-omics dataset, and patient/clinical dataset. The biomedical KG dataset covers relations between drugs, diseases, genes, pathways, and tissues. The multi-omics dataset includes data about genome, transcriptome, phenome, interactome, proteome, and drug, while the patient dataset contains patient health record and genome data of patient. c, AI-assisted drug approaches for both disease-centric and target-centric drug repurposing. d, Experimental validation of AI-prioritized repurposable drugs. The AI-prioritized candidate drugs can be validated using in vitro and in vivo experiments to ensure the efficacy and safety. Subsequently, the candidate drugs with ideal efficacy and safety profiles will be moved to clinical trial on patients with similar diseases, symptoms, and genetic profiles under the precision medicine hypothesis.

2. Target-centric and disease-centric drug repurposing

There are two major ways to do drug repurposing: target-centric drug repurposing and disease-centric drug repurposing. In a recent analysis of repurposed drugs, the majority of repurposed drugs are developed from disease-centric repurposing7. The premise of disease-centric drug repurposing is that the same drug can be applied to treat different diseases when these diseases share similar biological pathways, symptoms, or traits7. A key step in conducting disease-centric drug repurposing is to identify underlying biological mechanisms of action for the target disease which are homologous to the original disease that a drug treats8. Sildenafil is an oral medication to treat pulmonary arterial hypertension and erectile dysfunction9,10. Sildenafil reduces breakdown of cyclic guanosine monophosphate (cGMP) via inhibiting cGMP specific phosphodiesterase type 5 (PDE5)11. Recent works suggest sildenafil has potential in AD treatment12. By studying AD endophenotype disease modules within protein-protein interaction (PPI) networks, sildenafil stands out as a novel candidate drug for AD, along with reduced incidence of AD in real-world patient data13. In vivo investigations imply sildenafil helps brining back cognitive function of AD mouse by reducing activity of glycogen synthase kinase 3β (GSK3β) and cyclin-dependent kinase 5 (CDK5), as well as bringing up the level of brain-derived neurotrophic factor (BDNF)14. Studies using AD patient induced pluripotent stem cells (iPSC) also indicate that sildenafil targets AD related genes and pathways15.

On the other hand, target-centric drug repurposing assumes that the same target protein is associated with different diseases, so the drug that inhibits or activates this target protein has the potential to treat these diseases. Metformin is the first line treatment for type 2 diabetes via reducing hepatic glucose production and internal absorption of glucose while improving insulin sensitivity16. Metformin activates a cellular energy sensor, AMP-activated protein kinase (AMPK) which is considered to be a significant pathway in atrial fibrillation (AF) through transcriptomic-based network analysis17. Semaglutide is a glucagon-like peptide 1 (GLP-1) receptor agonist (RA) that controls type 2 diabetes, treat obesity, and reduce risk of cardiovascular diseases1820. Semaglutide promotes secretion of insulin from pancreatic beta cells and cuts production of glucagon from pancreatic alpha cells, which reduces fasting and postprandial plasma glucose21. Administration of GLP-1 receptor agonists is reported to lower the rewarding effect of alcohol and drugs, thus mediating substance abuse22. GLP-1 receptor may also be associated with neurodegenerative diseases, such as AD, via insulin signaling23. Studies show semaglutide demonstrates neuroprotection in rat model via reducing inflammation and apoptosis24. These successful examples of drug repurposing have offered effective strategies to rapid development of potential treatment for human disease using disease-centric or target-centric approaches (Figure 1).

Yet, existing data, including genomics, transcriptomics, proteomics, longitudinal real-world data, has not yet been fully utilized and integrated for disease-modifying drug repurposing for human disease. Systematic characterization and identification of underlying pathobiology could provide a foundation for identifying disease-modifying targets and repurposable medicine. Integration of the genome, transcriptome, proteome and human interactome from diverse populations is essential for such identification using computational models as discussed below.

3. Artificial Intelligence (AI) for Drug Repurposing

3.1. Artificial Intelligence (AI)-ready datasets

Effective drug discovery pipelines require multi-modal biomedical data, such as genomics, transcriptomics, proteomics, metabolomics, imaging, biofluid markers, real-world patient data (Figure 1), and in-vivo and in vitro validation in both animal and human models. These rich databases offer necessary data for developing AI/ML models or tools for drug repurposing. There are 5 different types of databases (Table 1) which can be leveraged in drug repurposing: (1) chemoinformatic database, where structures, formulas, and other properties of molecules are included; (2) bioinformatics database, where structure, function, and additional information of proteins and genes are provided; (3) systems biology database, where molecular reactions and interactions with diseases are stored; (4) multi-omics database, where details of diseases, and genes and mutations associated with diseases are deposited, and (5) pharmacological database, where interactions and bindings between drugs and targets are present. Commonly used chemoinformatic databases (Table 1) include ChEMBL25, PubChem26, and DrugBank27. Protein and gene databases include UniProt28, PDB29, AlphaFold Protein Structure Database30, GenBank31, and Ensembl32. Systems biology and pathway databases include KEGG33 and Reactome34. Disease-gene databases include DISEASES35 and DisGeNET36. Drug-target interaction databases include BindingDB37 and PDBbind38. These comprehensive databases offer rich AI-ready datasets to build and evaluate various AI/ML models for drug repurposing/repositioning studies.

Table 1.

The list of the AI/ML ready databases and selected tools for drug repurposing.

Database name Database type Description Website
ChEMBL25 Chemoinformatic A database including chemical and genomic information of bioactive compounds. https://www.ebi.ac.uk/chembl/
DrugBank27 A database with drugs, drug targets, and clinical information. https://go.drugbank.com/
PubChem26 A database of molecule properties, structure, and clinical information. https://pubchem.ncbi.nlm.nih.gov/
BindingDB37 Drug-target A database containing information of protein-ligand binding affinity. https://www.bindingdb.org/
PDBbind38 A database of experimentally determined binding affinity data of protein-ligand binding. http://www.pdbbind.org.cn/
Diseases35 Disease-gene A database of disease-gene associations curated from literature, mutation data, and genome-wide association studies. https://diseases.jensenlab.org/
DisGeNet123 A database of genomic and human diseases. https://disgenet.com/
GeneCards124 A database of genomic, genetic, clinical, and functional information. https://www.genecards.org/
Open Targets125 A database of identification and prioritization of potential genes related with diseases. https://platform.opentargets.org/
Epic Cosmos Electronic health record (EHR) A database of EHR data from Epic system. https://cosmos.epic.com/
TriNetX A database of EHR data from multiple EHR systems, including claims and mortality data. https://trinetx.com/solutions/real-world-datasets/
AlphaFold Protein Structure Database30 Protein-gene A database of AlphaFold predicted protein structures. https://alphafold.ebi.ac.uk/
Ensembl126 A database of annotated genomes information. https://www.ensembl.org/index.html
GenBank31 A database of annotated genetic sequences. https://www.ncbi.nlm.nih.gov/genbank/
PDB127 A database of experimentally validated protein crystal 3D structures. https://www.rcsb.org/
UniProt28 A database of protein sequences and functions. https://www.uniprot.org/
Interactome INSIDER128 Systems biology and pathway A protein-protein interaction database with genomic variant information. http://interactomeinsider.yulab.org/
KEGG33 A database of genes and pathways. https://www.genome.jp/kegg/
Reactome129 A database of human pathways and genes. https://reactome.org/
STRING130 A database of functional protein-protein association networks https://string-db.org/

3.2. Machine Learning (ML) and Deep Learning (DL) techniques

Machine learning (ML) and deep learning (DL) are subfields of AI, which utilize different algorithms to learn pattern from input data39. ML and DL algorithms can be categorized to three major types: (1) supervised learning, where label information is required; (2) unsupervised learning, where no label information is needed; and (3) semi-supervised learning, where labels for only a small portion of the data are necessary. ML algorithms usually use structured or vector data. Molecular fingerprints, which capture chemical features of molecules, are commonly used as input of ML algorithms. Circular fingerprints40, molecular access system (MACCS) fingerprints41, and PubChem fingerprints42 are widely adopted to represent structures and chemical properties of molecules. DL methods can be applied to unstructured data, such as simplified molecular input line entry system (SMILES)43 formulas, amino acid sequences, molecular images, 3D structures, and electronic health records (EHRs). Drug-target interactions (DTIs) are primarily used in target-centric drug repurposing (Figure 1). Several early studies combined fingerprints and ML models to predict DTIs, which is a fundamental step in target-centric drug repurposing4446. Beyond using molecular fingerprints, similarities between drugs or targets can also be used as input to classical ML algorithms. In this case, similarities between chemical structure, side effect profile, amino acid sequence, and gene expression responses of drugs and targets can be measured using different metrics, followed by generating similarity matrices to recognize potential repurposable drugs and their corresponding targets for diseases47. Jacob et al.48, Bleakley et al.49, and Mordelet et al.50 developed several models based on classical ML techniques to leverage similarity matrices target-centric drug repurposing.

Sequence models are user-friendly as they do not require hard work to obtain fingerprints or similarity matrices. Smiles2vec51, DeepSMILES52, CHEM-BERT53, and SMILES-BERT54 are several notable examples which leverage sequence data. Sequence models can achieve promising results, yet they lack information about the 3D structure of molecules and proteins which is essential to the function of drugs and targets. Molecular image data provides 2-dimensional (2D) information about molecular structure, leading to better performance. DEEPScreen55 and Chemception56 are early examples that use molecular images in drug discovery. ImageMol exploits the power of pretraining to improve accuracy and generalizability57. There are two primary ways to encode 3D structures: using nodes and edges in geometric data to represent atoms and connections or saving molecules from different angles as a video. A notable example of this is VideoMol58, where each molecule is encoded in a video to elucidate the dynamics of molecules. Another trend in DL techniques is going multi-modality, where distinct modalities of data, such as sequence data and structure data, are fused into one DL model to provide comprehensive understanding of drugs and targets59. MRL-Mol60 and MMELON61 provide pilot studies in combining diverse modalities of molecules, while CLEAN-Contact62 achieves better performance by incorporates different modalities of proteins.

3.3. Network-based approaches

Network-based techniques are largely used in disease-centric drug repurposing, as they use information derived from disease-related networks to find repurposable drugs for a specific disease63. Network-based techniques use graph algorithms to analyze graph data which consists of nodes and edges. In particular, nodes can represent either drugs, targets, patients, genes, or pathways, while edges represent relationships between different entities. Widely used networks include the protein-protein interactome (PPI) network, which helps understanding how protein functions within cells, and the drug-drug-disease network, which offers insight into how different drugs can interact with each other and cure diseases. Network-based techniques stand out at learning undiscovered relations between drugs and their potential targets, and diseases with their target proteins. A systems pharmacology-based network approach was developed in order identify hydroxychloroquine, an anti-malarial and anti-rheumatic drug, as a potential repurposable drug to reduce risk of cardiovascular diseases (CVDs) 64. More recently, a genome-wide positioning systems network (GPSnet)65 was developed to aim at human PPI networks with patients’ DNA and RNA sequences mapped. Ouabain, a cardiac arrhythmia and heart failure drug, was identified with anti-tumor activity in lung adenocarcinoma. DeepDTnet66 achieves high accuracy by integrating 15 different types of networks and predicts topotecan, a topoisomerase inhibitor, to be a potential therapy against multiple sclerosis (MS) by inhibiting human retinoic-acid-receptor-related orphan receptor-gamma t (ROR-γt). By concentrating on endophenotype, an intermediate characteristic on the pathway between the genotype and disorder, a medicine network found sildenafil as a potential treatment to AD13.

3.4. Clinical trial emulation from real-world patient databases

Randomized controlled trials (RCT) are the gold standard for drug development67. However, due to the stringent definition of eligibility criteria, the number of eligible trial participants is typically small. This makes the trial cohort “ideal” rather than representative of real-world patient populations. Real-world data (RWD) are comprised of practice-based observations from real-world patients and thus provide a valuable resource for population-based validation of drug repurposing68. However, RWD is challenged by confounding factors, like sex, race, and lack of detailed clinical, biomarker, genetic information, socioeconomic, and other unknown factors. Recently there have been attempts to replicate the treatment effects obtained from RCT in RWD with the help of AI techniques. RWD cohorts with features like those of the trial participants were identified, and if individual data were available for RCT participants then (weighted) propensity score matching approaches can be applied. Due to complicated confounding factors in RWD (e.g., high-dimensionality and temporality), classical ML approaches are unable to estimate the propensity scores with high accuracy. Deep learning approaches can resolve this problem through the target trial emulation method69,70. This method can be applied in large-scale insurance claims and found zolpidem (a FDA-approved anti-insomnia medicine) as a potential drug for slowing progression of Parkinson’s Dementia71. Using a deep learning framework on emulating clinical trials from RWD72, they identified 14 drug candidates which reduced risk of for AD in patient subgroups with specific clinical features. Under high-throughput target trial emulation, the team identified five top-ranked drugs (pantoprazole, gabapentin, atorvastatin, fluticasone, and omeprazole) originally intended for other indications with potential benefits for AD patients73.

Due to the complex confounding situations in RWD, conventional propensity score estimation approaches based on logistic regression may not be able to accurately estimate the probability of treatment and censoring. Advanced AI approaches may achieve this goal. Recent studies have evaluated three types of strategies for the representation of EHR data for predictions: (1) convolutional neural network (CNN)-based methods, which represents EHR of each patient as an event by time matrix and perform a series of one-side convolution, activation plus pooling, followed by a final fully connected layer to perform the predictions74; (2) recurrent neural network (RNN)-based methods, which represents each individual’s EHR as an event sequence and can be used to train a RNN-based prediction model75; and (3) graph-based methods, which represents the EHR as a heterogeneous information network with patients and clinical events as nodes and event co-occurrence as edges, applying graph neural network based approaches to perform the predictions76. Yet, a recent study showed that the deep learning-based propensity score model did not necessarily outperform logistic regression-based methods in confounding factor justification73. In addition, data missingness is another feature of noisy RWD. To handle missingness not at random, researchers can consider selection model-based methods (e.g., outcome-dependent sampling in longitudinal outcomes77). In addition, we can also utilize DL-based imputation methods such as the last observation carry-forward78 and generative-adversarial nets (GANs)79.

3.5. Biomedical Knowledge Graph for drug repurposing hypothesis generation

With the rich, relevant, and high-quality knowledge found in the literature, we can both validate the drug repurposing hypotheses generated from RWD and generate knowledge based repurposing hypotheses on their own. A recent study first generated embeddings for the entities (drugs, diseases, or genes/proteins) and relations (i.e., drug-disease associations) using knowledge graph embedding methods such as DGL-KE80, suitable for large-scale data. These embedding methods can generate vector-based representations for both entities and relations in the same embedded semantic space, such that the repurposing hypotheses can be generated according to the ranking scores of <drug, relation, disease> triples evaluated with vector-based similarities. Using a graph foundation model for zero-shot drug repurposing from a large biomedical knowledge graph, a graph neural network and metric learning module with high accuracy were developed to rank drugs as potential indications and contraindications for 17,080 diseases, including a large number of rare diseases81.

Another work developed a comprehensive biomedical knowledge graph concerning COVID-1982. This COVID-19 knowledge graph consists of 15 million edges and 39 different types of relations, including connections between drugs, diseases, genes, anatomies, pharmacologic classes, and gene expression. Similar to the previous work, an embedding model was developed to extract representation vectors for <head entity, relation type, tail entity>. A deep learning model trained on this knowledge graph identified more than 40 potential repurposable drugs for COVID-19. Additional enrichment analysis was conducted to validate the predicted repurposable drugs which shows high confidence in treating COVID-19. Such biomedical knowledge graphs accompanied with highly accurate deep learning models are important AI methods in advancing drug repurposing.

3.6. Clinical and experimental validation of AI-based drug repurposing

Experimental and clinical validation are crucial steps after repurposing drugs using AI models, as such validations could confirm their real-world accuracy and reliability. Primary approaches to conduct experimental validation include using induced pluripotent stem cells (iPSCs) derived models and animal models68, while clinical validation may use electronic health record (EHR) data or health insurance claim data (Figure 1). In a recent work where sildenafil was repurposed as a potential AD treatment, AD patient iPSC-derived neurons treated with sildenafil showed significant reduction in phosphorylated-tau181 (p-tau 181), an early biomarker of AD pathology13. Another work conducted additional experiments to validate the effectiveness of sildenafil on the reduction of both p-tau 181 and p-tau 20515. Additionally, enrichment analysis of differentially expressed genes (DEGs) between control and sildenafil treated groups displayed a neuroprotective effect. Another work which repurposed metformin as a therapeutic treatment for atrial fibrillation (AF) leveraged human iPSC-derived atrial-like cardiomyocytes (a-iCMs) to validate in silico predictions17. Critical cardiovascular-related markers were significantly upregulated accompanied by an increase in expression of several known markers by metformin associated with low expression in AF. In another study, AD patient iPSC-derived microglia were treated with ketorolac, a moderate to severe pain treatment by downregulating the Type-I interferon signaling, mechanistically supporting potential beneficial of ketorolac in reducing incidence of AD via targeting disease-relevant microglia83.

4. Applications of AI Approaches to Drug Repurposing

Critically, drug repurposing depends on efficient searching of the vast drug space, for which the optimal approach is rapidly evolving. As drug repurposing is a complex process involving many steps, multimodal machine learning tools can significantly reduce the time and cost of drug development. For instance, multimodal machine learning approaches84,85 improve accuracy of patient subphenotyping during clinical trial design by assembling neuroimaging, genetic, and multi-omics profiling data. With the help of deep learning, effective representations can be learned for different data modalities86,87, which can then be fused by simple concatenation or more complicated nonlinear transformation88 to perform downstream tasks such as molecular design, pharmacokinetics property evaluation and optimization. AI/ML tools has become a leading technique in expediting and reducing the cost of the drug development. We next turned to utilize four types of human challenging complex diseases (including AD, cancer, COVID-19, and cardiovascular diseases) to illustrate how AI accelerates the finding of treatments by repurposing approved drugs.

There are 50 repurposable drug trials (40 unique repurposed agents) based on the latest Alzheimer’s disease drug development pipeline 202489. To expand drug repurposing efforts, a recent machine learning-based framework, named DRIAD (Drug Repurposing In AD), has been developed to quantify the potential associations between AD biological processes and linked genetic datasets, thereby prioritizing drug candidates for repurposing90. DRIAD prioritized baricitinib as a candidate AD drug and baricitinib is being tested in an open-label, biomarker-driven basket trial (ClinicalTrials.gov Identifier: NCT05189106) in people with AD and Amyotrophic lateral sclerosis (ALS).

Another recently developed tool, AlzGPS, is a systems biology platform with over 100 multi-omics datasets that capture molecular profiles underlying AD pathobiology91. This tool enables network-based prioritization of potential targets for AD drug repurposing91. NETTAG is a network topology-based deep learning framework to identify disease-associated genes for AD and prioritize candidate drugs and targets. Using NETTAG92, the team successfully identified gemfibrozil (an approved lipid regulator) is significantly associated with reduced risk of AD compared to simvastatin using an active-comparator design from real-world patient data. A deep learning methodology (deepDTnet) was developed for new target identification and drug repurposing in a heterogeneous drug-gene-disease network embedded by 15 types of chemical, genomic, phenotypic, and cellular network profiles66 (Table 1). Trained on 732 U.S. FDA-approved small molecule drugs, deepDTnet shows high accuracy in identifying novel molecular targets for known drugs, outperforming previously published state-of-the-art methodologies. Importantly, the authors experimentally validated deepDTnet-predicted topotecan as a candidate repurposable drug for multiple sclerosis by targeting human retinoic-acid-receptor-related orphan receptor-gamma t (ROR-γt)66.

Using multimodal analysis of single-cell/nucleus RNA-sequencing data from AD patient brains, two approved asthma drugs (fluticasone and mometasone) were found to be significantly associated with reduced likelihood of AD by targeting AD-associated microglia93. Via analysis of real-world electronic insurance record data from 7.2 million patients from the IBM® MarketScan® Medicare Supplemental Database, two FDA-approved p300/CBP inhibitors, salsalate or diflunisal, were found to be associated with decreased incidence of AD, and neuroprotective efficacy was also validated in mice94,95. Using an endophenotype-based in silico network medicine approach13, a team showed that sildenafil usage was significantly associated with reduced likelihood of AD, and further validated the findings using an AD patient induced Pluripotent Stem Cell (iPSC)-derived neuron model13. Another team demonstrated that bumetanide (an FDA-approved oral diuretic) provides a potential treatment for apolipoprotein APOE4-related AD using in silico approaches combining experimental and real-world evidence96. These prototypical examples illustrate how AI-based multi-omics approaches combined with real-world patient databases and experimental approaches can rapidly identify potential treatments for AD (Figure 1). The details of computational methods for AD drug repurposing can be found in recent reviews46.

Cancer indicates challenging diseases that are caused by malignant tumors involving uncontrolled growth of cells in the body97. Cancer is the second biggest contributor to the death in the U.S. with lung cancer as the deadliest cancer98. A major method to apply AI to repurpose drugs for cancer is leveraging cancer cell line models and patient-derived primary cells99. A ML algorithm based on drug response to human breast cancer cell line finds 16 from 28 drugs have significantly better performance in treating human breast cancer. Additionally, electronic health records (EHRs) of patients can be used for predicting drug effectiveness on patients with cancer. For example, investigations have been made to discover potential drugs that can lower cancer mortality by using EHRs from Mayo Clinic and Vanderbilt University Medical Center100.

COVID-19 is an infectious disease caused by SARS-CoV-2 virus, which became a world pandemic101. In the early stage of the pandemic, AI is a key technique in discovering drugs repurposable to the treatment of COVID-19102. Studies has been focused on inhibiting SARS-CoV-2 spike proteins, a key essential in the infection of virus, and reducing inflammation and immune response caused by the disease, which plays a vital role in the mortality of patients. A network-based approach identified 16 drugs and 3 drug combinations with the potential to treat COVID-19103. ImageMol has been used to find inhibitors of SARS-CoV-2 spike proteins57. Remarkably, melatonin, a hormone that regulates sleep, and dexamethasone, a glucocorticoid agonist, were proposed to reduce inflammation and immune response, and toremifene was presented as a potential SARS-CoV-2 spike protein inhibitor82,104,105.

Cardiovascular disease (CVD) refers to a group of diseases and disorders concerning heart and blood vessels106. CVD is the primary reason of death around the globe. Main symptoms of CVD include chest pain, palpitations, and shortness of breath. By studying human protein-protein interactome, carbamazepine, a drug to control seizures, is believed to increase the risk of CVD, while hydroxychloroquine, a malaria treatment, is considered to lower the risk of CVD64. By incorporating information about adverse effects of drugs, approved drugs have been repurposed for CVD treatment while lowering potential adverse effects107.

5. Discussion, Perspective, and Future Directions

In this review, we briefly summarizied and discussed existing AI techniques and their applications to repurposing drugs in several challenging diseases. Despite substantial advancements in AI recently, the research circles of AI and biomedicine still face significant challenges. One formidable challenge is that the selection of hyperparameters, such as learning rate and size of hidden layers, can substantially affect the model’s performance. Early strategies to find optimal hyperparameters mainly use random search and grid search, which can be computationally expensive for complex AI models with large hyperparameter space108. Recent development of automated machine learning (AutoML) incorporating neural architecture search (NAS) provides a viable solution for this issue109. Even though an accurate AI model is trained with high-quality labeled datasets, experimental validation of drug discovery results is still required to confirm.

Recent growth in large language models (LLMs) provides heartening solutions for advancing drug repurposing through the integration of heterogeneous data sources. These AI models can have hundreds of billions to trillions of parameters while trained on innumerable data and documents110. Such training equips LLMs with the ability to find relations between drugs, diseases, targets, and pathways. Yet LLMs trained on generic data are not immune to hallucinations and biased responses111. Retrieval-augmented generation (RAG) is a method where a specialized dataset is used to supplement LLMs to reduce hallucinations and biased responses. Moreover, AI agents, emerging AI systems designed to automatically carry out tasks and complete goals with limited human interaction, are another promising direction to drug discovery112,113. A future AI agentic system based on LLMs could execute different kinds of drug discovery inquiries by leveraging multiple state-of-the-art drug repurposing tools, such as AlphaFold3114 and NETTAG92 (Figure 1). By integrating biomedical knowledge graphs (KGs) into the LLMs of AI agents, patients, physicians, and researchers can ask any questions regarding drugs and diseases. For example, physicians can ask the AI agent to find a repurposable drug for a patient, as the patient may display severe adverse effects on existing drugs for the disease.

AI-based drug repurposing communities are leading big data and open science. Data used for training AI models must be annotated by domain experts to ensure the quality of AI models before effectively advancing the drug development process for human complex diseases. The application of AI solutions to drug repurposing is possible via cross-disciplinary teamwork and cooperation to address current gaps and challenges, such as analysis of clinical trial data and application of trial outcomes to patients who could benefit greatly from repurposable drugs. Altogether, AI is an indispensable aspect of the future of drug repurposing and precision medicine for human challenging diseases.

There are several intimidating challenges in building the current AI-ready datasets. First, multi-omics and clinical data are generated from heterogeneous patient samples across different laboratories and health care systems. Harmonization of clinical and multi-omics data plays a crucial role in securing the quality of AI models in real-world drug repurposing and data harmonization is challenging in for most basic and clinical scientists. Limited data sharing is another hurdle in AI-based drug repurposing. Bio-pharmaceutical companies have generated massive data compared to academic institutions, which cannot be shared due to various intellectual property issues. Because of the complexity of human diseases, heterogeneous data sets covering genomic, cellular, clinical, and behavioral aspects are required to advance disease understanding115,116. Collaborations across different entities and institutions therefore need to incorporate diverse patient populations, as well as participants with diverse disease characteristics117, using AI technologies.

Data security is another key to inclusiveness and maintaining the trust of AI technologies. For example, Fast Healthcare Interoperability Resource (FHIR)118 for clinical data transmission, are crucial to data credibility and confidentiality. Another potential solution is federated learning119, which protects individual patient data through collective learning from multiple local sites without transferring the original raw data. Another important direction is to improve transparency and interpretability of AI models so that scientists can understand the decision-making process during drug repurposing in order to insure proper assessments of usability and potential failure cases, such as Explainable AI technologies92. Thus, model interpretability can be contextualized120, with different levels of model transparency required for different applications. Model interpretation methods have various potential uses121. For example, knowledge distillation is a popular interpretation technique that aims at advancing a secondary interpretable model to approximate the prediction results of the first model122. These models could be vulnerable to adversarial attacks, which could manipulate the model explanations deliberately and make them unreliable. Deriving robust, reliable, and secure model interpretations is therefore an important future research direction.

Table 2.

Commonly used data resources to develop AI drug repurposing tools.

Name Description Link
CHEM-BERT53 A molecular sequence model that predicts molecular properties and drug targets. https://github.com/HyunSeobKim/CHEM-BERT
MMELON61 A foundation multi-modal model that integrates different modalities of small molecules to predict molecular properties and drug targets. https://github.com/BiomedSciAI/biomed-multi-view
ImageMol57 A foundation model that trained on molecule images to predict drug targets and molecule properties. https://github.com/ChengF-Lab/ImageMol
VideoMol58 A foundation model that converts molecule images to videos to study molecule properties and activities. https://github.com/ChengF-Lab/VideoMol
deepDR131 A network-based model combining 10 types of networks to repurpose drugs. https://github.com/ChengF-Lab/deepDR
deepDTnet66 A network-based model combining 15 types of networks to identify drug targets. https://github.com/ChengF-Lab/deepDTnet
LISA-CPI132 A molecular image model combined with target 3D structure information to predict drug target binding activity. https://github.com/ChengF-Lab/LISA-CPI
Chemception56 A molecular image model predicts activity on drug targets. https://github.com/Abdulk084/Chemception
MolCLR133 A molecular structure-based model that can predict molecular properties and drug targets. https://github.com/yuyangw/MolCLR

Funding:

This work was primarily supported by the National Institute on Aging (NIA) under Award Number U01AG073323, R21AG083003, R01AG066707, R01AG076448, R01AG082118, RF1AG082211, and RF1NS133812 to F.C., and by the National Heart, Lung, and Blood Institute (NHLBI) under award number P01HL158502 to M.C.

Footnotes

Competing interests.

The authors don’t have any competing interests.

References

  • 1.Zhou Y, Wang F, Tang J, Nussinov R, and Cheng F (2020). Artificial intelligence in COVID-19 drug repurposing. Lancet Digit Health 2, e667–e676. 10.1016/S2589-7500(20)30192-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Zhou Y, Hou Y, Shen J, Huang Y, Martin W, and Cheng F (2020). Network-based drug repurposing for novel coronavirus 2019-nCoV/SARS-CoV-2. Cell Discov 6, 14. 10.1038/s41421-020-0153-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zhou Y, Liu Y, Gupta S, Paramo MI, Hou Y, Mao C, Luo Y, Judd J, Wierbowski S, Bertolotti M, et al. (2023). A comprehensive SARS-CoV-2-human protein-protein interactome reveals COVID-19 pathobiology and potential host therapeutic targets. Nat Biotechnol 41, 128–139. 10.1038/s41587-022-01474-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Fang J, Pieper AA, Nussinov R, Lee G, Bekris L, Leverenz JB, Cummings J, and Cheng F (2020). Harnessing endophenotypes and network medicine for Alzheimer’s drug repurposing. Med Res Rev 40, 2386–2426. 10.1002/med.21709. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Paranjpe MD, Taubes A, and Sirota M (2019). Insights into Computational Drug Repurposing for Neurodegenerative Disease. Trends Pharmacol Sci 40, 565–576. 10.1016/j.tips.2019.06.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ballard C, Aarsland D, Cummings J, O’Brien J, Mills R, Molinuevo JL, Fladby T, Williams G, Doherty P, Corbett A, and Sultana J (2020). Drug repositioning and repurposing for Alzheimer disease. Nat Rev Neurol 16, 661–673. 10.1038/s41582-020-0397-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Parisi D, Adasme MF, Sveshnikova A, Bolz SN, Moreau Y, and Schroeder M (2020). Drug repositioning or target repositioning: A structural perspective of drug-target-indication relationship for available repurposed drugs. Computational and Structural Biotechnology Journal 18, 1043–1055. 10.1016/j.csbj.2020.04.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Boller F, Mizutani T, Roessmann U, and Gambetti P (1980). Parkinson disease, dementia, and alzheimer disease: Clinicopathological correlations. Annals of Neurology 7, 329–335. 10.1002/ana.410070408. [DOI] [PubMed] [Google Scholar]
  • 9.Galiè N, Ghofrani Hossein A, Torbicki A, Barst Robyn J, Rubin Lewis J, Badesch D, Fleming T, Parpia T, Burgess G, Branzi A, et al. Sildenafil Citrate Therapy for Pulmonary Arterial Hypertension. New England Journal of Medicine 353, 2148–2157. 10.1056/NEJMoa050010. [DOI] [PubMed] [Google Scholar]
  • 10.Goldstein I, Lue Tom F, Padma-Nathan H, Rosen Raymond C, Steers William D, and Wicker Pierre A Oral Sildenafil in the Treatment of Erectile Dysfunction. New England Journal of Medicine 338, 1397–1404. 10.1056/NEJM199805143382001. [DOI] [PubMed] [Google Scholar]
  • 11.Andersson KE (2018). PDE5 inhibitors – pharmacology and clinical applications 20 years after sildenafil discovery. British Journal of Pharmacology 175, 2554–2565. 10.1111/bph.14205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Sanders O (2020). Sildenafil for the Treatment of Alzheimer’s Disease: A Systematic Review. Journal of Alzheimer’s Disease Reports 4, 91–106. 10.3233/ADR-200166. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Fang J, Zhang P, Zhou Y, Chiang C-W, Tan J, Hou Y, Stauffer S, Li L, Pieper AA, Cummings J, and Cheng F (2021). Endophenotype-based in silico network medicine discovery combined with insurance record data mining identifies sildenafil as a candidate drug for Alzheimer’s disease. Nature Aging 1, 1175–1188. 10.1038/s43587-021-00138-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Cuadrado-Tejedor M, Hervias I, Ricobaraza A, Puerta E, Pérez-Roldán JM, García-Barroso C, Franco R, Aguirre N, and García-Osta A (2011). Sildenafil restores cognitive function without affecting β-amyloid burden in a mouse model of Alzheimer’s disease. Br J Pharmacol 164, 2029–2041. 10.1111/j.1476-5381.2011.01517.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gohel D, Zhang P, Gupta AK, Li Y, Chiang CW, Li L, Hou Y, Pieper AA, Cummings J, and Cheng F (2024). Sildenafil as a Candidate Drug for Alzheimer’s Disease: Real-World Patient Data Observation and Mechanistic Observations from Patient-Induced Pluripotent Stem Cell-Derived Neurons. J Alzheimers Dis 98, 643–657. 10.3233/JAD-231391. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Rena G, Hardie DG, and Pearson ER (2017). The mechanisms of action of metformin. Diabetologia 60, 1577–1585. 10.1007/s00125-017-4342-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lal JC, Mao C, Zhou Y, Gore-Panter SR, Rennison JH, Lovano BS, Castel L, Shin J, Gillinov AM, Smith JD, et al. (2022). Transcriptomics-based network medicine approach identifies metformin as a repurposable drug for atrial fibrillation. Cell Reports Medicine 3. 10.1016/j.xcrm.2022.100749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Bergmann NC, Davies MJ, Lingvay I, and Knop FK (2023). Semaglutide for the treatment of overweight and obesity: A review. Diabetes, Obesity and Metabolism 25, 18–35. 10.1111/dom.14863. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Andersen A, Knop FK, and Vilsbøll T (2021). A Pharmacological and Clinical Overview of Oral Semaglutide for the Treatment of Type 2 Diabetes. Drugs 81, 1003–1030. 10.1007/s40265-021-01499-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Marso Steven P, Bain Stephen C, Consoli A, Eliaschewitz Freddy G, Jódar E, Leiter Lawrence A, Lingvay I, Rosenstock J, Seufert J, Warren Mark L, et al. Semaglutide and Cardiovascular Outcomes in Patients with Type 2 Diabetes. New England Journal of Medicine 375, 1834–1844. 10.1056/NEJMoa1607141. [DOI] [PubMed] [Google Scholar]
  • 21.Shah M, and Vella A (2014). Effects of GLP-1 on appetite and weight. Reviews in Endocrine and Metabolic Disorders 15, 181–187. 10.1007/s11154-014-9289-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Shirazi RH, Dickson SL, and Skibicka KP (2013). Gut Peptide GLP-1 and Its Analogue, Exendin-4, Decrease Alcohol Intake and Reward. PLOS ONE 8, e61965. 10.1371/journal.pone.0061965. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Hölscher C (2018). Novel dual GLP-1/GIP receptor agonists show neuroprotective effects in Alzheimer’s and Parkinson’s disease models. Neuropharmacology 136, 251–259. 10.1016/j.neuropharm.2018.01.040. [DOI] [PubMed] [Google Scholar]
  • 24.Yang X, Feng P, Zhang X, Li D, Wang R, Ji C, Li G, and Hölscher C (2019). The diabetes drug semaglutide reduces infarct size, inflammation, and apoptosis, and normalizes neurogenesis in a rat model of stroke. Neuropharmacology 158, 107748. 10.1016/j.neuropharm.2019.107748. [DOI] [PubMed] [Google Scholar]
  • 25.Gaulton A, Bellis LJ, Bento AP, Chambers J, Davies M, Hersey A, Light Y, McGlinchey S, Michalovich D, Al-Lazikani B, and Overington JP (2012). ChEMBL: a large-scale bioactivity database for drug discovery. Nucleic Acids Research 40, D1100–D1107. 10.1093/nar/gkr777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kim S, Chen J, Cheng T, Gindulyte A, He J, He S, Li Q, Shoemaker BA, Thiessen PA, Yu B, et al. (2023). PubChem 2023 update. Nucleic Acids Research 51, D1373–D1380. 10.1093/nar/gkac956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Knox C, Wilson M, Klinger, Christen M, Franklin M, Oler E, Wilson A, Pon A, Cox J, Chin NE, Strawbridge, Seth A, et al. (2024). DrugBank 6.0: the DrugBank Knowledgebase for 2024. Nucleic Acids Research 52, D1265–D1275. 10.1093/nar/gkad976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.The UniProt C (2019). UniProt: a worldwide hub of protein knowledge. Nucleic Acids Research 47, D506–D515. 10.1093/nar/gky1049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Berman HM, Westbrook J, Feng Z, Gilliland G, Bhat TN, Weissig H, Shindyalov IN, and Bourne PE (2000). The Protein Data Bank. Nucleic acids research 28, 235–242. 10.1093/nar/28.1.235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Varadi M, Anyango S, Deshpande M, Nair S, Natassia C, Yordanova G, Yuan D, Stroe O, Wood G, Laydon A, et al. (2022). AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Research 50, D439–D444. 10.1093/nar/gkab1061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, and Sayers EW (2013). GenBank. Nucleic Acids Research 41, D36–D42. 10.1093/nar/gks1195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Cunningham F, Allen JE, Allen J, Alvarez-Jarreta J, Amode MR, Armean, Irina M, Austine-Orimoloye O, Azov, Andrey G, Barnes I, Bennett R, et al. (2022). Ensembl 2022. Nucleic Acids Research 50, D988–D995. 10.1093/nar/gkab1049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kanehisa M, and Goto S (2000). KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research 28, 27–30. 10.1093/nar/28.1.27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Gillespie M, Jassal B, Stephan R, Milacic M, Rothfels K, Senff-Ribeiro A, Griss J, Sevilla C, Matthews L, Gong C, et al. (2022). The reactome pathway knowledgebase 2022. Nucleic Acids Research 50, D687–D692. 10.1093/nar/gkab1028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Grissa D, Junge A, Oprea TI, and Jensen LJ (2022). Diseases 2.0: a weekly updated database of disease–gene associations from text mining and data integration. Database 2022, baac019. 10.1093/database/baac019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Piñero J, Ramírez-Anguita JM, Saüch-Pitarch J, Ronzano F, Centeno E, Sanz F, and Furlong LI (2020). The DisGeNET knowledge platform for disease genomics: 2019 update. Nucleic Acids Research 48, D845–D855. 10.1093/nar/gkz1021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Liu T, Lin Y, Wen X, Jorissen RN, and Gilson MK (2007). BindingDB: a web-accessible database of experimentally determined protein–ligand binding affinities. Nucleic Acids Research 35, D198–D201. 10.1093/nar/gkl999. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wang R, Fang X, Lu Y, Yang C-Y, and Wang S (2005). The PDBbind Database:  Methodologies and Updates. Journal of Medicinal Chemistry 48, 4111–4119. 10.1021/jm048957q. [DOI] [PubMed] [Google Scholar]
  • 39.Jordan MI, and Mitchell TM (2015). Machine learning: Trends, perspectives, and prospects. Science 349, 255–260. 10.1126/science.aaa8415. [DOI] [PubMed] [Google Scholar]
  • 40.Glem RC, Bender A, Arnby CH, Carlsson L, Boyer S, and Smith J (2006). Circular fingerprints: flexible molecular descriptors with applications from physical chemistry to ADME. IDrugs 9, 199–204. [PubMed] [Google Scholar]
  • 41.Durant JL, Leland BA, Henry DR, and Nourse JG (2002). Reoptimization of MDL Keys for Use in Drug Discovery. Journal of Chemical Information and Computer Sciences 42, 1273–1280. 10.1021/ci010132r. [DOI] [PubMed] [Google Scholar]
  • 42.Fernández-de Gortari E, García-Jacas CR, Martinez-Mayorga K, and Medina-Franco JL (2017). Database fingerprint (DFP): an approach to represent molecular databases. Journal of Cheminformatics 9, 9. 10.1186/s13321-017-0195-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Weininger D (1988). SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of chemical information and computer sciences 28, 31–36. [Google Scholar]
  • 44.Wang Y-C, Yang Z-X, Wang Y, and Deng N-Y (2010). Computationally Probing Drug-Protein Interactions Via Support Vector Machine. Letters in Drug Design & Discovery 7, 370-378. 10.2174/157018010791163433. [DOI] [Google Scholar]
  • 45.Ahn S, Lee SE, and Kim M. h. (2022). Random-forest model for drug–target interaction prediction via Kullback–Leibler divergence. Journal of Cheminformatics 14, 67. 10.1186/s13321-022-00644-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Shi H, Liu S, Chen J, Li X, Ma Q, and Yu B (2019). Predicting drug-target interactions using Lasso with random forest based on evolutionary information and chemical structure. Genomics 111, 1839–1852. 10.1016/j.ygeno.2018.12.007. [DOI] [PubMed] [Google Scholar]
  • 47.Mitchell JBO (2001). The Relationship between the Sequence Identities of Alpha Helical Proteins in the PDB and the Molecular Similarities of Their Ligands. Journal of Chemical Information and Computer Sciences 41, 1617–1622. 10.1021/ci010364q. [DOI] [PubMed] [Google Scholar]
  • 48.Jacob L, and Vert J-P (2008). Protein-ligand interaction prediction: an improved chemogenomics approach. Bioinformatics 24, 2149–2156. 10.1093/bioinformatics/btn409. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Bleakley K, Biau G, and Vert J-P (2007). Supervised reconstruction of biological networks with local models. Bioinformatics 23, i57–i65. 10.1093/bioinformatics/btm204. [DOI] [PubMed] [Google Scholar]
  • 50.Mordelet F, and Vert J-P (2008). SIRENE: supervised inference of regulatory networks. Bioinformatics 24, i76–i82. 10.1093/bioinformatics/btn273. [DOI] [PubMed] [Google Scholar]
  • 51.Goh GB, Hodas NO, Siegel C, and Vishnu A (2017). Smiles2vec: An interpretable general-purpose deep neural network for predicting chemical properties. arXiv preprint arXiv:1712.02034. [Google Scholar]
  • 52.O’Boyle N, and Dalke A (2018). DeepSMILES: an adaptation of SMILES for use in machine-learning of chemical structures.
  • 53.Kim H, Lee J, Ahn S, and Lee JR (2021). A merged molecular representation learning for molecular properties prediction with a web-based service. Scientific Reports 11, 11028. 10.1038/s41598-021-90259-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Wang S, Guo Y, Wang Y, Sun H, and Huang J (2019). SMILES-BERT: Large Scale Unsupervised Pre-Training for Molecular Property Prediction. Proceedings of the 10th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. Association for Computing Machinery. [Google Scholar]
  • 55.Rifaioglu AS, Nalbat E, Atalay V, Martin MJ, Cetin-Atalay R, and Doğan T (2020). DEEPScreen: high performance drug–target interaction prediction with convolutional neural networks using 2-D structural compound representations. Chemical science 11, 2531–2557. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Goh GB, Siegel C, Vishnu A, Hodas NO, and Baker N (2017). Chemception: a deep neural network with minimal chemistry knowledge matches the performance of expert-developed QSAR/QSPR models. arXiv preprint arXiv:1706.06689. [Google Scholar]
  • 57.Zeng X, Xiang H, Yu L, Wang J, Li K, Nussinov R, and Cheng F (2022). Accurate prediction of molecular properties and drug targets using a self-supervised image representation learning framework. Nature Machine Intelligence 4, 1004–1016. 10.1038/s42256-022-00557-6. [DOI] [Google Scholar]
  • 58.Xiang H, Zeng L, Hou L, Li K, Fu Z, Qiu Y, Nussinov R, Hu J, Rosen-Zvi M, Zeng X, and Cheng F (2024). A molecular video-derived foundation model for scientific drug discovery. Nature Communications 15, 9696. 10.1038/s41467-024-53742-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Ngiam J, Khosla A, Kim M, Nam J, Lee H, and Ng AY (2011). Multimodal deep learning. pp. 689–696. [Google Scholar]
  • 60.Yang Y, Wang Z, Ahadian P, Jerger A, Zucker J, Feng S, Cheng F, and Guan Q (2024). A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction. Proceedings of the Great Lakes Symposium on VLSI 2024. Association for Computing Machinery. [Google Scholar]
  • 61.Suryanarayanan P, Qiu Y, Sethi S, Mahajan D, Li H, Yang Y, Eyigoz E, Saenz AG, Platt DE, and Rumbell TH (2024). Multi-view biomedical foundation models for molecule-target and property prediction. arXiv preprint arXiv:2410.19704. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Yang Y, Jerger A, Feng S, Wang Z, Brasfield C, Cheung MS, Zucker J, and Guan Q (2024). Improved enzyme functional annotation prediction using contrastive learning with structural inference. Communications Biology 7, 1690. 10.1038/s42003-024-07359-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Cheng F, Desai RJ, Handy DE, Wang R, Schneeweiss S, Barabasi AL, and Loscalzo J (2018). Network-based approach to prediction and population-based validation of in silico drug repurposing. Nat Commun 9, 2691. 10.1038/s41467-018-05116-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Cheng F, Desai RJ, Handy DE, Wang R, Schneeweiss S, Barabási A-L, and Loscalzo J (2018). Network-based approach to prediction and population-based validation of in silico drug repurposing. Nature Communications 9, 2691. 10.1038/s41467-018-05116-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Cheng F, Lu W, Liu C, Fang J, Hou Y, Handy DE, Wang R, Zhao Y, Yang Y, Huang J, et al. (2019). A genome-wide positioning systems network algorithm for in silico drug repurposing. Nature Communications 10, 3476. 10.1038/s41467-019-10744-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Zeng X, Zhu S, Lu W, Liu Z, Huang J, Zhou Y, Fang J, Huang Y, Guo H, Li L, et al. (2020). Target identification among known drugs by deep learning from heterogeneous networks. Chemical Science 11, 1775–1797. 10.1039/C9SC04336E. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Cummings J, Zhou Y, Lee G, Zhong K, Fonseca J, and Cheng F (2024). Alzheimer’s disease drug development pipeline: 2024. Alzheimers Dement (N Y) 10, e12465. 10.1002/trc2.12465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Fang J, Zhang P, Zhou Y, Chiang CW, Tan J, Hou Y, Stauffer S, Li L, Pieper AA, Cummings J, and Cheng F (2021). Endophenotype-based in silico network medicine discovery combined with insurance record data mining identifies sildenafil as a candidate drug for Alzheimer’s disease. Nat Aging 1, 1175–1188. 10.1038/s43587-021-00138-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Hernan MA, and Robins JM (2016). Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. Am J Epidemiol 183, 758–764. 10.1093/aje/kwv254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Sanchez P, Voisey JP, Xia T, Watson HI, O’Neil AQ, and Tsaftaris SA (2022). Causal machine learning for healthcare and precision medicine. R Soc Open Sci 9, 220638. 10.1098/rsos.220638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Laifenfeld D, Yanover C, Ozery-Flato M, Shaham O, Rosen-Zvi M, Lev N, Goldschmidt Y, and Grossman I (2021). Emulated Clinical Trials from Longitudinal Real-World Data Efficiently Identify Candidates for Neurological Disease Modification: Examples from Parkinson’s Disease. Front Pharmacol 12, 631584. 10.3389/fphar.2021.631584. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Lee S, Liu R, Cheng F, and Zhang P (2025). A Deep Subgrouping Framework for Precision Drug Repurposing via Emulating Clinical Trials on Real-world Patient Data. KDD; 2025, 2347–2358. 10.1145/3690624.3709418. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Zang C, Zhang H, Xu J, Zhang H, Fouladvand S, Havaldar S, Cheng F, Chen K, Chen Y, Glicksberg BS, et al. (2023). High-throughput target trial emulation for Alzheimer’s disease drug repurposing with real-world data. Nat Commun 14, 8180. 10.1038/s41467-023-43929-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Yang Z, Huang Y, Jiang Y, Sun Y, Zhang Y-J, and Luo P (2018). Clinical Assistant Diagnosis for Electronic Medical Record Based on Convolutional Neural Network. Scientific Reports 8, 6329. 10.1038/s41598-018-24389-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Rasmy L, Nigo M, Kannadath BS, Xie Z, Mao B, Patel K, Zhou Y, Zhang W, Ross A, Xu H, and Zhi D (2022). Recurrent neural network models (CovRNN) for predicting outcomes of patients with COVID-19 on admission to hospital: model development and validation using electronic health record data. The Lancet Digital Health 4, e415–e425. 10.1016/S2589-7500(22)00049-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Choi E, Xu Z, Li Y, Dusenberry M, Flores G, Xue E, and Dai A (2020). Learning the Graphical Structure of Electronic Health Records with Graph Convolutional Transformer. Proceedings of the AAAI Conference on Artificial Intelligence 34, 606–613. 10.1609/aaai.v34i01.5400. [DOI] [Google Scholar]
  • 77.Schildcrout JS, and Heagerty PJ (2008). On outcome-dependent sampling designs for longitudinal binary response data with time-varying covariates. Biostatistics 9, 735–749. 10.1093/biostatistics/kxn006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Shao J, and Zhong B (2003). Last observation carry-forward and last observation analysis. Statistics in Medicine 22, 2429–2441. 10.1002/sim.1519. [DOI] [PubMed] [Google Scholar]
  • 79.Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, and Bengio Y (2014). Generative adversarial nets. Advances in neural information processing systems 27. [Google Scholar]
  • 80.Zheng D, Song X, Ma C, Tan Z, Ye Z, Dong J, Xiong H, Zhang Z, and Karypis G (2020). DGL-KE: Training Knowledge Graph Embeddings at Scale. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery. [Google Scholar]
  • 81.Huang K, Chandak P, Wang Q, Havaldar S, Vaid A, Leskovec J, Nadkarni GN, Glicksberg BS, Gehlenborg N, and Zitnik M (2024). A foundation model for clinician-centered drug repurposing. Nat Med 30, 3601–3613. 10.1038/s41591-024-03233-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Zeng X, Song X, Ma T, Pan X, Zhou Y, Hou Y, Zhang Z, Li K, Karypis G, and Cheng F (2020). Repurpose Open Data to Discover Therapeutics for COVID-19 Using Deep Learning. Journal of Proteome Research 19, 4624–4636. 10.1021/acs.jproteome.0c00316. [DOI] [PubMed] [Google Scholar]
  • 83.Xu J, Song W, Xu Z, Danziger MM, Karavani E, Zang C, Chen X, Li Y, Paz IMR, Gohel D, et al. (2025). Single-microglia transcriptomic transition network-based prediction and real-world patient data validation identifies ketorolac as a repurposable drug for Alzheimer’s disease. Alzheimers Dement 21, e14373. 10.1002/alz.14373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Venugopalan J, Tong L, Hassanzadeh HR, and Wang MD (2021). Multimodal deep learning models for early detection of Alzheimer’s disease stage. Sci Rep 11, 3254. 10.1038/s41598-020-74399-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Koutsouleris N, Dwyer DB, Degenhardt F, Maj C, Urquijo-Castro MF, Sanfelici R, Popovic D, Oeztuerk O, Haas SS, Weiske J, et al. (2021). Multimodal Machine Learning Workflows for Prediction of Psychosis in Patients With Clinical High-Risk Syndromes and Recent-Onset Depression. JAMA Psychiatry 78, 195–209. 10.1001/jamapsychiatry.2020.3604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Zhavoronkov A, Ivanenkov YA, Aliper A, Veselov MS, Aladinskiy VA, Aladinskaya AV, Terentiev VA, Polykovskiy DA, Kuznetsov MD, Asadulaev A, et al. (2019). Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat Biotechnol 37, 1038–1040. 10.1038/s41587-019-0224-x. [DOI] [PubMed] [Google Scholar]
  • 87.Sun M, Zhao S, Gilvary C, Elemento O, Zhou J, and Wang F (2020). Graph convolutional networks for computational drug development and discovery. Brief Bioinform 21, 919–935. 10.1093/bib/bbz042. [DOI] [PubMed] [Google Scholar]
  • 88.Born J, and Manica M (2021). Trends in Deep Learning for Property-driven Drug Design. Curr Med Chem. 10.2174/0929867328666210729115728. [DOI] [PubMed] [Google Scholar]
  • 89.Cummings J, Zhou Y, Lee G, Zhong K, Fonseca J, and Cheng F (2024). Alzheimer’s disease drug development pipeline: 2024. Alzheimer’s & Dementia: Translational Research & Clinical Interventions 10, e12465. 10.1002/trc2.12465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Rodriguez S, Hug C, Todorov P, Moret N, Boswell SA, Evans K, Zhou G, Johnson NT, Hyman BT, Sorger PK, et al. (2021). Machine learning identifies candidates for drug repurposing in Alzheimer’s disease. Nat Commun 12, 1033. 10.1038/s41467-021-21330-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Zhou Y, Fang J, Bekris LM, Kim YH, Pieper AA, Leverenz JB, Cummings J, and Cheng F (2021). AlzGPS: a genome-wide positioning systems platform to catalyze multi-omics for Alzheimer’s drug discovery. Alzheimers Res Ther 13, 24. 10.1186/s13195-020-00760-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Xu J, Mao C, Hou Y, Luo Y, Binder JL, Zhou Y, Bekris LM, Shin J, Hu M, Wang F, et al. (2022). Interpretable deep learning translation of GWAS and multi-omics findings to identify pathobiology and drug repurposing in Alzheimer’s disease. Cell Reports 41. 10.1016/j.celrep.2022.111717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Xu J, Zhang P, Huang Y, Zhou Y, Hou Y, Bekris L, Lathia J, Chiang CW, Li L, Pieper A, et al. (2021). Multimodal single-cell/nucleus RNA sequencing data analysis uncovers molecular networks between disease-associated microglia and astrocytes with implications for drug repurposing in Alzheimer’s disease. Genome Res 31, 1900–1912. 10.1101/gr.272484.120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Shin MK, Vazquez-Rosa E, Koh Y, Dhar M, Chaubey K, Cintron-Perez CJ, Barker S, Miller E, Franke K, Noterman MF, et al. (2021). Reducing acetylated tau is neuroprotective in brain injury. Cell 184, 2715–2732 e2723. 10.1016/j.cell.2021.03.032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Min SW, Chen X, Tracy TE, Li Y, Zhou Y, Wang C, Shirakawa K, Minami SS, Defensor E, Mok SA, et al. (2015). Critical role of acetylation in tau-mediated neurodegeneration and cognitive deficits. Nat Med 21, 1154–1162. 10.1038/nm.3951. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Taubes A, Nova P, Zalocusky AK, Kosti I, Bicak M, Zilberter YM, Hao Y, Yoon S, Oskotsky t., Pineda S, Chen B, Jones AE, Choudhary K, Grone B, Balestra EM, Chaudhry F, Paranjpe I, Freitas J, Koutsodendris N, Chen N, Wang C, Chang W, An A, Glicksberg SB, Sirota M, Huang Y (2021). Experimental and real-world evidence supporting the computational repurposing of bumetanide for APOE4-related Alzheimer’s disease. Nature Aging 18, 932–947. 10.1038/s43587-021-00122-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Yahya EB, and Alqadhi AM (2021). Recent trends in cancer therapy: A review on the current state of gene delivery. Life Sciences 269, 119087. 10.1016/j.lfs.2021.119087. [DOI] [PubMed] [Google Scholar]
  • 98.Siegel RL, Giaquinto AN, and Jemal A (2024). Cancer statistics, 2024. CA: A Cancer Journal for Clinicians 74, 12–49. 10.3322/caac.21820. [DOI] [PubMed] [Google Scholar]
  • 99.Tanoli Z, Vähä-Koskela M, and Aittokallio T (2021). Artificial intelligence, machine learning, and drug repurposing in cancer. Expert Opinion on Drug Discovery 16, 977–989. 10.1080/17460441.2021.1883585. [DOI] [PubMed] [Google Scholar]
  • 100.Xu H, Aldrich MC, Chen Q, Liu H, Peterson NB, Dai Q, Levy M, Shah A, Han X, Ruan X, et al. (2015). Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortality. Journal of the American Medical Informatics Association 22, 179–191. 10.1136/amiajnl-2014-002649. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Velavan TP, and Meyer CG (2020). The COVID-19 epidemic. Trop Med Int Health 25, 278–280. 10.1111/tmi.13383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Zhou Y, Wang F, Tang J, Nussinov R, and Cheng F (2020). Artificial intelligence in COVID-19 drug repurposing. The Lancet Digital Health 2, e667–e676. 10.1016/S2589-7500(20)30192-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Zhou Y, Hou Y, Shen J, Huang Y, Martin W, and Cheng F (2020). Network-based drug repurposing for novel coronavirus 2019-nCoV/SARS-CoV-2. Cell Discovery 6, 14. 10.1038/s41421-020-0153-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Zhou Y, Hou Y, Shen J, Mehra R, Kallianpur A, Culver DA, Gack MU, Farha S, Zein J, Comhair S, et al. (2020). A network medicine approach to investigation and population-based validation of disease manifestations and drug repurposing for COVID-19. PLOS Biology 18, e3000970. 10.1371/journal.pbio.3000970. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Martin WR, and Cheng F (2020). Repurposing of FDA-Approved Toremifene to Treat COVID-19 by Blocking the Spike Glycoprotein and NSP14 of SARS-CoV-2. Journal of Proteome Research 19, 4670–4677. 10.1021/acs.jproteome.0c00397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Nabel Elizabeth G Cardiovascular Disease. New England Journal of Medicine 349, 60–72. 10.1056/NEJMra035098. [DOI] [PubMed] [Google Scholar]
  • 107.Paci P, Fiscon G, Conte F, Wang R-S, Handy DE, Farina L, and Loscalzo J (2022). Comprehensive network medicine-based drug repositioning via integration of therapeutic efficacy and side effects. npj Systems Biology and Applications 8, 12. 10.1038/s41540-022-00221-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Jiménez ÁB, Lázaro JL, and Dorronsoro JR (2007). Finding Optimal Model Parameters by Discrete Grid Search. In Innovations in Hybrid Intelligent Systems, Corchado E, Corchado JM, and Abraham A, eds. (Springer Berlin; Heidelberg: ), pp. 120–127. 10.1007/978-3-540-74972-1_17. [DOI] [Google Scholar]
  • 109.He X, Zhao K, and Chu X (2021). AutoML: A survey of the state-of-the-art. Knowledge-Based Systems 212, 106622. 10.1016/j.knosys.2020.106622. [DOI] [Google Scholar]
  • 110.Sandmann S, Riepenhausen S, Plagwitz L, and Varghese J (2024). Systematic analysis of ChatGPT, Google search and Llama 2 for clinical decision support tasks. Nature Communications 15, 2050. 10.1038/s41467-024-46411-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Zhang Y, Li Y, Cui L, Cai D, Liu L, Fu T, Huang X, Zhao E, Zhang Y, and Chen Y (2023). Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv preprint arXiv:2309.01219 2. [Google Scholar]
  • 112.Swanson K, Wu W, Bulaong NL, Pak JE, and Zou J (2024). The virtual lab: AI agents design new SARS-CoV-2 nanobodies with experimental validation. bioRxiv, 2024.2011. 2011.623004. [Google Scholar]
  • 113.Su X, Wang Y, Gao S, Liu X, Giunchiglia V, Clevert D-A, and Zitnik M (2024). Knowledge Graph Based Agent for Complex, Knowledge-Intensive QA in Medicine. arXiv preprint arXiv:2410.04660. [Google Scholar]
  • 114.Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, Ronneberger O, Willmore L, Ballard AJ, and Bambrick J (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Nebel RA, Aggarwal NT, Barnes LL, Gallagher A, Goldstein JM, Kantarci K, Mallampalli MP, Mormino EC, Scott L, Yu WH, et al. (2018). Understanding the impact of sex and gender in Alzheimer’s disease: A call to action. Alzheimers Dement 14, 1171–1183. 10.1016/j.jalz.2018.04.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Mehta KM, Yaffe K, Perez-Stable EJ, Stewart A, Barnes D, Kurland BF, and Miller BL (2008). Race/ethnic differences in AD survival in US Alzheimer’s Disease Centers. Neurology 70, 1163–1170. 10.1212/01.wnl.0000285287.99923.3c. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Raman R, Quiroz YT, Langford O, Choi J, Ritchie M, Baumgartner M, Rentz D, Aggarwal NT, Aisen P, Sperling R, and Grill JD (2021). Disparities by Race and Ethnicity Among Adults Recruited for a Preclinical Alzheimer Disease Trial. JAMA Netw Open 4, e2114364. 10.1001/jamanetworkopen.2021.14364. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Braunstein ML (2018). Healthcare in the Age of Interoperability: The Promise of Fast Healthcare Interoperability Resources. IEEE Pulse 9, 24–27. 10.1109/MPUL.2018.2869317. [DOI] [PubMed] [Google Scholar]
  • 119.Xu J, Glicksberg BS, Su C, Walker P, Bian J, and Wang F (2020). Federated Learning for Healthcare Informatics. J Healthc Inform Res, 1–19. 10.1007/s41666-020-00082-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Wang F, Kaushal R, and Khullar D (2020). Should Health Care Demand Interpretable Artificial Intelligence or Accept “Black Box” Medicine? Ann Intern Med 172, 59–60. 10.7326/M19-2548. [DOI] [PubMed] [Google Scholar]
  • 121.Martin W, Sheynkman G, Lightstone FC, Nussinov R, and Cheng F (2021). Interpretable artificial intelligence and exascale molecular dynamics simulations to reveal kinetics: Applications to Alzheimer’s disease. Curr Opin Struct Biol 72, 103–113. 10.1016/j.sbi.2021.09.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Alkhulaifi A, Alsahli F, and Ahmad I (2021). Knowledge distillation in deep learning and its applications. PeerJ Comput Sci 7, e474. 10.7717/peerj-cs.474. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Piñero J, Bravo À, Queralt-Rosinach N, Gutiérrez-Sacristán A, Deu-Pons J, Centeno E, García-García J, Sanz F, and Furlong LI (2017). DisGeNET: a comprehensive platform integrating information on human disease-associated genes and variants. Nucleic Acids Research 45, D833–D839. 10.1093/nar/gkw943. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Stelzer G, Rosen N, Plaschkes I, Zimmerman S, Twik M, Fishilevich S, Stein TI, Nudel R, Lieder I, Mazor Y, et al. (2016). The GeneCards Suite: From Gene Data Mining to Disease Genome Sequence Analyses. Current Protocols in Bioinformatics 54, 1.30.31–31.30.33. 10.1002/cpbi.5. [DOI] [PubMed] [Google Scholar]
  • 125.Ochoa D, Hercules A, Carmona M, Suveges D, Gonzalez-Uriarte A, Malangone C, Miranda A, Fumis L, Carvalho-Silva D, Spitzer M, et al. (2021). Open Targets Platform: supporting systematic drug–target identification and prioritisation. Nucleic Acids Research 49, D1302–D1310. 10.1093/nar/gkaa1027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Howe KL, Achuthan P, Allen J, Allen J, Alvarez-Jarreta J, Amode MR, Armean IM, Azov AG, Bennett R, Bhai J, et al. (2021). Ensembl 2021. Nucleic Acids Research 49, D884–D891. 10.1093/nar/gkaa942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Bank PD (1971). Protein data bank. Nature New Biol 233, 10–1038. [Google Scholar]
  • 128.Meyer MJ, Beltran JF, Liang S, Fragoza R, Rumack A, Liang J, Wei X, and Yu H (2018). Interactome INSIDER: a structural interactome browser for genomic studies. Nat Methods 15, 107–114. 10.1038/nmeth.4540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, et al. (2018). The Reactome Pathway Knowledgebase. Nucleic Acids Research 46, D649–D655. 10.1093/nar/gkx1132. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, et al. (2019). STRING v11: protein-protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Res 47, D607–D613. 10.1093/nar/gky1131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Zeng X, Zhu S, Liu X, Zhou Y, Nussinov R, and Cheng F (2019). deepDR: a network-based deep learning approach to in silico drug repositioning. Bioinformatics 35, 5191–5198. 10.1093/bioinformatics/btz418. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Yang Y, Qiu Y, Hu J, Rosen-Zvi M, Guan Q, and Cheng F (2024). A deep learning framework combining molecular image and protein structural representations identifies candidate drugs for pain. Cell Reports Methods 4. 10.1016/j.crmeth.2024.100865. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Wang Y, Wang J, Cao Z, and Barati Farimani A (2022). Molecular contrastive learning of representations via graph neural networks. Nature Machine Intelligence 4, 279–287. 10.1038/s42256-022-00447-x. [DOI] [Google Scholar]

RESOURCES