Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Apr 12;16:17018. doi: 10.1038/s41598-026-47765-3

Advancing target discovery through disease-specific integration of multi-modal target identification models and comprehensive benchmarking system

Howell Leung 1,#, Chengchen Duan 1,#, Wenhao Gou 1,#, Jianjiu Chen 1, Ying Xin 1, Zetian Zheng 1, Vladimir Naumov 2, David Gennert 3, Man Zhang 4, Alex Aliper 2, Feng Ren 1,4, Evgeny Izumchenko 5, Frank W Pun 1,, Alex Zhavoronkov 1,2,3,4,
PMCID: PMC13230693  PMID: 41968138

Abstract

Target identification is crucial for drug development. AI-driven approaches leveraging multi-omics and computational modeling can accelerate this process. However, integrating multi-modal data for disease-specific target identification and predicting translational potential remains challenging. Moreover, the absence of a systematic evaluation framework for model performance limits confidence in target reliability. This study presents a unified framework combining machine learning-based target identification with comprehensive benchmarking. We first developed Target Identification Pro (TargetPro), a disease-specific model spanning 38 diseases across oncology, metabolic, immune, fibrotic, and neurological categories. TargetPro shows strong predictive performance for clinical-stage targets and reveals disease-specific patterns, underscoring the need for tailored target detection models. We next created Target Identification Benchmark (TargetBench 1.0) to assess target identification systems, including large language models, based on their ability to recover established targets and find high-quality novel candidates. This integrated approach offers a streamlined strategy to evaluate target discovery models, ultimately improving drug development efficiency.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-47765-3.

Subject terms: Cancer, Computational biology and bioinformatics, Drug discovery

Introduction

Drug development is hindered by high costs, inherent challenges, and a striking failure rate, as up to 90% of candidates entering clinical trials do not achieve regulatory approval13. Many of these failures can be traced to the earliest stages of drug discovery, particularly the selection of biological targets that later prove to be less effective or more toxic than anticipated4. Given that the total cost of developing a single new drug may reach several billion US dollars, early identification of the most promising targets and de-risking their development pathway are critical for the pharmaceutical industry and, ultimately, for public health4.

A notable effort for optimizing target identification is AstraZeneca’s five ‘R’s framework, comprising the right target, right patient, right tissue, right safety, and right commercial potential5. Introduced after an analysis of drug development projects conducted between 2005 and 2010, this framework places the ‘right target’ as its cornerstone, emphasizing targets with strong mechanistic links to disease and supported by robust biological and genetic evidence. Implementation of the 5R framework increased the success rate of early development (from candidate drug nomination to phase III trial completion) from 4% in 2005–2010 to 19% in 2012–20166.

In parallel, the emergence of artificial intelligence (AI)-driven platforms has revolutionized target identification. Traditional methods, often relying on expert experience and limited datasets, suffer from inefficiencies and high attrition rates7. The AI-assisted target discovery approach uses machine learning and bioinformatics to integrate and analyze vast datasets encompassing genomics, proteomics, and literature, providing a systematic, evidence-based approach to target prioritization8. Examples of AI-driven target identification platforms include ConVERGE by Verge Genomics (See Related Link 1), TargetMATCH by Owkin (See Related Link 2), Open Targets (developed through a partnership among the European Bioinformatics Institute, the Wellcome Sanger Institute, and GlaxoSmithKline)911, and PandaOmics by Insilico Medicine12. Notably, several targets prioritized by PandaOmics, which integrates multiple omics and text scores, have been validated in both preclinical and clinical studies1316. These platforms underscore AI’s ability to bridge data-driven hypothesis generation and clinical translation, enhancing decision-making in drug discovery. Additionally, Large Language Models (LLMs) are being increasingly used in target selection. By analyzing vast amounts of scientific literature, LLMs can uncover hidden connections between genes, diseases, and targets. For example, BioGPT has been applied to discover potential dual-purpose targets that are relevant to both aging and age-related diseases, identifying novel anti-aging targets17. ChatGPT is also being explored for predicting protein domains, uncovering drug-binding pockets, and assisting in structure prediction18.

Despite these advances, two fundamental gaps continue to limit the full potential of AI in drug discovery. The first relates to the challenges of using multi-model platforms for disease-specific target identification and objective-driven prediction. Given the heterogeneous mechanisms underlying different diseases, it is unrealistic to expect a single model to provide a one-size-fits-all strategy for target identification across all disease contexts. However, for end users, selecting the optimal combination of target identification models is challenging due to the complexities of the underlying data sources and opaque assumptions inherent in AI-based approaches. This highlights the need for a validated, off-the-shelf framework for selecting and weighting multiple models tailored to specific disease contexts. In addition, current target identification systems lack the ability to predict a target’s potential to advance to clinical development, an objective that is important to end users. Developing machine learning models that can generate such predictions could assist human experts in making high-stakes decisions about which targets to prioritize for further development.

The second critical gap is the lack of standardized evaluation methods for these predictive tools. The rapid and heterogeneous proliferation of AI platforms has outpaced the development of common systems to evaluate their performance and reliability, especially for a specific task area. This ‘benchmarking gap’ slows scientific progress and industrial adoption, as it becomes difficult to compare different models, understand their relative strengths, or build the necessary confidence to integrate their outputs into multi-billion-dollar R&D decisions. Existing benchmarks, such as BETA, have been developed to address specific computational tasks like drug-target interaction prediction, but they may not fully capture the multifaceted biological and clinical considerations essential for selecting a viable therapeutic target for development19. Above all, a dependable and transparent evaluation system is essential for building trust and ensuring the responsible integration of AI into modern medicine.

This study introduces a framework designed to address these gaps by integrating a machine learning workflow, PandaOmics Target Identification Pro (TargetPro), along with a dedicated evaluation system, Target Identification Benchmark (TargetBench 1.0). Together, these components form a strategy to improve the accuracy of target discovery and validate the reliability of target predictions (Fig. 1).

Fig. 1.

Fig. 1

Overview of TargetBench 1.0 for drug target identification models benchmarking.

First, we developed TargetPro by implementing a disease-specific, objective-driven machine learning workflow designed to effectively utilize the multidimensional omics and text data derived from the PandaOmics platform12. TargetPro was trained using a meticulously curated set of targets that have entered the clinical stages (defined as Phase 1, Phase 2, Phase 3, and Launched) for a given disease. By using these targets as training labels, TargetPro learns the complex patterns and features associated with successful translation, enabling it to predict targets with a higher probability of clinical-stage associations. Second, we established TargetBench 1.0, a system designed to systematically evaluate the performance of any target discovery model and platform. This benchmark aligns with the practical goal of identifying targets with the potential to advance to therapeutic application, providing a clear and relevant measure of predictive power.

We applied this framework to 38 representative diseases, selected based on their prevalence, incidence, and data availability, spanning oncology20, immune-related21, metabolic22, fibrotic-related23, and neurological disorders24. TargetPro generated 38 disease-specific target lists, prioritizing those with the highest potential for clinical advancement. These lists were then compared against outputs from other target identification models (including LLMs) using TargetBench 1.0, which provides a comprehensive assessment across multiple dimensions, including precision, structural data availability, druggability potential, repurposing opportunities, biological relevance, bioassay accessibility, and gene modulator availability. Collectively, this integrated approach enhances the accuracy of target discovery, establishes confidence in computational predictions, and ultimately accelerates the development of novel therapies.

Materials and methods

Data collection and preprocessing

Representative disease selection

We selected a comprehensive set of 38 diseases to represent a broad range of pathological mechanisms and therapeutic areas, guided by their prevalence, incidence, and data availability. These diseases were manually curated and systematically categorized into five major pathophysiological categories: oncology, metabolic disorder, immune-related disease, fibrotic-related disease, and neurological disease. This diverse selection enables robust and generalizable comparison of drug target characteristics across distinct biological contexts. The specific diseases included in each category are summarized in Table 1.

Table 1.

Summary of the 38 diseases.

Disease group Diseases # of clinical targets1 for
a disease (range)
Class imbalance ratio2 for a disease (range)

Oncology

(N = 18)

Non-small cell lung carcinoma, breast cancer, prostate cancer, colorectal cancer, thyroid cancer, gastric cancer, bladder cancer, endometrial cancer, cervical cancer, acute leukemia, melanoma, renal cancer, liver cancer, ovarian cancer, myeloma, pancreatic cancer, chronic myelogenous leukemia, head and neck cancer 92–453 42–209

Metabolic Disorder

(N = 5)

Obesity, hypertension, Type 2 diabetes mellitus, hyperlipidemia, chronic kidney disease 71–180 106–271

Immune-related Disease

(N = 6)

Inflammatory bowel disease, Rheumatoid arthritis, Systemic lupus erythematosus, psoriasis, Type I Diabetes Mellitus, scleroderma 75–257 74–256

Fibrotic-related Disease

(N = 3)

Idiopathic pulmonary fibrosis, atherosclerosis, osteoarthritis 98–165 116–196

Neurological Disease

(N = 6)

Alzheimer’s disease, Parkinson’s disease, Amyotrophic lateral sclerosis, stroke, epilepsy, Huntington’s disease) 54–222 86–356

1Clinical targets refer to targets with drugs in clinical or launched stages. 2Class imbalance ratio is calculated as the ratio of preclinical targets (negative class) to clinical targets (positive class). For example, breast cancer has 453 drug targets within 19,291 protein coding genes, the ratio is (19,291–453)/453 = 42.

For each of the 38 diseases, omics datasets were obtained from PandaOmics for analysis, an established target identification platform12 that implements a standardized pipeline to automatically retrieve and normalize omics data from diverse publicly available repositories (including TCGA, GEO, ArrayExpress, and Proteomic Identification Database [PRIDE]). All datasets undergo manual curation to ensure accurate annotation of disease and control samples, thereby supporting the reliability of downstream analyses. For the 38 selected diseases, PandaOmics provides 631 pre-validated, human-curated ‘golden datasets’ (spanning microarray, RNA-Seq, and proteomics) involving 737 disease-control comparisons. To enhance the robustness of TargetPro score calculation, we compiled a manually curated set of high-quality, publicly available multi-omics human data, adding 137 datasets and 146 comparisons across the same 38 selected diseases. These datasets were selected based on clear clinical annotations comparing healthy human samples (controls) with untreated patient groups (cases) and prioritized newly published work (2024 onward) that is not already covered in the platform’s default data, thereby establishing an up-to-date evaluation environment without compromising data quality (all dataset details provided in Supplementary Data 1). The ‘golden’ together with these additional datasets are collectively referred to as ‘golden-plus datasets’ below.

Clinical-stage targets annotation

For each disease, the analysis encompassed 19,291 protein-coding genes annotated in the HGNC database (See Related Link 3). A manually curated disease-drug-target list was compiled from the latest public sources (e.g., ClinicalTrials.gov) and was subsequently reviewed and verified by our internal team of biologists. This manually curated list was used to assign each target to a specific stage of drug development (Preclinical, Phase 1, Phase 2, Phase 3, and Launched). Preclinical targets are defined as all protein-coding genes, excluding those in Phase 1–3 clinical trials or currently launched. Drugs with unspecified targets or without a clear mechanism of action were excluded. If multiple drugs targeted the same protein, the protein was assigned the highest development stage among its associated drugs.

Target identification model scores

We computed 22 target identification model scores for 733,058 gene-disease pairs (38 disease × 19,291 protein-coding genes) using PandaOmics algorithm, consisting of 12 omic-based and 10 text-driven measures. The omic-based scores include expression, pathways, causal inference, network neighbors, interactome community, knockouts, disease submodules, overexpression, matrix factorization, heterogeneous graph walk, mutated submodules, and mutations, while the text-driven scores include attention score, credible attention index, grant funding, mean hirsch, impact factor, evidence, grant size, attention spike, funding per publication, and trend. All generated scores were used as input features for machine learning.

Machine learning

Model training for TargetPro

To identify the most effective model for therapeutic target prioritization, we benchmarked five machine learning approaches for high-dimensional tabular data: CatBoost25, Elastic Net26, LightGBM27, Random Forest28, and XGBoost29. Models were trained and evaluated on 38 disease-specific datasets, covering 19,291 protein-coding genes, using 22 target-identification model scores as input features. Performance was accessed using Area Under the Precision-Recall Curve (AUPRC), F1-score, and Precision at top K (Supplementary Data 2 and 3). XGBoost was selected as the final model architecture based on its superior balance of precision and recall.

We adopted a Cost-Sensitive Learning strategy30,31 to address a fundamental challenge in drug discovery, where definitive labels (i.e., targets confirmed to be irrelevant to a disease) are often unattainable. In our setup, targets with clinically developed drugs were defined as the ‘positive’ class, while all remaining (unlabeled) targets, representing a mixture of true negatives and currently undiscovered positives, were treated as the ‘negative’ class.

To mitigate the bias introduced by potential latent positives in the negative class and to address the significant class imbalance, we applied an inverse class-weighting scheme that prioritizes correct classification of sparse positive cases. This approach aligns with the sparsity assumption in drug discovery32, where true targets are rare, and the majority of the proteome represents a negative set. Class weights (w) were calculated inversely proportional to their frequencies:

graphic file with name d33e437.gif
graphic file with name d33e440.gif

Where Inline graphic and Inline graphicare weights to positive and negative class; Inline graphic and Inline graphic are numbers of observations in positive and negative classes, respectively.

A distinct model was trained for each indication, recognizing that different biological patterns govern target success in different therapeutic areas. Model training and hyperparameter tuning for each of these disease-specific models were conducted using a nested 5-fold cross-validation approach (Fig. 3a). The entire dataset was first partitioned into five equal folds. The main process was then repeated five times, with each iteration using four folds for model training and holding out the remaining fold as a final, unseen test set. Within each training repetition, a separate inner cross-validation loop was performed exclusively on the training data for hyperparameter tuning. The optimal hyperparameters were subsequently used to train a single model on the entire training set, which then made predictions on the held-out test fold. Finally, the out-of-sample predictions from all five folds were combined to assess the model’s overall performance using the Area Under the Precision-Recall Curve (AUPRC), a metric that is particularly well-suited for evaluating classifiers on imbalanced datasets.

Fig. 3.

Fig. 3

(a) Workflow for TargetPro model training. The process begins with a manually curated ‘golden dataset’, representing the most relevant and available data for a given disease. This dataset is partitioned into five stratified splits based on whether targets have associated drugs in clinical trials. A nested 5-fold cross-validation is then employed, where an inner loop performs hyperparameter tuning with class upweighting. A TargetPro model is then trained with the optimal parameters to generate predictions on held-out testing data. Finally, the out-of-sample predictions from all five folds are combined to evaluate the overall generalizability of the model. (b) Performance comparison of the integrated TargetPro model against each of the 22 individual omics and text models.

Feature importance analysis using SHAP framework

We implemented the Shapley Additive Explanations (SHAP) framework33 to provide an interpretable and consistent feature importance analysis. Global feature importance was determined by calculating the mean absolute SHAP value for each feature, representing the average magnitude of its contribution to the model’s predictions across the entire dataset. To facilitate a higher-level analysis, individual features were subsequently aggregated into three categories based on their data source and nature: Text, Omics (Static), and Omics (Dynamic). This allows the assessment of the relative influence of different data modalities on model performance. The Text category included features derived from publication and funding metadata, such as attention, credible attention index, grant funding, mean hirsch, impact factor, evidence, grant size, attention spike, funding per publication, and trend. The Omics (Static) category comprises features whose values are calculated independently of the golden and user’s proprietary datasets. These features, representing stable biological states and pre-computed network properties, included matrix factorization, heterogeneous graph walk, mutated submodules, and mutations. In contrast, the Omics (Dynamic) category consists of features whose scores are directly affected by the specific input of omics data for an analysis. These features capture dynamic or latent biological signals and include expression, pathways, causal inference, network neighbors, interactome community, knockouts, overexpression, and disease submodules.

External validation

To assess model generalizability, we evaluated the models trained on the internal “golden” dataset using our “plus” datasets, which were independent and excluded from model training and optimization. To ensure a robust evaluation, we filtered the “plus” cohort to include only diseases with ≥ 5 independent datasets (Supplementary Data 1). This yielded 12 diseases: Acute Leukemia, Alzheimer’s disease, Bladder Cancer, Cervical Cancer, Chronic Kidney Disease, Head and Neck Cancer, Huntington’s Disease, Hyperlipidemia, Obesity, Ovarian Cancer, Prostate Cancer, and Stroke. Performance on this unseen data was evaluated using AUPRC.

Benchmarking

Setup for model benchmarking

To benchmark different target identification models, we conducted a comparative evaluation of the therapeutic target lists generated by each model. To ensure a fair and meaningful comparison, the experimental setup for each model was tailored to mimic its typical real-world application and maximize its performance. For the LLM evaluation, a standardized prompt (see Prompt for generating therapeutic targets) was used to query each model, including BioGPT, Grok3, Grok4, DeepSeek-R1, Claude-Opus-4, GPT4.1, o3, and GPT5 (Supplementary Data 4), for a list of targets for each of the 38 diseases. To account for the stochastic nature of these models, the process was repeated five times per disease. We prompted LLM to nominate a defined number of genes per iteration and then aggregated results across all runs for the final analysis. BioGPT was run and queried locally on a MacBook Pro (M4), the Grok series models were queried through their official application programming interface (API), and all other LLMs were queried via the Azure API. For Open Targets, we retrieved disease-specific target lists utilizing their default ranking, which integrates genetic, omics, and text-based evidence. The evaluation protocol for TargetPro was specifically designed to simulate a user performing a new meta-analysis by integrating additional, user-selected datasets with the foundational ‘golden dataset’. For each disease, we augmented the baseline ‘golden dataset’ with additional disease-specific omics data, creating an enhanced ‘golden-plus dataset’. TargetPro then generated new target prediction scores using the combined dataset. See details of the datasets and prediction process in Supplementary Data 1 and Supplementary Fig. 1.

Prompt for generating therapeutic targets

To elicit a ranked list of potential therapeutic targets for each disease, the following standardized prompt was provided to each LLM, with {disease_name} and {top_k} parameters adjusted for each query:

‘For {disease_name}, generate a complete list of {top_k} high-confidence drug targets, where each target is a gene name. High-confidence targets are defined as: Genes with drugs in clinical-stage development (e.g., approved or in trials) for {disease_name}, AND/OR preclinical targets supported by substantial evidence of their role in {disease_name} and their potential as drug targets. You should rank these {top_k} targets by confidence level.

Your response should be a json file ONLY. The JSON file should have {top_k} elements only, with each element only storing 1 gene name, ordered by their rank. Ignore the alternative name of each gene in your output. Ignore any unrelated content.’

Assessment of robust target identification models

From the perspective of target biologists and drug hunters, confidence in a model’s output is initially established by the presence of well-validated targets34. Accordingly, we use precision at top K as the primary performance metric, defined as the proportion of known clinical targets, namely targets with drugs in clinical development, among the top K targets prioritized by the model per disease.

graphic file with name d33e500.gif

Here, K represents the number of model-prioritized targets for a disease, and is set to the number of known clinical targets for a disease (Supplementary Data 2).

However, we acknowledge the inherent paradox in drug discovery: while retrieving known targets validates a model’s accuracy, the discovery of promising novel targets often holds greater translational value. Therefore, a comprehensive evaluation must be extended to assess the intrinsic quality of all identified targets, whether known or novel. This ensures that the model is not only accurate but also capable of generating candidates with a high probability of success in the development pipeline.

Beyond precision at top K, we also evaluated the quality of each model’s novel predictions. We focused on novel targets, defined as those ranked within top K but not yet in clinical development, as this reflects the practical need in reality to prioritize a limited number of high-potential candidates. These novel targets were assessed using a suite of metrics designed to quantify key characteristics relevant for drug development. We established a multi-faceted set of criteria, sourced from expert-curated databases, to ensure the robustness and clinical relevance of the identified targets. These criteria assess targets across several critical dimensions, including their clinical development stage, scientific validation, and safety profiles derived from clinical trial data and biological essentiality assessments.

First, the availability of protein crystal structures was confirmed through the Protein Data Bank, as this information is crucial for facilitating subsequent drug design to improve potency and selectivity. Next, a target’s druggability was systematically evaluated using information from multiple resources (Citeline, DGIdb & DrugBank). Gene targets with drug records are considered druggable, prioritizing those with a high likelihood of successful pharmacological modulation to ensure chemical and biological feasibility. To assess therapeutic adaptability and de-risk development, we identified targets with approved drugs in other indications by referencing comprehensive drug development and clinical trial databases, leveraging their established safety and efficacy profiles. To connect targets to established disease-specific biological mechanisms, their biological relevance was measured by the number of pathways they share with existing clinical-stage targets, based on data from Reactome. The level of experimental validation for each potential target was quantified by the number of available bioassays cataloged in PubChem, with a higher number of associated assays indicating a more robust evidence base and greater development potential. Finally, we used MedChemExpress (See Related Link 4) to obtain the number of available gene modulators, which included small molecules, peptides, and monoclonal antibodies, but excluded siRNA.

Use of AI in manuscript preparation

During the preparation of this manuscript, AI tools were utilized to assist with writing and editing. The role of these AI tools was limited to proofreading for grammatical errors, improving sentence structure, and enhancing overall clarity and readability. The authors carefully reviewed and revised all AI-generated suggestions to ensure scientific accuracy and retain full responsibility for the final content of this publication.

Results

Characteristics predictive features

We applied our target identification and benchmarking framework to 766 curated foundational datasets spanning 38 diseases across five major therapeutic areas: oncology, metabolic, immune-related, fibrotic-related, and neurological diseases (Supplementary Data 1). Among the 38 selected diseases, considerable heterogeneity was observed in both the number of clinical-stage targets (used as positive labels in model training) and in the ratio of clinical-stage to preclinical targets (used as negative labels and defined as all protein-coding genes excluding those in Phase 1–3 clinical trials or currently launched) (Table 1). The oncology group exhibited the highest number of clinical-stage targets (ranging from 92 to 453), whereas other areas, such as metabolic diseases, showed fewer targets (ranging from 71 to 180). Class imbalance between positive and negative labels also varied substantially, from 1:42 in oncology to 1:356 in neurological diseases. These observations underscore the sparsity of validated targets among certain disease areas and the striking variability in target landscapes across different pathologies.

We next analyzed 22 distinct feature scores that capture target-disease associations based on omics data, scientific literature, clinical records, and grants (see Materials & Methods for details). Across the different stages of clinical development, most scores increased progressively from the ‘preclinical’ stage into Phase 1, and then plateaued (Figs. 2a,b). Grouping targets into ‘preclinical’ and ‘clinical’ categories revealed a clear separation between these stages (Fig. 2c, Wilcoxon rank-sum test, two-tailed, FDR-adjusted P-values < 0.05), indicating that the scores capture information relevant to entry into clinical development.

Fig. 2.

Fig. 2

Omics and text scores across clinical development stages. All plots are generated using the combined scores from all 38 diseases to provide a global overview. (ab) Trend plots of the 22 omics and text scores across distinct stages of clinical development, from ‘preclinical’ to ‘launched’. (c) Bar plots comparing the mean of scores for targets grouped as ‘preclinical’ versus ‘clinical’ (Phase 1 and beyond). All models exhibit significant differences (Wilcoxon rank-sum test, two-tailed, FDR-adjusted P-values < 0.05).

Workflow implementation and TargetPro development

Given the distinct feature patterns differentiating preclinical and clinical-stage targets, we hypothesized that a machine learning model could be trained to identify targets with a higher likelihood of advancing to the clinical stage. The model assumes that targets with a high predictive score (TargetPro score) share key characteristics with known clinical-stage targets and therefore have a high probability of clinical-stage association. To construct this model, we implemented a systematic machine learning workflow that integrates 22 omics and text scores and uses developmental stage labels assigned to each gene as ground truth for supervised learning (Fig. 3a; see Materials and Methods for full details).

With this foundation, we trained a distinct TargetPro model for each therapeutic area using an XGBoost algorithm29. To optimize performance and account for the inherent class imbalance between clinical-stage and preclinical targets, we incorporated an upweighting strategy and conducted disease-specific hyperparameter tuning. This disease-specific design is a core feature of our workflow, recognizing that distinct biological patterns govern target progression in different therapeutic areas. The framework’s flexibility enables it to capture these nuances while maintaining rigorous and consistent methodological standards. To support interpretability, the model was configured to produce weights for each input score, allowing examination of the features driving predictions and comparison of their contributions across therapeutic areas.

To ensure a robust and unbiased assessment, model performance was evaluated using a nested 5-fold cross-validation strategy, yielding predictions for all genes. TargetPro achieved a significantly higher Area Under the Precision-Recall Curve (AUPRC) than any single omics and text score used as a baseline (AUPRCs: 0.295 vs. 0.015–0.22, based on paired T-tests at the disease level with FDR-adjusted P-values < 0.05, Fig. 3b). Given the low positive class prevalence of 1%, the baseline AUPRC expected from a random guess is 0.0135. Crucially, this performance generalized to an independent external validation cohort of 12 diseases (average AUPRC: 0.269, Supplementary Data 3). Accordingly, TargetPro demonstrates superior predictive power within the clinically relevant precision-recall space.

Disease-specific patterns of feature importance

To assess model drivers, we analyzed feature importance across the five disease groups (Fig. 4). At the individual feature level, matrix factorization and attention score were universally the most impactful, highlighting their core contribution to the trained disease-specific TargetPro models.

Fig. 4.

Fig. 4

Features importance of TargetPro models across five disease groups, only features with importance ≥ 2% are labelled.

However, the relative importance of these top features varied by disease context. Matrix factorization was the single most important feature in oncology (22.84%) and fibrotic (17.34%) models. In contrast, attention score was the leading contributor in immune (15.73%), metabolic (17.08%), and neurological (16.07%) models. Other features also showed context-specific utility, such as heterogeneous graph walk in neurological (11.56%) and fibrotic (12.32%) models. Additionally, the knockout score demonstrated a clear contribution pattern, ranking highest in oncology (8.53%), followed in descending order by fibrotic (6.89%), neurological (6.07%), metabolic (4.96%), and immune models (4.20%).

When features were aggregated into higher-level categories, a clear interplay between data sources became apparent. Text derived features were highly influential and represented the largest single category of importance in the immune (52.80%), metabolic (47.00%), and neurological (42.60%) models. The oncology and fibrotic models, however, demonstrated a more balanced reliance on all three feature categories, with nearly equal contributions from text, omics (Static), and omics (Dynamic) sources (see Materials and Methods). Omics-based features contributed substantially across all groups and accounted for the largest share of importance in several models. In particular, omics features collectively accounted for 64.90% and 63.73% of the total importance in the oncology and fibrotic models, respectively, surpassing the contribution from text-based features.

Model performance assessment in target identification

To address the lack of standardized methodologies for evaluating target identification models, we developed TargetBench 1.0. This platform rigorously assesses the performance of various target prediction models while simulating their real-world application environment (see Materials and Methods). Complete target lists for all 38 diseases across the evaluated models are provided in Supplementary Data 5 and 6. TargetBench 1.0 is accessible at https://www.targetbench.org/, enabling users to systematically benchmark their own target identification models using the metrics described below and demonstrated in Supplementary Fig. 2.

We first evaluated model performance by assessing the ability to retrieve established clinical targets, which represent the most reliable ground truth and provide confidence to biologists and drug developers in real-world applications. The was done by measuring the proportion of known clinical targets among the top-ranked predictions from each platform (Fig. 5a). TargetPro achieved an overall precision at top K of 71.6%, corresponding to a 1.7–5.5 fold improvement over the tested LLMs, which showed precision at top K values ranging from 13.1 to 42.3%. TargetPro also significantly outperformed Open Targets, which scored just under 20%. Notably, the strong performance was not confined to a single domain but remained consistently high across diverse therapeutic areas, including oncology, metabolic, immune, fibrotic, and neurological diseases.

Fig. 5.

Fig. 5

Precision at (a) top K Targets and (b) top K Percentage, where K is defined as the number of established clinical stage targets of a specific disease. Data of LLMs are presented as the mean from five repeated queries (N = 5). To calculate precision, TargetPro generated target prediction scores using ‘golden-plus datasets’ (detailed above), mimicking the real-world usage in which analyses combine ‘golden datasets’ with additional user-selected datasets.

Next, we examined how the number of requested targets, a critical consideration from the user’s perspective, influences precision (Fig. 5b). TargetPro consistently demonstrated the highest performance across different percentages of K. Its precision showed a slight gradual decline, starting above 95% when users requested 20% of K targets, and remaining as high as 71.6% when 100% of K was considered. In contrast, all other evaluated models, including the suite of LLMs and Open Targets, exhibited a much steeper decline in precision as the number of requested targets increased. Additionally, newer LLM model versions showed improved performance, with a positive trend observed in both the Grok series and OpenAI models. For instance, some of the higher-performing LLMs (such as o3 and GPT5) started with a precision of > 70% but dropped by over 25% points when 100% K was requested. Open Targets started at more moderate precision levels (< 50%) and also showed a consistent decline, while BioGPT exhibited the lowest precision across all evaluated percentages of K.

While metrics such as precision at top K are useful for evaluating a model’s ability to recover known clinical targets, the real-world utility in drug discovery often lies in identifying promising novel candidates. We therefore focused our analysis on characteristics critical for advancing novel targets through the early stages of the discovery pipeline. Specifically, we examined the novel targets (defined as the top K predictors after excluding known clinical-stage targets) from each model and assessed their druggability, biological relevance, and potential for drug repurposing.

First, in modern therapeutic development, the availability of a known target structure represents a crucial accelerator. TargetPro performs strongly in this regard, with 95.7% of its proposed targets having an available 3D structure in the Protein Data Bank, conferring a clear advantage over other high-performing LLMs, which range from 60.3 to 91.3%. This structural advantage, together with stable and uniformly high performance across all therapeutic areas (91.7–99.6%), underscores TargetPro’s reliability in prioritizing structurally enabled targets suitable for downstream drug design campaigns (Fig. 6a). In addition to structural availability, target druggability further enhances the likelihood of successful therapeutic development, particularly for small-molecule modalities. TargetPro outperforms other LLMs by identifying a higher proportion of druggable targets with clinical evidence (86.5% vs. 38.8–75.0%, Fig. 6b). Furthermore, we evaluated the repurposing potential of novel targets, defined as the percentage of targets with approved drugs in other indications. TargetPro demonstrated a remarkable advantage, with 46% of its targets meeting this criterion, significantly exceeding all other platforms whose outputs typically ranged from 17.0 to 27.5% (Fig. 6c).

Fig. 6.

Fig. 6

Multi-dimensional evaluation of novel targets selected by different target identification models: (a) the percentage of targets with available protein crystal structures; (b) the percentage of targets classified as druggable, supported by clinical evidence; (c) the percentage of targets with approved drugs in other indications; (d) the biological relevance, measured by pathway overlap with clinical-stage targets; (e) the average number of available bioassays; and (f) the average number of available gene modulators. Data of LLMs are presented as the mean from five repeated queries (N = 5).

Beyond druggability, biological relevance to the disease is also essential for target prioritization. To assess this aspect, we examined pathway overlap with established clinical-stage targets and found that TargetPro demonstrated the highest degree of relevance, with an average of 108 overlapping pathways (Fig. 6d). We next evaluated the feasibility of experimental validation. Targets identified by TargetPro were associated with a substantially greater number of available bioassays (averaging over 500), representing a more than 1.4-fold increase compared with any other model (Fig. 6e). Finally, we assessed the availability of gene modulators, which can facilitate the development of new drugs. Targets identified by TargetPro had a higher average number of modulators than those prioritized by LLMs (13.8 vs. 6.1–9.7, Fig. 6f).

Discussion

AI-driven target identification platforms and LLMs have been gradually adopted to enhance the efficiency of target discovery. However, two long-standing challenges remain. First, applying multiple-model strategies for disease-specific target identification is difficult, as no established framework exists for selecting or optimally integrating these models. Second, transparent and systematic approaches for evaluating the performance of target identification platforms are lacking. This study introduces a comprehensive framework designed to address these fundamental challenges. Specifically, our proposed solution comprises two main components: TargetPro, which automatically determines optimal feature combinations to enhance target prediction, and TargetBench 1.0 system, which comprehensively evaluates model’s predictive performance.

This integrated approach is designed to generate actionable intelligence, defined as reliable, well-characterized, and prioritized target hypotheses that can guide subsequent experimental work36. The central premise is that predictive accuracy and the ability to retrieve evaluative clinical targets are not separate objectives but rather complementary aspects of the same process. By developing these components in parallel (Fig. 1), we establish a positive feedback loop in which better models motivate the development of more sophisticated benchmarks. In turn, these benchmarks provide the confidence required to deploy the models in real-world drug discovery pipelines, ultimately aiming to reduce the high failure rates that currently affect drug development.

TargetPro successfully integrated omics and text signals

The design of the TargetPro component is guided by a detailed analysis of the characteristics and limitations underlying the data. This analysis confirmed that the data describing individual targets contain meaningful signals of clinical potential. Specifically, a positive association was observed between each model and a target progression through clinical development, with all 22 features showing statistically significant differences between preclinical and clinical-stage targets (Fig. 2). This finding is consistent with the established understanding that genetic associations and pathway evidence for preclinical targets are often poorly characterized (supported by limited evidence), whereas clinical-stage targets are extensively studied and typically have well-defined mechanisms of action37. However, the utility of individual scores for predicting success between clinical phases appears limited, as most scores tend to plateau once a target enters clinical development. Consequently, while these scores are valuable for prioritization based on historical success patterns, they are not suitable for predicting the likelihood of progression through different clinical stages. This limitation likely reflects the fact that later stage advancements depends not only on the target’s intrinsic biology but is also heavily influenced by external factors such as compound efficacy, trial design, human-specific safety profiles, regulatory considerations, and even commercial priorities4.

We therefore framed the problem as a binary classification task, distinguishing preclinical from clinical-stage targets, and trained machine learning models to predict a target’s overall potential for entry into the clinical development pipeline. The effectiveness of this strategy is demonstrated by model performance, with trained TargetPro models achieving an AUPRC of approximately 0.3, significantly outperforming any individual omics and text scores (Fig. 3B). These results indicate that a well-trained machine learning model can effectively integrate individually weak signals to enhance drug target discovery. This observation is consistent aligns with a growing body of evidence showing that multi-omics data integration can reveal complex biological patterns that are not detectable from any single data type alone38. In addition to improved predictive performance, analysis of feature contributions provides insight into the mechanisms underlying target prioritization. The SHAP analysis presented in Fig. 4 reveals that TargetPro’s decision-making is nuanced and context-dependent, with feature importance varying across disease groups. These results indicate that the model does not rely on simple or, fixed rules but instead captures biologically relevant, disease-specific patterns. For instance, the contributions of omics scores vary across disease types, with the strongest effects observed in oncology, followed by fibrotic, neurological, and metabolic diseases. In contrast, immune-related disorders exhibit a weaker impact from omics data, likely reflecting the systemic and highly adaptive nature of immune responses. These findings align with established biological hallmarks of each disease group. Cancer is characterized by genome instability, frequent mutations, and widespread transcriptional dysregulation, making omics information particularly informative for discovering anti-cancer targets39. Conversely, the complex and dynamic interplay between immune cells, cytokines, and environmental factors may constrain the predictive utility of omics-based scores in immune-related diseases40.

Furthermore, several features consistently rank highly in importance across all disease groups. These include: (i) matrix factorization (9.59–22.84%), a metric for hidden gene–disease associations, (ii) attention score (12.04–17.08%), derived from text mining of scientific literature, and (iii) heterogeneous graph walk (4.55–12.32%), which models relationships by exploring paths within a complex network of biological entities. These findings suggest that TargetPro effectively integrates biological signals derived from omics data with evidence extracted from human knowledge captured from publications, grant applications, and other relevant scientific literature. Such integration yields predictions that are more robust and reliable than those derived from any single data type alone, and generates more concrete hypotheses from target ranking, which is crucial for building confidence in the model among key stakeholders, including drug discovery scientists, potential licensing partners, and regulators.

TargetBench 1.0 accurately benchmarking different platforms

The rapid growth of AI models for drug discovery has created an environment where tools are often evaluated using different datasets and performance metrics, making direct and fair comparisons difficult41. This lack of standardization slows scientific progress and complicates the selection of the most suitable tool for specific research needs. To address this gap, our framework introduces TargetBench 1.0 (Fig. 1), a system designed for rigorous, reproducible, and transparent evaluation of target identification models. TargetBench 1.0 assesses models using criteria that are directly relevant to real-world drug development, such as the ability to retrieve known clinical targets and the drug development potential of newly proposed targets.

A key contribution of this work is the robust comparison of the specialized TargetPro model against a range of LLMs, including the latest general-purpose models such as GPT-5 and Grok4, as well as the domain-specific BioGPT42. This benchmarking was conducted under conditions designed to reflect real-world applications and provides critical insights into the relative performance and applicability of various target discovery models (refer to Materials and Methods for details). TargetPro achieved a top K precision of nearly 70% across all disease areas, substantially outperforming all tested LLMs (ranging 8–40%), and this advantage was preserved regardless of the number of targets requested (Fig. 5). The decline in LLM precision with longer requested target lists reflects a typical behavior of generative models, which tend to be more accurate when prompted for a few top candidates but become less reliable when producing extended lists that include more uncertain predictions. This performance difference is not a critique of LLMs, which have demonstrated remarkable capabilities in reasoning and scientific problem-solving43. Rather, it underscores a fundamental difference in model design and training. TargetPro is a supervised learning model explicitly trained to integrate multiple data modalities for a defined prediction task. In contrast, general-purpose LLMs are primarily pre-trained on vast volumes of unstructured text and are not inherently designed to discover informative weights for the specific categories used in this study without substantial task specific fine-tuning12,44. This underscores the importance of selecting the right tool for a given task and emphasizes the continued value of specialized models for data-intensive scientific applications. Notably, performance improvement aligns with the problem-solving capacities of general-purpose LLMs, as reflected in TargetBench results for the Grok series and OpenAI models (Figs. 5 and 6), further validating the sensitivity and reliability of our designed benchmarking criteria.

The ultimate goal of target identification is not merely to rediscover known targets but to uncover novel candidates. The AI-derived score generated by TargetPro serves as a powerful initial filter, enriching the candidate pool for targets with a higher intrinsic probability of progressing through clinical development (Fig. 5). However, refining a broad list of candidates into a focused set of actionable targets requires the integration of additional measures of ‘development potential’. These include assessment of druggability and supporting experimental evidence, both of which are critical for reducing attrition at later stages of development45. Importantly, a target can be biologically relevant yet remain difficult to modulate with available technologies.

The analysis presented in Fig. 6 evaluates model development potential using several criteria that are critical for generating ‘actionable intelligence’. While TargetPro consistently outperforms other models across most metrics, these measures provide more than a simple comparison of performance between models and disease groups. They also reveal distinct attributes of ‘good’ novel targets across therapeutic areas. The relative contribution of different evidence types varies by disease context, suggesting that target validation strategies should be context-dependent and disease-specific. For example, in a well-established area such as oncology, high-quality targets tend to have strong biological relevance and a large number of available bioassays that support rapid experimental validation (Figs. 6d,e)46,47. In contrast, in high-attrition fields such as neurological and metabolic diseases, the analysis highlights drug repurposing potential as a particularly important de-risking strategy (Fig. 6c)48.

To enable deployment of such evaluations across the drug discovery field, TargetBench 1.0 has been publicly accessible at https://www.targetbench.org/ for independent model assessment. This platform allows researchers to upload a ranked target list generated by their own proprietary models or discovery strategies. It then computes the same key performance metrics used in this study (Figs. 5 and 6) and presents the results graphically (Supplementary Fig. 2). User results are displayed alongside the benchmarked performance of the models reported here, enabling a direct comparison with established baselines.

In addition to benchmarking, the unified framework generates a ‘High Confidence Target List’ for each of the 38 diseases in the database (Supplementary Data 5). These lists are prioritized according to the AI score produced by the TargetPro model and validated through the TargetBench 1.0 system, providing actionable intelligence to guide drug discovery teams in selecting targets for costly and time-intensive experimental validation pipelines36,49.

Limitations and conclusion

While this framework represents a significant step forward, it is important to acknowledge its limitations. The analysis was performed on 38 diseases and does not encompass all human pathologies. However, TargetBench is designed to support regular updates, enabling continuous expansion of disease categories as new data becomes available. In addition, the definition of a ‘successful’ target in the training set was derived from historical clinical trial data, which may introduce inherent biases, as it captures attrition driven not only by scientific factors but also by commercial and strategic considerations. Consequently, TargetPro scores reflect prioritization based on learned historical associations and should not be interpreted as direct predictions of a target’s future clinical success or stage progression. To mitigate the retrospective nature of the dataset, TargetBench 1.0 incorporates alternative evaluative metrics, such as predicted druggability, the availability of approved drugs for other indications, and the number of bioassays for validation, providing a more holistic assessment of target development potential. To validate these predictive capabilities over time, we propose a ‘time capsule’ approach, in which current model predictions are archived and subsequently compared with future clinical outcomes, allowing objective assessment of target identification model performance longitudinally.

We recognize that, within the current feature set, text-based features are particularly susceptible to temporal leakage. Metrics such as grant funding or literature volume often surge after a target enters clinical trials, potentially reflecting ‘popularity’ rather than distinct biological signals. This limitation also applies to LLMs, which are constrained by the information available at the time of their training. To address this issue, TargetBench supports prospective benchmarking against newly emerged clinical data. This capability enables assessment of whether models can identify targets prior to widespread recognition, thereby distinguishing genuine predictive signal from retrospective popularity bias.

We further acknowledge the potential concern in the comparative analysis between TargetPro and LLMs. TargetPro superior performance may be partially attributable to its specialized training for target-prediction, whereas the baseline LLMs evaluated in this study were not fine-tuned for this specific task. However, this comparison reflects current research practices where users typically employ general-purpose LLMs directly for hypothesis generation. Additionally, TargetBench provides a flexible platform to facilitate more granular comparisons, such as evaluation of user-fine-tuned models. For instance, TargetPro outperforms a fine-tuned model based on the BioGPT framework trained on biomedical texts for target discovery17. Moreover, as general purpose LLMs continue to improve in research settings50, this further reinforces the relevance of our evaluation design.

Looking ahead, our framework provides a clear path for continued development. The workflow we applied to construct TargetPro is not restricted to the current 22 omics and text models and can be readily extended to additional multidimensional data sources. For example, models from the Open Targets platform (which currently rely on default rankings), could be incorporated to generate similarly enhanced predictive models11,51,52. The underlying strategy of integrating multiple weak signals to predict a defined outcome is broadly applicable. By adapting the input features and the target variable, this workflow could be adapted to address other critical challenges in drug development, such as prediction of patient response to therapy, identification of biomarkers for clinical trial stratification, or estimation of phase-specific success probabilities.

Finally, this framework can be integrated into fully automated, closed-loop discovery systems, often referred to as ‘self-driving labs’53. In this setting, TargetPro can function as the ‘AI brain’ to identify and prioritize promising targets, while results from automated in vitro knockout or inhibition experiments are fed back into the model for retraining and improvement, creating a continuous cycle of hypothesis generation, testing, and learning. Our framework provides two core software components required for such systems: an AI model for target prioritization and an evaluation module for assessing experimental outcomes. Together, these components establish a dynamic and extensible platform that supports a more AI-guided and automated paradigm for drug discovery.

Related links

  1. Verge Genomics: https://www.vergegenomics.com/approach.

  2. TargetMATCH: https://www.owkin.com/targetmatch.

  3. HGNC database: https://www.genenames.org/.

  4. MedChemExpress: https://www.medchemexpress.com.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (469.4KB, pdf)
Supplementary Material 2 (47.3KB, pdf)
Supplementary Material 3 (48.4KB, xlsx)
Supplementary Material 4 (10.5KB, xlsx)
Supplementary Material 5 (5.7KB, xlsx)
Supplementary Material 6 (8.5KB, xlsx)
Supplementary Material 7 (9.8MB, xlsx)
Supplementary Material 8 (10.4MB, xlsx)

Acknowledgements

We thank Ms. Elizaveta Ekimova for her technical assistance with figure design.

Author contributions

F.W.P., F.R., and A.Z. conceived the project.H.L., C.D., W.G., J.C., Y.X. and Z.Z. designed the framework and conducted the analyses.H.L., W.G., and J.C. prepared the figures.F.W.P. and A.Z. supervised the project.H.L. and C.D. wrote the manuscript.H.L., C.D., W.G., J.C., Y.X., Z.Z., V.N., D.G., M.Z., A.A., F.R., E.I., F.W.P. and A.Z. participated in manuscript review and editing.

Data availability

All datasets used in this study are publicly available and are detailed in Supplementary Data 1.

Code availability

The custom code used to generate the results in this study is proprietary intellectual property. It is available to researchers for non-commercial academic use upon reasonable request to the corresponding authors.

Declarations

Competing interests

All authors except E.I are affiliated with Insilico Medicine: Insilico Medicine is a global clinical-stage commercial generative artificial intelligence company with several hundred patents, pending patent applications, and commercially available software.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally to this work: Howell Leung, Chengchen Duan and Wenhao Gou.

Contributor Information

Frank W. Pun, Email: frank.pun@insilico.com

Alex Zhavoronkov, Email: alex@insilico.com.

References

  • 1.Dowden, H. & Munro, J. Trends in clinical success rates and therapeutic focus. Nat. Rev. Drug Discov.18, 495–496. 10.1038/d41573-019-00074-z (2019). [DOI] [PubMed] [Google Scholar]
  • 2.Hay, M., Thomas, D. W., Craighead, J. L., Economides, C. & Rosenthal, J. Clinical development success rates for investigational drugs. Nat. Biotechnol.32, 40–51. 10.1038/nbt.2786 (2014). [DOI] [PubMed] [Google Scholar]
  • 3.Mullard, A. Parsing clinical success rates. Nat. Rev. Drug Discov.15, 447. 10.1038/nrd.2016.136 (2016). [DOI] [PubMed] [Google Scholar]
  • 4.Sun, D., Gao, W., Hu, H. & Zhou, S. Why 90% of clinical drug development fails and how to improve it?. Acta Pharm. Sin. B12, 3049–3062. 10.1016/j.apsb.2022.02.002 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Cook, D. et al. Lessons learned from the fate of AstraZeneca’s drug pipeline: A five-dimensional framework. Nat. Rev. Drug Discov.13, 419–431. 10.1038/nrd4309 (2014). [DOI] [PubMed] [Google Scholar]
  • 6.Morgan, P. et al. Impact of a five-dimensional framework on R&D productivity at AstraZeneca. Nat. Rev. Drug Discov. 17, 167–181. 10.1038/nrd.2017.244 (2018). [DOI] [PubMed] [Google Scholar]
  • 7.Pun, F. W., Ozerov, I. V. & Zhavoronkov, A. AI-powered therapeutic target discovery. Trends Pharmacol. Sci.44, 561–572. 10.1016/j.tips.2023.06.010 (2023). [DOI] [PubMed] [Google Scholar]
  • 8.Paul, D. et al. Artificial intelligence in drug discovery and development. Drug Discov. Today26, 80–93. 10.1016/j.drudis.2020.10.010 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Carvalho-Silva, D. et al. Open Targets Platform: new developments and updates two years on. Nucleic Acids Res.47, 1056–1065. 10.1093/nar/gky1133 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ghoussaini, M. et al. Open Targets Genetics: systematic identification of trait-associated genes using large-scale genetics and functional genomics. Nucleic Acids Res.49, 1311–1320. 10.1093/nar/gkaa840 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ochoa, D. et al. Open Targets Platform: supporting systematic drug-target identification and prioritisation. Nucleic Acids Res.49, 1302–1310. 10.1093/nar/gkaa1027 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Kamya, P. et al. Pandaomics: An AI-driven platform for therapeutic target and biomarker discovery. J. Chem. Inf. Model.64, 3961–3969. 10.1021/acs.jcim.3c01619(2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Pun, F. W. et al. A comprehensive AI-driven analysis of large-scale omic datasets reveals novel dual-purpose targets for the treatment of cancer and aging. Aging Cell.22, 14017. 10.1111/acel.14017 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Ren, F. et al. A small-molecule TNIK inhibitor targets fibrosis in preclinical and clinical models. Nat. Biotechnol.43, 63–75. 10.1038/s41587-024-02143-0 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tang, Q. et al. AI-Driven Robotics Laboratory Identifies Pharmacological TNIK Inhibition as a Potent Senomorphic Agent. Aging Dis.10.14336/ad.2024.1492 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Xu, Z. et al. A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: A randomized phase 2a trial. Nat. Med.10.1038/s41591-025-03743-2 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zagirova, D. et al. Biomedical generative pre-trained based transformer language model for age-related disease target discovery. Aging15, 9293–9309. 10.18632/aging.205055 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Chakraborty, C., Bhattacharya, M. & Lee, S. S. Artificial intelligence enabled ChatGPT and large language models in drug target discovery, drug discovery, and development. Mol. Ther. Nucleic Acids33, 866–868. 10.1016/j.omtn.2023.08.009 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Zong, N. et al. BETA: a comprehensive benchmark for computational drug-target prediction. Brief. Bioinform10.1093/bib/bbac199 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Siegel, R. L., Giaquinto, A. N. & Jemal, A. Cancer statistics, 2024. CA Cancer J. Clin.74, 12–49. 10.3322/caac.21820 (2024). [DOI] [PubMed] [Google Scholar]
  • 21.Reynolds, J. A. & Putterman, C. Progress and unmet needs in understanding fundamental mechanisms of autoimmunity. J. Autoimmun.137, 102999. 10.1016/j.jaut.2023.102999 (2023). [DOI] [PubMed] [Google Scholar]
  • 22.Chew, N. W. S. et al. The global burden of metabolic disease: Data from 2000 to 2019. Cell Metab.35, 414-428 413. 10.1016/j.cmet.2023.02.003 (2023). [DOI] [PubMed] [Google Scholar]
  • 23.Mutsaers, H. A. M., Merrild, C., Nørregaard, R. & Plana-Ripoll, O. The impact of fibrotic diseases on global mortality from 1990 to 2019. J. Transl Med.21, 818. 10.1186/s12967-023-04690-7 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Savelieff, M. G., Noureldein, M. H. & Feldman, E. L. Systems biology to address unmet medical needs in neurological disorders. Methods Mol. Biol.2486, 247–276. 10.1007/978-1-0716-2265-0_13 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V. & Gulin, A. CatBoost: unbiased boosting with categorical features. arXiv 10.48550/arXiv.1706.09516 (2017).
  • 26.Zou, H. & Hastie, T. Regularization and variable selection via the elastic net. J. Roy. Stat. Soc.67, 301–320. 10.1111/j.1467-9868.2005.00503.x (2005). [Google Scholar]
  • 27.Ke, G. et al. LightGBM: a highly efficient gradient boosting decision tree. Association Comput. Mach.10.5555/3294996.3295074 (2017). [Google Scholar]
  • 28.Breiman, L. Random forests. Mach. Learn.45, 5–32. 10.1023/A:1010933404324 (2001). [Google Scholar]
  • 29.Chen, T. & Guestrin, C. XGBoost: A scalable tree boosting system. Association Comput. Mach.10.1145/2939672.2939785 (2016). [Google Scholar]
  • 30.Elkan, C. The foundations of cost-sensitive learning. Association Comput. Mach.10.5555/1642194.1642224 (2001). [Google Scholar]
  • 31.He, H. & Garcia, E. A. Learning from imbalanced data. IEEE10.1109/TKDE.2008.239 (2009). [Google Scholar]
  • 32.Oprea, T. I. Unexplored therapeutic opportunities in the human genome. Nat. Rev. Drug Discov.10.1038/nrd.2018.14 (2018). [DOI] [PubMed] [Google Scholar]
  • 33.Ponce-Bobadilla, A. V., Schmitt, V., Maier, C. S., Mensing, S. & Stodtmann, S. Practical guide to SHAP analysis: Explaining supervised machine learning model predictions in drug development. Clin. Transl Sci.17, 70056. 10.1111/cts.70056 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Stoeger, T., Gerlach, M., Morimoto, R. I. & Nunes Amaral, L. A. Large-scale investigation of the reasons why potentially important genes are ignored. PLoS Biol.16, 2006643. 10.1371/journal.pbio.2006643 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Saito, T. & Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One10.1371/journal.pone.0118432 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Ocana, A. et al. Integrating artificial intelligence in drug discovery and early drug development: a transformative approach. Biomark. Res.13, 45. 10.1186/s40364-025-00758-2 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Nelson, M. R. et al. The support of human genetic evidence for approved drug indications. Nat. Genet.47, 856–860. 10.1038/ng.3314 (2015). [DOI] [PubMed] [Google Scholar]
  • 38.Picard, M., Scott-Boyer, M. P., Bodein, A., Périn, O. & Droit, A. Integration strategies of multi-omics data for machine learning analysis. Comput. Struct. Biotechnol. J.19, 3735–3746. 10.1016/j.csbj.2021.06.030 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hanahan, D. Hallmarks of cancer: New dimensions. Cancer Discov.12, 31–46. 10.1158/2159-8290.Cd-21-1059 (2022). [DOI] [PubMed] [Google Scholar]
  • 40.Mangino, M., Roederer, M., Beddall, M. H., Nestle, F. O. & Spector, T. D. Innate and adaptive immune traits are differentially affected by genetic and environmental factors. Nat. Commun.8, 13850. 10.1038/ncomms13850 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Ferreira, F. J. N. & Carneiro, A. S. AI-Driven Drug Discovery: A Comprehensive Review. ACS Omega. 10, 23889–23903. 10.1021/acsomega.5c00549 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Luo, R. et al. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Brief. Bioinform10.1093/bib/bbac409 (2022). [DOI] [PubMed] [Google Scholar]
  • 43.Jiang, J. et al. Benchmarking large language models on multiple tasks in bioinformatics nlp with prompting. arXiv 10.48550/arXiv.2503.04013 (2025).
  • 44.Zheng, Y. et al. Large language models in drug discovery and development: From disease mechanisms to clinical trials. arXiv10.48550/arXiv.2409.04481 (2024).41031081 [Google Scholar]
  • 45.Hughes, J. P., Rees, S., Kalindjian, S. B. & Philpott, K. L. Principles of early drug discovery. Br. J. Pharmacol.162, 1239–1249. 10.1111/j.1476-5381.2010.01127.x (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Berger, M. F. & Mardis, E. R. The emerging clinical relevance of genomics in cancer medicine. Nat. Rev. Clin. Oncol.15, 353–365 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Creixell, P. et al. Pathway and network analysis of cancer genomes. Nat. Methods. 12, 615–621. 10.1038/nmeth.3440 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Cummings, J. L. et al. Drug repurposing for Alzheimer’s disease and other neurodegenerative disorders. Nat. Commun.16, 1755. 10.1038/s41467-025-56690-4 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Abdelsayed, M., Kort, E. J., Jovinge, S. & Mercola, M. Repurposing drugs to treat cardiovascular disease in the era of precision medicine. Nat. Rev. Cardiol.19, 751–764. 10.1038/s41569-022-00717-6 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Othman, Z. K. et al. Advancing drug discovery and development through GPT models: a review on challenges, innovations and future prospects. Intelligence-Based Med.10.1016/j.ibmed.2025.100233 (2025). [Google Scholar]
  • 51.Buniello, A. et al. Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Res.53, 1467–1475. 10.1093/nar/gkae1128 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Koscielny, G. et al. Open Targets: a platform for therapeutic target identification and validation. Nucleic Acids Res.45, 985–994. 10.1093/nar/gkw1055 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Fehlis, Y., Mandel, P., Crain, C., Liu, B. & Fuller, D. Accelerating drug discovery with artificial: A whole-lab orchestration and scheduling system for self-driving labs. arXiv10.48550/arXiv.2504.00986 (2025). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (469.4KB, pdf)
Supplementary Material 2 (47.3KB, pdf)
Supplementary Material 3 (48.4KB, xlsx)
Supplementary Material 4 (10.5KB, xlsx)
Supplementary Material 5 (5.7KB, xlsx)
Supplementary Material 6 (8.5KB, xlsx)
Supplementary Material 7 (9.8MB, xlsx)
Supplementary Material 8 (10.4MB, xlsx)

Data Availability Statement

All datasets used in this study are publicly available and are detailed in Supplementary Data 1.

The custom code used to generate the results in this study is proprietary intellectual property. It is available to researchers for non-commercial academic use upon reasonable request to the corresponding authors.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES