Skip to main content
ACS Omega logoLink to ACS Omega
. 2026 Feb 11;11(7):11875–11889. doi: 10.1021/acsomega.5c10773

Exploration of Novel Chemical Spaces to Discover JAK1 Inhibitors: An Ensemble Docking-Guided Deep Learning Approach

Baddipadige Raju 1, Gera Narendra 1, Bineet Kumar Mohanta 1, Anshuman Dixit 1,*
PMCID: PMC12947007  PMID: 41768660

Abstract

Janus kinase 1 (JAK1) is a key regulator of cytokine signaling and a validated therapeutic target in autoimmune, inflammatory, and oncological disorders. However, existing JAK inhibitors such as Tofacitinib and Ruxolitinib are limited by their narrow pyrrolo­[2,3-d]­pyrimidine scaffold, leading to poor isoform selectivity, JAK3 cross-reactivity, and dose-limiting toxicity. Expanding the chemical space for JAK1 inhibition while achieving higher selectivity therefore represents a critical challenge in drug discovery. To overcome these limitations, we developed a deep learning (DL) based virtual screening framework (VS) that explicitly integrates protein flexibility with a billion-scale chemical exploration. Eight high-resolution JAK1 crystal structures were employed to capture conformational diversity of the ATP-binding pocket. Ensemble docking scores derived from these structures were used to train a deep neural network (DNN) classifier on rigorously curated data sets. The model was applied to over 1.1 billion commercially available compounds from the ZINC database, identifying 131,730 high-confidence candidates. Redocking analysis confirmed that 57% of these compounds consistently surpassed a stringent activity threshold across all receptor conformations, underscoring the robustness of the approach. Scaffold-based analysis of the top 10% candidates revealed 7652 unique chemotypes, with only 13 overlapping with scaffolds of known JAK1 inhibitors, highlighting the substantial novelty of the predicted chemical space. Furthermore, physicochemical and ADME filtering enriched for candidates with favorable drug-like properties. By explicitly embedding receptor flexibility into a scalable artificial intelligence framework, this study establishes a generalizable strategy for kinase-targeted drug discovery and opens new opportunities for selective JAK1 inhibitor development.


graphic file with name ao5c10773_0013.jpg


graphic file with name ao5c10773_0011.jpg

1. Introduction

Janus kinase 1 (JAK1) is a nonreceptor tyrosine kinase that plays a central role in cytokine signaling via the JAK-STAT (Signal Transducer and Activator of Transcription) pathway, , a fundamental mechanism linking extracellular immune stimuli to transcriptional responses. , JAK1 transduce signals from key cytokines such as IL-2, IL-6, and IL-10, regulating immune homeostasis, inflammation, and host defense. , Upon cytokine binding, JAK1 undergoes autophosphorylation and activates downstream STAT proteins, initiating gene expression programs critical for immune function. − Aberrant JAK1-STAT signaling has been implicated in autoimmune and inflammatory disorders including, rheumatoid arthritis (RA), inflammatory bowel disease (IBD), and atopic dermatitis (AD), , as well as malignancies such as breast, lung, colorectal, and liver cancers. These broad associations have established JAK1 as a high-value therapeutic target. However, the first generation of JAK inhibitors (e.g., Tofacitinib, Ruxolitinib, Baricitinib) exhibited limited isoform selectivity, leading to unintended inhibition of JAK2/JAK3 and adverse events such as cytopenias, thromboses, and heightened infection risk. − These safety concerns prompted regulatory scrutiny and spurred the development of more selective agents. Second-generation JAK1-selective inhibitors such as Upadacitinib and Filgotinib improved therapeutic indices but continue to suffer from partial selectivity, dose-limiting off-target effects, and metabolic instability. , Most contemporary efforts remain confined to ligand-centric or scaffold-repurposing strategies that inadequately explore novel chemical space. Given the vastness of theoretical chemical space (>1060 molecules), traditional screening and drug design strategies typically limited to libraries of <106 molecules represent a narrow sampling, constraining innovation.

Recent advances in AI and machine learning (ML) offer powerful tools to transcend these limitations. ML models trained on molecular fingerprints and physicochemical descriptors have shown promise in predicting bioactivity and drug-likeness, − but many are constrained by biased data sets of known ligands. Moreover, prior billion-scale virtual screening (VS) efforts largely neglect protein flexibility, relaying on single crystal structures for docking, and frequently employ random or biased selection of training and test sets, reducing model generalizability and hit diversity. , A summary of previous studies utilizing limited JAK1 inhibitors is provided in Table . In parallel, deep docking methods and ML-augmented scoring functions have enabled the rapid screening of ultralarge libraries, reaching billion-compound scales. − Yet, these workflows often rely on chemically unrepresentative sampling, leading to poor coverage of chemical diversity and diminished hit prioritization. , To address these limitations, we developed a chemically informed, AI-driven pipeline specifically for the discovery of novel JAK1 inhibitors. Our approach integrates ensemble docking across multiple JAK1 structures to explicitly capture JAK1 flexibility, combined with deep neural network models (DNN) trained on diverse, rigorously curated data sets. Unlike conventional workflows, our method employs unsupervised clustering and chemical space analysis to align the diversity of training and test data sets thereby improving generalizability and reducing selection bias.

1. Summary of Studies That Employed Limited JAK1 Inhibitor Datasets for the Development of ML Models.

author target data size type ML/DL method source
Wang et al. JAK1 ∼16 K ligand-based + pharmacophore ML + docking
Bu et al. Pan-JAK ∼7.4 K traditional descriptors ensemble ML
Yang et al. JAK1 ∼3 K known inhibitors only random forest, SVM

The pipeline comprises three key stages: (i) compound selection based on structural diversity, (ii) high-resolution ensemble-docking to refine and rank predicted binders (iii) development of DNN models trained on ensemble docking scores to classify JAK1 binding affinity, and applied to billion-scale library. This hierarchical strategy identified structurally novel, high-confidence hits with minimal similarity to known JAK1 inhibitors, yet exhibited superior binding affinity compared to the cocrystallized ligands of multiple JAK1 crystal structures. Importantly, this study represents the first deep learning (DL) based VS framework focused on JAK1 providing both a practical route for discovering potential JAK1 binders and a generalizable template applicable to other kinase families, particularly where conformational dynamics critically influence ligand-binding and selectivity. This framework has broad implications for AI-guided drug discovery targeting kinases and other proteins with flexible active sites, enhancing robustness, novelty, and translational relevance.

In the context of chronic inflammatory and autoimmune diseases, where safety and long-term tolerability are essential, our findings provide both a promising new class of candidate JAK1 binders and a blueprint for rational, AI-enabled exploration of untapped chemical space.

2. Materials and Methods

2.1. Data Collection and Preparation

The ZINC20 [https://files.docking.org/zinc20-ML/] database, comprising approximately 1.1 billion commercially available and annotated small molecules, is one of the most widely utilized ultralarge chemical libraries in VS and drug discovery. This data set was specifically employed to support deep docking-based VS for novel JAK1 binders. The complete data set includes 100 preprocessed SMI (Simplified Molecular Input Line Entry System) files, each containing approximately 10 million unique small molecule structures. To facilitate DL modeling, Morgan fingerprints were generated for these compounds. These fingerprints represent molecular structures based on circular substructures, where each atom’s environment is encoded within a specified radius, capturing relevant topological and chemical information. The generated fingerprints were divided into 100 separate text files, each corresponding to one of the original SMI files. This structured format enabled efficient handling and high-throughput processing of the chemical space, ultimately supporting our goal of identifying structurally diverse and previously unexplored scaffolds for selective JAK1 inhibition.

2.2. Generation of Train and Test Data Sets by Stratified Clustering

To develop DL models with robust and generalizable performance for JAK1 inhibitor prediction, two independent and chemically diverse data sets training and test sets were constructed using a stratified clustering-based sampling approach. This strategy addresses the limitations of simple random sampling, which often yields structurally redundant and chemically biased subsets, reducing their effectiveness in large-scale VS. From each file, 10,000 molecules were randomly sampled, resulting in an initial pool of 1 million seed molecules. Although the initial sampling was random, it was performed at scale to ensure broad and unbiased coverage of chemical space, serving as a representative basis for subsequent diversity-driven clustering. Each molecule was encoded using 1024-bit Morgan fingerprints (radius = 2) to capture its topological and structural features. Due to the high dimensionality of these fingerprints, Principal Component Analysis (PCA) was employed to reduce the feature space to 50 components, preserving the majority of structural variance while enhancing computational efficiency. The PCA-transformed molecular fingerprints obtained from one million representative compounds were subjected to MiniBatchKMeans clustering, yielding 3000 distinct clusters that delineated chemically coherent subspaces within the sampled chemical landscape. The cluster number (k = 3000) was empirically optimized to provide an optimal trade-off between maximizing chemical diversity and maintaining computational tractability. For large-scale virtual screening, a stringent probability threshold of 0.99 was imposed on the deep learning classifier to ensure high predictive precision while effectively minimizing false-positive identifications. This unsupervised clustering step ensured inclusion of both common and rare chemotypes while mitigating sampling bias. To extend this clustering strategy to the entire database, all 1.1 billion molecules were projected into the learned PC space and assigned to the nearest of the 3000 clusters using the trained KMeans model. From each cluster, up to 1000 molecules were selected based on their proximity to the cluster centroid in PCA space, resulting in a master data set of 3.0 million structurally diverse compounds. This step ensured that no single cluster dominated the data set, maintaining balanced chemical diversity. By integrating unsupervised clustering, dimensionality reduction, and stratified sampling, this methodology facilitated the construction of DL ready data sets that are chemically representative, nonredundant, and scalable, enabling the discovery of structurally novel JAK1 binders with high confidence.

2.3. Docking

Eight high-resolution crystal structures of human JAK1 (PDB IDs: 3EYG, 3EYH, 4I5C, 4IVD, 6AAH, 6SM8, 6SMB, and 6W8L) were used to capture the conformational flexibility of the kinase active i.e., ATP-binding site. − Structures were retrieved from the Protein Data Bank and prepared by protonation at pH 7.4, and any missing side chains were modeled necessary using make receptor module available in Open Eye software. The binding pocket in each PDB was defined with respect to the cocrystallized ligand to ensure biologically relevant docking. RMSD calculations were performed on Cα atoms of the active site residues (5 Å radius) to quantify the structural deviation across the ensemble. Pocket volumes were estimated using the molecular surface of the active site residues, providing a proxy for the accessible pocket space for ligand binding. Ligand preparation was carried out using OMEGA v4.1.2.0 (OpenEye Scientific Software), where both cocrystal ligands and external screening compounds were processed to generate up to 200 low-energy conformers per molecule, thereby accounting for conformational flexibility and ensuring broad chemical space sampling while minimizing redundancy. Docking was performed using FRED (Fast Rigid Exhaustive Docking) in OEDOCKING v4.2.1.0, which systematically places ligand conformers into the binding pocket, and poses were evaluated using the Chemgauss4 scoring function to maintain consistency and reproducibility across all PDBs. To explicitly account for protein flexibility, a redocking experiment was performed in which each cocrystallized ligand was docked into the seven alternate JAK1 structures, and the resulting docking scores were averaged across all eight PDBs to obtain ensemble docking profiles, thus reflecting conformational variability of the binding site rather than reliance on a single static structure. The ensemble docking scores were then used to define an activity-based classification scheme: molecules with docking scores ≤−12 kcal/mol were labeled as binders, those with docking scores above −9 kcal/mol were labeled as nonbinders, while compounds falling between −12 and −9 kcal/mol were considered part of a gray area and excluded from training to reduce misclassification and enhance label reliability. This three-tier strategy ensured a balanced and biologically meaningful data set while avoiding the pitfalls of extreme class imbalance commonly observed in structure-based data sets.

2.4. Deep Learning Model Development and Validation on Hard Test Sets

A deep neural network (DNN) classifier was built using the curated data set to distinguish JAK1 binders from nonbinders. The network comprised fully connected layers (1024 → 512 → 128 → 32 neurons) with SWISH (Sigmoid Linear Unit) and GELU (Gaussian Error Linear Unit) activation functions in hidden layers to enhance nonlinear feature learning capacity. A sigmoid activation was applied in the output layer for binary classification. The network was compiled with the Adam optimizer (learning rate = 0.0001) and trained using binary cross-entropy as the loss function. To mitigate overfitting and improve generalization, batch normalization was applied after each dense layer, and dropout was incorporated with decreasing rates (0.3–0.2) across successive layers. Model performance was rigorously assessed using an independent validation set. Standard classification metrics accuracy, sensitivity (recall), specificity, precision, and F1-score, and MCC were computed using established eqs – to provide an interpretable and quantitative assessment of predictive reliability. To further assess model generalizability, three independent “hard” test sets were created using Tanimoto cutoffs of 0.3, 0.4, and 0.5 relative to the training set. These sets consisted exclusively of molecules structurally dissimilar to the training data, ensuring that model performance on these data sets reflects the ability to generalize to novel chemical scaffolds. Performance on these challenging test sets provided a stringent evaluation of DNN robustness and reliability in ultralarge-scale VS.

Sensitivity=TPTP+FN 1
Specificity=TNTN+FP 2
Precision=TPTP+FP 3
Accuracy=TP+TNTP+TN+FP+FN 4
F1=2×Precision×SensitivityPrecision+Sensitivity 5
MCC=(TP×TN)−(FN×FP)√(TP+FN)(TP+FP)(TN+FN)(TN+FP) 6

2.5. Virtual Screening

Following the development and validation of two DNN classification models employing SWISH and GELU activation functions, these models were applied to virtually screen the ZINC20 database, comprising over one billion commercially available small molecules. Compounds previously used in training and testing were excluded to prevent data leakage and ensure unbiased predictions. The DNN models were trained to classify compounds as potential JAK1 binders using docking-score-guided labels. Both models demonstrated an enrichment ratio of ∼2, indicating moderate early recognition of true positives. For large-scale VS, only high-confidence hits compounds with predicted probabilities ≥0.99 were selected for downstream evaluation. While enrichment factor (EF) provides a measure of early enrichment in model performance, using EF alone may not always justify selecting extremely high-confidence hits, as EF reflects ranking performance rather than absolute probability calibration. Therefore, the high-confidence threshold was chosen to prioritize precision and minimize false positives during the screening of an ultralarge chemical space. Shortlisted compounds were subsequently subjected to high-precision FRED docking against the JAK1 kinase domain to assess binding affinities and interaction profiles. This integrated workflow, combining ML-driven classification with structure-based docking, enabled the efficient identification and prioritization of novel JAK1 inhibitor candidates. A schematic representation of the VS workflow is shown in Figure .

1.

1

Schematic representation of the detailed DNN model development and virtual screening (VS) pipeline.

2.6. Chemical Space Analysis

To evaluate the structural novelty and chemical diversity of the screened compounds, a comprehensive analysis was performed comparing the identified hits with known JAK1 inhibitors. Molecular similarity was quantified using the Tanimoto coefficient, computed from Morgan fingerprints generated via RDKit. Compounds with low similarity scores were considered structurally novel. Complementing this, Bemis–Murcko scaffold analysis was conducted to assess scaffold novelty by comparing the core frameworks of screened hits with those of known inhibitors. Key physicochemical properties including molecular weight (MW), logarithm of the partition coefficient (ALogP), topological polar surface area (TPSA), hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and quantitative estimate of drug-likeness (QED) were calculated to evaluate drug-likeness, permeability, and bioavailability potential. To visualize structural diversity, high-dimensional molecular fingerprint data were projected into two dimensions using Uniform Manifold Approximation and Projection (UMAP). Two separate UMAP visualizations were generated: (i) to map the distribution and clustering of all screened hits within chemical space, and (ii) to directly compare the identified hits with known JAK1 inhibitors, highlighting regions of novelty and overlap. Additionally, heatmaps of pairwise Tanimoto similarity scores between screened hits and known inhibitors were generated to provide an intuitive representation of structural similarity patterns. Compounds showing high structural novelty and superior docking scores relative to reference JAK1 cocrystallized ligands were prioritized for further validation. To further assess the reliability of the DNN models, these high-confidence hits were subjected to ensemble docking across eight distinct JAK1 crystal structures, each cocrystallized with a known inhibitor. 3D conformers were generated using OMEGA, and docking was performed with FRED consistently across all structures to ensure methodological comparability and evaluate binding consistency across multiple biologically relevant protein conformations.

2.7. Drug Likeness Filtering

To ensure the drug-like nature of the identified hit compounds, drug-likeness filtering was carried out using the OpenEye FILTER tool with the BlockBuster filter. The BlockBuster filter is a comprehensive set of medicinal chemistry rules derived from the analysis of highly successful orally administered drugs, typically those generating over one billion dollars in annual sales. The filtering criteria included thresholds for molecular weight (130–781 Da), heavy atom count (9–55), rotatable bonds (0–16), rigid bonds (4–55), hydrogen bond donors (HBD) (0–9), hydrogen bond acceptors (HBA) (0–13), XLogP (−3.0 to 6.85), and topological polar surface area (TPSA) (0–205 Å2), among others. Additionally, compounds containing undesirable or reactive functionalities such as aldehydes, Michael acceptors, or PAINS (pan-assay interference compounds) were flagged and excluded. Filtering was implemented using Python and OpenEye FILTER, and only compounds that met all BlockBuster criteria were retained for downstream analysis. To further prioritize structurally diverse and high-quality candidates, the filtered hits were subjected to clustering using HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), based on 2D UMAP projections of their molecular fingerprints. This strategy ensured the final selection encompassed both high predicted activity and broad scaffold diversity, facilitating a balanced and representative subset for subsequent experimental or computational validation.

3. Results and Discussion

3.1. Chemical Space Analysis of Generated Train and Test Sets

The original data set of 3 million compounds was split into training (70%) and test (30%) sets using a stratified clustering approach, ensuring balanced representation of chemical diversity across subsets. To evaluate the structural distribution and overlap, 1024-bit Morgan fingerprints (radius = 2) were projected into two dimensions using UMAP, a nonlinear dimensionality reduction technique. As shown in Figure , the resulting UMAP plot reveals a highly interspersed and well-overlapped distribution of training (red) and test (blue) compounds across chemical space, indicating substantial structural similarity between the two sets. This overlap suggests that the test compounds largely reside within the applicability domain (AD) defined by the training data. UMAP embeddings were generated from binary Morgan fingerprints (1024 bits, radius = 2) using Jaccard distance, with parameters n_neighbors = 15 and min_dist = 0.1. In addition to this qualitative visualization, the applicability domain was quantitatively assessed by computing the maximum Tanimoto similarity of each screened compound to the training set, providing an independent and complementary measure of chemical space coverage (Supplementary File 1: Figure S1). Satisfying this condition is crucial for model reliability, as predictions made outside the AD are prone to reduced accuracy and questionable interpretability. The observed dense and uniform clustering across data sets indicates that the stratified clustering-based data set construction successfully preserved structural diversity while avoiding data set-specific chemical biases. Importantly, this approach goes beyond conventional random or scaffold-based splitting methods by integrating unsupervised clustering with PCA-guided stratification, ensuring a balanced chemical space representation in each subset. Such a strategy enables the development of DL models that are robust and generalizable, capable of making meaningful predictions on unseen, yet chemically relevant, compounds. The novelty of this method lies in its scalable and diversity-preserving data set generation pipeline, tailored for ultralarge chemical libraries. By confirming that the test and validation compounds are embedded within the learned chemical domain of the training data, this analysis validates the relevance, representativeness, and trustworthiness of the model’s predictive performance. Overall, the UMAP-based chemical space visualization strongly supports the integrity of the data set splitting strategy and reinforces the scientific soundness of the modeling workflow.

2.

2

UMAP visualization of the chemical space distribution for the training (70%) and test (30%) sets. The observed overlap and comparable coverage of both sets uniformly capture the chemical diversity, ensuring representative sampling and reducing bias, which is critical for reliable model training and validation.

3.2. Deep Docking

Eight high-resolution JAK1 crystal structures (PDB IDs: 3EYG, 3EYH, 4I5C, 4IVD, 6AAH, 6SM8, 6SMB, and 6W8L) were employed to capture the conformational heterogeneity of the ATP-binding site and to support an ensemble docking strategy. Our RMSD analysis of the 5Å ligand-proximal pocket (Cα atoms) shows a spectrum of deviations relative to the reference 3EYG structure, ranging from 0.453 Å for 3EYH, indicating near-identical geometry, to 1.435 Å for 6SMB, reflecting significant structural rearrangement. Detailed RMSD values and corresponding pocket volumes are provided in Supplementary File 1, Table S1. These deviations are further supported by pocket surface measurements: the calculated areas span from 1968 (3EYH) to 6752 Å2 (6SM8), demonstrating meaningful variability in pocket size and shape. Importantly, some structures (4I5C, 6AAH, 6SMB) have one residue missing in the active site, which accounts for minor discrepancies in RMSD or pocket volume, but does not compromise the overall ensemble representation. Together, the RMSD values and pocket volumes substantiate that the selected ensemble captures both small-scale and moderate conformational flexibility in the JAK1 active site. This ensures that docking and virtual screening studies sample a realistic range of binding site conformations, rather than being biased by a single static structure. In summary, the data indicate that the ensemble provides comprehensive coverage of active site flexibility, and the observed variations in pocket geometry and volume are consistent with the intrinsic plasticity of the JAK1 kinase domain. Each cocrystallized ligand was redocked across all eight receptor conformations, generating the cross-structural docking score matrix shown in Figure . Furthermore, the mean docking score for self-docking was −13.56 kcal/mol, while the off-diagonal mean docking score was −13.18 kcal/mol, demonstrating the minimal bias toward any specific crystal structure and robust pose transferability across JAK1 conformations. From this heatmap it was identified that there is clear difference in binding affinities across eight JAK1 structures illustrating the conformational flexibility of the structures. Mean docking of each cocrystal ligand against all structures was calculated and details of the scores were represented in (Figure S2). As shown in the Figure S2, all eight cocrystal ligands achieved mean docking scores of ≤−12 kcal/mol, indicating consistently high binding affinity across JAK1 structures. Based on this observation, −12 kcal/mol selected as a biologically meaningful cutoff for defining JAK1 binders, and −9 kcal/mol as the threshold for nonbinders. Receptor-specific docking profiles further supported this choice, with mean ligand scores ranging from −12.45 to −13.95 kcal/mol across all structures and nearly all cocrystal ligands docking at or below the −12 kcal/mol benchmark. Applying this validated protocol to the full compound library produced the docking-score distributions shown in Figure . We thank the reviewer for highlighting the importance of justifying the active and inactive docking score thresholds. To justify the docking score thresholds used to define active (≤−12 kcal/mol) and inactive (>−9 kcal/mol) compounds, which were initially derived from redocking of cocrystallized JAK1 ligands, we conducted sensitivity analyses using alternative active cutoffs of −11 and −10 kcal/mol across all eight JAK1 crystal structures. At the stringent −12 kcal/mol cutoff, the number of “gray-zone” compounds (scores between −12 and −9 kcal/mol) ranged between 10,297 and 28,600 compounds for different PDB structures. Relaxing the active cutoff to −11 and −10 kcal/mol progressively reduced the number of molecules to 5604–16,745 (−11 kcal/mol) compounds and 5168–12,583 (−10 kcal/mol) compounds, respectively. Though these cutoffs (−11 and −10 kcal/mol) increase the number of nominal actives, they also incorporate moderately binding and energetically ambiguous compounds, thereby reducing label stringency. In contrast, the ≤−12 kcal/mol cutoff consistently isolates the strongest binders while explicitly separating intermediate cases into a well-defined gray zone. Per-PDB docking score distributions (Supplementary File 2 and Table ) show clear separation between active, inactive, and gray-zone populations, supporting the robustness of the selected thresholds across all receptor conformations. Activity labels were carefully assigned to ensure balanced representation of binder and nonbinder molecules across both the training and test sets. The composition of these data sets including Morgan fingerprints, average docking scores and activity annotations, is available in the Zenodo data repository at 10.5281/zenodo.17310240. For the putative active class (≤−12 kcal/mol), most molecules clustered between −14 and −12 kcal/mol, whereas the putative inactive class (>−9 kcal/mol) primarily fell between −9 and −7 kcal/mol, with smaller tails toward stronger or weaker affinities. Care was taken while generating data sets to maintain class distribution in train (94,270 actives and 94,270 inactives) and test data (actives 40,005 and 40,005 inactives) sets to improve the model generalizability. Molecules in the intermediate or “gray zone” (−12 to −9 kcal/mol) were also represented, providing a nuanced activity spectrum. The largest fraction in both sets (∼167,000 in training and ∼72,000 in test) occupied the −8 to −13 kcal/mol range, ensuring comprehensive coverage of binding affinities. This balanced distribution, achieved through stratified clustering-based sampling, preserved both chemical diversity and binding potential across subsets. Collectively, these results demonstrate that integrating multiple receptor conformations with systematic redocking of cocrystal ligands enables the establishment of a scientifically grounded activity threshold and yields high-quality, well-balanced data sets suitable for training DL models to predict JAK1 binders. To further assess potential scoring-function and receptor-conformation bias, a decoy-based enrichment analysis was performed using two independent JAK1 crystal structures (PDB IDs: 3EYH and 3EYG) and a chemically diverse, Murcko scaffold-controlled validation set comprising 100 experimentally characterized JAK1 inhibitors and 2000 putative inactive decoys. All compounds were docked using the same FRED/Chemgauss4 protocol applied throughout this study (Supplementary File 2 and Table ). Docking scores demonstrated good discrimination between actives and decoys (ROC-AUC = 0.75) with strong early enrichment (EF1% > 14), and comparable performance was observed across both receptor conformations (Figure S3). Together, these results indicate that the docking protocol captures biologically meaningful binding signals rather than artifacts of a single receptor structure or scoring function, thereby supporting its use for defining activity labels in the absence of experimental affinity measurements for screened compounds.

3.

3

Heatmap of docking score distribution for eight JAK1 cocrystal ligands across eight JAK1 crystal structures. Each cell in the 8 × 8 matrix represents the docking score of a ligand-structure pair, illustrating both self-docking and cross-docking performance. The diagonal (self-docking) and off-diagonal (cross-docking) elements exhibit mean docking scores of (−13.56 and −13.18 kcal/mol, respectively), indicating the consistent docking performance across the JAK1 conformations.

4.

4

Bar graph showing the distribution of active and inactive compounds in the training and test sets. Blue bars represent the training set, while orange bars represent the test set. The comparable proportions confirm balanced data set partitioning, minimizing class imbalance during model training and evaluation.

2. Generalization Performance of the GELU Model on Hard Test Sets Defined by Tanimoto Similarity Cutoffs of 0.3, 0.4, and 0.5 .

test_set accuracy sensitivity specificity precision F1_score MCC
hard test-0.30 0.90 0.93 0.88 0.90 0.91 0.81
hard test-0.40 0.96 0.97 0.95 0.92 0.94 0.91
hard test-0.50 0.97 0.97 0.97 0.92 0.95 0.93
a

The results highlight the model’s ability to retain predictive power under increasing levels of structural dissimilarity.

3.3. Deep Docking Model Based Virtual Screening

The deep docking classification models were constructed using DNN trained for 100 epochs with a batch size of 256. To improve generalization and reduce overfitting, L1 (1 × 10–5) and L2 (1 × 10–6) regularization were applied to the dense layers. Training dynamics were optimized using EarlyStopping (patience = 10) to halt convergence and ReduceLROnPlateau (factor = 0.5, patience = 10) to adaptively lower the learning rate during plateaus in validation loss. The Details of the model architecture and implementation codes are available on GitHub at https://github.com/computational-biology-lab/DeepDocking. Model performance was evaluated using six classification metrics as mentioned in methodology section. Results for GELU and Swish-activated DNNs (Figure and Table S2) demonstrated uniformly high performance across cross-validation, training, and independent test sets, with all metrics exceeding 0.98. All metrics except MCC were nearly identical between the two activations, underscoring their stability and generalization. MCC, a more stringent measure of classification reliability, provided the key differentiator: on the independent test set, the GELU-based model achieved an MCC of 0.9880, slightly higher than 0.9873 for the Swish-based model, indicating a marginal but meaningful edge in predictive balance. To rigorously evaluate extrapolative generalization, hard test sets were constructed using Tanimoto similarity on Morgan fingerprints (radius = 2, 1024 bits), without using activity labels at any stage of data splitting. For each candidate test molecule, the maximum similarity to any compound in the training set was calculated, and only molecules below predefined similarity thresholds were retained, ensuring strict chemical dissimilarity from the training data. Using thresholds of ≤0.50, ≤0.40, and ≤0.30 resulted in hard test sets comprising 517, 293, and 54 molecules, respectively, with positive fractions ranging from 0.27 to 0.54. All splits retained both active and inactive compounds, enabling robust performance evaluation under increasingly challenging extrapolative conditions. As summarized in Table , the GELU-based model maintained strong performance across Tanimoto thresholds ranging from 0.30 to 0.50. At a Tanimoto cutoff of 0.30, the model achieved an accuracy of 0.90, sensitivity of 0.93, specificity of 0.88, and an MCC of 0.81. Performance improved at higher thresholds, reaching 0.96 accuracy and an MCC of 0.91 at 0.40, and achieving the best overall balance at 0.50 (accuracy 0.97, sensitivity 0.97, specificity 0.97, MCC 0.93). Confusion matrices and data set sizes for each hard test set are provided in Table to enable detailed inspection of classification outcomes. To assess threshold-independent discrimination, ROC curves were generated for all three hard test sets. As shown in Figure S4, the model achieved AUROC values of 0.997, 0.996, and 0.983 for the Tanimoto ≤0.50, ≤0.40, and ≤0.30 test sets, respectively, demonstrating robust predictive performance even under substantial chemical dissimilarity. These results demonstrate that the GELU-based DNN effectively generalizes to structurally novel compounds, with consistently high MCC values (>0.81) confirming robust discrimination of actives from inactives in challenging chemical space. Importantly, the threshold for defining strong binders was derived from the redocking benchmark, where ligands reproducing crystallographic binding poses consistently exhibited docking scores ≤−12 kcal/mol. This cutoff was therefore adopted as the reference criterion for identifying high-affinity compounds in downstream screening. To evaluate potential circularity arising from docking-derived activity labels, an orthogonal validation strategy was employed using three independent JAK1 crystal structures (PDB IDs: 4E5W, 4EI4, and 6DBN) that were excluded from ensemble docking and model development. The top 500 model-ranked compounds and a control set of 500 lower-ranked compounds were independently docked into each held-out receptor conformation. Across all three structures, top-ranked compounds consistently achieved docking scores below −12 kcal/mol and did not populate the intermediate gray zone (−12 to −9 kcal/mol), whereas a small subset of lower-ranked compounds fell within this range. Violin plot (Figure S5) analyses revealed tighter score distributions and stronger binding tendencies for top-ranked compounds, while Spearman rank correlation analysis (Figure S6) confirmed preservation of relative compound ranking across unseen receptor conformations. These findings demonstrate that the model generalizes beyond its training labels and captures robust binding patterns, thereby addressing concerns related to circularity or information leakage. The validated GELU and SWISH based models were subsequently applied for prospective large-scale VS. A stringent probability threshold of 0.99, combined with an enrichment factor (EF) of 2, was employed to prioritize the most confident predictions. This pipe line yielded 131,730 compounds as hits, which were further subjected to chemical space analysis using reported JAK1 inhibitors. Docking scores for each structure are provided in the Supporting Information (Supplementary File 2 and Table ).

5.

5

Heatmap representation of performance metrics for the (A) GELU and (B) SWISH based DNN models across the training, validation, and test sets. Color intensity reflects the relative magnitude of each metric, enabling comparison of model performance across data sets.

3. Confusion Matrix Statistics for Hard Test Sets Constructed at Different Tanimoto Similarity Thresholds.

data set data set size threshold TP FP TN FN
hard test-0.30 54 Tanimoto ≤ 0.30 28 5 20 1
hard test-0.40 293 Tanimoto ≤ 0.40 99 10 182 2
hard test-0.50 517 Tanimoto ≤ 0.50 140 14 361 2

3.4. Chemical Space Analysis

The UMAP, a cutting-edge nonlinear dimensionality reduction algorithm, was employed to systematically assess the structural diversity and novelty of the top 10% VS hits relative to a comprehensive collection of 8356 reported JAK1 inhibitors curated from ChEMBL, BindingDB, and PubChem database. In this analysis, UMAP was applied to both the screened and reported sets, yielding the 2D projection depicted in Figure . The resultant chemical space map demonstrates a pronounced separation between the screened compounds, annotated as “ZINC” (blue circles), and the literature-reported inhibitors, marked as “REPORTED” (orange squares). The convex hulls encompassing each cluster further accentuate this divergence, with the screened compounds densely populating a more compact and discrete region, whereas the reported inhibitors exhibit a broader, more heterogeneous spread across chemical space. This marked spatial segregation highlights the ability of VS approach, guided by ML models and docking-based prioritization, to traverse previously underexplored regions of JAK1-relevant chemical space. The dense clustering of the screened compounds suggests a focused chemical series, potentially characterized by shared structural motifs and optimized physicochemical properties selected during computational triage. In contrast, the diversity observed among the reported inhibitors likely reflects the amalgamation of varied chemical scaffolds derived from disparate drug discovery efforts. Collectively, the UMAP visualization underscores the novelty and structural diversity of the identified screening hits, supporting the assertion that these molecules possess scaffold innovation distinct from existing JAK1 inhibitors. This differentiation is a key driver for advancing compounds with unique profiles that may offer advantages in terms of potency, selectivity, reduced resistance liabilities, and intellectual property landscape, thereby enhancing the probability of successful lead optimization and progression in the drug discovery pipeline. The Tanimoto similarity heatmap (Figure S7), constructed from pairwise molecular fingerprints of both reported JAK1 inhibitors and ZINC-screened compounds (top 10%), reveals a predominance of low similarity values across the data set, indicating substantial structural diversity within and between both groups. The lack of highly clustered regions, combined with the diagonal self-similarity observed at a value of 1.0, confirms limited scaffold overlap and minimal redundancy among the screened molecules and known inhibitors. These findings parallel the insights from the UMAP analysis, collectively underscoring that the VS process identified compounds occupying novel regions of chemical space. This pronounced chemical diversity not only validates the design strategy but also enhances the prospects for discovering innovative JAK1 binders with unique pharmacological profiles and reduced resistance or prior-art liabilities.

6.

6

UMAP plot showing the chemical space distribution of reported JAK1 inhibitors (orange points) and the top 10% hits identified from VS (blue points). Each point represents a compound, and the plot illustrates the overlap and coverage between known inhibitors and newly identified hits, highlighting the chemical diversity captured during screening.

In addition to the UMAP-based clustering analysis, a comprehensive examination of fundamental molecular descriptors was conducted to further evaluate and compare the chemical novelty and drug-likeness of the top 10% ZINC docked hits against the set of reported JAK1 inhibitors. As depicted in Figure A,B, the top docked hits displayed lower median values for both hydrogen bond acceptors (HBA) and hydrogen bond donors (HBD) relative to the reported inhibitors, indicating reduced polarity a property conducive to improved membrane permeability and oral absorption. This finding aligns with established pharmacokinetic principles, as molecules with fewer hydrogen-bonding groups tend to more readily cross cellular membranes. The lipophilicity profile, represented by the LogP distribution (Figure C), revealed that the ZINC hits exhibit slightly lower LogP values compared to reported inhibitors, suggesting enhanced aqueous solubility and potentially diminished nonspecific binding, while still preserving adequate membrane permeability. Assessment of molecular weight (MW) distributions (Figure D) demonstrated that both data sets are predominantly composed of compounds within the optimal range for oral bioavailability (generally <500 Da), with the ZINC hits showing a slightly narrower and lower MW range, further supporting favorable drug-like properties. A comparison of the Quantitative Estimate of Drug-likeness (QED) scores (Figure E) showcased that the top docked hits maintain comparable or marginally higher median QED values compared to reported inhibitors, reinforcing their alignment with key criteria commonly associated with successful drug candidates. Furthermore, evaluation of topological polar surface area (TPSA, Figure F) indicated that the majority of the ZINC compounds possess lower TPSA values, with most confined below the critical threshold of 140 Å2, which is predictive of efficient intestinal absorption and overall oral bioavailability. Taken together, this multifaceted analysis of HBA, HBD, LogP, MW, QED, and TPSA provides robust evidence that the top 10% docked hits not only explore novel and diversified chemical scaffolds but also consistently satisfy critical drug-likeness filters, positioning them as highly promising candidates for continued optimization and development as potential JAK1 binders.

7.

7

Comparative chemical space analysis of reported JAK1 inhibitors and the top 10% predicted JAK1 hits. Key physicochemical descriptors including (A) HBA, (B) HBD, (C) LogP, (D) MolWt, (E) QED, and (F) TPSA are shown to highlight differences in drug-like properties and potential bioavailability between the two sets.

3.5. Murko Scaffold Analysis

Murcko scaffold analysis was conducted to evaluate the structural diversity and novelty of the top 10% ranked compounds identified through our deep docking workflow. The Bemis–Murcko approach focuses on the core scaffolds of molecules by stripping away side chains, allowing for meaningful comparison based on fundamental ring systems and linkers. This type of analysis is especially useful in early stage drug discovery, where identifying new chemotypes is crucial for expanding chemical space and uncovering novel mechanisms of action. To assess novelty, we extracted Murcko scaffolds from the top hits and compared them to those found among 8356 known JAK1 inhibitors. As illustrated in Figure A, the Venn diagram reveals that 13 (13) scaffolds were common between the two sets, despite a high total number of unique scaffolds (2879 in reported inhibitors vs 7652 in JAK1 hits).

8.

8

Murcko scaffold comparison of reported JAK1 inhibitors with top 10% hits. (A) Overlapping Murcko scaffolds identified in top 10% and reported JAK1 inhibitors. (B) Unique scaffolds identified in both data sets. (C) Top 10 common Murcko scaffolds identified in both data sets.

This strikingly low overlap demonstrates the ability of our deep docking strategy to identify structurally distinct chemotypes beyond those previously associated with JAK1 inhibition. The unique scaffolds identified in each data sets was illustrated in Figure B. Additionally, the top 10 most frequent scaffolds from each group were analyzed and visualized in Figures C. The dominant scaffolds in the top hits show no structural resemblance to those prevalent among reported JAK1 inhibitors, further supporting the chemical novelty of our identified hits. Altogether, this analysis underscores the strength of our scaffold-based screening approach in uncovering novel chemical matter, providing new opportunities for JAK1-targeted drug design.

3.6. Ensemble Docking

To validate the robustness of the DNN model predictions and reconcile the use of a stringent selection threshold (P ≥ 0.99) with observed enrichment behavior, we analyzed early enrichment and precision-recall performance across multiple operating points. Because the independent test set contains equal numbers of active and inactive compounds (50% prevalence), the theoretical maximum enrichment factor is bounded at EFmax = 2.0. As shown in Table S3, the model consistently achieves enrichment factors approaching this theoretical limit across all early Top-N retrieval depths, indicating near-optimal early ranking of active compounds. Evaluation across probability thresholds (Table S4) further shows that enrichment remains stable while precision and recall follow the expected trade-off: lower thresholds improve recall, whereas the high-confidence cutoff (P ≥ 0.99) maximizes precision, making it well suited for large-scale virtual screening. To further assist the robustness, all 131,730 high-confidence hits (≥0.99) were docked against eight distinct JAK1 crystal structures, and compounds were reranked based on average docking scores. Of these, 74,878 molecules (57%) achieved scores ≤−12 kcal/mol the threshold used during model development demonstrating strong enrichment of high-affinity binders. Figure A shows the docking score distributions of the top 10% ligand hits across the eight JAK1 structures. The violin plots capture both spread and density, with scores spanning −8 to −18 kcal/mol. Most ligands clustered tightly between −12 and −16 kcal/mol, with a few outliers, highlighting differences in binding affinities across conformational states. Figure B presents the aggregated average docking scores for the top 10% ligands, revealing a dominant range between −13.5 and −15.5 kcal/mol, indicative of consistent strong binding across all structures. Several key observations emerged: (i) all JAK1 structures displayed median docking scores better than −12 kcal/mol, confirming strong binding potential of the top ligands. (ii) Broader score distributions with longer tails were observed across PDBs, suggesting accommodation of diverse ligand binding modes. (iii) Incorporating multiple protein conformations introduced flexibility-aware docking, enabling identification of ligands stable across structural ensembles. Overall, the consistent binding of top ligands across conformationally distinct JAK1 structures underscores the predictive robustness of the deep docking model. This approach effectively captures receptor flexibility and enriches for diverse, high-affinity chemotypes.

9.

9

(A) Docking score fluctuations of the top 10% JAK1 hits across multiple JAK1 crystal structures, illustrating the variability in binding affinity due to structural differences. (B) Distribution of the average docking scores for the top 10% hits, highlighting the overall binding trends and comparative potency within the selected compound set.

3.7. Post-Filtering Clustering and Diversity Analysis of Top-Ranked Docking Hits

Following the initial deep docking experiment, the top-ranked compounds from the docking experiment was subjected to rigorous structural and physicochemical filtering using the Blockbuster tool from the OpenEye software suite. This step aimed to retain molecules with optimal drug-like properties and synthetic tractability. Substructure-based exclusions were also applied to remove reactive or undesirable functionalities such as nitroso, acid halide, and alkyl halide groups, as well as known aggregators and metal-containing compounds. After applying these filters, 578 compounds were excluded, yielding 13,152 molecules from the top 10% docked set that fulfilled Blockbuster drug-like criteria ass detailed in the Supporting Information (Supplementary File 2 and Table 4). To assess and preserve structural diversity within this refined set, a two-step unsupervised learning approach was employed: UMAP for dimensionality reduction, followed by Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) as shown in Figure . This figure visualizes the chemical space coverage and diversity of the top 10% of docked molecules after applying stringent physicochemical and structural filters. “The large HDBSCAN cluster labeled −1 represents low-density regions of chemical space rather than a single dominant chemotype, capturing structurally diverse molecules that do not form high-confidence density clusters. Similarly, the largest defined cluster comprises multiple closely related scaffold families with overlapping physicochemical features. The intermixing of these clusters reflects a continuous chemical landscape rather than algorithmic artifacts. In contrast, smaller peripheral clusters correspond to distinct, well-separated chemotypes. Collectively, these results indicate that the filtering and ranking strategy retained both enriched scaffold families and substantial chemical diversity.” Using UMAP for dimensionality reduction, followed by HDBSCAN clustering, the 2D plot reveals that most molecules aggregate into several major clusters with dense cores, representing prevalent chemotypes or scaffold families. A number of smaller, peripheral clusters and dispersed points reflect less common or unique chemical entities. The distinct separation and distribution of clusters indicate a broad exploration of chemical space, with both core chemical scaffolds and structurally novel compounds retained among the top hits. These results confirm that the selection process maintained substantial molecular diversity, which is essential for robust drug discovery and avoiding redundancy in chemical matter.

10.

10

UMAP projection of 13,512 compounds filtered through the Blockbuster workflow, visualizing their chemical space in two dimensions. Subsequent HDBSCAN clustering identifies densely populated regions, highlighting structural similarities and distinct compound subgroups within the data set.

4. Conclusions

This study presents the first DL-driven VS framework tailored explicitly for JAK1 binding, uniting ensemble docking, receptor flexibility, and billion-scale chemical exploration to discover novel JAK1 binders. By leveraging chemically diverse, stratified data sets with high fidelity ensemble docking across eight crystal structures, the approach captures conformational dynamics often neglected in large-scale AI screens. The resulting DNN classifier demonstrated exceptional predictive power and identified structurally diverse high-affinity ligands, more than half of which maintained strong binding across all JAK1 conformations. Scaffold analysis revealed substantial novelty, with minimal overlap to known JAK1 inhibitors, highlighting the discovery of untapped chemical space. Further drug-likeness filtering and clustering identified distinct chemotypes with favorable physicochemical properties, underscoring the translational potential of these candidates. Beyond advancing JAK1-targeted drug discovery, this framework establishes a scalable, generalizable paradigm for AI-guided exploration of ultralarge chemical space. By integrating receptor flexibility, chemically informed data set design, and DL augmented prioritization, our approach addresses long-standing challenges of chemical redundancy, poor isoform selectivity, and limited structural diversity. Importantly, the methodological blueprint we describe is extensible to other kinase families and flexible protein targets, offering a transformative route for rational, data-driven drug discovery at unprecedented scale.

Supplementary Material

ao5c10773_si_001.pdf (794.5KB, pdf)
ao5c10773_si_002.xlsx (1.8MB, xlsx)

Acknowledgments

The authors gratefully acknowledge the High-Performance Computing Facility at the Institute of Life Sciences for providing the computational infrastructure and support essential for this research.

All data supporting this study are publicly available. The R implementation codes and model architectures are openly accessible via GitHub at https://github.com/computational-biology-lab/DeepDocking. The training, test, and hard test data sets, including molecular fingerprints and activity annotations, are deposited in the Zenodo repository at 10.5281/zenodo.17310240.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.5c10773.

  • All supplementary figures and tables referenced in the manuscript (PDF)

  • Four worksheets providing additional docking and data set details: per-PDB docking score distributions across eight JAK1 crystal structures, demonstrating the rationale for the selected active (≤−12 kcal/mol), inactive (>−9 kcal/mol), and gray-zone docking score thresholds through sensitivity analyses using alternative cutoffs (Worksheet 1); the result of a decoy-based enrichment analysis performed on two independent JAK1 structures (PDB IDs: 3EYH and 3EYG) using a Murcko scaffold-controlled validation set (Worksheet 2); docking scores for all compounds across each JAK1 receptor conformation (Worksheet 3); and compound filtering results, including exclusion statistics and the final selection of drug-like molecules satisfying Blockbuster criteria (Worksheet 4). Additional data set compositions, molecular fingerprints, docking scores, and activity annotations are available via the Zenodo repository (DOI: 10.5281/zenodo.17310240) (XLSX)

B.R.: Conceptualization, writingoriginal draft and editing, methodology, and results; G.N.: review and editing, graphical abstract; B.K.M.: review and editing; and A.D.: review and editing, supervision, and resource.

This work was supported by the Ministry of Earth Sciences (MoES), Government of India, under the Deep Ocean Mission (DOM) project (F. No. MoES/PAMC/DOM/181/2023 (E-14616). Additional funding was provided by the Council of Scientific and Industrial Research (CSIR), India.

This manuscript is an original work, with all data and methods reported accurately and in sufficient detail for reproducibility. Where prior work has been used, it has been properly cited and permissions obtained. The material has not been published previously and is not under consideration elsewhere. All authors have contributed substantially and accept full responsibility for the content.

The authors declare no competing financial interest.

References

  1. Wei T., Lu M., Yao S., Hong Y., Yang J., Zhang M., Yin Y., Han Y., Li Q., Wang Z., Wang Y., Tong Z., Zhou Y., Dai W., Yu Y., Sun S., Yang Y., Li N., Shi Z.. Insight into Janus kinases specificity: From molecular architecture to cancer therapeutics. MedComm-Oncol. 2024;3:e69. doi: 10.1002/mog2.69. [DOI] [Google Scholar]
  2. Rane S. G., Reddy E. P.. Janus kinases: components of multiple signaling pathways. Oncogene. 2000;19:5662–5679. doi: 10.1038/sj.onc.1203925. [DOI] [PubMed] [Google Scholar]
  3. Parveen S., Fatma M., Mir S. S., Dermime S., Uddin S.. JAK-STAT signaling in autoimmunity and cancer. Immunotargets Ther. 2025:523–554. doi: 10.2147/ITT.S485670. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Azhar, M. ; Khan, A. H. ; Nadeem, A. Y. ; Fatima, A. ; Al-Suhaimi, E. A. ; Shehzad, A. . Cellular receptors and cell signaling. In Cell Signaling; CRC Press, 2025; pp 46–63. [Google Scholar]
  5. Yamaoka K., Saharinen P., Pesu M., Holt V. E. III, Silvennoinen O., O’Shea J. J.. The janus kinases (JAKs) Genome Biol. 2004;5:253. doi: 10.1186/gb-2004-5-12-253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Purohit M., Gupta G., Afzal O., Altamimi A. S. A., Alzarea S. I., Kazmi I., Almalki W. H., Gulati M., Kaur I. P., Singh S. K., Dua K.. Janus kinase/signal transducers and activator of transcription (JAK/STAT) and its role in lung inflammatory disease. Chem. Biol. Interact. 2023;371:110334. doi: 10.1016/j.cbi.2023.110334. [DOI] [PubMed] [Google Scholar]
  7. Villarino A. V., Kanno Y., O’Shea J. J.. Mechanisms and consequences of Jak-STAT signaling in the immune system. Nat. Immunol. 2017;18:374–384. doi: 10.1038/ni.3691. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Morris R., Kershaw N. J., Babon J. J.. The molecular details of cytokine signaling via the JAK/STAT pathway. Protein Sci. 2018;27:1984–2009. doi: 10.1002/pro.3519. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Chaimowitz N. S., Smith M. R., Forbes Satter L. R.. JAK/STAT defects and immune dysregulation, and guiding therapeutic choices. Immunol. Rev. 2024;322:311–328. doi: 10.1111/imr.13312. [DOI] [PubMed] [Google Scholar]
  10. Ptacek J., Hawtin R. E., Sun D., Louie B., Evensen E., Mittleman B. B., Cesano A., Cavet G., Bingham C. O., Cofield S. S., Curtis J. R., Danila M. I., Raman C., Furie R. A., Genovese M. C., Robinson W. H., Levesque M. C., Moreland L. W., Nigrovic P. A., Shadick N. A., O’Dell J. R., Thiele G. M., Clair E. W. S., Striebich C. C., Hale M. B., Khalili H., Batliwalla F., Aranow C., Mackay M., Diamond B., Nolan G. P., Gregersen P. K., Bridges S. L., Srivastava A.. Diminished cytokine-induced Jak/STAT signaling is associated with rheumatoid arthritis and disease activity. PLoS One. 2021;16:e0244187. doi: 10.1371/journal.pone.0244187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Muromoto R., Oritani K., Matsuda T.. Current understanding of the role of tyrosine kinase 2 signaling in immune responses. World J. Biol. Chem. 2022;13:1. doi: 10.4331/wjbc.v13.i1.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Yasuda T., Fukada T., Nishida K., Nakayama M., Matsuda M., Miura I., Dainichi T., Fukuda S., Kabashima K., Nakaoka S., Bin B. H., Kubo M., Ohno H., Hasegawa T., Ohara O., Koseki H., Wakana S., Yoshida H.. Hyperactivation of JAK1 tyrosine kinase induces stepwise, progressive pruritic dermatitis. J. Clin. Invest. 2016;126:2064–2076. doi: 10.1172/JCI82887. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Furth P. A.. STAT signaling in different breast cancer sub-types. Mol. Cell. Endocrinol. 2014;382:612–615. doi: 10.1016/j.mce.2013.03.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Virtanen A., Spinelli F. R., Telliez J. B., O’Shea J. J., Silvennoinen O., Gadina M.. JAK inhibitor selectivity: new opportunities, better drugs? Nat. Rev. Rheumatol. 2024;20:649–665. doi: 10.1038/s41584-024-01153-1. [DOI] [PubMed] [Google Scholar]
  15. McLornan D. P., Khan A. A., Harrison C. N.. Immunological consequences of JAK inhibition: friend or foe? Curr. Hematol. Malig. Rep. 2015;10:370–379. doi: 10.1007/s11899-015-0284-z. [DOI] [PubMed] [Google Scholar]
  16. Wlassits R., Müller M., Fenzl K. H., Lamprecht T., Erlacher L.. JAK-Inhibitors-a story of success and adverse events. Open Access Rheumatol. Res. Rev. 2024:43–53. doi: 10.2147/OARRR.S436637. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Clarke B., Yates M., Adas M., Bechman K., Galloway J.. The safety of JAK-1 inhibitors. Rheumatology (Oxford) 2021;60:ii24–ii30. doi: 10.1093/rheumatology/keaa895. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Biggioggero M., Becciolini A., Crotti C., Agape E., Favalli E. G.. Upadacitinib and filgotinib: the role of JAK1 selective inhibition in the treatment of rheumatoid arthritis. Drugs Context. 2019;8:212595. doi: 10.7573/dic.212595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Lopez Perez K., López-López E., Soulage F., Felix E., Medina-Franco J. L., Miranda-Quintana R. A.. Growth vs Diversity: A Time-Evolution Analysis of the Chemical Space. J. Chem. Inf. Model. 2025;65:6788–6796. doi: 10.1101/2025.02.18.638937. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Karthikeyan A., Priyakumar U. D.. Artificial intelligence: machine learning for chemical sciences. J. Chem. Sci. 2022;134:2. doi: 10.1007/s12039-021-01995-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Wang Z., Sun L., Xu Y., Liang P., Xu K., Huang J.. Discovery of novel JAK1 inhibitors through combining machine learning, structure-based pharmacophore modeling and bio-evaluation. J. Transl. Med. 2023;21:579. doi: 10.1186/s12967-023-04443-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Yang Z., Tian Y., Kong Y., Zhu Y., Yan A.. Classification of JAK1 inhibitors and SAR research by machine learning methods. Artif. Intell. Life Sci. 2022;2:100039. doi: 10.1016/j.ailsci.2022.100039. [DOI] [Google Scholar]
  23. Bu Y., Gao R., Zhang B., Zhang L., Sun D.. CoGT: ensemble machine learning method and its application on JAK inhibitor discovery. ACS Omega. 2023;8:13232–13242. doi: 10.1021/acsomega.3c00160. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Cavasotto C. N., Abagyan R. A.. Protein flexibility in ligand docking and virtual screening to protein kinases. J. Mol. Biol. 2004;337:209–225. doi: 10.1016/j.jmb.2004.01.003. [DOI] [PubMed] [Google Scholar]
  25. Lee J., Hao Nguyen C., Mamitsuka H.. Beyond rigid docking: deep learning approaches for fully flexible protein-ligand interactions. Brief. Bioinform. 2025;26:bbaf454. doi: 10.1093/bib/bbaf454. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Gentile F., Yaacoub J. C., Gleave J., Fernandez M., Ton A.-T., Ban F., Stern A., Cherkasov A.. Artificial intelligence-enabled virtual screening of ultra-large chemical libraries with deep docking. Nat. Protoc. 2022;17:672–697. doi: 10.1038/s41596-021-00659-2. [DOI] [PubMed] [Google Scholar]
  27. Baidya A. T., Goswami A. K., Das B., Darreh-Shori T., Kumar R.. AI-enabled ultra-large virtual screening identifies potential inhibitors of choline acetyltransferase for theranostic purposes. ACS Chem. Neurosci. 2024;15:4156–4170. doi: 10.1021/acschemneuro.4c00361. [DOI] [PubMed] [Google Scholar]
  28. Tang M., Wen C., Lin J., Chen H., Ran T.. Discovery of novel A2AR antagonists through deep learning-based virtual screening. Artif. Intell. Life Sci. 2023;3:100058. doi: 10.1016/j.ailsci.2023.100058. [DOI] [Google Scholar]
  29. Miyao T., Jasial S., Bajorath J., Funatsu K.. Evaluation of different virtual screening strategies on the basis of compound sets with characteristic core distributions and dissimilarity relationships. J. Comput.-Aided Mol. Des. 2019;33:729–743. doi: 10.1007/s10822-019-00218-8. [DOI] [PubMed] [Google Scholar]
  30. Rogers D., Hahn M.. Extended-connectivity fingerprints. J. Chem. Inf. Model. 2010;50:742–754. doi: 10.1021/ci100050t. [DOI] [PubMed] [Google Scholar]
  31. Pandey, K. K. ; Shukla, D. . Stratified sampling-based data reduction and categorization model for big data mining. In Int. Conf. Commun. Intell. Syst.; Springer, 2019; pp 107–122. [Google Scholar]
  32. Williams N. K., Bamert R. S., Patel O., Wang C., Walden P. M., Wilks A. F., Fantino E., Rossjohn J., Lucet I. S.. Dissecting specificity in the Janus kinases: the structures of JAK-specific inhibitors complexed to the JAK1 and JAK2 protein tyrosine kinase domains. J. Mol. Biol. 2009;387:219–232. doi: 10.1016/j.jmb.2009.01.041. [DOI] [PubMed] [Google Scholar]
  33. Su Q., Banks E., Bebernitz G., Bell K., Borenstein C. F., Chen H., Chuaqui C. E., Deng N., Ferguson A. D., Kawatkar S., Grimster N. P., Ruston L., Lyne P. D., Read J. A., Peng X., Pei X., Fawell S., Tang Z., Throner S., Vasbinder M. M., Wang H., Winter-Holt J., Woessner R., Wu A., Yang W., Zinda M., Kettle J. G.. Discovery of (2R)-N-[3-[2-[(3-methoxy-1-methyl-pyrazol-4-yl) amino] pyrimidin-4-yl]-1H-indol-7-yl]-2-(4-methylpiperazin-1-yl) propenamide (AZD4205) as a potent and selective janus kinase 1 inhibitor. J. Med. Chem. 2020;63:4517–4527. doi: 10.1021/acs.jmedchem.9b01392. [DOI] [PubMed] [Google Scholar]
  34. Zak M., Hurley C. A., Ward S. I., Bergeron P., Barrett K., Balazs M., Blair W. S., Bull R., Chakravarty P., Chang C., Crackett P., Deshmukh G., DeVoss J., Dragovich P. S., Eigenbrot C., Ellwood C., Gaines S., Ghilardi N., Gibbons P., Gradl S., Gribling P., Hamman C., Harstad E., Hewitt P., Johnson A., Johnson T., Kenny J. R., Koehler M. F. T., Bir Kohli P., Labadie S., Lee W. P., Liao J., Liimatta M., Mendonca R., Narukulla R., Pulk R., Reeve A., Savage S., Shia S., Steffek M., Ubhayakar S., van Abbema A., Aliagas I., Avitabile-Woo B., Xiao Y., Yang J., Kulagowski J. J.. Identification of C-2 hydroxyethyl imidazopyrrolopyridines as potent JAK1 inhibitors with favorable physicochemical properties and high selectivity over JAK2. J. Med. Chem. 2013;56:4764–4785. doi: 10.1021/jm4004895. [DOI] [PubMed] [Google Scholar]
  35. Hurley C. A., Blair W. S., Bull R. J., Chang C., Crackett P. H., Deshmukh G., Dyke H. J., Fong R., Ghilardi N., Gibbons P., Hewitt P. R., Johnson A., Johnson T., Kenny J. R., Kohli P. B., Kulagowski J. J., Liimatta M., Lupardus P. J., Maxey R. J., Mendonca R., Narukulla R., Pulk R., Ubhayakar S., van Abbema A., Ward S. I., Waszkowycz B., Zak M.. Novel triazolo-pyrrolopyridines as inhibitors of Janus kinase 1. Bioorg. Med. Chem. Lett. 2013;23:3592–3598. doi: 10.1016/j.bmcl.2013.04.018. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ao5c10773_si_001.pdf (794.5KB, pdf)
ao5c10773_si_002.xlsx (1.8MB, xlsx)

Data Availability Statement

All data supporting this study are publicly available. The R implementation codes and model architectures are openly accessible via GitHub at https://github.com/computational-biology-lab/DeepDocking. The training, test, and hard test data sets, including molecular fingerprints and activity annotations, are deposited in the Zenodo repository at 10.5281/zenodo.17310240.


Articles from ACS Omega are provided here courtesy of American Chemical Society

RESOURCES