Abstract
Ligand-based virtual screening (LBVS) seeks strong early enrichment when searching ultra-large libraries, but practical screening often relies on 1D/2D descriptions while 3D information is expensive and uncertain due to conformer generation and alignment. We propose GNN-MA, a retrieval-style pairwise scoring model for query–candidate molecular pairs that uses molecular graphs as a unified representation. Built on intra-graph message passing, GNN-MA adds cross-graph attention to learn atom-level soft alignment that focuses on key substructures relevant to activity matching, and introduces a bond-to-atom semantic aggregation module to better exploit chemical bond cues for similarity scoring. The framework uses 2D molecular graphs derived from SMILES for retrieval-style matching and does not rely on explicit 3D conformational modeling or alignment. Experiments on DUD-E and LIT-PCBA show that GNN-MA achieves competitive overall discrimination (ROC-AUC) and, relative to its ablated variants, provides consistent gains in early-enrichment metrics (EF@1–5%) on DUD-E, while on LIT-PCBA the improvements are more target-dependent. The learned atom-level soft alignment also provides a qualitative interpretability cue in case studies. Throughput benchmarks suggest that GNN-MA is most suitable as a re-ranking/refinement model after a fast prefiltering stage.
Keywords: virtual screening, scaffold hopping, molecular representation, graph attention, early enrichment
1. Introduction
Ligand-based virtual screening (LBVS) is a widely used prioritization strategy in early-stage drug discovery [1,2,3,4,5,6]. Its goal is to identify potential actives as early as possible from ultra-large candidate libraries [7,8,9,10,11]. Because experimental follow-up typically covers only a small top-ranked subset of compounds (often 0.1–5%), early enrichment often reflects practical screening value better than overall ranking performance [2,12,13,14,15,16]. Meanwhile, inputs in real screening workflows are frequently incomplete and heterogeneous: in most cases only 1D/2D descriptions are available, and 3D-related information is not always accessible; even when it is, conformer generation, conformer selection, and spatial alignment introduce extra cost and uncertainty [13,17,18,19,20,21,22]. Therefore, in this work we focus on a 2D graph-based formulation that avoids explicit 3D conformational modeling while still aiming at stable early enrichment under realistic screening conditions [1,2,6,7,8,9,10,11,15,16].
Existing LBVS methods can be roughly grouped into three categories [1,2,5,6,23]. The first category consists of traditional approaches built on 2D fingerprints and similarity measures, such as MACCS [24] and ECFP4 [25], which are computationally efficient and easy to scale; however, these representations inevitably compress structural information and struggle to explicitly capture fine-grained correspondences between a query and a candidate molecule, so early enrichment can be limited when the two molecules differ substantially in scaffold, or when activity is driven by only a few critical local fragments [2,6,23,24,25,26,27]. The second category includes 3D shape- or pharmacophore-matching methods (e.g., PheSA [19] and ROSHAMBO [20]), which can describe spatial similarity more directly; yet they typically rely on upstream procedures such as conformer generation and alignment, increasing computational cost and making performance sensitive to conformer quality and input settings, which in turn restricts their stable use on ultra-large libraries or in scenarios where 3D information is missing [9,12,13,17,18,19,20,28]. The third category comprises deep learning approaches to molecular representation learning [29,30,31,32,33,34,35,36,37,38,39,40], especially graph neural networks (GNNs) and models that incorporate 3D geometric information (e.g., SchNet and DimeNet++), which can learn richer structural semantics; nevertheless, many of these methods are primarily designed for single-molecule property modeling and pay insufficient attention to the query–candidate retrieval-style matching required by LBVS [3,41,42,43,44,45,46]. As a result, there is still clear room to improve how we explicitly model local correspondences and cross-molecule interactions—without relying on rigid alignment—so that the learning objective better serves early-enrichment-oriented screening and provides chemically meaningful rationales for prioritization [1,2,23,26,27,47,48,49].
To address these issues, we explicitly formulate LBVS as a query–candidate pairwise scoring-and-ranking problem and propose GNN-MA. The method uses the molecular graph as a unified representation: it first learns structural semantic embeddings via intra-molecular message passing, and then introduces a cross-graph attention mechanism to model inter-molecular interactions explicitly at the atom level. This enables atom-wise cross-molecule matching that highlights the key substructures most responsible for activity transfer and ranking decisions, providing a qualitative cue for interpreting which local fragments contribute to the matching score in case studies [3,26,44,45,46,48,49,50]. Rather than claiming novelty from the use of cross-graph attention alone, our main technical contribution lies in adapting alignment-aware graph comparison specifically to LBVS. In particular, the proposed framework combines bond-aware representation enhancement through edge fusion and bond-to-atom aggregation with a ranking-oriented training objective that is explicitly designed to improve early enrichment in target-specific virtual screening. In addition, we enhance atom representations through semantic aggregation of chemical bond information, strengthening the expressiveness of similarity estimation [41,42,43,48,50]. Importantly, GNN-MA operates on standard 2D molecular graph representations and does not depend on rigid 3D alignment procedures [2,6,13,17,18,19,20,42,43].
The main contributions of this work are as follows:
We cast ligand-based virtual screening as a retrieval-style query–candidate pairwise scoring task, aligning model learning more closely with early-enrichment-oriented screening objectives.
We propose an alignment-aware graph matching mechanism based on cross-graph attention to capture atom-level correspondences between molecules in a fully 2D representation space.
We design a bond-aware and ranking-oriented learning framework by combining edge fusion, bond-to-atom aggregation, and a within-target ranking constraint for improved early retrieval.
We provide a practical empirical study on DUD-E and LIT-PCBA with macro-aware aggregation views, per-target analysis, and efficiency evaluation for shortlist re-ranking.
2. Results and Discussion
2.1. Overview of Evaluation
We evaluated GNN-MA on two widely used ligand-based virtual screening (LBVS) benchmarks, DUD-E and LIT-PCBA. Performance was assessed using ROC-AUC [51] for overall discrimination and early enrichment EF@k% (k = 1, 2, 5, 10, 20) for top-ranked retrieval quality.
To avoid ambiguous reporting and to address target-size imbalance, we report results under four complementary aggregation views (definitions in Section 3.6): (i) global average (pooled), (ii) macro-average, (iii) weighted macro-average, and (iv) macro-target statistics. Because global average results can be disproportionately influenced by a few large targets, the main text focuses on macro-average, weighted macro-average, and macro-target evidence, whereas the corresponding global average summaries, detailed per-target tables, and supplementary statistical results are provided in the Supporting Information.
2.2. Results on DUD-E
2.2.1. Overall Discriminative Performance (ROC-AUC)
Figure 1 compares ROC-AUC on DUD-E across six models under the macro-average and weighted macro-average views. GNN-MA achieves the best overall discriminative performance, while GNN-MA-intra shows a clear drop, indicating that cross-graph interaction contributes positively to discrimination. GMN [52], Siamese_GNN [53], and DeepChem provide strong learning-based baselines but remain below GNN-MA, whereas ECFP4 performs substantially worse. The close agreement between macro and weighted macro AUC indicates that the ranking of methods is stable under these macro-based summaries.
Figure 1.
ROC-AUC comparison for DUD-E under macro-average and weighted macro-average views for each model.
2.2.2. Early Enrichment Performance (EF@k%) on DUD-E
Table 1 and Table 2 show that GNN-MA delivers stable early enrichment performance on DUD-E in both weighted macro-average and macro-average views. Its advantage is not limited to a single cutoff: the model remains strong from EF@1% through EF@20%, indicating that the gain is sustained across practical shortlist sizes rather than confined to only the very top-ranked molecules. The gap relative to GNN-MA-intra is consistent across all reported cutoffs, supporting the contribution of cross-graph interaction to retrieval quality. Compared with GMN, Siamese_GNN, and DeepChem, GNN-MA stays at the top or within the top tier under both aggregation views, whereas ECFP4 remains clearly inferior.
Table 1.
DUD-E early enrichment EF@k% under the weighted macro-average view.
| Model | EF@1% | EF@2% | EF@5% | EF@10% | EF@20% |
|---|---|---|---|---|---|
| GNN-MA | 35.52 | 34.54 | 19.42 | 9.89 | 4.98 |
| GNN-MA-intra | 17.51 | 15.52 | 10.94 | 7.24 | 4.28 |
| DeepChem | 35.30 | 32.45 | 17.79 | 9.39 | 4.85 |
| GMN | 34.24 | 32.16 | 18.16 | 9.63 | 4.94 |
| Siamese_GNN | 33.39 | 31.39 | 18.04 | 9.60 | 4.93 |
| ECFP4 | 19.35 | 14.05 | 7.77 | 4.73 | 2.99 |
Table 2.
DUD-E early enrichment EF@k% under the macro-average view.
| Model | EF@1% | EF@2% | EF@5% | EF@10% | EF@20% |
|---|---|---|---|---|---|
| GNN-MA | 34.38 | 33.20 | 19.26 | 9.86 | 4.98 |
| GNN-MA-intra | 15.74 | 14.61 | 10.20 | 6.89 | 4.14 |
| DeepChem | 33.81 | 31.55 | 17.45 | 9.34 | 4.83 |
| GMN | 32.58 | 30.91 | 17.91 | 9.59 | 4.93 |
| Siamese_GNN | 32.03 | 30.53 | 17.75 | 9.57 | 4.93 |
| ECFP4 | 19.05 | 14.00 | 8.08 | 4.79 | 3.00 |
2.3. Results on LIT-PCBA
2.3.1. Discriminative Performance (ROC-AUC)
Figure 2 indicates that GNN-MA maintains competitive discriminative performance on the more challenging LIT-PCBA benchmark and retains a clear advantage over GNN-MA-intra. GMN, Siamese_GNN, and DeepChem also remain competitive, whereas ECFP4 performs less favorably. Given the substantial variation in target sizes in LIT-PCBA, we focus primarily on macro-based summaries in the main text.
Figure 2.
ROC-AUC comparison for LIT-PCBA under the macro-average and weighted macro-average views for each model.
2.3.2. Early Enrichment Performance (EF@k%) on LIT-PCBA
Table 3 and Table 4 indicate that on the more heterogeneous LIT-PCBA benchmark, the benefit of GNN-MA is most pronounced at very early cutoffs. Under both weighted macro-average and macro-average summaries, the largest margins over GNN-MA-intra appear at EF@1–5%, which is particularly relevant for top-priority retrieval in practical screening. This pattern is consistent with the stronger target-size imbalance of LIT-PCBA, where macro-based summaries are more informative than pooled results and performance is more target-dependent than on DUD-E. Relative to GMN, Siamese_GNN, and DeepChem, GNN-MA remains highly competitive across cutoffs, while ECFP4 again shows clearly weaker enrichment.
Table 3.
LIT-PCBA early enrichment EF@k% under the weighted macro-average view.
| Model | EF@1% | EF@2% | EF@5% | EF@10% | EF@20% |
|---|---|---|---|---|---|
| GNN-MA | 55.94 | 32.10 | 15.47 | 8.60 | 4.65 |
| GNN-MA-intra | 9.89 | 6.97 | 3.36 | 2.21 | 1.87 |
| DeepChem | 6.91 | 6.86 | 3.95 | 2.65 | 1.96 |
| GMN | 36.66 | 23.95 | 12.06 | 7.49 | 4.40 |
| Siamese_GNN | 33.50 | 20.68 | 11.02 | 6.58 | 3.88 |
| ECFP4 | 5.74 | 5.39 | 2.33 | 1.78 | 1.42 |
Table 4.
LIT-PCBA early enrichment EF@k% under the macro-average view.
| Model | EF@1% | EF@2% | EF@5% | EF@10% | EF@20% |
|---|---|---|---|---|---|
| GNN-MA | 44.55 | 29.64 | 14.95 | 8.21 | 4.57 |
| GNN-MA-intra | 6.98 | 4.95 | 2.9 | 2.3 | 1.84 |
| DeepChem | 4.61 | 4.46 | 3.26 | 2.14 | 1.64 |
| GMN | 33.73 | 21.82 | 12.44 | 7.47 | 4.31 |
| Siamese_GNN | 26.8 | 18.68 | 10.32 | 6.47 | 3.86 |
| ECFP4 | 3.62 | 2.73 | 2.29 | 1.84 | 1.51 |
2.4. Per-Target Consistency Analysis (Macro)
To further examine whether the observed performance gains could be consistently reflected across targets, we analyzed the per-target differences between GNN-MA and several baseline models, as illustrated in Figure 3. For each target, ROC-AUC and EF@1% were computed independently on the target-specific test set, and the differences between GNN-MA and the corresponding baselines were then calculated. The resulting ΔAUC and ΔEF@1% values therefore indicate how much GNN-MA improves or decreases performance relative to a baseline for individual targets.
Figure 3.
Per-target ΔAUC and ΔEF@1% distributions.
Figure 3A,B show the distributions of per-target improvements for DUD-E. where green denotes comparisons against DeepChem and the intra-graph variant (GNN-MA-intra), and blue denotes comparisons against GMN and Siamese_GNN. Figure 3A presents ROC-AUC improvements, and Figure 3B presents EF@1%
The distributions are generally shifted toward positive values when comparing GNN-MA with DeepChem and the intra-graph variant (GNN-MA-intra), indicating that the performance gains are broadly observed across targets rather than driven by a small number of cases. In contrast, the differences relative to GMN and Siamese_GNN are smaller, suggesting that these graph-pair baselines remain competitive while GNN-MA maintains favorable target-level consistency.
Figure 3C,D present the corresponding results for LIT-PCBA. Because target sizes in LIT-PCBA are highly imbalanced (see Supplementary Materials Table S2), per-target analysis is particularly informative. The distributions suggest a generally positive trend in EF@1% relative to the baselines, while ROC-AUC remains broadly competitive at the target level.
2.5. Efficiency and Scalability
In large-scale virtual screening, inference efficiency is as critical as accuracy. To quantify deployability in large-library retrieval scenarios, we measured throughput under different CPU/GPU settings and separated preprocessing time from model inference time (Table 5).
Table 5.
Throughput benchmark protocol and reporting scheme.
| Dimension | Level (Paper Wording) |
Notation/Unit | Reporting Convention (Journal-Ready Wording) |
|---|---|---|---|
| Pipeline scope | Forward-only | Forward-only | Measures model forward pass only (including necessary tensor transfer/synchronization), excluding data loading and graph construction; reflects the upper bound of model compute efficiency. |
| End-to-end | E2E | Includes molecule → graph construction/feature encoding + data loading + forward; reflects deployment-relevant wall-clock cost. | |
| Compound preparation (E2E) | Cache | Cache | Graph/features are precomputed and cached (e.g., files in the LMDB directory). E2E primarily covers loading + forward. |
| Online | Online | Graph construction/encoding is performed on the fly. E2E covers construction/encoding + loading + forward. | |
| Hardware | GPU | GPU | Reports GPU model, memory, FP16/AMP setting, batch size, num_workers, etc. |
| CPU | CPU | Reports CPU model, thread count, MKL/OMP setting, batch size, etc. | |
| Scalability outputs | 100 K/1 M time | T(105), T(106) | Fixed query: T(N) = N/throughput. Variable queries: T(Q,N) = Q × N/(pairs/s). |
| Per-pair latency | ms/pair | t-pair = 1/(pairs/s) (converted to milliseconds). |
The results show that (as shown in Table 6), in the Forward-only (pure forward inference) setting, GPU reached 280.2 pairs/s, far higher than CPU’s 9.1 pairs/s, indicating that the forward computation is highly accelerator-friendly. In the end-to-end (E2E) evaluation, enabling caching boosted GPU throughput to 231.6 pairs/s, notably higher than 161.4 pairs/s under the Online mode, suggesting that preprocessing and data pipeline overheads are key factors for E2E performance. On CPU, E2E throughput was about 8–9 pairs/s, close to its Forward-only results, implying that CPU E2E is mainly constrained by overall compute.
Table 6.
Throughput and scalability results (CPU/GPU, Forward-only vs. E2E, Cache/Online).
| Pipeline Scope | Compound Prep | Hardware | Throughput (pairs/s) |
Per-Pair Latency (ms/pair) |
100 K Time T (105) |
1 M Time T (106) |
|---|---|---|---|---|---|---|
| Forward-only | — | GPU | 280.2 | 3.5690 | 356.9 | 3568.9 |
| Forward-only | — | CPU | 9.1 | 110.0933 | 10,989.0 | 109,890.1 |
| E2E | Cache | GPU | 231.6 | 4.3183 | 431.8 | 4317.8 |
| E2E | Cache | CPU | 8.3 | 120.5607 | 12,048.2 | 120,481.9 |
| E2E | Online | GPU | 161.4 | 6.1955 | 619.5 | 6195.5 |
| E2E | Online | CPU | 8.17 | 122.4379 | 12,243.8 | 122,437.9 |
Overall, caching and batching can substantially improve end-to-end throughput, supporting the use of GNN-MA as a refinement model for shortlist re-ranking in large-library screening workflows.
2.6. Soft-Alignment Visualization
To qualitatively illustrate the matching patterns learned by the model, we present a soft-alignment case study using a query–candidate pair from the KAT2A target in DUD-E. Figure 4A visualizes attention weights between atom pairs: instead of spreading across the entire matrix, high weights concentrate in a few regions, suggesting the model tends to lock onto key fragments during matching.
Figure 4.
Soft-alignment visualization example: attention heatmap and atom-level matching links.
Figure 4B illustrates a structural alignment sketch: the highest-weight atom correspondences from Figure 4A are mapped back onto the molecular structures and annotated with links, where darker/more solid lines indicate higher weights. Different atom colors follow the standard molecular visualization convention and denote different element types. The links cluster around a few crucial local regions rather than scattering randomly. This pattern indicates that cross-graph attention highlights a small number of locally focused correspondences, which can be qualitatively inspected to understand which fragments contribute most to the matching score in this example.
2.7. Summary
In summary, this chapter provides a systematic evaluation of GNN-MA from five aspects: overall performance, early enrichment, cross-target robustness, inference efficiency, and qualitative alignment visualization. Taken together, the results indicate that GNN-MA combines solid discriminative performance with stronger very-early retrieval behavior, while the per-target and efficiency analyses further clarify where these advantages are most evident and how the model may be used in practical re-ranking settings. Meanwhile, per-target analysis and the visualization case study provide qualitative support for the model’s learned matching behavior. Throughput benchmarking also suggests practical deployability, and indicates that caching and batching can markedly improve end-to-end screening efficiency.
3. Materials and Methods
3.1. Study Overview and Problem Definition
We formulated ligand-based virtual screening (LBVS) as a query–candidate pairwise scoring task. Given a query molecule and a candidate molecule under the same target, we learned a differentiable scoring function so that active compounds in the candidate library receive higher scores and are ranked earlier.
Training data were constructed in pairs: ligand–active pairs were treated as positives, and ligand–decoy/inactive pairs as negatives. The model output a continuous matching score to rank candidates, and we reported ROC-AUC and early-enrichment metrics (EF@k%) during testing.
3.2. Datasets and Preprocessing
3.2.1. Datasets and Molecular Representation
We evaluated our method on two standard LBVS benchmarks: DUD-E [14] and LIT-PCBA [16]. Both datasets are organized by target, and each target contains active molecules and inactive molecules. DUD-E originally included 102 targets; after removing 2 targets that failed parsing, we retained 100 targets. LIT-PCBA contains 15 targets, and all were used in this study.
Molecules were represented as 2D heavy-atom graphs derived from SMILES. We did not impose a hard limit on the number of atoms; variable-sized molecules were handled via in-batch padding and masking.
3.2.2. Data Split and Evaluation Protocol
To ensure fairness, reproducibility, and leakage-free evaluation, we adopted a unified molecule-level splitting and evaluation protocol.
-
(1)
Molecule-level split. For each target t, molecules were split into train/validation/test subsets with a ratio of 8:1:1 at the molecule level. The split satisfies disjointness:
| (1) |
| (2) |
A fixed random seed was used and the same split was reused across all experiments. All models in this work were trained and evaluated under an identical partition.
-
(2)
Pair construction. Within each target, ligand–active pairs form positive samples and ligand–decoy/inactive pairs form negative samples. Pair construction strictly respected subset boundaries, preventing pair-level leakage. During evaluation, all molecules used for scoring were taken exclusively from the test split of each target. Pairwise scores were computed only between molecules within the same target-specific test subset, and the resulting scores were ranked to compute ROC-AUC and EF@k%.
-
(3)
Evaluation views. During testing, ranking was performed independently within each target. We report: (i) pooled metrics that aggregate per-target statistics; and (ii) per-target metrics and robustness analysis. Formal metric definitions are provided in Section 3.6.
3.3. Molecular Graph Representation and Features
We represented a molecule as a graph , where is the atom feature matrix (with as the node feature dimension), and is the bond feature tensor (with as the edge feature dimension). To reflect the incomplete and heterogeneous inputs in real-world LBVS, we adopted a dimension-agnostic feature organization strategy: we used 2D topology and atom/bond attributes as the input features throughout all experiments. The specific atom and bond features we used are listed in Table 7 and Table 8.
Table 7.
Atom features included.
| No. | Feature Category | Meaning/Examples | Notes (Task Relevance) |
|---|---|---|---|
| 1 | Atom type | C, O, N, S, etc. | Basic chemical composition that influences molecular properties and interaction patterns. |
| 2 | Topological position | On the main scaffold/on a side chain | Distinguishes the structural core from substituents, affecting conformational flexibility and local interactions. |
| 3 | Scaffold type | Aliphatic scaffold, aromatic scaffold, etc. | Reflects global structural framework differences, often associated with hydrophobicity, rigidity, and overall molecular shape. |
| 4 | Aromaticity | Whether the atom belongs to an aromatic ring | Aromaticity affects electron distribution and pi–pi interactions, which are frequently linked to bioactivity. |
| 5 | Ring membership | Whether the atom is part of any ring system | Ring membership impacts rigidity, geometry, and accessibility, thereby influencing structural matching in screening. |
| 6 | Pharmacophoric features | H-bond donor, H-bond acceptor, aromatic ring, hydrophobic site, etc. | Captures key structural motifs closely related to ligand–target interactions and biological activity. |
Table 8.
Bond features included.
| No. | Feature Category | Meaning/Examples | Notes (Task Relevance) |
|---|---|---|---|
| 1 | Bond type | Single, double, triple, aromatic | Determines connectivity strength and geometric/electronic structure; fundamental to molecular topology and reactivity. |
| 2 | Aromatic bond | Whether the bond is aromatic | Indicates conjugated/pi systems, affecting electron delocalization and molecular recognition (e.g., pi–pi interactions). |
| 3 | Conjugation | Whether the bond is conjugated | Related to electron delocalization, influencing polarity, stability, and interaction patterns. |
| 4 | In-ring bond | Whether the bond is within a ring | Ring bonds constrain molecular geometry and rigidity, affecting structural matching during screening. |
3.4. GNN-MA Model Architecture
As illustrated in Figure 5, GNN-MA comprises four components: (i) Intra-Graph Message Passing, (ii) Cross-Graph Attention for Soft Alignment, (iii) Edge Fusion and Bond-to-Atom Aggregation, and (iv) graph-level pooling followed by an MLP to output the final pairwise matching score.
Figure 5.
Architecture of GNN-MA.
3.4.1. Intra-Molecular Message Passing
To learn intra-molecular structural semantics, GNN-MA performs -th layer message passing on the query graph and the candidate graph separately. Given the (-1)th layer node embedding and the edge (bond) representation between nodes and , node aggregates messages from its neighborhood as follows:
| (3) |
The node embeddings are then updated via the following function:
| (4) |
Here, and are learnable mappings (e.g., multilayer perceptrons) that integrate neighborhood atom and bond information into the node representation. To improve training stability and representation capacity, we further apply residual connections and normalization between layers.
3.4.2. Cross-Graph Attention for Soft Alignment
Independent encoding of each molecule is typically insufficient to extract the pairwise matching cues required for retrieval. Therefore, GNN-MA builds upon the intra-molecular representations with a cross-graph attention module that explicitly captures atom- and bond-level interactions between the query and candidate, producing an explicit soft-alignment matrix.
Let and , denoting the atom representations of the two compounds after intra-molecular message passing, where and are the numbers of atoms and is the embedding dimension. We use scaled dot-product attention to compute the cross-graph relevance between the -th atom in the query molecule and the atom in the candidate molecule :
| (5) |
Here, and are learnable parameters, and is the scaling factor used in the attention mechanism.
For each query atom , we normalize across all candidate atoms by applying a softmax along the candidate atom dimension, yielding the soft alignment weights from query to candidate:
| (6) |
Based on these weights, query atom aggregates information from the candidate molecule to obtain a cross-graph contextual representation:
| (7) |
To enhance the symmetry of alignment and the complementarity of information, we similarly compute the reverse attention and obtained .
Analogously, we also construct cross-graph attention at the bond level: the correlation computation, normalization, and cross-graph aggregation follow the same procedure as the atom-level attention, except that the inputs are changed from atom representations to bond representations.
3.4.3. Edge Fusion and Bond-to-Atom Aggregation
The atom-level interaction representations produced by cross-graph attention provide evidence of soft correspondences between the query and candidate. Building on this, GNN-MA introduces a two-stage structural enhancement prior to node updating, namely edge-level fusion → edge-to-node aggregation: we first fuse information at the bond (edge) level to obtain enhanced edge representations, and then aggregate these enhanced edge features to adjacent atoms. In this way, bond semantics are injected into atom representations and subsequently contribute to the final scoring.
-
(1)
Edge-level Fusion:
For each chemical bond, we construct a fused edge representation by combining the cross-graph-updated representations of its two incident atoms with the original bond feature:
| (8) |
Here, is a learnable mapping that integrates node semantics and bond information to produce an enhanced bond representation.
-
(2)
Edge-to-node aggregation
After obtaining the fused edge representations, we aggregate information from incident edges for each atom to form an edge-aggregation vector:
| (9) |
Then, we inject the aggregated bond information into the atom representation through a learnable transformation and an update operation:
| (10) |
Here, is a learnable mapping that transfers bond-level information to atom-level representations, facilitating subsequent fusion with atom features.
This design is particularly relevant to LBVS, where subtle local bond environments can affect functional similarity even when global graph topology appears similar.
3.4.4. Similarity Scoring and Ranking
For graph-level pooling, we fuse the molecular representations obtained from intra-graph convolution, cross-graph attention, and edge-to-node aggregation via a residual combination, and then apply a readout function to obtain graph-level embeddings for scoring:
| (11) |
For pairwise scoring, we combine the two graph-level embeddings and feed them into a multilayer perceptron to obtain the final matching score:
| (12) |
Here, || denotes feature concatenation.
First, the model outputs a matching score for each query–candidate molecular pair. During training, we used the binary cross-entropy loss as the primary supervision signal for active/decoy classification. To further emphasize early enrichment, we added a within-target batch-wise ranking constraint. Specifically, each mini-batch was constructed from molecules belonging to the same target. Within the current mini-batch, negative pairs are ranked according to their predicted similarity scores, and the top K highest-scoring negatives are selected as the hardest negatives (K = 10). The ranking term then encourages the scores of positive pairs to be higher than those of these selected hard negatives. In this work, the final objective is defined as:
| (13) |
where Lbce denotes the binary cross-entropy loss and Lrank denotes the batch-wise pairwise ranking loss. To stabilize training, the ranking term is activated only after a short warm-up period: λ = 0 for the first two epochs and λ = 0.05 for the remaining epochs.
Compared with a pure pairwise classification objective, this ranking-oriented design better reflects the practical goal of LBVS, namely prioritizing active candidates at the top of the ranked list.
3.5. Baselines and Ablation Settings
To assess how a general 2D molecular graph classifier performs at this task, we chose DeepChem’s GraphConvModel [46,54], GMN [52], Siamese_GNN [53], and ECFP4 [25] as representative comparison methods. GraphConvModel serves as a general molecular graph classification baseline, learning structural representations directly from 2D molecular graphs and producing a binary classification score. GMN and Siamese_GNN provide pairwise graph-matching and similarity-learning baselines, respectively, while ECFP4 serves as a classical fingerprint-based method. On both DUD-E and LIT-PCBA, all baselines were evaluated using the same data split and evaluation pipeline as our main model, and we report both pooled and per-target AUC and EF@k%. We also built an ablated variant, GNN-MA-intra. It matches GNN-MA exactly in data splits, loss functions, and hyperparameter settings, but removes the cross-graph attention module, ensuring a fair and reproducible comparison.
3.6. Evaluation Metrics and Statistical Reporting
3.6.1. ROC-AUC
Overall ranking quality was evaluated using the area under the receiver operating characteristic curve [51] (ROC-AUC). ROC-AUC measures the probability that a randomly selected active molecule receives a higher predicted score than a randomly selected inactive molecule. It is threshold-independent and reflects global discriminative ability.
3.6.2. Early Enrichment: EF@k%
To assess early retrieval performance under realistic screening constraints, we report the enrichment factor EF@k% [55] for k = 1, 2, 5, 10, and 20. For target , let denote the number of test molecules and the number of test actives. After ranking molecules by predicted score in descending order, the top-k% cutoff is defined as:
| (14) |
Let denote the number of actives within the top molecules. The target-wise enrichment factor is defined as:
| (15) |
We report with particular emphasis on EF@1%, which reflects very early enrichment.
3.6.3. Aggregation Views Across Targets
Because targets can differ substantially in candidate-set size and active ratio, we report four complementary aggregation views:
-
(1)
Global average (pooled)
Predictions from all targets are pooled together before computing the metric. This view provides a single overall summary and aligns with pooled reporting conventions in many prior works. However, it can be dominated by a few large targets.
-
(2)
Macro-average
Metrics are computed independently for each target and then averaged with equal weights across targets. This view reflects cross-target robustness by treating each target equally.
-
(3)
Weighted macro-average
Per-target metrics are averaged with weights proportional to the number of evaluated molecules (or candidates) per target. This view lies between macro-average and global average and helps assess whether aggregate results are driven by a target-size imbalance.
-
(4)
Macro-target statistics
Metrics are analyzed at the target level to assess consistency across targets. We summarize per-target performance using distributions, win/loss/tie counts, and paired statistical tests on per-target differences.
3.6.4. Per-Target Improvement and Significance Testing
For a given metric M (ROC-AUC or EF@1%), we defined the per-target improvement over a baseline as
| (16) |
We estimated uncertainty of the mean improvement using a nonparametric bootstrap over targets (B = 20,000 resamples) and report percentile-based 95% confidence intervals.
To test whether improvements are systematically positive across targets, we applied a two-sided Wilcoxon signed-rank test to the set of per-target differences {}. We additionally report win/loss/tie counts, where a win indicates > 0, a loss indicates < 0, and a tie indicates = 0.
In the main text, EF tables are reported under macro-average and weighted macro-average views, whereas global average metrics are included as supplementary results.
3.7. Implementation Details and Throughput Benchmark
We implemented the model in PyTorch 2.6.0+cu126 and used Adam optimizer [56] for optimization. Key hyperparameters were: batch size = 32; epochs = 20; learning rate = 1 × 103; weight decay = 1 × 10−4; dropout = 0.2; warm-up = 2; model size: hidden dim = 64; message passing layers = 3; gradient clipping max-norm = 5.0.
Hardware and software environment: OS = Windows 10 (10.0.26100); Python 3.11.9; PyTorch 2.6.0+cu126; CUDA 12.6; GPU = NVIDIA GeForce RTX 4060 Ti; CPU = Intel Core i5-14600K; RAM = 32 GB.
For the ranking loss, we set K = 10 hardest negatives per batch and λ = 0.05 after a two-epoch warm-up period. These values were kept fixed across all experiments.
4. Conclusions
We propose GNN-MA, a soft molecular alignment approach for ligand-based virtual screening (LBVS). Using molecular graphs as a unified representation, GNN-MA learns atom-level correspondences via cross-graph attention and, together with intra-graph message passing and bond-to-atom aggregation, explicitly injects cross-molecule matching evidence into similarity scoring and provides a qualitative interpretability cue in case studies.
Experiments on DUD-E and LIT-PCBA support the effectiveness of the proposed alignment-aware design: cross-graph interaction improves retrieval quality relative to the intra-graph variant, particularly in the early-ranking regime, while the accompanying per-target and throughput analyses suggest that the model is best suited for shortlist refinement rather than brute-force large-library screening. Throughput benchmarking further suggests strong deployment potential, showing that caching and batching can substantially improve end-to-end screening efficiency.
Future work will focus on more robust early-ranking objectives, improved negative sampling strategies, and validation on additional external benchmarks.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/molecules31060991/s1, Figure S1: DUD-E cumulative contribution curve; Figure S2: LIT-PCBA cumulative contribution curve; Figure S3: Global ROC curves for all compared models on DUD-E; Figure S4: Global ROC curves for all compared models on LIT-PCBA; Table S1: DUD-E per-target statistics; Table S2: LIT-PCBA per-target statistics; Table S3: Win/loss/tie statistics; Table S4: Global ROC-AUC results; Table S5: Global EF results; Table S6: Statistical significance of per-target improvements; Table S7: Per-target ROC-AUC and EF@1% values for all models on DUD-E and LIT-PCBA.
Author Contributions
Conceptualization, K.L.; methodology, K.L.; software, K.L.; validation, Z.Z.; formal analysis, K.L.; investigation, K.L.; resources, K.L.; data curation, R.S.; writing—original draft preparation, K.L.; writing—review and editing, K.L.; visualization, K.L.; supervision, D.W.; project administration, K.L.; funding acquisition, D.W. All authors have read and agreed to the published version of the manuscript.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The benchmark datasets used in this study, DUD-E and LIT-PCBA, are publicly available. All data used to generate the figures and tables in this manuscript are available from the corresponding author upon reasonable request (if applicable). The training and evaluation code for GNN-MA is available in a public repository: https://github.com/BobbyLiukeling/GNN-MA (accessed on 13 March 2026).
Conflicts of Interest
The authors declare no conflict of interest.
Funding Statement
This research was supported by multiple funding sources: the Sichuan Higher Education Research Association Project (SZJJ2024YB-006).
Footnotes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
References
- 1.Tao C., Wang Y., Li X., Liu M. Advancements in Ligand-Based Virtual Screening through the Synergistic Integration of Graph Neural Networks and Expert-Crafted Descriptors. J. Chem. Inf. Model. 2025;65:4898–4905. doi: 10.1021/acs.jcim.5c00822. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Sciabola S., Torella R., Nagata A., Boehm M. Critical Assessment of State-of-the-Art Ligand-Based Virtual Screening Methods. Mol. Inform. 2022;41:e2200103. doi: 10.1002/minf.202200103. [DOI] [PubMed] [Google Scholar]
- 3.Du Y., Xing L., Zhang J., Chen Y. Graph Neural Networks in Modern AI-Aided Drug Discovery. Chem. Rev. 2025;125:10001–10103. doi: 10.1021/acs.chemrev.5c00461. [DOI] [PubMed] [Google Scholar]
- 4.Lim J.H., Sankararaman S., Iliopoulos-Tsoutsouvas C., Ernst M., Buelens F., Bissantz C., De Fabritiis G., Dominy B., Elliott P., Jiang D., et al. Modeling protein-ligand interactions for drug discovery in the era of deep learning. Chem. Soc. Rev. 2025;54:11141–11183. doi: 10.1039/d5cs00415b. [DOI] [PubMed] [Google Scholar]
- 5.Thaingtamtanha T., Ravichandran R., Gentile F. On the application of artificial intelligence in virtual screening. Expert Opin. Drug Discov. 2025;20:845–857. doi: 10.1080/17460441.2025.2508866. [DOI] [PubMed] [Google Scholar]
- 6.Zhang X., Zhang H., Zhang Y., Wu H., Li J., Wang X. A review of deep learning methods for ligand based drug virtual screening. Fundam. Res. 2024;4:715–737. doi: 10.1016/j.fmre.2024.02.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Zhou G., Rusnac D.-V., Park H., Canzani D., Nguyen H.M., Stewart L., Bush M.F., Nguyen P.T., Wulff H., Yarov-Yarovoy V., et al. An artificial intelligence accelerated virtual screening platform for drug discovery. Nat. Commun. 2024;15:7761. doi: 10.1038/s41467-024-52061-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Lyu J., Wang S., Balius T.E., Singh I., Levit A., Moroz Y.S., O’Donnell T.J., Tao P., Irwin J.J., Shoichet B.K. Modeling the expansion of virtual screening libraries. Nat. Chem. Biol. 2023;19:712–718. doi: 10.1038/s41589-022-01234-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Michino M., Beautrait A., Boyles N.A., Nadupalli A., Dementiev A., Sun S., Ginn J., Baxt L., Suto R., Bryk R., et al. Shape-Based Virtual Screening of a Billion-Compound Library Identifies Mycobacterial Lipoamide Dehydrogenase Inhibitors. ACS Bio Med Chem Au. 2023;3:507–515. doi: 10.1021/acsbiomedchemau.3c00046. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Roggia M., Natale B., Amendola G., Di Maro S., Cosconati S. Streamlining Large Chemical Library Docking with Artificial Intelligence: The PyRMD2Dock Approach. J. Chem. Inf. Model. 2024;64:2143–2149. doi: 10.1021/acs.jcim.3c00647. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Irwin J.J., Tang K.G., Young J., Dandarchuluun C., Wong B.R., Khurelbaatar M., Moroz Y.S., Mayfield J., Sayle R.A. ZINC-22: A Free Multi-Billion-Scale Database of Tangible Compounds for Ligand Discovery. J. Chem. Inf. Model. 2023;63:1166–1176. doi: 10.1021/acs.jcim.2c01253. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hawkins P.C.D., Skillman A.G., Nicholls A. Comparison of Shape-Matching and Docking as Virtual Screening Tools. J. Med. Chem. 2007;50:74–82. doi: 10.1021/jm0603365. [DOI] [PubMed] [Google Scholar]
- 13.Jiang Z., Xu J., Yan A., Wang L. A comprehensive comparative assessment of 3D molecular similarity tools in ligand-based virtual screening. Brief. Bioinform. 2021;22:bbab231. doi: 10.1093/bib/bbab231. [DOI] [PubMed] [Google Scholar]
- 14.Mysinger M.M., Carchia M., Irwin J.J., Shoichet B.K. Directory of useful decoys, enhanced (DUD-E): Better ligands and decoys for better benchmarking. J. Med. Chem. 2012;55:6582–6594. doi: 10.1021/jm300687e. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Romero E., Ajenjo A., Roos M., Lima A. Efficient decoy selection to improve virtual screening using machine learning models. J. Cheminformatics. 2025;17:165. doi: 10.1186/s13321-025-01107-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Tran-Nguyen V., Jacobus D.P., Rodriguez S., Madan R., Tessier P.M., Rapp M., Irwin J.J., Shoichet B.K. LIT-PCBA: An Unbiased Data Set for Machine Learning and Virtual Screening. J. Chem. Inf. Model. 2020;60:4275–4283. doi: 10.1021/acs.jcim.0c00155. [DOI] [PubMed] [Google Scholar]
- 17.Nguyen T.-T., Matsui T., Kaneko S., Nishi F., Shimizu K. Employing Molecular Conformations for Ligand-Based Virtual Screening with Equivariant Graph Neural Network and Deep Multiple Instance Learning. Molecules. 2023;28:5982. doi: 10.3390/molecules28165982. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Vázquez J., García R., Llinares P., Luque F.J., Herrero E. On the relevance of query definition in the performance of 3D ligand-based virtual screening. J. Comput. Aided Mol. Des. 2024;38:18. doi: 10.1007/s10822-024-00561-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Wahl J. PheSA: An Open-Source Tool for Pharmacophore-Enhanced Shape Alignment. J. Chem. Inf. Model. 2024;64:5944–5953. doi: 10.1021/acs.jcim.4c00516. [DOI] [PubMed] [Google Scholar]
- 20.Atwi R., Wang Y., Sciabola S., Antoszewski A. ROSHAMBO: An Open-Source Molecular Alignment and 3D Similarity Scoring Tool. J. Chem. Inf. Model. 2024;64:8098–8104. doi: 10.1021/acs.jcim.4c01225. [DOI] [PubMed] [Google Scholar]
- 21.Zhao H. The Science and Art of Structure-Based Virtual Screening. ACS Med. Chem. Lett. 2024;15:436–440. doi: 10.1021/acsmedchemlett.4c00093. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Lima A.G., Penteado A.B., de Jesus J.G., de Paula V.J.R., Ferraz W.R., Trossini G.H.G. Structure-Based Virtual Screening: Successes and Pitfalls. J. Braz. Chem. Soc. 2024;35:e-20240112. doi: 10.21577/0103-5053.20240112. [DOI] [Google Scholar]
- 23.López-Pérez K., Avellaneda-Tamayo J.F., Chen L., López-López E., Juárez-Mercado K.E., Medina-Franco J.L., Miranda-Quintana R.A. Molecular similarity: Theory, applications, and perspectives. Artif. Intell. Chem. 2024;2:100077. doi: 10.1016/j.aichem.2024.100077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Durant J.L., Leland B.A., Henry D.R., Nourse J.G. Reoptimization of MDL keys for use in drug discovery. J. Chem. Inf. Comput. Sci. 2002;42:1273–1280. doi: 10.1021/ci010132r. [DOI] [PubMed] [Google Scholar]
- 25.Rogers D., Hahn M. Extended-connectivity fingerprints. J. Chem. Inf. Model. 2010;50:742–754. doi: 10.1021/ci100050t. [DOI] [PubMed] [Google Scholar]
- 26.Coupry D.E., Pogany P. Application of deep metric learning to molecular graph similarity. J. Cheminformatics. 2022;14:11. doi: 10.1186/s13321-022-00595-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Huang J., Qian W., Li Y., Li S., Xu Y., Pan Y., Zhao S. GESim: Ultrafast graph-based molecular similarity calculation via von Neumann graph entropy. J. Cheminformatics. 2025;17:57. doi: 10.1186/s13321-025-01003-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Bolcato G., Heid E., Boström J. On the Value of Using 3D Shape and Electrostatic Similarities in Deep Generative Methods. J. Chem. Inf. Model. 2022;62:1388–1398. doi: 10.1021/acs.jcim.1c01535. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Correia J., Capela J., Rocha M. Deepmol: An automated machine and deep learning framework for computational chemistry. J. Cheminformatics. 2024;16:136. doi: 10.1186/s13321-024-00937-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Li D., Feng J., Zhou Y., Wang G., Chen X., Jin X., Jiang M., Yi H., Tang J. Drug–Target Interaction Prediction Using Graph Neural Networks and Contact Maps. J. Med. Chem. 2022;65:10691–10706. [Google Scholar]
- 31.Zariquiey F.S., Galvelis R., Gallicchio E., Chodera J.D., Markland T.E., De Fabritiis G. Enhancing Protein-Ligand Binding Affinity Predictions Using Neural Network Potentials. J. Chem. Inf. Model. 2024;64:1481–1485. doi: 10.1021/acs.jcim.3c02031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Sahu M.K., Nayak A.K., Hailemeskel B., Eyupoglu O.E. Exploring Recent Updates on Molecular Docking: Types, Method, Application, Limitation & Future Prospects. Int. J. Pharm. Res. Allied Sci. 2024;13:24–40. doi: 10.51847/une9jqjucl. [DOI] [Google Scholar]
- 33.To V.-T., Van Nguyen P.-C., Truong G.-B., Phan T.-M., Phan T.-L., Fagerberg R., Stadler P.F., Truong T.N. KGG: Knowledge-Guided Graph Self-Supervised Learning to Enhance Molecular Property Predictions. J. Chem. Inf. Model. 2025;65:9443–9458. doi: 10.1021/acs.jcim.5c01068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Li Y., Hou L.-X., Yi H.-C., You Z.-H., Chen S.-H., Zheng J., Yuan Y., Mi C.-G. MOLGAECL: Molecular Graph Contrastive Learning via Graph Auto-Encoder Pretraining and Fine-Tuning Based on Drug-Drug Interaction Prediction. J. Chem. Inf. Model. 2025;65:3104–3116. doi: 10.1021/acs.jcim.5c00043. [DOI] [PubMed] [Google Scholar]
- 35.Sypetkowski M., Wenkel F., Poursafaei F., Dickson N., Suri K., Fradkin P., Beaini D. On the Scalability of GNNs for Molecular Graphs. arXiv. 2024 doi: 10.48550/arXiv.2404.11568.2404.11568 [DOI] [Google Scholar]
- 36.Wang H.W. Prediction of protein-ligand binding affinity via deep learning models. Brief. Bioinform. 2024;25:bbae081. doi: 10.1093/bib/bbae081. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Yang C., Chen E.A., Zhang Y. Protein–Ligand Docking in the Machine-Learning Era. Molecules. 2022;27:4568. doi: 10.3390/molecules27144568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Abbas M.K.G., Rassam A., Karamshahi F., Abunora R., Abouseada M. The Role of AI in Drug Discovery. ChemBioChem. 2024;25:e202300816. doi: 10.1002/cbic.202300816. [DOI] [PubMed] [Google Scholar]
- 39.Zhang H., Saravanan K.M., Zhang J.Z.H. Structure-Based Virtual Screening and Molecular Dynamics Simulations of FTO Inhibitors for the Treatment of Acute Myeloid Leukemia. Molecules. 2023;28:4691. doi: 10.3390/molecules28124691. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Song B., Zhang J., Liu Y., Liu Y., Jiang J., Yuan S., Zhen X., Liu Y. A systematic review of molecular representation learning foundation models. Brief. Bioinform. 2026;27:bbaf703. doi: 10.1093/bib/bbaf703. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wang Z., Feng Z., Li Y., Li B., Wang Y., Sha C., He M., Li X. BatmanNet: Bi-branch masked graph transformer autoencoder for molecular representation. Brief. Bioinform. 2023;25:bbad400. doi: 10.1093/bib/bbad400. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Du W., Yang X., Wu D., Ma F., Zhang B., Bao C., Huo Y., Jiang J., Chen X., Wang Y. Fusing 2D and 3D molecular graphs as unambiguous molecular descriptors for conformational and chiral stereoisomers. Brief. Bioinform. 2022;24:bbac560. doi: 10.1093/bib/bbac560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Fang X., Liu L., Lei J., He D., Zhang S., Zhou J., Wang F., Wu H., Wang H. Geometry-enhanced molecular representation learning for property prediction. Nat. Mach. Intell. 2022;4:127–134. doi: 10.1038/s42256-021-00438-4. [DOI] [Google Scholar]
- 44.Ross J., Belgodere B., Chenthamarakshan V., Padhi I., Mroueh Y., Das P. Large-scale chemical language representations capture molecular structure and properties. Nat. Mach. Intell. 2022;4:1256–1264. doi: 10.1038/s42256-022-00580-7. [DOI] [Google Scholar]
- 45.Wang Y., Wang J., Cao Z., Farimani A.B. Molecular contrastive learning of representations via graph neural networks. Nat. Mach. Intell. 2022;4:279–287. doi: 10.1038/s42256-022-00447-x. [DOI] [Google Scholar]
- 46.Wu Z., Ramsundar B., Feinberg E.N., Gomes J., Geniesse C., Pappu A.S., Leswing K., Pande V. MoleculeNet: A benchmark for molecular machine learning. Chem. Sci. 2018;9:513–530. doi: 10.1039/C7SC02664A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Atsango A.O., Diamant N., Lu Z., Biancalani T., Scalia G., Chuang K.V. A 3D-Shape Similarity-based Contrastive Approach to Molecular Representation Learning. arXiv. 20222211.02130 [Google Scholar]
- 48.Méndez-Lucio O., Nicolaou C.A., Earnshaw B. MolE: A foundation model for molecular graphs using disentangled attention. Nat. Commun. 2024;15:9431. doi: 10.1038/s41467-024-53751-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Huang J., Qi X., Zhang Y., Liu Z., Yang Z. SG-ATT: A Sequence Graph Cross-Attention Representation Architecture for Molecular Property Prediction. Molecules. 2024;29:492. doi: 10.3390/molecules29020492. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Wang Y., Guo Y., Zhou Y., Li M., Zhao W., Liu H. KnoMol: A Knowledge-Enhanced Graph Transformer for Molecular Property Prediction. J. Chem. Inf. Model. 2024;64:7337–7348. doi: 10.1021/acs.jcim.4c01092. [DOI] [PubMed] [Google Scholar]
- 51.Fawcett T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006;27:861–874. doi: 10.1016/j.patrec.2005.10.010. [DOI] [Google Scholar]
- 52.Li Y., Gu C., Dullien T., Vinyals O., Kohli P. Graph Matching Networks for Learning the Similarity of Graph Structured Objects. arXiv. 2019 doi: 10.48550/arXiv.1904.12787.1904.12787 [DOI] [Google Scholar]
- 53.Altalib M.K., Salim N. Similarity-Based Virtual Screen Using Enhanced Siamese Deep Learning Methods. ACS Omega. 2022;7:4769–4786. doi: 10.1021/acsomega.1c04587. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Ramsundar B., Eastman P., Walters P., Pande V., Leswing K., Wu Z. Deep Learning for the Life Sciences: Applying Deep Learning to Genomics, Microscopy, Drug Discovery, and More. O’Reilly Media; Sebastopol, CA, USA: 2019. [Google Scholar]
- 55.Truchon J.-F., Bayly C.I. Evaluating Virtual Screening Methods: Good and Bad Metrics for the “Early Recognition” Problem. J. Chem. Inf. Model. 2007;47:488–508. doi: 10.1021/ci600426e. [DOI] [PubMed] [Google Scholar]
- 56.Kingma D.P., Ba J. Adam: A Method for Stochastic Optimization. arXiv. 20141412.6980 [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The benchmark datasets used in this study, DUD-E and LIT-PCBA, are publicly available. All data used to generate the figures and tables in this manuscript are available from the corresponding author upon reasonable request (if applicable). The training and evaluation code for GNN-MA is available in a public repository: https://github.com/BobbyLiukeling/GNN-MA (accessed on 13 March 2026).





