Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Mar 26;45(3):e70026. doi: 10.1002/minf.70026

Interpretable and Scalable Similarity Metrics for DNA‐Encoded Library Design Using Generative Topographic Mapping

Louis Plyer 1, Alexey A Orlov 1, Tagir N Akhmetshin 1, Erik Yeghyan 1, Fanny Bonachera 1, Dragos Horvath 1, Alexandre Varnek 1,✉
PMCID: PMC13019124  PMID: 41883078

Abstract

The growing number and size of DNA‐encoded libraries (DELs), together with the vast space of possible DEL designs, demand interpretable and scalable criteria for selecting which libraries to construct and screen against a given target. An ideal target‐focused DEL shows both strong similarity with an active reference compound collection and high intra‐DEL diversity. Chemography with Generative Topographic Mapping (GTM) was shown to be a promising approach for selecting DELs, offering both intuitive visualization and fast quantitative analysis scalable to thousands of DEL designs. This is achieved by defining each library by a “stand‐alone” vector, the comparison of which precludes costly pairwise inter‐molecular similarity calculations. However, the extent to which such “stand‐alone” (SA) approaches in general, and GTM‐derived SA metrics in particular, recover DELs that are reference‐proximal and chemically diverse as evaluated by conventional compound pair‐matching (CP) metrics in the initial descriptor space remains insufficiently characterized. In this article, the comparative analysis of the Morgan count fingerprint‐based chemical‐library similarity versus GTM‐derived metrics, using 100 diverse DEL subsets and a reference set of compounds tested against cyclin‐dependent kinase 2 (CDK2) from ChEMBL, was performed. GTM‐based SA metrics provide robust approximations for “gold standard” molecular descriptor space CP metrics for DEL selection: Spearman rank correlations fall in the 0.6–0.7 range. Our results demonstrate that GTM helps to identify DELs that best span the reference space according to same “gold standard” molecular descriptor space metrics: SA GTM‐driven rankings of libraries achieve enrichment factors at 5% (EF5%) of 4–12 (in terms of finding “gold standard” top libraries within the 5% best ranked by GTM)—always picking 2 out of the top 3 libraries. The accompanying two‐dimensional landscapes make intra‐ and interlibrary diversity visually accessible, supporting rapid, interpretable screening of alternative DEL designs. Collectively, these results position GTM as an efficient tool for chemical‐library similarity assessment and target‐focused DEL selection.

Keywords: chemical libraries, chemical space, chemography, dimensionality reduction, Generative Topographic Mapping


Generative Topographic Mapping (GTM) enables rapid, computationally efficient, and interpretable selection of DNA‐encoded libraries (DELs) with comprehensive reference chemical space coverage, supported by visual landscapes that elucidate intra‐ and interlibrary diversity.

graphic file with name MINF-45-e70026-g005.jpg

1. Introduction

In recent years, there has been a rapid growth in the use of combinatorial libraries, driven by advances in synthetic chemistry and automation that allow generating vast collections of compounds [1, 2]. These libraries greatly expand the range of chemical space (CS) that can be explored, potentially increasing the likelihood of finding molecules with useful biological activity [3]. A notable extension of this concept is the development of DNA‐encoded libraries [4] (DELs), where each compound is linked to a unique DNA sequence that serves as a molecular “barcode.” This approach enables the synthesis, pooling, and concurrent screening of libraries comprising up to billions of compounds in a single experiment, with high‐throughput sequencing used to identify binders after affinity selections. DEL technology provides an efficient, cost‐effective platform for hit discovery that complements conventional high‐throughput screening [4].

Although the number of DEL‐compatible reactions is limited, the broad availability of diverse chemical building blocks enables the design of thousands of unique DELs. This, in turn, necessitates clear criteria for selecting which DELs to construct and screen against a given target [5]. In practice, this design problem can be reframed as one of interlibrary comparison, wherein the objective is to assess the extent to which two libraries occupy overlapping regions of CS: a designed focused DEL should share the space of known bioactive compounds with a minimal size. Calculating interlibrary overlap or “similarity” is, however, not a trivial problem. Ideally, given a “reference” library L to be matched, the query library Q should propose at least one molecule within the “similarity radius” (molecular similarity score deemed sufficient for Neighborhood Behavior compliance [6]) around each reference compound [5]. However, in general, this “gold standard” class of compound pair‐based (CP) comparative approaches scales as |L| × |Q|, with |.| denoting respective library sizes [7, 8, 9, 10, 11, 12]. While numerous accelerations have been proposed—e.g., tree‐based indexes applied to low‐dimensional latent representations of compounds [13], and approximate nearest‐neighbor (ANN) methods for high‐dimensional fingerprints [14, 15]—these introduce their own drawbacks: tree structures degrade in higher dimensions, whereas ANN indexes trading recall for latency, and can incur substantial memory and index‐build overheads. More recently, instant similarity indices have been proposed to compute the mean [16] and dispersion [17] of pairwise similarities between large sets with linear scaling, avoiding explicit enumeration. However, these accelerations are derived for specific CP summary statistics and might require specific adaptation for other statistics. It is therefore practical to treat each library as a stand‐alone (SA) chemical object, represented by a single descriptor vector—typically the centroid in the chosen descriptor space—so that interlibrary comparisons become computationally instantaneous [8, 10, 11, 12, 18, 19, 20, 21]. Since the construction effort of the library vectors would scale proportionally to their sizes, such SA comparison approaches favorably scale up as |L| + |Q|—but are they accurate enough to reach the same conclusions as CP scores?

Among other methods, Generative Topographic Mapping (GTM) has emerged as a particularly well‐studied tool for analyzing chemical‐library similarity, including DELs [5, 22, 23, 24, 25, 26]. GTM is a nonlinear dimensionality‐reduction approach that can be viewed as a Gaussian mixture model (GMM) whose component means are constrained to lie on a nonlinear manifold obtained by mapping a regular two‐dimensional latent grid through radial‐basis functions. This provides a probabilistic framework and a convenient representation of chemical libraries as cumulated responsibility vectors, reflecting their density and property ‘traces’ on the map. Such vectors may be further normalized by the library size to produce normalized cumulated responsibility vectors, or centroids in responsibility space (RC). Additionally, responsibility vectors can be transformed into responsibility patterns (RP) through cell‐based clustering, offering an alternative one‐vector summary of a library. However, since the GTM involves nonlinear dimensionality reduction, it can introduce distortions that affect how well similarity metrics in the reduced space reflect those in the original high‐dimensional descriptor space [27, 28, 29]. While some studies report concordance between these spaces [26], others find moderate correlations [22], underscoring the need for systematic benchmarking against high‐dimensional baselines and cautious interpretation of map‐based similarity.

The goal of this article is to compare compound‐pair (CP) standard metrics with favorably‐scaling SA alternatives in the context of DEL selection. Their ability to prioritize libraries for synthesis and screening against a specific target by maximizing both similarity to—and coverage of—the target‐relevant CS will be assessed. This will be probed on hand of a set of 100 diverse DELs and a reference set of compounds tested against cyclin‐dependent kinase 2 (CDK2) from ChEMBL, using Morgan count fingerprints as initial molecular descriptor space. Two CP scores, mean pairwise similarity (MP) and coverage fraction (CF) and three SA scores, responsibility vector centroid (RC), responsibility pattern coverage (RPC), and weighted RP coverage (wRPC), are compared in terms of Spearman rank correlations. Alternatively, we compute enrichment factors (EF) measuring how often libraries top‐ranked by SA metrics contain those top‐ranked as similar by the standard CP procedures.

2. Methods

2.1. GTM

GTM is a nonlinear dimensionality reduction technique that models the structure of high‐dimensional data by embedding a flexible, continuous manifold into the input space [30]. This manifold is defined as a linear combination of Gaussian radial basis functions (RBFs) centered on a regular grid in a latent low‐dimensional (typically 2D) space. During unsupervised training, the manifold is fitted to a representative subset of compounds encoded as vectors of molecular descriptors by maximizing the log‐likelihood of the data under the GTM model. This process adjusts the manifold to traverse the densest regions of CS, effectively capturing its intrinsic structure. It can be considered as a constrained GMM in which the components share a common variance, and the means are constrained to lie on a smooth two‐dimensional manifold [30].

The mapping process begins with the insertion of a flexible, continuous surface—referred to as a manifold—into the multidimensional CS. The manifold is defined as a linear combination of GRBFs, which centers form a square grid in the latent space.

y(x)=WΦ (1)

where W is the trainable weight matrix, and Φ is a matrix of evaluations of each RBF on each node.

During unsupervised GTM training, the manifold is inserted into the densest region of the CS, guided by a representative training set of compounds, called the frameset. The frameset items are then projected onto it, and the probability density of a data point with coordinates tn being associated with a node k of coordinates xk on the manifold, expressed as

p(tn|xk, W, β)=(β2π)−D/2e−β2‖yk−tn‖2 (2)

where D is the dimensionality of the data space and β is the inverse of the variance of the distribution. If integrated over all nodes of the manifold, we obtain the probability density (or likelihood) of a compound n to be projected into the manifold:

p(tn| W, β)=1K∑k=1Kp(tn|xk, W, β) (3)

In practice, the natural logarithm of the likelihood is preferred.

LLhn(W,β)=ln (p(tn|W,β)) (4)

The LLh of the full frameset serves as an objective function to be maximized, to find the optimal shape of the manifold:

LLh(W,β)=∑n=1NLLhn(W,β) (5)

The EM (expectation–maximization) algorithm is used to find the optimal weight matrix W. The E‐step updates the posterior probabilities (i.e., responsibilities) R kn :

Rkn=p(tn|xk, W, β)∑kp(tn|xk, W, β) (6)

The M‐step updates the weight matrix and the variance β−1. The optimization runs until the convergence of the LLh.

2.2. Data Collection and Preprocessing

All possible DEL designs were prepared from 70 709 building blocks using eDesigner as described earlier [5]. This led to 27 371 individual designs. The designs (7988) with the number of compounds not less than 10 000 were kept. The MaxMin algorithm based on Hamming distance was then used to select the diverse pool of 100 DELs based on their libDesign vectors. Subsamples of the size 10 000 were prepared and standardized according to an in‐house protocol as previously described [31, 32]. Duplicate compounds were removed. The subset of 1780 compounds tested against cyclin dependent kinase 2 (CDK2, ChEMBL target ID CHEMBL301) was retrieved from ChEMBL version 33 database [33], prepared according to an in‐house protocol as previously described [31, 32]. In addition to that, the frame set of 25 000 random compounds was retrieved from ChEMBL for the frame set and prepared according to an in‐house protocol as previously described [31, 32].

2.3. Descriptor Calculation

Morgan count fingerprints [34] were used to encode chemical compounds. Morgan count fingerprints are circular fingerprints that capture the presence and frequency of substructures within a molecule by encoding its atomic environments up to a certain radius. Morgan count fingerprints were calculated using the RDKit (v.2022.09.5) library [35]. For Morgan count fingerprints radius 2 and fingerprint size 1024 were used.

2.4. Optimization of GTM Parameters

The GTM parameters were optimized on the frameset compounds encoded as count‐based Morgan fingerprints (see above). The optimization was performed using the Optuna library (v.4.0.0) with the Tree‐structured Parzen Estimator (TPE) sampler. The following ranges were considered: number of nodes (30–50), difference between the number of nodes and the number of basis functions (15–25), basis function width (1–10), and regularization coefficient (0.1–1000). Optimizing the difference between nodes and basis functions, rather than the absolute number of basis functions, ensures that the latter never equals or exceeds the number of nodes. The number of basis functions was then computed as the square of this difference. Optimization was performed to maximize entropy over the map, as described previously [26]. The five best diverse solutions were inspected visually using a 3D PCA projection of the manifold, which guided the final parameter selection. The manifold with the following parameters was used: 45 x 45 nodes, 15 x 15 basis functions, basis function width of 5 and regularization coefficient of 500.

2.5. Similarity Metrics

Two principal approaches that can be used for the analysis of chemical libraries are: one based on assessing similarity between libraries exploring the pairwise similarity between library compounds (introduced as “CP metrics”) and another that represents a chemical library as a stand‐alone single descriptor vector (introduced as “SA scores”).

2.5.1. CP Metrics

In addition to mean pairwise similarity, other metrics were incorporated, including mean maximum similarity measures (Smax), and weighted and nonweighted k‐nearest neighbor similarity/distance (wkNN and kNN) variants [9, 36, 37, 38, 39]. The S max metric identifies the mean closest distance between compounds of two libraries. Nearest neighbor‐based methods incorporate local neighborhood information. The number of k was chosen to be equal to 20. The equations for corresponding equations are given below.

Note that two distance measures were employed using the equations below: the Soergel distance (S), defined as the complement of the Tanimoto similarity (T), and the Euclidean distance (E). To facilitate a more conventional presentation of the results, Soergel distances were converted to Tanimoto similarity values (T = 1−S) throughout the text, unless stated otherwise.

2.5.2. MP Similarity/Distance

This metric computes the average distance between all possible pairs of compounds from libraries L and DEL. It reflects the overall dissimilarity between the two sets, with lower values indicating greater similarity:

MP(L,DEL)=1|L||DEL|∑i∈L∑j∈DELDij (7)

where L and DEL represent the libraries (molecule sets), |L| is the number of compounds in set L, and D ij is the distance (or dissimilarity) between compound i in L and compound j in DEL.

2.5.3. Mean Maximum Similarity (Mean Minimum Distance, Smax)

For each compound in library L, this metric identifies the distance to its single closest neighbor in set DEL. These minimum distances are then averaged across all compounds in L. A lower mean minimum distance suggests that DEL provides good coverage of L:

Smax(L,DEL)=1|L|∑i∈Lminj∈DELDij (8)

2.5.4. kNN Distance

Instead of considering only the single closest neighbor, this metric averages the distances between each compound in L and its k nearest neighbors in DEL:

kNN(L,DEL)=1|L|∑i∈L∑j∈Nk(i,DEL)minDij (9)

where N k (i, DEL) is the set of k‐nearest neighbors of compound i in DEL.

2.5.5. wkNN Distance

This metric extends kNN by weighting neighbors according to their distance: closer neighbors contribute more strongly than distant ones. The weights are assigned using an exponential function, normalized across the k neighbors. This approach emphasizes local similarity while still considering multiple neighbors:

wkNN(L,DEL)=1|L|∑i∈L∑j∈Nk(i,DEL)wijDij,wij=exp(−Dij)∑l∈Nk(i,DEL)exp(−Dil) (10)

2.5.6. SA (Centroid)‐Based Similarity Metrics

Centroid‐based comparisons were performed by summarizing each library into a single centroid point. The centroid was computed as the arithmetic mean of the descriptor vectors or responsibility vectors of all compounds (Equations (11) and (12)). Soergel and Euclidean distances were subsequently calculated between these centroids to provide a global measure of library similarity. Soergel distances were further converted to Tanimoto similarity values (T = 1−S) to provide a more conventional presentation of the results.

2.5.7. Similarity between Centroids in the Initial Descriptor Space (C) and Responsibility Space (RC)

The centroid in the initial descriptor space was calculated according to Equation (11) :

Centr→=1|L|∑nx→n (11)

where Centr→ represents the centroid of the descriptor vectors, x→n is a molecular fingerprint vector, n is the index of a molecule in the library, and |L| is the total number of data points in the library.

With responsibility vectors R instead of molecular descriptors x, the above equations returns the normalized cumulated responsibility vector RespCentr→ Equation (12):

RespCentr→=1|L|∑nr→n (12)

where RespCentr→ represents the centroid of the responsibility vectors, r→n is a responsibility vector, n is the index of a molecule in the library, and |L| is the total number of data points in the library.

Distance/similarity metrics are then applied to these representations, yielding the metrics denoted as C (fingerprint‐based) and RC (responsibility‐based).

2.5.8. Metrics for CS Coverage Estimation

2.5.8.1. CS Coverage (CF)

Chemical space coverage was estimated as the ratio of reference compounds that had at least one neighboring compound from the DEL library in the CS, given a predefined Tanimoto similarity threshold (Equation (13)):

CF(L,D|τ)=1|L|∑l∈Lδ(l,τ)where δ(l,τ)={1 if ∃ m ∈ D with T(l,m)>τ0 otherwise (13)

where CF(L,D|τ) denotes the fraction of reference compounds L that are covered by the DEL D at Tanimoto similarity threshold τ; |L| is the number of compounds in the reference set L; T(l, m) is the Tanimoto similarity of compounds l and m (computed on their fingerprints); and δ(l,τ) is the indicator function equal to 1 when the condition holds and 0 otherwise.

In this work, a similarity threshold of τ = 0.65 was chosen, as it provides a clear separation between libraries in terms of coverage and corresponds to values frequently used for similar types of molecular fingerprints [40, 41, 42, 43]

2.5.8.2. RPC

The RPC was computed as described previously [5, 23, 44]. RPC is defined as the fraction of RPs encountered in the reference library which are also present in the DEL. In RPC, all RPs, even those represented by a single molecule, are equally weighted.

2.5.8.3. wRPC

The wRPC was computed as described previously [5, 23]. In order to emphasize the relative importance of RPs incarnated by many compounds, the wRPC score weighs RP counts by the number of times they were seen in the reference library.

The similarity metrics and their corresponding acronyms used in this work are summarized in Table 1.

TABLE 1.

Overview of the similarity metrics used in this work and their corresponding acronyms.

Acronym Description
CP metrics
MP Mean pairwise similarity
S max Mean maximum similarity
kNN k‐Nearest neighborhood distance
wkNN Weighted k‐nearest neighborhood distance
CF Coverage fraction in the initial descriptor space
SA metrics
C Similarity or distance between centroid representations in the initial Morgan count fingerprints descriptor space used for similarity calculations.
RC Similarity or distance between normalized cumulated responsibility vectors.
RPC Responsibility pattern coverage
wRPC Weighted responsibility pattern coverage

Notes: If the metric name begins with the letter “R,” they are derived from responsibility vectors obtained with the GTM. Otherwise, the metrics are calculated based on Morgan fingerprint counts.

2.6. Chemical Library Diversity Analysis Using Consensus Diversity Plots

In addition to the aforementioned metrics, library diversity was assessed with the Consensus Diversity Plots (CDPs) framework [45], a previously described [29]. In brief, CDPs characterize diversity along two axes: scaffold diversity (y‐axis) and the distribution of fingerprint‐based chemical similarity (x‐axis). Scaffold diversity was quantified using F50—the fraction of scaffolds needed to recover 50% of the database [45]—computed from RDKit‐derived Bemis–Murcko scaffolds. F50 was then plotted against the median intralibrary similarity calculated using Morgan count fingerprints.

2.7. Workflow

The following workflow was used for the comparative analysis of the GTM and the initial descriptor space‐based selection of DELs best covering the reference space (Figure 1). Starting from 70 709 building blocks, DEL designs were pre‐enumerated with eDesigner (27 371 libDesigns) and filtered to retain only libraries with ≥10 000 compounds (7988 libDesigns). A diverse panel of 100 DELs was selected via MaxMin on libDesign vectors, and 10 000‐compound subsamples were prepared for each. The subsample size was chosen to reduce computational time, as previous work had shown that it preserves the GTM‐based similarity metrics of the full libraries [5, 23]. Duplicates were removed, and 25 000 compounds were randomly sampled from the union of deduplicated subsamples to define the DEL frameset. In parallel, 25 000 ChEMBL compounds and a CDK2 reference set (1780 compounds) were retrieved. A GTM was then trained on the combined 50 000‐compound frameset (25K DEL + 25K ChEMBL). The reference set and all libraries were projected onto the GTM manifold, and similarity/coverage metrics were computed and compared between the GTM space and the original descriptor space to guide DEL selection.

FIGURE 1.

FIGURE 1

Workflow for comparison of the GTM and the initial space similarity metrics for DEL selection used in this work. (1) Pre‐enumeration of DEL designs with eDesigner from the provided building blocks. (2) Filtering out designs with <10 000 compounds. (3) Selection a diverse panel of 100 DELs using MaxMin on libDesign vectors. (4) Selecting 25 000 random compounds for the frameset. (5) Retrieval of a frameset and a reference set from ChEMBL. (6) Building the GTM using a combined dataset (50 000 compounds DELs + ChEMBL). (7) Projecting the reference library and all the compounds to the manifold. (8) Calculate and compare metrics across spaces (GTM vs. the initial descriptor space).

3. Results

3.1. Chemical Library Similarity in the Descriptor Space and Its Correlation with the GTM‐Based Metrics

The chemical libraries used in this study consisted of 100 diverse DEL subsets. They varied in chemical diversity and as assed by the CDPs methodology in the interlibrary pairwise‐distance distributions (Supporting Information: Figure S1). Across the sets investigated, DELs built with broad, DNA‐robust transformations (e.g., amide coupling, reductive amination, DEL‐E‐G) are markedly more diverse than those driven by narrower heterocyclizations (e.g., Larock/oxadiazole; see Supporting Information: Figure S1 DEL‐A‐D, Supporting Information: Table S1).

Interlibrary similarity between the DELs and the CDK2 library was computed using the seven metrics in both the original descriptor space and the GTM latent space (Figure 2, Supporting Information: Figures S3,S4). In the initial space, the widely used mean pairwise similarity (MP) correlated with other metrics that quantify how well a DEL covers the reference space—namely, nearest‐neighbor‐based metrics (kNN, wkNN, Figure 2a; Supporting Information: Figures S2,S3) and CF. However, the correlation of MP with CF is imperfect, revealing a potential bias of mean pairwise similarity: it can be inflated when a focused DEL is highly similar to compounds from one region of the reference set, while failing to reflect the absence of analogs for other chemotypes in a more diverse reference subset. Indeed, numerous DELs showing rather high mean pairwise similarity do not provide any effective coverage under the chosen threshold (Figure 2b). For example, DEL‐C—ranked third by mean pairwise Tanimoto similarity—covered less than 10% of the reference library at the specified T c threshold (Figure 2b–d).

FIGURE 2.

FIGURE 2

(a) Spearman‐rank correlation matrix of interlibrary similarity metrics. Pixel color encodes the pair‐wise Spearman ρ (scale bar, right). Thick white grid‐lines mark the transition from metrics built on the GTM responsibility vectors (RC, RPC, wRPC) from those built on features (all others). Axis labels replace the full metric names; their meanings are: MP, mean‐pairwise similarity; S max, mean maximal similarity to the closest neighbor; kNN, k‐nearest‐neighbor similarity (k = 20); CF, % of covered compounds at threshold (T c = 0.65); wRPC, weighted RP‐GTM distance; C and RC, centroid‐based distance. (b,c) Scatter plots showing MP versus coverage (% covered, CF) and versus Tanimoto coefficient calculated for RC vectors, respectively. (d,e) Scatter plots for wRPC versus coverage (CF), MP (f) scatter plots for RC versus coverage (CF). Black dots are all observations; red markers (DEL‐A to DEL‐D) highlight examples of DELs discussed in the main text. For plots (b–d) the marker size scales with the number of DEL compounds that contribute to covering the reference‐library compounds (i.e., those that are neighbors of the reference compounds in the initial descriptor space at the chosen similarity threshold Tc = 0.65). Marker color encodes the MP Tanimoto similarity among those DEL compounds, with the color bar indicating low (gray) to high similarity (green). Tanimoto similarity was used as a metric for plots (a–f).

GTM‐derived metrics (designated with R further standing for responsibility vectors on which metrics were calculated) showed significant Spearman's rank correlation with their initial‐space counterparts, consistent with previous observations [22, 26] (Figure 2a; Supporting Information: Figure S2). The choice of distance measure (Euclidean or Soergel) also affected library rankings (Supporting Information: Figure S2). In the discussion that follows, emphasis is placed on Tanimoto‐based similarities for clarity, and the presentation is organized to contrast CP with SA.

As expected, GTM‐derived metrics exhibited stronger correlations with coverage in the initial descriptor space than with MP. For example, the three libraries with the largest coverage in the initial space (>30%; DEL‐A, DEL‐B, and DEL‐D) were among the top‐10‐ranked by both the RPC metric and the GTM centroid‐responsibility vector (Figure 2a,c,d). On contrary, the DEL‐C library, which ranked among the top three by MP similarity in the initial space, was not a top scorer either in the coverage metric (Figure 2b) or in the GTM‐based metrics (Figure 2c–f). While scaffolds vary across different DELs, the most frequent scaffolds within each individual DEL are rather similar and differ mainly in their substitution patterns (Supporting Information: Figure S4); moreover, DEL compounds show varying degree of similarity to their closest CDK2 neighbors in the initial CS (Supporting Information: Figure S5), suggesting that further optimization or diversification of building blocks may be beneficial. To assess the impact of library size on GTM‐based analysis, DELs A–G were subsampled up to one million compounds and RC, CF, and MP metrics were evaluated across varying subset sizes (Supporting Information: Figure S6). Subsets of 10 000 compounds adequately approximate RC and MP, while CF increases with subset size. Importantly, the relative ranking of DELs is preserved for subsamples of comparable size, suggesting that for large DELs, optimization of building blocks can be an effective strategy to reduce library size while maintaining maximal efficiency. On average, GTM‐derived metrics enabled selection of top‐scoring libraries according to both mean pairwise similarity and initial‐space coverage, with EF5% values of 4–8 (Table 2; Supporting Information: Table S2).

TABLE 2.

EF between evaluation and reference metrics at top‐5% threshold.

Eval metric Ref metric Spearman rho EF@5%
RC CF 0.69 12
MP 0.59 8
wRPC CF 0.67 8
MP 0.54 4

Notes: Evaluation metrics: Tc centroids for responsibility vectors (RC) and weighted RP coverage (wRPC). Reference metrics: coverage fraction (CF) and mean pairwise Tc (MP).

Discrepancies between the rankings were examined using GTM‐based visualizations for the four libraries (DEL‐A–D): the three top‐scoring by MP and DEL‐D, which ranked among the top three for both CF and the GTM weighted responsibility‐pattern metric. It should be noted that “centroid‐based” approaches in initial descriptor space and GTM responsibility spaces are conceptually very different. In initial descriptor space, this amounts to comparing the “average” descriptor vectors of two libraries—in case of the herein used Morgan fingerprints (circular fragment counts)—the centroid vectors encode the average level of presence of each fragment over the library. This potentially triggers a huge information loss—two libraries may converge toward similar average fragment population values, all while containing perfectly different chemotypes. To intuitively illustrate the problem, consider a library containing a mixture of aliphatic amines and phenols versus a library containing aliphatic alcohols and benzylamines. Both centroid fingerprints may reveal equivalent populations of “aromatic”, “aliphatic”, “hydroxyl,” and “amino” fragments—but this hides the fact that in this case fragment covariance is not at all the same within the two collections. Reversely, if compared libraries are combinatorial constructs based on a common “scaffold” with varying ornaments, the centroid in molecular descriptor space will be the center of a narrow cluster of highly related molecules—and, as such, a meaningful descriptor of the “average molecule” characterizing that library. This is the case for small, cyclization‐driven DELs—by contrast, large amide bond‐forming libraries are too vast to allow for an “average amide” to meaningfully represent them. By contrast, responsibility‐based centroid vectors are reflecting average GTM node populations—and these are typically associated to various global chemotypes [46]. RC approaches would have no problem to distinguish between the two example libraries discussed above. It is not surprising to witness the RC approach showing its highest correlation ranks to neighborhood‐sensitive CP approaches (S max, kNN, wkNN, and to a lightly lesser extent CF, which may be tributary to the empirical choice of the similarity cutoff τ).

However, as a general trend, Figure 2 also shows that fundamentally, all the metrics in the (SA and CP, initial or responsibility space confounded) correlate very well. This is not trivial, and basically means that, on the overall, the different DELs tend to each address very well defined and specific CS areas. Calculating the distances/overlaps between these appears to lead to rather reproducible rankings, no matter what exact formalism is employed. This may no longer be true in a more challenging setup with libraries that mutually overlap a lot—for example, distinct subsets of a same large DEL, based on partially or perfectly orthogonal sets of building blocks. In particular, cases where two libraries happen to converge around roughly a same centroid, all while having a low degree of overlap are rarely met in the current study—see following subchapter—and this may be a key factor explaining why CP and SA metrics in descriptor space are nearly perfectly correlated.

Furthermore, within this picture of overall agreement of various library similarity measures, the other clear trend expectedly and easily visible in Figure 2 is that the most significant loss in correlation is a gap of ~0.3 Spearman coefficient units between initial descriptor‐based and responsibility‐based scores. This is clearly the effect of the unavoidable information loss upon chemographic projection—which may be compensated by the benefits of using chemography to enhance our understanding of inter‐library relationships on hand of intuitive, readable maps.

3.2. Comparing Chemical Libraries on the GTM

In addition to providing quantitative metrics for estimating similarity between chemical libraries, GTM offers illustrative visualizations of the latent CS (Figure 3). The density landscape of the combined 100 DEL subsets reveals two areas (top left and top right) on the map which were not occupied by any DEL compounds. This is in line with previous studies showing that some zones of CS related, e.g., to natural products are hard to cover by DELs [5, 23, 24]. Accordingly, the CDK2 reference space is also not fully covered by DEL subsets (Figure 3).

FIGURE 3.

FIGURE 3

Density and class landscapes for the reference chemical library CDK2 and 100 diverse DEL subsets. The left panel (“CDK2”) shows density landscapes for the CDK2 reference library. The panel “100 DELS” shows density landscape for 100 DELs. The right panel (“CDK2 vs 100DELs”) illustrates a class landscape comparing the reference library to the 100 DELs subsets. Class probability indicates the proportion of 100 DELs (1) and CDK2 (2) compounds: red zones indicate DEL‐dominant regions, blue zones indicate reference‐dominant regions, and intermediate colors represent mixed zones.

Density and class landscapes provide an illustrative means to compare libraries and to contextualize them against a reference set. For example, Figure 4 summarizes how four DEL libraries (DEL‐A to DEL‐D) distribute across the CDK2 reference CS (see also Supporting Information: Figure S7). DEL‐A, DEL‐B, and DEL‐D each populate multiple, noncontiguous areas and extend into zones that are also represented by the reference, suggesting broader and more complementary coverage. In contrast, DEL‐C ‐ despite exhibiting MP comparable to the other subsets ‐ collapses into a single localized region (albeit a central one!) with limited spread, implying restricted diversity around a narrow chemotype. Taken together, landscapes indicate that DEL‐A, DEL‐B, and DEL‐D provide superior coverage of CDK2‐relevant space and are more promising starting points for follow‐up screening than DEL‐C.

FIGURE 4.

FIGURE 4

Density and class landscapes for 4 selected DEL subsets and the reference chemical library CDK2. The top panel shows the density landscapes for 4 selected DELs (see also the density landscape of the reference library on Figure 3). The bottom panel illustrates class landscapes comparing the reference library to the DEL subsets. Class probability indicates the proportion of DEL (1) and CDK2 (2) compounds: red zones indicate DEL‐dominant regions, blue zones indicate reference‐dominant regions, and intermediate colors represent mixed zones. DEL‐A, ‐B, and ‐D occupy multiple, noncontiguous regions, indicating broader coverage, whereas DEL‐C—despite comparable Tanimoto (T c) similarity to the reference—populates a single, localized zone. Given their combination of reference similarity and wider spatial coverage, DEL‐A, ‐B, and ‐D are preferred over DEL‐C for follow‐up selection.

4. Discussion

Dimensionality reduction techniques are widely used to transform complex, high‐dimensional chemical data into human‐comprehensible illustrative plots. Beyond their role in data visualization, these techniques find applications in diverse tasks, including de novo design [47, 48], combinatorial library optimization [23, 49], and library comparison [24]. In the context of library comparison, since the generation of low‐dimensional embeddings is accompanied by information loss [29], which may reduce the precision of similarity metrics, it is essential to assess the extent to which this loss influences the validity of conclusions drawn from similarity analyses.

In line with previous studies [22, 26], GTM‐derived similarity metrics were found to significantly correlate with descriptor‐based metrics, with Spearman correlation values in the range of 0.6–0.8. However, in this work an additional analysis was performed to assess not only the correlation between the two spaces, but also to determine which metrics are most informative for selecting the best DELs, since values can be biased when relying, for instance, solely on mean pairwise Tanimoto similarity. GTM‐based metrics were found to be particularly efficient for prioritizing DELs for follow‐up analysis, as they enable the selection of libraries that balance high similarity to the reference set with broad coverage of its CS, while also preserving intra‐DEL diversity.

The present study is based on a single biological target, CDK2. CDK2 was selected because it offers a chemically diverse set of compounds that can potentially be covered by the synthetic chemistry accessible through DEL libraries. Characteristics of the reference CS, including chemotype composition and overall diversity, may influence both compound‐pair–based and GTM‐derived similarity relationships; therefore, a systematic evaluation of ChEMBL target coverage across a range of DEL designs is planned for future studies.

GTM landscapes provide an effective tool for assessing chemical diversity within and across libraries. By examining the density, spread, and overlap of libraries on these maps, patterns of structural diversity become evident: regions of high and low density can be distinguished, structural clusters detected, and areas of CS that are over‐ or under‐represented identified. Importantly, this visualization also indicates whether interlibrary similarity is confined to a single region of CS or distributed across several regions. Thus, visual insights complement quantitative similarity metrics, providing a more nuanced understanding of interlibrary relationships, potential biases, and the structural uniqueness of different libraries. In addition to the metrics considered in this study, alternative set‐based approaches—particularly those grounded, for instance, in optimal transport theory—constitute a promising direction for future work. Notedly, GTM offers computational advantages for such methods by simplifying the implementation of distribution comparison techniques.

Collectively, these properties position GTM as a useful framework for chemical library similarity analysis. It enables rapid, coarse‐grained selection of chemical libraries that cover the target CS, which can subsequently be examined in detail using metrics computed in the initial descriptor space. As libraries expand in scale and complexity, its dual capacity for efficient interlibrary comparison and detailed visual insight makes it valuable for diversity assessment, rational library design, and prioritization.

5. Conclusions

This study demonstrates that GTM provides an efficient framework for DEL analysis. GTM‐derived metrics enable selection of DELs that balance high similarity to the reference set with broad coverage of its CS while obviating costly pairwise calculations and thereby potentially supporting thousands of DEL designs. GTM not only supports SA‐type scores in broad concordance with fingerprint‐based CP measures (Spearman ρ ≈ 0.6–0.8; EF5% = 4–8), but furthermore provides intuitive visualizations of library diversity—revealing clusters, coverage gaps, and distribution patterns beyond the reach of numerical metrics. It also helps identify DELs that optimally span the reference space, achieving EF5% values of 4–12 and recovering 2 of the top 3 libraries in the top 5% according to both coverage and mean pairwise similarity. Together, these results position GTM as an efficient tool for chemical‐library similarity analysis and DEL prioritization. By combining fingerprint‐based metrics with GTM‐derived measures, both rigorous numerical comparisons and visual insights can be achieved, enabling more informed diversity assessment, rational design, and selection of libraries in the era of ultralarge chemical collections.

Supporting Information

Additional supporting information can be found online in the Supporting Information section. Supporting Fig. S1: Consensus Diversity Plot (CDP) for 100 DNA‐encoded libraries (DELs) and CDK2 library. Each point is one library, placed by its median intra‐library molecular similarity (x‐axis; Morgan count fingerprints) and scaffold diversity (y‐axis; F50 from Bemis–Murcko scaffolds). Red‐outlined points mark the highlighted sets (DEL‐A‐G) and the CDK2‐focused library. Dashed lines show cohort medians for quick, relative comparison. Color (left to right panels) encodes % covered compounds at Tc threshold 0.65, mean pairwise Tc, and Smax, respectively. DEL‐A‐D were designed from far fewer building blocks than DEL‐E‐G (on the order of a few dozen to hundreds vs. many hundreds to a few thousand BBs across synthesis cycles, which is consistent with their higher median similarity and lower scaffold diversity see (see also Supplementary Table S1). Supporting Fig. S2: Spearman‐rank correlation matrix of inter‐library similarity metrics. Pixel colour encodes the pair‐wise Spearman ρ (scale bar, right). Thick white grid‐lines mark the transition from metrics built on the GTM responsibility vectors (R) from those built on features. Axis labels replace the full metric names; their meanings are: MP –mean‐pairwise distance; Smax –mean maximal similarity to the closest neighor; kNN –k‐nearest‐neighbour similarity (k=20); wkNN –weighted‐kNN similarity (k=20); CF –% of covered compounds at threshold (Tc = 0.65); RPC –responsibility pattern‐based coverage; wRPC –weighted RP‐GTM coverage, C –centroid‐based distance. E and T denote the Euclidean distance and the Tanimoto similarity (Tc), respectively. To avoid sign flips when mixing distance‐ and similarity‐based measures, correlations were computed on comparable scales: Euclidean distance was correlated with 1 ‐ Tc (a dissimilarity), whereas RPC, wRPC, and CF—being similarity measures—were correlated with Tc.Supporting Fig. S3: Pairwise scatter plots relating DEL coverage (CF) to similarity‐ and GTM‐based metrics; axes use abbreviated metric names. Axis labels replace the full metric names; their meanings are: MP –mean‐pairwise distance; Smax –mean maximal similarity to the closest neighor; kNN –k‐nearest‐neighbour similarity (k = 20); CF –% of covered compounds at threshold (Tc = 0.65); wRPC –weighted responsibility pattern coverage, C and RC –centroid‐based distance. (a) Smax vs % covered (CF). (b) C vs % covered (CF). (c) C vs RC (d) kNN vs % covered (CF). (e) kNN vs TM_MP. (f) wRPC vs kNN. Black dots are all observations; red markers (DEL‐A to DEL‐D) highlight the DELs discussed in the main text. For the scatter plots, marker area scales with the number of DEL compounds contributing to coverage of the reference‐library compounds (neighbors of the reference set in the original descriptor space at Tc = 0.65). Marker color encodes the mean pairwise similarity (Tc) among those DEL compounds (color bar, grey to green corresponding to low→high). The horizontal green bar in (e) indicates the TM_MP axis range; the bottom color bar reports the pairwise Tc for DEL compounds covering CDK2 (Tc ≥ 0.65). Supporting Fig. S4: Bemis–Murcko scaffolds were computed for all SMILES in DEL‐A‐D using RDKit. Molecules were grouped by unique scaffold and ranked by frequency of occurrence within each library. The top 3 most frequent scaffolds were selected, and all corresponding molecules were used for subsequent R‐group decomposition. Supporting Fig. S5: For each DEL‐A‐D, the closest DEL compound‐CDK2 compound pair was identified using the Tanimoto coefficient (Tc). Each panel shows the CDK2 compound (left) and its matched DEL compound (right), with the corresponding Tanimoto similarity reported in the legend. 13C denotes the attachment point to the DNA tag. Supporting Fig. S6: Investigation of DEL ranking as a function of a subsample size. Each DEL was enumerated up to 1 million compounds (with three independent repetitions) or fully enumerated when feasible. The RC, CF, and MP metrics were then calculated for progressively larger subsets of each DEL (the subset size is denoted by N). The coverage fraction normalized by the number of molecules in the DEL subset (CF/N) was included for comparative analysis. In addition to DELs A–G (Supplementary Figure S1), DEL‐I‐the fifth largest DEL investigated (5327280 compounds) was included in the comparison. Supporting Fig. S7: Cumulated density and class landscapes for DEL A‐D and the reference chemical library CDK2. The right panel shows the cumulated density landscapes for 4 selected DELs. The left panel illustrates class landscapes comparing the reference library to the DEL subsets. Class probability indicates the proportion of DEL (1) and CDK2 (2) compounds: red zones indicate DEL‐dominant regions, blue zones indicate reference‐dominant regions, and intermediate colors represent mixed zones. Supporting Table S1: DEL Building Blocks and Reactions Summary for selected DELs. Supporting Table S2: Overlap enrichment factors (EF) between evaluation and reference metrics at multiple top‐k thresholds. EF is computed as the observed overlap divided by the random expectation (expected overlap = k2/n), where k is the number of items in the top x% and n is the total number of items. EF@1%, EF@3%, EF@5%, and EF@10% denote enrichment for the overlap between the top x% of each pair of ranked lists. Values are rounded to two decimals. Evaluation metrics: Tc centroids for responsibility vectors RC) and weighted RP coverage (wRPC). Reference metrics: coverage fraction (CF) and mean pairwise Tc (MP).

Author Contributions

T.N.A., A.A.O., and L.P. contributed to the development of ChemographyKit. L.P. and A.A.O. performed machine learning, interpreted the data, and contributed to manuscript writing. F.B. and E.Y. prepared ChEMBL datasets used in this study. D.H., and A.V. provided overarching guidance, conceived and planned the research, and supervised the overall project. All authors critically analyzed the data, revised the manuscript, and approved the final version.

Conflict of Interest

The authors declare no conflicts of interest.

Supporting information

Supplementary Material

Acknowledgments

The authors thank Dr. Arkadii Lin and Dr. Yuliana Zabolotna for their contributions to the development of the initial versions of the functions for density and classification landscape building, as well as GTM‐derived metric calculation in ChemographyKit.

Open access publication funding provided by COUPERIN CY26.

Data Availability Statement

All datasets and Python code required to reproduce this study are available at https://github.com/Laboratoire‐de‐Chemoinformatique/SCOPE‐DEL. The GPU‐accelerated implementation of the GTM algorithm—along with utilities to build density and classification landscapes and to compute GTM‐derived similarity metrics—is provided in the companion library ChemographyKit at https://github.com/Laboratoire‐de‐Chemoinformatique/ChemographyKit. The datasets are also provided under https://zenodo.org/records/17158359.

References

  • 1. Warr W. A., Nicklaus M. C., Nicolaou C. A., and Rarey M., “Exploration of Ultralarge Compound Collections for Drug Discovery,” Journal of Chemical Information and Modeling 62, no. 9 (2022): 2021–2034, 10.1021/acs.jcim.2c00224. [DOI] [PubMed] [Google Scholar]
  • 2. Furka Á., “Forty Years of Combinatorial Technology,“ Drug Discovery Today 27, no. 10 (2022): 103308, 10.1016/j.drudis.2022.06.008. [DOI] [PubMed] [Google Scholar]
  • 3. Collie G. W., Clark M. A., Keefe A. D., et al., “Screening Ultra‐Large Encoded Compound Libraries Leads to Novel Protein–Ligand Interactions and High Selectivity,” Journal of Medicinal Chemistry 67, no. 2 (2024): 864–884, 10.1021/acs.jmedchem.3c01861. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Gironda‐Martínez A., Donckele E. J., Samain F., and Neri D., “DNA‐Encoded Chemical Libraries: A Comprehensive Review With Succesful Stories and Future Challenges,” ACS Pharmacology & Translational Science 4, no. 4 (2021):1265–1279, 10.1021/acsptsci.1c00118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Pikalyova R., Zabolotna Y., Volochnyuk D. M., Horvath D., Marcou G., and Varnek A., “Exploration of the Chemical Space of DNA‐Encoded Libraries,” Molecular Informatics 41, no. 6 (2022), e2100289, 10.1002/minf.202100289. [DOI] [PubMed] [Google Scholar]
  • 6. Horvath D., and Jeandenans C., “Neighborhood Behavior of in Silico Structural Spaces With Respect to in Vitro Activity Spaces−A Novel Understanding of the Molecular Similarity Principle in the Context of Multiple Receptor Binding Profiles,” Journal of Chemical Information and Computer Sciences 43, no. 2 (2003): 680–690, 10.1021/ci025634z. [DOI] [PubMed] [Google Scholar]
  • 7. Lessel U., and Lemmen C., “Comparison of Large Chemical Spaces,” ACS Medicinal Chemistry Letters 10, no. 10 (2019): 1504–1510, 10.1021/acsmedchemlett.9b00331. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Bae B., Bae H., and Nam H., “LOGICS: Learning Optimal Generative Distribution for Designing De Novo Chemical Structures,” Journal of Cheminformatics 15 (2023): 77, 10.1186/s13321-023-00747-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Lžičař M., and Gamouh H., “CHEESE: 3D Shape and Electrostatic Virtual Screening in a Vector Space,“ ChemRxiv (2024), 10.26434/chemrxiv-2024-cswth. [DOI] [Google Scholar]
  • 10. Rácz A., Dunn T. B., Bajusz D., Kim T. D., Miranda‐Quintana R. A., and Héberger K., “Extended Continuous Similarity Indices: Theory and Application for QSAR Descriptor Selection,” Journal of Computer‐Aided Molecular Design 36, no. 3 (2022): 157–173, 10.1007/s10822-022-00444-7. [DOI] [PubMed] [Google Scholar]
  • 11. Miranda‐Quintana R. A., Rácz A., Bajusz D., and Héberger K., “Extended Similarity Indices: The Benefits of Comparing More Than Two Objects Simultaneously. Part 2: Speed, Consistency, Diversity Selection,” Journal of Cheminformatics 13, (2021): 33, 10.1186/s13321-021-00504-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Miranda‐Quintana R. A., Bajusz D., Rácz A., and Héberger K., “Extended Similarity Indices: The Benefits of Comparing More than Two Objects Simultaneously. Part 1: Theory and Characteristics†.,“ Journal of Cheminformatics 13,  (2021): 32, 10.1186/s13321-021-00505-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Kirchoff K. E., Wellnitz J., Hochuli J. E., et al., “Utilizing Low‐Dimensional Molecular Embeddings for Rapid Chemical Similarity Search,” Advances in Information Retrieval 14609 (2024): 34–49, 10.1007/978-3-031-56060-6_3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Probst D., and Reymond J.‐L., “A Probabilistic Molecular Fingerprint for Big Data Settings,” Journal of Cheminformatics 10, no. 1 (2018): 66, 10.1186/s13321-018-0321-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Malkov Y. A., and Yashunin D. A., “Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs,” IEEE Transactions on Pattern Analysis and Machine Intelligence 42, no. 4 (2020):824–836. 10.1109/TPAMI.2018.2889473. [DOI] [PubMed] [Google Scholar]
  • 16. López‐Pérez K., Kim T. D., and Miranda‐Quintana R. A., “iSIM: Instant Similarity,” Digital Discovery 3, no. 6 (2024): 1160–1171, 10.1039/D4DD00041B. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Lopez‐Perez K., Zhao B., and Miranda‐Quintana R. A., “iSIM‐Sigma: Efficient Standard Deviation Calculation for Molecular Similarity,” Journal of Chemical Information and Modeling 65, no. 13 (2025): 6797–6808, 10.1021/acs.jcim.5c00894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Fernández‐de Gortari E., García‐Jacas C. R., Martinez‐Mayorga K., and Medina‐Franco J. L., “Database Fingerprint (DFP): An Approach to Represent Molecular Databases,” Journal of Cheminformatics 9 (2017): 9, 10.1186/s13321-017-0195-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Shemetulskis N. E., Weininger D., Blankley C. J., Yang J. J., and Humblet C., “Stigmata: An Algorithm To Determine Structural Commonalities in Diverse Datasets,” Journal of Chemical Information and Computer Sciences 36, no. 4 (1996): 862–871, 10.1021/ci950169. [DOI] [PubMed] [Google Scholar]
  • 20. Sánchez‐Cruz N., and Medina‐Franco J. L., “Statistical‐Based Database Fingerprint: Chemical Space Dependent Representation of Compound Databases,” Journal of Cheminformatics 10 (2018): 55, 10.1186/s13321-018-0311-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Polykovskiy D., Zhebrak A., Sanchez‐Lengeling B., et al., “Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models,“ Frontiers in Pharmacology 11 (2020): 565644, 10.3389/fphar.2020.565644. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Kireeva N., Baskin I. I., Gaspar H. A., Horvath D., Marcou G., and Varnek A., “Generative Topographic Mapping (GTM): Universal Tool for Data Visualization, Structure‐Activity Modeling and Dataset Comparison,” Molecular Informatics 31, no. 3–4 (2012), 301–312, 10.1002/minf.201100163. [DOI] [PubMed] [Google Scholar]
  • 23. Pikalyova R., Zabolotna Y., Horvath D., Marcou G., and Varnek A., “Chemical Library Space: Definition and DNA‐Encoded Library Comparison Study Case,” Journal of Chemical Information and Modeling 63, no. 13 ( 2023): 4042–4055, 10.1021/acs.jcim.3c00520. [DOI] [PubMed] [Google Scholar]
  • 24. Pikalyova R., Zabolotna Y., Horvath D., Marcou G., and Varnek A., “Meta‐GTM: Visualization and Analysis of the Chemical Library Space,” Journal of Chemical Information and Modeling 63, no. 17 (2023): 5571–5582, 10.1021/acs.jcim.3c00719. [DOI] [PubMed] [Google Scholar]
  • 25. Lin A., Beck B., Horvath D., Marcou G., and Varnek A., “Diversifying Chemical Libraries With Generative Topographic Mapping,” Journal of Computer‐Aided Molecular Design 34, no. 7 (2020): 805–815, 10.1007/s10822-019-00215-x. [DOI] [PubMed] [Google Scholar]
  • 26. Gaspar H. A., Baskin I. I., Marcou G., Horvath D., and Varnek A., “Chemical Data Visualization and Analysis With Incremental Generative Topographic Mapping: Big Data Challenge,” Journal of Chemical Information and Modeling 55, no. 1 (2015): 84–94, 10.1021/ci500575y. [DOI] [PubMed] [Google Scholar]
  • 27. Marx V., “Seeing Data as t‐SNE and UMAP Do,” Nature Methods 21, no. 6 (2024): 930–933, 10.1038/s41592-024-02301-x. [DOI] [PubMed] [Google Scholar]
  • 28. Huang H., Wang Y., Rudin C., and Browne E. P., “Towards a Comprehensive Evaluation of Dimension Reduction Methods for Transcriptomic Data Visualization,” Communications Biology 5 (2022): 719, 10.1038/s42003-022-03628-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Orlov A. A., Akhmetshin T. N., Horvath D., Marcou G., and Varnek A., “From High Dimensions to Human Insight: Exploring Dimensionality Reduction for Chemical Space Visualization,” Molecular Informatics 44 (2025): e202400265, 10.1002/minf.202400265. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Bishop C. M., Svensén M., and Williams C. K. I., “GTM: The Generative Topographic Mapping,“Neural Computation 10, no. 1 (1998): 215–234, 10.1162/089976698300017953. [DOI] [Google Scholar]
  • 31. Casciuc I., Zabolotna Y., Horvath D., Marcou G., Bajorath J., and Varnek A., “Virtual Screening With Generative Topographic Maps: How Many Maps Are Required?,” Journal of Chemical Information and Modeling 59, no. 1 (2019): 564–572, 10.1021/acs.jcim.8b00650. [DOI] [PubMed] [Google Scholar]
  • 32. Lin A., Horvath D., Marcou G., Beck B., and Varnek A., “Multi‐Task Generative Topographic Mapping in Virtual Screening,” Journal of Computer‐Aided Molecular Design 33, no. 3 (2019), 331–343, 10.1007/s10822-019-00188-x. [DOI] [PubMed] [Google Scholar]
  • 33. Gaulton A., Bellis L. J., Bento A. P., et al., “ChEMBL: A Large‐Scale Bioactivity Database for Drug Discovery,” Nucleic Acids Research 40, no. D1 (2011): D1100–D1107, 10.1093/nar/gkr777. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Morgan H. L., “The Generation of a Unique Machine Description for Chemical Structures‐A Technique Developed at Chemical Abstracts Service,” Journal of Chemical Documentation 5, no. 2 (1965): 107–113, 10.1021/c160017a018. [DOI] [Google Scholar]
  • 35. Landrum G., rdkit/rdkit: March 3, 2025 (Q1 2025) Release (Release March 3, 2025), (Zenodo, 2025), 10.5281/zenodo.15605628. [DOI] [Google Scholar]
  • 36. López‐Pérez K., Avellaneda‐Tamayo J. F., Chen L., et al., “Molecular Similarity: Theory, Applications, and Perspectives,” Artificial Intelligence Chemistry 2, no. 2 (2024): 100077, 10.1016/j.aichem.2024.100077. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Willett P., “Similarity Methods in Chemoinformatics,” Annual Review of Information Science and Technology 43, no. 1 (2009): 1–117, 10.1002/aris.2009.1440430108. [DOI] [Google Scholar]
  • 38. Todeschini R., Ballabio D., and Consonni V., “Distances and Similarity Measures in Chemometrics and Chemoinformatics,” Encyclopedia of Analytical Chemistry (John Wiley & Sons, Ltd., 2020), 1–40. [Google Scholar]
  • 39. Leach A. R. and Gillet V. J., “Similarity Methods,” An Introduction To Chemoinformatics (Springer Netherlands, 2007): 99–117. [Google Scholar]
  • 40. Keiser M. J., Roth B. L., Armbruster B. N., Ernsberger P., Irwin J. J., and Shoichet B. K., “Relating Protein Pharmacology by Ligand Chemistry,” Nature Biotechnology 25, no. 2 (2007): 197–206, 10.1038/nbt1284. [DOI] [PubMed] [Google Scholar]
  • 41. Wang Z., Liang L., Yin Z., and Lin J., “Improving Chemical Similarity Ensemble Approach in Target Prediction,” Journal of Cheminformatics 8 (2016): 20, 10.1186/s13321-016-0130-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Lopez Perez K., López‐López E., Soulage F., Felix E., Medina‐Franco J. L., and Miranda‐Quintana R. A., “Growth Vs Diversity: A Time‐Evolution Analysis of the Chemical Space,” Journal of Chemical Information and Modeling 65, no. 13 (2025): 6788–6796, 10.1021/acs.jcim.5c00347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. López Pérez K., Jung V., Chen L., Huddleston K., and Miranda‐Quintana R. A., “BitBIRCH: Efficient Clustering of Large Molecular Libraries,” Digital Discovery 4, no. 4 (2025): 1042–1051, 10.1039/D5DD00030K. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Klimenko K., Marcou G., Horvath D., and Varnek A., “Chemical Space Mapping and Structure–Activity Analysis of the ChEMBL Antiviral Compound Set,” Journal of Chemical Information and Modeling 56, no. 8 (2016): 1438–1454, 10.1021/acs.jcim.6b00192. [DOI] [PubMed] [Google Scholar]
  • 45. González‐Medina M., Prieto‐Martínez F. D., Owen J. R., and Medina‐Franco Jé L., “Consensus Diversity Plots: A Global Diversity Analysis of Chemical Libraries,” Journal of Cheminformatics 8 (2016): 63, 10.1186/s13321-016-0176-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Kayastha S., Horvath D., Gilberg E., Gütschow M., Bajorath J., and Varnek A., “Privileged Structural Motif Detection and Analysis Using Generative Topographic Maps,” Journal of Chemical Information and Modeling 57, no. 5 (2017): 1218–1232, 10.1021/acs.jcim.7b00128. [DOI] [PubMed] [Google Scholar]
  • 47. Sattarov B., Baskin I. I., Horvath D., Marcou G., Bjerrum E. J., and Varnek A., “De Novo Molecular Design by Combining Deep Autoencoder Recurrent Neural Networks With Generative Topographic Mapping,” Journal of Chemical Information and Modeling 59, no. 3 (2019): 1182–1196, 10.1021/acs.jcim.8b00751. [DOI] [PubMed] [Google Scholar]
  • 48. Pikalyova K., Akhmetshin T., Orlov A., et al., “Design of Highly Potent Anti‐Biofilm, Antimicrobial Peptides Using Explainable Artificial Intelligence,“ Journal of Chemical Information and Modeling 66, no. 1 (2025): 744–755, 10.1021/acs.jcim.5c01992. [DOI] [PubMed] [Google Scholar]
  • 49. Fitzgerald P. R., Dixit A., Zhang C., Mobley D. L., and Paegel B. M., “Building Block‐Centric Approach to DNA‐Encoded Library Design,” Journal of Chemical Information and Modeling 64, no. 12 (2024): 4661–4672, 10.1021/acs.jcim.4c00232. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Data Availability Statement

All datasets and Python code required to reproduce this study are available at https://github.com/Laboratoire‐de‐Chemoinformatique/SCOPE‐DEL. The GPU‐accelerated implementation of the GTM algorithm—along with utilities to build density and classification landscapes and to compute GTM‐derived similarity metrics—is provided in the companion library ChemographyKit at https://github.com/Laboratoire‐de‐Chemoinformatique/ChemographyKit. The datasets are also provided under https://zenodo.org/records/17158359.


Articles from Molecular Informatics are provided here courtesy of Wiley

RESOURCES