Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jul 7;66(14):8239–8250. doi: 10.1021/acs.jcim.6c00939

Mapping Evolution of Molecules across Biochemistry with Assembly Theory

Sebastian Pagel 1, Abhishek Sharma 1, Leroy Cronin 1,*
PMCID: PMC13417883  PMID: 42411597

Abstract

Evolution is often understood through genetic mutations driving changes in an organism’s fitness, but there is potential to extend this understanding beyond the genetics. We propose that natural productscomplex molecules central to Earth’s biochemistrycan be used to uncover evolutionary mechanisms beyond genes. By applying assembly theory (AT), which views selection as a process not limited to biological systems, we can map and measure evolutionary forces in these molecules. AT enables the exploration of the assembly space of natural products, demonstrating how the principles of evolution apply to these complex chemical structures, selecting vastly improbable and complex molecules from a vast space of possibilities. By comparing natural products with a broader molecular database, we can assess the degree of evolutionary contingency, providing insight into how molecular novelty emerges and persists. This approach not only quantifies evolutionary selection at the molecular level but also offers a new avenue for drug discovery by exploring the molecular assembly spaces of natural products. Our method provides a fresh perspective on measuring the evolutionary processes both shaping and being read out by the molecular imprint of selection.


graphic file with name ci6c00939_0010.jpg


graphic file with name ci6c00939_0008.jpg

Introduction

The theory of evolution allows us to elucidate the processes by which distinct populations of species arise, finely tailored to occupy specific biological niches within a given biosphere. If there are sufficient resources, the species survives and reproduces until superseded by a better-adapted species or until the resources diminish and this represents the process of natural selection. While evolutionary theory describes the change of organisms on a biological level, the underlying mechanistic dynamics of evolutionary and pre-evolutionary processes at the molecular scale remain to be defined. The key physical phenomenon ubiquitous at all scales from molecules to complex evolutionary architectures such as cells is selection. Previously, various approaches have been utilized to describe Darwinian evolution such as the quasi-species model based on a physical chemistry framework, evolutionary game theory , and minimal cellular models such as the Chemoton. Most of these models probe evolutionary dynamics by defining key processes at the molecular scale such as replication-mutation leading to natural selection by processes such as cooperation, competition etc. However, these models cannot give full mechanistic insight into the process of selection within the vast chemical space starting at the fundamental molecular scale leading to the emergence of various biochemical processes including autocatalytic sets, replicators, and metabolic pathways. Quantifying the degree of selection required at the molecular scale to create biologically relevant molecules is a complex and open problem.

Recently, Assembly Theory (AT) was utilized to distinguish biological from nonbiological samples, showing that complex molecules if observed in high abundance can act as biosignatures for life on Earth, and this represents a generalized approach to search for life. − The observation of complex molecules is the outcome of evolutionary processes that occur within all biological life as we know it. As an example, within the Earth’s biosphere, complex molecules can either be functional as is NADP+, which is constantly being synthesized, transformed and regenerated, or as secondary metabolites, which are often adaptive characters essential for an organism’s survival and ecological success, with their distribution reflecting a complex history of coevolution. − Well-known examples of plant secondary metabolites can be categorized as terpenoids, phenolic compounds, sulfur-containing compounds, and alkaloids for example papaverine or paclitaxel. The production of these complex molecules appears to be driven by the evolutionary dynamics within the ecosystem, whereby these secondary metabolites can have a causal influence on other living systems. While primary metabolites such as amino acids, organic acids etc. function as all kinds of building blocks in cellular processes, secondary metabolites , interact with external agents and stresses − giving them selective dominance which could lead to higher chances of survival. Consequently, secondary metabolites accumulate in these organisms, and the environment they inhabit. The combination of complexity and accumulation (high copy number in the framework of AT) makes secondary metabolites the premier biosignatures and markers of selection since they allow for an almost unambiguous proof of chemical causality only enforced by a strongly selective evolutionary machine, as observed on Earth. While selected reaction steps of complex biomolecules may occur spontaneously, the series of steps required to construct a complex molecule in high abundance remains highly unlikely. Thus, such complex molecules would not spontaneously arise, but their synthesis pathways have been selected within a population of organisms in each environment. Indeed, the production of natural products can be driven by natural selection, where they can provide an advantage to the organism, such as defending against predators, attracting pollinators, or competing with other species in a molecular arms race.

The quest to find a generalized definition of what life would look like on macroscopic and molecular scales remains an open research question without much consensus. As a potential solution, AT was further extended to quantify and generalize selection in physical systems beyond natural selection as defined in biological systems. AT introduces the concept of an object (such as a molecule) as an entity which is finite, distinguishable, persists in time, and is breakable such that the set of constraints to construct it are quantifiable. Thus, the complexity of an object is quantified by defining an intrinsic measure called the assembly index (a) which is the shortest number of steps to construct an object in the absence of any physical constraints. For an ensemble of objects, we have defined an integrated quantity Assembly (A) of the ensemble which signifies the total amount of selection necessary to produce a set of observed objects, quantified using eq :

A=∑i=1Neai(ni−1NT) 1

where a i is the assembly index of i th object, n i is its copy number or concentration and N T the total number of objects in the ensemble. The combination of assembly index and copy number quantifies within the large combinatorial space, how selection process leads to the discovery of new objects and among them the system’s capacity to produce a high copy number of specific objects.

Herein, we present a detailed exploration of the application of Assembly Theory (AT) to the study of natural products, with the aim of quantifying the evolutionary processes that shape complex biochemical systems. Through our analysis, we demonstrate that the assembly spaces of natural products, as derived from the COCONUT database, represent highly selective outcomes of evolutionary processes. By comparing these assembly spaces with those of all known molecules, we quantify the degree of selection required to generate the observed molecular diversity on Earth. We also explore how a systematic reduction in the constraints controlling these assembly spaces can lead to the generation of novel molecular structures under alternate selection pressures, that is giving new constraints not limited by the previous ones. This approach is particularly significant for drug discovery, as it allows us to reconstruct and explore drug-like molecules that retain key features of natural products, which have evolved to interact with biological systems. By mimicking the fragment utilization and construction step distribution observed in natural products, we can generate new chemical spaces that are enriched in molecules with desirable drug-like properties, while also incorporating evolutionary selected features inherent to natural products. Previous fragment-based approaches have provided valuable libraries of biologically prevalidated, sp3-rich fragments, but they generally rely on predefined fragments or substructure libraries, making the explored chemical space dependent on the fragments selected a priori. , In contrast, our approach derives molecular assembly pathways directly from the molecular graph, allowing molecular construction to be quantified and assembly spaces to be compared without imposing a prior choice of fragments. This enables us to retain information about how molecular substructures are recursively used across natural products, while systematically relaxing historical constraints to explore new molecular spaces under alternative selection pressures. Our work provides a robust framework for understanding the intersection of selection, evolution, and molecular complexity, offering new insights into the origins and potential future trajectories of biochemical diversity, particularly in the context of drug discovery.

To use AT to explore natural product space, we first need to define the shortest pathway to construct the object as the assembly pathway on which the assembly index (molecular assembly index in the case of molecules) can be quantified. This is important because, in the absence of the knowledge of the mechanistic insights through which the object has been created, the assembly pathway represents the informational constraints along the construction pathway required to build the object. This approach, using assembly pathways, allows all the observed objects to be detected in a similar way such that the bias that emerges is due to the selective processes can be estimated by quantifying the historical contingency associated with the construction processes. Thus, the assembly pathways over an ensemble of objects constitute the combinatorial assembly spaces, see Figure .

1.

1

Concept of assembly spaces. (a) The figure shows the expansion of assembly spaces where Assembly Possible (A P , shown in black) represents all physically plausible molecules, Assembly Contingent (A C , shown in blue) represents all known molecules (such as PubChem database), and Assembly Observed (A O , Assembly Observed, shown in red) represents natural products (such as COCONUT database). The living threshold represents the complexity threshold over which the observed molecule in high abundance signifies the presence of life. (b) Example of an expanding assembly space of an ensemble where at each assembly depth, observed and contingent objects exist. The assembly spaces A P and A C represent combinatorial spaces of physically plausible objects because of undirected and directed processes toward higher complexity such that A C is a subspace of A P (A C ⊆ A P ). The assembly space A O (A O ⊆ A C ) represents the space of observed objects (or experimentally measured) which are constructed as an outcome of selection and emerged in high copy number.

Building on the concept of an assembly space, selection is defined in AT as a temporal process leading to the transition from undirected to directed exploration dynamics. We can consider that a physical process represented by forward dynamics, in the assembly space, selection can be observed in time by measuring the selectivity parameter alpha (α) which is equal to one for an undirected process and drops below one for a directed process. Here we quantify the amount of selection required to construct the assembly space of molecules representing biochemical processes occurring on planet Earth. Given the lack of temporal information on the evolution of biochemical processes, we postulate that it is possible to use the molecules in the natural products database as an ensemble of complex objects which represent the Assembly Observed. Assuming that the database samples most of the biochemical processes and has been observed in high enough abundance, we assume that exploration occurred over a long evolutionary time scale, and the Assembly Observed of natural products gives a good representation of the natural products produced by evolution which is captured in the space of the Assembly Contingent and represented as A NP – that is the space of natural products. The Assembly Possible represents all physically plausible molecules and expands exponentially, and hence is computationally intractable. Instead for simplicity, we consider all the molecules found in the PubChem database to represent physically plausible molecules as represented as A M such that the space of natural products is a subspace of the space of all physically possible molecules. The degree of selection is quantified by comparing features of assembly spaces of all possible molecules and natural products, see Figure . We then extend the methodology to explore new chemical spaces by reducing the constraints within the assembly spaces and reconstructing these spaces to give novel molecules as an outcome of alternate selection pressure.

Significance of Assembly Spaces

The assembly space of natural products represents a set of complex and highly selective molecules whose physical construction pathways are too complex to be known as they include evolutionary processes over a very long time scale. At such complexity, the presence of the complex object itself at higher abundance represents the presence of evolutionary processes indicating life. In the context of assembly spaces, the shortest construction pathway quantifies the minimal required information to build an object, where each step along the pathway a → a + 1, i.e. where the assembly index goes up by one indicating the addition of a constraint. This approach captures the process of selection from the assembly pool to combine molecular substructures in a specific way which quantifies the presence of the evolutionary processes which drive biochemical systems. These construction pathways define the lower bound of the free energy required to construct the molecule with connectivity as the only known constraint and do not explicitly consider other physical constraints such as bond strength, environmental factors etc. The potential configurational space of selection and combination of objects is extremely large, and each step along this pathway adds constraints or contingency which captures the effect of external factors such as physical selection due to the influence of evolutionary dynamics. One can argue that in molecular assembly spaces with bonds as building blocks, there is a weak but finite selectivity where specific bonding constraints are favored. By comparing the spaces, Assembly Possible for all molecules (A M ) and Assembly Contingent for natural products (A NP ), this weak selectivity can be decoupled from selectivity introduced by the evolutionary processes. Assuming a lack of knowledge of temporal information for complex molecule formation as well as any intermediate molecular substructures observed, each step along the assembly pathway quantifies information regarding selection and combination representing the effect of the presence of an evolutionary process in biology. These effects can be quantified more precisely if temporal information on intermediate species within the joint assembly space has been observed.

Assembly Depth and Joint Assembly Spaces

The assembly index quantifies the number of steps in a serial construction process, where all the subobjects (also referred as contingent objects) are required to be constructed sequentially. However, the order of construction steps along the shortest path is not unique if two independent construction processes have been used. As an example, the string AABB can be constructed in a serial process by two possible pathways A → AA → AAB → AABB and B → BB → ABB → AABB. These serial pathways require an arbitrary ordering of steps, even though the subobjects AA and BB can be constructed independently. Assembly depth captures this possibility of concurrent construction, assigning the building blocks A and B an assembly depth of 0 gives d(AA) = d(BB) = 1, and joining these two independently constructed subobjects gives d(AABB) = max (1,1) + 1 = 2. Thus, assembly depth measures the longest dependent chain of construction steps, rather than the total number of serial joining operations. In this sense, the assembly depth of an object defines a lower bound on the assembly index (a) when independent construction processes can occur concurrently. As an example, Figure a shows an assembly pathway of a molecule (acetylsalicylic acid) as an object and bonds as building blocks. In the case of molecules, the assembly index (a) is also referred to as Molecular Assembly (MA). If the assembly pathway constitutes the addition of building blocks to one linearly growing object, the MA is equivalent to the assembly depth (d). If instead, the assembly pathway connects two or more distinct objects which are not building blocks and can be constructed independently, the MA will differ from the assembly depth. In cases where two or more objects were constructed on different branches, but by the same number of joining operations, their assembly depth will be identical. Thus, the introduction of assembly depth removes the need to assign an arbitrary order to the joining operations in an assembly pathway in cases where multiple equivalent solutions exist, by allowing concurrent processes. This is particularly useful for forward dynamics, where future combinatorial spaces need to be generated based on the existing information within the assembly spaces. Practically, the MA is defined as the number of joining operations to construct a given molecule from its assembly pathway. The assembly depth in contrast is assigned recursively starting from the building blocks (assembly depth 0). Each subsequent object’s assembly depth is calculated according to eq .

d=max(df1,df2)+1 2

where d f1 and d f2 are the assembly depths of the two fragments that were connected to build a new object. Per definition, the assembly depth of building blocks is set to 0. In the case of multiple (not building block) fragments with the same assembly depth within an assembly pathway (existence of concurrent processes), the MA is not equal to the assembly depth. For example, the MA of acetylsalicylic acid is 8, but the assembly depth is 6, see Figure a (see SI Section 1 for details on the assembly pathway and assembly depth calculations).

2.

2

Assembly space of a single molecule and joint assembly space of an ensemble of molecules. (a) Assembly pathway of acetylsalicylic acid (MA = 8, d = 6). The unique bonds found in acetylsalicylic acid act as the building blocks and are used together with the molecular substructures (both in green boxes) to construct acetylsalicylic acid (red box). (b) Depiction of the Joint Assembly Space (JAS) of 24 molecules from the COCONUT database with a molecular assembly index between 7 and 10. The assembly fragments are ordered by their assembly depth. Each fragment is represented by a node in the JAS, where the size of the node represents the number of times the fragment was used during the construction process of the Assembly Space. Nodes that are colored in light green represent building blocks and contingent objects (here molecular substructures) that have been used to construct the observed objects (natural products; red nodes). One observed molecule pathway is shown in dark green with the molecular (sub)­structures shown next to each respective node. (c) Conceptual depiction of the assembly space of molecules. The outermost cone depicts the outer bounds of the Assembly Possible space, which contains all theoretical possible molecules that can be constructed from an assembly process obeying the laws of physics. The x-axis shows the distribution of novelty/diversity of observed objects at a given assembly step relative to an arbitrary reference molecule. Within Assembly Possible lies the assembly space of natural products (in green) which are the products of evolutionary dynamics of living systems. The yellow cones depict assembly spaces diverging from the assembly space of natural products at different assembly steps.

In this study, we used the natural product database (COCONUT database) to generate the JAS as a model representing the outcomes of evolutionary processes found across biochemistry, where the JAS of 24 natural product molecules selected from the COCONUT database is shown in Figure b. The JAS of all natural products signifies the presence of molecular evolution defined by the physical and chemical constraints given on Earth and the biological constraints of the evolutionary processes that produced all life on Earth. With 407,270 unique natural products, the COCONUT database is one of the most comprehensive databases of molecules that are a direct result of the evolutionary processes on Earth, accurately representing the chemical distribution found in biological life on Earth.

The JAS of natural products represents the assembly space of physically plausible molecules that emerged as an outcome of selection by the biological processes. This represents construction and selection processes leading to the observation of natural products utilizing a common pool of building blocks and molecular intermediates and captures the effect of contingency in the assembly space, where the shared pathways represent utilization of common molecular fragments and building blocks to create distinct molecules. An intuitive example of a molecular fragment adding contingency between many assembly pathways of natural products is the phenyl ring, which once constructed could be recursively utilized to create complex aromatic molecules such as phenolic compounds.

Loss of Contingency and Reconstruction of Assembly Spaces

The presence of contingency in the JAS is observed by the utilization of subobjects or molecular fragments in a recursive way to construct the observed molecules and represents the information existing as subobjects that is accessible to construct the molecules. Given a JAS of some observed molecules, the existing contingency can be reduced in a restricted way by removing objects and subobjects up to a certain assembly depth which we define as contingency loss. The remaining constraints can be used to construct new molecules different from the initial JAS. This approach is particularly important in exploring the assembly spaces defined by Assembly Possible constrained by the remaining Assembly Contingent. Figure c shows the pictorial representation of assembly spaces A M (Assembly of all Possible Molecules) and A NP (Assembly Contingent or Observed) together with emerging novel spaces by losing contingency (represented by ω) and reconstructing assembly spaces with alternate selection pressure.

Here, alternate selection pressure means sampling contingent objects and constructing physically plausible molecules with different rules compared to natural products. The key hypothesis is the higher the loss of contingency within the assembly space, the more novel molecules sharing the limited contingency can be generated with varying selection pressure. To study the influence of contingency loss on the emergence of novelty, an algorithm for the construction of molecules from a set of molecular fragments (i.e., in this case, an assembly space) was developed, see Figure . Given an arbitrary JAS or an isolated assembly pathway, molecules were reconstructed from the available fragments, by first removing the ω topmost levels of assembly depth (for reconstruction from JAS or assembly pathway only) as an outcome of the removal of historical contingency. From the truncated JAS (JASω), molecules are then reconstructed by selecting a random fragment from the topmost assembly depth of JASω, and iteratively sampling other available fragments (compare Figure a or SI Sections 4 and 7 for details) with valence rules as the only constraint. In principle, the number of construction steps (n steps ) for the reconstruction of molecules can be set to any integer value. In this work, n steps was either set to a fixed integer value (for reconstruction from a single assembly pathway) or sampled according to the distribution of number of construction steps as observed by the observed objects in a JAS. Alternatively to starting the construction of molecules from a predefined JAS, it may also start from a set of building blocks (S) to iteratively build up a new JAS only constrained by the initial building blocks (Figure b).

3.

3

Workflow for reconstructing and filtering molecules. (a) Reconstruction of molecules from a Joint Assembly Space. To reconstruct molecules (and a JAS) after contingency loss ω, all fragments with an assembly depth (d) larger than d max – ω are removed from the initial JAS. To start the reconstruction process, a fragment with assembly depth d max – ω is randomly selected. The number of reconstruction steps is then sampled according to the distribution of construction steps in the original JAS. For each reconstruction step, a second fragment is sampled from the truncated JAS (JAS dmax‑ω). To connect the two fragments, first, the number of connections to be formed is sampled. When two fragments have been successfully connected, the resulting object is added to the JAS. This process is repeated until n steps reconstruction steps have been completed. (b) When constructing molecules from a set of predefined building blocks, first a building block is randomly chosen from the set of building blocks. The number of construction steps (n steps ) is defined by the user or application. For each construction step, a second building block from the set of building blocks is chosen. Fragments are connected as described for a). If the fragments were successfully connected, the new resulting fragment is added to the set of building blocks. The process is repeated until n steps have been completed. (c) To improve the physical validity of the resulting molecules an optional filtering pipeline was developed (see SI Sections 4, 5, and 7 for details).

In either case, an optional three-step filtering pipeline was implemented to reject chemically valid, but physically implausible molecules generated (Figure c). In the filtering pipeline, molecules are first checked against a set of forbidden substructures as defined in MOLGEN. Conformers of molecules that pass this stage were generated and subsequently geometry optimized. Molecules for which either no valid conformer could be generated, or the geometry optimization did not converge for any of the generated conformers were discarded. In the last stage, bond lengths and angles in passing molecules were compared against the distribution of the respective bond lengths and angles found in reported molecular structures from the CCDC database (SI Section 5 for details). The novelty of constructed molecules compared to the observed molecules in the original JAS or assembly pathway is then calculated by the Dice-Similarity between the Morgan-Fingerprints of two molecules (eq ; SI Section 6).

s=2|xf1×xf2||xf1|2+|xf2|2 3

where s is the similarity between the Morgan Fingerprint between two molecules f 1 and f 2, x f1 and x f2 are the Morgan Fingerprints of the respective molecules.

To investigate the characteristic novelty (defined as mean maximal similarity between two sets of molecules) of reconstructed molecules from the assembly pathway of a single molecule, 10,000 molecules were reconstructed from the assembly pathway of Brefelamide (an aromatic amide isolated from Dictyostelium cellular slime molds see Figure a) with increasing contingency loss ω from 1 to 10 (Figure b). In good agreement with general intuition, the mean similarity between reconstructed molecules from the remaining object in the assembly pathway and Brefelamide decreases with an increase in contingency loss ω from approximately 0.8 (ω = 1) to approximately 0.15 (ω = 10). Notably, from contingency loss ω = 6 to ω = 7 the similarity of reconstructed molecules relative to Brefelamide is increasing.

4.

4

Reconstruction of molecules from the assembly pathway of a single molecule. (a) Assembly pathway of the secondary metabolite Brefelamide ordered by the assembly depth of the fragments within the pathway. (b) Boxplot representation of the similarity (eq ) of molecules reconstructed from the partial assembly path of Brefelamide, was calculated, by first removing ω = {1, ···,10} contingency from the assembly path and reconstructing 10,000 molecules from the remaining fragments up to the assembly depth of Brefelamide. The most similar reconstructed molecules after ω contingency loss steps are displayed after each respective bar plot. Error bars indicate the standard deviation of the similarity. (c) Mean maximal pairwise Dice similarity, also calculated using eq , among the 10,000 reconstructed molecules generated for each value of ω. The error bars represent the standard deviation of the mean maximal pairwise similarity.

We hypothesize that this effect arises from the presence of phenol at assembly depth 3. After contingency loss ω = 7, phenol becomes the fragment with the highest remaining assembly depth and therefore serves as the starting point for the reconstruction of new molecules. Because three phenol substructures are present in Brefelamide, the retention or loss of this fragment strongly affects the similarity of the reconstructed molecules to the parent structure. To further assess the diversity of the reconstructed molecules, we calculated the mean maximal pairwise Dice similarity among molecules generated at each level of contingency loss using the same similarity measure defined in eq (Figure c). Whereas Figure b compares each reconstructed molecule to Brefelamide, Figure c compares reconstructed molecules to one another. Together, these analyses show that increasing contingency loss generally reduces similarity to Brefelamide and increases the diversity of the reconstructed molecules, although retained high-depth fragments can produce local deviations from this trend. Thus, contingency loss progressively relaxes the constraints of the original assembly pathway while still allowing retained fragments to influence the structures generated.

Generation of Novelty with Contingency Loss and Selection Pressure

The generation of novelty in chemical spaces was studied based on the JAS of 211731 natural products from a COCONUT database (JASNP), with an assembly depth of up to 20 (d max ; SI Section 2). To investigate the generation of novelty in newly generated JASs with respect to the parent JAS, an increasing amount of contingency (ω = {2,6,10,14}) was removed from JASNP by removing objects (molecules and molecular fragments on their assembly pathways) with an assembly depth of d max – ω as described in SI Section 3.1. Starting from the resulting truncated JAS (JASω), the same number of molecules (observed objects) as removed through the contingency loss were reconstructed as described above (and SI Section 4). The divergence of the reconstructed JAS (JASNP*) relative to JASNP was calculated using eq (compare Figure ; SI Section 6 for more details). The divergence of two JASs is thus represented as the mean maximal similarity between molecules (and molecular substructures) of the same assembly depth and calculated for each assembly depth individually.

5.

5

Distribution of Assembly Depth and Assembly Index JASNP and divergence of reconstructed JAS after contingency loss with alternate selection pressure. (a) Distribution of assembly index and assembly depth of the 211731 natural products with an assembly depth of up to 20 from which JASNP was constructed. (b) Divergence (θ) of reconstructed JASs after contingency loss ω = {2,6,10,14} relative to JASNP calculated as described above. The JAS was either reconstructed to mimic JASNP (COCONUTmimic; blue) or diverge from JASNP (COCONUTdivergent; green). (c) Mean of molecular properties of molecules in COCONUT JAS, and the reconstructed JASs with increasing contingency loss.

While JASNP* was reconstructed to mimic the construction process of the natural product JAS (JASNP) by setting the sampling weights in the molecule generation algorithm according to the distributions in JASNP (SI Sections 3 and 7) an alternate selection pressure can be induced by tuning these sampling weights. In principle, the sampling weights can easily be adjusted to favor certain functional groups, atomic distributions, or any other chemical feature. Since only limited knowledge of the mechanistic leading to JASNP is available the sampling weights for the fragment selection (SI Sections 3 and 7) were set to be uniform over all available fragments to model an alternate selection pressure. This represents a random exploration of the chemical space, contingent only on the previously explored fragments in JASω (SI Section 8.1). In both cases (reconstruction mimicking JASNP and random reconstruction) the divergence of the resulting JAS (JASNP*) increases with increasing contingency loss ω relative to JASNP (see Figure ). While for smaller contingency loss (ω = {2,6}) the divergence of the resulting JASsNP* relative to JASNP is almost identical, an increasing difference of the divergence can be observed for ω = {10,14} and increasing assembly depth when comparing random reconstruction vs reconstruction mimicking the dynamics in JASNP. Thus, the divergence of JASs, and the construction of molecules using fragments that exist in the evolutionary history of JASs, might be used to construct novel chemical spaces. Additional analysis has been performed of common molecular properties with contingency loss for reconstructed molecules (see Figure c).

Quantification of Selection in Earth’s Chemical Space of Natural Products

The exploration of an assembly space as an outcome of biological selection becomes highly restricted with increase in assembly depth due to informational constraints introduced by biological processes. Here, we quantify the overall selection by estimating how the exploration ratio of the assembly space of natural products with respect to all molecules scales with assembly depth by defining their relationship as r = k e –βd , where r is the exploration ratio, k is the fixed ratio of the number of building blocks in natural products and all molecules, β is the characteristic constant representing selection and d is the assembly depth. The molecular space was modeled using the JAS of more than 70 million molecules extracted from the PubChem database (JASPC) and the contingent assembly space as JAS of natural products (JASNP; see SI Sections 2 and 9 for more details on JASNP and JASPC). JASNP was constructed from all natural products with an assembly depth of up to 25 from COCONUT database. Here, JASNP and JASPC represent A NP and A M respectively where A NP is a subspace of A M as described above. To quantify the selection that was required to construct A NP within A M , all unique objects grouped by assembly depth were aggregated for both JASPC and JASNP. The exploration ratio defined as N NP /N PC of JASNP relative to JASPC was then calculated by taking the ratio of unique objects in JASNP (N NP ) to JASPC (N PC ) for objects of each assembly depth individually (see Figure c,d). Additionally, the exploration ratio of reconstructed JAS (JASNP*) after contingency loss from JASNP relative to JASPC was calculated (N NP*/N PC ). This was done by removing contingency from JASNP as described above by either removing objects with an assembly depth larger than d max – ω where ω = 23 (see Figure b,c), or removing the objects with an assembly depth larger than d max – ω with ω = {17,19,21} as well as the objects exclusively appearing on their assembly pathways (see Figure a,d). In this case, the ratio k was adjusted to match the ratio of building blocks after contingency loss. Subsequently, molecules were reconstructed as described above (see SI Section 4 for details).

6.

6

Exploration of JASPC by JASNP and selectively reconstructed JASs. (a) Schematic representation of a JAS from which contingency is removed by removing observed molecules and the fragments on their pathways and subsequent reconstruction of the JAS (left to right). (b) Schematic representation of JAS from which contingency is removed by removing objects with an assembly depth larger than d max and subsequent reconstruction of the JAS (left to right). (c) Exploration ratio (N NP /N PC or N NP*/N PC ; the ratio of unique objects per assembly depth relative to JASPC) of JASPC by the JAS of natural products (JASNP; gray bars) and reconstructed JASs thereof after contingency loss ω = 23 by removing objects as described in b). The fragment selection for the reconstruction of new molecules was altered by adjusting the exponential scaling factor s = {0.5,1,5} (see eq 3 in the SI). The exploration ratio N X /N P was estimated by an exponential function. The characteristic constant (β) of the exploration ratio by JASNP was 0.499 (purple) and 0.279, 0.374, and 0.508 of the JAS with s = {0.5,1,5} respectively (red, blue and green). (d) N NP /N PC of JASPC by JASNP (gray bars) and reconstructed JASs after contingency as described in a) with contingency loss ω = {17,19,21}. The exploration ratio was estimated by an exponential function. The characteristic constant (β) of the exploration ratios was 0.603, 0.647, and 0.688 of the JAS with ω = {17,19,21} respectively (red, blue and green). The value of β for JASNP (purple) is equivalent to c) since these represent the same data.

In the latter case, the reconstruction of molecules was controlled by adjusting the exponential scaling factor s for the fragment selection s = {0.5,1,5} as described in SI Section 7. In short, adjusting this scaling factor results in a distribution shift where s = 0 represents a uniform distribution and increasing s, increases the likelihood of a fragment being sampled that was preferentially selected in the original JASNP.

As previously defined, the exploration ratio of JASNP as well as all reconstructed JASNP* was estimated with the exponential function k e –βd where d is the assembly depth, the ratio k ≈ 0.12, and β is a characteristic constant quantifying the selection induced to attain a JAS subspace (see Figure c,d and SI Section 10). In case a JAS is fully explored β will be 0. Thus, as the value of β increases, the process must become increasingly selective in order to construct the corresponding JAS subspace. β estimates the amount of selection required to construct the JAS of natural products (JASNP) within the JAS on a planetary scale (JASPC) to be 0.499. The reconstruction of JASs from JASNP with the same amount of contingency loss (ω = 23) but differing exponential scaling factor for the fragment selection s = {0.5,1,5} results in increasing exploration ratios, with decreasing s (β of 0.279, 0.374, and 0.508, respectively; Figure c). Likewise, removing an increasing amount of contingency (ω = {17,19,21}) with constant s results in a decreasing exploration ratio with β of 0.603, 0.647, and 0.688, respectively (Figure d). It is important to note that β is not a process-specific marker of evolution by itself, but an ensemble-level measure of constrained exploration whose interpretation depends on the origin of the data set and the reference ensemble used for comparison, which in our case is selection as an outcome of evolutionary processes.

Constructing a Drug-like Chemical Space from the Contingency of Natural Products

In principle, every JAS that is the subspace of another JAS can be constructed from this parent JAS if the exact dynamics and selection conditions are known. As an example, the JAS of natural products was constructed on Earth within Assembly Possible as defined by all known molecules. While we do not have exact knowledge of the dynamics and selection conditions leading to the natural products observed on Earth, the chemical space distribution of JAS of natural products can be reconstructed by mimicking the fragment utilization and construction step distribution observed during the reconstruction of molecules. Here, we postulate that this concept can be extended to construct drug-like molecules from the JAS of natural products since many drug molecules are derived from natural products which we consider to be a subspace of Assembly Contingent and Assembly Possible (see Figure a). The JAS of drug molecules (JASD) was modeled with 10656 small molecules from the ChEMBL database with an Assembly Depth of up to 20 (SI Section 2). As natural products naturally evolved to interact with biological organelles and biomolecules, fragments in the JAS of natural products may thus contain promising chemical substructures for novel drug-like compounds that are yet to be explored. By generating a JAS of drug-like molecules from the JAS of natural products, we thus may be able to generate molecules containing both classically drug-like features, but also chemical features produced by evolution and selection.

7.

7

Construction of drug-like molecule from the contingency of natural products. (a) Conceptual depiction of the JAS of drug molecules (JASD) in comparison to the JASNP and JASPC. (b) Similarity of constructed molecules from the contingency of natural products with fragment connection rules of drug molecules (JASD; orange), or random reconstruction (green) to the drug-like molecules in JASD. (c) 7 example molecules labeled with QED-score out of 10,000 molecules that were reconstructed from the JAS of natural products after contingency loss ω = 15 (down to assembly depth 5). In addition to the previously described substructure filters (see above), PAINS substructures were used to filter out molecules with known problematic substructures in drug-screening assays.

To test this hypothesis, contingency was removed from the JAS of natural products by deleting molecular fragments and molecules with an assembly depth larger than 5. From the remaining fragments, 10,000 molecules were reconstructed as described above (see SI Section 8.2) while applying fragment connection rules, as observed in JASD. Additionally to the filtering steps described previously, PAINS substructure filters were used to filter out molecules commonly considered to be problematic in drug-screening campaigns. This strategy yielded molecules more similar to molecules observed in JASD than molecules constructed without these rules (Figure b). QED scores were calculated for the reconstructed molecules with 7 exemplary molecules shown in Figure c showcasing a strategy for exploring novel drug-like molecules. Although the potential of drug-like molecules reconstructed from the JAS of natural products remains to be further explored and developed, this represents a potential workflow to explore chemical spaces guided by the constraints of evolution and selection that lead to physically observed molecules.

Conclusions

In this work, we have expanded the conceptual framework of AT by introducing the significance of causal contingency in assembly spaces and its application to molecular spaces where the mechanisms of chemical reaction networks are too complex, occur over long time scales and are unknown. In the evolutionary systems, these contingent effects indicate the selection processes with the vast combinatorial universe which can be quantified. As an application of AT, we used natural products and all molecules’ databases as observed and physically plausible molecular ensembles, generated their assembly pathways, and created joint assembly spaces which quantify the overall informational constraints required to construct them. By introducing the concept of contingency loss with the joint assembly spaces and developing a cheminformatic engine to filter physically plausible molecules, we demonstrated and quantified how novelty emerges in molecular space by losing causal contingency. Using this approach, we expanded and quantified the properties of assembly spaces of natural products creating physically plausible molecules in the presence of alternate selection pressure and at different degree of contingency loss. This approach is applicable to understanding the emergence of complex molecular observables in the presence of alternate reaction networks of the scale of biologically relevant evolutionary systems. We then compared the joint assembly spaces of natural products and all molecules’ database to quantify the degree of selection emerging from the evolutionary processes on Earth and further expanded to explore drug-like molecules utilizing contingent space of natural products. This approach provides a novel framework to understand and quantify selection processes emerging from complex evolutionary networks and exploring alternate selection pressures and their observables beyond Earth biochemistry.

Supplementary Material

Acknowledgments

We acknowledge financial support from the John Templeton Foundation (grant nos. 61184 and 62231), the Engineering and Physical Sciences Research Council (EPSRC) (grant nos. EP/L023652/1, EP/R01308X/1, EP/S019472/1, and EP/P00153X/1), the Breakthrough Prize Foundation and NASA (Agnostic Biosignatures award no. 80NSSC18K1140), MINECO (project CTQ2017-87392-P), and the European Research Council (ERC) (project 670467 SMART-POM). We like to acknowledge Amit Kahana and Szymon Łukaszyk for the feedback on the manuscript.

Glossary

Abbreviations

The following abbreviations were used to describe the Joint Assembly Spaces explored in this work:

JAS PC

the JAS of all known molecules modeled by over 70 million molecules from the PubChem database

JAS NP

the JAS of natural products as an outcome of the evolutionary machinery of Earth modeled by approximately 211 thousand molecules from the COCONUT database

JAS NP*

JASs reconstructed from JASNPafter contingency loss

JAS D

the JAS of drug molecules modeled by 10656 small molecules from the ChEMBL database

JAS ω

a JAS after loss of contingency (ω)

Exploration ratios were abbreviated as follows:

N NP / N PC

the exploration ratio of JASPCby the objects in JASNP

N NP */ N PC

the exploration ratio of JASPCby the objects in JASNP*

N D / N PC

the exploration ratio of JASPCby the objects in JASD

N D / N NP

the exploration ratio of JASNPby the objects in JASD

All the codes used to perform analysis and generate figures in the manuscript and Supporting Information are available at https://github.com/croningp/molecular_spaces. The code to generate the raw data is available on GitHub, but the raw data generated, ca. 700 GB, is also available upon request.

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.6c00939.

  • Details generating the assembly pathways, joint assembly spaces, molecular generation pipeline, molecular filter description, divergence estimation, influence of molecular generation hyperparameters, reconstruction of molecules, database analysis, and quantification of exploration in joint assembly spaces (PDF)

L.C. conceived the idea and research plan together with A.S and S.P. A.S. built the concept and approach to utilize the methods of Assembly Theory for quantifying selection together with L.C. and S.P. S.P. and A.S. ran the assembly calculations on the databases. S.P. developed cheminformatics tools and performed all the calculations on joint assembly spaces, molecular reconstruction, and quantification of selection. A.S. and L.C. mentored S.P. All the authors wrote the manuscript.

All the molecular assembly calculations on molecules from the COCONUT and PubChem databases were performed using AssemblyGo and AssemblyCpp codes. All the pathways and further data analysis were performed using Python3 using standard libraries.

The authors declare no competing financial interest.

References

  1. Darwin, C. The Origin of Species: By Means of Natural Selection, or the Preservation of Favoured Races in the Struggle for Life; Cambridge University Press: Cambridge, 2009. [Google Scholar]
  2. Eigen M., McCaskill J., Schuster P.. Molecular quasi-species. J. Phys. Chem. 1988;92:6881–6891. doi: 10.1021/j100335a010. [DOI] [Google Scholar]
  3. Nowak M. A.. Five Rules for the Evolution of Cooperation. Science. 2006;314:1560–1563. doi: 10.1126/science.1133755. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Nowak, M. A. Evolutionary Dynamics: Exploring the Equations of Life; Belknap Press of Harvard Univ. Press: Cambridge, Mass., 2006. [Google Scholar]
  5. Gánti, T. The Principles of Life; Oxford University Press: Oxford, 2006. [Google Scholar]
  6. Manrubia S.. The simple emergence of complex molecular function. Philos. Trans. R. Soc. A. 2022;380:20200422. doi: 10.1098/rsta.2020.0422. [DOI] [PubMed] [Google Scholar]
  7. Lincoln T. A., Joyce G. F.. Self-Sustained Replication of an RNA Enzyme. Science. 2009;323:1229–1232. doi: 10.1126/science.1167856. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Patzke V., Kiedrowski G. V.. Self replicating systems. Arkivoc. 2007;2007:293–310. doi: 10.3998/ark.5550190.0008.522. [DOI] [Google Scholar]
  9. Voet, D. ; Voet, J. G. . Biochemistry; John Wiley & Sons: Hoboken, NJ, 2011. [Google Scholar]
  10. Marth J. D.. A unified vision of the building blocks of life. Nat. Cell Biol. 2008;10:1015–1015. doi: 10.1038/ncb0908-1015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Marshall S. M., Mathis C., Carrick E., Keenan G., Cooper G. J. T., Graham H., Craven M., Gromski P. S., Moore D. G., Walker S. I., Cronin L.. et al. Identifying molecules as biosignatures with assembly theory and mass spectrometry. Nat. Commun. 2021;12:3033. doi: 10.1038/s41467-021-23258-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Jirasek M.. et al. Investigating and Quantifying Molecular Complexity Using Assembly Theory and Spectroscopy. ACS Cent. Sci. 2024;10:1054–1064. doi: 10.1021/acscentsci.4c00120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Sharma A.. et al. Assembly theory explains and quantifies selection and evolution. Nature. 2023;622:321–328. doi: 10.1038/s41586-023-06600-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Fisher, A. B. ; Zhang, Q. . NADPH AND NADPH OXIDASE. in Encyclopedia of Respiratory Medicine; Elsevier, 2006. 77–84 10.1016/B0-12-370879-6/00252-0. [DOI] [Google Scholar]
  15. Zhang Y., Deng T., Sun L., Landis J. B., Moore M. J., Wang H., Wang Y., Hao X., Chen J., Li S., Xu M., Puno P. T., Raven P. H., Sun H.. et al. Phylogenetic patterns suggest frequent multiple origins of secondary metabolites across the seed-plant ‘tree of life’. Natl. Sci. Rev. 2021;8:nwaa105. doi: 10.1093/nsr/nwaa105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Elhamouly N. A., Hewedy O. A., Zaitoon A., Miraples A., Elshorbagy O. T., Hussien S., El-Tahan A., Peng D.. et al. The hidden power of secondary metabolites in plant-fungi interactions and sustainable phytoremediation. Front. Plant Sci. 2022;13:1044896. doi: 10.3389/fpls.2022.1044896. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Keller N. P.. Fungal secondary metabolism: regulation, function and drug discovery. Nat. Rev. Microbiol. 2019;17:167–180. doi: 10.1038/s41579-018-0121-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Guerriero G.. et al. Production of Plant Secondary Metabolites: Examples, Tips and Suggestions for Biotechnologists. Genes. 2018;9:309. doi: 10.3390/genes9060309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Ashrafi S.. et al. Papaverine: A Miraculous Alkaloid from Opium and Its Multimedicinal Application. Molecules. 2023;28:3149. doi: 10.3390/molecules28073149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Tew, K. D. Paclitaxel. in xPharm: The Comprehensive Pharmacology Reference; Elsevier, 2007. 1–5 10.1016/B978-008055232-3.62356-6. [DOI] [Google Scholar]
  21. Ashraf, M. A. et al. Environmental Stress and Secondary Metabolites in Plants. in Plant Metabolites and Regulation Under Environmental Stress; Elsevier, 2018. 153–167 10.1016/B978-0-12-812689-9.00008-X. [DOI] [Google Scholar]
  22. Demain A. L.. Microbial production of primary metabolites. Naturwissenschaften. 1980;67:582–587. doi: 10.1007/BF00396537. [DOI] [PubMed] [Google Scholar]
  23. Herbert, R. B. The Biosynthesis of Secondary Metabolites; Springer Netherlands: Dordrecht, 1981. 10.1007/978-94-009-5833-3. [DOI] [Google Scholar]
  24. Vining L. C.. FUNCTIONS OF SECONDARY METABOLITES. Annu. Rev. Microbiol. 1990;44:395–427. doi: 10.1146/annurev.mi.44.100190.002143. [DOI] [PubMed] [Google Scholar]
  25. Mahajan M., Kuiry R., Pal P. K.. Understanding the consequence of environmental stress for accumulation of secondary metabolites in medicinal and aromatic plants. Journal of Applied Research on Medicinal and Aromatic Plants. 2020;18:100255. doi: 10.1016/j.jarmap.2020.100255. [DOI] [Google Scholar]
  26. Coutinho I. D.. et al. Flooded soybean metabolomic analysis reveals important primary and secondary metabolites involved in the hypoxia stress response and tolerance. Environmental and Experimental Botany. 2018;153:176–187. doi: 10.1016/j.envexpbot.2018.05.018. [DOI] [Google Scholar]
  27. Hashim A. M.. et al. Oxidative Stress Responses of Some Endemic Plants to High Altitudes by Intensifying Antioxidants and Secondary Metabolites Content. Plants. 2020;9:869. doi: 10.3390/plants9070869. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Greiner L. C.. et al. Pseudonatural Products for Chemical Biology and Drug Discovery. J. Med. Chem. 2025;68:14137–14170. doi: 10.1021/acs.jmedchem.5c00643. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Over B.. et al. Natural-product-derived fragments for fragment-based ligand discovery. Nature Chem. 2013;5:21–28. doi: 10.1038/nchem.1506. [DOI] [PubMed] [Google Scholar]
  30. Kim S.. et al. PubChem. 2023 update. Nucleic Acids Res. 2023;51:D1373–D1380. doi: 10.1093/nar/gkac956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Sorokina M., Merseburger P., Rajan K., Yirik M. A., Steinbeck C.. COCONUT online: Collection of Open Natural Products database. J. Cheminform. 2021;13:2. doi: 10.1186/s13321-020-00478-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Gugisch, R. ; et al. Chapter 6 - MOLGEN 5.0, A Molecular Structure Generator. In Advances in Mathematical Chemistry and Applications; Basak, S. C. ; Restrepo, G. ; Villaveces, J. L. , eds.; Bentham Science Publishers, 2015; pp 113–138 10.1016/B978-1-68108-198-4.5000. [DOI] [Google Scholar]
  33. Groom C. R., Bruno I. J., Lightfoot M. P., Ward S. C.. The Cambridge Structural Database. Acta Crystallogr. B Struct Sci. Cryst. Eng. Mater. 2016;72:171–179. doi: 10.1107/S2052520616003954. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Kikuchi H.. et al. Isolation and Synthesis of a New Aromatic Compound, Brefelamide, from Dictyostelium Cellular Slime Molds and Its Inhibitory Effect on the Proliferation of Astrocytoma Cells. J. Org. Chem. 2005;70:8854–8858. doi: 10.1021/jo051352x. [DOI] [PubMed] [Google Scholar]
  35. Baell J. B., Holloway G. A.. New Substructure Filters for Removal of Pan Assay Interference Compounds (PAINS) from Screening Libraries and for Their Exclusion in Bioassays. J. Med. Chem. 2010;53:2719–2740. doi: 10.1021/jm901137j. [DOI] [PubMed] [Google Scholar]
  36. Seet I., Patarroyo K. Y., Siebert G., Walker S. I., Cronin L.. Rapid Exploration of the Assembly Chemical Space of Molecular Graphs. J. Chem. Inf. Model. 2025;65:13203–13214. doi: 10.1021/acs.jcim.5c01964. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

All the codes used to perform analysis and generate figures in the manuscript and Supporting Information are available at https://github.com/croningp/molecular_spaces. The code to generate the raw data is available on GitHub, but the raw data generated, ca. 700 GB, is also available upon request.


Articles from Journal of Chemical Information and Modeling are provided here courtesy of American Chemical Society

RESOURCES