Abstract
Reticular materials have come to the fore of chemistry with exceptional potential in applications ranging from CO2 capture and chemical separations to catalysis and drug delivery. However, due to the vast combinatorial space of molecular building blocks that can form these materials, designing high-performing reticular materials for applications remains a considerable challenge. Here, we present a computational approach that combines a library of molecular fragments suitable for constructing organic building units, template-based reassembly, and evolutionary optimization to accelerate the discovery of reticular materials. Applied to metal–organic polyhedra (MOPs), this approach produces a design space of nearly 800,000 MOP configurations. A genetic algorithm (GA) based on the molecular fragments is shown to be effective at rapidly identifying optimal MOPs within this space, demonstrated through optimizing cavity properties for host–guest applications and CO2 interaction energies estimated by machine-learning-accelerated simulations. An important component of our approach is that it is fully ontologized and integrated within The World Avatar, forming part of a broader, interoperable knowledge model for the discovery of reticular materials.


1. Introduction
In recent decades reticular chemistry has become established as a general approach to designing materials with targeted properties. − In this paradigm, materials are constructed through the programmed assembly of molecular building blocks into higher-order structures with exemplary reticular materials including metal–organic frameworks (MOFs), covalent organic frameworks (COFs), and the discrete analogue of MOFs, metal–organic polyhedra (MOPs). A core attraction of reticular chemistry is the predictable assembly which enables fine control over the material properties through modifying the underlying building blocks, the potential of this being demonstrated in a diverse array of applications, including carbon capture, − atmospheric water harvesting, − drug delivery, , and battery materials. −
Despite the advantages for rational design, searching for optimal reticular materials is still a considerable challenge. A primary reason for this is the expansive design space that arises from combining possible chemical building units (CBUs), which are instances of molecular building blocks, and assembly models, that define the topology and connectivity. The multitude of possible combinations precludes an entirely experimental approach even with high-throughput methods, and emphasizes the importance of developing computational methods that can rapidly explore the chemical space and return promising candidates for further study. The potential of computational methods to accelerate materials discovery is well recognized and there has been considerable focus on reticular materials, including screening known materials for applications and to gain insights into property trends − as well as enumerating extensive libraries of hypothetical materials. − More recently, machine learning methods have gained prominence, accelerating screening efforts and advancing inverse design through generative models capable of proposing materials or building blocks given a composition or desired properties. − While generative models have considerable potential for interrogating structure–property relationships at scale and thereby progressing rational design, their application to reticular materials is still nascent, typically exhibiting limited explainability and requiring extensive training data sets to reliably generate valid materials. Consequently, these methods have largely been restricted to classes of materials with abundant data, particularly MOFs. −
An alternative approach for expanding the accessible chemical space of materials, which is especially applicable in cases with limited data, is to use fragmentation methods, where materials are decomposed into constituent fragments that are then recombined into new materials. The fragmentation approach has been successfully used in computational-aided design of materials, including those based on molecular cages, − organic semiconductors, , and molecules for singlet fission. , The results from these studies illustrate the key benefit of the fragmentation approach: using experimentally derived fragments ensures generated structures are chemically valid (i.e., correct bonding and valency) and maximizes their synthesizability. Fragments are also advantageous in terms of optimizing material properties since they are readily encoded as genes, which when combined with a suitable fitness function, can be optimized by evolutionary algorithms. ,,−
Previously our group applied fragmentation methods to enumerate the immediate chemical space of MOPs by combining experimentally reported CBUs and assembly models. , As part of The World Avatar, a universal framework to connect and utilize data across domains, this data set containing 2,217 novel MOPs was integrated into a knowledge graph through the OntoMOPs ontology, which defined the relationships between CBUs, assembly models, MOPs, and properties, allowing for advanced analysis. − Although OntoMOPs successfully captures the core concepts in MOP assembly, the focus on reported CBUs is not well-suited for incorporating hypothetical building units, and thus limits exploration of larger design spaces.
The purpose of this paper is to take the fragmentation approach for MOPs further by introducing a strategy that expands the reticular chemistry space through deconstructing reported organic CBUs into molecular fragments and reassembling them based on templates (Figure ). We demonstrate this approach by enumerating the FragMOPs data set, comprising a total of 98,098 CBUs suitable for MOPs. The combination of these organic CBUs with four assembly models and seven metal CBUs represents a chemical space containing nearly 800,000 MOP configurations, which is captured ontologically through extensions of OntoMOPs. We further present results for optimizing MOP properties within this space through a genetic algorithm (GA) approach that is shown to effectively guide exploration toward MOPs with optimal host–guest properties. In particular, we apply the GA to optimize cavity volumes for a C60 guest and combine the GA with simulations accelerated by a machine learning potential in order to identify MOPs that are estimated to have strong interactions with CO2. This study thus provides not only a method for systematically expanding the chemical space of reticular materials in a chemically meaningful way, but further allows for explainable and effective optimization toward high-performing materials for key host–guest applications in chemical capture, catalysis, and separations.
1.

Computational approaches to exploring the MOP chemical space. Previously, we created the OntoMOPs data set by enumerating all valid combinations of experimentally reported CBUs and assembly models. By contrast, in the FragMOPs approach described here reported organic CBUs are decomposed into molecular fragments and then recombined into new CBUs that can be assembled into orders of magnitude more MOPs using the same metal units and assembly models.
2. Methods
2.1. Molecular Fragmentation
The initial library of organic CBUs was compiled by merging those from the previously reported OntoMOPs data set with the HEALED data set, which contains CBUs derived from experimental MOF structures. Each of the collected CBUs was then processed to fragments by cleaving all exocyclic, non-hydrogen single bonds and replacing each fragmentation point with dummy atoms at either end to denote available bonding sites. This approach was selected to favor cross-coupling sites, however, it is important to note that different fragmentation approaches, such as BRICS, could yield significantly different fragments and design spaces. To allow for subsequent, controlled functionalization, fragmentation points that resulted in fragments with one bonding site, corresponding to a side chain group, were replaced by a hydrogen atom on the parent fragment. Following this, the unique fragments were collected based on canonical SMILES and were categorized according to the number of bonding sites: fragments with one bonding site were marked as side chain fragments, those with two as linker fragments, and those with three or more as node fragments. An additional category for binding group fragments was manually defined based on the carboxylate and pyrazolate binding groups in the OntoMOPs data set, with the defined fragments excluded from the categorization. All final fragments were instantiated in the knowledge graph and linked to data detailing their type, SMILES, extracted geometry, and chemical properties (see Figure S.2).
2.2. CBU Assembler
Given a CBU template and a set of fragments containing at least a binding group fragment and optionally linker and node fragments, the CBU assembler first checks that each position in the template is satisfied and, if so, then orders the linker fragments based on the designated sequence in the template (Figure ). The fragments are then joined starting from the binding group, appending the linker fragments in sequence, and then either appending a second binding group at the end, or if a node fragment is provided, appending copies of the partially assembled structure to each bonding site on the node fragment. The orientations of linker fragments are controlled by setting which bonding site will be connected first, represented as a vector with values of 0 or 1 for each linker. For rapidly enumerating the space SMILES-based assembly is supported. When a geometry is required, an alternative protocol is invoked that operates on the geometries recorded with each fragment. In this geometry-based assembly, a local coordinate system is determined for the current fragment using the selected bonding site dummy atom and the two nearest neighboring atoms. A transformation is then calculated to align this coordinate system with a similarly defined local system on a second fragment (or partially assembled structure) at an appropriate separation and orientation. Effectively, this results in mapping the dummy atom of one bonding site onto the neighboring atom of the other and vice versa. The bonding site dummy atoms are then removed and a new single bond is created between the atoms neighboring the bonding sites of each fragment. At this point, a distance matrix is calculated to check for overlaps between nonbonding atoms. If so, the torsion angle around the new bond is incremented iteratively until either the overlaps are eliminated or the angle exceeds 360 deg, in which case the assembly is failed. Alternatively, through interfacing with RDKit, an option is provided to energy minimize the assembled CBU geometry using the MMFF94 force field. Finally, to make the CBUs able to function as intended, the binding groups are aligned, either to each other in the case of 2-linear CBU templates, or otherwise to a plane that contains the center point of each binding group. For enumerating CBU combinations from a list of fragments and a CBU template, the combinatorial product of the sets of all fragments accepted at each position is determined. By default, if asymmetric linkers are present all unique orientations are attempted for each combination. Alternatively, if the symmetric constraint is enabled, the orientations of asymmetric linkers with more than one position are constrained to alternate. To generate MOPs from the created organic CBUs we use our MOP assembler reported previously, which is a purely geometric approach that positions the building units according to an assembly model. ,
2.

Overview of the relationship between a CBU and its underlying template, shown for a representative example. The template consists of fragment slots that specify both fragment type and sequence position. The position guides how fragments are connected together to form the complete CBU, with linker fragments assigned explicit positions, whereas binding group and node fragments occupy the terminal sequence positions. To realize a CBU requires providing an appropriate fragment for each slot and, in the case of linkers, setting the orientation (0 or 1). An insert below illustrates the linker orientations. Legend: l = linear, c = cyclic, a = acyclic, ★ = bonding site.
2.3. Genetic Algorithm
To efficiently explore and optimize within the FragMOPs design space, we developed a GA adapted to describe the MOP components. In the implementation, each candidate MOP is represented by a chromosome, consisting of genes encoding its assembly model, metal CBU, the molecular fragments occupying each position of the organic CBU template, and, where relevant, the orientations of linker fragments in the template. This chromosome is defined in practice as a sequence of values (i.e. a vector), where, for example, the value at the first position selects the assembly model from a predefined list. The CBU template is not included in the chromosome. Consequently, the lists of assembly models, fragments, and metal CBUs needs to be filtered for each template beforehand, and the length of the chromosome varies depending on the number of fragment slots the CBU template contains. Nevertheless, with the relatively small number of templates, it is simpler and more efficient to parallelize over the templates. This also greatly reduces dimensionality, thereby improving convergence, and allows exploring CBU topologies independently.
To begin the optimization an initial population of chromosomes is generated randomly. Thereafter, at each generation, the ranking of individuals is determined using a defined fitness function, which incorporates properties calculated for the structure returned by the MOP assembler. Our implementation includes elitism so that the top ranked individuals are copied to the next generation. The remainder of the population is created through tournament selection, gene crossover, and mutation, with an option allowing for duplicate chromosomes to be replaced with random ones. In the optimizations reported we use a population of 30 individuals evolved for 30 generations. The mutation probability was set at 0.15, the crossover probability at 0.9, and the elitism proportion was 0.05, with all parameters remaining constant throughout. Further details for each optimization are available in the supplementary files.
2.4. Knowledge Graph and The World Avatar Integration
FragMOPs extends the framework established in OntoMOPs by fully encoding the concepts and relationships between fragments, CBUs, CBU templates, and MOPs within a semantic knowledge graph. This structured representation enables automated reasoning for data validation, constraint enforcement, and complex querying. For example, within the CBU templates, encoded relationships define which fragment types are valid at each position and impose constraints such as atom count, cyclicity, and linearity.
In line with our previous work on reticular materials, FragMOPs is implemented as part of The World Avatar (TWA), a dynamic ecosystem of interconnected data and knowledge models built on FAIR principles. This integration links FragMOPs to other chemical knowledge bases, including OntoSpecies, an ontology describing chemical species. Moreover, TWA hosts a diverse suite of knowledge models, such as OntoCompChem and OntoPESScan, creating future opportunities to interoperate with complementary domains of chemistry knowledge and beyond. Through this integration, FragMOPs contributes to our ongoing work within The World Avatar to develop a holistic knowledge model for reticular materials discovery. −
3. Results
3.1. FragMOPs Chemical Space
Fragmentation of the organic CBUs from the combined data set produced 71 unique fragments following manual curation to prioritize simple and general fragments. This set comprises two binding group fragments (carboxylate and pyrazolate), two node fragments with three bonding sites, five nonlinear noncyclic linkers, 24 side-chain fragments, 20 linear cyclic linkers, and 18 linear noncyclic linkers. A further 24 asymmetric linear linker fragments were created through performing single C–H substitutions on a benzene linear linker with each of the 24 side-chain fragments.
With this fragment library in hand, a series of complementary CBU templates were defined based on 2-linear, 2-bent, and 3-planar CBUs in the OntoMOPs data set, where the notation refers to the number and relative arrangement of the binding site vectors (see Figure S.1). In total eight CBU templates were defined: four 2-linear templates, two 2-bent, and two 3-planar (see Table S.1). These were selected as representative examples that capture common topologies observed in MOPs, providing a set that is diverse yet amenable to detailed analysis for demonstrating the FragMOPs approach. In all cases the templates were defined to accept a single binding fragment such that CBUs with mixed binding groups would not be generated.
To assess the chemical space enclosed by these fragments and templates, we then enumerated all valid CBUs using the SMILES assembly method with all linker orientations allowed. In total this returned 98,098 unique SMILES and was completed in just 138 s on a single CPU core. Compared to the previous OntoMOPs data set that contained 101 organic CBUs, 19 of which align with the defined CBU templates, this represents a 971-fold increase in available CBUs and a 5,163-fold greater sampling of the mutual design space. The considerable expansion in the chemical space is illustrated by a t-SNE projection of Morgan fingerprints calculated for the previous CBUs and the FragMOPs CBUs (Figure A), with the two hemispheres observed in the plot found to correspond to the two binding groups (Figure S.4).
3.
Expansion and characterization of the FragMOPs chemical space. A t-SNE embedding of 2048-bit Morgan fingerprints depicts the chemical space of organic CBUs from the previous OntoMOPs data set and the newly generated FragMOPs CBUs (A) annotated with examples of commercially available CBUs present only in the FragMOPs data set and their type. Distributions of inner-sphere (B) and window (C) diameter for 202 MOPs from the previous data set and 3000 randomly sampled FragMOPs with mutual assembly models illustrates the broader and more continuous range of properties in the FragMOPs data set.
To further characterize this expanded chemical space, we computed several molecular property distributions. The distribution of molecular weight for the FragMOPs CBUs is found to be shifted significantly to a higher mass average of 615.21 g mol–1 compared to the previous organic CBUs, which had an average of 332.73 g mol–1 (Figure S.5). While this is a significant increase, it is largely an expected consequence from the combinatorial increase in CBUs containing two to three linker fragments, which are relatively less common in the previous data set. Additionally, the FragMOPs data set contains heavy fragments, such as iodide, that are not present in the OntoMOPs data set.
We next evaluated synthetic accessibility using the SAscore metric. For the FragMOPs CBUs we find that the mean SAscore is higher than that for the previous organic CBUs, 3.06 to 2.36 (Figure S.6). Interestingly, the maximum SAscore is lower for FragMOPs (5.86 vs 6.50), with the higher value for the previous set of CBUs arising from pyrogallolarene- and metal-containing species outside the FragMOPs design space. Excluding these, the maximum drops to 3.07. Given that a SAscore of around 3 corresponds to the average for commercially available compounds, this analysis indicates that the FragMOPs CBUs are broadly synthetically accessible. Indeed, a search of the ZINC in-stock database identified 100 commercially available CBUs, of which only 16 are present in the previous OntoMOPs data set (full list detailed in Section S.5.3).
For exploring the diversity of MOPs accessible through FragMOPs, we next implemented a workflow using the geometry-based protocol of the CBU assembler without energy optimization to generate geometries which could be input to the MOP assembler. The other inputs required for the MOP assembler are an assembly model and a metal CBU. For the assembly models, we selected three assembly models that are commonly reported experimentally with the CBU templates and feature varying stoichiometries: (3-pyramidal)2(2-bent)3, (3-pyramidal)4(2-linear)6, and (3-pyramidal)4(3-planar)4. Metal CBUs for the 3-pyramidal sites, which form the vertices in these assembly models, were extracted from the OntoMOPs data set, yielding seven CBUs. Under these conditions, the total number of possible MOP configurations is estimated to be 791,056. In comparison to the 202 MOPs from the previous OntoMOPs data set within the same chemical space, this represents increasing the design space by in excess of three orders of magnitude. From this multitude of configurations, we assembled a subset of 1,000 MOPs per assembly model (3,000 total) using randomly selected metal and organic CBUs. For each MOP we calculated the cavity properties including the maximum inner sphere diameter and window diameter to compare with the 202 MOPs from the previous data set. These results indicate that, despite differences in the underlying chemical spaces, the shared CBU templates yield distributions with comparable overall statistics. The average inner-sphere diameters are 10.5 Å for the previous data set and 12.9 Å for the FragMOPs data set, while the average window diameters are 6.6 Å and 7.9 Å, respectively. The FragMOPs data set exhibits larger maximum inner-sphere and window diameters (27.8 Å and 21.2 Å) compared to those of the previous data set (19.3 Å and 14.3 Å), reflecting the inclusion of larger fragments derived from the HEALED data set. Although the overall statistics are similar, the more continuous coverage observed in the FragMOPs distributions highlights its enhanced tunability of cavity properties across the assembly models.
3.2. MOP Optimization
Having created a large and scalable design space for MOPs, we next sought to explore optimizing within the space to identify promising candidates for applications. The fragment-based approach lends itself naturally to optimization with a genetic algorithm, and therefore to assess the potential of such an approach, we examined two applications: optimizing MOP pore volumes for a C60 guest and optimizing average CO2 interaction energies.
3.2.1. Cavity Structure
For host–guest applications a crucial factor is the complementarity between the shape of the guest and the cavity of the host. In the case of C60, we thus chose to optimize the spherical volume of the cavity to be 1.05× the van der Waals volume of C60, which resulted in a target volume of 625 Å3. While we do not expect this is a complete descriptor of host performance, it does represent an inexpensive target to demonstrate optimizing the MOP cavity properties from the CBU fragments. To favor synthetically accessible and structurally simple candidates, the fitness function combined three weighted terms: deviation from the target cavity volume, SAscore, and the number of side chain fragments. The weights were fixed with a relative ratio of 1:10:10. In cases where the only information is the atom connectivity, determining the last term would require significant processing similar to the fragmentation protocol used initially. However, with the FragMOPs knowledge graph, this can be simply queried from the CBU.
With the fitness function determined, we then set the search space to consist of the three assembly models from above, (3-pyramidal)2(2-bent)3, (3-pyramidal)4(2-linear)6, and (3-pyramidal)4(3-planar)4, and three representative metal CBUs, [V3O2(OH)2(HCO2)3], [Zr3O(OH)3(C5H5)3], and [V6O6(OCH3)9(SO4)], corresponding approximately to small, medium, and large sized CBUs. The organic CBU space was restricted by enforcing that the CBUs be (pseudo)symmetric to maintain equivalent coordination environments, which is desirable synthetically to avoid forming mixtures of isomers. Overall, with these settings the search space contains an estimated 51,100 organic CBUs and a total of 153,300 MOP configurations.
The optimization to the target inner sphere volume within the defined search space aggregated across assembly models is shown in Figure . These show that high-fitness candidates with cavity volumes close to the target are generated even within the random initialization, and convergence to optimal solutions is achieved in only a few generations. The fittest MOPs achieve the desired characteristics in terms of cavity size and synthetic accessibility, with inner sphere volumes between 614 and 621 Å3 and SAscore values between 1.57 and 2.59. The population averages also rapidly converge to the target volume, the (3-pyramidal)4(2-linear)6 and (3-pyramidal)4(3-planar)4 converging in around 10 generations while the (3-pyramidal)2(2-bent)3 model takes slightly longer at approximately 15 generations. Analyzing the individual runs (Figure S.14) we expect this is due to template 1 for the 2-bent CBU only having relatively few configurations that reach the target volume, and hence any mutations or crossover with other genes leads to lower fitness individuals. The other CBU templates exhibit the opposite, most of the genes yield MOPs with larger inner sphere volumes, leading to the population averages for the (3-pyramidal)4(2-linear)6 and (3-pyramidal)4(3-planar)4 starting and remaining above the target volume.
4.
Genetic algorithm optimization of the MOP internal cavity volume to accommodate C60. An example gene for a CBU template containing three distinct fragments, two orientated (A), the convergence of the GA trajectories showing the volume of the best ranked and population average at each generation aggregated across the three assembly models (B-D), and the top ranked MOP found for each assembly model encapsulating C60 (E-G). The C60 atoms are colored orange and hydrogen atoms are omitted for clarity.
Overall, only 3,814 MOP configurations, corresponding to 2.5% of the total search space, were evaluated to identify the optimized structures. This efficiency strongly indicates the effectiveness of GA optimization combined with FragMOPs.
3.2.2. CO2 Interaction Energy
As a more complex demonstration of optimization within the FragMOPs space, we next searched for MOPs exhibiting strong CO2 interactions, which could be potential candidates for carbon capture applications. We focused on the chemical space surrounding a known CO2-binding MOP, MOP-1, which is assembled from a 2-linear organic CBU containing three symmetrically arranged linker fragments, the [Zr3O(OH)3(C5H5)3] metal CBU, and the (3-pyramidal)2(2-linear)3 assembly model. Extracting all organic CBUs matching this template from the FragMOPs data set yielded 5,912 unique molecules. Combining these organic units with the same metal CBU and assembly model resulted in 936 successfully assembled MOPs, the remainder being removed due to atom overlaps arising from the narrow geometry of the assembly model.
To estimate the CO2 interaction energies, we employed a Widom insertion approach, involving Monte Carlo sampling of accessible CO2 positions around each MOP and calculating the interaction energy as U int = U MOP+CO2 – U MOP – U CO2 . This is averaged over all insertions to yield the ensemble average, ⟨U int⟩. To balance efficiency and accuracy in these simulations, energy calculations were performed with the small UMA model, a universal machine learning interatomic potential (MLIP) trained across multiple domains, including importantly CO2 adsorption in MOFs. , The high structural and chemical similarity between MOFs and MOPs, provides confidence the model will retain semiquantitative accuracy when applied to our data set. To assess this, we compared isosteric heat of adsorptions (Q st) calculated using the UMA model against experimentally reported values for two literature MOPs (including MOP-1) with resolved crystal structures, obtaining a mean absolute error of 2.79 kJ mol–1 and correct relative ranking (Table S.3). To further probe the model’s accuracy, we sampled a set of CO2 insertions and recalculated them using density functional theory (DFT) at the B97–3c level. The UMA model achieves a strong correlation with the resulting interaction energies (R 2 = 0.996) and an MAE of 0.072 eV (Figure S.17), equivalent to the errors observed for MOFs previously. While this limited benchmark does not completely establish transferability, it indicates that the UMA model is sufficiently reliable to reproduce adsorption trends within the defined search space, and thus, is suitable for guiding the optimization of CO2 interactions. Considering this, we proceeded to calculate ⟨U int⟩ values for all 936 MOPs.
Because these simulations were performed on isolated MOPs, we further assessed whether the resulting ⟨U int⟩ values are representative of material performance. Using crystal structure prediction (CSP) methods, we generated solid-state structures with consistent packing for a subset of the MOPs and computed Q st values for the corresponding crystals. A regression between the isolated and solid-state results gave a strong correlation (R 2 = 0.958, Figure S.15), confirming that the isolated ⟨U int⟩ values serve as a reliable proxy for material performance. Undoubtedly, varying the crystal packing will significantly affect gas adsorption properties. However, it is reasonable to expect that MOPs with the same shape and metal CBU will have similar crystal energy landscapes, this expectation being supported by experimental studies, which often show similar crystal packing for related MOPs. Therefore, optimizing interaction energy with the isolated MOP will most likely lead to higher material performance.
The distribution of average CO2 interaction energies across the 936 assembled MOPs is shown in Figure . The calculated ⟨U int⟩ values range from −0.0828 to −0.326 eV, with a mean of −0.128 eV. MOP-1 exhibits a value of −0.154 eV, and in total 57 MOPs are predicted to have stronger average interactions with CO2. While absolute rankings are subject to the uncertainty of the ML potential, analyzing the interaction energies across the data set reveals clear chemical trends that suggest general guiding principles. For example, averaging ⟨U int⟩ for each fragment across all CBU occurrences shows that more functionalized fragments, especially those involving nitrogen heterocycles, consistently lead to more stable CO2 interactions. Interestingly, the top-ranked fragment is the sulfonate group present in MOP-1, though its limited frequency (two MOPs) could be biasing this result. At the other end, fragments with aliphatic or halogen substituents are found to yield the weakest interactions on average. Additionally, comparing the results for the two binding groups reveals that MOPs featuring the pyrazolate fragment have slightly more stable average ⟨U int⟩ than those featuring the carboxylate fragment, which is found to be statistically significant (Figure S.24). Overall, we find these insights align closely with prior experimental and computational studies that have demonstrated more polarized functionalities tend to increase affinity for CO2 due to stabilizing interactions with the CO2 quadrupole. −
5.
Results from optimizing average CO2 interaction energy for (3-pyramidal)2(2-linear)3 MOPs using organic CBUs containing three linker fragments. The distribution of average CO2 interaction energy for all 936 valid MOPs (A) and the trajectory from the GA optimization (B) are shown with reference to the reported CO2 MOP host, MOP-1. The CBUs estimated to have the minimum, maximum and median ⟨U int⟩ (C) and the four fragments with the lowest and highest average ⟨U int⟩ (D) as well as the average ⟨U int⟩ and standard deviation for both binding fragments (E) are highlighted.
The results from genetic optimization applied to ⟨U int⟩ are also summarized in Figure . Using a fitness function containing solely ⟨U int⟩, The GA rapidly converges to the global minimum of the search space within seven generations, and the population mean surpasses MOP-1 after only a further two generations. Considering that many of the chromosomes do not successfully assemble, it is remarkable the algorithm can identify high-performing candidates rapidly in such a fragmented space. Together, these results underscore the synergy in combining the FragMOPs framework with machine-learning–accelerated simulations to efficiently discover and rationalize promising MOPs for gas capture.
4. Discussion
The FragMOPs approach presented here has enabled the exploration of a chemical space several orders of magnitude larger than that previously created through direct enumeration of known CBUs and assembly models. This substantial expansion was achieved while maintaining realistic molecules through the selection of simple fragments, templates derived from known CBUs, and enforcing constraints on fragment type and symmetry. Moreover, encoding the molecular fragments as genes was found to be an effective method to optimize the assembled MOPs toward desired properties. These results are a significant advancement yet there is considerable scope to develop and extend the FragMOPs approach further, both in terms of chemical space and methods.
First, this study focused on a select number of assembly models and complementary CBU templates and as such there is clear potential to expand the data set through adding more assembly models and CBU templates. Indeed, the OntoMOPs data set includes 20 assembly models, and therefore, if we extrapolate the results we achieved with the selected four assembly models, it is reasonable to expect that the FragMOPs approach could yield millions of unique MOP configurations. This could be further amplified by adding more fragments, which can also expand the diversity of chemistry and functionalities in the resulting MOPs beyond the relatively simple, singly functionalized linkers that were used here. While generative models have an advantage over fragmentation approaches in that they can explore unconstrained chemical spaces beyond the training data, the results presented clearly illustrate that with sufficient fragments the chemical diversity of the CBUs generated is considerable. Moreover, rather than choosing one or the other, we expect that fragmentation and generative models can be combined synergistically to explore large chemical spaces with a high level of control.
An important consideration when expanding the chemical space is to ensure that the geometries of generated CBUs translate to realistic MOP geometries. In this study, the geometries of fragments were kept rigid and the only flexibility was in the torsions between fragments, which were adjusted to eliminate overlaps between nonbonding atoms. While this guarantees the CBUs will have suitable geometries to function as intended, not adapting the conformation within the MOP environment can lead to collisions between the CBUs, which in the results presented were simply marked as failed assemblies and removed from further analysis. Consequently, with the current approach this can lead to a number of valid MOP configurations not being successfully assembled. For example, in creating the data set of 3000 MOPs in Section we recorded 181 failed MOP assemblies. Improving the computational construction of supramolecular assemblies has been demonstrated in previous studies through applying geometry optimization methods, and we expect this would also be effective in the MOP assembler, though careful consideration will be required for how the metal CBUs are modeled as these are not described well by standard force fields such as MMFF94. ,,, Including geometry optimizations would also yield more realistic cavity environments, and thereby, improved downstream evaluation for host–guest applications.
A further refinement that could improve the quality of the results is a stronger focus on the synthesizability, both at the level of the CBU and MOP. While we attempted to incorporate some features of this through assessing the SAscore of CBUs and adding penalties to the fitness function that guided the genetic optimization toward simpler CBUs, these approaches have limitations. For example, SAscore is only an approximate estimate of synthesizability and is less effective for differentiating molecules with similar scores, as is the case with the majority of the FragMOPs CBUs. A potential approach to improve the assessment of CBU synthesizability would be to include retrosynthesis methods to estimate the synthesizability through the number of steps required to reach commercially available reactants, which has been demonstrated to effectively guide reinforcement learning. By contrast, accounting for MOP synthesizability will require including estimates of the formation energy, which are important not only for indicating if a MOP is thermodynamically favored to form, but also, in cases where multiple assemblies are possible, which will be dominant. These calculations can be complex with results that vary depending on synthesis conditions, however, previous studies have shown the effectiveness of including these aspects in modeling supramolecular assemblies. , If possible to efficiently geometry optimize the MOP structures, simpler metrics, such as the strain energy of the organic CBU, could also be used as a proxy for synthesizability. Furthermore, with our approach the integration of the knowledge graph provides an opportunity to involve semantic reasoning based on experimental observations and heuristics in predicting MOP synthesizability.
5. Conclusions
Developing efficient and effective methods to computationally screen materials for advanced applications is a critical area of research toward accelerating materials discovery. In this study we have presented a method for generating an extensive chemical space of organic building units suitable for reticular chemistry based on assembling molecular fragments. We demonstrated applying this method to explore MOPs resulting in the FragMOPs data set containing 98,098 CBUs and encapsulating a configurational space of 791,056 MOPs across four assembly models and seven metal units. Furthermore, we showed that a genetic algorithm applied to the molecular fragments enables effective optimization of the MOP properties, with examples illustrating the optimization of cavity size and CO2 interaction energies estimated from machine-learning-accelerated simulations. While the current implementation has limitations in single CBU conformations and synthesizability information, the FragMOPs approach provides a robust foundation for future developments, including expanded fragment libraries, integrated conformational and retrosynthesis analysis, and thermodynamic evaluations. Overall, this work establishes a scalable, explainable, and data-efficient strategy for the accelerated rational design of high-performing reticular materials for applications.
Supplementary Material
Acknowledgments
This research was supported by the National Research Foundation, Prime Minister’s Office, Singapore under its Campus for Research Excellence and Technological Enterprise (CREATE) programme. This project has received funding from the European Union’s Horizon Europe research and innovation programme under grants 101058732 (JIDEP), 101074004 (C2IMPRESS), and 101188248 (CLIMATE-ADAPT4EOSC). This work was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service (www.csd3.cam.ac.uk), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (capital grant EP/T022159/1), and DiRAC funding from the Science and Technology Facilities Council (www.dirac.ac.uk). M. Kraft gratefully acknowledges the support of the Alexander von Humboldt Foundation and the Massachusetts Institute of Technology. S. D. Rihm acknowledges financial support from Fitzwilliam College, Cambridge, and the Cambridge Trust. For the purpose of open access, the author has applied a Creative Commons Attribution (CC BY) licence to any Author Accepted Manuscript version arising.
All codes and ontologies developed are available on GitHub under MIT license: www.github.com/TheWorldAvatar/MOPTools.
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.5c02956.
Generic shapes of the CBUs and MOP assembly models used in this study; diagram of the FragMOPs ontology; illustration of the progression from the definition of a CBU template to an assembled CBU; summary of fragment types and sequence positions for each CBU template defined; summary of the number of unique CBUs assembled for each template; t-SNE plot; distributions of molecular weight; three highest molecular weight CBUs; and other figures (PDF)
Conceptualization, P.W.V.B.; data curation, P.W.V.B.; methodology, P.W.V.B., S.D.R.; investigation, P.W.V.B.; writing – original draft, P.W.V.B.; writing – review and editing, P.W.V.B., S.D.R, S.M., J.A., and M.K.; funding acquisition, M.K.; supervision, S.M., J.A., and M.K. All authors have read and agreed to the published version of the manuscript.
The authors declare no competing financial interest.
References
- Yaghi O. M., O’Keeffe M., Ockwig N. W., Chae H. K., Eddaoudi M., Kim J.. Reticular Synthesis and the Design of New Materials. Nature. 2003;423:705–714. doi: 10.1038/nature01650. [DOI] [PubMed] [Google Scholar]
- Jiang H., Alezi D., Eddaoudi M.. A Reticular Chemistry Guide for the Design of Periodic Solids. Nature Reviews Materials. 2021;6:466–487. doi: 10.1038/s41578-021-00287-y. [DOI] [Google Scholar]
- Yaghi O. M.. Reticular ChemistryConstruction, Properties, and Precision Reactions of Frameworks. J. Am. Chem. Soc. 2016;138:15507–15509. doi: 10.1021/jacs.6b11821. [DOI] [PubMed] [Google Scholar]
- Ding M., Flaig R. W., Jiang H.-L., Yaghi O. M.. Carbon Capture and Conversion Using Metal–Organic Frameworks and MOF-based Materials. Chem. Soc. Rev. 2019;48:2783–2828. doi: 10.1039/C8CS00829A. [DOI] [PubMed] [Google Scholar]
- Liang W., Bhatt P. M., Shkurenko A., Adil K., Mouchaham G., Aggarwal H., Mallick A., Jamal A., Belmabkhout Y., Eddaoudi M.. A Tailor-Made Interpenetrated MOF with Exceptional Carbon-Capture Performance from Flue Gas. Chem. 2019;5:950–963. doi: 10.1016/j.chempr.2019.02.007. [DOI] [Google Scholar]
- Chen B., Fan D., Pinto R. V., Dovgaliuk I., Nandi S., Chakraborty D., García-Moncada N., Vimont A., McMonagle C. J., Bordonhos M., Al Mohtar A., Cornu I., Florian P., Heymans N., Daturi M., De Weireld G., Pinto M., Nouar F., Maurin G., Mouchaham G., Serre C.. et al. A Scalable Robust Microporous Al-MOF for Post-Combustion Carbon Capture. Advanced Science. 2024;11:2401070. doi: 10.1002/advs.202401070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gosselin A. J., Decker G. E., McNichols B. W., Baumann J. E., Yap G. P., Sellinger A., Bloch E. D.. Ligand-Based Phase Control in Porous Zirconium Coordination Cages. Chem. Mater. 2020;32:5872–5878. doi: 10.1021/acs.chemmater.0c01965. [DOI] [Google Scholar]
- Zeng Y., Zou R., Zhao Y.. Covalent Organic Frameworks for CO2 Capture. Adv. Mater. 2016;28:2855–2873. doi: 10.1002/adma.201505004. [DOI] [PubMed] [Google Scholar]
- D’Alessandro D. M., Smit B., Long J. R.. Carbon Dioxide Capture: Prospects for New Materials. Angew. Chem., Int. Ed. 2010;49:6058–6082. doi: 10.1002/anie.201000431. [DOI] [PubMed] [Google Scholar]
- Liu Y., Wang Z. U., Zhou H.-C.. Recent Advances in Carbon Dioxide Capture with Metal-Organic Frameworks. Greenhouse Gases: Science and Technology. 2012;2:239–259. doi: 10.1002/ghg.1296. [DOI] [Google Scholar]
- Zheng Z., Nguyen H. L., Hanikel N., Li K. K.-Y., Zhou Z., Ma T., Yaghi O. M.. High-Yield, Green and Scalable Methods for Producing MOF-303 for Water Harvesting from Desert Air. Nat. Protoc. 2023;18:136–156. doi: 10.1038/s41596-022-00756-w. [DOI] [PubMed] [Google Scholar]
- Hanikel N., Kurandina D., Chheda S., Zheng Z., Rong Z., Neumann S. E., Sauer J., Siepmann J. I., Gagliardi L., Yaghi O. M.. MOF Linker Extension Strategy for Enhanced Atmospheric Water Harvesting. ACS Central Science. 2023;9:551–557. doi: 10.1021/acscentsci.3c00018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu W., Yaghi O. M.. Metal–Organic Frameworks for Water Harvesting from Air, Anywhere, Anytime. ACS central science. 2020;6:1348–1354. doi: 10.1021/acscentsci.0c00678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu M., Yang Y.. Metal–Organic Framework (MOF)-Based Drug/Cargo Delivery and Cancer Therapy. Adv. Mater. 2017;29:1606134. doi: 10.1002/adma.201606134. [DOI] [PubMed] [Google Scholar]
- Samanta S. K., Moncelet D., Briken V., Isaacs L.. Metal–Organic Polyhedron Capped with Cucurbit [8] Uril Delivers Doxorubicin to Cancer Cells. J. Am. Chem. Soc. 2016;138:14488–14496. doi: 10.1021/jacs.6b09504. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Song G., Shi Y., Jiang S., Pang H.. Recent Progress in MOF-derived Porous Materials as Electrodes for High-Performance Lithium-Ion Batteries. Adv. Funct. Mater. 2023;33:2303121. doi: 10.1002/adfm.202303121. [DOI] [Google Scholar]
- Pu X., Jiang B., Wang X., Liu W., Dong L., Kang F., Xu C.. High-Performance Aqueous Zinc-Ion Batteries Realized by MOF Materials. Nano-micro letters. 2020;12:152. doi: 10.1007/s40820-020-00487-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen T., Wang F., Cao S., Bai Y., Zheng S., Li W., Zhang S., Hu S., Pang H.. In Situ Synthesis of MOF-74 Family for High Areal Energy Density of Aqueous Nickel–Zinc Batteries. Adv. Mater. 2022;34:2201779. doi: 10.1002/adma.202201779. [DOI] [PubMed] [Google Scholar]
- Avci G., Erucar I., Keskin S.. Do New MOFs Perform Better for CO2 Capture and H2 Purification? Computational Screening of the Updated MOF Database. ACS Appl. Mater. Interfaces. 2020;12:41567–41579. doi: 10.1021/acsami.0c12330. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhao G.. et al. CoRE MOF DB: A Curated Experimental Metal-Organic Framework Database with Machine-Learned Properties for Integrated Material-Process Screening. Matter. 2025;8:102140. doi: 10.1016/j.matt.2025.102140. [DOI] [Google Scholar]
- Kancharlapalli S., Snurr R. Q.. High-Throughput Screening of the CoRE-MOF-2019 Database for CO2 Capture from Wet Flue Gas: A Multi-Scale Modeling Strategy. ACS Appl. Mater. Interfaces. 2023;15:28084–28092. doi: 10.1021/acsami.3c04079. [DOI] [PubMed] [Google Scholar]
- Chung Y. G., Camp J., Haranczyk M., Sikora B. J., Bury W., Krungleviciute V., Yildirim T., Farha O. K., Sholl D. S., Snurr R. Q.. Computation-Ready, Experimental Metal–Organic Frameworks: A Tool to Enable High-Throughput Screening of Nanoporous Crystals. Chem. Mater. 2014;26:6185–6192. doi: 10.1021/cm502594j. [DOI] [Google Scholar]
- Wilmer C. E., Leaf M., Lee C. Y., Farha O. K., Hauser B. G., Hupp J. T., Snurr R. Q.. Large-Scale Screening of Hypothetical Metal–Organic Frameworks. Nature Chem. 2012;4:83–89. doi: 10.1038/nchem.1192. [DOI] [PubMed] [Google Scholar]
- Nandy A., Yue S., Oh C., Duan C., Terrones G. G., Chung Y. G., Kulik H. J.. A Database of Ultrastable MOFs Reassembled from Stable Fragments with Machine Learning Models. Matter. 2023;6:1585–1603. doi: 10.1016/j.matt.2023.03.009. [DOI] [Google Scholar]
- Colón Y. J., Gómez-Gualdrón D. A., Snurr R. Q.. Topologically Guided, Automated Construction of Metal–Organic Frameworks and Their Evaluation for Energy-Related Applications. Cryst. Growth Des. 2017;17:5801–5810. doi: 10.1021/acs.cgd.7b00848. [DOI] [Google Scholar]
- Boyd P. G., Woo T. K.. A Generalized Method for Constructing Hypothetical Nanoporous Materials of Any Net Topology from Graph Theory. CrystEngComm. 2016;18:3777–3792. doi: 10.1039/C6CE00407E. [DOI] [Google Scholar]
- Lee S., Kim B., Cho H., Lee H., Lee S. Y., Cho E. S., Kim J.. Computational Screening of Trillions of Metal–Organic Frameworks for High-Performance Methane Storage. ACS Appl. Mater. Interfaces. 2021;13:23647–23654. doi: 10.1021/acsami.1c02471. [DOI] [PubMed] [Google Scholar]
- Yao Z., Sánchez-Lengeling B., Bobbitt N. S., Bucior B. J., Kumar S. G. H., Collins S. P., Burns T., Woo T. K., Farha O. K., Snurr R. Q., Aspuru-Guzik Á.. Inverse Design of Nanoporous Crystalline Reticular Materials with Deep Generative Models. Nature Machine Intelligence. 2021;3:76–86. doi: 10.1038/s42256-020-00271-1. [DOI] [Google Scholar]
- Kang Y., Kim J.. ChatMOF: An Artificial Intelligence System for Predicting and Generating Metal-Organic Frameworks Using Large Language Models. Nat. Commun. 2024;15:4705. doi: 10.1038/s41467-024-48998-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Badrinarayanan S., Magar R., Antony A., Meda R. S., Barati Farimani A.. MOFGPT: Generative Design of Metal–Organic Frameworks Using Language Models. J. Chem. Inf. Model. 2025;65:9049–9060. doi: 10.1021/acs.jcim.5c01625. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Inizan, T. J. ; Yang, S. ; Kaplan, A. ; Lin, Y. -h. ; Yin, J. ; Mirzaei, S. ; Abdelgaid, M. ; Alawadhi, A. H. ; Cho, K. ; Zheng, Z. ; Cubuk, E. D. ; Borgs, C. ; Chayes, J. T. ; Persson, K. A. ; Yaghi, O. M. . System of Agentic AI for the Discovery of Metal-Organic Frameworks. 2025; 10.48550/arXiv.2504.14110.
- Park J., Lee Y., Kim J.. Multi-Modal Conditional Diffusion Model Using Signed Distance Functions for Metal-Organic Frameworks Generation. Nat. Commun. 2025;16:34. doi: 10.1038/s41467-024-55390-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Merchant A., Batzner S., Schoenholz S. S., Aykol M., Cheon G., Cubuk E. D.. Scaling Deep Learning for Materials Discovery. Nature. 2023;624:80–85. doi: 10.1038/s41586-023-06735-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tarzia A., Lewis J. E. M., Jelfs K. E.. High-Throughput Computational Evaluation of Low Symmetry Pd2L4 Cages to Aid in System Design. Angew. Chem., Int. Ed. 2021;60:20879–20887. doi: 10.1002/anie.202106721. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lewis J. E. M., Tarzia A., White A. J. P., Jelfs K. E.. Conformational Control of Pd2L4 Assemblies with Unsymmetrical Ligands. Chemical Science. 2020;11:677–683. doi: 10.1039/C9SC05534G. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tarzia A., Jelfs K. E.. Unlocking the Computational Design of Metal–Organic Cages. Chem. Commun. 2022;58:3717–3730. doi: 10.1039/D2CC00532H. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Young T. A., Gheorghe R., Duarte F.. Cgbind: A Python Module and Web App for Automated Metallocage Construction and Host–Guest Characterization. J. Chem. Inf. Model. 2020;60:3546–3557. doi: 10.1021/acs.jcim.0c00519. [DOI] [PubMed] [Google Scholar]
- Cheng C. Y., Campbell J. E., Day G. M.. Evolutionary Chemical Space Exploration for Functional Materials: Computational Organic Semiconductor Discovery. Chemical science. 2020;11:4922–4933. doi: 10.1039/D0SC00554A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Johal, J. ; Day, G. . Exploring Organic Chemical Space for Materials Discovery Using Crystal Structure Prediction-Informed Evolutionary Optimisation. 2025; 10.26434/chemrxiv-2025-v8692. [DOI] [PMC free article] [PubMed]
- Worakul T., Laplaza R., Blaskovits J. T., Corminboeuf C.. Generative Design of Singlet Fission Materials Leveraging a Fragment-Oriented Database. Chem. Sci. 2025;16:17956–17969. doi: 10.1039/D5SC03184B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blaskovits J. T., Laplaza R., Vela S., Corminboeuf C.. Data-Driven Discovery of Organic Electronic Materials Enabled by Hybrid Top-Down/Bottom-Up Design. Adv. Mater. 2024;36:2305602. doi: 10.1002/adma.202305602. [DOI] [PubMed] [Google Scholar]
- Schaufelberger L., Blaskovits J. T., Laplaza R., Jorner K., Corminboeuf C.. Inverse Design of Singlet-Fission Materials with Uncertainty-Controlled Genetic Optimization. Angew. Chem., Int. Ed. 2025;64:e202415056. doi: 10.1002/anie.202415056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Berardo E., Turcani L., Miklitz M., Jelfs K. E.. An Evolutionary Algorithm for the Discovery of Porous Organic Cages. Chemical Science. 2018;9:8513–8527. doi: 10.1039/C8SC03560A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Le T. C., Winkler D. A.. Discovery and Optimization of Materials Using Evolutionary Approaches. Chem. Rev. 2016;116:6107–6132. doi: 10.1021/acs.chemrev.5b00691. [DOI] [PubMed] [Google Scholar]
- Nigam A., Pollice R., Friederich P., Aspuru-Guzik A.. Artificial design of organic emitters via a genetic algorithm enhanced by a deep neural network. Chem. Sci. 2024;15:2618–2639. doi: 10.1039/D3SC05306G. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hibbert D. B.. Genetic algorithms in chemistry. Chemometrics and Intelligent Laboratory Systems. 1993;19:277–293. doi: 10.1016/0169-7439(93)80028-G. [DOI] [Google Scholar]
- Jennings P. C., Lysgaard S., Hummelshøj J. S., Vegge T., Bligaard T.. Genetic algorithms for computational materials discovery accelerated by machine learning. npj Comput. Mater. 2019;5:46. doi: 10.1038/s41524-019-0181-4. [DOI] [Google Scholar]
- Kim C., Batra R., Chen L., Tran H., Ramprasad R.. Polymer design using genetic algorithm and machine learning. Comput. Mater. Sci. 2021;186:110067. doi: 10.1016/j.commatsci.2020.110067. [DOI] [Google Scholar]
- Keser M., Stupp S. I.. Genetic algorithms in computational materials science and engineering: simulation and design of self-assembling materials. Computer Methods in Applied Mechanics and Engineering. 2000;186:373–385. doi: 10.1016/S0045-7825(99)00392-8. [DOI] [Google Scholar]
- Cao R., Zhang C.-R., Liu X.-M., Gong J.-J., Zhang M.-L., Liu Z.-J., Wu Y.-Z., Chen H.-S.. Molecular design of organic photovoltaic donors and non-fullerene acceptors: a combined machine learning and genetic algorithm approach. J. Mater. Chem. C. 2025;13:12150–12168. doi: 10.1039/D5TC02014J. [DOI] [Google Scholar]
- Kondinski A., Menon A., Nurkowski D., Farazi F., Mosbach S., Akroyd J., Kraft M.. Automated Rational Design of Metal–Organic Polyhedra. J. Am. Chem. Soc. 2022;144:11713–11728. doi: 10.1021/jacs.2c03402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bai J., Rihm S. D., Kondinski A., Saluz F., Deng X., Brownbridge G., Mosbach S., Akroyd J., Kraft M.. Twa: The World Avatar Python Package for Dynamic Knowledge Graphs and Its Application in Reticular Chemistry. Digital Discovery. 2025;4:2123–2135. doi: 10.1039/D5DD00069F. [DOI] [Google Scholar]
- Rihm S. D., Tran D. N., Kondinski A., Pascazio L., Saluz F., Deng X., Mosbach S., Akroyd J., Kraft M.. Natural Language Access Point to Digital Metal–Organic Polyhedra Chemistry in The World Avatar. Data-Centric Eng. 2025;6:e22. doi: 10.1017/dce.2025.12. [DOI] [Google Scholar]
- Kondinski A., Oyarzún-Aravena A. M., Rihm S. D., Bai J., Mosbach S., Akroyd J., Kraft M.. Automated Assembly Modelling of Metal-Organic Polyhedra. Eur. J. Inorg. Chem. 2025;28:e202500115. doi: 10.1002/ejic.202500115. [DOI] [Google Scholar]
- Rihm S. D., Saluz F., Kondinski A., Bai J., Butler P. W. V., Mosbach S., Akroyd J., Kraft M.. Extraction of Chemical Synthesis Information Using the World Avatar. Digital . Discovery. 2025;4:2893–2909. doi: 10.1039/D5DD00183H. [DOI] [Google Scholar]
- Gibaldi M., Kwon O., White A., Burner J., Woo T. K.. The HEALED SBU Library of Chemically Realistic Building Blocks for Construction of Hypothetical Metal–Organic Frameworks. ACS Appl. Mater. Interfaces. 2022;14:43372–43386. doi: 10.1021/acsami.2c13100. [DOI] [PubMed] [Google Scholar]
- Degen J., Wegscheid-Gerlach C., Zaliani A., Rarey M.. On the Art of Compiling and Using ’Drug-Like’ Chemical Fragment Spaces. ChemMedChem. 2008;3:1503–1507. doi: 10.1002/cmdc.200800178. [DOI] [PubMed] [Google Scholar]
- Tosco P., Stiefl N., Landrum G.. Bringing the MMFF Force Field to the RDKit: Implementation and Validation. J. Cheminf. 2014;6:37. doi: 10.1186/s13321-014-0037-3. [DOI] [Google Scholar]
- Lim, M. Q. ; Wang, X. ; Inderwildi, O. ; Kraft, M. . Intelligent Decarbonisation: Can. Artificial Intelligence and Cyber-Physical Systems Help Achieve Climate Mitigation Targets?; Springer International Publishing, 2022; pp 39–53. [Google Scholar]
- Pascazio L., Rihm S., Naseri A., Mosbach S., Akroyd J., Kraft M.. Chemical Species Ontology for Data Integration and Knowledge Discovery. J. Chem. Inf. Model. 2023;63:6569–6586. doi: 10.1021/acs.jcim.3c00820. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Krdzavac N., Mosbach S., Nurkowski D., Buerger P., Akroyd J., Martin J., Menon A., Kraft M.. An Ontology and Semantic Web Service for Quantum Chemistry Calculations. J. Chem. Inf. Model. 2019;59:3154–3165. doi: 10.1021/acs.jcim.9b00227. [DOI] [PubMed] [Google Scholar]
- Menon A., Pascazio L., Nurkowski D., Farazi F., Mosbach S., Akroyd J., Kraft M.. OntoPESScan: An Ontology for Potential Energy Surface Scans. ACS omega. 2023;8:2462–2475. doi: 10.1021/acsomega.2c06948. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ertl P., Schuffenhauer A.. Estimation of Synthetic Accessibility Score of Drug-like Molecules Based on Molecular Complexity and Fragment Contributions. J. Cheminf. 2009;1:8. doi: 10.1186/1758-2946-1-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Irwin J. J., Tang K. G., Young J., Dandarchuluun C., Wong B. R., Khurelbaatar M., Moroz Y. S., Mayfield J., Sayle R. A.. ZINC20a Free Ultralarge-Scale Chemical Database for Ligand Discovery. J. Chem. Inf. Model. 2020;60:6065–6073. doi: 10.1021/acs.jcim.0c00675. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xing W.-H., Li H.-Y., Dong X.-Y., Zang S.-Q.. Robust Multifunctional Zr-based Metal–Organic Polyhedra for High Proton Conductivity and Selective CO2 Capture. Journal of Material Chemistry A. 2018;6:7724–7730. doi: 10.1039/C8TA00858B. [DOI] [Google Scholar]
- Wood, B. M. et al. UMA: A Family of Universal Models for Atoms. 2025; 10.48550/arXiv.2506.23971. [DOI]
- Sriram, A. ; Brabson, L. M. ; Yu, X. ; Choi, S. ; Abdelmaqsoud, K. ; Moubarak, E. ; de Haan, P. ; Löwe, S. ; Brehmer, J. ; Kitchin, J. R. ; Welling, M. ; Zitnick, C. L. ; Ulissi, Z. ; Medford, A. J. ; Sholl, D. S. . The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture. 2025; 10.48550/arXiv.2508.03162. [DOI]
- Torrisi A., Bell R. G., Mellot-Draznieks C.. Functionalized MOFs for Enhanced CO2 Capture. Cryst. Growth Des. 2010;10:2839–2841. doi: 10.1021/cg100646e. [DOI] [Google Scholar]
- Das A., D’Alessandro D. M.. Tuning the Functional Sites in Metal–Organic Frameworks to Modulate CO 2 Heats of Adsorption. CrystEngComm. 2015;17:706–718. doi: 10.1039/C4CE01341G. [DOI] [Google Scholar]
- Gładysiak A., Song A.-Y., Vismara R., Waite M., Alghoraibi N. M., Alahmed A. H., Younes M., Huang H., Reimer J. A., Stylianou K. C.. Enhanced Carbon Dioxide Capture from Diluted Streams with Functionalized Metal–Organic Frameworks. JACS Au. 2024;4:4527–4536. doi: 10.1021/jacsau.4c00923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Alsmail N. H., Suyetin M., Yan Y., Cabot R., Krap C. P., Lü J., Easun T. L., Bichoutskaia E., Lewis W., Blake A. J., Schröder M.. et al. Analysis of High and Selective Uptake of CO2 in an Oxamide-Containing {Cu2 (OOCR) 4}-Based Metal–Organic Framework. Chem. – Eur. J. 2014;20:7317–7324. doi: 10.1002/chem.201304005. [DOI] [PubMed] [Google Scholar]
- Dautzenberg E., Li G., De Smet L. C.. Aromatic Amine-Functionalized Covalent Organic Frameworks (COFs) for CO2/N2 Separation. ACS Appl. Mater. Interfaces. 2023;15:5118–5127. doi: 10.1021/acsami.2c17672. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang H., Peng J., Li J.. Ligand Functionalization in Metal–Organic Frameworks for Enhanced Carbon Dioxide Adsorption. Chem. Rec. 2016;16:1298–1310. doi: 10.1002/tcr.201500307. [DOI] [PubMed] [Google Scholar]
- Emerson A. J., Chahine A., Batten S. R., Turner D. R.. Synthetic Approaches for the Incorporation of Free Amine Functionalities in Porous Coordination Polymers for Enhanced CO2 Sorption. Coord. Chem. Rev. 2018;365:1–22. doi: 10.1016/j.ccr.2018.02.012. [DOI] [Google Scholar]
- Turcani L., Tarzia A., Szczypiński F. T., Jelfs K. E.. Stk: An Extendable Python Framework for Automated Molecular and Supramolecular Structure Assembly and Discovery. J. Chem. Phys. 2021;154:214102. doi: 10.1063/5.0049708. [DOI] [PubMed] [Google Scholar]
- Turcani L., Berardo E., Jelfs K. E.. Stk: A Python Toolkit for Supramolecular Assembly. J. Comput. Chem. 2018;39:1931–1942. doi: 10.1002/jcc.25377. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Guo J., Schwaller P.. Directly Optimizing for Synthesizability in Generative Molecular Design Using Retrosynthesis Models. Chemical Science. 2025;16:6943–6956. doi: 10.1039/D5SC01476J. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yadav A. K., Gładysiak A., Wolpert E. H., Ganose A. M., Samel-Garloff B., Koley D., Jelfs K. E., Stylianou K. C.. Solvatomorphic Diversity Dictates the Stability and Solubility of Metal–Organic Polyhedra. Chemical Science. 2025;16:2589–2599. doi: 10.1039/D4SC05037A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Santolini V., Tribello G. A., Jelfs K. E.. Predicting Solvent Effects on the Structure of Porous Organic Molecules. Chem. Commun. 2015;51:15542–15545. doi: 10.1039/C5CC05344G. [DOI] [PubMed] [Google Scholar]
- Livas C. G., Trikalitis P. N., Froudakis G. E.. MOFSynth: A Computational Tool toward Synthetic Likelihood Predictions of MOFs. J. Chem. Inf. Model. 2024;64:8193–8200. doi: 10.1021/acs.jcim.4c01298. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All codes and ontologies developed are available on GitHub under MIT license: www.github.com/TheWorldAvatar/MOPTools.



