Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2022 Oct 31;62(22):5525–5535. doi: 10.1021/acs.jcim.2c00690

CAT: A Compound Attachment Tool for the Construction of Composite Chemical Compounds

Bas van Beek #, Juliette Zito β,†, Lucas Visscher #,*, Ivan Infante †,‡,¶,*
PMCID: PMC9976287  PMID: 36314636

Abstract

graphic file with name ci2c00690_0008.jpg

The continuous improvement of computer architectures allows for the simulation of molecular systems of growing sizes. However, such calculations still require the input of initial structures, which are also becoming increasingly complex. In this work, we present CAT, a Compound Attachment Tool (source code available at https://github.com/nlesc-nano/CAT) and Python package for the automatic construction of composite chemical compounds, which supports the functionalization of organic, inorganic, and hybrid organic–inorganic materials. The CAT workflow consists in defining the anchoring sites on the reference material, usually a large molecular system denoted as a scaffold, and on the molecular species that are attached to it, i.e., the ligands. Usually, ligands are pre-optimized in a conformation biased toward more linear structures to minimize interligand(s) steric interactions, a bias that is important when multiple ligands are attached onto the scaffold. The resulting superstructure(s) are then stored in various formats that can be used afterward in quantum chemical calculations or classical force field-based simulations.

1. Introduction

Advances in computer hardware architectures are opening new routes to allow simulations of large composite chemical systems. Meanwhile, recent breakthroughs in the application of machine learning algorithms in computational chemistry, for example, for the prediction of chemico-physical properties of chemical structures, require efficient preparation and execution of thousands of calculations to generate a sufficient amount of training data.1−4 For all such atomistic simulations, a recurring task is the generation of initial structures consisting of Cartesian coordinates of all atoms in the system. Without these structures, no calculation is possible, and in many cases, this step forms a critical bottleneck in the workflow of designing and running the simulation on a supercomputer.5,6 In particular, the functionalization with organic molecules of nano- or bulk materials, whether they are organic, inorganic, or mixed organic–inorganic, can turn in a complex task because structures are more varied and less predictable by low-level theories. In these cases, there is a pressing need for fast and simple tools to aid in structure generation.

Usually, building initial structures for the nano- or bulk materials to be functionalized, also named as scaffolds from now on, is nontrivial; however, there are already tools available to tackle their preparation. In the case of purely inorganic scaffolds, their (crystal) structure can be generated using software tools like the pymatgen libraries7 and the Atomic Simulation Environment (ASE),8 which are able to build 2D and 3D structures of arbitrary ion combinations. These periodic structures, as well as nonperiodic cuts of a number of unit cells, can be translated into input formats for a range of quantum chemistry codes such as VASP,9−11 Quantum Espresso,12,13 ADF,14 Gaussian,15etc. Computational results produced on the basis of these structures can subsequently be stored in databases such as the Materials Project,16 NOMAD,17 AiiDA,18 or ioChem-BD19 to allow for data mining and machine learning applications. Nowadays, all these databases offer facilities to quickly browse the properties of a plethora of inorganic materials. In the case of organic scaffolds, such as dyes,20 proteins,21,22 and catalysts,23 there are three common ways to generate atomistic structures. The simplest is to manually construct the desired structure with the aid of a graphical user interface (e.g., ADF-GUI,14 Molden,24 Avogadro,25etc.). This method works well when a few medium-sized molecules made of dozens of atoms need to be studied and has the advantage that the researcher has full control over the generated structures. The second method is to first (automatically) extract molecular structures from available databases such as PubChem,26 GDB-13,27 or QM7.28 These databases typically only provide molecules in SMILES29,30 or SELFIES31 formats meaning that these molecules still need to be converted into three-dimensional structural models with the aid of cheminformatics tools such as RDKit32 or OpenBabel.33 The third approach is to not select a particular molecule beforehand but instead design one with desired properties via generative adversarial networks,34 a machine learning-based architecture nowadays mostly employed for drug discovery.35−37 Finally, in the case of hybrid (organic–inorganic) material scaffolds, we rely either on building the desired structures manually using a graphical user interface or by mining available datasets built on purpose for a specific set of materials. A well-known example of hybrid materials with a large chemical space is represented by metal–organic frameworks (MOFs), where the growing variety of 3D porous MOFs and MOF-like frameworks38,39 has endeavored the creation of constantly updated machine-readable databases40,41 to store the available crystal structures.42 Another class of mixed materials is represented by hybrid perovskite materials of ABX3 composition, characterized by an inorganic network of corner-sharing BX6 octahedra (where B is a metal and X is a halogen) with organic A cations occupying the voids in between.43 Typical examples include CH3NH3PbI3 and CH(NH2)2SnI3, which employ respectively the methylammonium and formamidinium as A cations.

In most cases, if not all, the scaffolds described above are functionalized by a plethora of organic ligands introducing additional challenges, since the combination of organic molecules and inorganic materials leads to an explosion in the number of possible combinations and relevant structures that need to be built. Examples of functionalized materials include inorganic surfaces terminated with organic surfactants such as in self-assembled monolayers (SAMs)44 and colloidal nanoparticles.45,46 The latter represents a wide class of novel materials with promising optoelectronic and/or catalytic properties.47,48 Examples are nanocrystal semiconductors consisting of group II-VI, III-V, and IV-VI elements (examples being HgTe, InP, and PbS, respectively),49−51 metal halide perovskites,52 and nanoparticles consisting of purely metallic systems, such as gold, silver, and platinum.53−55 A key role played by the ligands here is stabilizing the inorganic scaffolds in organic solvents, preventing their dissolution.56 The suitability of a certain ligand for this purpose depends on factors such as interligand packing, magnitude of the scaffold/ligand binding strength, and ligand/solvent interactions.57−59 These criteria, augmented by practical considerations such as the ease of synthesis and material costs, in principle define an optimization target for ligand passivation. In the experimental practice, however, only a limited set of ligands is currently employed on a regular basis. For example, in colloidal nanocrystals, oleic acid and oleylamine are the most known passivating agents due their availability and low cost. Nevertheless, there remains a large unexplored space of potentially interesting ligands available within datasets such as Pubchem,26 GDB-13,27 and QM7.28

As the vast chemical space prohibits high-throughput experimental screening, an (initial) computational screening provides an attractive alternative. To facilitate such screening and allow for the exploration of many ligand-passivated scaffolds, it is desirable to have a tool that can automate the preparation of starting structures for these complex composite systems. While a number of such tools have recently emerged, e.g., to obtain molecular models of organometallic species,60−63 no counterparts are currently available for more generalized scaffolds. We thus have developed a Python library called CAT (Compound Attachment Tool), which is intended for the construction of composite chemical compounds resulting from the functionalization of organic, inorganic, or mixed organic–inorganic materials with organic molecules. CAT has already been introduced, in an early prototype form, for the automatic generation of new dyes’ structures obtained by addition of one or more ligands (alcohols, amines, thiophenes, etc.) to an organic scaffold (1,4,5,8-naphthalenediimide).64 To disclose the full potential and general applicability of our tool, we decided to apply it on ligand-passivated colloidal nanocrystals and metal–organic frameworks, i.e., on the functionalization of composite scaffolds. To facilitate the construction of all these composite chemical compounds, CAT possesses several features that we will discuss in more detail below. Features include, among others, automatic ligand functional group recognition, partial surface passivation of the scaffold, and minimization of interligand steric interactions via a biased conformational search.

2. Results and Discussion

2.1. CAT Workflow

The user of CAT should provide information about the type of structure that is to be generated, which includes the scaffold and the ligands that are to be attached to it. These user-specified settings are provided via a YAML file65 (Figure S1), a human-readable format aimed at data serialization. The two most important options to be provided are the input_cores and input_ligands keywords, the former representing a common moiety or scaffold (e.g. a CdSe nanocrystal) that will be functionalized with the latter (e.g. a carboxylic acid or its carboxylate base conjugate) at specific anchoring points. The dummy atoms serving as anchoring sites should generally be provided by the user. If unspecified, CAT will either default to chlorine for the scaffold or a wide range of common functional groups for the ligand, including, among others, hydroxides, amines, and phosphines. In the model used for illustration, the core dummy atoms are the chlorides, indicated by a green color (Figure 1). If multiple cores and/or ligands are provided, all possible combinations between a single scaffold and a ligand will be automatically constructed.

Figure 1.

Figure 1

Schematic overview of the general CAT workflow.

As we have constructed our software on top of the PLAMS66 library for automating molecular simulations, we can import structures in a variety of formats. These formats include, for instance, SMILES strings,29,30 Protein Databank format (PDB), XYZ, etc.(67) Regardless of the input, the structures are first converted to PLAMS Molecule objects, which are then internally used by CAT for representing the various chemical compounds. For the organic ligands, there is the advantage of being able to work with compact string notation schemes such as SYBYL,68 InChI,69 or the aforementioned SMILES format.29,30 As explained in the Introduction, for the scaffolds, the structure generation is cumbersome because it usually represents a nontrivial structure that the user attempts to functionalize. Examples are inorganic semiconductor nanocrystals (as is the focus of our current work), but similar considerations hold for proteins, enzymes, dyes, and metal–organic frameworks. For this reason, one generally has the preconstructed scaffold (also called the core in the following when the method is applied to semiconductor nanocrystals) and passes its full structural information in the form of XYZ or PDB data. In the case of semiconductor nanocrystals, the builder NanoCrystal,70 which generates nanostructures with desired facets based on the Wulff construction method,71 can serve as a good starting point for generating initial scaffold structures.

Just as important as the actual structures is the definition of anchoring sites (i.e., specific atoms) that determines how the scaffolds and ligands can be combined into a single structure. The specification of scaffold and ligand anchor atoms will be discussed in detail in their respective sections, including their relationship with the split and anchor options displayed in Figure S1.

2.2. Distribution of the Scaffold Anchoring Sites

Starting from a preprocessed scaffold (core), the core anchoring sites are user-specified dummy atoms that define where on the core’s surface the ligands can be attached. These dummy atoms are removed after the ligand is attached and can either be enumerated explicitly as a set of atomic indices or indicated by specific atomic symbols that should be considered anchoring sites.

The latter approach of treating all (dummy) atoms of a particular type as anchoring sites has the advantage of greatly simplifying the input, especially if one is interested in a full passivation of the scaffold surface at specific anchoring points. Whether such full passivation is appropriate for a particular scaffold depends on several factors, including the ligand size, the spacing between the anchoring sites on the scaffold, and the mutual attraction between the scaffold and the ligand. Due to above-mentioned caveats, a few schemes are available to automate the partial passivation of the scaffold (Figure 2a). This is especially important in cases where many ligands functionalize a scaffold, a typical case encountered in colloidal nanocrystals. For partial surface passivation, the full set of possible anchor atoms is reduced to a subset (of user-specified size) that can be distributed according to one or more of the following criteria:

  • Uniform: Maximization of the (weighted) nearest-neighbor distance.

  • Cluster: Minimization of the (weighted) nearest-neighbor distance.

  • Random: Atoms are picked at random.

Figure 2.

Figure 2

(a) Examples of a “uniform”, “cluster”, and “random” distribution of anchoring points (dummy atoms represented in purple color) on the surface that will be converted into organic ligands. The ligand surface coverage is set to 33%. The color code is provided to improve visual clarity between the various facets. The green color schematically indicates the atoms (whatever their type is) of the cubic lattice that will remain unchanged when dummy atoms are converted into organic ligands. (b) Core polyhedron representation. The vector p herein represents the shortest path between the atoms i and j along the surface.

Note that one can freely combine the three distributions, allowing for a large degree of customization. For example, one could create a uniform distribution of n-sized clusters or a clustered/uniform distribution mixed with a degree of randomness (see Figure 2a).

As mentioned previously, the general idea behind the “uniform” and “cluster” distribution is, respectively, to maximize or minimize the nearest-neighbor distances. While exploring all possible combinations would in principle allow one to find a global optimum, such an approach will readily become prohibitively expensive as the number of combinations grows exponentially. Rather than attempting a full optimization, we therefore employ an iterative algorithm in which each to-be-passivated anchoring site is defined by maximizing/minimizing, respectively for “uniform”/“cluster” distributions, the weighted distance with respect to all previously picked anchoring sites (eq 1). As no previously picked anchor exists during the first iteration, the algorithm instead defaults to whichever anchor maximizes/minimizes the weighted distance with respect to all other anchoring atoms in the initial superset.

2.2. 1

In eq 1, the variable Inline graphic represents the distance matrix constructed from all-atom pairs in the initial superset of n atoms. Inline graphic is a vector of unique atomic indices representing up to n atoms. For the “cluster” distribution, the argmax operation, as is used for ai = 0, is substituted for argmin, while argmax is used for all remaining elements. The default weighting factor, f(x) = e–x, ensures that nearest neighbors enter the equation with the largest weight.

By default, the distance matrix utilized in eq 1 represents the shortest paths between anchor–atom pairs through space. While this through-space definition is reasonable, especially with the increased weight of nearest neighbors, the distance between anchoring sites is in practice constrained by the scaffold surface. A more realistic representation of the distance criterion can therefore be desirable, i.e., one based on the length of paths over the scaffold surface that connect these sites. Such a definition can be implemented by approximating the core as a polyhedron, constructed using SciPy’s implementation of the Qhull convex hull algorithm,72 and constraining all allowed paths to a graph defined by its edges. As is illustrated in Figure 2b, the result is a distance matrix that represents the shortest path along a scaffold’s (approximate) surface, rather than the shortest path through space.

All the scaffold–ligand structures obtained in this way can be used later on (after attachment of the ligands) as a starting point for geometry optimizations to verify which of the distribution is more stable, as well as employed as initial frames for molecular dynamic simulations, for example, for the replica exchange type of runs.

2.3. Ligand Anchors: Functional Group Recognition

Like their counterparts in the core, ligand anchors are defined by sites that are to be attached to the core’s anchors. In contrast to the definition of the core anchors, which require the use of dummy atoms, the ligand anchoring sites are located within functional groups and can thus readily be identified. Once a functional group (e.g,, a carboxylate) and an anchoring site within the aforementioned group (i.e., the oxyanion) are defined, all the required information is available for the automatic identification (by searching for a characteristic pattern of atoms and bonds) of ligand anchoring sites for a given ligand specified in the input (see the “optional/ligand/anchor” option in Figure S1).

The specification of functional groups is done using SMILES strings,29,30 which provides a compact representation of organic systems. The user inputs desired functional groups in SMILES format to be anchored to the scaffold. Then, the user-provided ligand database is screened (by matching connectivity patterns with the aid of the RDKit library29,32) to select those ligands that contain the specified functional groups. These candidate ligands are henceforth referred to as “proto-ligands” and contain at least one of the desired functional groups and possibly more. A single copy is subsequently created of a proto-ligand for each valid functional group that it contains. These unique combinations of proto-ligand and functional group define a “ligand”, which is passed further along the workflow. For example, if we look for thiol and carboxylic acid functional groups, mercaptopropionic acid as a proto-ligand will produce two proper ligands, each with a different anchoring site (Figure 3). Depending on the user input, the functional groups can be used as anchors without modification or be deprotonated beforehand, e.g., turning them into thiolate or carboxylate anions in the mercaptopropionic acid example (see the “optional/ligand/split” option in Figure S1).

Figure 3.

Figure 3

Example of automated functional group discovery in mercaptopropionic acid.

2.4. Ligand Anchors: Biased Ligand Conformer Optimization

While SMILES strings contain all ingredients for constructing a particular molecular configuration, they do not contain information about the optimal three-dimensional conformation. In practice, this problem can generally be alleviated by using one of the numerous available tools for finding/approaching the global energy minimum of a ligand, for example, with the versatile GFN-xTB tight-binding methods.73,74 A challenge particular to the passivation of scaffolds with many anchoring sites is, however, that one does not deal with a single isolated ligand. This may render a minimum energy structure less suitable for use in ligand–scaffold combinations due to unfavorable ligand–ligand and ligand–scaffold interactions. A ligand conformational search must therefore somehow account for the expected interactions with the neighboring ligands and the scaffold. This is illustrated in Figure 4a, where the ligand’s particular global minimum (left) will result in severe interligand steric interactions once placed upon the scaffold’s surface.

Figure 4.

Figure 4

(a) Example conformations of a branched decanoic acid. (b) Example of a biased ligand conformation optimization where each branching group is separately linearized and then put together in the most linear configuration.

There are therefore two main goals when deciding upon a suitable ligand conformation. The first one is that the resulting scaffold–ligand structures are stable even when all ligands have passivated the scaffold, i.e., one should not employ a ligand conformation that will tear the system apart via repulsive interligand steric interactions. The second goal is that the conformational search is fast, also for large ligands. To deal with these challenges, a custom conformational search algorithm has been implemented (Figure 4b), which biases toward the creation of linear conformations that have the advantage of minimizing interligand steric interactions and prevent the collapse of the ligand on the scaffold. This aspect is particularly important for colloidal nanocrystals or systems alike, where ligands densely pack the scaffold surface and usually tend to elongate to minimize interligand steric repulsion. While not necessarily a global minimum, such elongated structures do typically yield low-energy conformations for the ligand–nanocrystal structure.

Our algorithm to produce linear conformers consists of the following six steps:

  • 1.

    The (potentially) branched ligand is fragmented into linear components and capped with hydrogen atoms. Fragments are constructed such that the anchor atom is in the largest fragment.

  • 2.

    The individual fragments are put in anti-periplanar conformations (Figure 4b). An exception is made whenever the anchor atom is part of a dihedral-defining bond, in which case, a syn-periplanar conformation is adopted.

  • 3.

    The fragments are recombined one by one, removing the capping atoms and reforming the previously broken bonds. A set of three candidate conformers is generated by rotating the newly reformed bond, with a subsequent geometry optimization using an empirical universal force field (UFF)75 to ensure that each conformer corresponds to a stationary point.

  • 4.

    The (weighted) perpendicular distance of each atom with respect to a central ligand vector is evaluated (eq 3 below), which is used as a measure for the “linearity” or the complementary “bulkiness” character of a conformation. The conformation that minimizes this weighted distance is optimized without further constraints and passed onto the next step. The definition of the ligand vector will be discussed in more detail in the next section.

  • 5.

    Steps 3 and 4 are repeated until all fragments are reattached.

  • 6.

    In cases where higher quality geometries are desired, the conformational search can be augmented with a final geometry optimization at a higher level of theory by interfacing CAT with an external engine. For example, one could optimize the ligand with a CHARMM-based force field76−79 using parameters automatically generated by MATCH,80 or if even greater accuracy is desired, one could choose a tight binding (e.g., GFNn-xTB)73,74 or density functional calculation.81,82 All these external engines can be used if their executables are already pre-installed and, when required, licensed.

2.5. Ligand/Core Attachment

The general concept behind the scaffold–ligand passivation is similar to its previously described counterpart for organic systems:64 vectors and anchoring sites are defined for the scaffold and ligand, which are then used for rotating and translating the ligand(s) to their final position. A key difference between various scaffold types is, however, the presence or absence of covalent bonds with the organic ligands. Covalent bonds indeed play a key role in defining vectors for the cores. In the case of ionic inorganic nanocrystal scaffolds, where covalent bonds with the ligands are absent, the surface is generally defined by well-defined crystal facets; thus, the vectors of the inorganic core can be defined as those perpendicular to such facets. To identify the surface, CAT utilizes an idealized convex hull constructed from all scaffold anchors.72 The advantage of this approach is a guarantee of the surface’s convexity, the latter ensuring that all inorganic core vectors are either parallel or diverging. The assumption herein is that the surface can be reasonably represented by a convex hull, which might not hold for strongly concave surfaces or more exotically shaped cores (e.g. a torus). In such situations, the resulting regions of varying concavity give rise to converging inorganic core vectors, the latter being detrimental to interligand steric interactions once the surface is passivated with ligands.

On the other hand, for the vast majority of organic scaffolds where the interaction with the ligands is covalent, the procedure to define appropriate scaffold vectors is generally redundant. In that case, the presence of pre-existing bonds allows for easy identification of a scaffold vector, more specifically as the bond between the scaffold anchor and its direct neighbor. This simpler approach was already used by some of us in the construction of functionalized diimides for dye-sensitized photo-electrochemical cells.64

The ligand (unit-)vector Inline graphic is determined by minimizing the (weighted) perpendicular distance (eq 2) of all atoms in the ligand. The origin is herein defined by its functional group-specific anchor atom, ri, representing a vector pointing from the anchor atom to atom i (eq 3).

2.5. 2
2.5. 3

Note that as vlig is both a parameter and a result, a guess is thus required for its initial value, which is then iterated until self-consistency is achieved. This initial trial vector is herein defined as the vector pointing from the ligand anchor atom to the mean position of all ligand atoms. The difference between the optimized and initial trial vectors is shown in Figure 5, which illustrates the already high quality of the trial vector as compared to the optimized vector.

Figure 5.

Figure 5

Difference in ligand rotation before (teal-colored oleate; eq 3) and after (gray-colored oleate) optimizing the ligand vector. vlig is defined by the red horizontal line in both cases.

After all vectors have been defined, the scaffold and the ligand(s) are combined into a single structure by aligning the vectors with each other. This is done in a sequential manner with each ligand being rotated along its vector’s axis such that the distance with respect to all previously defined neighboring ligands is maximized, thus minimizing steric interactions. A warning will be issued if any atoms between ligands are closer than 1.4 Å, indicating that further geometry optimizations are desirable. Ultimately, one obtains one or more ligand-passivated scaffolds. This procedure is identical for all scaffolds, be it a nanocrystal, a metal–organic framework, or an organic dye, with differences only in the definition of the scaffold vector as explained above.

2.6. Postprocessing

After assembling the ligand-passivated scaffolds, several optional workflows are available in CAT, ranging from a simple (constrained) geometry optimization to a workflow for computing ligand dissociation energies with the help of external engines such as AMS. At the end of each calculation, the resulting structures are exported to a user-specified format (e.g. XYZ or PDB), which can then be used for further calculations.

2.7. Examples

To demonstrate the efficiency of our tools for the preparation of realistic composite models and their transferability to a variety of different scaffold/ligand combinations, we demonstrate below the use of CAT for (i) the passivation of a nanocrystal surface by organic surfactants and (ii) the functionalization of a metal–organic framework (MOF) cavity with amino acids.

2.7.1. Example 1

We start by illustrating the construction of a Cs2AgInCl6 double perovskite nanocrystal capped with both (cationic) oleylammonium and (anionic) phenylacetate ligands. This system is meant as a computational model for the Bi-doped Cs2Ag1–xNaxInCl6 nanocrystals synthesized by Zhang et al.(83) and characterized, according to nuclear magnetic resonance analysis, by a mixed oleylammonium/phenylacetate ligand shell (occupying respectively 54% of surface Cs sites/82% of surface Cl sites). In the first step, we built a cubic double perovskite nanocrystal model of about 6.0 nm in side by cutting a cubic Cs2AgInCl6 bulk system along the (100) planes, leaving Cs and Cl on the surface. As-cut, this nanostructure presents a stoichiometry of Cs2197Ag864In864Cl5616 corresponding to an excess of positive charge when each ion is considered in its more stable electronic configuration (i.e., Cs+, Ag+, In3+, and Cl–). Consistent with previous works,84,85 we compensated this excess by removing 37 Cs ions from the surface, leading to a charge balanced Cs2160Ag864In864Cl5616 stoichiometry (Figure 6a). In a second step, we prepared the preprocessed nanocrystal core to carry out the partial surface passivation, i.e., to define the positions where the cationic and anionic ligands should be attached. More specifically, we substituted 54% of the Cs surface atoms by dummy atom A and 82% of the Cl surface atoms by dummy atom B in a uniform manner (Figure 6b) using the dedicated CAT recipe: replace_surface (see the Python script in the Supporting Information, Figure S3). We then prepared the YAML input for CAT (Figure S4) by providing the above-mentioned inorganic core, covered by dummy atoms A and B, in an XYZ format. Here, the structure of the oleylammonium ligands is expressed as a SMILES string and replaces dummy atoms of type A. Moreover, the optional multi_ligand section is added and allows one to include the phenylacetate ligands, also expressed in the SMILES format, that replace dummy atoms of type B. Finally, we performed the multiligand attachment procedure by running the CAT workflow (YAML script executed in about 61 s on a single CPU), resulting in the double perovskite nanocrystal model passivated by a two-component ligand shell (Figure 6c), in line with the composition of the NCs synthesized in the experiments.83 This model was successfully employed as an initial structure to run classical MD simulations.83

Figure 6.

Figure 6

Top panel: Cs2AgInCl6 cubic nanocrystal model (a) enclosed by Cs-Cl-terminated facets, (b) after uniform replacement of 54% of the surface Cs by dummy atom A and 82% of the surface Cl by dummy atom B, and (c) after attachment of oleylammonium and phenylacetate ligands. The inset provides a closer look at the nanocrystal surface. The two ligands are also depicted alone to highlight the color code. Bottom panel: (d) (In) MIL-68-NH2 metal–organic framework model with the anchor site indicated by the arrow, (e) 2D (top) and 3D (bottom) representation of the Fmoc-Ala-OH ligand with the anchor site evidenced by a star, and (f) final (In) MIL-68-NH-Ala-Fmoc compound.

2.7.2. Example 2

In the second example, we aim to attach the protected enantiopure d-alanine, denoted as Fmoc-Ala-OH, to the (In) MIL-68-NH2 metal–organic framework to simulate the complex (In) MIL-68-NH-Ala-Fmoc compound obtained experimentally by Aguado et al.(86)via covalent postmodification (condensation). Following the guidelines provided by the computational study of Todorova et al.,87 we initially generated the metal–organic framework model from the isostructural (In) MIL-6888 by recasting it to its primitive cell (thereby facilitating the further grafting of ligands into a single hexagonal channel) before introducing an amino group to each benzene dicarboxylate linker. Here, we then identified one target NH2 group for the attachment procedure and labeled the H that will serve as the anchor site with dummy atom C (Figure 6d). As in the previous case, we inserted the preprocessed scaffold, its anchoring atom, and the SMILES string representing the Fmoc-Ala-OH ligand into the YAML input file (Figure S5). To account for the concave nature of the hosting pores in metal–organic frameworks, opposed to the convex surface of nanocrystals, we additionally invert the core vectors by setting their alignment to sphere_invert in the optional input section. Finally, to deal with the condensation reaction, we add more specialized information about the ligand anchor site to its input section: this includes the SMILES string of the functional group (carboxylic group here) as well as the indices of the atoms to be anchored (C) and removed (−OH) (Figure 6e).

By providing all these settings to CAT software (YAML script executed in about half a second on a single CPU), we finally obtained a (In) MIL-68-NH-Ala-Fmoc model (Figure 6f). This model also proves to be a suitable initial structure for further ab initio calculations, for example, for geometry optimizations using density functional theory.

3. Conclusions

In this work, we have developed a Python library named Compound Attachment Tool (CAT) for the automatic construction of composite chemical compounds, with an emphasis on inorganic nanocrystals passivated with organic ligands, although the library can be used for any type of system. The workflow consists of four distinct steps: the first step consists in defining the scaffold anchors by a user-provided set of dummy atoms. In the second step, complementary ligand anchor sites are identified by SMILES-based functional group recognition, allowing for parsing and filtering of a large set of (potential) ligands. In the third step, the ligand conformation is optimized to minimize interligand steric interactions when multiple ligands are present. This is accomplished by favoring anti-periplanar conformations and identifying (relatively) compact orientations of side chains. In the fourth and final step, the ligands are combined with the scaffold using the previously defined scaffold/ligand anchoring sites as attachment sites. The ligands are herein placed perpendicular with respect to the scaffold surface, and each ligand is furthermore rotated to maximize the distance with respect to neighboring ligands, thus minimizing interligand steric interactions. The resulting structure is the ligand-passivated (organic, inorganic, or hybrid organic–inorganic) scaffold. In addition to construction of the composite structural models themselves, various postprocessing workflows are available, e.g., ligand dissociation energy calculations.

Software and Data Availability

The CAT 1.0.0 source code is available on GitHub (https://github.com/nlesc-nano/CAT/tree/1.0.0), while binary distributions can be found on PyPi (https://pypi.org/project/nlesc-CAT/1.0.0/). Both are available under the GNU Lesser General Public License (version 3). All data referenced in Section 2.7 was generated using CAT 1.0.0, with the respective input and output files included in the Supporting Information.

Acknowledgments

The authors acknowledge The Netherlands Organization of Scientific Research (NWO) for financial support through the Computational Sciences for Energy Research (CSER) Joint CSER & eScience Research Programme 2017 grant with number 680-91-086. The computational work was carried out on the Dutch national e-infrastructure with the support of the SURFCooperative.

Glossary

Abbreviations

CAT

Compound Attachment Tool

ASE

Atomic Simulation Environment

VASP

Vienna Ab initio Simulation Package

ADF

Amsterdam Density Functional program

NOMAD

Novel Materials Discovery

GDB-13

Generated Database with up to 13 atoms

SMILES

Simplified Molecular-Input Line-Entry System

SELFIES

Self-Referencing Embedded Strings

YAML

YAML Ain’t Markup Language

PLAMS

Python Library for Automating Molecular Simulation

PDB

Protein Databank format

InChI

International Chemical Identifier

GFN-xTB

Geometry Frequency Noncovalent Extended Tight Binding

MOF

metal–organic framework

PyPi

The Python Packaging Index

GNU

GNU’s Not Unix!

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.2c00690.

  • Example YAML input files and illustrations of partially passivated nanocrystals (PDF)

  • Input file for the Cs2AgInCl6 double perovskite nanocrystal model (TXT)

  • YAML input file for the Cs2AgInCl6 double perovskite nanocrystal model (XYZ)

  • Output file for Cs2AgInCl6 double perovskite nanocrystal model (XYZ)

  • Input file for (In) MIL-68-NH2 (TXT)

  • YAML input file for (In) MIL-68-NH2 (XYZ)

  • Output file for (In) MIL-68-NH2 (XYZ)

  • README file (TXT)

Author Contributions

The manuscript was written through contributions of all authors. All authors have given approval to the final version of the manuscript.

The authors declare no competing financial interest.

Supplementary Material

ci2c00690_si_001.pdf (271.4KB, pdf)
ci2c00690_si_002.txt (402B, txt)
ci2c00690_si_003.xyz (538.3KB, xyz)
ci2c00690_si_004.xyz (2.5MB, xyz)
ci2c00690_si_005.txt (337B, txt)
ci2c00690_si_006.xyz (6.3KB, xyz)
ci2c00690_si_007.xyz (8.7KB, xyz)
ci2c00690_si_008.txt (381B, txt)

References

  1. Butler K. T.; Davies D. W.; Cartwright H.; Isayev O.; Walsh A. Machine Learning for Molecular and Materials Science. Nature 2018, 559, 547–555. 10.1038/s41586-018-0337-2. [DOI] [PubMed] [Google Scholar]
  2. Liu Y.; Zhao T.; Ju W.; Shi S. Materials Discovery and Design Using Machine Learning. J. Mater. 2017, 3, 159–177. 10.1016/j.jmat.2017.08.002. [DOI] [Google Scholar]
  3. Gómez-Bombarelli R.; Aguilera-Iparraguirre J.; Hirzel T. D.; Duvenaud D.; Maclaurin D.; Blood-Forsythe M. A.; Chae H. S.; Einzinger M.; Ha D.-G.; Wu T.; Markopoulos G.; Jeon S.; Kang H.; Miyazaki H.; Numata M.; Kim S.; Huang W.; Hong S. I.; Baldo M.; Adams R. P.; Aspuru-Guzik A. Design of Efficient Molecular Organic Light-Emitting Diodes by a High-Throughput Virtual Screening and Experimental Approach. Nat. Mater. 2016, 15, 1120–1127. 10.1038/nmat4717. [DOI] [PubMed] [Google Scholar]
  4. Hachmann J.; Olivares-Amaya R.; Atahan-Evrenk S.; Amador-Bedolla C.; Sánchez-Carrera R. S.; Gold-Parker A.; Vogt L.; Brockway A. M.; Aspuru-Guzik A. The Harvard Clean Energy Project: Large-Scale Computational Screening and Design of Organic Photovoltaics on the World Community Grid. J. Phys. Chem. Lett. 2011, 2, 2241–2251. 10.1021/jz200866s. [DOI] [Google Scholar]
  5. Engel T. Basic Overview of Chemoinformatics. J. Chem. Inf. Model. 2006, 46, 2267–2277. 10.1021/ci600234z. [DOI] [PubMed] [Google Scholar]
  6. Schwab C. H.Conformational Analysis and Searching. In Handbook of Chemoinformatics; Wiley-VCH Verlag Gmb H, 2008; pp 262–301. 10.1002/9783527618279.ch9b. [DOI] [Google Scholar]
  7. Ong S. P.; Richards W. D.; Jain A.; Hautier G.; Kocher M.; Cholia S.; Gunter D.; Chevrier V. L.; Persson K. A.; Ceder G. Python Materials Genomics (Pymatgen): A Robust, Open-Source Python Library for Materials Analysis. Comput. Mater. Sci. 2013, 68, 314–319. 10.1016/j.commatsci.2012.10.028. [DOI] [Google Scholar]
  8. Larsen A. H.; Mortensen J. J.; Blomqvist J.; Castelli I. E.; Christensen R.; Dułak M.; Friis J.; Groves M. N.; Hammer B.; Hargus C.; Hermes E. D.; Jennings P. C.; Jensen P. B.; Kermode J.; Kitchin J. R.; Kolsbjerg E. L.; Kubal J.; Kaasbjerg K.; Lysgaard S.; Maronsson J. B.; Maxson T.; Olsen T.; Pastewka L.; Peterson A.; Rostgaard C.; Schiøtz J.; Schütt O.; Strange M.; Thygesen K. S.; Vegge T.; Vilhelmsen L.; Walter M.; Zeng Z.; Jacobsen K. W.; et al. J. Phys. Condens. Matter 2017, 29, 273002. 10.1088/1361-648X/aa680e. [DOI] [PubMed] [Google Scholar]
  9. Kresse G.; Hafner J. Ab Initio Molecular Dynamics for Liquid Metals. Phys. Rev. B 1993, 47, 558–561. 10.1103/PhysRevB.47.558. [DOI] [PubMed] [Google Scholar]
  10. Kresse G.; Furthmüller J. Efficiency of Ab-Initio Total Energy Calculations for Metals and Semiconductors Using a Plane-Wave Basis Set. Comput. Mater. Sci. 1996, 6, 15–50. 10.1016/0927-0256(96)00008-0. [DOI] [PubMed] [Google Scholar]
  11. Kresse G.; Furthmüller J. Efficient Iterative Schemes for Ab Initio Total-Energy Calculations Using a Plane-Wave Basis Set. Phys. Rev. B 1996, 54, 11169–11186. 10.1103/PhysRevB.54.11169. [DOI] [PubMed] [Google Scholar]
  12. Giannozzi P.; Baroni S.; Bonini N.; Calandra M.; Car R.; Cavazzoni C.; Ceresoli D.; Chiarotti G. L.; Cococcioni M.; Dabo I.; Corso A. D.; de Gironcoli S.; Fabris S.; Fratesi G.; Gebauer R.; Gerstmann U.; Gougoussis C.; Kokalj A.; Lazzeri M.; Martin-Samos L.; Marzari N.; Mauri F.; Mazzarello R.; Paolini S.; Pasquarello A.; Paulatto L.; Sbraccia C.; Scandolo S.; Sclauzero G.; Seitsonen A. P.; Smogunov A.; Umari P. Wentzcovitch, R. M. QUANTUM ESPRESSO: A Modular and Open-Source Software Project for Quantum Simulations of Materials. J. Phys. Condens. Matter 2009, 21, 395502. 10.1088/0953-8984/21/39/395502. [DOI] [PubMed] [Google Scholar]
  13. Giannozzi P.; Andreussi O.; Brumme T.; Bunau O.; Nardelli M. B.; Calandra M.; Car R.; Cavazzoni C.; Ceresoli D.; Cococcioni M.; Colonna N.; Carnimeo I.; Corso A. D.; de Gironcoli S.; Delugas P.; DiStasio R. A.; Ferretti A.; Floris A.; Fratesi G.; Fugallo G.; Gebauer R.; Gerstmann U.; Giustino F.; Gorni T.; Jia J.; Kawamura M.; Ko H.-Y.; Kokalj A.; Küçükbenli E.; Lazzeri M.; Marsili M.; Marzari N.; Mauri F.; Nguyen N. L.; Nguyen H.-V.; Otero-de-la-Roza A.; Paulatto L.; Poncé S.; Rocca D.; Sabatini R.; Santra B.; Schlipf M.; Seitsonen A. P.; Smogunov A.; Timrov I.; Thonhauser T.; Umari P.; Vast N.; Wu X.; Baroni S. Advanced Capabilities for Materials Modelling with Quantum ESPRESSO. J. Phys. Condens. Matter 2017, 29, 465901. 10.1088/1361-648X/aa8f79. [DOI] [PubMed] [Google Scholar]
  14. te Velde G.; Bickelhaupt F. M.; Baerends E. J.; Fonseca Guerra C.; van Gisbergen S. J. A.; Snijders J. G.; Ziegler T. Chemistry with ADF. J. Comput. Chem. 2001, 22, 931–967. 10.1002/jcc.1056. [DOI] [Google Scholar]
  15. Frisch M. J.; Trucks G. W.; Schlegel H. B.; Scuseria G. E.; Robb M. A.; Cheeseman J. R.; Scalmani G.; Barone V.; Petersson G. A.; Nakatsuji H.; Li X.; Caricato M.; Marenich A. V.; Bloino J.; Janesko B. G.; Gomperts R.; Mennucci B.; Hratch D. J.. Gaussian 16, Revision C.01; Gaussian Inc.: Wallingford CT: 2016. [Google Scholar]
  16. Jain A.; Ong S. P.; Hautier G.; Chen W.; Richards W. D.; Dacek S.; Cholia S.; Gunter D.; Skinner D.; Ceder G.; Persson K. A. Commentary: The Materials Project: A Materials Genome Approach to Accelerating Materials Innovation. APL Mater. 2013, 1, 11002. 10.1063/1.4812323. [DOI] [Google Scholar]
  17. Draxl C.; Scheffler M. NOMAD: The FAIR Concept for Big Data-Driven Materials Science. MRS Bull. 2018, 43, 676–682. 10.1557/mrs.2018.208. [DOI] [Google Scholar]
  18. Huber S. P.; Zoupanos S.; Uhrin M.; Talirz L.; Kahle L.; Häuselmann R.; Gresch D.; Müller T.; Yakutovich A. V.; Andersen C. W.; Ramirez F. F.; Adorf C. S.; Gargiulo F.; Kumbhar S.; Passaro E.; Johnston C.; Merkys A.; Cepellotti A.; Mounet N.; Marzari N.; Kozinsky B.; Pizzi G. AiiDA 1.0, a Scalable Computational Infrastructure for Automated Reproducible Workflows and Data Provenance. Sci. Data 2020, 7, 1–18. 10.1038/s41597-020-00638-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Álvarez-Moreno M.; De Graaf C.; López N.; Maseras F.; Poblet J. M.; Bo C. Managing the Computational Chemistry Big Data Problem: The IoChem-BD Platform. J. Chem. Inf. Model. 2015, 55, 95–103. 10.1021/ci500593j. [DOI] [PubMed] [Google Scholar]
  20. Olivares-Amaya R.; Amador-Bedolla C.; Hachmann J.; Atahan-Evrenk S.; Sánchez-Carrera R. S.; Vogt L.; Aspuru-Guzik A. Accelerated Computational Discovery of High-Performance Materials for Organic Photovoltaics by Means of Cheminformatics. Energy Environ. Sci. 2011, 4, 4849. 10.1039/c1ee02056k. [DOI] [Google Scholar]
  21. Sousa S. F.; Fernandes P. A.; Ramos M. J. Protein-Ligand Docking: Current Status and Future Challenges. Proteins Struct. Funct. Bioinf. 2006, 65, 15–26. 10.1002/prot.21082. [DOI] [PubMed] [Google Scholar]
  22. Huang S.-Y.; Zou X. Advances and Challenges in Protein-Ligand Docking. Int. J. Mol. Sci. 2010, 11, 3016–3034. 10.3390/ijms11083016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Motz R. N.; Lopato E. M.; Connell T. U.; Bernhard S. High-Throughput Screening of Earth-Abundant Water Reduction Catalysts toward Photocatalytic Hydrogen Evolution. Inorg. Chem. 2021, 60, 774–781. 10.1021/acs.inorgchem.0c02790. [DOI] [PubMed] [Google Scholar]
  24. Schaftenaar G.; Noordik J. H. Molden: A Pre- and Post-Processing Program for Molecular and Electronic Structures. J. Comput.-Aided Mol. Des. 2000, 14, 123–134. 10.1023/A:1008193805436. [DOI] [PubMed] [Google Scholar]
  25. Hanwell M. D.; Curtis D. E.; Lonie D. C.; Vandermeersch T.; Zurek E.; Hutchison G. R. Avogadro: An Advanced Semantic Chemical Editor, Visualization, and Analysis Platform. J. Chem. 2012, 4, 476. 10.1016/j.aim.2014.05.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Kim S.; Chen J.; Cheng T.; Gindulyte A.; He J.; He S.; Li Q.; Shoemaker B. A.; Thiessen P. A.; Yu B.; Zaslavsky L.; Zhang J.; Bolton E. E. PubChem in 2021: New Data Content and Improved Web Interfaces. Nucleic Acids Res. 2021, 49, D1388–D1395. 10.1093/nar/gkaa971. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Blum L. C.; Reymond J.-L. 970 Million Druglike Small Molecules for Virtual Screening in the Chemical Universe Database GDB-13. J. Am. Chem. Soc. 2009, 131, 8732–8733. 10.1021/ja902302h. [DOI] [PubMed] [Google Scholar]
  28. Rupp M.; Tkatchenko A.; Müller K. R.; Von Lilienfeld O. A. Fast and Accurate Modeling of Molecular Atomization Energies with Machine Learning. Phys. Rev. Lett. 2012, 108, 1–5. 10.1103/PhysRevLett.108.058301. [DOI] [PubMed] [Google Scholar]
  29. Weininger D. SMILES, a Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules. J. Chem. Inf. Model. 1988, 28, 31–36. 10.1021/ci00057a005. [DOI] [Google Scholar]
  30. Weininger D.; Weininger A.; Weininger J. L. SMILES. 2. Algorithm for Generation of Unique SMILES Notation. J. Chem. Inf. Comput. Sci. 1989, 29, 97–101. 10.1021/ci00062a008. [DOI] [Google Scholar]
  31. Krenn M.; Häse F.; Nigam A.; Friederich P.; Aspuru-Guzik A. Self-Referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation. Mach. Learn. Sci. Technol. 2020, 1, 45024. 10.1088/2632-2153/aba947. [DOI] [Google Scholar]
  32. Landrum G.; Tosco P.; Kelley B.; Ric; sriniker; gedeck; Vianello, R.; Schneider, N.; Kawashima, E.; Dalke, A.; N, D.; Cosgrove, D.; Cole, B.; Swain, M.; Turk, S.; Savelyev, A.; Jones, G.; Vaucher, A.; Wójcikowski, M.; Take, I.; Probst, D.; Ujihara, K.; Scalfani, V. F.; Godin, G.; Pahl, A.; Berenger, F.; JLVarjo; strets123; JP; Gavid, D. RDKit. 2022.
  33. O’Boyle N. M.; Banck M.; James C. A.; Morley C.; Vandermeersch T.; Hutchison G. R. Open Babel: An Open Chemical Toolbox. Aust. J. Chem. 2011, 3, 33. 10.1186/1758-2946-3-33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Creswell A.; White T.; Dumoulin V.; Arulkumaran K.; Sengupta B.; Bharath A. A. Generative Adversarial Networks: An Overview. IEEE Signal Process. Mag. 2018, 35, 53–65. 10.1109/MSP.2017.2765202. [DOI] [Google Scholar]
  35. Lavecchia A. Machine-Learning Approaches in Drug Discovery: Methods and Applications. Drug Discovery Today 2015, 20, 318–331. 10.1016/j.drudis.2014.10.012. [DOI] [PubMed] [Google Scholar]
  36. Stephenson N.; Shane E.; Chase J.; Rowland J.; Ries D.; Justice N.; Zhang J.; Chan L.; Cao R. Survey of Machine Learning Techniques in Drug Discovery. Curr. Drug Metab. 2019, 20, 185–193. 10.2174/1389200219666180820112457. [DOI] [PubMed] [Google Scholar]
  37. Vamathevan J.; Clark D.; Czodrowski P.; Dunham I.; Ferran E.; Lee G.; Li B.; Madabhushi A.; Shah P.; Spitzer M.; Zhao S. Applications of Machine Learning in Drug Discovery and Development. Nat. Rev. Drug Discov. 2019, 18, 463–477. 10.1038/s41573-019-0024-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Furukawa H.; Cordova K. E.; O’Keeffe M.; Yaghi O. M. The Chemistry and Applications of Metal-Organic Frameworks. Science 2013, 341, 1230444. 10.1126/science.1230444. [DOI] [PubMed] [Google Scholar]
  39. Slater A. G.; Cooper A. I. Function-Led Design of New Porous Materials. Science 2015, 348, aaa8075. 10.1126/science.aaa8075. [DOI] [PubMed] [Google Scholar]
  40. Chung Y. G.; Camp J.; Haranczyk M.; Sikora B. J.; Bury W.; Krungleviciute V.; Yildirim T.; Farha O. K.; Sholl D. S.; Snurr R. Q. Computation-Ready, Experimental Metal-Organic Frameworks: A Tool to Enable High-Throughput Screening of Nanoporous Crystals. Chem. Mater. 2014, 26, 6185–6192. 10.1021/cm502594j. [DOI] [Google Scholar]
  41. Chung Y. G.; Haldoupis E.; Bucior B. J.; Haranczyk M.; Lee S.; Zhang H.; Vogiatzis K. D.; Milisavljevic M.; Ling S.; Camp J. S.; Slater B.; Siepmann J. I.; Sholl D. S.; Snurr R. Q. Advances, Updates, and Analytics for the Computation-Ready, Experimental Metal-Organic Framework Database: CoRE MOF 2019. J. Chem. Eng. Data 2019, 64, 5985–5998. 10.1021/acs.jced.9b00835. [DOI] [Google Scholar]
  42. Groom C. R.; Bruno I. J.; Lightfoot M. P.; Ward S. C. The Cambridge Structural Database. Acta Crystallogr. Sect. B Struct. Sci. Cryst. Eng. Mater. 2016, 72, 171–179. 10.1107/S2052520616003954. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Saparov B.; Mitzi D. B. Organic-Inorganic Perovskites: Structural Versatility for Functional Materials Design. Chem. Rev. 2016, 116, 4558–4596. 10.1021/acs.chemrev.5b00715. [DOI] [PubMed] [Google Scholar]
  44. Ulman A. Formation and Structure of Self-Assembled Monolayers. Chem. Rev. 1996, 4, 1533–1554. 10.1021/cr9502357. [DOI] [PubMed] [Google Scholar]
  45. Englebienne P.; Hoonacker A.; Verhas M.; Khlebtsov N. Advances in High-Throughput Screening: Biomolecular Interaction Monitoring in Real-Time with Colloidal Metal Nanoparticles. Comb. Chem. High Throughput Screening 2003, 6, 777–787. 10.2174/138620703771826955. [DOI] [PubMed] [Google Scholar]
  46. Emory S. R.; Nie S. Screening and Enrichment of Metal Nanoparticles with Novel Optical Properties. J. Phys. Chem. B 1998, 102, 493–497. 10.1021/jp9734033. [DOI] [Google Scholar]
  47. Narayanan R.; El-Sayed M. A. Some Aspects of Colloidal Nanoparticle Stability, Catalytic Activity, and Recycling Potential. Top. Catal. 2008, 47, 15–21. 10.1007/s11244-007-9029-0. [DOI] [Google Scholar]
  48. Reichenberger S.; Marzun G.; Muhler M.; Barcikowski S. Perspective of Surfactant-Free Colloidal Nanoparticles in Heterogeneous Catalysis. ChemCatChem 2019, 11, 4489–4518. 10.1002/cctc.201900666. [DOI] [Google Scholar]
  49. Xiao G.; Wang Y.; Ning J.; Wei Y.; Liu B.; Yu W. W.; Zou G.; Zou B. Recent Advances in IV–VI Semiconductor Nanocrystals: Synthesis, Mechanism, and Applications. RSC Adv. 2013, 3, 8104. 10.1039/c3ra23209c. [DOI] [Google Scholar]
  50. Wells R. L.; Gladfelter W. L. Pathways to Nanocrystalline III-V (13-15) Compound Semiconductors. J. Cluster Sci. 1997, 8, 217–238. 10.1023/A:1022684024708. [DOI] [Google Scholar]
  51. Fu H.; Tsang S. W. Infrared Colloidal Lead Chalcogenide Nanocrystals: Synthesis, Properties, and Photovoltaic Applications. Nanoscale 2012, 4, 2187–2201. 10.1039/c2nr11836j. [DOI] [PubMed] [Google Scholar]
  52. Stranks S. D.; Snaith H. J. Metal-Halide Perovskites for Photovoltaic and Light-Emitting Devices. Nat. Nanotechnol. 2015, 10, 391–402. 10.1038/nnano.2015.90. [DOI] [PubMed] [Google Scholar]
  53. Sardar R.; Funston A. M.; Mulvaney P.; Murray R. W. Gold Nanoparticles: Past, Present, and Future. Langmuir 2009, 25, 13840–13851. 10.1021/la9019475. [DOI] [PubMed] [Google Scholar]
  54. Zhang X. F.; Liu Z. G.; Shen W.; Gurunathan S. Silver Nanoparticles: Synthesis, Characterization, Properties, Applications, and Therapeutic Approaches. Int. J. Mol. Sci. 2016, 17, 1534. 10.3390/ijms17091534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Peng Z.; Yang H. Designer Platinum Nanoparticles : Control of Shape , Composition in Alloy, Nanostructure and Electrocatalytic Property. Nano Today 2009, 4, 143–164. 10.1016/j.nantod.2008.10.010. [DOI] [Google Scholar]
  56. Zhang Y.; Clapp A. Overview of Stabilizing Ligands for Biocompatible Quantum Dot Nanocrystals. Sensors 2011, 11, 11036–11055. 10.3390/s111211036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Boles M. A.; Ling D.; Hyeon T.; Talapin D. V. The Surface Science of Nanocrystals. Nat. Mater. 2016, 15, 141–153. 10.1038/nmat4526. [DOI] [PubMed] [Google Scholar]
  58. Fischer S. A.; Crotty A. M.; Kilina S. V.; Ivanov S. A.; Tretiak S. Passivating Ligand and Solvent Contributions to the Electronic Properties of Semiconductor Nanocrystals. Nanoscale 2012, 4, 904–914. 10.1039/C2NR11398H. [DOI] [PubMed] [Google Scholar]
  59. Weidman M. C.; Nguyen Q.; Smilgies D.-M.; Tisdale W. A. Impact of Size Dispersity, Ligand Coverage, and Ligand Length on the Structure of PbS Nanocrystal Superlattices. Chem. Mater. 2018, 30, 807–816. 10.1021/acs.chemmater.7b04322. [DOI] [Google Scholar]
  60. Foscato M.; Venkatraman V.; Occhipinti G.; Alsberg B. K.; Jensen V. R. Automated Building of Organometallic Complexes from 3D Fragments. J. Chem. Inf. Model. 2014, 54, 1919–1931. 10.1021/ci5003153. [DOI] [PubMed] [Google Scholar]
  61. Ioannidis E. I.; Gani T. Z. H.; Kulik H. J. MolSimplify: A Toolkit for Automating Discovery in Inorganic Chemistry. J. Comput. Chem. 2016, 2106–2117. 10.1002/jcc.24437. [DOI] [PubMed] [Google Scholar]
  62. Nandy A.; Duan C.; Janet J. P.; Gugler S.; Kulik H. J. Strategies and Software for Machine Learning Accelerated Discovery in Transition Metal Chemistry. Ind. Eng. Chem. Res. 2018, 57, 13973–13986. 10.1021/acs.iecr.8b04015. [DOI] [Google Scholar]
  63. Sobez J. G.; Reiher M. Molassembler: Molecular Graph Construction, Modification, and Conformer Generation for Inorganic and Organic Molecules. J. Chem. Inf. Model. 2020, 60, 3884–3900. 10.1021/acs.jcim.0c00503. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Belić J.; van Beek B.; Menzel J. P.; Buda F.; Visscher L. Systematic Computational Design and Optimization of Light Absorbing Dyes. J. Phys. Chem. A 2020, 124, 6380–6388. 10.1021/acs.jpca.0c04506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Ben-Kiki O.; Evans C.; Döt Net I.. YAML Ain’t Markup Language. 2005.
  66. Handzlik M.Python Library for Automating Molecular Simulation; SCM: Amsterdam: 2019. [Google Scholar]
  67. Bernstein F. C.; Koetzle T. F.; Williams G. J. B.; Meyer E. F. Jr.; Brice M. D.; Rodgers J. R.; Kennard O.; Shimanouchi T.; Tasumi M. The Protein Data Bank: A Computer-Based Archival File for Macromolecular Structures. J. Mol. Biol. 1977, 112, 535–542. 10.1016/S0022-2836(77)80200-3. [DOI] [PubMed] [Google Scholar]
  68. Ash S.; Cline M. A.; Homer R. W.; Hurst T.; Smith G. B. SYBYL Line Notation (SLN): A Versatile Language for Chemical Structure Representation †. J. Chem. Inf. Comput. Sci. 1997, 37, 71–79. 10.1021/ci960109j. [DOI] [Google Scholar]
  69. Heller S.; McNaught A.; Stein S.; Tchekhovskoi D.; Pletnev I. InChI - the Worldwide Chemical Structure Identifier Standard. Aust. J. Chem. 2013, 5, 7. 10.1186/1758-2946-5-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Chatzigoulas A.; Karathanou K.; Dellis D.; Cournia Z. NanoCrystal: A Web-Based Crystallographic Tool for the Construction of Nanoparticles Based on Their Crystal Habit. J. Chem. Inf. Model. 2018, 58, 2380–2386. 10.1021/acs.jcim.8b00269. [DOI] [PubMed] [Google Scholar]
  71. Wulff G. Zur Frage Der Geschwindigkeit Des Wachsthums Und Der Auflösung Der Krystallflächen. Z. Kristallogr. 1901, 34, 449–530. 10.1524/zkri.1901.34.1.449. [DOI] [Google Scholar]
  72. Barber C. B.; Dobkin D. P.; Huhdanpaa H. The Quickhull Algorithm for Convex Hulls. ACM Trans. Math. Softw. 1996, 22, 469–483. 10.1145/235815.235821. [DOI] [Google Scholar]
  73. Grimme S.; Bannwarth C.; Shushkov P. A Robust and Accurate Tight-Binding Quantum Chemical Method for Structures, Vibrational Frequencies, and Noncovalent Interactions of Large Molecular Systems Parametrized for All Spd-Block Elements ( Z = 1–86). J. Chem. Theory Comput. 2017, 13, 1989–2009. 10.1021/acs.jctc.7b00118. [DOI] [PubMed] [Google Scholar]
  74. Bannwarth C.; Ehlert S.; Grimme S. GFN2-XTB—An Accurate and Broadly Parametrized Self-Consistent Tight-Binding Quantum Chemical Method with Multipole Electrostatics and Density-Dependent Dispersion Contributions. J. Chem. Theory Comput. 2019, 15, 1652–1671. 10.1021/acs.jctc.8b01176. [DOI] [PubMed] [Google Scholar]
  75. Rappé A. K.; Casewit C. J.; Colwell K. S.; Goddard W. A. III; Skiff W. M. UFF, a Full Periodic Table Force Field for Molecular Mechanics and Molecular Dynamics Simulations. J. Am. Chem. Soc. 1992, 114, 10024–10035. 10.1021/ja00051a040. [DOI] [Google Scholar]
  76. Brooks B. R.; Bruccoleri R. E.; Olafson B. D.; States D. J.; Swaminathan S.; Karplus M. CHARMM: A Program for Macromolecular Energy, Minimization, and Dynamics Calculations. J. Comput. Chem. 1983, 4, 187–217. 10.1002/jcc.540040211. [DOI] [Google Scholar]
  77. Brooks B. R.; Brooks C. L.; Mackerell A. D.; Nilsson L.; Petrella R. J.; Roux B.; Won Y.; Archontis G.; Bartels C.; Boresch S.; Caflisch A.; Caves L.; Cui Q.; Dinner A. R.; Feig M.; Fischer S.; Gao J.; Hodoscek M.; Im W.; Kuczera K.; Lazaridis T.; Ma J.; Ovchinnikov V.; Paci E.; Pastor R. W.; Post C. B.; Pu J. Z.; Schaefer M.; Tidor B.; Venable R. M.; Woodcock H. L.; Wu X.; Yang W.; York D. M.; Karplus M. CHARMM: The Biomolecular Simulation Program. J. Comput. Chem. 2009, 30, 1545–1614. 10.1002/jcc.21287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Vanommeslaeghe K.; MacKerell A. D. Jr. Automation of the CHARMM General Force Field (CGenFF) I: Bond Perception and Atom Typing. J. Chem. Inf. Model. 2012, 52, 3144–3154. 10.1021/ci300363c. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Vanommeslaeghe K.; Raman E. P.; MacKerell A. D. Jr. Automation of the CHARMM General Force Field (CGenFF) II: Assignment of Bonded Parameters and Partial Atomic Charges. J. Chem. Inf. Model. 2012, 52, 3155–3168. 10.1021/ci3003649. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Yesselman J. D.; Price D. J.; Knight J. L.; Brooks C. L. III MATCH: An Atom-Typing Toolset for Molecular Mechanics Force Fields. J. Comput. Chem. 2012, 33, 189–202. 10.1002/jcc.21963. [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Burke K. Perspective on Density Functional Theory. J. Chem. Phys. 2012, 136, 150901. 10.1063/1.4704546. [DOI] [PubMed] [Google Scholar]
  82. Kohn W.; Sham L. J. Self-Consistent Equations Including Exchange and Correlation Effects. Phys. Rev. 1965, 140, A1133–A1138. 10.1103/PhysRev.140.A1133. [DOI] [Google Scholar]
  83. Zhang B.; Wang M.; Ghini M.; Melcherts A. E. M.; Zito J.; Goldoni L.; Infante I.; Guizzardi M.; Scotognella F.; Kriegel I.; De Trizio L.; Manna L. Colloidal Bi-Doped Cs2Ag1– x NaxInCl6 Nanocrystals: Undercoordinated Surface Cl Ions Limit Their Light Emission Efficiency. ACS Mater. Lett. 2020, 1442–1449. 10.1021/acsmaterialslett.0c00359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  84. Giansante C.; Infante I. Surface Traps in Colloidal Quantum Dots: A Combined Experimental and Theoretical Perspective. J. Phys. Chem. Lett. 2017, 8, 5209–5215. 10.1021/acs.jpclett.7b02193. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Bodnarchuk M. I.; Boehme S. C.; Ten Brinck S.; Bernasconi C.; Shynkarenko Y.; Krieg F.; Widmer R.; Aeschlimann B.; Günther D.; Kovalenko M. V.; Infante I. Rationalizing and Controlling the Surface Structure and Electronic Passivation of Cesium Lead Halide Nanocrystals. ACS Energy Lett. 2019, 4, 63–74. 10.1021/acsenergylett.8b01669. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. Canivet J.; Aguado S.; Bergeret G.; Farrusseng D. Amino Acid Functionalized Metal-Organic Frameworks by a Soft Coupling-Deprotection Sequence. Chem. Commun. 2011, 47, 11650–11652. 10.1039/c1cc15541e. [DOI] [PubMed] [Google Scholar]
  87. Todorova T. K.; Rozanska X.; Gervais C.; Legrand A.; Ho L. N.; Berruyer P.; Lesage A.; Emsley L.; Farrusseng D.; Canivet J.; Mellot-Draznieks C. Molecular Level Characterization of the Structure and Interactions in Peptide-Functionalized Metal–Organic Frameworks. Chem. – Eur. J. 2016, 22, 16531–16538. 10.1002/chem.201603255. [DOI] [PubMed] [Google Scholar]
  88. Volkringer C.; Meddouri M.; Loiseau T.; Guillou N.; Marrot J.; Férey G.; Haouas M.; Taulelle F.; Audebrand N.; Latroche M. The Kagomé Topology of the Gallium and Indium Metal-Organic Framework Types with a MIL-68 Structure: Synthesis, XRD, Solid-State NMR Characterizations, and Hydrogen Adsorption. Inorg. Chem. 2008, 47, 11892–11901. 10.1021/ic801624v. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

ci2c00690_si_001.pdf (271.4KB, pdf)
ci2c00690_si_002.txt (402B, txt)
ci2c00690_si_003.xyz (538.3KB, xyz)
ci2c00690_si_004.xyz (2.5MB, xyz)
ci2c00690_si_005.txt (337B, txt)
ci2c00690_si_006.xyz (6.3KB, xyz)
ci2c00690_si_007.xyz (8.7KB, xyz)
ci2c00690_si_008.txt (381B, txt)

Articles from Journal of Chemical Information and Modeling are provided here courtesy of American Chemical Society

RESOURCES