Abstract
Coarse-grained (CG) molecular dynamics is a powerful tool for simulating the collective behavior of biomolecules. However, the structural information lost during coarse-graining prevents the CG configurations from being more widely useful (e.g., for ligand binding). Regenerating the lost all-atom coordinates, or backmapping, is an unmet challenge for protein CG at resolutions lower than one coarse-grain site or bead per amino acid residue. This low resolution is computationally necessary to simulate many protein complexes including viruses like SARS-CoV-2 and HIV-1. We propose MSBack, a method to backmap highly CG proteins using a diffusion model for the all-atom coordinates constrained to fit the CG coordinates. This diffusion process works by perturbing a known all-atom structure and does not require retraining. We show that this stochastically generates a distribution of α-carbon traces that match the CG coordinates. By combining this with physics-based methods for smaller-length backmapping, we fully backmap a mature HIV-1 capsid bound with the small molecule inositol hexakisphosphate at 1 Å resolution.


Introduction
All-atom molecular dynamics (AAMD) is a powerful tool for simulation the collective behavior of proteins and other biomolecules. − However, it requires computation of forces on every atom at every time step in the simulation and thus is currently limited to microsecond to millisecond time scales and nanometer length scales. Coarse-graining helps to solve this problem by grouping atoms together into “beads”, reducing the computational cost, including “bottom-up” coarse-graining that is derived from the atomistic scale interactions. , Coarse-grained (CG) molecular dynamics (MD) has found many applications in biomolecular systems and specifically in proteins, whose structures are critical to their functions (examples can be found in refs and for the HIV-1 virus), although other biomolecules such as lipids and their interactions with proteins are also critically important. − Optimal mapping of the CG sites of “beads” is also an active area of research with new methods creating more complex mapping functions. − Models can vary in resolution, that is how many amino acids are mapped to a single bead and, in some cases, parts of the protein structure are unmapped but their effects are accounted for by hidden state variables in the Ultra-Coarse-Grained (UCG) models. − Thus, there is currently a diverse set of methodologies for creating CG models proteins, leading to an even more diverse set of CG models.
One drawback of coarse-graining, however, can be the loss of information during the coarse-graining process. This means that the final structure after the simulation does not include the details found in an AAMD simulation. But, recovering this information can allow for iterative simulations between different length scales and verification of certain results from CG force fields and is thus a powerful tool in multiscale modeling. , The process of recovering this information is known as backmapping and it amounts to sampling from a distribution of all-atom protein structures which fit the CG mapping function. In principle, this problem can be solved by biasing AAMD such that it matches aspects of the CG mapping; however, such an approach requires running additional AAMD simulations which may be so slow as to defeat the purpose.
Recently, it has been shown that diffusion models based on neural network potentials can be used to approximate distributions of all-atom protein structures. − Many state-of-the-art backmapping methodologies are based on this principle or other machine learning techniques. However, low-resolution CG models and models where certain amino acids are unmapped make backmapping a highly challenging task as current methodologies are primarily aimed at models which are not very coarse. Examples include one bead to one amino acid mappings of CG proteins ,− or one bead to one monomer mappings of polymers. , Some methods even suggest creating the CG model in such a way that it is easily backmapped, , which unfortunately constrains the nature of the CG model. By contrast, we propose MSBack, a method to backmap highly CG proteins using a diffusion model constrained to produce all-atom backbone structures which match the original CG mapping function.
MSBack differs from others because it utilizes AA structures which are available from the original training of the CG model. Our diffusion process starts from an AA structure instead of Gaussian noise and utilizes a pretrained diffusion model with an adjusted noise schedule that perturbs but does not destroy the initial structure. We note that this does not require any retraining of the diffusion model. The α carbon traces we create are both accurate (matches the CG coordinates) and diverse (stochastically generates many different, yet reasonable α carbon traces). Once that is achieved, software from previous work is sufficient to finish the backmapping process, including small molecules and lipids. , We will focus here as examples the backmapping of two different and very relevant CG proteins, the HIV-1 capsid (CA) protein and the SARS-CoV-2 membrane (M) protein. In the case of the HIV capsid protein, we will focus on backmapping a widely used CG model that does not map all the AA residues to the CG model, meaning that the other methods listed in the prior paragraphs would not be applicable. In the case of the M protein, we will focus on CG models using a KMC-CG mapping with an average resolution of 10 residues per CG bead, i.e., such a low resolution for which previous methods would not at all perform well. We will show that the present method can backmap this protein very well to within 1 Å and that small variations in ability to backmap can be used to compare different CG models. Finally, we provide an example of “heroic backmapping” by starting from the UCG HIV-1 capsid model to backmap a nearly complete HIV-1 capsid involving 1200 proteins, including the small bound IP6 anion from CG simulations.
Overview of the Approach
We define the linear CG mapping function, M( r ), as a set of linear mapping functions on the fine-grain or all-atom coordinates, r , following previous work, ,,
| 1 |
| 2 |
The CG model is composed of N CG beads indexed with the capital letter I, while the all-atom model is composed of n atoms indexed by the lower-case letter i. The matrix c Ii contains the information relating the two configurations and it is assumed that there is only one nonzero value in each column, i.e., no atom is mapped to more the one CG bead and each row I normalizes to unity. This creates a many-to-one mapping where many different all-atom states map to the same CG state. The backmapping function, M –1( R N ), is a one-to-many map, which samples from the probability distribution of the n AA coordinates, r , given the N CG coordinates, R , such that
| 3 |
where β = 1/k B T. This equation has 2 key features. The first is the product of delta functions which enforces that the backmapped coordinates must perfectly match the CG coordinates. The second is the exponential term which enforces that these coordinates should follow the Boltzmann distribution of the energies of the different states.
We introduce here a backward diffusion process performed on the all-atom coordinates to sample P( r | R ). We define the backward diffusion process in the standard way following previous work, ,, (see Supporting Information for more details)
| 4 |
Here t̅ signifies that this is the backward diffusion process, β t̅ is the variance-preserving noise schedule (and note that it is not related to the inverse temperature in eq ), and Bt̅ is the Brownian noise (also called the Wiener Process) at time t̅. The energy of the configuration is log P t̅ ( r ). However, we want the gradient of the log of the probability given the CG coordinates, ∇ r log P t̅ ( r|R ). Using Bayes” theorem and following previous work, , this reduces to a convenient form with two terms, such that
| 5 |
The first term should be the same as the gradient of U( r ) in eq , while the second term functions like the product of delta functions in eq to enforce the constraint of the CG coordinates.
By using a previously trained denoising-diffusion score model, S θ( r , t̅) the Chroma diffusion model, we already have an approximate solution for the first term and does not require any further training.
| 6 |
The score model was trained on a log linear noise schedule from a signal-to-noise ratio of (10–7, 1013.5). It was obtained by applying an affine transformation to a denoiser, which is trained to predict the know structure of a given protein based on the current noisy structure during the diffusion trajectory (see original publications for more detail. ,
The only remaining term is the probability of the given CG coordinates, R , for the current all-atom coordinates, r . In principle, this probability P t̅ ( R | r ), is the mapping product from eq , , which is either 0 or 1 if the mapped all-atom coordinates match the given CG coordinates. In practice, we utilize a soft and a hard constraint. For the soft constraint, we calculate an energy (log P t̅ ( R | r )), which is automatically differentiated within the neural network. As shown in eq below, we compute the root-mean-square deviation between the mapped all-atom and CG coordinates (RMSD) and apply a softplus activation function shifted by 2 Å. This is so that the structure is not forced to match the CG coordinates during this step, even though the constraint is applied for all times t. We then add a weighting factor of 10 times the signal-to-noise ratio, SNR, clamped at 5 (see Supporting Information for more details).
| 7 |
For the hard constraint, we define a linear transformation in eq ,
| 8 |
which forces the AA coordinates to match the CG coordinates. It is applied during the forward pass of the neural network. Thus, the hard and soft constraints are both applied within the neural network that calculates S θ( r , t̅) following the standard protocol in Chroma.
Our strategy for applying these constraints and the overall diffusion process is shown in Figure and described as follows: We start by aligning the CG coordinates to an all-atom structure via rigid body translation and diffusion. The all-atom coordinates can come from the training data used to build the CG model and thus one can even choose the all-atom structures which initially minimize the difference in these structures. We then run the diffusion process with a signal-to-noise ratio increasing in a log–linear fashion for 1000 steps from 1 to 103 with the soft constraint applied. We then run for another 1000 steps with only the hard constraint applied and the signal-to-noise ratio varied from 103 to 104 in a log–linear fashion. At this point, the hard constraint has enforced that the RMSD between the target CG structure and mapped all-atom structure to be 0, while running with the soft constraint has generated structural diversity. This mirrors the two key features of eq . Running only at higher signal-to-noise ratios preserves much of the information found in the initial all-atom structure. Although the Chroma model can be used to generate the all backbone and side chain atoms, other methods for backmapping from a single α carbon, ,, create more reasonable structures and thus we use only the α carbon trace. Finally, we perform a very quick energy minimization of the whole structure in vacuum with no restraints using Gromacs. This is done only to test the structural soundness of the protein (there is no time for larger scale rearrangement) without the influence of other, stochastically placed molecules and quantify it using the RMSD before and after minimization. Protein structure is highly dependent on the surrounding environment, especially in the case of IDPs which expose a high fraction of residues to water and membrane proteins which may be dependent on the membrane to fold. Thus, in real applications, this step should be replaced by energy minimization in water or a lipid membrane as is normally done in preparing an AAMD simulation before further temperature and pressure equilibration. Generation of the lipid membrane from a CG simulation will be discussed later in the text.
1.
Method of backmapping. Coordinates for a CG model of the HIV-1 CA protein are shown in Cyan. The CG model is an α carbon level model with about 40% of the protein chain unmapped. They are overlaid with the all-atom structure shown with beads at the α carbon positions and lines for the side chains if they exist for each step in the diffusion process. Each step is shown with a different color. The mapped coordinates match the target coordinates (RMSD 0.0) after the diffusion process while the unmapped degrees of freedom (unmapped alpha carbons) have new configurations. Side chain addition causes the α carbon trace to be slightly disturbed during energy minimization.
Distribution of Generated Coordinates from Unmapped Amino Acids
We now investigate the distribution of structures we are sampling by running MSBack many times on the HIV-1 CA protein for the same CG targets as shown in Figure . As discussed previously, MSBack works by first applying only the soft constraint which makes the average root-mean-square deviation, ⟨RMSD⟩, between the generated structures and the target coordinates a value of 1.1 Å. The generalized diversity score, DIV, is a metric with a value between 0 and 1, with 1 meaning the generated structures are far from the target or reference structure (see Supporting Information for full definitions). Here, DIV is nearly 0 meaning that the target structure would be an average structure within the distribution of generated structures. The unmapped coordinates have a much higher RMSD (⟨RMSDu⟩ = 1.8 Å), since they are not constrained. Furthermore, the unmapped regions of the protein are specifically loop regions with little secondary structure that should be more dynamic relative to the mapped regions (see Figure ). After applying the hard constraint there is no diversity at all in the mapped regions since there is a one-to-one mapping between alpha carbons and CG beads. The RMSD of the unmapped regions has decreased only slightly, likely due to regularizing the positions where they start. As expected, the first step is generating all the diversity, because it has the highest noise level. We also see that specific loops have varying ⟨RMSD⟩ with long flexible loops being able to develop a higher ⟨RMSD⟩. This suggests that the diffusion process is accurately reflecting the intrinsic flexibility of different regions in the diversity of structures it generates.
2.
Procedure and results for sampling the stochastic backmapping of HIV-1 CA protein. (Left) To assess how the stochastic backmapping procedure creates diverse structures we align the all-atom protein structure the a CG target and then run the diffusion steps many times taking statistics during the process. (Right) After completion, RMSDs for protein sections are given as ⟨RMSDx–y⟩ where x and y are the beginning and ending residue index of the section. Three different generated structures as shown with these sections opaque while all other sections are translucent. ⟨RMSDU⟩ refers to all unmapped amino acids.
Comparison to “True” All-Atom Structures
We then investigated if these results would hold for a completely different CG mapping, which also presents significant backmapping challenges. This is a CG model of the SARS-CoV-2 membrane (M) protein mapped using KMC-CG at an average resolution of 10 amino acids to 1 CG bead. The KMC-CG model allows for amino acids which are nonconsecutive in the protein sequence to be mapped to the same CG bead and allows for a varying number of amino acids to be mapped to the same bead. In some cases, this is 20 amino acids while in some cases it is 4. The M protein has 2 almost entirely solved structures, the long form and compact form as shown at the top of Figure . Each structure has small unsolved regions near the N and C-terminals, which are likely disordered. We included these regions as predicted by AlphaFold (see Supporting Information for details) in the structure but although they are not technically solved, there are likely many different reasonable conformations for these regions. We could have not mapped these regions to the CG representation like in the case of HIV-1 CA; however, we already investigated that case and we wanted to include some flexibility in the CG models we test in Figure . Thus, there are two structures (long and compact) which we consider to be “true”. In Figure , we used our diffusion process to take all-atom structures from one form and diffuse them to the other form using only the CG coordinates. We then compare the diffused structure to the “true” all-atom reference structure. We found that MSBack was able to create structures which matched the CG coordinates perfectly and were visually indistinguishable from the “true” all atom structures as shown at the bottom of Figure . This set of structures did contain some diversity with an average generated RMSD with respect to the alpha carbons, ⟨RMSDCα⟩, of 1.2 Å and DIV of .2 for the long form diffused to the compact form. For the compact form diffused to the long form the numbers were almost identical. In both cases, one can visually see enhanced diversity at the terminal regions, reflective of intrinsic flexibility. Combined with the results for HIV CA protein, this shows the present method is creating a reasonable distribution of all-atom structures for two CG models which are quite complex to backmap in different ways.
3.
Procedure and results for sampling the stochastic backmapping of SARS-CoV-2 M Protein. (Top) The M protein has crystal structures for the long and compact form, both of which can be CG using the same mapping rule. (Left) The CG coordinates for the compact form of the M protein are used as targets for the diffusion process starting from the all-atom structure of the long form. The resultant structures are nearly indistinguishable from the all-atom structures of the long form yet they are structurally diverse. (Right) The reverse process where the CG coordinates for the long form of the M protein are used as targets for the diffusion process starting from the all-atom structure of the compact form.
4.
Backmapping applied to CG trajectories of length T = 11 (configurations taken every 106 CG timesteps) using two different models of the compact SARS-CoV-2 M protein. (A) The process of backmapping each of the trajectories including side chain addition using two different diffusion models for all 22 CG states. (B) Trajectory snapshots for the hENM models using two different cutoff distances. (C) The RMSD at alignment is always ∼8 Å for both models and thus the models are indistinguishable based on this measurement. (D) The RMSD after energy minimization shows performance near that of the reference all-atom structures expect for the model with the lower cutoff distance in the second half of the trajectory.
Evaluation of CG Models
Since we showed that MSBack works well for “true” CG coordinates we wondered how it would perform when the quality of the CG coordinates is decreased or even unknown. To assess this, we created two different heteroelastic network (hENM) models for the compact form of the M protein. The only difference in the hENM models was the cutoff distance where CG beads are no longer bonded, which was 50 Å for one model and 30 Å for the other as shown in Figure B. Although, we were not completely sure of the CG coordinate quality created by these models, we expected them both to be reasonable and that the model with a 30 Å cutoff might perform worse. We ran both models for 10 million timesteps, while backmapping the CG coordinates every 1 million timesteps. The backmapping included adding the side chains and other backbone atoms using two different machine learning methods, DiAMoNDBack and Flowback. This was done to ensure that any problems with backmapping were not a function of how the side chain and additional backbone atoms were applied. We then finished the backmapping process by energy minimizing the structures with no α carbon restraints (as shown in Figure ) and measuring the RMSD postminimization. This RMSD is the metric for backmapping success. In Figure C, we show that the alignment RMSD or the difference between the initial CG coordinates and the CG coordinates generated by the model is nearly constant across both trajectories and approximately 8 Å, which is very similar to the difference between the compact and long forms. Thus, any difference between the backmapping of the “true” CG coordinates and hENM-generated CG coordinates is due to the quality of the coordinates themselves. Figure D shows that backmapping the true coordinates gave a postminimization RMSD of 0.55 Å, which we believe is quite acceptable for a backmapping method and a notable improvement over existing methods. The average RMSD over the course of the 50 Å cutoff trajectory was only slightly higher at 0.61 Å and not dependent on the side chain method. For the 30 Å cutoff model, it was also slightly higher on average at 0.70 Å, which is line with what we would have predicted and suggests that backmapping is a reasonable way to evaluate a CG model. Furthermore, the structures were significantly worse in the second half of the trajectory, which would also be in line with our expectation that it may take some time for the CG model to drift into poor configurations. Overall, these results suggest that the backmapping is sensitive to the quality of the CG coordinates and that this can be a useful way to evaluate the quality of the model that produced the CG coordinates.
Backmapping Large Multi-protein Complexes
Finally, we wanted to test MSBack’s ability to generate large structures made of many backmapped proteins. We backmapped a partial HIV-1 capsid with 954 CA proteins taken from a CG simulation (see Supporting Information for details). A concern was that this structure and other large structures with protein–protein interfaces will have structural clashes since the diffusion process is applied independently as shown in Figure A. To avoid this issue, the side chain and other backbone atoms are added using OpenMSCG reduced nonbonded energy minimization (RNEM) starting from the alpha-carbon traces as has been used in previous backmapping strategies. By using this strategy, we can also include CG small molecules and 4-bead lipids as has been shown previously. , This includes the inositol hexakisphosphate (IP6) anions, which are essential for capsid formation, ,, and part of the CG capsid assembly model. They bind at the pores of CA pentamers and hexamers as shown in yellow in Figure B and can be seen in the all-atom structures in Figure C–E as well. Parameters for IP6 come from the CHARMM General Force field. , As shown in Figure E, this strategy creates good contacts between the R18 residues at the pore of the pentamer and IP6, which is known to be the interaction that stabilizes IP6 binding. This structure has an RMSD of 1.1 Å, which is only slightly higher than the best structures in Figure . This is a promising result given the propensity for clashes in such a tightly packed structure. In Figures S3 and S4, we show that using a diffusion model like Flowback can lead to lower RMSD structures but often creates clashes when the structures are large. A hybrid method that generates side chain and additional backbone atoms with Flowback and then fixes the clashes and adds all small molecules with RNEM achieves a slightly lower postminimization RMSD, but there are less R18-IP6 side chain contacts.
5.
Process for backmapping a nearly complete CG HIV-1 capsid with bound inositol hexakisophosphatehexakisphosphate (IP6). (A) Procedures for backmapping large configurations of proteins and other molecules. (B) The CG structure for the partial capsid with 954 CA monomers. CA hexamers are shown in green, CA pentamers in red, disordered oligomers in gray, and IP6 in yellow. (C) The all-atom structure with N-terminal domains shown in line representations, C-terminal domains in gray surface representations and van der Waals representations of IP6. A clear vacancy is seen in the same location as the corresponding gray region in (B). One hexamer with bound IP6 at the pore is circled and shown in more detail in (D). (E) The R18 side chains at the pore are shown to be interacting strongly with the IP6 at the hexamer pore from (D).
Discussion
We have demonstrated that a MSBack can be used to backmap highly CG representations of proteins to all-atom structures which match the CG coordinates with high precision. The degrees of freedom lost in the CG mapping follow a reasonably diverse distribution in line with realistic protein conformations. This method is functional to the point where we can differentiate between 2 different CG force fields and accurately detect which gives more realistic CG configurations of the all-atom structure. It can be successfully integrated with several existing solutions for completing the backmapping process from the α carbon trace. We find that reduced nonbonded energy minimization in OpenMSCG is an effective way to avoid clashes for very large systems of backmapped proteins and to add all-atom structures from backmapped lipids and small molecules. In the future, we will be able to backmap highly complex CG protein/lipid/small molecule simulations such as the HIV-1 capsid docked in the pore of the nuclear membrane and bound with small molecules. These systems will likely be much too large to simulate in all-atom MD for long periods of time; however, we can focus in on regions of interest and run all-atom simulations starting from configurations that can only be generated via CG MD. Furthermore, we can use this model to iteratively simulate between different CG models. For instance, one can envision a self-assembling CG model of a protein complex and a different model of the complex trained for mechanical robustness with completely unrelated CG mapping functions (this is the current case for the HIV-1 capsid). We could now use the self-assembling model to create polymorphic structures, backmap them to all-atom, and then forward map them to a completely different CG model which is more mechanically accurate to see differences in morphology from the first model affect mechanical properties in the second model.
We also anticipate that our backmapping strategy may be useful in other problems where there is limited information about protein structure. If this limited information about a structure is a differentiable function of the all-atom coordinates, then one can define a residual that can be minimized via a similar constraint on the diffusion process. One example is the interpretation of cryo-EM densities , by using the difference between those densities and the densities from the all-atom structure as the residual which allows a diffusion process to find a close-to-ideal all-atom structure. However, these examples were mostly limited to single proteins with less than 40 residues, while cryo-EM is being used to produce structures at the length scales of hundreds of nm that include nonprotein densities, such as the HIV-1 capsid bound in the nuclear pore. As we have shown, MSBack is able to handle these types of structures. Another example could be an elastic model of a large protein assembly with a certain strain distribution that can be calculated precisely from the all-atom coordinates. One issue could be the need to define a linear transformation of the coordinates as the hard constraint, although even without the hard constraint we show that reasonable, if imperfect structures can be found.
Overall, what has been accomplished in this work is the combination of data-driven tools for sampling protein structures under constraint with other physics-based tools for backmapping proteins and small molecules. We anticipate that all of these tools will continue to improve by becoming more accurate with respect to protein structures, increasing the diversity of molecules they can model, and becoming more computationally efficient. Thus, we expect the specific tools for state-of-the-art backmapping to continue to evolve. However, by thinking of backmapping as an inverse design problem using highly CG structures as constraints to generate an accurate distribution of possible all-atom conformations promises to continue to be a successful strategy.
Methods
CG Simulations of the HIV-1 CA Protein
The model of HIV-1 CA assembly comes from previous work with the addition of a model of the RNP, which is fully explained in the Supporting Information. A virion-like boundary condition is also created by placing immobile particles on the surface of a 144 nm sphere to match the approximate size of the virion. The area per particle is 12 nm2 /particle and they interact uniformly with all other species at 4.4 nm. The sphere contains 1500 CA dimers, 3400 inert crowders, 1042 IP6, and one RNP. All CGMD simulations were performed using LAMMPS (21 July 2020). All simulations were performed in a constant NVT ensemble using a Langevin thermostat at 300 K with a damping period of 100 ps. The fraction of active CA monomers ([CA]+/[CA]) was 0.15, following previous work. With CG time step τ = 10 fs, state switching interval of 5 × 105 τ corresponds to 5 ns of CG time. Therefore, [CA]+/[CA]− UCG state switching was attempted every 5 × 105 τ.
AA Simulations of the SARS-CoV-2 M Protein
To obtain full-length compact and long forms of M protein, we used AlphaFold2-multimer to predict the full-length structures using the cryo-EM structures of compact (PDB 8ctk ) and long forms (PDB 7vgr ), respectively. The full-length M structures (RMSD ∼ 2 Å compared to cryo-EM structures) were then used to perform simulations. All all-atom MD simulations of the M protein were performed using GROMACS2022. The crystal structure of the compact form of the M Protein (PDB 8ctk) was used for the simulation. CHARMM-GUI was used to generate a 30 nm x 30 nm lipid bilayer membrane (35% DPPC, 35% DLIPC, 30% CHL1) and insert M protein. − TIP3P water is added to the system to generate the solvated structure and the water on both sides of the membrane is approximately 4.5 nm thick. We neutralized the system with a NaCl concentration of 150 mM. The periodic boundary condition was active across the three dimensions of the simulation box. The CHARMM36 force field was used with the particle mesh Ewald sum method for long-range electrostatic interactions. After energy minimization with the steepest descent algorithm, the system was pre-equilibrated at 1 atm and 310 K (the constant NVT ensemble followed by the constant NPT ensemble). The production runs were performed for 500 ns with a 2 fs time step in the constant NPT ensemble using Parrinello–Rahman pressure coupling and Nose-Hoover temperature coupling. The last 450 ns trajectory was used for CG modeling. The process of CG model selection and parametrization can be found in the Supporting Information.
CG Simulations of the SARS-CoV-2 M Protein
All CG MD simulations were prepared and simulated in the LAMMPS MD software. All CG MD simulations were performed using a CG MD time step of 15 fs, and the equations of motions were integrated with the Velocity Verlet algorithm. All simulations were periodic in all three directions. The temperature of the system was maintained using Langevin thermostat at 310 K. Simulation trajectory snapshots were saved every 1000 CG MD timesteps.
Side Chain Generation with RNEM and Full Energy Minimization
In the main text, we used OpenMSCG reduced nonbonded energy minimization (RNEM) to generate the side chains. This is done following previous work by starting with the alpha carbons and then placing the stochastically placing the rest of the atoms near the corresponding α carbon. Energy minimization occurs via a modified AA force field (Charmm36) where nonbonded interactions are replaced with cosine functions to guarantee numerical stability and then additional torsions are added to maintain isomeric properties of the peptide backbone. During this process we add a harmonic potential as a positional restraint on the alpha carbons. The force constant is the default 1000 kJ/mol/nm2 in Figure . At the completion of the RNEM step, the RMSD with respect to the alpha carbons given by the diffusion model is 0.7 Å. The rest of the 1.1 Å comes from the energy minimization step with no restraints and the full all-atom force field. Other strategies are explored in the Supporting Information. Energy minimization is performed in Gromacs using the steepest descent algorithm and the Charmm36 force field. We energy minimize in vacuum (without water, ions, or any other solvent) with no restraints of any kind for 10,000 steps with early exit if the maximum force goes below the EM tolerance of 10 kJ/mol/nm although this rarely occurs.
RMSD Definitions
There are many different RMSD’s used in this manuscript. Without any other subscripts or markings the RMSD is defined as the RMSD between the mapped all-atom coordinates and the target CG coordinates as
| 9 |
When specified in the text this RMSD can also be with respect to an all-atom reference state. We may also state that at what point in the backmapping process this is done. For example, the alignment RMSD is performed before any diffusion takes place and the Post-Minimization RMSD is performed after energy minimization. To denote an average over an ensemble of all-atom structures, we add ⟨ ⟩ to make ⟨RMSD⟩. Mathematically,
| 10 |
To denote this ensemble is a trajectory we add the subscript T, ⟨RMSD⟩T and to denote this is an average over both the trajectory and multiple methods for side chain addition we add the subscript T∗SC, ⟨RMSD⟩T∗SC. A direct subscript of the RMSD informs how the RMSD was taken and implies that this an RMSD of the all-atom α carbon coordinates with respect to other generated all-atom α carbon coordinates. For example, the subscript U denotes that only the unmapped part of the protein (the amino acids which do not correspond to any CG beads). Thus, ⟨RMSDU⟩ denotes the average RMSD between all generated structures for the unmapped part of the protein. Instead of U, as series of residues is also often used. For example, RMSD100–105 is the RMSD calculated just on residues 100, 101, 102, 103, 104, and 105. The subscript gen here denotes that the RMSD is taken for all unique pairs of structures which were both generated by MSBack.
| 11 |
Diversity Scores
Ideally, the all-atom configurations we create would sample the Boltzmann distribution of structures given the constraints of the CG coordinates as defined in the first equation. However, an exact quantification of this is difficult and instead other metrics are often used in place. The generated RMSD is a good measure of how different the structures are from one another however a high generated diversity may mean that the structures are inaccurate. In contrast, the RMSD with respect to a reference is a good measure of whether the generated structures are accurate. These two are combined into one metric, which is the diversity score, DIV. This metric seeks to measure whether the reference structure is an average structure within the generated distribution. If it is far from the generated distribution, then ⟨RMSD⟩ will be significantly larger than ⟨RMSDgen⟩ , while the two will be identical if the reference structure is simply an average structure within the generated distribution. Given the definition below
| 12 |
The case where the reference structure is an average structure in the distribution corresponds to DIV approaching 0. If DIV is large, then the reference structure is not a part of the generated distribution.
Supplementary Material
Acknowledgments
We thank Dr. Sahithya Sridharan Iyer for her work making the LAMMPS code run on GPUs for the UCG simulations of the HIV-1 capsid. Research reported in this publication was supported in part by the Behavior of HIV in Viral Environments (B-HIVE) Center of the National Institute of Allergy and Infectious Diseases (NIAID) of the National Institutes of Health (NIH) under award number U54AI170855, and in part by the NIH NIAID award number R01AI178850. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. C.W. also acknowledges the support of an NIH Ruth L. Kirschstein National Research Service Award (NRSA) Postdoctoral Fellowship under award number F32AI186429 and in part from a Chicago Center for Theoretical Chemistry (CCTCh) Postdoctoral Fellowship. Y.W. acknowledges the support of an Eric and Wendy Schmidt Artificial Intelligence in Science Postdoctoral Fellowship and prior to that a CCTCh Postdoctoral Fellowship. Computational resources were provided by the University of Chicago Research Computing Center (RCC).
MSBack will soon be available in a branch of OpenMSCG (https://software.rcc.uchicago.edu/mscg/). For now, all necessary code is available in the following repositories. Code for the diffusion process is available here (https://github.com/waltmann1/MyChroma). Code for the alignment and manipulation of CG and all-atom structures is available here (https://github.com/waltmann1/MSBack-Build). LAMMPS simulation data and initial structure files for the CG simulations are available here (https://github.com/waltmann1/RNP/tree/main/run_system/RNP_model_3_integrase_memb).
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jctc.5c00459.
Additional details on the coarse-grained models which were backmapped and intricacies of the diffusion process including the relative success of other possible protocols (PDF)
The authors declare no competing financial interest.
References
- Hollingsworth S. A., Dror R. O.. Molecular Dynamics Simulation for All. Neuron. 2018;99(6):1129–1143. doi: 10.1016/j.neuron.2018.08.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schlick T., Portillo-Ledesma S.. Biomolecular modeling thrives in the age of technology. Nat. Comput. Sci. 2021;1(5):321–331. doi: 10.1038/s43588-021-00060-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brooks C. L. 3rd, MacKerell A. D. Jr, Post C. B., Nilsson L.. Biomolecular dynamics in the 21st century. Biochim. Biophys. Acta, Gen. Subj. 2024;1868(2):130534. doi: 10.1016/j.bbagen.2023.130534. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jin J., Pak A. J., Durumeric A. E. P., Loose T. D., Voth G. A.. Bottom-up Coarse-Graining: Principles and Perspectives. J. Chem. Theory Comput. 2022;18(10):5759–5791. doi: 10.1021/acs.jctc.2c00643. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noid W. G.. Perspective: Advances, Challenges, and Insight for Predictive Coarse-Grained Models. J. Phys. Chem. B. 2023;127(19):4174–4207. doi: 10.1021/acs.jpcb.2c08731. [DOI] [PubMed] [Google Scholar]
- Hurdait A., Voth G. A.. HIV-1 Capsid Shape, Orientation, and Entropic Elasticity Regulate Translocation into the Nuclear Pore Complex. Proc. Nat. Acad. Sci. U.S.A. 2024;121(4):e2313737121. doi: 10.1073/pnas.2313737121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gupta M., Pak A. J., Voth G. A.. Critical mechanistic features of HIV-1 viral capsid assembly. Sci. Adv. 2023;9(1):eadd7434. doi: 10.1126/sciadv.add7434. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jarin Z., Tsai F. C., Davtyan A., Pak A. J., Bassereau P., Voth G. A.. Unusual Organization of I-BAR Proteins on Tubular and Vesicular Membranes. Biophys. J. 2019;117(3):553–562. doi: 10.1016/j.bpj.2019.06.025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pak A. J., Dannenhoffer-Lafage T., Madsen J. J., Voth G. A.. Systematic Coarse-Grained Lipid Force Fields with Semiexplicit Solvation via Virtual Sites. J. Chem. Theory Comput. 2019;15(3):2087–2100. doi: 10.1021/acs.jctc.8b01033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sahrmann P. G., Voth G. A.. Enhancing the Assembly Properties of Bottom-Up Coarse-Grained Phospholipids. J. Chem. Theory Comput. 2024;20(22):10235–10246. doi: 10.1021/acs.jctc.4c00905. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z., Pfaendtner J., Grafmüller A., Voth G. A.. Defining coarse-grained representations of large biomolecules and biomolecular complexes from elastic network models. Biophys. J. 2009;97(8):2327–2337. doi: 10.1016/j.bpj.2009.08.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z.. Systematic methods for defining coarse-grained maps in large biomolecules. Adv. Exp. Med. Biol. 2015;827:33–48. doi: 10.1007/978-94-017-9245-5_4. [DOI] [PubMed] [Google Scholar]
- Zhu Y., Zhao X., Xiang C., Liu X., Li J.. Evaluation of Essential Dynamics and Fixed-Length Coarse Graining for Multidomain Proteins. J. Phys. Chem. B. 2024;128(21):5147–5156. doi: 10.1021/acs.jpcb.3c08198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wu Z., Zhang Y., Zhang J. Z., Xia K., Xia F.. Determining Optimal Coarse-Grained Representation for Biomolecules Using Internal Cluster Validation Indexes. J. Comput. Chem. 2020;41(1):14–20. doi: 10.1002/jcc.26070. [DOI] [PubMed] [Google Scholar]
- Wu J., Xue W., Voth G. A.. K-Means Clustering Coarse-Graining (KMC-CG): A Next Generation Methodology for Determining Optimal Coarse-Grained Mappings of Large Biomolecules. J. Chem. Theory Comput. 2023;19(23):8987–8997. doi: 10.1021/acs.jctc.3c01053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Foley T. T., Shell M. S., Noid W. G.. The impact of resolution upon entropy and information in coarse-grained models. J. Chem. Phys. 2015;143(24):243104. doi: 10.1063/1.4929836. [DOI] [PubMed] [Google Scholar]
- Dama J. F., Jin J., Voth G. A.. The Theory of Ultra-Coarse-Graining. 3. Coarse-Grained Sites with Rapid Local Equilibrium of Internal States. J. Chem. Theory Comput. 2017;13(3):1010–1022. doi: 10.1021/acs.jctc.6b01081. [DOI] [PubMed] [Google Scholar]
- Davtyan A., Dama J. F., Sinitskiy A. V., Voth G. A.. The Theory of Ultra-Coarse-Graining. 2. Numerical Implementation. J. Chem. Theory Comput. 2014;10(12):5265–5275. doi: 10.1021/ct500834t. [DOI] [PubMed] [Google Scholar]
- Dama J. F., Sinitskiy A. V., McCullagh M., Weare J., Roux B., Dinner A. R., Voth G. A.. The Theory of Ultra-Coarse-Graining. 1. General Principles. J. Chem. Theory Comput. 2013;9(5):2466–2480. doi: 10.1021/ct4000444. [DOI] [PubMed] [Google Scholar]
- Kim S.. Backmapping with Mapping and Isomeric Information. J. Phys. Chem. B. 2023;127(49):10488–10497. doi: 10.1021/acs.jpcb.3c05593. [DOI] [PubMed] [Google Scholar]
- Yu A., Lee E. M. Y., Briggs J. A. G., Ganser-Pornillos B. K., Pornillos O., Voth G. A.. Strain and rupture of HIV-1 capsids during uncoating. Proc. Natl. Acad. Sci. U. S. A. 2022;119(10):e2117781119. doi: 10.1073/pnas.2117781119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hocky G. M., Dannenhoffer-Lafage T., Voth G. A.. Coarse-Grained Directed Simulation. J. Chem. Theory Comput. 2017;13(9):4593–4603. doi: 10.1021/acs.jctc.7b00690. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mahmoud A. H., Masters M., Lee S. J., Lill M. A.. Accurate Sampling of Macromolecular Conformations Using Adaptive Deep Learning and Coarse-Grained Representation. J. Chem. Inf. Model. 2022;62(7):1602–1617. doi: 10.1021/acs.jcim.1c01438. [DOI] [PubMed] [Google Scholar]
- Zheng S., He J., Liu C., Shi Y., Lu Z., Feng W., Ju F., Wang J., Zhu J., Min Y.. et al. Predicting equilibrium distributions for molecular systems with deep learning. Nat. Mach. Intell. 2024;6:558–567. doi: 10.1038/s42256-024-00837-3. [DOI] [Google Scholar]
- Watson J. L., Juergens D., Bennett N. R., Trippe B. L., Yim J., Eisenach H. E., Ahern W., Borst A. J., Ragotte R. J., Milles L. F.. et al. De novo design of protein structure and function with RFdiffusion. Nature. 2023;620(7976):1089–1100. doi: 10.1038/s41586-023-06415-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jones M. S., Shmilovich K., Ferguson A. L.. DiAMoNDBack: Diffusion-Denoising Autoregressive Model for Non-Deterministic Backmapping of Cα Protein Traces. J. Chem. Theory Comput. 2023;19(21):7908–7923. doi: 10.1021/acs.jctc.3c00840. [DOI] [PubMed] [Google Scholar]
- Shmilovich K., Stieffenhofer M., Charron N. E., Hoffmann M.. Temporally Coherent Backmapping of Molecular Trajectories From Coarse-Grained to Atomistic Resolution. J. Phys. Chem. A. 2022;126(48):9124–9139. doi: 10.1021/acs.jpca.2c07716. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jones, M. ; Khanna, S. ; Ferguson, A. . Flowback: A Flow-matching Approach for the Generative Backmapping of Macromolecules; OpenReview Sponsors, 2024. [Google Scholar]
- Christofi E., Bačová P., Harmandaris V. A.. Physics-Informed Deep Learning Approach for Reintroducing Atomic Detail in Coarse-Grained Configurations of Multiple Poly(lactic acid) Stereoisomers. J. Chem. Inf. Model. 2024;64(6):1853–1867. doi: 10.1021/acs.jcim.3c01870. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li W., Burkhart C., Polińska P., Harmandaris V., Doxastakis M.. Backmapping coarse-grained macromolecules: An efficient and versatile machine learning approach. J. Chem. Phys. 2020;153(4):041101. doi: 10.1063/5.0012320. [DOI] [PubMed] [Google Scholar]
- Chennakesavalu S., Toomer D. J., Rotskoff G. M.. Ensuring thermodynamic consistency with invertible coarse-graining. J. Chem. Phys. 2023;158(12):124126. doi: 10.1063/5.0141888. [DOI] [PubMed] [Google Scholar]
- Ingraham J. B., Baranov M., Costello Z., Barber K. W., Wang W., Ismail A., Frappier V., Lord D. M., Ng-Thow-Hing C., Van Vlack E. R.. et al. Illuminating protein space with a programmable generative model. Nature. 2023;623(7989):1070–1078. doi: 10.1038/s41586-023-06728-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Renner N., Kleinpeter A., Mallery D. L., Albecka A., Rifat Faysal K. M., Böcking T., Saiardi A., Freed E. O., James L. C.. HIV-1 is dependent on its immature lattice to recruit IP6 for mature capsid assembly. Nat. Struct. Mol. Biol. 2023;30(3):370–382. doi: 10.1038/s41594-022-00887-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Peng Y., Pak A. J., Durumeric A. E. P., Sahrmann P. G., Mani S., Jin J., Loose T. D., Beiter J., Voth G. A.. OpenMSCG: A Software Tool for Bottom-Up Coarse-Graining. J. Phys. Chem. B. 2023;127(40):8537–8550. doi: 10.1021/acs.jpcb.3c04473. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Grime J. M. A., Dama J. F., Ganser-Pornillos B. K., Woodward C. L., Jensen G. J., Yeager M., Voth G. A.. Coarse-grained simulation reveals key features of HIV-1 capsid self-assembly. Nat. Commun. 2016;7:11568. doi: 10.1038/ncomms11568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noid W. G., Chu J. W., Ayton G. S., Krishna V., Izvekov S., Voth G. A., Das A., Andersen H. C.. The multiscale coarse-graining method. I. A rigorous bridge between atomistic and coarse-grained models. J. Chem. Phys. 2008;128(24):244114. doi: 10.1063/1.2938860. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noid W. G., Liu P., Wang Y., Chu J. W., Ayton G. S., Izvekov S., Andersen H. C., Voth G. A.. The multiscale coarse-graining method. II. Numerical implementation for coarse-grained molecular models. J. Chem. Phys. 2008;128(24):244115. doi: 10.1063/1.2938857. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson B. D. O.. Reverse-Time Diffusion Equation Models. Stochastic. Process. Appl. 1982;12:313–326. doi: 10.1016/0304-4149(82)90051-5. [DOI] [Google Scholar]
- Song, Y. ; Sohl-Dickstein, J. ; Kingma, D. P. ; Kumar, A. ; Ermon, S. ; Poole, B. . Score-Based Generative Modeling through Stochastic Differential Equations. arXiv, 2021, 10.48550/arXiv.2011.13456. [DOI] [Google Scholar]
- Liu, Y. ; Chen, M. ; Lin, G. . Backdiff: a diffusion model for generalized transferable protein backmapping. arXiv, 2023, 10.48550/arXiv.2310.01768. [DOI] [Google Scholar]
- Chung, H. ; Kim, J. ; Mccann, M. T. ; Klasky, M. L. ; Ye, J. C. . Diffusion Posterior Sampling for General Noisy Inverse Problems. arXiv, 2022, 10.48550/arXiv.2209.14687. [DOI] [Google Scholar]
- Abraham M. J., Murtola T., Schulz R., Páll S., Smith J. C., Hess B., Lindahl E.. GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX. 2015;1-2:19–25. doi: 10.1016/j.softx.2015.06.001. [DOI] [Google Scholar]
- Dolan K. A., Dutta M., Kern D. M., Kotecha A., Voth G. A., Brohawn S. G.. Structure of SARS-CoV-2 M protein in lipid nanodiscs. Elife. 2022;11:e81702. doi: 10.7554/eLife.81702. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z., Nomura N., Muramoto Y., Ekimoto T., Uemura T., Liu K., Yui M., Kono N., Aoki J., Ikeguchi M.. et al. Structure of SARS-CoV-2 membrane protein essential for virus assembly. Nat. Commun. 2022;13(1):4399. doi: 10.1038/s41467-022-32019-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A.. et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lyman E., Pfaendtner J., Voth G. A.. Systematic multiscale parameterization of heterogeneous elastic network models of proteins. Biophys. J. 2008;95(9):4183–4192. doi: 10.1529/biophysj.108.139733. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dick R. A., Mallery D. L., Vogt V. M., James L. C.. IP6 Regulation of HIV Capsid Assembly, Stability, and Uncoating. Viruses. 2018;10(11):640. doi: 10.3390/v10110640. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dick R. A., Zadrozny K. K., Xu C., Schur F. K. M., Lyddon T. D., Ricana C. L., Wagner J. M., Perilla J. R., Ganser-Pornillos B. K., Johnson M. C.. et al. Inositol phosphates are assembly co-factors for HIV-1. Nature. 2018;560(7719):509–512. doi: 10.1038/s41586-018-0396-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vanommeslaeghe K., MacKerell A. D. Jr. Automation of the CHARMM General Force Field (CGenFF) I: Bond Perception and Atom Typing. J. Chem. Inf. Model. 2012;52(12):3144–3154. doi: 10.1021/ci300363c. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vanommeslaeghe K., Prabhu R. E., MacKerell A. D. Jr. Automation of the CHARMM General Force Field (CGenFF) II: Assignment of Bonded Parameters and Partial Atomic Charges. J. Chem. Inf. Model. 2012;52(12):3155–3168. doi: 10.1021/ci3003649. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu A., Lee E. M. Y., Jin J., Voth G. A.. Atomic-scale characterization of mature HIV-1 capsid stabilization by inositol hexakisphosphate (IP. Sci. Adv. 2020;6(38):abc6465. doi: 10.1126/sciadv.abc6465. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu Y., Yu Z., Lindsay R. J., Lin G., Chen M., Sahoo A., Hanson S. M.. ExEnDiff: An Experiment-guided Diffusion model for protein conformational Ensemble generation. bioRxiv. 2024;2024:616517. doi: 10.1101/2024.10.04.616517. [DOI] [Google Scholar]
- Yu Z., Liu Y., Lin G., Jiang W., Chen M.. ESMAdam: a plug-and-play all-purpose protein ensemble generator. bioRxiv. 2025;2025:633818. doi: 10.1101/2025.01.19.633818. [DOI] [Google Scholar]
- Kreysing J. P., Heidari M., Zila V., Cruz-Leon S., Obarska-Kosinska A., Laketa V., Rohleder L., Welsch S., Kofinger J., Turonova B., Hummer G.. et al. Passage of the HIV capsid cracks the nuclear pore. Cell. 2025;188(4):930–943.e21. doi: 10.1016/j.cell.2024.12.008. [DOI] [PubMed] [Google Scholar]
- Plimpton S.. Fast Parallel Algorithms for Short-Range Molecular Dynamics. J. Comput. Phys. 1995;117(1):1–19. doi: 10.1006/jcph.1995.1039. [DOI] [Google Scholar]
- Vanden-Eijnden E., Ciccotti G.. Second-order integrators for Langevin equations with holonomic constraints. Chem. Phys. Lett. 2006;429(1–3):310–316. doi: 10.1016/j.cplett.2006.07.086. [DOI] [Google Scholar]
- Evans R., O’Neill M., Pritzel A., Antropova N., Senior A., Green T., Žídek A., Bates R., Blackwell S., Yim J.. et al. Protein complex prediction with AlphaFold-Multimer. bioRxiv. 2022;2021:463034. doi: 10.1101/2021.10.04.463034. [DOI] [Google Scholar]
- Lee J., Cheng X., Swails J. M., Yeom M. S., Eastman P. K., Lemkul J. A., Wei S., Buckner J., Jeong J. C., Qi Y.. et al. CHARMM-GUI Input Generator for NAMD, GROMACS, AMBER, OpenMM, and CHARMM/OpenMM Simulations Using the CHARMM36 Additive Force Field. J. Chem. Theory Comput. 2016;12(1):405–413. doi: 10.1021/acs.jctc.5b00935. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jo S., Kim T., Iyer V. G., Im W.. CHARMM-GUI: a web-based graphical user interface for CHARMM. J. Comput. Chem. 2008;29(11):1859–1865. doi: 10.1002/jcc.20945. [DOI] [PubMed] [Google Scholar]
- Brooks B. R., Brooks C. L. 3rd, Mackerell A. D. Jr, Nilsson L., Petrella R. J., Roux B., Won Y., Archontis G., Bartels C., Boresch S.. et al. CHARMM: the biomolecular simulation program. J. Comput. Chem. 2009;30(10):1545–1614. doi: 10.1002/jcc.21287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jorgensen W. L., Chandrasekhar J., Madura J. D., Impey R. W., Klein M. L.. Comparison of Simple Potential Functions for Simulating Liquid Water. J. Chem. Phys. 1983;79(2):926–935. doi: 10.1063/1.445869. [DOI] [Google Scholar]
- Huang J., Rauscher S., Nawrocki G., Ran T., Feig M., de Groot B. L., Grubmuller H., MacKerell A. D. Jr. CHARMM36m: an improved force field for folded and intrinsically disordered proteins. Nat. Methods. 2017;14(1):71–73. doi: 10.1038/nmeth.4067. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
MSBack will soon be available in a branch of OpenMSCG (https://software.rcc.uchicago.edu/mscg/). For now, all necessary code is available in the following repositories. Code for the diffusion process is available here (https://github.com/waltmann1/MyChroma). Code for the alignment and manipulation of CG and all-atom structures is available here (https://github.com/waltmann1/MSBack-Build). LAMMPS simulation data and initial structure files for the CG simulations are available here (https://github.com/waltmann1/RNP/tree/main/run_system/RNP_model_3_integrase_memb).





