Abstract
Here we present a series of tutorials demonstrating the use of various methods which integrate structural mass spectrometry (MS) data with computational protein structure prediction methods. We give usage examples of widely used modeling frameworks, including Rosetta-based approaches (ab initio modeling, comparative modeling, and protein-protein docking) and deep learning methods such as AlphaFold2. We then describe strategies for incorporating covalent labeling, ion mobility, and surface-induced dissociation MS data into these workflows through Rosetta scoring terms and specialized applications. Finally, we provide instructions on calculating structural metrics, such as solvent accessibility, collision cross sections, and energy-resolved MS data and comparing them to actual MS data. We also introduce new PyRosetta implementations of the PARCS algorithm and the SID_ERMS_Rescore application. Together, these tutorials provide a comprehensive framework for integrating computational modeling with structural MS to enhance protein structure prediction.
Introduction
Accurate protein structures are essential for understanding biological function, molecular interactions, and reaction mechanisms. Structural mass spectrometry (MS) techniques can provide valuable structural information, such as residue-level solvent exposure, overall protein shape, and connectivity between subunits in protein complexes. However, MS measurements alone cannot fully resolve a protein’s three-dimensional structure. Computational methods provide a framework for predicting protein structures directly from primary amino acid sequences, while MS-derived data has the potential to provide experimental constraints which can bias these methods toward native-like conformations. When integrated with computational approaches, ranging from physics-based sampling methods to modern deep learning networks, MS data can guide and refine structural models, improving accuracy and reliability of resulting protein structures.1,2
Despite these advances, integrating structural MS data with computational modeling remains challenging for many researchers. The growing number of available prediction tools, scoring methods, and analysis tools, combined with increasingly complex MS datasets, can make workflows difficult to implement and interpret. This is particularly true for users without extensive computational training. In practice, improper use of experimental restraints, insufficient conformational sampling, and overinterpretation of unvalidated prediction models can lead to structural conclusions that are not fully supported by the available data.
Our group has developed a set of computational algorithms that enable automated structure elucidation from different MS data types. In this review, we provide practical tutorials describing how structural mass spectrometry data can be integrated with these computational structure prediction workflows. These tutorials are intended to provide practical guidance for researchers interested in incorporating structural mass spectrometry data into computational protein structure prediction workflows. In particular, the manuscript is primarily designed for users who may be familiar with structural MS techniques but have limited experience implementing computational modeling approaches, such as Rosetta or AlphaFold2. By presenting step-by-step examples on how to use not only the tools themselves but also how to then integrate MS data, we aim to make these workflows more accessible and easier to incorporate within in-house pipelines.
While many of the computational methods are individually well documented, it can be challenging to determine what tools are best for a given task and their output can be difficult to interpret by a new user. The tutorial format is motivated to provide clear instructions on how to implement both computational structure prediction methods and how to incorporate experimental guidance (Figure 1). The examples presented in this manuscript are therefore intended not only to demonstrate software usage, but also to provide context for how these methods are commonly applied in integrative structural biology. The first section introduces commonly used computational structure prediction tools, including Rosetta-based approaches such as ab initio modeling, comparative modeling, and protein-protein docking, as well as deep learning-based methods such as AlphaFold2. In addition to describing these methods, we highlight practical considerations for selecting appropriate workflows, preparing inputs, interpreting scores, and understanding common limitations associated with MS-guided modeling. The second section focuses on incorporating MS-derived information into structural modeling through Rosetta score terms and dedicated rescoring applications that leverage data from techniques such as covalent labeling, ion mobility, and surface-induced dissociation. Particular emphasis is placed on the role of rescoring with experimental data to help distinguish native-like conformations from unrealistic predictions, which can be a limitation of conformational sampling tools like Rosetta. The third section describes how to calculate structural metrics commonly used in MS-guided modeling, including solvent exposure measures, collision cross sections, and energy resolved mass spectrometry data, using both Rosetta and PyRosetta. These structural metrics are not only important in the context of rescoring but serve as comparison to experimentally derived quantities for more direct interpretations. Together, these tutorials aim to provide researchers with guidance for combining computational modeling and structural mass spectrometry to improve protein structure prediction and analysis. We have also developed new PyRosetta versions of our PARCS algorithm and SID_ERMS_Rescore application which are presented for the first time in this tutorial.
Figure 1.

Overview of the integration of structural mass spectrometry data with computational protein structure prediction workflows. Experimental structural mass spectrometry (MS) measurements discussed include covalent labeling (CL-MS), ion mobility (IM-MS), and surface induced-dissociation (SID-MS). Data from these experiments can be processed to determine structural information such natural log of protection factor (lnPF) or modification percentage, protein collision cross sections, and energy-resolved mass spectrometry profiles. This information can then be incorporated into computational methods such as Rosetta and AlphaFold2 to improve prediction accuracy. The numbered panels illustrate the general procedure from experimental MS measurements (1-2), computational model generation (3), and MS-guided model evaluation (4).
Section 1. General Structure Prediction Tutorials
This section covers how to generate computational structures for proteins and protein complexes using traditional methods with the Rosetta software suite (such as Ab initio modeling, comparative/homology modeling, and protein-protein docking). The modern deep learning technique AlphaFold2 will also be explained. Each subsection covers different methods, and it is assumed the necessary software has already been installed and configured. Importantly, both Rosetta and AlphaFold2 require more than introductory computer experience to use effectively. Users should be familiar with working in Linux-based operating systems, and performing basic scripting or coding tasks (e.g., Python). Rosetta and AlphaFold can be installed following the instructions indicated on their respective GitHub pages: https://github.com/RosettaCommons/rosetta and https://github.com/google-deepmind/alphafold. Rosetta can be installed natively on Mac/Linux, whereas Windows is not officially supported and generally requires the use of a Linux virtual machine such as Windows Subsystem for Linux (WSL) or dual booting. AlphaFold is supported only on Linux systems.
Both Rosetta and AlphaFold2 require substantial computational resources, and performance is strongly dependent on available hardware. Rosetta calculations are primarily CPU-based and can often be run on modern multi-core desktops or laptops for smaller jobs; however, large-scale modeling or docking calculations are best performed on high-performance computing (HPC) systems or dedicated lab workstations with high core counts and large memory capacity. As a practical minimum, users should ideally have access to a Linux/macOS system with at least 16-32 GB RAM and multiple CPU cores for standard Rosetta workflows. AlphaFold2 inference can technically be performed on CPUs, but runtimes are often prohibitively long. Therefore, the use of a modern NVIDIA GPU with substantial VRAM (typically ≥ 16 GB, with larger proteins benefiting from 24-48 GB VRAM) is highly recommended. In practice, AlphaFold2 is most commonly run on GPU-enabled HPC clusters or dedicated GPU workstations. For users without access to such hardware to run AlphaFold2, an alternative implementation called ColabFold3 can be run through Google Colab notebooks in a web browser, although extensive use may require paid GPU access. All tutorial files are located in Supplementary File 1 called ‘tutorials.zip’, with subdirectories corresponding to each section. Many of the settings we use are recommended, but we encourage users to explore the documentation of these methods if further customization is desired.
Notes:
- Files, options, and commands are indicated in the following format:
~/rosetta/source/bin/relax.linuxgccrelease
-
When using PDBs with Rosetta, it may be necessary to clean these files. This entails removing all non-protein atom lines, such as header information or ligands and waters. Rosetta has a ‘clean_pdb.py’ script, but it should be used carefully since it will renumber the PDB and skip residues with zero occupancy. A PDB can be cleaned with this script using the following command:
> python ~/rosetta/tools/protein_tools/scripts/clean_pdb.py 1PGA.pdb A
This will generate a cleaned PDB containing only chain A. Multiple chains can be specified by joining them, such as ‘AB’ or ‘CDEF’.
- Alternatively, the following grep command will take ‘SOME.pdb’ and only transfer lines beginning with ‘ATOM’ to a new PDB.
> grep “^ATOM” SOME.pdb > SOME.clean.pdb
-
The full name of Rosetta applications can vary depending on the OS it was installed in and the compiler that was used. The naming convention has this format:
~/Rosetta/source/bin/application.<os><compiler><mode>
- Where:
- <os> is the operating system Rosetta was compiled on (Linux or MacOS)
- <compiler> is the compiler used (gcc, clang, etc.)
- <mode> is the mode chosen during compiling (release, debug).
For this tutorial, all Rosetta applications were installed on Linux with GCC and release mode. Thus, they will all have this format:~/Rosetta/source/bin/application.linuxgccrelease
1.1. Rosetta: Ab initio Modeling
Ab initio (or de novo) folding simulations were among the earliest Rosetta methods developed to solve the “protein folding problem”: given a protein’s amino acid sequence, predict the three-dimensional structure it folds into.4 Rosetta is a software suite of multiple algorithms, including de novo structure prediction algorithms.5,6 When generating structures with Rosetta AbinitioRelax, a structure is predicted from its sequence alone without the use of homologous structures. This method is more reliable for small proteins (sequence length < 150 amino acids).
Rosetta’s strategy is physics-informed but strongly guided by fragment assembly, knowledge-based score functions, and Monte Carlo sampling. Short structural fragments of several residues, usually 3-mer and 9-mer segments, are selected based on known deposited structures in the PDB and repeatedly inserted to explore conformational space. These fragment combinations are refined and relaxed to identify low energy folds. In this tutorial, we will walk through generating fragment files, preparing the necessary inputs, and running Rosetta to produce de novo structural models.
Note: Files for this tutorial are located in the folder labeled ‘1_1_ab_initio/’. This tutorial is based on the ab initio structure prediction tutorial provided in the official RosettaCommons documentation: https://docs.rosettacommons.org/docs/latest/application_documentation/structure_prediction/abinitio.
For this tutorial, we will generate de novo structures for the B1 domain of streptococcal Protein G (GB1) using Rosetta’s AbinitioRelax. This protein has an experimental structure (1PGA), but sequences for unknown structures can be obtained from the UniProt website (https://www.uniprot.org/).7,8 Create a directory called ‘1_1_ab_initio/native/’ and download the PDB and FASTA for 1PGA and move/rename them to ‘1_1_abinitio/native/1pga_A.fasta’ and ‘1_1_ab_initio/native/1pga_A.pdb’. The following are steps to generate ab initio models for GB1:
- Generate 3-mer and 9-mer fragments.
- 3-mer and 9-mer fragment files are required for structure generation, and can be acquired either by using the Robetta server at https://robetta.bakerlab.org, or by installing additional software and using a combination of PSIPred and Rosetta’s fragment picking algorithm (https://docs.rosettacommons.org/demos/latest/public/fragment_picking/README).9 In this tutorial, we will use the webserver.
- Access the Robetta fragment generation server at https://robetta.bakerlab.org, and create an account.
- Login, and click ‘Structure Prediction’ -> ‘Submit’, and then enter a ‘Target Name’, and either enter or upload the sequence for GB1. Under ‘Optional’, select ‘AB’ for ab initio, and submit the job.
- After the jobs are finished, download the fragment files labeled ‘rb_11_26_695188_686015_ab_t000__robetta.200.3mers’ and ‘rb_11_26_695188_686015_ab_t000__robetta.200.9mers’. Create a directory called ‘fragments/’ and move the files into the directory. The server can be used to generate ab initio models, but we are simply using it as a vehicle to get fragments. The server is not recommended for large volumes of ab initio structures, which is necessary for proper conformational sampling.
- Rename the files to ‘1pga_frags3’ and ‘1pga_frags9’, respectively.
- Prepare input files for AbinitioRelax. Example files are located in ‘1_1_ab_initio/ab_initio_files/’.
- Create and enter a directory called ‘ab_initio_files/’. For Rosetta jobs that have a lot of options or ‘flags’, flag files can be created which list options for the application. Create a flags file called ‘1pga_abinitio_flags’ and input the following, making sure to keep spacing consistent:
-abinitio
-relax
-in
-file
-fasta ~/1_1_ab_initio/native/1pga_A.fasta
-frag3 ~/1_1_ab_initio/fragments/1pga_frags3
-frag9 ~/1_1_ab_initio/fragments/1pga_frags9
-native ~/1_1_ab_initio/native/1pga_A.pdb
-out
-pdb
-file
-silent
-nstruct 10
- The ‘-abinitio’ and ‘-relax’ flags specify to AbinitioRelax to run ab initio modeling, and a short relax.
- The ‘-in’ section specifies the FASTA sequence, the 3mer/9mer fragment files, and the crystal structure for an RMSD comparison. ‘-native’ is an optional flag, and we only specify it so that we could compare our models to the native structure.
- The ‘-out’ flags specify that PDBs should be generated as well as a score file called ‘score.fsc’ which contains Rosetta scores and RMSDs for each of the models
- ‘-nstruct’ indicates the number of models to generate.
- Run this command in the terminal to run AbinitioRelax.
> ~/rosetta/source/bin/AbinitioRelax.linuxgccrelease -database ~/rosetta/database @1pga_abinitio_flags.
- This job should take < 20 minutes if no errors occur. We only generate 10 structures for this tutorial, but generally you should run thousands to tens of thousands of models for a real run to get a proper sampling of the conformational space.
- The score file ‘score.fsc’ is formatted in the following way:
- Rosetta’s score function comprises several terms, which are listed in this score file. The total score listed under the score column labeled ‘score’ represents the combined score of all these terms.
- Since ‘-in:file:native’ was specified (here this is shown using colon-separated namespaces, but above it is given in the option file as hierarchical options, they are equivalent ways of specifying the native file), an RMSD is calculated of the ab initio structure to the native and is listed under ‘rms’.
- The column named ‘description’ indicates the name of the ab initio structure. In this example, I’m only showing one exmaple. Since 10 structures were generated (specified with ‘-nstruct 10’), there would be separate scores, RMSDs, and descriptions for each.
SCORE: score fa_atr fa_rep fa_sol fa_intra_rep fa_intra_sol_xover4 lk_ball_wtd fa_elec pro_close hbond_sr_bb hbond_lr_bb hbond_bb_sc hbond_sc dslf_fa13 omega fa_dun p_aa_pp yhh_planarity ref rama_prepro Filter_Stage2_aBefore Filter_Stage2_bQuarter Filter_Stage2_cHalf Filter_Stage2_dEnd co rms maxsub clashes_total clashes_bb time description
SCORE: −114.296 −262.539 30.027 181.042 0.593 11.059 −7.285 −90.571 0.000 −19.080 −13.884 −6.606 −5.067 0.000 4.672 60.979 −12.559 0.014 15.660 −0.752 0.000 0.000 0.000 0.000 15.898 12.978 37.000 0.000 0.000 24.000 S_00000001
- The best-scoring or lowest-scoring structure is typically viewed as the most energetically favorable and should be treated as the predicted structure (unless rescoring, see Section 2).
- Common errors which users should be aware of include:
- Incorrect Rosetta and file paths, incorrectly generated fragment files (using a different method which has a different format than what is expected, wrong sequence), and incorrect formatting of the options file.
1.2. Rosetta: Comparative/Homology Modeling
If a target sequence has homologous proteins with experimental structures, comparative or homology modeling is a better technique for structure prediction than ab initio modeling. In comparative modeling, homologous proteins serve as templates onto which the target sequence is aligned, allowing Rosetta to infer the overall fold of the unknown structure from evolutionarily related proteins.
In Rosetta’s homology modeling pipeline, a single or multiple templates can be incorporated to capture complementary structural information. After aligning sequences, Rosetta builds an initial hybrid model by combining template segments and filling in gaps, such as insertions or unresolved regions using fragment-based sampling.10 Resulting models are then refined through iterative cycles of side-chain packing and all-atom minimization. The output is a set of energetically favorable structural models that reflect both homologous templates and Rosetta’s scoring function. In this tutorial, we will explore how to obtain structural templates, perform sequence alignments, and generate structures with homology modeling.
Note: Files for this tutorial are located in the folder labeled ‘1_2_homology’. This tutorial is based on the comparative modeling structure prediction tutorial provided in the official RosettaCommons documentation: https://docs.rosettacommons.org/docs/latest/application_documentation/structure_prediction/RosettaCM.
For this tutorial, we will generate models for the B1 domain of streptococcal Protein G (GB1) using homology modeling with RosettaScripts, an xml-scriptable interface for conducting custom Rosetta protocols. This protein has a crystal structure (1PGA), but sequences for proteins without high-resolution experimental structures can be obtained from the UniProt website (https://www.uniprot.org/). Create a directory called ‘1_2_homology_modeling/native/’ and download the PDB and FASTA for 1PGA and move/rename them to ‘1_2_homology_modeling/native/1pga_A.fasta’ and ‘1_2_homology_modeling/native/1pga_A.pdb’. The following are steps to generate homology models for GB1:
- Generate 3-mer and 9-mer fragments.
- 3-mer and 9-mer fragment files are required for structure generation, and can be acquired either by using the Robetta server at https://robetta.bakerlab.org, or by installing additional software and using a combination of PSIPred and Rosetta’s fragment picking algorithm (Instructions are available at https://docs.rosettacommons.org/demos/latest/public/fragment_picking/README). In this tutorial, we will use the webserver.
- Access the Robetta fragment generation server at https://robetta.bakerlab.org, and create an account.
- Login, and click ‘Structure Prediction’ -> ‘Submit’, and then enter a ‘Target Name’, and either enter or upload the sequence for GB1. Under ‘Optional’, select ‘AB’ for ab initio, and submit the job.
- After the jobs are finished, download the fragment files labeled ‘rb_11_26_695188_686015_ab_t000__robetta.200.3mers’ and ‘rb_11_26_695188_686015_ab_t000__robetta.200.9mers’. Create a directory called ‘fragments/’ and move the files into the directory.
- Rename the files to ‘1pga_frags3’ and ‘1pga_frags9’, respectively.
- Selection of structural templates.
- PDBs of homologous proteins to your target sequence can be identified using the National Center for Biotechnology Information (NCBI) website, https://blast.ncbi.nlm.nih.gov/Blast.cgi, via Protein Blast.11
- Navigate to the website and click ‘Protein BLAST’. This will open a new window where you can specify your query settings.
- Under ‘Enter Query Sequence’, enter your target FASTA sequence.
- If you do not know the exact sequence of your protein, FASTA sequences are available for proteins with no experimental structures at https://www.uniprot.org/.
- Give the job a title, and under ‘Choose Search Set’ and ‘Database’, click the dropdown menu and change the database to ‘Protein Data Bank proteins (pdb)’
- Click ‘BLAST’ to run the job. The search can take up to several minutes for larger proteins.
- The program returns a list of proteins ranked by those with the greatest match to the query sequence. Select a template structure with the highest query coverage and percent identity. Since this protein has several PDB entries, we use PDB ID 6KMC_A as a more realistic template, which has a coverage of 98% and percent identity of 85%. Multiple templates can be used, but for this tutorial we only use one.
- Create a directory called ‘template’ and download the PDB file(s) for the selected template(s) and prepare them. Name the PDB ‘6kmc_A.pdb’.
- Generating alignments of structural templates.
- Use Clustal Omega to generate sequence alignments for the template(s) and target sequence. Navigate to the Clustal Omega webpage, https://www.ebi.ac.uk/jdispatcher/msa/clustalo, and paste the contents of both the target and template FASTAs in the ‘Input sequence’ box.12
- The alignments, by default, will be in the ‘ClustalW’ format:
CLUSTAL O(1.2.4) multiple sequence alignment
1pga_A
-MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE
56
6kmc_A
MDTYKLILNGKTLKGETTTEAVDAAHAEKVFKHYANEHGVHGHWTYDPETKTFTVTE
57
*********************** ******:***::**.*.**** ********
- These need to be modified to be in the Grishin format, which is expected by Rosetta. To reformat into the Grishin format, the first line should start with ‘##” followed by the target name (‘1pga_A’) and the template name (‘6kmc_A’) from the alignment. The next line should be a single ‘#’. The next line after that should just be ‘scores_from_program: 0’, where the score does not matter for the subsequent steps, so it is set to 0. The last lines should correspond to the sequence alignments from the Clustal file in the order of target and template, where each line starts with 0.
## 1pga_A 6kmc_A
#
scores_from_program: 0
0 -MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE
0 MDTYKLILNGKTLKGETTTEAVDAAHAEKVFKHYANEHGVHGHWTYDPETKTFTVTE
- Create a directory called ‘alignment/’ and copy the Clustal alignments into a new file called ‘6kmc.grishin’ and modify the contents to match the expected format.
- Map the target sequence onto the template PDB with Rosetta’s partial threading application. This ‘threading’ application maps a target amino acid sequence onto an existing protein structure template, replacing the template residues with those in the target sequence. Here is the command to run this application
> ~/rosetta/source/bin/partial_thread.linuxgccrelease -in:file:fasta
~/1_2_homology/native/1pga_A.fasta -in:file:alignment
~/1_2_homology/alignment/6kmc.grishin -in:file:template_pdb
~/1_2_homology/template/6kmc_A.pdb -ignore_unrecognized_res
- This will generate a PDB file, ‘6kmc_A.pdb.pdb’. Rename this file to ‘target_on_6kmc_A.pdb’.
- Use RosettaScripts to generate homology models.
- Create and enter a directory called ‘homology_modeling/’.
- Create an options file, ‘1pga.options’, which contains all the flags to be used by the program:
-in:file:fasta ~/1_2_homology/native/1pga_A.fasta
-parser:protocol homology_modeling.xml
-relax:minimize_bond_angles
-relax:minimize_bond_lengths
-relax:jump_move true
-default_max_cycles 200
-relax:min_type lbfgs_armijo_nonmonotone
-relax:jump_move true
-use_bicubic_interpolation
-hybridize:stage1_probability 1.0
- Create an XML script formatted like the following, making sure to change the relevant paths to your files. The files ‘stage1.wts’ and ‘stage2.wts’ are located in ‘~/1_2-homology/homology_modeling/’:
<ROSETTASCRIPTS>
<TASKOPERATIONS>
</TASKOPERATIONS>
<SCOREFXNS>
<ScoreFunction name=“stage1”
weights=“~1_2_homology/homology_modeling/stage1.wts”
symmetric=“0”>
<Reweight scoretype=“atom_pair_constraint”
weight=“1”/>
</ScoreFunction>
<ScoreFunction name=“stage2”
weights=“~1_2_homology/homology_modeling/stage2.wts”
symmetric=“0”>
<Reweight scoretype=“atom_pair_constraint”
weight=“0.5”/>
</ScoreFunction>
<ScoreFunction name=“fullatom” weights=“ref2015_cart”
symmetric=“0”>
<Reweight scoretype=“atom_pair_constraint”
weight=“0.5”/>
</ScoreFunction>
</SCOREFXNS>
<FILTERS>
</FILTERS>
<MOVERS>
<Hybridize name=“hybridize” stage1_scorefxn=“stage1”
stage2_scorefxn=“stage2” fa_scorefxn=“fullatom” batch=“1”
stage1_increase_cycles=“1” stage2_increase_cycles=“1”
linmin_only=“1”>
<Fragments three_mers=“~1_2_homology/fragments/1pga_frags3”
nine_mers=“~1_2_homology/fragments/1pga_frags9”/>
<Template pdb=“~1_2_homology/alignment/target_on_6kmc_A.pdb”
cst_file=“AUTO” weight=“1.000”/>
</Hybridize>
</MOVERS>
<APPLY_TO_POSE>
</APPLY_TO_POSE>
<PROTOCOLS>
<Add mover=“hybridize”/>
</PROTOCOLS>
</ROSETTASCRIPTS>
- Run the homology modeling protocol:
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease
@~/1_2_homology/homology_modeling/1pga.options -nstruct 10
- This will generate 10 structures and a score file called ‘score.sc’. These files can be relaxed with this command:
- We give an example of only one relaxed structure, but this should be done for all generated structures
> ~/rosetta/source/bin/relax.linux.gccrelease -in:file:s S_0001.pdb -nstruct 1
- The relaxed structures can be scored using the following command:
- ‘-in:file:native’ specifies a reference structure for RMSDs, but can be ignored if one is not available.
- The final scores will be in ‘1pga_hm_0001_0001_scores.sc’.
> ~/rosetta/source/bin/score.default.linuxgccrelease -s S_0001_0001.pdb -in:file:native ~/1_2_homology/native/1pga_A.pdb -out:file:scorefile
1pga_hm_0001_0001_scores.sc
- The output ‘1pga_hm_0001_0001_scores.sc’ score file will have this format:
- Rosetta’s score function comprises several terms, which are listed in this score file. The total score listed under the score column labeled ‘score’ represents the combined score of all these terms.
- Since ‘-in:file:native’ was specified, an RMSD is calculated of the relaxed structure to the native and is listed under ‘rms’.
- The column named ‘description’ indicates the name of the relaxed structure. If N relaxed structures were generated (specified with ‘-nstruct N’), there would be separate scores, RMSDs, and descriptions for each.
SCORE: score fa_atr fa_rep fa_sol fa_intra_rep fa_intra_sol_xover4 lk_ball_wtd fa_elec pro_close hbond_sr_bb hbond_lr_bb hbond_bb_sc hbond_sc dslf_fa13 omega fa_dun p_aa_pp yhh_planarity ref rama_prepro allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −203.286 −298.931 28.311 189.429 0.622 9.975 −5.540 −124.523 0.000 −16.673 −25.522 −8.452 −6.530 0.000 0.828 61.874 −17.116 0.170 15.660 −6.868 1.283 0.982 0.911 1.000 1.000 1.000 1.000 0.000 56.000 56.000 0.698 S_0001_0001_0001
- Common errors which users should be aware of include:
- Incorrect Rosetta and file paths, incorrectly generated fragment files (using a different method which has a different format than what is expected, incorrect sequence), not formatting alignments into the Grishin format, and incorrect formatting of the options file.
1.3. Rosetta: Protein-Protein Docking with RosettaDock
Protein complexes can be generated with Rosetta’s protein-protein docking protocol, RosettaDock.13 RosettaDock predicts a protein complex by starting with two partners of either monomeric or subcomplex units. Unlike ab initio modeling or homology modeling, docking assumes that each subunit is already folded and focuses on predicting how the partners interact.
RosettaDock operates in two main phases: a low-resolution (centroid) docking and a high-resolution (full-atom) refinement. In the first stage, Rosetta performs coarse sampling of rigid-body orientations and translations between the docking partners. The residues in the proteins are represented in reduced-atom or centroid form, which allows for large and rapid movements to explore larger portions of the conformational landscape. Monte Carlo sampling identifies promising complexes guided by an interface score. Promising poses are refined in the next stage with full atom representation using side-chain repacking, interface optimization, and rigid-body minimization. The result of both stages is a set of docked complexes representing potential binding modes. In this tutorial, we will explore how to generate protein complexes using protein-protein docking with RosettaDock.
Note: Files for this tutorial are located in the folder labeled ‘1_3_rosettadock/’. This tutorial is based on the protein-protein docking protocol tutorial provided in the official RosettaCommons documentation: https://docs.rosettacommons.org/docs/latest/application_documentation/docking/docking-protocol
For this tutorial, we will be trying to predict the Actin/GS1 heterodimer by redocking the A and G chain from a crystal structure with PDB ID 1YAG.14 If structures of subunits/subcomplexes are not known, these can be generated using any of the other techniques discussed in this section (1.1, 1.2, 1.5). This tutorial will also generate backbone ensembles for additional backbone sampling, but this step is not necessary. Backbone perturbations enable a larger sampling of potential binding interfaces, since RosettaDock otherwise treats docking partners as rigid objects. By iteratively sampling perturbed backbone conformations for each subunit, ensemble docking increases conformational diversity and can improve the identification of better-scoring interfaces. Create a directory called ‘1_3_rosettadock/native/’ and download the PDB and FASTA for 1YAG and move/rename them to ‘1_3_rosettadock/native/1YAG_AG.fasta’ and ‘1_3_rosettadock/native/1YAG_AG.pdb’. The following are steps to generate docked structures for Actin/GS1 heterodimer:
- Prepare a PDB file containing generated subunit/subcomplex models.
- RosettaDock is not meant for proper global docking, where a larger conformational space is explored since the location of the interface is not known. RosettaDock is closer to local docking, so some prior knowledge of the interface is necessary. A complex of a homologous protein, or an AlphaFold generated complex could serve as a starting structure. For this tutorial, we will be using the 1YAG PDB.
- Generation of perturbed backbones for each docking partner. Four types of perturbation will be used: relax, normal mode analysis (nma), sheer, and backrub.15,16 For this tutorial, only one of each perturbation type will be generated. It is recommended to do more, with at least five. The following steps should be performed for each docking partner, which in this case is chain ‘A’ and chain ‘G’. Separate the ‘1YAG_AG.pdb’ into two PDBs, each containing either chain ‘A’ (‘1YAG_A.pdb’) and chain ‘G’ (‘1YAG_G.pdb’). Create these directories:
‘1_3_rosettadock/perturbations/backrub/’,
‘1_3_rosettadock/perturbations/relax/’, ‘1_3_rosettadock/perturbations/nma/’, and ‘1_3_rosettadock/perturbations/shear/’.
- Backrub:
- Inside of ‘1_3_rosettadock/perturbations/backrub/’, run this command for each chain:
> ~/rosetta/source/bin/backrub.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_A.pdb -nstruct 1 -backrub:mc_kt 0.6 -backrub:ntrials 200 -out:prefix backrub_
> ~/rosetta/source/bin/backrub.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_G.pdb -nstruct 1 -backrub:mc_kt 0.6 -backrub:ntrials 200 -out:prefix backrub_
- This will generate several PDBs from different stages of the protocol: ‘backrub_1YAG_A_0001_last.pdb’, ‘backrub_1YAG_A_0001_low.pdb’, and ‘backrub_1YAG_A_0001.pdb’ for chain ‘A’ and ‘backrub_1YAG_G_0001_last.pdb’, ‘backrub_1YAG_G_0001_low.pdb’, and ‘backrub_1YAG_G_0001.pdb’ for chain ‘G’. The final structure which should be used is ‘backrub_1YAG_A_0001.pdb’ and ‘backrub_1YAG_G_0001.pdb’. A score file called ‘backrub_score.sc’ is also produced.
- Relax:
- Inside of ‘1_3_rosettadock/perturbations/relax/’, run this command for each chain:
> ~/rosetta/source/bin/relax.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_A.pdb -relax:quick -nstruct 1 -out:prefix relax_
> ~/rosetta/source/bin/relax.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_G.pdb -relax:quick -nstruct 1 -out:prefix relax_
- This will generate a PDB ‘relax_1YAG_A_0001.pdb’, ‘relax_1YAG_G_0001.pdb’ and a score file ‘relax_score.sc’.
- NMA:
- NMA is performed using RosettaScripts (as described in the previous section). Here is the full XML:
<ROSETTASCRIPTS>
<SCOREFXNS>
<ScoreFunction name=“r15_cart” weights=“ref2015_cart.wts” />
</SCOREFXNS>
<RESIDUE_SELECTORS>
</RESIDUE_SELECTORS>
<TASKOPERATIONS>
</TASKOPERATIONS>
<FILTERS>
</FILTERS>
<MOVERS>
<NormalModeRelax name=“nma” cartesian=“true” centroid=“false”
scorefxn=“r15_cart”
nmodes=“5” mix_modes=“true” pertscale=“1.0” randomselect=“false”
relaxmode=“relax”
nsample=“1” cartesian_minimize=“false”
/>
</MOVERS>
<APPLY_TO_POSE>
</APPLY_TO_POSE>
<PROTOCOLS>
<Add mover=“nma” />
</PROTOCOLS>
<OUTPUT scorefxn=“r15_cart” />
</ROSETTASCRIPTS>
- The XML file can be used without any changes. Inside of ‘1_3_rosettadock/perturbations/nma/’, run these commands for each chain:
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_A.pdb -nstruct 1 -parser:protocol
~/1_3_rosettadock/perturbations/xmls/nma.xml -out:prefix nma_
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_G.pdb -nstruct 1 -parser:protocol
~/1_3_rosettadock/perturbations/xmls/nma.xml -out:prefix nma_
- This will generate a PDB ‘nma_1YAG_A_0001.pdb’, ‘nma_1YAG_G_0001.pdb’ and a score file ‘nma_score.sc’.
- Shear:
- Shear is performed using RosettaScripts. Here is the full XML for the protocol:
<ROSETTASCRIPTS>
<SCOREFXNS>
<ScoreFunction name=“r15_cart” weights=“ref2015_cart.wts” />
</SCOREFXNS>
<RESIDUE_SELECTORS>
</RESIDUE_SELECTORS>
<TASKOPERATIONS>
</TASKOPERATIONS>
<FILTERS>
</FILTERS>
<MOVERS>
<Shear name=“full_shear” scorefxn=“r15_cart” temperature=“0.2” nmoves=“10”
angle_max=“360.0” preserve_detailed_balance=“0”/>
</MOVERS>
<APPLY_TO_POSE>
</APPLY_TO_POSE>
<PROTOCOLS>
<Add mover=“full_shear” />
</PROTOCOLS>
<OUTPUT scorefxn=“r15_cart” />
</ROSETTASCRIPTS>
- The XML file can be used without any changes. Inside of ‘1_3_rosettadock/perturbations/shear/’, run these commands for each chain:
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_A.pdb -nstruct 1 -parser:protocol
~/1_3_rosettadock/perturbations/xmls/shear.xml -out:prefix shear_
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_G.pdb -nstruct 1 -parser:protocol
~/1_3_rosettadock/perturbations/xmls/shear.xml -out:prefix shear_
- This will generate a PDB ‘shear_1YAG_A_0001.pdb’, ‘shear_1YAG_G_0001.pdb’ and a score file ‘shear_score.sc’.
- Create a directory called ‘ensembles/’ and prepare ensemble files for each docking partner, containing the paths to all of the perturbed structures. Call them ‘ensemble_list1’ and ‘ensemble_list2’ which correspond to partner 1 and 2, respectively. Here is an example for ‘ensemble_list1’ for 1YAG_A:
~/1_3_rosettadock/perturbations/backrub/backrub_1YAG_A_0001.pdb
~/1_3_rosettadock/perturbations/relax/relax_1YAG_A_0001.pdb
~/1_3_rosettadock/perturbations/nma/nma_1YAG_A_0001.pdb
~/1_3_rosettadock/perturbations/shear/shear_1YAG_A_0001.pdb
- Prepack the perturbed backbone ensembles. Repacking/pre-packing optimizes the sidechain rotamers of the protein partners before docking, typically while keeping the backbone fixed and reducing steric clashes. This step is recommended because individual subunit structures often contain side-chain conformations that are incompatible at the docking interface, which can cause the scoring function to incorrectly penalize otherwise reasonable docking poses. Create and enter a directory called ‘1_3_rosettadock/prepacking/’
- Prior to docking, it is recommended to repack the sidechains, to allow for better interface optimization. The following command prepacks both ensembles:
> ~/rosetta/source/bin/docking_prepack_protocol.linuxgccrelease -in:file:s
~/1_3_rosettadock/native/1YAG_AG.pdb -nstruct 1 -partners A_G -ensemble1
~/1_3_rosettadock/ensembles/ensemble_list1 -ensemble2
~/1_3_rosettadock/ensembles/ensemble_list2 -ex1 -ex2aro -out:prefix ppk_ensemble_
- ‘-in:file:s’ is used to give the structure which represents the starting orientation of the docking partners. As mentioned earlier, we are using the 1YAG crystal structure for simplicity, but this starting structure can be manually created or generated using a method such as AlphaFold.
- ‘-ensemble1’ and ‘-ensemble2’ specify the paths to the ensembles for each docking partner.
- ‘-partners’ indicates which chain(s) correspond to each docking partner. For multiple chains, they would be given as ‘AB_CD’. Make sure the partner 1 corresponds to ensemble 1 and likewise for partner 2.
- ‘-ex1’ and ‘-ex2aro’ are additional recommended flags which add additional sidechain rotamers.
- This will produce several PDBs, including temporary PDBs for each partner, but the most important PDB, used as the input for the next stage, is the prepacked input structure ‘ppk_ensemble_1YAG_AG_0001.pdb’. This structure contains information derived from all the perturbed docking partner structures. There is also a score file called ‘ppk_ensemble_score.sc’. This also results in updated ensemble files, which will be named ‘ensemble_list1.ensemble’ and ‘ensemble_list2.ensemble’, which are located in ‘~/1_3_rosettadock/ensembles/’. These ensembles should be used in the next stage.
- Docking the perturbed ensembles. Create and enter a directory called ‘1_3_rosettadock/docking/’.
- Several different flags can be specified that result in using different parts of the docking protocol. A full list of flags can be found in the Rosetta documentation. Some that are most commonly used: ‘-randomize1’ and/or ‘-randomize2’ which randomize the orientation of the docking partner 1 and/or 2, ‘-spin’ which spins a second docking partner around the center of mass of the first partner, and ‘-ex1’ and ‘-ex2aro’ which add additional sidechain rotamers.
- Using the prepacked input structure and ensembles, run the following command to dock 10 structures:
~/rosetta/source/bin/docking_protocol.linuxgccrelease -in:file:s
~/1_3_rosettadock/prepacking/ppk_ensemble_1YAG_AG_0001.pdb -partners A_G -nstruct 10 -ensemble1 ~/1_3_rosettadock/ensembles/ensemble_list1.ensemble -ensemble2 ~/1_3_rosettadock/ensembles/ensemble_list2.ensemble -randomize1 -ex1 -ex2aro -out:prefix dock_
- After it finishes, there will be 10 docked PDBs and a score file called ‘dock_score.sc’. As with all score-based Rosetta applications, these methods are meant to generate thousands to tens of thousands of models, where the lowest-scoring is viewed as the best prediction.
- Common errors which users should be aware of include:
- Incorrect Rosetta and file paths, syntax errors in option files and XML scripts, or incorrect chain names. A common error specific to RosettaDock is encountered when performing ensemble docking and the number of residues differs between the initial structure and the perturbed structures. This can occur if the perturbed structures are processed in some way. Thus, the user will be faced with an error due to mismatch of residues.
1.4. Rosetta: Protein-Protein Docking with SymDock
SymDock extends Rosetta’s docking framework to protein assemblies with defined symmetry, such as homo-dimers, homo-trimers, homo-tetramers, etc.17 Biological systems often have proteins with large, stable symmetric structures, and SymDock utilizes symmetry constraints to efficiently sample and refine these multimeric interactions. In symmetric docking, Rosetta simultaneously optimizes all subunits under specified symmetry operations. Instead of treating each partner independently, like with RosettaDock, SymDock models the entire oligomeric state by moving a single asymmetric unit and propagating those motions across all symmetry-related subunits. This approach dramatically reduces search space and enforces symmetric interfaces. The SymDock protocol mirrors the two stages in RosettaDock’s protocol: a low-resolution symmetric sampling stage and a high-resolution symmetric refinement stage. In this tutorial, we will outline how to symmetrically dock protein complexes using SymDock.
Note: Files for this tutorial are located in the folder labeled ‘1_4_symdock’. This tutorial is based on the symmetric protein-protein docking protocol tutorial provided in the official RosettaCommons documentation: https://docs.rosettacommons.org/docs/latest/application_documentation/docking/sym-dock.
As with RosettaDock, SymDock requires knowing how many subunits are present in the complex. Generally, it is also useful to know the symmetry, but if that is unknown, there are limited possible symmetries for a certain number of subunits.18 For this tutorial, we will be predicting the β2-microglobuin homodimer by performing symmetric docking with SymDock using the A chain from the PDB 2F8O.19 If structures of subunits/subcomplexes are not known, these can be generated using any of the other techniques discussed in this section (1.1, 1.2, and 1.5). This tutorial will also generate backbone ensembles for additional backbone sampling, but this step is not necessary. Create a directory called ‘1_4_symdock/native/’. The following are steps to generate symmetrically docked structures for the β2-microglobuin homodimer:
- Download the 2F8O PDB from https://www.rcsb.org/structure/2F8O, delete all contents except for chain A and save it as ‘~/1_4_symdock/native/2F8O_A.pdb’.
- SymDock takes a single subunit and generates a complex by copying that file and placing subunits into the specified symmetry. As a result, only a single input chain is needed.
- Generation of perturbed backbones for the docking subunit. Four types of perturbation will be used: relax, normal mode analysis (nma), sheer, and backrub. For this tutorial, only one perturbation of type will be generated. It is recommended to do more, with at least five. Create these directories: ‘1_4_symdock/perturbations/backrub/’, ‘1_4_symdock/perturbations/relax/’, ‘1_4_symdock/perturbations/nma/’, and ‘1_4_symdock/perturbations/shear/’.
- Backrub:
- Inside of ‘1_4_symdock/perturbations/backrub/’, run this command for each chain:
> ~/rosetta/source/bin/backrub.linuxgccrelease -in:file:s
~/1_4_symdock/native/2F8O_A.pdb -nstruct 1 -backrub:mc_kt 0.6
-backrub:ntrials 200 -out:prefix backrub_
- This will generate several PDBs, where ‘backrub_2F8O_A_0001.pdb’ is the one to use, and score file ‘backrub_score.sc’.
- Relax:
- Inside of ‘1_4_symdock/perturbations/relax/’, run this command for each chain:
> ~/rosetta/source/bin/relax.linuxgccrelease -in:file:s
~/1_4_symdock/native/2F8O_A.pdb -relax:quick -nstruct 1 -out:prefix relax_
- This will generate a PDB ‘relax_2F8O_A_0001.pdb’ and a score file ‘relax_score.sc’.
- NMA:
- NMA is performed using RosettaScripts (as described in the previous section). This XML is shown in the previous section, but the file is located in ‘~/1_4_symdock/perturbations/xmls/nma.xml.
- The XML file can be used without any changes. Inside of ‘1_4_symdock/perturbations/nma/’, run these commands for each chain:
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_4_symdock/native/2F8O_A.pdb -nstruct 1 -parser:protocol
~/1_4_symdock/perturbations/xmls/nma.xml -out:prefix nma_
- This will generate a PDB ‘nma_2F8O_A_0001.pdb’ and a score file ‘nma_score.sc’.
- Shear:
- Shear is performed using RosettaScripts. his XML is shown in the previous section, but the file is located in ‘~/1_4_symdock/perturbations/xmls/shear.xml’.
- The XML file can be used without any changes. Inside of ‘1_4_symdock/perturbations/shear/’, run these commands for each chain:
> ~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -in:file:s
~/1_4_symdock/native/2F8O_A.pdb -nstruct 1 -parser:protocol
~/1_4_symdock/perturbations/xmls/shear.xml -out:prefix shear_
- This will generate a PDB ‘shear_2F8O_A_0001.pdb’ and a score file ‘shear_score.sc’.
- Unlike for RosettaDock, where you generate ensemble files for both partners, only a single ensemble list needs to be generated. Create and enter a directory called ‘1_4_symdock/ensemble/’. Create a file called ‘ensemble_list1’ with the following contents:
~/1_4_symdock/perturbations/backrub/backrub_2F8O_A_0001.pdb
~/1_4_symdock/perturbations/relax/relax_2F8O_A_0001.pdb
~/1_4_symdock/perturbations/nma/nma_2F8O_A_0001.pdb
~/1_4_symdock/perturbations/shear/shear_2F8O_A_0001.pdb
- Generation of symmetry definition (symdef) file. Example files are located in ‘1_4_symdock/symmetry/’.
- Symmetry definitions can be made either by using a template PDB, or by using ‘make_symmdef_file_denovo.py’. If a homologous structure is known, it is recommended to follow this method: https://docs.rosettacommons.org/docs/latest/application_documentation/utilities/make-symmdef-file. For this tutorial, we will use the de novo script. With this method, SymDock is limited to cyclical and dihedral symmetry.
- The symmetry type for a homodimer is C2, which is also the symmetry of the crystal structure, so we will generate a C2 symmetry file. Create a directory called ‘1_4_symdock/symmetry/’ and run this command to generate C2 symmetry definition:
> ~/rosetta/source/src/apps/public/symmetry/make_symmdef_file_denovo.py -symm_type cn -nsub 2
- Where ‘-symm_type’ is either cyclical (cn) or dihedral (dn)
- ‘-nsub’ is the number of subunits.
- This will produce the following:
symmetry_name c2
subunits 2
recenter
number_of_interfaces 1
E = 2*VRT0001 + 1*(VRT0001:VRT0002)
anchor_residue COM
virtual_transforms_start
start −1,0,0 0,1,0 0,0,0
rot Rz 2
virtual_transforms_stop
connect_virtual JUMP1 VRT0001 VRT0002
set_dof BASEJUMP x(50) angle_x(0:360) angle_y(0:360) angle_z(0:360)
- Save this to a file called ‘C2.symm’.
- Docking the perturbed ensembles. Create a directory called ‘1_4_symdock/docking/’.
- Several different flags can be specified that result in using different parts of the docking protocol. A full list of flags for SymDock can be found in the Rosetta documentation. Run the following command in ‘1_4_symdock/docking/’ which has some the recommended flags:
> ~/rosetta/source/bin/SymDock.linuxgccrelease -in:file:l
~/1_4_symdock/ensemble/ensemble_list1 -in:file:native
~/1_4_symdock/native/2F8O_AB.pdb -symmetry:symmetry_definition
~/1.4_symdock/symmetry/C2.symm -nstruct 10 -
symmetry:initialize_rigid_body_dofs -symmetry:symmetric_rmsd -ex1 -ex2aro
-out:prefix dock_symmetric_ -detect_disulf false -ignore_unrecognized_res
- ‘-in:file:l’ is used to specify a list of perturbed PDBs, instead of the -ensemble flags with RosettaDock
- ‘-in:file:native’ is used for calculating RMSDs if a reference structure is available, and ‘-symmetry:symmetric_rmsd’ specifies to calculate symmetric RMSD, which accounts for differences in chain naming.
- ‘-symmetry:symmetry_definition’ specifies the symmetry file created in step 4.
- ‘-symmetry:initialize_rigid_body_dofs’ initialize the rigid body configuration according to the symmetry file.
- ‘-detect_disulf false’ prevents Rosetta from trying to create disulfides.
- After it finishes, there will be 10 docked PDBs for each perturbation type and a score file called ‘dock_symmetric_score.sc’. As with all score-based Rosetta applications, these methods are meant to generate thousands to tens of thousands of models, where the lowest-scoring is viewed as the best prediction.
- Common errors which users should be aware of include:
- Incorrect Rosetta and file paths, syntax errors in option files and XML scripts, or incorrect chain names. A common error specific to SymDock is the number of subunits is not possible for a given symmetry, which would occur if using the ‘make_symmdef_file_denovo.py’ script.For example, if you give two subunits and specify ‘D2’ symmetry, this will result in an error.
1.5. AlphaFold: Monomer/Multimer Predictions with AlphaFold2
AlphaFold2 is a deep-learning neural network developed by DeepMind that has revolutionized the modern field of protein structure prediction.20 In 2021, AlphaFold2 demonstrated unparalleled improvements in accuracy when compared to other existing methods at the time. Although newer methods have been released, such as AlphaFold3, AlphaFold2 still remains a commonly used method for structure prediction for both monomers and multimers.
AlphaFold2 utilizes co-evolutionary information present in multiple sequence alignments (MSAs) of homologous sequences. Co-evolutionary behavior of residues typically is indicative of close-proximity and interactions. AlphaFold identifies these patterns through a transformer-based network to construct a three-dimensional structure. AlphaFold2, like ab initio modeling, only requires a sequence from which it generates MSAs, and also has a template track for leveraging homologous structures. In this tutorial, we present how to use AlphaFold via the command line for predicting both monomers and multimers.
Note: Files for this tutorial are located in the folder labeled ‘1_5_alphafold’. This tutorial is based on AlphaFold2’s documentation: https://github.com/google-deepmind/alphafold.
The installation of AlphaFold2 can be intimidating, but the official documentation provides rather thorough instructions. If access is possible, many high-performance computing clusters (HPCs) usually have AlphaFold2 already installed. For running AlphaFold2, only a FASTA file is needed for running, although there are many custom options/settings. For this tutorial, it is assumed the user has a working docker installation of AlphaFold2. This section will cover using AlphaFold2.3.2 to generate structures for both the monomer GB1 protein and homodimer Actin/GS1 protein:
- Predicting a structure of GB1 monomer:
- Create and enter the directory ‘1_5_alphafold/fastas/’ and create a fasta called ‘1PGA_A.fasta’. Enter the following in the FASTA file:
>1PGA_A
MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE
- Create and enter the directory ‘1_5_alphafold/output_monomer/’. Inside the directory, run this command:
> python3 docker/run_docker.py \
--fasta_paths=~/1_5_alphafold/fasta/1PGA_A.fasta \
--max_template_date=2025-11-16 \
--model_preset=monomer \
--data_dir=$DOWNLOAD_DIR \
--output_dir=~/1_5_alphafold/output_monomer/
- Where ‘--fasta_paths is the FASTA file. Multiple FASTAs can be run sequentially, they would be listed such as ‘--fasta_paths=PROT1.fasta,PROT2.fasta,PROT3.fasta’.
- ‘--max_template_date’ specifies that AlphaFold2 will filter for only proteins deposited before the designated date.
- ‘--model_preset’ indicates whether to run monomer or multimer AlphaFold2.
- ‘--data_dir’ is the path to all of the AlphaFold databases (downloaded during installation).
- ‘--output_dir’ is the path to where all of the output files are generated and saved.
- A directory called ‘1PGA_A/’ should have been created within ‘output_monomer/’. This directory will have the computed MSAs, unrelaxed structures, relaxed structures, ranked structures, raw model outputs, prediction metadata, and section timings. Five different versions, or models, of AlphaFold2 were originally, trained, each representing different training/validation configurations and hyperparameters. By default, AlphaFold2 generates predictions using each of these models, and then ranks them by network confidence. The following are a breakdown of the most relevant files:
- ‘msas/’ will contain computed MSAs generated with different genetic tools. Note: These files were removed from the tutorial files due to size.
- ‘features.pkl’ is a pickle file containing NumPy arrays with the input and output features which are used by the models.
- ‘unrelaxed_model_{#}_pred_0.pdb’ and ‘relaxed_model_{#}_pred_0.pdb’ are PDBs containing predicted structures before and after a short Amber relaxation. Which of the 5 AlphaFold models the PDB was generated with is indicated by ‘{#}’.
- ‘ranked_{#}.pdb’ are PDBs of the predicted structures renamed according to the models confidence, pLDDT. The top ranked structure, with the highest pLDDT, is ‘ranked_0.pdb’.
- ‘ranking_debug.json’ contains the pLDDT values of each of the structures, as well as the mapping to which original PDB corresponds to each ranked structure.
- ‘relax_metrics.json’ and ‘timings.json’ are JSON files which contain relax metrics and the timing information for the AlphaFold run, respectively.
- Pickle files are generated for each produced structure and are called ‘result_model_{#}_pred_0.pkl’, where ‘{#}’ corresponds to the model number. These contain dictionaries with information such as distograms, per-residue pLDDT scores, predicted TM-score, and predicted pairwise aligned errors. Note: These files were removed from the tutorial files due to size.
- The following are additional options that can be specified:
- ‘--models_to_relax=all’, ‘--models_to_relax=best’, and ‘--models_to_relax=none’ to perform an Amber relax on all the generated structures, the top ranked structure, or none of the structures, respectively. The default is only to relax the best structure.
- GPU relax can be specified by setting ‘--enable_gpu_relax=true’, which is by default, or it can be set to CPU relax with ‘--enable_gpu_relax=false’.
- If you are rerunning AlphaFold with the same output directory, and MSAs have already been computed in ‘msas/’, then MSA generation can be skipped by passing ‘--use_precomputed_msas=true’.
- Predicting a structure of Actin/GS1 heterodimer:
- For multimers, the major difference is specifying the different chains in the FASTA file and choosing the multimer model preset.
- Enter the directory ‘1_5_alphafold/fastas/’ and create a FASTA file called ‘1YAG_AG.fasta’. Enter the following in the FASTA file:
>1YAG_A
MDSEVAALVIDNGSGMCKAGFAGDDAPRAVFPSIVGRPRHQGIMVGMGQKDSYVGDEAQSKRGILTLRYPIEHGIVTNWDDMEKIWHHTFYNELRVAPEEHPVLLTEAPMNPKSNREKMTQIMFETFNVPAFYVSIQAVLSLYSSGRTTGIVLDSGDGVTHVVPIYAGFSLPHAILRIDLAGRDLTDYLMKILSERGYSFSTTAEREIVRDIKEKLCYVALDFEQEMQTAAQSSSIEKSYELPDGQVITIGNERFRAPEALFHPSVLGLESAGIDQTTYNSIMKCDVDVRKELYGNIVMSGGTTMFPGIAERMQKEITALAPSSMKVKIIAPPERKYSVWIGGSILASLTTFQQMWISKQEYDESGPSIVHHKCF
>1YAG_G
MVVEHPEFLKAGKEPGLQIWRVEKFDLVPVPTNLYGDFFTGDAYVILKTVQLRNGNLQYDLHYWLGNECSQDESGAAAIFTVQLDDYLNGRAVQHREVQGFESATFLGYFKSGLKYKKGGVASGF
- Create and enter the directory ‘1_5_alphafold/output_multimer/’. Inside the directory, run this command:
> python3 docker/run_docker.py \
--fasta_paths=~/1_5_alphafold/fasta/1YAG_AG.fasta \
--max_template_date=2025-11-16 \
--model_preset=multimer \
--data_dir=$DOWNLOAD_DIR \
--output_dir=~/1_5_alphafold/output_multimer/
- ‘--model_preset=multimer’ indicates to use the multimer model to generate a complex.
- All the other flags are the same as the monomer flags.
- The output files will be the same as for monomer, and ‘ranked_0.pdb’ is the top structure. The ‘msas/’ folder will be different however. It should contain subdirectories with separate MSAs for each protein chain, as well as a ‘chain_id_map.json’ file which contains the mapping.
- To generate multiple structures for each model using multiple seeds, you can specify ‘--num_multimer_prediction_per_model=N’, where ‘N’ is the number of structures to generate per model. If ‘N=5’, then there would be 5 structures per model for a total of 25 structures.
- Common errors which users should be aware of include:
- Incorrect file paths (for inputs, databases, and outputs), the FASTA file is not formatted correctly, and potential issues resulting from improper installation of the network.
The installation of AlphaFold and its corresponding databases can be challenging, and performance can be substantially impacted if GPU access is not available. An alternative way to generate structures is to use ColabFold (see https://github.com/sokrypton/ColabFold), an AlphaFold implementation using Google Collaboratory with a faster homology search using MMseqs2.3,21 ColabFold can be used through a Google Collaboratory or Jupyter notebook. This notebook can be run locally, or through Google Collaboratory itself. The notebook consists of individual code chunks that are run consecutively. Several chunks contain well-documented options for user input, such as the target FASTA, MSA options, and different model settings. Structures for both 1PGA and 1YAG can be generated with ColabFold by specifying settings in the notebook similar to what we used in the command line. We recommended trying and comparing to the AlphaFold2 predictions. In AlphaFold2, a ‘seed’ is an initialized value used to control stochastic parts of network. Different seeds will result in structural differences between proteins due to this inherent randomness.
Section 2. Rescoring with Rosetta-MS Score Terms and Applications.
Rosetta is an integrative modeling platform for which numerous applications and custom score terms have been developed to incorporate diverse forms of data, including sparse experimental restraints, into structural modeling. In this section, we describe tutorials for a range of Rosetta applications and score terms that leverage data from structural mass spectrometry to enhance Rosetta’s modeling accuracy and scoring capabilities. These methods extract structural characteristics and experimental insights from mass spectrometry (MS) measurements and apply them to both monomeric and complex protein structure prediction.
The approaches presented here fall into two general categories. In some cases, new Rosetta energy methods have been developed, allowing data-derived score terms to be applied automatically during Rosetta’s standard scoring procedures (hrf_ms_labeling, covalent_labeling_fa, hrf_dyanmics, depc_ms, hrf_GBM, and IM_Score). In other cases, self-contained Rosetta applications are used to evaluate pre-generated structural models by calculating novel score terms externally (cl_complex_rescore, SID_Rescore, SID_ERMS_Rescore). Although both approaches yield equivalent results, the choice between implementing a new score term or a standalone application typically reflects considerations of computational efficiency and ease of integration within Rosetta.
For monomeric structure prediction, structural mass spectrometry data from covalent labeling and ion mobility experiments have been used to develop dedicated score terms, along with additional Rosetta structural analysis scripts described in Sections 3.1 and 3.2 for ab initio and homology modeling.22–30 Covalent labeling methods, including hydroxyl radical footprinting (HRF) and diethyl pyrocarbonate (DEPC) labeling, generally rely on comparing residue-level solvent exposure metrics calculated from computational models with experimentally measured labeling rates. Ion mobility mass spectrometry, by contrast, involves comparing experimentally determined collisional cross-section (CCS) values with CCS values calculated from structural models using the Rosetta application Projection Approximation using Rough Circular Shapes (PARCS; see Section 3.3). For protein complex prediction, we describe methods that utilize covalent labeling and surface-induced dissociation (SID) mass spectrometry.31–36 In differential covalent labeling experiments, changes in residue labeling upon complex formation, correlating to interface proximity, are compared against predicted inter-residue or interface distances. For non-differential covalent labeling, when a labeling experiment can be performed only on a complex, a similar approach to monomers is performed. In SID-based approaches, energy-resolved mass spectrometry (ERMS) distributions are used to evaluate complex topology and subunit connectivity.
Final note: all the MS-guided structure prediction algorithms presented here are Rosetta-based. While being an active focus of the Lindert lab, no AlphaFold-based MS guided prediction algorithms have been developed yet.
2.1. Rescoring with covalent labeling score terms and applications: hrf_ms_labeling, covalent_labeling_fa, hrf_dynamics, depc_ms, cl_complex_rescore, hrf_GBM
Note: Files for this tutorial are in the directory labeled ‘2_1_cl_methods/’ . Tutorials for these methods were adapted from the original papers. Additional documentation for hrf_dynamics and depc_ms is located here: https://docs.rosettacommons.org/docs/latest/rosetta_basics/scoring/ms_expdata_score_terms.
Covalent labeling (CL) is a structural mass spectrometry technique which identifies exposed residues on a protein surface. In a CL experiment, proteins are exposed to a labeling reagent in solution, which depending on the reagent, will label residues specifically or non-specifically proportional to their solvent accessibility. Residues which experience higher solvent interactions will typically experience greater labeling. The methods discussed in this section use experimentally obtained CL to identify which residues are more exposed to solvent, which can be used to evaluate agreement with computational structures. The two labeling reagents with tutorials in this section are hydroxyl radical footprinting (HRF) and diethylpyrocarbonate (DEPC).
For this tutorial, each subsection focuses on a different covalent labeling method. Most of these methods are intended for monomeric proteins (hrf_ms_labeling, covalent_labeling_fa, hrf_dynamics, depc_ms, and hrf_GBM); however, cl_complex_rescore is designed for use with protein complexes.22,23,25–27,35 Each tutorial will provide example proteins and experimental data suitable for each method.
2.1.1. hrf_ms_labeling, covalent_labeling_fa
For this tutorial, we will be presenting how to use the hrf_ms_labeling and covalent_labeling_fa score terms. These are both score terms for scoring computationally generated structures with experimental insights derived from covalent labeling experiments. covalent_labeling_fa is an updated version of hrf_ms_labeling, which accounts for full-atom representations of residues and uses a slightly different functional form.26,27 The steps to using these methods are identical, prior to indicating the scoring weights. We will be using reported hydroxyl radical footprinting rate constants reported by Xie et al. for myoglobin, with a native reference PDB of 1DWR.37,38 These rates have been converted into protection factors (PFs), which when expressed on a natural logarithmic scale, have been shown to correlate with the solvent exposure of given residue.39 This method is designed for using PFs, and will not work as intended with raw labeling rates.
Create a directory, ‘2_1_1_hrf_ms_labeling/native/’, and download the PDB of 1DWR from www.rcsb.org, remove non ‘ATOM’ and ‘TER’ records, and then move it to a new directory, and name the file ‘1DWR_A.pdb’.
- Create a directory called ‘2_1_1_hrf_ms_labeling/cl_data/’, enter the directory and create a text file called ‘myoglobin_lnpf.txt’. Inside the text file, we will put the per-residue lnPFs derived from HRF-FPOP data. Paste the following in the text file.
- The first column corresponds to residue numbers, and the second is the lnPFs. The first row is optional header for labeling columns but is ignored by Rosetta.
- Make sure the numbering of residues in the text file corresponds to the numbering of residues in the FASTA file.
#Residue_ID
LNPF
7
4.0943445622
11
4.0989003788
13
3.017009672
14
4.3705979389
18
2.2462321564
21
3.6018680771
24
6.429719478
27
3.5922526184
30
2.9957322736
31
5.5745707432
33
5.6347896032
36
3.3386770247
37
2.9187712324
38
3.1027043931
40
4.1408645779
42
4.1997050779
67
3.9386912525
68
2.3566523143
69
4.605170186
72
4.4003757733
77
4.3640081292
113
2.0989861378
116
2.190793687
119
3.0739844705
120
0.8439700703
123
5.6347896032
136
3.7641028754
137
3.6018680771
138
5.4116460519
139
6.8738537273
147
0.7207987119
149
2.9957322736
151
3.7629874263
Structures for myoglobin can be generated using any of the monomeric structure prediction methods discussed in section 1 (1.1, 1.2, 1.5). For this tutorial, we will simply rescore the 1DWR crystal structure for myoglobin.
- hrf_ms_labeling and covalent_labeling_fa are custom score terms and can be called when running Rosetta’s score application. Create and enter a directory called ‘2_1_1_hrf_ms_labeling/rescoring/’. Run the following command for rescoring with hrf_ms_labeling.
- ‘-in:file:s’ specifies the input PDB.
- ‘-score:weights’ ‘hrf_ms_labeling.wts’ indicates to use the hrf_ms_labeling score term and its weights.
- ‘-in:file:hrf_ms_labeling’ indicates the text file with the per-residue lnPFs.
- ‘-centroid_input’ indicates how to calculate neighbor count in the centroid representation. More information about neighbor count is provided in section 3.1.
- ‘-out:file:scorefile’ sets the name of the output score file. New scores will automatically be appended to score files with the same name, so you do not need to make separate score files if scoring multiple structures.
> ~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database
-in:file:s ../native/1DWR_A.pdb -score:weights hrf_ms_labeling.wts -
in:file:hrf_ms_labeling ../cl_data/myoglobin_lnpf.txt –centroid_input -
out:file:scorefile 1DWR_hrf_ms_labeling_score.sc
- The output ‘1DWR_hrf_ms_labeling_score.sc’ will have the following format if using hrf_ms_labeling.
- The value in the ‘score’ column is the value from the hrf_ms_labeling term. The value in the ‘hrf_ms_labeling’ column is the same as that in the ‘score’ column but is labeled separately to make parsing easier. To obtain a new total score for structures generated with Rosetta, add the hrf_ms_labeling score to the ‘score value from the AbinitioRelax ‘score.fsc’. You could also calculate the normal Rosetta score by rerunning the command without the ‘-score:weights hrf_ms_labeling.wts’ and ‘-in:file_hrf_ms_labeling’ flags and then add the score from hrf_ms_labeling. For rescoring of AlphaFold structures, you could score with and without the hrf_ms_labeling weights and add them, or you could combine with AlphaFold’s confidence.
SCORE: score hrf_ms_labeling allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −84.596 −84.596 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 152.000 152.000 0.000 1DWR_A_0001
- For rescoring with covalent_labeling_fa, run the following command. Note: centroid_input is not specified since we are running Rosetta with full-atom.
- ‘-score:weights covalent_labeling_fa.wts’ indicates to use the covalent_labeling_fa score term.
- The file containing the experimental per-residue lnPFs is specified with a new flag, ‘-score:covalent_labeling_fa_input’. The same file that was used in step 5 is used here. Note: this was changed in the new energy method development.
- ‘-out:file:scorefile’ sets the name of the output score file. New scores will automatically be appended to score files with the same name, so you do not need to make separate score files if scoring multiple structures.
> ~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database
-in:file:s ../native/1DWR_A.pdb -score:weights covalent_labeling_fa.wts -
score:covalent_labeling_fa_input ../cl_data/myoglobin_lnpf.txt -
out:file:scorefile 1DWR_covalent_labeling_fa_rescore_score.sc
- The output ‘1DWR_covalent_labeling_fa_score.sc’ will have the following format if using covalent_labeling_fa.
- Since this method uses the full-atom representation, additional score terms are outputted in the score file. The value in the ‘score’ column is the value from the covalent_labeling_fa term (‘−132.838 in the example). All other values in this file should be ignored. To obtain a new total score for structures generated with Rosetta, add the covalent_labeling_fa score to the ‘score value from the AbinitioRelax ‘score.fsc’. You could also calculate the normal Rosetta score by rerunning the command without the ‘-score:weights covalent_labeling_fa.wts’ and ‘-score:covalent_labeling_fa_input’ flags and then add the score from covalent_labeling_fa. For rescoring of AlphaFold structures, you could score with and without the covalent_labeling_fa weights and add them, or you could combine with AlphaFold’s confidence.
SCORE: score fa_atr fa_rep fa_sol fa_intra_rep fa_intra_sol_xover4 lk_ball_wtd fa_elec pro_close hbond_sr_bb hbond_lr_bb hbond_bb_sc hbond_sc dslf_fa13 omega fa_dun p_aa_pp yhh_planarity ref covalent_labeling_fa rama_prepro allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −132.838 −889.874 122.800 605.285 1.638 35.582 −28.999 −247.955 17.746 −100.916 −3.361 −9.569 −3.877 0.000 14.045 334.765 −15.229 0.000 35.047 0.000 0.033 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 152.000 152.000 0.000 1DWR_A_0001
Typically, rescoring is performed for multiple models, often thousands or tens of thousands. The best scoring structure, i.e. the structure with the lowest Rosetta Energy Units (REU) after rescoring is the structure indicated by the method to best reflect the experimental data, and likely the most native-like.
2.1.2. hrf_dynamics
For this tutorial, we will be presenting how to use the hrf_dynamics score term, as well as employing additional mover models to increase sampling of side-chain flexibility and dynamics. hrf_dynamics is a score term similar to those discussed in 2.1.1, but differs in that residue dynamic effects were accounted for during development.22 In this tutorial, we will be using PFs derived from hydroxyl radical footprinting rate constants reported by Xie et al for myoglobin.38
Create a directory, ‘2_1_2_hrf_dynamics/native/’, download the PDB of 1DWR from www.rcsb.org, remove non ‘ATOM’ and ‘TER’ records, and then move it to new directory labeling it ‘1DWR_A.pdb’.
- Create a directory called ‘2_1_2_hrf_dynamics/cl_data’, enter the directory and create a text file called ‘myoglobin_WYFHL_lnpf.txt’. Inside the text file, we will put the per-residue lnPFs derived from HRF-FPOP data from 2.1.1. It was found from a benchmark of various proteins that a subset of residues with moderate reactivity to hydroxyl radicals, W, F, Y, H, and L, give the best scoring results. It is recommended to only use those residues for this method. Paste the following in the text file.
- The first column corresponds to residue numbers (and of only the W, F, Y, H, and L residue subset), and the second is the lnPFs. The first row is an optional header for labeling columns but is ignored by Rosetta.
- Make sure the numbering of residues in the text file corresponds to the numbering of residues in the FASTA file.
#Residue_ID
LNPF
7
4.0943445622
11
4.0989003788
14
4.3705979389
24
6.429719478
33
5.6347896032
36
3.3386770247
40
4.1408645779
69
4.605170186
72
4.4003757733
116
2.190793687
119
3.0739844705
123
5.6347896032
137
3.6018680771
138
5.4116460519
149
2.9957322736
151
3.7629874263
Structures for myoglobin can be generated using any of the monomeric structure prediction methods discussed in section 1 (1.1, 1.2, 1.5). For this tutorial, we will simply rescore the 1DWR crystal structure.
- Create and enter a directory called ‘2_1_2_hrf_dynamics/rescoring/’. Create a flags file, called ‘1DWR_hrf_dynamics_flags’, which contains the options for the score application. When running a Rosetta application, flags can be contained into a text file. This makes running many commands with identical flags more efficient. Enter these contents into the file:
- ‘-hrf_dynamics_input’ indicates the text file with the CL data.
- ‘-weights’ indicates to use the hrf_dynamics score term weights.
- If the use of a flags file is not preferred, these flags can be specified as -‘score:hrf_dyanmics_input ../cl_data/myoglobin_WYFHL_lnpf.txt’ and ‘-score:weights hrf_dynamics.wts’. The notation is slightly different, but the indentation of ‘-hrf_dynamics_input’ and ‘-weights’ in the flags file under ‘-score’ is equivalent to the command line option, and indicates they are sub-flags of score. This formatting of the flags file is handled by Rosetta’s internal logic.
-score
-hrf_dynamics_input ../cl_data/myoglobin_WYFHL_lnpf.txt
-weights hrf_dynamics.wts
- Run the following command.
- ‘-in:file:s’ specifies the input PDB.
- ‘-out:file:scorefile’ sets the name of the output score file. New scores will automatically be appended to score files with the same name, so you do not need to make separate score files for each model if scoring multiple structures.
- The use of a flags file, or file containing different commands or options for the target application, is indicated to the application with an ‘@’ symbol followed by the name of the flags file.
> ~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database
-in:file:s ../native/1DWR_A.pdb @1DWR_hrf_dynamics_flags -
out:file:scorefile 1dwr_hrf_dynamics_score.sc
- The output ‘1dwr_hrf_dynamics_rescore_score.sc’ will have the following format if using hrf_dynamics.
- The value in the ‘score’ column is the value from the hrf_dynamics term. To obtain a new total score for structures generated with Rosetta, add the hrf_dynamics score to the ‘score’ value from the AbinitioRelax ‘score.fsc’. You could also calculate the normal Rosetta score by rerunning the command without the ‘@1DWR_hrf_dynamics_flags’ flag file and then add the score from hrf_dynamics. For rescoring of AlphaFold structures, you could score with and without the hrf_dynamics weights and add them, or you could combine with AlphaFold’s confidence.
SCORE: score hrf_dynamics allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −188.678 −188.678 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 152.000 152.000 0.000 1DWR_A_0001
- In the hrf_dynamics manuscript, mover models were generated for a subset of the top 20 scoring models (after rescoring) to sample additional side chain dynamics. A combination of Rosetta’s NormalModeRelax and FastRelax movers were used to mimic side chain flexibility sampled with molecular dynamics simulations. For this tutorial, we will use the 1DWR crystal structure, and generate mover models. Create a directory called ‘mover_models’ and copy the 1DWR_A.pdb file to that directory. If generating structures via another method such as AbinitioRelax, a Python script is provided called ‘extract_top_five.py’ which will identify the top 5 lowest scoring models and copy them to your ‘mover_models’ directory. Run this command to generate 5 mover models.
- ‘-s’ tells Rosetta the name of the structure to use for model generation.
- ‘-nstruct’ controls the number of models generated with the mover ensemble. In the original source publication, they recommend 30, but for this example we will do 5.
- ‘-parser:protocol’ flag is used to indicate which flexibility protocol to use.
>~/rosetta/source/bin/rosetta_scripts.linuxgccrelease -s 1DWR_A.pdb -nstruct 5 -parser:protocol
~/rosetta/source/scripts/rosetta_scripts/public/flexibility/nma_relax.xml -overwrite
- After the structures have been generated, rescore the mover models with the score application by running the following command for each output structure.
- ‘-in:file:native’ can be specified if an RMSD to a reference structure is desired.
>~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database -
in:file:s 1DWR_A_0001.pdb -out:file:scorefile
1dwr_mover_models_rosetta_score.sc
- Now repeat, but with hrf_dynamics. Copy the ‘1dwr_hrf_dynamics_flags’ from ‘2_1_2_hrf_dynamics/rescoring’ into the current working directory. For each of the output structures, run the following command to calculate scores with hrf_dynamics.
>~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database -
in:file:s 1DWR_A_0001.pdb @1dwr_hrf_dynamics_flags
1dwr_mover_models_hrf_dynamics_score.sc
- Analyze the results from the score files.
- Extract the ‘score’ values and ‘description’ entries from the ‘1dwr_mover_models_rosetta_score.sc’ file. Description entries are the names of the output files with the ‘.pdb’ file suffix. These entries usually contain a number attached as an additional suffix, such as ‘_0001’ to indicate it is the first of a potential series of structures.
- Extract the ‘hrf_dynamics’ values and ‘description’ entries from the ‘1dwr_mover_models_hrf_dynamics_score.sc’.
- Sum the ‘score and ‘hrf_dynamics’ scores for each model to determine the model’s total score.
Typically, rescoring is performed for multiple models, often thousands or tens of thousands. The best scoring structure, i.e. the structure with the lowest Rosetta Energy Units (REU) after rescoring, considering both the AbinitioRelax models and mover models, is the structure indicated by the method to best reflect the experimental data, and likely the most native-like.
2.1.3. depc_ms
For this tutorial, we will be presenting how to use the depc_ms score term.23 This method provides an additional score term for scoring agreement of computational structures to experimental DEPC covalent labeling data. depc_ms accounts for full-atom representations of residues. This method differs from the previously described methods in this section since it is designed for use with DEPC covalent labeling, whereas the previous sections were meant for use with HRF-based covalent labeling. We will be using DEPC labeling extents reported by Zhou et al. for myoglobin, with a native reference PDB of 1DWR.40 This method is designed for DEPC covalent labeling and will not work as intended with other labels.
Create a directory, ‘2_1_3_depc_ms_labeling/native/’, download the PDB of 1DWR from www.rcsb.org , remove non ‘ATOM’ and ‘TER’ records, and then move it to a new directory labeling it ‘1DWR_A.pdb’.
- Create a directory called ‘2_1_3_depc_ms/cl_data/’, enter the directory and create a text file called ‘myoglobin_depc.txt’. In the original source publication, K, H, S, T, and Y residues were assigned a status as being either labeled or unlabeled, with residues having measured labeling extents greater than 0.01% being assigned as labeled. The depc_ms score term requires experimental input to have this binary classification. Inside the text file, we will put the per-residue labeling classifications derived from the source DEPC data. Paste the following in the text file.
- The first column corresponds to residue numbers, and the second is the label classification, where L is for labeled, and U is for unlabeled.
- Make sure the numbering of residues in the text file corresponds to the numbering of residues in the FASTA file.
3
L
16
L
24
L
34
L
36
L
39
U
42
L
45
L
50
L
51
L
56
L
58
L
62
L
63
L
64
L
66
L
70
U
77
L
78
L
79
L
81
L
82
L
87
L
92
U
93
L
95
L
98
L
102
L
103
L
108
L
113
L
116
L
117
L
118
L
119
L
132
U
145
L
146
L
147
L
Structures for myoglobin can be generated using any of the monomeric structure prediction methods discussed in section 1 (1.1, 1.2, 1.5). For this tutorial, we will simply rescore the 1DWR crystal structure.
- Create and enter a directory called ‘2_1_3_depc_ms/rescoring/’. Create a flags file, called ‘1DWR_depc_ms_flags’, which contains the options for the score application. Enter these contents into the file:
- ‘-depc_ms indicates the text file with the DEPC data, and ‘-weights’ indicates to use the depc_ms score term weights.
-score
-depc_ms_input ../cl_data/myoglobin_depc.txt
-weights depc_ms.wts
- Run the following command to score the structure with depc_ms.
- ‘-in:file:s’ specifies the input PDB.
- ‘-out:file:scorefile’ sets the name of the output score file. New scores will automatically be appended to score files with the same name, so you do not need to make separate score files for each model if scoring multiple structures.
> ~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database
-in:file:s ../native/1DWR_A.pdb @1DWR_depc_ms_flags -out:file:scorefile
1DWR_depc_ms_score.sc
- The output ‘1DWR_depc_ms_score.sc’ will have the following format if using depc_ms.
- The value in the ‘score’ column is the value from the depc_ms term. To obtain a new total score for structures generated with Rosetta, add the depc_ms score to the ‘score’ value from the AbinitioRelax ‘score.fsc’. You could also calculate the normal Rosetta score by rerunning the command without the ‘@1DWR_depc_ms_flags’ flag file and then add the score from depc_ms. For rescoring of AlphaFold structures, you could score with and without the depc_ms weights and add them, or you could combine with AlphaFold’s confidence.
SCORE: score depc_ms allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −102.816 −102.816 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 152.000 152.000 0.000 1DWR_A_0001
Typically, rescoring is performed for multiple models, often thousands or tens of thousands. The best scoring structure, i.e. the structure with the lowest Rosetta Energy Units (REU) after rescoring, considering both the AbinitioRelax models and mover models, is the structure indicated by the method to best reflect the experimental data, and likely the most native-like.
2.1.4. cl_complex_rescore
For this tutorial, we will be investigating how to rescore docked protein complexes with the cl_complex_rescore Rosetta application. The input into this application are per-residue labeling data for the unbound and bound states of complexes. It then calculates the modification percentage (the percentage change in labeling), and then compares those percentages with calculated interface distances of docked complexes to calculate penalty scores.35 This method is the only covalent labeling method in this section that is a self-contained application (as opposed to a score term). As a model system, we will be using hydroxyl radical footprinting labeling rates for the heterodimer of actin bound with gelsolin segment 1 (actin/gs1) which were reported by Guan et al.41 However, the cl_complex_rescore Rosetta application can use any form of numerical labeling data, as long as values exist for both the monomeric and complex form.
Structures for both the actin and gelsolin chain can be generated with any of the monomeric methods previously described (1.1, 1.2, 1.5). These subunits can then be docked using RosettaDock (1.3), but for simplicity, we will simply use the crystal structure of the full complex (PDB 1YAG) for tutorial rescoring. Create a directory, ‘~/2_1_4_cl_complex_rescore/native/’, download the PDB of 1YAG from www.rcsb.org , remove non ‘ATOM’ and ‘TER’ records, and then move it to the new directory labeling it ‘1YAG_AG.pdb’.
- Create a directory called ‘~/2_1_4_cl_complex_rescore/cl_data/’, enter the directory and create a text file called ‘actin_gs1_cl_data.txt’. The covalent labeling data available for actin/gs1 are oxidative rate constants. Since these are used to calculate a percentage in modification difference, we do not convert them to PFs like in previous sections with hydroxyl radical footprinting data. Paste the following in the text file.
- The first column corresponds to residue numbers, the second is the unbound state labeling measurement or degree of modification, the third is the bound state labeling measurement or degree of modification, and the final column indicates which chain the residue belongs to. For this example, only data for the ‘A’ chain is used. For larger homo-oligomeric complexes, multiple chains can be specified. If, for example, a residue on identical chains ‘A’, B’, and ‘C’ has been labeled, you would specify ‘ABC’ in the chain column. The first row is an optional header for labeling columns but is ignored by Rosetta.
- Make sure the numbering of residues in the text file corresponds to the numbering of residues in the FASTA file.
#Residue_Index
Unbound_Modif.
Bound_Modif.
Chains
119
14
10
A
123
14
10
A
124
14
10
A
143
14
10
A
322
8.2
4.3
A
325
8.2
4.3
A
337
33
12
A
346
33
12
A
349
33
12
A
352
33
12
A
355
33
12
A
161
9.8
6.1
A
164
9.8
6.1
A
166
9.8
6.1
A
374
8.5
6.1
A
375
8.5
6.1
A
40
33
35
A
44
33
35
A
47
33
35
A
53
0.69
0.95
A
67
0.4
0.46
A
69
7.3
7.3
A
73
7.3
7.3
A
79
7.3
7.3
A
82
7.3
7.3
A
87
2.85
3.25
A
88
2.85
3.25
A
90
2.85
3.25
A
91
2.85
3.25
A
101
0.5
0.65
A
102
0.5
0.65
A
112
0.5
0.65
A
190
10
9.2
A
332
0.78
0.69
A
333
0.78
0.69
A
362
0.72
0.72
A
367
0.72
0.72
A
371
0.72
0.72
A
21
0.7
0.64
A
227
7.7
6.9
A
243
0.72
1.05
A
-
Create and enter a directory called ‘~/2_1_4_cl_complex_rescore/rescoring/’. cl_complex_rescore is meant to be run with the entire set of docked complexes due to normalization during rescoring. Typically, you would create a text file called ‘paths_to_docked_pdbs.txt’ which contains the paths of all the docked complexes. It would have this format:
~/2_1_4_cl_complex_rescore/docking/dock_ppk_ensemble_1YAG_AG_0001_0001.pdb
~/2_1_4_cl_complex_rescore/docking/dock_ppk_ensemble_1YAG_AG_0001_0002.pdb
~/2_1_4_cl_complex_rescore/docking/dock_ppk_ensemble_1YAG_AG_0001_0003.pdb
~/2_1_4_cl_complex_rescore/docking/dock_ppk_ensemble_1YAG_AG_0001_0004.pdb
~/2_1_4_cl_complex_rescore/docking/dock_ppk_ensemble_1YAG_AG_0001_0005.pdb
For simplicity, we will simply run the application for just the ‘1YAG_AG.pdb’. Create a text file called ‘path_to_native_pdb.txt’ and enter the following path:~/2_1_4_cl_complex_rescore/docking/native/1YAG_AG.pdb
- Once the text file has been generated, the models can be rescored using the cl_complex_rescore application by running the following command.
- ‘-in:file:l’ indicates to the application that a file containing PDB paths is being given.
- ‘-cl_data’ is the text file with the per-residue labeling data formatted as specified in step 3.
- ‘-interface’ indicates which chains comprise each docking partner. For oligomers with multi-chain subunits, chains of each docking partner should be combined. For example, if you docked chains ‘A’ and ‘B’ with ‘C’ and ‘D’, you would specify ‘AB_CD’.
>~/rosetta/source/bin/cl_complex_rescore.linuxgccrelease -in:file:l
paths_to_native_pdb.txt -cl_file ../cl_data/actin_gs1_cl_data.txt -
interface A_G
- After successfully running the application, a ‘cl_score.out’ file will be created. The first column specifies the model, the second is the raw penalty score for the model, and the third is the weighted and normalized penalty score. It should have this format:
Model
Raw_Penalty
Weighted_CL_ScoreTerm
../native/1YAG_AG.pdb
4.87576
65
The ‘Weighted_CL_ScoreTerm’ is just the penalty score calculated for each specified structure. For actual use, it should be added to RosettaDock’s interface score (I_sc), which is available in the output ‘dock_score.sc’ file which is generated when docking subunits. This new score should be used to identify the best scoring structure from a set of Rosetta-generated models.
2.1.5. hrf_GBM
This tutorial details how to score protein models in Rosetta using the hrf_GBM score term. In this approach, experimental HRF data, together with additional sequence-derived features, are provided as input to a gradient boosting decision tree regression model to predict expected residue neighbor counts for labeled residues. This algorithm improved on simple linear regression models used to derive neighbor counts from HRF data (see sections 2.1.1–2.1.3).22–24,26,27 These expected neighbor counts are then compared with the neighbor counts observed in a structural model to assess agreement between the experimental data and predicted structure.25 For this tutorial, we will also be using the myoglobin HRF protection factors from Xie et al.38
In addition to the target protein FASTA sequence and associated HRF data, the gradient boosting model requires additional input generated by external programs. This tutorial utilizes publicly available web servers to generate these features whenever possible. While convenient and accessible, these web servers are not readily automated. Advanced users who need to process many proteins are advised to compile their own versions of the programs mentioned below for a more streamlined workflow. The scripts required for feature generation and processing, as well as example input and output files, are available in the ‘~/2_1_5_hrf_GBM/’ directory.
Create a directory ‘~/2_1_5_hrf_GBM/native/’ and download the 1DWR PDB structure, as described in Section 2.1.1. The FASTA sequence for 1DWR should also be downloaded at this stage.
- Create a directory called ‘~/2_1_5_hrf_GBM/hrf/’. Prepare the experimental HRF data for neighbor count prediction.
- The gradient boosting model requires experimental data to be stored in a CSV file with the following columns:
- Column 1: Residue number (ensure that the residue numbering matches the FASTA sequence indexing rather than the residue numbering in the PDB file, if the two differ).
- Column 2: Three-letter amino acid abbreviation
- Column 3: Residue labeling rate (or labeling extent). Note that this differs from the ln(PF) values used in Section 2.1.1
- Create ‘1DWR_A_HRF.csv’ and add the experimental data to the file:
#Residue_ID
Residue_AA
HRF_Rate
7
TRP
0.29
8
GLN
0.2
9
GLN
0.2
11
LEU
0.073
12
ASN
0.1
13
VAL
0.093
14
TRP
0.22
18
GLU
0.073
21
ILE
0.12
24
HIS
0.015
26
GLN
0.019
27
GLU
0.019
30
ILE
0.22
31
ARG
0.011
33
PHE
0.04
36
HIS
0.33
37
PRO
0.054
38
GLU
0.031
40
LEU
0.07
42
LYS
0.033
66
THR
0.008
67
VAL
0.037
68
VAL
0.18
69
LEU
0.044
70
THR
0.16
72
LEU
0.054
77
LYS
0.028
113
HIS
1.14
116
HIS
1.04
119
HIS
0.43
120
PRO
0.43
123
PHE
0.04
125
ALA
0.21
126
ASP
0.098
128
GLN
0.22
131
MET
1.06
136
GLU
0.016
137
LEU
0.12
138
PHE
0.05
139
ARG
0.003
147
LYS
1.07
149
LEU
0.22
151
PHE
0.26
- Use the PSIPRED Workbench to predict per-residue secondary structure and disorder.
- Visit the PSIPRED Workbench at http://bioinf.cs.ucl.ac.uk/psipred/
- Set the input data type to ‘Sequence Data’ and select the ‘PSIPRED 4.0’ and ‘DISOPRED3’ under ‘Popular Analyses’.
- Upload the 1DWR FASTA sequence and submit the job.
- After the job is completed, download the ‘.ss2’ format PSIPRED output and the .comb format DISOPRED output, create a new directory ‘~/2_1_5_hrf_GBM/example_files’, and move the files to this directory. Rename the files to ‘1DWR_A.ss2’ and ‘1DWR_A_disopred.out’, respectively.
-
Generate the PSSM profile for 1DWR using PSI-BLAST with the following command:
Note: This step must be run using a local installation of the NCBI BLAST+ suite and a local copy of the nr (non-redundant) protein sequence database, as web-based PSI-BLAST servers do not produce output in a format compatible with the gradient boosting model and downstream feature-processing scripts.>~/psiblast -db nr -query 1DWR_A.fasta -inclusion_ethresh 0.001-out_pssm
1YMB_A.chk -out_ascii_pssm 1DWR_A.pssm_A -num_iterations 3 -num_alignments
0 -num_descriptions 5000
- Search the 1DWR sequence against the PDB using BLAST to identify homologs for side-chain environment (SCE) calculations.
- Navigate to https://blast.ncbi.nlm.nih.gov/Blast.cgi and upload the FASTA sequence.
- Select ‘blastp’ as the algorithm, and ‘Protein Data Bank proteins (pdb)’ as the desired database. Leave all other parameters as default.
- After the job is completed, download the hit table in CSV format and move it to ‘~/2_1_5_hrf_GBM/example_files/’. Rename the file to ‘1DWR_A_BLAST_HITS.csv’.
- Run ‘process_hits.py’ to download the BLAST hits and perform SCE calculations
>~/python 2_1_5_hrf_GBM/scripts/process_hits.py 1DWR_A
- Run the R script ‘feat_generation.R’ to generate additional features derived from peptide properties using the Peptides package. Use the following command:
>~/Rscript 2_1_5_hrf_GBM/scripts/feat_generation.R 1DWR A
- Confirm that all required input files are present in ‘~/2_1_5_hrf_GBM/example_files/’ and are named as follows:
- 1DWR_A.fasta
- 1DWR_A.ss2
- 1DWR_A_disopred.out
- 1DWR_A.pssm_A
- 1DWR_A_mapped.xlsx
- 1DWR_A_HRF.csv
- 1DWR_A_R_features.csv
- Run the following commands to format the input features and predict labeled-residue neighbor counts:
- ‘NC_regressor.pkl’ is a pickle file which contains the GBM regressor model.
>~/python 2_1_5_hrf_GBM_method/scripts/generate_features.py 1DWR A
>~/python 2_1_5_hrf_GBM_method/scripts/predict_NC.py 1DWR A
~/2_1_5_hrf_GBM/model/NC_regressor.pkl
- Examine the output file.
- The file ‘1DWR_A_predicted_NC.out’ contains residue number in the first column and predicted neighbor count in the second column. The file is automatically filtered to include only predicted neighbor counts for a subset of residues used for rescoring in Rosetta, WYFHLIR, which were determined to have the best rescoring results.
#1DWR_A predicted neighbor counts
7
7.04401948485227
11
6.658934372870389
14
8.68839427346131
21
5.675338674226502
24
8.836893720278962
30
7.276444307574219
31
5.7287215422704305
33
8.710358890425546
36
6.5405225162125715
40
6.801744535764449
69
8.639377686388134
72
8.402227512935422
113
5.416812662361045
116
4.753147771944863
119
5.793147651909494
123
8.471197377085108
137
6.174787391861471
138
7.3630890634082125
139
6.5303675249979865
149
6.172254643899492
151
7.140211421364629
-
Generate or select a structure for scoring.
Structures for myoglobin can be generated using any of the monomeric structure prediction methods discussed in Section 1 (Sections 1.1, 1.2, and 1.5). For this tutorial, however, we will simply score the 1DWR myoglobin crystal structure.
-
Score the structure using the hrf_GBM score term.
The hrf_GBM score term is called in a similar manner to the score term described in Sections 2.1.1 and 2.1.2. Create and enter a directory called ‘2_1_5_hrf_GBM/rescore/’, then run the following command to score the crystal structure with hrf_GBM:> ~/rosetta/source/bin/score.linuxgccrelease -database ~/rosetta/database
-in:file:s ../1DWR_A.pdb -score:weights hrf_gbm.wts -score:hrf_gbm_input
../1DWR_A_predicted_NC.out -out:file:scorefile 1DWR_A_hrf_GBM_score.sc
- Analyze the resulting score file.
- In ‘1DWR_A_hrf_GBM_score.sc’, the ‘score’ and ‘hrf_gbm’ columns both correspond to the hrf_GBM score assigned to the structure. Since this command evaluates only the hrf_GBM term, the value in the score column does not represent the overall Rosetta energy of the model. To calculate the total score, rerun the Rosetta ‘score’ application without the hrf_GBM-specific flags to obtain the standard Rosetta ref15 score, and then add the hrf_GBM score to the score value in the resulting score file.
SCORE: score hrf_gbm allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: −179.293 −179.293 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 152.000 152.000 0.000 1DWR_A_0001
2.2. Rescoring with SID applications: SID_Rescore, SID_ERMS_Rescore
Note: Files for this tutorial are located in the folder labeled ‘2_2_sid_methods/’. Tutorials for these methods were adapted from the original papers.
Mass spectrometry surface-induced dissociation (SID) is an ion activation method pioneered by Dr. Vicki Wysocki and coworkers for the study of protein complexes, in which protein complexes are first multiply charged by a soft ionization method and then transferred into the gas phase in conditions that preserve the native quaternary structure. The gas-phase complexes are accelerated toward a rigid, high-mass surface resulting in a collision event. Upon impact with the surface, protein complexes can break apart into subcomplexes and individual subunits, with fragmentation patterns which reflect the connectivity and topology of the native complex.42–46 In SID experiments, appearance energy (AE) of a given fragment, defined as the collision energy at which a specified amount of that fragment is first observed at 10% intensity, has been shown to correlate with relative strength of the interfaces within a complex.47 SID_Rescore is a Rosetta application that integrates experimental SID data with protein-protein docking predictions by comparing experimental AEs with those predicted for a computationally docked model.31,32 Energy-Resolved Mass Spectrometry (ERMS) profiles describe the abundance of the subunits/subcomplexes generated by collisions at different acceleration energies. SID_ERMS_Rescore is a Rosetta application where ERMS are simulated for each predicted model and compared directly to experimental ERMS curves to rescore docked complexes.34 SID_Rescore and SID_ERMS_Rescore are among the few computational methods developed specifically to incorporate SID measurements into computational structure prediction. By integrating appearance energies and energy-resolved dissociation data with docking prediction, these approaches allow the topology and interface information obtained from SID experiments to be used directly in protein-protein complex modeling. The tutorials presented here provide a comprehensive set of practical examples with instructions on how these tools can be implemented within experimental workflows.
For this tutorial, two subsections focus on using both SID_Rescore and SID_ERMS_Rescore, respectively. Each tutorial will provide example proteins and relevant data suitable for each method.
2.2.1. SID_Rescore
For this tutorial, we will be presenting how to rescore docked protein complexes with the SID_Rescore Rosetta application.31 The input into this application are PDB(s), the interface of interest for the complex, the number of intra-chain contacts, and the SID appearance energy of that interface. From this information, computationally docked complexes are rescored by comparing experimental and simulated AEs, resulting in a new SID_Score that is combined with the Rosetta score. For this example, we will be using choleragenoid, the B subunit pentamer of the cholera toxin, with PDB ID 1FGB.48 From this pentamer, we will only use the interface between the ‘D’ and ‘E’ chains. This interface was experimentally determined to have an appearance energy of 205 eV.47
Structures for individual subunits can be generated with any of the monomeric methods previously described (1.1, 1.2, 1.5). These subunits can then be docked using RosettaDock or SymDock (1.3, 1.4), but for simplicity, we will simply use the D and E chains of the 1FGB crystal structure as the dimer complex to score. Create a directory, ‘~/2_2_1_sid_rescore/native/’, and download the PDB of 1FGB from www.rcsb.org , remove non ‘ATOM’ and ‘TER’ records, remove all chains besides D and E, and then move it to the new directory, naming it ‘1FGB_DE.pdb’.
- Create a directory called ‘~/2_2_1_sid_rescore/rescore/’. Run the following command to calculate the SID_Rescore for 1FGB_DE.pdb.
- ‘-in:file:s’ specifies the input file. For rescoring multiple structures, a text file containing individual paths to each structure can be given with ‘-in:file:l’.
- ‘-AE’ specifies the interface appearance energy and should be given in eV.
- ‘-interface’ specifies the chains comprising each subunit/subcomplex of the interface. Chains should be specified in the order they appear in the PDB. For complexes with multi-chain subcomplexes, chains in each partner should be combined. So, for example, for a tetramer with two subcomplexes AB and CD, the interface would be specified as ‘AB_CD’.
- ‘-n_ints’ is the number of interfaces between subunits/subcomplexes.
- ‘-out:file:o’ is the name of the output score file.
- ‘-native’ can be specified if a RMSD to the reference structures is desired.
~/rosetta/source/bin/SID_rescore.linuxgccrelease -in:file:s
../native/1FGB_DE.pdb -AE 205 -interface D_E -n_ints 1 -out:file:o
1FGB_DE_sid_rescore.out
- The output will have the following format. ‘AE_pred’ is the predicted appearance energy, ‘Rosetta_score’ is the Rosetta score, ‘SID_score’ is the score from the SID score term, and ‘Rosetta_SID_score’ is the addition of the weighted SID score term to the Rosetta score. If a reference structure is specified with ‘-native’, then calculated RMSDs will be in the ‘RMSD’ column.
Pose_number AE_pred Rosetta_score SID_score Rosetta_SID_score RMSD
1 466.732 −116.65 2.694 −100.486 0
2.2.2. SID_ERMS_Rescore
For this tutorial, we will be presenting how to rescore docked protein complexes with the SID_ERMS_Rescore application.34 The application has been converted into a PyRosetta script and is included with this tutorial in ‘~/tutorials/scripts/SID_ERMS_Rescore_pyrosetta.py’. PyRosetta is a python-based interface to the Rosetta molecular modeling suite. Instructions for installing PyRosetta can be found here: https://www.pyrosetta.org/downloads.49 This application takes in a text file of PDBs, a text file describing the complex type and complex interfaces and symmetry, and experimental ERMS data. For this example, we will be using the full structure of choleragenoid, the B subunit pentamer of the cholera toxin, with PDB ID 1FGB.48
Structures for individual subunits can be generated with any of the monomeric methods previously described (1.1, 1.2, 1.5). These subunits can then be docked using RosettaDock or SymDock (1.3, 1.4), but for simplicity, we will simply use the full 1FGB crystal structure as the complex to score. Create a directory, ‘~/2_2_2_sid_erms_rescore/native/’, download the PDB of 1FGB from www.rcsb.org , remove non ‘ATOM’ and ‘TER’ records, and then move it to the new directory labeling it ‘1FGB_DEFGH.pdb’. Create a text file called ‘1FGB_file_list.txt’ which contains the path to this file. Normally, you would run the script with all the structures you generated, but for this example we are simply performing it on the native structure.
- Prior to running SID_ERMS_Rescore, we will relax the crystal structure to obtain important interface information provided by Rosetta. This should be repeated for any structures not docked with Rosetta. Run the following command to relax the ‘1FGB_DEFGH.pdb’ structure.
- ‘-in:file:s’ specifies the input structure.
- ‘-relax:fast’ indicates the fast relax protocol, which samples low-energy backbone and side-chain conformations through packing and minimizing.
>~/rosetta/source/bin/relax.linuxgccrelease -
in:file:s 1FGB_DEFGH.pdb -relax:fast
- Create the directory ‘~/2_2_2_sid_erms_rescore/erms/’. Inside this directory, create a CSV file called ‘ERMS_1FGB.csv’ containing the following experimental ERMS data:
- Each row corresponds to different acceleration energy. The first column corresponds to the acceleration energy and subsequent columns correspond to each different possible oligomeric state. For this example, it would be the precursor pentamer intensity, tetramer intensity, trimer intensity, dimer intensity, and monomer intensity.
0 1 0 0 0 0
220 1 0 0 0 0
440 0.65087 0.06387 0.07346 0.07615 0.13564
660 0.05629 0.21897 0.22785 0.24108 0.2558
880 0.01377 0.14718 0.27548 0.30199 0.2616
1100 0.00574 0.09317 0.25712 0.3138 0.33018
1320 0.00446 0.06339 0.20672 0.29055 0.43488
1540 0.00393 0.0468 0.15422 0.24546 0.54959
1760 0.00534 0.03601 0.11345 0.18793 0.65727
1980 0.01634 0.02817 0.08048 0.13861 0.73639
- Create a directory called ‘~/2_2_2_sid_erms_rescore/rescore/’. Inside this directory, create a file called ‘c5_complex_type’. SID_ERMS_Rescore requires a complex definition file. This file defines both the oligomeric state, as well as the different interfaces. The first line should contain all chains in the structure. Subsequent lines define symmetric interfaces, with each entry in the line corresponding to a unique interface of specific symmetric type. Since 1FGB is a C5 complex, paste the following into ‘c5_complex_type’.
- For more information about making the definition file, such as for complexes with Dihedral symmetry, see https://docs.rosettacommons.org/docs/latest/application_documentation/Application-Documentation/SID_ERMS_prediction.33
A B C D E
A_B B_C C_D D_E E_A
- Run the following command to use SID_ERMS_Rescore for scoring 1FGB_DEFGH.pdb via PyRosetta.
- ‘--pdb_filelist’ specifies a text file containing paths and files for all structures to be rescored with SID_ERMS_Rescore.
- ‘--symm_filename’ is the input file which contains the information on the number of subunits and connectivities.
- ‘--erms’ specifies the path to the ERMS data.
- ‘--output_path’ specifies the path and name of the output file.
>~ python SID_ERMS_Rescore_pyrosetta.py --pdb_filelist
~/2_2_2_sid_erms_rescore/native/1FGB_file_list.txt --symm_filename
c5_complex_type --erms ~/2_2_2_sid_erms_rescore/erms/ERMS_1FGB.csv --
output_path 1FGB_SID_ERMS_Score.out
- The output ‘1FGB_SID_ERMS_Score.out’ will have the following format. The ‘Structure’ column contains which particular file was scored, the ‘RMSE’ columns contains the ‘RMSE’ of the predicted ERMS for the given structure to the experimental ERMS, the ‘ref15score’ column contains the Rosetta score, and the ‘Rescore’ column contains the Rosetta score plus the score from SID_ERMS_Rescore.
Structure RMSE ref15score Rescore
relax_native_1fgb.pdb 0.2409359611630499 −1411.9971097021507 −703.0593871539096
2.3. Rescoring with IM score terms: IM_Score
Note: Files for this tutorial are located in the folder labeled ‘2_3_im_score/’. The tutorial for this method was adapted from the original paper.
Ion mobility (IM) is a structural mass spectrometry technique in which ions are transferred through an inert gas at constant pressure and temperature under the influence of an electric field. When applied to proteins, IM enables separation based on molecular shape and size. From IM measurements, the rotationally averaged collision cross section (CCS) of a protein can be determined, which reflects the extent of momentum exchange between the buffer gas and the ion.50 Consequently, CCS values can be incorporated into integrative modeling to identify native-like protein conformations among computationally generated structures.28,29
In this section, we provide a tutorial on using an IM-guided Rosetta score term to rescore a crystal structure of ubiquitin with experimental CCS.
For this tutorial, we will be presenting how to use the IM_Score score term with Rosetta’s ‘score’ application. We will use ubiquitin, with PDB ID 1UBQ, as an example.
IM_Score can be used to rescore monomeric structures derived from any of the methods described in 1.1, 1.2, and 1.5. For simplicity, to demonstrate the functionality of this score term, we will use the crystal structure of ubiquitin with PDB ID 1UBQ.51 Create a directory called ‘2_3_im_score/native/’. Enter this directory and download the PDB from https://www.rcsb.org, remove non ‘ATOM’ and ‘TER’ records, and then move it to the new directory labeling it ‘1UBQ_A.pdb’.
- Make and enter a directory called ‘2_3_im_score/rescore/’. CCS for ubiquitin has been experimentally measured as 930 Å2 with a helium buffer gas.52 Run the following command to score ‘1UBQ_A.pdb’ with IM_Score.
- ‘-in:file:s’ specifies the input file. For a list of structures, use ‘-in:file:l’.
- ‘-ccs_nrots’ is an optional flag that specifies the number of random rotations for the CCS calculations. The value must be given as an integer; the default is 300.
- ‘-ccs_prad’ is an optional flag for specifying the radius of the buffer gas (in Angstrom). Defaults are 1.0 and 1.81 for helium and nitrogen, respectively.
- ‘-ccs_exp’ specifies the experimental CCS (in Å2). The corresponding value for ‘-ccs_prad’ should be specified depending on the buffer gas.
- ‘-score:patch ccs_imms.wts_patch’ specified to include the IM score function during scoring.
>~/rosetta/source/bin/score.linuxgccrelease -in:file:s ../native/1UBQ_A.pdb
-ccs_nrots 300 -ccs_prad 1.0 -ccs_exp 930 -score:patch ccs_imms.wts_patch
- An output file is generated called ‘default.sc’ which contains the IM_Score score under the column called ‘ccs_imms’ (“100.000”). Unlike the previous score terms, the ‘score’ column (“132.678”) already contains the Rosetta score plus the score from IM_Score. It will have this content:
SCORE: score fa_atr fa_rep fa_sol fa_intra_rep fa_intra_sol_xover4 lk_ball_wtd fa_elec pro_close hbond_sr_bb hbond_lr_bb hbond_bb_sc hbond_sc dslf_fa13 omega fa_dun p_aa_pp yhh_planarity ref ccs_imms rama_prepro allatom_rms gdtmm gdtmm1_1 gdtmm2_2 gdtmm3_3 gdtmm4_3 gdtmm7_4 irms maxsub maxsub2.0 rms description
SCORE: 132.678 −397.647 57.039 242.952 1.777 16.826 −8.756 −113.091 2.383 −18.828 −23.132 −7.389 −1.549 0.000 1.713 288.599 −12.808 0.000 11.884 100.000 −7.297 0.000 1.000 1.000 1.000 1.000 1.000 1.000 0.000 76.000 76.000 0.000 1UBQ_0001
Section 3. Calculating various structural metrics with Python/PyRosetta
Many of the mass spectrometry methods which have been integrated within Rosetta utilize relationships between different structural metrics calculated from computational structures with experimental data. This section will describe some of the different metrics and how to use them. For covalent labeling, the two most common solvent exposure metrics are neighbor count and solvent accessible surface area (SASA). Neighbor count is defined as the number of neighboring residues within close proximity to a target residue, while SASA is the accessible surface area of a target residue exposed to solvent. For ion mobility, projection approximation via a grid-based calculation of rough circular shapes (PARCS) aims to estimate a collision cross section (CCS) of a protein. And finally for SID, energy-resolved mass spectrometry (ERMS), which shows the relative abundance of precursors and fragments over a range of acceleration energies, can be simulated from structures of protein complexes. The following tutorials will cover using each of these methods within Rosetta, as well as with PyRosetta if available.
3.1. Neighbor Count
Per-residue neighbor count is calculated by quantifying the number of residues surrounding a target residue. Two versions of the neighbor count were developed: a low-resolution centroid representation, in which all atoms of a residue side chain are represented by a single point, and full-atom representation, in which all atoms are explicitly defined.26 Neighbor count can be computed either by considering all residues within a sphere centered on the target residue (spherical mode) or by considering only residues located within a cone (conical mode).
In the cone-version of the method, the cone is centered along the vector defined by the target residue’s alpha carbon () and beta carbon (), where the angle determines the width of the cone. For a target residue and a neighbor residue , the neighbor count is calculated using two sigmoidal functions: a distance-dependent term and an angle-dependent term . The distance function ) (Equation 3.1) is calculated using the distance between the atom of residue and the centroid (CEN) of residue . In the full-atom representation, the atom is used instead of the centroid, with the atom substituted for glycine residues. Two parameters, and , control the steepness and midpoint of the function, respectively; by default, and . The angular function (Equation 3.2) is calculated using the angle (in radians) formed by the vectors CENj-Cαi-CENj (or for the full-atom representation, with used for glycine). This function is parameterized by and , which define the steepness and midpoint angle of the cone. By default, and . The final neighbor count contribution of residue on residue is computed as the product of the distance and angular terms (Equation 3.3).
In the sphere-version of the method, the neighbor count calculation only considers the distance function, ), which is consistent with the description provided above.
| 3.1 |
| 3.2 |
| 3.3 |
Note: Files for this tutorial are located in the folder labeled ‘3_1_neighbor_count/’. The tutorial for this method was adapted from the original paper. Additional documentation can be found at https://docs.rosettacommons.org/docs/latest/application_documentation/Application-Documentation/PerResidueSolventExposure.
In Rosetta, neighbor count can be calculated for a cleaned PDB (see the beginning of section 1) using the per_residue_solvent_exposure application. For this tutorial, we will be calculating neighbor count for the PDB of the B1 domain of streptococcal Protein G (GB1), with PDB ID 1PGA.7
Create a directory called ‘3_1_neighbor_count/’, download and move the 1PGA PDB from https://www.rcsb.org/structure/1PGA to this directory. Clean the PDB, by removing all header information and only keep ‘ATOM’ records.
- The following command is for calculating spherical neighbor count with the centroid representation.
- ‘-in:file:s’ specifies the input PDB.
- ‘-in:file:centroid’ indicates to Rosetta to read the PDB into the centroid representation.
- ‘-centroid_version’ indicates the application to calculate neighbor count with the centroid representation.
- ‘-solvent_exposure:method sphere’ indicates to calculate spherical neighbor count.
- ‘-dist_steepness’ and ‘-dist_midpoint’ are for setting and , respectively.
>~/rosetta/source/bin/per_residue_solvent_exposure.linuxgccrelease -
in:file:s 1PGA.pdb -in:file:centroid -centroid_version -
solvent_exposure:method sphere -dist_midpoint 9.0 -dist_steepness 1.0
- The following command is for calculating conical neighbor count with the centroid representation.
- ‘-solvent_exposure:method cone’ indicates to the application to calculate conical neighbor count.
- ‘-angle_steepness’ and ‘-angle_midpoint’ are for setting and , respetively.
>~/rosetta/source/bin/per_residue_solvent_exposure.linuxgccrelease -
in:file:s 1PGA.pdb -in:file:centroid -centroid_version -
solvent_exposure:method cone -dist_midpoint 9.0 -dist_steepness 1.0 -
angle_midpoint 1.5708 -angle_steepness 6.2832
- The following command is for calculating spherical neighbor count with full-atom representation.
- By default, the application calculates full-atom neighbor count, so removing both ‘-in:file:centroid’ and ‘-centroid_version’ flags will default to full-atom.
- The application also supports an additional flag with full-atom representation, ‘-neighbor_closest_atom’, which uses the closest atom of residue j rather than . This is not a recommended option, and is not used in this example.
>~/rosetta/source/bin/per_residue_solvent_exposure.linuxgccrelease -
in:file:s 1PGA.pdb -solvent_exposure:method sphere -dist_midpoint 9.0 -
dist_steepness 1.0
- The following command is for calculating conical neighbor count with full-atom representation.
- By default, the application calculates full-atom neighbor count, so removing both ‘-in:file:centroid’ and ‘-centroid_version’ flags will default to full-atom.
- The application also supports an additional flag with full-atom representation, ‘-neighbor_closest_atom’, which uses the closest atom of residue j rather than .
>~/rosetta/source/bin/per_residue_solvent_exposure.linuxgccrelease -
in:file:s 1PGA.pdb -solvent_exposure:method sphere -dist_midpoint 9.0 -
dist_steepness 1.0 -angle_midpoint 1.5708 -angle_steepness 6.2832
- An output file called ‘NeighborCount_default.out’ is produced, with the following format.
- The first column contains the residue type, the second is the residue number, and the third is the calculated neighbor count.
Residue_Type
Residue_Number
Neighbor_Count
M
1
8.86078
T
2
10.1108
Y
3
15.7881
K
4
15.5919
L
5
17.6711
I
6
15.0335
L
7
16.9441
N
8
13.4778
G
9
13.5977
K
10
8.39982
3.2. Solvent Accessible Surface Area
Solvent-accessible surface area (SASA) quantifies the extent to which a residue or atom is exposed to solvent and is commonly used as a measure of burial and packing in protein structures. In this tutorial, we will be using a Rosetta application, per_residue_sc_sasa, which uses the full-atom representation. This application uses the LeGrand-Merz method, which provides an efficient approximation of solvent accessibility through precomputed atomic surface sampling.
The LeGrand method estimates SASA by representing each atom with a fixed set of surface dots distributed over a sphere defined by the atom’s van der Waals radius plus a solvent probe radius.53 For a given atom , each surface dot is evaluated for occlusion by neighbor atoms. A dot is considered inaccessible if it lies within the expanded radius of another atom; otherwise, it contributes to the solvent-accessible surface area. The total SASA is calculated as the sum of accessible surface contributions from all atoms in the structure. per_residue_sc_sasa calculates per-residue SASA, which is obtained by summing the solvent-accessible surface areas of all atoms in a target residue. This application calculates the total SASA, which is the raw sum of per-atom contributions, and relative SASA, which is normalized by the SASA of a residue in an extended conformation, such as Gly-X-Gly, which generally represents the maximum possible exposure.
Note: Files for this tutorial are located in the folder labeled ‘3_2_sasa/’. The tutorial for this method was adapted from https://docs.rosettacommons.org/docs/latest/application_documentation/Application-Documentation/PerResidueSolventExposure.
In Rosetta, SASA can be calculated for a cleaned PDB using the per_residue_sc_sasa application. For this tutorial, we will be calculating the SASA for the PDB of the B1 domain of streptococcal Protein G (GB1), with PDB ID 1PGA.
Create a directory called ‘3_2_sasa/’, download and move the 1PGA PDB from https://www.rcsb.org/structure/1PGA to this directory. Clean the PDB, by removing all header information and only keep ‘ATOM’ records.
- The following is an example command for calculating SASA of 1PGA.
- Based on how the per_residue_sc_sasa application is written, an output file is not generated, the results are simply printed to screen. As a result, our example command contains ‘> 1pga_sasa.out’ at the end to pipe the results into a file called ‘1pga_sasa.out’.
- ‘-sasa:probe_radius sets the radius of the probe. The default is the radius of water, 1.4.
>~/rosetta/source/bin/per_residue_sc_sasa.linuxgccrelease -in:file:s
1PGA.pdb -sasa:probe_radius 1.4 > 1pga_sasa.out
- The output will have the following format:
- There will be some information that is output during the initialization of the application. Only lines which start with ‘apps.pilot.jkleman contain the SASA information.
- Each line corresponds to a residue. ‘resn indicates residue number, followed by the chain and residue type. ‘sasa: is the total SASA, and ‘rel sasa: is the relative SASA.
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 1 A MET sasa: 98.4457 rel sasa: 0.615286
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 2 A THR sasa: 73.2913 rel sasa: 0.718542
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 3 A TYR sasa: 1.34041 rel sasa: 0.00716798
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 4 A LYS sasa: 81.4962 rel sasa: 0.488001
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 5 A LEU sasa: 0 rel sasa: 0
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 6 A ILE sasa: 58.95 rel sasa: 0.421072
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 7 A LEU sasa: 4.35609 rel sasa: 0.0317963
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 8 A ASN sasa: 75.4589 rel sasa: 0.667778
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 9 A GLY sasa: 0 rel sasa: 0
apps.pilot.jkleman.per_residue_sc_sasa: PDB resn: 10 A LYS sasa: 146.78 rel sasa: 0.878925
3.3. PARCS
Projection Approximation using Rough Circular Shapes (PARCS) is a method implemented in Rosetta to estimate the collision cross section (CCS) of proteins from three-dimensional structures.28 CCS values measured by ion mobility (IM) experiments reflect the effective size and shape of an ion as it traverses an inert buffer gas under a weak electric field. PARCS enables the incorporation of such experimental size restraints into structure prediction and evaluation protocols by computing CCS directly from atomic coordinates.
The PARCS algorithm takes three-dimensional atomic coordinates as input and computes CCS through a projection-based approximation.28,29,54,55 The structure is first randomly rotated, and for each rotation, orthogonal two-dimensional projections are generated in the , , and planes. Each projection is mapped onto a two-dimensional grid with a resolution of 1 Å2 per grid cell. The projected structure is centered on the grid, which extends 5 Å beyond the most extreme projected atom in each direction. For each atom in a projection, the grid cell corresponding to the atom’s projected position is filled. Additional surrounding cells are also filled to approximate the atom’s cross-sectional area using a rough circular shape. The radius of this circular projection is determined by the sum of the projected atom’s effective radius and the buffer gas radius. Effective atomic radii of 1.91 Å are used for heavy atoms. Buffer gas radii of 1.0 Å for helium and 1.82 Å for nitrogen are applied depending on experimental conditions. The circular approximation is discretized using eight additional grid points positioned at 45° intervals around the central grid cell. This procedure is repeated for all atoms in the structure, resulting in a filled two-dimensional grid representing the molecular projection.
The projection area () is calculated by summing the areas of all occupied grid cells. From the , , and projections for each random rotation, three projection areas are obtained (, , and ). The CCS of the structure (), shown in Equation 3.4, is then acquired from the average area of the total number of projections (N = 3R, where R is the total number of random rotations).
| 3.4 |
Note: Files for this tutorial are located in the folder labeled ‘3_3_parcs/’. The tutorial for this method was adapted from the original paper.
In this tutorial, we will explain how to use the Rosetta application parcs_ccs_calc to calculate CCS from a protein structure. We will be predicting CCS for the PDB of the B1 domain of streptococcal Protein G (GB1), with PDB ID 1PGA.7
Create a directory called ‘3_3_parcs/’, download and move the 1PGA PDB from https://www.rcsb.org/structure/1PGA to this directory. Clean the PDB, by removing all header information and only keep ‘ATOM’ records.
- The following is an example command for using PARCS to calculate CCS of 1PGA.
- ‘-in:file:s’ specifies the input PDB file. A list of structures can be specified with in:file:l’.
- ‘-ccs_nrots’ is the number of random rotations performed during the calculation. It must be an integer; the default is 300.
- ‘-ccs_prad’ is the radius of the buffer gas probe. The default is 1.0 Å for helium gas. For nitrogen gas, use 1.81 Å.
>~/rosetta/source/bin/parcs_ccs_calc.linuxgccrelease -database
~/rosetta/database -in:file:s 1PGA.pdb -ccs_nrots 300 -ccs_prad 1.81
-out:file:o 1pga_nitrogen_ccs.out
- The output file contains two columns. One is for file names, and the other is CCS_PARCS. If given a list of files, each row will contain the CCS (in Å2) of a given structure. The following is the general format.
File_Name
CCS_PARCS
1PGA.pdb
799.81
PARCS is also available to use on the Rosetta Online Server that Includes Everyone (ROSIE) webserver.29,56 To use, navigate to https://rosie.graylab.jhu.edu/ and click on the [parcs-ccs-calc] option. A GitHub account is needed to access the server. Once logged in, you can upload a maximum of 100 PDB files, specify the input number of random rotations and probe radius, give the job a name and a job description, and then click ‘upload and queue job’. After the job is finished, you can download the results.
A PyRosetta script for PARCS has also been developed and is provided in ‘~/tutorials/scripts/PARCS_pyrosetta.py’. The commands to run the PyRosetta version of PARCS are the same as the C++ version, just with a python call to the script instead of calling the compiled Rosetta application. Additionally, a GIF file is provided at ‘~/tutorials/3_metrics/3_3_parcs/parcs.gif’ which illustrates how the PARCS algorithm works.
3.4. SID-ERMS
Surface-induced dissociation (SID) mass spectrometry captures how protein complexes dissociate into subcomplexes as a function of collision energy. In SID, a native protein complex is ionized and then accelerated into a collision surface where the conversion of collision energy into internal energy breaks up the complex into subunits/subcomplexes, depending on different interface strengths. The resulting breakdown of complexes is measured by tandem mass spectrometry and an energy profile of the relative intensities of each subcomplex type can be formed by sampling different acceleration energies, which is a type of data called energy-resolved mass spectrometry (ERMS). These ERMS data reflect the distribution of subunits/subcomplexes across a range of acceleration energies and can be used to infer interface strengths. SID_ERMS is a Rosetta application designed to predict the ERMS data of a given input protein complex by relating interface features with the probability of interface cleavage as acceleration energy increases.33 For each interface, a probability function (3.5) is calculated which defines the probability an interface breaks () based on the interface strength (, midpoint of the curve) and acceleration energy (). A steepness parameter () determines the sharpness or softness of the breakage threshold. (3.6) values are determined based on the observed correlations of particular interface features: interface surface area () and energy per interface residue (). , , and are optimal weights. Different values for A are used depending on oligomeric state and if a breakage cutoff is defined.
| 3.5 |
| 3.6 |
Note: Files for this tutorial are located in the folder labeled ‘3_4_sid_erms/’. The tutorial for this method was adapted from the original paper. Additional documentation can be found at https://docs.rosettacommons.org/docs/latest/application_documentation/Application-Documentation/SID_ERMS_prediction.
In this tutorial, we will describe how to use SID_ERMS to simulate ERMS profiles for choleragenoid, the B subunit pentamer of the cholera toxin, with PDB ID 1FGB.48
Create a directory called ‘3_4_sid_erms/’, download and move the 1FGB PDB from https://www.rcsb.org/structure/1FGB to this directory. Clean the PDB, by removing all header information and only keep ‘ATOM’ and ‘TER’ records. Save it as ‘1FGB_DEFGH.pdb’.
- The application requires a file which defines information about the complex type and the complex interfaces and symmetry. Create a file called ‘c5_complex_type’ and paste the following into the file.
- The first line of the file contains the letters for chain IDs present in the complex, with spaces between them.
- The subsequent lines contain specific interfaces for each unique symmetric interface type. Interfaces should be written as ‘Chain1_Chain2’. For example, the interface of chains A and B would be written as ‘A_B’. Unique symmetric interface types refer to identical chains in the structure whose interface are identical. A complex with C5 symmetry has 5 interfaces between 5 subunits (A, B, C, D, E), and each are identical. Therefore, there would only be one line with each interface, A_B, B_C, C_D, D_E, and E_A. For additional information, see https://docs.rosettacommons.org/docs/latest/application_documentation/Application-Documentation/SID_ERMS_prediction.
- 1FGB by default contains chains D, E, F, G, and H.
D E F G H
D_E E_F F_G G_H H_D
- Prior to running SID_ERMS, we will relax the crystal structure to obtain important interface information provided by Rosetta. This should be repeated for any structures not docked with Rosetta. Run the following command to relax the ‘1FGB_DEFGH.pdb’ structure.
- ‘-in:file:s’ specifies the input structure.
- ‘-relax:fast’ indicates the fast relax protocol, which samples low-energy backbone and side-chain conformations through packing and minimizing.
>~/rosetta/source/bin/relax.linuxgccrelease -in:file:s 1FGB_DEFGH.pdb -
relax:fast
- A new file, ‘1FGB_DEFGH_0001.pdb’ is generated. The following are all the options for running SID_ERMS_prediction.
- ‘-complex_type’ is the input file which contains the information on the number of subunits and connectivities.
- ‘-RMSE’ is an optional flag used to calculate RMSE of the prediction result. If specified, one must provide an experimental ERMS csv with ‘-ERMS’.
- ‘-ERMS’ is an optional flag which specifies the name of the experimental ERMS data file with the first column containing appearance energies. If ‘-RMSE’ is not specified, only the first column will be used to set the appearance energies. If ‘-RMSE’ is specified, the experimental ERMS data in the other columns starting with the second will be used to calculate RMSE of the predicted ERMS. If -ERMS is not specified, appearance energies will be automatically determined based on values, where they will range from 0 to , in 10 steps.
- ‘-B_vals’ and ‘-in:file:s’ are used for specifying the B values necessary for calculating the ERMS. To manually set the B values, use ‘-B_vals’ for specifying a file containing the interface strengths of each unique interface in the structure. The input file should contain the same number of rows corresponding to the number of unique interfaces in the file specified with ‘-complex-type’, and the same order. Only input one value per interface. To calculate the B values from an input structure, specify ‘-in:file:s’.
- ‘-steepness’ is an optional value for manually setting the steepness of the probability function. The default value is recommended.
- ‘-breakage_cutoff’ is an optional flag for specifying the breakage cutoff if it is known.
- ‘-out:file:o’ is an optional flag for creating an output file containing the ERMS prediction results.
- Run the following command to predict ERMS for the C5 homopentamer of the relaxed ‘1FGB_DEFGH_0001.pdb’ using the experimental ERMS file generated in 2_2_2. We provide experimental ERMS data to calculate RMSE, but is an optional flag and not necessary. If ERMS data is not given, then acceleration energies will be generated by the algorithm.
>~/rosetta/source/bin/SID_ERMS_prediction.linuxgccrelease -database
~/rosetta/database/ -complex_type c5_complex_type -out:file:o
ERMS_1FGB_predicted.out -in:file:s 1FGB_DEFGH_0001.pdb -ERMS ERMS_1FGB.txt
-RMSE
- The following is the contents of the output file with the simulated ERMS (‘ERMS_1FGB_predicted.out’). The first column corresponds to acceleration energies, and the subsequent columns correspond to the ERMS data of the pentamer, tetramer, trimer, dimer, and monomer, respectively.
0
1
0
0
0
0
220
0.946
0.0208
0.015
0.0112
0.007
440
0.861
0.0328
0.0546
0.0368
0.0148
660
0.722
0.0744
0.0828
0.076
0.0448
880
0.451
0.1408
0.1422
0.1488
0.1172
1100
0.256
0.132
0.1812
0.2292
0.2016
1320
0.102
0.1032
0.1878
0.2732
0.3338
1540
0.026
0.0392
0.1236
0.2968
0.5144
1760
0.004
0.0168
0.06
0.2708
0.6484
1980
0
0.0056
0.0294
0.1796
0.7854
- As a second example, run the following command to predict ERMS for the C5 homopentamer of the relaxed ‘1FGB_DEFGH_0001.pdb’ without using the experimental ERMS. The acceleration energies are different than the previous example since now, the application generates its own.
>~/rosetta/source/bin/SID_ERMS_prediction.linuxgccrelease -database
~/rosetta/database/ -complex_type c5_complex_type -out:file:o
ERMS_1FGB_predicted_no_RMSE.out -in:file:s 1FGB_DEFGH_0001.pdb
- The following is the contents of the output file with the simulated ERMS (‘ERMS_1FGB_predicted_no_RMSE.out’). The first column corresponds to acceleration energies, and the subsequent columns correspond to the ERMS data.
0
1
0
0
0
0
216.255
0.921
0.0256
0.0252
0.0168
0.0114
432.51
0.793
0.0608
0.0606
0.058
0.0276
648.764
0.619
0.1112
0.1086
0.0968
0.0644
865.019
0.38
0.1448
0.1686
0.1588
0.1478
1081.27
0.188
0.1208
0.189
0.2688
0.2334
1297.53
0.06
0.0688
0.1584
0.3056
0.4072
1513.78
0.02
0.0296
0.105
0.2808
0.5646
1730.04
0.003
0.0152
0.051
0.2228
0.708
1946.29
0
0.0032
0.0246
0.176
0.7962
2162.55
0
0
0.0102
0.104
0.8858
Practical Considerations and Limitations
While the workflows presented in this tutorial provide practical approaches for integrating structural MS data with computational modeling, several considerations are important when interpreting results. Different computational protein structure prediction methods are appropriate for different modeling problems. Rosetta ab initio modeling is generally more suitable for smaller proteins without homologous structures, whereas comparative/homology modeling is preferable when suitable templates are available. RosettaDock and SymDock are protein-protein docking algorithms and require approximate knowledge of interaction interfaces or symmetry relationships. Rosetta is a physics- and knowledge-based sampling method, where users need to generate thousands of structures to sample enough feasible conformations. AlphaFold2 is a deep learning neural network that can often generate accurate protein structures for monomers (and sometimes multimers) from only an amino acid sequence. However, AlphaFold2 may struggle with proteins that have not been previously studied (meaning fewer sequence homologs), large protein complexes, and dynamic proteins which adopt multiple conformations.57,58
Structural MS measurements provide indirect structural information and are recommended to be used to evaluate agreement between experimental signals and computational models. For example, covalent labeling experiments are generally interpreted in terms of relative solvent accessibility, ion mobility provides information related to overall molecular shape, and SID provides information regarding complex topology and connectivity. However, these data types are inherently sparse and cannot be used to completely generate an atomically accurate structure.
Workflows using Rosetta applications require adequate conformational sampling. Generating only a limited number of structures may not sufficiently explore conformational space and can reduce the likelihood of realistic predictions. Computational cost should also be considered when selecting workflows. Rosetta protocols are primarily CPU-based and can require substantial computational time for docking or ab initio simulations, whereas AlphaFold2 performance is strongly dependent on GPU availability and memory resources. As discussed in Section 1, access to high-performance computing resources can substantially improve runtimes for many of these methods.
Finally, users should carefully verify input formatting and residue numbering consistency between experimental datasets and structural models. Many Rosetta applications are sensitive to file formatting, alignment formatting, and residue indexing, which are common sources of error during setup and analysis.
Conclusions
Structural mass spectrometry offers a diverse set of experimental measurements that report on protein topology, solvent accessibility, and connectivity. Although these structural insights can provide valuable information for a target protein, MS alone cannot fully resolve its structure. Computational structure prediction has become an essential tool in structural biology, and the integration of experimental data provides an important strategy for improving prediction accuracy and confidence. By incorporating MS-derived information into computational methods, researchers can reduce conformational sampling space, evaluate structures more effectively, and obtain structures that better reflect experimental observed structural features.
In this review, we presented a series of tutorials that guide users through common computational structure prediction workflows and demonstrate how structural mass spectrometry data can be integrated into these approaches. We first introduced several widely used structure prediction tools, including Rosetta-based modeling methods and AlphaFold2. We then described Rosetta applications and score terms that incorporate MS-derived restraints for rescoring and evaluating predicted structures. Finally, we outlined methods for calculating structural metrics commonly used to interpret MS experiments, including solvent exposure, collision cross sections, and simulated dissociation behavior. Additionally, we developed novel PyRosetta implementations of PARCS and SID_ERMS_Rescore applications. Together, these tutorials aim to provide practical guidance for researchers interested in combining computational modeling with structural mass spectrometry. As both computational predictions methods and MS-based experimental techniques continue to advance, integrative approaches that combine these complementary sources of information are likely to play an increasingly important role in accurate protein structure determination and analysis.
Recent advances in deep learning-based prediction methods have substantially improved the accuracy and accessibility of protein structure prediction. Recent work has demonstrated that sparse experimental data can be incorporated into network architectures such as AlphaFold2 to improve prediction accuracy and guide structural refinement.59–62 Although structural MS data have not yet been widely integrated into these frameworks, techniques such as covalent labeling, ion mobility, and surface-induced dissociation provide experimentally derived information on solvent accessibility, overall shape, and subunit connectivity that could serve as valuable complementary restraints for future AI-based modeling approaches.
At the same time, important challenges remain in computational structural biology, particularly for modeling protein dynamics, transient interactions, ligand-induced conformational changes, and large macromolecular assemblies. Many structural MS techniques are especially well suited to probe these properties experimentally, providing information that can help evaluate, refine, and interpret computational predictions. Future progress in the field will likely depend on increasingly integrative workflows that combine predictive modeling with orthogonal experimental measurements. This may include tighter incorporation of MS-derived measurements into deep learning-based prediction pipelines, improved approaches for modeling conformational heterogeneity and structural ensembles, and more streamlined workflows for interpreting diverse MS datasets alongside computational models. Together, these developments will continue to strengthen the role of integrative MS-guided modeling in studying biomolecular systems that remain difficult to characterize using any single structural method alone.
Supplementary Material
A ZIP archive, called ‘tutorials.zip’, is provided as a supplementary file containing all tutorial files described in the manuscript. Subdirectories correspond to the individual sections and provide the example files, settings, and workflows used in the tutorials.
Acknowledgments
The authors would like to thank the members of the Lindert group for useful discussions. We thank SM Bargeen Alam Turzo for providing a GIF depicting the PARCS algorithm. We also thank our outstanding mass spectrometry collaborators, Vicki Wysocki, Sophie Harvey, Lisa Jones, James Prell, Joshua Sharp, and Richard Vachet, for their support in developing novel computational algorithms for protein structure elucidation from MS data. This work was supported by NIH (RM1-GM149374 and R01-GM127267 to S.L.).
References
- 1.Seffernick JT & Lindert S. Hybrid methods for combined experimental and computational determination of protein structure. J. Chem. Phys. 153, 240901 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Biehn SE & Lindert S. Protein Structure Prediction with Mass Spectrometry Data. Annu. Rev. Phys. Chem. 73, 1–19 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Mirdita M et al. ColabFold: making protein folding accessible to all. Nat. Methods 19, 679–682 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Simons KT, Kooperberg C, Huang E & Baker D. Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions. J. Mol. Biol. 268, 209–225 (1997). [DOI] [PubMed] [Google Scholar]
- 5.Bradley P, Misura KMS & Baker D. Toward high-resolution de novo structure prediction for small proteins. Science 309, 1868–1871 (2005). [DOI] [PubMed] [Google Scholar]
- 6.Rohl CA, Strauss CEM, Misura KMS & Baker D. Protein structure prediction using Rosetta. Methods Enzymol. 383, 66–93 (2004). [DOI] [PubMed] [Google Scholar]
- 7.Gallagher T, Alexander P, Bryan P & Gilliland GL. Two crystal structures of the B1 immunoglobulin-binding domain of streptococcal protein G and comparison with NMR. Biochemistry 33, 4721–4729 (1994). [PubMed] [Google Scholar]
- 8.The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 53, D609–D617 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Kim DE, Chivian D & Baker D. Protein structure prediction and analysis using the Robetta server. Nucleic Acids Res. 32, W526–531 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Song Y et al. High-Resolution Comparative Modeling with RosettaCM. Structure 21, 1735–1742 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Altschul SF et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 25, 3389–3402 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Sievers F et al. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Mol. Syst. Biol. 7, 539 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Gray JJ et al. Protein-protein docking with simultaneous optimization of rigid-body displacement and side-chain conformations. J. Mol. Biol. 331, 281–299 (2003). [DOI] [PubMed] [Google Scholar]
- 14.Vorobiev S et al. The structure of nonvertebrate actin: Implications for the ATP hydrolytic mechanism. Proc. Natl. Acad. Sci. 100, 5760–5765 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Marze NA, Roy Burman SS, Sheffler W & Gray JJ. Efficient flexible backbone protein-protein docking for challenging targets. Bioinformatics 34, 3461–3469 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Chaudhury S & Gray JJ. Conformer selection and induced fit in flexible backbone protein-protein docking using computational and NMR ensembles. J. Mol. Biol. 381, 1068–1087 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.André I, Bradley P, Wang C & Baker D. Prediction of the structure of symmetrical protein assemblies. Proc. Natl. Acad. Sci. 104, 17656–17661 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ahnert SE, Marsh JA, Hernández H, Robinson CV & Teichmann SA. Principles of assembly reveal a periodic table of protein complexes. Science 350, aaa2245 (2015). [DOI] [PubMed] [Google Scholar]
- 19.Eakin CM, Berman AJ & Miranker AD. A native to amyloidogenic transition regulated by a backbone trigger. Nat. Struct. Mol. Biol. 13, 202–208 (2006). [DOI] [PubMed] [Google Scholar]
- 20.Jumper J et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Steinegger M & Söding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026–1028 (2017). [DOI] [PubMed] [Google Scholar]
- 22.Biehn SE & Lindert S. Accurate protein structure prediction with hydroxyl radical protein footprinting data. Nat. Commun. 12, 341 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Biehn SE, Limpikirati P, Vachet RW & Lindert S. Utilization of Hydrophobic Microenvironment Sensitivity in Diethylpyrocarbonate Labeling for Protein Structure Prediction. Anal. Chem. 93, 8188–8195 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Biehn SE, Picarello DM, Pan X, Vachet RW & Lindert S. Accounting for Neighboring Residue Hydrophobicity in Diethylpyrocarbonate Labeling Mass Spectrometry Improves Rosetta Protein Structure Prediction. J. Am. Soc. Mass Spectrom. 33, 584–591 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Day EH & Lindert S. Extracting Residue Solvent Exposure from Covalent Labeling Data with Machine Learning: A Hybrid Approach for Protein Structure Prediction. J. Am. Soc. Mass Spectrom. 36, 1336–1347 (2025). [DOI] [PubMed] [Google Scholar]
- 26.Aprahamian ML & Lindert S. Utility of Covalent Labeling Mass Spectrometry Data in Protein Structure Prediction with Rosetta. J. Chem. Theory Comput. 15, 3410–3424 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Aprahamian ML, Chea EE, Jones LM & Lindert S. Rosetta Protein Structure Prediction from Hydroxyl Radical Protein Footprinting Mass Spectrometry Data. Anal. Chem. 90, 7721–7729 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Turzo SBA et al. Protein shape sampled by ion mobility mass spectrometry consistently improves protein structure prediction. Nat. Commun. 13, 4377 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Turzo SMBA, Seffernick JT, Lyskov S & Lindert S. Predicting ion mobility collision cross sections using projection approximation with ROSIE-PARCS webserver. Brief. Bioinform. 24, bbad308 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Khaje NA et al. Validated determination of NRG1 Ig-like domain structure by mass spectrometry coupled with computational modeling. Commun. Biol. 5, 452 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Seffernick JT, Harvey SR, Wysocki VH & Lindert S. Predicting Protein Complex Structure from Surface-Induced Dissociation Mass Spectrometry Data. ACS Cent. Sci. 5, 1330–1341 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Seffernick JT, Canfield SM, Harvey SR, Wysocki VH & Lindert S. Prediction of Protein Complex Structure Using Surface-Induced Dissociation and Cryo-Electron Microscopy. Anal. Chem. 93, 7596–7605 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Seffernick JT et al. Simulation of Energy-Resolved Mass Spectrometry Distributions from Surface-Induced Dissociation. Anal. Chem. 94, 10506–10514 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Bolz RM et al. Energy Resolved Mass Spectrometry Data from Surfaced Induced Dissociation Improves Prediction of Protein Complex Structure. Anal. Chem. 97, 2375–2383 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Drake ZC, Seffernick JT & Lindert S. Protein complex prediction using Rosetta, AlphaFold, and mass spectrometry covalent labeling. Nat. Commun. 13, 7846 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Drake ZC, Fowler AG, Blum AA & Lindert S. Enhanced Protein Complex Prediction via Rosetta, AlphaFold, and Nondifferential Covalent Labeling Mass Spectrometry. J. Phys. Chem. B 129, 6489–6497 (2025). [DOI] [PubMed] [Google Scholar]
- 37.Chu K et al. Structure of a ligand-binding intermediate in wild-type carbonmonoxy myoglobin. Nature 403, 921–923 (2000). [DOI] [PubMed] [Google Scholar]
- 38.Xie B, Sood A, Woods RJ & Sharp JS. Quantitative Protein Topography Measurements by High Resolution Hydroxyl Radical Protein Footprinting Enable Accurate Molecular Model Selection. Sci. Rep. 7, 4552 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Huang W, Ravikumar KM, Chance MR & Yang S. Quantitative mapping of protein structure by hydroxyl radical footprinting-mediated structural mass spectrometry: a protection factor analysis. Biophys. J. 108, 107–115 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Zhou Y & Vachet RW. Increased protein structural resolution from diethylpyrocarbonate-based covalent labeling and mass spectrometric detection. J. Am. Soc. Mass Spectrom. 23, 708–717 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Guan J-Q, Almo SC, Reisler E & Chance MR. Structural reorganization of proteins revealed by radiolysis and mass spectrometry: G-actin solution structure is divalent cation dependent. Biochemistry 42, 11992–12000 (2003). [DOI] [PubMed] [Google Scholar]
- 42.Wysocki VH, Joyce KE, Jones CM & Beardsley RL. Surface-Induced Dissociation of Small Molecules, Peptides, and Non-covalent Protein Complexes. J. Am. Soc. Mass Spectrom. 19, 190–208 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wysocki VH, Ding J-M, Jones JL, Callahan JH & King FL. Surface-induced dissociation in tandem quadrupole mass spectrometers: A comparison of three designs. J. Am. Soc. Mass Spectrom. 3, 27–32 (1992). [DOI] [PubMed] [Google Scholar]
- 44.Zhou M & Wysocki VH. Surface Induced Dissociation: Dissecting Noncovalent Protein Complexes in the Gas phase. Acc. Chem. Res. 47, 1010–1018 (2014). [DOI] [PubMed] [Google Scholar]
- 45.Zhou M, Dagan S & Wysocki VH. Protein subunits released by surface collisions of noncovalent complexes: nativelike compact structures revealed by ion mobility mass spectrometry. Angew. Chem. 51, 4336–4339 (2012). [DOI] [PubMed] [Google Scholar]
- 46.Stiving AQ et al. Surface-Induced Dissociation: An Effective Method for Characterization of Protein Quaternary Structure. Anal. Chem. 91, 190–209 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Harvey SR et al. Relative interfacial cleavage energetics of protein complexes revealed by surface collisions. Proc. Natl. Acad. Sci. 116, 8143–8148 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Zhang RG et al. The 2.4 A crystal structure of cholera toxin B subunit pentamer: choleragenoid. J. Mol. Biol. 251, 550–562 (1995). [DOI] [PubMed] [Google Scholar]
- 49.Chaudhury S, Lyskov S & Gray JJ. PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta. Bioinformatics 26, 689–691 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Dodds JN & Baker ES. Ion Mobility Spectrometry: Fundamental Concepts, Instrumentation, Applications, and the Road Ahead. J. Am. Soc. Mass Spectrom. 30, 2185–2195 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Vijay-Kumar S, Bugg CE & Cook WJ. Structure of ubiquitin refined at 1.8 A resolution. J. Mol. Biol. 194, 531–544 (1987). [DOI] [PubMed] [Google Scholar]
- 52.Stiving AQ, Jones BJ, Ujma J, Giles K & Wysocki VH. Collision Cross Sections of Charge-Reduced Proteins and Protein Complexes: A Database for Collision Cross Section Calibration. Anal. Chem. 92, 4475–4483 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Le Grand SM & Merz KM Jr. Rapid approximation to molecular surface area via the use of Boolean logic and look-up tables. J. Comput. Chem. 14, 349–352 (1993). [Google Scholar]
- 54.Howard JB, Narayanasamy A & Lindert S. Improving Protein Structure Prediction Using Integrative Cryo-EM and Ion Mobility Mass Spectrometry Modeling. bioRxiv 2026.02.07.704481 (2026) doi: 10.64898/2026.02.07.704481. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Narayanasamy A et al. Ion Mobility Mass Spectrometry Guided Modeling with AlphaFold and Rosetta Improves Protein Complex Structure Prediction. BioRxiv Prepr. Serv. Biol 2026.02.16.706193 (2026) doi: 10.64898/2026.02.16.706193. [DOI] [Google Scholar]
- 56.Moretti R, Lyskov S, Das R, Meiler J & Gray JJ. Web-accessible molecular modeling with Rosetta: The Rosetta Online Server that Includes Everyone (ROSIE). Protein Sci. Publ. Protein Soc. 27, 259–268 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Meng Q, Guo F & Tang J. Improved structure-related prediction for insufficient homologous proteins using MSA enhancement and pre-trained language model. Brief. Bioinform. 24, bbad217 (2023). [DOI] [PubMed] [Google Scholar]
- 58.Gaudreault F, Corbeil CR & Sulea T. Enhanced antibody-antigen structure prediction from molecular docking using AlphaFold2. Sci. Rep. 13, 15107 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Drake ZC, Day EH, Toth PD & Lindert S. Deep-learning structure elucidation from single-mutant deep mutational scanning. Nat. Commun. 16, 6874 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Wu T, Stein RA, Kao T-Y, Brown B & Mchaourab HS. Modeling protein conformational ensembles by guiding AlphaFold2 with Double Electron Electron Resonance (DEER) distance distributions. Nat. Commun. 16, 7107 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Stahl K, Graziadei A, Dau T, Brock O & Rappsilber J. Protein structure prediction with in-cell photo-crosslinking mass spectrometry and deep learning. Nat. Biotechnol. 41, 1810–1819 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Stahl K. et al. Modelling protein complexes with crosslinking mass spectrometry and deep learning. Nat. Commun. 15, 7866 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
