Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Jul 16.
Published in final edited form as: J Mol Biol. 2025 Feb 11;437(15):168996. doi: 10.1016/j.jmb.2025.168996

ModelArchive: A Deposition Database for Computational Macromolecular Structural Models

Gerardo Tauriello 1,, Andrew M Waterhouse 1,, Juergen Haas 1, Dario Behringer 1, Stefan Bienert 1, Thomas Garello 1, Torsten Schwede 1,*
PMCID: PMC13370677  NIHMSID: NIHMS2171745  PMID: 39947281

Abstract

A wide range of applications in life science research benefit from the availability of three-dimensional structures of biological macromolecules as they provide valuable insights into their molecular function. Recent advances in structure prediction techniques have made it possible to generate high quality computational macromolecular structural models for almost all known proteins. In this context, ModelArchive (https://modelarchive.org/) serves as a deposition database for computational models, complementing the Protein Data Bank (PDB) and PDB-IHM, which require experimental data, and specialised databases such as the AlphaFold DB. ModelArchive contains over 600,000 models contributed by researchers using a variety of modelling techniques. It supports single biological macromolecules and complexes, including any combination of polymers and small molecules. Each deposited model can be referenced in manuscripts using an immutable accession code provided by ModelArchive. Depositors are required to provide a minimal set of information about the modelling process and the expected accuracy of the resulting model, enabling scientific reproducibility and maximising the potential reuse of the models. The vast majority of models in ModelArchive use the ModelCIF format which includes coordinates and metadata, allows for programmatic validation of the models, and makes the models interoperable with structures obtained from other sources such as the PDB. The ModelArchive web service provides access to the models and search queries. Model findability is also provided in external services either through APIs or by importing data from ModelArchive.

Keywords: ModelArchive, FAIR databases, structural biology, macromolecular structure prediction, ModelCIF

Introduction

Proteins, DNA, and RNA are indispensable players in all biological processes, and their functions are inextricably linked to their three-dimensional structure. The determination of structures has heretofore been conducted mainly through experimental means. However, recent developments in computational methods have resulted in notable advancements in the accurate prediction of three-dimensional macromolecular structures. The most recent illustration of these advancements is the awarding of the 2024 Nobel Prize in Chemistry to protein structure prediction and specifically for the development of AlphaFold2.1 Computed structure models (CSMs) can nowadays be produced with sufficient accuracy to often rival those determined by experiment, and enable new possibilities for research where experimental structures are unavailable or impractical to obtain.

The growing utility of CSMs poses new challenges in terms of reproducibility and reusability. Unlike experimental structures, in the scientific literature CSMs are not always systematically deposited or annotated. This lack of standardisation hinders the validation, reuse, and integration of CSMs into downstream studies. As the scientific community increasingly uses CSMs for tasks such as protein function determination, drug design, protein engineering, and variant analysis, the need for dedicated repositories to store, share, and document these models has become apparent.

Since its establishment in 1971, the Protein Data Bank (PDB) has functioned as a unified global repository for three-dimensional structures of biological macromolecules obtained through experimental techniques, including X-ray crystallography, nuclear magnetic resonance (NMR), and, more recently, cryo-electron microscopy (EM).2 However, since 2006, the PDB has not archived structures determined through computational modelling,3 resulting in CSMs stored in undefined locations, in incompatible formats, and lacking essential metadata. In a workshop held in 2008, leading international scientists involved with structure modelling investigated the utilisation of CSMs in biomedical research4 and formulated a set of recommendations for the archiving and sharing of CSMs. The workshop emphasised the value of enhancing the structural coverage of protein sequences with CSMs, the necessity for data standards and best practices for the publication and dissemination of models, and the pivotal importance of providing model accuracy estimates for CSMs. Additionally, it advocated for the establishment of an archive for CSMs that could not be deposited in the PDB.

Consequently, an initial iteration of the ModelArchive was developed as part of the Protein Structure Initiative (PSI) Structural Biology Knowledgebase.5 The archive incorporated CSMs that had been stored in the PDB prior to 2006 and has been accepting new depositions since 2013. ModelArchive has been designed with the specific purpose of depositing structural models that are not based on experimental data. This provides a complementary resource to the PDB, which serves the same function for experimental structures, and to PDB-IHM,6 which is dedicated to integrative structures.7

In addition to the ModelArchive, other data resources have been established to facilitate large-scale access to CSMs generated through a specific automated prediction method. These include ModBase8 and the SWISS-MODEL Repository,9 which contain CSMs for millions of proteins generated by homology modelling using Modeller and SWISS-MODEL, respectively. These models are continuously updated to reflect the latest advances in the modelling methods and the availability of input data. More recently, the AlphaFold Protein Structure Database10 and the ESM Metagenomic Atlas11 have been established. The AlphaFold database contains over 200 million proteins modelled using AlphaFold2, while the ESM atlas contains 772 million proteins modelled using ESMFold. Despite their widespread use, these repositories are limited in scope. For example, the models are all monomers, they exclude either all or large parts of viral proteins, and they are limited to a maximum sequence length. In addition, they are tied to specific modelling pipelines and do not allow external deposition, limiting their scope for wider CSM integration.

The Biological Structure Model Archive (BSM-Arc)12 and general-purpose repositories such as Dryad, FigShare, Zenodo, and the Open Science Framework (OSF) provide researchers with alternative storage options for the sharing of CSMs. However, these platforms lack the structural metadata and standardisation required for effective reuse and interoperability of the contained CSMs, thereby reducing their potential impact.

CSM interoperability and reuse remain a topic of ongoing research. For example, the Research Data Management Kit13,14 provides best-practice guidelines for the administration of computational data, whereas the ModelCIF format15 offers a standardised data structure for the description of CSMs. The 3D-Beacons network16 makes a further contribution by incorporating federated queries to multiple structural model resources, thus providing a unified platform for accessing and visualising models. These initiatives demonstrate the progress made in supporting computational models, yet they underline the need for a centralised repository that allows for depositions and that adheres to the FAIR (Findable, Accessible, Interoperable, Reusable) principles for CSMs.

The ModelArchive addresses this need by providing a dedicated platform for the deposition, curation, and sharing of CSMs. By enabling FAIR access to models and emphasising metadata standardisation, ModelArchive ensures that CSMs remain accessible and reusable for future scientific endeavours. Here we describe the design, features, and current content of ModelArchive (https://modelarchive.org/), demonstrating its potential to enhance reproducibility and accelerate discoveries in the structural biology community.

Results and Discussion

Data deposition pipeline

The deposition pipeline (Figure 1) for ModelArchive ensures that Computed Structure Models (CSMs) are collected, curated and archived for scientific reproducibility and reuse. The process consists of three steps. First, molecular coordinates and metadata are collected to describe the model generation process. The data is then processed, curated, and validated by the ModelArchive team. Finally, the model is securely stored with a unique accession code and made available either publicly or with restricted access for review.

Figure 1.

Figure 1.

Schematic representation of the data deposition pipeline for ModelArchive, showing the iterative processing of data collected for modelling results and the archiving of an entry. Curation and validation by the ModelArchive team ensures that the provided metadata is sufficient for scientific reproducibility and has potential for reuse.

ModelArchive supports the same range of molecular entities as the PDB, such as single molecules and complexes containing proteins, RNA, DNA, carbohydrates or small molecules. However, complementary to PDB or PDB-IHM, which require experimental data, ModelArchive only accepts computationally generated models where no experimental input is used in modelling. Certain types of structural data are outside the scope of ModelArchive, including molecular dynamics trajectories and protein ensembles for intrinsically disordered proteins, the latter already covered by the PED database.17

Deposited models in ModelArchive are categorised into three different types based on their format and metadata structure. Legacy pre-2006 models have been migrated from the PDB to preserve historical data and are stored in the legacy PDB format. Individual ModelArchive entries currently contain free text descriptions of the metadata, allowing individual models to be reproduced and evaluated while maintaining simplicity for depositors. For these entries, coordinates are stored in the standard wwPDB PDBx/mmCIF format. Ongoing efforts will ensure that these models use the ModelCIF format in the future. Large model sets already use the ModelCIF format, which is an extension of PDBx/mmCIF. This stores the metadata alongside the model coordinates in a standardised data structure. This allows for programmatic validation and the handling of large model sets, both during deposition and for further reuse of the models. Further details of the formats used for data import can be found in the supplementary material. The use of standard file formats facilitates integration with structural data from other sources. Standardised metadata ensures that deposited models can be checked for suitability for a wide range of applications in structural biology.

To ensure scientific reproducibility and maximum potential for reuse, depositors are required to provide metadata covering several essential aspects. First, the metadata must clearly describe the molecular content of the model, including descriptions of the molecular entities and cross-references to resources such as UniProtKB or NCBI’s protein database. Second, it must document the modelling process in sufficient detail, including the software used, the specific steps performed and the input data, so that others can reproduce the model. Third, submitters must provide estimates of the expected accuracy of the model. Modern modelling methods are capable not only of generating accurate structures, but also of reliably assessing their own accuracy, both globally and locally, to identify potential inaccuracies in regions of interest. Typically, the accuracy estimates provided predict the expected similarity to the unknown correct structure according to metrics such as LDDT,18 TM-score19 or DockQ.20 In addition, depositors provide information to improve the usability and discoverability of their models. This includes a concise title, a description of the purpose of the model, the type of model (e.g. homology-based or de novo), an illustrative image, the list of authors, optional funding information to acknowledge supporting grants, and an optional zip file containing supplementary data.

The ModelArchive team validates each deposition to ensure that it meets both syntactic and semantic standards. Syntactic validation confirms that all required data is included and correctly formatted, while semantic validation ensures consistency between the metadata and the model itself, for example by checking the match between the modelled sequence and its stated source. If discrepancies or missing information are identified, the submission is returned to the submitter with feedback for correction. While model confidence is visible as part of the metadata, it is not used as a criterion for rejection. This approach ensures that the repository maintains high standards without excluding potentially useful models.

Each model deposited in ModelArchive is assigned a unique accession code, such as “ma-jd-viral-22025”, which resolves to a stable URL for citation in publications (e.g. https://www.modelarchive.org/doi/10.5452/ma-jd-viral-22025). Depositors can choose to make their models publicly available immediately, or to delay public release until after publication, while providing password-protected access for peer review. Once a model has been accepted, depositors can request its release and provide citation details to link the model to the publication for which it was created. To ensure that models do not remain private or uncited, we send annual reminders to all depositors with such entries. The latest information on the deposition process can be found on the ModelArchive help page (https://modelarchive.org/help).

Overview of current content

By the end of 2024, there were 618,491 models publicly available in ModelArchive. Each model has a unique accession code and the vast majority of models (615,828) are grouped into model sets using the ModelCIF format and have been added in the last four years. The depositions come from 279 depositors from 43 different countries. With the exception of pre-2006 legacy models migrated from the PDB, all models have been deposited by the creators of the models. Table 1 lists the current content, including each grouped model set. Model sets have a single accession code that represents the entire set (e.g. ma-jd-viral) and acts as a prefix to the accession code of each individual model (e.g. ma-jd-viral-22025). Historical growth of ModelArchive can be found in the supplementary material.

Table 1.

Publicly visible content in ModelArchive by the end of 2024, including pre-2006 legacy models, individual non-ModelCIF ModelArchive entries, and all model sets with ModelCIF formatted files identified by their ModelArchive accession code listed in chronological order.

Type Models Description

pre-2006 1,373 Legacy models migrated from the PDB
non-ModelCIF 1,290 Individual entries with free-text metadata descriptions
Identifier Models Description

ma-bak-cepc 1,106 Yeast protein dimers21
ma-coffe-slac 41,932 Freshwater sponge proteins22
ma-tbvar3d 19 M. tuberculosis complexes related to antibiotic resistance
ma-ornl-sphdiv 25,134 Peat moss (Sphagnum) proteins23
ma-asfv-asfvg 197 African swine fever virus proteins24
ma-t3vr3 957 Human cancer related protein dimers25
ma-low-csi 929 Human protein dimers26
ma-ombbaf2 441 OMBB models for AlphaFold2 benchmarking27
ma-rap-bacsu 167 Bacillus subtilis dimer and trimer complexes28
ma-rap-alink 1,510 Models for AlphaLink benchmarking29
ma-tur-clump 4,224 Human isoforms from RefSeq to assess variants30
ma-kul-lams 55 LN-lamininopathy proteins and complexes31
ma-saps 21 Phytoplasma effector models32
ma-nmpfamsdb 80,585 Novel metagenome protein families33
ma-fesnov 389,522 Functionally & evolutionarily significant novel gene families34
ma-jd-viral 67,715 Eukaryotic virus proteins35
ma-dm-prc 742 Structural prediction screen for PRC complexes36
ma-dm-hisrep 268 Structural prediction screen for histone complexes37
ma-osf-ppp2r2a 273 PP2A-B55 protein phosphatase design38
ma-denv 31 Dengue virus proteins

The oldest models available (ma-c1ewu, ma-cdw44, ma-ceo7o, ma-cwc6z) are from a 1978 paper that generated homology models of relaxin based on a known insulin structure.39 Today, most depositions are de novo models generated with a variant of AlphaFold. The first set of models based on ModelCIF (ma-bak-cepc) was added in late 2021 for a study in which 1,106 core eukaryotic binary complexes were identified and modelled with AlphaFold2.21 The complexes have been deposited in ModelArchive to make them available to the wider community, as the structures could provide insights into the biological function of proteins involved in key processes of eukaryotic cells. As a result, we observe a number of studies that have used these models.4042

The models in ModelArchive are complementary to the AlphaFold database by including protein–protein complexes (ma-bak-cepc, ma-tbvar3d, ma-t3vr3, ma-low-csi, ma-rap-bacsu, ma-kul-lams, ma-dm-prc, ma-dm-hisrep, ma-osf-ppp2r2a), protein–ligand complexes (ma-tbvar3d), viral proteins (ma-asfv-asfvg, ma-jd-viral) and relevant subsets of long proteins (ma-kul-lams). Furthermore, UniProt and therefore the AlphaFold database do not contain all known or possible proteins, and some isoforms, organisms, results of metagenomics studies and designed proteins may be missing. ModelArchive contains models of representatives of metagenomic protein families (ma-nmpfamsdb, ma-fesnov), designed proteins (ma-osf-ppp2r2a) and other proteins missing in UniProt (ma-coffe-slac, ma-ornl-sphdiv, ma-tur-clump). In addition, models can be deposited to support the evaluation of modelling methods, as has been done for the benchmarking of AlphaFold2 on outer membrane beta-barrels (ma-ombbaf2) and for the benchmarking of AlphaLink (ma-rap-alink).

An advantage of ModelArchive and ModelCIF is that it can be used to describe modified modelling pipelines, including those with manual intervention, while still maintaining an interoperable set of results. For example, for a study of viral proteins,35 models of 67,715 eukaryotic viral proteins were deposited (ma-jd-viral). These were generated with ColabFold43 using MMseqs244 to obtain multiple sequence alignments (MSAs), which were used as input for AlphaFold2.1 In addition, a custom search database based on viral proteins in RefSeq was used for most models to improve the MSAs. These aspects can be captured in ModelCIF while maintaining the same type of output and model confidence estimates (i.e. pLDDT, pTM and PAE) as for other models generated with AlphaFold2.

Data distribution

The ModelArchive website (https://modelarchive.org/) acts as the entry point for depositing, browsing and accessing CSMs. Several indexing technologies are used to search and retrieve metadata about them. The primary database is PostgreSQL, which supports accession code lookups and simple text queries for ModelArchive users, as well as searches initiated by the Swiss Bioinformatics Resource Portal, Expasy.45 Once a ModelArchive entry is cited in a journal, we register its Digital Object Identifier (DOI) to make it accessible as https://doi.org/10.5452/ma-xxxxxx. Links between ModelArchive entries and journal articles are also tracked as external links in Europe PMC.

The Python web framework Django sits in front of the database, handling authentication for accessing released and unreleased depositions, and managing user and admin accounts throughout the deposition pipeline. Static content, including items requiring authorisation, is served by Nginx. Django’s REST framework provides metadata for editing and viewing entries. ModelArchive’s architecture was designed to accommodate updates to its front-end framework. It currently uses the flexible and powerful component-based JavaScript library, React.

Data in ModelArchive is distributed to external services (Figure 2), improving visibility and discoverability. It also allows integration with experimental structures and CSMs from other sources thanks to the interoperable data standards. For entries in ModelArchive linked to UniProtKB, metadata and structural coordinates can be retrieved using the 3D-Beacons API.16 This API is used, for example, in the SWISS-MODEL Repository9 and the PDBe-KB.46 The RCSB PDB includes selected ModelArchive models with metadata extracted from ModelCIF. ViralZone links to ModelArchive for viral proteins with available CSMs.47 Foldseek includes ModelArchive models in the BFMD database for structure similarity searches.48 These external services provide additional query capabilities which are particularly useful when CSMs are queried, accessed and combined with experimental data, as is the case in the SWISS-MODEL Repository, PDBe-KB, the RCSB PDB and Foldseek. Further architectural details can be found in the supplementary material.

Figure 2.

Figure 2.

Data distribution showing how entries in ModelArchive are also made available in external services. Interoperable data standards ensure that the models can be readily integrated with structural data from other sources, including experimental structures and computed structure models.

By default, ModelArchive models are available under a permissive licence. Some methods, like AlphaFold 3,49 do not allow this due to restrictive terms of use. We allow such models to be deposited in ModelArchive with the specific terms of use explicitly displayed, but they cannot be accessed via APIs requiring a more permissive licence.

Conclusion and Perspectives

Through its rigorous yet flexible deposition pipeline, ModelArchive provides a robust platform for the preservation and sharing of computationally generated structural models that are not based on experimental data. By adhering to the FAIR principles and emphasising metadata quality, it enables researchers to effectively access, evaluate and reuse these models. By using standardised metadata and formats, ModelArchive enables the seamless integration of computational models into existing research workflows, complementing experimental data.

ModelArchive is a major contributor and user of the ModelCIF format. The ModelArchive team has been actively assisting depositors to convert their data into valid sets of ModelCIF files. In parallel with the existing tool and library support for ModelCIF,15 we are now working on a data harvesting system that will be compatible with PDB-IHM,6 which will allow depositors to interactively generate valid ModelCIF files for individual ModelArchive entries. The ModelCIF format currently covers macromolecular structure predictions. We are working to extend the format to cover more specific details of other modelling applications, including predicted protein–ligand interactions, macromolecular complexes, predictions of different protein conformational states and results of protein design studies.

It will remain a challenge for ModelArchive to keep up with the rapid developments in the field. As the prediction of individual protein chains is already generally very accurate, the focus is shifting to assessing the accuracy of a model in functionally critical regions, e.g. interaction sites with other molecules, as this is relevant for many applications. This in turn requires new metrics for assessing the accuracy and reliability of predictions to be incorporated into ModelCIF and ModelArchive.

Widespread adoption of ModelArchive will require outreach to promote best practices and encourage CSM deposition and we will work with publishers and funding agencies to establish appropriate guidelines. Advances in AI, such as AlphaFold, are making CSMs increasingly accurate and transformative for applications such as protein function determination, drug design, protein engineering and variant analysis. By providing a platform for accessible and reusable models, ModelArchive amplifies the impact of these advances. As computational modelling continues to evolve, ModelArchive will remain critical to fostering reproducibility, collaboration and innovation, ensuring that these models drive progress in both basic and applied life sciences.

Supplementary Material

1

Acknowledgements

The authors would like to thank all the researchers worldwide who have deposited their CSMs in ModelArchive and used CSMs from ModelArchive in their research. We are also grateful to the members of the wwPDB ModelCIF working group for creating a data format suitable for archiving CSMs. We thank the sciCORE centre for scientific computing (https://scicore.unibas.ch) at the University of Basel for providing computing resources and system administration support to run ModelArchive. Icons from https://feathericons.com/ were used for the figures. This work was supported by funding from NIH and National Institute of General Medical Sciences (U01 GM093324–01), ELIXIR (3D-BioInfo), the SIB Swiss Institute of Bioinformatics (https://www.sib.swiss), the Biozentrum of the University of Basel (https://www.biozentrum.unibas.ch), and the swissuniversities Open Science Programme (Open Research Data project “ModelArchive”).

Appendix A. Supplementary material

Supplementary material to this article can be found online at https://doi.org/10.1016/j.jmb.2025.168996.

Footnotes

DECLARATION OF COMPETING INTEREST

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Declaration of Generative AI and AI-assisted technologies in the writing process

During the preparation of this work the authors used ChatGPT and DeepL Write in order to improve the writing in terms of clarity, readability, and language. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

CRediT authorship contribution statement

Gerardo Tauriello: Writing – review & editing, Writing – original draft, Visualization, Supervision, Software, Project administration, Funding acquisition, Data curation, Conceptualization. Andrew M. Waterhouse: Writing – review & editing, Writing – original draft, Visualization, Software, Data curation. Juergen Haas: Writing – review & editing, Supervision, Project administration, Data curation, Conceptualization. Dario Behringer: Writing – review & editing, Software, Data curation. Stefan Bienert: Writing – review & editing, Software, Data curation. Thomas Garello: Writing – review & editing, Visualization, Software, Data curation. Torsten Schwede: Writing – review & editing, Supervision, Project administration, Funding acquisition, Conceptualization.

This article is part of a special issue entitled: ‘Computation Resources (2025)’ published in Journal of Molecular Biology.

References

  • 1.Jumper J, Evans R, Pritzel A, Green T, et al. , (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.wwPDB consortium, (2019). Protein Data Bank: the single global archive for 3D macromolecular structure data. Nucleic Acids Res. 47, D520–D528. 10.1093/nar/gky949. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Berman HM, Burley SK, Chiu W, Sali A, et al. , (2006). Outcome of a workshop on archiving structural models of biological macromolecules. Structure 14, 1211–1217. 10.1016/j.str.2006.06.005. [DOI] [PubMed] [Google Scholar]
  • 4.Schwede T, Sali A, Honig B, Levitt M, et al. , (2009). Outcome of a workshop on applications of protein models in biomedical research. Structure 17, 151–159. 10.1016/j.str.2008.12.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Gabanyi MJ, Adams PD, Arnold K, Bordoli L, et al. , (2011). The structural biology knowledgebase: a portal to protein structures, sequences, functions, and methods. J. Struct. Funct. Genomics 12, 45–54. 10.1007/s10969-011-9106-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Vallat B, Webb B, Fayazi M, Voinea S, et al. , (2021). New system for archiving integrative structures. Acta Crystallogr. Sect. D Struct. Biol. 77, 1486–1496. 10.1107/S2059798321010871. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Berman HM, Adams PD, Bonvin AA, Burley SK, et al. , (2019). Federating structural models and data: outcomes from a workshop on archiving integrative structures. Structure 27, 1745–1759. 10.1016/j.str.2019.11.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Pieper U, Webb BM, Dong GQ, Schneidman-Duhovny D, et al. , (2014). ModBase, a database of annotated comparative protein structure models and associated resources. Nucleic Acids Res. 42, D336–D346. 10.1093/nar/gkt1144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Bienert S, Waterhouse A, de Beer TAP, Tauriello G, et al. , (2017). The SWISS-MODEL repository-new features and functionality. Nucleic Acids Res. 45, D313–D319. 10.1093/nar/gkw1132. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Varadi M, Bertoni D, Magana P, Paramval U, et al. , (2024). AlphaFold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 52, D368–D375. 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lin Z, Akin H, Rao R, Hie B, et al. , (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130. 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 12.Bekker G-J, Kawabata T, Kurisu G, (2020). The biological structure model archive (BSM-Arc): an archive for in silico models and simulations. Biophys. Rev. 12, 371–375. 10.1007/s12551-020-00632-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Tauriello G, Sillitoe I, Bordin N, Orengo C, et al. (n.d.). Your domain: Structural Bioinformatics. https://rdmkit.elixir-europe.org/structural_bioinformatics (accessed November 28, 2024). [Google Scholar]
  • 14.ELIXIR, Research Data Management Kit (n.d.). A deliverable from the EU-funded ELIXIR-CONVERGE project (grant agreement 871075). https://rdmkit.elixir-europe.org. [Google Scholar]
  • 15.Vallat B, Tauriello G, Bienert S, Haas J, et al. , (2023). ModelCIF: an extension of PDBx/mmCIF data representation for computed structure models. J. Mol. Biol. 435, 168021. 10.1016/j.jmb.2023.168021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Varadi M, Nair S, Sillitoe I, Tauriello G, et al. , (2022). 3D-Beacons: decreasing the gap between protein sequences and structures through a federated network of protein structure data resources. GigaScience 11 10.1093/gigascience/giac118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ghafouri H, Lazar T, Del Conte A, Tenorio Ku LG, et al. , (2024). PED in 2024: improving the community deposition of structural ensembles for intrinsically disordered proteins. Nucleic Acids Res. 52, D536–D544. 10.1093/nar/gkad947. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Mariani V, Biasini M, Barbato A, Schwede T, (2013). lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics 29, 2722–2728. 10.1093/bioinformatics/btt473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Zhang Y, Skolnick J, (2004). Scoring function for automated assessment of protein structure template quality. Proteins 57, 702–710. 10.1002/prot.20264. [DOI] [PubMed] [Google Scholar]
  • 20.Basu S, Wallner B, (2016). DockQ: a quality measure for protein-protein docking models. PLoS One 11, e0161879. 10.1371/journal.pone.0161879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Humphreys IR, Pei J, Baek M, Krishnakumar A, et al. , (2021). Computed structures of core eukaryotic protein complexes. Science 374eabm4805. 10.1126/science.abm4805. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Ruperti F, Papadopoulos N, Musser JM, Mirdita M, et al. , (2023). Cross-phyla protein annotation by structural prediction and alignment. GenomeBiol 24, 113. 10.1186/s13059-023-02942-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Gao M, Coletti M, Davidson RB, Prout R, et al. , (2022). Proteome-scale deployment of protein structure prediction workflows on the Summit supercomputer. In: 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE. 10.1109/ipdpsw55747.2022.00045. [DOI] [Google Scholar]
  • 24.Spinard E, Azzinaro P, Rai A, Espinoza N, et al. , (2022). Complete structural predictions of the proteome of African swine fever virus strain Georgia 2007. Microbiol. Resour. Announc. 11, e0088122. 10.1128/mra.00881-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Zhang J, Pei J, Durham J, Bos T, Cong Q, (2022). Computed cancer interactome explains the effects of somatic mutations in cancers. ProteinSci 31, e4479. 10.1002/pro.4479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Bartolec TK, Vázquez-Campos X, Norman A, Luong C, et al. , (2023). Cross-linking mass spectrometry discovers, evaluates, and corroborates structures and protein-protein interactions in the human cell. PNAS 120, e2219418120. 10.1073/pnas.2219418120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Topitsch A, Schwede T, Pereira J, (2024). Outer membrane β-barrel structure prediction through the lens of AlphaFold2. Proteins 92, 3–14. 10.1002/prot.26552. [DOI] [PubMed] [Google Scholar]
  • 28.O’Reilly FJ, Graziadei A, Forbrig C, Bremenkamp R, et al. , (2023). Protein complexes in cells by AI-assisted structural proteomics. Mol. Syst. Biol. 19, e11544. 10.15252/msb.202311544. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Stahl K, Graziadei A, Dau T, Brock O, Rappsilber J, (2023). Protein structure prediction with in-cell photo-crosslinking mass spectrometry and deep learning. Nature Biotechnol. 41, 1810–1819. 10.1038/s41587-023-01704-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Ng JK, Chen Y, Akinwe TM, Heins HB, et al. , (2024). Proteome-wide assessment of clustering of missense variants in neurodevelopmental disorders versus cancer. medRxiv. 10.1101/2024.02.02.24302238. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Kulczyk AW, (2023). Artificial intelligence and the analysis of cryo-EM data provide structural insight into the molecular mechanisms underlying LN-lamininopathies. Sci. Rep. 13, 17825. 10.1038/s41598-023-45200-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Correa Marrero M, Capdevielle S, Huang W, Al-Subhi AM, et al. , (2024). Protein interaction mapping reveals widespread targeting of development-related host transcription factors by phytoplasma effectors. Plant. J. 117, 1281–1297. 10.1111/tpj.16546. [DOI] [PubMed] [Google Scholar]
  • 33.Pavlopoulos GA, Baltoumas FA, Liu S, Selvitopi O, et al. , (2023). Unraveling the functional dark matter through global metagenomics. Nature 622, 594–602. 10.1038/s41586-023-06583-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Rodríguez Del Río Á, Giner-Lamia J, Cantalapiedra CP, Botas J, et al. , (2024). Functional and evolutionary significance of unknown genes from uncultivated taxa. Nature 626, 377–384. 10.1038/s41586-023-06955-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Nomburg J, Doherty EE, Price N, Bellieny-Rabelo D, et al. , (2024). Birth of protein folds and functions in the virome. Nature 633, 710–717. 10.1038/s41586-024-07809-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Shafiq TA, Yu J, Feng W, Zhang Y, et al. , (2024). Genomic context- and H2AK119 ubiquitination-dependent inheritance of human Polycomb silencing. Sci. Adv. 10, eadl4529. 10.1126/sciadv.adl4529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Yu J, Zhang Y, Fang Y, Paulo JA, et al. , (2024). A replisome-associated histone H3-H4 chaperone required for epigenetic inheritance. Cell 187, 5010–5028.e24. 10.1016/j.cell.2024.07.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Kruse T, Garvanska DH, Varga JK, Garland W, et al. , (2024). Substrate recognition principles for the PP2A-B55 protein phosphatase. Sci. Adv. 10, eadp5491. 10.1126/sciadv.adp5491. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Isaacs N, James R, Niall H, Bryant-Greenwood G, et al. , (1978). Relaxin and its structural relationship to insulin. Nature 271, 278–281. 10.1038/271278a0. [DOI] [PubMed] [Google Scholar]
  • 40.Fielden LF, Busch JD, Merkt SG, Ganesan I, et al. , (2023). Central role of Tim17 in mitochondrial presequence protein translocation. Nature 621, 627–634. 10.1038/s41586-023-06477-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Pasquini M, Grosjean N, Hixson KK, Nicora CD, et al. , (2022). Zng1 is a GTP-dependent zinc transferase needed for activation of methionine aminopeptidase. Cell Rep. 39, 110834. 10.1016/j.celrep.2022.110834. [DOI] [PubMed] [Google Scholar]
  • 42.Badonyi M, Marsh JA, (2022). Large protein complex interfaces have evolved to promote cotranslational assembly. Elife 11 10.7554/eLife.79602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Mirdita M, Schütze K, Moriwaki Y, Heo L, et al. , (2022). ColabFold: making protein folding accessible to all. Nature Methods 19, 679–682. 10.1038/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Mirdita M, Steinegger M, Söding J, (2019). MMseqs2 desktop and local web server app for fast, interactive sequence searches. Bioinformatics 35, 2856–2858. 10.1093/bioinformatics/bty1057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Duvaud S, Gabella C, Lisacek F, Stockinger H, et al. , (2021). Expasy, the Swiss Bioinformatics Resource Portal, as designed by its users. Nucleic Acids Res. 49, W216–W227. 10.1093/nar/gkab225. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.PDBe-KB consortium, (2022). PDBe-KB: collaboratively defining the biological context of structural data. Nucleic Acids Res 50, D534–D542. 10.1093/nar/gkab988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.De Castro E, Hulo C, Masson P, Auchincloss A, et al. , (2024). ViralZone 2024 provides higher-resolution images and advanced virus-specific resources. Nucleic Acids Res. 52, D817–D821. 10.1093/nar/gkad946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Kim W, Mirdita M, Karin EL, Gilchrist CLM, et al. , (2024). Rapid and sensitive protein complex alignment with Foldseek-Multimer. bioRxiv2024.04.14.589414. 10.1101/2024.04.14.589414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. , (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES