Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2023 Oct 18;93(1):400–410. doi: 10.1002/prot.26614

Challenges in bridging the gap between protein structure prediction and functional interpretation

Mihaly Varadi 1, Maxim Tsenkov 1, Sameer Velankar 1,
PMCID: PMC11623436  PMID: 37850517

Abstract

The rapid evolution of protein structure prediction tools has significantly broadened access to protein structural data. Although predicted structure models have the potential to accelerate and impact fundamental and translational research significantly, it is essential to note that they are not validated and cannot be considered the ground truth. Thus, challenges persist, particularly in capturing protein dynamics, predicting multi‐chain structures, interpreting protein function, and assessing model quality. Interdisciplinary collaborations are crucial to overcoming these obstacles. Databases like the AlphaFold Protein Structure Database, the ESM Metagenomic Atlas, and initiatives like the 3D‐Beacons Network provide FAIR access to these data, enabling their interpretation and application across a broader scientific community. Whilst substantial advancements have been made in protein structure prediction, further progress is required to address the remaining challenges. Developing training materials, nurturing collaborations, and ensuring open data sharing will be paramount in this pursuit. The continued evolution of these tools and methodologies will deepen our understanding of protein function and accelerate disease pathogenesis and drug development discoveries.

Keywords: AI‐based structure prediction, AlphaFold, functional interpretation, PDB, protein structures, structural bioinformatics


Abbreviations

CAD

contact area difference

CAID

critical assessment of intrinsic disorder

EM

electron microscopy

FAIR

findable, accessible, interoperable, reusable

GDT

global distance test

IDR

intrinsically disordered region

MSA

multiple sequence alignments

NMR

nuclear magnetic resonance

PAE

predicted aligned error

PDB

protein data bank

pLDDT

predicted local distance difference test

PPI

protein–protein interaction

1. INTRODUCTION

The advent of AI‐based protein structure prediction tools, such as AlphaFold 2, 1 RoseTTAFold2, 2 ESMFold, 3 OmegaFold, 4 OpenFold, 5 and UniFold, 6 has brought forth a transformative era in life sciences. 7 , 8 With their unprecedented ability to predict protein structures with remarkable accuracy, these tools have provided a great tool to accelerate research in numerous fields, from drug discovery 9 and structure determination 10 , 11 , 12 to bioinformatics 13 , 14 , 15 , 16 , 17 and synthetic biology. 18 , 19 Undoubtedly, we have stepped into an era where the impact of structural biology has massively broadened, reaching more researchers and domains than ever before. 20 , 21

Although the successes are evident from the notable surge in the adoption of predicted structure models across scientific disciplines, it is easy to overlook their limitations. Experts in structural biology laid out the principles for correctly working with predicted models, the shortcomings and limitations in their application, and how they can be misused. 8 , 22 , 23 , 24 Terwilliger et al. 12 demonstrate how to streamline and expedite the protein structure determination process with predicted models and use them as valuable hypotheses for expanding research questions. Similarly, Lowe 24 expressed that knowledge of a protein structure is rarely a significant obstacle in advancing drug discovery efforts. Instead, it forms the basis for generating insights and guiding the design of novel compounds. Predicted structure models are not meant to serve as replacements for understanding protein function, as several limitations still need to be addressed, even in this new era of structural biology.

While we revel in this new age of abundant structural data, we must acknowledge and grapple with its challenges. Amongst them is the inherent limitation of next‐generation prediction tools in accurately predicting multi‐chain assemblies. Many proteins exist as multimers in their functional state or interact dynamically with other protein molecules, complicating the prediction of their three‐dimensional structure and subsequent functional interpretation. Predicted structure models also lack several components typically associated with a protein and essential for its function or fold. For example, they do not encompass the presence of various ligands they are often associated with, such as DNA, RNA, lipids, ions, and cofactors, thereby limiting their representation in their native state. 25 There is also the absence of co‐ and post‐translational modifications, such as protein glycosylation, phosphorylation, acetylation, and several other covalent modifications. 26

Additionally, a significant fraction of proteins and protein regions are intrinsically flexible, undergoing conformational changes as a part of their function. 27 Most current prediction tools cannot capture this dynamic nature of proteins, often leading to static representations that might not accurately depict their biological reality. Furthermore, modern structure predictors cannot accurately predict mutations' structural effects. 22 This limitation potentially restricts their applicability in areas like disease modeling, where understanding the structural implications of mutations is crucial.

Albeit the extensive challenges and limitations mentioned herein, there is still plenty of room for cautious optimism. It falls on the research community to infer functions from predicted structures. 28 Indeed, protein structures, while serving as powerful tools in our scientific toolbox, are fundamentally coordinates in space. Yet, these structures can become more powerful tools for answering scientific questions with the necessary biological context from domain annotations and molecular context from ligands, cofactors, and ions. 29 While our ability to fully understand and interpret the function of proteins, even those with accurately predicted structures continuously improves, one of the fundamental limitations of current AI‐based tools in structural biology is their inability to provide a comprehensive functional understanding based merely on a structure. Thus, while the predicted structures can help us better grasp protein function within certain limits, a protein's form alone is insufficient. We require additional biological and molecular context layers to tease apart the complex web of protein function.

As we venture further into this exciting era of structural biology, it becomes increasingly crucial to address this functional inference challenge. As a scientific community, we must develop strategies and scalable tools to help us bridge this gap between structure and function. By doing so, we can fully harness the potential of the vast trove of predicted structures, allowing us to gain deeper insights into the intricate workings of life at the molecular level.

One of the most groundbreaking aspects of these advancements is the democratization of protein structure data. AlphaFold, ESMFold and other tools are not only available as software but have also been used to launch new, dedicated databases, such as the AlphaFold Protein Structure Database 30 and the ESM Metagenomic Atlas, 3 which offer open access to a staggering wealth of more than 900 million predicted protein structures, immensely expanding our collective structural biology resource, potentially catalyzing significant strides in our understanding of life at the molecular level.

Another important consideration is that as predicted protein structures reach a wider audience, many researchers new to the field may need assistance with correctly interpreting protein structure data. This consideration emphasizes the need for comprehensive training materials and resources to prevent potential misinterpretations of the models. Establishing these data resources has been a significant effort, and for these to remain relevant and useful to the broader scientific community, they must continue to evolve to meet the needs across diverse research domains.

Finally, from a research data management point of view, another challenge lies in the necessity for a standardized way to access models from different data providers. Given the dominance of resources like AlphaFold DB and ESM Metagenomic Atlas, there is a risk that more minor predictors, which may excel in specific niches, could be overlooked. For example, a suite of tools predicting structures of proteins that make up various components of the adaptive immune system 31 or a curated set of predicted protein structures from the ABC transmembrane protein family 32 are valuable datasets which might be harder to find. A uniform system for accessing and comparing models across platforms could help alleviate this issue, bringing us to the role of initiatives such as the 3D‐Beacons Network, which seeks to provide a centralized platform for accessing protein structure models from various resources. 33 Moreover, the common understanding between these data resource providers on making quality assessment metrics accessible is another consideration that makes this network valuable.

In the unfolding landscape of protein structure prediction, several critical challenges persist that need addressing. This perspective will delve into these hurdles, illuminating potential solutions and directions for further development in the field. Here, we aim to address these challenges, facilitating dialogue and fostering collaborative efforts to navigate this exciting new era of protein structure prediction.

1.1. Predicting multi‐chain structures: Current challenges

Despite the recent advances in single‐chain protein structure prediction, the challenge of accurately predicting multi‐chain structures remains a significant hurdle in structural biology. 34 , 35 Understanding the function of proteins that operate through transient or stable macromolecular interactions necessitates access to quaternary structures. 36 , 37 , 38 , 39 However, the number of experimentally resolved complexes is underrepresented since it is estimated that only 5% of human PPIs are structurally characterized. 40

To tackle this challenge, the research community repurposed tools like AlphaFold2, to predict the structure of multimeric protein complexes, even though it was not initially designed to model such assemblies. 41 , 42 AlphaFold‐Multimer, 43 designed to predict macromolecular complexes, operates at higher accuracy but still lags behind the accuracy of single‐chain models. It was also reported that the accuracy of predicted multimeric complexes declines with an increasing number of constituent structures. 44 This degradation in performance arises from the escalating challenge of discerning coevolution with the addition of more protein chains as it increases the possible pairings of sequences from individual multiple sequence alignments (MSAs). 42 , 44 , 45

It is crucial to emphasize that given the lower accuracy of multi‐chain models, integrating additional experimental data becomes essential for validating these models (Figure 1), which underscores the usefulness of the predicted models in generating hypotheses and the indispensable role of experimental data in validating the hypothesis to enhance our understanding of protein function. By using the combination of experimental data and predicted models, the research community has predicted pairs of proteins on a small scale to structurally visualize protein–protein interactions (PPIs) supported by experimental data. 40 , 44 Other research groups have used predicted models as subcomponents to resolve behemoth assemblies, like the nuclear protein complex, 46 guided by electron microscopy (EM) data.

FIGURE 1.

FIGURE 1

Validation of multimeric predictions using low‐resolution experimental data. This conceptual figure shows that external validation of computationally predicted models is generally preferable, but it becomes crucial when predicting the structure of multimeric assemblies. Experimental data such as small‐angle scattering and cross‐linking provide valuable information on the overall shape and the proximity between specific residues.

A series of studies harness the sensitivity and depth of information from proteomics data; crosslinking and co‐fractionation mass spectrometry overcome the formidable computational scales of structurally characterizing PPIs. 35 Cross‐linking data has proven invaluable for validating predicted assemblies. 47 Tüting et al. (2023) similarly centered on uncovering protein communities and exploring inter‐protein interactions by using mass spectrometry‐derived proteomics data to identify protein assemblies and methodically fit AlphaFold models. 17 The authors have the added advantage over multi‐chain predictors in that they can unambiguously capture near‐native states of protein complexes. Others in this research community have used the same methods in reverse, using databases with experimentally‐supported interactions to identify pairs of interacting proteins, structurally elucidate the interactome, and validate the predictions with crosslinking data. 40 The approaches outlined offer an innovative solution to some of the limitations faced by multi‐chain protein structure tools like AlphaFold‐Multimer. 17 , 47 By utilizing proteomics data for identifying and validating protein communities, they can circumvent some ambiguities that often arise when predicting protein complex structures without prior knowledge.

Similarly, nuclear magnetic resonance (NMR) data has been utilized to corroborate multi‐chain model predictions. Notably, intermolecular contacts identified by Nuclear Overhauser Effect spectroscopy were combined with AlphaFold2‐Multimer predictions to generate an AZUL:UBA model structure. 48

New developments in bioinformatics approaches have the potential to improve the accuracy of multi‐chain model predictions. Earlier methods for pairing interacting partners in MSAs were not yet optimal, and new methods might more reliably determine which potential interacting partners are orthologues that maintain the targeted interaction, compared to paralogues that do not. 49

These advancements in multi‐chain prediction hold substantial practical implications for large‐scale interactome predictions. 50 The AlphaPulldown package, a Python tool for PPI screens using AlphaFold‐Multimer, 51 exemplifies this potential. Given the availability of accurate‐enough multi‐chain models, it offers a means for researchers to identify which proteins may interact based on their respective confidence metrics. In another proof‐of‐concept, Banhos Danneskiold‐Sams E et al. devised a computational approach to predict high‐confidence cell‐surface receptors for various ligands through structural binding prediction, which could greatly expand our understanding of cell–cell communication and holds broad applicability in biology. 52

While the prediction of multi‐chain protein structures poses significant challenges, recent advancements and ongoing research offer hope for substantial progress. By integrating experimental data and refining prediction models, we inch closer toward a better understanding of macromolecular interactions.

1.2. Protein conformational states and flexibility: Unmet needs in structural predictions

AI‐based protein structure prediction tools typically generate static models that reflect a stable conformation of a given protein sequence. Proteins often have more than one distinct conformation, and the dynamic nature of proteins, particularly the alternation between active and inactive states, is crucial for their biological function. 26 Calpain, a calcium‐dependent protease and a potential therapeutic target in cancer, oscillates between distinct conformations depending on their environment or binding status. 53 Other proteins, like hexokinase, adopt different conformations based on their interaction with sugar molecules. In such cases, AlphaFold tends to predict a single conformation. Discerning whether this model represents an active or inactive state is only possible with additional context (Figure 2). In some databases, like PDBe‐KB, 29 it is possible to superpose the AlphaFold models to distinct conformations observed in the experimentally determined structures available in the PDB for a protein of interest.

FIGURE 2.

FIGURE 2

Comparing AlphaFold models to distinct conformations in the PDB.

AlphaFold and most similar new‐generation protein structure prediction tools consistently predict a single conformation. It might be unclear which conformation it is if a protein has more than one distinct, biologically relevant conformation. In the example above, the adenylate kinase‐4 protein adopts significantly different conformations depending on its interaction with ATP. Based on data from the PDB, AlphaFold predicts the ATP‐bound conformation.

Inherently, AlphaFold 2 provides only a single conformation per protein sequence, but according to studies, there might be approaches to alleviate this limitation. A recent study by Stein and Mchaourab 54 presents a novel method using in silico mutagenesis and MSAs to direct AlphaFold 2, facilitating it to yield alternate protein conformations. This approach might imply that by manipulating the input, the model might generate multiple meaningful structures, hinting at its adaptability and potential to represent the inherent flexibility of proteins, although more work is required to validate if the predicted conformations are biologically relevant. In a different approach, using the latest network of RosettaFold2, 2 the work by Stein and Mchaourab 54 raises the idea of combining the right components from different deep learning network architectures. The next generation of these computational methods may switch between predicting highly accurate protein structures with multiple conformations from evolutionarily informative MSAs or responsibly scaling to predict single conformations from large metagenomic datasets. 3 , 55 Such a modular approach could benefit from new developments such as the Distributional Graphormer, a deep learning‐based tool designed to predict the equilibrium distribution of molecular systems, offering a more efficient alternative to traditional, computationally intensive methods like molecular dynamics simulation. 56

An intriguing observation is that the AI‐based new prediction methodologies often predict the conditional fold of proteins that exhibit intrinsically disordered regions (IDRs) in their free state. IDRs can adopt temporary secondary structures upon interacting with their macromolecular partners or under various physiological conditions. Notably, these folded regions often have high confidence scores, with pLDDT values exceeding 70, suggesting a more complex relationship between intrinsic disorder and lower pLDDT scores than previously understood. 57 A recent study has systematically characterized human IDRs concerning pLDDT and its correlation to conditional folding, where 15% are predicted with high confidence. 58 Based on their findings, AlphaFold could accurately identify 88% of intrinsically disordered proteins known to fold upon binding. Their work highlights impressive precision, considering this protein class is underrepresented in the AlphaFold training dataset. Indeed, as demonstrated by the latest critical assessment of intrinsic disorder predictors, although AlphaFold is not the best predictor of intrinsic disorder propensities, its prediction accuracy is already high. The main distinction among state‐of‐the‐art predictors lies in their runtime rather than accuracy. 59 Despite these advancements, accurately modeling intrinsically disordered protein ensembles remains a formidable challenge. The achievement of this goal would significantly impact drug discovery, specifically in the targeting of IDPs with drug molecules based on structure models. 60 This unmet need underscores the continued necessity for development and refinement in structural prediction in the context of conformational flexibility.

1.3. Assessing the quality and reliability of predicted structures

Quality assessment of predicted protein structures is essential to their practical application across a broad range of biological and biomedical research. It is crucial to understand the accuracy of the predictions and the areas of a model where confidence is high or low. 61 , 62 , 63 , 64 , 65 , 66 , 67 , 68 Several metrics, notably the global distance test, 69 template modeling score, 70 and superposition‐independent scores like local distance difference test, 71 residue–residue contact area difference (CAD), 72 and SphereGrinder 73 have been developed to evaluate the quality of these structures by comparing them to experimentally‐resolved structures, providing valuable indications of their reliability.

AlphaFold and other predictive tools generate predicted quality metrics as output, aiding users in interpreting the predicted protein structures. AlphaFold, for instance, provides two metrics: per‐residue confidence scores (pLDDT) and Predicted Aligned Error (PAE). 1 The pLDDT score provides an estimate of the confidence of the network in the individual residue in the predicted model, with scores below 50 often indicating regions with low confidence and those above 90 suggesting a high degree of confidence. 74 On the other hand, PAE estimates the predicted error estimate in the relative position of residue pairs.

However, while these metrics offer valuable tools in the interpretation of the predicted model, they have limitations. They may not accurately predict the level of uncertainty in regions of intrinsic disorder or flexible loop regions, underscoring the need for additional lines of evidence to validate these models. 7 In a large‐scale analysis, Ruff and Pappu 75 outlined that regions with a low pLDDT score average do not always necessarily indicate a failure of AlphaFold, but rather a reflection of the conformational heterogeneity. Later analysis found that this holds for specific categories of intrinsic disorder, such as flexible linkers and entropic chains, but IDRs that adopt a stable secondary structure when bound to an interaction partner are generally predicted with higher (>70) pLDDT scores. 57 In another analysis, Monzon et al. 76 used AlphaFold to predict the structures of short sequences comprising AntiFam, which stores Pfam families for protein sequences derived from spurious open reading frames. The authors initially attempted to confirm these protein sequences did not fold into familiar globular proteins for their quality control purposes. Instead, they discovered a negative correlation between sequence length and average pLDDT score, where shorter sequences displayed higher average pLDDT scores, which might indicate that this potential bias is important to consider when using predicted structures for short sequences. The AntiFam resource may serve as a valuable negative control dataset to assess the accuracy of pLDDT, including metrics from other deep learning methods.

Experimental data are crucial in validating predicted structures. 7 Predictions can be compared with experimental structures determined by X‐ray crystallography, NMR spectroscopy, cryo‐EM, as well as low resolution and specialist techniques such as cross‐linking data and small‐angle scattering information to gauge their accuracy at different levels ( 35 , 77 ; Q. 78 ). These experimental techniques can validate the predictions and offer insights into protein dynamics, providing a level of detail that predictions alone often cannot capture. Integrating multiple lines of evidence when assessing the quality and reliability of predicted structures is thus of utmost importance, with validation by comparing computational predictions to experimental data being the preferred option.

Indeed, while many methodologies are emerging to gauge the quality of predicted protein structures, more must be done in developing tools that evaluate the robustness of these prediction methods. 79 , 80 Techniques that employ adversarial strategies, 81 , 82 commonly utilized in fields like computer vision and natural language processing, 83 , 84 are conspicuously absent in our current arsenal of evaluation techniques.

Revisiting the effective integration of structural proteomics and Alphafold‐Multimer, a significant advancement is integrating structural proteomics data with PAE scores, as visualized in a recently released PAE viewer. 85 Such innovations pave the way for enriching the functional capabilities of visualizing and interacting with predicted protein complexes. Crucially, the introduction of associated crosslinking data exemplifies how to evaluate the reliability of complexes using the PAE plot and, thus, potentially offers a standard for others to adopt.

There is a pressing need for the continued development and standardization of tools and metrics to assess predicted protein structures' quality and reliability and the underlying methods that generate them. 67 , 86 An emphasis should be on adapting current methods and developing new ones for estimating the model accuracy of multi‐subunit protein complexes. 2 , 3 , 4 , 5 , 6 Ongoing efforts in this direction will foster trust in computational predictions and provide the necessary metric for their careful usage across diverse fields, thus extending our understanding of protein structure and function.

1.4. Understanding the consequences of mutations on protein stability and function

The study of how mutations influence protein stability and function carries critical implications across diverse scientific domains, including disease research, drug discovery, and protein engineering. 7 , 8 , 19 Mutations can profoundly affect a protein's properties 87 , 88 and its structure, instigating significant conformational changes and even the manifestation of diseases. 89 , 90

Alterations in protein stability resulting from mutations can disrupt the intricate network of internal contacts, leading to changes in protein folding. These modifications may result in the protein's misfolding, often associated with deleterious outcomes such as neurodegenerative disorders. 91 The functional implications of mutations are just as pivotal. By altering active sites or modifying ligand binding sites, allosteric sites and mutations can significantly affect a protein's interaction with other molecules, causing detrimental functional changes. For instance, a mutation in the active site of an enzyme can disrupt its catalytic activity, impacting the metabolic pathway in which the enzyme is involved. 92 , 93

Computational models generally do not directly support the prediction of the impacts of mutations on protein structures. 1 , 22 However, the Baker group has recently showcased a refined method for mutation impact prediction. 94 They leveraged RoseTTAFold Joint, integrating sequence and structural data, which amplified the model's capacity to help researchers understand the protein mutational landscapes and potentially enhance precision in mutation effect predictions. However, such a perspective alone could be restrictive since single‐point mutations' impact is only sometimes evident from studying the corresponding single protein in isolation. 40 , 87 , 95 However, other tools like FoldX and Missense3D can use predicted protein structures as input and may produce improved predictions in the context of structural stability changes as a consequence of mutations. 96 , 97

An example of a successful application of predicted protein models in characterizing pathogenic mutations with no experimentally‐determined structures or close homologs was demonstrated recently. 95 Among the hundreds of proteins studied, 80% of pathogenic mutations were found close to predicted functional sites, such as ligand‐binding and conserved sites and protein–protein interfaces in predicted structure models. Altogether, by placing missense mutations within a structural context, their pathogenic nature could more effectively be deciphered. This approach underscores the crucial role of structures in understanding missense mutations and highlights the importance of predicted structure models, in the absence of experimentally‐resolved structures. 98 Another example of this approach is AlphScore, 99 a cutting‐edge pathogenicity prediction score derived from AlphaFold. This tool is calibrated on key feature classes, including pLDDT, physicochemical descriptors, amino acid networks, and solvent accessibility. When used with prevalent mutation predictors, AlphScore delivers an enriched, multi‐dimensional perspective of mutation impacts, enhancing our understanding of their consequences on protein function and stability. These approaches highlight the importance of having a rich ecosystem of complementary computational tools to address the limitations of individual methods to help tackle challenging biological problems.

1.5. The role of interdisciplinary collaborations and training in advancing structural biology

In an era where computational and experimental biologists can collaboratively interrogate biological systems, using predicted structures has become an accessible and versatile tool across the life sciences for understanding protein function. Recent advances in accurate protein structure prediction have paved the way for such investigations, even in fields like metagenomics that previously did not significantly utilize structural data. 3

As protein structures permeate a broader range of scientific investigations, supporting new users in navigating this rich resource becomes imperative. It falls on the shoulders of structure prediction tool providers and structure data resources to ensure that users understand the best practices of working with protein structures and know the limitations, which might be unique to specific prediction software or more generic.

A crucial aspect of this education is the understanding and consideration of available confidence metrics associated with modeled structures (Figure 3). As discussed in a previous section, users of AlphaFold models should be well‐versed with metrics such as the pLDDT local confidence metric and the predicted aligned error (PAE) confidence metric. The pLDDT score can help identify regions of the predicted model that are predicted with low confidence, which in many cases represent IDRs but can also reflect limitations in the input data, for example, shallow multiple sequence alignment. The PAE provides insights into the network's confidence in the spatial relationships between residue pairs, which can help identify domains, offering snapshots of possible conformations and suggesting potential flexibility. Both metrics are essential in making informed interpretations and decisions based on predicted structures. Complexities arise in certain situations, such as membrane proteins, where flexible loops may occupy spatial positions that typically represent the membrane location. Understanding such intricacies is fundamental to accurately working with these data.

FIGURE 3.

FIGURE 3

Most frequently used confidence metrics. The most frequently used confidence metrics of the new AI‐based protein structure prediction tools are the pLDDT and PAE scores. The pLDDT scores are local confidence metrics, while the PAE scores give information about the confidence in the relative position of residue pairs.

While some training materials are available from data providers and academic groups, there remains a need for a more centralized, systematic library of structural biology training materials that encompass computational and experimental models and their use.

1.6. The power of data sharing and integration in protein structure predictions

The impact of biological data, including predicted structures, relies significantly on the accessibility and integration of data. Ensuring findable, accessible, interoperable, and reusable (FAIR) access to both computationally‐predicted and experimentally determined protein structures from diverse providers enables the broader scientific community to tackle specific biological questions.

Currently, three primary databases, the ESM Metagenomic Atlas (600million+), the AlphaFold Protein Structure Database (200million+), and SWISS‐MODEL (>2 million) 100 provide predicted structure models. However, more specialized databases exist that can provide more pertinent models for specific problems (Table 1). For example, AlphaFill 25 offers the AlphaFold database models with obligatory ligand molecules. The ModelArchive 101 contains invaluable datasets of predicted proteins, including complexes, while the Protein Ensemble Database 102 offers conformational ensembles derived using experimental data with computational modeling, of intrinsically disordered proteins. The challenge users face is to keep track of these diverse data providers and discover the models that best meet their needs. Initiatives such as the 3D‐Beacons Network significantly democratize access to protein structures by providing FAIR, standardized data access mechanisms for macromolecular structure files from various data providers while also striving to provide standard validation metrics. 33

TABLE 1.

Data resources of predicted protein structures.

Data resources Description

3D‐Beacons Network

https://3d‐beacons.org

Unified access to 220 million + experimentally derived and predicted protein structures and their quality metrics

AlphaFold Protein Structure Database

https://alphafold.ebi.ac.uk

214 million + predicted monomeric protein structures

AlphaFill

https://alphafill.eu/

1 million AlphaFold models with transplanted small molecules

CHESS Human Protein Structure Database

https://www.isoform.io/

230k + predicted human protein structures

ESM Metagenomic Atlas

https://esmatlas.com/

600 million + predicted monomeric protein structures

ModelArchive

https://www.modelarchive.org/

74k + predicted structures

Protein Ensemble Database

https://proteinensemble.org/

300k + ensemble conformations

Protein Data Bank in Europe—Knowledge Base

https://pdbe‐kb.org

80 million + structure‐based annotations

SWISS‐MODEL Repository

https://swissmodel.expasy.org/repository

2.3 million + predicted protein structures

From a broader perspective, it is vital to enriching these structures, whether predicted or experimentally determined, with biological context to enhance their interpretation and utilization. This goal can be achieved by annotating structures with structural, functional, and biophysical attributes, as the PDBe‐KB database aims to facilitate, or by providing services that simplify the transfer of these annotations between models. Several tools assist users with searching and superposing structures from databases such as the Protein Data Bank and the AlphaFold database. Structural clustering of all structure models in the AlphaFold database has been demonstrated as a powerful resource for studying protein function and evolution across the Tree of Life. 15 Foldseek (van Kempen et al., 2023) and DALI offer structure‐based search functionality and, together with other initiatives, help illuminate “functional darkness” in the UniProt and AlphaFold databases, allowing for the exploration of novel protein families and structural folds. 14

2. CONCLUSION

AI‐based protein structure prediction software, notably highlighted by tools like AlphaFold 2, RoseTTAFold, ESMFold, and many others, have ushered in a new era in (structural) biology. These tools can identify previously unknown protein families and folds, thereby broadening our understanding of the evolutionary landscape of proteins. 103

Several studies exemplify how predicted structures can effectively address real biological challenges. For example, in a recent study, Huang et al. generated predicted structure models for the entire deaminase family of proteins using AlphaFold 2 and performed structure‐based clustering. 16 This approach identified new members of existing clades and revealed that many proteins within this family were not traditional double‐stranded DNA cytidine deaminases. Impressively, they were able to engineer one of the minor proteins of the new clade members, enabling successful cytosine base editing in soybean plants, demonstrating the potential of these methods to drive advances in practical applications such as agriculture.

We must, however, acknowledge the challenges that lie ahead. While we now have access to an unprecedented amount of structural data, interpreting the function of proteins based solely on their form is not trivial. There is a critical need for additional context from experimental data, domain annotations, molecular interactions, and the dynamics of protein flexibility. Nonetheless, the structural biology and bioinformatics community, with decades of experience in working with protein structures, is well‐positioned to meet these challenges. The responsibility of structure data providers is to aid researchers across the life sciences to use this rich dataset to accelerate fundamental and translational research.

The advancements in AI‐driven structural prediction mark the onset of a golden age in structural biology. As we navigate this exciting landscape, we remain optimistic about the potential of these tools to revolutionize our understanding of protein structure and function and to advance our ability to address complex biological problems. The future of structural biology is undoubtedly bright, and we look forward to the discoveries it holds.

AUTHOR CONTRIBUTIONS

Mihaly Varadi: Conceptualization; investigation; writing – review and editing; writing – original draft; project administration; supervision. Maxim Tsenkov: Writing – original draft; investigation; writing – review and editing; validation. Sameer Velankar: Funding acquisition; conceptualization.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflict of interest.

PEER REVIEW

The peer review history for this article is available at https://www.webofscience.com/api/gateway/wos/peer-review/10.1002/prot.26614.

ACKNOWLEDGMENT

We would like to acknowledge funding from Google DeepMind. Open Access funding enabled and organized by Projekt DEAL.

Varadi M, Tsenkov M, Velankar S. Challenges in bridging the gap between protein structure prediction and functional interpretation. Proteins. 2025;93(1):400‐410. doi: 10.1002/prot.26614

Mihaly Varadi and Maxim Tsenkov contributed equally to this study.

DATA AVAILABILITY STATEMENT

Data sharing is not applicable to this article as no new data were created or analyzed in this study.

REFERENCES

  • 1. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583‐589. doi: 10.1038/s41586-021-03819-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Baek M, Anishchenko I, Humphreys IR, Cong Q, Baker D, DiMaio F. Efficient and accurate prediction of protein structure using RoseTTAFold2. (p. 2023.05.24.542179). bioRxiv 2023. doi: 10.1101/2023.05.24.542179 [DOI]
  • 3. Lin Z, Akin H, Rao R, et al. Evolutionary‐scale prediction of atomic‐level protein structure with a language model. Science. 2023;379(6637):1123‐1130. doi: 10.1126/science.ade2574 [DOI] [PubMed] [Google Scholar]
  • 4. Wu R, Ding F, Wang R, et al. High‐resolution de novo structure prediction from primary sequence. (p. 2022.07.21.500999). bioRxiv 2022. doi: 10.1101/2022.07.21.500999 [DOI]
  • 5. Ahdritz G, Bouatta N, Kadyan S, et al. OpenFold: retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization. (p. 2022.11.20.517210) 2022. doi: 10.1101/2022.11.20.517210.bioRxiv [DOI] [PMC free article] [PubMed]
  • 6. Li Z, Liu X, Chen W, et al. Uni‐fold: An open‐source platform for developing protein folding models beyond AlphaFold (p. 2022.08.04.502811). bioRxiv. 2022. doi: 10.1101/2022.08.04.502811 [DOI]
  • 7. Akdel M, Pires DEV, Pardo EP, et al. A structural biology community assessment of AlphaFold2 applications. Nat Struct Mol Biol. 2022;29(11):1056‐1067. doi: 10.1038/s41594-022-00849-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Borkakoti N, Thornton JM. AlphaFold2 protein structure prediction: implications for drug discovery. Curr Opin Struct Biol. 2023;78:102526. doi: 10.1016/j.sbi.2022.102526 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Ren F, Ding X, Zheng M, et al. AlphaFold accelerates artificial intelligence powered drug discovery: efficient discovery of a novel CDK20 small molecule inhibitor. Chem Sci. 2023;14(6):1443‐1452. doi: 10.1039/D2SC05709C [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Noone DP, Dijkstra DJ, van der Klugt TT, et al. PTX3 structure determination using a hybrid cryoelectron microscopy and AlphaFold approach offers insights into ligand binding and complement activation. Proc Natl Acad Sci. 2022;119(33):e2208144119. doi: 10.1073/pnas.2208144119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Sleutel M, Pradhan B, Volkov AN, Remaut H. Structural analysis and architectural principles of the bacterial amyloid curli. Nat Commun. 2023;14(1):2822. doi: 10.1038/s41467-023-38204-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Terwilliger TC, Liebschner D, Croll TI, et al. AlphaFold predictions are valuable hypotheses, and accelerate but do not replace experimental structure determination. (p. 2022.11.21.517405). bioRxiv. 2023. doi: 10.1101/2022.11.21.517405 [DOI] [PMC free article] [PubMed]
  • 13. Dobson L, Szekeres LI, Gerdán C, Langó T, Zeke A, Tusnády GE. TmAlphaFold database: membrane localization and evaluation of AlphaFold2 predicted alpha‐helical transmembrane protein structures. Nucleic Acids Res. 2023;51(D1):D517‐D522. doi: 10.1093/nar/gkac928 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Durairaj J, Waterhouse AM, Mets T, et al. What is hidden in the darkness? Deep‐learning assisted large‐scale protein family curation uncovers novel protein families and folds (p. 2023.03.14.532539). bioRxiv. 2023. doi: 10.1101/2023.03.14.532539 [DOI]
  • 15. Hernandez IB, Yeo J, Jänes J, et al. Clustering predicted structures at the scale of the known protein universe. (p. 2023.03.09.531927). bioRxiv 2023. doi: 10.1101/2023.03.09.531927 [DOI] [PMC free article] [PubMed]
  • 16. Huang J, Lin Q, Fei H, et al. Discovery of new deaminase functions by structure‐based protein clustering. (p. 2023.05.21.541555). bioRxiv 2023. doi: 10.1101/2023.05.21.541555 [DOI]
  • 17. Tüting C, Schmidt L, Skalidis I, Sinz A, Kastritis PL. Enabling cryo‐EM density interpretation from yeast native cell extracts by proteomics data and AlphaFold structures. Proteomics. 2023;17:2200096. doi: 10.1002/pmic.202200096 [DOI] [PubMed] [Google Scholar]
  • 18. Jendrusch M, Korbel JO, Sadiq SK. AlphaDesign: a de novo protein design framework based on AlphaFold. (p. 2021.10.11.463937). bioRxiv 2021. doi: 10.1101/2021.10.11.463937 [DOI]
  • 19. Watson JL, Juergens D, Bennett NR, et al. Broadly applicable and accurate protein design by integrating structure prediction networks and diffusion generative models (p. 2022.12.09.519842). bioRxiv. 2022. doi: 10.1101/2022.12.09.519842 [DOI]
  • 20. Demarchi B, Stiller J, Grealy A, et al. Ancient proteins resolve controversy over the identity of Genyornis eggshell. Proc Natl Acad Sci. 2022;119(43):e2109326119. doi: 10.1073/pnas.2109326119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Yates AD, Allen J, Amode RM, et al. Ensembl genomes 2022: An expanding genome resource for non‐vertebrates. Nucleic Acids Res. 2022;50(D1):D996‐D1003. doi: 10.1093/nar/gkab1007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Buel GR, Walters KJ. Can AlphaFold2 predict the impact of missense mutations on structure? Nat Struct Mol Biol. 2022;29(1):1‐2. doi: 10.1038/s41594-021-00714-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Callaway E. What's next for AlphaFold and the AI protein‐folding revolution. Nature. 2022;604(7905):234‐238. doi: 10.1038/d41586-022-00997-5 [DOI] [PubMed] [Google Scholar]
  • 24. Lowe D. Why AlphaFold won't revolutionise drug discovery. 2022. ChemistryWorld. https://www.chemistryworld.com/opinion/why-alphafold-wont-revolutionise-drug-discovery/4016051.article
  • 25. Hekkelman ML, de Vries I, Joosten RP, Perrakis A. AlphaFill: enriching AlphaFold models with ligands and cofactors. Nat Methods. 2023;20(2):205‐213. doi: 10.1038/s41592-022-01685-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Bagdonas H, Fogarty CA, Fadda E, Agirre J. The case for post‐predictional modifications in the AlphaFold protein structure database. Nat Struct Mol Biol. 2021;28(11):869‐870. doi: 10.1038/s41594-021-00680-9 [DOI] [PubMed] [Google Scholar]
  • 27. Evans R, Ramisetty S, Kulkarni P, Weninger K. Illuminating intrinsically disordered proteins with integrative structural biology. Biomolecules. 2023;13(1):124. doi: 10.3390/biom13010124 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Bordin N, Dallago C, Heinzinger M, et al. Novel machine learning approaches revolutionize protein knowledge. Trends Biochem Sci. 2023;48(4):345‐359. doi: 10.1016/j.tibs.2022.11.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. PDBe‐KB consortium . PDBe‐KB: collaboratively defining the biological context of structural data. Nucleic Acids Res. 2022;50(D1):D534‐D542. doi: 10.1093/nar/gkab988 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Varadi M, Anyango S, Deshpande M, et al. AlphaFold protein structure database: massively expanding the structural coverage of protein‐sequence space with high‐accuracy models. Nucleic Acids Res. 2022;50(D1):D439‐D444. doi: 10.1093/nar/gkab1061 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Abanades B, Wong WK, Boyles F, Georges G, Bujotzek A, Deane CM. ImmuneBuilder: deep‐learning models for predicting the structures of immune proteins. Commun Biol. 2023;6(1):575. doi: 10.1038/s42003-023-04927-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Tordai H, Suhajda E, Sillitoe I, Nair S, Varadi M, Hegedus T. Comprehensive collection and prediction of ABC transmembrane protein structures in the AI era of structural biology. Int J Mol Sci. 2022;23(16):8877. doi: 10.3390/ijms23168877 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Varadi M, Nair S, Sillitoe I, et al. 3D‐beacons: decreasing the gap between protein sequences and structures through a federated network of protein structure data resources. GigaScience. 2022;11:giac118. doi: 10.1093/gigascience/giac118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Bouatta N, AlQuraishi M. Structural biology at the scale of proteomes. Nat Struct Mol Biol. 2023;30(2):129‐130. doi: 10.1038/s41594-023-00924-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Träger T, Kastritis PL. Cracking the code of cellular protein–protein interactions: Alphafold and whole‐cell crosslinking to the rescue. Mol Syst Biol. 2023;19(4):e11587. doi: 10.15252/msb.202311587 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Mosca R, Céol A, Aloy P. Interactome3D: adding structural details to protein networks. Nat Methods. 2013;10(1):47‐53. doi: 10.1038/nmeth.2289 [DOI] [PubMed] [Google Scholar]
  • 37. Mosca R, Céol A, Stein A, Olivella R, Aloy P. 3did: a catalog of domain‐based interactions of known three‐dimensional structure. Nucleic Acids Res. 2014;42(D1):D374‐D379. doi: 10.1093/nar/gkt887 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Wang X, Wei X, Thijssen B, Das J, Lipkin SM, Yu H. Three‐dimensional reconstruction of protein networks provides insight into human genetic disease. Nat Biotechnol. 2012;30(2):159‐164. doi: 10.1038/nbt.2106 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Zhang QC, Petrey D, Deng L, et al. Structure‐based prediction of protein–protein interactions on a genome‐wide scale. Nature. 2012;490(7421):556‐560. doi: 10.1038/nature11503 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Burke DF, Bryant P, Barrio‐Hernandez I, et al. Towards a structurally resolved human protein interaction network. Nat Struct Mol Biol. 2023;30(2):216‐225. doi: 10.1038/s41594-022-00910-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Bryant P, Pozzati G, Elofsson A. Improved prediction of protein‐protein interactions using AlphaFold2. Nat Commun. 2022;13(1):1265. doi: 10.1038/s41467-022-28865-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Gao M, Nakajima An D, Parks JM, Skolnick J. AF2Complex predicts direct physical interactions in multimeric proteins with deep learning. Nat Commun. 2022;13(1):1744. doi: 10.1038/s41467-022-29394-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Evans R, O'Neill M, Pritzel A, et al. Protein complex prediction with AlphaFold‐multimer (p. 2021.10.04.463034). bioRxiv . 2022. doi: 10.1101/2021.10.04.463034 [DOI]
  • 44. Bryant P, Pozzati G, Zhu W, Shenoy A, Kundrotas P, Elofsson A. Predicting the structure of large protein complexes using AlphaFold and Monte Carlo tree search. (p. 2022.03.12.484089). bioRxiv 2022. doi: 10.1101/2022.03.12.484089 [DOI] [PMC free article] [PubMed]
  • 45. Bryant P. Deep learning for protein complex structure prediction. Curr Opin Struct Biol. 2023;79:102529. doi: 10.1016/j.sbi.2023.102529 [DOI] [PubMed] [Google Scholar]
  • 46. Mosalaganti S, Obarska‐Kosinska A, Siggel M, et al. AI‐based structure prediction empowers integrative structural analysis of human nuclear pores. Science. 2022;376(6598):eabm9506. doi: 10.1126/science.abm9506 [DOI] [PubMed] [Google Scholar]
  • 47. O'Reilly FJ, Graziadei A, Forbrig C, et al. Protein complexes in cells by AI‐assisted structural proteomics. Mol Syst Biol. 2023;19(4):e11544. doi: 10.15252/msb.202311544 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Buel GR, Chen X, Myint W, et al. E6AP AZUL interaction with UBQLN1/2 in cells, condensates, and an AlphaFold‐NMR integrated structure. Structure. 2023;31(4):395‐410.e6. doi: 10.1016/j.str.2023.01.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Chen B, Xie Z, Qiu J, Ye Z, Xu J, Tang J. Improved the heterodimer protein complex prediction with protein language models. Brief Bioinform. 2023;24(4):bbad221. doi: 10.1093/bib/bbad221 [DOI] [PubMed] [Google Scholar]
  • 50. Schweke H, Levin T, Pacesa M, et al. An atlas of protein homo‐oligomerization across domains of life [preprint]. BioRxiv. 2023. doi: 10.1101/2023.06.09.544317 [DOI] [PubMed] [Google Scholar]
  • 51. Yu D, Chojnowski G, Rosenthal M, Kosinski J. AlphaPulldown—a python package for protein‐protein interaction screens using AlphaFold‐Multimer. Bioinformatics. 2023;39(1):btac749. doi: 10.1093/bioinformatics/btac749 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Banhos Danneskiold‐Samsøe N, Kavi D, Jude KM, et al. Rapid and Accurate Deorphanization of Ligand‐Receptor Pairs Using AlphaFold. bioRxiv: The Preprint Server for Biology, 2023.03.16.531341 2023. doi: 10.1101/2023.03.16.531341 [DOI]
  • 53. Shapovalov I, Harper D, Greer PA. Calpain as a therapeutic target in cancer. Expert Opin Ther Targets. 2022;26(3):217‐231. doi: 10.1080/14728222.2022.2047178 [DOI] [PubMed] [Google Scholar]
  • 54. Stein RA, Mchaourab HS. SPEACH_AF: sampling protein ensembles and conformational heterogeneity with Alphafold2. PLoS Comput Biol. 2022;18(8):e1010483. doi: 10.1371/journal.pcbi.1010483 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Roney JP, Ovchinnikov S. State‐of‐the‐art estimation of protein model accuracy using AlphaFold. Phys Rev Lett. 2022;129(23):238101. doi: 10.1103/PhysRevLett.129.238101 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Zheng S, He J, Liu C, et al. Towards predicting equilibrium distributions for molecular systems with deep learning [preprint]. arXiv. 2023. doi: 10.48550/ARXIV.2306.05445 [DOI]
  • 57. Piovesan D, Monzon AM, Tosatto SCE. Intrinsic protein disorder, conditional folding and AlphaFold2 [preprint]. BioRxiv. 2022. doi: 10.1101/2022.03.03.482768 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Alderson TR, Pritišanac I, Kolarić Đ, Moses AM, Forman‐Kay JD. Systematic identification of conditionally folded intrinsically disordered regions by AlphaFold2. (p. 2022.02.18.481080). bioRxiv 2023. doi: 10.1101/2022.02.18.481080 [DOI] [PMC free article] [PubMed]
  • 59. Necci M, Piovesan D, CAID Predictors , DisProt Curators , Tosatto SCE. Critical assessment of protein intrinsic disorder prediction. Nat Methods. 2021;18(5):472‐481. doi: 10.1038/s41592-021-01117-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Saurabh S, Nadendla K, Purohit SS, Sivakumar PM, Cetinel S. Fuzzy drug targets: disordered proteins in the drug‐discovery realm. ACS Omega. 2023;8(11):9729‐9747. doi: 10.1021/acsomega.2c07708 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Cozzetto D, Kryshtafovych A, Ceriani M, Tramontano A. Assessment of predictions in the model quality assessment category. Proteins. 2007;69(Suppl 8):175‐183. doi: 10.1002/prot.21669 [DOI] [PubMed] [Google Scholar]
  • 62. Cozzetto D, Kryshtafovych A, Tramontano A. Evaluation of CASP8 model quality predictions. Proteins. 2009;77(Suppl 9):157‐166. doi: 10.1002/prot.22534 [DOI] [PubMed] [Google Scholar]
  • 63. Kryshtafovych A, Barbato A, Fidelis K, Monastyrskyy B, Schwede T, Tramontano A. Assessment of the assessment: evaluation of the model quality estimates in CASP10. Proteins. 2014;82(Suppl 2(0 2)):112‐126. doi: 10.1002/prot.24347 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. Kryshtafovych A, Barbato A, Monastyrskyy B, Fidelis K, Schwede T, Tramontano A. Methods of model accuracy estimation can help selecting the best models from decoy sets: assessment of model accuracy estimations in CASP11. Proteins. 2016;84(Suppl 1):349‐369. doi: 10.1002/prot.24919 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Kryshtafovych A, Fidelis K, Tramontano A. Evaluation of model quality predictions in CASP9. Proteins. 2011;79(Suppl 10):91‐106. doi: 10.1002/prot.23180 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Kryshtafovych A, Monastyrskyy B, Fidelis K, Schwede T, Tramontano A. Assessment of model accuracy estimations in CASP12. Proteins. 2018;86(Suppl 1):345‐360. doi: 10.1002/prot.25371 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Kwon S, Won J, Kryshtafovych A, Seok C. Assessment of protein model structure accuracy estimation in CASP14: old and new challenges. Proteins. 2021;89(12):1940‐1948. doi: 10.1002/prot.26192 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68. Won J, Baek M, Monastyrskyy B, Kryshtafovych A, Seok C. Assessment of protein model structure accuracy estimation in CASP13: challenges in the era of deep learning. Proteins. 2019;87(12):1351‐1360. doi: 10.1002/prot.25804 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Zemla A. LGA: a method for finding 3D similarities in protein structures. Nucleic Acids Res. 2003;31(13):3370‐3374. doi: 10.1093/nar/gkg571 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Zhang Y, Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins. 2004;57(4):702‐710. doi: 10.1002/prot.20264 [DOI] [PubMed] [Google Scholar]
  • 71. Mariani V, Biasini M, Barbato A, Schwede T. lDDT: a local superposition‐free score for comparing protein structures and models using distance difference tests. Bioinformatics. 2013;29(21):2722‐2728. doi: 10.1093/bioinformatics/btt473 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Olechnovič K, Kulberkytė E, Venclovas C. CAD‐score: a new contact area difference‐based function for evaluation of protein structural models. Proteins. 2013;81(1):149‐162. doi: 10.1002/prot.24172 [DOI] [PubMed] [Google Scholar]
  • 73. Antczak PLM, Ratajczak T, Lukasiak P, Blazewicz J. SphereGrinder—reference structure‐based tool for quality assessment of protein structural models. IEEE Int Conf Bioinform Biomed. 2015;2015:665‐668. doi: 10.1109/BIBM.2015.7359765 [DOI] [Google Scholar]
  • 74. Tunyasuvunakool K, Adler J, Wu Z, et al. Highly accurate protein structure prediction for the human proteome. Nature. 2021;596(7873):590‐596. doi: 10.1038/s41586-021-03828-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Ruff KM, Pappu RV. AlphaFold and implications for intrinsically disordered proteins. J Mol Biol. 2021;433(20):167208. doi: 10.1016/j.jmb.2021.167208 [DOI] [PubMed] [Google Scholar]
  • 76. Monzon V, Haft DH, Bateman A. Folding the unfoldable: using AlphaFold to explore spurious proteins. Bioinform Adv. 2022;2(1):vbab043. doi: 10.1093/bioadv/vbab043 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. Daccache D, Jonge ED, Liloku P, et al. Evolutionary conservation of the structure and function of meiotic Rec114‐Mei4 and Mer2 complexes. (p. 2022.12.16.520760). bioRxiv 2023. doi: 10.1101/2022.12.16.520760 [DOI] [PMC free article] [PubMed]
  • 78. Zhang Q, Pavanello L, Potapov A, Bartlam M, Winkler GS. Structure of the human Ccr4‐not nuclease module using x‐ray crystallography and electron paramagnetic resonance spectroscopy distance measurements. Protein Sci. 2022;31(3):758‐764. doi: 10.1002/pro.4262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79. Alkhouri I, Jha S, Beckus A, et al. On the robustness of AlphaFold: a COVID‐19 case study. (arXiv: 2301.04093) 2023. doi: 10.48550/arXiv.2301.04093 [DOI]
  • 80. Jha SK, Ramanathan A, Ewetz R, Velasquez A, Jha S. Protein folding neural networks are not robust. (arXiv: 2109.04460) 2021. doi: 10.48550/arXiv.2109.04460 [DOI]
  • 81. Goodfellow IJ, Shlens J, Szegedy C. Explaining and harnessing adversarial examples. arXiv: 1412.6572 2015. doi: 10.48550/arXiv.1412.6572 [DOI]
  • 82. Szegedy C, Zaremba W, Sutskever I, et al. Intriguing properties of neural networks. arXiv: 1312.6199 2014. doi: 10.48550/arXiv.1312.6199 [DOI]
  • 83. Akhtar N, Mian A. Threat of adversarial attacks on deep learning in computer vision: a survey. IEEE Access. 2018;6:14410‐14430. doi: 10.1109/ACCESS.2018.2807385 [DOI] [Google Scholar]
  • 84. Qiu S, Liu Q, Zhou S, Huang W. Adversarial attack and defense technologies in natural language processing: a survey. Neurocomputing. 2022;492:278‐307. doi: 10.1016/j.neucom.2022.04.020 [DOI] [Google Scholar]
  • 85. Elfmann C, Stülke J. PAE viewer: a webserver for the interactive visualization of the predicted aligned error for multimer structure predictions and crosslinks. Nucleic Acids Res. 2023;51:W404‐W410. doi: 10.1093/nar/gkad350 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Pereira J, Simpkin AJ, Hartmann MD, Rigden DJ, Keegan RM, Lupas AN. High‐accuracy protein structure prediction in CASP14. Proteins. 2021;89(12):1687‐1699. doi: 10.1002/prot.26171 [DOI] [PubMed] [Google Scholar]
  • 87. David A, Sternberg MJE. The contribution of missense mutations in core and rim residues of protein–protein interfaces to human disease. J Mol Biol. 2015;427(17):2886‐2898. doi: 10.1016/j.jmb.2015.07.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88. Kucukkal TG, Petukh M, Li L, Alexov E. Structural and physico‐chemical effects of disease and non‐disease nsSNPs on proteins. Curr Opin Struct Biol. 2015;32:18‐24. doi: 10.1016/j.sbi.2015.01.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89. Hartl FU. Protein misfolding diseases. Annu Rev Biochem. 2017;86(1):21‐26. doi: 10.1146/annurev-biochem-061516-044518 [DOI] [PubMed] [Google Scholar]
  • 90. Petukh M, Kucukkal TG, Alexov E. On human disease‐causing amino acid variants: statistical study of sequence and structural patterns. Hum Mutat. 2015;36(5):524‐534. doi: 10.1002/humu.22770 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91. Soto C, Pritzkow S. Protein misfolding, aggregation, and conformational strains in neurodegenerative diseases. Nat Neurosci. 2018;21(10):1332‐1340. doi: 10.1038/s41593-018-0235-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92. Hollingsworth LR, Sharif H, Griswold AR, et al. DPP9 sequesters the C terminus of NLRP1 to repress inflammasome activation. Nature. 2021;592(7856):778‐783. doi: 10.1038/s41586-021-03350-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93. Lei Y, Han P, Chen Y, et al. Protein arginine methyltransferase 3 promotes glycolysis and hepatocellular carcinoma growth by enhancing arginine methylation of lactate dehydrogenase a. Clin Transl Med. 2022;12(1):e686. doi: 10.1002/ctm2.686 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94. Mansoor S, Baek M, Juergens D, Watson JL, Baker D. Accurate mutation effect prediction using RoseTTAFold [preprint]. BioRxiv. 2022. doi: 10.1101/2022.11.04.515218 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95. Sen N, Anishchenko I, Bordin N, et al. Characterizing and explaining the impact of disease‐associated mutations in proteins without known structures or structural homologs. Brief Bioinform. 2022;23(4):bbac187. doi: 10.1093/bib/bbac187 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96. Delgado J, Radusky LG, Cianferoni D, Serrano L. FoldX 5.0: working with RNA, small molecules and a new graphical interface. Bioinformatics. 2019;35(20):4168‐4169. doi: 10.1093/bioinformatics/btz184 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97. Khanna T, Hanna G, Sternberg MJE, David A. Missense3D‐DB web catalogue: An atom‐based analysis and repository of 4M human protein‐coding genetic variants. Hum Genet. 2021;140(5):805‐812. doi: 10.1007/s00439-020-02246-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98. Ittisoponpisan S, Islam SA, Khanna T, Alhuzimi E, David A, Sternberg MJE. Can predicted protein 3D structures provide reliable insights into whether missense variants are disease associated? J Mol Biol. 2019;431(11):2197‐2212. doi: 10.1016/j.jmb.2019.04.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99. Schmidt A, Röner S, Mai K, Klinkhammer H, Kircher M, Ludwig KU. Predicting the pathogenicity of missense variants using features derived from AlphaFold2. Bioinformatics. 2023;39(5):btad280. doi: 10.1093/bioinformatics/btad280 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100. Waterhouse A, Bertoni M, Bienert S, et al. SWISS‐MODEL: homology modelling of protein structures and complexes. Nucleic Acids Res. 2018;46(W1):W296‐W303. doi: 10.1093/nar/gky427 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101. Schwede T, Sali A, Honig B, et al. Outcome of a workshop on applications of protein models in biomedical research. Structure. 2009;17(2):151‐159. doi: 10.1016/j.str.2008.12.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102. Lazar T, Martínez‐Pérez E, Quaglia F, et al. PED in 2021: a major update of the protein ensemble database for intrinsically disordered proteins. Nucleic Acids Res. 2021;49(D1):D404‐D411. doi: 10.1093/nar/gkaa1021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103. Bordin N, Sillitoe I, Nallapareddy V, et al. AlphaFold2 reveals commonalities and novelties in protein structure space for 21 model organisms. Commun Biol. 2023;6(1):160. doi: 10.1038/s42003-023-04488-9 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing is not applicable to this article as no new data were created or analyzed in this study.


Articles from Proteins are provided here courtesy of Wiley

RESOURCES