Skip to main content
Computational and Structural Biotechnology Journal logoLink to Computational and Structural Biotechnology Journal
. 2025 May 16;27:1975–1997. doi: 10.1016/j.csbj.2025.05.009

Multimeric protein interaction and complex prediction: Structure, dynamics and function

Da Lu a,b, Shuhong Yu a,b, Yixiang Huang a,b, Xinqi Gong a,b,
PMCID: PMC12149419  PMID: 40496891

Abstract

Understanding the structure, interactions, dynamics, and functions of multimeric protein complexes is essential for studying multimeric protein complexes, with broad implications for disease mechanisms and drug design, and other areas of biomedical research. Although remarkable achievements have been made in monomer prediction in recent years, protein multimers prediction remains a crucial yet challenging area due to their complex structures, diverse physicochemical properties, and limited experimental data. This review encompasses recent advancements in multimer research, providing an overview of classical concepts and methodologies and the key differences from monomer prediction methods. It further explores state-of-the-art advances in CASP16, including predictions of unknown stoichiometries, supercomplexes, conformational ensembles. This review also delves into the contributions of AlphaFold2 & 3 to multimer prediction, highlighting both the successes and limitations, particularly in handling functional protein-protein interactions and dynamical conformations. Recent deep learning methods and their applications in multimer interaction analysis and quality assessment are discussed, along with insights into future research directions, such as improving prediction accuracy, enabling functional interpretation of protein–protein interactions, and reconstructing protein mechanisms.

Keywords: Protein multimer prediction, Protein dynamics, Protein function, Protein-protein interaction, Quality assessment, Deep learning, AlphaFold2 & 3

1. Introduction

Proteins are organic macromolecules that play pivotal roles in diverse biological processes, including enzyme-catalyzed reactions [1], signal transduction [2], immune responses [3], molecular transport [4], and other essential cellular functions. Proteins are composed of 20 standard amino acids, which are linked by peptide bonds to form polypeptide chains. In 1973, Anfinsen proposed his seminal hypothesis [5], [6] that the sequence of amino acids in a protein determines its primary structure [7], and consequently dictates its secondary structure (alpha-helix, beta-sheet, random coil, etc.), as well as its tertiary structure. Notably, most proteins do not function as individual monomers but rather form multimeric assemblies to carry out their biological roles [8], [9], [10]. The study of protein multimers is fundamental to basic biological research and holds great practical value for drug development and disease treatment.

The research paradigm for protein multimers has evolved from experimental studies to an increasing integration of computational methods. In December 2024, the UniProt database contained 254 million amino acid sequences [11], while the Protein Data Bank (PDB) [12] (https://www.rcsb.org/) has released just over 220,000 protein structures, with approximately 115,000 of these being structures of protein multimers or complexes. Fig. 1 illustrates the annual growth in the number of released protein structures and the multimeric protein structures. The three main experimental techniques currently employed to determine protein structures are: X-ray crystallography, nuclear magnetic resonance (NMR) spectroscopy, and electron microscopy. In recent years, electron microscopy, particularly cryogenic electron microscopy/tomography (cryo-EM/ET), has undergone rapid advancement [13], culminating in the award of the Nobel Prize in Chemistry in 2017 for its role in enabling the resolution of complex biomolecular structures. However, experimental techniques are often resource-intensive and time-consuming endeavors. Therefore, the advancement of computational approaches for accurate multimeric protein prediction has become imperative.

Fig. 1.

Fig. 1

The overall situation of protein structures in the PDB database and the annual deposits of multimeric protein structures is shown. a. The protein structures in the figure are categorized into protein-related, pure protein, and multimeric protein complexes. b. The figure shows the annual deposit count of multimeric protein structures. By combining both figures, it can be observed that the number of protein structures is increasing rapidly, and multimeric proteins account for more than half of the total number of proteins. c. In the figure, Above the timeline are key events and below the timeline are four emerging types of new methods: docking-based, feature-based, end-to-end, and pre-trained models.

In silico methods have become fundamental tools for accurately predicting protein complex structures facilitating experimental data. Indeed, the physicochemical properties of protein multimers are significantly more intricate than those of monomeric proteins [14], owing to their diverse subunit interactions [15] and the dynamic nature of their quaternary structures [16]. Thus, computational methods that reveal the folding, assembly, and dynamic interactions of multimers have emerged as an essential focus in protein structure prediction, with techniques such as protein-protein docking [17], molecular dynamics simulations [18], and energy landscape sampling [19], [20] gradually playing an increasingly important role.

From a physicochemical perspective, protein complexes are composed of separate peptide chains, interacting and docking specifically with one another. Hydrogen bonds, hydrophobic contacts, and short-range van der Waals forces are particularly important for maintaining complex stability [14], [21]. Additionally, long-range non-covalent interactions, such as electrostatic effects (e.g., ππ stacking, salt bridges [22], [23]), ionic bonds [24], and hydrogen bonds [25], are also crucial for the correct assembly and stability of the complex. Currently, novel prediction methods aim to integrate the mentioned physicochemical information with the genetic information of biomolecules.

In the first end-to-end model by AlQuraishi [26], AlphaFold2 [27], AlphaFold-Multimer [28] and other contemporaneous approaches [29], [30], the focus has primarily been on incorporating MSA (Multiple Sequence Alignment) [31] information and residual geometric contacts. Subsequently, some methods have expanded to consider inter-chain geometric and co-evolutionary information [27], [32], [33], [34]. Nevertheless, accurate multimer prediction remains challenging due to factors like structural stability [35], binding affinity [36], and conformational flexibility [37], which are not adequately addressed in monomeric models.

Since the 15th CASP(Critical Assessment of Structure Prediction) [38], protein structure prediction has undergone a gradual shift towards the prediction of protein-protein interactions and the modeling of protein complexes [39]. CAPRI(Critical Assessment of PRedicted Interactions) [40] and CASP have collaboratively orchestrated 6 evaluation experiments, propelling advancements in this domain. Presently, in CASP16 [41], several methods have been proposed to more accurately account for protein-protein interactions and the complex assembly processes in protein complexes.

Research on multimeric proteins holds significant importance across multiple related fields. Understanding and predicting interactions within protein complexes are essential for uncovering biological mechanisms in protein-protein interactions [42]. The interactions between subunits in protein complexes not only influence their stability but also determine their functional roles in cellular processes [17], [43]. In the area of protein structure assessment, the structure of multimeric complexes involves not only the folding of individual subunits but also the precise docking and dynamic interactions between them. Thus, evaluating the structural stability, affinity, and functionality of these complexes is a major challenge for current computational methods [44], [45]. In protein design, the development of proteins capable of self-assembling into specific functional complexes has become a center in synthetic biology and nanotechnology [46], [47], [48]. In drug discovery, designing molecules that target protein-protein interaction interfaces remains challenging due to the inherent flexibility and dynamic nature of these regions, which hinders the development of effective binders [49]. Furthermore, experimental studies have shown that the functional complexity of multimers in processes such as signal transduction [50] and immune recognition [51] exceeds the scope of monomer-based research. The advancement of more accurate multimer prediction and analysis methods holds the potential to drive significant progress in protein design and drug development.

The content of this review is structured as follows: Section 2 provides an overview of the fundamental principles and methodologies in multimer prediction, summarizing the theoretical foundations and application frameworks of current approaches. Section 3 offers a systematic review of the development, applications and key technologies in multimer prediction about structure, function and dynamics. Section 4 discusses the quality assessment of multimer predictions, highlighting the strengths and limitations of existing methods in terms of accuracy and practical applicability. The objective of this review is to provide a comprehensive analysis of advancements in multimeric protein prediction. The review emphasizes the characteristics of multimeric protein prediction, as well as the strengths and limitations of current methodologies. This article reviews the new methods in multimeric protein prediction research related-with AlphaFold, protein function analysis, protein-protein interactions predictions, protein dynamics prediction, the application of related databases, and analyzes the current challenges and future development directions.

2. Principles and foundational methods in multimeric prediction

Protein complexes are multi-subunit structures formed through specific interactions between two or more peptide chains, playing crucial roles in key physiological processes [52]. Within the interface region of protein complexes, multiple hydrogen bonds and various electrostatic interactions can typically be observed, ensuring high affinity and specific assembly [53], [54], [55]. A deeper understanding of the mechanisms behind these interactions enhances our knowledge of protein complex stability and function [56], while also offering a theoretical basis and practical direction for developing computational methods involving molecular docking and assembly [57], [58]. Furthermore, insights into these computational mechanisms can yield key experimental and computational data, serving as valuable references for constructing more accurate biomolecular predictive models. As shown in Fig. 1c, the major events and emerging methods in protein complex prediction are highlighted. While deep learning methods have made significant advancements, classical approaches remain fundamental to multimer prediction, and the features they consider still offer valuable insights for improving prediction accuracy.

Indeed, protein structure prediction encompasses not only the folding and function prediction of monomers but also the docking, assembly, and interaction prediction of polymers. However, protein multimer prediction and monomer prediction differ significantly in many ways, mainly in terms of data differences, prediction targets, kinetic considerations, relative chain positions, and quality assessments. The ensuing discussion will address these discrepancies and their ramifications on prediction methodologies.

2.1. Differences in data for monomer vs multimer prediction

In protein monomer structure prediction, methods such as I-TASSER [59], trRosetta [60], RaptorX [61], FALCON [62], Quark [63], SPARKS [64], and others are representative template-based modeling approaches. In contrast, co-evolution prediction methods are commonly used in the context of complex prediction. However, the availability of experimental data for protein multimer is relatively limited. The primary focus of protein monomer prediction is on the single-chain folding process, wherein the model's emphasis is on the manner in which the amino acid sequence dictates the three-dimensional structure and their stability. In fact, the prediction of multimeric proteins entails a substantial increase in the complexity of the problem. Multimers generally comprise multiple monomers, and their structure prediction may encompass the monomer folding state [65], the assembly state of the complex [66], [67], the interaction interface between monomers [68], spatial symmetry [69], and the dynamic behavior of subunits [67]. The success of multimer prediction hinges not only on the accurate folding of individual monomers but also on the optimization of their relative positions to facilitate binding through suitable interfaces, thereby forming a stable complex [70], [71]. Although recent studies have shown that de novo design of protein binders and complexes is feasible [72], [73], data availability still limits progress for some types of proteins, which include transmembrane or membrane-associated complexes, conformationally flexible proteins, and transient interaction complexes [74], [75], [76], [77], [78].

In comparison with the field of monomer prediction, the paucity of experimental data engenders substantial challenges in directly adapting monomer prediction models to the prediction of polymers. The formation of a polymer is frequently accompanied by substantial conformational changes and adaptive adjustments, underscoring the necessity for the consideration of dynamic interactions between monomers and their flexible properties in the prediction process [67]. These factors not only affect the stability of the complex but may also play a key role in the function of the complex during biological processes. For instance, the complex may undergo adaptive conformational changes in disparate biological environments, and this flexible behavior is critical to its stability and function [79], [80]. Therefore, incorporating the dynamics and flexibility of the complex into the prediction process is imperative. For this reason, a multifaceted approach involving the integration of high-resolution experimental data (e.g. X-ray crystallography and electron microscopy data) and computational methods (e.g., molecular dynamics simulations and coevolutionary modeling) is essential to obtain more comprehensive structural information and to compensate for the paucity of data.

2.2. Sequence-based coevolutionary analysis

In the process of biological evolution, certain residue pairs appear more frequently, and sometimes, when mutations occur, the amino acid positions can be exchanged to maintain structural and functional stability. This phenomenon is known as coevolution. When two residues interact, a multiple sequence alignment of their homologous sequences often reveals that if one position undergoes a mutation, the corresponding position in the other residue is likely to experience a mutation as well. Early coevolution analysis considered the correlations between residue pairs in an unsupervised manner. Some methods establish a full probabilistic model for all residue positions and then attempt to remove the influence of indirect correlations to avoid the limitations of local models. Other models use Markov Random Fields (MRF) to model multiple sequence alignments, learning coevolution information from a set of similar sequences. This approach is commonly known as Direct Coupling Analysis (DCA) [81].

Coevolutionary analysis is one of the important methods in protein complex prediction, which can reveal potential interaction sites and interface regions between protein chains [82], [83], [84]. By studying the sequence mutations patterns in different species, coevolutionary analysis can identify the interdependence between amino acid residues within and outside the protein chain [85]. In the study of protein complexes, unlike the homology information obtained by analyzing a single amino acid chain separately, coevolutionary-based method uses template information of multi-chain interactions by integrating the coordinated changes between the chains in the complex.

In recent years, methods based on deep learning and machine learning have made important progress in the field of coevolutionary analysis. For example, tools such as AlphaFold-Multimer [86], AF2Complex [29] and RoseTTAFoldNA [87] effectively predict the interactions of protein chains by incorporating coevolutionary information, thereby significantly improving the accuracy and efficiency of complex structure prediction. These methods not only predict the single-chain structure of proteins, but also reveal the interactions between different chains, further optimizing the prediction of the three-dimensional structure of complexes. Functional-related protein peptide chains usually show coordinated changes during evolution [88], [89]. Mutations in certain amino acid residues are often closely related to changes in other residues [90]. This coordinated change is usually due to the spatial proximity of these residues and their joint action in protein function.

Through high-throughput sequence comparison and statistical analysis, coevolutionary analysis can identify these interdependent amino acid residues in multiple protein sequences, thereby inferring potential interaction sites and their possible interface regions [91], [92]. When predicting multimers, coevolutionary analysis not only helps to identify potential interaction interfaces, but also provides important supporting data for the construction of contact matrices. For example, by identifying coevolutionary amino acid pairs between different protein monomers, EVcomplex2 [93] can speculate which residues are located in possible interface regions. This information not only helps to optimize the construction of contact matrices, but also provides a reliable basis for predicting the assembly of protein complexes.

2.3. Structure-based contact prediction

Contact matrices are important tools for predicting the assembly of protein complexes [94], which help predict the assembly pattern and stability of the complex by capturing the contact information of amino acid residues between protein monomers. In particular, contact matrices can reveal which amino acid residues are likely to contact each other, thereby supporting further prediction of protein complex assembly. Therefore, contact matrices have been widely used to predict the structure of complex protein complexes and polymers. By analyzing contact matrices, possible contact points can be speculated between protein monomers and provide data support for further structure prediction.

In CASP12, the residue contact prediction algorithm based on residual networks, RaptorX-contact [95], proposed by Xu et al., won first place in the residue contact prediction track, demonstrating the efficient performance of deep learning methods in protein residue contact prediction. In 2018, the team applied RaptorX-contact to inter-chain residue contact prediction for heterodimeric complexes (RaptorX-ComplexContact) [96], using a training set and model from monomer residue contact prediction. They simply concatenated the multiple sequence alignments of the two chains and predicted the contact map of the heterodimer, similar to monomer residue contact prediction. This work employed two MSA concatenation strategies based on phylogenetic trees and genomic distances, a strategy later used in AlphaFold-Multimer [86]. In 2021, following ComplexContact, Xu et al. incorporated protein language model information along with atomic, residue, and surface features to develop Glinter [97], a method for predicting contact maps of protein dimers (including both homodimers and heterodimers). Jianlin Cheng's team designed several algorithms for inter-chain residue contact prediction in heterodimers. One such approach, DeepInteract [98], is a geometry-based deep learning algorithm that leverages protein geometric information derived from monomer structures. Another method, CDpred [99], utilizes attention mechanisms and integrates monomer distance matrices, co-evolutionary analysis, and protein language model features. These two works significantly advanced heterodimer residue contact prediction.

With the development of protein language models, some recent methods have explored the use of embeddings from ESM [100] or MSA Transformer [101] to improve inter-chain residue contact prediction in protein complexes, such as PGT [102] and DRN-1D2D_Inter [103]. For homodimers or multiprotein complexes, where sequence alignment concatenation is not an issue, DeepHomo [104] and its improved version DeepHomo2.0[73], as well as DRcon [105], developed by Jianlin Cheng team, extract MSA information from a single chain and incorporate intra-monomer contact maps or additional protein language model representations to predict inter-chain contacts in homodimers or homomultimers. Compared to heterodimeric complexes, homodimer residue contact prediction generally achieves higher accuracy. This aligns with the broader observation in complex structure prediction that homomeric complexes are easier to predict than heteromeric ones. Due to the high noise in sequence alignments of protein complexes, some methods avoid using MSA-based co-evolutionary information. To address this issue, PDII [106] was proposed, a method based on image inpainting, which relies solely on intra-chain contact maps without requiring multiple sequence alignment (MSA) data. Instead, it learns inter-chain contacts purely from intra-chain interaction information. PDII does not require concatenated MSA data and takes only the internal contact maps of monomeric proteins as input, eliminating the need for additional physicochemical features. Furthermore, this model is not dependent on the format of input structures, demonstrating strong robustness to both bound and unbound protein structures. Additionally, it is capable of handling both homodimers and heterodimers effectively.

In addition to revealing potential contact areas, contact matrices also provide important reference data for further structure optimization and polypeptide assembly studies, which is critical for understanding the stability of polypeptide aggregates [107], assembly patterns [108], and potential interface regions [109].

2.4. Molecular docking and assembly

Molecular docking [110] is a fundamental and critical instrument in the prediction of protein multimer, and it is extensively utilized in the prediction of binding modes of protein complexes. Molecular docking assists in predicting how protein monomers bind and potentially form stable complexes by arranging the relative spatial positions, conformational changes, and interactions between them [111]. Specifically, molecular docking methods principally rely on calculating the binding interface between protein monomers, evaluating the binding energy under different conformations, and screening out the most likely binding modes. The application of energy scoring functions [112] in molecular docking provides insights into the binding mechanisms of protein monomers. The confirmation of the final assembly pattern sometimes necessitates additional simulations, such as molecular dynamics simulations, to ensure the stability of the complex.

Molecular docking methods can be categorized into two distinct classifications: rigid docking [113], [114] and flexible docking [115], [116]. Rigid docking operates under the assumption that the protein monomers remain unchanged during the assembly process. This method is well-suited for protein systems that undergo minor structural changes; however, it may not be applicable to systems with significant conformational changes [117]. Conversely, flexible docking permits the protein structure to undergo conformational changes during assembly, thereby simulating more complex protein multimers. However, this approach entails a substantial computational cost [115], [116].

Mutimer assembly methods center on the manner in which protein monomers aggregate into stable complex through reasonable spatial arrangements, interactions, and potential symmetry requirements [118]. In comparison to the prediction of monomer structure, multimer assembly is a more intricate process that encompasses factors such as the binding mode between monomers, spatial arrangement, and interface optimization. Common multimer assembly methods include simulations based on physical models and algorithm-based optimization methods such as Monte Carlo simulations and molecular dynamics simulations [119]. Monte Carlo simulations explore possible assembly modes through large-scale conformational sampling, while molecular dynamics simulations can be used to evaluate the dynamic changes during assembly and further analyze the stability of different assembly modes.

2.5. Learning-based methods

In recent years, deep learning and machine learning methods have emerged as significant tools in the field of protein complex prediction. The success of AlphaFold 2 represents significant progress, especially the excellent performance in CASP14 highlights the ability of its deep learning-based model in predicting the 3D structures of complex proteins with high accuracy [120]. This advancement not only greatly facilitates the research process in structural biology, but also provides strong support for application areas such as drug discovery and bioengineering. Nevertheless, the model still faces the dependence on single-domain structures and the limitations of cofactor prediction, which are challenges that limit its scope of application in more complex biological systems. To address these issues, the US-SOMO database provides hydrodynamic parameters based on the AlphaFold structure, which provides an important reference for scientists to assess the reliability of the predicted structure in biological solutions [121]. In addition, the AlphaFill algorithm compensates for the lack in structure prediction by integrating experimentally confirmed small molecules into the AlphaFold model, advancing the study of biomolecular functions and thus enhancing the understanding of biomolecular interactions and functions [122]. In the field of drug discovery, the MISATO database provides researchers with abundant data on protein-ligand interactions, which facilitates structure-based drug development efforts, making the design and screening of new drugs more efficient and precise [123]. In recent years, deep learning methods for cryo-EM/ET have also advanced rapidly. Deep learning tools like Topaz are widely used in cryo-EM/ET data processing, such as improving particle picking. Meanwhile, some studies have attempted to use AlphaFold to help build 3D structures from cryo-EM maps. For instance, CryoFEM [124] combines AlphaFold predictions with feature enhancement to support model construction, which offers a possible new strategy for interpreting cryo-EM data.

AlphaFold3 [125] further improves the accuracy of structural modeling of proteins and their interactions, especially in the joint structure prediction of multiple biomolecules, showing its great potential in modeling biologically complex systems. This new model employs an updated diffusion-based architecture, which significantly improves the prediction accuracy of protein-ligand, protein-nucleic acid, and antibody-antigen interactions, and provides new perspectives for our understanding of the complex interactions among these biomolecules. Meanwhile, the combination of deep learning techniques provides new ideas for phosphorylation site prediction, and PhosAF surpasses the accuracy of existing methods by integrating sequence and structural information, thus opening up new paths for target identification and functional modulation in drug design [126]. In terms of conformational prediction, AF-Cluster effectively identifies different conformations of variant proteins by clustering multiple sequence pairs through sequence similarity, providing important insights into understanding the impact of point mutations. These conformational changes not only affect the stability and function of proteins, but may also play a key role in the development of diseases [127]. In the field of protein design, RFdiffusion [47] demonstrates the potential of deep learning for a wide range of applications in designing novel functional proteins and successfully generates a variety of symmetric structures and metal-binding proteins, which provides new ideas for the development of functional proteins. TERMinator and COORDinator, which combine neural networks and Potts models, further enhance the accuracy and reliability of structure-based protein design, proving the importance of computational methods in protein engineering [128]. With the rapid development of bioinformatics, protein language models (pLMs) are redefining the understanding of protein sequences. The protTrans [129] has successfully captured the biophysical features of protein sequences by training autoregressive and autocoder models on large-scale amino acid data, and excelled in the prediction of secondary structure and subcellular localization, which provides new tools and ideas. The newly proposed xTrimoPGLM model realizes the balance between protein understanding and generation tasks through an innovative pre-training framework, demonstrating remarkable versatility and providing a broad exploration space for future research [130].

The algorithm that combines sequence pretraining models with structural prediction modules can also be used for protein structure modeling, as the following examples: ESMFold [131], HelixFold-single [132], OmegaFold [133], trRosettaX-Single [134], and RGN2 [135]. These algorithms demonstrate excellent prediction performance for orphan proteins (proteins with almost no homologous sequences) or artificially designed proteins. The employment of sequence protein structure prediction models has been demonstrated to reduce the time required for constructing multiple sequence alignments, thereby enhancing the efficiency of protein structure prediction. Furthermore, studies have explored the application of protein language models for mutation prediction, including ESM1v [136] and ProtT5 [129]. The Ntranos team [137] employed single-sequence language models to analyze all proteins in the human genome, predicting the effects of approximately 450 million potential missense mutations. Their findings demonstrated excellent performance and showcased potential in areas such as pathogenic mutation prediction, deep mutation scanning analysis, and isoform-specific predictions.

2.6. Protein-protein interaction analysis

Early protein complex inter-chain contact prediction primarily focused on predicting interacting residue pairs [138], which involved assessing whether high-scoring residue pairs (e.g., top 5, 10, or 50) formed interfaces. This includes a series of methods for predicting protein complex interaction residue pairs based on probabilistic models, machine learning, and deep learning. Earlier work [139] considered the distinction between interface and non-interface residues in terms of their physical, chemical, and structural properties, and found that computational and statistical methods could be used to predict interacting residues. In this study [140], the authors proposed several properties of surface residues, such as conservation, amino acid preference, hydrophobicity, and solvent accessibility, and developed three geometric representations of residues: external contact area (ECA) with other residues, external void area (EVA), and internal contact area (ICA). They used statistical models to score residue pairs, pioneering research on predicting interacting residue pairs. Building on this, another study [141] proposed a protein-protein interaction residue pair prediction method that integrates multiple machine learning techniques, improving prediction accuracy compared to other machine learning approaches. Additionally, a series of works utilized Long Short-Term Memory (LSTM) networks to predict interaction residue pairs in heterodimers, trimers, and tetramers [142], [143], [144], [136], [145], [146]. In these methods, geometric features were used to describe residue properties, and the LSTM approach was improved (incorporating attention mechanisms [136], graph neural networks [145], etc.). The methods were tested on trimer and tetramer datasets, providing new insights for research on multimer complex structure prediction algorithms.

The MaSIF framework utilizes geometric deep learning to extract interaction fingerprints from the surface of protein molecules, emphasizing the importance of specific biomolecular interactions, which provides new ideas and methods for understanding the interaction mechanisms between biomolecules [147]. Aiming at the limitations of lattice-based protein representation, the new end-to-end learning framework simplifies the application of deep learning methods in protein science by dynamically computing and sampling molecular surfaces, which significantly improves the performance of interaction site identification and protein-protein interaction prediction, and promotes the further development of structural biology [148]. Effective protein representation learning is crucial for predicting protein function and structure. This proposed 3D structure-based pre-training method and unsupervised comparative learning framework fully demonstrate the importance of geometric deep learning in protein structure representation, which lays the foundation for accurate prediction of protein function and structure [149], [150]. In addition, for the problem of feature capturing of protein structures, novel convolutional operators and hierarchical pooling operators are introduced to specifically deal with the multilevel structures of proteins, which surpass the performance of existing learning algorithms and provide more powerful tools and methods for protein engineering and design [151].

Indeed, protein interaction analysis is a critical step in predicting the assembly of polypeptides. By studying the interaction network and functional connections between proteins, such as STRING [152], PINA [153], it is possible to reveal the interaction sites that play a key role in the protein aggregation process. Beyond mere physical contact between amino acid residues, protein interactions encompass a wide array of biological functions, including regulation of enzymatic activity [154], and spatial positioning within the cell [155], [156]. These interactions influence the assembly and stability of polypeptides by altering the protein conformation or the interaction network.

By accurately identifying and analyzing these interaction sites, the understanding of protein multimer interactions and assembly processes can be enhanced, leading to the formation of stable polymers and the execution of specific biological functions. Notably, these interactions influence not only the assembly process of the polymer but also its functional localization and interactions within the cell. The elucidation of these interactions can offer significant structural insights, including the stability, assembly patterns, and biological functions of protein complexes.

2.7. Classification of quality assessment

The critical step in protein complex prediction is quality assessment and verification, a process that is crucial to ensuring the reliability of the prediction. Conventional evaluation metrics encompass the template score (TM-score) [157], the interface root mean square deviation (iRMSD) [158], and comprehensive evaluation metrics (e.g., DockQ) [44]. TM-score [157] is employed to assess the similarity between the overall model of the complex and the experimental structure, while the iRMSD primarily quantifies the structural accuracy of the complex interaction interface. DockQ [44], a comprehensive metric, integrates the quality of the overall structure and interface to provide a score that assists researchers in evaluating the reliability of prediction results. Meanwhile, it is particularly evident in the context of large-scale and highly complex protein aggregation predictions, where the integration of emerging computational and experimental assessment methods has been shown to enhance prediction accuracy and ensure the scientific validity and credibility of research outcomes.

By employing a comprehensive approach that integrates experimental data and computational models, researchers can identify the most probable polymer models and furnish theoretical underpinnings to support subsequent experimental verification and functional analysis. Moreover, the acquisition of experimental data is frequently constrained by factors such as resolution and experimental methodologies, thereby placing augmented demands on the effectiveness and precision of computational models. Future research directions may include further optimizing existing deep learning models, especially in terms of combining high-quality experimental data to correct and improve prediction results.

3. Multimeric structure, dynamics, and function prediction

The landscape of computational protein prediction has been fundamentally reshaped by the advent of AlphaFold2 [27] and AlphaFold3 [125]. In 2024, the Nobel Prize in Chemistry [159] was bestowed upon John Jumper, Demis Hassabis, and David Baker for their seminal contributions to this domain of research, as exemplified by the development of the computational tools AlphaFold and Rosetta [160], [161]. While significant advancements have been made in the prediction of protein multimer complexes, existing models continue to face challenges. For instance, while AlphaFold demonstrates proficiency in predicting monomeric proteins, its precision remains constrained for large complexes and structures comprising disordered regions [28]. Moreover, incorporating dynamic behavior and functional attributes into structural predictions remains an unresolved and complex challenge. This section will address recent advancements in research as outlined in CASP16, including breakthroughs in protein structure, function, and dynamics. It will also examine the application and limitations of AlphaFold in terms of structure, function, and dynamics. For reader convenience, Table 1 provides a concise summary of some methods discussed in this review.

Table 1.

Softwares for Multimeric Protein Prediction.

Purpose Software Description Last Update URL or source code
Structure Prediction AlphaFold3 [125] Highly accurate biomolecular structure predictions are generated, covering proteins, DNA, RNA, ligands, and ions, along with chemical modifications in proteins and nucleic acids. 2024 https://deepmind.google/technologies/alphafold/alphafold-server/
AlphaFoldDB [299] In collaboration with EMBL's European Bioinformatics Institute (EMBL-EBI), the release of over 200 million freely accessible AlphaFold protein structure predictions provides a broad range of coverage of UniProt entries, including the complete human proteome and those of 47 other key organisms relevant to research and global health. 2025 https://alphafold.com/
AF2Complex [29] AF2Complex is a computational approach that extends AlphaFold to predict protein-protein interactions and higher-order protein complexes from arbitrary protein sequences. 2022 https://colab.research.google.com/github/FreshAirTonight/af2complex/blob/main/notebook/AF2Complex_notebook.ipynb
AF3Complex [300] AF3Complex is a model equipped with the same improvements as AF2Complex, along with a novel method for excluding ligands, built on AlphaFold 3. 2025 https://github.com/Jfeldman34/AF3Complex
ColabFold [301] ColabFold is an open-source platform that accelerates protein structure and complex prediction by integrating the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. 2023 https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/AlphaFold2.ipynb
RosettaFold [87] RoseTTAFold offers a relatively fast and accurate deep learning-based method for protein modeling, featuring an interactive interface that supports custom sequence alignments, constraints, local fragments, and more, with capabilities for modeling multi-chain complexes via paired MSAs or comparative modeling, along with options for large-scale sampling. 2024 https://robetta.bakerlab.org/
PreStoi [162] A novel approach integrating AlphaFold3 predictions, homologous template data, and template-based stoichiometry adjustment. 2025 https://github.com/jianlin-cheng/prestoi
AF_unmasked [171] An enhanced version of AlphaFold integrates experimental data with AI-based prediction, enabling researchers to input partial structural information and refine it through the AI for modeling large, complex proteins. 2025 https://colab.research.google.com/github/clami66/AF_unmasked/blob/notebook/notebooks/AF_unmasked.ipynb
EModelX [165] An automated cryo-EM protein complex structure modeling method that combines AlphaFold with sequence-guided modeling through cross-modal alignments between cryo-EM maps and protein sequences. 2024 https://bio-web1.nscc-gz.cn/app/EModelX
SymProFold [172] A pipeline that combines AlphaFold predictions with general symmetry considerations to predict symmetric protein assemblies 2024 https://github.com/symprofold/SymProFold_code
APPRAISE [173] Automated Pair-wise Peptide-Receptor binding model AnalysIs for Screening Engineered proteins (APPRAISE) is a method that predicts the receptor binding propensity of engineered proteins based on high-precision protein structure prediction tools, such as AlphaFold2-multimer. 2024 https://colab.research.google.com/github/xz-ding/APPRAISE/blob/main/Colab_APPRAISE.ipynb
UM-TBM [30] A hierarchical approach to protein structure prediction and structure-based function annotation, also known as ‘Zhang-Server’ or ‘I-TASSER’. 2024 https://zhanggroup.org/I-TASSER/
MoLPC2 [302] A pipeline for modeling large protein complexes without knowing the stoichiometry using AlphaFold2, and Monte Carlo Tree Search. 2024 https://colab.research.google.com/github/patrickbryant1/MoLPC/blob/master/MoLPC.ipynb
DMFold [182] A deep learning-based approach to protein complex structure and function prediction built on deep multiple sequence alignments. 2024 https://zhanggroup.org/DMFold/
Phenix [183] A comprehensive software package for macromolecular structure determination using crystallographic (X-ray, neutron and electron) and electron cryo-microscopy data, providing access to protein structure prediction with AlphaFold for users of the Phenix software. 2024 https://www.phenix-online.org/
ImmuneBuilder [51] A set of deep learning models trained to accurately predict the structure of antibodies (ABodyBuilder2), nanobodies (NanoBodyBuilder2) and T-Cell receptors (TCRBuilder2). 2024 https://neurosnap.ai/service/Immune%20Builder
DeepMSA2 [181] A hierarchical approach to create high-quality multiple sequence alignments(MSAs) for monomer and multimer proteins. 2024 https://zhanggroup.org/DeepMSA/
ESMPair [175] Identification of interologs of a complex using protein language models. 2023 https://pypi.org/project/fair-esm/
Protein-Protein Interaction Prediction
DeepInter [109] A deep learning framework to predict inter-protein residue-residue contacts of protein complexes by a triangle-aware protein language model. 2023 http://huanglab.phys.hust.edu.cn/DeepInter/
DeepTMP [215] A deep transfer learning framework for predicting the inter-chain residue-residue contacts of transmembrane protein complexes. 2023 http://huanglab.phys.hust.edu.cn/DeepTMP/
ColabDock [17] A general framework adapting deep learning structure prediction models to integrate experimental restraints of different forms and sources without further large-scale retraining or fine tuning. 2024 https://neurosnap.ai/service/ColabDock
DiffPALM [217] DiffPALM (Differentiable Pairing using Alignment-based Language Models) is an AI-based approach that can significantly advance the prediction of interacting protein sequences. 2024 https://github.com/Bitbol-Lab/DiffPALM?tab=readme-ov-file
SpatialPPIv2 [303] An advanced graph-neural-network-based model that predicts PPIs using large language models to embed sequence features and graph attention networks to capture structural information. 2025 https://colab.research.google.com/github/ohuelab/SpatialPPIv2/blob/main/demo/SpatialPPIv2_Colab_Example.ipynb
PF2PI [219] Predicting protein functions within a single species. 2024 https://github.com/cookie-roki/PF2PI
IDBindT5 [221] A novel machine learning (ML) model trained to specifically predict binding regions in intrinsically disordered proteins or regions (IDPs/IDPRs). 2024 https://github.com/jahnl/binding_in_disorder
LazyAF [220] A Google Colaboratory-based pipeline which integrates the existing ColabFold BATCH to streamline the process of medium-scale protein-protein interaction prediction. 2024 https://github.com/ThomasCMcLean/LazyAF
FlyPredictome [86] An online platform that integrates AFM-based predictions with additional information such as gene expression correlations and subcellular localization predictions. 2024 https://www.flyrnai.org/tools/fly_predictome/web/
HyperGraph–Complex [234] A method based on hypergraph variational autoencoder that can capture expressive features from protein sequences without feature engineering, while also considering topological properties in PPI networks, to predict protein complexes. 2024 https://github.com/LiDlab/HyperGraphComplex

Protein-Protein Interaction Database
SAbDab [304] A database containing all the antibody structures available in the PDB, annotated and presented in a consistent fashion. 2025 https://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/sabdab
STRING [305] A database and tool for exploring and analyzing protein-protein interactions and functions. 2025 https://string-db.org/
Biogrid [306] a biomedical interaction repository with data compiled through comprehensive curation efforts. 2025 https://thebiogrid.org/
HLA3DB [238] An auto-curated repository containing the pHLA structures encompassing class I HLA alleles. 2025 https://hla3db.research.chop.edu
PrePPI [239] a structure-informed database of predicted protein-protein interactions. 2023 https://honig.c2b2.columbia.edu/preppi
ABAG-docking [242] A comprehensive antibody-antigen structure dataset for complex docking prediction, named the ABAG-docking benchmark, which contains various docking difficulty-level cases. 2024 https://github.com/Zhaonan99/Antibody-antigen-complex-structure-benchmark-dataset
Predictomes [240] A classifier-curated database of AlphaFold-modeled protein-protein interactions 2025 https://predictomes.org/

Antibody-Antigen Prediction
IgFold [226] A fast deep learning method for antibody structure prediction. 2023 https://pypi.org/project/igfold/
H3-OPT [227] A toolkit for predicting the 3D structures of monoclonal antibodies and nanobodies. 2024 https://github.com/chdcg/H3-OPT
dyMEAN [307] An end-to-end full-atom model for E (3)-equivariant antibody design given the epitope and the incomplete sequence of the antibody. 2023 https://app.tamarind.bio/dymean

Protein Dynamics Prediction PAthreader [255] Predicting multiple protein folding pathways by exploiting evolutionary and folding information from common ancestors. 2024 http://zhanglab-bioinf.com/PAthreader/
Cfold [216] A new version of AlphaFold on the conformational split of the PDB to generate alternative conformations. 2024 https://colab.research.google.com/github/patrickbryant1/Cfold/blob/master/Cfold.ipynb
DynamicBind [257] A deep learning method that employs equivariant geometric diffusion networks to construct a smooth energy landscape, promoting efficient transitions between different equilibrium states. 2024 https://neurosnap.ai/service/DynamicBind
ARTINA [258] A deep learning-based method for NMR spectra analysis, resonance assignment, and protein structure determination by NMR spectroscopy. 2022 https://github.com/PiotrKlukowski/ARTINA
GeoDock [259] A novel multi-track iterative transformer network designed to address limitations in conventional protein-protein docking algorithms and existing deep learning methods. 2023 https://colab.research.google.com/github/Graylab/GeoDock/blob/main/GeoDock.ipynb
AF-cluster [127] A method using clustering of multiple-sequence alignments by sequence similarity to enable AlphaFold2 to sample alternative conformational states of metamorphic proteins with high confidence. 2024 https://colab.research.google.com/github/HWaymentSteele/AF_Cluster/blob/main/AF_cluster_in_colabdesign.ipynb
Protein Dynamics Database
MobiDB [266] A curated biological database designed to offer a centralized resource for annotations of intrinsic protein disorder. 2024 https://mobidb.org/
DisProt [267] A major manually curated repository of intrinsically disordered proteins, both for structural and functional aspects. 2024 https://www.disprot.org/
ATLAS [308] A database of standardized all-atom molecular dynamics simulations, accompanied by their analysis in the form of interactive diagrams and trajectory visualization. 2024 https://www.dsimb.inserm.fr/ATLAS

Estimation of Model Accuracy GCPNet-EMA [185] A geometric message passing neural network referred to as the geometry-complete perceptron network for protein structure EMA. 2024 http://gcpnet-ema.missouri.edu/
ComplexQA [186] A deep graph learning approach for protein complex structure assessment. 2023 https://github.com/Cao-Labs/ComplexQA
DProQA [45] A gated-graph transformer model for end-to-end protein complex structure's quality evaluation. 2023 https://github.com/jianlin-cheng/DProQA
VoloIF-jury [170] A method integrated in software system FTDMP for running docking experiments and scoring/ranking multimeric models. 2024 https://github.com/kliment-olechnovic/ftdmp
af-analysis [187] A python package for the analysis of AlphaFold protein structure predictions. 2025 https://pypi.org/project/af-analysis/
MULTICOM4 [309] An advanced protein structure prediction system built on AlphaFold2 and 3. 2025 https://github.com/BioinfoMachineLearning/MULTICOM4
ModFOLD-dock2 [291] A method includes integrated stoichiometry prediction for quaternary structures and improved sampling and scoring, leading to high performance in continuous independent benchmarks such as CAMEO. 2025 https://www.reading.ac.uk/bioinf/ModFOLDdock/
DockQ [44] A auality measure for protein, nucleic acids and small molecule docking models 2024 https://neurosnap.ai/service/DockQ
DeepUMQA3 [292] A web server for accurate assessment of interface residue accuracy in protein complexes 2023 http://zhanglab-bioinf.com/DeepUMQA_server/

3.1. Multimeric structure prediction

This section focuses on the challenges presented in CASP16 [41] and new advances based on AlphaFold, especially in the areas of multimeric complex prediction, unknown stoichiometric ratios, and the handling of very large complexes. While AlphaFold2 excels in single protein prediction, it still faces difficulties in large-scale complex prediction. AlphaFold3 has made progress in multimeric complex prediction, but dealing with dynamic and flexible structures is still a pressing issue. In addition, AlphaFold2 & 3 provide powerful tools in protein interaction and complex prediction, but how to combine experimental data to improve prediction accuracy is still a current research focus.

3.1.1. CASP16

As the CASP (Critical Assessment of Structure Prediction) competition progresses, CASP16 has presented numerous novel challenges [41]. Firstly, the unknown stoichiometry problem encompasses protein complexes for which the ratio of components is not predetermined, thereby rendering structure prediction more arduous. PreStoi [162] was proposed by Jianlin Cheng et al. that combines AlphaFold3 with homologous template data for stoichiometry prediction, which generates candidate stoichiometries, builds and ranks models using AlphaFold3, and further refines predictions with template-based information when available. In CASP16, PreStoi outperformed others, achieving 71.4% top-1 accuracy and 92.9% top-3 accuracy, demonstrating the complementary strengths of AlphaFold3 and template-based predictions. This approach is suitable for protein complexes lacking stoichiometry data. Secondly, the field has recently identified the structure prediction of super-large complexes as a prominent research priority. Notably, the structure prediction of certain large multi-subunit protein complexes is contingent not only on the availability of high-resolution data, but also on the utilization of more precise algorithmic support. Conformational ensemble prediction [163], [164] is also becoming increasingly important. The selection of the optimal solution from multiple possible conformations remains a significant challenge in structure prediction. The advent of technologies such as AlphaFold has begun to illuminate a path to resolving these issues; however, further integration of theory and experiment remains imperative.

3.1.2. AlphaFold-related computational structure prediction

AlphaFold, a groundbreaking protein structure prediction tool, has made significant progress in recent years. The emergence of AlphaFold 2 has greatly improved the accuracy of protein structure prediction, while AlphaFold 3 has further advanced the prediction of multimer protein complexes [27], [125]. However, challenges remain in predicting multimer proteins based on sequence and structure. Additionally, AlphaFold can be integrated with traditional experimental techniques, such as X-ray crystallography and cryo-EM, to verify and complement the predictions [124], [165]. Case studies highlighting the successes and limitations of this experimental combination offer valuable insights into the potential and bottlenecks of the method. Further work is needed to overcome the current gaps in the model and better integrate computational and experimental approaches.

AlphaFold-Multimer excels at structure prediction of chemical factor and chemical factor receptor complexes, accurately reproducing the binding direction and providing insights into the recognition sites of chemical factors [166]. Another study demonstrated the potential of AlphaFold to handle complexity by combining AlphaFold with symmetry-aware docking simulations to successfully predict the high-precision structure of a complex symmetric assembly [167]. The AlphaFold database (AFDB), which has expanded to include more than 214 million predicted structures, has greatly advanced structural biology and enhanced the usability of related research [168]. By leveraging the AlphaFold database, researchers are able to explore the cross-impact of different fields and address the ensuing data processing challenges.

To improve its success rate in modeling full-length protein sequences, the researchers generated a high-confidence set of protein-protein interaction types to guide the prediction process and constructed a validated DDI reference set to optimize modeling accuracy [169]. In addition, the AF2Complex method, which uses the neural network model AlphaFold2 to directly predict the structure of a protein in a complex without retraining, has performed well in predicting the structure of a complex, especially in the case of pairwise comparison of unrelated sequences, and its accuracy has exceeded that of traditional protein-protein docking strategies [29]. Another new study proposes a prediction protocol that combines model sampling and interface scoring, which significantly improves model accuracy by adapting to multiple sequence pairs and using the VoloIF-jury scoring method, and demonstrates the potential of structure sampling to provide a solid foundation for understanding and applying protein complexes [170]. In addition, methods such as AF_unmasked [171] and EModelX [165] further optimize the prediction capabilities of AlphaFold2, provide methods for integrating experimental data, and enhance the ability to model complex multimers, generating more high-confidence structures. SymProFold [172] also uses the high accuracy of AlphaFold to derive symmetric assemblies, opening up new opportunities for functional exploration. Meanwhile, the APPRAISE [173] has significantly improved the efficiency of protein engineering through structural modeling, quickly evaluating the targeted binding tendencies of engineered proteins. In addition, the improved success rate of AlphaFold2 in molecular replacement method prediction modeling [174] has provided a new perspective for structural biology, demonstrating the importance of AlphaFold in the experimental phase.

Despite the significant progress made by AlphaFold-Multimer, its accuracy still relies on high-quality multiple sequence alignments (MSAs). To this end, ESMPair [175] was proposed, which identifies complex interactors through protein language modeling. The generated MSA outperforms the default method of AlphaFold in low-confidence predictions, with a 10.7% improvement in DockQ's Top-5 best scores. Combining multiple MSA generation methods can improve the accuracy of complex structure prediction by 22%. The proposed ultra-fast end-to-end protein structure prediction method significantly reduces the pre-processing requirements of multiple sequence alignment (MSA), making it possible to achieve large-scale 3D protein modeling with lower hardware requirements [176].

AlphaFold3 [125] introduces a diffusion-based architecture that significantly improves the accuracy of structure prediction for proteins, nucleic acids, and small-molecule complexes, especially surpassing existing tools in protein-ligand and antibody-antigen prediction. In addition, an AI-guided pipeline that combines experimental and computational tools was proposed [177], with the aim of identifying and validating protein–protein interaction (PPI) targets and accelerating early drug discovery. High-confidence interactions of SARS-CoV-2 were discovered by analyzing PPI assay data using machine learning, and a compound that effectively inhibits the interaction of the NSP10-NSP16 complex was discovered by performing large-scale virtual drug screening using VirtualFlow [178]. These advances have not only promoted research on the prediction of protein complex structures, but also provided new tools and strategies for drug discovery.

Several new features and optimizations were introduced in CASP15 [40], significantly improving the prediction of important polymeric proteins. In the CASP15-CAPRI experiment, about 40% of the targets produced high-quality models, a significant improvement over the 8% two years ago. This progress is largely attributed to the widespread use of AlphaFold2 and its variants, as well as in-depth analysis of model quality. The PEZYFoldings team excelled in the single-domain protein category, improving the accuracy of protein structure prediction using an enhanced sequence similarity search and deep learning optimization model [179]. The method emphasizes the importance of evolutionarily related sequences and points out the contribution of human intervention in the final result, especially in the analysis of complex subunit interfaces. Meanwhile, the Wallner group has significantly improved the quality of the prediction of the structure of oligomers by applying multiple sampling and discarding layer strategies to AlphaFold, successfully generating 274,289 models, and the proportion of high-quality models has increased significantly [180]. The UM-TBM and Zheng teams have significantly improved the accuracy of protein monomer and complex structure prediction using a new multiple sequence alignment generation protocol and spatial constraints based on an attention network [30]. In addition, the DeepMSA2 pipeline significantly improves the prediction accuracy of protein tertiary and quaternary structures using iterative alignment searches of genome and metagenome sequence databases [181]. The DMFold [182] algorithm has been used to successfully construct reliable four-level coevolution structures, advancing research on the structural prediction of protein complexes. Meanwhile, the Phenix software suite integrates AlphaFold predictions, allowing automatic model construction and optimization from amino acid sequences and data [183]. It has been shown that deep learning methods like AlphaFold2 perform superior in most cases, but are still insufficient when predicting non-helical secondary structures and solvent-exposed peptides [184]. The ImmuneBuilder model has demonstrated its potential in predicting the structure of immunoproteins, providing new ideas for the biotherapeutic application of antibodies and nanobodies [51].

With the development of tools such as deep learning, generative models, and graph neural networks, new methods have emerged to improve the accuracy and assessment quality of protein structure prediction. GCPNet-EMA(Geometric Completeness Perceiver Network) uses geometric information to improve the accuracy estimation of protein structure models through a geometric message passing neural network. Studies have shown that it is 47% faster than existing three-layer structure assessment methods and improves accuracy by 10% [185]. DProQA is a novel gated neighborhood modulation mapping transformer for assessing the quality of the three-dimensional structure of protein complexes [45]. It performed well in the CASP15 experiment and became the third-best single model quality assessment method. For the quality assessment of interface residues, ComplexQA [186] is based on a deep graph neural network that combines information about the order of residues in three-dimensional space to provide a local quality score. Test results show that the method outperforms other existing methods on multiple high-quality structure modeling datasets, especially on difficult targets. In addition, VoloIF-jury [170] proposes a new protein complex structure prediction protocol that targets sub-unit interaction interfaces through model sampling and scoring, achieving significantly better results than the standard AlphaFold-Multimer pipeline in CASP15, emphasizing the importance of improving structure sampling. To address the limitations of AlphaFold-Multimer in polypeptide prediction, a new scoring method, pDockQ2 [187], was proposed to evaluate the quality of each interface in the polypeptide. The research analysis highlighted the differences in the performance of some complexes on different evaluation metrics on datasets with reduced homology. The introduction of the scoring function mpDockQ [188], combined with AlphaFold and Monte Carlo tree search, can effectively predict the structure of large protein complexes and evaluate the completeness of the assembly and the accuracy of the prediction, especially showing high accuracy on symmetric complexes.

3.1.3. Experimental analysis of AlphaFold-based applications

In the study of protein interactions and complex structure prediction, the combination of continuous improvement of AlphaFold and experimental verification is driving scientific progress. For example, the success of AlphaFold2 has opened up new possibilities for the application of molecular replacement (MR) methods in crystal structure resolution. Most structures can be solved by AF2 predictions, although there are still some cases that require experimental phasing methods [174]. The integration of AlphaFold's structure prediction with biophysical experiments facilitates the systematic identification and validation of protein interaction interfaces, thereby providing novel targets for drug discovery [189]. The implementation of the AF_unmasked [171] method overcomes the limitations of AlphaFold in large-scale complex prediction and effectively integrates experimental information to generate high-quality structures.

AlphaFold-Multimer and its derivatives have been utilized in numerous studies to elucidate the intricacies of biomolecules and their functions. An X-ray crystal structure of an anti-MHC-I monoclonal antibody has been employed to discern the epitope, underscoring the constraints of computational models in accurately identifying antibody binding sites [190]. Similarly, in the study of bacterial proteins, AlphaFold-Multimer effectively captured multiple aggregation states, highlighting the underestimation of protein structure complexity by traditional prediction methods [191]. Further studies have elucidated the synergistic role of DNA-PK and TRF2 in telomere maintenance, and the AlphaFold model has yielded novel insights into the interaction with Rad50 [192]. In addition, the AlphaFold-predicted structure of the mitotic-specific kinase Mek1 in yeast revealed the specific binding of the acidic loop to the substrate, thereby verifying its functional importance [193], while cryo-EM structural analysis further demonstrated the dynamic mechanism of these complexes, such as the role of the Mnx protein complex in manganese biomineralization [194].

Cryo-EM/ET is also important in protein structure determination. However, current automated cryo-EM modeling methods often rely on prior chain separation, which can lead to the accumulation of errors [165]. Therefore, EModelX has successfully improved the accuracy of protein complex modeling to an atomic level by performing cross-modal alignment of cryo-EM maps and protein sequences for sequence-guided fully automated modeling [165]. DeepTracer-Refine enhances the accuracy of cryo-EM structure prediction, automates the refinement of AlphaFold structures, and improves modeling efficiency [195]. Finally, AlphaFold-Multimer significantly improves the performance of protein complex prediction through denoising represented by multiple sequence alignment [196]. The combination of these studies provides important insights into understanding the function and dynamics of proteins.

In the field of yeast biology, research endeavors have elucidated the significance of conserved contact motifs in the assembly of the tetrameric Schizothoracin complex. It has been instrumental in refining the octameric model through the utilization of AlphaFold-Multimer [197] and it has demonstrated that the advent of AlphaFold has expedited the systematic identification of protein interaction interfaces. In this study, researchers employed AlphaFold to generate predicted protein complex structures and integrated these predictions with biophysical experiments, such as surface plasmon resonance and yeast two-hybrid, to validate the interactions [189]. By screening an important set of protein-protein interaction pairs and analyzing the structures generated by AlphaFold, the team successfully identified potential interaction interfaces, thereby demonstrating the high accuracy of AlphaFold in predicting protein interaction interfaces. Furthermore, the study demonstrated the capability of AlphaFold-Multimer in predicting interactions between plant pathogens, offering novel insights into the comprehension of plant defense mechanisms [198]. In a study on the Kv2.1 potassium channel, AlphaFold-Multimer constructed a tetramer model and analyzed the binding site of the inhibitor, thereby providing a theoretical foundation for the design of novel Kv2.1 inhibitors [199]. These studies demonstrate that AlphaFold-Multimer not only deepens the understanding of biomolecular structures, but also promotes the development of drug design and biotechnology.

3.1.4. Limitations and problems in AlphaFold-based modeling

AlphaFold has made significant progress in the field of protein structure prediction, but some limitations still exist. First, while its predictions approach experimental accuracy for single-chain proteins, it performs poorly in modeling multimeric structures, especially in antibody-antigen and other complex multichain complexes, with low success rates [78], [187]. Second, AlphaFold's prediction results failed to fully consider environmental factors such as ligands and covalent modifications, which led to some high-confidence predictions that differed significantly from the experimental results in main chain and side chain conformations [200]. In addition, although deep learning methods such as AlphaFold2 perform well in many cases, they are less accurate in the prediction of non-helical secondary structures and solvent exposure of peptides, suggesting that their adaptation to peptide structures still needs to be improved [184]. Finally, the presence of topologically entangled linkage structures in complexes predicted using AlphaFold-Multimer shows the shortcomings of current models when dealing with complex topologies [201]. These limitations emphasize the importance of treating AlphaFold predictions as hypotheses rather than final results, and suggest that experimental validation be combined to confirm structural details [200].

3.2. Protein function prediction

The revolutionary progress of AlphaFold has significantly changed the field of protein structure prediction, especially in the prediction of the three-dimensional structure of monomeric proteins, which has achieved unprecedented success. However, with the deepening of research, the prediction of protein-protein interactions (PPI) and protein complex structures has become a key area of concern [202]. Driven by AlphaFold2 and its extended versions, the research on PPIs and protein complexes has made initial progress, but these fields still face many challenges.

3.2.1. Protein functional prediction and interaction analysis

There is a close link between the functional prediction of proteins and the analysis of their interactions [203], [204], [205]. The realization of protein function not only depends on its three-dimensional structure, but is also deeply influenced by its interaction network in the cell. In the cell, proteins usually do not function alone, but interact with other proteins or molecules to form functional complexes or participate in complex biological pathways [206]. For example, biological processes such as enzyme-catalyzing reactions [207] and gene regulation [208] involve protein interactions. Therefore, accurately predicting the function of a protein requires considering both its three-dimensional structure and its interactions with other biomolecules, especially other proteins.

With the advancement of genomics research, researchers can obtain a large amount of protein information from sequencing data, which provides rich input data for functional prediction. Classical functional prediction methods include homology sequence alignment [209] and structural alignment [210]. These methods usually infer the function of a new protein based on known functional or structural domains. However, with the improvement of computing power and the increase of data volume, modern functional prediction methods have begun to use machine learning and deep learning methods to train models to identify potential functional markers or patterns. These methods can process larger volumes of data and mine patterns in them to improve the accuracy of predictions.

Multimeric protein complexes with interactions not only have complex spatial conformations, but also involve many different types of interactions, such as protein-protein, protein-nucleic acid, and protein-lipid interactions [211], [212]. These interactions not only affect the stability of the protein [213], [214], but may also regulate its activity, location distribution, and function. Analyzing the interactions of multimer proteins provides a more detailed perspective when understanding complex biological processes. Technological advances in the study of multimer protein complexes (such as cryo-electron microscopy, X-ray crystallography and nuclear magnetic resonance) have enabled structural biologists to obtain higher-resolution structural data on the complexes, which has promoted a deeper understanding of their functions.

3.2.2. AlphaFold-related interaction prediction

With the release of AlphaFold2, researchers began to try to use its extended version, AlphaFold-Multimer, to predict the structure of protein complexes. The three-dimensional structure prediction of small protein complexes has become more accurate. However, when faced with more complex interaction systems, such as those involving multiple subunits or transmembrane protein complexes, there is still room for improvement in multimeric protein interaction prediction.

In the field of protein complex structure prediction, the introduction of deep learning methods in recent years has significantly improved the accuracy and robustness of predictions. For example, the DeepInter method successfully predicts residue-residue contacts in protein complexes through triangulation updating and self-attention mechanisms [109]. Similarly, DeepTMP effectively captures inter-chain contacts in membrane protein complexes by relying on deep transfer learning and pre-trained knowledge of non-transmembrane protein datasets, further improving the prediction accuracy [215]. Protein-ligand docking is a common tool in drug discovery and development for screening potential therapeutic drugs. Successful docking usually relies on high-quality protein structures, which are often considered to be fully or partially rigid structures [216]. To address this problem, researchers have developed an AI system that can directly predict the all-atom flexible structure of protein-ligand complexes from sequence information. The predicted confidence index (plDDT) can help distinguish between strong and weak binding ligands. In addition, ColabDock [17], as a general framework, can integrate experimental constraints from different sources to optimize structure prediction, demonstrating its high compatibility and accuracy between simulated and experimental data.

To better understand the structural prediction of protein complexes, the DMFold algorithm enhances the biomedical usefulness of the prediction by constructing accurate multiple sequence alignments (MSAs) and deriving interchain coevolutionary information [182]. This model [19] based on the interface energy distribution can predict the co-translational assembly pathway, which provides a new perspective for understanding the dynamic process of protein-protein interactions. The DiffPALM method uses masked language models, especially in AlphaFold-Multimer, to further improve the accuracy of predictions by pairing interacting protein sequences [217]. The SpatialPPI uses structural information generated by AlphaFold Multimer demonstrates the potential of deep learning in processing three-dimensional structural details [155]. AlphaPulldown is a Python package designed specifically for PPI screening. It simplifies the process of high-throughput modeling using AlphaFold-Multimer, is especially suitable for higher-order oligomers, and provides a friendly command line interface that integrates multiple confidence scores and graphical analysis functions [218]. The PF2PI method provides a new idea for protein function prediction by integrating AlphaFold2 data and PPI networks, significantly improving the effectiveness of function prediction and verifying its effectiveness [219]. The emergence of the LazyAF pipeline [220] has facilitated medium-scale predictions. This tool integrates the ColabFold BATCH software, simplifies the process of predicting protein interactions, and demonstrates its advantages in terms of accessibility and ease of use. Using LazyAF, the interaction set of 76 proteins encoded on the multi-drug resistant plasmid RK2 was successfully predicted, further verifying the effectiveness [220]. To identify protein binding sites in disordered regions, the researchers proposed a new machine learning model IDBindT5 [221], which focuses on predicting binding sites in intrinsically disordered protein regions. The model achieves a balanced accuracy of 57.2% using the embedding of the protein language model ProtT5 and is significantly superior to other advanced methods in terms of prediction speed. The progress of IDBindT5 not only demonstrates the potential of machine learning in analyzing the characteristics of disordered proteins, but also provides new tools and perspectives for whole proteome analysis. In short, these research advances not only promote in-depth theoretical research on the prediction of protein complex function, but also open up new paths for the practical application of protein interaction tools.

Although AlphaFold performs well in structure prediction, it still faces challenges in some cases, especially when the binding partner undergoes significant conformational changes. For this reason, AlphaFold could be combined with a physics-based replica exchange docking algorithm to improve the prediction accuracy of protein complexes. They found that integrating AlphaFold's confidence information can effectively estimate the flexibility of proteins and docking accuracy, and this method performs well when dealing with complex docking tasks [222]. In the study of protein-peptide interactions, understanding the molecular details between peptides and proteins is critical to the regulation of biological processes. By using AlphaFold-Multimer, researchers have improved peptide-protein docking and significantly improved the accuracy of docking predictions using a forced sampling method. This study [223] found that AlphaFold-Multimer successfully predicted 66 of 112 peptide-protein complexes with a quality that met or exceeded acceptable levels. More importantly, by randomly perturbing the weights of the neural network and forcing the network to explore more conformational space, the number of acceptable models was increased from 66 to 75, demonstrating the potential and flexibility of AlphaFold in peptide-protein docking.

Concurrently, accurate modeling of antibody-antigen complexes remains challenging due to two probable limitations. First, the scarcity of high-resolution experimental structures, especially for the highly variable CDR-H3 region, hampers precise modeling efforts [224]. Second, the lack of strong evolutionary constraints in antibody-antigen interactions reduces the effectiveness of MSA-based methods [225]. As a result, methods such as AlphaFold3, IgFold [226], ImmuneBuilder [51], and H3-OPT [227] still struggle to achieve application-level accuracy in CDR-H3 modeling, despite substantial progress in deep learning-driven feature extraction.

3.2.3. Protein-protein interactions predictions and analysis

Advances in protein-protein interaction (PPI) research, where understanding and identifying the function and composition of protein complexes is critical for life processes, disease diagnosis and drug development, have provided new perspectives and methods for drug discovery and protein engineering. However, there are many limitations to traditional experimental methods, making the development of effective computational methods to predict protein complexes all the more necessary. The key to PPI prediction lies in determining the interaction interface, affinity and function of proteins. How to accurately identify interaction partners [228] and incorporate protein dynamics [229] into computational prediction is one of the core challenges in the field of PPI research.

AlphaFold2 shows superiority in the structural prediction of heterodimeric protein complexes. By combining multiple sequence alignments, AlphaFold2 [27] generates acceptable quality models for 63% of the dimers and proposes a simple function to predict the DockQ score. AlphaFold2 can effectively distinguish interacting and non-interacting proteins and improve the accuracy of interaction identification. Meanwhile, analyzing the conformational stability of interacting residues in the binding interface reveals the importance of residue-residue interaction stability characteristics for improving the accuracy of protein-protein binding predictions. Analysis of 574 protein complexes shows that pre-stabilized residue interactions are characteristic of interface selectivity, a finding that may help in the future design of more effective binding proteins [230]. Furthermore, when using AlphaFold-Multimer [86] for in-depth protein-protein interaction studies, the researchers found that traditional metrics may have overlooked interactions in small interfaces or flexible regions. To this end, they introduced the local interaction score (LIS), which is based on the predicted alignment error (PAE) and can more sensitively detect protein-protein interactions, especially when flexibility and small interfaces are involved. By applying LIS to a large-scale Drosophila dataset, the researchers significantly improved the detection of direct interactions and launched the online platform FlyPredictome [86], which integrates AlphaFold-Multimer-based predictions with other bioinformatics to provide new insights into protein-protein interaction research. DeepInter [109] is a deep learning-based protein language model designed to predict residue contacts in protein complexes, taking into account residue contacts. By introducing a triplet perception mechanism, the model shows high accuracy and robustness on multiple datasets, especially in the prediction of heterodimeric protein complexes, overcoming the limitations of traditional methods.

As a new protein structure prediction software, AlphaFold3 has demonstrated its potential through improved protein-protein complex prediction. Studies have shown that AlphaFold3 has a correlation of 0.86 with the state-of-the-art model MT-TopLap when predicting the change in binding free energy caused by mutations, although it is slightly lower than the results using traditional PDB structures. This finding highlights the limitations of AlphaFold3 under specific conditions, especially when dealing with regions with inherent flexibility, where some structures have significant prediction errors [231]. In addition, predicting residue contacts is important for understanding the structural modeling of protein complexes. In terms of predicting interaction partners, MSA Pairing Transformer [232] effectively improves the accuracy of pairings between known interacting protein families through fine-tuned contrastive learning techniques. This method in particular shows superior performance when there are few paired sequences, and the resulting multiple sequence alignment contains a stronger coevolution signal, providing new possibilities for the coevolution analysis of protein-protein interactions.

In the discovery of protein-protein interactions, researchers used AlphaFold-pairs [233] to accurately map a large number of direct interactions between humans and viruses. AlphaFold-pairs successfully identified a variety of specific interactions, including the SARS-CoV-2 surface glycoprotein Spike, by evaluating the physical proximity between proteins. The results show that AlphaFold-pairs has great potential for future protein-protein interaction research [233]. The researchers proposed a framework called HyperGraphComplex [234], which is based on hypergraph learning. This framework combines protein sequences and PPI network topology, and can effectively capture expression features without feature engineering. Experimental results show that HyperGraphComplex outperforms many advanced methods in terms of prediction performance, and the predicted protein complexes have similar features to known complexes, further verifying its effectiveness in protein complex identification. These studies not only promote the development of the PPI field, but also demonstrate the great potential of combining deep learning and structure prediction in understanding and designing protein complexes, providing an important theoretical basis and practical tools for future research in drug discovery and protein engineering.

3.2.4. Applications of protein function prediction methods

In recent years, research on protein-protein interactions (PPI) has made significant progress in the field of protein function prediction, especially in the combination of computational methods and experimental techniques. Researchers have successfully mapped the PPI network of Saccharomyces cerevisiae by combining affinity enrichment with mass spectrometry analysis [235]. This high-throughput method not only provides a comprehensive network structure, but also reveals the importance of low-abundance complexes and demonstrates the high connectivity of yeast proteins. In studying the influence of the ribosome on the folding of co-translating proteins and the assembly of complexes [19], a variety of experimental and computational methods are combined to determine the key role of specific residues in the assembly process and propose a predictive model based on the interface energy distribution. A study work [236] evaluating the potential of AlphaFold in drug discovery, especially in the study of “dark” human proteins that have not been fully studied, shows that the use of confidence scores can significantly improve prediction accuracy. Furthermore, the successful targeting of a PPI by a loop peptide designed with AlphaFold demonstrates the promise of this new method for designing more effective loop peptide sequences [237].

Further improvement and expansion of the existing PPI database will help to improve the accuracy of computational predictions and provide stronger support for follow-up research, since experimental verification is still a necessary part of PPI research. The construction of the HLA3DB database [238] provides a new approach for the prediction of T cell epitope peptides and significantly improves the accuracy of modeling, which has important applications in the study of antigen immunogenicity. The updated PrePPI database [239] combines structural and non-structural evidence to predict interactions across the whole genome. It provides a comprehensive view of structural information through AlphaFold structures and a unique scoring function, which further advances PPI research. These studies not only demonstrate the effectiveness of computational methods in PPI prediction, but also highlight the importance of combining experimental and computational techniques in understanding the formation and function of protein complexes.

Accurate prediction of the structure of antibody-antigen complexes is critical for drug discovery and vaccine design. The launch of the Predictomes database has improved the classification and confidence assessment of PPI predictions by the AlphaFold model using machine learning techniques [240]. These resources provide important support for understanding the molecular mechanism of PPI and improving the accuracy of prediction, laying a foundation for future biomedical research.

In the study of antibody-antigen complexes, recent progress has mainly been reflected in improvements to structural prediction and modeling techniques. For example, community-based deep learning methods have been used to predict the structure of the antibody CDR H3 loop. This new neural network architecture can improve structural sampling and accuracy through information exchange between communities [241]. In addition, the combination of molecular docking methods with AlphaFold2 has provided new perspectives for the structural prediction of antibody-antigen complexes.

However, progress in antibody-antigen prediction field remains constrained by the limited availability of high-resolution experimental antibody-antigen complex structures, particularly those involving the highly diverse CDR-H3 region, which is notoriously difficult to model accurately [224]. Expanding high-quality training datasets, especially those that encompass various CDR conformations and binding modes, will be essential to drive further advances in antibody structure prediction and improve the robustness of computational docking approaches. Recently, the ABAG-docking benchmark dataset [242] has been established to provide a rich set of cases and standardized evaluation for antibody-antigen docking. Meanwhile, the performance of non-binding docking has been improved by improving the evaluation of docking models [243]. Cross-linking mass spectrometry (XL-MS) has also been applied to the modeling of antibody-antigen complexes. The understanding of binding interfaces has been enhanced by using cross-linking data as distance restraints [244]. Together, these studies have advanced the understanding of antibody-antigen interactions and provide an important foundation for the development of new treatments.

Protein function prediction is also widely used in experimental applications. A study combining whole-cell cross-linking with AI structure prediction provided an in-depth analysis of PPIs in cells and revealed complex interactions between bacterial proteins [245]. AlphaFold has also demonstrated its potential in discovering new protein interactions, highlighting the advantages and limitations of structure prediction [43]. Furthermore, AlphaFold provided new structural insights and facilitated progress in cancer drug design in a PPI study related to renal cell carcinoma [246]. Review study [247] has highlighted the importance of computational tools in PPI network research, with mass spectrometry being considered a key technique for PPI identification. Meanwhile, structural reports on the KCTD1 protein and its pathogenic mutations have revealed the role of this protein in various diseases [248]. The first report of the structure of the new Golm1 tetramer also provides new insights into its function in physiological and pathological processes [249]. The research method combining AlphaFold prediction with experimental verification not only systematically identifies multiple PPI interfaces, but also provides new tools and methods for future research [189], [250]. These studies have jointly promoted in-depth exploration in the field of PPI and opened new doors for broader research and application prospects.

In summary, although some progress has been made in PPI prediction and protein complex structure prediction after AlphaFold, future challenges still focus on the accuracy of multi-chain systems, the ability to incorporate dynamic information, and the difficulty of experimental verification.

3.3. Protein dynamics prediction

Protein dynamics analysis is an important bridge linking protein structure and function. Understanding the trajectory of protein movement and its impact on function is an important topic in structural biology research. Dynamics is important not only in revealing the relationship between protein structure and function, but also in helping us understand the molecular mechanisms of biological processes, especially in drug design and disease treatment.

In recent years, the study of protein conformations has attracted widespread attention, covering different problems in protein folding and its processes, as well as the challenges they face [251], the future application of artificial intelligence in the protein folding revolution [252], and the combination of research from physical and chemical rules to intracellular protein quality control mechanisms with artificial intelligence solutions [253]. Nevertheless, the problem of protein folding has not yet been fully solved, which provides a broad space for future research [254].

3.3.1. AlphaFold-related dynamics prediction

In protein conformation research, new methods and models are driving the understanding of protein dynamics and their functions. The PAthreader [255] method excels at identifying remote homologous structures and exploring protein folding pathways, significantly improving the accuracy of template recognition and outperforming traditional methods in structure modeling. The development of the Cfold [216] network has demonstrated its ability to efficiently explore the conformational space of monomeric proteins and accurately predict known alternative conformations. PANDA-3D [256] provides an effective protein function prediction tool using AlphaFold models, emphasizing the importance of structural information in functional annotation. Together, these studies have contributed to a deeper understanding of protein conformation and function, providing new perspectives and approaches for drug discovery. DynamicBind uses deep learning techniques and an isomeric geometric diffusion network to effectively predict the structure of ligand-specific protein-ligand complexes, demonstrating advanced performance in recovering ligand-specific conformations [257]. The emergence of this method has provided new tools for studying the dynamics of proteins under ligand regulation. Meanwhile, the combination of ARTINA and AlphaFold improves the accuracy of chemical shift assignment and reduces the need for experimental data, thereby improving the efficiency of large-scale protein research [258]. This comprehensive method provides greater reliability for NMR research. In addition, GeoDock [259] overcomes the limitations of traditional protein-protein docking methods by using a multi-track iterative transformation network to effectively capture the conformational changes caused by ligand binding, making high-throughput complex structure prediction more feasible. The interface energy distribution provides a strong basis for predicting co-translational assembly and reveals the complex dynamics of protein-protein interactions. For example, this study [19] combines selective ribosome profiling, imaging, mass spectrometry, molecular dynamics, and AlphaFold-Multimer modeling to reveal the key role of hotspots residues in co-translational assembly pathways and propose a model based on interface energy distribution as a predictive factor for co-translational assembly.

In terms of dynamics prediction, a method that combines the physical energy landscape with a deep learning method has been proposed to better predict the conformational motion of proteins [260]. By incorporating local energy frustration information into the model, this method can more accurately simulate the dynamics of proteins. In addition, statistical analysis of residue spacing reveals the flexibility of protein conformational sets, and a study [261] provides new insights into understanding how proteins change in different environments.

3.3.2. Intrinsically disordered proteins

Intrinsically disordered regions (IDRs) in proteins are an important component of protein structure. Despite their unstable structure, they play a vital role in many biological processes. Functional prediction and kinetic analysis of IDRs is currently one of the hot topics in protein function research. Intrinsically disordered proteins (IDPs) lack stable three-dimensional structures in organisms, but they may exhibit ordered conformations when binding to globular receptors. By studying intrinsic disordered regions, one can reveal their important functions in cell signaling, regulatory effects, and their relationship to disease.

To accurately model these IDP-receptor complexes, researchers evaluated the performance of three modeling tools, which showed that all methods can effectively identify general binding sites, but have difficulty capturing high-resolution biophysical interactions [262]. To improve the prediction of protein binding sites, the proposed machine learning model IDBindT5 [221] performed well in disordered regions, achieving a balanced accuracy of 57.2%, demonstrating the potential of using protein language models to explore disordered features. Meanwhile, the monomer and dimer structures of IDP Nvjp-1 were modeled by AlphaFold2, and it was found that the dimer structure is more rigid and exhibits consistency in pLDDT scores [263]. In addition, AlphaFoldDB has shown competitiveness in disordered region prediction, demonstrating the plasticity between disordered and ordered conformations [264]. The hierarchical chain growth (HCG) method was proposed as an effective alternative for computational exploration of the conformational space of flexible biomolecules, capable of combining experimental data and simulation results for the study of proteins associated with neurodegenerative diseases [265].

The MobiDB and DisProt databases provide knowledge integration for disordered proteins, significantly improving the quality and coverage of annotations [266], [267]. In addition, researchers have developed molecular models to generate conformational sets of IDRs and correlate them with cellular functions, amino acid sequences, and disease variants, providing new insights into understanding the biological roles of IDPs [268]. Together, these works have promoted a deeper understanding of endogenous disordered proteins and their functions, providing an important reference for future drug design and biological research.

3.3.3. Protein dynamics computational validation and analysis

In recent years, AlphaFold2 has made remarkable achievements in the field of protein structure prediction and has become one of the main tools for researchers. However, the limitations and scope of application of this tool still need to be discussed in depth [269]. Although AlphaFold2 can successfully predict the structure of small folded proteins, in some cases it may incorrectly predict the folding state of the protein with a high degree of confidence. For example, in this study [269], the structure of the small protein pro-IL-18 was shown, and it was found that AlphaFold2's prediction matched the structure of the mature cytokine rather than the experimentally determined structure of the protein precursor, indicating that even for small folded proteins, there is still a need to rely on experimental structure determination.

Furthermore, the study found that AlphaFold2's low confidence regions often highly overlap with intrinsically disordered regions (IDRs), emphasizing the importance of understanding these disordered regions and their functions [270]. When analyzing false protein sequences, the researchers found that most families had no evidence of globular structures, but in some specific cases, AlphaFold2's predictions could still identify proteins that may be real, providing a useful quality control tool for the AntiFam database [271].

Notably, AlphaFold2 not only predicts the three-dimensional structure of proteins, but also reveals the dynamic properties of proteins through the local distance difference test (pLDDT) score and the predicted alignment error (PAE) map. Studies have shown that AlphaFold2's PAE map is correlated with the distance change matrix of molecular dynamics simulation, which provides a new perspective for the dynamic analysis of proteins [272]. Further research has shown that, in combination with local contact modeling, the information returned by AlphaFold2 can be used for dynamic prediction at the level of individual residues. The cdsAF2 [273] can semi-quantitatively capture the dynamic characteristics observed in the experimental main chain NMR N-H S2 ordering parameter spectrum, thereby providing an important basis for understanding protein function.

These findings highlight the need for experimental validation and the complex interrelationship between sequence and structure when making computational structural predictions, especially when studying intrinsically disordered and short-sequence proteins. Overall, although AlphaFold2 shows great potential in structural prediction, its limitations and relevance to actual biological function still need to be fully considered in specific applications.

By re-sampling multiple sequence pairs, researchers have developed an innovative method that successfully predicts the relative conformational distributions of different proteins. This work [274] has an accuracy rate of over 80% and is particularly effective in qualitatively predicting the impact of mutations or evolution on the conformational landscape, demonstrating its potential for use in pharmacology and experimental result analysis. In addition, AlphaFold's structural models also excel in drug binding mode prediction, especially in G protein-coupled receptor targets, where their structural capture of the binding pocket surpasses traditional homology models. However, despite the accurate structural prediction of the binding pocket, the prediction accuracy of the ligand binding site has not been significantly improved, which poses new challenges for drug discovery research [275]. Conformational dynamics research has also been enhanced. The competitive modeling approach (CMA) was used to evaluate alternative conformations of multi-domain complexes and compare them with crystallographic and cryo-EM data, showing consistency and further revealing the conformational dynamics of multi-domain electron transfer complex function [276]. By combining AlphaFold with molecular dynamics simulations, researchers were able to efficiently sample the conformational dynamics of transport proteins, obtaining an unresolvable inward-open conformation that provides new insights into the understanding of transport function [277], [278]. In the study of protein function, the AlphaFold method combined with sequence clustering (AF-Cluster) also provides new ideas for predicting alternative states of transmembrane proteins. This method shows how to use structural and sequence information to improve the accuracy of conformation prediction, and reveals unrecognized conformational states in some protein families [127]. Finally, the application of AlphaFold also shows great potential in the field of drug discovery. Through virtual screening of more than 16 million compounds, researchers found that the structural models generated by AlphaFold performed well in the discovery of highly effective agonists and provided a reference for structure-based drug design [279], which demonstrates a new way of computational methods to improve the efficiency of drug screening.

With the in-depth study of the functional and structural relationships of proteins, AlphaFold and other deep learning tools are playing an increasingly important role in drug discovery and biological research. By combining experimental data with computational simulations, these methods provide new insights into the study of complex biological processes, especially in understanding the dynamic properties and functions of protein conformations. However, researchers still need to remain vigilant and ensure continued experimental validation when relying on computational predictions in order to better understand the biological roles of proteins and their potential impact on diseases.

4. Multimer quality assessment

Proteins regularly function as complexes within living organisms, therefore the precise prediction of their structures and functions is a prerequisite to grasping disease mechanisms, drug design, and enzymatic reactions. Assessment of multimeric protein prediction quality constitutes a critical aspect in improving the accuracy of predictive methodologies [280]. The demand for structural analysis and functional annotation is implicated in rapid advances in experimental techniques and computational methods, such as cryo-electron microscopy and AlphaFold. Simultaneously, the development of data-driven approaches, particularly deep learning techniques (e.g., Transformers, convolutional neural networks, residual networks, and graph neural networks), has resulted in considerable enhancements in the precision and efficiency of quality assessment methods [281]. The advent of these novel data-driven methodologies signifies a paradigm shift from conventional experiment-based assessments, thereby propelling the field toward new frontiers.

The variety of contemporary assessment methods for multimeric protein prediction is based on their distinct assumptions and requirements. These methods comprise structural assessments based on global structural topology, interface contacts, and the local environment of residues [41], as well as functional assessments derived from experimental techniques, structure-sequence pre-training, and knowledge graphs. Within the framework of quality assessment for multimeric protein predictions, two critical components are the estimation of model accuracy (EMA) for structural predictions and functional annotation assessment [282]. The structural assessments are multi-objective optimization problems, ranging from evaluating the applicability of prediction algorithms to verifying the accuracy of computational methods. The primary objective is to enhance the efficiency of obtaining high-quality structural and functional predictions, deriving applications such as functional design [283], antibody binding [224], and drug discovery and development [284].

CASP16 [41] introduces multi-dimensional scoring metrics to comprehensively assess the folding quality, interface contact and local accuracy of protein structures. The assessment of global structural quality was performed using metrics such as GDT_TS-like [285] and TM-score [157], in combination with statistical methods including Pearson, Spearman, and AUC. The interface quality was evaluated through the utilization of DockQ-wave [286] and QS scores [287], while local accuracy was assessed using metrics such as PatchDockQ, PatchQS, CAD, and lDDT [286]. EMA Methods combines various statistical metrics, including Pearson, Spearman, AUC, ROC, and employs Z-score ranking and scoring. In regard to CASP15 [38], CASP16 requires more intricate interface prediction of protein multimers, thus enhancing the assessment capability of multimer structure prediction. The scoring metrics encompass ICS (interface contact degree), ICP (interface positional accuracy), ICR (interface contact rate), IPS (interface surface quality), QS, lDDT, TM-score, RMSD, and DockQ [44], which comprehensively measure the structural, interfacial, and localization accuracies of protein multimers. The diversity of quality assessment methods and objectives in the prediction of multimeric proteins is indicative of both the challenges and the potential inherent in this field.

4.1. Multimer structural estimation of model accuracy

The evaluation of the structural model accuracy (EMA) is pivotal to protein prediction, as it is a measure of how well the prediction model matches the actual structure. In the assessment of multimeric structures, it is not sufficient to focus exclusively on the folding quality of the individual subunits [286]. Global topological assessment is a key component of this evaluation, providing confirmation of the spatial arrangement and relative positions of the subunits in the multimeric system, ensuring that the rational geometrical and symmetrical topology. Subsequently, the analysis of contact interfaces is employed to understand the interactions between subunits and to assess their contact areas and mutual stability, thus verifying the stability and rationality of the system. Lastly, the assessment of local patterns focuses on the stability and functionality of local regions in the multimeric system, revealing potential structural defects or dynamic changes. The comprehensive analysis of these three aspects provides a comprehensive evaluation framework for the accuracy of multimeric structures.

In addition, the EMA can be categorized into consensus methods, quasi-single model methods, and single model methods [282]. The consensus method relies on clustering multiple structural models to extract key information from the structural clusters. The quasi-single model method generates structural clusters from the model itself, while the single model method directly evaluates the three-dimensional structure of the model. The latter method requires accurate capture of the main features of sequence, structure, physicochemical properties, and both global and local topologies.

4.1.1. Assessment for global topology

When evaluating global topology from the perspective of protein structures, assessing the topological similarity between predicted models and experimental structures is vital for ensuring the biological reliability and practical applicability of these models.

In CASP16, GDT_TS-like and TM-score were primarily used as reference metrics for complex predictions. Empirically, a GDT_TS score around 60% indicates “correct folding,” suggesting a reasonable understanding of how the protein folds overall. Scores above 80% indicate a high correlation between the side chains and the model, while a TM-score above 0.7 indicates acceptable model quality, with high-quality complexes requiring a TM-score above 0.8. Recent advancements in the EMA methods for complexes were revealed in CASP16. Leading methods include the MULTICOM [288], Guijunlab-PAthreader & QA [255], [185], MIEnsembles [181], and others. Jianlin Cheng's MULTICOM, which includes MULTICOM_LLM, MULTICOM _ GATE, and MULTICOM, uses MMAlign-based pairwise similarity scores to predict model quality. MULTICOM_LLM evaluates the similarity of a model to all other models. MULTICOM_GATE builds a model similarity graph and uses a graph transformer (GATE) [45] to predict model quality, integrating multiple node features such as AlphaFold score, EnQA score, and GCPNet-EMA score. MULTICOM combines the GCPNet-EMA and GATE scores to predict the overall model quality. Guijunlab-PAthreader enhances protein structure prediction accuracy using multiple sequence alignments (MSA) and model quality assessment. By modeling the target protein with AlphaFold2, retrieving distant homologous structures via Foldseek [289] from the AlphaFold DB, and constructing high-quality MSAs, the method predicts multiple models using enhanced MSAs and homologous templates. The best model is selected through internal single-model quality assessment and AlphaFold2 self-evaluation. MIEnsembles [181] by Wei Zheng et al. selects the best model from DMFold as a reference, using TM-score for overall folding accuracy and DockQ for interface accuracy. Quality of decoys is assessed relative to this reference model.

Furthermore, with the development of deep learning, there are now numerous methods for scoring overall topology of protein complexes, such as DPROQ [45], DProQA, GCPNet-EMA, ComplexQA [186], and VoroIF-jury [170]. DProQA [45] incorporates a new gated neighborhood modulation graph transformer and was ranked third in single-model quality assessment during blind testing in CASP15. GCPNet-EMA [185] uses a geometric message-passing network, applied in methods like MULTICOM. ComplexQA [186] leverages residue-level substructures and sequence constraints in three-dimensional space to assess local quality of protein complex interfaces. VoroIF-jury [170] is used when AlphaFold fails to assemble a complete protein complex or generates unreliable results, scoring additional models built by docking monomers or subcomplexes, independent of AlphaFold's self-assessment scores.

In conclusion, recent advancements in the field of protein complex structure prediction, fueled by continuous innovation and deep learning techniques, highlight the increasing importance of evaluating topological similarity between predicted models and experimental structures. These evolving methods, with their diverse evaluation metrics, provide more comprehensive and accurate tools for protein complex structure prediction.

4.1.2. Assessment for contact interface

In protein complexes, the contact interface determines the functional mechanisms of the entire molecule. Accurate assessment of the contact interface of protein complexes includes two aspects: whether corresponding residues are in contact and the overall topological accuracy of the contact interface.

In CASP16, metrics such as DockQ-wave and QS scores are used to evaluate the relative differences in contact interfaces between predicted models and experimental structures. ModFOLDdock [290] by McGuffin, L. J. and colleagues performed well in CASP16's local evaluation. ModFOLDdock2 [291] significantly improved the accuracy of protein quaternary structure model quality assessment by combining multiple scoring methods, especially in predicting local interface residues. This tool is specifically designed for evaluating protein quaternary structure model quality, focusing on both global and local scoring, particularly the accuracy of interface residues. Compared to ModFOLDdock1, ModFOLDdock2 introduces new scoring methods and utilizes neural networks to predict local scores, thus optimizing quality assessment. Its server integrates 12 different scoring methods, which assess model quality from multiple dimensions, including global and local scores. Another version, ModFOLDdock2R, calculates overall fold accuracy through the average of lDDTJury, DockQ-waveJury, and VoroIF (weighted average of pcad scores), while the overall interface accuracy in column 3 is computed by averaging VoroMQA, DockQ-waveJury, and VoroIF (weighted average of pcad scores).

Yang Zhang et al.'s MQA_base [292] calculates residue-level lDDT and uses complete distance deviation maps and mask maps to extract scores for interface residues, whose training loss function includes: cross-entropy loss for distance deviation maps, binary cross-entropy loss for mask maps, L2 loss for lDDT, and L1 loss for TM-score.

4.1.3. Assessment for local mode

In CASP16, local accuracy evaluation is performed using metrics such as PatchDockQ, PatchQS, CAD, and lDDT. The EMA method combines multiple statistical measures. Among them, lDDT is widely used because it accurately describes the atomic environment surrounding each residue. In the self-assessment of Local Mode, methods such as ModFOLDdock2, GuijunLab-QA, and MQA have performed exceptionally well.

In MQA, DeepUMQA3 [292] is used, which is a deep neural network web server for evaluating the accuracy of protein complex interface residues. For the input complex structure, features are extracted from three levels: the entire complex, within a monomer, and between monomers. An improved deep residual neural network is used to predict the lDDT of each residue and the accuracy of interface residues. Meanwhile, DeepUMQA3 can evaluate the accuracy of all residues in the entire complex and distinguish between high-precision and low-precision residues. In the CASP15 test, DeepUMQA3 ranked first in blind testing for interface residue accuracy estimation. Under the lDDT measurement, the Pearson, Spearman, and AUC scores were 0.564, 0.535, and 0.755, respectively, outperforming the second-place method by 17.6%, 23.6%, and 10.9%.

4.2. Challenges of effective multimer quality assessment

Despite continuous progress in existing methods for evaluating the quality of polymers, a number of challenges remain [280]. First, the accuracy of prediction algorithms is still limited, especially for large polymer systems, and it is still difficult to accurately predict the interactions of all subunits and the global structure. Second, how to effectively integrate dynamic information (such as protein conformational changes) in polymer structure prediction is still an urgent problem to be solved. In addition, the difficulty of experimental verification is also an important factor limiting the effectiveness of the assessment of the quality of polymers. In particular, when the dataset is large or the complexity of the polymer is high, the efficiency and cost of experimental verification become bottlenecks. Therefore, how to further improve the accuracy of the prediction of the structure and function of polymers while reducing the dependence on experimental verification is an important direction for future research.

5. Conclusion

AlphaFold [27], [196] has undoubtedly revolutionized the field of protein structure prediction, significantly enhancing our understanding of protein biology. The release of AlphaFold2 marked a pivotal moment in structural biology, providing unprecedented insights into protein folding and structure determination. By generating atomistic three-dimensional structures from amino acid sequences, AlphaFold 2 has refined traditional methods and become a cornerstone of modern structural biology.

Deep learning models like AlphaFold 2 utilize vast datasets and complex neural networks to predict protein structures. These models leverage evolutionary information and contextual features to generate highly accurate three-dimensional structures, enhancing our comprehensive understanding of protein structures and uncovering details about previously unknown protein functions and interactions.

Despite the remarkable progress achieved, some limitations remain. One significant challenge is the modeling of non-globular and intrinsically disordered proteins. These regions, lacking stable secondary or tertiary structures, are difficult to model accurately. Although AlphaFold has made advances in this area, predicting these regions remains a challenge, emphasizing the need for further improvements in modeling flexible and disordered protein segments [270].

In addition, while AlphaFold has made significant strides in predicting individual protein structures, predicting protein-protein interactions (PPIs) presents additional difficulties. Recent advancements in protein complex structure prediction have leveraged deep learning to tackle these challenges [40]. Models for protein complexes, such as AlphaFold-Multimer, aim to improve the accuracy of protein complex predictions by providing better insights into interaction interfaces and binding modes [293].

In the field of PPI prediction, significant progress has been made through the application of deep learning. Tools such as AlphaPulldown and SpatialPPI use AlphaFold's predictions to incorporate structural information into interaction models, thereby enhancing the accuracy of PPI predictions. These models leverage AlphaFold's high-resolution structure predictions, making PPI predictions more accurate and reliable.

However, integrating AlphaFold predictions into PPI models requires careful consideration of several issues [294], [293]. One of the major challenges is accounting for dynamic conformational changes and ensuring that predictions reflect true biological interactions. To address these challenges, advanced methods are being developed that combine AlphaFold predictions with experimental data and employ integrative approaches to capture variability in protein interactions.

With the continuous development of computational methods and experimental techniques, quality assessment methods for predicting the structure and function of polymer proteins have also progressed significantly. These evaluation methods, including global topology, contact interfaces, and local modes, enable a more comprehensive assessment of the structural accuracy of polymers. Nevertheless, challenges remain, particularly in terms of accuracy and experimental verification. To address these issues, future progress will require innovative methods, optimized algorithms, and the integration of dynamic information to improve the efficiency and precision of quality assessments for polymers. These advancements will not only promote the development of protein science but will also open new research avenues in drug design, disease treatment, and biological research.

The advancements brought about by AlphaFold have had a profound impact on both basic research and practical applications such as drug discovery and protein design. Accurate protein structure prediction helps identify novel drug targets and design more effective therapeutic drugs [295], [296]. Moreover, the integration of AlphaFold with deep learning techniques has facilitated the development of advanced tools for novel protein design and sequence optimization, potentially revolutionizing therapeutic protein engineering and biomolecular design [297], [298].

To further advance protein structure prediction, the continuous expansion of high-resolution experimental structure data plays a critical priority—especially for structurally complex targets such as transmembrane/membrane protein complexes or conformationally flexible proteins, and transient interaction complexes, where structural coverage is still limited and thus constrains modeling accuracy and generalizability. Recent developments in structural biology have diversified the sources of structural data. For instance, NMR provides dynamic insights into systems with multiple conformations, while cryo-electron microscopy (cryo-ET) is increasingly capable of revealing the architectures of large macromolecular assemblies in near-native states. These advances offer new dimensions for understanding functional states and complex biomolecular interactions.

In particular, the success of current structure prediction methods is largely built on decades of sustained progress in structural biology. These experimental structures have not only supported the training of deep learning models such as AlphaFold, but also served as critical benchmarks for model evaluation and algorithm refinement. As the diversity and coverage of the experimental data continue to grow, structure prediction methods are expected to achieve further breakthroughs in accuracy, generalization, and their ability to handle dynamic and complex biological systems.

Future research efforts may focus on optimizing AlphaFold's capabilities and addressing its current limitations. Key areas for improvement include enhancing the modeling of flexible and disordered regions, refining predictions for protein complexes, and integrating AlphaFold's predictions with experimental data. As deep learning and computational methods continue to advance, they will further enhance the accuracy and applicability of protein structure predictions, unlocking new possibilities for scientific discoveries and medical applications.

AlphaFold has made significant strides in advancing our understanding of protein structures, reshaping the landscape of structural biology and computational modeling. Despite numerous breakthroughs, further research is necessary to overcome its limitations and fully realize its potential in predicting complex protein interactions and designing new drugs. Ongoing advancements in deep learning technologies and experimental validation will play a pivotal role in advancing this field and fostering scientific discoveries.

CRediT authorship contribution statement

Da Lu: Writing – review & editing, Writing – original draft, Investigation, Conceptualization. Shuhong Yu: Writing – review & editing. Yixiang Huang: Writing – review & editing. Xinqi Gong: Writing – review & editing.

Declaration of Competing Interest

All authors disclosed no relevant relationships.

Acknowledgement

This research was supported by Innovation Platform, Renmin University of China, Public Computing Cloud, Renmin University of China and Shenzhen Medical Research Fund (B2402038).

Contributor Information

Da Lu, Email: luda@ruc.edu.cn.

Shuhong Yu, Email: shuhongyu98@ruc.edu.cn.

Yixiang Huang, Email: 2021103691@ruc.edu.cn.

Xinqi Gong, Email: xinqigong@ruc.edu.cn.

References

  • 1.Ahmed F.H., Liu J.-W., Royan S., Warden A.C., Esquirol L., Pandey G., et al. Structural insights into the enzymatic breakdown of azomycin-derived antibiotics by 2-nitroimdazole hydrolase (nnha) Commun Biol. 2024;7(1):1676. doi: 10.1038/s42003-024-07336-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Enbar T., Hickmott J.W., Siu R., Gao D., Garcia-Flores E., Smart J., et al. Regionally distinct gfap promoter expression plays a role in off-target neuron expression following aav5 transduction. Sci Rep. 2024;14(1) doi: 10.1038/s41598-024-79124-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Li X., Xu M., Yang J., Zhou L., Liu L., Li M., et al. Nasal vaccination of triple-RBD scaffold protein with flagellin elicits long-term protection against SARS-CoV-2 variants including JN. 1. Signal Transduct Targeted Ther. 2024;9(1):114. doi: 10.1038/s41392-024-01822-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Li P., Zhu Z., Wang Y., Zhang X., Yang C., Zhu Y., et al. Substrate transport and drug interaction of human thiamine transporters SLC19A2/A3. Nat Commun. 2024;15(1) doi: 10.1038/s41467-024-55359-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Anfinsen C.B. Principles that govern the folding of protein chains. Science. 1973;181(4096):223–230. doi: 10.1126/science.181.4096.223. [DOI] [PubMed] [Google Scholar]
  • 6.Anfinsen C., Scheraga H. Experimental and theoretical aspects of protein folding. Adv Protein Chem. 1975;29:205–300. doi: 10.1016/s0065-3233(08)60413-1. [DOI] [PubMed] [Google Scholar]
  • 7.Schulz G.E., Schirmer R.H. Springer Science & Business Media; 2013. Principles of protein structure. [Google Scholar]
  • 8.Rha M.-S., Kim G., Lee S., Kim J., Jeong Y., Jung C.M., et al. SARS-CoV-2 spike-specific nasal-resident CD49a+ CD8+ memory T cells exert immediate effector functions with enhanced IFN-γ production. Nat Commun. 2024;15(1):8355. doi: 10.1038/s41467-024-52689-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Hao H., Yuan Y., Ito A., Eberand B.M., Tjondro H., Cielesh M., et al. FUT10 and FUT11 are protein O-fucosyltransferases that modify protein EMI domains. Nat Chem Biol. 2025:1–13. doi: 10.1038/s41589-024-01815-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Yang Y., Ding T., Cong Y., Luo X., Liu C., Gong T., et al. Interferon-induced transmembrane protein-1 competitively blocks ephrin receptor a2-mediated Epstein–barr virus entry into epithelial cells. Nat Microbiol. 2024:1–15. doi: 10.1038/s41564-024-01659-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Boutet E., Lieberherr D., Tognolli M., Schneider M., Bansal P., Bridge A.J., et al. UniProtKB/Swiss-Prot, the manually annotated section of the uniprot knowledgebase: how to use the entry view. Plant Bioinform: Methods Protoc. 2016:23–54. doi: 10.1007/978-1-4939-3167-5_2. [DOI] [PubMed] [Google Scholar]
  • 12.Berman H.M., Battistuz T., Bhat T.N., Bluhm W.F., Bourne P.E., Burkhardt K., et al. The protein data bank. Acta Crystallogr, D Biol Crystallogr. 2002;58(6):899–907. doi: 10.1107/s0907444902003451. [DOI] [PubMed] [Google Scholar]
  • 13.Cressey D., Callaway E. Cryo-electron microscopy wins chemistry Nobel. Nature. 2017;550(7675) doi: 10.1038/nature.2017.22738. [DOI] [PubMed] [Google Scholar]
  • 14.Kundrotas P.J., Alexov E. Electrostatic properties of protein-protein complexes. Biophys J. 2006;91(5):1724–1736. doi: 10.1529/biophysj.106.086025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lemaire S., Ferreira M., Claes Z., Derua R., Lake M., Van der Hoeven G., et al. Ppp1R2 stimulates protein phosphatase-1 through stabilisation of dynamic subunit interactions. Nat Commun. 2024;15(1):9822. doi: 10.1038/s41467-024-54256-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Guo S., Chang Y., Brun Y.V., Howell P.L., Burrows L.L., Liu J. Pily1 regulates the dynamic architecture of the type iv pilus machine in pseudomonas aeruginosa. Nat Commun. 2024;15(1):9382. doi: 10.1038/s41467-024-53638-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Feng S., Chen Z., Zhang C., Xie Y., Ovchinnikov S., Gao Y.Q., et al. Integrated structure prediction of protein–protein docking with experimental restraints using colabdock. Nat Mach Intell. 2024;6(8):924–935. [Google Scholar]
  • 18.Wang T., He X., Li M., Li Y., Bi R., Wang Y., et al. Ab initio characterization of protein molecular dynamics with ai2bmd. Nature. 2024:1–9. doi: 10.1038/s41586-024-08127-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Venezian J., Bar-Yosef H., Ben-Arie Zilberman H., Cohen N., Kleifeld O., Fernandez-Recio J., et al. Diverging co-translational protein complex assembly pathways are governed by interface energy distribution. Nat Commun. 2024;15(1):2638. doi: 10.1038/s41467-024-46881-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Abouzied A.S., Alqarni S., Younes K.M., Alanazi S.M., Alrsheed D.M., Alhathal R.K., et al. Structural and free energy landscape analysis for the discovery of antiviral compounds targeting the cap-binding domain of influenza polymerase pb2. Sci Rep. 2024;14(1) doi: 10.1038/s41598-024-69816-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Roth C.M., Neal B.L., Lenhoff A.M. Van der Waals interactions involving proteins. Biophys J. 1996;70(2):977–987. doi: 10.1016/S0006-3495(96)79641-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Aliste M.P., MacCallum J.L., Tieleman D.P. Molecular dynamics simulations of pentapeptides at interfaces: salt bridge and cation- π interactions. Biochemistry. 2003;42(30):8976–8987. doi: 10.1021/bi027001j. [DOI] [PubMed] [Google Scholar]
  • 23.Donald J.E., Kulp D.W., DeGrado W.F. Salt bridges: geometrically specific, designable interactions. Proteins, Struct Funct Bioinform. 2011;79(3):898–915. doi: 10.1002/prot.22927. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Sippel K.H., Quiocho F.A. Ion–dipole interactions and their functions in proteins. Protein Sci. 2015;24(7):1040–1046. doi: 10.1002/pro.2685. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hubbard R.E., Haider M.K. Hydrogen bonds in proteins: role and strength. Encycl Life Sci. 2010;1:1–6. [Google Scholar]
  • 26.AlQuraishi M. End-to-end differentiable learning of protein structure. Cell Syst. 2019;8(4):292–301. doi: 10.1016/j.cels.2019.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Bryant P., Pozzati G., Elofsson A. Improved prediction of protein-protein interactions using alphafold2. Nat Commun. 2022;13(1):1265. doi: 10.1038/s41467-022-28865-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Evans R., O'Neill M., Pritzel A., Antropova N., Senior A., Green T., et al. Protein complex prediction with alphafold-multimer. bioRxiv. 2021 2021–10. [Google Scholar]
  • 29.Gao M., Nakajima An D., Parks J.M., Skolnick J. Af2complex predicts direct physical interactions in multimeric proteins with deep learning. Nat Commun. 2022;13(1):1744. doi: 10.1038/s41467-022-29394-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zheng W., Wuyun Q., Freddolino P.L., Zhang Y. Integrating deep learning, threading alignments, and a multi-msa strategy for high-quality protein monomer and complex structure prediction in casp15. Proteins, Struct Funct Bioinform. 2023;91(12):1684–1703. doi: 10.1002/prot.26585. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Edgar R.C., Batzoglou S. Multiple sequence alignment. Curr Opin Struct Biol. 2006;16(3):368–373. doi: 10.1016/j.sbi.2006.04.004. [DOI] [PubMed] [Google Scholar]
  • 32.Baek M., Anishchenko I., Park H., Humphreys I.R., Baker D. Protein oligomer modeling guided by predicted interchain contacts in casp14. Proteins, Struct Funct Bioinform. 2021;89(12):1824–1833. doi: 10.1002/prot.26197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Quadir F., Roy R.S., Soltanikazemi E., Cheng J. Deepcomplex: a web server of predicting protein complex structures by deep learning inter-chain contact prediction and distance-based modelling. Front Mol Biosci. 2021;8 doi: 10.3389/fmolb.2021.716973. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Soltanikazemi E., Roy R.S., Quadir F., Giri N., Morehead A., Cheng J. Drlcomplex: reconstruction of protein quaternary structures using deep reinforcement learning. 2022. arXiv:2205.13594 arXiv preprint.
  • 35.Manicki M., Aydin H., Abriata L.A., Overmyer K.A., Guerra R.M., Coon J.J., et al. Structure and functionality of a multimeric human coq7: Coq9 complex. Mol Cell. 2022;82(22):4307–4323. doi: 10.1016/j.molcel.2022.10.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Cao L., Coventry B., Goreshnik I., Huang B., Sheffler W., Park J.S., et al. Design of protein-binding proteins from the target structure alone. Nature. 2022;605(7910):551–560. doi: 10.1038/s41586-022-04654-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Garabedian M.V., Su Z., Dabdoub J., Tong M., Deiters A., Hammer D.A., et al. Protein condensate formation via controlled multimerization of intrinsically disordered sequences. Biochemistry. 2022;61(22):2470–2481. doi: 10.1021/acs.biochem.2c00250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Alexander L.T., Durairaj J., Kryshtafovych A., Abriata L.A., Bayo Y., Bhabha G., et al. Protein target highlights in casp15: analysis of models by structure providers. Proteins, Struct Funct Bioinform. 2023;91(12):1571–1599. doi: 10.1002/prot.26545. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Elofsson A. Progress at protein structure prediction, as seen in CASP15. Curr Opin Struct Biol. 2023;80 doi: 10.1016/j.sbi.2023.102594. [DOI] [PubMed] [Google Scholar]
  • 40.Lensink M.F., Brysbaert G., Raouraoua N., Bates P.A., Giulini M., Honorato R.V., et al. Impact of alphafold on structure prediction of protein complexes: the casp15-capri experiment. Proteins, Struct Funct Bioinform. 2023;91(12):1658–1683. doi: 10.1002/prot.26609. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.CASP prediction center, Casp16. 2024. https://predictioncenter5.genomecenter.ucdavis.edu/casp16/index.cgi
  • 42.Huang Y., Wuchty S., Zhou Y., Zhang Z. Sgppi: structure-aware prediction of protein–protein interactions in rigorous conditions with graph convolutional network. Brief Bioinform. 2023;24(2) doi: 10.1093/bib/bbad020. [DOI] [PubMed] [Google Scholar]
  • 43.Homma F., Lyu J., van der Hoorn R.A. Using alphafold multimer to discover interkingdom protein–protein interactions. Plant J. 2024 doi: 10.1111/tpj.16969. [DOI] [PubMed] [Google Scholar]
  • 44.Basu S., Wallner B. Dockq: a quality measure for protein-protein docking models. PLoS ONE. 2016;11(8) doi: 10.1371/journal.pone.0161879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Chen X., Morehead A., Liu J., Cheng J. A gated graph transformer for protein complex structure quality assessment and its performance in casp15. Bioinformatics. 2023;39(Supplement_1) doi: 10.1093/bioinformatics/btad203. i308–i317. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Krishna R., Wang J., Ahern W., Sturmfels P., Venkatesh P., Kalvet I., et al. Generalized biomolecular modeling and design with rosettafold all-atom. Science. 2024;384(6693) doi: 10.1126/science.adl2528. [DOI] [PubMed] [Google Scholar]
  • 47.Watson J.L., Juergens D., Bennett N.R., Trippe B.L., Yim J., Eisenach H.E., et al. De novo design of protein structure and function with rfdiffusion. Nature. 2023;620(7976):1089–1100. doi: 10.1038/s41586-023-06415-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Ruffolo J.A., Madani A. Designing proteins with language models. Nat Biotechnol. 2024;42(2):200–202. doi: 10.1038/s41587-024-02123-4. [DOI] [PubMed] [Google Scholar]
  • 49.Pan X., Kortemme T. Recent advances in de novo protein design: principles, methods, and applications. J Biol Chem. 2021;296 doi: 10.1016/j.jbc.2021.100558. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Cunningham J.M., Koytiger G., Sorger P.K., AlQuraishi M. Biophysical prediction of protein–peptide interactions and signaling networks using machine learning. Nat Methods. 2020;17(2):175–183. doi: 10.1038/s41592-019-0687-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Abanades B., Wong W.K., Boyles F., Georges G., Bujotzek A., Deane C.M. Immunebuilder: deep-learning models for predicting the structures of immune proteins. Commun Biol. 2023;6(1):575. doi: 10.1038/s42003-023-04927-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Zahiri J., Emamjomeh A., Bagheri S., Ivazeh A., Mahdevar G., Tehrani H.S., et al. Protein complex prediction: a survey. Genomics. 2020;112(1):174–183. doi: 10.1016/j.ygeno.2019.01.011. [DOI] [PubMed] [Google Scholar]
  • 53.Grassmann G., Di Rienzo L., Gosti G., Leonetti M., Ruocco G., Miotto M., et al. Electrostatic complementarity at the interface drives transient protein-protein interactions. Sci Rep. 2023;13(1) doi: 10.1038/s41598-023-37130-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Tam J.Z., Palumbo T., Miwa J.M., Chen B.Y. Analysis of protein-protein interactions for intermolecular bond prediction. Molecules. 2022;27(19):6178. doi: 10.3390/molecules27196178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Jiang L., Lai L. CH⋯O hydrogen bonds at protein-protein interfaces⁎ 210. J Biol Chem. 2002;277(40):37732–37740. doi: 10.1074/jbc.M204514200. [DOI] [PubMed] [Google Scholar]
  • 56.Ó'Fágáin C. Protein chromatography: methods and protocols. Springer; 2023. Protein stability: enhancement and measurement; pp. 369–419. [Google Scholar]
  • 57.Kaledhonkar S., Fu Z., White H., Frank J. Protein complex assembly: methods and protocols. Methods Mol Biol. 2018;1764:59–71. doi: 10.1007/978-1-4939-7759-8_4. [DOI] [PubMed] [Google Scholar]
  • 58.Kaczor A.A. Springer Nature; 2024. Protein-protein docking: methods and protocols, vol. 2780. [Google Scholar]
  • 59.Yang J., Zhang Y. I-tasser server: new development for protein structure and function predictions. Nucleic Acids Res. 2015;43(W1):W174–W181. doi: 10.1093/nar/gkv342. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Du Z., Su H., Wang W., Ye L., Wei H., Peng Z., et al. The trrosetta server for fast and accurate protein structure prediction. Nat Protoc. 2021;16(12):5634–5651. doi: 10.1038/s41596-021-00628-9. [DOI] [PubMed] [Google Scholar]
  • 61.Källberg M., Margaryan G., Wang S., Ma J., Xu J. Raptorx server: a resource for template-based protein structure modeling. Protein Struct Predict. 2014:17–27. doi: 10.1007/978-1-4939-0366-5_2. [DOI] [PubMed] [Google Scholar]
  • 62.Wang C., Zhang H., Zheng W.-M., Xu D., Zhu J., Wang B., et al. Falcon@ home: a high-throughput protein structure prediction server based on remote homologue recognition. Bioinformatics. 2016;32(3):462–464. doi: 10.1093/bioinformatics/btv581. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Xu D., Zhang Y. Ab initio protein structure assembly using continuous structure fragments and optimized knowledge-based force field. Proteins, Struct Funct Bioinform. 2012;80(7):1715–1735. doi: 10.1002/prot.24065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Yang Y., Faraggi E., Zhao H., Zhou Y. Improving protein fold recognition and template-based modeling by employing probabilistic-based matching between predicted one-dimensional structural properties of query and corresponding native properties of templates. Bioinformatics. 2011;27(15):2076–2082. doi: 10.1093/bioinformatics/btr350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Chan H.S., Dill K.A. The protein folding problem. Phys Today. 1993;46(2):24–32. [Google Scholar]
  • 66.Levy E.D., Erba E.B., Robinson C.V., Teichmann S.A. Assembly reflects evolution of protein complexes. Nature. 2008;453(7199):1262–1265. doi: 10.1038/nature06942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Marsh J.A., Teichmann S.A. Structure, dynamics, assembly, and evolution of protein complexes. Annu Rev Biochem. 2015;84(1):551–575. doi: 10.1146/annurev-biochem-060614-034142. [DOI] [PubMed] [Google Scholar]
  • 68.Bahadur R., Zacharias M. The interface of protein-protein complexes: analysis of contacts and prediction of interactions. Cell Mol Life Sci. 2008;65:1059–1072. doi: 10.1007/s00018-007-7451-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Goodsell D.S., Olson A.J. Structural symmetry and protein function. Annu Rev Biophys Biomol Struct. 2000;29(1):105–153. doi: 10.1146/annurev.biophys.29.1.105. [DOI] [PubMed] [Google Scholar]
  • 70.Gray J.J., Moughon S., Wang C., Schueler-Furman O., Kuhlman B., Rohl C.A., et al. Protein–protein docking with simultaneous optimization of rigid-body displacement and side-chain conformations. J Mol Biol. 2003;331(1):281–299. doi: 10.1016/s0022-2836(03)00670-3. [DOI] [PubMed] [Google Scholar]
  • 71.Totrov M., Abagyan R. Flexible protein–ligand docking by global energy optimization in internal coordinates. Proteins, Struct Funct Bioinform. 1997;29(S1):215–220. doi: 10.1002/(sici)1097-0134(1997)1+<215::aid-prot29>3.3.co;2-i. [DOI] [PubMed] [Google Scholar]
  • 72.Bennett N.R., Coventry B., Goreshnik I., Huang B., Allen A., Vafeados D., Dauparas J., Baek M., Stewart L. Improving de novo protein binder design with deep learning. Nat Commun. 2023;14(8):2625. doi: 10.1038/s41467-023-38328-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Vázquez Torres S., Benard Valle M., Mackessy S.P., Menzies S.K., Casewell N.R., Ahmadi S., Burlet N.J., Muratspahić E., Sappington I., Overath M.D., Esperanza R.-de-T., Jann L., Esperanza R.-de-T., Andreas H.L., Kim B., Asim K.B., Alex K., Evans B., Iara A.C., Edouard P.C., Rebecca J.E., Justin D., Robert J.R., Arvind S.P., Mohamad A., Hannah L.H., Stacey R.G., Analisa M., Rebecca S., Lynda S., Lance S., Thomas J.A.F., Timothy P.J., David B. De novo designed proteins neutralize lethal snake venom toxins. Nature. 2025;639(8053):225–231. doi: 10.1038/s41586-024-08393-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Kiani Y.S., Jabeen I. Springer US; 2024. Challenges of protein-protein docking of the membrane proteins; pp. 203–255. [DOI] [PubMed] [Google Scholar]
  • 75.Szczepski K., Jaremko Ł. AlphaFold and what is next: bridging functional, systems and structural biology. Expert Rev Proteomics. 2025;22(2):45–58. doi: 10.1080/14789450.2025.2456046. [DOI] [PubMed] [Google Scholar]
  • 76.Erdős G., Dosztányi Z. Deep learning for intrinsically disordered proteins: from improved predictions to deciphering conformational ensembles. Curr Opin Struct Biol. 2024;89:102950. doi: 10.1016/j.sbi.2024.102950. [DOI] [PubMed] [Google Scholar]
  • 77.Veale C.G.L., Clarke D.J. Mass spectrometry-based methods for characterizing transient protein–protein interactions. Trends Chem. 2024;6(7):377–391. [Google Scholar]
  • 78.Yin R., Feng B.Y., Varshney A., Pierce B.G. Benchmarking alphafold for protein complex modeling reveals accuracy determinants. Protein Sci. 2022;31(8) doi: 10.1002/pro.4379. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Ruvinsky A.M., Kirys T., Tuzikov A.V., Vakser I.A. Structure fluctuations and conformational changes in protein binding. J Bioinform Comput Biol. 2012;10(02) doi: 10.1142/S0219720012410028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Campitelli P., Modi T., Kumar S., Ozkan S.B. The role of conformational dynamics and allostery in modulating protein evolution. Annu Rev Biophys. 2020;49(1):267–288. doi: 10.1146/annurev-biophys-052118-115517. [DOI] [PubMed] [Google Scholar]
  • 81.Morcos F., Pagnani A., Lunt B., Bertolino A., Marks D.S., Sander C., et al. Direct-coupling analysis of residue coevolution captures native contacts across many protein families. Proc Natl Acad Sci. 2011;108(49):E1293–E1301. doi: 10.1073/pnas.1111471108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.De Juan D., Pazos F., Valencia A. Emerging methods in protein co-evolution. Nat Rev Genet. 2013;14(4):249–261. doi: 10.1038/nrg3414. [DOI] [PubMed] [Google Scholar]
  • 83.Goh C.-S., Cohen F.E. Co-evolutionary analysis reveals insights into protein–protein interactions. J Mol Biol. 2002;324(1):177–192. doi: 10.1016/s0022-2836(02)01038-0. [DOI] [PubMed] [Google Scholar]
  • 84.Hopf T.A., Schärfe C.P., Rodrigues J.P., Green A.G., Kohlbacher O., Sander C., et al. Sequence co-evolution gives 3d contacts and structures of protein complexes. eLife. 2014;3 doi: 10.7554/eLife.03430. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Cheng F., Zhao J., Wang Y., Lu W., Liu Z., Zhou Y., et al. Comprehensive characterization of protein–protein interactions perturbed by disease mutations. Nat Genet. 2021;53(3):342–353. doi: 10.1038/s41588-020-00774-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Kim A.-R., Hu Y., Comjean A., Rodiger J., Mohr S.E., Perrimon N. Enhanced protein-protein interaction discovery via alphafold-multimer. bioRxiv. 2024 2024–02. [Google Scholar]
  • 87.Baek M., McHugh R., Anishchenko I., Jiang H., Baker D., DiMaio F. Accurate prediction of protein–nucleic acid complexes using rosettafoldna. Nat Methods. 2024;21(1):117–121. doi: 10.1038/s41592-023-02086-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Bepler T., Berger B. Learning the protein language: evolution, structure, and function. Cell Syst. 2021;12(6):654–669. doi: 10.1016/j.cels.2021.05.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Riedl S., Bilgen E., Agam G., Hirvonen V., Jussupow A., Tippl F., et al. Evolution of the conformational dynamics of the molecular chaperone hsp90. Nat Commun. 2024;15(1):8627. doi: 10.1038/s41467-024-52995-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Wu S., Tian C., Liu P., Guo D., Zheng W., Huang X., et al. Effects of sars-cov-2 mutations on protein structures and intraviral protein–protein interactions. J Med Virol. 2021;93(4):2132–2140. doi: 10.1002/jmv.26597. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Kim D., Noh M.H., Park M., Kim I., Ahn H., Ye D.-y., et al. Enzyme activity engineering based on sequence co-evolution analysis. Metab Eng. 2022;74:49–60. doi: 10.1016/j.ymben.2022.09.001. [DOI] [PubMed] [Google Scholar]
  • 92.Liang Z., Verkhivker G.M., Hu G. Integration of network models and evolutionary analysis into high-throughput modeling of protein dynamics and allosteric regulation: theory, tools and applications. Brief Bioinform. 2019;21(3):815–835. doi: 10.1093/bib/bbz029. https://academic.oup.com/bib/article-pdf/21/3/815/33227481/bbz029.pdf [DOI] [PubMed] [Google Scholar]
  • 93.Green A.G., Elhabashy H., Brock K.P., Maddamsetti R., Kohlbacher O., Marks D.S. Large-scale discovery of protein interactions at residue resolution using co-evolution calculated from genomic sequences. Nat Commun. 2021;12(1):1396. doi: 10.1038/s41467-021-21636-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Kryshtafovych A., Schwede T., Topf M., Fidelis K., Moult J. Critical assessment of methods of protein structure prediction (casp)—round xiv. Proteins, Struct Funct Bioinform. 2021;89(12):1607–1617. doi: 10.1002/prot.26237. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Wang S., Sun S., Li Z., Zhang R., Xu J. Accurate de novo prediction of protein contact map by ultra-deep learning model. PLoS Comput Biol. 2017;13(1) doi: 10.1371/journal.pcbi.1005324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Zeng H., Wang S., Zhou T., Zhao F., Li X., Wu Q., et al. Complexcontact: a web server for inter-protein contact prediction using deep learning. Nucleic Acids Res. 2018;46(W1):W432–W437. doi: 10.1093/nar/gky420. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Xie Z., Xu J. Deep graph learning of inter-protein contacts. Bioinformatics. 2022;38(4):947–953. doi: 10.1093/bioinformatics/btab761. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Morehead A., Chen C., Cheng J. Geometric transformers for protein interface contact prediction. 2021. arXiv:2110.02423 arXiv preprint.
  • 99.Guo Z., Liu J., Skolnick J., Cheng J. Prediction of inter-chain distance maps of protein complexes with 2d attention-based deep neural networks. Nat Commun. 2022;13(1):6963. doi: 10.1038/s41467-022-34600-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Rives A., Meier J., Sercu T., Goyal S., Lin Z., Liu J., et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci. 2021;118(15) doi: 10.1073/pnas.2016239118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Rao R.M., Liu J., Verkuil R., Meier J., Canny J., Abbeel P., et al. International conference on machine learning. PMLR; 2021. Msa transformer; pp. 8844–8856. [Google Scholar]
  • 102.Wu T., Huang H., Li J., Wang W., Gong X. 2022 IEEE international conference on bioinformatics and biomedicine (BIBM) IEEE; 2022. Inter-chain contact map prediction for protein complex based on graph attention network and triangular multiplication update; pp. 2143–2148. [Google Scholar]
  • 103.Si Y., Yan C. Improved inter-protein contact prediction using dimensional hybrid residual networks and protein language models. Brief Bioinform. 2023;24(2) doi: 10.1093/bib/bbad039. [DOI] [PubMed] [Google Scholar]
  • 104.Yan Y., Huang S.-Y. Accurate prediction of inter-protein residue–residue contacts for homo-oligomeric protein complexes. Brief Bioinform. 2021;22(5) doi: 10.1093/bib/bbab038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Roy R.S., Quadir F., Soltanikazemi E., Cheng J. A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers. Bioinformatics. 2022;38(7):1904–1910. doi: 10.1093/bioinformatics/btac063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Huang H., Zeng C., Gong X. 2021 IEEE international conference on bioinformatics and biomedicine (BIBM) IEEE; 2021. Inter-protein contact map generated only from intra-monomer by image inpainting; pp. 131–136. [Google Scholar]
  • 107.Housmans J.A., Wu G., Schymkowitz J., Rousseau F. A guide to studying protein aggregation. FEBS J. 2023;290(3):554–583. doi: 10.1111/febs.16312. [DOI] [PubMed] [Google Scholar]
  • 108.Mortuza S., Zheng W., Zhang C., Li Y., Pearce R., Zhang Y. Improving fragment-based ab initio protein structure assembly using low-accuracy contact-map predictions. Nat Commun. 2021;12(1):5011. doi: 10.1038/s41467-021-25316-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Lin P., Tao H., Li H., Huang S.-Y. Protein–protein contact prediction by geometric triangle-aware protein language models. Nat Mach Intell. 2023;5(11):1275–1284. [Google Scholar]
  • 110.Morris G.M., Lim-Wilby M. Molecular docking. Mol Model Proteins. 2008:365–382. doi: 10.1007/978-1-59745-177-2_19. [DOI] [PubMed] [Google Scholar]
  • 111.Tsuchiya Y., Yamamori Y., Tomii K. Protein–protein interaction prediction methods: from docking-based to ai-based approaches. Biophys Rev. 2022;14(6):1341–1348. doi: 10.1007/s12551-022-01032-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Shirali A., Stebliankin V., Karki U., Shi J., Chapagain P., Narasimhan G. A comprehensive survey of scoring functions for protein docking models. BMC Bioinform. 2025;26(1):25. doi: 10.1186/s12859-024-05991-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Matsuzaki Y., Uchikoga N., Ohue M., Akiyama Y. Rigid-docking approaches to explore protein–protein interaction space. Netw Biol. 2017:33–55. doi: 10.1007/10_2016_41. [DOI] [PubMed] [Google Scholar]
  • 114.Raval K., Ganatra T. Basics, types and applications of molecular docking: a review. IP Int J Compr Adv Pharmacol. 2022;7(1):12–16. [Google Scholar]
  • 115.Harmalkar A., Gray J.J. Advances to tackle backbone flexibility in protein docking. Curr Opin Struct Biol. 2021;67:178–186. doi: 10.1016/j.sbi.2020.11.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Andrusier N., Mashiach E., Nussinov R., Wolfson H.J. Principles of flexible protein–protein docking. Proteins, Struct Funct Bioinform. 2008;73(2):271–289. doi: 10.1002/prot.22170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Desta I.T., Porter K.A., Xia B., Kozakov D., Vajda S. Performance and its limits in rigid body protein-protein docking. Structure. 2020;28(9):1071–1081. doi: 10.1016/j.str.2020.06.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Wells J.N., Marsh J.A. Springer; New York, New York, NY: 2018. Experimental characterization of protein complex structure, dynamics, and assembly; pp. 3–27. [DOI] [PubMed] [Google Scholar]
  • 119.van der Heijden T., Dekker C. Monte Carlo simulations of protein assembly, disassembly, and linear motion on dna. Biophys J. 2008;95(10):4560–4569. doi: 10.1529/biophysj.108.135061. https://www.sciencedirect.com/science/article/pii/S0006349508785977 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Skolnick J., Gao M., Zhou H., Singh S. Alphafold 2: why it works and its implications for understanding the relationships of protein sequence, structure, and function. J Chem Inf Model. 2021;61(10):4827–4831. doi: 10.1021/acs.jcim.1c01114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Brookes E., Rocco M. A database of calculated solution parameters for the alphafold predicted protein structures. Sci Rep. 2022;12(1):7349. doi: 10.1038/s41598-022-10607-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Hekkelman M.L., de Vries I., Joosten R.P., Perrakis A. Alphafill: enriching alphafold models with ligands and cofactors. Nat Methods. 2023;20(2):205–213. doi: 10.1038/s41592-022-01685-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Siebenmorgen T., Menezes F., Benassou S., Merdivan E., Didi K., Mourão A.S.D., et al. Misato: machine learning dataset of protein–ligand complexes for structure-based drug discovery. Nat Comput Sci. 2024:1–12. doi: 10.1038/s43588-024-00627-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Dai X., Wu L., Yoo S., Liu Q. Integrating alphafold and deep learning for atomistic interpretation of cryo-em maps. Brief Bioinform. 2023;24(6) doi: 10.1093/bib/bbad405. [DOI] [PubMed] [Google Scholar]
  • 125.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature. 2024:1–3. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Yu Z., Yu J., Wang H., Zhang S., Zhao L., Shi S. Phosaf: an integrated deep learning architecture for predicting protein phosphorylation sites with alphafold2 predicted structures. Anal Biochem. 2024;690 doi: 10.1016/j.ab.2024.115510. [DOI] [PubMed] [Google Scholar]
  • 127.Wayment-Steele H.K., Ojoawo A., Otten R., Apitz J.M., Pitsawong W., Hömberger M., et al. Predicting multiple conformations via sequence clustering and alphafold2. Nature. 2024;625(7996):832–839. doi: 10.1038/s41586-023-06832-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Li A.J., Lu M., Desta I., Sundar V., Grigoryan G., Keating A.E. Neural network-derived Potts models for structure-based protein design using backbone atomic coordinates and tertiary motifs. Protein Sci. 2023;32(2) doi: 10.1002/pro.4554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Elnaggar A., Heinzinger M., Dallago C., Rehawi G., Wang Y., Jones L., et al. Prottrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2021;44(10):7112–7127. doi: 10.1109/TPAMI.2021.3095381. [DOI] [PubMed] [Google Scholar]
  • 130.Chen B., Cheng X., Li P., Geng Y.-a., Gong J., Li S., et al. xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein. 2024. arXiv:2401.06199 arXiv preprint. [DOI] [PubMed]
  • 131.Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 132.Fang X., Wang F., Liu L., He J., Lin D., Xiang Y., et al. Helixfold-single: Msa-free protein structure prediction by using protein language model as an alternative. 2022. arXiv:2207.13921 arXiv preprint.
  • 133.Wu R., Ding F., Wang R., Shen R., Zhang X., Luo S., et al. High-resolution de novo structure prediction from primary sequence. bioRxiv. 2022 2022–07. [Google Scholar]
  • 134.Wang W., Peng Z., Yang J. Single-sequence protein structure prediction using supervised transformer protein language models. Nat Comput Sci. 2022;2(12):804–814. doi: 10.1038/s43588-022-00373-3. [DOI] [PubMed] [Google Scholar]
  • 135.Chowdhury R., Bouatta N., Biswas S., Floristean C., Kharkar A., Roy K., et al. Single-sequence protein structure prediction using a language model and deep learning. Nat Biotechnol. 2022;40(11):1617–1623. doi: 10.1038/s41587-022-01432-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Hsu C., Verkuil R., Liu J., Lin Z., Hie B., Sercu T., et al. International conference on machine learning. PMLR; 2022. Learning inverse folding from millions of predicted structures; pp. 8946–8970. [Google Scholar]
  • 137.Brandes N., Goldman G., Wang C.H., Ye C.J., Ntranos V. Genome-wide prediction of disease variant effects with a deep protein language model. Nat Genet. 2023;55(9):1512–1522. doi: 10.1038/s41588-023-01465-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Sun D., Liu S., Gong X. Review of multimer protein–protein interaction complex topology and structure prediction. Chin Phys B. 2020;29(10) [Google Scholar]
  • 139.Yang Y., Gong X. A new probability method to understand protein-protein interface formation mechanism at amino acid level. J Theor Biol. 2018;436:18–25. doi: 10.1016/j.jtbi.2017.09.026. [DOI] [PubMed] [Google Scholar]
  • 140.Yang Y., Wang W., Lou Y., Yin J., Gong X. Geometric and amino acid type determinants for protein-protein interaction interfaces. Quant Biol. 2018;6(2):163–174. [Google Scholar]
  • 141.Wang W., Yang Y., Yin J., Gong X. Different protein-protein interface patterns predicted by different machine learning methods. Sci Rep. 2017;7(1) doi: 10.1038/s41598-017-16397-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Zhao Z., Gong X. Protein-protein interaction interface residue pair prediction based on deep learning architecture. IEEE/ACM Trans Comput Biol Bioinform. 2017;16(5):1753–1759. doi: 10.1109/TCBB.2017.2706682. [DOI] [PubMed] [Google Scholar]
  • 143.Zhao Z., Gong X. Proceedings of the 10th ACM international conference on bioinformatics, computational biology and health informatics. 2019. Trimer protein-protein complex interface interacting residue pairs prediction using deep learning approach; pp. 580–585. [Google Scholar]
  • 144.Liu J., Gong X. Attention mechanism enhanced lstm with residual architecture and its application for protein-protein interaction residue pairs prediction. BMC Bioinform. 2019;20:1–11. doi: 10.1186/s12859-019-3199-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Sun D., Gong X. Tetramer protein complex interface residue pairs prediction with lstm combined with graph representations. Biochim Biophys Acta, Proteins Proteomics. 2020;1868(11) doi: 10.1016/j.bbapap.2020.140504. [DOI] [PubMed] [Google Scholar]
  • 146.Lyu Y., He R., Hu J., Wang C., Gong X. Prediction of the tetramer protein complex interaction based on cnn and svm. Front Genet. 2023;14 doi: 10.3389/fgene.2023.1076904. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Gainza P., Sverrisson F., Monti F., Rodola E., Boscaini D., Bronstein M.M., et al. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nat Methods. 2020;17(2):184–192. doi: 10.1038/s41592-019-0666-6. [DOI] [PubMed] [Google Scholar]
  • 148.Sverrisson F., Feydy J., Correia B.E., Bronstein M.M. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. Fast end-to-end learning on protein surfaces; pp. 15272–15281. [Google Scholar]
  • 149.Zhang Z., Xu M., Jamasb A., Chenthamarakshan V., Lozano A., Das P., et al. Protein representation learning by geometric structure pretraining. 2022. arXiv:2203.06125 arXiv preprint.
  • 150.Hermosilla P., Schäfer M., Lang M., Fackelmann G., Vázquez P.P., Kozlíková B., et al. Intrinsic-extrinsic convolution and pooling for learning on 3d protein structures. 2020. arXiv:2007.06252 arXiv preprint.
  • 151.Hermosilla P., Ropinski T. Contrastive representation learning for 3d protein structures. 2022. arXiv:2205.15675 arXiv preprint.
  • 152.Szklarczyk D., Kirsch R., Koutrouli M., Nastou K., Mehryary F., Hachilif R., et al. The string database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):D638–D646. doi: 10.1093/nar/gkac1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153.Du Y., Cai M., Xing X., Ji J., Yang E., Wu J. Pina 3.0: mining cancer interactome. Nucleic Acids Res. 2021;49(D1):D1351–D1357. doi: 10.1093/nar/gkaa1075. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.Shukla V.K., Siemons L., Hansen D.F. Intrinsic structural dynamics dictate enzymatic activity and inhibition. Proc Natl Acad Sci. 2023;120(41) doi: 10.1073/pnas.2310910120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Hu W., Ohue M. Spatialppi: three-dimensional space protein-protein interaction prediction with alphafold multimer. Comput Struct Biotechnol J. 2024;23:1214–1225. doi: 10.1016/j.csbj.2024.03.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.Breckels L.M., Hutchings C., Ingole K.D., Kim S., Lilley K.S., Makwana M.V., et al. Advances in spatial proteomics: mapping proteome architecture from protein complexes to subcellular localizations. Cell Chem Biol. 2024;31(9):1665–1687. doi: 10.1016/j.chembiol.2024.08.008. [DOI] [PubMed] [Google Scholar]
  • 157.Zhang Y., Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins, Struct Funct Bioinform. 2004;57(4):702–710. doi: 10.1002/prot.20264. [DOI] [PubMed] [Google Scholar]
  • 158.Armougom F., Moretti S., Keduas V., Notredame C. The irmsd: a local measure of sequence alignment accuracy using structural information. Bioinformatics. 2006;22(14):e35–e39. doi: 10.1093/bioinformatics/btl218. [DOI] [PubMed] [Google Scholar]
  • 159.Abriata L.A. The Nobel prize in chemistry: past, present, and future of ai in biology. Commun Biol. 2024;7(1):1409. doi: 10.1038/s42003-024-07113-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160.Das R., Baker D. Macromolecular modeling with Rosetta. Annu Rev Biochem. 2008;77(1):363–382. doi: 10.1146/annurev.biochem.77.062906.171838. [DOI] [PubMed] [Google Scholar]
  • 161.Rohl C.A., Strauss C.E., Misura K.M., Baker D. Methods in enzymology, vol. 383. Elsevier; 2004. Protein structure prediction using Rosetta; pp. 66–93. [DOI] [PubMed] [Google Scholar]
  • 162.Liu J., Neupane P., Cheng J. Accurate prediction of protein complex stoichiometry by integrating alphafold3 and template information. bioRxiv. 2025 [Google Scholar]
  • 163.Wei G., Xi W., Nussinov R., Ma B. Protein ensembles: how does nature harness thermodynamic fluctuations for life? The diverse functional roles of conformational ensembles in the cell. Chem Rev. 2016;116(11):6516–6551. doi: 10.1021/acs.chemrev.5b00562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 164.Janson G., Valdes-Garcia G., Heo L., Feig M. Direct generation of protein conformational ensembles via machine learning. Nat Commun. 2023;14(1):774. doi: 10.1038/s41467-023-36443-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 165.Chen S., Zhang S., Fang X., Lin L., Zhao H., Yang Y. Protein complex structure modeling by cross-modal alignment between cryo-em maps and protein sequences. Nat Commun. 2024;15(1):8808. doi: 10.1038/s41467-024-53116-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 166.Urvas L., Chiesa L., Bret G., Jacquemard C., Kellenberger E. Benchmarking alphafold-generated structures of chemokine–chemokine receptor complexes. J Chem Inf Model. 2024 doi: 10.1021/acs.jcim.3c01835. [DOI] [PubMed] [Google Scholar]
  • 167.Jeppesen M., André I. Accurate prediction of protein assembly structure by combining alphafold and symmetrical docking. Nat Commun. 2023;14(1):8283. doi: 10.1038/s41467-023-43681-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 168.Varadi M., Velankar S. The impact of alphafold protein structure database on the fields of life sciences. Proteomics. 2023;23(17) doi: 10.1002/pmic.202200128. [DOI] [PubMed] [Google Scholar]
  • 169.Geist J.L., Lee C.Y., Strom J.M., de Jesús Naveja J., Luck K. Generation of a high confidence set of domain–domain interface types to guide protein complex structure predictions by alphafold. Bioinformatics. 2024;40(8) doi: 10.1093/bioinformatics/btae482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 170.Olechnovič K., Valančauskas L., Dapkūnas J., Venclovas Č. Prediction of protein assemblies by structure sampling followed by interface-focused scoring. Proteins, Struct Funct Bioinform. 2023;91(12):1724–1733. doi: 10.1002/prot.26569. [DOI] [PubMed] [Google Scholar]
  • 171.Mirabello C., Wallner B., Nystedt B., Azinas S., Carroni M. Unmasking alphafold to integrate experiments and predictions in multimeric complexes. Nat Commun. 2024;15(1):8724. doi: 10.1038/s41467-024-52951-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 172.Buhlheller C., Sagmeister T., Grininger C., Gubensäk N., Sleytr U.B., Usón I., et al. Symprofold: structural prediction of symmetrical biological assemblies. Nat Commun. 2024;15(1):8152. doi: 10.1038/s41467-024-52138-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 173.Ding X., Chen X., Sullivan E.E., Shay T.F., Gradinaru V. Fast, accurate ranking of engineered proteins by target-binding propensity using structure modeling. Mol Ther. 2024;32(6):1687–1700. doi: 10.1016/j.ymthe.2024.04.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 174.Keegan R.M., Simpkin A.J., Rigden D.J. The success rate of processed predicted models in molecular replacement: implications for experimental phasing in the alphafold era. Acta Crystallogr. 2024;80(11) doi: 10.1107/S2059798324009380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 175.Chen B., Xie Z., Qiu J., Ye Z., Xu J., Tang J. Improved the heterodimer protein complex prediction with protein language models. Brief Bioinform. 2023;24(4) doi: 10.1093/bib/bbad221. [DOI] [PubMed] [Google Scholar]
  • 176.Kandathil S.M., Greener J.G., Lau A.M., Jones D.T. Ultrafast end-to-end protein structure prediction enables high-throughput exploration of uncharacterized proteins. Proc Natl Acad Sci. 2022;119(4) doi: 10.1073/pnas.2113348119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 177.Trepte P., Secker C., Olivet J., Blavier J., Kostova S., Maseko S.B., et al. Ai-guided pipeline for protein–protein interaction drug discovery identifies a sars-cov-2 inhibitor. Mol Syst Biol. 2024;20(4):428–457. doi: 10.1038/s44320-024-00019-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 178.Or A., Zhang H., Freedman M.N. Virtualflow: decoupling deep learning models from the underlying hardware. Proc Mach Learn Syst. 2022;4:126–140. [Google Scholar]
  • 179.Oda T. Improving protein structure prediction with extended sequence similarity searches and deep-learning-based refinement in casp15. Proteins, Struct Funct Bioinform. 2023;91(12):1712–1723. doi: 10.1002/prot.26551. [DOI] [PubMed] [Google Scholar]
  • 180.Wallner B. Improved multimer prediction using massive sampling with alphafold in casp15. Proteins, Struct Funct Bioinform. 2023;91(12):1734–1746. doi: 10.1002/prot.26562. [DOI] [PubMed] [Google Scholar]
  • 181.Zheng W., Wuyun Q., Li Y., Zhang C., Freddolino P.L., Zhang Y. Improving deep learning protein monomer and complex structure prediction using deepmsa2 with huge metagenomics data. Nat Methods. 2024;21(2):279–289. doi: 10.1038/s41592-023-02130-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 182.Zheng W., Wuyun Q., Zhang Y. One step forward towards deep-learning protein complex structure prediction by precise multiple sequence alignment construction. Clin Transl Med. 2024;14(6) doi: 10.1002/ctm2.1689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 183.Poon B.K., Terwilliger T.C., Adams P.D. The phenix-alphafold webservice: enabling alphafold predictions for use in phenix. Protein Sci. 2024;33(5) doi: 10.1002/pro.4992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 184.McDonald E.F., Jones T., Plate L., Meiler J., Gulsevin A. Benchmarking alphafold2 on peptide structure prediction. Structure. 2023;31(1):111–119. doi: 10.1016/j.str.2022.11.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 185.Morehead A., Liu J., Cheng J. Protein structure accuracy estimation using geometry-complete perceptron networks. Protein Sci. 2024;33(3) doi: 10.1002/pro.4932. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 186.Zhang L., Wang S., Hou J., Si D., Zhu J., Cao R. Complexqa: a deep graph learning approach for protein complex structure assessment. Brief Bioinform. 2023;24(6) doi: 10.1093/bib/bbad287. [DOI] [PubMed] [Google Scholar]
  • 187.Zhu W., Shenoy A., Kundrotas P., Elofsson A. Evaluation of alphafold-multimer prediction on multi-chain protein complexes. Bioinformatics. 2023;39(7) doi: 10.1093/bioinformatics/btad424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 188.Bryant P., Pozzati G., Zhu W., Shenoy A., Kundrotas P., Elofsson A. Predicting the structure of large protein complexes using alphafold and Monte Carlo tree search. Nat Commun. 2022;13(1):6028. doi: 10.1038/s41467-022-33729-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 189.Lee C.Y., Hubrich D., Varga J.K., Schäfer C., Welzel M., Schumbera E., et al. Systematic discovery of protein interaction interfaces using alphafold and experimental validation. Mol Syst Biol. 2024;20(2):75–97. doi: 10.1038/s44320-023-00005-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 190.Boyd L.F., Jiang J., Ahmad J., Natarajan K., Margulies D.H. Experimental structures of antibody/mhc-i complexes reveal details of epitopes overlooked by computational prediction. J Immunol. 2024;212(8):1366–1380. doi: 10.4049/jimmunol.2300839. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 191.Marciano S., Dey D., Listov D., Fleishman S.J., Sonn-Segev A., Mertens H., et al. Protein quaternary structures in solution are a mixture of multiple forms. Chem Sci. 2022;13(39):11680–11695. doi: 10.1039/d2sc02794a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 192.Myler L.R., Toia B., Vaughan C.K., Takai K., Matei A.M., Wu P., et al. Dna-pk and the trf2 iddr inhibit mrn-initiated resection at leading-end telomeres. Nat Struct Mol Biol. 2023;30(9):1346–1356. doi: 10.1038/s41594-023-01072-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 193.Weng Q., Wan L., Straker G.C., Deegan T.D., Duncker B.P., Neiman A.M., et al. An acidic loop in the forkhead-associated domain of the yeast meiosis-specific kinase mek1 interacts with a specific motif in a subset of mek1 substrates. Genetics. 2024;228(1) doi: 10.1093/genetics/iyae106. iyae106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 194.Novikova I.V., Soldatova A.V., Moser T.H., Thibert S.M., Romano C.A., Zhou M., et al. Cryo-em structure of the mnx protein complex reveals a tunnel framework for the mechanism of manganese biomineralization. J Am Chem Soc. 2024;146(33):22950–22958. doi: 10.1021/jacs.3c06537. [DOI] [PubMed] [Google Scholar]
  • 195.Chen J., Zia A., Luo A., Meng H., Wang F., Hou J., et al. Enhancing cryo-em structure prediction with deeptracer and alphafold2 integration. Brief Bioinform. 2024;25(3) doi: 10.1093/bib/bbae118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 196.Bryant P., Noé F. Improved protein complex prediction with alphafold-multimer by denoising the msa profile. PLoS Comput Biol. 2024;20(7) doi: 10.1371/journal.pcbi.1012253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 197.Grupp B., Denkhaus L., Gerhardt S., Vögele M., Johnsson N., Gronemeyer T. The structure of a tetrameric septin complex reveals a hydrophobic element essential for nc-interface integrity. Commun Biol. 2024;7(1):48. doi: 10.1038/s42003-023-05734-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 198.Homma F., Huang J., van der Hoorn R.A. Alphafold-multimer predicts cross-kingdom interactions at the plant-pathogen interface. Nat Commun. 2023;14(1):6040. doi: 10.1038/s41467-023-41721-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 199.Wang X., Zhang X., Zhou J., Wang W., Wang X., Xu B. An in silico investigation of kv2. 1 potassium channel: model building and inhibitors binding sites analysis. Mol Inform. 2023;42(12) doi: 10.1002/minf.202300072. [DOI] [PubMed] [Google Scholar]
  • 200.Terwilliger T.C., Liebschner D., Croll T.I., Williams C.J., McCoy A.J., Poon B.K., et al. Alphafold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. Nat Methods. 2024;21(1):110–116. doi: 10.1038/s41592-023-02087-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 201.Hou Y., Xie T., He L., Tao L., Huang J. Topological links in predicted protein complex structures reveal limitations of alphafold. Commun Biol. 2023;6(1):1098. doi: 10.1038/s42003-023-05489-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 202.Soleymani F., Paquet E., Viktor H., Michalowski W., Spinello D. Protein–protein interaction prediction with deep learning: a comprehensive review. Comput Struct Biotechnol J. 2022;20:5316–5341. doi: 10.1016/j.csbj.2022.08.070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 203.Vazquez A., Flammini A., Maritan A., Vespignani A. Global protein function prediction from protein-protein interaction networks. Nat Biotechnol. 2003;21(6):697–700. doi: 10.1038/nbt825. [DOI] [PubMed] [Google Scholar]
  • 204.Deng M., Zhang K., Mehta S., Chen T., Sun F. Proceedings. IEEE computer society bioinformatics conference. IEEE; 2002. Prediction of protein function using protein-protein interaction data; pp. 197–206. [PubMed] [Google Scholar]
  • 205.Bernhofer M., Dallago C., Karl T., Satagopam V., Heinzinger M., Littmann M., et al. Predictprotein-predicting protein structure and function for 29 years. Nucleic Acids Res. 2021;49(W1):W535–W540. doi: 10.1093/nar/gkab354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 206.Burke D.F., Bryant P., Barrio-Hernandez I., Memon D., Pozzati G., Shenoy A., et al. Towards a structurally resolved human protein interaction network. Nat Struct Mol Biol. 2023;30(2):216–225. doi: 10.1038/s41594-022-00910-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 207.Zhang S.-T., Deng S.-K., Li T., Maloney M.E., Li D.-F., Spain J.C., et al. Discovery of the 1-naphthylamine biodegradation pathway reveals a broad-substrate-spectrum enzyme catalyzing 1-naphthylamine glutamylation. eLife. 2024;13 doi: 10.7554/eLife.95555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 208.Blackledge N.P., Klose R.J. The molecular principles of gene regulation by polycomb repressive complexes. Nat Rev Mol Cell Biol. 2021;22(12):815–833. doi: 10.1038/s41580-021-00398-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 209.Törönen P., Holm L. Pannzer—a practical tool for protein function prediction. Protein Sci. 2022;31(1):118–128. doi: 10.1002/pro.4193. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 210.Zheng W., Wuyun Q., Zhou X., Li Y., Freddolino P.L., Zhang Y. Lomets3: integrating deep learning and profile alignment for advanced protein template recognition and function annotation. Nucleic Acids Res. 2022;50(W1):W454–W464. doi: 10.1093/nar/gkac248. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 211.Jain A., Mittal S., Tripathi L.P., Nussinov R., Ahmad S. Host-pathogen protein-nucleic acid interactions: a comprehensive review. Comput Struct Biotechnol J. 2022;20:4415–4436. doi: 10.1016/j.csbj.2022.08.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 212.Herianto S., Subramani B., Chen B.-R., Chen C.-S. Recent advances in liposome development for studying protein-lipid interactions. Crit Rev Biotechnol. 2024;44(1):1–14. doi: 10.1080/07388551.2022.2111294. [DOI] [PubMed] [Google Scholar]
  • 213.Shoichet B.K., Baase W.A., Kuroki R., Matthews B.W. A relationship between protein stability and protein function. Proc Natl Acad Sci. 1995;92(2):452–456. doi: 10.1073/pnas.92.2.452. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 214.Mou K., Mukhtar F., Khan M.T., Darwish D.B., Peng S., Muhammad S., et al. Emerging mutations in nsp1 of sars-cov-2 and their effect on the structural stability. Pathogens. 2021;10(10):1285. doi: 10.3390/pathogens10101285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 215.Lin P., Yan Y., Tao H., Huang S.-Y. Deep transfer learning for inter-chain contact predictions of transmembrane protein complexes. Nat Commun. 2023;14(1):4935. doi: 10.1038/s41467-023-40426-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 216.Bryant P., Kelkar A., Guljas A., Clementi C., Noé F. Structure prediction of protein-ligand complexes from sequence information with umol. Nat Commun. 2024;15(1):4536. doi: 10.1038/s41467-024-48837-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 217.Lupo U., Sgarbossa D., Bitbol A.-F. Pairing interacting protein sequences using masked language modeling. Proc Natl Acad Sci. 2024;121(27) doi: 10.1073/pnas.2311887121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 218.Yu D., Chojnowski G., Rosenthal M., Kosinski J. Alphapulldown—a python package for protein–protein interaction screens using alphafold-multimer. Bioinformatics. 2023;39(1) doi: 10.1093/bioinformatics/btac749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 219.Li R., Jiao P., Li J. International conference on intelligent computing. Springer; 2024. Pf2pi: protein function prediction based on alphafold2 information and protein-protein interaction; pp. 278–289. [Google Scholar]
  • 220.McLean T.C. Lazyaf, a pipeline for accessible medium-scale in silico prediction of protein-protein interactions. Microbiology. 2024;170(7) doi: 10.1099/mic.0.001473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 221.Jahn L.R., Marquet C., Heinzinger M., Rost B. Protein embeddings predict binding residues in disordered regions. Sci Rep. 2024;14(1) doi: 10.1038/s41598-024-64211-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 222.Harmalkar A., Lyskov S., Gray J.J. Reliable protein-protein docking with alphafold, Rosetta, and replica-exchange. bioRxiv. 2023 doi: 10.7554/eLife.94029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 223.Johansson-Åkhe I., Wallner B. Improving peptide-protein docking with alphafold-multimer using forced sampling. Front Bioinform. 2022;2 doi: 10.3389/fbinf.2022.959160. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 224.Kim J., McFee M., Fang Q., Abdin O., Kim P.M. Computational and artificial intelligence-based methods for antibody development. Trends Pharmacol Sci. 2023;44(3):175–189. doi: 10.1016/j.tips.2022.12.005. [DOI] [PubMed] [Google Scholar]
  • 225.McCoy K.M., Ackerman M.E., Grigoryan G. A comparison of antibody–antigen complex sequence-to-structure prediction methods and their systematic biases. Protein Sci. 2024;33(9) doi: 10.1002/pro.5127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 226.Ruffolo J.A., Chu L.-S., Mahajan S.P., Gray J.J. Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies. Nat Commun. 2023;14(1):2389. doi: 10.1038/s41467-023-38063-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 227.Chen H., Fan X., Zhu S., Pei Y., Zhang X., Zhang X., et al. Accurate prediction of cdr-h3 loop structures of antibodies with deep learning. eLife. 2024;12 doi: 10.7554/eLife.91512. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 228.Richards A.L., Eckhardt M., Krogan N.J. Mass spectrometry-based protein–protein interaction networks for the study of human diseases. Mol Syst Biol. 2021;17(1) doi: 10.15252/msb.20188792. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 229.Susa K.J., Kruse A.C., Blacklow S.C. Tetraspanins: structure, dynamics, and principles of partner-protein recognition. Trends Cell Biol. 2024;34(6):509–522. doi: 10.1016/j.tcb.2023.09.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 230.Chauhan V.M., Pantazes R.J. Analysis of conformational stability of interacting residues in protein binding interfaces. Protein Eng Des Sel. 2023;36 doi: 10.1093/protein/gzad016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 231.Wee J., Wei G.-W. Evaluation of alphafold 3's protein–protein complexes for predicting binding free energy changes upon mutation. J Chem Inf Model. 2024;64(16):6676–6683. doi: 10.1021/acs.jcim.4c00976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 232.Hawkins-Hooker A., Cerigo D.B., Lupo U., Jones D., Paige B. ICML 2024 workshop on efficient and accessible foundation models for biological discovery. 2024. MSA pairing transformer: protein interaction partner prediction with few-shot contrastive learning. [Google Scholar]
  • 233.Poitras C., Lamontagne F., Grandvaux N., Song H., Pinard M., Coulombe B. High-accuracy mapping of human and viral direct physical protein-protein interactions using the novel computational system alphafold-pairs. bioRxiv. 2023 2023–08. [Google Scholar]
  • 234.Xia S., Li D., Deng X., Liu Z., Zhu H., Liu Y., et al. Integration of protein sequence and protein–protein interaction data by hypergraph learning to identify novel protein complexes. Brief Bioinform. 2024;25(4) doi: 10.1093/bib/bbae274. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 235.Michaelis A.C., Brunner A.-D., Zwiebel M., Meier F., Strauss M.T., Bludau I., et al. The social and structural architecture of the yeast protein interactome. Nature. 2023;624(7990):192–200. doi: 10.1038/s41586-023-06739-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 236.Binder J.L., Berendzen J., Stevens A.O., He Y., Wang J., Dokholyan N.V., et al. Alphafold illuminates half of the dark human proteins. Curr Opin Struct Biol. 2022;74 doi: 10.1016/j.sbi.2022.102372. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 237.Kosugi T., Ohue M. Design of cyclic peptides targeting protein–protein interactions using alphafold. Int J Mol Sci. 2023;24(17) doi: 10.3390/ijms241713257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 238.Gupta S., Nerli S., Kutti Kandy S., Mersky G.L., Sgourakis N.G. Hla3db: comprehensive annotation of peptide/hla complexes enables blind structure prediction of t cell epitopes. Nat Commun. 2023;14(1):6349. doi: 10.1038/s41467-023-42163-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 239.Petrey D., Zhao H., Trudeau S.J., Murray D., Honig B. Preppi: a structure informed proteome-wide database of protein–protein interactions. J Mol Biol. 2023;435(14) doi: 10.1016/j.jmb.2023.168052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 240.Schmid E.W., Walter J.C. Predictomes: a classifier-curated database of alphafold-modeled protein-protein interactions. bioRxiv. 2024 doi: 10.1016/j.molcel.2025.01.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 241.Woo H., Kim Y., Seok C. Protein loop structure prediction by community-based deep learning and its application to antibody cdr h3 loop modeling. PLoS Comput Biol. 2024;20(6) doi: 10.1371/journal.pcbi.1012239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 242.Zhao N., Han B., Zhao C., Xu J., Gong X. Abag-docking benchmark: a non-redundant structure benchmark dataset for antibody–antigen computational docking. Brief Bioinform. 2024;25(2) doi: 10.1093/bib/bbae048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 243.Gaudreault F., Corbeil C.R., Sulea T. Enhanced antibody-antigen structure prediction from molecular docking using alphafold2. Sci Rep. 2023;13(1) doi: 10.1038/s41598-023-42090-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 244.Di Ianni A., Di Ianni A., Cowan K., Barbero L.M., Sirtori F.R. Leveraging cross-linking mass spectrometry for modeling antibody–antigen complexes. J Proteome Res. 2024;23(3):1049–1061. doi: 10.1021/acs.jproteome.3c00816. [DOI] [PubMed] [Google Scholar]
  • 245.Träger T., Kastritis P.L. Cracking the code of cellular protein–protein interactions: alphafold and whole-cell crosslinking to the rescue. Mol Syst Biol. 2023;19(4) doi: 10.15252/msb.202311587. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 246.Pei J., Zhang J., Cong Q. Computational analysis of protein–protein interactions of cancer drivers in renal cell carcinoma. FEBS Open Bio. 2024;14(1):112–126. doi: 10.1002/2211-5463.13732. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 247.González-Avendaño M., López J., Vergara-Jaque A., Cerda O. The power of computational proteomics platforms to decipher protein-protein interactions. Curr Opin Struct Biol. 2024;88 doi: 10.1016/j.sbi.2024.102882. [DOI] [PubMed] [Google Scholar]
  • 248.Balasco N., Ruggiero A., Smaldone G., Pecoraro G., Coppola L., Pirone L., et al. Structural studies of kctd1 and its disease-causing mutant p20s provide insights into the protein function and misfunction. Int J Biol Macromol. 2024;277 doi: 10.1016/j.ijbiomac.2024.134390. [DOI] [PubMed] [Google Scholar]
  • 249.Bai W., Li B., Wu P., Li X., Huang X., Shi N., et al. The first structure of human golm1 coiled coil domain reveals an unexpected tetramer and highlights its structural diversity. Int J Biol Macromol. 2024;275 doi: 10.1016/j.ijbiomac.2024.133624. [DOI] [PubMed] [Google Scholar]
  • 250.Baryshev A., La Fleur A., Groves B., Michel C., Baker D., Ljubetič A., et al. Massively parallel measurement of protein–protein interactions by sequencing using mp3-seq. Nat Chem Biol. 2024:1–10. doi: 10.1038/s41589-024-01718-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 251.Chen S.-J., Hassan M., Jernigan R.L., Jia K., Kihara D., Kloczkowski A., et al. Protein folds vs. protein folding: differing questions, different challenges. Proc Natl Acad Sci. 2023;120(1) doi: 10.1073/pnas.2214423119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 252.Callaway E. What's next for the ai protein-folding revolution. Nature. 2022;604(7905):234–238. doi: 10.1038/d41586-022-00997-5. [DOI] [PubMed] [Google Scholar]
  • 253.Shimanovich U., Hartl F.U. 2024. Protein folding: from physico-chemical rules and cellular machineries of protein quality control to ai solutions. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 254.Moore P.B., Hendrickson W.A., Henderson R., Brunger A.T. The protein-folding problem: not yet solved. Science. 2022;375(6580):507. doi: 10.1126/science.abn9422. [DOI] [PubMed] [Google Scholar]
  • 255.Zhao K., Xia Y., Zhang F., Zhou X., Li S.Z., Zhang G. Protein structure and folding pathway prediction based on remote homologs recognition using pathreader. Commun Biol. 2023;6(1):243. doi: 10.1038/s42003-023-04605-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 256.Zhao C., Liu T., Wang Z. Panda-3d: protein function prediction based on alphafold models. NAR Genomics Bioinform. 2024;6(3) doi: 10.1093/nargab/lqae094. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 257.Lu W., Zhang J., Huang W., Zhang Z., Jia X., Wang Z., et al. DynamicBind: predicting ligand-specific protein-ligand complex structure with a deep equivariant generative model. Nat Commun. 2024;15(1):1071. doi: 10.1038/s41467-024-45461-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 258.Klukowski P., Riek R., Güntert P. Time-optimized protein nmr assignment with an integrative deep learning approach using alphafold and chemical shift prediction. Sci Adv. 2023;9(47) doi: 10.1126/sciadv.adi9323. eadi9323. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 259.Chu L.-S., Ruffolo J.A., Harmalkar A., Gray J.J. Flexible protein–protein docking with a multitrack iterative transformer. Protein Sci. 2024;33(2) doi: 10.1002/pro.4862. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 260.Guan X., Tang Q.-Y., Ren W., Chen M., Wang W., Wolynes P.G., et al. Predicting protein conformational motions using energetic frustration analysis and alphafold2. Proc Natl Acad Sci. 2024;121(35) doi: 10.1073/pnas.2410662121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 261.Audagnotto M., Czechtizky W., De Maria L., Käck H., Papoian G., Tornberg L., et al. Machine learning/molecular dynamic protein structure prediction approach to investigate the protein conformational ensemble. Sci Rep. 2022;12(1) doi: 10.1038/s41598-022-13714-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 262.Verburgt J., Zhang Z., Kihara D. Multi-level analysis of intrinsically disordered protein docking methods. Methods. 2022;204:55–63. doi: 10.1016/j.ymeth.2022.05.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 263.Guo H.-B., Huntington B., Perminov A., Smith K., Hastings N., Dennis P., et al. Alphafold2 modeling and molecular dynamics simulations of an intrinsically disordered protein. PLoS ONE. 2024;19(5) doi: 10.1371/journal.pone.0301866. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 264.Piovesan D., Monzon A.M., Tosatto S.C. Intrinsic protein disorder and conditional folding in alphafolddb. Protein Sci. 2022;31(11) doi: 10.1002/pro.4466. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 265.Pietrek L.M., Stelzl L.S., Hummer G. Structural ensembles of disordered proteins from hierarchical chain growth and simulation. Curr Opin Struct Biol. 2023;78 doi: 10.1016/j.sbi.2022.102501. [DOI] [PubMed] [Google Scholar]
  • 266.Piovesan D., Del Conte A., Clementel D., Monzon A.M., Bevilacqua M., Aspromonte M.C., et al. Mobidb: 10 years of intrinsically disordered proteins. Nucleic Acids Res. 2023;51(D1):D438–D444. doi: 10.1093/nar/gkac1065. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 267.Aspromonte M.C., Nugnes M.V., Quaglia F., Bouharoua A., Tosatto S.C., Piovesan D. Disprot in 2024: improving function annotation of intrinsically disordered proteins. Nucleic Acids Res. 2024;52(D1):D434–D441. doi: 10.1093/nar/gkad928. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 268.Tesei G., Trolle A.I., Jonsson N., Betz J., Knudsen F.E., Pesce F., et al. Conformational ensembles of the human intrinsically disordered proteome. Nature. 2024;626(8000):897–904. doi: 10.1038/s41586-023-07004-5. [DOI] [PubMed] [Google Scholar]
  • 269.Bonin J.P., Aramini J.M., Dong Y., Wu H., Kay L.E. Alphafold2 as a replacement for solution nmr structure determination of small proteins: not so fast! J Magn Res. 2024 doi: 10.1016/j.jmr.2024.107725. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 270.Ruff K.M., Pappu R.V. Alphafold and implications for intrinsically disordered proteins. J Mol Biol. 2021;433(20) doi: 10.1016/j.jmb.2021.167208. [DOI] [PubMed] [Google Scholar]
  • 271.Monzon V., Haft D.H., Bateman A. Folding the unfoldable: using alphafold to explore spurious proteins. Bioinform Adv. 2022;2(1) doi: 10.1093/bioadv/vbab043. vbab043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 272.Guo H.-B., Perminov A., Bekele S., Kedziora G., Farajollahi S., Varaljay V., et al. Alphafold2 models indicate that protein sequence determines both structure and dynamics. Sci Rep. 2022;12(1) doi: 10.1038/s41598-022-14382-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 273.Ma P., Li D.-W., Brüschweiler R. Predicting protein flexibility with alphafold. Proteins, Struct Funct Bioinform. 2023;91(6):847–855. doi: 10.1002/prot.26471. [DOI] [PubMed] [Google Scholar]
  • 274.Monteiro da Silva G., Cui J.Y., Dalgarno D.C., Lisi G.P., Rubenstein B.M. High-throughput prediction of protein conformational distributions with subsampled alphafold2. Nat Commun. 2024;15(1):2464. doi: 10.1038/s41467-024-46715-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 275.Karelina M., Noh J.J., Dror R.O. How accurately can one predict drug binding modes using alphafold models? eLife. 2023;12 doi: 10.7554/eLife.89386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 276.Urban P., Pompon D. Confrontation of alphafold models with experimental structures enlightens conformational dynamics supporting cyp102a1 functions. Sci Rep. 2022;12(1) doi: 10.1038/s41598-022-20390-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 277.Ohnuki J., Okazaki K.-i. Integration of alphafold with molecular dynamics for efficient conformational sampling of transporter protein nark. J Phys Chem B. 2024;128(31):7530–7537. doi: 10.1021/acs.jpcb.4c02726. [DOI] [PubMed] [Google Scholar]
  • 278.Ohnuki J., Jaunet-Lahary T., Yamashita A., Okazaki K.-i. Accelerated molecular dynamics and alphafold uncover a missing conformational state of transporter protein oxlt. J Phys Chem Lett. 2024;15(3):725–732. doi: 10.1021/acs.jpclett.3c03052. [DOI] [PubMed] [Google Scholar]
  • 279.Díaz-Holguín A., Saarinen M., Vo D.D., Sturchio A., Branzell N., Cabeza de Vaca I., et al. Alphafold accelerated discovery of psychotropic agonists targeting the trace amine–associated receptor 1. Sci Adv. 2024;10(32) doi: 10.1126/sciadv.adn1524. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 280.Tome D. Criteria and markers for protein quality assessment–a review. Br J Nutr. 2012;108(S2):S222–S229. doi: 10.1017/S0007114512002565. [DOI] [PubMed] [Google Scholar]
  • 281.Adhikari S., Schop M., de Boer I.J., Huppertz T. Protein quality in perspective: a review of protein quality metrics and their applications. Nutrients. 2022;14(5):947. doi: 10.3390/nu14050947. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 282.Liang F., Sun M., Xie L., Zhao X., Liu D., Zhao K., et al. Recent advances and challenges in protein complex model accuracy estimation. Comput Struct Biotechnol J. 2024;23:1824–1832. doi: 10.1016/j.csbj.2024.04.049. https://www.sciencedirect.com/science/article/pii/S2001037024001363 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 283.Pearce R., Zhang Y. Deep learning techniques have significantly impacted protein structure prediction and protein design. Curr Opin Struct Biol. 2021;68:194–207. doi: 10.1016/j.sbi.2021.01.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 284.Schauperl M., Denny R.A. Ai-based protein structure prediction in drug discovery: impacts and challenges. J Chem Inf Model. 2022;62(13):3142–3156. doi: 10.1021/acs.jcim.2c00026. [DOI] [PubMed] [Google Scholar]
  • 285.Zemla A. Lga: a method for finding 3d similarities in protein structures. Nucleic Acids Res. 2003;31(13):3370–3374. doi: 10.1093/nar/gkg571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 286.Studer G., Tauriello G., Schwede T. Assessment of the assessment—all about complexes. Proteins, Struct Funct Bioinform. 2023;91(12):1850–1860. doi: 10.1002/prot.26612. [DOI] [PubMed] [Google Scholar]
  • 287.Bertoni M., Kiefer F., Biasini M., Bordoli L., Schwede T. Modeling protein quaternary structure of homo- and hetero-oligomers beyond binary interactions by homology. Sci Rep. 2017;7(1) doi: 10.1038/s41598-017-09654-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 288.Liu J., Guo Z., Wu T., Roy R.S., Quadir F., Chen C., et al. Enhancing alphafold-multimer-based protein complex structure prediction with multicom in casp15. Commun Biol. 2023;6(1):1140. doi: 10.1038/s42003-023-05525-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 289.van Kempen M., Kim S.S., Tumescheit C., Mirdita M., Gilchrist C.L., Söding J., et al. Foldseek: fast and accurate protein structure search. bioRxiv. 2022 doi: 10.1038/s41587-023-01773-0. 2022–02. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 290.Edmunds N.S., Alharbi S.M., Genc A.G., Adiyaman R., McGuffin L.J. Estimation of model accuracy in casp15 using the m odfolddock server. Proteins, Struct Funct Bioinform. 2023;91(12):1871–1878. doi: 10.1002/prot.26532. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 291.McGuffin L.J., Alhaddad S.N., Behzadi B., Edmunds N.S., Genc A.G., Adiyaman R. Prediction and quality assessment of protein quaternary structure models using the MultiFOLD2 and ModFOLDdock2 servers. Nucleic Acids Res. 2025 doi: 10.1093/nar/gkaf336. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 292.Liu J., Liu D., Zhang G.-J. Deepumqa3: a web server for accurate assessment of interface residue accuracy in protein complexes. Bioinformatics. 2023;39(10) doi: 10.1093/bioinformatics/btad591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 293.Zhang J., Durham J., Cong Q. Revolutionizing protein–protein interaction prediction with deep learning. Curr Opin Struct Biol. 2024;85 doi: 10.1016/j.sbi.2024.102775. [DOI] [PubMed] [Google Scholar]
  • 294.Durham J., Zhang J., Humphreys I.R., Pei J., Cong Q. Recent advances in predicting and modeling protein–protein interactions. Trends Biochem Sci. 2023;48(6):527–538. doi: 10.1016/j.tibs.2023.03.003. [DOI] [PubMed] [Google Scholar]
  • 295.Kellici T.F., Hristozov D., Morao I. Ai-based protein structure predictions and their implications in drug discovery. Comput Drug Discov: Methods Appl. 2024;1:227–253. [Google Scholar]
  • 296.Frappier V., Keating A.E. Data-driven computational protein design. Curr Opin Struct Biol. 2021;69:63–69. doi: 10.1016/j.sbi.2021.03.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 297.Goverde C.A., Wolf B., Khakzad H., Rosset S., Correia B.E. De novo protein design by inversion of the alphafold structure prediction network. Protein Sci. 2023;32(6) doi: 10.1002/pro.4653. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 298.Chen L., Li Q., Nasif K.F.A., Xie Y., Deng B., Niu S., et al. Ai-driven deep learning techniques in protein structure prediction. Int J Mol Sci. 2024;25(15):8426. doi: 10.3390/ijms25158426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 299.Varadi M., Bertoni D., Magana P., Paramval U., Pidruchna I., Radhakrishnan M., et al. Alphafold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024;52(D1):D368–D375. doi: 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 300.Feldman J., Skolnick J. Af3complex yields improved structural predictions of protein complexes. bioRxiv. 2025 2025–02. [Google Scholar]
  • 301.Mirdita M., Schütze K., Moriwaki Y., Heo L., Ovchinnikov S., Steinegger M. Colabfold: making protein folding accessible to all. Nat Methods. 2022;19(6):679–682. doi: 10.1038/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 302.Chim H.Y., Elofsson A. Molpc2: improved prediction of large protein complex structures and stoichiometry using Monte Carlo tree search and alphafold2. Bioinformatics. 2024;40(6) doi: 10.1093/bioinformatics/btae329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 303.Hu W., Ohue M. Spatialppiv2: enhancing protein–protein interaction prediction through graph neural networks with protein language models. Comput Struct Biotechnol J. 2025 [Google Scholar]
  • 304.Dunbar J., Krawczyk K., Leem J., Baker T., Fuchs A., Georges G., et al. Sabdab: the structural antibody database. Nucleic Acids Res. 2014;42(D1):D1140–D1146. doi: 10.1093/nar/gkt1043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 305.Szklarczyk D., Nastou K., Koutrouli M., Kirsch R., Mehryary F., Hachilif R., et al. The string database in 2025: protein networks with directionality of regulation. Nucleic Acids Res. 2025;53(D1):D730–D737. doi: 10.1093/nar/gkae1113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 306.Oughtred R., Rust J., Chang C., Breitkreutz B.-J., Stark C., Willems A., et al. The biogrid database: a comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Sci. 2021;30(1):187–200. doi: 10.1002/pro.3978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 307.Kong X., Huang W., Liu Y. End-to-end full-atom antibody design. 2023. arXiv:2302.00203 arXiv preprint.
  • 308.Vander Meersche Y., Cretin G., Gheeraert A., Gelly J.-C., Galochkina T. Atlas: protein flexibility description from atomistic molecular dynamics simulations. Nucleic Acids Res. 2024;52(D1):D384–D392. doi: 10.1093/nar/gkad1084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 309.Liu J., Neupane P., Cheng J. Improving AlphaFold2 and 3-based protein complex structure prediction with MULTICOM4 in CASP16. bioRxiv. 2025 doi: 10.1002/prot.26850. 2025–03. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Computational and Structural Biotechnology Journal are provided here courtesy of AAAS Science Partner Journal Program

RESOURCES