Abstract
In this review, we discuss the various different types of learnable protein representations that have been used in computational biology, with a particular focus on representations that have been used in the paradigm of predicting drug-target affinity. We explore this from multiple perspectives: the source of protein information used, the training paradigms used in generating and applying such representations, and the types of (deep-learning-based) encoding or embedding methods that have been used to generate and operate on such representations. We focus on drug-target affinity due to its particular relevance and utility in the field of drug development and assessment, and we make suggestions for how drug-target affinity prediction methods development can be further improved by examining the current literature from the aforementioned perspectives. This survey thus serves as a valuable resource for researchers seeking to develop methods for predicting drug-target affinity by exploring how protein information has been used and could be used in effective ways to improve such predictions.
Keywords: Graph neural networks, Structural biology, Deep learning, Protein representation learning
Introduction
Proteins and protein representation learning
Proteins are generally considered to be one of the fundamental functional units of biology. They perform or facilitate a wide array of biological functions both within and outside the cell, and their structure is generally determined by the sequence of amino acids that comprise them. Proteins can fold into a variety of different, often-related conformations that can determine their stability, interaction with other biological elements, and even how they respond to various stresses, among many others. This relationship between sequence, structure, and function is a critical element for understanding biology [1, 2].
For these reasons, the field of computational biology often includes trying to make use of protein information in some form to predict a variety of protein properties, including structure and function. This comprises the broad field of protein representation learning, where researchers try to generate some meaningful representation of a protein (often in some lower-dimensional, continuous latent space) that can be used for a specific downstream task, for a generalized set of tasks, or a combination of both [3–5].
Drug-target affinity
In this review, we focus on the task of predicting drug-target affinity (DTA). In this paradigm, one attempts to predict the binding affinity of small molecules (“drugs”) with proteins (“targets”) by using information about those drugs and proteins. To do so, researchers have made use of various types of protein representations to use alongside a variety of molecular representations. DTA is particularly valuable in drug development and research in part because how tightly a drug binds to a protein can have various implications for the utility and value of that drug. The gold standard for determining DTA is through experimental screens, but this is often costly and impractical, especially when assaying a wide variety of possible drugs or when targeting a large array of possible proteins.
As such, computational DTA prediction methods have become increasingly popular in the past decades, especially in the past 5 years. To illustrate the growing popularity in these methods and their utility, we performed a PubMed search for “(drug-target affinity prediction) NOT review [publication type]” to get all articles that reference DTA prediction that were not reviews and plot the results for the past 20 years in Figure 1 (including papers in 2025 up until December 10, 2025). As can be seen, the field has become particularly more popular in the last 5 years, with over 500 articles being published in 2024 and well over 800 papers in 2025, showing increasing interest in the field. We do note that not all of these papers are papers that propose new DTA methods; in fact, the vast majority are papers that make use of DTA methods for other analyses, such as developing new drugs or assessing the utility of a new drug. This further highlights how useful DTA methods have become in the past few years alone, and further research into the field is warranted to continue maximizing their utility.
Fig. 1.

PubMed results count on a per-year basis from 2005–2025 for papers that mention "drug-target affinity prediction" that are not reviews. This includes many papers that make use of DTA prediction methods as well as papers that propose new DTA prediction methods. The search was performed on December 10, 2025 and includes papers up to that date for 2025
We focus on drug-target affinity in this review due to several observations: 1) that powerful and highly-enriched methods for generating protein representations have been developed and used for a variety of other tasks and paradigms; 2) DTA represents an underutilization of some of these increasingly powerful representations (as seen in other paradigms); and 3) the datasets for making use of these methods and representations are becoming increasingly accessible and as such the field of DTA prediction may be able to make use of them effectively.
It is important to note that drug-target affinity typically involves generating representations for proteins and ligands separately and that we focus on protein representations in this review because of the aforementioned improvements in protein representation learning in other paradigms. Specifically, we do not focus on ligand (also known as small-molecule or drug) representations as the field of small molecule representation learning is more nascent with a less broad set of representation options that have not already been used in DTA prediction. We direct readers to a review by McGibbon et al., particularly their discussion of learned representations [6]. We hope that future work in small molecule representations will also improve DTA prediction alongside improved protein representations.
Additionally, we note that we did not evaluate work that generates representations for a joint protein-ligand complex due to the increased difficulty in generating such complexes for arbitrary protein-ligand pairs and because the field of joint representation learning has a very different scope and set of considerations with fewer lessons that can be learned from protein representation learning; for those interested in learning more about representations protein-ligand complex, we direct the reader to work by Ragoza et al. [7] and Li et al. [8] who developed methods for representations of the joint complex and describe many other methods for the same.
Overview of our review
In this review, we discuss protein representations broadly with a focus on representations that have been used in DTA. We do so from the lens of understanding what protein modalities or information have been used as sources for such representations, the training paradigms that have been used in the past to generate protein representations, and what deep learning encoders have been used to actually make protein representations. We do all of this with a focus in DTA prediction and assess how the field could be further improved based on work done in general protein representation learning or based on our observed gaps in the field thus far.
Protein representations: a general categorization
Sources of protein information
To generate protein representations, researchers have made use of a variety of different sources of information that may be available about a protein. These include, in general order of complexity: (1) its amino acid sequence; (2) a contact map of the interaction of amino acid residues; (3) the actual 3D structure of the protein; (4) the dynamics of the protein structure over time or over multiple conformations. We provide a visualization of each of these in Figure 2.
Fig. 2.
An illustration of the modalities or information that can be utilized for a given protein. The pictured protein and all information thereof, including the multiple conformations, is from Protein DataBank ID 1IYT: the Alzheimer’s disease amyloid beta peptide, deposited by Crescenzi et al. [10]
The complexity of what information researchers choose to use often reflects the hierarchical organization of protein structure within biology. For example, the structure of a protein (and thus its contact map or 3D structure) can be determined by the amino acid sequence alone, though actually determining that structure from a sequence is technically an unsolved problem (albeit one that has been heavily advanced upon in recent years with the advent of tools like AlphaFold2 [9]). Essentially, as one advances from one-dimensional (sequence) to three-dimensional (3D structure) or four-dimensional (molecular dynamics) information, more complex information and internal relationships about the protein can be utilized, though this often requires increasing computational needs as the level of complexity used increases.
Sequences (1D)
The amino acid sequence of the protein reflects the primary structure of and usually uniquely identifies a protein; this usually consists of a simple text sequence of amino acid residues in the order that they are translated (such as “MVHLT...”).
The sequence of a protein is the most simple form of information that can usually be accessed easily by researchers for the purpose of prediction. Furthermore, the “alphabet” of proteins is effectively (barring nonstandard amino acids) only 20 amino acids, which is relatively small and can be much easier to work with computationally than other languages (such as English), though more complex approaches do sometimes define tokens in different ways such as defining motifs of specific amino acids in sequence or including tokens to represent phosphorylated residues or different peptide subunits. This approach to modeling protein sequences as an alphabet or language is often referred to as “protein language modeling” [3, 11].
Because of the ease of accessing sequence information (or generating it for a protein of interest) and the relative simplicity in working with it, the sequence is one of the most commonly-used pieces of information in generating various protein representations, and it is also commonly used in generating pretrained protein embeddings for downstream use. This includes methods like ESM-2, TAPE, and ProtBERT, among many others [12–14].
Contacts (2D)
The next level of complexity that is used by some researchers is a contact map, which represents the pairwise distances or contacts of individual amino acids or residues; this is essentially a two-dimensional projection of the protein’s structure and is usually represented by a symmetric square matrix A where elements represent either the distance or binary contact value (often computed from a threshold based on distance) between residues i and j.
The usage of contact map information is less common than the usage of sequence information; this may be due to it being harder to generate such contact maps than it is to simply acquire protein sequences, though many methods such as Pconsc4 do exist to create contact maps from protein sequences [15] that researchers have used to augment their protein representations. Notably, it is mathematically trivial to compute the contact map from a 3D protein structure, though this is not always done for a variety of reasons, discussed below.
Structures (3D)
The most complex information usually feasibly available for a protein is its actual 3D structure; this usually constitutes resolved coordinates of individual residues or atoms within the protein but can also include additional information such as known bonds between elements. Information that can usually be obtained from 3D structures can include features such as bond angles, the orientation of side chains, and the orientations of bonds or contacts.
This is less-used than the 1D or 2D sources of protein information that we discuss in this paper, primarily due to the lack of availability of protein structures for many proteins and due to the nontrivial nature of generating such structures either experimentally or computationally. However, recent advances in protein folding has made it much easier to access protein structures; through methods like AlphaFold2, predicted protein structures can be generated for almost any protein in the human proteome, though the quality of these predictions can vary substantially [9, 16].
It is important to note that, at this scale, 3D structural data are subject to certain geometric considerations. For example, a protein that is translated or rotated in 3D space is fundamentally the same protein, but certain features that may be derived from 3D coordinates may not be identical. For example, the direction of an edge between two residues may depend on rotation or translation. To address this, one can make use of models that are designed to be equivariant in 3D space to process such features (known as SE(3)-equivariance) or explicitly design features such that they do not rely on translation or rotation (are invariant to these transformations), such as using bond angles or the magnitude of vectors (distances) instead.
An example of a method that has made use of structural data in an equivariant way is PocketMiner [17], which makes use of protein 3D structures using GVP-GNNs, an equivariant graph neural network [18]. An example of a method that makes use of structural data by generating invariant features is 3DProtDTA [19], which instead makes use of certain dihedral bond angles. 3D structural protein data can also be processed by 3D convolutional neural networks that perform convolutions in 3D space, as in work done by Torng and Altman [20].
Dynamics (4D)
Finally, the highest form of complexity for a protein includes not just its 3D structure, but also how that 3D structure changes with time (molecular dynamics) or in different situations (conformation dynamics); this is usually presented as multiple sets of 3D coordinates for different conformations or different time points. The same considerations that apply to 3D structural data in terms of rotation and translation invariance usually apply here, as individual time points or conformations are essentially 3D projections of the 4D data.
This is the least-used of all types of protein information in most analyses, largely in part due to the high difficulty or infeasibility of generating such dynamics. Doing so often requires either experimentally resolving a protein multiple times or performing computationally-expensive molecular dynamics simulations, both of which can be challenging to do so practically and reasonably. Similarly to structures, though, recent advances have made it possible to generate some level of dynamics for some proteins; for example, AlphaFold2 has been shown to potentially, over the course of its stochastic generation process, produce some predictions that can be mapped to the conformational landscape of some proteins [21, 22], and other methods exist for predicting conformations such as Cfold [23].
Multimodal approaches
Some researchers also make use of multiple of these sources of possible information in a multimodal approach to generate representations. This usually involves generating separate or joint representations from sequence data in addition to contact or structure data, though there are a variety of other combinations that have been used. For example, Elia Venanzi et al. used sequence, structure, and dynamics data together to predict enzyme activity [24].
Other representations of proteins
Although not often employed in DTA, it is worth noting that a handful of other representations of protein structure and function have been used by the broader field. These sources of structural information—generally grouped together under the broader term of “protein signatures”—revolve around summarizing the overall structure of a protein as a more condensed sequence of motifs and other higher-order patterns [25]. The European Molecular Biology Laboratory (EMBL) groups signatures into four categories: Patterns, profiles, fingerprints, and hidden Markov models. While the details of these signature types are beyond the scope of this review, we direct the reader to EMBL’s tutorials [26] on the subject as well as the InterPro database [27], which is one of the most popular databases of protein signatures. Although useful for cataloging and classifying proteins, signatures are rarely used in DTA or similar representation learning approaches because they tend to be too low-resolution to capture nuanced details of binding pockets or other active sites.
Paradigms for generating and using representations
Here, we discuss various approaches for generating and using such representations. This includes paradigms such as supervised and unsupervised learning, as well as mixed paradigms such as transfer learning, fine-tuning, and semi-supervised learning.
Supervised (task-specific)
In the paradigm of supervised learning, protein representations are generated for a specific downstream task, such as predicting structure or function. In this case, each protein has a label of some kind (which can be a category, continuous value, or otherwise) that the representations are tuned to best map to. Fundamentally, the representations are learned to maximize the amount of information they contain that is relevant to predicting the downstream objective, and irrelevant information is often discarded or not captured effectively by such representations. In this paradigm, models are typically trained in an end-to-end manner where all of the parameters of a model, including the encoder and prediction head, are optimized at the same time. Occasionally, multiple related tasks may be trained at the same time to allow for representations to encode information that may be shared across tasks or to create more enriched representations. A visual overview of supervised learning can be found in Figure 3.
Fig. 3.
An overview of supervised learning as it may be applied to DTA, including the incorporation of possible drug information into the prediction
For example, supervised representations have been generated for predicting protein sequence from structure as in ProteinMPNN [28], predicting protein function from sequences [29], or predicting contact maps from sequence [30], among many others.
Unsupervised and self-supervised (generalized)
In the paradigm of unsupervised learning, protein representations are instead generated such that they provide an underlying latent map of the input proteins without a fixed label being present for each protein. The breadth and focus of the information contained within these representations depends in part on the method used, but such representations can often be used to cluster proteins in some way or even used as inputs to other tasks (further discussed in subsubsection "Transfer Learning and Fine-Tuning (Mixed)"). Self-supervised learning is a closely-related paradigm that is sometimes considered unsupervised learning wherein no prior labels exist and so instead labels are generated from the data itself (such as masking portions of the input data and training the model to predict those masked portions). In this paradigm, models are typically trained in an end-to-end manner where both an encoder and decoder are trained at the same time; subsequently, the trained encoder can be used to generate representations. A visual overview of both types of learning can be found in Figure 4.
Fig. 4.
An overview of unsupervised and self-supervised learning as it may be applied to generate protein representations for DTA. The top half of the figure represents an unsupervised paradigm where the encoder is trained to generate representations such that some decoder can reconstruct the original data. The bottom half represents a self-supervised paradigm where the encoder is trained to generate representations such that masked components (represented here by red boxes) of the input can be filled in by a decoder
For example, some such as Mansoor et al. have used autoencoders, which often conditions the representations to be good lower-dimensional representations that allow for maximal reconstruction of the input data [31]. Some generate representations by training a model to iteratively predict the next amino acid in a sequence such as UniRep [32] or by predicting masked parts of a sequence such as ESM Cambrian [33]. Others perform self-supervision through contrastive learning techniques, where representations are learned such that similar residues or proteins have similar representations while distinct residues or proteins are pushed apart in the embedding space, such as in S-PLM [34]. For example, Yu et al. used contrastive learning for enzyme function prediction to overcome a relative lack of functional annotations and saw improved performance [35]. Chen et al. were able to use self-supervised learning to improve protein representations through integration of structural information with sequence information by conditioning their model on residue pair distances and dihedral angles extracted from the structure [36].
Transfer learning and fine-tuning (mixed)
Transfer learning is a paradigm that usually blurs the line between other approaches, in that it often involved using a pretrained model to generate representations that are trained on a different task, then using those representations for a different task. While it is not necessarily a method of generating new representations, there are approaches to transfer learning that do create derivative representations that may be specialized for a different task. Transfer learning often allows researchers to take advantage of much larger or much more robust datasets than what may be available for their own task in generating good representations that contain useful information that they can then “transfer” to their own task. The original model may be trained using any of the paradigms mentioned here but is most commonly trained using a supervised, unsupervised, or self-supervised approach.
For example, ThermoMPNN makes use of features extracted from a pretrained version of ProteinMPNN to predict changes in stability of a protein after induced mutation [37], where ProteinMPNN was originally trained using a supervised inverse-folding task (using structure to predict sequence). In some cases such as in the work by Uzma et al., researchers train their own unsupervised model to generate embeddings based on a much larger dataset before using those embeddings in a downstream task [38].
Fine-tuning is a related paradigm to transfer learning and is considered a subtype of transfer learning by some, which is why we also expand on it here. Fine-tuning is often done by “unfreezing” all or some parameters in the pretrained model to allow those parameters to be reoptimized or optimized further using a different objective. This offers extended benefits in that it allows researchers to benefit from the training done on a different task as a sort of “pretraining” before further optimizing the representations for their own task.
For example, Schmirler et al. fine-tuned a variety of protein language models and found that they were able to improve performance on several downstream tasks compared to simple transfer learning alone [39], Dickson and Mofrad found that they were able to fine-tune language models to predict functional similarity based on Gene Ontology annotations [40], and Zeng et al. applied parameter-efficient fine-tuning to predict signal peptides for various proteins [41].
A visual overview of both types of learning can be found in Figure 5.
Fig. 5.
An overview of transfer learning and fine-tuning as it may be applied for DTA. The top half of the figure represents the pretraining phase, where an encoder is trained to generate representations for some initial task. Subsequently, the encoder is either frozen (as in transfer learning, represented by blue) or unfrozen (as in fine-tuning, represented by green) and used to generate embeddings that are then used for a different task. The green and blue boxes represent portions of the model that are trained in each paradigm, while the red box represents the pretraining phase where the encoder (along with its corresponding decoder) is trained on its original task
Semi-supervised (mixed)
In semi-supervised learning, one uses a combination of both labeled and unlabeled data to train a model. The most common example of semi-supervised learning is a case where the labeled data is used to generate an initial set of parameters for a model, and then the model is used to generate “pseudo-labels” for the unlabeled data, often picking high-confidence predictions for the purposes of generating pseudo-labels. The newly pseudo-labeled data are then used to continue training the model, and the pseudo-labels are sometimes then refined iteratively as the model is continually trained. A visual overview of self-supervised learning can be found in Figure 6.
Fig. 6.
An overview of self-supervised learning. In this diagram, a model is trained on a set of supervised labels and then used to predict pseudolabels for an unlabeled set of data, after which high-confidence labels are used to continue training the model on the remaining unlabeled data
For example, Moffat and Jones use semi-supervised learning to train a model to generate representations and labels for a large dataset of proteins based on their multiple sequence alignments (MSAs); they then use these labels generated using semi-supervised learning on a first model for a secondary training approach of a different model [42]. Dhanuka et al. do something similar by training autoencoders conditioned to specific protein functions and ontologies in an unsupervised manner, then use the reconstruction losses produces by those autoencoders to predict the actual function(s) of an arbitrary sequence [43].
Protein representations in drug-target affinity
In general, drug-target affinity is a regression task where labels exist for the datasets in question; the goal is to predict the binding affinity between a protein and a molecule using various possible information sets from both of them. For the purposes of this review, we focus on methods that use protein information independently of molecule information; that is, we disregard methods that operate on mixed protein-ligand information (such as the ligand-bound protein complex structure) mostly because these generally do not generate disentangled protein representations (instead generating representations of the mixed protein-ligand information) and rely on the existence of the protein-ligand complex, which is harder to generate (at present) than unbound (apo) protein structures. Most recent approaches for predicting DTA have used deep learning in some capacity, and so we narrow our focus to these as they have tended to produce much better performance than non-deep-learning methods.
Modalities used in creating protein representations in DTA
A wide variety of modalities have been used in DTA to produce protein representations. While our survey of the literature is not meant to be wholly comprehensive of the landscape of DTA methods, we do include at least one method for each protein modality that appears to have been used in some capacity to predict DTA that we were able to find. These include sequences, contact maps, and structures, encompassing 1D, 2D, and 3D data. We find that some methods make use of multiple modalities within these three categories as well. Interestingly, we note that no methods that we found to date make use of what we previously termed as “4D” data in the form of dynamics (either molecular or conformational dynamics). We discuss methods that make use of each of these modalities below when discussing encoders, as different encoders tend to be specialized in extracting features for certain modalities due to inherent inductive biases, and so discussing an encoder tends to also mean discussing the modality on which it most commonly operates (in the context of protein representation learning).
Encoders for generating protein representations in DTA
In drug-target affinity, protein representations have been generated using a variety of different encoders. We summarize what some existing approaches for predicting DTA use as encoders in Table 1; while this is not intended to be a comprehensive list, we have endeavored to identify methods that represent a wide breadth of different methods to generate protein representations. We also disregard methods that do not explicitly learn or fine-tune the protein representations for drug-target affinity (such as Affinity2Vec, which uses transfer learning on pretrained embeddings from ProtVec [44]) in favor of focusing on learnable representations specifically.
Table 1.
Table of Previous DTA Methods with Information on Protein Representations. “(PT)” before an encoder type in this table represents that the encoder is Pre-Trained
| Method | Training paradigm | Protein modalities | Encoder type | Comments |
|---|---|---|---|---|
| DeepDTA [45] | Supervised | Sequence | CNN | |
| GraphDTA [46] | Supervised | Sequence | CNN | |
| DGraphDTA [47] | Supervised | Contact | GNN | |
| DeepGLSTM [48] | Supervised | Sequence | BiLSTM | |
| AttentionDTA [49] | Supervised | Sequence | CNN | |
| MSGNN-DTA [50] | Supervised | Contact | GNN | |
| DeepCDA [51] | Supervised | Sequence | CNN + BiLSTM | |
| FusionDTA [52] | Transfer | Sequence | (PT) Transformer + BiLSTM | Used pretrained ESM-1B embeddings and processed further |
| TC-DTA [53] | Supervised | Sequence | Transformer | |
| HGTDP-DTA [54] | Supervised | Contact | GNN | |
| GLCN-DTA [55] | Supervised | Contact | GNN | |
| 3DProtDTA [19] | Supervised | Structure | GNN | |
| PocketDTA [56] | Transfer | Sequence | (PT) Transformer | Used pretrained ESM-2 embeddings |
| Supervised | Structure | GNN | Extracted and used only structures for binding pockets | |
| MDF-DTA [57] | Transfer | Sequence | (PT) CNN + (PT) Transformer(s) | Used pretrained ProtVec + ProtBERT + ESM-Fold embeddings |
| AttentionMGT-DTA [58] | Transfer | Sequence | (PT) Transformer | Used pretrained ESM-2 embeddings |
| Supervised | Structure | Graph Transformer | Used structures for mathematically-determined binding pockets only | |
| DTA-GTOmega [59] | Transfer | Sequence | (PT) Transformer | Used pretrained OmegaFold embeddings |
| Supervised | Contact | Graph Transformer | Computed contact map from predicted OmegaFold 3D structure | |
| SSM-DTA [60] | Semi-Supervised | Sequence | Transformer | Used masked language modeling for semi-supervision |
We describe these generation methods as “encoders” to distinguish them from the overarching model that is used to predict DTA, which is often more complex than the protein encoder alone due to also needing to consider drug information. We found that protein encoders such as convolutional neural networks, recurrent neural networks, graph neural networks, transformers, and graph transformers have all been used to varying degrees to generate protein representations, with varying input information from proteins being used. By their nature and design, different encoders tend to operate on specific modalities; we discuss each of these encoders and the modalities that they operate on in the sections below. We provide a simple summary of which encoders operate on what modalities as well as which DTA methods use each encoder in Table 2 as well as providing figures with visual summaries in each encoder’s subsection.
Table 2.
Encoders, modalities, and methods that use the given encoders. For the sake of brevity, we do not provide citations for every DTA method in this table; citations for each method can be found in Table 1. Abbreviations: CNN = convolutional neural network; RNN = recurrent neural network; BiLSTM = bidirectional long short-term memory; GNN = graph neural network
| Encoder | Modalities | DTA Methods |
|---|---|---|
| CNN | Sequence | DeepDTA, GraphDTA, AttentionDTA, DeepCDA, MDF-DTA |
| RNN (BiLSTM) | Sequence | DeepGLSTM, DeepCDA, FusionDTA |
| GNN | Contact, Structure | DGraphDTA, MSGNN-DTA, HGTDP-DTA, GLCN-DTA, 3DProtDTA, PocketDTA |
| Transformer | Sequence | FusionDTA, TC-DTA, PocketDTA, MDF-DTA, AttentionMGT-DTA, DTA-GTOmega, SSM-DTA |
| Graph Transformer | Contact, Structure | AttentionMGT-DTA, DTA-GTOmega |
Convolutional neural networks
Convolutional neural networks (CNNs) are a deep learning architecture that applies convolutional filters over an input space to produce embeddings [61]. CNNs have most commonly been applied in computer vision to analyze 2D (or even 3D) image data; however, in the context of protein representation learning for drug-target affinity, they are most commonly used in their 1D form to analyze amino acid sequences, and all of the methods discussed here apply CNNs solely to sequence data. Specifically, 1-dimensional convolutional filters are applied to the sequence (oftentimes an embedding or one-hot encoding of the sequence) to generate embeddings that capture local information at each element of a sequence. A visual representation of how CNNs process sequence data can be seen in Figure 7. We note that convolutional neural networks can be used to process 2D and 3D data as well and have been used for this purpose in other protein tasks, though we have not observed their usage in this regard to process contact maps or protein structures for the purposes of predicting DTA.
Fig. 7.
An overview of the convolutional neural network architecture as it might be applied to protein sequences. Convolutions create new representations for each amino acid based on a convolution of itself and adjacent amino acids
Several of the methods discussed use CNNs, all to process sequences. DeepDTA [45], GraphDTA [46], and AttentionDTA [49] all make use of 1D-CNNs to process protein sequences to generate embeddings that are then used downstream. DeepCDA [51] uses a 1D-CNN to generate initial embeddings that are then further processed by a BiLSTM later. MDF-DTA [57] uses a pretrained 1D-CNN model called ProtVec [44] as one of several pretrained models used to generate protein embeddings.
Recurrent neural networks (LSTM)
Recurrent neural networks (RNNs) are a deep learning architecture designed to operate on sequences, doing so by making use of a recurrent unit that update an internal hidden state as elements of a sequence are passed through it [62]. By continually updating the hidden state, an RNN can produce a representation of a sequence. Long short-term memory (LSTM) architectures are an extension of RNNs that seeks to improve the management of information over long “distances” within a sequence by enabling the architecture to choose whether to continually incorporate or forget certain information. Bidirectional LSTMs (BiLSTM) are LSTMs that operate on sequences in both directions (forward and backward). In protein representation learning for drug-target affinity, these are almost universally used to process amino acid sequences, and all RNNs that we found are universally BiLSTMs. A visual representation of how RNNs (particularly BiLSTMs) operate on sequence data can be seen in Figure 8.
Fig. 8.
An overview of the recurrent neural network architecture as it might be applied to protein sequences. For BiLSTMs, the protein sequence is processed both forward and backwards with each amino acid taking input from the previous amino acids (in forward or reverse order). The LSTM cell can contain different components based on the underlying architecture.
Of the methods discussed, several use RNNs in the form of BiLSTMs, all operating on sequences (initially or downstream). DeepGLSTM [48] uses a BiLSTM by itself to process amino acid sequences. DeepCDA [51] uses a BiLSTM to further process embeddings after initial generation by a CNN. FusionDTA [52] uses a BiLSTM to generate more elaborate embeddings from embeddings initially generated using a pretrained model.
Transformers (Sequence-Based)
Transformers are a deep learning architecture that processes tokenizable sequences by computing embedding vectors for individual tokens and then applying a multi-head attention mechanism to those tokens to update their representations. Over the training process, “attention” to important tokens is maximized to improve representations [63, 64]. In many cases, these transformers comprise “protein language models”, and in the context of protein representation learning for DTA, they almost universally operate on amino acid sequence data wherein the tokens are defined as different amino acid residues. A visual representation of how transformers compute embeddings from sequence data can be seen in Figure 9.
Fig. 9.
An overview of a transformer architecture as it might be applied to protein sequences. The tokenized protein sequence is projected into query (Q), key (K), and value (V) representations. An attention matrix is computed from the query and key (usually using scaled dot-product attention), then multiplied by the value projection to produce the downstream embedding where each amino acid has a representation based on its attention with all other amino acids.
Of the methods described, most use pretrained transformer models rather than training their own transformer, and in all cases the transformer is used to generate embeddings for sequences. Both PocketDTA [56] and AttentionMDT-DTA [58] use pretrained embeddings generated by ESM-2 [12]. DTA-GTOmega [59] uses pretrained embeddings produced by OmegaFold [65]. MDF-DTA [57] uses embeddings derived from pretrained transformer models in ProtBERT [14] and ESM-Fold [12]. FusionDTA [52] initially generates embeddings from ESM-1B [66], a pretrained transformer model. In two cases, a model does train a transformer itself to process protein sequences: TC-DTA [53] and SSM-DTA [60]. The reason why most models use pretrained transformers may be because of the general complexity of training transformers for protein sequences which can require large quantities of data for effective generalizability or high amounts of computational resources.
Graph neural networks
Graph neural networks (GNNs) are a deep learning architecture that operates on graphs containing nodes and edges, creating representations at each node typically through iterative message-passing, where a node is updated based on the embeddings of its neighbors (and sometimes embeddings of the edges connecting it to those neighbors); notably, there are various subclasses to GNNs where the update or message-passing framework varies [67, 68]. In protein representation learning for DTA, these are usually used to process contact data or 3D structural data where these modalities are reframed in the context of a graph. This is usually done by creating a graph where nodes are amino acid residues or atoms and edges are either based on bonds between residues/atoms or based on distances between nodes in 3D space. A visual representation of how graphs may be processed for GNNs and how GNNs construct new embeddings from contact maps or 3D structures can be seen in Figure 10.
Fig. 10.
A graph neural network architecture as it might be applied to either protein contact maps and 3D structure. The graph is typically constructed with nodes representing individual amino acids and edges being constructed based on distance; in some cases, this represents the whole protein while in others a subset of the protein may be modeled. For some GNNs, features may be put on each node or edge representing invariant features (such as bond angles or absolute distances) while others may make use of equivariant features. Graph convolutions are performed, usually leading to embeddings being updated based on neighboring nodes; multiple convolutions lead to each node aggregating information from nodes over a greater distance.
Many of the methods described use GNNs to process either contact maps or structures. DGraphDTA [47], MSGNN-DTA [50], HGTDP-DTA [54], and GLCN-DTA [55] all make use of GNNs to process contact maps. 3DProtDTA [19] and PocketDTA [56] both use GNNs to process structures; the former incorporates features that are invariant to rotation and translation while the latter makes use of equivariant graph neural networks on binding pocket graphs specifically (PocketDTA also separately processes sequences using a pretrained model).
Graph transformers
Graph transformers are a deep learning architecture that combine principles of GNNs and transformers in a variety of different possible ways, including but not limited to using edge information to condition attention calculations, incorporating attention into message-passing, or by incorporating graph spectral or positional embeddings into the computation of attention [69, 70]. These are generally used solely on contact or structure data as they require a graph structure of some sort, similar to GNNs, and the way that the graphs are constructed is often very similar to or identical to GNNs. A visual representation of this can be seen in Figure 11.
Fig. 11.
A graph transformer architecture as it might be applied to either protein contact maps and 3D structure. The graph is usually constructed in much the same way as it might be for a GNN (Figure 10), though the way the graph is processed is different, with each node usually attending to other nodes in a manner much like that of a transformer (Figure 9) with additional regularization based on the graph structure.
DTA-GTOmega [59] uses graph transformers to process contacts alongside a pretrained transformer to process sequences. AttentionMGT-DTA [58] uses a graph transformer to process structures alongside a pretrained transformer to process sequences; in doing so, it constructs rotation- and translation-invariant features for downstream use, such as the dihedral angles about the alpha carbons.
Combined methods
Some approaches combine multiple methods from above, and we describe them above in detail. Generally, multiple encoder methods are used to process multiple modalities, though at times researchers use the sequence-based methods sequentially (such as CNNs and transformers or RNNs and transformers) to further process embeddings generated from sequences. A simplified visual example of how methods may make use of multiple encoders in this way can be seen in Figure 12.
Fig. 12.
An example of a method that might make use of both sequence (1D) and contact map (2D) inputs by making use of a transformer to process sequence information and a GNN to process contact map information to generate separate embeddings. These separate embeddings are then fused in some way (most often through concatenation, though other methods exist) to produce a single embedding with information from both encoders for downstream use.
These methods (all also discussed above in their relevant sections) include PocketDTA [56], AttentionMGT-DTA [58], and DTA-GTOmega [59], which all make use of different encoders to process different input modalities (such as sequence and structure or sequence and contact). They also include DeepCDA [51], FusionDTA [52], and MDF-DTA [57], which all make use of one or more sequence-based encoders to process sequences multiple times.
Observations in drug-target affinity
Predicting drug-target affinity is a biomedical question that has become increasingly popular in the last several years, especially with the rise of increased access to protein-level information and improved methods for analyzing such data, including improvements in deep learning. We see that DTA prediction models are highly diverse, making use of a wide range of modalities, multiple different encoder types, and even show some variability in the training paradigms used to train these models. However, DTA lags behind in many of these regards when compared to the broader scope of protein representation learning.
Training paradigms for protein representations in DTA
The vast majority of the DTA models make use of supervised or transfer learning approaches; this is largely to be expected as predicting drug-target affinity is, by definition, a paradigm where each input (a protein and molecule pair) has a label (affinity) to predict, which lends itself to to end-to-end training in a supervised fashion. However, most drug-target affinity datasets are relatively small, and the use of pretrained models to generate initial embeddings can allow for embeddings trained on a more diverse array of proteins to be used, which contributes to transfer learning also being common.
Interestingly, we find only rare examples of semi-supervised learning (for example, SSM-DTA [60]) and were not able to find any examples of fine-tuning as applied to drug-target affinity, despite the fact that both of these paradigms may be similarly viable as transfer learning in that drug-target affinity is a labeled paradigm where the datasets tend to be small, and both of these paradigms can make use of unlabeled data or pretrained models to “augment” the dataset.
Possible challenges for training paradigms in DTA
One anticipated challenge in implementing semi-supervision may be in developing methods of assessing which labels can be considered high-confidence and determining the best set of predicted labels to propagate for further training. In DTA, the primary prediction is typically a continuous label of binding affinity, and determining the confidence of binding affinity labels can be challenging or poorly-defined in this paradigm as the vast majority of possible proteins and drugs likely do not bind to each other at all and so do not necessarily contribute new information to the model. Furthermore, semi-supervised learning does still rely on a large set of data but benefits from the idea that such data may be unlabeled rather than labeled; however, for many modalities in protein representations (particularly 2D contacts, 3D structure, and 4D dynamics data), even large quantities of unlabeled data have not previously been readily available.
With regards to fine-tuning, this is a common paradigm in other protein-related tasks, but its underutilization in DTA may be due to the computational complexity involved in performing fine-tuning with highly-enriched protein representations; for example, the smallest ESM-2 model is 6 million parameters while the largest is 15 billion parameters [12], and fine-tuning such models in a canonical manner (by unfreezing all of their parameters) may lead to unstable training dynamics or outstrip available computing resources during both training and evaluation.
Protein modality usage in DTA
Broadly, the most common modality/information used by DTA models is the 1D amino acid sequence, with the vact majority of the models found making use of the sequence in one way or another. The next most common modality used is the 2D contact map, with several modalities that make use of GNNs to process the protein graph making use of this data.
Relatively few methods made use of 3D structural data, though this seems to be becoming more popular; in the cases where they did, the methods made use of computed invariant features (such as bond angles) or focused specifically on binding pockets (usually determined by another method) rather than making full use of the entire protein structure.
Additionally, we were unable to find any methods that made use of 4D protein data in the form of multiple conformations or molecular dynamics data for the purposes of predicting DTA (where the protein is represented in an unbound state with multiple conformations or dynamics), though we do note that some recent work such as that done by Libouban et al. has made use of molecular dynamics data of protein-ligand interactions for predicting DTA with a combination of a CNN and RNN where the CNN operates over the 3D space and the RNN fuses information across different conformations or timepoints [71].
Possible challenges for protein modalities in DTA
There are two challenges that likely exist for utilizing different modalities in DTA. The first is the availability of such data - protein sequence data (1D) is the most common and most abundant modality, and many preexisting DTA datasets provide only protein sequences by default. In contrast, 2D contact map and 3D structure data is much less common both within and outside of DTA, and 4D protein dynamics or conformation data is more or less nonexistent.
The second challenge likely arises from the considerations that must be made to ensure that your encoder reflects some degree of known biology or is at least capable of taking advantage of the modality of choice. For sequence data, there are few considerations that need to be made beyond treating the amino acid sequence as an alphabet and determining the best method for tokenizing it into machine-readable form; while additional complexities can be included (such as incorporating evolutionary information through multiple sequence alignments or incorporating known amino acid properties), these are not required to be able to develop a DTA method that generates its representations from 1D protein sequence information.
In contrast, 2D contact map and 3D structural data requires increasingly complex and unique considerations. For 2D contact map data, one often needs to consider how to encode geometric information (if at all) and ensuring that contact maps provide enough information about the protein (as defining edges based on a cutoff may lead to slightly misspecified contact maps). For 3D structural data, one consideration is that it may be important to ensure that the utilized representation of the protein in 3D space produces the same result even if the input protein is rotated or translated in 3D space as these transformations do not impact the actual structure of the protein. 4D data considerations are still emerging as the paradigm becomes more popular, but this includes trying to ensure that different conformations of the same protein are treated appropriately and that information is not oversmoothed over different conformations or timepoints to the extent that important within-timepoint information is lost.
Such complexities and constraints can be hard to properly model and become increasingly difficult to manage as the amount of information increases (from 1D to 2D to 3D to 4D).
It is also important to note that as the data modality becomes more complex, so too does the method for analyzing it, and the choice of modality and encoder are often inherently intertwined. We discuss encoder-related considerations further in the encoder section, but these also impact the choice of protein modality as certain methods are often well-specified for certain modalities.
Protein encoder usage in DTA
As most methods make use of sequences, they expectedly make preferential use of the modalities that process sequences. Of these, transformers are the most common, with CNNs being the next most common and BiLSTMs being the least common (but still used in several cases). Many of these sequence-encoders, especially transformers, are pretrained models that are used in a transfer learning paradigm.
The next most-common encoder are graph neural networks, used by several models that make use of contact or structure data. Similarly, graph transformers are also used for both contact and structure data, though they are much less common than GNNs. In most cases, the GNNs or graph transformers operate on graphs where individual nodes are residues or amino acids and edges connect residues based on distance.
However, as discussed above, all of the graph-based methods that make use of structural data either use explicitly-invariant features (3DProtDTA) or subset the protein structure substantially by focusing on only binding pockets (PocketDTA) as discussed in subsubsection "Structures (3D)" , often alongside other modalities. This inspired our own work in CASTER-DTA, which uses an encoder that makes use of the full 3D protein data in an equivariant manner for predicting drug-target affinity [72]; we believe that DTA has great potential in being able to use such encoders and that this represents a next logical step for the field to advance into.
Furthermore, there exist many additional encoders not used in DTA that may be valuable in protein representations more broadly that may improve generalizability or allow for representations to more closely mirror biology. For example, Implicit Neural Networks (INNs; also known as Implicit Neural Representations) have been developed for protein representations and show great promise in being able to model not only 3D protein structures but also 4D protein dynamics as well. For example, Sun et al. [73] and Wu et al. [74] use INNs to model protein surfaces (and the dynamics thereof over multiple conformations through time) in continuous representation spaces that can then be used for a variety of downstream tasks.
Possible challenges for encoder usage in DTA
As we discussed in the corresponding section for modality selection and usage in DTA, many considerations thereof rely on encoder selection as well as certain encoders are well-suited for analyzing particular modalities. For example, while a GNN could theoretically be used to analyze sequence data by essentially making a line graph, such data is likely better analyzed by a CNN, RNN, or transformer. As such, the choice of encoder and modality are highly related.
This also means that the choice of encoder is subject to the same considerations as for modalities; however, additional considerations arise for encoders based on their ease of implementation. However, we also note that the perceived computational complexity of these models likely represents a barrier to entry for use of them as well.
From the ease of implementation standpoint, consider that sequence data can easily be analyzed by a wide variety of well-developed and well-established methods such as RNNs, CNNs, and transformers. All of these methods for analyzing sequence data can be easily implemented using a wide variety of deep learning libraries in multiple languages, and there is no shortage of existing models to iterate off of in this regard to make improvements. Similarly, contact map data has been consistently analyzed using CNNs or GNNs in several papers for DTA and otherwise, and these methods are relatively easy to implement using existing packages.
However, there are far fewer proven methods for encoding 3D and 4D data, and such methods are often more complex or bespoke, leading to them being more difficult to implement. For example, equivariant GNNs for analyzing 3D structural data, while they are becoming more popular, are by no means common, and existing frameworks for implementing them often require much more domain knowledge and debugging. This issue is further exacerbated when considering more complex 4D dynamics data which often requires a model that can understand the relationship between protein spatial relationships and temporal or conformational dynamics.
Intuitively, it might also seem that more complex models from a theoretical standpoint might also require increased computational complexity (often represented by the number of parameters that need to be trained for the model); however, from our experience, we find that this is not necessarily actually the case. More complex encoders for different modalities are capable of capturing more information, but the modality complexity to encoder complexity ratio to acquire the same information is not positively correlated; in fact, we find that it may even be negatively correlated. For example, some of the state-of-the-art encoders for sequence data (such as the ESM models) require millions or even billions of parameters to fully capture protein information from sequence data alone while state-of-the-art structure encoders can typically operate with fewer parameters for similar performance.
Essentially, the further a modality gets from the truth of a protein (essentially, along the canonical sequence-structure-function axis), the “harder” an encoder has to work (the number of parameters it needs to have) to learn the full array of relationships that lead to the underlying biology that allows the protein to have the function or activity that it does. Operating on more complex modalities that present more information about proteins from the outset also tends to improve generalization performance as less has to be inferred (during the learning process) by the model.
We encourage future researchers interested in entering the field not to be intimidated by the complexity of these models, both in implementation and in apparent computational requirements, as we find that more complex models capable of taking advantage of the modalities that they operate on effectively are able to perform just as well if not better even when they appear to be computationally small or simple from a parameterization standpoint (for example, equivariant GNNs for 3D structural data do not need to learn positional invariance as the equivariance is part of the model design), and this often reflects in their generalization capacity as well.
Conclusions and future directions
We discuss above some of the limitations that we have seen in our foray into the computational DTA prediction literature, speculate on why such limitations may exist, and elaborate on ways that these gaps may inform future directions in the field for further improving our approach to drug-target affinity prediction, with a particular focus on protein representations.
With regards to training paradigms, we note a relative lack of methods that make use of semi-supervised learning and fine-tuning of pretrained models in the literature. This may be, for semi-supervised learning, due to the increased complexity of setting up such training paradigms (requiring the acquisition of large amounts of unlabeled data) and, for fine-tuning, the increased computational resources needed to fine-tune certain pretrained models, which can be quite large in size. However, as improved finetuning methods such as parameter-efficient fine-tuning (PEFT) [75] continue to be developed and the availability of protein data continues to improve, both of these paradigms become increasingly more viable and likely to be implemented in DTA prediction.
We note from a modality-usage perspective that no methods to date in the literature appear to make use of multiple conformations or molecular dynamics data, what we describe as “4D” data. This is likely due to the generally-poor access to such data for a wide range of proteins; however, as we mention, newer methods may be able to automatically produce/predict multiple conformations of proteins [21–23] that may be able to be used for more effective DTA prediction. The development of encoders that may be able to make use of 4D data is still an evolving field, but one possible example may be temporal graph networks or other deep learning methods designed to operate on dynamic graphs [76], and one may be able to take pointers from methods such as [71] to develop methods for protein representation.
When considering encoder usage in addition to protein modality incorporation, we observe that few methods make use of protein structure as the sole modality, make use of only a portion of the protein structure, or use only invariant features. Using a single modality in the form of structure may lead to improved training dynamics; furthermore, making use of the entire protein structure in an equivariant manner may lead to improved performance of these methods as using invariant features or a subset of the protein can lead to a loss of relevant information and poorer generalization performance, particularly for proteins not in the dataset used to train the model. We note that previous methods outside of DTA such as PocketMiner [17] have been able to make effective use of 3D protein structural information in this way, and we sought to show the utility of using equivariant protein representations in our own recent work in CASTER-DTA [72], though this is still a nascent and evolving field.
To conclude, there are a wide array of protein representations that have been used in drug-target affinity, but there exists a large amount of potential in the field to make use of richer representations and more powerful encoder methods. In this paper, we review what we have seen done in DTA in detail to provide an overview of the field as a whole, which enables us to make observations about possible next steps for continued improvement in predicting DTA based on advances that have been made in protein representation learning in other paradigms, with the hope that our insights will lead to an advancement in DTA prediction as well as in the field of protein representation learning more broadly.
Acknowledgements
RK was supported by the National Human Genome Research Institute of the National Institutes of Health (T32HG000046). JDR was supported by the National Library of Medicine of the National Institutes of Health (R00LM013646). MDR was supported by the National Institute on Aging of the National Institutes of Health (U01AG066833). The funders played no role in study design, data collection, analysis and interpretation of data, or the writing of this manuscript.
Abbreviations
- DTA
Drug-target affinity
- CNN
Convolutional neural network
- GNN
Graph neural network
- RNN
Recurrent neural network
- INN (INR)
Implicit neural network (implicit neural representation)
- BiLSTM
Bidirectional long short-term memory (network)
- PEFT
Parameter-efficient fine-tuning
Author contributions
RK conceived the idea for this review and performed the initial literature search, wrote the initial draft of the paper, and edited the final draft of the paper. JDR provided supervision, additional literature review, contributed to the initial draft of the paper, and edited the final draft of the paper. MDR provided supervision, funding support, contributed to the initial draft of the paper, and edited the final draft of the paper.
Funding
RK was supported by the National Human Genome Research Institute of the National Institutes of Health (T32HG000046). JDR was supported by the National Library of Medicine of the National Institutes of Health (R00LM013646). MDR was supported by the National Institute on Aging of the National Institutes of Health (U01AG066833).
Data availability
No datasets were generated or analysed during the current study.
Code availability
No code was used or written for this paper.
Declarations
Competing interests
The authors declare no Conflict of interest.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Alberts B, Johnson A, Lewis J, Raff M, Roberts K, Walter P (2002) Analyzing Protein Structure and Function. Molecular Biology of the Cell. 4th Edition
- 2.Sadowski MI, Jones DT (2009) The sequence-structure relationship and protein function prediction. Curr Opin Struct Biol 19(3):357–362. 10.1016/j.sbi.2009.03.008 [DOI] [PubMed] [Google Scholar]
- 3.Bepler T, Berger B (2021) Learning the protein language: Evolution, structure, and function. Cell Syst 12(6):654–6693. 10.1016/j.cels.2021.05.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Bengio Y, Courville A, Vincent P (2013) Representation learning: A review and new perspectives. IEEE Trans Pattern Anal Mach Intell 35(8):1798–1828. 10.1109/TPAMI.2013.50 [DOI] [PubMed] [Google Scholar]
- 5.Detlefsen NS, Hauberg S, Boomsma W (2022) Learning meaningful representations of protein sequences. Nat Commun 13(1):1914. 10.1038/s41467-022-29443-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.McGibbon M, Shave S, Dong J, Gao Y, Houston DR, Xie J, Yang Y, Schwaller P, Blay V (2024) From intuition to AI: Evolution of small molecule representations in drug discovery. Brief Bioinform 25(1):422. 10.1093/bib/bbad422 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ragoza M, Hochuli J, Idrobo E, Sunseri J, Koes DR (2017) Protein-Ligand Scoring with Convolutional Neural Networks. J Chem Inf Model 57(4):942–957. 10.1021/acs.jcim.6b00740 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Li Z, Huang R, Xia M, Patterson TA, Hong H (2024) Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery. Biomolecules 14(1):72. 10.3390/biom14010072 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O et al (2021) Highly accurate protein structure prediction with AlphaFold. Nature 596(7873):583–589. 10.1038/s41586-021-03819-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Crescenzi O, Tomaselli S, Guerrini R, Salvadori S, D’Ursi AM, Temussi PA, Picone D (2002) Solution structure of the Alzheimer amyloid -peptide (1–42) in an apolar microenvironment. Eur J Biochem 269(22):5642–5648. 10.1046/j.1432-1033.2002.03271.x [DOI] [PubMed] [Google Scholar]
- 11.Nijkamp E, Ruffolo JA, Weinstein EN, Naik N, Madani A (2023) ProGen2: Exploring the boundaries of protein language models. Cell Syst 14(11):968–9783. 10.1016/j.cels.2023.10.002 [DOI] [PubMed] [Google Scholar]
- 12.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y, dos Santos Costa A, Fazel-Zarandi M, Sercu T, Candido S, Rives A (2023) Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379(6637):1123–1130. 10.1126/science.ade2574 [DOI] [PubMed] [Google Scholar]
- 13.Rao R, Bhattacharya N, Thomas N, Duan Y, Chen X, Canny J, Abbeel P, Song YS (2019) Evaluating Protein Transfer Learning with TAPE. Adv Neural Inf Process Syst 32:9689–9701 [PMC free article] [PubMed] [Google Scholar]
- 14.Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Wang Y, Jones L, Gibbs T, Feher T, Angerer C, Steinegger M, Bhowmik D, Rost B (2022) ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Trans Pattern Anal Mach Intell 44(10):7112–7127. 10.1109/TPAMI.2021.3095381 [DOI] [PubMed] [Google Scholar]
- 15.Michel M, Menéndez Hurtado D, Elofsson A (2018) PconsC4: Fast, accurate and hassle-free contact predictions. Bioinformatics 35(15):2677–2679. 10.1093/bioinformatics/bty1036 [DOI] [PubMed] [Google Scholar]
- 16.Lyu J, Kapolka N, Gumpper R, Alon A, Wang L, Jain MK et al (2024) AlphaFold2 structures guide prospective ligand discovery. Science 384(6702):6354. 10.1126/science.adn6354 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Meller A, Ward M, Borowsky J, Kshirsagar M, Lotthammer JM, Oviedo F, Ferres JL, Bowman GR (2023) Predicting locations of cryptic pockets from single protein structures using the PocketMiner graph neural network. Nat Commun 14(1):1177. 10.1038/s41467-023-36699-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Jing B, Eismann S, Soni PN, Dror RO (2021) Equivariant Graph Neural Networks for 3D Macromolecular Structure. arXiv. 10.48550/arXiv.2106.03843
- 19.Voitsitskyi T, Stratiichuk R, Koleiev I, Popryho L, Ostrovsky Z, Henitsoi P et al (2023) 3DProtDTA: A deep learning model for drug-target affinity prediction based on residue-level protein graphs. RSC Adv 13(15):10261–10272. 10.1039/D3RA00281K [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Torng W, Altman RB (2019) High precision protein functional site detection using 3D convolutional neural networks. Bioinformatics 35(9):1503–1512. 10.1093/bioinformatics/bty813 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Winski A, Ludwiczak J, Orlowska M, Madaj R, Kaminski K, Dunin-Horkawicz S (2024) AlphaFold2 captures the conformational landscape of the HAMP signaling domain. Prot Sci A Publ Prot Soc 33(1):4846. 10.1002/pro.4846 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Al-Masri C, Trozzi F, Lin S-H, Tran O, Sahni N, Patek M, Cichonska A, Ravikumar B, Rahman R (2023) Investigating the conformational landscape of AlphaFold2-predicted protein kinase structures. Bioinform Adv 3(1):129. 10.1093/bioadv/vbad129 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Bryant P, Noé F (2024) Structure prediction of alternative protein conformations. Nat Commun 15:7328. 10.1038/s41467-024-51507-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Venanzi NAE, Basciu A, Vargiu AV, Kiparissides A, Dalby PA, Dikicioglu D (2024) Machine Learning Integrating Protein Structure, Sequence, and Dynamics to Predict the Enzyme Activity of Bovine Enterokinase Variants. J Chem Inf Model 64(7):2681–2694. 10.1021/acs.jcim.3c00999 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Martin S, Roe D, Faulon J-L (2005) Predicting protein-protein interactions using signature products. Bioinformatics 21(2):218–226. 10.1093/bioinformatics/bth483 [DOI] [PubMed] [Google Scholar]
- 26.EMBL-EBI: Signature Types | Protein Classification. https://www.ebi.ac.uk/training/online/courses/protein-classification-intro-ebi-resources/what-are-protein-signatures/ Accessed 10 Apr 2025
- 27.Blum M, Andreeva A, Florentino LC, Chuguransky SR, Grego T, Hobbs E, Pinto BL, Orr A, Paysan-Lafosse T, Ponamareva I, Salazar GA, Bordin N, Bork P, Bridge A, Colwell L, Gough J, Haft DH, Letunic I, Llinares-López F, Marchler-Bauer A, Meng-Papaxanthos L, Mi H, Natale DA, Orengo CA, Pandurangan AP, Piovesan D, Rivoire C, Sigrist CJA, Thanki N, Thibaud-Nissen F, Thomas PD, Tosatto SCE, Wu CH, Bateman A (2025) InterPro: The protein sequence classification resource in 2025. Nucleic Acids Res 53(D1):444–456. 10.1093/nar/gkae1082 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Dauparas J, Anishchenko I, Bennett N, Bai H, Ragotte RJ, Milles LF, Wicky BIM, Courbet A, de Haas RJ, Bethel N, Leung PJY, Huddy TF, Pellock S, Tischer D, Chan F, Koepnick B, Nguyen H, Kang A, Sankaran B, Bera AK, King NP, Baker D (2022) Robust deep learning-based protein sequence design using ProteinMPNN. Science 378(6615):49–56. 10.1126/science.add2187 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Gelman S, Fahlberg SA, Heinzelman P, Romero PA, Gitter A (2021) Neural networks to learn protein sequence-function relationships from deep mutational scanning data. Proc Natl Acad Sci 118(48):2104878118. 10.1073/pnas.2104878118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Singh J, Litfin T, Singh J, Paliwal K, Zhou Y (2022) SPOT-Contact-LM: Improving single-sequence-based prediction of protein contact map using a transformer language model. Bioinformatics 38(7):1888–1894. 10.1093/bioinformatics/btac053 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Mansoor S, Baek M, Park H, Lee GR, Baker D (2024) Protein Ensemble Generation Through Variational Autoencoder Latent Space Sampling. J Chem Theory Comput 20(7):2689–2695. 10.1021/acs.jctc.3c01057 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Alley EC, Khimulya G, Biswas S, AlQuraishi M, Church GM (2019) Unified rational protein engineering with sequence-based deep representation learning. Nat Methods 16(12):1315–1322. 10.1038/s41592-019-0598-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.ESM Team (2024) ESM Cambrian: Revealing the Mysteries of Proteins with Unsupervised Learning. EvolutionaryScale Website
- 34.Wang D, Pourmirzaei M, Abbas UL, Zeng S, Manshour N, Esmaili F, Poudel B, Jiang Y, Shao Q, Chen J, Xu D (2025) S-PLM: Structure-Aware Protein Language Model via Contrastive Learning Between Sequence and Structure. Adv Sci 12(5):2404212. 10.1002/advs.202404212 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Yu T, Cui H, Li JC, Luo Y, Jiang G, Zhao H (2023) Enzyme function prediction using contrastive learning. Science 379(6639):1358–1363. 10.1126/science.adf2465 [DOI] [PubMed] [Google Scholar]
- 36.Chen CS, Zhou J, Wang F, Liu X, Dou D (2023) Structure-aware protein self-supervised learning. Bioinformatics 39(4):189. 10.1093/bioinformatics/btad189 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Dieckhaus H, Brocidiacono M, Randolph NZ, Kuhlman B (2024) Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc Natl Acad Sci 121(6):2314853121. 10.1073/pnas.2314853121 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Uzma M, U, Halim Z, (2023) Protein encoder: An autoencoder-based ensemble feature selection scheme to predict protein secondary structure. Exp Syst Appl. 10.1016/j.eswa.2022.119081 [Google Scholar]
- 39.Schmirler R, Heinzinger M, Rost B (2024) Fine-tuning protein language models boosts predictions across diverse tasks. Nat Commun 15(1):7407. 10.1038/s41467-024-51844-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Dickson A, Mofrad MRK (2024) Fine-tuning protein embeddings for functional similarity evaluation. Bioinformatics (Oxford England) 40(8):445. 10.1093/bioinformatics/btae445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Zeng S, Wang D, Jiang L, Xu D (2024) Parameter-efficient fine-tuning on large protein language models improves signal peptide prediction. Genome Res 34(9):1445–1454. 10.1101/gr.279132.124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Moffat L, Jones DT (2021) Increasing the accuracy of single sequence prediction methods using a deep semi-supervised learning framework. Bioinformatics 37(21):3744–3751. 10.1093/bioinformatics/btab491 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Dhanuka R, Tripathi A, Singh JP (2022) A Semi-Supervised Autoencoder-Based Approach for Protein Function Prediction. IEEE J Biomed Health Inform 26(10):4957–4965. 10.1109/JBHI.2022.3163150 [DOI] [PubMed] [Google Scholar]
- 44.Asgari E, Mofrad MRK (2015) ProtVec: A Continuous Distributed Representation of Biological Sequences. PLoS ONE 10(11):0141287. 10.1371/journal.pone.0141287 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Öztürk H, Özgür A, Ozkirimli E (2018) DeepDTA: Deep drug-target binding affinity prediction. Bioinformatics 34(17):821–829. 10.1093/bioinformatics/bty593 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Nguyen T, Le H, Quinn TP, Nguyen T, Le TD, Venkatesh S (2021) GraphDTA: Predicting drug-target binding affinity with graph neural networks. Bioinformatics 37(8):1140–1147. 10.1093/bioinformatics/btaa921 [DOI] [PubMed] [Google Scholar]
- 47.Jiang M, Li Z, Zhang S, Wang S, Wang X, Yuan Q, Wei Z (2020) Drug-target affinity prediction using graph neural network and contact maps. RSC Adv 10(35):20701–20712. 10.1039/D0RA02297G [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Mukherjee S, Ghosh M, Basuchowdhuri P (2022) DeepGLSTM, Deep Graph Convolutional Network and LSTM based approach for predicting drug-target binding affinity. pp 729–737. 10.1137/1.9781611977172.82
- 49.Zhao Q, Duan G, Yang M, Cheng Z, Li Y, Wang J (2023) AttentionDTA: Drug-Target Binding Affinity Prediction by Sequence-Based Deep Learning With Attention Mechanism. IEEE/ACM Trans Comput Biol Bioinf 20(2):852–863. 10.1109/TCBB.2022.3170365 [DOI] [PubMed] [Google Scholar]
- 50.Wang S, Song X, Zhang Y, Zhang K, Liu Y, Ren C, Pang S (2023) MSGNN-DTA: Multi-Scale Topological Feature Fusion Based on Graph Neural Networks for Drug-Target Binding Affinity Prediction. Int J Mol Sci 24(9):8326. 10.3390/ijms24098326 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Abbasi K, Razzaghi P, Poso A, Amanlou M, Ghasemi JB, Masoudi-Nejad A (2020) DeepCDA: Deep cross-domain compound-protein affinity prediction through LSTM and convolutional neural networks. Bioinformatics 36(17):4633–4642. 10.1093/bioinformatics/btaa544 [DOI] [PubMed] [Google Scholar]
- 52.Yuan W, Chen G, Chen CY-C (2022) FusionDTA: Attention-based feature polymerizer and knowledge distillation for drug-target binding affinity prediction. Brief Bioinform 23(1):506. 10.1093/bib/bbab506 [DOI] [PubMed] [Google Scholar]
- 53.Tang X, Zhou Y, Yang M, Li W (2024) TC-DTA: Predicting Drug-Target Binding Affinity With Transformer and Convolutional Neural Networks. IEEE Trans Nanobiosci 23(4):572–578. 10.1109/TNB.2024.3441590 [DOI] [PubMed] [Google Scholar]
- 54.Xiao X, Wang W, Xie J, Zhu L, Chen G, Li Z, Wang T, Xu M (2024) HGTDP-DTA Hybrid Graph-Transformer with Dynamic Prompt for Drug-Target Binding Affinity Prediction. 10.48550/arXiv:2406.17697
- 55.Qi H, Yu T, Yu W, Liu C (2024) Drug-target affinity prediction with extended graph learning-convolutional networks. BMC Bioinform 25:75. 10.1186/s12859-024-05698-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Zhao L, Wang H, Shi S (2024) PocketDTA: An advanced multimodal architecture for enhanced prediction of drug-target affinity from 3D structural data of target binding pockets. Bioinformatics 40(10):594. 10.1093/bioinformatics/btae594 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Ranjan A, Bess A, Alvin C, Mukhopadhyay S (2024) MDF-DTA: A Multi-Dimensional Fusion Approach for Drug-Target Binding Affinity Prediction. J Chem Inf Model 64(13):4980–4990. 10.1021/acs.jcim.4c00310 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Wu H, Liu J, Jiang T, Zou Q, Qi S, Cui Z, Tiwari P, Ding Y (2024) AttentionMGT-DTA: A multi-modal drug-target affinity prediction using graph transformer and attention mechanism. Neural Netw 169:623–636. 10.1016/j.neunet.2023.11.018 [DOI] [PubMed] [Google Scholar]
- 59.Quan L, Wu J, Jiang Y, Pan D, Qiang L (2024) DTA-GTOmega: Enhancing Drug-Target Binding Affinity Prediction with Graph Transformers Using OmegaFold Protein Structures. J Mol Biol. 10.1016/j.jmb.2024.168843 [DOI] [PubMed] [Google Scholar]
- 60.Pei Q, Wu L, Zhu J, Xia Y, Xie S, Qin T, Liu H, Liu T-Y, Yan, R (2023) SSM-DTA, Breaking the Barriers of Data Scarcity in Drug-Target Affinity Prediction 10.48550/arXiv:2206.09818 [DOI] [PubMed]
- 61.Li Z, Liu F, Yang W, Peng S, Zhou J (2022) A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Trans Neural Netw Learn Syst 33(12):6999–7019. 10.1109/TNNLS.2021.3084827 [DOI] [PubMed] [Google Scholar]
- 62.Yu Y, Si X, Hu C, Zhang J (2019) A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures. Neural Comput 31(7):1235–1270. 10.1162/neco_a_01199 [DOI] [PubMed] [Google Scholar]
- 63.Denecke K, May R, Rivera-Romero O (2024) Transformer Models in Healthcare: A Survey and Thematic Analysis of Potentials, Shortcomings and Risks. J Med Syst 48(1):23. 10.1007/s10916-024-02043-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Chandra A, Tünnermann L, Löfstedt T, Gratz R (2023) Transformer-based deep learning for predicting protein properties in the life sciences. Elife 12:82819. 10.7554/eLife.82819 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Wu R, Ding F, Wang R, Shen R, Zhang X, Luo S, Su C, Wu Z, Xie Q, Berger B, Ma J, Peng J (2022) High-Resolution de Novo Structure Prediction from Primary Sequence. 10.1101/2022.07.21.500999
- 66.Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, Guo D, Ott M, Zitnick CL, Ma J, Fergus R (2021) Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci 118(15):2016239118. 10.1073/pnas.2016239118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Zhang X-M, Liang L, Liu L, Tang M-J (2021) Graph Neural Networks and Their Current Applications in Bioinformatics. Front Genet 12:690049. 10.3389/fgene.2021.690049 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Li MM, Huang K, Zitnik M (2022) Graph representation learning in biomedicine and healthcare. Nat Biomed Eng 6(12):1353–1369. 10.1038/s41551-022-00942-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Shehzad A, Xia F, Abid S, Peng C, Yu S, Zhang D, Verspoor K (2024) Graph Transformers A Survey 10.48550/arXiv:2407.09777 [DOI] [PubMed]
- 70.Yun S, Jeong M, Kim R, Kang J, Kim HJ (2020) Graph Transformer, Networks 10.48550/arXiv:1911.06455
- 71.Libouban P-Y, Parisel C, Song M, Aci-Sèche S, Gómez-Tamayo JC, Tresadern G, Bonnet P (2025) Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction. Bioinformatics 41(8):429. 10.1093/bioinformatics/btaf429 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Kumar R, Romano JD, Ritchie MD (2025) CASTER-DTA: Equivariant graph neural networks for predicting drug-target affinity. Brief Bioinform 26(5):554. 10.1093/bib/bbaf554 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Sun D, Huang H, Li Y, Gong X, Ye, Q (2023) DSR, Dynamical surface representation as implicit neural networks for protein, vol ’23. Curran Associates Inc., Red Hook, NY, USA, pp 13873–13886
- 74.Wu F, Hu B, Li SZ (2025). Generalized implicit neural representations for dynamic molecular surface modeling, vol AAAI’25/IAAI’25/EAAI’25, vol. 39. AAAI Press, Washington, D.C., USA, pp 877–885. 10.1609/aaai.v39i1.32072 [Google Scholar]
- 75.Xu L, Xie H, Qin S-ZJ, Tao X, Wang FL (2023) Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models. A Critical Review and Assessment 10.48550/arXiv:2312.12148 [DOI] [PubMed]
- 76.Rossi E, Chamberlain B, Frasca F, Eynard D, Monti F, Bronstein M (2020) Temporal Graph Networks for Deep Learning on Dynamic Graphs 10.48550/arXiv:2006.10637
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.
No code was used or written for this paper.











