Skip to main content
mAbs logoLink to mAbs
. 2025 Sep 26;17(1):2562997. doi: 10.1080/19420862.2025.2562997

Accelerating antibody development: sequence and structure-based models for predicting developability properties via size exclusion chromatography

A N M Nafiz Abeer a,b,*, Mehdi Boroumand a, Isabelle Sermadiras c, Jenna G Caldwell d, Valentin Stanev a, Neil Mody d, Gilad Kaplan c, James Savery a, Rebecca Croasdale-Wood c, Maryam Pouryahya a,
PMCID: PMC12477876  PMID: 41004127

ABSTRACT

Experimental screening for biopharmaceutical developability properties typically relies on resource-intensive, and time-consuming assays such as size exclusion chromatography (SEC). This study highlights the potential of in silico models to accelerate the screening process by exploring sequence and structure-based machine learning techniques. Specifically, we compared surrogate models based on pre-computed features extracted from sequence and predicted structure with sequence-based approaches using protein language models (PLMs) like ESM-2. In addition to different end-to-end fine-tuning strategies for PLM, we have also investigated the integration of the structural information of the antibodies into the prediction pipeline through graph neural networks (GNN). We applied these different methods for predicting protein aggregation propensity using a dataset of approximately 1200 Immunoglobulin G (IgG1) molecules. Through this empirical evaluation, our study identifies the most effective in silico approach for predicting developability properties for SEC assays, thereby adding insights to existing screening efforts for accelerating the antibody development process.

KEYWORDS: Antibody structure, developability properties, graph neural network, protein language model, size exclusion chromatography, therapeutic antibodies

1. Introduction

Monoclonal antibodies (mAbs) are effective therapeutic proteins due to their high specificity, versatility, and efficacy in targeting a wide range of diseases. They enable targeted therapy with minimal off-target effects, applicable in oncology, autoimmune, infectious, cardiovascular, and metabolic diseases.1 Advances in antibody engineering, such as humanization and affinity maturation, have enhanced their clinical efficacy and reduced immunogenicity, making mAbs safer and therapeutically effective for long-term use.2 With over 145 antibody-based drugs approved by the FDA and many more in clinical trials, their impact on modern medicine is substantial and continues to grow.3,4

The developability properties of antibodies are crucial for their transition from early-stage discovery to large-scale manufacturing. Since the antibody design process prioritizes the binding activity in neutralizing the target protein, the resulting antibodies may end up with unfavorable biophysical attributes hindering its progression into the next development stages.5,6 Hence, it is extremely beneficial to screen antibodies not only for target specificity and binding affinity but also for physicochemical and biopharmaceutical properties such as solubility, stability, aggregation propensity, and manufacturability.7 These properties – which quantifies the developability of antibodies – significantly influence the antibody’s success in later development stages, including clinical trials and commercialization.8 Poor developability can lead to high attrition rates, increased costs, and extended timelines due to formulation challenges, instability, or immunogenicity.9 Systematic early evaluation of these properties can streamline development, advancing only the most promising candidates.10 This proactive approach also aids in designing robust manufacturing processes to produce high-quality therapeutic antibodies at scale, reducing late-stage failures and ensuring a more efficient path to commercialization.2,11

Various assays and methods, such as size exclusion chromatography (SEC)12 – which is the focus of this paper – dynamic light scattering (DLS),13 differential scanning calorimetry (DSC),14 and isoelectric focusing (IEF),15 are employed to measure the physicochemical attributes of antibodies, serving as proxy evaluations for their developability characteristics. These measurements are utilized to detect aggregates, assess conformational stability, evaluate chemical degradation, and determine both the isoelectric point and charge isoform content, all of which impact the molecule’s efficacy, solubility, and overall chemical stability.2,5,7 In this work, we have considered the SEC assay which offers unique advantages in the purification and characterization of biomolecules. This chromatographic method, which separates analytes based on molecular size, plays a critical role in ensuring the purity, efficacy, and safety of therapeutic proteins. SEC allows for the effective separation of desired therapeutic proteins from aggregates and impurities that may affect the drug’s safety and efficacy. Furthermore, SEC is crucial for the detailed characterization of protein biotherapeutics, providing insights into their molecular weight distribution, aggregation state, and stability, which are essential parameters for regulatory approval and clinical success.16,17 The technique’s ability to operate under mild conditions without altering the biological activity of the molecules makes it particularly valuable for the analysis of sensitive biologics.18,19

While the experimental assays like SEC provide valuable insights into the development of biologics, they are often time-consuming and costly. The high-throughput screening methods required to evaluate a large number of candidates can be resource-intensive, requiring significant investment in both equipment and materials.20 Machine learning and in silico approaches can significantly enhance the prediction of antibody developability by leveraging medium to high-throughput datasets. These methods aid in early-stage drug development, enabling the rapid selection of lead candidates and potentially reducing the need for extensive high-throughput analytical measurements. However, the limited amount of experimental assay datapoints poses a critical challenge in building a reliable surrogate model for predicting the developability property of interest. Existing efforts21–25 primarily involve processing the sequences (with or without 3D protein structure) to compute the protein descriptors that are utilized by machine learning models for prediction. In addition to the computational burden of generating the features from structures, the performance of this approach is sensitive to the feature selection process. The protein language models (PLMs) offer a faster and more efficient alternative to utilize the structure information of the proteins. Leveraging the large pool of protein sequences, the PLMs learn to incorporate the structural information implicitly into the sequence embedding.26 With the advancement of protein language models, there have been efforts27,28 to predict the protein property from the sequence embedding learned by the pre-trained PLM. Since the PLMs work on the sequence representations of the antibodies, this approach has the potential for an accurate and faster screening pipeline by removing the need for prediction of structure and processing of structural features. On the other hand, several works28–30 explicitly leverage the 3D protein structures along with the protein language model for the prediction of protein properties. In the absence of experimental structure data, this approach relies on either the protein folding tools like AlphaFold2,31 or homology modeling to build the structure from the protein sequence.

The performance of the above-mentioned approaches for predicting the protein properties varies across different types of assays. In this work, we have investigated the application of these different prediction approaches for the SEC assay. Specifically, we have considered the classification task of two developability properties – monomer content and difference in retention time of the IgG1 molecules to a reference sample. To summarize our work:

• We performed experiments with four prediction pipelines – one based on sequence and structure-based features, three others leveraging the protein language model and graph neural network – to select the best performing configuration (e.g. which features/PLM/GNN to use) under each pipeline.

• We have identified the best prediction strategy for two SEC properties of interest based on the hold-out test set performance by these four pipelines.

• Furthermore, we have assessed the impact of two different protein structure prediction tools on the developability prediction pipeline.

2. Methodology

2.1. High-performance size exclusion chromatography assay

HPSEC assays help determine the purity of protein samples by measuring the percentage of protein monomer (main product), HMWFs (higher molecular weight forms) and LMWFs (lower molecular weight forms). The chromatogram results from SEC assays typically display several peaks corresponding to detected species. Peaks appearing before the expected monomer peak are classified as HMWFs, while those appearing after are identified as LMWFs. This interpretation is crucial for understanding the composition and quality of the protein sample. The chromatograms also provide the retention time of peaks, which indicate the elution time of the protein in the chromatography columns. The SEC retention time of the monomer is primarily influenced by its molecular weight, but it is also affected by other molecular properties, including charge, hydrophobicity, and self-association,32 suggesting that SEC can offer a multifaceted view of the molecule’s behavior.5,33 To normalize retention times of monomer content across multiple studies utilized in this paper, ΔRT is considered, which compares the retention time of the monomer protein to a reference protein, NIP228, with a known SEC retention time of approximately 8.47 minutes. This normalization is essential for consistent analysis across different studies.19,34 In this work, we focus on predicting the percentage peak area for the monomer product and the delta retention time of the monomer peak. Since these assays are primarily used to distinguish problematic molecules from developable ones using well-defined thresholds, we concentrate mainly on a classification problem to differentiate between the two.

2.2. Problem statement

Experimental assays for screening antibodies, such as SEC, can characterize multiple attributes of the molecule under consideration (Section 2.1). For one of such properties, we assume to have a dataset D={s(i),o(i)} where s denotes the sequence representation of the antibody, i.e. both heavy and light chains. The attribute oR corresponds to the observation value directly obtained through the assay. During the screening stage, the sample is classified as desirable or problematic by comparison with a pre-defined specification, usually based on developability requirements. By denoting this process as fscreen we can get the binary label, y{0,1}, i.e. desirable/problematic for each sample as:

y(i)=fscreen(o(i))=1o(i)Iproblematic (1)

1 is an indicator function which labels the sample 1 (problematic) when the corresponding observation falls within the interval Iproblematic. This interval expectedly varies across different properties (Figure 1).

Figure 1.

Two subfigures: (a) and (b) – showing the histograms of monomer % and ∆ RT respectively for the antibodies in the dataset. Subfigure (a) has a vertical red line near 100%, and subfigure (b) has two vertical red lines centered around zero.

Data distribution of SEC monomer % and ΔRT. Problematic antibodies are characterized by lower monomer content and ΔRT values further from zero, as indicated by the red dashed lines in the figure.

Given the data with binary labels, Dbin={s(i),y(i)}, our goal is to build a classifier to identify an Immunoglobulin G (IgG1) molecule as a desirable (developable) or problematic sample from information embedded in s. Once we have such a classifier trained, this can serve as a surrogate for the experimental assay in the in silico screening process.

2.3. Dataset and train-test split

The SEC dataset, consisting of around 1200 IgG1 molecules, was collected from multiple internal studies. We considered the monomer content percentage and the ΔRT of the antibodies relative to a reference antibody, NIP228, as previously described. Duplicate samples, defined as multiple assay observations for an identical antibody sequence, were removed for each property. The processed dataset was then divided into training (90%) and test (10%) partitions. Due to the dataset’s collection from various studies, the sequences exhibit considerable diversity, with a median of over 100 mutations between pairs. Further details regarding sequence diversity are provided in the supplementary section. Despite this diversity, we still specifically chose the test set to remain diverse, based on clusters formed according to sequence similarities (Levenshtein distance with representative members of clusters chosen randomly35,36). This selection also ensured a similar stratified distribution of the two classes as observed in the train set which has 71% and 69% samples from class 0 for monomer % and ΔRT respectively. The test split serves as a hold-out set for evaluating the efficacy of the prediction pipelines (Section 3.2.2).

2.4. Prediction pipeline

The primary data modality in our study is the sequence representation. s=(sheavy,slight) of the antibody. All molecules in our dataset belong to the IgG1 subclass and share a similar constant region. This bias to similar constant region is introduced by the antibody design approach where the variable fragment sequences are lead optimized, followed by grafting to the known constant framework. Consequently, our predictions concentrate exclusively on the impact of sequences within the variable fragments (Fvs). Through the application of the AlphaFold2 (AF2)31 or similar protein folding tools, one can also predict the 3D structure folded from the sequence information alone, adding another modality to our prediction pipeline. While explicit incorporation of structure information provides richer information than the sequence for predicting protein properties, one also needs to consider possible errors propagated from the protein structure prediction tool’s inaccuracy.

In this work, we have considered the following four approaches (illustrated in Figure 2) for building the prediction network for the developability property of interest by leveraging different combinations of both modalities. Details of each pipeline are discussed in the subsequent sections.

Figure 2.

Four subfigures showing different components of each prediction workflow. In subfigure (a), the heavy and light chain sequences are input to structure prediction, followed by the protein descriptor generator to process sequence and structure-based features, which go through feature selection and are used in ML model. The subfigure (b) depicts prediction using sequence embedding derived from light and heavy chains using a protein language model. In subfigures (c) and (d), the graph neural network predicts the developability property from the k-NN residue level graph constructed from the antibody structure predicted using light and heavy chain sequences.

Developability property prediction workflow. All four pipelines except PLM (b) leverage the predicted 3D protein structure into their prediction network by either protein descriptor generator (a) or graph neural network (c and d). Note that GNN pipeline (c) utilizes the fixed node attributes which can include the embeddings from a pre-trained protein language model. In the PLM+GNN (d), we allow the PLM to be updated jointly with the GNN during training of the prediction network.

Sequence and Structure-based Features: An ML classifier predicts the target from the selected protein features, processed by a protein descriptor generator from the antibody sequence and its predicted structure.

PLM Pipeline: The pipeline utilizes only the sequence information through a combination of a protein language model (PLM) and a prediction head. The latter component is a shallow multilayer perceptron (MLP) network that predicts the property from concatenated heavy and light chain sequences embedding projected by the PLM.

GNN Pipeline: The graph neural network (GNN) predicts the property from the amino acid (AA) graph constructed from the predicted structure.

PLM + GNN Pipeline: This incorporates the residue embedding from the protein language model into the GNN pipeline as the node attributes of the AA graph and the combined network is trained jointly.

2.4.1. Prediction utilizing sequence and structural features

We used Schrödinger software to extract the molecular properties from the predicted AF2 static structure https://www.schrodinger.com/(Schrödinger, LLC). Schrödinger software offers an extensive suite of computational tools for molecular modeling, simulation, and extraction of protein descriptors. In our workflow, we utilized AF2 instead of Schrödinger’s homology modeling and employed Schrödinger solely for extracting molecular properties and biophysical features. Schrödinger provides an exhaustive list of patch-level sequence- and structure-based protein properties, including positive and negative charges, hydrophobicity, and aggregation propensities. To ensure we selected the most informative features for our final model, we utilized a combination of unsupervised and supervised feature selection Section 3.2.1. Finally, an Extra Trees classifier37 was applied to the selected features for predicting the developability properties.

2.4.2. PLM-based prediction network from sequence representation

The core idea of the PLM pipeline is to utilize the sequence embedding of an antibody sequence from the protein language model in combination with a simple prediction head network to classify the developability property. For the heavy chain and light chain in the antibody sequence, we consider separate instances of the protein language model, ϕPLM,H and ϕPLM,L respectively. We first compute the hidden states of each chain using corresponding instance of PLM. Specifically, we denote residue embedding matrix EH,EL as the last hidden state of all tokens of each chain. Note these embedding matrices can include the hidden state of special tokens added for the PLM. Next, we compute the sequence-level representation, i.e. eH,eL for each chain by pooling its embedding matrix. We have explored two aggregation techniques to do this – mean pooling where the residue embeddings (excluding any special tokens) are averaged and CLS pooling where embedding for the special token [CLS] is considered as the sequence-level embedding. We have considered [CLS] token since it is often used in end-to-end fine-tuning28 of the language models. Finally, we concatenate the sequence embeddings eH and eL and pass it to a shallow MLP network ϕMLP to predict the class label (i.e., the logits) for the sequence s.

EH=ϕPLM,H(sheavy)EL=ϕPLM,L(slight) (2)
eH=Pooling(EH)eL=Pooling(EL) (3)
yˆ=ϕMLP(eHeL) (4)

In case of general protein language model like ESM2,38,39 ϕPLM,H and ϕPLM,L are initialized at the pre-trained weights. One can also consider the antibody chain-specific pre-trained models like AbLang-1.40 When both instances are frozen at their pre-trained model weights, and we learn the prediction head ϕMLP, we denote this as “fixed PLM” approach in our work. On the other hand, we can also fine-tune those PLM instances while training the prediction head by allowing the gradient information backpropagated to the layers of the PLMs. We have investigated two ways of performing this fine-tuning – full parameter fine-tuning and low rank adaptation (LoRA)41 technique. The latter approach learns low-rank transformation matrices (for the attention networks in our work) that work in conjunction with pre-trained weights to adapt the model for a specific task.

2.4.3. GNN-Based prediction network leveraging structure

To leverage the protein structural information explicitly in the prediction pipeline, the GNN pipeline begins with the construction of the amino acid graph for the antibody sequence s from the predicted 3D structure. For each residue in the sequence s, the top k nearest residues are considered to be its neighbors based on the coordinates of Cα of the residues. Up to this step, the procedure is similar to the approaches of.28,30,42,43 In our work, however, we add a refinement step by pruning residues that lie outside a predefined local region. Hence, in the resulting AA graph G, the edge from node j to node i exists if residue j is in the k-nearest neighbors of residue i and distance of the edge, ||eji||2 is lower than rthr Å. In our work, we set rthr=9 which results in around 13 neighbors on average for the nodes of AA graphs in our dataset.

The m layers of GNN, denoted as ϕGNN projects the node attribute matrix XRL×Fi of AA graph G for sequence s with L residues into the hidden embeddings matrix HRL×Fh. Fi and Fh are the dimension of node attributes and hidden representation respectively. The prediction head ϕMLP transforms the pooled embedding from H into the class probability for the antibody sample.

H=ϕGNNX,G (5)
yˆ=ϕMLP(Pooling(H)) (6)

For the node attributes of the graph, we have considered the one-hot encoding, amino acid properties from,44,45 node embedding from the pre-trained variational graph autoencoder (VGAE) of46 and residue embedding from pre-trained ESM2 (8 M) model. In case of the VGAE and ESM2 (8 M), the weights of these models are frozen at the pre-trained values in the entire training process. To evaluate the effectiveness of GNN in processing the AA graph, three different GNNs are considered – geometric vector perception (GVP),43 graph attention network (GAT)47 and graph isomorphism network (GIN).48 The main difference between these three is that the GVP leverages the edge features in the message-passing updates while the other two do not. Also, the GVP considers the scalar and vector features for the nodes and edges where the scalar features of the nodes are based on dihedral angles and one of the four node attributes mentioned earlier. In case of other two GNNs, only the node attributes are utilized.

After 3 layers of GNN, the updated node embeddings are aggregated to get the global embedding for the entire graph. Here, we have considered the mean pooling technique which takes an average of node embeddings across all nodes of the graph. In summary, we have explored 4 choices of node attributes and 3 choices of GNN for each developability property.

2.4.4. Prediction network leveraging PLM and GNN

In terms of architectural similarity, the PLM+GNN pipeline is the same as the GNN-based prediction network where the node attributes of the amino acid residue graph are assigned by the pre-trained protein language model. While in the GNN pipeline, we only learn the GNN modules during training, the PLM+GNN pipeline additionally facilitates the adaptation of the parameters of the PLM leveraging the gradients at the node attributes backpropagated from the GNN modules as in.28 For an antibody sequence s, first the PLM module ϕPLM computes the node attribute matrix X. The rest of the steps are similar to the GNN pipeline where X is transformed to the prediction through ϕGNN and ϕMLP by using the corresponding amino acid graph G.

H=ϕGNNϕPLM(s),G (7)
yˆ=ϕMLP(Pooling(H)) (8)

For this approach, we have considered 3 PLMs – (ESM2 (8 M),38 AbLang-140 and AbLang-249). For the ESM2 (8 M), we allow the gradient flow through its all layers excluding the embedding layer. In the case of fine-tuning of AbLang-1 and AbLang-2, the first 6 layers are frozen at their pre-trained weights while the rest of the layers are updated based on the gradient backpropagated from the GNN module. In addition to the global mean pooling of the GNN pipeline, we have investigated the effectiveness of universal pooling30,50 over the mean pooling technique. Although the sequences in our work are not processed using the multiple sequence alignment (MSA) technique, our exploration attempts to analyze the robustness of this pooling technique for three different choices of GNN.

3. Result and discussion

3.1. Processing of 3D protein structure

In this work, we utilized two state-of-the-art computational tools for predicting protein structures: AlphaFold 2 and ImmuneBuilder. AlphaFold 2, developed by DeepMind, has revolutionized the field by accurately predicting protein structures from amino acid sequences, applicable across a wide range of proteins.31 ImmuneBuilder, on the other hand, is tailored for rapid and accurate predictions of immune protein structures, such as antibodies and T-cell receptors. Noted for its speed, ImmuneBuilder is significantly faster than AlphaFold 251 and more suitable for high-throughput screening of large numbers of antibodies. Therefore, we conducted a comprehensive comparison of the structure prediction models in Section 3.2.3.

3.2. Experiments

3.2.1. Selecting the best combination for each pipeline

We have explored different combinations for each of the four prediction pipelines discussed in Section 2.4. In each pipeline, we have run 10 trials of 10-fold cross-validation on the SEC dataset using different combinations and selected the best configuration based on the average prediction performance on the validation splits. In the Supplementary Information, we have provided pipeline-specific procedures and explained the corresponding cross-validation results in detail.

3.2.2. Results on test set

From the cross-validation experiments of Section 3.2.1, we have only identified the best combination for each pipeline. Next, we have trained each best combination with 10-fold cross-validation with the training split and evaluated the trained models’ performance on the hold-out test set. Specifically, we split the training data (90% of the SEC dataset) into 10-fold cross-validation splits, and we pick the model with best performance (accuracy) on the validation split in each of 10 folds. These 10 trained models for each pipeline are applied on the hold-out test set (Section 2.3) to measure the predictive performance. Tables 1 and 2 show the average and standard deviation of performance metrics across these 10 models for each pipeline corresponding to monomer percentage and ΔRT respectively.

Table 1.

Prediction performance on the hold-out test set for SEC monomer %. The majority class (label 0) predictor has an accuracy of 71%. ESM2 (8 M) model fine-tuned via LoRA technique produces the sequence embedding for “PLM” pipeline. For “GNN” pipeline, the embedding from pre-trained VGAE is used with 3 layers of GVPs. For “PLM+GNN,” a combination of AbLang1 and GVP is trained in an end-to-end fashion. The best-performing pipeline is highlighted for each metric based on their average performance over the hold-out test set for 10 models from 10-Fold cross-validation.

Prediction Pipeline Accuracy F1 Sensitivity Precision
PLM 0.75 (0.03) 0.54 (0.04) 0.52 (0.05) 0.58 (0.13)
GNN 0.74 (0.02) 0.49 (0.03) 0.44 (0.04) 0.58 (0.12)
PLM + GNN 0.75 (0.01) 0.50 (0.04) 0.43 (0.06) 0.61 (0.14)
Sequence + Structure based features 0.77 (0.01) 0.49 (0.04) 0.38 (0.05) 0.66 (0.15)
Table 2.

Prediction performance on the hold-out test set for SEC ΔRT. The majority class (label 0) predictor has an accuracy of 69%. The “PLM” pipeline utilizes the ESM2 (8 M) model obtained after the full parameter fine-tuning with the training set. For “GNN” pipeline, the embedding from pre-trained ESM2 (8 M) is used with 3 layers of GVPs. For “PLM+GNN,” a combination of AbLang1 and GVP is trained in an end-to-end fashion. The best-performing pipeline is highlighted for each metric based on their average performance over the hold-out test set for 10 models from 10-Fold cross-validation.

Prediction Pipeline Accuracy F1 Sensitivity Precision
PLM 0.77 (0.02) 0.60 (0.04) 0.56 (0.06) 0.66 (0.12)
GNN 0.80 (0.03) 0.64 (0.07) 0.58 (0.10) 0.68 (0.15)
PLM + GNN 0.78 (0.04) 0.61 (0.13) 0.57 (0.15) 0.58 (0.19)
Sequence + Structure based features 0.80 (0.01) 0.59 (0.03) 0.46 (0.05) 0.77 (0.12)

For both properties, the pipeline utilizing sequence and structure-based features achieves the highest accuracy while performing poorly in selecting problematic molecules (class label 1) as evidenced by the low value in sensitivity. The other three pipelines show slightly lower accuracy in SEC monomer percentage and out of these three, the PLM pipeline produces better F1 score and sensitivity. This result is particularly interesting for high-throughput screening since this PLM pipeline predicts the developability property from the sequences, making it much faster than “sequence + structure-based features” which requires computationally time-consuming feature processing by the Schrödinger suite. In the case of the SEC ΔRT, the GNN pipeline and the feature-based pipeline have similar accuracy, but the former shows better sensitivity. When the structural information is combined with the PLM approach resulting in PLM+GNN, the average performance stays similar to the PLM approach (within the standard deviation) but this causes larger variation in performance across different models learned in each of the 10 folds.

All four pipelines have a similar standard deviation in F1 score (and sensitivity) for the SEC monomer percentage property. On the other hand for SEC ΔRT, the GNN and PLM+GNN pipelines show comparatively larger fluctuations in similar performance metrics. For example, the standard deviations in sensitivity for these two pipelines are 0.10 and 0.15 respectively (Table 2) which are 2–3 times larger than the other two pipelines. Given that the PLM approach shows similar fluctuations as the pipeline with sequence and structure-based features, this higher variation in the performance metrics possibly originates from the sensitivity of GNN components to AA graphs seen in the 10-fold cross-validation training. Additionally, the difference in this trend between the monomer percentage and ΔRT indicates that the structural information of the antibody molecules may have more importance in the predictive performance of the latter property.

3.2.3. Ablation study for different structure prediction tools

The results in Tables 1 and 2 are for the experiments utilizing the predicted structure from AlphaFold2. Despite having a high structure prediction performance, AlphaFold2 can be a source of bottleneck in the high-throughput screening process due to its longer structure prediction time. In this section, we used a faster structure prediction tool, ImmuneBuilder51 for the three pipelines that use antibody structure in predicting the developability properties. By comparing with the previous results (from AlphaFold2), we assessed the robustness of the developability prediction pipeline to the predicted structures.

For three developability prediction pipelines – Sequence + Structure based features, GNN, PLM+GNN – we have repeated the hold-out test set experiment from Section 3.2.2. The models under each pipeline have the same architecture as in Section 3.2.2 but are trained with the datapoints where the structures are predicted via ImmuneBuilder. Tables 3 and 4 have these hold-out set performance metrics for monomer % and ΔRT respectively. We have also shown the results with AlphaFold2 (from Tables 1 and 2) for reference. For both properties, we observe a decline in the sensitivity (and F1 score) for Sequence + Structure based features and GNN pipelines when we replace AlphaFold2 with ImmuneBuilder as the antibody structure prediction tools. The performance of PLM+GNN pipeline remains similar for SEC monomer % while we see a more stable performance trend for SEC ΔRT. With AlphaFold2, the GNN pipeline is the best performing pipeline for ΔRT but has a relatively large performance variation. In the PLM+GNN pipeline for ΔRT, the ImmuneBuilder not only works as a faster structure prediction tool but provides a comparable performance consistently (0.62(0.04) vs 0.61(0.13) with AlphaFold2).

Table 3.

Variation in predictive performance for SEC monomer % due to different structure prediction tools. The neural network architectures for three prediction pipelines are the same as in Table 1.

Prediction Pipeline Structure Prediction Tool Accuracy F1 Sensitivity Precision
GNN AlphaFold2 0.74 (0.02) 0.49 (0.03) 0.44 (0.04) 0.58 (0.12)
ImmuneBuilder 0.75 (0.02) 0.46 (0.04) 0.38 (0.04) 0.59 (0.15)
PLM + GNN AlphaFold2 0.75 (0.01) 0.50 (0.04) 0.43 (0.06) 0.61 (0.14)
ImmuneBuilder 0.74 (0.02) 0.49 (0.05) 0.44 (0.08) 0.56 (0.15)
Sequence + Structure based features AlphaFold2 0.77 (0.01) 0.49 (0.04) 0.38 (0.05) 0.66 (0.15)
ImmuneBuilder 0.76 (0.01) 0.45 (0.04) 0.34 (0.04) 0.64 (0.16)
Table 4.

Variation in predictive performance for SEC ΔRT due to different structure prediction tools. The neural network architectures for three prediction pipelines are the same as in Table 2.

Prediction Pipeline Structure Prediction Tool Accuracy F1 Sensitivity Precision
GNN AlphaFold2 0.80 (0.03) 0.64 (0.07) 0.58 (0.10) 0.68 (0.15)
ImmuneBuilder 0.78 (0.03) 0.59 (0.07) 0.52 (0.10) 0.64 (0.16)
PLM + GNN AlphaFold2 0.78 (0.04) 0.61 (0.13) 0.57 (0.15) 0.58 (0.19)
ImmuneBuilder 0.79 (0.01) 0.62 (0.04) 0.55 (0.06) 0.71 (0.12)
Sequence + Structure based features AlphaFold2 0.80 (0.01) 0.59 (0.03) 0.46 (0.05) 0.77 (0.12)
ImmuneBuilder 0.79 (0.01) 0.54 (0.02) 0.40 (0.02) 0.81 (0.09)

4. Conclusion

Our work has explored four different prediction approaches for two developability properties from the SEC assays: monomer percentage and ΔRT. We aimed to provide promising in silico models that can be used to screen less developable molecules at an early stage without the need for time- and material-intensive experimental methods. We investigated high-throughput models that require only sequences (protein language models) as well as models that require protein structure prediction and feature extraction from predicted structures, which are computationally more demanding. Specifically, we examined whether protein language models and graph neural networks can effectively utilize sequence and structural information to achieve performance comparable to methods employing explicit structural and sequence features calculated by Schrödinger software. For both properties, the performance on the hold-out test set shows that the latter approach achieves higher accuracy but suffers from considerable degradation in its ability to identify molecules violating the developability constraint. The best-performing pipeline (in terms of F1 score) for monomer percentage is the protein language model (PLM) approach, which offers a promising opportunity for high-throughput screening due to its faster inference time from the sequence of the antibody chains. Further comparison of variations in the performance metrics for ΔRT and monomer percentage indicates the relative importance of structural information in predicting these two properties. Finally, we have identified (in Section 3.2.3) the potential of the PLM+GNN pipeline as a high-throughput screening tool for ΔRT, with ImmuneBuilder (replacing AlphaFold2) as the protein structure prediction tool.

Since the dataset used in our work was compiled from multiple projects, it naturally encompasses a diverse range of sequences. Nonetheless, in Lead Optimization (LO) projects, new antibodies may exhibit mutations in regions that were previously unmutated in the training data, and Lead Identification (LI) projects might yield sequences that are significantly different. Specifically, the prediction pipelines that use the predicted structure, e.g., sequence and structure-based features, PLM+GNN and GNN, is likely to be negatively impacted by the structure prediction model’s low sensitivity to the small number of mutations, which are shared by the similar sequences in the LO projects. However, the PLM-based pipeline may be more robust in such scenario, as PLMs are shown to be more effective in encoding the mutational change in protein fitness. Assessing the robustness of the prediction pipelines in handling this challenge of out-of-distribution data is an important area for further exploration. Given the cost of measuring developability properties, our dataset of approximately 1,200 antibodies is considered moderately large. For other new or expensive low-throughput assays, we might not have access to such a large dataset, which makes it challenging to directly leverage the protein language model and graph neural networks. In those cases, it would be interesting to see whether the trained models using the relatively large dataset can give an edge to learning the prediction networks for low-data assay through transfer learning technique.52 A few recent works53–55 explicitly incorporate the structural information into the training of the protein language models. Exploration of their applicability in developability prediction can provide further insights into the importance of explicit structural information in protein property prediction.

Supplementary Material

SI_mAbs.pdf

Acknowledgments

We would like to express our sincere gratitude to all those who have contributed to the success of this research project. Special thanks go to Jurgen Haas, Christopher Lloyd, Beverley Smith, Robert Calvert, Andrew Dippel, Bismark Amofah, and Tony Pham for providing the necessary resources and facilities to conduct this research.

Funding Statement

The author(s) reported there is no funding associated with the work featured in this article.

Disclosure statement

M.B., I.S., J.G.C, V.S., N.M., G.K., J.S., R.C.W. and M.P. are employees of AstraZeneca. A.N.M.N.A. was employed as an intern at AstraZeneca.

Author contributions

Conceptualization: M.P., A.N.M.N.A., M.B., V.S., J.G.C., N.M., G.K., R.C.W.; Methodology: A.N.M.N.A., M.P.; Data Processing: I.S., G.K., A.N.M.N.A., M.P.; Visualization: A.N.M.N.A., M.P.; Supervision: M.P., J.S., R.C.W.; Writing – Original Draft: A.N.M.N.A., M.P.; Writing – Review and Editing: A.N.M.N.A., M.P., N.M., J.G.C., V.S., M.B.

Supplementary Information

Supplemental data for this article can be accessed online at https://doi.org/10.1080/19420862.2025.2562997

References

  • 1.Kaplon H, Reichert JM.. Antibodies to watch in 2019. Mabs. 2019. Dec. 11(2):0 219–12. doi: 10.1080/19420862.2018.1556465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Saxena V, Panicucci R, Joshi Y, Garad S.. Developability assessment in pharmaceutical industry: an integrated group approach for selecting developable candidates. J Pharm Sci. 2009. June. 98(6):1962–1979 1962–1979, 10.1002/jps.21592. [DOI] [PubMed] [Google Scholar]
  • 3.Ecker DM, Jones SD, Levine HL. The therapeutic monoclonal antibody market. Mabs. 2015;7(1):0 9–14. doi: 10.4161/19420862.2015.989042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Strohl WR. Structure and function of therapeutic antibodies approved by the US FDA in 2023. Antibody Ther. 2024. ISSN 2516-4236. 03. 7(2):0 132–156. doi: 10.1093/abt/tbae007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jain T, Sun T, Durand S, Hall A, Houston NR, Nett JH, Sharkey B, Bobrowicz B, Caffry I, Yu Y. Biophysical properties of the clinical-stage antibody landscape. Proc Natl Acad Sci USA. 2017. Jan. 114(5):0 944–949. doi: 10.1073/pnas.1616408114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Jain T, Boland T, Vásquez M. Identifying developability risks for clinical progression of antibodies using high-throughput in vitro and in silico approaches. mAbs. 2023. 15(1): 2200540. 10.1080/19420862.2023.2200540 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Venkatesh S, Lipper RA. Role of the development scientist in compound lead selection and optimization. J Pharm Sci. 2000. Feb. 890(2):0 145–154. doi: 10.1002/(SICI)1520-6017(200002)89:2<145::AID-JPS2>3.0.CO;2-6. [DOI] [PubMed] [Google Scholar]
  • 8.Kola I, Landis J. Can the pharmaceutical industry reduce attrition rates? Nat Rev Drug Discov. 2004. ISSN 1474-1784. Aug. 3(8):0 711–716. doi: 10.1038/nrd1470. [DOI] [PubMed] [Google Scholar]
  • 9.Sun D, Yu LX, Hussain MA, Wall DA, Smith RL, Amidon GL. In vitro testing of drug absorption for drug ‘developability’ assessment: forming an interface between in vitro preclinical data and clinical outcome. Curr Opin Drug Discov Devel. 2004. Jan. 70(1):0 75–85. [PubMed] [Google Scholar]
  • 10.Serajuddin ATM. Salt formation to improve drug solubility. Adv Drug Delivery Rev. 2007. May. 59(7):0 603–616. doi: 10.1016/j.addr.2007.05.010. [DOI] [PubMed] [Google Scholar]
  • 11.Garad S. How to improve the bioavailability of poorly soluble drugs. Am Pharm Rev. 2004;70(1):0 80–85. [Google Scholar]
  • 12.Mori S, Barth HG. Size exclusion chromatography. Berlin Heidelberg New York: Springer-Verlag; 1999. [Google Scholar]
  • 13.Stetefeld J, McKenna SA, Patel TR. Dynamic light scattering: a practical guide and applications in biomedical sciences. Biophys Rev. 2016;8(4):0 409–427. doi: 10.1007/s12551-016-0218-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Johnson CM. Differential scanning calorimetry as a tool for protein folding and stability. Arch Biochem Biophys. 2013;531(1–2):0 100–109. doi: 10.1016/j.abb.2012.09.008. [DOI] [PubMed] [Google Scholar]
  • 15.Righetti PG. Isoelectric focusing: theory, methodology and application. Amsterdam, New York: Elsevier; 1983. [Google Scholar]
  • 16.D’Atri V, ImioÅ‚ek M, Quinn C, Finny A, Lauber M, Fekete S, Guillarme D. Size exclusion chromatography of biopharmaceutical products: from current practices for proteins to emerging trends for viral vectors, nucleic acids and lipid nanoparticles. J Chromatogr A. 2024. ISSN 0021-9673. 1722:0 464862. doi: 10.1016/j.chroma.2024.464862. [DOI] [PubMed] [Google Scholar]
  • 17.Fekete S, Beck A, Veuthey J-L, Guillarme D. Theory and practice of size exclusion chromatography for the analysis of protein aggregates. J Pharmaceut Biomed. 2014. ISSN 0731-7085. 101:0 161–173. doi: 10.1016/j.jpba.2014.04.011. JPBA Reviews 2014. [DOI] [PubMed] [Google Scholar]
  • 18.Chakrabarti A. Separation of monoclonal antibodies by analytical size exclusion chromatography. In: Böldicke T. editor. Antibody engineering, chapter 7. Rijeka: IntechOpen; 2018. 133–174. doi: 10.5772/intechopen.73321. [DOI] [Google Scholar]
  • 19.Hong P, Koza S, Bouvier ESP. A review size-exclusion chromatography for the analysis of protein biotherapeutics and their aggregates. J Liq Chromatogr Relat Technol. 2012. Nov. 35(20):0 2923–2950. doi: 10.1080/10826076.2012.743724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Balbach S, Korn C. Pharmaceutical evaluation of early development candidates “the 100 mg-approach”. Int J Pharm. 2004. May. 2750(1–2):0 1–12. doi: 10.1016/j.ijpharm.2004.01.034. [DOI] [PubMed] [Google Scholar]
  • 21.Bailly M, Mieczkowski C, Juan V, Metwally E, Tomazela D, Baker J, Uchida M, Kofman E, Raoufi F, Motlagh S. et al. Predicting Antibody Developability Profiles Through Early Stage Discovery Screening. mAbs. 2020;12(1):1743053. 10.1080/19420862.2020.1743053 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Park E, Izadi S. Molecular surface descriptors to predict antibody developability: sensitivity to parameters, structure models, and conformational sampling. mAbs. 2024;16(1):2362788. 10.1080/19420862.2024.2362788 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Rai BK, Apgar JR, Bennett EM. Low-data interpretable deep learning prediction of antibody viscosity using a biophysically meaningful representation. Sci Rep. 2023;13(1):0 2917. doi: 10.1038/s41598-023-28841-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Rollins ZA, Widatalla T, Cheng AC, Metwally E. Abmelt: learning antibody thermostability from molecular dynamics. Biophys J. 2024;123(17):2921–2933. doi: 10.1016/j.bpj.2024.06.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Waight AB, Prihoda D, Shrestha R, Metcalf K, Bailly M, Ancona M, Widatalla T, Rollins Z, Cheng AC, Bitton DA. et al. A machine learning strategy for the identification of key in silico descriptors and prediction models for IgG monoclonal antibody developability properties. mAbs. 2023;15:2248671. 10.1080/19420862.2023.2248671 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Rao R, Meier J, Sercu T, Ovchinnikov S, Rives A. Transformer protein language models are unsupervised structure learners. International Conference on Learning Representations. Vienna, Austria; 2021. [Google Scholar]
  • 27.Villegas-Morcillo A, Makrodimitris S, van Ham RC, Gomez AM, Sanchez V, Reinders MJ, Elofsson A. Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function. Bioinformatics. 2021;37(2):0 162–170. doi: 10.1093/bioinformatics/btaa701. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Wang Z, Combs SA, Brand R, Calvo MR, Xu P, Price G, Golovach N, Salawu EO, Wise CJ, Ponnapalli SP. Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction. Sci Rep. 2022;12(1):0 6832. doi: 10.1038/s41598-022-10775-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Rollins ZA, Widatalla T, Waight A, Cheng AC, Metwally E. Ablef: antibody language ensemble fusion for thermodynamically empowered property predictions. Bioinformatics. 2024. ISSN 1367-4811. 04. 400(5):0 btae268, 10.1093/bioinformatics/btae268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Widatalla T, Rollins Z, Chen M-T, Waight A, Cheng AC. Abprop: language and graph deep learning for antibody property prediction. ICML workshop on computational biology. Hawaiʻi, USA; 2023. [Google Scholar]
  • 31.Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):0 583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Giddings JC. Dynamics of chromatography principles and theory. (NY), (NY); Basel, Switzerland: Marcel Dekker, Inc; 1965. ISBN 9781315275871. 12. doi: 10.1201/9781315275871. [DOI] [Google Scholar]
  • 33.Arakawa T, Timasheff SN. The stabilization of proteins by osmolytes. Biophys J. 1985. Mar. 47(3):0 411–414. doi: 10.1016/S0006-3495(85)83932-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Podwojski K, Fritsch A, Chamrad DC, Paul W, Sitek B, Stühler K, Mutzel P, Stephan C, Meyer HE, Urfer W, et al. Retention time alignment algorithms for LC/MS data must consider non-linear shifts. Bioinformatics. 2009. Jan. 25(6):0 758–764. doi: 10.1093/bioinformatics/btp052. [DOI] [PubMed] [Google Scholar]
  • 35.Berger B, Waterman MS, Yu YW. Levenshtein distance, sequence comparison and biological database search. IEEE Trans Inf Theory. 2021;67(6):0 3287–3294. doi: 10.1109/TIT.2020.2996543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Levenshtein VI. Binary codes capable of correcting deletions, insertions, and reversals. Sov Phys Dokl. 1965;10:0 707–710. [Google Scholar]
  • 37.Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn. 2006;63(1):0 3–42. doi: 10.1007/s10994-006-6226-1. [DOI] [Google Scholar]
  • 38.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, dos Santos Costa A, Fazel-Zarandi M, Sercu T, et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. BioRxiv. 2022. [Google Scholar]
  • 39.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, Smetanin N, Verkuil R, Kabeli O, Shmueli Y, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):0 1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 40.Olsen TH, Moal IH, Deane CM, Lengauer T. Ablang: an antibody language model for completing antibody sequences. Bioinf Adv. 2022;2(1):0 vbac046, 10.1093/bioadv/vbac046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W. LoRA: low-rank adaptation of large language models. International Conference on Learning Representations; 2022. [Google Scholar]
  • 42.Ingraham J, Garg V, Barzilay R, Jaakkola T. Generative models for graph-based protein design. Adv Neural Inf Process Syst. 2019;32. [Google Scholar]
  • 43.Jing B, Eismann S, Suriana P, Townshend RJL, Dror R. Learning from protein structure with geometric vector perceptrons. International Conference on Learning Representations; 2021; Virtual. [Google Scholar]
  • 44.Gasteiger E. Expasy: the proteomics server for in-depth protein knowledge and analysis. Nucleic Acids Res. 2003;31(13):0 3784–3788. doi: 10.1093/nar/gkg563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Jamasb AR, Torné RV, Ma EJ, Du Y, Harris C, Huang K, Hall D, Lio P, Blundell TL. Graphein - a Python library for geometric deep learning and network analysis on biomolecular structures and interaction networks. In: Oh AH, Agarwal A, Belgrave DCho K (Curran Associates, Inc; ). editors. Advances in neural information processing systems; 35; 2022. p. 27153-–27167. [Google Scholar]
  • 46.Nguyen VTD, Hy TS. Multimodal pretraining for unsupervised protein representation learning. Biol Methods protoc. 2024;9(1). doi: 10.1093/biomethods/bpae043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Veličković P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph attention networks. International Conference on Learning Representations; Vancouver, Canada; 2018. [Google Scholar]
  • 48.Xu K, Hu W, Leskovec J, Jegelka S. How powerful are graph neural networks? International Conference on Learning Representations; Vancouver, Canada; 2018. [Google Scholar]
  • 49.Olsen TH, Moal IH, Deane C. Addressing the antibody germline bias and its effect on language models for improved antibody design. BioRxiv. 2024; 2024–02. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Navarin N, Van Tran D, Sperduti A. Universal readout for graph convolutional neural networks. 2019 international joint conference on neural networks (IJCNN); Budapest, Hungary; 2019. IEEE; p. 1–7. [Google Scholar]
  • 51.Abanades B, Wong WK, Boyles F, Georges G, Bujotzek A, Deane CM. Immunebuilder: deep-learning models for predicting the structures of immune proteins. Commun Biol. 2023;6(1):0 575. doi: 10.1038/s42003-023-04927-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Golinski AW, Schmitz ZD, Nielsen GH, Johnson B, Saha D, Appiah S, Hackel BJ, Martiniani S. Predicting and interpreting protein developability via transfer of convolutional sequence representation. ACS Synth Biol. 2023;12(9):0 2600–2615. doi: 10.1021/acssynbio.3c00196. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Hayes T, Rao R, Akin H, Sofroniew NJ, Oktay D, Lin Z, Verkuil R, Tran VQ, Deaton J, Wiggert M, et al. Simulating 500 million years of evolution with a language model. BioRxiv. 2024; 2024–2027. [DOI] [PubMed] [Google Scholar]
  • 54.Malherbe C, Ucar T. IgBLend: unifying 3D structures and sequences in antibody language models. BioRxiv. 2024; doi: 10.1101/2024.10.01.615796. [DOI] [Google Scholar]
  • 55.Sun Y, Shen Y. Structure-informed protein language models are robust predictors for variant effects. Human Genetics. 2024;144:1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

SI_mAbs.pdf

Articles from mAbs are provided here courtesy of Taylor & Francis

RESOURCES