ABSTRACT
To meet the widespread demand for subcutaneous delivery of antibody therapeutics, candidates with low viscosity, high solubility, and/or low aggregation propensity in concentrated formulations must be identified. Moreover, early identification of candidates with low self-association increases the likelihood of success at later stages of the development process. Here, we experimentally profile the self-association behavior of a panel of clinical-stage antibodies as a function of pH, excipient content, and antibody isotype. We find that acidic formulations (pH 5) with proline (200 mM) are most effective at suppressing self-association for both IgG1 and IgG4 variants. Moreover, our self-association measurements are correlated with antibody viscosity measurements and inversely correlated with antibody recovery after their concentration using membrane filters. Notably, we developed interpretable machine learning-based classifier and regressor models for predicting IgG1 and IgG4 self-association and demonstrated that they identify antibodies with favorable high-concentration properties. These findings are expected to improve the antibody development process by facilitating the identification of drug-like molecules during their discovery and optimization.
KEYWORDS: Aggregation, electrostatic, hydrophobic, isoelectric point, isotype, mAb, self-interaction, solubility, viscosity
Introduction
Monoclonal antibodies (mAbs) are widely used for therapeutic applications ranging from cancer to neurodegenerative diseases.1–3 The preferred and most convenient delivery route is subcutaneous delivery, which requires concentrated formulations to deliver efficacious doses.4,5 It is now commonplace for antibodies to be prepared in liquid formulations at >100 mg/mL, and multiple approved antibodies are formulated at higher concentrations. For example, dupilumab is formulated at 150 mg/mL,6 and belimumab is formulated at 200 mg/mL.7
Nevertheless, these concentrated formulations force antibodies together at average separation distances similar to or smaller than their own size, resulting in highly variable and difficult-to-predict levels of self-association.8–10 These self-interactions between closely packed antibodies in concentrated formulations are mediated by attractive charge, hydrophobic, and hydrophilic interactions.11–13 Antibodies with relatively high self-association have been linked to low solubility, high levels of aggregation, and high viscosity.14–20 Moreover, the isotype of antibodies has also been shown to strongly impact their biophysical properties, as IgG4s have lower isoelectric points than IgG1s, and IgG4 antibodies typically have a higher risk of biophysical liabilities, ranging from increased self-association10,14,21 to increased viscosity.18,21 In some cases, the biophysical liabilities of antibodies, such as self-association, can be mitigated using suitable formulations, including the addition of excipients like arginine or proline to disrupt attractive self-interactions.16,22–24
Due to these developability challenges, substantial efforts have been directed toward measuring these weak, reversible self-interactions. At early stages of antibody discovery and optimization, hundreds to thousands of candidates at relatively low concentrations (e.g., <0.1 mg/mL) and intermediate purities (e.g., one-step purified antibody) are generated, and multiple methods for measuring their interactions have been reported by us and others.25–29 For example, we have shown that antibody self-association at 0.01–0.05 mg/mL can be measured in physiological-like solution conditions [phosphate-buffered saline (PBS), pH 7.4] using affinity-capture self-interaction nanoparticle spectroscopy (AC-SINS) 27–29 and in a standard formulation condition (pH 6, 10 mM histidine) using charge-stabilized self-interaction nanoparticle spectroscopy (CS-SINS).29,30
Recent advances in computation and machine learning are being applied to the prediction of antibody self-interactions, which is notable because such predictions may reduce the need for experimental measurements. One important area of computational research is the identification of molecular features extracted from antibody sequences and structures that are most strongly correlated with biophysical properties of interest, such as antibody self-association and viscosity.31–38 Previous studies have shown that these properties are often linked to various types of charge and/or hydrophobicity properties, although these studies are typically based on relatively small datasets. A second important area of computational developability research is the training of models to predict antibody biophysical properties. Substantial progress in this direction has been made,8,33,39–45 although common challenges include a lack of model interpretability, overfitting, and/or difficulty generalizing beyond the training sets.
Here, we utilize CS-SINS to evaluate the self-association of a panel of clinical-stage antibodies, addressing five main questions. First, what are the most important molecular properties that govern antibody self-association in a common formulation condition used for therapeutic antibodies (pH 6, 10 mM histidine)? Second, how much does the antibody isotype impact self-association? Third, to what extent can antibody self-association be mitigated by altering formulation pH and/or adding excipients? Fourth, how accurately can CS-SINS be predicted based on sequence and predicted structural properties? Fifth, to what extent do CS-SINS predictions identify antibodies with favorable high-concentration properties, including low viscosity and high recovery after concentration? Herein, we address these questions by evaluating the self-association of a panel of IgG1 and IgG4 antibodies in different formulation conditions, training and applying machine learning models to predict self-association, and demonstrating that these predictions are correlated with high-concentration antibody properties (Figure 1).
Figure 1.

Overview of approach for predicting antibodies with low self-association and favorable biophysical properties at high concentration. First, antibody self-association is measured using charge-stabilized self-interaction nanoparticle spectroscopy (CS-SINS) at ultra-dilute concentrations (0.01 mg/mL mAb) in different formulation conditions. Next, antibody structural models are generated, and molecular features are extracted and analyzed. These features are also used to generate classifier or regressor models for predicting antibodies with high or low CS-SINS scores. Finally, the antibody self-association predictions were used to identify antibodies with low viscosity or high recovery after concentration.
Results
Evaluation of self-association of clinical-stage antibodies using CS-SINS
We first sought to evaluate the self-association behavior of a panel of clinical-stage antibodies that were reformatted on common IgG1 and IgG4 frameworks. Therefore, we grafted the variable regions of 23 clinical-stage antibodies onto a common human IgG1 framework (Figure S1 and Table S1) and measured their self-association using CS-SINS at 0.01 mg/mL and pH 6 (10 mM histidine; Figure 2A). Sixteen of the clinical-stage antibodies had a CS-SINS score > 0.35, corresponding to high self-association based on previous work.29 The diffusion interaction parameter (kD) was also measured at an antibody concentration of 2–10 mg/mL for 11 of the clinical-stage antibodies (Figure 2B). These results revealed a strong correlation between kD and CS-SINS values (R2 of 0.81), indicating that the two parameters capture similar phenomena. The kD value of a twelfth antibody – idactamab – was determined to be an outlier and omitted from this analysis (Figure S2). Interestingly, idactamab had the highest Fv isoelectric point (pI of 9.41) among the panel of clinical-stage antibodies.
Figure 2.

CS-SINS analysis of IgG1 antibodies at pH 6 (10 mM histidine) and correlation with dynamic light scattering measurements of diffusion interaction parameters. (A) CS-SINS measurements for a panel of IgGs with Fv regions from clinical-stage antibodies and the constant regions from a human IgG1 common framework. (B) Diffusion interaction parameter (kD) values for a subset of the antibodies compared to their CS-SINS measurements. The kD values were measured using dynamic light scattering. In (A), the data are averages, and the errors are standard deviations (n=6, independent experiments). In (B), the kD values are from a single experiment.
Next, we investigated the impact of pH and isotype on the antibody self-association behavior. Therefore, the variable regions of the antibodies were also grafted onto a common IgG4 framework (Figure S1 and Table S1), and CS-SINS was measured at pH 5 and 6 (Figure 3 and S3, Table S2). Most antibodies (22 of 23) were successfully generated as matched IgG1 and IgG4 variants.
Figure 3.

Self-association analysis of a panel of clinical-stage antibodies formatted as IgG1s and IgG4s. (A-B) CS-SINS measurements for (A) IgG1s and (B) IgG4s with common constant regions (regardless of the actual isotype of the clinical-stage antibodies) as a function of pH and the presence of 200 mM proline. The formulations were buffered using 10 mM sodium acetate (pH 5) or 10 mM histidine (pH 6). The results are averages of three to six independent experiments, and the error bars are standard deviations.
CS-SINS analysis of the paired IgG1 and IgG4 antibodies revealed that IgG1s generally exhibit lower self-association than IgG4s at both pH 5 and 6 (Figure S3). At pH 6, 17 IgG1s demonstrated lower self-association than IgG4s, and eight of these results were statistically significant (axatilimab, barecetamab, cabiralizumab, cetuximab, CNTO607, dupilumab, omalizumab, and pepinemab). At pH 5, 15 IgG1s displayed lower self-association than their corresponding IgG4 counterparts, and six of these differences were significant (barecetamab, cabiralizumab, narsoplimab, omalizumab, revdofilimab, and vedolizumab). Interestingly, three antibodies (barecetamab, cabiralizumab, and omalizumab) showed significant differences between the two isotypes at both pH values. Two antibodies exhibited significantly higher CS-SINS scores as an IgG1 than as an IgG4 at pH 5, despite showing the opposite behavior at pH 6 (axatilimab and pepinemab). Idactamab is the only antibody to exhibit significantly higher self-association as an IgG1 relative to IgG4 at both pHs.
It is also notable that 20 of 23 IgG1s have lower self-association at pH 5 than at pH 6 (Figure 3A), and all such differences were statistically significant (Figure S4). Similarly, 22 of 23 IgG4s showed lower self-association at pH 5 than at pH 6, with most of these results (20 of 23) being statistically significant. Notably, three IgG1 antibodies (mavezelimab, CNTO607, and romosozumab) and one IgG4 antibody (CNTO607) exhibited increased CS-SINS scores at pH 5 relative to pH 6. All these antibodies have low Fv isoelectric points (pIs of 4.33 for mavezelimab, 3.94 for CNTO607, and 4.98 for romosozumab), suggesting that both pH and antibody-specific electrostatic properties are important factors influencing antibody self-association.
After evaluating the effects of pH and isotype, we also investigated the impact of proline as an excipient on self-association (Figure 3). CS-SINS scores of clinical-stage antibodies were measured in the presence and absence of proline (200 mM) at pH 5 and 6. The presence of proline generally reduced CS-SINS for antibodies in both the IgG1 (Figure 3A) and IgG4 (Figure 3B) formats at both pH values. The pH 5 formulation with 200 mM proline was most effective at reducing antibody self-association for both isotypes. The lowest CS-SINS scores for 19 of 23 IgG1s and 21 of 23 IgG4s were in the reduced pH formulation with proline. The remaining four IgG1 antibodies had CS-SINS scores at pH 5 (200 mM proline) that were not statistically higher than those of the best-performing formulations. Finally, for two IgG4 antibodies (mavezelimab and CNTO607), the lowest CS-SINS scores were instead observed at pH 6 in the presence of proline. Regardless of isotype or pH, formulations with proline consistently displayed low CS-SINS scores.
Analysis of molecular features and machine learning models for predicting CS-SINS
We next sought to understand which molecular features contribute to antibody self-association using the CS-SINS measurements from this study (N = 23 IgG1 and 23 IgG4) and a previous study (N = 72 IgG1; pH 6, 10 mM histidine).8 Antibody Fv regions were modeled using ABodyBuilder2,46 and molecular features were extracted using Molecular Operating Environment (MOE) software (Figure 4 and Table S3). At pH 6, we found that electrostatic Fv properties, such as apparent charge and isoelectric point, exhibit a strong negative correlation with CS-SINS scores. Features that describe negatively charged surface patches, such as the number of negative patches in the Fv region and CDRs, were positively correlated with the CS-SINS scores. In contrast, features that describe positively charged surface patches exhibited less significant correlations than those that describe negatively charged patches. While electrostatic features showed the strongest correlations with self-association at pH 6, we also observed a significant positive correlation between hydrophobic moment and IgG1 CS-SINS scores. Hydrophobic moment was also positively correlated with CS-SINS scores of IgG4s in the absence and presence of proline, but the correlation was not significant in the latter case.
Figure 4.

Evaluation of antibody molecular features correlated with self-association. Spearman’s ρ correlations between antibody self-association measured by CS-SINS and antibody molecular features are reported for pH 6 in the absence and presence of 200 mM proline. IgG1 correlations are also shown for a combination of the data for the 23 IgG1s in this study and an additional 72 IgG1s reported in a previous study8. Positive correlations are shown in red, and negative correlations are shown in blue. Moreover, the types of features are denoted as hydrophobic (hyd.), positive (pos.), negative (neg.), ensemble dipole moment (ens. dipole moment), structure-based isoelectric point (pI 3D), van der Waals accessible surface area (ASA vdw), hydrophobic accessible surface area (ASA hyd), hydrophilic accessible surface area (ASA hph), hydrophobic moment (hyd. moment), and sequence-based isoelectric point [pI (seq)]. The statistical significance is reported as * (p<0.05), ** (p<0.01), and *** (p<0.001), and the underlined p-value symbols indicate statistical significance after correction for multiple comparisons (Bonferroni correction, n=39 features).
To further understand the impact of formulation conditions on the molecular determinants of self-association, we evaluated correlations between the CS-SINS scores and the same set of molecular features at pH 5 (Figure S5 and Table S3). Similarly to the correlations at pH 6, we found that Fv apparent charge and isoelectric point exhibit strong negative correlations with the CS-SINS scores (i.e., increased apparent charge or isoelectric point both reduce self-association), while negatively charged surface patches were positively correlated. These findings indicate that negative charge is a strong predictor of self-association across a range of weakly acidic pH values.
Using CS-SINS scores and molecular features for the set of clinical-stage antibodies at pH 6, we sought to develop machine learning models to predict self-association, including classifier (Figure 5) and regressor (Figure 6) models. For the decision tree classifier model, which was trained and tested on the set of variable regions for 95 IgG1s (23 from this study and 72 from a previous study8), antibodies with a CS-SINS score > 0.35 were defined to have high self-association (N = 47), while those with a CS-SINS score < 0.35 were defined to have low self-association (N = 48; Figure 5A). The data was split into training and testing sets, with 20% of the data held out as the testing set for model evaluation. The classifier was trained to predict the level of self-association, using leave-one-out cross-validation to select the final model. The classifier had relatively favorable performance metrics on both the training and testing sets (Figure 5B).
Figure 5.

Classification models for predicting the level of IgG1 self-association at pH 6. (A) The decision tree model was trained using 80% of 95 IgGs with CS-SINS measurements at pH 6 (10 mM histidine). Antibodies with high self-association were those with CS-SINS scores >0.35. This resulted in 48 low and 47 high self-association IgGs. IgG1s with predicted low self-association have blue “X” marks in their terminating decisions, while IgG1s with predicted high self-association have red “X” marks in their terminating decisions. (B) The performance metrics for the training set, holdout test set, and all antibodies.
Figure 6.

Regressor models for predicting the level of IgG1 and IgG4 self-association at pH 6. (A-C) The regressor model (support vector regressor, SVR) was trained using CS-SINS measurements at pH 6 (10 mM histidine) for 75 IgG1s and 19 IgG4s, and tested using a holdout set of 20 IgG1s and 4 IgG4s. The model performance relative to the experimental measurements is shown for (A) all IgGs (n=118), (B) IgG1s (n=95), and (C) IgG4s (n=23). (D) SHAP analysis of feature importance. The feature values in magenta indicate high feature values, and positive SHAP values indicate a positive contribution to CS-SINS (i.e., increase self-association), while cyan indicates low feature values, and negative SHAP values indicate a negative contribution to CS-SINS (i.e., reduce self-association). In (A-C), Spearman’s r value and corresponding p-value are reported.
Using the classifier model (Figure 5A), we identified two Fv features for predicting IgG1 self-association at pH 6. First, antibodies were evaluated based on their apparent charge. Of the 47 high CS-SINS antibodies (38:9 training:testing), 39 (31:8) exhibited an apparent charge < 0.52. This is consistent with our finding in Figure 4, which shows that negative charge is linked to high self-association. Next, antibodies were evaluated based on their zeta quadrupole moment. Of the 48 low CS-SINS antibodies, 40 (33:7 training:testing) had zeta quadrupole moments < 1.76 mV in addition to apparent charges > 0.52.
We also sought to develop regression models to predict the self-association levels of a total of 118 IgG1 and IgG4 antibodies at pH 6, including 46 IgG1 and IgG4 antibodies from this study and 72 IgG1 antibodies from a previous study8 (Figure 6). We first divided the dataset into four subgroups based on their CS-SINS scores, which included groups of antibodies with very low (0–0.20), low (0.20–0.35), intermediate (0.35–0.50), and high (>0.50) CS-SINS scores. Next, the data were split using a stratified train-test split, where 20% of the antibodies (N = 24) were held out for model testing. The features used for model training included combinations of 2–4 molecular features obtained from MOE, as well as a feature indicating the antibody isotype. Multiple regression algorithms were tested, and leave-one-out cross-validation was used to select the best-performing model for each algorithm and feature combination. Examples of the regression algorithms and feature combinations tested in this study are summarized in Table S4. The best-performing model was a Support Vector Regressor with a Radial Basis Function kernel, which performed similarly across different train/test splits (Table S5). In addition to the isotype feature, the best-performing regressor used four additional features: structure-based isoelectric point, zeta dipole moment, hydrophobic moment, and the largest ionic patch area. We repeated model training using features extracted from full-length homology models of the IgG1 and IgG4 antibodies, and observed similar, albeit weaker, model performance (Table S6).
We evaluated the correlation between experimentally measured and predicted values of CS-SINS (Figure 6A). A strong correlation was observed between the measured and predicted CS-SINS scores for the set of 118 antibodies, as determined by the Spearman correlation coefficient ( of 0.74). The correlation between the measured and predicted CS-SINS values for 95 IgG1 antibodies was weaker ( of 0.70), although still significant, compared to the correlation for the entire set of 118 antibodies (Figure 6B). We observed a robust correlation between measured and predicted CS-SINS scores of IgG4 antibodies ( of 0.96; Figure 6C). In addition, SHAP47,48 analysis revealed that the structure-based isoelectric point and zeta dipole moment were the most important Fv features that impacted CS-SINS predictions, while the hydrophobic moment was of intermediate importance, and the largest ionic patch area and isotype (IgG1 or IgG4) were the least important (Figure 6D). Overall, these results show that while both electrostatic (isoelectric point and zeta dipole moment) and hydrophobic (hydrophobic moment) features contribute to CS-SINS predictions, the electrostatic contribution is more important.
Antibody self-interactions are correlated with concentrated antibody properties
Next, we sought to evaluate the relationship between CS-SINS and multiple antibody solution properties at elevated concentrations. First, we assessed the relationship between IgG1 and IgG4 self-association and the recovery of antibodies after ultrafiltration (Figure 7). Both isotypes, IgG1 and IgG4, were tested at two initial quantities and concentrations, which were 1 mg (1 mg/mL) and 4 mg (4 mg/mL), to determine the maximum concentration and corresponding percent recovery after concentration. A total of 27 antibodies, including 13 IgG1s and 14 IgG4s, were first tested at the 1 mg scale, and among them, 6 IgG1s and 5 IgG4s were also tested at the 4 mg scale. Percent recovery was determined based on the experimental final volume and theoretical concentration.
Figure 7.

Antibodies with low self-association display high recovery after being concentrated using ultrafiltration. (A-B) IgG1 and IgG4 antibody solutions (pH 6, 10 mM histidine) were concentrated in ultrafiltration concentrators (10 kDa) using a total of 1 mg of antibody to achieve intermediate final concentrations of ~15-60 mg/mL (initial concentration of 1 mg/mL) or 4 mg of antibody to achieve high final concentrations of 100-170 mg/mL (initial concentration of 4 mg/mL), and the results were correlated with (A) experimental or (B) predicted CS-SINS measurements. (C) Spearman’s ρ correlations for % recovery after concentration with experimental and predicted CS-SINS scores, as well as with antibody molecular features for the entire panel of IgG1 and IgG4 antibodies at both intermediate and high antibody concentrations. All molecular features are Fv structural properties, and the Fv isoelectric point is structural. (D) Comparison of CS-SINS classifier model predictions and % recovery after concentration for the entire panel of IgG1 and IgG4 antibodies at intermediate and high concentrations. In (A), (B), and (D), the % recovery values are averages of two independent experiments, and the error bars are standard deviations. In (A), the CS-SINS values are averages of six independent experiments, and the error bars are standard deviations. In (C), the statistical significance is reported as * (p<0.05), *** (p<0.001), and **** (p<0.0001). In (D), the classifier used for the predictions is shown in Fig. 5, and the p-value was calculated using the Anderson-Darling test.
At the 1 mg scale, we observed one IgG1 (idactamab) that was poorly concentrated (15 mg/mL) and recovered (58%), while most IgG1s achieved ~25–32 mg/mL and recoveries of 72–91% (Table S7). For IgG4s, one (axatilimab) failed to concentrate and achieved low recovery (3%), while most IgG4s achieved ~31–48 mg/mL and recoveries of 74–96% (Table S8). At the 4 mg scale, the IgG1s achieved relatively high concentrations (114–145 mg/mL) and a wide range of recoveries (67–94%). For IgG4s at the 4 mg scale, we observed similar high concentrations (114–160 mg/mL) and recoveries (75–95%) as for IgG1s. Of note, bococizumab was the best recovered antibody at the 4 mg scale for the IgG1 format (94.2%), while ralpancizumab was the best recovered antibody in the IgG4 format (94.9% at the 4 mg scale).
We also directly compared the percent recovery after ultrafiltration with the CS-SINS measurements for measurements at both the 1 and 4 mg scales for both isotypes (Figure 7A). This analysis revealed a strong negative correlation between percent recovery and CS-SINS measurements, with a Spearman’s ρ of −0.73 (p-value < 0.0001). The antibody with the lowest CS-SINS score (0.098 for ralpancizumab IgG4) exhibited the highest percent recovery (95.7%). Moreover, axatilimab, which had high CS-SINS scores in both the IgG1 (CS-SINS of 0.88) and IgG4 (CS-SINS of 1.09) formats, had low percent recoveries for both isotypes (71% for IgG1 and 3% for IgG4).
We observed that the percent recovery after ultrafiltration was significantly correlated with the CS-SINS score for each isotype, regardless of the initial antibody concentration (Figure S6). Notably, we found that CS-SINS measurements were much better correlated with the percent recovery (ρ of −0.69 at 1 mg scale and −0.90 at 4 mg scale) than with the maximum antibody concentration (ρ of −0.28 at 1 mg scale and −0.58 at 4 mg scale).
Next, we calculated the correlation between the percent recovery and the predicted CS-SINS values using our regression model (Figure 7B). We observed a slightly weaker, though statistically significant, correlation between predicted CS-SINS and percent recovery (ρ of −0.50, p of 0.0016) relative to that for the measured CS-SINS values (ρ of −0.73, p-value < 0.0001). We also compared the correlation between the percent recovery and various molecular features (Figure 7C). We observed the strongest correlation for the experimentally measured CS-SINS score. Among the individual molecular features analyzed, the structure-based isoelectric point showed the strongest correlation with the percent recovery (ρ of 0.60). Finally, we compared the percent recovery values of antibodies predicted to have high and low self-association using the classifier model reported in Figure 5 (Figure 7D). The percent recovery was statistically higher for antibodies with predicted low CS-SINS (p-value of 0.0014).
We also investigated the relationship between viscosity at high antibody concentrations and the predicted CS-SINS properties (Figure 8). We first compared the CS-SINS predictions with the viscosity of a wild-type anti-PDGF-BB antibody, referred to as AB-001, and variants thereof previously reported49 (Figure 8A-D). These antibody variants have solvent-exposed mutations that alter surface properties, including those that disrupt negatively charged patches to improve viscosity. We used the viscosity data for the 38 antibody variants, which were interpolated or extrapolated in the original study to estimate the viscosities at 100 and 150 mg/mL in a formulation containing 20 mM histidine (pH 5.8). In general, we found that both our classifier and regressor models were able to differentiate between variants with low and high viscosity. At both 100 mg/mL (Figure 8A) and 150 mg/mL (Figure 8C), antibodies classified as having high self-association, based on the model in Figure 5, had significantly higher viscosity than those classified as low self-association (p-values of 7.4 × 10−14 at 100 mg/mL and 1.0 × 10−14 at 150 mg/mL). We observed similar trends when viscosity values were compared to regressor-predicted CS-SINS scores. In particular, we observed Spearman’s ρ values of 0.71 (p-value of 5.0 × 10−7) at 100 mg/mL (Figure 8B) and 0.72 (p-value of 4.4 × 10−7) at 150 mg/mL (Figure 8D). Moreover, the bimodal nature of the viscosity data appears to be linked to the bimodal distribution of the isoelectric points (Fig. S7).
Figure 8.

Antibodies with low predicted self-association display reduced viscosity in concentrated formulations. (A-D) Comparison of CS-SINS model predictions for a panel of mutants of a viscous parental IgG1 and viscosity measurements for antibody solutions (pH 5.8, 20 mM histidine)49 at (A-B) 100 mg/mL for the (A) classifier (cutoff of 20 cP) and (B) regressor models, and at (C-D) 150 mg/mL for the (C) classifier (cutoff of 20 cP) and (D) regressor models. (E-F) Comparison of CS-SINS model predictions and viscosity measurements for a panel of clinical-stage antibody solutions (pH 6, 10 mM histidine) at 150 mg/mL12 for the (E) classifier (21 IgG1 mAbs, cutoff of 30 cP) and (F) regressor models (21 IgG1 mAbs, 3 IgG4 mAbs). In (A), (C), and (E), the classifier used for the predictions is shown in Fig. 5, and the p-values were calculated using the Anderson-Darling test.
Finally, we compared predicted CS-SINS values with viscosity measurements for an additional set of 21 IgG1 and 3 IgG4 US Food and Drug Administration-approved mAbs, evaluated at 150 mg/mL in a buffer containing 10 mM histidine at pH 6 (Figure 8E,F).12 IgG1 antibodies classified as having high self-association based on the model in Figure 5 had a significantly higher viscosity than those classified as having low self-association (p-value of 0.0065; Figure 8E). Moreover, the viscosity values were also correlated with the regressor-predicted CS-SINS scores of 21 IgG1 mAbs and 3 IgG4 mAbs (ρ of 0.46, p-value of 2.3 × 10−2; Figure 8F). Overall, these findings for both ultrafiltration and viscosity reveal our experimental and computational methods are useful for identifying antibodies with different levels of risk for undesirable properties at elevated concentrations.
Discussion
Our results demonstrate that antibody isotype impacts antibody self-association, and these effects vary depending on the specific antibody or formulation condition. In general, we find that IgG1s have lower self-association than IgG4s (Figure S3). This finding is consistent with several previous studies, which have also reported that IgG4s typically exhibit higher self-association,10,14,21 higher opalescence,14,18 lower solubility,18 higher aggregation,10,18,50–53 and higher viscosity 18,21 compared to IgG1s. However, we identified exceptions in which IgG1s can display higher self-association than IgG4s, including axatilimab and pepinemab at pH 5, and idactamab at both pH 5 and 6 (Figure 3). This surprising finding is supported by the improved recovery after concentration of idactamab at pH 6 for the IgG4 isotype (91–94%; Table S8) relative to the IgG1 isotype (58–70%, Table S7). Interestingly, idactamab has the highest apparent Fv charge at pH 6 (+3.1) and pH 5 (+3.6) among the 23 antibodies, and is the only one without negatively charged Fv patches, while axatilimab has one of the largest solvent-exposed hydrophobic surface areas at pH 6 (Table S3). The unexpected results for pepinemab may be due to its unusual combination of a low isoelectric point and high hydrophobic moment.
Our findings that IgG1 and IgG4 antibodies displayed lower self-association at pH 5 relative to pH 6, and that this was not observed for all antibodies, deserve further consideration. First, the observation that reduced pH for formulation-relevant pH values (e.g., pH 5–6) generally reduces self-association – regardless of isotype – is consistent with previous reports that reduced pH reduces self-association,10,20,54–57 increases solubility,55 reduces aggregation,53,58 and reduces viscosity17,19,20,54 for many antibodies. Second, it is interesting that the antibodies with nonstandard pH-dependent self-association (i.e., lower pH increases self-association), namely, the IgG1 versions of mavezelimab, CNTO607, and romosozumab, and the IgG4 version of CNTO607, have relatively low Fv isoelectric points (pIs of 3.94 for CNTO607, 4.33 for mavezelimab, and 4.98 for romosozumab). This type of nonstandard behavior has been reported previously for antibody self-association,19,20 solubility,59 and viscosity.19 In particular, an IgG1 with an acidic isoelectric point was found to display nonstandard pH-dependent self-association and viscosity, as this antibody exhibited increased self-association and viscosity at lower pH values. In contrast, IgG1s in this previous study with basic isoelectric points displayed standard behavior,19 an observation that appears to be generally consistent with our findings.
The most effective formulation for reducing the self-association of both IgG1s and IgG4s tested in this work is pH 5 (10 mM acetate) and 200 mM proline. While proline is a less common amino acid additive in antibody formulations than arginine,60 it is used in the formulations of several approved drugs.61 For example, three notable examples – which are all IgG2 antibodies at similar acidic pH values (pH 4.8–5.2) – are formulated with high concentrations of proline, including brodalumab (208 mM proline, pH 4.8), tezepelumab (218 mM proline, pH 5.2), and evolocumab (217 mM proline, pH 5). Moreover, other antibodies formulated with proline include evinacumab (260 mM proline, pH 6; IgG4) and cemiplimab (130 mM proline, pH 6; IgG4 antibody). Even more interesting is that human IVIG products – including Hizentra and Privigen – are formulated in 250 mM proline (pH 4.6–5.2), highlighting the importance of this acidic formulation with a high concentration of proline.
The beneficial role of proline in antibody formulations has also been observed in multiple other studies. For example, analysis of an IgG1 antibody with a high isoelectric point (pI of 9.3) showed that a formulation with a similar proline concentration (250 mM proline) and pH (pH 6) also reduced antibody viscosity and aggregation relative to a corresponding formulation without proline.62 Likewise, a study of three IgG1 antibodies revealed that proline (100 mM) reduced self-association and viscosity at pH 5.5, while maintaining or even improving resistance to aggregation.16
It is also notable that our CS-SINS measurements, performed at ultra-dilute concentrations (0.01 mg/mL mAb), are correlated with the properties of antibodies when concentrated by orders of magnitude. Our observations that CS-SINS measurements are linked to viscosity measurements (Figure 8) are consistent with several previous reports.29,30,63 However, our observations about the connection between antibody self-association and recovery after ultrafiltration are relatively unique and deserve further consideration. Interestingly, the maximum concentrations achieved during the concentration process were not correlated with self-association measurements, unlike previous AC-SINS measurements for different antibodies, formulation conditions, and concentrators (30 kDa molecular weight cutoff in the previous study vs. 10 kDa in our study).27 In our study, this may suggest that antibodies at high concentrations adhere to the membranes and/or phase separate or aggregate near them, but these processes fail to sufficiently foul the membranes, especially due to the small membrane pores (10 kDa cutoff), resulting in little impact on the maximum obtainable concentrations. The fact that the percentage recovery of antibodies after concentration is inversely correlated with native self-association may suggest that antibodies phase separate into native-like precipitates on or near the membrane that are poorly reversible on the time scales of recovering and analyzing the concentrated proteins (minutes). More work is needed to evaluate both the generality of our ultrafiltration findings and the molecular origins of the differences in antibody recovery.
Our molecular feature analysis of CS-SINS measurements and related predictive models is also unique compared with previous reports and warrants further consideration. First, our previous work suggested that charge-based Fv properties are the most important molecular features governing IgG1 self-association,8 consistent with the findings in this work. However, we also now find – using an expanded dataset – that hydrophobic properties such as the hydrophobic moment are significantly and positively correlated with CS-SINS for IgG1s and IgG4s in the absence of 200 mM proline (Figure 4). While our classifier primarily used electrostatic properties to predict self-association (Figure 5), our regressor model used hydrophobic moment in addition to electrostatic properties and isotype for predicting CS-SINS (Figure 6). Overall, these findings reveal that both electrostatic and hydrophobic interactions mediate self-association measured using CS-SINS.
The expanded CS-SINS dataset used in this work, which contains 23 IgG1 and 23 IgG4 antibodies from this study in addition to 72 IgG1s from a previous study,8 resulted in an IgG1 classifier model with improved overall performance. This included improved balanced accuracy (86% in this study vs. 85% in the previous study), recall (89% vs. 80%), and F1 score (87% vs. 84%), while displaying modestly reduced precision (84% vs. 88%). Moreover, we now report a regressor model that not only predicts CS-SINS for IgG1s, but also for IgG4s. The model predictions, which are correlated with self-association (ρ of 0.74, Figure 6), viscosity (ρ of 0.46–0.72, Figure 8), and recovery after concentration (ρ of −0.50, Figure 7), should be useful for identifying antibodies with increased risk for high self-association and undesirable properties when concentrated, as well as for designing mutations that suppress self-association and associated developability challenges.
In conclusion, our results demonstrate that electrostatic interactions are the most important property to optimize for minimizing the self-association of both IgG1 and IgG4 antibodies prepared in common formulation conditions, thereby reducing viscosity and increasing recovery after concentration. While we observed that IgG4 antibodies typically display higher self-association than IgG1 antibodies, exceptions were observed, and the differences were most pronounced for relatively high self-association antibodies at pH 5. Moreover, we found that pH 5 (10 mM acetate) with 200 mM proline is an effective formulation for suppressing the self-association of most IgG1 and IgG4 antibodies, although notable exceptions were also observed. Finally, we demonstrated that CS-SINS measurements of antibody self-association, which are correlated with viscosity and recovery after ultrafiltration, could be predicted using machine learning models that capture the role of electrostatic and hydrophobic interactions, as well as antibody isotype, opening the door for their future use for improving antibody candidate selection and engineering.
Materials and methods
mAb production and purification
CHO-GS knockout cells were cultured and maintained as previously described.64 Briefly, the CHO cells were maintained in a proprietary DMEM-based medium with 8 mM L-glutamine (LM-Growth) (SAFC, St. Louis, MO) in shake flasks at 37°C and 8% CO2 and passaged every 3–4 days by dilution. The cells were maintained in culture for a minimum of three passages prior to transfection with the mAb expression vectors containing the light chain and the heavy chain genes controlled by the CMV (cytomegalovirus) promoter. Expression vectors were constructed by ATUM (Newark, CA, USA). The stable CHO mAb-expressing pools were generated by transfecting cells using FreeStyle™ Max reagent (Thermo Fisher Scientific, Grand Island, NY) with 95% mAb expression plasmid and 5% helper plasmid containing the piggyBac transposase gene.64 Cells were re-suspended in selection media (DMEM-based medium without L-glutamine) at a concentration of 0.3 × 106 cells/mL. Cultures were passaged twice a week under selection until recovery. Once recovered, the pools were subject to 14-day fed batch production. Cells were seeded at 1 × 106 cells/mL in 1 L volume of proprietary production media in vented shake flasks. Cultures were maintained at 37°C in an incubator with 6% CO2 and 80% humidity, followed by a temperature shift to 32°C on Day 6 post-inoculation. Multiple proprietary feeds were added during the production process. All cultures were harvested after 14 days.
Protein concentration (titer) was measured via analytical Protein A affinity chromatography, as previously described.64 Primary capture of expressed mAbs from conditioned and clarified cell culture media was facilitated by Protein A affinity chromatography on a MabSelect PrismA column on an ÄKTA pure 25 M system (Cytiva, Marlborough, MA). For all mAbs, the resulting Protein A captures (~990 mL) were subsequently pumped over HiLoad 50/60 Superdex 200pg (Cytiva, Marlborough, MA) prep size-exclusion chromatography (SEC) column that was equilibrated with 1x PBS (pH 7.4) at a linear velocity of ~30 cm/h. All fractions from single peak fractionation on the ÄKTA pure 25 M system were collected (45 mL; level of 50 mAU) in 50 mL Falcon tubes using the F9-C fraction collector.65 The percent purity of the SEC products was determined by running analytical SEC analysis on an Agilent 1260 (Agilent Technologies, Wilmington, DE) equipped with either an analytical grade TSKgel UP-SW3000 column (4.6 mm ×15 cm, 2 μm; Tosoh Bioscience LLC, King of Prussia, PA) or Zenix-C SEC 300 (4.6×300 mm; Sepax Technologies, Newark, DE). Final products (3–5 µg typically) were injected at a flow rate of 0.35 mL/min onto the column equilibrated with 1x PBS and 350 mM NaCl (pH 7.0). Purities of captured samples were determined based on the integration of resultant peaks following injection.
Finally, antibody bulk solutions were dialyzed into the final formulation (10-20 mM histidine at pH 6 or PBS at pH 7.4) using Slide-A-Lyzer Dialysis Cassettes (Thermo Fisher Scientific, Waltham, MA, USA). Dialysis with gentle stirring was carried out for 4 h at room temperature, followed by placement of the cassette into the fresh formulation and continued overnight at 4°C. Following dialysis, the antibodies were concentrated to the desired concentration using Amicon Ultra centrifugal concentrators (MilliporeSigma, Burlington, MA, USA). Concentrations were measured on a SoloVPE (C Technologies, Bridgewater, NJ, USA) using the respective antibody extinction coefficient at 280 nm. Final pH values were adjusted using 1 N HCl and/or 1 N NaOH. After preparation, all solutions were sterile filtered using 0.2 µm PES filters from MilliporeSigma.
Expression and purification of calibration antibodies
The six calibration antibodies (human IgG1s) were cloned into mammalian expression vectors and expressed transiently in HEK293 cells using F17 freestyle media. The transfected cells were incubated at 37°C with 5% CO2 for five days, following a 20% yeastolate feeding on the first day. The antibodies were then purified by Protein A chromatography, and the monomer peak was collected by high-performance liquid chromatography using a running buffer containing 200 mM arginine (PBS, pH 7.4). A buffer exchange was performed using desalting columns (Thermo Fisher Scientific, PI-89882), with the final condition in 10 mM histidine (pH 6) buffer.
Preparation of immunogold conjugates
Goat anti-human Fc-specific antibody (Jackson ImmunoResearch Laboratories, 109–005-008) was buffer exchanged twice using Zeba desalting columns (Thermo Fisher Scientific, PI-89882) and diluted to 0.8 mg/mL with buffer #1 (20 mM acetate buffer, pH 4.3), which had been filtered through a sterile 0.2 μm PES bottle-top filter (Fisher Scientific, 09–741-07). The antibody concentration was determined by measuring UV absorbance at 280 nm and using an extinction coefficient of 1.26 mL/(mg*cm). Polylysine (>70,000 MW; Sigma-Aldrich, P1274) was dissolved in water at 5 mg/mL, and then diluted to 2.67 mg/mL using buffer #1. To prepare 200 µL of immunogold conjugates, 97 µL of goat anti-human Fc-specific antibody (0.8 mg/mL) was mixed with 3 µL of polylysine (2.67 mg/mL), resulting in 100 μL of the capture antibody-polylysine solution. Next, 1.2 mL of 20 nm gold nanoparticles (Ted Pella, 15,705) was concentrated to 100 µL by centrifugation at 21,300 xg for 6 min using 1.5 mL centrifuge tubes (USA Scientific, 1615–5500). The supernatant (1100 µL) was removed gently without disturbing the pellet. The pellet was then resuspended in the remaining 100 µL of supernatant via pipetting up and down approximately five times. Next, the 100 µL of resuspended and concentrated gold nanoparticles was added to the 100 µL of capture antibody-polylysine solution, and rapidly mixed by pipetting up and down approximately four times. The immunogold conjugates were then kept at room temperature overnight and used the next day.
CS-SINS measurement and analysis
The first step in the CS-SINS measurements was the calibration process. Two standard antibodies, including a human polyclonal antibody (Jackson ImmunoResearch, 009000003) and the NIST mAb (Sigma-Aldrich, NIST8671), along with a panel of six clinical-stage antibodies (tocilizumab, cetuximab, evolocumab, denosumab, pembrolizumab, and omalizumab) on a common human IgG1 framework, were tested. Antibody samples were evaluated using a 384-well plate (Thermo Fisher Scientific, 12–565-506). Each well was prepared with 5 µL of immunogold conjugates and 45 µL of each human mAb (0.011 mg/mL), resulting in a final mAb concentration of 0.01 mg/mL (pH 6). The solution was pipetted up and down at least three times for rapid mixing. The plate was incubated at room temperature for 4 h, and then plasmon wavelengths (λp) were evaluated using a BioTek Neo2 plate reader, with measurements taken from 450 to 650 nm in 1 nm increments. To identify λp, the maximum point (around 40 points or the maximum absorbance) was fitted using a quadratic equation, and the first derivative was set to zero.
The calibration process was evaluated using two tests. To pass test #1, λp values must be below 534 nm for the human polyclonal antibody and 533 nm for the NIST mAb. For test #2, the CS-SINS scores for the panel of six clinical-stage antibodies were calculated using the following equation: (λp of a given mAb – λp of the mAb with the lowest λp value)/(λp of the mAb with the highest λp – λp of the mAb with the lowest λp value). Tocilizumab was the antibody with the lowest λp value, while omalizumab was the antibody with the highest λp value. A linear fit was applied between historical measurements and the new experimental data, minimizing the sum of ((1-slope) 2 + (intercept) 2) to evaluate the slope, intercept, and R2 values. To pass test #2, the three values (slope, intercept, and R2) needed to be within 10% of their ideal values. In the case in which both tests were passed, the CS-SINS scores for the antibodies were calculated as described for Test #2. The calibration process was performed at pH 6 (10 mM histidine) regardless of the solution conditions for each test experiment.
Dynamic light scattering
The diffusion interaction parameter (kD) was obtained using the DynaPro III (Wyatt Technologies, Santa Barbara, CA, US). Measurements were made at 25°C with a laser wavelength of 830 nm and a 163.5° scattering angle. Triplicate (30 μL) aliquots of each sample were loaded into a 384-well microplate (Aurora P/N ABM210100A). After loading samples, the plate was centrifuged at 1000 xg for 2 min to remove air bubbles. The kD values of samples with antibody concentrations of 2, 4, 6, 8, and 10 mg/mL were measured. At least eight scans of 5 s per scan were accumulated for each sample. The laser power was auto-attenuated during data collection. The method of cumulants was used to obtain the diffusion coefficient from the autocorrelation function.
Evaluation of antibody recovery after concentration
Ultrafiltration was performed using initial quantities of mAbs of either 1 mg (1 mL of 1 mg/mL) or 4 mg (1 mL of 4 mg/mL). Each mAb (500 µL) was loaded into Amicon ultrafiltration tubes (10 kDa MWCO; Merck, UFC#501096) and centrifuged for 40 min (4°C), and this process was repeated one additional time. Finally, the filters and the retentates were transferred into new tubes and subjected to a 10–40 min spin until there was no longer flow-through. The filters were then gently transferred and inverted into pre-weighed new tubes, and the concentrated samples were collected by centrifugation for 2 min at 1,000 xg. To measure the high-concentration samples, the mAbs were diluted sufficiently to ensure accurate measurements. The concentrations were determined by UV absorbance at 280 nm. The percent of recovery was calculated by dividing the final concentration by the theoretical maximum concentration, where the theoretical maximum concentration was determined by dividing the initial antibody mass (i.e., 1 or 4 mg) by the final volume of the retentate, as inferred from the mass of the retentate assuming the density of water.
Molecular feature analysis and model predictions
Five Fv structural models per antibody were generated using ABodyBuilder266. Partial charges were added for each structure using Molecular Operating Environment (MOE). All structures were titrated to both pH 5 and pH 6 using Protonate3D66 and energy minimized in MOE. Finally, protein properties were calculated and averaged for the five structural models per antibody, including at different pH values, for the feature analysis. The Spearman correlation coefficients between different features and CS-SINS measurements and their p-values were calculated using SciPy67,68 (version 1.15) package in Python (version 3.10).
The machine learning classifier was developed using the DecisionTreeClassifier scikit-learn69 package (version 1.2.2) in Python (version 3.10). The classifier models were trained to predict self-association using structure-based features (pH 6) of IgG1 antibodies extracted from MOE. The training data included 80% of the antibodies (76 IgG1s), and the test (hold-out) set included 20% of the IgG1s (19 IgG1s). Gini impurity was optimized during model training. Leave-one-out cross-validation was implemented via the GridSearchCV algorithm for hyperparameter tuning to select the best model with the maximum average validation accuracy.
The regression models were developed using scikit-learn (version 1.5.2) 69 in Python (version 3.9). The dataset used for the regression models included 95 IgG1 antibodies and 23 IgG4 antibodies. The antibodies were split into groups based on their experimentally determined CS-SINS score. The data was split into training and testing sets using the scikit-learn train_test_split function with 20% of the dataset held out for model testing, and stratification was used on the CS-SINS group to maintain balanced distribution in both sets. Combinations of 2–4 structure-based features calculated from MOE (pH 6) were used to train regression models using SupportVectorRegressor and KNeighborsRegressor. Features were selected such that each molecular property was correlated with experimentally determined CS-SINS values (Spearman’s ρ >0.2) and not strongly correlated with each other (Spearman’s ρ <0.6).
An additional feature was included to describe the IgG isotype, where 0 indicated an IgG1 isotype, and 1 indicated an IgG4 isotype. Additionally, DecisionTreeRegressor and RandomForestRegressor models were trained with all molecular features and the IgG isotype feature. For all model architectures and feature combinations, GridSearchCV was used for hyperparameter tuning with leave-one-out cross-validation to select the model with the minimum average root mean squared error. The model presented here had the lowest error among all architectures and feature combinations tested in the leave-one-out cross-validation analysis.
To evaluate the generalizability of the regression model, an additional four train-test splits were evaluated by changing the random seed used in the scikit-learn train_test_split function.69 Regression models were trained using the training set with the best-performing algorithm and feature set. The regression models were then evaluated using their corresponding test sets.
Full-length IgG1 and IgG4 antibody models were generated using the Antibody Modeler in MOE. Two antibody structures from the Protein Data Bank (1HZH and 5DK3) were used to model the Fc regions of IgG1 and IgG4, respectively. One antibody (crizanlizumab) was removed from the dataset because MOE automatically assigned it an IgG2 Fc region despite our attempts to model it as an IgG1. Partial charges were added to the structures, which were titrated to pH 6 using Protonate3D.66 The structures were then energy minimized. Finally, protein properties of the full IgG1 and IgG4 models were calculated at pH 6. Regression models were trained to predict CS-SINS using the full-length IgG protein properties. The same feature selection and training methodologies were employed as those used for the regression models based on Fv protein properties and the binary isotype feature.
Supplementary Material
Acknowledgments
The authors gratefully acknowledge critical support in the expression, purification, and characterization of the antibodies in this study from Anthony Ransdell, Robert Peery, John Herrington, and Melora Reed of Eli Lilly and Company and Kelley Graybeal and Linnea Schlerer of Eurofins. This work was supported by Eli Lilly and the Albert M. Mattocks Chair (to P.M.T).
P.T., B.J., and W.W. conceived and directed this project. N.K. and S.C. performed the biophysical experiments. R.K. directed the protein production. C.B. and H.C. performed computational analysis and developed the reported models. P.T., N.K., and C.B. wrote the manuscript with feedback from co-authors.
Funding Statement
The work was supported by the Eli Lilly and Company.
Abbreviations
- mAb
Monoclonal antibody
- IgG
Immunoglobulin G
- IgG1
Immunoglobulin G1
- IgG4
Immunoglobulin G4
- ML
Machine learning
- DLS
Dynamic light scattering
- kD
Diffusion interaction parameter
- SEC
Size-exclusion chromatography
- pI
Isoelectric point
- CS-SINS
Charge-stabilized self-interaction nanoparticle spectroscopy
- AC-SINS
Affinity-capture self-interaction nanoparticle spectroscopy
- SVR
Support vector regressor
- SVM
Support vector machine
- RF
Random forest
- MOE
Molecular Operating Environment
- R2
Coefficient of determination
- RMSE
Root mean square error
- AUC
Area under the curve
- ROC
Receiver operating characteristic
- CDR
Complementarity-determining region
- Fv
Fragment variable
- VH
Variable heavy chain
- VL
Variable light chain
- Fc
Fragment crystallizable
- Fab
Fragment antigen-binding
Disclosure statement
P.M.T. is a member of the scientific advisory boards or serves as a scientific advisor for Nabla Bio, Aureka Biotechnologies, Dualitas Therapeutics, CelineBio, and Metaphore Biotechnologies.
Data and code availability
The training data and code used in this manuscript are available in the Tessier lab GitHub repository at https://github.com/Tessier-Lab-UMich/cs-sins_models.
Supplementary material
Supplemental data for this article can be accessed online at https://doi.org/10.1080/19420862.2026.2663641.
References
- 1.Crescioli S, Kaplon H, Wang L, Visweswaraiah J, Kapoor V, Reichert JM.. Antibodies to watch in 2025. mAbs. 2025;17(1):2443538. doi: 10.1080/19420862.2024.2443538. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Buss NA, Henderson SJ, McFarlane M, Shenton JM, de Haan L.. Monoclonal antibody therapeutics: history and future. Curr Opin Pharmacol. 2012;12(5):615–21. doi: 10.1016/j.coph.2012.08.001. [DOI] [PubMed] [Google Scholar]
- 3.Carter PJ, Rajpal A. Designing antibodies as therapeutics. Cell. 2022;185(15):2789–2805. doi: 10.1016/j.cell.2022.05.029. [DOI] [PubMed] [Google Scholar]
- 4.Badkar AV, Gandhi RB, Davis SP, LaBarre MJ. Subcutaneous delivery of high-dose/volume biologics: current status and prospect for future advancements. Drug Des Devel Ther. 2021;15:159–170. doi: 10.2147/DDDT.S287323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Ren S. Current and emerging strategies for subcutaneous delivery of high-concentration and high-dose antibody therapeutics. J Pharm Sci. 2025;114(8):103877. doi: 10.1016/j.xphs.2025.103877. [DOI] [PubMed] [Google Scholar]
- 6.Li Z, Radin A, Li M, Hamilton JD, Kajiwara M, Davis JD, Takahashi Y, Hasegawa S, Ming JE, DiCioccio AT, et al. Pharmacokinetics, pharmacodynamics, safety, and tolerability of dupilumab in healthy adult subjects. Clin Pharmacol Drug Dev. 2020;9(6):742–755. doi: 10.1002/cpdd.798. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Cai WW, Fiscella M, Chen C, Zhong ZJ, Freimuth WW, Subich DC. Bioavailability, pharmacokinetics, and safety of belimumab administered subcutaneously in healthy subjects. Clin Pharmacol Drug Dev. 2013;2(4):349–357. doi: 10.1002/cpdd.54. [DOI] [PubMed] [Google Scholar]
- 8.Makowski EK, Wang T, Zupancic JM, Huang J, Wu L, Schardt JS, De Groot AS, Elkins SL, Martin WD, Tessier PM. Optimization of therapeutic antibodies for reduced self-association and non-specific binding via interpretable machine learning. Nat Biomed Eng. 2023;8(1):45–56. doi: 10.1038/s41551-023-01074-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Jain T, Sun T, Durand S, Hall A, Houston NR, Nett JH, Sharkey B, Bobrowicz B, Caffry I, Yu Y, et al. Biophysical properties of the clinical-stage antibody landscape. Proc Natl Acad Sci. 2017;114(5):944–949. doi: 10.1073/pnas.1616408114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Bailly M, Mieczkowski C, Juan V, Metwally E, Tomazela D, Baker J, Uchida M, Kofman E, Raoufi F, Motlagh S, et al. Predicting antibody developability profiles through early stage discovery screening. mAbs. 2020;12(1):1743053. doi: 10.1080/19420862.2020.1743053. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Cruz MA, Blanco M, Ekladious I. Mechanistic and predictive formulation development for viscosity mitigation of high-concentration biotherapeutics. mAbs. 2025;17(1):2550757. doi: 10.1080/19420862.2025.2550757. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lai P-K, Fernando A, Cloutier TK, Gokarn Y, Zhang J, Schwenger W, Chari R, Calero-Rubio C, Trout BL. Machine learning applied to determine the molecular descriptors responsible for the viscosity behavior of concentrated therapeutic antibodies. Mol Pharm. 2021;18(3):1167–1175. doi: 10.1021/acs.molpharmaceut.0c01073. [DOI] [PubMed] [Google Scholar]
- 13.Makowski EK, Wu L, Gupta P, Tessier PM. Discovery-stage identification of drug-like antibodies using emerging experimental and computational methods. mAbs. 2021;13(1):1895540. doi: 10.1080/19420862.2021.1895540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Kingsbury JS, Saini A, Auclair SM, Fu L, Lantz MM, Halloran KT, Calero-Rubio C, Schwenger W, Airiau CY, Zhang J, et al. A single molecular descriptor to predict solution behavior of therapeutic antibodies. Sci Adv. 2020;6(32):eabb0372. doi: 10.1126/sciadv.abb0372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Yadav S, Shire SJ, Kalonia DS. Factors affecting the viscosity in high concentration solutions of different monoclonal antibodies. J Pharm Sci. 2010;99(12):4812–4829. doi: 10.1002/jps.22190. [DOI] [PubMed] [Google Scholar]
- 16.Cloutier TK, Sudrik C, Mody N, Hasige SA, Trout BL. Molecular computations of preferential interactions of proline, arginine.HCl, and NaCl with IgG1 antibodies and their impact on aggregation and viscosity. mAbs. 2020;12(1):1816312. doi: 10.1080/19420862.2020.1816312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Liu J, Nguyen MDH, Andya JD, Shire SJ. Reversible self-association increases the viscosity of a concentrated monoclonal antibody in aqueous solution. J Pharm Sci. 2005;94(9):1928–1940. doi: 10.1002/jps.20347. [DOI] [PubMed] [Google Scholar]
- 18.Cain P, Huang L, Tang Y, Anguiano V, Feng Y. Impact of IgG subclass on monoclonal antibody developability. mAbs. 2023;15(1):2191302. doi: 10.1080/19420862.2023.2191302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Saito S, Hasegawa J, Kobayashi N, Kishi N, Uchiyama S, Fukui K. Behavior of monoclonal antibodies: relation between the second virial coefficient (B2) at low concentrations and aggregation propensity and viscosity at high concentrations. Pharm Res. 2012;29(2):397–410. doi: 10.1007/s11095-011-0563-x. [DOI] [PubMed] [Google Scholar]
- 20.Connolly BD, Petry C, Yadav S, Demeule B, Ciaccio N, Moore JMR, Shire SJ, Gokarn YR. Weak interactions govern the viscosity of concentrated antibody solutions: high-throughput analysis using the diffusion interaction parameter. Biophys J. 2012;103(1):69–78. doi: 10.1016/j.bpj.2012.04.047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Lai P-K, Ghag G, Yu Y, Juan V, Fayadat-Dilman L, Trout BL. Differences in human IgG1 and IgG4 S228P monoclonal antibodies viscosity and self-interactions: experimental assessment and computational predictions of domain interactions. mAbs. 2021;13(1):1991256. doi: 10.1080/19420862.2021.1991256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Xu AY, Castellanos MM, Mattison K, Krueger S, Curtis JE. Studying excipient modulated physical stability and viscosity of monoclonal antibody formulations using small-angle scattering. Mol Pharm. 2019;16(10):4319–4338. doi: 10.1021/acs.molpharmaceut.9b00687. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Prašnikar M, Žiberna MB, Ahlin Grabnar P. Targeting intermolecular interactions to reduce viscosity in monoclonal antibody formulations: a review. Int J Biol Macromol. 2025;327(Pt 2):147515. doi: 10.1016/j.ijbiomac.2025.147515. [DOI] [PubMed] [Google Scholar]
- 24.Xin L, Lan L, Mellal M, McChesney N, Vaughan R, Berdugo C, Li Y, Zhang J. Leveraging high-throughput analytics and automation to rapidly develop high-concentration mAb formulations: integrated excipient compatibility and viscosity screening. Antib Ther. 2024;7(4):335–350. doi: 10.1093/abt/tbae028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Liu Y, Caffry I, Wu J, Geng SB, Jain T, Sun T, Reid F, Cao Y, Estep P, Yu Y, et al. High-throughput screening for developability during early-stage antibody discovery using self-interaction nanoparticle spectroscopy. mAbs. 2014;6(2):483–492. doi: 10.4161/mabs.27431. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Phan S, Walmer A, Shaw EW, Chai Q. High-throughput profiling of antibody self-association in multiple formulation conditions by PEG stabilized self-interaction nanoparticle spectroscopy. mAbs. 2022;14(1):2094750. doi: 10.1080/19420862.2022.2094750. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Wu J, Schultz JS, Weldon CL, Sule SV, Chai Q, Geng SB, Dickinson CD, Tessier PM. Discovery of highly soluble antibodies prior to purification using affinity-capture self-interaction nanoparticle spectroscopy. Protein Eng Des Sel. 2015;28(10):403–414. doi: 10.1093/protein/gzv045. [DOI] [PubMed] [Google Scholar]
- 28.Sule SV, Dickinson CD, Lu J, Chow C-K, Tessier PM. Rapid analysis of antibody self-association in complex mixtures using immunogold conjugates. Mol Pharm. 2013;10(4):1322–1331. doi: 10.1021/mp300524x. [DOI] [PubMed] [Google Scholar]
- 29.Starr CG, Makowski EK, Wu L, Berg B, Kingsbury JS, Gokarn YR, Tessier PM. Ultradilute measurements of self-association for the identification of antibodies with favorable high-concentration solution properties. Mol Pharm. 2021;18(7):2744–2753. doi: 10.1021/acs.molpharmaceut.1c00280. [DOI] [PubMed] [Google Scholar]
- 30.Makowski EK, Chen H, Lambert M, Bennett EM, Eschmann NS, Zhang Y, Zupancic JM, Desai AA, Smith MD, Lou W, et al. Reduction of therapeutic antibody self-association using yeast-display selections and machine learning. mAbs. 2022;14(1):2146629. doi: 10.1080/19420862.2022.2146629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Agrawal NJ, Helk B, Kumar S, Mody N, Sathish HA, Samra HS, Buck PM, Li L, Trout BL. Computational tool for the early screening of monoclonal antibodies for their viscosities. mAbs. 2016;8(1):43–48. doi: 10.1080/19420862.2015.1099773. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Tomar DS, Li L, Broulidakis MP, Luksha NG, Burns CT, Singh SK, Kumar S. In-silico prediction of concentration-dependent viscosity curves for monoclonal antibody solutions. mAbs. 2017;9(3):476–489. doi: 10.1080/19420862.2017.1285479. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Waight AB, Prihoda D, Shrestha R, Metcalf K, Bailly M, Ancona M, Widatalla T, Rollins Z, Cheng AC, Bitton DA, et al. A machine learning strategy for the identification of key in silico descriptors and prediction models for IgG monoclonal antibody developability properties. mAbs. 2023;15(1):2248671. doi: 10.1080/19420862.2023.2248671. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Bashour H, Smorodina E, Pariset M, Zhong J, Akbar R, Chernigovskaya M, Lê Quý K, Snapkow I, Rawat P, Krawczyk K, et al. Biophysical cartography of the native and human-engineered antibody landscapes quantifies the plasticity of antibody developability. Commun Biol. 2024;7(1):922. doi: 10.1038/s42003-024-06561-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Yadav S, Laue TM, Kalonia DS, Singh SN, Shire SJ. The influence of charge distribution on self-association and viscosity behavior of monoclonal antibody solutions. Mol Pharm. 2012;9(4):791–802. doi: 10.1021/mp200566k. [DOI] [PubMed] [Google Scholar]
- 36.Park E, Izadi S. Molecular surface descriptors to predict antibody developability: sensitivity to parameters, structure models, and conformational sampling. mAbs. 2024;16(1):2362788. doi: 10.1080/19420862.2024.2362788. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Chennamsetty N, Voynov V, Kayser V, Helk B, Trout BL. Design of therapeutic proteins with enhanced stability. Proc Natl Acad Sci USA. 2009;106(29):11937–11942. doi: 10.1073/pnas.0904191106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Hebditch M, Roche A, Curtis RA, Warwicker J. Models for antibody behavior in hydrophobic interaction chromatography and in self-association. J Pharm Sci. 2019;108(4):1434–1441. doi: 10.1016/j.xphs.2018.11.035. [DOI] [PubMed] [Google Scholar]
- 39.Tomar DS, Singh SK, Li L, Broulidakis MP, Kumar S. In silico prediction of diffusion interaction parameter (kD), a key indicator of antibody solution behaviors. Pharm Res. 2018;35(10):193. doi: 10.1007/s11095-018-2466-6. [DOI] [PubMed] [Google Scholar]
- 40.Li B, Luo S, Wang W, Xu J, Liu D, Shameem M, Mattila J, Franklin MC, Hawkins PG, Atwal GS. ProperMab: an integrative framework for in silico prediction of antibody developability using machine learning. mAbs. 2025;17(1):2474521. doi: 10.1080/19420862.2025.2474521. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Makowski EK, Chen H-T, Wang T, Wu L, Huang J, Mock M, Underhill P, Pelegri-O’Day E, Maglalang E, Winters D, et al. Reduction of monoclonal antibody viscosity using interpretable machine learning. mAbs. 2024;16(1):2303781. doi: 10.1080/19420862.2024.2303781. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Rai BK, Apgar JR, Bennett EM. Low-data interpretable deep learning prediction of antibody viscosity using a biophysically meaningful representation. Sci Rep. 2023;13(1):2917. doi: 10.1038/s41598-023-28841-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Kalejaye LA, Chu J-M, Wu I-E, Amofah B, Lee A, Hutchinson M, Chakiath C, Dippel A, Kaplan G, Damschroder M, et al. Accelerating high-concentration monoclonal antibody development with large-scale viscosity data and ensemble deep learning. mAbs. 2025;17(1):2483944. doi: 10.1080/19420862.2025.2483944. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Wu I-E, Kalejaye L, Lai P-K. Machine learning models for predicting monoclonal antibody biophysical properties from molecular dynamics simulations and deep learning-based surface descriptors. Mol Pharm. 2025;22(1):142–153. doi: 10.1021/acs.molpharmaceut.4c00804. [DOI] [PubMed] [Google Scholar]
- 45.Anapindi KDB, Liu K, Wang W, Yu Y, He Y, Hsieh EJ, Huang Y, Tomazela D. Leveraging multi-modal feature learning for predictions of antibody viscosity. mAbs. 2025;17(1):2490788. doi: 10.1080/19420862.2025.2490788. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Abanades B, Wong WK, Boyles F, Georges G, Bujotzek A, Deane CM. Immunebuilder: deep-learning models for predicting the structures of immune proteins. Commun Biol. 2023;6(1):1–8. doi: 10.1038/s42003-023-04927-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Lundberg SM, Erion G, Chen H, DeGrave A, Prutkin JM, Nair B, Katz R, Himmelfarb J, Bansal N, Lee S-I. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell. 2020;2(1):56–67. doi: 10.1038/s42256-019-0138-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Lundberg S, Lee S-I. A unified approach to interpreting model predictions. arXiv. 2017. Nov 25. doi: 10.48550/arXiv.1705.07874. [DOI] [Google Scholar]
- 49.Apgar JR, Tam ASP, Sorm R, Moesta S, King AC, Yang H, Kelleher K, Murphy D, D’Antona AM, Yan G, et al. Modeling and mitigation of high-concentration antibody viscosity through structure-based computer-aided protein design. PLOS ONE. 2020;15(5):e0232713. doi: 10.1371/journal.pone.0232713. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Skamris T, Tian X, Thorolfsson M, Karkov HS, Rasmussen HB, Langkilde AE, Vestergaard B. Monoclonal antibodies follow distinct aggregation pathways during production-relevant acidic incubation and neutralization. Pharm Res. 2016;33(3):716–728. doi: 10.1007/s11095-015-1821-0. [DOI] [PubMed] [Google Scholar]
- 51.Neergaard MS, Nielsen AD, Parshad H, De Weert MV. Stability of monoclonal antibodies at high-concentration: head-to-head comparison of the IgG1 and IgG4 subclass. J Pharm Sci. 2014;103(1):115–127. doi: 10.1002/jps.23788. [DOI] [PubMed] [Google Scholar]
- 52.Ito T, Tsumoto K. Effects of subclass change on the structural stability of chimeric, humanized, and human antibodies under thermal stress. Protein Sci. 2013;22(11):1542–1551. doi: 10.1002/pro.2340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Tian X, Langkilde AE, Thorolfsson M, Rasmussen HB, Vestergaard B. Small-angle x-ray scattering screening complements conventional biophysical analysis: comparative structural and biophysical analysis of monoclonal antibodies IgG1, IgG2, and IgG4. J Pharm Sci. 2014;103(6):1701–1710. doi: 10.1002/jps.23964. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Chari R, Jerath K, Badkar AV, Kalonia DS. Long- and short-range electrostatic interactions affect the rheology of highly concentrated antibody solutions. Pharm Res. 2009;26(12):2607–2618. doi: 10.1007/s11095-009-9975-2. [DOI] [PubMed] [Google Scholar]
- 55.Pindrus M, Shire SJ, Kelley RF, Demeule B, Wong R, Xu Y, Yadav S. Solubility challenges in high concentration monoclonal antibody formulations: relationship with amino acid sequence and intermolecular interactions. Mol Pharm. 2015;12(11):3896–3907. doi: 10.1021/acs.molpharmaceut.5b00336. [DOI] [PubMed] [Google Scholar]
- 56.Sahin E, Grillo AO, Perkins MD, Roberts CJ. Comparative effects of pH and ionic strength on protein–protein interactions, unfolding, and aggregation for IgG1 antibodies. J Pharm Sci. 2010;99(12):4830–4848. doi: 10.1002/jps.22198. [DOI] [PubMed] [Google Scholar]
- 57.Kalonia C, Toprani V, Toth R, Wahome N, Gabel I, Middaugh CR, Volkin DB. Effects of protein conformation, apparent solubility, and protein–protein interactions on the rates and mechanisms of aggregation for an IgG1 monoclonal antibody. J Phys Chem B. 2016;120(29):7062–7075. doi: 10.1021/acs.jpcb.6b03878. [DOI] [PubMed] [Google Scholar]
- 58.Zheng JY, Janis LJ. Influence of pH, buffer species, and storage temperature on physicochemical stability of a humanized monoclonal antibody LA298. Int J Pharm. 2006;308(1):46–51. doi: 10.1016/j.ijpharm.2005.10.024. [DOI] [PubMed] [Google Scholar]
- 59.Casaz P, Boucher E, Wollacott R, Pierce BG, Rivera R, Sedic M, Ozturk S, Thomas WD Jr, Wang Y. Resolving self-association of a therapeutic antibody by formulation optimization and molecular approaches. mAbs. 2014;6(6):1533–1539. doi: 10.4161/19420862.2014.975658. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Ren S. Effects of arginine in therapeutic protein formulations: a decade review and perspectives. Antib Ther. 2023;6(4):265–276. doi: 10.1093/abt/tbad022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Ghosh I, Gutka H, Krause ME, Clemens R, Kashi RS. A systematic review of commercial high concentration antibody drug products approved in the US: formulation composition, dosage form design and primary packaging considerations. mAbs. 2023;15(1):2205540. doi: 10.1080/19420862.2023.2205540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Hung JJ, Dear BJ, Dinin AK, Borwankar AU, Mehta SK, Truskett TT, Johnston KP. Improving viscosity and stability of a highly concentrated monoclonal antibody solution with concentrated proline. Pharm Res. 2018;35(7):133. doi: 10.1007/s11095-018-2398-1. [DOI] [PubMed] [Google Scholar]
- 63.Makowski EK, Kinnunen PC, Huang J, Wu L, Smith MD, Wang T, Desai AA, Streu CN, Zhang Y, Zupancic JM, et al. Co-optimization of therapeutic antibody affinity and specificity using machine learning models that generalize to novel mutational space. Nat Commun. 2022;13(1):3788. doi: 10.1038/s41467-022-31457-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Rajendra Y, Balasubramanian S, Peery RB, Swartling JR, McCracken NA, Norris DL, Frye CC, Barnard GC. Bioreactor scale up and protein product quality characterization of piggyBac transposon derived CHO pools. Biotechnol Prog. 2017;33(2):534–540. doi: 10.1002/btpr.2447. [DOI] [PubMed] [Google Scholar]
- 65.Ransdell AS, Reed M, Herrington J, Cain P, Kelly RM. Creation of a versatile automated two-step purification system with increased throughput capacity for preclinical mAb material generation. Protein Expr Purif. 2023;207:106269. doi: 10.1016/j.pep.2023.106269. [DOI] [PubMed] [Google Scholar]
- 66.Labute P. Protonate3D: assignment of ionization states and hydrogen coordinates to macromolecular structures. Proteins. 2009;75(1):187–205. doi: 10.1002/prot.22234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, Burovski E, Peterson P, Weckesser W, Bright J, et al. Author correction: SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):352–352. doi: 10.1038/s41592-020-0772-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, Burovski E, Peterson P, Weckesser W, Bright J, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261–272. doi: 10.1038/s41592-019-0686-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–2830. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The training data and code used in this manuscript are available in the Tessier lab GitHub repository at https://github.com/Tessier-Lab-UMich/cs-sins_models.
