Skip to main content
BMC Bioinformatics logoLink to BMC Bioinformatics
. 2026 Feb 25;27:73. doi: 10.1186/s12859-026-06401-7

Enhancing CAR-T cell activity prediction via fine-tuning protein language models with generated CAR sequences

Kei Yoshida 1,✉, Shoji Hisada 1, Ryoichi Takase 2, Atsushi Okuma 1, Yoshihito Ishida 1, Taketo Kawara 1, Takuya Miura-Yamashita 1, Daisuke Ito 1
PMCID: PMC13040910  PMID: 41742038

Abstract

Background

Chimeric antigen receptor (CAR)-T cell therapy has shown remarkable success in treating hematological malignancies. However, several challenges remain, including limited efficacy against solid tumors, T cell exhaustion, and lack of T cell persistence, which have restricted its clinical efficacy across various indications. Sequence optimization of CAR constructs offers a promising strategy for enhancing the therapeutic efficacy of CAR-T cells. Recent advances in machine learning, particularly in protein language models (PLMs), have enabled the prediction of mutational effects based on sequence representations. However, applying PLMs to CARs is challenging because of the artificial nature of CARs and the absence of comprehensive CAR sequence databases.

Results

We developed a computational framework for predicting CAR-T cell activity by fine-tuning ESM-2 with CAR sequences generated using sequence augmentation. The CAR sequences were constructed through the in silico recombination of the homologous domains of the CARs, enabling a task-specific adaptation of the model. To evaluate the prediction performance, we experimentally assessed the cytotoxicity of CAR-T cells expressing mutated CAR variants and compared the results with model predictions. Our results demonstrated that fine-tuning ESM-2 significantly improved the prediction performance of CAR-T cell activity. Furthermore, we showed that training parameters such as sequence diversity, number of training steps, and model size substantially influenced prediction performance.

Conclusions

Our findings highlight the potential of combining sequence augmentation with fine-tuning of PLMs to advance data-driven CAR-T cell design.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12859-026-06401-7.

Keywords: CAR-T cell, Protein language models, Fine-tuning, Data augmentation, Sequence-based prediction

Background

Chimeric antigen receptor (CAR)-T cell therapy holds great promise in revolutionizing cancer treatment. CARs are synthetic proteins that combine distinct functional domains to recognize and eliminate target cells [1]. To date, seven CAR-T cell therapies have received FDA approval for hematological malignancies [2, 3]. However, CAR-T cell therapy has shown limited efficacy against solid tumors [4, 5], partly due to its reduced persistence and attenuated antitumor effects [6, 7]. Addressing these challenges is essential for enhancing the therapeutic efficacy of CAR-T cells and expanding their applications in cancer treatment.

Among the various readouts of CAR-T cell function, cytotoxicity is regarded as a critical and direct indicator of the ability to rapidly and efficiently eliminate target cells. Cytotoxicity reflects the short-term activity of CAR-T cell, and a trade-off exists between short-term activity and in vivo persistence [8, 9]. In this context, the two approved CD19 CARs, axicabtagene ciloleucel (axi-cel) and tisagenlecleucel (tisa-cel), have been frequently compared. Tisa-cel elicits weaker signals upon CD19 antigen detection but shows better in vivo persistence [10, 11]. Conversely, axi-cel is potently activated by the antigen, leading to greater target cell elimination; however, its presence in patients is comparatively limited [12, 13]. A comparative study has shown that axi-cel exhibited higher efficacy for diffuse large B-cell lymphoma (DLBCL) than tisa-cel [14]. Therefore, cytotoxicity could serve as a key indicator for the clinical efficacy of CAR-T cell therapies.

Sequence optimization of CAR constructs represents a fundamental approach for improving the therapeutic performance of CAR-T cells [15]. One approach involves the introduction of point mutations into the CAR constructs, which can considerably affect their function and efficacy. For instance, mutations in both the single-chain fragment variable (scFv) domain, responsible for modulating antigen binding affinity, and the intracellular domains, which control CAR-T cell activation and cytotoxicity, can enhance antitumor activity [16–19]. Therefore, designing CAR sequences with the desired therapeutic outcomes requires the systematic introduction of diverse mutations across the CAR domains and a comprehensive evaluation of their effects.

Machine learning (ML) techniques are increasingly recognized as powerful tools for predicting the mutational effects of various proteins [20–22]. Most ML models learn the relationship between numerical representations derived from amino acid sequences and their functional characteristics [23]. In particular, protein language models (PLMs) have attracted significant attention because of their ability to extract meaningful representations from amino acid sequences [24].

PLMs are ML models that are often based on deep learning architectures employing techniques similar to those utilized in natural language processing [25–28]. Trained by millions of amino acid sequences, PLMs can capture the evolutionary patterns and structural relationships inherent to natural proteins. Fine-tuning the parameters in PLMs is a commonly beneficial approach for obtaining meaningful sequence representations for specific proteins. Fine-tuning involves adapting pretrained PLMs to particular tasks by further training them on smaller task-specific datasets [29–31]. For example, evolutionary fine-tuning (evotuning) is an unsupervised learning approach that uses homologous sequences of target proteins [25]. Fine-tuned models obtained via evotuning can capture evolutionary relationships and functionally relevant patterns in amino acid sequences, thereby enhancing their utility in downstream prediction tasks [32–34]. In another study, Vincoff et al. demonstrated that fine-tuning a PLM using sequences of fusion oncoproteins, a class of chimeric proteins collected from databases, led to superior performance compared with baseline embeddings in various fusion-specific tasks [35]. However, it is challenging to directly utilize this evotuning approach for the analysis of CAR sequences because CARs are synthetic proteins. Additionally, the absence of focused CAR sequence databases and the limited availability of reported CAR sequences further complicate the fine-tuning of PLMs for CARs. These factors hinder the accurate prediction of CAR T-cell activity using PLMs.

In this study, we generated unlabeled CAR sequences through sequence augmentation and developed prediction models for CAR T-cell activity by fine-tuning PLMs with the generated sequences. To generate a large and diverse set of unlabeled CAR sequences, homologous sequences from individual domains of a target CAR region (Fig. 1a) —specifically, CD28 and CD3ζ—were randomly combined. The generated sequences were then used to fine-tune ESM2, a PLM developed by Meta-AI that is available in various model sizes [36]. To validate whether our methodology was effective in predicting CAR-T cell activity, numerous CAR mutants were constructed by introducing point mutations into the wild-type sequence (Additional File 1: Table S1), and the cytotoxicity of CAR-T cells expressing these mutants was experimentally assessed. Subsequently, cytotoxicity was predicted based on the CAR sequences using the fine-tuned ESM2 (Fig. 1b). Our results demonstrate that fine-tuning ESM2 significantly improves the prediction performance of CAR-T cell activity. Furthermore, we showed that optimizing the fine-tuning conditions, such as the sequence diversity, number of training steps, and choice of ESM2 model size, can maximize the fine-tuning efficacy and enhance the prediction performance.

Fig. 1.

Fig. 1

Schematic outline of the methodology. a Diagram illustrating the target CAR region analyzed in this study. The focus is placed on the hinge domain (HD), transmembrane domain (TD), and intracellular domain (ICD), which are derived from CD28 and CD3ζ. b Workflow for predicting CAR-T cell activity. CAR mutants are generated by introducing point mutations into the target CAR region, and the cytotoxic activity of CAR-T cells is experimentally assessed. Additionally, unlabeled CAR sequences are constructed using the wild-type sequence as a template. These generated sequences are utilized to fine-tune PLMs, and the fine-tuned PLMs are subsequently employed to train the predictive model

Methods

Sequence augmentation

We utilized the HMMER software (https://www.ebi.ac.uk/Tools/hmmer/) [37] to search for homologous sequences of CD28 (position: 115–220) and CD3ζ (position: 52–164) using the phmmer algorithm. The E-value threshold was set to 1e-4 and UniProt was used as the database. After removing duplicates, sequences that were longer than × 1.25 or shorter than × 0.75 the length of the wild-type were filtered out. Subsequently, the sequences were converted into numerical vectors using one-hot encoding, and k-means clustering was performed for each domain. The number of clusters (K) was determined to ensure that at least 25% of the total sequences were included in the clusters containing wild-type sequences. In this study, K was set to 4 for CD28 and 3 for CD3ζ. The sequences in the high-diversity group were generated by randomly sampling and combining sequences from each cluster of CD28 and CD3ζ, whereas the sequences in the low-diversity group were generated by sampling and combining sequences from clusters containing the wild-type sequences of CD28 and CD3ζ. In this study, the high-diversity group comprised 5,500 sequences and the low-diversity group comprised 5,250 sequences. This scheme is illustrated in Additional File 1, Fig. S1.

CAR mutants

We used anti-CD19 FMC63-CAR, comprising FMC63 scFv, CD28, and CD3ζ domains, as the template sequence, referred to as 28z, and another FMC63-CAR, containing FMC63 scFv, CD8, 4-1BB, and CD3ζ domains, as the reference sequence, referred to as BBz. Random point mutations, up to a total of 10, were introduced into the hinge, transmembrane, and intracellular domains of CD28 in the template sequence. The number of mutations was determined to ensure that the mutation rate within the target regions was approximately 10%. Consequently, 382 CAR mutants were acquired.

Cell culture

Two donors of primary human CD8 T cells (STEMCELL Technologies, Vancouver, BC, Canada) were cultured in X-VIVO15 medium (LONZA, Basel, Switzerland) with 5% human serum (Sigma-Aldrich, St. Louis, MO, USA), 100 μM 2-mercaptoethanol (FUJIFILM-Wako Pure Chemical Corporation, Osaka, Japan), 25 U/mL human IL-2 (Peprotech, Rocky Hill, NJ, USA), 10 ng/mL human IL-7 (Miltenyi Biotec, Gladbach, Germany), and 10 ng/mL human IL-15 (Miltenyi Biotec).

Transduction

Double strand DNA (eBlocks Gene Fragments, Integrated DNA Technologies, Coralville, IA, USA) encoding CAR mutant sequences were inserted into the pHR-SFFV-P2A-EGFP plasmid vectors using the Gibson Assembly method [38] to easily monitor transduction efficiency by expressing EGFP simultaneously with CAR. Next, Escherichia coli cells transformed with plasmid vectors were cultured for vector expansion. Vector purification was performed using the Miniprep kit (Qiagen, Hilden, Germany). HEK293FT cells (Thermo Fisher Scientific, Waltham, MA, USA) were transfected with the pHR-GOI vector and packaging plasmid mix (pMD2.G, pCMVR8.74, and pAdvantage) to produce an incomplete replication lentivirus vector. The viral supernatant was harvested after 2 d. Human CD8+ T cells were stimulated with Human T-activator CD3/CD28 Dynabeads (Thermo Fisher Scientific) at a 1:1 cell-to-bead ratio, cultured for 24 h, and mixed with the viral supernatant. After culturing for 2 d, the beads were removed from the T cell culture using a Magnum FLX magnet plate (ALPAQUA, Beverly, MA, USA). After 3 d, transduction efficiency was measured by detecting GFP-expressing cell rate using flow cytometer (CytoFLEX S, Beckman Coulter, Brea, CA, USA). Transduced cells were used for cytotoxic assays.

Cytotoxicity assay

The cytotoxicity assays were performed using a bioluminescence-based method. Briefly, Nalm-Luc cells were seeded into 4 wells of a 384-well round-bottom plate (Sumitomo Bakelite, Tokyo, Japan). Subsequently, CAR-T cells were added at an effector-to-target (E:T) ratio of 1:8. After co-culturing for 23 h, the luciferin reagent (Promega, Madison, WI, USA) was added at half the volume of the total culture medium, and luminescence was measured using a microplate reader (SpectraMax i3x, Molecular Devices, Sunnyvale, CA, USA). Target cell cytotoxicity was determined using the following formula:

graphic file with name d33e414.gif

Negative control wells contained only Nalm-Luc cells, while CAR-T cells genetically engineered to express wild-type CAR were added to positive control wells at an E:T ratio of 2:1. Luminescence was measured in 16 wells for each control, and the median luminescence value was used to calculate cytotoxicity. The cytotoxicity values for each mutant from four wells were converted to median values for each donor, and thaverage of these median values was calculated.

Curve fitting

Curve fitting was performed using a four-parameter Gompertz model [39]. Cytotoxic activity ((y)) was fitted as a function of transduction efficiency ((x)) using the template sequence 28z and the reference sequence BBz. The following formula was used:

graphic file with name d33e426.gif

where (A), (B), (C), and (D) are the parameters of the model, each of which plays a role in controlling the shape and position of the curve. The goodness of fit was evaluated based on the R2 value.

Statistical analysis

Dunnett’s multiple comparison test was performed to assess the significance of the differences in Spearman’s rank correlation coefficient in each split of the fivefold cross-validation between the pre-trained and fine-tuned ESM2. Dunnett’s test was executed utilizing the “scipy” package (v1.11.2) in Python (v3.10.3).

Multi-dimensional scaling analysis

Multi-dimensional scaling (MDS), a statistical technique used to reduce the dimensionality of data, was used to visualize the generated sequences. Prior to MDS analysis, 2% of the sequences from both the high- and low-diversity groups were randomly sampled. The sequences were subsequently transformed into numerical vectors using one-hot encoding. The MDS analysis was conducted utilizing the “scikit-learn” package (v1.3.1) in Python (v3.10.3).

Fine-tuning

Full-model fine-tuning was conducted using masked language modeling (MLM). MLM is a training approach that predicts masked tokens in an input sequence based on its surrounding context. It is an unsupervised learning method, meaning it does not require labeled data and relies solely on the input sequence. In this study, 15% of the tokens in a sequence were randomly selected for prediction. Among the selected tokens, 80% were replaced with a [MASK] token, 10% were replaced with a random token, and the remaining 10% were left unchanged prior to prediction. The PyTorch framework [40] and Huggingface transformers library [41] were used to implement the code for loading and fine-tuning the ESM-2 model. Two datasets, comprising high- and low-diversity groups, were used for fine-tuning, with an 80/20 split for training and validation, respectively. An AdamW optimizer was used with a learning rate of 5e-6 for 50 epochs, and the batch size was set to 128. Fine tuning was performed using eight NVIDIA V100 graphical processing units (GPUs). The fine-tuned models were saved every five epochs and used in the downstream task. The fine-tuning was conducted utilizing the “torch” package (v1.13.1) in Python (v3.10.3).

Pseudo log-likelihood

Wang and Cho proposed a pseudo-log-likelihood (PLL) for a masked language [42]. The PLL was calculated for each mutable position (i) in the target sequence (x) when masked using parameters (θ), and the average logarithmic likelihood of all positions was computed. PLL was calculated using the following formula:

graphic file with name d33e466.gif

L is the length of the protein. In this study, we calculated the PLLs of 382 CAR mutant sequences using pre-trained ESM2-35 M and fine-tuned ESM2-35 M with high- and low-diversity groups and then averaged the results.

Sequence representation

A sequence containing L amino acids was input into the pretrained PLMs, resulting in a numerical matrix. The numerical matrix obtained from each layer of the PLMs was (L + 2) × H. The additional two positions account for the special [CLS] and [EOS] tokens added at the beginning and end of the sequence, respectively. H represents the hidden size, which corresponds to the size of the individual embedding in each PLM (Additional File 1: Table S2). Fixed-size sequence features were extracted from the [CLS] token of the last layer.

Prediction model

Prediction models were constructed using sequence representations to predict CAR-T cell cytotoxicity. In this study, we employed a ridge regression model, which is a type of linear regression that includes a regularization term. Nested cross-validation was used to assess the performance of the prediction model. Briefly, prior to prediction, the dataset was divided into five splits, with four used as training data and the remaining used as validation data, resulting in the creation of five models (fivefold cross-validation in the outer loop). The training data were further divided into three splits to optimize the hyperparameter of the Ridge model, regularization strength (α) (threefold cross-validation in the inner loop). During the optimization process, two of the three splits in the inner loop were used as training data to build the models by varying the value of α over the range of 10–6 to 106, while the remaining split data were used as validation data to evaluate the model performance. This process was repeated for each of the three splits, and the average prediction performance on the inner loop validation data was calculated. The value of α that yielded the best prediction performance was subsequently used to predict the outer loop validation data. Prediction scores, which represent the correlation between the predicted and actual values, were calculated for the five models. We utilized the “scikit-learn” package (v1.3.1) in Python (v3.10.3) for model implementation.

Prediction score

Two prediction scores were used to evaluate the prediction models. First, we calculated Spearman’s correlation coefficient between the estimated and actual values to assess the overall trend of the predictions. Second, we calculated the recall at K (Recall@K) and precision at K (Precision@K) to evaluate whether the high-cytotoxicity CAR mutants could be accurately predicted. Briefly, Recall@K evaluated how effectively the model retrieved all the high-cytotoxicity CARs, whereas Precision@K evaluated how many of the top K predictions were high-cytotoxicity CARs. In this study, we labeled the top 25% cytotoxicity values in the validation data as high-cytotoxicity CARs. Both Recall@K and Precision@K were calculated for K = 5 and K = 10 using the following formulas:

graphic file with name d33e494.gif
graphic file with name d33e497.gif

Results

Sequence augmentation for fine-tuning PLMs: high-diversity and low-diversity sequence groups

To fine-tune pretrained PLMs to predict CAR T-cell activity, a large number of CAR sequences are required. Therefore, numerous unlabeled CAR sequences have been generated via sequence augmentation. In this study, we focused on a chimeric region composed of CD28 and CD3ζ within a target CAR sequence. In the sequence augmentation process, homologous sequences for each protein domain were collected using the HMMER software [37] and combined randomly. Previous studies have shown that human 4-1BB and CD3 subunits are functional in mouse T cells [43–45]. Additionally, CARs with CD3ζ replaced by subunits of CD3 also exhibit functionality [46]. Therefore, we hypothesized that homologous sequences could be used in this context. Our sequence augmentation methodology facilitates the creation of a diverse set of sequences based on different combination patterns. It is well known that the diversity of training data has an impact on prediction results, and a high level of diversity is expected to contribute to improved performance [36]. Therefore, we established two sequence groups with distinct characteristics: a high-diversity group and a low-diversity group (details provided in the Methods section and Table 1). To visualize the resulting sequence patterns, sequences were one-hot encoded and then subjected to dimensionality reduction using multi-dimensional scaling (MDS) (Fig. 2a). The high-diversity group was distributed across the entire sequence space, whereas the low-diversity group clustered around the wild-type sequence. Additionally, the distribution of sequence lengths and Levenshtein distances [47] from the wild-type sequence are shown in Fig. 2b and c. The Levenshtein distance denotes the similarity between two sequences and is defined as the minimum number of single-character edits (insertions, deletions, or substitutions) required to transform one sequence into another. Compared with the low-diversity group, the high-diversity group exhibited a broader range of sequence lengths and lower similarities to the wild-type sequence. These results confirm the successful establishment of two sequence groups characterized by distinct levels of diversity.

Table 1.

Comparison of generated sequences across diversity levels

Diversity Number of sequences Length Levenshtein distance
Mean Max Min Mean Max Min
High 5500 211 237 170 87 153 0
Low 5250 218 226 180 45 91 0

Fig. 2.

Fig. 2

Analysis of generated CAR sequences. a MDS plot of the generated sequences, comparing the high-diversity and low-diversity groups. Orange points represent sequences in the high-diversity group, while blue points represent sequences in the low-diversity group. The filled contours depict the kernel density estimate (KDE). The yellow star denotes the wild-type sequence (i.e., the original sequence prior to the introduction of mutations). b Distribution of sequence lengths. c Levenshtein distances from the wild-type sequence

Data acquisition of CAR-T cell activity associated with mutated CAR sequences

To evaluate the impact of fine-tuning on the prediction of CAR-T cell activity, we obtained data on mutated CAR sequences and their cytotoxicity. Initially, these mutated CAR sequences were acquired by introducing point mutations into hinge, transmembrane, and intracellular domains in a wild-type CAR sequence, resulting in a total of 382 CAR mutants. The number of mutations per sequence ranged from 1 to 10, with the majority containing 10 mutations (Fig. 3a). To assess the spatial distribution of mutations, we analyzed the mutation positions and number of mutations for all mutants (Fig. S2a). The mutations appear to be roughly uniformly distributed, and this result indicates that the mutations in the CAR mutants represent uniform coverage rather than a representative sampling. Therefore, we considered that the impact of selection bias on model performance is negligible. Subsequently, the cytotoxicity of CAR T cells expressing the CAR mutants was investigated. Cytotoxicity was measured using T cells from two donors, and a positive correlation between transduction efficiency and cytotoxicity was experimentally observed across both donors (Fig. S3). The cytotoxicity averaged across two donors of the 382 mutants was distributed over the entire range, with that of the wild-type at approximately 50% (Fig. 3b). To further investigate whether specific mutations contribute to the improvement of cytotoxic activity, mutants harboring individual mutations at each position were collected, and their cytotoxic activity was evaluated (Fig. S2b). Compared to the wild-type, the cytotoxic activity of mutants at each position was generally reduced, indicating that mutations at specific amino acid positions are unlikely to significantly enhance cytotoxic activity. Thus, we successfully acquired CAR mutants and their cytotoxicity data, enabling the evaluation of their fine-tuning effects.

Fig. 3.

Fig. 3

Visualization of the CAR mutant dataset. a Distribution of the number of mutations per sequence. The numbers displayed above each bar represent the count, with no sequences containing 4 to 8 mutations. b Distribution of cytotoxicity across a total of 382 CAR mutants. The vertical gray line denotes the average cytotoxicity of all mutants, and the vertical orange line denotes the cytotoxicity of the wild-type sequence

Fine-tuning ESM2 and predicting cytotoxicity of CAR-T cells using generated CAR sequences

To determine the optimal model size of ESM2 for fine-tuning, we evaluated four pre-trained ESM2 models (Additional File 1: Table S2) and compared their predictive performance for the cytotoxicity of CAR mutants. For prediction, we chose ridge regression, a linear regression model that incorporates regularization to reduce overfitting, thereby enhancing its performance on high-dimensional data such as amino acid sequence representation [33]. Among the evaluated models, ESM2-35 M exhibited the best prediction performance, whereas the other models consistently showed lower prediction performances (Additional File 1: Fig. S4). Consequently, we concluded that ESM2-35 M was the optimal model for developing a higher-performing prediction model through fine-tuning.1

Next, we predicted the cytotoxicity using a fine-tuned ESM2-35 M model. Fine-tuning was executed for up to 50 epochs on sequences from both the high- and low-diversity groups. The loss and perplexity against the training and validation data gradually decreased toward 50 epochs in both groups (Additional File 1: Fig. S5). The relationship between the number of epochs and the Spearman's rank correlation coefficient is illustrated in Fig. 4a and b. In both groups, the prediction performance peaked within the initial training steps, with the high-diversity group peaking at 10 epochs and the low-diversity group peaking at 5 epochs. The prediction models were defined as ESM2-35MH/e10 and ESM2-35ML/e5, respectively. Moreover, ESM2-35MH/e10 led to a significant improvement in prediction performance, increasing by 20.2% (from 0.318 to 0.458) (Fig. 4c, Table 2), with a performance improvement observed across all splits in the cross-validation (Additional File 1: Fig. S6).

Fig. 4.

Fig. 4

Prediction performance of fine-tuned ESM2-35 M. a, b The relationship between the number of epochs and Spearman’s rank correlation coefficient. The points represent the average values across five models generated by cross-validation, while the shaded regions denote the standard deviation. Red arrows highlight the maximum value within the range of 5–50 epochs. c Comparison of Spearman’s rank correlation coefficients between the pre-trained and fine-tuned ESM2-35 M models. The results for the fine-tuned ESM2-35 M are obtained at the maximum point indicated in (a). White circles represent individual values for each cross-validation model. The thick central line represents the mean value, while the thin upper and lower lines denote the mean ± standard deviation (SD). Statistical comparisons are performed using Dunnett’s test (p < 0.01, n.s. = not significant). d, e Recall@K and Precision@K for K = 5 and K = 10

Table 2.

Summary of prediction performance using pre-trained and fine-tuned ESM2-35 M

Model Epoch Spearman ρ Recall@K Precision@K
Mean SD Top5 Top10 Top5 Top10
ESM2-35 M – 0.381 0.0264 0.084 0.188 0.320 0.360
ESM2-35MH/e10 10 0.458 0.0149 0.134 0.249 0.520 0.480
ESM2-35ML/e5 5 0.413 0.0494 0.073 0.166 0.280 0.320

In addition, a binary evaluation was performed to validate the capability of ESM2-35MH/e10 to predict highly cytotoxic CARs. The accurate prediction of top-ranked data is crucial for designing proteins with enhanced activity [48]. Specifically, the top 25% of cytotoxicity values in the validation data were classified as high-cytotoxicity CARs, and both Recall@K and Precision@K scores were calculated. These metrics are commonly used to evaluate the ranking and retrieval tasks. Recall@K indicates how effectively the model retrieved all high-cytotoxicity CARs, whereas Precision@K signifies how many of the top K predictions were high-cytotoxicity CARs. The evaluation results for K = 5 and K = 10 revealed that both Recall@K and Precision@K improved with ESM2-35MH/e10 (Fig. 4d, e). These findings demonstrate that fine-tuning with sequences from the high-diversity group enhanced the prediction performance of CAR-T cell cytotoxicity.

Discussion

PLMs are valuable tools for extracting meaningful representations from amino acid sequences. Fine-tuning pre-trained PLMs on task-specific datasets can enhance prediction performance in various bioinformatic applications. Nevertheless, fine-tuning PLMs for predicting CAR-T cell activity poses notable challenges given the artificial nature of CARs and the lack of focused CAR sequence databases. In this study, we generated unlabeled CAR sequences and investigated the effect of fine-tuning ESM2 using these sequences for downstream prediction tasks. Evotuning has been shown to be an effective approach for fine-tuning PLMs, enhancing their performance through training on homologous sequences [25]. Our sequence augmentation method, which involves a combination of homologous sequences, applies this approach. Although CARs are synthetic proteins, homologs of the domains in CARs have shown CAR-T cell functions [43, 46]. This supports the rationale that incorporating evolutionary sequence diversity can be an effective strategy.

The ESM-2 model used was primarily developed for predicting the three-dimensional structures of proteins [49]. However, sequence representations derived from ESM-2 are applicable to various downstream prediction tasks, including activity, binding, organismal fitness, and stability [50]. In this study, smaller-sized pre-trained ESM-2 achieved notable predictive performance (ρ = 0.381) (Table 2). Therefore, ESM-2 is likely to have also been effective in predicting CAR sequences, which involves signaling functions, and we determined ESM-2 to be a suitable base model for this study.

For sequence augmentation, clustering was performed using the characteristics of homologous sequences collected from the database, and two types of sequences were created by randomly combining them from distinct clusters. This methodology can generate a large and diverse set of sequences by adjusting the number of clusters and the manner in which they are combined. Consequently, our approach holds the potential for broader applications in the analysis of other multidomain proteins in future research.

Among the pretrained ESM2 models utilized in this study, ESM2-35 M exhibited the highest prediction performance for our specific task (Additional File 1: Fig. S4). The size of PLMs is directly linked to their performance in downstream tasks, with larger models generally exhibiting superior predictive capabilities [49, 51]. Nonetheless, the advantage of scaling up the model size depends on task complexity; smaller models may provide sufficient performance for certain tasks [32]. In our study, ESM2-35 M was the most appropriate choice for predicting CAR T-cell activity. These results indicate that task-specific requirements and data considerations are critical for determining the optimal model size. In addition, ESM2-35MH/e10 exhibited a statistically significant improvement in the prediction performance owing to fine-tuning. In contrast, the other models not only exhibited lower prediction performance but also failed to show any improvement through fine-tuning (Additional File 1: Fig. S7, Table S3). These results indicate the significance of selecting an appropriate model size for PLMs to effectively utilize the benefits of fine-tuning.

Only the high-diversity group significantly enhanced the predictive performance for CAR-T cell cytotoxicity (Fig. 4c). This result is consistent with a previous report hypothesizing that a high diversity of training data improves the robustness of the model [52]. Given the absence of CAR sequences in the pre-training data of ESM2, the high diversity of sequences may have enabled the model to gain insight into a broader range of sequence–function relationships. These findings suggest that adjusting sequence diversity for fine-tuning according to each downstream task is essential for maximizing prediction performance. Furthermore, in contrast to the prediction performance, both ESM-35MH/e10 and ESM-35ML/e5 showed similar log-likelihood values for the CAR mutants (Additional File 1: Fig. S8). A previous study highlighted that pretraining performance does not consistently correlate with downstream task performance [53]. Accordingly, while training data diversity does not notably impact log-likelihood scores of CARs, it plays a crucial role in improving CAR-T cell activity prediction.

During the fine-tuning process of ESM2-35 M using the high-diversity group, prediction performance peaked at 10 epochs (Fig. 4a). Beyond this point, the performance gradually decreased and the variance increased with additional epochs, indicating overfitting in certain cross-validation splits during the later stages of fine-tuning. The limited availability of training data and the use of mask-replacement tasks may contribute to overfitting [54]. Therefore, excessive fine-tuning of small datasets could compromise generalization, whereas selecting an optimal number of training steps has the potential to enhance prediction performance while minimizing overfitting.

To obtain cytotoxicity data from diverse CAR mutants, point mutations were introduced into the hinge, transmembrane, and intracellular domains of CD28, while no mutations were introduced into the intracellular domain CD3ζ. CD3ζ is a crucial protein for T-cell activation, and all currently approved CAR-T cell products incorporate wild-type-derived CD3ζ [2, 3]. From the perspective of data acquisition for machine learning, the random mutation generation method employed in this study is effective for producing diverse sequences. However, this approach poses the risk of enriching sequences with zero functional fitness, potentially introducing biases into the training data [55]. Based on these considerations, we considered that introducing mutations into CD3ζ was not desirable.

Among the CAR mutants evaluated for cytotoxic activity, more than half exhibited reduced cytotoxic activity compared to the wild-type (Fig. 3b). However, a subset of mutants demonstrated higher cytotoxic activity than the wild-type. An analysis of the mutation sites in the top 10 mutants with the highest cytotoxic activity (top 10) and the bottom 10 mutants with the lowest cytotoxic activity (bottom 10) revealed slight differences in the number of mutations in the transmembrane domain (Fig. S9). However, the mutations appear to be distributed across the entire sequence in both groups. Moreover, when looking at the presence or absence of mutations in the three signaling motifs (YMNM, PRRP, and PYAP) within the intracellular domain of CD28 [56]. Mouse CAR T cells with mutations in CD28 signaling motifs exhibit enhanced survival and function [18]. The both high and low activity mutants were found to harbor multiple mutations in this region (Fig. S9). While some mutations appear to be unique to the top 10 mutants, no specific mutations were identified as contributing directly to the improved cytotoxic activity (Fig. S2b). Instead, cytotoxic activity seems to be influenced by the synergistic combination of multiple mutations, highlighting the highly complex nature of the underlying mechanisms.

One technical limitation in this study is the inability to account for the influence of transduction efficiency in the evaluation of cytotoxic activity. The training data used for construction of prediction model inherently reflects the influence of transduction efficiency. To focus exclusively on improving cytotoxic activity through sequence optimization, it is important to account for the effects of transduction efficiency. However, we found that the relationship between transduction efficiency and cytotoxic activity differs between two types of second-generation CAR constructs (Fig. S10). Consequently, applying a universal correction approach, such as fitting curves based on wild-type controls, would be challenging. Moving forward, the development of correction methods specifically tailored to each construct will be necessary.

Another technical limitation is that the CAR constructs used in the experiments were limited to those incorporating the FMC63, CD28, and CD3ζ domains. While this construct is widely used and extensively studied, the findings may not be directly applicable to other CAR designs, particularly third-generation CARs or constructs comprising three or more protein domains. Furthermore, no experiments were performed using a different scFv targeting another antigen; therefore, the applicability of our findings to such cases remains unclear. The increased structural and functional complexity of such CARs could significantly impact the prediction performance of our computational model, requiring additional validation with a broader range of CAR constructs.

ML model-based optimization not only identifies the best sequences based on a trained sequence-to-function but also helps guide multiple rounds of computational predictions and experimental measurements [57]. In this context, we evaluated the prediction performance for top-ranked data included in the training dataset and assessed the extrapolation ability of the model to determine the applicability to novel proteins design (Fig. 4d, e). However, experimental validation of novel variants designed using this model was not conducted, and presents scope for future research.

Furthermore, the evaluation of CAR-T cells in this study primarily focused on cytotoxic activity, which, while essential, represents only one aspect of CAR-T cell functionality. Additional parameters, such as stemness and persistence, were not experimentally validated. To achieve a more comprehensive characterization of CAR-T cell activity, future studies should aim to incorporate multiple functional parameters into the predictive framework and validate their accuracy experimentally. Expanding the scope of the model to address these additional factors would enhance its utility and applicability in clinical and research settings.

Adjusting affinity through point mutations in the scFv can enhance therapeutic efficacy [9, 10]. Additionally, PLMs for scFv hit maturation are valuable for identifying candidate CARs [58]. Our approach remained applicable even as the number of target domains increased, suggesting the potential for future research to optimize full-length CAR, including the scFv region.

Conclusions

Our study demonstrated that sequence augmentation with different diversities, coupled with fine-tuning of the ESM2 model, specifically limited to small-scale models, can significantly improve the prediction performance of CAR-T cell cytotoxicity. In the future, combining our methodology with high-throughput screening of CAR mutants could facilitate the efficient design of novel CAR sequences, thereby enabling the enhancement and selection of CAR-T cell candidates.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

The authors would like to thank Satoru Mimura for technical assistance, Kenichi Hironaka, Tomohiko Okuda, Hiroya Ijima, Yoshihiro Osakabe, and Akinori Asahara for helpful discussion. Many thanks go to my supervisors, Shizu Takeda and Hiroko Hanzawa, for advice, encouragement, and support.

Abbreviations

CAR

Chimeric antigen receptor

scFv

Single-chain fragment variable

ML

Machine learning

PLMs

Protein language models

PLL

Pseudo-log-likelihood

MDS

Multi-dimensional scaling

Author contributions

KY conceptualized the method, developed the methodology, conducted the analysis, and wrote the main manuscript. SH contributed to conceptualizing the method and critically revised the manuscript. RT validated the programming codes and revised the manuscript. AO provided the resources essential for conducting the experiments and contributed to writing parts of the main manuscript. YI, TK, TM-Y, and DI assisted with the experimental design and provided critical feedback to improve the manuscript. All authors read and approved the final manuscript.

Funding

This study was funded by Hitachi, Ltd.

Data availability

The dataset supporting the conclusions of this study is included in this article (Additional File 2). Furthermore, the data analyzed and generated in this study, along with the analysis code, are available at https://github.com/keiyoshi91/CART-Activity-Prediction.

Declarations

Ethics approval and consent to participate

Ethics approval was not required for this study as it does not involve human participants, animals, or sensitive data requiring ethical review.

Consent for publication

Not applicable. This study does not involve human participants, animals, or other data requiring consent for publication.

Competing interests

The authors declare no competing interests.

Footnotes

1

Since our focus is on constructing CAR-specific prediction model and not revealing the correlation between performance and model sizes, we leave comprehensive analysis for future work.

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Sadelain M, Brentjens R, Rivière I. The basic principles of chimeric antigen receptor design. Cancer Discov. 2013;3:388–98. 10.1158/2159-8290.CD-12-0548. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Bhaskar ST, Dholaria B, Savani BN, Sengsayadeth S, Oluwole O. Overview of approved CAR-T products and utility in clinical practice. Clin Hematol Int. 2024;6:100–6. 10.46989/001c.124277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Bouchkouj N, Przepiorka D, Fashoyin-Aje LA. Obecabtagene autoleucel for B-cell acute lymphoblastic leukemia. JAMA. 2025. 10.1001/jama.2024.28312. [DOI] [PubMed] [Google Scholar]
  • 4.Korell F, Berger TR, Maus MV. Understanding CAR T cell-tumor interactions: paving the way for successful clinical outcomes. Med. 2022;3:538–64. 10.1016/j.medj.2022.05.001. [DOI] [PubMed] [Google Scholar]
  • 5.Albelda SM. Car T cell therapy for patients with solid tumours: key lessons to learn and unlearn. Nat Rev Clin Oncol. 2024;21:47–66. 10.1038/s41571-023-00832-4. [DOI] [PubMed] [Google Scholar]
  • 6.Park JH, Rivière I, Gonen M, Wang X, Sénéchal B, Curran KJ, et al. Long-term follow-up of CD19 CAR therapy in acute lymphoblastic leukemia. N Engl J Med. 2018;378:449–59. 10.1056/NEJMoa1709919. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Cappell KM, Kochenderfer JN. Long-term outcomes following CAR T cell therapy: what we know so far. Nat Rev Clin Oncol. 2023;20:359–71. 10.1038/s41571-023-00754-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Cho W, Liu JY, Beckett AN, Lunger JC, Xu P, Ho K, et al. A continuous landscape of signaling encodes a corresponding landscape of CAR T cell phenotype. 2025. 10.1101/2025.06.05.658149.
  • 9.Long AH, Haso WM, Shern JF, Wanhainen KM, Murgai M, Ingaramo M, et al. 4-1BB costimulation ameliorates T cell exhaustion induced by tonic signaling of chimeric antigen receptors. Nat Med. 2015;21:581–90. 10.1038/nm.3838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Bishop MR, Dickinson M, Purtill D, Barba P, Santoro A, Hamad N, et al. Second-line tisagenlecleucel or standard care in aggressive B-Cell lymphoma. N Engl J Med. 2022;386:629–39. 10.1056/NEJMoa2116596. [DOI] [PubMed] [Google Scholar]
  • 11.Maziarz RT, Bishop MR, Tam CS, Borchmann P, Worel N, McGuirk JP, et al. Five-year analysis of the JULIET trial of tisagenlecleucel in patients with relapsed/refractory large B-cell lymphoma. J Clin Oncol. 2025. 10.1200/JCO-25-00507. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Neelapu SS, Locke FL, Bartlett NL, Lekakis LJ, Miklos DB, Jacobson CA, et al. Axicabtagene ciloleucel CAR T-cell therapy in refractory large b-cell lymphoma. N Engl J Med. 2017;377:2531–44. 10.1056/NEJMoa1707447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Westin JR, Oluwole OO, Kersten MJ, Miklos DB, Perales M-A, Ghobadi A, et al. Survival with axicabtagene ciloleucel in large B-cell lymphoma. N Engl J Med. 2023;389:148–57. 10.1056/NEJMoa2301665. [DOI] [PubMed] [Google Scholar]
  • 14.Bachy E, Le Gouill S, Di Blasi R, Sesques P, Manson G, Cartron G, et al. A real-world comparison of tisagenlecleucel and axicabtagene ciloleucel CAR T cells in relapsed or refractory diffuse large B cell lymphoma. Nat Med. 2022;28:2145–54. 10.1038/s41591-022-01969-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Mei A. Engineering next-generation chimeric antigen receptor-T cells_ recent breakthroughs and remaining challenges in design and screening of novel chimeric antigen receptor variants. Curr Opin Biotechnol. 2024. 10.1016/j.copbio.2024.103223. [DOI] [PubMed] [Google Scholar]
  • 16.Di Roberto RB, Castellanos-Rueda R, Frey S, Egli D, Vazquez-Lombardi R, Kapetanovic E, et al. A functional screening strategy for engineering chimeric antigen receptors with reduced on-target, off-tumor activation. Mol Ther. 2020;28:2564–76. 10.1016/j.ymthe.2020.08.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Sharma P, Marada VVVR, Cai Q, Kizerwetter M, He Y, Wolf SP, et al. Structure-guided engineering of the affinity and specificity of CARs against Tn-glycopeptides. Proc Natl Acad Sci U S A. 2020;117:15148–59. 10.1073/pnas.1920662117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Boucher JC, Li G, Kotani H, Cabral ML, Morrissey D, Lee SB, et al. CD28 costimulatory domain–targeted mutations enhance chimeric antigen receptor T-cell function. Cancer Immunol Res. 2021;9:62–74. 10.1158/2326-6066.CIR-20-0253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Feucht J. Calibration of CAR activation potential directs alternative T cell fates and therapeutic potency. Nat Med. 2019;25:82–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Greenhalgh JC, Fahlberg SA, Pfleger BF, Romero PA. Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production. Nat Commun. 2021;12:5825. 10.1038/s41467-021-25831-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wu Z, Kan SBJ, Lewis RD, Wittmann BJ, Arnold FH. Machine learning-assisted directed protein evolution with combinatorial libraries. Proc Natl Acad Sci U S A. 2019;116:8852–8. 10.1073/pnas.1901979116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Saito Y, Oikawa M, Sato T, Nakazawa H, Ito T, Kameda T, et al. Machine-learning-guided library design cycle for directed evolution of enzymes: the effects of training data composition on sequence space exploration. ACS Catal. 2021;11:14615–24. 10.1021/acscatal.1c03753. [Google Scholar]
  • 23.Yang KK, Wu Z, Bedbrook CN, Arnold FH. Learned protein embeddings for machine learning. Bioinformatics. 2018;34:2642–8. 10.1093/bioinformatics/bty178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Schwartz AS, Hannum GJ, Dwiel ZR, Smoot ME, Grant AR, Knight JM, et al. Deep semantic protein representation for annotation, discovery, and engineering. 2018. 10.1101/365965.
  • 25.Alley EC, Khimulya G, Biswas S, AlQuraishi M, Church GM. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods. 2019;16:1315–22. 10.1038/s41592-019-0598-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci U S A. 2021;118:e2016239118. 10.1073/pnas.2016239118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Wang Y, Jones L, et al. ProtTrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44:7112–27. 10.1109/TPAMI.2021.3095381. [DOI] [PubMed] [Google Scholar]
  • 28.Ferruz N, Schmidt S, Höcker B. ProtGPT2 is a deep unsupervised language model for protein design. Nat Commun. 2022;13:4348. 10.1038/s41467-022-32007-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wang D, Ye F, Zhou H. On pre-trained language models for antibody. 2023. 10.48550/arXiv.2301.12112.
  • 30.Sledzieski S, Kshirsagar M, Baek M, Dodhia R, Lavista Ferres J, Berger B. Democratizing protein language models with parameter-efficient fine-tuning. Proc Natl Acad Sci U S A. 2024;121:e2405840121. 10.1073/pnas.2405840121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Schmirler R, Heinzinger M, Rost B. Fine-tuning protein language models boosts predictions across diverse tasks. Nat Commun. 2024;15:7407. 10.1038/s41467-024-51844-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Gordon C, Lu AX, Abbeel P. Protein language model fitness is a matter of preference. 2024. 10.1101/2024.10.03.616542.
  • 33.Biswas S. Low-N protein engineering with data-efficient deep learning. Nat Methods. 2021. 10.1038/s41592-021-01100-y. [DOI] [PubMed] [Google Scholar]
  • 34.Yamaguchi H, Saito Y. Evotuning protocols for transformer-based variant effect prediction on multi-domain proteins. Brief Bioinform. 2021;22:bbab234. 10.1093/bib/bbab234. [DOI] [PubMed] [Google Scholar]
  • 35.Vincoff S, Goel S, Kholina K, Pulugurta R, Vure P, Chatterjee P. FusOn-pLM: a fusion oncoprotein-specific language model via adjusted rate masking. Nat Commun. 2025;16:1436. 10.1038/s41467-025-56745-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–30. 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 37.Eddy SR. Accelerated profile HMM searches. PLoS Comput Biol. 2011;7:e1002195. 10.1371/journal.pcbi.1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Gibson DG, Young L, Chuang R-Y, Venter JC, Hutchison CA, Smith HO. Enzymatic assembly of DNA molecules up to several hundred kilobases. Nat Methods. 2009;6:343–5. 10.1038/nmeth.1318. [DOI] [PubMed] [Google Scholar]
  • 39.Tjørve E, Tjørve KMC. A unified approach to the Richards-model family for use in growth analyses: why we need only two model forms. J Theor Biol. 2010;267:417–25. 10.1016/j.jtbi.2010.09.008. [DOI] [PubMed] [Google Scholar]
  • 40.Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. PyTorch: an imperative style, high-performance deep learning library. in: advances in neural information processing systems. Curran Associates, Inc.; 2019.
  • 41.Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: state-of-the-art natural language processing. In: Liu Q, Schlangen D, editors. Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations. Online: Association for Computational Linguistics; 2020. p. 38–45. 10.18653/v1/2020.emnlp-demos.6.
  • 42.Wang A, Cho K. BERT has a mouth, and it must speak: BERT as a Markov random field language model. 2019. 10.48550/arXiv.1902.04094.
  • 43.Wang L-CS, Lo A, Scholler J, Sun J, Majumdar RS, Kapoor V, et al. Targeting fibroblast activation protein in tumor stroma with chimeric antigen receptor T cells can inhibit tumor growth and augment host immunity without severe toxicity. Cancer Immunol Res. 2014. 10.1158/2326-6066.CIR-13-0027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Ueda O, Wada NA, Kinoshita Y, Hino H, Kakefuda M, Ito T, et al. Entire CD3ε, δ, and γ humanized mouse to evaluate human CD3–mediated therapeutics. Sci Rep. 2017;7:45839. 10.1038/srep45839. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Zhang R, Zhang J, Zhou X, Zhao A, Yu C. The establishment and application of CD3E humanized mice in immunotherapy. Exp Anim. 2022;71:442–50. 10.1538/expanim.22-0012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Velasco Cárdenas RM-H, Brandl SM, Meléndez AV, Schlaak AE, Buschky A, Peters T, et al. Harnessing CD3 diversity to optimize CAR T cells. Nat Immunol. 2023;24:2135–49. 10.1038/s41590-023-01658-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Levenshtein V. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady. 1965.
  • 48.Zhang Y, Chen Y, Wang C, Lo C-C, Liu X, Wu W, et al. ProDCoNN: protein design using a convolutional neural network. Proteins. 2020;88:819–29. 10.1002/prot.25868. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. 2022. 10.1101/2022.07.20.500902.
  • 50.Notin P, Kollasch AW, Ritter D, Niekerk L van, Paul S, Spinner H, et al. ProteinGym: large-scale benchmarks for protein design and fitness prediction. 2023. 10.1101/2023.12.07.570727.
  • 51.Chen B, Cheng X, Li P, Geng Y, Gong J, Li S, et al. xTrimoPGLM: unified 100B-scale pre-trained transformer for deciphering the language of protein. 2024. 10.48550/arXiv.2401.06199.
  • 52.Gontijo-Lopes R, Smullin SJ, Cubuk ED, Dyer E. Affinity and diversity: quantifying mechanisms of data augmentation. 2020. 10.48550/arXiv.2002.08973.
  • 53.Li F-Z, Amini AP, Yue Y, Yang KK, Lu AX. Feature reuse and scaling: understanding transfer learning with protein language models. 2024. 10.1101/2024.02.05.578959.
  • 54.Fournier Q, Vernon RM, Sloot A van der, Schulz B, Chandar S, Langmead CJ. Protein language models: is scaling necessary? 2024. 10.1101/2024.09.23.614603.
  • 55.Wittmann BJ, Yue Y, Arnold FH. Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst. 2021;12:1026-1045.e7. 10.1016/j.cels.2021.07.008. [DOI] [PubMed] [Google Scholar]
  • 56.Isakov N, Altman A. PKC-theta-mediated signal delivery from the TCR/CD28 surface receptors. Front Immunol. 2012. 10.3389/fimmu.2012.00273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Hie BL, Yang KK. Adaptive machine learning for protein engineering. Curr Opin Struct Biol. 2022;72:145–52. 10.1016/j.sbi.2021.11.002. [DOI] [PubMed] [Google Scholar]
  • 58.Janocha K, Ling A, Godson A, Lampi Y, Bornschein S, Hammerla NY. Harnessing preference optimisation in protein LMs for hit maturation in cell therapy. 2024. 10.48550/arXiv.2412.01388.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The dataset supporting the conclusions of this study is included in this article (Additional File 2). Furthermore, the data analyzed and generated in this study, along with the analysis code, are available at https://github.com/keiyoshi91/CART-Activity-Prediction.


Articles from BMC Bioinformatics are provided here courtesy of BMC

RESOURCES