Skip to main content
PLOS Neglected Tropical Diseases logoLink to PLOS Neglected Tropical Diseases
. 2025 Apr 29;19(4):e0012985. doi: 10.1371/journal.pntd.0012985

Transformer-based deep learning enables improved B-cell epitope prediction in parasitic pathogens: A proof-of-concept study on Fasciola hepatica

Rui-Si Hu 1,2, Kui Gu 3,*, Muhammad Ehsan 4, Sayed Haidar Abbas Raza 5, Chun-Ren Wang 6,*
Editor: Aysegul Taylan Ozkan7
PMCID: PMC12064019  PMID: 40300022

Abstract

Background

The identification of B-cell epitopes (BCEs) is fundamental to advancing epitope-based vaccine design, therapeutic antibody development, and diagnostics, such as in neglected tropical diseases caused by parasitic pathogens. However, the structural complexity of parasite antigens and the high cost of experimental validation present certain challenges. Advances in Artificial Intelligence (AI)-driven protein engineering, particularly through machine learning and deep learning, offer efficient solutions to enhance prediction accuracy and reduce experimental costs.

Methodology/Principal findings

Here, we present deepBCE-Parasite, a Transformer-based deep learning model designed to predict linear BCEs from peptide sequences. By leveraging a state-of-the-art self-attention mechanism, the model achieved remarkable predictive performance, achieving an accuracy of approximately 81% and an AUC of 0.90 in both 10-fold cross-validation and independent testing. Comparative analyses against 12 handcrafted features and four conventional machine learning algorithms (GNB, SVM, RF, and LGBM) highlighted the superior predictive power of the model. As a case study, deepBCE-Parasite predicted eight BCEs from the leucine aminopeptidase (LAP) protein in Fasciola hepatica proteomic data. Dot-blot immunoassays confirmed the specific binding of seven synthetic peptides to positive sera, validating their IgG reactivity and demonstrating the model’s efficacy in BCE prediction.

Conclusions/Significance

deepBCE-Parasite demonstrates excellent performance in predicting BCEs across diverse parasitic pathogens, offering a valuable tool for advancing the design of epitope-based vaccines, antibodies, and diagnostic applications in parasitology.

Author Summary

Antigen-antibody interactions are critical events in the humoral immune response, facilitating the recognition and neutralization of invasive parasites. BCEs, defined as surface-exposed clusters of amino acids recognized by B-cell receptors or antibodies, play a critical role in initiating a humoral immune response. This study focuses on the identification of parasite BCEs, which serve as promising targets for the development of vaccines, therapeutic antibodies, and diagnostic tools. To this end, we developed a deep learning model, termed deepBCE-Parasite, which was rigorously benchmarked against and integrated with traditional machine learning models. By leveraging state-of-the-art AI techniques, these models enable rapid and precise BCE identification directly from amino acid sequences, rendering it particularly suitable for large-scale epitope screening. As a proof-of-concept, we applied these AI-driven models to predict BCEs in F. hepatica, a globally distributed parasite responsible for fascioliasis, a neglected tropical disease. Utilizing available proteomics data of this trematode species, we identified peptides exhibiting high specificity for antibody binding. This work highlights the potential of AI in advancing epitope prediction within parasitology, providing a rapid, scalable, and cost-effective strategy for discovering immune targets.

Introduction

Parasitic diseases, caused by both established and emerging infections, remain a significant global health challenge, particularly neglected tropical diseases (NTDs) associated with foodborne parasites [1]. These parasites often cause opportunistic infections in humans and livestock, imposing not only significant public health burdens on humans but also considerable and often underestimated economic losses in the livestock industry [2]. Notable examples of such NTDs include fascioliasis, caused by Fasciola hepatica and Fasciola gigantica, which pose significant economic and public health challenges globally [3,4]. Although conventional broad-spectrum antiparasitic drugs are effective in controlling or eradicating these parasites, several limitations and challenges hinder their long-term efficacy and safety. Among these, the growing emergence of drug resistance represents as a major hurdle, driven by factors such as genomic evolution, improper drug administration, environmental factors, and limited availability of novel therapeutics [5,6]. Consequently, antibody-based therapies and vaccine development are increasingly being explored as promising alternatives for the prevention and control of parasitic infections.

Upon host invasion, parasite and its molecular components act as antigens that trigger host immune response. Some specific antibodies recognize BCEs, which are short amino acids sequences on antigenic proteins, thereby mediating humoral immunity against pathogens [7]. Nevertheless, the identification of BCE remains challenging due to the complex interactions between antibody variable regions and epitope sequences [8]. BCEs can be categorized as either conformational or linear epitopes [9]. Conformational epitopes comprise discontinuous amino acid residues that are spatially adjacent through protein folding, requiring structural modeling for precise prediction. In contrast, linear epitopes consist of continuous amino acid sequences that maintain their antigenicity even under denatured or partially unfolded conditions. Traditional in silico approaches with respect to the prediction of parasite BCEs predominantly depend on bioinformatics screening coupled with experimental immunogenicity validation [1013]. Although computational methodologies have advanced significantly, the identification of BCEs using large-scale datasets has not been thoroughly investigated in parasitic immunology.

Machine learning and deep learning approaches have been extensively applied in the prediction of BCEs. For conformational epitopes, various algorithms, including Support Vector Machine (SVM), Adaptive Boosting (AdaBoost), Logistic Regression (LR), eXtreme Gradient Boosting (XGBoost), and pre-trained protein structure models (e.g., ESM-2), have demonstrated exceptional predictive accuracy [1420]. Regarding linear BCEs, sequence-based approaches employing Recurrent Neural Networks (RNN), Language Models (LM), Extremely Randomized Tree (ERT), Gradient Boosting (GB), Random Forest (RF), and SVM have been extensively implemented [2127]. In the field of parasitology, the post-genomic era has witnessed two significant developments: the completion of genomic sequencing projects for diverse protozoan and helminthic parasites [28,29], and considerable advancements in multi-omics technologies, particularly transcriptomics and proteomics [3032]. These progressions have notably expanded parasite datasets, providing unprecedented opportunities to integrate AI models with multi-omics data, thereby enhancing prediction accuracy while reducing experimental costs.

In this study, we developed a Transformer-based deep learning model to predict linear BCEs in parasites. Trained on experimentally validated datasets from diverse parasite species, the model achieved an accuracy of 80.97% and an AUC of 0.9 on an independent test set. To benchmark its performance, we compared it with traditional machine learning approaches using 12 feature extraction techniques and four classifiers. Furthermore, we applied our models to proteomic data from the liver fluke Fasciola hepatica, predicting eight peptide sequences derived from the leucine aminopeptidase (LAP) protein. These peptides were synthesized and experimentally validated via dot-blot immunoassays, revealing that seven exhibited specific binding to antibodies in F. hepatica-positive ovine sera. Our findings demonstrate the efficacy of Transformer-based and machine learning models in predicting BCEs in parasites, providing valuable insights for the development of antibody-based therapeutics, peptide vaccines, and immunodiagnostic tools.

Methods

Dataset preparation of parasites

Although numerous prediction tools for BCEs have been developed (Table 1), none have been specifically trained on parasitic datasets, thus limiting their applicability to parasite-specific antigens. Moreover, the effective AI model training necessitates a well-curated dataset comprising both positive and negative samples. In this study, positive and negative samples of parasite BCEs were systematically collected from the Immune Epitope Database (IEDB) [33] and Bcipep database [34] (retrieved in April 2024). Initially, 8,128 positive BCE sequences were obtained. To mitigate redundancy and enhance dataset diversity for robust model training, sequence clustering was performed using CD-HIT v4.8.1 [35] with a similarity threshold of 0.8, yielding a refined dataset of 5,752 unique positive peptide sequences. The IEDB provides an experimentally validated negative dataset of parasite BCEs, which were extracted, deduplicated using the CD-HIT software, and numerically balanced to match the positive sample set. The final dataset was divided into training and test sets in an 80:20 ratio, ensuring rigorous model training and evaluation.

Table 1. A summary of bioinformatics tools for conformational and linear BCE prediction.

Epitope Predictor Method ACC AUC SE SP MCC Reference
Conformational CBTOPE SVM 0.87 - 0.83 0.90 0.73 Ansari et al., [14]
epitope3D Adaboost 0.70 0.78 - - 0.55 da Silva et al., [15]
DiscoTope-3.0 XGBoost - 0.81 - - 0.52 Høie et al., [16]
SEMA 2.0 ESM-2 - 0.78 0.70 - 0.23 Ivanisenko et al., [17]
EPSVR SVR - 0.60 - - - Liang et al., [18]
ElliPro Thornton 0.84 0.73 0.60 0.86 - Ponomarenko et al., [19]
SEPPA 3.0 LR 0.67 0.75 - - - Zhou et al., [20]
Linear LBCEPred RF 0.87 0.93 0.86 0.87 0.73 Alghamdi et al., [21]
BepiPred3 LM 0.69 0.76 0.70 - 0.31 Clifford et al., [22]
EpiDope DNN 0.77 0.67 - - - Collatz et al., [23]
iBCE-EL ERT, GB 0.73 0.79 0.74 0.72 0.46 Manavalan et al., [24]
DeepLBCEPred CNN, LSTM 0.77 - 0.78 0.75 0.54 Qi et al., [25]
ABCpred RNN 0.66 - 0.67 0.65 0.32 Saha et al., [26]
SVMTriP SVM - 0.70 0.80 - - Yao et al., [27]

Deep learning architecture

Embedding.

The model input consists of amino acid sequences, where each amino acid is represented as a dense vector through an embedding vector. This layer comprises two components: amino acid embedding and positional encoding.

  • (1)

    Amino acid embedding:

Each amino acid residue is mapped to a low-dimensional vector using a learned embedding matrix EV×dk, where V=24 (the vocabulary size, representing the 20 standard amino acids plus 4 special tokens) and dk =256 (the embedding dimension). The embedding matrix is optimized during training via backpropagation to capture semantic relationships between amino acids.

  • (2)

    Positional encoding:

To encode sequence order information, positional encoding is added to the amino acid embeddings. The encoding for position p is computed as:

PE(p)2i=sin(p100002idk), PE(p)2i+1=cos(p100002idk)

Here, p denotes the position of the amino acid in the sequence, and 2i, 2i+1 represents even and odd dimensions of the positional encoding vector. The final embedding for each residue is the sum of its amino acid embedding and positional encoding:

X=AminoAcidEmbedding(a)+PositionalEncoding(p)

Encoder-decoder architecture.

  • (1)

    Multi-head attention mechanism:

An encoder-decoder architecture is built around a multi-head attention mechanism, which enables the model to focus on different parts of the input sequence simultaneously. This mechanism is implemented in both the encoder and decoder layers to capture complex dependencies within the sequences. For each attention head, queries (Q), keys (K), and values (V) are computed through learned linear transformations:

Q=WQ×X,  K=WK×X,  V=WV×X

Here, X represents the input embeddings, and WQ, WK, and WV are learned weight matrices.

The attention scores are computed as:

Attention(Q,K,V)=softmax(QKTdk)V

where dk=32 is the dimensionality of the keys. The scaling factor dk ensures stable gradients during training. The multi-head attention mechanism applies this process across 8 parallel attention heads, allowing the model to capture diverse aspects of the input sequence. The outputs from all heads are concatenated and projected back into the model’s dimensional space (dk=256) using a learned linear transformation.

  • (2)

    Feed-forward network (FFN):

Both the encoder and decoder layers contain a FFN, which consists of two linear transformations with a ReLU activation function in between:

FFN(x)=max(0,xW1+b1)W2+b2

Here, W1256×2048 and W2256×2048 are learned weight matrices, and b1, b2 are bias terms. The FFN expands the hidden dimension to dff=2048 to enhance the model’s capacity for nonlinear representation.

  • (3)

    Feature optimization block (FOB):

The FOB refines learned features using a 1D convolutional layer (kernel size 3, padding 1), followed by ReLU activation and dropout (p=0.3). The convolutional layer extracts local features, while dropout reduces overfitting. The output is projected back to the model’s dimensional space (dk=256) via a fully connected layer.

  • (4)

    Encoder layer:

The encoder layer consists of 2 sub-layers, namely multi-head attention and FFN. The input is first processed by the multi-head attention mechanism, followed by the FNN. Each sub-layer is accompanied by a residual connection and layer normalization:

EncoderOutput1=Norm(FFN(MultiHeadAttention(x)))

In addition, we incorporate a FOB to further refine the learned features using convolutional layers and dropout.

EncoderOutput=FOB(EncoderOutput1)
  • (5)

    Decoder layer:

The decoder layer has a structure similar to the encoder but includes an additional encoder-decoder attention mechanism. This mechanism enables the decoder to focus on the encoder’s output while processing the target sequence, integrating information from both the input and target sequences to generate predictions.

Model output.

The final output from the decoder is passed through an adaptive average pooling layer, followed by a fully connected (FC) layer and a ReLU activation. Dropout is applied to prevent overfitting:

Output=ReLU(FC(Pool(DecoderOutput)))

Design of traditional machine learning models

Comparing the performance of different methods in the presence of a large number of handcrafted features is a complex task. To tackle this, we selected twenty representative statistical features as the foundation for our prediction and analysis. These features were derived from classical protein sequence encoding methods, including Amino Acid Composition (AAC), Adaptive Skip Dipeptide Composition (ASDC), Composition of k-Spaced Amino Acid Group Pairs (CKSAAGP), Composition of k-Spaced Amino Acid Pairs (CKSAAP), Dipeptide Deviation from Expected Mean (DDE), Di-Peptide Composition (DPC), Grouped Amino Acid Composition (GAAC), Grouped Di-Peptide Composition (GDPC), Grouped Tri-Peptide Composition (GTPC), Pseudo-Amino Acid Composition (PAAC), Quasi Sequence Order (QSO), and Sequence Order Coupling Number (SOCN). These features were extracted using the iLearnPlus tool [36].

To reduce noise in the feature matrix and negative effects on model performance, we employed the Light Gradient-Boosting Machine (LGBM) algorithm to rank all extracted features based on their importance. The proven capability of LGBM in feature selection makes it an ideal choice for enhancing model performance in high-dimensional tasks [37,38]. From this ranking, the top 200 features were selected. Subsequently, Recursive Feature Elimination (RFE) was applied to further fine-tune the selection, thereby improving both model performance and robustness.

For the classification task, we utilized four classic classification algorithms that are widely adopted in biological sequence analysis and prediction: SVM, RF, LGBM, and Gaussian Naïve Bayes (GNB). These algorithms were chosen due to their demonstrated effectiveness to handle complex, high-dimensional datasets, such as protein sequences, while maintaining robustness and generalizability. To implement and compare the performance of these features and classifiers, we used the scikit-learn (https://scikit-learn.org/stable/) library to execute machine learning algorithms.

Model evaluation metrics

We assessed the performance of all models using four standard metrics: accuracy (ACC), sensitivity (SE), specificity (SP), and the Matthews correlation coefficient (MCC) [3941].

&ACC=TP+TNTP+FP+TN+FN ×100%&&SE=TPTP+FN ×100%&&&SP=TNTN+FP × 100%&&MCC=(TP×TN)(FP×FN)(TP+FP)×(TN+FN)×(TP+FN)×(TN+FP)&&

In the aforementioned formulas, “True Positive” (TP) denotes the number of instances correctly classified as positive (i.e., accurately predicted BCEs), while “True Negative” (TN) represents the number of instances correctly classified as negative (i.e., accurately predicted non-BCEs). Conversely, “False Positive” (FP) indicates the number of negative instances erroneously classified as positive (i.e., non-BCEs misclassified as BCEs), and “False Negative” (FN) refers to the number of positive instances erroneously classified as negative (i.e., BCEs misclassified as non-BCEs). ACC indicates the proportion of correctly predicted instances, including both true positives and true negatives. SE measures the proportion of correctly predicted positives among all positive instances, while SP reflects the proportion of correctly predicted negatives among all negative instances. MCC provides a comprehensive assessment of model performance by considering all four prediction categories (TP, FP, TN, FN), offering a balanced assessment. Additionally, the Area Under the Curve (AUC) refers to the area under the Receiver Operating Characteristic (ROC) curve, which illustrates the model’s ability to distinguish between positive and negative samples, with values closer to 1 indicating better classification performance.

A case study for wet-lab evaluation

Peptide identification in the liver fluke Fasciola hepatica.

To validate and evaluate the applicability of our models in parasitic research, we used F. hepatica as a case study. The raw proteomic sequencing data of F. hepatica were retrieved from the iProX database (https://www.iprox.cn/; Project ID: IPX0002165000). This dataset comprises mass spectrometry (MS)-based proteomic data from four developmental stages of F. hepatica: (1) Metacercaria, obtained from Galba pervia snails cultured under suitable conditions for 30–45 days, during which cercariae emerged and underwent cyst formation; (2) Juvenile fluke, isolated from the livers of sheep artificially infected with metacercariae for 28 days; (3) Immature fluke, recovered from the livers of sheep infected for 59 days; and (4) Adult fluke, extracted from the livers of sheep infected for 118 days [42]. The protein sequences of F. hepatica were downloaded from the UniProt database on October 12, 2024, using the keyword “Fasciola hepatica”. Analysis of raw data was conducted using MaxQuant v2.6.5.0 [43], extracting key parameters from the proteinGroups file, including peptide counts, label-free quantification (LFQ) intensity values, sequence coverage, molecular weight, and FASTA headers. To facilitate cross-sample comparisons, LFQ intensity values were normalized using Z-score standardization.

Model prediction of potential BCEs.

Based on MS-based proteomics data, we applied our models to predict potential linear BCEs of F. hepatica. To enhance prediction robustness, we conducted an intersection analysis across multiple models, retaining only epitopes consistently identified by different approaches. Furthermore, to elucidate the biological relevance of the predicted BCEs, we performed subcellular localization analysis of their source proteins using DeepLoc 2.0 [44] and assessed protein-protein interaction networks for key proteins via the STRING database [45].

Synthesis, antigenicity, and experimental validation of BCEs via dot-blot immunoassays.

To experimentally validate the predicted BCEs, eight peptides derived from the LAP protein (Uniport ID: A0A345G0S4) were synthesized, alongside a human peptide sequence (RSRTPSLPTPPTREP) used as a control. All peptides were chemically synthesized by Sangon Biotech Co., Ltd. (Shanghai, China), with the purity of the synthesized peptides confirmed to exceed 95% through high-performance liquid chromatography (HPLC) and MS analyses. Dot-blot immunoassays were performed on each nitrocellulose membrane (purchased from Beyotime Biotech Co., Ltd., Shanghai, China), with three peptides (10 μg of each) spotted per membrane. These membranes were air-dried, blocked with PBS containing 2% non-fat milk, and incubated overnight at 4°C with F. hepatica-positive and negative (control) ovine serum antibodies (diluted 1:400). Following washing, the membranes were incubated at room temperature for 1 hour with horseradish peroxidase (HRP)-conjugated rabbit anti-ovine IgG (H+L) (AS029, diluted 1:3000; ABclonal Biotech Co., Ltd., Wuhan, China). The binding signals were detected using the SuperPico ECL Master Mix (Vazyme Biotech Co., Ltd., Nanjing, China) and visualized with an enhanced chemiluminescence (ECL) detection system. Signal images were captured using a digital imaging platform. All peptides were tested in three independent experimental repetitions, and only peptides that consistently exhibited positive dot signals across all three replicates were confirmed as true BCEs specific to IgG antibodies.

Results

Design of the deep learning pipeline

We proposed a deep learning pipeline, named deepBCE-parasite, based on a custom Transformer architecture for predicting BCEs in parasite proteins. This model was trained on a balanced dataset containing 5,752 positive and 5,752 negative BCEs, with amino acid features encoded using amino acid embedding and positional encoding. The architecture employs a binary classification framework using an encoder-decoder structure (Fig 1A). The innovation is the integration of a multi-head attention mechanism within this encoder-decoder framework, which enables efficient feature extraction and context-aware sequence modeling (Fig 1B). Model hyperparameters include a vocabulary size of 24, an embedding dimension of 256, two encoder layers, eight attention heads, and a feedforward hidden layer dimension of 2,048. Detailed parameter settings are provided in S1 Table. This architecture offers a scalable and robust solution for epitope identification in complex parasite protein sequences.

Fig 1. Schematic representation of the study workflow.

Fig 1

(A) Overview of the proposed framework, highlighting a comparative architecture between the Transformer-based model (upper section) and the traditional machine learning model (lower section). (B) A detailed depiction of the Transformer-based model architecture, emphasizing its structure comprising 2 encoder layers and 8 attention heads. The model processes input amino acid sequences via an embedding layer, positional encoding, multi-head self-attention, and feature optimization modules, before classifying the sequences through a fully connected layer.

The model was trained with a batch size of 512, utilizing 80% of the data for training and 20% for testing in each fold. Before training, as shown in S1 Fig, we conducted a comprehensive analysis of amino acid frequencies and peptide length distributions in both the training and test datasets. The results indicated that positive samples exhibited elevated frequencies of amino acids D, E, I, K, N, Q, and Y compared to negative samples, while amino acids A, G, H, L, P, R, S, T, and V occurred less frequently in positive samples. Furthermore, analysis of length distribution demonstrated that positive samples predominantly contained peptides ranging from 10 to 15 amino acids in length, whereas negative samples were primarily 15 to 20 amino acids long. Upon completion of training, deepBCE-parasite achieved an average performance on the training set, with an ACC of 81.91%, an AUC of 0.89, a SE of 71.74%, a SP of 82.09%, and a MCC of 0.64. Rigorous evaluation via 10-fold cross-validation on the balanced BCE dataset yielded an average performance on the independent test set, with an ACC of 80.97%, an AUC of 0.90, a SE of 69.64%, a SP of 89.26%, and a MCC of 0.61. These results are summarized in Table 2.

Table 2. Performance of the Transformer-based deep learning model for predicting BCEs in parasites across training and independent test datasets.

Kernel 10-Fold cross validation Independent testing
ACC (%) AUC SE (%) SP (%) MCC ACC (%) AUC SE (%) SP (%) MCC
K=1 82.64 0.90 72.41 82.87 0.65 81.44 0.90 69.95 89.84 0.62
K=2 81.95 0.89 71.93 81.97 0.64 80.78 0.90 70.31 88.45 0.60
K=3 81.32 0.89 70.67 81.96 0.63 80.33 0.89 67.70 89.57 0.60
K=4 82.38 0.90 72.36 82.40 0.65 80.93 0.90 69.60 89.23 0.61
K=5 81.55 0.89 71.32 81.78 0.63 79.98 0.89 67.70 88.97 0.59
K=6 82.05 0.89 71.61 82.50 0.64 81.59 0.90 69.95 90.10 0.62
K=7 82.23 0.89 72.35 82.11 0.65 81.18 0.90 69.36 89.84 0.61
K=8 81.74 0.89 71.98 81.51 0.64 81.23 0.90 72.21 87.84 0.61
K=9 81.38 0.89 70.83 81.92 0.63 80.68 0.89 68.77 89.40 0.60
K=10 81.88 0.89 71.90 81.86 0.64 81.59 0.90 70.90 89.40 0.62
Mean 81.91 0.89 71.74 82.09 0.64 80.97 0.90 69.64 89.26 0.61

ACC, accuracy; AUC, area under the curve; SE, sensitivity; SP, specificity; MCC, Matthews correlation coefficient

Selection of conventional machine learning models based on handcrafted features

Following the model establishment of the deep learning pipeline, we benchmarked its performance against traditional machine learning models to identify the most appropriate approaches for subsequent BCE screening. As shown in Fig 1A, this study also incorporated a traditional machine learning pipeline, which evaluated commonly used algorithms, including SVM, RF, LGBM, and GNB. To ensure a fair comparison, the same training and testing datasets were used across both deep learning and machine learning. A total of twelve handcrafted features were utilized in this study, including AAC, ASDC, CKSAAGP, CKSAAP, DDE, DPC, GAAC, GDPC, GTPC, PAAC, QSO, and SOCN, which were then applied to the four machine learning algorithms. Hyperparameters for the classifiers were optimized via Grid Search, with the detailed settings provided in S1 Table.

The performance results showed that features derived from the deep learning model consistently outperform those handcrafted features across all four algorithms, yielding a superior ROC curve (AUC=0.90) and highlighting the enhanced evaluation performance of deep learning over machine learning methods (S2 Fig). In addition to AUC, we evaluated several performance metrics, including SP, SE, and MCC. Notably, the RF classifier consistently achieved the highest classification performance compared to GNB, LGBM, and SVM, particularly when employing the CKSAAP, DDE, and DPC feature descriptors (S3 Fig). Specifically, for the independent test set, RF attained SP values of 81.67%, 84.71%, and 83.23% for CKSAAP, DDE, and DPC, respectively (Fig 2A), with corresponding SE values of 74.28%, 75.76%, and 75.07% (Fig 2B), and MCC values of 0.56, 0.61, and 0.58 (Fig 2C). Further analysis of the AUC values across the four classifiers showed the following results: for CKSAAP, the AUC values for GNB, LGBM, RF, and SVM were 0.71, 0.75, 0.78, and 0.73, respectively (Fig 2D); for DDE, the AUC values were 0.70, 0.76, 0.80, and 0.73 (Fig 2E); and for DPC, the AUC values were 0.69, 0.74, 0.79, and 0.73 (Fig 2F). Additional performance metrics for machine learning methods are compiled in S2 Table. These findings further validate the superior performance of the RF classifier, confirming its selection as the optimal predictive models.

Fig 2. Comparison of deep learning feature (deepFeature) with eight top traditional handcrafted features (AAC, ASDC, CKSAAGP, CKSAAP, DDE, DPC, GDPC, and GTPC).

Fig 2

(A-C) Radar plots showing SP, SE, and MCC values for each feature. (D-F) ROC curve comparing the deep learning model (deepBCE) with four traditional machine learning algorithms (GNB, LGBM, RF, and SVM) on CKSAAP, DDE, and DPC features. The x-axis represents False Positive Rate (FPR), and the y-axis represents True Positive Rate (TPR).

Benchmarking custom and existing state-of-the-art models on dual test sets

To ensure a fair comparison between our models and existing linear BCE prediction models, we randomly partitioned two independent test sets, namely test1 and test2. As shown in Table 3 and Fig 3, deepBCE and BCE_RF are specialized models trained exclusively on experimentally validated datasets derived from the human and veterinary parasites, whereas EpiDope and BepiPred3 are general-purpose models trained on a wide range of BCE datasets. In the benchmark evaluation, BepiPred3 achieved remarkable SE with values of 99.46% and 98.31%, but its SP was notably low, reaching only 2.01% and 2.64%. This suggests that BepiPred3 is highly effective in identifying positive BCEs but performs inadequately in distinguishing negative BCEs. In contrast, although EpiDope surpassed BCE_RF in terms of AUC, it underperformed in other critical metrics, including SE and SP. Among our proposed models, deepBCE exhibited superior performance in discriminating positive samples and overall classification accuracy, while BCE_RF demonstrated slightly better capability in identifying negative samples and overall predictive performance. Overall, when evaluated in various metrics, the deepBCE and BCE_RF models, trained on parasite-specific BCE datasets, consistently surpassed EpiDope and BepiPred3, thereby highlighting their superior adaptability and effectiveness for parasite-related prediction tasks.

Table 3. Performance comparison of our models against existing state-of-the-art models.

Model Dataset ACC (%) AUC SE (%) SP (%) MCC
deepBCE test1 82.02 0.72 86.80 77.19 0.64
BCE_RF test1 81.20 0.55 79.20 85.22 0.65
EpiDope test1 59.95 0.64 66.73 53.10 0.20
BepiPred3 test1 50.95 0.57 99.46 2.01 0.07
deepBCE test2 81.34 0.73 86.47 76.27 0.63
BCE_RF test2 80.22 0.54 78.38 85.76 0.64
EpiDope test2 57.49 0.61 65.23 50.26 0.16
BepiPred3 test2 48.87 0.53 98.31 2.64 0.03

Fig 3. Performance comparison between our models and existing models on benchmark datasets.

Fig 3

(A) Results on Test1 dataset. (B) Results on Test2 dataset.

Proteomics-based bioinformatics analysis and prediction of BCEs in the liver fluke Fasciola hepatica

We evaluated the accuracy and utility of our predictive models for identifying BCEs using F. hepatica as a case study. Our models were applied to predict BCEs within F. hepatica proteins using proteomic data and peptide fragments. The complex life cycle of F. hepatica involves an intermediate snail host and a definitive mammalian host, typically cattle or sheep, although humans also act as definitive hosts (Fig 4A) [46]. Leveraging raw proteomic data, we characterized the proteome and expression profiles of F. hepatica across four developmental stages: metacercariae, juvenile fluke (28 days post-infection [dpi] in sheep), immature fluke (59 dpi), and adult fluke (118 dpi). Detailed protein identifications and corresponding expression profiles are provided in S3 Table. Based on these results, the following observations can be drawn.

Fig 4. Proteomic and bioinformatics analysis of F. hepatica across four developmental stages.

Fig 4

(A) The lifecycle of Fasciola spp., demonstrating the developmental stages within the intermediate snail host, environmental phases, and definitive mammalian/human host. The parasite image was adapted from Servier Medical Art (https://smart.servier.com/). (B) Principal component analysis (PCA) of protein expression profiles in metacercaria, juvenile fluke (28 dpi), immature fluke (58 dpi), and adult fluke (118 dpi). (C) Distribution of subcellular localization for proteins capable of generating B-cell epitopes (BCEs), highlighting their potential immunogenic properties. (D) Heatmap depicting protein expression dynamics across the four stages, with color-coded subcellular localization categories shown on the right. (E) Protein-protein interaction network of the leucine aminopeptidase (LAP) protein, generated using the STRING database, emphasizing key interactions and its potential as a vaccine candidate.

The most recent F. hepatica protein reference from the UniProt database was utilized to guide our MS analysis, resulting the identification of 2,165 proteins with a Q-value ≤ 0.01. PCA revealed pronounced differences in expression profiles across developmental stages, with the most substantial divergence observed between the metacercarial stage and the mammalian-host stages (Fig 4B). Additionally, the clustering of biological replicates across stages highlighted the high quality and reproducibility of the sequencing data, establishing a robust foundation for investigating the distribution of BCEs within parasitic proteins.

The deep learning model, combined with RF-based machine learning models, was employed to predict BCEs of F. hepatica. This integrated approach facilitates the identification of high-confidence BCEs in subsequent bioinformatics analysis, with the results summarized in S4 Table. To investigate the pre-cleavage subcellular localization of these source proteins, we systematically mapped their distribution across ten subcellular regions: nucleus, cytoplasm, extracellular space, mitochondria, cell membrane, endoplasmic reticulum, chloroplast, Golgi apparatus, lysosome/vacuole, and peroxisome. Our results indicate that the majority of BCE-source proteins were localized to the cytoplasm (61.24%), followed by mitochondria (10.90%), extracellular space (8.11%), nucleus (7.24%), cell membrane (5.88%), and endoplasmic reticulum (4.83%), with the remaining regions collectively accounting for 1.80% (Fig 4C). Moreover, heatmap clustering demonstrated that proteins carrying BCEs showed the highest expression levels during the metacercariae (MC) stage, while significantly lower expression levels were observed in the mammalian stages (JF, IF, and AF) (Fig 4D).

Here, we selected four proteins: glutathione S-transferase (GST), leucine aminopeptidase (LAP), annexin (ANX), and β-actin, and conducted three-dimensional structure prediction using AlphaFold3 [47], followed by domain analysis using the InterPro web service [48]. As shown in Fig 5, the GST protein contains the GST_NTER and GST_C_Mu domains, yielding seven BCEs; the LAP protein harbors a Peptidase_M17 domain, resulting in eight BCEs; the ANX protein includes four ANNEXIN_2 domains, producing ten BCEs; and the β-actin protein contains a single domain, generating four BCEs. Notably, in prior animal studies, the LAP protein demonstrated an 89% vaccine efficacy in sheep [49]. Protein interaction predictions further revealed that the LAP protein interacts with 27 high-confidence proteins, each of which also carries BCEs (Fig 4E), suggesting promising pathways for the exploration of F. hepatica immune targets.

Fig 5. Predicted 3D structural models of potential F. hepatica vaccine molecules, highlighting the spatial distribution of putative BCEs.

Fig 5

(A–D) Structural models for Glutathione transferase, leucine aminopeptidase, annexin, and β-actin, respectively. The predicted template modelling (pTM) score from AlphaFold3, which reflects prediction quality, is provided in the legend (higher value indicates greater model reliability). Predicted functional domains are color-coded, while non-domain regions are shown in gray.

Dot-blot immunoassays

Structural domain prediction of the LAP protein identified the presence of a Peptidase_M17 domain, confirming its classification as a canonical member of the M17 family (Fig 5B). To assess its evolutionary conservation, a multiple sequence alignment was conducted using MAFFT [50], incorporating LAP protein sequences from diverse trematode species, including Fasciola gigantica, Clonorchis sinensis, Opisthorchis felineus, Schistosoma japonicum, and Schistosoma mansoni. The alignment indicated a high degree of conservation in predicted BCEs between F. hepatica and F. gigantica. In contrast, comparisons with other trematode species showed that only the sequence “DIGGSDPER” was fully conserved, while other predicted BCEs exhibited significant amino acid variability (Fig 6A).

Fig 6. Sequence alignment and experimental validation of the LAP protein and its putative BCEs.

Fig 6

(A) Multiple sequence alignment of F. hepatica LAP with five other trematode species. Predicted BCEs identified by our AI models are highlighted with red boxes. (B) Dot-blot immunoassay validation of predicted BCEs using serum from F. hepatica-positive animal, with F. hepatica-negative serum and a synthetic human peptide as controls.

To experimentally validate the predicted BCEs, eight peptides derived from the LAP protein was chemically synthesized and subjected to dot-blot immunoassays. Immunoreactivity analysis revealed that seven of the predicted BCEs (P2, P3, P4, P5, P6, P7, and P8) produced distinct dot signals in F. hepatica-positive ovine serum, while no dot signals were observed in synthetic human peptide controls and F. hepatica-negative ovine serum (Fig 6B and Table 4). In contrast, P1 failed to generate a detectable signal, suggesting a false-positive prediction by the model. Among the validated BCEs, P2, P4, P5, P6, and P8 were located within the predicted structural domain, whereas P3 and P7 were situated outside the domain. These findings demonstrate that the LAP protein exhibits multi-site antigenicity and has the potential to elicit host immune responses through distinct structural regions.

Table 4. Experimental validation of predicted BCEs derived from the LAP protein using dot-blot immunoassay.

Peptide name Sequence Position on LAP Dot-blot
P1 GKTEYEDILQCNNLPSSATPR 439-459 -
P2 DIGGSDPER 187-195 +
P3 LIYSPTGALNTDTADIR 67-83 +
P4 ISDMAEISTVR 420-430 +
P5 GITYDTGGADIK 270-281 +
P6 EDYEFNR 432-438 +
P7 EDTKGSDELVLR 162-173 +
P8 IVDYLK 201-206 +
Control RSRTPSLPTPPTREP NA -

+, presence of a dot signal; -, absence of a dot signal

Discussion

The identification of eukaryotic antigenic peptides derived from parasitic proteins previously relied on machine learning approaches that utilize handcrafted feature extraction, such as LGBM and stack-based models [5153]. Although these approaches have demonstrated good prediction accuracy, their dependence on manual feature engineering often limits their ability to fully capture the complexity of antigenic peptide biology. In contrast, Transformer-based deep learning approaches have quickly emerged as a powerful alternative for protein sequence analysis, offering significant advantages over other deep learning methods, such as RNN-based models, which are constrained by sequential processing and struggle to model long-range dependencies [5456]. By leveraging the Transformer architecture, the developed model effectively extracts hierarchical features that integrate amino acid properties and positional information through self-attention mechanisms. This approach not only captures both local and long-range residue interactions but also enables efficient parallel processing of large, variable-length sequences without the need for extensive preprocessing.

In this study, by integrating Transformer-based models with traditional machine learning models and leveraging publicly available F. hepatica proteomic data, potential BCEs were predicted and experimentally validated for those targeted by the LAP protein. Notably, seven of the eight synthesized peptides were confirmed as immunoreactive BCEs, demonstrating the high predictive accuracy and computational efficiency of our models in screening epitope candidates from large-scale MS-based datasets. The LAP protein was selected for its well-documented potential as a vaccine candidate that has been supported by extensive experimental evidence [5761]. Our findings reveal the sequence characteristics and evolutionary relationships of LAP-driven BCEs, providing a bioinformatics foundation for peptide vaccine and diagnostic tool development. Additionally, proteins such as GST, ANX, and β-actin were identified as potential vaccine candidates based on their immunoprotective effects [62]. However, further experimental validation is required to confirm the antibody-binding efficacy of the BCEs derived from these candidate proteins.

Importantly, our model is not restricted to the liver fluke F. hepatica but can be broadly applied to predict BCEs across a wide range of human and veterinary parasites, including both protozoan and helminth species. For example, in veterinary parasitology, this approach enables high-throughput screening of potential BCEs, allowing synthesized peptides to be combined with adjuvants and evaluated for their immunoprotective potential against pathogens [63]. Furthermore, integrating BCEs into diagnostic assays offers the potential to enhance the accuracy of infection detection, allowing for the differentiation between past and active infections, as well as the evaluation of vaccine efficacy. Although the current models demonstrate considerable predictive ability, further advancements are needed to improve the prediction of BCEs and to better capture complex structural interactions.

Our study only focuses on the development of linear BCE prediction models that leverage the linear sequence features of parasitic peptides to identify potential BCEs. Linear epitopes consist of contiguous amino acid sequences, which offer an advantage in model development due to their direct reliance on sequence features [64]. As a result, these epitopes are relatively straightforward to classify and experimentally validate. Consequently, linear epitope predictions are often prioritized in vaccine development and diagnostic design, particularly for rapid antigenic peptide identification [6569]. However, this study does not address conformational epitope prediction, thus limiting its applicability in natural biological systems where BCEs typically adopt three-dimensional (3D) conformational structures. Conformational epitopes, formed by non-contiguous amino acid residues, depend on the overall protein folding and spatial exposure [70,71]. This dependence introduces additional challenges to epitope prediction, as it requires modeling complex structural interactions.

With advancements in AI-facilitated protein structure prediction, future studies on parasitic infections could benefit from protein models that integrate both linear and conformational epitope recognition. Open-source tools like AlphaFold, combined with structure determination techniques like X-ray crystallography and cryo-electron microscopy, would provide high-accuracy structures [72]. These advancements enable the identification of conformational epitopes, which are critical for simulating real-world antigen-antibody interactions. Future research should explore the development of multimodal predictive frameworks that integrate 3D structural information with sequence-based approach to improve the prediction of parasite-specific BCEs. Additionally, state-of-the-art deep learning architectures, such as 3D convolutional neural networks or graph neural networks [73,74], should be explored in the future to identify critical contact sites and surface-exposed features within protein structures.

Conclusions

In this study, we developed a Transformer-based deep learning model alongside several traditional machine learning models for predicting BCEs in parasitic protein. Performance comparisons demonstrated that the Transformer-based model achieved superior predictive capability, with an independent test performance characterized by an ACC of 80.97%, an AUC of 0.90, a SE of 69.64%, a SP of 89.26%, and a MCC of 0.61. Among the traditional machine learning models, the RF algorithm performed best, particularly when using the DDE feature descriptor, yielding an ACC of 80.23%, an AUC of 0.88, a SE of 75.76%, a SP of 84.71%, and a MCC of 0.61. In addition, this study leveraged proteomic data from F. hepatica, identifying seven BCEs through dot-blot immunoassay targeting the LAP protein, which specifically bind to F. hepatica-positive ovine serum antibodies, underscoring their potential as vaccine antigen candidates. The developed models show versatility, suitable for predicting BCEs in various parasitic pathogens, including protozoa and helminths. To enhance predictive accuracy and robustness in experimental validation, we recommend combining multiple models for BCE prediction in parasitic peptides.

Supporting information

S1 Fig. Amino acid frequency distribution and length statistics of peptides.

(A-B) Amino acid frequency in positive and negative samples from the training and testing datasets. (C-D) Peptide length distribution in positive and negative samples from the training and testing datasets.

(TIF)

pntd.0012985.s001.tif (1.1MB, tif)
S2 Fig. Comparison of ROC curve between deep learning and 12 handcrafted features using four machine learning algorithms.

(TIF)

pntd.0012985.s002.tif (1.5MB, tif)
S3 Fig. Radar chart comparing the performance of deep learning features (deepFeature) with 12 handcrafted features across four machine learning algorithms.

The chart highlights their respective performance in SP, SE, and MCC metrics.

(TIF)

pntd.0012985.s003.tif (1.5MB, tif)
S1 Table. Parameter settings for deep learning and machine learning models used in this study.

(XLSX)

pntd.0012985.s004.xlsx (9.6KB, xlsx)
S2 Table. Performance of conventional machine learning models with handcrafted features for B-cell epitope prediction in parasites.

(XLSX)

pntd.0012985.s005.xlsx (23.1KB, xlsx)
S3 Table. Bioinformatic analysis of proteomic features and expression profiles of F. hepatica across four developmental stages.

(XLSX)

pntd.0012985.s006.xlsx (364.7KB, xlsx)
S4 Table. The prediction results of B-cell epitope in F. hepatica using deep learning and machine learning models.

(XLSX)

pntd.0012985.s007.xlsx (128.3KB, xlsx)

Data Availability

The MaxQuant search results, based on entries from the UniProt database, can be downloaded from the Zenodo repository at https://doi.org/10.5281/zenodo.14065139. The source code for the deepBCE-Parasite software is available on GitHub at https://github.com/RuiSiHu/deepBCE-Parasite and is also archived in the Zenodo repository at https://doi.org/10.5281/zenodo.14907455.

Funding Statement

This work was supported by the National Natural Science Foundation of China (Grant No. 62202083 to RSH), the China Postdoctoral Science Foundation (Grant No. 2022M710618 to RSH), and the High-level Talent Research Start-up Project of Sichuan University of Arts and Science (Grant No. 2024GCC16Z to RSH). The funders had no role in study design, data collection, and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Torgerson PR, Devleesschauwer B, Praet N, Speybroeck N, Willingham AL, Kasuga F, et al. World Health Organization estimates of the global and regional disease burden of 11 foodborne parasitic diseases, 2010: a data synthesis. PLoS Med. 2015;12(12):e1001920. doi: 10.1371/journal.pmed.1001920 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Morgan ER, Aziz N-AA, Blanchard A, Charlier J, Charvet C, Claerebout E, et al. 100 questions in livestock helminthology research. Trends Parasitol. 2019;35(1):52–71. doi: 10.1016/j.pt.2018.10.006 [DOI] [PubMed] [Google Scholar]
  • 3.Cwiklinski K, O’Neill SM, Donnelly S, Dalton JP. A prospective view of animal and human Fasciolosis. Parasite Immunol. 2016;38(9):558–68. doi: 10.1111/pim.12343 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Mas-Coma S. Epidemiology of fascioliasis in human endemic areas. J Helminthol. 2005;79(3):207–16. doi: 10.1079/joh2005296 [DOI] [PubMed] [Google Scholar]
  • 5.Hodgkinson JE, Kaplan RM, Kenyon F, Morgan ER, Park AW, Paterson S, et al. Refugia and anthelmintic resistance: Concepts and challenges. Int J Parasitol Drugs Drug Resist. 2019;10:51–7. doi: 10.1016/j.ijpddr.2019.05.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Pramanik PK, Alam MN, Roy Chowdhury D, Chakraborti T. Drug resistance in protozoan parasites: an incessant wrestle for survival. J Glob Antimicrob Resist. 2019;18:1–11. doi: 10.1016/j.jgar.2019.01.023 [DOI] [PubMed] [Google Scholar]
  • 7.Getzoff ED, Tainer JA, Lerner RA, Geysen HM. The chemistry and mechanism of antibody binding to protein antigens. Adv Immunol. 1988;43:1–98. doi: 10.1016/s0065-2776(08)60363-6 [DOI] [PubMed] [Google Scholar]
  • 8.Potocnakova L, Bhide M, Pulzova LB. An Introduction to B-cell epitope mapping and in silico epitope prediction. J Immunol Res. 2016;2016:6760830. doi: 10.1155/2016/6760830 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Van Regenmortel MHV. What is a B-cell epitope?. Methods Mol Biol. 2009;524:3–20. doi: 10.1007/978-1-59745-450-6_1 [DOI] [PubMed] [Google Scholar]
  • 10.Bastos LM, Macêdo AG Jr, Silva MV, Santiago FM, Ramos ELP, Santos FAA, et al. Toxoplasma gondii-derived synthetic peptides containing B- and T-cell epitopes from GRA2 protein are able to enhance mice survival in a model of experimental toxoplasmosis. Front Cell Infect Microbiol. 2016;6:59. doi: 10.3389/fcimb.2016.00059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Buffoni L, Garza-Cuartero L, Pérez-Caballero R, Zafra R, Javier Martínez-Moreno F, Molina-Hernández V, et al. Identification of protective peptides of Fasciola hepatica-derived cathepsin L1 (FhCL1) in vaccinated sheep by a linear B-cell epitope mapping approach. Parasit Vectors. 2020;13(1):390. doi: 10.1186/s13071-020-04260-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Garza-Cuartero L, Geurden T, Mahan SM, Hardham JM, Dalton JP, Mulcahy G. Antibody recognition of cathepsin L1-derived peptides in Fasciola hepatica-infected and/or vaccinated cattle and identification of protective linear B-cell epitopes. Vaccine. 2018;36(7):958–68. doi: 10.1016/j.vaccine.2018.01.020 [DOI] [PubMed] [Google Scholar]
  • 13.Mu Y, Gordon CA, Olveda RM, Ross AG, Olveda DU, Marsh JM, et al. Identification of a linear B-cell epitope on the Schistosoma japonicum saposin protein, SjSAP4: Potential as a component of a multi-epitope diagnostic assay. PLoS Negl Trop Dis. 2022;16(7):e0010619. doi: 10.1371/journal.pntd.0010619 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Ansari HR, Raghava GP. Identification of conformational B-cell epitopes in an antigen from its primary sequence. Immunome Res. 2010;6:6. doi: 10.1186/1745-7580-6-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.da Silva BM, Myung Y, Ascher DB, Pires DEV. epitope3D: a machine learning method for conformational B-cell epitope prediction. Brief Bioinform. 2022;23(1):bbab423. doi: 10.1093/bib/bbab423 [DOI] [PubMed] [Google Scholar]
  • 16.Høie MH, Gade FS, Johansen JM, Würtzen C, Winther O, Nielsen M, et al. DiscoTope-3.0: improved B-cell epitope prediction using inverse folding latent representations. Front Immunol. 2024;15:1322712. doi: 10.3389/fimmu.2024.1322712 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ivanisenko NV, Shashkova TI, Shevtsov A, Sindeeva M, Umerenkov D, Kardymon O. SEMA 2.0: web-platform for B-cell conformational epitopes prediction using artificial intelligence. Nucleic Acids Res. 2024;52(W1):W533–9. doi: 10.1093/nar/gkae386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Liang S, Zheng D, Standley DM, Yao B, Zacharias M, Zhang C. EPSVR and EPMeta: prediction of antigenic epitopes using support vector regression and multiple server results. BMC Bioinformatics. 2010;11:381. doi: 10.1186/1471-2105-11-381 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ponomarenko J, Bui H-H, Li W, Fusseder N, Bourne PE, Sette A, et al. ElliPro: a new structure-based tool for the prediction of antibody epitopes. BMC Bioinformatics. 2008;9:514. doi: 10.1186/1471-2105-9-514 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhou C, Chen Z, Zhang L, Yan D, Mao T, Tang K, et al. SEPPA 3.0-enhanced spatial epitope prediction enabling glycoprotein antigens. Nucleic Acids Res. 2019;47(W1):W388–94. doi: 10.1093/nar/gkz413 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Alghamdi W, Attique M, Alzahrani E, Ullah MZ, Khan YD. LBCEPred: a machine learning model to predict linear B-cell epitopes. Brief Bioinform. 2022;23(3):bbac035. doi: 10.1093/bib/bbac035 [DOI] [PubMed] [Google Scholar]
  • 22.Clifford JN, Høie MH, Deleuran S, Peters B, Nielsen M, Marcatili P. BepiPred-3.0: Improved B-cell epitope prediction using protein language models. Protein Sci. 2022;31(12):e4497. doi: 10.1002/pro.4497 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Collatz M, Mock F, Barth E, Hölzer M, Sachse K, Marz M. EpiDope: a deep neural network for linear B-cell epitope prediction. Bioinformatics. 2021;37(4):448–55. doi: 10.1093/bioinformatics/btaa773 [DOI] [PubMed] [Google Scholar]
  • 24.Manavalan B, Govindaraj RG, Shin TH, Kim MO, Lee G. iBCE-EL: a new ensemble learning framework for improved linear B-cell epitope prediction. Front Immunol. 2018;9:1695. doi: 10.3389/fimmu.2018.01695 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Qi Y, Zheng P, Huang G. DeepLBCEPred: A Bi-LSTM and multi-scale CNN-based deep learning method for predicting linear B-cell epitopes. Front Microbiol. 2023;14:1117027. doi: 10.3389/fmicb.2023.1117027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Saha S, Raghava GPS. Prediction of continuous B-cell epitopes in an antigen using recurrent neural network. Proteins. 2006;65(1):40–8. doi: 10.1002/prot.21078 [DOI] [PubMed] [Google Scholar]
  • 27.Yao B, Zhang L, Liang S, Zhang C. SVMTriP: a method to predict antigenic epitopes using support vector machine to integrate tri-peptide similarity and propensity. PLoS One. 2012;7(9):e45152. doi: 10.1371/journal.pone.0045152 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Howe KL, Bolt BJ, Shafie M, Kersey P, Berriman M. WormBase ParaSite - a comprehensive resource for helminth genomics. Mol Biochem Parasitol. 2017;215:2–10. doi: 10.1016/j.molbiopara.2016.11.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Warrenfeltz S, Basenko EY, Crouch K, Harb OS, Kissinger JC, Roos DS, et al. EuPathDB: the eukaryotic pathogen genomics database resource. Methods Mol Biol. 2018;1757:69–113. doi: 10.1007/978-1-4939-7737-6_5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Cwiklinski K, Dalton JP. Advances in Fasciola hepatica research using “omics” technologies. Int J Parasitol. 2018;48(5):321–31. doi: 10.1016/j.ijpara.2017.12.001 [DOI] [PubMed] [Google Scholar]
  • 31.Hu R-S, Zhang F-K, Elsheikha HM, Ma Q-N, Ehsan M, Zhao Q, et al. Proteomic profiling of the liver, hepatic lymph nodes, and spleen of buffaloes infected with Fasciola gigantica. Pathogens. 2020;9(12):982. doi: 10.3390/pathogens9120982 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Hu R-S, Zhang F-K, Ma Q-N, Ehsan M, Zhao Q, Zhu X-Q. Transcriptomic landscape of hepatic lymph nodes, peripheral blood lymphocytes and spleen of swamp buffaloes infected with the tropical liver fluke Fasciola gigantica. PLoS Negl Trop Dis. 2022;16(3):e0010286. doi: 10.1371/journal.pntd.0010286 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Vita R, Mahajan S, Overton JA, Dhanda SK, Martini S, Cantrell JR, et al. The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res. 2019;47(D1):D339–43. doi: 10.1093/nar/gky1006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Saha S, Bhasin M, Raghava GPS. Bcipep: a database of B-cell epitopes. BMC Genomics. 2005;6:79. doi: 10.1186/1471-2164-6-79 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Fu L, Niu B, Zhu Z, Wu S, Li W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012;28(23):3150–2. doi: 10.1093/bioinformatics/bts565 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Chen Z, Zhao P, Li C, Li F, Xiang D, Chen Y-Z, et al. iLearnPlus: a comprehensive and automated machine-learning platform for nucleic acid and protein sequence analysis, prediction and visualization. Nucleic Acids Res. 2021;49(10):e60. doi: 10.1093/nar/gkab122 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Bolón-Canedo V, Sánchez-Maroño N, Alonso-Betanzos A. Feature selection for high-dimensional data. Prog Artif Intell. 2016;5(2):65–75. doi: 10.1007/s13748-015-0080-y [DOI] [Google Scholar]
  • 38.Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W. Lightgbm: A highly efficient gradient boosting decision tree. Advance neural information processing systems. n.d.;30:3146–54. [Google Scholar]
  • 39.Jiao S, Ye X, Ao C, Sakurai T, Zou Q, Xu L. Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites. Bioinformatics. 2023;39(11):btad627. doi: 10.1093/bioinformatics/btad627 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Jiao S, Ye X, Sakurai T, Zou Q, Liu R. Integrated convolution and self-attention for improving peptide toxicity prediction. Bioinformatics. 2024;40(5):btae297. doi: 10.1093/bioinformatics/btae297 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Yan K, Lv H, Guo Y, Peng W, Liu B. sAMPpred-GAT: prediction of antimicrobial peptide by graph attention network and predicted peptide structure. Bioinformatics. 2023;39(1):btac715. doi: 10.1093/bioinformatics/btac715 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Xu J, Wu L, Sun Y, Wei Y, Zheng L, Zhang J, et al. Proteomics and bioinformatics analysis of Fasciola hepatica somatic proteome in different growth phases. Parasitol Res. 2020;119(9):2837–50. doi: 10.1007/s00436-020-06833-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Cox J, Mann M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat Biotechnol. 2008;26(12):1367–72. doi: 10.1038/nbt.1511 [DOI] [PubMed] [Google Scholar]
  • 44.Thumuluri V, Almagro Armenteros JJ, Johansen AR, Nielsen H, Winther O. DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Res. 2022;50(W1):W228–34. doi: 10.1093/nar/gkac278 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Szklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):D638–46. doi: 10.1093/nar/gkac1000 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Hu R-S, Zhang X-X, Ma Q-N, Elsheikha HM, Ehsan M, Zhao Q, et al. Differential expression of microRNAs and tRNA fragments mediate the adaptation of the liver fluke Fasciola gigantica to its intermediate snail and definitive mammalian hosts. Int J Parasitol. 2021;51(5):405–14. doi: 10.1016/j.ijpara.2020.10.009 [DOI] [PubMed] [Google Scholar]
  • 47.Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630(8016):493–500. doi: 10.1038/s41586-024-07487-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Blum M, Andreeva A, Florentino LC, Chuguransky SR, Grego T, Hobbs E, et al. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res. 2025;53(D1):D444–56. doi: 10.1093/nar/gkae1082 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Piacenza L, Acosta D, Basmadjian I, Dalton JP, Carmona C. Vaccination with cathepsin L proteinases and with leucine aminopeptidase induces high levels of protection against fascioliasis in sheep. Infect Immun. 1999;67(4):1954–61. doi: 10.1128/IAI.67.4.1954-1961.1999 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013;30(4):772–80. doi: 10.1093/molbev/mst010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Charoenkwan P, Schaduangrat N, Pham NT, Manavalan B, Shoombuatong W. Pretoria: An effective computational approach for accurate and high-throughput identification of CD8+ t-cell epitopes of eukaryotic pathogens. Int J Biol Macromol. 2023;238:124228. doi: 10.1016/j.ijbiomac.2023.124228 [DOI] [PubMed] [Google Scholar]
  • 52.Hu R-S, Hesham AE-L, Zou Q. Machine learning and its applications for protozoal pathogens and protozoal infectious diseases. Front Cell Infect Microbiol. 2022;12:882995. doi: 10.3389/fcimb.2022.882995 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Hu R-S, Wu J, Zhang L, Zhou X, Zhang Y. CD8TCEI-EukPath: a novel predictor to rapidly identify CD8+ T-cell epitopes of eukaryotic pathogens using a hybrid feature selection approach. Front Genet. 2022;13:935989. doi: 10.3389/fgene.2022.935989 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Chandra A, Tünnermann L, Löfstedt T, Gratz R. Transformer-based deep learning for predicting protein properties in the life sciences. Elife. 2023;12:e82819. doi: 10.7554/eLife.82819 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Le NQK. Leveraging transformers-based language models in proteome bioinformatics. Proteomics. 2023;23(23–24):e2300011. doi: 10.1002/pmic.202300011 [DOI] [PubMed] [Google Scholar]
  • 56.Mswahili ME, Jeong Y-S. Transformer-based models for chemical SMILES representation: A comprehensive literature review. Heliyon. 2024;10(20):e39038. doi: 10.1016/j.heliyon.2024.e39038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Acosta D, Cancela M, Piacenza L, Roche L, Carmona C, Tort JF. Fasciola hepatica leucine aminopeptidase, a promising candidate for vaccination against ruminant fasciolosis. Mol Biochem Parasitol. 2008;158(1):52–64. doi: 10.1016/j.molbiopara.2007.11.011 [DOI] [PubMed] [Google Scholar]
  • 58.Checa J, Salazar C, Goyeche A, Rivera M, Silveira F, Maggioli G. A promising new target to control fasciolosis: Fasciola hepatica leucine aminopeptidase 2. Vet Parasitol. 2023;320:109959. doi: 10.1016/j.vetpar.2023.109959 [DOI] [PubMed] [Google Scholar]
  • 59.Hernández-Guzmán K, Sahagún-Ruiz A, Vallecillo AJ, Cruz-Mendoza I, Quiroz-Romero H. Construction and evaluation of a chimeric protein made from Fasciola hepatica leucine aminopeptidase and cathepsin L1. J Helminthol. 2016;90(1):7–13. doi: 10.1017/S0022149X14000686 [DOI] [PubMed] [Google Scholar]
  • 60.Ortega-Vargas S, Espitia C, Sahagún-Ruiz A, Parada C, Balderas-Loaeza A, Villa-Mancera A, et al. Moderate protection is induced by a chimeric protein composed of leucine aminopeptidase and cathepsin L1 against Fasciola hepatica challenge in sheep. Vaccine. 2019;37(24):3234–40. doi: 10.1016/j.vaccine.2019.04.067 [DOI] [PubMed] [Google Scholar]
  • 61.Salazar C, Tort JF, Carmona C. Design of a peptide-carrier vaccine based on the highly immunogenic Fasciola hepatica leucine aminopeptidase. Methods Mol Biol. 2020;2137:191–204. doi: 10.1007/978-1-0716-0475-5_14 [DOI] [PubMed] [Google Scholar]
  • 62.Toet H, Piedrafita DM, Spithill TW. Liver fluke vaccines in ruminants: strategies, progress and future opportunities. Int J Parasitol. 2014;44(12):915–27. doi: 10.1016/j.ijpara.2014.07.011 [DOI] [PubMed] [Google Scholar]
  • 63.Ehsan M, Hu R-S, Liang Q-L, Hou J-L, Song X, Yan R, et al. Advances in the development of anti-Haemonchus contortus vaccines: challenges, opportunities, and perspectives. Vaccines (Basel). 2020;8(3):555. doi: 10.3390/vaccines8030555 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Bahrami AA, Payandeh Z, Khalili S, Zakeri A, Bandehpour M. Immunoinformatics: in silico approaches and computational design of a multi-epitope, immunogenic protein. Int Rev Immunol. 2019;38(6):307–22. doi: 10.1080/08830185.2019.1657426 [DOI] [PubMed] [Google Scholar]
  • 65.Baptista B de O, Souza ABL de, Oliveira LS de, Souza HADS de, Barros JP de, Queiroz LT de, et al. B-cell epitope mapping of the Plasmodium falciparum malaria vaccine candidate GMZ2.6c in a naturally exposed population of the brazilian amazon. Vaccines (Basel). 2023;11(2):446. doi: 10.3390/vaccines11020446 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Durante IM, La Spina PE, Carmona SJ, Agüero F, Buscaglia CA. High-resolution profiling of linear B-cell epitopes from mucin-associated surface proteins (MASPs) of Trypanosoma cruzi during human infections. PLoS Negl Trop Dis. 2017;11(9):e0005986. doi: 10.1371/journal.pntd.0005986 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Guedes RLM, Rodrigues CMF, Coatnoan N, Cosson A, Cadioli FA, Garcia HA, et al. A comparative in silico linear B-cell epitope prediction and characterization for South American and African Trypanosoma vivax strains. Genomics. 2019;111(3):407–17. doi: 10.1016/j.ygeno.2018.02.017 [DOI] [PubMed] [Google Scholar]
  • 68.Javadi Mamaghani A, Fathollahi A, Spotin A, Ranjbar MM, Barati M, Aghamolaie S, et al. Candidate antigenic epitopes for vaccination and diagnosis strategies of Toxoplasma gondii infection: A review. Microb Pathog. 2019;137:103788. doi: 10.1016/j.micpath.2019.103788 [DOI] [PubMed] [Google Scholar]
  • 69.Mendes TA de O, Reis Cunha JL, de Almeida Lourdes R, Rodrigues Luiz GF, Lemos LD, dos Santos ARR, et al. Identification of strain-specific B-cell epitopes in Trypanosoma cruzi using genome-scale epitope prediction and high-throughput immunoscreening with peptide arrays. PLoS Negl Trop Dis. 2013;7(10):e2524. doi: 10.1371/journal.pntd.0002524 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Cretich M, Gori A, D’Annessa I, Chiari M, Colombo G. Peptides for infectious diseases: from probe design to diagnostic microarrays. Antibodies (Basel). 2019;8(1):23. doi: 10.3390/antib8010023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Durairaj R, Pageat P, Bienboire-Frosini C. Impact of semiochemicals binding to Fel d 1 on Its 3D conformation and predicted B-cell epitopes using computational approaches. Int J Mol Sci. 2023;24(14):11685. doi: 10.3390/ijms241411685 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Sela-Culang I, Kunik V, Ofran Y. The structural basis of antibody-antigen recognition. Front Immunol. 2013;4:302. doi: 10.3389/fimmu.2013.00302 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Ramakrishnan G, Baakman C, Heijl S, Vroling B, van Horck R, Hiraki J, et al. Understanding structure-guided variant effect predictions using 3D convolutional neural networks. Front Mol Biosci. 2023;10:1204157. doi: 10.3389/fmolb.2023.1204157 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Strokach A, Becerra D, Corbi-Verge C, Perez-Riba A, Kim PM. Fast and flexible protein design using deep graph neural networks. Cell Syst. 2020;11(4):402–411.e4. doi: 10.1016/j.cels.2020.08.016 [DOI] [PubMed] [Google Scholar]
PLoS Negl Trop Dis. doi: 10.1371/journal.pntd.0012985.r002

Decision Letter 0

jong-Yil Chai, Aysegul Taylan Ozkan

13 Jan 2025

Improved B-cell epitope identification in parasitic pathogens through a Transformer-driven model: Proof-of-concept in Fasciola hepatica

PLOS Neglected Tropical Diseases

Dear Dr. Hu,

Thank you for submitting your manuscript to PLOS Neglected Tropical Diseases. After careful consideration, we feel that it has merit but does not fully meet PLOS Neglected Tropical Diseases's publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript within 60 days Mar 14 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosntds@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pntd/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

* A rebuttal letter that responds to each point raised by the editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.

* A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

* An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, competing interests statement, or data availability statement, please make these updates within the submission form at the time of resubmission. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

We look forward to receiving your revised manuscript.

Kind regards,

Aysegul Taylan Ozkan, M.D., Ph.D.,

Academic Editor

PLOS Neglected Tropical Diseases

Jong-Yil Chai

Section Editor

PLOS Neglected Tropical Diseases

Shaden Kamhawi

co-Editor-in-Chief

PLOS Neglected Tropical Diseases

orcid.org/0000-0003-4304-636XX

Paul Brindley

co-Editor-in-Chief

PLOS Neglected Tropical Diseases

orcid.org/0000-0003-1765-0002

Journal Requirements:

1) Please ensure that the CRediT author contributions listed for every co-author are completed accurately and in full.

At this stage, the following Authors/Authors require contributions: Rui-Si Hu, Kui Gu, Muhammad Ehsan, Sayed Haidar Abbas Raza, and Chun-Ren Wang. Please ensure that the full contributions of each author are acknowledged in the "Add/Edit/Remove Authors" section of our submission form.

The list of CRediT author contributions may be found here: https://journals.plos.org/plosntds/s/authorship#loc-author-contributions

2) Please ensure that all Table files have corresponding citations and legends within the manuscript. Currently, Table 2 in your submission file inventory does not have an in-text citation. Please include the in-text citation of the table.

3) Tables should not be uploaded as individual files. Please remove these files and include the Tables in your manuscript file as editable, cell-based objects. For more information about how to format tables, see our guidelines:

https://journals.plos.org/plosntds/s/tables 

4) Some material included in your submission may be copyrighted. According to PLOSu2019s copyright policy, authors who use figures or other material (e.g., graphics, clipart, maps) from another author or copyright holder must demonstrate or obtain permission to publish this material under the Creative Commons Attribution 4.0 International (CC BY 4.0) License used by PLOS journals. Please closely review the details of PLOSu2019s copyright requirements here: PLOS Licenses and Copyright. If you need to request permissions from a copyright holder, you may use PLOS's Copyright Content Permission form.

Please respond directly to this email and provide any known details concerning your material's license terms and permissions required for reuse, even if you have not yet obtained copyright permissions or are unsure of your material's copyright compatibility. Once you have responded and addressed all other outstanding technical requirements, you may resubmit your manuscript within Editorial Manager. 

Potential Copyright Issues:

i) Figures 1a, and 3a. Please confirm whether you drew the images / clip-art within the figure panels by hand. If you did not draw the images, please provide (a) a link to the source of the images or icons and their license / terms of use; or (b) written permission from the copyright holder to publish the images or icons under our CC BY 4.0 license. Alternatively, you may replace the images with open source alternatives. See these open source resources you may use to replace images / clip-art:

- https://commons.wikimedia.org

- https://openclipart.org/.

5) Please amend your detailed Financial Disclosure statement. This is published with the article. It must therefore be completed in full sentences and contain the exact wording you wish to be published.

1) State the initials, alongside each funding source, of each author to receive each grant. For example: "This work was supported by the National Institutes of Health (####### to AM; ###### to CJ) and the National Science Foundation (###### to AM)."

2) State what role the funders took in the study. If the funders had no role in your study, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

6) Your current Financial Disclosure states, "This work was supported by the National Natural Science Foundation of China [62202083]; the China Postdoctoral Science Foundation [2022M710618]; and the High-level Talent Research Start-up Project of Sichuan University of Arts and Science [WL085267]. ".

However, your funding information on the submission form indicates receiving one fund. Please ensure that the funders and grant numbers match between the Financial Disclosure field and the Funding Information tab in your submission form. Note that the funders must be provided in the same order in both places as well.                                                  . 

Please indicate by return email the full and correct funding information for your study and confirm the order in which funding contributions should appear. Please be sure to indicate whether the funders played any role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript.

7) Please amend the label of Figure S3 in the online submission form as it is currently uploaded as Figure S2.

Comments for the authors:

Please note that the reviews are uploaded as attachments.

Reviewers' Comments:

Reviewer's Responses to Questions

Key Review Criteria Required for Acceptance?

As you describe the new analyses required for acceptance, please consider the following:

Methods

-Are the objectives of the study clearly articulated with a clear testable hypothesis stated?

-Is the study design appropriate to address the stated objectives?

-Is the population clearly described and appropriate for the hypothesis being tested?

-Is the sample size sufficient to ensure adequate power to address the hypothesis being tested?

-Were correct statistical analysis used to support conclusions?

-Are there concerns about ethical or regulatory requirements being met?

Reviewer #1: Please see the more details annex file.

Reviewer #2: Authors have presented a well-executed study of predicting B-cell epitopes in F. hepatica using a deep learning model developed via transformer driven approach. Detailed comments are provided in the report.

Reviewer #3: The Methods section provides a comprehensive explanation of the Transformer architecture and methodology, but its technical density may overwhelm non-technical readers. For readers without a deep learning background, in order to improve accessibility and clarity, simplify the explanation of the Transformer architecture and consider including a flowchart or visual chart to illustrate key components such as attention layers and embeddings.

Also, please expand on the data preprocessing steps by describing how the dataset was balanced (e.g., oversampling) and detailing the processes for removing duplicates or noise.

Provide justification for the chosen hyperparameters, such as stating, "The choice of two encoder layers and eight attention heads was informed by prior studies and optimization experiments." Additionally, clarify the validation strategy by explicitly mentioning whether k-fold CV or an independent test set was used, and discuss measures taken to prevent overfitting. Finally, consider rephrasing overly technical sentences for broader readability. For instance, revise "The model utilized self-attention mechanisms to capture sequence dependencies" to "The model's attention mechanisms effectively identified patterns within the sequence data." These changes will make the section more engaging and accessible while maintaining its technical rigor.

Reviewer #4: This manuscript by Hu et al describes the development of a deep learning model, deepBCE-Parasite, which is applied in the identification of linear B-cell epitopes in parasitic pathogens, with a focus on Fasciola hepatica, a parasite of major health and economic importance. The authors have applied the Transformer architecture in designing a model that predicts BCEs from amino acid sequences using state-of-the-art self-attention mechanisms. In contrast, the deep learning approach significantly outperforms traditional machine learning methods such as Support Vector Machines and Random Forests with up to an accuracy of 81%, though robust performance across testing metrics is shown. The authors were able to use proteomic data from F. hepatica and predict eight BCEs derived from the leucine aminopeptidase protein, a known vaccine candidate, as a form of demonstration of its application. Experimental confirmation using dot-blot assays has been done on these, which shows the proof of concept that deep learning can help speed up BCE identification by reducing dependency on expensive and time-consuming experimental methods. It would be helpful in the development of vaccines, therapeutic antibody design, and diagnostics against neglected tropical diseases. In their paper, focusing on linear epitopes, the authors feel their future work needs to be directed towards conformational epitopes for better biological applicability. This manuscript is well written and I have a few suggestions for the improvement of the manuscript.

1. While the authors tested the model on Fasciola hepatica, its utility across other parasitic species was not thoroughly explored. The authors could test the model on other protozoan and helminthic parasites to demonstrate its generalizability and robustness in predicting BCEs for a wider range of pathogens.

2. Experimental validation was limited to eight peptides from a single protein. This small sample size may not fully reflect the model's predictive accuracy or its applicability to diverse proteomes. The authors should acknowledge this limitation.

3. While the model outperformed traditional machine learning approaches, it was not compared to state-of-the-art epitope prediction tools like DiscoTope or BepiPred-3.0. The authors could benchmark deepBCE-Parasite against other advanced models to establish its competitive edge and address any potential limitations in its methodology.

4. The authors should clearly state the limitations of their study.

Reviewer #5: This article comprehensively addresses the development of both deep learning and traditional machine learning models for predicting linear B cell epitopes (BCEs) in parasitic organisms, focusing on protozoa and helminths. The manuscript successfully demonstrates the benefits of Transformer-based architectures (in particular, its encoder-decoder approach with multi-head attention) over traditional methods, while providing an insightful case study on Fasciola hepatica. The validation of leucine aminopeptidase (LAP)-derived epitopes in the lab lends additional credence to the computational predictions.

1.Some details—particularly hyperparameter choices and rationale—could be expanded for reproducibility. For instance, justify why certain embedding dimensions or numbers of heads were selected.

Comparison Across Multiple Models

2. Add more discussion on how these findings could accelerate or integrate into wider vaccine and diagnostic test pipelines in veterinary or human parasitology.

3. Increase the clarity around experimental methods (e.g., how peptides were chosen for validation and how many repetitions were performed). If possible, include quantification (intensity measurements) in a table or figure to show how signal intensities compare across peptides.

4. Provide a clearer roadmap of how the proposed pipeline can be extended to 3D data and conformational epitopes in future work. A brief discussion on potential synergy with AlphaFold or related structural modeling tools would be beneficial.

**********

Results

-Does the analysis presented match the analysis plan?

-Are the results clearly and completely presented?

-Are the figures (Tables, Images) of sufficient quality for clarity?

Reviewer #1: Please see the more details annex file.

Reviewer #2: Work is well-executed and explained but figures of the main manuscript lack clarity, tremendously. High resolution images are needed.

Reviewer #3: In the Results section, while performance metrics such as AUC and MCC are provided, their significance in the context of vaccine design is not clearly articulated.

For example, why does achieving an AUC of 0.90 matter, and how does it advance the field? Similarly, although the figures and tables are informative, do the captions provide enough detail to ensure clarity and proper interpretation for the reader?

To enhance this section, qualitative examples would help demonstrate the model's real-world utility. For instance, highlighting specific epitopes that the Transformer model successfully identified but traditional methods missed would strengthen the narrative. A brief case study showing the model’s practical application could further illustrate its value.

Moreover, explaining why the Transformer outperforms traditional methods like SVM or Random Forest would add depth to the discussion—what aspects of the model contribute to its superior performance?

Finally, acknowledging any limitations of the model, such as specific types of epitopes it struggles to predict, or potential dataset biases, would provide a more balanced perspective.

Other comments for experiment validation

Why did the authors choose Fasciola hepatica as the research subject to validate the effectiveness of the model? Also, the authors highlighted four proteins, namely Glutathione transferase, Leucine aminopeptidase, Annexin, and β-Actin, to present the results. However, in the context of Fasciola hepatica research, these proteins have not demonstrated particularly ideal immune protection in livestock animals, even though they may exhibit some immunogenic effects. Additionally, can the deep learning and machine learning models developed in this study be applied to research on other parasites? And explain how this method could be applied to other parasitic pathogens or even other classes of pathogens (e.g., viral or bacterial epitopes)?

Reviewer #4: Yes, the analysis aligns with the outlined plan. The manuscript describes the development and validation of the deepBCE-Parasite model.

Yes, the results are generally well-organized and clearly presented

Yes the figures are of good quality.

Reviewer #5: Reiterate or cross-reference the original analysis plan in the Results section to show how each step (e.g., cross-validation, feature selection, benchmarking) was carried out exactly as proposed.

**********

Conclusions

-Are the conclusions supported by the data presented?

-Are the limitations of analysis clearly described?

-Do the authors discuss how these data can be helpful to advance our understanding of the topic under study?

-Is public health relevance addressed?

Reviewer #1: Please see the more details annex file.

Reviewer #2: Authors have addressed the limitations of the study at the current stage.

Reviewer #3: The conclusions are well supported by the data presented.

Reviewer #4: Yes, the conclusions are supported by the data.

The authors acknowledge some limitations, such as the model's focus on linear epitopes and the lack of conformational epitope prediction, which reduces its applicability to real-world antigen-antibody interactions. However, they do not fully discuss other potential shortcomings, such as dataset bias, limited validation across diverse parasites, or the relatively small scope of experimental validation.

Yes. The authors emphasize the significance of their model in advancing epitope-based vaccine design, antibody therapeutics, and diagnostics for parasitic diseases.

Yes. The authors address the public health implications of their work, noting that Fasciola hepatica is a globally significant zoonotic parasite and emphasizing the potential of the model to aid in vaccine development and diagnostic tools for neglected tropical diseases.

Reviewer #5: Strengthen the connection to public health by explicitly mentioning how more accurate BCE prediction can reduce infection rates, improve diagnostics, or inform surveillance programs.

**********

Editorial and Data Presentation Modifications?

Use this section for editorial suggestions as well as relatively minor modifications of existing data that would enhance clarity. If the only modifications needed are minor and/or editorial, you may wish to recommend “Minor Revision” or “Accept”.

Reviewer #1: Please see the more details annex file.

Reviewer #2: Authors have provided the background of current study and highlighted the limitations/challenges followed by desired future actions. The work can be accepted after addressing some minor concerns:

• Highlight the importance of LGBM algorithm along with supporting references, for better understanding of the readers.

• Authors should also discuss other limitations/challenges (in addition to rising drug-resistance) associated with the use of ‘conventional antiparasitic broad-spectrum drugs’, in the current scenario.

• Authors should highlight the possible advantages of incorporating artificial intelligence tools in such practices, to support or compliment experimental studies.

Reviewer #3: Hu et al., presents an innovative approach to B-cell epitope prediction using a Transformer-based deep learning model, specifically focusing on parasitic pathogens. While the study demonstrates strong technical merit and relevance to vaccine design, several issues require attention to improve the overall clarity, methodological rigor, and presentation. The overall writing and logic are smooth, but the content and spelling throughout the text need careful checking. Some information need explained and enhance clarified. I wish to recommend Minor Revision.

Reviewer #4: Minor Revision

Reviewer #5: (No Response)

**********

Summary and General Comments

Use this section to provide overall comments, discuss strengths/weaknesses of the study, novelty, significance, general execution and scholarship. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. If requesting major revision, please articulate the new experiments that are needed.

Reviewer #1: Please see the more details annex file.

Reviewer #2: Manuscript is very well written.

Reviewer #3: Hu et al., presents an innovative approach to B-cell epitope prediction using a Transformer-based deep learning model, specifically focusing on parasitic pathogens. While the study demonstrates strong technical merit and relevance to vaccine design, several issues require attention to improve the overall clarity, methodological rigor, and presentation. Below, I provide detailed comments and suggestions for each section of the manuscript.

In the abstract part, the author provides a concise overview of the study, including the problem addressed, methodology, and results. However, it is overly technical and does not clearly articulate the novelty and significance of the research. For example, while metrics such as AUC are mentioned, their relevance is not explained. Additionally, the concluding statement lacks an impact-oriented summary of how the findings could contribute to vaccine design. I have some suggestions that require authors to clarify and revise them:

First, please simplify the technical jargon to make it accessible to a broader audience. For instance, it perhaps more suitable for replacing "Transformer-driven model leveraging self-attention mechanisms" with "a novel deep learning model that improves prediction accuracy using advanced attention techniques.

Second, clearly state the novelty of the approach, such as its ability to outperform traditional machine learning models in specific metrics.

Lastly, can add a sentence on the broader implications of the study, e.g., "This study provides a robust tool for identifying B-cell epitopes, which could significantly enhance the development of vaccines for parasitic diseases."

The introduction provides a strong foundation on epitope prediction and underscores the importance of B-cell epitopes in vaccine development. However, the narrative is disrupted by an excessive focus on datasets and existing tools early on, and while the limitations of traditional methods are acknowledged, the specific research gap that this study addresses remains unclear. Reorganizing the introduction could significantly improve its clarity and flow. Thus, it is recommended to start with the significance of B-cell epitopes in vaccine design, followed by the challenges of existing prediction methods, such as limited accuracy and reliance on hand-crafted features. Moreover, clearly articulate the research gap and objective by stating, "To address these challenges, we developed a Transformer-based model to improve the accuracy and applicability of B-cell epitope prediction." Avoid redundancy, such as repeating the mention of the IEDB database, and briefly highlight the relevance of parasitic diseases like Fasciola hepatica to contextualize the study's focus. This approach will create a more engaging and logically structured introduction.

The Methods section provides a comprehensive explanation of the Transformer architecture and methodology, but its technical density may overwhelm non-technical readers. For readers without a deep learning background, in order to improve accessibility and clarity, simplify the explanation of the Transformer architecture and consider including a flowchart or visual chart to illustrate key components such as attention layers and embeddings.

Also, please expand on the data preprocessing steps by describing how the dataset was balanced (e.g., oversampling) and detailing the processes for removing duplicates or noise.

Provide justification for the chosen hyperparameters, such as stating, "The choice of two encoder layers and eight attention heads was informed by prior studies and optimization experiments." Additionally, clarify the validation strategy by explicitly mentioning whether k-fold CV or an independent test set was used, and discuss measures taken to prevent overfitting. Finally, consider rephrasing overly technical sentences for broader readability. For instance, revise "The model utilized self-attention mechanisms to capture sequence dependencies" to "The model's attention mechanisms effectively identified patterns within the sequence data." These changes will make the section more engaging and accessible while maintaining its technical rigor.

In the Results section, while performance metrics such as AUC and MCC are provided, their significance in the context of vaccine design is not clearly articulated.

For example, why does achieving an AUC of 0.90 matter, and how does it advance the field? Similarly, although the figures and tables are informative, do the captions provide enough detail to ensure clarity and proper interpretation for the reader?

To enhance this section, qualitative examples would help demonstrate the model's real-world utility. For instance, highlighting specific epitopes that the Transformer model successfully identified but traditional methods missed would strengthen the narrative. A brief case study showing the model’s practical application could further illustrate its value.

Moreover, explaining why the Transformer outperforms traditional methods like SVM or Random Forest would add depth to the discussion—what aspects of the model contribute to its superior performance?

Finally, acknowledging any limitations of the model, such as specific types of epitopes it struggles to predict, or potential dataset biases, would provide a more balanced perspective.

Other comments for experiment validation

Why did the authors choose Fasciola hepatica as the research subject to validate the effectiveness of the model? Also, the authors highlighted four proteins, namely Glutathione transferase, Leucine aminopeptidase, Annexin, and β-Actin, to present the results. However, in the context of Fasciola hepatica research, these proteins have not demonstrated particularly ideal immune protection in livestock animals, even though they may exhibit some immunogenic effects. Additionally, can the deep learning and machine learning models developed in this study be applied to research on other parasites? And explain how this method could be applied to other parasitic pathogens or even other classes of pathogens (e.g., viral or bacterial epitopes)?

(6)The overall writing and logic are smooth, but the content and spelling throughout the text need careful checking. For example, on line 241, the abbreviation for Gaussian Naive Bayes should be GNB.

Reviewer #4: Minor Revision

Reviewer #5: Overall, the study’s design, methodology, and results are presented cohesively. Once the minor revisions—particularly regarding clarity on experimental methods, dataset annotation, and limitations—are addressed, the manuscript will be suitable for publication. The revised paper would serve as a valuable resource for researchers working on vaccine design and diagnostic test development in parasitology and related fields.

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

Reviewer #4: No

Reviewer #5: Yes:  Simna Saraswathi Prasannakumari

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

Figure resubmission:

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. If there are other versions of figure files still present in your submission file inventory at resubmission, please replace them with the PACE-processed versions.

Reproducibility:

?>

Attachment

Submitted filename: reviewcomments.docx

pntd.0012985.s008.docx (15.6KB, docx)
Attachment

Submitted filename: Comments.docx

pntd.0012985.s009.docx (18.2KB, docx)
PLoS Negl Trop Dis. doi: 10.1371/journal.pntd.0012985.r004

Decision Letter 1

jong-Yil Chai, Aysegul Taylan Ozkan

13 Mar 2025

Dear Dr Hu,

We are pleased to inform you that your manuscript 'Transformer-based deep learning enables improved B-cell epitope prediction in parasitic pathogens: A proof-of-concept study on Fasciola hepatica' has been provisionally accepted for publication in PLOS Neglected Tropical Diseases.

Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow up email. A member of our team will be in touch with a set of requests.

Please note that your manuscript will not be scheduled for publication until you have made the required changes, so a swift response is appreciated.

IMPORTANT: The editorial review process is now complete. PLOS will only permit corrections to spelling, formatting or significant scientific errors from this point onwards. Requests for major changes, or any which affect the scientific understanding of your work, will cause delays to the publication date of your manuscript.

Should you, your institution's press office or the journal office choose to press release your paper, you will automatically be opted out of early publication. We ask that you notify us now if you or your institution is planning to press release the article. All press must be co-ordinated with PLOS.

Thank you again for supporting Open Access publishing; we are looking forward to publishing your work in PLOS Neglected Tropical Diseases.

Best regards,

Aysegul Taylan Ozkan, M.D., Ph.D.,

Academic Editor

PLOS Neglected Tropical Diseases

Jong-Yil Chai

Section Editor

PLOS Neglected Tropical Diseases

Shaden Kamhawi

co-Editor-in-Chief

PLOS Neglected Tropical Diseases

orcid.org/0000-0003-4304-636XX

Paul Brindley

co-Editor-in-Chief

PLOS Neglected Tropical Diseases

orcid.org/0000-0003-1765-0002

***********************************************************

p.p1 {margin: 0.0px 0.0px 0.0px 0.0px; line-height: 16.0px; font: 14.0px Arial; color: #323333; -webkit-text-stroke: #323333}span.s1 {font-kerning: none

Reviewer's Responses to Questions

Key Review Criteria Required for Acceptance?

As you describe the new analyses required for acceptance, please consider the following:

Methods

-Are the objectives of the study clearly articulated with a clear testable hypothesis stated?

-Is the study design appropriate to address the stated objectives?

-Is the population clearly described and appropriate for the hypothesis being tested?

-Is the sample size sufficient to ensure adequate power to address the hypothesis being tested?

-Were correct statistical analysis used to support conclusions?

-Are there concerns about ethical or regulatory requirements being met?

Reviewer #2: The manuscript is acceptable in the current form comprises of all the above aspects.

Reviewer #3: YES

Reviewer #4: The authors responded satisfactorily to all the comments and made the required changes in the manuscript.

**********

Results

-Does the analysis presented match the analysis plan?

-Are the results clearly and completely presented?

-Are the figures (Tables, Images) of sufficient quality for clarity?

Reviewer #2: Yes.

Reviewer #3: YES

Reviewer #4: The authors responded satisfactorily to all the comments and made the required changes in the manuscript.

**********

Conclusions

-Are the conclusions supported by the data presented?

-Are the limitations of analysis clearly described?

-Do the authors discuss how these data can be helpful to advance our understanding of the topic under study?

-Is public health relevance addressed?

Reviewer #2: Yes.

Reviewer #3: YES

Reviewer #4: The authors responded satisfactorily to all the comments and made the required changes in the manuscript.

**********

Editorial and Data Presentation Modifications?

Use this section for editorial suggestions as well as relatively minor modifications of existing data that would enhance clarity. If the only modifications needed are minor and/or editorial, you may wish to recommend “Minor Revision” or “Accept”.

Reviewer #2: Accept.

Reviewer #3: Accepted

Reviewer #4: The authors responded satisfactorily to all the comments and made the required changes in the manuscript.

**********

Summary and General Comments

Use this section to provide overall comments, discuss strengths/weaknesses of the study, novelty, significance, general execution and scholarship. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. If requesting major revision, please articulate the new experiments that are needed.

Reviewer #2: The manuscript has been significantly improved after the revision and can be accepted in the current form.

Reviewer #3: Accepted

Reviewer #4: The authors responded satisfactorily to all the comments and made the required changes in the manuscript.

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #2: No

Reviewer #3: No

Reviewer #4: No

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Amino acid frequency distribution and length statistics of peptides.

    (A-B) Amino acid frequency in positive and negative samples from the training and testing datasets. (C-D) Peptide length distribution in positive and negative samples from the training and testing datasets.

    (TIF)

    pntd.0012985.s001.tif (1.1MB, tif)
    S2 Fig. Comparison of ROC curve between deep learning and 12 handcrafted features using four machine learning algorithms.

    (TIF)

    pntd.0012985.s002.tif (1.5MB, tif)
    S3 Fig. Radar chart comparing the performance of deep learning features (deepFeature) with 12 handcrafted features across four machine learning algorithms.

    The chart highlights their respective performance in SP, SE, and MCC metrics.

    (TIF)

    pntd.0012985.s003.tif (1.5MB, tif)
    S1 Table. Parameter settings for deep learning and machine learning models used in this study.

    (XLSX)

    pntd.0012985.s004.xlsx (9.6KB, xlsx)
    S2 Table. Performance of conventional machine learning models with handcrafted features for B-cell epitope prediction in parasites.

    (XLSX)

    pntd.0012985.s005.xlsx (23.1KB, xlsx)
    S3 Table. Bioinformatic analysis of proteomic features and expression profiles of F. hepatica across four developmental stages.

    (XLSX)

    pntd.0012985.s006.xlsx (364.7KB, xlsx)
    S4 Table. The prediction results of B-cell epitope in F. hepatica using deep learning and machine learning models.

    (XLSX)

    pntd.0012985.s007.xlsx (128.3KB, xlsx)
    Attachment

    Submitted filename: reviewcomments.docx

    pntd.0012985.s008.docx (15.6KB, docx)
    Attachment

    Submitted filename: Comments.docx

    pntd.0012985.s009.docx (18.2KB, docx)
    Attachment

    Submitted filename: PNTD-D-24-01725-Response to Reviewers.docx

    pntd.0012985.s010.docx (63KB, docx)

    Data Availability Statement

    The MaxQuant search results, based on entries from the UniProt database, can be downloaded from the Zenodo repository at https://doi.org/10.5281/zenodo.14065139. The source code for the deepBCE-Parasite software is available on GitHub at https://github.com/RuiSiHu/deepBCE-Parasite and is also archived in the Zenodo repository at https://doi.org/10.5281/zenodo.14907455.


    Articles from PLOS Neglected Tropical Diseases are provided here courtesy of PLOS

    RESOURCES