Skip to main content
Journal of Genetic Engineering & Biotechnology logoLink to Journal of Genetic Engineering & Biotechnology
. 2026 Apr 26;24(2):100692. doi: 10.1016/j.jgeb.2026.100692

Structural and immunogenic evaluation of silk proteins from Bombyx mori using advanced bioinformatics and deep learning for biomaterials applications

Andrés Puerta-González a, Alejandro Soto-Ospina b, Yuliet Montoya Osorio c, John Bustamante-Osorno c, Lina María Salazar-Peláez a,
PMCID: PMC13137201  PMID: 42309597

Abstract

Silk from Bombyx mori is a premier biomaterial. Yet, comprehensive structural and immunoinformatic characterization of its protein components—fibroin subunits (FibH, FibL, P25) and five sericin isoforms—remains incomplete, hindering rational design of biocompatible medical devices. We integrated AlphaFold3 structural prediction with multi-algorithm immunoinformatic profiling (VaxiJen, NetMHCpan, BepiPred, AllerTOP, ToxinPred) to establish structure-immunogenicity relationships across the silk proteome. Structural modeling revealed that FibL and P25 adopt well-defined architectures (pTM = 0.85), whereas FibH exhibited low confidence (ipTM/pTM = 0.28), reflecting its intrinsically disordered pre-assembly state. Sericins displayed predominantly disordered conformations that undergo partial ordering upon complexing with FibL-P25-Cu2+, supporting a disorder-to-order templating mechanism for fiber assembly. Immunoinformatic analysis revealed striking antigenic heterogeneity: P25 emerged as uniquely hypoimmunogenic, with subthreshold antigenicity (VaxiJen: 0.395), a single strong MHC-I binder for a single HLA allele, and minimal MHC-II reactivity. Conversely, FibL and FibH showed the highest potential for inducing CD8+ and CD4+ cell responses among the fibroin subunits, respectively. Ser-4 and Ser-1 exhibited a broad MHC coverage, presenting ≥ 10 strong binders in 81% and 30% of MHC-I alleles, respectively, as well as ≥ 4 high-priority peptides across all 27 and 6 tested MHC-II alleles, respectively. Allergenicity prediction classified FibH, FibL, Ser-2, and Ser-5 as probable allergens. All proteins were non-toxic. These findings challenge the paradigm that sericins exclusively drive silk immunogenicity, revealing instead an HLA-dependent risk profile dominated by FibL, FibH, Ser-4, and Ser-1. This computational framework provides a rational foundation for engineering hypoimmunogenic silk variants through P25 enrichment, epitope deletion, or HLA-matched biomaterial selection.

Keywords: Protein structure prediction, Supervised machine learning, Immunoinformatics profile, Epitope mapping, Silkworm, Fibroin, Sericin

1. Introduction

Silk, a natural animal-derived fiber, is integral to the world economy and has notable applications in modern medicine. The proteins that comprise silk, primarily fibroin and sericin, are biopolymers synthesized by diverse arthropods, including silkworms and spiders.1 Silk produced by Bombyx mori silkworms is notable for its distinctive combination of physical and chemical properties. These attributes render them exceptionally well-suited for pioneering applications in advanced materials in several scenarios.2, 3

The physicochemical characteristics, bioactivity, biocompatibility, availability, and low production costs of silk proteins have attracted significant attention for potential biomedical applications. Fibroin is widely used to manufacture biocompatible materials for sutures, tissue-engineering scaffolds, and controlled drug-delivery systems.4 Sericin has gained increasing attention for its versatile applications across multiple fields, including biomedicine, particularly in drug delivery, tissue engineering, and cell culture media, as well as in the cosmetic, textile, and food industries, owing to its antioxidant, moisturizing, and antimicrobial properties, and mitogenic effect.5, 6

Current understanding of fibroin structure indicates that it is composed of a heavy chain (FibH), a light chain (FibL), and a fibrohexamerin glycoprotein (P25). These components are organized in a precise 6:6:1 ratio and are held together by an intricate network of covalent and non-covalent interactions.1 In B. mori, silk fibroin folds into crystalline β-sheet domains that are responsible for the mechanical strength of silk, whereas the amorphous regions provide flexibility.7, 8 In contrast, sericins are a group of hydrophilic glycoproteins that act as adhesive matrix-binding silk fibers in cocoons. They have a lower structural order than fibroin, consisting of a mixture of random coils and secondary conformations, and are enriched in polar amino acids, with serine being one of their principal components.9, 10

Despite these advances, the three-dimensional organization of sericins remains poorly understood, and their conformational transitions under physiological conditions have not been well-characterized. Moreover, the dynamic interactions between silk fiber components remain largely unknown. These gaps hinder the rational design of biomaterials and underscore the need for predictive approaches, such as deep learning-based structural modeling, to elucidate the influence of sequence variations on the structural and functional properties of proteins. Furthermore, identifying distant homologs of proteins is a key approach in structural bioinformatics, as it enables the discovery of the biological functions and unique properties of proteins with low sequence similarity but high structural conservation. Addressing these knowledge gaps will expand the potential of silk proteins for biomedical applications.

Therefore, accurately predicting the tertiary structures of fibroin and sericin proteins is essential for advancing their applications in biotechnology and biomedicine. Traditionally, two main computational strategies have been employed for protein structure modeling: template-based modeling (TBM) and template-free modeling (TFM). TBM utilizes known protein structures as templates, assuming significant sequence similarity between the target protein and the template, and employs prominent methods such as I-TASSER and MODELLER.11, 12 Conversely, TFM constructs three-dimensional models de novo without relying on structural templates, using fragment assembly and energy-based optimization methods, such as Rosetta.13 Recently, the integration of deep learning approaches has revolutionized this field, dramatically improving the accuracy and efficiency of protein structure prediction.14, 15, 16, 17

One of the most significant milestones in computational structural biology was the development of AlphaFold2 by Jumper et al. in 2021.15 This deep learning–based system has revolutionized protein modeling by achieving near-experimental accuracy, even for distant homologs, by integrating evolutionary, phylogenetic, and sequence data. The release of AlphaFold3 marked a significant advancement, improving the prediction speed and enhancing the accuracy of protein complex modeling. By refining the neural network architecture and incorporating richer evolutionary and structural datasets, AlphaFold3 delivers more robust and detailed representations of molecular assemblies than its predecesor.18, 19 These developments are particularly relevant to silk protein research, where precise structural predictions are key to uncovering molecular mechanisms and enabling the rational design of biomaterials with customized physicochemical, mechanical, and functional properties, as previously mentioned.

Additionally, understanding the immunogenic potential of silk proteins is not only a matter of safety but also a prerequisite for optimizing their biomedical applications and performance. Therefore, predicting antigenic epitopes is fundamental for anticipating and mitigating undesirable immune responses that could compromise the therapeutic success of silk-derived materials. However, obtaining experimental, high-resolution structural, and immunological data is challenging and costly, particularly for proteins with large sizes, repetitive motifs, or intrinsically disordered regions, such as FibH and sericins. The convergence of computational structural biology and immunoinformatics offers an innovative and efficient approach for predicting the immunogenicity of specific peptides through epitope mapping for T and B lymphocytes, with a high probability of presentation by the host's Human Leukocyte Antigen (HLA) molecules.20 By leveraging these tools, the present study performed a comprehensive in silico characterization of fibroin and sericin proteins, integrating structural modeling with multi-step immunological predictions encompassing T- and B-cell epitope identification, MHC-binding affinity analysis, and allergenicity and toxicity evaluation.21, 22

In summary, although considerable progress has been made in elucidating the molecular architecture and biomedical potential of silk proteins, critical knowledge gaps persist regarding the three-dimensional organization of sericin proteins, structural implications of fibroin–sericin interactions, and immunogenic impact of sequence variability. Addressing these gaps is essential for advancing the rational design of safe, silk-based biomaterials. Consequently, this study integrates structural prediction and immunoinformatics methodologies to characterize silk protein interactions at the structural level and elucidate their antigenic landscape, thereby establishing a computational framework for designing biocompatible silk-derived materials with controlled immunogenicity and expanded translational potential in medical biotechnology.

2. Methodology

2.1. Sequence characterization of fibroin subunits and sericin proteins from Bombyx mori and comparative analysis with distant homologs

Reference sequences of the three subunits of fibroin —Fibroin Heavy Chain (FibH), Fibroin Light Chain (FibL), and fibrohexamerin (P25)— and the five sericins (Ser-1, Ser-2, Ser-3, Ser-4, and Ser-5) were retrieved from the KAIKObase database,23 Protein Data Bank (PDB), Research Collaboratory for Structural Bioinformatics (RCSB PDB), and specialized literature.24, 25 These proteins were characterized based on their amino acid composition, molecular weight, isoelectric point (pI), and hydrophobicity index (GRAVY, Grand Average of Hydropathicity),26 which were calculated using the ProteinAnalysis library from Bio.SeqUtils.ProtParam.

Various strategies were employed to analyze distant homologs using the Basic Local Alignment Search Tool (BLAST).27 BLAST databases were constructed from protein sequences reported in the National Center for Biotechnology Information (NCBI), SwissProt, and PDB databases, downloaded from the latest versions of each. Databases were generated using the makeblastdb command with the −dbtype prot parameter.27 Different BlastP strategies were applied to identify distant homologs (Supplemental Methods 1). The first strategy involved a standard search with an E-value threshold of 0.001 and four processing threads. The second strategy increased sensitivity by setting the E-value to 0.01, reducing the word size to two, and using four threads. The third strategy involved low-complexity filtering, using the −seg yes option to eliminate regions prone to false positives while maintaining an E-value of 0.001 and employing four threads per search. The fourth strategy employed the BLOSUM45 substitution matrix, which is suitable for detecting distant homologies, with an E-value of 0.001 and 11 threads.

Additionally, Position-Specific Iterated BLAST (PSI-BLAST) strategies, which incorporate iterations and Position-Specific Scoring Matrices (PSSMs), were employed to enhance the detection of sequences that standard BLAST searches may not identify. Four different strategies were tested: (i) a standard search with an E-value threshold of 0.001 using four threads; (ii) a high-sensitivity search obtained by relaxing the threshold to an E-value of 0.01 and reducing the word size to two, also using four threads; (iii) a search including low-complexity filtering with the option −seg yes, using an E-value of 0.001 and four threads; and (iv) a search with the BLOSUM45 substitution matrix, using an E-value of 0.001, and 11 threads. Moreover, DELTA-BLAST was used to enhance homolog detection by leveraging specific conserved-domain databases, with an E-value of 0.001, four threads, and a specified output formatting.28

The databases used as subjects were the NCBI BLAST nr (non-redundant) database, which includes protein sequences and, in some cases, translated nucleotide sequences from a wide range of organisms and sources, integrating data from multiple repositories, such as GenBank,28 RefSeq,29 PDB,30 and SwissProt.19 Additionally, the BLAST pdbaa database, which contains protein sequences derived from three-dimensional structures stored in the PDB, was used.

After generating results from the implemented BLAST strategies, a filtering step was applied, selecting sequences with a percentage identity greater than 20% and a subject length exceeding 70% of the query length. These criteria are widely accepted for distant homolog analysis because a percentage identity above 20% suggests a significant homologous relationship between sequences,31 and ensuring that the subject length covers at least 70% of the query sequence increases confidence in the identified homology.32 This approach balances sensitivity and specificity by identifying sufficiently similar homologs that cover a substantial portion of the sequence, thereby avoiding false positives and ensuring the biological relevance of the results. Filtered tables were used to download sequences in the FASTA format from NCBI.

2.2. Structural characterization of fibroin subunits and sericin proteins and their distant homologs

AlphaFold3.0 (latest algorithm, 2024 version) was employed to predict the tertiary structures of fibroin subunits and sericin proteins, as well as the structures of their distant homologs identified through the BLAST and PSI-BLAST strategies described in Section 2.1. Furthermore, AlphaFold3.0 was used to explore the protein–protein interactions among fibroin and sericin proteins of B. mori, providing a detailed analysis of their molecular associations and potential conformational changes.18 To investigate whether the model geometry could be improved, AlphaFold3.0 outputs were further subjected to structural refinement with ModRefiner33 and FG-MD,34 and pre- and post-refinement models were compared using the same quality indicators to identify any systematic improvement attributable to the refinement.

AIUPred, which combines energy-based estimations with deep learning techniques, was used in parallel to characterize intrinsically disordered regions (IDRs) computationally.35 This approach surpasses traditional biophysical models by employing advanced neural network architectures, including a transformer network for energy estimation and a feedforward network for computing disorder propensity.35 Additionally, the AIUPred's conditional disorder prediction feature was used to identify regions whose structural states depend on environmental redox conditions. For all general disorder predictions, residues with disorder propensity scores > 0.5 were classified as disordered.

To identify specific functional binding sites within IDRs, the complementary tool AIUPred-binding was employed.36 This predictor, integrated within the same web server, applies an innovative transfer-learning strategy based on the AIUPred architecture. It leverages a mathematical representation of structural energies, referred to as energy embedding, derived from the AIUPred transformer and combines it with AlphaMissense pathogenicity scores, which serve as a proxy for identifying functionally critical regions. Residues with predicted binding scores greater than 0.5 were considered functionally disordered binding sites.

Charge–hydropathy analysis was performed to distinguish ordered from intrinsically disordered proteins based on their net charge and average hydrophobicity, using charge–hydropathy plots as described by Dunker et al. (2013).37 Finally, the Disorder Enhanced Phosphorylation Predictor (DEPP), which is trained on experimentally validated phosphorylation sites, was applied to assess the relationship between structural disorder and phosphorylation. The predictor achieved accuracies of 76.0 ± 0.3% for serine, 81.3 ± 0.3% for threonine, and 83.3 ± 0.3% for tyrosine residues, highlighting the link between intrinsic disorder and post-translational modification potential.37, 38 However, our analyses did not detect post-translational modifications, such as phosphorylation of serines, threonines, or tyrosines, in any of the proteins analyzed.

2.3. Computational identification of immunogenic epitopes

2.3.1. Antigenicity prediction of silk proteins

To evaluate the immunogenic potential of fibroin subunits and sericin proteins, an antigenic prediction score was calculated using the VaxiJen 2.0 server,39 a computational antigen prediction tool that does not rely on sequence alignment.

2.3.2. Cytotoxic T lymphocyte (CD8+) and helper lymphocyte (CD4+) epitope prediction

Cytotoxic T lymphocyte (CTL) epitopes were predicted using two complementary computational approaches to ensure comprehensive coverage and robust prediction. Initially, the NetCTL (Neural network-based prediction of CTL epitopes) web server (version 1.2) was used to simulate the intracellular antigen presentation pathway by integrating three major predictive components: (i) proteasomal cleavage of the protein, (ii) peptide transport efficiency mediated by the TAP (Transporter associated with Antigen Processing) molecule, and (iii) binding affinity to MHC (Major Histocompatibility Complex) class I molecules. Based on the combined results of these steps, the algorithm assigns a final score to each 9-mer peptide, ranking its probability of functioning as a potential epitope.40

Subsequently, the T-cell prediction class I (TC1) suite, available on the Next-Generation IEDB Tools platform (https://nextgen-tools.iedb.org), was used to refine and expand the prediction process of peptide binding to MHC class I molecules.21 This platform integrates advanced algorithms that simulate the full antigen presentation pathway, including proteasomal cleavage, TAP-mediated transport, peptide elution, MHC class I binding, and epitope immunogenicity. Peptide–MHC class I binding affinity was evaluated using NetMHCpan-4.1, a state-of-the-art neural network–based model trained on experimental binding affinity and mass spectrometry–eluted ligand data to enhance predictive accuracy. Nine amino acid peptides were analyzed, and each was assigned a median binding percentile score (MBPS) relative to the natural peptide dataset. Following the recommended thresholds, peptides with MBPS < 0.5% were classified as strong binders, whereas those with MBPS > 0.5 and < 2% were classified as weak binders.22

In parallel, peptide binding to MHC class II molecules was predicted using NetMHCIIpan-4.0.22 This tool offers comprehensive coverage across all class II loci, enabling the reliable prediction of 15-mer peptides. Based on the MBPS, peptides were categorized as strong binders (<2%) or weak binders (<10%).41 All analyses were conducted using the Immune Epitope Database (IEDB) Analysis Resource, a comprehensive suite of tools dedicated to epitope prediction and analysis, which complements the manually curated IEDB repository of experimentally validated immune epitopes.21

2.3.3. B-cell epitope prediction

Both linear and conformational B-cell epitopes were predicted using the BepiPred-3.0 server (https://services.healthtech.dtu.dk/services/BepiPred-3.0/). This state-of-the-art method is powered by protein language models (LMs), which infer structural and antigenic features directly from the primary amino acid sequence, achieving a markedly higher predictive accuracy than earlier approaches.42 The algorithm assigns a probability score to each residue in the protein sequence, indicating its likelihood of being an epitope.42 For the final epitope identification, the default classification threshold recommended by the tool (0.1512) was applied, as this cutoff maximizes the Matthews Correlation Coefficient (MCC) on the validation datasets, ensuring an optimal balance between sensitivity and specificity.42

2.3.4. Allergenicity and toxicity evaluation of predicted epitopes

The allergenicity and toxicity of epitopes were evaluated using the AllerTOP v.243, 44 and ToxinPred v3.0 servers.45 AllerTOP v.2 considers factors such as hydrophobicity, molecular size, and helical propensity to determine allergenicity. The ToxinPred 3.0 server was used for the toxicity prediction. This method uses a hybrid (ensemble) approach that integrates a machine learning (ML) classifier with motif-based analysis. The best-performing model combined the Extra Tree (ET) classifier, based on compositional features such as amino acid composition (AAC) and dipeptide composition (DPC), with the MERCI motif-scoring algorithm. This hybrid ET + MERCI model achieved the highest predictive accuracy, with an Area Under Curve (AUC) of 0.98 in the independent validation dataset, demonstrating excellent reliability in distinguishing toxic from nontoxic protein sequences.

2.4. Data analysis

All analyses were performed in Python (v3.12.9). Data preprocessing and numerical analyses were conducted using Pandas (v1.5.3) and NumPy (v1.26.4). Figures were generated with Matplotlib (v3.10.0) and Seaborn (v0.13.2).

3. Results

3.1. Sequence characterization of fibroin subunits and sericin proteins from Bombyx mori and comparative analysis with distant homologs

The percentage composition of amino acid residues in a protein influences the subsequent levels of organization of the molecule and its intrinsic properties. Table 1 summarizes the relative abundance of each amino acid residue in the eight silk proteins analyzed, based on reference sequences from KAIKObase, PDB, RCSB PDB, and specialized literature. The fibroin light chain (FibL) is abundant in alanine (Ala, 14.12%), serine (Ser, 9.54%), and glycine (Gly, 8.40%), which are the three most abundant amino acid residues. Fibroin heavy chain (FibH) displayed exceptionally high proportions of glycine (Gly, 45.83%), alanine (Ala, 30.18%), and serine (Ser, 12.09%). Fibrohexamerin (P25) was enriched in leucine (Leu, 10.0%) and alanine (Ala, 7.27%), but its threonine (Thr, 5.45%) and cysteine (Cys, 4.09%) contents were higher than those of the other fibroin subunits.

Table 1.

Relative abundance (percentage) of amino acid residues in silk proteins.

FibL FibH P25 Ser-1 Ser-2 Ser-3 Ser-4 Ser-5
Ala 14.12 30.18 7.27 6.98 2.44 5.19 5.55 0.86
Cys 1.15 0.10 4.09 0.41 0.00 0.39 0.44 0.35
Asp 6.49 0.48 6.36 7.31 8.78 5.35 5.37 12.85
Glu 1.91 0.57 3.18 2.55 14.89 4.41 9.03 10.68
Phe 3.05 0.55 6.82 0.74 1.78 0.39 2.16 1.37
Gly 8.40 45.83 4.09 10.76 6.44 11.96 7.44 10.22
His 1.91 0.10 4.09 1.40 1.33 0.79 5.42 0.71
Ile 8.02 0.25 6.82 1.40 1.22 0.31 5.42 2.73
Lys 1.91 0.23 3.64 3.94 14.44 6.22 9.82 20.19
Leu 7.63 0.13 10.0 1.89 2.89 0.55 4.32 1.72
Met 0.76 0.08 0.91 0.16 0.22 0.08 0.31 0.20
Asn 6.87 0.38 5.45 5.18 6.67 6.69 4.98 5.01
Pro 3.44 0.27 5.45 0.82 0.78 0.16 3.04 6.43
Gln 5.73 0.19 2.73 2.47 3.44 7.40 5.64 1.01
Arg 3.82 0.29 5.91 4.77 6.44 2.68 2.51 5.11
Ser 9.54 12.09 6.36 31.39 16.33 43.59 9.47 10.27
Thr 3.05 0.89 5.45 8.87 5.11 2.75 9.60 7.14
Val 7.25 1.94 5.45 3.78 3.33 0.71 5.73 2.18
Trp 0.76 0.21 1.36 0.82 0.00 0.00 0.53 0.30
Tyr 4.20 5.27 4.55 4.35 3.44 0.39 3.22 0.66

Amino acid residues are displayed using a three-letter code. Abbreviations: FibL, Fibroin Light Chain; FibH, Fibroin Heavy Chain; P25, fibrohexamerin. Ser-1, sericin 1; Ser-2, sericin 2; Ser-3, sericin 3; Ser-4, sericin 4; Ser-5, sericin 5.

The amino acid composition varies markedly among different sericin proteins. Serine (Ser) was the most abundant residue in Ser-1 (31.39%), Ser-2 (16.33%), and Ser-3 (43.59%), although it was also highly represented in Ser-4 (9.47%) and Ser-5 (10.27%). Ser-1 additionally shows high levels of glycine (Gly, 10.76%) and threonine (Thr, 8.87%), while Ser-2 is also characterized by elevated amounts of glutamic acid (Glu, 14.89%) and lysine (Lys, 14.44%). Ser-3 was particularly rich in glycine (Gly, 11.96%) and glutamine (Gln, 7.40%), whereas Ser-4 contains high levels of lysine (Lys, 9.82%), threonine (Thr, 9.60%), and glutamic acid (Glu, 9.03%). Ser-5 exhibits a compositional profile dominated by lysine (Lys, 20.19%), aspartic acid (Asp, 12.85%), glutamic acid (Glu, 10.68%), and glycine (Gly, 10.22%).

On the other hand, BLAST analysis successfully identified distant homologs for each silk protein (Supplemental Fig. 1). After filtering, we obtained 101 homologs for FibL, 12 for FibH, and a notably large set of 168 for P25. These proteins span a wide range of sequence identities, clustering primarily between 20% and 60%. In contrast, sericins yielded a much smaller number of homologs: six, six, three, four, and 22 were identified for Ser-1, Ser-2, Ser-3, Ser-4, and Ser-5, respectively. Moreover, the homologs of Ser-1, Ser-2, Ser-3, and Ser-4 showed sequence identities greater than 78%. Notably, Ser-5 contrasted sharply with the other sericins by exhibiting low identity values (20–32%).

Supplemental Fig. 2 illustrates the molecular weight (kDa, log10 scale) distribution of the fibroin subunits and sericin variants, highlighting significant heterogeneity across the silk proteome. The boxplot for FibL shows a tightly clustered distribution in the 25–30 kDa range. In stark contrast, FibH exhibited the most significant variability, with molecular weights spanning approximately 350–800 kDa. The P25 glycoprotein has a median molecular weight of 25–30 kDa, similar to FibL, but shows much wider dispersion due to numerous outliers. Among the sericins, Ser-1, Ser-2, and Ser-3 show remarkable consistency, with their molecular weights tightly grouped around 120 kDa. Ser-4, however, is represented by a single, much higher value of approximately 250 kDa. Notably, Ser-5 showed a broader distribution, with most values spanning 75–400 kDa, and an outlier exceeding 1,000 kDa.

The isoelectric point (pI) is a key determinant of protein solubility, electrostatic behavior in aqueous environments, and stability. As shown in Supplemental Fig. 3, the silk proteins exhibited distinct pI distributions. FibL displayed remarkable variability, with two clusters clearly differentiated: one grouped around a strongly acidic pI (4–5), while the other was grouped within 6.5–8.5 pI values. Similarly, FibH and P25 exhibited a wide pI distribution, with minimum and maximum values at the acidic (4–6) and basic (8–10) extremes. In contrast, most sericins exhibited a comparatively uniform electrostatic profile: Ser-1, Ser-2, and Ser-3 clustered tightly within the acidic pI range of 5.5–6.5, and Ser-4, represented by a single value near pI 6.2, followed the same trend. Notably, Ser-5 deviated from this pattern, displaying the widest pI dispersion among sericins, with values spanning from strongly acidic 4.0 to highly basic 10 and a broad interquartile range.

The grand average of hydropathy (GRAVY) index quantifies the relative hydrophobicity (positive values) or hydrophilicity (negative values) of a protein. As illustrated in Supplemental Fig. 4, the distribution of GRAVY scores revealed a clear functional distinction between the two prominent silk protein families. The fibroin subunits FibL, FibH, and P25 showed GRAVY values clustering around zero, even though P25 displayed the widest dispersion in the negative value zone. The sericin proteins consistently exhibited strongly negative GRAVY scores (< −0.5), confirming their predominantly hydrophilic nature. For instance, sericin-2 showed extremely low, narrowly distributed values (approximately −2.0). Notably, Ser-5 maintained a strongly hydrophilic profile but showed the most significant variability among the sericins, with a broad interquartile range (–2.4 to –1.3) and two outliers located at −2.9 and −0.7, respectively.

3.2. Structural characterization of fibroin subunits and sericin proteins and their distant homologs

Structural predictions for silk proteins were obtained using AlphaFold3. This report presents complementary confidence metrics that allow readers to assess predicted structures at the local, global, and interface levels using model-internal estimates of standard structural accuracy metrics. The pLDDT (predicted Local Distance Difference Test)46 provides per-residue confidence, helping distinguish well-defined folded segments from regions likely to be flexible or intrinsically disordered. The pTM (predicted TM-score) summarizes global topological reliability on a 0–1 scale, indicating whether the overall architecture of a protein (or complex) is expected to be correct, largely independent of fine local details. For complexes, ipTM (interface-predicted TM-score) focuses on inter-chain packing, estimating how trustworthy the relative positioning and contacts between subunits are; thus, a low ipTM suggests that the interaction geometry may be incorrect even when individual chains appear well modeled.47, 48 To rank multimeric predictions, AlphaFold3 combines these signals using a weighted confidence score (commonly 0.8 × ipTM + 0.2 × pTM), prioritizing interface accuracy because correct biological interpretation of complexes depends primarily on how partners assemble rather than on the fold of each isolated component.

The models for FibL, FibH, and P25 fibrohexamerin are shown in Fig. 1. FibL was characterized by a predominantly α-helical architecture, with regions of very high confidence (pLDDT > 90), combined with some flexible segments of low confidence (70 > pLDDT ≥ 50), consistent (ipTM/pTM = 0.85) (Fig. 1A). In contrast, FibH revealed an extensive β-sheet–rich structure, but with overall low confidence (pLDDT between 50–70) and only a few short segments modeled with very low reliability (pLDDT < 50), reflecting its large size and structural complexity (ipTM/pTM = 0.28) (Fig. 1B). P25 showed a compact fold dominated by β-sheets, with a substantial portion of the structure modeled with very high confidence (pLDDT > 90, dark blue) and confident regions (90 > pLDDT > 70, light blue), whereas some peripheral loops exhibited lower confidence (yellow and orange) (Fig. 1C). Consistent with this residue-level pattern, the predicted global accuracy for P25 was high (pTM = 0.85), supporting the overall plausibility of the modeled topology. The ipTM was reported as unavailable, indicating that an interface score was not computed under the evaluated setting; therefore, our interpretation of P25 focuses on the global fold and local confidence distribution rather than on any protein–protein interface geometry.

Fig. 1.

Fig. 1

Structural predictions of fibroin subunits from B. mori with AlphaFold3. Predicted 3D structures of the fibroin light chain (FibL) (A), fibroin heavy chain (FibH) (B), and fibrohexamerin (P25) (C). The model confidence is color-coded by pLDDT score: very high (pLDDT > 90, dark blue), confident (90 > pLDDT > 70, light blue), low (70 > pLDDT > 50, yellow), and very low (pLDDT < 50, red). FibL mainly exhibits α-helices with mixed confidence regions (ipTM/pTM = 0.85). FibH revealed an extensive β-sheet arrangement but with a predominantly low-confidence prediction (ipTM/pTM = 0.28). P25 exhibited a predominance of β-sheets, with high-confidence regions dominating the fold (pTM = 0.85). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Fig. 2 shows the AlphaFold3 structural predictions for the sericin proteins. Overall, sericins display partially ordered cores surrounded by extensive intrinsically disordered regions (IDRs). Ser-1 presents a β-sheet–enriched core but with very low prediction confidence (pLDDT < 50; ipTM/pTM = 0.20), reflecting its highly disordered nature (Fig. 2A). Ser-2 exhibited a predominance of α-helices, modeled with regions of confident reliability (90 > pLDDT > 70), interspersed with poorly defined segments (pLDDT < 70; ipTM/pTM = 0.20) (Fig. 2B). Ser-3 contains long α-helical stretches but is predicted to have low-to-very-low confidence (pLDDT < 70; ipTM/pTM = 0.14), suggesting extensive flexibility (Fig. 2C). In contrast, Ser-4 presented a better-defined fold with both α-helices and β-sheets, including segments modeled with confident reliability (90 > pLDDT > 70), although large regions remained poorly predicted (pLDDT < 70; ipTM/pTM = 0.28) (Fig. 2D). Ser-5 was predicted to harbor a compact, well-defined region composed of α-helices modeled with very high local confidence (pLDDT > 90), although embedded within extensive very-low-confidence regions (predominantly pLDDR < 50), consistent with a structured core surrounded by large IDRs. The low global confidence (pTM = 0.28; ipTM not applicable) suggests limited overall model accuracy and substantial conformational flexibility across the full-length protein (Fig. 2E).

Fig. 2.

Fig. 2

Structural predictions of sericin proteins from B. mori with AlphaFold3. Predicted 3D structures of sericin 1 (A), sericin 2 (B), sericin 3 (C), sericin 4 (D), and sericin 5 (E). The model confidence is color-coded by pLDDT score: very high (pLDDT > 90, dark blue), confident (90 > pLDDT > 70, light blue), low (70 > pLDDT > 50, yellow), and very low (pLDDT < 50, red). Sericin 1 shows a β-sheet–enriched core but with very low-confidence predictions (pLDDT < 50; ipTM/pTM = 0.20). Sericin 2 exhibited a predominance of α-helices modeled with regions of confident prediction (90 > pLDDT > 70), interspersed with poorly defined segments (pLDDT < 70; ipTM/pTM = 0.20). Sericin 3 contains long α-helical stretches but is mainly modeled with low-to-very-low confidence (pLDDT < 70; ipTM/pTM = 0.14). Sericin 4 displays both α-helices and β-sheets, with some regions of confidence (90 > pLDDT > 70) but extensive disorder (pLDDT < 70; ipTM/pTM = 0.28). Sericin 5 was predicted to harbor a compact, well-defined region composed of α-helices with very high local confidence (pLDDT > 90), embedded within extensive very low-confidence regions (predominantly pLDDR < 50), consistent with a structured core surrounded by large IDRs (pTM = 0.28; ipTM not applicable). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Consistent with the nature of FibH, FibL, P25, and Ser-1 to Ser-5, the predicted structures displayed pronounced confidence variability along the sequences, with recurrent low-confidence regions corresponding to low-complexity segments likely reflecting flexible or intrinsically disordered behavior. Accordingly, our structural interpretation focused on comparatively higher-confidence portions and avoided over-interpreting low-confidence repetitive regions. Refinement with ModRefiner and FG-MD did not yield consistent improvements over the original AlphaFold3 models, as no systematic enhancement was observed across the evaluated quality indicators. Therefore, we report these structures as informed structural hypotheses for describing global trends, while explicitly acknowledging the inherent limitations of single-conformation predictions for highly repetitive silk proteins.

Intrinsically disordered proteins (IDPs) or regions (IDRs) are segments that lack a stable three-dimensional conformation under physiological conditions but often acquire a structure upon binding to partners or under specific environmental cues.37 These regions are crucial for molecular recognition, flexibility, and regulatory processes in many biomolecules.37 The prediction of IDRs in silk proteins using AIUPred and AIUPred-binding revealed significant differences among fibroins. FibH was identified as intrinsically disordered, exhibiting stable regions only at its N- and C-termini (Fig. 3A), and its structural conformation depended on the redox state (Fig. 3B). In contrast, no significant interaction points leading to structural stabilization were detected in FibL (Supplemental Fig. 5A-B) or P25 (Supplemental Fig. 5C-D).

Fig. 3.

Fig. 3

Prediction of structural disorder, interaction-prone regions, and redox sensitivity in the fibroin heavy chain (FibH) using AIUPred. For the figure on the left (A), the solid blue line represents the predicted disorder probability along the amino acid sequence, with values above the 0.5 threshold (blue dashed line) indicating intrinsically disordered regions, whereas the solid orange line indicates potential structural stabilization points via protein–protein interactions. For the figure on the right side (B), the blue and orange solid lines show disorder predictions under reducing (redox minus) and oxidizing (redox plus) conditions, respectively. FibH is predicted to be predominantly disordered, with stable segments only in the N- and C-terminal regions. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Ser-1 (Fig. 4A-B), Ser-2 (Fig. 4C-D), and Ser-3 (Fig. 4E-F) exhibited IDP behavior, with stabilization points associated with potential protein–protein interactions (left side figures, solid orange lines) and redox-dependent conformational changes (right side figures). In contrast, Ser-4 (Fig. 4G-H) exhibited two large intrinsically disordered regions that may adopt stable secondary structures upon binding, whereas ordered segments were identified between positions 10 and 300 and 1300–1700. Ser-5 (Fig. 4I-J) showed a marked transition from a relatively ordered N-terminal segment (approximately positions 1–300) to an extensive intrinsically disordered region starting around 320–350, where the disorder scores remained persistently high (often approaching 1.0), suggesting a dense distribution of potential interaction-prone motifs within the long-disordered region.

Fig. 4.

Fig. 4

Prediction of structural disorder, interaction-prone regions, and redox sensitivity of sericin proteins using AIUPred. Sericin 1 (A-B), sericin 2 (C-D), sericin 3 (E-F), sericin 4 (G-H), and sericin 5 (I-J). For the figures on the left, the solid blue line represents the predicted disorder probability along the amino acid sequence, with values above the 0.5 threshold (blue dashed line) indicating intrinsically disordered regions, whereas the solid orange line indicates potential structural stabilization points via protein–protein interactions. For the figures on the right side, the blue and orange solid lines show disorder predictions under reducing (redox minus) and oxidizing (redox plus) conditions, respectively. Sericin 1, 2, and 3 exhibited IDP behaviors. Sericin 4 exhibited two large intrinsically disordered regions, whereas ordered segments were identified between positions 10 and 300 and 1300–1700. Sericin 5 exhibited only an ordered segment in the N-terminal region (positions 1–300). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Complementary analyses using AIUPred-binding corroborated AIUPred's results, supporting the reliability of both tools in identifying IDRs. Charge–hydropathy analysis, which classifies proteins based on the balance between net charge and hydrophobicity, revealed that fibroins and sericins fall into distributions consistent with their expected biochemical characteristics: fibroins cluster near the order–disorder threshold due to their repetitive motifs, and sericins are classified as predominantly disordered because of their polar and hydrophilic composition.

On the other hand, considering that β-sheet formation in silk fibroins has been reported to be efficiently promoted by copper ions through histidine coordination, with an estimated stoichiometry of 3.4 Cu2+ ions per histidine residue,49 we selected copper as the metal ion for the structural modeling. This choice was made because, among the metal ions reported for fiber formation in the silk glands of the moth (e.g., Zn2+, Na+, Mg2+, Ca2+, and K+), the available literature is typically presented in terms of broader concentration/weight ranges rather than residue-level stoichiometries. Accordingly, fibroin models were constructed by incorporating the corresponding number of cooper ions to represent this specific coordination effect. The FibL chain, composed of 262 amino acids, contained approximately 1.91% histidine residues, corresponding to a total of five histidines. Based on the estimated interaction ratio of 3.4 Cu2+ ions per histidine residue, this structure could theoretically coordinate 17 Cu2+ ions (Supplemental Fig. 6).

The FibH chain comprises 5,259 amino acids; however, because AlphaFold3 can model up to 5,000 residues, FibH was analyzed using a 5,000-residue truncation (Fig. 5). Within this modeled region, histidines account for ∼0.10% of the sequence (n = 5), which, based on the reported coordination ratio, supports the inclusion of 17 Cu2+ ions. To minimize potential bias introduced by the length constraint, we generated two complementary Cu2+-bound variants: one lacking 259 residues at the C-terminus (Fig. 5A), and another lacking 259 residues at the N-terminus (Fig. 5B). In both cases, the global confidence remained low and changed only modestly (pTM = 0.29; ipTM = 0.32–0.33), consistent with the known size, repetitiveness, and conformational heterogeneity of fibroin heavy chains. Nonetheless, the Cu2+-containing models displayed a more coherent appearance of the structured core, with clearer β-sheet features relative to the Cu2+-free baseline (Fig. 1B), suggesting a limited but detectable stabilization compatible with metal-mediated ordering.

Fig. 5.

Fig. 5

Prediction of truncated fibroin heavy chain interactions using AlphaFold3. FibH lacking 259 residues at the C-terminus (A) or at the N-terminus (B). FibH N-terminally truncated, with copper ions and FibL (C). FibH C-terminally truncated, with copper ions and FibL (D). We specified a complex comprising FibH and 17 Cu2+ ions, consistent with the estimated stoichiometry derived from the histidine content (3.4 Cu2+:5His). In both cases of FibH truncated modeling, the global confidence remained low and changed only modestly (pTM = 0.29; ipTM = 0.32–0.33). However, the Cu2+-containing models displayed a more coherent appearance of the structured core, with clearer β-sheet features relative to the Cu2+-free baseline (Fig. 1B). Abbreviations: Fibroin heavy chain, FibH; fibroin heavy chain, FibL.

Finally, when truncated FibH (17 Cu2+) was modeled together with FibL, the outcome depended on the truncation context. The N-terminally truncated model showed an increase in global confidence (pTM = 0.32) and, importantly, a substantial gain in interface confidence (ipTM = 0.69) (Fig. 5C), supporting a more reliably defined FibH–FibL association under these conditions. In contrast, the C-terminally truncated complex yielded only a modest global improvement (pTM = 0.31) and a low interface metric (ipTM = 0.32) (Fig. 5D), indicating that while association is still plausible, its relative docking is less well constrained, consistent with the expectation that sequence context and terminal completeness can influence how confidently AlphaFold3 resolves large, partially ordered assemblies.

The P25 glycoprotein consists of 220 amino acid residues, nine of which are histidine residues. Based on the estimated interaction ratio of 3.4 copper ions per histidine, the protein was predicted to coordinate 31 Cu2+ ions (Supplemental Fig. 7). AlphaFold3 predictions for P25 alone yielded a high-confidence fold (pTM = 0.85), characterized by a well-defined, structured core, dominated by β-sheets- with only a short, low-confidence tail (Fig. 1C). Upon incorporation of 31 Cu2+ ions, the global confidence decreased slightly (pTM = 0.79), although it remained high. In contrast, an explicit interface metric became available (ipTM = 0.48), consistent with a more complex model in which cooper coordination introduces additional degrees of freedom, increasing conformational variability (i.e., a larger set of feasible relative positions and geometries). Importantly, the overall architecture of the structured core was preserved.

According to Inoue et al. (2000), fibroin assembles in a molar ratio of six FibH, six FibL, and one P25 molecule.1 Based on this stoichiometry, we modeled the interactions among these proteins in AlphaFold3, excluding FibH due to its size constraints, and we incorporated copper ions in proportion to the total histidine content of the modeled subunits. However, in a fully stoichiometric representation of the six FibL and one P25 assembly, the expected Cu2+ load would be ≈133 ions in total. Nonetheless, AlphaFold3 currently supports modeling up to 50 ions, preventing the direct inclusion of the full complement. We therefore adopted a conservative and chemically motivated approximation by including 48 Cu2+ ions, corresponding to the combined estimated coordination capacity of a single FibL (17 Cu2+; five histidines) plus one P25 (31 Cu2+; nine histidines), thereby retaining the most plausible high-affinity coordination sites while remaining within the platform limits.

Consistent with the high-confidence fold predicted for FibL (pTM = 0.81) and the similarly reliable structure obtained for P25 (pTM = 0.85), we constructed a complex comprising six FibL chains and one P25 (Fig. 6A). The predicted confidence for the assembly was low (pTM = 0.31; ipTM = 0.21; 0.8·ipTM + 0.2·pTM = 0.23), indicating that the global arrangement and inter-chain packing are weakly constrained and may admit multiple plausible configurations. When the same complex was modeled with 48 Cu2+ ions (Fig. 6B), confidence increased only modestly (pTM = 0.32; ipTM = 0.23; 0.8·ipTM + 0.2·pTM = 0.25). Visually, the Cu2+-containing model showed closer clustering of structured elements, particularly α-helical regions, suggesting that copper coordination may promote incremental local consolidation of the assembly without markedly altering the predicted global architecture.

Fig. 6.

Fig. 6

Structural modeling of the fibroin light chain (FibL)–fibrohexamerin (P25) – copper complex. Assembly of six FibL subunits and one P25 glycoprotein modeled using AlphaFold3 (A). The same complex was modeled with 48 incorporated Cu2+ ions (red) according to the histidine coordination ratio (B). The presence of copper ions promotes a more ordered, compact arrangement, thereby enhancing the organization of α-helical regions. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

In addition to modeling the structural complex of six FibL subunits, one P25 glycoprotein, and 48 Cu2+ ions, we explored the interaction of this assembly with sericin proteins, which are characterized by a high degree of intrinsic disorder (Fig. 2, Fig. 3, Fig. 4). As shown in Fig. 7, sericins exhibited structural ordering upon interaction with the FibL–P25–Cu complex. Specifically, Ser-1 and Ser-4 (Fig. 7A and D, respectively) adopted α-helical and parallel β-sheet conformations; Ser-2 (Fig. 7B) primarily adopted an α-helical structure; and Ser-3 (Fig. 7C) formed α-helices and antiparallel β-sheets. Across the four docking experiments, the global confidence of the resulting assemblies remained low (pTM ≈ 0.30–0.31) with modest interface scores (ipTM ≈ 0.24–0.27), indicating that the relative placement of sericins on the FibL–P25–Cu scaffold was only weakly constrained, which was an expected outcome given the high disorder content of sericins and their likely conformational plasticity. Notably, Ser-4 yielded the highest interface metric (ipTM = 0.27), consistent with its proposed role as a fibroin-proximal sericin layer in fourth instar.25 In contrast, Ser-2 showed slightly lower interface confidence (ipTM = 0.24), consistent with a more external, highly flexible adhesive function in fourth- and fifth-instar larvae. Despite the limited ipTM values, the models consistently suggest that contact with the fibroin complex can promote local structuring within sericins (folding-upon-binding), supporting a mechanism in which intrinsically disordered sericin segments acquire a secondary structure at the interface while remaining essentially dynamic elsewhere.

Fig. 7.

Fig. 7

Structural modeling of sericin proteins with the fibroin light chain (FibL)–fibrohexamerin (P25)–copper complex. Sericin 1 (Ser-1) (A), sericin 2 (Ser-2) (B), sericin 3 (Ser-3) (C), and sericin 4 (Ser-4) (D) are shown interacting with a complex composed of six FibL subunits, one P25 glycoprotein, and 48 Cu2+ ions (red). Upon interaction, the intrinsically disordered regions (IDRs) of sericins undergo conformational ordering, forming α-helices and β-sheet structures. Ser-1 and Ser-4 adopt α-helices and parallel β-sheets, Ser-2 primarily assumes α-helical conformations, and Ser-3 forms α-helices and antiparallel β-sheets. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Ser-5 showed evidence of induced ordering, consistent with its reported β-sheet propensity and its role as a mechanically relevant adhesive component in non-cocoon silk.25 Based on the proposed multilayer organization of fourth-instar silk,25 in which Ser-4 occupies the fibroin-proximal layer, and Ser-2 forms the outermost adhesive layer, we modeled Ser-5 interactions specifically with Ser-4 and Ser-2 rather than directly with the hydrophobic fibroin core. In both complexes, Ser-5–Ser-4 (Fig. 8A) displayed a slight increase in global confidence (pTM 0.28 → 0.30; ipTM = 0.26), whereas Ser-5–Ser-2 (Fig. 8B) showed a marked improvement relative to the Ser-2 monomer (pTM 0.16 → 0.28; ipTM = 0.21). Despite these gains, the low ipTM values and the prevalence of long low-pLDDT segments indicate that both interfaces are weakly constrained and likely conformationally heterogeneous, with interactions mediated by limited structured patches embedded within extensive intrinsically disordered regions. Notably, a joint Ser-5/Ser-4/Ser-2 assembly could not be modeled because its combined size exceeds AlphaFold3′s practical input limits; nonetheless, a complete ternary model is expected to better capture cooperative constraints across layers and may yield improved confidence metrics and more clearly defined interaction geometries.

Fig. 8.

Fig. 8

Structural modeling of sericin 5 (Ser-5) with sericin 4 (Ser-4) or sericin 2 (Ser-2). Based on the proposed multilayer organization of fourth-instar silk25, in which Ser-4 occupies the fibroin-proximal layer, and Ser-2 forms the outermost adhesive layer, we modeled Ser-5 interactions specifically with Ser-4 (A) or Ser-2 (B) rather than directly with the hydrophobic fibroin core. Ser-5–Ser-4 (A) displayed a slight increase in global confidence (pTM 0.28 → 0.30; ipTM = 0.26), whereas Ser-5–Ser-2 (B) showed a marked improvement relative to the Ser-2 monomer (pTM 0.16 → 0.28; ipTM = 0.21). An Ser-5/Ser-4/Ser-2 assembly could not be modeled because its combined size exceeds AlphaFold3′s practical input limits.

3.3. Computational identification of immunogenic epitopes

3.3.1. Antigenicity prediction of silk proteins

The antigenic potential of the silk proteins was evaluated using VaxiJen v2.0 software, which classifies proteins as highly antigenic (score > 0.8), moderately antigenic (score 0.5–0.8), or weakly antigenic (score < 0.5). FibH and Ser-3 reached a maximum score of 1.0, indicating a strong potential to elicit an immune response. Similarly, Ser-1 (0.858), Ser-2 (0.815), and Ser-5 (0.897) fell within the highly antigenic range. FibL (0.637) and Ser-4 (0.656) were classified as moderately antigenic, suggesting a lower but still significant immunogenic potential. P25, with a score of 0.395, was the only protein categorized as having low antigenicity, suggesting a minimal likelihood of triggering immune recognition.

3.3.2. Cytotoxic T lymphocyte (CD8+) and helper lymphocyte (CD4+) epitope prediction

The prediction of cytotoxic T lymphocyte (CTL, CD8+) epitopes revealed marked differences in the distribution of supertype alleles among fibroins (Table 2). FibH showed a strong bias toward alleles A26 (with 225 predicted epitopes), B58 (n = 48), and B62 (n = 265), which far exceeded those observed in FibL (A26 = 9, B58 = 8, B62 = 14) and P25 (A26 = 6; B58 = 8; B62 = 13).

Table 2.

Comparative analysis of the predicted cytotoxic T-lymphocyte (CTL) epitope counts for fibroin subunits distributed across 12 major HLA supertypes.

Supertype alleles FibL FibH P25
A1 9 26 6
A2 10 3 7
A3 7 6 4
A24 8 10 9
A26 9 225 6
B7 6 1 7
B8 0 3 6
B27 6 8 9
B39 7 5 7
B44 2 3 3
B58 8 48 8
B62 14 265 13

Each value corresponds to the number of unique 9-mer peptides predicted by the NetCTL 1.2 server. Abbreviations: FibL, Fibroin Light Chain; FibH, Fibroin Heavy Chain; P25, fibrohexamerin.

Sericin proteins exhibited greater heterogeneity in allele distribution (Table 3). Ser-1, Ser-4, and Ser-5 exhibited the highest number of predicted epitopes across 10 HLA supertypes, with the highest numbers for A1, A26, B27, and B62. Notably, Ser-4 showed the highest total number of CTL epitopes for these supertypes: 210. Ser-3 also showed notable enrichment, especially for A3 (n = 51), B27 (n = 18), and B62 (n = 21). Ser-2 displayed moderate counts across several alleles (A1/B27 = 18, A3 = 19, and B62 = 20).

Table 3.

Comparative analysis of the predicted cytotoxic T-lymphocyte (CTL) epitope counts for sericin proteins distributed across 12 major HLA supertypes.

Supertype alleles Ser-1 Ser-2 Ser-3 Ser-4 Ser-5
A1 64 18 7 52 59
A2 8 4 2 27 14
A3 46 19 51 55 16
A24 10 3 2 25 8
A26 26 7 4 55 11
B7 7 5 2 22 34
B8 2 5 2 30 19
B27 20 18 18 29 32
B39 12 7 2 41 9
B44 6 3 0 35 10
B58 24 7 4 28 10
B62 38 20 21 74 22

Each value corresponds to the number of unique 9-mer peptides predicted by the NetCTL 1.2 server. Abbreviations: Ser-1, sericin 1; Ser-2, sericin 2; Ser-3, sericin 3; Ser-4, sericin 4; Ser-5, sericin 5.

CTL epitopes were predicted using the Next-Generation IEDB Tools platform. The NetMHCpan-4.1 algorithm was used to evaluate the binding affinity of 9-mer peptides to MHC class I molecules, enabling the identification and prioritization of peptides with the highest potential for effective antigen presentation. The binding affinities of different silk protein-derived peptides were assessed using the median binding percentile score (MBPS) as a comparative metric. Based on this parameter, the number of peptides per allele, classified as strong or weak binders, was determined. For MHC class I, peptides were considered strong binders when the median binding percentile was < 0.5, and weak binders when it ranged from 0.5 to 2.0. In contrast, for MHC class II (helper T-cell response), strong binders were defined as MBPS < 2.0, whereas weak binders were defined as MBPS between 2.0 and 10.0.

For fibroin proteins, we observed that FibL (Fig. 9A) and FibH (Fig. 9B) contain several peptides that bind strongly to MHC class I molecules. Notably, in FibL, the HLA-A*26:01 (n = 93), HLA-A*30:02 (n = 255), and HLA-B*35:01 (n = 130) alleles exhibited the highest number of strong-binding peptides, whereas in FibH, the maximum number of strong-binding peptides corresponded to the HLA-A*68:02 allele (n = 9). In contrast, P25 displayed only two strong-binding peptides, both of which were associated with HLA-A*02:01 (not shown). Supplemental Table 1 shows a list of predicted MHC class I-restricted cytotoxic T-cell epitopes from B. mori silk proteins (NetMHCpan-4.1, IEDB platform).

Fig. 9.

Fig. 9

Distribution of MHC-I-binding peptides in fibroin subunits. The horizontal bar chart compares the number of predicted peptides (x-axis) identified across 27 HLA alleles for MHC-I (y-axis). HLA alleles were arranged alphabetically in ascending order, allowing for direct comparison of the binding profiles across fibroin subunit-derived peptides. The bars were segmented by binding strength, with strong (dark blue) and weak (light blue) binders distinguished. Fibroin light chain (A) exhibited the highest number of strong-binders for MCH-I alleles (HLA-A*26:01 = 93, HLA-A*30:02 = 255, and HLA-B*35:01 = 130). Fibroin heavy chain (B) showed the maximum number of strong-binding peptides corresponded to the HLA-A*68:02 allele (n = 9). In contrast, P25 displayed only two strong-binding peptides, both of which were associated with HLA-A*02:01 (not shown). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Regarding MHC class II, FibL (Fig. 10A) exhibited 17 alleles containing between one and four strong-binding peptides, whereas FibH (Fig. 10B) showed the highest number of predicted binders for an allele, with 160 strong-binding peptides interacting with the HLA-DQA1*05:01/DQB1*03:01 allele. P25 (Fig. 10C) exhibited six alleles, each containing between one and three strong-binding peptides. Supplemental Table 2 shows a list of predicted MHC class II-restricted T-cell epitopes from B. mori silk proteins (NetMHCpan-4.1, IEDB platform).

Fig. 10.

Fig. 10

Distribution of MHC-II-binding peptides in fibroin subunits. The horizontal bar chart compares the number of predicted peptides (x-axis) identified across 27 HLA alleles for MHC-II (y-axis). HLA alleles were arranged alphabetically in ascending order, allowing for direct comparison of binding profiles across fibroin subunit-derived peptides. The bars were segmented into strong (dark orange) and weak (light orange) binders. Fibroin light chain (FibL) (A), fibroin heavy chain (FibH) (B), and P25 (fibrohexamerin) (C). FibL exhibited 17 alleles containing between one and four strong-binding peptides, whereas FibH showed the highest number of strong-binders for an MHC-II allele (HLA-DQA1*05:01/DQB1*03:01 = 160). P25 exhibited six alleles, each containing between one and three strong-binding peptides. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

For the sericin proteins, in the case of MHC class I, Ser-1 (Fig. 11A) showed the highest binding diversity, with alleles HLA-A*26:01, HLA-A*30:01, HLA-A*33:01, HLA-A*68:01, HLA-B*15:01, and HLA-B*35:01, each containing more than 10 strong-binding peptides, and HLA-A*01:01 and HLA-A*30:02 containing more than 20 and 40 strong-binding peptides, respectively. For Ser-2 (Fig. 11B), HLA-A*30:01 was the only one that displayed more than 10 strong-binding peptides. Ser-3 (Fig. 11C), HLA-A*03:01 exhibited over 10 strong-binding peptides, whereas HLA-A*11:01 and HLA-A*30:01 showed over 30 strong-binding peptides. Ser-4 (Fig. 11D) exhibited the broadest reactivity, with 22 alleles containing > 10 strong-binding peptides, of which seven contained more than 20 strong-binding peptides (e.g., HLA-A*30:02, HLA-B*35:01, HLA-B*40:01). For Ser-5 (Fig. 11E), the strongest-binding signal was dominated by HLA-A11:01, which showed > 30 strong-binding peptides. Notably, HLA-A*01:01, HLA-A*03:01, HLA-A*11:01, HLA-A*30:01, and HLA-*30:02 alleles consistently displayed high numbers of strong binding peptides between sericins. Supplemental Table 1 shows a list of predicted MHC class I–restricted T-lymphocyte epitopes from B. mori silk proteins (NetMHCpan-4.1, IEDB platform).

Fig. 11.

Fig. 11

Distribution of MHC-I-binding peptides in sericin proteins. The horizontal bar chart compares the number of predicted peptides (x-axis) identified across 27 HLA-MHC II alleles (y-axis). HLA alleles were arranged alphabetically in ascending order, allowing for direct comparison of binding profiles across sericin-derived peptides. Each bar was segmented by the binding strength, distinguishing between strong (dark blue) and weak (light blue) binders. Sericin 1 (A), sericin 2 (B), sericin 3 (C), sericin 4 (D), and sericin 5 (E). Sericin 4 exhibited the broadest reactivity, with seven alleles containing more than 20 strong-binding peptides. Also, HLA-A*01:01, HLA-A*03:01, HLA-A*11:01, HLA-A*30:01, and HLA-*30:02 alleles consistently displayed high numbers of strong binding peptides for MHC-I alleles between sericins. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

For MHC class II, Ser-1 (Fig. 12A) showed six alleles with more than five strong-binding peptides, whereas Ser-2 (Fig. 12B), Ser-3 (Fig. 12C), and Ser-5 (Fig. 12E) did not contain any alleles with more than five strong-binding peptides. In contrast, Ser-4 (Fig. 12D) demonstrated a markedly higher number of predicted strong-binding peptides, with 17 alleles presenting more than 10 strong-binding peptides. Supplemental Table 2 shows a list of predicted MHC class II-restricted T-cell epitopes from B. mori silk proteins (NetMHCpan-4.1, IEDB platform).

Fig. 12.

Fig. 12

Distribution of MHC-II-binding peptides in sericin proteins. The horizontal bar chart compares the number of predicted peptides (x-axis) identified across 27 HLA-MHC II alleles (y-axis). HLA alleles are arranged alphabetically in ascending order, allowing for direct comparison of binding profiles across sericin-derived peptides. Each bar was segmented by the binding strength, distinguishing between strong (dark orange) and weak (light orange) binders. Sericin 1 (A), sericin 2 (B), sericin 3 (C), sericin 4 (D), and sericin 5 (E). Sericin 4 exhibited the highest number of predicted strong-binding peptides, with 17 alleles for MCH-II presenting more than 10 strong-binding peptides. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Once the strong-binding peptides were identified, the total processing score for each MHC-I allele was analyzed to assess the overall efficiency of the antigen-processing pathway. This composite score integrates three key components: (i) the likelihood of proteasomal cleavage generating a specific peptide, (ii) the efficiency of TAP-mediated transport into the endoplasmic reticulum, and (iii) the binding affinity of the peptide for MHC class I molecules. Fig. 13, Fig. 14 illustrate the distributions of total processing scores across all alleles for the fibroin and sericin proteins, respectively. In these plots, the blue dots highlight peptides with scores above the median, indicating a higher probability of triggering CTL responses.

Fig. 13.

Fig. 13

Distribution of the total processing score for strong binders of fibroin proteins across MHC class I alleles. Boxplots illustrate the distribution of the total processing score for fibroin subunit-derived peptides, Fibroin light chain (FibL) (A) and fibroin heavy chain (FibH) (B), predicted as strong binders (median binding percentile < 0.5), grouped by the MHC class I allele. Each box represents the interquartile range (IQR), with the central line indicating the median value. Blue dots highlight peptides with processing total scores above the median, indicating a higher probability of efficiently completing the antigen-processing pathway and eliciting a CTL response. FibL exhibited the highest number of strong-binding peptides across 14 of the 27 HLA alleles, while FibH showed 36 strong-binding peptides distributed across 21 HLA alleles. P25, in contrast, displayed a single strong-binding peptide associated with HLA-A*02:01 (not shown). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Fig. 14.

Fig. 14

Distribution of the total processing score for strong binders of sericin proteins across MHC class I alleles. Boxplots illustrate the distribution of the total processing scores for sericin-derived peptides for sericin 1 (A), sericin 2 (B), sericin 3 (C), sericin 4 (D), and sericin 5 (E) predicted as strong binders (median binding percentile < 0.5), grouped by the MHC class I allele. Each box represents the interquartile range (IQR), with the central line indicating the median value. Blue dots highlight peptides with processing total scores above the median, indicating a higher probability of efficiently completing the antigen-processing pathway and eliciting a CTL response. For sericin proteins, a higher number of strong-binding peptides with processing total scores above the median was observed across at least the third part of the alleles. Sericin 4 exhibited the broadest distribution of strong-binding peptides with processing total scores above the median: at least four strong-binding peptides per allele across all tested alleles. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

Among the fibroins, FibL (Fig. 13A) exhibited the highest number of strong-binding peptides across 14 of the 27 HLA alleles, while FibH (Fig. 13B) showed 36 strong-binding peptides distributed across 21 HLA alleles. P25, in contrast, displayed a single strong-binding peptide associated with HLA-A*02:01 (not shown). For sericin proteins, a higher number of strong-binding peptides with processing total scores above the median was observed across at least the third part of the alleles. Specifically, Ser-1 (Fig. 14A) exhibited this pattern in 23 alleles, whereas Ser-2 (Fig. 14B) showed this pattern in 15 alleles, but with relatively few strong-binding peptides per allele. Ser-3 (Fig. 14C) displayed strong-binding peptides with high processing scores in 11 alleles. Ser-4 (Fig. 14D) exhibited the broadest distribution of strong-binding peptides with processing total scores above the median: at least four strong-binding peptides per allele across all tested alleles. Ser-5 had an average of fewer than three strong-binding peptides per allele across 24 alleles.

From the strongly binding peptides with processing total scores above the median, we examined how processing likelihood relates to predicted T-cell immunogenicity using scatter plots. This analysis aimed to identify peptides with the highest likelihood of being recognized by T-cell receptors (TCRs), assuming successful binding to MHC molecules and consequently triggering an immune response. As shown in Fig. 15, FibL (Fig. 15A) and FibH (Fig. 15B) contained two and four strong-binding peptides, respectively, which were predicted to have the potential to elicit T-cell activation. Similarly, among the sericin proteins, Ser-1 (Fig. 16A) exhibited three strong-binding peptides, whereas Ser-2 (Fig. 16B), Ser-3 (Fig. 16C), Ser-4 (Fig. 16D), and Ser-5 (Fig. 16E) presented four strong-binding peptides. Collectively, these patterns support the idea that FibL, FibH, and sericins harbor multiple candidate epitopes that combine favorable processing likelihood with positive immunogenicity and are thus more consistently poised for T-cell recognition when efficiently processed and presented by MHC molecules.

Fig. 15.

Fig. 15

Correlation between the total processing score and immunogenicity score for predicted strong-binding class I epitopes in fibroin proteins. Scatter plots depict the relationship between the processing total score (X-axis) and immunogenicity score (Y-axis) for peptides identified as strong binders (median binding percentile > 0.5) in fibroin light chain (A), and fibroin heavy chain (B). Point color reflects immunogenicity, and point size tracks the magnitude of the processing score, helping highlight peptides that are both more likely to be generated/presented and more likely to be recognized by T-cell receptors once displayed on MHC molecules (numbered points). The distribution of points provides an integrated view of the antigen-processing efficiency and predicted immunogenic potential of each fibroin subunit.

Fig. 16.

Fig. 16

Correlation between the total processing score and immunogenicity score for predicted strong-binding class I epitopes in sericin proteins. Scatter plots depict the relationship between the processing total score (X-axis) and the immunogenicity score (Y-axis) for peptides identified as strong binders (median binding percentile > 0.5) in sericin 1 (A), sericin 2 (B), sericin 3 (C), sericin 4 (D), and sericin 5 (E). Point color reflects immunogenicity, and point size tracks the magnitude of the processing score, helping highlight peptides that are both more likely to be generated/presented and more likely to be recognized by T-cell receptors once displayed on MHC molecules (numbered points). This distribution illustrates the combined landscape of antigen-processing potential and predicted immunogenicity across the sericin protein family.

For MHC class II analysis, data generated with NetMHCIIpan-4.022 were further examined beyond the count of strong-binding peptides (median binding percentile < 2.0), as shown in Fig. 10, Fig. 12. Specifically, the percentile rank of the MHCII-NP cleavage probability was used to assess the likelihood of proteolytic cleavage leading to peptide generation. Only peptides with percentile ranks below 25.0 were considered, as lower values indicate a higher probability of cleavage compared to alternative peptide sites. This metric provides an additional layer of insight into antigen processing efficiency, complementing binding predictions by estimating the peptides that are most likely to be naturally produced and presented by MHC class II molecules.

Table 4 shows the distribution of high-priority predicted MHC-II peptides for proteolytic cleavage across silk proteins and HLA alleles. For the FibL protein, seven alleles contained two high-priority peptides, whereas nine alleles exhibited only one high-priority peptide. For FibH, 16 alleles contained at least one high-priority peptide, whereas the HLA-DQA1*05:01/DQB1*03:01 allele exhibited the highest number: 113. For the P25 glycoprotein, one allele (HLA-DRB3*02:02) displayed two high-priority peptides, whereas four alleles included a single high-priority peptide.

Table 4.

Distribution of high-priority predicted MHC-II peptides for proteolytic cleavage across silk proteins and HLA alleles.

Allele FibL FibH P25 Ser-1 Ser-2 Ser-3 Ser-4 Ser-5
HLA-DPA1*01:03/DPB1*02:01 0 0 0 1 0 0 7 0
HLA-DPA1*01:03/DPB1*04:01 1 0 0 1 0 0 4 0
HLA-DPA1*02:01/DPB1*01:01 0 0 0 1 1 0 7 0
HLA-DPA1*02:01/DPB1*05:01 0 0 0 2 1 0 12 0
HLA-DPA1*02:01/DPB1*14:01 0 0 0 3 0 1 6 0
HLA-DPA1*03:01/DPB1*04:02 1 0 0 2 0 0 10 0
HLA-DQA1*01:01/DQB1*05:01 2 0 0 4 2 0 5 0
HLA-DQA1*01:02/DQB1*06:02 2 2 0 9 2 1 14 0
HLA-DQA1*03:01/DQB1*03:02 0 3 0 4 3 0 8 0
HLA-DQA1*04:01/DQB1*04:02 1 4 0 8 4 0 17 0
HLA-DQA1*05:01/DQB1*02:01 0 2 0 4 3 0 6 0
HLA-DQA1*05:01/DQB1*03:01 2 113 0 3 0 0 11 0
HLA-DRB1*01:01 1 1 0 1 0 0 11 0
HLA-DRB1*03:01 1 0 1 1 2 0 6 0
HLA-DRB1*04:01 1 3 0 9 3 0 12 0
HLA-DRB1*04:05 1 1 0 6 2 1 13 0
HLA-DRB1*07:01 0 1 0 4 0 0 6 0
HLA-DRB1*08:02 2 0 0 2 2 0 20 0
HLA-DRB1*09:01 0 5 0 4 1 1 13 0
HLA-DRB1*11:01 2 0 0 0 1 0 14 0
HLA-DRB1*12:01 0 0 0 1 1 0 6 0
HLA-DRB1*13:02 1 1 1 3 0 0 11 0
HLA-DRB1*15:01 2 3 0 2 0 0 3 0
HLA-DRB3*01:01 1 1 1 7 3 0 4 0
HLA-DRB3*02:02 2 7 2 6 2 0 9 0
HLA-DRB4*01:01 0 1 0 3 3 0 4 0
HLA-DRB5*01:01 0 3 1 2 3 0 13 0

Abbreviations: FibL, Fibroin Light Chain; FibH, Fibroin Heavy Chain; P25, fibrohexamerin. Ser-1, sericin 1; Ser-2, sericin 2; Ser-3, sericin 3; Ser-4, sericin 4; Ser-5, sericin 5.

For Ser-1, only one analyzed allele lacks high-priority peptides. Notably, HLA-DQA1*01:02/DQB1*06:02 and HLA-DRB1*04:01 alleles showed the highest numbers, with nine high-priority peptides each. In the case of Ser-2, 18 alleles exhibited at least one high-priority peptide, with HLA-DQA1*04:01/DQB1*04:02 showing the highest number: 4. For Ser-3, four alleles exhibited only one high-priority peptide. Ser-4 displayed the broadest distribution, with all 27 alleles presenting three or more high-priority peptides (an average of nine per allele). Among these, HLA-DQA1*04:01/DQB1*04:02 and HLA-DRB1*08:02 stood out, with 17 and 20 high-priority peptides, respectively. In contrast, Ser-5 showed no high-priority MHC-II peptides in any of the 27 allele combinations analyzed.

Finally, Supplemental Table 3 provides a detailed view of high-priority predicted MHC-II peptides across silk proteins and HLA alleles.

3.3.3. B cell epitope prediction

Linear antigenicity analysis of silk proteins (Supplemental Fig. 8) revealed highly variable epitope prediction profiles. The structural proteins FibL (Supplemental Fig. 8A), FibH (Supplemental Fig. 8B), and P25 (Supplemental Fig. 8C) exhibited distinct antigenicity patterns. They identified multiple regions with probability scores exceeding the 0.1512 threshold, indicating a high density of antigenic ‘hotspots’. Sericin proteins also showed complex and extended antigenicity profiles (Supplemental Fig. 8D-G), except for Ser-5, which did not exhibit a detectable antigenicity profile above the defined threshold across the analyzed sequence.

3.3.4. Allergenicity and toxicity evaluation of predicted epitopes

Allergenicity predictions using AllerTOP v.2.1 indicated differential profiles between fibroin subunits and sericin proteins (Table 5). FibL and FibH were both classified as probable allergens, similar to known allergenic proteins such as maltogenic α-amylase and a 30 kDa fragment of a velvet grass allergen, respectively. In contrast, P25 and Ser-1, Ser-3, and Ser-4 were classified as probable non-allergens, showing the closest similarity to proteins such as aldehyde oxidase homologs (P25), corneodesmosin (Ser-1 and Ser-3), and chromodomainhelicase-DNA-binding protein 9 (Ser-4). Notably, Ser-2 and Ser-5 were predicted to be probable allergens, showing similarity to the human PC4- and SFRS1-interacting protein.

Table 5.

AllerTOP v.2.1 predicts the allergenicity of fibroin subunits and sericin proteins.

Protein Most similar protein Classification based on the most similar protein with AllerTOP v.2
FibL sp|P19531|AMYM_GEOSE Maltogenic alpha-amylase OS = Geobacillus stearothermophilus GN = amyM PE = 1 SV = 2 Probable ALLERGEN
FibH gi|543482|pir||S38291 30 K allergen − velvet grass (fragment) Probable ALLERGEN
P25 tr|Q7DM89|Q7DM89_SOLLC Aldehyde oxidase 1 homolog (Fragment) OS=Solanum lycopersicum GN = TAO1 PE = 1 SV = 1 Probable NON-ALLERGEN
Ser-1 sp|Q15517|CDSN_HUMAN Corneodesmosin OS=Homo sapiens GN=CDSN PE = 1 SV = 3 Probable NON-ALLERGEN
Ser-2 sp|O75475|PSIP1_HUMAN PC4 and SFRS1-interacting protein OS=Homo sapiens GN=PSIP1 PE = 1 SV = 1 Probable ALLERGEN
Ser-3 sp|Q15517|CDSN_HUMAN Corneodesmosin OS=Homo sapiens GN=CDSN PE = 1 SV = 3 Probable NON-ALLERGEN
Ser-4 sp|Q3L8U1|CHD9_HUMAN Chromodomain-helicase-DNA-binding protein 9 OS=Homo sapiens GN=CHD9 PE = 1 SV = 2 Probable NON-ALLERGEN
Ser-5 sp|O75475|PSIP1_HUMAN PC4 and SFRS1-interacting protein OS=Homo sapiens GN=PSIP1 PE = 1 SV = 1 Probable ALLERGEN

Allergenicity assessment was performed using AllerTOP v.2, based on an auto-cross-covariance (ACC) transformation and classification with a k-nearest neighbors (k = 1) algorithm. The analysis provided the most similar protein match from the training dataset, along with its corresponding classification as either a probable allergen or a non-allergen. Abbreviations: FibL, Fibroin Light Chain; FibH, Fibroin Heavy Chain; P25, fibrohexamerin. Ser-1, sericin 1; Ser-2, sericin 2; Ser-3, sericin 3; Ser-4, sericin 4; Ser-5, sericin 5.

All analyzed silk proteins (Table 6) were consistently predicted to be nontoxic using ToxinPred 3.0. The machine learning (ML) scores ranged from 0.13 to 0.33, indicating uniformly low toxicity potential. All proteins matched a nontoxic motif (MERCI Score −ve = −0.50), resulting in a hybrid score of 0.00 for most sequences. P25 and Ser-5 also showed alignment with a toxic motif (MERCI Score +ve = 0.50), and their final hybrid score remained low (<0.35). This nontoxic classification is further supported by the positive predictive values (PPVs), which were 0.00 for nearly all proteins and <0.30 for P25 and Ser-5.

Table 6.

ToxinPred v3.0 prediction of fibroin and sericin proteins.

Subject ML score MERCI score (+ve) MERCI score (−ve) Hybrid score Prediction PPV
FibL 0.20 0.00 −0.50 0.00 Non-Toxin 0.00
FibH 0.15 0.00 −0.50 0.00 Non-Toxin 0.00
P25 0.33 0.50 −0.50 0.33 Non-Toxin 0.27
Ser-1 0.15 0.00 −0.50 0.00 Non-Toxin 0.00
Ser-2 0.25 0.00 −0.50 0.00 Non-Toxin 0.00
Ser-3 0.13 0.00 −0.50 0.31 Non-Toxin 0.00
Ser-4 0.15 0.00 −0.50 0.00 Non-Toxin 0.00
Ser-5 0.19 0.50 −0.50 0.29 Non-Toxin 0.10

Results from the ToxinPred 3.0 prediction module showing the ML Score, MERCI score (+ve), MERCI score (−ve), Hybrid Score, Prediction, and Positive Predictive Value (PPV) for silk proteins. Predictions were generated using a hybrid model implemented on the server, which integrates motif analysis (MERCI) with the best-performing machine learning classifier (Extra Tree, ET) based on amino acid composition (AAC) and dipeptide composition (DPC). All proteins were consistently classified as nontoxic under both the individual and hybrid scoring strategies. Abbreviations: FibL, Fibroin Light Chain; FibH, Fibroin Heavy Chain; P25, fibrohexamerin. Ser-1, sericina 1; Ser-2, sericina 2; Ser-3, sericina 3; Ser-4, sericina 4; Ser-5, sericina 5.

Therefore, ToxinPred 3.0 consistently classified all analyzed silk proteins, including FibL, FibH, P25, and Ser-1 to -5, as nontoxic. Both the machine learning and hybrid scores remained well below the established toxicity threshold (0.38), leading to a final prediction of ‘non-toxic’ for every sequence. When these toxicity predictions are considered alongside the allergenicity results, the findings suggest that while silk proteins pose no toxicological risks, specific components, particularly FibL, FibH, Ser-2, and Ser-5, may still harbor potential allergenic properties.

4. Discussion

The emergence of AI-driven structural prediction tools, such as AlphaFold3, has revolutionized the analysis of protein structure; however, integrating these structural insights with computational immunology for biomaterial design remains largely underexplored. While silk-based biomaterials have demonstrated medical utility,2 the structural characterization of their protein components presents experimental challenges. Furthermore, a fundamental paradox persists: despite degumming protocols removing sericins, traditionally considered the primary immunogenic fraction, adverse reactions still occur, indicating that the molecular basis of silk immunogenicity is more complex than previously thought. In this study, we performed a comprehensive in silico characterization integrating AlphaFold3 structural prediction with multi-algorithm immunoinformatic profiling of eight major B. mori silk proteins. Our structural analysis revealed that FibL and fibrohexamerin (P25) adopt well-defined, ordered architectures, whereas FibH and sericins exhibited significant conformational plasticity, likely undergoing templated transitions during hierarchical assembly. Crucially, immunogenic profiling revealed striking heterogeneity: P25 displayed minimal immunogenicity, while FibL and specific sericins (Ser-4 and Ser-1) emerged as highly antigenic with dense epitope distributions across multiple HLA supertypes. These findings fundamentally challenge the ‘sericins-only’ paradigm and provide a mechanistic explanation for the immunogenicity in degummed silk, establishing a rational foundation for designing HLA-matched, biocompatible biomaterials.

To understand the molecular drivers of the structural plasticity and immunogenic heterogeneity of silk components, we analyzed the fundamental compositional and biophysical differences between fibroin and sericin proteins. Our sequence analysis confirmed that FibH’s extreme Gly-Ala enrichment (45.83% and 30.18%) generates the repetitive (GAGAGX)n motifs that drive antiparallel β-sheet crystallite formation. These crystalline motifs are responsible for silk's tensile strength,7, 50, 51 a finding consistent with established X-ray diffraction studies and solid-state nuclear magnetic resonance (ss-NMR).52, 53, 54 In contrast, FibL exhibits a non-repetitive, acidic sequence that lacks these crystalline motifs, suggesting a role centered on structural flexibility and complex stabilization rather than mechanical rigidity. Our structural prediction supports its function as a chaperone-like protein, maintaining the stability of the FibH-FibL complex through the conserved disulfide bond (Cys-c20 of FibH and Cys-172 of FibL), which is essential for the FibH6:FibL6:P251 molar ratio assembly within the posterior silk gland lumen.1, 55, 56 Similarly, P25 displayed the highest cysteine content (4.09%) and stands as the only N-glycosylated component, maintaining fibroin secretory globule morphology in the posterior silk gland lumen.1, 57 Conversely, sericins exhibited dramatically different profiles, characterized by serine enrichment (up to 43.59%) and a high polar/charged content that confers significant hydrophilicity (GRAVY < −0.5). This is functionally consistent with their role as water-soluble coating proteins that maintain fiber hydration.58, 59 At the same time, their abundant hydroxyl groups facilitate O-glycosylation and modulate the critical solubility transitions required during silk assembly and spinning.60, 61

Beyond primary sequence divergence, the biophysical parameters of silk proteins underscore their functional specialization. FibH exhibited extraordinary molecular weight heterogeneity (350–800 kDa), mirroring the diversity seen in Ser-5 (75–400 kDa). This variation likely reflects two distinct biological mechanisms: genetic polymorphism within the massive repetitive domains and proteolytic processing of FibH, consistent with reported size variations across B. mori strains,62, 63 and extensive alternative splicing in the case of Ser-5. In contrast, the FibL/P25 subunits (25–30 kDa) and the other sericins (120–250 kDa) displayed consistent molecular weights, reflecting their roles as structural anchors. Our analysis also revealed a striking pI distribution: FibL, FibH, and P25 showed pI bimodality (4–5 versus 6.5–9.5), potentially reflecting isoforms with distinct roles in silk gland compartmentalization and fiber pH-dependent assembly.64 Conversely, sericins display acidic pI values (5.5–6.5), with the notable exception of Ser-5, which shows a broad pI range (4–10). This predominantly negative charge at physiological pH promotes electrostatic repulsion that prevents premature aggregation, a property exploitable for the rational design of pH-responsive materials. Finally, the GRAVY index distribution (near-zero values for fibroins versus strongly negative values for sericins) confirms the hydrophobic-hydrophilic dichotomy essential for the core–shell architecture of silk fibers.65 These physicochemical insights provide a roadmap for optimizing recombinant expression systems and downstream purification strategies, such as isoelectric precipitation and ion-exchange chromatography, for industrial-scale silk production in biotechnology applications.

These biophysical traits connect with the conformational landscapes revealed by AlphaFold3, which showed a spectrum of structural order with profound implications for silk assembly. FibL exhibited a predominantly α-helical architecture with very high confidence (pLDDT > 90, ipTM/pTM = 0.85), consistent with its chaperone-like role in preventing FibH retention in the endoplasmic reticulum through disulfide bonding.66 Similarly, P25 displayed a well-defined globular β-sheet fold with very high to high confidence (pTM = 0.85), supporting its stabilizing function in the fibroin complex.66 In contrast, the model for FibH revealed extensive β-sheet propensity but paradoxically low confidence (pLDDT 50–70, ipTM/pTM = 0.28). This apparent discordance reflects a fundamental limitation of AlphaFold3 when processing the native-length FibH sequence, which exceeds 5,000 residues composed of highly repetitive, low-complexity Gly-Ala motifs (GAGAGX)n.15, 18 Such extreme sequence repetition lies outside the distribution of protein architectures in AlphaFold's training data (the PDB), which is dominated by globular proteins with well-defined hydrophobic cores and limited internal symmetry. Consequently, the algorithm struggles to confidently predict the extended, crystalline antiparallel β-sheet architecture that experimental techniques (X-ray diffraction, solid-state NMR) have definitively established for spun silk fibers.53, 54, 67 Critically, these low confidence scores likely capture a biological reality: FibH exists as an intrinsically disordered precursor before fiber assembly, undergoing disorder-to-order transitions only during spinning through mechanochemical stimuli (shear stress, dehydration, pH gradient). The crystalline fraction ultimately adopts a lamellar structure with repetitive folding patterns, featuring β-turns every eighth amino acid and antipolar arrangement of side chains.68 Thus, AlphaFold's uncertainty appropriately reflects the conformational heterogeneity of unassembled FibH in solution, while experimental crystallographic data captures its post-assembly, mechanically-induced final state. This divergence underscores the complementary nature of computational prediction and experimental validation for mapping the complete conformational trajectory from disordered monomers to hierarchical silk architectures.

Complementing the AlphaFold predictions, IUPred analysis revealed substantial intrinsic disorder across the silk proteome, with profound functional implications for fiber assembly. FibH behaved as a predominantly disordered protein, with stable structural regions confined strictly to the N- and C-termini. This is consistent with the established roles of these termini in pH-dependent self-assembly and disulfide bonding during silk fiber maturation.54, 62, 69 The high disorder propensity of FibH's repetitive core aligns with the mechanochemical model of spinning, where crystalline β-sheets form through disorder-to-order transitions triggered by mechanical shear, dehydration, and the steep pH gradient (8.2 → 4.8) characteristic of the silk gland, as previously mentioned.54, 64, 70, 71, 72 Ser-1, 2, and 3 displayed classic IDP behavior, with stabilization points likely linked to protein–protein interactions and redox-dependent conformational changes, consistent with their role in modulating long-range conformational control in fibroin.10, 73 Notably, Ser-4 displayed a distinct modularity, with two large IDRs flanking ordered segments (positions 10–300 and 1300–1700), whereas Ser-5 showed a single ordered N-terminal segment (positions 1–300). This modular architecture suggests that these sericins function as multifaceted adaptors, potentially enabling dual roles in both fiber coating and stabilizing the liquid silk dope during storage.

The coordination of metal ions, particularly Cu2+, has been proposed to influence fiber assembly through interactions of fibroin chains with histidine residues, as Cu2+ concentrations increase progressively along the silk gland's secretory pathway.49, 74, 75, 76 In parallel, our analyses suggest that the structural conformation of these disordered proteins is sensitive to the redox state, which is closely linked to their ability to coordinate metals and form cysteine bridges. Incorporating a Cu2+ coordination modeling (3.4:1 Cu2+:His ratio)49 revealed a potential templating mechanism. While the complete molar complex should, in theory, involve 133 Cu2+ ions, AlphaFold3 is currently limited to modeling a maximum of 50 ions per prediction. Consequently, we selected a stoichiometry of 48 Cu2+ ions, based on the available coordination sites for a representative FibL1-P251 heterodimer. This FibL1-P251-Cu2+48 complex exhibited enhanced α-helical organization, supporting the hypothesis that metal ions may serve as architectural organizers during fibroin complex assembly in the silk gland, as the environment transitions from the posterior-middle to the anterior-middle sections of the silk gland.49, 76 Most remarkably, the interaction of this metalloprotein complex with intrinsically disordered sericins induced dramatic conformational ordering: Ser-1 and 4 adopted α-helical and parallel β-sheet conformations, Ser-3 formed α-helices and antiparallel β-sheets, and Ser-2 adopted primarily α-helical structure. This disorder-to-order transition upon complex binding suggests that the structured FibL-P25-Cu2+ scaffold acts as a nucleating center for conformational organization in otherwise disordered sericins, analogous to the induced fit seen in chaperone-mediated folding.77, 78

Furthermore, the observed global confidence of the resulting FibL6-P25- Cu2+48-Ser assemblies remained low (pTM ≈ 0.30–0.31) with modest interface scores (ipTM ≈ 0.24–0.27). Significantly, this contrasts with the high- confidence predictions obtained for FibL and P25 individually (pTM > 0.80), indicating that the reduced confidence upon sericin binding does not stem from the same algorithmic limitations observed with FibH's extreme sequence repetition. Instead, we interpret these persistently low scores as evidence of genuine conformational heterogeneity in the bound state, a hallmark of ‘fuzzy complexes’, in which intrinsically disordered proteins retain structural plasticity even when complexed with structured partners.79 Three converging lines of evidence support this interpretation: (i) sericins exhibit extensive IDRs (demonstrated by IUPred analysis) that undergo partial ordering upon metalloprotein complex binding, adopting α-helical and β-sheet elements while retaining low-pLDDT segments (indicating local disorder); (ii) the induced secondary structures vary between sericin paralogs (Ser-1/4: α-helix + parallel β-sheet; Ser-3: α-helix + antiparallel β-sheet; Ser-2: predominantly α-helix), suggesting multiple binding modes rather than a single rigid conformation; and (iii) the consistently low ipTM values indicate that the sericin-fibroin interface lacks the tight, geometrically-defined complementarity characteristic of obligate complexes, instead relying on distributed, multivalent contacts mediated by short structured patches embedded within disordered regions. This conformational flexibility is functionally consistent with the proposed adhesive role of sericins within the silk coating.10 Rather than forming stable, ‘lock-and-key’ interactions, the sericin-fibroin interface likely employs dynamic sampling of multiple conformational substates, enabling adaptation to chemically heterogeneous surfaces, a binding mechanism observed in other adhesive proteins such as mussel foot proteins.80, 81 In this framework, adhesion emerges not from a single high-affinity interaction, but from the collective contributions of numerous weak, transient contacts distributed across a large conformational ensemble.82 Therefore, we argue that the low pTM and ipTM values observed in these complexes are not indicative of prediction failure but rather reflect an adhesive mechanism that exploits structural disorder as a functional feature. Experimental validation using hydrogen–deuterium exchange mass spectrometry (HDX-MS) or single-molecule FRET could directly test this hypothesis by mapping the conformational dynamics of sericins in bound and unbound states.83, 84, 85

The observed disorder-to-order transitions likely arise from multiple cooperative mechanisms: (i) intermolecular disulfide bonding between FibH (Cys-c20) and FibL (Cys-172) forming covalent FibL-FibH heterodimers; (ii) non-covalent hydrophobic interactions between P25 and FibH, or FibH-FibL complex; (iii) hydrogen-bonding networks between the carboxyl or amide side groups of fibroin chains and the serine-rich sericins, (iv) pH-driven protonation state changes through the silk gland that eliminate electrostatic repulsions while promoting solvent exclusion,63, 73 and (v) the hydrophobic effect induced by mechanical stress and dehydration. Likewise, Cu2+ ions may function as architectural organizers by: (i) crosslinking histidine residues, creating nucleation sites for conformational ordering, (ii) stabilizing transition states through coordination geometry constraints, and (iii) modulating local electrostatic environments to favor folding. This mechanism parallels the role of metallochaperones in copper-binding proteins, in which metal coordination precedes structural maturation.86, 87 In this way, the FibL-P25-Cu2+ complex could act as a structural template for FibH and sericins, enabling a hierarchical assembly: disordered synthesis in the posterior gland, metal-mediated nucleation upon FibL-P25 binding in the middle gland, and final crystallization via pH drop and shear in the anterior gland.

Integrating these structural insights with immunoinformatic profiling provides a molecular rationale for the reported immunogenicity of silk biomaterials. This phenomenon has long been a subject of debate in the field of clinical application. Historically, reports of foreign-body reactions to silk sutures raised significant concerns about their biocompatibility,88, 89, 90 leading to an apparent dichotomy in experimental studies: inflammatory responses were attributed almost exclusively to residual sericin proteins, while properly degummed fibroin was thought to exhibit minimal immunogenicity.91, 92 However, recent investigations have partially exonerated sericin from its reputation as the sole immunogenic component, demonstrating that purified sericin biomaterials elicit only mild innate immune responses and exhibit negligible allergenicity as standalone materials.93, 94, 95 Our comprehensive immunoinformatic analysis, using complementary prediction algorithms, revealed pronounced antigenic heterogeneity with potential for HLA-dependent risk stratification. Critically, it is essential to distinguish between antigenicity, the potential for recognition by the adaptive immune system, and immunogenicity, the capacity to elicit a full immune response.39, 96 While all immunogenic substances are antigenic, not all antigenic substances possess the context-dependent properties required for immunogenicity.39, 96 This principle guided our multi-tiered analysis. VaxiJen 2.0 antigenicity scoring identified P25 as the only protein with sub-threshold antigenicity (0.395), whereas FibH and Ser-1, 2, 3, and 5 emerged as highly antigenic (0.8–1.0), and FibL and Ser-4 were classified as moderately antigenic (0.5–0.8). While VaxiJen provides initial prioritization of antigenic potential, determining which of these antigenic proteins can achieve immunogenicity requires deeper analysis of MHC binding affinity, proteasomal processing efficiency, and TAP transport compatibility, properties captured through NetCTL 1.2 and NetMHCpan-4.1 epitope prediction.

Integration of NetCTL 1.240 and NetMHCpan-4.122 epitope predictions -accounting for binding affinity, proteasomal cleavage probability, and TAP transport efficiency22-, revealed a pronounced HLA class I supertype bias and established a clear immunogenic risk hierarchy for CTL responses. FibH exhibited extreme epitope enrichment within specific supertypes, with 225 epitopes in A26, 48 in B58, and 265 in B62, representing the highest epitopes per-supertype densities across the silk proteome. In contrast, FibL displayed a different distribution pattern: while exhibiting lower epitope counts per-supertype, it accumulated the highest total number of strong-binding peptides (n = 478) when summing across its three primary restriction alleles (HLA-A*26:01, A*30:02, and B*35:01), indicating a narrower but still substantial HLA coverage. Sericin-1 and Ser-4 showed the strongest CTL signatures across multiple supertypes (A01, A3, A26, B27, B62), with Ser-4 emerging as the highest-risk component: 7 of 27 tested alleles contained ≥ 20 strong-binding peptides, and all 27 alleles presented ≥ 4 strong binders. Ser-1 exhibited a more restricted profile, with 2 alleles containing ≥ 20 peptides. Across the sericin family (Ser-1 through Ser-5), consistent epitope enrichment was observed in HLA-A*01:01, A*03:01, A*11:01, A*30:01, and A*30:02 alleles, though individual sericin isoforms varied in overall epitope burden. In stark contrast, P25 exhibited only two strong-binding peptides restricted to HLA-A*02:01, positioning it at the opposite end of the immunogenic spectrum. Analysis of total processing scores, to assess the overall efficiency of the antigen-processing pathway, highlighted a clear hierarchy: Ser-4 displayed a robust processable epitope burden, with > 4 strong-binding peptides across all 27 tested alleles, even though sericins exhibited strong-binding peptides with processing total scores above the median in at least three-quarters of the alleles. Furthermore, FibH showed high coverage across 21 of 27 alleles, while FibL exhibited the highest number of strong-binding peptides across 14 of the 27 HLA alleles. In stark contrast, P25 exhibited only one processable epitope for HLA-A*02:01. Immunogenicity scoring predicting TCR recognition probability and high-priority epitope analysis,21, 22 integrating MHC-II binding affinity with endosomal processing efficiency (percentile rank < 25.0), revealed a clear MHC-II-restricted immunogenic hierarchy. Among fibroin subunits, FibH displayed the broadest MHC-II coverage (16 of 27 alleles with ≥ 1 high-priority peptide) and the highest single-allele density (113 high-priority peptides for HLA-DQA1*05:01/DQB1*03:01). FibL showed moderate coverage (9 of 27 alleles with ≥ 1 high-priority peptide, including 7 alleles presenting ≥ 2 high- priority peptides). In contrast, P25 exhibited minimal MHC-II reactivity, with only one allele (HLA-DRB3*02:02) presenting 2 high-priority peptides. Ser-4 exhibited the most robust CD4+ T-cell signature, with an average of 9 high-priority peptides across all 27 evaluated alleles, followed closely by Ser-1 (≥1 high-priority epitope in 26 of 27 alleles). These data not only reinforce the known immunogenicity of sericins, identifying Ser-4 and Ser-1 as the primary drivers of CD8+ and CD4+ T-cell responses, but also demonstrate that FibH and FibL possess significant potential to elicit similar adaptive responses. These findings suggest that residual fibroin immunogenicity, particularly from FibH and FibL, could contribute to the variable biocompatibility outcomes observed with silk-based biomaterials,88, 89, 90 challenging the assumption that degumming alone is sufficient to eliminate immunogenic risk. Finally, these findings highlight that individuals with specific high-risk HLA supertypes are more prone to activation of adaptive response upon exposure to silk protein-containing biomaterials. In contrast, individuals with underrepresented HLA types may exhibit greater tolerance, potentially explaining the historical inconsistency in reports of silk biocompatibility.

Linear B-cell epitope prediction using BepiPred algorithms revealed complex antigenic landscapes across silk proteins (except for Ser-5), with multiple regions surpassing the 0.1512 probability threshold indicative of antibody recognition sites capable of eliciting IgM or IgG responses.42 Concurrently, sericins displayed extended epitope-dense areas consistent with their abundance of surface-exposed polar residues (Ser, Thr, Glu, Lys) that serve as primary B-cell recognition determinants.97, 98, 99 On the other hand, allergenicity assessment using AllerTOP 2.044 yielded findings with significant clinical relevance: FibL and FibH were classified as probable allergens, exhibiting structural/physicochemical similarities to known allergens, including maltogenic α-amylase and a 30 kDa fragment of a velvet grass allergen, respectively, the latter being explicitly associated with Type I allergic reactions.100, 101 Ser-2 and Ser-5 were also classified as probable allergens. However, interpreting these predictions requires caution, as the structural similarities identified by the algorithm may not directly translate into clinical allergenicity, particularly when they involve mammalian proteins subject to immunological tolerance. In striking contrast, P25, along with Ser- 1, 3, and 4, was classified as a non-allergen. Importantly, ToxinPred 3.045 analysis classified all silk proteins as non-toxic (ML scores 0.13–0.33, well below the 0.38 toxicity threshold). These findings provide computational evidence that immunogenic concerns pertain exclusively to adaptive immune recognition, rather than innate cytotoxicity. This distinction fundamentally distinguishes silk proteins from microbial toxins that cause direct cellular damage.

The strikingly low immunogenicity of P25 across all predictive metrics arises from a synergy of structural and biochemical features that collectively minimize immune recognition. First, P25′s compact β-sheet architecture (pLDDT > 90, ipTM/pTM = 0.85) combined with extensive disulfide bonding (4.09% cysteine content) confers high structural stability, which likely resists endosomal proteolysis.102, 103, 104 This resistance reduces the efficiency of MHC class II peptide generation, as proteins with extensive disulfide networks and structural rigidity limit cathepsin accessibility to cleavage sites.102, 103, 104 Additionally, P25′s compact globular fold may sequester potential linear epitopes within the hydrophobic core, limiting their accessibility to B-cell receptor recognition, which requires surface-exposed epitopes on the native folded protein. Second, P25 is the only N-glycosylated silk protein, with glycan attachment sites predicted at multiple asparagine residues at positions 69, 113, and 133.105 These bulky carbohydrate moieties likely create a 'glycan shield' that sterically occludes underlying peptide epitopes from immune surveillance. This mechanism parallels the immune evasion strategies of viral envelope proteins, such as influenza hemagglutinin and HIV gp120, which exploit glycosylation to mask antigenic sites from antibody recognition.106, 107 The convergence of structural compactness, disulfide rigidity, and N-glycosylation supports P25 as a candidate hypoimmunogenic component suitable for enrichment in biocompatible silk formulations.

In stark contrast to the hypoimmunogenic profile of P25, FibL, Fib-H, Ser-4, and Ser-1 emerge as the primary immunological hurdles in silk-based biomaterials, albeit through distinct molecular mechanisms. FibL, despite its well-defined α-helical architecture (pLDDT > 90, ipTM/pTM = 0.85) and compact globular fold, poses a significant immunogenic risk due to a focused HLA-restriction pattern. Our analysis revealed that FibL concentrates 478 strong-binding epitopes within three specific alleles (HLA-A*26:01, A*30:02, B*35:01), creating a narrow but high-density epitope distribution. This focused restriction profile has critical clinical implications: while FibL may exhibit lower population-level immunogenicity than more broadly reactive components, it poses a disproportionately high risk for individuals carrying these specific HLA haplotypes, a crucial consideration for personalized, HLA-matched biomaterial selection. Furthermore, FibL demonstrated MHC-II reactivity across 9 of 27 tested alleles, suggesting the potential to activate CD4+ T cells in a clinically relevant subset of the population. Importantly, FibL's ordered structure, while conferring some resistance to complete proteolytic degradation, still provides surface-accessible epitopes for B-cell recognition and sufficient cathepsin cleavage sites for MHC-II antigen presentation.108, 109, 110 FibH, in contrast to FibL's focused restriction, exhibits immunogenic potency driven by extraordinary molecular scale and repetitive architecture, resulting in pan-reactive HLA supertype coverage. Its massive surface area statistically increases the likelihood of harboring MHC-binding motifs across diverse HLA backgrounds, as confirmed by our identification of > 500 strong- binding epitopes distributed across HLA-A*26, B*58, and B*62 supertypes, representing the broadest epitope distribution among all fibroin components. Unlike both P25′s sequestered core and FibL's compact fold, FibH's extended conformation and low-complexity repetitive domains facilitate rapid endosomal processing and efficient MHC-II loading.108, 109, 110 Furthermore, the highly repetitive (GAGAGX)n motifs may induce ‘multivalent“ immune responses; the recurrence of identical or near-identical epitopes along the primary sequence can promote cross-linking of B-cell receptors, a potent activation signal that often bypasses tolerance thresholds.111 This risk could be exacerbated by crystalline structures within degummed silk, as analogous microcrystalline biomaterials trigger persistent foreign body responses and macrophage recruitment, providing co-stimulatory signals that can transform antigenic peptides into chronic inflammatory reactions.112 However, direct experimental evidence linking FibH crystallites to adjuvant activity remains to be established.

Parallelly, Ser-4 exhibits the broadest immunogenic reactivity among all silk proteins, driven by its compositional heterogeneity and modular architecture, for both cytotoxic T-lymphocyte (CD8+) and helper T-cell (CD4+) responses, enabling generation of peptides with diverse physicochemical properties that bind multiple MHC-I supertypes (A01, A03, A26, B39, B62) and, remarkably, all 27 tested MHC-II alleles (≥3 high-priority peptides for clevage each). Its architecture, two large IDRs (positions 10–300 and 1300–1700) interspersed with ordered segments, enhances this profile through multiple mechanisms: compositional heterogeneity provides diverse proteasomal cleavage motifs generating 8–11 amino acid MHC-I-compatible peptides,113, 114 while extensive IDRs offer abundant cathepsin-accessible sites for efficient generation of 13–25 amino acid MHC-II peptides.108, 109, 110 Notably, order–disorder transitions may serve as preferential proteolytic sites, further amplifying peptide repertoire diversity.115 Ser-1 follows a similar immunogenic mechanism but with more restricted HLA coverage: while presenting ≥ 1 high-priority MHC-II peptide in 26 of 27 alleles (nearly as comprehensive as Ser-4), it exhibits narrower MHC-I reactivity (only two alleles with ≥ 20 epitopes for supertypes compared to Ser-4′s seven alleles). This pattern suggests that Ser-1 may predominantly drive CD4+ helper T-cell responses with limited CD8+ cytotoxic activation, whereas Ser-4 may pose a risk to comprehensive adaptive immunity.

These comprehensive epitope maps enable rational design strategies for biocompatible silk-based biomaterials: (1) patient-specific HLA-matched biomaterial selection, wherein individuals carrying HLA supertypes associated with elevated immunogenic risk (e.g., A1, A3, A26, B58, B62 for MHC-I-restricted CD8+ responses; specific HLA-DR and HLA-DQ alleles showing high reactivity to FibH and Ser-4 for CD4+ responses) would receive formulations depleted of high-risk components116, 117, 118; (2) P25-enriched formulations that maximize biocompatibility by leveraging fibrohexamerin’s intrinsically hypoimmunogenic profile; (3) structure-guided deletion or substitution mutagenesis targeting predicted high-affinity T-cell and B-cell epitopes in FibH and Ser-4, with the critical constraint that FibH's repetitive (GAGAGX)n crystalline domains essential for mechanical strength must be preserved; (4) site-specific PEGylation119 to shield surface-exposed B-cell epitopes, though this strategy does not prevent intracellular antigen processing and MHC-mediated T-cell activation; and (5) potential incorporation of T-regulatory cell epitopes derived from tolerogenic proteins120 to induce active immune tolerance, though this approach remains experimental mainly for biomaterial applications. Among these strategies, P25 enrichment and HLA-matched component selection represent immediately translatable approaches based on the present computational predictions, whereas epitope engineering and Treg incorporation require additional experimental validation.

Despite the comprehensive nature of this computational analysis, several limitations must be acknowledged. First, while AlphaFold3 represents a paradigm shift in structural biology, its performance on silk proteins revealed both informative constraints and genuine limitations. For FibH, the low confidence scores appropriately reflect the intrinsically disordered nature of the pre-assembly precursor; however, the algorithm cannot predict the mechanically induced crystalline β-sheet architecture that forms during fiber spinning.18, 121, 122, 123 Characterizing the solution-state conformational ensemble would require circular dichroism spectroscopy, small-angle X-ray scattering (SAXS), and HDX-MS, whereas the post-assembly crystalline state is best validated through X-ray fiber diffraction and ss-NMR, techniques that have already established the antiparallel β-sheet signatures in spun fibers.52, 53, 83, 124, 125, 126, 127 Similarly, our interpretation of low-confidence FibL-P25-Cu2+-sericin unit as ‘fuzzy complex’79 remains a hypothesis. Direct experimental validation via HDX-MS, advanced NMR for IDPs, or cryo-electron microscopy of complexed states is necessary to distinguish genuine conformational heterogeneity from algorithmic uncertainty.126, 127, 128 Second, our Cu2+ coordination model is a theoretical approximation constrained by AlphaFold3′s maximum of ligands per prediction. While we prioritized copper due to its well-documented stoichiometry and concentration gradients in the silk gland,49, 74, 75, 76 other ions (Zn2+, Ca2+, Mg2+) may also contribute to assembly. However, the current literature lacks stoichiometric data at the residue level. Rigorous biochemical validation using XAS, electron paramagnetic resonance (EPR), and inductively coupled plasma mass spectrometry (ICP-MS) is essential to confirm these coordination sites, stoichiometries, and their functional roles in conformational ordering.129, 130 Third, immunoinformatic predictions inherently oversimplify in vivo complexity, as they do not account for antigen-processing kinetics, immunological tolerance mechanisms, or the influence of the local tissue microenvironment on immune response activation.131, 132 The absence of experimental validation using HLA-transgenic mouse models remains a critical gap for future investigations.133 Finally, our exclusive focus on linear epitopes may overlook conformational epitopes arising from the quaternary structure of native fibers or biomaterial scaffolds.134 Nonetheless, these computational predictions provide a rational foundation for prioritizing experimental validation efforts and guiding initial biomaterial design strategies.

The implications of this study extend far beyond silk proteins, establishing a computational framework for immunological risk stratification of protein-derived biomaterials before costly experimental validation. By integrating AlphaFold3 structural prediction with multi-algorithm epitope mapping, we demonstrated that immunogenicity is a multifaceted property amenable to rational deconstruction and prospective engineering. Our findings challenge the historical paradigm that attributes silk immunogenicity exclusively to sericins, revealing that FibH and Ser-4 harbor the most extensive epitope repertoires, whereas P25 emerges as a uniquely hypoimmunogenic component suitable for next-generation biocompatible formulations.

This work opens several high-impact translational avenues: (1) patient-specific HLA- matched silk biomaterial selection to minimize rejection risk; (2) structure-guided epitope deletion, substitution mutagenesis, or insertion of immunomodulatory sequences in immunogenic proteins to generate hypoimmunogenic recombinant variants; (3) chimeric silk proteins combining FibH's mechanical properties with P25′s biocompatibility profile; (4) P25-enriched formulations that exploit fibrohexamerin's intrinsic glycan shielding and structural compactness. Experimental validation through peptide-MHC binding assays, T-cell activation studies, and HLA-transgenic mouse implantation models will be essential to translate these computational predictions into clinically viable biomaterials.

Looking forward, the integration of artificial intelligence-driven structural biology with experimental immunopeptidomics and genome-editing technologies enables systematic engineering of immunologically optimized biomaterials that retain structural functionality while minimizing adaptive immune recognition.135, 136 Ultimately, this study exemplifies how computational biology transforms empirical trial-and-error into hypothesis-driven, mechanistically informed engineering strategies that bridge molecular structure and clinical biocompatibility.

5. Conclusions

This study integrates AlphaFold3 structural prediction with multi-algorithm immunoinformatic analysis to establish structure-immunogenicity relationships across the Bombyx mori silk proteome, providing a computational blueprint for rational biomaterial design. Structural modeling revealed that FibL and P25 adopt well-defined, high-confidence architectures (pTM = 0.85), whereas FibH exhibits low confidence (ipTM/pTM = 0.28), consistent with its intrinsically disordered pre-assembly state despite strong propensity. Sericins displayed predominantly disordered conformations that undergo partial ordering upon complexing with FibL-P25-Cu2+, supporting a disorder- to-order templating model for hierarchical fiber assembly.

Immunoinformatic profiling revealed pronounced antigenic heterogeneity. P25 emerged as minimally immunogenic, with sub-threshold antigenicity (VaxiJen: 0.395) and only one MHC-I strong binder restricted to a single HLA allele and minimal MHC-II reactivity. Conversely, FibL and FibH showed the highest potential for inducing CD8+ and CD4+ cell responses among the fibroin subunits, respectively. Ser-4 and Ser-1 exhibited a broad MHC coverage, presenting ≥ 10 strong binders in 81% and 30% of MHC-I alleles, respectively, as well as ≥ 4 high-priority peptides across all 27 and 6 tested MHC-II alleles, respectively. Allergenicity prediction classified FibH, FibL, Ser-2, and Ser-5 as probable allergens. However, all proteins were predicted to be non-toxic, indicating that immunogenic concerns stem from adaptive immune recognition rather than innate cytotoxicity.

These computational predictions challenge the paradigm that sericins alone drive silk immunogenicity, identifying FibL, FibH, Ser-4, and Ser-1 as the primary immunogenic drivers, with HLA-dependent risk profiles. The integration of structural modeling and epitope mapping provides a rational framework for engineering hypoimmunogenic silk variants. However, experimental validation remains essential to translate these predictions into clinically viable, biocompatible silk biomaterials. This computational pipeline is extensible to other protein-based biomaterials for which immunogenicity poses a translational barrier.

Ethical approval

Institutional review board approval was not required.

Funding sources

This work was supported by the General System of Royalties of Colombia [BPIN code: 2020000100190].

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

This manuscript was prepared, including the ethical and transparent use of generative artificial intelligence tools (Perplexity Pro and Claude 4.5 by Anthropic). These tools were employed solely to assist in outlining the manuscript’s structure of the introduction and discussion sections, as well as to improve the manuscript's legibility, without generating, analyzing, or interpreting any scientific data.

All content produced with AI assistance was critically reviewed, edited, and validated by the author(s) to ensure scientific accuracy and originality, and compliance with Elsevier’s policy and the COPE guidelines for responsible use of artificial intelligence in scientific writing.

CRediT authorship contribution statement

Andrés Puerta-González: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Investigation, Software, Formal analysis, Validation, Visualization, Data curation, Writing – original draft, Writing – review & editing. Alejandro Soto-Ospina: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Formal analysis, Software, Formal analysis, Validation, Visualization, Writing – original draft, Writing – review & editing. Yuliet Montoya Osorio: Writing – review & editing, Funding acquisition, Conceptualization, Funding acquisition, Writing – review & editing. John Bustamante-Osorno: Writing – review & editing, Funding acquisition, Conceptualization, Funding acquisition, Writing – review & editing. Lina María Salazar-Peláez: Writing – review & editing, Writing – original draft, Visualization, Supervision, Methodology, Investigation, Conceptualization, Methodology, Investigation, Visualization, Writing – original draft, Writing – review & editing, Supervision.

Declaration of competing interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Bustamante-Osorno John reports financial support was provided by Colombian General System of Royalties. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgement

The authors acknowledge the assistance of the Directorate of Science, Technology, and Innovation at CES University in completing this work, as well as the Ministry of Science, Technology, and Innovation and the General System of Royalties for providing funding for this study.

Footnotes

Appendix A

Supplementary data for this article can be found online. Supplemental Methods 1: Analysis of distant homologs. Supplemental Fig. 1. Distribution of identity percentages among distant homologs of silk proteins. Supplemental Fig. 2. Molecular weight distribution of fibroin subunits and sericin proteins. Supplemental Fig. 3. Isoelectric point (pI) distribution of fibroin subunits and sericin proteins. Supplemental Fig. 4. GRAVY (Grand Average of Hydropathy) index distribution of fibroins and sericins. Supplemental Fig. 5. Prediction of intrinsically disordered regions (IDRs) in the fibroin light chain (FibL) (A-B) and P25 (C-D) using AIUPred. Supplemental Fig. 6. Prediction of fibroin light chain (FibL) interactions using 17 Cu2+ ions. Supplemental Fig. 7. Prediction of P25 interactions with 31 Cu2+ ions. Supplemental Fig. 8. B-cell epitope prediction profile for B. mori silk proteins using BepiPred-3.0. Supplemental Table 1. Predicted MHC class I-restricted cytotoxic T-cell epitopes from Bombyx mori silk proteins (NetMHCpan-4.1, IEDB platform). Supplemental Table 2. Predicted MHC class II-restricted T-lymphocytes epitopes from Bombyx mori silk proteins (NetMHCpan-4.1, IEDB platform). Supplemental Table 3. Details of high-priority predicted MHC-II peptides across silk proteins and HLA alleles. Supplementary data to this article can be found online at https://doi.org/10.1016/j.jgeb.2026.100692.

Contributor Information

Andrés Puerta-González, Email: APUERTAG@ucesedu.onmicrosoft.com.

Alejandro Soto-Ospina, Email: jsoto72@unilasallista.edu.co.

Yuliet Montoya Osorio, Email: yuliet.montoya@upb.edu.co.

John Bustamante-Osorno, Email: john.bustamante@upb.edu.co.

Lina María Salazar-Peláez, Email: lmsalazar@ces.edu.co.

Appendix A. Supplementary data

The following are the Supplementary data to this article:

Supplementary Data
mmc3.docx (26KB, docx)
Supplementary Figure 1

Distribution of identity percentages among distant homologs of silk proteins. Boxplots represent the percentage of identity values (%) obtained for homologous sequences identified through BLAST and PSI-BLAST searches against the NCBI, SwissProt, and PDB databases. Each dot corresponds to an individual homologous sequence retrieved after applying the filtering criteria (greater than 20% identity and alignment coverage of at least 70% of the query length). The orange horizontal lines within the boxes indicate the median, the boxes represent the interquartile range (IQR), and the whiskers extend 1.5 times the IQR from the top and bottom of the box. The analysis included 101 homologs of fibroin light chain (FibL), 12 of fibroin heavy chain (FibH), 168 of fibrohexamerin (P25), and six, six, three, four, and 22 homologs of sericin 1 (Ser-1), sericin 2 (Ser-2), sericin 3 (Ser-3), sericin 4 (Ser-4), and sericin 5 (Ser-5), respectively. Fibroins spanned a wide range of sequence identities, clustering primarily in the 20–60% identity range. Ser-1 to Ser-4 showed sequence identities > 78%, whereas Ser-5 exhibited low identity values (20–32%).

mmc4.pdf (29.1KB, pdf)
Supplementary Figure 2

Molecular weight distribution of fibroin subunits and sericin proteins. Boxplots represent the molecular weight (kDa, log₁₀ scale) of homologous sequences retrieved for the fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericin 1 to 5 (Ser-1 to Ser-5). Values were derived from reference sequences and distant homologs identified using the BLAST and PSI-BLAST searches. Each dot corresponds to a homologous sequence that meets the filtering criteria (greater than 20% identity and greater than 70% coverage). The median is indicated by the orange horizontal line within each box; the boxes represent the interquartile range (IQR), and the whiskers extend 1.5 times the IQR from the top and bottom of the box. FibL showed consistent clustering at 25–30 kDa, whereas FibH exhibited broad variability (350–800 kDa) in its molecular weight. P25 displayed clustered values between 25 and 30 kDa, similar to those of FibL, but with higher dispersion. Ser-1 to Ser-3 showed values ranging from 100 to 140 kDa. Ser-4 was represented by a single sequence (~250 kDa), whereas Ser-5 values were clustered around 75-400 kDa, with an outlier around 1,000 kDa.

mmc5.pdf (32.5KB, pdf)
Supplementary Figure 3

Isoelectric point (pI) distribution of fibroin subunits and sericin proteins. The boxplots illustrate the predicted pI values for homologous sequences of the fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericins 1 to 5 (Ser-1 to Ser-5). Each point represents an individual homolog, while the boxplot displays the median (horizontal line inside the box), interquartile range (the box), and whiskers (extending to 1.5 times the interquartile range). A clear distinction was observed between fibroin subunits and sericins. FibL, FibH, and P25 exhibited broad variability with two clusters clearly differentiated, and with FibH displaying the widest pI distribution, spanning from acidic to basic values (pI 4–10). In contrast, Ser-1 to Ser-4 showed much narrower, more consistent pI ranges, with pI values tightly clustered between 5.5 and 6.5. However, Ser-5 displayed the widest pI dispersion among the sericins, with values spanning from strongly acidic 4.0 to highly basic 10.

mmc6.pdf (25.5KB, pdf)
Supplementary Figure 4

GRAVY (Grand Average of Hydropathy) index distribution of fibroins and sericins. Boxplots represent GRAVY values for homologous sequences of fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericin 1 to 5 (Ser-1 to Ser-5). Data were derived from reference sequences and distant homologs identified through BLAST and PSI-BLAST searches against the NCBI, SwissProt, and PDB databases, applying filtering thresholds (> 20 % identity and > 70 % coverage). Each dot represents an individual homologous sequence. Medians are shown as horizontal lines inside each box, interquartile ranges (IQR) are represented by the boxes, and whiskers extend to 1.5× IQR. Fibroin subunits clustered around neutrality (-0.5 to 0.5). All sericins displayed strongly negative GRAVY values (< -0.5), reflecting their hydrophilic character, with Ser-5 showing the most negative values (-2.4 to -1.3).

mmc7.pdf (26.5KB, pdf)
Supplementary Figure 5

Prediction of intrinsically disordered regions (IDRs) in the fibroin light chain (FibL) (A-B) and P25 (C-D) using AIUPred. For the figures on the left, the solid blue line represents the predicted disorder probability along the amino acid sequence, with values above the 0.5 threshold (blue dashed line) indicating intrinsically disordered regions, whereas the solid orange line indicates potential structural stabilization points via protein-protein interactions. For the figures on the right side, the blue and orange solid lines show disorder predictions under reducing (redox minus) and oxidizing (redox plus) conditions, respectively. FibL and P25 did not exceed these values.

mmc8.pdf (412.3KB, pdf)
Supplementary Figure 6

Prediction of fibroin light chain (FibL) interactions using 17 Cu2+ ions. To explore metal coordination, we modeled the structure of FibL with copper ions using AlphaFold3. We specified a complex comprising FibL and 17 Cu²⁺ ions, consistent with the estimated stoichiometry derived from the histidine content (3.4 Cu²⁺:5His).

mmc9.jpg (150.6KB, jpg)
Supplementary Figure 7

Prediction of P25 interactions with 31 Cu2+ ions. To explore metal coordination, we modeled the P25 tertiary structure with AlphaFold3 and specified a complex comprising P25 and 31 Cu²⁺ ions, consistent with the estimated stoichiometry derived from the histidine content (3.4 Cu²⁺:9His).

mmc10.jpg (228.9KB, jpg)
Supplementary Figure 8

B-cell epitope prediction profile for B. mori silk proteins using BepiPred-3.0. The plot illustrates the epitope probability (Y-axis) for each amino acid residue along its sequence position (X-axis). The solid blue line (epitope probability) represents the antigenicity score calculated by the BepiPred-3.0 server, while the green dashed line (prediction threshold) indicates the recommended classification cutoff of 0.1512. Blue-shaded regions (predicted epitope regions) mark the sequence segments where the score exceeds this threshold, visually identifying the predicted linear epitopes. This threshold corresponds to the value that maximizes the Matthews Correlation Coefficient (MCC) in the tool’s validation dataset, ensuring an optimal balance between sensitivity and specificity in epitope prediction. Fibroin light chain (A), fibroin heavy chain (B), P25 (C), sericin 1 (D), sericin 2 (E), sericin 3 (F), and sericin 4 (G). Silk proteins exhibited distinct antigenicity patterns displaying multiple regions with probability scores that clearly exceeded the 0.1512 threshold, indicating a high density of antigenic “hotspots”, except for sericin 5, which did not exhibit a detectable antigenicity profile above the defined threshold across the analyzed sequence (not shown).

mmc11.pdf (1.4MB, pdf)
Supplementary Methods
mmc12.docx (26KB, docx)
Supplementary Table 1
mmc13.xlsx (41.7MB, xlsx)
Supplementary Table 2
mmc1.xlsx (5.5MB, xlsx)
Supplementary Table 3
mmc2.xlsx (39.4KB, xlsx)

References

  • 1.Inoue S., Tanaka K., Arisaka F., Kimura S., Ohtomo K., Mizuno S. Silk fibroin of Bombyx mori is secreted, assembling a high molecular mass elementary unit consisting of H-chain, L-chain, and P25, with a 6:6:1 molar ratio. J Biol Chem. 2000;275(51):40517–40528. doi: 10.1074/jbc.M006897200. [DOI] [PubMed] [Google Scholar]
  • 2.Altman G.H., Diaz F., Jakuba C., et al. Silk-based biomaterials. Biomaterials. 2003;24(3):401–416. doi: 10.1016/s0142-9612(02)00353-8. [DOI] [PubMed] [Google Scholar]
  • 3.Guo C., Zhang J., Wang X., Nguyen A.T., Liu X.Y., Kaplan D.L. Comparative study of strain‐dependent structural changes of silkworm silks: insight into the structural origin of strain‐stiffening. Small. 2017;13(47) doi: 10.1002/smll.201702266. [DOI] [PubMed] [Google Scholar]
  • 4.Ma L., Dong W., Lai E., Wang J. Silk fibroin-based scaffolds for tissue engineering. Front Bioeng Biotechnol. 2024;12 doi: 10.3389/fbioe.2024.1381838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Arango M.C., Montoya Y., Peresin M.S., Bustamante J., Álvarez-López C. Silk sericin as a biomaterial for tissue engineering: a review. Int J Polym Mater Polym Biomater. 2021;70(16):1115–1129. doi: 10.1080/00914037.2020.1785454. [DOI] [Google Scholar]
  • 6.Liu J., Shi L., Deng Y., et al. Silk sericin-based materials for biomedical applications. Biomaterials. 2022;287 doi: 10.1016/j.biomaterials.2022.121638. [DOI] [PubMed] [Google Scholar]
  • 7.Zhou C.Z., Confalonieri F., Jacquet M., Perasso R., Li Z.G., Janin J. Silk fibroin: structural implications of a remarkable amino acid sequence. Proteins. 2001;44(2):119–122. doi: 10.1002/prot.1078. [DOI] [PubMed] [Google Scholar]
  • 8.Lefèvre T., Rousseau M.E., Pézolet M. Protein secondary structure and orientation in silk as revealed by Raman spectromicroscopy. Biophys J. 2007;92(8):2885–2895. doi: 10.1529/biophysj.106.100339. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Saad M., El-Samad L.M., Gomaa R.A., Augustyniak M., Hassan M.A. A comprehensive review of recent advances in silk sericin: extraction approaches, structure, biochemical characterization, and biomedical applications. Int J Biol Macromol. 2023;250 doi: 10.1016/j.ijbiomac.2023.126067. [DOI] [PubMed] [Google Scholar]
  • 10.Aad R., Dragojlov I., Vesentini S. Sericin protein: structure, properties, and applications. J Funct Biomater. 2024;15(11):322. doi: 10.3390/jfb15110322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Roy A., Kucukural A., Zhang Y. I-TASSER: a unified platform for automated protein structure and function prediction. Nat Protoc. 2010;5(4):725–738. doi: 10.1038/nprot.2010.5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Janson G., Grottesi A., Pietrosanto M., Ausiello G., Guarguaglini G., Paiardini A. Revisiting the “satisfaction of spatial restraints” approach of MODELLER for protein homology modeling. PLoS Comput Biol. 2019;15(12) doi: 10.1371/journal.pcbi.1007219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Yarov-Yarovoy V., Schonbrun J., Baker D. Multipass membrane protein structure prediction using Rosetta. Proteins. 2006;62(4):1010–1025. doi: 10.1002/prot.20817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Baek M., DiMaio F., Anishchenko I., et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373(6557):871–876. doi: 10.1126/science.abj8754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Jumper J., Evans R., Pritzel A., et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Tunyasuvunakool K., Adler J., Wu Z., et al. Highly accurate protein structure prediction for the human proteome. Nature. 2021;596(7873):590–596. doi: 10.1038/s41586-021-03828-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang S., Sun S., Li Z., Zhang R., Xu J. Accurate De Novo prediction of protein contact map by ultra-deep learning model. PLoS Comput Biol. 2017;13(1) doi: 10.1371/journal.pcbi.1005324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Abramson J., Adler J., Dunger J., et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630(8016):493–500. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Varadi M., Bertoni D., Magana P., et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024;52(D1):D368–D375. doi: 10.1093/nar/gkad1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Bulla A.C.S., Sbano da Silva A., Prado Sereno B., Dias M.F.R., Leal da Silva M. Computational methods in immunoinformatics: epitope discovery and diagnostic applications. ACS Omega. 2025;10(39):44816–44839. doi: 10.1021/acsomega.5c05538. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Yan Z., Kim K., Kim H., et al. Next-generation IEDB tools: a platform for epitope prediction and analysis. Nucleic Acids Res. 2024;52(W1):W526–W532. doi: 10.1093/nar/gkae407. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Reynisson B., Alvarez B., Paul S., Peters B., Nielsen M. NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Res. 2020;48(W1):W449–W454. doi: 10.1093/nar/gkaa379. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Shimomura M., Minami H., Suetsugu Y., et al. KAIKObase: an integrated silkworm genome database and data mining tool. BMC Genomics. 2009;10:486. doi: 10.1186/1471-2164-10-486. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Dong Z., Guo K., Zhang X., et al. Identification of Bombyx mori sericin 4 protein as a new biological adhesive. Int J Biol Macromol. 2019;132:1121–1130. doi: 10.1016/j.ijbiomac.2019.03.166. [DOI] [PubMed] [Google Scholar]
  • 25.Guo K., Zhang X., Zhao D., et al. Identification and characterization of sericin5 reveals non-cocoon silk sericin components with high β-sheet content and adhesive strength. Acta Biomater. 2022;150:96–110. doi: 10.1016/j.actbio.2022.07.021. [DOI] [PubMed] [Google Scholar]
  • 26.Kyte J., Doolittle R.F. A simple method for displaying the hydropathic character of a protein. J Mol Biol. 1982;157(1):105–132. doi: 10.1016/0022-2836(82)90515-0. [DOI] [PubMed] [Google Scholar]
  • 27.Korf I, Yandell M, Bedell J. BLAST: an essential guide to the basic local alignment search tool. 1. ed. O’Reilly; 2003.
  • 28.Sayers E.W., Cavanaugh M., Clark K., et al. GenBank 2024 update. Nucleic Acids Res. 2024;52(D1):D134–D137. doi: 10.1093/nar/gkad903. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Haft D.H., Badretdin A., Coulouris G., et al. RefSeq and the prokaryotic genome annotation pipeline in the age of metagenomes. Nucleic Acids Res. 2024;52(D1):D762–D769. doi: 10.1093/nar/gkad988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Burley S.K., Bhatt R., Bhikadiya C., et al. Updated resources for exploring experimentally-determined PDB structures and computed structure models at the RCSB protein data bank. Nucleic Acids Res. 2025;53(D1):D564–D574. doi: 10.1093/nar/gkae1091. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Hamamsy T., Morton J.T., Blackwell R., et al. Protein remote homology detection and structural alignment using deep learning. Nat Biotechnol. 2024;42(6):975–985. doi: 10.1038/s41587-023-01917-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Mistry J., Finn R.D., Eddy S.R., Bateman A., Punta M. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucleic Acids Res. 2013;41(12):e121. doi: 10.1093/nar/gkt263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Xu D., Zhang Y. Improving the physical realism and structural accuracy of protein models by a two-step atomic-level energy minimization. Biophys J. 2011;101(10):2525–2534. doi: 10.1016/j.bpj.2011.10.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang J., Liang Y., Zhang Y. Atomic-level protein structure refinement using fragment-guided molecular dynamics conformation sampling. Structure. 2011;19(12):1784–1795. doi: 10.1016/j.str.2011.09.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Erdős G., Dosztányi Z. AIUPred: combining energy estimation with deep learning for the enhanced prediction of protein disorder. Nucleic Acids Res. 2024;52(W1):W176–W181. doi: 10.1093/nar/gkae385. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Erdős G., Deutsch N., Dosztányi Z. AIUPred - binding: energy embedding to identify disordered binding regions. J Mol Biol. 2025;437(15) doi: 10.1016/j.jmb.2025.169071. [DOI] [PubMed] [Google Scholar]
  • 37.Dunker A.K., Babu M.M., Barbar E., et al. What’s in a name? why these proteins are intrinsically disordered: why these proteins are intrinsically disordered. Intrinsically Disord Proteins. 2013;1(1) doi: 10.4161/idp.24157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Romero P., Obradovic Z., Dunker A.K. Natively disordered proteins: functions and predictions. Appl Bioinformatics. 2004;3(2–3):105–113. doi: 10.2165/00822942-200403020-00005. [DOI] [PubMed] [Google Scholar]
  • 39.Doytchinova I.A., Flower D.R. VaxiJen: a server for prediction of protective antigens, tumour antigens and subunit vaccines. BMC Bioinf. 2007;8:4. doi: 10.1186/1471-2105-8-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Larsen M.V., Lundegaard C., Lamberth K., Buus S., Lund O., Nielsen M. Large-scale validation of methods for cytotoxic T-lymphocyte epitope prediction. BMC Bioinf. 2007;8:424. doi: 10.1186/1471-2105-8-424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Nilsson J.B., Kaabinejadian S., Yari H., et al. Accurate prediction of HLA class II antigen presentation across all loci using tailored data acquisition and refined machine learning. Sci Adv. 2023;9(47) doi: 10.1126/sciadv.adj6367. eadj6367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Clifford J.N., Høie M.H., Deleuran S., Peters B., Nielsen M., Marcatili P. BepiPred-3.0: improved B-cell epitope prediction using protein language models. Protein Sci Publ Protein Soc. 2022;31(12):e4497. doi: 10.1002/pro.4497. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Dimitrov I., Flower D.R., Doytchinova I. AllerTOP–a server for in silico prediction of allergens. BMC Bioinf. 2013;14 doi: 10.1186/1471-2105-14-S6-S4. Suppl 6(Suppl 6):S4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Dimitrov I., Bangov I., Flower D.R., Doytchinova I. AllerTOP vol 2–a server for in silico prediction of allergens. J Mol Model. 2014;20(6):2278. doi: 10.1007/s00894-014-2278-5. [DOI] [PubMed] [Google Scholar]
  • 45.Rathore A.S., Choudhury S., Arora A., Tijare P., Raghava G.P.S. ToxinPred 3.0: an improved method for predicting the toxicity of peptides. Comput Biol Med. 2024;179 doi: 10.1016/j.compbiomed.2024.108926. [DOI] [PubMed] [Google Scholar]
  • 46.Mariani V., Biasini M., Barbato A., Schwede T. lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics. 2013;29(21):2722–2728. doi: 10.1093/bioinformatics/btt473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Zhang Y., Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins. 2004;57(4):702–710. doi: 10.1002/prot.20264. [DOI] [PubMed] [Google Scholar]
  • 48.Xu J., Zhang Y. How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics. 2010;26(7):889–895. doi: 10.1093/bioinformatics/btq066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zhou L., Chen X., Shao Z., Huang Y., Knight D.P. Effect of metallic ions on silk formation in the Mulberry silkworm. Bombyx Mori J Phys Chem b. 2005;109(35):16937–16945. doi: 10.1021/jp050883m. [DOI] [PubMed] [Google Scholar]
  • 50.Numata K., Cebe P., Kaplan D.L. Mechanism of enzymatic degradation of beta-sheet crystals. Biomaterials. 2010;31(10):2926–2933. doi: 10.1016/j.biomaterials.2009.12.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Cheng Y., Koh L.D., Li D., Ji B., Han M.Y., Zhang Y.W. On the strength of β-sheet crystallites of Bombyx mori silk fibroin. J R Soc Interface. 2014;11(96) doi: 10.1098/rsif.2014.0305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Marsh R.E., Corey R.B., Pauling L. An investigation of the structure of silk fibroin. Biochim Biophys Acta. 1955;16:1–34. doi: 10.1016/0006-3002(55)90178-5. [DOI] [PubMed] [Google Scholar]
  • 53.Takahashi Y., Gehoh M., Yuzuriha K. Structure refinement and diffuse streak scattering of silk (Bombyx mori) Int J Biol Macromol. 1999;24(2–3):127–138. doi: 10.1016/s0141-8130(98)00080-4. [DOI] [PubMed] [Google Scholar]
  • 54.Moreno-Tortolero R.O., Luo Y., Parmeggiani F., et al. Molecular organization of fibroin heavy chain and mechanism of fibre formation in Bombyx mori. Commun Biol. 2024;7(1):786. doi: 10.1038/s42003-024-06474-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Yamaguchi K., Kikuchi Y., Takagi T., et al. Primary structure of the silk fibroin light chain determined by cDNA sequencing and peptide analysis. J Mol Biol. 1989;210(1):127–139. doi: 10.1016/0022-2836(89)90295-7. [DOI] [PubMed] [Google Scholar]
  • 56.Zafar M.S., Belton D.J., Hanby B., Kaplan D.L., Perry C.C. Functional material features of Bombyx mori silk light versus heavy chain proteins. Biomacromolecules. 2015;16(2):606–614. doi: 10.1021/bm501667j. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zabelina V., Takasu Y., Sehadova H., et al. Mutation in Bombyx mori fibrohexamerin (P25) gene causes reorganization of rough endoplasmic reticulum in posterior silk gland cells and alters morphology of fibroin secretory globules in the silk gland lumen. Insect Biochem Mol Biol. 2021;135 doi: 10.1016/j.ibmb.2021.103607. [DOI] [PubMed] [Google Scholar]
  • 58.Sparkes J., Holland C. The rheological properties of native sericin. Acta Biomater. 2018;69:234–242. doi: 10.1016/j.actbio.2018.01.021. [DOI] [PubMed] [Google Scholar]
  • 59.Välisalmi T., Linder M.B. The ratio of fibroin to sericin in the middle silk gland of Bombyx mori and its correlation with the extensional behavior of the silk dope. Protein Sci Publ Protein Soc. 2024;33(3):e4907. doi: 10.1002/pro.4907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Asakura T., Ohgo K., Ishida T., Taddei P., Monti P., Kishore R. Possible implications of serine and tyrosine residues and intermolecular interactions on the appearance of silk I structure of Bombyx mori silk fibroin-derived synthetic peptides: high-resolution 13C cross-polarization/magic-angle spinning NMR study. Biomacromolecules. 2005;6(1):468–474. doi: 10.1021/bm049487k. [DOI] [PubMed] [Google Scholar]
  • 61.Teramoto H., Kakazu A., Yamauchi K., Asakura T. Role of hydroxyl side chains in Bombyx mori silk sericin in stabilizing its solid structure. Macromolecules. 2007;40(5):1562–1569. doi: 10.1021/ma062604e. [DOI] [Google Scholar]
  • 62.Zhou C.Z., Confalonieri F., Medina N., et al. Fine organization of Bombyx mori fibroin heavy chain gene. Nucleic Acids Res. 2000;28(12):2413–2419. doi: 10.1093/nar/28.12.2413. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Lu W., Lan X., Zhang T., Sun H., Ma S., Xia Q. Precise characterization of Bombyx mori fibroin heavy chain gene using Cpf1-based enrichment and oxford nanopore technologies. Insects. 2021;12(9):832. doi: 10.3390/insects12090832. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Domigan L.J., Andersson M., Alberti K.A., et al. Carbonic anhydrase generates a pH gradient in Bombyx mori silk glands. Insect Biochem Mol Biol. 2015;65:100–106. doi: 10.1016/j.ibmb.2015.09.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Bini E., Knight D.P., Kaplan D.L. Mapping domain structures in silks from insects and spiders related to protein assembly. J Mol Biol. 2004;335(1):27–40. doi: 10.1016/j.jmb.2003.10.043. [DOI] [PubMed] [Google Scholar]
  • 66.Inoue S., Tanaka K., Tanaka H., et al. Assembly of the silk fibroin elementary unit in endoplasmic reticulum and a role of L-chain for protection of alpha1,2-mannose residues in N-linked oligosaccharide chains of fibrohexamerin/P25. Eur J Biochem. 2004;271(2):356–366. doi: 10.1046/j.1432-1033.2003.03934.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Asakura T., Okushita K., Williamson M.P. Analysis of the structure of Bombyx mori silk fibroin by NMR. Macromolecules. 2015;48(8):2345–2357. doi: 10.1021/acs.macromol.5b00160. [DOI] [Google Scholar]
  • 68.Asakura T., Williamson M.P. A review on the structure of Bombyx mori silk fibroin fiber studied using solid-state NMR: an antipolar lamella with an 8-residue repeat. Int J Biol Macromol. 2023;245 doi: 10.1016/j.ijbiomac.2023.125537. [DOI] [PubMed] [Google Scholar]
  • 69.He Y.X., Zhang N.N., Li W.F., et al. N-terminal domain of bombyx mori fibroin mediates the assembly of silk in response to pH decrease. J Mol Biol. 2012;418(3–4):197–207. doi: 10.1016/j.jmb.2012.02.040. [DOI] [PubMed] [Google Scholar]
  • 70.Hu X., Kaplan D., Cebe P. Determining beta-sheet crystallinity in fibrous proteins by thermal analysis and infrared spectroscopy. Macromolecules. 2006;39(18):6161–6170. doi: 10.1021/ma0610109. [DOI] [Google Scholar]
  • 71.Liu M., Millard P., Urch H., et al. Microencapsulation of high‐content actives using biodegradable silk materials. Small. 2022;18(31) doi: 10.1002/smll.202201487. [DOI] [PubMed] [Google Scholar]
  • 72.Wang X., Tan X., Liu Q., et al. Fiber formation and mechanical properties of Bombyx mori silk are regulated by vacuolar-type ATPase. ACS Biomater Sci Eng. 2021;7(12):5532–5540. doi: 10.1021/acsbiomaterials.1c01230. [DOI] [PubMed] [Google Scholar]
  • 73.Kunz R.I., Brancalhão R.M.C., de Ribeiro L., FC, Natali M.R.M. Silkworm sericin: properties and biomedical applications. Biomed Res Int. 2016;2016 doi: 10.1155/2016/8175701. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Zhou L., Chen X., Shao Z., Zhou P., Knight D.P., Vollrath F. Copper in the silk formation process of Bombyx mori silkworm. FEBS Lett. 2003;554(3):337–341. doi: 10.1016/S0014-5793(03)01184-0. [DOI] [PubMed] [Google Scholar]
  • 75.Zong X.H., Zhou P., Shao Z.Z., et al. Effect of pH and copper(II) on the conformation transitions of silk fibroin based on EPR, NMR, and Raman spectroscopy. Biochemistry. 2004;43(38):11932–11941. doi: 10.1021/bi049455h. [DOI] [PubMed] [Google Scholar]
  • 76.Brookstein O., Shimoni E., Eliaz D., Kaplan-Ashiri I., Carmel I., Shimanovich U. Metal ions guide the production of silkworm silk fibers. Nat Commun. 2024;15(1):6671. doi: 10.1038/s41467-024-50879-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Breydo L., Uversky V.N. Role of metal ions in aggregation of intrinsically disordered proteins in neurodegenerative diseases. Met Integr Biometal Sci. 2011;3(11):1163–1180. doi: 10.1039/c1mt00106j. [DOI] [PubMed] [Google Scholar]
  • 78.Saibil H. Chaperone machines for protein folding, unfolding and disaggregation. Nat Rev Mol Cell Biol. 2013;14(10):630–642. doi: 10.1038/nrm3658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Tompa P., Fuxreiter M. Fuzzy complexes: polymorphism and structural disorder in protein–protein interactions. Trends Biochem Sci. 2008;33(1):2–8. doi: 10.1016/j.tibs.2007.10.003. [DOI] [PubMed] [Google Scholar]
  • 80.Petrone L., Kumar A., Sutanto C.N., et al. Mussel adhesion is dictated by time-regulated secretion and molecular conformation of mussel adhesive proteins. Nat Commun. 2015;6:8737. doi: 10.1038/ncomms9737. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Waite J.H. Mussel adhesion - essential footwork. J Exp Biol. 2017;220(Pt 4):517–530. doi: 10.1242/jeb.134056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Fasting C., Schalley C.A., Weber M., et al. Multivalency as a chemical organization and action principle. Angew Chem Int Ed. 2012;51(42):10472–10498. doi: 10.1002/anie.201201114. [DOI] [PubMed] [Google Scholar]
  • 83.Konermann L., Pan J., Liu Y.H. Hydrogen exchange mass spectrometry for studying protein structure and dynamics. Chem Soc Rev. 2011;40(3):1224–1234. doi: 10.1039/c0cs00113a. [DOI] [PubMed] [Google Scholar]
  • 84.Schuler B., Hofmann H. Single-molecule spectroscopy of protein folding dynamics—expanding scope and timescales. Curr Opin Struct Biol. 2013;23(1):36–47. doi: 10.1016/j.sbi.2012.10.008. [DOI] [PubMed] [Google Scholar]
  • 85.Borgia A., Borgia M.B., Bugge K., et al. Extreme disorder in an ultrahigh-affinity protein complex. Nature. 2018;555(7694):61–66. doi: 10.1038/nature25762. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Rosenzweig A.C. Copper delivery by metallochaperone proteins. Acc Chem Res. 2001;34(2):119–128. doi: 10.1021/ar000012p. [DOI] [PubMed] [Google Scholar]
  • 87.Palm-Espling M.E., Niemiec M.S., Wittung-Stafshede P. Role of metal in folding and stability of copper proteins in vitro. Biochim Biophys Acta BBA - Mol Cell Res. 2012;1823(9):1594–1603. doi: 10.1016/j.bbamcr.2012.01.013. [DOI] [PubMed] [Google Scholar]
  • 88.Kurosaki S., Otsuka H., Kunitomo M., Koyama M., Pawankar R., Matumoto K. Fibroin allergy. IgE mediated hypersensitivity to silk suture materials. Nihon Ika Daigaku Zasshi. 1999;66(1):41–44. doi: 10.1272/jnms.66.41. [DOI] [PubMed] [Google Scholar]
  • 89.Cook K.A., Kelso J.M. Surgery-related contact dermatitis: a review of potential irritants and allergens. J Allergy Clin Immunol Pract. 2017;5(5):1234–1240. doi: 10.1016/j.jaip.2017.03.001. [DOI] [PubMed] [Google Scholar]
  • 90.Ceremsak J.J., Kahwash B.M., Budnick H., Kent D.T. Silk suture hypersensitivity masquerading as recurrent hypoglossal nerve stimulator infections. OTO Open. 2025;9(1) doi: 10.1002/oto2.70058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Zhang Y., Shi N., He L., et al. Silk sericin activates mild immune response and increases antibody production. J Biomed Nanotechnol. 2021;17(12):2433–2443. doi: 10.1166/jbn.2021.3206. [DOI] [PubMed] [Google Scholar]
  • 92.Wang Y., Liang Y., Huang J., et al. Proteomic analysis of silk fibroin reveals diverse biological function of different degumming processing from different origin. Front Bioeng Biotechnol. 2022;9 doi: 10.3389/fbioe.2021.777320. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Jiao Z., Song Y., Jin Y., et al. In vivo characterizations of the immune properties of sericin: an ancient material with emerging value in biomedical applications. Macromol Biosci. 2017;17(12) doi: 10.1002/mabi.201700229. [DOI] [PubMed] [Google Scholar]
  • 94.Ode Boni B.O., Bakadia B.M., Osi A.R., et al. Immune response to silk sericin–fibroin composites: potential immunogenic elements and alternatives for immunomodulation. Macromol Biosci. 2022;22(1) doi: 10.1002/mabi.202100292. [DOI] [PubMed] [Google Scholar]
  • 95.Tian Z., Chen H., Zhao P. Compliant immune response of silk-based biomaterials broadens application in wound treatment. Front Pharmacol. 2025;16 doi: 10.3389/fphar.2025.1548837. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Ilinskaya A.N., Dobrovolskaia M.A. Understanding the immunogenicity and antigenicity of nanomaterials: past, present and future. Toxicol Appl Pharmacol. 2016;299:70–77. doi: 10.1016/j.taap.2016.01.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Ramaraj T., Angel T., Dratz E.A., Jesaitis A.J., Mumey B. Antigen-antibody interface properties: composition, residue interactions, and features of 53 non-redundant structures. Biochim Biophys Acta. 2012;1824(3):520–532. doi: 10.1016/j.bbapap.2011.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Kringelum J.V., Nielsen M., Padkjær S.B., Lund O. Structural analysis of B-cell epitopes in antibody:protein complexes. Mol Immunol. 2013;53(1–2):24–34. doi: 10.1016/j.molimm.2012.06.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Zheng W, Ruan J, Hu G, Wang K, Hanlon M, Gao J. Analysis of conformational B-cell epitopes in the antibody-antigen complex using the depth function and the convex hull. Muyldermans S, ed. PLOS ONE. 2015;10(8):e0134835. doi:10.1371/journal.pone.0134835. [DOI] [PMC free article] [PubMed]
  • 100.Schramm G., Petersen A., Bufe A., Schlaak M., Becker W.M. Identification and Characterization of the Major Allergens of Velvet Grass (Holcus lanatus), Hol I 1 and Hol I 5. Int Arch Allergy Immunol. 1996;110(4):354–363. doi: 10.1159/000237328. [DOI] [PubMed] [Google Scholar]
  • 101.Schramm G., Bufe A., Petersen A., Schlaak M., Becker W.M. Molecular and immunological characterization of group V allergen isoforms from velvet grass pollen (Holcus lanatus) Eur J Biochem. 1998;252(2):200–206. doi: 10.1046/j.1432-1327.1998.2520200.x. [DOI] [PubMed] [Google Scholar]
  • 102.Jensen P.E. Reduction of disulfide bonds during antigen processing: evidence from a thiol-dependent insulin determinant. J Exp Med. 1991;174(5):1121–1130. doi: 10.1084/jem.174.5.1121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Watts C. Capture and processing of exogenous antigens for presentation on MHC molecules. Annu Rev Immunol. 1997;15:821–850. doi: 10.1146/annurev.immunol.15.1.821. [DOI] [PubMed] [Google Scholar]
  • 104.Hastings K.T., Cresswell P. Disulfide reduction in the endocytic pathway: immunological functions of gamma-interferon-inducible lysosomal thiol reductase. Antioxid Redox Signal. 2011;15(3):657–668. doi: 10.1089/ars.2010.3684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Inoue S., Tanaka K., Arisaka F., Kimura S., Ohtomo K., Mizuno S. Silk fibroin of Bombyx mori is secreted, assembling a high m olecular mass elementary unit consisting of H-chain, L-chain, and P25 , with a 6: 6: 1 molar ratio. J Biol Chem. 2000;275(51):40517–40528. doi: 10.1074/jbc.M006897200. [DOI] [PubMed] [Google Scholar]
  • 106.Skehel J.J., Wiley D.C. Receptor binding and membrane fusion in virus entry: the influenza hemagglutinin. Annu Rev Biochem. 2000;69:531–569. doi: 10.1146/annurev.biochem.69.1.531. [DOI] [PubMed] [Google Scholar]
  • 107.Crispin M., Ward A.B., Wilson I.A. Structure and immune recognition of the HIV glycan shield. Annu Rev Biophys. 2018;47:499–523. doi: 10.1146/annurev-biophys-060414-034156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.So T., Ito H.O., Koga T., Watanabe S., Ueda T., Imoto T. Depression of T-cell epitope generation by stabilizing hen lysozyme. J Biol Chem. 1997;272(51):32136–32140. doi: 10.1074/jbc.272.51.32136. [DOI] [PubMed] [Google Scholar]
  • 109.Uversky V.N. Intrinsic disorder-based protein interactions and their modulators. Curr Pharm Des. 2013;19(23):4191–4213. doi: 10.2174/1381612811319230005. [DOI] [PubMed] [Google Scholar]
  • 110.Roche P.A., Furuta K. The ins and outs of MHC class II-mediated antigen processing and presentation. Nat Rev Immunol. 2015;15(4):203–216. doi: 10.1038/nri3818. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Bachmann M.F., Zinkernagel R.M. Neutralizing antiviral B cell responses. Annu Rev Immunol. 1997;15:235–270. doi: 10.1146/annurev.immunol.15.1.235. [DOI] [PubMed] [Google Scholar]
  • 112.Meinel L., Hofmann S., Karageorgiou V., et al. The inflammatory responses to silk films in vitro and in vivo. Biomaterials. 2005;26(2):147–155. doi: 10.1016/j.biomaterials.2004.02.047. [DOI] [PubMed] [Google Scholar]
  • 113.Rock K.L., Goldberg A.L. Degradation of cell proteins and the generation of MHC class I-presented peptides. Annu Rev Immunol. 1999;17:739–779. doi: 10.1146/annurev.immunol.17.1.739. [DOI] [PubMed] [Google Scholar]
  • 114.Kloetzel P.M. Antigen processing by the proteasome. Nat Rev Mol Cell Biol. 2001;2(3):179–187. doi: 10.1038/35056572. [DOI] [PubMed] [Google Scholar]
  • 115.Fontana A., de Laureto P.P., Spolaore B., Frare E., Picotti P., Zambonin M. Probing protein structure by limited proteolysis. Acta Biochim Pol. 2004;51(2):299–321. [PubMed] [Google Scholar]
  • 116.Weber C.A., Mehta P.J., Ardito M., Moise L., Martin B., De Groot A.S. T cell epitope: friend or foe? Immunogenicity of biologics in context. Adv Drug Deliv Rev. 2009;61(11):965–976. doi: 10.1016/j.addr.2009.07.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Jawa V., Cousens L.P., Awwad M., Wakshull E., Kropshofer H., De Groot A.S. T-cell dependent immunogenicity of protein therapeutics: preclinical assessment and mitigation. Clin Immunol. 2013;149(3):534–555. doi: 10.1016/j.clim.2013.09.006. [DOI] [PubMed] [Google Scholar]
  • 118.Jawa V., Terry F., Gokemeijer J., et al. T-cell dependent immunogenicity of protein therapeutics pre-clinical assessment and mitigation-updated consensus and review 2020. Front Immunol. 2020;11:1301. doi: 10.3389/fimmu.2020.01301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Shi D., Beasock D., Fessler A., et al. To PEGylate or not to PEGylate: Immunological properties of nanomedicine’s most popular component, polyethylene glycol and its alternatives. Adv Drug Deliv Rev. 2022;180 doi: 10.1016/j.addr.2021.114079. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Scotland B.L., Shaw J.R., Dharmaraj S., et al. Cell and biomaterial delivery strategies to induce immune tolerance. Adv Drug Deliv Rev. 2023;203 doi: 10.1016/j.addr.2023.115141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Wee J., Wei G.W. Evaluation of AlphaFold 3’s protein–protein complexes for predicting binding free energy changes upon mutation. J Chem Inf Model. 2024;64(16):6676–6683. doi: 10.1021/acs.jcim.4c00976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Lin P.Y., Huang S.C., Chen K.L., et al. Analysing protein complexes in plant science: insights and limitation with AlphaFold 3. Bot Stud. 2025;66(1):14. doi: 10.1186/s40529-025-00462-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Malhotra Y., John J., Yadav D., et al. Advancements in protein structure prediction: a comparative overview of AlphaFold and its derivatives. Comput Biol Med. 2025;188 doi: 10.1016/j.compbiomed.2025.109842. [DOI] [PubMed] [Google Scholar]
  • 124.Pesce F., Lindorff-Larsen K. Refining conformational ensembles of flexible proteins against small-angle x-ray scattering data. Biophys J. 2021;120(22):5124–5135. doi: 10.1016/j.bpj.2021.10.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Cordeiro TN, Herranz-Trillo F, Urbanek A, et al. Structural characterization of highly flexible proteins by small-angle scattering. In: Chaudhuri B, Muñoz IG, Qian S, Urban VS, eds. Biological small angle scattering: techniques, strategies and tips. Vol 1009. Advances in experimental medicine and biology. Springer Singapore; 2017:107-129. doi:10.1007/978-981-10-6038-0_7. [DOI] [PubMed]
  • 126.Maiti S., Singh A., Maji T., Saibo N.V., De S. Experimental methods to study the structure and dynamics of intrinsically disordered regions in proteins. Curr Res Struct Biol. 2024;7 doi: 10.1016/j.crstbi.2024.100138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Brotzakis Z.F., Zhang S., Murtada M.H., Vendruscolo M. AlphaFold prediction of structural ensembles of disordered proteins. Nat Commun. 2025;16(1):1632. doi: 10.1038/s41467-025-56572-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Kashyap R., Walsh N., Deveryshetty J., et al. Cryo-EM captures the coordination of asymmetric electron transfer through a di-copper site in DPOR. Nat Commun. 2025;16(1):3866. doi: 10.1038/s41467-025-59158-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Mertens H.D.T., Svergun D.I. Combining NMR and small angle X-ray scattering for the study of biomolecular structure and dynamics. Arch Biochem Biophys. 2017;628:33–41. doi: 10.1016/j.abb.2017.05.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Horitani M., Kusubayashi K., Oshima K., Yato A., Sugimoto H., Watanabe K. X-ray crystallography and electron paramagnetic resonance spectroscopy reveal active site rearrangement of cold-adapted inorganic pyrophosphatase. Sci Rep. 2020;10(1):4368. doi: 10.1038/s41598-020-61217-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Bulla A.C.S., Sbano da Silva A., Prado Sereno B., Dias M.F.R., Leal da Silva M. Computational me thods in immunoinformatics : epitope dis covery and diagnostic applications. ACS Omega. 2025;10(39):44816–44839. doi: 10.1021/acsomega.5c05538. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Mukherjee R., Verma N., Bhagat A., Verma C.K. Immunoinformatics: expanding frontiers and emerging tools in bioinformatics for immunology. Indian J Physiol Pharmacol. 2025:1–8. doi: 10.25259/IJPP_379_2024. [DOI] [Google Scholar]
  • 133.Villanueva-Flores F., Sanchez-Villamil J.I., Garcia-Atutxa I. AI-driven epitope prediction: a systematic review, comparative analysis, and practical guide for vaccine development. npj Vaccines. 2025;10(1):207. doi: 10.1038/s41541-025-01258-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Majumder N., Bhattacharjee M., Spagnoli G.C., Ghosh S. Immune response profiles induced by silk-based biomaterials: a journey from ‘immunogenicity’ towards ‘immuno-compatibility. J Mater Chem B. 2024;12(38):9508–9523. doi: 10.1039/D4TB01231C. [DOI] [PubMed] [Google Scholar]
  • 135.Bruno P.M. The genomics revolution comes to the immunopeptidome. Genes Immun. 2023;25(3):256–258. doi: 10.1038/s41435-023-00244-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Zhou X., Renauer P.A., Zhou L., Fang S., Chen S. Applications of CRISPR technology in cellular immunotherapy. Immunol Rev. 2023;320(1):199–216. doi: 10.1111/imr.13241. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Data
mmc3.docx (26KB, docx)
Supplementary Figure 1

Distribution of identity percentages among distant homologs of silk proteins. Boxplots represent the percentage of identity values (%) obtained for homologous sequences identified through BLAST and PSI-BLAST searches against the NCBI, SwissProt, and PDB databases. Each dot corresponds to an individual homologous sequence retrieved after applying the filtering criteria (greater than 20% identity and alignment coverage of at least 70% of the query length). The orange horizontal lines within the boxes indicate the median, the boxes represent the interquartile range (IQR), and the whiskers extend 1.5 times the IQR from the top and bottom of the box. The analysis included 101 homologs of fibroin light chain (FibL), 12 of fibroin heavy chain (FibH), 168 of fibrohexamerin (P25), and six, six, three, four, and 22 homologs of sericin 1 (Ser-1), sericin 2 (Ser-2), sericin 3 (Ser-3), sericin 4 (Ser-4), and sericin 5 (Ser-5), respectively. Fibroins spanned a wide range of sequence identities, clustering primarily in the 20–60% identity range. Ser-1 to Ser-4 showed sequence identities > 78%, whereas Ser-5 exhibited low identity values (20–32%).

mmc4.pdf (29.1KB, pdf)
Supplementary Figure 2

Molecular weight distribution of fibroin subunits and sericin proteins. Boxplots represent the molecular weight (kDa, log₁₀ scale) of homologous sequences retrieved for the fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericin 1 to 5 (Ser-1 to Ser-5). Values were derived from reference sequences and distant homologs identified using the BLAST and PSI-BLAST searches. Each dot corresponds to a homologous sequence that meets the filtering criteria (greater than 20% identity and greater than 70% coverage). The median is indicated by the orange horizontal line within each box; the boxes represent the interquartile range (IQR), and the whiskers extend 1.5 times the IQR from the top and bottom of the box. FibL showed consistent clustering at 25–30 kDa, whereas FibH exhibited broad variability (350–800 kDa) in its molecular weight. P25 displayed clustered values between 25 and 30 kDa, similar to those of FibL, but with higher dispersion. Ser-1 to Ser-3 showed values ranging from 100 to 140 kDa. Ser-4 was represented by a single sequence (~250 kDa), whereas Ser-5 values were clustered around 75-400 kDa, with an outlier around 1,000 kDa.

mmc5.pdf (32.5KB, pdf)
Supplementary Figure 3

Isoelectric point (pI) distribution of fibroin subunits and sericin proteins. The boxplots illustrate the predicted pI values for homologous sequences of the fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericins 1 to 5 (Ser-1 to Ser-5). Each point represents an individual homolog, while the boxplot displays the median (horizontal line inside the box), interquartile range (the box), and whiskers (extending to 1.5 times the interquartile range). A clear distinction was observed between fibroin subunits and sericins. FibL, FibH, and P25 exhibited broad variability with two clusters clearly differentiated, and with FibH displaying the widest pI distribution, spanning from acidic to basic values (pI 4–10). In contrast, Ser-1 to Ser-4 showed much narrower, more consistent pI ranges, with pI values tightly clustered between 5.5 and 6.5. However, Ser-5 displayed the widest pI dispersion among the sericins, with values spanning from strongly acidic 4.0 to highly basic 10.

mmc6.pdf (25.5KB, pdf)
Supplementary Figure 4

GRAVY (Grand Average of Hydropathy) index distribution of fibroins and sericins. Boxplots represent GRAVY values for homologous sequences of fibroin light chain (FibL), fibroin heavy chain (FibH), fibrohexamerin (P25), and sericin 1 to 5 (Ser-1 to Ser-5). Data were derived from reference sequences and distant homologs identified through BLAST and PSI-BLAST searches against the NCBI, SwissProt, and PDB databases, applying filtering thresholds (> 20 % identity and > 70 % coverage). Each dot represents an individual homologous sequence. Medians are shown as horizontal lines inside each box, interquartile ranges (IQR) are represented by the boxes, and whiskers extend to 1.5× IQR. Fibroin subunits clustered around neutrality (-0.5 to 0.5). All sericins displayed strongly negative GRAVY values (< -0.5), reflecting their hydrophilic character, with Ser-5 showing the most negative values (-2.4 to -1.3).

mmc7.pdf (26.5KB, pdf)
Supplementary Figure 5

Prediction of intrinsically disordered regions (IDRs) in the fibroin light chain (FibL) (A-B) and P25 (C-D) using AIUPred. For the figures on the left, the solid blue line represents the predicted disorder probability along the amino acid sequence, with values above the 0.5 threshold (blue dashed line) indicating intrinsically disordered regions, whereas the solid orange line indicates potential structural stabilization points via protein-protein interactions. For the figures on the right side, the blue and orange solid lines show disorder predictions under reducing (redox minus) and oxidizing (redox plus) conditions, respectively. FibL and P25 did not exceed these values.

mmc8.pdf (412.3KB, pdf)
Supplementary Figure 6

Prediction of fibroin light chain (FibL) interactions using 17 Cu2+ ions. To explore metal coordination, we modeled the structure of FibL with copper ions using AlphaFold3. We specified a complex comprising FibL and 17 Cu²⁺ ions, consistent with the estimated stoichiometry derived from the histidine content (3.4 Cu²⁺:5His).

mmc9.jpg (150.6KB, jpg)
Supplementary Figure 7

Prediction of P25 interactions with 31 Cu2+ ions. To explore metal coordination, we modeled the P25 tertiary structure with AlphaFold3 and specified a complex comprising P25 and 31 Cu²⁺ ions, consistent with the estimated stoichiometry derived from the histidine content (3.4 Cu²⁺:9His).

mmc10.jpg (228.9KB, jpg)
Supplementary Figure 8

B-cell epitope prediction profile for B. mori silk proteins using BepiPred-3.0. The plot illustrates the epitope probability (Y-axis) for each amino acid residue along its sequence position (X-axis). The solid blue line (epitope probability) represents the antigenicity score calculated by the BepiPred-3.0 server, while the green dashed line (prediction threshold) indicates the recommended classification cutoff of 0.1512. Blue-shaded regions (predicted epitope regions) mark the sequence segments where the score exceeds this threshold, visually identifying the predicted linear epitopes. This threshold corresponds to the value that maximizes the Matthews Correlation Coefficient (MCC) in the tool’s validation dataset, ensuring an optimal balance between sensitivity and specificity in epitope prediction. Fibroin light chain (A), fibroin heavy chain (B), P25 (C), sericin 1 (D), sericin 2 (E), sericin 3 (F), and sericin 4 (G). Silk proteins exhibited distinct antigenicity patterns displaying multiple regions with probability scores that clearly exceeded the 0.1512 threshold, indicating a high density of antigenic “hotspots”, except for sericin 5, which did not exhibit a detectable antigenicity profile above the defined threshold across the analyzed sequence (not shown).

mmc11.pdf (1.4MB, pdf)
Supplementary Methods
mmc12.docx (26KB, docx)
Supplementary Table 1
mmc13.xlsx (41.7MB, xlsx)
Supplementary Table 2
mmc1.xlsx (5.5MB, xlsx)
Supplementary Table 3
mmc2.xlsx (39.4KB, xlsx)

Articles from Journal of Genetic Engineering & Biotechnology are provided here courtesy of Academy of Scientific Research and Technology, Egypt

RESOURCES