Skip to main content
Frontiers in Microbiology logoLink to Frontiers in Microbiology
. 2026 Aug 3;17:1839420. doi: 10.3389/fmicb.2026.1839420

TPT: a compact CNN-transformer encoder for efficient microbial small protein modeling

Fang Sheng 1,2,†, Junhe Zhang 3,†, Chengkai Zhu 1,*
PMCID: PMC13478124  PMID: 42609539

Abstract

Introduction

Microbial small proteins, encoded by small open reading frames (smORFs), play essential roles in antimicrobial activity, metabolic regulation, and signaling pathways. However, their short length and rapid evolutionary rate present significant challenges for computational modeling.

Methods

We introduce TinyProteinTransformer (TPT), a lightweight CNN-Transformer hybrid encoder pretrained on the Global Microbial smORF Catalog (GMSC; >280 million smORF families). TPT integrates multi-scale convolutional filters to capture local sequence motifs with Transformer layers for broader contextual modeling, and is trained jointly with masked language modeling and contrastive learning to capture both residue-level and sequence-level representations. We evaluated TPT on six downstream peptide/protein classification benchmarks spanning antimicrobial peptides (AMP), toxic peptides (TOX), bacteriocins (BCN), anti-CRISPR proteins (Acr), quorum-sensing peptides (QSP), and cell-penetrating peptides (CPP) using frozen-encoder linear probing.

Results

On the AMP and TOX tasks, TPT (103 M parameters) matched ESM2-150 M in predictive performance (AUC 0.930 vs. 0.928 for AMP; 0.930 vs. 0.925 for TOX) while achieving 4.3-fold faster inference. Across the evaluated benchmarks, TPT showed competitive performance compared with larger protein language models. Ablation experiments identified the contrastive objective as an important contributor to representation quality: its removal reduced AUC across all six downstream tasks and produced performance patterns consistent with MLM-only pretraining.

Conclusion

These results suggest that, within the evaluated benchmarks, effective smORF representation may benefit from pretraining objectives and inductive biases tailored to short, rapidly evolving sequences, rather than from model scale alone. TPT therefore provides a compact and computationally efficient encoder with competitive performance for microbial peptide analysis.

Keywords: deep learning, microproteins, peptide prediction, protein language model, smORFs

1. Introduction

Small proteins, also referred to as peptides encoded by small open reading frames (smORFs), are generally considered to consist of 100 amino acids or fewer. These peptides play crucial biological roles in antimicrobial activity, metabolic regulation, intercellular signaling, and adaptation to environmental stressors (Duval and Cossart, 2017; Sberro et al., 2019; Storz et al., 2014). In eukaryotes, smORF-encoded small proteins have attracted growing attention for their roles in diverse biological processes and disease mechanisms, including embryogenesis (Pauli et al., 2014; Savard et al., 2006; Hao et al., 2025), stress responses (Andreev et al., 2015; Khitun et al., 2019), muscle development (Anderson et al., 2015; Bi et al., 2017; Matsumoto et al., 2017), immune and inflammatory responses (Bhatta et al., 2020; Jackson et al., 2018; Nichols et al., 2024), and tumorigenesis (Huang et al., 2021; Lyu et al., 2023; Zhang et al., 2025). Recent advances in high-throughput genome sequencing and ribosome profiling have demonstrated that smORFs are far more abundant and functionally diverse than previously recognized (Ingolia et al., 2009; Miravet-Verde et al., 2019; Weaver et al., 2019). However, their short length, rapid evolutionary turnover, and sparse domain organization pose unique challenges for computational characterization (Couso and Patraquim, 2017).

Earlier peptide prediction methods, such as AMPlify (Li et al., 2022) and AMPScannerV2 (Veltri et al., 2018), mainly relied on task-specific supervised learning frameworks for antimicrobial peptide prediction. Although effective within individual tasks, these non-pretrained models are typically trained on limited labeled datasets and may have reduced generalizability across diverse peptide functions. Protein language models (PLMs) have proven highly effective in learning representations from large unlabeled protein sequence corpora. In particular, models like ESM (Lin et al., 2023), ProtTrans (ProtBERT, ProtT5) (Elnaggar et al., 2021), and MSA Transformer (Rao et al., 2021) have demonstrated strong performance in structure prediction, mutation effect prediction, and functional annotation. However, these models are predominantly trained on databases such as UniRef50 and UniRef90 (Suzek et al., 2007), which largely consist of long, evolutionarily conserved proteins. Consequently, their architectural and pretraining assumptions are often oriented toward long-range attention, global structural dependencies, and evolutionary alignment signals, which may not be optimally suited to the highly local, motif-driven signatures characteristic of smORF-encoded small proteins. More recently, models like MCWS-Transformers (Ranjan et al., 2023) have further highlighted the importance of modeling local sequence contexts in protein representation learning. These developments suggest that compact architectures with enhanced local motif extraction may be particularly useful for microbial smORF-encoded small protein modeling.

In this study, we introduce TinyProteinTransformer (TPT), a lightweight CNN–Transformer hybrid encoder specifically designed for microbial smORF-encoded small proteins (Figure 1). The main contributions of this study are as follows: (1) We propose a compact CNN-Transformer architecture that combines multi-scale convolutional filters with Transformer-based contextual modeling to capture local motifs and sequence-level dependencies in short peptides. (2) We develop a smORF-adapted pretraining strategy on the GMSC dataset by integrating masked language modeling with contrastive learning. (3) We evaluate TPT across multiple downstream peptide/protein classification tasks and further examine its key components through ablation, hidden-dimension, representation, and efficiency analyses. Through this combination of architecture and training strategy, TPT achieves competitive performance while maintaining high computational efficiency across the evaluated downstream tasks, compared with established protein language models such as the ESM2 and Prot series.

Figure 1.

Diagram illustrating a transformer-based model architecture for protein analysis. Panel a displays input embeddings processed through dropout, layer normalization, and stacked transformer blocks containing multi-head attention and feedforward layers. Panel b depicts downstream tasks such as masked language modeling, contrastive learning, with corresponding loss functions including cross entropy and InfoNCE.

Architecture of TinyProteinTransformer and its self-supervised training objectives. (a) Model architecture. The TinyProteinTransformer (TPT) combines multi-scale convolutional filters with deep Transformer encoder blocks to model both local motifs and global contextual features in small proteins. Input sequences (≤128 aa) are embedded with learned token and positional encodings, passed through four parallel 1D convolutions (kernel sizes 3/5/7/9), and fused via residual projection. The resulting features are processed by stacked Transformer layers, and a learned attention-pooling module aggregates residue representations into a fixed-length embedding. (b) Pretraining objectives. TPT is trained with two complementary self-supervised tasks: masked language modeling (MLM), where 15% of residues are masked and predicted using cross-entropy loss; and contrastive learning, where two augmented views of each peptide are encoded and optimized with an InfoNCE loss to bring paired representations together while separating negatives. The combined objectives shape both token-level and sequence-level representations.

2. Materials and methods

2.1. Pretraining dataset

TinyProteinTransformer (TPT) was pretrained on representative sequences from the Global Microbial smORF Catalog (GMSC) reported by Duan et al. (2024). The GMSC was constructed from 63,410 assembled metagenomes and 87,920 isolate microbial genomes from 75 countries, and contains 287,926,875 non-redundant smORF families clustered at 90% amino acid identity and 90% coverage (GMSC10.90). In this study, we used the officially released version of GMSC10.90 without additional processing. All sequences were truncated to a maximum length of 128 residues.

To assess potential data leakage between the GMSC pretraining corpus and the downstream benchmarks, we performed a two-level sequence similarity analysis across all six downstream datasets. First, exact string matching was applied by streaming all pretraining sequences and checking against the complete set of benchmark sequences. Second, near-identity search was conducted using MMseqs2 easy-search (Steinegger and Söding, 2018) at 90% sequence identity and 80% minimum coverage (−-min-seq-id 0.90 -c 0.80 --cov-mode 0), corresponding to the redundancy threshold used in the original GMSC construction. Exact match rates and near-identity overlap rates per dataset are reported in Supplementary Table S1.

2.2. Model architecture

The model employs pre-layer normalization (norm_first = True) for training stability; CNN kernel sizes of 3, 5, 7, and 9 were chosen to span typical functional motif lengths observed in short proteins; dropout of 0.05 was set conservatively given the large pretraining corpus; and the contrastive temperature τ = 0.1 and weighting coefficient λ = 0.05 follow standard InfoNCE practice, with λ set low to ensure MLM serves as the dominant training signal. Key architectural and training hyperparameters are listed in Supplementary Table S2.

2.2.1. Input representation

For each amino acid sequence input to the model with a length less than 128, it is tokenized using a custom vocabulary: 20 canonical amino acids, ambiguity symbols, masking token, and a padding token. Tokens are then mapped to a 640-dimensional embedding via a learned lookup table (Equation 1):

ei=E(xi)∈ℝ640,x=(x1,x2,…,xL),L≤128 (1)

To encode residue positions, a trainable positional embedding is used (Equation 2):

pi=P(i)∈ℝ640 (2)

The positional embedding and token embeddings are summed to each other, and sequences shorter than 128 residues are padded (Equation 3):

H0(i)=ei+pi (3)

2.2.2. Multi-scale convolutional motif encoder

A parallel multi-scale 1D convolutional module captures local sequence motifs complementary to global Transformer self-attention, with kernel sizes 3, 5, 7, and 9 spanning the range of typical short-protein functional motif lengths. The number of output channels per branch is 256, and the activation function is ReLU. For each kernel size K (Equation 4–5):

Ck=ReLU(Conv1Dk(H0)),whereCk∈ℝL×256 (4)
H1=LayerNorm(H0+Wproj[C3;C5;C7;C9]+bproj)∈ℝL×640 (5)

2.2.3. Transformer layers

Global sequence semantics are modeled using a 20-layer Transformer encoder (PyTorch nn. TransformerEncoder), with layer normalization applied in advance to improve stability. Stacking 20 such layers yields the final contextual representation (Equation 6):

Henc=TransformerEncoder(H1)∈ℝL×640 (6)

2.2.4. Attention-based global pooling

To condense the residue-level representations into a fixed-length embedding suitable for contrastive learning and downstream tasks, we introduced a learned attention pooling operator. Each position i is assumed to have a scalar importance score (Equation 7):

si=w⊤hi,w∈ℝ640 (7)

These scores are converted into normalized attention weights through softmax (Equation 8):

αi=exp(si)∑j=1Lexp(sj) (8)

The final sequence embedding is the weighted average (Equation 9):

v=∑i=1Lαihi∈ℝ640 (9)

This allows the encoder to emphasize residue encoding functional motifs while suppressing padding and low-information positions.

2.2.5. Gated attention

To improve the selectivity of global context modeling, we incorporate a head-specific gated self-attention into each Transformer encoder layer (Vaswani et al., 2017; Qiu et al., 2025). The attention output (Equation 10):

Y∈ℝL×d (10)

It is modulated by a sigmoid gate computed from the input hidden states (Equation 11):

G=σ(HWg),Wg∈ℝd×H (11)

Where H denotes the number of attention heads. The attention output is reshaped into a head-wise form, multiplied by the gating scores, and then merged. This gating is applied head-wise and was tested as an optional variant (TPT_gated) in downstream tasks; the base model uses standard multi-head attention. We hypothesized that head-specific gating might improve selectivity on tasks dominated by localized functional motifs.

2.3. Self-supervised Pretraining

We pre-trained TinyProteinTransformer in a purely self-supervised fashion using masked language modeling at the residue level and contrastive representation learning at the sequence level.

2.3.1. Masked language modeling

To learn residue-level semantics and short-range patterns, we adopted a Bert-style masked language modeling target for the small protein sequence (Devlin et al., 2019). Each input peptide was tokenized and filled to a fixed length of 128 residues. The token set includes 20 typical amino acids, the MASK token, and the PAD token. During the loss calculation process, the filling position is ignored.

To construct the mask prediction target, we combine a span mask and a per-token mask.

For each sequence, we aim to choose 15% of the residues as the targets for masking. The model performs span masking with a 30% probability, where short consecutive fragments (2–5 residues) are removed and replaced with <MASK> tokens. This encourages the model to learn how to reconstruct local motifs and short functional patterns. Under the remaining 70% probability, token-masking positions are selected at the same 15% masking rate.

During the masking period, the replaced positions are recorded, and all unmasked positions are assigned an ignore tag to prevent any loss. The obtained mask sequence is then passed through the encoder to obtain the representations. The linear prediction head projects each hidden state back into the vocabulary space, generating a distribution over all possible amino acid tokens for each position.

The MLM loss was computed using standard cross-entropy (Equation 12):

ℒMLM=1∣ℳ∣∑i∈ℳCE(z^i,xi) (12)

Where ℳ denotes the set of masked residues, xi is the original amino acid at position i, and z^i is the model’s predicted distribution.

The encoder learns to infer the identity of masked residues from their surroundings, thereby being forced to capture motif patterns and short-range dependencies.

2.3.2. Contrastive learning

To better learn the complete features of peptides, we added sequence-level supervision to the training process.

The first type of augmentation is random deletion. A small continuous segment of amino acids (usually 2 to 5) in the peptide will be deleted with a 30% probability. This operation imitates small, naturally occurring truncations in peptides.

The second transformation is conservative substitution, which randomly selects an amino acid position with a 15% probability and replaces it with another amino acid with similar chemical properties. For instance, hydrophobic amino acids (such as I, V, L), aromatic ones (F, W, Y), or positively charged ones (K, R, H). It imitates common mutations in nature that do not alter their core functions or structures. After processing, each sequence is padded to 128 tokens, so that each original peptide chain corresponds to two input data points for contrastive learning.

Following InfoNCE (van den Oord et al., 2018), L2-normalized embeddings from the two augmented views are compared using dot-product similarity at temperature τ = 0.1, with same-sequence pairs as positives and all other in-batch pairs as negatives (Equation 13):

ℒcon=−1B∑i=1Blogexp(sim(i(1)z,i(2)z)/τ)∑j=1Bexp(sim(i(1)z,j(2)z)/τ) (13)

In each training batch, there are both mask sequences for training the mask language model (MLM) and pairs of augmented sequences for contrastive learning. We train the entire model in an end-to-end manner. The final training loss is the weighted sum of these two losses: the total loss is equal to the MLM loss plus λ multiplied by the contrastive loss (where λ takes a value of 0.05). This setting makes the MLM loss the primary training signal while allowing the contrastive loss to serve as a regularization term, helping optimize the model’s global embedding space.

These two training objectives share the encoder parameters and the attention pooling module. Therefore, the model can obtain complementary supervisory information.

2.3.3. Pretraining procedure

All pre-training experiments were conducted using the representative peptide sequence of GMSC10.90. The end-to-end optimization of the TinyProteinTransformer encoder was performed using the AdamW optimizer (Loshchilov and Hutter, 2019) with a learning rate of 2e-5. We used the autocast and GradScaler modules of PyTorch (Paszke et al., 2019) for mixed-precision training to enhance memory efficiency and throughput on NVIDIA GPUs (two NVIDIA RTX 5090 s). When multiple GPUs are available, the model is wrapped in a data-parallel manner, synchronizing gradients across devices at each optimization step.

In each training step, the model simultaneously generates MLM token predictions and embeds them into two augmented sequences. The total loss is calculated as the weighted sum of the MLM and InfoNCE losses (weight = 0.05).

A global random seed of 42 was set for reproducibility. No validation split was used during self-supervised pre-training; convergence was monitored via training loss curves (Figure 2). The best model checkpoint was selected based on the lowest average loss across epochs.

Figure 2.

Line chart comparing masked language modeling (MLM) loss in blue and contrastive loss in red over five thousand training steps, with both losses starting high and sharply decreasing, contrastive loss approaching zero while MLM loss stabilizes above one.

Retraining loss dynamics of TinyProteinTransformer on the GMSC90 dataset. Training curves for masked language modeling (MLM) and contrastive loss. Both losses exhibit a rapid initial decrease followed by stabilization (note the axis scaling in the plateau region).

Three epochs of pre-training were conducted on the entire GMSC dataset. Due to the large number of sequences and mixed self-supervised targets in each epoch, this was sufficient to enable the model to reach a stable loss platform. The resulting pre-trained model serves as the initialization for all downstream classification and representation learning evaluations.

2.4. Downstream evaluation

We performed supervised binary classification on six benchmark datasets spanning functionally distinct classes of microbial and therapeutic peptides: the AMPlify antimicrobial peptide (AMP) dataset, the ToxMSRC toxic peptide (TOX) dataset, the BaPreS bacteriocin (BCN) dataset, the PaCRISPR anti-CRISPR (Acr) dataset, the Quorumpeps quorum sensing peptide (QSP) dataset, and the pLM4CPPs cell-penetrating peptide (CPP) dataset.

For antimicrobial peptide prediction, we used the final non-redundant benchmark dataset released by the AMPlify framework (Li et al., 2022). This dataset contains experimentally verified AMP sequences derived from public peptide databases, including APD3 and DADP, as well as non-AMP sequences selected from UniProtKB/Swiss-Prot. Redundancy reduction was performed in the original AMPlify study using CD-HIT (Li and Godzik, 2006) at 90% AAI. We used the published benchmark without additional processing, comprising 4,173 AMP sequences and 4,173 non-AMP sequences.

For toxicity prediction, we used the benchmark dataset established in the ToxMSRC framework (Zhang et al., 2025). Toxic peptides were collected from CSM-Toxin (Morozov et al., 2023), ToxinPred2, and ATSE (Wei et al., 2021). Non-toxic sequences were obtained from UniProt using keyword-based exclusion of toxin- and antimicrobial-related entries. The combined dataset was then filtered for inconsistent annotations and clustered at 90% AAI with CD-HIT. We used the released benchmark, comprising 2,138 toxic and 5,375 non-toxic peptides.

For bacteriocin prediction, we used the benchmark dataset from BaPreS (Akhter and Miller, 2023), comprising experimentally validated bacteriocin sequences from the BAGEL and BACTIBASE databases, along with non-bacteriocin sequences from the RMSCNN dataset; redundancy was removed using CD-HIT at 90% sequence identity. The dataset was restricted to sequences no longer than 100 amino acids to align with the definition of small proteins, yielding 242 sequences (229 bacteriocin and 13 non-bacteriocin) for evaluation.

For anti-CRISPR (Acr), we used the benchmark dataset released with PaCRISPR (Wang et al., 2020), which includes experimentally characterized Acr proteins sourced from anti-CRISPRdb, together with non-Acr negative sequences derived from Acr-containing phages and bacterial mobile genetic elements. After length filtering, 294 sequences (47 Acr and 247 non-Acr) were retained.

For quorum-sensing peptide prediction, we used the benchmark dataset employed by Rajput et al. (2015), comprising 220 experimentally verified quorum-sensing peptides sourced from the Quorumpeps database and 220 non-QSP negative sequences, totalling 440 balanced sequences.

For cell-penetrating peptide prediction, we used the benchmark dataset established in the pLM4CPPs framework (Kumar et al., 2025), integrating experimentally validated CPPs from CPPsite2.0, C2Pred, CellPPD, MLCPP 2.0, and KELM-CPPpred. After redundancy removal, the dataset comprises 1,399 CPP and 4,080 non-CPP.

2.5. Frozen-encoder downstream evaluation

For each pre-trained encoder (including TinyProteinTransformer and baselines), we froze the encoder parameters and extracted fixed-dimensional sequence embeddings. For TinyProteinTransformer, embeddings were obtained from the learned attention pooling layer (640 dimensions). For ESM-series models and ProtBERT, we used mean pooling over the last layer representations, since they do not provide attention pooling.

A single linear classification head (nn. Linear(embedding_dim, 2)) was trained from scratch for each encoder. The AdamW optimizer (learning rate = 1e-3, weight decay = 0.01) was used with cross-entropy loss. Training ran for exactly 5 epochs without early stopping or learning-rate scheduling, with a batch size of 64. To ensure reliable and unbiased performance estimates, we employed 5-fold stratified cross-validation (StratifiedKFold from scikit-learn, shuffle = True, random_state = 42), preserving class proportions in each fold. In each fold, 80% of the data were used to train the linear head, and 20% were used for validation.

Sequence representations were pre-computed in batches using frozen encoders to avoid redundant forward passes. Evaluation metrics included accuracy (ACC), macro F1-score, and area under the ROC curve (AUC), computed using scikit-learn functions. Metrics were averaged across the 5 folds, with standard deviations reported as ± values. All experiments fixed a global random seed of 42 (via the seed_everything function) to ensure full reproducibility across random splits, model initialization, and data shuffling.

2.6. Comparison with non-pretrained prediction methods

To contextualize TPT’s pretrain-then-probe performance against non-pretrained supervised baselines, we further evaluated three sequence classifiers trained end-to-end from random initialization on the AMP benchmarks: AMP Scanner v2 (Veltri et al., 2018), comprising an embedding layer, a single 1D convolutional layer (64 filters, kernel size 16), a max-pooling layer, and an LSTM layer (100 units); AMPlify (Li C. et al., 2022), comprising a bidirectional LSTM (512 units per direction) followed by 32-head scaled dot-product attention and a context attention pooling layer; and a Fast-MCWS Transformer (Mahala et al., 2025), comprising a convolutional front-end with average pooling followed by three parallel windowed self-attention blocks at 25, 50, and 100% of the pooled sequence length. All three models were implemented in PyTorch and trained under identical conditions: AdamW optimizer, learning rate 1 × 10−3, batch size 32–64, maximum 200 epochs with early stopping monitored on validation accuracy (patience = 10), and the same 5-fold stratified splits (StratifiedKFold, random_state = 42) as the frozen-encoder evaluation. No pretrained weights were loaded. AMP Scanner v2 and AMPlify used one-hot or integer-encoded amino acid sequences as input; Fast-MCWS used a 4-mer dictionary tokenization scheme built from the downstream dataset sequences. This setup reflects each method’s intended supervised use case and enables direct comparison of end-to-end training against TPT’s pretrain-then-probe paradigm.

2.7. Ablation studies

To evaluate the contribution of key architectural components and training objectives, we constructed several TPT variants by removing or replacing one component at a time. TPT-NoCL was trained using only the masked language modeling objective, with the InfoNCE contrastive loss removed. TPT-MeanPool replaced the learned attention-pooling module with mean pooling over all non-padding residue positions. TPT-NoCNN removed the parallel multi-scale convolutional branch and passed token embeddings directly into the Transformer encoder. In addition, TPT-Gated was evaluated as a gated-attention variant to examine the effect of head-specific sigmoid gating.

All ablation variants were evaluated using the identical downstream protocol as the full model: frozen encoder embeddings extracted via attention pooling (or mean pooling for TPT-MeanPool), a single linear classification head trained for 5 epochs (AdamW, lr = 1 × 10−3, weight decay = 0.01), and 5-fold stratified cross-validation across all four benchmarks. Results are reported as mean ± standard deviation of ACC, macro F1, and AUC across folds.

2.8. Representation visualization

To qualitatively compare how different pre-trained encoders construct peptide representations, we performed a two-dimensional t-SNE visualization of the toxicity dataset (Van der Maaten and Hinton, 2008). For each sequence, a fixed-length embedding is obtained using the frozen encoder. The t-SNE (n_components = 2, perplexity = 30, learning_rate = 200, random_state = 42) implemented in scikit-learn was used to reduce the dimensionality of each model independently.

3. Results

3.1. Pretraining loss dynamics

TPT was pretrained on microbial small-protein sequences from the GMSC10.90 catalog, using a joint objective that combined masked language modeling (MLM) with a contrastive loss. The MLM term encouraged the model to recover masked residues from the local sequence context. Meanwhile, the contrastive term aligned augmented views of the same sequence in the embedding space. Figure 2 shows the two loss curves over 50,000 training steps. Both terms decreased rapidly during early training and then converged to stable plateaus. The MLM loss settled at approximately 1.8, with only minor late-stage fluctuations. The contrastive loss fell to consistently low values, and we observed no signs of representation collapse. Together, these dynamics indicate that the two objectives were optimized in a stable manner and provided a well-behaved starting point for downstream evaluation.

3.2. Overall performance across six downstream tasks

We systematically evaluated our TPT and TPT_gated models against three pretrained baselines (ESM2-35 M, ESM2-150 M, and ProtBERT). The six tasks covered functionally distinct classes of microbial peptides and proteins: antimicrobial peptides (AMP), toxic peptides (TOX), bacteriocins (BCN), anti-CRISPR proteins (Acr), quorum-sensing peptides (QSP), and cell-penetrating peptides (CPP), assessed using 5-fold stratified cross-validation. All encoders were frozen and evaluated with a linear classification head, so that performance reflected the quality of pretrained representations. Accuracy, F1, and AUC are reported in Supplementary Table S7.

Because ESM2 and ProtBERT were evaluated using mean-pooled residue embeddings, we also included TPT-MeanPooling in Supplementary Table S7 as a common mean-pooling comparison. This variant replaced the learned attention-pooling module with mean pooling while keeping the remaining TPT encoder unchanged. TPT-MeanPooling remained competitive with the baseline PLMs across the evaluated tasks, suggesting that TPT performance was not solely driven by the attention-pooling module.

On AMP and TOX benchmarks, the TPT showed competitive performance compared with the pretrained baseline models. For AMP, TPT obtained the highest mean AUC (0.930 ± 0.008), but this did not differ significantly from ESM2-150 M (0.928 ± 0.004; p = 0.68). A similar pattern held for TOX, where TPT had the highest mean AUC (0.930 ± 0.006), followed closely by ESM2-150 M (0.925 ± 0.009; p = 0.12).

On BCN, all models reached accuracy between 0.94 and 0.97. ESM2-150 M obtained the highest mean AUC (0.955 ± 0.040) and TPT_gated the second-highest (0.947 ± 0.045), but the two were again statistically indistinguishable (p = 0.77). Notably, TPT_gated achieved the best F1 (0.983 ± 0.015). The wide confidence intervals on this task reflect its small, unbalanced benchmark and limited resolution between comparisons.

On QSP and CPP, TPT also showed strong performance. For QSP, TPT achieved the best overall results among all compared models, with the highest ACC (0.868 ± 0.027), F1 score (0.872 ± 0.031), and AUC (0.923 ± 0.021). For CPP, TPT again achieved the highest overall performance, with an ACC of 0.901 ± 0.008, F1 score of 0.793 ± 0.022, and AUC of 0.938 ± 0.009. TPT_gated showed comparable performance on CPP, with an AUC of 0.937 ± 0.009, indicating that both TPT variants retained effective on these peptide-related tasks. Across the six tasks, the two TPT variants showed complementary behavior: the base TPT model performed strongly on AMP, TOX, QSP, and CPP, whereas TPT_gated performed particularly well on BCN.

The Acr task revealed a clear and significant separation between the two model families. The three pretrained baselines collapsed to majority-class prediction, with an F1 of 0.000 and an AUC at or below chance (ESM2-35 M 0.463, ESM2-150 M 0.428, ProtBERT 0.322). In contrast, TPT and TPT_gated retained predictive signal, achieving AUCs of 0.696 ± 0.071 and 0.692 ± 0.098, respectively, with non-trivial F1 scores. The advantage of TPT over all baselines was consistent across all five folds and statistically significant; for example, p = 0.015 against ESM2-150 M and p = 0.0004 against ProtBERT. The same pattern was evident in the embedding geometry. On Acr, all three baselines yielded negative silhouette coefficients (−0.034 to −0.050), indicating that samples lay closer to the opposite class on average, whereas TPT retained a positive coefficient. Silhouette values for the evaluated tasks are reported in Supplementary Table S3. One possible explanation for this robustness is the contrastive component of pretraining. By encouraging distinct sequences to occupy more separable regions in the embedding space, contrastive learning may help preserve minority-class signals that are difficult to capture using MLM-only objectives under strongly imbalanced settings. This interpretation is further supported by the ablation results (Supplementary Table S8).

To further evaluate TPT against task-specific SOTA methods, we compared it with two representative state-of-the-art AMP prediction models, AMPlify and AMPScannerV2, using the same AMP benchmark dataset and evaluation protocol. As shown in Table 1, TPT achieved the best overall performance among the compared methods. This additional comparison further supports the effectiveness of TPT for AMP prediction.

Table 1.

Comparison of TPT with representative non-pretrained methods on the AMP prediction task.

Model Training setting ACC F1 AUC
TPT-FullFinetune Full fine-tuning 0.9161 ± 0.0084 0.9157 ± 0.0079 0.9733 ± 0.0038
TPT-Gated-FullFinetune Full fine-tuning 0.9099 ± 0.0068 0.9102 ± 0.0075 0.9692 ± 0.0061
AMPScannerV2 Trained from scratch 0.9035 ± 0.0022 0.9011 ± 0.0024 0.9555 ± 0.0020
AMPlify Trained from scratch 0.8766 ± 0.0241 0.8752 ± 0.0238 0.9449 ± 0.0174
FastMCWS-FromScratch Trained from scratch 0.8663 ± 0.0100 0.8651 ± 0.0087 0.9406 ± 0.0045

All models were retrained and evaluated on the same AMP benchmark dataset using the same data splits and evaluation protocol. AMPlify and AMPScannerV2 were included as representative AMP-specific prediction methods, and FastMCWS-FromScratch was included as an MCWS-based comparison. ACC, F1 score, and AUC are reported as mean ± standard deviation across five folds where available.

3.3. Contribution of each component

To assess the contribution of key model components, we compared the full TPT model with ablation variants removing contrastive learning, replacing attention pooling with mean pooling, removing the multi-scale CNN branch, or using gated attention (Supplementary Table S8). Removing attention pooling or the CNN branch led to moderate performance reductions across most tasks, as reflected by the AMP AUC decreasing from 0.929 to 0.916 and 0.917, respectively. In contrast, removing the contrastive objective caused the largest performance drop, with AUC decreasing from 0.929 to 0.869 on AMP (p = 0.0001), from 0.955 to 0.716 on BCN (p = 0.007), and from 0.733 to 0.528 on Acr (p = 0.024). These results suggest that local motif extraction and adaptive pooling contribute to TPT performance, while contrastive learning provides a particularly important sequence-level signal for learning discriminative representations, especially in imbalanced or low-resource tasks.

3.4. Impact of hidden dimension

To determine the optimal representational capacity for TPT, we evaluated the effect of the Transformer’s hidden dimension size. We trained variants of the model with hidden dimensions of 320, 480, 640, and 800, keeping all other hyperparameters constant, and assessed their downstream classification performance (Supplementary Table S9).

The results revealed a clear performance peak at a hidden dimension of 480 and 640. On the challenging and highly imbalanced Acr task, increasing the dimension from 320 to 640 substantially improved the mean AUC from 0.561 ± 0.123 to a peak of 0.733 ± 0.040. However, further expanding the capacity to 800 resulted in a performance drop (AUC 0.631 ± 0.094), suggesting the onset of overfitting on the rare minority class. A similar pattern emerged on the AMP benchmark, where the 640-dimensional model achieved the highest mean AUC of 0.930 ± 0.003.

For the toxicity prediction (TOX) task, performance was remarkably stable across all tested capacities, with mean AUCs consistently plateauing between 0.925 and 0.928. This indicates that the task is relatively insensitive to model scaling once a baseline representation capacity is reached. While a smaller dimension (480) yielded a higher mean AUC on the small bacteriocin (BCN) dataset, the 640-dimensional variant maintained an exceptionally high F1 score (0.964 ± 0.020) well within the margin of the dataset’s high variance. Consequently, a hidden dimension of 640 was established as the default configuration for TPT. It possesses sufficient capacity to resolve complex sequence features without over-parameterizing and destabilizing the learning dynamics on limited microbial datasets.

3.5. Model size, inference cost, and accuracy

Having observed that TPT achieved competitive performance across the evaluated benchmarks, we next examined its model size and inference cost (Figure 3). TPT contains 103 M parameters, fewer than ESM2-150 M (150 M) and ProtBERT (420 M), but more than ESM2-35 M (35 M). Among the baseline models, ESM2-150 M generally performed better than ESM2-35 M, indicating that model scale can still be beneficial. However, under the current evaluation setting, TPT achieved competitive performance while maintaining a smaller parameter count than ESM2-150 M and ProtBERT (Figure 3a).

Figure 3.

Two scatter plots compare different protein language models on mean AUC. Panel a shows mean AUC versus model parameters, with TPT and TPT-Gated achieving the highest mean AUC around 0.87 with fewer parameters, surpassing ESM2-150M, ESM2-35M, and ProtBERT. Panel b shows mean AUC versus latency, indicating TPT and TPT-Gated balance high mean AUC and low latency, outperforming ESM2-150M and ProtBERT in efficiency. Legends identify model shapes and colors.

Parameter and inference efficiency of TPT. Model performance was evaluated on the six downstream prediction task. (a) Mean AUC plotted against model parameter size. (b) Mean AUC plotted against inference latency. Latency is reported as milliseconds per sample. All encoders were evaluated on the same NVIDIA RTX 5090 GPU.

We further compared inference latency across models (Figure 3b; Supplementary Table S5). TPT encoded a single sequence in approximately 6.6 ms on one NVIDIA RTX 5090 GPU, compared with approximately 12 ms for ESM2-35 M, 16.5 ms for ProtBERT, and 28.6 ms for ESM2-150 M. This lower inference cost may be related to the compact CNN–Transformer architecture and lightweight sequence-processing pipeline used in TPT. Taken together, these results suggest that, within the evaluated benchmarks, TPT provides a favorable balance between predictive performance, model size, and inference efficiency, rather than indicating that model scale is unimportant in general.

3.6. Representation, geometric and class structure

To further examine the representation structure learned by different encoders, we projected frozen encoder embeddings into two dimensions using t-SNE (Figure 4). The complete t-SNE projections for the remaining model–task combinations are provided in Supplementary Figure S1. On the toxicity task, TPT showed a relatively clear visual separation between toxic and non-toxic peptides, whereas the baseline models displayed more fragmented projections with greater class overlap. On the AMP task, all models showed noticeable overlap between AMP and non-AMP sequences, indicating that two-dimensional t-SNE projections alone may not fully capture the discriminative structure of the embedding space. To provide a quantitative complement to the t-SNE visualization, we further calculated silhouette coefficients for each model across downstream tasks (Supplementary Table S3). The silhouette coefficients did not always parallel the linear-probe AUC results. For AMP and TOX, TPT achieved the highest AUCs despite having lower silhouette coefficients than some baseline models. In contrast, on the Acr task, TPT was the only model with a positive silhouette coefficient, whereas all baseline models showed negative values. These results suggest that t-SNE and silhouette analyses provide complementary information about embedding geometry and should be interpreted together with downstream classification metrics.

Figure 4.

Four t-SNE scatter plots compare protein representation clustering for AMP and TOX classes using TinyProteinTransformer and ESM2_150M models, with red and blue dots indicating two distinct sample categories in each plot.

t-SNE visualization of peptide embeddings across different downstream tasks. Two-dimensional t-SNE projections of embeddings were generated using TinyProteinTransformer and ESM2-150M. (a) t-SNE visualization for the AMP task. (b) t-SNE visualization for the TOX task. Red and blue dots indicate the two class labels in each task.

4. Discussion

The results suggest that modeling of microbial small proteins may benefit from architectural approaches that differ from those used in large protein language models (PLMs). Most existing PLMs are trained primarily on long, evolutionarily conserved, and domain-rich proteins (Elnaggar et al., 2021; Lin et al., 2023). In contrast, smORFs are characterized by rapid evolutionary turnover, relatively sparse domain organization, and frequent reliance on short functional motifs (Durrant and Bhatt, 2021; Miravet-Verde et al., 2019; Storz et al., 2014). These features may limit the effectiveness of architectures designed mainly to capture long-range dependencies and deep evolutionary signals.

To address this mismatch, we developed TinyProteinTransformer (TPT), which combines convolution-based motif extraction with lightweight Transformer-based contextual modeling. This design is motivated by previous studies suggesting that CNNs can effectively detect local biochemical patterns relevant to peptide function (Veltri et al., 2018), while Transformer layers provide a flexible mechanism for modeling broader sequence context. The combination of these two inductive biases offers a plausible explanation for the favorable efficiency–performance trade-off observed for TPT. Furthermore, contrastive learning, employed in this work largely as a regularizer, is known to induce functional clustering even without homology (Lu et al., 2020), which is consistent with the coherent geometric structure revealed in the TPT embeddings. We also note that alternative approaches have been proposed to introduce locality into protein sequence modeling. For example, MCWS-Transformers employ multi-context-window self-attention to capture local sequence context, while subsequent CNN + MCWS architectures combine convolutional operations with localized attention mechanisms for efficient protein function prediction (Ranjan et al., 2023; Mahala et al., 2025). These studies further highlight the importance of explicitly modeling local sequence patterns in protein representation learning.

This work also has broader implications for protein representation learning in microbial genomics. As GMSC and other large-scale metagenomic resources continue to expand the number of reported microbial peptides, the domain-adapted encoder may become increasingly important. In high-throughput microbiome studies, where millions of candidate smORFs may need to be screened (Steinegger and Söding, 2018), parameter-efficient models provide a practical alternative to multi-billion-parameter PLMs. The performance of TPT suggests that the effective representation of smORFs may depend not only on model scale, but also on architectural priors tailored to the sequence and biochemical characteristics of short microbial proteins.

The ablation experiments further support the contribution of the main TPT components. Removing attention pooling or the CNN module caused moderate performance reductions, suggesting that adaptive residue-level aggregation and local motif extraction both contribute to the learned representations. In contrast, removing the contrastive objective led to the most consistent decline across the four downstream tasks, particularly on the imbalanced Acr task, where AUC approached the chance level. These findings suggest that contrastive learning provides an important sequence-level discriminative signal that complements the residue-level MLM.

The effect of gated attention was also task-dependent across the four downstream benchmarks. Compared with the base TPT model, TPT_gated did not provide consistent improvements on AMP, TOX, or Acr, but showed better performance on the BCN task. This suggests that gated attention is not universally beneficial, and its utility may depend on the type of sequence signal required by each task. One possible explanation is that gating may be beneficial when functional information is concentrated in a limited set of informative residues or local motifs (Rathore et al., 2024), as may occur in some bacteriocin-related sequences. In contrast, for tasks such as AMP prediction, where activity may depend more strongly on distributed physicochemical properties such as net charge and hydrophobicity (Hancock and Sahl, 2006; Wimley, 2010), gating may attenuate globally informative signals. These findings indicate that gated attention should be considered a task-dependent module rather than a generally superior replacement for standard attention.

Silhouette coefficients and AUC scores conveyed complementary but distinct aspects of representation geometry. On AMP and TOX tasks, TPT achieved the highest AUC values while yielding lower silhouette coefficients than the baselines (Supplementary Table S3). This reflects a known dissociation between cluster compactness and linear separability: the contrastive objective may produce embeddings well-organized for linear discrimination without forming compact globular clusters. On the Acr task, all three baselines yielded negative silhouette coefficients, indicating that their embeddings placed Acr sequences closer to the opposite class on average, consistent with their collapse under linear probing. TPT retained the only positive silhouette coefficient on this task, aligning with its substantially higher AUC.

Several limitations of this study should be acknowledged. (1) Functional annotations for smORFs remain relatively limited, constraining the size and diversity of available benchmark datasets. (2) Our downstream evaluation was conducted using a frozen-encoder linear-probing framework to enable fair comparisons across models. Although this strategy provides a standardized evaluation setting, it may not fully reflect the achievable performance under end-to-end fine-tuning. (3) While TPT integrates both local and global sequence information, it does not currently incorporate structural supervision, which has recently shown favorable effects in large-scale PLMs (Hartout et al., 2025; Lin et al., 2023; Wu et al., 2024; Zhang et al., 2024). (4) Potential sequence overlap between the GMSC10.90 pretraining corpus and downstream benchmark datasets should also be considered. Although exact sequence overlap was limited overall, near-identity overlap showed dataset-dependent variation and was relatively higher in several small benchmarks, such as Acr and BCN (Supplementary Table S1). These overlaps are partly expected because smORF-encoded peptides are short and the GMSC corpus is extremely large, making near-identical short peptide sequences difficult to completely avoid. Because TPT was pretrained in a self-supervised manner without downstream task labels, this does not represent supervised label leakage. Nevertheless, sequence-level homology between pretraining and downstream datasets may still influence performance estimates. Future work should therefore explore the integration of geometric or structure-aware training objectives, as well as broader evaluation across additional classes of smORF-associated functions, including regulatory, signaling, and membrane-associated micropeptides.

Taken together, our findings suggest that effective modeling of microbial smORFs depends not only on model scale, but also on inductive biases tailored to the sequence characteristics of short, rapidly evolving proteins. By integrating local motif extraction with Transformer-based contextual modeling, TPT provides a compact and computationally efficient framework for learning microbial peptide representations. While further validation across a broader range of smORF-associated functions is needed, this work highlights the potential of specialized protein encoders as scalable tools for peptide analysis in microbial genomics.

5. Conclusion

In conclusion, this study demonstrates that accurate modeling of microbial smORFs requires architectures informed by their distinct biological properties, rather than directly transferring designs optimized for longer proteins. Protein sequence modeling stands to benefit from domain- and length-specialized architectures rather than relying solely on ever-larger general-purpose models. While training universal PLMs incurs high computational cost and inference latency, compact, domain-adapted encoders such as TPT offer a more practical and scalable solution for large-scale peptide analysis. By combining local motif extraction with robust context modeling, TPT provides an efficient, biologically motivated framework for advancing small protein research in microbial genomics.

Acknowledgments

We would like to thank Yiqian Duan for her valuable assistance and supportive discussions during the course of this study.

Funding Statement

The author(s) declared that financial support was not received for this work and/or its publication.

Footnotes

Edited by: Junfeng Su, Xi'an University of Architecture and Technology, China

Reviewed by: Nandan Kumar, Kansas State University, United States

Ashish Ranjan, C.V. Raman Global University, India

Data availability statement

The original contributions presented in the study are included in the article/Supplementary material. The code and data used to reproduce the analyses are available at https://github.com/F4NG66/TPT. Further inquiries can be directed to the corresponding author.

Author contributions

FS: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Visualization, Writing – original draft. JZ: Validation, Visualization, Writing – original draft. CZ: Methodology, Project administration, Supervision, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmicb.2026.1839420/full#supplementary-material

Table_1.DOCX (13.1KB, DOCX)
Table_2.DOCX (17.9KB, DOCX)
Table_3.DOCX (11.8KB, DOCX)
Table_4.XLSX (19KB, XLSX)
Table_5.DOCX (12.5KB, DOCX)
Table_6.DOCX (2.1MB, DOCX)

References

  1. Akhter S., Miller J. H. (2023). BaPreS: a software tool for predicting bacteriocins using an optimal set of features. BMC Bioinformatics 24:313. doi: 10.1186/s12859-023-05330-z, [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Anderson D. M., Anderson K. M., Chang C. L., Makarewich C. A., Nelson B. R., McAnally J. R., et al. (2015). A micropeptide encoded by a putative long noncoding RNA regulates muscle performance. Cell 160, 595–606. doi: 10.1016/j.cell.2015.01.009, [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Andreev D. E., O'Connor P. B., Fahey C., Kenny E. M., Terenin I. M., Dmitriev S. E., et al. (2015). Translation of 5′ leaders is pervasive in genes resistant to eIF2 repression. eLife 4:e03971. doi: 10.7554/eLife.03971, [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Bhatta A., Atianand M., Jiang Z., Crabtree J., Blin J., Fitzgerald K. A. (2020). A mitochondrial micropeptide is required for activation of the Nlrp3 inflammasome. J. Immunol. 204, 428–437. doi: 10.4049/jimmunol.1900791, [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Bi P., Ramírez-Martínez A., Li H., Cannavino J., McAnally J. R., Shelton J. M., et al. (2017). Control of muscle formation by the fusogenic micropeptide myomixer. Science 356, 323–327. doi: 10.1126/science.aam9361, [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Couso J. P., Patraquim P. (2017). Classification and function of small open reading frames. Nat. Rev. Mol. Cell Biol. 18, 575–589. doi: 10.1038/nrm.2017.58, [DOI] [PubMed] [Google Scholar]
  7. Devlin J., Chang M.-W., Lee K., Toutanova K. (2019) BERT: Pre-Training of deep Bidirectional Transformers for Language Understanding Proceedings of NAACL-HLT 2019 Minneapolis, MN: ACL; 4171–4186 [Google Scholar]
  8. Duan Y., Santos-Júnior C. D., Schmidt T. S., Fullam A., de Almeida B. L. S., Zhu C., et al. (2024). A catalog of small proteins from the global microbiome. Nat. Commun. 15:7563. doi: 10.1038/s41467-024-51894-6, [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Durrant M. G., Bhatt A. S. (2021). Automated prediction and annotation of small open reading frames in microbial genomes. Cell Host Microbe 29, 121–131.e4. doi: 10.1016/j.chom.2020.11.002, [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Duval M., Cossart P. (2017). Small bacterial and phagic proteins: an updated view on a rapidly moving field. Curr. Opin. Microbiol. 39, 81–88. doi: 10.1016/j.mib.2017.09.010, [DOI] [PubMed] [Google Scholar]
  11. Elnaggar A., Heinzinger M., Dallago C., Rehawi G., Wang Y., Jones L., et al. (2021). ProtTrans: toward understanding the language of life through self-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, 7112–7127. doi: 10.1109/TPAMI.2021.3095381, [DOI] [PubMed] [Google Scholar]
  12. Hancock R. E., Sahl H. G. (2006). Antimicrobial and host-defense peptides as new anti-infective therapeutic strategies. Nat. Biotechnol. 24, 1551–1557. doi: 10.1038/nbt1267, [DOI] [PubMed] [Google Scholar]
  13. Hao Z., Wu Y., Huang Y., Zhang M., Liu Y., Li Y., et al. (2025). Microprotein PLUM encoded by Lin28b uORF is a cytoplasmic determinant of pluripotency and embryonic development. Nat. Commun. 16:10324. doi: 10.1038/s41467-025-66297-4, [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Hartout P., Chen D., Pellizzoni P., Oliver C., Borgwardt K. (2025). Endowing protein language models with structural knowledge. Bioinformatics 41:btaf582. doi: 10.1093/bioinformatics/btaf582, [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Huang N., Li F., Zhang M., Zhou H., Chen Z., Ma X., et al. (2021). An upstream open reading frame in phosphatase and tensin homolog encodes a circuit breaker of lactate metabolism. Cell Metab. 33, 128–144.e9. doi: 10.1016/j.cmet.2020.12.008, [DOI] [PubMed] [Google Scholar]
  16. Ingolia N. T., Ghaemmaghami S., Newman J. R. S., Weissman J. S. (2009). Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science 324, 218–223. doi: 10.1126/science.1168978, [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Jackson R., Kroehling L., Khitun A., Bailis W., Jarret A., York A. G., et al. (2018). The translation of non-canonical open reading frames controls mucosal immunity. Nature 564, 434–438. doi: 10.1038/s41586-018-0794-7, [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Khitun A., Ness T. J., Slavoff S. A. (2019). Small open reading frames and cellular stress responses. Mol. Omics 15, 108–116. doi: 10.1039/c8mo00283e, [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Kumar N., Du Z., Li Y. (2025). pLM4CPPs: Protein Language Model-Based Predictor for Cell Penetrating Peptides. Journal of Chemical Information and Modeling. 65, 1128–1139. doi: 10.1021/acs.jcim.4c01338, [DOI] [PubMed] [Google Scholar]
  20. Li W., Godzik A. (2006). Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 22, 1658–1659. doi: 10.1093/bioinformatics/btl158, [DOI] [PubMed] [Google Scholar]
  21. Li C., Sutherland D., Hammond S. A., Yang C., Taho F., Bergman L., et al. (2022). AMPlify: attentive deep learning model for discovery of novel antimicrobial peptides effective against WHO priority pathogens. BMC Genomics. 23:77. doi: 10.1186/s12864-022-08310-4, [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., et al. (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130. doi: 10.1126/science.ade2574, [DOI] [PubMed] [Google Scholar]
  23. Loshchilov I., Hutter F. (2019) Decoupled weight decay regularization. In: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA. OpenReview.net.
  24. Lu A. X., Zhang H., Ghassemi M., Moses A. (2020). Self-supervised contrastive learning of protein representations by mutual information maximization. bioRxiv. doi: 10.1101/2020.09.04.283929 [DOI] [Google Scholar]
  25. Lyu Y., Tan B., Li L., Liang R., Lei K., Wang K., et al. (2023). A novel protein encoded by circUBE4B promotes progression of esophageal squamous cell carcinoma by augmenting MAPK/ERK signaling. Cell Death Dis. 14:346. doi: 10.1038/s41419-023-05865-2, [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Mahala A., Ranjan A., Priyadarshini R., Vikram R., Dansena P. (2025). A fast (CNN+MCWS-transformer) based architecture for protein function prediction. Stat. Appl. Genet. Mol. Biol. 24:20240027. doi: 10.1515/sagmb-2024-0027, [DOI] [PubMed] [Google Scholar]
  27. Matsumoto A., Pasut A., Matsumoto M., Yamashita R., Fung J., Monteleone E., et al. (2017). mTORC1 and muscle regeneration are regulated by the LINC00961-encoded SPAR polypeptide. Nature 541, 228–232. doi: 10.1038/nature21034, [DOI] [PubMed] [Google Scholar]
  28. Miravet-Verde S., Ferrar T., Espadas-García G., Mazzolini R., Gharrab A., Sabido E., et al. (2019). Unraveling the hidden universe of small proteins in bacterial genomes. Mol. Syst. Biol. 15:e8290. doi: 10.15252/msb.20188290, [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Morozov V., Rodrigues C. H. M., Ascher D. B. (2023). CSM-toxin: a web-server for predicting protein toxicity. Pharmaceutics 15:431. doi: 10.3390/pharmaceutics15020431, [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Nichols C., Do-Thi V. A., Peltier D. C. (2024). Noncanonical microprotein regulation of immunity. Mol. Ther. 32, 2905–2929. doi: 10.1016/j.ymthe.2024.05.021, [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Paszke A., Gross S., Massa F., Lerer A., Bradbury J., Chanan G., et al. (2019). PyTorch: an imperative style, high-performance deep learning library. In: Adv. Neural Inf. Process. Syst 32, 8024–8035. doi: 10.48550/arXiv.1912.01703 [DOI] [Google Scholar]
  32. Pauli A., Norris M. L., Valen E., Chew G. L., Gagnon J. A., Zimmerman S., et al. (2014). Toddler: an embryonic signal that promotes cell movement via Apelin receptors. Science 343:1248636. doi: 10.1126/science.1248636 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Qiu Z., Zhang Y., Liu H., Wang X. (2025). Gated attention for large language models: non-linearity, sparsity, and attention-sink-free. arXiv preprint arXiv:2505.06708. doi: 10.48550/arXiv.2505.06708 [DOI]
  34. Rao R. M., Liu J., Verkuil R., Meier J., Canny J., Abbeel P., et al. (2021). MSA Transformer. Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research 139, 8844–8856. [Google Scholar]
  35. Rajput A., Gupta A. K., Kumar M. (2015). Prediction and analysis of quorum sensing peptides based on sequence features. PLoS One 10:e0120066. doi: 10.1371/journal.pone.0120066, [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Ranjan A., Fahad M. S., Fernández-Baca D., Tripathi S., Deepak A. (2023). MCWS-transformers: towards an efficient modeling of protein sequences via multi context-window based scaled self-attention. IEEE/ACM Trans. Comput. Biol. Bioinform. 20, 1188–1199. doi: 10.1109/TCBB.2022.3173789, [DOI] [PubMed] [Google Scholar]
  37. Rathore A. S., Choudhury S., Arora A., Tijare P., Raghava G. P. S. (2024). ToxinPred 3.0: an improved method for predicting the toxicity of peptides. Comput. Biol. Med. 179:108926. doi: 10.1016/j.compbiomed.2024.108926, [DOI] [PubMed] [Google Scholar]
  38. Savard J., Marques-Souza H., Aranda M., Tautz D. (2006). A segmentation gene in tribolium produces a polycistronic mRNA that codes for multiple conserved peptides. Cell 126, 559–569. doi: 10.1016/j.cell.2006.05.053, [DOI] [PubMed] [Google Scholar]
  39. Sberro H., Fremin B. J., Zlitni S., Edfors F., Greenfield N., Snyder M. P., et al. (2019). Large-scale analyses of human microbiomes reveal thousands of small, novel genes. Cell 178, 1245–1259.e14. doi: 10.1016/j.cell.2019.07.016, [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Steinegger M., Söding J. (2018). Clustering huge protein sequence sets in linear time. Nat. Commun. 9:2542. doi: 10.1038/s41467-018-04964-5, [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Storz G., Wolf Y. I., Ramamurthi K. S. (2014). Small proteins can no longer be ignored. Annu. Rev. Biochem. 83, 753–777. doi: 10.1146/annurev-biochem-070611-102400, [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Suzek B. E., Huang H., McGarvey P., Mazumder R., Wu C. H. (2007). UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics 23, 1282–1288. doi: 10.1093/bioinformatics/btm098, [DOI] [PubMed] [Google Scholar]
  43. van den Oord A., Li Y., Vinyals O. (2018). Representation learning with contrastive predictive coding. arXiv preprint, arXiv:1807.03748. doi: 10.48550/arXiv.1807.03748 [DOI]
  44. Van der Maaten L., Hinton G. (2008). Visualizing data using t-SNE. J. Mach. Learn. Res. 9, 2579–2605. [Google Scholar]
  45. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A. N., et al. (2017). Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17), Long Beach, CA, USA. In: Adv. Neural Inf. Process. Syst, Red Hook, NY, USA: Curran Associates Inc. 6000–6010. doi: 10.48550/arXiv.1706.0376 [DOI] [Google Scholar]
  46. Veltri D., Kamath U., Shehu A. (2018). Deep learning improves antimicrobial peptide recognition. Bioinformatics 34, 2740–2747. doi: 10.1093/bioinformatics/bty179, [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Wang J., Dai W., Li J., Xie R., Dunstan R. A., Stubenrauch C., et al. (2020). PaCRISPR: a server for predicting and visualizing anti-CRISPR proteins. Nucleic Acids Res. 48, W348–W357. doi: 10.1093/nar/gkaa432, [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Weaver J., Mohammad F., Buskirk A. R., Storz G. (2019). Identifying small proteins by ribosome profiling with stalled initiation complexes. MBio 10:e02819-18. doi: 10.1128/mBio.02819-18, [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Wei L., Ye X., Xue Y., Sakurai T., Wei L. (2021). ATSE: a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural network and attention mechanism. Brief. Bioinform. 22:bbab041. doi: 10.1093/bib/bbab041, [DOI] [PubMed] [Google Scholar]
  50. Wimley W. C. (2010). Describing the mechanism of antimicrobial peptide action with the interfacial activity model. ACS Chem. Biol. 5, 905–917. doi: 10.1021/cb1001558, [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Wu K. E., Yang K. K., van den Berg R., Alamdari S., Zou J. Y., Lu A. X., et al. (2024). Protein structure generation via folding diffusion. Nat. Commun. 15:1059. doi: 10.1038/s41467-024-45051-2, [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Zhang T., Li Z., Li J., Peng Y. (2025). Small open reading frame-encoded microproteins in cancer: identification, biological functions and clinical significance. Mol. Cancer 24:105. doi: 10.1186/s12943-025-02278-x, [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Zhang Z., Lu J., Chenthamarakshan V., Lozano A., Das P., Tang J. (2024). Structure-informed protein language model. arXiv [Preprint]. doi: 10.48550/arXiv.2402.05856 [DOI] [Google Scholar]
  54. Zhang S., Ren J., Liang Y. (2025). An innovative peptide toxicity prediction model based on multi-scale convolutional neural network and residual connection. Bioinformatics 41:btaf462. doi: 10.1093/bioinformatics/btaf462, [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table_1.DOCX (13.1KB, DOCX)
Table_2.DOCX (17.9KB, DOCX)
Table_3.DOCX (11.8KB, DOCX)
Table_4.XLSX (19KB, XLSX)
Table_5.DOCX (12.5KB, DOCX)
Table_6.DOCX (2.1MB, DOCX)

Data Availability Statement

The original contributions presented in the study are included in the article/Supplementary material. The code and data used to reproduce the analyses are available at https://github.com/F4NG66/TPT. Further inquiries can be directed to the corresponding author.


Articles from Frontiers in Microbiology are provided here courtesy of Frontiers Media SA

RESOURCES