Abstract
Background
Neoantigens—tumor-specific peptides generated by somatic mutations—are central targets of effective anticancer T cell immunity and underpin the clinical success of immune checkpoint blockade and personalized cancer vaccines. Advances in high-throughput sequencing, immunopeptidomics, and artificial intelligence (AI) have transformed neoantigen discovery from tailored experimental workflows into scalable, computational pipelines. However, accurately identifying the small subset of tumor mutations that yield processed, presented, and immunogenic epitopes remains a major bottleneck.
Methods
This review summarizes how AI is reshaping neoantigen discovery, from somatic variant calling, HLA typing, and peptide processing to peptide–MHC binding, presentation, and T cell recognition. We first outline the immunobiological foundations of antigen presentation, emphasizing class I and II peptide-binding grooves and their allele-specific motifs, then describe AI workflows that integrate somatic mutation calling, HLA typing, transcriptomics, and immunopeptidomics to nominate candidate neoepitopes. We highlight recent AI-driven tools for presentation and immunogenicity prediction, integrative pipelines that support personal and shared neoantigen targeting, and early clinical applications in vaccination and T cell therapies.
Results
AI-driven models trained on eluted ligand datasets substantially outperform affinity-only predictors for peptide presentation across diverse HLA alleles and populations. Consortium-scale benchmarking demonstrates that integrating features of antigen processing, presentation, and TCR recognition can eliminate the majority of non-immunogenic candidates while retaining clinically relevant neoepitopes. Immunopeptidomics provides essential ground truth, revealing that only a small fraction of genomically predicted candidates are naturally presented and uncovering noncanonical antigen sources, including splice variants, post-translational modifications, and noncoding regions. Integrative pipelines now support both personal (private) and shared (public) neoantigen prioritization, enabling translational applications such as personalized vaccines and TCR-based therapies.
Conclusions
AI-guided neoantigen discovery is now clinically actionable, enabled by immunopeptidomics and deep learning models. Despite significant progress, key challenges remain, including limited class II prediction accuracy, incomplete coverage of rare HLA alleles, tumor heterogeneity, and the need for standardized benchmarking and validation. Anchoring computational predictions to mass spectrometry–derived ligands and incorporating tumor evolution and immune escape mechanisms will be critical for improving target selection. Continued integration of AI, proteogenomics, and clinical data is poised to accelerate the development of effective, precision neoantigen-based cancer immunotherapies.
Keywords: Artificial intelligence, Cancer, Immunotherapy, MHC, Neoantigen
Introduction
Over the past decade, cancer immunotherapy has established T cells as central effectors in the control of diverse human malignancies [1–3]. Many effective T cell responses are directed against neoantigens, novel peptide sequences created by tumor-specific mutations that allow the immune system to distinguish cancer cells from normal tissues [1]. Recent technological innovations permit systematic characterization of patient-specific neoantigen responses, and emerging data suggest that recognition of such neoantigens is a major factor in the activity of clinical immunotherapies [1]. These observations position neoantigen load as a candidate biomarker for response and provide a compelling rationale to develop therapeutic strategies that selectively enhance T cell reactivity against this class of antigens [1, 4, 5].
Neoantigens arise from somatic alterations such as point mutations, insertions-deletions, gene fusions or aberrant splicing, and because they are absent from the normal proteome, they offer highly tumor-selective targets for immunotherapy [1, 4, 6]. Proof-of-concept trials of individualized neoantigen vaccines in melanoma demonstrated that RNA mutanome vaccines can be manufactured based on each patient’s tumor mutational repertoire, elicit T cell responses against multiple vaccine neo-epitopes and, in some cases, drive durable clinical responses when administered alone or with checkpoint blockade [4]. Broader clinical experience summarized in recent reviews indicates that personalized cancer vaccines can trigger broad-based antitumor immunity and are now being evaluated across a spectrum of solid tumors [7, 8]. In parallel, large-scale immunogenomic studies have shown that tumor mutational burden and predicted neoantigen load correlate with outcome to immune checkpoint inhibitors across multiple cancer types, underscoring the clinical relevance of neoantigen-directed immunity [1, 5].
Realizing the therapeutic promise of neoantigens requires accurately identifying, among thousands of tumor mutations, the small subset that yields processed, presented, and T cell–recognizable epitopes [9–12]. The accurate identification and prioritization of antigenic peptides have emerged as a central bottleneck for the development of personalized cancer immunotherapies [9]. Contemporary neoantigen discovery pipelines therefore integrate whole-exome or whole-genome sequencing of matched tumor-normal samples, transcriptomics, high-resolution HLA typing, in silico modeling of antigen processing and peptide–MHC (pMHC) presentation, and experimental validation using patient-derived T cells [9–12]. Complementing genomic approaches, mass spectrometry (MS)-based immunopeptidomics now enables direct profiling of HLA ligand landscapes in primary tumors, uncovering both canonical and non-canonical antigen sources and providing high-confidence ligands that serve as ground truth for model training [9–11].
Machine learning (ML) and deep learning (DL) models now underpin key steps in neoantigen discovery, including prediction of peptide–MHC binding and presentation, often in a pan-allelic fashion that covers the extensive polymorphism of human HLA [11, 12]. Models trained directly on large immunopeptidome datasets significantly improve class I epitope prediction across diverse alleles and human populations [11]. Consortium-scale benchmarking by the Tumor Neoantigen Selection Alliance (TESLA) further demonstrated that integrating peptide features associated with antigen processing and T cell recognition can filter out the vast majority of non-immunogenic peptides while retaining a substantial fraction of true targets [6, 12]. At the same time, TESLA and subsequent analyses highlight that most candidate peptides prioritized by current algorithms are non-immunogenic in functional assays, emphasizing persistent gaps in our understanding and modeling of antigen processing, presentation and TCR recognition [6, 12].
AI is therefore central to scaling neoantigen discovery from non-standardized experimental efforts to reproducible, clinically deployable pipelines. In this review, we first discuss the structure of major histocompatibility complex (MHC) molecules and how polymorphic pockets shape peptide binding and T cell recognition, then outline current neoantigen discovery workflows from somatic variant calling through integrative pipelines and immunopeptidomics, emphasizing where AI tools are already reshaping practice. Finally, we examine emerging strategies for prioritizing personal versus shared neoantigens, summarize the evolving clinical landscape of neoantigen-targeted therapies, and highlight outstanding challenges and future directions for deploying AI-guided neoantigen discovery in precision oncology.
We conducted a comprehensive and systematic search to identify relevant research articles using the keywords “neoantigen/neoepitope identification” and “artificial intelligence.” The primary objective of this study was to collect and compare AI-based methods for neoantigen prediction across multiple stages, ranging from somatic mutation detection to the generation of immune responses through the recognition of neoantigen–MHC complexes by T cell receptors (TCRs). Articles were not selected based on publication year, as the aim was to evaluate and compare both earlier and more recent models in terms of their modeling approaches and predictive performance. Studies published in languages other than English were excluded. Furthermore, only studies employing AI-based techniques, including ML and DL, and applying these methods across various cancer types were included to enable a comparative assessment of model performance and applicability across different tumor contexts.
The major histocompatibility complex
MHC molecules are divided into two main groups: class I and class II. Class I MHC molecules are present on the surface of all nucleated cells. Conversely, MHC class II molecules present peptides primarily on APCs such as dendritic cells, macrophages, and B cells [13, 14]. In order to perform their physiological role, MHC proteins must initially bind with peptide antigens [14]. Peptides that occupy the peptide-binding groove of surface-expressed MHC molecules are generated through specific proteolytic processes [15]. MHC class I typically presents endogenous (cytosolic) peptides to CD8+ cytotoxic T cells, whereas MHC class II typically presents exogenous (extracellular) peptides to CD4+ helper T cells [14, 16]. The pathways involved in antigen processing and presentation for both MHC class I and class II are outlined in Fig. 1.
Fig. 1.

MHC class I and class II molecules utilize different antigen processing pathways. a) MHC class I molecules display peptides that mainly originate from proteins produced inside the cell, whether from the host or invading pathogens. These proteins are broken down into peptides by the proteasome, after which the peptides are transported into the endoplasmic reticulum via TAP (transporters associated with antigen processing) to be loaded onto MHC class I molecules. b) Unlike MHC class I, MHC class II molecules present proteins that enter the cell via endocytosis. During their maturation, they are prevented from binding self-peptides in the endoplasmic reticulum by the invariant chain (Ii). These MHC II–Ii complexes pass through the Golgi and reach the late endosome, where the invariant chain is degraded into a smaller fragment called CLIP (class II–associated invariant chain peptide). CLIP is then displaced from the MHC–CLIP complex and substituted with an antigenic peptide. c) Dendritic cells possess the ability to internalize antigens from other cells and subsequently cross-present them to CD8+ cytotoxic T lymphocytes. The reliance of this process on TAP suggests that cross-presentation requires the redirection of cellular antigens into the classical MHC class I pathway, although the precise mechanisms underlying this redirection remain unclear. Typically, these antigens are also processed through the MHC class II pathway, enabling recognition by CD4+ helper T cells
Decoding MHC class I: Insights into polymorphic pockets
MHC class I molecules consist of two parts: a heavy (α) chain paired noncovalently with β2-microglobulin. The α1 and α2 domains of the heavy chain come together to create the peptide-binding groove [17]. This groove is characterized by its closed-ended structure, which accommodates peptides typically 8–11 amino acids in length, though the precise length can vary based on MHC allele and antigen characteristics [18].
The specificity of peptide binding is primarily governed by anchor residues located at specific positions within the peptide sequence, such as P2 and P9. These anchor residues interact with complementary pockets in the MHC class I groove, contributing significantly to the stability and affinity of the peptide-MHC complex [13]. Insights into peptide binding preferences of specific MHC class I alleles facilitate the design of peptide vaccines aimed at eliciting robust CD8+ T cell responses against viral infections or tumor antigens [19]. Moreover, these insights inform strategies for enhancing antigen presentation efficiency in immunotherapies targeting cancer and other diseases [14].
Polymorphism in class I molecules is concentrated in the α1 and α2 domains that form the peptide-binding groove; this polymorphism sculpts pocket architecture and determines allele-specific peptide motifs. Among the classical loci, HLA-B is the most polymorphic and contributes disproportionately to population peptide-presentation diversity, although HLA-A and HLA-C are also highly variable. These allelic differences give rise to supertypes (groups of alleles with overlapping peptide specificities) but also to highly allele-specific preferences that shape T cell responses and population immunity [20].
Individual MHC-I proteins balance allele-specific anchors (P2, PΩ and sometimes P1/P3/P5) with broader tolerance at other peptide positions. Class I motifs often rely on two or three dominant anchors, producing narrower length restrictions than class II but still permitting many sequence variants to bind the same allele. Subdominant (secondary) anchor residues and peptide flanking residues can modulate affinity and TCR recognition but generally have less impact on primary MHC binding than the canonical anchors [21].
In TCR recognition of peptide-MHC I complexes, the CDR2 loops of both alpha and beta chains make contact with only the MHC molecules, whereas the CDR1 and CDR3 loops make contact with both peptide and MHC. Additionally, the V domains of alpha and beta chains are adjacent to the N-terminal and C-terminal of the bound peptide, respectively. The binding region on TCR is almost flat and sometimes has a hole; in contrast, MHC surfaces have two apexes, and the TCR binds between them in interaction with peptide-MHC. Auxiliary systems such as CD3 and CD8 co-receptors on T cells generate a signal to the T cell during peptide-MHC I recognition, which causes T cell signaling and triggers a cytotoxic response and then kills the virally infected or otherwise abnormal cell [17, 18].
Decoding MHC class II: Insights into polymorphic pockets
Antigens are internalized into cells through processes like phagocytosis, pinocytosis, endocytosis, and autophagy, and are delivered to late endosomes. Within these endosomes, they are processed by proteolytic enzymes such as cathepsins and thiol oxidoreductase. MHC-II molecules are loaded with peptides in a specialized late endosomal–lysosomal antigen-processing compartment [22].
MHC-II molecules are dimers composed of an alpha (α) and a beta (β) polypeptide chain. These glycoproteins have an immunoglobulin-like domain (α2 and β2) close to the cell membrane, as well as a peptide-binding cleft positioned farther from the membrane (α1 and β1), which includes the most polymorphic sites of HLA class II molecules. This peptide binding region (PBR) is essential for the molecule’s function [23]. The peptide-binding groove of MHC class II is structured with two alpha-helices above a beta-pleated sheet, designed to specifically bind short peptides. The amino acid sequence around the binding site is highly variable, enhancing the molecule’s ability to bind a wide range of peptides. The ends of the antigen-binding cleft of class II molecules contain small residues, such as glycine or valine, which do not restrict the size of the peptides. Thus, unlike MHC-I, MHC-II molecules possess an open-ended binding groove that can fit longer peptides, generally 13–18 amino acids in length, and occasionally up to 30 amino acids [23, 24].
The specificity of peptide binding in MHC-II molecules is influenced by the interaction between anchor residues (side chains of amino acids) of the peptides and the pockets within the MHC grooves. These anchor residues form electrostatic, van der Waals, and hydrogen bonds with the pockets, stabilizing the peptide-MHC complex. Conserved residues in the MHC-II molecule form hydrogen bonds with the peptide’s backbone, whereas the variable residues create specialized pockets that accommodate the peptide’s side chains. Some peptides have charged residues that interact with the alpha helix. Antigenic peptides that successfully bind to the floor of the MHC class II groove have a specific conserved secondary structure resembling a polyproline chain. This polyproline structure is characterized by its openness and lack of internal hydrogen bonding [25].
MHC-II molecules exhibit broad specificity, facilitated by five cavities that accommodate the side chains of the attached peptide molecules. These pockets vary in size and hydrophobicity, allowing MHC-II molecules to bind a diverse array of peptides. Within the binding groove of MHC class II molecules, there are typically three to four critical anchor pockets, including positions like P1, P4, P6, P7, and occasionally P9, whereas MHC class I molecules have only two major pockets. Secondary or minor anchor residues influence peptide-MHC binding affinity (BA), but their effect is comparatively subtle. Each of these pockets, usually located in beta strands, in MHC class II plays a distinct role in peptide binding and allele specificity. MHC class II proteins generally bind a narrower range of proteins compared to MHC class I proteins. This is because MHC class II molecules have more major anchor pockets, which increases the specificity for peptide binding [26, 27]. The P1 pocket is crucial for anchoring peptides into the HLA class II cleft. It typically accommodates large, hydrophobic residues. P6 and P7 pockets are involved in determining the allele specificity of MHC-II molecules. The P4 and P9 pockets play a dual role by determining antigen specificity and influencing how the antigenic peptide interacts with the groove. The P4 and P7 pockets are composed exclusively of the β chain, whereas P1, P6, and P9 are shaped by contributions from both the α and β chains [26, 27].
Polymorphic amino acids are grouped within the peptide-binding groove of MHC-II molecules, accounting for the allele-specific nature of peptide-binding motifs. This results in strong preferences for certain side chains at specific positions in the peptide, while other positions can accommodate a wide variety of side chains. The primary anchor residues differ among alleles, for example: P1 in HLA-DR1 and P4 in HLA-DR3, while P7 and P9 also contribute to binding and may add extra stabilization. Only a few peptide residues, such as the side chains at P2, P5, and P8, extend outward from the groove to interact with T cell receptors. The amino acid identities at these positions generally do not influence the peptide’s binding affinity to MHC-II unless they undergo specific post-translational chemical modifications [27].
MHC-I and MHC-II molecules present peptides using both sequence-specific and sequence-independent features, reflecting allele-specific and shared properties, respectively. Sequence-dependent interactions are key for determining which peptides fit into the binding groove [23]. A network of sequence-independent hydrogen bonds connects the peptide backbone to conserved residues of MHC-II, extending along the entire peptide. This differs from MHC-I, where conserved hydrogen bonds mainly involve the peptide’s N- and C-terminal ends [25].
Analysis of the peptide-binding region of HLA class II molecules reveals a concentration of polymorphism activity in PBR, specifically in pockets. Among the HLA class II genes, HLA-DRB1 is the most variable. The α-chain of HLA-DR remains constant, which allows the β-chain to undergo more variation. The binding process begins with hydrophobic side chains fitting into the P1 pocket. Amino acids occupying the P6 pocket also strongly affect preferences at P9. Allelic differences can involve up to 56 amino acids, with most of this variation localized in the peptide-binding region [27]. It should be noted that each of the pocket variants has different amino acid sequences which affect their specificity for the peptide residues.
Post-translational modifications (PTMs) like glycosylation, phosphorylation, and acylation are crucial for regulating protein function in numerous physiological processes. These modified epitopes differ in immunogenicity from unmodified peptides. PTMs can serve as anchor points when peptides are presented by MHC molecules. They can also affect protein processing, thereby altering peptide presentation. When interacting with specific MHC molecules, modified peptides display unique characteristics that impact their immunogenicity. For example, in HLA-DR1, phosphate groups engage the peptide-binding groove through water-mediated interactions. Phosphorylated peptides restricted by MHC-II can also modify peptide structure and influence T cell receptor recognition [23].
Studies have shown that peptide flanking residues (PFRs), which lie outside the main binding groove, can affect both peptide-MHC binding (in a sequence-independent way) and T cell recognition. Extended peptides that are shorter than optimal can enhance their binding to MHC class II molecules. In particular, the residue at position P-1 can increase measurable peptide affinity for the MHC groove. The ideal peptide length for MHC class II binding is around 18–20 amino acids, while adding residues beyond this length tends to reduce or abolish binding [24].
Understanding interactions within pMHC complexes and the roles of specific pockets in the MHC-II groove is crucial for developing targeted immunotherapies and vaccines, thereby enhancing our ability to combat infectious diseases and autoimmune conditions. Each population has its own restricted alleles, which are crucial for designing vaccines and immunotherapies. Understanding the specificities and polymorphisms of MHC-II molecules can aid in developing targeted treatments that leverage the immune system’s natural capabilities. By identifying the distinct peptide-binding motifs and anchor pockets of MHC-II molecules, researchers can design more effective vaccines and therapies that elicit robust immune responses [27].
Neoantigen identification workflow
Neoantigens, a subset of tumor-specific antigens (TSAs), are unique, non-self-peptides generated by tumor-acquired genetic alterations and presented on MHC molecules on tumor cells that can be recognized by immune cells [17, 18]. In addition, antigen-presenting cells (APCs) can acquire tumor-derived proteins and process them into peptide fragments for loading onto MHC molecules [13]. The complexes of tumor-specific pMHC are identified by cognate T cell receptors. Such neoantigen-specific T cells can elicit antitumor immune responses in patients [19]. Because neoantigens are not presented during thymic selection, T cells specific for them typically escape central tolerance, enabling a neoantigen-specific repertoire with potentially high functional avidity [28, 29].
Neoantigen discovery pipelines can be divided into several phases. (i) Matched tumor–normal DNA sequencing profiles the somatic variant landscape. (ii) Tumor RNA sequencing (RNA-seq) quantifies gene- and allele-specific expression and reveals transcript-derived alterations, including aberrant splicing, gene fusions, intron retention, and RNA editing. (iii) High-resolution HLA class I/II genotyping, together with tumor HLA status (e.g., loss of heterozygosity, downregulation, or allele-specific expression), constrains the feasible set of presenting molecules. (iv) Antigen processing and presentation are modeled computationally, encompassing proteasomal cleavage to nominate peptides that can be generated from mutant proteins, TAP-mediated transport to estimate endoplasmic reticulum import, and pMHC binding and complex stability to prioritize epitopes with a high likelihood of cell-surface display. (v) Putative immunogenicity is estimated using TCR recognition models that integrate epitope foreignness, self-similarity, TCR-facing residue features, repertoire evidence, and tumor immune context. (vi) Finally, immunopeptidomics (MS-based HLA ligandome profiling) provides empirical observations of presented peptides, enabling refinement and validation of presentation rules and calibration of in silico predictors. In the following sections, we detail each step.
Somatic mutation calling
The body consists of somatic and germline cells. Mutations that occur in somatic cells throughout a person’s lifetime affect only that individual. In contrast, mutations in germline cells can be passed to subsequent generations, serving as the basis for species evolution and leading to hereditary diseases [30, 31]. Detection of variant alleles by using next-generation sequencing (NGS) technology is a common method for identifying tumor neoantigens. Whole-exome sequencing is conducted on tumor tissue and corresponding normal tissue to extract tumor-specific mutations, and the expression level of these mutations is determined by integrating RNA sequencing data [32–34]. The generation of tumor neoantigens originates from somatic genomic mutations, including single-nucleotide variants (SNVs), insertions and deletions (INDELs), fusion genes [35], and splice variants [34].
Two principal strategies have been established for the identification of neoantigen epitopes. The immunogenomic approach employs NGS data to computationally generate virtual peptidomes through in silico prediction methods, whereas the immunopeptidomic strategy utilizes MS to directly characterize peptides presented by MHC molecules [36]. In addition, several TCR-guided neoantigen discovery methods have recently emerged, enabling the systematic mapping and validation of immunogenic neoantigens [19].
Immunogenomic research has advanced rapidly through the application of NGS to compare genetic alterations between tumor and matched normal tissues. The first critical step in identifying potential neoantigens from NGS data typically involves detecting tumor-specific genomic aberrations using whole-exome sequencing (WES) of paired tumor and normal DNA. Integration of RNA sequencing with WES enables assessment of whether mutant alleles are transcriptionally expressed within the tumor. Moreover, RNA-seq provides additional layers of biological insight, including information on copy number variations, microbial contamination, transposable element activity, cellular composition, and the presence of candidate neoantigens [37, 38].
While immunogenomic analyses predict millions of mutation-derived neoantigens, most are not confirmed at the protein level [39, 40]. MS-based immunopeptidomics serves as the gold standard for validating neoantigens by directly identifying MHC-bound peptides, including those arising from post-translational modifications [41–46], non-coding RNA, or proteasome splicing features often missed by DNA- or RNA-only approaches [47–49]. Integrating MS with NGS and developing user-friendly tools that combine genomic, transcriptomic, and proteomic data are key to improving neoantigen discovery for cancer immunotherapy.
The process of mutation calling starts with quality control of the sequencing reads and aligning them to a reference genome using FastQC screen [50] and BWA [51], respectively. After alignment, the mutation calling step must carefully differentiate true somatic variants from sequencing errors, artifacts introduced during sample preparation, and germline mutations. Numerous software tools have been developed to tackle the key challenges associated with this process. A wide variety of somatic mutation caller tools are listed in Table 1.
Table 1.
Somatic variant calling tools
| Model/Year | Algorithm | Mutation Type | Key Features/Web and Code Accessibility | Ref. |
|---|---|---|---|---|
|
samtools/bcftools 2009 |
Compressed block indexing algorithm (BGZF) combined with hierarchical binning indexing | Basic variant calling |
Versatile in structure, space-efficient, optimized for rapid random access, and used as the standard format for releasing alignments from the 1000 Genomes Project. |
[52] |
|
GATK 2011 |
Haplotype analysis |
Multi-step process, including a germline genotyper |
Unified analytic framework for variation discovery and three-phase conceptual pipeline including preprocessing (read mapping, duplicate removal, local realignment, base quality score recalibration), variant discovery, variant refinement. |
[53] |
|
VarScan 2 2012 |
Heuristic and statistical classification |
Somatic SNVs and INDELs, germline variants, copy number variants, loss of heterozygosity (LOH) |
Reads data from tumor–normal samples simultaneously. For germline variants, it showed high concordance with SNP arrays (99.56%). For somatic mutations, it achieved high sensitivity (94.7%) and a strong true-positive rate (89%). |
[54] |
|
SomaticSniper 2012 |
Bayesian genotype contrast | Somatic SNVs |
Direct tumor-normal comparison. Identifies systematic errors and introduces empirical/statistical filters. |
[55] |
|
FreeBayes 2012 |
Bayesian haplotype-based | SNPs, INDELs, and MNPs |
Overcomes the problem of sequences aligning to multiple genomic locations. |
[56] |
|
Lofreq 2012 |
Poisson-binomial model-based | Somatic SNVs |
More sensitive than ad hoc or model-based methods, achieving higher sensitivity without loss of specificity. Validated to detect rare variants down to 0.05% frequency. |
[57] |
|
EBCall 2013 |
Empirical Bayesian framework |
Somatic SNVs, INDELs |
Effectively detects low-frequency somatic mutations (<10%). |
[58] |
|
Shimmer 2013 |
Probabilistic, hypothesis-driven statistical method | Somatic SNVs |
Highly effective on heterogeneous or contaminated tumor samples. |
[59] |
|
Seurat 2013 |
Bayesian analysis with beta-binomial distributions | Somatic SNVs, small INDELs, LOH, and SVs |
Extends to RNA-seq data for detecting allelic imbalance events in annotated transcripts. |
[60] |
|
Virmid 2013 |
Bayesian inference with the estimated joint genotype probability matrix |
Somatic SNVs |
Outperforms other tools in detecting somatic mutations in highly contaminated samples. |
[61] |
|
RADIA 2014 |
Heuristic and integrative algorithm | Somatic SNVs from DNA and matched RNA |
Integration of RNA and DNA. High accuracy. Simulation-based validation. |
[62] |
|
Platypus 2014 |
mapping + assembly + haplotype-based Bayesian | SNP, INDELs, and complex polymorphisms |
High sensitivity and specificity for SNPs. |
[63] |
|
SMUFIN 2014 |
Quaternary sequence tree-based, direct read comparison algorithm | Somatic SNVs, INDELs, and SVs |
Defines complex chromosomal rearrangements (chromoplexy, chromothripsis) at base-pair resolution. |
[64] |
|
Abra 2014 |
Assembly-based realigner | Somatic INDELs |
Employs a fast and adaptable localized de novo assembly, followed by global realignment, to improve the accuracy of read mapping. |
[65] |
|
cgpPindel 2015 |
Split-read mapping algorithm | Somatic INDELs |
Optimized for somatic indels. Scalable execution. Uses post-hoc filtering. |
[66] |
|
CaVEMan 2016 |
Expectation-Maximization (EM) | Somatic SNVs |
Produces high-quality somatic substitution calls with high recall and positive predictive value. Uses post-hoc filtering. |
[67] |
|
MuSE 2016 |
Markov substitution model | Somatic SNVs |
Builds a sample-specific error model and sets tiered cutoffs to balance sensitivity and specificity. |
[68] |
|
VarDict 2016 |
Local realignment and consensus building | Somatic SNV, MNV, INDELs, germline variants, LOH, complex variants, and SVs |
Ultra-deep sequencing scalability. Supports amplicon-aware variant calling and can detect PCR artifacts. |
[69] |
|
Scalpel 2016 |
Microassembly using self-tuning de Bruijn graphs | Somatic INDELs |
Highly accurate in repetitive regions, supporting single-sample, de novo, and somatic analysis modes. |
[70] |
|
Strelka2 2018 |
Bayesian model for continuous allele frequencies | Somatic SNVs and small INDELs, germline variants |
Estimation of insertion/deletion error parameters for each sample using a mixture model. Uses an enhanced somatic variant model that corrects for normal-sample contamination to improve detection in liquid tumor data. |
[71] |
|
SvABA 2018 |
Assembly-based | Somatic INDELs and SVs |
Shows broad sensitivity for indels and structural variants (SVs), especially 20–300 bp variants and complex rearrangements (viral integrations). Combines assembly and alignment signals. Low computational burden and is suitable for large-scale data. |
[72] |
|
Lancet 2018 |
Local-assembly-based using colored de Bruijn graphs to jointly analyze tumor and normal reads | Somatic SNVs, INDELs |
High precision and robust quality scoring for prioritizing somatic variants. |
[73] |
|
TNScope 2018 |
Haplotype-based | Somatic SNVs, INDELs, SVs |
Hybrid approach combining haplotype-based variant detection with ML filtration. |
[74] |
|
MuTect 2 2019 |
Local assembly of haplotypes and pair-HMM Read-to-Haplotype realignment | Somatic SNVs and INDELs |
Employs probabilistic models for genotyping and filtering that perform effectively across all sequencing depths and can operate with or without a matched normal sample. https://github.com/broadinstitute/gatk/releases/download/4.2.2.0/gatk-4.2.2.0.zip |
[75] |
|
NeuSomatic 2019 |
CNN |
First DL-based somatic SNV detection |
Tumor/normal alignment information summarized into matrices capturing genomic context. Learns deep feature representations directly from raw data. |
[76] |
|
DNN-Boost 2021 |
Ensemble of Deep Neural Network (DNN) and XGBoost | Somatic variants |
Somatic mutation identification of tumor-only whole-exome sequencing data. Distinguishes somatic mutations from germline variants. Shows consistent accuracy and F1-score improvements on external datasets. |
[77] |
|
VarNet 2022 |
CNN | Somatic SNVs, INDELs |
Outperforms existing tools and even the ensemble method. Learns effectively from weakly labeled datasets. |
[78] |
|
Dragen 2023 |
Haplotype-based | Somatic SNVs, INDELs |
Flexible architecture. Built-in noise models. Joint tumor–normal analysis. FPGA acceleration. |
[79] |
|
DeepSom 2023 |
CNN | Somatic SNVs, INDELs |
Tumor-only somatic variant caller. Combines Mutect2 candidate generation, gnomAD-based filtering, and DL classification. |
[80] |
|
DeepSomatic 2024 |
CNN | Somatic SNVs, INDELs |
Outperforms existing tools across Illumina and long-read platforms. Adaptable models for tumor-only, WES, and FFPE data. |
[81] |
|
OncoTOP 2024 |
Combines realDcaller2 and Mutect2 | Somatic SNVs, INDELs |
Tumor-only somatic variant caller. Combines realDcaller2 and Mutect2. Predicts their germline or somatic origins, and evaluates clinically relevant biomarkers. NA |
[82] |
|
VarNet-T 2025 |
CNN | Somatic SNVs, INDELs |
Novel method for tumor-only somatic variant calling. Offers improved accuracy for clinical applications such as assessing tumor mutational burden (TMB) for immunotherapy and DNA mismatch repair (MMR) deficiency for PARP inhibitor therapies. NA |
[83] |
|
ClairS-TO 2025 |
Ensemble of two (Affirmative and Negational) neural networks | Somatic SNVs, INDELs |
Long-read tumor-only somatic variant caller. Pre- and post-filtering enhancements. Addresses limitations of real tumor data. |
[84] |
SNV: Single Nucleotide Variant. Indel: Insertion and Deletion. MNP: Multi-Nucleotide Polymorphism. HMM: Hidden Markov Model. SV: Structural Variant
Community reference tumor–normal DNA pairs and curated call sets from SEQC2 enable standardized evaluation of WGS/WES somatic callers across platforms and sites, with raw data and truth sets openly available [85]. In parallel, the ICGC–TCGA DREAM Somatic Mutation Calling Challenge provides simulated tumor genomes and crowd-sourced leaderboards to compare callers under controlled conditions [86]. Analysis of The Cancer Genome Atlas (TCGA) database has facilitated the characterization of 933,954 expressed neoantigens across 20 distinct solid tumor types. These neoantigens are derived from 893,960 somatic mutations, exhibiting a variable median frequency across different cancer types. A limited subset of neoantigens, specifically 24, including those originating from mutations within key driver genes such as PIK3CA, RAS, and BRAF, is observed to be shared among a minimum of 5% of patients across various cancer types or within the same cancer [47, 87]. The distribution of mutation types and sample sizes across these cancer types is summarized in Fig. 2.
Fig. 2.

Distribution of mutation types across different cancer types from the Cancer Genome Atlas (TCGA). The radial bar plot displays the number of distinct mutations identified per cancer type (indicated on the periphery, with sample sizes in parentheses). Bars are segmented by mutation type, including missense, frameshift, in-frame deletion, in-frame insertion, and splice variants, as indicated in the legend. SKCM, skin cutaneous melanoma; LUAD, lung adenocarcinoma; LUSC, lung squamous cell carcinoma; BLCA, bladder urothelial carcinoma; UCEC, uterine corpus endometrial carcinoma; COAD, colon adenocarcinoma; STAD, stomach adenocarcinoma; DLBC, lymphoid neoplasm diffuse large B cell lymphoma; HNSC, head and neck squamous cell carcinoma; PAAD, pancreatic adenocarcinoma; CESC, cervical and endocervical cancers; ACC, adrenocortical carcinoma; READ, rectum adenocarcinoma; LIHC, liver hepatocellular carcinoma; SARC, sarcoma; UCS, uterine carcinosarcoma; TGCT, testicular germ cell tumor; CHOL, cholangiocarcinoma; BRCA, breast invasive carcinoma; GBM, glioblastoma multiforme; KICH, kidney chromophobe; THYM, thymoma; UVM, uveal melanoma; LGG, brain lower-grade glioma; PRAD, prostate adenocarcinoma; KIRP, kidney renal papillary cell carcinoma; THCA, thyroid carcinoma; PCPG, pheochromocytoma and paraganglioma; KIRC, kidney renal clear-cell carcinoma; MESO, mesothelioma
Recommendations: A wide range of somatic variant calling tools is available, and their selection should be guided by study design, data characteristics, and analytical goals. For tumor–normal paired sequencing, established methods such as MuTect2 (most widely used) and Strelka2 (time-efficient) are generally preferred due to their improved performance [88], whereas tumor-only analyses benefit from specialized approaches such as DeepSom, VarNet-T, and ClairS-TO that incorporate advanced filtering or learning strategies to distinguish somatic from germline variants. The choice of tool should also depend on the mutation type of interest: for somatic SNV detection, commonly used methods include LoFreq, MuSE, SomaticSniper, and Lancet, whereas INDEL detection is often better supported by tools such as VarScan 2 and Pindel. In addition, MuTect2 and Strelka provide robust performance across both SNVs and INDELs [89]. Whereas tools like SvABA and SMUFIN are better suited for structural variants and complex rearrangements. For detecting low-frequency or subclonal mutations, highly sensitive methods such as LoFreq and EBCall are advantageous, whereas tools like VarDict perform well on ultra-deep or amplicon sequencing data. In samples with significant tumor heterogeneity or contamination, probabilistic approaches such as Virmid and Shimmer may offer improved robustness. When multi-omics data are available, integrative tools such as RADIA and Seurat can enhance variant confidence by incorporating RNA evidence. More broadly, conventional statistical and Bayesian frameworks, including GATK and FreeBayes, remain widely used due to their interpretability and robustness, while DL-based methods such as NeuSomatic, VarNet, and DeepSomatic have demonstrated improved performance in complex or noisy datasets, albeit with higher computational demands. Finally, no single tool is universally optimal; therefore, combining multiple callers or adopting ensemble strategies is often recommended to improve accuracy and reduce false-positive rates.
HLA typing
HLA typing is foundational to neoantigen identification because peptide–MHC binding must be predicted in each patient’s autologous HLA context, enabling recognition of tumor-specific “non-self” peptides that drive effective immunotherapy responses [1, 90]. Standard pipelines integrate tumor/normal exome (±RNA-seq) with high-resolution HLA genotypes to nominate mutant peptides predicted to bind class I/II molecules, an approach validated by MS-guided discovery and vaccination readouts in preclinical models [91]. These typed, patient-specific predictions underpin personalized vaccines; in melanoma, a multi-peptide vaccine targeting up to 20 predicted neoantigens was feasible, safe and immunogenic, inducing polyfunctional CD4+/CD8+ responses with durable disease control in several patients [4, 90]. HLA-aware analyses of tumor evolution also refine target selection; allele-specific HLA loss of heterozygosity occurs in ~40% of NSCLC and biases binding toward the lost allele, so jointly assessing HLA typing and HLA LOH helps avoid non-presented targets and improves prioritization [92]. Collectively, accurate HLA genotyping, paired with sequencing, binding prediction and immunopeptidomics, enables precise neoantigen nomination, guides vaccine and T cell therapy design, and supports biomarker discovery linking neoantigen burden and clinical response [1, 90].
The HLA region is hyperpolymorphic; for example, the IPD-IMGT/HLA database now catalogs > 35,000 curated alleles and serves as the official WHO repository with regular quarterly releases [93]. This database is the authoritative, highly curated repository for HLA allele sequences and nomenclature and underpins benchmark studies [91]. Short-read HLA typing tools are typically assessed against this reference and gold-standard methods; for example, xHLA achieves 99–100% four-digit accuracy across class I and II loci on WGS data within minutes, and graph-based HLA*LA attains high accuracy on exome and low-coverage WGS [94, 95].
Inferring HLA alleles from standard NGS is challenging because extensive polymorphism and inter-gene homology complicate read mapping [96]. Alignment-based tools achieve high accuracy on short reads; OptiType reports ~97% overall accuracy from unenriched WGS/WES/RNA data [97]. Graph-aware methods model allelic diversity explicitly, for example HLA*PRG attains accuracies comparable to sequence-based typing (SBT) on high-quality WGS [96], and HLA*LA extends graph alignment to both short- and long-read data, reporting ~99% on WGS, ~93% on exome, and ~98% on long-read datasets [95]. A broader graph framework, HISAT-genotype, integrates > 14 M variants and includes an HLA module that matches or exceeds laboratory assays on benchmarks [98].
Data type influences performance and reporting resolution. On exomes, accuracy is typically lower and coverage less uniform than on WGS; prior work recommends high fragment lengths and ≥30× depth for robust calls [96]. Consistent with this, HLA*LA reports ~93% mean accuracy on exome versus ~99% on WGS and supports long-read inputs (~98%), facilitating full-length allele inference when needed [95]. For RNA-seq, ArcasHLA attains 100% two-field accuracy for class I and >99.7% for class II on benchmark sets, although post-transcriptional effects and amplification bias can complicate inference [99]. Assembly-based approaches such as Kourami enable discovery of novel alleles from high-coverage WGS but show reduced performance at lower coverage and on WES [100]. In very large cohorts with standardized exome capture, direct exome-based calling at multi-field resolution (e.g., HLA-HD in UK Biobank WES) scales effectively and increases power for association studies, though some loci (e.g., DQA1) exhibit lower coverage [101]. Representative HLA typing tools, their input data types, and performance characteristics are listed in Table 2.
Table 2.
HLA typing tools
| Model/Year | Architecture | Input Data | Performance/MHC Class/Key Features/Code Accessibility | Ref. |
|---|---|---|---|---|
|
HLAminer (2012) |
Alignment-based and targeted assembly | NGS reads (whole genome, whole transcriptome/RNA-Seq, exome) |
Direct-alignment mode reached ~80% sens. / 78% spec. Class I Does not require specialized HLA enrichment, compatible with various sequencing data. |
[102] |
|
seq2HLA 2012 |
Alignment-based read-mapping and statistical inference | RNA-seq |
NA Class I & II Determines locus-specific expression level. Provides a confidence (p-value) for each HLA call. http://tron-mainz.de/tron-facilities/computational-medicine/seq2HLA/ |
[103] |
|
SOAP-HLA 2013 |
Alignment-based haplotype assembly | Aligned BAMs, WGS |
NA Class I & II Integrated one-step pipeline for variant detection and HLA typing. High coverage, low bias. Long DNA fragments (~500 bp). |
[104] |
|
ATHLATES 2013 |
Assembly-based | WES |
Overall concordance rate of 99% Class I & II Assembly-based, not alignment-based, reconstructs exon sequences directly from reads; automated allele pair inference. |
[105] |
|
HLAforest 2013 |
Alignment-based hierarchical read weighting | RNA-seq |
NA Class I & II Suitable for longer-read data. Supports paired-end reads to maximize phasing information. |
[106] |
|
OptiType 2014 |
Integer Linear Programming (ILP)-based optimization |
RNA-Seq, Exome, WGS |
Accuracy: 97.1% (4-digit) Class I Uses exons 2 and 3 plus flanking introns as references, reconstructing missing introns using phylogenetic methods, fast, accurate. |
[97] |
|
POLYSOLVER 2015 |
Alignment-guided, model-based probabilistic | WES |
Accuracy: 97% Class I Model-based design; somatic mutation detection; custom reference construction. |
[107] |
|
HLAreporter 2015 |
De novo assembly-based approach | WES |
NA Class I & II Uses a comprehensive reference panel; zero mismatch tolerance for both assembly and allele matching ensures high reliability. |
[108] |
|
HLA-VBSeq 2015 |
Alignment-based and Variational Bayesian (VB) inference optimization | WGS |
Accuracy: 99.95% Class I & II No need for primer design for HLA loci; does not rely on allele frequency data or prior population information; achieves full (8-digit) HLA typing resolution. |
[109] |
|
HLA-HD 2017 |
Dictionary-based mapping + weighted read-count scoring | NGS data |
High-coverage WES data: 100%; Low-coverage: 91.0% Class I & II Unrestricted use of all exons for typing; considers variation inside and outside G-DOMAIN; supports up to 6-digit precision. |
[110] |
|
xHLA 2017 |
Translated-read strategy with exhaustive MSA expansion and iterative allele-set refinement | 30× WGS BAM |
99–100% four-digit for both class I and II Class I & II Minute-scale runtime on a desktop. |
[94] |
|
HLAScan 2017 |
Alignment-based mapping to IMGT/HLA with a read-distribution score function | WGS, WES, targeted sequencing |
96.9% overall Class I & II Read-distribution scoring to cut false positives; unique-read phasing; supports up to six-digit typing; exome and target-panel compatible. |
[111] |
|
HLAProfiler 2017 |
k-mer profiling (Kraken-based gene filtering; competitive pairwise scoring) | RNA-seq FASTQ |
>99% accuracy at two-field (biological and simulated) Class I & II Excels on rare/novel/partial alleles via allele-refinement (e.g., 68% of novel alleles with correct protein or exact CDS); can work with as few as ~1,000 filtered HLA reads; extendable to KIR genes. |
[112] |
|
PHLAT 2018 |
Alignment-based, probabilistic, and pairwise optimization | RNA-seq, WES, amplicon FASTQ |
NA Class I & II Supports a broad range of read lengths, coverage, and targeted amplicon sequencing data; can output supporting BAM format. |
[113] |
|
Kourami 2018 |
Graph-guided assembly using modified partial-order graphs (POGs) | High-coverage WGS |
>98% accuracy for known alleles Classical class I & II Novel allele discovery, fast, moderate memory. |
[100] |
|
HLA*LA (HLA*PRG successor) 2019 |
Graph-based: linear alignments projected onto a population reference graph (PRG) | WGS, WES, ONT/PacBio long reads, assemblies |
99% (WGS); 93% (WES); 98% for long-read WGS/targeted Classical class I & II Broad input support; graph model derived from PRG lineage. |
[95] |
|
HISAT-genotype 2019 |
Graph FM-index (HISAT2) + genotype genome; guided k-mer assembly | WGS (short reads) |
NA Class I & II Novel-allele discovery via guided k-mer assembly; can assemble full-length alleles (exons+introns) from typical WGS; fast and memory-efficient alignments thanks to graph. Exactly matches known alleles for six HLA genes in 17 Platinum Genomes; “matches or exceeds” lab assays. |
[98] |
|
ArcasHLA 2020 |
Alignment-based RNA-seq HLA typing that uses Kallisto pseudoalignment plus EM quantification | RNA-seq (paired-end; also, single-end tested) |
100% (Class I, two-field) and >99.7% (Class II) on 1000 G benchmark Class I & II Partial-allele typing option; population-specific priors; works with single-end or paired-end RNA-seq; integrates easily in BAM-based pipelines; fast; works on metatranscriptomes. |
[99] |
WES: Whole Exome Sequencing. WGS: Whole Genome Sequencing
Platform choice should reflect study aims. When full-gene phasing is essential, long-read sequencing can type complete alleles, but in one recent comparison of classical methods, PacBio long reads offered little protein-coding advantage over MiSeq, whereas MiSeq provided superior scalability and cost-effectiveness [114]. For tumor–normal cohorts, researchers additionally assessed allele-specific HLA loss of heterozygosity using LOHHLA, a copy-number tool that has revealed HLA LOH in ~40% of NSCLC and refines neoantigen prediction [92].
Recommendations: Tool selection depends on availability of data type (WES or RNA), dataset size, sequencing depth, and computational capacity. For instance, RNA-seq–based tools such as seq2HLA, HLAforest, HLAProfiler, and ArcasHLA are well suited for studies that also require expression quantification, whereas WES or WGS data may be better paired with high-accuracy tools like OptiType, HLA-HD, xHLA, or HLA-VBSeq. If novel allele discovery or high-resolution typing is a priority, graph-based or assembly-driven methods such as Kourami, HLA*LA, or HISAT-genotype offer clear advantages. Users working with limited computational resources may prefer faster, lightweight approaches like xHLA or ArcasHLA, while those prioritizing maximum accuracy and completeness should consider tools with demonstrated high concordance rates and multi-digit resolution. If resources allow, use OptiType and POLYSOLVER for MHC-I and HLA-HD for MHC-II, or combine multiple tools for best results [115].
From peptide processing to MHC-peptide-TCR complex
The identification of neoantigens, tumor-specific peptides capable of eliciting immune responses, is pivotal for advancing cancer immunotherapy. The process leading from peptide generation to T cell recognition and the formation of the MHC–peptide–TCR complex is highly regulated and occurs through several sequential steps: proteolytic cleavage of proteins into peptides, peptide transport into the endoplasmic reticulum via the Transporter Associated with Antigen Processing (TAP) in the MHC class I pathway, binding of peptides to MHC molecules, presentation of these complexes on the cell surface, and ultimately, recognition by T cells. Each step is critical for determining which peptides are presented to T cells, and computational tools, particularly those leveraging AI, have significantly enhanced prediction accuracy.
The experimental data required for modeling each of these steps have been collected by researchers through various laboratory assays and are available in publicly accessible databases. For proteasomal cleavage prediction, datasets are derived from in vitro digestion experiments and MS analyses that identify cleavage sites within proteins, providing insight into how peptides are generated inside cells. TAP transport prediction relies on peptide translocation assays, often measuring BA or transport efficiency of peptides across the endoplasmic reticulum membrane via the TAP complex. Peptide–MHC binding and presentation prediction requires high-quality BA data (IC50 values), eluted ligand (EL) datasets identified through MS, and T cell activation assays that confirm immunogenicity. These diverse experimental data types are systematically collected and annotated in specialized databases such as IEDB [116], SYFPEITHI [117], TSNAdb [39], TumorAgDB1.0 [118], TANTIGEN 2.0 [119], MHCBN 4.0 [120], AntiJen [121], and MHCPEP [122], which integrate biochemical binding data, naturally processed ligands, and immunological assay results. Together, these datasets enable the development and benchmarking of computational models that aim to accurately predict antigen processing and immune recognition. The databases, along with the types of data available in each, are presented in Table 3.
Table 3.
Databases for neoantigen prediction and immunogenicity analysis
| Task | Database | Data Type | Web Accessibility | Ref. |
|---|---|---|---|---|
| Somatic mutation calling | The Cancer Genome Atlas (TCGA) | WGS, WES, and gene expression data | https://portal.gdc.cancer.gov/ | [123] |
| HLA typing | IMGT | HLA sequences | https://www.imgt.org/ | [124] |
| IPD-HLA | MHC sequences of different species | https://www.ebi.ac.uk/ipd/mhc/ | [125] | |
| Antigen processing and presentation (Proteasomal cleavage, TAP transport, peptide–MHC binding, immunogenicity) | Immune Epitope Database (IEDB) | Experimentally validated epitopes and immune responses | https://www.iedb.org/ | [116] |
| SYFPEITHI | MHC class I and II ligands, peptide motifs | http://www.syfpeithi.de/ | [117] | |
| TSNAdb |
Predicted and experimentally validated tumor-specific neoantigen |
https://pgx.zju.edu.cn/tsnadb1/ | [39] | |
| TumorAgDB1.0 | Curated tumor antigens | https://tumoragdb.com.cn/#/home | [118] | |
| TANTIGEN 2.0 | T cell epitopes and HLA ligands | http://projects.met-hilab.org/tadb | [119] | |
| MHCBN 4.0 | Peptides interacting with TAP and MHC | https://webs.iiitd.edu.in/raghava/mhcbn/index.html | [120] | |
| AntiJen | Quantitative peptide binding data (TAP, MHC, TCR–MHC) | https://www.ddg-pharmfac.net/antijen/AntiJen/antijenhomepage.htm | [121] | |
| MHCPEP | MHC class I and II ligands | http://wehih.wehi.edu.au/mhcpep | [122] | |
| Immunogenicity and TCR recognition | Immune Epitope Database (IEDB) | Antibody and T cell epitope data | https://www.iedb.org/ | [116] |
| McPAS-TCR | Disease-associated TCR sequences | https://friedmanlab.weizmann.ac.il/McPAS-TCR/ | [126] | |
| VDJdb | TCR sequences with known antigen specificities | https://vdjdb.cdr3.net/ | [127] | |
| 10x Genomics | Single-cell immune profiling (TCR sequences, gene expression) | https://www.10xgenomics.com/datasets | [128] | |
| TCRdb | TCR sequences | http://bioinfo.life.hust.edu.cn/TCRdb | [129] | |
| PIRD | TCR and immunoglobulin sequences across species | https://db.cngb.org/pird/ | [130] | |
| ImmuneCODE™ | TCR sequences linked to antigen specificity (infectious diseases) | https://clients.adaptivebiotech.com/pub/covid-2020 | [131] | |
| TBAdb | TCRs targeting specific antigens/diseases (subset of PIRD) | https://gitlab.com/immunomind/immunarch/raw/dev-0.5.0/private/TBAdb.xlsx | [NA] |
This section reviews the molecular mechanisms underlying each step and highlights AI-based tools (Fig. 3) that facilitate neoantigen identification, with applications in personalized immunotherapy.
Fig. 3.

Developed computational tools for antigen processing, MHC binding, and T cell recognition prediction for MHC class I and class II. The figure illustrates the antigen processing and presentation pathways for both MHC class I (left) and MHC class II (right) molecules, highlighting their interaction with CD8+ and CD4+ T cells, respectively. Various computational tools are categorized according to their functional roles, including cleavage prediction, TAP transport prediction, peptide–MHC BA prediction, peptide–MHC presentation prediction, immunogenicity (IM) prediction, and TCR–pMHC binding/IM prediction. Tools associated with the MHC class I pathway include those for proteasomal cleavage, TAP transport, peptide–MHC binding, and IM prediction, whereas tools for the MHC class II pathway emphasize endosomal processing, peptide binding, and IM prediction. The green boxes represent prediction tools designed for individual steps of the pathway, whereas the red boxes indicate tools that either integrate multiple steps within a single pathway (MHC-I or MHC-II) or predict a step across both pathways (MHC-I and MHC-II)
Proteasomal Cleavage. The proteasome, a multi-subunit protease complex, degrades intracellular proteins into peptides, generating potential ligands for MHC class I molecules. This process is highly specific, with cleavage sites determined by the proteasome’s catalytic subunits, which differ between constitutive proteasomes (expressed in all cells) and immunoproteasomes (induced in antigen-presenting cells under inflammatory conditions). The immunoproteasome, upregulated by interferon-γ, favors peptides with hydrophobic or basic C-termini, suitable for MHC class I binding. Furthermore, the immunoproteasome exhibits limited specificity, implying that not every potential mutated peptide will be generated during protein degradation [132]. Additionally, not all peptides produced by the proteasome will reach the necessary cellular compartments to potentially interact with HLA proteins [36]. Accurate prediction of cleavage sites is essential for identifying peptides that progress in the MHC class I pathway. Algorithms developed in this field are typically trained using either in vitro proteasome digestion data or in vivo MHC-I and -II ligand elution data [133].
Several computational tools, many employing AI, predict proteasomal cleavage sites by modeling sequence preferences. NetChop uses artificial neural networks (ANNs) trained on in vitro and in vivo cleavage data, achieving approximately 70% accuracy for C-terminal cleavage sites [134]. Pepsickle, a more recent tool, employs a gradient-boosted classifier, offering superior performance (AUC of 0.821 for constitutive and 0.789 for immunoproteasome) and the ability to differentiate between proteasome types [135]. PCPS utilizes an n-gram-based ML approach, outperforming NetChop in sensitivity (0.89 vs. 0.79) for immunoproteasome cleavage [136]. PCleavage applies support vector machines (SVMs), achieving a Matthews correlation coefficient of 0.54 for constitutive proteasome data [137]. PAProC uses an evolutionary algorithm, suitable for human and yeast proteasomes [138], while MAPPP relies on a kinetic model, though it is less flexible and not AI-based [139]. A random forest-based model by Li et al. (2012) provides additional ML-driven predictions, primarily for constitutive proteasomes [140]. For predicting cleavage sites of proteins in the MHC class II presentation pathway, tools such as PepCleaveCD4 [141] and MHCII-NP [142] show promise, but they require further improvement to accurately predict IM. Key features and reported performance of these proteasomal cleavage prediction tools are summarized in Table 4.
Table 4.
Cleavage site prediction tools
| Model/Year | Architecture | Supported Protease Type | Performance | MHC Class/Key Features/Web and Code Accessibility | Ref. |
|---|---|---|---|---|---|
|
PAProC 2001 |
Evolutionary algorithm | Human, Yeast | Accuracy = 0.82 |
I Evolves algorithms for diverse proteasome types. |
[138] |
|
MAPPP 2003 |
Kinetic model | Constitutive | NA |
I Modeling kinetic processes for cleavage prediction. NA |
[139] |
|
NetChop-3.1 2005 |
ANN | Constitutive, Immunoproteasome | AUC = 0.80 |
I Accurate prediction based on neural networks. |
[134] |
|
ProteaSMM 2005 |
Matrix-based | Constitutive, Immunoproteasome | AUC = 0.76 for i20S |
I Integrated proteasomal cleavage, TAP transport, and MHC I binding predictions. |
[143] |
|
PCleavage 2005 |
SVM | Constitutive, Immunoproteasome | AUC = 0.79 |
I Support vector machines for robust predictions. |
[137] |
|
Li et al. 2012 |
RF | Constitutive (likely) | Accuracy = 0.85 |
I Random forests for versatile prediction models. NA |
[140] |
|
PepCleaveCD4 2013 |
SVM |
Endosomal proteases: cathepsins L and S |
AUC = 0.85 |
II Predictor highlights the effect of secondary structure on CD4+ T cell epitope preprocessing. |
[141] |
|
MHCII-NP 2018 |
Enrichment and depletion-based score | Lysosomal proteases | AUC = 0.767 |
II Incorporates cleavage and binding motifs into prediction of MHC-II ligands. |
[142] |
|
PCPS 2020 |
N-gram-based | Constitutive, Immunoproteasome |
Sensitivity = 0.88 Specificity = 0.57 |
I Uses n-gram models to predict cleavage sites. |
[136] |
|
Pepsickle 2021 |
Gradient-boosted classifier | Constitutive, Immunoproteasome | AUC = 0.87 |
I Utilizes ML for improved accuracy. |
[135] |
Recommendations: When selecting a cleavage site prediction tool from those listed in Table 4, users should first consider the type of protease and MHC class relevant to their study, as most tools are specialized either for MHC class I (proteasomal cleavage) or class II (endosomal/lysosomal processing). For MHC class I applications, widely used and reliable options such as NetChop-3.1, PCleavage, and Pepsickle offer solid performance, with newer ML-based methods like Pepsickle generally providing improved accuracy and flexibility. If users require integrated antigen-processing predictions, tools like ProteaSMM are advantageous because they combine cleavage, TAP transport, and MHC binding steps. For MHC class II studies, more specialized tools such as PepCleaveCD4 and MHCII-NP should be preferred, as they are designed to capture endosomal protease activity and peptide presentation pathways.
TAP Transport. TAP moves peptides that are 8–16 amino acids long from the cytosol into the ER, where they become available for loading onto MHC class I molecules. TAP preferentially transports peptides with specific sequence motifs, influencing the epitope repertoire. Predicting TAP binding affinity is crucial for identifying peptides likely to proceed to MHC loading.
AI-based tools enhance TAP binding predictions. TAPPred uses a cascade SVM approach, achieving a correlation coefficient of 0.88 in jack-knife validation [144]. PREDTAP combines ANNs and hidden Markov models (HMMs), offering robust predictions with an AUC greater than 0.85 [145]. TAPREG, also SVM-based, predicts binding affinities for peptides of variable lengths (8–16 residues), with a maximum correlation of 0.89 [146]. The IEDB TAP prediction tool employs a matrix-based approach, providing a non-AI alternative [147]. SVMTAP applies SVMs with a focus on peptide sequence features, achieving high specificity [148]. Combining TAP transport with MHC-I binding improves epitope identification, and cascade SVM models report a strong correlation with measured transport [144, 147]. DeepTAP, a recent DL tool, uses bidirectional gated recurrent units (BiGRUs) to capture sequential dependencies, outperforming traditional ML methods [149]. Representative TAP-binding prediction methods and their main characteristics are listed in Table 5.
Table 5.
TAP binding affinity prediction tools
| Model/Year | Architecture | MHC Class | Performance | Key Features/Code Accessibility | Ref. |
|---|---|---|---|---|---|
|
IEDB TAP 2003 |
Matrix-based | I | AUC = 0.932 on HLA-A0201 |
Simple scoring matrix derived from experimental binding data. |
[147] |
|
TAPPred 2004 |
Cascade SVM | I | Correlation r = 0.88 |
Combines sequence plus 33 amino-acid physicochemical features; high accuracy via multi-level SVM cascade. |
[144] |
|
PREDTAP 2006 |
ANN + HMM | I | AUC > 0.85 |
Integrates neural network and hidden Markov modeling for TAP affinity. |
[145] |
|
TAPREG 2010 |
SVM | I | Pearson r = 0.89 |
SVM-based TAP binding prediction with residue-based encoding. |
[146] |
|
TAP Hunter 2010 |
SVM | I | AUC = 0.85 |
Using feature vectors from the N- and C-terminal positions of TAP ligands. |
[150] |
|
DeepTAP 2023 |
BiGRU RNN | I |
Spearman r = 0.91 Pearson r = 0.89 |
DL captures sequential context; improved precision for top-ranked neoantigens. |
[149] |
|
CLTAP 2025 |
LLM with contrastive learning | I | NA |
Integrates Contrastive Learning (CL) with a contextual co-attention mechanism. |
[151] |
Recommendations: For straightforward analyses or rapid screening, classical tools such as IEDB TAP and TAPPred provide reliable performance with simple implementations, making them suitable for users with limited computational resources or those integrating TAP prediction into larger antigen-processing pipelines. For more advanced applications, especially in neoantigen discovery or precision immunotherapy, newer ML- and DL-based approaches such as DeepTAP offer improved predictive power by capturing sequence context more effectively, while emerging models like CLTAP may further enhance performance through advanced architectures like contrastive learning and attention mechanisms. Overall, combining TAP prediction with upstream (cleavage) and downstream (MHC binding) analyses, and optionally cross-validating results with multiple tools, can significantly improve the reliability of antigen presentation predictions.
Peptide–MHC binding/presentation. The binding of peptides to MHC class I or II molecules determines which peptides are presented to T cells. MHC class I binds short peptides (8–10 amino acids) in a closed groove, while MHC class II accommodates longer peptides (13–25 amino acids) in an open-ended groove. High BA is a primary determinant of immunogenicity, as only stable binders are likely to be presented effectively. AI-driven tools, particularly those using DL, have revolutionized BA predictions by modeling complex peptide-MHC interactions across diverse alleles.
Accurate prediction of peptide-MHC BA is essential for neoantigen identification in cancer immunotherapy. Numerous computational methods have been created, leveraging ML and DL to model these interactions. Early tools like NetMHC used ANNs to predict binding affinities based on sequence data [152]. The introduction of pan-specific models, such as NetMHCpan, marked a significant advancement by enabling predictions for alleles not included in the training set, thus broadening applicability to diverse populations [153]. The advent of DL has further enhanced prediction accuracy. Convolutional neural networks (CNNs), as employed in tools like ConvMHC and MHCflurry, capture local patterns in peptide sequences [154, 155], while recurrent neural networks (RNNs) and long short-term memory (LSTM) networks, featured in MHCnuggets and USMPep, handle sequential dependencies effectively [156, 157]. More sophisticated architectures, including attention mechanisms and transformer models (e.g., DeepHLApan, ImmunoBERT), focus on critical residues and contextual information within sequences [158, 159]. Beyond BA, some tools predict peptide presentation or immunogenicity directly. For instance, MARIA incorporates gene expression and cleavage predictions to estimate the likelihood of MHC class II presentation, providing a more comprehensive assessment for neoantigen identification [160].
Recent tools such as BigMHC, DeepNeo, HLAthena, PRIME 2.0, NetMHCpan-4.1, and Class I Immunogenicity (IEDB) further advance the field by integrating diverse data types and advanced architectures to predict BA, EL presentation, and IM. These tools differ in their methodologies, training data, and predictive capabilities, offering varied strengths for specific applications [11, 161–165]. In Table 6, we provide a detailed comparison of these tools, including their architectures, training data, prediction types, and web accessibility, along with short explanations highlighting their main methodologies and features, and differences from other models.
Table 6.
Peptide-MHC binding affinity and presentation prediction tools
| Model/Year/Ref. | Architecture | Training Data/Output | MHC Class /Performance | Key Features/Web and Code Accessibility |
|---|---|---|---|---|
|
Parker et al., 1994 [166] |
Additive motif/matrix model |
BA BA |
Class I NA |
First fully quantitative per-position coefficient table for HLA-A2; formalizes the independent binding of side-chains hypothesis and shows it holds for most peptides (only 3/83 require side-chain interactions). NA |
|
SYFPEITHI 1999 [117] |
Motif-based, hand-crafted position-specific scoring matrix per allele |
200 peptide motifs and 2,000 peptide sequences Motif score and rank |
Class I NA |
Early motif-based system built on experimentally eluted ligands and known T cell epitopes only (no synthetic binders). Encodes anchor and auxiliary anchor residues, preferred and unfavorable AAs via integer weights. |
|
ProPred 2001 [167] |
Classical matrix / PSSM-based predictor |
BA BA |
Class II NA |
One of the earliest widely-used web servers for DR-restricted epitope prediction; covers 51 DR alleles; simple threshold-based interface. |
|
CTLPred 2004 [168] |
Quantitative Matrix, ANN, SVM |
BA BA |
Class I Accuracy ≈ 70% |
Independent of exact HLA restriction → can be used when HLA is unknown; Implemented as a practical web server, widely used in early vaccine-design pipelines. |
|
RANKPEP 2004 [169] |
PSSM-based scorer |
BA BA |
Class I & II AUC > 0.8 for Class I; AUC > 0.7 for Class II |
Unified PSSM framework for MHCI and MHCII; supports MSA-based variability masking so only conserved regions are scanned; integrates proteasomal cleavage prediction; provides allele-specific PSSM binding thresholds (PSBT). |
|
ARB 2005 [170] |
Average Relative Binding (ARB) coefficient matrices |
BA BA |
Class I & II AUC ≈ 0.80 |
Provides direct IC₅₀ prediction rather than only scores, enabling multi-allele / multi-length global ranking; covers multiple species and peptide sizes (8–11 for class I, ≥9 for class II) with systematic cross-validation. |
|
IEDB SMM 2005 [171] |
SMM |
BA BA |
Class I NA |
Fully quantitative inputs and outputs; built-in handling of experimental noise; robust treatment of bounded measurements. |
|
MHCPred 2.0 2006 [172] |
QSAR-based regression models (additive method) |
BA BA |
Class I & II NA |
Provides quantitative affinity predictions (not just classifier scores); supports both human and mouse alleles; allows specifying anchor positions; built on explicit QSAR models that can help interpret residue contributions. Frequently used as a QSAR-style comparator in MHC-binding studies. |
|
SMM-Align 2007 [173] |
Stabilization Matrix Alignment |
BA BA |
Class II Mean AUC ~0.73–0.76 |
Hybrid alignment + motif-learning method. |
|
SMMPMBEC 2009 [174] |
SMM |
BA BA |
Class I Average AUC for SMMP M B E c vs SMM is 0.860 vs 0.836 |
Penalizes opposite charge substitutions strongly (due to adverse effects on peptide–MHC binding), unlike BLOSUM62. |
|
NN-align 2009 [175] |
ANN |
BA BA |
Class II AUC = 0.810 |
Simultaneously infers binding core and affinity; explicitly encodes peptide-flanking residues and peptide length; corrects redundancy bias by scaling backprop step size by the size of the core-redundancy cluster (Hobohm-1 clustering); uses ensemble averaging over multiple architectures and random initializations. |
|
PickPocket 1.1 2009 [176] |
PSMM |
BA BA |
Class I AUC = 0.895 |
Works well with limited training data, unlike artificial neural networks. Especially effective for non-human species. |
|
MultiRTA 2010 [177] |
Regularized Thermodynamic Average |
BA BA |
Class II AUC = 0.783 |
Extends the earlier single-allele RTA model to 430 HLA-DR/DP allotypes; demonstrates strong leave-one-allele-out performance and good generalization to overlapping-peptide epitope data sets. |
|
NetCTLpan 2010 [178] |
Integrated pathway model (ANN) |
SYFPEITHI NetCTLpan score = pathway presentation likelihood |
Class I AUC ≈ 0.976–0.982 |
Pan-specific CTL epitope prediction for any HLA class I molecule with known sequence; integrates antigen processing (cleavage + TAP) with binding; Supports 8, 9, 10, and 11-mer epitopes. |
|
MULTIPRED2 2011 [179] |
ANN |
BA BA |
Class I & II NA |
Large-scale screening across alleles, supertypes and full genotypes; maps promiscuous T cell epitope hotspots in whole proteomes; exploits strong pan-specific predictors without re-training; convenient for vaccine design / neoantigen prioritization. |
|
PAComplex 2011 [180] |
Structural modeling |
BA, TCR–pMHC structure, complexes and experimental data pMHC binding model, pTCR binding model |
Class I NA |
Detects homologous peptide antigens. |
|
TEPITOPEpan 2012 [181] |
Pan-specific PSSM |
BA BA |
Class II Best average AUC ≈0.739 |
Competitive AUC with NetMHCIIpan on many alleles; robust performance even when training data are restricted; best among compared methods in predicting exact binding cores for 20 peptide–MHC complexes. |
|
IEDB-AR-Consensus 2012 [182] |
Consensus-based |
BA BA |
Class I AUC = 0.96 |
Large quantitative peptide–MHC binding datasets (IC50) across HLA-A/B, primate and mouse alleles. |
|
NetMHCcons 1.1 2012 [183] |
Consensus |
BA BA |
Class I NA |
Dynamically combines NetMHCpan, NetMHC, and PickPocket depending on allele characterization and similarity to known alleles. |
|
OWA-PSSM 2013 [184] |
PSSM |
BA BA |
Class II AUC = 0.747 |
Extends TEPITOPE’s pocket-profile framework to 879 HLA-DR alleles using OWA weights over pocket pseudo-sequence similarities; does not require large allele-specific training sets. NA |
|
NetMHCstabpan 1.0 2016 [185] |
ANN |
pMHC-I half-life measurements data BS |
Class I NA |
Shows that combining stability and affinity predictions significantly improves T cell epitope prediction. https://services.healthtech.dtu.dk/services/NetMHCstabpan-1.0/ |
|
NetMHC-4.0 2016 [152] |
Gapped sequence alignment using ANN |
BA, EL BA, EL |
Class I AUC > 0.9 |
Integrates BA and EL data into a unified model for improved prediction of peptide presentation and binding. Accurately captures peptide length preferences. Performs well on both neoantigen identification and natural ligand datasets. |
|
PSSMHCpan 2017 [186] |
Position-Specific Scoring Matrix (PSSM)-based model |
BA BA |
Class I AUC = 0.94 |
Combines allele-specific and pan-specific PSSMs using HLA sequence similarity; supports broad peptide lengths (8–25) and very large HLA coverage (4,896 alleles). Designed for high-throughput neoantigen discovery. |
|
NNAlign − 2.1 2017 [187] |
ANN |
BA, EL BA, EL |
Class I & II AUC = 0.81 |
Works with any biological sequence type (protein, DNA, RNA); generates gapped sequence alignments; aligns sequences of variable lengths to identify motifs. |
|
MSIntrinsic 2017 [188] |
ANN |
EL EL |
Class I AUC = 0.99 |
Discovery of sequence motifs; improved quantification of the roles of gene expression and proteasomal processing. NA |
|
HLA-CNN 2017 [189] |
CNN |
BA BA |
Class I AUC = 66.7% |
Treating peptides as sentences and AAs as words. HLA-Vec captures physicochemical properties; demonstrated use on the entire human proteome to highlight highly self-binding 9-mers. |
|
ConvMHC 2017 [154] |
DCNN |
BA BA |
Class I AUC > 0.9 |
Leverages convolutional layers to detect positional motifs in peptide sequences. One of the earliest deep learning architectures for peptide–MHC binding prediction. |
|
DeepMHC 2017 [190] |
DCNN |
BA BA |
Class I NA |
DL for improved feature extraction; robust for MHC-I binding. |
|
AI-MHC 2018 [191] |
DCNN |
BA BA |
Class I & II NA |
Supports both MHC classes; DL enhances binding prediction. |
|
EDGE 2019 [192] |
DCNN |
EL EL |
Class I NA |
Trained on tumor MS data; high accuracy for neoantigen identification. |
|
MARIA 2019 [160] |
RNN |
EL EL |
Class II AUC = 0.89 |
Diverse training data; integrates gene expression and cleavage data. |
|
CNN-NF 2019 [193] |
DCNN |
BA BA |
Class I NA |
Focuses on local sequence features; improves BA prediction. |
|
DeepLigand 2019 [194] |
Deep language model (ELMo) with residual network (DL) |
BA BA |
Class I NA |
Language model approach; robust for MHC-I binding prediction. |
|
PUFFIN 2019 [195] |
Deep residual network |
BA, EL BA |
Class I & II NA |
Uncertainty quantification; supports both MHC classes. |
|
NeonMHC2 2019 [196] |
Ensemble of CNNs |
EL EL |
Class II NA |
Ensemble approach; high accuracy for MHC-II presentation. |
|
DeepSeqPan 2019 [197] |
DCNN |
BA, EL BA, EL |
Class I NA |
DL for pan-specific MHC-I binding; high generalization. |
|
ACME 2019 [198] |
Attention-based CNNs |
BA BA |
Class I NA |
Attention mechanism improves focus on key residues. |
|
ForestMHC 2019 [199] |
RF |
EL EL |
Class I AUC = 0.73 |
ForestMHC scores correlate monotonically (not linearly) with IC50 values, suggesting peptide presentation isn’t solely dependent on chemical affinity. |
|
MHCherryPan 2020 [200] |
LSTM, CNN |
BA, EL BA, EL |
Class I NA |
Pan-specific; combines LSTM and CNN for robust predictions. NA |
|
MHCnuggets 2020 [201] |
LSTM networks and GRUs |
BA BA |
Class I & II AUC (human) = 0.64 |
Handles sequential dependencies; supports both MHC classes. |
|
USMPep 2020 [157] |
AWD LSTM with embedding layer |
BA BA |
Class I & II NA |
Universal sequence model; effective for diverse alleles. |
|
IConMHC 2020 [202] |
Deep CNN |
BA BA |
Class I AUC on common vs rare alleles: 0.834 vs 0.789 |
Encodes pairwise amino-acid interaction properties rather than one-hot/BLOSUM; handles non-9mers by gapping/trimming to multiple 9-mer views and takes max-affinity; pan-allele, predicts rare alleles with no training data. NA |
|
MHCAttnNet 2020 [203] |
Bi-LSTM encoder with an attention mechanism |
BA BA |
Class I and II AUC-PRC = 94.18% |
Attention yields interpretable heatmaps over amino acids and reduces candidate MHC trigrams from ~9,251 to ~258, focusing wet-lab validation. Handles variable-length peptides via Bi-LSTM. |
|
MHCflurry 2.0 2020 [204] |
ANN |
BA, EL BA, EL |
Class I AUC = 0.877–0.985 |
Improved over MHCflurry; integrates MS data for better predictions. |
|
NetMHCpan 4.1 2020 [164] |
an ensemble of 100 single-layer neural networks |
BA, EL BA, EL |
Class I NA |
Improved over NetMHCpan-4.0; better for HLA-B and HLA-C; pan-specific. |
|
HLAthena 2020 [11] |
Single-layer neural network |
IEDB EL |
Class I HLAthena (MSiCEB model): up to 78% recall |
Integrates gene expression and peptide cleavage information. |
|
BERTMHC 2021 [205] |
BERT-based with multiple instance learning |
BA, EL BA, EL |
Class II NA |
BERT-based; high performance for MHC-II binding. |
|
DeepAttentionPan 2021 [206] |
DL with attention mechanism |
BA, EL BA, EL |
Class I NA |
Attention-based; improved pan-specific predictions. |
|
DeepNetBim 2021 [207] |
DL with network analysis |
BA, IM BA, IM |
Class I NA |
Integrates immunogenicity; network analysis enhances predictions. |
|
SHERPA 2021 [208] |
Gradient boosting decision trees |
EL EL |
Class I NA |
High performance for eluted ligands; robust for neoantigens. NA |
|
MATHLA 2021 [209] |
Bidirectional LSTM with attention |
BA, EL BA, EL |
Class I NA |
Bidirectional LSTM; focuses on key sequence features. |
|
ImmunoBERT 2021 [159] |
BERT-based |
BA, EL BA, EL |
Class I NA |
BERT-based; high generalization for MHC-I binding. |
|
Anthem 2021 [210] |
NB, XGBoost, LR, NN, SVM, DT, RF |
BA BA |
Class I AUC = from 0.888 to 0.981 |
Combines scoring function-based methods (PWM/PSSM) with ML. |
|
NMER 2021 [211] |
RF |
EL EL |
Class I AUC ≈ 0.91 |
Integrates MHCflurry1.6 binding/processing/presentation scores, NetMHCpan-2.8 binding, NetMHCstabpan-1.0 stability, proteasomal C-terminal cleavage, TAP transport scores, etc. |
|
RBM-MHC 2021 [212] |
Restricted Boltzmann Machine (RBM) |
EL EL |
Class I AUC = 0.991 |
Score presentability of (neo)antigens without HLA information. Visualize HLA-binding motifs and binding modes for data exploration and feature discovery. Classify peptides by HLA restriction with minimal annotation. |
|
TransPHLA 2022 [213] |
Simplified Transformer with masked multi-head self-attention |
BA BA |
Class I AUC = 0.926 |
TransPHLA for binding + AOMP module that uses attention scores to propose higher-affinity peptide mutations for a target HLA; provides allele/length-specific attention heatmaps; pan-allele support for A/B/C and 8–14mers. |
|
DeepSeqPanII 2022 [214] |
RNN with attention |
BA, EL BA, EL |
Class II NA |
Attention-based; tailored for MHC-II binding prediction. |
|
MHCRoBERTa 2022 [215] |
Transfer learning with BERT |
BA, EL BA, EL |
Class I NA |
Transfer learning enhances performance; robust for MHC-I. |
|
FIONA 2022 [216] |
Flexible NN architecture |
BA, EL BA, EL |
Class II NA |
Flexible architecture; optimized for MHC-II binding. |
|
HLApollo 2022 [217] |
Transformer model (DL) |
BA, EL BA, EL |
Class II NA |
Transformer-based; high accuracy for MHC-II. NA |
|
HLAB 2022 [218] |
ProtBert, BiLSTM, UMAP, LR/SVM/XGBoost |
BA BA |
Class I AUC = 0.9891 |
Cascaded protein LM (ProtBert) + BiLSTM embeddings; UMAP dimensionality reduction; filter and wrapper feature selection (T-test, W-test, RF, LR-RFE, SVM-RFE); evaluates seven classifiers and picks the best per task. |
|
DeepMHCII 2022 [219] |
CNN |
BA BA |
Class II AUC = 0.77 |
Identifying the binding core and pinpointing key binding pockets. |
|
IEPAPI 2023 [220] |
Transformer-based feature extraction |
EL, IM EL, IM |
Class I NA |
Two-task design (EL→IM) mirrors biology; attention visualizes HLA-restricted motifs; uses HLA pseudo-sequence (34 positions); EL-pretraining then freezes for IM; supports variable peptide lengths (8–11). |
|
MixMHC2pred-2.0 2023 [221] |
Deep motif deconvolution with NNs |
EL EL |
Class II AUC (human) = 0.94 |
Motif deconvolution; high specificity for MHC-II ligands. |
|
CapsNet-MHC 2023 [222] |
Capsule neural networks |
BA BA |
Class I AUC: Higher than baselines (specific values in Fig. 7a) |
Captures pairwise features; high interpretability via dynamic routing. |
|
DeepMHCI 2023 [223] |
Anchor position-aware DL |
BA, EL BA, EL |
Class I NA |
Anchor-focused; improves prediction for MHC-I binding. |
|
TLimmuno2 2023 [224] |
LSTM, Transfer learning |
BA BA, IM |
Class II NA |
Transfer learning; high performance for MHC-II binding. |
|
NetMHCIIpan-4.3 2023 [225] |
ANN |
BA, EL BA, EL |
Class II NA |
Enhances prediction accuracy for all HLA class II molecules by integrating locus-specific immunopeptidomics data. https://services.healthtech.dtu.dk/services/NetMHCIIpan-4.3/ |
|
TransMHCII 2023 [226] |
PLM embedding and an image classifier |
BA BA |
Class II AUC = 0.908 |
Represents multi-class classification model for MHC-II binding prediction. |
|
MHC2AffyPred 2023 [227] |
Random forest regression on structural fingerprints |
BA BA |
Class II NA |
Structure-based ML approach; shows consistently higher correlations than NetMHCIIpan-3.2 and MHCII3D on shared benchmarks and is demonstrated on SARS-CoV-2 epitope affinity prediction. |
|
STMHCpan 2023 [228] |
Star-Transformer |
EL EL |
Class I AUC = 0.9499 |
Replaced the fully connected structure with star topology using a lightweight Star-Transformer, reducing model complexity while enhancing predictive performance. Incorporated an attention mechanism to enhance the architecture and performance of a DL model for prediction. |
|
MHCSeqNet2 2024 [156] |
DCNN, GRU and LSTM |
BA, EL BA, EL |
Class I AUC = 0.99 (on MS data); AUC = 0.54 (on T cell epitopes) |
Uses sub-word-level peptide features and a 3D structure embedding; generalizes to unseen alleles; outperforms NetMHCpan on some datasets. |
|
Graph-pMHC 2024 [229] |
GNN (utilizes Alphafold2-multimer-derived graph adjacency matrices) |
EL EL |
Class II NA |
Better evaluation via GO-based splitting, assess antibody immunogenicity risk. |
|
RPEMHC 2024 [230] |
CNN |
BA BA |
Class I & II AUC (MHCI) = 0.866, AUC (MHCII) = 0.759 |
Uses residue–residue pair encoding; captures critical interaction information between the molecules. |
|
TripHLApan 2024 [231] |
Triple coding matrix, BiGRU, Attention model, Transfer learning |
BA BA |
Class I & II AUC = 0.958 |
Optimizes allele sequence extraction by focusing on interaction sites and correlating HLA binding motifs with coding strategies. |
|
HLAPepBinder 2024 [232] |
RF |
BA BA |
Class I Accuracy = 0.90–0.91 |
Consensus pipeline using nine predictors (ANN, Consensus, NetMHCpan BA, NetMHCpan EL, SMM, SMMPMBEC, PickPocket, NetMHCcons, NetMHCstabpan), high generalizability to unseen HLA subtypes, addressing the negative data problem. |
|
MUNIS 2025 [233] |
LSTM, ESM-2 |
EL EL |
Class I AUC = 0.980 |
Predicts immunodominance hierarchies in HIV and EBV more effectively. 10.5281/zenodo.14219509 |
|
OnmiMHC 2025 [234] |
1D-CNN-LSTM and 2D-CNN |
BA, EL BA |
Class I & II MHC-I PR-AUC = 0.854 MHC-II PR-AUC = 0.606 |
Multimodal feature fusion; combination of 2D + 1D convolutional kernels; iterative data preprocessing. |
|
pMHChat 2025 [235] |
LLMs and Deep Hypergraph Learning |
BA, PDB BA |
Class II AUC = 0.7676 |
BA prediction through illuminating the residue contact profiling of the binding surface. |
|
MixMHCpred3.0 2025 [236] |
Neural network |
EL EL |
Class I LOA AUC ≥ 0.98 |
Predicts ligands for MHC-I alleles without known ligands, as well as MHC-I alleles across species. |
|
PHLA-SiNet 2025 [237] |
Siamese neural network |
BA BA |
Class I NA |
Higher sensitivity without substantial degradation of other metrics. |
Current presentation predictors are trained on EL datasets in addition to affinity data, for instance, NetMHCpan-4.0 integrates both to capture length preferences and increase accuracy, and MHCflurry-2.0 couples binding and antigen-processing models into an integrated presentation score [153, 204]. Large mono-allelic EL compendia (e.g., HLAthena: >185,000 peptides across 95 alleles) drive improved PPV and enable validation in tumor cell lines [11]. Tissue-resolved, benign-organ immunopeptidomes (HLA Ligand Atlas) contextualize tumor ligandomes and offer broad background sets for benchmarking [238].
Recommendations: Based on extensive benchmarking studies, several models have been identified as top-performing tools for peptide–MHC prediction. For MHC class I, NetMHCpan-4.1 and MixMHCpred, and for MHC class II, NetMHCIIpan-4.1 and MixMHC2pred, are widely regarded as state-of-the-art and industry standards for predicting pMHC binding affinity and presentation [239, 240]. MHCflurry 2.0 demonstrates comparable performance. These tools leverage artificial neural networks and integrate both BA and EL data, enabling them to predict not only binding but also natural antigen processing and presentation. They also provide broad HLA allele coverage, including rare alleles. Additionally, integrative tools such as NetCTLpan, which combine multiple biological processes including peptide binding, antigen processing, and presentation, can deliver more biologically realistic predictions, particularly for applications like neoantigen discovery and vaccine design. Finally, combining predictions from multiple tools is often a robust strategy to mitigate model-specific biases.
Immunogenicity and TCR recognition. Recognition of a presented pMHC occurs when T cell receptors identify the pMHC as foreign (non-self). This recognition activates T cells, leading to their expansion and the targeted elimination of cancer cells displaying the detected pMHCs. Predicting pMHC recognition is a key step in neoantigen prediction, since not all presented pMHCs provoke an immune response.
Approaches to predicting pMHC recognition can be divided into two categories: those that account for specific TCR information and those that do not. In TCR-focused strategies, the goal is to model the binding relationship between individual TCRs and pMHCs. These methods frequently employ protein sequence representations combined with ML architectures such as convolutional or recurrent neural networks [241–243]. Training typically relies on curated repositories including the Immune Epitope Database (IEDB) [116], McPAS-TCR [126], and VDJdb [127] that provide experimentally confirmed records of TCR–pMHC interactions. TESLA assembled a multi-center IM dataset and model that removes the vast majority of non-immunogenic peptides at high precision, providing a community benchmark for recognition-aware pipelines [12]. Public epitope resources anchor comparative evaluations across antigen processing and presentation. IEDB functions as a central repository that aggregates experimentally measured epitopes and exposes them through a public, searchable interface, supporting systematic reuse in computational studies [116]. Its scope spans antibody, T cell, and MHC binding contexts, allowing modelers to align tasks to immunological readouts and to stratify benchmarks by assay type, organism, or antigen class within a single harmonized framework [116]. The TESLA consortium coordinated a multi-group, end-to-end evaluation in which participants predicted immunogenic epitopes from shared tumor sequencing data, after which a large candidate set was tested for T cell binding; the resulting model excluded the overwhelming majority of non-immunogenic peptides at high precision and reproduced performance in an independent cohort, defining rigorous criteria for pipeline comparison [12]. On the receptor side, VDJdb has expanded the number of TCR sequences with known cognate antigens, introduced compact motif sets suitable for training specificity predictors, and enabled batch annotation of repertoire sequencing data, supporting generalization tests on realistic repertoires [127]. Complementarily, McPAS-TCR offers a manually curated catalogue of pathology-linked TCR sequences numbering over five thousand entries, furnishing disease-labeled sets for benchmarking recognition models that aim to recover pathology-specific repertoires or antigen associations [126].
Most ML models for TCR–pMHC prediction are based on amino acid sequence encoding. Input sequences are numerically encoded into vector embeddings, using methods such as one-hot encoding, BLOSUM (evolutionary similarity) [244], AA-index [245], or Atchley factors [246] (physicochemical properties). They use the antigenic peptide, the MHC, and the TCR sequence as inputs. The MHC-peptide-TCR binding predictors, including ERGO-II, ImRex, DLpTCR, TITAN, NetTCR-2.2, and pMTnet, differ in their architectural choices and training datasets [247–252]. BERT-based models, such as STAPLER, TABR-BERT, and EPIC-TRACE, treat protein sequences as a biological language, generating contextual embeddings of TCRs, peptides, and MHCs to predict recognition [253–255]. Extending this idea, ESM-based methods leverage large protein language models trained on massive sequence datasets (ESM-1b, ESM-2) to capture both evolutionary and structural features; examples like TCR-ESM often benefit from transfer learning to improve generalization [256]. In contrast, AlphaFold-driven approaches emphasize structure, predicting 3D TCR–pMHC docking complexes and refining predictions with scoring functions or graph neural networks, as seen in tools like TCRmodel2 and NetTCR-struc [257, 258]. Together, these strategies highlight complementary strengths; BERT and ESM models capture sequence-level features, whereas AlphaFold-based methods provide structural insight into binding interfaces. Nonetheless, all approaches remain challenged by limited paired TCR–pMHC data and the vast diversity of TCR clonotypes. Table 7 provides a comparative overview of the available tools.
Table 7.
TCR-pMHC binding and IM prediction tools
| Model/Year/Ref. | Architecture | Training Data Source/Output | MHC Class/Performance | Key Features/Web and Code Accessibility |
|---|---|---|---|---|
|
POPISK 2011 [259] |
SVM and weighted degree string kernel |
MHCPEP, SYFPEITHI, IEDB IM |
Class I AUC = 0.74 |
Only sequence information used in the model. |
|
PAAQD 2013 [260] |
RF |
MHCPEP, SYFPEITHI, IEDB IM |
Class I AUC = 0.72 |
Uses quantum-derived physicochemical features describing molecular similarity. |
|
Class I Immunogenicity (IEDB) 2013 [165] |
Combination of the enrichment scores and position weights |
IEDB, publications IM |
Class I NA |
It excludes anchor positions P2/P9 to avoid binding biases. |
|
EpiToolkit 2015 [261] |
Multiple immunoinformatics tools |
EpiToolKit itself does not train new models Ranked epitope sets |
Class I & II |
Full vaccine-design workbench covering all steps from allele selection and epitope conservation to prediction. |
|
TepiTool 2016 [262] |
Multiple ML and matrix-based MHC binding predictors |
IEDB-curated quantitative peptide–MHC BA and EL datasets IC50 and percentile rank |
Class I & II TepiTool itself is an interface; underlying IEDB MHC binding predictors typically achieve ROC AUC > 0.9 in benchmark evaluations |
Step-by-step 6-step wizard; exposes IEDB “recommended” settings by default; supports epitope scanning, allele selection, peptide length/overlap control, and automatic selection of top-ranked epitope candidates across MHC I and II. |
|
GLIPH 2017 [263] |
Similarity-based clustering of similar TCRs |
Publications, PDB Motifs, CDR3 clusters, antigen-specific TCR groups |
Class I NA |
Finds TCRs with shared antigen specificity by combining motif discovery, sequence similarity, structural contact probability, and immunological features such as V-gene bias and clonal expansion. |
|
TCRdist 2017 [264] |
Distance-based (BLOSUM62 matrix) clustering of similar TCRs |
In-house dataset TCRdist matrices, clusters, motifs, epitope-specific TCR assignment |
Class I NA |
Generalizes Simpson’s diversity index by including similarity between receptors, not only identity. Classifies new TCRs to epitopes based on density of nearby receptors in TCRdist space. |
|
ITCell 2018 [265] |
Template-based |
Multiplex Substrate Profiling by Mass Spectrometry (MSP-MS), IEDB, PDB Antigen cleavage sites, BA, IM |
Class II NA |
Integrative structure-based approach that incorporates antigen cleavage by proteases, MHCII presentation, and TCR recognition. |
|
Structure 2019 [266] |
ANN |
PDB, publications, IEDB IM |
Class I AUC = 0.60 |
Structure-based IM prediction. NA |
|
Tcell_predictor 2019 [267] |
RF |
Publication, TANTIGEN IM |
Class I AUC = 0.65–0.73 |
Protein expression level and charge differences introduced by the mutation significantly improves the ability to identify immunogenic neo-epitopes. |
|
TCRex 2019 [268] |
RF |
McPAS-TCR, VDJdb, ImmuneCODE TP binding |
Class I & II NA |
A stringent BPR threshold is applied to reduce false positives. Enables annotation of entire TCR repertoires with predicted epitope targets. |
|
DeepHLApan 2019 [158] |
RNN (BiGRU) with an attention layer |
IEDB BA, IM |
Class I BA AUC > 0.9 |
Good performance on clinical data. |
|
Smith et al. 2019 [269] |
Gradient boosting model |
210 class I and 68 class II predicted neoantigens with IFNγ ELISpot readouts (in-house dataset) IM |
Class I NA |
Single model predicts neoantigen/mHA immunogenicity directly from peptide-intrinsic features; extensive feature engineering. |
|
TTAgP 1.0 2019 [270] |
RF |
IEDB, TANTIGEN IM |
Class I NA |
TTAgP uses a comprehensive set of peptide features. |
|
POTN 2020 [271] |
SVM |
IEDB, peptide database, SYFPEITHI IM |
Class I NA |
Physicochemical properties (position-specific, especially at P3). https://www.frontiersin.org/articles/10.3389/fimmu.2020.02193/full#supplementary-material |
|
INeo-Epp 2020 [272] |
RF |
IEDB, publications IM |
Class I AUC = 0.81 |
Eliminate false positive predicted peptides. |
|
iTTCA-Hybrid 2020 [273] |
SVM, RF |
IEDB, TANTIGEN IM |
Class I AUC = 0.783 |
Hybrid feature representation. Balanced training through SMOTE. |
|
DeepImmuno-CNN and -GAN 2021 [274] |
CNN for IM prediction + GAN for sequence generation |
IEDB IM |
Class I ROC AUC ≈ 0.85 and PR AUC ≈ 0.81 |
Integrates beta-binomial modeling of immunogenic potential with an HLA-contextual CNN; explicitly models peptide–MHC pairs; interpretable via residue occlusion analysis to highlight key antigen positions. |
|
TCRGP 2021 [275] |
Gaussian processes (GP) |
VDJdb, publication Classification probability (indicating whether a given TCR recognizes a given epitope) |
NA AUC = 0.863 |
Leveraging CDR sequence information from both TCRα and TCRβ chains and automatically learning which CDRs are most important for each epitope. |
|
pMTnet 2021 [252] |
LSTM and autoencoder combined through dense network |
McPAS-TCR, VDJdb, PIRD, NetMHCpan, publications TpM binding |
Class I AUC > 0.8 |
Overcomes limitations of clustering methods like GLIPH and TCRdist. |
|
ImRex 2021 [248] |
CNN |
VDJdb TpM binding |
Class I & II NA |
Define a feature representation using pairwise amino acid interaction maps instead of separate embeddings. |
|
TITAN 2021 [250] |
Attention-based NNs pretrained on BindingDB |
VDJdb, publications TpM binding |
NA AUC = 0.78 |
Reformulates the task as compound–protein interaction prediction by representing epitopes as SMILES strings, enabling data augmentation and pretraining on large datasets like BindingDB. |
|
ERGO-II 2021 [247] |
LSTM, AE, MLP |
McPAS-TCR, VDJdb TpM binding |
Class I & II NA |
Uses large-scale TCR-peptide dictionaries. |
|
DLpTCR 2021 [249] |
Ensemble DL framework from FCN, CNN and ResNet |
TetTCR-seq, VDJdb Likelihood of TpM binding |
NA AUC = 0.8564 to 0.9227 |
Shows high accuracy on independent datasets, even when using only a single TCR chain. |
|
TCRAI 2021 [276] |
ANN using ICON (Integrative COntext-specific Normalization) |
VDJdb, McPAS-TCR, 10x Genomics TpM binding |
Class I & II AUC > 0.90 |
ICON solves the low signal-to-noise problem in high-throughput pMHC binding. Multi omics driven background estimation. |
|
TCRMatch 2021 [277] |
BLOSUM62 similarity matrix |
IEDB Similar TCRs, similarity scores, predicted epitopes |
NA PRAUC = 0.737 |
Designed specifically for TCR similarity. Predicts epitope specificity by matching to known receptors. |
|
iTTCA-RF 2021 [278] |
RF |
IEDB, TANTIGEN IM |
Class I AUC = 0.78 |
Hybrid feature representation. Applied MRMD (Maximum Relevance Minimum Redundancy) to rank features. |
|
NeoScore 2022 [279] |
Logistic regression-based | TESLA |
Class I AUC = 0.60–0.83 for four independent test set |
Significant association of the NeoScore with survival in response to immune checkpoint inhibition. https://bordene.shinyapps.io/MHCI_neoantigen_prioritization/ |
|
AttnTAP 2022 [280] |
Dual-input framework combining Attn-BiLSTM and Attn-MLP |
VDJdb, McPAS-TCR TP binding |
NA AUC = 0.83–0.89 |
Avoids overfitting by simplifying architecture. Performs well on unseen TCR–peptide predictions. |
|
ATM-TCR 2022 [281] |
Multi-head self-attention network |
VDJdb, McPAS-TCR, IEDB TP binding |
Class I Recall = 0.71 |
Captures biological contextual relationships shaped by sequence arrangement. Improves out-of-sample performance. |
|
Seq2Neo 2022 [282] |
CNN |
IEDB IM, BA |
Class I Accuracy = 0.75; Precision = 0.96 |
One-stop pipeline from raw FASTQ/BAM to neoepitope features; supports SNVs, INDELs, and gene fusions. |
|
PRIME2.0 2023 [163] |
Neural Network |
Publications, IEDB BA, IM |
Class I AUC = 0.69 |
High-quality training data, correct peptide length modeling, aromatic residue enrichment—especially tryptophan. |
|
epiTCR 2023 [283] |
Random forest and BLOSUM62 encoding |
IEDB, TBAdb [223], VDJdb, McPAS-TCR, and 10x [224] TpM binding |
NA AUC = 0.969 |
Trained on a large and diverse dataset; achieves higher sensitivity without sacrificing specificity. |
|
DePTH 2023 [284] |
Combination of a CNN and a dense layer |
Publications TCR-HLA associations |
Class I & II AUC = 0.64 to 0.69 |
Makes predictions for any TCR-HLA pairs then possible to study rare HLAs. |
|
EPIC-TRACE 2023 [255] |
Convolution and multi-head attention using ProtBERT embedding |
VDJdb, IEDB TpM binding |
NA AUC = 0.69 |
Uses the full TCR information. |
|
MIX-TPI 2023 [285] |
CNN and self-attention fusion layer |
VDJdb, ImmuneCODE, IEDB, McPAS TpM binding |
NA NA |
Multimodal framework to combine sequence-based and physicochemical features. |
|
PanPep 2023 [286] |
Meta learning and Neural Turing machine |
IEDB, VDJdb, PIRD, McPas-TCR TP binding |
NA AUC = 0.67 |
Allows accurate prediction of TCR binding to unseen peptides. |
|
POP-UP TCR 2023 [287] |
RF |
PDB, IMGT-numbered structure files from the Structural T Cell Receptor Database (STCRDab), McPAS TCR TP binding |
Class I AUC = 0.5 |
Models trained using only TCR beta chains perform well. |
|
STAPLER 2023 [253] |
BERT |
IEDB, Francis, 10 × 35, McPas, VDJdb, publications TpM binding |
NA NA |
Amino Acid masking in fine tuning. |
|
TABR-BERT 2023 [254] |
BERT |
TCRdb, IEDB, McPAS, VDJdb, PIRD TpM binding |
NA AUC = 0.84 |
Leverages large-scale unlabeled databases and superior performance on unseen epitopes. |
|
TCR-ESM 2023 [256] |
ESM1v |
Datasets from netTCR2.0, ERGO II, pMTnet TpM binding |
NA NA |
Embeddings from protein language models improved model efficiency and performance over traditional methods like CNNs or AEs. |
|
TCR-Pred 2023 [288] |
SAR using MNA with MultiPASS |
VDJdb, McPASTCR, IEDB TpM binding |
NA NA |
Atom-centered substructural MNA descriptors instead of the traditional amino acid (AA) one-letter codes. |
|
TCRmodel2 2023 [257] |
AlphaFold v2.3.0-based DL |
AlphaFold’s database TpM binding |
NA NA |
Higher accuracy than AlphaFold for TCR–pMHC structures, runs faster, no need for fine-tuning or templates. |
|
iTCep 2023 [289] |
Deep CNN with fusion features |
McPAS-TCR, VDJdb, IEDB TP binding |
Class I AUC = 0.86 and 0.91 on two datasets |
DL framework for TCR–epitope recognition using fusion of a novel amino-acid pair propensity (AAPP) interaction map with one-hot encoding; strong generalization to unseen peptides and imbalanced data. |
|
BERTrand-peptide 2023 [290] |
Transformer (BERT) model with MLM |
VDJdb, McPAS, TBAdb, 10x Genomics, publication TP binding |
Class I AUROC ≈ 0.69 |
BERT model over concatenated peptide+TCR β CDR3 sequences with token, position, and type embeddings; MLM pre-training on synthetic peptide. |
|
NeoRanking 2023 [291] |
Voting classifier (LR and XGBoost) |
NCI, TESLA, HiTIDE (in-house dataset) IM |
Class I NA |
Integrates diverse biological features especially BA, stability, RNA expression, immunopeptidome evidence, and oncogenicity. |
|
DeepNeo-v2 2023 [162] |
CNN |
IEDB BA, IM |
Class I & II AUC (MHC I) = 0.76 AUC (MHC II) = 0.80 |
Designed to be simple, fast, and user-friendly through a publicly accessible web service, enabling broad adoption. |
|
BigMHC 2023 [161] |
An ensemble of seven pan-allelic deep neural networks (LSTM) |
MHCflurry-2.0, NetMHCpan-4.1, PRIME-1.0, PRIME-2.0 EL, IM |
Class I EL AUC = 0.58 IM AUC = 0.52–0.55 |
BigMHC-EL (presentation model) and BigMHC-IM (immunogenicity model) created by transfer learning. |
|
VitTCR 2024 [242] |
Vision Transformer (ViT) |
IEDB, VDJdb, McPAS, TCRdb, 10x TP binding |
NA NA |
Integrates a positional bias weight matrix (PBWM) with output of 3-dimensional numeric tensor named AtchleyMaps. |
|
MixTCRpred 2024 [241] |
Transformer |
VDJdb, IEDB, McPAS, 10x, and publications TpM binding |
Class I & II AUC = 0.891 |
A quality control resource for analyzing single-cell TCR sequencing data. |
|
NetTCR-2.2 2024 [251] |
CNN |
IEDB, VDJdb, 10x TpM binding |
Class I AUC = 0.8476 |
Combining pan-specific and peptide-specific models with similarity-based predictions. |
|
HeteroTCR 2024 [292] |
Heterogeneous GNN |
IEDB, VDJdb, McPAS-TCR TpM binding |
NA NA |
First GNN-based TpM binding predictor, superior performance compared to SotA. |
|
CATCR 2024 [293] |
Convolutional-self-attention (hybrid framework of CATCR-D: discriminative model and CATCR-G: generative model) |
VDJdb, IEDB, McPAS-TCR TpM binding |
NA AUC = 0.89 |
Integrates sequence and structure-based data using RCM (residue contact matrix). |
|
ImmugenX 2024 [294] |
A modular transformer-based protein language model | BA, EL, stability, IM, TpM binding |
Class I AUC = 0.666 |
Peptide-MHC multitask pretraining, interpretable, fast stability predictions. 10.5281/zenodo.13850954 |
|
MHLAPre 2024 [295] |
Transformer-based Meta-Learning (MHLAPre-IM, MHLAPre-TT |
IEDB IM, TpM binding |
Class I MHLAPre TT performance score = 0.8953 |
TCR-pHLA binding (MHLAPre-TT) transfer-learn from pHLA immunogenicity (MHLAPre-IM). Dynamic sampling of support/query sets to reduce bias. BLOSUM62 is employed to encode antigenic peptides and HLA alleles. |
|
Sa-TTCA 2024 [296] |
SVM |
IEDB, TANTIGEN 2.0 IM |
Class I NA |
Includes a principled framework for statistical evaluation of selected features to improve reliability of feature selection. |
|
MHLAPre 2024 [295] |
DL + Meta Learning + Transfer Learning |
IEDB, VDJdb, McPAS-TCR BA, IM |
Class I NA |
Meta-learning improved performance and generalization in epitope IM prediction. |
|
NetTCR-struc 2025 [258] |
Geometric Vector Perceptron Graph Neural Network (GVP-GNN) |
IEDB, VDJdb, 10x, RCSB Docking quality score and binder vs. non-binder classification |
Class I AUC₀.₁ = 0.487–0.6 |
Distinguishes binding vs. non-binding complexes in zero-shot settings. Struggles when the structural models are inaccurate. |
|
TRAP 2025 [243] |
Contrastive learning framework |
VDJdb, McPASTCR, IEDB, publication TpM binding |
NA AUC = 0.75 |
Captures both cross-reactivity and specificity among TCRs. |
|
LightCTL 2025 [297] |
Contrastive learning framework with context-aware prompt module (CAPM) |
McPAS-TCR, YFV TpM binding |
Class I AUC = from 0.7767 to 0.9445 on eight independent datasets |
CAPM was developed to identify key features linked to T cell activation, antigen recognition, and specific diseases by weighing the significance of different extracted feature maps. |
|
deepAntigen 2025 [298] |
Graph Convolutional Network |
IEDB, publications, STCRDab, MHC Motif Atlas BA, IM |
Class I & II AUC = 0.71 |
T cell antigen identification at the atomic level. |
|
UniPMT 2025 [299] |
GNN |
BigMHC, PanPep, DLpTCR, NetTCR, ERGO, pMTnet, IEDB TpM binding, TP binding, BA |
Class I AUC = 0.7214 |
Handles peptide–MHC–TCR, peptide–MHC, and peptide–TCR binding prediction within a single model. |
|
NeoTImmuML 2025 [300] |
A weighted ensemble model of LightGBM, XGBoost, and RF |
TumorAgDB1.0, latest publications (2024–2025) IM |
Class I & II AUC = 0.885 |
Built on an upgraded curated database (TumorAgDB2.0). Multi-model ML evaluation. Model interpretability using SHAP. |
TpM: TCR-peptide-MHC, TP: TCR-Peptide, Attn-BiLSTM: attention-based bi-directional LSTM
Recommendations: IM prediction tools can be broadly classified into two categories based on their input requirements. The first category includes methods that do not require TCR sequence information; tools such as PRIME2.0 and DeepNeo fall into this group and can be recommended due to their relatively improved predictive performance. In contrast, the second category comprises approaches that rely on the availability of TCR sequence data to predict IM; in this context, tools such as pMTnet, TITAN, and NetTCR-2.2 leverage deep neural networks and large-scale repertoire datasets to model TCR specificity. When analyzing large-scale TCR repertoire data or identifying antigen-specific T cell clusters, similarity- and clustering-based tools such as GLIPH and TCRdist provide interpretable and computationally efficient solutions. More recent transformer and protein language model-based approaches, including MixTCRpred and ImmugenX, offer improved generalization to unseen epitopes but may require substantial computational resources and careful validation. Overall, no single model is universally optimal; therefore, users are encouraged to select tools based on their specific application and, where possible, adopt ensemble or complementary strategies to improve robustness and reduce bias.
Integrative pipelines
Integrative neoantigen pipelines provide a reproducible, end-to-end in silico workflow that unifies somatic variant discovery, expression assessment, HLA typing, peptide processing and binding prediction, and IM or TCR-recognition analyses and, in some cases, integrates immunopeptidomic data to generate neoantigen predictions and prioritize patient-specific targets (Table 8).
Table 8.
Integrative tools for neoantigen prediction and prioritization
| Model/Year | Epitope Prediction Model | MHC Class | Input Data/Key Features/Web and Code Accessibility | Ref. |
|---|---|---|---|---|
|
INTEGRATE-neo 2016 |
NetMHC4.0 | Class I |
The pipeline consists of three steps: gene fusion peptide prediction, HLA allele prediction, and neoantigen identification. Its first step uses the human reference genome (FASTA), gene models (GenePred), and INTEGRATE-predicted gene fusions (BEDPE) to generate fusion peptide sequences Stand-alone module for gene fusion neoantigen discovery. |
[301] |
|
FRED2 2016 |
NetMHC(pan)-(I/II), NetChop, NetCTL, PickPocket | Class I & II |
FASTA file as input Unified Python API for epitope prediction pipelines, including polymorphic proteins; consistent I/O for many predictors; supports epitope selection and vaccine design. |
[302] |
|
pVAC-Seq 2016 |
NetMHC 3.4 | Class I |
WGS of tumor-normal pairs and RNA-seq as input End-to-end genome-guided neoantigen discovery from NGS data. Integrates HLA typing, somatic variant calling, epitope prediction, binding-based filtering, and expression/coverage filters; Designed for clinical-grade workflows for personalized cancer vaccines. |
[303] |
|
neoepitope 2017 |
NetMHCcons v1.1 | Class I |
WGS and RNA-seq Analyzes neoepitopes arising from somatic missense mutations and gene fusions. |
[304] |
|
MuPeXI 2017 |
NetMHCpan3.0 | Class I |
Somatic mutation calls (VCF file), HLA types, and optionally tumor gene expression Similar functionality to pVac-Seq but with enhanced usability; provides mutant allele frequency when the variant calls come from MuTect or MuTect2. |
[305] |
|
TIminer 2017 |
NetMHCpan-3.0 | Class I |
RNA-seq FASTQ and somatic mutation data (VCF) HLA typing, class I neoantigen prediction, immune infiltration characterization, and tumor immunogenicity quantification from NGS. |
[306] |
|
CloudNeo 2017 |
NetMHCpan-3.0 | Class I |
VCF file (for mutations) and bam file (for HLA typing) as inputs Cloud-based CWL workflow; end-to-end patient-specific neoantigen identification on the cloud; integrates HLA typing and NetMHCpan binding prediction. |
[307] |
|
Vaxrank 2017 |
NetMHC, NetMHCpan, NetMHCcons, MHCflurry, web-based predictors through IEDB | Class I |
Tumor Mutations (VCF), Tumor RNA-Seq (BAM), Patient MHC Alleles as input Vaxrank is under the Apache 2.0 open source license and can also be installed from the Python Package Index. |
[308] |
|
Epidisco 2017 |
NetMHCcons | Class I |
Tumor/normal RNAseq Typed EDSL, modular workflows, cloud backends, reproducibility, bioinformatics tool catalog. |
[309] |
|
Neopepsee 2018 |
Locally weighted naive Bayes (LNB) | Class I |
RNA-Seq FASTQ format as input Automates full NGS-based workflow; uses 9 key features (IC50, percentile rank, NetCTLpan scores, T cell IM score, hydrophobicity, polarity/charge, DAI, AAPP, similarity to known pathogenic epitopes). |
[310] |
|
Rubinsteyn1 et al. 2018 |
NetMHCpan | Class I |
WGS/WES of tumor-normal pairs and RNA-seq as input End-to-end workflow from FASTQ → somatic variants → RNA support → HLA typing → Vaxrank-based prioritization. |
[311] |
|
pTuneos 2019 |
RF | Class I |
Tumor/normal WES/WGS and RNA-seq as input Refined IM score shown to be a pan-cancer survival marker and better predictor of ICI response than TMB / simple neoantigen load. |
[312] |
|
Antigen.garnish 2019 |
netMHCI/II, netMHCI/IIpan, MHCnuggets and MHCflurry | Class I & II |
VCFs, peptide sequences, cDNA transcripts Dissimilarity to human proteome is a strong predictor. Supports human and mouse data; in addition to IM, it also predicts clinical outcomes. |
[313] |
|
NeoPredPipe 2019 |
netMHCpan | Class I & II |
Single and multi-region variant call format (VCF) files Multi-sample input, integration of clonal architecture and immunogenicity (recognition potential). |
[314] |
|
ScanNeo 2019 |
NetMHC, NetMHCpan | Class I |
RNA-Seq data in BAM format aligned as input Specifically targets indel neoantigens from RNA-seq, complementing DNA-based pipelines; uses RNA evidence to focus on expressed, frame-shifted peptides. |
[315] |
|
Neoepiscope 2020 |
MHCnuggets, MHCflurry, NetMHCpan, NetMHCIIpan |
Class I & II |
Tumor/normal DNA-seq input Key strength is explicit multi-variant phasing (germline + somatic) and support for custom references; open-source MIT-licensed command-line tool. |
[316] |
|
NeoFuse 2020 |
MHCflurry | Class I |
RNA-seq FASTQ input In fusion-caller benchmarking, Arriba + STAR-Fusion gave the best performance in validated fusions while limiting the number of called fusions. Focus on fusion neoantigens; fully containerized (docker and singularity); single-sample workflow; automatically annotates each candidate with binding and expression metrics. |
[317] |
|
OpenVax 2020 |
NetMHCpan 4.0, NetMHCcons 1.1, SMM, SMMPMBEC | Class I |
Tumor/normal DNA-seq and tumor RNA-seq as input Dockerized end-to-end pipeline. |
[318] |
|
pVACtools 2020 |
NetMHC, NetMHCpan, MHCflurry | Class I & II |
Somatic mutations (VCFs), RNA-seq End-to-end neoantigen workflow including vector design (pVACvector) and interactive review (pVACview); extensive documentation, tutorials, Docker images and PyPI package. Modular framework (pVACseq, pVACbind, pVACfuse, pVACsplice, pVACvector, pVACview). |
[319] |
|
neoANT-HILL 2020 |
IEDB, MHCflurry | Class I & II |
Handles WES/WGS and RNA-seq as input Dockerized pipeline with Flask GUI; supports RNA-seq–only workflows; GUI; computes DAI and immune-cell composition. |
[320] |
|
TruNeo 2020 |
NetMHCpan, MHCflurry | Class I |
Tumor and normal WES FASTQ, RNA-seq FASTQ as input Recall of immunogenic neoantigens among top-10 predictions = 52.63%. Explicitly models many biological steps (binding, processing, expression, tumor heterogeneity, clonality, HLA LOH). |
[321] |
|
NeoFox 2021 |
NetMHCpan, MixMHCpred, PHBR-I, NetMHCIIpan, MixMHC2pred, PHBR-II | Class I & II |
Neoantigen candidate sequence, its corresponding WT sequence and gene name as input NeoFox is for annotation, not classification; easy to use python package (NEOantigen Feature toolbOX); integrates many predictors and features; designed to plug into existing pipelines. |
[322] |
|
TSNAD 2021 |
DeepHLApan1.1 | Class I |
WGS/WES of tumor-normal pairs and RNA-seq Supports multiple reference genome versions for mutation calling and gene fusion analysis. |
[323] |
|
TSAFinder 2022 |
netMHCpan4.0 | Class I |
RNAseq FASTQ files for matched tumor and control RNA-seq only pipeline; translates every RNA-seq read into all possible 8–11mer peptides; HLA typing from RNA-seq. |
[324] |
|
nextNEOpi 2022 |
pVACseq, NetMHCpan and MHCflurry, NeoFuse | Class I & II |
FASTQ or BAM of normal-tumor WES or WGS and optionally RNA-seq Quantifies patient- and neoepitope-specific attributes linked to tumor immunogenicity and therapy response. |
[325] |
|
ProGeo-Neo v2.0 2022 |
NetMHCpan, NetMHCIIpan | Class I & II |
FASTQ of paired normal-tumor WGS/WES and tumor RNA-seq data Presents a proteogenomic approach that combines HLA presentation analysis with direct identification of mutant peptides through MS data. |
[326] |
|
TSNAdb v2.0 2023 |
DeepHLApan, MHCflurry, NetMHCpan 4.0 | Class I |
Somatic SNVs, INDELs, and gene fusions from TCGA; experimentally validated neoantigens Stringent multi-tool criteria to reduce false positives; coverage of three mutation types (SNVs, INDELs, fusions) with per-mutation neoantigen counts; identification of shared neoantigens recurring across patients. |
[327] |
|
PGNneo 2023 |
NetMHCpan-4.1 | Class I |
RNA-seq profiles and MS datasets as input Extends neoantigen discovery to noncoding regions via proteogenomics; four modules (noncoding variant calling and HLA typing, peptide extraction and DB construction, variant peptide ID via MS, neoantigen prediction and prioritization); designed to reduce false positives by requiring peptide-level MS support. |
[328] |
|
LENS 2023 |
NetMHCpan, NetMHCstabpan, MHCFlurry, DeepHLAPan | Class I |
FASTQ of paired normal-tumor DNA and RNA-sequencing data Modular, extensible, multi-workflow neoantigen prediction platform built on top of Nextflow DSL2, broader tumor antigen coverage. https://gitlab.com/landscape-of-effective-neoantigens-software |
[329] |
|
Neo-intline 2023 |
NetMHCpan 4.0, NetMHCIIpan, Uses similarity to known T cell epitopes (IEDB) via BLASTP | Class I & II |
Full simulation of T cell epitope presentation, including proteasome processing, TAP transport, MHC binding, TCR recognition, and a unified scoring system producing ranked neoantigen candidates. |
[330] |
|
NeoHunter 2024 |
NetMHCpan, NetMHCstabpan, ERGO | Class I |
RNA-seq and/or WES/WGS Detects not only SNV- and indel-derived neoantigens but also gene fusion- and aberrant splicing-derived neoantigens; predicts TCR recognition both indirectly (via agretopicity, foreignness) and directly (DL on TCR–pMHC data). |
[331] |
|
ImmuneMirror 2024 |
RF | Class I & II |
FASTQ of matched normal-tumor WES and tumor bulk RNA-seq (optional) Generates a visual analysis report for each sample; AUC = 0.87. |
[332] |
ImmuneMirror is an ML-driven integrative pipeline that embeds a balanced random-forest model, trained on experimentally validated neopeptides, within a stand-alone workflow and web server for neoantigen prediction and prioritization [332]. The ImmuneMirror workflow processes matched tumor-normal exome data, optionally combined with tumor RNA-seq, to call somatic variants, infer HLA class I and II alleles, evaluate microsatellite instability, and compute composite neoantigen scores, which were associated with clinical outcomes in large gastrointestinal cancer cohorts [332]. Proteogenomic pipelines such as ProGeo-Neo v2.0 extend this concept by integrating whole-genome or exome sequencing, RNA-seq and LC-MS/MS immunopeptidomics and by supporting both MHC class I and II neoantigen prediction in a single one-stop software environment [326]. By allowing detection of multiple mutation classes, including single-nucleotide variants, small insertions or deletions, frameshifts and gene fusions, while filtering peptide candidates against MS-confirmed mutant peptides, ProGeo-Neo v2.0 enriches for high-confidence, naturally presented neoepitopes [326]. NeoFuse focuses specifically on gene-fusion-derived neoantigens, providing an RNA-seq-based pipeline that predicts fusion transcripts, translates them into chimeric peptides, infers patient MHC class I types, and scores peptide-MHC binding to nominate fusion neoepitopes [317]. Because fusion neoantigens can generate immunogenic neoantigens that mediate anticancer immune responses and expand the neoantigen repertoire, NeoFuse helps extend integrative approaches beyond point mutations and small indels [317]. nextNEOpi builds on existing modules such as pVACseq and NeoFuse to offer a comprehensive workflow that accepts tumor-normal WES/WGS (and optional RNA-seq) as input and jointly predicts SNV-, indel- and fusion-derived class I and II neoepitopes [325]. In addition to neoepitope prediction, nextNEOpi incorporates tumor purity and other tumor-intrinsic features, enabling downstream association of neoantigen burden and quality with immune contexture and therapy response [325]. NeoHunter further generalizes integrative neoantigen discovery by providing flexible software for systematically detecting candidate neoantigens from sequencing data and combining support for multiple mutation types with HLA-binding and TCR-recognition models [331].
Collectively, these approaches integrate the individual components described in previous sections into practical, scalable workflows that can be deployed in translational studies and, with appropriate validation, in personalized neoantigen vaccine design.
Recommendations: Users should begin by evaluating their available datasets and research goals, such as whether they have RNA-seq alone or paired genomic and transcriptomic data, and which types of neoantigens (SNVs, indels, or fusions) they want to investigate. For general use, well-documented and end-to-end pipelines like pVACtools are a good starting point, while more specialized tools may be better for specific research needs. It is also advisable to combine multiple prediction methods and include additional evidence like gene expression to improve accuracy. Finally, users should consider ease of use, computational requirements, and reproducibility when selecting a tool.
Immunopeptidomics and AI-driven neoantigen discovery
MS-based immunopeptidomics directly identifies MHC-associated peptides eluted from tumor HLA molecules, yielding large catalogs of patient-presented ligands for neoantigen discovery [40, 49, 333].
Integration of genomic sequencing with HLA ligandome profiling has enabled direct identification of mutated peptide ligands on native tumor tissue and demonstration of their immunogenicity in patients [40, 49, 91]. Pioneering proteogenomic studies combined exome/transcriptome sequencing with MS to discover neo-epitopes, showing that only a small fraction of variants are confirmed at the ligand level [91].
Analyses of the tumor antigenic repertoire emphasize that immunopeptidomics complements genomics by empirically defining the set of tumor-associated antigens and neoantigens visible to T cells [41, 49]. Large-scale immunopeptidome datasets from mono-allelic model systems and diverse tumor types have become foundational training resources for AI-based peptide–MHC presentation models [11, 49, 188, 192, 208]. In mono-allelic systems, Abelin et al. profiled HLA-associated peptidomes and showed that incorporating MS-derived ligands into prediction algorithms improves epitope prediction compared with affinity-based methods [188]. Several studies combined hundreds of thousands of MS-identified ligands with ML architectures, substantially improving positive predictive value for class I ligand prediction and enabling improved neoantigen prioritization across diverse alleles [11, 192]. Building on these resources, Pyke et al. integrated large tumor immunopeptidomes with composite presentation models to demonstrate “precision neoantigen discovery” in clinically relevant cohorts [208].
State-of-the-art MS-trained models such as HLAthena, MHCflurry 2.0, BigMHC, and MaNeo integrate peptide sequence and HLA context, achieving higher accuracy than binding-affinity-only predictors in cross-validation and benchmarking studies [11, 161, 204, 208, 334]. These composite presentation scores correlate with tumor immunopeptidomes and improve identification of clinically relevant neoantigens in benchmarking studies and patient cohorts [161, 192, 208, 335]. For HLA class II, the multimodal MARIA framework integrates MS-identified ligands with expression and cleavage features to improve neoantigen identification performance and enrich for CD4+ responses [160].
Proteogenomic workflows extend AI-based immunopeptidomics by searching spectra against customized protein databases that include variant peptides, alternative splice isoforms, gene fusions, and non-coding translational events [48, 49, 91, 326, 333]. This strategy enables detection of non-canonical tumor antigens, including phosphopeptides, splice-derived peptides and cryptic antigens from noncoding regions, which can elicit strong T cell responses [40, 48, 49, 91]. In particular, these antigen sources correspond to multiple classes of entries summarized in Table 9, including gene fusion–derived neoantigens (e.g., FusionGDB [336], TumorFusions [337], FusionNeoAntigen [338]), RNA splicing–derived isoforms (e.g., FLIBase [339]), and non-coding region–derived peptides (e.g., IEAtlas [340]), highlighting their relevance for systematic cataloging of non-SNV cancer neoantigens. Software such as ProGeo-Neo v2.0 and pVACtools implement end-to-end pipelines in which MS-confirmed ligands are combined with AI-based predictors to filter, score and prioritize neoantigen candidates for vaccines or T cell therapies, while best-practice guidelines and resources like the HLA Ligand Atlas address current bottlenecks such as incomplete ligandomes and under-representation of rare HLA alleles [47, 49, 238, 319, 326, 335].
Table 9.
Databases cataloging non-SNV cancer neoantigens
| Neoantigen Source | Category | Database | Description | Web Accessibility | Ref. |
|---|---|---|---|---|---|
| DNA alterations | Gene fusion | FusionGDB | Compiles ~48K cancer fusion genes and provides functional annotations including ORF prediction, protein domain retention, sequences, and clinical relevance to support biomarker and therapeutic discovery in cancer. | https://compbio.uth.edu/FusionGDB/ | [336] |
| Gene fusion | TumorFusions | A TCGA-based database of ~20,000 gene fusions across cancers, validated by genomic rearrangements and enriched with functional annotations for prioritizing cancer-relevant fusion events. | http://www.tumorfusions.org/ | [337] | |
| Gene fusion | FusionNeoAntigen | A comprehensive resource for fusion gene-derived neoantigens, including fusion protein sequences, HLA binding predictions, and potential vaccine or CAR-T targets. | https://compbio.uth.edu/FusionNeoAntigen | [338] | |
| RNA aberrations | Alternative splicing | FLIBase | A long-read–based transcriptome database that catalogs full-length splice isoforms, including many novel and tumor-specific transcripts with potential neoantigen relevance. | http://www.flibase.org/ | [339] |
| Non-Coding genomic regions | IEAtlas | A database of HLA-presented immunogenic epitopes derived from non-coding regions, integrating MS data and immunogenicity features to support cancer vaccine and immunotherapy research. | http://bio-bigdata.hrbmu.edu.cn/IEAtlas | [340] | |
| Transposable elements | TEITbase | A database of transposable element (TE)-initiated transcripts across 33 cancer types, including 6,203 tumor-specific transcripts, onco-exaptation events, and candidate tumor-specific antigens derived from TE activation. | http://teitbase.medbioinfo.org/ | [341] | |
| Post-translational modifications (PTMs) | Modified tumor antigens / PTM-derived HLA ligands | caAtlas | caAtlas is a human cancer immunopeptidome resource that catalogs MHC/HLA-presented peptides identified from cancer immunopeptidomic datasets, including modified peptides and PTM-associated tumor antigens, supporting prioritization of candidate peptides for immunogenicity testing and cancer immunotherapy development. | http://www.zhang-lab.org/caatlas/ | [342] |
Collectively, integrating WES/RNA-seq with MS/MS triangulates evidence toward “valid HLA-restricted tumor antigens” and guides rigorous target prioritization for vaccines, TCR-T, and other antigen-directed therapies [41]. Clinically, MS-defined ligands, including post-translationally modified phosphopeptides, have advanced into first-in-human vaccination, demonstrating safety and immunogenicity in patients with high-risk melanoma [43].
Personal neoantigens vs. shared neoantigens
Personal (private) neoantigens arise uniquely in individual tumors or subclones. Their mosaic expression contributes to intra-tumor heterogeneity, complicating durable immune targeting. Vaccines based on these neoantigens require complex workflows—tumor sequencing, epitope prediction, immunopeptidomics, and personalized GMP manufacturing—making them costly, labor-intensive, and prone to false positives in epitope prediction. While clinical results have been encouraging, challenges include immune evasion by subclones lacking the targeted antigens and uncertainty about the proportion of tumor cells that must express the neoantigen for therapeutic efficacy [343]. In contrast, shared (public) neoantigens derive from recurrent driver mutations (e.g., TP53, KRAS, BRAF, EGFR, HER2, PI3K, TARP) present across many patients. Unlike personal neoantigens, they are typically clonal, avoid issues of mosaicism, and are less susceptible to immune escape. They can serve as “off-the-shelf” immunotherapies with broader applicability, provided their epitopes are appropriately presented in the patient’s HLA context. Similar to personal neoantigens, they are tumor-specific and generally pose minimal risk of autoimmunity [344, 345]. Targeting shared neoantigens in cancer immunotherapy remains difficult but offers considerable promise due to their tumor specificity and wide therapeutic relevance.
Over 150 clinical trials investigating neoantigen-based immuno-oncotherapies have been launched [346]. Notably, a clinical trial investigating TARP-based (shared neoantigen) vaccination in prostate cancer demonstrated significant tumor growth reduction in the majority of patients [347]. Similarly, a phase 1 clinical trial tested TP53 and KRAS vaccines in patients with advanced solid tumors, showing encouraging results [348].
Li et al. [349] quantified Tumor Mutational Burden (TMB) across 30 cancer types. Skin cutaneous melanoma (SKCM) showed the highest proportion of high-TMB tumors (49.4%). Lung cancers ranked next, with lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC) displaying 36.9% and 28.1% high-TMB cases, respectively (Fig. 4). In tumors with a low mutational burden, vaccination based on personalized neoantigens may have limited efficacy. Instead, the use of shared neoantigens represents a potential therapeutic strategy.
Fig. 4.

Comparison of tumor mutational burden (TMB) across different cancer types. The bar plot illustrates the proportion of patients with high TMB in various cancer types. The highest rates of elevated TMB are observed in skin cutaneous melanoma (SKCM, 49%), lung adenocarcinoma (LUAD, 37%), and lung squamous cell carcinoma (LUSC, 28%), followed by bladder carcinoma (BLCA, 26%) and uterine corpus endometrial carcinoma (UCEC, 23%). In contrast, several cancer types, including kidney renal papillary cell carcinoma (KIRP), thyroid carcinoma (THCA), pheochromocytoma and paraganglioma (PCPG), kidney renal clear cell carcinoma (KIRC), and mesothelioma (MESO), show negligible or absent frequencies of high TMB. This comparison highlights the heterogeneity of mutational load among distinct cancer types, which may have implications for immunotherapy response
Challenges and future directions
Despite rapid progress in AI-guided neoantigen discovery, method performance and interoperability still lag the field’s clinical ambition; notably, there remains no consensus analysis workflow and few established best practices, underscoring the need for standardized datasets and reporting across the pipeline from variant calling to immunogenicity testing [335].
Consortium-scale benchmarking highlights both promise and limitations; for instance the TESLA study showed that integrating presentation- and recognition-relevant features filtered out 98% of non-immunogenic peptides with a precision above 0.70, while noting that, prior to the consortium effort, there were no common reference datasets to systematically compare prediction methods [12].
Data choice and modeling remain core bottlenecks. Training directly on endogenous HLA ligands from tumors and incorporating contextual features can markedly improve performance; EDGE increased the positive predictive value of HLA antigen prediction by up to ninefold [192], but the magnitude of gain varies by allele, tumor type, and evaluation regime. Likewise, large mono-allelic immunopeptidome resources plus integrative modeling (HLAthena) achieved a 1.5-fold improvement in positive predictive value and correctly identified > 75% of HLA-bound peptides observed in patient-derived tumor cell lines [11], advancing allele coverage yet still leaving gaps in rare alleles and class II presentation. Mono-allelic immunopeptidomics coupled to neural networks outperformed affinity-trained algorithms, illustrating the value of training on EL data and offering a scalable strategy to learn endogenous presentation rules [188].
Validation and antigen source completeness are parallel challenges. Foundational proteogenomic studies showed that among missense-derived candidates only a small fraction of predicted binders were confirmed by MS, highlighting the need for rigorous, prospective orthogonal validation at scale [91]. Moreover, ~90% of tumor-specific antigens in one screen arose from allegedly noncoding regions and would have been missed by standard exome-centric pipelines, arguing for systematic inclusion of noncanonical ORFs, aberrant transcripts, and proteasome-spliced peptides in training and inference [48]. Tumor evolution compounds these issues; allele-specific HLA loss occurs in ~40% of NSCLCs and is associated with subclonal neoantigen burden, reinforcing the importance of modeling immune escape and intratumoral heterogeneity during prioritization [92].
Accordingly, near-term priorities are clear: harmonized benchmarking, expanded and higher-accuracy HLA class II typing and prediction, and first-class support for alternative antigen sources and post-translationally generated peptides within pipelines [12, 335]. Clinically, personalized neoantigen vaccines have demonstrated feasibility, safety, and immunogenicity in melanoma, providing a scaffold on which to prospectively test whether AI-enhanced selection improves true clinical benefit, and to co-develop adaptive trial designs with TCR therapies and checkpoint blockade [4, 90].
Conclusion
AI-driven neoantigen discovery has evolved into a clinically testable paradigm, enabled by large immunopeptidomics datasets and DL predictors that now inform vaccine design [11, 90, 192]. Mono-allelic MS datasets and MS-informed models have improved presentation prediction and recovery of tumor-displayed ligands [11]. Likewise, DL trained on tumor immunopeptidomes markedly increases positive predictive value for antigen presentation [192]. Consortium-scale benchmarking has begun to clarify parameters of immunogenicity, enabling triage of the vast majority of non-immunogenic peptides at practical precision [11, 12].
Despite rapid progress, there is still no consensus end-to-end approach, and key gaps persist, including improving HLA class II typing accuracy, expanding support for diverse antigen sources, and incorporating clinical response data to refine prediction. Practically, mature pipelines should integrate somatic variant calling, precise HLA typing, peptide processing, and peptide–MHC binding prediction before prioritization and validation.
Anchoring pipelines to ground truth via direct MS of native tumor tissue remains essential to isolate bona fide neoepitopes. Early trials show that personalized neoantigen vaccines are feasible, safe, and immunogenic, eliciting broad T cell responses in melanoma [90]. Key priorities now include robust class II modeling and expansion beyond coding SNVs to noncanonical antigen sources, alongside convergence toward standardized, end-to-end workflows for clinical utility. Coupling presentation and recognition modeling with tumor-evolution constraints, such as frequent HLA loss of heterozygosity, should sharpen prioritization of clinically actionable targets. The field’s trajectory is clear: standardized, MS-informed, and recognition-aware AI models, coupled with rigorous validation, can accelerate rational, personalized vaccines with the ultimate aim of improving patient outcomes.
Acknowledgements
This study was supported by Tehran University of Medical Sciences Grant No. 1404-18-148-94336.
Abbreviations
- AA-index
Amino acid index
- AI
Artificial intelligence
- ANN
Artificial neural network
- APC
Antigen-presenting cell
- AUC
Area under the (ROC) curve
- BiGRU
Bidirectional gated recurrent unit
- BLOSUM
BLOcks SUbstitution Matrix
- BWA
Burrows–Wheeler Aligner
- CDR
Complementarity-determining region
- CLIP
Class II–associated invariant chain peptide
- CNN
Convolutional neural network
- ER
Endoplasmic reticulum
- ESM
Evolutionary Scale Modeling
- GBM
Glioblastoma multiforme
- HLA
Human leukocyte antigen
- HMM
Hidden Markov model
- ICGC
International Cancer Genome Consortium
- IEDB
Immune Epitope Database
- INDEL
Insertion and deletion
- IPD-IMGT/HLA
Immuno Polymorphism Database – International ImMunoGeneTics/Human Leukocyte Antigen database
- LC-MS/MS
Liquid chromatography–tandem mass spectrometry
- LOH
Loss of heterozygosity
- LSTM
Long short-term memory
- MHC
Major histocompatibility complex
- ML
Machine learning
- MS
Mass spectrometry
- MS/MS
Tandem mass spectrometry
- NGS
Next-generation sequencing
- NSCLC
Non-small-cell lung cancer
- ORF
Open reading frame
- PBR
Peptide-binding region
- PFR
Peptide flanking residue
- PPV
Positive predictive value
- PTM
Post-translational modification
- pMHC
Peptide–major histocompatibility complex
- RNA-seq
RNA sequencing
- RNN
Recurrent neural network
- SNV
Single-nucleotide variant
- SVM
Support vector machine
- TAP
Transporter Associated with Antigen Processing
- TARP
TCR gamma alternate reading frame protein
- TCGA
The Cancer Genome Atlas
- TCR
T cell receptor
- TCR-T
T cell receptor–engineered T cell therapy
- TESLA
Tumor Neoantigen Selection Alliance
- TMB
Tumor mutational burden
- TSA
Tumor-specific antigen
- WES
Whole-exome sequencing
- WGS
Whole-genome sequencing
- WHO
World Health Organization
Author contributions
A.B: main idea and conceptualization, review of literature, analysis and interpretation, writing—original draft preparation, Figure preparation; S.G: review of literature, writing—original draft preparation; F.F: writing—original draft preparation; E.E: review of literature, writing—original draft preparation; A.A: writing—original draft preparation; A.E., B.N., K.K., G.K. and M.M: writing—review and editing. All authors have read and approved the final manuscript.
Funding
Not applicable.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Schumacher TN, Schreiber RD. Neoantigens in cancer immunotherapy. Science (New York, NY) [Internet]. 2015 [cited 2025 Dec 12];348(6230):69–74. Available from: https://doi.org/10.1126/science.aaa4971; https://pubmed.ncbi.nlm.nih.gov/25838375/. [DOI] [PubMed] [Google Scholar]
- 2.Raskov H, Orhan A, Christensen JP, Gögenur I. Cytotoxic CD8+ T cells in cancer and cancer immunotherapy. Br J Cancer. 2021 [cited 2025 Dec 12];124(2):359–67. Available from: https://doi.org/10.1038/s41416-020-01048-4; https://pubmed.ncbi.nlm.nih.gov/32929195/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Shifrut E, Carnevale J, Tobin V, Roth TL, Woo JM, Bui CT, et al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell. 2018;175(7):1958–71.e15. 10.1016/j.cell.2018.10.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Sahin U, Derhovanessian E, Miller M, Kloke BP, Simon P, Löwer M, et al. Personalized RNA mutanome vaccines mobilize poly-specific therapeutic immunity against cancer. Nature [Internet]. 2017 [cited 2025 Dec 12];547. Available from: https://pubmed.ncbi.nlm.nih.gov/28678784/. [DOI] [PubMed]
- 5.Samstein RM, Lee CH, Shoushtari AN, Hellmann, Shen R, Janjigian YY, et al. Tumor mutational load predicts survival after immunotherapy across multiple cancer types. Nat Genet. 2019 [cited 2025 Dec 12];51(2):202–06. Available from: https://doi.org/10.1038/s41588-018-0312-8; https://pubmed.ncbi.nlm.nih.gov/30643254/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Zhang Z, Lu M, Qin Y, Gao W, Tao L, Su W, et al. Neoantigen: a new breakthrough in tumor immunotherapy. Front Immunol. 2021 [cited 2025 Dec 12];12. Available from: https://doi.org/10.3389/fimmu.2021.672356; https://pubmed.ncbi.nlm.nih.gov/33936118/. [DOI] [PMC free article] [PubMed]
- 7.Sahin U, Türeci Ö. Personalized vaccines for cancer immunotherapy. Science (New York, NY) [Internet]. 2018 [cited 2025 Dec 12];359(6382):1355–60. Available from: https://doi.org/10.1126/science.aar7112; https://pubmed.ncbi.nlm.nih.gov/29567706/. [DOI] [PubMed] [Google Scholar]
- 8.Blass E, Ott PA. Advances in the development of personalized neoantigen-based therapeutic cancer vaccines. Nat Rev Clin Oncol. 2021 [cited 2025 Dec 12];18(4):215–29. Available from: https://doi.org/10.1038/s41571-020-00460-2; https://pubmed.ncbi.nlm.nih.gov/33473220/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Huber F, Arnaud M, Stevenson BJ, Michaux J, Benedetti F, Thevenet J, et al. A comprehensive proteogenomic pipeline for neoantigen discovery to advance personalized cancer immunotherapy. Nat Biotechnol [Internet]. 2025 [cited 2025 Dec 12];43. Available from: https://pubmed.ncbi.nlm.nih.gov/39394480/. [DOI] [PMC free article] [PubMed]
- 10.Meng W, Schreiber RD, Lichti CF. Recent advances in immunopeptidomic-based tumor neoantigen discovery. Adv Immunol [Internet]. 2023 [cited 2025 Dec 12];160. Available from: https://pubmed.ncbi.nlm.nih.gov/38042584/. [DOI] [PMC free article] [PubMed]
- 11.Sarkizova S, Klaeger S, Le PM, Li LW, Oliveira G, Keshishian H, et al. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nat Biotechnol [Internet]. 2020 [cited 2025 Dec 12];38. Available from: https://pubmed.ncbi.nlm.nih.gov/31844290/. [DOI] [PMC free article] [PubMed]
- 12.Wells DK, van Buuren MM, Dang KK, Hubbard-Lucey VM, Sheehan KCF, Campbell KM, et al. Key parameters of tumor epitope immunogenicity revealed through a consortium approach improve neoantigen prediction. Cell [Internet]. 2020 [cited 2025 Nov 25];183. Available from: https://pubmed.ncbi.nlm.nih.gov/33038342/. [DOI] [PMC free article] [PubMed]
- 13.Kallingal A, Olszewski M, Maciejewska N, Brankiewicz W, Baginski M. Cancer immune escape: the role of antigen presentation machinery. J Cancer Res Clin Oncol. 2023 [cited 2025 Dec 12];149(10):8131–41. Available from: https://doi.org/10.1007/s00432-023-04737-8; https://pubmed.ncbi.nlm.nih.gov/37031434/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Vyas JM, Van der Veen AG, Ploegh HL. The known unknowns of antigen processing and presentation. Nat Rev Immunol. 2008 [cited 2025 Dec 12];8(8):607–18. Available from: https://doi.org/10.1038/nri2368; https://pubmed.ncbi.nlm.nih.gov/18641646/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Gannage M, Münz C. MHC presentation via autophagy and how viruses escape from it. Semin Immunopathol. 2010 [cited 2025 Dec 12];32(4):373–81. Available from: https://doi.org/10.1007/s00281-010-0227-7; https://pubmed.ncbi.nlm.nih.gov/20857294/. [DOI] [PubMed] [Google Scholar]
- 16.Neefjes J, Jongsma ML, Paul P, Bakke O. Towards a systems understanding of MHC class I and MHC class II antigen presentation. Nat Rev Immunol. 2011 [cited 2025 Dec 12];11(12):823–36. Available from: https://doi.org/10.1038/nri3084; https://pubmed.ncbi.nlm.nih.gov/22076556/. [DOI] [PubMed] [Google Scholar]
- 17.Balasubramanian A, John T, Asselin-Labat ML. Regulation of the antigen presentation machinery in cancer and its implication for immune surveillance. Biochemical Soc Trans [Internet]. 2022 [cited 2025 Dec 12];50(2):825–37. Available from: https://doi.org/10.1042/BST20210961; https://pubmed.ncbi.nlm.nih.gov/35343573/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Jiang T, Shi T, Zhang H, Hu J, Song Y, Wei J, et al. Tumor neoantigens: from basic research to clinical applications. J Hematol Oncol. 2019 [cited 2025 Dec 12];12(1). Available from: https://doi.org/10.1186/s13045-019-0787-5; https://pubmed.ncbi.nlm.nih.gov/31492199/. [DOI] [PMC free article] [PubMed]
- 19.Xie N, Shen G, Gao W, Huang Z, Huang C, Fu L. Neoantigens: promising targets for cancer therapy. Sig Transduct Target Ther. 2023 [cited 2025 Dec 12];8(1). Available from: https://doi.org/10.1038/s41392-022-01270-x; https://pubmed.ncbi.nlm.nih.gov/36604431/. [DOI] [PMC free article] [PubMed]
- 20.Goulder PJR, Watkins DI. Impact of MHC class I diversity on immune control of immunodeficiency virus replication. Nat Rev Immunol. 2008 [cited 2025 Dec 12];8(8):619–30. Available from: https://doi.org/10.1038/nri2357; https://pubmed.ncbi.nlm.nih.gov/18617886/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Wieczorek M, Abualrous ET, Sticht J, Álvaro-Benito M, Stolzenberg S, Noé F, et al. Major histocompatibility complex (MHC) class I and MHC class II proteins: conformational plasticity in antigen presentation. Front Immunol. 2017 [cited 2025 Dec 12];8. Available from: https://doi.org/10.3389/fimmu.2017.00292; https://pubmed.ncbi.nlm.nih.gov/28367149/. [DOI] [PMC free article] [PubMed]
- 22.Roche PA, Furuta K. The ins and outs of MHC class II-mediated antigen processing and presentation. Nat Rev Immunol. 2015 [cited 2025 Dec 12];15(4):203–16. Available from: https://doi.org/10.1038/nri3818; https://pubmed.ncbi.nlm.nih.gov/25720354/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Liu J, Gao GF. Major histocompatibility complex: interaction with peptides. eLS. John Wiley & Sons, Ltd; 2011. [Google Scholar]
- 24.O’Brien C, Flower DR, Feighery C. Peptide length significantly influences in vitro affinity for MHC class II molecules. Immunome Res. 2008;4(1):6. 10.1186/1745-7580-4-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.McFarland BJ, Katz JF, Beeson C, Sant AJ. Energetic asymmetry among hydrogen bonds in MHC class II⋅peptide complexes. Proc Natl Acad Sci USA. 2001 [cited 2025 Dec 12];98(16):9231–36. Available from: https://doi.org/10.1073/pnas.151131498; https://pubmed.ncbi.nlm.nih.gov/11470892/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Painter CA, Stern LJ. Conformational variation in structures of classical and non-classical MHCII proteins and functional implications. Immunological Rev [Internet]. 2012 [cited 2025 Dec 12];250(1):144–57. Available from: https://doi.org/10.1111/imr.12003; https://pubmed.ncbi.nlm.nih.gov/23046127/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Sarri CA, Giannoulis T, Moutou KA, Mamuris Z. HLA class II peptide-binding-region analysis reveals funneling of polymorphism in action. Immunol Lett [Internet]. 2021 [cited 2025 Dec 12];238. 75–95. Available from: https://doi.org/10.1016/j.imlet.2021.07.005; https://pubmed.ncbi.nlm.nih.gov/34329645/. [DOI] [PubMed] [Google Scholar]
- 28.Peng M, Mo Y, Wang Y, Wu P, Zhang Y, Xiong F, et al. Neoantigen vaccine: an emerging tumor immunotherapy. Mol Cancer. 2019 [cited 2025 Dec 13];18(1). Available from: https://doi.org/10.1186/s12943-019-1055-6; https://pubmed.ncbi.nlm.nih.gov/31443694/. [DOI] [PMC free article] [PubMed]
- 29.Wirth TC, Kühnel F. Neoantigen targeting—dawn of a New Era in cancer immunotherapy? Front Immunol. 2017 [cited 2025 Dec 13];8. Available from: https://doi.org/10.3389/fimmu.2017.01848; https://pubmed.ncbi.nlm.nih.gov/29312332/. [DOI] [PMC free article] [PubMed]
- 30.Moore L, Cagan A, Coorens THH, Neville MDC, Sanghvi R, Sanders MA, et al. The mutational landscape of human somatic and germline cells. Nature [Internet]. 2021 [cited 2025 Dec 13];597(7876):381–86. Available from: https://doi.org/10.1038/s41586-021-03822-7; https://pubmed.ncbi.nlm.nih.gov/34433962/. [DOI] [PubMed] [Google Scholar]
- 31.Veltman JA, Brunner HG. De Novo mutations in human genetic disease. Nat Rev Genet. 2012 [cited 2025 Dec 13];13(8):565–75. Available from: https://doi.org/10.1038/nrg3241; https://pubmed.ncbi.nlm.nih.gov/22805709/. [DOI] [PubMed] [Google Scholar]
- 32.Robbins PF, Lu YC, El-Gamil M, Li YF, Gross C, Gartner J, et al. Mining exomic sequencing data to identify mutated antigens recognized by adoptively transferred tumor-reactive T cells. Nat Med. 2013 [cited 2025 Dec 13];19(6):747–52. Available from: https://doi.org/10.1038/nm.3161; https://pubmed.ncbi.nlm.nih.gov/23644516/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Rosenberg SA, Restifo NP. Adoptive cell transfer as personalized immunotherapy for human cancer. Science (New York, NY) [Internet]. 2015 [cited 2025 Dec 13];348(6230):62–68. Available from: https://doi.org/10.1126/science.aaa4967; https://pubmed.ncbi.nlm.nih.gov/25838374/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Lang F, Schrörs B, Löwer M, Türeci Ö, Sahin U. Identification of neoantigens for individualized therapeutic cancer vaccines. Nat Rev Drug Discov. 2022 [cited 2025 Dec 13];21(4):261–82. Available from: https://doi.org/10.1038/s41573-021-00387-y; https://pubmed.ncbi.nlm.nih.gov/35105974/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Smith CC, Selitsky SR, Chai S, Armistead PM, Vincent BG, Serody JS. Alternative tumour-specific antigens. Nat Rev Cancer. 2019 [cited 2025 Dec 13];19(8):465–78. Available from: https://doi.org/10.1038/s41568-019-0162-4; https://pubmed.ncbi.nlm.nih.gov/31278396/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Gopanenko AV, Kosobokova EN, Kosorukov VS. Main strategies for the identification of neoantigens. Cancers [Internet]. 2020 [cited 2025 Dec 13];12(10):2879. Available from: https://doi.org/10.3390/cancers12102879; https://pubmed.ncbi.nlm.nih.gov/33036391/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Thind AS, Monga I, Thakur PK, Kumari P, Dindhoria K, Krzak M, et al. Demystifying emerging bulk RNA-Seq applications: the application and utility of bioinformatic methodology. Briefings Bioinf [Internet]. 2021 [cited 2025 Dec 13];22(6). Available from: https://doi.org/10.1093/bib/bbab259; https://pubmed.ncbi.nlm.nih.gov/34329375/. [DOI] [PubMed]
- 38.Liu XS, Mardis ER. Applications of immunogenomics to cancer. Cell [Internet]. 2017 [cited 2025 Dec 13];168(4):600–12. Available from: https://doi.org/10.1016/j.cell.2017.01.014; https://pubmed.ncbi.nlm.nih.gov/28187283/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Wu J, Zhao W, Zhou B, Su Z, Gu X, Zhou Z, et al. Tsnadb: a database for tumor-specific neoantigens from Immunogenomics data analysis. Genomics Proteomics Bioinf [Internet]. 2018 [cited 2025 Dec 13];16(4):276–82. Available from: https://doi.org/10.1016/j.gpb.2018.06.003; https://pubmed.ncbi.nlm.nih.gov/30223042/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Bassani-Sternberg M, Bräunlein E, Klar R, Engleitner T, Sinitcyn P, Audehm S, et al. Direct identification of clinically relevant neoepitopes presented on native human melanoma tissue by mass spectrometry. Nat Commun. 2016 [cited 2025 Dec 13];7(1). Available from: https://doi.org/10.1038/ncomms13404; https://pubmed.ncbi.nlm.nih.gov/27869121/. [DOI] [PMC free article] [PubMed]
- 41.Haen SP, Löffler MW, Rammensee HG, Brossart P. Towards new horizons: characterization, classification and implications of the tumour antigenic repertoire. Nat Rev Clin Oncol. 2020 [cited 2025 Dec 13];17(10):595–610. Available from: https://doi.org/10.1038/s41571-020-0387-x; https://pubmed.ncbi.nlm.nih.gov/32572208/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Mohammed F, Cobbold M, Zarling AL, Salim M, Barrett-Wilt GA, Shabanowitz J, et al. Phosphorylation-dependent interaction between antigenic peptides and MHC class I: a molecular basis for the presentation of transformed self. Nat Immunol. 2008;9(11):1236. 10.1038/ni.1660. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Engelhard VH, Obeng RC, Cummings KL, Petroni GR, Ambakhutwala AL, Chianese-Bullock KA, et al. MHC-restricted phosphopeptide antigens: preclinical validation and first-in-humans clinical trial in participants with high-risk melanoma. J Immunother Cancer. 2020 [cited 2025 Dec 13];8(1):e000262. Available from: https://doi.org/10.1136/jitc-2019-000262; https://pubmed.ncbi.nlm.nih.gov/32385144/. [DOI] [PMC free article] [PubMed]
- 44.Meyer VS, Drews O, Günder M, Hennenlotter J, Rammensee HG, Stevanovic S. Identification of natural MHC class II presented phosphopeptides and tumor-derived MHC class I phospholigands. J Proteome Res [Internet]. 2009 [cited 2025 Dec 13];8. Available from: https://pubmed.ncbi.nlm.nih.gov/19415920/. [DOI] [PubMed]
- 45.Zarling AL, Polefrone JM, Evans AM, Mikesh LM, Shabanowitz J, Lewis ST, et al. Identification of class I MHC-associated phosphopeptides as targets for cancer immunotherapy. Proc Natl Acad Sci USA. 2006 [cited 2025 Dec 13];103(40):14889–94. Available from: https://doi.org/10.1073/pnas.0604045103; https://pubmed.ncbi.nlm.nih.gov/17001009/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Engelhard VH, Altrich-Vanlith M, Ostankovitch M, Zarling AL. Post-translational modifications of naturally processed MHC-binding epitopes. Curr Opin Immunol [Internet]. 2006 [cited 2025 Dec 13];18(1):92–97. Available from: https://doi.org/10.1016/j.coi.2005.11.015; https://pubmed.ncbi.nlm.nih.gov/16343885/. [DOI] [PubMed] [Google Scholar]
- 47.Zhou C, Zhu C, Liu Q. Toward in silico identification of tumor neoantigens in immunotherapy. Trends Mol Med [Internet]. 2019 [cited 2025 Dec 13];25(11):980–92. Available from: https://doi.org/10.1016/j.molmed.2019.08.001; https://pubmed.ncbi.nlm.nih.gov/31494024/. [DOI] [PubMed] [Google Scholar]
- 48.Laumont CM, Vincent K, Hesnard L, Audemard É, Bonneil É, Laverdure JP, et al. Noncoding regions are the main source of targetable tumor-specific antigens. Sci Transl Med. 2018 [cited 2025 Dec 13];10(470). Available from: https://doi.org/10.1126/scitranslmed.aau5516; https://pubmed.ncbi.nlm.nih.gov/30518613/. [DOI] [PubMed]
- 49.Zhang X, Qi Y, Zhang Q, Liu W. Application of mass spectrometry-based MHC immunopeptidome profiling in neoantigen identification for tumor immunotherapy. Biomed Pharmacother. 2019 [cited 2025 Dec 13];120:109542. Available from: https://doi.org/10.1016/j.biopha.2019.109542; https://pubmed.ncbi.nlm.nih.gov/31629254/. [DOI] [PubMed] [Google Scholar]
- 50.Wingett SW, Andrews S. FastQ screen: a tool for multi-genome mapping and quality control. F1000Research [Internet]. 2018 [cited 2025 Dec 13];7:1338. Available from: https://doi.org/10.12688/f1000research.15931.2; https://pubmed.ncbi.nlm.nih.gov/30254741/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Li H, Durbin R. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics (Oxford, England) [Internet]. 2010 [cited 2025 Dec 13];26(5):589–95. Available from: https://doi.org/10.1093/bioinformatics/btp698; https://pubmed.ncbi.nlm.nih.gov/20080505/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The sequence alignment/map format and SAMtools. Bioinformatics (Oxford, England) [Internet]. 2009 [cited 2025 Dec 14];25(16):2078–79. Available from: https://doi.org/10.1093/bioinformatics/btp352; https://pubmed.ncbi.nlm.nih.gov/19505943/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet. 2011 [cited 2025 Dec 14];43(5):491–98. Available from: https://doi.org/10.1038/ng.806; https://pubmed.ncbi.nlm.nih.gov/21478889/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Koboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, et al. VarScan, 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 2012;22(3):568–76. 10.1101/gr.129684.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Larson DE, Harris CC, Chen K, Koboldt DC, Abbott TE, Dooling DJ, et al. SomaticSniper: identification of somatic point mutations in whole genome sequencing data. Bioinformatics. 2012;28(3):311–17. 10.1093/bioinformatics/btr665. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Garrison E, Marth G. Haplotype-based variant detection from short-read sequencing [Internet]. arXiv [q-bio.GN]. 2012. Available from: 10.48550/ARXIV.1207.3907.
- 57.Wilm A, Aw PPK, Bertrand D, Yeo GHT, Ong SH, Wong CH, et al. LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets. Nucleic Acids Res. 2012;40(22):11189–201. 10.1093/nar/gks918. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Shiraishi Y, Sato Y, Chiba K, Okuno Y, Nagata Y, Yoshida K, et al. An empirical Bayesian framework for somatic mutation detection from cancer genome sequencing data. Nucleic Acids Res [Internet]. 2013 [cited 2025 Dec 14];41(7):e89–89. Available from: https://doi.org/10.1093/nar/gkt126; https://pubmed.ncbi.nlm.nih.gov/23471004/. [DOI] [PMC free article] [PubMed]
- 59.Hansen NF, Gartner JJ, Mei L, Samuels Y, Mullikin JC. Shimmer: detection of genetic alterations in tumors using next-generation sequence data. Bioinformatics. 2013;29(12):1498–503. 10.1093/bioinformatics/btt183. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Christoforides A, Carpten JD, Weiss GJ, Demeure MJ, Von Hoff DD, Craig DW. Identification of somatic mutations in cancer through Bayesian-based analysis of sequenced genome pairs. BMC Genomics. 2013;14(1):302. 10.1186/1471-2164-14-302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Kim S, Jeong K, Bhutani K, Lee J, Patel A, Scott E, et al. Virmid: accurate detection of somatic mutations with sample impurity inference. Genome Biol. 2013;14(8):R90. 10.1186/gb-2013-14-8-r90. [DOI] [PMC free article] [PubMed]
- 62.Radenbaugh AJ, Ma S, Ewing A, Stuart JM, Collisson EA, Zhu J, et al. RADIA: RNA and DNA integrated analysis for somatic mutation detection. PLoS One. 2014;9(11):e111516. 10.1371/journal.pone.0111516. [DOI] [PMC free article] [PubMed]
- 63.Rimmer A, Phan H, Mathieson I, Iqbal Z, Twigg SRF, Wilkie AOM, et al. Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications. Nat Genet. 2014 [cited 2025 Dec 14];46(8):912–18. Available from: https://doi.org/10.1038/ng.3036; https://pubmed.ncbi.nlm.nih.gov/25017105/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Moncunill V, Gonzalez S, Beà S, Andrieux LO, Salaverria I, Royo C, et al. Comprehensive characterization of complex structural variations in cancer by directly comparing genome sequence reads. Nat Biotechnol. 2014 [cited 2025 Dec 14];32(11):1106–12. Available from: https://doi.org/10.1038/nbt.3027; https://pubmed.ncbi.nlm.nih.gov/25344728/. [DOI] [PubMed] [Google Scholar]
- 65.Mose LE, Wilkerson MD, Hayes DN, Perou CM, Parker JS. ABRA: improved coding indel detection via assembly-based realignment. Bioinformatics. 2014;30(19):2813–15. 10.1093/bioinformatics/btu376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Raine KM, Hinton J, Butler AP, Teague JW, Davies H, Tarpey P, et al. cgpPindel: identifying somatically acquired insertion and deletion events from paired end sequencing. Curr Protoc Bioinf. 2015;52(1):15.7.1–15.7.12. 10.1002/0471250953.bi1507s52. [DOI] [PMC free article] [PubMed]
- 67.Jones D, Raine KM, Davies H, Tarpey PS, Butler AP, Teague JW, et al. cgpCavemanwrapper: simple execution of CaVEMan in order to detect somatic single nucleotide variants in NGS data. Curr Protoc Bioinf. 2016;56(1):.15.10.1–15.10.18. 10.1002/cpbi.20. [DOI] [PMC free article] [PubMed]
- 68.Fan Y, Xi L, Hughes DST, Zhang J, Zhang J, Futreal PA, et al. MuSE: accounting for tumor heterogeneity using a sample-specific error model improves sensitivity and specificity in mutation calling from sequencing data. Genome Biol. 2016;17(1):178. 10.1186/s13059-016-1029-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Lai Z, Markovets A, Ahdesmaki M, Chapman B, Hofmann O, McEwen R, et al. VarDict: a novel and versatile variant caller for next-generation sequencing in cancer research. Nucleic Acids Res. 2016;44(11):e108. 10.1093/nar/gkw227. [DOI] [PMC free article] [PubMed]
- 70.Fang H, Bergmann EA, Arora K, Vacic V, Zody MC, Iossifov I, et al. Indel variant analysis of short-read sequencing data with scalpel. Nat Protoc. 2016 [cited 2025 Dec 14];11(12):2529–48. Available from: https://doi.org/10.1038/nprot.2016.150; https://pubmed.ncbi.nlm.nih.gov/27854363/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kim S, Scheffler K, Halpern AL, Bekritsky MA, Noh E, Källberg M, et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat Methods. 2018;15(8):591–94. 10.1038/s41592-018-0051-x. [DOI] [PubMed] [Google Scholar]
- 72.Wala JA, Bandopadhayay P, Greenwald NF, O’Rourke R, Sharpe T, Stewart C, et al. SvABA: genome-wide detection of structural variants and indels by local assembly. Genome Res. 2018;28(4):581–91. 10.1101/gr.221028.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Narzisi G, Corvelo A, Arora K, Bergmann EA, Shah M, Musunuri R, et al. Genome-wide somatic variant calling using localized colored de Bruijn graphs. Commun Biol. 2018;1(1):20. 10.1038/s42003-018-0023-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Freed D, Pan R, Aldana R. Tnscope: accurate detection of somatic mutations with haplotype-based variant candidate detection and machine learning filtering [Internet]. bioRxiv. bioRxiv. 2018. Available from: http://biorxiv.org/lookup/doi/10.1101/250647.
- 75.Benjamin D, Sato T, Cibulskis K, Getz G, Stewart C, Lichtenstein L. Calling somatic SNVs and indels with Mutect2 [Internet]. bioRxiv. bioRxiv. 2019. Available from: http://biorxiv.org/lookup/doi/10.1101/861054.
- 76.Sahraeian SME, Liu R, Lau B, Podesta K, Mohiyuddin M, Lam HYK. Deep convolutional neural networks for accurate somatic mutation detection. Nat Commun. 2019;10(1):1041. 10.1038/s41467-019-09027-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Maruf FA, Pratama R, Song G. DNN-Boost: somatic mutation identification of tumor-only whole-exome sequencing data using deep neural network and XGBoost. J Bioinform Comput Biol. 2021 [cited 2025 Dec 5];19(06). Available from: https://doi.org/10.1142/S0219720021400175; https://www.worldscientific.com/worldscinet/jbcb. [DOI] [PubMed]
- 78.Krishnamachari K, Lu D, Swift-Scott A, Yeraliyev A, Lee K, Huang W, et al. Accurate somatic variant detection using weakly supervised deep learning. Nat Commun. 2022;13(1):4248. 10.1038/s41467-022-31765-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Scheffler K, Catreux S, O’Connell T, Jo H, Jain V, Heyns T, et al. Somatic small-variant calling methods in Illumina DRAGENTM secondary analysis [Internet]. bioRxiv. bioRxiv. 2023. Available from: http://biorxiv.org/lookup/doi/10.1101/2023.03.23.534011.
- 80.Vilov S, Heinig M. DeepSom: a CNN-based approach to somatic variant calling in WGS samples without a matched normal. Bioinformatics [Internet]. 2023;39(1). Available from: 10.1093/bioinformatics/btac828. [DOI] [PMC free article] [PubMed]
- 81.Park J, Cook DE, Chang P-C, Kolesnikov A, Brambrink L, Mier JC, et al. DeepSomatic: accurate somatic small variant discovery for multiple sequencing technologies [Internet]. bioRxiv. 2024. Available from: 10.1101/2024.08.16.608331. [DOI] [PubMed]
- 82.Li H, Meng L, Wang H, Cui L, Sheng H, Zhao P, et al. Precise identification of somatic and germline variants in the absence of matched normal samples. Briefings Bioinf [Internet]. 2024 [cited 2025 Dec 14];26(1). Available from: https://doi.org/10.1093/bib/bbae677; https://pubmed.ncbi.nlm.nih.gov/39737564/. [DOI] [PMC free article] [PubMed]
- 83.Krishnamachari K, Ngyuen B, An H, Skanderup AJ. Identifying somatic mutations from tumor-only sequencing using deep learning. [cited 3 Nov 2025]. Available from: https://openreview.net/pdf?id=UaKdIbTduH.
- 84.Chen L, Zheng Z, Su J, Yu X, Wong AOK, Zhang J, et al. ClairS-TO: a deep-learning method for long-read tumor-only somatic small variant calling. Nat Commun. 2025;16(1):9630. 10.1038/s41467-025-64547-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Fang LT, Zhu B, Zhao Y, Chen W, Yang Z, Kerrigan L, et al. Establishing community reference samples, data and call sets for benchmarking cancer mutation detection using whole-genome sequencing. Nat Biotechnol. 2021;39(9):1151–60. 10.1038/s41587-021-00993-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Ewing AD, Houlahan KE, Hu Y, Ellrott K, Caloian C, Yamaguchi TN, et al. Combining tumor genome simulation with crowdsourcing to benchmark somatic single-nucleotide-variant detection. Nat Methods. 2015;12(7):623–30. 10.1038/nmeth.3407. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Charoentong P, Finotello F, Angelova M, Mayer C, Efremova M, Rieder D, et al. Pan-cancer immunogenomic analyses reveal genotype-immunophenotype relationships and predictors of response to checkpoint blockade. Cell Rep [Internet]. 2017 [cited 2025 Dec 13];18(1):248–62. Available from: https://doi.org/10.1016/j.celrep.2016.12.019; https://pubmed.ncbi.nlm.nih.gov/28052254/. [DOI] [PubMed] [Google Scholar]
- 88.Chen Z, Yuan Y, Chen X, Chen J, Lin S, Li X, et al. Systematic comparison of somatic variant calling performance among different sequencing depth and mutation frequency. Sci Rep. 2020;10(1):3501. 10.1038/s41598-020-60559-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Guille A, Adélaïde J, Finetti P, Andre F, Birnbaum D, Mamessier E, et al. A benchmarking study of individual somatic variant callers and voting-based ensembles for whole-exome sequencing. Briefings Bioinf. 2024;26(1):bbae697. 10.1093/bib/bbae697. [DOI] [PMC free article] [PubMed]
- 90.Ott PA, Hu Z, Keskin DB, Shukla SA, Sun J, Bozym DJ, et al. An immunogenic personal neoantigen vaccine for patients with melanoma. Nature. 2017;547(7662):217–21. 10.1038/nature22991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Yadav M, Jhunjhunwala S, Phung QT, Lupardus P, Tanguay J, Bumbaca S, et al. Predicting immunogenic tumour mutations by combining mass spectrometry and exome sequencing. Nature [Internet]. 2014 [cited 2025 Dec 13];515(7528):572–76. Available from: https://doi.org/10.1038/nature14001; https://pubmed.ncbi.nlm.nih.gov/25428506/. [DOI] [PubMed] [Google Scholar]
- 92.McGranahan N, Rosenthal R, Hiley CT, Rowan AJ, Watkins TBK, Wilson GA, et al. Allele-specific HLA loss and immune escape in lung cancer evolution. Cell [Internet]. 2017 [cited 2025 Dec 13];171. Available from: https://pubmed.ncbi.nlm.nih.gov/29107330/. [DOI] [PMC free article] [PubMed]
- 93.Barker DJ, Maccari G, Georgiou X, Cooper MA, Flicek P, Robinson J, et al. The IPD-IMGT/HLA database. Nucleic Acids Res [Internet]. 2023 [cited 2025 Dec 13];51. Available from: https://pubmed.ncbi.nlm.nih.gov/36350643/. [DOI] [PMC free article] [PubMed]
- 94.Xie C, Yeo ZX, Wong M, Piper J, Long T, Kirkness EF, et al. Fast and accurate HLA typing from short-read next-generation sequence data with xHLA. Proc Natl Acad Sci USA. 2017 [cited 2025 Dec 13];114(30):8059–64. Available from: https://doi.org/10.1073/pnas.1707945114; https://pubmed.ncbi.nlm.nih.gov/28674023/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Dilthey AT, Mentzer AJ, Carapito R, Cutland C, Cereb N, Madhi SA, et al. HLA*LA—HLA typing from linearly projected graph alignments. Bioinformatics (Oxford, England) [Internet]. 2019 [cited 2025 Dec 13];35(21):4394–96. Available from: https://doi.org/10.1093/bioinformatics/btz235; https://pubmed.ncbi.nlm.nih.gov/30942877/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Dilthey AT, Gourraud P-A, Mentzer AJ, Cereb N, Iqbal Z, McVean G. High-accuracy HLA type inference from whole-genome sequencing data using population reference graphs. PLoS Comput Biol. 2016;12(10):e1005151. 10.1371/journal.pcbi.1005151. [DOI] [PMC free article] [PubMed]
- 97.Szolek A, Schubert B, Mohr C, Sturm M, Feldhahn M, Kohlbacher O. OptiType: precision HLA typing from next-generation sequencing data. Bioinformatics. 2014;30(23):3310–16. 10.1093/bioinformatics/btu548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019 [cited 2025 Dec 13];37(8):907–15. Available from: https://doi.org/10.1038/s41587-019-0201-4; https://pubmed.ncbi.nlm.nih.gov/31375807/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Orenbuch R, Filip I, Comito D, Shaman J, Pe’er I, Rabadan R. arcasHLA: high-resolution HLA typing from RNAseq. Bioinformatics. 2020;36(1):33–40. 10.1093/bioinformatics/btz474. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Lee H, Kingsford C. Kourami: graph-guided assembly for novel human leukocyte antigen allele discovery. Genome Biol. 2018;19(1):16. 10.1186/s13059-018-1388-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Butler-Laporte G, Farjoun J, Nakanishi T, Lu T, Abner E, Chen Y, et al. HLA allele-calling using multi-ancestry whole-exome sequencing from the UK Biobank identifies 129 novel associations in 11 autoimmune diseases. Commun Biol. 2023;6(1):1113. 10.1038/s42003-023-05496-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Warren RL, Choe G, Freeman DJ, Castellarin M, Munro S, Moore R, et al. Derivation of HLA types from shotgun sequence datasets. Genome Med. 2012;4(12):95. 10.1186/gm396. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Boegel S, Löwer M, Schäfer M, Bukur T, de Graaf J, Boisguérin V, et al. HLA typing from RNA-Seq sequence reads. Genome Med [Internet]. 2012 [cited 2025 Dec 15];4(12). Available from: https://doi.org/10.1186/gm403; https://pubmed.ncbi.nlm.nih.gov/23259685/. [DOI] [PMC free article] [PubMed]
- 104.Cao H, Wu J, Wang Y, Jiang H, Zhang T, Liu X, et al. An integrated tool to study MHC region: accurate SNV detection and HLA genes typing in human MHC region using targeted high-throughput sequencing. PLoS One. 2013;8(7):e69388. 10.1371/journal.pone.0069388. [DOI] [PMC free article] [PubMed]
- 105.Liu C, Yang X. Using exome and amplicon-based sequencing data for high-resolution HLA typing with ATHLATES. Methods Mol Biol. 2018;1802:203–13. [DOI] [PubMed] [Google Scholar]
- 106.Kim HJ, Pourmand N. HLA haplotyping from RNA-seq data using hierarchical read weighting. PLoS One. 2013;8(6):e67885. 10.1371/journal.pone.0067885. [DOI] [PMC free article] [PubMed]
- 107.Shukla SA, Rooney MS, Rajasagi M, Tiao G, Dixon PM, Lawrence MS, et al. Comprehensive analysis of cancer-associated somatic mutations in class I HLA genes. Nat Biotechnol. 2015 [cited 2025 Dec 15];33(11):1152–58. Available from: https://doi.org/10.1038/nbt.3344; https://pubmed.ncbi.nlm.nih.gov/26372948/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Huang Y, Yang J, Ying D, Zhang Y, Shotelersuk V, Hirankarn N, et al. Hlareporter: a tool for HLA typing from next generation sequencing data. Genome Med. 2015;7(1):25. 10.1186/s13073-015-0145-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Nariai N, Kojima K, Saito S, Mimori T, Sato Y, Kawai Y, et al. HLA-VBSeq: accurate HLA typing at full resolution from whole-genome sequencing data. BMC Genomics. 2015;16(S2):S7. 10.1186/1471-2164-16-S2-S7. [DOI] [PMC free article] [PubMed]
- 110.Kawaguchi S, Higasa K, Shimizu M, Yamada R, Matsuda F. HLA-HD: an accurate HLA typing algorithm for next-generation sequencing data. Hum Mutat. 2017;38(7):788–97. 10.1002/humu.23230. [DOI] [PubMed] [Google Scholar]
- 111.Ka S, Lee S, Hong J, Cho Y, Sung J, Kim H-N, et al. Hlascan: genotyping of the HLA region using next-generation sequencing data. BMC Bioinf. 2017;18(1):258. 10.1186/s12859-017-1671-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Buchkovich ML, Brown CC, Robasky K, Chai S, Westfall S, Vincent BG, et al. Hlaprofiler utilizes k-mer profiles to improve HLA calling accuracy for rare and common alleles in RNA-seq data. Genome Med. 2017;9(1):86. 10.1186/s13073-017-0473-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Bai Y, Wang D, Fury W. PHLAT: inference of high-resolution HLA types from RNA and whole exome sequencing. Methods Mol Biol. 2018;1802:193–201. [DOI] [PubMed] [Google Scholar]
- 114.Mentzer AJ, Dilthey AT, Pollard M, Gurdasani D, Karakoc E, Carstensen T, et al. High-resolution African HLA resource uncovers HLA-DRB1 expression effects underlying vaccine response. Nat Med. 2024;30(5):1384–94. 10.1038/s41591-024-02944-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Claeys A, Merseburger P, Staut J, Marchal K, Van den Eynden J. Benchmark of tools for in silico prediction of MHC class I and class II genotypes from NGS data. BMC Genomics. 2023;24(1):247. 10.1186/s12864-023-09351-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Vita R, Mahajan S, Overton JA, Dhanda SK, Martini S, Cantrell JR, et al. The immune epitope database (IEDB): 2018 update. Nucleic Acids Res [Internet]. 2019 [cited 2025 Dec 14];47(D1):D339–43. Available from: https://doi.org/10.1093/nar/gky1006; https://pubmed.ncbi.nlm.nih.gov/30357391/. [DOI] [PMC free article] [PubMed]
- 117.Rammensee H, Bachmann J, Emmerich NP, Bachor OA, Stevanović S. SYFPEITHI: database for MHC ligands and peptide motifs. Immunogenetics. 1999;50(3–4):213–19. 10.1007/s002510050595. [DOI] [PubMed] [Google Scholar]
- 118.Shao Y, Gao Y, Wu L-Y, Ge S-G, Wen P-B. TumorAgDB1.0: tumor neoantigen database platform. Database (Oxford) [Internet]. 2025;2025. Available from: 10.1093/database/baaf010. [DOI] [PMC free article] [PubMed]
- 119.Zhang G, Chitkushev L, Olsen LR, Keskin DB, Brusic V. TANTIGEN 2.0: a knowledge base of tumor T cell antigens and epitopes. BMC Bioinf. 2021;22(S8):40. 10.1186/s12859-021-03962-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Lata S, Bhasin M, Raghava GPS. MHCBN 4.0: a database of MHC/TAP binding peptides and T-cell epitopes. BMC Res Notes. 2009;2(1):61. 10.1186/1756-0500-2-61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Toseland CP, Clayton DJ, McSparron H, Hemsley SL, Blythe MJ, Paine K, et al. AntiJen: a quantitative immunology database integrating functional, thermodynamic, kinetic, biophysical, and cellular data. Immunome Res. 2005;1(1):4. 10.1186/1745-7580-1-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Brusic V, Rudy G, Harrison LC. MHCPEP: a database of MHC-binding peptides. Nucleic Acids Res. 1994;22(17):3663–65. 10.1093/nar/22.17.3663. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Cancer Genome Atlas Research Network, Weinstein JN, Collisson EA, Mills GB, Shaw KRM, Ozenberger BA, et al. The cancer genome Atlas pan-cancer analysis project. Nat Genet. 2013;45(10):1113–20. 10.1038/ng.2764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Robinson J, Barker DJ, Georgiou X, Cooper MA, Flicek P, Marsh SGE. IPD-IMGT/HLA database. Nucleic Acids Res. 2020;48:D948–55. [DOI] [PMC free article] [PubMed]
- 125.Robinson J, Waller MJ, Parham P, de Groot N, Bontrop R, Kennedy LJ, et al. IMGT/HLA and IMGT/MHC: sequence databases for the study of the major histocompatibility complex. Nucleic Acids Res. 2003;31(1):311–14. 10.1093/nar/gkg070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126.Tickotsky N, Sagiv T, Prilusky J, Shifrut E, Friedman N. McPAS-TCR: a manually curated catalogue of pathology-associated T cell receptor sequences. Bioinformatics. 2017;33(18):2924–29. 10.1093/bioinformatics/btx286. [DOI] [PubMed] [Google Scholar]
- 127.Bagaev DV, Vroomans RMA, Samir J, Stervbo U, Rius C, Dolton G, et al. Vdjdb in 2019: database extension, new analysis infrastructure and a T-cell receptor motif compendium. Nucleic Acids Res [Internet]. 2020 [cited 2025 Dec 14];48(D1):D1057–62. Available from: https://doi.org/10.1093/nar/gkz874; https://pubmed.ncbi.nlm.nih.gov/31588507/. [DOI] [PMC free article] [PubMed]
- 128.Website [Internet]. Available from: https://www.10xgenomics.com/.
- 129.Chen S-Y, Yue T, Lei Q, Guo A-Y. Tcrdb: a comprehensive database for T-cell receptor sequences with powerful search function. Nucleic Acids Res. 2021;49(D1):D468–74. 10.1093/nar/gkaa796. [DOI] [PMC free article] [PubMed]
- 130.Zhang W, Wang L, Liu K, Wei X, Yang K, Du W, et al. PIRD: Pan immune repertoire database. Bioinformatics. 2020;36(3):897–903. 10.1093/bioinformatics/btz614. [DOI] [PubMed] [Google Scholar]
- 131.Nolan S, Vignali M, Klinger M, Dines JN, Kaplan IM, Svejnoha E, et al. A large-scale database of T-cell receptor beta sequences and binding associations from natural and synthetic exposure to SARS-CoV-2. Front Immunol. 2025;16:1488851. 10.3389/fimmu.2025.1488851. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 132.Sijts EJAM, Kloetzel PM. The role of the proteasome in the generation of MHC class I ligands and immune responses. Cell Mol Life Sci. 2011;68(9):1491–502. 10.1007/s00018-011-0657-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133.Calis JJ, Reinink P, Keller C, Kloetzel PM, Keşmir C. Role of peptide processing predictions in T cell epitope identification: contribution of different prediction programs. Immunogenet [Internet]. 2015 [cited 2025 Dec 13];67(2):85–93. Available from: https://doi.org/10.1007/s00251-014-0815-0; https://pubmed.ncbi.nlm.nih.gov/25475908/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Nielsen M, Lundegaard C, Lund O, Keşmir C. The role of the proteasome in generating cytotoxic T-cell epitopes: insights obtained from improved predictions of proteasomal cleavage. Immunogenetics. 2005;57(1–2):33–41. 10.1007/s00251-005-0781-7. [DOI] [PubMed] [Google Scholar]
- 135.Weeder BR, Wood MA, Li E, Nellore A, Thompson RF. Pepsickle rapidly and accurately predicts proteasomal cleavage sites for improved neoantigen identification. Bioinformatics (Oxford, England) [Internet]. 2021 [cited 2025 Dec 13];37. Available from: https://pubmed.ncbi.nlm.nih.gov/34478497/. [DOI] [PubMed]
- 136.Gomez-Perosanz M, Ras-Carmona A, Lafuente EM, Reche PA. Identification of CD8+ T cell epitopes through proteasome cleavage site predictions. BMC Bioinf. 2020;21(S17):484. 10.1186/s12859-020-03782-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 137.Bhasin M, Raghava GPS. Pcleavage: an SVM based method for prediction of constitutive proteasome and immunoproteasome cleavage sites in antigenic sequences. Nucleic Acids Res. 2005;33:W202–7. 10.1093/nar/gki587. [DOI] [PMC free article] [PubMed]
- 138.Nussbaum AK, Kuttler C, Hadeler KP, Rammensee HG, Schild H. PAProC: a prediction algorithm for proteasomal cleavages available on the WWW. Immunogenetics. 2001;53(2):87–94. 10.1007/s002510100300. [DOI] [PubMed] [Google Scholar]
- 139.Hakenberg J, Nussbaum AK, Schild H, Rammensee H-G, Kuttler C, Holzhütter H-G, et al. MAPPP: MHC class I antigenic peptide processing prediction. Appl Bioinf. 2003;2(3):155–58. [PubMed] [Google Scholar]
- 140.Li B-Q, Cai Y-D, Feng K-Y, Zhao G-J. Prediction of protein cleavage site with feature selection by random forest. PLoS One. 2012;7(9):e45854. 10.1371/journal.pone.0045854. [DOI] [PMC free article] [PubMed]
- 141.Hoze E, Tsaban L, Maman Y, Louzoun Y. Predictor for the effect of amino acid composition on CD4+ T cell epitopes preprocessing. J Immunol Methods. 2013;391(1–2):163–73. 10.1016/j.jim.2013.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 142.Paul S, Karosiene E, Dhanda SK, Jurtz V, Edwards L, Nielsen M, et al. Determination of a predictive cleavage motif for eluted major histocompatibility complex class II ligands. Front Immunol. 2018;9:1795. 10.3389/fimmu.2018.01795. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Tenzer S, Peters B, Bulik S, Schoor O, Lemmel C, Schatz MM, et al. Modeling the MHC class I pathway by combining predictions of proteasomal cleavage, TAP transport and MHC class I binding. Cellular Mol Life Sci: CMLS [Internet]. 2005 [cited 2025 Dec 15];62(9):1025–37. Available from: https://doi.org/10.1007/s00018-005-4528-2; https://pubmed.ncbi.nlm.nih.gov/15868101/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 144.Bhasin M, Raghava GPS. Analysis and prediction of affinity of TAP binding peptides using cascade SVM. Protein Sci. 2004;13(3):596–607. 10.1110/ps.03373104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 145.Zhang GL, Petrovsky N, Kwoh CK, August JT, Brusic V. PRED(TAP): a system for prediction of peptide binding to the human transporter associated with antigen processing. Immunome Res. 2006;2(1):3. 10.1186/1745-7580-2-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 146.Diez-Rivero CM, Chenlo B, Zuluaga P, Reche PA. Quantitative modeling of peptide binding to TAP using support vector machine. Proteins [Internet]. 2010 [cited 2025 Dec 13];78(1):63–72. Available from: https://doi.org/10.1002/prot.22535; https://pubmed.ncbi.nlm.nih.gov/19705485/. [DOI] [PubMed] [Google Scholar]
- 147.Peters B, Bulik S, Tampe R, Van Endert PM, Holzhütter H-G. Identifying MHC class I epitopes by predicting the TAP transport efficiency of epitope precursors. J Immunol. 2003;171(4):1741–49. 10.4049/jimmunol.171.4.1741. [DOI] [PubMed] [Google Scholar]
- 148.Dönnes P, Kohlbacher O. Integrated modeling of the major events in the MHC class I antigen processing pathway. Protein Sci. 2005;14(8):2132–40. 10.1110/ps.051352405. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 149.Zhang X, Wu J, Baeza J, Gu K, Zheng Y, Chen S, et al. DeepTAP: an RNN-based method of TAP-binding peptide prediction in the selection of tumor neoantigens. Comput Biol Med. 2023;164:107247. 10.1016/j.compbiomed.2023.107247. [DOI] [PubMed] [Google Scholar]
- 150.Lam TH, Mamitsuka H, Ren EC, Tong JC. TAP Hunter: a SVM-based system for predicting TAP ligands using local description of amino acid sequence. Immunome Res [Internet]. 2010 [cited 2025 Dec 15];6(Suppl 1). Available from: https://doi.org/10.1186/1745-7580-6-S1-S6; https://pubmed.ncbi.nlm.nih.gov/20875157/. [DOI] [PMC free article] [PubMed]
- 151.Zhu L, Chen W, Yang S. CLTAP: a TAP-binding peptide prediction method using pre-trained large language-embedding models with contrastive learning to enhance contextual co-attention mechanism. Expert Syst Appl. 2025;285:127991. 10.1016/j.eswa.2025.127991. [Google Scholar]
- 152.Andreatta M, Nielsen M. Gapped sequence alignment using artificial neural networks: application to the MHC class I system. Bioinformatics (Oxford, England) [Internet]. 2016 [cited 2025 Dec 13];32(4):511–17. Available from: https://doi.org/10.1093/bioinformatics/btv639; https://pubmed.ncbi.nlm.nih.gov/26515819/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 153.Jurtz V, Paul S, Andreatta M, Marcatili P, Peters B, Nielsen M. NetMHCpan-4.0: improved peptide–MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data. J Immunol. 2017 [cited 2025 Dec 13];199(9):3360–68. Available from: https://doi.org/10.4049/jimmunol.1700893; https://pubmed.ncbi.nlm.nih.gov/28978689/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 154.Han Y, Kim D. Deep convolutional neural networks for pan-specific peptide-MHC class I binding prediction. BMC Bioinf. 2017;18(1):585. 10.1186/s12859-017-1997-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 155.O’Donnell TJ, Rubinsteyn A, Bonsack M, Riemer AB, Laserson U, Hammerbacher J. Mhcflurry: open-source class I MHC binding affinity prediction. Cell Syst. 2018;7(1):129–32.e4. 10.1016/j.cels.2018.05.014. [DOI] [PubMed] [Google Scholar]
- 156.Wongklaew P, Sriswasdi S, Chuangsuwanich E. MHCSeqNet2—improved peptide-class I MHC binding prediction for alleles with low data. Bioinformatics [Internet]. 2024;40(1). Available from: 10.1093/bioinformatics/btad780. [DOI] [PMC free article] [PubMed]
- 157.Vielhaben J, Wenzel M, Samek W, Strodthoff N. Usmpep: universal sequence models for major histocompatibility complex binding affinity prediction. BMC Bioinf. 2020;21(1):279. 10.1186/s12859-020-03631-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 158.Wu J, Wang W, Zhang J, Zhou B, Zhao W, Su Z, et al. DeepHLApan: a deep learning approach for neoantigen prediction considering both HLA-Peptide binding and immunogenicity. Front Immunol. 2019;10:2559. 10.3389/fimmu.2019.02559. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 159.Gasser H-C, Bedran G, Ren B, Goodlett D, Alfaro J, Rajan A. Interpreting BERT architecture predictions for peptide presentation by MHC class I proteins [Internet]. arXiv [q-bio.QM]. 2021. Available from: 10.48550/ARXIV.2111.07137.
- 160.Chen B, Khodadoust MS, Olsson N, Wagar LE, Fast E, Liu CL, et al. Predicting HLA class II antigen presentation through integrated deep learning. Nat Biotechnol. 2019;37(11):1332–43. 10.1038/s41587-019-0280-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 161.Albert BA, Yang Y, Shao XM, Singh D, Smit KN, Anagnostou V, et al. Deep neural networks predict class I major histocompatibility complex epitope presentation and transfer learn neoepitope immunogenicity. Nat Mach Intell. 2023 [cited 2025 Dec 13];5(8):861–72. Available from: https://doi.org/10.1038/s42256-023-00694-6; https://pubmed.ncbi.nlm.nih.gov/37829001/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 162.Kim JY, Bang H, Noh S-J, Choi JK. DeepNeo: a webserver for predicting immunogenic neoantigens. Nucleic Acids Res. 2023;51(W1):W134–40. 10.1093/nar/gkad275. [DOI] [PMC free article] [PubMed]
- 163.Gfeller D, Schmidt J, Croce G, Guillaume P, Bobisse S, Genolet R, et al. Improved predictions of antigen presentation and TCR recognition with MixMHCpred2.2 and PRIME2.0 reveal potent SARS-CoV-2 CD8+ T-cell epitopes. Cell Syst [Internet]. 2023 [cited 2025 Dec 13];14(1):72–83.e5. Available from: https://doi.org/10.1016/j.cels.2022.12.002; https://pubmed.ncbi.nlm.nih.gov/36603583/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 164.Reynisson B, Alvarez B, Paul S, Peters B, Nielsen M. NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Res. 2020;48(W1):W449–54. 10.1093/nar/gkaa379. [DOI] [PMC free article] [PubMed]
- 165.Calis JJ, Maybeno M, Greenbaum JA, Weiskopf D, De Silva AD, Sette A, et al. Properties of MHC class I presented peptides that enhance immunogenicity. PLoS Comput Biol. 2013 [cited 2025 Dec 13];9(10):e1003266. Available from: https://doi.org/10.1371/journal.pcbi.1003266; https://pubmed.ncbi.nlm.nih.gov/24204222/. [DOI] [PMC free article] [PubMed]
- 166.Parker KC, Bednarek MA, Coligan JE. Scheme for ranking potential HLA-A2 binding peptides based on independent binding of individual peptide side-chains. J Immunol. 1994 [cited 2025 Dec 15];152(1):163–75. Available from: https://doi.org/10.4049/jimmunol.152.1.163; https://pubmed.ncbi.nlm.nih.gov/8254189/. [PubMed] [Google Scholar]
- 167.Singh H, Raghava GP. ProPred: prediction of HLA-DR binding sites. Bioinformatics. 2001;17(12):1236–37. 10.1093/bioinformatics/17.12.1236. [DOI] [PubMed] [Google Scholar]
- 168.Bhasin M, Raghava GPS. Prediction of CTL epitopes using QM, SVM and ANN techniques. Vaccine. 2004;22(23–24):3195–204. 10.1016/j.vaccine.2004.02.005. [DOI] [PubMed] [Google Scholar]
- 169.Reche PA, Glutting JP, Zhang H, Reinherz EL. Enhancement to the RANKPEP resource for the prediction of peptide binding to MHC molecules using profiles. Immunogenet [Internet]. 2004 [cited 2025 Dec 15];56(6). Available from: https://doi.org/10.1007/s00251-004-0709-7; https://pubmed.ncbi.nlm.nih.gov/15349703/. [DOI] [PubMed]
- 170.Bui HH, Sidney J, Peters B, Sathiamurthy M, Sinichi A, Purton KA, et al. Automated generation and evaluation of specific MHC binding predictive tools: ARB matrix applications. Immunogenet [Internet]. 2005 [cited 2025 Dec 15];57(5):304–14. Available from: https://doi.org/10.1007/s00251-005-0798-y; https://pubmed.ncbi.nlm.nih.gov/15868141/. [DOI] [PubMed] [Google Scholar]
- 171.Peters B, Sette A. Generating quantitative models describing the sequence specificity of biological processes with the stabilized matrix method. BMC Bioinf [Internet]. 2005 [cited 2025 Dec 15];6(1). Available from: https://doi.org/10.1186/1471-2105-6-132; https://pubmed.ncbi.nlm.nih.gov/15927070/. [DOI] [PMC free article] [PubMed]
- 172.Guan P, Hattotuwagama CK, Doytchinova IA, Flower DR. Mhcpred 2.0: an updated quantitative T-cell epitope prediction server. Appl Bioinf. 2006;5(1):55–61. 10.2165/00822942-200605010-00008. [DOI] [PubMed] [Google Scholar]
- 173.Nielsen M, Lundegaard C, Lund O. Prediction of MHC class II binding affinity using SMM-align, a novel stabilization matrix alignment method. BMC Bioinf. 2007;8(1):238. 10.1186/1471-2105-8-238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 174.Kim Y, Sidney J, Pinilla C, Sette A, Peters B. Derivation of an amino acid similarity matrix for peptide: MHC binding and its application as a Bayesian prior. BMC Bioinf [Internet]. 2009 [cited 2025 Dec 15];10(1). Available from: https://doi.org/10.1186/1471-2105-10-394; https://pubmed.ncbi.nlm.nih.gov/19948066/. [DOI] [PMC free article] [PubMed]
- 175.Nielsen M, Lund O. NN-align. An artificial neural network-based alignment algorithm for MHC class II peptide binding prediction. BMC Bioinf. 2009;10(1):296. 10.1186/1471-2105-10-296. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 176.Zhang H, Lund O, Nielsen M. The PickPocket method for predicting binding specificities for receptors based on receptor pocket similarities: application to MHC-peptide binding. Bioinformatics. 2009;25(10):1293–99. 10.1093/bioinformatics/btp137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 177.Bordner AJ, Mittelmann HD. MultiRTA: a simple yet reliable method for predicting peptide binding affinities for multiple class II MHC allotypes. BMC Bioinf. 2010;11(1):482. 10.1186/1471-2105-11-482. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 178.Stranzl T, Larsen MV, Lundegaard C, Nielsen M. NetCTLpan: pan-specific MHC class I pathway epitope predictions. Immunogenetics. 2010;62(6):357–68. 10.1007/s00251-010-0441-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 179.Zhang GL, DeLuca DS, Keskin DB, Chitkushev L, Zlateva T, Lund O, et al. MULTIPRED2: a computational system for large-scale identification of peptides predicted to bind to HLA supertypes and alleles. J Immunol Methods. 2011;374(1–2):53–61. 10.1016/j.jim.2010.11.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 180.Liu I-H, Lo Y-S, Yang J-M. Pacomplex: a web server to infer peptide antigen families and binding models from TCR–pMHC complexes. Nucleic Acids Res. 2011;39(suppl_2):W254–60. 10.1093/nar/gkr434. [DOI] [PMC free article] [PubMed]
- 181.Zhang L, Chen Y, Wong H-S, Zhou S, Mamitsuka H, Zhu S. Tepitopepan: extending TEPITOPE for peptide binding prediction covering over 700 HLA-DR molecules. PLoS One. 2012;7(2):e30483. 10.1371/journal.pone.0030483. [DOI] [PMC free article] [PubMed]
- 182.Kim Y, Ponomarenko J, Zhu Z, Tamang D, Wang P, Greenbaum J, et al. Immune epitope database analysis resource. Nucleic Acids Res [Internet]. 2012 [cited 2025 Nov 24];40(W1):W525–30. Available from: https://doi.org/10.1093/nar/gks438; https://pubmed.ncbi.nlm.nih.gov/22610854/. [DOI] [PMC free article] [PubMed]
- 183.Karosiene E, Lundegaard C, Lund O, Nielsen M. NetMHCcons: a consensus method for the major histocompatibility complex class I predictions. Immunogenet [Internet]. 2012 [cited 2025 Dec 15];64(3):177–86. Available from: https://doi.org/10.1007/s00251-011-0579-8; https://pubmed.ncbi.nlm.nih.gov/22009319/. [DOI] [PubMed] [Google Scholar]
- 184.Shen WJ, Zhang S, Wong HS. An effective and effecient peptide binding prediction approach for a broad set of HLA-DR molecules based on ordered weighted averaging of binding pocket profiles. Proteome Sci [Internet]. 2013 [cited 2025 Dec 15];11(S1). Available from: https://doi.org/10.1186/1477-5956-11-S1-S15; https://pubmed.ncbi.nlm.nih.gov/24565049/. [DOI] [PMC free article] [PubMed]
- 185.Rasmussen M, Fenoy E, Harndahl M, Kristensen AB, Nielsen IK, Nielsen M, et al. Pan-specific prediction of peptide-MHC class I complex stability, a correlate of T cell immunogenicity. J Immunol. 2016;197(4):1517–24. 10.4049/jimmunol.1600582. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 186.Liu G, Li D, Li Z, Qiu S, Li W, Chao CC, et al. Pssmhcpan: a novel PSSM-based software for predicting class I peptide-HLA binding affinity. Giga Sci. 2017 [cited 2025 Dec 15];6(5). Available from: https://doi.org/10.1093/gigascience/gix017; https://pubmed.ncbi.nlm.nih.gov/28327987/. [DOI] [PMC free article] [PubMed]
- 187.Nielsen M, Andreatta M. Nnalign: a platform to construct and evaluate artificial neural network models of receptor–ligand interactions. Nucleic Acids Res. 2017;45(W1):W344–9. 10.1093/nar/gkx276. [DOI] [PMC free article] [PubMed]
- 188.Abelin JG, Keskin DB, Sarkizova S, Hartigan CR, Zhang W, Sidney J, et al. Mass spectrometry profiling of HLA-associated peptidomes in mono-allelic cells enables more accurate epitope prediction. Immunity. 2017;46(2):315–26. 10.1016/j.immuni.2017.02.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 189.Vang YS, Xie X. HLA class I binding prediction via convolutional neural networks. Bioinformatics (Oxford, England) [Internet]. 2017 [cited 2025 Dec 15];33. Available from: https://pubmed.ncbi.nlm.nih.gov/28444127/. [DOI] [PubMed]
- 190.Hu J, Liu Z. DeepMHC: deep convolutional neural networks for high-performance peptide-MHC binding affinity prediction [Internet]. bioRxiv. bioRxiv; 2017. Available from: http://biorxiv.org/lookup/doi/10.1101/239236.
- 191.Sidhom J-W, Pardoll D, Baras A. AI-MHC: an allele-integrated deep learning framework for improving class I & class II HLA-binding predictions [Internet]. bioRxiv. bioRxiv. 2018. Available from: http://biorxiv.org/lookup/doi/10.1101/318881.
- 192.Bulik-Sullivan B, Busby J, Palmer CD, Davis MJ, Murphy T, Clark A, et al. Deep learning using tumor HLA peptide mass spectrometry datasets improves neoantigen identification. Nat Biotechnol [Internet]. 2018 [cited 2025 Dec 14]. Available from: https://pubmed.ncbi.nlm.nih.gov/30556813/. [DOI] [PubMed]
- 193.Zhao T, Cheng L, Zang T, Hu Y. Peptide-major histocompatibility complex class I binding prediction based on deep learning with novel feature. Front Genet. 2019 [cited 2025 Dec 15];10. Available from: https://doi.org/10.3389/fgene.2019.01191; https://pubmed.ncbi.nlm.nih.gov/31850062/. [DOI] [PMC free article] [PubMed]
- 194.Zeng H, Gifford DK. DeepLigand: accurate prediction of MHC class I ligands using peptide embedding. Bioinformatics. 2019;35(14):i278–83. 10.1093/bioinformatics/btz330. [DOI] [PMC free article] [PubMed]
- 195.Zeng H, Gifford DK. Quantification of Uncertainty in peptide-MHC binding prediction improves high-affinity peptide selection for therapeutic design. Cell Syst [Internet]. 2019 [cited 2025 Dec 15];9(2):159–66.e3. Available from: https://doi.org/10.1016/j.cels.2019.05.004; https://pubmed.ncbi.nlm.nih.gov/31176619/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 196.Abelin JG, Harjanto D, Malloy M, Suri P, Colson T, Goulding SP, et al. Defining HLA-II ligand processing and binding rules with mass spectrometry enhances cancer epitope prediction. Immunity. 2021;54(2):388. 10.1016/j.immuni.2020.12.005. [DOI] [PubMed] [Google Scholar]
- 197.Liu Z, Cui Y, Xiong Z, Nasiri A, Zhang A, Hu J. DeepSeqPan, a novel deep convolutional neural network model for pan-specific class I HLA-peptide binding affinity prediction. Sci Rep. 2019;9(1):794. 10.1038/s41598-018-37214-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 198.Hu Y, Wang Z, Hu H, Wan F, Chen L, Xiong Y, et al. ACME: pan-specific peptide-MHC class I binding prediction through attention-based deep neural networks. Bioinformatics (Oxford, England) [Internet]. 2019 [cited 2025 Dec 15];35(23):4946–54. Available from: https://doi.org/10.1093/bioinformatics/btz427; https://pubmed.ncbi.nlm.nih.gov/31120490/. [DOI] [PubMed] [Google Scholar]
- 199.Boehm KM, Bhinder B, Raja VJ, Dephoure N, Elemento O. Predicting peptide presentation by major histocompatibility complex class I: an improved machine learning approach to the immunopeptidome. BMC Bioinf [Internet]. 2019 [cited 2025 Dec 15];20(1). Available from: https://doi.org/10.1186/s12859-018-2561-z; https://pubmed.ncbi.nlm.nih.gov/30611210/. [DOI] [PMC free article] [PubMed]
- 200.Xie X, Han Y, Zhang K. MHCherryPan: a novel pan-specific model for binding affinity prediction of class I HLA-peptide. Int J Data Min Bioinform. 2020;24(3):201. 10.1504/IJDMB.2020.112850. [Google Scholar]
- 201.Shao XM, Bhattacharya R, Huang J, Sivakumar IKA, Tokheim C, Zheng L, et al. High-throughput prediction of MHC class I and II neoantigens with MHCnuggets. Cancer Immunol Res [Internet]. 2020 [cited 2025 Dec 15];8(3):396–408. Available from: https://doi.org/10.1158/2326-6066.CIR-19-0464; https://pubmed.ncbi.nlm.nih.gov/31871119/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 202.Pei B, Hsu Y-H. IConMHC: a deep learning convolutional neural network model to predict peptide and MHC-I binding affinity. Immunogenetics. 2020;72(5):295–304. 10.1007/s00251-020-01163-9. [DOI] [PubMed] [Google Scholar]
- 203.Venkatesh G, Grover A, Srinivasaraghavan G, Rao S. MHCAttnNet: predicting MHC-peptide bindings for MHC alleles classes I and II using an attention-based deep neural model. Bioinformatics. 2020;36(Supplement_1):i399–406. 10.1093/bioinformatics/btaa479. [DOI] [PMC free article] [PubMed]
- 204.O’Donnell TJ, Rubinsteyn A, Laserson U. Mhcflurry 2.0: improved pan-allele prediction of MHC class I-Presented peptides by incorporating antigen processing. Cell Syst. 2020;11(4):418–19. 10.1016/j.cels.2020.09.001. [DOI] [PubMed] [Google Scholar]
- 205.Cheng J, Bendjama K, Rittner K, Malone B. BERTMHC: improved MHC–peptide class II interaction prediction with transformer and multiple instance learning. Bioinformatics. 2021;37(22):4172–79. 10.1093/bioinformatics/btab422. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 206.Jin J, Liu Z, Nasiri A, Cui Y, Louis SY, Zhang A, et al. Deep learning pan-specific model for interpretable MHC-I peptide binding prediction with improved attention mechanism. Proteins [Internet]. 2021 [cited 2025 Dec 15];89. Available from: https://pubmed.ncbi.nlm.nih.gov/33594723/. [DOI] [PubMed]
- 207.Yang X, Zhao L, Wei F, Li J. DeepNetBim: deep learning model for predicting HLA-epitope interactions based on network analysis by harnessing binding and immunogenicity information. BMC Bioinf. 2021;22(1):231. 10.1186/s12859-021-04155-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 208.Pyke RM, Mellacheruvu D, Dea S, Abbott C, Zhang SV, Phillips NA, et al. Precision neoantigen discovery using large-scale immunopeptidomes and composite modeling of MHC peptide presentation. Mol Cellular Proteomics. 2023 [cited 2025 Dec 14];22(4):100506. Available from: https://doi.org/10.1016/j.mcpro.2023.100506; https://pubmed.ncbi.nlm.nih.gov/36796642/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 209.Ye Y, Wang J, Xu Y, Wang Y, Pan Y, Song Q, et al. MATHLA: a robust framework for HLA-peptide binding prediction integrating bidirectional LSTM and multiple head attention mechanism. BMC Bioinf. 2021;22(1):7. 10.1186/s12859-020-03946-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 210.Mei S, Li F, Xiang D, Ayala R, Faridi P, Webb GI, et al. Anthem: a user customised tool for fast and accurate prediction of binding between peptides and HLA class I molecules. Brief Bioinform [Internet]. 2021;22(5). Available from: 10.1093/bib/bbaa415. [DOI] [PubMed]
- 211.Gartner JJ, Parkhurst MR, Gros A, Tran E, Jafferji MS, Copeland A, et al. A machine learning model for ranking candidate HLA class I neoantigens based on known neoepitopes from multiple human tumor types. Nat Cancer. 2021;2(5):563–74. 10.1038/s43018-021-00197-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 212.Bravi B, Tubiana J, Cocco S, Monasson R, Mora T, Walczak AM. RBM-MHC: a semi-supervised machine-learning method for sample-specific prediction of antigen presentation by HLA-I alleles. Cell Syst. 2021;12(2):195–202.e9. 10.1016/j.cels.2020.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 213.Chu Y, Zhang Y, Wang Q, Zhang L, Wang X, Wang Y, et al. A transformer-based model to predict peptide-HLA class I binding and optimize mutated peptides for vaccine design. Nat Mach Intell. 2022;4(3):300–11. 10.1038/s42256-022-00459-7. [Google Scholar]
- 214.Liu Z, Jin J, Cui Y, Xiong Z, Nasiri A, Zhao Y, et al. DeepSeqPanII: an interpretable recurrent neural network model with attention mechanism for peptide-HLA class II binding prediction. IEEE/ACM Trans Comput Biol Bioinform. 2022;19(4):2188–96. 10.1109/TCBB.2021.3074927. [DOI] [PubMed] [Google Scholar]
- 215.Wang F, Wang H, Wang L, Lu H, Qiu S, Zang T, et al. MHCRoBERTa: pan-specific peptide–MHC class I binding prediction through transfer learning with label-agnostic protein sequences. Briefings Bioinf. 2022;23(3). Available from: 10.1093/bib/bbab595. [DOI] [PubMed]
- 216.Xu S, Wang X, Fei C. A highly effective system for predicting MHC-II epitopes with immunogenicity. Front Oncol. 2022;12:888556. 10.3389/fonc.2022.888556. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 217.Thrift WJ, Lounsbury NW, Broadwell Q, Heidersbach A, Freund E, Abdolazimi Y, et al. Hlapollo: a superior transformer model for pan-allelic peptide-MHC-I presentation prediction, with diverse negative coverage, deconvolution and protein language features [Internet]. bioRxiv. 2022. Available from: http://biorxiv.org/lookup/doi/10.1101/2022.12.08.519673.
- 218.Zhang Y, Zhu G, Li K, Li F, Huang L, Duan M, et al. HLAB: learning the BiLSTM features from the ProtBert-encoded proteins for the class I HLA-peptide binding prediction. Briefings Bioinf. 2022;23(5). Available from: 10.1093/bib/bbac173. [DOI] [PMC free article] [PubMed]
- 219.You R, Qu W, Mamitsuka H, Zhu S. DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction. Bioinformatics. 2022;38:i220–8. [DOI] [PMC free article] [PubMed]
- 220.Deng J, Zhou X, Zhang P, Cheng W, Liu M, Tian J. IEPAPI: a method for immune epitope prediction by incorporating antigen presentation and immunogenicity. Brief Bioinform [Internet]. 2023;24(4). Available from: 10.1093/bib/bbad171. [DOI] [PubMed]
- 221.Racle J, Guillaume P, Schmidt J, Michaux J, Larabi A, Lau K, et al. Machine learning predictions of MHC-II specificities reveal alternative binding mode of class II epitopes. Immunity [Internet]. 2023 [cited 2025 Dec 15];56. Available from: https://pubmed.ncbi.nlm.nih.gov/37023751/. [DOI] [PubMed]
- 222.Kalemati M, Darvishi S, Koohi S. CapsNet-MHC predicts peptide-MHC class I binding based on capsule neural networks. Commun Biol. 2023;6(1):492. 10.1038/s42003-023-04867-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 223.Qu W, You R, Mamitsuka H, Zhu S. DeepMHCI: an anchor position-aware deep interaction model for accurate MHC-I peptide binding affinity prediction. Bioinformatics [Internet]. 2023;39(9). Available from: 10.1093/bioinformatics/btad551. [DOI] [PMC free article] [PubMed]
- 224.Wang G, Wu T, Ning W, Diao K, Sun X, Wang J, et al. TLimmuno2: predicting MHC class II antigen immunogenicity through transfer learning. Briefings Bioinf. 2023;24(3). Available from: 10.1093/bib/bbad116. [DOI] [PubMed]
- 225.Nilsson JB, Kaabinejadian S, Yari H, Kester MGD, van Balen P, Hildebrand WH, et al. Accurate prediction of HLA class II antigen presentation across all loci using tailored data acquisition and refined machine learning. Sci Adv. 2023;9(47):eadj 6367. 10.1126/sciadv.adj6367. [DOI] [PMC free article] [PubMed]
- 226.Yu X, Negron C, Huang L, Veldman G. TransMHCII: a novel MHC-II binding prediction model built using a protein language model and an image classifier. Antib Ther. 2023;6(2):137–46. 10.1093/abt/tbad011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 227.Jani SP, Kumar SP, Mangukia N, Patel SK, Pandya HA, Rawal RM. MHC2AffyPred: a machine-learning approach to estimate affinity of MHC class II peptides based on structural interaction fingerprints. Proteins. 2023;91:277–89. 10.1002/prot.26428. [DOI] [PubMed] [Google Scholar]
- 228.Ye Z, Li S, Mi X, Shao B, Dai Z, Ding B, et al. Stmhcpan, an accurate star-Transformer-based extensible framework for predicting MHC I allele binding peptides. Brief Bioinform [Internet]. 2023;24(3). Available from: 10.1093/bib/bbad164. [DOI] [PubMed]
- 229.Thrift WJ, Perera J, Cohen S, Lounsbury NW, Gurung HR, Rose CM, et al. Graph-pMHC: graph neural network approach to MHC class II peptide presentation and antibody immunogenicity. Briefings Bioinf. 2024;25(3). Available from: 10.1093/bib/bbae123. [DOI] [PMC free article] [PubMed]
- 230.Wang X, Wu T, Jiang Y, Chen T, Pan D, Jin Z, et al. RPEMHC: improved prediction of MHC-peptide binding affinity by a deep learning approach based on residue-residue pair encoding. Bioinformatics [Internet]. 2024;40(1). Available from: 10.1093/bioinformatics/btad785. [DOI] [PMC free article] [PubMed]
- 231.Wang M, Lei C, Wang J, Li Y, Li M. TripHLApan: predicting HLA molecules binding peptides based on triple coding matrix and transfer learning. Brief Bioinform [Internet]. 2024;25(3). Available from: 10.1093/bib/bbae154. [DOI] [PMC free article] [PubMed]
- 232.Saadat M, Zare-Mirakabad F, Masoudi-Nejad A, Baradaran MF, Hosseinkhan N. HLAPepBinder: an ensemble model for the prediction of HLA-Peptide binding. Iran J Biotechnol. 2024;22(4):e3927. 10.30498/ijb.2024.459448.3927. [DOI] [PMC free article] [PubMed]
- 233.Wohlwend J, Nathan A, Shalon N, Crain CR, Tano-Menka R, Goldberg B, et al. Deep learning enhances the prediction of HLA class I-presented CD8+ T cell epitopes in foreign pathogens. Nat Mach Intell. 2025;7(2):232–43. 10.1038/s42256-024-00971-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 234.Jian F, Cai H, Chen Q, Pan X, Feng W, Yuan Y. OnmiMHC: a machine learning solution for UCEC tumor vaccine development through enhanced peptide-MHC binding prediction. Front Immunol. 2025 [cited 2025 Dec 15];16. Available from: https://doi.org/10.3389/fimmu.2025.1550252; https://pubmed.ncbi.nlm.nih.gov/40092998/. [DOI] [PMC free article] [PubMed]
- 235.Ma J, Wang Z, Tong C, Yang Q, Zhang L, Liu H. pMhchat, characterizing the interactions between major histocompatibility complex class II molecules and peptides with large language models and deep hypergraph learning. Briefings Bioinf. 2025;26(4):26. Available from: 10.1093/bib/bbaf321. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 236.Tadros DM, Racle J, Gfeller D. Predicting MHC-I ligands across alleles and species: how far can we go? Genome Med. 2025;17(1):25. 10.1186/s13073-025-01450-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 237.Nazarloo M, Saadat M, Zare-Mirakabad F. PHLA-SiNet: a novel peptide-HLA binding prediction model using heterogeneous siamese neural networks. Comput Biol Med. 2025;197:111017. 10.1016/j.compbiomed.2025.111017. [DOI] [PubMed] [Google Scholar]
- 238.Marcu A, Bichmann L, Kuchenbecker L, Kowalewski DJ, Freudenmann LK, Backert L, et al. HLA ligand Atlas: a benign reference of HLA-presented peptides to improve T-cell-based cancer immunotherapy. J Immunother Cancer. 2021 [cited 2025 Dec 13];9(4):e002071. Available from: https://doi.org/10.1136/jitc-2020-002071; https://pubmed.ncbi.nlm.nih.gov/33858848/. [DOI] [PMC free article] [PubMed]
- 239.Zhao W, Sher X. Systematically benchmarking peptide-MHC binding predictors: from synthetic to naturally processed epitopes. PLoS Comput Biol. 2018;14(11):e1006457. 10.1371/journal.pcbi.1006457. [DOI] [PMC free article] [PubMed]
- 240.Yang Y, Wei Z, Cia G, Song X, Pucci F, Rooman M, et al. MHCII-peptide presentation: an assessment of the state-of-the-art prediction methods. Front Immunol. 2024;15:1293706. 10.3389/fimmu.2024.1293706. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 241.Croce G, Bobisse S, Moreno DL, Schmidt J, Guillame P, Harari A, et al. Deep learning predictions of TCR-epitope interactions reveal epitope-specific chains in dual alpha T cells. Nat Commun. 2024;15(1):3211. 10.1038/s41467-024-47461-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 242.Jiang M, Yu Z, Lan X. VitTCR: a deep learning method for peptide recognition prediction. iScience. 2024;27(5):109770. 10.1016/j.isci.2024.109770. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 243.Ge J, Wang J, Ye Q, Pan L, Kang Y, Shen C, et al. TRAP: a contrastive learning-enhanced framework for robust TCR–pMHC binding prediction with improved generalizability. Chem Sci. 2025;16(22):9881. 10.1039/D4SC08141B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 244.Henikoff S, Henikoff JG. Amino acid substitution matrices from protein blocks. Proc Natl Acad Sci USA. 1992 [cited 2025 Dec 14];89(22):10915–19. Available from: https://doi.org/10.1073/pnas.89.22.10915; https://pubmed.ncbi.nlm.nih.gov/1438297/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 245.Kawashima S, Ogata H, Kanehisa M. Aaindex: amino acid index database. Nucleic Acids Res [Internet]. 1999 [cited 2025 Dec 14];27(1):368–69. Available from: https://doi.org/10.1093/nar/27.1.368; https://pubmed.ncbi.nlm.nih.gov/9847231/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 246.Atchley WR, Zhao J, Fernandes AD, Drüke T. Solving the protein sequence metric problem. Proc Natl Acad Sci USA. 2005 [cited 2025 Dec 14];102(18):6395–400. Available from: https://doi.org/10.1073/pnas.0408677102; https://pubmed.ncbi.nlm.nih.gov/15851683/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 247.Springer I, Tickotsky N, Louzoun Y. Contribution of T cell receptor alpha and beta CDR3, MHC typing, V and J genes to peptide binding prediction. Front Immunol. 2021;12:664514. 10.3389/fimmu.2021.664514. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 248.Moris P, De Pauw J, Gielis S, De Neuter N, Bittremieux W, Ogunjimi B, et al. Current challenges for unseen-epitope TCR interaction prediction and a new perspective derived from image classification. Briefings Bioinf [Internet]. 2021 [cited 2025 Dec 14];22(4). Available from: https://doi.org/10.1093/bib/bbaa318; https://pubmed.ncbi.nlm.nih.gov/33346826/. [DOI] [PMC free article] [PubMed]
- 249.Xu Z, Luo M, Lin W, Xue G, Wang P, Jin X, et al. DLpTCR: an ensemble deep learning framework for predicting immunogenic peptide recognized by T cell receptor. Briefings Bioinf [Internet]. 2021 [cited 2025 Dec 14];22(6). Available from: https://doi.org/10.1093/bib/bbab335; https://pubmed.ncbi.nlm.nih.gov/34415016/. [DOI] [PubMed]
- 250.Weber A, Born J, Rodriguez MM. TITAN: T-cell receptor specificity prediction with bimodal attention networks. Bioinformatics (Oxford, England) [Internet]. 2021 [cited 2025 Dec 14];37(Supplement_1):i237–44. Available from: https://doi.org/10.1093/bioinformatics/btab294; https://pubmed.ncbi.nlm.nih.gov/34252922/. [DOI] [PMC free article] [PubMed]
- 251.Jensen MF, Nielsen M. Enhancing TCR specificity predictions by combined pan- and peptide-specific training, loss-scaling, and sequence similarity integration. eLife [Internet]. 2024 [cited 2025 Dec 14];12. Available from: https://doi.org/10.7554/eLife.93934; https://pubmed.ncbi.nlm.nih.gov/38437160/. [DOI] [PMC free article] [PubMed]
- 252.Lu T, Zhang Z, Zhu J, Wang Y, Jiang P, Xiao X, et al. Deep learning-based prediction of the T cell receptor-antigen binding specificity. Nat Mach Intel [Internet]. 2021 [cited 2025 Dec 14];3. Available from: https://pubmed.ncbi.nlm.nih.gov/36003885/. [DOI] [PMC free article] [PubMed]
- 253.Kwee BPY, Messemaker M, Marcus E, Oliveira G, Scheper W, Wu C, et al. STAPLER: efficient learning of TCR-peptide specificity prediction from full-length TCR-peptide data [Internet]. bioRxiv. 2023. Available from: http://biorxiv.org/lookup/doi/10.1101/2023.04.25.538237.
- 254.Zhang J, Ma W, Yao H. Accurate TCR-pMHC interaction prediction using a BERT-based transfer learning method. Briefings Bioinf. 2023;25(1). Available from: 10.1093/bib/bbad436. [DOI] [PMC free article] [PubMed]
- 255.Korpela D, Jokinen E, Dumitrescu A, Huuhtanen J, Mustjoki S, Lähdesmäki H. EPIC-TRACE: predicting TCR binding to unseen epitopes using attention and contextualized embeddings. Bioinformatics (Oxford, England) [Internet]. 2023 [cited 2025 Dec 14];39(12). Available from: https://doi.org/10.1093/bioinformatics/btad743; https://pubmed.ncbi.nlm.nih.gov/38070156/. [DOI] [PMC free article] [PubMed]
- 256.Yadav S, Vora DS, Sundar D, Dhanjal JK. TCR-ESM: employing protein language embeddings to predict TCR-peptide-MHC binding. Comput Struct Biotechnol J. 2024;23:165–73. 10.1016/j.csbj.2023.11.037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 257.Yin R, Ribeiro-Filho HV, Lin V, Gowthaman R, Cheung M, Pierce BG. TCRmodel2: high-resolution modeling of T cell receptor recognition using deep learning. Nucleic Acids Res. 2023;51(W1):W569–76. 10.1093/nar/gkad356. [DOI] [PMC free article] [PubMed]
- 258.Deleuran SN, Nielsen M. NetTCR-struc, a structure driven approach for prediction of TCR-pMHC interactions. Front Immunol. 2025;16:1616328. 10.3389/fimmu.2025.1616328. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 259.Tung C-W, Ziehm M, Kämper A, Kohlbacher O, Ho S-Y. POPISK: T-cell reactivity prediction using support vector machines and string kernels. BMC Bioinf. 2011;12(1):446. 10.1186/1471-2105-12-446. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 260.Saethang T, Hirose O, Kimkong I, Tran VA, Dang XT, Nguyen LAT, et al. PAAQD: predicting immunogenicity of MHC class I binding peptides using amino acid pairwise contact potentials and quantum topological molecular similarity descriptors. J Immunol Methods. 2013;387(1–2):293–302. 10.1016/j.jim.2012.09.016. [DOI] [PubMed] [Google Scholar]
- 261.Schubert B, Brachvogel H-P, Jürges C, Kohlbacher O. EpiToolKit—a web-based workbench for vaccine design. Bioinformatics. 2015;31(13):2211–13. 10.1093/bioinformatics/btv116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 262.Paul S, Sidney J, Sette A, Peters B. TepiTool: a pipeline for computational prediction of T cell epitope candidates. Curr Protoc Immunol. 2016;114(1):.18.19.1–18.19.24. 10.1002/cpim.12. [DOI] [PMC free article] [PubMed]
- 263.Glanville J, Huang H, Nau A, Hatton O, Wagar LE, Rubelt F, et al. Identifying specificity groups in the T cell receptor repertoire. Nature [Internet]. 2017 [cited 2025 Dec 15];547(7661):94–98. Available from: https://doi.org/10.1038/nature22976; https://pubmed.ncbi.nlm.nih.gov/28636589/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 264.Dash P, Fiore-Gartland AJ, Hertz T, Wang GC, Sharma S, Souquette A, et al. Quantifiable predictive features define epitope-specific T cell receptor repertoires. Nature [Internet]. 2017 [cited 2025 Nov 27];547(7661):89–93. Available from: https://doi.org/10.1038/nature22383; https://pubmed.ncbi.nlm.nih.gov/28636592/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 265.Schneidman-Duhovny D, Khuri N, Dong GQ, Winter MB, Shifrut E, Friedman N, et al. Predicting CD4 T-cell epitopes based on antigen cleavage, MHCII presentation, and TCR recognition. PLoS One [Internet]. 2018 [cited 2025 Dec 15];13. Available from: https://pubmed.ncbi.nlm.nih.gov/30399156/. [DOI] [PMC free article] [PubMed]
- 266.Riley TP, Keller GLJ, Smith AR, Davancaze LM, Arbuiso AG, Devlin JR, et al. Structure based prediction of neoantigen immunogenicity. Front Immunol. 2019 [cited 2025 Dec 15];10. Available from: https://doi.org/10.3389/fimmu.2019.02047; https://pubmed.ncbi.nlm.nih.gov/31555277/. [DOI] [PMC free article] [PubMed]
- 267.Besser H, Yunger S, Merhavi-Shoham E, Cohen CJ, Louzoun Y. Level of neo-epitope predecessor and mutation type determine T cell activation of MHC binding peptides. J Immunother Cancer [Internet]. 2019 [cited 2025 Nov 24];7(1). Available from: https://doi.org/10.1186/s40425-019-0595-z; https://pubmed.ncbi.nlm.nih.gov/31118084/. [DOI] [PMC free article] [PubMed]
- 268.Gielis S, Moris P, Bittremieux W, De Neuter N, Ogunjimi B, Laukens K, et al. Detection of enriched T cell epitope specificity in full T cell receptor sequence repertoires. Front Immunol. 2019 [cited 2025 Nov 27];10. Available from: https://doi.org/10.3389/fimmu.2019.02820; https://pubmed.ncbi.nlm.nih.gov/31849987/. [DOI] [PMC free article] [PubMed]
- 269.Smith CC, Chai S, Washington AR, Lee SJ, Landoni E, Field K, et al. Machine-learning prediction of tumor antigen immunogenicity in the selection of therapeutic epitopes. Cancer Immunol Res. 2019;7(10):1591–604. 10.1158/2326-6066.CIR-19-0155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 270.Jf BL, Herrera BL, Farias JG. TTAgP 1.0: a computational tool for the specific prediction of tumor T cell antigens. Comput Biol Chem [Internet]. 2019 [cited 2025 Nov 30];83:107103. Available from: https://doi.org/10.1016/j.compbiolchem.2019.107103; https://pubmed.ncbi.nlm.nih.gov/31437642/. [DOI] [PubMed] [Google Scholar]
- 271.Meng Q, Wu Y, Sui X, Meng J, Wang T, Lin Y, et al. POTN: a human leukocyte antigen-A2 immunogenic peptides screening model and its applications in tumor antigens prediction. Front Immunol. 2020 [cited 2025 Nov 30];11. Available from: https://doi.org/10.3389/fimmu.2020.02193; https://pubmed.ncbi.nlm.nih.gov/33133063/. [DOI] [PMC free article] [PubMed]
- 272.Wang G, Wan H, Jian X, Li Y, Ouyang J, Tan X, et al. Ineo-epp: a novel T-Cell HLA class-I immunogenicity or neoantigenic epitope prediction method based on sequence-related amino acid features. Biomed Res Int. 2020;2020(1):5798356. 10.1155/2020/5798356. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 273.Charoenkwan P, Nantasenamat C, Hasan MM, Shoombuatong W. iTTCA-Hybrid: improved and robust identification of tumor T cell antigens by utilizing hybrid feature representation. Analytical Biochem. 2020;599:113747. 10.1016/j.ab.2020.113747. [DOI] [PubMed] [Google Scholar]
- 274.Li G, Iyer B, Prasath VBS, Ni Y, Salomonis N. DeepImmuno: deep learning-empowered prediction and generation of immunogenic peptides for T-cell immunity. Briefings Bioinf. 2021;22(6):22. Available from: 10.1093/bib/bbab160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 275.Jokinen E, Huuhtanen J, Mustjoki S, Heinonen M, Lähdesmäki H. Predicting recognition between T cell receptors and epitopes with TCRGP. PLoS Comput Biol. 2021;17(3):e1008814. 10.1371/journal.pcbi.1008814. [DOI] [PMC free article] [PubMed]
- 276.Zhang W, Hawkins PG, He J, Gupta NT, Liu J, Choonoo G, et al. A framework for highly multiplexed dextramer mapping and prediction of T cell receptor sequences to antigen specificity. Sci Adv [Internet]. 2021 [cited 2025 Nov 27];7(20). Available from: https://doi.org/10.1126/sciadv.abf5835; https://pubmed.ncbi.nlm.nih.gov/33990328/. [DOI] [PMC free article] [PubMed]
- 277.Chronister WD, Crinklaw A, Mahajan S, Vita R, Koşaloğlu-Yalçın Z, Yan Z, et al. Tcrmatch: predicting T-Cell receptor specificity based on sequence similarity to Previously characterized receptors. Front Immunol. 2021;12:640725. 10.3389/fimmu.2021.640725. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 278.Jiao S, Zou Q, Guo H, Shi L. iTTCA-RF: a random forest predictor for tumor T cell antigens. J Transl Med. 2021;19(1):449. 10.1186/s12967-021-03084-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 279.Borden ES, Ghafoor S, Buetow KH, LaFleur BJ, Wilson MA, Hastings KT. NeoScore Integrates characteristics of the neoantigen: MHC class I interaction and expression to accurately prioritize immunogenic Neoantigens. J Immunol. 2022 [cited 2025 Nov 25];208(7):1813–27. Available from: https://doi.org/10.4049/jimmunol.2100700; https://pubmed.ncbi.nlm.nih.gov/35304420/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 280.Xu Y, Qian X, Tong Y, Li F, Wang K, Zhang X, et al. AttnTAP: a dual-input framework incorporating the attention mechanism for accurately predicting TCR-peptide binding. Front Genet. 2022 [cited 2025 Nov 27];13. Available from: https://doi.org/10.3389/fgene.2022.942491; https://pubmed.ncbi.nlm.nih.gov/36072653/. [DOI] [PMC free article] [PubMed]
- 281.Cai M, Bang S, Zhang P, Lee H. ATM-TCR: TCR-Epitope binding affinity prediction using a multi-head self-attention model. Front Immunol. 2022;13:893247. 10.3389/fimmu.2022.893247. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 282.Diao K, Chen J, Wu T, Wang X, Wang G, Sun X, et al. Seq2Neo: a comprehensive pipeline for cancer neoantigen immunogenicity prediction. Int J Mol Sci [Internet]. 2022;23. Available from: 10.3390/ijms231911624. [DOI] [PMC free article] [PubMed]
- 283.Pham M-D, Nguyen T-N, Tran LS, Nguyen Q-T, Nguyen T-P, Pham TMQ, et al. epiTCR: a highly sensitive predictor for TCR-peptide binding. Bioinformatics [Internet]. 2023;39(5). Available from: 10.1093/bioinformatics/btad284. [DOI] [PMC free article] [PubMed]
- 284.Liu S, Bradley P, Sun W. Neural network models for sequence-based TCR and HLA association prediction. PLoS Comput Biol. 2023 [cited 2025 Dec 15];19(11):e1011664. Available from: https://doi.org/10.1371/journal.pcbi.1011664; https://pubmed.ncbi.nlm.nih.gov/37983288/. [DOI] [PMC free article] [PubMed]
- 285.Yang M, Huang Z-A, Zhou W, Ji J, Zhang J, He S, et al. MIX-TPI: a flexible prediction framework for TCR-pMHC interactions based on multimodal representations. Bioinformatics [Internet]. 2023;39(8). Available from: 10.1093/bioinformatics/btad475. [DOI] [PMC free article] [PubMed]
- 286.Gao Y, Gao Y, Fan Y, Zhu C, Wei Z, Zhou C, et al. Pan-peptide Meta learning for T-cell receptor-antigen binding recognition. Nat Mach Intell. 2023;5(3):236–49. 10.1038/s42256-023-00619-3. [Google Scholar]
- 287.Tickotsky N. POP-UP TCR: prediction of previously unseen paired TCR-pMHC [Internet]. bioRxiv. bioRxiv. 2023. Available from: http://biorxiv.org/lookup/doi/10.1101/2023.09.28.560071.
- 288.Smirnov AS, Rudik AV, Filimonov DA, Lagunin AA. TCR-Pred: a new web-application for prediction of epitope and MHC specificity for CDR3 TCR sequences using molecular fragment descriptors. Immunology. 2023;169(4):447–53. 10.1111/imm.13641. [DOI] [PubMed] [Google Scholar]
- 289.Zhang Y, Jian X, Xu L, Zhao J, Lu M, Lin Y, et al. iTcep: a deep learning framework for identification of T cell epitopes by harnessing fusion features. Front Genet. 2023;14:1141535. 10.3389/fgene.2023.1141535. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 290.Myronov A, Mazzocco G, Król P, Plewczynski D. Bertrand-peptide: TCR binding prediction using Bidirectional encoder representations from transformers augmented with random TCR pairing. Bioinformatics [Internet]. 2023;39. Available from: 8). 10.1093/bioinformatics/btad468. [DOI] [PMC free article] [PubMed]
- 291.Müller M, Huber F, Arnaud M, Kraemer AI, Altimiras ER, Michaux J, et al. Machine learning methods and harmonized datasets improve immunogenic neoantigen prediction. Immunity. 2023;56(11):2650–63.e6. 10.1016/j.immuni.2023.09.002. [DOI] [PubMed] [Google Scholar]
- 292.Yu Z, Jiang M, Lan X. HeteroTCR: a heterogeneous graph neural network-based method for predicting peptide-TCR interaction. Commun Biol. 2024;7(1):684. 10.1038/s42003-024-06380-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 293.Ji H, Wang X-X, Zhang Q, Zhang C, Zhang H-M. Predicting TCR sequences for unseen antigen epitopes using structural and sequence features. Briefings Bioinf. 2024;25(3). Available from: 10.1093/bib/bbae210. [DOI] [PMC free article] [PubMed]
- 294.O’Brien H, Salm M, Morton LT, Szukszto M, O’Farrell F, Boulton C, et al. A modular protein language modelling approach to immunogenicity prediction. PLoS Comput Biol. 2024 [cited 2025 Nov 30];20(11):e1012511. Available from: https://doi.org/10.1371/journal.pcbi.1012511; https://pubmed.ncbi.nlm.nih.gov/39527593/. [DOI] [PMC free article] [PubMed]
- 295.Xu L, Yang Q, Dong W, Li X, Wang K, Dong S, et al. Meta learning for mutant HLA class I epitope immunogenicity prediction to accelerate cancer clinical immunotherapy. Briefings Bioinf. 2024;26(1):bbae625. 10.1093/bib/bbae625. [DOI] [PMC free article] [PubMed]
- 296.Tran T-O, Le NQK. Sa-TTCA: an SVM-based approach for tumor T-cell antigen classification using features extracted from biological sequencing and natural language processing. Comput Biol Med. 2024;174:108408. 10.1016/j.compbiomed.2024.108408. [DOI] [PubMed] [Google Scholar]
- 297.Ye F, Chen M, Huang Y, Zhang R, Li X, Wang X, et al. LightCTL: lightweight contrastive TCR-pMHC specificity learning with context-aware prompt. Briefings Bioinf. 2025;26(3). Available from: 10.1093/bib/bbaf246. [DOI] [PMC free article] [PubMed]
- 298.Que J, Xue G, Wang T, Jin X, Wang Z, Cai Y, et al. Identifying T cell antigen at the atomic level with graph convolutional network. Nat Commun [Internet]. 2025 [cited 2025 Dec 15];16(1). Available from: https://doi.org/10.1038/s41467-025-60461-6; https://pubmed.ncbi.nlm.nih.gov/40467559/. [DOI] [PMC free article] [PubMed]
- 299.Zhao Y, Yu J, Su Y, Shu Y, Ma E, Wang J, et al. A unified deep framework for peptide-major histocompatibility complex-T cell receptor binding prediction. Nat Mach Intell. 2025;7(4):650–60. 10.1038/s42256-025-01002-0. [Google Scholar]
- 300.Shao Y, Ge S, Dong R, Ji W, Qin C, Wen P. NeoTImmuML: a machine learning-based prediction model for human tumor neoantigen immunogenicity. Front Immunol. 2025 [cited 2025 Nov 30];16. Available from: https://doi.org/10.3389/fimmu.2025.1681396; https://pubmed.ncbi.nlm.nih.gov/41200173/. [DOI] [PMC free article] [PubMed]
- 301.Zhang J, Mardis ER, Maher CA. INTEGRATE-neo: a pipeline for personalized gene fusion neoantigen discovery. Bioinformatics. 2016;33(4):555. 10.1093/bioinformatics/btw674. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 302.Schubert B, Walzer M, Brachvogel HP, Szolek A, Mohr C, Kohlbacher O. FRED 2: an immunoinformatics framework for Python. Bioinformatics (Oxford, England) [Internet]. 2016 [cited 2025 Dec 1];32(13):2044–46. Available from: https://doi.org/10.1093/bioinformatics/btw113; https://pubmed.ncbi.nlm.nih.gov/27153717/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 303.Hundal J, Carreno BM, Petti AA, Linette GP, Griffith OL, Mardis ER, et al. pVAC-seq: a genome-guided in silico approach to identifying tumor neoantigens. Genome Med. 2016;8(1):11. 10.1186/s13073-016-0264-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 304.Chang TC, Carter RA, Li Y, Li Y, Wang H, Edmonson MN, et al. The neoepitope landscape in pediatric cancers. Genome Med [Internet]. 2017 [cited 2025 Nov 25];9(1). Available from: https://doi.org/10.1186/s13073-017-0468-3; https://pubmed.ncbi.nlm.nih.gov/28854978/. [DOI] [PMC free article] [PubMed]
- 305.Bjerregaard AM, Nielsen M, Hadrup SR, Szallasi Z, Eklund AC. MuPeXI: prediction of neo-epitopes from tumor sequencing data. Cancer Immunol Immunother. 2017 [cited 2025 Nov 28];66(9):1123–30. Available from: https://doi.org/10.1007/s00262-017-2001-3; https://pubmed.ncbi.nlm.nih.gov/28429069/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 306.Tappeiner E, Finotello F, Charoentong P, Mayer C, Rieder D, Trajanoski Z. Timiner: NGS data mining pipeline for cancer immunology and immunotherapy. Bioinformatics (Oxford, England) [Internet]. 2017 [cited 2025 Dec 1];33(19):3140–41. Available from: https://doi.org/10.1093/bioinformatics/btx377; https://pubmed.ncbi.nlm.nih.gov/28633385/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 307.Bais P, Namburi S, Gatti DM, Zhang X, Chuang JH. CloudNeo: a cloud pipeline for identifying patient-specific tumor neoantigens. Bioinformatics (Oxford, England) [Internet]. 2017 [cited 2025 Dec 1];33(19):3110–12. Available from: https://doi.org/10.1093/bioinformatics/btx375; https://pubmed.ncbi.nlm.nih.gov/28605406/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 308.Rubinsteyn A, Hodes I, Kodysh J, Hammerbacher J. Vaxrank: a computational tool for designing personalized cancer vaccines [Internet]. bioRxiv. bioRxiv. 2017. Available from: http://biorxiv.org/lookup/doi/10.1101/142919.
- 309.Mondet S, Aksoy BA, Rozenberg L, Hodes I, Hammerbacher J. Bioinformatics workflow management with the wobidisco ecosystem [Internet]. bioRxiv. bioRxiv. 2017. Available from: http://biorxiv.org/lookup/doi/10.1101/213884.
- 310.Kim S, Kim HS, Kim E, Lee MG, Shin E-C, Paik S, et al. Neopepsee: accurate genome-level prediction of neoantigens by harnessing sequence and amino acid immunogenicity information. Ann Oncol. 2018;29(4):1030–36. 10.1093/annonc/mdy022. [DOI] [PubMed] [Google Scholar]
- 311.Rubinsteyn A, Kodysh J, Hodes I, Mondet S, Aksoy BA, Finnigan JP, et al. Computational pipeline for the PGV-001 neoantigen vaccine trial. Front Immunol. 2018;8:301047. 10.3389/fimmu.2017.01807. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 312.Zhou C, Wei Z, Zhang Z, Zhang B, Zhu C, Chen K, et al. pTuneos: prioritizing tumor neoantigens from next-generation sequencing data. Genome Med. 2019;11(1):67. 10.1186/s13073-019-0679-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 313.Richman LP, Vonderheide RH, Rech AJ. Neoantigen Dissimilarity to the self-Proteome predicts immunogenicity and response to immune checkpoint blockade. Cell Syst [Internet]. 2019 [cited 2025 Nov 28];9(4):375–82.e4. Available from: https://doi.org/10.1016/j.cels.2019.08.009; https://pubmed.ncbi.nlm.nih.gov/31606370/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 314.Schenck RO, Lakatos E, Gatenbee C, Graham TA, Anderson ARA. NeoPredPipe: high-throughput neoantigen prediction and recognition potential pipeline. BMC Bioinf. 2019;20(1):264. 10.1186/s12859-019-2876-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 315.Wang TY, Wang L, Alam SK, Hoeppner LH, Yang R. ScanNeo: identifying indel-derived neoantigens using RNA-Seq data. Bioinformatics (Oxford, England) [Internet]. 2019 [cited 2025 Dec 1];35(20):4159–61. Available from: https://doi.org/10.1093/bioinformatics/btz193; https://pubmed.ncbi.nlm.nih.gov/30887025/. [DOI] [PubMed] [Google Scholar]
- 316.Wood MA, Nguyen A, Struck AJ, Ellrott K, Nellore A, Thompson RF. Neoepiscope improves neoepitope prediction with multivariant phasing. Bioinformatics (Oxford, England) [Internet]. 2020 [cited 2025 Dec 1];36(3):713–20. Available from: https://doi.org/10.1093/bioinformatics/btz653; https://pubmed.ncbi.nlm.nih.gov/31424527/. [DOI] [PubMed] [Google Scholar]
- 317.Fotakis G, Rieder D, Haider M, Trajanoski Z, Finotello F. NeoFuse: predicting fusion neoantigens from RNA sequencing data. Bioinformatics. 2020;36(7):2260–61. 10.1093/bioinformatics/btz879. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 318.Kodysh J, Rubinsteyn A. OpenVax: an open-source computational pipeline for cancer neoantigen prediction. Methods Mol Biol (Clifton, NJ) [Internet]. 2020 [cited 2025 Nov 29];2120. Available from: https://pubmed.ncbi.nlm.nih.gov/32124317/. [DOI] [PubMed]
- 319.Hundal J, Kiwala S, McMichael J, Miller CA, Xia H, Wollam AT, et al. pVactools: a computational toolkit to identify and visualize cancer neoantigens. Cancer Immunol Res. 2020;8(3):409–20. 10.1158/2326-6066.CIR-19-0401. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 320.Coelho ACMF, Fonseca AL, Martins DL, Lins PBR, da Cunha LM, de Souza SJ. neoANT-HILL: an integrated tool for identification of potential neoantigens. BMC Med Genomics. 2020;13(1):30. 10.1186/s12920-020-0694-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 321.Tang Y, Wang Y, Wang J, Li M, Peng L, Wei G, et al. TruNeo: an integrated pipeline improves personalized true tumor neoantigen identification. BMC Bioinf. 2020;21(1):532. 10.1186/s12859-020-03869-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 322.Lang F, Riesgo-Ferreiro P, Löwer M, Sahin U, Schrörs B. NeoFox: annotating neoantigen candidates with neoantigen features. Bioinformatics. 2021;37(22):4246–47. 10.1093/bioinformatics/btab344. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 323.Zhou Z, Wu J, Ren J, Chen W, Zhao W, Gu X, et al. TSNAD v2.0: a one-stop software solution for tumor-specific neoantigen detection. Comput Struct Biotechnol J. 2021 [cited 2025 Nov 28];19:4510–16. Available from: https://doi.org/10.1016/j.csbj.2021.08.016; https://pubmed.ncbi.nlm.nih.gov/34471496/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 324.Sharpnack MF, Johnson TS, Chalkley R, Han Z, Carbone D, Huang K, et al. Tsafinder: exhaustive tumor-specific antigen detection with RNAseq. Bioinformatics (Oxford, England) [Internet]. 2022 [cited 2025 Dec 2];38(9):2422–27. Available from: https://doi.org/10.1093/bioinformatics/btac116; https://pubmed.ncbi.nlm.nih.gov/35191489/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 325.Rieder D, Fotakis G, Ausserhofer M, René G, Paster W, Trajanoski Z, et al. nextNeopi: a comprehensive pipeline for computational neoantigen prediction. Bioinformatics. 2022;38(4):1131–32. 10.1093/bioinformatics/btab759. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 326.Liu C, Zhang Y, Jian X, Tan X, Lu M, Ouyang J, et al. ProGeo-neo v2.0: a One-stop software for neoantigen prediction and filtering based on the proteogenomics strategy. Genes [Internet]. 2022 [cited 2025 Dec 14];13(5):783. Available from: https://doi.org/10.3390/genes13050783; https://pubmed.ncbi.nlm.nih.gov/35627168/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 327.Wu J, Chen W, Zhou Y, Chi Y, Hua X, Wu J, et al. Tsnadb v2.0: the updated version of tumor-specific neoantigen database. Genomics Proteomics Bioinf [Internet]. 2023 [cited 2025 Dec 2];21(2):259–66. Available from: https://doi.org/10.1016/j.gpb.2022.09.012; https://pubmed.ncbi.nlm.nih.gov/36209954/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 328.Tan X, Xu L, Jian X, Ouyang J, Hu B, Yang X, et al. Pgnneo: a proteogenomics-based neoantigen prediction pipeline in noncoding regions. Cells [Internet]. 2023 [cited 2025 Dec 2];12(5):782. https://doi.org/10.3390/cells12050782; https://pubmed.ncbi.nlm.nih.gov/36899918/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 329.Vensko SP, Olsen K, Bortone D, Smith CC, Chai S, Beckabir W, et al. LENS: landscape of effective neoantigens software. Bioinformatics (Oxford, England) [Internet]. 2023 [cited 2025 Nov 29];39(6). Available from: https://doi.org/10.1093/bioinformatics/btad322; https://pubmed.ncbi.nlm.nih.gov/37184881/. [DOI] [PMC free article] [PubMed]
- 330.Li B, Jing P, Zheng G, Pi C, Zhang L, Yin Z, et al. Neo-intline: integrated pipeline enables neoantigen design through the in-silico presentation of T-cell epitope. Sig Transduct Target Ther. 2023;8(1):397. 10.1038/s41392-023-01644-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 331.Ma T, Zhao Z, Li H, Wei L, Zhang X. NeoHunter: flexible software for systematically detecting neoantigens from sequencing data. Quant Biol. 2024;12(1):70–84. 10.1002/qub2.28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 332.Chuwdhury GS, Guo Y, Chiang C-L, Lam K-O, Kam N-W, Liu Z, et al. ImmuneMirror: a machine learning-based integrative pipeline and web server for neoantigen prediction. Briefings Bioinf. 2024;25(2). Available from: 10.1093/bib/bbae024. [DOI] [PMC free article] [PubMed]
- 333.Kote S, Pirog A, Bedran G, Alfaro J, Dapic I. Mass spectrometry-based identification of MHC-Associated peptides. Cancers [Internet]. 2020 [cited 2025 Dec 14];12(3):535. Available from: https://doi.org/10.3390/cancers12030535; https://pubmed.ncbi.nlm.nih.gov/32110973/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 334.Cai Y, Gong M, Zeng M, Leng F, Lv D, Guo J, et al. Immunopeptidomics-guided discovery and characterization of neoantigens for personalized cancer immunotherapy. Sci Adv. 2025;11(21):eadv 6445. 10.1126/sciadv.adv6445. [DOI] [PMC free article] [PubMed]
- 335.Richters MM, Xia H, Campbell KM, Gillanders WE, Griffith OL, Griffith M. Best practices for bioinformatic characterization of neoantigens for clinical utility. Genome Med [Internet]. 2019 [cited 2025 Dec 14];11(1). Available from: https://doi.org/10.1186/s13073-019-0666-2; https://pubmed.ncbi.nlm.nih.gov/31462330/. [DOI] [PMC free article] [PubMed]
- 336.Kim P, Tan H, Liu J, Lee H, Jung H, Kumar H, et al. FusionGDB 2.0: fusion gene annotation updates aided by deep learning. Nucleic Acids Res. 2022;50(D1):D1221–30. 10.1093/nar/gkab1056. [DOI] [PMC free article] [PubMed]
- 337.Hu X, Wang Q, Tang M, Barthel F, Amin S, Yoshihara K, et al. TumorFusions: an integrative resource for cancer-associated transcript fusions. Nucleic Acids Res. 2018;46(D1):D1144–9. 10.1093/nar/gkx1018. [DOI] [PMC free article] [PubMed]
- 338.Kumar H, Luo R, Wen J, Yang C, Zhou X, Kim P. FusionNeoAntigen: a resource of fusion gene-specific neoantigens. Nucleic Acids Res. 2024;52(D1):D1276–88. 10.1093/nar/gkad922. [DOI] [PMC free article] [PubMed]
- 339.Shi Q, Li X, Liu Y, Chen Z, He X. Flibase: a comprehensive repository of full-length isoforms across human cancers and tissues. Nucleic Acids Res. 2024;52(D1):D124–33. 10.1093/nar/gkad745. [DOI] [PMC free article] [PubMed]
- 340.Cai Y, Lv D, Li D, Yin J, Ma Y, Luo Y, et al. Ieatlas: an atlas of HLA-presented immune epitopes derived from non-coding regions. Nucleic Acids Res. 2023;51(D1):D409–17. 10.1093/nar/gkac776. [DOI] [PMC free article] [PubMed]
- 341.Zhang Y, She J, Hu X, Jin Y, Tao C, Du M, et al. Teitbase: a database for transposable element (TE)-initiated transcripts in human cancers. Database (Oxford) [Internet]. 2026;2026. Available from: 10.1093/database/baag025. [DOI] [PMC free article] [PubMed]
- 342.Yi X, Liao Y, Wen B, Li K, Dou Y, Savage SR, et al. caAtlas: an immunopeptidome atlas of human cancer. iScience. 2021;24(10):103107. 10.1016/j.isci.2021.103107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 343.Manoutcharian K, Guzman VJ, Gevorkian G. Neoantigen cancer vaccines: real opportunity or another illusion? Arch Immunol Ther Exp. 2021 [cited 2025 Dec 14];69(1). Available from: https://doi.org/10.1007/s00005-021-00615-8; https://pubmed.ncbi.nlm.nih.gov/33909124/. [DOI] [PMC free article] [PubMed]
- 344.Ward JP, Gubin MM, Schreiber RD. The role of neoantigens in naturally occurring and therapeutically induced immune responses to cancer. Adv Immunol [Internet]. 2016 [cited 2025 Dec 14];130. Available from: https://pubmed.ncbi.nlm.nih.gov/26922999/. [DOI] [PMC free article] [PubMed]
- 345.Goloudina A, Le Chevalier F, Authié P, Charneau P, Majlessi L. Shared neoantigens for cancer immunotherapy. Mol Ther Oncol [Internet]. 2025 [cited 2025 Dec 14];33(2):200978. Available from: https://doi.org/10.1016/j.omton.2025.200978; https://pubmed.ncbi.nlm.nih.gov/40256120/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 346.Niemi JVL, Sokolov AV, Schiöth HB. Neoantigen vaccines; clinical trials, classes, indications, adjuvants and combinatorial treatments. Cancers [Internet]. 2022 [cited 2025 Dec 14];14(20):5163. Available from: https://doi.org/10.3390/cancers14205163; https://pubmed.ncbi.nlm.nih.gov/36291947/. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 347.Wood LV, Fojo A, Roberson BD, Hughes MS, Dahut W, Gulley JL, et al. TARP vaccination is associated with slowing in PSA velocity and decreasing tumor growth rates in patients with stage D0 prostate cancer. Oncoimmunology [Internet]. 2016 [cited 2025 Dec 14];5(8):e1197459. Available from: https://doi.org/10.1080/2162402X.2016.1197459; https://pubmed.ncbi.nlm.nih.gov/27622067/. [DOI] [PMC free article] [PubMed]
- 348.Rappaport AR, Kyi C, Lane M, Hart MG, Johnson ML, Henick BS, et al. A shared neoantigen vaccine combined with immune checkpoint blockade for advanced metastatic solid tumors: phase 1 trial interim results. Nat Med. 2024;30(4):1013–22. 10.1038/s41591-024-02851-9. [DOI] [PubMed] [Google Scholar]
- 349.Li M, Gao X, Wang X. Identification of tumor mutation burden-associated molecular and clinical features in cancer by analyzing multi-omics data. Front Immunol. 2023 [cited 2025 Dec 14];14. Available from: https://doi.org/10.3389/fimmu.2023.1090838; https://pubmed.ncbi.nlm.nih.gov/36911742/. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.
