Skip to main content
Journal of Translational Medicine logoLink to Journal of Translational Medicine
. 2026 Jul 6;24:924. doi: 10.1186/s12967-026-08535-x

AI-driven neoantigen identification: a comprehensive review from somatic variant calling to T cell recognition

Atefeh Bakhshian 1,2,3, Sajjad Ghorghanlu 1, Fereshteh Fallah Atanaki 2, Elham Erfani Ezadyar 1,2,3, Azadeh Ashkiyan 1,3, Ali Etemadi 1, Babak Negahdari 1, Kaveh Kavousi 2, Gholamali Kardar 1,3, Mohammadali Mazloomi 1,
PMCID: PMC13378337  PMID: 42410432

Abstract

Background

Neoantigens—tumor-specific peptides generated by somatic mutations—are central targets of effective anticancer T cell immunity and underpin the clinical success of immune checkpoint blockade and personalized cancer vaccines. Advances in high-throughput sequencing, immunopeptidomics, and artificial intelligence (AI) have transformed neoantigen discovery from tailored experimental workflows into scalable, computational pipelines. However, accurately identifying the small subset of tumor mutations that yield processed, presented, and immunogenic epitopes remains a major bottleneck.

Methods

This review summarizes how AI is reshaping neoantigen discovery, from somatic variant calling, HLA typing, and peptide processing to peptide–MHC binding, presentation, and T cell recognition. We first outline the immunobiological foundations of antigen presentation, emphasizing class I and II peptide-binding grooves and their allele-specific motifs, then describe AI workflows that integrate somatic mutation calling, HLA typing, transcriptomics, and immunopeptidomics to nominate candidate neoepitopes. We highlight recent AI-driven tools for presentation and immunogenicity prediction, integrative pipelines that support personal and shared neoantigen targeting, and early clinical applications in vaccination and T cell therapies.

Results

AI-driven models trained on eluted ligand datasets substantially outperform affinity-only predictors for peptide presentation across diverse HLA alleles and populations. Consortium-scale benchmarking demonstrates that integrating features of antigen processing, presentation, and TCR recognition can eliminate the majority of non-immunogenic candidates while retaining clinically relevant neoepitopes. Immunopeptidomics provides essential ground truth, revealing that only a small fraction of genomically predicted candidates are naturally presented and uncovering noncanonical antigen sources, including splice variants, post-translational modifications, and noncoding regions. Integrative pipelines now support both personal (private) and shared (public) neoantigen prioritization, enabling translational applications such as personalized vaccines and TCR-based therapies.

Conclusions

AI-guided neoantigen discovery is now clinically actionable, enabled by immunopeptidomics and deep learning models. Despite significant progress, key challenges remain, including limited class II prediction accuracy, incomplete coverage of rare HLA alleles, tumor heterogeneity, and the need for standardized benchmarking and validation. Anchoring computational predictions to mass spectrometry–derived ligands and incorporating tumor evolution and immune escape mechanisms will be critical for improving target selection. Continued integration of AI, proteogenomics, and clinical data is poised to accelerate the development of effective, precision neoantigen-based cancer immunotherapies.

Keywords: Artificial intelligence, Cancer, Immunotherapy, MHC, Neoantigen

Introduction

Over the past decade, cancer immunotherapy has established T cells as central effectors in the control of diverse human malignancies [13]. Many effective T cell responses are directed against neoantigens, novel peptide sequences created by tumor-specific mutations that allow the immune system to distinguish cancer cells from normal tissues [1]. Recent technological innovations permit systematic characterization of patient-specific neoantigen responses, and emerging data suggest that recognition of such neoantigens is a major factor in the activity of clinical immunotherapies [1]. These observations position neoantigen load as a candidate biomarker for response and provide a compelling rationale to develop therapeutic strategies that selectively enhance T cell reactivity against this class of antigens [1, 4, 5].

Neoantigens arise from somatic alterations such as point mutations, insertions-deletions, gene fusions or aberrant splicing, and because they are absent from the normal proteome, they offer highly tumor-selective targets for immunotherapy [1, 4, 6]. Proof-of-concept trials of individualized neoantigen vaccines in melanoma demonstrated that RNA mutanome vaccines can be manufactured based on each patient’s tumor mutational repertoire, elicit T cell responses against multiple vaccine neo-epitopes and, in some cases, drive durable clinical responses when administered alone or with checkpoint blockade [4]. Broader clinical experience summarized in recent reviews indicates that personalized cancer vaccines can trigger broad-based antitumor immunity and are now being evaluated across a spectrum of solid tumors [7, 8]. In parallel, large-scale immunogenomic studies have shown that tumor mutational burden and predicted neoantigen load correlate with outcome to immune checkpoint inhibitors across multiple cancer types, underscoring the clinical relevance of neoantigen-directed immunity [1, 5].

Realizing the therapeutic promise of neoantigens requires accurately identifying, among thousands of tumor mutations, the small subset that yields processed, presented, and T cell–recognizable epitopes [912]. The accurate identification and prioritization of antigenic peptides have emerged as a central bottleneck for the development of personalized cancer immunotherapies [9]. Contemporary neoantigen discovery pipelines therefore integrate whole-exome or whole-genome sequencing of matched tumor-normal samples, transcriptomics, high-resolution HLA typing, in silico modeling of antigen processing and peptide–MHC (pMHC) presentation, and experimental validation using patient-derived T cells [912]. Complementing genomic approaches, mass spectrometry (MS)-based immunopeptidomics now enables direct profiling of HLA ligand landscapes in primary tumors, uncovering both canonical and non-canonical antigen sources and providing high-confidence ligands that serve as ground truth for model training [911].

Machine learning (ML) and deep learning (DL) models now underpin key steps in neoantigen discovery, including prediction of peptide–MHC binding and presentation, often in a pan-allelic fashion that covers the extensive polymorphism of human HLA [11, 12]. Models trained directly on large immunopeptidome datasets significantly improve class I epitope prediction across diverse alleles and human populations [11]. Consortium-scale benchmarking by the Tumor Neoantigen Selection Alliance (TESLA) further demonstrated that integrating peptide features associated with antigen processing and T cell recognition can filter out the vast majority of non-immunogenic peptides while retaining a substantial fraction of true targets [6, 12]. At the same time, TESLA and subsequent analyses highlight that most candidate peptides prioritized by current algorithms are non-immunogenic in functional assays, emphasizing persistent gaps in our understanding and modeling of antigen processing, presentation and TCR recognition [6, 12].

AI is therefore central to scaling neoantigen discovery from non-standardized experimental efforts to reproducible, clinically deployable pipelines. In this review, we first discuss the structure of major histocompatibility complex (MHC) molecules and how polymorphic pockets shape peptide binding and T cell recognition, then outline current neoantigen discovery workflows from somatic variant calling through integrative pipelines and immunopeptidomics, emphasizing where AI tools are already reshaping practice. Finally, we examine emerging strategies for prioritizing personal versus shared neoantigens, summarize the evolving clinical landscape of neoantigen-targeted therapies, and highlight outstanding challenges and future directions for deploying AI-guided neoantigen discovery in precision oncology.

We conducted a comprehensive and systematic search to identify relevant research articles using the keywords “neoantigen/neoepitope identification” and “artificial intelligence.” The primary objective of this study was to collect and compare AI-based methods for neoantigen prediction across multiple stages, ranging from somatic mutation detection to the generation of immune responses through the recognition of neoantigen–MHC complexes by T cell receptors (TCRs). Articles were not selected based on publication year, as the aim was to evaluate and compare both earlier and more recent models in terms of their modeling approaches and predictive performance. Studies published in languages other than English were excluded. Furthermore, only studies employing AI-based techniques, including ML and DL, and applying these methods across various cancer types were included to enable a comparative assessment of model performance and applicability across different tumor contexts.

The major histocompatibility complex

MHC molecules are divided into two main groups: class I and class II. Class I MHC molecules are present on the surface of all nucleated cells. Conversely, MHC class II molecules present peptides primarily on APCs such as dendritic cells, macrophages, and B cells [13, 14]. In order to perform their physiological role, MHC proteins must initially bind with peptide antigens [14]. Peptides that occupy the peptide-binding groove of surface-expressed MHC molecules are generated through specific proteolytic processes [15]. MHC class I typically presents endogenous (cytosolic) peptides to CD8+ cytotoxic T cells, whereas MHC class II typically presents exogenous (extracellular) peptides to CD4+ helper T cells [14, 16]. The pathways involved in antigen processing and presentation for both MHC class I and class II are outlined in Fig. 1.

Fig. 1.

Fig. 1

MHC class I and class II molecules utilize different antigen processing pathways. a) MHC class I molecules display peptides that mainly originate from proteins produced inside the cell, whether from the host or invading pathogens. These proteins are broken down into peptides by the proteasome, after which the peptides are transported into the endoplasmic reticulum via TAP (transporters associated with antigen processing) to be loaded onto MHC class I molecules. b) Unlike MHC class I, MHC class II molecules present proteins that enter the cell via endocytosis. During their maturation, they are prevented from binding self-peptides in the endoplasmic reticulum by the invariant chain (Ii). These MHC II–Ii complexes pass through the Golgi and reach the late endosome, where the invariant chain is degraded into a smaller fragment called CLIP (class II–associated invariant chain peptide). CLIP is then displaced from the MHC–CLIP complex and substituted with an antigenic peptide. c) Dendritic cells possess the ability to internalize antigens from other cells and subsequently cross-present them to CD8+ cytotoxic T lymphocytes. The reliance of this process on TAP suggests that cross-presentation requires the redirection of cellular antigens into the classical MHC class I pathway, although the precise mechanisms underlying this redirection remain unclear. Typically, these antigens are also processed through the MHC class II pathway, enabling recognition by CD4+ helper T cells

Decoding MHC class I: Insights into polymorphic pockets

MHC class I molecules consist of two parts: a heavy (α) chain paired noncovalently with β2-microglobulin. The α1 and α2 domains of the heavy chain come together to create the peptide-binding groove [17]. This groove is characterized by its closed-ended structure, which accommodates peptides typically 8–11 amino acids in length, though the precise length can vary based on MHC allele and antigen characteristics [18].

The specificity of peptide binding is primarily governed by anchor residues located at specific positions within the peptide sequence, such as P2 and P9. These anchor residues interact with complementary pockets in the MHC class I groove, contributing significantly to the stability and affinity of the peptide-MHC complex [13]. Insights into peptide binding preferences of specific MHC class I alleles facilitate the design of peptide vaccines aimed at eliciting robust CD8+ T cell responses against viral infections or tumor antigens [19]. Moreover, these insights inform strategies for enhancing antigen presentation efficiency in immunotherapies targeting cancer and other diseases [14].

Polymorphism in class I molecules is concentrated in the α1 and α2 domains that form the peptide-binding groove; this polymorphism sculpts pocket architecture and determines allele-specific peptide motifs. Among the classical loci, HLA-B is the most polymorphic and contributes disproportionately to population peptide-presentation diversity, although HLA-A and HLA-C are also highly variable. These allelic differences give rise to supertypes (groups of alleles with overlapping peptide specificities) but also to highly allele-specific preferences that shape T cell responses and population immunity [20].

Individual MHC-I proteins balance allele-specific anchors (P2, PΩ and sometimes P1/P3/P5) with broader tolerance at other peptide positions. Class I motifs often rely on two or three dominant anchors, producing narrower length restrictions than class II but still permitting many sequence variants to bind the same allele. Subdominant (secondary) anchor residues and peptide flanking residues can modulate affinity and TCR recognition but generally have less impact on primary MHC binding than the canonical anchors [21].

In TCR recognition of peptide-MHC I complexes, the CDR2 loops of both alpha and beta chains make contact with only the MHC molecules, whereas the CDR1 and CDR3 loops make contact with both peptide and MHC. Additionally, the V domains of alpha and beta chains are adjacent to the N-terminal and C-terminal of the bound peptide, respectively. The binding region on TCR is almost flat and sometimes has a hole; in contrast, MHC surfaces have two apexes, and the TCR binds between them in interaction with peptide-MHC. Auxiliary systems such as CD3 and CD8 co-receptors on T cells generate a signal to the T cell during peptide-MHC I recognition, which causes T cell signaling and triggers a cytotoxic response and then kills the virally infected or otherwise abnormal cell [17, 18].

Decoding MHC class II: Insights into polymorphic pockets

Antigens are internalized into cells through processes like phagocytosis, pinocytosis, endocytosis, and autophagy, and are delivered to late endosomes. Within these endosomes, they are processed by proteolytic enzymes such as cathepsins and thiol oxidoreductase. MHC-II molecules are loaded with peptides in a specialized late endosomal–lysosomal antigen-processing compartment [22].

MHC-II molecules are dimers composed of an alpha (α) and a beta (β) polypeptide chain. These glycoproteins have an immunoglobulin-like domain (α2 and β2) close to the cell membrane, as well as a peptide-binding cleft positioned farther from the membrane (α1 and β1), which includes the most polymorphic sites of HLA class II molecules. This peptide binding region (PBR) is essential for the molecule’s function [23]. The peptide-binding groove of MHC class II is structured with two alpha-helices above a beta-pleated sheet, designed to specifically bind short peptides. The amino acid sequence around the binding site is highly variable, enhancing the molecule’s ability to bind a wide range of peptides. The ends of the antigen-binding cleft of class II molecules contain small residues, such as glycine or valine, which do not restrict the size of the peptides. Thus, unlike MHC-I, MHC-II molecules possess an open-ended binding groove that can fit longer peptides, generally 13–18 amino acids in length, and occasionally up to 30 amino acids [23, 24].

The specificity of peptide binding in MHC-II molecules is influenced by the interaction between anchor residues (side chains of amino acids) of the peptides and the pockets within the MHC grooves. These anchor residues form electrostatic, van der Waals, and hydrogen bonds with the pockets, stabilizing the peptide-MHC complex. Conserved residues in the MHC-II molecule form hydrogen bonds with the peptide’s backbone, whereas the variable residues create specialized pockets that accommodate the peptide’s side chains. Some peptides have charged residues that interact with the alpha helix. Antigenic peptides that successfully bind to the floor of the MHC class II groove have a specific conserved secondary structure resembling a polyproline chain. This polyproline structure is characterized by its openness and lack of internal hydrogen bonding [25].

MHC-II molecules exhibit broad specificity, facilitated by five cavities that accommodate the side chains of the attached peptide molecules. These pockets vary in size and hydrophobicity, allowing MHC-II molecules to bind a diverse array of peptides. Within the binding groove of MHC class II molecules, there are typically three to four critical anchor pockets, including positions like P1, P4, P6, P7, and occasionally P9, whereas MHC class I molecules have only two major pockets. Secondary or minor anchor residues influence peptide-MHC binding affinity (BA), but their effect is comparatively subtle. Each of these pockets, usually located in beta strands, in MHC class II plays a distinct role in peptide binding and allele specificity. MHC class II proteins generally bind a narrower range of proteins compared to MHC class I proteins. This is because MHC class II molecules have more major anchor pockets, which increases the specificity for peptide binding [26, 27]. The P1 pocket is crucial for anchoring peptides into the HLA class II cleft. It typically accommodates large, hydrophobic residues. P6 and P7 pockets are involved in determining the allele specificity of MHC-II molecules. The P4 and P9 pockets play a dual role by determining antigen specificity and influencing how the antigenic peptide interacts with the groove. The P4 and P7 pockets are composed exclusively of the β chain, whereas P1, P6, and P9 are shaped by contributions from both the α and β chains [26, 27].

Polymorphic amino acids are grouped within the peptide-binding groove of MHC-II molecules, accounting for the allele-specific nature of peptide-binding motifs. This results in strong preferences for certain side chains at specific positions in the peptide, while other positions can accommodate a wide variety of side chains. The primary anchor residues differ among alleles, for example: P1 in HLA-DR1 and P4 in HLA-DR3, while P7 and P9 also contribute to binding and may add extra stabilization. Only a few peptide residues, such as the side chains at P2, P5, and P8, extend outward from the groove to interact with T cell receptors. The amino acid identities at these positions generally do not influence the peptide’s binding affinity to MHC-II unless they undergo specific post-translational chemical modifications [27].

MHC-I and MHC-II molecules present peptides using both sequence-specific and sequence-independent features, reflecting allele-specific and shared properties, respectively. Sequence-dependent interactions are key for determining which peptides fit into the binding groove [23]. A network of sequence-independent hydrogen bonds connects the peptide backbone to conserved residues of MHC-II, extending along the entire peptide. This differs from MHC-I, where conserved hydrogen bonds mainly involve the peptide’s N- and C-terminal ends [25].

Analysis of the peptide-binding region of HLA class II molecules reveals a concentration of polymorphism activity in PBR, specifically in pockets. Among the HLA class II genes, HLA-DRB1 is the most variable. The α-chain of HLA-DR remains constant, which allows the β-chain to undergo more variation. The binding process begins with hydrophobic side chains fitting into the P1 pocket. Amino acids occupying the P6 pocket also strongly affect preferences at P9. Allelic differences can involve up to 56 amino acids, with most of this variation localized in the peptide-binding region [27]. It should be noted that each of the pocket variants has different amino acid sequences which affect their specificity for the peptide residues.

Post-translational modifications (PTMs) like glycosylation, phosphorylation, and acylation are crucial for regulating protein function in numerous physiological processes. These modified epitopes differ in immunogenicity from unmodified peptides. PTMs can serve as anchor points when peptides are presented by MHC molecules. They can also affect protein processing, thereby altering peptide presentation. When interacting with specific MHC molecules, modified peptides display unique characteristics that impact their immunogenicity. For example, in HLA-DR1, phosphate groups engage the peptide-binding groove through water-mediated interactions. Phosphorylated peptides restricted by MHC-II can also modify peptide structure and influence T cell receptor recognition [23].

Studies have shown that peptide flanking residues (PFRs), which lie outside the main binding groove, can affect both peptide-MHC binding (in a sequence-independent way) and T cell recognition. Extended peptides that are shorter than optimal can enhance their binding to MHC class II molecules. In particular, the residue at position P-1 can increase measurable peptide affinity for the MHC groove. The ideal peptide length for MHC class II binding is around 18–20 amino acids, while adding residues beyond this length tends to reduce or abolish binding [24].

Understanding interactions within pMHC complexes and the roles of specific pockets in the MHC-II groove is crucial for developing targeted immunotherapies and vaccines, thereby enhancing our ability to combat infectious diseases and autoimmune conditions. Each population has its own restricted alleles, which are crucial for designing vaccines and immunotherapies. Understanding the specificities and polymorphisms of MHC-II molecules can aid in developing targeted treatments that leverage the immune system’s natural capabilities. By identifying the distinct peptide-binding motifs and anchor pockets of MHC-II molecules, researchers can design more effective vaccines and therapies that elicit robust immune responses [27].

Neoantigen identification workflow

Neoantigens, a subset of tumor-specific antigens (TSAs), are unique, non-self-peptides generated by tumor-acquired genetic alterations and presented on MHC molecules on tumor cells that can be recognized by immune cells [17, 18]. In addition, antigen-presenting cells (APCs) can acquire tumor-derived proteins and process them into peptide fragments for loading onto MHC molecules [13]. The complexes of tumor-specific pMHC are identified by cognate T cell receptors. Such neoantigen-specific T cells can elicit antitumor immune responses in patients [19]. Because neoantigens are not presented during thymic selection, T cells specific for them typically escape central tolerance, enabling a neoantigen-specific repertoire with potentially high functional avidity [28, 29].

Neoantigen discovery pipelines can be divided into several phases. (i) Matched tumor–normal DNA sequencing profiles the somatic variant landscape. (ii) Tumor RNA sequencing (RNA-seq) quantifies gene- and allele-specific expression and reveals transcript-derived alterations, including aberrant splicing, gene fusions, intron retention, and RNA editing. (iii) High-resolution HLA class I/II genotyping, together with tumor HLA status (e.g., loss of heterozygosity, downregulation, or allele-specific expression), constrains the feasible set of presenting molecules. (iv) Antigen processing and presentation are modeled computationally, encompassing proteasomal cleavage to nominate peptides that can be generated from mutant proteins, TAP-mediated transport to estimate endoplasmic reticulum import, and pMHC binding and complex stability to prioritize epitopes with a high likelihood of cell-surface display. (v) Putative immunogenicity is estimated using TCR recognition models that integrate epitope foreignness, self-similarity, TCR-facing residue features, repertoire evidence, and tumor immune context. (vi) Finally, immunopeptidomics (MS-based HLA ligandome profiling) provides empirical observations of presented peptides, enabling refinement and validation of presentation rules and calibration of in silico predictors. In the following sections, we detail each step.

Somatic mutation calling

The body consists of somatic and germline cells. Mutations that occur in somatic cells throughout a person’s lifetime affect only that individual. In contrast, mutations in germline cells can be passed to subsequent generations, serving as the basis for species evolution and leading to hereditary diseases [30, 31]⁠. Detection of variant alleles by using next-generation sequencing (NGS) technology is a common method for identifying tumor neoantigens. Whole-exome sequencing is conducted on tumor tissue and corresponding normal tissue to extract tumor-specific mutations, and the expression level of these mutations is determined by integrating RNA sequencing data ⁠ [3234]. The generation of tumor neoantigens originates from somatic genomic mutations, including single-nucleotide variants (SNVs), insertions and deletions (INDELs), fusion genes [35], ⁠ and splice variants [34].

Two principal strategies have been established for the identification of neoantigen epitopes. The immunogenomic approach employs NGS data to computationally generate virtual peptidomes through in silico prediction methods, whereas the immunopeptidomic strategy utilizes MS to directly characterize peptides presented by MHC molecules [36]. In addition, several TCR-guided neoantigen discovery methods have recently emerged, enabling the systematic mapping and validation of immunogenic neoantigens [19].

Immunogenomic research has advanced rapidly through the application of NGS to compare genetic alterations between tumor and matched normal tissues. The first critical step in identifying potential neoantigens from NGS data typically involves detecting tumor-specific genomic aberrations using whole-exome sequencing (WES) of paired tumor and normal DNA. Integration of RNA sequencing with WES enables assessment of whether mutant alleles are transcriptionally expressed within the tumor. Moreover, RNA-seq provides additional layers of biological insight, including information on copy number variations, microbial contamination, transposable element activity, cellular composition, and the presence of candidate neoantigens [37, 38].

While immunogenomic analyses predict millions of mutation-derived neoantigens, most are not confirmed at the protein level [39, 40]. MS-based immunopeptidomics serves as the gold standard for validating neoantigens by directly identifying MHC-bound peptides, including those arising from post-translational modifications [4146], non-coding RNA, or proteasome splicing features often missed by DNA- or RNA-only approaches [4749]. Integrating MS with NGS and developing user-friendly tools that combine genomic, transcriptomic, and proteomic data are key to improving neoantigen discovery for cancer immunotherapy.

The process of mutation calling starts with quality control of the sequencing reads and aligning them to a reference genome using FastQC screen [50] and BWA [51], respectively. After alignment, the mutation calling step must carefully differentiate true somatic variants from sequencing errors, artifacts introduced during sample preparation, and germline mutations. Numerous software tools have been developed to tackle the key challenges associated with this process. A wide variety of somatic mutation caller tools are listed in Table 1.

Table 1.

Somatic variant calling tools

Model/Year Algorithm Mutation Type Key Features/Web and Code Accessibility Ref.

samtools/bcftools

2009

Compressed block indexing algorithm (BGZF) combined with hierarchical binning indexing Basic variant calling

Versatile in structure, space-efficient, optimized for rapid random access, and used as the standard format for releasing alignments from the 1000 Genomes Project.

http://samtools.sourceforge.net/

[52]

GATK

2011

Haplotype analysis

Multi-step process, including

a germline genotyper

Unified analytic framework for variation discovery and three-phase conceptual pipeline including preprocessing (read mapping, duplicate removal, local realignment, base quality score recalibration), variant discovery, variant refinement.

http://www.broadinstitute.org/gatk/

https://github.com/broadinstitute/gatk

[53]

VarScan 2

2012

Heuristic and statistical classification

Somatic SNVs and INDELs,

germline variants, copy

number variants, loss of heterozygosity (LOH)

Reads data from tumor–normal samples simultaneously. For germline variants, it showed high concordance with SNP arrays (99.56%). For somatic mutations, it achieved high sensitivity (94.7%) and a strong true-positive rate (89%).

https://github.com/dkoboldt/varscan

[54]

SomaticSniper

2012

Bayesian genotype contrast Somatic SNVs

Direct tumor-normal comparison.

Identifies systematic errors and introduces empirical/statistical filters.

http://gmt.genome.wustl.edu/somatic-sniper/current/

https://github.com/genome/somatic-sniper

[55]

FreeBayes

2012

Bayesian haplotype-based SNPs, INDELs, and MNPs

Overcomes the problem of sequences aligning to multiple genomic locations.

https://github.com/freebayes/freebayes/

[56]

Lofreq

2012

Poisson-binomial model-based Somatic SNVs

More sensitive than ad hoc or model-based methods, achieving higher sensitivity without loss of specificity. Validated to detect rare variants down to 0.05% frequency.

http://sourceforge.net/projects/lofreq/

https://csb5.github.io/lofreq/

[57]

EBCall

2013

Empirical

Bayesian framework

Somatic SNVs, INDELs

Effectively detects low-frequency somatic mutations (<10%).

https://github.com/friend1ws/EBCall

[58]

Shimmer

2013

Probabilistic, hypothesis-driven statistical method Somatic SNVs

Highly effective on heterogeneous or contaminated tumor samples.

http://www.github.com/nhansen/Shimmer

[59]

Seurat

2013

Bayesian analysis with beta-binomial distributions Somatic SNVs, small INDELs, LOH, and SVs

Extends to RNA-seq data for detecting allelic imbalance events in annotated transcripts.

https://sites.google.com/site/seuratsomatic

https://github.com/tgen/seurat

[60]

Virmid

2013

Bayesian inference with the estimated joint genotype

probability matrix

Somatic SNVs

Outperforms other tools in detecting somatic mutations in highly contaminated samples.

http://sourceforge.net/projects/virmid

[61]

RADIA

2014

Heuristic and integrative algorithm Somatic SNVs from DNA and matched RNA

Integration of RNA and DNA. High accuracy. Simulation-based validation.

https://github.com/aradenbaugh/radia/

[62]

Platypus

2014

mapping + assembly + haplotype-based Bayesian SNP, INDELs, and complex polymorphisms

High sensitivity and specificity for SNPs.

http://www.well.ox.ac.uk/platypus

[63]

SMUFIN

2014

Quaternary sequence tree-based, direct read comparison algorithm Somatic SNVs, INDELs, and SVs

Defines complex chromosomal rearrangements (chromoplexy, chromothripsis) at base-pair resolution.

http://cg.bsc.es/smufin/download

https://github.com/smufin/smufin-core

[64]

Abra

2014

Assembly-based realigner Somatic INDELs

Employs a fast and adaptable localized de novo assembly, followed by global realignment, to improve the accuracy of read mapping.

https://github.com/mozack/abra

[65]

cgpPindel

2015

Split-read mapping algorithm Somatic INDELs

Optimized for somatic indels. Scalable execution. Uses post-hoc filtering.

https://github.com/cancerit/cgpPindel

[66]

CaVEMan

2016

Expectation-Maximization (EM) Somatic SNVs

Produces high-quality somatic substitution calls with high recall and positive predictive value. Uses post-hoc filtering.

https://github.com/cancerit/CaVEMan/releases

[67]

MuSE

2016

Markov substitution model Somatic SNVs

Builds a sample-specific error model and sets tiered cutoffs to balance sensitivity and specificity.

http://bioinformatics.mdanderson.org/main/MuSE

https://github.com/danielfan/MuSE

[68]

VarDict

2016

Local realignment and consensus building Somatic SNV, MNV, INDELs, germline variants, LOH, complex variants, and SVs

Ultra-deep sequencing scalability. Supports amplicon-aware variant calling and can detect PCR artifacts.

https://github.com/AstraZeneca-NGS/VarDict

[69]

Scalpel

2016

Microassembly using self-tuning de Bruijn graphs Somatic INDELs

Highly accurate in repetitive regions, supporting single-sample, de novo, and somatic analysis modes.

http://scalpel.sourceforge.net/

[70]

Strelka2

2018

Bayesian model for continuous allele frequencies Somatic SNVs and small INDELs, germline variants

Estimation of insertion/deletion error parameters for each sample using a mixture model. Uses an enhanced somatic variant model that corrects for normal-sample contamination to improve detection in liquid tumor data.

https://github.com/Illumina/strelka

[71]

SvABA

2018

Assembly-based Somatic INDELs and SVs

Shows broad sensitivity for indels and structural variants (SVs), especially 20–300 bp variants and complex rearrangements (viral integrations). Combines assembly and alignment signals. Low computational burden and is suitable for large-scale data.

https://github.com/walaj/svaba

[72]

Lancet

2018

Local-assembly-based using colored de Bruijn graphs to jointly analyze tumor and normal reads Somatic SNVs, INDELs

High precision and robust quality scoring for prioritizing somatic variants.

https://github.com/nygenome/lancet

[73]

TNScope

2018

Haplotype-based Somatic SNVs, INDELs, SVs

Hybrid approach combining haplotype-based variant detection with ML filtration.

https://www.sentieon.com/

[74]

MuTect 2

2019

Local assembly of haplotypes and pair-HMM Read-to-Haplotype realignment Somatic SNVs and INDELs

Employs probabilistic models for genotyping and filtering that perform effectively across all sequencing depths and can operate with or without a matched normal sample.

https://github.com/broadinstitute/gatk/releases/download/4.2.2.0/gatk-4.2.2.0.zip

[75]

NeuSomatic

2019

CNN

First DL-based

somatic SNV detection

Tumor/normal alignment information summarized into matrices capturing genomic context. Learns deep feature representations directly from raw data.

https://github.com/bioinform/neusomatic

[76]

DNN-Boost

2021

Ensemble of Deep Neural Network (DNN) and XGBoost Somatic variants

Somatic mutation identification of tumor-only whole-exome sequencing data. Distinguishes somatic mutations from germline variants. Shows consistent accuracy and F1-score improvements on external datasets.

https://github.com/firdaaminy/DNN-Boost

[77]

VarNet

2022

CNN Somatic SNVs, INDELs

Outperforms existing tools and even the ensemble method. Learns effectively from weakly labeled datasets.

https://github.com/skandlab/VarNet

[78]

Dragen

2023

Haplotype-based Somatic SNVs, INDELs

Flexible architecture. Built-in noise models. Joint tumor–normal analysis. FPGA acceleration.

https://basespace.illumina.com/

[79]

DeepSom

2023

CNN Somatic SNVs, INDELs

Tumor-only somatic variant caller.

Combines Mutect2 candidate generation, gnomAD-based filtering, and DL classification.

https://github.com/heiniglab/DeepSom

[80]

DeepSomatic

2024

CNN Somatic SNVs, INDELs

Outperforms existing tools across Illumina and long-read platforms. Adaptable models for tumor-only, WES, and FFPE data.

https://github.com/google/deepsomatic

[81]

OncoTOP

2024

Combines realDcaller2 and Mutect2 Somatic SNVs, INDELs

Tumor-only somatic variant caller. Combines realDcaller2 and Mutect2. Predicts their germline or somatic origins, and evaluates clinically relevant biomarkers.

NA

[82]

VarNet-T

2025

CNN Somatic SNVs, INDELs

Novel method for tumor-only somatic variant calling. Offers improved accuracy for clinical applications such as assessing tumor mutational burden (TMB) for immunotherapy and DNA mismatch repair (MMR) deficiency for PARP inhibitor therapies.

NA

[83]

ClairS-TO

2025

Ensemble of two (Affirmative and Negational) neural networks Somatic SNVs, INDELs

Long-read tumor-only somatic variant caller. Pre- and post-filtering enhancements. Addresses limitations of real tumor data.

https://github.com/HKU-BAL/ClairS-TO

[84]

SNV: Single Nucleotide Variant. Indel: Insertion and Deletion. MNP: Multi-Nucleotide Polymorphism. HMM: Hidden Markov Model. SV: Structural Variant

Community reference tumor–normal DNA pairs and curated call sets from SEQC2 enable standardized evaluation of WGS/WES somatic callers across platforms and sites, with raw data and truth sets openly available [85]. In parallel, the ICGC–TCGA DREAM Somatic Mutation Calling Challenge provides simulated tumor genomes and crowd-sourced leaderboards to compare callers under controlled conditions [86]. Analysis of The Cancer Genome Atlas (TCGA) database has facilitated the characterization of 933,954 expressed neoantigens across 20 distinct solid tumor types. These neoantigens are derived from 893,960 somatic mutations, exhibiting a variable median frequency across different cancer types. A limited subset of neoantigens, specifically 24, including those originating from mutations within key driver genes such as PIK3CA, RAS, and BRAF, is observed to be shared among a minimum of 5% of patients across various cancer types or within the same cancer [47, 87]. The distribution of mutation types and sample sizes across these cancer types is summarized in Fig. 2.

Fig. 2.

Fig. 2

Distribution of mutation types across different cancer types from the Cancer Genome Atlas (TCGA). The radial bar plot displays the number of distinct mutations identified per cancer type (indicated on the periphery, with sample sizes in parentheses). Bars are segmented by mutation type, including missense, frameshift, in-frame deletion, in-frame insertion, and splice variants, as indicated in the legend. SKCM, skin cutaneous melanoma; LUAD, lung adenocarcinoma; LUSC, lung squamous cell carcinoma; BLCA, bladder urothelial carcinoma; UCEC, uterine corpus endometrial carcinoma; COAD, colon adenocarcinoma; STAD, stomach adenocarcinoma; DLBC, lymphoid neoplasm diffuse large B cell lymphoma; HNSC, head and neck squamous cell carcinoma; PAAD, pancreatic adenocarcinoma; CESC, cervical and endocervical cancers; ACC, adrenocortical carcinoma; READ, rectum adenocarcinoma; LIHC, liver hepatocellular carcinoma; SARC, sarcoma; UCS, uterine carcinosarcoma; TGCT, testicular germ cell tumor; CHOL, cholangiocarcinoma; BRCA, breast invasive carcinoma; GBM, glioblastoma multiforme; KICH, kidney chromophobe; THYM, thymoma; UVM, uveal melanoma; LGG, brain lower-grade glioma; PRAD, prostate adenocarcinoma; KIRP, kidney renal papillary cell carcinoma; THCA, thyroid carcinoma; PCPG, pheochromocytoma and paraganglioma; KIRC, kidney renal clear-cell carcinoma; MESO, mesothelioma

Recommendations: A wide range of somatic variant calling tools is available, and their selection should be guided by study design, data characteristics, and analytical goals. For tumor–normal paired sequencing, established methods such as MuTect2 (most widely used) and Strelka2 (time-efficient) are generally preferred due to their improved performance [88], whereas tumor-only analyses benefit from specialized approaches such as DeepSom, VarNet-T, and ClairS-TO that incorporate advanced filtering or learning strategies to distinguish somatic from germline variants. The choice of tool should also depend on the mutation type of interest: for somatic SNV detection, commonly used methods include LoFreq, MuSE, SomaticSniper, and Lancet, whereas INDEL detection is often better supported by tools such as VarScan 2 and Pindel. In addition, MuTect2 and Strelka provide robust performance across both SNVs and INDELs [89]. Whereas tools like SvABA and SMUFIN are better suited for structural variants and complex rearrangements. For detecting low-frequency or subclonal mutations, highly sensitive methods such as LoFreq and EBCall are advantageous, whereas tools like VarDict perform well on ultra-deep or amplicon sequencing data. In samples with significant tumor heterogeneity or contamination, probabilistic approaches such as Virmid and Shimmer may offer improved robustness. When multi-omics data are available, integrative tools such as RADIA and Seurat can enhance variant confidence by incorporating RNA evidence. More broadly, conventional statistical and Bayesian frameworks, including GATK and FreeBayes, remain widely used due to their interpretability and robustness, while DL-based methods such as NeuSomatic, VarNet, and DeepSomatic have demonstrated improved performance in complex or noisy datasets, albeit with higher computational demands. Finally, no single tool is universally optimal; therefore, combining multiple callers or adopting ensemble strategies is often recommended to improve accuracy and reduce false-positive rates.

HLA typing

HLA typing is foundational to neoantigen identification because peptide–MHC binding must be predicted in each patient’s autologous HLA context, enabling recognition of tumor-specific “non-self” peptides that drive effective immunotherapy responses [1, 90]. Standard pipelines integrate tumor/normal exome (±RNA-seq) with high-resolution HLA genotypes to nominate mutant peptides predicted to bind class I/II molecules, an approach validated by MS-guided discovery and vaccination readouts in preclinical models [91]. These typed, patient-specific predictions underpin personalized vaccines; in melanoma, a multi-peptide vaccine targeting up to 20 predicted neoantigens was feasible, safe and immunogenic, inducing polyfunctional CD4+/CD8+ responses with durable disease control in several patients [4, 90]. HLA-aware analyses of tumor evolution also refine target selection; allele-specific HLA loss of heterozygosity occurs in ~40% of NSCLC and biases binding toward the lost allele, so jointly assessing HLA typing and HLA LOH helps avoid non-presented targets and improves prioritization [92]. Collectively, accurate HLA genotyping, paired with sequencing, binding prediction and immunopeptidomics, enables precise neoantigen nomination, guides vaccine and T cell therapy design, and supports biomarker discovery linking neoantigen burden and clinical response [1, 90].

The HLA region is hyperpolymorphic; for example, the IPD-IMGT/HLA database now catalogs > 35,000 curated alleles and serves as the official WHO repository with regular quarterly releases [93]. This database is the authoritative, highly curated repository for HLA allele sequences and nomenclature and underpins benchmark studies [91]. Short-read HLA typing tools are typically assessed against this reference and gold-standard methods; for example, xHLA achieves 99–100% four-digit accuracy across class I and II loci on WGS data within minutes, and graph-based HLA*LA attains high accuracy on exome and low-coverage WGS [94, 95].

Inferring HLA alleles from standard NGS is challenging because extensive polymorphism and inter-gene homology complicate read mapping [96]. Alignment-based tools achieve high accuracy on short reads; OptiType reports ~97% overall accuracy from unenriched WGS/WES/RNA data [97]. Graph-aware methods model allelic diversity explicitly, for example HLA*PRG attains accuracies comparable to sequence-based typing (SBT) on high-quality WGS [96], and HLA*LA extends graph alignment to both short- and long-read data, reporting ~99% on WGS, ~93% on exome, and ~98% on long-read datasets [95]. A broader graph framework, HISAT-genotype, integrates > 14 M variants and includes an HLA module that matches or exceeds laboratory assays on benchmarks [98].

Data type influences performance and reporting resolution. On exomes, accuracy is typically lower and coverage less uniform than on WGS; prior work recommends high fragment lengths and ≥30× depth for robust calls [96]. Consistent with this, HLA*LA reports ~93% mean accuracy on exome versus ~99% on WGS and supports long-read inputs (~98%), facilitating full-length allele inference when needed [95]. For RNA-seq, ArcasHLA attains 100% two-field accuracy for class I and >99.7% for class II on benchmark sets, although post-transcriptional effects and amplification bias can complicate inference [99]. Assembly-based approaches such as Kourami enable discovery of novel alleles from high-coverage WGS but show reduced performance at lower coverage and on WES [100]. In very large cohorts with standardized exome capture, direct exome-based calling at multi-field resolution (e.g., HLA-HD in UK Biobank WES) scales effectively and increases power for association studies, though some loci (e.g., DQA1) exhibit lower coverage [101]. Representative HLA typing tools, their input data types, and performance characteristics are listed in Table 2.

Table 2.

HLA typing tools

Model/Year Architecture Input Data Performance/MHC Class/Key Features/Code Accessibility Ref.

HLAminer

(2012)

Alignment-based and targeted assembly NGS reads (whole genome, whole transcriptome/RNA-Seq, exome)

Direct-alignment mode reached ~80% sens. / 78% spec.

Class I

Does not require specialized HLA enrichment, compatible with various sequencing data.

https://www.bcgsc.ca/resources/software/hlaminer

https://github.com/bcgsc/HLAminer

[102]

seq2HLA

2012

Alignment-based read-mapping and statistical inference RNA-seq

NA

Class I & II

Determines locus-specific expression level. Provides a confidence (p-value) for each HLA call.

http://tron-mainz.de/tron-facilities/computational-medicine/seq2HLA/

https://github.com/TRON-Bioinformatics/seq2HLA

[103]

SOAP-HLA

2013

Alignment-based haplotype assembly Aligned BAMs, WGS

NA

Class I & II

Integrated one-step pipeline for variant detection and HLA typing. High coverage, low bias. Long DNA fragments (~500 bp).

http://soap.genomics.org.cn/SOAP-HLA.html

https://github.com/adefelicibus/soap-hla

[104]

ATHLATES

2013

Assembly-based WES

Overall concordance rate of 99%

Class I & II

Assembly-based, not alignment-based, reconstructs exon sequences directly from reads; automated allele pair inference.

https://github.com/cliu32/athlates

[105]

HLAforest

2013

Alignment-based hierarchical read weighting RNA-seq

NA

Class I & II

Suitable for longer-read data. Supports paired-end reads to maximize phasing information.

https://code.google.com/p/hlaforest/

https://github.com/FNaveed786/hlaforest

[106]

OptiType

2014

Integer Linear Programming (ILP)-based optimization

RNA-Seq, Exome,

WGS

Accuracy: 97.1% (4-digit)

Class I

Uses exons 2 and 3 plus flanking introns as references, reconstructing missing introns using phylogenetic methods, fast, accurate.

https://github.com/FRED-2/OptiType

[97]

POLYSOLVER

2015

Alignment-guided, model-based probabilistic WES

Accuracy: 97%

Class I

Model-based design; somatic mutation detection; custom reference construction.

http://www.broadinstitute.org/cancer/cga/polysolver

https://github.com/jason-weirather/hla-polysolver

[107]

HLAreporter

2015

De novo assembly-based approach WES

NA

Class I & II

Uses a comprehensive reference panel; zero mismatch tolerance for both assembly and allele matching ensures high reliability.

http://paed.hku.hk/genome/

[108]

HLA-VBSeq

2015

Alignment-based and Variational Bayesian (VB) inference optimization WGS

Accuracy: 99.95%

Class I & II

No need for primer design for HLA loci; does not rely on allele frequency data or prior population information; achieves full (8-digit) HLA typing resolution.

http://nagasakilab.csml.org/hla

https://github.com/oicr-gsi/hlavbseq

[109]

HLA-HD

2017

Dictionary-based mapping + weighted read-count scoring NGS data

High-coverage WES data: 100%; Low-coverage: 91.0%

Class I & II

Unrestricted use of all exons for typing; considers variation inside and outside G-DOMAIN; supports up to 6-digit precision.

https://www.genome.med.kyoto-u.ac.jp/HLA-HD/

https://github.com/TRON-Bioinformatics/tronflow-hla-hd

[110]

xHLA

2017

Translated-read strategy with exhaustive MSA expansion and iterative allele-set refinement 30× WGS BAM

99–100% four-digit for both class I and II

Class I & II

Minute-scale runtime on a desktop.

https://github.com/humanlongevity/HLA

[94]

HLAScan

2017

Alignment-based mapping to IMGT/HLA with a read-distribution score function WGS, WES, targeted sequencing

96.9% overall

Class I & II

Read-distribution scoring to cut false positives; unique-read phasing; supports up to six-digit typing; exome and target-panel compatible.

http://www.genomekorea.com/display/tools/HLA_SCAN

[111]

HLAProfiler

2017

k-mer profiling (Kraken-based gene filtering; competitive pairwise scoring) RNA-seq FASTQ

>99% accuracy at two-field (biological and simulated)

Class I & II

Excels on rare/novel/partial alleles via allele-refinement (e.g., 68% of novel alleles with correct protein or exact CDS); can work with as few as ~1,000 filtered HLA reads; extendable to KIR genes.

https://github.com/ExpressionAnalysis/HLAProfiler

[112]

PHLAT

2018

Alignment-based, probabilistic, and pairwise optimization RNA-seq, WES, amplicon FASTQ

NA

Class I & II

Supports a broad range of read lengths, coverage, and targeted amplicon sequencing data; can output supporting BAM format.

https://sites.google.com/site/phlatfortype

https://github.com/chrisamiller/docker-phlat

[113]

Kourami

2018

Graph-guided assembly using modified partial-order graphs (POGs) High-coverage WGS

>98% accuracy for known alleles

Classical class I & II

Novel allele discovery, fast, moderate memory.

https://github.com/Kingsford-Group/kourami

[100]

HLA*LA (HLA*PRG successor)

2019

Graph-based: linear alignments projected onto a population reference graph (PRG) WGS, WES, ONT/PacBio long reads, assemblies

99% (WGS); 93% (WES); 98% for long-read WGS/targeted

Classical class I & II

Broad input support; graph model derived from PRG lineage.

https://github.com/DiltheyLab/HLA-LA

[95]

HISAT-genotype

2019

Graph FM-index (HISAT2) + genotype genome; guided k-mer assembly WGS (short reads)

NA

Class I & II

Novel-allele discovery via guided k-mer assembly; can assemble full-length alleles (exons+introns) from typical WGS; fast and memory-efficient alignments thanks to graph. Exactly matches known alleles for six HLA genes in 17 Platinum Genomes; “matches or exceeds” lab assays.

https://github.com/DaehwanKimLab/hisat2

[98]

ArcasHLA

2020

Alignment-based RNA-seq HLA typing that uses Kallisto pseudoalignment plus EM quantification RNA-seq (paired-end; also, single-end tested)

100% (Class I, two-field) and >99.7% (Class II) on 1000 G benchmark

Class I & II

Partial-allele typing option; population-specific priors; works with single-end or paired-end RNA-seq; integrates easily in BAM-based pipelines; fast; works on metatranscriptomes.

https://github.com/RabadanLab/arcasHLA

[99]

WES: Whole Exome Sequencing. WGS: Whole Genome Sequencing

Platform choice should reflect study aims. When full-gene phasing is essential, long-read sequencing can type complete alleles, but in one recent comparison of classical methods, PacBio long reads offered little protein-coding advantage over MiSeq, whereas MiSeq provided superior scalability and cost-effectiveness [114]. For tumor–normal cohorts, researchers additionally assessed allele-specific HLA loss of heterozygosity using LOHHLA, a copy-number tool that has revealed HLA LOH in ~40% of NSCLC and refines neoantigen prediction [92].

Recommendations: Tool selection depends on availability of data type (WES or RNA), dataset size, sequencing depth, and computational capacity. For instance, RNA-seq–based tools such as seq2HLA, HLAforest, HLAProfiler, and ArcasHLA are well suited for studies that also require expression quantification, whereas WES or WGS data may be better paired with high-accuracy tools like OptiType, HLA-HD, xHLA, or HLA-VBSeq. If novel allele discovery or high-resolution typing is a priority, graph-based or assembly-driven methods such as Kourami, HLA*LA, or HISAT-genotype offer clear advantages. Users working with limited computational resources may prefer faster, lightweight approaches like xHLA or ArcasHLA, while those prioritizing maximum accuracy and completeness should consider tools with demonstrated high concordance rates and multi-digit resolution. If resources allow, use OptiType and POLYSOLVER for MHC-I and HLA-HD for MHC-II, or combine multiple tools for best results [115].

From peptide processing to MHC-peptide-TCR complex

The identification of neoantigens, tumor-specific peptides capable of eliciting immune responses, is pivotal for advancing cancer immunotherapy. The process leading from peptide generation to T cell recognition and the formation of the MHC–peptide–TCR complex is highly regulated and occurs through several sequential steps: proteolytic cleavage of proteins into peptides, peptide transport into the endoplasmic reticulum via the Transporter Associated with Antigen Processing (TAP) in the MHC class I pathway, binding of peptides to MHC molecules, presentation of these complexes on the cell surface, and ultimately, recognition by T cells. Each step is critical for determining which peptides are presented to T cells, and computational tools, particularly those leveraging AI, have significantly enhanced prediction accuracy.

The experimental data required for modeling each of these steps have been collected by researchers through various laboratory assays and are available in publicly accessible databases. For proteasomal cleavage prediction, datasets are derived from in vitro digestion experiments and MS analyses that identify cleavage sites within proteins, providing insight into how peptides are generated inside cells. TAP transport prediction relies on peptide translocation assays, often measuring BA or transport efficiency of peptides across the endoplasmic reticulum membrane via the TAP complex. Peptide–MHC binding and presentation prediction requires high-quality BA data (IC50 values), eluted ligand (EL) datasets identified through MS, and T cell activation assays that confirm immunogenicity. These diverse experimental data types are systematically collected and annotated in specialized databases such as IEDB [116], SYFPEITHI [117], TSNAdb [39], TumorAgDB1.0 [118], TANTIGEN 2.0 [119], MHCBN 4.0 [120], AntiJen [121], and MHCPEP [122], which integrate biochemical binding data, naturally processed ligands, and immunological assay results. Together, these datasets enable the development and benchmarking of computational models that aim to accurately predict antigen processing and immune recognition. The databases, along with the types of data available in each, are presented in Table 3.

Table 3.

Databases for neoantigen prediction and immunogenicity analysis

Task Database Data Type Web Accessibility Ref.
Somatic mutation calling The Cancer Genome Atlas (TCGA) WGS, WES, and gene expression data https://portal.gdc.cancer.gov/ [123]
HLA typing IMGT HLA sequences https://www.imgt.org/ [124]
IPD-HLA MHC sequences of different species https://www.ebi.ac.uk/ipd/mhc/ [125]
Antigen processing and presentation (Proteasomal cleavage, TAP transport, peptide–MHC binding, immunogenicity) Immune Epitope Database (IEDB) Experimentally validated epitopes and immune responses https://www.iedb.org/ [116]
SYFPEITHI MHC class I and II ligands, peptide motifs http://www.syfpeithi.de/ [117]
TSNAdb

Predicted and experimentally validated

tumor-specific neoantigen

https://pgx.zju.edu.cn/tsnadb1/ [39]
TumorAgDB1.0 Curated tumor antigens https://tumoragdb.com.cn/#/home [118]
TANTIGEN 2.0 T cell epitopes and HLA ligands http://projects.met-hilab.org/tadb [119]
MHCBN 4.0 Peptides interacting with TAP and MHC https://webs.iiitd.edu.in/raghava/mhcbn/index.html [120]
AntiJen Quantitative peptide binding data (TAP, MHC, TCR–MHC) https://www.ddg-pharmfac.net/antijen/AntiJen/antijenhomepage.htm [121]
MHCPEP MHC class I and II ligands http://wehih.wehi.edu.au/mhcpep [122]
Immunogenicity and TCR recognition Immune Epitope Database (IEDB) Antibody and T cell epitope data https://www.iedb.org/ [116]
McPAS-TCR Disease-associated TCR sequences https://friedmanlab.weizmann.ac.il/McPAS-TCR/ [126]
VDJdb TCR sequences with known antigen specificities https://vdjdb.cdr3.net/ [127]
10x Genomics Single-cell immune profiling (TCR sequences, gene expression) https://www.10xgenomics.com/datasets [128]
TCRdb TCR sequences http://bioinfo.life.hust.edu.cn/TCRdb [129]
PIRD TCR and immunoglobulin sequences across species https://db.cngb.org/pird/ [130]
ImmuneCODE™ TCR sequences linked to antigen specificity (infectious diseases) https://clients.adaptivebiotech.com/pub/covid-2020 [131]
TBAdb TCRs targeting specific antigens/diseases (subset of PIRD) https://gitlab.com/immunomind/immunarch/raw/dev-0.5.0/private/TBAdb.xlsx [NA]

This section reviews the molecular mechanisms underlying each step and highlights AI-based tools (Fig. 3) that facilitate neoantigen identification, with applications in personalized immunotherapy.

Fig. 3.

Fig. 3

Developed computational tools for antigen processing, MHC binding, and T cell recognition prediction for MHC class I and class II. The figure illustrates the antigen processing and presentation pathways for both MHC class I (left) and MHC class II (right) molecules, highlighting their interaction with CD8+ and CD4+ T cells, respectively. Various computational tools are categorized according to their functional roles, including cleavage prediction, TAP transport prediction, peptide–MHC BA prediction, peptide–MHC presentation prediction, immunogenicity (IM) prediction, and TCR–pMHC binding/IM prediction. Tools associated with the MHC class I pathway include those for proteasomal cleavage, TAP transport, peptide–MHC binding, and IM prediction, whereas tools for the MHC class II pathway emphasize endosomal processing, peptide binding, and IM prediction. The green boxes represent prediction tools designed for individual steps of the pathway, whereas the red boxes indicate tools that either integrate multiple steps within a single pathway (MHC-I or MHC-II) or predict a step across both pathways (MHC-I and MHC-II)

Proteasomal Cleavage. The proteasome, a multi-subunit protease complex, degrades intracellular proteins into peptides, generating potential ligands for MHC class I molecules. This process is highly specific, with cleavage sites determined by the proteasome’s catalytic subunits, which differ between constitutive proteasomes (expressed in all cells) and immunoproteasomes (induced in antigen-presenting cells under inflammatory conditions). The immunoproteasome, upregulated by interferon-γ, favors peptides with hydrophobic or basic C-termini, suitable for MHC class I binding. Furthermore, the immunoproteasome exhibits limited specificity, implying that not every potential mutated peptide will be generated during protein degradation [132]. Additionally, not all peptides produced by the proteasome will reach the necessary cellular compartments to potentially interact with HLA proteins [36]. Accurate prediction of cleavage sites is essential for identifying peptides that progress in the MHC class I pathway. Algorithms developed in this field are typically trained using either in vitro proteasome digestion data or in vivo MHC-I and -II ligand elution data [133].

Several computational tools, many employing AI, predict proteasomal cleavage sites by modeling sequence preferences. NetChop uses artificial neural networks (ANNs) trained on in vitro and in vivo cleavage data, achieving approximately 70% accuracy for C-terminal cleavage sites [134]. Pepsickle, a more recent tool, employs a gradient-boosted classifier, offering superior performance (AUC of 0.821 for constitutive and 0.789 for immunoproteasome) and the ability to differentiate between proteasome types [135]. PCPS utilizes an n-gram-based ML approach, outperforming NetChop in sensitivity (0.89 vs. 0.79) for immunoproteasome cleavage [136]. PCleavage applies support vector machines (SVMs), achieving a Matthews correlation coefficient of 0.54 for constitutive proteasome data [137]. PAProC uses an evolutionary algorithm, suitable for human and yeast proteasomes [138], while MAPPP relies on a kinetic model, though it is less flexible and not AI-based [139]. A random forest-based model by Li et al. (2012) provides additional ML-driven predictions, primarily for constitutive proteasomes [140]. For predicting cleavage sites of proteins in the MHC class II presentation pathway, tools such as PepCleaveCD4 [141] and MHCII-NP [142] show promise, but they require further improvement to accurately predict IM. Key features and reported performance of these proteasomal cleavage prediction tools are summarized in Table 4.

Table 4.

Cleavage site prediction tools

Model/Year Architecture Supported Protease Type Performance MHC Class/Key Features/Web and Code Accessibility Ref.

PAProC

2001

Evolutionary algorithm Human, Yeast Accuracy = 0.82

I

Evolves algorithms for diverse proteasome types.

http://www.paproc.de/

[138]

MAPPP

2003

Kinetic model Constitutive NA

I

Modeling kinetic processes for cleavage prediction.

NA

[139]

NetChop-3.1

2005

ANN Constitutive, Immunoproteasome AUC = 0.80

I

Accurate prediction based on neural networks.

https://services.healthtech.dtu.dk/services/NetChop-3.1/

[134]

ProteaSMM

2005

Matrix-based Constitutive, Immunoproteasome AUC = 0.76 for i20S

I

Integrated proteasomal cleavage, TAP transport, and MHC I binding predictions.

http://www.mhc-pathway.net

[143]

PCleavage

2005

SVM Constitutive, Immunoproteasome AUC = 0.79

I

Support vector machines for robust predictions.

http://www.imtech.res.in/raghava/pcleavage/

[137]

Li et al.

2012

RF Constitutive (likely) Accuracy = 0.85

I

Random forests for versatile prediction models.

NA

[140]

PepCleaveCD4

2013

SVM

Endosomal

proteases: cathepsins L and S

AUC = 0.85

II

Predictor highlights the effect of secondary structure on CD4+ T cell epitope preprocessing.

http://peptibase.cs.biu.ac.il/PepCleave_cd4/

[141]

MHCII-NP

2018

Enrichment and depletion-based score Lysosomal proteases AUC = 0.767

II

Incorporates cleavage and binding motifs into prediction of MHC-II ligands.

http://tools.iedb.org/mhciinp

[142]

PCPS

2020

N-gram-based Constitutive, Immunoproteasome

Sensitivity = 0.88

Specificity = 0.57

I

Uses n-gram models to predict cleavage sites.

http://imed.med.ucm.es/Tools/pcps/

[136]

Pepsickle

2021

Gradient-boosted classifier Constitutive, Immunoproteasome AUC = 0.87

I

Utilizes ML for improved accuracy.

https://github.com/pdxgx/pepsickle

[135]

Recommendations: When selecting a cleavage site prediction tool from those listed in Table 4, users should first consider the type of protease and MHC class relevant to their study, as most tools are specialized either for MHC class I (proteasomal cleavage) or class II (endosomal/lysosomal processing). For MHC class I applications, widely used and reliable options such as NetChop-3.1, PCleavage, and Pepsickle offer solid performance, with newer ML-based methods like Pepsickle generally providing improved accuracy and flexibility. If users require integrated antigen-processing predictions, tools like ProteaSMM are advantageous because they combine cleavage, TAP transport, and MHC binding steps. For MHC class II studies, more specialized tools such as PepCleaveCD4 and MHCII-NP should be preferred, as they are designed to capture endosomal protease activity and peptide presentation pathways.

TAP Transport. TAP moves peptides that are 8–16 amino acids long from the cytosol into the ER, where they become available for loading onto MHC class I molecules. TAP preferentially transports peptides with specific sequence motifs, influencing the epitope repertoire. Predicting TAP binding affinity is crucial for identifying peptides likely to proceed to MHC loading.

AI-based tools enhance TAP binding predictions. TAPPred uses a cascade SVM approach, achieving a correlation coefficient of 0.88 in jack-knife validation [144]. PREDTAP combines ANNs and hidden Markov models (HMMs), offering robust predictions with an AUC greater than 0.85 [145]. TAPREG, also SVM-based, predicts binding affinities for peptides of variable lengths (8–16 residues), with a maximum correlation of 0.89 [146]. The IEDB TAP prediction tool employs a matrix-based approach, providing a non-AI alternative [147]. SVMTAP applies SVMs with a focus on peptide sequence features, achieving high specificity [148]. Combining TAP transport with MHC-I binding improves epitope identification, and cascade SVM models report a strong correlation with measured transport [144, 147]. DeepTAP, a recent DL tool, uses bidirectional gated recurrent units (BiGRUs) to capture sequential dependencies, outperforming traditional ML methods [149]. Representative TAP-binding prediction methods and their main characteristics are listed in Table 5.

Table 5.

TAP binding affinity prediction tools

Model/Year Architecture MHC Class Performance Key Features/Code Accessibility Ref.

IEDB TAP

2003

Matrix-based I AUC = 0.932 on HLA-A0201

Simple scoring matrix derived from experimental binding data.

http://tools.iedb.org/processing/

[147]

TAPPred

2004

Cascade SVM I Correlation r = 0.88

Combines sequence plus 33 amino-acid physicochemical features; high accuracy via multi-level SVM cascade.

http://www.imtech.res.in/raghava/tappred/

http://bioinformatics.uams.edu/mirror/tappred/

[144]

PREDTAP

2006

ANN + HMM I AUC > 0.85

Integrates neural network and hidden Markov modeling for TAP affinity.

http://antigen.i2r.a-star.edu.sg/predTAP

[145]

TAPREG

2010

SVM I Pearson r = 0.89

SVM-based TAP binding prediction with residue-based encoding.

http://imed.med.ucm.es/Tools/tapreg

[146]

TAP Hunter

2010

SVM I AUC = 0.85

Using feature vectors from

the N- and C-terminal positions of TAP ligands.

http://datam.i2r.a-star.edu.sg/taphunter/

[150]

DeepTAP

2023

BiGRU RNN I

Spearman r = 0.91

Pearson r = 0.89

DL captures sequential context; improved precision for top-ranked neoantigens.

https://pgx.zju.edu.cn/deeptap

https://github.com/zjupgx/deeptap

[149]

CLTAP

2025

LLM with contrastive learning I NA

Integrates Contrastive Learning (CL) with a contextual co-attention mechanism.

https://github.com/ChenWeiCCZU/CLTAP

[151]

Recommendations: For straightforward analyses or rapid screening, classical tools such as IEDB TAP and TAPPred provide reliable performance with simple implementations, making them suitable for users with limited computational resources or those integrating TAP prediction into larger antigen-processing pipelines. For more advanced applications, especially in neoantigen discovery or precision immunotherapy, newer ML- and DL-based approaches such as DeepTAP offer improved predictive power by capturing sequence context more effectively, while emerging models like CLTAP may further enhance performance through advanced architectures like contrastive learning and attention mechanisms. Overall, combining TAP prediction with upstream (cleavage) and downstream (MHC binding) analyses, and optionally cross-validating results with multiple tools, can significantly improve the reliability of antigen presentation predictions.

Peptide–MHC binding/presentation. The binding of peptides to MHC class I or II molecules determines which peptides are presented to T cells. MHC class I binds short peptides (8–10 amino acids) in a closed groove, while MHC class II accommodates longer peptides (13–25 amino acids) in an open-ended groove. High BA is a primary determinant of immunogenicity, as only stable binders are likely to be presented effectively. AI-driven tools, particularly those using DL, have revolutionized BA predictions by modeling complex peptide-MHC interactions across diverse alleles.

Accurate prediction of peptide-MHC BA is essential for neoantigen identification in cancer immunotherapy. Numerous computational methods have been created, leveraging ML and DL to model these interactions. Early tools like NetMHC used ANNs to predict binding affinities based on sequence data [152]. The introduction of pan-specific models, such as NetMHCpan, marked a significant advancement by enabling predictions for alleles not included in the training set, thus broadening applicability to diverse populations [153]. The advent of DL has further enhanced prediction accuracy. Convolutional neural networks (CNNs), as employed in tools like ConvMHC and MHCflurry, capture local patterns in peptide sequences [154, 155], while recurrent neural networks (RNNs) and long short-term memory (LSTM) networks, featured in MHCnuggets and USMPep, handle sequential dependencies effectively [156, 157]. More sophisticated architectures, including attention mechanisms and transformer models (e.g., DeepHLApan, ImmunoBERT), focus on critical residues and contextual information within sequences [158, 159]. Beyond BA, some tools predict peptide presentation or immunogenicity directly. For instance, MARIA incorporates gene expression and cleavage predictions to estimate the likelihood of MHC class II presentation, providing a more comprehensive assessment for neoantigen identification [160].

Recent tools such as BigMHC, DeepNeo, HLAthena, PRIME 2.0, NetMHCpan-4.1, and Class I Immunogenicity (IEDB) further advance the field by integrating diverse data types and advanced architectures to predict BA, EL presentation, and IM. These tools differ in their methodologies, training data, and predictive capabilities, offering varied strengths for specific applications [11, 161165]. In Table 6, we provide a detailed comparison of these tools, including their architectures, training data, prediction types, and web accessibility, along with short explanations highlighting their main methodologies and features, and differences from other models.

Table 6.

Peptide-MHC binding affinity and presentation prediction tools

Model/Year/Ref. Architecture Training Data/Output MHC Class /Performance Key Features/Web and Code Accessibility

Parker et al.,

1994

[166]

Additive motif/matrix model

BA

BA

Class I

NA

First fully quantitative per-position coefficient table for HLA-A2; formalizes the independent binding of side-chains hypothesis and shows it holds for most peptides (only 3/83 require side-chain interactions).

NA

SYFPEITHI

1999

[117]

Motif-based, hand-crafted position-specific scoring matrix per allele

200 peptide motifs and 2,000 peptide sequences

Motif score and rank

Class I

NA

Early motif-based system built on experimentally eluted ligands and known T cell epitopes only (no synthetic binders). Encodes anchor and auxiliary anchor residues, preferred and unfavorable AAs via integer weights.

https://www.syfpeithi.de/

ProPred

2001

[167]

Classical matrix / PSSM-based predictor

BA

BA

Class II

NA

One of the earliest widely-used web servers for DR-restricted epitope prediction; covers 51 DR alleles; simple threshold-based interface.

http://crdd.osdd.net/raghava/propred/

CTLPred

2004

[168]

Quantitative Matrix, ANN, SVM

BA

BA

Class I

Accuracy ≈ 70%

Independent of exact HLA restriction → can be used when HLA is unknown; Implemented as a practical web server, widely used in early vaccine-design pipelines.

http://crdd.osdd.net/raghava/ctlpred/

RANKPEP

2004

[169]

PSSM-based scorer

BA

BA

Class I & II

AUC > 0.8 for Class I; AUC > 0.7 for Class II

Unified PSSM framework for MHCI and MHCII; supports MSA-based variability masking so only conserved regions are scanned; integrates proteasomal cleavage prediction; provides allele-specific PSSM binding thresholds (PSBT).

http://imed.med.ucm.es/Tools/rankpep.html

ARB

2005

[170]

Average Relative Binding (ARB) coefficient matrices

BA

BA

Class I & II

AUC ≈ 0.80

Provides direct IC₅₀ prediction rather than only scores, enabling multi-allele / multi-length global ranking; covers multiple species and peptide sizes (8–11 for class I, ≥9 for class II) with systematic cross-validation.

http://www.epitope.liai.org:8080/matrix

IEDB SMM

2005

[171]

SMM

BA

BA

Class I

NA

Fully quantitative inputs and outputs; built-in handling of experimental noise; robust treatment of bounded measurements.

http://www.mhc-pathway.net/smm

MHCPred 2.0

2006

[172]

QSAR-based regression models (additive method)

BA

BA

Class I & II

NA

Provides quantitative affinity predictions (not just classifier scores); supports both human and mouse alleles; allows specifying anchor positions; built on explicit QSAR models that can help interpret residue contributions. Frequently used as a QSAR-style comparator in MHC-binding studies.

http://www.jenner.ac.uk/MHCPred/

SMM-Align

2007

[173]

Stabilization Matrix Alignment

BA

BA

Class II

Mean AUC ~0.73–0.76

Hybrid alignment + motif-learning method.

https://services.healthtech.dtu.dk/services/NetMHCII-2.3/

SMMPMBEC

2009

[174]

SMM

BA

BA

Class I

Average AUC for SMMP M B E c vs SMM is 0.860 vs 0.836

Penalizes opposite charge substitutions strongly (due to adverse effects on peptide–MHC binding), unlike BLOSUM62.

http://www.mhc-pathway.net/smmpmbec

NN-align

2009

[175]

ANN

BA

BA

Class II

AUC = 0.810

Simultaneously infers binding core and affinity; explicitly encodes peptide-flanking residues and peptide length; corrects redundancy bias by scaling backprop step size by the size of the core-redundancy cluster (Hobohm-1 clustering); uses ensemble averaging over multiple architectures and random initializations.

http://www.cbs.dtu.dk/services/NetMHCII-2.0

PickPocket 1.1

2009

[176]

PSMM

BA

BA

Class I

AUC = 0.895

Works well with limited training data, unlike artificial neural networks. Especially effective for non-human species.

https://services.healthtech.dtu.dk/services/PickPocket-1.1/

MultiRTA

2010

[177]

Regularized Thermodynamic Average

BA

BA

Class II

AUC = 0.783

Extends the earlier single-allele RTA model to 430 HLA-DR/DP allotypes; demonstrates strong leave-one-allele-out performance and good generalization to overlapping-peptide epitope data sets.

http://bordnerlab.org/MultiRTA

NetCTLpan

2010

[178]

Integrated pathway model (ANN)

SYFPEITHI

NetCTLpan score = pathway presentation likelihood

Class I

AUC ≈ 0.976–0.982

Pan-specific CTL epitope prediction for any HLA class I molecule with known sequence; integrates antigen processing (cleavage + TAP) with binding; Supports 8, 9, 10, and 11-mer epitopes.

https://www.cbs.dtu.dk/services/NetCTLpan/

MULTIPRED2

2011

[179]

ANN

BA

BA

Class I & II

NA

Large-scale screening across alleles, supertypes and full genotypes; maps promiscuous T cell epitope hotspots in whole proteomes; exploits strong pan-specific predictors without re-training; convenient for vaccine design / neoantigen prioritization.

http://cvc.dfci.harvard.edu/multipred2/

PAComplex

2011

[180]

Structural modeling

BA, TCR–pMHC structure, complexes and experimental data

pMHC binding model, pTCR binding model

Class I

NA

Detects homologous peptide antigens.

http://PAcomplex.life.nctu.edu.tw

TEPITOPEpan

2012

[181]

Pan-specific PSSM

BA

BA

Class II

Best average AUC ≈0.739

Competitive AUC with NetMHCIIpan on many alleles; robust performance even when training data are restricted; best among compared methods in predicting exact binding cores for 20 peptide–MHC complexes.

http://www.biokdd.fudan.edu.cn/Service/TEPITOPEpan/

IEDB-AR-Consensus

2012

[182]

Consensus-based

BA

BA

Class I

AUC = 0.96

Large quantitative peptide–MHC binding datasets (IC50) across HLA-A/B, primate and mouse alleles.

http://tools.iedb.org/mhci/

NetMHCcons 1.1

2012

[183]

Consensus

BA

BA

Class I

NA

Dynamically combines NetMHCpan, NetMHC, and PickPocket depending on allele characterization and similarity to known alleles.

https://services.healthtech.dtu.dk/services/NetMHCcons-1.1/

OWA-PSSM

2013

[184]

PSSM

BA

BA

Class II

AUC = 0.747

Extends TEPITOPE’s pocket-profile framework to 879 HLA-DR alleles using OWA weights over pocket pseudo-sequence similarities; does not require large allele-specific training sets.

NA

NetMHCstabpan 1.0

2016

[185]

ANN

pMHC-I half-life measurements data

BS

Class I

NA

Shows that combining stability and affinity predictions significantly improves T cell epitope prediction.

https://services.healthtech.dtu.dk/services/NetMHCstabpan-1.0/

NetMHC-4.0

2016

[152]

Gapped sequence alignment using ANN

BA, EL

BA, EL

Class I

AUC > 0.9

Integrates BA and EL data into a unified model for improved prediction of peptide presentation and binding. Accurately captures peptide length preferences. Performs well on both neoantigen identification and natural ligand datasets.

https://services.healthtech.dtu.dk/services/NetMHC-4.0/

PSSMHCpan

2017

[186]

Position-Specific Scoring Matrix (PSSM)-based model

BA

BA

Class I

AUC = 0.94

Combines allele-specific and pan-specific PSSMs using HLA sequence similarity; supports broad peptide lengths (8–25) and very large HLA coverage (4,896 alleles). Designed for high-throughput neoantigen discovery.

https://github.com/BGI2016/PSSMHCpan

NNAlign − 2.1

2017

[187]

ANN

BA, EL

BA, EL

Class I & II

AUC = 0.81

Works with any biological sequence type (protein, DNA, RNA); generates gapped sequence alignments; aligns sequences of variable lengths to identify motifs.

https://services.healthtech.dtu.dk/services/NNAlign-2.1/

MSIntrinsic

2017

[188]

ANN

EL

EL

Class I

AUC = 0.99

Discovery of sequence motifs; improved quantification of the roles of gene expression and proteasomal processing.

NA

HLA-CNN

2017

[189]

CNN

BA

BA

Class I

AUC = 66.7%

Treating peptides as sentences and AAs as words. HLA-Vec captures physicochemical properties; demonstrated use on the entire human proteome to highlight highly self-binding 9-mers.

https://github.com/uci-cbcl/HLA-bind

ConvMHC

2017

[154]

DCNN

BA

BA

Class I

AUC > 0.9

Leverages convolutional layers to detect positional motifs in peptide sequences. One of the earliest deep learning architectures for peptide–MHC binding prediction.

https://github.com/aidanbio/convmhc

DeepMHC

2017

[190]

DCNN

BA

BA

Class I

NA

DL for improved feature extraction; robust for MHC-I binding.

http://mleg.cse.sc.edu/deepMHC/

AI-MHC

2018

[191]

DCNN

BA

BA

Class I & II

NA

Supports both MHC classes; DL enhances binding prediction.

https://baras.pathology.jhu.edu/AI-MHC/index.html

EDGE

2019

[192]

DCNN

EL

EL

Class I

NA

Trained on tumor MS data; high accuracy for neoantigen identification.

http://massive.ucsd.edu/

MARIA

2019

[160]

RNN

EL

EL

Class II

AUC = 0.89

Diverse training data; integrates gene expression and cleavage data.

https://maria.stanford.edu/

CNN-NF

2019

[193]

DCNN

BA

BA

Class I

NA

Focuses on local sequence features; improves BA prediction.

https://github.com/zty2009/MHC-I

DeepLigand

2019

[194]

Deep language model (ELMo) with residual network (DL)

BA

BA

Class I

NA

Language model approach; robust for MHC-I binding prediction.

https://github.com/gifford-lab/DeepLigand

PUFFIN

2019

[195]

Deep residual network

BA, EL

BA

Class I & II

NA

Uncertainty quantification; supports both MHC classes.

https://github.com/gifford-lab/PUFFIN

NeonMHC2

2019

[196]

Ensemble of CNNs

EL

EL

Class II

NA

Ensemble approach; high accuracy for MHC-II presentation.

https://neonmhc2.org/

DeepSeqPan

2019

[197]

DCNN

BA, EL

BA, EL

Class I

NA

DL for pan-specific MHC-I binding; high generalization.

https://github.com/pcpLiu/DeepSeqPan

ACME

2019

[198]

Attention-based CNNs

BA

BA

Class I

NA

Attention mechanism improves focus on key residues.

https://github.com/HYsxe/ACME

ForestMHC

2019

[199]

RF

EL

EL

Class I

AUC = 0.73

ForestMHC scores correlate monotonically (not linearly) with IC50 values, suggesting peptide presentation isn’t solely dependent on chemical affinity.

https://github.com/kmboehm/ForestMHC

MHCherryPan

2020

[200]

LSTM, CNN

BA, EL

BA, EL

Class I

NA

Pan-specific; combines LSTM and CNN for robust predictions.

NA

MHCnuggets

2020

[201]

LSTM networks and GRUs

BA

BA

Class I & II

AUC (human) = 0.64

Handles sequential dependencies; supports both MHC classes.

https://github.com/KarchinLab/mhcnuggets

USMPep

2020

[157]

AWD LSTM with embedding layer

BA

BA

Class I & II

NA

Universal sequence model; effective for diverse alleles.

https://github.com/nstrodt/USMPep

IConMHC

2020

[202]

Deep CNN

BA

BA

Class I

AUC on common vs rare alleles: 0.834 vs 0.789

Encodes pairwise amino-acid interaction properties rather than one-hot/BLOSUM; handles non-9mers by gapping/trimming to multiple 9-mer views and takes max-affinity; pan-allele, predicts rare alleles with no training data.

NA

MHCAttnNet

2020

[203]

Bi-LSTM encoder with an attention mechanism

BA

BA

Class I and II

AUC-PRC = 94.18%

Attention yields interpretable heatmaps over amino acids and reduces candidate MHC trigrams from ~9,251 to ~258, focusing wet-lab validation. Handles variable-length peptides via Bi-LSTM.

https://github.com/gopuvenkat/MHCAttnNet

MHCflurry 2.0

2020

[204]

ANN

BA, EL

BA, EL

Class I

AUC = 0.877–0.985

Improved over MHCflurry; integrates MS data for better predictions.

https://github.com/openvax/mhcflurry

NetMHCpan 4.1

2020

[164]

an ensemble of 100 single-layer neural networks

BA, EL

BA, EL

Class I

NA

Improved over NetMHCpan-4.0; better for HLA-B and HLA-C; pan-specific.

https://services.healthtech.dtu.dk/services/NetMHCpan-4.1/

HLAthena

2020

[11]

Single-layer neural network

IEDB

EL

Class I

HLAthena (MSiCEB model): up to 78% recall

Integrates gene expression and peptide cleavage information.

http://hlathena.tools/

BERTMHC

2021

[205]

BERT-based with multiple instance learning

BA, EL

BA, EL

Class II

NA

BERT-based; high performance for MHC-II binding.

https://bertmhc.privacy.nlehd.de/

https://github.com/s6juncheng/BERTMHC

DeepAttentionPan

2021

[206]

DL with attention mechanism

BA, EL

BA, EL

Class I

NA

Attention-based; improved pan-specific predictions.

https://github.com/jjin49/DeepAttentionPan

DeepNetBim

2021

[207]

DL with network analysis

BA, IM

BA, IM

Class I

NA

Integrates immunogenicity; network analysis enhances predictions.

https://github.com/Li-Lab-SJTU/DeepNetBim

SHERPA

2021

[208]

Gradient boosting decision trees

EL

EL

Class I

NA

High performance for eluted ligands; robust for neoantigens.

NA

MATHLA

2021

[209]

Bidirectional LSTM with attention

BA, EL

BA, EL

Class I

NA

Bidirectional LSTM; focuses on key sequence features.

https://github.com/MATHLAtools/

ImmunoBERT

2021

[159]

BERT-based

BA, EL

BA, EL

Class I

NA

BERT-based; high generalization for MHC-I binding.

https://github.com/hcgasser/ImmunoBERT

Anthem

2021

[210]

NB, XGBoost, LR, NN, SVM, DT, RF

BA

BA

Class I

AUC = from 0.888 to 0.981

Combines scoring function-based methods (PWM/PSSM) with ML.

http://anthem.erc.monash.edu/

https://github.com/17shutao/Anthem

NMER

2021

[211]

RF

EL

EL

Class I

AUC ≈ 0.91

Integrates MHCflurry1.6 binding/processing/presentation scores, NetMHCpan-2.8 binding, NetMHCstabpan-1.0 stability, proteasomal C-terminal cleavage, TAP transport scores, etc.

https://github.com/JaredJGartner/SB_neoantigen_Models

RBM-MHC

2021

[212]

Restricted Boltzmann Machine

(RBM)

EL

EL

Class I

AUC = 0.991

Score presentability of (neo)antigens without HLA information. Visualize HLA-binding motifs and binding modes for data exploration and feature discovery. Classify peptides by HLA restriction with minimal annotation.

https://github.com/bravib/rbm-mhc

TransPHLA

2022

[213]

Simplified Transformer with masked multi-head self-attention

BA

BA

Class I

AUC = 0.926

TransPHLA for binding + AOMP module that uses attention scores to propose higher-affinity peptide mutations for a target HLA; provides allele/length-specific attention heatmaps; pan-allele support for A/B/C and 8–14mers.

https://github.com/a96123155/TransPHLA-AOMP

DeepSeqPanII

2022

[214]

RNN with attention

BA, EL

BA, EL

Class II

NA

Attention-based; tailored for MHC-II binding prediction.

https://github.com/pcpLiu/DeepSeqPanII

MHCRoBERTa

2022

[215]

Transfer learning with BERT

BA, EL

BA, EL

Class I

NA

Transfer learning enhances performance; robust for MHC-I.

https://github.com/FuxuWang/MHCRoBERTa

FIONA

2022

[216]

Flexible NN architecture

BA, EL

BA, EL

Class II

NA

Flexible architecture; optimized for MHC-II binding.

http://therarna.cn/fiona.html

HLApollo

2022

[217]

Transformer model (DL)

BA, EL

BA, EL

Class II

NA

Transformer-based; high accuracy for MHC-II.

NA

HLAB

2022

[218]

ProtBert, BiLSTM, UMAP, LR/SVM/XGBoost

BA

BA

Class I

AUC = 0.9891

Cascaded protein LM (ProtBert) + BiLSTM embeddings; UMAP dimensionality reduction; filter and wrapper feature selection (T-test, W-test, RF, LR-RFE, SVM-RFE); evaluates seven classifiers and picks the best per task.

https://www.healthinformaticslab.org/supp/resources.php

DeepMHCII

2022

[219]

CNN

BA

BA

Class II

AUC = 0.77

Identifying the binding core and pinpointing key binding pockets.

https://github.com/yourh/DeepMHCII

IEPAPI

2023

[220]

Transformer-based feature extraction

EL, IM

EL, IM

Class I

NA

Two-task design (EL→IM) mirrors biology; attention visualizes HLA-restricted motifs; uses HLA pseudo-sequence (34 positions); EL-pretraining then freezes for IM; supports variable peptide lengths (8–11).

https://github.com/ddd9898/IEPAPI

MixMHC2pred-2.0

2023

[221]

Deep motif deconvolution with NNs

EL

EL

Class II

AUC (human) = 0.94

Motif deconvolution; high specificity for MHC-II ligands.

http://mixmhc2pred.gfellerlab.org/

https://github.com/GfellerLab/MixMHC2pred

CapsNet-MHC

2023

[222]

Capsule neural networks

BA

BA

Class I

AUC: Higher than baselines (specific values in Fig. 7a)

Captures pairwise features; high interpretability via dynamic routing.

https://github.com/s7776d/CapsNet-MHC

DeepMHCI

2023

[223]

Anchor position-aware DL

BA, EL

BA, EL

Class I

NA

Anchor-focused; improves prediction for MHC-I binding.

https://github.com/ZhuLab-Fudan/DeepMHCI

TLimmuno2

2023

[224]

LSTM, Transfer learning

BA

BA, IM

Class II

NA

Transfer learning; high performance for MHC-II binding.

https://github.com/XSLiuLab/TLimmuno2

NetMHCIIpan-4.3

2023

[225]

ANN

BA, EL

BA, EL

Class II

NA

Enhances prediction accuracy for all HLA class II molecules by integrating locus-specific immunopeptidomics data.

https://services.healthtech.dtu.dk/services/NetMHCIIpan-4.3/

TransMHCII

2023

[226]

PLM embedding and an image

classifier

BA

BA

Class II

AUC = 0.908

Represents multi-class classification model for MHC-II binding prediction.

https://github.com/xinyu-dev/TransMHCII

MHC2AffyPred

2023

[227]

Random forest regression on structural fingerprints

BA

BA

Class II

NA

Structure-based ML approach; shows consistently higher correlations than NetMHCIIpan-3.2 and MHCII3D on shared benchmarks and is demonstrated on SARS-CoV-2 epitope affinity prediction.

https://github.com/SiddhiJani/MHC2AffyPred

STMHCpan

2023

[228]

Star-Transformer

EL

EL

Class I

AUC = 0.9499

Replaced the fully connected structure with star topology using a lightweight Star-Transformer, reducing model complexity while enhancing predictive performance. Incorporated an attention mechanism to enhance the architecture and performance of a DL model for prediction.

https://github.com/Luckysoutheast/STMHCPan.git

MHCSeqNet2

2024

[156]

DCNN, GRU and LSTM

BA, EL

BA, EL

Class I

AUC = 0.99 (on MS data); AUC = 0.54 (on T cell epitopes)

Uses sub-word-level peptide features and a 3D structure embedding; generalizes to unseen alleles; outperforms NetMHCpan on some datasets.

https://github.com/cmb-chula/MHCSeqNet2

Graph-pMHC

2024

[229]

GNN (utilizes

Alphafold2-multimer-derived graph adjacency matrices)

EL

EL

Class II

NA

Better evaluation via GO-based splitting, assess antibody immunogenicity risk.

https://github.com/Genentech/gpmhc

RPEMHC

2024

[230]

CNN

BA

BA

Class I & II

AUC (MHCI) = 0.866, AUC (MHCII) = 0.759

Uses residue–residue pair encoding; captures critical interaction information between the molecules.

https://github.com/lennylv/RPEMHC

TripHLApan

2024

[231]

Triple coding matrix, BiGRU, Attention model, Transfer learning

BA

BA

Class I & II

AUC = 0.958

Optimizes allele sequence extraction by focusing on interaction sites and correlating HLA binding motifs with coding strategies.

https://github.com/CSUBioGroup/TripHLApan.git

HLAPepBinder

2024

[232]

RF

BA

BA

Class I

Accuracy = 0.90–0.91

Consensus pipeline using nine predictors (ANN, Consensus, NetMHCpan BA, NetMHCpan

EL, SMM, SMMPMBEC, PickPocket, NetMHCcons,

NetMHCstabpan), high generalizability to unseen HLA subtypes, addressing the negative data problem.

https://github.com/CBRC-lab/HLAPepBinder

MUNIS

2025

[233]

LSTM, ESM-2

EL

EL

Class I

AUC = 0.980

Predicts immunodominance hierarchies in HIV and EBV more effectively.

10.5281/zenodo.14219509

OnmiMHC

2025

[234]

1D-CNN-LSTM and 2D-CNN

BA, EL

BA

Class I & II

MHC-I PR-AUC = 0.854

MHC-II PR-AUC = 0.606

Multimodal feature fusion; combination of 2D + 1D convolutional kernels; iterative data preprocessing.

https://github.com/caihaihua057200/OnmiMHC

pMHChat

2025

[235]

LLMs and Deep Hypergraph Learning

BA, PDB

BA

Class II

AUC = 0.7676

BA prediction through illuminating the residue contact profiling of the binding surface.

https://github.com/jianiM/pMHChat

MixMHCpred3.0

2025

[236]

Neural network

EL

EL

Class I

LOA AUC ≥ 0.98

Predicts ligands for MHC-I alleles without known ligands, as well as MHC-I alleles across species.

https://github.com/GfellerLab/MixMHCpred

PHLA-SiNet

2025

[237]

Siamese neural network

BA

BA

Class I

NA

Higher sensitivity without substantial degradation of other metrics.

https://github.com/CBRC-lab/PHLA_SiNet

Current presentation predictors are trained on EL datasets in addition to affinity data, for instance, NetMHCpan-4.0 integrates both to capture length preferences and increase accuracy, and MHCflurry-2.0 couples binding and antigen-processing models into an integrated presentation score [153, 204]. Large mono-allelic EL compendia (e.g., HLAthena: >185,000 peptides across 95 alleles) drive improved PPV and enable validation in tumor cell lines [11]. Tissue-resolved, benign-organ immunopeptidomes (HLA Ligand Atlas) contextualize tumor ligandomes and offer broad background sets for benchmarking [238].

Recommendations: Based on extensive benchmarking studies, several models have been identified as top-performing tools for peptide–MHC prediction. For MHC class I, NetMHCpan-4.1 and MixMHCpred, and for MHC class II, NetMHCIIpan-4.1 and MixMHC2pred, are widely regarded as state-of-the-art and industry standards for predicting pMHC binding affinity and presentation [239, 240]. MHCflurry 2.0 demonstrates comparable performance. These tools leverage artificial neural networks and integrate both BA and EL data, enabling them to predict not only binding but also natural antigen processing and presentation. They also provide broad HLA allele coverage, including rare alleles. Additionally, integrative tools such as NetCTLpan, which combine multiple biological processes including peptide binding, antigen processing, and presentation, can deliver more biologically realistic predictions, particularly for applications like neoantigen discovery and vaccine design. Finally, combining predictions from multiple tools is often a robust strategy to mitigate model-specific biases.

Immunogenicity and TCR recognition. Recognition of a presented pMHC occurs when T cell receptors identify the pMHC as foreign (non-self). This recognition activates T cells, leading to their expansion and the targeted elimination of cancer cells displaying the detected pMHCs. Predicting pMHC recognition is a key step in neoantigen prediction, since not all presented pMHCs provoke an immune response.

Approaches to predicting pMHC recognition can be divided into two categories: those that account for specific TCR information and those that do not. In TCR-focused strategies, the goal is to model the binding relationship between individual TCRs and pMHCs. These methods frequently employ protein sequence representations combined with ML architectures such as convolutional or recurrent neural networks [241243]. Training typically relies on curated repositories including the Immune Epitope Database (IEDB) [116], McPAS-TCR [126], and VDJdb [127] that provide experimentally confirmed records of TCR–pMHC interactions. TESLA assembled a multi-center IM dataset and model that removes the vast majority of non-immunogenic peptides at high precision, providing a community benchmark for recognition-aware pipelines [12]. Public epitope resources anchor comparative evaluations across antigen processing and presentation. IEDB functions as a central repository that aggregates experimentally measured epitopes and exposes them through a public, searchable interface, supporting systematic reuse in computational studies [116]. Its scope spans antibody, T cell, and MHC binding contexts, allowing modelers to align tasks to immunological readouts and to stratify benchmarks by assay type, organism, or antigen class within a single harmonized framework [116]. The TESLA consortium coordinated a multi-group, end-to-end evaluation in which participants predicted immunogenic epitopes from shared tumor sequencing data, after which a large candidate set was tested for T cell binding; the resulting model excluded the overwhelming majority of non-immunogenic peptides at high precision and reproduced performance in an independent cohort, defining rigorous criteria for pipeline comparison [12]. On the receptor side, VDJdb has expanded the number of TCR sequences with known cognate antigens, introduced compact motif sets suitable for training specificity predictors, and enabled batch annotation of repertoire sequencing data, supporting generalization tests on realistic repertoires [127]. Complementarily, McPAS-TCR offers a manually curated catalogue of pathology-linked TCR sequences numbering over five thousand entries, furnishing disease-labeled sets for benchmarking recognition models that aim to recover pathology-specific repertoires or antigen associations [126].

Most ML models for TCR–pMHC prediction are based on amino acid sequence encoding. Input sequences are numerically encoded into vector embeddings, using methods such as one-hot encoding, BLOSUM (evolutionary similarity) [244], AA-index [245], or Atchley factors [246] (physicochemical properties). They use the antigenic peptide, the MHC, and the TCR sequence as inputs. The MHC-peptide-TCR binding predictors, including ERGO-II, ImRex, DLpTCR, TITAN, NetTCR-2.2, and pMTnet, differ in their architectural choices and training datasets [247252]. BERT-based models, such as STAPLER, TABR-BERT, and EPIC-TRACE, treat protein sequences as a biological language, generating contextual embeddings of TCRs, peptides, and MHCs to predict recognition [253255]. Extending this idea, ESM-based methods leverage large protein language models trained on massive sequence datasets (ESM-1b, ESM-2) to capture both evolutionary and structural features; examples like TCR-ESM often benefit from transfer learning to improve generalization [256]. In contrast, AlphaFold-driven approaches emphasize structure, predicting 3D TCR–pMHC docking complexes and refining predictions with scoring functions or graph neural networks, as seen in tools like TCRmodel2 and NetTCR-struc [257, 258]. Together, these strategies highlight complementary strengths; BERT and ESM models capture sequence-level features, whereas AlphaFold-based methods provide structural insight into binding interfaces. Nonetheless, all approaches remain challenged by limited paired TCR–pMHC data and the vast diversity of TCR clonotypes. Table 7 provides a comparative overview of the available tools.

Table 7.

TCR-pMHC binding and IM prediction tools

Model/Year/Ref. Architecture Training Data Source/Output MHC Class/Performance Key Features/Web and Code Accessibility

POPISK

2011

[259]

SVM and weighted degree string kernel

MHCPEP, SYFPEITHI, IEDB

IM

Class I

AUC = 0.74

Only sequence information used in the model.

http://iclab.life.nctu.edu.tw/POPISK.

PAAQD

2013

[260]

RF

MHCPEP, SYFPEITHI, IEDB

IM

Class I

AUC = 0.72

Uses quantum-derived physicochemical features describing molecular similarity.

http://pirun.ku.ac.th/%7Efsciiok/PAAQD.rar

Class I Immunogenicity (IEDB)

2013

[165]

Combination of the enrichment scores and position weights

IEDB, publications

IM

Class I

NA

It excludes anchor positions P2/P9 to avoid binding biases.

http://tools.iedb.org/immunogenicity/

EpiToolkit

2015

[261]

Multiple immunoinformatics tools

EpiToolKit itself does not train new models

Ranked epitope sets

Class I & II

Full vaccine-design workbench covering all steps from allele selection and epitope conservation to prediction.

http://www.epitoolkit.de

TepiTool

2016

[262]

Multiple ML and matrix-based MHC binding predictors

IEDB-curated quantitative peptide–MHC BA and EL datasets

IC50 and percentile rank

Class I & II

TepiTool itself is an interface; underlying IEDB MHC binding predictors typically achieve ROC AUC > 0.9 in benchmark evaluations

Step-by-step 6-step wizard; exposes IEDB “recommended” settings by default; supports epitope scanning, allele selection, peptide length/overlap control, and automatic selection of top-ranked epitope candidates across MHC I and II.

https://tools.iedb.org/tepitool/

GLIPH

2017

[263]

Similarity-based clustering of similar TCRs

Publications, PDB

Motifs, CDR3 clusters, antigen-specific TCR groups

Class I

NA

Finds TCRs with shared antigen specificity by combining motif discovery, sequence similarity, structural contact probability, and immunological features such as V-gene bias and clonal expansion.

https://github.com/immunoengineer/gliph

TCRdist

2017

[264]

Distance-based (BLOSUM62 matrix) clustering of similar TCRs

In-house dataset

TCRdist matrices, clusters, motifs, epitope-specific TCR assignment

Class I

NA

Generalizes Simpson’s diversity index by including similarity between receptors, not only identity.

Classifies new TCRs to epitopes based on density of nearby receptors in TCRdist space.

https://github.com/phbradley/tcr-dist

ITCell

2018

[265]

Template-based

Multiplex Substrate Profiling by Mass Spectrometry (MSP-MS), IEDB, PDB

Antigen cleavage sites, BA, IM

Class II

NA

Integrative structure-based approach that incorporates antigen cleavage by proteases, MHCII presentation, and TCR recognition.

http://salilab.org/itcel

https://github.com/salilab/itcell

Structure

2019

[266]

ANN

PDB, publications, IEDB

IM

Class I

AUC = 0.60

Structure-based IM prediction.

NA

Tcell_predictor

2019

[267]

RF

Publication, TANTIGEN

IM

Class I

AUC = 0.65–0.73

Protein expression level and charge differences introduced by the mutation significantly improves the ability to identify immunogenic neo-epitopes.

http://peptibase.cs.biu.ac.il/Tcell_predictor/

TCRex

2019

[268]

RF

McPAS-TCR, VDJdb, ImmuneCODE

TP binding

Class I & II

NA

A stringent BPR threshold is applied to reduce false positives.

Enables annotation of entire TCR repertoires with predicted epitope targets.

https://tcrex.biodatamining.be/

DeepHLApan

2019

[158]

RNN (BiGRU) with an attention layer

IEDB

BA, IM

Class I

BA AUC > 0.9

Good performance on clinical data.

http://biopharm.zju.edu.cn/deephlapan

https://github.com/jiujiezz/deephlapan

Smith et al.

2019

[269]

Gradient boosting model

210 class I and 68 class II predicted neoantigens with IFNγ ELISpot readouts (in-house dataset)

IM

Class I

NA

Single model predicts neoantigen/mHA immunogenicity directly from peptide-intrinsic features; extensive feature engineering.

https://github.com/vincentlaboratories/neoag

TTAgP 1.0

2019

[270]

RF

IEDB, TANTIGEN

IM

Class I

NA

TTAgP uses a comprehensive set of peptide features.

https://github.com/bio-coding/TTAgP

POTN

2020

[271]

SVM

IEDB, peptide database, SYFPEITHI

IM

Class I

NA

Physicochemical properties (position-specific, especially at P3).

https://www.frontiersin.org/articles/10.3389/fimmu.2020.02193/full#supplementary-material

INeo-Epp

2020

[272]

RF

IEDB, publications

IM

Class I

AUC = 0.81

Eliminate false positive predicted peptides.

http://www.biostatistics.online/ineo-epp/neoantigen.php

iTTCA-Hybrid

2020

[273]

SVM, RF

IEDB, TANTIGEN

IM

Class I

AUC = 0.783

Hybrid feature representation.

Balanced training through SMOTE.

http://camt.pythonanywhere.com/iTTCA-Hybrid

DeepImmuno-CNN and -GAN

2021

[274]

CNN for IM prediction + GAN for sequence generation

IEDB

IM

Class I

ROC AUC ≈ 0.85 and PR AUC ≈ 0.81

Integrates beta-binomial modeling of immunogenic potential with an HLA-contextual CNN; explicitly models peptide–MHC pairs; interpretable via residue occlusion analysis to highlight key antigen positions.

https://github.com/frankligy/DeepImmuno

TCRGP

2021

[275]

Gaussian processes (GP)

VDJdb, publication

Classification probability (indicating whether a given TCR recognizes a given epitope)

NA

AUC = 0.863

Leveraging CDR sequence information from both TCRα and TCRβ chains and automatically learning which CDRs are most important for each epitope.

https://github.com/emmijokinen/TCRGP

pMTnet

2021

[252]

LSTM and autoencoder combined through dense network

McPAS-TCR, VDJdb, PIRD, NetMHCpan, publications

TpM binding

Class I

AUC > 0.8

Overcomes limitations of clustering methods like GLIPH and TCRdist.

https://dbai.biohpc.swmed.edu/pmtnet/index.php

https://github.com/tianshilu/pMTnet

ImRex

2021

[248]

CNN

VDJdb

TpM binding

Class I & II

NA

Define a feature representation using pairwise amino acid interaction maps instead of separate embeddings.

https://github.com/pmoris/ImRex

TITAN

2021

[250]

Attention-based NNs pretrained on BindingDB

VDJdb, publications

TpM binding

NA

AUC = 0.78

Reformulates the task as compound–protein interaction prediction by representing epitopes as SMILES strings, enabling data augmentation and pretraining on large datasets like BindingDB.

https://github.com/PaccMann/TITAN

ERGO-II

2021

[247]

LSTM, AE, MLP

McPAS-TCR, VDJdb

TpM binding

Class I & II

NA

Uses large-scale TCR-peptide dictionaries.

http://tcr2.cs.biu.ac.il/

https://github.com/IdoSpringer/ERGO-II

DLpTCR

2021

[249]

Ensemble DL framework from FCN, CNN and ResNet

TetTCR-seq, VDJdb

Likelihood of TpM binding

NA

AUC = 0.8564 to 0.9227

Shows high accuracy on independent datasets, even when using only a single TCR chain.

http://jianglab.org.cn/DLpTCR/

https://github.com/jiangBiolab/DLpTCR

TCRAI

2021

[276]

ANN using ICON (Integrative COntext-specific Normalization)

VDJdb, McPAS-TCR, 10x Genomics

TpM binding

Class I & II

AUC > 0.90

ICON solves the low signal-to-noise problem in high-throughput pMHC binding. Multi omics driven background estimation.

https://github.com/regeneron-mpds/TCRAI

TCRMatch

2021

[277]

BLOSUM62 similarity matrix

IEDB

Similar TCRs, similarity scores, predicted epitopes

NA

PRAUC = 0.737

Designed specifically for TCR similarity.

Predicts epitope specificity by matching to known receptors.

http://tools.iedb.org/tcrmatch/

https://www.github.com/IEDB/TCRMatch

iTTCA-RF

2021

[278]

RF

IEDB, TANTIGEN

IM

Class I

AUC = 0.78

Hybrid feature representation.

Applied MRMD (Maximum Relevance Minimum Redundancy) to rank features.

http://lab.malab.cn/%7Eacy/iTTCA

NeoScore

2022

[279]

Logistic regression-based TESLA

Class I

AUC = 0.60–0.83 for four independent test set

Significant association of the NeoScore with survival in

response to immune checkpoint inhibition.

https://bordene.shinyapps.io/MHCI_neoantigen_prioritization/

AttnTAP

2022

[280]

Dual-input framework combining Attn-BiLSTM and Attn-MLP

VDJdb, McPAS-TCR

TP binding

NA

AUC = 0.83–0.89

Avoids overfitting by simplifying architecture. Performs well on unseen TCR–peptide predictions.

https://github.com/Bioinformatics7181/AttnTAP/

ATM-TCR

2022

[281]

Multi-head self-attention network

VDJdb, McPAS-TCR, IEDB

TP binding

Class I

Recall = 0.71

Captures biological contextual relationships shaped by sequence arrangement. Improves out-of-sample performance.

https://github.com/Lee-CBG/ATM-TCR

Seq2Neo

2022

[282]

CNN

IEDB

IM, BA

Class I

Accuracy = 0.75; Precision = 0.96

One-stop pipeline from raw FASTQ/BAM to neoepitope features; supports SNVs, INDELs, and gene fusions.

https://github.com/XSLiuLab/Seq2Neo

PRIME2.0

2023

[163]

Neural Network

Publications, IEDB

BA, IM

Class I

AUC = 0.69

High-quality training data, correct peptide length modeling, aromatic residue enrichment—especially tryptophan.

http://prime.gfellerlab.org/

epiTCR

2023

[283]

Random forest and

BLOSUM62 encoding

IEDB, TBAdb [223], VDJdb, McPAS-TCR, and 10x [224]

TpM binding

NA

AUC = 0.969

Trained on a large and diverse dataset; achieves higher sensitivity without sacrificing specificity.

https://github.com/ddiem-ri-4D/epiTCR

DePTH

2023

[284]

Combination

of a CNN and a dense layer

Publications

TCR-HLA associations

Class I & II

AUC = 0.64 to 0.69

Makes predictions for any TCR-HLA pairs then possible to study rare

HLAs.

https://github.com/Sun-lab/DePTH

EPIC-TRACE

2023

[255]

Convolution and

multi-head attention using ProtBERT embedding

VDJdb, IEDB

TpM binding

NA

AUC = 0.69

Uses the full TCR information.

https://github.com/DaniTheOrange/EPIC-TRACE

MIX-TPI

2023

[285]

CNN and self-attention fusion layer

VDJdb, ImmuneCODE, IEDB, McPAS

TpM binding

NA

NA

Multimodal framework to combine sequence-based and physicochemical features.

https://github.com/Wolverinerine/MIX-TPI

PanPep

2023

[286]

Meta learning and Neural Turing machine

IEDB, VDJdb, PIRD,

McPas-TCR

TP binding

NA

AUC = 0.67

Allows accurate prediction of TCR binding to unseen peptides.

https://github.com/bm2-lab/PanPep

POP-UP TCR

2023

[287]

RF

PDB, IMGT-numbered structure files from the Structural T Cell Receptor Database (STCRDab), McPAS TCR

TP binding

Class I

AUC = 0.5

Models trained using only TCR beta chains perform well.

https://github.com/NiliTicko/POP-UP-TCR

STAPLER

2023

[253]

BERT

IEDB,

Francis, 10 × 35, McPas, VDJdb, publications

TpM binding

NA

NA

Amino Acid masking in fine tuning.

https://github.com/schumacherlab/STAPLER

TABR-BERT

2023

[254]

BERT

TCRdb, IEDB, McPAS, VDJdb, PIRD

TpM binding

NA

AUC = 0.84

Leverages large-scale unlabeled databases and superior performance on unseen epitopes.

https://github.com/Freshwind-Bioinformatics/TABR-BERT

TCR-ESM

2023

[256]

ESM1v

Datasets from netTCR2.0, ERGO II, pMTnet

TpM binding

NA

NA

Embeddings from protein language models improved model efficiency and performance over traditional methods like CNNs or AEs.

http://tcresm.dhanjal-lab.iiitd.edu.in/

https://GitHub.com/dhanjal-lab/tcr-esm

TCR-Pred

2023

[288]

SAR using MNA with

MultiPASS

VDJdb, McPASTCR, IEDB

TpM binding

NA

NA

Atom-centered substructural MNA descriptors instead of the traditional amino acid (AA) one-letter codes.

http://way2drug.com/TCR-pred/

TCRmodel2

2023

[257]

AlphaFold v2.3.0-based DL

AlphaFold’s

database

TpM binding

NA

NA

Higher accuracy than AlphaFold for TCR–pMHC structures, runs

faster, no need for fine-tuning or templates.

https://tcrmodel.ibbr.umd.edu/

https://github.com/piercelab/tcrmodel2

iTCep

2023

[289]

Deep CNN with fusion features

McPAS-TCR, VDJdb, IEDB

TP binding

Class I

AUC = 0.86 and 0.91 on two datasets

DL framework for TCR–epitope recognition using fusion of a novel amino-acid pair propensity (AAPP) interaction map with one-hot encoding; strong generalization to unseen peptides and imbalanced data.

http://biostatistics.online/iTCep/

BERTrand-peptide

2023

[290]

Transformer (BERT) model with MLM

VDJdb, McPAS, TBAdb, 10x Genomics, publication

TP binding

Class I

AUROC ≈ 0.69

BERT model over concatenated peptide+TCR β CDR3 sequences with token, position, and type embeddings; MLM pre-training on synthetic peptide.

https://github.com/SFGLab/bertrand

NeoRanking

2023

[291]

Voting classifier (LR and XGBoost)

NCI, TESLA, HiTIDE (in-house dataset)

IM

Class I

NA

Integrates diverse biological features especially BA, stability, RNA expression, immunopeptidome evidence, and oncogenicity.

https://github.com/bassanilab/NeoRanking

DeepNeo-v2

2023

[162]

CNN

IEDB

BA, IM

Class I & II

AUC (MHC I) = 0.76

AUC (MHC II) = 0.80

Designed to be simple, fast, and user-friendly through a publicly accessible web service, enabling broad adoption.

https://deepneo.net/

https://github.com/kaistomics/DeepNeo

BigMHC

2023

[161]

An ensemble of seven pan-allelic deep neural networks (LSTM)

MHCflurry-2.0, NetMHCpan-4.1, PRIME-1.0, PRIME-2.0

EL, IM

Class I

EL AUC = 0.58

IM AUC = 0.52–0.55

BigMHC-EL (presentation model) and BigMHC-IM (immunogenicity model) created by transfer learning.

https://github.com/KarchinLab/bigmhc

VitTCR

2024

[242]

Vision Transformer (ViT)

IEDB, VDJdb, McPAS, TCRdb, 10x

TP binding

NA

NA

Integrates a positional bias

weight matrix (PBWM) with output of 3-dimensional numeric tensor named AtchleyMaps.

https://github.com/Jiang-Mengnan/VitTCR

MixTCRpred

2024

[241]

Transformer

VDJdb, IEDB, McPAS, 10x, and publications

TpM binding

Class I & II

AUC = 0.891

A quality control resource for analyzing single-cell TCR sequencing data.

https://github.com/GfellerLab/MixTCRpred

NetTCR-2.2

2024

[251]

CNN

IEDB, VDJdb, 10x

TpM binding

Class I

AUC = 0.8476

Combining pan-specific and peptide-specific models with similarity-based predictions.

https://services.healthtech.dtu.dk/services/NetTCR-2.2/

https://github.com/mnielLab/NetTCR-2.2

HeteroTCR

2024

[292]

Heterogeneous GNN

IEDB, VDJdb, McPAS-TCR

TpM binding

NA

NA

First GNN-based TpM binding predictor, superior performance compared to SotA.

https://github.com/yuzilan/HeteroTCR

CATCR

2024

[293]

Convolutional-self-attention (hybrid framework of CATCR-D: discriminative model and CATCR-G: generative model)

VDJdb, IEDB, McPAS-TCR

TpM binding

NA

AUC = 0.89

Integrates sequence and structure-based data using RCM (residue contact matrix).

https://github.com/FreudDolce/CTTCR/

ImmugenX

2024

[294]

A modular transformer-based protein language model BA, EL, stability, IM, TpM binding

Class I

AUC = 0.666

Peptide-MHC multitask pretraining, interpretable, fast stability predictions.

10.5281/zenodo.13850954

MHLAPre

2024

[295]

Transformer-based Meta-Learning (MHLAPre-IM, MHLAPre-TT

IEDB

IM, TpM binding

Class I

MHLAPre TT performance score = 0.8953

TCR-pHLA binding (MHLAPre-TT) transfer-learn from pHLA immunogenicity (MHLAPre-IM). Dynamic sampling of support/query sets to reduce bias. BLOSUM62 is employed to encode antigenic peptides and HLA alleles.

https://github.com/ChanganMakeYi/MHLAPre

Sa-TTCA

2024

[296]

SVM

IEDB, TANTIGEN 2.0

IM

Class I

NA

Includes a principled framework for statistical evaluation of selected features to improve reliability of feature selection.

https://github.com/khanhlee/sa-ttca

MHLAPre

2024

[295]

DL + Meta Learning + Transfer Learning

IEDB, VDJdb, McPAS-TCR

BA, IM

Class I

NA

Meta-learning improved performance and generalization in epitope IM prediction.

https://github.com/ChanganMakeYi/MHLAPre

NetTCR-struc

2025

[258]

Geometric Vector Perceptron Graph Neural Network (GVP-GNN)

IEDB, VDJdb, 10x, RCSB

Docking quality score and binder vs. non-binder classification

Class I

AUC₀.₁ = 0.487–0.6

Distinguishes binding vs. non-binding complexes in zero-shot settings.

Struggles when the structural models are inaccurate.

https://github.com/mnielLab/NetTCR-struc

TRAP

2025

[243]

Contrastive learning framework

VDJdb, McPASTCR, IEDB, publication

TpM binding

NA

AUC = 0.75

Captures both cross-reactivity and specificity among TCRs.

https://github.com/gejingxuan/TRAP

LightCTL

2025

[297]

Contrastive learning framework with context-aware prompt module (CAPM)

McPAS-TCR, YFV

TpM binding

Class I

AUC = from 0.7767 to 0.9445 on eight independent datasets

CAPM was developed to identify key features linked to T cell activation, antigen recognition, and specific diseases by weighing the significance of different extracted feature maps.

https://github.com/YYYYYeFei/LightCTL.git

deepAntigen

2025

[298]

Graph Convolutional Network

IEDB, publications, STCRDab, MHC Motif Atlas

BA, IM

Class I & II

AUC = 0.71

T cell antigen identification at the atomic level.

https://github.com/JiangBioLab/deepAntigen

UniPMT

2025

[299]

GNN

BigMHC, PanPep, DLpTCR, NetTCR, ERGO, pMTnet, IEDB

TpM binding, TP binding, BA

Class I

AUC = 0.7214

Handles peptide–MHC–TCR, peptide–MHC, and peptide–TCR binding prediction within a single model.

https://github.com/ethanmock/UniPMT

NeoTImmuML

2025

[300]

A weighted ensemble model of LightGBM, XGBoost, and RF

TumorAgDB1.0, latest publications (2024–2025)

IM

Class I & II

AUC = 0.885

Built on an upgraded curated database (TumorAgDB2.0). Multi-model ML evaluation. Model interpretability using SHAP.

https://tumoragdb.com.cn

TpM: TCR-peptide-MHC, TP: TCR-Peptide, Attn-BiLSTM: attention-based bi-directional LSTM

Recommendations: IM prediction tools can be broadly classified into two categories based on their input requirements. The first category includes methods that do not require TCR sequence information; tools such as PRIME2.0 and DeepNeo fall into this group and can be recommended due to their relatively improved predictive performance. In contrast, the second category comprises approaches that rely on the availability of TCR sequence data to predict IM; in this context, tools such as pMTnet, TITAN, and NetTCR-2.2 leverage deep neural networks and large-scale repertoire datasets to model TCR specificity. When analyzing large-scale TCR repertoire data or identifying antigen-specific T cell clusters, similarity- and clustering-based tools such as GLIPH and TCRdist provide interpretable and computationally efficient solutions. More recent transformer and protein language model-based approaches, including MixTCRpred and ImmugenX, offer improved generalization to unseen epitopes but may require substantial computational resources and careful validation. Overall, no single model is universally optimal; therefore, users are encouraged to select tools based on their specific application and, where possible, adopt ensemble or complementary strategies to improve robustness and reduce bias.

Integrative pipelines

Integrative neoantigen pipelines provide a reproducible, end-to-end in silico workflow that unifies somatic variant discovery, expression assessment, HLA typing, peptide processing and binding prediction, and IM or TCR-recognition analyses and, in some cases, integrates immunopeptidomic data to generate neoantigen predictions and prioritize patient-specific targets (Table 8).

Table 8.

Integrative tools for neoantigen prediction and prioritization

Model/Year Epitope Prediction Model MHC Class Input Data/Key Features/Web and Code Accessibility Ref.

INTEGRATE-neo

2016

NetMHC4.0 Class I

The pipeline consists of three steps: gene fusion peptide prediction, HLA allele prediction, and neoantigen identification. Its first step uses the human reference genome (FASTA), gene models (GenePred), and INTEGRATE-predicted gene fusions (BEDPE) to generate fusion peptide sequences

Stand-alone module for gene fusion neoantigen discovery.

https://github.com/ChrisMaherLab/INTEGRATE-Neo

[301]

FRED2

2016

NetMHC(pan)-(I/II), NetChop, NetCTL, PickPocket Class I & II

FASTA file as input

Unified Python API for epitope prediction pipelines, including polymorphic proteins; consistent I/O for many predictors;

supports epitope selection and vaccine design.

http://fred-2.github.io/

https://github.com/FRED-2/Fred2

[302]

pVAC-Seq

2016

NetMHC 3.4 Class I

WGS of tumor-normal pairs and RNA-seq as input

End-to-end genome-guided neoantigen discovery from NGS data. Integrates HLA typing, somatic variant calling, epitope prediction, binding-based filtering, and expression/coverage filters; Designed for clinical-grade workflows for personalized cancer vaccines.

https://github.com/griffithlab/pVAC-Seq

[303]

neoepitope

2017

NetMHCcons v1.1 Class I

WGS and RNA-seq

Analyzes neoepitopes arising from somatic missense mutations and gene fusions.

https://github.com/zhanglabstjude/neoepitope

[304]

MuPeXI

2017

NetMHCpan3.0 Class I

Somatic mutation calls (VCF file), HLA types, and optionally tumor gene expression

Similar functionality to pVac-Seq but with enhanced usability;

provides mutant allele frequency when the variant calls come from MuTect or MuTect2.

http://www.cbs.dtu.dk/services/MuPeXI

https://github.com/ambj/MuPeXI

[305]

TIminer

2017

NetMHCpan-3.0 Class I

RNA-seq FASTQ and somatic mutation data (VCF)

HLA typing, class I neoantigen prediction, immune infiltration characterization, and tumor immunogenicity quantification from NGS.

http://icbi.i-med.ac.at/software/timiner/timiner.shtml

[306]

CloudNeo

2017

NetMHCpan-3.0 Class I

VCF file (for mutations) and bam file (for HLA typing) as inputs

Cloud-based CWL workflow; end-to-end patient-specific neoantigen identification on the cloud; integrates HLA typing and NetMHCpan binding prediction.

https://github.com/TheJacksonLaboratory/CloudNeo

[307]

Vaxrank

2017

NetMHC, NetMHCpan, NetMHCcons, MHCflurry, web-based predictors through IEDB Class I

Tumor Mutations (VCF), Tumor RNA-Seq (BAM), Patient MHC Alleles as input

Vaxrank is under the Apache 2.0 open source license and can also be installed from the Python Package Index.

https://www.github.com/hammerlab/vaxrank

[308]

Epidisco

2017

NetMHCcons Class I

Tumor/normal RNAseq

Typed EDSL, modular workflows, cloud backends, reproducibility, bioinformatics tool catalog.

https://github.com/hammerlab/wobidisco

[309]

Neopepsee

2018

Locally weighted naive Bayes (LNB) Class I

RNA-Seq FASTQ format as input

Automates full NGS-based workflow; uses 9 key features (IC50, percentile rank, NetCTLpan scores, T cell IM score, hydrophobicity, polarity/charge, DAI, AAPP, similarity to known pathogenic epitopes).

https://sourceforge.net/projects/neopepsee

[310]

Rubinsteyn1 et al.

2018

NetMHCpan Class I

WGS/WES of tumor-normal pairs and RNA-seq as input

End-to-end workflow from FASTQ → somatic variants → RNA support → HLA typing → Vaxrank-based prioritization.

https://github.com/openvax/neoantigen-vaccine-pipeline

[311]

pTuneos

2019

RF Class I

Tumor/normal WES/WGS and RNA-seq as input

Refined IM score shown to be a pan-cancer survival marker and better predictor of ICI response than TMB / simple neoantigen load.

https://github.com/bm2-lab/pTuneos

[312]

Antigen.garnish

2019

netMHCI/II, netMHCI/IIpan, MHCnuggets and MHCflurry Class I & II

VCFs, peptide sequences, cDNA transcripts

Dissimilarity to human proteome is a strong predictor.

Supports human and mouse data; in addition to IM, it also predicts clinical outcomes.

https://github.com/immune-health/antigen.garnish

[313]

NeoPredPipe

2019

netMHCpan Class I & II

Single and multi-region variant call format (VCF) files

Multi-sample input, integration of clonal architecture and immunogenicity (recognition potential).

https://github.com/MathOnco/NeoPredPipe

[314]

ScanNeo

2019

NetMHC, NetMHCpan Class I

RNA-Seq data in BAM format aligned as input

Specifically targets indel neoantigens from RNA-seq, complementing DNA-based pipelines; uses RNA evidence to focus on expressed, frame-shifted peptides.

https://github.com/ylab-hi/ScanNeo

[315]

Neoepiscope

2020

MHCnuggets,

MHCflurry,

NetMHCpan,

NetMHCIIpan

Class I & II

Tumor/normal DNA-seq input

Key strength is explicit multi-variant phasing (germline + somatic) and support for custom references; open-source MIT-licensed command-line tool.

https://github.com/pdxgx/neoepiscope

[316]

NeoFuse

2020

MHCflurry Class I

RNA-seq FASTQ input

In fusion-caller benchmarking, Arriba + STAR-Fusion gave the best performance in validated fusions while limiting the number of called fusions. Focus on fusion neoantigens; fully containerized (docker and singularity); single-sample workflow; automatically annotates each candidate with binding and expression metrics.

https://icbi.i-med.ac.at/NeoFuse/

https://github.com/icbi-lab/NeoFuse

[317]

OpenVax

2020

NetMHCpan 4.0, NetMHCcons 1.1, SMM, SMMPMBEC Class I

Tumor/normal DNA-seq and tumor RNA-seq as input

Dockerized end-to-end pipeline.

https://github.com/openvax/neoantigen-vaccine-pipeline

[318]

pVACtools

2020

NetMHC, NetMHCpan, MHCflurry Class I & II

Somatic mutations (VCFs), RNA-seq

End-to-end neoantigen workflow including vector design (pVACvector) and interactive review (pVACview); extensive documentation, tutorials, Docker images and PyPI package. Modular framework (pVACseq, pVACbind, pVACfuse, pVACsplice, pVACvector, pVACview).

https://github.com/griffithlab/pVACtools

[319]

neoANT-HILL

2020

IEDB, MHCflurry Class I & II

Handles WES/WGS and RNA-seq as input

Dockerized pipeline with Flask GUI; supports RNA-seq–only workflows; GUI; computes DAI and immune-cell composition.

https://github.com/neoanthill/neoANT-HILL

[320]

TruNeo

2020

NetMHCpan, MHCflurry Class I

Tumor and normal WES FASTQ, RNA-seq FASTQ as input

Recall of immunogenic neoantigens among top-10 predictions = 52.63%. Explicitly models many biological steps (binding, processing, expression, tumor heterogeneity, clonality, HLA LOH).

https://github.com/yucebio/TruNeo

[321]

NeoFox

2021

NetMHCpan, MixMHCpred, PHBR-I, NetMHCIIpan, MixMHC2pred, PHBR-II Class I & II

Neoantigen candidate sequence, its

corresponding WT sequence and gene name as input

NeoFox is for annotation, not classification; easy to use python package (NEOantigen Feature toolbOX); integrates many predictors and features; designed to plug into existing pipelines.

https://github.com/TRON-Bioinformatics/neofox

[322]

TSNAD

2021

DeepHLApan1.1 Class I

WGS/WES of tumor-normal pairs and RNA-seq

Supports multiple reference genome versions for mutation calling and gene fusion analysis.

http://biopharm.zju.edu.cn/tsnad/

https://github.com/jiujiezz/tsnad

[323]

TSAFinder

2022

netMHCpan4.0 Class I

RNAseq FASTQ files for matched tumor and control

RNA-seq only pipeline; translates every RNA-seq read into all possible 8–11mer peptides; HLA typing from RNA-seq.

https://github.com/RNAseqTSA

[324]

nextNEOpi

2022

pVACseq, NetMHCpan and MHCflurry, NeoFuse Class I & II

FASTQ or BAM of normal-tumor WES or WGS and optionally RNA-seq

Quantifies patient- and neoepitope-specific attributes linked to tumor immunogenicity and therapy response.

https://github.com/icbi-lab/nextNEOpi

[325]

ProGeo-Neo v2.0

2022

NetMHCpan, NetMHCIIpan Class I & II

FASTQ of paired normal-tumor WGS/WES and tumor RNA-seq data

Presents a proteogenomic approach that combines HLA presentation analysis with direct identification of mutant peptides through MS data.

https://github.com/kbvstmd/ProGeo-neo2.0

[326]

TSNAdb v2.0

2023

DeepHLApan, MHCflurry, NetMHCpan 4.0 Class I

Somatic SNVs, INDELs, and gene fusions from TCGA; experimentally validated neoantigens

Stringent multi-tool criteria to reduce false positives; coverage of three mutation types (SNVs, INDELs, fusions) with per-mutation neoantigen counts; identification of shared neoantigens recurring across patients.

https://pgx.zju.edu.cn/tsnadb/

[327]

PGNneo

2023

NetMHCpan-4.1 Class I

RNA-seq profiles and MS datasets as input

Extends neoantigen discovery to noncoding regions via proteogenomics; four modules (noncoding variant calling and HLA typing, peptide extraction and DB construction, variant peptide ID via MS, neoantigen prediction and prioritization); designed to reduce false positives by requiring peptide-level MS support.

https://github.com/tanxiaoxiu/PGNneo

[328]

LENS

2023

NetMHCpan, NetMHCstabpan, MHCFlurry, DeepHLAPan Class I

FASTQ of paired normal-tumor DNA and RNA-sequencing data

Modular, extensible, multi-workflow neoantigen prediction platform built on top of Nextflow DSL2, broader tumor antigen coverage.

https://gitlab.com/landscape-of-effective-neoantigens-software

[329]

Neo-intline

2023

NetMHCpan 4.0, NetMHCIIpan, Uses similarity to known T cell epitopes (IEDB) via BLASTP Class I & II

Full simulation of T cell epitope presentation, including proteasome processing, TAP transport, MHC binding, TCR recognition, and a unified scoring system producing ranked neoantigen candidates.

https://github.com/zoolie/neoantigenPre

[330]

NeoHunter

2024

NetMHCpan, NetMHCstabpan, ERGO Class I

RNA-seq and/or WES/WGS

Detects not only SNV- and indel-derived neoantigens but also gene fusion- and aberrant splicing-derived neoantigens; predicts TCR recognition both indirectly (via agretopicity, foreignness) and directly (DL on TCR–pMHC data).

https://github.com/XuegongLab/NeoHunter

[331]

ImmuneMirror

2024

RF Class I & II

FASTQ of matched normal-tumor WES and tumor bulk RNA-seq (optional)

Generates a visual analysis report for each sample; AUC = 0.87.

https://github.com/weidai2/ImmuneMirror

http://immunemirror.hku.hk/App/

[332]

ImmuneMirror is an ML-driven integrative pipeline that embeds a balanced random-forest model, trained on experimentally validated neopeptides, within a stand-alone workflow and web server for neoantigen prediction and prioritization [332]. The ImmuneMirror workflow processes matched tumor-normal exome data, optionally combined with tumor RNA-seq, to call somatic variants, infer HLA class I and II alleles, evaluate microsatellite instability, and compute composite neoantigen scores, which were associated with clinical outcomes in large gastrointestinal cancer cohorts [332]. Proteogenomic pipelines such as ProGeo-Neo v2.0 extend this concept by integrating whole-genome or exome sequencing, RNA-seq and LC-MS/MS immunopeptidomics and by supporting both MHC class I and II neoantigen prediction in a single one-stop software environment [326]. By allowing detection of multiple mutation classes, including single-nucleotide variants, small insertions or deletions, frameshifts and gene fusions, while filtering peptide candidates against MS-confirmed mutant peptides, ProGeo-Neo v2.0 enriches for high-confidence, naturally presented neoepitopes [326]. NeoFuse focuses specifically on gene-fusion-derived neoantigens, providing an RNA-seq-based pipeline that predicts fusion transcripts, translates them into chimeric peptides, infers patient MHC class I types, and scores peptide-MHC binding to nominate fusion neoepitopes [317]. Because fusion neoantigens can generate immunogenic neoantigens that mediate anticancer immune responses and expand the neoantigen repertoire, NeoFuse helps extend integrative approaches beyond point mutations and small indels [317]. nextNEOpi builds on existing modules such as pVACseq and NeoFuse to offer a comprehensive workflow that accepts tumor-normal WES/WGS (and optional RNA-seq) as input and jointly predicts SNV-, indel- and fusion-derived class I and II neoepitopes [325]. In addition to neoepitope prediction, nextNEOpi incorporates tumor purity and other tumor-intrinsic features, enabling downstream association of neoantigen burden and quality with immune contexture and therapy response [325]. NeoHunter further generalizes integrative neoantigen discovery by providing flexible software for systematically detecting candidate neoantigens from sequencing data and combining support for multiple mutation types with HLA-binding and TCR-recognition models [331].

Collectively, these approaches integrate the individual components described in previous sections into practical, scalable workflows that can be deployed in translational studies and, with appropriate validation, in personalized neoantigen vaccine design.

Recommendations: Users should begin by evaluating their available datasets and research goals, such as whether they have RNA-seq alone or paired genomic and transcriptomic data, and which types of neoantigens (SNVs, indels, or fusions) they want to investigate. For general use, well-documented and end-to-end pipelines like pVACtools are a good starting point, while more specialized tools may be better for specific research needs. It is also advisable to combine multiple prediction methods and include additional evidence like gene expression to improve accuracy. Finally, users should consider ease of use, computational requirements, and reproducibility when selecting a tool.

Immunopeptidomics and AI-driven neoantigen discovery

MS-based immunopeptidomics directly identifies MHC-associated peptides eluted from tumor HLA molecules, yielding large catalogs of patient-presented ligands for neoantigen discovery [40, 49, 333].

Integration of genomic sequencing with HLA ligandome profiling has enabled direct identification of mutated peptide ligands on native tumor tissue and demonstration of their immunogenicity in patients [40, 49, 91]. Pioneering proteogenomic studies combined exome/transcriptome sequencing with MS to discover neo-epitopes, showing that only a small fraction of variants are confirmed at the ligand level [91].

Analyses of the tumor antigenic repertoire emphasize that immunopeptidomics complements genomics by empirically defining the set of tumor-associated antigens and neoantigens visible to T cells [41, 49]. Large-scale immunopeptidome datasets from mono-allelic model systems and diverse tumor types have become foundational training resources for AI-based peptide–MHC presentation models [11, 49, 188, 192, 208]. In mono-allelic systems, Abelin et al. profiled HLA-associated peptidomes and showed that incorporating MS-derived ligands into prediction algorithms improves epitope prediction compared with affinity-based methods [188]. Several studies combined hundreds of thousands of MS-identified ligands with ML architectures, substantially improving positive predictive value for class I ligand prediction and enabling improved neoantigen prioritization across diverse alleles [11, 192]. Building on these resources, Pyke et al. integrated large tumor immunopeptidomes with composite presentation models to demonstrate “precision neoantigen discovery” in clinically relevant cohorts [208].

State-of-the-art MS-trained models such as HLAthena, MHCflurry 2.0, BigMHC, and MaNeo integrate peptide sequence and HLA context, achieving higher accuracy than binding-affinity-only predictors in cross-validation and benchmarking studies [11, 161, 204, 208, 334]. These composite presentation scores correlate with tumor immunopeptidomes and improve identification of clinically relevant neoantigens in benchmarking studies and patient cohorts [161, 192, 208, 335]. For HLA class II, the multimodal MARIA framework integrates MS-identified ligands with expression and cleavage features to improve neoantigen identification performance and enrich for CD4+ responses [160].

Proteogenomic workflows extend AI-based immunopeptidomics by searching spectra against customized protein databases that include variant peptides, alternative splice isoforms, gene fusions, and non-coding translational events [48, 49, 91, 326, 333]. This strategy enables detection of non-canonical tumor antigens, including phosphopeptides, splice-derived peptides and cryptic antigens from noncoding regions, which can elicit strong T cell responses [40, 48, 49, 91]. In particular, these antigen sources correspond to multiple classes of entries summarized in Table 9, including gene fusion–derived neoantigens (e.g., FusionGDB [336], TumorFusions [337], FusionNeoAntigen [338]), RNA splicing–derived isoforms (e.g., FLIBase [339]), and non-coding region–derived peptides (e.g., IEAtlas [340]), highlighting their relevance for systematic cataloging of non-SNV cancer neoantigens. Software such as ProGeo-Neo v2.0 and pVACtools implement end-to-end pipelines in which MS-confirmed ligands are combined with AI-based predictors to filter, score and prioritize neoantigen candidates for vaccines or T cell therapies, while best-practice guidelines and resources like the HLA Ligand Atlas address current bottlenecks such as incomplete ligandomes and under-representation of rare HLA alleles [47, 49, 238, 319, 326, 335].

Table 9.

Databases cataloging non-SNV cancer neoantigens

Neoantigen Source Category Database Description Web Accessibility Ref.
DNA alterations Gene fusion FusionGDB Compiles ~48K cancer fusion genes and provides functional annotations including ORF prediction, protein domain retention, sequences, and clinical relevance to support biomarker and therapeutic discovery in cancer. https://compbio.uth.edu/FusionGDB/ [336]
Gene fusion TumorFusions A TCGA-based database of ~20,000 gene fusions across cancers, validated by genomic rearrangements and enriched with functional annotations for prioritizing cancer-relevant fusion events. http://www.tumorfusions.org/ [337]
Gene fusion FusionNeoAntigen A comprehensive resource for fusion gene-derived neoantigens, including fusion protein sequences, HLA binding predictions, and potential vaccine or CAR-T targets. https://compbio.uth.edu/FusionNeoAntigen [338]
RNA aberrations Alternative splicing FLIBase A long-read–based transcriptome database that catalogs full-length splice isoforms, including many novel and tumor-specific transcripts with potential neoantigen relevance. http://www.flibase.org/ [339]
Non-Coding genomic regions IEAtlas A database of HLA-presented immunogenic epitopes derived from non-coding regions, integrating MS data and immunogenicity features to support cancer vaccine and immunotherapy research. http://bio-bigdata.hrbmu.edu.cn/IEAtlas [340]
Transposable elements TEITbase A database of transposable element (TE)-initiated transcripts across 33 cancer types, including 6,203 tumor-specific transcripts, onco-exaptation events, and candidate tumor-specific antigens derived from TE activation. http://teitbase.medbioinfo.org/ [341]
Post-translational modifications (PTMs) Modified tumor antigens / PTM-derived HLA ligands caAtlas caAtlas is a human cancer immunopeptidome resource that catalogs MHC/HLA-presented peptides identified from cancer immunopeptidomic datasets, including modified peptides and PTM-associated tumor antigens, supporting prioritization of candidate peptides for immunogenicity testing and cancer immunotherapy development. http://www.zhang-lab.org/caatlas/ [342]

Collectively, integrating WES/RNA-seq with MS/MS triangulates evidence toward “valid HLA-restricted tumor antigens” and guides rigorous target prioritization for vaccines, TCR-T, and other antigen-directed therapies [41]. Clinically, MS-defined ligands, including post-translationally modified phosphopeptides, have advanced into first-in-human vaccination, demonstrating safety and immunogenicity in patients with high-risk melanoma [43].

Personal neoantigens vs. shared neoantigens

Personal (private) neoantigens arise uniquely in individual tumors or subclones. Their mosaic expression contributes to intra-tumor heterogeneity, complicating durable immune targeting. Vaccines based on these neoantigens require complex workflows—tumor sequencing, epitope prediction, immunopeptidomics, and personalized GMP manufacturing—making them costly, labor-intensive, and prone to false positives in epitope prediction. While clinical results have been encouraging, challenges include immune evasion by subclones lacking the targeted antigens and uncertainty about the proportion of tumor cells that must express the neoantigen for therapeutic efficacy [343]. In contrast, shared (public) neoantigens derive from recurrent driver mutations (e.g., TP53, KRAS, BRAF, EGFR, HER2, PI3K, TARP) present across many patients. Unlike personal neoantigens, they are typically clonal, avoid issues of mosaicism, and are less susceptible to immune escape. They can serve as “off-the-shelf” immunotherapies with broader applicability, provided their epitopes are appropriately presented in the patient’s HLA context. Similar to personal neoantigens, they are tumor-specific and generally pose minimal risk of autoimmunity [344, 345]. Targeting shared neoantigens in cancer immunotherapy remains difficult but offers considerable promise due to their tumor specificity and wide therapeutic relevance.

Over 150 clinical trials investigating neoantigen-based immuno-oncotherapies have been launched [346]. Notably, a clinical trial investigating TARP-based (shared neoantigen) vaccination in prostate cancer demonstrated significant tumor growth reduction in the majority of patients [347]. Similarly, a phase 1 clinical trial tested TP53 and KRAS vaccines in patients with advanced solid tumors, showing encouraging results [348].

Li et al. [349] quantified Tumor Mutational Burden (TMB) across 30 cancer types. Skin cutaneous melanoma (SKCM) showed the highest proportion of high-TMB tumors (49.4%). Lung cancers ranked next, with lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC) displaying 36.9% and 28.1% high-TMB cases, respectively (Fig. 4). In tumors with a low mutational burden, vaccination based on personalized neoantigens may have limited efficacy. Instead, the use of shared neoantigens represents a potential therapeutic strategy.

Fig. 4.

Fig. 4

Comparison of tumor mutational burden (TMB) across different cancer types. The bar plot illustrates the proportion of patients with high TMB in various cancer types. The highest rates of elevated TMB are observed in skin cutaneous melanoma (SKCM, 49%), lung adenocarcinoma (LUAD, 37%), and lung squamous cell carcinoma (LUSC, 28%), followed by bladder carcinoma (BLCA, 26%) and uterine corpus endometrial carcinoma (UCEC, 23%). In contrast, several cancer types, including kidney renal papillary cell carcinoma (KIRP), thyroid carcinoma (THCA), pheochromocytoma and paraganglioma (PCPG), kidney renal clear cell carcinoma (KIRC), and mesothelioma (MESO), show negligible or absent frequencies of high TMB. This comparison highlights the heterogeneity of mutational load among distinct cancer types, which may have implications for immunotherapy response

Challenges and future directions

Despite rapid progress in AI-guided neoantigen discovery, method performance and interoperability still lag the field’s clinical ambition; notably, there remains no consensus analysis workflow and few established best practices, underscoring the need for standardized datasets and reporting across the pipeline from variant calling to immunogenicity testing [335].

Consortium-scale benchmarking highlights both promise and limitations; for instance the TESLA study showed that integrating presentation- and recognition-relevant features filtered out 98% of non-immunogenic peptides with a precision above 0.70, while noting that, prior to the consortium effort, there were no common reference datasets to systematically compare prediction methods [12].

Data choice and modeling remain core bottlenecks. Training directly on endogenous HLA ligands from tumors and incorporating contextual features can markedly improve performance; EDGE increased the positive predictive value of HLA antigen prediction by up to ninefold [192], but the magnitude of gain varies by allele, tumor type, and evaluation regime. Likewise, large mono-allelic immunopeptidome resources plus integrative modeling (HLAthena) achieved a 1.5-fold improvement in positive predictive value and correctly identified > 75% of HLA-bound peptides observed in patient-derived tumor cell lines [11], advancing allele coverage yet still leaving gaps in rare alleles and class II presentation. Mono-allelic immunopeptidomics coupled to neural networks outperformed affinity-trained algorithms, illustrating the value of training on EL data and offering a scalable strategy to learn endogenous presentation rules [188].

Validation and antigen source completeness are parallel challenges. Foundational proteogenomic studies showed that among missense-derived candidates only a small fraction of predicted binders were confirmed by MS, highlighting the need for rigorous, prospective orthogonal validation at scale [91]. Moreover, ~90% of tumor-specific antigens in one screen arose from allegedly noncoding regions and would have been missed by standard exome-centric pipelines, arguing for systematic inclusion of noncanonical ORFs, aberrant transcripts, and proteasome-spliced peptides in training and inference [48]. Tumor evolution compounds these issues; allele-specific HLA loss occurs in ~40% of NSCLCs and is associated with subclonal neoantigen burden, reinforcing the importance of modeling immune escape and intratumoral heterogeneity during prioritization [92].

Accordingly, near-term priorities are clear: harmonized benchmarking, expanded and higher-accuracy HLA class II typing and prediction, and first-class support for alternative antigen sources and post-translationally generated peptides within pipelines [12, 335]. Clinically, personalized neoantigen vaccines have demonstrated feasibility, safety, and immunogenicity in melanoma, providing a scaffold on which to prospectively test whether AI-enhanced selection improves true clinical benefit, and to co-develop adaptive trial designs with TCR therapies and checkpoint blockade [4, 90].

Conclusion

AI-driven neoantigen discovery has evolved into a clinically testable paradigm, enabled by large immunopeptidomics datasets and DL predictors that now inform vaccine design [11, 90, 192]. Mono-allelic MS datasets and MS-informed models have improved presentation prediction and recovery of tumor-displayed ligands [11]. Likewise, DL trained on tumor immunopeptidomes markedly increases positive predictive value for antigen presentation [192]. Consortium-scale benchmarking has begun to clarify parameters of immunogenicity, enabling triage of the vast majority of non-immunogenic peptides at practical precision [11, 12].

Despite rapid progress, there is still no consensus end-to-end approach, and key gaps persist, including improving HLA class II typing accuracy, expanding support for diverse antigen sources, and incorporating clinical response data to refine prediction. Practically, mature pipelines should integrate somatic variant calling, precise HLA typing, peptide processing, and peptide–MHC binding prediction before prioritization and validation.

Anchoring pipelines to ground truth via direct MS of native tumor tissue remains essential to isolate bona fide neoepitopes. Early trials show that personalized neoantigen vaccines are feasible, safe, and immunogenic, eliciting broad T cell responses in melanoma [90]. Key priorities now include robust class II modeling and expansion beyond coding SNVs to noncanonical antigen sources, alongside convergence toward standardized, end-to-end workflows for clinical utility. Coupling presentation and recognition modeling with tumor-evolution constraints, such as frequent HLA loss of heterozygosity, should sharpen prioritization of clinically actionable targets. The field’s trajectory is clear: standardized, MS-informed, and recognition-aware AI models, coupled with rigorous validation, can accelerate rational, personalized vaccines with the ultimate aim of improving patient outcomes.

Acknowledgements

This study was supported by Tehran University of Medical Sciences Grant No. 1404-18-148-94336.

Abbreviations

AA-index

Amino acid index

AI

Artificial intelligence

ANN

Artificial neural network

APC

Antigen-presenting cell

AUC

Area under the (ROC) curve

BiGRU

Bidirectional gated recurrent unit

BLOSUM

BLOcks SUbstitution Matrix

BWA

Burrows–Wheeler Aligner

CDR

Complementarity-determining region

CLIP

Class II–associated invariant chain peptide

CNN

Convolutional neural network

ER

Endoplasmic reticulum

ESM

Evolutionary Scale Modeling

GBM

Glioblastoma multiforme

HLA

Human leukocyte antigen

HMM

Hidden Markov model

ICGC

International Cancer Genome Consortium

IEDB

Immune Epitope Database

INDEL

Insertion and deletion

IPD-IMGT/HLA

Immuno Polymorphism Database – International ImMunoGeneTics/Human Leukocyte Antigen database

LC-MS/MS

Liquid chromatography–tandem mass spectrometry

LOH

Loss of heterozygosity

LSTM

Long short-term memory

MHC

Major histocompatibility complex

ML

Machine learning

MS

Mass spectrometry

MS/MS

Tandem mass spectrometry

NGS

Next-generation sequencing

NSCLC

Non-small-cell lung cancer

ORF

Open reading frame

PBR

Peptide-binding region

PFR

Peptide flanking residue

PPV

Positive predictive value

PTM

Post-translational modification

pMHC

Peptide–major histocompatibility complex

RNA-seq

RNA sequencing

RNN

Recurrent neural network

SNV

Single-nucleotide variant

SVM

Support vector machine

TAP

Transporter Associated with Antigen Processing

TARP

TCR gamma alternate reading frame protein

TCGA

The Cancer Genome Atlas

TCR

T cell receptor

TCR-T

T cell receptor–engineered T cell therapy

TESLA

Tumor Neoantigen Selection Alliance

TMB

Tumor mutational burden

TSA

Tumor-specific antigen

WES

Whole-exome sequencing

WGS

Whole-genome sequencing

WHO

World Health Organization

Author contributions

A.B: main idea and conceptualization, review of literature, analysis and interpretation, writing—original draft preparation, Figure preparation; S.G: review of literature, writing—original draft preparation; F.F: writing—original draft preparation; E.E: review of literature, writing—original draft preparation; A.A: writing—original draft preparation; A.E., B.N., K.K., G.K. and M.M: writing—review and editing. All authors have read and approved the final manuscript.

Funding

Not applicable.

Data availability

No datasets were generated or analysed during the current study.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from Journal of Translational Medicine are provided here courtesy of BMC

RESOURCES