PepCL (Peptide-MHC Continual Learning), a continual learning framework for updating peptide-MHC predictors with new assay data while explicitly preserving prior MS knowledge, is introduced and it is demonstrated that PepCL allows MHCPrime to learn previously unseen, assay-specific information while preventing catastrophic forgetting that is typically observed with conventional fine-tuning.
Abstract
Understanding peptide-major histocompatibility complex (MHC) class I binding is critical for effective vaccine and immunotherapy design but is a combinatorially complex challenge for which prediction models have become essential. MHC ligands are typically identified at scale via untargeted mass spectrometry (MS), and this has built a strong base for peptide-MHC model training. However, MS incompletely captures the vast peptide-MHC space due to technical, sampling, and biological biases. Although recently developed experimental assays have queried such blind spots yielding complementary information, existing peptide-MHC predictors have not yet incorporated these orthogonal data and are not designed to be updated as new data are generated. Here, we introduce PepCL (Peptide-MHC Continual Learning), a continual learning framework for updating peptide-MHC predictors with new assay data while explicitly preserving prior MS knowledge. To enable PepCL, we also develop MHCPrime, a new state-of-the-art pan-allelic peptide-MHC prediction model, trained on publicly available MS data, that can be effectively updated under our framework. We demonstrate that PepCL allows MHCPrime to learn previously unseen, assay-specific information while preventing catastrophic forgetting that is typically observed with conventional fine-tuning. We evaluate PepCL and MHCPrime in a variety of biological contexts, including infectious disease and cancer, and show improved peptide-MHC prediction that transfers across alleles for broader applicability in clinical settings. Overall, our results establish PepCL as a flexible framework for extending the utility of peptide-MHC models by improving their predictive performance as immunopeptidomics assays continue to evolve and new data become available.
Identifying which peptides bind major histocompatibility complex (MHC) molecules is central to vaccine design, neoantigen prioritization, and precision immunotherapy. Existing deep learning predictors largely encode amino acids as discrete symbols, thereby missing the residue-level chemistry driving molecular recognition. Performance also tends to degrade under class imbalance, for rare alleles, and on peptide– MHC combinations outside the training distribution. We developed CALFP-MHC, a framework that encodes each amino acid as a set of complementary cheminformatics fingerprints capturing functional groups, atomic connectivity, and substructural features, and combines positional encoding with supervised contrastive pre-training to organize the latent space by binding class before fine-tuning a binary classifier. Peptide–MHC interactions are modeled through a hybrid convolutional-transformer backbone. In a large-scale computational benchmark covering ∼18.7 million peptide–MHC pairs across 112 HLA class I and 53 class II alleles, CALFP-MHC achieved AUCs of 0.93-0.97 and PPVs of 0.66–0.94. Critically, performance remained above AUC 0.90 even at a 200:1 negative-to-positive ratio, where competing tools frequently collapsed toward chance. On independent experimental data containing 3,627 class I and 520 class II MS/MS-confirmed ligands and 570 validated neoantigens, the model maintained strong discrimination, correctly prioritizing immunogenic peptides and MHC-presented ligands. Attention and integrated-gradient analyses recovered established anchor positions (P2 and PΩ for class I, P1, P4, P6, and P9 for class II) and highlighted chemically interpretable functional groups consistent with known binding determinants. CALFP-MHC demonstrates that grounding residue representations in molecular chemistry, rather than sequence symbols alone, improves both robustness and interpretability in peptide–MHC binding prediction.
My-Diem Nguyen Pham, T. Ho, H. Nguyen et al.· bioRxiv· 0 citations
Abstract Motivation Peptide-MHC II binding drives adaptive immunity, yet discovery of novel binder peptides remains challenging due to open binding grooves of MHC-II that accommodate variable-length peptides. While discriminative models perform well, they are unfeasible for generation via enumeration due to vast peptide space (2013≈8×1016 for peptides of length 13 amino acids). Generative AI approaches could accelerate binder design to enable vaccines targeted to particular MHC-II alleles or optimize other peptide chemical properties. Results We introduce PepGen, the first protein language model for MHC II peptide generation building on Generalized Language Modeling. PepGen conditions on alleles, arbitrary partial peptides including putative TCR-interacting motifs, and continuous binding affinity. Across multiple benchmarks including infilling and de novo generation, PepGen outperformed frequency sampling, Gibbs clustering, and autoregressive baselines. Adjusted log-probabilities enable good classification performance. Experimental validation confirmed that the SARS-CoV-2 peptide TEGALNTPKDHIGTR binding the HLA-DQA101:03-DQB106:03 allele can be redesigned to bind the HLA-DQA101:02-DQB105:02 allele. PepGen generated three putative TCR-motif-preserving binders gaining up to 70% of original MFI. Overall, PepGen provides scalable, motif-constrained MHC II peptide redesign and de novo generation, validated through thorough benchmarks and functional assays. Availability and implementation Code and Data are available at https://github.com/DaniTheOrange/PepGen.
Dani Korpela, A. Dumitrescu, Martin Stražar et al.· Bioinformatics· 0 citations
Accurate prediction of peptide–MHC (pMHC) binding is central to immunogenicity assessment, yet many existing predictors are trained and evaluated on narrow allele sets and restricted peptide-lengths. Here, we present MHChron, a unified pMHC binding prediction framework predicated on systematic data curation, meticulous engineering of dataset balance and diversity, and rigorous evaluation through careful splits controlling for data leakage. We assemble one of the most diverse pMHC training dataset reported to date, integrating publicly available binding data across a broad allele coverage (class I n=214, class II n=98) and peptide length range (from 8 to 36 residues). Using a focused and carefully sampled subset of this dataset, we train complementary sequence-based and structure-aware models and test them under increasingly stringent generalisation regimes. Both models achieve consistently strong performance, outperforming the evaluated state-of-the-art predictors despite being trained on numerically fewer data points. Notably, the structure-aware model did not consistently surpass the sequence-based model, except under the most demanding setting of extrapolation to unseen allele clusters, suggesting that performance gains stem primarily from dataset diversity and rigorous evaluation rather than architectural complexity. Sequence-based MHChron is released with reproducible installation and an automated whole-protein screening pipeline, enabling broad and practical use.
Marta Chronowska, Eugene Shrimpton-Phoenix, Tadas Kluonis· bioRxiv· 0 citations
The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.
Smita Krishnaswamy, J. Rocha, Hiren Madhu et al.· Journal of Immunology· 0 citations
T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.
Pilar Ballesteros-Cuartero, J. Lund, Morten Nielsen· bioRxiv· 1 citation· ⚡1
PepChem, a deep learning model utilizing novel, molecular-level peptide representations that enable predictions for sidechain modifications, bridges the critical gap in PTM-aware immune recognition prediction, with immediate applications in autoimmunity, cancer, and infectious disease.
A. Dumitrescu, Dani Korpela, Adrian M. Bebenek et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.