Skip to content
Open access

Peptide-HLA II interaction prediction for post-translationally modified peptides

Aug 2026 · bioRxiv · 0 citations · 37 references
Biology

TL;DR

PepChem, a deep learning model utilizing novel, molecular-level peptide representations that enable predictions for sidechain modifications, bridges the critical gap in PTM-aware immune recognition prediction, with immediate applications in autoimmunity, cancer, and infectious disease.

Abstract

CD4+ T cells recognize peptides presented by human leukocyte antigen (HLA) II, implementing a fundamental mediation mechanism of the adaptive immune system. Although post-translational modifications (PTMs) alter immune responses, PTM-peptide-HLA interaction prediction remains challenging due to data scarcity resulting from substoichiometric levels of PTMs. To overcome this, we developed PepChem, a deep learning model utilizing novel, molecular-level peptide representations that enable predictions for sidechain modifications. Using monoallelic datasets that we reanalyze for PTMs of interest, we show accurate predictions on PTMs that were unseen during training. Furthermore, we introduce a novel training protocol that improves PTM-peptide generalization compared to conventional methods. We predict and experimentally validate citrullination-induced binding increase of rheumatoid arthritis (RA)-linked peptides to HLA II risk allele DRB1*04:01. This framework bridges the critical gap in PTM-aware immune recognition prediction, with immediate applications in autoimmunity, cancer, and infectious disease.

Read PDF

Similar papers

Aug 2026

A predictive multiscale framework for post-translational modification-dependent peptide-MHC class I binding.

A unified structure-energy-dynamics model explaining how Post-translational modifications function as atomic-level chemical switches in antigen presentation is established, establishing a unified structure-energy-dynamics model explaining how PTMs function as atomic-level chemical switches in antigen presentation.

Xiao Yao, Yue Gang, Yao-Yue Zhang et al. · 0 citations
Sep 2026

Deciphering T-cell receptor-antigen recognition through interpretable residue-level interaction modeling.

Accurate identification of interactions between T-cell receptors (TCRs) and antigenic peptides presented by major histocompatibility complex (MHC) molecules is essential for advancing precision immunotherapy. However, existing approaches often exhibit limited generalization to unseen peptides and struggle to capture the complex interaction patterns underlying immune recognition. Here, we present TCR-IFNet, a biologically informed deep learning framework for interpretable TCR-peptide interaction prediction. The model integrates global contextual representations from protein language models with local motif refinement via a gated convolutional module. To model cross-sequence dependencies, we introduce a Fast Kolmogorov-Arnold Network (FastKAN)-based cross-attention mechanism for nonlinear interaction modeling, together with a bilinear attention network to aggregate residue-level features into compact interface representations. Evaluation across multiple settings indicates that TCR-IFNet achieves competitive performance compared with existing methods, with higher AUPRC observed on both antigen-specific and healthy-sourced datasets, as well as improved results on independent test sets. The model also shows consistent generalization to unseen peptides under different negative sampling strategies. In addition, TCR-IFNet provides biologically meaningful interpretability by identifying key residue-level interaction patterns consistent with structural binding interfaces. Collectively, these findings demonstrate that TCR-IFNet provides a robust and generalizable computational framework for characterizing TCR-peptide interactions.

Wen-Yu Xi, Ruheng Wang, Xiu-Cai Ye et al. · 0 citations
Open access Aug 2026

PepGen: conditional generation of peptides for MHC binding

Abstract Motivation Peptide-MHC II binding drives adaptive immunity, yet discovery of novel binder peptides remains challenging due to open binding grooves of MHC-II that accommodate variable-length peptides. While discriminative models perform well, they are unfeasible for generation via enumeration due to vast peptide space (2013≈8×1016 for peptides of length 13 amino acids). Generative AI approaches could accelerate binder design to enable vaccines targeted to particular MHC-II alleles or optimize other peptide chemical properties. Results We introduce PepGen, the first protein language model for MHC II peptide generation building on Generalized Language Modeling. PepGen conditions on alleles, arbitrary partial peptides including putative TCR-interacting motifs, and continuous binding affinity. Across multiple benchmarks including infilling and de novo generation, PepGen outperformed frequency sampling, Gibbs clustering, and autoregressive baselines. Adjusted log-probabilities enable good classification performance. Experimental validation confirmed that the SARS-CoV-2 peptide TEGALNTPKDHIGTR binding the HLA-DQA101:03-DQB106:03 allele can be redesigned to bind the HLA-DQA101:02-DQB105:02 allele. PepGen generated three putative TCR-motif-preserving binders gaining up to 70% of original MFI. Overall, PepGen provides scalable, motif-constrained MHC II peptide redesign and de novo generation, validated through thorough benchmarks and functional assays. Availability and implementation Code and Data are available at https://github.com/DaniTheOrange/PepGen.

Dani Korpela, A. Dumitrescu, Martin Stražar et al. · 0 citations
Open access Aug 2026

CALFP-MHC: Interpretable Pan-Allelic Prediction of Peptide-MHC Binding and Presentation Using Chemically Grounded Fingerprints and Contrastive Learning

Identifying which peptides bind major histocompatibility complex (MHC) molecules is central to vaccine design, neoantigen prioritization, and precision immunotherapy. Existing deep learning predictors largely encode amino acids as discrete symbols, thereby missing the residue-level chemistry driving molecular recognition. Performance also tends to degrade under class imbalance, for rare alleles, and on peptide– MHC combinations outside the training distribution. We developed CALFP-MHC, a framework that encodes each amino acid as a set of complementary cheminformatics fingerprints capturing functional groups, atomic connectivity, and substructural features, and combines positional encoding with supervised contrastive pre-training to organize the latent space by binding class before fine-tuning a binary classifier. Peptide–MHC interactions are modeled through a hybrid convolutional-transformer backbone. In a large-scale computational benchmark covering ∼18.7 million peptide–MHC pairs across 112 HLA class I and 53 class II alleles, CALFP-MHC achieved AUCs of 0.93-0.97 and PPVs of 0.66–0.94. Critically, performance remained above AUC 0.90 even at a 200:1 negative-to-positive ratio, where competing tools frequently collapsed toward chance. On independent experimental data containing 3,627 class I and 520 class II MS/MS-confirmed ligands and 570 validated neoantigens, the model maintained strong discrimination, correctly prioritizing immunogenic peptides and MHC-presented ligands. Attention and integrated-gradient analyses recovered established anchor positions (P2 and PΩ for class I, P1, P4, P6, and P9 for class II) and highlighted chemically interpretable functional groups consistent with known binding determinants. CALFP-MHC demonstrates that grounding residue representations in molecular chemistry, rather than sequence symbols alone, improves both robustness and interpretability in peptide–MHC binding prediction.

My-Diem Nguyen Pham, T. Ho, H. Nguyen et al. · 0 citations
Open access Aug 2026

Label Noise Limits TCR-pMHC Specificity Prediction: Improved Performance Through AlphaFold3-Based Structural Modeling and Data Denoising

T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.

Pilar Ballesteros-Cuartero, J. Lund, Morten Nielsen · 1 citation · ⚡1
Open access Jul 2026

ImmunoFoundation: A Multimodal Deep Learning Approach to Immunogenicity Prediction 2310036

The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.

Smita Krishnaswamy, J. Rocha, Hiren Madhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.