Jul 2026· Clinical Cancer Research· Vol 32, pp. A028-A028· 0 citations
TL;DR
An ensemble approach (enFoldX) that leverages structure prediction models such as AlphaFold3 to build sensitive binding predictors and outperforms the current co-folding methods which rely on predictions from the single top ranked structure.
Abstract
Adaptive immunity relies on T-cell receptor (TCR) recognition of non-self epitopes, short peptides presented by the Major Histocompatibility Complex (MHC) on the cell surface. Accurate computational prediction of TCR-epitope binding would unlock the development of targeted immunotherapies, such as cancer vaccines and TCR T cell therapies, while simultaneously deepening our fundamental understanding of self/nonself discrimination, pathogen recognition, and autoimmunity. We created an ensemble approach (enFoldX) that leverages structure prediction models such as AlphaFold3 to build sensitive binding predictors. enFoldX can distinguish T cell reactivity between peptides that differ by a single amino acid substitution, as needed for cancer neoantigens. enFoldX utilizes a customized highly parallelized workflow which allows us to produce ensembles of predicted protein structures at scale and train classifiers to infer reactivity based on distributions of engineered structure features and alignment confidence metrics. While state-of-the-art sequence-based approaches we evaluated could predict well for observed TCRs and epitopes close in sequence to training data, their applicability to novel sequences was limited. Conversely, our ensemble approach is the only model that showed true generalizability to novel datasets and even across species. Moreover, our ensemble approach outperforms the current co-folding methods which rely on predictions from the single top ranked structure. By leveraging the entire protein universe at scale, structure ensembles therefore enable classifiers that reflect physical free energies, providing a tractable path towards TCR T therapy design at the sensitivity required for cancer neoantigen discrimination and imparting lessons for a wide array of complex binding problems.
Olga Lyudovyk, Jonathan A. Levine, Melissa Pathil, Stephen Martis, Yuval Elhanati, Vinod P. Balachandran, Quaid Morris, Benjamin D. Greenbaum. enFoldX: AI classification of AlphaFold3-derived structural ensembles enables T cell specificity prediction [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A028.
Accurate identification of interactions between T-cell receptors (TCRs) and antigenic peptides presented by major histocompatibility complex (MHC) molecules is essential for advancing precision immunotherapy. However, existing approaches often exhibit limited generalization to unseen peptides and struggle to capture the complex interaction patterns underlying immune recognition. Here, we present TCR-IFNet, a biologically informed deep learning framework for interpretable TCR-peptide interaction prediction. The model integrates global contextual representations from protein language models with local motif refinement via a gated convolutional module. To model cross-sequence dependencies, we introduce a Fast Kolmogorov-Arnold Network (FastKAN)-based cross-attention mechanism for nonlinear interaction modeling, together with a bilinear attention network to aggregate residue-level features into compact interface representations. Evaluation across multiple settings indicates that TCR-IFNet achieves competitive performance compared with existing methods, with higher AUPRC observed on both antigen-specific and healthy-sourced datasets, as well as improved results on independent test sets. The model also shows consistent generalization to unseen peptides under different negative sampling strategies. In addition, TCR-IFNet provides biologically meaningful interpretability by identifying key residue-level interaction patterns consistent with structural binding interfaces. Collectively, these findings demonstrate that TCR-IFNet provides a robust and generalizable computational framework for characterizing TCR-peptide interactions.
Wen-Yu Xi, Ruheng Wang, Xiu-Cai Ye et al.· International Journal of Bio...· 0 citations
T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.
Pilar Ballesteros-Cuartero, J. Lund, Morten Nielsen· bioRxiv· 1 citation· ⚡1
Abstract Cross-immunity, defined as the ability of T-cells to recognize multiple antigen peptide-major histocompatibility complexes, is a fundamental feature of adaptive immunity. However, the prediction of different peptide epitopes that can be recognized by the same T-cell receptor remains challenging. Currently, artificial intelligent (AI)-based machine learning (ML) methods can be successfully used for pattern recognition in epitope molecular space by detecting the functional similarity between peptide sequences. In this study, using literature-based experimental data, we examined ML-based binary classification models trained on small datasets to predict the activity of nine-amino-acid-long peptides. Our results suggest that the consensus function of well-established similarity matrix-based representations and structural-based descriptors of epitopes yields better performance because representation-specific noises are reduced and individual model weaknesses are partially compensated. We also sought to determine the extent to which the predictive power of the applied AIs procedure depended on the physicochemical content of the descriptor set during the training process. In addition, challenging the models, we applied them to an independent experimental dataset to examine the effects of diverse laboratory conditions on a regulated biological measurement. In summary, applying a consensus function can capture the biological complexity of cross-reactivity at the binary classification level, even when applied to relatively small datasets.
V. Resch, László Tóth, Anita Rácz et al.· Briefings in Bioinformatics· 0 citations
A lightweight post-hoc filter that requires no re-docking and is directly compatible with existing AF3 prediction pipelines and transferable to other diffusion-based complex predictors, providing a practical quality-assurance layer for antibody epitope mapping in early-stage drug discovery.
The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.
Smita Krishnaswamy, J. Rocha, Hiren Madhu et al.· Journal of Immunology· 0 citations
Identifying which peptides bind major histocompatibility complex (MHC) molecules is central to vaccine design, neoantigen prioritization, and precision immunotherapy. Existing deep learning predictors largely encode amino acids as discrete symbols, thereby missing the residue-level chemistry driving molecular recognition. Performance also tends to degrade under class imbalance, for rare alleles, and on peptide– MHC combinations outside the training distribution. We developed CALFP-MHC, a framework that encodes each amino acid as a set of complementary cheminformatics fingerprints capturing functional groups, atomic connectivity, and substructural features, and combines positional encoding with supervised contrastive pre-training to organize the latent space by binding class before fine-tuning a binary classifier. Peptide–MHC interactions are modeled through a hybrid convolutional-transformer backbone. In a large-scale computational benchmark covering ∼18.7 million peptide–MHC pairs across 112 HLA class I and 53 class II alleles, CALFP-MHC achieved AUCs of 0.93-0.97 and PPVs of 0.66–0.94. Critically, performance remained above AUC 0.90 even at a 200:1 negative-to-positive ratio, where competing tools frequently collapsed toward chance. On independent experimental data containing 3,627 class I and 520 class II MS/MS-confirmed ligands and 570 validated neoantigens, the model maintained strong discrimination, correctly prioritizing immunogenic peptides and MHC-presented ligands. Attention and integrated-gradient analyses recovered established anchor positions (P2 and PΩ for class I, P1, P4, P6, and P9 for class II) and highlighted chemically interpretable functional groups consistent with known binding determinants. CALFP-MHC demonstrates that grounding residue representations in molecular chemistry, rather than sequence symbols alone, improves both robustness and interpretability in peptide–MHC binding prediction.
My-Diem Nguyen Pham, T. Ho, H. Nguyen et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.