Adaptive immunity relies on T-cell receptor (TCR) recognition of peptides presented by the major histocompatibility complex (pMHC). Accurate prediction of TCR:pMHC binding pairs from sequence data remains a longstanding challenge in computational immunology, limiting the development of precision immunotherapies like cancer vaccines and adoptive cell therapies. Here, we present enFoldX (ensemble of Folded compleXes), a structure-based approach leveraging biophysical characterization of AlphaFold3-generated ensembles to classify TCR:pMHC sequence pairs as cognate versus non-cognate. Unlike previous methods reliant on only sequence data or a single, static predicted structure, enFoldX extracts features from an entire generated ensemble with a custom focus on the biophysical binding interface. Our model distinguishes T cell reactivity between peptides differing by a single amino acid substitution, the resolution required for cancer neoantigens, and generalizes to unseen peptides, MHCs, and TCRs, a major objective for artificial intelligence (AI) in immunology. Our performance on these crucial tasks demonstrates that diverse, structural sampling of biophysical interactions over an ensemble is fundamental for accurate AI-driven binding predictions and offers lessons for efficient future data generation to improve models. Our findings therefore offer a scalable framework to accelerate therapeutic binder design, and we provide access to a publicly available code repository.
T-cell immunity acts as a major defense system against controlling viral infections in vertebrates. During viral entry, innate immune cells degrade the viral proteins (antigens) and present them on their surface via Major Histocompatibility (MHC) proteins. T-cell receptors (TCRs) recognize these antigens/peptides presented by MHC (pMHC), initiating a T-cell mediated immune response. Despite its significance, the mechanism by which pMHC-TCR binding triggers T-cell activation remains unclear. In this study, we employed an integrative computational approach combining Bioinformatics, Molecular Dynamics (MD) simulations, and Machine Learning (ML) to identify viral epitopes as potential vaccine candidates. We performed large-scale all-atom and coarse-grained MD simulations on MHC-peptide-TCR complexes embedded into dendritic and T-cells, for which experimental immunogenicity data is available. One hundred fifty such systems are simulated for 1 {\mu}s each to capture the conformational and dynamical changes that underlie T-cell activation. Our ML model (DynamiT), trained on simulation-derived structural and dynamical features extracted from 2500 time points, revealed key determinants responsible for T-cell activation with an accuracy of 73.3%. Notably, we have identified the bending of the TCR transmembrane region, major dynamic motions of the TCR{\alpha} constant region and the buried surface area at the pMHC and TCR interface as critical factors influencing immune response initiation. Our approach unravels the mechanism of T-cell mediated immune response and helps ML-guided screening of viral epitopes for vaccine development.
Jaya Vasavi Pamidimukkala, Roshan Balaji, N. Bhatt et al.· 0 citations
The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on major histocompatibility complex class I (MHC-I) molecules, collectively termed peptide–MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting cancer neoantigens, accurately predicting which peptides bind to MHC alleles is critical. Current computational methods for pMHC-I binding prediction fall into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and pMHC binding energetics. Although sequence-based methods are widely used, their performance depends on the size and quality of the training data. Structure-based approaches, by contrast, can generalize better across diverse MHC alleles, but they traditionally depend on identifying a single global minimum-energy conformation, an assumption that may be inadequate for the promiscuous binding of MHC-I molecules. To address these limitations, we developed STRUMP-I (STRUcture-based pMHC Prediction for class I), a novel pMHC-I binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine learning features. In the standard benchmark set, STRUMP-I achieved performance comparable to state-of-the-art sequence-based models overall and showed a clear advantage for alleles with limited or imbalanced representation. Furthermore, STRUMP-I complemented sequence-based methods by removing method-specific false positives and improving precision, with a more favorable precision–recall tradeoff than AF-FT. These evaluations reinforced the value of STRUMP-I as a structure-informed prioritization method, particularly for underrepresented alleles and as a high-precision post-prediction filter.
Adam Voshall, Jeongjun Chae, Honglan Li et al.· Computational and Structural...· 0 citations
ABSTRACT Monoclonal antibody-derived chimeric antigen receptors (CAR) that mimic T‑cell receptors (TCR) on binding to peptide-major histocompatibility complex (pMHC) and activating T‑cell functions hold great promise for the development of effective immunotherapy. This study applied AlphaFold 3 (AF3) to model the quaternary structures of TCR mimic antibody variable fragment (TCRm Fv) complexed with NY-ESO-1157–165/HLA-A*02:01. Benchmark study suggested that reliable TCRm Fv–pMHC structures were achieved by AF3 with high confidence. Toward NY-ESO-1/A2-specific CAR clones, AF3 prediction revealed their intermolecular geometry resembling the canonical TCR engagement, and binding avidity tests confirmed the peptide-dependent recognition of pMHC. Molecular interactions between Fv and the peptide antigen and its binding groove on MHC were further pinpointed, and the predicted epitopes and paratopes were validated by mutagenesis studies. Aiming to generate TCRm Fv mutants of improved affinity, AF3 aided designs with larger Fv–pHLA contact surface areas and more hydrogen bonds toward the peptide antigen were successfully identified. However, these AF3‑designed mutants failed to deliver enhancement on binding avidity or functional potency in experimental tests. Overall, AF3 is a highly valuable tool for TCRm Fv–pMHC complex modeling, and its combination with force field-based methods will be desirable to aid optimization tasks.
Abhijit Sarkar, Sara Michael, Zening Wang et al.· mAbs· 0 citations
T cell antigen-specific immunity depends on pairwise interactions between T cell receptors and peptide- MHC, yet isolating the TCR-pMHC pairs that drive productive engagement remains a major obstacle for antigen-specific therapeutics and for decoding TCR specificity. We overcome this by co-encoding TCR and pMHC in a single founder cell, then clonally expanding it inside a semi-permeable capsule so that genetically identical daughter cells engage in trans. T cell activation, rather than binding affinity, is used to sort cells with functional pairs, and a single PCR on the clone’s linked genomic library captures both partners. This platform, LINC-seq, recovered known cognate pairs from pooled libraries at up to 95% accuracy and performed simultaneous, library-on-library deep mutational scanning of both partners. Wild- type clonotypes ranked among the top-enriched sequences in complex mixtures, and the screens resolved co-evolutionary epistasis and cross-reactivity rules inaccessible to one-sided mutagenesis. The approach generalizes to any receptor-ligand pair whose trans-engagement drives a reporter.
L. Liu, Seung Won Shin, Kevin M. Joslin et al.· bioRxiv· 0 citations
T cell receptor (TCR) binding to peptides presented by major histocompatibility complex (MHC) molecules is a key step in T cell activation, and forms the basis of adaptive immunity. Predicting this specificity is therefore essential to developing effective TCR-based immunotherapies and vaccines. Despite its clinical relevance, predicting TCR-pMHC specificity for previously unseen peptides remains an open problem, with structural modeling so far the only strategy showing any predictive power in this setting. In this study, we find that this limited performance is substantially driven by label noise in the data used to train and evaluate these methods, an effect that has so far been largely underexplored. Using an AlphaFold3-based pipeline adapted for TCR-pMHC structural modeling, we achieve state-of-the-art specificity prediction, outperforming AlphaFold2.3-based and sequence based methods, and performing at par with the leading Immrep2025 competition submission. Combining this pipeline with a cluster-based denoising algorithm, we show that removing mislabeled points from a large specificity dataset increased binder ranking accuracy by more than 70% relative to the full dataset. Together, these results highlight label noise as a major factor limiting the performance that any method in this field can achieve, and show that combining structural modeling with label denoising substantially improves TCR-pMHC specificity prediction, making such approaches an attractive complement to current sequence-based approaches for refining TCR target selection.
Pilar Ballesteros-Cuartero, J. Lund, Morten Nielsen· bioRxiv· 1 citation· ⚡1
T cell receptor (TCR) repertoire diversity enables the orchestration of antigen-specific immune responses against the vast space of possible pathogenic peptides. Identifying TCR/antigen specificity from the large TCR repertoire and antigen space is crucial for biomedical research.
We introduce copepodTCR, an open-access tool for the design and interpretation of high-throughput experimental assays to determine TCR/antigen specificity. copepodTCR implements a combinatorial peptide pooling scheme for efficient experimental testing of T cell responses against large overlapping peptide libraries, useful for “deorphaning” TCRs of unknown specificity. The scheme detects experimental errors and, coupled with a hierarchical Bayesian model for unbiased results interpretation, identifies the response-eliciting peptide for a TCR of interest out of hundreds of peptides tested using a simple experimental set-up.
Using in silico simulation, we demonstrated the applicability of our design scheme and the sensitivity of our results evaluation across varied experimental layouts and range of TCR-peptide activation signals. We experimentally validated our approach on a library of 253 overlapping peptides covering the SARS-CoV-2 spike protein, split across 12 pools. A single stimulation with combinatorial pools identified the correct epitope of two TCRs with known specificity and then deorphanized two SARS-CoV-2 associated TCRs shared among a large cohort of COVID-19 patients.
In conclusion, copepodTCR enables efficient and accurate mapping of TCR—peptide specificities through optimized combinatorial peptide pooling coupled with Bayesian inference. Beyond deorphanizing TCRs from established cell lines, we anticipate that copepodTCR can facilitate primary T cells deorphanization using single-sequencing as a read out, due to optimization of the pooling scheme, rational assignment of peptides and robust error-correction.
Simons Center for Quantitative Biology at Cold Spring Harbor Laboratory; Starr Centennial Scholarship; US National Institutes of Health Grants U01AI150747, R01AI136514, S10OD028632-01, 1R01AI167862; Simons Pivot Fellowship; National Natural Science Foundation of China Grant 62331002, National Science Foundation grant PHY-2210452.
Computational and Systems Immunology (COMP)
V. Kovaleva, David J Pattinson, Guanchen He et al.· Journal of Immunology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.