LM-QASAS provides a rapid, high-precision platform for monitoring humoral immunity against emerging threats by mapping sequences into a high-dimensional semantic embedding space, and identifies functionally convergent clusters of sequences that are semantically similar and exhibit transient expansion upon immune stimulation.
Abstract
The B-cell receptor (BCR) repertoire serves as a historical record of immunological events. However, deciphering antigen-specific sequences from this vast dataset remains a challenge, particularly for novel pathogens where prior knowledge is absent. While time-course analysis methods such as QASAS have proven effective for tracking immune responses, they rely on existing antibody databases, limiting their applicability to emerging diseases. To overcome this limitation, we introduce LM-QASAS, a reference-free computational framework that integrates antibody language models with repertoire dynamics. By mapping sequences into a high-dimensional semantic embedding space, LM-QASAS identifies functionally convergent clusters of sequences that are semantically similar and exhibit transient expansion upon immune stimulation. In healthy individuals vaccinated with SARS-CoV-2 mRNA vaccines, our method identified spike-specific sequences with over 90% purity, significantly outperforming methods based on simple sequence identity or abundance. Leave-one-out cross-validation demonstrated that LM-QASAS could accurately reconstruct immune dynamics in unseen individuals without external references. Conversely, the method showed limited sensitivity in an influenza vaccine cohort, revealing that the approach is most effective under conditions of robust clonal expansion (high signal-to-noise ratio), such as those induced by mRNA vaccines. LM-QASAS provides a rapid, high-precision platform for monitoring humoral immunity against emerging threats.
Human antibody diversification, achieved through gene selection and somatic hypermutation (SHM), is critical for protecting against diverse pathogens. This study investigates whether specific immune responses possess distinct receptor sequence patterns that differentiate them from the general immune repertoire. Utilizing data from an anti-SARS-CoV-2 vaccination study, we analyzed two properties of SARS-CoV-2 specific memory B-cells and compared them to the background immune repertoire. Driven by somatic hypermutation (SHM), B cells exhibit a highly dynamic nature. Consequently, groups sharing a direct lineage from a common progenitor are defined as B-cell clones. First, we studied substitution survival - the number of clones to survive amino acid substitutions across the variable region of the B cell receptor (BCR). Second, we analyzed clonal amino acid trimer usage patterns across the BCR gene to gain insight into prevalent genomic motifs found in different immune sub-repertoires. We demonstrated that these two metrics can effectively cluster and distinguish SARS-CoV-2 specific B cell responses. Furthermore, we observed that SARS-CoV-2-specific B-cells show an increased tendency to utilize and conserve a specific CDR2 motif derived from the VH3–30 gene and its alleles. Beyond identifying a specific germline motif related to SARS-CoV-2-specific B-cells response, our findings demonstrate that our novel analysis pipeline can successfully identify signatures of specific immune responses. We therefore suggest that using the methods described here could be key for the study of the substrate of B-cell selection and protective immunity in other vaccine and pathogen responses.
Daniel G. Fridman, Lena Israitel, Areen Shtewe et al.· Frontiers in Immunology· 0 citations
It is shown that conventional random train-test splits inflate retrieval accuracy by 15–28 percentage points owing to clonal leakage, and clone-aware benchmarking as a practical standard of comparison is established, defining the strengths and limitations of frozen zero-shot PLM embeddings for immune receptor retrieval and establishing clone-aware benchmarking as a practical standard of comparison.
Suyue Wang, Ayan Sengupta, Song-Ling Li et al.· bioRxiv· 0 citations
Accurate identification of interactions between T-cell receptors (TCRs) and antigenic peptides presented by major histocompatibility complex (MHC) molecules is essential for advancing precision immunotherapy. However, existing approaches often exhibit limited generalization to unseen peptides and struggle to capture the complex interaction patterns underlying immune recognition. Here, we present TCR-IFNet, a biologically informed deep learning framework for interpretable TCR-peptide interaction prediction. The model integrates global contextual representations from protein language models with local motif refinement via a gated convolutional module. To model cross-sequence dependencies, we introduce a Fast Kolmogorov-Arnold Network (FastKAN)-based cross-attention mechanism for nonlinear interaction modeling, together with a bilinear attention network to aggregate residue-level features into compact interface representations. Evaluation across multiple settings indicates that TCR-IFNet achieves competitive performance compared with existing methods, with higher AUPRC observed on both antigen-specific and healthy-sourced datasets, as well as improved results on independent test sets. The model also shows consistent generalization to unseen peptides under different negative sampling strategies. In addition, TCR-IFNet provides biologically meaningful interpretability by identifying key residue-level interaction patterns consistent with structural binding interfaces. Collectively, these findings demonstrate that TCR-IFNet provides a robust and generalizable computational framework for characterizing TCR-peptide interactions.
Wen-Yu Xi, Ruheng Wang, Xiu-Cai Ye et al.· International Journal of Bio...· 0 citations
Abstract Cross-immunity, defined as the ability of T-cells to recognize multiple antigen peptide-major histocompatibility complexes, is a fundamental feature of adaptive immunity. However, the prediction of different peptide epitopes that can be recognized by the same T-cell receptor remains challenging. Currently, artificial intelligent (AI)-based machine learning (ML) methods can be successfully used for pattern recognition in epitope molecular space by detecting the functional similarity between peptide sequences. In this study, using literature-based experimental data, we examined ML-based binary classification models trained on small datasets to predict the activity of nine-amino-acid-long peptides. Our results suggest that the consensus function of well-established similarity matrix-based representations and structural-based descriptors of epitopes yields better performance because representation-specific noises are reduced and individual model weaknesses are partially compensated. We also sought to determine the extent to which the predictive power of the applied AIs procedure depended on the physicochemical content of the descriptor set during the training process. In addition, challenging the models, we applied them to an independent experimental dataset to examine the effects of diverse laboratory conditions on a regulated biological measurement. In summary, applying a consensus function can capture the biological complexity of cross-reactivity at the binary classification level, even when applied to relatively small datasets.
V. Resch, László Tóth, Anita Rácz et al.· Briefings in Bioinformatics· 0 citations
Revising all immune receptor records to produce resolved, standardized, and analysis-ready receptor data will enable researchers to seamlessly query large-scale repertoires for receptors with experimentally verified specificity in the IEDB, link orphan sequences to known targets, and support cross-repository studies of receptor-epitope pairs and their relationship to health and disease.
Lonneke Scheffer, Eve Richardson, R. Vita et al.· Journal of Immunology· 0 citations
The human immune system serves as a ledger of environmental exposures, infections, and self-reactivity and there is a there is a critical need for high-throughput tools capable of mapping the entire human antibody reactome (all the antibody-antigen interactions). However, capturing the entire “immune history” can be difficult as conventional assays are limited to a narrow set of defined antigens.
Molecular Indexing of Proteins by Self-Assembly (MIPSA) technology employs a single-pot process to create libraries with antigens are covalently linked to unique DNA barcodes and decoded via Next-Generation Sequencing. This multiplexed, sample-sparing approach enables the simultaneous profiling of tens of thousands of interactions via core libraries: (1) HuSIGHT, covering the human proteome with 353,034 peptides and 16,516 full-length proteins; (2) VirSIGHT, covering all viruses known to infect humans with 285,669 peptides; and (3) EnviroSIGHT, covering 1,207 genera of allergens, microbial antigens, and toxins with 242,884 peptides.
Using commercial monoclonal antibodies and serum from patients with confirmed autoimmune diseases, the inclusion of both peptides and full-length proteins increases the detection of autoantibody targets on HuSIGHT. Profiling with VirSIGHT effectively distinguished “common” viral signatures from unique exposure histories. Furthermore, EnviroSIGHT allowed for the discrimination of distinct allergic profiles and geographical locations based on endemic pathogen exposure of the participants. We validated the platform’s robustness against technical and clinical confounders.
The MIPSA technology offers a scalable, high-resolution solution for uncovering immune insights offering a comprehensive view of immune history, this platform facilitates the discovery of novel biomarkers, infectious triggers of chronic conditions, the characterization of host-pathogen interactions, and the development of more effective therapeutic interventions for human health.
n/a
Technological Innovations in Immunology (TECH)
K. Shaw-Saliba· Journal of Immunology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.