Skip to content
Open access

The Immune Epitope Database: Revised Receptor Data and Integration with the Adaptive Immune Receptor Repertoire Knowledge Commons 2306862

Jul 2026 · Journal of Immunology · 0 citations

TL;DR

Revising all immune receptor records to produce resolved, standardized, and analysis-ready receptor data will enable researchers to seamlessly query large-scale repertoires for receptors with experimentally verified specificity in the IEDB, link orphan sequences to known targets, and support cross-repository studies of receptor-epitope pairs and their relationship to health and disease.

Abstract

The Immune Epitope Database (IEDB, iedb.org) is a freely available resource that catalogs experimentally defined immune epitopes. Concurrently, the IEDB records ∼190,000 T cell receptors and ∼5,000 antibodies with experimentally verified epitope specificity. Because these receptors have been manually curated from 3,300 references spanning decades, reported data and nomenclature can be inconsistent, posing challenges for computational analyses. To support interoperability and integration with community resources such as the Adaptive Immune Receptor Repertoire Knowledge Commons (AKC), we are revising all immune receptor records to produce resolved, standardized, and analysis-ready receptor data. We developed a computational pipeline that employs IgBLAST for V/D/J gene assignment, ANARCII for identification of Complementarity Determining Regions (CDRs), and tidytcells to standardize author-reported gene names. We furthermore extended tidytcells to validate and standardize CDR3 sequences based on reported V/J gene usage and to support antibody data. Crucially, the pipeline also flags anomalous data for targeted re-curation by expert curators. The reprocessed receptor dataset contains V/D/J gene names that are correctly formatted and mapped to existing reference genes, and CDR3 sequences are consistently represented up to their conserved anchor residues. Improved anomaly detection allowed us to identify and correct anomalous receptor records from hundreds of studies. These revisions increase data quality and improve interoperability, as exemplified by integration with the AKC. This integration will enable researchers to seamlessly query large-scale repertoires for receptors with experimentally verified specificity in the IEDB, link orphan sequences to known targets, and support cross-repository studies of receptor-epitope pairs and their relationship to health and disease. The IEDB is funded by NIAID contract 75N93019C00001. The AIRR Knowledge Commons is supported by a U24 (U24I177622) from the NIAID. Computational and Systems Immunology (COMP)

Read PDF

Similar papers

Review Open access Aug 2026

The Cancer Epitope Database and Analysis Resource (CEDAR): current capabilities and future directions

CEDAR’s current capabilities, report on progress in curation, database development, and tool availability, and outline the opportunities and challenges ahead for expanding its scope and utility to the cancer research community are described.

Zeynep Koşaloğlu-Yalçın, Ibel Carri, Daniel Marrama et al. · 0 citations
Open access Jul 2026

Adaptive Immune Receptor Repertoire Knowledge Commons: data harmonization 2309859

Building a knowledgebase that integrates multiple data repositories requires the concepts, relationships, and data schemas/formats to be harmonized across those repositories. One technique is designing a common data model (CDM) that encompasses all concepts and relationships, transforming the data into the CDM, and utilizing ontologies to provide shared semantics. The Adaptive Immune Receptor Repertoire Knowledge Commons (AKC) is a publicly accessible repository that integrates data and knowledge about 1) adaptive immune receptors (AIRs) and AIR repertoires from the AIRR Data Commons, 2) AIR germline allele, genotype, haplotype, and population genetic data from the OGRDB and VDJbase, and 3) AIR specificity data from the IEDB and IRAD. We designed a CDM for the AKC based upon the Ontology for Biomedical Investigations, a community standard for scientific data integration, and we used the LinkML data modeling language for implementation. The AKC provides a consistent CDM for study, subject, and sample information; immune exposures and other study events; sample collection; assays and processing; and data processing and analysis workflows. The CDM also provides adaptive immunity domain knowledge for chains, receptors, antigens, epitopes, and MHC/HLA. LinkML’s flexible data modeling language allows for existing data standards, such as the AIRR Standards, and ontologies from the OBO Foundry to be directly incorporated. The foundation of the AKC is data integrated from these community-supported repositories and harmonized around a CDM based on widely adopted ontologies and data standards. The AKC assembles the critical mass of data required to develop highly accurate predictive algorithms for long-standing questions of critical importance (e.g., predicting AIR specificity, determining the contribution of AIR germline polymorphisms to disease propensity) and to ask questions across a large and diverse set of subjects with a variety of health and disease phenotypes. The research described is supported by the National Institute of Allergy and Infectious Diseases of the National Institutes of Health under award number U24AI177622. Computational and Systems Immunology (COMP)

Scott Christley, Felix Breden, Kevin A. Burns et al. · 0 citations
Open access Apr 2026

LM-QASAS: reference-free identification of antigen-specific sequences from the BCR repertoire using antibody language models

LM-QASAS provides a rapid, high-precision platform for monitoring humoral immunity against emerging threats by mapping sequences into a high-dimensional semantic embedding space, and identifies functionally convergent clusters of sequences that are semantically similar and exhibit transient expansion upon immune stimulation.

Genki Masuda, Y. Funakoshi, S. Iizumi et al. · 0 citations
Open access Jul 2026

A germline shortcut in protein language model retrieval of adaptive immune receptors

It is shown that conventional random train-test splits inflate retrieval accuracy by 15–28 percentage points owing to clonal leakage, and clone-aware benchmarking as a practical standard of comparison is established, defining the strengths and limitations of frozen zero-shot PLM embeddings for immune receptor retrieval and establishing clone-aware benchmarking as a practical standard of comparison.

Suyue Wang, Ayan Sengupta, Song-Ling Li et al. · 0 citations
Jul 2026

ImmuneSpace in 2026: A Centralized Repository for Curated Human Immune-Profiling Data 2310296

ImmuneSpace (immunespace.org) is a freely accessible database that hosts curated human immune-profiling data from a wide range of studies. It was created as the central data repository of the Human Immunology Project Consortium (HIPC), a multi-center NIH-funded program to characterize the diverse states of the human immune system and its regulation, using consistently formatted data [PMID: 23648045]. The overarching goal is to investigate human immune perturbations using state-of-the-art systems-level profiling technologies and innovative methodologies, and to make these data available to the scientific community, accessible for both humans and machines. ImmuneSpace hosts data related to immunological exposure, demographics, cytokine profiling, cytometry, and neutralizing antibody assays. It builds on the ImmPort data model, a long-term archive of research and clinical data for the NIH [PMID: 29485622], but implements additional standardization and normalization rules, following the HIPC Data Standards initiative [PMID: 22343568, 26861911, 31272390, 32283555]. ImmuneSpace is continuously updated with data and features. Each study undergoes enhanced curation to ensure consistent use of ontology-based terminology, enabling more efficient queries both within and across studies. Study components stored in different repositories, unparsed raw data, and computationally inaccessible elements are integrated through manual curation, such as study timelines and author-determined ‘immune signatures’. Recently, we have added the ‘Finder Feature’ which simplifies complex searches through a hierarchical, tree-based interface that allows users to visually browse, search with autocomplete, or select entire branches of categorized filters while providing definitions, synonyms, and ontology links. The HIPC Project and the ImmuneSpace platform demonstrate the feasibility and benefits of a structured approach to representing human immunological studies to elucidate system-level phenomena. U01 AI167892 Computational and Systems Immunology (COMP)

Kerstin Westendorf, M. Kojima, James A. Overton et al. · 0 citations
Open access Jul 2026

Decoding the Antibody Reactome at Scale with Novel Antigen Display Technology to Decipher Immune History 2335216

The human immune system serves as a ledger of environmental exposures, infections, and self-reactivity and there is a there is a critical need for high-throughput tools capable of mapping the entire human antibody reactome (all the antibody-antigen interactions). However, capturing the entire “immune history” can be difficult as conventional assays are limited to a narrow set of defined antigens. Molecular Indexing of Proteins by Self-Assembly (MIPSA) technology employs a single-pot process to create libraries with antigens are covalently linked to unique DNA barcodes and decoded via Next-Generation Sequencing. This multiplexed, sample-sparing approach enables the simultaneous profiling of tens of thousands of interactions via core libraries: (1) HuSIGHT, covering the human proteome with 353,034 peptides and 16,516 full-length proteins; (2) VirSIGHT, covering all viruses known to infect humans with 285,669 peptides; and (3) EnviroSIGHT, covering 1,207 genera of allergens, microbial antigens, and toxins with 242,884 peptides. Using commercial monoclonal antibodies and serum from patients with confirmed autoimmune diseases, the inclusion of both peptides and full-length proteins increases the detection of autoantibody targets on HuSIGHT. Profiling with VirSIGHT effectively distinguished “common” viral signatures from unique exposure histories. Furthermore, EnviroSIGHT allowed for the discrimination of distinct allergic profiles and geographical locations based on endemic pathogen exposure of the participants. We validated the platform’s robustness against technical and clinical confounders. The MIPSA technology offers a scalable, high-resolution solution for uncovering immune insights offering a comprehensive view of immune history, this platform facilitates the discovery of novel biomarkers, infectious triggers of chronic conditions, the characterization of host-pathogen interactions, and the development of more effective therapeutic interventions for human health. n/a Technological Innovations in Immunology (TECH)

K. Shaw-Saliba · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.