Skip to content
Preprint

TransNRank: Towards Accurate Neoantigen Ranking with Transformer

Aug 2026 · 0 citations
Computer Science

TL;DR

A novel deep learning framework based on Transformer, coined as TransNRank, that streamlines the prediction pipeline but also sets a new state-of-the-art for neoantigen discovery, with broad implications for accurate immuno-oncology.

Abstract

Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression and XGBoost fail to model long-range dependencies and contextual relationships within peptide features, therefore the performance of neoantigen positive recall rate is limited. In this paper, we present a novel deep learning framework based on Transformer, coined as TransNRank. By leveraging the self-attention mechanism, our model captures both local and global feature contexts, enabling more accurate recognition of immunogenic neoantigens. A positive-aware training objective is utilized to handle the class imbalance problem, assigning more weights to those few positive samples. Extensive experiments are performed on NCI, TESLA and HiTIDE datasets. Notably, our TransNRank can push the upper bound top 20 recall rate of neoantigen prediction from 46.9% (45 from 96) to 53.1% (51 from 96), while reducing the training epochs from 200 epochs to 20 epochs. Furthermore, we analyze the features contribution based on TransNRank and find that the mutation at anchor and TCGA expression level play an unexpected important role in neoantigen prediction, and removing insignificant features to reduce the input dimensionality of peptides does not drastically impair the overall performance of the model. Our paradigm not only streamlines the prediction pipeline but also sets a new state-of-the-art for neoantigen discovery, with broad implications for accurate immuno-oncology.

View source

Similar papers

Open access Sep 2026

VDJdb in 2026: boosting T-cell receptor recognition evidence using paratope embeddings and AI-based structure prediction.

We present a significant update to VDJdb, introducing substantial enhancements to both the data content and the technical infrastructure. The integration of steadily accumulating T-cell receptor (TCR):epitope recognition data, together with advances in high-throughput experimental techniques, has expanded the landscape...

D. Luppov, Anna E. Koneva, Dmitry V. Bagaev et al. · 0 citations
Open access Sep 2026

HSSynergy: scale-aware hierarchical attention for interpretable drug synergy prediction

HSSynergy, a hierarchical substructure-aware deep learning framework for predicting anticancer drug synergy, introduces a Scale-Aware Masked Attention mechanism that enforces precise layer-wise alignment, and utilizes hierarchical grouping with mask constraints to achieves same-scale focusing while shielding against cr...

Yue-Hua Feng, Shu-Cong Zhang, Xiao-Ying Yan et al. · 0 citations
Open access Aug 2026

Toward Imbalanced Molecular Property Regression: A Benchmark Study and Interval-Aware Mixture of Experts

Molecular property prediction is a key task in AI-driven drug discovery, yet the prevalence and impact of label imbalances in molecular property regression remain poorly understood. Through a systematic benchmark of widely used molecular property data sets, we show that target values are often highly imbalanced and tha...

Y. Sun, Yu Shi, Alana Deng et al. · 0 citations
#protein folding Open access Sep 2026

AcrSeek: Metric Learning with Hybrid Negative Mining for Anti-CRISPR Protein Detection under Extreme Class Imbalance.

A metric-learning framework that trains a projection head over a frozen protein language model (PLM) encoder with hybrid negative mining with hybrid negative mining under a joint triplet–focal objective that remains structurally robust under extreme class imbalance and may also serve other protein-function detection pr...

Chan-seok Jeong · 0 citations
#graph neural networks Open access Sep 2026

Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction

Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-un...

F. Midjani, S. Hashemi, Fatemeh Keshtkar et al. · 0 citations
Oct 2026

DeepRAGIL-2: a retrieval-augmented protein language model framework for sensitive and accurate prediction of IL-2-inducing peptides.

Interleukin-2 (IL-2) is a pleiotropic cytokine central to T-cell activation, proliferation, and immunological tolerance, yet existing computational tools for predicting IL-2-inducing peptides suffer from near-zero sensitivity under realistic class-imbalanced conditions, rendering them impractical for positive-class dis...

Juan Peter Timothy Yuune, Van-The Le, Yu-Yen Ou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.