Skip to content

Abstract P15: Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML

Jul 2026 · Cancer Research · 0 citations

TL;DR

A context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML, which establishes the model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights.

Abstract

While language models extract linguistic structures from text, similar approaches can uncover biological rules from genetic patterns. Though these methods have shown promise in single-cell analysis, bulk transcriptomics remains underexplored despite offering distinct clinical advantages including preserved tissue-level information, higher sequencing depth, and cost-effectiveness. Here, we present a transformer-based foundation model leveraging transcriptomic profiles from over 30,000 diverse bulk RNA samples, including normal tissues and various cancer types. Unlike conventional language models, our model incorporates specialized modules for modelling pairwise gene interactions through a dual representation system that captures both gene-level features and their higher-order relationships. Our model shows robust performance across multiple downstream applications. It achieves zero-shot accuracy of 78.81% in cancer classification without fine-tuning and outperforms existing approaches in cancer stages prediction through simple fine-tuning. Notably, it can extract critical gene interaction networks without relying on prior biological knowledge. More importantly, we leverage it to introduce dynamic interpretations to static bulk transcriptomic data, successfully modelling logical gene regulation rules with 91.07% overall accuracy—reaching 100% for rules related to key genes like GATA2 and SCL. With the context-specific modelling ability, it also identifies, for example, transcriptional dynamics in normal haematopoiesis and dysregulated circuits during transition to leukemic states. We further demonstrate clinical utility in predicting patient response to first induction chemotherapy (AUROC=0.75) in acute myeloid leukemia, a challenging task due to patient and mechanism heterogeneity. Through our novel response-directed feature-space gradient ascent approach, we identify patient-specific gene expression modifications that could computationally redirect resistant phenotypes toward responsive ones, revealing potential therapeutic targets aligned with individual patients' clinical features. These results establish our model as a powerful framework to extract useful information from bulk transcriptomics data and has potential applications in precision medicine by connecting computational predictions with biological insights. Yi Chai, Yang Li, Jianbiao Zhou, Wee Joo Chng, Yang Zhang. Context-aware foundation model of bulk transcriptomics for interpretable analysis of transcriptional dynamics and treatment response in AML [abstract]. In: Proceedings of Frontiers in Cancer Science 2025; 2025 Nov 5-7; Singapore. Philadelphia (PA): AACR; Cancer Res 2026;86(13_Suppl):Abstract nr P15.

View source

Similar papers

Open access Jul 2026

FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations

A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).

Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al. · 0 citations
Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations
Open access Jul 2026

Decoding cancer circulating transcriptomic signatures with language models

GeneLLM is presented, a Transformer-based model that directly processes the nucleotide sequences of human-mapped cfRNA reads to identify cancer-indicative signatures and allows accurate cancer classification from plasma biopsies, suggesting that sequence-level modelling of plasma cfRNA can capture diagnostically relevant information beyond annotation-dependent approaches.

Siwei Deng, Lei Sha, Yongcheng Jin et al. · 0 citations
Preprint Aug 2026

Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models

A representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks shows that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

Nirjhor Datta, Swakkhar Shatabda, M. S. Rahman · 0 citations
Open access Jul 2026

Knowledge-guided contextual gene set analysis with large language models

Abstract Motivation Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. Results We introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis. Availability and Implementation The demo website is publicly available at https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/cGSA/, while the data and code can be accessed at https://github.com/ncbi-nlp/cGSA.

Zhizheng Wang, Chi-Ping Day, Chih-Hsuan Wei et al. · 1 citation · ⚡1
Open access Jul 2026

Annotation-free phenotype prediction using knowledge-augmented clustering from single-cell RNA sequencing data

Evaluated across three public scRNA-seq datasets, scCap consistently outperforms baseline models in predictive accuracy and identifies disease-associated subpopulations previously reported in the literature without relying on predefined cell-type annotations.

Janghyun Noh, Yoobin Shin, M. Kim et al. · 0 citations