Skip to content
Open access

LLM-powered Functional Gene Set Summarization with genesetGPT

Jul 2026 · bioRxiv · 0 citations · 61 references
Biology Medicine

TL;DR

GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.

Abstract

Transcriptomics datasets generated using next-generation sequencing techniques such as single cell RNA-sequencing (scRNA-seq) and spatially-resolved transcriptomics (SRT) allow researchers to study patterns in gene expression across celltypes, temporal processes, and spatial organization at ever-higher resolutions and depths. scRNA-seq analyses produce gene expression profiles and celltype-specific gene sets that require annotation to provide biological meaning, a process that has traditionally relied on the manual interpretations of clinical scientists. Similarly, SRT experiments typically require subjective, time-consuming annotation of spatial domains. Recent advances in large language model (LLM) methods offer opportunities to assist in the interpretation of such datasets. Many current LLM-based approaches aim to annotate transcriptomics-derived gene sets by integrating information from publicly available and online biological resources. While these approaches can be effective, they often struggle when presented with weakly-related or fully uncorrelated genes, sometimes inferring and justifying biological relationships that are not supported by existing literature. Additionally, the quality of LLM-generated interpretations is dependent on the provision of appropriate biological context and careful prompt design, both of which can present significant barriers to effective use. To address these limitations we propose genesetGPT, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale. genesetGPT is implemented as an open-source Python package available for download at https://github.com/jr-leary7/genesetGPT.

Read PDF

Similar papers

Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations
Open access Aug 2026

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes

Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Haijun Tang, Hening Li, Qing-Hao Zhao et al. · 0 citations
Open access Jul 2026

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.

L. Mboning, Maciej Dlugosz, Marek Kokot et al. · 0 citations
Preprint Aug 2026

Uncovering Cellular Resolution in scRNAseq via Unbiased Cell and Gene Network Analysis

Conventional annotation of single-cell RNA-sequencing (scRNA-seq) data relies heavily on manual, marker-based thresholding, an approach that can obscure subtle transcriptomic gradients and collapse functionally distinct cell states into broad, heterogeneous populations. Here we apply the Gaussian multi-Graphical Model (GmGM) framework, which jointly infers cell-cell and gene-gene dependency structure from a single scRNA-seq data matrix, to a 10x Genomics PBMC dataset. Ten independent GMGM-Leiden clustering runs were integrated into a robust consensus partition using a soft cluster ensemble approach and benchmarked against reference cell-type annotations. This strategy yielded stable cluster partitions that resolve biologically meaningful sub-populations not distinguished by the reference annotation. In parallel, for each cluster, gene co-expression modules were extracted from the fitted model via consensus Leiden clustering across resolutions, evaluated using standard network metrics, and validated functionally with the Network Enrichment Analysis Test (NEAT), which confirmed non-random enrichment signal. A module-scoring procedure linked network topology to per-cell, per-cluster expression signatures, and a novel extension of GmGM, recovering a shared cell-cell network together with population-specific gene networks in a single model run, was demonstrated in a case study on the CD4+ T-cell population. These results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.

O. Lanzetta, L. Cutillo, Bailey Andrew et al. · 0 citations
Open access Aug 2026

pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.

Tommaso Giacomello, S. Mazzara, Gennaro Abbruzzese et al. · 0 citations
Review Open access Jul 2026

Single-cell long-read transcriptomics: from technologies to biological insights

Abstract Single-cell long-read transcriptomics (scLR-seq) extends single-cell analysis beyond gene abundance by resolving full-length transcript structures in individual cells. It can directly interrogate isoform usage, alternative splicing, and transcription start and end site selection, thereby revealing regulatory variation that is often obscured by short-read measurements. In this review, we examine the experimental and computational foundations of scLR-seq, including platform selection, library design, cell barcode and unique molecular identifier recovery, transcript discovery, and isoform quantification. We discuss how these choices influence the reliability of downstream biological interpretation, and summarize emerging insights into isoform usage, alternative splicing, transcription start and end site selection, allele-specific expression, fusion transcripts, transposable element-derived transcripts, and RNA modifications. Finally, we highlight applications of scLR-seq in diverse biological systems, such as the immune system, neural development, and tumor microenvironments, and consider future opportunities and challenges in integrating multi-omics data to decode cellular programs and disease evolution.

Zengjun Ren, Wenteng Liu, Jianhua Yin et al. · 0 citations