MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.
Abstract
Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals—including sex differences in killifish liver—and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex-and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer–based analysis of next-generation sequencing data.
Conventional annotation of single-cell RNA-sequencing (scRNA-seq) data relies heavily on manual, marker-based thresholding, an approach that can obscure subtle transcriptomic gradients and collapse functionally distinct cell states into broad, heterogeneous populations. Here we apply the Gaussian multi-Graphical Model (GmGM) framework, which jointly infers cell-cell and gene-gene dependency structure from a single scRNA-seq data matrix, to a 10x Genomics PBMC dataset. Ten independent GMGM-Leiden clustering runs were integrated into a robust consensus partition using a soft cluster ensemble approach and benchmarked against reference cell-type annotations. This strategy yielded stable cluster partitions that resolve biologically meaningful sub-populations not distinguished by the reference annotation. In parallel, for each cluster, gene co-expression modules were extracted from the fitted model via consensus Leiden clustering across resolutions, evaluated using standard network metrics, and validated functionally with the Network Enrichment Analysis Test (NEAT), which confirmed non-random enrichment signal. A module-scoring procedure linked network topology to per-cell, per-cluster expression signatures, and a novel extension of GmGM, recovering a shared cell-cell network together with population-specific gene networks in a single model run, was demonstrated in a case study on the CD4+ T-cell population. These results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.
O. Lanzetta, L. Cutillo, Bailey Andrew et al.· 0 citations
Summary High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and Implementation Source code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary information Supplementary data are available at Bioinformatics online.
Tommaso Giacomello, S. Mazzara, Gennaro Abbruzzese et al.· bioRxiv· 0 citations
Knowledge of RNA sequences, expression, splicing, isoforms, structure and modifications is central for understanding and targeting cellular processes. Revolutionary single-cell and spatial transcriptomics technologies-for example, as deployed by consortia such as the Human Cell Atlas-partially capture this diversity and generate cellular profiles that expand at petabyte scale each year1-5. Yet researchers cannot search sequences across these datasets: standard pipelines do not scale or rely on references, retaining only gene or isoform counts, whereas accessing raw sequences requires collecting, downloading and processing millions of large files. Here we present Malva, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts. The continuously expanding Malva Index currently comprises around 74 million cells from thousands of experiments in health and disease. Malva enables reference-free discovery-researchers can, for example, identify cell types and predict cell-cell similarity directly from sequence composition. Building on Malva's speed and accuracy, we demonstrate how Malva can be flexibly connected to state-of-the-art neural networks and how to execute complex searches and enable automated analyses. Malva transforms single-cell atlases from static gene count tables into dynamic, sequence-resolved resources that may help to bridge human-machine reasoning about biology.
D. León-Periñán, Nikos Karaiskos, N. Rajewsky· Nature· 0 citations
GenesetGPT is proposed, an efficient, LLM-based framework that emphasizes both curated biological context and iterative prompt construction, thus enabling realistic summarization of heterogeneous gene sets at scale.
Jack R. Leary, Samantha Pattey, Rhonda L. Bacher· bioRxiv· 0 citations
GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.
Xiaodan Zhang, S. Paithankar, Jing Pu et al.· bioRxiv· 0 citations
This PhD thesis establishes a comprehensive framework for advancing non-invasive cancer diagnostics through the characterization and combinatorial analysis of small non-coding RNAs (sncRNAs), specifically microRNA isoforms (isomiRs) and tRNA-derived fragments (tRFs). Liquid biopsy offers a minimally invasive alternative to traditional tissue biopsies, allowing for real-time disease monitoring via biomolecules like circulating tumor cells, extracellular vesicles (EVs), and tumor-educated platelets (TEPs). While canonical miRNAs are established biomarkers, this work demonstrates that isomiRs (variants resulting from alternative processing) and tRFs arising from tRNA cleavage represent a richer, underutilized reservoir of disease-specific signals.
To address the significant bioinformatics challenges and lack of standardized pipelines in the field, this thesis presents miRGalaxy. miRGalaxy is a novel, open-source, Galaxy-based framework designed for interactive and in-depth sequencing data analysis, enabling researchers without extensive computational backgrounds to identify and assess the differential expression of individual isomiR species.
The clinical utility of these sncRNAs was explored through multiple omics studies. In pancreatic ductal adenocarcinoma (PDAC), multi-omics profiling of TEPs revealed profound changes in the biological repertoire, including significantly high activity in RNA splicing and mRNA processing. A key finding was the downregulation of SPARC transcripts in PDAC platelets, which was strongly correlated with negative regulation by specific isomiRs such as miR-29a-3p and miR-22-3p.
Further characterization of the small RNA landscape in non-small-cell lung cancer (NSCLC) focused on "Platelet Dust" (PD) platelet-derived extracellular vesicles. PD was found to be significantly more enriched with miRNAs (~80–82%) compared to general EVs (~66%), which are relatively more enriched in tRNAs. This highlights PD as a more informative and distinct biomarker source for distinguishing cancer patients from healthy controls.
The culmination of this research is a combinatorial analysis of miRNAs, isomiRs, and tRFs from plasma EVs in colorectal and prostate cancer. By leveraging the synergistic effects of these different RNA species, the study achieved a diagnostic accuracy and Area Under the Curve (AUC) of approximately 80%. This approach proves more effective than analyzing single RNA types in isolation, providing a robust statistical framework for improved cancer management.
In conclusion, this thesis demonstrates that the integration of refined isomiR and tRF profiles within targeted liquid biopsy populations (TEPs and PD), supported by advanced bioinformatics tools like miRGalaxy, significantly enhances the accuracy and sensitivity of cancer diagnosis. These findings pave the way for more precise preventive screening and personalized oncology.