GeneLLM is presented, a Transformer-based model that directly processes the nucleotide sequences of human-mapped cfRNA reads to identify cancer-indicative signatures and allows accurate cancer classification from plasma biopsies, suggesting that sequence-level modelling of plasma cfRNA can capture diagnostically relevant information beyond annotation-dependent approaches.
Abstract
Current liquid biopsy methods for multi-cancer detection using plasma cell-free RNA (cfRNA, short RNA fragments circulating in blood that can reflect disease states) typically rely on gene annotations, which can overlook signals from unannotated or repetitive genomic regions. We present GeneLLM, a Transformer-based model that directly processes the nucleotide sequences of human-mapped cfRNA reads to identify cancer-indicative signatures. By bypassing gene-level quantification, the model retains signals from transcriptomic dark matter. The model learns latent pseudo-biomarkers (prototype representations from aggregated cfRNA read embeddings) that serve as discriminative features for cancer classification, rather than corresponding to explicit genomic sequences. Here we show that, in a multi-centre cohort, GeneLLM achieves ROC-AUC values ranging from 0.9250 to 0.9962 across several cancers, while maintaining comparable performance at one-sixth of the typical sequencing depth. These results suggest that sequence-level modelling of plasma cfRNA can capture diagnostically relevant information beyond annotation-dependent approaches, enabling more cost-efficient and scalable cancer screening. Cell-freeRNA (cfRNA) can be a non-invasive and cost-effective biomarker for cancer therapy and clinical outcomes, but its analysis remains challenging. Here, the authors develop GeneLLM, a cfRNA-based large language model that processes raw cfRNA data and allows accurate cancer classification from plasma biopsies.
This PhD thesis establishes a comprehensive framework for advancing non-invasive cancer diagnostics through the characterization and combinatorial analysis of small non-coding RNAs (sncRNAs), specifically microRNA isoforms (isomiRs) and tRNA-derived fragments (tRFs). Liquid biopsy offers a minimally invasive alternative to traditional tissue biopsies, allowing for real-time disease monitoring via biomolecules like circulating tumor cells, extracellular vesicles (EVs), and tumor-educated platelets (TEPs). While canonical miRNAs are established biomarkers, this work demonstrates that isomiRs (variants resulting from alternative processing) and tRFs arising from tRNA cleavage represent a richer, underutilized reservoir of disease-specific signals.
To address the significant bioinformatics challenges and lack of standardized pipelines in the field, this thesis presents miRGalaxy. miRGalaxy is a novel, open-source, Galaxy-based framework designed for interactive and in-depth sequencing data analysis, enabling researchers without extensive computational backgrounds to identify and assess the differential expression of individual isomiR species.
The clinical utility of these sncRNAs was explored through multiple omics studies. In pancreatic ductal adenocarcinoma (PDAC), multi-omics profiling of TEPs revealed profound changes in the biological repertoire, including significantly high activity in RNA splicing and mRNA processing. A key finding was the downregulation of SPARC transcripts in PDAC platelets, which was strongly correlated with negative regulation by specific isomiRs such as miR-29a-3p and miR-22-3p.
Further characterization of the small RNA landscape in non-small-cell lung cancer (NSCLC) focused on "Platelet Dust" (PD) platelet-derived extracellular vesicles. PD was found to be significantly more enriched with miRNAs (~80–82%) compared to general EVs (~66%), which are relatively more enriched in tRNAs. This highlights PD as a more informative and distinct biomarker source for distinguishing cancer patients from healthy controls.
The culmination of this research is a combinatorial analysis of miRNAs, isomiRs, and tRFs from plasma EVs in colorectal and prostate cancer. By leveraging the synergistic effects of these different RNA species, the study achieved a diagnostic accuracy and Area Under the Curve (AUC) of approximately 80%. This approach proves more effective than analyzing single RNA types in isolation, providing a robust statistical framework for improved cancer management.
In conclusion, this thesis demonstrates that the integration of refined isomiR and tRF profiles within targeted liquid biopsy populations (TEPs and PD), supported by advanced bioinformatics tools like miRGalaxy, significantly enhances the accuracy and sensitivity of cancer diagnosis. These findings pave the way for more precise preventive screening and personalized oncology.
Liquid biopsy has emerged as a transformative approach in oncology, offering minimally invasive means for cancer detection, molecular characterization, and longitudinal disease monitoring. Among the diverse biosources available for liquid biopsy, tumor-educated platelets (TEPs) have garnered substantial interest as a rich and dynamic source of RNA-based biomarkers. Unlike circulating tumor DNA, which may present with low mutant allele fractions in early-stage disease, platelets offer abundant and relatively stable RNA that can be isolated from routine blood draws. Platelets, though anucleate, harbor megakaryocyte-derived messenger RNA and possess the capacity for RNA processing, enabling them to generate diverse transcriptomic repertoires. Importantly, platelets can sequester tumor-derived RNA from the circulation and through contact with tumor cells, producing disease-specific RNA signatures that can be captured through RNA sequencing and analysed using machine learning algorithms. Pan-cancer studies have demonstrated that TEP profiles can distinguish cancer patients from healthy controls with high accuracy, identify the primary site of tumor origin, and detect actionable molecular alterations. Disease-specific investigations have further validated TEP-based diagnostics across multiple solid tumor types, including non-small cell lung cancer, glioblastoma, colorectal cancer, ovarian cancer, pancreatic cancer, and sarcoma. Beyond diagnosis, TEP RNA signatures exhibit dynamic changes during treatment, supporting their application in monitoring therapeutic response and detecting disease progression. Nevertheless, critical challenges remain, including protocol sensitivity, pre-analytical confounding, and the need for rigorous prospective validation. This narrative review comprehensively examines the biological foundations of platelet tumor-RNA sequestration, synthesizes evidence on diagnostic and monitoring performance across cancer types, discusses technical platforms and computational methodologies, addresses limitations and negative findings, compares TEPs with other liquid biopsy modalities, and delineates future research priorities necessary to translate this promising approach into clinical practice.
S. D., S. P, J. M K et al.· Research Journal of Pharmaco...· 0 citations
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
Manoj M Wagle, Yongheng Wang, Soham Samanta et al.· bioRxiv· 0 citations
Background Circulating microbial DNA (cmDNA) has been proposed as a non-invasive cancer biomarker, but most evidence comes from cancer-sequencing datasets not designed for microbial analysis and lacking contamination controls. Whether reported signatures reflect biology or artifact is unclear in low-biomass specimens, where standard taxonomic pipelines are prone to systematic error. Methods In a tightly controlled pilot study of metastatic castration-resistant prostate cancer, we profiled plasma cell-free DNA (cfDNA) and buffy-coat genomic DNA (gDNA) from two patients and two healthy volunteers alongside mock blood-draw and reagent controls, each with and without host-DNA depletion. Reads were classified with Kraken2/Bracken and, independently, with the marker-gene classifier MetaPhlAn. As informatics controls, reads were per-base shuffled to randomize nucleotide order while preserving read length and guanine-cytosine (GC) content, and purely synthetic reads were generated from a four-base process matched only to an aggregate GC target; both were classified identically. Genus abundances were regressed against Kraken2 database k-mer representation and against GC content. Results Across 40 samples, Kraken2 reported several thousand genera, samples clustered by specimen type in principal-coordinate analysis (PCoA), and pooled genus counts correlated strongly with a published cancer-microbiome catalog (The Cancer Genome Atlas lung adenocarcinoma, TCGA-LUAD; Spearman ρ = 0.81 over 282 shared genera), a pattern readily interpreted as biological signal. However, these observations were also made in per-base shuffling, which preserves GC content and length but destroys all biological sequence: shuffled reads were still abundantly classified, still clustered by specimen type, and still correlated with the catalog (ρ ≈ 0.7), as did every sample group, including pure reagent controls. Genus counts scaled tightly with each genus’s k-mer representation in the Kraken2 database on real (r² = 0.74) and shuffled (r² = 0.85) reads, and the same dependence appeared in the independent published cohort. Purely synthetic reads carrying no information beyond an aggregate GC target reproduced much of the cross-cohort agreement (synthetic TCGA-LUAD ρ = 0.61 versus 0.81 for real reads; significant in 27 of 33 TCGA cancers), and replicate shuffles of a low-GC versus a high-GC plasma sample, for which the true difference is zero, produced spurious significant differences in about 46% of genera. Regressing observed counts against the shuffled baseline left 23 genera above the artifact floor at 5% false discovery rate (FDR), nearly all known kit contaminants, control-enriched viruses, or very-low-abundance taxa; a four-criterion validity filter reduced thousands of Kraken2 genera to a single defensible candidate, Klebsiella. Conclusions Much of the apparent cmDNA structure, including its agreement with a published cancer-microbiome catalog, is explained by base composition and reference-database architecture rather than authentic biology, and short-read k-mer pipelines cannot separate the two on their own. We find little positive evidence of an authentic circulating microbial signal, though our small sample cannot prove its absence. To limit false discovery in low-biomass metagenomics, we recommend specimen-matched negative controls, corroboration with a conservative second classifier, per-base shuffling (with GC-matched synthetic reads as a stricter floor), and GC-aware analysis. Graphical Abstract Pipeline pitfalls in blood microbial-DNA assays. cmDNA from low-biomass blood is vulnerable to three errors across the workflow (collection, processing, library preparation and sequencing, informatics): environmental (non-blood) contamination, kit and reagent contamination, and taxonomic misclassification. The corresponding safeguards are mock and specimen-matched negative controls, a literature sweep for known contaminants, and corroboration of k-mer output with a marker-gene classifier and a shuffle-based artifact control.
Daisy Fry Brumit, Daniel Bsteh, Shan Sun et al.· bioRxiv· 0 citations
Accurate prognosis in lung adenocarcinoma (LUAD) requires integration of high-dimensional transcriptomic profiles with compact but clinically stable patient covariates. Naïve fusion strategies allow the high-variance RNA-seq modality to dominate learned representations, suppressing clinical signal. We present Cooperative Modular Representation Learning (CMRL), an uncertainty-gated multimodal framework that dynamically regulates inter-modality information flow based on sample-level epistemic uncertainty estimated via Evidential Deep Learning (EDL). Each modality encoder produces a latent embedding and a scalar uncertainty score; an adaptive communication gate controls how much each module updates its representation from messages sent by the other module. A Variational Information Bottleneck (VIB) on the transcriptomic encoder further suppresses noise in the high-dimensional genomic latent space. CMRL is evaluated via 5-fold stratified cross-validation on 490 TCGA-LUAD patients with matched RNA-seq (504 features) and clinical data. It achieves a concordance index (C-index) of 0.732 ± 0.024, AUROC of 0.772 ± 0.019, and AUPRC of 0.773 ± 0.056 for 3-year survival prediction, outperforming a concatenation-fusion baseline (C-index 0.656), RNA-only (0.711), and clinical-only (0.670) variants, as well as several published LUAD survival models including CustOmics (0.625) and a whole-slide imaging method (0.675). An ablation study confirms that the uncertainty gate and evidential heads each contribute independently to the gain. Calibration analysis yields an Expected Calibration Error of 0.122, and uncertainty-stratified evaluation shows that low-uncertainty patients achieve AUROC 0.795 versus 0.681 for high-uncertainty patients, providing interpretable evidence that the gate mechanism is functioning as intended.
Sundus M. Jasim, Nabil Hezil, A. Bouridane et al.· bioRxiv· 0 citations