In conclusion, integrated coding-noncoding analysis is established as a strategic approach for discovering functional ncRNAs from transcriptomic sequencing data.
Abstract
Although the human genome encodes a vast repertoire of noncoding RNAs that regulate gene expression, the noncoding genome remains underexplored due to technical challenges. Specifically, during transcriptomic sequencing data alignment, the overlap between noncoding and coding loci can create ambiguous read alignments that are subsequently discarded from downstream analysis. For this reason, most of the noncoding genome is excluded from standard genomic annotations used for sequencing alignment. To address this challenge and enable concurrent profiling of the coding and noncoding transcriptome, we systematically integrated standard coding (GENCODE) and noncoding (LncBook) genome annotations, preserving coding gene annotations and removing overlapping noncoding regions. The resulting integrated genome annotation expanded the number of annotated noncoding genes from 40,785 to 138,296 while preserving all coding genes and reducing ambiguous read assignment. To evaluate the utility of our integrated genome annotation for uncovering novel, biologically relevant noncoding RNAs (ncRNAs), we realigned CD138-positive bulk RNA-seq (N = 942) and CD138-negative single-cell RNA-seq (N = 478) data from the MMRF CoMMpass study, generating a comprehensive coding-noncoding atlas of the myeloma bone marrow microenvironment with noncoding genes representing 51% of highly variable genes and displaying significant cell type specificity. Tumor expression profiling based on this integrated profiling identified 15 clusters, including two enriched for amp(1q21) or t(4;14) and associated with shorter progression-free survival (PFS). Differential expression and systematic filtering yielded 19 candidate high-risk ncRNAs, including previously uncharacterized ENSG00000310209, which was associated with poor PFS (HR = 1.141, P = 0.0025), increased IRF4 activity, Wnt pathway activation, CCL5 signaling, and the accumulation of anergic-like CD8+ T cells. These findings establish integrated coding-noncoding analysis as a strategic approach for discovering functional ncRNAs from transcriptomic sequencing data.
The analysis of long-read RNA sequencing (RNA-seq) data can extend current annotations across the genomes, transcriptomes, and proteomes of even the most annotated species such as
Homo sapiens
. Long-read sequencing (LRS) technologies like Oxford Nanopore Technologies' (ONTs) nanopore sequencing can capture full-le...
Kristina Santucci, Yu-Ning Cheng, Yu-Lan Gao et al.· Frontiers in Cardiovascular...· 0 citations
How genetic variation regulates long non-coding RNA (lncRNA) expression across brain cell types remains poorly understood. A major barrier is the resolution–power trade-off between bulk and single-nucleus expression quantitative trait locus (eQTL) studies. Here, we quantify gene expression from cortical RNA-seq of 2443...
Yu-Ran Jia, Li Chen, Li-Yang Song et al.· Nature Communications· 0 citations
It is proposed that the upregulated lncRNA ENSG00000265613 may enhance malignancy by stabilizing the RNA target ENSG00000582008 in luminal A breast cancer, particularly given its established role in oncogenesis.
C. Guda, Sankarasubramanian Jagadesan, Avinash M. Veerappa· Methods in molecular biology· 0 citations
Genetic alterations are closely associated with prostate cancer development and progression, but the RNA-derived transcriptional allele landscape of prostate cancer and cancer stem cells (CSCs) remains poorly understood. To characterize cancer-associated transcriptional alleles, we analyzed bulk and single-cell RNA seq...
Wen-Yang Hu, Ran-Li Lu, M. Maienschein-Cline et al.· Biomolecules· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.