Discovery and annotation of putative genes and novel RNA splice-isoforms in the human heart
Abstract
The analysis of long-read RNA sequencing (RNA-seq) data can extend current annotations across the genomes, transcriptomes, and proteomes of even the most annotated species such as Homo sapiens . Long-read sequencing (LRS) technologies like Oxford Nanopore Technologies' (ONTs) nanopore sequencing can capture full-length transcripts, making the task of read-to-isoform assignment less challenging. This enables both the discovery of new RNA splice isoforms as well as a more accurate measure of their individual expression. RNA expression from unannotated regions of the genome can also lead to the discovery of putative (candidate novel) genes. Computational methods such as protein-coding potential assessment and ORF sequence prediction allow for the detection of new coding sequences (CDS), which can expand proteome annotations. This study seeks to apply these methods for the discovery of new RNA isoforms deriving from both annotated and putative genes in the human cardiac transcriptome. Recent LRS of prized Genotype Tissue Expression (GTEx) project human samples of the atrial appendage, left ventricle, and skeletal muscle were re-analysed with a new analytical pipeline that seeks to robustly discover these novel features. As a result, 54 intergenic and 112 antisense putative genes were detected. 491 novel RNA isoforms were discovered, 173 of which were transcribed from putative genes. 142 of the novel RNA isoforms were determined to have coding potential, and the remaining 349 determined to have limited or no coding potential with lengths greater than 200 nt, deeming them long non-coding RNAs (lncRNAs). Of the messenger RNA (mRNA) isoforms, 56 had low homology to known proteins, suggesting they may encode new protein isoforms or truncated proteins, and 24 had no detected homology to known proteins. Overall, this study has characterised new genes, RNA molecules, and protein isoforms, and has explored their potential function through various sequence and expression-based analyses, including the detection of conserved sequence and protein domains, cis-regulatory potential analysis, and weighted gene co-expression network analysis. Such analyses consistently demonstrated the potential for these discovered features to have functional roles in the human cardiac system. By re-analysing existing datasets, this study has revealed information that was not captured in previous analyses, resulting in more complete and accurate reference annotations of the human genome.